CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment
Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns. While existing LLM safety guardrails excel in English or multilingual settings, they lack adaptation to Chinese-specific regulatory policies, cultural context and linguistic nuances, failing to support fine-grained risk classification for diverse deployment needs. In this paper, we introduce a 5-macro, 31-micro category fine-grained risk taxonomy for Chinese scenarios, and build CHI
Record details
Published: 13 June 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
AlignAtt4LLM: Fast AlignAtt for Decoder-Only LLMs at IWSLT 2026 Simultaneous Speech Translation Task
arXiv · 2 June 2026
Geopolitical alignment: Endorsement effects in large language models
arXiv · 10 July 2026
Faithful or evasive? An empirical study on translation norm preferences of Chinese and American LLMs in Chinese official political and policy discourse
Frontiers in Artificial Intelligence · 10 August 2026
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
arXiv · 18 March 2026
China works on AI safety benchmark as regulators target large model risks
SCMP Tech (HK/CN) · 13 July 2026
Huawei’s Ren Zhengfei backs Tau Scaling Law to beat US sanctions
SCMP Tech (HK/CN) · 24 July 2026
How to cite this record
ethics.ai (13 June 2026), “CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment,” evidence record 1020, https://ethics.ai/record/1020 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.