Rules or Character? Scaling Laws for AI Safety Design
Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at training time, with rule enforcement (e.g., output filters, safety classifiers), which blocks harmful outputs at inference time, yet little formal analysis exists on how their optimal balance should change as deployment scales increase. We introduce a stylized comparative-statics model that parameterizes saf
Record details
Published: 13 August 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems
arXiv · 13 August 2026
Synthetic Persona Pretraining: Alignment from Token Zero
arXiv · 13 August 2026
ProME: Prototype-Margin Environments with Repair-Aware Selection for Group-Robust Learning
arXiv cs.LG · 13 August 2026
Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds
arXiv · 13 August 2026
From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion
arXiv · 13 August 2026
Data augmentation in multimodal frameworks: a survey
Artificial Intelligence Review · 14 August 2026
How to cite this record
ethics.ai (13 August 2026), “Rules or Character? Scaling Laws for AI Safety Design,” evidence record 19167, https://ethics.ai/record/19167 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.