A Scalable Approach to Evaluating Moral Sensitivity in LLMs
Moral sensitivity is the ability to identify the morally relevant features of a decision situation and use them as the basis for action. It is the foundation of broader moral competence: any other moral reasoning capabilities will be irrelevant if an agent lacks sensitivity to the relevant facts. In this paper, we offer a new evaluation of LLM moral sensitivity and in doing so, we address and resolve a central problem in AI alignment research: how to scale behavioural evaluations beyond expensiv
Record details
Published: 3 July 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Overloading Large Vision-Language Models for Jailbreaking
arXiv red teaming query · 3 July 2026
CONTRA: Red-Teaming Configurations of Personalizable Agents
arXiv red teaming query · 3 July 2026
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
arXiv · 3 July 2026
Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring
arXiv · 2 July 2026
VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment
arXiv · 2 July 2026
Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents
arXiv red teaming query · 4 July 2026
How to cite this record
ethics.ai (3 July 2026), “A Scalable Approach to Evaluating Moral Sensitivity in LLMs,” evidence record 293, https://ethics.ai/record/293 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.