Alignment midtraining for animals
We investigate the robustness of value alignment via midtraining with synthetic documents, using animal compassion as a value that is both important in its own right and orthogonal to existing alignment efforts. To evaluate compassionate reasoning, we develop and publicly release Animal Norms In Moral Assessment (ANIMA), a 26-question evaluation spanning 13 ethical dimensions, publicly available as a dataset and Inspect evaluation. On ANIMA, training with 3000 documents achieves 77% compared to
Record details
Published: 21 March 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Characterizing Linear Alignment Across Language Models
arXiv · 19 March 2026
Too much of a good thing? Entrepreneurial orientation and the non-linear governance effects of SaaS platforms
arXiv · 22 March 2026
Why AI Readiness Is an Organizational Learning Problem, Not a Technology Purchase
arXiv · 22 March 2026
Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models
arXiv · 23 March 2026
VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models
arXiv · 18 March 2026
Beyond Grading Accuracy: Exploring Alignment of TAs and LLMs
arXiv · 17 March 2026
How to cite this record
ethics.ai (21 March 2026), “Alignment midtraining for animals,” evidence record 6932, https://ethics.ai/record/6932 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.