Item Response Theory for AI Safety
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. We fit IRT models to eight safety ben
Record details
Published: 5 August 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 6 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming
arXiv red teaming query · 5 August 2026
Rectifying Geometric Misalignment: Online Source-Free Adaptation for Class-Imbalanced EEG
arXiv cs.LG · 5 August 2026
Scaling Inherently Interpretable Language Models
HuggingFace Daily Papers · 5 August 2026
Decoupling Perception from Description: Computation-Grounded Representation Alignment between Multivariate Time Series and Language
arXiv cs.LG · 5 August 2026
Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment
arXiv red teaming query · 5 August 2026
RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists
arXiv cs.AI · 5 August 2026
How to cite this record
ethics.ai (5 August 2026), “Item Response Theory for AI Safety,” evidence record 16650, https://ethics.ai/record/16650 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.