Models That Know How Evaluations Are Designed Score Safer
The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such as hypothetical scenarios, as a source of verbalized evaluation awareness and subsequent behavioral shift. In this paper, we investigate a potential explanation of this phenomenon: evaluation meta-knowledge, defined as parametric knowledge about the structural traits that characterize evaluations. Similar to dataset con
Record details
Published: 27 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
GS-FUSE: Granger-Supervised Gated Fusion and Multi-Granularity Alignment for Event-Driven Financial Forecasting
arXiv · 27 May 2026
Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning
arXiv · 28 May 2026
Rationalize: Shared Semantic Reasoning for Human-AI Alignment
arXiv · 28 May 2026
Human-Alignment, Calibration, and Activation Patterns in Large Language Model Uncertainty
arXiv · 29 May 2026
Algorithmic Fragility and Persona Bias in LLM-Generated Autistic Communication
arXiv · 26 May 2026
What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness
arXiv · 29 May 2026
How to cite this record
ethics.ai (27 May 2026), “Models That Know How Evaluations Are Designed Score Safer,” evidence record 3582, https://ethics.ai/record/3582 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.