LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the validity of safety and alignment benchmarks. We propose LURE (Live-Usage Replay Evaluations), a method for constructing deployment-like evaluations by replaying realistic agentic interaction trajectories and appending evaluation prompt at the end. We also introduce an automated pipeline for measuring evaluation realism, combining detection of verba
Record details
Published: 8 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
arXiv · 8 April 2026
IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents
arXiv · 6 April 2026
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
arXiv · 6 April 2026
The Sustainability Gap in Robotics: A Large-Scale Survey of Sustainability Awareness in 50,000 Research Articles
arXiv · 9 April 2026
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
arXiv · 9 April 2026
From Safety Risk to Design Principle: Peer-Preservation in Multi-Agent LLM Systems and Its Implications for Orchestrated Democratic Discourse Analysis
arXiv · 9 April 2026
How to cite this record
ethics.ai (8 April 2026), “LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness,” evidence record 6220, https://ethics.ai/record/6220 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.