Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds
Reward hacking, where AI systems exploit misspecified objectives to achieve high reward without satisfying intended goals, remains a central challenge in AI safety. Yet most known instances have been discovered post hoc in frontier systems where controlled study is impractical. We adapt the AI Safety Gridworlds framework into a text-based evaluation suite that reformulates classic reinforcement learning safety tasks for language-based agents. Across frontier and mid-scale models, we find that sp
Record details
Published: 13 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
arXiv · 13 June 2026
Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation
arXiv · 13 June 2026
Regulating the Machine Contributor: Governance and Policy Alignment in Open Source
arXiv · 12 June 2026
CAMI: Cost-Aware Agent-Guided Multi-Indexing for Semantic Retrieval
arXiv · 14 June 2026
Agentic Framework for Deep Learning workload migration via In-Context Learning
arXiv · 14 June 2026
An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios
arXiv · 15 June 2026
How to cite this record
ethics.ai (13 June 2026), “Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds,” evidence record 1021, https://ethics.ai/record/1021 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.