Questionnaire Responses Do not Capture the Safety of AI Agents
As AI systems advance in capabilities, measuring their safety and alignment to human values is becoming paramount. A fast-growing field of AI research is devoted to developing such assessments. However, most current advances therein may be ill-suited for assessing AI systems across real-world deployments. Standard methods prompt large language models (LLMs) in a questionnaire-style to describe their values or behavior in hypothetical scenarios. By focusing on unaugmented LLMs, they fall short of
Record details
Published: 15 March 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Emotional Cost Functions for AI Safety: Teaching Agents to Feel the Weight of Irreversible Consequences
arXiv · 15 March 2026
PA3: Policy-Aware Agent Alignment through Chain-of-Thought
arXiv · 15 March 2026
Cryptographic Runtime Governance for Autonomous AI Systems: The Aegis Architecture for Verifiable Policy Enforcement
arXiv · 15 March 2026
Directional Embedding Smoothing for Robust Vision Language Models
arXiv · 16 March 2026
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
arXiv · 16 March 2026
OMNIFLOW: A Physics-Grounded Multimodal Agent for Generalized Scientific Reasoning
arXiv · 16 March 2026
How to cite this record
ethics.ai (15 March 2026), “Questionnaire Responses Do not Capture the Safety of AI Agents,” evidence record 7217, https://ethics.ai/record/7217 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.