From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents
The integration of external tools has transitioned LLM agents from passive responders to autonomous systems. However, current benchmarks prioritize execution success, neglecting self-awareness capability, the ability to discern whether a problem requires necessary external resources or can be solved via internal parametric knowledge. To address this, we introduce KAPRO (Knowing-Acting Quadrant PRObe), a framework that evaluates cognitive-behavioral alignment by decoupling an agent's metacognitiv
Record details
Published: 9 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment
arXiv · 9 June 2026
Making Foresight Actionable: Repurposing Representation Alignment in World Action Models
arXiv · 10 June 2026
CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy
arXiv · 10 June 2026
Should LLM Agents Decide in Social Simulations? Comparing Finite-State and LLM-Based Decision Policies
arXiv · 10 June 2026
From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation
arXiv · 10 June 2026
Agent Economics: An Entropy-Controlled Pluralistic Alignment Framework for Preventing Artificial Hivemind in Autonomous Agents
arXiv · 8 June 2026
How to cite this record
ethics.ai (9 June 2026), “From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents,” evidence record 1181, https://ethics.ai/record/1181 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.