To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents
LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models from three families show high call accuracy but much lower no-call accuracy, leaving overall accuracy in the 55%-70% range. We trace this to an Intrinsic Bias Hypothesis (IBH): the call/no-call decision mapping carries an activation-independent call offset, so the model favors call even at activation parity. Using Sparse Autoencoders (SAEs), we rec
Record details
Published: 16 May 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Healthcare · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture
arXiv · 29 April 2026
GeoMind: An Agentic Workflow for Lithology Classification with Reasoned Tool Invocation
arXiv · 23 April 2026
First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows
arXiv · 20 April 2026
AgentFairBench: Do LLM Agents Discriminate When They Act?
arXiv · 15 June 2026
AD-CARE: A Guideline-grounded, Modality-agnostic LLM Agent for Real-world Alzheimer's Disease Diagnosis with Multi-cohort Assessment, Fairness Analysis, and Reader Study
arXiv · 26 March 2026
Front-End Ethics for Sensor-Fused Health Conversational Agents: An Ethical Design Space for Biometrics
arXiv · 14 March 2026
How to cite this record
ethics.ai (16 May 2026), “To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents,” evidence record 4234, https://ethics.ai/record/4234 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.