Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents
Adaptive test-time compute for LLM agents aims to invoke extra computation only when it improves performance. Existing methods typically use confidence-, uncertainty-, or difficulty-based gates, assuming a fixed direction from the gating signal through compute need to the value of computation. This makes gating a utility-calibration problem: gating signals should align with whether extra computation improves the final outcome over the base policy. We show that this alignment is unstable: the sam
Record details
Published: 7 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Containment Verification: AI Safety Guarantees Independent of Alignment
arXiv · 9 May 2026
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
arXiv · 5 May 2026
Guided Streaming Stochastic Interpolant Policy
arXiv · 11 May 2026
On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment
arXiv · 12 May 2026
AI Alignment via Incentives and Correction
arXiv · 2 May 2026
Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents
arXiv · 28 April 2026
How to cite this record
ethics.ai (7 May 2026), “Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents,” evidence record 4798, https://ethics.ai/record/4798 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.