TraceGuard: Structured Multi-Dimensional Monitoring as a Collusion-Resistant Control Protocol
AI control protocols use monitors to detect attacks by untrusted AI agents, but standard single-score monitors face two limitations: they miss subtle attacks where outputs look clean but reasoning is off, and they collapse to near-zero safety when the monitor is the same model as the agent (collusion). We present TraceGuard, a structured multi-dimensional monitoring protocol that evaluates agent actions across five dimensions -- goal alignment, constraint adherence, reasoning coherence, safety a
Record details
Published: 5 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Coupled Control, Structured Memory, and Verifiable Action in Agentic AI (SCRAT -- Stochastic Control with Retrieval and Auditable Trajectories): A Comparative Perspective from Squirrel Locomotion and Scatter-Hoarding
arXiv · 3 April 2026
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
arXiv · 6 April 2026
IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents
arXiv · 6 April 2026
Cognitive Comparability and the Limits of Governance: Evaluating Authority Under Radical Capability Asymmetry
arXiv · 3 April 2026
DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning
arXiv · 3 April 2026
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
arXiv · 8 April 2026
How to cite this record
ethics.ai (5 April 2026), “TraceGuard: Structured Multi-Dimensional Monitoring as a Collusion-Resistant Control Protocol,” evidence record 6356, https://ethics.ai/record/6356 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.