SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing
LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a guard model that labels each proposed action as safe or unsafe, but this binary view conflates two distinct decisions: whether the action is harmful in itself, and whether it is appropriate given the user's context. It also operates at the granularity of action categories rather than individual instances, producing routine interruptions that erode a
Record details
Published: 15 July 2026
Source: arXiv cs.AI
Category: Research
Topics: Agents & autonomy · Environment
Retrieved: 16 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Agile perceptive multi-skill locomotion for quadrupedal robots in the wild
arXiv cs.AI · 15 July 2026
UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
arXiv cs.AI · 15 July 2026
From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception
arXiv cs.AI · 15 July 2026
DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
arXiv cs.HC · 15 July 2026
AgentSociety 2: An Integrated Research Environment for Executable Social Science
arXiv cs.CY · 15 July 2026
Object-centric diffusion policies for real-world robotic-arm imitation learning
Frontiers in Robotics and AI · 15 July 2026
How to cite this record
ethics.ai (15 July 2026), “SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing,” evidence record 10919, https://ethics.ai/record/10919 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.