Measuring Safety Alignment Effects in Autonomous Security Agents
Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Single-turn refusal benchmarks cannot answer this question: security agents must inspect repositories, call tools, and produce vulnerability evidence inside authorized sandboxes. We present a trace-based benchmark of 30 local vulnerability-analysis tasks with fixed tools, deterministic success predicates, redaction rules, and grounding checks, and com
Record details
Published: 19 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
arXiv · 17 May 2026
Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains
arXiv · 19 May 2026
Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay
arXiv · 19 May 2026
ContextFlow: Hierarchical Task-State Alignment for Long-Horizon Embodied Agents
arXiv · 19 May 2026
DEFLECT: Temporal Counterfactual Preference Learning for Delay-Robust Asynchronous VLAs
arXiv · 19 May 2026
REBAR: Reference Ethical Benchmark for Autonomy Readiness
arXiv · 18 May 2026
How to cite this record
ethics.ai (19 May 2026), “Measuring Safety Alignment Effects in Autonomous Security Agents,” evidence record 4029, https://ethics.ai/record/4029 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.