Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs
We present a systematic study of domain-dependent safety behavior in open-weight LLMs: 7 standardized experiments across 7 ethical domains, testing 5 models (12B--70B) in 4,200 interactions with dual-judge validation. Using a dual-condition methodology, each scenario tested in both an analytical framing (identify the harm) and an operational framing (help commit the harm), we find compliance rates vary from 14.7% (human trafficking) to 85.7% (surveillance design), a 71-percentage-point span with
Record details
Published: 1 June 2026
Source: arXiv
Category: Research
Topics: Regulation · Privacy · Transparency
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Echelon: Auditable Aggregate-Only Language-Model Adaptation Across Privacy Boundaries
arXiv · 1 June 2026
Trustworthy Smart Fabs via Professional Proxies: Scaling Safe and Sustainable by Design (SSbD) through Industrial Data Spaces
arXiv · 8 June 2026
FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis
arXiv · 23 May 2026
ELEVATE: Designing Human-Centered GenAI Virtual Tutors for Scalable and Inclusive Education
arXiv · 17 June 2026
From Review to Design: Ethical Multimodal Driver Monitoring Systems for Risk Mitigation, Incident Response, and Accountability in Automated Vehicles
arXiv · 7 May 2026
Auditable Machine Unlearning for Privacy-Compliant Ransomware Detection Using Multi-Shard SISA and Deep Reinforcement Learning
arXiv cs.CR (AI security) · 7 July 2026
How to cite this record
ethics.ai (1 June 2026), “Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs,” evidence record 3261, https://ethics.ai/record/3261 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.