PRISM Risk Signal Framework: Hierarchy-Based Red Lines for AI Behavioral Risk
Current approaches to AI safety define red lines at the case level: specific prompts, specific outputs, specific harms. This paper argues that red lines can be set more fundamentally -- at the level of value, evidence, and source hierarchies that govern AI reasoning. Using the PRISM (Profile-based Reasoning Integrity Stack Measurement) framework, we define a taxonomy of 27 behavioral risk signals derived from structural anomalies in how AI systems prioritize values (L4), weight evidence types (L
Record details
Published: 13 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
AI Integrity: A New Paradigm for Verifiable AI Governance
arXiv · 13 April 2026
Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?
arXiv · 13 April 2026
FlowCoMotion: Text-to-Motion Generation via Token-Latent Flow Modeling
arXiv · 13 April 2026
Efficient Training for Cross-lingual Speech Language Models
arXiv · 13 April 2026
A Proposed Biomedical Data Policy Framework to Reduce Fragmentation, Improve Quality, and Incentivize Sharing in Indian Healthcare in the era of Artificial Intelligence and Digital Health
arXiv · 13 April 2026
An ontological approach to foster the convergence, interoperability and operationalization of frameworks for Trustworthy AI
arXiv · 13 April 2026
How to cite this record
ethics.ai (13 April 2026), “PRISM Risk Signal Framework: Hierarchy-Based Red Lines for AI Behavioral Risk,” evidence record 5961, https://ethics.ai/record/5961 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.