Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?
Mechanistic interpretability is often motivated for alignment auditing, where a model's verbal explanations can be absent, incomplete, or misleading. Yet many evaluations do not control whether black-box prompting alone can recover the target behavior, so apparent gains from white-box tools may reflect elicitation rather than internal signal; we call this the elicitation confounder. We introduce Pando, a model-organism benchmark that breaks this confound via an explanation axis: models are train
Record details
Published: 13 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Transparency
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
A Proposed Biomedical Data Policy Framework to Reduce Fragmentation, Improve Quality, and Incentivize Sharing in Indian Healthcare in the era of Artificial Intelligence and Digital Health
arXiv · 13 April 2026
A Framework for Longitudinal Health AI Agents
arXiv · 13 April 2026
LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines
arXiv · 14 April 2026
From Helpful to Trustworthy: LLM Agents for Pair Programming
arXiv · 11 April 2026
From Reactive to Proactive: A Multi-Regulatory Empirical Analysis of 480 AI Incidents and a Data-Driven Governance Compliance Framework
arXiv · 10 April 2026
Decentralized autonomous organization and blockchain-based incentivization framework for community-based facilities management
arXiv · 16 April 2026
How to cite this record
ethics.ai (13 April 2026), “Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?,” evidence record 5963, https://ethics.ai/record/5963 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.