Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis
Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used. This letter proposes a general framework in order to conduct task-conditional faithfulness audit. It compares self-reported reliance, intervention-derived behavioral reliance, and preregistered engineering importance. The framework first registers task-specific evidence requirements and compares them with
Record details
Published: 27 July 2026
Source: arXiv cs.AI
Category: Research
Topics: Healthcare · Transparency
Retrieved: 28 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings
arXiv cs.AI · 27 July 2026
KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability
arXiv cs.AI · 27 July 2026
Fairness Interventions in Classification: A Study on AI Explainability
arXiv cs.CY · 28 July 2026
Principles and Guidelines for Randomized Controlled Trials in AI Evaluation
arXiv cs.CY · 28 July 2026
Auditing Institutional Heterogeneity for Generative AI in Patient Education: A Large-Scale Study of 102 US Transplant Handbooks
arXiv cs.CY · 28 July 2026
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
arXiv cs.AI · 28 July 2026
How to cite this record
ethics.ai (27 July 2026), “Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis,” evidence record 14044, https://ethics.ai/record/14044 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.