Explaining Black-Box Language Models: Learning to Optimize Linguistically-Structured Word Subsets
As deep language models (DLMs) are increasingly deployed in high-stakes domains such as healthcare, understanding their decision rationale becomes paramount for ensuring trust, safety, and accountability. However, achieving this vital level of interpretability is particularly challenging when these DLMs operate as black-box systems (e.g., via APIs), where access to internal model states (e.g., parameters, gradients) is restricted. Despite numerous efforts, existing explanation methods often fail
Record details
Published: 7 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Healthcare · Transparency
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Graph-of-Differences: Anatomy-Structured Difference Alignment for Medical Image Re-Identification
arXiv · 19 June 2026
Cohort-Anchored Foundation Models for Electronic Health Records: From Risk Scores to Auditable Peer Cohorts
arXiv · 20 June 2026
Uncertainty Reasoning with Large Language Models for Explainable Disease Diagnosis
arXiv · 25 May 2026
Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant
arXiv · 23 June 2026
SAM-Sode: Towards Faithful Explanations for Tiny Bacteria Detection
arXiv · 20 May 2026
Do Vision Models Truly Forget? New Findings from Representation-Level Certification of Visual Unlearning in Vertical Federated Learning
arXiv · 19 May 2026
How to cite this record
ethics.ai (7 June 2026), “Explaining Black-Box Language Models: Learning to Optimize Linguistically-Structured Word Subsets,” evidence record 1302, https://ethics.ai/record/1302 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.