E-TCAV: Formalizing Penultimate Proxies for Efficient Concept Based Interpretability
TCAV (Testing with Concept Activation Vectors) is an interpretability method that assesses the alignment between the internal representations of a trained neural network and human-understandable, high-level concepts. Though effective, TCAV suffers from significant computational overhead, inter-layer disagreement of TCAV scores, and statistical instability. This work takes a step toward addressing these challenges by introducing E-TCAV, a framework for efficient approximation of TCAV scores, whic
Record details
Published: 11 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
arXiv · 11 May 2026
Positive Alignment: Artificial Intelligence for Human Flourishing
arXiv · 11 May 2026
ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design
arXiv · 11 May 2026
Phoenix-VL 1.5 Medium Technical Report
arXiv · 11 May 2026
Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance
arXiv · 11 May 2026
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
arXiv · 11 May 2026
How to cite this record
ethics.ai (11 May 2026), “E-TCAV: Formalizing Penultimate Proxies for Efficient Concept Based Interpretability,” evidence record 4571, https://ethics.ai/record/4571 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.