Investigating Concept Alignment Using Implausible Category Members
Developing AI systems with a human-like understanding of everyday concepts is a key step towards developing safe, reliable systems whose behavior makes sense to humans. When probing concept understanding, asking questions about plausible category members (e.g., "Is a car a vehicle?") is likely to recall patterns in the model's vast training data. We pursue an alternative strategy, characterizing the boundaries of conceptual categories by asking about implausible category members (e.g., "Is an ol
Record details
Published: 20 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Representation Alignment Rests on Linear Structure
arXiv · 22 May 2026
Semantic Generative Tuning for Unified Multimodal Models
arXiv · 18 May 2026
The Hidden Cost of Contextual Sycophancy: an AI Literacy Intervention in Human-AI Collaboration
arXiv · 18 May 2026
Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling
arXiv · 18 May 2026
Robust Fuzzy Multi-view Learning under View Conflict
arXiv · 23 May 2026
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
arXiv · 24 May 2026
How to cite this record
ethics.ai (20 May 2026), “Investigating Concept Alignment Using Implausible Category Members,” evidence record 3952, https://ethics.ai/record/3952 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.