Pigeonholing: Bad prompts hurt models to collapse and make mistakes
While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing." **Unintentionally bad** contexts can happen without malicious jailbreaking intents: For example, a user asks the model to justify an incorrect math theorem or fails to correct the model's buggy code. Specifically, we investigate ``pigeonholing" in two scenarios: (1) when the user suggests a solution, a
Record details
Published: 23 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
ASALT: Adaptive State Alignment for Lateral Transfer in Multi-agent Reinforcement Learning
arXiv · 23 June 2026
What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics
arXiv · 23 June 2026
The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs
arXiv · 21 June 2026
AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing
arXiv · 25 June 2026
Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning
arXiv · 26 June 2026
Editorial Alignment: A Participatory Approach to Engaging Editorial Expertise in LLM-mediated Knowledge Dissemination
arXiv · 18 June 2026
How to cite this record
ethics.ai (23 June 2026), “Pigeonholing: Bad prompts hurt models to collapse and make mistakes,” evidence record 661, https://ethics.ai/record/661 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.