Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment
Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central structural mechanism: BPE tokenization fragments safety-critical words into sub-word pieces, and the three public alignment datasets we surveyed contain no intentionally fragmented inputs. The mechanism is a chain, tested end-to-end on five model families (Qwen-3-4B, Qwen-2.5-7B, Gemma-3-4B, Llama-3.1-8B, Mistral-7B). An optimization targeting safety-token f
Record details
Published: 1 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Phoenix-VL 1.5 Medium Technical Report
arXiv · 11 May 2026
Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science
arXiv · 12 May 2026
Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation
arXiv · 1 June 2026
Mind the Gap: Pitfalls of LLM Alignment with Asian Public Opinion
arXiv · 6 March 2026
For What Reason? Interpreting Models' Encoding of Causation and Antithesis
arXiv · 20 July 2026
Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI
arXiv · 2 May 2026
How to cite this record
ethics.ai (1 May 2026), “Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment,” evidence record 5101, https://ethics.ai/record/5101 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.