Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
Large language models increasingly mediate decisions that turn on moral judgement, yet a growing body of evidence shows that their implicit preferences are not culturally neutral. Existing cultural alignment methods either require per-country preference data and fine-tuning budgets or assume white-box access to model internals that commercial APIs do not expose. In this work, we focus on this realistic black-box, public-data-only regime and observe that within-country sociodemographic disagreeme
Record details
Published: 11 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
TrajPrism: A Multi-Task Benchmark for Language-Grounded Urban Trajectory Understanding
arXiv · 11 May 2026
Interpretability Can Be Actionable
arXiv · 11 May 2026
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
arXiv · 11 May 2026
The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime
arXiv · 11 May 2026
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
arXiv · 11 May 2026
Leveraging RAG for Training-Free Alignment of LLMs
arXiv · 11 May 2026
How to cite this record
ethics.ai (11 May 2026), “Training-Free Cultural Alignment of Large Language Models via Persona Disagreement,” evidence record 4540, https://ethics.ai/record/4540 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.