Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning
Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models and their misalignment with natural language usage. These misalignments are thought to partly originate in the preference-learning stage, e.g. Reinforcement Learning from Human Feedback, which generally makes models more useful but simultaneously may introduce systematic lexical bias. In terms of lexical behavior, this is visible in a model's pref
Record details
Published: 29 May 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Improving Visual Representation Alignment Generation with GRPO
arXiv · 30 May 2026
Silent Failures in Federated Personalization of Foundation Models
arXiv · 31 May 2026
When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop
arXiv · 28 May 2026
TriAlign: Towards Universal Truth Consistency in Personalized LLM Alignment
arXiv · 1 June 2026
Algorithmic Fragility and Persona Bias in LLM-Generated Autistic Communication
arXiv · 26 May 2026
DeGRe: Dense-supervised Generative Reranking for Recommendation
arXiv · 25 May 2026
How to cite this record
ethics.ai (29 May 2026), “Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning,” evidence record 3414, https://ethics.ai/record/3414 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.