Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
Multilingual Large Language Models (LLMs) struggle with cross-lingual tasks due to data imbalances between high-resource and low-resource languages, as well as monolingual bias in pre-training. Existing methods, such as bilingual fine-tuning and contrastive alignment, can improve cross-lingual performance, but they often require extensive parallel data or suffer from instability. To address these challenges, we introduce a Cross-Lingual Mapping Task during the pre-training phase, which enhances
Record details
Published: 12 April 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool Invocations
arXiv · 13 April 2026
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
arXiv · 13 April 2026
Retrieval-Augmented Generation Must Move Beyond Factual Grounding to Represent Diverse Opinions
arXiv · 13 April 2026
SOAR: Self-Correction for Optimal Alignment and Refinement in Diffusion Models
arXiv · 14 April 2026
Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment
arXiv · 14 April 2026
Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition
arXiv · 10 April 2026
How to cite this record
ethics.ai (12 April 2026), “Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance,” evidence record 6003, https://ethics.ai/record/6003 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.