Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free quality estimation models and is gated by language identification. We then linearly interpolate the supervised fine-tuning (SFT) and reinforcement learning (RL) model checkpoints to obtain MiLMMT-46-v1.0. Across 46 languages, the re
Record details
Published: 11 August 2026
Source: arXiv cs.AI
Category: Research
Topics: Regulation
Retrieved: 12 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
HuggingFace Daily Papers · 10 August 2026
A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona Problem
arXiv · 11 August 2026
SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense
arXiv red teaming query · 11 August 2026
Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement
arXiv · 11 August 2026
IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning
arXiv cs.LG · 11 August 2026
Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers
arXiv cs.AI · 11 August 2026
How to cite this record
ethics.ai (11 August 2026), “Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation,” evidence record 18662, https://ethics.ai/record/18662 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.