Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free quality estimation models and is gated by language identification. We then linearly interpolate the supervised fine-tuning (SFT) and reinforcement learning (RL) model checkpoints to obtain MiLMMT-46-v1.0. Across 46 languages, the re
Record details
Published: 10 August 2026
Source: HuggingFace Daily Papers
Category: Research
Topics: Regulation
Retrieved: 12 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
arXiv cs.AI · 11 August 2026
Social Media as a Driver of Obesity in Children and Adolescents (Aged 6-18 Years): It Is Time for Regulatory Action
JMIR (Journal of Medical Internet Research) · 10 August 2026
NIH limits funding for research on the health effects of public policy
Nature Machine Intelligence · 10 August 2026
Mismatch Matters: On-Policy Distillation Beyond Token Agreement
arXiv cs.AI · 10 August 2026
Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation
arXiv cs.AI · 10 August 2026
SR-OPSD: Self-Referenced On-Policy Self-Distillation
arXiv cs.AI · 10 August 2026
How to cite this record
ethics.ai (10 August 2026), “Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation,” evidence record 18376, https://ethics.ai/record/18376 (originally published by HuggingFace Daily Papers).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.