Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration
Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. Traditional post-hoc temperature scaling is inherently domain-dependent: a temperature fitted on one domain does not generalize across domains. This motivates us to modify model parameters during training to improve calibration. We propose maximizing the entropy of predictive distributions as the calibration objective, which directly targets overconfidence by discouraging overly concentrated predic
Record details
Published: 7 August 2026
Source: arXiv cs.LG
Category: Research
Topics: Safety & alignment
Retrieved: 10 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits
arXiv · 7 August 2026
People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe
arXiv · 7 August 2026
Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation
arXiv · 7 August 2026
Representation Handoffs for OpenArm-Based Laboratory Mobile Manipulation
arXiv · 7 August 2026
The Anatomy of a Prompt Injection: A Component Model for Structured Analysis
arXiv red teaming query · 7 August 2026
MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents
arXiv · 7 August 2026
How to cite this record
ethics.ai (7 August 2026), “Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration,” evidence record 17937, https://ethics.ai/record/17937 (originally published by arXiv cs.LG).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.