EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
Gradient-based preference optimization methods for large language model (LLM) alignment suffer from preference collapse, converging to narrow behavioral modes while neglecting preference diversity. We introduce EvoPref, a multi-objective evolutionary algorithm that maintains populations of Low-Rank Adaptation (LoRA) adapters optimized across helpfulness, harmlessness, and honesty objectives using Non-dominated Sorting Genetic Algorithm II (NSGA-II) selection with archive-based diversity preserva
Record details
Published: 10 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Brain-LLM Alignment Tracks Training Data, Not Typology
arXiv · 21 May 2026
Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography
arXiv · 21 May 2026
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
arXiv · 20 April 2026
SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters
arXiv · 1 May 2026
Flag Varieties: A Geometric Framework for Deep Network Alignment
arXiv · 11 May 2026
Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions
arXiv · 11 May 2026
How to cite this record
ethics.ai (10 May 2026), “EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent,” evidence record 4607, https://ethics.ai/record/4607 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.