Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training
Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution strategies (ES) enable memory-efficient full-parameter post-training without backpropagation and can eventually match the performance of gradient-based reinforcement learning (RL). However, resource-constrained settings typically offer only a few GPUs, so the high GPU-hour requirements of ES translate into prohibitively long training times. To addres
Record details
Published: 3 August 2026
Source: arXiv cs.AI
Category: Research
Topics: Agents & autonomy
Retrieved: 4 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Antares: Foundation Models for Agentic Vulnerability Localization
arXiv cs.AI · 3 August 2026
Human-Centered Reflections on Care Robots: A Comparative Study of Caregiver Perspectives
arXiv · 3 August 2026
Faster-WAM: Do World Action Models Need Deep Action Modules?
arXiv cs.AI · 3 August 2026
SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents
arXiv cs.AI · 3 August 2026
Agentic Incident Response through Digital Twin-Enhanced Multiscale Planning
arXiv cs.AI · 3 August 2026
KC-Agent: A Dual-Process Cognitive Architecture for Efficient ML Model Improvement
arXiv cs.AI · 3 August 2026
How to cite this record
ethics.ai (3 August 2026), “Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training,” evidence record 16100, https://ethics.ai/record/16100 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.