SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
Current LLM-based agents demonstrate strong performance in episodic task execution but remain constrained by static toolsets and episodic amnesia, failing to accumulate experience across task boundaries. This paper formalizes the Self-Evolving Agent (SEA) from the perspective of digital embodiment and continuous cross-task evolution, introduces the Evolutionary Flywheel as its minimal sufficient architecture, and presents SEA-Eval -- the first benchmark designed specifically for evaluating SEAs.
Record details
Published: 10 April 2026
Source: arXiv
Category: Research
Topics: Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
PilotBench: A Benchmark for General Aviation Agents with Safety Constraints
arXiv · 10 April 2026
Aligned Agents, Biased Swarm: Measuring Bias Amplification in Multi-Agent Systems
arXiv · 10 April 2026
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
arXiv · 10 April 2026
Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition
arXiv · 10 April 2026
ECM Contracts: Contract-Aware, Versioned, and Governable Capability Interfaces for Embodied Agents
arXiv · 10 April 2026
MPAC: A Multi-Principal Agent Coordination Protocol for Interoperable Multi-Agent Collaboration
arXiv · 10 April 2026
How to cite this record
ethics.ai (10 April 2026), “SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment,” evidence record 6081, https://ethics.ai/record/6081 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.