How Far Are We From True Auto-Research?
Recent auto-research systems can produce complete papers, but feasibility is not the same as quality, and the field still lacks a systematic study of how good agent-generated papers actually are. We introduce ResearchArena, a minimal scaffold that lets off-the-shelf agents (Claude Code using Opus 4.6, Codex using GPT-5.4, and Kimi Code using K2.5) carry out the full research loop themselves (ideation, experimentation, paper writing, self-refinement) under only lightweight guidance. Across 13 com
Record details
Published: 18 May 2026
Source: arXiv
Category: Research
Topics: Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Towards Human-Level Book-Writing Capability
arXiv · 16 May 2026
Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute
arXiv · 14 May 2026
Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows
arXiv · 4 June 2026
IceBreaker for Conversational Agents: Breaking the First-Message Barrier with Personalized Starters
arXiv · 20 April 2026
IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO
arXiv · 22 June 2026
The Persistent Vulnerability of Aligned AI Systems
arXiv · 31 March 2026
How to cite this record
ethics.ai (18 May 2026), “How Far Are We From True Auto-Research?,” evidence record 4075, https://ethics.ai/record/4075 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.