ContextWeave: A Real-World Workflow Benchmark
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, wit
Record details
Published: 5 August 2026
Source: arXiv
Category: Research
Topics: Privacy · Agents & autonomy
Retrieved: 6 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent
arXiv cs.AI · 5 August 2026
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
HuggingFace Daily Papers · 5 August 2026
When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
arXiv cs.CY · 5 August 2026
Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators
arXiv cs.HC · 6 August 2026
Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints
arXiv · 6 August 2026
From 2D image synthesis to 3D scene generation: a comprehensive review of synthetic data for agricultural vision
Artificial Intelligence Review · 31 July 2026
How to cite this record
ethics.ai (5 August 2026), “ContextWeave: A Real-World Workflow Benchmark,” evidence record 16655, https://ethics.ai/record/16655 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.