MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents
We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, multi-turn tasks where agents must ground progressively arriving visual inputs into executable tool calls while handling realistic conversational phenomena (goal revisions, error corrections, state mutations). An automated scenario generation pipeline produces
Record details
Published: 13 July 2026
Source: arXiv cs.AI
Category: Research
Topics: Agents & autonomy · Environment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
PalmClaw: A Native On-Device Agent Framework for Mobile Phones
HuggingFace Daily Papers · 13 July 2026
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
HuggingFace Daily Papers · 13 July 2026
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
HuggingFace Daily Papers · 13 July 2026
ERR@HRI 3.0 Challenge: Multimodal Detection of Errors and Anticipation in Human-Robot Interactions
arXiv cs.HC · 13 July 2026
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy
arXiv · 13 July 2026
From Tools to Teacher-Built Teammates: No-Code Pedagogical Plugin Authoring with LearnAdapt Agentic Studio and PedOS 1.1 Lumina
arXiv cs.CY · 14 July 2026
How to cite this record
ethics.ai (13 July 2026), “MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents,” evidence record 3009, https://ethics.ai/record/3009 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.