VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}etrieval \textbf{A}gents), a benchmark of over $8{,}000$ executable APIs across $62$ domains with tasks spanning three settings of increasing difficulty: diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source rea
Record details
Published: 12 August 2026
Source: arXiv cs.AI
Category: Research
Topics: Agents & autonomy
Retrieved: 13 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence
arXiv cs.AI · 12 August 2026
Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents
arXiv · 12 August 2026
DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation
arXiv · 12 August 2026
AI isn’t ready to research itself
Nature Machine Intelligence · 12 August 2026
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
arXiv cs.AI · 12 August 2026
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS
arXiv cs.AI · 12 August 2026
How to cite this record
ethics.ai (12 August 2026), “VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies,” evidence record 19024, https://ethics.ai/record/19024 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.