SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters
AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermediate state between steps and inflating end-to-end latency by 3-8x. We argue that this request-level abstraction is fundamentally mismatched to compound AI workloads, and propose a shift to program-level scheduling: treating the entire agent workflow (not individual inference calls) as the first-class schedulable unit. We present SAGA, a distributed
Record details
Published: 1 May 2026
Source: arXiv
Category: Research
Topics: Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
arXiv · 10 May 2026
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
arXiv · 20 April 2026
Brain-LLM Alignment Tracks Training Data, Not Typology
arXiv · 21 May 2026
Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography
arXiv · 21 May 2026
AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning
arXiv · 1 May 2026
Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines
arXiv · 1 May 2026
How to cite this record
ethics.ai (1 May 2026), “SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters,” evidence record 5123, https://ethics.ai/record/5123 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.