On-Policy and Off-Policy Learning for Large Action Spaces
This thesis studies policy learning in interactive systems where an agent observes a context, selects an action from a very large set, and receives partial feedback. The main framework is contextual bandits, with two paradigms: on-policy learning, where the agent interacts sequentially with the environment and minimizes regret, and off-policy learning, where it learns from logged data collected by a logging policy. In large action spaces, both settings face major challenges: inefficient explorat
Record details
Published: 30 July 2026
Source: arXiv cs.AI
Category: Research
Topics: Regulation · Agents & autonomy · Environment
Retrieved: 31 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
arXiv · 30 July 2026
MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation
arXiv cs.AI · 31 July 2026
Harness-G: A Graph-Structured Harness for Search Agents
HuggingFace Daily Papers · 29 July 2026
Beyond Component Testing: Validating Agentic AI Systems
arXiv cs.AI · 31 July 2026
Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models
arXiv · 29 July 2026
Constitutional governance for societies of AI agents in the built environment: a research agenda
arXiv cs.CY · 28 July 2026
How to cite this record
ethics.ai (30 July 2026), “On-Policy and Off-Policy Learning for Large Action Spaces,” evidence record 15223, https://ethics.ai/record/15223 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.