PilotBench: A Benchmark for General Aviation Agents with Safety Constraints
As Large Language Models (LLMs) advance toward embodied AI agents operating in physical environments, a fundamental question emerges: can models trained on text corpora reliably reason about complex physics while adhering to safety constraints? We address this through PilotBench, a benchmark evaluating LLMs on safety-critical flight trajectory and attitude prediction. Built from 708 real-world general aviation trajectories spanning nine operationally distinct flight phases with synchronized 34-c
Record details
Published: 10 April 2026
Source: arXiv
Category: Research
Topics: Agents & autonomy · Environment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
arXiv · 10 April 2026
Semantic Rate-Distortion for Bounded Multi-Agent Communication: Capacity-Derived Semantic Spaces and the Communication Cost of Alignment
arXiv · 10 April 2026
ACF: A Collaborative Framework for Agent Covert Communication under Cognitive Asymmetry
arXiv · 9 April 2026
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
arXiv · 9 April 2026
Learning Without Losing Identity: Capability Evolution for Embodied Agents
arXiv · 9 April 2026
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
arXiv · 12 April 2026
How to cite this record
ethics.ai (10 April 2026), “PilotBench: A Benchmark for General Aviation Agents with Safety Constraints,” evidence record 6082, https://ethics.ai/record/6082 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.