What Could the Agent See at 19:05? Generating Temporal Enterprise Scenarios from Real Research and Replaying Them to Evaluate Agents
Enterprise AI agents act across many apps whose data changes continuously, so an answer is correct only relative to what data existed and who could see it at the moment it was asked. Offline evaluation today grades against a single static snapshot, effectively the end of the episode. So, it can only evaluate one situation, the final one, even though every earlier moment of the episode is a different situation that invites its own realistic questions with its own correct answers. Recreating each
Record details
Published: 2 August 2026
Source: arXiv cs.HC
Category: Research
Topics: Agents & autonomy
Retrieved: 4 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
When Collaboration Becomes a Trigger: Collective Evidence-Threshold Backdoors in Multi-Agent Systems
arXiv cs.LG · 2 August 2026
From AI Technical Debt to Agentic Technical Debt: A Systematic Mapping of Root Causes and Manifestations in Agentic AI Systems
arXiv cs.CR (AI security) · 2 August 2026
IR222: From Cyberwar to Killer Robots: Emerging Technology and International Security
LSE Data Science Institute · 2 August 2026
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
HuggingFace Daily Papers · 1 August 2026
MiniWorld: Democratizing the Training of Video World Models from Scratch
HuggingFace Daily Papers · 1 August 2026
Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics
arXiv cs.LG · 2 August 2026
How to cite this record
ethics.ai (2 August 2026), “What Could the Agent See at 19:05? Generating Temporal Enterprise Scenarios from Real Research and Replaying Them to Evaluate Agents,” evidence record 16126, https://ethics.ai/record/16126 (originally published by arXiv cs.HC).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.