Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation
End-to-end vision-language navigation (VLN) with causal vision-language models can map instructions and egocentric observations directly to actions, but standard behavior cloning supervises only the next action and does not explicitly train the policy state to be predictive of future visual outcomes. We first ask a diagnostic question: if the policy is given an expert-trajectory future image as privileged input at training and testing time, is that additional visual evidence useful for choosing
Record details
Published: 20 July 2026
Source: arXiv
Category: Research
Topics: Regulation · Healthcare
Retrieved: 21 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Clinical Audit Logs as Multi-Axial Traces of Care Delivery
arXiv cs.CY · 20 July 2026
A Diagnostic Framework for AI Agent Behavior
arXiv cs.CY · 21 July 2026
Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development
arXiv · 21 July 2026
Digital Transformation in Health Care: Are We on the Right Track?
JMIR (Journal of Medical Internet Research) · 21 July 2026
A Diagnostic Framework for AI Agent Behavior
arXiv · 19 July 2026
Balancing public health and individual autonomy: a study of Chinas vaccination policy
Journal of Medical Ethics (BMJ) · 22 July 2026
How to cite this record
ethics.ai (20 July 2026), “Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation,” evidence record 11971, https://ethics.ai/record/11971 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.