FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also follow human instructions about how those tasks should be executed. However, existing robot datasets usually pair trajectories with coarse goal-level language, leaving execution-critical details such as active arm, approach direction, and contact region unspecified. This limits steerable policy learning and robotic video understanding. We introduce FineVLA, an open framework for action-aligne
Record details
Published: 26 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Systems
arXiv · 26 May 2026
OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation
arXiv · 25 May 2026
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
arXiv · 28 May 2026
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
arXiv · 22 May 2026
Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness
arXiv · 22 May 2026
REBAR: Reference Ethical Benchmark for Autonomy Readiness
arXiv · 18 May 2026
How to cite this record
ethics.ai (26 May 2026), “FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies,” evidence record 3660, https://ethics.ai/record/3660 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.