Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs
Tool learning enables LLMs to invoke external tools to accomplish tasks. Prior studies have demonstrated the effectiveness of a hierarchical structure: a high-level policy handles global planning and decomposes tasks into manageable sub-tasks, and a low-level policy focuses on invoking tools to solve these sub-tasks. However, these works typically optimize the high-level and low-level policies separately, leading to planner-executor misalignment and limiting LLM performance on tool-use tasks. In
Record details
Published: 8 June 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts
arXiv · 9 June 2026
PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation
arXiv · 7 June 2026
RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations
arXiv · 7 June 2026
The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In
arXiv · 6 June 2026
CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy
arXiv · 10 June 2026
Should LLM Agents Decide in Social Simulations? Comparing Finite-State and LLM-Based Decision Policies
arXiv · 10 June 2026
How to cite this record
ethics.ai (8 June 2026), “Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs,” evidence record 1260, https://ethics.ai/record/1260 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.