Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation
Object pose estimation is a fundamental task in 3D vision with applications in robotics, AR/VR, and scene understanding. We address the challenge of category-level 9-DoF pose estimation (6D pose + 3Dsize) from RGB-D input, without relying on CAD models during inference. Existing depth-only methods achieve strong results but ignore semantic cues from RGB, while many RGB-D fusion models underperform due to suboptimal cross-modal fusion that fails to align semantic RGB cues with 3D geometric repres
Record details
Published: 29 March 2026
Source: arXiv
Category: Research
Topics: Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
MoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement
arXiv · 29 March 2026
Learning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning
arXiv · 29 March 2026
Umwelt Engineering: Designing the Cognitive Worlds of Linguistic Agents
arXiv · 29 March 2026
A Reference Architecture for Agentic Hybrid Retrieval in Dataset Search
arXiv · 28 March 2026
The Novelty Bottleneck: A Framework for Understanding Human Effort Scaling in AI-Assisted Work
arXiv · 28 March 2026
Grounding Social Perception in Intuitive Physics
arXiv · 28 March 2026
How to cite this record
ethics.ai (29 March 2026), “Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation,” evidence record 6626, https://ethics.ai/record/6626 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.