Home | Publications | LGD+26

Object Pose Transformer: Unifying Unseen Object Pose Estimation

MCML Authors

Abstract

Learning model-free object pose estimation for unseen instances remains a central problem in 3D vision. Existing methods are typically separated into two paradigms: category-level approaches predict absolute poses in a canonical space but depend on predefined categories, while relative pose methods estimate transformations across views but cannot recover absolute pose from a single image. In this work, we propose Object Pose Transformer (OPT-Pose), a unified feedforward framework that connects these paradigms through task factorization within a single model. OPT-Pose jointly predicts depth, point maps, camera parameters, and normalized object coordinates (NOCS) from RGB inputs, enabling both category-level absolute SA(3) pose estimation and unseen-object relative SE(3) pose estimation. The method uses contrastive object-centric latent embeddings to achieve canonicalization without semantic labels at inference time, and represents geometry using point maps in camera space to support multi-view reasoning. Through cross-frame feature interaction and shared object embeddings, the model exploits relative geometric consistency across views to improve absolute pose estimation and reduce ambiguity in single-view predictions. In addition, OPT-Pose is camera-agnostic, learning camera intrinsics during inference and supporting optional depth input for metric scale recovery, while remaining fully functional in RGB-only settings. Extensive experiments on NOCS, HouseCat6D, Omni6DPose, and Toyota-Light show state-of-the-art performance on both absolute and relative pose estimation tasks within a single unified architecture.

inproceedings LGD+26


Findings @CVPR 2026

Findings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Denver, CO, USA, Jun 03-07, 2026. To be published. Preprint available.
Conference logo

Authors

W. Li • L. Garattoni • F. Despinoy • N. NavabB. Busam

Links

URL

Research Areas

 B1 | Computer Vision

 C1 | Medicine

BibTeXKey: LGD+26

Back to Top