Home | Publications | XHH+26

Pose Anything Anywhere: Model-Free Object Poses From Arbitrary References

MCML Authors

Abstract

Estimating the 6D pose of unseen objects is a fundamental yet challenging problem for open-world robotics and embodied perception. Model-based methods are accurate but depend on CAD assets or heavy onboarding, while most model-free approaches are still limited to pairwise single-anchor matching and thus fail under occlusion and large viewpoint changes with low query-reference overlap. Therefore, we present PANY, a unified model-free framework that seamlessly supports both RGB and RGB-D inputs, operates on one or sparse pose-free reference views, and generalizes effectively to novel objects. Built on a multi-view transformer geometry backbone, PANY moves beyond pairwise matching by learning view-consistent geometry and cross-view alignment cues that remain stable under wide baselines and limited overlap. When additional unposed assist views are available, PANY aggregates them via pose-graph canonical registration to increase geometric coverage and reinforce the final pose. Extensive experiments show that PANY achieves state-of-the-art performance across multiple benchmarks, substantially outperforming existing model-free methods, improving pose accuracy by +12% on YCB-V and over +20% on LM-O. Furthermore, PANY consistently performs well under both single-reference and sparse-reference settings, demonstrating strong robustness in real-world environments.

inproceedings XHH+26


ECCV 2026

19th European Conference on Computer Vision. Malmö, Sweden, Sep 08-12, 2026. To be published. Preprint available.
Conference logo
A* Conference

Authors

H. Xu • J. Hu • J. Huang • B. Zhong • P. K. Yu • N. NavabB. Busam • S. Ilic

Links

arXiv

Research Areas

 B1 | Computer Vision

 C1 | Medicine

BibTeXKey: XHH+26

Back to Top