Home | Publications | HHX26

SOS! : A Streamlined Object-Conditional Transformer for Model-Free Segmentation

MCML Authors

Abstract

Foundation segmentation models excel at generating high-quality, class-agnostic masks, but they struggle to associate these proposals with specific target objects. This semantic gap severely hinders their deployment in downstream applications like robotic manipulation, which demand precise unseen objects segmentation. Existing approaches attempt to resolve this by relying on exhaustive 3D object model priors, inherently introducing prohibitive computational overhead and complex, multi-stage pipelines. To address these limitations, we propose SOS (Streamlined Object-conditional Transformer for model-free Segmentation). SOS completely eliminates the reliance on 3D models, requiring only a single reference image per target object. Central to our framework is a novel Object-Conditional Transformer that learns identity-anchored queries, unifying mask generation and target identification into a single feed-forward pass. This streamlined design drastically improves both structural and computational efficiency. Extensive evaluations across multiple benchmarks demonstrate that SOS establishes a new state-of-the-art for model-free unseen objects segmentation, delivering accurate and high-efficiency performance.

inproceedings HHX26


BMVC 2026

37th British Machine Vision Conference. Lancaster, UK, Nov 23-26, 2026. To be published. Preprint available.
Conference logo
A Conference

Authors

J. Hu • J. Huang • H. Xu • P. K. Yu • N. Navab • B. Busam • S. Ilic

Links

arXiv GitHub

Research Areas

 B1 | Computer Vision

 C1 | Medicine

BibTeXKey: HHX26

Back to Top