Home | Publications | KMR+26

MAGneT-3D: Monocular and Domain-Generalizable Temporal 3D Detection

MCML Authors

Abstract

Monocular temporal 3D detection aims to detect objects in 3D, given a monocular video. Query-based 3D detectors unify detection and cross-view association, but their learnable queries fit the spatial distribution of the training data (e.g., field-of-view). We show that this issue is especially severe when these models are applied to monocular video, hindering generalization to unseen datasets and environments. To address this limitation, we introduce MAGneT-3D, the first method for domain-generalized monocular temporal 3D object detection. Instead of relying on static learnable queries, we propose a Domain-Robust Anchor Generator (DRAG) approach that adaptively derives 3D proposals during inference. To further enable domain generalization, we propose a Temporal Refinement and Identity Merging (TRIM) strategy, reducing dependence on specific 3D proposals. To enable comprehensive domain-generalization evaluation, we establish a cross-dataset benchmark spanning nuScenes, Waymo, Lyft, and ONCE. Under zero-shot domain shifts, MAGneT-3D outperforms all baselines, improving NDS from 12.1% to 18.6% while also increasing in-domain accuracy.

inproceedings KMR+26


DRIVEX @ECCV 2026

6th Workshop on Foundation Models for Autonomous Driving at the 19th European Conference on Computer Vision. Malmö, Sweden, Sep 08-12, 2026. Oral Presentation.

Authors

M. Kotb • J. Meier • C. Reich • O. Dhaouadi • L. Denninger • D. Cremers

Links

URL GitHub

Research Area

 B1 | Computer Vision

BibTeXKey: KMR+26

Back to Top