Home | Publications | SBB+26b

Show Me Examples: Inferring Visual Concepts From Image Sets

MCML Authors

Abstract

Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current models fail to infer shared concepts from sets of example images and apply them to new inputs. We introduce Visual Concept Inference from Sets (VICIS), a task that evaluates this capability. Given a small context set of images sharing a concept and a query image, the model must generate new images that preserve the context-defined concept while remaining consistent with the query. We show that state-of-the-art VLMs perform poorly on this task, often ignoring the visual context or defaulting to biased generations. To address this gap, we propose a training framework and architecture that learn to infer visual concepts from image sets and extract concept-specific embeddings from queries. Experiments on synthetic data and large-scale ImageNet/WordNet data show that our model generates more accurate and diverse outputs and generalizes to unseen concepts and modalities such as sketches.

inproceedings SBB+26b


ECCV 2026

19th European Conference on Computer Vision. Malmö, Sweden, Sep 08-12, 2026. To be published. Preprint available.
Conference logo
A* Conference

Authors

N. StrackeK. Bauer • S. A. Baumann • M. A. Bautista • J. Susskind • B. Ommer

Links

arXiv

Research Area

 B1 | Computer Vision

BibTeXKey: SBB+26b

Back to Top