Home | News

05.03.2026

Teaser image to Foundations of Diffusion: One Map for Images and Text

Foundations of Diffusion: One Map for Images and Text

MCML Research Insight – With Vincent Pauline, Tobias Höppe, Andrea Dittadi, and Stefan Bauer

From hyper-realistic video generation to protein design, Diffusion Models are the engine behind the current wave of Generative AI. But if you try to understand how diffusion models actually work, you quickly run into heavy math. Most explanations focus only on images. If you’re working with text, biology, or graphs, you’re often left alone with scattered papers and technical jargon.

«Diffusion models are central to generative AI, yet most introductions only cover Euclidean data and seldom clarify their connection to discrete-state analogues.»


Andrea Dittadi

MCML Junior Member

This is where the MCML authors Vincent Pauline, Tobias Höppe, Andrea Dittadi, and Stefan Bauer, along with collaborators Kirill Neklyudov, Alexander Tong, step in. In their paper the authors ask: Can we explain all diffusion models — for images and text — in one clean, unified way?

Problem Statement: Two Worlds, One Theory?

Diffusion models are often taught in pieces. The theory around it is fragmented into:

  • The Euclidean Bias: Most introductions explain the image case, because images live in smooth, continuous space and the math is well-established.
  • The Discrete Gap: For text, graphs, or biological sequences, the explanations are scattered across different papers and use different notation. As a result, learners often bounce between beginner-friendly blogs and very technical textbooks and struggle to see that the same core idea connects both.

This creates a “knowledge silo” where researchers expert in image generation struggle to translate their intuition to sequence modeling, and newcomers are overwhelmed by the mathematical terminology.

Young researcher reading Foundations of Diffusion Models in General State Spaces: A Self-Contained Introduction.

Highlights: The Visual “Cheat Code” for Diffusion

A big reason this handbook is easier to follow is how it is written. The authors show us that continuous and discrete diffusion match.

1. Parallel Presentation

Key ideas are presented side-by-side so you can instantly see what corresponds to what:

🔵 Blue boxes: the “image version” (continuous space)

🔴 Red boxes: the “text/sequence version” (discrete space)

🟡 Yellow boxes: ideas that apply to both

This means we can read the paper like a guidebook: follow everything end-to-end, or focus on the parts that match our background.


2. A Roadmap for Every Reader

The paper is structured so different readers can jump in at the right depth:

🟢 Introductory: A “Diffusion Primer” that builds intuition on why variational inference works, leading into the discrete-time formulation and training objectives without needing heavy prerequisites.

🟠 Advanced: Already know DDPMs? Jump straight to continuous-time theory: SDEs, CTMCs, Fokker-Planck & master equations.

🟣 Expert: presents the most general “all-in-one” framework for diffusion across different data types.

Visual roadmap of the manuscript.

3. The Unifying Lens: The Infinitesimal Generator

At the heart of the paper (Section 7) lies the mathematical unification. The authors utilize the Markov process infinitesimal generator to show that continuous diffusion for images and discrete diffusion for text are just special cases of the same formalism.

Instead of two separate theories, we get:

  • one way to describe how data gets corrupted over time,
  • one way to describe how to reverse that corruption,
  • and one common training view (through an ELBO-style objective) that works across settings.

Why This Matters

By providing a single, coherent story that connects VAEs, score-based models, latent diffusion, and flow matching, this work lowers the barrier to entry for discrete generative modeling. It empowers researchers to port the same techniques from the image-generation domain directly to discrete problems like language modeling and biology.

Unified perspective on di!usion models in continuous and discrete state spaces.

Further Reading & Reference

Interested in mastering the math of diffusion?

V. PaulineT. Höppe • K. Neklyudov • A. Tong • S. BauerA. Dittadi
Foundations of Diffusion Models in General State Spaces: A Self-Contained Introduction.
Preprint (Dec. 2025). arXiv

Share Your Research!


Get in touch with us!

Are you an MCML Junior Member and interested in showcasing your research on our blog?

We’re happy to feature your work—get in touch with us to present your paper.

#blog #research #bauer-s

Related

Link to MCML Welcomes Student Delegation from HEC Montréal

22.07.2026

MCML Welcomes Student Delegation From HEC Montréal

MCML welcomed students from HEC Montréal for discussions on AI research, ethics, and international academic collaboration.

Read more
Link to Timo Heiß Receives Best Student Paper Award at XAI 2026

21.07.2026

Timo Heiß Receives Best Student Paper Award at XAI 2026

Timo Heiß receives the Best Student Paper Award at XAI 2026 for research on improving feature effect estimation in explainable AI.

Read more
Link to The Learning Rate Does More Than Set the Pace

21.07.2026

The Learning Rate Does More Than Set the Pace

New ICML 2026 research by Gitta Kutyniok and her team shows how learning rates balance competing biases that shape neural network generalization.

Read more
Link to How fair does AI seem in job interviews?

20.07.2026

How Fair Does AI Seem in Job Interviews?

A new study by Enkelejda Kasneci examines how applicants perceive fairness in AI-driven job interviews.

Read more
Link to TUM develops intelligent 3D twins of crime scenes

20.07.2026

TUM Develops Intelligent 3D Twins of Crime Scenes

AI-powered 3D twins enable investigators to analyze virtual crime scenes through interactive object recognition and scene understanding.

Read more
Back to Top