TECH FLOW Svět Androida
← Back to the stream
kuleshov-group.github.io · picked by Petr Mišák · 33d ago

How to Build a Diffusion Language Model

Source preview: How to Build a Diffusion Language Model
AI summary

The text introduces diffusion language models and their building blocks. The authors explain how diffusion-based generation differs from autoregressive approaches by generating text in parallel through iterative refinement of the entire sequence at once, and describe techniques for masking, iterative improvement, and training, with examples of recent implementations such as Mercury 2, Gemma Diffusion, and Nemotron Diffusion.

The summary is written by AI from the source; it isn’t the newsroom’s opinion. For details, read the source.

5 people have already opened the source

Tip author’s note

If you've never heard of diffusion AI models, don't worry. But it seems this approach is the future. Current AI models are already smart enough today; what they're lacking is an efficient harness and speed in generating responses. And speed is crucial, because nobody wants to wait. So how do diffusion models do it? Simply put, they don't generate word by word—they generate essentially everything at once.

AI questions & answers
What are the main advantages of diffusion models over autoregressive approaches?

Diffusion models generate sequences in parallel through iterative refinement from an initial approximation, enabling error correction, faster parallel generation, and bidirectional context. Autoregressive models generate left-to-right one token at a time, preventing error correction and being computationally slower.

How does masking work in diffusion language models?

Masking randomly replaces a fraction of tokens with masks in a sequence, and a bidirectional transformer is trained to reconstruct them. Generation proceeds iteratively: the model fills in the masked positions, then a subset of predictions are re-masked with fewer masks in each round until convergence to clean text.

How does the diffusion approach differ from classical Gaussian noise used in image generation?

Classical diffusion for images uses continuous Gaussian noise that can be added to continuous data. For discrete text tokens, this is problematic, so masking is used instead—replacing tokens with placeholder symbols.

When did diffusion language models become competitive with autoregressive models?

The turning point came in 2024 when diffusion models achieved comparable quality to autoregressive models. By 2026, diffusion LLMs are a reality with releases from leading labs including Inception Labs, Google, and NVIDIA.

Questions and answers are written by AI about the topic, not taken from the source; they aren’t the newsroom’s opinion.
Related from the stream
Mentions
  • Volodymyr Kuleshov
  • Marianne Arriola
  • Yair Schiff
  • Guanghan Wang
  • Cornell University
  • Inception
  • Mercury 2
  • Gemma Diffusion
  • Nemotron Diffusion
  • BERT