AI EncyclopediaArchitectures & models
Diffusion Model
A generative model that learns to create data - typically images - by learning to reverse a process that gradually adds noise to training data, starting from pure noise and denoising it step by step into a coherent output.
In plain English
Picture teaching a model to un-blur a photo that has been slowly turned into static, over and over, until it becomes an expert at removing noise. Run that skill starting from pure random static, guided by a text description, and the 'un-blurring' produces a brand-new image matching the description instead of restoring an original.
Technical explanation
Diffusion models perform two processes: a forward process that progressively destroys training data by adding noise, and a learned reverse process that recovers data by removing it step by step. Text-to-image variants like Stable Diffusion train the noising and denoising chain jointly with the text description, so the model learns to associate specific words with the visual structure it is learning to reconstruct. Stable Diffusion, released by Stability AI in 2022, operates in a compressed latent space rather than raw pixel space, applying noise to and recovering from that compressed representation for efficiency.
Why it matters
Diffusion is the dominant approach behind current text-to-image and increasingly text-to-video generation, and unlike some earlier proprietary approaches, Stability AI open-sourced both Stable Diffusion's weights and source code, which materially broadened who could build on the technique.
Real-world example
DALL-E (OpenAI), Stable Diffusion (Stability AI) and Midjourney are the three most widely used diffusion-based image generators, each trained on the same underlying noise-and-denoise principle with different training data, tuning and interfaces.
Common misunderstanding
That diffusion models 'collage' or retrieve pieces of real training images to assemble an output. They generate pixels from learned statistical patterns via the denoising process - there is no lookup or copy-paste step, even though outputs can sometimes closely resemble training data in specific, studied cases.
Something wrong here?
Every entry is hand-researched and hand-written by Nakoda. If a fact is stale, a source has changed or a definition needs sharpening, tell us and we will check it.

