Models · reviewed · reviewed Oct 5, 2026 · 5 min
How does diffusion generate an image?
Diffusion training teaches a model to predict structure or noise from corrupted examples. Sampling starts from noise and repeatedly uses a fixed denoiser to move toward a sample. The clean original is available when constructing training examples; it is not supplied when generating a new image.
Denoising is learned from examples
Start with a clean training image and deliberately add noise. Because the corruption was constructed, the training procedure knows both the clean example and the noise that was added. It can ask a network to predict the noise from the corrupted image and its noise level, compare the prediction with the known target, and update the network’s parameters.
Repeated across examples and noise levels, this teaches a denoiser to recognize useful structure under corruption. Predicting noise is one formulation; diffusion models can also parameterize related quantities, such as a clean-image estimate. A denoiser is not simply a generic blur filter.
The following example uses a six-by-six X or ring and a fixed noise draw. At signal fraction 1, the input is clean. Reducing the signal fraction mixes in more noise. The procedure still knows the original because it constructed this particular training example.
Explore the mechanism
Add noise to a known training example
- Signal fraction α
- 1.00
- Noise fraction 1 − α
- 0.00
The clean example and the added noise are available when constructing the training target. Sampling starts under a different condition.
Two authored 6×6 patterns and a fixed Gaussian draw (seed 7) illustrate corruption. Pixel values are computed; display colors alone are clipped to 0–1. This constructs a noisy example but does not train a denoiser or measure image quality.
The color display clips values to its visible range; the underlying calculation does not. These small patterns illustrate the corruption equation, not photographs or a trained model. Moving the slider constructs inputs at different noise levels; it does not learn parameters.
Sampling begins without an original
Now remove the clean image from the input. Generation starts from a random noise sample. At each scheduled noise level, the fixed denoiser estimates useful structure or residual noise; a sampling rule uses that estimate to construct the next, less noisy state.
The current sample changes. The learned parameters remain fixed. Several denoising calls therefore represent repeated inference, not several training updates.
To make that difference inspectable, the reverse experiment uses an analytical denoiser instead of a neural network. Its entire prior—the set of clean structures it knows—is two equally likely authored patterns, X and ring. It compares the current noisy state with both possibilities and computes a weighted clean estimate. No original image is passed to this sampler.
Choose a seed and take seven denoising steps. Watch the noisy sample and the estimated clean structure separately. Resetting the same seed repeats the same deterministic trajectory.
Explore the mechanism
Generate from noise with a fixed denoiser
- Denoising steps
- 0 / 7
- Signal fraction α
- 0.01
Estimate clean structure → estimate residual noise → move to the next signal fraction.
Inspect the two-pattern prior
Posterior weights within this toy prior: X 36.3%; ring 63.7%. These weights compare two authored patterns, not the factual correctness or quality of a real image.
The initial sample comes from seed 1. No target pattern is selected for reverse sampling. Its denoiser can only favor structures already present in its fixed two-pattern prior.
This local sampler starts from seeded Gaussian noise and receives no original image. An analytical denoiser knows only an equally likely X/ring prior; it is not a trained neural network. Seven deterministic DDIM-style updates change the sample, never the denoiser. Values remain unclipped in calculations; colors alone are bounded.
The generated structure comes from the initial noise interacting with that fixed prior. It cannot invent a face, landscape, or a third pattern: those possibilities do not exist in this tiny denoiser. This is the central limitation of the demo, stated before interpreting its pictures as generation.
The update follows a deterministic DDIM-style rule. Other diffusion samplers may inject additional randomness or use different schedules. The existence of several valid sampling rules is why “number of denoising steps” alone does not specify how an image was generated.
Why real systems can work in a latent space
A full-resolution image has many numerical values. Latent diffusion moves the repeated denoising work into a more compact learned representation.
An encoder maps training images to latents, numerical representations that preserve useful image structure. A denoising model learns and samples in that space. A decoder maps the resulting latent sample back to visible pixels.
The sequence is therefore sample latent noise → denoise latent state repeatedly → decode an image. Decoding does not mean fetching the original training picture. It transforms a generated representation into pixels, with limitations introduced by both the latent model and the encoder/decoder.
Our widget works directly on 36 pixel values. It has no encoder, decoder, or learned compression. Keeping those absent makes the sampling mechanism visible; it does not make the widget a miniature implementation of a production text-to-image model.
A text prompt conditions the process
For text-to-image generation, the system can encode a prompt and supply that representation to the denoiser. In the latent diffusion architecture, cross-attention connects conditioning information with the denoising network. The prompt influences predictions throughout sampling.
Conditioning is not a guarantee that the final image satisfies every instruction. A requested count, relationship, or readable label can be wrong even when the picture is attractive. Assess the generated image itself; the presence of words in the prompt does not establish their successful realization.
Likewise, a model that understands an input image and produces text is solving a different problem from a model generating pixels. The multimodal LLM article follows image evidence into language; this article follows a noisy generative state into an image.
What actually changes when you generate again?
In this toy, changing the seed changes the initial noise, while the denoiser and schedule stay fixed. Changing the forward-demo pattern has no effect on reverse sampling because the demos share no target image or sampling state.
In a real system, the checkpoint, conditioning, seed, sampling algorithm, schedule, resolution, and decoder can all affect the result. More steps add denoiser work but do not establish a universal quality improvement. Repeating a seed is a useful comparison only when the other relevant conditions are controlled.
Inspect the corruption and deterministic sampling equations
Let α be the cumulative signal fraction, x₀ a clean sample, and ε a Gaussian noise vector. The corrupted input is x = √α x₀ + √(1−α) ε.
The toy denoiser assigns equal prior probability to X and ring. For each pattern, it computes the Gaussian likelihood of the current x under that corruption rule, normalizes the two likelihoods, and returns their weighted mean as cleanEstimate.
The residual estimate is ε̂ = (x − √α cleanEstimate) / √(1−α). The next sample is √αNext cleanEstimate + √(1−αNext) ε̂. No original-image argument appears in this reverse update.
The seven transitions use signal fractions 0.01, 0.05, 0.15, 0.35, 0.65, 0.85, 0.97, 1. At the last transition the residual coefficient becomes zero. Sampling then stops; it does not evaluate a division by zero at signal fraction 1. These large teaching steps and the two-pattern prior are not a real image-quality benchmark.
Sources
Sources and further reading
- 01Denoising Diffusion Probabilistic ModelsHo, Jain, and Abbeel · research · published Jun 19, 2020 · source checked Oct 5, 2026
Forward Gaussian corruption, noise-prediction training, and iterative sampling separate known training targets from generation without an original.
- 02Denoising Diffusion Implicit ModelsSong, Meng, and Ermon · research · published Oct 6, 2020 · source checked Oct 5, 2026
Deterministic sampling with a fixed denoiser supplies the update rule used by the disclosed analytical two-pattern teaching prior.
- 03High-Resolution Image Synthesis with Latent Diffusion ModelsRombach et al. · research · published Dec 20, 2021 · source checked Oct 5, 2026
Learned image encoding, denoising in latent space, decoding, and cross-attention conditioning explain the text-to-image path.
