Diffusion Models from Scratch
Build a noise-prediction problem, derive its posterior on a known mixture, and run real reverse steps. The examples separate denoiser training, sampling, and conditional guidance.
01Corrupt samples with a known process
A learns from deliberately corrupted data. The corruption process is known, so training can create a noisy input and a target without asking a human to label each noise pattern. Generation later uses a learned predictor repeatedly, starting from noise and moving toward a sample.
For a variance-preserving Gaussian process, one forward step scales the previous sample and adds independent Gaussian noise. Repeated steps have a closed-form marginal, so training can jump directly to a selected noise level.[1]
x_t = √ᾱ_t × x₀ + √(1 − ᾱ_t) × ε
x₀ is a clean sample; ε is a standard normal draw. β_t is the per-step noise variance. ᾱ_t is the remaining cumulative signal power, not the per-step coefficient.
Conditioned on a fixed x₀, x_t has mean √ᾱ_t x₀ and variance 1−ᾱ_t per coordinate. That noise variance is not the variance of the entire data distribution. If clean data has its own spread, its contribution remains after scaling.
The examples use an explicit one-dimensional data distribution: an equal mixture of Gaussian components centered at −2 and +2, each with variance 0.16. At noise level ᾱ, each noisy component has mean ±2√ᾱ and variance 0.16ᾱ+1−ᾱ. These are computed toy parameters, chosen so the two clean modes are visibly distinct.
A noise schedule controls which intermediate distributions the model sees. Nichol and Dhariwal introduce an offset cosine schedule. The widget uses a simpler cosine-squared signal-power curve to span the full display range; it is explicitly not their exact offset-and-clipped discrete schedule.[2]
02A denoising mean can lie between plausible samples
Given a noisy observation, several clean values may explain it. The averages those possibilities according to their conditional probability. For squared-error prediction, that mean minimizes the expected loss. It is not generally the most probable clean value.
For a Gaussian clean component with mean μ and variance τ², write s=√ᾱ and σ²=1−ᾱ. Its noisy variance is s²τ²+σ². Completing the square in the product of its Gaussian prior and likelihood gives a clean posterior mean:
For the mixture, weight each component's posterior mean by its Bayes responsibility: its noisy likelihood times its prior mixture weight, normalized across components.
At x_t=0 with equal component weights, both modes are equally plausible. Their posterior means are opposite, so the mixture's conditional mean is zero. A predicted zero is therefore a correct squared-error answer even when zero is an unlikely clean sample. The widget displays the posterior density and this mean separately.
The corresponding expected noise follows from the forward equation: E[ε|x_t]=(x_t−sE[x₀|x_t])/σ for σ>0. It also relates to the score of the noisy density: E[ε|x_t]=−σ∇log p_t(x_t). Song and Ermon develop generation with scores of perturbed distributions; here the score is evaluated directly from the known mixture.[3]
That connection explains why a local denoising predictor can guide generation even though its clean estimate can average incompatible possibilities. Generation follows a sequence of updates through different noise levels. It does not finish by taking one heavily corrupted sample's posterior mean.
03Train a predictor using known noise targets
A common noise-prediction objective samples clean data, a timestep, and Gaussian noise, constructs x_t, and trains ε̂θ(x_t,t) toward the sampled noise. The target is available because the training pipeline generated it. Ho et al.'s simplified objective uses a squared-error noise-prediction loss.[1]
The expectation averages over data, noise, and the selected training timesteps. The network receives the noisy sample and a representation of the timestep or noise level; it does not receive the clean target sample.
A predictor cannot in general recover the individual sampled noise exactly from x_t alone. Clean data is unknown, and different clean/noise combinations can yield the same observation. At the population optimum under squared error, it predicts conditional expected noise. A neural model approximates that function from finite training data.
The next widget actually trains a much smaller model: ε̂=w x_t+b. It takes full-batch gradient steps on 64 fixed generated pairs at one selected noise level. A line cannot express the mixture's nonlinear conditional expectation. Its loss can fall without reaching zero or reproducing the analytic denoiser used later.
In a multi-level neural model, the same input value can require different predictions at different noise levels. The time or noise-level conditioning disambiguates them. A training loss averaged across levels can also hide a poor region of the schedule; evaluate prediction behavior by noise level as well as in aggregate.
Changing the predicted quantity or timestep weighting changes the optimization problem even when algebraic conversions exist between predictions. Small denoising loss by itself does not establish sample quality, diversity, or performance under a different sampler.
04Generate by stepping toward less noise
Sampling starts with a draw from a simple distribution at a very noisy level. A reverse update uses the denoiser to produce the next, less noisy state. Some samplers introduce fresh random noise between steps. The deterministic shown here uses η=0 and adds none.[4]
x_next = √ᾱ_next × x̂₀ + √(1 − ᾱ_next) × ε̂
The next level has larger signal power. This is the deterministic DDIM formula for selected noise levels. The clean estimate and noise direction both come from the current prediction.
"Deterministic" describes the trajectory after initialization. Different initial noise draws can produce different outputs. A stochastic reverse sampler also draws noise at intermediate steps. The original DDPM algorithm also handles its final step specially, without adding a new random draw there.[1]
This demo starts at ᾱ=10⁻⁴, where the data contribution is small but nonzero. Standard normal initialization approximates that noisy marginal. A finite-step DDIM path with the exact conditional-mean predictor still has discretization and initialization differences; the demo does not claim an exact draw from the clean mixture. Both the predictor and sampler have to be considered when judging results.
In an image model, x_t is an array of pixels or latent coordinates, and a neural network predicts an array with the same shape. The scalar algebra applies coordinatewise, but the predictor uses information across coordinates to estimate a joint structure. The one-dimensional mixture isolates the sampling calculation without claiming image-generation capability.
05Conditional predictions and guidance scale
A conditional model receives an additional signal, such as a class or a text representation. combines its conditional prediction with an unconditional prediction. During the original method's training, the condition is omitted on a fraction of examples so the model can supply both.[5]
g=0 uses the unconditional predictor; g=1 uses the conditional predictor; g>1 extrapolates. Ho and Salimans use a w convention equivalent to g=1+w.
The term "classifier-free" means no separately trained classifier is required for this guidance mechanism. It does not mean no conditioning information exists. In the widget, the unconditional distribution is the equal two-mode mixture; the condition selects the +2 component. Both predictors are analytic, so the effect of combining them can be calculated directly.
At each noise level, the guided score is the gradient of g log p_cond+(1−g)log p_uncond. Where normalizable, that corresponds to a density proportional to p_cond^g p_uncond^(1−g) at that level. This local identity does not by itself prove that a finite guided sampler produces that density at its endpoint.[5]
Larger guidance can alter diversity and move predictions into regions that differ from ordinary conditional sampling. It needs evaluation for the intended model and task.
06Noise, clean sample, and velocity targets
A denoiser can predict ε, x₀, or a different combination of them. For a variance-preserving process, define s=√ᾱ and σ=√(1−ᾱ). The velocity target is v=sε−σx₀. It is not the physical velocity of an object.[6]
x₀ = s × x_t − σ × v
ε = σ × x_t + s × v
s²+σ²=1, so these equations form an invertible rotation of the pair (x₀, ε). The epsilon-to-clean conversion instead divides by s.
For exact targets, all these descriptions agree. For imperfect predictions, their errors transform differently with the noise level. An epsilon-prediction error is multiplied by −σ/s during clean reconstruction, which can amplify error in the reconstructed clean estimate. Salimans and Ho discuss parameterization and training-loss weighting together; "these targets are algebraically related" does not make all training objectives equivalent.[6]
Record the target convention with the model weights and schedule. A sampler expecting epsilon predictions cannot simply interpret a velocity output as epsilon. Conversion requires the current signal and noise coefficients. Matching tensor shapes will not detect this semantic mistake.
07A deterministic DDIM step in two languages
The complete programs implement the same deterministic step. Their executable checks supply an exact noise target to check the update algebra, including a final clean endpoint and the velocity reconstruction identity. The code does not contain a trained denoising model.
#include <cassert>
#include <cmath>
struct Step { double clean; double next; };
Step ddimStep(double noisy, double predictedNoise,
double alphaBar, double nextAlphaBar) {
assert(alphaBar > 0 && alphaBar < 1);
assert(nextAlphaBar >= alphaBar && nextAlphaBar <= 1);
// eta=0: no fresh random draw enters this reverse update.
// Invert the known corruption using this level's prediction.
const double clean = (noisy-std::sqrt(1-alphaBar)*predictedNoise)
/ std::sqrt(alphaBar);
// Reuse that prediction as the deterministic direction at the next level.
const double next = std::sqrt(nextAlphaBar)*clean
+ std::sqrt(1-nextAlphaBar)*predictedNoise;
return {clean,next};
}
int main() {
const double clean = 2.0, noise = -0.8;
const double alphaBar = 0.4, nextAlphaBar = 0.7;
const double noisy = std::sqrt(alphaBar)*clean+std::sqrt(1-alphaBar)*noise;
const Step result = ddimStep(noisy,noise,alphaBar,nextAlphaBar);
assert(std::abs(result.clean-clean) < 1e-12);
const double expected = std::sqrt(nextAlphaBar)*clean
+ std::sqrt(1-nextAlphaBar)*noise;
assert(std::abs(result.next-expected) < 1e-12);
const Step final = ddimStep(noisy,noise,alphaBar,1.0);
assert(std::abs(final.next-clean) < 1e-12); // Final noise coefficient is zero.
const double velocity = std::sqrt(alphaBar)*noise-std::sqrt(1-alphaBar)*clean;
const double fromVelocity = std::sqrt(alphaBar)*noisy-std::sqrt(1-alphaBar)*velocity;
assert(std::abs(fromVelocity-clean) < 1e-12);
}
struct Step { clean: f64, next: f64 }
fn ddim_step(noisy: f64, predicted_noise: f64,
alpha_bar: f64, next_alpha_bar: f64) -> Step {
assert!(alpha_bar > 0.0 && alpha_bar < 1.0);
assert!(next_alpha_bar >= alpha_bar && next_alpha_bar <= 1.0);
// eta=0: no fresh random draw enters this reverse update.
// Invert the known corruption using this level's prediction.
let clean = (noisy-(1.0-alpha_bar).sqrt()*predicted_noise)/alpha_bar.sqrt();
// Reuse that prediction as the deterministic direction at the next level.
let next = next_alpha_bar.sqrt()*clean
+ (1.0-next_alpha_bar).sqrt()*predicted_noise;
Step { clean,next }
}
fn main() {
let clean = 2.0; let noise = -0.8;
let alpha_bar: f64 = 0.4; let next_alpha_bar: f64 = 0.7;
let noisy = alpha_bar.sqrt()*clean+(1.0-alpha_bar).sqrt()*noise;
let result = ddim_step(noisy,noise,alpha_bar,next_alpha_bar);
assert!((result.clean-clean).abs() < 1e-12);
let expected = next_alpha_bar.sqrt()*clean+(1.0-next_alpha_bar).sqrt()*noise;
assert!((result.next-expected).abs() < 1e-12);
let final_step = ddim_step(noisy,noise,alpha_bar,1.0);
assert!((final_step.next-clean).abs() < 1e-12); // Final noise coefficient is zero.
let velocity = alpha_bar.sqrt()*noise-(1.0-alpha_bar).sqrt()*clean;
let from_velocity = alpha_bar.sqrt()*noisy-(1.0-alpha_bar).sqrt()*velocity;
assert!((from_velocity-clean).abs() < 1e-12);
}
A full sampler needs a model call at each selected level, consistent time conditioning, the model's expected prediction convention, and the same data or latent scaling used in training. A stochastic variant also needs the specified variance and independent random draws. Reusing one unrelated variance formula with another sampler is a model change.
The program accepts noisy current levels with 0<ᾱ<1. Positivity prevents division by zero in clean reconstruction; the upper limit excludes an already clean current sample. The widget uses a small positive starting floor and handles the final next level as one. An implementation that evaluates a singular endpoint must define the limiting behavior rather than divide by zero.
08Separate predictor error from sampler error
Test a sampler first with exact supplied targets and a known distribution. Then test the learned predictor at individual noise levels. Finally evaluate generated outputs. A correct reverse formula cannot compensate for an incorrect time index, wrong normalization, or a prediction expressed in the wrong convention.
The posterior-density widget computes Bayes conditioning. The trained line fits a finite fixed dataset. The sampler evaluates reverse steps with a known analytic predictor. Testing these components separately can reveal errors hidden by the final animation.
A sample can look plausible while failing to preserve a particular original, satisfy a requested condition, or represent the full data distribution. Evaluate the property actually needed, rather than treating the denoising formula as a guarantee.
09What's next
AI Inference and Quantization covers running a trained model: floating-point stability, integer representations, memory traffic, and sampling from output probabilities. Each changes a different part of the deployed system.
10Sources
The cited equations define the algorithms. Widget readouts report the displayed toy arrays and settings; they are not hardware benchmarks.
- Jonathan Ho, Ajay Jain, Pieter Abbeel, 2020. Denoising Diffusion Probabilistic Models. Gaussian forward marginals, noise-prediction training, and reverse sampling.
- Alexander Quinn Nichol and Prafulla Dhariwal, 2021. Improved Denoising Diffusion Probabilistic Models. Noise scheduling, including the offset cosine schedule.
- Yang Song and Stefano Ermon, 2019. Generative Modeling by Estimating Gradients of the Data Distribution. Scores of noise-perturbed distributions and iterative generation.
- Jiaming Song, Chenlin Meng, Stefano Ermon, 2021. Denoising Diffusion Implicit Models. Deterministic reverse steps and sampling on selected noise levels.
- Jonathan Ho and Tim Salimans, 2022. Classifier-Free Diffusion Guidance. Combining conditional and unconditional denoising predictions.
- Tim Salimans and Jonathan Ho, 2022. Progressive Distillation for Fast Sampling of Diffusion Models. The velocity parameterization and reconstruction identities.