AI tutorialsMighty Professional
Build a Language Model · Making it useful

Fine-tuning and Alignment from Scratch

A pretrained model continues text. Turning it into something that answers, follows instructions and prefers good answers to bad ones is a second training problem with its own data, losses and failure modes. This page runs each stage on a small model in the browser: supervised fine-tuning with and without loss masking, LoRA against full fine-tuning, a reward model fit to comparisons, KL-regularized policy optimization with a reward model that is wrong about one response, DPO, and group-relative advantages.

Time~60 minLevelIntermediatePrereqsTrain a Language Model for the model being tuned; Reinforcement Learning for policy gradients; Probability for KL divergence and the sigmoid.StackC++20 & Rust · browser demos
◂ Build a Language ModelPhase 4 · Making it usefulNext · Inference and quantization ▸

01From predictor to assistant

Pretraining produces a model of text as it occurs. The assistant behaviour people use is added afterwards, in stages that each change what the model optimizes. Llama 3's instruction-tuned models, to take one documented example, were produced from the pretrained ones with followed by reinforcement learning with human feedback, on public instruction data plus over ten million human-annotated examples.[1] The stages share the pretrained weights and the next-token machinery; what differs is the data and the loss.

The costs are not symmetric with pretraining. Stanford's Alpaca fine-tuned a 7-billion-parameter model on 52,000 instruction-following examples generated from a stronger model for under $500 of API calls, with three epochs at a learning rate of 2e-5 and batch 128, on one machine with four 80 GB A100s.[2] The preference stages need comparisons rather than demonstrations: tens of thousands of human judgments between two model outputs for summarization, and sets of chosen-versus-rejected dialogue turns for helpfulness and harmlessness.[3][4] Each stage below is run on the one-block character model of the previous page, pretrained on sentences about a farm and then taught something else.

02Supervised fine-tuning

SFT is pretraining's loss on chosen data. Each example is a prompt and a completion; they are concatenated, tokenized, shifted by one, and scored with the same cross-entropy as before. The one change that matters is which positions count. By default TRL computes the loss on the completion tokens only, and for conversations an option restricts it to the assistant's turns; both work by setting the target at every other position to an ignore index that the loss and its gradient skip, −100 in PyTorch's cross-entropy and −1 in nanoGPT's.[5][6] Padding is masked the same way, and packing several short examples into one sequence keeps the batch full.[5]

LSFT = −(1/|C|) Σt ∈ C log pθ(yt+1 | x, y≤t)

C is the set of positions that count. Prompt tokens still pass through the model as context; they just produce no gradient.

Fine-tune on question–answer pairs with and without prompt masking

The base model, pretrained on 7,200 characters of farm sentences, is fine-tuned on examples of the form "Q: the cat? A: the mat." where each animal has a fixed answer. Each Step is ten iterations of batch 8. Both runs are scored only on completion characters, which start at 1.29 nats because the farm model already predicts "the " well. With the prompt masked the completion loss is 0.21 after 20 iterations and 0.04 after 50; without masking it is 0.45 and 0.11, and it stays higher at every logged point, because more than half of the target positions in every unmasked batch are question characters. The per-character panel shows the masked model's loss on one example: up to 11 nats on the question characters it was never trained to predict, and close to zero on every answer character.

03Low-rank adapters

Full fine-tuning updates every weight and so needs optimizer state for every weight: for a 7-billion-parameter model in mixed precision, about 112 GB on the 16-bytes-per-parameter count, and a full copy of the model per task. freezes the pretrained matrix and trains a low-rank correction instead.[7] The reference implementation initializes A with the same uniform draw a linear layer gets and B with zeros, so the adapted model begins exactly as the pretrained one, and scales the product by α/r.[7] At inference the product is added into the weight once and the adapter disappears: no extra latency.[8]

W = W0 + (α/r)·A·B ∂L/∂A = (α/r)·∂L/∂W·Bᵀ ∂L/∂B = (α/r)·Aᵀ·∂L/∂W

The adapter has r·(in + out) parameters against in·out. Its gradients are the full matrix gradient projected through the other factor, so a backward pass that already computes ∂L/∂W needs two small matrix products more. PEFT's rsLoRA option scales by α/√r instead.[9]

Teach the farm model about spacecraft: full fine-tuning against LoRA

Both runs start from the same pretrained weights and see the same 300 iterations of space-domain batches; α is set to 2r so the scaling is 2. At rate 0.01, full fine-tuning takes the space loss from 3.7 to about 0.5 while the farm loss climbs from 1.1 past 4.4, worse than the uniform 3.5: the old domain is gone. LoRA at rank 4 trains 1,024 adapter parameters on the four attention and MLP matrices of the 4,160-parameter model, reaches about 1.35 on the new domain and holds the old one near 2.5. Rank trades one against the other: rank 1 ends at 2.0 new and 2.1 old, rank 16 at 0.86 new and 3.0 old. With a 4,160-parameter base there is little capacity to spare, which is why the gap between full and LoRA here is far larger than the GLUE and E2E results in the LoRA README, where adapting 0.8M of 125M or 0.35M of 355M parameters matches or beats full fine-tuning.[8]

04Preferences and the reward model

Demonstrations show the model what to say; comparisons tell it which of two things it said was better. A preference example is a prompt with a chosen and a rejected completion, which is the format TRL's trainers consume and the format of the HH-RLHF and summarization datasets.[10][4][3] A reward model turns comparisons into a score. Under the Bradley–Terry model the probability that the chosen completion is preferred is the sigmoid of the score difference, and the reward model is trained to maximize the log-likelihood of the observed labels.[11]

P(y⁺ ≻ y⁻ | x) = σ(rθ(x, y⁺) − rθ(x, y⁻)) LRM = −log σ(rθ(x, y⁺) − rθ(x, y⁻))

Only differences of rewards enter the likelihood, so adding any constant to every score leaves the loss unchanged: the model is underdetermined by a constant, and TRL offers a penalty on the mean score to pin it.[11]

Fit Bradley–Terry scores to sampled comparisons

Eight responses have hidden true scores from +1.8 to −1.5. Each comparison picks two at random and labels the winner through the Bradley–Terry probability of the true scores, so even perfect scores agree with only 77% of fresh labels; a scorer cannot beat the noise in the labels. With 100 clean comparisons the fitted scores correlate about 0.90 with the truth and agree with 75% of fresh labels; 400 comparisons reach 0.99 and the 77% ceiling; 20 comparisons give 0.5 and 65%. Flipping 20% of labels mostly compresses the fitted range and costs little agreement at 400 comparisons but ruins the fit at 20. Scores start at 2 to show the constant: the pairwise gradient never moves the mean, so without the centering penalty the fitted mean stays at exactly 2, and with it (coefficient 0.1 on the squared mean here) the mean is pulled to 0.01 within the 400 steps.

05RLHF: reward minus a KL penalty

With a reward model in hand, adjusts the policy to produce completions the reward model scores highly, while charging for every nat it drifts from the model it started as. The objective is the expected reward minus β times the KL divergence from the reference policy. The implementation OpenAI released with its 2019 paper on fine-tuning language models from human preferences adds the penalty per token: each generated token contributes −kl_coef·(log π − log π_ref) to its reward, the reward model's score is added at the final token, and the sum is what PPO maximizes.[12] The coefficient was 0.15 with an adaptive target of 6 nats for stylistic tasks and 0.01 to 0.03 for summarization, with a note in the launch script that 0.01 was too low.[13]

maxπ Ey∼π[r(y)] − β·KL(π ‖ πref) ⇒ π*(y) = πref(y)·exp(r(y)/β) / Z

The closed form follows from setting the derivative of Σ π(y)·r(y) − β Σ π(y)·log(π(y)/πref(y)) to a constant under the constraint Σ π = 1. Over a finite set of candidate completions it can be computed exactly, which is what the widget does; for a language model the set is every possible string and the optimum has to be approached by gradient steps.

Optimize a policy over six responses against a reward model that is wrong about one

Six candidate responses to one prompt; the reference policy gives them probabilities 0.35 down to 0.05. The true rewards are 0, 0.5, 1, 1.5, 2 and −1: response F is bad. The reward model agrees on A to E but scores F at +4, the kind of error a model makes on a response it rarely saw in the comparison data, since F is the rarest under the reference. Each Step is ten exact gradient steps; there is no sampling noise. At β = 0.5 the policy puts 94% of its mass on F within 100 steps: the optimized reward climbs from 0.8 to 3.9, the true reward falls from 0.56 to −0.86, and the KL reaches 2.7 nats. At β = 2 the policy stays within 0.2 nats of the reference, F gets 21%, and the true reward is 0.49, still below the reference's 0.56. Switching to the true reward shows what the penalty buys when the reward is right: β = 0.5 reaches a true reward of 1.54 at 0.9 nats. Over six responses the gradient steps land on the closed-form optimum every time, which is the check that the update is right; a language model has no such check.

Two more pieces of the released implementation are worth knowing. Rewards are whitened, scaled to unit variance over the batch, before PPO sees them, and a KL controller multiplies the coefficient by 1 + clip(KL/target − 1, −0.2, 0.2)·n/horizon after each batch, so a run that drifts too fast is charged more.[12] Later trainers moved the penalty around: RLOO folds it into the scalar reward with no gradient through it and uses the mean reward of the other completions as a baseline, and GRPO adds it to the loss as a differentiable term.[14][15]

06DPO: the same optimum without the reward model

Solving the closed form for the reward gives r(y) = β·log(π*(y)/πref(y)) + β·log Z. Substituting that into the Bradley–Terry probability makes the log Z terms cancel, because both completions share the prompt, and leaves a likelihood that mentions only the policy and the reference. Maximizing it on the comparison data is : a classification loss on the policy, no reward model, no sampling.[16][17]

LDPO = −log σ( β·( log πθ(y⁺)/πref(y⁺) − log πθ(y⁻)/πref(y⁻) ) ) r̂(y) = β·log(πθ(y)/πref(y))

The gradient on one pair scales both log-ratios by β·σ(−β·margin): large while the policy still gets the pair wrong, vanishing once the margin is wide. Over a finite set the softmax normalizer cancels in the margin, so each pair moves the chosen logit up and the rejected one down by the same amount. Over sequences each token has its own normalizer, so the update also moves the other logits at every position of both completions. The reference code adds label smoothing for noisy labels, which mixes in the flipped pair's loss with weight ε, and the README suggests β between 0.1 and 0.5 and running SFT first so the pairs are in-distribution.[16][18]

DPO on sampled comparisons reaches the RLHF optimum

The same six responses and reference as above, with comparisons labelled from the true rewards through Bradley–Terry; no reward model is fit. Each Step is 25 gradient steps on the pairs. The policy's logits start at log πref and move only where the pairs push them, and the implicit rewards β·log(π/πref) line up with the centered true rewards as the pairs accumulate. With 1,000 pairs the ordering is right and the largest probability gap to the closed-form optimum is 0.04. With 200 pairs, about 13 per pair of responses, the sampled labels put D above E and the gap is 0.22; with 50 the ordering is wrong in several places. Flipping 20% of labels compresses the implicit rewards the way it compressed the reward model's scores, and the policy stays closer to the reference. β plays the same role as in the KL-regularized objective: at 0.2 the 1,000-pair policy puts 84% on E, the best response under the true reward, against 45% at 0.5.

07Online RL with verifiable rewards

When the reward can be computed, by checking a final answer or running tests, the reward model goes away and the loop becomes: sample G completions per prompt from the current policy, score each, and push up the ones that scored above their peers. GRPO normalizes each completion's reward by the mean and standard deviation of its group, (r − mean)/std, and uses that as the advantage in a PPO-style clipped objective, with an optional differentiable KL term estimated per token as ratio − log ratio − 1; TRL sets that term's coefficient to zero by default and notes the recent runs that dropped it.[15] RLOO uses the mean of the other G − 1 rewards as the baseline instead and folds the KL penalty into the reward.[14]

Âi = (ri − mean(r)) / std(r) LGRPO = −(1/Σ|oi|) Σi Σt [ ρi,t·Âi − β·DKL ]

Dividing by the standard deviation makes the signal independent of the reward scale, and makes it undefined when every completion in the group scores the same; implementations set the advantage to zero then, and TRL logs the fraction of such groups. Dividing by std also changes how hard and easy prompts are weighted, which TRL's documentation flags with a reference and a switch to turn it off.[15]

Group-relative advantages and the groups that carry no signal

Rewards are 1 or 0, as a correctness checker gives. In a group of 8 with two successes the advantages are +1.73 for each success and −0.58 for each failure; with one success they are +2.65 and −0.38: the rarer outcome gets the larger push. A group where all 8 agree has zero standard deviation and contributes nothing, and the right panel shows how often that happens: at p = 0.9 it is 43% of groups for G = 8 and 66% for G = 4, which is why hard prompts that the model nearly always fails and easy ones it nearly always solves both drop out of training until G is raised or the prompts are filtered.

08Two checked programs

The first program puts a rank-2 adapter on an 8 × 6 linear layer, derives the adapter gradients from the full-weight gradient, checks all 28 of them against centered finite differences, and asserts that the merged weight gives the same output as the unmerged two-path forward. The second builds the preference pipeline over six responses: 2,000 comparisons sampled through Bradley–Terry from hidden rewards, a reward model fit by gradient descent with the centering penalty, the closed-form KL-regularized optimum, exact policy gradient reaching it, and DPO reaching the same policy from the raw pairs. Both run in well under a second.

LoRA on one linear layer: gradients checked, merge verified
// LoRA on one linear layer: the adapter's gradients, checked against finite differences, and the merge identity.
// Build: g++ -std=c++20 -O2 -Wall -Wextra -Wpedantic -Werror lora.cpp && ./a.out
#include <cassert>
#include <cmath>
#include <cstdio>
#include <vector>

struct Rng {  // xorshift32 with a fixed seed
    unsigned state = 11;
    double next() { state ^= state << 13; state ^= state >> 17; state ^= state << 5; return state / 4294967296.0; }
};

constexpr int kIn = 8, kOut = 6, kRank = 2, kRows = 4;
constexpr double kAlpha = 4.0, kScaling = kAlpha / kRank;  // loralib: scaling = lora_alpha / r

// y[rows][out] = x[rows][in] · W[in][out]
std::vector<double> matmul(const std::vector<double>& left, const std::vector<double>& right, int rows, int inner, int columns) {
    std::vector<double> out(std::size_t(rows) * columns, 0.0);
    for (int row = 0; row < rows; ++row) for (int i = 0; i < inner; ++i) for (int c = 0; c < columns; ++c) out[row * columns + c] += left[row * inner + i] * right[i * columns + c];
    return out;
}
// The adapted weight the model actually multiplies by: W₀ + scaling · A · B.
std::vector<double> mergedWeight(const std::vector<double>& base, const std::vector<double>& A, const std::vector<double>& B) {
    std::vector<double> product = matmul(A, B, kIn, kRank, kOut);
    std::vector<double> merged(base);
    for (std::size_t i = 0; i < merged.size(); ++i) merged[i] += kScaling * product[i];
    return merged;
}
// Half the squared error between x·W and the targets, summed over the batch.
double loss(const std::vector<double>& x, const std::vector<double>& weight, const std::vector<double>& targets) {
    std::vector<double> y = matmul(x, weight, kRows, kIn, kOut);
    double total = 0;
    for (std::size_t i = 0; i < y.size(); ++i) total += 0.5 * (y[i] - targets[i]) * (y[i] - targets[i]);
    return total;
}

int main() {
    Rng rng;
    auto uniform = [&](std::size_t size, double bound) { std::vector<double> values(size); for (double& value : values) value = (rng.next() * 2 - 1) * bound; return values; };
    std::vector<double> x = uniform(kRows * kIn, 1.0), targets = uniform(kRows * kOut, 1.0), base = uniform(kIn * kOut, 0.5);
    // loralib initialization: A ~ kaiming_uniform(a=√5), which is uniform in ±1/√in; B = 0, so the adapted layer starts equal to the base layer.
    std::vector<double> A = uniform(kIn * kRank, 1.0 / std::sqrt(double(kIn))), B(kRank * kOut, 0.0);
    assert(mergedWeight(base, A, B) == base);
    // Give B some content so the gradient with respect to A is not identically zero.
    for (double& value : B) value = (rng.next() * 2 - 1) * 0.3;

    // Forward through the merged weight, then the gradient of the loss with respect to that weight: xᵀ · (y − t).
    std::vector<double> merged = mergedWeight(base, A, B), y = matmul(x, merged, kRows, kIn, kOut);
    std::vector<double> weightGrad(kIn * kOut, 0.0);
    for (int row = 0; row < kRows; ++row) for (int i = 0; i < kIn; ++i) for (int o = 0; o < kOut; ++o) weightGrad[i * kOut + o] += x[row * kIn + i] * (y[row * kOut + o] - targets[row * kOut + o]);
    // Chain rule through W = W₀ + s·A·B: ∂L/∂A = s · ∂L/∂W · Bᵀ, ∂L/∂B = s · Aᵀ · ∂L/∂W. W₀ gets no update at all.
    std::vector<double> gradA(kIn * kRank, 0.0), gradB(kRank * kOut, 0.0);
    for (int i = 0; i < kIn; ++i) for (int o = 0; o < kOut; ++o) for (int k = 0; k < kRank; ++k) {
        gradA[i * kRank + k] += kScaling * weightGrad[i * kOut + o] * B[k * kOut + o];
        gradB[k * kOut + o] += kScaling * A[i * kRank + k] * weightGrad[i * kOut + o];
    }

    // Check every adapter entry against a centered finite difference of the loss.
    double worst = 0;
    auto check = [&](std::vector<double>& values, const std::vector<double>& analytic) {
        for (std::size_t i = 0; i < values.size(); ++i) {
            const double original = values[i], h = 1e-6;
            values[i] = original + h; const double plus = loss(x, mergedWeight(base, A, B), targets);
            values[i] = original - h; const double minus = loss(x, mergedWeight(base, A, B), targets);
            values[i] = original;
            const double numeric = (plus - minus) / (2 * h);
            worst = std::max(worst, std::fabs(numeric - analytic[i]) / std::max(1e-9, std::fabs(numeric) + std::fabs(analytic[i])));
        }
    };
    check(A, gradA); check(B, gradB);
    std::printf("adapter gradient check: worst relative error %.2e over %zu entries\n", worst, A.size() + B.size());
    assert(worst < 1e-6);

    // Merge identity: the two-path forward x·W₀ + s·(x·A)·B equals the single merged multiply, so inference pays nothing extra after merging.
    std::vector<double> twoPath = matmul(x, base, kRows, kIn, kOut), inner = matmul(x, A, kRows, kIn, kRank), outer = matmul(inner, B, kRows, kRank, kOut);
    double largestGap = 0;
    for (std::size_t i = 0; i < twoPath.size(); ++i) largestGap = std::max(largestGap, std::fabs(twoPath[i] + kScaling * outer[i] - y[i]));
    assert(largestGap < 1e-12);
    std::printf("trainable: %d adapter values versus %d in the full weight (%.0f%%); merge gap %.1e\n", kRank * (kIn + kOut), kIn * kOut, 100.0 * kRank * (kIn + kOut) / (kIn * kOut), largestGap);
    return 0;
}
// LoRA on one linear layer: the adapter's gradients, checked against finite differences, and the merge identity.
// Build: rustc --edition 2021 -O -D warnings lora.rs && ./lora
struct Rng { state: u32 } // xorshift32 with a fixed seed
impl Rng { fn next(&mut self) -> f64 { self.state ^= self.state << 13; self.state ^= self.state >> 17; self.state ^= self.state << 5; self.state as f64 / 4294967296.0 } }

const IN: usize = 8; const OUT: usize = 6; const RANK: usize = 2; const ROWS: usize = 4;
const ALPHA: f64 = 4.0; const SCALING: f64 = ALPHA / RANK as f64; // loralib: scaling = lora_alpha / r

// y[rows][out] = x[rows][in] · W[in][out]
fn matmul(left: &[f64], right: &[f64], rows: usize, inner: usize, columns: usize) -> Vec<f64> {
    let mut out = vec![0.0; rows * columns];
    for row in 0..rows { for i in 0..inner { for c in 0..columns { out[row * columns + c] += left[row * inner + i] * right[i * columns + c]; } } }
    out
}
// The adapted weight the model actually multiplies by: W₀ + scaling · A · B.
fn merged_weight(base: &[f64], a: &[f64], b: &[f64]) -> Vec<f64> {
    let product = matmul(a, b, IN, RANK, OUT);
    base.iter().zip(&product).map(|(w, p)| w + SCALING * p).collect()
}
// Half the squared error between x·W and the targets, summed over the batch.
fn loss(x: &[f64], weight: &[f64], targets: &[f64]) -> f64 {
    matmul(x, weight, ROWS, IN, OUT).iter().zip(targets).map(|(y, t)| 0.5 * (y - t) * (y - t)).sum()
}

fn main() {
    let mut rng = Rng { state: 11 };
    let mut uniform = |size: usize, bound: f64| (0..size).map(|_| (rng.next() * 2.0 - 1.0) * bound).collect::<Vec<f64>>();
    let (x, targets, base) = (uniform(ROWS * IN, 1.0), uniform(ROWS * OUT, 1.0), uniform(IN * OUT, 0.5));
    // loralib initialization: A ~ kaiming_uniform(a=√5), which is uniform in ±1/√in; B = 0, so the adapted layer starts equal to the base layer.
    let mut a = uniform(IN * RANK, 1.0 / (IN as f64).sqrt());
    let mut b = vec![0.0; RANK * OUT];
    assert!(merged_weight(&base, &a, &b) == base);
    // Give B some content so the gradient with respect to A is not identically zero.
    for value in b.iter_mut() { *value = (rng.next() * 2.0 - 1.0) * 0.3; }

    // Forward through the merged weight, then the gradient of the loss with respect to that weight: xᵀ · (y − t).
    let merged = merged_weight(&base, &a, &b);
    let y = matmul(&x, &merged, ROWS, IN, OUT);
    let mut weight_grad = vec![0.0; IN * OUT];
    for row in 0..ROWS { for i in 0..IN { for o in 0..OUT { weight_grad[i * OUT + o] += x[row * IN + i] * (y[row * OUT + o] - targets[row * OUT + o]); } } }
    // Chain rule through W = W₀ + s·A·B: ∂L/∂A = s · ∂L/∂W · Bᵀ, ∂L/∂B = s · Aᵀ · ∂L/∂W. W₀ gets no update at all.
    let (mut grad_a, mut grad_b) = (vec![0.0; IN * RANK], vec![0.0; RANK * OUT]);
    for i in 0..IN { for o in 0..OUT { for k in 0..RANK {
        grad_a[i * RANK + k] += SCALING * weight_grad[i * OUT + o] * b[k * OUT + o];
        grad_b[k * OUT + o] += SCALING * a[i * RANK + k] * weight_grad[i * OUT + o];
    } } }

    // Check every adapter entry against a centered finite difference of the loss.
    let mut worst = 0.0f64;
    for which in 0..2 {
        let count = if which == 0 { a.len() } else { b.len() };
        for i in 0..count {
            let h = 1e-6;
            let original = if which == 0 { a[i] } else { b[i] };
            if which == 0 { a[i] = original + h; } else { b[i] = original + h; }
            let plus = loss(&x, &merged_weight(&base, &a, &b), &targets);
            if which == 0 { a[i] = original - h; } else { b[i] = original - h; }
            let minus = loss(&x, &merged_weight(&base, &a, &b), &targets);
            if which == 0 { a[i] = original; } else { b[i] = original; }
            let numeric = (plus - minus) / (2.0 * h);
            let analytic = if which == 0 { grad_a[i] } else { grad_b[i] };
            worst = worst.max((numeric - analytic).abs() / (numeric.abs() + analytic.abs()).max(1e-9));
        }
    }
    println!("adapter gradient check: worst relative error {:.2e} over {} entries", worst, a.len() + b.len());
    assert!(worst < 1e-6);

    // Merge identity: the two-path forward x·W₀ + s·(x·A)·B equals the single merged multiply, so inference pays nothing extra after merging.
    let two_path = matmul(&x, &base, ROWS, IN, OUT);
    let outer = matmul(&matmul(&x, &a, ROWS, IN, RANK), &b, ROWS, RANK, OUT);
    let largest_gap = two_path.iter().zip(&outer).zip(&y).map(|((base_out, low_rank), merged_out)| (base_out + SCALING * low_rank - merged_out).abs()).fold(0.0, f64::max);
    assert!(largest_gap < 1e-12);
    println!("trainable: {} adapter values versus {} in the full weight ({:.0}%); merge gap {:.1e}", RANK * (IN + OUT), IN * OUT, 100.0 * (RANK * (IN + OUT)) as f64 / (IN * OUT) as f64, largest_gap);
}
Reward model, closed-form optimum, policy gradient and DPO on six responses
// Preference optimization over a finite set of responses: a Bradley–Terry reward model, the KL-regularized
// optimum in closed form, exact policy-gradient ascent toward it, and DPO reaching the same policy from the pairs.
// Build: g++ -std=c++20 -O2 -Wall -Wextra -Wpedantic -Werror preference.cpp && ./a.out
#include <algorithm>
#include <cassert>
#include <cmath>
#include <cstdio>
#include <vector>

struct Rng {  // xorshift32 with a fixed seed
    unsigned state = 5;
    double next() { state ^= state << 13; state ^= state >> 17; state ^= state << 5; return state / 4294967296.0; }
};
struct Comparison { int chosen, rejected; };

double sigmoid(double value) { return 1 / (1 + std::exp(-value)); }
std::vector<double> softmax(const std::vector<double>& logits) {
    const double maximum = *std::max_element(logits.begin(), logits.end());
    std::vector<double> out(logits.size()); double total = 0;
    for (std::size_t i = 0; i < logits.size(); ++i) { out[i] = std::exp(logits[i] - maximum); total += out[i]; }
    for (double& value : out) value /= total;
    return out;
}
// Labels drawn through the Bradley–Terry model from the true rewards: P(i preferred over j) = σ(r_i − r_j).
std::vector<Comparison> sampleComparisons(const std::vector<double>& rewards, int count, Rng& rng) {
    std::vector<Comparison> pairs;
    while (int(pairs.size()) < count) {
        const int first = int(rng.next() * rewards.size()), second = int(rng.next() * rewards.size());
        if (first == second) continue;
        const bool firstWins = rng.next() < sigmoid(rewards[first] - rewards[second]);
        pairs.push_back({firstWins ? first : second, firstWins ? second : first});
    }
    return pairs;
}
double expectedValue(const std::vector<double>& policy, const std::vector<double>& values) {
    double total = 0; for (std::size_t i = 0; i < policy.size(); ++i) total += policy[i] * values[i]; return total;
}
double klDivergence(const std::vector<double>& policy, const std::vector<double>& reference) {
    double total = 0; for (std::size_t i = 0; i < policy.size(); ++i) if (policy[i] > 0) total += policy[i] * std::log(policy[i] / reference[i]); return total;
}

int main() {
    // Six candidate responses to one prompt: the reference model's probabilities and the true (hidden) rewards.
    const std::vector<double> reference{0.40, 0.30, 0.15, 0.10, 0.04, 0.01}, trueRewards{0.0, 0.5, 1.0, 1.5, 2.0, 4.0};
    const double beta = 0.5;
    Rng rng;
    const std::vector<Comparison> training = sampleComparisons(trueRewards, 2000, rng), heldOut = sampleComparisons(trueRewards, 2000, rng);

    // 1. Reward model: gradient descent on −log σ(r_chosen − r_rejected) with a small penalty on the mean, which pins the constant Bradley–Terry cannot determine.
    std::vector<double> fitted(reference.size(), 0.0);
    for (int step = 0; step < 400; ++step) {
        std::vector<double> gradient(fitted.size(), 0.0);
        for (const Comparison& pair : training) {
            const double miss = 1 - sigmoid(fitted[pair.chosen] - fitted[pair.rejected]);  // 1 − P(label) under the current scores
            gradient[pair.chosen] -= miss / training.size(); gradient[pair.rejected] += miss / training.size();
        }
        double mean = 0; for (double value : fitted) mean += value / fitted.size();
        for (std::size_t i = 0; i < fitted.size(); ++i) fitted[i] -= 0.5 * (gradient[i] + 0.01 * 2 * mean / fitted.size());
    }
    int agree = 0; for (const Comparison& pair : heldOut) agree += fitted[pair.chosen] > fitted[pair.rejected];
    std::printf("reward model: held-out agreement %.3f\n", double(agree) / heldOut.size());
    assert(double(agree) / heldOut.size() > 0.7);
    for (std::size_t i = 1; i < fitted.size(); ++i) assert(fitted[i] > fitted[i - 1]);  // the true ordering is recovered

    // 2. The KL-regularized optimum in closed form: π*(y) ∝ π_ref(y) · exp(r(y) / β).
    std::vector<double> optimumLogits(reference.size());
    for (std::size_t i = 0; i < reference.size(); ++i) optimumLogits[i] = std::log(reference[i]) + fitted[i] / beta;
    const std::vector<double> optimum = softmax(optimumLogits);

    // 3. Exact policy gradient on E_π[r] − β·KL(π‖π_ref), starting from the reference: reaches the same policy.
    std::vector<double> logits(reference.size());
    for (std::size_t i = 0; i < reference.size(); ++i) logits[i] = std::log(reference[i]);
    for (int step = 0; step < 3000; ++step) {
        const std::vector<double> policy = softmax(logits);
        std::vector<double> score(policy.size());
        for (std::size_t i = 0; i < policy.size(); ++i) score[i] = fitted[i] - beta * std::log(policy[i] / reference[i]);
        const double baseline = expectedValue(policy, score);
        for (std::size_t i = 0; i < policy.size(); ++i) logits[i] += 0.5 * policy[i] * (score[i] - baseline);
    }
    const std::vector<double> ascended = softmax(logits);
    double gap = 0; for (std::size_t i = 0; i < ascended.size(); ++i) gap = std::max(gap, std::fabs(ascended[i] - optimum[i]));
    std::printf("policy gradient versus closed form: largest probability gap %.1e; KL to reference %.3f, expected reward %.3f\n", gap, klDivergence(ascended, reference), expectedValue(ascended, fitted));
    assert(gap < 1e-4);

    // 4. DPO: no reward model, no sampling. Each pair moves the chosen logit up and the rejected one down by β·σ(−β·margin),
    //    where the margin is the difference of log(π/π_ref); the softmax normalizer cancels in that difference.
    std::vector<double> dpoLogits(reference.size());
    for (std::size_t i = 0; i < reference.size(); ++i) dpoLogits[i] = std::log(reference[i]);
    for (int step = 0; step < 3000; ++step) {
        const std::vector<double> policy = softmax(dpoLogits);
        std::vector<double> gradient(policy.size(), 0.0);
        for (const Comparison& pair : training) {
            const double margin = std::log(policy[pair.chosen] / reference[pair.chosen]) - std::log(policy[pair.rejected] / reference[pair.rejected]);
            const double weight = beta * sigmoid(-beta * margin);
            gradient[pair.chosen] -= weight / training.size(); gradient[pair.rejected] += weight / training.size();
        }
        for (std::size_t i = 0; i < dpoLogits.size(); ++i) dpoLogits[i] -= 0.5 * gradient[i];
    }
    const std::vector<double> dpoPolicy = softmax(dpoLogits);
    double dpoGap = 0; for (std::size_t i = 0; i < dpoPolicy.size(); ++i) dpoGap = std::max(dpoGap, std::fabs(dpoPolicy[i] - optimum[i]));
    std::printf("DPO versus closed form: largest probability gap %.3f; implicit rewards", dpoGap);
    for (std::size_t i = 0; i < dpoPolicy.size(); ++i) std::printf(" %+.2f", beta * std::log(dpoPolicy[i] / reference[i]));  // β·log(π/π_ref): the reward DPO learned, up to a constant
    std::printf("\n");
    assert(dpoGap < 0.05);  // the same pairs, through DPO, give the same policy the two-stage pipeline gives
    return 0;
}
// Preference optimization over a finite set of responses: a Bradley–Terry reward model, the KL-regularized
// optimum in closed form, exact policy-gradient ascent toward it, and DPO reaching the same policy from the pairs.
// Build: rustc --edition 2021 -O -D warnings preference.rs && ./preference
struct Rng { state: u32 } // xorshift32 with a fixed seed
impl Rng { fn next(&mut self) -> f64 { self.state ^= self.state << 13; self.state ^= self.state >> 17; self.state ^= self.state << 5; self.state as f64 / 4294967296.0 } }
struct Comparison { chosen: usize, rejected: usize }

fn sigmoid(value: f64) -> f64 { 1.0 / (1.0 + (-value).exp()) }
fn softmax(logits: &[f64]) -> Vec<f64> {
    let maximum = logits.iter().cloned().fold(f64::NEG_INFINITY, f64::max);
    let exponentials: Vec<f64> = logits.iter().map(|value| (value - maximum).exp()).collect();
    let total: f64 = exponentials.iter().sum();
    exponentials.iter().map(|value| value / total).collect()
}
// Labels drawn through the Bradley–Terry model from the true rewards: P(i preferred over j) = σ(r_i − r_j).
fn sample_comparisons(rewards: &[f64], count: usize, rng: &mut Rng) -> Vec<Comparison> {
    let mut pairs = Vec::new();
    while pairs.len() < count {
        let (first, second) = ((rng.next() * rewards.len() as f64) as usize, (rng.next() * rewards.len() as f64) as usize);
        if first == second { continue; }
        let first_wins = rng.next() < sigmoid(rewards[first] - rewards[second]);
        pairs.push(Comparison { chosen: if first_wins { first } else { second }, rejected: if first_wins { second } else { first } });
    }
    pairs
}
fn expected_value(policy: &[f64], values: &[f64]) -> f64 { policy.iter().zip(values).map(|(p, v)| p * v).sum() }
fn kl_divergence(policy: &[f64], reference: &[f64]) -> f64 { policy.iter().zip(reference).filter(|(p, _)| **p > 0.0).map(|(p, q)| p * (p / q).ln()).sum() }

fn main() {
    // Six candidate responses to one prompt: the reference model's probabilities and the true (hidden) rewards.
    let reference: [f64; 6] = [0.40, 0.30, 0.15, 0.10, 0.04, 0.01];
    let true_rewards: [f64; 6] = [0.0, 0.5, 1.0, 1.5, 2.0, 4.0];
    let beta = 0.5;
    let mut rng = Rng { state: 5 };
    let training = sample_comparisons(&true_rewards, 2000, &mut rng);
    let held_out = sample_comparisons(&true_rewards, 2000, &mut rng);

    // 1. Reward model: gradient descent on −log σ(r_chosen − r_rejected) with a small penalty on the mean, which pins the constant Bradley–Terry cannot determine.
    let mut fitted = vec![0.0; reference.len()];
    for _ in 0..400 {
        let mut gradient = vec![0.0; fitted.len()];
        for pair in &training {
            let miss = 1.0 - sigmoid(fitted[pair.chosen] - fitted[pair.rejected]); // 1 − P(label) under the current scores
            gradient[pair.chosen] -= miss / training.len() as f64; gradient[pair.rejected] += miss / training.len() as f64;
        }
        let mean = fitted.iter().sum::<f64>() / fitted.len() as f64;
        for i in 0..fitted.len() { fitted[i] -= 0.5 * (gradient[i] + 0.01 * 2.0 * mean / fitted.len() as f64); }
    }
    let agree = held_out.iter().filter(|pair| fitted[pair.chosen] > fitted[pair.rejected]).count();
    println!("reward model: held-out agreement {:.3}", agree as f64 / held_out.len() as f64);
    assert!(agree as f64 / held_out.len() as f64 > 0.7);
    for i in 1..fitted.len() { assert!(fitted[i] > fitted[i - 1]); } // the true ordering is recovered

    // 2. The KL-regularized optimum in closed form: π*(y) ∝ π_ref(y) · exp(r(y) / β).
    let optimum_logits: Vec<f64> = (0..reference.len()).map(|i| reference[i].ln() + fitted[i] / beta).collect();
    let optimum = softmax(&optimum_logits);

    // 3. Exact policy gradient on E_π[r] − β·KL(π‖π_ref), starting from the reference: reaches the same policy.
    let mut logits: Vec<f64> = reference.iter().map(|p| p.ln()).collect();
    for _ in 0..3000 {
        let policy = softmax(&logits);
        let score: Vec<f64> = (0..policy.len()).map(|i| fitted[i] - beta * (policy[i] / reference[i]).ln()).collect();
        let baseline = expected_value(&policy, &score);
        for i in 0..policy.len() { logits[i] += 0.5 * policy[i] * (score[i] - baseline); }
    }
    let ascended = softmax(&logits);
    let gap = ascended.iter().zip(&optimum).map(|(a, b)| (a - b).abs()).fold(0.0, f64::max);
    println!("policy gradient versus closed form: largest probability gap {:.1e}; KL to reference {:.3}, expected reward {:.3}", gap, kl_divergence(&ascended, &reference), expected_value(&ascended, &fitted));
    assert!(gap < 1e-4);

    // 4. DPO: no reward model, no sampling. Each pair moves the chosen logit up and the rejected one down by β·σ(−β·margin),
    //    where the margin is the difference of log(π/π_ref); the softmax normalizer cancels in that difference.
    let mut dpo_logits: Vec<f64> = reference.iter().map(|p| p.ln()).collect();
    for _ in 0..3000 {
        let policy = softmax(&dpo_logits);
        let mut gradient = vec![0.0; policy.len()];
        for pair in &training {
            let margin = (policy[pair.chosen] / reference[pair.chosen]).ln() - (policy[pair.rejected] / reference[pair.rejected]).ln();
            let weight = beta * sigmoid(-beta * margin);
            gradient[pair.chosen] -= weight / training.len() as f64; gradient[pair.rejected] += weight / training.len() as f64;
        }
        for i in 0..dpo_logits.len() { dpo_logits[i] -= 0.5 * gradient[i]; }
    }
    let dpo_policy = softmax(&dpo_logits);
    let dpo_gap = dpo_policy.iter().zip(&optimum).map(|(a, b)| (a - b).abs()).fold(0.0, f64::max);
    print!("DPO versus closed form: largest probability gap {:.3}; implicit rewards", dpo_gap);
    for i in 0..dpo_policy.len() { print!(" {:+.2}", beta * (dpo_policy[i] / reference[i]).ln()); } // β·log(π/π_ref): the reward DPO learned, up to a constant
    println!();
    assert!(dpo_gap < 0.05); // the same pairs, through DPO, give the same policy the two-stage pipeline gives
}
What's intentionally missing

A language model behind the reward and policies: the second program's "completions" are six labelled items, so expectations and the normalizer Z are exact sums where a real run has samples and per-token log-probabilities. Dropout on the adapter, the MergedLinear handling of fused QKV weights, PPO's value network, clipping and generalized advantage estimation, reward whitening, the adaptive KL controller, and length normalization of sequence log-probabilities. The losses and gradients are the ones the cited trainers compute.

09Where alignment training misleads

A rising reward is not a better model

The KL widget's optimized reward climbs smoothly while the true reward falls below the reference's. A reward model is a fit to comparisons from a particular distribution of outputs, and the policy's job is to move away from that distribution. Monitor the KL, cap it, and re-evaluate with fresh human or checker judgments on the policy's own outputs, not with the reward model that trained it.

Full fine-tuning on a narrow dataset forgets. The LoRA widget's full run ends with a farm loss above the uniform baseline after 300 iterations on space text. Mixing in pretraining data, lowering the rate, or constraining the update's rank all trade new-domain fit for retention; which trade is right depends on whether the old behaviour is still wanted.

Preference labels are noisy and finite, and the fit degrades gracefully only up to a point. The reward widget shows 20% flipped labels costing almost nothing at 400 comparisons and everything at 20. Report the held-out agreement of a reward model against the agreement the labels' own noise permits, which the widget estimates at 77%, rather than against 100%.

Group-relative methods silently waste compute on prompts that are too easy or too hard for the current policy. Track the fraction of zero-variance groups, as TRL logs it, and adjust the prompt mix or G when it climbs.[15]

10What's next

A fine-tuned model still has to run somewhere. Inference and Quantization covers what a forward pass costs at generation time, where the memory goes, and how weights are compressed to fewer bits without losing the behaviour this page trained in.

11Sources

The reference implementations and trainer documentation define the losses, scalings and defaults quoted; the dataset READMEs and model cards define the data. Widget readouts report values computed in the page from the embedded corpora and sampled comparisons.

  1. Meta, 2024. Meta Llama 3 model card. Instruction-tuned versions produced with supervised fine-tuning and reinforcement learning with human feedback, and fine-tuning data of public instruction datasets plus over 10 million human-annotated examples.
  2. Rohan Taori et al., 2023. Stanford Alpaca: An Instruction-following LLaMA Model, repository README. 52K instruction-following examples generated from text-davinci-003 for under $500, and the fine-tuning hyperparameters: batch 128, learning rate 2e-5, 3 epochs, no weight decay for the 7B model.
  3. OpenAI, 2020. summarize-from-feedback, repository README (paper: arXiv:2009.01325). 64,832 human summary comparisons on TL;DR, with a supervised baseline, a reward model and an RL-fine-tuned policy.
  4. Anthropic, 2022. hh-rlhf, dataset README (paper: arXiv:2204.05862). Each line a chosen and a rejected text; helpfulness data from base models, rejection sampling against an early preference model, and an iterated online process.
  5. Hugging Face. SFT Trainer, TRL documentation. Token-level cross-entropy with a one-token shift, padding ignored through index −100, loss on completion tokens only by default for prompt-completion data, assistant-only loss for conversations, and example packing.
  6. Andrej Karpathy. nanoGPT, model.py. Cross-entropy with ignore_index = −1, the convention the page's kernel uses to mask prompt tokens.
  7. Edward J. Hu et al. loralib/layers.py, reference implementation of LoRA (paper: arXiv:2106.09685). Scaling = lora_alpha / r, A initialized with kaiming_uniform(a=√5) and B with zeros, the forward pass adding (x·Aᵀ)·Bᵀ·scaling, and merging the product into the frozen weight on eval().
  8. Edward J. Hu et al. LoRA: Low-Rank Adaptation of Large Language Models, repository README. Trainable parameter counts and scores: RoBERTa base 125M fully fine-tuned versus 0.8M with LoRA at a higher GLUE average, GPT-2 Medium 354.92M versus 0.35M with higher E2E and DART BLEU; no added inference latency after merging; MergedLinear for fused QKV projections.
  9. Hugging Face. peft/tuners/lora/layer.py. Scaling α/r by default and α/√r with rsLoRA; A as a Linear(in, r) and B as a Linear(r, out).
  10. Hugging Face. Dataset formats and types, TRL documentation. The prompt-completion and preference (prompt, chosen, rejected) formats, in standard and conversational form.
  11. Hugging Face. Reward Trainer, TRL documentation. The Bradley–Terry model P(y⁺ ≻ y⁻) = σ(r(y⁺) − r(y⁻)), the negative log-likelihood loss, the undetermined constant, and the centering penalty with coefficient 1e-2.
  12. OpenAI, 2019. lm-human-preferences, train_policy.py (paper: arXiv:1909.08593). The per-token penalty −kl_coef · (log π − log π_ref) added to the reward with the score at the final token, reward whitening, and the adaptive KL controller that multiplies the coefficient by 1 + clip(KL/target − 1, ±0.2) · n/horizon.
  13. OpenAI, 2019. lm-human-preferences, launch.py. KL coefficients of 0.15 with an adaptive target of 6 nats for the stylistic tasks and 0.01 to 0.03 for summarization, with a note that 0.01 was too low.
  14. Hugging Face. RLOO Trainer, TRL documentation. The leave-one-out baseline over G completions, the KL penalty folded into the scalar reward without a gradient through it, and the clipped importance ratio for repeated updates.
  15. Hugging Face. GRPO Trainer, TRL documentation. G completions per prompt, the advantage (r − mean)/std, the KL estimator ratio − log ratio − 1, the clipped surrogate for repeated updates, β = 0 by default, and the fraction of groups with zero reward standard deviation as a logged metric.
  16. Eric Mitchell et al. direct-preference-optimization, trainers.py, reference implementation (paper: Rafailov et al., 2023, arXiv:2305.18290). The loss −log σ(β·(h⁺ − h⁻)) with h the policy-to-reference log ratio, label smoothing for noisy preferences, the IPO variant, and the implicit rewards β·(log π − log π_ref).
  17. Hugging Face. DPO Trainer, TRL documentation. The DPO objective, the sigmoid loss as a Bradley–Terry classifier, the logged implicit rewards and margins, and the observation that in practice the loss mostly suppresses dispreferred completions.
  18. Eric Mitchell et al. direct-preference-optimization, repository README. SFT first so preference data is in-distribution, β of 0.1 to 0.5 as a starting range, and label smoothing as the assumed fraction of flipped preferences.