AI tutorialsMighty Professional
Build a Language Model · Training at depth

Evaluating Models from Scratch

A falling training loss proves that the optimizer works. Whether the model works is a separate question with its own machinery: data the model never saw, metrics that match the decision being made, a baseline to beat, probabilities that mean what they say, and an honest error bar. This page builds each piece on seeded data so every number can be checked, and shows the ways each one is commonly fooled.

Time~50 minLevelIntermediatePrereqsProbability for expectations and cross-entropy; Linear Models for overfitting.StackC++20 & Rust · browser demos
◂ Build a Language ModelPhase 2 · Training at depthNext · Embeddings and vector search ▸

01Three splits and the leak between them

The error a model makes on its own training data says little about the error it will make on new data; the generalization error is what matters, and it is estimated on examples held out from training.[1] Three splits carry three jobs: the training set fits the parameters, the validation set chooses between models and settings, and the test set is used once, at the end, because every choice made against it makes its score optimistic.[1][2]

is any route by which held-out information reaches the fit. The obvious one is tuning on the test set. The subtle one is preprocessing: a scaler, a feature selector or a vocabulary fit on the full dataset before the split has seen the test examples, and the model inherits that knowledge. scikit-learn's own illustration uses features and labels that are pure noise: selecting the most label-correlated features on all the data and then splitting yields accuracy far above chance on a problem where no skill is possible.[2]

Select features before or after the split, on pure noise

One hundred examples with coin-flip labels and Gaussian features, so the true accuracy of any classifier is 50%. Choosing the ten features most correlated with the label on all hundred examples, then evaluating on a held-out half, scores in the 70s and 80s; the selector found the features that happen to correlate with the held-out labels by chance. Selecting on the training half only lands near 50%. More candidate features make the leak worse, because there are more chance correlations to find.

When the data are scarce, k-fold cross-validation reuses every example for validation once and averages the folds. For classification with rare classes, stratified folds keep each class's proportion the same in every fold, so no fold ends up with none of a class.[3] The leakage rule still applies inside each fold: every preprocessing step is fit on that fold's training part alone.[2]

02Generalization and early stopping

Training error and validation error part ways in two ways. Both high is underfitting. Training error far below validation error is overfitting.[1] The linear models page showed the second one as a function of model complexity; it is also a function of training time. A flexible model fits the clean structure of its data first and the noise later, so the validation error falls, bottoms out and rises while the training error keeps falling.[4]

Early stopping turns that curve into a regularizer: monitor validation error each epoch and stop when it has not improved by more than a small margin for a while.[4] Because the validation set now chooses the stopping epoch, it is being used for model selection, and the final number must come from the test set.

Train past the point of best validation error

Each Step runs 50 epochs of gradient descent on a degree-9 polynomial with 16 training points. With seed 1 the validation error bottoms out near 0.22 around epoch 90 and climbs past 1.7 by epoch 3,000 while the training error keeps falling toward 0.09; the right panel shows the fit growing wiggles between the training points. The yellow dot is where early stopping would have stopped. Gradient descent from zero weights fits the low-order shape first, which is why the curve has a minimum at all.

The decomposition behind this is the bias–variance trade-off: an estimator's mean squared error is its squared bias plus its variance plus the noise in the target. A rigid model has high bias; a flexible one fit to a small sample has high variance; the noise term is the floor no model crosses.[5] Early stopping, weight decay and more data all act on the variance term.

03Classification metrics and the baseline

A classifier that outputs a score becomes a decision only once a threshold is chosen, and every metric below depends on that choice except the last. From the confusion matrix: precision is the share of predicted positives that are right, recall is the share of actual positives found, and F1 is their harmonic mean.[6] Accuracy counts both kinds of correct answer and is easily fooled: on a set that is 90% negative, predicting "negative" for everything scores 90%, which is why scikit-learn ships dummy classifiers that predict the most frequent class as the sanity baseline.[6]

P = TP / (TP + FP),   R = TP / (TP + FN),   F₁ = 2PR / (P + R)

Sweeping the threshold from high to low traces the , true-positive rate against false-positive rate. Its area, the AUC, equals the probability that a random positive scores above a random negative, a threshold-free measure of ranking.[6]

Move the threshold, change every metric but one

Six hundred seeded scores: negatives from N(0, 1), positives shifted right by the separation. Moving the threshold trades precision for recall and moves the yellow point along the ROC curve; the AUC does not move, and the readout computes it two ways, by trapezoids under the curve and as the fraction of positive–negative pairs ranked correctly, which agree. Switch to 10% positives: at threshold 1 the accuracy is 85.0% while the majority-class baseline is 90.7%, so the classifier is below the baseline on accuracy while its recall is 84% and its precision 37%. Accuracy is the wrong summary for that problem.

04Calibration

A model that reports 0.9 should be right about nine times in ten when it says so. That property is , and it is independent of accuracy: the ordering of scores can be perfect while the numbers attached to them are all too confident.[7] A reliability diagram makes it visible, and the expected calibration error compresses it to one number.[7][8]

Modern networks trained with cross-entropy tend to come out overconfident, and temperature scaling is a simple post-hoc repair: divide the logits by one scalar T chosen on held-out data, which leaves every argmax, and so the accuracy, unchanged.[8] Log loss and the Brier score are proper scoring rules: they reward calibration and discrimination together, which accuracy and ECE each measure alone.[6]

Over-sharpen a calibrated model, then fix it with one temperature

Two thousand seeded examples whose labels are drawn from a known probability; a model that reports that probability is calibrated by construction. Sharpening its logits by 2.5 leaves accuracy at 70.9% and makes its probabilities too extreme in both directions: the purple bars (the observed fraction of positives in each bin) sit above the green mean predicted probability in the bins below 0.5 and below it in the bins above 0.5; ECE is 0.13 and log loss 0.66. Set T to the fitted value near 2.3 and the bars meet, ECE drops to about 0.02 and log loss to 0.56, with accuracy unchanged. Temperature cannot fix a model whose errors are not a uniform over- or under-confidence.

05Perplexity and its context

A language model's natural metric is the cross-entropy of the probability page applied to text: the average negative log-likelihood of each token given the ones before it. is its exponential.[9]

PPL(X) = exp(−(1/t) Σi log p(xi | x<i))

Two things change the number without changing the model. It is per token, so a tokenizer that splits text into more pieces yields a different value for the same text; bits per character, the total negative log-likelihood in bits divided by the number of characters, is the tokenizer-independent form. And it depends on how much context each token is given: a model with a fixed window evaluated on non-overlapping chunks scores the first tokens of each chunk with little or no context, which raises the perplexity; a sliding window with a stride gives every token a long context at the cost of more forward passes.[9]

Score a sentence under a trigram model with and without context

A character trigram model with add-one smoothing, trained on a 200-character text, scores each character given up to two previous ones. With unlimited context the first sentence has perplexity 10.28, 3.36 bits per character. Non-overlapping windows of 8 restart the context every 8 characters: the spikes in the bars are those positions, and the perplexity rises to 11.41. Windows of 2 raise it to 15.1. The sliding variant with stride 4 gives every character its full two-character context and matches the unlimited figure, with the same model throughout.

06Benchmarks, contamination and error bars

Public benchmarks are fixed test sets with a scoring rule: MMLU is multiple-choice questions over 57 subjects with a 25% random baseline; GSM8K's test split is 1,319 grade-school word problems scored on the final number; HumanEval scores generated code by running unit tests.[10][11][12] Because the test sets are public, they leak into training corpora. The decontamination procedure in the lm-evaluation-harness follows GPT-3's Appendix C with n fixed at 13 (GPT-3 used 8 to 13 depending on the dataset): a test item counts as contaminated when a 13-gram of it appears in the training data, and a clean score excludes those items.[13]

For code, one sample per problem understates what a model can do, so HumanEval reports pass@k: the probability that at least one of k samples passes. Generating exactly k samples per problem and checking whether any passes gives a high-variance estimate; the HumanEval evaluation code instead takes n ≥ k samples, counts the c that pass and uses the unbiased estimator 1 − C(n − c, k)/C(n, k); it skips pass@k when n < k, where no unbiased estimate exists.[12]

Every benchmark score is a sample mean over a finite set of questions, so it has the standard error of a sample mean. For a score p over n questions, treat each question as a coin with bias p: the standard error is √(p(1 − p)/n), and a 95% interval is roughly ±2 of those. A one-point difference between two models on a thousand questions is inside the noise.

Put an error bar on a benchmark score

The curve is the exact binomial distribution of the measured score, computed from the true accuracy and the question count. At 1,000 questions and 70% a run lands within about ±2.8 points of the truth 95% of the time, and a model that is really 2 points worse outscores the better one on a fresh thousand-question set about 17% of the time. At 100 questions that becomes 38%, close to a coin flip. The approximation treats questions as independent; correlated questions make the real noise larger.

07The metrics, checked

These complete programs implement the confusion-matrix metrics, the ROC AUC by two methods, expected calibration error, the pass@k estimator and perplexity from a list of token log-likelihoods, and assert hand-computed values for each: the two AUC methods agree on a ranked example, pass@k matches a direct combinatorial count, and ECE matches a two-bin hand calculation.

Evaluation metrics, complete programs
#include <algorithm>
#include <cassert>
#include <cmath>
#include <cstddef>
#include <vector>
struct Scored { double score; int label; };
struct Metrics { double precision, recall, f1, accuracy; };
Metrics thresholdMetrics(const std::vector<Scored>& data, double threshold) {
    int truePositive = 0, falsePositive = 0, falseNegative = 0, trueNegative = 0;
    for (const Scored& example : data) {
        const bool predicted = example.score >= threshold;
        if (predicted && example.label) ++truePositive; else if (predicted) ++falsePositive;
        else if (example.label) ++falseNegative; else ++trueNegative;
    }
    const double precision = double(truePositive) / (truePositive + falsePositive), recall = double(truePositive) / (truePositive + falseNegative);
    return {precision, recall, 2 * precision * recall / (precision + recall), double(truePositive + trueNegative) / data.size()};
}
// AUC as the fraction of (positive, negative) pairs the scores rank correctly; ties count one half.
double rankAuc(const std::vector<Scored>& data) {
    double concordant = 0; int pairs = 0;
    for (const Scored& positive : data) if (positive.label)
        for (const Scored& negative : data) if (!negative.label) {
            ++pairs; concordant += positive.score > negative.score ? 1.0 : positive.score == negative.score ? 0.5 : 0.0;
        }
    return concordant / pairs;
}
// AUC by the trapezoid rule under the ROC curve traced by sweeping the threshold downward.
double trapezoidAuc(std::vector<Scored> data) {
    std::sort(data.begin(), data.end(), [](const Scored& left, const Scored& right) { return left.score > right.score; });
    int positives = 0; for (const Scored& example : data) positives += example.label;
    const int negatives = int(data.size()) - positives;
    double area = 0, previousFpr = 0, previousTpr = 0; int truePositive = 0, falsePositive = 0;
    for (const Scored& example : data) {
        if (example.label) ++truePositive; else ++falsePositive;
        const double fpr = double(falsePositive) / negatives, tpr = double(truePositive) / positives;
        area += (fpr - previousFpr) * (tpr + previousTpr) / 2; previousFpr = fpr; previousTpr = tpr;
    }
    return area;
}
// Expected calibration error over equal-width confidence bins.
double expectedCalibrationError(const std::vector<double>& confidences, const std::vector<int>& correct, int bins) {
    std::vector<double> sumConfidence(bins, 0), sumCorrect(bins, 0); std::vector<int> count(bins, 0);
    for (std::size_t index = 0; index < confidences.size(); ++index) {
        const int bin = std::min(bins - 1, int(confidences[index] * bins));
        ++count[bin]; sumConfidence[bin] += confidences[index]; sumCorrect[bin] += correct[index];
    }
    double error = 0;
    for (int bin = 0; bin < bins; ++bin) if (count[bin])
        error += double(count[bin]) / confidences.size() * std::abs(sumCorrect[bin] / count[bin] - sumConfidence[bin] / count[bin]);
    return error;
}
// HumanEval's unbiased pass@k from n samples of which c passed: 1 − C(n−c, k) / C(n, k).
double passAtK(int samples, int passed, int k) {
    if (samples - passed < k) return 1.0;
    double product = 1;
    for (int index = samples - passed + 1; index <= samples; ++index) product *= 1.0 - double(k) / index;
    return 1.0 - product;
}
double perplexity(const std::vector<double>& logProbabilities) {
    double total = 0;
    for (double value : logProbabilities) total += value;
    return std::exp(-total / logProbabilities.size()); // exp of the mean negative log-likelihood.
}
int main() {
    const std::vector<Scored> data{{0.9, 1}, {0.8, 1}, {0.7, 0}, {0.6, 1}, {0.4, 0}, {0.3, 0}, {0.2, 1}, {0.1, 0}};
    const Metrics metrics = thresholdMetrics(data, 0.5); // Predicted positive: 0.9, 0.8, 0.7, 0.6 → TP 3, FP 1, FN 1, TN 3.
    assert(std::abs(metrics.precision - 0.75) < 1e-12 && std::abs(metrics.recall - 0.75) < 1e-12);
    assert(std::abs(metrics.f1 - 0.75) < 1e-12 && std::abs(metrics.accuracy - 0.75) < 1e-12);
    assert(std::abs(rankAuc(data) - 12.0 / 16.0) < 1e-12);                    // 12 of the 16 positive–negative pairs are ranked correctly.
    assert(std::abs(trapezoidAuc(data) - rankAuc(data)) < 1e-12);          // Both methods agree.
    const std::vector<double> confidences{0.95, 0.95, 0.95, 0.95, 0.55, 0.55};
    const std::vector<int> correct{1, 1, 1, 0, 1, 0};             // 0.95 bin: 75% right; 0.55 bin: 50% right.
    assert(std::abs(expectedCalibrationError(confidences, correct, 10) - (4.0 / 6 * 0.20 + 2.0 / 6 * 0.05)) < 1e-12);
    assert(std::abs(passAtK(10, 3, 1) - 0.3) < 1e-12);                      // One sample: the plain pass rate.
    assert(std::abs(passAtK(10, 3, 5) - (1.0 - 21.0 / 252.0)) < 1e-12);     // C(7,5) = 21 all-failing draws of C(10,5) = 252.
    assert(std::abs(perplexity({std::log(0.5), std::log(0.5), std::log(0.5)}) - 2.0) < 1e-12); // Coin-flip tokens: perplexity 2.
}
struct Scored { score: f64, label: i32 }
struct Metrics { precision: f64, recall: f64, f1: f64, accuracy: f64 }
fn threshold_metrics(data: &[Scored], threshold: f64) -> Metrics {
    let (mut true_positive, mut false_positive, mut false_negative, mut true_negative) = (0, 0, 0, 0);
    for example in data {
        let predicted = example.score >= threshold;
        if predicted && example.label == 1 { true_positive += 1; } else if predicted { false_positive += 1; }
        else if example.label == 1 { false_negative += 1; } else { true_negative += 1; }
    }
    let precision = true_positive as f64 / (true_positive + false_positive) as f64;
    let recall = true_positive as f64 / (true_positive + false_negative) as f64;
    Metrics { precision, recall, f1: 2.0 * precision * recall / (precision + recall), accuracy: (true_positive + true_negative) as f64 / data.len() as f64 }
}
// AUC as the fraction of (positive, negative) pairs the scores rank correctly; ties count one half.
fn rank_auc(data: &[Scored]) -> f64 {
    let (mut concordant, mut pairs) = (0.0, 0);
    for positive in data.iter().filter(|example| example.label == 1) {
        for negative in data.iter().filter(|example| example.label == 0) {
            pairs += 1;
            concordant += if positive.score > negative.score { 1.0 } else if positive.score == negative.score { 0.5 } else { 0.0 };
        }
    }
    concordant / pairs as f64
}
// AUC by the trapezoid rule under the ROC curve traced by sweeping the threshold downward.
fn trapezoid_auc(data: &[Scored]) -> f64 {
    let mut sorted: Vec<&Scored> = data.iter().collect();
    sorted.sort_by(|left, right| right.score.partial_cmp(&left.score).unwrap());
    let positives = data.iter().filter(|example| example.label == 1).count() as f64;
    let negatives = data.len() as f64 - positives;
    let (mut area, mut previous_fpr, mut previous_tpr, mut true_positive, mut false_positive) = (0.0, 0.0, 0.0, 0.0, 0.0);
    for example in sorted {
        if example.label == 1 { true_positive += 1.0; } else { false_positive += 1.0; }
        let (fpr, tpr) = (false_positive / negatives, true_positive / positives);
        area += (fpr - previous_fpr) * (tpr + previous_tpr) / 2.0;
        previous_fpr = fpr; previous_tpr = tpr;
    }
    area
}
// Expected calibration error over equal-width confidence bins.
fn expected_calibration_error(confidences: &[f64], correct: &[i32], bins: usize) -> f64 {
    let (mut sum_confidence, mut sum_correct, mut count) = (vec![0.0; bins], vec![0.0; bins], vec![0; bins]);
    for (index, &confidence) in confidences.iter().enumerate() {
        let bin = ((confidence * bins as f64) as usize).min(bins - 1);
        count[bin] += 1; sum_confidence[bin] += confidence; sum_correct[bin] += correct[index] as f64;
    }
    let mut error = 0.0;
    for bin in 0..bins {
        if count[bin] > 0 { error += count[bin] as f64 / confidences.len() as f64 * (sum_correct[bin] / count[bin] as f64 - sum_confidence[bin] / count[bin] as f64).abs(); }
    }
    error
}
// HumanEval's unbiased pass@k from n samples of which c passed: 1 − C(n−c, k) / C(n, k).
fn pass_at_k(samples: i32, passed: i32, k: i32) -> f64 {
    if samples - passed < k { return 1.0; }
    let product: f64 = (samples - passed + 1..=samples).map(|index| 1.0 - k as f64 / index as f64).product();
    1.0 - product
}
fn perplexity(log_probabilities: &[f64]) -> f64 {
    (-log_probabilities.iter().sum::<f64>() / log_probabilities.len() as f64).exp() // exp of the mean negative log-likelihood.
}
fn main() {
    let data = [Scored { score: 0.9, label: 1 }, Scored { score: 0.8, label: 1 }, Scored { score: 0.7, label: 0 }, Scored { score: 0.6, label: 1 },
                Scored { score: 0.4, label: 0 }, Scored { score: 0.3, label: 0 }, Scored { score: 0.2, label: 1 }, Scored { score: 0.1, label: 0 }];
    let metrics = threshold_metrics(&data, 0.5); // Predicted positive: 0.9, 0.8, 0.7, 0.6 → TP 3, FP 1, FN 1, TN 3.
    assert!((metrics.precision - 0.75).abs() < 1e-12 && (metrics.recall - 0.75).abs() < 1e-12);
    assert!((metrics.f1 - 0.75).abs() < 1e-12 && (metrics.accuracy - 0.75).abs() < 1e-12);
    assert!((rank_auc(&data) - 12.0 / 16.0).abs() < 1e-12);                   // 12 of the 16 positive–negative pairs are ranked correctly.
    assert!((trapezoid_auc(&data) - rank_auc(&data)).abs() < 1e-12);         // Both methods agree.
    let confidences = [0.95, 0.95, 0.95, 0.95, 0.55, 0.55];
    let correct = [1, 1, 1, 0, 1, 0];                                       // 0.95 bin: 75% right; 0.55 bin: 50% right.
    assert!((expected_calibration_error(&confidences, &correct, 10) - (4.0 / 6.0 * 0.20 + 2.0 / 6.0 * 0.05)).abs() < 1e-12);
    assert!((pass_at_k(10, 3, 1) - 0.3).abs() < 1e-12);                       // One sample: the plain pass rate.
    assert!((pass_at_k(10, 3, 5) - (1.0 - 21.0 / 252.0)).abs() < 1e-12);      // C(7,5) = 21 all-failing draws of C(10,5) = 252.
    let half = 0.5_f64.ln();
    assert!((perplexity(&[half, half, half]) - 2.0).abs() < 1e-12);         // Coin-flip tokens: perplexity 2.
}
What's intentionally missing

Multi-class averaging (macro, micro, weighted), confidence intervals around each metric, bootstrap resampling for paired model comparisons, the striding logic for perplexity over long texts, and the sandboxed execution that makes pass@k safe to compute on generated code.

08Where evaluations mislead

Report every score with its baseline and its error bar

Report what a trivial predictor scores on the same split and metric, and the standard error of the score. A 92% that sits next to a 91% majority baseline with ±3 points of noise is a different finding from a 92% next to 50% with ±0.5.

Validation scores drift upward with every experiment run against them, because the best of many noisy comparisons is biased high. The test set exists to catch that; use it once, and if it has been used more than once, say so.

Perplexities from different papers are comparable only when the tokenizer, the evaluation text and the context protocol match. A lower number from a model with a larger vocabulary, or evaluated with a sliding window, may reflect the protocol rather than the model.[9]

Benchmark contamination comes from the training data and does not show in the score. For a model trained on a web crawl, assume public test sets were in the crawl unless the training data were filtered against them, and prefer scores on held-out sets created after the training cutoff.[13]

09What's next

The rest of the series is about language. Embeddings and Vector Search turns words and documents into vectors whose dot products mean something, trains them from co-occurrence, and searches them at scale; it is the representation the transformer reads and the retrieval system that feeds it.

10Sources

The cited documentation defines the metrics and protocols; the benchmark repositories define their own scoring. Widget readouts report values computed in the page from seeded data.

  1. Aston Zhang, Zachary C. Lipton, Mu Li, Alexander J. Smola, 2023. Dive into Deep Learning, Section 3.6: Generalization, Cambridge University Press. Training versus generalization error, underfitting and overfitting, the validation set, and not touching the test set.
  2. scikit-learn developers. Common pitfalls and recommended practices. Data leakage, fitting preprocessing on training data only, and the random-feature selection example.
  3. scikit-learn developers. Cross-validation: evaluating estimator performance. k-fold cross-validation and stratified folds.
  4. Aston Zhang et al., 2023. Dive into Deep Learning, Section 5.5: Generalization in Deep Learning. Networks fitting clean data before noise, and early stopping on validation error.
  5. Aston Zhang et al., 2023. Dive into Deep Learning, Appendix: Statistics. The bias–variance decomposition of an estimator's mean squared error and the interpretation of confidence intervals.
  6. scikit-learn developers. Metrics and scoring. Precision, recall and F-measures, the ROC curve and AUC, log loss and the Brier score as proper scoring rules, and dummy estimators as baselines.
  7. scikit-learn developers. Probability calibration. Reliability diagrams and the meaning of a calibrated classifier.
  8. Geoff Pleiss. temperature_scaling, reference code for Guo et al. (2017), On Calibration of Modern Neural Networks. Overconfidence of modern networks, temperature scaling as a single learned divisor of the logits, and the binned expected calibration error.
  9. Hugging Face Transformers contributors. Perplexity of fixed-length models. The definition as exponentiated mean negative log-likelihood, and the dependence on non-overlapping versus strided evaluation windows.
  10. Dan Hendrycks et al., 2021. Measuring Massive Multitask Language Understanding, repository. Multiple-choice tasks with a 25% random baseline; the repository's categories.py maps the 57 subjects.
  11. Karl Cobbe et al., 2021. GSM8K, repository. 8.5K grade-school problems split into 7.5K training and 1K test (1,319 lines in the released test.jsonl), scored on the final numeric answer.
  12. Mark Chen et al., 2021. HumanEval, repository. Execution-based scoring and the unbiased pass@k estimator 1 − C(n − c, k)/C(n, k).
  13. EleutherAI. lm-evaluation-harness: Decontamination. Contamination as 13-gram overlap with training data, following Appendix C of the GPT-3 paper, and clean-score reporting.