AI tutorials Mighty Professional
The Series 路 Artificial Intelligence

Build a Language Model

This is the map of a sixteen-tutorial series: the mathematics, the models, the training machinery and the deployment engineering behind a language model, in the order that lets each page build on the last, from a dot product to a character-level GPT trained in the browser and then tuned, quantized and served. C++ and Rust throughout, every widget computing what it shows. Read it straight through, or drop in at the box you already need.

FormatA hub plus 16 standalone tutorials PathFoundations to a trained model, then alignment and serving LevelBeginner to senior, per page StackC++ & Rust 路 browser demos

01How to read this

This is a collection, not one long page. Every box below is its own deep tutorial, held to the same standard as the rest of the site: a hard concept gets a live widget that computes what it draws, every claim is cited to a paper, a reference implementation or a vendor document, and the C++ and Rust sit side by side. The series puts the pages in order and wires them together.

The order is a dependency graph, not a straight line. Prerequisites come first, and each tutorial links back to what it assumes and forward to what comes next, so you can also arrive sideways: land on Fine-tuning from a search, walk its prerequisites back through the capstone and reinforcement learning, then follow its "next" into inference. Every page in the series is written, and each card below links to it.

You do not have to start at the top

If matrices and softmax are old news, start at Linear Models or Neural Networks. If you only came for transformers, serving, or diffusion, those pages stand on their own and say what they assume. The graph in 搂3 shows what each page leans on, so you can read just the spine that gets you where you want to go.

02The whole system

A language model is layers of a different kind. At the bottom is the mathematics every page uses: vectors, matrices, probability distributions and the cross-entropy that scores a prediction. On that sit the models, from a linear classifier to a transformer, and the machinery that trains them: automatic differentiation, optimizers, schedules and the scaling arithmetic that decides how big to build. Evaluation sits beside training and keeps it honest. Above the trained model come the stages that make it useful, fine-tuning and alignment, and finally the engineering that runs it for many users at once.

The rule that keeps this buildable is the same as for an engine: a layer may lean down, never up. The serving page assumes a trained, quantized model; the training page assumes a differentiable one; nothing below assumes anything above. Click a layer to jump to its tutorials.

Deployment Quantization 路 KV cache 路 Batching 路 Speculative decoding Alignment Reinforcement learning 路 SFT 路 LoRA 路 Reward models 路 RLHF 路 DPO The model Embeddings 路 Attention 路 The transformer block 路 A trained GPT Training & evaluation Initialization 路 Optimizers 路 Precision 路 Scaling 路 Held-out metrics 路 Benchmarks Learning Linear models 路 Gradient descent 路 Neural networks 路 Autodiff Mathematics Linear algebra 路 Probability 路 Information theory

Dependencies point downward only. Real systems blur the lines (the serving page's speculative decoder is a sampling question, and alignment is a training question), but keeping the graph acyclic is what lets you learn and test one layer at a time.

03The path

Reading order is a graph. The widget below shows every page in the series and the prerequisites each one lists, from the two mathematics pages through the trained model to fine-tuning and serving. Click a page, or pick one from the list, to trace its prerequisites.

Click any page to light up the chain of prerequisites it assumes you already know.

Click a page: its prerequisite chain lights up wave by wave, back to the foundations, so you can see what to read first, and the readout links to the tutorial. Transitive edges are left out (Training at Depth needs Neural Networks, but it already reaches it through Automatic Differentiation). Neural Networks states no prerequisite of its own; the series places it after Linear Models because the pages after it assume gradient descent and the cross-entropy loss that Linear Models introduces.

04The curriculum

The whole series, grouped into phases. Each phase ends somewhere useful, and the milestone is where the pieces become a model that writes.

Phase 0

Foundations

The mathematics every later page uses, with the notation the field uses.
Phase 1

Learning from data

From a line through points to a network that computes its own gradients.
Phase 2

Training at depth

What keeps a deep network trainable, and how to tell whether it learned anything.
Phase 3

Language

Turning tokens into vectors, and vectors into a model that mixes them.
Milestone

A trained language model

The point where the pieces become a model that writes.
Phase 4

Making it useful

From a model that continues text to one that answers and prefers good answers.
Phase 5

Deployment

Running the model, and running it for many users at once.
Adjacent

Other model families

The same training machinery applied to images. No page in the spine depends on them.

The Build a Game Engine series is the site's other curriculum; its IEEE-754 Floating Point page is the one the training page's precision section leans on.

05The milestone

A pile of techniques is not a model. One point in the series is where the pieces have to fit together and run, and where you end up with something that writes.

The capstone takes the embedding table from Phase 3, the attention block from Transformers, the backward passes from Automatic Differentiation written out by hand, the AdamW, warmup and cosine schedule from Training at Depth, and the perplexity yardsticks from Evaluation, and trains a one-block GPT on sixteen kilobytes of Shakespeare inside the page, in about a minute of play. The complete C++ and Rust trainers assert that the loss starts at ln V, that every hand-derived gradient matches finite differences, and that the validation loss halves. Everything after it, fine-tuning, inference and serving, starts from the model it produces.

06How each tutorial holds the bar

Every tutorial in this series is held to the same gates. The goal is that a principal AI engineer can read any page and find no claim to push back on.

The reference spine for the whole series is Dive into Deep Learning[1] for the models and optimizers, Bishop's Pattern Recognition and Machine Learning[2] and Mathematics for Machine Learning[3] for the foundations, nanoGPT and llm.c[4] for the model that is built, the TRL and PEFT documentation[5] for alignment, and the vLLM and TensorRT-LLM design documents[6] for serving. Each tutorial researches its own topic deeper from there.

  1. Aston Zhang, Zachary C. Lipton, Mu Li, Alexander J. Smola. Dive into Deep Learning. Cambridge University Press, 2023. d2l.ai. The chapters on linear regression, softmax regression, numerical stability, optimization, language models, attention and word embeddings underpin the model and training pages.
  2. Christopher M. Bishop. Pattern Recognition and Machine Learning. Springer, 2006. Free PDF from the author. The probability, maximum-likelihood and linear-model foundations.
  3. Marc Peter Deisenroth, A. Aldo Faisal, Cheng Soon Ong. Mathematics for Machine Learning. Cambridge University Press, 2020. mml-book.github.io. The linear-algebra and matrix-decomposition foundations.
  4. Andrej Karpathy. nanoGPT and llm.c. The GPT-2-shaped model, its training loop, configurations and reported results that the capstone reproduces at toy scale.
  5. Hugging Face. TRL documentation and PEFT, with Microsoft's LoRA reference implementation and the DPO reference implementation. The losses, scalings and defaults the alignment page quotes.
  6. vLLM contributors, vLLM documentation, and NVIDIA, TensorRT-LLM documentation. Paged attention, prefix caching, in-flight batching, chunked prefill and speculative decoding as implemented.

Good places to start