AI Inference and Quantization from Scratch
Run a trained function without losing track of its numerical meaning. Compute quantization errors, verify integer arithmetic, estimate memory limits, and inspect how a decoding rule changes probabilities.
01Preserve the computation before optimizing it
Inference evaluates a trained function on new inputs. Its deployed behavior includes preprocessing, model parameters, numerical kernels, and output interpretation. Reducing storage or changing a sampling rule affects different parts of that system. Begin with a trusted reference result so those changes can be measured separately.
A numerical problem can appear before quantization. Logits are often converted to probabilities with softmax. Directly exponentiating large positive logits can overflow. Subtracting their maximum leaves the exact probabilities unchanged while limiting all exponential arguments to zero or below.[1]
A common shift cancels between numerator and denominator. The formula assumes at least one finite valid logit; masks and non-finite inputs need explicit handling.
For exact arithmetic, adding a common constant to all logits changes none of the probabilities. In floating point, the stable shifted formula avoids positive-exponential overflow, though it cannot recover distinctions already lost when extremely large scores were rounded. Very negative shifted logits can underflow to zero; that is a different effect from an invalid infinity divided by infinity.
Compare intermediate arrays when validating a port. Equal final class labels can hide significant probability changes. Conversely, a tiny rounding difference can flip a nearly tied argmax without indicating a large numerical error. Inspect both numerical differences and the task outcome.
02An integer code needs a scale and zero point
stores a real-valued array using integers and parameters that interpret them. The integer code alone is not the real number. Two arrays with identical codes but different scales represent different values.[2]
reconstructed = scale × (integer − zero_point)
scale is positive; zero_point is an integer representing real zero. Rounding and clipping introduce different kinds of error. The widget uses unsigned codes; signed codes are another convention.
If the code range has 2ᵇ distinct integers, a common min/max choice uses scale=(high−low)/(2ᵇ−1). Include real zero when choosing the range, then choose an integer zero point. Rounding the zero point can slightly move the represented endpoints. Within range, scalar rounding error is at most roughly half a step; outside range, clipping error can be much larger.[2]
The slider changes the representable interval for a fixed array. A wide interval avoids clipping but spreads a fixed number of codes farther apart. A narrow interval resolves small values more finely while saturating the large ones. The displayed mean squared error measures this particular array; it does not report model accuracy.
The rounding rule at an exact half-step must match between implementations. The reference code and widgets use round-to-nearest with ties away from zero before adding the zero point. Other deployed formats can use ties-to-even. Neither rule should be inferred solely from the word "round."
Symmetric quantization simplifies the zero point, but can leave range unused when values are strongly asymmetric. The most useful representation depends on the tensors and supported kernels, not just the nominal bit count.
03One range can waste resolution in another channel
A per-tensor scale covers an entire array. A per-channel scale covers a chosen channel axis independently. The latter can preserve a small channel's resolution when another channel contains much larger values, at the cost of extra scale parameters and a kernel that supports their use.[3]
The next example has two weight rows, treated as output channels. One row contains a large coefficient. The other contains values around a tenth. With a shared symmetric scale, the large row determines the step size for both. With separate row scales, the small row gets a finer representation. This is a direct reconstruction comparison, not a prediction that a given model will retain its accuracy.
Transformer activation outliers introduce another issue. Dettmers et al.'s LLM.int8 method combines vector-wise quantization with a higher-precision path for outlier dimensions. That is more specific than "cast every tensor to int8." SmoothQuant instead applies an equivalent rescaling that moves some range difficulty from activations into weights before quantization.[4][5]
s is a positive channel-wise rescaling vector. The exact real-valued matrix product is unchanged; subsequent quantization errors can differ. Which axis is rescaled must match the matrix multiplication.
These methods solve different problems from choosing one scale per output row in the toy widget. Output-channel weight scales and input-feature activation outliers refer to different axes. Name the tensor and axis before claiming that "per-channel quantization" addresses an observed error.
selects fixed activation parameters from representative data. ONNX Runtime distinguishes static calibration from dynamic activation quantization, which computes parameters during execution. Dynamic ranges add work; fixed ranges can misrepresent inputs outside calibration. Evaluate the chosen path on held-out deployment-like inputs.[3]
04Integer products must account for zero points
A quantized dot product can multiply integer-centered codes and accumulate them in a wider integer. Subtract the two zero points before multiplication, then apply the product of the scales. Multiplying raw codes interprets the zero-point offsets as signal.[2]
real_output ≈ input_scale × weight_scale × accumulator
This reference returns a real number. A fully integer pipeline can rescale the accumulator into the next tensor's integer format, add bias in compatible units, and apply clipping.
Because accumulation includes many products, accumulator width matters. With unsigned 8-bit codes and arbitrary unsigned zero points, centered magnitudes can be as large as 255. A length-K dot therefore has the conservative magnitude bound K×255² before bias. Choosing int32 is safe only within an appropriate workload bound; it is not a guarantee for unbounded vector length.
The omitted bias and output rescaling are not incidental in a deployed integer-only network. Bias must use units compatible with the product scale; output quantization needs its own scale and zero point. Framework graph transformations and hardware kernels must agree on where each conversion occurs. Jacob et al. describe these pieces as part of an integer inference pipeline.[2]
A weight-only quantized model may unpack low-bit weights inside a floating-point computation while leaving activations in floating point. An all-int8 matrix multiply uses a different arithmetic path. A saved model's nominal bit count does not establish which path ran, how much unpacking it required, or whether the kernel became faster.
05Estimate the bottleneck with a stated workload
The compares two optimistic limits: arithmetic throughput and memory bandwidth. Arithmetic intensity connects them. State which memory boundary and bytes are counted; a cache-level estimate and device-memory estimate are different models.[6]
time ≥ max(operation_count / peak, byte_count / bandwidth)
The inequalities are optimistic bounds. Launch overhead, dependencies, scheduling, conversion work, and incomplete hardware utilization can make actual latency larger.
For a dense M×K weight matrix applied to B input rows, count 2BMK arithmetic operations when a multiply and an add count separately. If the weights are fetched once for the whole batch, increasing B reuses them and increases work per transferred weight byte. The model also counts input reads and output writes. It omits intermediates and assumes those transfers occur once.
At batch one, a large matrix-vector operation can have little reuse of each loaded weight. Batching gives a matrix-matrix operation more reuse, but may increase request waiting time. Throughput, time to first output, and per-request latency are different measurements. The bound above models execution of a specified layer; it does not include batching queue delay.
Decoder inference also stores attention state. As derived in the transformer article, a dense KV payload occupies 2×layers×batch×KV-heads×tokens×head-width×bytes-per-entry bytes. Reducing weight storage does not automatically reduce this cache. The cache can grow as sequences grow, and full-context attention still reads previous keys and values.[7]
06A sampling rule changes the output distribution
After a language model produces logits, a decoding rule selects the next token. Temperature T>0 gives softmax(logits/T). Lower T makes a unique maximum more dominant; higher T flattens differences. Exact T=0 needs a separate greedy rule rather than division by zero.[8]
sorts tokens by probability and retains the smallest leading set whose cumulative mass reaches a threshold p. Then it renormalizes that set and samples. A token crossing the threshold is retained. This makes the candidate count depend on the distribution rather than a fixed top-k count.[8]
Lowering temperature does not make incorrect logits correct. A confidently preferred token can still be wrong. Restricting the tail changes diversity and likelihood of low-probability choices; it is not a factuality certificate. Holtzman et al. study text-generation behavior under decoding strategies, not a universal optimum for every application.
Record the order of transformations. Applying a cutoff before temperature scaling can retain a different token set from scaling first. Handle ties deterministically when reproducibility matters, and verify that at least one token remains. The model's learned probabilities and the sampler's final probabilities are separate arrays.
07An affine integer dot product with assertions
The complete programs below quantize four values using unsigned 8-bit codes and return a real-valued dot product. Their executable checks verify equality to the reconstructed-vector product, preservation of real zero, and the error on the displayed original input. The fixed four-term workload bounds the accumulator by 4×255²=260,100 before bias. Codes are held in widened integer slots for this readable reference, rather than packed bytes.
#include <algorithm>
#include <array>
#include <cassert>
#include <cmath>
#include <cstdint>
struct Quantized { std::array<std::int32_t,4> codes; double scale; std::int32_t zero; };
Quantized quantize(const std::array<double,4>& values, double low, double high) {
assert(std::isfinite(low) && std::isfinite(high));
assert(low <= 0 && high >= 0 && high > low);
const double scale = (high-low)/255.0;
assert(std::isfinite(scale) && scale > 0);
// Include zero exactly; its code is rounded with the same tie convention.
const std::int32_t zero = static_cast<std::int32_t>(std::round(-low/scale));
Quantized result{{},scale,zero};
for (int index = 0; index < 4; ++index) {
assert(std::isfinite(values[index]));
// Clip while still floating point so huge finite inputs cannot overflow an integer cast.
const double code = std::clamp(std::round(values[index]/scale)+zero,0.0,255.0);
result.codes[index] = static_cast<std::int32_t>(code);
}
return result;
}
double dot(const Quantized& left, const Quantized& right) {
std::int32_t accumulator = 0;
for (int index = 0; index < 4; ++index)
// Codes are already widened; centering removes offsets before multiplication.
accumulator += (left.codes[index]-left.zero)*(right.codes[index]-right.zero);
return left.scale*right.scale*accumulator; // Return a real value for this reference.
}
int main() {
const auto left = quantize({-1,0.3,2,0.1},-2,3);
const auto right = quantize({0.5,-0.2,1,1.2},-1,2);
double reconstructed = 0;
for (int index = 0; index < 4; ++index)
reconstructed += left.scale*(left.codes[index]-left.zero)
* right.scale*(right.codes[index]-right.zero);
assert(std::abs(dot(left,right)-reconstructed) < 1e-12);
const auto zeros = quantize({0,0,0,0},-2,3);
assert(dot(zeros,right) == 0); // Exact real zero survives the affine representation.
assert(std::abs(dot(left,right)-1.56) < 0.04); // Float dot is 1.56 for this input.
const auto extremes = quantize({1e100,-1e100,0,1},-2,3);
assert(extremes.codes[0] == 255 && extremes.codes[1] == 0); // Clipping precedes conversion.
const auto ties = quantize({0.5,-0.5,0,0},-127.5,127.5);
assert(ties.codes[0]-ties.zero == 1 && ties.codes[1]-ties.zero == -1);
}
struct Quantized { codes: [i32;4], scale: f64, zero: i32 }
fn quantize(values: [f64;4], low: f64, high: f64) -> Quantized {
assert!(low.is_finite() && high.is_finite());
assert!(low <= 0.0 && high >= 0.0 && high > low);
let scale = (high-low)/255.0;
assert!(scale.is_finite() && scale > 0.0);
// Include zero exactly; its code is rounded with the same tie convention.
let zero = (-low/scale).round() as i32;
let mut result = Quantized { codes: [0;4],scale,zero };
for (index,value) in values.iter().enumerate() {
assert!(value.is_finite());
// Clip while still floating point so huge finite inputs cannot overflow integer addition.
let code = ((value/scale).round()+f64::from(zero)).clamp(0.0,255.0);
result.codes[index] = code as i32;
}
result
}
fn dot(left: &Quantized, right: &Quantized) -> f64 {
let mut accumulator: i32 = 0;
for index in 0..4 {
// Codes are already widened; centering removes offsets before multiplication.
accumulator += (left.codes[index]-left.zero)*(right.codes[index]-right.zero);
}
left.scale*right.scale*f64::from(accumulator) // Real-valued reference result.
}
fn main() {
let left = quantize([-1.0,0.3,2.0,0.1],-2.0,3.0);
let right = quantize([0.5,-0.2,1.0,1.2],-1.0,2.0);
let mut reconstructed = 0.0;
for index in 0..4 {
reconstructed += left.scale*f64::from(left.codes[index]-left.zero)
* right.scale*f64::from(right.codes[index]-right.zero);
}
assert!((dot(&left,&right)-reconstructed).abs() < 1e-12);
let zeros = quantize([0.0;4],-2.0,3.0);
assert_eq!(dot(&zeros,&right),0.0); // Exact real zero survives affine coding.
assert!((dot(&left,&right)-1.56).abs() < 0.04); // Float dot is 1.56 for this input.
let extremes = quantize([1e100,-1e100,0.0,1.0],-2.0,3.0);
assert!(extremes.codes[0] == 255 && extremes.codes[1] == 0); // Clipping precedes conversion.
let ties = quantize([0.5,-0.5,0.0,0.0],-127.5,127.5);
assert!(ties.codes[0]-ties.zero == 1 && ties.codes[1]-ties.zero == -1);
}
This code defines the reference calculation. It does not pack four-bit codes, use vector instructions, fuse bias and activation, or implement output requantization. A model deployment needs those choices to match its graph and selected hardware backend. The browser's numerical demo similarly shows representations and results, not an optimized int8 throughput test.
Keep representative tensors and end-to-end examples from the unquantized reference. First compare reconstruction and layer-output differences. Then evaluate the actual task on held-out data. Finally benchmark the deployed execution path under the intended batch sizes and sequence lengths. Numerical fidelity, task quality, and performance answer different questions.
08Measure the executed model
A smaller serialized model can still run a higher-precision fallback. Inspect the selected backend and operators when measuring a deployment. If a range or axis is wrong, compare intermediate values before tuning the final decoding rule.
Report benchmark conditions: hardware, runtime version, batch size, prompt and generated lengths, warmup, and whether transfers and tokenization are included. Separate memory occupancy from bytes transferred. A tensor can remain allocated while being read repeatedly; the two numbers have different implications.
The widgets compute exact errors for displayed arrays. A model needs its own accuracy and robustness evaluation, including representative inputs that stress calibration ranges. Acceptable error depends on the downstream task.
09What's next
The six-part Artificial Intelligence series now connects neural-network training to image models, attention, learned decisions, generation, and inference. Return to the foundation trainer to change a learned function, or to the transformer cache demo to compare execution without changing its mathematical result.
10Sources
The cited equations define the algorithms. Widget readouts report the displayed toy arrays and settings; they are not hardware benchmarks.
- Pierre Blanchard, Desmond J. Higham, Nicholas J. Higham, 2019 arXiv preprint. Accurate Computation of the Log-Sum-Exp and Softmax Functions. Shifted exponential formulas and floating-point behavior.
- Benoit Jacob et al., 2018. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. Affine quantization, zero points, integer accumulators, and rescaling.
- ONNX Runtime contributors, quantization documentation. Quantize ONNX models. Calibration, static and dynamic quantization, and error diagnosis.
- Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer, 2022. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. Outlier-sensitive quantization and mixed-precision decomposition.
- Guangxuan Xiao et al., 2023. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. Equivalent channel rescaling between activations and weights.
- Samuel Williams, Andrew Waterman, David Patterson, 2008 technical report, published 2009. Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures. Arithmetic intensity and compute versus bandwidth bounds.
- Hugging Face Transformers contributors, cache documentation. How caching works. Cache payload dimensions and reuse during generation.
- Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, Yejin Choi, 2020. The Curious Case of Neural Text Degeneration. Temperature, probability-tail truncation, and nucleus sampling.