All tutorials Mighty Professional
Build a Game Engine ยท Animation

Vertex Animation
Textures

Bake a simulation into pixels, sample the pixels from a vertex shader, and the same animation plays on thousands of mesh instances with almost no per-instance CPU work. The technique is called Vertex Animation Textures (VAT), and it's one of the main ways crowds, swarms, shattering glass, and splashing fluid ship cheaply in current engines. The tutorial starts with one mesh and one texture, works up to the four production VAT modes (soft body, rigid body, fluid, sprite), and live-decodes the texture in your browser.

Time~45 min LevelJunior to mid graphics / tech-artist PrereqsYou can read HLSL or GLSL. You've heard of UVs and vertex shaders. HardwareKnowledge that the GPU has a vertex shader and a texture sampler.
โ—‚ Build a Game Engine Phase 8 ยท Animation Next ยท Skeletal Animation & Skinning โ–ธ

01Why a VAT

A skeletal mesh is the standard way to animate a character. Each frame, the CPU samples the character's animation clips for every bone in its hierarchy (dozens to a few hundred bones, depending on the rig), builds a palette of skinning matrices, and hands the palette to the GPU, where the vertex shader skins each vertex by blending the bones it's weighted to. It has been the default since the late 1990s and is how most player characters ship. The catch is that all of it is per-character CPU work every frame, and it grows with bone count and with every blend, IK solve and foot-placement pass layered on top. In most engines' default skeletal path each character is also at least one draw call, because its bone palette is per-instance state. Put a thousand animated characters on screen and the CPU usually runs out of frame time well before the GPU does.

A sidesteps the bone palette entirely. Bake the animation once (every vertex position at every frame) into the pixels of a texture. Ship the mesh as a static mesh. Replay the animation in the vertex shader by sampling the texture at (vertex_id, current_frame) and using the sampled value as the vertex position. There is no skeleton at runtime and no matrix palette; the only per-instance state is a time offset. With instancing, the CPU issues one draw call for the whole crowd.

What you'll have by the end

A working soft-body VAT vertex shader in HLSL and GLSL, and a clear picture of how the rigid-body, fluid, and sprite modes change it. By the end you'll know why Houdini's exporter ships a mesh plus several textures, why The Matrix Awakens's distant pedestrians are static meshes, why texture-size limits cap vertex count ร— frame count, and why the rotation channel of a rigid-body VAT is a quaternion instead of a matrix.

Three jobs VAT does well

The widget below races the CPU cost of skeletal animation against VAT as character count scales. Drag the slider; the skeletal curve is roughly linear in characters because the CPU has to evaluate one rig per character, while the VAT curve is flat because the per-vertex work runs on the GPU:

Live ยท CPU cost as crowd scales
skeletal frame cost
ยทยทยท
VAT frame cost
ยทยทยท
speedup
ยทยทยท
Skeletal cost is modeled as a flat 0.8 ยตs per bone per character per frame on one thread, lumping clip sampling, blending, pose build and palette upload together. That constant is an assumption, not a measurement: real engines vary several-fold with rig, animation graph and threading. VAT is modeled as a small fixed cost (texture binding plus one instanced draw) plus a tiny per-instance term. Only the shapes carry over to a real engine: linear versus flat.

02A short history of putting animation in a texture

The pattern of using a texture as a database that the vertex shader can read is older than the name "VAT". A quick tour, because the current shape only makes sense in context:

2007
Bryan Dudash, "Skinned Instancing," NVIDIA SDK 10.[6] An early published version of the idea. A DirectX 10 demo packs every bone matrix of every frame of every animation into one big texture, reads it in the vertex shader with per-instance data indexed by SV_InstanceID, and runs 9,547 instanced characters at about 34 fps on a GeForce 8800 GTX. Republished as "Animated Crowd Rendering," GPU Gems 3, Chapter 2[7]. The vertex shader is reading the animation out of a texture; it's still skeletal under the hood, but the texture-as-animation-database shape is here.
2007
Crysis vegetation.[8] Tiago Sousa's GPU Gems 3 chapter on Crysis vegetation describes procedural wind animation done entirely in the vertex shader, with per-vertex bending parameters (edge stiffness, per-leaf phase, overall stiffness) stored in vertex colors. It isn't a texture and it isn't a recording, but the division of labor is the same: per-vertex animation data ships with the mesh, and the vertex shader does the animating.
2017
Unreal's Pivot Painter 2.0[9] (in use by 2017). Epic's MAXScript stores each leaf's and branch's pivot point, direction vector, bounds size and inherited-motion data in textures. The shader reads those and rotates each element around its own pivot in response to wind: vertex animation, baked offline, replayed by a vertex shader. Epic's docs note it can be combined with Unreal's Vertex Animation Tool, a 3ds Max script that bakes per-vertex position offsets and normals into textures: VAT under another name.
2017
SideFX's Vertex Animation Textures export ROP.[21] Shipped in Houdini's Game Development Toolset (later renamed SideFX Labs). Luiz Kruel's March 2017 tutorial already shows all four modes the tool still has: soft body (constant topology), rigid body, fluid (changing topology), and sprite. The layout that the rest of this tutorial uses settles here: X axis is vertex, Y axis is frame, RGB is bounds-normalized position.
2020
Labs VAT 2.0.[10] The Labs rewrite of the ROP, with updated Unreal and Unity material functions. Its RBD mode stores each piece's per-frame quaternion rotation in a rotation texture. The same year, the Half-Life: Alyx Workshop Tools document Houdini VAT import for Source 2[2].
2021
Labs VAT 3.0.[4] The current SideFX node, released in August 2021. It deprecates the normal texture in favor of a rotation texture that describes both normal and tangent, which makes tangent-space normal maps work on deforming meshes; adds interframe interpolation that uses exported velocities (curved trajectories, fast-spinning rigid pieces); and exports up to 9 channels of custom data. The modes are renamed Soft-Body Deformation, Rigid-Body Dynamics, Dynamic Remeshing, and Particle Sprites.
2022
Unreal Engine 5 City Sample.[1] The Matrix Awakens tech demo (December 2021) fills a city with thousands of simulated pedestrians; Epic releases the City Sample project built from it alongside UE5's launch in April 2022. Distant pedestrians are vertex-animated static meshes baked by the AnimToTexture plugin, released with the sample and included with the engine from UE 5.1[15].
2025
OpenVAT.[5] A free, open-source Blender add-on (GPL-3.0, with its engine-side shader templates under MIT) that bakes VAT from anything that deforms in Blender's viewport, including geometry nodes, modifiers, shape keys, and simulations. It writes 8- or 16-bit PNG and half-float EXR, a JSON file with the bounds, and ships decoders for Unity, Unreal, Godot, and TikTok Effect House.
Today
VAT alongside skeletal animation, not instead of it. Engines commonly run both: skeletal animation (CPU pose evaluation, then GPU skinning in the vertex shader or in a compute-shader pass that writes a skinned vertex buffer other passes can reuse) for characters that need blending, IK and interaction, and VAT for crowds and baked simulations.

The texture-as-database idea dates to at least 2007, and Houdini's four-mode exporter (soft / rigid / fluid / sprite) to 2017. Engine support now reaches well beyond Unreal and Unity: Cocos Creator, for one, documents importing Houdini VAT bakes[19], and Blender has an open-source baker.

03The core idea: a texture is just a 2D array

Forget for a moment that textures are usually pictures. To the GPU, a texture is a 2D array of small numeric tuples (RGB or RGBA). The sampler hardware does nothing more than: given (u, v) coordinates, return the tuple at that location. There's no rule that the tuple has to be a color. It can be a position, a normal, a quaternion, or anything else four floats can carry.

The whole VAT trick is in choosing what to store and where:

That's the entire idea. The rest of this tutorial is about details: how to encode positions so the precision is acceptable, how to handle normals and tangents, what to do when topology changes per frame, how to compress per-piece rigid rotations into a quaternion, and how to keep the textures small enough to ship.

Live ยท The texture, the mesh, the read
A 16-vertex banner cloth, 32 frames long, baked to a 16 ร— 32 position texture. Move the playhead and the mesh on the right is reconstructed by sampling row frame of the texture. The shader does no math beyond a texture fetch and a remap.
What's SV_VertexID and why does the shader need to know which vertex it is?

When the GPU rasterizes a mesh, the vertex shader runs once per vertex. The shader sees the per-vertex inputs you bound (position, normal, UVs, etc.), but it has no idea which vertex it's currently shading unless you tell it.

SV_VertexID is a system-generated integer in HLSL (OpenGL GLSL spells it gl_VertexID, Vulkan GLSL gl_VertexIndex) that tells the vertex shader which vertex it is processing[17]. It's a one-line way to ask "which vertex am I?" without baking the answer into the mesh.

The catch: for an indexed draw, SV_VertexID is the index value fetched from the index buffer[17], so it names a slot in the vertex buffer as the engine stored it. A vertex shared by three triangles gets the same ID all three times, which is what VAT needs; the problem is the slot numbering itself. Importers routinely split vertices at UV seams and hard edges and reorder them for cache efficiency, so slot i is often not the vertex the bake called i. The portable answer is to bake each vertex's texture coordinate into a spare UV channel, which travels with the vertex through any reordering. Houdini's VAT exporter does exactly this with its uv2 attribute[10]; Wildlife Studios' Unity implementation reads SV_VertexID directly, which works as long as the import keeps the bake's vertex order[3].

04The texture layout, in detail

A VAT texture is a 2D array of vertex states, and most implementations lay it out the same way. The layout matters mostly because of what the GPU's bilinear filter does along each axis: blending neighboring frames is useful, blending neighboring vertices is not.

eq. 1 ยท vertex index โ†’ texture U coordinate u = ( vertex_id + 0.5 ) รท vertex_count

The + 0.5 is essential. Without it, sampling at u = i / N lands on the seam between pixels i โˆ’ 1 and i; bilinear filtering averages two unrelated vertices and the mesh stretches between neighbors. With it, the sample lands on the center of pixel i, the filter has no neighbor to blend with, and you get exact values.

For the frame coordinate, the engine drives time. The V coordinate is:

eq. 2 ยท time โ†’ texture V coordinate v = ( t ร— framerate + 0.5 ) รท frame_count

Same half-pixel offset, for the same reason. The fractional part of t ร— framerate is what gives the bilinear filter useful work: it blends between frame j and frame j+1, so the animation interpolates smoothly even though it's only sampled at the bake rate.

Step through the grid in the widget. Each square is one pixel; the texture is a literal database keyed on (vertex, frame). Hover (or tap) anywhere to see which vertex and which frame that pixel represents:

Live ยท The grid you're sampling
pixels stored
ยทยทยท
at this v
ยทยทยท
Hover or tap the grid to read off a pixel. The amber line is the playhead. Scrub it across a row boundary and a point sampler switches to the next frame's row; the blend readout is the fractional frame a bilinear sampler returns instead, v ร— frames โˆ’ 0.5 (eq. 2 inverted), because row centers sit at half-pixel offsets.
Mipmaps are not your friend

Position data has no spatial meaning between adjacent columns: vertex 5 and vertex 6 are neighbors in vertex order, not in space. A mipmapped texture would average them and produce nonsense, and lower mips would also average frames together. Author and bind VAT textures without mipmaps. A sampler can't filter one axis and not the other, so there are two workable setups: point sampling with frame interpolation done by hand (two fetches, ยง13), or bilinear filtering with U landing exactly on texel centers, so the horizontal filter weight is zero and only the frames blend[3].

05Encoding positions: the bounding-box trick

A vertex position is three floats in world (or object) space. They can be anything: (123.4, โˆ’5.6, 7.8), (0.001, 1234.5, 0.0), anything. Texture channels traditionally are not arbitrary floats. The most common pixel format, R8G8B8A8_UNORM, only holds values in [0, 1] at 8-bit precision. Even R16G16B16A16_FLOAT, which can hold raw coordinates, loses precision as magnitudes grow, so vertices far from the origin get coarse steps.

The fix is the bounding-box remap. Compute the min and max of every vertex's position over every frame of the animation. Now every position fits in a known cube. Remap each component into [0, 1] by lerping against that cube, write the result to the texture, ship the min and max as shader constants, and decode on the way out:

eq. 3 ยท the encode pixel = ( position โˆ’ bbox_min ) รท ( bbox_max โˆ’ bbox_min )

Per-component subtract and divide. The output is in [0, 1] by construction, which is exactly what the texture format wants.

eq. 4 ยท the decode (run in the vertex shader) position = pixel ร— ( bbox_max โˆ’ bbox_min ) + bbox_min

Two constants per axis is six floats, typically passed as two float4 uniforms (Houdini's VAT 3.0 instead embeds the bounds in the exported mesh's own bounding box and derives them at runtime[4]). The decode is one multiply-add per vertex.

How much precision do you actually get?

A bounding box of side length L mapped onto an 8-bit channel gives L / 255 units per quantization step. For a character standing 2 m tall, that's 7.8 mm per step in each axis, enough to be visible as a per-frame "swimming" or "popping" of vertices, especially during slow animations where the eye has time to register the steps[11]. 16-bit half-float is not fixed point: it has a 10-bit stored mantissa (11 effective bits), so the step size scales with the value. The worst case, near the top of the [0, 1] range, is a step of L / 2048, about 1 mm in the same 2 m box, shrinking toward the box minimum. In practice that's enough to make the swimming invisible; if you want uniform precision, the 16-bit fixed-point split below gives ~30 ยตm everywhere in the box.

The three common storage choices, with their tradeoffs:

Format Bits/channel Precision in a 2 m box Bytes/vertex/frame Notes
R8G8B8A8_UNORM8~7.8 mm4Smallest, visibly quantized. Acceptable for large props at a distance.
Split 8-bit (two textures)16 (two 8-bit lerps)~30 ยตm8Two RGB8 textures storing high and low bytes per axis; reassembled in the shader. Works on hardware without HDR sampling.
R16G16B16A16_FLOAT16~1 mm worst case (step scales with value)8A common default: EXR on disk, half float in VRAM. Houdini's VAT 3.0 docs recommend HDR output when the budget allows[4].
R32G32B32A32_FLOAT32~0.1 ยตm16Rarely needed for bounds-normalized positions. OpenVAT offers 32-bit EXR for storing absolute, unnormalized values[5].
BC6Hblock-compresseddepends on each 4ร—4 block (lossy)1HDR block compression, 8ร— smaller than half float. Each 4ร—4 block stores endpoints plus 3- or 4-bit per-texel indices, and a VAT block spans 4 unrelated vertices ร— 4 frames, so the error depends on vertex order and can be large. Test per asset.

The split-8-bit format deserves a closer look because it shows up on mobile and WebGL targets where float textures aren't reliably supported[3]. The trick is to take a 16-bit value and split it across two channels:

eq. 5 ยท split-byte encode value โ†’ ( high, low )

Decode is value = high ร— 256 + low. Wildlife Studios' implementation splits each normalized coordinate across two 8-bit channels (its encoder works in base 255 rather than 256, a detail of that code) and packs X-high, X-low, Y-high into one RGB8 texture and Y-low, Z-high, Z-low into a second[3]. Fiddly, but it runs on any device that can sample an 8-bit texture from the vertex shader.

The widget shows the encoding live. Pick a position with the sliders and toggle the format; it shows the encoded pixel, the decoded position, and the round-trip error. 8-bit snaps to steps of 2 m / 255, about 7.8 mm, so the error can reach about 3.9 mm per axis:

Live ยท Bounding-box encode
encoded pixel
ยทยทยท
decoded position
ยทยทยท
round-trip error
ยทยทยท
Slide the position. The decoded value is the encoded value re-expanded through the same bounding box. At 8-bit the round-trip error runs to a few millimeters in this 2 m box; at 16-bit half it stays under about 0.5 mm, and shrinks further for encoded values near 0, where half floats are denser. At this zoom both errors are smaller than a pixel, which is why 8-bit swimming shows up in close-ups and slow motion rather than everywhere.
Why include every frame in the bounding box, not just the rest pose?

The bounding box has to enclose every position the animation ever produces, including extreme deformations. If you measure the box from the rest pose and the animation later swings vertices outside it, those vertices encode at values greater than 1, which a UNORM texture clamps to 1. The result is a vertex that visibly stops at the box wall.

Houdini's VAT exporter computes the bounds over every frame (VAT 3.0 embeds them in the exported mesh's bounds[4]), and OpenVAT writes per-channel min/max to a JSON sidecar[5]. A hand-built exporter has to do the same scan.

06Decoding in the shader

The shader-side decode is small. Here it is as an HLSL vertex shader:

vat_decode.hlsl ยท vertex shader entry point
// Uniforms supplied by the engine, the same for every vertex of every
// instance using this VAT asset.
cbuffer VatConstants {
  float3 bboxMin;            // minimum corner of the position bounding box
  float3 bboxMax;            // maximum corner of the position bounding box
  float  frameCount;         // total number of frames stored in the texture
  float  bakeFramerate;      // the rate the simulation was sampled at (e.g. 30 fps)
  float  vertexCount;        // total number of mesh vertices
  float  currentTime;        // elapsed seconds since the animation began (engine clock)
  float4x4 worldViewProjection;  // object space to clip space
};

// Bind with bilinear filtering and no mipmaps. Samplers can't filter one axis
// and not the other, but uv1.x lands exactly on a texel center, so the
// horizontal weight is zero: only V (frames) blends, which interpolates
// between frames for free.
Texture2D<float4> positionTexture;
SamplerState vatSampler;

struct VertexInput {
  float3 position : POSITION;     // the static-mesh rest pose; mostly ignored at runtime
  float2 uv0      : TEXCOORD0;    // material UVs
  float2 uv1      : TEXCOORD1;    // VAT lookup: uv1.x is normalized vertex index + half-pixel
};

struct VertexOutput {
  float4 clipPosition : SV_POSITION;
  float2 materialUv   : TEXCOORD0;
};

VertexOutput VatVertexShader(VertexInput input) {
  // 1. Compute the V coordinate: which row of the texture are we sampling?
  // frameIndexFloat is fractional: e.g. 12.37 means "between frame 12 and 13".
  // Bilinear filtering on V will give us the lerp between those two rows for free.
  float frameIndexFloat = fmod(currentTime * bakeFramerate, frameCount);
  float sampleV = (frameIndexFloat + 0.5) / frameCount;

  // 2. The U coordinate was pre-baked into uv1.x by the exporter. It already
  // includes the half-pixel offset, so we use it directly.
  float2 sampleUv = float2(input.uv1.x, sampleV);

  // 3. Read the normalized position from the texture. Vertex shaders have no
  // screen-space derivatives, so the level must be explicit: mip 0.
  float3 normalizedPosition = positionTexture.SampleLevel(
      vatSampler, sampleUv, 0).rgb;

  // 4. Un-remap through the stored bounding box. This is the inverse of the
  // bake-time normalize.
  float3 objectSpacePosition = lerp(bboxMin, bboxMax, normalizedPosition);

  // 5. Transform to clip space the usual way. From here it's a normal vertex shader.
  VertexOutput output;
  output.clipPosition = mul(worldViewProjection, float4(objectSpacePosition, 1));
  output.materialUv   = input.uv0;
  return output;
}

Five lines of math: one texture sample, one lerp, and the standard projection. The shader doesn't know whether the animation was a cloth sim or a baked character cycle; in soft-body mode both are the same shape of texture and the same five lines.

Why SampleLevel, not Sample

The Sample intrinsic picks a mip level from screen-space UV derivatives, which only exist in pixel shaders, so HLSL rejects it in a vertex shader. SampleLevel takes the mip level as an explicit argument; pass 0 and you get the base level. GLSL's texture(...) does compile in a vertex shader and uses the base level there; textureLod(..., 0.0) says so explicitly. For VAT the level question is moot anyway, since the lookup coordinate is a vertex index, not a screen-space gradient.

The GLSL version

The same math in GLSL for Vulkan. The half-pixel offset on V is the same; the texture lookup uses textureLod instead of SampleLevel:

vat_decode.glsl ยท GLSL 4.5 vertex shader
#version 450

layout(location=0) in vec3 inPosition;
layout(location=1) in vec2 inUv0;
layout(location=2) in vec2 inUv1;     // VAT lookup, normalized vertex index in .x

layout(binding=0) uniform VatConstants {
  vec3  bboxMin;
  vec3  bboxMax;
  float frameCount;
  float bakeFramerate;
  float vertexCount;
  float currentTime;
  mat4  worldViewProjection;
};

layout(binding=1) uniform sampler2D positionTexture;

layout(location=0) out vec2 outUv;

void main() {
  float frameIndexFloat = mod(currentTime * bakeFramerate, frameCount);
  float sampleV         = (frameIndexFloat + 0.5) / frameCount;
  vec2  sampleUv        = vec2(inUv1.x, sampleV);

  // textureLod forces mip 0 so the GPU doesn't pick a useless mipmap.
  vec3 normalizedPosition = textureLod(positionTexture, sampleUv, 0.0).rgb;
  vec3 objectSpacePosition = mix(bboxMin, bboxMax, normalizedPosition);

  gl_Position = worldViewProjection * vec4(objectSpacePosition, 1.0);
  outUv       = inUv0;
}

07Normals: three approaches and their tradeoffs

A vertex moves; its normal usually has to move with it. A cloth waves and the lighting on each face changes. A character bends an arm and the normals around the elbow rotate with the skin. Skipping them (keeping the static mesh's rest-pose normals) makes the lighting look wrong: the geometry moves, but the shading stays frozen in the rest pose.

VAT has three ways to ship per-frame normals, in increasing order of fidelity (and cost):

1. Compress into the position texture's alpha channel

A unit normal has three components but only two degrees of freedom (it lives on the unit sphere), so a spherical projection (octahedral encoding is a common choice) turns it into two small numbers that fit in the one spare channel: the alpha of the position texture. One texture, one sample, and the normal comes along with the position.

Houdini's exporter offers this as an option that stores the normals in the position texture's alpha channel, and documents it as a lossy compression that produces "medium-quality normals"[4]. The loss shows most on glossy surfaces, where small normal errors move the highlights; matte materials hide it better.

2. A separate normal texture

Author a second texture, same vertex ร— frame layout, RGB = normal ร— 0.5 + 0.5 (the standard remap from [-1, 1] to [0, 1]). 8-bit RGB is usually enough because only the direction matters (Wildlife's mobile pipeline stores its normals this way[3]), though very smooth, glossy surfaces can show banding at 8 bits.

Two textures and two samples per vertex, but the normal decodes independently of the position. Houdini's VAT 2.0 could export this separate normal texture[10]; 3.0 deprecates it in favor of the rotation texture below[4].

3. A rotation texture with tangent

Per-frame normals are enough to light the surface but not enough to do tangent-space normal mapping. A normal map (the kind made from a high-poly bake) is defined in the tangent space of the surface, and that tangent space rotates as the surface deforms. Without the tangent, the normal map's details rotate wrong: a brick wall lit in tangent space ends up with its bricks appearing to slide as the wall flexes.

VAT 3.0 ships a rotation texture that describes both the normal and the tangent of each vertex at each frame[4], stored as a rotation (a quaternion). The shader rebuilds the per-frame tangent frame from it, so tangent-space normal maps and their highlights follow the deforming surface.

Object-space normal maps are a useful shortcut

An object-space normal map stores detail normals as absolute directions in the mesh's rest-pose frame, so it needs no tangents. On a deforming mesh those directions have to be carried along with the surface, and the per-frame VAT normal supplies that: rotate the map's normal by the rotation that takes the rest-pose normal to the current one (an approximation, since it ignores twist around the normal). OpenVAT's docs recommend combining an object-space normal map with the animated VAT normal rather than using tangent-space maps[5].

Why three options?

The options trade memory for quality. A 4096 ร— 1024 half-float position texture is already 32 MiB, and a second texture of the same size doubles that. A hero cloth gets the rotation texture; a distant crowd member gets normals-in-alpha; small debris can skip per-frame normals and compute flat ones per pixel from screen-space derivatives instead.

08Frame interpolation

A 60 fps game running a 30 fps bake samples in between frames half the time. The simplest answer is to round to the nearest frame, which produces visibly stepped animation. The standard answer is to linearly interpolate between two adjacent frames, which the GPU's bilinear filter does for free when you sample with the half-pixel-offset V coordinate from ยง4.

The math, written explicitly:

eq. 6 ยท frame blending position = (1 โˆ’ ฮฑ) ร— Pfloor + ฮฑ ร— Pceil

With bilinear filtering on V, the GPU does this for you automatically when the sample coordinate falls between two pixel centers. With point filtering on V, you read both Pfloor and Pceil by hand and lerp.

The widget plays a deliberately sparse bake (4 fps, 16 frames) so the differences are easy to see. Point sampling holds each baked frame, so the motion steps. Linear interpolation is continuous but has a corner at every baked frame, where the velocity jumps. Cubic (Catmull-Rom, four taps) passes through the same baked samples with continuous velocity. The right panel plots each mode's curve through the samples:

Live ยท Frame blending
Linear interpolation blends two adjacent frames, so its curve is a polyline through the baked samples, and the corners read as jerks when the bake is sparse relative to the motion. Cubic reads four frames per vertex (four fetches instead of one hardware-filtered fetch) and rounds the corners off. Neither recovers motion the bake never captured, and the playback rate doesn't change that: faster playback walks the same curve faster. Houdini's VAT 3.0 goes after the same corners differently, exporting velocities so the shader can follow a curved path between two frames[4].

Looping cleanly

Most baked cycles loop. With WRAP (REPEAT) addressing and a texture exactly frame_count rows tall, the bilinear filter blends the last row into row 0 as V passes the last row's center, which is the correct loop transition. Two setups break it. Extra rows below the animation (padding, or several clips stacked in one texture) make the filter blend into the wrong row, and CLAMP addressing holds the last pose for a frame and then jumps. The usual fixes are to interpolate by hand with an explicit wrap, as ยง13 does, or to append a copy of frame 0 after the last frame and stop the playhead one frame early, so the loop transition is stored in the texture itself.

Don't lerp across cuts

If the bake has cuts (say, the character jumps to a different pose between frames 30 and 31, or two clips sit back to back in one texture), interpolating across the cut produces an in-between pose that never existed. Stop the interpolation at cuts: clamp V inside each clip's rows, as Wildlife Studios does for its stacked clips[3], or point-sample assets whose frames don't correspond. Topology-changing bakes (ยง11) have this problem on every frame, since vertex i in one frame has nothing to do with vertex i in the next.

09Soft-body VAT: the simple case

Soft-body mode is what we've been describing so far. The mesh has stable topology (the same vertex IDs from start to finish), and only positions (and optionally normals) change per frame. Cloth, jiggle, banner waves, character cycles where you don't need IK, deforming foliage. The pipeline is:

  1. Bake. Houdini, Blender, or Unreal samples the simulation at the chosen rate (typically 30 fps) and writes the per-frame position to row j, column i, of the position texture. Normals (or rotations) go to a second texture. The bounding box is computed across all frames.
  2. Export the mesh. The exporter writes a static mesh with the same vertex count as the simulation, in rest pose, with the normalized vertex ID written into UV1.x. Material UVs, vertex colors, and other per-vertex attributes pass through unchanged.
  3. Sample at runtime. The shader does the ยง6 decode. The vertex shader's only job beyond a standard transform is one texture lookup and one lerp.

The widget plays back a 32-frame banner cloth. The bake is fixed, so the cloth only knows the motion it was baked with; playback speed and the position texture's precision are the runtime knobs:

Live ยท Soft-body cloth
The cloth is 16 ร— 12 vertices, baked at 30 fps for 32 frames (a loop of about 1.07 s). At 4-bit precision the vertices visibly snap between steps, most clearly at slow playback. At 8-bit the steps in this small view are under a pixel, which is why 8-bit swimming shows up on large or close-up meshes rather than small ones; 16-bit steps are far smaller again. The mesh is identical in every case; only the position texture's precision changes.

What it costs

Per instance: no draw call of its own (instances share one draw and the same texture binding), just the mesh's vertices. A 192-vertex cloth with a 192 ร— 32 half-float position texture costs 192 ร— 32 ร— 8 = 49,152 bytes (48 KiB) of texture data, shared across every cloth on screen. The GPU runs one vertex shader per vertex per instance; the per-vertex cost is one sample and one lerp on top of the standard transform.

Compared with real-time cloth (a physics simulation per instance) or cloth rigged to many bones, this is far cheaper, at the cost of every cloth playing the same baked motion. For props in the environment that's fine; for hero cloth you'd ship simulated cloth instead.

10Rigid-body VAT: when pieces rotate as units

A shattering window is a different problem. The mesh is made of fragments: discrete pieces, each rigid, each moving and rotating as a unit. Storing per-vertex positions for thousands of pieces' worth of vertices is wasteful: every vertex on a fragment carries the same translation and rotation, so most of the bits are redundant.

Rigid-body VAT (RBD) factors the data the way a physics engine would:

eq. 7 ยท rigid-body vertex transform P = Tpiece + Qpiece ยท voffset

The dot is a quaternion-vector rotation: roughly 30 floating-point ops in the standard two-cross-product form, about double the ~15-op matrix-vector product of a 3ร—3 rotation matrix. Quaternions don't win on per-vertex math; they win on storage (4 floats vs 9, or vs 12 for an affine 3ร—4) and on interpolation.

Why quaternions instead of matrices

Three reasons. First, storage: 4 floats vs 9 (or 12 for an affine row). Second, interpolation: SLERP between two quaternions produces a clean rotation; lerping two matrices produces non-rigid intermediates that have to be re-orthogonalized. Bilinear filtering of the rotation texture lerps the four components; normalize the result and you have nlerp (normalized linear interpolation), close to slerp for the small per-frame rotations of a typical bake. That holds only if the exporter keeps consecutive frames' quaternions in the same hemisphere: q and โˆ’q are the same rotation, and a sign flip between frames sends the blend through zero. Pieces that spin too far between frames defeat both; VAT 3.0's angular-velocity interpolation option exists for them[4]. Third, repair: a quaternion knocked off unit length by filtering or quantization is fixed with one normalize, while a matrix needs re-orthogonalization.

Houdini's exporter keeps the per-frame rotations in a rotation texture and ships each piece's rest pivot on the mesh: VAT 2.0 put the pivots in vertex color[10], and VAT 3.0 moves them to UV channels, with accuracy settings that choose between two 16-bit channels of raw pivots, two 16-bit channels of encoded pivots, and two 32-bit channels[4].

Live ยท Rigid-body shatter
Each piece has one column of position data (its pivot per frame) and one column of rotation data (its quaternion per frame). The mesh carries each vertex's piece lookup coordinate and the piece's rest pivot. The total data scales with pieces ร— frames, not vertices ร— frames. A 6,000-vertex shatter with 24 pieces and 60 frames stores 24 ร— 60 = 1,440 texels in each of its two textures, not 6,000 ร— 60 = 360,000.
The pivot is baked into the mesh, not the texture

Each piece's rest-pose pivot (the point around which the piece rotates) has to be exposed to the shader somehow. The common answer is to write it into spare UV channels as a per-piece constant (every vertex on piece 17 carries the same values; Houdini's VAT 3.0 uses two UV channels for this[4]). The shader then forms the offset on the fly: voffset = restPosition โˆ’ restPivot. Some exporters bake the subtraction in and store voffset as the vertex position, which saves a subtraction and gives up the rest pose.

11Fluid VAT: when topology changes per frame

A water splash or a smoke wisp doesn't have a stable mesh. Marching cubes (or whatever isosurface extractor the sim uses) produces a different vertex count and a different connectivity every frame. Vertex IDs from frame 12 are meaningless in frame 13. The whole "column = vertex" assumption collapses.

Fluid VAT handles this with two changes from the soft-body design:

  1. The shipped mesh is a pool of triangles sized for the busiest frame. A pool vertex has no identity across frames; each frame decides where it goes, and the triangles a frame doesn't need collapse out of sight.
  2. A lookup table adds one indirection. Houdini's exporter writes a lookup-table texture that maps each pool vertex, per frame, to where its data sits in the position texture[4]. The shader reads the lookup, then the position: two dependent fetches instead of one, and no assumption that vertex i means the same thing in two frames.

Material UVs change every frame too. Houdini's VAT 3.0 calls this mode Dynamic Remeshing and requires the input geometry to be UV-unwrapped on every frame in the source DCC, or generates UVs automatically[4]; Snap's Lens Studio guide lists the same mode as Fluid[12].

Memory grows quickly, because the pool is sized for the worst frame and the lookup table is often several times larger than the animation textures themselves[4]. It is still the standard way to replay a remeshing sim from a static mesh. The main alternative, streaming the sim frame by frame as a geometry cache imported from Alembic, costs far more disk space and streaming bandwidth.

Live ยท Fluid splash
A 2D stand-in for a fluid bake: a surface profile plus droplets that appear and vanish, so the geometry has no stable topology. Fluid VAT pays for that with a worst-case vertex count and a per-frame lookup-table indirection. The runtime cost is the same every frame; the bake pays for remeshing and unwrapping every frame at export time.

12Sprite VAT: just points

The fourth and simplest mode. There's no deforming surface; there are particles, each one a position. The position texture stores one column per particle, and the exported mesh holds one small card per particle. The vertex shader reads the particle's column of the position texture, moves the card there, optionally turns it to face the camera, and you have an animated particle system.

Houdini's Labs VAT names this "Particle Sprites" mode and lets you pick the card shape (square, triangle, hexagon, or custom)[4]. The textures stay modest: 2,000 particles over 200 frames is 400,000 texels, about 3.1 MiB at half float. The whole effect (sparks, snow, fireflies) ships as one draw call and one texture.

In practice, sprite VAT competes with classical particle systems where the engine evaluates particle motion at runtime. Sprite VAT loses the ability to react to gameplay (a particle can't ricochet off geometry), but wins on cost (no per-frame simulation at all, on the CPU or the GPU) and reproducibility (the bake is identical every play).

13A complete shader, with normals and frame blending

Putting ยง5 through ยง8 together: a soft-body VAT vertex shader that decodes positions, blends frames by hand (so the result doesn't depend on the sampler's filter or address mode), and decodes per-frame normals from a second texture. About 60 lines of HLSL.

vat_full.hlsl ยท soft-body with blended frames and per-frame normals
// Engine-supplied uniforms. Constant across every vertex of every instance
// using this VAT asset.
cbuffer VatConstants {
  float3   bboxMin;            // position bounding box minimum corner
  float3   bboxMax;            // position bounding box maximum corner
  float    frameCount;         // total number of frames stored
  float    bakeFramerate;      // sample rate of the bake (typically 30)
  float    currentTime;        // elapsed seconds since the animation began
  float4x4 worldViewProjection;
};

// Position texture: each pixel encodes (x, y, z) normalized into the bounding box.
// Bound with mipmaps disabled and point sampling so we control filtering manually.
Texture2D<float4> positionTexture;

// Normal texture: each pixel encodes a unit normal as (n * 0.5 + 0.5).
Texture2D<float4> normalTexture;

SamplerState pointSampler;

struct VertexInput {
  float3 restPosition : POSITION;     // rest-pose position (unused at runtime)
  float2 materialUv   : TEXCOORD0;    // UVs for diffuse / normal mapping
  float2 vatLookup    : TEXCOORD1;    // vatLookup.x = normalized vertex ID (with +0.5 baked in)
};

struct VertexOutput {
  float4 clipPosition  : SV_POSITION;
  float2 materialUv    : TEXCOORD0;
  float3 objectNormal  : TEXCOORD1;   // object space; see below
};

// Sample two adjacent frames and lerp between them by hand, so the result doesn't
// depend on the sampler's filter or address mode: frameHi wraps to row 0
// explicitly at the loop point, even if the texture has padding rows below
// the last frame.
float3 SamplePositionBlended(float vertexU, float frameIndexFloat) {
  float  frameLo    = floor(frameIndexFloat);
  float  frameHi    = fmod(frameLo + 1.0, frameCount);   // wraps cleanly at the loop point
  float  blendAlpha = frac(frameIndexFloat);

  float  vLo        = (frameLo + 0.5) / frameCount;
  float  vHi        = (frameHi + 0.5) / frameCount;

  float3 normalizedLo = positionTexture.SampleLevel(pointSampler,
                                float2(vertexU, vLo), 0).rgb;
  float3 normalizedHi = positionTexture.SampleLevel(pointSampler,
                                float2(vertexU, vHi), 0).rgb;

  float3 normalized   = lerp(normalizedLo, normalizedHi, blendAlpha);
  return lerp(bboxMin, bboxMax, normalized);
}

// Same as above but for the normal texture. Normals don't need a bounding box;
// they decode with the standard *2 - 1 inverse remap.
float3 SampleNormalBlended(float vertexU, float frameIndexFloat) {
  float frameLo    = floor(frameIndexFloat);
  float frameHi    = fmod(frameLo + 1.0, frameCount);
  float blendAlpha = frac(frameIndexFloat);

  float vLo = (frameLo + 0.5) / frameCount;
  float vHi = (frameHi + 0.5) / frameCount;

  float3 encodedLo = normalTexture.SampleLevel(pointSampler,
                          float2(vertexU, vLo), 0).rgb;
  float3 encodedHi = normalTexture.SampleLevel(pointSampler,
                          float2(vertexU, vHi), 0).rgb;

  // lerp(encoded, encoded) then *2 - 1 is the same as lerping the decoded values,
  // since both are affine. Normalize because lerped unit vectors aren't unit anymore.
  float3 encoded   = lerp(encodedLo, encodedHi, blendAlpha);
  return normalize(encoded * 2.0 - 1.0);
}

VertexOutput VatVertexShader(VertexInput input) {
  float  frameIndexFloat = fmod(currentTime * bakeFramerate, frameCount);
  float  vertexU         = input.vatLookup.x;

  float3 objectPosition  = SamplePositionBlended(vertexU, frameIndexFloat);
  float3 objectNormal    = SampleNormalBlended(vertexU, frameIndexFloat);

  VertexOutput output;
  output.clipPosition = mul(worldViewProjection, float4(objectPosition, 1));
  output.materialUv   = input.materialUv;
  output.objectNormal = objectNormal;     // pixel shader rotates it to world space
  return output;
}
What's intentionally missing

This implementation is meant to read clearly, not be every-feature complete.

14VAT vs skeletal animation: when to use which

Both approaches survive in production because they optimize for different things:

Axis Skeletal animation Vertex animation textures
CPU cost per character Grows with bone count and animation-graph complexity (clip sampling, blending, IK, palette build and upload). Linear in characters. Near zero per instance: a shared texture binding and a per-instance time offset.
Draw calls Usually at least one per character in engines' default path, since the bone palette is per-character state. Instanced skinning from a shared palette buffer is possible, as Dudash's 2007 demo shows[6]. One instanced draw per mesh and material; every instance shares the texture.
VRAM per asset The mesh plus per-clip bone tracks, which scale with bone count and clip length, not with vertex count. Depends on vertex ร— frame product. A 5,000-vertex character with 60 frames at half float is 5,000 ร— 60 ร— 8 bytes โ‰ˆ 2.3 MiB per clip.
Runtime flexibility Full: any blend tree, IK, foot placement, ragdoll. Animations compose arbitrarily. Limited: play the bake, scrub the bake, lerp between two baked frames. No procedural composition.
Authoring complexity Rig + skin weight setup. Familiar workflow for character animators. Bake-time setup in the source DCC (Houdini/Blender/Unreal). Has to be re-baked when the animation changes.
Memory growth pattern Animation data is linear in clip length and independent of vertex count. Linear in both clip length and mesh size. Long animations on dense meshes run into texture-size limits quickly.
Supports topology changes No (mesh topology is fixed; vertex weights are per-vertex). Yes, in fluid mode (ยง11), with a memory penalty for worst-case vertex count.
Supports IK / runtime modification Yes. No. The animation is what was baked.

The City Sample's crowd is the standard example: the closest pedestrians are fully rigged MetaHumans, and the more distant ones are vertex-animated static meshes generated from them[1]. Near the camera, faces, foot placement and reactions need the rig; farther out, the static meshes keep the per-character CPU cost near zero, which is what lets the city hold thousands of pedestrians.

A useful rule of thumb

If the character will ever be controlled by the player, or interact with the player, or appear in a cutscene close to the camera, ship it skeletal. If the character is one of many, far away, deterministic, and visually identical to its peers, ship it VAT. If you're not sure, ship both and switch at a distance threshold, as the City Sample does.

15The memory wall: 8K ร— 8K and what fits

VAT memory scales as vertex_count ร— frame_count ร— bytes_per_vertex, and the product grows fast. Texture dimensions are capped too. Direct3D 11-class GPUs support 16,384 ร— 16,384, but the guaranteed minimums elsewhere are lower: Vulkan requires only 4,096 (8,192 from Vulkan 1.4)[22], and Wildlife's mobile pipeline worked inside OpenGL ES 2.0's 2,048[3]. A SideFX Labs developer described 8K as the practical limit for most game engines[13]. Because exporters wrap long vertex lists onto extra rows, the cap binds on vertices ร— frames rather than on either one alone.

Concretely, a half-float position texture at 8,192 ร— 8,192 is 8,192 ร— 8,192 ร— 8 bytes = 512 MiB, far beyond any per-asset budget. Real assets are orders of magnitude smaller:

Asset Vertices Frames Format Texture size
Banner cloth (loop)19232R16G16B16A1648 KiB
Crowd pedestrian (one clip)2,04864R16G16B16A161.0 MiB
Crowd pedestrian (BC6H)2,04864BC6H128 KiB
Shattering building (RBD)128 pieces120R16G16B16A16, position + rotation240 KiB
Fluid splash (worst-case mesh)8,19296R16G16B16A166.0 MiB
Hero cloth (4-second loop)4,096120R16G16B16A163.75 MiB

The widget below lets you dial in vertex count, frame count, and format, lays the vertices out the way exporters do (wrapping onto extra rows past 8,192), and reports the texture size and whether it fits inside an 8K ร— 8K cap. However many instances play the asset, the texture is stored once:

Live ยท Memory budget calculator
texture dims
ยทยทยท
texture size
ยทยทยท
fits 8K cap
ยทยทยท
Vertices run along X up to 8,192 and wrap onto extra rows beyond that, so each frame can take several rows and the height grows with vertices ร— frames. Houdini's exporter can optionally pad textures to powers of two[4], which some older engines and compression paths required; padding a 5,000-vertex row to 8,192 would waste 39% of the memory. BC6H sizes round each dimension up to a multiple of 4.

Compression: BC6H buys you 8ร—

Half-float RGBA at 8 bytes per texel is the safe choice but not the cheap one. BC6H compresses HDR data in the 4ร—4 blocks of the BCn family at 16 bytes per block, 1 byte per texel: an 8ร— reduction in VRAM[14]. The catch is specific to VAT. Each block stores one or two pairs of endpoint colors and a 3- or 4-bit index per texel that picks a point on the line between them[14], so it works best when a block's 16 texels are similar. A photo's neighboring pixels usually are. A VAT block holds 4 vertices ร— 4 frames, and consecutive vertex indices can sit anywhere on the mesh, so the error depends on vertex order and can be far larger than on image data. Test each asset, and order vertices so that neighbors in the texture are neighbors on the mesh before counting on it.

BC6H also has no alpha channel, so normals packed into the position alpha can't ride along[14], and it needs Direct3D 11-class hardware; mobile GPUs generally use ASTC for compressed HDR data instead.

16Try it yourself

The playground below runs a JavaScript port of the VAT decode against a procedurally-generated bake. The library is exposed as MPGVat; you can adjust vertex count, frame count, playback rate, position precision, and frame interpolation mode and watch the same mesh respond. Press Run (or Ctrl+Enter / Cmd+Enter). The right pane animates the result:

โ–ธ playground.js ยท live VAT decode in your browser

Switch the precision to '8-bit' and re-run: the peak round-trip error rises from a fraction of a millimeter to a few millimeters, about a pixel at this zoom. Switch the interpolation to 'point' and the motion steps at the bake's 30 fps; raise speedX above 2 and whole baked frames get skipped between display frames.

17How Unreal does it

Unreal Engine 5 has two common VAT paths.

City Sample's two-tier crowd

The instructive part of the City Sample is how the two representations compose. The closest pedestrians are fully rigged MetaHumans driven by animation blueprints; the more distant ones are vertex-animated static meshes generated from the same characters[1]. The crowd runs on Unreal's Mass framework[1], which chooses each agent's representation by distance: a full actor near the camera, an instanced static mesh farther out. The swap holds up because the baked meshes play the same animations the rig does.

Niagara mesh particles

Niagara, Unreal's particle system, can render static-mesh particles. With a VAT material on the mesh, each particle plays the baked animation: flocks of birds, schools of fish, swarms of insects. Passing a per-particle time offset to the material (for example through Niagara's dynamic material parameters) keeps the swarm from moving in lockstep, and the whole swarm renders as instanced draws of one mesh.

18How Unity does it

Unity doesn't ship an official AnimToTexture equivalent, so the VAT story is a mosaic of community tools and engine-supplied building blocks.

At runtime, SideFX's shaders and the OpenVAT packages cover what AnimToTexture's material functions do in Unreal; the gap is in tooling and editor integration.

19Pitfalls and how to spot them

VAT failures usually look like the animation, but wrong. The common ones:

sRGB on the position texture

VAT textures encode non-color data: positions, normals, quaternions. They must be flagged as linear in the engine's texture import settings, never sRGB. This bites 8-bit formats such as PNG and TGA; float formats have no sRGB variant. If the engine applies the sRGB-to-linear conversion to a VAT texture, every encoded value comes back pushed toward 0 (0.5 returns as about 0.21), so the mesh squashes toward the minimum corner of its bounding box. The fix is one checkbox in the importer, and the symptom is unmistakable once you've seen it.

Mipmaps generated by default

Most engines generate mipmaps automatically on import, and mips on a VAT texture average unrelated vertices together. SampleLevel(..., 0) alone doesn't make that safe: texture streaming or a lower texture-quality setting can drop the top mip, and then level 0 is a downsampled level, so the mesh collapses into a smear on some machines and not others. Disable mip generation (and streaming) for VAT textures at import.

Bilinear filtering on X

Bilinear filtering is useful along Y (frames) and harmful along X (vertices), where it averages adjacent vertex columns and makes vertices slide toward their neighbors. Samplers can't filter one axis and not the other, so either point-sample and blend frames by hand (ยง13), or keep bilinear and make sure U lands exactly on texel centers with the half-texel offset (ยง4), which zeroes the horizontal filter weight.

Tangent-space normal maps that look right at rest, slide during animation

A tangent-space normal map is defined relative to each vertex's tangent frame, and under VAT deformation that frame rotates. If the shader keeps using the rest-pose tangents, the map's detail is lit as if the surface hadn't moved, and it appears to slide across the deforming surface. The two fixes are (a) ship a rotation texture and reconstruct the per-frame tangent frame (VAT 3.0's approach[4]) or (b) use an object-space normal map and rotate its normals by the per-frame VAT normal's rotation from rest, as OpenVAT recommends[5].

Wrong bounding box

If the bake's bounding box doesn't enclose every vertex of every frame, the out-of-range vertices clamp to the box wall and visibly stop. The symptom is a deformation that looks correct until a peak frame, when part of the mesh appears stuck on a plane. The fix is to recompute the bounding box across all frames and re-export; SideFX VAT and OpenVAT both do this automatically, but a hand-built exporter is easy to get wrong.

Loop seam

With CLAMP addressing, or with padding or another clip below the last frame, the bilinear filter can't blend the last frame into frame 0: the pose holds or blends into the wrong row, then jumps. The symptom is one bad frame at the loop point. Use WRAP with a texture exactly one loop tall, interpolate by hand with an explicit wrap, or append a copy of frame 0 after the last frame (ยง8).

Single-instance time on a crowd

If every instance reads the same global time uniform, every instance plays the same frame, and a crowd of pedestrians all walks in synchronized lockstep. The fix is to bind a per-instance time offset (Unreal exposes this through PerInstanceCustomData, Unity through per-instance material properties or Entities Graphics overrides) and add it to the global time before computing the V coordinate. A random offset anywhere within the clip's length breaks the lockstep without needing more bakes.

Lerp before decode, or after?

Some shaders lerp the two raw frame samples and then run the bounding-box decode; others decode each sample first and lerp the results. For positions it makes no difference: the decode is affine, so the two orderings produce identical values, and once the samples are in shader registers they're plain floats (no UNORM clamping applies to intermediates). Lerp-then-decode is still the better habit: it runs the multiply-add once instead of twice, and it matches the code path you need for data where ordering does matter, like normals and quaternions, which get renormalized after the blend, not before (ยง13).

20Where to go from here

VAT is a small, well-understood technique. Once you have the pattern in your head, the practical learning is reading other people's exporters and shader graphs to see the variations.

Read these tools

Read these references

The final exam

Five questions on the whole tutorial. If you can answer all five without scrolling back, you've got the fundamentals.

21Sources & further reading

Numbered citations refer to the superscripts above. Everything below is freely available on the open web or linked from a vendor's documentation page.

A note on originality

The prose, code, CSS, and interactive demos on this page are original writing. The bounding-box remap and the (vertex ร— frame) layout follow SideFX's VAT exporter [4]; the precursor "texture-as-bone-database" pattern is from Dudash's NVIDIA work [6][7]. The four-mode taxonomy (soft body, rigid body, fluid, sprite) is Houdini's convention [21][4]; OpenVAT [5] covers the open-source Blender path.

  1. Epic Games. City Sample Project Unreal Engine Demonstration. Unreal Engine documentation. dev.epicgames.com. The City Sample built from The Matrix Awakens: fully rigged MetaHumans for the closest crowd characters, vertex-animated static meshes generated from them for the distant ones, and MassEntity for the crowd simulation.
  2. Valve. Half-Life: Alyx Workshop Tools โ€” Modeling / Houdini Vertex Animation. Valve Developer Community. developer.valvesoftware.com. The Source 2 documentation for importing Houdini-baked VAT, including the texture conventions.
  3. Vasconcelos, L. O., with Andre Sato. (2020). Texture Animation: Applying Morphing and Vertex Animation Techniques. Wildlife Studios Tech Blog. medium.com. Mobile VAT in Unity for Tennis Clash crowds: an SV_VertexID lookup, bilinear frame blending with texel-centered U, RGB24 normal and split-position encodings, OpenGL ES 2.0's 2,048 texture limit, clamped stacked clips, and a 2,000-instance test on a Samsung Galaxy S6.
  4. SideFX. Labs Vertex Animation Textures 3.0 render node. Houdini documentation. sidefx.com. The current SideFX exporter; documents the four modes (Soft-Body Deformation, Rigid-Body Dynamics, Dynamic Remeshing, Particle Sprites), the rotation texture, the lookup table, pivot storage, interpolation and encoding options.
  5. sharpen3d. OpenVAT โ€” Vertex Animation Toolkit. openvat.org; github.com/sharpen3d/openvat. GPL-3.0 Blender-native VAT baker with MIT-licensed engine templates for Unity, Unreal, Godot, and Effect House; PNG and EXR output, JSON bounds metadata, and its guidance on object-space normal maps.
  6. Dudash, B. (2007). Skinned Instancing. NVIDIA SDK 10 whitepaper. PDF. An early published use of a texture as the bone-matrix database read from a vertex shader; the structural precursor of texture-driven crowd animation.
  7. Dudash, B. (2007). Animated Crowd Rendering. GPU Gems 3, Chapter 2. developer.nvidia.com. The republished version of the 2007 whitepaper; 9,547 instanced animated characters at about 34 fps on a GeForce 8800 GTX.
  8. Sousa, T. (2007). Vegetation Procedural Animation and Shading in Crysis. GPU Gems 3, Chapter 16. developer.nvidia.com. Per-vertex wind-bending parameters in vertex colors, animated in the vertex shader.
  9. Epic Games. Pivot Painter Tool 2.0. Unreal Engine documentation. dev.epicgames.com. Stores per-leaf and per-branch pivots, direction vectors, and bounds in textures; notes it can be combined with the Vertex Animation Tool.
  10. SideFX. Labs Vertex Animation Textures 2.0 render node. Houdini documentation. sidefx.com. The 2020 version: Soft, Rigid, Fluid, and Sprite modes, the uv2 lookup attribute, pivots in vertex color, and packed or separate normals. Superseded by 3.0.
  11. Dimitrov, S. (2021). Vertex Animation Texture (VAT). stoyan3d.wordpress.com. Walkthrough of the encoding choices including the bounding-box normalization and the 8-bit precision tradeoff.
  12. Snap Inc. Vertex Animation Textures Guide. Lens Studio documentation. developers.snap.com. Lens Studio's VAT pipeline; documents the four modes (Softbody, Rigidbody, Fluid, Sprite) and their texture outputs.
  13. SideFX. vertex animation texture limit? SideFX Forums. sidefx.com/forum. 2019 thread in which a SideFX Labs developer calls an 8K texture the limit for most game engines and suggests splitting large meshes.
  14. Microsoft Learn. BC6H Texture Block Compression. learn.microsoft.com. The HDR BCn format: RGB half-float data, 16 bytes per 4ร—4 block (1 byte per texel), endpoint pairs plus 3- or 4-bit per-texel indices, no alpha channel.
  15. Farrow, J. (2024). Animation Textures Part 2: Using the AnimToTexture plugin. Unrealcode.net. unrealcode.net. A UE 5.4 walkthrough of the plugin (an Experimental engine plugin, off by default), its vertex and bone modes, and the bone position and rotation textures it writes.
  16. Bonjour Interactive Lab. Unity3D-VATUtils. GitHub. github.com/Bonjour-Interactive-Lab. Utilities that adapt SideFX Labs' VAT 3.0 shaders for HDRP's Visual Effect Graph, covering all four modes.
  17. Microsoft Learn. Using System-Generated Values (VertexID). Direct3D 11 documentation. learn.microsoft.com. The system-generated vertex ID; for indexed draws it is the index value read from the index buffer.
  18. Verstraete, S. (2021). Vertex Animation Textures in Unreal. SideFX tutorial. sidefx.com/tutorials. Walkthrough of the Houdini โ†’ Unreal VAT workflow with the Labs 3.0 node and the Unreal material functions.
  19. Cocos Creator. Vertex Animation Texture (VAT). Cocos documentation. docs.cocos.com. Another engine's VAT implementation, useful for cross-checking the shape of the pipeline against Unity / Unreal.
  20. keijiro. HdrpVatExample. GitHub. github.com/keijiro/HdrpVatExample. HDRP Shader Graphs for Houdini VAT's soft, rigid, fluid, and sprite modes, with Visual Effect Graph support for sprite bakes.
  21. Kruel, L. (2017). Game Tools | Vertex Animation Textures. SideFX tutorial. sidefx.com/tutorials. The March 2017 VAT export ROP in Houdini's Game Development Toolset, already with soft, rigid, fluid, and sprite modes.
  22. The Khronos Group. Vulkan Specification: Required Limits. docs.vulkan.org. maxImageDimension2D: 4,096 in core Vulkan, 8,192 in Vulkan 1.4 and Roadmap 2022.

See also