Vertex Animation
Textures
Bake a simulation into pixels, sample the pixels from a vertex shader, and the same animation plays on thousands of mesh instances with almost no per-instance CPU work. The technique is called Vertex Animation Textures (VAT), and it's one of the main ways crowds, swarms, shattering glass, and splashing fluid ship cheaply in current engines. The tutorial starts with one mesh and one texture, works up to the four production VAT modes (soft body, rigid body, fluid, sprite), and live-decodes the texture in your browser.
01Why a VAT
A skeletal mesh is the standard way to animate a character. Each frame, the CPU samples the character's animation clips for every bone in its hierarchy (dozens to a few hundred bones, depending on the rig), builds a palette of skinning matrices, and hands the palette to the GPU, where the vertex shader skins each vertex by blending the bones it's weighted to. It has been the default since the late 1990s and is how most player characters ship. The catch is that all of it is per-character CPU work every frame, and it grows with bone count and with every blend, IK solve and foot-placement pass layered on top. In most engines' default skeletal path each character is also at least one draw call, because its bone palette is per-instance state. Put a thousand animated characters on screen and the CPU usually runs out of frame time well before the GPU does.
A sidesteps the bone palette entirely. Bake the animation once (every vertex position at every frame) into the pixels of a texture. Ship the mesh as a static mesh. Replay the animation in the vertex shader by sampling the texture at (vertex_id, current_frame) and using the sampled value as the vertex position. There is no skeleton at runtime and no matrix palette; the only per-instance state is a time offset. With instancing, the CPU issues one draw call for the whole crowd.
A working soft-body VAT vertex shader in HLSL and GLSL, and a clear picture of how the rigid-body, fluid, and sprite modes change it. By the end you'll know why Houdini's exporter ships a mesh plus several textures, why The Matrix Awakens's distant pedestrians are static meshes, why texture-size limits cap vertex count ร frame count, and why the rotation channel of a rigid-body VAT is a quaternion instead of a matrix.
Three jobs VAT does well
- Crowds. The driving use case. In The Matrix Awakens and the City Sample project built from it, the pedestrians nearest the camera are fully rigged MetaHumans and the more distant ones are vertex-animated static meshes generated from them[1]. The baker behind those static meshes, the AnimToTexture plugin, ships with the engine as an Experimental plugin[15].
- Baked simulations. A Houdini cloth, particle, or rigid-body sim has no rig to import. The mesh changes per frame in ways no skeleton can describe. Bake the frames into pixels and the game engine can replay the sim without knowing it was ever a sim. Half-Life: Alyx's Workshop Tools document importing Houdini VAT bakes into Source 2[2], and Wildlife Studios described the mobile version behind the crowds in Tennis Clash, with a test scene of more than 2,000 animated instances on a low-end Samsung Galaxy S6[3].
- Effects that aren't rigs. Shatters, splashes, foliage gusts, flock-of-birds Niagara emitters. Anything where the motion is a recording, not a rule. SideFX's Labs VAT 3.0 node[4] is the encoder most engine-side VAT shaders target; OpenVAT[5] is a Blender-native alternative released in 2025.
The widget below races the CPU cost of skeletal animation against VAT as character count scales. Drag the slider; the skeletal curve is roughly linear in characters because the CPU has to evaluate one rig per character, while the VAT curve is flat because the per-vertex work runs on the GPU:
02A short history of putting animation in a texture
The pattern of using a texture as a database that the vertex shader can read is older than the name "VAT". A quick tour, because the current shape only makes sense in context:
SV_InstanceID, and runs 9,547 instanced characters at about 34 fps on a GeForce 8800 GTX. Republished as "Animated Crowd Rendering," GPU Gems 3, Chapter 2[7]. The vertex shader is reading the animation out of a texture; it's still skeletal under the hood, but the texture-as-animation-database shape is here.
The texture-as-database idea dates to at least 2007, and Houdini's four-mode exporter (soft / rigid / fluid / sprite) to 2017. Engine support now reaches well beyond Unreal and Unity: Cocos Creator, for one, documents importing Houdini VAT bakes[19], and Blender has an open-source baker.
03The core idea: a texture is just a 2D array
Forget for a moment that textures are usually pictures. To the GPU, a texture is a 2D array of small numeric tuples (RGB or RGBA). The sampler hardware does nothing more than: given (u, v) coordinates, return the tuple at that location. There's no rule that the tuple has to be a color. It can be a position, a normal, a quaternion, or anything else four floats can carry.
The whole VAT trick is in choosing what to store and where:
- Each pixel = one vertex at one moment in time. Picture a grid: the X axis is the index of the vertex (0 to N-1), the Y axis is the frame of the animation (0 to F-1). The pixel at (i, j) contains the position of vertex i at frame j.
- The mesh becomes a static mesh. The triangles never change; only the positions of their corners do. You ship the mesh in some neutral pose (often the first frame of the animation, but any pose works) and let the shader move its vertices around.
- The vertex shader reads its own position out of the texture. A vertex needs to know which vertex it is to know where to sample. Houdini's exporter stores it in a spare UV channel of the mesh (its
uv2attribute), which survives whatever vertex reordering the engine's importer does[10]; the built-inSV_VertexIDsystem value works instead when the engine keeps the vertex buffer in bake order. Either way, the shader has its column coordinate. The frame coordinate comes from the engine, usually (time ร framerate) mod frame_count.
That's the entire idea. The rest of this tutorial is about details: how to encode positions so the precision is acceptable, how to handle normals and tangents, what to do when topology changes per frame, how to compress per-piece rigid rotations into a quaternion, and how to keep the textures small enough to ship.
What's SV_VertexID and why does the shader need to know which vertex it is?
When the GPU rasterizes a mesh, the vertex shader runs once per vertex. The shader sees the per-vertex inputs you bound (position, normal, UVs, etc.), but it has no idea which vertex it's currently shading unless you tell it.
SV_VertexID is a system-generated integer in HLSL (OpenGL GLSL spells it gl_VertexID, Vulkan GLSL gl_VertexIndex) that tells the vertex shader which vertex it is processing[17]. It's a one-line way to ask "which vertex am I?" without baking the answer into the mesh.
The catch: for an indexed draw, SV_VertexID is the index value fetched from the index buffer[17], so it names a slot in the vertex buffer as the engine stored it. A vertex shared by three triangles gets the same ID all three times, which is what VAT needs; the problem is the slot numbering itself. Importers routinely split vertices at UV seams and hard edges and reorder them for cache efficiency, so slot i is often not the vertex the bake called i. The portable answer is to bake each vertex's texture coordinate into a spare UV channel, which travels with the vertex through any reordering. Houdini's VAT exporter does exactly this with its uv2 attribute[10]; Wildlife Studios' Unity implementation reads SV_VertexID directly, which works as long as the import keeps the bake's vertex order[3].
04The texture layout, in detail
A VAT texture is a 2D array of vertex states, and most implementations lay it out the same way. The layout matters mostly because of what the GPU's bilinear filter does along each axis: blending neighboring frames is useful, blending neighboring vertices is not.
- Columns are vertices. A texture's X axis runs from 0 to (width โ 1). Each column holds one vertex's data across every frame. When the vertex count exceeds the maximum texture width, exporters wrap the vertices onto several rows per frame.
- Rows are frames. Each row holds one frame of the animation, with every vertex's data in order along the row. A 64-frame animation produces a texture 64 pixels tall (or 64 ร rows-per-frame when the vertices wrap).
- The mesh carries a normalized vertex ID in a spare UV channel. UV0 usually carries the material's texture mapping, so VAT puts its lookup coordinate in the next free channel (UV1 in engine numbering, which Houdini calls
uv2, or a later one if lightmaps take it). The X coordinate is the vertex index, normalized to [0, 1] with a half-pixel offset to land on pixel centers.
The + 0.5 is essential. Without it, sampling at u = i / N lands on the seam between pixels i โ 1 and i; bilinear filtering averages two unrelated vertices and the mesh stretches between neighbors. With it, the sample lands on the center of pixel i, the filter has no neighbor to blend with, and you get exact values.
For the frame coordinate, the engine drives time. The V coordinate is:
Same half-pixel offset, for the same reason. The fractional part of t ร framerate is what gives the bilinear filter useful work: it blends between frame j and frame j+1, so the animation interpolates smoothly even though it's only sampled at the bake rate.
Step through the grid in the widget. Each square is one pixel; the texture is a literal database keyed on (vertex, frame). Hover (or tap) anywhere to see which vertex and which frame that pixel represents:
Position data has no spatial meaning between adjacent columns: vertex 5 and vertex 6 are neighbors in vertex order, not in space. A mipmapped texture would average them and produce nonsense, and lower mips would also average frames together. Author and bind VAT textures without mipmaps. A sampler can't filter one axis and not the other, so there are two workable setups: point sampling with frame interpolation done by hand (two fetches, ยง13), or bilinear filtering with U landing exactly on texel centers, so the horizontal filter weight is zero and only the frames blend[3].
05Encoding positions: the bounding-box trick
A vertex position is three floats in world (or object) space. They can be anything: (123.4, โ5.6, 7.8), (0.001, 1234.5, 0.0), anything. Texture channels traditionally are not arbitrary floats. The most common pixel format, R8G8B8A8_UNORM, only holds values in [0, 1] at 8-bit precision. Even R16G16B16A16_FLOAT, which can hold raw coordinates, loses precision as magnitudes grow, so vertices far from the origin get coarse steps.
The fix is the bounding-box remap. Compute the min and max of every vertex's position over every frame of the animation. Now every position fits in a known cube. Remap each component into [0, 1] by lerping against that cube, write the result to the texture, ship the min and max as shader constants, and decode on the way out:
Per-component subtract and divide. The output is in [0, 1] by construction, which is exactly what the texture format wants.
Two constants per axis is six floats, typically passed as two float4 uniforms (Houdini's VAT 3.0 instead embeds the bounds in the exported mesh's own bounding box and derives them at runtime[4]). The decode is one multiply-add per vertex.
How much precision do you actually get?
A bounding box of side length L mapped onto an 8-bit channel gives L / 255 units per quantization step. For a character standing 2 m tall, that's 7.8 mm per step in each axis, enough to be visible as a per-frame "swimming" or "popping" of vertices, especially during slow animations where the eye has time to register the steps[11]. 16-bit half-float is not fixed point: it has a 10-bit stored mantissa (11 effective bits), so the step size scales with the value. The worst case, near the top of the [0, 1] range, is a step of L / 2048, about 1 mm in the same 2 m box, shrinking toward the box minimum. In practice that's enough to make the swimming invisible; if you want uniform precision, the 16-bit fixed-point split below gives ~30 ยตm everywhere in the box.
The three common storage choices, with their tradeoffs:
| Format | Bits/channel | Precision in a 2 m box | Bytes/vertex/frame | Notes |
|---|---|---|---|---|
| R8G8B8A8_UNORM | 8 | ~7.8 mm | 4 | Smallest, visibly quantized. Acceptable for large props at a distance. |
| Split 8-bit (two textures) | 16 (two 8-bit lerps) | ~30 ยตm | 8 | Two RGB8 textures storing high and low bytes per axis; reassembled in the shader. Works on hardware without HDR sampling. |
| R16G16B16A16_FLOAT | 16 | ~1 mm worst case (step scales with value) | 8 | A common default: EXR on disk, half float in VRAM. Houdini's VAT 3.0 docs recommend HDR output when the budget allows[4]. |
| R32G32B32A32_FLOAT | 32 | ~0.1 ยตm | 16 | Rarely needed for bounds-normalized positions. OpenVAT offers 32-bit EXR for storing absolute, unnormalized values[5]. |
| BC6H | block-compressed | depends on each 4ร4 block (lossy) | 1 | HDR block compression, 8ร smaller than half float. Each 4ร4 block stores endpoints plus 3- or 4-bit per-texel indices, and a VAT block spans 4 unrelated vertices ร 4 frames, so the error depends on vertex order and can be large. Test per asset. |
The split-8-bit format deserves a closer look because it shows up on mobile and WebGL targets where float textures aren't reliably supported[3]. The trick is to take a 16-bit value and split it across two channels:
Decode is value = high ร 256 + low. Wildlife Studios' implementation splits each normalized coordinate across two 8-bit channels (its encoder works in base 255 rather than 256, a detail of that code) and packs X-high, X-low, Y-high into one RGB8 texture and Y-low, Z-high, Z-low into a second[3]. Fiddly, but it runs on any device that can sample an 8-bit texture from the vertex shader.
The widget shows the encoding live. Pick a position with the sliders and toggle the format; it shows the encoded pixel, the decoded position, and the round-trip error. 8-bit snaps to steps of 2 m / 255, about 7.8 mm, so the error can reach about 3.9 mm per axis:
Why include every frame in the bounding box, not just the rest pose?
The bounding box has to enclose every position the animation ever produces, including extreme deformations. If you measure the box from the rest pose and the animation later swings vertices outside it, those vertices encode at values greater than 1, which a UNORM texture clamps to 1. The result is a vertex that visibly stops at the box wall.
Houdini's VAT exporter computes the bounds over every frame (VAT 3.0 embeds them in the exported mesh's bounds[4]), and OpenVAT writes per-channel min/max to a JSON sidecar[5]. A hand-built exporter has to do the same scan.
06Decoding in the shader
The shader-side decode is small. Here it is as an HLSL vertex shader:
// Uniforms supplied by the engine, the same for every vertex of every // instance using this VAT asset. cbuffer VatConstants { float3 bboxMin; // minimum corner of the position bounding box float3 bboxMax; // maximum corner of the position bounding box float frameCount; // total number of frames stored in the texture float bakeFramerate; // the rate the simulation was sampled at (e.g. 30 fps) float vertexCount; // total number of mesh vertices float currentTime; // elapsed seconds since the animation began (engine clock) float4x4 worldViewProjection; // object space to clip space }; // Bind with bilinear filtering and no mipmaps. Samplers can't filter one axis // and not the other, but uv1.x lands exactly on a texel center, so the // horizontal weight is zero: only V (frames) blends, which interpolates // between frames for free. Texture2D<float4> positionTexture; SamplerState vatSampler; struct VertexInput { float3 position : POSITION; // the static-mesh rest pose; mostly ignored at runtime float2 uv0 : TEXCOORD0; // material UVs float2 uv1 : TEXCOORD1; // VAT lookup: uv1.x is normalized vertex index + half-pixel }; struct VertexOutput { float4 clipPosition : SV_POSITION; float2 materialUv : TEXCOORD0; }; VertexOutput VatVertexShader(VertexInput input) { // 1. Compute the V coordinate: which row of the texture are we sampling? // frameIndexFloat is fractional: e.g. 12.37 means "between frame 12 and 13". // Bilinear filtering on V will give us the lerp between those two rows for free. float frameIndexFloat = fmod(currentTime * bakeFramerate, frameCount); float sampleV = (frameIndexFloat + 0.5) / frameCount; // 2. The U coordinate was pre-baked into uv1.x by the exporter. It already // includes the half-pixel offset, so we use it directly. float2 sampleUv = float2(input.uv1.x, sampleV); // 3. Read the normalized position from the texture. Vertex shaders have no // screen-space derivatives, so the level must be explicit: mip 0. float3 normalizedPosition = positionTexture.SampleLevel( vatSampler, sampleUv, 0).rgb; // 4. Un-remap through the stored bounding box. This is the inverse of the // bake-time normalize. float3 objectSpacePosition = lerp(bboxMin, bboxMax, normalizedPosition); // 5. Transform to clip space the usual way. From here it's a normal vertex shader. VertexOutput output; output.clipPosition = mul(worldViewProjection, float4(objectSpacePosition, 1)); output.materialUv = input.uv0; return output; }
Five lines of math: one texture sample, one lerp, and the standard projection. The shader doesn't know whether the animation was a cloth sim or a baked character cycle; in soft-body mode both are the same shape of texture and the same five lines.
SampleLevel, not SampleThe Sample intrinsic picks a mip level from screen-space UV derivatives, which only exist in pixel shaders, so HLSL rejects it in a vertex shader. SampleLevel takes the mip level as an explicit argument; pass 0 and you get the base level. GLSL's texture(...) does compile in a vertex shader and uses the base level there; textureLod(..., 0.0) says so explicitly. For VAT the level question is moot anyway, since the lookup coordinate is a vertex index, not a screen-space gradient.
The GLSL version
The same math in GLSL for Vulkan. The half-pixel offset on V is the same; the texture lookup uses textureLod instead of SampleLevel:
#version 450 layout(location=0) in vec3 inPosition; layout(location=1) in vec2 inUv0; layout(location=2) in vec2 inUv1; // VAT lookup, normalized vertex index in .x layout(binding=0) uniform VatConstants { vec3 bboxMin; vec3 bboxMax; float frameCount; float bakeFramerate; float vertexCount; float currentTime; mat4 worldViewProjection; }; layout(binding=1) uniform sampler2D positionTexture; layout(location=0) out vec2 outUv; void main() { float frameIndexFloat = mod(currentTime * bakeFramerate, frameCount); float sampleV = (frameIndexFloat + 0.5) / frameCount; vec2 sampleUv = vec2(inUv1.x, sampleV); // textureLod forces mip 0 so the GPU doesn't pick a useless mipmap. vec3 normalizedPosition = textureLod(positionTexture, sampleUv, 0.0).rgb; vec3 objectSpacePosition = mix(bboxMin, bboxMax, normalizedPosition); gl_Position = worldViewProjection * vec4(objectSpacePosition, 1.0); outUv = inUv0; }
07Normals: three approaches and their tradeoffs
A vertex moves; its normal usually has to move with it. A cloth waves and the lighting on each face changes. A character bends an arm and the normals around the elbow rotate with the skin. Skipping them (keeping the static mesh's rest-pose normals) makes the lighting look wrong: the geometry moves, but the shading stays frozen in the rest pose.
VAT has three ways to ship per-frame normals, in increasing order of fidelity (and cost):
1. Compress into the position texture's alpha channel
A unit normal has three components but only two degrees of freedom (it lives on the unit sphere), so a spherical projection (octahedral encoding is a common choice) turns it into two small numbers that fit in the one spare channel: the alpha of the position texture. One texture, one sample, and the normal comes along with the position.
Houdini's exporter offers this as an option that stores the normals in the position texture's alpha channel, and documents it as a lossy compression that produces "medium-quality normals"[4]. The loss shows most on glossy surfaces, where small normal errors move the highlights; matte materials hide it better.
2. A separate normal texture
Author a second texture, same vertex ร frame layout, RGB = normal ร 0.5 + 0.5 (the standard remap from [-1, 1] to [0, 1]). 8-bit RGB is usually enough because only the direction matters (Wildlife's mobile pipeline stores its normals this way[3]), though very smooth, glossy surfaces can show banding at 8 bits.
Two textures and two samples per vertex, but the normal decodes independently of the position. Houdini's VAT 2.0 could export this separate normal texture[10]; 3.0 deprecates it in favor of the rotation texture below[4].
3. A rotation texture with tangent
Per-frame normals are enough to light the surface but not enough to do tangent-space normal mapping. A normal map (the kind made from a high-poly bake) is defined in the tangent space of the surface, and that tangent space rotates as the surface deforms. Without the tangent, the normal map's details rotate wrong: a brick wall lit in tangent space ends up with its bricks appearing to slide as the wall flexes.
VAT 3.0 ships a rotation texture that describes both the normal and the tangent of each vertex at each frame[4], stored as a rotation (a quaternion). The shader rebuilds the per-frame tangent frame from it, so tangent-space normal maps and their highlights follow the deforming surface.
An object-space normal map stores detail normals as absolute directions in the mesh's rest-pose frame, so it needs no tangents. On a deforming mesh those directions have to be carried along with the surface, and the per-frame VAT normal supplies that: rotate the map's normal by the rotation that takes the rest-pose normal to the current one (an approximation, since it ignores twist around the normal). OpenVAT's docs recommend combining an object-space normal map with the animated VAT normal rather than using tangent-space maps[5].
Why three options?
The options trade memory for quality. A 4096 ร 1024 half-float position texture is already 32 MiB, and a second texture of the same size doubles that. A hero cloth gets the rotation texture; a distant crowd member gets normals-in-alpha; small debris can skip per-frame normals and compute flat ones per pixel from screen-space derivatives instead.
08Frame interpolation
A 60 fps game running a 30 fps bake samples in between frames half the time. The simplest answer is to round to the nearest frame, which produces visibly stepped animation. The standard answer is to linearly interpolate between two adjacent frames, which the GPU's bilinear filter does for free when you sample with the half-pixel-offset V coordinate from ยง4.
The math, written explicitly:
With bilinear filtering on V, the GPU does this for you automatically when the sample coordinate falls between two pixel centers. With point filtering on V, you read both Pfloor and Pceil by hand and lerp.
The widget plays a deliberately sparse bake (4 fps, 16 frames) so the differences are easy to see. Point sampling holds each baked frame, so the motion steps. Linear interpolation is continuous but has a corner at every baked frame, where the velocity jumps. Cubic (Catmull-Rom, four taps) passes through the same baked samples with continuous velocity. The right panel plots each mode's curve through the samples:
Looping cleanly
Most baked cycles loop. With WRAP (REPEAT) addressing and a texture exactly frame_count rows tall, the bilinear filter blends the last row into row 0 as V passes the last row's center, which is the correct loop transition. Two setups break it. Extra rows below the animation (padding, or several clips stacked in one texture) make the filter blend into the wrong row, and CLAMP addressing holds the last pose for a frame and then jumps. The usual fixes are to interpolate by hand with an explicit wrap, as ยง13 does, or to append a copy of frame 0 after the last frame and stop the playhead one frame early, so the loop transition is stored in the texture itself.
If the bake has cuts (say, the character jumps to a different pose between frames 30 and 31, or two clips sit back to back in one texture), interpolating across the cut produces an in-between pose that never existed. Stop the interpolation at cuts: clamp V inside each clip's rows, as Wildlife Studios does for its stacked clips[3], or point-sample assets whose frames don't correspond. Topology-changing bakes (ยง11) have this problem on every frame, since vertex i in one frame has nothing to do with vertex i in the next.
09Soft-body VAT: the simple case
Soft-body mode is what we've been describing so far. The mesh has stable topology (the same vertex IDs from start to finish), and only positions (and optionally normals) change per frame. Cloth, jiggle, banner waves, character cycles where you don't need IK, deforming foliage. The pipeline is:
- Bake. Houdini, Blender, or Unreal samples the simulation at the chosen rate (typically 30 fps) and writes the per-frame position to row j, column i, of the position texture. Normals (or rotations) go to a second texture. The bounding box is computed across all frames.
- Export the mesh. The exporter writes a static mesh with the same vertex count as the simulation, in rest pose, with the normalized vertex ID written into UV1.x. Material UVs, vertex colors, and other per-vertex attributes pass through unchanged.
- Sample at runtime. The shader does the ยง6 decode. The vertex shader's only job beyond a standard transform is one texture lookup and one lerp.
The widget plays back a 32-frame banner cloth. The bake is fixed, so the cloth only knows the motion it was baked with; playback speed and the position texture's precision are the runtime knobs:
What it costs
Per instance: no draw call of its own (instances share one draw and the same texture binding), just the mesh's vertices. A 192-vertex cloth with a 192 ร 32 half-float position texture costs 192 ร 32 ร 8 = 49,152 bytes (48 KiB) of texture data, shared across every cloth on screen. The GPU runs one vertex shader per vertex per instance; the per-vertex cost is one sample and one lerp on top of the standard transform.
Compared with real-time cloth (a physics simulation per instance) or cloth rigged to many bones, this is far cheaper, at the cost of every cloth playing the same baked motion. For props in the environment that's fine; for hero cloth you'd ship simulated cloth instead.
10Rigid-body VAT: when pieces rotate as units
A shattering window is a different problem. The mesh is made of fragments: discrete pieces, each rigid, each moving and rotating as a unit. Storing per-vertex positions for thousands of pieces' worth of vertices is wasteful: every vertex on a fragment carries the same translation and rotation, so most of the bits are redundant.
Rigid-body VAT (RBD) factors the data the way a physics engine would:
- One column per piece, not one column per vertex. Each fragment is one entry per frame. The position texture stores the piece's pivot translation. A separate rotation texture stores its quaternion orientation.
- The mesh carries the piece's lookup coordinate and its rest pivot. Every vertex on fragment 17 carries the same coordinate into column 17, plus the piece's rest-pose pivot. The vertex's offset from the pivot is its rest position minus that pivot (see the callout below for where the subtraction happens).
- The shader rotates and translates. Read the piece's rotation and translation, apply the rotation to the vertex's pivot-relative offset, add the translation. The vertex ends up in the right place; the fragment is rigid; the cost is one quaternion-rotate per vertex.
The dot is a quaternion-vector rotation: roughly 30 floating-point ops in the standard two-cross-product form, about double the ~15-op matrix-vector product of a 3ร3 rotation matrix. Quaternions don't win on per-vertex math; they win on storage (4 floats vs 9, or vs 12 for an affine 3ร4) and on interpolation.
Why quaternions instead of matrices
Three reasons. First, storage: 4 floats vs 9 (or 12 for an affine row). Second, interpolation: SLERP between two quaternions produces a clean rotation; lerping two matrices produces non-rigid intermediates that have to be re-orthogonalized. Bilinear filtering of the rotation texture lerps the four components; normalize the result and you have nlerp (normalized linear interpolation), close to slerp for the small per-frame rotations of a typical bake. That holds only if the exporter keeps consecutive frames' quaternions in the same hemisphere: q and โq are the same rotation, and a sign flip between frames sends the blend through zero. Pieces that spin too far between frames defeat both; VAT 3.0's angular-velocity interpolation option exists for them[4]. Third, repair: a quaternion knocked off unit length by filtering or quantization is fixed with one normalize, while a matrix needs re-orthogonalization.
Houdini's exporter keeps the per-frame rotations in a rotation texture and ships each piece's rest pivot on the mesh: VAT 2.0 put the pivots in vertex color[10], and VAT 3.0 moves them to UV channels, with accuracy settings that choose between two 16-bit channels of raw pivots, two 16-bit channels of encoded pivots, and two 32-bit channels[4].
Each piece's rest-pose pivot (the point around which the piece rotates) has to be exposed to the shader somehow. The common answer is to write it into spare UV channels as a per-piece constant (every vertex on piece 17 carries the same values; Houdini's VAT 3.0 uses two UV channels for this[4]). The shader then forms the offset on the fly: voffset = restPosition โ restPivot. Some exporters bake the subtraction in and store voffset as the vertex position, which saves a subtraction and gives up the rest pose.
11Fluid VAT: when topology changes per frame
A water splash or a smoke wisp doesn't have a stable mesh. Marching cubes (or whatever isosurface extractor the sim uses) produces a different vertex count and a different connectivity every frame. Vertex IDs from frame 12 are meaningless in frame 13. The whole "column = vertex" assumption collapses.
Fluid VAT handles this with two changes from the soft-body design:
- The shipped mesh is a pool of triangles sized for the busiest frame. A pool vertex has no identity across frames; each frame decides where it goes, and the triangles a frame doesn't need collapse out of sight.
- A lookup table adds one indirection. Houdini's exporter writes a lookup-table texture that maps each pool vertex, per frame, to where its data sits in the position texture[4]. The shader reads the lookup, then the position: two dependent fetches instead of one, and no assumption that vertex i means the same thing in two frames.
Material UVs change every frame too. Houdini's VAT 3.0 calls this mode Dynamic Remeshing and requires the input geometry to be UV-unwrapped on every frame in the source DCC, or generates UVs automatically[4]; Snap's Lens Studio guide lists the same mode as Fluid[12].
Memory grows quickly, because the pool is sized for the worst frame and the lookup table is often several times larger than the animation textures themselves[4]. It is still the standard way to replay a remeshing sim from a static mesh. The main alternative, streaming the sim frame by frame as a geometry cache imported from Alembic, costs far more disk space and streaming bandwidth.
12Sprite VAT: just points
The fourth and simplest mode. There's no deforming surface; there are particles, each one a position. The position texture stores one column per particle, and the exported mesh holds one small card per particle. The vertex shader reads the particle's column of the position texture, moves the card there, optionally turns it to face the camera, and you have an animated particle system.
Houdini's Labs VAT names this "Particle Sprites" mode and lets you pick the card shape (square, triangle, hexagon, or custom)[4]. The textures stay modest: 2,000 particles over 200 frames is 400,000 texels, about 3.1 MiB at half float. The whole effect (sparks, snow, fireflies) ships as one draw call and one texture.
In practice, sprite VAT competes with classical particle systems where the engine evaluates particle motion at runtime. Sprite VAT loses the ability to react to gameplay (a particle can't ricochet off geometry), but wins on cost (no per-frame simulation at all, on the CPU or the GPU) and reproducibility (the bake is identical every play).
13A complete shader, with normals and frame blending
Putting ยง5 through ยง8 together: a soft-body VAT vertex shader that decodes positions, blends frames by hand (so the result doesn't depend on the sampler's filter or address mode), and decodes per-frame normals from a second texture. About 60 lines of HLSL.
// Engine-supplied uniforms. Constant across every vertex of every instance // using this VAT asset. cbuffer VatConstants { float3 bboxMin; // position bounding box minimum corner float3 bboxMax; // position bounding box maximum corner float frameCount; // total number of frames stored float bakeFramerate; // sample rate of the bake (typically 30) float currentTime; // elapsed seconds since the animation began float4x4 worldViewProjection; }; // Position texture: each pixel encodes (x, y, z) normalized into the bounding box. // Bound with mipmaps disabled and point sampling so we control filtering manually. Texture2D<float4> positionTexture; // Normal texture: each pixel encodes a unit normal as (n * 0.5 + 0.5). Texture2D<float4> normalTexture; SamplerState pointSampler; struct VertexInput { float3 restPosition : POSITION; // rest-pose position (unused at runtime) float2 materialUv : TEXCOORD0; // UVs for diffuse / normal mapping float2 vatLookup : TEXCOORD1; // vatLookup.x = normalized vertex ID (with +0.5 baked in) }; struct VertexOutput { float4 clipPosition : SV_POSITION; float2 materialUv : TEXCOORD0; float3 objectNormal : TEXCOORD1; // object space; see below }; // Sample two adjacent frames and lerp between them by hand, so the result doesn't // depend on the sampler's filter or address mode: frameHi wraps to row 0 // explicitly at the loop point, even if the texture has padding rows below // the last frame. float3 SamplePositionBlended(float vertexU, float frameIndexFloat) { float frameLo = floor(frameIndexFloat); float frameHi = fmod(frameLo + 1.0, frameCount); // wraps cleanly at the loop point float blendAlpha = frac(frameIndexFloat); float vLo = (frameLo + 0.5) / frameCount; float vHi = (frameHi + 0.5) / frameCount; float3 normalizedLo = positionTexture.SampleLevel(pointSampler, float2(vertexU, vLo), 0).rgb; float3 normalizedHi = positionTexture.SampleLevel(pointSampler, float2(vertexU, vHi), 0).rgb; float3 normalized = lerp(normalizedLo, normalizedHi, blendAlpha); return lerp(bboxMin, bboxMax, normalized); } // Same as above but for the normal texture. Normals don't need a bounding box; // they decode with the standard *2 - 1 inverse remap. float3 SampleNormalBlended(float vertexU, float frameIndexFloat) { float frameLo = floor(frameIndexFloat); float frameHi = fmod(frameLo + 1.0, frameCount); float blendAlpha = frac(frameIndexFloat); float vLo = (frameLo + 0.5) / frameCount; float vHi = (frameHi + 0.5) / frameCount; float3 encodedLo = normalTexture.SampleLevel(pointSampler, float2(vertexU, vLo), 0).rgb; float3 encodedHi = normalTexture.SampleLevel(pointSampler, float2(vertexU, vHi), 0).rgb; // lerp(encoded, encoded) then *2 - 1 is the same as lerping the decoded values, // since both are affine. Normalize because lerped unit vectors aren't unit anymore. float3 encoded = lerp(encodedLo, encodedHi, blendAlpha); return normalize(encoded * 2.0 - 1.0); } VertexOutput VatVertexShader(VertexInput input) { float frameIndexFloat = fmod(currentTime * bakeFramerate, frameCount); float vertexU = input.vatLookup.x; float3 objectPosition = SamplePositionBlended(vertexU, frameIndexFloat); float3 objectNormal = SampleNormalBlended(vertexU, frameIndexFloat); VertexOutput output; output.clipPosition = mul(worldViewProjection, float4(objectPosition, 1)); output.materialUv = input.materialUv; output.objectNormal = objectNormal; // pixel shader rotates it to world space return output; }
This implementation is meant to read clearly, not be every-feature complete.
- No tangent-frame reconstruction. Adding tangent-space normal maps requires a third texture (or a rotation quaternion) and an extra
mulby the per-frame tangent basis. The VAT 3.0 rotation-texture path covers this. - No animation indexing. The shader plays one bake; a real implementation has an animation-state-machine equivalent that selects between several baked clips stacked vertically in the same texture, with offsets per clip.
- No per-instance time offset. To desynchronize an instanced crowd, each instance needs its own
currentTime; in practice that's a per-instance attribute the engine passes through. - No rigid-body path. Add a piece-ID UV, a rotation texture, and the ยง10 transform.
- No world-space normal. The shader passes the object-space normal through; lighting needs it rotated by the instance's world matrix (its inverse-transpose if the instance has non-uniform scale).
- No texture-size management. Every instance shares one texture, but every vertex of every instance still fetches from it, so a large VAT costs texture-cache bandwidth. Keeping textures small (fewer frames, lower precision where it holds up) is the production lever.
14VAT vs skeletal animation: when to use which
Both approaches survive in production because they optimize for different things:
| Axis | Skeletal animation | Vertex animation textures |
|---|---|---|
| CPU cost per character | Grows with bone count and animation-graph complexity (clip sampling, blending, IK, palette build and upload). Linear in characters. | Near zero per instance: a shared texture binding and a per-instance time offset. |
| Draw calls | Usually at least one per character in engines' default path, since the bone palette is per-character state. Instanced skinning from a shared palette buffer is possible, as Dudash's 2007 demo shows[6]. | One instanced draw per mesh and material; every instance shares the texture. |
| VRAM per asset | The mesh plus per-clip bone tracks, which scale with bone count and clip length, not with vertex count. | Depends on vertex ร frame product. A 5,000-vertex character with 60 frames at half float is 5,000 ร 60 ร 8 bytes โ 2.3 MiB per clip. |
| Runtime flexibility | Full: any blend tree, IK, foot placement, ragdoll. Animations compose arbitrarily. | Limited: play the bake, scrub the bake, lerp between two baked frames. No procedural composition. |
| Authoring complexity | Rig + skin weight setup. Familiar workflow for character animators. | Bake-time setup in the source DCC (Houdini/Blender/Unreal). Has to be re-baked when the animation changes. |
| Memory growth pattern | Animation data is linear in clip length and independent of vertex count. | Linear in both clip length and mesh size. Long animations on dense meshes run into texture-size limits quickly. |
| Supports topology changes | No (mesh topology is fixed; vertex weights are per-vertex). | Yes, in fluid mode (ยง11), with a memory penalty for worst-case vertex count. |
| Supports IK / runtime modification | Yes. | No. The animation is what was baked. |
The City Sample's crowd is the standard example: the closest pedestrians are fully rigged MetaHumans, and the more distant ones are vertex-animated static meshes generated from them[1]. Near the camera, faces, foot placement and reactions need the rig; farther out, the static meshes keep the per-character CPU cost near zero, which is what lets the city hold thousands of pedestrians.
If the character will ever be controlled by the player, or interact with the player, or appear in a cutscene close to the camera, ship it skeletal. If the character is one of many, far away, deterministic, and visually identical to its peers, ship it VAT. If you're not sure, ship both and switch at a distance threshold, as the City Sample does.
15The memory wall: 8K ร 8K and what fits
VAT memory scales as vertex_count ร frame_count ร bytes_per_vertex, and the product grows fast. Texture dimensions are capped too. Direct3D 11-class GPUs support 16,384 ร 16,384, but the guaranteed minimums elsewhere are lower: Vulkan requires only 4,096 (8,192 from Vulkan 1.4)[22], and Wildlife's mobile pipeline worked inside OpenGL ES 2.0's 2,048[3]. A SideFX Labs developer described 8K as the practical limit for most game engines[13]. Because exporters wrap long vertex lists onto extra rows, the cap binds on vertices ร frames rather than on either one alone.
Concretely, a half-float position texture at 8,192 ร 8,192 is 8,192 ร 8,192 ร 8 bytes = 512 MiB, far beyond any per-asset budget. Real assets are orders of magnitude smaller:
| Asset | Vertices | Frames | Format | Texture size |
|---|---|---|---|---|
| Banner cloth (loop) | 192 | 32 | R16G16B16A16 | 48 KiB |
| Crowd pedestrian (one clip) | 2,048 | 64 | R16G16B16A16 | 1.0 MiB |
| Crowd pedestrian (BC6H) | 2,048 | 64 | BC6H | 128 KiB |
| Shattering building (RBD) | 128 pieces | 120 | R16G16B16A16, position + rotation | 240 KiB |
| Fluid splash (worst-case mesh) | 8,192 | 96 | R16G16B16A16 | 6.0 MiB |
| Hero cloth (4-second loop) | 4,096 | 120 | R16G16B16A16 | 3.75 MiB |
The widget below lets you dial in vertex count, frame count, and format, lays the vertices out the way exporters do (wrapping onto extra rows past 8,192), and reports the texture size and whether it fits inside an 8K ร 8K cap. However many instances play the asset, the texture is stored once:
Compression: BC6H buys you 8ร
Half-float RGBA at 8 bytes per texel is the safe choice but not the cheap one. BC6H compresses HDR data in the 4ร4 blocks of the BCn family at 16 bytes per block, 1 byte per texel: an 8ร reduction in VRAM[14]. The catch is specific to VAT. Each block stores one or two pairs of endpoint colors and a 3- or 4-bit index per texel that picks a point on the line between them[14], so it works best when a block's 16 texels are similar. A photo's neighboring pixels usually are. A VAT block holds 4 vertices ร 4 frames, and consecutive vertex indices can sit anywhere on the mesh, so the error depends on vertex order and can be far larger than on image data. Test each asset, and order vertices so that neighbors in the texture are neighbors on the mesh before counting on it.
BC6H also has no alpha channel, so normals packed into the position alpha can't ride along[14], and it needs Direct3D 11-class hardware; mobile GPUs generally use ASTC for compressed HDR data instead.
16Try it yourself
The playground below runs a JavaScript port of the VAT decode against a procedurally-generated bake. The library is exposed as MPGVat; you can adjust vertex count, frame count, playback rate, position precision, and frame interpolation mode and watch the same mesh respond. Press Run (or Ctrl+Enter / Cmd+Enter). The right pane animates the result:
Switch the precision to '8-bit' and re-run: the peak round-trip error rises from a fraction of a millimeter to a few millimeters, about a pixel at this zoom. Switch the interpolation to 'point' and the motion steps at the bake's 30 fps; raise speedX above 2 and whole baked frames get skipped between display frames.
17How Unreal does it
Unreal Engine 5 has two common VAT paths.
- AnimToTexture plugin.[15] Epic's path for character VAT. It bakes a Skeletal Mesh plus a set of Animation Sequences into a Static Mesh and animation textures, in one of two modes. Vertex mode stores per-vertex positions and normals, as in this tutorial. Bone mode stores per-frame bone positions and rotations, with each vertex's bone weights baked alongside, so the material skins every vertex from the textures: Dudash's 2007 layout moved into a material, with one small texture set shared by every mesh on that skeleton. Material functions do the sampling and decoding. The plugin came out of the City Sample in 2022, whose distant pedestrians are vertex-animated static meshes[1]; it has been included with the engine since UE 5.1, and in 5.4 it's an Experimental plugin that's off by default[15].
- Houdini + Labs VAT 3.0.[4] The usual path for everything else: baked sims, cloth, splashes, shatters. Bake in Houdini, import the mesh and textures into Unreal, and build the material on the Unreal-side VAT material functions that SideFX Labs publishes in its own repository; the Houdini Engine plugin isn't involved[18].
City Sample's two-tier crowd
The instructive part of the City Sample is how the two representations compose. The closest pedestrians are fully rigged MetaHumans driven by animation blueprints; the more distant ones are vertex-animated static meshes generated from the same characters[1]. The crowd runs on Unreal's Mass framework[1], which chooses each agent's representation by distance: a full actor near the camera, an instanced static mesh farther out. The swap holds up because the baked meshes play the same animations the rig does.
Niagara mesh particles
Niagara, Unreal's particle system, can render static-mesh particles. With a VAT material on the mesh, each particle plays the baked animation: flocks of birds, schools of fish, swarms of insects. Passing a per-particle time offset to the material (for example through Niagara's dynamic material parameters) keeps the swarm from moving in lockstep, and the whole swarm renders as instanced draws of one mesh.
18How Unity does it
Unity doesn't ship an official AnimToTexture equivalent, so the VAT story is a mosaic of community tools and engine-supplied building blocks.
- Houdini + Labs VAT 3.0. Same as the Unreal story: the Houdini-baked mesh and textures import into Unity, and SideFX Labs' Unity shaders handle the decode[4].
- OpenVAT.[5] The open-source Blender baker ships a Unity package alongside its Unreal, Godot and Effect House decoders.
- Shader Graph. A common pattern builds the decode in Shader Graph (UV1 and time in, displaced position and normal out), so it drops into any URP or HDRP graph. Keijiro Takahashi's HdrpVatExample is a compact reference with one graph per Houdini mode, plus Visual Effect Graph support for sprite bakes[20], and Bonjour Interactive Lab's Unity3D-VATUtils adapts SideFX's VAT 3.0 shaders for HDRP's Visual Effect Graph[16].
- DOTS / Entities Graphics. Unity's data-oriented runtime for large instance counts. Entities Graphics supports per-entity material property overrides, which is where a VAT's per-instance time offset goes.
At runtime, SideFX's shaders and the OpenVAT packages cover what AnimToTexture's material functions do in Unreal; the gap is in tooling and editor integration.
19Pitfalls and how to spot them
VAT failures usually look like the animation, but wrong. The common ones:
sRGB on the position texture
VAT textures encode non-color data: positions, normals, quaternions. They must be flagged as linear in the engine's texture import settings, never sRGB. This bites 8-bit formats such as PNG and TGA; float formats have no sRGB variant. If the engine applies the sRGB-to-linear conversion to a VAT texture, every encoded value comes back pushed toward 0 (0.5 returns as about 0.21), so the mesh squashes toward the minimum corner of its bounding box. The fix is one checkbox in the importer, and the symptom is unmistakable once you've seen it.
Mipmaps generated by default
Most engines generate mipmaps automatically on import, and mips on a VAT texture average unrelated vertices together. SampleLevel(..., 0) alone doesn't make that safe: texture streaming or a lower texture-quality setting can drop the top mip, and then level 0 is a downsampled level, so the mesh collapses into a smear on some machines and not others. Disable mip generation (and streaming) for VAT textures at import.
Bilinear filtering on X
Bilinear filtering is useful along Y (frames) and harmful along X (vertices), where it averages adjacent vertex columns and makes vertices slide toward their neighbors. Samplers can't filter one axis and not the other, so either point-sample and blend frames by hand (ยง13), or keep bilinear and make sure U lands exactly on texel centers with the half-texel offset (ยง4), which zeroes the horizontal filter weight.
Tangent-space normal maps that look right at rest, slide during animation
A tangent-space normal map is defined relative to each vertex's tangent frame, and under VAT deformation that frame rotates. If the shader keeps using the rest-pose tangents, the map's detail is lit as if the surface hadn't moved, and it appears to slide across the deforming surface. The two fixes are (a) ship a rotation texture and reconstruct the per-frame tangent frame (VAT 3.0's approach[4]) or (b) use an object-space normal map and rotate its normals by the per-frame VAT normal's rotation from rest, as OpenVAT recommends[5].
Wrong bounding box
If the bake's bounding box doesn't enclose every vertex of every frame, the out-of-range vertices clamp to the box wall and visibly stop. The symptom is a deformation that looks correct until a peak frame, when part of the mesh appears stuck on a plane. The fix is to recompute the bounding box across all frames and re-export; SideFX VAT and OpenVAT both do this automatically, but a hand-built exporter is easy to get wrong.
Loop seam
With CLAMP addressing, or with padding or another clip below the last frame, the bilinear filter can't blend the last frame into frame 0: the pose holds or blends into the wrong row, then jumps. The symptom is one bad frame at the loop point. Use WRAP with a texture exactly one loop tall, interpolate by hand with an explicit wrap, or append a copy of frame 0 after the last frame (ยง8).
Single-instance time on a crowd
If every instance reads the same global time uniform, every instance plays the same frame, and a crowd of pedestrians all walks in synchronized lockstep. The fix is to bind a per-instance time offset (Unreal exposes this through PerInstanceCustomData, Unity through per-instance material properties or Entities Graphics overrides) and add it to the global time before computing the V coordinate. A random offset anywhere within the clip's length breaks the lockstep without needing more bakes.
Lerp before decode, or after?
Some shaders lerp the two raw frame samples and then run the bounding-box decode; others decode each sample first and lerp the results. For positions it makes no difference: the decode is affine, so the two orderings produce identical values, and once the samples are in shader registers they're plain floats (no UNORM clamping applies to intermediates). Lerp-then-decode is still the better habit: it runs the multiply-add once instead of twice, and it matches the code path you need for data where ordering does matter, like normals and quaternions, which get renormalized after the blend, not before (ยง13).
20Where to go from here
VAT is a small, well-understood technique. Once you have the pattern in your head, the practical learning is reading other people's exporters and shader graphs to see the variations.
Read these tools
- SideFX Labs VAT 3.0.[4] The implementation most engine-side shaders target. The Houdini node ships as an open-source HDA in the SideFXLabs repository, and reading its network covers the encoding choices above.
- OpenVAT.[5] The Blender-native baker: GPL-3.0, with MIT-licensed engine templates for Unity, Unreal, Godot and Effect House. Smaller than the SideFX equivalent.
- Unreal AnimToTexture plugin. Included with UE 5.1 and later; in UE 5.4 its source sits under
Engine/Plugins/Experimental/AnimToTexture[15]. Its material functions show both the vertex and the bone decode. - Bonjour Interactive Lab's Unity3D-VATUtils.[16] Utilities that adapt SideFX's VAT 3.0 shaders (all four modes) for HDRP's Visual Effect Graph.
Read these references
- Dudash, B. (2007). Skinned Instancing, NVIDIA SDK[6]; republished as GPU Gems 3 Chapter 2[7]. An early published version of the texture-as-animation-database idea.
- Vasconcelos, L. O. (2020). Texture Animation: Applying Morphing and Vertex Animation Techniques.[3] The Wildlife Studios mobile case study: RGB24 encodings, the OpenGL ES 2.0 texture limit, stacked clips.
- Dimitrov, S. (2021). Vertex Animation Texture (VAT).[11] A clear walkthrough of the encoding tradeoffs.
- Valve. Half-Life: Alyx Workshop Tools / Houdini Vertex Animation.[2] Source 2's documented VAT pipeline. Useful as a sanity check that the pattern is engine-portable.
The final exam
Five questions on the whole tutorial. If you can answer all five without scrolling back, you've got the fundamentals.
21Sources & further reading
Numbered citations refer to the superscripts above. Everything below is freely available on the open web or linked from a vendor's documentation page.
The prose, code, CSS, and interactive demos on this page are original writing. The bounding-box remap and the (vertex ร frame) layout follow SideFX's VAT exporter [4]; the precursor "texture-as-bone-database" pattern is from Dudash's NVIDIA work [6][7]. The four-mode taxonomy (soft body, rigid body, fluid, sprite) is Houdini's convention [21][4]; OpenVAT [5] covers the open-source Blender path.
- Epic Games. City Sample Project Unreal Engine Demonstration. Unreal Engine documentation. dev.epicgames.com. The City Sample built from The Matrix Awakens: fully rigged MetaHumans for the closest crowd characters, vertex-animated static meshes generated from them for the distant ones, and MassEntity for the crowd simulation.
- Valve. Half-Life: Alyx Workshop Tools โ Modeling / Houdini Vertex Animation. Valve Developer Community. developer.valvesoftware.com. The Source 2 documentation for importing Houdini-baked VAT, including the texture conventions.
- Vasconcelos, L. O., with Andre Sato. (2020). Texture Animation: Applying Morphing and Vertex Animation Techniques. Wildlife Studios Tech Blog. medium.com. Mobile VAT in Unity for Tennis Clash crowds: an SV_VertexID lookup, bilinear frame blending with texel-centered U, RGB24 normal and split-position encodings, OpenGL ES 2.0's 2,048 texture limit, clamped stacked clips, and a 2,000-instance test on a Samsung Galaxy S6.
- SideFX. Labs Vertex Animation Textures 3.0 render node. Houdini documentation. sidefx.com. The current SideFX exporter; documents the four modes (Soft-Body Deformation, Rigid-Body Dynamics, Dynamic Remeshing, Particle Sprites), the rotation texture, the lookup table, pivot storage, interpolation and encoding options.
- sharpen3d. OpenVAT โ Vertex Animation Toolkit. openvat.org; github.com/sharpen3d/openvat. GPL-3.0 Blender-native VAT baker with MIT-licensed engine templates for Unity, Unreal, Godot, and Effect House; PNG and EXR output, JSON bounds metadata, and its guidance on object-space normal maps.
- Dudash, B. (2007). Skinned Instancing. NVIDIA SDK 10 whitepaper. PDF. An early published use of a texture as the bone-matrix database read from a vertex shader; the structural precursor of texture-driven crowd animation.
- Dudash, B. (2007). Animated Crowd Rendering. GPU Gems 3, Chapter 2. developer.nvidia.com. The republished version of the 2007 whitepaper; 9,547 instanced animated characters at about 34 fps on a GeForce 8800 GTX.
- Sousa, T. (2007). Vegetation Procedural Animation and Shading in Crysis. GPU Gems 3, Chapter 16. developer.nvidia.com. Per-vertex wind-bending parameters in vertex colors, animated in the vertex shader.
- Epic Games. Pivot Painter Tool 2.0. Unreal Engine documentation. dev.epicgames.com. Stores per-leaf and per-branch pivots, direction vectors, and bounds in textures; notes it can be combined with the Vertex Animation Tool.
-
SideFX. Labs Vertex Animation Textures 2.0 render node. Houdini documentation. sidefx.com. The 2020 version: Soft, Rigid, Fluid, and Sprite modes, the
uv2lookup attribute, pivots in vertex color, and packed or separate normals. Superseded by 3.0. - Dimitrov, S. (2021). Vertex Animation Texture (VAT). stoyan3d.wordpress.com. Walkthrough of the encoding choices including the bounding-box normalization and the 8-bit precision tradeoff.
- Snap Inc. Vertex Animation Textures Guide. Lens Studio documentation. developers.snap.com. Lens Studio's VAT pipeline; documents the four modes (Softbody, Rigidbody, Fluid, Sprite) and their texture outputs.
- SideFX. vertex animation texture limit? SideFX Forums. sidefx.com/forum. 2019 thread in which a SideFX Labs developer calls an 8K texture the limit for most game engines and suggests splitting large meshes.
- Microsoft Learn. BC6H Texture Block Compression. learn.microsoft.com. The HDR BCn format: RGB half-float data, 16 bytes per 4ร4 block (1 byte per texel), endpoint pairs plus 3- or 4-bit per-texel indices, no alpha channel.
- Farrow, J. (2024). Animation Textures Part 2: Using the AnimToTexture plugin. Unrealcode.net. unrealcode.net. A UE 5.4 walkthrough of the plugin (an Experimental engine plugin, off by default), its vertex and bone modes, and the bone position and rotation textures it writes.
- Bonjour Interactive Lab. Unity3D-VATUtils. GitHub. github.com/Bonjour-Interactive-Lab. Utilities that adapt SideFX Labs' VAT 3.0 shaders for HDRP's Visual Effect Graph, covering all four modes.
- Microsoft Learn. Using System-Generated Values (VertexID). Direct3D 11 documentation. learn.microsoft.com. The system-generated vertex ID; for indexed draws it is the index value read from the index buffer.
- Verstraete, S. (2021). Vertex Animation Textures in Unreal. SideFX tutorial. sidefx.com/tutorials. Walkthrough of the Houdini โ Unreal VAT workflow with the Labs 3.0 node and the Unreal material functions.
- Cocos Creator. Vertex Animation Texture (VAT). Cocos documentation. docs.cocos.com. Another engine's VAT implementation, useful for cross-checking the shape of the pipeline against Unity / Unreal.
- keijiro. HdrpVatExample. GitHub. github.com/keijiro/HdrpVatExample. HDRP Shader Graphs for Houdini VAT's soft, rigid, fluid, and sprite modes, with Visual Effect Graph support for sprite bakes.
- Kruel, L. (2017). Game Tools | Vertex Animation Textures. SideFX tutorial. sidefx.com/tutorials. The March 2017 VAT export ROP in Houdini's Game Development Toolset, already with soft, rigid, fluid, and sprite modes.
-
The Khronos Group. Vulkan Specification: Required Limits. docs.vulkan.org.
maxImageDimension2D: 4,096 in core Vulkan, 8,192 in Vulkan 1.4 and Roadmap 2022.