A shader reads like ordinary code that handles one pixel at a time, but a GPU rarely runs it that way. The hardware collects invocations into small batches, gives the whole batch a single program counter, and advances every invocation through the same instruction together. This arrangement is called SIMD, and the batch is called a wavefront or a warp.
That design explains both the throughput of a GPU and its most common performance surprise. When every invocation in a batch follows the same path, the hardware does a large amount of arithmetic for very little control overhead. When invocations need different paths, the hardware runs each path in turn and idles the lanes that took the other one.
This page follows the vertex and fragment shaders overview, which counts how often each stage runs. Here the question is how those runs are executed, why a shader branch can cost more than the branch itself, and how the same grouping shapes texture access and fragment derivatives.
Why GPUs Batch Shader Invocations
A CPU core is built to move one instruction stream as fast as possible. A large share of its area goes to branch prediction, out-of-order scheduling, and cache, all in service of a single stream. A GPU takes the opposite trade. It has thousands of narrow arithmetic units and very little control logic per unit, because the work it sees is the same program applied to millions of independent inputs.
If every shader invocation carried its own instruction fetch and decode, the control logic would cost more than the math. Batching solves that. A small group of invocations is scheduled as one unit, so a single fetch and decode feeds the whole group. The shader source still reads as scalar code operating on one invocation’s variables, but the compiler and hardware turn it into vector operations where each lane holds one invocation’s data.
What Is a Wavefront (or Warp)?
The batch has several names. NVIDIA calls it a warp, which is 32 invocations. AMD calls it a wavefront, which was 64 invocations on older architectures and is 32 on many newer ones. Vulkan and HLSL expose the same idea as a subgroup or wave, and the size can be queried at runtime rather than assumed. This page uses wavefront as the generic term.
Inside a wavefront, each invocation keeps its own registers and its own values, but all of them share a single program counter. At any moment the wavefront sits at one instruction. When that instruction is r0 = r0 * r1, every lane multiplies its own r0 by its own r1. There is no per-lane branch hardware that can send one lane to a different instruction while its neighbors stay put.
The wavefront is also the unit of scheduling. A GPU core keeps many wavefronts in flight and switches between them to hide latency, so when one wavefront waits on a texture fetch, another can issue arithmetic. A shader’s performance therefore depends not only on its instruction count but on how many wavefronts the core can keep resident.
SIMD: One Instruction, Many Lanes
SIMD stands for single instruction, multiple data, and that is the precise description of the model. One instruction is issued for the whole group, and it acts on many data lanes in parallel. On a 32-wide datapath, one issued instruction does up to 32 invocations’ worth of work. Some architectures have a datapath narrower than the wavefront, so a 64-lane wavefront issues over several cycles, but the wavefront remains the scheduling unit.
The practical consequence is a cost model. You count wavefront instructions, not source lines, and you watch how many lanes are active for each one. A branch-free shader that every lane executes is the efficient case: each instruction retires with all lanes busy, so the work per instruction is at its maximum.
This also explains why shader code often avoids if for small per-invocation choices. A ternary such as a ? b : c is usually compiled into a select instruction that every lane evaluates for itself, with no masking and no divergence. The compiler makes that choice when both sides are cheap. The next section shows what happens when it cannot.
Why GPU Threads Diverge
Divergence happens when lanes in one wavefront need to run different instructions. A branch whose condition depends on per-invocation data, such as a UV coordinate or a sampled value, is the usual cause.
Because there is one program counter, the hardware cannot run both paths at the same time. It splits the wavefront by the condition, runs one path while masking the lanes that belong to the other, then runs the other path while masking the first group. A masked lane still occupies its position in the wavefront: it consumes the issue slot and register bandwidth without producing a useful result. After the last path finishes, the masks clear and the lanes reconverge on the instruction following the branch.
The cost of a divergent branch is therefore the sum of the paths that any lane takes, measured in full wavefront instructions. A lane having a short path does not help if another lane in the same wavefront has a long one.
The explorer makes that accounting visible. The 32 cells are the lanes of one wavefront, and their uv.x increases from left to right. Dragging the threshold changes which lanes satisfy uv.x < t. Each row below is one instruction issued to the wavefront. Blue cells are active lanes on the cheap path, orange cells are active lanes on the expensive path, and pale cells are masked lanes.
At either end of the drag, every lane takes the same path. The wavefront issues only that path, all lanes stay busy, and the bar fills. In the middle, the lanes split. Both paths are now issued, and during each one the other group sits idle. Every lane does the work of one path, yet the wavefront pays for both. The bar measures the fraction of issued lane slots that did useful work, and the gap is the price of the split.
Written out, the utilization of a set of issued instructions is
where is the number of active lanes on instruction , is the wavefront width (32 or 64), and is the number of instructions issued. When every lane is active, each equals and is 1. A split wavefront drops lanes out of some of the , so falls below 1 even though grew.
Reconvergence and Predication
Reconvergence is the point after a branch where the wavefront’s lanes are active together again. The hardware tracks it with a small stack of masks: it pushes the mask for the fall-through path, runs one side, pops, runs the other side, then restores the full mask at the reconvergence point. That stack is why nested branches can nest their divergence.
Two details matter in practice. First, short branches often disappear. When both sides are cheap and free of side effects, the compiler turns the branch into predicated instructions or a select, and every lane evaluates the whole expression. The wavefront stays converged, though it now does the work of both sides. That is usually a win because it avoids masking and reconvergence overhead.
Second, uniform branches are free. If the condition is the same for every lane in the wavefront, there is nothing to split. A branch on a uniform value, one that is constant across the draw, or on a value that is constant per wavefront, follows one path with every lane active. This is why it helps to separate a per-object or per-draw decision from a per-fragment one.
Divergent Loops and Early Exits
Loops with a data-dependent trip count diverge in a way that is easy to miss. All lanes enter the loop together, but lanes that finish early are masked while the others keep iterating. The wavefront runs for as many iterations as the slowest lane needs, so the loop cost follows the maximum trip count across the wavefront rather than the average.
Ray marching through a signed distance field is a familiar case. Along one wavefront, some rays hit the surface in five steps while others need forty, and the wavefront pays for forty. Early exits, break, and discard all behave the same way. They remove lanes from future instructions, but the instruction is still issued for the lanes that remain, and masking the others does not make it cheaper.
A common mitigation is to cap the loop at a uniform maximum and let every lane finish, which bounds the worst case. Another is to sort or partition the work so that lanes in the same wavefront have similar trip counts, which lowers the average maximum.
Quads, Derivatives, and Helper Invocations
Fragment shaders add one more layer to the lane grouping. Fragment invocations are packed into 2 by 2 quads, and the four positions of a quad are always shaded together as a group. The reason is screen-space derivatives. Expressions such as dFdx, dFdy, and fwidth estimate how a value changes across the screen by comparing it with the neighboring lanes in the quad, so the hardware needs all four positions present to answer.
The consequence is that a partly covered quad still shades all four positions. The covered lanes write their results, and the uncovered ones run as helper invocations: they execute the shader, but their color and depth are discarded. A triangle that covers a single pixel at the edge of the screen still keeps four lanes busy, and a scene full of tiny or thin triangles pays that overhead everywhere.
Moving the geometry edge in the explorer changes the coverage. Blue pixels are covered and useful. Orange pixels sit in a quad with at least one covered pixel, so the hardware shades them as helpers. Empty pixels are in quads with no coverage, and those quads are never scheduled. The bar compares the number of covered pixels with the number of lane slots the quads reserve. Slide the edge toward an extreme and the covered area shrinks faster than the boundary length, so the helper share grows.
The quad rule also constrains derivatives. If dFdx is called inside a branch that only some lanes take, the neighboring lane may be masked and the derivative is undefined. That is why derivative instructions belong in uniform control flow. The same rule explains why discard does not save work: it stops a fragment from writing, but it does not free the lane, and a shader that may discard typically disables early depth testing so that more fragments run to the end.
Memory Access Across Lanes
The shared nature of a wavefront shapes memory too. A load instruction issues for all lanes at once. If the lanes request nearby addresses, the memory system merges them into one or two transactions. If the addresses are scattered, the request splits into many transactions and the wavefront waits longer for each one.
Fragment shaders usually get that merging for free, because the lanes of a wavefront cover a compact patch of screen and sample the same texture at nearby coordinates. The merging breaks down when the address is driven by a discontinuous value, such as a lookup table indexed by a per-pixel identifier, a texture chosen by a branch, or UVs warped by a nonlinear function. A distance field is a friendly case: nearby lanes read nearby texels. A random-access indirection is the opposite.
Writing SIMD-Friendly Shaders
None of this makes every branch a problem. Divergence is a cost, and it deserves attention only where the work is long or the lanes split often. A few habits keep the wavefront busy:
- Branch on values that are uniform across the wavefront. A per-draw flag, a material index from a uniform, or the same loop bound for every lane keeps the lanes converged.
- Replace small per-lane choices with branchless math.
mix,step,clamp, andmin/maxcompile to instructions that every lane runs in parallel. - Bound loops with a uniform maximum. A fixed iteration count keeps the worst case predictable and stops a wavefront from dragging along its slowest lane.
- Keep derivative calls in uniform control flow.
dFdx,dFdy, and texture sampling that relies on implicit derivatives all need neighboring lanes active. - Keep memory addresses close together. Prefer contiguous indexing and stable UVs over per-lane indirection.
These are guides, not laws. A short divergent branch over cheap operations is often faster than the branchless rewrite that evaluates both sides for every lane. The wavefront model tells you where to look first when a shader is unexpectedly slow.
Summary
GPUs execute shaders by batching invocations into wavefronts of 32 or 64 lanes that share one program counter and run SIMD-style, one instruction across every lane. That gives high throughput for the common case where all lanes do the same thing. When lanes need different control flow, the hardware serializes the paths and masks the idle lanes, so the cost becomes the sum of the paths and lane utilization drops. Fragment shaders add the 2 by 2 quad rule for derivatives, which shades helper lanes at partly covered quads. The same sharing applies to memory: coherent addresses merge into few transactions, scattered ones do not. Write shaders so the lanes agree, and the hardware does the rest.