Compute Shaders vs Fragment Shaders: When to Use Each

ENNLESPT-BR


A fragment shader is a pixel factory. It runs once for every covered screen sample during rasterization, and its output lands in an image at the resolution of the render target. A compute shader is a general thread launcher. It runs once for every thread you ask for in a dispatch, and its output can land anywhere you can write: a buffer, an image, several of them at once, or nowhere at all. That single difference in where the work comes from explains most of the practical advice about which one to reach for.

This page compares compute shaders against the raster stages, answers when a compute shader is worth its extra setup, and shows a case where the same operation is meaningfully faster as compute. The interactive vertex and fragment shader tutorial builds the raster baseline that everything here contrasts against: vertices go in, rasterization produces fragments, and a fragment shader finishes the pixels.

What a Compute Shader Is

Every shader stage we have looked at so far receives its work from geometry. A vertex shader gets one vertex at a time. A fragment shader gets one fragment at a time, and the rasterizer decides how many fragments exist. A compute shader breaks that pattern completely. It has no vertices, no triangles, and no rasterizer. The program is launched by a dispatch call that names how many workgroups to run, and each workgroup is a small block of threads. Inside the shader, the current thread looks up its own identifier and decides what to do with it.

A minimal compute shader that updates particle positions looks like this:

#version 450
layout(local_size_x = 64) in;

layout(std430, binding = 0) buffer Positions { vec4 pos[]; };
layout(std430, binding = 1) readonly buffer Velocities { vec4 vel[]; };

uniform float dt;

void main() {
    uint i = gl_GlobalInvocationID.x;
    if (i >= pos.length()) return;   // guard against the last partial group
    pos[i].xyz += vel[i].xyz * dt;
}

The local_size_x = 64 declares that each workgroup contains 64 threads. The CPU side dispatches ceil(particleCount / 64) groups, and the shader runs once per particle. Nothing here knows about the screen. If you have 500,000 particles, you can dispatch over 500,000 items directly, even though a 1080p frame only has about two million pixel samples and a small object might cover a few thousand of them.

The identifier gl_GlobalInvocationID is the compute equivalent of a fragment’s screen position or a vertex’s index. It is the thread’s coordinates inside the whole dispatch, so it is the natural thing to use as an array index. Storage buffers and images, declared with binding slots, replace the vertex attributes and texture samplers of the raster stages. Compute shaders also get shared memory, a small scratchpad visible only inside one workgroup, and barriers that let the whole group synchronize. Those two features are the reason compute can outperform a fragment pass on some problems. The threads themselves still run as SIMD lanes, so the wavefront model in how GPUs execute shaders applies here as well, and they are worth understanding before the comparison table below.

How a Fragment Shader Gets Its Work

The contrast becomes concrete when you look at where a fragment shader’s invocation count comes from. A draw call submits triangles. The rasterizer walks each triangle in screen space and generates a fragment for every sample the triangle covers. The fragment count is therefore a product of three things you may not fully control at compile time: how much of the screen the geometry covers, the render resolution, and the sample count for antialiasing or supersampling.

A fragment shader also cannot communicate with its neighbors. Each invocation sees interpolated attribute values and can sample textures, but it has no way to hand a value to the fragment next to it. If a stencil or blur needs the same neighbor value in several nearby outputs, each output fetches it separately from global memory.

Drag the orange box below to see the fragment count change while the compute count stays put. The render resolution buttons subdivide the render target without touching the buffer, and the buffer buttons change the compute dispatch without touching the screen.

Render resolution
Compute buffer

The left side changes when you drag, because fragment work is a function of coverage. The right side does not, because a compute dispatch is a function of the buffer you declare. That decoupling is the whole reason compute exists. Some workloads simply are not images.

Fragment Shader vs Compute Shader: The Core Differences

The fastest way to compare the two stages is to line up the same questions for each.

Question Fragment shader Compute shader
What is one invocation? One covered pixel sample One thread in a workgroup
Who decides the invocation count? Rasterizer, from coverage and resolution Your dispatch dimensions
Where does the work index come from? Screen-space raster position gl_GlobalInvocationID and friends
What are the inputs? Interpolated varyings, textures, uniforms Buffers, images, uniforms, IDs
Where does output go? The fragment’s own spot in a render target Any writable buffer, image, or texture index
Can one thread write many scattered locations? No, one output per thread Yes, within buffer access rules
Can threads exchange data? No Yes, through shared memory and barriers
Can threads coordinate ordering? No Yes, through barriers and atomics
What fixed-function hardware helps it? Depth test, early-z, blending, stencil, MSAA None
Typical cost driver Coverage times overdraw Dispatch size times per-thread work

Two rows carry most of the weight. The output row means a fragment shader is confined to writing its own pixel, while a compute thread can append to a list, increment a counter, or write to any index it computes. The coordination row means compute threads in the same workgroup can share data and synchronize, which is how they avoid redundant fetches. Neither of those is a small optimization detail. They change which algorithms are expressible at all.

When To Use a Compute Shader

The short version of the answer is this: use a compute shader when the work is not tied to pixels, when results must be scattered, when threads benefit from sharing data, or when the GPU needs to generate its own workload. The rest of this section unpacks each case.

The Work Is Not Shaped Like an Image

Particles, cloth vertices, rigid bodies, foliage animation, culling tests, sorting, and prefix sums are all lists of independent items. A fragment shader can only reach them by pretending the list is a texture and giving each item a pixel, which couples the list length to the render resolution and forces you to allocate targets that are mostly wasted when the list is small or truncated when it is large. Compute maps one thread to one item with no intermediate image and no resolution limit.

You Need To Scatter, Append, Or Count

Some operations produce an unpredictable number of outputs. A stream compaction pass writes the survivors of a test into a tightly packed buffer. A particle system spawns children. A histogram bumps one of several bins. These are scatter or reduction patterns. A fragment shader cannot write to an arbitrary index and cannot safely increment a shared counter, because it has no atomic access to a global counter and no ordering guarantees between invocations. A compute shader can use buffer atomics to append, or a two-pass prefix sum to compute output slots for every item before writing.

Neighbors Reuse Data

This is the case where compute is not just more general but actually faster. Consider a blur or any stencil: each output reads a small neighborhood of inputs. The same input value is read by many outputs, so a fragment shader fetches it from global memory many times. A compute workgroup can instead load one tile of input plus a one-cell halo into shared memory, synchronize, and then serve every neighborhood read from that shared scratchpad. Global memory traffic drops from outputs times stencil size to roughly one read per input value.

Stencil

Toggle between the two modes and watch the global read counter. In fragment mode every output samples its own nine values from global memory, and the heat map shows the interior cells being fetched many times over. In compute mode the workgroup loads the tile and its halo once, the barrier makes that data visible to the whole group, and every output reads from shared memory instead. Switching the stencil from 3x3 to 5x5 widens the gap sharply, because the redundant reads grow with the stencil area while the cooperative load only grows along the tile edges.

You Want To Generate Work On The GPU

Compute shaders can write into the structures that later draws read. A culling pass reads a mesh list, tests bounding volumes against the camera, appends the visible meshes to a draw list, and atomically increments the draw count. The graphics pass then issues one indirect draw that reads that count. The CPU never learns which objects were visible, which removes a round trip back to the host and lets the whole pipeline scale with GPU throughput. This style of GPU-driven rendering is difficult to express with fragment shaders because there is no render target shaped like a command buffer.

You Want To Avoid Render Target Bookkeeping

Compute shaders read and write buffers directly, which is convenient for intermediate data that is not really an image: bone matrices, light clusters, tile lists, or dense temporary structures. Working through textures instead would mean choosing a pixel format that can represent the data, allocating mip levels you do not want, and inserting layout transitions between passes. Buffers avoid that overhead and let you store structs with a natural layout.

When a Fragment Shader Is the Better Choice

Compute is not automatically faster. When the work genuinely maps one-to-one to output pixels, the raster path is usually the simpler and often the faster option, because the fixed-function hardware around the fragment shader does work you would otherwise reimplement.

A fragment shader gets coverage for free. Samples outside the triangle are never shaded, and with early depth testing, samples hidden behind closer geometry are skipped before the shader runs at all. A compute dispatch over the same screen shades everything you asked for, including the parts that would have been hidden, unless you add your own visibility test. For an opaque scene with heavy overdraw, that is a large amount of wasted work.

A fragment shader also gets blending, depth testing, stencil tests, and multisample resolve from dedicated hardware. Reproducing alpha blending, depth writes, or MSAA coverage in a compute pass means writing those rules by hand, usually with atomics or ordering that the hardware gives you for free. If an effect needs any of those, a fragment pass is the natural home.

Finally, a fragment shader needs no synchronization scaffolding. There are no workgroups, no shared memory, no barriers, and no separate memory barrier inserted before the next pass reads the result. A simple full-screen color grade, a tonemap, or a material evaluation is less code and fewer bugs as a fragment shader. A full-screen compute dispatch over exactly the visible pixels costs roughly the same per pixel as the fragment version, so with no algorithmic advantage on the table, choose the simpler stage.

Vertex Shader vs Compute Shader

The same reasoning applies to the vertex stage. A vertex shader runs once per input vertex, and the hardware that follows it does a great deal of work that compute would have to reproduce. The vertex shader outputs a clip-space position and interpolated varyings, and rasterization connects the dots: it assembles triangles, interpolates values across them, and clips anything outside the view. None of that exists in a compute dispatch.

That division is why the vertex stage is the right place for the ordinary case: take each vertex, apply the model, view, and projection transforms, write out normals and texture coordinates, and let the raster path handle the rest. As long as every vertex is transformed once and immediately, there is nothing to gain from moving the math to compute.

Compute earns its place when the vertex-stage contract is too rigid. A compute pre-pass can run skeletal skinning on a buffer of vertices, write the results into a second buffer, and let the draw read positions from that buffer instead of the original attributes. The same pattern handles morph targets, vertex animation, and GPU-driven meshlets. It also enables operations the vertex stage cannot express, such as compacting a vertex list, reusing one transformation across several draws, or generating geometry counts that a later indirect draw consumes.

The modern middle ground is the mesh shader, which exposes the generation of meshlets directly and can replace both a compute producer and a vertex consumer for some pipelines. If you want the broad stage taxonomy behind that trend, the overview of the graphics pipeline places geometry, tessellation, and compute relative to the vertex-to-fragment path.

Choosing Between Them in Practice

A short decision path handles most real cases:

  1. Does the output have one value per output pixel, with no cross-thread sharing? Use a fragment shader and keep the fixed-function pipeline benefits.
  2. Is the workload a flat list of items, such as particles, lights, or triangles? Use a compute shader and index the list by thread ID.
  3. Does the operation append, compact, histogram, or otherwise scatter into a variable number of slots? Use a compute shader with atomics or a prefix sum.
  4. Does each output read a neighborhood of inputs? Use a compute shader with shared memory when the reuse is high enough to beat the overhead.
  5. Does a later draw depend on results the GPU just produced? Use compute to build the buffers and an indirect or vertex-pulling draw to consume them.
  6. Does the effect need blending, depth, stencil, or MSAA? Prefer the fragment shader unless you have a concrete reason to reimplement those rules.

The recurring theme is that compute buys you freedom and costs you convenience. Take the freedom when the algorithm needs it, and take the convenience when it does not.

A Worked Example: A Separable Blur

A blur is a nice test case because both stages can do it, and the difference is purely in the memory strategy. The fragment version runs two passes, one horizontal and one vertical, and each output fetches the nine samples it needs from the source texture:

// Fragment shader, run once per output pixel for each of the two passes.
uniform sampler2D src;
uniform vec2 direction;   // (1, 0) or (0, 1), scaled by the texel size
in vec2 uv;
out vec4 fragColor;

void main() {
    vec4 sum = vec4(0.0);
    for (int i = -4; i <= 4; i++) {
        sum += texture(src, uv + direction * float(i));
    }
    fragColor = sum / 9.0;
}

Each output performs nine texture fetches, so the whole blur costs eighteen fetches per pixel across the two passes. Neighboring outputs fetch almost exactly the same values as each other, and the hardware cache absorbs some of that, but the shader still issues every fetch.

The compute version keeps one workgroup’s input in shared memory. A workgroup of 16x16 threads handles a 16x16 tile of output. It loads the tile plus four columns of halo on each side, which is what the horizontal pass needs, and then every output reads its nine samples from the shared tile instead of from the texture:

#version 450
layout(local_size_x = 16, local_size_y = 16) in;

layout(rgba16f, binding = 0) uniform readonly image2D src;
layout(rgba16f, binding = 1) uniform writeonly image2D dst;

const int RADIUS = 4;
const int TILE = 16;
const int SHARED_WIDTH = TILE + 2 * RADIUS;

shared vec4 tile[TILE][SHARED_WIDTH];

void main() {
    ivec2 local = ivec2(gl_LocalInvocationID.xy);
    ivec2 interior = ivec2(gl_WorkGroupID.xy) * TILE + local;
    ivec2 size = imageSize(src);

    // Every thread loads the interior column it owns.
    bool inside = interior.x < size.x && interior.y < size.y;
    tile[local.y][local.x + RADIUS] = inside ? imageLoad(src, interior) : vec4(0.0);

    // Threads on the left edge load the left halo columns.
    if (local.x < RADIUS) {
        ivec2 left = ivec2(interior.x - RADIUS, interior.y);
        tile[local.y][local.x] =
            (left.x >= 0 && interior.y < size.y) ? imageLoad(src, left) : vec4(0.0);
    }

    // Threads on the right edge load the right halo columns.
    if (local.x >= TILE - RADIUS) {
        ivec2 right = ivec2(interior.x + RADIUS, interior.y);
        tile[local.y][local.x + 2 * RADIUS] =
            (right.x < size.x && interior.y < size.y) ? imageLoad(src, right) : vec4(0.0);
    }

    // Wait until the whole workgroup has finished loading.
    barrier();

    vec4 sum = vec4(0.0);
    for (int i = -RADIUS; i <= RADIUS; i++) {
        sum += tile[local.y][local.x + RADIUS + i];
    }

    if (inside) {
        imageStore(dst, interior, sum / float(2 * RADIUS + 1));
    }
}

The mapping between the two coordinate spaces is the part worth reading slowly. Shared column local.x + RADIUS holds the thread’s own interior texel, so local.x alone is four texels to the left. The halo loads fill the outer four columns on each side, and after barrier() the group’s shared memory holds a continuous strip of input that the whole 16x16 tile can sample. Global fetches per workgroup drop from 16 * 16 * 9 to 16 * (16 + 8), and the savings only appear because the threads cooperate.

Two details keep this correct. The barrier() synchronizes invocations within one workgroup only; it says nothing about other workgroups, which is why the halo must be loaded rather than read from a neighbor’s tile. And the bounds checks matter because the last workgroup in each direction can run past the image edges.

The vertical pass is the same idea with the roles of x and y swapped. With a buffer instead of an image, remember to insert glMemoryBarrier with the storage or image access bit between a compute dispatch and a later pass that consumes its output. The GPU does not order memory writes for you across invocations, so the barrier is what makes the result visible. This example is deliberately small: production blurs vectorize the loads, use textureGather where available, and often handle both axes in one pass with two shared buffers.

Edge Cases and Common Mistakes

A few pitfalls show up again and again when moving work between the two stages.

Dispatching a size that does not divide evenly by the workgroup size. The final workgroup still runs a full complement of threads, so every compute shader needs a bounds guard like the if (i >= pos.length()) return; above. Forgetting it causes writes past the end of the buffer and reads of shifted data.

Treating barrier() as a global sync. It synchronizes a workgroup and nothing more. Cross-workgroup communication requires separate dispatches or atomics, never a shared-memory exchange.

Overlooking the memory barrier between passes. Image and buffer writes from compute are not visible to later stages until a memory barrier with the correct access bits is issued. Missing it produces results that look correct on one GPU and corrupt on another.

Assuming an invocation order. Compute threads run in any order and may be batched however the hardware likes. Counters need atomics, and any algorithm that depends on thread zero finishing first needs a real synchronization primitive.

Reading and writing the same resource in one dispatch without care. Overlapping reads and writes are undefined unless you separate them by dispatch or use atomics. Ping-pong buffers avoid the issue.

Choosing a workgroup size casually. Too small wastes scheduling slots, and too large can hurt occupancy when the shader uses a lot of registers or shared memory. Sizes like 64 or 128 threads are common starting points.

Expecting compute to be faster by default. For a full-screen pass with the same per-pixel work, a compute dispatch is roughly on par with a fragment shader, and it gives up early depth rejection and fixed-function blending. Compute wins when the work is not pixel-shaped, when it needs reuse or scatter, or when it lets you skip pixels a fragment pass would have shaded.

Summary

Fragment shaders and compute shaders differ in where their work comes from. A fragment shader’s invocation count is set by rasterized coverage and resolution, its output is pinned to its own pixel, and it shades each sample in isolation. A compute shader’s invocation count is set by a dispatch you declare, its output can go anywhere, and its threads can share data and synchronize inside a workgroup.

Reach for compute when the workload is a list rather than an image, when you need to scatter or append, when neighbor reuse would otherwise mean repeated global fetches, or when one pass must generate work for another. Stay with the fragment stage when the work maps cleanly to pixels and you can lean on coverage tests, early depth rejection, and blending. The vertex stage follows the same logic: keep it for the ordinary transform-and-rasterize case, and move to compute only when the vertex contract is too rigid.

The graphics pipeline overview covers the raster stages that all of this compares against, and a good next step after this page is to look at procedural noise functions, which are exactly the kind of input that compute passes are often built to generate.