Last updated on

Computational Photography: HDR, Burst Denoising, Super-Resolution

ENNLESPT-BR


Point a camera at a room with a bright window and a dim corner and your eye takes in both at once. The sensor cannot. It records brightness over a limited range, so a single exposure has to choose: hold the window, or hold the corner, but rarely both. Whatever the frame leaves out is gone, and the photograph becomes one answer to a question the scene asked with much more detail.

Computational photography sidesteps that choice by capturing several measurements instead of one, then combining them in software. The frames can differ in exposure, in timing, or in where they sit on the sensor. The final photo is reconstructed from that pile of evidence rather than read straight off a single sensor readout.

Three of these techniques show up on almost every phone. HDR keeps more of a scene’s brightness range. Burst denoising suppresses grain in low light. Super-resolution recovers fine detail from tiny frame-to-frame shifts. They are related, but each one solves a different problem and fails in its own way, so this page keeps them separate and then brings them back together.

The wider field is mapped well in Mobile Computational Photography: A Tour, a useful survey if you want the broader landscape of burst photography, denoising, and super-resolution on phones. This page stays with the three techniques its demos can show, so the visuals carry the explanation instead of the page turning into a survey.

If you want the filtering intuition behind frame averaging, convolution and filtering in images and signals is a good companion. And if you want the optics floor that no amount of software removes, diffraction and the Airy disk explains why a small lens blurs a point of light no matter how clever the processing gets.

HDR: Fitting a Bright Scene into One Photo

A sensor records brightness over a limited span, from the faintest tone that still rises above its own noise to the point where its pixels fill up and further light changes nothing. That span is measured in stops, where one stop is a doubling of brightness:

DR=log⁡2(Lmax⁡Lmin⁡)\mathrm{DR} = \log_2\left(\frac{L_{\max}}{L_{\min}}\right)

Lmax⁡L_{\max} is the brightest level the sensor records before it saturates, Lmin⁡L_{\min} is the dimmest level that survives above the noise floor, and the ratio is how many doublings of scene brightness fit into a single frame. A phone sensor typically covers something like 10 to 12 stops. A room with a window can span 18 or more, because the window is a small patch that is hundreds of times brighter than the wall beside it.

One exposure therefore has to spend its limited range somewhere in particular. Center the exposure low, on the shadows, and the bright end clips to flat white. Center it high, on the highlights, and the dark end collapses into black. The information disappears because a wide scene is being squeezed through a narrow window. The sensor is behaving correctly; it simply has a limited capacity.

High dynamic range, or HDR, works around the squeeze by taking several exposures at different centers and keeping the useful part of each. The demo below makes that concrete. Scene brightness runs from dark on the left to bright on the right, and the blue handle marks where the exposure sits. Drag it toward the dark end and the window and the crisp sign are the first values to clip; drag it toward the bright end and the lamp and the coat are the first to sink into black. The meter reads out how many of the five key zones the exposure loses, so the tradeoff appears as a number instead of an impression.

The HDR merge panel beside it stays closer to the full range because three captures, centered at different points, cover more of the timeline between them. For each brightness zone the merge keeps whichever capture handled it best, so a zone clipped in one frame is still intact in another. Drag the handle across the whole timeline and the merge holds up, while the single exposure keeps trading one end of the scene for the other.

Tone map

Tone Mapping Is a Separate Decision

The merged capture now holds more range than any screen can show. A display covers roughly 8 stops, well short of the scene, so the wide capture still has to be squeezed into it. That squeeze is tone mapping, and the tone-map buttons expose it directly. Balanced compresses the range gently and keeps the midtones even. High contrast pushes the curve harder, deepening shadows and brightening highlights while still staying inside the display limit.

Both choices happen after the capture is finished. That order is the part casual descriptions often blur together. The screen keeps its own limited range either way. HDR instead preserves more scene information first, then lets the render step choose how to fit that information into the display. Capture recovers the scene. Tone mapping picks the look.

Google’s HDR+ work follows the same split. Its writeup starts from the visible symptoms in a hard scene, including blur, noise, and blown highlights, then shows the cleaner result after a burst of raw frames is aligned and merged. Adobe’s Project Indigo describes the recipe in plain language: underexpose a little, capture several frames, merge them, and only then tone map the result into a natural look. Both sources treat capture and display as two problems, which is exactly the distinction the tone-map buttons are there to make.

Burst Denoising: Many Short Exposures Beat One Long One

Low light breaks a different part of the pipeline. The scene may fit inside the display range, but each frame is grainy because the sensor collected very little light. Light arrives as discrete photons, and in a short exposure the count landing on a single pixel is small enough that its random ups and downs become visible. That randomness is shot noise, and it is strongest wherever the light is weakest.

A single photon count is a poor measurement, but several independent counts can be averaged. If each of MM frames records the same scene with independent noise of spread σ\sigma, the average carries noise

σavg=σM\sigma_{\text{avg}} = \frac{\sigma}{\sqrt{M}}

so the signal-to-noise ratio improves with the square root of the frame count. Going from 1 frame to 4 cuts the noise in half, and going from 4 to 16 cuts it in half again. The gains are real but they shrink, which is why an eight-frame shot looks clearly cleaner than a four-frame shot without looking twice as clean.

There is a second reason phones reach for a burst instead of one long exposure. A long exposure gathers the same total light, but a handheld camera drifts during it, and any drift smears the scene into a blur. Many short frames can each stay sharp, and software can line them up before averaging. Short frames also clip highlights less, because each one can be exposed for the shadows rather than the brightest part of the scene. The trade is that alignment now has to work.

Alignment Is the Hard Part

The explorer below shows both sides of a burst: the stack of noisy captures along the bottom, the merged preview built from them, and a magnified patch of the dark road where the grain is easiest to see. Start with the scene still and step the burst from 2 to 4 to 8. The preview clears up each time while the individual frames stay just as grainy, so the improvement comes entirely from combining more measurements. The jump from 2 to 4 is easy to see, and the jump from 4 to 8 is smaller, which is the square-root trend made visible.

Now change the motion setting. With small motion the preview still cleans up, but the moving subject starts to leave a faint trail. With large motion that trail dominates and the merge looks less like denoising and more like a smear. The average only cancels noise when the scene stays put. Once the subject moves, every frame records it in a different place, and the average keeps all of those positions at once. That is ghosting, and it is the burst method’s characteristic failure.

Burst sizeScene motion

Burst photography therefore has a precondition that is easy to overlook: the frames have to be alignable. Registration estimates how the scene shifted between captures before averaging is trusted. Where the estimate is good, averaging behaves like a practical noise filter. Where it fails, motion becomes a trail instead of a subject. Convolution and filtering covers the basic idea of combining repeated evidence into a cleaner result, and burst photography adds registration in front of it.

Adobe’s Project Indigo gives the same rule of thumb as the equation above: aligned frames reduce noise roughly with the square root of the frame count. That is why a four-frame night shot is a clear win, an eight-frame shot is a smaller one, and a hundred-frame shot is nowhere near a hundred times better.

Super-Resolution: Recovering Detail from Subpixel Shifts

Super-resolution tackles a third problem. The exposure may be fine and the noise manageable, but the sensor still samples the world on a fixed grid of pixels. Any detail finer than the spacing between pixel centers is missed, and enlarging a single frame only interpolates between the samples already there. Interpolation can make the picture bigger, but it cannot recreate detail that was never measured.

A second frame helps only if it lands on a different part of that grid. If the camera shifts by a fraction of a pixel between captures, the new frame samples points that fall between the old ones. Register a handful of frames with slightly different subpixel offsets and the combined set of samples is denser than any single frame, which gives the reconstruction more to work with. For a one-dimensional signal, this is the whole idea in one line. Frame kk records

yk[n]=f(nΔ+δk)y_k[n] = f(n\Delta + \delta_k)

where ff is the fine scene pattern, Δ\Delta is the pixel pitch, and δk\delta_k is the subpixel shift of that frame. When every offset is zero, each frame reports the same samples and yky_k adds no information. When the offsets differ, the frames cover different points and together constrain ff more tightly. Two frames offset by half a pixel sample twice as densely along the shifted axis, and four offset by a quarter pixel sample four times as densely. Photographers call the result computational zoom, because the extra detail comes from measured samples rather than from inflating the pixels you already have.

What Subpixel Diversity Buys

The explorer puts the true detail beside a single frame and beside a reconstruction from several frames. The lettering is drawn with strokes about one sensor pixel wide, so a single frame catches only scattered pieces of each stroke. Switch the frame count between 1, 4, and 8, then change the registration case. With no shift, extra frames repeat the same sample positions, so the reconstruction plateaus at what a single grid could already capture and the letters stay ragged. With subpixel shift, the frames land at different points and the reconstruction fills in the missing pieces, so the letters resolve cleanly as more well-placed frames join in. With a moving subject, the frames disagree about where the strokes belong and the reconstruction smears instead of sharpening. The useful middle case is the one that matters in practice, because subpixel shifts smaller than one pixel are exactly what natural hand tremor supplies for free.

FramesRegistration

How Super-Resolution Differs from Sharpening

Sharpening and super-resolution both make an image look crisper, which makes them easy to confuse. Sharpening boosts edges that are already present in the samples. It raises local contrast and cannot manufacture detail the sensor never measured, no matter how strong the filter setting. Super-resolution takes the opposite route: it first gathers new samples at positions between the original pixel centers, then reconstructs a finer image from the denser set. One is a contrast operation on existing data. The other is an estimation problem over more data. That distinction is why a super-resolved photo can separate fine texture that a sharpened single frame simply does not contain, and why the technique depends on registration quality rather than on a sharpening slider.

Google’s handheld multi-frame super-resolution paper ties the method to natural hand tremor, the small motion photographers usually try to avoid. On a phone that tremor is the feature, because it supplies the subpixel offsets while the software supplies the alignment. The catch mirrors burst denoising. Lock the frames and there is no diversity to exploit. Move them too far and the offsets are too large to register. The technique lives in the narrow band between those two extremes.

What the Three Techniques Share

Side by side, the three techniques reveal a common shape. Each one collects more evidence than a single frame holds, aligns that evidence, fuses it, and then renders a final image for the screen. What differs is what each method is trying to preserve and what breaks it.

Technique What it preserves What it needs Main failure
HDR Brightness range Several exposures covering different zones Clipping when a zone is still uncovered
Burst denoising Texture without grain Static, alignable frames Ghosting when the subject moves
Super-resolution Spatial detail Subpixel diversity plus accurate registration A result that softens instead of sharpening

The failure modes follow the same pattern as the methods. A technique that leans on alignment fails when the scene moves. A technique that leans on sampling diversity fails when the frames repeat the same samples. A technique that leans on a range of exposures fails when no capture covers a zone. In each case the broken assumption is visible on screen, which is why these methods are easier to understand as a family of reconstruction problems than as a bag of camera modes.

The field reaches well beyond these three cases. MIT Media Lab’s Coded Computational Photography overview shows that exposure, aperture, motion, wavelength, and illumination can all be deliberately patterned to make later reconstruction easier. The ambition grows, but the lesson stays the same. A camera builds a measurement, and software turns that measurement into an image. Keep that model in mind and the demos become easier to read, because every control is an assumption the algorithm needs in order to work. When the assumption holds, the image improves, and when it breaks, the failure shows up almost immediately.