Skip to content
scroll to zoom · drag to pan

The transform benchmark

Measured 2026-08-17 against https://api.nitida.gofuture.space and https://8ok.uk, tenant demo, on the five documentation photographs — five different chromatic characters, so no conclusion here rests on one kind of image. Their licences are at the bottom.

The headline: a stored ladder makes /t/ compress twice

Section titled “The headline: a stored ladder makes /t/ compress twice”

/t/ takes a documented shortcut: when a stored variant already covers the requested size, the server decodes that instead of the master, because decoding a 640px sm for a 320px thumb is ~9× cheaper than decoding a 3840px xl. It is a real and sensible optimisation — and it means the bytes handed to the encoder have already been through a WebP pass.

So: the same photograph, uploaded twice. Once with the full ladder (["original","thumb","sm","md","lg","xl"]), once with ["original"] and nothing else. Same request on both: format=webp,quality=75,width=800.

The shortcut fires. On all 5 photographs the laddered asset answered with X-Transform-Source: md — the 1280 px WebP — and the original-only asset with X-Transform-Source: original. That is the difference under test, and nothing else: the two assets are the same pixels (the upload deduplicates by sha256, so the second copy carries a few bytes appended after the JPEG EOI marker, which every decoder ignores).

photo ladder → source bytes PSNR vs master SSIM vs master original-only → source bytes PSNR SSIM
autumn md 103.3 KB 30.76 dB 0.9455 original 112.7 KB 33.77 dB 0.9656
glacier md 61.9 KB 34.06 dB 0.9335 original 67.7 KB 35.73 dB 0.9511
night md 11.7 KB 38.35 dB 0.9405 original 12.7 KB 39.49 dB 0.9495
snow md 27.1 KB 35.38 dB 0.9834 original 29.3 KB 38.00 dB 0.9892
spices md 51.3 KB 34.74 dB 0.9304 original 56.6 KB 36.41 dB 0.9499

Both sides are compared against the master, never against each other — two outputs measured against one another tell you they differ, not which is worse.

Yes, it re-compresses, and it is a straight loss — 5 times out of 5. The doubly-compressed file is smaller (by 7.3 – 9.4 %) and worse on both metrics: up to −3.01 dB PSNR and −0.0200 SSIM.

That combination is the signature of generation loss, not of a saving. The file shrank because the first pass had already thrown the fine detail away, so the second encoder had less to describe. You are not buying bytes with quality — you are losing both, because the bytes you “saved” bought you nothing you asked for.

The shortcut is gated on the request carrying an explicit numeric quality — auto-quality has to probe the full-resolution master, so it cannot take it. Third arm, same laddered asset, same width, quality simply omitted. “pinned” = quality=75; “unpinned” = no quality in the DSL.

photo pinned bytes pinned PSNR pinned SSIM unpinned bytes unpinned PSNR unpinned SSIM unpinned source
autumn 103.3 KB 30.76 0.9455 105.5 KB 33.26 0.9611 original
glacier 61.9 KB 34.06 0.9335 63.6 KB 35.42 0.9472 original
night 11.7 KB 38.35 0.9405 11.4 KB 39.02 0.9457 original
snow 27.1 KB 35.38 0.9834 27.8 KB 37.44 0.9882 original
spices 51.3 KB 34.74 0.9304 53.8 KB 36.08 0.9463 original

Dropping quality moved every request back to original — one pass — and recovered most of the loss for almost nothing: the byte change ranges from -2.6 % to +4.7 % (1 of 5 came back smaller), while SSIM improves on 5 of 5. If you have the ladder and you want the master’s quality, do not pin quality.

One photo (spices, 6000×4000 master), format=webp,quality=75, both assets. source is the variant the encoder actually decoded.

width delivered px with ladder source original only source ladder vs plain
400 400 px 18.0 KB sm 19.8 KB original -9.1 %
480 480 px 23.1 KB sm 25.9 KB original -10.5 %
600 600 px 30.6 KB sm 35.7 KB original -14.2 %
640 640 px 43.6 KB sm 39.3 KB original 10.9 %
800 800 px 51.3 KB md 56.6 KB original -9.4 %
960 960 px 63.3 KB md 72.1 KB original -12.2 %
1080 1080 px 73.8 KB md 85.0 KB original -13.2 %
1200 1200 px 85.2 KB md 98.9 KB original -13.8 %
1280 1280 px 110.9 KB md 112.7 KB original -1.5 %
1440 1440 px 118.6 KB lg 134.8 KB original -12.0 %
1600 1600 px 136.4 KB lg 157.4 KB original -13.4 %
1920 1920 px 207.6 KB lg 203.8 KB original 1.9 %
2560 2560 px 290.0 KB xl 321.3 KB original -9.7 %

Every width delivers exactly the pixels it promises — the ladder is a fixed list and an off-ladder number is a 400 at the edge, not a resize.

Same photo, width=1200, both assets.

quality ladder bytes ladder SSIM original-only bytes original-only SSIM original-only PSNR
60 70.3 KB 0.9090 80.5 KB 0.9375 36.00 dB
70 80.2 KB 0.9165 91.9 KB 0.9460 36.72 dB
75 85.2 KB 0.9196 98.9 KB 0.9505 37.15 dB
80 105.7 KB 0.9313 123.6 KB 0.9621 38.45 dB
85 132.7 KB 0.9395 158.6 KB 0.9744 40.06 dB
90 177.6 KB 0.9469 212.1 KB 0.9828 42.03 dB

What each quality step costs, per hundredth of SSIM (single-pass arm — the honest one):

step extra bytes extra SSIM bytes per 0.01 SSIM
60 → 70 +11.5 KB +0.0085 13.5 KB
70 → 75 +6.9 KB +0.0045 15.3 KB
75 → 80 +24.7 KB +0.0117 21.2 KB
80 → 85 +35.0 KB +0.0123 28.5 KB
85 → 90 +53.5 KB +0.0084 63.6 KB

The curve is the answer to “where does it stop being worth it”: the marginal price of quality rises from 13.5 KB per SSIM hundredth at the bottom of the range to 63.6 KB at the top. Note what that is and is not: SSIM is an objective similarity metric, not an eye. It tells you where the encoder stops buying fidelity with bytes. Whether a human can still see it at 1200 px on a given screen is a perceptual question this benchmark does not answer — see what is still open.

fit request delivered bytes
cover format=webp,quality=75,width=800,height=800,fit=cover,gravity=auto 800×800 55.8 KB
inside format=webp,quality=75,width=800,height=800,fit=inside 800×533 51.3 KB

cover fills the box and crops to do it; inside fits the box and keeps the whole frame, so the delivered height is whatever the aspect ratio gives.

There is no upscale. Asking for more width than the source has returns the source’s width, quietly and with a 200.

master asked delivered bytes source
laddered, 6000 px 3840 3840 px 647.6 KB xl
small master, 1200 px 1920 1200 px 64.6 KB original
small master, 1200 px 3840 1200 px 64.6 KB original

Put widths in your srcSet that your masters can sustain, or the browser picks a candidate that does not exist at that size and downloads a smaller image than it thinks it is getting.

Same photo, width=800, quality=75:

format status content-type bytes
webp 200 image/webp 51.3 KB
jpeg 200 image/jpeg 59.8 KB
png 200 image/png 877.6 KB
avif 200 image/avif 73.5 KB

The platform standardises on WebP. The DSL accepts the others; the numbers above are bytes only. Note in particular that AVIF is heavier than WebP here at the same nominal quality — which is a statement about the two encoders’ quality scales, not proof that AVIF compresses worse: nothing on this page measured AVIF’s fidelity. Read the row as “the number you type means different things in different encoders”, and do not switch format on a byte count alone.

  • If your assets have the preset ladder (the recommended upload), an explicit quality puts a second lossy pass between your master and your visitor. Either drop quality and let auto-quality read the master, or accept the loss knowingly for the cheaper decode.
  • If your assets are ["original"] only, none of this applies: every transform comes off the master, one pass.
  • The widths that match a stored variant exactly (640, 1280, 1920) are the worst case — there is no downscale between the two passes to hide the first one’s artefacts, and at 640 and 1920 the result is heavier as well.

Stated so nobody mistakes silence for a result:

  1. Perceptibility. PSNR and SSIM are objective metrics. Nobody sat down and compared the two 800 px files side by side at 100 % on a calibrated display, and no perceptual metric (SSIMULACRA2, butteraugli) was run. “Where the difference stops being visible” is therefore reported as where the metric plateaus, which is not the same claim.
  2. Sample size per case. The double-compression A/B is n = 5 photographs and every one agreed. The width, quality, fit, ceiling and format sweeps are n = 1 — one photograph, chosen for being the busiest of the five. Treat the shape of those curves as the finding and the exact kilobytes as one data point.
  3. Cold-path timing. Server-Timing is on every response and was not collected; this page is about bytes and quality, not latency.
  4. Whether the shortcut should exist. It is a large decode saving and this page does not price it. It measures what it costs, so that the trade can be made on purpose.

Nothing here needs a privileged key. The probe runs against the public API with an ordinary runtime key on an ordinary tenant — the same two things credentials gives you — so you can re-run the whole thing on your own account:

Terminal window
export NITIDA_KEY="…" # your runtime key
export NITIDA_TENANT_CODE="…" # your tenant code

What the probe does, in order:

  1. Uploads the same photograph twice — once with the full ladder (["original","thumb","sm","md","lg","xl"]), once with ["original"] and nothing else.
  2. Requests the identical transform DSL from both and records status, content-type, bytes and the variant the response was decoded from.
  3. Sweeps width, quality, fit and format against the busiest of the five photographs, and computes PSNR/SSIM between the two arms.

The first run uploads both sides of the A/B; re-runs deduplicate and only re-measure. The cold path can answer 302 to the canonical /v/ object — the probe follows it and retries until it has bytes, because a naïve fetch records 0 KB and calls it a measurement.

All five from Pexels under the Pexels License (verified 2026-08-17: free to use, attribution not required — credited anyway).

photo character photographer · source
autumn warm / golden — backlit yellow and orange, almost no cool hue Ekaterina Swiss · Pexels 11696183
glacier cool / blue — deep cyan ice against near-black rock Lars H Knudsen · Pexels 13314282
night very dark — a low-key twilight scene, everything below mid-grey Andrea Bova · Pexels 5871563
snow very light — a high-key scene, almost no dark pixels at all Perati Wattanawikkarn · Pexels 30471757
spices saturated — cones of red, orange and yellow powder, no neutrals Mike van Schoonderwalt · Pexels 5504603