Why a Single Image Quality Score Out of 100 Lies to You
Upload a photo to almost any quality checker and you get a number. 73. 81. 64 out of 100.
Now try this: take a photo that is perfectly exposed, completely free of noise, shot on a 24-megapixel camera — and totally out of focus. Run it through the same checker.
You will probably get something in the seventies.
That number is not wrong because the maths is broken. It is wrong because of what it did with the maths. Three of the four things it measured came back excellent, one came back terrible, and it averaged them. The photo is unusable, and the score says “pretty good”.
The averaging problem
Image quality is not one property. It is several independent ones that happen to be measurable on the same file:
- Focus — is anything in the frame crisply resolved?
- Noise — how much grain sits in the flat areas?
- Compression — is there visible JPEG block structure?
- Clipping — how much of the image is blown to pure white or crushed to pure black?
These do not trade off against each other. A photo is not “more usable” because its exposure is great when its focus is gone. You cannot spend good noise performance to buy back sharpness. They are not commensurable quantities, and averaging quantities that are not commensurable is how you get a number that describes nothing.
Worse, the average is actively misleading in the one case that matters most. A file with a single severe problem and three perfect scores lands in the same place as a file that is mediocre at everything. Those two files need completely different responses — reshoot versus re-export — and the score cannot tell you which one you are holding.
The right headline is the worst dimension, not the mean. If focus is Poor, the answer is Poor, and the word “focus” appears in the first sentence.
The second problem: “quality” is not a property of a file
Even per-dimension, “quality” only means something relative to a purpose.
Film grain is a defect in a product photograph and a deliberate choice in a portrait. Shallow depth of field is a blurred background or a focus miss depending entirely on where the subject is. A night scene is supposed to be dark. A gradient background is supposed to have no detail in it.
A checker that does not know this will confidently tell a photographer that their carefully lit portrait is “out of focus” because most of the frame is soft. That is not a small inaccuracy — it is the tool telling someone their deliberate work is a mistake, which is the fastest way to get a tool closed and never reopened.
The defensible version does three things instead:
- Measures per region, not per frame. Focus should be judged on the sharpest part of the image. If any region resolves crisply, the photograph is focused; the soft background is depth of field.
- Reports measurements, not judgements. “Noise is about 4 grey levels in the flat areas” is a fact. “This photo is noisy” is an opinion with a threshold hidden inside it.
- Declines to answer when it cannot. Which brings us to the interesting part.
When the honest answer is nothing at all
Some checks simply do not work on some images, and the failure is silent unless you build for it.
The compression check. JPEG blocking is detected by comparing the pixel steps across 8-pixel block boundaries with the steps between them. On a photograph this works beautifully. On a screenshot or a chart, it collapses: the flat fills compress to genuinely flat, so the boundary steps stay near zero no matter how hard the file is squeezed, while the hard text edges dominate the comparison. Measured across JPEG quality 95 down to 8, a photographic image moves from 1.29 to 13.59 — an enormous, clean signal. A flat-fill graphic over the same range wobbles between 1.11 and 1.27 and never leaves “no blocking detected”.
So on a screenshot, that check has to say nothing. Reporting “no compression artefacts” would be a clean bill of health the measurement cannot support.
The focus check. A plain gradient — a clear sky, a studio backdrop — has almost no edge energy, for the honest reason that it contains no edges. Run a naive sharpness metric over it and you get a number indistinguishable from a badly defocused photo.
The two are separable, but only if you look. Measured on a natural image, even absurd defocus bottoms out around 0.20 on a normalised edge-energy scale (Gaussian blur σ=8 gives 0.280, σ=16 gives 0.203). A linear gradient reads 0.097, and a uniform field produces no measurable region at all. There is a wide gap, and the honest thing to do inside it is stop.
What about “was this upscaled?”
This is the check everyone wants, and it is the one we could not make work honestly.
The idea is appealing: a genuine 4K capture carries real detail all the way to its pixel limit, while a 1080p file stretched to 4K has a ceiling where the interpolation stopped inventing. Measure where the detail runs out and you can tell someone their file has the real resolution of something a quarter the size.
Two methods, both measured:
Spectral cutoff — compare energy in the top octave of the frequency spectrum to the octave below. On a natural image the ratio came out at 0.091; the same image upscaled 2× gave 0.045, and 4× gave 0.027. A real separation. But film grain over the original scored 0.816, and a night scene scored 0.290 — both far above the untouched file. Noise is broadband. It fills exactly the frequencies the check is looking at, so the metric ends up measuring grain, not detail.
Round-trip idempotence — an already-upscaled file should barely change when you downscale and upscale it again, because there is nothing left to lose. Measured: native 0.554, upscaled-2× 0.332, upscaled-4× 0.697. The most degraded file scored higher than the original, and both grainy variants sat at 0.90 regardless of whether they had been upscaled at all.
Neither separates upscaling from noise or from ordinary variation between scenes. So the check does not ship. A tool that tells a photographer their grainy night shot “has the detail of 1080p” is not a useful tool, however interesting the underlying idea is.
What this means for you
If you are using a quality checker:
- Distrust a single number. Ask what it averaged. If one dimension was terrible and the score is fine, the score is hiding it.
- Look for per-dimension results. A tool that will not show you the breakdown is a tool that cannot defend the summary.
- Treat “cannot tell” as a feature. A checker that always has an answer has an answer for cases it cannot measure, and you have no way to know which ones those are.
- Remember what one file can and cannot tell you. No-reference analysis is measuring a file with nothing to compare it against. It can find blur, noise, blocking and clipping. It cannot tell you whether the colours are right, whether the crop is good, or what the file looked like before someone edited it.
That last point is the real limit. If you still have the original, comparing the two directly will tell you exactly what an export, an edit or a platform’s re-encode changed — pixel for pixel — which is a different and much more precise question than “how good is this file”.
Check one image free: DiffALL’s image quality checker measures focus, noise, compression and clipping, names the weakest one, and tells you when a check does not apply. No score out of 100, no sign-up.
Got the original too? Compare the two images for SSIM, PSNR and a pixel-level heatmap of exactly what changed.
Stop hunting for differences by hand. DiffALL spots every change between any two files — automatically.
Compare your files — free