← Back to the vault

Lab Note · updated 2026-07-10

The QC-Aware Report

A QC-aware report doesn't just present a result — it surfaces the evidence that the result is trustworthy, structured by the four nested levels of quality and ending in an explicit decision. Its job is to make a number un-trustable on sight when it shouldn't be trusted.

TopicsQC · Report · Cell Painting · Digital Pathology · Light-Sheet 3D/4D · Spatial Omics · no-reference IQA · percent-replicating
A symmetric, mandala-like lattice of fine pale-cyan lines and small nodes on a dark navy field, forming nested concentric star-and-diamond tiers around a dark center, with the innermost tier glowing faint amber.

A normal report presents the answer: a profile, a slide-level score, a count matrix. It says nothing about whether the answer should be believed. A result computed on an out-of-focus field, a batch-confounded plate, or a mis-registered tile looks identical to a good one — the failure is invisible at exactly the moment a decision is made, and invisibility at that moment is the whole problem. The number is not wrong in a way anyone can see; it is wrong in a way that only its provenance would reveal, and provenance is precisely what a conventional report discards.

A QC-aware report inverts that default. It surfaces the evidence of trustworthiness alongside the result, so a reader can see why to trust it — or see immediately that they shouldn't. It is not a longer report or a report with more charts; it is a report whose structure mirrors the structure of quality itself, and whose final act is an explicit decision rather than a silent hand-off. This is the reporting discipline Fovea Lab builds into every pipeline it engineers, because a pipeline that cannot say why to believe its own output is not finished — it is only running.

1. Four nested levels of quality

Structure the report by the four levels from the quality model, because evidence at one level says nothing about the levels above it. The levels nest: each depends on the one beneath it and is blind to it. Reporting only the bottom level while implying the top is the single most common way a trustworthy-looking number turns out to be untrustworthy.

  • Image-space — per-image physics: focus, signal-to-noise, illumination uniformity, saturation, edge artifacts. Where no clean reference image exists — which is the normal case in microscopy — these are flagged with no-reference IQA, quality estimates computed from the image alone. Image-space QC is necessary and cheap, which is exactly why it is so often mistaken for the whole story.
  • Run / sample — consistency across the acquisition rather than within a single frame: field-to-field intensity drift, plate and batch structure, z-depth signal decay, stage-registration stability. This is the level at which the distinction between provenance metadata — what was acquired, when, on which instrument, under which settings — and quality metadata — whether those acquisitions are mutually consistent — earns its keep; conflating the two leaves a run that is fully documented and quietly inconsistent 1.
  • Readout — the fidelity of the quantity the platform actually reports: segmentation accuracy, percent-replicating, transcript assignment, slide-level score. This is the only level the science consumes, and the only one whose metric must be chosen to reflect the domain interest rather than whichever number is convenient to compute 3.
  • Decision — the operational verdict the report must emit: reportable, needs-review, or reject-and-reacquire. A report that stops before the decision has pushed the hardest judgment onto whoever reads it last, usually with the least context.

2. The report is a loop, not a document

A QC-aware report ends in that decision, traces it back through the three levels of evidence beneath it, and carries the lineage (attached to the output) so the verdict travels with the result rather than living in a reviewer's memory. The lineage is what makes a verdict auditable: a reportable stamp with no trail beneath it is an assertion, not evidence.

The two negative verdicts are not dead ends — they are edges back into the pipeline. A verdict of reject-and-reacquire routes back to acquisition: the sample is re-imaged, because no downstream correction recovers signal that was never captured. A verdict of needs-review routes back to QC: a human, or a second method, adjudicates the ambiguous case before it is allowed to become a number. Only reportable exits forward. Drawn out, the report is a small control system — evidence gates a decision, and the decision closes back onto the acquisition that produced it (Fig. 1) — not a document that is written once and filed. Treating it as a loop is also what keeps the report honest over time: a station that keeps routing back to acquisition is telling you something about the instrument, not just the sample.

3. Where it breaks

The general failure is reporting a single level and implying the rest — almost always image-space sharpness standing in for readout truth. It is worth naming why this is seductive: image-space QC is the easiest level to measure and the most visually reassuring, so a crisp field becomes a proxy for a correct result even though the two are separated by every processing step in between. Methodological review of machine learning in medical imaging found this pattern to be systemic rather than incidental — validation that measures something adjacent to the claim being made, and is therefore silent about the claim itself 2.

The specific form the gap takes is platform-dependent, and each is a distinct, nameable failure mode:

  • A Cell Painting report showing sharp fields but hiding plate-level batch structure passes a confounded assay: every image is in focus, and the biology is an artifact of which plate a sample happened to sit on.
  • A digital pathology report with acceptable global focus but no per-tile QC ships slide-level scores computed partly from blurred, folded, or pen-marked regions a model should never have embedded — the aggregate looks fine because the failures are local.
  • A light-sheet report citing a post-deconvolution sharpness figure, with no readout-level segmentation check, certifies structure that the deconvolution may have synthesized rather than recovered — a sharper image of something that was not there.
  • A spatial-omics report that omits registration error lets a spatially corrupted count matrix look clean at the readout, because the numbers are internally consistent even while they are assigned to the wrong locations.

In every case the report is not lying; it is silent at the level that matters, and silence reads as assent.

4. Choosing the metric is itself the hard part

Even a report that reaches the readout level can fail there, because the readout metric can be the wrong metric. The Metrics Reloaded framework exists precisely because the field's default metric choices routinely fail to reflect the biological or clinical interest they are meant to certify: a segmentation score can be excellent while the readout it feeds is meaningless, and an aggregate can look strong while systematically hiding the small-object or rare-class errors that the actual question turns on 3. A QC-aware report therefore does not merely display a readout number — it displays a readout number chosen to move when the thing you care about moves, and it says which number that is and why. Picking the metric is not a formatting detail downstream of the science; it is the science, made measurable, and a report that hides the choice hides the most consequential judgment in the pipeline.

5. The operating stance

Report the decision. Show the evidence at every level beneath it. Never let a green light at one level imply green lights above it. That discipline is the operational face of verified outputs — breadth, because it insists the whole system be shown rather than a single flattering slice of it, and trust, because it makes belief conditional on evidence that travels with the result. The goal is not a report that always says reportable; it is a report that makes a number un-trustable on sight when it shouldn't be trusted.

Warning

A report that shows only the result is asking for blind trust. A QC-aware report shows the evidence and the decision — so a result that shouldn't be trusted is un-trustable on sight.

The QC-aware reporting loop: evidence gates a decision, and the decision routes back to acquisition.
Fig. 1The QC-aware reporting loop: evidence gates a decision, and the decision routes back to acquisition.

Bibliography

  1. 1Huisman et al., A Perspective on Microscopy Metadata: Data Provenance and Quality Control (2019)
  2. 2Varoquaux & Cheplygina, ML for Medical Imaging: Methodological Failures (npj Digital Medicine, 2022)
  3. 3Maier-Hein, Reinke et al., Metrics Reloaded: Recommendations for Image Analysis Validation (Nature Methods, 2024)