.. _image_checking: ********************** Reviewing image checks ********************** The integration tests compare every image a run produces against a saved expected image. This page explains how to review what comes back. Why images differ ================= Two images can differ for very different reasons: * **Cosmetic.** Upgrading matplotlib changes text metrics slightly. Figures are saved with ``bbox_inches="tight"``, so a plot can come out a pixel wider and every anti-aliased edge lands a fraction of a pixel away. Nothing about the science changed. * **Real.** A panel is missing, a title was dropped, a contour moved, or a printed statistic changed. These need a person to look at them. A plain pixel-by-pixel comparison cannot tell these apart. Visually identical plots routinely differ in 2-20% of their pixels after a matplotlib upgrade, so comparing raw pixels reports essentially every image as a failure. In the 2026-08-04 weekly test that meant 1274 of 1280 MPAS-Analysis images were listed as failures with nothing to distinguish them. Severity ======== Each image pair is therefore given a severity, and failures are sorted worst first. Work down the list and stop when the differences stop mattering. .. list-table:: :header-rows: 1 :widths: 15 60 25 * - Severity - Meaning - What to do * - ``STRUCTURAL`` - The figure changed size enough that a panel came or went. - Always investigate. * - ``MAJOR`` - A large part of the plot looks different. - Always investigate. * - ``MODERATE`` - A visible part of the plot looks different. - Investigate. * - ``MINOR`` - Slightly different, or a small isolated change such as a printed number. - Skim. * - ``NEGLIGIBLE`` - The images differ, but only in ways that look like rendering noise. - Nothing, though a sample is rendered so you can confirm that. * - ``IDENTICAL`` - The images do not differ at all. - Nothing. * - ``MISSING`` - The image was never created. - Always investigate. ``IDENTICAL`` and ``NEGLIGIBLE`` images do not fail the test. They are counted separately in the summary, because "did not change" and "changed in a way that looks cosmetic" are different claims, and only the second is a judgement. What gets written ================= Alongside the existing outputs, each task's diff directory gets: ``severity_report.txt`` The ranked list. Start here. It opens with counts, then a grouping by likely cause, then every image needing review, worst first. ``image_scores.json`` The same information as raw numbers, one entry per image. Useful for re-examining thresholds without re-running the comparison. ``cosmetic_sample//image_diff_grid.pdf`` The most-different images that were called cosmetic, so that verdict can be checked rather than trusted. Only the first twenty are rendered; ``severity_report.txt`` lists all of them. They are the ones closest to the line, so if the worst are genuinely cosmetic the rest are too. Reviewing efficiently ===================== Most failures share a handful of root causes -- a single upstream change can alter hundreds of images identically. ``severity_report.txt`` therefore groups failures by cause before listing them: .. code-block:: none Grouped by cause (check one example from each): 381 STRUCTURAL figure size changed a lot 302 MAJOR figure height changed 53 MAJOR figure size changed slightly 28 MODERATE same size, content differs 9 MINOR small isolated difference Check one example from each group rather than every image. If the example is a genuine regression, the whole group almost certainly is too. These labels describe what was measured, not a diagnosis -- the check knows how the images differ, not why. They are useful because a single upstream change tends to alter many images the same way, so images sharing a label usually share a cause. In past runs "figure size changed a lot" has turned out to be a missing panel in one task and added axis labels in another. A note on small changes ======================= Severity is based on how much of the picture looks different, and that is not the same as how much it matters. A single wrong digit in a ``Mean -0.18`` label is about 13 pixels, far too small to affect an area-based score, but it means the data changed. Those are found by a separate check that looks for small, sharply-defined differences, and reported as ``MINOR`` with the cause "small isolated difference". Do not skip that group just because it ranks low -- it is the one place where a low rank does not mean a small problem.