Found while reading backend/app/api/analysis.py and the Analysis page components line by line at 591234b, and by running the module's own functions over a real library (416 LIGHT frames, 3 nights, 2 rigs with plate scales of about 0.99 and 2.94 arcsec per pixel). Grouped by how visible the fault is.
Wrong output a user sees
1. The histogram drops the maximum value (analysis.py:542)
The last bin's upper edge is accumulated as v_min + i * bin_width + bin_width, which is usually not exactly v_max in binary floating point. When it lands below v_max, the test i == n_bins - 1 and v == b_end fails and the frame at the maximum falls in no bin.
Observed: v_min 0.0, v_max 2.6245, 10 bins, bin_width 0.26244999999999996, last edge 2.6244999999999994. Bin counts sum to 415 of 416 and the last bin reads 0 instead of 1.
Suggested fix: use v_max as the last bin's upper edge, or test i == n_bins - 1 and v >= b_start.
2. The Compare verdict prints percentages above 100 (analysis.py:890 to 896)
pct_diff divides by med_a whenever it is non-zero, but the sentence always names the smaller group as "X% lower" than the larger. When the smaller median is under half the larger, the sentence reads over 100 percent lower.
Observed on three of three real comparisons: two rigs, HFR in arcsec, 1.772 against 3.997, reads "126% lower" (actual 55.7%); two filters, 1.944 against 3.879, reads "100% lower"; an FWHM comparison reads "111% lower".
Suggested fix: divide by the larger of the two medians, which is the one the sentence compares against.
3. The correlation verdict assumes lower Y is always better (CorrelationChart.tsx:55 to 63)
A rising slope is called a "negative impact" and a falling one "improves". detected_stars and sky_quality are in the Y list and are higher-is-better, so the sentence is backwards for both. TimeSeriesTab.tsx:18 already keeps a HIGHER_IS_BETTER set that the correlation verdict could share.
4. The Time Series tooltip names an arbitrary target on a multi-target night (analysis.py:712 to 713)
tid_list = list(g["target_ids"]) then tnames.get(tid_list[0]) takes the first element of a set. On a night with two targets the label is one of them, chosen by hash order, and can differ between processes. Two of the three real nights carry two targets.
Suggested fix: show the name only when the night has one target, otherwise "Mixed" or the joined names.
5. Raw pixel HFR is plotted across rigs with different plate scales
_PIXEL_METRICS has one reader, get_compare (analysis.py:843). Correlation, distribution, box plot, time series and matrix plot median_hfr in pixels even when the filter spans two optical trains. On the library above, pixel HFR ranks the two rigs the opposite way from arcsec HFR (1.36 against 1.79 px; 4.00 against 1.77 arcsec). The existing "not comparable across optical trains" note is easy to miss. A visible warning when the selected frames span more than one arcsec_per_pixel value would help, or an arcsec option on those tabs.
6. Help text that does not match the page
- The Matrix popover describes a heatmap of one metric across two categorical axes ("integration time per channel", "telescope by filter"). The tab is a Pearson matrix over metric pairs.
- The Time Series popover says data is "either per frame or aggregated by session or night". The endpoint ignores
granularity; so do box plot, matrix and compare. Only correlation and the histogram honour it, and nothing in the UI says so.
- Several popovers call the date range "capture date".
_apply_filters filters session_date (analysis.py:167).
Statistical or consistency issues
7. The "95% CI" band is too narrow on small samples (analysis.py:305)
The multiplier is 1.96 if n > 30 else 2.0. The two-sided 95 percent t value at n = 4 is 4.30, so the band is under half its correct width there while still labelled "95% CI". A small t table, or relabelling the band as indicative, would fix the claim.
8. _pearson_r answers 0.0 for three different conditions (analysis.py:239 to 250)
n below 3, a constant X and a constant Y all return exactly 0.0, which is also a valid result for uncorrelated data. The page cannot distinguish "no correlation" from "not enough data" or "this column is constant". The matrix path uses PostgreSQL corr() (:791), which returns NULL for a constant column, so the two paths disagree on the same data. Returning None and rendering "n/a" would align them.
9. Box plot "By Month" groups by capture_date while every other grouping and the date filter use session_date (analysis.py:597 against :167)
A frame taken after midnight on the last night of a month lands in the next month's box while its session belongs to the previous one.
10. "Not enough data" behaves two ways on one page (analysis.py:529, :650 to 652, :883)
Distribution raises HTTP 400 under 2 values and compare raises 400 under 4 per group, while the box plot silently drops a group under 4. A user who filters down to 3 frames per filter sees an empty chart with no reason.
11. Skewness is computed from already-rounded inputs (analysis.py:546 to 550)
stats.mean and stats.std_dev come from _compute_summary_stats rounded to 6 decimals, then get cubed. Harmless in practice, but the last digits differ from a full-precision computation.
Performance
12. _is_outlier_iqr re-sorts both axes once per point (analysis.py:226 to 236, called at :446)
For n points this is 2n sorts of n elements. The quartiles do not depend on the point being tested, so they can be computed once outside the comprehension with identical flags. At the 100,000 frames the comment at :457 anticipates, this is 200,000 sorts of 100,000 elements per request.
Dead or unreachable code
13. Dead branch in /distribution (analysis.py:561 to 563)
_compute raises HTTPException rather than returning None, so the caller's if data is None check never fires.
14. The first "not comparable" sentence in Compare is unreachable
The backend sets both arcsec medians only inside the comparable branch (analysis.py:910 to 913), so comparable == false always arrives with both null, and the frontend's crossTrainVerdict branch that reads the two medians can never run. The backend's own verdict string for that case is built and then discarded by the frontend.
15. No ORDER BY on any statement feeding a figure
Row order is whatever PostgreSQL returns. The builtin float sum() results, which points survive the 5,000 point downsample, and the emission order of session groups all depend on it, so two identical requests can differ in the last digits or in which points are drawn. ORDER BY session_date, id would make results reproducible.
Found while reading
backend/app/api/analysis.pyand the Analysis page components line by line at591234b, and by running the module's own functions over a real library (416 LIGHT frames, 3 nights, 2 rigs with plate scales of about 0.99 and 2.94 arcsec per pixel). Grouped by how visible the fault is.Wrong output a user sees
1. The histogram drops the maximum value (
analysis.py:542)The last bin's upper edge is accumulated as
v_min + i * bin_width + bin_width, which is usually not exactlyv_maxin binary floating point. When it lands belowv_max, the testi == n_bins - 1 and v == b_endfails and the frame at the maximum falls in no bin.Observed:
v_min0.0,v_max2.6245, 10 bins,bin_width0.26244999999999996, last edge 2.6244999999999994. Bin counts sum to 415 of 416 and the last bin reads 0 instead of 1.Suggested fix: use
v_maxas the last bin's upper edge, or testi == n_bins - 1 and v >= b_start.2. The Compare verdict prints percentages above 100 (
analysis.py:890to896)pct_diffdivides bymed_awhenever it is non-zero, but the sentence always names the smaller group as "X% lower" than the larger. When the smaller median is under half the larger, the sentence reads over 100 percent lower.Observed on three of three real comparisons: two rigs, HFR in arcsec, 1.772 against 3.997, reads "126% lower" (actual 55.7%); two filters, 1.944 against 3.879, reads "100% lower"; an FWHM comparison reads "111% lower".
Suggested fix: divide by the larger of the two medians, which is the one the sentence compares against.
3. The correlation verdict assumes lower Y is always better (
CorrelationChart.tsx:55to63)A rising slope is called a "negative impact" and a falling one "improves".
detected_starsandsky_qualityare in the Y list and are higher-is-better, so the sentence is backwards for both.TimeSeriesTab.tsx:18already keeps aHIGHER_IS_BETTERset that the correlation verdict could share.4. The Time Series tooltip names an arbitrary target on a multi-target night (
analysis.py:712to713)tid_list = list(g["target_ids"])thentnames.get(tid_list[0])takes the first element of a set. On a night with two targets the label is one of them, chosen by hash order, and can differ between processes. Two of the three real nights carry two targets.Suggested fix: show the name only when the night has one target, otherwise "Mixed" or the joined names.
5. Raw pixel HFR is plotted across rigs with different plate scales
_PIXEL_METRICShas one reader,get_compare(analysis.py:843). Correlation, distribution, box plot, time series and matrix plotmedian_hfrin pixels even when the filter spans two optical trains. On the library above, pixel HFR ranks the two rigs the opposite way from arcsec HFR (1.36 against 1.79 px; 4.00 against 1.77 arcsec). The existing "not comparable across optical trains" note is easy to miss. A visible warning when the selected frames span more than onearcsec_per_pixelvalue would help, or an arcsec option on those tabs.6. Help text that does not match the page
granularity; so do box plot, matrix and compare. Only correlation and the histogram honour it, and nothing in the UI says so._apply_filtersfilterssession_date(analysis.py:167).Statistical or consistency issues
7. The "95% CI" band is too narrow on small samples (
analysis.py:305)The multiplier is
1.96 if n > 30 else 2.0. The two-sided 95 percent t value at n = 4 is 4.30, so the band is under half its correct width there while still labelled "95% CI". A small t table, or relabelling the band as indicative, would fix the claim.8.
_pearson_ranswers 0.0 for three different conditions (analysis.py:239to250)n below 3, a constant X and a constant Y all return exactly 0.0, which is also a valid result for uncorrelated data. The page cannot distinguish "no correlation" from "not enough data" or "this column is constant". The matrix path uses PostgreSQL
corr()(:791), which returns NULL for a constant column, so the two paths disagree on the same data. ReturningNoneand rendering "n/a" would align them.9. Box plot "By Month" groups by
capture_datewhile every other grouping and the date filter usesession_date(analysis.py:597against:167)A frame taken after midnight on the last night of a month lands in the next month's box while its session belongs to the previous one.
10. "Not enough data" behaves two ways on one page (
analysis.py:529,:650to652,:883)Distribution raises HTTP 400 under 2 values and compare raises 400 under 4 per group, while the box plot silently drops a group under 4. A user who filters down to 3 frames per filter sees an empty chart with no reason.
11. Skewness is computed from already-rounded inputs (
analysis.py:546to550)stats.meanandstats.std_devcome from_compute_summary_statsrounded to 6 decimals, then get cubed. Harmless in practice, but the last digits differ from a full-precision computation.Performance
12.
_is_outlier_iqrre-sorts both axes once per point (analysis.py:226to236, called at:446)For n points this is 2n sorts of n elements. The quartiles do not depend on the point being tested, so they can be computed once outside the comprehension with identical flags. At the 100,000 frames the comment at
:457anticipates, this is 200,000 sorts of 100,000 elements per request.Dead or unreachable code
13. Dead branch in
/distribution(analysis.py:561to563)_computeraisesHTTPExceptionrather than returningNone, so the caller'sif data is Nonecheck never fires.14. The first "not comparable" sentence in Compare is unreachable
The backend sets both arcsec medians only inside the comparable branch (
analysis.py:910to913), socomparable == falsealways arrives with both null, and the frontend'scrossTrainVerdictbranch that reads the two medians can never run. The backend's own verdict string for that case is built and then discarded by the frontend.15. No
ORDER BYon any statement feeding a figureRow order is whatever PostgreSQL returns. The builtin float
sum()results, which points survive the 5,000 point downsample, and the emission order of session groups all depend on it, so two identical requests can differ in the last digits or in which points are drawn.ORDER BY session_date, idwould make results reproducible.