Skip to content

Latest commit

 

History

History
733 lines (605 loc) · 99.4 KB

File metadata and controls

733 lines (605 loc) · 99.4 KB

Backlog

Derived from session work, uncommitted changes, and codebase state. Last updated: 2026-08-12 (added 6 items from a ~/Downloads classification audit — 28 files inspected against their actual contents, which found 6 misclassifications; three defect classes were fixed in the same session and the six items below are what was not addressed. The audit's own lesson: two of the three fixes were mischaracterised on first report and had to be corrected after measurement. "PDF text extraction returned nothing" was not an extraction bug at all — a filename rule returning ("business","legal") at full confidence early-exited the cheap wave, so no text was ever read; and "no confidence floor on renaming" was the wrong instrument, since every correct rename also scored 1–2% against a softmax floor of 1.06% — the real defect was a calibrated ratio gate that skipped itself whenever the margin was None, which is precisely the ungated OCR-override case. Read a symptom stat before trusting it: the 0/28 extracted-text counter that prompted the false diagnosis is itself an undercount, now recorded below. A third correction: removing the GPS-presence travel rule was reported as a likely broad regression and measured as a single row. Prior update 2026-08-11 (closed 2 items — Lint debt flake8 src/ scripts/ tests/ 540 → 498 → 0 via a mechanical black pass over 109 files (f905953) plus the 72 findings black cannot fix (a8e1d72); and E2E dashboard assertions, now driven by DASHBOARD_CARDS/DASHBOARD_STATS fixtures (d97be75, 1013673). Two lessons worth carrying: an enumerated worklist in this file is a snapshot, not a contract — the lint item's 540 had already decayed to 498 and its per-code breakdown no longer matched; and when an item names one hardcoded list, look for its siblings — the E2E item blamed the count assertion alone, but the same card set was hand-enumerated in three places, so the fifth card had no navigation test either. Both residuals recorded on the E2E item were then closed in 231019a — the .tech-badge filters (the footer has 8 badges; the test asserted 3 with no count, so 5 were uncovered) and the .feature-card:nth-child(1..4) stagger that left the 5th card at animation-delay: 0s. Both residual write-ups were themselves wrong in a way worth noting: the badge one called the gap "lower-stakes" without checking how many badges existed, and the CSS one asserted _site/index.html is generated by content runs when it is the committed source (copy_to_site.sh skips it by name). The stagger needed a computed-style assertion because :nth-child enumerations are invisible to ordinary E2E coverage — no locator, navigation or text assertion can see a missing animation delay. Prior entry: closed HEIC OCR _PREPROCESS_AVAILABLE bug via /backlog-implementer — 7899442; one recall-measurement residual remains open, and two notes were appended to that item — the np.asarray guard shipped with the fix is unreachable given the module's numpy-before-PIL import order, and f8aa49b's DocTRResult Protocol introduced a mypy no-any-return that the Stop hook caught. Also closed the scripts/d1/schema.sql drift item — tests/unit/test_d1_schema_drift.py regenerates from Base.metadata and diffs against the committed file, wired into a schema-drift job alongside the new style gate in .github/workflows/checks.yml (make lint / make schema-check are the identical local commands). Both gates were verified to fail, not just to pass — a misformatted function trips black, an unused import that black accepts trips flake8, and a model column added without regenerating trips the schema check, which make d1-schema then repairs. Repo-state caveat: 195a5b5 committed a docs/BACKLOG.md that links to docs/changelog/2.3.0/CHANGELOG.md and docs/CHANGELOG.md while both were still untracked, and 2.2.0's additions unstaged — until those are committed, HEAD's backlog points at migration destinations that exist only in the working tree. Earlier same day: migrated 11 Done items out via /backlog-migrate — 6 to docs/changelog/2.2.0/CHANGELOG.md, 5 to the new docs/changelog/2.3.0/CHANGELOG.md, with docs/CHANGELOG.md added as the version index. Four of the migrated items carried open residuals that the migration left recorded only as changelog prose; they were re-opened the same day as four standalone items — near-dupe reconcile integration + synthetic scale numbers, the unvalidated DEFAULT_SIMILARITY_THRESHOLD = 0.85, the 164-row calibration corpus, and the HEIC OCR verification gaps (recall never measured; docTR _PREPROCESS_AVAILABLE gate now fixed in 7899442). The general hazard is worth remembering: a "Done" item whose residuals live in bullets is fully migratable by shape, and the residuals leave the worklist with it — a changelog records what shipped, not what remains. Separately, the P2 /git-commit-smart blanket-staging item was moved to ~/.claude/docs/BACKLOG.md as GCS1 — the defect is in the skill, not this project). Prior update 2026-08-11 (closed 2 items via /backlog-implementer — P4 package.json version/test:unit script in 04a9ba7, and the P2 HEIC-OCR decode hand-off in e9fb0a8+41ee326+1c32fe7. The HEIC item is Done for the decode path only: two residuals were found during verification and appended to it — recall was never measured on text-bearing HEICs, and the docTR branch was gated on _PREPROCESS_AVAILABLE; the gate was removed in 7899442). Prior update 2026-08-11 (added 3 items from a test-maintenance session — hardcoded E2E dashboard card assertions, stale package.json version + missing unit-test script, and /git-commit-smart blanket-staging another session's work in the shared checkout; no code written for any of them beyond the 5-line dashboard.spec.ts fix already committed in 58f7659). Prior update 2026-08-10 (added 4 items from the facebookresearch library audit — faiss near-dupe index + SSCD descriptors as two halves of one feature, nevergrad joint weight search, and a PE-Core-vs-ViT-B-32 backbone A/B; library viability, licensing, and py3.14/arm64 wheel availability verified on this machine, no code written for any of them). Prior update 2026-07-25 (migrated 3 Done items to docs/changelog/2.2.0/CHANGELOG.md; resolved the census-gazetteer setup gap — scripts/download_census_names.py verified + documented in QUICK_START/CLAUDE.md; shipped the --ocr-doctr-fallback config flag for the P2 docTR-fallback gate, closing the OCR-bound item's last work item). Prior update 2026-07-18 (added Repo Snapshot — repomix token census, top-churn gitlog, and the uncommitted InteriorSignal→SceneSignal retirement inventory; added identity-detection license-back item + partial fix — corrective lenses keyword added to ID_KEYWORDS (restrictions/endorsements trialed then dropped after backtest showed insurance-doc collision), front/back fixtures, 23 tests pass; added redact_pii.py barcode/alphabetic-PII blind-spot item — OCR-token redaction silently no-ops on ID barcodes + health terms; added trained graphic-vs-photograph probe item — opaque AI graphics/logos leak past the cheap GraphicDetectionSignal gate (code path since closed by c327877 — graphic scene-probe class wired end-to-end, pending corpus + retrain); corrected PHOTO_PROPERTY_CONFIDENCE item post-f6488b9 — two-signal case resolved, residual is probe-absence only; probe now health-checked; fixed person-name false positive — ambiguous Census given names (summer/spring/autumn/winter, month names, virtue words) were auto_accepted when paired with a Census surname; new _AMBIGUOUS_GIVEN_NAMES hard rule + 41-test suite; closed redact_pii.py barcode item — cv2 barcode+QR detection, --redact-terms flag, barcode_unredacted manifest field, non-zero exit, 27-test suite).

Open Items

Identity detection misses driver-license backs (front-side keyword list)

IdentityDocumentSignal / the legacy _classify_identification_document both delegate to detect_identity_document (src/scoring/signals/identity_document.py:95), which fires only when ocr_text contains one of the ID_KEYWORDS (:44). The original 14 keywords were all front-side terms ("passport", "driver license", "date of birth", "surname"…). A photographed license back carries none of them — its OCR text is class/restriction/endorsement fields plus barcodes — so it was never detected and fell through to MIME/neither instead of personal/identification.

Surfaced 2026-07-18 while handling a real Texas license back (PXL_20220607_234355242.MP.jpg): OCR text was "CLASS: C-Single… / HAZMAT / REST: A - With corrective lenses / END: NONE / Directive to physician / Emergency Contact / Allergic reaction to drugs / TEXAS ROADSIDE ASSISTANCE" — zero ID_KEYWORDS hits. (Same file that exposed the redact_pii.py barcode blind spot above.)

Partially fixed 2026-07-18: added one license-back keyword to ID_KEYWORDS — "corrective lenses" (near-unique license restriction) — appended last so any front-side keyword still wins the reported-matched_keyword slot. ("restrictions" and "endorsements" were trialed then dropped after the backtest below showed they collide with insurance-document language; only corrective lenses was kept.) Two fixtures + tests added (test_signal_identity_document.py: DRIVER_LICENSE_TEXT front, LICENSE_BACK_TEXT back); 23 tests pass. Flows to both engines via the shared core.

Status: Open — partial. Residual gaps below. Priority: P3 Source: manual license handling during scene-probe corpus seeding, 2026-07-18

  • Unrestricted backs still undetectable. A license back with no restriction/endorsement (only barcodes + "NONE") has no keyword hook at all — OCR-keyword detection fundamentally can't see it. Robust license-back detection needs a barcode/PDF417 presence cue (cf. the redact_pii.py item — a PDF417 is itself a strong ID signal) or a trained ID-image probe, not more keywords.
  • "restrictions" / "endorsements" trialed and dropped (backtest 2026-07-18). Both are generic enough to appear on insurance/contract/benefits documents and would file to personal/identification at ID_KEYWORD_CONFIDENCE(0.85). Backtest (results/file_organization.db, 237 files / 58 with OCR): corrective lenses 0 matches, restrictions 0 matches, endorsements 2 matches — both insurance PDFs (My Documents _ USAA.pdf, Property_Insurance.pdf, "Declarations Page and endorsements…"), i.e. the exact collision predicted. No live false positive only because both are DigitalDocument and the ID signal is image-only gated (applies_to: is_image) — incidental protection that fails for a photographed insurance card. Decision: dropped both; kept only corrective lenses (clean, ID-specific, 0 collisions). If broader license-back coverage is later needed, reintroduce behind a corroborating-token requirement rather than as bare keywords. Corpus is small (spot-check, not comprehensive).
  • ID_KEYWORDS is now front+back mixed. If the reported matched_keyword is ever used to sub-classify (front vs back), the flat list won't distinguish them — would need tagging.

Trained graphic-vs-photograph probe (opaque full-bleed graphics leak to photos_*)

GraphicDetectionSignal (src/scoring/signals/graphic_detection.py) is a cheap pre-CLIP pixel gate tuned for transparent, small, square icon assets: its four additive heuristics are has_alpha (+0.40), is_square_icon (square & ≤512px, +0.30), small_palette (≤64 colors in a 64² thumbnail, +0.25), and asset_path (parent dir in the asset-folder set, +0.15), needing ≥ _MIN_RASTER_CONFIDENCE(0.35) to vote. It has no capability for opaque, high-res, flat-design graphics — logos on a solid background and text marketing posters, exactly what ChatGPT/AI image tools emit. Those score ≈ 0 on all four cues (no alpha, >512px, anti-aliased gradients/text push distinct colors >64, parent folder is ChatGPT not an asset dir), fall below 0.35, and leak through to the photos_chatgpt filename fallback (or photos_other).

Reproduced (shadow run, 2026-07-18) on ~/Documents/Media/Photos/ChatGPT, visually verified:

  • ChatGPTImageNov10,2025,02_32_53PM.png — "InventoryAI" brand logo (1024², opaque navy) → misfiled photos_chatgpt (both engines agreed; both wrong).
  • ChatGPTImageAug30,2025,03_41_57PM.png — "GOT A VISION FOR A HEALTHIER AUSTIN?" text poster (1536×1024, opaque cream) → misfiled photos_chatgpt.
  • Contrast: ChatGPTImageNov10,2025,02_32_56PM.png (busy 2×2 icon grid) did route to graphics_other — so detection currently fires only on busy multi-icon layouts, not single logos or text posters. Inconsistent by construction.

The durable fix mirrors the interior-detection precedent (docs/reviews/INTERIOR_DETECTION_DURABLE_FIX_ANALYSIS.md): replace/augment the pixel heuristics with a trained linear probe over the frozen ViT-B-32 embeddings the pipeline already caches — a graphic (or binary graphic-vs-photograph) class — exactly as InteriorSignal (src/scoring/signals/interior.py, results/interior_probe.joblib) did for the zero-shot interior gate. CLIP embeddings do separate graphics from photographs even though CLIP zero-shot labels don't (the whole reason GraphicDetectionSignal avoided CLIP). Landing chosen and shipped: a dedicated graphic class (index 4) in the scene probe (scripts/prototype_scene_probe.py, corpus results/scene_labels/) — see the runtime bullet below; the logo + poster above are ideal graphic/ positives (boundary rules: results/scene_labels/README.md, graphic-vs-photograph split added 38ddfd5).

Status: Open — narrowed 2026-07-27, labeling policy decided 2026-08-11. Corpus expanded 34 → 307 via Crello; real-world graphic recall 0.67 → 0.76. The promo-panel/dashboard boundary is now settled (content rule — see the 2026-08-11 bullets at the end of this item), which unblocks the next corpus increment. Remaining: pull data-viz into graphic/ and functional-UI/document screenshots into neither/, then retrain and re-eval. Priority: P3 Source: ChatGPT shadow-scorer investigation, 2026-07-18 (results/scoring_shadow.jsonl; agreement-set manual review)

  • Do not chase it with more pixel heuristics. Opaque AI-generated graphics are pixel-indistinguishable from photos by palette/size/alpha; only the learned embedding separates them.

  • Keep the cheap gate. GraphicDetectionSignal correctly and cheaply catches transparent/square icon assets pre-CLIP; the probe is an additional heavy-tier voter for the opaque-graphic case it can't see, gated on CLIP availability with graceful no-op (same pattern as InteriorSignal).

  • Corpus dependency. Needs labeled graphic/neither positives (target 150–300/class per the scene-probe README); results/scene_labels/ is seeded and actively being labeled — a dedicated graphic/ class dir now exists — but every class is still below target (2026-07-18 snapshot: neither 46, interior 44, exterior 21, graphic 8, place 3), so the probe is not yet trainable.

  • Regression guard. Measure false positives on genuine photos (product still-lifes, staged interiors) before deploying — a graphic probe that over-fires would pull real photos out of photos_*.

  • First eval baseline (2026-07-18, gather --label-dirs + eval, 3-fold CV, n=122). Corpus at eval time: neither 42, interior 44, exterior 21, place 3, graphic 12. Graphic class: precision 1.00 / recall 0.42 / F1 0.59 — confusion row graphic → [7 neither, 0, 0, 0, 5 graphic] (7/12 graphics still fall back to neither → leak to photos_*, the original bug). Classic starved-minority signature: high precision, low recall — corpus volume is the fix, not tuning. Deploy-safety confirmed: at the default SCENE_MIN_PROB=0.5 graphic precision holds at 1.00 across the 0.5→0.7 sweep, so a live 5-class probe would add no photo→graphics false positives (safe-but-partial: misses ~60% of graphics, no regression on them). Headline metrics are inflated (macro ROC-AUC 0.993 / acc 0.918): place n=3 is meaningless and near-duplicate images (two coffee-station photos, multiple Integrity banners) leak across CV folds — do not read as deployment-ready. Decision: do not train/commit a .joblib yet — an artifact activates the registry swap, and graphic recall 0.42 isn't worth shipping. Hold until graphic/≈150 and place/ is real.

  • Runtime landing path — code steps done (c327877, 2026-07-18). SceneSignal (src/scoring/signals/scene.py) + the artifact-gated registry swap landed in 5f6db5b; c327877 wired the 5th class end-to-end: "graphic": 4 in SCENE_CLASSES + _POSITIVE_NAMES (so gather picks up results/scene_labels/graphic/), runtime mapping SCENE_CATEGORY["graphic"] → ("media", "graphics_other") (resolves to Media/Graphics/Other, the same target GraphicDetectionSignal emits), SCENE_SCHEMA["graphic"] → ImageObject, _INT_CLASS_NAMES[4]. Back-compat: a pre-graphic 4-class artifact still loads — SceneSignal ignores classes absent from SCENE_CATEGORY, so nothing misroutes before retraining. 16/16 scene tests pass (new test pins graphic-argmax → media/graphics_other).

  • Trained and shipped (2026-07-18, supersedes the "do not train yet" hold above). Corpus was expanded the same day (Places365 sampling for place/exterior/interior + Business-folder graphics + Media-filtered --db-neither) to 835 rows — neither 118, interior 178, exterior 158, place 347, graphic 33 — clearing the hold's place-is-fake blocker. Final 5-fold eval: accuracy 0.92; graphic P 0.73–0.82 / R 0.60–0.67 / F1 ~0.70. User decision: train now, accept conservative graphic recall — the miss mode is benign (a missed graphic gets no scene vote and falls through to GraphicDetectionSignal + other signals, i.e. today's behavior). scene_probe.joblib trained + committed (6f61449); swap completion followed (interior.py deleted, photo_composition interior vote retired, scene_probe health feature, W_SCENE backtest: −20% flips 4 / +20% flips 0 on the 202-row replay). Still open here: graphic recall. Corpus volume from pure graphics (logos, posters, flat illustrations — local sources are tapped out; a public logo/infographic dataset is the realistic path), and a labeling-policy call on promo-panels-with-embedded-UI (semantically graphics, visually half-screenshot — currently the main graphic↔neither confusion).

  • Corpus expanded from Crello, 2026-07-27 — and the headline number is misleading. scripts/download_crello_graphics.py pulled 273 flat-design previews from cyberagent/crello across 14 format subtypes (logo, poster, flyer, ad creative, certificate, coupon…), filtered to templates with no ImageElement (those embed photographs) and one per cluster_index (near-duplicate variants would leak across CV folds). Class went 34 → 307, inside the 150–300 target band. Aggregate 5-fold CV: graphic P 0.73 → 0.97, R 0.67 → 0.96, F1 0.70 → 0.97; overall accuracy 0.90 → 0.93. Do not read that as a 0.97. Per-source recall: Crello 272/273 = 0.996, hand-collected 25/33 = 0.758. Crello templates are near-trivially separable, so they inflate the aggregate; the honest real-world gain is 0.67 → 0.76. Images are gitignored (graphic/crello_*) — CyberAgent conditions use on the VistaCreate ToS and does not redistribute source files, so the script is the reproduction path (same arrangement as download_census_names.py).

  • The residual error is one coherent subtype, not general scarcity. 7 of the 8 missed hand-collected graphics are biz_dashboard_* images and the 8th is an "AI Market Map" infographic — all flat-design data-viz, all routed to neither. Visually confirmed: these are non-photographic flat vector charts, so graphic is the correct label and the probe is genuinely wrong. Crello supplies posters/logos/ads and almost no data-viz, which is exactly why it didn't fix them. Next increment should target dashboards/infographics/charts, not more posters. InfographicVQA is the ideal content match but is research/education-only (blocked — see ~/.claude/.../memory/dataset-license-constraints.md); DomainNet infograph (51,605 images, one-line TFDS pull) is the license-viable candidate, pending verification of DomainNet's terms at ai.bu.edu.

  • This also sharpens the promo-panel labeling call. DECIDED 2026-08-11 — the content rule: the pixels decide, not the framing. Synthetic imagery is graphic/ even when the image is a screenshot; live interface chrome in the frame does not move it to neither/. Written into results/scene_labels/README.md as three bullets, replacing the both-ways ambiguity ("data-viz → graphic" vs "UI screenshots → neither") that made the README undecidable for the same image. Enrico's precedent argued the other way and was not followed: the graphic class exists to stop non-photographic imagery reaching photos_*, and a dashboard is not a photograph. No image was relabelled — all 12 contested biz_* files keep graphic/, so the 0.76 real-world recall figure remains comparable across the decision rather than being mechanically improved by moving the errors.

  • The rule left one line undecided, so the README resolves it explicitly: text documents and functional UI (settings, terminals, editors, forms, chat) stay in neither/. A literal "synthetic → graphic/" would swallow them, and they belong on the OCR/screenshot path, not in Media/Graphics. The discriminator inside synthetic content is imagery vs text/controls. A decision that resolves an ambiguity can introduce a new one at its own edge — state the edge in the same commit or the item reopens under a different name.

  • The risk now runs the other way, and the corpus is not shaped for it. This rule trains the probe toward "UI layout → graphic", so the new failure mode is genuine UI screenshots and document scans being pulled into Media/Graphics. neither/ is the class that must teach that distinction and it is by far the smallest — 73, against graphic 308 and place 347. So the next corpus increment is no longer just DomainNet infograph into graphic/; it must add functional-UI and document screenshots to neither/, and the retrain must report neither recall and the neither → graphic confusion cell.

  • Correction to the bullet above: the residual errors are not "one coherent subtype". Inspecting the images rather than the filenames, the 12 hand-collected biz_* files are at least four different things: a pure flat vector illustration of chart panels (biz_dashboard_9f2885b8, no interface at all), a chart-dominant real analytics view (biz_dashboard_26478eb2), composed promo panels of headline + bullets + product mockup (fba0e47a, biz_20260107_dashboard), and web captures carrying live chat-widget chrome (882118b3). Only the last two were ever contested; the first is a plain graphic the probe misses on volume alone. "They're all biz_dashboard_*" is a filename observation being read as a content finding — which also means a wholesale relabel of that prefix would have been wrong in either direction.

  • Encoding confound found and fixed (scripts/normalize_scene_corpus.py). The corpus had class partly recoverable from file metadata alone: place/ was 100% JPEG at 256px (Places365), hand-collected graphic/ mostly PNG at ~1536px. A logistic regression on encoding metadata only (format, dimensions, aspect, size, bytes-per-pixel — no pixels) scored 0.561 vs a 0.327 majority baseline, lift +0.234. Normalizing every class through one encoder (JPEG q90, longest side 256 — above CLIP's 224 input, so nothing the model sees is lost) cuts it to 0.411, lift +0.084. Ablation shows the remainder is aspect ratio (+0.139 alone) and compressibility (+0.072); aspect cannot reach the probe because CLIP center-crops to square, and compressibility is genuine content signal (flat design compresses better than photographs). Re-eval on the normalized tree holds graphic recall at 297/306, confirming the Crello gain is content, not encoding artifact. Output tree is gitignored and regenerable.

Content pipeline is OCR-bound (gate OCR on text-likelihood)

Profiling organize-files content (unified scorer, dry-run) on 2026-07-17 showed the workflow is ~85% OCR-bound (torch.conv2d in the easyocr CRAFT + docTR detection CNNs ≈ 67% of self-time; CLIP is negligible). This session shipped P2/P3/P5 (docTR-fallback gate, screenshot double-OCR dedup, CLIP text-embedding memoization) — conv2d call count dropped 1189→528 (−56%). All planned work items are shipped; the item stays listed only to monitor the P2 gate's faint-text recall tradeoff in live use.

Status: Done pending monitoring — P1 shipped 2026-07-18 (gate on by default at K=3); P2 gate shipped, and its escape hatch (--ocr-doctr-fallback) shipped 2026-07-25. Remaining: watch for faint-text misses in live runs; no code work open. Priority: P3 Source: content-classification profiling + P2/P3/P5 optimization session, 2026-07-17

  1. Gate OCR on text-likelihood (P1) — Done. OCR_CLIP_GATE_TOPK = 3 constant added to src/scoring/weights.py; gate enabled by default at K=3 in ContentOrganizer, ContentBasedFileOrganizer, and the CLI. Opt out with --ocr-clip-topk 0 (or ocr_clip_topk=None); FileContext._skip_ocr_by_clip_gate treats K=0/None as disabled. Eval: K=3 → 100% text recall, ~35% of photos skip OCR. Gate fails open when CLIP is unavailable. 15 unit tests added to tests/unit/scoring/test_context.py::TestClipOcrGate. Reusable tooling: scripts/profile_pipeline.py (hot-path profiler) + scripts/eval_ocr_gate.py (folder-labeled gate eval).

  2. P2 docTR-fallback gate — recall tradeoff to monitor. The shipped gate (extract_ocr_with_confidence: skip the docTR fallback when easyocr cleanly finds no text) was eval'd over 7 text images at varying difficulty: 1/7 recall loss — very-low-contrast text (easyocr's detector found nothing; docTR would have caught it). Clean, dark-mode, and rotated text were all gate-safe. So P2 trades a rare miss on near-invisible text for eliminating the docTR fallback. For a screenshot/photo-dominated 265k-file library this is very likely a net win, but it is a real behavior change — put it behind a config flag or revert if faint-text recall matters. Config flag shipped 2026-07-25: --ocr-doctr-fallback (store_true, default off = gate on; constant OCR_FORCE_DOCTR_FALLBACK in src/scoring/weights.py) forces the docTR pass after a clean easyocr negative. Plumbed force_doctr_fallback through extract_ocr_with_confidence, TextExtractor, ContentOrganizer (all 3 call sites incl. the FileContext ocr_provider), ContentBasedFileOrganizer, ContentInputs, and the CLI — same pattern as --ocr-clip-topk. 5 gate tests in tests/unit/test_shared.py::TestDoctrFallbackGate; CLI-inputs contract + integration suites pass.

DB↔filesystem provenance drift (original_path / current_path integrity)

A 2026-07-26 audit of results/file_organization.db (495 rows) found 50 rows whose current_path no longer resolved on disk. Repairing them surfaced two distinct integrity problems, neither of which is data loss, and both of which mislead any future reader of the DB (including the calibration oracle in docs/architecture/scoring-calibration-20260726.md).

Status: Path integrity resolved — 7 dead rows pruned (reconcile --prune-missing), 42 of the 43 reverted-run rows re-organized and healed (item 1), the 3 category-less Events/Burning Flipside/ rows now resolve via the looped entity-segment strip (item 3, 2026-07-27), and the flutter_auth.txt straggler repaired 2026-08-11 (item 3). 0 of 488 rows now claim a current_path that does not exist while their file is live. Open: only the 25-row staging-dir provenance record (item 2), which has no honest repair and stands as a rule rather than a task. Priority: P3 Source: DB path audit, 2026-07-26 (same session as the sprite naming-trap fix)

  1. Reverted organize run leaves rows claiming destinations that never persisted (43 rows). A batch run on 2026-06-27 04:07:41 moved 43 files from ~/Desktop/Uncategorized (42) and ~/Desktop (1) into ~/Documents/..., persisted them with status=ORGANIZED, and the moves were later reverted — the files are back at their original_path and byte-size identical to their DB records (43/43 verified), while every current_path points at a Documents path that does not exist. Persistence is gated on if not dry_run (file_processor.py), so this is not a dry-run artifact: the moves really happened and were undone afterwards. GraphStore.prune_missing_files correctly refuses to touch these (it requires both paths gone), so they survive every prune and permanently misreport where those 43 files live. Options when someone picks this up: repair current_path → original_path (record reality, and revisit status), or re-run organize-files content over ~/Desktop/Uncategorized so the rows become true. The latter is attractive because these 43 are exactly the corpus that motivated the weak-shape sprite fix — 19 of them are photos the DB still labels game_assets/sprites, which the fixed classifier now routes to media/photos_*. RESOLVED 2026-07-26 by taking that option: organize-files content --source ~/Desktop/Uncategorized re-organized 42 of the 43 (the 43rd is item 3 below). Because add_file keys on generate_id(original_path) and the files were back at their original paths, all 42 rows updated in place — no duplicates. The sprite fix held: zero files routed to game_assets/sprites; 36 photos went to media/photos_{social,other}, the logos to Organization/Integrity Studio, and Burning_Flipside_Map.pdf to Events/Burning Flipside/.
  2. Agent temp-staging dirs must never be the organize source (25 rows). 25 rows record original_path inside ~/.claude/jobs/e184540b/tmp/interior_apply_stage/ — an agent job's temp staging directory, which still exists but is empty. 23 of the 25 files are fine on disk; 2 were among 6 unrecoverable Media/Interiors renders pruned on 2026-07-26. For all 25 the recorded provenance is an ephemeral path, and the true pre-staging origin is unrecoverable, so there is no honest repair — only this rule: an agent run must not organize files out of a temp/staging dir into the production graph store, because original_path then permanently records a path that ceases to exist and provenance is destroyed. Stage into a durable location, or organize from the user's real source directory.
  3. Residual rows the 2026-07-26 sweep deliberately left (4 rows).
    • flutter_auth.txt — the 43rd reverted-run file, sitting at ~/Desktop root rather than ~/Desktop/Uncategorized, so it was outside the source the re-organize covered. FIXED 2026-08-11 — but by neither route this item proposed, because both were wrong. The row is now current_path=NULL, status=PENDING, organized_at=NULL, with organization_reason recording the reverted run; the file is untouched at ~/Desktop/flutter_auth.txt (23,596 bytes, matching the row).
      • reconcile has no path-repair operation. This item named one as if it existed. cmd_reconcile implements exactly --set-category, --prune-missing and --backfill-categories — none of which can rewrite a path. Check the subcommand before citing it as the cheap route.
      • The --source ~/Desktop --limit 1 pass was not viable either. BatchProcessor.scan_directory is rglob("*"), so that source is 751 files recursively, not the 10 at the top level, and --limit 1 takes an arbitrary rglob entry rather than the intended file — a scattergun that would move unrelated Desktop files into ~/Documents. There is no single-file --source.
      • current_path=NULL rather than = original_path, which is how this item phrased the option. add_file never sets current_path, so NULL is the never-organized shape, and the current_path or original_path readers (graph_store.py:1288, :1364) resolve it to the Desktop path either way. Writing the original path in instead would make the four sites that treat a non-NULL current_path as "organized" (:988, :1021, :1448) report a file that was never filed as filed — trading a wrong path for a wrong status.
      • Verified that NULL does not make it prunable: prune_missing_files requires both paths gone, and original_path exists, so reconcile --prune-missing still reports the same 6 unrecoverable rows and leaves this one alone.
      • Its business/planning category edge was deliberately left in place — the edge records the classification, and this repair is about path provenance. A file at ~/Desktop root has no taxonomy ancestor, so resolve_taxonomy_folder would report it unresolved rather than propose a better one.
    • 3 rows under Events/Burning Flipside/ have no category edge FIXED 2026-07-27. Folder lookup moved to resolve_taxonomy_folder (graph_store.py): exact match first, then trailing segments stripped in a loop until an ancestor is in the taxonomy, so entity-named folders resolve at any depth (Events/{Name}/2026/maps → Events, Media/Interiors/{Prop}/{Room} → Media/Interiors). A parent reached by stripping that declares no subcategory is filed under the generic bucket — Events/* → events/other — while an exact match keeps the bare category, so a file sitting directly in Events/ still reads events. Live dry-run: the 3 Burning Flipside rows now resolve events/other, 0 unresolved of 3 orphaned (was 3 unresolved). 3 new tests in tests/unit/test_category_identity.py (deep nesting, exact-match preservation) + the existing entity test updated; 224 storage/pipeline tests pass. Applied to the live DB 2026-07-27 (backup …bak-20260727_192729): 3 edges attached, category-less rows 0 of 488 — closing the orphan count 125 → 3 → 0. As-intended consequence, documented in CLAUDE.md and the method docstring: the pair follows the folder and nothing else, so two copies of one document in two trees get two individually-correct edges (Documents/Events/Burning Flipside/… → events/other; Documents/Personal/Events/…_300dpi.png → personal/events), and a category query returns a subset of the document family. That is filing, not drift — the backfill must not infer which copy is canonical. Likewise Organization/{Name} → organization/vendors follows the taxonomy's declaration of that folder as the vendor/partner root.

scripts/d1/schema.sql silently drifts from the model

scripts/d1/schema.sql is generated from Base.metadata by scripts/d1/generate_schema.py, and its own header says "the output file is the authoritative D1 schema — do not hand-edit it. Edit src/storage/models.py instead, then re-run this script." Nothing enforces the second half. Regenerating it on 2026-07-27 for the file_count removal revealed it had been stale across three earlier model changes that nobody regenerated:

  • ix_categories_full_path was still plain and ix_categories_name still UNIQUE — i.e. the D1 schema still carried the exact index shape whose UNIQUE-on-name bug dropped category edges for 26% of rows (fixed locally 2026-07-26, never regenerated).
  • people was missing review_status, detection_confidence, validation_scores, validated_at and ix_people_review_status (the person-validation work).
  • file_categories was missing signal_evidence (UNIFIED_SCORING_PLAN §5.4).

A D1 load against that schema would have failed on the missing columns or silently recreated the fixed UNIQUE bug in the mirror. Nothing consumes those columns today — workers/file-org-api has no reference to any of them — so no live breakage resulted, but the drift is unbounded and invisible until someone deploys.

Options: a test that regenerates into a temp file and asserts it matches the committed copy (cheapest, catches it in CI); a pre-commit hook on src/storage/models.py; or generating it at deploy time and not committing it at all. The first is the obvious fit — the repo already has hand-rolled golden-file comparisons (tests/unit/golden/).

Status: Done 2026-08-11 — the first option shipped (tests/unit/test_d1_schema_drift.py). Priority: P3 (no live consumer today; blocks or corrupts a future D1 deploy) Source: file_count cache removal, 2026-07-27

  • Shipped as the option this item recommended: the test regenerates in memory (no temp file needed — generate() already returns the string) and diffs against the committed copy. Local: make schema-check / npm run schema:check. CI: the schema-drift job in .github/workflows/checks.yml, so it fails on the commit that causes the drift rather than at deploy.
  • Verified against the exact historical failure, not just written: adding a column to the Person model without regenerating fails with schema.sql is stale, naming the first differing line and the line-count delta. Then make d1-schema repairs it and the check goes green — the full detect → repair → pass loop was exercised end to end.
  • The failure message has to carry the fix, because the fix is not guessable. A byte-diff of two ~10 KB SQL blobs is unreadable, so the assertion reports the first differing line plus Run \python scripts/d1/generate_schema.py` and commit the result.` A drift guard whose output does not say how to regenerate gets "fixed" by hand-editing the generated file — which the header already forbids and which this test also catches.
  • Two guards protect the guard itself. test_generate_is_deterministic pins the table/index sort ordering: a non-deterministic generator would make the drift test flake, and a flaky test gets deleted rather than fixed. test_generator_does_not_write_on_import pins the if __name__ == "__main__" guard — without it, importing the module during collection would rewrite schema.sql and the drift test would pass by silently repairing the very thing it exists to detect.
  • CI installs requirements-schema-check.txt (SQLAlchemy + pydantic + pytest), verified sufficient in a clean venv — importing src.storage.models pulls no torch, Pillow or doctr, so the job stays cheap. Unlike the lint pins, these are deliberately unbounded: the check asserts generator and file agree, and both move together under a SQLAlchemy upgrade, so a DDL-rendering change should surface as a regenerate-and-commit rather than hide behind a pin.
  • Not chosen: the pre-commit hook. It would not survive a fresh clone (hooks are not versioned), and the existing pre-commit hook in this repo deliberately only warns — making this one blocking would be inconsistent with it. CI is the durable surface.

Persisted-text PII redaction is medical-only and best-effort

organize-files content redacts files.extracted_text and schema_data["text"] for files classified medical (src/analyzers/text_redaction.py, added 2026-07-26). Two deliberate limits are worth revisiting rather than forgetting:

  1. Category-gated, not content-gated. MEDICAL_TEXT_REDACTION_CATEGORIES is {"medical"}, so a health document the scorer files elsewhere stores raw text. This is not hypothetical: during the medical ingestion the scorer wanted personal/contacts for 5 of the 11 files and uncategorized for another — only the folder-truth category override kept them inside the redaction gate. Consider gating on detected content (medical vocabulary / KIE field classes) rather than the winning category, and extending to the other sensitive categories the QUICK_START tips call out.
  2. Only numeric/email/known-name PII is masked. Emails, digit runs of 2+, and the decision's detected person names are masked; alphabetic PII outside that set — conditions, medications, third-party names the entity detector missed — survives verbatim. Same known limitation as scripts/redact_pii.py's raster path, and the reason --no-db remains the guidance for genomics and similar.

Also unquantified: how many organized files are missing from the DB entirely. The Medical/ audit found 12 of 17 on-disk files untracked (filed manually or by the DB-free name/type organizers, which persist nothing by design). No census was run over the rest of ~/Documents, so the true coverage of the graph is unknown.

Partially fixed 2026-07-26 (3b90a08). The category gate was widened:

  • MEDICAL_TEXT_REDACTION_CATEGORIES → TEXT_REDACTION_CATEGORIES (same value {"medical"}); backward-compatible alias retained.
  • New TEXT_REDACTION_SUBCATEGORY_PAIRS = {("personal", "identification"), ("personal", "records")}.
  • New should_redact_text(category, subcategory) helper checks both sets.
  • _persist_to_graph_store now uses should_redact_text instead of the bare category check.
  • 12 tests (6 new) in tests/unit/test_text_redaction.py.

Still open: content-based gating (a medical document misclassified as uncategorized still stores raw text); alphabetic PII (conditions, medications) is not masked; untracked-file census over ~/Documents has not been run.

Status: Partially fixed — category pairs extended; content-based gating and alphabetic PII remain. Priority: P3 Source: medical text-redaction work + Medical/ folder audit, 2026-07-26

Lint debt: pre-existing flake8 findings

mypy is clean repo-wide as of 2026-07-26 (2dbc148), but flake8 src/ scripts/ tests/ reports 540 findings across the tree (config: .flake8, max-line-length 100, extend-ignore E203/W503). Every finding encountered during the type pass was verified pre-existing (checked against git stash), so this is long-standing debt, not new.

By code: E111 indentation-not-a-multiple-of-four (249), E501 long lines (151), F401 unused imports (44), E402 module-import-not-at-top (21), E128/E131 continuation-line indent (22), F841 unused locals (15), E114/E302/W293 (23), plus a scatter including E741 ambiguous variable name 'l' (tests/integration/test_schema_org_export_e2e.py).

By file, the debt is concentrated — five files hold ~55% of it: src/api/schema_org_models.py (162), scripts/d1/export_to_d1.py (49), src/cost_roi_calculator.py (30), tests/unit/test_image_metadata.py (28), scripts/image_content_analyzer.py (26).

The E111/E114 bulk is 2-space-indented code that predates the project's 4-space convention, so black would fix most of it mechanically — but reformatting schema_org_models.py (the Pydantic API surface) and export_to_d1.py in one sweep is a large diff that should land on its own, not mixed with behaviour changes. Suggested order: run black over the five concentrated files, then clear F401/F841 (safe deletions), then hand-wrap the residual E501s. E402 in scripts/profile_pipeline.py and file_organizer_content_based.py is deliberate (sys.path setup precedes imports) and should get # noqa: E402 rather than reordering.

Status: Done 2026-08-11 — flake8 src/ scripts/ tests/ is at 0. Priority: P3 (no behaviour at stake; mechanical but a large diff) Source: mypy cleanup pass, 2026-07-26

  • The count had already drifted before the work started: 540 → 498. The per-code breakdown above is the 2026-07-26 census and does not match what was actually fixed. Treat an enumerated worklist in this file as a snapshot, not a contract — re-run the tool before sizing the job.
  • f905953 — the black pass. 109 files, 498 → 72. Cleared E111 (254), E128 (18), E114 (9), E302 (8), W293 (6), E131 (4), E305 (3), E261 (2) and one each of E306/E303/E129/E124; E501 150 → 32. pyproject.toml already set black's line-length to 100, matching .flake8, so the two agree on E501 — no config change was needed.
  • a8e1d72 — the 72 black cannot fix, in the order suggested above. F401 (15) deleted; E402 (19) + F811 (1) took # noqa; E741 (2), F841, F541, E231 fixed directly. E501 (32) hand-wrapped.
  • The regex splits were verified, not eyeballed. Seven long patterns in entity_detector.py were split with implicit raw-string concatenation (the idiom that file already used for its "Name on Card" pattern). All 35 compiled patterns across company/people/relationship/sentence were snapshotted before and after and compared equal. A regex split is a behaviour change until proven otherwise — a mis-placed break inside a character class or escape is silent and would only surface as a classification miss.
  • Seven E501s were caused by a trailing # type: ignore[...], not by the code. Hoisting the ignored expression to its own line (or parenthesising the import with the ignore on the opening line) keeps mypy clean, which is the proof the ignore still lands on the line that reports the error. Moving one to the wrong line silently disables it and leaves mypy green only until the underlying error moves.
  • E402 sites are all deliberate bootstraps — sys.path setup or pytest module stubs that must precede their imports — so # noqa per line, as this item originally called for. Reordering them would break the bootstrap.
  • Verification: 2427 passed / 2 skipped (unit), 192 (integration), 14 (performance), mypy clean on 148 files, organize-files health 12/12, compileall clean. scripts/d1/schema.sql regenerates unchanged, confirming the models.py edits do not reach Base.metadata.
  • Nothing prevents the next drift. Gate shipped — .github/workflows/checks.yml (job lint) runs black --check --diff + flake8 over src/ scripts/ tests/ on every push and PR, with make lint / npm run lint as the identical local command so a green run locally means a green CI job. Verified to actually fail, both legs independently: a misformatted-but-valid function exits 2 on black, and an unused import that black accepts exits 2 on flake8 — neither tool subsumes the other.
  • The version bounds are the subtle part, and an unpinned gate would have rotted. black changes its stable style once a year at the first release of a new major, so a plain pip install black in CI would eventually demand a reformat of correctly-formatted code and fail on an untouched tree. Bounded <27 in both requirements-lint.txt (what CI installs) and the dev extra; keep them in step.
  • CI installs requirements-lint.txt, never pip install -e ".[dev]" — the base dependencies pull torch via python-doctr, which would dominate a lint job. This is also why mypy is not in the gate: it needs the full dependency set to resolve imports, so it stays a local check.
  • Still ungated: scripts/d1/schema.sql regeneration. Closed the same day — it now has its own schema-drift job in the renamed .github/workflows/checks.yml (the file was lint.yml for one commit; it grew a second job, so the name no longer fit).

Near-dupe report is not wired into reconcile, and its scale numbers are synthetic

organize-files find-duplicates shipped 2026-08-10 (src/similarity/, 30 tests) and stands alone: src/cli.py:455 registers cmd_find_duplicates, and nothing in cmd_reconcile consumes its groups. That leaves the read-side view unbuilt for the case that motivated the feature. reconcile --backfill-categories deliberately answers "where does this file sit", never "what is this file", so it files a split document family to two different categories by design (Documents/Events/Burning Flipside/PlacementMap.pdf → events/other; Documents/Personal/Events/PlacementMap_300dpi.png → personal/events). Both edges are individually correct, and the split is intended — but today it is also silent. A near-dupe report is what would make it visible.

Status: Open — feature shipped, integration and scale verification did not. Priority: P3 Source: faiss/SSCD near-dupe build, 2026-08-10; re-opened 2026-08-11 after /backlog-migrate moved the parent items to docs/changelog/2.3.0/CHANGELOG.md with these residuals recorded only as changelog prose

  • End-to-end verification was on 8, 3 and 3 files. No real full-corpus run has been done. Everything below the small-corpus level is projection.
  • The 265k figures are synthetic, and one of them is a schedule risk. IndexFlatIP at 265,000 × 512 measured 543 MB resident with 5,000 queries k=5 in 1.13 s (→ full self-kNN ≈ 60 s brute force) — but on random vectors, not descriptors. The cost that matters is the descriptor pass: ~45 ms/image implies roughly 3.3 CPU-hours to describe 265k files cold. Budget that before promising a full-corpus run; it is the reason this has not simply been run once.
  • Do not reach for IVF/PQ when you do wire it. Exact flat search is already fast enough at this corpus size; an approximate index would add recall tuning for no gain. Recorded because "265k vectors" reads like an ANN problem and is not one.
  • Integration shape is a read-side view, not a scorer signal. Consuming groups inside reconcile is the open work. A SimilarityIndexSignal inside the unified scorer is a separate, later question needing its own weight calibration — do not conflate them.

find-duplicates similarity threshold is unvalidated on real data

DEFAULT_SIMILARITY_THRESHOLD = 0.85 (src/similarity/constants.py:58) was chosen to keep a review queue readable, not measured against a labelled duplicate set. The only evidence behind it is synthetic: transforms of a single image put the true group at 0.926 with no false pairs. That establishes the transform survives the descriptor, and says nothing about precision or recall across the real corpus.

Status: Open — number is a placeholder wearing a constant's clothes. Priority: P3 Source: SSCD descriptor work, 2026-08-10; re-opened 2026-08-11 (see the item above for why)

  • The failure modes are asymmetric. Too low floods the review queue and the feature gets ignored; too high silently reports nothing and looks like a clean corpus. The second is the dangerous one because it is indistinguishable from success.
  • A labelled set is cheap here — the split-document families the backfill already produces are known-positive pairs, and find-duplicates output on the real corpus supplies candidates to hand-judge. This does not need a public benchmark.
  • First-page-only PDF rasterisation bounds what can be validated. Multi-page documents match on their cover, which is right for "same document" and wrong for "same content buried on page 7" — so a labelled set should record which kind of match it is testing.

Scoring calibration is corpus-bound: 164 labelled non-media rows

make weight-search (nevergrad, joint search over 19 priors + both thresholds) ran 2026-08-11 across NGOpt/CMA/TwoPointsDE at budgets 120–250 and found nothing: train non-media agreement never moved off the shipped 59/164, and every best candidate was flat or −1 on the holdout. The correct reading is not "the weights are optimal" — it is that 488 rows / 164 non-media with a stored label cannot resolve the difference, and the oracle itself carries manual corrections and pre-unified placements.

Status: Open — the binding constraint is labelled data, not optimiser budget. Priority: P3 Source: nevergrad joint weight search, 2026-08-11; re-opened 2026-08-11 (see the near-dupe item above for why)

  • More budget is the wrong lever and will look like progress. A larger budget on this corpus buys more overfitting to oracle noise, and the holdout guard will keep reporting flat-or-−1 — which is the guard working, not a tuning problem.
  • Media rows are excluded for a reason. They disagree for replay-fidelity reasons rather than weight reasons, so the non-media slice stays the reported objective; an optimiser scored on the overall slice chases fidelity artifacts.
  • Adding labels means adding correct ones. Some stored labels are simply wrong — inspect winning_signals and OCR text on actual rows before treating a disagreement as a signal defect. A larger corpus of the same biased oracle does not help.
  • make golden (43/43) stays the gate for any re-tune this eventually enables, and any accepted change still ships as a calibration doc + backtest report. The script never writes weights.py by design.

make calibrate and the goldens are structurally blind to PhotoCompositionSignal

CLAUDE.md names make calibrate + make golden as the gate any scoring change must pass. For one of the 19 registered signals that gate cannot fail, because both harnesses construct the registry with image_analyzer=None — scripts/backtest_scoring.py:486 and tests/integration/test_unified_scoring_golden.py:121, the latter with the explicit comment "PhotoCompositionSignal gates itself off". applies_to requires a non-None analyzer, so the signal never runs on any row, and signal_wins in results/backtest_report.json confirms it: photo_composition is absent entirely across all 488 replayed rows while the other 18 signals accumulate wins.

Found 2026-08-11 running the gate on the people-gate rank conversion (d699f45, d5c308c). Before/after agreement was byte-identical — 269/488 = 0.5512, decision states 469/15/4 — measured against a worktree at the pre-change commit using the same post-backfill DB. That looks like a clean no-regression result and is not one: the harness cannot observe the code that changed.

Status: Open — the gate passes unconditionally for this signal, so a real regression in it would ship green. Priority: P3 Source: make calibrate run for the people-gate rank conversion, 2026-08-11

  • The dangerous property is that it fails silently in the safe-looking direction. A gate that errors gets fixed. A gate that returns "no change" for a change it cannot see gets believed — and the more confident the surrounding process (backup taken, before/after compared, EXIT=0, 45 goldens green), the more convincing the wrong conclusion. Every one of those steps was performed here and none of them could have caught it.
  • image_analyzer=None is not obviously wrong at either call site, which is why it survived: the replay has no live analyzer and the goldens want hermeticity. The fix is not to wire a real ImageContentAnalyzer — it would need real pixels and CLIP at replay time. It is to inject a stub analyzer driven from stored data, the same shape as make_screenshot_ocr_stub (backtest_scoring.py) already uses for ScreenshotOcrSignal. That precedent exists in the same file.
  • A stub needs a stored has_people/has_faces to replay from, and nothing records one today. files.image_classification holds CLIP label scores, not the cascade's verdict, and the cascade is the primary path (of 15 route changes measured over 400 images, 13 were faces-only and 2 were both — zero were CLIP-rank-only). So closing this needs a persisted face-detection result first; file_categories.signal_evidence is the natural home now that it is readable via GraphStore.get_signal_evidence (b6daa39).
  • Second, narrower blindness in the same context: image_metadata_provider is never wired. FIXED 2026-08-11. build_context now passes make_image_metadata_provider(row), which reads EXIF datetime + GPS from the row's on-disk file. 11 tests in tests/unit/test_backtest_scoring.py (TestOnDiskPath, TestImageMetadataProvider, TestBuildContextImageMetadata); goldens 45/45, lint and mypy clean.
    • The item understated the blindness: it disabled two signals, not one. ClipVisionSignal's GPS→GPS_TRAVEL_CATEGORY upgrade was named, but MediaHeuristicSignal reads the same dict for both its GPS travel branch and its EXIF-datetime "personal photos" branch (media_heuristic.py:202) — and that is where all three observed changes actually came from. When an item names one consumer of a shared value, grep for the others.
    • The stored-column route is dead, which is not obvious from the schema. files has exif_datetime, gps_latitude and gps_longitude, so reconstructing from persisted data looks like the hermetic option — but all three are empty on all 488 rows, so that provider would return {} for the whole corpus and the gate would stay just as blind while looking fixed. Reading the file is the only route; _ReplaySceneSignal already set the precedent.
    • It resolves current_path first, falling back to original_path. The replay classifies under original_path (pre-move), so a provider that read the path it is handed would find nothing for every file that was actually filed — the same trap scene_path_lookup exists to avoid. Pinned by test_reads_the_resolved_file_not_the_path_it_is_handed.
    • Deliberately not get_metadata_summary. That reverse-geocodes GPS through Nominatim, which would put a network call in the calibration gate; no signal reads location_name, year, month, date_str or text_metadata, so the provider returns only datetime and gps_coordinates. It does reproduce the PNG "Creation Time" fallback — 0 rows take that path today, but .png is a PHOTO_EXTENSION, so the branch is reachable and would otherwise diverge silently.
    • Measured effect: 3 of 488 decisions change, and agreement drops 263 → 262. Only 3 of 374 resolvable image rows carry any EXIF at all. The −1 is the oracle being wrong, verified row by row rather than assumed: IMG_8550.HEIC moves media/photos_other → media/interiors_other, which is exactly what the HEIC mtime-names item predicts a --force pass should produce ("the lost GPS degraded media_heuristic to photos_other"). The license back PXL_20220607_234355242.MP.jpg moves to personal/identification, matching its disk location and one of its two stored edges, but scores as no change because the oracle pair for that row is the known-wrong game_assets/sprites. So 2 rows got more correct and 0 got less correct while the headline number went down — do not read this metric without opening the rows.
    • Baselines in this file decay; measure your own. The −1 above is against a same-session run with the provider stubbed out, not against the 269/488 recorded earlier in this item — d5c308c and f486eba landed in between, and the current unstubbed number is 268/488 through summarize_replay's own accounting, which filters differently again.
  • Do not read this as "the other 18 are covered". It says only that they are not disabled by construction. Signals gated on text or OCR still depend on what the row happened to store, which is the separate fidelity caveat the calibration item already records for media rows.
  • Until it is closed, changes to PhotoCompositionSignal or ImageContentAnalyzer need their own evidence — the rank conversion was validated instead by unit tests, two-direction mutation tests, and corpus measurements over 300–400 images where every behaviour-changed file was judged individually. Say so explicitly in the commit, because "calibrate passed" will otherwise be read as covering it.

A/B trial: Perception Encoder (PE-Core) vs ViT-B-32 as the CLIP backbone

scripts/shared/clip_utils.py:75 pins MODEL_NAME = "ViT-B-32" with OpenAI pretrained weights. Meta's Perception Encoder (facebookresearch/perception_models, Apache-2.0 for both code and weights) is a substantially stronger CLIP-family model — and is already present in the installed open_clip 3.3.0 as a first-class architecture with a meta pretrained tag (PE-Core-B-16, PE-Core-S-16-384, PE-Core-T-16-384, PE-Core-L-14-336, PE-Core-bigG-14-448), verified 2026-08-10. So the model swap itself is a one-line change; everything expensive is downstream of it.

Status: Open — trial not started. Availability + config verified 2026-08-10, nothing measured on this corpus. Priority: P3 — largest available quality lever on the vision path, and the largest disruption. Do not swap without the A/B. Source: facebookresearch library audit, 2026-08-10

  • This is not a free swap — three things break at once. (1) PE-Core-B-16 is 1024-d vs 512-d for ViT-B-32, so .cache/clip_embeddings_v2/ is invalidated (bump to _v3; the cache is already versioned for exactly this). (2) results/scene_probe.joblib is trained on 512-d ViT-B-32 features and must be retrained — the scene corpus is the gating cost, not the model. (3) ViT-B/16 at 224px is ~4× the vision FLOPs of B/32 (196 vs 49 patches), directly against the banked OCR/CLIP cost work.
  • The 512-d PE variants are not the cheap escape. PE-Core-S-16-384 and PE-Core-T-16-384 keep embed_dim 512 (so the cache dimensionality survives) but run at 384px — more compute than B-16@224, not less. If the trial is motivated by cost, none of these win; if it is motivated by accuracy, evaluate B-16 and accept the retrain.
  • The OCR gate moves too, and it is the sneaky one. OCR_CLIP_GATE_TOPK = 3 is a ranking heuristic over CLIP labels, calibrated on ViT-B-32 softmax behaviour (the flat-softmax finding: absolute probabilities are near-uniform, only rank separates text from photos). A different backbone changes that ranking distribution. Re-run scripts/eval_ocr_gate.py as part of the A/B — a backbone that improves classification while degrading the gate could be a net cost regression.
  • Measure with the tools that already exist. scripts/profile_pipeline.py for the wall/per-file and CLIP-vs-OCR hotspot split, scripts/backtest_scoring.py for decision agreement, make golden as the correctness gate, eval_ocr_gate.py for the gate. Requires a backfill_clip_scores.py re-run under the new backbone before the replay is meaningful — the replay serves CLIP/scene votes from the cache.
  • Reported zero-shot gap is large enough to justify the trial (PE-Core-B16 ≈ 78 vs ViT-B-32/openai ≈ 63 ImageNet zero-shot, upstream numbers, not measured here) — but upstream benchmark deltas have repeatedly failed to predict behaviour on this corpus. Treat them as the reason to run the A/B, never as its result.
  • Related, deliberately out of scope: DINOv2 (Apache-2.0) has stronger linear-probe features and would be the better SceneSignal input, but has no text tower and so cannot replace the zero-shot ClipVisionSignal. That is a second-embedding question, not a backbone swap. MetaCLIP (CC-BY-NC 4.0) and ImageBind (CC-BY-NC-SA 4.0) are license-blocked — non-commercial binds the employer, same finding as the ImageNet/LLD/InfographicVQA call.

E2E dashboard assertions hardcode the card set (drifts every time a card is added)

tests/e2e/dashboard.spec.ts:42-56 asserts toHaveCount(4) against .feature-card plus four explicit .card-title filters. _site/index.html gained a fifth card ("Residence Galleries") in 205bd7e; the count assertion then failed while the title assertions stayed silent — they covered 4 of the 5 cards, so the new card shipped with zero content coverage and only the count noticed it existed. Numbers were corrected in 58f7659 (4→5, added the missing Residence Galleries title assert), but the shape is unchanged: the next card added breaks the suite again and lands untested until someone hand-edits the spec.

Status: Done 2026-08-11 (d97be75 cards, 1013673 stats bar + exact matching). Priority: P3 Source: npm run test:e2e failure triage, 2026-08-11

  • The failure is amplified 5×. The assertion lives in one test but runs under chromium/firefox/webkit/mobile-chrome/mobile-safari, so a single stale literal reports as 5 failures and reads like a systemic break rather than one number.
  • Cheapest durable fix: hoist the expected titles into one fixture array, then assert toHaveCount(EXPECTED.length) and loop the title checks. Adding a card becomes a one-line data change that cannot leave a card uncovered.
  • Keep both assertion kinds. The count is the only thing that catches an added card; the title filters only catch a removed or renamed one. Dropping either reintroduces a blind spot.
  • The card set has changed without human intent before, which is the real case for the count assertion: the resolved copy_to_site.sh item records a content run silently deleting this same "Residence Galleries" card from _site/index.html on 2026-07-26, orphaning _site/residence_gallery.html from navigation. That clobber path is fixed (1d3b262), so the count is not currently load-bearing against it — but it was the only assertion that would have caught it.
  • Shipped d97be75: DASHBOARD_CARDS ({ title, href }, DOM order) now drives the count, the titles, and the navigation tests, which are generated per entry.
  • This item understated the gap: navigation was uncovered too. The write-up above says only the count noticed the fifth card. In fact the four should navigate to … tests were also hand-written, so "Residence Galleries" had no navigation test at all — the loop took the spec 4 → 5 nav tests without anyone adding one. When an item names one hardcoded list, check for its siblings before calling the fix complete; the same card set was enumerated by hand in three places, not one.
  • Both guards were verified to actually bite, in both directions: a phantom sixth fixture entry fails toHaveCount and its generated nav test; a title renamed with the length unchanged lets the count pass and fails the text assertion. Without the second check the loop could have been vacuous — a green suite asserting nothing.
  • hasText/getByText are substring matches, and that was a live hole, not a theoretical one (fixed 1013673). Injecting the strict prefix 'Residence Galler' into the committed d97be75 spec left it green. Replaced with toHaveText(array), which matches each element's full text, and the list in order — so DOM order is now enforced rather than asserted in a comment. The same injection now fails.
  • The stats bar had the identical shape and got the same treatment in 1013673: DASHBOARD_STATS replaces toHaveCount(4) plus four page-wide getByText calls, now scoped to .stat-label.
  • Why both assertion kinds survive, restated precisely: the counts run over .feature-card/.stat-item (containers) while the text runs over .card-title/.stat-label (their children). A container with no label, or a stray label outside a container, fails exactly one. toHaveText alone would not catch either.
  • Still hand-enumerated: the three .tech-badge filters. Fixed 231019a, and the bullet above undercounted the gap. It called this "lower-stakes"; the footer actually has 8 badges, the test asserted 3 by substring with no count assertion at all, so 5 were entirely uncovered and a removed badge failed nothing. TECH_BADGES + toHaveText(array) now pins the set exactly and in order. Verified by deleting the PyTorch badge from the page: fails now, passed before. A partial assertion with no count reads as coverage and is not — the card set at least had a count telling the truth about how many existed.
  • Unfixed and outside the test suite: _site/index.html has the same bug in CSS. Fixed 231019a, and the "regenerated by content runs" claim above was wrong. _site/index.html is the committed dashboard source: copy_to_site.sh skips it explicitly ("_site/index.html is the maintained source") and update_site_data.py only rewrites .stat-value numbers and the Last Updated date via targeted regexes. So a direct edit is safe and there is no generator to push the fix into.
  • The CSS fix alone would have preserved the defect's shape, so the test is the fix. Rules now run past the current card count, but new test every feature card has its own entry-animation delay is what makes that safe: it reads computed animationDelay per card and compares against (index + 1) * CARD_STAGGER_STEP_SECONDS derived from DASHBOARD_CARDS. Verified against the pre-fix CSS — reports 0s where 0.5s was expected, i.e. it catches the exact bug that shipped in 205bd7e.
  • Why this one hid for so long: it is presentation-only state no other assertion can observe. A missing animation-delay breaks no locator, no navigation, no text. :nth-child enumerations are invisible to ordinary E2E coverage and need an explicit computed-style assertion, which is the transferable lesson — the footer badge list and the card set were both catchable by normal assertions; this was not.
  • A count-agnostic CSS formulation was considered and rejected. sibling-index() is not available in webkit/firefox, which this suite runs, and an inline --card-index custom property is ruled out by the project's no-inline-styles convention. Per-card rules are therefore unavoidable today; revisit if sibling-index() lands across the matrix.

CLIP scores are a softmax over raw cosines — no logit scaling

CLIPClassifier._similarities (scripts/shared/clip_utils.py:302-303) computes sims = img_emb @ txt_norm.T then sims.softmax(dim=0), omitting the model's trained logit_scale (≈100 for ViT-B-32). Softmaxing unscaled cosines (~0.15–0.35) is necessarily near-uniform: over the 94-label photo vocab the floor is 1.064% and winners land at 1.13–1.16%, i.e. ~8% above chance. This is the "flat-softmax" property already relied on in several places (W_CLIP as a ranking-only signal, CLIP_OCR_FALLBACK_THRESHOLD, _SCREENSHOT_OCR_KEYWORD_THRESHOLD, graphic_detection's 0.077 floor, the OCR_CLIP_GATE_TOPK note under the PE-Core item) — this entry records its cause and why it was not fixed.

Status: Open — deliberately not fixed; the relative-margin gate in 4b56759 works around it instead. Priority: P3 Source: label-accuracy investigation of organize-files content, 2026-08-11

  • Applying logit_scale is the principled fix and a wide blast radius. It would make absolute confidences meaningful for the first time, and simultaneously invalidate every threshold calibrated against the flat distribution — at least the five gates listed above plus CLIP_REFINEMENT_MIN_CONFIDENCE/ACCEPT_CONFIDENCE (0.15/0.30), which no current score can ever reach, so refinement is presently dead code on the photo profile. Sequence it as its own change with make calibrate + goldens, never bundled.
  • The workaround deliberately does not touch the distribution. CLIP_MIN_LABEL_MARGIN_RATIO = 1.012 gates on top-1/top-2 ratio, which is scale-invariant and so survives a later logit-scale fix unchanged (the ratio of two softmax outputs changes, but the ordering-based gate keeps its meaning). Contained, but it only suppresses bad names — it cannot produce good ones.
  • Confirm against the backbone A/B before investing. The PE-Core trial above would change this distribution anyway; doing both at once makes neither attributable.
  • The blast radius above was incomplete: four more thresholds, in src/analyzers/image_analyzer.py, and they were already dead. Found 2026-08-11 while converting the people gate to rank. _PEOPLE_SCORE_LOW_THRESHOLD (0.15), _PEOPLE_SCORE_THRESHOLD (0.2), _INTERIOR_SCORE_THRESHOLD (0.3) and _SCREENSHOT_SCORE_THRESHOLD (0.4) are all measured unreachable — the global max across all labels is 0.1003 over 39 images against a 1/N floor of 0.0909, so they miss by 1.5×–4.0×. has_people had silently degraded to has_faces alone with two no-op clauses around it, and is_home_interior_no_people is permanently False.
  • This is where the git history of the flat softmax actually starts, and the cause is a refactor, not a design choice. 681c5ed (2025-12-09) scored via outputs.logits_per_image.softmax(dim=1) — logits_per_image is logit_scale × cosine — and the inline comments read as genuine percentages ("30% confidence threshold", "20% confidence"). 33264df (2026-03-27) swapped transformers for sentence-transformers and, in its own words, removed "~60 lines of manual tokenization, logit_scale arithmetic, and processor boilerplate" while claiming to "keep exact public API". The API was kept; the value distribution behind it was not. Two days later ff36546 extracted the now-unreachable numbers into named constants, which made them look deliberately calibrated. A refactor that preserves signatures can still invalidate every threshold written against the values, and nothing in the type system or the test suite notices.
  • Restoring logit_scale silently reactivates them — quantified. On the same 25 photos, people_score goes from 0.0903–0.0958 (0/25 above 0.15) to 0.0000–0.8896 (7/25 above 0.15). So the fix would wake a branch that has not executed since March and start routing ~28% of photos toward media/photos_social. That is a larger behaviour change than anything in the five gates originally listed here, and it arrives with no error and no failing test.
  • Two of the four are now immunised; two are not. The people and screenshot gates were converted to rank tests in d699f45 — softmax is order-preserving, so they are invariant to the scaling and will not move when it is restored. _INTERIOR_SCORE_THRESHOLD was deliberately left dead (converting it would resurrect a heuristic SceneSignal replaced on 2026-07-18 and re-route photos to property_management/other), and _PEOPLE_SCORE_THRESHOLD lives in the caller-less is_home_interior_no_people. Both still need handling in the same change that applies logit_scale — decide them deliberately rather than letting the scaling decide for you.
  • The test suite passed throughout the dead period, and that is the transferable lesson. tests/unit/test_image_analyzer.py fed synthetic scores of 0.5 and 0.175 to a 0.15 threshold — values the real distribution cannot produce (max 0.1003). Green tests over a branch live data could never take. When a threshold's fixtures are hand-written numbers, the suite proves the comparison works, never that anything can reach it.

all_scores is scored from unprefixed prompts and can rank differently than the decision

classify_image (scripts/shared/clip_classification.py) picks the winning category from score_embedding(emb, labels, "a photo of ") but then returns all_scores from score_embedding(emb, labels, "") — a second, unprefixed scoring pass. The two disagree in practice: PXL_20260723_231733185.jpg was categorized living room while all_scores' argmax was dining room, and IMG_9421.HEIC was backyard against an all_scores argmax of patio. So any consumer that reads all_scores to explain, validate, or re-derive a decision is reading a distribution that did not make it.

Status: Done 2026-08-11 — d4efff5. Both distributions are now explicitly named in CLIPResult. Priority: P3 Source: margin-gate implementation, 2026-08-11

  • Shipped (d4efff5) — second option: keep both, name them distinctly. CLIPResult gains decision_scores (the "a photo of " prefixed pass that chose the winner) and raw_scores (the unprefixed pass). all_scores becomes a deprecated @property alias for raw_scores — carries the same value it always held, so external callers whose behaviour is unchanged; zero in-repo readers remain. Every in-repo call site was migrated to the explicit name.
  • Consumer audit:
    • result["top_scores"] in rename_images.py — migrated to clip_result.decision_scores. Top-5 now reflects the same distribution that picked the label, consistent with the winning category and confidence.
    • files.image_classification column via scripts/backfill_clip_scores.py — no computation change needed. CLIP_CATEGORY_PROMPTS are full prompts built by _make_clip_prompt (they already embed "a photo of " per label). prompt_prefix="" + CLIP_CATEGORY_PROMPTS IS the decision distribution — semantically equivalent to what the production _clip_scores_for_context path does. Added a comment in backfill_clip_scores.py making this explicit. Renamed local variable to decision_results. No mixed-provenance consequence: all existing rows were already written using the decision distribution; nothing stored was the unprefixed-label distribution. No re-run needed.
    • Calibration harness (backtest_scoring.py) — reads image_classification (full-prompt decision distribution, consistent with production ClipVisionSignal). The oracle's known limitations (164-row biased corpus) remain unchanged; cross-reference the "Scoring calibration is corpus-bound" item. The backtest does not replay the renamer's decision_scores/raw_scores — those are separate label sets for a separate pipeline.
  • 8 new tests in TestCLIPResultScoreFields — field shape, alias contract, different-ranking guarantee, classify_image two-call structure with correct per-pass prefix.

Screenshot renamer is ungated pending a labelled margin eval

RenamerProfile.min_label_margin defaults to 1.0 (disabled) and only PHOTO_PROFILE sets it to CLIP_MIN_LABEL_MARGIN_RATIO. The threshold was calibrated on 8 hand-labelled photos and does not transfer as-is: on a random 20-file sample from ~/Documents/Media/Photos/Screenshots, 75% fall below 1.012 while their chosen label still matches the folder they were previously filed into (terminal→Terminal, settings→Settings, at ratios as low as 1.0000). Enabling the gate there on today's evidence would suppress three-quarters of screenshot renames for no demonstrated accuracy gain.

Status: Open — intentionally left at 1.0; needs its own labelled corpus. Priority: P3 Source: margin-gate calibration, 2026-08-11

  • The folder agreement is circular evidence — those folders were produced by this same classifier, so it cannot distinguish "the label is right" from "the label is consistently wrong". A real eval needs hand labels, as the photo threshold got.
  • Screenshots mostly bypass the CLIP label anyway. analyze_image prefers an OCR title snippet (Screenshot_<title>) and only falls through to generate_clip_filename when no line qualifies, so the gate would apply to a minority of screenshots — measure that slice, not the whole folder.
  • A smaller vocab may want a different number, not the same one: 36 labels put the uniform floor at 2.78% versus the photo vocab's 1.06%.

HEIC OCR recall was never measured

The HEIC decode hand-off shipped 2026-08-11 (e9fb0a8 + 41ee326 + 1c32fe7), and the spurious _PREPROCESS_AVAILABLE gate that made docTR still raise without the preprocessing deps was removed the same day (7899442, 12 tests in tests/unit/test_ocr_heic_decode.py). One residual remains, and it is a measurement gap rather than a defect: the path is verified by shape, never by output.

Status: Done — measurement tooling shipped c266ace + 8cd1c7a (2026-08-11). A labelled HEIC corpus must still be assembled and the tool run; see the final residual below. Priority: P3 Source: /backlog-implementer hand-off verification, 2026-08-11; re-opened 2026-08-11 after /backlog-migrate moved the parent item to docs/changelog/2.3.0/CHANGELOG.md with these residuals recorded only as changelog prose

  • The docTR branch is gated on an unrelated capability flag. Fixed 7899442: removed _PREPROCESS_AVAILABLE from the condition in _run_image_ocr. The decode uses _load_rgb (PIL only), not preprocess_for_ocr; _load_rgb's except Exception guard returns None when PIL is absent, so the HEIC early-return fires instead of falling through to DocumentFile.from_images. 12 tests (2 new regression guards).
  • The np.asarray guard comment was misleading. Fixed c266ace (2026-08-11): the old comment said "PIL available but numpy not installed" — a state that cannot occur because numpy is imported before PIL in the same try block. Replaced with a defence-in-depth note explaining the import-order coupling and why reordering would make the guard reachable. Import-order note also added to the _PREPROCESS_AVAILABLE block so the coupling is visible at both ends. Test docstring updated to match. The guard itself is kept: it is tested via monkeypatch and costs nothing to retain.
  • A type regression rode in on the follow-up commit. f8aa49b added the DocTRResult Protocol and return hints, which tightened _run_image_ocr's signature without constraining the value satisfying it: _get_predictor() is untyped (docTR is an optional import), so predictor(doc) is Any and flowed straight to a DocTRResult | None return — mypy no-any-return, caught by the Stop hook, not by the commit's own checks. Fixed by annotating at the assignment (result: DocTRResult = predictor(doc)). Adding a Protocol to an optional-dependency module needs a check at the boundary where the untyped value enters, or the annotation is decorative.
  • Recall was never measured — tooling was missing. Fixed c266ace + 8cd1c7a (2026-08-11): scripts/eval_heic_ocr_recall.py runs both easyocr and docTR against a directory of HEIC images, bypassing the CLIP gate so raw backend recall is measured independently. Supports optional ground-truth JSON (basename→bool), per-file verbose output, and JSON results export. Run: python scripts/eval_heic_ocr_recall.py --source DIR [--ground-truth labels.json] [--json results/heic_recall.json]. The key design point: the gate bypass is intentional — measuring with the gate on would skip most HEIC photos and conflate gate decisions with decode recall.
  • One open residual: the tool has not been run. A labelled corpus of text-bearing HEICs (screenshots of text, HEIC document scans) must be assembled and the script run before a number can be reported. The precondition for that run — that OCR reaches the backends on HEIC input — was never verified before the decode fix; the script makes it verifiable.
  • Measure on text-bearing HEICs specifically, or measure nothing. The CLIP OCR gate (--ocr-clip-topk, K=3) can skip OCR before either reader is reached, so a naive before/after on a photo-heavy corpus shows no change and reads as "the fix did nothing". The script bypasses the gate for exactly this reason.
  • What the fix does buy, unmeasured: screenshot_ocr, text_content and kie_structured become available as voters on HEIC screenshots and document scans, which previously classified CLIP-only. Harmless for camera photos, which is most HEICs.

HEIC files already filed keep their mtime-derived names

The EXIF fix in 4b56759 corrects new runs, but files organized before it retain filenames built from file mtime (download time) rather than capture time, and _maybe_rename_image will not revisit them: is_generic_filename is False for an already-descriptive name like 20260516_kitchen.heic, so the rename step returns early on every subsequent run. Confirmed on the two HEICs under ~/Documents — Media/Photos/Other/20260516_kitchen.heic was actually captured 2024-04-04, a 2-year error frozen into the name; Media/Photos/Products/20240426_fabric_sofa.heic happens to be correct.

Status: Name fixed 2026-08-11 — 20260516_kitchen.heic → 20240404_kitchen.heic, matching its DateTimeOriginal of 2024:04:04 18:44:35. A full rescan of ~/Documents confirms 0 of 2 HEICs now carry a stem date disagreeing with EXIF. Open: the category residual only (first bullet below). Priority: P4 Source: post-fix verification sweep, 2026-08-11

  • Scope was rescanned, not assumed, and the item's count held. ~/Documents contains exactly 2 HEIC/HEIF files; only the kitchen one mismatched. Worth doing because an enumerated count in this file is a snapshot — the lint item's 540 had already decayed to 498 before anyone acted on it.
  • The DB carried the path in three fields, not one. files.current_path is the obvious one, but files.schema_data and schema_metadata.schema_json both embed it too, and a rename that updates only current_path leaves two stale JSON blobs pointing at a file that no longer exists. Found by sweeping every column of every table for the old basename rather than by editing the column named in this item. Do that sweep before any manual path repair — reconcile still has no path-repair operation to do it for you (same finding as the flutter_auth.txt straggler).
  • schema_data.filePath is historical and was deliberately left stale-looking. It reads ~/Downloads/20260516_kitchen.heic while files.original_path reads ~/Downloads/IMG_8550.HEIC — the two disagree correctly, because the photo-profile renamer works in place, so the file was renamed in Downloads and only then moved. filePath snapshots the pre-move path at schema-generation time (file_processor.py:275); original_path records the name before the renamer touched it. A blanket find-and-replace across the JSON would have rewritten it into a path that never existed. Only live pointers were updated: schema_data.name/contentUrl/url and schema_metadata.contentUrl (its name stays IMG_8550.HEIC, the original filename).
  • Verified no new provenance drift, by differencing against the backup. Unresolvable current_path rows went 7 → 6, the repaired one being exactly this file and the remaining 6 being the known-unrecoverable set. Backup at results/file_organization.db.bak-heicrename-20260811 (reference it by exact name — these backups sort by neither name nor mtime).
  • Still open: the category is pre-fix. The row is still media/photos_other because the lost GPS degraded media_heuristic; its own stored schema description reads "an interior room (5% confident)", and the calibrate-blindness item independently measured this exact file moving to media/interiors_other once EXIF is available. Deliberately not fixed here: there is no single-file --source, and BatchProcessor.scan_directory is rglob("*"), so the --force pass this item proposed would sweep all 25 entries of Media/Photos/Other, re-filing 23 unrelated files to fix one. Fixing the name without moving the file is the contained half; the re-file wants either a single-file source or a hand move.
  • Scale check before building anything: at two files this was a manual rename, as this item originally judged. A repair pass is only worth writing if a bulk HEIC import lands, in which case the rule is "re-derive the date prefix from EXIF when the stem's leading date disagrees with DateTimeOriginal" — the rescan snippet used here is that rule, and it is three lines.

--dry-run creates destination directories

ContentOrganizer.get_destination_path calls dest_dir.mkdir(parents=True, exist_ok=True) (src/organizers/content_organizer.py) before any dry-run check, so a preview run writes empty folders into the target tree. Confirmed by birth time: ~/Documents/Media/Photos/Travel was created at 00:08:23 by a --dry-run --no-db pass and left empty, alongside Media/Place from an earlier one.

Status: Open — untouched. Priority: P3 Source: 28-file ~/Downloads classification audit, 2026-08-12

  • The contract says otherwise. --dry-run prints "no files were moved" and QUICK_START tells users to always preview first, so the one command documented as safe is the one quietly mutating the target tree.
  • Empty dirs are not harmless here. resolve_collision and the taxonomy both read the filesystem, and reconcile --backfill-categories derives categories from what is on disk — a directory created by a preview is indistinguishable from one earned by a real filing.
  • Fix shape: thread dry_run into get_destination_path, or hoist the mkdir to the move site in FileProcessor (shutil.move is already guarded by if not dry_run). The latter is smaller and puts directory creation next to the only thing that needs it.

"Files with extracted text" undercounts every image

The run summary reports Files with extracted text: 0/28 on a batch where OCR demonstrably ran — one file's OCR text drove an ↪ OCR fallback: financial_statements (2%) decision in the same run. The counter reads the extracted_text slot returned by _detect_file_category_unified, which is ctx.text_if_loaded or ""; image OCR routes through ctx.ensure_ocr into ocr_if_loaded, a different slot. So images can never contribute to the count regardless of how much text was read.

Status: Open — reporting only, no misclassification. Priority: P4 Source: 28-file ~/Downloads classification audit, 2026-08-12

  • It cost real diagnostic time. The 0/28 was the reason a PDF-extraction bug was initially reported that did not exist — the actual defect was an early-exit short-circuit. A stat that reads "extraction is broken" when extraction is fine sends the next reader down the same path.
  • Check the persistence path separately before "fixing" the counter. _last_file_ocr_text is stashed for persistence, so files.extracted_text may well be populated for images while the summary says zero; confirm which of the two is actually wrong before changing either.

Travel classification is unvalidated — the backtest corpus has no travel photos

TRAVEL_MIN_DISTANCE_KM = 100 and the home-distance gate that replaced GPS-presence detection (2026-08-12) are unmeasured. The 488-row replay corpus predicted media/photos_travel for exactly one row before the change and zero after; agreement was 268/488 both ways. The harness therefore shows no regression and provides no positive evidence.

Status: Open — the change shipped on mechanism plus three hand-checked files, not on corpus evidence. Priority: P3 Source: home-location travel gate, 2026-08-12

  • The radius is a judgment call nobody has tested. At 100 km, San Antonio (118 km from Austin) is travel and Round Rock (28 km) is not. Nothing establishes that boundary is where a user would draw it; a labelled home-vs-trip slice would.
  • Needs photos from outside the home metro to test at all. Every geotagged row in the current corpus appears to be local, which is why the distance gate is invisible to it. Sampling a trip's worth of photos is the cheapest way to get signal.
  • Overrides exist and are untested end-to-end. FILE_ORGANIZE_HOME_COORDINATES / FILE_ORGANIZE_TRAVEL_RADIUS_KM / FILE_ORGANIZE_HOME_NAME are exercised by unit tests only; a malformed-coordinate override silently falls back to the Austin default, which is correct but unobserved in a real run.
  • Second homes and long stays break the single-point model. One coordinate pair cannot express "home is Austin, but also three weeks a year in X"; every photo from a repeated destination reads as travel forever.

Photo CLIP vocabulary has no object-level labels

A photo of a wall-mounted spice rack was labelled cabinet at a top-1/top-2 ratio of 1.0075 — below CLIP_MIN_LABEL_MARGIN_RATIO, so the rename was correctly withheld and the file went to Media/Photos/Other. The gate did its job; the vocabulary is the limit. "Spice rack" is the obviously correct description and no label in the 94-prompt photo vocab can express it, so the best available answer is a near-tie between furniture words.

Status: Open — correct behaviour given the vocabulary, wrong outcome for the user. Priority: P3 Source: 28-file ~/Downloads classification audit, 2026-08-12

  • This is the failure mode the margin gate converts, not removes. Withholding turns a wrong name into no name. That is the right trade, but a file that could have been 20240426_spice_rack.heic still lands as IMG_9436.HEIC in the generic bucket.
  • Adding labels is not free. The vocab is a softmax denominator — every prompt added lowers the uniform floor and shifts every ratio, so CLIP_MIN_LABEL_MARGIN_RATIO (calibrated on 8 photos at the current vocab size) would need recalibrating. See the "CLIP scores are a softmax over raw cosines" item.
  • Measure the gap before expanding. Count how many withheld renames are vocabulary misses versus genuine ambiguity; if most withholds are near-ties between two correct-ish labels, more labels will not help.

Duplicate-copy filename rule fires on ordinary camera filenames

classify_by_filename_patterns returns ("skip", "duplicate") for any stem matching _\d{8}_\d{6}$ — intended for timestamped copies, but IMG_20190327_172114 (the standard Android camera convention) matches it exactly. Scoring that stem directly yields skip/duplicate at confidence 1.0, early-exiting before any content signal runs.

Status: Open — currently masked, not fixed. Priority: P3 Source: 28-file ~/Downloads classification audit, 2026-08-12

  • It is masked by rename ordering, which is luck. FileProcessor._maybe_rename_image runs before detect_file_category and passes the descriptive name as display_path, so the classifier usually sees 20190327_lake.jpg and never the camera stem. Any path that classifies without renaming first — a non-generic filename, a renamer failure, a direct API caller — hits the rule.
  • The camera-timestamp fix does not cover it. CAMERA_TIMESTAMP_STEM_PATTERN is anchored at both ends and guards the sprite rules; the duplicate rule uses a suffix match and runs earlier, so IMG_/PXL_-prefixed stems still reach it.
  • Verify the intended shape first. A genuine timestamped copy is <original>_<YYYYMMDD>_<HHMMSS>, i.e. the timestamp is a suffix on an existing name. Requiring a non-empty, non-camera-prefix stem before the timestamp would separate the two cases.

Dry-run reports show colliding destinations that a live run would resolve

Two pairs of files in the 28-file run were reported with identical destinations (20260723_bathroom.jpg, 20260723_bedroom.jpg). resolve_collision only tests dest_path.exists(), which is never true during a dry run because nothing is written.

Status: Open — reporting artifact; live runs are correct. Priority: P4 Source: 28-file ~/Downloads classification audit, 2026-08-12

  • Live behaviour was verified, not assumed. get_destination_path and shutil.move run in the same loop iteration in FileProcessor, so the first file is on disk before the second's destination is computed and the suffix is applied. No data loss.
  • The report still misleads. Two of the four colliding files were genuinely different photos (a bedroom from two angles), so a reader reasonably concludes one would overwrite the other.
  • Fix shape: have the dry-run path track planned destinations in a set and resolve against it, so the preview shows the names a real run would produce.

Repo Snapshot — 2026-07-18

Recorded from the npm run repomix regeneration (docs/repomix/, gitignored) and git status on the primary checkout. Everything under "Uncommitted working tree" below is not yet committed — reconcile the affected open items when that work lands.

Token census (docs/repomix/token-tree.txt)

  • Full repo pack (repomix.xml) ≈ 983k tokens; compressed ≈ 604k; git-ranked ≈ 931k; docs-only ≈ 131k.
  • src/ ~550k, but ~332k of that is one data dir — src/classifiers/data/census_names/ (surnames.txt alone 315k tokens; given_names.txt 16k). Hand-written src/ code is ≈ 218k.
  • Other buckets: tests/ ~252k (unit ~200k, of which tests/unit/scoring/ signal tests ~43k), scripts/ ~117k (shared/ ~41k; filename_classifier.py 15.4k is the largest script module), docs/ ~110k (largest single doc: reviews/SCRIPTS_SRC_DUPLICATION_AUDIT.md 24.2k).
  • Largest source files: organizers/content_organizer.py 16.5k, storage/models.py 12.4k, storage/graph_store.py 12.2k, cli.py 8.5k, cost_roi_calculator.py 7.6k. Largest test: test_content_organizer.py 16.3k.
  • surnames.txt added to .gitignore and untracked (git rm --cached, staged) 2026-07-18. The file stays on disk locally, but once the deletion commits, fresh clones will lack it — person_name_validator depends on it at runtime, so setup needs scripts/download_census_names.py (or equivalent) to materialize it. Resolved 2026-07-25: the script exists and is committed (e50ae06), documented in QUICK_START.md §1 and the CLAUDE.md Quick Start, and verified to reproduce the shipped gazetteer exactly (162,254 Census rows → 92,357 surnames at SURNAME_MIN_COUNT=200; the script's stale "~25k" comment was corrected). person_name_validator degrades gracefully when the file is absent (_load_gazetteer → None) and organize-files health reports the gazetteer layer.

Top-churn git history (docs/repomix/gitlog-top20.txt)

  • 6f61449 (2026-07-18) — added 16 images to results/scene_labels/graphic/ (graphic-class corpus labeling for the scene probe; see the graphic-probe item).
  • eceb166 (2026-07-17) — touched census_names/surnames.txt; 5f7f563 (2026-07-01) — removed a license image (biometric PII) from results/test_set_augmentation/.
  • Six older commits (2025-12 → 2026-02) are all _site/ dashboard churn (metadata.json regenerations, metadata_viewer_backup.html a11y fix).
  • Note: both 2026-07 refactor: commits are auto-typed wrong — they are data/corpus changes, not refactors.

Uncommitted working tree — InteriorSignal → SceneSignal retirement in flight

git status 2026-07-18: 10 modified + 2 staged deletions (12 files, +373/−508):

  • Deleted (staged): src/scoring/signals/interior.py, tests/unit/scoring/test_signal_interior.py — InteriorSignal retired. SceneSignal already registers unconditionally in committed registry.py (no-ops when the artifact is absent).
  • scripts/backtest_scoring.py — the artifact-gated SceneSignal/InteriorSignal sweep-row swap collapsed to a plain ("W_SCENE", W_SCENE, "scene") row.
  • src/health_check.py + tests/unit/test_health_check.py — interior_probe health feature replaced by scene_probe (checks results/scene_probe.joblib; retrain hint now points at prototype_scene_probe.py).
  • src/pipeline/file_processor.py + tests/unit/test_pipeline.py — _IMAGE_SCHEMA_TYPES widened to every SceneSignal @type (House, Place, Accommodation join the Room family) so scene photos keep image metadata instead of falling through to DigitalDocument.
  • tests/unit/scoring/test_signal_photo_composition.py — records that PhotoCompositionSignal's interior (is_property_mgmt) vote was retired 2026-07-18; interior detection now belongs to SceneSignal exclusively.
  • tests/unit/scoring/test_registry.py — TestSceneSwap replaced by TestSceneSlot (unconditional registration; artifact pinned absent for hermeticity only).
  • Also touched: src/organizers/content_organizer.py (formatting only), tests/unit/test_content_organizer.py, tests/unit/scoring/test_signal_scene.py.

On disk but untracked: results/scene_probe.joblib (34k, trained 2026-07-18 18:37). Scene-label corpus is far past the counts recorded in the graphic-probe item: interior 178, exterior 158, place 347, neither 73, graphic 34 (item snapshot was 44/21/3/46/8–12).

Open items affected when this lands:

  • Trained graphic-vs-photograph probe — the "data-only" gap has largely closed (graphic corpus 34 and climbing; a 5-class scene_probe.joblib exists untracked). The item's "do not commit an artifact without eval + backtest" guard still applies.