Derived from session work, uncommitted changes, and codebase state.
Last updated: 2026-08-12 (added 6 items from a ~/Downloads classification audit — 28 files inspected against their actual contents, which found 6 misclassifications; three defect classes were fixed in the same session and the six items below are what was not addressed. The audit's own lesson: two of the three fixes were mischaracterised on first report and had to be corrected after measurement. "PDF text extraction returned nothing" was not an extraction bug at all — a filename rule returning ("business","legal") at full confidence early-exited the cheap wave, so no text was ever read; and "no confidence floor on renaming" was the wrong instrument, since every correct rename also scored 1–2% against a softmax floor of 1.06% — the real defect was a calibrated ratio gate that skipped itself whenever the margin was None, which is precisely the ungated OCR-override case. Read a symptom stat before trusting it: the 0/28 extracted-text counter that prompted the false diagnosis is itself an undercount, now recorded below. A third correction: removing the GPS-presence travel rule was reported as a likely broad regression and measured as a single row. Prior update 2026-08-11 (closed 2 items — Lint debt flake8 src/ scripts/ tests/ 540 → 498 → 0 via a mechanical black pass over 109 files (f905953) plus the 72 findings black cannot fix (a8e1d72); and E2E dashboard assertions, now driven by DASHBOARD_CARDS/DASHBOARD_STATS fixtures (d97be75, 1013673). Two lessons worth carrying: an enumerated worklist in this file is a snapshot, not a contract — the lint item's 540 had already decayed to 498 and its per-code breakdown no longer matched; and when an item names one hardcoded list, look for its siblings — the E2E item blamed the count assertion alone, but the same card set was hand-enumerated in three places, so the fifth card had no navigation test either. Both residuals recorded on the E2E item were then closed in 231019a — the .tech-badge filters (the footer has 8 badges; the test asserted 3 with no count, so 5 were uncovered) and the .feature-card:nth-child(1..4) stagger that left the 5th card at animation-delay: 0s. Both residual write-ups were themselves wrong in a way worth noting: the badge one called the gap "lower-stakes" without checking how many badges existed, and the CSS one asserted _site/index.html is generated by content runs when it is the committed source (copy_to_site.sh skips it by name). The stagger needed a computed-style assertion because :nth-child enumerations are invisible to ordinary E2E coverage — no locator, navigation or text assertion can see a missing animation delay. Prior entry: closed HEIC OCR _PREPROCESS_AVAILABLE bug via /backlog-implementer — 7899442; one recall-measurement residual remains open, and two notes were appended to that item — the np.asarray guard shipped with the fix is unreachable given the module's numpy-before-PIL import order, and f8aa49b's DocTRResult Protocol introduced a mypy no-any-return that the Stop hook caught. Also closed the scripts/d1/schema.sql drift item — tests/unit/test_d1_schema_drift.py regenerates from Base.metadata and diffs against the committed file, wired into a schema-drift job alongside the new style gate in .github/workflows/checks.yml (make lint / make schema-check are the identical local commands). Both gates were verified to fail, not just to pass — a misformatted function trips black, an unused import that black accepts trips flake8, and a model column added without regenerating trips the schema check, which make d1-schema then repairs. Repo-state caveat: 195a5b5 committed a docs/BACKLOG.md that links to docs/changelog/2.3.0/CHANGELOG.md and docs/CHANGELOG.md while both were still untracked, and 2.2.0's additions unstaged — until those are committed, HEAD's backlog points at migration destinations that exist only in the working tree. Earlier same day: migrated 11 Done items out via /backlog-migrate — 6 to docs/changelog/2.2.0/CHANGELOG.md, 5 to the new docs/changelog/2.3.0/CHANGELOG.md, with docs/CHANGELOG.md added as the version index. Four of the migrated items carried open residuals that the migration left recorded only as changelog prose; they were re-opened the same day as four standalone items — near-dupe reconcile integration + synthetic scale numbers, the unvalidated DEFAULT_SIMILARITY_THRESHOLD = 0.85, the 164-row calibration corpus, and the HEIC OCR verification gaps (recall never measured; docTR _PREPROCESS_AVAILABLE gate now fixed in 7899442). The general hazard is worth remembering: a "Done" item whose residuals live in bullets is fully migratable by shape, and the residuals leave the worklist with it — a changelog records what shipped, not what remains. Separately, the P2 /git-commit-smart blanket-staging item was moved to ~/.claude/docs/BACKLOG.md as GCS1 — the defect is in the skill, not this project). Prior update 2026-08-11 (closed 2 items via /backlog-implementer — P4 package.json version/test:unit script in 04a9ba7, and the P2 HEIC-OCR decode hand-off in e9fb0a8+41ee326+1c32fe7. The HEIC item is Done for the decode path only: two residuals were found during verification and appended to it — recall was never measured on text-bearing HEICs, and the docTR branch was gated on _PREPROCESS_AVAILABLE; the gate was removed in 7899442). Prior update 2026-08-11 (added 3 items from a test-maintenance session — hardcoded E2E dashboard card assertions, stale package.json version + missing unit-test script, and /git-commit-smart blanket-staging another session's work in the shared checkout; no code written for any of them beyond the 5-line dashboard.spec.ts fix already committed in 58f7659). Prior update 2026-08-10 (added 4 items from the facebookresearch library audit — faiss near-dupe index + SSCD descriptors as two halves of one feature, nevergrad joint weight search, and a PE-Core-vs-ViT-B-32 backbone A/B; library viability, licensing, and py3.14/arm64 wheel availability verified on this machine, no code written for any of them). Prior update 2026-07-25 (migrated 3 Done items to docs/changelog/2.2.0/CHANGELOG.md; resolved the census-gazetteer setup gap — scripts/download_census_names.py verified + documented in QUICK_START/CLAUDE.md; shipped the --ocr-doctr-fallback config flag for the P2 docTR-fallback gate, closing the OCR-bound item's last work item). Prior update 2026-07-18 (added Repo Snapshot — repomix token census, top-churn gitlog, and the uncommitted InteriorSignal→SceneSignal retirement inventory; added identity-detection license-back item + partial fix — corrective lenses keyword added to ID_KEYWORDS (restrictions/endorsements trialed then dropped after backtest showed insurance-doc collision), front/back fixtures, 23 tests pass; added redact_pii.py barcode/alphabetic-PII blind-spot item — OCR-token redaction silently no-ops on ID barcodes + health terms; added trained graphic-vs-photograph probe item — opaque AI graphics/logos leak past the cheap GraphicDetectionSignal gate (code path since closed by c327877 — graphic scene-probe class wired end-to-end, pending corpus + retrain); corrected PHOTO_PROPERTY_CONFIDENCE item post-f6488b9 — two-signal case resolved, residual is probe-absence only; probe now health-checked; fixed person-name false positive — ambiguous Census given names (summer/spring/autumn/winter, month names, virtue words) were auto_accepted when paired with a Census surname; new _AMBIGUOUS_GIVEN_NAMES hard rule + 41-test suite; closed redact_pii.py barcode item — cv2 barcode+QR detection, --redact-terms flag, barcode_unredacted manifest field, non-zero exit, 27-test suite).
IdentityDocumentSignal / the legacy _classify_identification_document both delegate to detect_identity_document (src/scoring/signals/identity_document.py:95), which fires only when ocr_text contains one of the ID_KEYWORDS (:44). The original 14 keywords were all front-side terms ("passport", "driver license", "date of birth", "surname"…). A photographed license back carries none of them — its OCR text is class/restriction/endorsement fields plus barcodes — so it was never detected and fell through to MIME/neither instead of personal/identification.
Surfaced 2026-07-18 while handling a real Texas license back (PXL_20220607_234355242.MP.jpg): OCR text was "CLASS: C-Single… / HAZMAT / REST: A - With corrective lenses / END: NONE / Directive to physician / Emergency Contact / Allergic reaction to drugs / TEXAS ROADSIDE ASSISTANCE" — zero ID_KEYWORDS hits. (Same file that exposed the redact_pii.py barcode blind spot above.)
Partially fixed 2026-07-18: added one license-back keyword to ID_KEYWORDS — "corrective lenses" (near-unique license restriction) — appended last so any front-side keyword still wins the reported-matched_keyword slot. ("restrictions" and "endorsements" were trialed then dropped after the backtest below showed they collide with insurance-document language; only corrective lenses was kept.) Two fixtures + tests added (test_signal_identity_document.py: DRIVER_LICENSE_TEXT front, LICENSE_BACK_TEXT back); 23 tests pass. Flows to both engines via the shared core.
Status: Open — partial. Residual gaps below. Priority: P3 Source: manual license handling during scene-probe corpus seeding, 2026-07-18
- Unrestricted backs still undetectable. A license back with no restriction/endorsement (only barcodes + "NONE") has no keyword hook at all — OCR-keyword detection fundamentally can't see it. Robust license-back detection needs a barcode/PDF417 presence cue (cf. the
redact_pii.pyitem — a PDF417 is itself a strong ID signal) or a trained ID-image probe, not more keywords. "restrictions"/"endorsements"trialed and dropped (backtest 2026-07-18). Both are generic enough to appear on insurance/contract/benefits documents and would file topersonal/identificationatID_KEYWORD_CONFIDENCE(0.85). Backtest (results/file_organization.db, 237 files / 58 with OCR):corrective lenses0 matches,restrictions0 matches,endorsements2 matches — both insurance PDFs (My Documents _ USAA.pdf,Property_Insurance.pdf, "Declarations Page and endorsements…"), i.e. the exact collision predicted. No live false positive only because both areDigitalDocumentand the ID signal is image-only gated (applies_to: is_image) — incidental protection that fails for a photographed insurance card. Decision: dropped both; kept onlycorrective lenses(clean, ID-specific, 0 collisions). If broader license-back coverage is later needed, reintroduce behind a corroborating-token requirement rather than as bare keywords. Corpus is small (spot-check, not comprehensive).ID_KEYWORDSis now front+back mixed. If the reportedmatched_keywordis ever used to sub-classify (front vs back), the flat list won't distinguish them — would need tagging.
GraphicDetectionSignal (src/scoring/signals/graphic_detection.py) is a cheap pre-CLIP pixel gate tuned for transparent, small, square icon assets: its four additive heuristics are has_alpha (+0.40), is_square_icon (square & ≤512px, +0.30), small_palette (≤64 colors in a 64² thumbnail, +0.25), and asset_path (parent dir in the asset-folder set, +0.15), needing ≥ _MIN_RASTER_CONFIDENCE(0.35) to vote. It has no capability for opaque, high-res, flat-design graphics — logos on a solid background and text marketing posters, exactly what ChatGPT/AI image tools emit. Those score ≈ 0 on all four cues (no alpha, >512px, anti-aliased gradients/text push distinct colors >64, parent folder is ChatGPT not an asset dir), fall below 0.35, and leak through to the photos_chatgpt filename fallback (or photos_other).
Reproduced (shadow run, 2026-07-18) on ~/Documents/Media/Photos/ChatGPT, visually verified:
ChatGPTImageNov10,2025,02_32_53PM.png— "InventoryAI" brand logo (1024², opaque navy) → misfiledphotos_chatgpt(both engines agreed; both wrong).ChatGPTImageAug30,2025,03_41_57PM.png— "GOT A VISION FOR A HEALTHIER AUSTIN?" text poster (1536×1024, opaque cream) → misfiledphotos_chatgpt.- Contrast:
ChatGPTImageNov10,2025,02_32_56PM.png(busy 2×2 icon grid) did route tographics_other— so detection currently fires only on busy multi-icon layouts, not single logos or text posters. Inconsistent by construction.
The durable fix mirrors the interior-detection precedent (docs/reviews/INTERIOR_DETECTION_DURABLE_FIX_ANALYSIS.md): replace/augment the pixel heuristics with a trained linear probe over the frozen ViT-B-32 embeddings the pipeline already caches — a graphic (or binary graphic-vs-photograph) class — exactly as InteriorSignal (src/scoring/signals/interior.py, results/interior_probe.joblib) did for the zero-shot interior gate. CLIP embeddings do separate graphics from photographs even though CLIP zero-shot labels don't (the whole reason GraphicDetectionSignal avoided CLIP). Landing chosen and shipped: a dedicated graphic class (index 4) in the scene probe (scripts/prototype_scene_probe.py, corpus results/scene_labels/) — see the runtime bullet below; the logo + poster above are ideal graphic/ positives (boundary rules: results/scene_labels/README.md, graphic-vs-photograph split added 38ddfd5).
Status: Open — narrowed 2026-07-27, labeling policy decided 2026-08-11. Corpus expanded
34 → 307 via Crello; real-world graphic recall 0.67 → 0.76. The promo-panel/dashboard
boundary is now settled (content rule — see the 2026-08-11 bullets at the end of this item),
which unblocks the next corpus increment. Remaining: pull data-viz into graphic/ and
functional-UI/document screenshots into neither/, then retrain and re-eval.
Priority: P3
Source: ChatGPT shadow-scorer investigation, 2026-07-18 (results/scoring_shadow.jsonl; agreement-set manual review)
-
Do not chase it with more pixel heuristics. Opaque AI-generated graphics are pixel-indistinguishable from photos by palette/size/alpha; only the learned embedding separates them.
-
Keep the cheap gate.
GraphicDetectionSignalcorrectly and cheaply catches transparent/square icon assets pre-CLIP; the probe is an additional heavy-tier voter for the opaque-graphic case it can't see, gated on CLIP availability with graceful no-op (same pattern asInteriorSignal). -
Corpus dependency. Needs labeled graphic/neither positives (target 150–300/class per the scene-probe README);
results/scene_labels/is seeded and actively being labeled — a dedicatedgraphic/class dir now exists — but every class is still below target (2026-07-18 snapshot: neither 46, interior 44, exterior 21, graphic 8, place 3), so the probe is not yet trainable. -
Regression guard. Measure false positives on genuine photos (product still-lifes, staged interiors) before deploying — a graphic probe that over-fires would pull real photos out of
photos_*. -
First eval baseline (2026-07-18,
gather --label-dirs+eval, 3-fold CV, n=122). Corpus at eval time: neither 42, interior 44, exterior 21, place 3, graphic 12. Graphic class: precision 1.00 / recall 0.42 / F1 0.59 — confusion rowgraphic → [7 neither, 0, 0, 0, 5 graphic](7/12 graphics still fall back toneither→ leak tophotos_*, the original bug). Classic starved-minority signature: high precision, low recall — corpus volume is the fix, not tuning. Deploy-safety confirmed: at the defaultSCENE_MIN_PROB=0.5graphic precision holds at 1.00 across the 0.5→0.7 sweep, so a live 5-class probe would add no photo→graphics false positives (safe-but-partial: misses ~60% of graphics, no regression on them). Headline metrics are inflated (macro ROC-AUC 0.993 / acc 0.918):placen=3 is meaningless and near-duplicate images (two coffee-station photos, multiple Integrity banners) leak across CV folds — do not read as deployment-ready. Decision: do nottrain/commit a.joblibyet — an artifact activates the registry swap, and graphic recall 0.42 isn't worth shipping. Hold untilgraphic/≈150 andplace/is real. -
Runtime landing path — code steps done (
c327877, 2026-07-18).SceneSignal(src/scoring/signals/scene.py) + the artifact-gated registry swap landed in5f6db5b;c327877wired the 5th class end-to-end:"graphic": 4inSCENE_CLASSES+_POSITIVE_NAMES(sogatherpicks upresults/scene_labels/graphic/), runtime mappingSCENE_CATEGORY["graphic"] → ("media", "graphics_other")(resolves toMedia/Graphics/Other, the same targetGraphicDetectionSignalemits),SCENE_SCHEMA["graphic"] → ImageObject,_INT_CLASS_NAMES[4]. Back-compat: a pre-graphic 4-class artifact still loads —SceneSignalignores classes absent fromSCENE_CATEGORY, so nothing misroutes before retraining. 16/16 scene tests pass (new test pins graphic-argmax →media/graphics_other). -
Trained and shipped (2026-07-18, supersedes the "do not train yet" hold above). Corpus was expanded the same day (Places365 sampling for place/exterior/interior + Business-folder graphics + Media-filtered
--db-neither) to 835 rows — neither 118, interior 178, exterior 158, place 347, graphic 33 — clearing the hold'splace-is-fake blocker. Final 5-fold eval: accuracy 0.92; graphic P 0.73–0.82 / R 0.60–0.67 / F1 ~0.70. User decision: train now, accept conservative graphic recall — the miss mode is benign (a missed graphic gets no scene vote and falls through toGraphicDetectionSignal+ other signals, i.e. today's behavior).scene_probe.joblibtrained + committed (6f61449); swap completion followed (interior.py deleted,photo_compositioninterior vote retired,scene_probehealth feature, W_SCENE backtest: −20% flips 4 / +20% flips 0 on the 202-row replay). Still open here: graphic recall. Corpus volume from pure graphics (logos, posters, flat illustrations — local sources are tapped out; a public logo/infographic dataset is the realistic path), and a labeling-policy call on promo-panels-with-embedded-UI (semantically graphics, visually half-screenshot — currently the main graphic↔neither confusion). -
Corpus expanded from Crello, 2026-07-27 — and the headline number is misleading.
scripts/download_crello_graphics.pypulled 273 flat-design previews fromcyberagent/crelloacross 14formatsubtypes (logo, poster, flyer, ad creative, certificate, coupon…), filtered to templates with noImageElement(those embed photographs) and one percluster_index(near-duplicate variants would leak across CV folds). Class went 34 → 307, inside the 150–300 target band. Aggregate 5-fold CV: graphic P 0.73 → 0.97, R 0.67 → 0.96, F1 0.70 → 0.97; overall accuracy 0.90 → 0.93. Do not read that as a 0.97. Per-source recall: Crello 272/273 = 0.996, hand-collected 25/33 = 0.758. Crello templates are near-trivially separable, so they inflate the aggregate; the honest real-world gain is 0.67 → 0.76. Images are gitignored (graphic/crello_*) — CyberAgent conditions use on the VistaCreate ToS and does not redistribute source files, so the script is the reproduction path (same arrangement asdownload_census_names.py). -
The residual error is one coherent subtype, not general scarcity. 7 of the 8 missed hand-collected graphics are
biz_dashboard_*images and the 8th is an "AI Market Map" infographic — all flat-design data-viz, all routed toneither. Visually confirmed: these are non-photographic flat vector charts, sographicis the correct label and the probe is genuinely wrong. Crello supplies posters/logos/ads and almost no data-viz, which is exactly why it didn't fix them. Next increment should target dashboards/infographics/charts, not more posters. InfographicVQA is the ideal content match but is research/education-only (blocked — see~/.claude/.../memory/dataset-license-constraints.md); DomainNetinfograph(51,605 images, one-line TFDS pull) is the license-viable candidate, pending verification of DomainNet's terms at ai.bu.edu. -
This also sharpens the promo-panel labeling call.DECIDED 2026-08-11 — the content rule: the pixels decide, not the framing. Synthetic imagery isgraphic/even when the image is a screenshot; live interface chrome in the frame does not move it toneither/. Written intoresults/scene_labels/README.mdas three bullets, replacing the both-ways ambiguity ("data-viz →graphic" vs "UI screenshots →neither") that made the README undecidable for the same image. Enrico's precedent argued the other way and was not followed: the graphic class exists to stop non-photographic imagery reachingphotos_*, and a dashboard is not a photograph. No image was relabelled — all 12 contestedbiz_*files keepgraphic/, so the 0.76 real-world recall figure remains comparable across the decision rather than being mechanically improved by moving the errors. -
The rule left one line undecided, so the README resolves it explicitly: text documents and functional UI (settings, terminals, editors, forms, chat) stay in
neither/. A literal "synthetic →graphic/" would swallow them, and they belong on the OCR/screenshot path, not inMedia/Graphics. The discriminator inside synthetic content is imagery vs text/controls. A decision that resolves an ambiguity can introduce a new one at its own edge — state the edge in the same commit or the item reopens under a different name. -
The risk now runs the other way, and the corpus is not shaped for it. This rule trains the probe toward "UI layout → graphic", so the new failure mode is genuine UI screenshots and document scans being pulled into
Media/Graphics.neither/is the class that must teach that distinction and it is by far the smallest — 73, againstgraphic308 andplace347. So the next corpus increment is no longer just DomainNetinfographintographic/; it must add functional-UI and document screenshots toneither/, and the retrain must reportneitherrecall and theneither → graphicconfusion cell. -
Correction to the bullet above: the residual errors are not "one coherent subtype". Inspecting the images rather than the filenames, the 12 hand-collected
biz_*files are at least four different things: a pure flat vector illustration of chart panels (biz_dashboard_9f2885b8, no interface at all), a chart-dominant real analytics view (biz_dashboard_26478eb2), composed promo panels of headline + bullets + product mockup (fba0e47a,biz_20260107_dashboard), and web captures carrying live chat-widget chrome (882118b3). Only the last two were ever contested; the first is a plain graphic the probe misses on volume alone. "They're allbiz_dashboard_*" is a filename observation being read as a content finding — which also means a wholesale relabel of that prefix would have been wrong in either direction. -
Encoding confound found and fixed (
scripts/normalize_scene_corpus.py). The corpus had class partly recoverable from file metadata alone:place/was 100% JPEG at 256px (Places365), hand-collectedgraphic/mostly PNG at ~1536px. A logistic regression on encoding metadata only (format, dimensions, aspect, size, bytes-per-pixel — no pixels) scored 0.561 vs a 0.327 majority baseline, lift +0.234. Normalizing every class through one encoder (JPEG q90, longest side 256 — above CLIP's 224 input, so nothing the model sees is lost) cuts it to 0.411, lift +0.084. Ablation shows the remainder is aspect ratio (+0.139 alone) and compressibility (+0.072); aspect cannot reach the probe because CLIP center-crops to square, and compressibility is genuine content signal (flat design compresses better than photographs). Re-eval on the normalized tree holds graphic recall at 297/306, confirming the Crello gain is content, not encoding artifact. Output tree is gitignored and regenerable.
Profiling organize-files content (unified scorer, dry-run) on 2026-07-17 showed the workflow is ~85% OCR-bound (torch.conv2d in the easyocr CRAFT + docTR detection CNNs ≈ 67% of self-time; CLIP is negligible). This session shipped P2/P3/P5 (docTR-fallback gate, screenshot double-OCR dedup, CLIP text-embedding memoization) — conv2d call count dropped 1189→528 (−56%). All planned work items are shipped; the item stays listed only to monitor the P2 gate's faint-text recall tradeoff in live use.
Status: Done pending monitoring — P1 shipped 2026-07-18 (gate on by default at K=3); P2 gate shipped, and its escape hatch (--ocr-doctr-fallback) shipped 2026-07-25. Remaining: watch for faint-text misses in live runs; no code work open.
Priority: P3
Source: content-classification profiling + P2/P3/P5 optimization session, 2026-07-17
-
Gate OCR on text-likelihood (P1) — Done.
OCR_CLIP_GATE_TOPK = 3constant added tosrc/scoring/weights.py; gate enabled by default at K=3 inContentOrganizer,ContentBasedFileOrganizer, and the CLI. Opt out with--ocr-clip-topk 0(orocr_clip_topk=None);FileContext._skip_ocr_by_clip_gatetreatsK=0/Noneas disabled. Eval: K=3 → 100% text recall, ~35% of photos skip OCR. Gate fails open when CLIP is unavailable. 15 unit tests added totests/unit/scoring/test_context.py::TestClipOcrGate. Reusable tooling:scripts/profile_pipeline.py(hot-path profiler) +scripts/eval_ocr_gate.py(folder-labeled gate eval). -
P2 docTR-fallback gate — recall tradeoff to monitor. The shipped gate (
extract_ocr_with_confidence: skip the docTR fallback when easyocr cleanly finds no text) was eval'd over 7 text images at varying difficulty: 1/7 recall loss — very-low-contrast text (easyocr's detector found nothing; docTR would have caught it). Clean, dark-mode, and rotated text were all gate-safe. So P2 trades a rare miss on near-invisible text for eliminating the docTR fallback. For a screenshot/photo-dominated 265k-file library this is very likely a net win, but it is a real behavior change — put it behind a config flag or revert if faint-text recall matters. Config flag shipped 2026-07-25:--ocr-doctr-fallback(store_true, default off = gate on; constantOCR_FORCE_DOCTR_FALLBACKinsrc/scoring/weights.py) forces the docTR pass after a clean easyocr negative. Plumbedforce_doctr_fallbackthroughextract_ocr_with_confidence,TextExtractor,ContentOrganizer(all 3 call sites incl. the FileContextocr_provider),ContentBasedFileOrganizer,ContentInputs, and the CLI — same pattern as--ocr-clip-topk. 5 gate tests intests/unit/test_shared.py::TestDoctrFallbackGate; CLI-inputs contract + integration suites pass.
A 2026-07-26 audit of results/file_organization.db (495 rows) found 50 rows whose
current_path no longer resolved on disk. Repairing them surfaced two distinct
integrity problems, neither of which is data loss, and both of which mislead any
future reader of the DB (including the calibration oracle in
docs/architecture/scoring-calibration-20260726.md).
Status: Path integrity resolved — 7 dead rows pruned (reconcile --prune-missing),
42 of the 43 reverted-run rows re-organized and healed (item 1), the 3 category-less
Events/Burning Flipside/ rows now resolve via the looped entity-segment strip
(item 3, 2026-07-27), and the flutter_auth.txt straggler repaired 2026-08-11 (item 3).
0 of 488 rows now claim a current_path that does not exist while their file is
live. Open: only the 25-row staging-dir provenance record (item 2), which has no
honest repair and stands as a rule rather than a task.
Priority: P3
Source: DB path audit, 2026-07-26 (same session as the sprite naming-trap fix)
- Reverted organize run leaves rows claiming destinations that never persisted (43 rows).
A batch run on 2026-06-27 04:07:41 moved 43 files from
~/Desktop/Uncategorized(42) and~/Desktop(1) into~/Documents/..., persisted them withstatus=ORGANIZED, and the moves were later reverted — the files are back at theiroriginal_pathand byte-size identical to their DB records (43/43 verified), while everycurrent_pathpoints at a Documents path that does not exist. Persistence is gated onif not dry_run(file_processor.py), so this is not a dry-run artifact: the moves really happened and were undone afterwards.GraphStore.prune_missing_filescorrectly refuses to touch these (it requires both paths gone), so they survive every prune and permanently misreport where those 43 files live. Options when someone picks this up: repaircurrent_path→original_path(record reality, and revisitstatus), or re-runorganize-files contentover~/Desktop/Uncategorizedso the rows become true. The latter is attractive because these 43 are exactly the corpus that motivated the weak-shape sprite fix — 19 of them are photos the DB still labelsgame_assets/sprites, which the fixed classifier now routes tomedia/photos_*. RESOLVED 2026-07-26 by taking that option:organize-files content --source ~/Desktop/Uncategorizedre-organized 42 of the 43 (the 43rd is item 3 below). Becauseadd_filekeys ongenerate_id(original_path)and the files were back at their original paths, all 42 rows updated in place — no duplicates. The sprite fix held: zero files routed togame_assets/sprites; 36 photos went tomedia/photos_{social,other}, the logos toOrganization/Integrity Studio, andBurning_Flipside_Map.pdftoEvents/Burning Flipside/. - Agent temp-staging dirs must never be the organize source (25 rows).
25 rows record
original_pathinside~/.claude/jobs/e184540b/tmp/interior_apply_stage/— an agent job's temp staging directory, which still exists but is empty. 23 of the 25 files are fine on disk; 2 were among 6 unrecoverableMedia/Interiorsrenders pruned on 2026-07-26. For all 25 the recorded provenance is an ephemeral path, and the true pre-staging origin is unrecoverable, so there is no honest repair — only this rule: an agent run must not organize files out of a temp/staging dir into the production graph store, becauseoriginal_paththen permanently records a path that ceases to exist and provenance is destroyed. Stage into a durable location, or organize from the user's real source directory. - Residual rows the 2026-07-26 sweep deliberately left (4 rows).
FIXED 2026-08-11 — but by neither route this item proposed, because both were wrong. The row is nowflutter_auth.txt— the 43rd reverted-run file, sitting at~/Desktoproot rather than~/Desktop/Uncategorized, so it was outside the source the re-organize covered.current_path=NULL,status=PENDING,organized_at=NULL, withorganization_reasonrecording the reverted run; the file is untouched at~/Desktop/flutter_auth.txt(23,596 bytes, matching the row).reconcilehas no path-repair operation. This item named one as if it existed.cmd_reconcileimplements exactly--set-category,--prune-missingand--backfill-categories— none of which can rewrite a path. Check the subcommand before citing it as the cheap route.- The
--source ~/Desktop --limit 1pass was not viable either.BatchProcessor.scan_directoryisrglob("*"), so that source is 751 files recursively, not the 10 at the top level, and--limit 1takes an arbitrary rglob entry rather than the intended file — a scattergun that would move unrelated Desktop files into~/Documents. There is no single-file--source. current_path=NULLrather than= original_path, which is how this item phrased the option.add_filenever setscurrent_path, so NULL is the never-organized shape, and thecurrent_path or original_pathreaders (graph_store.py:1288,:1364) resolve it to the Desktop path either way. Writing the original path in instead would make the four sites that treat a non-NULLcurrent_pathas "organized" (:988,:1021,:1448) report a file that was never filed as filed — trading a wrong path for a wrong status.- Verified that NULL does not make it prunable:
prune_missing_filesrequires both paths gone, andoriginal_pathexists, soreconcile --prune-missingstill reports the same 6 unrecoverable rows and leaves this one alone. - Its
business/planningcategory edge was deliberately left in place — the edge records the classification, and this repair is about path provenance. A file at~/Desktoproot has no taxonomy ancestor, soresolve_taxonomy_folderwould report it unresolved rather than propose a better one.
3 rows underFIXED 2026-07-27. Folder lookup moved toEvents/Burning Flipside/have no category edgeresolve_taxonomy_folder(graph_store.py): exact match first, then trailing segments stripped in a loop until an ancestor is in the taxonomy, so entity-named folders resolve at any depth (Events/{Name}/2026/maps→Events,Media/Interiors/{Prop}/{Room}→Media/Interiors). A parent reached by stripping that declares no subcategory is filed under the generic bucket —Events/*→events/other— while an exact match keeps the bare category, so a file sitting directly inEvents/still readsevents. Live dry-run: the 3 Burning Flipside rows now resolveevents/other, 0 unresolved of 3 orphaned (was 3 unresolved). 3 new tests intests/unit/test_category_identity.py(deep nesting, exact-match preservation) + the existing entity test updated; 224 storage/pipeline tests pass. Applied to the live DB 2026-07-27 (backup…bak-20260727_192729): 3 edges attached, category-less rows 0 of 488 — closing the orphan count 125 → 3 → 0. As-intended consequence, documented in CLAUDE.md and the method docstring: the pair follows the folder and nothing else, so two copies of one document in two trees get two individually-correct edges (Documents/Events/Burning Flipside/…→events/other;Documents/Personal/Events/…_300dpi.png→personal/events), and a category query returns a subset of the document family. That is filing, not drift — the backfill must not infer which copy is canonical. LikewiseOrganization/{Name}→organization/vendorsfollows the taxonomy's declaration of that folder as the vendor/partner root.
scripts/d1/schema.sql is generated from Base.metadata by
scripts/d1/generate_schema.py, and its own header says "the output file is the
authoritative D1 schema — do not hand-edit it. Edit src/storage/models.py instead, then
re-run this script." Nothing enforces the second half. Regenerating it on 2026-07-27 for
the file_count removal revealed it had been stale across three earlier model changes
that nobody regenerated:
ix_categories_full_pathwas still plain andix_categories_namestill UNIQUE — i.e. the D1 schema still carried the exact index shape whose UNIQUE-on-namebug dropped category edges for 26% of rows (fixed locally 2026-07-26, never regenerated).peoplewas missingreview_status,detection_confidence,validation_scores,validated_atandix_people_review_status(the person-validation work).file_categorieswas missingsignal_evidence(UNIFIED_SCORING_PLAN §5.4).
A D1 load against that schema would have failed on the missing columns or silently
recreated the fixed UNIQUE bug in the mirror. Nothing consumes those columns today —
workers/file-org-api has no reference to any of them — so no live breakage resulted, but
the drift is unbounded and invisible until someone deploys.
Options: a test that regenerates into a temp file and asserts it matches the committed
copy (cheapest, catches it in CI); a pre-commit hook on src/storage/models.py; or
generating it at deploy time and not committing it at all. The first is the obvious fit —
the repo already has hand-rolled golden-file comparisons (tests/unit/golden/).
Status: Done 2026-08-11 — the first option shipped (tests/unit/test_d1_schema_drift.py).
Priority: P3 (no live consumer today; blocks or corrupts a future D1 deploy)
Source: file_count cache removal, 2026-07-27
- Shipped as the option this item recommended: the test regenerates in memory (no temp
file needed —
generate()already returns the string) and diffs against the committed copy. Local:make schema-check/npm run schema:check. CI: theschema-driftjob in.github/workflows/checks.yml, so it fails on the commit that causes the drift rather than at deploy. - Verified against the exact historical failure, not just written: adding a column to
the
Personmodel without regenerating fails withschema.sql is stale, naming the first differing line and the line-count delta. Thenmake d1-schemarepairs it and the check goes green — the full detect → repair → pass loop was exercised end to end. - The failure message has to carry the fix, because the fix is not guessable. A
byte-diff of two ~10 KB SQL blobs is unreadable, so the assertion reports the first
differing line plus
Run \python scripts/d1/generate_schema.py` and commit the result.` A drift guard whose output does not say how to regenerate gets "fixed" by hand-editing the generated file — which the header already forbids and which this test also catches. - Two guards protect the guard itself.
test_generate_is_deterministicpins the table/index sort ordering: a non-deterministic generator would make the drift test flake, and a flaky test gets deleted rather than fixed.test_generator_does_not_write_on_importpins theif __name__ == "__main__"guard — without it, importing the module during collection would rewriteschema.sqland the drift test would pass by silently repairing the very thing it exists to detect. - CI installs
requirements-schema-check.txt(SQLAlchemy + pydantic + pytest), verified sufficient in a clean venv — importingsrc.storage.modelspulls no torch, Pillow or doctr, so the job stays cheap. Unlike the lint pins, these are deliberately unbounded: the check asserts generator and file agree, and both move together under a SQLAlchemy upgrade, so a DDL-rendering change should surface as a regenerate-and-commit rather than hide behind a pin. - Not chosen: the pre-commit hook. It would not survive a fresh clone (hooks are not
versioned), and the existing
pre-commithook in this repo deliberately only warns — making this one blocking would be inconsistent with it. CI is the durable surface.
organize-files content redacts files.extracted_text and schema_data["text"] for
files classified medical (src/analyzers/text_redaction.py, added 2026-07-26). Two
deliberate limits are worth revisiting rather than forgetting:
- Category-gated, not content-gated.
MEDICAL_TEXT_REDACTION_CATEGORIESis{"medical"}, so a health document the scorer files elsewhere stores raw text. This is not hypothetical: during the medical ingestion the scorer wantedpersonal/contactsfor 5 of the 11 files anduncategorizedfor another — only the folder-truth category override kept them inside the redaction gate. Consider gating on detected content (medical vocabulary / KIE field classes) rather than the winning category, and extending to the other sensitive categories the QUICK_START tips call out. - Only numeric/email/known-name PII is masked. Emails, digit runs of 2+, and the
decision's detected person names are masked; alphabetic PII outside that set —
conditions, medications, third-party names the entity detector missed — survives
verbatim. Same known limitation as
scripts/redact_pii.py's raster path, and the reason--no-dbremains the guidance for genomics and similar.
Also unquantified: how many organized files are missing from the DB entirely. The
Medical/ audit found 12 of 17 on-disk files untracked (filed manually or by the DB-free
name/type organizers, which persist nothing by design). No census was run over the
rest of ~/Documents, so the true coverage of the graph is unknown.
Partially fixed 2026-07-26 (3b90a08). The category gate was widened:
MEDICAL_TEXT_REDACTION_CATEGORIES→TEXT_REDACTION_CATEGORIES(same value{"medical"}); backward-compatible alias retained.- New
TEXT_REDACTION_SUBCATEGORY_PAIRS = {("personal", "identification"), ("personal", "records")}. - New
should_redact_text(category, subcategory)helper checks both sets. _persist_to_graph_storenow usesshould_redact_textinstead of the bare category check.- 12 tests (6 new) in
tests/unit/test_text_redaction.py.
Still open: content-based gating (a medical document misclassified as
uncategorized still stores raw text); alphabetic PII (conditions, medications) is
not masked; untracked-file census over ~/Documents has not been run.
Status: Partially fixed — category pairs extended; content-based gating and alphabetic PII remain.
Priority: P3
Source: medical text-redaction work + Medical/ folder audit, 2026-07-26
mypy is clean repo-wide as of 2026-07-26 (2dbc148), but flake8 src/ scripts/ tests/
reports 540 findings across the tree (config: .flake8, max-line-length 100,
extend-ignore E203/W503). Every finding encountered during the type pass was verified
pre-existing (checked against git stash), so this is long-standing debt, not new.
By code: E111 indentation-not-a-multiple-of-four (249), E501 long lines (151),
F401 unused imports (44), E402 module-import-not-at-top (21), E128/E131
continuation-line indent (22), F841 unused locals (15), E114/E302/W293 (23),
plus a scatter including E741 ambiguous variable name 'l'
(tests/integration/test_schema_org_export_e2e.py).
By file, the debt is concentrated — five files hold ~55% of it:
src/api/schema_org_models.py (162), scripts/d1/export_to_d1.py (49),
src/cost_roi_calculator.py (30), tests/unit/test_image_metadata.py (28),
scripts/image_content_analyzer.py (26).
The E111/E114 bulk is 2-space-indented code that predates the project's 4-space
convention, so black would fix most of it mechanically — but reformatting
schema_org_models.py (the Pydantic API surface) and export_to_d1.py in one sweep is a
large diff that should land on its own, not mixed with behaviour changes. Suggested order:
run black over the five concentrated files, then clear F401/F841 (safe deletions),
then hand-wrap the residual E501s. E402 in scripts/profile_pipeline.py and
file_organizer_content_based.py is deliberate (sys.path setup precedes imports) and
should get # noqa: E402 rather than reordering.
Status: Done 2026-08-11 — flake8 src/ scripts/ tests/ is at 0.
Priority: P3 (no behaviour at stake; mechanical but a large diff)
Source: mypy cleanup pass, 2026-07-26
- The count had already drifted before the work started: 540 → 498. The per-code breakdown above is the 2026-07-26 census and does not match what was actually fixed. Treat an enumerated worklist in this file as a snapshot, not a contract — re-run the tool before sizing the job.
f905953— the black pass. 109 files, 498 → 72. Cleared E111 (254), E128 (18), E114 (9), E302 (8), W293 (6), E131 (4), E305 (3), E261 (2) and one each of E306/E303/E129/E124; E501 150 → 32.pyproject.tomlalready set black's line-length to 100, matching.flake8, so the two agree on E501 — no config change was needed.a8e1d72— the 72 black cannot fix, in the order suggested above. F401 (15) deleted; E402 (19) + F811 (1) took# noqa; E741 (2), F841, F541, E231 fixed directly. E501 (32) hand-wrapped.- The regex splits were verified, not eyeballed. Seven long patterns in
entity_detector.pywere split with implicit raw-string concatenation (the idiom that file already used for its "Name on Card" pattern). All 35 compiled patterns across company/people/relationship/sentence were snapshotted before and after and compared equal. A regex split is a behaviour change until proven otherwise — a mis-placed break inside a character class or escape is silent and would only surface as a classification miss. - Seven E501s were caused by a trailing
# type: ignore[...], not by the code. Hoisting the ignored expression to its own line (or parenthesising the import with the ignore on the opening line) keeps mypy clean, which is the proof the ignore still lands on the line that reports the error. Moving one to the wrong line silently disables it and leaves mypy green only until the underlying error moves. - E402 sites are all deliberate bootstraps —
sys.pathsetup or pytest module stubs that must precede their imports — so# noqaper line, as this item originally called for. Reordering them would break the bootstrap. - Verification: 2427 passed / 2 skipped (unit), 192 (integration), 14
(performance), mypy clean on 148 files,
organize-files health12/12,compileallclean.scripts/d1/schema.sqlregenerates unchanged, confirming themodels.pyedits do not reachBase.metadata. Nothing prevents the next drift.Gate shipped —.github/workflows/checks.yml(joblint) runsblack --check --diff+flake8oversrc/ scripts/ tests/on every push and PR, withmake lint/npm run lintas the identical local command so a green run locally means a green CI job. Verified to actually fail, both legs independently: a misformatted-but-valid function exits 2 on black, and an unused import that black accepts exits 2 on flake8 — neither tool subsumes the other.- The version bounds are the subtle part, and an unpinned gate would have rotted.
blackchanges its stable style once a year at the first release of a new major, so a plainpip install blackin CI would eventually demand a reformat of correctly-formatted code and fail on an untouched tree. Bounded<27in bothrequirements-lint.txt(what CI installs) and thedevextra; keep them in step. - CI installs
requirements-lint.txt, neverpip install -e ".[dev]"— the base dependencies pull torch viapython-doctr, which would dominate a lint job. This is also whymypyis not in the gate: it needs the full dependency set to resolve imports, so it stays a local check. Still ungated:Closed the same day — it now has its ownscripts/d1/schema.sqlregeneration.schema-driftjob in the renamed.github/workflows/checks.yml(the file waslint.ymlfor one commit; it grew a second job, so the name no longer fit).
organize-files find-duplicates shipped 2026-08-10 (src/similarity/, 30 tests) and stands alone: src/cli.py:455 registers cmd_find_duplicates, and nothing in cmd_reconcile consumes its groups. That leaves the read-side view unbuilt for the case that motivated the feature. reconcile --backfill-categories deliberately answers "where does this file sit", never "what is this file", so it files a split document family to two different categories by design (Documents/Events/Burning Flipside/PlacementMap.pdf → events/other; Documents/Personal/Events/PlacementMap_300dpi.png → personal/events). Both edges are individually correct, and the split is intended — but today it is also silent. A near-dupe report is what would make it visible.
Status: Open — feature shipped, integration and scale verification did not.
Priority: P3
Source: faiss/SSCD near-dupe build, 2026-08-10; re-opened 2026-08-11 after /backlog-migrate moved the parent items to docs/changelog/2.3.0/CHANGELOG.md with these residuals recorded only as changelog prose
- End-to-end verification was on 8, 3 and 3 files. No real full-corpus run has been done. Everything below the small-corpus level is projection.
- The 265k figures are synthetic, and one of them is a schedule risk.
IndexFlatIPat 265,000 × 512 measured 543 MB resident with 5,000 queries k=5 in 1.13 s (→ full self-kNN ≈ 60 s brute force) — but on random vectors, not descriptors. The cost that matters is the descriptor pass: ~45 ms/image implies roughly 3.3 CPU-hours to describe 265k files cold. Budget that before promising a full-corpus run; it is the reason this has not simply been run once. - Do not reach for IVF/PQ when you do wire it. Exact flat search is already fast enough at this corpus size; an approximate index would add recall tuning for no gain. Recorded because "265k vectors" reads like an ANN problem and is not one.
- Integration shape is a read-side view, not a scorer signal. Consuming groups inside
reconcileis the open work. ASimilarityIndexSignalinside the unified scorer is a separate, later question needing its own weight calibration — do not conflate them.
DEFAULT_SIMILARITY_THRESHOLD = 0.85 (src/similarity/constants.py:58) was chosen to keep a review queue readable, not measured against a labelled duplicate set. The only evidence behind it is synthetic: transforms of a single image put the true group at 0.926 with no false pairs. That establishes the transform survives the descriptor, and says nothing about precision or recall across the real corpus.
Status: Open — number is a placeholder wearing a constant's clothes. Priority: P3 Source: SSCD descriptor work, 2026-08-10; re-opened 2026-08-11 (see the item above for why)
- The failure modes are asymmetric. Too low floods the review queue and the feature gets ignored; too high silently reports nothing and looks like a clean corpus. The second is the dangerous one because it is indistinguishable from success.
- A labelled set is cheap here — the split-document families the backfill already produces are known-positive pairs, and
find-duplicatesoutput on the real corpus supplies candidates to hand-judge. This does not need a public benchmark. - First-page-only PDF rasterisation bounds what can be validated. Multi-page documents match on their cover, which is right for "same document" and wrong for "same content buried on page 7" — so a labelled set should record which kind of match it is testing.
make weight-search (nevergrad, joint search over 19 priors + both thresholds) ran 2026-08-11 across NGOpt/CMA/TwoPointsDE at budgets 120–250 and found nothing: train non-media agreement never moved off the shipped 59/164, and every best candidate was flat or −1 on the holdout. The correct reading is not "the weights are optimal" — it is that 488 rows / 164 non-media with a stored label cannot resolve the difference, and the oracle itself carries manual corrections and pre-unified placements.
Status: Open — the binding constraint is labelled data, not optimiser budget. Priority: P3 Source: nevergrad joint weight search, 2026-08-11; re-opened 2026-08-11 (see the near-dupe item above for why)
- More budget is the wrong lever and will look like progress. A larger budget on this corpus buys more overfitting to oracle noise, and the holdout guard will keep reporting flat-or-−1 — which is the guard working, not a tuning problem.
- Media rows are excluded for a reason. They disagree for replay-fidelity reasons rather than weight reasons, so the non-media slice stays the reported objective; an optimiser scored on the overall slice chases fidelity artifacts.
- Adding labels means adding correct ones. Some stored labels are simply wrong — inspect
winning_signalsand OCR text on actual rows before treating a disagreement as a signal defect. A larger corpus of the same biased oracle does not help. make golden(43/43) stays the gate for any re-tune this eventually enables, and any accepted change still ships as a calibration doc + backtest report. The script never writesweights.pyby design.
CLAUDE.md names make calibrate + make golden as the gate any scoring change must pass. For one of the 19 registered signals that gate cannot fail, because both harnesses construct the registry with image_analyzer=None — scripts/backtest_scoring.py:486 and tests/integration/test_unified_scoring_golden.py:121, the latter with the explicit comment "PhotoCompositionSignal gates itself off". applies_to requires a non-None analyzer, so the signal never runs on any row, and signal_wins in results/backtest_report.json confirms it: photo_composition is absent entirely across all 488 replayed rows while the other 18 signals accumulate wins.
Found 2026-08-11 running the gate on the people-gate rank conversion (d699f45, d5c308c). Before/after agreement was byte-identical — 269/488 = 0.5512, decision states 469/15/4 — measured against a worktree at the pre-change commit using the same post-backfill DB. That looks like a clean no-regression result and is not one: the harness cannot observe the code that changed.
Status: Open — the gate passes unconditionally for this signal, so a real regression in it would ship green.
Priority: P3
Source: make calibrate run for the people-gate rank conversion, 2026-08-11
- The dangerous property is that it fails silently in the safe-looking direction. A gate that errors gets fixed. A gate that returns "no change" for a change it cannot see gets believed — and the more confident the surrounding process (backup taken, before/after compared, EXIT=0, 45 goldens green), the more convincing the wrong conclusion. Every one of those steps was performed here and none of them could have caught it.
image_analyzer=Noneis not obviously wrong at either call site, which is why it survived: the replay has no live analyzer and the goldens want hermeticity. The fix is not to wire a realImageContentAnalyzer— it would need real pixels and CLIP at replay time. It is to inject a stub analyzer driven from stored data, the same shape asmake_screenshot_ocr_stub(backtest_scoring.py) already uses forScreenshotOcrSignal. That precedent exists in the same file.- A stub needs a stored
has_people/has_facesto replay from, and nothing records one today.files.image_classificationholds CLIP label scores, not the cascade's verdict, and the cascade is the primary path (of 15 route changes measured over 400 images, 13 were faces-only and 2 were both — zero were CLIP-rank-only). So closing this needs a persisted face-detection result first;file_categories.signal_evidenceis the natural home now that it is readable viaGraphStore.get_signal_evidence(b6daa39). Second, narrower blindness in the same context:FIXED 2026-08-11.image_metadata_provideris never wired.build_contextnow passesmake_image_metadata_provider(row), which reads EXIF datetime + GPS from the row's on-disk file. 11 tests intests/unit/test_backtest_scoring.py(TestOnDiskPath,TestImageMetadataProvider,TestBuildContextImageMetadata); goldens 45/45, lint and mypy clean.- The item understated the blindness: it disabled two signals, not one.
ClipVisionSignal's GPS→GPS_TRAVEL_CATEGORYupgrade was named, butMediaHeuristicSignalreads the same dict for both its GPS travel branch and its EXIF-datetime "personal photos" branch (media_heuristic.py:202) — and that is where all three observed changes actually came from. When an item names one consumer of a shared value, grep for the others. - The stored-column route is dead, which is not obvious from the schema.
fileshasexif_datetime,gps_latitudeandgps_longitude, so reconstructing from persisted data looks like the hermetic option — but all three are empty on all 488 rows, so that provider would return{}for the whole corpus and the gate would stay just as blind while looking fixed. Reading the file is the only route;_ReplaySceneSignalalready set the precedent. - It resolves
current_pathfirst, falling back tooriginal_path. The replay classifies underoriginal_path(pre-move), so a provider that read the path it is handed would find nothing for every file that was actually filed — the same trapscene_path_lookupexists to avoid. Pinned bytest_reads_the_resolved_file_not_the_path_it_is_handed. - Deliberately not
get_metadata_summary. That reverse-geocodes GPS through Nominatim, which would put a network call in the calibration gate; no signal readslocation_name,year,month,date_strortext_metadata, so the provider returns onlydatetimeandgps_coordinates. It does reproduce the PNG "Creation Time" fallback — 0 rows take that path today, but.pngis aPHOTO_EXTENSION, so the branch is reachable and would otherwise diverge silently. - Measured effect: 3 of 488 decisions change, and agreement drops 263 → 262. Only 3 of 374 resolvable image rows carry any EXIF at all. The −1 is the oracle being wrong, verified row by row rather than assumed:
IMG_8550.HEICmovesmedia/photos_other→media/interiors_other, which is exactly what the HEIC mtime-names item predicts a--forcepass should produce ("the lost GPS degradedmedia_heuristictophotos_other"). The license backPXL_20220607_234355242.MP.jpgmoves topersonal/identification, matching its disk location and one of its two stored edges, but scores as no change because the oracle pair for that row is the known-wronggame_assets/sprites. So 2 rows got more correct and 0 got less correct while the headline number went down — do not read this metric without opening the rows. - Baselines in this file decay; measure your own. The −1 above is against a same-session run with the provider stubbed out, not against the 269/488 recorded earlier in this item —
d5c308candf486ebalanded in between, and the current unstubbed number is 268/488 throughsummarize_replay's own accounting, which filters differently again.
- The item understated the blindness: it disabled two signals, not one.
- Do not read this as "the other 18 are covered". It says only that they are not disabled by construction. Signals gated on text or OCR still depend on what the row happened to store, which is the separate fidelity caveat the calibration item already records for media rows.
- Until it is closed, changes to
PhotoCompositionSignalorImageContentAnalyzerneed their own evidence — the rank conversion was validated instead by unit tests, two-direction mutation tests, and corpus measurements over 300–400 images where every behaviour-changed file was judged individually. Say so explicitly in the commit, because "calibrate passed" will otherwise be read as covering it.
scripts/shared/clip_utils.py:75 pins MODEL_NAME = "ViT-B-32" with OpenAI pretrained weights. Meta's Perception Encoder (facebookresearch/perception_models, Apache-2.0 for both code and weights) is a substantially stronger CLIP-family model — and is already present in the installed open_clip 3.3.0 as a first-class architecture with a meta pretrained tag (PE-Core-B-16, PE-Core-S-16-384, PE-Core-T-16-384, PE-Core-L-14-336, PE-Core-bigG-14-448), verified 2026-08-10. So the model swap itself is a one-line change; everything expensive is downstream of it.
Status: Open — trial not started. Availability + config verified 2026-08-10, nothing measured on this corpus. Priority: P3 — largest available quality lever on the vision path, and the largest disruption. Do not swap without the A/B. Source: facebookresearch library audit, 2026-08-10
- This is not a free swap — three things break at once. (1)
PE-Core-B-16is 1024-d vs 512-d forViT-B-32, so.cache/clip_embeddings_v2/is invalidated (bump to_v3; the cache is already versioned for exactly this). (2)results/scene_probe.joblibis trained on 512-dViT-B-32features and must be retrained — the scene corpus is the gating cost, not the model. (3) ViT-B/16 at 224px is ~4× the vision FLOPs of B/32 (196 vs 49 patches), directly against the banked OCR/CLIP cost work. - The 512-d PE variants are not the cheap escape.
PE-Core-S-16-384andPE-Core-T-16-384keepembed_dim512 (so the cache dimensionality survives) but run at 384px — more compute than B-16@224, not less. If the trial is motivated by cost, none of these win; if it is motivated by accuracy, evaluate B-16 and accept the retrain. - The OCR gate moves too, and it is the sneaky one.
OCR_CLIP_GATE_TOPK = 3is a ranking heuristic over CLIP labels, calibrated onViT-B-32softmax behaviour (the flat-softmax finding: absolute probabilities are near-uniform, only rank separates text from photos). A different backbone changes that ranking distribution. Re-runscripts/eval_ocr_gate.pyas part of the A/B — a backbone that improves classification while degrading the gate could be a net cost regression. - Measure with the tools that already exist.
scripts/profile_pipeline.pyfor the wall/per-file and CLIP-vs-OCR hotspot split,scripts/backtest_scoring.pyfor decision agreement,make goldenas the correctness gate,eval_ocr_gate.pyfor the gate. Requires abackfill_clip_scores.pyre-run under the new backbone before the replay is meaningful — the replay serves CLIP/scene votes from the cache. - Reported zero-shot gap is large enough to justify the trial (PE-Core-B16 ≈ 78 vs
ViT-B-32/openai≈ 63 ImageNet zero-shot, upstream numbers, not measured here) — but upstream benchmark deltas have repeatedly failed to predict behaviour on this corpus. Treat them as the reason to run the A/B, never as its result. - Related, deliberately out of scope:
DINOv2(Apache-2.0) has stronger linear-probe features and would be the betterSceneSignalinput, but has no text tower and so cannot replace the zero-shotClipVisionSignal. That is a second-embedding question, not a backbone swap.MetaCLIP(CC-BY-NC 4.0) andImageBind(CC-BY-NC-SA 4.0) are license-blocked — non-commercial binds the employer, same finding as the ImageNet/LLD/InfographicVQA call.
tests/e2e/dashboard.spec.ts:42-56 asserts toHaveCount(4) against .feature-card plus four explicit .card-title filters. _site/index.html gained a fifth card ("Residence Galleries") in 205bd7e; the count assertion then failed while the title assertions stayed silent — they covered 4 of the 5 cards, so the new card shipped with zero content coverage and only the count noticed it existed. Numbers were corrected in 58f7659 (4→5, added the missing Residence Galleries title assert), but the shape is unchanged: the next card added breaks the suite again and lands untested until someone hand-edits the spec.
Status: Done 2026-08-11 (d97be75 cards, 1013673 stats bar + exact matching).
Priority: P3
Source: npm run test:e2e failure triage, 2026-08-11
- The failure is amplified 5×. The assertion lives in one test but runs under
chromium/firefox/webkit/mobile-chrome/mobile-safari, so a single stale literal reports as 5 failures and reads like a systemic break rather than one number. - Cheapest durable fix: hoist the expected titles into one fixture array, then assert
toHaveCount(EXPECTED.length)and loop the title checks. Adding a card becomes a one-line data change that cannot leave a card uncovered. - Keep both assertion kinds. The count is the only thing that catches an added card; the title filters only catch a removed or renamed one. Dropping either reintroduces a blind spot.
- The card set has changed without human intent before, which is the real case for the count assertion: the resolved
copy_to_site.shitem records a content run silently deleting this same "Residence Galleries" card from_site/index.htmlon 2026-07-26, orphaning_site/residence_gallery.htmlfrom navigation. That clobber path is fixed (1d3b262), so the count is not currently load-bearing against it — but it was the only assertion that would have caught it. - Shipped
d97be75:DASHBOARD_CARDS({ title, href }, DOM order) now drives the count, the titles, and the navigation tests, which are generated per entry. - This item understated the gap: navigation was uncovered too. The write-up above says only the count noticed the fifth card. In fact the four
should navigate to …tests were also hand-written, so "Residence Galleries" had no navigation test at all — the loop took the spec 4 → 5 nav tests without anyone adding one. When an item names one hardcoded list, check for its siblings before calling the fix complete; the same card set was enumerated by hand in three places, not one. - Both guards were verified to actually bite, in both directions: a phantom sixth fixture entry fails
toHaveCountand its generated nav test; a title renamed with the length unchanged lets the count pass and fails the text assertion. Without the second check the loop could have been vacuous — a green suite asserting nothing. hasText/getByTextare substring matches, and that was a live hole, not a theoretical one (fixed1013673). Injecting the strict prefix'Residence Galler'into the committedd97be75spec left it green. Replaced withtoHaveText(array), which matches each element's full text, and the list in order — so DOM order is now enforced rather than asserted in a comment. The same injection now fails.- The stats bar had the identical shape and got the same treatment in
1013673:DASHBOARD_STATSreplacestoHaveCount(4)plus four page-widegetByTextcalls, now scoped to.stat-label. - Why both assertion kinds survive, restated precisely: the counts run over
.feature-card/.stat-item(containers) while the text runs over.card-title/.stat-label(their children). A container with no label, or a stray label outside a container, fails exactly one.toHaveTextalone would not catch either. Still hand-enumerated: the threeFixed.tech-badgefilters.231019a, and the bullet above undercounted the gap. It called this "lower-stakes"; the footer actually has 8 badges, the test asserted 3 by substring with no count assertion at all, so 5 were entirely uncovered and a removed badge failed nothing.TECH_BADGES+toHaveText(array)now pins the set exactly and in order. Verified by deleting thePyTorchbadge from the page: fails now, passed before. A partial assertion with no count reads as coverage and is not — the card set at least had a count telling the truth about how many existed.Unfixed and outside the test suite:Fixed_site/index.htmlhas the same bug in CSS.231019a, and the "regenerated by content runs" claim above was wrong._site/index.htmlis the committed dashboard source:copy_to_site.shskips it explicitly ("_site/index.htmlis the maintained source") andupdate_site_data.pyonly rewrites.stat-valuenumbers and the Last Updated date via targeted regexes. So a direct edit is safe and there is no generator to push the fix into.- The CSS fix alone would have preserved the defect's shape, so the test is the fix. Rules now run past the current card count, but new test
every feature card has its own entry-animation delayis what makes that safe: it reads computedanimationDelayper card and compares against(index + 1) * CARD_STAGGER_STEP_SECONDSderived fromDASHBOARD_CARDS. Verified against the pre-fix CSS — reports0swhere0.5swas expected, i.e. it catches the exact bug that shipped in205bd7e. - Why this one hid for so long: it is presentation-only state no other assertion can observe. A missing
animation-delaybreaks no locator, no navigation, no text.:nth-childenumerations are invisible to ordinary E2E coverage and need an explicit computed-style assertion, which is the transferable lesson — the footer badge list and the card set were both catchable by normal assertions; this was not. - A count-agnostic CSS formulation was considered and rejected.
sibling-index()is not available in webkit/firefox, which this suite runs, and an inline--card-indexcustom property is ruled out by the project's no-inline-styles convention. Per-card rules are therefore unavoidable today; revisit ifsibling-index()lands across the matrix.
CLIPClassifier._similarities (scripts/shared/clip_utils.py:302-303) computes sims = img_emb @ txt_norm.T then sims.softmax(dim=0), omitting the model's trained logit_scale (≈100 for ViT-B-32). Softmaxing unscaled cosines (~0.15–0.35) is necessarily near-uniform: over the 94-label photo vocab the floor is 1.064% and winners land at 1.13–1.16%, i.e. ~8% above chance. This is the "flat-softmax" property already relied on in several places (W_CLIP as a ranking-only signal, CLIP_OCR_FALLBACK_THRESHOLD, _SCREENSHOT_OCR_KEYWORD_THRESHOLD, graphic_detection's 0.077 floor, the OCR_CLIP_GATE_TOPK note under the PE-Core item) — this entry records its cause and why it was not fixed.
Status: Open — deliberately not fixed; the relative-margin gate in 4b56759 works around it instead.
Priority: P3
Source: label-accuracy investigation of organize-files content, 2026-08-11
- Applying
logit_scaleis the principled fix and a wide blast radius. It would make absolute confidences meaningful for the first time, and simultaneously invalidate every threshold calibrated against the flat distribution — at least the five gates listed above plusCLIP_REFINEMENT_MIN_CONFIDENCE/ACCEPT_CONFIDENCE(0.15/0.30), which no current score can ever reach, so refinement is presently dead code on the photo profile. Sequence it as its own change withmake calibrate+ goldens, never bundled. - The workaround deliberately does not touch the distribution.
CLIP_MIN_LABEL_MARGIN_RATIO = 1.012gates on top-1/top-2 ratio, which is scale-invariant and so survives a later logit-scale fix unchanged (the ratio of two softmax outputs changes, but the ordering-based gate keeps its meaning). Contained, but it only suppresses bad names — it cannot produce good ones. - Confirm against the backbone A/B before investing. The PE-Core trial above would change this distribution anyway; doing both at once makes neither attributable.
- The blast radius above was incomplete: four more thresholds, in
src/analyzers/image_analyzer.py, and they were already dead. Found 2026-08-11 while converting the people gate to rank._PEOPLE_SCORE_LOW_THRESHOLD(0.15),_PEOPLE_SCORE_THRESHOLD(0.2),_INTERIOR_SCORE_THRESHOLD(0.3) and_SCREENSHOT_SCORE_THRESHOLD(0.4) are all measured unreachable — the global max across all labels is 0.1003 over 39 images against a 1/N floor of 0.0909, so they miss by 1.5×–4.0×.has_peoplehad silently degraded tohas_facesalone with two no-op clauses around it, andis_home_interior_no_peopleis permanently False. - This is where the git history of the flat softmax actually starts, and the cause is a refactor, not a design choice.
681c5ed(2025-12-09) scored viaoutputs.logits_per_image.softmax(dim=1)—logits_per_imageislogit_scale × cosine— and the inline comments read as genuine percentages ("30% confidence threshold", "20% confidence").33264df(2026-03-27) swapped transformers for sentence-transformers and, in its own words, removed "~60 lines of manual tokenization,logit_scalearithmetic, and processor boilerplate" while claiming to "keep exact public API". The API was kept; the value distribution behind it was not. Two days laterff36546extracted the now-unreachable numbers into named constants, which made them look deliberately calibrated. A refactor that preserves signatures can still invalidate every threshold written against the values, and nothing in the type system or the test suite notices. - Restoring
logit_scalesilently reactivates them — quantified. On the same 25 photos,people_scoregoes from 0.0903–0.0958 (0/25 above 0.15) to 0.0000–0.8896 (7/25 above 0.15). So the fix would wake a branch that has not executed since March and start routing ~28% of photos towardmedia/photos_social. That is a larger behaviour change than anything in the five gates originally listed here, and it arrives with no error and no failing test. - Two of the four are now immunised; two are not. The people and screenshot gates were converted to rank tests in
d699f45— softmax is order-preserving, so they are invariant to the scaling and will not move when it is restored._INTERIOR_SCORE_THRESHOLDwas deliberately left dead (converting it would resurrect a heuristicSceneSignalreplaced on 2026-07-18 and re-route photos toproperty_management/other), and_PEOPLE_SCORE_THRESHOLDlives in the caller-lessis_home_interior_no_people. Both still need handling in the same change that applieslogit_scale— decide them deliberately rather than letting the scaling decide for you. - The test suite passed throughout the dead period, and that is the transferable lesson.
tests/unit/test_image_analyzer.pyfed synthetic scores of 0.5 and 0.175 to a 0.15 threshold — values the real distribution cannot produce (max 0.1003). Green tests over a branch live data could never take. When a threshold's fixtures are hand-written numbers, the suite proves the comparison works, never that anything can reach it.
classify_image (scripts/shared/clip_classification.py) picks the winning category from score_embedding(emb, labels, "a photo of ") but then returns all_scores from score_embedding(emb, labels, "") — a second, unprefixed scoring pass. The two disagree in practice: PXL_20260723_231733185.jpg was categorized living room while all_scores' argmax was dining room, and IMG_9421.HEIC was backyard against an all_scores argmax of patio. So any consumer that reads all_scores to explain, validate, or re-derive a decision is reading a distribution that did not make it.
Status: Done 2026-08-11 — d4efff5. Both distributions are now explicitly named in CLIPResult.
Priority: P3
Source: margin-gate implementation, 2026-08-11
- Shipped (
d4efff5) — second option: keep both, name them distinctly.CLIPResultgainsdecision_scores(the"a photo of "prefixed pass that chose the winner) andraw_scores(the unprefixed pass).all_scoresbecomes a deprecated@propertyalias forraw_scores— carries the same value it always held, so external callers whose behaviour is unchanged; zero in-repo readers remain. Every in-repo call site was migrated to the explicit name. - Consumer audit:
result["top_scores"]inrename_images.py— migrated toclip_result.decision_scores. Top-5 now reflects the same distribution that picked the label, consistent with the winning category and confidence.files.image_classificationcolumn viascripts/backfill_clip_scores.py— no computation change needed.CLIP_CATEGORY_PROMPTSare full prompts built by_make_clip_prompt(they already embed"a photo of "per label).prompt_prefix=""+CLIP_CATEGORY_PROMPTSIS the decision distribution — semantically equivalent to what the production_clip_scores_for_contextpath does. Added a comment inbackfill_clip_scores.pymaking this explicit. Renamed local variable todecision_results. No mixed-provenance consequence: all existing rows were already written using the decision distribution; nothing stored was the unprefixed-label distribution. No re-run needed.- Calibration harness (
backtest_scoring.py) — readsimage_classification(full-prompt decision distribution, consistent with productionClipVisionSignal). The oracle's known limitations (164-row biased corpus) remain unchanged; cross-reference the "Scoring calibration is corpus-bound" item. The backtest does not replay the renamer'sdecision_scores/raw_scores— those are separate label sets for a separate pipeline.
- 8 new tests in
TestCLIPResultScoreFields— field shape, alias contract, different-ranking guarantee,classify_imagetwo-call structure with correct per-pass prefix.
RenamerProfile.min_label_margin defaults to 1.0 (disabled) and only PHOTO_PROFILE sets it to CLIP_MIN_LABEL_MARGIN_RATIO. The threshold was calibrated on 8 hand-labelled photos and does not transfer as-is: on a random 20-file sample from ~/Documents/Media/Photos/Screenshots, 75% fall below 1.012 while their chosen label still matches the folder they were previously filed into (terminal→Terminal, settings→Settings, at ratios as low as 1.0000). Enabling the gate there on today's evidence would suppress three-quarters of screenshot renames for no demonstrated accuracy gain.
Status: Open — intentionally left at 1.0; needs its own labelled corpus. Priority: P3 Source: margin-gate calibration, 2026-08-11
- The folder agreement is circular evidence — those folders were produced by this same classifier, so it cannot distinguish "the label is right" from "the label is consistently wrong". A real eval needs hand labels, as the photo threshold got.
- Screenshots mostly bypass the CLIP label anyway.
analyze_imageprefers an OCR title snippet (Screenshot_<title>) and only falls through togenerate_clip_filenamewhen no line qualifies, so the gate would apply to a minority of screenshots — measure that slice, not the whole folder. - A smaller vocab may want a different number, not the same one: 36 labels put the uniform floor at 2.78% versus the photo vocab's 1.06%.
The HEIC decode hand-off shipped 2026-08-11 (e9fb0a8 + 41ee326 + 1c32fe7), and the spurious _PREPROCESS_AVAILABLE gate that made docTR still raise without the preprocessing deps was removed the same day (7899442, 12 tests in tests/unit/test_ocr_heic_decode.py). One residual remains, and it is a measurement gap rather than a defect: the path is verified by shape, never by output.
Status: Done — measurement tooling shipped c266ace + 8cd1c7a (2026-08-11). A labelled HEIC corpus must still be assembled and the tool run; see the final residual below.
Priority: P3
Source: /backlog-implementer hand-off verification, 2026-08-11; re-opened 2026-08-11 after /backlog-migrate moved the parent item to docs/changelog/2.3.0/CHANGELOG.md with these residuals recorded only as changelog prose
The docTR branch is gated on an unrelated capability flag.Fixed7899442: removed_PREPROCESS_AVAILABLEfrom the condition in_run_image_ocr. The decode uses_load_rgb(PIL only), notpreprocess_for_ocr;_load_rgb'sexcept Exceptionguard returns None when PIL is absent, so the HEIC early-return fires instead of falling through toDocumentFile.from_images. 12 tests (2 new regression guards).TheFixednp.asarrayguard comment was misleading.c266ace(2026-08-11): the old comment said "PIL available but numpy not installed" — a state that cannot occur because numpy is imported before PIL in the same try block. Replaced with a defence-in-depth note explaining the import-order coupling and why reordering would make the guard reachable. Import-order note also added to the_PREPROCESS_AVAILABLEblock so the coupling is visible at both ends. Test docstring updated to match. The guard itself is kept: it is tested via monkeypatch and costs nothing to retain.- A type regression rode in on the follow-up commit.
f8aa49badded theDocTRResultProtocol and return hints, which tightened_run_image_ocr's signature without constraining the value satisfying it:_get_predictor()is untyped (docTR is an optional import), sopredictor(doc)isAnyand flowed straight to aDocTRResult | Nonereturn —mypyno-any-return, caught by the Stop hook, not by the commit's own checks. Fixed by annotating at the assignment (result: DocTRResult = predictor(doc)). Adding a Protocol to an optional-dependency module needs a check at the boundary where the untyped value enters, or the annotation is decorative. Recall was never measured — tooling was missing.Fixedc266ace+8cd1c7a(2026-08-11):scripts/eval_heic_ocr_recall.pyruns both easyocr and docTR against a directory of HEIC images, bypassing the CLIP gate so raw backend recall is measured independently. Supports optional ground-truth JSON (basename→bool), per-file verbose output, and JSON results export. Run:python scripts/eval_heic_ocr_recall.py --source DIR [--ground-truth labels.json] [--json results/heic_recall.json]. The key design point: the gate bypass is intentional — measuring with the gate on would skip most HEIC photos and conflate gate decisions with decode recall.- One open residual: the tool has not been run. A labelled corpus of text-bearing HEICs (screenshots of text, HEIC document scans) must be assembled and the script run before a number can be reported. The precondition for that run — that OCR reaches the backends on HEIC input — was never verified before the decode fix; the script makes it verifiable.
- Measure on text-bearing HEICs specifically, or measure nothing. The CLIP OCR gate (
--ocr-clip-topk, K=3) can skip OCR before either reader is reached, so a naive before/after on a photo-heavy corpus shows no change and reads as "the fix did nothing". The script bypasses the gate for exactly this reason. - What the fix does buy, unmeasured:
screenshot_ocr,text_contentandkie_structuredbecome available as voters on HEIC screenshots and document scans, which previously classified CLIP-only. Harmless for camera photos, which is most HEICs.
The EXIF fix in 4b56759 corrects new runs, but files organized before it retain filenames built from file mtime (download time) rather than capture time, and _maybe_rename_image will not revisit them: is_generic_filename is False for an already-descriptive name like 20260516_kitchen.heic, so the rename step returns early on every subsequent run. Confirmed on the two HEICs under ~/Documents — Media/Photos/Other/20260516_kitchen.heic was actually captured 2024-04-04, a 2-year error frozen into the name; Media/Photos/Products/20240426_fabric_sofa.heic happens to be correct.
Status: Name fixed 2026-08-11 — 20260516_kitchen.heic → 20240404_kitchen.heic, matching its
DateTimeOriginal of 2024:04:04 18:44:35. A full rescan of ~/Documents confirms 0 of 2 HEICs
now carry a stem date disagreeing with EXIF. Open: the category residual only (first bullet below).
Priority: P4
Source: post-fix verification sweep, 2026-08-11
- Scope was rescanned, not assumed, and the item's count held.
~/Documentscontains exactly 2 HEIC/HEIF files; only the kitchen one mismatched. Worth doing because an enumerated count in this file is a snapshot — the lint item's 540 had already decayed to 498 before anyone acted on it. - The DB carried the path in three fields, not one.
files.current_pathis the obvious one, butfiles.schema_dataandschema_metadata.schema_jsonboth embed it too, and a rename that updates onlycurrent_pathleaves two stale JSON blobs pointing at a file that no longer exists. Found by sweeping every column of every table for the old basename rather than by editing the column named in this item. Do that sweep before any manual path repair —reconcilestill has no path-repair operation to do it for you (same finding as theflutter_auth.txtstraggler). schema_data.filePathis historical and was deliberately left stale-looking. It reads~/Downloads/20260516_kitchen.heicwhilefiles.original_pathreads~/Downloads/IMG_8550.HEIC— the two disagree correctly, because the photo-profile renamer works in place, so the file was renamed inDownloadsand only then moved.filePathsnapshots the pre-move path at schema-generation time (file_processor.py:275);original_pathrecords the name before the renamer touched it. A blanket find-and-replace across the JSON would have rewritten it into a path that never existed. Only live pointers were updated:schema_data.name/contentUrl/urlandschema_metadata.contentUrl(itsnamestaysIMG_8550.HEIC, the original filename).- Verified no new provenance drift, by differencing against the backup. Unresolvable
current_pathrows went 7 → 6, the repaired one being exactly this file and the remaining 6 being the known-unrecoverable set. Backup atresults/file_organization.db.bak-heicrename-20260811(reference it by exact name — these backups sort by neither name nor mtime). - Still open: the category is pre-fix. The row is still
media/photos_otherbecause the lost GPS degradedmedia_heuristic; its own stored schema description reads "an interior room (5% confident)", and the calibrate-blindness item independently measured this exact file moving tomedia/interiors_otheronce EXIF is available. Deliberately not fixed here: there is no single-file--source, andBatchProcessor.scan_directoryisrglob("*"), so the--forcepass this item proposed would sweep all 25 entries ofMedia/Photos/Other, re-filing 23 unrelated files to fix one. Fixing the name without moving the file is the contained half; the re-file wants either a single-file source or a hand move. - Scale check before building anything: at two files this was a manual rename, as this item
originally judged. A repair pass is only worth writing if a bulk HEIC import lands, in which case
the rule is "re-derive the date prefix from EXIF when the stem's leading date disagrees with
DateTimeOriginal" — the rescan snippet used here is that rule, and it is three lines.
ContentOrganizer.get_destination_path calls dest_dir.mkdir(parents=True, exist_ok=True) (src/organizers/content_organizer.py) before any dry-run check, so a preview run writes empty folders into the target tree. Confirmed by birth time: ~/Documents/Media/Photos/Travel was created at 00:08:23 by a --dry-run --no-db pass and left empty, alongside Media/Place from an earlier one.
Status: Open — untouched.
Priority: P3
Source: 28-file ~/Downloads classification audit, 2026-08-12
- The contract says otherwise.
--dry-runprints "no files were moved" and QUICK_START tells users to always preview first, so the one command documented as safe is the one quietly mutating the target tree. - Empty dirs are not harmless here.
resolve_collisionand the taxonomy both read the filesystem, andreconcile --backfill-categoriesderives categories from what is on disk — a directory created by a preview is indistinguishable from one earned by a real filing. - Fix shape: thread
dry_runintoget_destination_path, or hoist themkdirto the move site inFileProcessor(shutil.moveis already guarded byif not dry_run). The latter is smaller and puts directory creation next to the only thing that needs it.
The run summary reports Files with extracted text: 0/28 on a batch where OCR demonstrably ran — one file's OCR text drove an ↪ OCR fallback: financial_statements (2%) decision in the same run. The counter reads the extracted_text slot returned by _detect_file_category_unified, which is ctx.text_if_loaded or ""; image OCR routes through ctx.ensure_ocr into ocr_if_loaded, a different slot. So images can never contribute to the count regardless of how much text was read.
Status: Open — reporting only, no misclassification.
Priority: P4
Source: 28-file ~/Downloads classification audit, 2026-08-12
- It cost real diagnostic time. The
0/28was the reason a PDF-extraction bug was initially reported that did not exist — the actual defect was an early-exit short-circuit. A stat that reads "extraction is broken" when extraction is fine sends the next reader down the same path. - Check the persistence path separately before "fixing" the counter.
_last_file_ocr_textis stashed for persistence, sofiles.extracted_textmay well be populated for images while the summary says zero; confirm which of the two is actually wrong before changing either.
TRAVEL_MIN_DISTANCE_KM = 100 and the home-distance gate that replaced GPS-presence detection (2026-08-12) are unmeasured. The 488-row replay corpus predicted media/photos_travel for exactly one row before the change and zero after; agreement was 268/488 both ways. The harness therefore shows no regression and provides no positive evidence.
Status: Open — the change shipped on mechanism plus three hand-checked files, not on corpus evidence. Priority: P3 Source: home-location travel gate, 2026-08-12
- The radius is a judgment call nobody has tested. At 100 km, San Antonio (118 km from Austin) is travel and Round Rock (28 km) is not. Nothing establishes that boundary is where a user would draw it; a labelled home-vs-trip slice would.
- Needs photos from outside the home metro to test at all. Every geotagged row in the current corpus appears to be local, which is why the distance gate is invisible to it. Sampling a trip's worth of photos is the cheapest way to get signal.
- Overrides exist and are untested end-to-end.
FILE_ORGANIZE_HOME_COORDINATES/FILE_ORGANIZE_TRAVEL_RADIUS_KM/FILE_ORGANIZE_HOME_NAMEare exercised by unit tests only; a malformed-coordinate override silently falls back to the Austin default, which is correct but unobserved in a real run. - Second homes and long stays break the single-point model. One coordinate pair cannot express "home is Austin, but also three weeks a year in X"; every photo from a repeated destination reads as travel forever.
A photo of a wall-mounted spice rack was labelled cabinet at a top-1/top-2 ratio of 1.0075 — below CLIP_MIN_LABEL_MARGIN_RATIO, so the rename was correctly withheld and the file went to Media/Photos/Other. The gate did its job; the vocabulary is the limit. "Spice rack" is the obviously correct description and no label in the 94-prompt photo vocab can express it, so the best available answer is a near-tie between furniture words.
Status: Open — correct behaviour given the vocabulary, wrong outcome for the user.
Priority: P3
Source: 28-file ~/Downloads classification audit, 2026-08-12
- This is the failure mode the margin gate converts, not removes. Withholding turns a wrong name into no name. That is the right trade, but a file that could have been
20240426_spice_rack.heicstill lands asIMG_9436.HEICin the generic bucket. - Adding labels is not free. The vocab is a softmax denominator — every prompt added lowers the uniform floor and shifts every ratio, so
CLIP_MIN_LABEL_MARGIN_RATIO(calibrated on 8 photos at the current vocab size) would need recalibrating. See the "CLIP scores are a softmax over raw cosines" item. - Measure the gap before expanding. Count how many withheld renames are vocabulary misses versus genuine ambiguity; if most withholds are near-ties between two correct-ish labels, more labels will not help.
classify_by_filename_patterns returns ("skip", "duplicate") for any stem matching _\d{8}_\d{6}$ — intended for timestamped copies, but IMG_20190327_172114 (the standard Android camera convention) matches it exactly. Scoring that stem directly yields skip/duplicate at confidence 1.0, early-exiting before any content signal runs.
Status: Open — currently masked, not fixed.
Priority: P3
Source: 28-file ~/Downloads classification audit, 2026-08-12
- It is masked by rename ordering, which is luck.
FileProcessor._maybe_rename_imageruns beforedetect_file_categoryand passes the descriptive name asdisplay_path, so the classifier usually sees20190327_lake.jpgand never the camera stem. Any path that classifies without renaming first — a non-generic filename, a renamer failure, a direct API caller — hits the rule. - The camera-timestamp fix does not cover it.
CAMERA_TIMESTAMP_STEM_PATTERNis anchored at both ends and guards the sprite rules; the duplicate rule uses a suffix match and runs earlier, soIMG_/PXL_-prefixed stems still reach it. - Verify the intended shape first. A genuine timestamped copy is
<original>_<YYYYMMDD>_<HHMMSS>, i.e. the timestamp is a suffix on an existing name. Requiring a non-empty, non-camera-prefix stem before the timestamp would separate the two cases.
Two pairs of files in the 28-file run were reported with identical destinations (20260723_bathroom.jpg, 20260723_bedroom.jpg). resolve_collision only tests dest_path.exists(), which is never true during a dry run because nothing is written.
Status: Open — reporting artifact; live runs are correct.
Priority: P4
Source: 28-file ~/Downloads classification audit, 2026-08-12
- Live behaviour was verified, not assumed.
get_destination_pathandshutil.moverun in the same loop iteration inFileProcessor, so the first file is on disk before the second's destination is computed and the suffix is applied. No data loss. - The report still misleads. Two of the four colliding files were genuinely different photos (a bedroom from two angles), so a reader reasonably concludes one would overwrite the other.
- Fix shape: have the dry-run path track planned destinations in a set and resolve against it, so the preview shows the names a real run would produce.
Recorded from the npm run repomix regeneration (docs/repomix/, gitignored) and git status on the primary checkout. Everything under "Uncommitted working tree" below is not yet committed — reconcile the affected open items when that work lands.
- Full repo pack (
repomix.xml) ≈ 983k tokens; compressed ≈ 604k; git-ranked ≈ 931k; docs-only ≈ 131k. src/~550k, but ~332k of that is one data dir —src/classifiers/data/census_names/(surnames.txtalone 315k tokens;given_names.txt16k). Hand-writtensrc/code is ≈ 218k.- Other buckets:
tests/~252k (unit ~200k, of whichtests/unit/scoring/signal tests ~43k),scripts/~117k (shared/~41k;filename_classifier.py15.4k is the largest script module),docs/~110k (largest single doc:reviews/SCRIPTS_SRC_DUPLICATION_AUDIT.md24.2k). - Largest source files:
organizers/content_organizer.py16.5k,storage/models.py12.4k,storage/graph_store.py12.2k,cli.py8.5k,cost_roi_calculator.py7.6k. Largest test:test_content_organizer.py16.3k. surnames.txtadded to.gitignoreand untracked (git rm --cached, staged) 2026-07-18. The file stays on disk locally, but once the deletion commits, fresh clones will lack it —person_name_validatordepends on it at runtime, so setup needsscripts/download_census_names.py(or equivalent) to materialize it. Resolved 2026-07-25: the script exists and is committed (e50ae06), documented in QUICK_START.md §1 and the CLAUDE.md Quick Start, and verified to reproduce the shipped gazetteer exactly (162,254 Census rows → 92,357 surnames atSURNAME_MIN_COUNT=200; the script's stale "~25k" comment was corrected).person_name_validatordegrades gracefully when the file is absent (_load_gazetteer→ None) andorganize-files healthreports thegazetteerlayer.
6f61449(2026-07-18) — added 16 images toresults/scene_labels/graphic/(graphic-class corpus labeling for the scene probe; see the graphic-probe item).eceb166(2026-07-17) — touchedcensus_names/surnames.txt;5f7f563(2026-07-01) — removed a license image (biometric PII) fromresults/test_set_augmentation/.- Six older commits (2025-12 → 2026-02) are all
_site/dashboard churn (metadata.jsonregenerations,metadata_viewer_backup.htmla11y fix). - Note: both 2026-07
refactor:commits are auto-typed wrong — they are data/corpus changes, not refactors.
git status 2026-07-18: 10 modified + 2 staged deletions (12 files, +373/−508):
- Deleted (staged):
src/scoring/signals/interior.py,tests/unit/scoring/test_signal_interior.py—InteriorSignalretired.SceneSignalalready registers unconditionally in committedregistry.py(no-ops when the artifact is absent). scripts/backtest_scoring.py— the artifact-gatedSceneSignal/InteriorSignalsweep-row swap collapsed to a plain("W_SCENE", W_SCENE, "scene")row.src/health_check.py+tests/unit/test_health_check.py—interior_probehealth feature replaced byscene_probe(checksresults/scene_probe.joblib; retrain hint now points atprototype_scene_probe.py).src/pipeline/file_processor.py+tests/unit/test_pipeline.py—_IMAGE_SCHEMA_TYPESwidened to every SceneSignal @type (House,Place,Accommodationjoin the Room family) so scene photos keep image metadata instead of falling through toDigitalDocument.tests/unit/scoring/test_signal_photo_composition.py— records thatPhotoCompositionSignal's interior (is_property_mgmt) vote was retired 2026-07-18; interior detection now belongs to SceneSignal exclusively.tests/unit/scoring/test_registry.py—TestSceneSwapreplaced byTestSceneSlot(unconditional registration; artifact pinned absent for hermeticity only).- Also touched:
src/organizers/content_organizer.py(formatting only),tests/unit/test_content_organizer.py,tests/unit/scoring/test_signal_scene.py.
On disk but untracked: results/scene_probe.joblib (34k, trained 2026-07-18 18:37). Scene-label corpus is far past the counts recorded in the graphic-probe item: interior 178, exterior 158, place 347, neither 73, graphic 34 (item snapshot was 44/21/3/46/8–12).
Open items affected when this lands:
- Trained graphic-vs-photograph probe — the "data-only" gap has largely closed (graphic corpus 34 and climbing; a 5-class
scene_probe.joblibexists untracked). The item's "do not commit an artifact without eval + backtest" guard still applies.