Skip to content

[GPT QA] Calibration apply re-queries every eligible song to recover paths already loaded by state discovery #207

Description

@xiaden

Finding

Starting the supported calibration-apply workflow performs an N+1 song lookup over the entire eligible calibration population before any calibration work begins. State discovery already materializes each matching SongStateCandidate with its song record, but get_uncalibrated_tagged_song_ids() discards those records and returns only identities; LibraryQueryMixin.get_paths_needing_calibration() then calls db.library.get_song(locator) once for every identity solely to recover each path.

User/System Journey

After calibration generation completes, the composition root can automatically trigger TaggingService.start_apply_calibration_background(). The managed task enters _run_apply_calibration() -> tag_library() -> library_service.get_paths_needing_calibration() before apply_calibration_wf() starts processing files. For each enabled library, get_uncalibrated_tagged_song_ids() loads the processed-state candidates and the not-calibrated-state candidates, intersects their semantic identities, and returns those identities. The processed-state candidates already carry candidate.song, but the caller then loops every returned identity through db.library.get_song(locator) to obtain song.path.

For N eligible songs, calibration startup therefore performs N additional point reads after the state queries have already materialized song records. A 50,000-song library requiring a first calibration or bulk recalibration incurs about 50,000 avoidable database round trips before the actual per-file calibration work begins. Multiple enabled libraries preserve the same per-song amplification over the combined eligible population.

Files in scope

  • nomarr/services/domain/library_svc/query.py — get_paths_needing_calibration() collects semantic identities, then calls db.library.get_song() once per identity.
  • nomarr/components/library/library_song_state_comp.py — _state_candidates() returns SongStateCandidate objects containing song records; get_uncalibrated_tagged_song_ids() discards those records after intersecting processed and not-calibrated state membership.
  • nomarr/services/domain/tagging_svc/apply.py — tag_library() is the production calibration-apply caller and pays the full lookup cost before invoking apply_calibration_wf().

Expected behavior

Calibration eligibility discovery should produce the paths or song records needed by the apply workflow without issuing a persistence lookup per eligible song. After the bounded state-membership queries, path projection should be in-memory or use a bounded batch lookup if semantic re-resolution is required.

Actual behavior

The state query materializes eligible song records, reduces them to identities, and the service immediately re-resolves every identity through an individual get_song() call. Database round trips therefore grow linearly with every song requiring calibration in addition to the state-discovery work; at 50,000 eligible songs this adds about 50,000 point queries before calibration processing starts.

Correction direction

Preserve the already-materialized song/path data across the calibration eligibility boundary, or provide a batch projection that resolves all eligible identities in bounded persistence work. Keep semantic library scoping and the processed + not-calibrated intersection invariant intact; remove only the per-song re-fetch required to reconstruct paths. Add query-count coverage showing that calibration-startup persistence calls do not scale one-for-one with eligible song count.

Metadata

QA-Agent: ChatGPT
Reviewed-SHA: 3913c63
Branch: feat/develop-branch-migration
Severity: Medium
Category: performance
Review-Mode: performance-head
QA-Fingerprint: xiaden/nomarr|performance|calibration-apply-path-discovery|per-eligible-song-path-reresolution

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    category:performanceExcessive CPU, GPU, memory, I/O, API, database, or scaling cost.qa:chatgptIssues discovered by automated ChatGPT QA review.severity:mediumMaterial defect with meaningful impact but bounded scope or a viable workaround.source:performance-qaOriginally discovered by the scheduled performance QA review.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions