Skip to content

rank identity candidates before counting their repositories - #15

Merged
nepinhum merged 1 commit into
masterfrom
rank-before-counting
Sep 16, 2026
Merged

nepinhum merged 1 commit into
masterfrom
rank-before-counting

Conversation

@nepinhum

Copy link
Copy Markdown
Owner

identity_candidates() computed the distinct repository count for every identity in the index, then threw all but twenty away. That count is the expensive one: EXPLAIN QUERY PLAN gives it MULTI-INDEX OR and a temp B-tree for the DISTINCT while the two counts it is ranked by are covering-index lookups.

Rank on the cheap counts first, ask for the repository count only for the twenty rows that survive. Same columns, same ordering, same values.

The old shape got cheaper as identities grew, because the per identity count then walked fewer commits each. Its worst case is a small team with a long history which is gitlife's subject.

fixture before after
200 identities, 200k commits 0.168–0.199 s 0.079–0.082 s
2000 identities, 200k commits 0.029–0.030 s 0.018 s

summary end to end, on an index built from a 20k commit repository with 50 authors: 0.03–0.04 s to 0.01–0.02 s.

Closes #9

@nepinhum
nepinhum merged commit 381b6ad into master Sep 16, 2026
0 of 2 checks passed
@nepinhum
nepinhum deleted the rank-before-counting branch September 16, 2026 18:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf: identity candidate query scales poorly with indexed identities/history

1 participant