What we found
Ran jev-review's codebase-scan mode (npm run review:codebase) against a real 1,471-file TypeScript repo. Screening worked correctly — it flagged 1,609 real signals above the 0.7 threshold (spot-checked several: genuinely under-tested files). But the run produced zero final findings.
Root cause
MAX_FOLLOW_UPS = 8 in src/domain/config.ts is a flat constant, independent of codebase size or signal count. With 1,609 signals above threshold, only the top 8 were ever passed to the locate/evidence stage — 1,601 flagged signals were silently discarded, with no warning in the tool's own output that this happened.
Of the 8 that were inspected, the evidence-location step returned noMatch for 6 (confidence 0.18–0.61) and found a weak candidate region for the other 2, both below MIN_LOCATION_CONFIDENCE = 0.55. So all 8 were also discarded, at a different stage — leaving the run with nothing to report despite screening having done its job correctly.
We also noticed the code's own design intent (comments around the follow-up budget) suggests test-gap findings — "a fact about a whole file... no line to locate" — are meant to bypass the slot competition and be listed in full rather than compete for one of the 8 slots. That exemption does not appear to fire in codebase-scan mode (or codebase-scan routes through a different path than the one carrying it) — worth checking against src/review/codebase-judgments.ts (or wherever that mode's pipeline lives).
Numbers, for scale
- 1,471 files scanned
- 3,049 Jev API calls, 8,201,780 total tokens (7,918,564 in / 283,216 out)
- ~214 seconds wall clock
- 1,609 signals passed the screen threshold
- 8 signals actually inspected (the hardcoded cap)
- 0 final findings
Suggested fix
Scale MAX_FOLLOW_UPS with either the flagged-signal count or repo size (e.g. min(signalCount, someUpperBound) rather than a flat 8), and/or confirm the test-gap exemption from the slot budget is actually reachable in review:codebase mode. Happy to share the exact repro (a disposable local repo scan) if useful — didn't want to attach 8M-token debug logs here.
Secondary note (debug logging)
Separately: TYPESAFE_LOG_LEVEL=debug writes the SDK's raw request/response JSON via console.debug, which Node routes to stdout — the same stream review:changes/review:codebase use for their own JSON report. Turning on debug logging to see token usage corrupts the tool's own report output. Might be worth routing debug logs to stderr.
Diff-review mode (review:changes) worked cleanly for us on a small real diff — cheap (~4.4K tokens, ~1s) and correctly found nothing wrong with a clean bug-fix commit. This issue is specifically about codebase-scan mode.
What we found
Ran
jev-review's codebase-scan mode (npm run review:codebase) against a real 1,471-file TypeScript repo. Screening worked correctly — it flagged 1,609 real signals above the 0.7 threshold (spot-checked several: genuinely under-tested files). But the run produced zero final findings.Root cause
MAX_FOLLOW_UPS = 8insrc/domain/config.tsis a flat constant, independent of codebase size or signal count. With 1,609 signals above threshold, only the top 8 were ever passed to the locate/evidence stage — 1,601 flagged signals were silently discarded, with no warning in the tool's own output that this happened.Of the 8 that were inspected, the evidence-location step returned
noMatchfor 6 (confidence 0.18–0.61) and found a weak candidate region for the other 2, both belowMIN_LOCATION_CONFIDENCE = 0.55. So all 8 were also discarded, at a different stage — leaving the run with nothing to report despite screening having done its job correctly.We also noticed the code's own design intent (comments around the follow-up budget) suggests test-gap findings — "a fact about a whole file... no line to locate" — are meant to bypass the slot competition and be listed in full rather than compete for one of the 8 slots. That exemption does not appear to fire in codebase-scan mode (or codebase-scan routes through a different path than the one carrying it) — worth checking against
src/review/codebase-judgments.ts(or wherever that mode's pipeline lives).Numbers, for scale
Suggested fix
Scale
MAX_FOLLOW_UPSwith either the flagged-signal count or repo size (e.g.min(signalCount, someUpperBound)rather than a flat 8), and/or confirm the test-gap exemption from the slot budget is actually reachable inreview:codebasemode. Happy to share the exact repro (a disposable local repo scan) if useful — didn't want to attach 8M-token debug logs here.Secondary note (debug logging)
Separately:
TYPESAFE_LOG_LEVEL=debugwrites the SDK's raw request/response JSON viaconsole.debug, which Node routes to stdout — the same streamreview:changes/review:codebaseuse for their own JSON report. Turning on debug logging to see token usage corrupts the tool's own report output. Might be worth routing debug logs to stderr.Diff-review mode (
review:changes) worked cleanly for us on a small real diff — cheap (~4.4K tokens, ~1s) and correctly found nothing wrong with a clean bug-fix commit. This issue is specifically about codebase-scan mode.