Generate an AI-ready knowledge graph of any codebase — so an AI agent (or a new developer) can read one file and start building immediately.
repokg extracts everything that can be known deterministically about a repo —
module inventory, internal import graph, every branch classified against every PR
(merged / squash-merged / abandoned / stale), contributor stats, CI/Docker/Helm/Make
surface — and renders it as:
KNOWLEDGE_GRAPH.md— a single human/AI-readable document with a mermaid architecture graph, module tables, branch & PR catalog, timeline, and ops inventory..repokg/kg.json— the same graph, machine-readable.
The semantic layer (module purposes, data-flow narratives, project eras, gotchas)
can't be produced by static analysis without guessing — so repokg is
agent-first: it emits .repokg/prompts/enrich.md, a rigorous prompt any AI coding
agent (Claude Code, Cursor, Copilot Workspace…) executes to verify-and-fill the
narrative sections, writing .repokg/narratives.json. Re-render and the knowledge graph is
complete. No API keys, no LLM dependency in the tool itself.
pipx install repokg # or: pip install repokg
# from source:
pipx install git+https://github.com/NehharShah/repokgRequirements: Python ≥ 3.9, git. Optional: gh (logged
in) for the PR/branch cross-reference — without it the knowledge graph still builds, minus PR data.
cd your-repo
repokg # = generate: scan + prompts + renderOutput:
.repokg/kg.json # machine-readable knowledge graph
.repokg/prompts/enrich.md # hand this to your AI agent
KNOWLEDGE_GRAPH.md # the knowledge graph document
Then, in your AI agent of choice:
Follow the instructions in .repokg/prompts/enrich.md
The agent explores the code, writes .repokg/narratives.json, and runs
repokg render — KNOWLEDGE_GRAPH.md now carries verified purposes, data flows,
timeline eras, and gotchas alongside the deterministic structure.
| Command | Effect |
|---|---|
repokg scan [path] |
Extract structure → .repokg/kg.json |
repokg prompts [path] |
Write the enrichment prompt |
repokg render [path] |
kg.json (+ narratives.json) → KNOWLEDGE_GRAPH.md |
repokg generate [path] |
All three (default) |
repokg inject [path] |
Wire the knowledge graph into CLAUDE.md / AGENTS.md / Cursor rules (--diff for dry run) |
repokg audit [path] |
Show every inferred conclusion with confidence + evidence (--json for machines) |
repokg clean [path] |
Remove everything repokg authored — never touches your content (--diff for dry run) |
repokg check [path] |
Exit 1 if the knowledge graph is stale vs HEAD (CI-friendly) |
repokg diff [path] |
Report what changed between two graphs; exit 1 if the shape changed |
Flags: --out DIR (default <repo>/.repokg), --md FILE (default <repo>/KNOWLEDGE_GRAPH.md),
--exclude PATTERN (repeatable), --no-github, --no-cache, --pr-limit N, --diff, --json,
--from KG.JSON, --to KG.JSON, --from-ref REF, --to-ref REF,
--format text|json|md, --no-renames.
check answers "is the graph stale?". diff answers what actually changed:
$ repokg diff .
comparing two graphs of 46e6f96f1a2b (the same commit)
modules +1 ~1
+ billing Python, 4 files, 212 lines
~ api loc 980 -> 1044
edges +1
+ api -> billing (Python) 3 imports
languages ~1
~ Python files 120 -> 124, loc 18400 -> 18612
shape changed: modules, edges (exit 1)
That new edge is the point — "this adds a dependency from api/ to billing/,
intentional?" is the review comment a raw file diff will not give you.
With no arguments it compares the graph the last scan left in .repokg/kg.json
against a fresh scan. The stored graph is never written over — it is the
baseline, so saving over it would destroy the answer to the next question.
(.repokg/cache.json is still updated, since it records what each file
contained rather than what the graph concluded, and is what keeps the scan
feeding the diff fast. repokg clean removes it either way.) Point --from and
--to at saved graphs to compare two of them without scanning at all.
--format md emits a report ready to paste into a PR comment, and --format json prints the whole delta keyed by record, uncapped, with scan progress moved
to stderr so it stays pipeable.
--from-ref and --to-ref scan the repo as it was at a git ref — a branch,
tag, sha or anything git rev-parse accepts:
repokg diff . --from-ref main --to-ref HEAD # what this branch changes
repokg diff . --from-ref v0.4.1 --to-ref v0.4.3
repokg diff . --from-ref main # main vs your working treeEach ref is laid out with git worktree add --detach into a throwaway
directory, scanned, and removed — including when the scan fails, since a leaked
worktree stays registered in .git and would show up in every later git worktree list. A ref scan writes no cache: the cache keys on HEAD and on file
mtimes, and a checkout of an old commit matches neither, so it is switched off
rather than allowed to overwrite the one your real scans depend on.
Refs get their own flags instead of being accepted by --from/--to because a
file named main and a branch named main are both legal, and guessing which
you meant is the kind of inference this project avoids everywhere else.
Two things worth knowing. Comparing a ref against your working tree is
asymmetric — the working tree holds untracked and ignored files that no commit
does — so the report says so in a note rather than leaving you to deduce it from
a surprising addition. And a ref scan is always cold by nature, so pass
--no-github when you do not need the PR list; it is the slowest part of a
scan and it cancels out on both sides anyway.
In CI, note that actions/checkout fetches depth 1 by default. A shallow clone
does not contain the base branch, so ref diffs need fetch-depth: 0 — repokg
detects the shallow case and says so rather than failing obscurely.
Exit codes follow diff(1) and git diff --exit-code: 0 unchanged,
1 the shape changed, 2 error. Shape means the membership of modules,
edges, languages and the ops surface — a module or dependency appearing or
disappearing, or a module switching primary language. LOC drift, import counts,
branch tips and PR states are all reported but deliberately do not move the exit
code, because they change on essentially every commit and a gate that fired
every time would be switched off within a week.
Not everything that differs is reported, either. A branch's tip, date and commit subject move whenever anyone pushes, so comparing them would bury the transitions that matter — a branch going stale, or merging — under noise from unrelated work. A section recorded by only one of the two graphs is skipped and noted rather than reported as wholly added, since a version gap is not an architectural change.
A module is identified by its path, so moving one reads as a removal plus an addition. Pairing those back up is the diff's only heuristic, and it says so:
modules R1
R lib -> shared high, 1 of 1 imports across unmoved modules preserved
The evidence is the import neighbourhood, not size similarity — a moved
module keeps its dependencies, whereas two modules having a similar line count
is a coincidence waiting to happen. Confidence is high when every import to
and from the unmoved parts of the graph survived, medium when most did or the
directory name is unchanged, and low when the name is the only thing matching.
An ambiguous pairing — a module split in two, or two candidates for one move —
is not reported at all, and a language change is never a rename.
The pair stays in added and removed as well, so a consumer that distrusts
the pairing can ignore it and see exactly what it would have seen otherwise.
--no-renames turns it off. A rename still exits 1: no dependency changed, but
every path a doc, an agent or a CODEOWNERS entry referenced did.
scan caches what it extracted from each file in .repokg/cache.json and replays
it for files that have not changed, so a re-scan only parses what moved:
$ repokg scan .
cache: cold (no cache yet) — parsed 431 files
$ repokg scan .
cache: replayed 431 of 431 files, parsed 0
$ vim src/app.py && repokg scan .
cache: replayed 430 of 431 files, parsed 1
A file is replayed only when git has not flagged it — git diff against the
commit the cache was written at, plus git status for anything dirty, staged, or
untracked — and its size and mtime still match what was recorded. The second
check is what covers the files git cannot speak for: anything in .gitignore
that repokg still walks, submodule contents, or a directory that is not a git
checkout at all.
The cache is an optimization and nothing else. Output is identical either way
(a test enforces byte-for-byte equality), a missing or unreadable or unusable
cache degrades to a full scan and says so on stdout, and --no-cache forces one.
repokg clean removes it with everything else.
One caveat, shared with every build tool that trusts stat: a file rewritten
with its size and modification time preserved is invisible to both checks. Use
--no-cache if you have reason to think that happened.
Synthetic mixed Python/TypeScript/Java monorepos, one laptop, best of three — indicative only, not a benchmark suite (see #17):
| files | LOC | edges | cold | warm | speedup | cache.json |
|---|---|---|---|---|---|---|
| 300 | 30k | 21 | 144 ms | 121 ms | 1.2× | 0.1 MB |
| 1,500 | 228k | 111 | 336 ms | 129 ms | 2.6× | 0.3 MB |
| 6,000 | 1.2M | 450 | 1,316 ms | 176 ms | 7.5× | 1.0 MB |
The ratio grows with the repo because the saving is essentially the whole parse cost, while the warm floor does not shrink. That floor is worth knowing before optimising further: at 6,000 files it is roughly 100 ms of git metadata (branches, contributors, merge classification — none of which touches the cache) and 30 ms of walking, stat'ing and writing the output. Parsing, the part the cache removes, was around 1,150 ms of the cold scan.
Common noise (node_modules, .git, build output, …) is skipped automatically.
For repo-specific noise — fixtures, snapshots, vendored trees, generated docs —
add globs on the command line or in a committed .repokgignore at the repo root:
repokg scan --exclude '*fixtures' --exclude 'docs/gen'# .repokgignore — one glob per line, same semantics as --exclude
*fixtures
*.snap
packages/*/gen
Patterns are matched (fnmatch) against repo-relative paths; matching
directories are pruned wholesale and matching files dropped, so modules, import
edges, and ops all inherit the exclusion. * crosses /: *fixtures matches
at any depth, fixtures only at the root. CLI patterns and .repokgignore are
unioned. Exclusions are never silent — scan prints what it dropped, kg.json
records the patterns and counts, and repokg audit carries an uncertainty note.
Most of the graph is measured fact. The parts that are heuristics are labeled
as findings with confidence and evidence, surfaced by repokg audit:
[git]
trunk = master high detected via origin/HEAD symref
integration = staging medium matched a well-known integration branch name
[modules]
4 flagged generated low path-name heuristic; verify before excluding
repokg diff carries the same discipline. Its one heuristic — pairing a removed
module with an added one to call it a move — reports the confidence it matched
at and the evidence for it inline, refuses ambiguous pairings outright, and can
be turned off with --no-renames.
Agent-written narratives.json is schema-validated before rendering — malformed
enrichment fails loudly with errors precise enough for the agent to self-correct.
(Findings/confidence design inspired by RepoCanon.)
repokg inject adds a managed block (delimited by
<!-- repokg:begin/end -->, idempotent, never touches your hand-written
content) pointing agents at KNOWLEDGE_GRAPH.md:
CLAUDE.md(Claude Code) — updated if presentAGENTS.md(the cross-tool agent standard) — updated if present, created if no agent file exists at all.github/copilot-instructions.md(Copilot) — updated if present.cursor/rules/repokg.mdc(Cursor, withalwaysApply: true) — created if.cursor/rules/exists; falls back to legacy.cursorrules
Keep it fresh in CI, either with the action:
- uses: actions/checkout@v4
- uses: NehharShah/repokg@v0.4.5That fails the job when the committed graph no longer matches HEAD. Set
strict: false to annotate a warning and pass instead, which is the honest
setting while you are adopting this. The stale output is set either way, so a
following step can react without re-running anything:
- uses: NehharShah/repokg@v0.4.5
id: kg
with:
strict: false
- if: steps.kg.outputs.stale == 'true'
run: echo "${{ steps.kg.outputs.result }}" >> "$GITHUB_STEP_SUMMARY"The action installs repokg from its own checkout rather than from PyPI, so
@v0.4.5 runs repokg 0.4.5 by construction — there is no second version to
keep in sync and no way for the two to drift apart.
Check mode needs a committed graph. It compares the HEAD recorded in
.repokg/kg.json against the actual HEAD, so a repo that gitignores
.repokg/ has nothing to compare and will report stale on every run. Commit
.repokg/kg.json and KNOWLEDGE_GRAPH.md if you want this check to mean
anything. Or without the action:
- run: pipx run repokg check . || echo "::warning::KNOWLEDGE_GRAPH.md is stale"mode: comment reports what a PR changes architecturally, as a PR comment and
a job summary. It needs nothing committed:
on: pull_request
permissions:
contents: read
pull-requests: write # or the comment cannot be posted
jobs:
diff:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0 # the base branch has to be in the clone
- uses: NehharShah/repokg@v0.4.5
with:
mode: commentOne comment per PR, found by a hidden marker and edited in place, so a long branch does not accumulate a running commentary of its own history. A PR with no structural change gets no comment at all — but if an earlier push already posted one, it is rewritten rather than left asserting a finding the latest push undid.
It compares the base branch against HEAD as two commits, not against the
working tree. Both sides are then built the same way, so nothing in the report
is an artefact of how it was produced. changed is true only when the
shape moved; LOC drift alone leaves it false and posts nothing.
The report is also written to the job summary, and report-path points at a
file holding it — for uploading as an artifact or sending somewhere else.
Read the file, not $GITHUB_STEP_SUMMARY: GitHub gives every step its own
summary file, so a later step cannot see what this one wrote.
A shape change never fails the job — a PR that adds a module is the normal
case, not an error. Only repokg failing to run does, which is what the
fetch-depth: 0 above prevents: a shallow clone has no base branch, and
repokg says so by name. Posting the comment is best-effort — a PR from a fork
gets a read-only token, so that case warns and leaves the report in the job
summary rather than failing a contributor's build.
Or surface the architectural change a PR makes, which is what the three exit codes are for — a mistyped path must not read as a new dependency:
- uses: actions/checkout@v4
with:
fetch-depth: 0 # ref diffs need history; the default is depth 1
- run: |
rc=0
pipx run repokg diff . --no-github \
--from-ref "origin/${{ github.base_ref }}" --format md > diff.md || rc=$?
if [ "$rc" -gt 1 ]; then exit "$rc"; fi # 2 = the diff failed, not the graph
if [ "$rc" -eq 1 ]; then cat diff.md >> "$GITHUB_STEP_SUMMARY"; fiThat compares the base branch against the checkout, so it reports what the PR
changes rather than what has happened since someone last ran a scan. Drop the
--from-ref to compare against a committed graph instead.
KNOWLEDGE_GRAPH.md itself also lists any agent-context files it found, so an agent landing on the knowledge graph discovers your rules — and vice versa.
| Area | How |
|---|---|
| Branch classification | git for-each-ref + --merged ancestry vs the integration branch (auto-detects staging/develop), cross-referenced with every PR's head ref via gh — distinguishes true merges from squash-merges from abandoned work |
| PR catalog | gh pr list --state all — open / merged / closed-unmerged, full appendix table |
| Module inventory | Filesystem walk with LOC per directory, language detection, generated-code flagging |
| Import graph | Go: import blocks resolved against go.mod module paths · Python: stdlib ast incl. relative imports · JS/TS: import/require resolution — relative paths, tsconfig/jsconfig paths + baseUrl aliases (nearest config wins), npm/yarn/pnpm workspace package names · Rust: use declarations resolved against Cargo crate names (cross-crate) and src/ module trees (intra-crate) · Java/Kotlin: imports resolved by longest prefix against package declarations. Directory→directory edges with counts |
| Ops surface | CI workflow names, Dockerfiles, compose files, Helm charts, Makefile targets, config/docs/test/migration dirs |
| Timeline | Merged PRs grouped by month with conventional-commit scope frequencies (replaced by agent-written eras after enrichment) |
Because the enrichment quality depends on reading the code, and your coding agent
already has the repo open, tools to search it, and your permission model. A prompt it
can execute beats a second LLM integration with its own keys, costs, and context limits.
The contract between tool and agent is one JSON file (narratives.json) with a fixed
schema — everything else stays deterministic and reproducible.
- JS/TS: relative imports, tsconfig/jsconfig
paths/baseUrlaliases and workspace package names are resolved;extendschains are not followed (a leaf config without its own aliases is skipped rather than shadowing the root's), and packageexportsmaps are not modeled — subpath imports fall back to the package dir. Alias imports whose targets ground nowhere are counted in an uncertainty note. - Fork PRs: a fork PR whose head branch name matches a local branch will be linked to it (GitHub's API reports bare head refs).
- Python: packages are discovered at the repo root and under
src/; deeper monorepo layouts (packages/*/src/…) get file-level edges only. - Rust:
usedeclarations only — macro-generated imports, re-export chains, and[dependencies] path = …(non-workspace) crates are not resolved;crate::paths ground only in module dirs/files that exist. - Java/Kotlin: explicit imports only — same-package references (no import needed) and fully-qualified inline names produce no edges; when a package is declared only in test roots, edges resolve there.
- Branch
aheadcounts use one batched git call on git ≥ 2.41, with a per-branch fallback on older git. - Incremental scans cache per-file facts, not repo metadata:
go.mod,Cargo.toml,tsconfig.json,package.jsonandpnpm-workspace.yamlare re-read every scan (one per package, against one read per source file).
- Rust import graph
- Java / Kotlin import graphs
-
--excludeglob patterns +.repokgignore - Incremental scan cache for large monorepos
-
repokg diff— structural diff between two scans, or two git refs -
llms.txtemission alongside KNOWLEDGE_GRAPH.md - tsconfig
pathsalias + workspace package resolution - PyPI release + GitHub Action (staleness check, PR structural diffs)
pip install -e .
python -m unittest discover -s tests -vNo runtime dependencies — stdlib only.
All work goes through issue → branch (issue-<N>/<desc>) → PR → review → squash-merge to main.
See CONTRIBUTING.md for the workflow and the ground rules
(zero deps, findings for heuristics, clean reversibility).
MIT