Skip to content

About

Score your team's AI-assisted development maturity from git history + agent config, against published industry research. 9 dimensions, PDF report, quarterly baselines.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

ai-dev-maturity

How good is your team actually getting at building software with AI — and what to fix next.

selftest License Python Claude Code Stars

platforms assertions deps offline

English · 简体中文

Report cover with 9-dimension radar chart

Sample report on synthetic data. Report language follows your input — chart and section labels are overridable via labels in report.json.


Reads your git history and agent configuration, scores your team on 9 dimensions against published industry research, and produces a shareable PDF report with a concrete, landable improvement plan.

Works with any coding agent — Codex, Cursor, Copilot, Gemini CLI, Cline, Amp, Claude Code. Scanning and rendering are pure Python (stdlib only) and call no model API; the scoring step is done by whichever agent you already use. Cross-tool entry point: AGENTS.md; Claude Code users read SKILL.md (same content).

Works on any git project — single repo or multi-repo workspace; GitHub, GitLab, Gitee, CODING or Bitbucket. No CI integration required, nothing uploaded anywhere: it reads your local clone.

Why this exists

Every team adopting AI coding tools eventually asks "are we actually good at this, or just busy?" The usual answers are poor ones — vendor benchmarks measure someone else's team, and "lines written by AI" measures activity, not outcome.

This answers a narrower, more useful question: is your way of working with AI built to survive scale, and where will it break first? It scores the harness — specs, context, gates, review, measurement — not the people.

What you get

Output Content
scan.json ~40 metrics per repo: throughput, AI attribution, PR size, test ratio, CI gates, agent config, spec-layer health
Markdown analysis Evidence → 9-dimension score → three-wave plan, every item with a concrete landing point and an acceptance metric
HTML + A4 PDF Radar, trend lines, bars, donuts, roadmap. Print-ready, with automatic pagination self-check
Baseline Stored per project, so the next run diffs against it (--compare)
Evidence page with charts

One snapshot will lie to you

--trend 12 buckets the same git history by month, so you can see throughput and rework move together instead of reading a single 90-day average.

Throughput vs rework over 13 months

Illustrative, synthetic data.

Read the two lines together: throughput up with rework flat is real gain; throughput up with rework up means you are buying speed with rework, and the only question is how much of the gain survives.

This is not hypothetical. On the project this tool was first built for, the 90-day snapshot supported an optimistic reading. The 14-month view said the opposite — rework in two long-lived repos had gone from a 10–20% baseline to over 60%, both in the same month. Same data, opposite conclusion, purely because of window length.

On the "J-curve". DORA and the productivity-economics literature describe technology adoption as a J-curve — things get worse before they get better, and the dip is paid in verification cost. It is a useful metaphor, not a standard: there is no agreed measurement, no threshold, and no benchmark dataset for it. Do not report "we are on the upslope" as if it were a grade. If you want something standardised to compare against, that is DORA's four keys (deployment frequency, lead time, change failure rate, time to restore) — a decade of research and real benchmarks behind them. The value of the chart above is narrower and more reliable: it stops a short window from fooling you.

Rework has no pass mark. fix/(fix+feat) is a composition ratio, not a quality score. A mature product in maintenance should be mostly fixes; a greenfield project mostly features; a team in a pre-release stabilisation push will spike, and that is the system working. The number is only meaningful against your own history, with lifecycle held roughly constant — and even then it is a prompt to go ask what happened, never a verdict. It also depends on commit-message discipline: check what share of commits are classifiable at all before trusting a trend.

Quickstart

Clone it anywhere:

git clone https://github.com/yuna78/ai_dev_maturity
python3 -m playwright install chromium     # only needed for PDF rendering

Claude Code users: to have it auto-discovered, clone to ~/.claude/skills/ai-collab-maturity — that directory name is what Claude Code looks for, even though the repo is named ai_dev_maturity. Any other agent: no specific path needed — point it at the repo and have it read AGENTS.md.

Easiest install: hand the repo URL to whatever coding agent you already use and say "install this skill".

Then, from any git project:

S=~/.claude/skills/ai-collab-maturity/scripts

# 1. Sanity-check what it detected (branch, merge convention)
python3 $S/scan.py --root . --detect-only

# 2. Scan (+ monthly trend data). Self-check warnings go to stderr, so they
#    stay visible and never contaminate scan.json
python3 $S/scan.py --root . --trend 12 > scan.json

# 3. Score and write up: in Claude Code, ask it to "assess our AI collaboration
#    maturity". The skill reads scan.json, scores each dimension, writes report.json

# 4. Render HTML + PDF (report.json comes from step 3)
python3 $S/build_report.py report.json --out .

Step 3 — turning scan.json into judgement — is the agent's job, not something a script can do for you. Open the project and ask your agent to assess our AI collaboration maturity; it reads AGENTS.md / SKILL.md and walks through scoring against the rubric.

Pre-scan self-check

Every run checks the handful of things most likely to make a report silently wrong. Hits go to stderr (so they stay visible when stdout is redirected) and into scan.json as warnings[] plus a per-repo repos[].warnings:

code Condition Why it matters
stale_remote origin/<branch> HEAD older than 30 days Remote hasn't moved — usually a missing git fetch / git push
unpushed Local HEAD ahead of origin/<branch> Those commits are outside the scan; throughput and AI attribution read low
empty_window Zero commits in window, but the branch has commits Every throughput metric will be 0 — wrong branch or too narrow a window
guessed_branch origin/HEAD doesn't point at the chosen branch The mainline was guessed by activity and may be wrong
no_remote No origin/<branch>; scanning a local branch Different basis than a reviewed remote mainline

Why this deserves a table: in any of these cases scan.py still emits a perfectly well-formed scan.json in which every number is zero. Without the warnings, anything downstream scores it as "this team doesn't use AI" — on a report that looks immaculate.

For CI or the cautious: --strict exits 1 when any error-level warning fires.

python3 $S/scan.py --root . --strict > scan.json

Comparing quarters:

python3 $S/scan.py --compare baseline-2026Q1.json scan.json
Three-wave roadmap

The 9 dimensions

Intent & specs · Context engineering · Feed-forward constraints · Feedback & verification · Multi-agent orchestration · Scaling human oversight · Measurement & economics · Security by default · Adoption beyond engineering

Scored 1–5, where 3 is the watershed: a mechanism that exists but relies on discipline scores 3; one enforced deterministically by a hook or CI scores 4; enforced and measured scores 5. Anchors and evidence sources in references/maturity-rubric.md.

Grounded in Anthropic's 2026 Agentic Coding Trends Report · DORA's ROI of AI-assisted Software Development · Thoughtworks Technology Radar Vol.34 (cognitive debt) · OpenAI's harness engineering · Microsoft's CLI-agent rollout study · Stack Overflow Developer Survey — summarised in references/paradigms-2026.md, with a refresh checklist, because the summary has a shelf life.

What this is not

  • Not an ROI calculator. A high score means you are likely doing it right, not that it paid off this quarter. How to measure returns yourself, and which metrics will mislead you: references/measuring-roi.md.
  • Not a performance review. The report names individuals — who writes, who merges, who signs commits with AI. Concentration of merge authority is a real risk signal and anonymising it hides the finding. But commit share is an activity metric, and reading it as productivity is the exact mistake the rubric warns against. Decide who sees the report before you run it.
  • Not a vendor benchmark. The only target that matters is your own previous baseline.

Known limits

  • AI attribution is a floor, not a rate. Only some tools leave a Co-Authored-By trailer. The scanner also mines branch names (codex/*) and reports both — a wide gap means a tool in use is invisible to the trailer count.
  • Manual checks exist. Server-side branch protection, real tool usage, rework counts and non-engineering adoption cannot be read from git. --manual manual.json keeps those answers across quarters instead of losing them in a chat log.
  • Self-check the scanner's own output. spec.time_source_warning, an unlinked count equal to 0 or to the total, a wide ai_by_branch / ai_by_tool gap — suspiciously tidy numbers are usually a rule that failed to match, not reality.

Privacy

Everything runs locally against your own clone. Nothing is uploaded, no telemetry, no network calls except the optional Chromium download for PDF rendering.

Requirements

git · python3 3.9+ (standard library only for scanning) · playwright + chromium for PDF · pdf2image optional, adds orphan-page detection to the layout self-check.

Development

python3 scripts/selftest.py -v    # 36 assertions across 5 git platforms

Run it after touching any detection rule. It builds throwaway repos with GitHub / GitLab / Gitee / CODING / Bitbucket merge conventions and asserts the scanner reads them correctly.

License

MIT

About

Score your team's AI-assisted development maturity from git history + agent config, against published industry research. 9 dimensions, PDF report, quarterly baselines.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages