Skip to content

Latest commit

 

History

46 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

mped-skills

Claude Code skills and subagents for building software. There is a project lifecycle here, running from a fuzzy idea through to steered autonomous execution, plus the agents that argue about the design and check the result.

Each skill is a procedure with its failure modes written down: the trap that cost a day, the check that has to run, the claim you are not allowed to make. Several ship executable scripts, so the logic lives in a file that can be linted and run. Most of them are built around refusing to let an agent certify its own work. Gates fail loudly, audits report what they did not check, and the stop conditions prefer an honest "blocked" over a confident wrong answer.

The sections below cover the main entries. The repo ships more, so browse skills/ and agents/; every one carries its own frontmatter description.

Getting started

Read before you copy. These skills encode opinions about how work should be done, so adopting one you have not read means inheriting a workflow you did not choose. Open the SKILL.md and start with the description in its frontmatter. Where a skill ends with an anti-patterns section, about half of them do, read that next; it is the fastest way to see what the skill refuses to do and where it will push back on you.

These are heavily biased towards Nix. Many of the skills reach for nix-shell, a flake.nix or a shell.nix to get their tooling, and some of the agents look for one before deciding how to run a build. Installing Nix is highly advisable; without it several of these will tell you a command is missing and stop. To see which ones apply to you:

grep -rl 'nix-shell\|flake.nix\|shell.nix' skills/ agents/

The lifecycle skills also assume a backlog CLI for durable task state. They were written against Backlog.md, which is where the conventions come from: task files on disk, stable task IDs, labels, and --plain output so an agent can read a board without opening a TUI. Phases 2 and 3 refer to it only as "the tracker CLI", so another tool offering the same operations would fit, but nothing has been tried against one.

A skill you know well is worth several you have merely installed. Most of these are prescriptive enough that using one effectively means knowing roughly what it is going to do next.


Agents

mped-architect

Greenfield work, and it has opinions. Fundamentals first, reproducible infrastructure (Git, Nix, declarative over ad-hoc), libraries over frameworks ("frameworks are captivity"), fail-fast error handling, minimal state, and a standing preference for Git over a database when Git will do.

It cares little about syntax and a lot about scalability, maintenance and patterns. Do not bring it a formatting argument.

Reach for it when the shape of a thing is still undecided and you want a position argued rather than a menu of options. It is also the reviewer of record in several of the loops here, which is deliberate: the same opinions that pick a design are the ones that notice when an implementation has quietly drifted off it.

swe-gardener

Brownfield work, and it has no house style of its own. It adapts to whatever codebase it is handed, which makes it the counterpart to mped-architect once the architecture is already decided and the job is to work inside it without leaving a mess.

Its method is dual-altitude, oscillating deliberately between architecture (data flow, ownership, abstractions) and code-near detail, with an explicit trigger for switching. Friction down in the syntax, such as proliferating conditionals or an adapter that exists only to reshape data, is treated as evidence of a problem one level up, so it goes and looks.

qa-test-runner

Runs the gate and reports actual numbers. It finds the project's own tooling (justfile recipes, shell.nix or flake.nix), executes lint, build and test, and returns errors and warnings with counts. It does not guess at commands.

It exists so that no other agent grades its own homework, and so that the main implementer or orchestrator context window stays insulated from spammy test output. Build logs are long and mostly noise; the orchestrator should get the verdict and the counts, not the scrollback. It is read-only by design, which also keeps it available in situations where an editing agent balks.

compass-project-direction-consultant

Advises on what to build next and in what order. It is read-only and advisory, and it never selects.

It sweeps a backlog by headline and classifies each task on a growth ladder: verification harness, then the hardest cornerstones, then integration, then regression insurance, then refactoring, then breadth. It digs into the top candidates and returns a ranked recommendation with its reasoning exposed for you to disagree with.

The failure it exists to catch is cornerstone-avoidance, where easy hardening tasks get cleared while an unblocked, load-bearing, probably-wrong core task waits because it is hard. It also carries a domain-coherence lens: where the design's concepts or boundaries are muddy, that ambiguity is a sequencing risk rather than a naming nitpick.


Skills

The lifecycle trilogy

phase1-prd-grill → phase2-backlog-snowball → phase3-backlog-ralph take a project from a fuzzy idea to steered autonomous execution. Phases 1 and 2 are greenfield. Phase 3 works on either.

They hand off through named artifacts, so each runs alone if you already have the previous stage's output.

phase1-prd-grill

For articulating a project vision. Greenfield. It turns a fuzzy idea into a PRD by interviewing you until you are satisfied, not until the model is, asking pointed questions and surfacing weak assumptions along the way. It is explicitly not a yes-man.

Its distinctive deliverable is the irreversibility map: which surfaces cannot be cheaply walked back once shipped (on-disk formats, data-integrity write paths, migrations, published APIs) and which are still tentative. Phase 2 turns that map into task labels and phase 3 turns it into review gates, so a missing map leaves the later phases deep-reviewing everything or nothing.

One round of phase1-prd-grill: elicit, draft or refine the PRD, spawn a read-only adversarial reviewer, grill the human, commit, and repeat until the human says it is good enough.

phase2-backlog-snowball

Ramps a backlog from the vision. Greenfield.

Design-for-test comes first. You define what good and bad look like before planning the work, and only then is the backlog populated: deeply if the project is firm, or one wave at a time if it is experimental, with a re-plan task terminating each wave.

It is functionality-first by doctrine. Feature tasks get lean acceptance criteria, irreversible labels are applied only where deep review is earned, there is a UX/e2e journey task every five or so, and exhaustive hardening is consolidated at the wave's end against surfaces that have stopped moving. Hardening code that is still tentative is planned waste, since the wave may rewrite it.

phase2-backlog-snowball: design the test grounding first, branch on firm versus experimental, populate the backlog through the tracker CLI, intersperse journeys and consolidate hardening, then pass a parallel read-only review gate.

phase3-backlog-ralph

Executes a backlog Ralph-style, but with steering rather than a dumb while true loop. Greenfield or brownfield alike; by this point it is just a backlog and a codebase.

Selection walks the dependency graph over a priority ladder, so the next task is the highest-leverage unblocked one, not simply the next ID. Each task gets one heavily-onboarded implementer subagent. The review gate is tiered: lightweight for ordinary feature cycles, deep multi-reviewer only for irreversible surfaces and scheduled checkpoints. Deep-reviewing everything is the slow loop this replaces, and reviewing nothing is the naive one.

Underneath sits honest-failure discipline. "Done" requires every acceptance criterion checked and the gate green, and a correct "blocked, here is the exact cause, here are the filed prerequisites" counts as a successful cycle.

One cycle of phase3-backlog-ralph: the orchestrator briefs one implementer, folds its report back, runs a light or deep gate, routes gate-breaking findings back for fix-up and the rest to the hardening wave, records, and loops.

readme-improver

Audits a README against a rubric distilled from the canonical guides: the four questions a scanner needs answered in thirty seconds, cognitive funneling, overload delegated to docs/, stale content moved to cruft/ rather than deleted. Every finding cites where, why it matters, and one concrete fix. If it cannot propose the fix, it is noise and gets dropped.

Its distinctive mode is --drift. A project is rarely one document, and README, sub-READMEs, CLAUDE.md, docs/ and spec notes rot at different rates. It fans out one read-only agent per document, builds an axis × document matrix, then verifies against the code, so a finding can name the stale document instead of reporting that two docs differ and leaving you to pick.

This README was written by it, including the part where it caught three errors in its own first draft.

refactor-smeller

An adversarial audit for the refactoring debt that survives review and compounds later. It starts from the assumption that the code reads fine on a first pass and contains drift anyway. Eight categories: oversized files, fragile citations, silent-sibling defects, test discipline, defensive-construct rot, doc-versus-code drift, split-induced second-order drift, and enumerated-disclosure.

The split-induced category is the one that pays. When a file is split, the person doing it sees only the file being moved. Inline comments get promoted to module-level docs where their staleness becomes more visible, and every citation elsewhere in the tree pointing at the old path goes stale silently. Neither is visible during the split itself.

Two rules carry more weight than the category list. Every finding must propose one concrete remediation, and if it cannot, it is noise and gets dropped. A finding that rests on a claim about the code must be verified by grep before it is emitted, and dropped when the grep disagrees with it.

It reports and does not fix. An automatic fix would compound the silent-sibling defect it hunts for, because now the fix is silent too.

explainer-video

Produces a narrated, animated 1080p explainer video from a topic: manuscript, local TTS speech, seekable-HTML graphics, then a combined mp4. No cloud services and no API keys.

Reproducible on x86_64-linux via a Nix flake, with virtual-clock frame capture so a given scene renders identically every time. What usually makes generated video un-reviewable is that you cannot diff two runs.

android-app-reverser

A phased, backlog-driven pipeline for turning an Android APK into a production-quality Rust CLI and API library: extraction, decompilation, architecture recovery, auth flow, then the client implementation.

The discipline is what keeps a reverse-engineering project from sprawling. All artifacts land under re/, every tangent becomes a backlog task, and one delegated agent handles one task at a time. A PRD and an agent-instructions file act as the persistence layer, so a fresh agent picks up cold without re-deriving what was already learned.

xterm-tmux-control

Drives graphical terminals through tmux for testing TUI applications: start an xterm, run commands in it, capture screen state, take rapid snapshots of a startup sequence.

It also documents one trap worth the price of admission. tmux send-keys is variadic, so when a multi-line command gets collapsed onto one line, the following commands become extra key arguments and are typed into the remote session as concatenated garbage. Chaining with ; fixes it. The failure is silent: send-keys exits 0, and the damage only shows up in the captured pane, where it reads as the application misbehaving.


Contributing

Personal tooling, developed in the open. Issues and pull requests welcome, but expect opinions attached.

License

MIT — see LICENSE.

About

Public subset of my Claude Code skills and agents

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages