I’m an AI solutions engineer and research scientist building systems that make autonomous behavior observable, replayable, and independently verifiable.
My work sits where agent infrastructure meets experimental science: record what happened, challenge the evaluator, preserve the evidence, and make every important claim reproducible by someone else.
My upstream work is discovered automatically across public repositories outside my account and included only when GitHub also lists me as a contributor. Merged PRs are the accepted-work measure; GitHub-indexed commits are a separate cached attribution signal. Stars and forks describe repository reach, not personal credit.
- MoneyPrinterTurbo: GitHub-listed contributor with 60 GitHub-indexed commits and 60 merged PRs; +4,782 / −452 accepted lines across 138 changed files. Repository reach: 126,744 stars and 19,791 forks.
- Ruflo: GitHub-listed contributor with 89 GitHub-indexed commits and 51 merged PRs; +6,041 / −801 accepted lines across 195 changed files. Repository reach: 73,464 stars and 8,724 forks.
- Agentic-QE: GitHub-listed contributor with 84 GitHub-indexed commits and 42 merged PRs; +11,868 / −2,105 accepted lines across 284 changed files. Repository reach: 486 stars and 96 forks.
- CowAgent: GitHub-listed contributor with 11 GitHub-indexed commits and 11 merged PRs; +1,430 / −140 accepted lines across 25 changed files. Repository reach: 47,158 stars and 10,375 forks.
Last verified 2026-09-29 UTC · visual + evidence refreshed every 30 minutes by GitHub Actions · machine-readable evidence
🛡️ Gradia GuardProof-bound evidence for AI execution. Records hash-chained decisions and actions, verifies them offline, and makes the boundary of every claim explicit.
|
An evaluation immune system. Finds reward-hacking exploits, isolates their causal slice, and tests whether scorer patches truly converge.
|
|
An evidence-graded RL environment for multi-touch enterprise negotiation: 3,510 graded episodes, 13 models, and a fully reproducible methodology.
|
Regression memory for coding agents. Seal behavioral claims once; check them over MCP or CI before a silent regression reaches history.
|
Agentic Power asks a practical question: how much accepted skilled work can one hour of human direction command?
I design for that multiplier, but I won’t publish one without exposing its scope, evidence, assumptions, and uncertainty. My operating model is to delegate meaningful responsibility, compose the right tools and context, verify the result, then use the evidence to improve the next run.
AP ≈ 61.3× (operator-calibrated provisional estimate): approximately 8,718.6 skilled Human-Equivalent Hours, or 218.0 engineer-weeks, divided by 142.1 operator-estimated human-direction hours. The transparent scenario range is 38.2× to 99.4×.
The evidence base since January 2025 combines 164 merged upstream PRs with 6,532 authored, non-merge default-branch commits across 109 currently accessible repositories: personal, private, employer, and open-source alike. Work such as Gradia is included only as a redacted aggregate: no repository names, commit messages, code, links, employer, or client details are published.
The conservative GitHub activity proxy produces 1,378.5 direction hours because it assigns attention to individual commits and repository-months. Operator recall is 1.5–2.0 active direction hours per week; across 81.2 weeks, that calibrates the denominator to 121.8–162.4 hours, with 142.1 as the midpoint.
Direction includes active briefing, steering, reviewing, correcting, and coordinating. It excludes agent runtime and waiting. The calibration is operator-estimated rather than reconstructed from time logs, so this remains a transparent scenario, not a completed Full Evidence Audit.
Framework and formula · calculation evidence · public upstream evidence and visuals refreshed every 30 minutes; redacted repository snapshot retained until a private read credential is available
- Direct: goals, constraints, acceptance criteria, and escalation rules stay explicit.
- Orchestrate: models, specialist agents, reusable skills, tools, memory, and loops become one working system.
- Verify: Gradia Guard, ProofSeal, and replayable evaluations separate completed work from plausible-looking activity.
- Improve: Wind Tunnel and Reward Loop turn measured failure into a better next run.
Framework concept by Dr. Mark Allen / HeroForge.AI. Visual and operating-model adaptation are original to this profile.
Can we trust the record? → proof-bound execution evidence
Can we trust the score? → adversarial evaluator stress testing
Can we reproduce the result? → deterministic replay + committed artifacts
Can an agent remember? → regression memory before commit
The thread through all of it is simple: an AI system should be able to show its work without asking you to trust the AI system.
- Gradia Universes — proof-carrying, interruption-capable synthetic agent worlds with deterministic replay.
- Gradia Reward Loop — oracle-witnessed reward-hacking experiments and a replay-verified paired-GRPO diagnostic.
I build communities as deliberately as systems: Chapter Lead for AI Tinkerers Charlotte and Regional Ambassador for the Agentics Foundation, connecting builders who are actively shipping agentic AI.
If you care about agent reliability, evaluation integrity, or evidence-first AI, explore the work and compare notes.
Explore Gradia → Research record → Follow the work →
Based in Charlotte · working in public



