Skip to content
View rudycelekli's full-sized avatar

Highlights

  • Pro

Block or report rudycelekli

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
rudycelekli/README.md
Rudy Celekli — building AI systems that can prove what happened

Gradia ORCID Follow

Intelligence is cheap. Evidence is the product.

I’m an AI solutions engineer and research scientist building systems that make autonomous behavior observable, replayable, and independently verifiable.

My work sits where agent infrastructure meets experimental science: record what happened, challenge the evaluator, preserve the evidence, and make every important claim reproducible by someone else.

The proof loop: observe, record, replay, verify

Open-source impact, verified

My upstream work is discovered automatically across public repositories outside my account and included only when GitHub also lists me as a contributor. Merged PRs are the accepted-work measure; GitHub-indexed commits are a separate cached attribution signal. Stars and forks describe repository reach, not personal credit.

GitHub-verified contribution statistics for MoneyPrinterTurbo, Ruflo, Agentic-QE, CowAgent

Last verified 2026-09-29 UTC · visual + evidence refreshed every 30 minutes by GitHub Actions · machine-readable evidence

The proof stack

🛡️ Gradia Guard

Proof-bound evidence for AI execution. Records hash-chained decisions and actions, verifies them offline, and makes the boundary of every claim explicit.

TypeScript Python AG-UI Cryptographic evidence

An evaluation immune system. Finds reward-hacking exploits, isolates their causal slice, and tests whether scorer patches truly converge.

Python Causal testing Offline replay DOI

An evidence-graded RL environment for multi-touch enterprise negotiation: 3,510 graded episodes, 13 models, and a fully reproducible methodology.

TypeScript Python RLVR Evaluation

Regression memory for coding agents. Seal behavioral claims once; check them over MCP or CI before a silent regression reaches history.

JavaScript MCP CI Deterministic testing

Agentic power, engineered

Agentic Power asks a practical question: how much accepted skilled work can one hour of human direction command?

I design for that multiplier, but I won’t publish one without exposing its scope, evidence, assumptions, and uncertainty. My operating model is to delegate meaningful responsibility, compose the right tools and context, verify the result, then use the evidence to improve the next run.

Live Agentic Power snapshot

Operator-calibrated provisional Agentic Power since 2025 with a month-by-month timeline

AP ≈ 61.3× (operator-calibrated provisional estimate): approximately 8,718.6 skilled Human-Equivalent Hours, or 218.0 engineer-weeks, divided by 142.1 operator-estimated human-direction hours. The transparent scenario range is 38.2× to 99.4×.

The evidence base since January 2025 combines 164 merged upstream PRs with 6,532 authored, non-merge default-branch commits across 109 currently accessible repositories: personal, private, employer, and open-source alike. Work such as Gradia is included only as a redacted aggregate: no repository names, commit messages, code, links, employer, or client details are published.

The conservative GitHub activity proxy produces 1,378.5 direction hours because it assigns attention to individual commits and repository-months. Operator recall is 1.5–2.0 active direction hours per week; across 81.2 weeks, that calibrates the denominator to 121.8–162.4 hours, with 142.1 as the midpoint.

Direction includes active briefing, steering, reviewing, correcting, and coordinating. It excludes agent runtime and waiting. The calibration is operator-estimated rather than reconstructed from time logs, so this remains a transparent scenario, not a completed Full Evidence Audit.

Framework and formula · calculation evidence · public upstream evidence and visuals refreshed every 30 minutes; redacted repository snapshot retained until a private read credential is available

Agentic operating model: delegate, orchestrate, verify, and improve

  • Direct: goals, constraints, acceptance criteria, and escalation rules stay explicit.
  • Orchestrate: models, specialist agents, reusable skills, tools, memory, and loops become one working system.
  • Verify: Gradia Guard, ProofSeal, and replayable evaluations separate completed work from plausible-looking activity.
  • Improve: Wind Tunnel and Reward Loop turn measured failure into a better next run.

Framework concept by Dr. Mark Allen / HeroForge.AI. Visual and operating-model adaptation are original to this profile.

What I’m exploring

Can we trust the record?      → proof-bound execution evidence
Can we trust the score?       → adversarial evaluator stress testing
Can we reproduce the result?  → deterministic replay + committed artifacts
Can an agent remember?        → regression memory before commit

The thread through all of it is simple: an AI system should be able to show its work without asking you to trust the AI system.

Working set

TypeScript Python Node.js MCP GitHub Actions Research

In the lab

  • Gradia Universes — proof-carrying, interruption-capable synthetic agent worlds with deterministic replay.
  • Gradia Reward Loop — oracle-witnessed reward-hacking experiments and a replay-verified paired-GRPO diagnostic.

Building the human layer

I build communities as deliberately as systems: Chapter Lead for AI Tinkerers Charlotte and Regional Ambassador for the Agentics Foundation, connecting builders who are actively shipping agentic AI.


Building AI that can survive contact with reality.

If you care about agent reliability, evaluation integrity, or evidence-first AI, explore the work and compare notes.

Explore Gradia →    Research record →    Follow the work →

Based in Charlotte · working in public

Popular repositories Loading

  1. proofseal proofseal Public

    Regression memory for coding agents — seal how your repo behaves, then check over MCP (or in CI) that an edit didn't regress it: pass/drift/regressed/missing. With seal history, bisection, and stal…

    JavaScript 2

  2. Papers-Literature-ML-DL-RL-AI Papers-Literature-ML-DL-RL-AI Public

    Forked from tirthajyoti/Papers-Literature-ML-DL-RL-AI

    Highly cited and useful papers related to machine learning, deep learning, AI, game theory, reinforcement learning

  3. data data Public

    Forked from jswansburg/data

    Data and code behind the articles and graphics at FiveThirtyEight

    Jupyter Notebook

  4. dr-streamlit-rudy dr-streamlit-rudy Public

    Forked from datarobot/dr-apps

    Python

  5. effinai effinai Public

  6. nextjs-subscription-payments nextjs-subscription-payments Public

    Forked from vercel/nextjs-subscription-payments

    Clone, deploy, and fully customize a SaaS subscription application with Next.js.

    TypeScript