Benchmarks for Kogen, a coding-agent harness. The repository records comparisons with Codex and Grok Build, model and effort settings, harness designs, implementation languages, correctness, and reported usage costs. See the public glossary for arm, model, and analysis terms.
Open source, closed contribution. Source code is available under the Apache License 2.0. This repository does not accept issues or pull requests and is not a public support channel. You may study and adapt it under the license.
- FINDINGS.md is the family-by-family summary of what the public evidence does and does not establish.
- Hypothesis register lists H01–H153 and links each proposition to its family page.
- Round register lists experiment pages and their status; each round page is the authority for that run's history and limits.
- Publication validation records the final validator output, current blockers, and bounded claims safe to publish.
- Public glossary defines arm, model, and analysis terms.
- Verification guide gives the commands and limits for recomputing published records and numbers.
- Rerun guide explains task grading and the current state of the execution kit.
- Ubuntu host setup is the one-command installer for a fresh Ubuntu 24.04 x86_64 worker. Its fresh-machine verification is pending the EU rebuild; see the rerun guide for the manual login and admission controls.
- Round status summarizes current round states.
- Credits records acknowledgements.
Start with the topic in FINDINGS.md, then follow its H IDs to the linked family page. The family page points to relevant round receipts. To look up a known round or study ID directly, use the round register or the machine-readable round index. Task definitions are in tasks/; official outcomes and captured deliveries are indexed separately in results/.
Use the hypothesis register for family ownership and evidence status, the family page for hypothesis-level synthesis, and each round page for run history. cells.jsonl contains official outcome rows; the run-record index lists one captured-delivery JSONL per round, including deliveries without an official grade, plus an unassigned.jsonl partition for records without an unambiguous round tag. Other large evidence families are also indexed per round. Schema 1.2 uses compact missing-reason codes listed in the legend. These exports answer different questions and must not be treated as interchangeable. Study design, execution, grading, and interpretation are documented in METHOD.md; required cell fields and their schema are in STANDARD.md.
From the repository root, use the standard-library-only scripts in reproduce/README.md. The primary commands are:
python3 reproduce/export_results.py
python3 reproduce/build_records.py
python3 reproduce/build_grade_join.py
python3 reproduce/validate_repo.pyThe historical audit command can rewrite round audit documents and the validation summary; read the reproduction notes before running it.
- METHOD.md: study design, execution, grading, and interpretation
- STANDARD.md: required cell records and the schema
- Vendored Kogen specification: pinned to v1.2 at commit
1118f7fc0dbb042af8c8de2ffd1b85768cf9e2a0. Upstream v1.3-draft exists, but it is not the version cited in EVIDENCE-MAP.md. - rounds/: round records, outcomes, and status labels
- decisions/: analysis rulings
- categories/: results grouped by topic
- tasks/: task records and public-prompt coverage
- results/: official outcome export and separately captured standard records
For scored studies, VALID, CONFOUNDED, INVALID, and INTERIM describe outcome validity. WITHDRAWN and NOT-RUN describe lifecycle; PRE-REGISTERED identifies a design with no scored cells. DESCRIPTIVE, EXPLORATORY, and CONFIRMATORY describe the analysis. Round pages identify whether the design was registered before execution; historical records may be retrospective. Early rounds are exploratory, some samples are small, and task mixes and venues differ.
Results are counted on task hidden suites; stack lint and typecheck results are separate diagnostics. Task grading materials ship under tasks/<id>/hidden/ and tasks/<id>/grader/ and remain outside the agent-visible workspace during a run (PRIVATE.md, rerun guide).
For record verification, requirements are Python 3 (standard library only); no model access, credentials, or benchmark hosts are needed. Rerunning tasks has additional Linux, Bubblewrap, and operator-owned Codex access requirements described in RERUN.md.
git clone https://github.com/KogenAI/kogen-bench.git
cd kogen-bench
python3 reproduce/export_results.py
python3 reproduce/build_records.py
python3 reproduce/build_grade_join.pyexport_results.py rebuilds cells.csv, cells.jsonl, unmapped.json, and export-report.json. build_records.py rebuilds the indexed per-round Standard records from committed indexed evidence. The historical audit command also rewrites round audit documents and the validation summary; see reproduction notes before using it.
The research website is built from site/ for https://bench.kogen.dev. See its README for build inputs, private preview mode, and deployment settings. The source records above remain the evidence authority.
python3 reproduce/validate_repo.pyThe validator checks round and task indexes, result exports, required record fields, privacy rules, and authored relative links. It excludes archival source excerpts under sources/ except for the source index. It does not inspect remote configuration. See VERIFY.md for the full recomputation commands and limits.
The source is licensed under the Apache License 2.0. Product identity is described in BRANDING.md.
Upstream code is unmodified; its contents are the upstream authors'. Source files identified as upstream copies remain as published, including their hostnames, deployment paths, public addresses, and test-fixture values such as DNS test IPs. This applies to 37signals LLC's Fizzy, Fizzy SaaS, and Writebook, and to Agents on Rails, commissioned by the Rails Foundation and built by Evil Martians. Their attributions and licence terms are listed in CREDITS.md, NOTICE, and LICENSES/.
Publication edits to Kogen-produced artifacts are recorded in PUBLICATION-MANIFEST.json, with separate as-run and published hashes and the reason for each change. See SECURITY-NOTES.md for the retained public commit identity and the upstream-content exception.
Forks and public deployments must use their own name, logo, visual identity, domain, content, credentials, and data. Kogen brand assets are not licensed under Apache-2.0. You may refer to Kogen only as reasonably necessary to describe the origin of the software.