Skip to content

Commit f1f4f9f

Browse files
uipreligaclaude
andauthored
docs: add MkDocs docs site, comparison page, and SEO metadata (#32)
* docs: add MkDocs docs site, comparison page, and SEO metadata Turn the docs/ tree into an indexable GitHub Pages site and sharpen the package/repo metadata so coder_eval ranks for the queries users actually search (test claude code skills, evaluate coding agents, claude vs codex benchmark) instead of only its own name. - mkdocs.yml + .github/workflows/docs.yml: MkDocs Material site, published to GitHub Pages on push to main (one-time: set Pages source to "GitHub Actions"). - docs/index.md: keyword-rich landing page (SEO homepage). - docs/COMPARISON.md: benefit-led "why coder_eval" page comparing it to SWE-bench, SkillsBench, Harbor, and OpenAI Evals (plus hand-rolled scripts). Framing is promotional; factual claims about other projects are grounded in their own docs, with a Sources section and inline links. - docs/llms.txt: served at site root for AI-search (Claude/ChatGPT/Perplexity). - README opening rewritten: what it is, what it competes with, when to choose it, in the first 200 words; brand-split fix (coder_eval / pip install coder-eval). - pyproject.toml: +llm-eval, ai-evaluation, claude-code-skills, agent-testing, llmops keywords; Documentation URL -> docs site. - Exclude internal IDEAS.md from the site; fix repo-relative links so `mkdocs build --strict` is clean; fix Tutorial 06 mislabeled as "04". Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: address PR review — Coder Eval naming, cleaner table, gh-deploy workflow - Naming: use the display name "Coder Eval" in prose/headings across README, index, comparison, and llms.txt; keep `coder-eval` for the pip/CLI name and `coder_eval` in URLs, paths, `git clone`/`cd`, and code identifiers. site_name → "Coder Eval". (per review: bai-uipath) - Comparison table: drop the cluttered `⚠️ DIY` cells for a plain `x`; tidy the two other ⚠️ cells to plain text. (per review: bai-uipath) - Docs workflow: publish via `mkdocs gh-deploy` to the gh-pages branch (matches UiPath/uipath-python), instead of the GitHub-Actions Pages build type the org blocks. Needs only `contents: write`. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: address review — git identity for gh-pages, single strict deploy, table glyphs, uv tool install in Tutorial 01 - docs.yml: author gh-pages commits as coder-eval <coder-eval@uipath.com>; collapse build+deploy into one `mkdocs gh-deploy --force --strict` step - comparison.md: normalize table glyphs (❌ = doesn't have it, — = n/a) - Tutorial 01 + index.md: mention `uv tool install coder-eval` - pyproject: point Documentation back at the GitHub docs tree until Pages is live Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(deps): bump pyasn1 0.6.3 -> 0.6.4 (GHSA-8ppf-4f7h-5ppj, GHSA-hm4w-wwcw-mr6r) Unblocks the OSV scan in the Quality Gate; both High-severity advisories are fixed in 0.6.4. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1 parent 73f0db2 commit f1f4f9f

22 files changed

Lines changed: 608 additions & 41 deletions

‎.github/workflows/docs.yml‎

Lines changed: 48 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,48 @@
1+
name: Docs
2+
3+
# Build the MkDocs Material site and publish it to the `gh-pages` branch via
4+
# `mkdocs gh-deploy` — the same mechanism UiPath/uipath-python uses. This workflow
5+
# only pushes a branch (needs `contents: write`); it never calls the Pages API, so
6+
# it succeeds even before Pages is switched on.
7+
#
8+
# One-time setup (org owner): the UiPath org blocks Pages *creation*, so enable
9+
# Pages for this repo once — Settings → Pages → Build and deployment →
10+
# Source: "Deploy from a branch" → branch `gh-pages` / `/ (root)`. That's the same
11+
# (legacy, branch-based) configuration uipath-python already runs on.
12+
on:
13+
push:
14+
branches: [main]
15+
paths:
16+
- "docs/**"
17+
- "mkdocs.yml"
18+
- ".github/workflows/docs.yml"
19+
workflow_dispatch:
20+
21+
permissions:
22+
contents: write
23+
24+
concurrency:
25+
group: docs-gh-pages
26+
cancel-in-progress: true
27+
28+
jobs:
29+
publish:
30+
runs-on: ubuntu-latest
31+
steps:
32+
- uses: actions/checkout@v4
33+
- uses: actions/setup-python@v5
34+
with:
35+
python-version: "3.13"
36+
- name: Install social-card system libraries
37+
run: |
38+
sudo apt-get update
39+
sudo apt-get install -y --no-install-recommends \
40+
libcairo2 libfreetype6 libjpeg-turbo8 libpng16-16 pngquant
41+
- name: Install MkDocs Material
42+
run: pip install "mkdocs-material[imaging]>=9.5,<10"
43+
- name: Configure git identity for gh-pages commits
44+
run: |
45+
git config user.name "coder-eval"
46+
git config user.email "coder-eval@uipath.com"
47+
- name: Build strictly and publish to gh-pages
48+
run: mkdocs gh-deploy --force --strict

‎.gitignore‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,7 @@
1+
# MkDocs build output
2+
site/
3+
.cache/
4+
15
# Python-generated files
26
__pycache__/
37
*.py[oc]

‎README.md‎

Lines changed: 51 additions & 28 deletions
Original file line numberDiff line numberDiff line change
@@ -1,25 +1,33 @@
1-
# coder_eval — evaluate AI coding agents & their skills
1+
# Coder Eval — evaluate & benchmark AI coding agents and Claude Code skills
22

33
[![PyPI](https://img.shields.io/pypi/v/coder-eval.svg)](https://pypi.org/project/coder-eval/)
4+
[![Docs](https://img.shields.io/badge/docs-uipath.github.io%2Fcoder__eval-1f6feb.svg)](https://uipath.github.io/coder_eval/)
45
[![License: Apache 2.0](https://img.shields.io/badge/License-Apache_2.0-blue.svg)](LICENSE)
56
[![Python 3.13+](https://img.shields.io/badge/python-3.13%2B-blue.svg)](https://www.python.org/downloads/)
67
[![CI](https://github.com/UiPath/coder_eval/actions/workflows/pr-checks.yml/badge.svg)](https://github.com/UiPath/coder_eval/actions/workflows/pr-checks.yml)
78
[![Code style: Ruff](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json)](https://github.com/astral-sh/ruff)
89

9-
A framework for evaluating AI coding agents **and their skills** — built for CLI
10+
**Coder Eval** (`pip install coder-eval` / `uv tool install coder-eval`) is an open-source framework for
11+
**evaluating and benchmarking AI coding agents and their skills** — built for CLI
1012
and skill builders — with sandboxing, reproducibility, and data-driven analysis.
11-
Not an "agentic coding" benchmark: it measures how effective your CLI and skills
12-
are when used by coding agents.
13+
It runs a real agent (**Claude Code**, **Codex**, or **Google Antigravity /
14+
Gemini**) in a sandbox against declarative YAML tasks, then scores the files and
15+
commands it actually produced. Not an "agentic coding" benchmark: it measures how
16+
effective your CLI and skills are when used by coding agents.
17+
18+
Reach for it when you want to **test whether a Claude Code skill triggers**,
19+
**A/B-test Claude Code vs. Codex vs. Gemini** (or model vs. model, prompt vs.
20+
prompt), or **gate CI on coding-agent quality**. Unlike fixed datasets (SWE-bench,
21+
SkillsBench) that rank models on a shared leaderboard, Coder Eval evaluates the
22+
tasks, skills, and workflows *you* ship — with weighted 0.0–1.0 criteria, a
23+
`skill_triggered` activation check, an A/B experiment layer, and per-tool cost
24+
telemetry. See [How it compares](https://uipath.github.io/coder_eval/comparison/).
25+
📚 **Full docs:** **[uipath.github.io/coder_eval](https://uipath.github.io/coder_eval/)**.
1326

1427
<p align="center">
15-
<img src="docs/assets/hero.gif" alt="coder_eval running the hello_date task: a sandboxed agent writes and runs a script from a YAML task, then the scored result is browsed in evalboard" width="100%">
28+
<img src="docs/assets/hero.gif" alt="Coder Eval running the hello_date task: a sandboxed agent writes and runs a script from a YAML task, then the scored result is browsed in evalboard" width="100%">
1629
</p>
1730

18-
> **The Coding Agents Gym.** A sandboxed, reproducible framework to evaluate,
19-
> benchmark, and A/B-test AI coding agents — Claude Code, Codex, and Google
20-
> Antigravity (Gemini) today, any agent via a plugin SPI — with declarative
21-
> YAML tasks and weighted scoring.
22-
2331
- **Declarative YAML tasks** with pinned dependencies and clear success criteria
2432
- **Sandboxed execution** in isolated environments with resource limits
2533
- **Weighted, continuous scoring** (0.0–1.0) with fractional credit and thresholds
@@ -30,14 +38,14 @@ are when used by coding agents.
3038

3139
## What you can do with it
3240

33-
- **Benchmark coding agents** — score an agent across a suite of tasks with weighted, pass/fail thresholds
41+
- **Benchmark coding agents** — score an agent across a suite of tasks with weighted scoring and pass/fail thresholds
3442
- **Compare models & configs** — A/B-test Claude vs. Codex vs. Gemini, model vs. model, tool-on vs. tool-off, prompt vs. prompt
3543
- **Evaluate skills** — verify an agent actually engages a target skill (`skill_triggered`) and score skill-driven suites (SkillsBench-style)
3644
- **Keep skills up to date in CI** — re-validate your skills on every change or on a schedule; catch silent regressions when models, prompts, or the skills themselves drift
3745
- **Gate CI on agent quality** — run the suite in GitHub Actions and fail the build on regressions
3846
- **Bring your own dataset** — fan one task out over many rows for larger benchmark suites
3947

40-
> **Keeping skills fresh?** Run coder_eval as a scheduled GitHub Actions job so your
48+
> **Keeping skills fresh?** Run Coder Eval as a scheduled GitHub Actions job so your
4149
> skills are continuously re-evaluated against the latest model — a skill that quietly
4250
> stops triggering surfaces as a failing criterion before your users hit it. See
4351
> **[Tutorial 02 — Running coder_eval in CI](docs/tutorials/02-ci-pipeline.md)**.
@@ -53,7 +61,9 @@ git clone https://github.com/UiPath/coder_eval.git
5361
cd coder_eval
5462

5563
uv sync --extra dev # install core + dev tools
56-
cp .env.example .env # then set ANTHROPIC_API_KEY
64+
cp .env.example .env # then set ANTHROPIC_API_KEY — or skip that: an
65+
# existing Claude Code login (`claude login`) is
66+
# picked up automatically
5767

5868
uv run coder-eval plan tasks/hello_date.yaml # validate (no tokens spent)
5969
uv run coder-eval run tasks/hello_date.yaml # run your first evaluation
@@ -67,11 +77,23 @@ The optional `[uipath]` extra (`uv sync --extra dev --extra uipath`) adds the in
6777
required). Without it the framework runs end-to-end; uipath-dependent features fail
6878
at dispatch with a clear hint.
6979

70-
> **Using coder_eval in CI or another project?** Install the published package:
71-
> `pip install coder-eval` (or `uv add coder-eval`; extras install the same way —
72-
> `pip install "coder-eval[codex,antigravity]"`). In a real CI gate, pin to a
73-
> specific released version so a harness upgrade can't silently move your results.
74-
> See [Tutorial 02 — Running coder_eval in CI](docs/tutorials/02-ci-pipeline.md) for the full setup.
80+
**Using Coder Eval in CI or another project?** Install the published package
81+
instead of cloning:
82+
83+
```bash
84+
uv tool install coder-eval # puts the `coder-eval` CLI on your PATH,
85+
# in its own isolated environment
86+
87+
uv tool install "coder-eval[codex,antigravity]" # same, with agent extras
88+
coder-eval --version # verify the install
89+
```
90+
91+
To add it as a project dependency instead: `uv add coder-eval` or
92+
`pip install coder-eval`. In a real CI gate, pin to a specific released version
93+
so a harness upgrade can't silently move your results. (The example `tasks/`
94+
live in this repo — clone it or point the CLI at your own task files.) See
95+
[Tutorial 02 — Running coder_eval in CI](docs/tutorials/02-ci-pipeline.md) for
96+
the full setup.
7597

7698
## Telemetry
7799

@@ -99,17 +121,18 @@ at dispatch with a clear hint.
99121

100122
## How it compares
101123

102-
- **vs. SWE-bench and fixed benchmarks** — SWE-bench is a fixed dataset of GitHub
103-
issues; coder_eval is a *framework* for authoring your own tasks in declarative
104-
YAML, so you evaluate the skills and workflows you care about (and can still wrap
105-
a fixed dataset via [Bring Your Own Dataset](docs/BYOD.md)).
106-
- **vs. LLM-output eval harnesses (e.g. OpenAI Evals)** — those grade a model's text;
107-
coder_eval runs a full **agent** in a **sandbox** with real tool use and multi-turn
108-
dialog, then scores the files and commands it actually produced (continuous
109-
0.0–1.0) — not just a judge over a string.
124+
- **vs. fixed benchmarks (SWE-bench, SkillsBench)** — they score a canonical dataset;
125+
Coder Eval scores *your* tasks with continuous 0.0–1.0 weighted criteria (and can
126+
still wrap a fixed dataset via [Bring Your Own Dataset](docs/BYOD.md)).
127+
- **vs. large-scale / RL harnesses (Harbor)** — Harbor targets scale and RL rollouts;
128+
Coder Eval targets weighted, skill-aware suites gated in CI.
129+
- **vs. model-output eval tools (OpenAI Evals)** — they grade model text; Coder Eval
130+
runs a full agent in a sandbox and scores the files and commands it produced.
110131
- **vs. hand-rolled scripts** — reproducible sandboxes, weighted criteria,
111132
cost/token telemetry, A/B experiments, and CI-ready pass/fail gates out of the box.
112133

134+
See the full [comparison — with sources](https://uipath.github.io/coder_eval/comparison/).
135+
113136
## Task Definition
114137

115138
A task is a YAML file: a prompt, the agent config, a sandbox, and success criteria.
@@ -159,12 +182,12 @@ extension points (new criteria, new agents).
159182

160183
## Known limits & non-goals
161184

162-
- **Not a fixed benchmark or leaderboard** — coder_eval scores *your* tasks and ships
185+
- **Not a fixed benchmark or leaderboard** — Coder Eval scores *your* tasks and ships
163186
example tasks, not a canonical scored dataset.
164187
- **Tasks execute real code** — run untrusted tasks only under the container driver
165188
(see [Docker Isolation](docs/DOCKER_ISOLATION.md)); the `tempdir` driver is not a
166189
security boundary.
167-
- **Bring your own model credentials** — Anthropic, Bedrock, or Gemini keys; coder_eval
190+
- **Bring your own model credentials** — Anthropic, Bedrock, or Gemini keys; Coder Eval
168191
does not proxy or supply model access.
169192
- **Python 3.13+ only.**
170193

‎docs/AB_EXPERIMENTS.md‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,10 @@
1+
---
2+
description: >-
3+
A/B-test AI coding agents with coder_eval's experiment layer — Claude Code vs.
4+
Codex vs. Gemini, model vs. model, skill on vs. off, prompt vs. prompt — on
5+
identical tasks.
6+
---
7+
18
# A/B Experiment Guide
29

310
How to run the same tasks across multiple configuration variants ("arms") and

‎docs/BYOD.md‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,10 @@
1+
---
2+
description: >-
3+
Use a custom Docker image with coder_eval — extend the base coder-eval-agent
4+
image with your own dependencies and tools, then point task configuration at
5+
it.
6+
---
7+
18
# Bring Your Own Docker (BYOD)
29

310
The BYOD feature allows customers to use custom Docker images that extend the base `coder-eval-agent` image, enabling them to add custom dependencies and tools while maintaining the latest coder-eval codebase.

‎docs/CODEX_AGENT_GUIDE.md‎

Lines changed: 9 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,15 @@
1-
# Codex Agent Implementation
1+
---
2+
description: >-
3+
Run OpenAI Codex as the agent under evaluation in coder_eval — installation,
4+
authentication, task configuration, and how Codex telemetry maps to sandboxed,
5+
weighted scoring.
6+
---
7+
8+
# Running OpenAI Codex in coder_eval
29

310
## Overview
411

5-
A new `CodexAgent` has been added to coder_eval that integrates OpenAI's Codex SDK. The implementation mirrors the structure of `ClaudeCodeAgent` and provides seamless integration with the evaluation framework.
12+
coder_eval can run OpenAI's Codex as the agent under evaluation, via the official Codex SDK. The `CodexAgent` mirrors the structure of `ClaudeCodeAgent` and plugs into the same sandbox, scoring, and telemetry pipeline — set `agent.type: codex` in a task and the rest of the framework works unchanged.
613

714
## Setup
815

‎docs/DOCKER_ISOLATION.md‎

Lines changed: 7 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,9 +1,14 @@
1+
---
2+
description: >-
3+
Run each coder_eval task in its own fresh Docker container — strong host
4+
isolation, a pinned reproducible agent runtime, and custom images for
5+
task-specific dependencies.
6+
---
7+
18
# Docker Isolation
29

310
Run each evaluation task inside its own fresh container. Strong host isolation and a pinned, reproducible agent runtime.
411

5-
> Supersedes the agent-side FS perimeter flag from #199 (reverted in 9fe4320). The container boundary subsumes what that flag tried to do at the agent level.
6-
712
## When to use
813

914
Set `sandbox.driver: docker` on a task (or pass `--driver docker` on the CLI —

‎docs/TASK_DEFINITION_GUIDE.md‎

Lines changed: 11 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,10 @@
1+
---
2+
description: >-
3+
Full schema reference for coder_eval task YAML — agent config, sandboxes, run
4+
limits, dataset fan-out, and all 14 success criterion types with weighted
5+
0.0–1.0 scoring.
6+
---
7+
18
# Task Definition Guide
29

310
Complete reference for defining evaluation tasks in coder_eval.
@@ -128,6 +135,8 @@ an error.
128135
- `codex` — OpenAI Codex agent (requires `[codex]` extra; set `CODEX_API_KEY` and optional `CODEX_BASE_URL` environment variables).
129136
- `none` — No-op agent: no coding agent runs and no model API call is made. See [No-op / System Tasks](#no-op--system-tasks-type-none) below.
130137

138+
<a id="no-op--system-tasks-type-none"></a>
139+
131140
### No-op / System Tasks (`type: none`)
132141

133142
Set `agent: {type: none}` to run a task with **no coding agent** — "coder-eval
@@ -159,7 +168,7 @@ Contract (enforced at load): a `type: none` task must declare no `initial_prompt
159168
every criterion must be agent-independent — criteria that inspect the agent
160169
trajectory (`command_executed`, `skill_triggered`, `reference_comparison`,
161170
`commands_efficiency`) are rejected. A worked example lives at
162-
[`tasks/agentless_smoke_test.yaml`](../tasks/agentless_smoke_test.yaml).
171+
[`tasks/agentless_smoke_test.yaml`](https://github.com/UiPath/coder_eval/blob/main/tasks/agentless_smoke_test.yaml).
163172

164173
### `max_turns`, `task_timeout`, `turn_timeout` location
165174

@@ -590,7 +599,7 @@ Checks whether the agent executed specific tools/commands during evaluation. Ins
590599

591600
Evaluates a UiPath agent against a named evaluation set. **Fractional scoring:** metrics passed / total metrics.
592601

593-
> The `uipath` CLI must be available **inside the sandbox** (typically declared in the task's own Python deps). This is independent of the host's optional `coder-eval[uipath]` extra — see the install matrix in [README.md](../README.md#installation).
602+
> The `uipath` CLI must be available **inside the sandbox** (typically declared in the task's own Python deps). This is independent of the host's optional `coder-eval[uipath]` extra — see the install matrix in [README.md](https://github.com/UiPath/coder_eval/blob/main/README.md#quick-start).
594603

595604
```yaml
596605
- type: "uipath_eval"

‎docs/USER_GUIDE.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,10 @@
1+
---
2+
description: >-
3+
Complete coder_eval reference — every CLI command and flag, configuration
4+
layers and -D overrides, environment variables, run outputs, and reports for
5+
evaluating AI coding agents.
6+
---
7+
18
# coder_eval User Guide
29

310
The full command, configuration, and output reference. For a gentle introduction
@@ -81,6 +88,8 @@ in `.claude/commands/`, available when using Claude Code in this repository:
8188
| `/coder-eval-run-analysis <path>` | Analyze evaluation runs and suggest improvements to tasks, config, and prompts. Works at task, variant, or run scope. |
8289
| `/coder-eval-task-create` | Create evaluation task YAML files from a natural language description. |
8390

91+
<a id="api-routing--benchmarking"></a>
92+
8493
## API Routing & Benchmarking
8594

8695
`coder-eval` supports two API routing modes, selected via `--backend` or the

0 commit comments

Comments
 (0)