You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
docs: add MkDocs docs site, comparison page, and SEO metadata (#32)
* docs: add MkDocs docs site, comparison page, and SEO metadata
Turn the docs/ tree into an indexable GitHub Pages site and sharpen the
package/repo metadata so coder_eval ranks for the queries users actually
search (test claude code skills, evaluate coding agents, claude vs codex
benchmark) instead of only its own name.
- mkdocs.yml + .github/workflows/docs.yml: MkDocs Material site, published to
GitHub Pages on push to main (one-time: set Pages source to "GitHub Actions").
- docs/index.md: keyword-rich landing page (SEO homepage).
- docs/COMPARISON.md: benefit-led "why coder_eval" page comparing it to
SWE-bench, SkillsBench, Harbor, and OpenAI Evals (plus hand-rolled scripts).
Framing is promotional; factual claims about other projects are grounded in
their own docs, with a Sources section and inline links.
- docs/llms.txt: served at site root for AI-search (Claude/ChatGPT/Perplexity).
- README opening rewritten: what it is, what it competes with, when to choose
it, in the first 200 words; brand-split fix (coder_eval / pip install coder-eval).
- pyproject.toml: +llm-eval, ai-evaluation, claude-code-skills, agent-testing,
llmops keywords; Documentation URL -> docs site.
- Exclude internal IDEAS.md from the site; fix repo-relative links so
`mkdocs build --strict` is clean; fix Tutorial 06 mislabeled as "04".
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: address PR review — Coder Eval naming, cleaner table, gh-deploy workflow
- Naming: use the display name "Coder Eval" in prose/headings across README,
index, comparison, and llms.txt; keep `coder-eval` for the pip/CLI name and
`coder_eval` in URLs, paths, `git clone`/`cd`, and code identifiers. site_name
→ "Coder Eval". (per review: bai-uipath)
- Comparison table: drop the cluttered `⚠️ DIY` cells for a plain `x`; tidy the
two other ⚠️ cells to plain text. (per review: bai-uipath)
- Docs workflow: publish via `mkdocs gh-deploy` to the gh-pages branch (matches
UiPath/uipath-python), instead of the GitHub-Actions Pages build type the org
blocks. Needs only `contents: write`.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: address review — git identity for gh-pages, single strict deploy, table glyphs, uv tool install in Tutorial 01
- docs.yml: author gh-pages commits as coder-eval <coder-eval@uipath.com>;
collapse build+deploy into one `mkdocs gh-deploy --force --strict` step
- comparison.md: normalize table glyphs (❌ = doesn't have it, — = n/a)
- Tutorial 01 + index.md: mention `uv tool install coder-eval`
- pyproject: point Documentation back at the GitHub docs tree until Pages is live
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(deps): bump pyasn1 0.6.3 -> 0.6.4 (GHSA-8ppf-4f7h-5ppj, GHSA-hm4w-wwcw-mr6r)
Unblocks the OSV scan in the Quality Gate; both High-severity advisories
are fixed in 0.6.4.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
<imgsrc="docs/assets/hero.gif"alt="coder_eval running the hello_date task: a sandboxed agent writes and runs a script from a YAML task, then the scored result is browsed in evalboard"width="100%">
28
+
<imgsrc="docs/assets/hero.gif"alt="Coder Eval running the hello_date task: a sandboxed agent writes and runs a script from a YAML task, then the scored result is browsed in evalboard"width="100%">
16
29
</p>
17
30
18
-
> **The Coding Agents Gym.** A sandboxed, reproducible framework to evaluate,
19
-
> benchmark, and A/B-test AI coding agents — Claude Code, Codex, and Google
20
-
> Antigravity (Gemini) today, any agent via a plugin SPI — with declarative
21
-
> YAML tasks and weighted scoring.
22
-
23
31
-**Declarative YAML tasks** with pinned dependencies and clear success criteria
24
32
-**Sandboxed execution** in isolated environments with resource limits
25
33
-**Weighted, continuous scoring** (0.0–1.0) with fractional credit and thresholds
@@ -30,14 +38,14 @@ are when used by coding agents.
30
38
31
39
## What you can do with it
32
40
33
-
-**Benchmark coding agents** — score an agent across a suite of tasks with weighted, pass/fail thresholds
41
+
-**Benchmark coding agents** — score an agent across a suite of tasks with weighted scoring and pass/fail thresholds
34
42
-**Compare models & configs** — A/B-test Claude vs. Codex vs. Gemini, model vs. model, tool-on vs. tool-off, prompt vs. prompt
35
43
-**Evaluate skills** — verify an agent actually engages a target skill (`skill_triggered`) and score skill-driven suites (SkillsBench-style)
36
44
-**Keep skills up to date in CI** — re-validate your skills on every change or on a schedule; catch silent regressions when models, prompts, or the skills themselves drift
37
45
-**Gate CI on agent quality** — run the suite in GitHub Actions and fail the build on regressions
38
46
-**Bring your own dataset** — fan one task out over many rows for larger benchmark suites
39
47
40
-
> **Keeping skills fresh?** Run coder_eval as a scheduled GitHub Actions job so your
48
+
> **Keeping skills fresh?** Run Coder Eval as a scheduled GitHub Actions job so your
41
49
> skills are continuously re-evaluated against the latest model — a skill that quietly
42
50
> stops triggering surfaces as a failing criterion before your users hit it. See
43
51
> **[Tutorial 02 — Running coder_eval in CI](docs/tutorials/02-ci-pipeline.md)**.
Copy file name to clipboardExpand all lines: docs/BYOD.md
+7Lines changed: 7 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,3 +1,10 @@
1
+
---
2
+
description: >-
3
+
Use a custom Docker image with coder_eval — extend the base coder-eval-agent
4
+
image with your own dependencies and tools, then point task configuration at
5
+
it.
6
+
---
7
+
1
8
# Bring Your Own Docker (BYOD)
2
9
3
10
The BYOD feature allows customers to use custom Docker images that extend the base `coder-eval-agent` image, enabling them to add custom dependencies and tools while maintaining the latest coder-eval codebase.
Copy file name to clipboardExpand all lines: docs/CODEX_AGENT_GUIDE.md
+9-2Lines changed: 9 additions & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,8 +1,15 @@
1
-
# Codex Agent Implementation
1
+
---
2
+
description: >-
3
+
Run OpenAI Codex as the agent under evaluation in coder_eval — installation,
4
+
authentication, task configuration, and how Codex telemetry maps to sandboxed,
5
+
weighted scoring.
6
+
---
7
+
8
+
# Running OpenAI Codex in coder_eval
2
9
3
10
## Overview
4
11
5
-
A new `CodexAgent` has been added to coder_eval that integrates OpenAI's Codex SDK. The implementation mirrors the structure of `ClaudeCodeAgent` and provides seamless integration with the evaluation framework.
12
+
coder_eval can run OpenAI's Codex as the agent under evaluation, via the official Codex SDK. The `CodexAgent` mirrors the structure of `ClaudeCodeAgent` and plugs into the same sandbox, scoring, and telemetry pipeline — set `agent.type: codex` in a task and the rest of the framework works unchanged.
Copy file name to clipboardExpand all lines: docs/DOCKER_ISOLATION.md
+7-2Lines changed: 7 additions & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,9 +1,14 @@
1
+
---
2
+
description: >-
3
+
Run each coder_eval task in its own fresh Docker container — strong host
4
+
isolation, a pinned reproducible agent runtime, and custom images for
5
+
task-specific dependencies.
6
+
---
7
+
1
8
# Docker Isolation
2
9
3
10
Run each evaluation task inside its own fresh container. Strong host isolation and a pinned, reproducible agent runtime.
4
11
5
-
> Supersedes the agent-side FS perimeter flag from #199 (reverted in 9fe4320). The container boundary subsumes what that flag tried to do at the agent level.
6
-
7
12
## When to use
8
13
9
14
Set `sandbox.driver: docker` on a task (or pass `--driver docker` on the CLI —
@@ -590,7 +599,7 @@ Checks whether the agent executed specific tools/commands during evaluation. Ins
590
599
591
600
Evaluates a UiPath agent against a named evaluation set. **Fractional scoring:** metrics passed / total metrics.
592
601
593
-
> The `uipath` CLI must be available **inside the sandbox** (typically declared in the task's own Python deps). This is independent of the host's optional `coder-eval[uipath]` extra — see the install matrix in [README.md](../README.md#installation).
602
+
> The `uipath` CLI must be available **inside the sandbox** (typically declared in the task's own Python deps). This is independent of the host's optional `coder-eval[uipath]` extra — see the install matrix in [README.md](https://github.com/UiPath/coder_eval/blob/main/README.md#quick-start).
Copy file name to clipboardExpand all lines: docs/USER_GUIDE.md
+9Lines changed: 9 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,3 +1,10 @@
1
+
---
2
+
description: >-
3
+
Complete coder_eval reference — every CLI command and flag, configuration
4
+
layers and -D overrides, environment variables, run outputs, and reports for
5
+
evaluating AI coding agents.
6
+
---
7
+
1
8
# coder_eval User Guide
2
9
3
10
The full command, configuration, and output reference. For a gentle introduction
@@ -81,6 +88,8 @@ in `.claude/commands/`, available when using Claude Code in this repository:
81
88
|`/coder-eval-run-analysis <path>`| Analyze evaluation runs and suggest improvements to tasks, config, and prompts. Works at task, variant, or run scope. |
82
89
|`/coder-eval-task-create`| Create evaluation task YAML files from a natural language description. |
83
90
91
+
<aid="api-routing--benchmarking"></a>
92
+
84
93
## API Routing & Benchmarking
85
94
86
95
`coder-eval` supports two API routing modes, selected via `--backend` or the
0 commit comments