Skip to content

Commit a28bbad

Browse files
uipreligaclaude
andcommitted
docs(stub): make the Pages stub and package metadata agent-agnostic
Bring the GitHub Pages stub in line with the reworked README: drop "Claude Code skills" from the title, name all four harnesses (OpenCode was missing entirely), and lead with the "Playwright for coding agents" tagline in both the meta description and the visible lead paragraph. The old lead also read "agents and their Claude Code skills", which parsed wrong even for a Claude-only framing. The stub's own comment says its description mirrors the package metadata, so pyproject's description and keywords gain OpenCode too — otherwise that comment stops being true the moment the stub changes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016v9L8gXbq4MzH6V9JZPqcb
1 parent 18858cb commit a28bbad

2 files changed

Lines changed: 8 additions & 7 deletions

File tree

‎.github/pages-stub/index.html‎

Lines changed: 6 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -3,7 +3,7 @@
33
<head>
44
<meta charset="utf-8" />
55
<meta name="viewport" content="width=device-width, initial-scale=1" />
6-
<title>Coder Eval — evaluate AI coding agents and Claude Code skills</title>
6+
<title>Coder Eval — evaluate AI coding agents and their skills</title>
77
<!--
88
The description is here for link previews and for anyone reading source,
99
not for ranking: this page is noindex (see below), so search engines will
@@ -12,7 +12,7 @@
1212
-->
1313
<meta
1414
name="description"
15-
content="Coder Eval is an open-source framework for evaluating and benchmarking AI coding agents and Claude Code skills — sandboxed runs of Claude Code, Codex, and Gemini against declarative YAML tasks, with weighted scoring and CI gates. Documentation: coder-eval.com/docs"
15+
content="Coder Eval is Playwright for coding agents: an open-source, agent-agnostic framework for evaluating and benchmarking AI coding agents and their skills — sandboxed runs of Claude Code, Codex, Antigravity (Gemini), or OpenCode against declarative YAML tasks, with weighted scoring and CI gates. Documentation: coder-eval.com/docs"
1616
/>
1717

1818
<!--
@@ -225,9 +225,10 @@
225225
-->
226226
<h1 class="sr-only">Coder Eval</h1>
227227
<p class="lead">
228-
An open-source framework for evaluating and benchmarking AI coding agents and their Claude
229-
Code skills: it runs a real agent — Claude Code, Codex, or Gemini — in a sandbox against
230-
declarative YAML tasks, then scores the files and commands the agent actually produced.
228+
<strong>Playwright for coding agents.</strong> An open-source, agent-agnostic framework for
229+
evaluating and benchmarking AI coding agents and their skills: it runs a real agent — Claude
230+
Code, Codex, Antigravity (Gemini), or OpenCode — in a sandbox against declarative YAML
231+
tasks, then scores the files and commands the agent actually produced.
231232
</p>
232233
<p class="notice">
233234
The documentation has moved to

‎pyproject.toml‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,15 +1,15 @@
11
[project]
22
name = "coder-eval"
33
version = "0.11.6"
4-
description = "Evaluate, benchmark, and A/B-test AI coding agents (Claude Code, Codex, Gemini/Antigravity) with sandboxed, reproducible YAML task suites."
4+
description = "Evaluate, benchmark, and A/B-test AI coding agents (Claude Code, Codex, Gemini/Antigravity, OpenCode) with sandboxed, reproducible YAML task suites."
55
readme = "README.md"
66
license = "Apache-2.0"
77
requires-python = ">=3.13"
88
authors = [{ name = "UiPath", email = "coder-eval@uipath.com" }]
99
keywords = [
1010
"ai", "llm", "agent", "coding-agent", "evaluation", "eval", "evals",
1111
"benchmark", "swe-bench", "claude", "claude-code", "codex", "anthropic",
12-
"gemini", "antigravity", "sandbox", "code-generation", "agent-evaluation",
12+
"gemini", "antigravity", "opencode", "sandbox", "code-generation", "agent-evaluation",
1313
"llm-evaluation", "llm-eval", "ai-evaluation", "skills-evaluation",
1414
"claude-skills", "claude-code-skills", "agent-skills", "skillsbench",
1515
"agent-testing", "llmops",

0 commit comments

Comments
 (0)