Skip to content

Commit fdb3bc1

Browse files
semantic-releasegithub-actions[bot]
authored andcommitted
chore(release): 0.12.5 [skip ci]
1 parent f3db499 commit fdb3bc1

6 files changed

Lines changed: 42 additions & 5 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 37 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,43 @@
22

33
<!-- version list -->
44

5+
## v0.12.5 (2026-09-22)
6+
7+
### Bug Fixes
8+
9+
- **agents**: Address PR #191 code review findings on the Delegate agent
10+
([#191](https://github.com/UiPath/coder_eval/pull/191),
11+
[`b5d8ba8`](https://github.com/UiPath/coder_eval/commit/b5d8ba8b9cecbb7847628d7d2ef88612a10aed73))
12+
13+
- **agents**: Delegate token usage was silently zero on every live turn
14+
([#191](https://github.com/UiPath/coder_eval/pull/191),
15+
[`b5d8ba8`](https://github.com/UiPath/coder_eval/commit/b5d8ba8b9cecbb7847628d7d2ef88612a10aed73))
16+
17+
- **criteria**: Close scoring-integrity and credential-disclosure gaps in system_one_judge
18+
([#192](https://github.com/UiPath/coder_eval/pull/192),
19+
[`f3db499`](https://github.com/UiPath/coder_eval/commit/f3db4997e042d969494e6eaaf7bcb7da1c039b26))
20+
21+
### Code Style
22+
23+
- **criteria**: Use deferred annotations in system_one_judge so the TYPE_CHECKING import reads as
24+
used ([#192](https://github.com/UiPath/coder_eval/pull/192),
25+
[`f3db499`](https://github.com/UiPath/coder_eval/commit/f3db4997e042d969494e6eaaf7bcb7da1c039b26))
26+
27+
### Features
28+
29+
- **agents**: Port the Delegate agent as a built-in
30+
([#191](https://github.com/UiPath/coder_eval/pull/191),
31+
[`b5d8ba8`](https://github.com/UiPath/coder_eval/commit/b5d8ba8b9cecbb7847628d7d2ef88612a10aed73))
32+
33+
- **agents**: Port the Delegate agent from coder_eval_uipath as a built-in
34+
([#191](https://github.com/UiPath/coder_eval/pull/191),
35+
[`b5d8ba8`](https://github.com/UiPath/coder_eval/commit/b5d8ba8b9cecbb7847628d7d2ef88612a10aed73))
36+
37+
- **criteria**: Add system_one_judge, a typed-rubric grader on a System One model
38+
([#192](https://github.com/UiPath/coder_eval/pull/192),
39+
[`f3db499`](https://github.com/UiPath/coder_eval/commit/f3db4997e042d969494e6eaaf7bcb7da1c039b26))
40+
41+
542
## v0.12.4 (2026-09-18)
643

744
### Bug Fixes

‎action.yml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -25,7 +25,7 @@ inputs:
2525
version:
2626
description: coder-eval version to install from PyPI, or "local" to install from the action checkout
2727
required: false
28-
default: "0.12.4" # <-- kept in sync with releases by release.yml
28+
default: "0.12.5" # <-- kept in sync with releases by release.yml
2929
extras:
3030
description: >-
3131
Comma-separated coder-eval extras (`codex`, `antigravity,litellm`), composed

‎plugins/coder-eval/.claude-plugin/plugin.json‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -2,7 +2,7 @@
22
"$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
33
"name": "coder-eval",
44
"displayName": "Coder Eval",
5-
"version": "0.12.4",
5+
"version": "0.12.5",
66
"description": "Test whether your Claude Code skills actually trigger, and benchmark any coding agent — author, run, and analyze the eval suite that proves it, locally or as a CI gate.",
77
"author": { "name": "UiPath", "email": "coder-eval@uipath.com", "url": "https://github.com/UiPath/coder_eval" },
88
"homepage": "https://coder-eval.com",

‎pyproject.toml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
[project]
22
name = "coder-eval"
3-
version = "0.12.4"
3+
version = "0.12.5"
44
description = "Evaluate, benchmark, and A/B-test AI coding agents (Claude Code, Codex, Gemini/Antigravity, OpenCode, Pi, Delegate) with sandboxed, reproducible YAML task suites."
55
readme = "README.md"
66
license = "Apache-2.0"

‎src/coder_eval/__init__.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,3 @@
11
"""coder_eval - A framework for evaluating AI coding agents."""
22

3-
__version__ = "0.12.4"
3+
__version__ = "0.12.5"

‎uv.lock‎

Lines changed: 1 addition & 1 deletion
Some generated files are not rendered by default. Learn more about customizing how changed files appear on GitHub.

0 commit comments

Comments
 (0)