Skip to content

Commit b3bba2b

Browse files
jiyangzhakshayliveclaude
authored
feat(criteria): accept glob patterns in criterion path fields (#65)
* feat(criteria): accept glob patterns in criterion path fields A criterion `path` is a literal, so a task must hardcode where an artifact lands. When the prompt does not pin that location — a scaffolding tool that creates a wrapper directory the agent names itself — a correct artifact in an unexpected directory scores 0.0 on the path alone, while a sibling criterion that discovers the file reports the run as valid. Resolve `path` through `Sandbox.resolve_files`, which expands a pattern containing `*`, `?`, or `[` against the sandbox root and leaves a literal path untouched. Every path-based criterion inherits it: file_exists, file_contains, file_matches_regex, file_check, json_check, import_check. `file_exists` passes on at least one match. Content reads require exactly one match and otherwise raise, listing every match — an ambiguous pattern is captured as a scored-0.0 result rather than silently grading one arbitrary file. Matches are sorted for determinism and directories are dropped. * fix(criteria): make glob path resolution literal-first and ignore-filtered Review follow-ups on the glob-in-`path` seam. Two of them changed what a run scores. Literal-first resolution. `Path.glob` turns `[...]` into a character class, so sniffing for `*?[` reinterpreted a plain filename as a pattern: a real `report[2024].json` resolved to a `report2.json` decoy and was graded silently, and `logs[1]` — which exists — globbed to nothing, flipping `file_exists` to 0.0 for unchanged agent output. `resolve_files` now probes the literal path first and only expands when it does not exist. This is not just hand-written YAML: dataset fan-out substitutes `${row.<field>}` into criterion paths, so a row value carrying a metacharacter injected glob semantics per-row. Ignore-pattern filtering. Expansion walked the whole sandbox root, which holds harness-created content the agent never authored — `.venv` (created inside the root for any task with a `python:` block), copied template trees, `node_modules` — and pathlib descends into dotdirs. `**/*.json` could pass off a vendored `package.json`, or hard-fail on ambiguity that has nothing to do with the agent. Glob matches now filter through the same `get_ignore_patterns` / `should_ignore_path` used for template copying, on the sandbox-relative path. A segment the pattern names literally is an explicit opt-in and survives, so `dist/**/*.js` still grades `dist`; `ignore_patterns: ["!dist"]` un-ignores a segment a wildcard discovers. Also: - `reference_comparison.agent_file` read `sandbox_dir` directly and bypassed the seam, so path semantics differed per criterion. It now routes through `get_file_content` like every other path field. - The graded file is echoed as `resolved: <path>` in criterion details — with exactly-one semantics, which file was picked is most of the signal. - The ambiguity error caps its listing at 10 matches with `+N more`; it is persisted to task.json and injected into judge prompts. - `json_check.path` / `json_schema`, `classification_match.path` and `agent_file` carried pre-glob field descriptions. - The guide listed `import_check`, which is not a criterion type. - `test_matches_are_sorted` asserted against `sorted()` of its own output. New lint rule CE032 fails a criterion checker that joins a path onto `sandbox.sandbox_dir` instead of using the seam — the mechanically detectable root cause of the `agent_file` drift. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Akshaya Shanbhogue <akshaya.shanbhogue@uipath.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent ce006c1 commit b3bba2b

14 files changed

Lines changed: 544 additions & 27 deletions

‎docs/TASK_DEFINITION_GUIDE.md‎

Lines changed: 25 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -21,6 +21,7 @@ Complete reference for defining evaluation tasks in Coder Eval.
2121
- [Template Sources](#template-sources)
2222
- [Success Criteria](#success-criteria)
2323
- [Continuous Scoring](#continuous-scoring)
24+
- [Glob patterns in path](#glob-patterns-in-path)
2425
- [file_exists](#file_exists)
2526
- [file_contains](#file_contains)
2627
- [file_check](#file_check)
@@ -670,6 +671,30 @@ score mattered.
670671

671672
**Weighted score:** `weighted_score = sum(score * weight) / sum(weight)` — calculated regardless for quality assessment.
672673

674+
### Glob patterns in `path`
675+
676+
Every sandbox-relative path field accepts a glob — `path` on `file_exists`, `file_contains`, `file_matches_regex`, `file_check`, `json_check` and `classification_match`, `json_schema` on `json_check`, and `agent_file` on `reference_comparison`. Use one when the prompt does not pin where the file lands — a scaffolding tool that creates a wrapper directory the agent names itself, for example.
677+
678+
```yaml
679+
- type: "file_contains"
680+
path: "**/*.flow" # matches any depth under the sandbox root
681+
includes: ['"core.logic.decision"']
682+
description: "flow wires a Decision node"
683+
```
684+
685+
Rules:
686+
687+
- **A path that exists is never treated as a pattern.** A literal `path` behaves exactly as before, including one containing `*`, `?`, or `[` — a real file named `report[2024].json` is graded as itself, not as a character class that would match `report2.json`. Globbing only kicks in when the literal path does not exist.
688+
- **Glob matches skip ignored directories.** Expansion runs over the live sandbox root, which also holds harness-created content the agent never wrote (`.venv` for any task with a `python:` block, `node_modules`, `dist`, `build`, `__pycache__`, …), so matches are filtered through the same [`ignore_patterns`](#sandbox-configuration) set used for template copying. A segment your pattern names *literally* is an opt-in and survives, so `dist/**/*.js` still grades `dist`; to un-ignore a directory a wildcard has to discover, use the negation escape hatch — `ignore_patterns: ["!dist"]`.
689+
- Matches are sorted, and directories are skipped.
690+
- `file_exists` passes when the glob matches **at least one** file.
691+
- Content checks require the glob to match **exactly one** file. An ambiguous glob scores 0.0 and reports the matches (first 10, then `+N more`) rather than silently grading one of them — narrow the pattern.
692+
- When a glob resolves, the file that was actually graded is echoed in the criterion's `details` as `resolved: <path>`.
693+
694+
Prefer a glob over a hardcoded path whose leading directory the task prompt never specifies: a correct artifact in an unexpected directory otherwise scores 0.0 on the path alone. Glob away only the segment the prompt leaves free, though — if the free part is an unknown wrapper directory, `**/<Name>.flow` stays unique where a blanket `**/*.flow` turns exactly-one into a hard 0.0 the moment a second flow file exists.
695+
696+
> **Dataset note:** `${row.<field>}` substitution runs over `success_criteria` string leaves, so a row value containing `*`, `?`, or `[` lands inside `path`. Literal-first resolution means such a path still grades the real file when it exists; it falls back to glob expansion only when it does not.
697+
673698
### `file_exists`
674699

675700
Checks if a file exists. **Binary scoring.**

‎src/coder_eval/criteria/file_check.py‎

Lines changed: 7 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -47,21 +47,27 @@ def _check_impl(
4747
has_includes = len(criterion.includes) > 0
4848
has_excludes = len(criterion.excludes) > 0
4949
has_patterns = len(criterion.patterns) > 0
50+
resolved = sandbox.resolved_path_label(criterion.path)
5051

5152
# 2. Pure existence check (no sub-checks specified)
5253
if not has_includes and not has_excludes and not has_patterns:
54+
details = f"File '{criterion.path}' exists"
55+
if resolved:
56+
details += f" (resolved: {resolved})"
5357
return CriterionResult(
5458
criterion_type=criterion.type,
5559
description=criterion.description,
5660
score=1.0,
57-
details=f"File '{criterion.path}' exists",
61+
details=details,
5862
)
5963

6064
# 3. Read file content
6165
content = sandbox.get_file_content(criterion.path)
6266

6367
scores: list[float] = []
6468
details_parts: list[str] = []
69+
if resolved:
70+
details_parts.append(f"Resolved: {resolved}")
6571

6672
# 4a. Includes score
6773
if has_includes:

‎src/coder_eval/criteria/file_contains.py‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -71,6 +71,9 @@ def _check_impl(
7171

7272
# Build details
7373
details_parts = []
74+
resolved = sandbox.resolved_path_label(criterion.path)
75+
if resolved:
76+
details_parts.append(f"Resolved: {resolved}")
7477
details_parts.append(f"Includes: {includes_found}/{includes_total} found")
7578
if criterion.excludes:
7679
excludes_absent = len(criterion.excludes) - sum(1 for exc in criterion.excludes if exc in content)

‎src/coder_eval/criteria/file_exists.py‎

Lines changed: 6 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -37,9 +37,14 @@ def _check_impl(
3737
exists = sandbox.file_exists(criterion.path)
3838
score = 1.0 if exists else 0.0
3939

40+
details = f"File '{criterion.path}' {'exists' if exists else 'does not exist'}"
41+
resolved = sandbox.resolved_path_label(criterion.path)
42+
if resolved:
43+
details += f" (resolved: {resolved})"
44+
4045
return CriterionResult(
4146
criterion_type=criterion.type,
4247
description=criterion.description,
4348
score=score,
44-
details=f"File '{criterion.path}' {'exists' if exists else 'does not exist'}",
49+
details=details,
4550
)

‎src/coder_eval/criteria/file_matches_regex.py‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -78,6 +78,10 @@ def _check_impl(
7878
matched_text = match.group()[:100]
7979
details = f"Pattern '{criterion.pattern}' found but should not be present (matched: '{matched_text}')"
8080

81+
resolved = sandbox.resolved_path_label(criterion.path)
82+
if resolved:
83+
details += f" (resolved: {resolved})"
84+
8185
return CriterionResult(
8286
criterion_type=criterion.type,
8387
description=criterion.description,

‎src/coder_eval/criteria/json_check.py‎

Lines changed: 7 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -91,18 +91,24 @@ def _check_impl(
9191

9292
has_schema = criterion.json_schema is not None
9393
has_assertions = len(criterion.assertions) > 0
94+
resolved = sandbox.resolved_path_label(criterion.path)
9495

9596
# 3. Pure validity check
9697
if not has_schema and not has_assertions:
98+
details = f"'{criterion.path}' is valid JSON"
99+
if resolved:
100+
details += f" (resolved: {resolved})"
97101
return CriterionResult(
98102
criterion_type=criterion.type,
99103
description=criterion.description,
100104
score=1.0,
101-
details=f"'{criterion.path}' is valid JSON",
105+
details=details,
102106
)
103107

104108
scores: list[float] = []
105109
details_parts: list[str] = []
110+
if resolved:
111+
details_parts.append(f"Resolved: {resolved}")
106112

107113
# 4. Schema validation (gates assertions — if schema fails, skip assertions)
108114
if has_schema:

‎src/coder_eval/criteria/reference_comparison.py‎

Lines changed: 6 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -62,18 +62,18 @@ def _check_impl(
6262
error="Sandbox not initialized",
6363
)
6464

65-
# Load agent code
66-
agent_path = sandbox.sandbox_dir / criterion.agent_file
67-
if not agent_path.exists():
65+
# Load agent code through the shared path seam, so `agent_file` resolves
66+
# (glob expansion, ignore filtering, exactly-one) like every other
67+
# sandbox-relative criterion path.
68+
try:
69+
agent_code = sandbox.get_file_content(criterion.agent_file)
70+
except FileNotFoundError:
6871
return CriterionResult(
6972
criterion_type="reference_comparison",
7073
description=criterion.description,
7174
score=0.0,
7275
error=f"Agent file not found: {criterion.agent_file}",
7376
)
74-
75-
try:
76-
agent_code = agent_path.read_text(encoding="utf-8")
7777
except Exception as e:
7878
return CriterionResult(
7979
criterion_type="reference_comparison",

‎src/coder_eval/models/criteria.py‎

Lines changed: 24 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -321,7 +321,9 @@ class FileExistsCriterion(BaseSuccessCriterion):
321321
"""
322322

323323
type: Literal["file_exists"] = "file_exists"
324-
path: str = Field(description="Path to the file that must exist")
324+
path: str = Field(
325+
description="Path to the file that must exist; a glob pattern passes when it matches at least one file"
326+
)
325327

326328

327329
class FileContainsCriterion(BaseSuccessCriterion):
@@ -331,7 +333,7 @@ class FileContainsCriterion(BaseSuccessCriterion):
331333
"""
332334

333335
type: Literal["file_contains"] = "file_contains"
334-
path: str = Field(description="Path to the file to check")
336+
path: str = Field(description="Path to the file to check; may be a glob matching exactly one file")
335337
includes: list[str] = Field(description="List of strings that must be present in the file")
336338
excludes: list[str] | None = Field(default=None, description="List of strings that must NOT be present in the file")
337339

@@ -404,7 +406,7 @@ class FileMatchesRegexCriterion(BaseSuccessCriterion):
404406
"""
405407

406408
type: Literal["file_matches_regex"] = "file_matches_regex"
407-
path: str = Field(description="Path to the file to check")
409+
path: str = Field(description="Path to the file to check; may be a glob matching exactly one file")
408410
pattern: str = Field(description="Regex pattern that must match somewhere in the file")
409411
must_match: bool = Field(default=True, description="If True, pattern must match; if False, pattern must NOT match")
410412
flags: int = Field(default=0, description="Regex flags (e.g., re.IGNORECASE=2, re.MULTILINE=8, re.DOTALL=16)")
@@ -777,8 +779,13 @@ class JsonCheckCriterion(BaseSuccessCriterion):
777779
"""
778780

779781
type: Literal["json_check"] = "json_check"
780-
path: str = Field(description="Path to the JSON file (relative to sandbox root)")
781-
json_schema: str | None = Field(default=None, description="Path to JSON Schema file (relative to sandbox root)")
782+
path: str = Field(
783+
description="Path to the JSON file (relative to sandbox root); may be a glob matching exactly one file"
784+
)
785+
json_schema: str | None = Field(
786+
default=None,
787+
description="Path to JSON Schema file (relative to sandbox root); may be a glob matching exactly one file",
788+
)
782789
assertions: list[JMESPathAssertion] = Field(
783790
default_factory=list, description="JMESPath assertions to evaluate against the parsed JSON"
784791
)
@@ -809,7 +816,9 @@ class FileCheckCriterion(BaseSuccessCriterion):
809816
"""
810817

811818
type: Literal["file_check"] = "file_check"
812-
path: str = Field(description="Path to the file to check (relative to sandbox root)")
819+
path: str = Field(
820+
description="Path to the file to check (relative to sandbox root); may be a glob matching exactly one file"
821+
)
813822
includes: list[str] = Field(default_factory=list, description="Strings that must be present in the file")
814823
excludes: list[str] = Field(default_factory=list, description="Strings that must NOT be present in the file")
815824
patterns: list[RegexPattern] = Field(
@@ -841,7 +850,9 @@ class ReferenceComparisonCriterion(BaseSuccessCriterion):
841850
type: Literal["reference_comparison"] = "reference_comparison"
842851

843852
# Required fields
844-
agent_file: str = Field(description="Path to agent's generated file (relative to sandbox root)")
853+
agent_file: str = Field(
854+
description="Path to agent's generated file (relative to sandbox root); may be a glob matching exactly one file"
855+
)
845856

846857
comparison_method: Literal["ast", "token", "complexity"] = Field(
847858
default="ast",
@@ -1033,7 +1044,12 @@ class ClassificationMatchCriterion(BaseSuccessCriterion):
10331044
"""
10341045

10351046
type: Literal["classification_match"] = "classification_match"
1036-
path: str = Field(description="Path to the file (relative to sandbox) containing the agent's predicted label")
1047+
path: str = Field(
1048+
description=(
1049+
"Path to the file (relative to sandbox) containing the agent's predicted label; "
1050+
"may be a glob matching exactly one file"
1051+
)
1052+
)
10371053
expected_label: str = Field(description="Ground-truth label for this row")
10381054
allowed_labels: list[str] = Field(
10391055
min_length=1,

‎src/coder_eval/sandbox.py‎

Lines changed: 114 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -66,6 +66,27 @@
6666
".wget-hsts",
6767
)
6868

69+
# Characters that make a criterion `path` eligible for glob expansion. Eligible,
70+
# not automatic: `Sandbox.resolve_files` tries the literal path first.
71+
_GLOB_METACHARACTERS = "*?["
72+
73+
# Cap on how many matches an ambiguity error enumerates. The message is
74+
# persisted to task.json and injected into judge prompts, so an unbounded
75+
# listing over a wide pattern is a real payload.
76+
_MAX_LISTED_MATCHES = 10
77+
78+
79+
def _is_glob(path: str) -> bool:
80+
"""Return whether ``path`` contains a glob metacharacter."""
81+
return any(c in path for c in _GLOB_METACHARACTERS)
82+
83+
84+
def _format_matches(matches: list[Path], root: Path) -> str:
85+
"""Render matches as sandbox-relative paths, truncated to a bounded list."""
86+
listed = ", ".join(str(p.relative_to(root)) for p in matches[:_MAX_LISTED_MATCHES])
87+
remaining = len(matches) - _MAX_LISTED_MATCHES
88+
return f"{listed}, +{remaining} more" if remaining > 0 else listed
89+
6990

7091
def _grant_read_traverse(root: Path) -> None:
7192
"""Recursively apply ``chmod a+rX`` semantics under ``root``.
@@ -1093,38 +1114,121 @@ def run_command(self, command: str, timeout: float | int | None = None) -> tuple
10931114
# needs filesystem access beyond the sandbox root (e.g., reading installed packages,
10941115
# system headers). Path traversal protection is handled at the agent permission level.
10951116

1117+
def resolve_files(self, path: str) -> list[Path]:
1118+
"""Resolve a criterion ``path`` to the sandbox files it addresses.
1119+
1120+
A path that names an existing file or directory resolves to itself,
1121+
**even when it contains a glob metacharacter** — a real file called
1122+
``report[2024].json`` is graded as itself rather than reinterpreted as
1123+
a character class that would silently match ``report2.json``. Only when
1124+
the literal does not exist is a path containing ``*``, ``?`` or ``[``
1125+
expanded against the sandbox root, so a criterion can address a file
1126+
whose exact location the task prompt does not pin — e.g. ``**/*.flow``
1127+
matches a scaffolded wrapper directory the agent was free to name.
1128+
1129+
Glob matches are filtered through the sandbox's ignore patterns
1130+
(``.venv``, ``node_modules``, ``dist``, … — see
1131+
:func:`~coder_eval.resources.get_ignore_patterns`), because the sandbox
1132+
root holds harness-created content the agent never authored and
1133+
grading off it is neither fair nor deterministic. Only path segments
1134+
the glob *discovered* are filtered: a segment the pattern names
1135+
literally (``dist/**/*.js``) is an explicit opt-in and survives.
1136+
Matches are sorted so grading is deterministic, and directories are
1137+
dropped so a glob cannot resolve to something unreadable.
1138+
1139+
Args:
1140+
path: Relative path or glob pattern
1141+
1142+
Returns:
1143+
Sorted matching files; empty when nothing matches
1144+
"""
1145+
if not self.sandbox_dir:
1146+
return []
1147+
1148+
# Literal first: an existing path is never reinterpreted as a pattern.
1149+
candidate = self.sandbox_dir / path
1150+
if candidate.exists():
1151+
return [candidate]
1152+
1153+
if not _is_glob(path):
1154+
return []
1155+
1156+
patterns = get_ignore_patterns(self.config.ignore_patterns)
1157+
pinned = {segment for segment in path.split("/") if segment and not _is_glob(segment)}
1158+
1159+
matches: list[Path] = []
1160+
for match in self.sandbox_dir.glob(path):
1161+
if not match.is_file():
1162+
continue
1163+
discovered = [part for part in match.relative_to(self.sandbox_dir).parts if part not in pinned]
1164+
if discovered and should_ignore_path(Path(*discovered), patterns):
1165+
continue
1166+
matches.append(match)
1167+
1168+
return sorted(matches)
1169+
1170+
def resolved_path_label(self, path: str) -> str | None:
1171+
"""Sandbox-relative path a glob resolved to, for grading transparency.
1172+
1173+
With exactly-one-match semantics on content reads, *which* file was
1174+
graded is most of the signal. Returns ``None`` for a literal path
1175+
(nothing was inferred) and for a pattern that did not resolve to
1176+
exactly one file.
1177+
1178+
Args:
1179+
path: Relative path or glob pattern
1180+
1181+
Returns:
1182+
Sandbox-relative path of the single match, or ``None``
1183+
"""
1184+
if not self.sandbox_dir or not _is_glob(path):
1185+
return None
1186+
1187+
matches = self.resolve_files(path)
1188+
if len(matches) != 1:
1189+
return None
1190+
1191+
return str(matches[0].relative_to(self.sandbox_dir))
1192+
10961193
def get_file_content(self, path: str) -> str:
10971194
"""Read the content of a file in the sandbox.
10981195
10991196
Args:
1100-
path: Relative path to the file
1197+
path: Relative path to the file, or a glob pattern matching exactly
1198+
one file
11011199
11021200
Returns:
11031201
File content as string
11041202
11051203
Raises:
11061204
RuntimeError: If sandbox is not set up
1107-
FileNotFoundError: If file doesn't exist
1205+
FileNotFoundError: If nothing matches ``path``
1206+
ValueError: If a glob matches more than one file
11081207
"""
11091208
if not self.sandbox_dir:
11101209
raise RuntimeError("Sandbox not set up")
11111210

1112-
file_path = self.sandbox_dir / path
1113-
return file_path.read_text(encoding="utf-8")
1211+
matches = self.resolve_files(path)
1212+
if not matches:
1213+
raise FileNotFoundError(f"No file matches '{path}' in the sandbox")
1214+
if len(matches) > 1:
1215+
raise ValueError(
1216+
f"Pattern '{path}' matches {len(matches)} files — refusing to guess which to grade: "
1217+
+ _format_matches(matches, self.sandbox_dir)
1218+
)
1219+
1220+
return matches[0].read_text(encoding="utf-8")
11141221

11151222
def file_exists(self, path: str) -> bool:
11161223
"""Check if a file exists in the sandbox.
11171224
11181225
Args:
1119-
path: Relative path to the file
1226+
path: Relative path to the file, or a glob pattern
11201227
11211228
Returns:
1122-
True if file exists, False otherwise
1229+
True if at least one file matches, False otherwise
11231230
"""
1124-
if not self.sandbox_dir:
1125-
return False
1126-
1127-
return (self.sandbox_dir / path).exists()
1231+
return bool(self.resolve_files(path))
11281232

11291233
def list_files(self, path: str = ".") -> list[str]:
11301234
"""List files in a directory within the sandbox.

0 commit comments

Comments
 (0)