Skip to content

fix(truthbench): keep SystemRoot for the Windows sandbox child; fail closed when it cannot start - #431

Merged
kevincostner17 merged 1 commit into
mainfrom
fix/truthbench-sandbox-windows-env
Sep 15, 2026
Merged

kevincostner17 merged 1 commit into
mainfrom
fix/truthbench-sandbox-windows-env

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Summary

This fixes the TruthBench generated-code sandbox on Windows with CPython ≤ 3.10. It also makes the sandbox fail closed when its child interpreter cannot start.

Root cause

verify_generated_code runs generated code in a child interpreter with a deliberately minimal environment, and that environment dropped SystemRoot. On Windows, CPython ≤ 3.10 seeds hash randomization with CryptGenRandom, which needs SystemRoot. Without it the child died at startup with Fatal Python error: _Py_HashRandomization_Init: failed to get random numbers. CPython 3.11+ uses BCryptGenRandom and is unaffected.

Every sandbox case then reported an ordinary generated code exited 1:

  • the timeout, PII-canary and input-overwrite checks never saw the generated code;
  • negative tests that only look for a failure still passed.

This showed up on windows-latest / py3.9 (run 34995394282).

Behaviour change

  • Windows environment: on Windows (os.name == "nt") the sandbox child keeps SystemRoot. The rest of the environment is still scrubbed, and the POSIX environment is unchanged.
  • Start marker: the harness prints a start marker first. A child that exits without it gives a GeneratedCodeResult with infrastructure_failure set, no execute stage and no post-run checks, instead of a normal "exited N" verdict.
  • Run result: run_release raises TruthBenchRunError for such a result, so python -m benchmarks.truthbench run exits 2 (infrastructure failure), including under --check-regressions. It no longer grades the generated_code_sandbox gate on a child that never ran.
  • Unchanged: signal exits still use the existing native-crash retry, and timeouts work as before.

Default-output changes

None (dev tooling only). Run artifacts, the schema and baseline.json are unchanged, and the sandbox's reported stdout does not include the marker.

Tests

  • Environment:
    • _sandbox_env keeps SystemRoot on nt, under either spelling, without needing real Windows.
    • The POSIX environment is exactly unchanged.
    • The Windows environment reaches subprocess.run through verify_generated_code.
  • Child that never starts: a child that dies with the startup error, or exits 0 without running the harness, is an infrastructure failure. It is not retried, has no execute stage and gets no "generated code exited" verdict.
  • Negative controls: the runtime poison, timeout, stdout canary and input overwrite tests now also assert that the child started.
  • CLI: a sandbox child that cannot start exits 2 with INFRASTRUCTURE FAILURE, and no artifacts are written.
  • Existing stubs: existing stubbed-child tests emit the start marker.

Verification

  • ruff check .: clean.
  • pytest tests/truthbench: 300 passed on Python 3.12 and 300 passed on Python 3.9.
  • python -m benchmarks.truthbench run --profile release --backends pandas,polars,duckdb --repeats 2 --check on macOS: exit 0, with all 48 gates passing, including generated_code_sandbox.
  • The Windows py3.9 confirmation comes from CI; this was not run on a real Windows machine locally.

On Windows with CPython 3.10 and earlier, the generated-code sandbox child
died at startup with "Fatal Python error: _Py_HashRandomization_Init:
failed to get random numbers": its scrubbed environment dropped
SystemRoot, which CryptGenRandom needs to seed hash randomization
(3.11+ uses BCryptGenRandom and is unaffected). Every sandbox case then
reported an ordinary "generated code exited 1", so the timeout, canary
and input-overwrite checks never observed the generated code while the
negative tests still saw a failure and passed.

Keep SystemRoot in the child environment on nt only; the POSIX
environment is unchanged.

Also fail closed when the sandbox child cannot start: the harness now
prints a start marker first, and a child that exits without it yields a
GeneratedCodeResult with infrastructure_failure set, no "execute" stage,
and none of the post-run checks. run_release raises TruthBenchRunError
on such a result, so the CLI exits 2 (infrastructure failure) instead of
grading the generated_code_sandbox gate. Signal exits keep the existing
native-crash retry path.
@coderabbitai

coderabbitai Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 0ccacc6d-6a7e-4819-94a8-3e3908f43d88


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

@kevincostner17
kevincostner17 merged commit 2d94d64 into main Sep 15, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant