Conversation
Windows still defaults stdout to cp1252, which turns the routing summary's separators into `?` and makes Chinese prompt output unprintable.
Agent commands were built as POSIX shell strings and handed to `create_subprocess_shell`, which is cmd.exe on Windows, so `mkdir -p` failed on the first call. Process control used `os.killpg`/`SIGKILL`/`SIGHUP`, none of which exist there, so every timeout and Ctrl+C raised AttributeError. Rather than translate the shell strings per platform, remove the shell from the agent path entirely. Everything it was doing is a real subprocess argument: cd <ws> && ... -> cwd= VAR=value <cmd> -> env= layered onto the scrubbed parent environment ... < prompt.md -> the prompt is written to the child's stdin mkdir -p / chmod -> native pathlib calls on the local environment Commands are argv lists now, so `shlex.quote` is gone and no model-supplied value (model id, path, MCP config) can reach a shell parser. `Environment.exec` stays as an explicit escape hatch for callers that genuinely want a shell. Also fixed, all found by running it for real on Windows: - process_group: platform split, CREATE_NEW_PROCESS_GROUP plus `taskkill /T`, and named operations instead of POSIX-only signal numbers. - MAX_PATH: run directories reach 260 characters easily, and the failure is deceptive because mkdir succeeds on the shorter parent while open() on the file inside raises FileNotFoundError. New utils/paths applies the \\?\ prefix, but only to paths that actually exceed the limit. - screenshot(): PowerShell on Windows, screencapture on macOS, the existing X11 tools on Linux. - Deny rules for the claude_code backend now also emit the drive-letter form, so harness-owned paths are actually hidden from the auditor on Windows. Adds tests over LocalEnvironment, the process-group primitives and adapter argv construction, none of which had any coverage before. Test doubles for agent CLIs move to tests/fake_cli, which ships a .cmd shim on Windows because CreateProcess cannot run a `#!/bin/sh` script. BREAKING: CommandAgentAdapter takes argv=/env= instead of command_template=. `EpisodeResult.metadata["command"]` is now the argv list rather than a string; existing readers already accept both.
`_ensure_dir_nofollow`, `_atomic_bytes_write`, `_append_jsonl`, the worker-log open, the dashboard's approval and artifact readers and the manager's event append all require O_NOFOLLOW, O_DIRECTORY and dir_fd, so the run tree could not be written on Windows at all. Each grows a Windows branch that keeps what the platform can express -- reparse-point refusal, atomic replace, the hard-link check before truncation, and the same event-id sequencing -- and documents the one guarantee that cannot be reproduced without directory descriptors: the check-then-use window is not anchored. POSIX behaviour is untouched. Stop/abort go through the process-group helpers instead of os.killpg with raw signal numbers, and name the two signals so a receipt still distinguishes a stop from an abort on a platform that has no SIGKILL.
Creating a symlink needs SeCreateSymbolicLinkPrivilege, which an ordinary account only holds under Developer Mode, so the no-follow tests failed on the fixture rather than on the behaviour they cover. They now skip with a reason where the privilege is missing and run unchanged where it exists. Only the test body is wrapped: pytest's own tmp_path bookkeeping also makes symlinks, and swallowing that would skip the entire suite. The Content-Disposition fixture also asked for a filename containing a quote, which Windows reserves outright; that half of the fixture now only runs on POSIX.
Agent CLIs are launched as plain subprocesses now, so command construction behaves identically on every platform, and run directories escape MAX_PATH automatically.
A run died the moment the provider answered 429, overloaded, or dropped the
connection, which on a long-horizon task throws away hours of work for a
condition that clears on its own.
`_run_role_episode` now backs off exponentially for the two failure kinds that
are actually transient -- rate_limit and network -- and retries: 60s doubling to
a 900s cap, up to 8 attempts or 2 hours total, tunable through
LH_HARNESS_PROVIDER_RETRY_{MAX_ATTEMPTS,BASE_SECONDS,CAP_SECONDS,MAX_TOTAL_SECONDS}
with MAX_ATTEMPTS=0 restoring the old fail-fast behaviour. Terminal kinds
(authentication, quota, model_unavailable, timeout) return immediately as
before; retrying those wastes time or masks a real hang. The wait is visible:
an `agent_runtime_retry` event and a `role_retry` progress record carry the
attempt, the delay and the provider's own message, and a cancel during backoff
still returns a cancelled episode.
Classification also learns the subscription-limit shapes that were previously
read as success: 529, a session/usage limit result whose `is_error` is false but
whose `api_error_status` is 429, and a `rate_limit_event` record whose status is
"rejected".
A Python 3.13+ venv Scripts\python.exe is a launcher that starts the real interpreter as a child process. The supervisor records the launcher pid from Popen, while the worker compares it to os.getpid() of the real interpreter, so every supervised run died with "reservation belongs to another process". The worker identity on Windows now includes os.getppid() (the launcher), and the supervisor_pid lineage check accepts a live recorded parent there, since the grandparent is not portably reachable. POSIX behaviour is unchanged.
Three test files added in 0.1.7 regressed the Windows test run: - test_resume.py and test_resume_routes.py monkeypatch os.killpg in their supervisor fixtures; the attribute does not exist on Windows, so every test in both files errored at setup. The patch now tolerates the absence and also stubs the Windows delivery helper the supervisor uses instead. - test_agent_registry.py wrote `#!/bin/sh` probe stubs, which CreateProcess cannot execute; they now go through tests/fake_cli with Python bodies, which builds the same sh launcher on POSIX. - test_reasoning_effort_chain.py read adapter.command_template, the shell string this branch replaced with argv lists; the assertions now inspect adapter.argv. The expected tokens are unchanged. No test semantics change; the full suite now passes on Windows (405 passed, 42 skipped - the documented symlink-privilege skips).
Dropping the shell took a side effect with it: `sh -c` rewrites `PWD` from `getcwd()` when it starts, so `cd <workspace> && ...` kept the variable and the real working directory in agreement for free. Executing argv directly passes `cwd=` to the OS and leaves the launcher's `PWD` inherited unchanged. That is not cosmetic, because not everything asks the OS. OpenCode resolves the directory its tools operate in from `PWD`, so an agent launched this way worked in whatever directory the operator happened to run `lh-harness` from -- writing files outside the workspace the run promised to contain them in, while the transcript kept naming plausible paths and the auditor confirmed them there. `PWD` is now derived from the `cwd` the child actually gets, in the local environment and in the supervisor's worker launch alike, and `OLDPWD` is dropped rather than blanked so `cd -` behaves like a fresh shell instead of failing. Verified end to end on Linux with all three installed backends launched from a directory outside the workspace: OpenCode now stays inside it, Claude Code and the DeepSeek Harness CLI are unaffected, as they read the working directory from the OS.
`test_deny_rules_cover_drive_letter_paths` guarded its Windows-only assertion by
looking for a colon in `Path("C:/runs/logs").resolve()`. On POSIX that string is
a *relative* path, so it resolves to `<cwd>/C:/runs/logs` -- which contains a
colon as well, and the guard was therefore true everywhere. The Windows branch
ran on Linux and failed there, the one red test in an otherwise green suite.
|
Ran this branch on Linux (Fedora/Nobara, kernel 7.1.4, Python 3.14, uv) to check the POSIX side of the shell-free rewrite. Two things came out of it, both now fixed on the branch (e3ac679, 732ecff). 1. One red test on Linux — the guard, not the product code.
2. A real regression: agents could work outside the workspace.
Reproduced with the harness launched from a directory outside the workspace, task "run pwd, then create marker.txt in the current working directory":
Confirmed the mechanism directly: The fix derives Live end-to-end on Linux, all three backends installed here, each launched from outside its workspace, all
Codex is not installed on this machine, so that backend is untested here. Windows re-verification of these two commits is still to come. |
The dsh npm launcher is a .CMD shim, so its arguments travel through cmd.exe and its 8191-character command-line limit. The headless runner takes the task as a positional argument, and a role prompt is far larger than the limit, so every DeepSeek episode on Windows died in 0.3s with "The command line is too long" before reaching the provider. The headless runner resolves its `task` config from the command line, but a later --patch layer may override that row with a literal - the same override mechanism the runner already uses for the model. On Windows the prompt now rides in the existing per-episode patch file as a JSON-escaped (valid YAML) scalar, and the positional argument shrinks to a fixed placeholder that only satisfies the non-empty check. POSIX keeps passing the prompt as the positional argument, exactly as verified on Linux. Verified against dsh 0.1.0-rc.7 on Windows with a 19 KB prompt: the override is honoured and the episode completes.
|
Completing the picture from the Windows side (Windows 10, Python 3.14 provisioned by uv — no system Python, so Test suite: 409 passed, 42 skipped, 0 failed. The skips are the documented symlink-fixture skips ( Live verification, all three installed backends:
The dsh task cannot travel as a command-line argument on Windows. The npm launcher is a With that, both install shapes (system Python and uv-provisioned launcher venv) and all three backends are verified live on Windows, and the suite is green on both platforms. |
|
Linux verification of Suite: 449 passed, 2 skipped, 0 failed. The POSIX path is untouched in practice, not just by inspection. In a live dsh-backed run the per-episode patch files contain exactly the rows they contained before — Live end-to-end (dsh 0.1.0-rc.7, harness launched from a directory outside the workspace): One observation from these runs that is about the model, not this PR: Remaining gaps, being addressed next on the branch:
Codex remains the one backend with no live verification anywhere — the CLI isn't installed on either machine. The argv rewrite touches its adapter the same way as the other three, and both live-only surprises so far (OpenCode's |
The patch-layer task route was guarded by `os.name` in the runner and in its test alike, so a POSIX run of the suite never executed the Windows branch at all -- and that branch is a contract with dsh's patch precedence, not with the OS: if dsh ever stopped preferring the patch override, agents would silently receive the placeholder string as their task, and only a Windows machine could have noticed. The delivery is now a parameter (`task_via_patch`, defaulting to the platform rule), the placeholder is a named constant, and a parametrised test drives both routes on both platforms with a 9KB prompt full of YAML-hostile characters: the patch route must keep the prompt off the command line and carry it as one exactly-escaped scalar, the positional route must keep the patch file free of the `headless-runner` row. Behaviour at the defaults is unchanged and still covered by the existing test.
The PR's own history is the argument: the branch was verified green on each platform by hand, and each round of hand-verification still found something the other platform could not see (a POSIX-only test guard, a cmd.exe-only command-line limit). A matrix of ubuntu + windows at both ends of requires-python (3.10 and 3.14) makes that check automatic for every push and pull request. The suite needs no Node toolchain -- the Web bundle is a packaging artifact -- so the job is checkout, setup-python, `pip install -e ".[test]"`, pytest. The Windows symlink fixtures skip themselves on runners without SeCreateSymbolicLinkPrivilege, which is expected and green.
The first admin-privileged CI run caught three Windows bugs the local test machine could not see, because SeCreateSymbolicLinkPrivilege changed which tests actually ran and a busier pid space changed what a stray probe could hit. The severe one: two code paths asked "is this process alive" with `os.kill(pid, 0)`. That is the POSIX idiom -- and a kill switch on Windows, where os.kill can only deliver console control events and unconditionally calls TerminateProcess for every other signal value, zero included. The supervisor's `_is_alive` poll would kill the very worker it was checking on (or a pid-reuse victim), and the venv-launcher owner check in `--supervised` adoption would kill the supervisor whose reservation it was validating. On the CI runner, a status poll probing a test's fake pid terminated a live runner process and took pytest down with it (`KeyboardInterrupt`, suite dead after 27 tests). Both sites now use `process_alive`, a real observe-only probe: `tasklist` on Windows, `os.kill(pid, 0)` where it actually means "probe", with pid <= 0 rejected outright (POSIX group semantics, Windows idle process). Also from that run: * The Windows atomic-write branch refused a symlinked *destination*, where the POSIX branch's rename atomically evicts the link and installs a private regular file. Refusing is strictly weaker -- the crash-report writer swallows the OSError and the planted link survives for a later, less careful writer -- and os.replace already operates on the name, never the target. The refusal is gone; Windows now evicts exactly like POSIX. * The supervisor's idempotency lock on Windows surfaced the shared helper's "secure control-bus locking" message; it now reports "secure supervisor locking" like the POSIX branch, keeping the message a caller can match on. New cross-platform tests cover all three: the probe must see a live child and leave it running (the old Windows code kills it right there), must see a dead one, must reject nonsense pids; the atomic write must evict a symlinked destination without touching the file it pointed at.
Two corrections to the previous commit's own additions, both caught by the first Windows CI round that got past the liveness-probe kill: * ``_run_quiet`` reached ``tasklist``/``taskkill`` through ``subprocess.run``, which resolves ``Popen`` through the module namespace at call time -- and the supervisor tests stub ``subprocess.Popen`` with a fake worker that is no context manager, so every resume-path status refresh on Windows died with ``AttributeError: __enter__``. The probes now hold the real ``Popen`` captured at import, the same pattern service.py already documents for its worker spawn. * The both-routes dsh test drove the *positional* delivery with a 9KB multi-line prompt on every platform. On Windows that argument cannot fit through the .CMD shim -- the exact limitation the patch route exists to bypass -- so the test was asserting the platform bug away. The big hostile prompt now exercises the patch route only; the positional case proves just the route itself with a plain prompt.
|
Follow-up on the "remaining gaps" list from the previous comment — all three are now closed on the branch, plus one correction to my own earlier analysis. CI is on and fully green ( Getting there took three rounds, and the Windows rounds each caught something the hand-testing could not see. The GitHub runner executes as admin, so it holds
Correction to my earlier model observation. The Linux suite at head: 455 passed, 2 skipped. Codex remains the one backend never verified live on either platform. |
…run" This reverts commit cbf36a8.
|
Re-verified the new tip (
|
|
One more Windows round, this time targeting the platform's classic trouble spots — Unicode, path limits, process cleanup, and lock contention. Six scenarios, all live runs with the
Together with the earlier reports this closes out the Windows verification matrix: suite green on both platforms and both privilege shapes, three backends live, both Python install shapes, Unicode/MAX_PATH/timeout/concurrency edges exercised for real. |
…ands The windows-latest lanes exercise platform support this branch does not carry: it is based on a main whose supervisor still calls os.killpg and whose agent stubs are #!/bin/sh scripts, so those lanes fail on known pre-existing breakage rather than on anything in this change. AMAP-ML#57 brings the Windows support together with the full two-platform matrix; when it merges, its version of this workflow supersedes this one.
…ands The windows-latest lanes exercise platform support this branch does not carry: it is based on a main whose supervisor still calls os.killpg and whose agent stubs are #!/bin/sh scripts, so those lanes fail on known pre-existing breakage rather than on anything in this change. AMAP-ML#57 brings the Windows support together with the full two-platform matrix; when it merges, its version of this workflow supersedes this one.
…ands The windows-latest lanes exercise platform support this branch does not carry: it is based on a main whose supervisor still calls os.killpg and whose agent stubs are #!/bin/sh scripts, so those lanes fail on known pre-existing breakage rather than on anything in this change. AMAP-ML#57 brings the Windows support together with the full two-platform matrix; when it merges, its version of this workflow supersedes this one.
…ands The windows-latest lanes exercise platform support this branch does not carry: it is based on a main whose supervisor still calls os.killpg and whose agent stubs are #!/bin/sh scripts, so those lanes fail on known pre-existing breakage rather than on anything in this change. AMAP-ML#57 brings the Windows support together with the full two-platform matrix; when it merges, its version of this workflow supersedes this one.
…ands The windows-latest lanes exercise platform support this branch does not carry: it is based on a main whose supervisor still calls os.killpg and whose agent stubs are #!/bin/sh scripts, so those lanes fail on known pre-existing breakage rather than on anything in this change. AMAP-ML#57 brings the Windows support together with the full two-platform matrix; when it merges, its version of this workflow supersedes this one.
|
This PR was auto-closed when I cleaned up the head branches in my fork. The Windows work lives on and is even exercised by CI (windows-latest in the matrix) in my fork's |
What this does
Makes LongHorizon-Harness work on Windows out of the box:
uv tool install,lh-harness run, the Web workbench, stop/abort, and the full test suite — no WSL, no shell shims.Windows was previously unusable at several independent layers: the secure no-follow filesystem walk (
O_NOFOLLOW/O_DIRECTORY/dir_fd/fcntl.flockare POSIX-only) refused to create the run tree at startup; agent commands were POSIX shell strings executed bycmd.exe; process control relied onos.killpg/SIGKILL/SIGHUP/start_new_session; screenshots had no Windows path.Changes
cwd=,env=, prompt over stdin). This removes the platform-specific shell entirely, and as a side effect no model-supplied value (model id, path, MCP config) can reach a shell parser anymore.CommandAgentAdaptertakesargv=/env=instead ofcommand_template=.control_bus, worker log, dashboard approval/artifact readers, event append: reparse-point refusal, atomic replace andmsvcrtbyte locks keep what the platform can express; POSIX behaviour is untouched.CREATE_NEW_PROCESS_GROUP+taskkill /Treplacekillpg; named terminate/force operations instead of raw signal numbers, so stop/abort receipts stay meaningful on a platform with noSIGKILL.\\?\prefix only when they need it (utils/paths.py).Graphics.CopyFromScreen.Scripts\python.exelaunches the real interpreter as a child, so the supervised-run owner checks now accept the recorded launcher (parent) pid on Windows.SeCreateSymbolicLinkPrivilege; CLI stubs go through a cross-platformfake_clihelper (#!/bin/shscripts cannot be executed byCreateProcess); the 0.1.7 test additions that monkeypatchos.killpgor readcommand_templateare adapted.Transient provider failures back off and retry— not Windows-specific, so it moved to its own PR: Wait out transient provider failures instead of aborting the run #59 (added here as a commit and reverted, keeping this PR's diff Windows-only).Testing
claude_codebackend: CLI runs toResult: completewith clean audits; Web workbench end-to-end (create run → live status/events → completion approval → resolve → final report); stop/abort; round artifacts and trajectories in the dashboard.