Skip to content

Windows support: run the harness natively on Windows - #57

Closed
TON14 wants to merge 16 commits into
AMAP-ML:mainfrom
TON14:support-windows
Closed

TON14 wants to merge 16 commits into
AMAP-ML:mainfrom
TON14:support-windows

Conversation

@TON14

@TON14 TON14 commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

What this does

Makes LongHorizon-Harness work on Windows out of the box: uv tool install, lh-harness run, the Web workbench, stop/abort, and the full test suite — no WSL, no shell shims.

Windows was previously unusable at several independent layers: the secure no-follow filesystem walk (O_NOFOLLOW/O_DIRECTORY/dir_fd/fcntl.flock are POSIX-only) refused to create the run tree at startup; agent commands were POSIX shell strings executed by cmd.exe; process control relied on os.killpg/SIGKILL/SIGHUP/start_new_session; screenshots had no Windows path.

Changes

  • Run agent CLIs without a shell — commands are argv lists now (cwd=, env=, prompt over stdin). This removes the platform-specific shell entirely, and as a side effect no model-supplied value (model id, path, MCP config) can reach a shell parser anymore. CommandAgentAdapter takes argv=/env= instead of command_template=.
  • Windows branches for the 0.1.4 hardening layer — control_bus, worker log, dashboard approval/artifact readers, event append: reparse-point refusal, atomic replace and msvcrt byte locks keep what the platform can express; POSIX behaviour is untouched.
  • Process groups — CREATE_NEW_PROCESS_GROUP + taskkill /T replace killpg; named terminate/force operations instead of raw signal numbers, so stop/abort receipts stay meaningful on a platform with no SIGKILL.
  • MAX_PATH — run directories easily exceed 260 chars; long paths get the \\?\ prefix only when they need it (utils/paths.py).
  • Screenshots on Windows via PowerShell Graphics.CopyFromScreen.
  • UTF-8 console output on legacy code pages.
  • Venv-launcher pids — on Python 3.13+ a venv Scripts\python.exe launches the real interpreter as a child, so the supervised-run owner checks now accept the recorded launcher (parent) pid on Windows.
  • Tests run on Windows — symlink fixtures skip without SeCreateSymbolicLinkPrivilege; CLI stubs go through a cross-platform fake_cli helper (#!/bin/sh scripts cannot be executed by CreateProcess); the 0.1.7 test additions that monkeypatch os.killpg or read command_template are adapted.
  • Transient provider failures back off and retry — not Windows-specific, so it moved to its own PR: Wait out transient provider failures instead of aborting the run #59 (added here as a commit and reverted, keeping this PR's diff Windows-only).

Testing

  • Full suite on Windows 10 (Python 3.14, uv): 405 passed, 42 skipped (the documented symlink-privilege skips), 0 failed.
  • Live end-to-end on Windows with the claude_code backend: CLI runs to Result: complete with clean audits; Web workbench end-to-end (create run → live status/events → completion approval → resolve → final report); stop/abort; round artifacts and trajectories in the dashboard.
  • Both install shapes verified: a full Python install, and a machine with no system Python where uv provisions Python 3.13+ (the venv-launcher case).
  • POSIX code paths are untouched by design; the suite passes unchanged there.

TON14 added 10 commits August 20, 2026 16:11
Windows still defaults stdout to cp1252, which turns the routing summary's
separators into `?` and makes Chinese prompt output unprintable.
Agent commands were built as POSIX shell strings and handed to
`create_subprocess_shell`, which is cmd.exe on Windows, so `mkdir -p` failed on
the first call. Process control used `os.killpg`/`SIGKILL`/`SIGHUP`, none of
which exist there, so every timeout and Ctrl+C raised AttributeError.

Rather than translate the shell strings per platform, remove the shell from the
agent path entirely. Everything it was doing is a real subprocess argument:

  cd <ws> && ...        -> cwd=
  VAR=value <cmd>       -> env= layered onto the scrubbed parent environment
  ... < prompt.md       -> the prompt is written to the child's stdin
  mkdir -p / chmod      -> native pathlib calls on the local environment

Commands are argv lists now, so `shlex.quote` is gone and no model-supplied
value (model id, path, MCP config) can reach a shell parser. `Environment.exec`
stays as an explicit escape hatch for callers that genuinely want a shell.

Also fixed, all found by running it for real on Windows:

- process_group: platform split, CREATE_NEW_PROCESS_GROUP plus `taskkill /T`,
  and named operations instead of POSIX-only signal numbers.
- MAX_PATH: run directories reach 260 characters easily, and the failure is
  deceptive because mkdir succeeds on the shorter parent while open() on the
  file inside raises FileNotFoundError. New utils/paths applies the \\?\
  prefix, but only to paths that actually exceed the limit.
- screenshot(): PowerShell on Windows, screencapture on macOS, the existing
  X11 tools on Linux.
- Deny rules for the claude_code backend now also emit the drive-letter form,
  so harness-owned paths are actually hidden from the auditor on Windows.

Adds tests over LocalEnvironment, the process-group primitives and adapter argv
construction, none of which had any coverage before. Test doubles for agent
CLIs move to tests/fake_cli, which ships a .cmd shim on Windows because
CreateProcess cannot run a `#!/bin/sh` script.

BREAKING: CommandAgentAdapter takes argv=/env= instead of command_template=.
`EpisodeResult.metadata["command"]` is now the argv list rather than a string;
existing readers already accept both.
`_ensure_dir_nofollow`, `_atomic_bytes_write`, `_append_jsonl`, the worker-log
open, the dashboard's approval and artifact readers and the manager's event
append all require O_NOFOLLOW, O_DIRECTORY and dir_fd, so the run tree could not
be written on Windows at all.

Each grows a Windows branch that keeps what the platform can express --
reparse-point refusal, atomic replace, the hard-link check before truncation,
and the same event-id sequencing -- and documents the one guarantee that cannot
be reproduced without directory descriptors: the check-then-use window is not
anchored. POSIX behaviour is untouched.

Stop/abort go through the process-group helpers instead of os.killpg with raw
signal numbers, and name the two signals so a receipt still distinguishes a stop
from an abort on a platform that has no SIGKILL.
Creating a symlink needs SeCreateSymbolicLinkPrivilege, which an ordinary
account only holds under Developer Mode, so the no-follow tests failed on the
fixture rather than on the behaviour they cover. They now skip with a reason
where the privilege is missing and run unchanged where it exists. Only the test
body is wrapped: pytest's own tmp_path bookkeeping also makes symlinks, and
swallowing that would skip the entire suite.

The Content-Disposition fixture also asked for a filename containing a quote,
which Windows reserves outright; that half of the fixture now only runs on
POSIX.
Agent CLIs are launched as plain subprocesses now, so command construction
behaves identically on every platform, and run directories escape MAX_PATH
automatically.
A run died the moment the provider answered 429, overloaded, or dropped the
connection, which on a long-horizon task throws away hours of work for a
condition that clears on its own.

`_run_role_episode` now backs off exponentially for the two failure kinds that
are actually transient -- rate_limit and network -- and retries: 60s doubling to
a 900s cap, up to 8 attempts or 2 hours total, tunable through
LH_HARNESS_PROVIDER_RETRY_{MAX_ATTEMPTS,BASE_SECONDS,CAP_SECONDS,MAX_TOTAL_SECONDS}
with MAX_ATTEMPTS=0 restoring the old fail-fast behaviour. Terminal kinds
(authentication, quota, model_unavailable, timeout) return immediately as
before; retrying those wastes time or masks a real hang. The wait is visible:
an `agent_runtime_retry` event and a `role_retry` progress record carry the
attempt, the delay and the provider's own message, and a cancel during backoff
still returns a cancelled episode.

Classification also learns the subscription-limit shapes that were previously
read as success: 529, a session/usage limit result whose `is_error` is false but
whose `api_error_status` is 429, and a `rate_limit_event` record whose status is
"rejected".
A Python 3.13+ venv Scripts\python.exe is a launcher that starts the real
interpreter as a child process. The supervisor records the launcher pid from
Popen, while the worker compares it to os.getpid() of the real interpreter,
so every supervised run died with "reservation belongs to another process".

The worker identity on Windows now includes os.getppid() (the launcher), and
the supervisor_pid lineage check accepts a live recorded parent there, since
the grandparent is not portably reachable. POSIX behaviour is unchanged.
Three test files added in 0.1.7 regressed the Windows test run:

- test_resume.py and test_resume_routes.py monkeypatch os.killpg in their
  supervisor fixtures; the attribute does not exist on Windows, so every
  test in both files errored at setup. The patch now tolerates the absence
  and also stubs the Windows delivery helper the supervisor uses instead.
- test_agent_registry.py wrote `#!/bin/sh` probe stubs, which CreateProcess
  cannot execute; they now go through tests/fake_cli with Python bodies,
  which builds the same sh launcher on POSIX.
- test_reasoning_effort_chain.py read adapter.command_template, the shell
  string this branch replaced with argv lists; the assertions now inspect
  adapter.argv. The expected tokens are unchanged.

No test semantics change; the full suite now passes on Windows
(405 passed, 42 skipped - the documented symlink-privilege skips).
Dropping the shell took a side effect with it: `sh -c` rewrites `PWD` from
`getcwd()` when it starts, so `cd <workspace> && ...` kept the variable and the
real working directory in agreement for free. Executing argv directly passes
`cwd=` to the OS and leaves the launcher's `PWD` inherited unchanged.

That is not cosmetic, because not everything asks the OS. OpenCode resolves the
directory its tools operate in from `PWD`, so an agent launched this way worked
in whatever directory the operator happened to run `lh-harness` from -- writing
files outside the workspace the run promised to contain them in, while the
transcript kept naming plausible paths and the auditor confirmed them there.

`PWD` is now derived from the `cwd` the child actually gets, in the local
environment and in the supervisor's worker launch alike, and `OLDPWD` is
dropped rather than blanked so `cd -` behaves like a fresh shell instead of
failing. Verified end to end on Linux with all three installed backends
launched from a directory outside the workspace: OpenCode now stays inside it,
Claude Code and the DeepSeek Harness CLI are unaffected, as they read the
working directory from the OS.
`test_deny_rules_cover_drive_letter_paths` guarded its Windows-only assertion by
looking for a colon in `Path("C:/runs/logs").resolve()`. On POSIX that string is
a *relative* path, so it resolves to `<cwd>/C:/runs/logs` -- which contains a
colon as well, and the guard was therefore true everywhere. The Windows branch
ran on Linux and failed there, the one red test in an otherwise green suite.
@TON14

TON14 commented Aug 20, 2026

Copy link
Copy Markdown
Contributor Author

Ran this branch on Linux (Fedora/Nobara, kernel 7.1.4, Python 3.14, uv) to check the POSIX side of the shell-free rewrite. Two things came out of it, both now fixed on the branch (e3ac679, 732ecff).

1. One red test on Linux — the guard, not the product code.

test_deny_rules_cover_drive_letter_paths decided whether to run its Windows-only assertion by looking for a colon in Path("C:/runs/logs").resolve(). On POSIX that is a relative path, so it resolves to <cwd>/C:/runs/logs — which also contains a colon. The guard was true on every platform, so the Windows branch ran on Linux and failed there. Now it asks os.name == "nt". Suite on Linux: 444 passed, 2 skipped, 1 failed before; 449 passed, 2 skipped after (the new tests below included).

2. A real regression: agents could work outside the workspace.

sh -c rewrites PWD from getcwd() when it starts, so the old cd <workspace> && ... command strings kept PWD and the actual working directory in agreement for free. Executing argv directly passes cwd= to the OS and leaves the launcher's PWD inherited unchanged — and not everything asks the OS. OpenCode resolves the directory its tools operate in from PWD.

Reproduced with the harness launched from a directory outside the workspace, task "run pwd, then create marker.txt in the current working directory":

pwd inside the agent where marker.txt landed
main workspace workspace
this branch (before the fix) launch directory launch directory
this branch (after the fix) workspace workspace

Confirmed the mechanism directly: subprocess.run(["opencode", ...], cwd=X) with a stale PWD makes OpenCode's bash tool report the stale directory; setting PWD=X makes it report X. The failure mode is quiet — the transcript keeps naming plausible paths and the auditor confirms the file at the path it was really written to, so the run reports complete.

The fix derives PWD from the cwd the child actually gets, in LocalEnvironment._spawn and in the supervisor's worker launch alike, and drops OLDPWD instead of blanking it so cd - behaves like a fresh shell. Regression tests cover PWD matching cwd, OLDPWD not being inherited, and PWD being left alone when no cwd is imposed.

Live end-to-end on Linux, all three backends installed here, each launched from outside its workspace, all Result: complete with clean audits and every file inside the workspace:

  • claude_code (Claude Code 2.1.235)
  • opencode (OpenCode 1.18.19)
  • deepseek_harness (dsh 0.1.0-rc.7)

Codex is not installed on this machine, so that backend is untested here. Windows re-verification of these two commits is still to come.

The dsh npm launcher is a .CMD shim, so its arguments travel through
cmd.exe and its 8191-character command-line limit. The headless runner
takes the task as a positional argument, and a role prompt is far larger
than the limit, so every DeepSeek episode on Windows died in 0.3s with
"The command line is too long" before reaching the provider.

The headless runner resolves its `task` config from the command line,
but a later --patch layer may override that row with a literal - the
same override mechanism the runner already uses for the model. On
Windows the prompt now rides in the existing per-episode patch file as
a JSON-escaped (valid YAML) scalar, and the positional argument shrinks
to a fixed placeholder that only satisfies the non-empty check. POSIX
keeps passing the prompt as the positional argument, exactly as
verified on Linux.

Verified against dsh 0.1.0-rc.7 on Windows with a 19 KB prompt: the
override is honoured and the episode completes.
@TON14

TON14 commented Aug 20, 2026

Copy link
Copy Markdown
Contributor Author

Completing the picture from the Windows side (Windows 10, Python 3.14 provisioned by uv — no system Python, so Scripts\python.exe is the 3.13+ venv launcher; the counterpart of the Linux report above).

Test suite: 409 passed, 42 skipped, 0 failed. The skips are the documented symlink-fixture skips (SeCreateSymbolicLinkPrivilege is only held under Developer Mode). The two commits from the Linux round (732ecff, e3ac679) pass here unchanged, including the new PWD tests.

Live verification, all three installed backends:

  • claude_code — full CLI runs to Result: complete with clean audits, and the Web workbench end to end: create a run from the API, live status/events, the completion approval gate, resolve, stop, round artifacts and trajectories in the dashboard.
  • opencode 1.18.19 — reproduced the PWD scenario from the Linux report (harness launched from a directory outside the workspace): the agent worked inside the workspace, marker.txt landed there with byte-exact content. The fix holds on Windows too.
  • deepseek_harness (dsh 0.1.0-rc.7) — found and fixed one more Windows-only failure, now on the branch as ed6d772:

The dsh task cannot travel as a command-line argument on Windows. The npm launcher is a .CMD shim, so its arguments go through cmd.exe and its 8191-character limit; a role prompt is far larger, and every DeepSeek episode died in 0.3s with The command line is too long before reaching the provider. The headless runner resolves its task config from the command line, but a later --patch layer may override that row with a literal — the same mechanism the runner already uses to pin the model. On Windows the prompt now rides in the existing per-episode patch file as a JSON-escaped (valid YAML) scalar, and the positional argument shrinks to a fixed placeholder that only satisfies the non-empty check. POSIX keeps passing the prompt as the positional, exactly as verified on Linux. Confirmed with a 19 KB prompt directly against the runner, and end to end: a dsh-backed run now completes on Windows with a clean audit (Result: complete).

With that, both install shapes (system Python and uv-provisioned launcher venv) and all three backends are verified live on Windows, and the suite is green on both platforms.

@TON14

TON14 commented Aug 20, 2026

Copy link
Copy Markdown
Contributor Author

Linux verification of ed6d772 (the dsh patch-layer commit), same machine as the earlier Linux report.

Suite: 449 passed, 2 skipped, 0 failed.

The POSIX path is untouched in practice, not just by inspection. In a live dsh-backed run the per-episode patch files contain exactly the rows they contained before — agent-default-model / provider / model — with no headless-runner row, and the prompt still travels as the positional argument.

Live end-to-end (dsh 0.1.0-rc.7, harness launched from a directory outside the workspace): Result: complete, marker.txt created inside the workspace with the exact content. So the commit changes nothing observable on Linux while fixing the Windows-only 8191-char failure.

One observation from these runs that is about the model, not this PR: deepseek-v4-flash as the manager frequently fails to emit a valid Next: route — in one run 2 of 4 rounds went to next=invalid, and a 3-round budget burned entirely on it (max_rounds_exhausted). The same pattern shows on OpenCode's free-tier model. Worth keeping in mind when sizing --max-rounds for runs driven by small models.

Remaining gaps, being addressed next on the branch:

  • the Windows branch of the dsh runner is currently exercised only on Windows (both the code and its test gate on os.name), so a Linux run never executes the patch-task path — making the branch injectable so both variants are tested on both platforms;
  • the task-via---patch contract relies on dsh's layer precedence (an rc-versioned CLI) — the cross-platform test above is the tripwire for that;
  • no CI on the repo: adding a GitHub Actions matrix (ubuntu + windows) so both platforms' suites run on every push.

Codex remains the one backend with no live verification anywhere — the CLI isn't installed on either machine. The argv rewrite touches its adapter the same way as the other three, and both live-only surprises so far (OpenCode's PWD, dsh's 8191 limit) were invisible to unit tests, so a live Codex run before merge would be valuable if anyone has it installed.

@TON14
TON14 marked this pull request as draft August 20, 2026 16:00
TON14 added 4 commits August 20, 2026 19:03
The patch-layer task route was guarded by `os.name` in the runner and in its
test alike, so a POSIX run of the suite never executed the Windows branch at
all -- and that branch is a contract with dsh's patch precedence, not with the
OS: if dsh ever stopped preferring the patch override, agents would silently
receive the placeholder string as their task, and only a Windows machine could
have noticed.

The delivery is now a parameter (`task_via_patch`, defaulting to the platform
rule), the placeholder is a named constant, and a parametrised test drives
both routes on both platforms with a 9KB prompt full of YAML-hostile
characters: the patch route must keep the prompt off the command line and
carry it as one exactly-escaped scalar, the positional route must keep the
patch file free of the `headless-runner` row. Behaviour at the defaults is
unchanged and still covered by the existing test.
The PR's own history is the argument: the branch was verified green on each
platform by hand, and each round of hand-verification still found something
the other platform could not see (a POSIX-only test guard, a cmd.exe-only
command-line limit). A matrix of ubuntu + windows at both ends of
requires-python (3.10 and 3.14) makes that check automatic for every push and
pull request.

The suite needs no Node toolchain -- the Web bundle is a packaging artifact --
so the job is checkout, setup-python, `pip install -e ".[test]"`, pytest. The
Windows symlink fixtures skip themselves on runners without
SeCreateSymbolicLinkPrivilege, which is expected and green.
The first admin-privileged CI run caught three Windows bugs the local test
machine could not see, because SeCreateSymbolicLinkPrivilege changed which
tests actually ran and a busier pid space changed what a stray probe could
hit.

The severe one: two code paths asked "is this process alive" with
`os.kill(pid, 0)`. That is the POSIX idiom -- and a kill switch on Windows,
where os.kill can only deliver console control events and unconditionally
calls TerminateProcess for every other signal value, zero included. The
supervisor's `_is_alive` poll would kill the very worker it was checking on
(or a pid-reuse victim), and the venv-launcher owner check in `--supervised`
adoption would kill the supervisor whose reservation it was validating. On
the CI runner, a status poll probing a test's fake pid terminated a live
runner process and took pytest down with it (`KeyboardInterrupt`, suite dead
after 27 tests). Both sites now use `process_alive`, a real observe-only
probe: `tasklist` on Windows, `os.kill(pid, 0)` where it actually means
"probe", with pid <= 0 rejected outright (POSIX group semantics, Windows
idle process).

Also from that run:

* The Windows atomic-write branch refused a symlinked *destination*, where
  the POSIX branch's rename atomically evicts the link and installs a
  private regular file. Refusing is strictly weaker -- the crash-report
  writer swallows the OSError and the planted link survives for a later,
  less careful writer -- and os.replace already operates on the name, never
  the target. The refusal is gone; Windows now evicts exactly like POSIX.

* The supervisor's idempotency lock on Windows surfaced the shared helper's
  "secure control-bus locking" message; it now reports "secure supervisor
  locking" like the POSIX branch, keeping the message a caller can match on.

New cross-platform tests cover all three: the probe must see a live child
and leave it running (the old Windows code kills it right there), must see
a dead one, must reject nonsense pids; the atomic write must evict a
symlinked destination without touching the file it pointed at.
Two corrections to the previous commit's own additions, both caught by the
first Windows CI round that got past the liveness-probe kill:

* ``_run_quiet`` reached ``tasklist``/``taskkill`` through ``subprocess.run``,
  which resolves ``Popen`` through the module namespace at call time -- and
  the supervisor tests stub ``subprocess.Popen`` with a fake worker that is
  no context manager, so every resume-path status refresh on Windows died
  with ``AttributeError: __enter__``. The probes now hold the real ``Popen``
  captured at import, the same pattern service.py already documents for its
  worker spawn.

* The both-routes dsh test drove the *positional* delivery with a 9KB
  multi-line prompt on every platform. On Windows that argument cannot fit
  through the .CMD shim -- the exact limitation the patch route exists to
  bypass -- so the test was asserting the platform bug away. The big hostile
  prompt now exercises the patch route only; the positional case proves just
  the route itself with a plain prompt.
@TON14

TON14 commented Aug 20, 2026

Copy link
Copy Markdown
Contributor Author

Follow-up on the "remaining gaps" list from the previous comment — all three are now closed on the branch, plus one correction to my own earlier analysis.

CI is on and fully green (.github/workflows/tests.yml, commit eec3250): ubuntu + windows × Python 3.10/3.14. The current head passes all four jobs — the first fully automated green run on both platforms.

Getting there took three rounds, and the Windows rounds each caught something the hand-testing could not see. The GitHub runner executes as admin, so it holds SeCreateSymbolicLinkPrivilege — the symlink-hardening tests that skip on an ordinary Windows account actually ran for the first time, and a busier pid space changed what a stray probe could hit. Found and fixed (c5f75ea, 6a0bde7):

  • os.kill(pid, 0) is a liveness probe on POSIX and a kill switch on Windows — os.kill there can only deliver console control events and unconditionally calls TerminateProcess for any other signal value, zero included. The supervisor's _is_alive poll would kill the very worker it was checking on, and the --supervised owner check would kill the supervisor it was validating. On CI, a status poll probing a test's fake pid terminated a live runner process and took pytest down with it (KeyboardInterrupt, suite dead after 27 tests). Both sites now use a dedicated observe-only probe (tasklist on Windows), with cross-platform regression tests — including "the probe must leave the probed process running", which the old code fails by killing it mid-assert.
  • The Windows atomic-write branch refused a symlinked destination where POSIX rename atomically evicts the link and installs a private regular file. Refusing is strictly weaker — the crash-report writer swallowed the OSError and the planted link survived. Windows now evicts exactly like POSIX.
  • The supervisor idempotency lock on Windows surfaced the shared helper's "secure control-bus locking" message instead of its own "secure supervisor locking"; normalized.
  • Both dsh task deliveries (positional and patch-layer) are now exercised on every platform via an injectable parameter (7a67eb1) — the patch route is a contract with dsh's layer precedence, and a contract only Windows machines could test would rot quietly.

Correction to my earlier model observation. The next=invalid rounds I attributed to deepseek-v4-flash are not a model-quality problem — they are a parser gap. The manager prompt displays every route inside backticks ("then exactly one route: Next: gui, Next: cli, ..."), the model copies that formatting literally, and parse_role_manager_next_step strips * (so bold **Next: cli** is accepted) but not backticks — so a fully valid plan ending in a code-span-wrapped route scores invalid, and a 3-round run burned its whole budget on formatting the harness's own instruction taught the model. That bug exists on main and is out of this PR's scope; a one-line fix with tests is ready and will go up separately.

Linux suite at head: 455 passed, 2 skipped. Codex remains the one backend never verified live on either platform.

@TON14

TON14 commented Aug 20, 2026

Copy link
Copy Markdown
Contributor Author

Re-verified the new tip (3e1bcd9) on the real Windows 10 machine — the non-admin, no-Developer-Mode counterpart of the privileged CI runner, so the two environments now cover both privilege shapes.

  • Suite: 414 passed, 43 skipped, 0 failed. Every skip is an explicit platform skip (symlink privilege, POSIX-only locking/tempfile paths, the macOS /var alias).
  • The process_alive fix is confirmed by live behaviour, not just tests: a supervised Web run survived several minutes of 10-second status polls that previously went through os.kill(pid, 0) — the venv-launcher adoption path included — and completed with a clean audit (report: complete).
  • Bonus coverage from a mis-scoped task on my side: the executor searched, declined to invent the missing file, the auditor returned blocked/needs_revision, and the blocking gate surfaced in the Web API and resolved correctly on Windows. The escalation path works end to end there too.
  • The retry revert keeps the branch strictly Windows-scoped; suite and live runs are unaffected by it.

@TON14

TON14 commented Aug 20, 2026

Copy link
Copy Markdown
Contributor Author

One more Windows round, this time targeting the platform's classic trouble spots — Unicode, path limits, process cleanup, and lock contention. Six scenarios, all live runs with the claude_code backend on the current tip, all green, no code changes needed:

Scenario What it exercises Result
Workspace path with spaces + Cyrillic (lh тест папка) Path handling, argv/task Unicode, file content round-trip ✅ complete, byte-exact Cyrillic content
Workspace ~250 chars deep, run tree beyond MAX_PATH The \\?\ long-path layer (utils/paths.py) in production, not just tests ✅ complete
Executor timeout (agent told to wait 300 s, --cli-executor-timeout 45) taskkill /T tree cleanup and graceful round failure ✅ killed at ~50 s, zero orphaned processes (checked via tasklist/WMI), run concluded cleanly with a final reply
Two runs created simultaneously through the Web supervisor .supervisor.lock and the msvcrt byte-lock fallbacks under real contention ✅ both created in the same second, ran in parallel, both report: complete
Mixed-script everything: path lh-中文-тест мешанина, file отчёт 报告 file.txt, trilingual content Unicode across path, filename, and content at once ✅ complete, byte-exact
--prompt-language zh The full Chinese prompt pipeline through Windows files and console ✅ complete, audits complete/clean/aligned, Chinese-named output file correct

Together with the earlier reports this closes out the Windows verification matrix: suite green on both platforms and both privilege shapes, three backends live, both Python install shapes, Unicode/MAX_PATH/timeout/concurrency edges exercised for real.

@TON14
TON14 marked this pull request as ready for review August 20, 2026 17:53
TON14 added a commit to TON14/LongHorizon-Harness that referenced this pull request Aug 20, 2026
…ands

The windows-latest lanes exercise platform support this branch does not
carry: it is based on a main whose supervisor still calls os.killpg and
whose agent stubs are #!/bin/sh scripts, so those lanes fail on known
pre-existing breakage rather than on anything in this change. AMAP-ML#57 brings
the Windows support together with the full two-platform matrix; when it
merges, its version of this workflow supersedes this one.
TON14 added a commit to TON14/LongHorizon-Harness that referenced this pull request Aug 20, 2026
…ands

The windows-latest lanes exercise platform support this branch does not
carry: it is based on a main whose supervisor still calls os.killpg and
whose agent stubs are #!/bin/sh scripts, so those lanes fail on known
pre-existing breakage rather than on anything in this change. AMAP-ML#57 brings
the Windows support together with the full two-platform matrix; when it
merges, its version of this workflow supersedes this one.
@TON14 TON14 closed this Sep 16, 2026
@TON14
TON14 deleted the support-windows branch September 16, 2026 16:08
TON14 added a commit to TON14/LongHorizon-Harness that referenced this pull request Sep 17, 2026
…ands

The windows-latest lanes exercise platform support this branch does not
carry: it is based on a main whose supervisor still calls os.killpg and
whose agent stubs are #!/bin/sh scripts, so those lanes fail on known
pre-existing breakage rather than on anything in this change. AMAP-ML#57 brings
the Windows support together with the full two-platform matrix; when it
merges, its version of this workflow supersedes this one.
TON14 added a commit to TON14/LongHorizon-Harness that referenced this pull request Sep 17, 2026
…ands

The windows-latest lanes exercise platform support this branch does not
carry: it is based on a main whose supervisor still calls os.killpg and
whose agent stubs are #!/bin/sh scripts, so those lanes fail on known
pre-existing breakage rather than on anything in this change. AMAP-ML#57 brings
the Windows support together with the full two-platform matrix; when it
merges, its version of this workflow supersedes this one.
TON14 added a commit to TON14/LongHorizon-Harness that referenced this pull request Sep 17, 2026
…ands

The windows-latest lanes exercise platform support this branch does not
carry: it is based on a main whose supervisor still calls os.killpg and
whose agent stubs are #!/bin/sh scripts, so those lanes fail on known
pre-existing breakage rather than on anything in this change. AMAP-ML#57 brings
the Windows support together with the full two-platform matrix; when it
merges, its version of this workflow supersedes this one.
@TON14

TON14 commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

This PR was auto-closed when I cleaned up the head branches in my fork. The Windows work lives on and is even exercised by CI (windows-latest in the matrix) in my fork's main: https://github.com/TON14/LongHorizon-Harness. Happy to rebase and reopen if upstream activity resumes — thanks for the project!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant