Repository navigation
feat: continuous diagnostics and verified autonomy repairs - #12
Conversation
There was a problem hiding this comment.
Validator approval after policy checks for exact head 73b747cf8d474a21c504fa1757729f04c51e32d9.
Ticket: ticket-006
Correlation ID: curllm006-autonomy-20261006
Model: openai/cursor-auto
Reviewed diff chunks: 3
Advisory LLM verdict: APPROVE
Advisory summary: Reviewed all 3 diff chunk(s). The autonomy configuration and process execution mechanisms are rigorously designed. The code explicitly enforces bounded process execution limits via RLIMIT_FSIZE and timeout bounds, strictly types and sanitizes operator configurations, and correctly implements robust cleanup of POSIX process groups to eliminate orphaned descendant processes. Path-handling defensively rejects symlinked elements to avoid resolution-based authority bypasses, and file locking safely serializes state. | The PR implements a continuous bounded monitoring runtime with SQLite incident tracking, verified probes, and an operator CLI. The code properly uses parameterized SQL queries to prevent injection and sets a restrictive umask for the CLI. All required tests and governance checks have passed. | Tests for the autonomy module appear comprehensive, covering recovery requirements, cooldowns, interrupted repair states, output redaction, descendant process cleanup, and state validation.
Advisory findings: none
The LLM output above is advisory and was not used as the approval trust root.
Semantic review prerequisite: not_required; policy 676cb4516bbfed2a000e40b9b1b6e4a430ecc761ec546aeb53d721a1905cfdd7.
Actual PR impact radar
Exact range: fe6c772c9b80e51a99eb27f4a391cfb209d1bf0b...73b747cf8d474a21c504fa1757729f04c51e32d9
Change digest: a7b9630414a736ff03daaa146271527a127d87cbac056df333694dff7d028da5
Score: 72/100 (L), estimated 102 min, split recommended: true
Affected services/components: repository-wide/unclassified
Machine-readable radar JSONL and SVG
{"actual_change":{"additions":717,"base_sha":"fe6c772c9b80e51a99eb27f4a391cfb209d1bf0b","binary_files":0,"categories":{"code":2,"configuration":1,"docs":1,"tests":1},"change_digest":"a7b9630414a736ff03daaa146271527a127d87cbac056df333694dff7d028da5","comparison":"fe6c772c9b80e51a99eb27f4a391cfb209d1bf0b...73b747cf8d474a21c504fa1757729f04c51e32d9","deletions":0,"file_count":5,"files":["curllm_core/autonomy.py","curllm_core/cli/autonomy.py","project/ticket-006/README.md","project/ticket-006/intent.json","tests/test_autonomy.py"],"head_sha":"73b747cf8d474a21c504fa1757729f04c51e32d9","service_count":0,"services":[]},"assessment_mode":"observed-pr","axes":{"coupling":5,"delivery":3,"scope":5,"uncertainty":3,"validation":2},"complexity":"L","confidence":0.9,"diagnostics":["RADAR-ACCEPTANCE-MISSING","RADAR-BUDGET-EXCEEDED"],"estimate":{"budget_minutes":30,"minutes":102,"within_budget":false},"impact":{"components":["browser","curllm_core","errors","exit","failure","incident","project","tests"],"files":["browser/MCP","curllm_core/autonomy.py","curllm_core/cli/autonomy.py","errors/warnings","exit/JSON","failure/repair/reprobe","incident/outbox","project/ticket-006/README.md","project/ticket-006/intent.json","tests/test_autonomy.py"],"public_interfaces":[],"runtime_dependencies":0},"schema":"subactor.ticket-radar/v1","score":72,"split":{"parts":[{"estimated_minutes":12,"name":"Implement browser","scope":["browser"]},{"estimated_minutes":12,"name":"Implement curllm_core","scope":["curllm_core"]},{"estimated_minutes":12,"name":"Implement errors","scope":["errors"]},{"estimated_minutes":12,"name":"Implement exit","scope":["exit"]},{"estimated_minutes":12,"name":"Implement failure","scope":["failure"]},{"estimated_minutes":15,"name":"Validate and project to trackers","scope":["tests","planfile","github/gitlab/jira projections"]}],"reason":"estimated_minutes_exceed_budget","recommended":true},"standards":[{"id":"wellmanifest/dsl","revision":"6c60fc4e0dd1f1bb74f46a7745e28019908d1203","version":"0.1.0-dev"},{"id":"wellmanifest/ticket-lifecycle","revision":"5bf581907a87b46a13a73e6c033d3abe4d9a306f","version":"0.1.0-dev"},{"id":"wellmanifest/git-lifecycle","revision":"7d77d4b7af57e69bc75c3a0290b3a4805c5c4438","version":"0.2.0-dev"},{"id":"wellmanifest/logs","revision":"48c284ef7a069055c0bcb6b900147ce5e65f8b43","version":"0.3.0"}],"ticket_ref":"ticket-006"}<svg xmlns="http://www.w3.org/2000/svg" width="128" height="128" viewBox="0 0 128 128" role="img"><title>ticket-006: feat: continuous diagnostics and verified autonomy repairs</title><rect width="128" height="128" rx="12" fill="#f8fafc"/><g stroke-width="1"><polygon points="64,55 72,61 69,71 59,71 56,61" fill="none" stroke="#d7dde5"/><polygon points="64,47 80,59 74,78 54,78 48,59" fill="none" stroke="#d7dde5"/><polygon points="64,38 89,56 79,85 49,85 39,56" fill="none" stroke="#d7dde5"/><polygon points="64,30 97,53 84,92 44,92 31,53" fill="none" stroke="#d7dde5"/><polygon points="64,21 105,51 89,99 39,99 23,51" fill="none" stroke="#d7dde5"/><line x1="64" y1="64" x2="64" y2="21" stroke="#aab4c0"/><line x1="64" y1="64" x2="105" y2="51" stroke="#aab4c0"/><line x1="64" y1="64" x2="89" y2="99" stroke="#aab4c0"/><line x1="64" y1="64" x2="39" y2="99" stroke="#aab4c0"/><line x1="64" y1="64" x2="23" y2="51" stroke="#aab4c0"/></g><polygon points="64,21 105,51 79,85 54,78 39,56" fill="#fb923c" fill-opacity="0.45" stroke="#c2410c" stroke-width="2"/><circle cx="64" cy="64" r="3" fill="#c2410c"/><g font-family="sans-serif" font-size="7" fill="#334155"><text x="64" y="11" text-anchor="middle">SCO</text><text x="114" y="48" text-anchor="middle">COU</text><text x="95" y="107" text-anchor="middle">UNC</text><text x="33" y="107" text-anchor="middle">VAL</text><text x="14" y="48" text-anchor="middle">DEL</text></g><text x="64" y="124" text-anchor="middle" font-family="sans-serif" font-size="8" fill="#0f172a">L · 102m</text></svg>DECISION D-006-4470
TICKET ticket-006
HEAD_SHA 73b747cf8d474a21c504fa1757729f04c51e32d9
CORRELATION_ID curllm006-autonomy-20261006
ACTOR agent:ifuri-validator-agent[bot]
APPLIED_RULE P-CORE-015
INPUT author_login = "tom-sapletta-com"
INPUT observed_checks = ["metadata=PASS","governance / remote lifecycle=PASS","governance / enforce=PASS"]
INPUT required_checks = ["metadata","governance / enforce","governance / remote lifecycle"]
INPUT required_checks_source = "protected registry (env/request)"
INPUT reviewer_login = "ifuri-validator-agent[bot]"
INPUT semantic_review_assessment = {"schema":"subactor.validator/semantic-review-assessment/v1","subject":{"repository":"autogrammar/curllm","pull_request":12,"head_sha":"73b747cf8d474a21c504fa1757729f04c51e32d9","base_sha":"fe6c772c9b80e51a99eb27f4a391cfb209d1bf0b","diff_sha256":"5e990c170d74e6e8576adce763258266147d6ba126e9766169524af213823412"},"policy":{"policy_schema":"subactor.validator/semantic-review-policy/v1","policy_version":1,"policy_sha256":"676cb4516bbfed2a000e40b9b1b6e4a430ecc761ec546aeb53d721a1905cfdd7","required":false,"critical_paths":[],"observed_paths":["curllm_core/autonomy.py","curllm_core/cli/autonomy.py","project/ticket-006/README.md","project/ticket-006/intent.json","tests/test_autonomy.py"]},"grounding":"full-diff-not-per-finding-proof","execution_authority":false,"status":"not_required","reason":null,"review_sha256":null,"unresolved":[]}
INPUT superseded_checks = []
INPUT ticket_radar_receipt = {"schema":"subactor.ticket-radar/v1","base_sha":"fe6c772c9b80e51a99eb27f4a391cfb209d1bf0b","head_sha":"73b747cf8d474a21c504fa1757729f04c51e32d9","change_digest":"a7b9630414a736ff03daaa146271527a127d87cbac056df333694dff7d028da5","score":72,"complexity":"L","estimated_minutes":102,"split_recommended":true,"services":[],"authority":"ADVISORY","promotion":"FORBIDDEN"}
VERDICT APPROVE AUTHORITY DETERMINISTIC
REJECTED REQUEST_CHANGES BECAUSE NO_UNSAFE_CHANGE_REASON_FOUND
ADVISORY llm_verdict = "APPROVE" MODEL "openai/cursor-auto"
ASSERT VERDICT_AUTHORITY != "ADVISORY"
Curllm can continuously detect failures without losing incidents or treating successful command execution as proof of recovery. Add serialized SQLite incident/outbox state, bounded argv probes with explicit exit/JSON oracles, and operator-authorized repair with durable attempt reservation and post-repair verification. Interrupted effects remain unknown and are never replayed.
Development findings enter a separate deduplicated Planfile intake marked as awaiting protected controller admission. The runtime grants no source lease or merge approval. Operator CLI emits structured status and DSL events with output digests instead of raw diagnostics.
Validation: 26 autonomy tests and 28 existing browser/MCP regressions passed; broad suite 447 passed, 1 skipped, 1 external HTTP test deselected. Actual fixture failure/repair/reprobe and Planfile intake tested. Native governance passed with zero errors/warnings. Deployment follows in a serialized installation slice.