Conversation
The check at github_action_audit.py:263 compared the result field against
the exact string 'Prompt is too long'. The API returns that marker with its
token accounting appended, for example:
Prompt is too long · the request is ~1139713 tokens (limit 1000000) but
the maximum allowed is 1000000
Because the equality never held, PROMPT_TOO_LONG was never returned and the
retry-without-diff branch was unreachable. The run instead fell through to
_extract_security_findings, which found a 'result' key, failed to parse the
error text as JSON, and returned the empty structure with review_completed
set to False while run_security_audit reported success.
Match on the prefix instead. The str() keeps the guard total for malformed
payloads: the previous equality could never raise, and startswith on a
non-string would.
Adds three tests: the real message with token counts, the bare message as a
regression guard, and four unrelated errors that must not trigger the retry.
One of them contains the phrase without beginning with it, which is why the
check matches a prefix rather than a substring.
3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
SimpleClaudeRunner.run_security_auditdetects the prompt-too-long condition by comparing theresultfield against the exact stringPrompt is too long. The API returns that marker with its token accounting appended, so the comparison never holds,PROMPT_TOO_LONGis never returned, and the retry-without-diff branch is unreachable.Symptom
Verbatim from the
resultfield of a production GitHub Actions run:The scan reports success with zero findings. Nothing was analyzed.
Root cause
claudecode/github_action_audit.py:263The marker is a prefix of the message, not the whole of it, so the equality is never true.
PROMPT_TOO_LONG(l.264) is never returned, which makes the retry block at l.590-595 dead code.What happens instead, in order:
_extract_security_findings(l.274).'result' in claude_outputis true, so the error text is handed to the JSON parser.findings: [],review_completed: False.run_security_auditreturns success.Fix
One line. The
str(..., '')keeps the guard total for malformed payloads — the previous equality could never raise, andstartswithon a non-string would.Prefix rather than substring is deliberate:
The output prompt is too long to displaycontains the phrase but is a different condition and must not trigger the retry.Testing
Baseline before the change: 173 passed. After: 176 passed. Run as CI does, with
pytest claudecode.Three tests added to
claudecode/test_claude_runner.py— this path had no coverage at all before:Prompt is too longmessage, as a regression guardRate limit exceeded,Invalid API key,overloaded_error, andThe output prompt is too long to displayDemonstrated in red. Reverting the one-line fix while keeping the tests fails the first one, and the failure output shows the whole chain:
assert True is Falseis step 6; the log line above it is steps 3 and 4. The other two tests pass with and without the fix — they are regression guards, not the discriminator.Relationship to #82 and #80
#82 takes a more complete approach to the 406 than anything here, and this PR does not touch that path — no overlap in the diffs. The two are complementary, and I think #82 needs this line to fully close #80.
With #82 merged and this line unchanged, the 406 stops killing the process but returns as the same silent zero through a different door:
_extract_security_findings:'result' in claude_outputis true, the JSON parse of the error message fails, and the empty structure is returned withfindings: []andreview_completed: False.run_security_auditreturns success.action.ymldecides onfindings_countand never readsreview_completed.The result is a green check with zero findings and no analysis performed — the symptom #80 reports.
Marked
Refs #80rather thanFixes #80, since closing that issue takes #82 as well.Separate observation, not part of this PR
action.ymlgates onfindings_countand contains no reference toreview_completed(zero matches in the file). Since_extract_security_findingssetsreview_completed: Falseon every failure path while still returning a zero count, any failure that yields zero findings is indistinguishable from a clean scan at the action level.This PR fixes one route into that state. The general property — that the signal for "nothing was reviewed" exists in the payload but is never consumed by the action — seemed worth flagging separately for maintainers to weigh.
https://claude.ai/code/session_01VQyxybenLFB1GryQQ7XSmw