Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
91 changes: 62 additions & 29 deletions claude_sdk/agent_safety.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -62,33 +62,49 @@ rules:
descriptions before passing them down.

- id: CSDK-103
title: AgentDefinition sets permissionMode to bypassPermissions
title: AgentDefinition sets permissionMode to bypassPermissions with a broad tool set
severity: high
confidence: 0.9
language: python
applies_to:
- claude_agent_definition
scope: agent
match:
agent_kwarg_value:
kwarg: permissionMode
value: bypassPermissions
all:
- agent_kwarg_value:
kwarg: permissionMode
value: bypassPermissions
- any:
- agent_kwarg_missing:
- tools
- agent_grants_builtin_tool:
- Bash
- Write
- Edit
- NotebookEdit
- WebFetch
- WebSearch
- Agent
- Task
explanation: >
This AgentDefinition sets permissionMode="bypassPermissions", which
disables the SDK's interactive approval gate for every tool the subagent
runs. Subagents are dispatched autonomously by a lead agent in response to
model-generated task descriptions, so a prompt-injected or mis-scoped task
reaches the subagent's tools — including Bash, Write, and Edit — with no
human in the loop and no per-call confirmation. This is the single
highest-impact Claude-specific misconfiguration: it removes the one control
that stands between model output and real side effects.
runs, and either omits tools entirely (which inherits every tool
available to subagents) or lists a side-effecting or exfiltration-
capable one. Unlike allowed_tools on ClaudeAgentOptions, an
AgentDefinition's tools list genuinely restricts the subagent — a tool
left out isn't in its session at all — so this combination is the
actual highest-impact shape: subagents are dispatched autonomously by a
lead agent in response to model-generated task descriptions, so a
prompt-injected or mis-scoped task reaches Bash, Write, Edit, or similar
with no human in the loop and no per-call confirmation.
fix: >
Remove permissionMode="bypassPermissions" unless this subagent runs in a
fully sandboxed, non-interactive context you control (e.g. CI with no
secrets or network). If autonomous operation is required, pair it with a
PreToolUse hook that allowlists exactly the commands and paths the subagent
may touch, and prefer a restrictive mode ("default"/"acceptEdits") scoped
to a minimal tool set.
secrets or network). If autonomous operation is required, scope tools to
the minimal read-only set the subagent actually needs (e.g. Read, Grep,
Glob) so bypassPermissions has nothing side-effecting to bypass approval
for, and prefer a restrictive mode ("default"/"acceptEdits") otherwise.

- id: CSDK-104
title: Claude subagent is granted filesystem-write built-ins
Expand Down Expand Up @@ -144,33 +160,50 @@ rules:
validate the task descriptions the lead agent passes down.

- id: CSDK-120
title: TypeScript AgentDefinition sets permissionMode to bypassPermissions
title: TypeScript AgentDefinition sets permissionMode to bypassPermissions with a broad tool set
severity: high
confidence: 0.9
language: typescript
applies_to:
- claude_agent_definition
scope: agent
match:
agent_kwarg_value:
kwarg: permissionMode
value: bypassPermissions
all:
- agent_kwarg_value:
kwarg: permissionMode
value: bypassPermissions
- any:
- agent_kwarg_missing:
- tools
- agent_grants_builtin_tool:
- Bash
- Write
- Edit
- NotebookEdit
- WebFetch
- WebSearch
- Agent
- Task
explanation: >
This TypeScript AgentDefinition sets permissionMode: "bypassPermissions",
which disables the SDK's interactive approval gate for every tool the agent
runs. Agents are dispatched autonomously in response to model-generated task
descriptions, so a prompt-injected or mis-scoped task reaches the agent's
tools — including Bash, Write, and Edit — with no human in the loop and no
per-call confirmation. This is the single highest-impact Claude-specific
misconfiguration: it removes the one control that stands between model output
and real side effects.
which disables the SDK's interactive approval gate for every tool the
agent runs, and either omits tools entirely (which inherits every tool
available to subagents) or lists a side-effecting or exfiltration-
capable one. An AgentDefinition's tools field genuinely restricts the
agent — a tool left out isn't in its session at all — so this
combination is the actual highest-impact shape: agents are dispatched
autonomously in response to model-generated task descriptions, so a
prompt-injected or mis-scoped task reaches Bash, Write, Edit, or similar
with no human in the loop and no per-call confirmation.
fix: >
Remove permissionMode: "bypassPermissions" unless this agent runs in a
fully sandboxed, non-interactive context you control (e.g. CI with no
secrets or network). If autonomous operation is required, pass
allowedTools or disallowedTools in the AgentDefinition constructor to
restrict the tool surface, and prefer a safe default permission mode
("default" or "acceptEdits") scoped to a minimal tool set.
secrets or network). If autonomous operation is required, scope tools
to the minimal read-only set the agent actually needs (e.g. Read, Grep,
Glob) — or pass disallowedTools naming what it must never call — so
bypassPermissions has nothing side-effecting to bypass approval for,
and prefer a safe default permission mode ("default" or "acceptEdits")
otherwise.

- id: CSDK-121
title: TypeScript AgentDefinition is granted the Bash tool
Expand Down
62 changes: 54 additions & 8 deletions claude_sdk/repo.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -34,29 +34,75 @@ rules:
Reserve bypassPermissions for disposable sandboxes, never a shared repo.

- id: CSDK-202
title: Session permission mode bypasses approvals
title: Session permission mode bypasses approvals with no tool deny-list
severity: high
confidence: 0.9
applies_to:
- claude_sdk
scope: repo
match:
repo_claude_options_permission_mode_is:
- bypassPermissions
repo_claude_options_mode_without_kwarg:
modes:
- bypassPermissions
kwarg: disallowed_tools
explanation: >
A ClaudeAgentOptions(...) in this project sets permission_mode to
"bypassPermissions", which turns off Claude Code's approval prompts for
that session. Every tool the agent can call — file writes, shell
commands, network fetches — then runs with no human in the loop, so a
that session, and the same construction sets no disallowed_tools.
allowed_tools does not help here — it only auto-approves tools, it does
not restrict which ones can run, so an empty or narrow allowed_tools
list alongside bypassPermissions is not a mitigation. Every tool the
agent can call — file writes, shell commands, network fetches — then
runs with no human in the loop and nothing bounding the surface, so a
single prompt-injection or model error becomes an unguarded action. This
is the in-code, session-level form of the settings.json defaultMode
bypass (CSDK-201), and is where most applications actually enable it.
fix: >
Drop permission_mode="bypassPermissions" from the ClaudeAgentOptions(...)
call, or set it to "default" so tool calls prompt (or "acceptEdits" to
auto-approve only file edits while still gating shell and network access).
Reserve bypassPermissions for disposable sandboxes, never code that runs
on a developer's or user's machine.
auto-approve only file edits while still gating shell and network
access). If bypassPermissions is genuinely required, pass
disallowed_tools= naming at minimum shell execution and any tool that
reaches the network or credentials — disallowed_tools denies matching
calls in every permission mode, including bypassPermissions, so it is
the one control that still bounds the surface. Reserve bypassPermissions
for disposable sandboxes, never code that runs on a developer's or
user's machine.

- id: CSDK-206
title: Session bypasses approvals with a deny-list that still leaves a broad surface
severity: medium
confidence: 0.6
applies_to:
- claude_sdk
scope: repo
match:
all:
- repo_claude_options_permission_mode_is:
- bypassPermissions
- not:
repo_claude_options_mode_without_kwarg:
modes:
- bypassPermissions
kwarg: disallowed_tools
explanation: >
A ClaudeAgentOptions(...) in this project sets permission_mode to
"bypassPermissions" and the same construction sets disallowed_tools, so
the named tools are denied in every permission mode including
bypassPermissions — a real, SDK-enforced bound on the tool surface, not
just an auto-approve list. The residual risk is what disallowed_tools
does not name: allowed_tools only auto-approves and does not restrict,
so every tool not on the deny-list still runs with no human approval
step at all under bypassPermissions. A deny-list is allow-by-default —
it is only as good as the completeness of what it excludes.
fix: >
Review the disallowed_tools list against every tool this session can
reach and confirm it denies shell execution and anything that touches
the network or credentials, not just an obvious tool or two. If the
session's actual tool needs are narrow, prefer naming them explicitly
with allowed_tools alongside a non-bypass permission_mode ("default" or
"acceptEdits") instead of relying on bypassPermissions plus a deny-list
to cover everything else.

- id: CSDK-204
title: Claude Agent SDK session sets no explicit max_turns limit
Expand Down
137 changes: 86 additions & 51 deletions claude_skill/skill_quality_text.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -20,29 +20,34 @@ rules:
- claude_skill
scope: skill
match:
any:
- skill_name_has_text:
- crypto
- encrypt
- decrypt
- sign
- hash
- certificate
- signing
- cipher
- asymmetric
- symmetric
- skill_description_has_text:
- crypto
- encrypt
- decrypt
- sign
- hash
- certificate
- signing
- cipher
- asymmetric
- symmetric
skill_text_matches:
fields: [name, description]
terms:
- crypto
- cryptographic
- encrypt
- decrypt
- cipher
- certificate
- keypair
- private key
- sign
- signature
- hash
- hmac
- asymmetric
- symmetric
exclude_context:
- example
- for instance
- e.g.
- such as
- sample
- documented
- documentation
- see the
- refer to
- placeholder
explanation: >
This skill's name or description claims a cryptographic operation —
encrypting, decrypting, signing, hashing, or certificate handling — but
Expand All @@ -51,6 +56,9 @@ rules:
signature, or a hardcoded key all look like working code right up until
they fail under attack, and a skill invoked without those specifics gives
Claude no constraint to hold to when it writes the actual crypto code.
A mention that is clearly an example or a pointer to documentation (e.g.
"for instance, hash a string" or "see the crypto docs") does not count —
only a claim that the skill itself performs the operation does.
fix: >
State the exact primitive the skill uses (e.g. "AES-256-GCM via the
`cryptography` library", "Ed25519 signing"), where keys or certificates
Expand All @@ -66,36 +74,63 @@ rules:
- claude_skill
scope: skill
match:
any:
- skill_body_has_text:
- password
- secret
- token
- ssn
- credit card
- pii
- personal data
- sensitive
- confidential
- private key
- skill_description_has_text:
- password
- secret
- token
- ssn
- credit card
- pii
- personal data
- sensitive
- confidential
- private key
skill_text_matches:
fields: [description, body]
terms:
- password
- secret
- credential
- token
- api key
- ssn
- social security
- credit card
- pii
- personal data
- sensitive data
- confidential
- private key
require_context:
- read
- write
- store
- save
- send
- transmit
- collect
- process
- handle
- log
- extract
- parse
- redact
- encrypt
- upload
- post
- fetch
- retrieve
- access
exclude_context:
- lives in
- live in
- configured in
- see the
- documentation
- example
- for instance
- e.g.
explanation: >
This skill's body or description names a sensitive data class —
passwords, tokens, SSNs, credit card numbers, PII, or private keys —
without stating what it does with that data beyond naming it. Without an
explicit data-minimization boundary and a stated purpose, there is no way
to tell whether the skill only touches the fields the task actually
requires or whether it retains what it reads.
passwords, tokens, SSNs, credit card numbers, PII, or private keys — in
the same sentence as an operational verb (reads, stores, sends,
processes, ...), so it claims to actually handle that data rather than
merely mention it, without stating what it does with that data beyond
naming it. Without an explicit data-minimization boundary and a stated
purpose, there is no way to tell whether the skill only touches the
fields the task actually requires or whether it retains what it reads. A
sentence that only states where a credential already lives (e.g. "API
keys live in GitHub Actions secrets") does not count — only a claim that
the skill itself operates on the data does.
fix: >
Declare exactly which sensitive fields the skill reads, why it needs
them, and how long it keeps them (ideally: not at all, beyond the current
Expand Down
2 changes: 1 addition & 1 deletion manifest.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -11,4 +11,4 @@
#
# This file is metadata, not a rule: the engine's loader skips manifest.yaml
# when walking the pack for policy files.
schema_version: 17
schema_version: 18
Loading