Skip to content

Grant executor capabilities in upgrade-database.yml; fall back when the eBPF counter load fails - #427

Merged
vincent10400094 merged 2 commits into
mainfrom
fix/upgrade-db-executor-capabilities
Oct 8, 2026
Merged

vincent10400094 merged 2 commits into
mainfrom
fix/upgrade-db-executor-capabilities

Conversation

@vincent10400094

Copy link
Copy Markdown
Member

Problem

deploy/ansible/upgrade-database.yml activates the candidate executor and starts it, but never grants it the file capabilities that the executor role sets (executor_capabilities). Two production executors then failed to create their eBPF maps. In packet_counter = "auto" mode they exited with initialize packet counter: failed to initialize packet count: packet counter cleanup unconfirmed instead of falling back, and systemd restarted them about 9,000 times.

Cause

  1. Deployment. The getcap/setcap logic existed in two copies, in the executor role and in rollout-executors.yml. upgrade-database.yml had no copy. The two copies had also drifted: the rollout always ran setcap and skipped it under executor_ambient_caps, while the role compared against getcap first and ignored executor_ambient_caps.
  2. Executor. newBPFCount let only EPERM, EACCES and EINVAL load failures fall back. Every other load failure was tagged ErrCleanupUnconfirmed, which makes the factory stop the executor. Without CAP_BPF, cilium/ebpf on kernel 7.0 reports the map create as prealloc maps not supported (requires >= v4.6). That message comes from cilium's own feature probe, which was also refused, and it carries no errno. So the load failure became fatal. Yet a load attaches nothing: attach runs only after a successful load.

Fix

Deployment (deploy/ansible)

  • New tasks/executor-capabilities.yml holds the one copy of the logic. It runs getcap (also in check mode), then runs setcap for exactly executor_capabilities only when the set differs. It does nothing when executor_enable_bpf is false.
  • Three callers import it:
    • Executor role: still warns on a refused setcap, and still notifies Restart executor when the set changes.
    • rollout-executors.yml: still fails before the drain.
    • upgrade-database.yml: new. It grants the set to the staged candidate before the service is stopped, and a refusal fails the host while the service and the database are still untouched.
  • One behaviour change to note: under executor_ambient_caps=true the rollout now grants the file capabilities. Before this PR it skipped them. The role and verify.yml already required them in that case, so the rollout was the one that had drifted.
  • deploy/README.md and the upgrade-database header comment are updated.

Executor (internal/executor/ratelimit/ebpf)

  • Every load failure now counts as a clean rollback, so auto mode falls back to the userspace counter. Loading creates no TCX link and, without pin options, nothing that outlives the process. Any descriptors cilium/ebpf may not have released count no packet and close when the process exits.
  • The safety property is unchanged. A failed attach whose release fails still carries ErrCleanupFailed and stops the executor.
  • When the process lacks CAP_BPF or CAP_NET_ADMIN (and CAP_SYS_ADMIN), the load error now names the missing capabilities and points to setcap. It matches os.ErrPermission, so the reported fallback reason is not_permitted. The new error replaces the old text at the front, for example executor lacks CAP_BPF, CAP_NET_ADMIN (grant the executor binary executor_capabilities with setcap); loader reported: ... MEMLOCK may be too low ....
  • The doc comment on ErrCleanupUnconfirmed now says that a load failure does not carry it.

Tests

  • deploy/test/provisioner-check.sh passes in full: render checks, then 10/10 tests in test_rollout.py. New tests:
    • upgrade-database.yml runs setcap with the exact set on the staged candidate, before stop, upgrade and start, on each host.
    • It skips setcap when getcap already reports the set (in a different order).
    • It grants nothing with executor_enable_bpf=false.
    • A refused setcap fails before the service is stopped, leaves the database unchanged, and leaves the second host untouched.
    • The rollout grants the set before stop and start on each host.
  • I also checked separately that notify on import_tasks reaches the role's handler.
  • Go:
    • go build / go vet for ./internal/... ./cmd/... pass on darwin and with GOOS=linux.
    • New unit tests:
      • every load failure, including the prealloc maps not supported shape, falls back without attaching;
      • a failed attach release stays fatal;
      • the factory falls back with the right reason, and stays fatal for ErrCleanupFailed and ErrCleanupUnconfirmed;
      • missing-capability attribution, including CAP_SYS_ADMIN covering both and CAP_PERFMON not required.
    • The ratelimit, ebpf and cleanup tests pass on Linux in a container.
    • In an unprivileged container the real load now reports executor lacks CAP_BPF, CAP_NET_ADMIN .... In a privileged one it loads.

Fixes #414

🤖 Generated with Claude Code

vincent10400094 and others added 2 commits October 7, 2026 17:12
upgrade-database.yml activated and started the candidate executor
without the file capabilities the executor role grants, so until a full
deploy followed the executor ran without CAP_BPF, CAP_PERFMON,
CAP_NET_ADMIN and CAP_NET_RAW.

Move the getcap/setcap logic into tasks/executor-capabilities.yml and
import it from the executor role, rollout-executors.yml and
upgrade-database.yml so the three cannot drift. upgrade-database.yml
grants the set to the staged candidate before stopping the service; a
refusal stops there with the service and database untouched, as in the
rollout. The role keeps warning instead of failing, and still restarts
the executor when the set changes. The rollout now also skips setcap
when the binary already carries the set, and grants it under
executor_ambient_caps as the role and verify.yml already expected.

test_rollout.py runs upgrade-database.yml against the fixture hosts and
checks that setcap precedes stop/upgrade/start, is skipped when the set
is present or BPF is disabled, and that a refusal stops before the
service is touched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
In packet_counter = "auto" mode a load failure other than EPERM, EACCES
or EINVAL was tagged ErrCleanupUnconfirmed, so the executor exited with
"packet counter cleanup unconfirmed" and systemd restarted it in a loop.
Without CAP_BPF, cilium/ebpf can report the map create as "prealloc maps
not supported" (its refused feature probe), which carries no errno.

A load attaches nothing: attach runs only after a successful load, and
loading creates no TCX link or pin. Whatever descriptors cilium/ebpf may
not have released count no packet and close with the process, so every
load failure is now a clean rollback and the factory falls back. A
failed attach whose release fails still refuses to fall back.

When the process lacks CAP_BPF or CAP_NET_ADMIN (and CAP_SYS_ADMIN), the
load error now names the missing capabilities and matches
os.ErrPermission, so the fallback reason is not_permitted rather than a
MEMLOCK or prealloc hint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vincent10400094
vincent10400094 merged commit 453b1e7 into main Oct 8, 2026
17 checks passed
@vincent10400094
vincent10400094 deleted the fix/upgrade-db-executor-capabilities branch October 8, 2026 00:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

upgrade-database.yml starts the new executor without file capabilities; executor crash-loops instead of falling back

1 participant