Skip to content

feat(broker): add [broker] max_open_files and explain descriptor exha… - #14

Merged
pbosetti merged 1 commit into
v2from
claude/broker-client-limit-r9bzea
Aug 28, 2026
Merged

pbosetti merged 1 commit into
v2from
claude/broker-client-limit-r9bzea

Conversation

@pbosetti

Copy link
Copy Markdown
Owner

…ustion

Fleet size is bounded by the broker's open-file limit, not by anything in libzmq. Every connected agent holds two of the broker's descriptors for as long as it stays connected -- its publisher on the XSUB frontend, its subscriber on the XPUB backend -- plus a third while it fetches its settings. With the usual soft RLIMIT_NOFILE of 1024 and ~30 descriptors of broker overhead, that caps a fleet at roughly 495 agents.

Past that ceiling the broker failed in the least helpful way possible. libzmq's tcp_listener_t::accept() counts EMFILE/ENFILE among the errnos it tolerates (src/tcp_listener.cpp:193-199 in the pinned v4.3.5), so it raised ZMQ_EVENT_ACCEPT_FAILED -- which nothing listened for -- and refused every new agent in silence, while the still-readable listener re-fired the failing accept as fast as it could poll. The only thing that printed was the service-discovery thread, whose once-a-second getifaddrs() and broadcast sockets are the broker's sole timed descriptor allocations and so failed first. The result was a broker that looked healthy while turning agents away, reporting "getifaddrs failed: Too many open files" -- pointing at discovery rather than at the limit actually responsible.

Three changes, all backwards-compatible:

  • New src/detail/fd_limit.hpp resolves, plans and applies the limit. The decision logic is pure -- it takes the observed limits as data and makes no syscalls -- so every branch, including the platforms a given build is not running on, is unit-testable. Only rlim_cur is ever raised, since raising rlim_max needs CAP_SYS_RESOURCE; the target is capped by kern.maxfilesperproc on macOS and fs.nr_open on Linux, both of which setrlimit rejects values above even when rlim_max is RLIM_INFINITY. Windows reports the concept as not applicable rather than pretending: libzmq uses SOCKET handles there, and _setmaxstdio() governs only stdio streams.

  • [broker] max_open_files raises the soft limit at startup, before anything binds. Unset (the default) leaves it untouched and only reports it, warning when raising would help; 0 asks for as much as the process is permitted; a value above the hard limit is clamped with a warning, and a negative one is ignored with a warning rather than refused, matching io_threads' posture.

  • The broker now watches its bound sockets for ZMQ_EVENT_ACCEPT_FAILED and reports a refused agent explicitly -- endpoint, limit in force, the agent count it allows, and how to raise it -- rate-limited, and latched per failure class so a routine ECONNABORTED cannot consume the one-shot descriptor-exhaustion explanation. SocketMonitor gains LinkEvent:: AcceptFailed, LinkState::last_event_value (the errno libzmq reports) and a monotonic accept_failures counter; an accept failure deliberately does not move LinkStatus, since it describes the listener rather than any link.

  • Any discovery socket error caused by a full descriptor table now says so, and the advertising loop reports an unchanged failure at most once every 30 seconds instead of every second.

The shipped systemd template now sets LimitNOFILE=65536; without it a unit inherits systemd's DefaultLimitNOFILE, which is that same 1024 on most distributions. mads doctor gains a check reporting the limit in force and whether it leaves room for the fleet.

…ustion

Fleet size is bounded by the broker's open-file limit, not by anything in
libzmq. Every connected agent holds two of the broker's descriptors for as
long as it stays connected -- its publisher on the XSUB frontend, its
subscriber on the XPUB backend -- plus a third while it fetches its settings.
With the usual soft RLIMIT_NOFILE of 1024 and ~30 descriptors of broker
overhead, that caps a fleet at roughly 495 agents.

Past that ceiling the broker failed in the least helpful way possible.
libzmq's tcp_listener_t::accept() counts EMFILE/ENFILE among the errnos it
tolerates (src/tcp_listener.cpp:193-199 in the pinned v4.3.5), so it raised
ZMQ_EVENT_ACCEPT_FAILED -- which nothing listened for -- and refused every
new agent in silence, while the still-readable listener re-fired the failing
accept as fast as it could poll. The only thing that printed was the
service-discovery thread, whose once-a-second getifaddrs() and broadcast
sockets are the broker's sole timed descriptor allocations and so failed
first. The result was a broker that looked healthy while turning agents
away, reporting "getifaddrs failed: Too many open files" -- pointing at
discovery rather than at the limit actually responsible.

Three changes, all backwards-compatible:

- New src/detail/fd_limit.hpp resolves, plans and applies the limit. The
  decision logic is pure -- it takes the observed limits as data and makes no
  syscalls -- so every branch, including the platforms a given build is not
  running on, is unit-testable. Only rlim_cur is ever raised, since raising
  rlim_max needs CAP_SYS_RESOURCE; the target is capped by
  kern.maxfilesperproc on macOS and fs.nr_open on Linux, both of which
  setrlimit rejects values above even when rlim_max is RLIM_INFINITY.
  Windows reports the concept as not applicable rather than pretending:
  libzmq uses SOCKET handles there, and _setmaxstdio() governs only stdio
  streams.

- [broker] max_open_files raises the soft limit at startup, before anything
  binds. Unset (the default) leaves it untouched and only reports it, warning
  when raising would help; 0 asks for as much as the process is permitted; a
  value above the hard limit is clamped with a warning, and a negative one is
  ignored with a warning rather than refused, matching io_threads' posture.

- The broker now watches its bound sockets for ZMQ_EVENT_ACCEPT_FAILED and
  reports a refused agent explicitly -- endpoint, limit in force, the agent
  count it allows, and how to raise it -- rate-limited, and latched per
  failure class so a routine ECONNABORTED cannot consume the one-shot
  descriptor-exhaustion explanation. SocketMonitor gains LinkEvent::
  AcceptFailed, LinkState::last_event_value (the errno libzmq reports) and a
  monotonic accept_failures counter; an accept failure deliberately does not
  move LinkStatus, since it describes the listener rather than any link.

- Any discovery socket error caused by a full descriptor table now says so,
  and the advertising loop reports an unchanged failure at most once every 30
  seconds instead of every second.

The shipped systemd template now sets LimitNOFILE=65536; without it a unit
inherits systemd's DefaultLimitNOFILE, which is that same 1024 on most
distributions. `mads doctor` gains a check reporting the limit in force and
whether it leaves room for the fleet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TpZK4EKXZooaAXr1ASFHwo
@pbosetti
pbosetti merged commit c534dbb into v2 Aug 28, 2026
2 checks passed
@codecov

codecov Bot commented Aug 28, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 90.71429% with 13 lines in your changes missing coverage. Please review.
✅ Project coverage is 85.3%. Comparing base (b279d67) to head (14d7531).
⚠️ Report is 6 commits behind head on v2.

Files with missing lines Patch % Lines
src/doctor_checks.cpp 86.1% 5 Missing ⚠️
src/socket_monitor.cpp 28.5% 5 Missing ⚠️
src/detail/fd_limit.hpp 96.8% 3 Missing ⚠️
Additional details and impacted files
@@           Coverage Diff           @@
##              v2     #14     +/-   ##
=======================================
+ Coverage   85.1%   85.3%   +0.2%     
=======================================
  Files         39      42      +3     
  Lines       4537    4753    +216     
=======================================
+ Hits        3863    4059    +196     
- Misses       674     694     +20     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants