Skip to content

feat: next-gen - #89

Open
definfo wants to merge 15 commits into
masterfrom
next-gen
Open

definfo wants to merge 15 commits into
masterfrom
next-gen

Conversation

@definfo

@definfo definfo commented Sep 13, 2026

Copy link
Copy Markdown

No description provided.

- prepareLogger: JSONFormatter to stdout (12-factor XI, event stream);
  logstash hook Fatal -> Warn so a dead sink never takes down the scheduler
- executorInvokeWorker.RunSync: stamp a short sync_id per sync attempt and
  thread it through every log line + into executor.RunOnce, so all lines of
  one run correlate in the stream
- rename logrus field "worker" -> "repo" on the two worker loggers
  (executorInvokeWorker, ExternalWorker); Prometheus label "worker" and
  manager's "manager" field unchanged for metric/API compatibility
Pin go-version, refresh flake.lock/gomod2nix and module deps.
Replace RLIMIT_AS pre/post-hook utility with ephemeral cgroup v2 job isolation.

Each sync attempt now runs in its own cgroup:
memory.max is enforced with swap disabled and oom.group set, so a job exceeding
its budget is killed atomically instead of silently spilling to swap.
The child enters the cgroup atomically via clone3 (CgroupFD).

Per-attempt telemetry (duration, memory.peak, oom_kill) is collected from
the cgroup and accumulated into a per-repo EWMA, exposed via new Prometheus
gauges (lug_job_*).
Expose POST /lug/v1/admin/worker/:name/abort. Cancellation propagates
through the worker context and kills the complete cgroup process tree.
Gate queued launches on memory PSI and learned per-repo peak-memory EWMA,
while retaining concurrent_limit as a hard cap.
Grants reserve estimated memory until workers return idle
…rence

Document all REST endpoints including the new abort endpoint, Prometheus
metrics (sync counters, resource telemetry gauges, admission verdicts),
structured log events, cgroup v2 requirements for rootful Docker / rootless
Podman / bare systemd, elastic admission control, per-repo resource controls,
development workflow with cgroup-aware test execution, and the manual sync
triggering workaround.
…etails

New REST endpoints:

  GET  /lug/v1/admin/queue             — running workers (with cgroup/PID
       info) and the pending launch queue in FIFO order.

  POST /lug/v1/admin/worker/:name/sync — insert/lift the named worker to the
       head of the pending queue and launch immediately if capacity allows.
       Returns 409 if the worker is already syncing (abort first).

  GET  /lug/v1/admin/worker/:name/job  — live job detail: cgroup path, main
       PID, start time, and real-time cgroup stats (memory.current,
       memory.peak, memory.max, pids.current). Includes an nsenter attach
       hint for interactive debugging.

Implementation: queue inspection and manual-sync requests are routed through
dedicated channels into the manager Run() goroutine, so pending-queue access
is race-free without adding a mutex. Active job tracking is published from
shellScriptExecutor (mutex-protected) via the new jobInspector interface and
surfaced in worker.Status.ActiveJob.

Non-Linux: ReadCgroupStats and jobCgroup.Path stubs compile on all platforms;
ActiveJob.CgroupPath is empty when cgroup isolation is unavailable.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant