What problem would this solve?
A module runs only when something happens: startup, an [[events]] hook, or an action someone clicks. That is the right default, but it leaves no way to notice that nothing is happening.
Concrete case: agent-board, a scheduler module that starts and merges task lanes from task.* events. Its worst failures are boards that stop: work is ready, lane slots are free, and no event will ever come, because the thing that should have fired didn't. We had five of these in one day. Each looked like a busy board that had gone quiet, and a person noticed each one 5 to 15 minutes later. The module can say what's wrong in one short run, if something runs it.
Today the workarounds all sit outside Luvus:
- a systemd user timer per session running the module's CLI (what we ship now,
agent-board install-module --stall-timer 5m);
- a cron line;
- a Luvus automation, which starts an agent for every run, far too heavy for a check this small.
What we'd like
A declared timer in the manifest, run like an event hook:
[[timers]]
id = "stall-check"
every = "5m"
command = ["agent-board", "stall", "--notify"]
It would get the same environment as an event hook (LUVUS_SESSION, LUVUS_BIN_PATH, the module context) and log to module log the same way. A tick that finds the previous run still going would be skipped. It would run per session, since a module is linked into each session's server.
Relation to #182
Resident services (#182) would cover this too, since a resident process can keep its own clock. A declared timer is smaller: no long-lived process, no stdin protocol, and nothing pins a Running slot. It fits modules that need a heartbeat and keep no state. If #182 lands first, we'd use it.
Alternatives considered
The ones above. Binding the check to a frequent event such as pane.agent_status_changed doesn't work, because a stalled board fires none, and it costs a process per status change on every board.
What problem would this solve?
A module runs only when something happens: startup, an
[[events]]hook, or an action someone clicks. That is the right default, but it leaves no way to notice that nothing is happening.Concrete case: agent-board, a scheduler module that starts and merges task lanes from
task.*events. Its worst failures are boards that stop: work is ready, lane slots are free, and no event will ever come, because the thing that should have fired didn't. We had five of these in one day. Each looked like a busy board that had gone quiet, and a person noticed each one 5 to 15 minutes later. The module can say what's wrong in one short run, if something runs it.Today the workarounds all sit outside Luvus:
agent-board install-module --stall-timer 5m);What we'd like
A declared timer in the manifest, run like an event hook:
It would get the same environment as an event hook (
LUVUS_SESSION,LUVUS_BIN_PATH, the module context) and log tomodule logthe same way. A tick that finds the previous run still going would be skipped. It would run per session, since a module is linked into each session's server.Relation to #182
Resident services (#182) would cover this too, since a resident process can keep its own clock. A declared timer is smaller: no long-lived process, no stdin protocol, and nothing pins a
Runningslot. It fits modules that need a heartbeat and keep no state. If #182 lands first, we'd use it.Alternatives considered
The ones above. Binding the check to a frequent event such as
pane.agent_status_changeddoesn't work, because a stalled board fires none, and it costs a process per status change on every board.