make post asserts pipeline_idle on four apps. Two of them do not publish it, and one of those publishes a different gauge that answers a different question. So post is permanently red on dfe-transform-elastic, and has been reporting it as a failure to read rather than as a difference of convention.
What each app actually serves
Read off a live stack just now:
| app |
pipeline_idle |
pipeline_ready |
| dfe-transform-vrl |
yes (3 lines) |
absent |
| dfe-archiver |
yes (3 lines) |
absent |
| dfe-transform-elastic |
absent |
pipeline_ready 1 |
| dfe-receiver |
absent |
absent |
| dfe-loader |
absent |
absent |
IDLE_APPS in post.py:231 names archiver, fetcher, transform-elastic and transform-vrl. Two of the four hold the convention it asserts.
Neither side is obviously wrong
pipeline_idle is scalo's, not each app's -- scalo-rs/src/lifecycle/mod.rs:206 sets it from the lifecycle idle gate, and no consumer hand-rolls it. An app that arms that gate gets the gauge; one that does not, does not.
transform-elastic does not simply lack it. It pushes its own scaling signals on a five-second floor, and says why at src/service.rs:49-54:
/// A batch pushes them as it finishes too, so a busy partition signals at the
/// batch cadence and this is the floor for an idle one -- KEDA still sees the
/// lag on a topic that has stopped producing.
That is a deliberate approach to the same problem, not an omission.
Why I have not just widened the check
The obvious patch is to accept pipeline_ready where pipeline_idle is missing. That would be wrong, and it is worth saying why: ready is not inert. pipeline_ready 1 says the app can take work. pipeline_idle 1 says it currently has none. An app can be ready and busy, and the post check exists to catch a component that was handed a config and is chewing on something it should not be.
Substituting one for the other makes the run green by asking an easier question. The e2e suite's own docstring forbids exactly that move -- "An assertion is never softened to make a run green" -- and it applies here.
So this is a decision, not a patch
Which signal is the fleet's contract for "ready and inert"?
pipeline_idle everywhere. Every pipeline app arms scalo's idle gate, and post keeps its assertion as written. Cleanest, but it asks transform-elastic to take on a gate it deliberately did without, and that repo is another team's.
- Per-app, declared.
IDLE_APPS carries which gauge each app answers with, and post asserts the right one per app. Honest about the split, but it makes post the place the split is recorded, which is the wrong home for a fleet convention.
- A new signal that means inert, on both. Most work, and the only one that leaves a single question with a single answer.
I would take 1 if transform-elastic's team is willing, and 2 only as a stopgap with 1 on the tracker -- but the call is not mine to make in their repo.
Either way, post should stop reporting this as "cannot be read". It can be read; it says something else.
Not the whole story
receiver and loader publish neither gauge and are not in IDLE_APPS, which is probably right -- a listener and a continuous consumer are never meaningfully idle. Worth confirming rather than assuming, since it is the same question one layer along.
make postassertspipeline_idleon four apps. Two of them do not publish it, and one of those publishes a different gauge that answers a different question. So post is permanently red ondfe-transform-elastic, and has been reporting it as a failure to read rather than as a difference of convention.What each app actually serves
Read off a live stack just now:
pipeline_idlepipeline_readypipeline_ready 1IDLE_APPSinpost.py:231names archiver, fetcher, transform-elastic and transform-vrl. Two of the four hold the convention it asserts.Neither side is obviously wrong
pipeline_idleis scalo's, not each app's --scalo-rs/src/lifecycle/mod.rs:206sets it from the lifecycle idle gate, and no consumer hand-rolls it. An app that arms that gate gets the gauge; one that does not, does not.transform-elastic does not simply lack it. It pushes its own scaling signals on a five-second floor, and says why at
src/service.rs:49-54:That is a deliberate approach to the same problem, not an omission.
Why I have not just widened the check
The obvious patch is to accept
pipeline_readywherepipeline_idleis missing. That would be wrong, and it is worth saying why: ready is not inert.pipeline_ready 1says the app can take work.pipeline_idle 1says it currently has none. An app can be ready and busy, and the post check exists to catch a component that was handed a config and is chewing on something it should not be.Substituting one for the other makes the run green by asking an easier question. The e2e suite's own docstring forbids exactly that move -- "An assertion is never softened to make a run green" -- and it applies here.
So this is a decision, not a patch
Which signal is the fleet's contract for "ready and inert"?
pipeline_idleeverywhere. Every pipeline app arms scalo's idle gate, and post keeps its assertion as written. Cleanest, but it asks transform-elastic to take on a gate it deliberately did without, and that repo is another team's.IDLE_APPScarries which gauge each app answers with, and post asserts the right one per app. Honest about the split, but it makes post the place the split is recorded, which is the wrong home for a fleet convention.I would take 1 if transform-elastic's team is willing, and 2 only as a stopgap with 1 on the tracker -- but the call is not mine to make in their repo.
Either way, post should stop reporting this as "cannot be read". It can be read; it says something else.
Not the whole story
receiver and loader publish neither gauge and are not in
IDLE_APPS, which is probably right -- a listener and a continuous consumer are never meaningfully idle. Worth confirming rather than assuming, since it is the same question one layer along.