The gap
Both guardrail ceilings are enforced by counting rows and then writing, with nothing holding the count still in between.
_check_enforced_limit in src/gateway/services/guardrail_credential_service.py counts the enabled stored definitions and refuses an eleventh.
_check_guardrail_limit in src/gateway/services/tenancy/organization_guardrail_service.py does the same for MAX_GUARDRAILS_PER_ORGANIZATION.
Two writes that enable a definition at the same time both pass the count and both commit, so a deployment can end up above the ceiling. Nothing repairs it afterwards: the count is only consulted on the next write, which is then refused, so the extra rows stay enabled until an operator disables one.
Why it was left
Raised while reviewing #1296 and deliberately not fixed there. Two reasons:
- What it bounds is latency, not money. Each enabled definition is one more sequential check in front of a request. Eleven instead of ten is slower, not incorrect, and no budget or tenant boundary is crossed. A reservation race would have been fixed in place.
- Fixing one of the two would split them. The two services use the same shape on purpose. Closing the race properly means a table-level lock, a serializable transaction, or a counter row, and whichever is chosen belongs in both at once rather than in the one that a guardrail PR happened to touch.
What a fix looks like
Pick one mechanism and apply it to both call sites. A SELECT ... FOR UPDATE over the enabled rows is the smallest option; an advisory lock keyed on the table avoids holding row locks across the write. Either way the count and the write end up in one transaction, and both services should get the same treatment and share a test that runs two concurrent enables.
Pointers
The gap
Both guardrail ceilings are enforced by counting rows and then writing, with nothing holding the count still in between.
_check_enforced_limitinsrc/gateway/services/guardrail_credential_service.pycounts the enabled stored definitions and refuses an eleventh._check_guardrail_limitinsrc/gateway/services/tenancy/organization_guardrail_service.pydoes the same forMAX_GUARDRAILS_PER_ORGANIZATION.Two writes that enable a definition at the same time both pass the count and both commit, so a deployment can end up above the ceiling. Nothing repairs it afterwards: the count is only consulted on the next write, which is then refused, so the extra rows stay enabled until an operator disables one.
Why it was left
Raised while reviewing #1296 and deliberately not fixed there. Two reasons:
What a fix looks like
Pick one mechanism and apply it to both call sites. A
SELECT ... FOR UPDATEover the enabled rows is the smallest option; an advisory lock keyed on the table avoids holding row locks across the write. Either way the count and the write end up in one transaction, and both services should get the same treatment and share a test that runs two concurrent enables.Pointers
_check_enforced_limit.MAX_ENFORCED_GUARDRAILS(guardrail credential service) andMAX_GUARDRAILS_PER_ORGANIZATION.