Skip to content

Run a workspace's stored guardrail definitions on every request #1287

Description

@dpoulopoulos

The gap

Three stacked PRs built a guardrail Otari runs itself, and stopped one step short of using it.

An operator can define a guardrail, store its credential, switch it on and test it. None of that changes what happens to a request. docs/guardrails.md says so outright:

Nothing on the request path reads these rows yet, so storing a definition still does not change how a request behaves.

The proof in code: GuardrailRunner.check() (src/gateway/services/guardrail_runner.py) has zero callers. Everything on the request path still goes over HTTP to the sidecar container at guardrails_url.

What this adds

A definition that is stored, enabled and scoped to a workspace checks the input of every inference request from that workspace, before the provider is called. The caller sends nothing; there is no guardrails field to fill in.

The design splits one question into two, and keeping them apart is what keeps the change small:

Question Answered by
Which guardrails apply to this request? the database, scoped on the request's workspace
How is one run? GuardrailRunner, which already built them all at startup

So GuardrailRunner needs no change at all. knows() and check() already exist and already do the right thing. Nothing is built while a request waits.

Schema

Three columns on guardrail_credentials, and one scope table.

  • mode: block or monitor. What to do when the guardrail flags the input.
  • on_unavailable: block or allow. What Otari does when no verdict came back at all.
  • applies_to_all_workspaces: a flag, with guardrail_credential_workspaces holding the explicit list otherwise. Both copy organization_guardrails and organization_guardrail_workspaces column for column.

A guardrail is unavailable when the vendor API failed or timed out, or the answer was malformed. An explicit inconclusive verdict (valid: null) is not unavailable and never blocks.

The permissive value is spelled allow rather than the legacy monitor used by the request-body and organization fields, because a guardrail that never answered monitors nothing, and the decision is Otari's rather than the guardrail's. One mapping turns allow into monitor where a GuardrailConfig is built; the published request-body contract is not renamed.

Request path

The workspace's definitions become one more mandate layer on the existing merge, applied after the caller, the organization and the routing policy. That reuses the union by profile with stricter-wins, so a caller who also names the profile is deduped rather than checked twice and cannot weaken it. Every refusal shape is the existing one: 403 guardrail_violation, 502 fail closed, and the X-Otari-Guardrails header for monitor mode.

run_input_guardrails gains one branch: run in process when the runner knows the profile, else POST to the sidecar as today. A stored definition beats a sidecar profile of the same name.

Known limits, deliberate

  • A definition that failed to build is skipped, so the request goes through unchecked. To stop that being silent the read endpoint reports whether the runner holds the profile, and the dashboard row says so.
  • Enabled rows are capped at ten, because each one is another sequential vendor call in front of a request. The checks run sequentially; making them concurrent touches the sidecar path too and belongs in its own change.
  • Pass-through routes (embeddings, images, audio) do not go through the shared pre-dispatch seam and stay unchecked.

Stack

Last PR in the #1211, #1244, #1256 stack. Based on feat-guardrail-dashboard.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions