A Kubernetes operator that ingests signals from security/observability tools (Trivy, and more later), uses an LLM to enrich them into ranked hypotheses, and proposes remediation as GitOps pull requests — with bounded LLM spend and a published record of its own accuracy.
📖 Full docs, architecture, and the case for Candor →
Status: early build, shipped in slices — see docs/design.md for exactly
what's done.
What sets Candor apart:
- Bounded LLM cost, by construction — an LLM call happens once per distinct cluster-state fingerprint, ever, not once per reconcile. A hard budget ceiling degrades to deterministic-only findings on exhaustion rather than silently spending past it.
- GitOps-native remediation — the default write path is a pull request against your GitOps repo, never a direct cluster mutation.
- Published accuracy — every finding is re-checked automatically, and its outcome (resolved, still present, recurred) is a queryable Prometheus metric, not a claim in a README.
- Suppression that resurfaces on its own — mute a fingerprint with a required reason and optional expiry; if the underlying content actually changes, it stops matching and comes back automatically. Nothing to remember to clean up.
- Untrusted-input safe — all ingested telemetry is treated as data, never instructions. Model output is grammar-constrained to a fixed action catalog; it can't emit a free-form action.
- Every write path has a brake, not just the LLM - a pull request only opens once the fix is independently, mechanically verified, and even then it's bounded by a per-namespace cap, a cluster-wide cap, and a panic switch that pauses it everywhere with no redeploy.
What you'd expect from a tool in this space, done properly:
- Provider-based signal ingestion (Trivy today; the interface is provider-agnostic)
- Fully Kubernetes-native —
kubectl get findings, no separate UI or database to run - Prometheus metrics and a ready-made Grafana dashboard, shipped in the Helm chart
- A generic webhook sink and periodic digest — one JSON payload shape, works with Slack, Teams, PagerDuty, or anything that can receive a POST
- Every release is signed, with a CycloneDX SBOM and build provenance attached
helm install candor oci://ghcr.io/teerakarna/charts/candor --version <version> \
--namespace candor-system --create-namespaceOpt a namespace in:
kubectl apply -n <your-namespace> -f - <<EOF
apiVersion: candor.dev/v1alpha1
kind: SignalPolicy
metadata:
name: default
spec:
providers: [trivy]
minSeverity: HIGH
EOFIf Trivy Operator is already scanning that
namespace, you'll see results as soon as it produces a VulnerabilityReport:
kubectl get findings -n <your-namespace>That's it — deterministic findings work with zero further configuration. See Configuration below to enable LLM enrichment, a budget ceiling, suppression, notifications, and GitOps pull request remediation.
LLM enrichment (ranked hypotheses on each Finding) is opt-in - without an API key it's simply not
active, not degraded:
kubectl create secret generic candor-llm --namespace candor-system \
--from-literal=anthropic-api-key=<your key>Uses Anthropic's structured outputs
so enrichment output is grammar-constrained to the hypotheses schema - see
internal/llm/anthropic and SECURITY.md for why that
matters given the input is untrusted scanner data. Override the model with
CANDOR_LLM_MODEL (default: claude-sonnet-5) via manager.envOverrides in the Helm chart's
values.
apiVersion: candor.dev/v1alpha1
kind: SignalPolicy
metadata:
name: default
spec:
providers: [trivy]
minSeverity: HIGH
budget:
maxCalls: 50
windowSeconds: 86400On exhaustion, enrichment degrades to deterministic-only findings for the rest of the window.
Mute a specific Finding by its exact content fingerprint:
kubectl get finding <name> -o jsonpath='{.status.fingerprint}'
kubectl apply -f - <<EOF
apiVersion: candor.dev/v1alpha1
kind: Suppression
metadata:
name: known-false-positive
spec:
fingerprint: "<the fingerprint above>"
reason: "known false positive - tracked in TICKET-123"
# expiresAt: "2026-12-31T00:00:00Z" # omit for indefinite
EOFMatching is exact by construction: if the underlying signal's content actually changes, its fingerprint changes too, the Suppression no longer matches, and the Finding resurfaces on its own
- there's nothing to remember to delete or update.
Set spec.webhook.url on a SignalPolicy to get a JSON POST on every Finding created, resolved,
or recurred in that namespace, plus a periodic digest (CANDOR_DIGEST_INTERVAL, default 24h)
tallying Findings by verification outcome and severity:
apiVersion: candor.dev/v1alpha1
kind: SignalPolicy
metadata:
name: team-a
spec:
providers: [trivy]
webhook:
url: https://example.com/hooks/candorIt's a plain JSON POST with no vendor-specific formatting - point it at whatever turns JSON into a Slack/Teams/PagerDuty message (a relay, a low-code webhook, etc.).
The inbound counterpart to the section above: Trivy works out of the box because it's CRD-shaped,
but SonarQube, Falco, and most other scanners aren't. Enable the generic receiver on a namespace's
SignalPolicy, and any tool that can POST an event on its own can feed Candor - no bespoke
provider needed.
kubectl create secret generic webhook-secret --namespace <your-namespace> \
--from-literal=secret=<a random HMAC signing key>apiVersion: candor.dev/v1alpha1
kind: SignalPolicy
metadata:
name: default
spec:
providers: [webhook]
webhookReceiver:
secretRef:
name: webhook-secretPOST Candor's own normalized envelope to
http://candor-webhook-receiver-service.candor-system:9444/webhook/<namespace>/<signalpolicy-name>
(the Service both the Helm chart and the kustomize base ship), signed with
X-Candor-Signature: sha256=<hex-encoded HMAC-SHA256 of the raw body>:
BODY='{"severity":"HIGH","kind":"Deployment","name":"api","summary":"SQL injection risk","id":"sonarqube-issue-123"}'
SIG=$(printf '%s' "$BODY" | openssl dgst -sha256 -hmac "<the secret above>" | awk '{print "sha256="$NF}')
curl -X POST "http://candor-webhook-receiver-service.candor-system:9444/webhook/<your-namespace>/default" \
-H "X-Candor-Signature: $SIG" -d "$BODY"id is the sender's own stable identifier for the underlying issue (a SonarQube issue key, a Falco
rule+resource combination) - repeated deliveries with the same id update one Finding rather than
create a new one each time, the same identity guarantee every other provider gets for free from its
own CRD. A tool whose native webhook payload doesn't already match this shape needs a small
transform in front - the same relay-in-front pattern the outbound sink above uses, just running the
other direction.
No webhookReceiver configured means that namespace's endpoint always rejects, regardless of what's
posted to it - unlike every other guardrail in Candor, there's no conservative default here to fall
back to. The controller needs the same namespaced-Role opt-in as gitOpsRepo below to read the
Secret - no cluster-wide Secret access, ever.
Candor's default write path is a pull request against your GitOps repo, never a direct cluster
mutation. It only fires when a fix is independently, mechanically verifiable - today that means a
Trivy VulnerabilityReport where every vulnerability agrees on one fixedVersion - regardless of
what the LLM recommends. No such fix, no PR.
GitHub and GitLab (SaaS or self-hosted) are both supported - set gitOpsRepo.provider to pick one
(defaults to github) and gitOpsRepo.host if it's a self-hosted GitLab or a GitHub Enterprise
instance rather than the public SaaS API.
The controller deliberately has no cluster-wide access to Secrets, so enabling this needs two
things in the namespace: a token for the chosen provider, and a Role granting the controller's
ServiceAccount get on that one Secret specifically.
kubectl create secret generic github-token --namespace <your-namespace> \
--from-literal=token=<a GitHub token with contents + pull-request write access>
kubectl apply -n <your-namespace> -f - <<EOF
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: candor-read-github-token
rules:
- apiGroups: [""]
resources: ["secrets"]
resourceNames: ["github-token"]
verbs: ["get"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: candor-read-github-token
subjects:
- kind: ServiceAccount
name: candor-controller-manager # matches your Helm release name + "-controller-manager"
namespace: candor-system
roleRef:
kind: Role
name: candor-read-github-token
apiGroup: rbac.authorization.k8s.io
EOFThen point a SignalPolicy at the repo to patch:
apiVersion: candor.dev/v1alpha1
kind: SignalPolicy
metadata:
name: default
spec:
providers: [trivy]
gitOpsRepo:
provider: github # optional, defaults to github - set to "gitlab" for GitLab
# host: https://ghe.example.com # optional - self-hosted GitHub Enterprise or GitLab only
owner: your-org
repo: gitops-demo
baseBranch: main # optional, defaults to main
path: apps/api/values.yaml
yamlPath: image.tag
secretRef:
name: github-token
pullRequestBudget: # optional - omitting it still applies a conservative
maxPullRequests: 5 # built-in default (1/24h), never "unlimited"
windowSeconds: 86400For GitLab, owner is the namespace the project lives in (a group, or group/subgroup for one
nested in a subgroup), repo is the project name, and secretRef should point at a personal or
project access token with api scope rather than a GitHub token - the RBAC and Secret-creation
steps above are otherwise identical, just naming a GitLab token instead.
The opened PR (or, for GitLab, merge request) carries the finding, its ranked hypotheses, and confidence in the description - no new surface to learn beyond reading a normal PR/MR.
SignalPolicy.spec.pullRequestBudget bounds one namespace. OperatingPolicy - Candor's one
cluster-scoped resource - bounds the whole cluster, independently and in addition: a namespace
comfortably under its own budget can still be refused once the cluster-wide ceiling is spent by
everyone else combined.
It's also the panic switch. Flip mode to Audit to stop every ProposePullRequest action across
every namespace immediately, with no redeploy - checked before any budget is even touched, so it's
a genuine full stop, not "still counted but not executed."
apiVersion: candor.dev/v1alpha1
kind: OperatingPolicy
metadata:
name: default # exactly one is meaningful cluster-wide; a second is flagged, not merged
spec:
pullRequestRateLimit:
maxPullRequests: 5
windowSeconds: 86400
mode: Active # set to "Audit" to pause ProposePullRequest cluster-wideUnlike every other budget in Candor, no OperatingPolicy existing at all means ProposePullRequest
is off cluster-wide, not "unlimited" - there's nowhere to persist a count, so allowing anyway would
just be an uncounted, unenforced brake. Create even an empty one to turn the action path on at the
conservative built-in default.
Ship a pre-built dashboard (LLM calls, enrichment skipped by reason, verification transitions, budget usage, cost-avoidance ratio) as a ConfigMap the kube-prometheus-stack Grafana sidecar auto-discovers:
helm upgrade --install candor oci://ghcr.io/teerakarna/charts/candor \
--set prometheus.enabled=true --set grafanaDashboard.enabled=trueEach release attaches a versioned install.yaml
(all resources, generated fresh at release time — not a stale copy on main):
kubectl apply -f https://github.com/teerakarna/candor/releases/download/<tag>/install.yamlBoth install paths are produced by the same release pipeline — nothing hand-built or committed to
main, so what you install is always a specific, versioned, signed release. One content
difference: NetworkPolicy support is currently Helm-only (--set networkPolicy.enabled=true), since
config/network-policy/ isn't wired into config/default/kustomization.yaml yet (tracked as
#64), so a kubectl apply -f install.yaml install
gets no NetworkPolicy at all.
- go version v1.24.6+
- docker version 17.03+
- kubectl version v1.11.3+
- Access to a Kubernetes v1.11.3+ cluster
make dev-up # creates (or reuses) a Kind cluster, builds the image, deploys Candor
make dev-status # kubectl get pods -n candor-system
make dev-down # tear the cluster downmake dev-up is safe to re-run after a code change - it rebuilds the image and redeploys onto the
same cluster. It also installs the pinned Trivy VulnerabilityReport CRD
(test/crd/), so you can hand-apply a SignalPolicy and a fake VulnerabilityReport and watch a
Finding come out the other end without a real Trivy Operator running. This is a separate,
persistent cluster from the one make test-e2e creates and destroys automatically around itself.
make docker-build docker-push IMG=<some-registry>/candor:tag
make install # CRDs
make deploy IMG=<some-registry>/candor:tag
kubectl apply -k config/samples/ # sample SignalPolicyIf you hit RBAC errors, make sure you're logged in with sufficient cluster privileges.
To tear back down:
kubectl delete -k config/samples/
make undeploy
make uninstallThe Helm chart source lives under charts/chart/ (generated via kubebuilder edit --plugins=helm/v2-alpha --output-dir=charts, regenerate the same way after changing the API or
RBAC). Published to the OCI registry on every tagged release, alongside the signed image.
See CONTRIBUTING.md — in particular the design constraints PRs must respect (bounded LLM calls, read-only cluster access by default, no free-form model-selected actions).
NOTE: Run make help for more information on all potential make targets
More information can be found via the Kubebuilder Documentation
Copyright 2026 Albert Asawaroengchai.
Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.