Repository navigation
Harden OpenShell eval pipeline: judge-safe model routing, gateway alias, best-effort cleanup - #110
Open
shricharan-ks wants to merge 11 commits into
Open
shricharan-ks wants to merge 11 commits into
shricharan-ks wants to merge 11 commits into
Conversation
Port the verified OIDC refresh, USER.md fixture, pinned SAW image, and published_brief gate onto main. Add same-namespace NetworkPolicy templates that match the canonical Pipeline name or app.kubernetes.io/part-of=abevalflow so ad-hoc Pipeline copies no longer time out on gateway preflight. Co-authored-by: Cursor <cursoragent@cursor.com>
Depth-1 --branch fails for bare SHAs; fall back to full clone + checkout so clean-source pins work the same way as the submission revision. Co-authored-by: Cursor <cursoragent@cursor.com>
…ervice Selects the agent VMI directly so Pipeline defaults need no deployment namespace. The gateway server leaf includes 'openshell' as a DNS SAN, letting the evaluate task keep OPENSHELL_GATEWAY_INSECURE=false. Co-Authored-By: Claude Code <noreply@anthropic.com>
- openai/ route for rits/zai-org/GLM-5-3-Flash (agent model) - judge-glm-5-3 hosted_vllm route: LLM judges call /v1/messages (Anthropic dialect); with an openai/ model LiteLLM converts that to the upstream /responses API which answers 403, so every judge errors. hosted_vllm stays on /chat/completions. Select with aeh-judge-model-override=judge-glm-5-3. Also drop the hardcoded gz-forge-eval namespace so the ConfigMap applies anywhere. Co-Authored-By: Claude Code <noreply@anthropic.com>
Probes every dependency an eval pod needs: mlflow, litellm, minio, k8s api, github, pypi, github releases/raw, and TCP reachability of the agent/integ gateways and postgres. Deliberately not labelled part-of=abevalflow so NetworkPolicy sees it as an ordinary Tekton pod and reports real connectivity. Co-Authored-By: Claude Code <noreply@anthropic.com>
- image-registry.openshift-image-registry openshift/cli instead of registry.redhat.io (avoids pull failures without a redhat registry secret on the pipeline SA) - --request-timeout on oc calls, 3 retries with backoff - onError: continue so the finally task never fails the PipelineRun Co-Authored-By: Claude Code <noreply@anthropic.com>
… alias - llm-model / aeh-model-override default to the Flash model; the submission's model entry carries the 128000-token budget it needs - aeh-judge-model-override defaults to judge-glm-5-3 (hosted_vllm route; a plain openai/ judge model 403s on /v1/messages) - openshell-gateway-endpoint default becomes https://openshell:17670, the alias Service from config/forge-saw/openshell-eval-alias.yaml, so the Pipeline applies to any namespace - example PipelineRun: short-name endpoints (namespace-agnostic, drops the gz-forge-eval hostAliases pin), main revisions, and timeouts sized for 15-minute cases (3h pipeline / 2h30m tasks) Co-Authored-By: Claude Code <noreply@anthropic.com>
…ardening Extends the OpenShell profile tests for the carried forge-nommen changes: Flash model defaults, judge-glm-5-3 override, namespace- agnostic example run, gateway alias Service selector, LiteLLM Flash/judge routes, and best-effort PVC cleanup. Co-Authored-By: Claude Code <noreply@anthropic.com>
Pinned namespaces (ab-eval-flow, gz-forge-eval) made `oc apply -n <ns>` fail with a namespace mismatch for anyone deploying the OpenShell profile into their own workspace (e.g. forge-nommen). Remove the namespace fields so the manifests land in whatever namespace they are applied to; the example PipelineRun already uses short service names that resolve in the run namespace. Co-Authored-By: Claude Code <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why this is needed
The OpenShell OpenClaw profile (
eval-engine=aeh_openshell_openclaw) evaluates the Forge "Chief of Staff" agent end-to-end: it stages a submission workspace, boots the agent in an OpenShell sandbox on a KubeVirt VM, drives scene-based cases (morning-briefing, analysis-panel), scores them with deterministic + LLM judges, and publishes results to MLflow. Running it in a per-user namespace (verification was done inforge-nommen) exposed five gaps that made runs either fail outright or silently lose all LLM-judge signal. This PR closes those gaps; none of them are covered by #108/#109.Problems earlier → how this fixes them
mean_rewardunusable. The judge client (score.py) only speaks the Anthropic dialect (POST /v1/messages). For anopenai/model, LiteLLM converts that to the upstream's/responsesAPI, which the upstream vLLM gateway rejects with403 Authentication parameters missing.openai/route.judge-glm-5-3LiteLLM route (hosted_vllm/) that keeps judge calls on/chat/completions, plusaeh-judge-model-override=judge-glm-5-3default.llm-modelwasrits/zai-org/glm-5-3(reasoning model), while the deployed agent runs GLM-5-3-Flash. Evals were slower and didn't measure the model users actually get.rits/zai-org/GLM-5-3-Flash(pipeline + example run +aeh-model-override).openshell-gateway-endpointembedded another namespace's DNS and didn't validate against the gateway certificate, so the same Pipeline failed when created anywhere else.openshell-saw-agentVMI in the workspace namespace; there was no stable local name for it.openshell(port 17670) selecting the VMI; default endpointhttps://openshell:17670, valid wherever the alias is installed.bitnami/kubectlfrom Docker Hub (rate-limited pulls) with no request timeout — a hung API call heldcleanupuntil the finally budget expired and a failed delete left the PVC behind; a failed finally marks the PipelineRun failed.image-registry.openshift-image-registry.svc:5000/openshell/cli),--request-timeout=20s, 3 retries,onError: continue.oc apply -n <ns> -f …failed for anyone but the original namespace. Manifests pinnednamespace: ab-eval-flow/gz-forge-eval, so applying into a per-user namespace returned "the namespace from the provided object does not match".http://litellm:4000,https://openshell:17670) that resolve in whatever namespace the run is created in.Additionally, failures were slow to diagnose because nothing could tell which dependency (litellm, mlflow, minio, k8s API, github, pypi, agent gateway, postgres) was unreachable — hence the standalone probe task below.
Changes, one by one
feat(forge-saw): add stable namespace-local openshell gateway alias Service— newconfig/forge-saw/openshell-eval-alias.yaml. A selector Service for theopenshell-saw-agentVMI on port 17670. Justification: gives every namespace a stable, certificate-validhttps://openshell:17670endpoint without cross-namespace DNS or cert pinning; installed with the stack, referenced by the Pipeline default.feat(litellm): add GLM-5-3-Flash routes for agent and judge traffic—config/litellm/configmap.yamladdsrits/zai-org/GLM-5-3-Flash(openai/ lane, agent traffic) andjudge-glm-5-3(hosted_vllm lane, judge traffic) over the sameINFERENCE_ENDPOINT_URL/INFERENCE_API_KEYenv. Justification: the dialect-safe judge route is the only way the Anthropic-dialect judge client gets served by this OpenAI-dialect upstream (problem 1); the Flash route matches the deployed agent (problem 2).feat(pipeline): add eval-stack-probe task— newpipeline/tasks/eval-stack-probe.yaml: HTTP probes (mlflow, litellm, minio, k8s API, github, pypi, local wheels) + TCP probes (agent gateway, integration gateway, postgres). Deliberately not labelledpart-ofso it can run without the pipeline's NetworkPolicy allowances. Justification: one TaskRun replaces reading thousands of lines of evaluate logs to find a dead dependency; it is a triage tool, not a pipeline stage.fix(cleanup): make PVC cleanup best-effort and use in-cluster CLI image—pipeline/tasks/post/cleanup_pvc.yaml. Justification: problem 4 — cleanup must never fail the run or leak PVCs; the in-cluster registry removes the Docker Hub dependency.feat(pipeline): default to GLM-5-3-Flash with judge route and gateway alias—pipeline/pipelines/ci-pipeline-openshell.yaml+pipeline/runs/openshell-openclaw-pipelinerun.yaml. Justification: wires problems 1–3 into the Pipeline defaults so a plainoc create -fof the example run works in any namespace; keeps explicit timeout budgets (pipeline 3h / tasks 2h30m / finally 15m) and thepart-oflabel required by the canonical NetworkPolicies.test: cover Flash defaults, judge route, alias Service, and cleanup hardening—tests/test_openshell_pipeline_profile.py(10 tests). Justification: locks in every default this PR changes — judge route must stay hosted_vllm, example run must stay namespace-agnostic, cleanup must stay best-effort — so a future edit can't silently regress them. Also asserts the evaluate step takes the model key from$(params.llm-api-key)(no namespace-secret dependency, per Fix OpenShell evaluation Secret dependency #108).chore: drop hardcoded namespaces from deployable manifests— 8 files, one line each. Justification: problem 5; makes every manifest apply cleanly with-n <any-namespace>.Dependencies
3232787); until Port Forge OpenShell reliability fixes onto current main #109 merges, this PR's diff also shows its 4 commits. After Port Forge OpenShell reliability fixes onto current main #109 merges, GitHub will recompute this PR to show only the 7 commits above.$(params.llm-api-key)instead of reading a namespace secret; this PR's tests pin that behavior.main) — OpenClaw brief-reader/gateway fixes the pipeline's harness revision default (main) requires.openshell-mtls/forge-ai-gateway-ca/openshell-credentials/openshell-oidc-credentials, theopenshellalias Service, and the agent VMI runningghcr.io/rh-forge/openclaw-saw-agent@sha256:b47b92a6b3fd03327c1f2093a5c28aba0fdf3cb620e9154335688900191fe2b9.Verification
scharan-verify-hardening-2(namespaceforge-nommen, 2026-10-09): all 5 tasks Succeeded; recommendationpass, mean_reward 0.7000 (threshold 0.5); 12 judges × 2 cases with 0 errors (previously all LLM judges errored); scorecard 6/6 gates; MLflow experiment recorded with traces + judge feedback; the hardened finally task deleted the run's PVC.sana-morning-briefing-pr105-single-*,-collector-diag-*) scored 0.00–0.35 and failed theaeh-low-scoregate.tests/test_openshell_pipeline_profile.py— 10/10 pass.forge-nommenverified identical to this branch (all 44 Pipeline param defaults, every task step script, litellm configmap, alias Service selector).Ops note (out of scope, for the record)
During verification the run-namespace LiteLLM lost its upstream key (the k8s secret holding it had been deleted; only the old pods still had it in memory). #108 already removed the eval task's dependency on that secret; the key for the namespace's LiteLLM deployment itself was restored from the vault-injected value on the running
forge-ai-gatewaypod. Keeping that namespace secret in sync with Vault is an ops concern outside this repo.🤖 Generated with Claude Code