Skip to content

Use Defang's managed model provider (litellm) instead of the pinned gateway - #131

Merged
lionello merged 1 commit into
mainfrom
fix/litellm-model-provider
Sep 5, 2026
Merged

lionello merged 1 commit into
mainfrom
fix/litellm-model-provider

Conversation

@defangdevs

Copy link
Copy Markdown
Contributor

Why

ask.defang.io is still down after #115 and #129. The remaining cause is in the gateway, not our app.

compose.yaml pinned image: defangio/openai-access-gateway. That gateway builds its Bedrock inferenceConfig from temperature, maxTokens and topP, and both parameters carry schema defaults (temperature 1.0, top_p 1.0). So it always sends both — even when the client sends neither — and every Claude 4.5+ model rejects that:

ValidationException ... `temperature` and `top_p` cannot both be specified for this model.

#129 removed top_p from our app, and the production log confirms the app now posts only temperature — yet Bedrock still returned the both-specified error, because the gateway re-added it. No change on our side could have fixed this.

Defang has since moved to a managed model provider (litellm), which is what x-defang-llm is meant to give you now. We were pinning our own image and missing it — fixupLLM only adjusts an image whose name already ends in /litellm, so a pinned image is deployed as-is. Confirmed against the running task: the llm container image is defangio/openai-access-gateway.

What changes

Replace the hand-rolled llm service with a top-level models: entry and declare the dependency from app. Defang then generates the provider service itself.

The CLI renders it as:

llm:
  image: litellm/litellm:v1.92.0
  command: [--drop_params, --model, bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0,
            --alias, us.anthropic.claude-haiku-4-5-20251001-v1:0]
  environment: {AWS_REGION: us-east-1, LITELLM_MASTER_KEY: networkisalreadyprivate}
  ports: [{target: 4000, mode: host}]
  x-defang-llm: true

and wires app up automatically:

MODEL: bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0
OPENAI_API_KEY: networkisalreadyprivate   # matches LITELLM_MASTER_KEY
OPENAI_BASE_URL: http://llm:4000/v1

This fixes the outage because litellm forwards only the parameters the caller actually sent — it does not substitute a default top_p — so with #129 the request carries temperature alone. --drop_params additionally drops anything a given model does not accept, which makes the whole class of failure much less likely to recur on the next model swap.

It also means we stop maintaining a pinned gateway image and follow whatever Defang ships.

Two details worth reviewing

  • OPENAI_BASE_URL is pinned explicitly rather than injected. The injected value is http://llm:4000/v1/ with a trailing slash, and the openai 0.x client concatenates api_base + "/chat/completions", which would request /v1//chat/completions. Defang only injects when the variable is absent, so the explicit value wins.
  • OPENAI_API_KEY is removed from the app service. It has to be, for the same reason: Defang injects the litellm master key only if the variable is not already set, and a mismatch would mean 401s. The app used that key solely to reach the gateway — embeddings run locally via sentence-transformers — so nothing else needs it. The now-unused OPENAI_API_KEY entry in deploy.yaml's config-env-vars is harmless and left alone; happy to strip it in a follow-up.

Validation

defang compose config --provider aws exits 0 with no warnings, and the rendered output above is taken from it. Not deployable from here — merging deploys it. I'll verify the live site once it lands.

Refs #115, #129

…ateway

The hand-pinned defangio/openai-access-gateway always sends both
temperature and topP to Bedrock, which every Claude 4.5+ model rejects,
and no client-side change can prevent it.
@lionello
lionello merged commit a2287b5 into main Sep 5, 2026
6 checks passed
@lionello
lionello deleted the fix/litellm-model-provider branch September 5, 2026 00:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants