Use Defang's managed model provider (litellm) instead of the pinned gateway - #131
Merged
Merged
Conversation
…ateway The hand-pinned defangio/openai-access-gateway always sends both temperature and topP to Bedrock, which every Claude 4.5+ model rejects, and no client-side change can prevent it.
lionello
approved these changes
Sep 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
ask.defang.io is still down after #115 and #129. The remaining cause is in the gateway, not our app.
compose.yamlpinnedimage: defangio/openai-access-gateway. That gateway builds its BedrockinferenceConfigfromtemperature,maxTokensandtopP, and both parameters carry schema defaults (temperature1.0,top_p1.0). So it always sends both — even when the client sends neither — and every Claude 4.5+ model rejects that:#129 removed
top_pfrom our app, and the production log confirms the app now posts onlytemperature— yet Bedrock still returned the both-specified error, because the gateway re-added it. No change on our side could have fixed this.Defang has since moved to a managed model provider (litellm), which is what
x-defang-llmis meant to give you now. We were pinning our own image and missing it —fixupLLMonly adjusts an image whose name already ends in/litellm, so a pinned image is deployed as-is. Confirmed against the running task: thellmcontainer image isdefangio/openai-access-gateway.What changes
Replace the hand-rolled
llmservice with a top-levelmodels:entry and declare the dependency fromapp. Defang then generates the provider service itself.The CLI renders it as:
and wires
appup automatically:This fixes the outage because litellm forwards only the parameters the caller actually sent — it does not substitute a default
top_p— so with #129 the request carriestemperaturealone.--drop_paramsadditionally drops anything a given model does not accept, which makes the whole class of failure much less likely to recur on the next model swap.It also means we stop maintaining a pinned gateway image and follow whatever Defang ships.
Two details worth reviewing
OPENAI_BASE_URLis pinned explicitly rather than injected. The injected value ishttp://llm:4000/v1/with a trailing slash, and theopenai0.x client concatenatesapi_base + "/chat/completions", which would request/v1//chat/completions. Defang only injects when the variable is absent, so the explicit value wins.OPENAI_API_KEYis removed from the app service. It has to be, for the same reason: Defang injects the litellm master key only if the variable is not already set, and a mismatch would mean 401s. The app used that key solely to reach the gateway — embeddings run locally via sentence-transformers — so nothing else needs it. The now-unusedOPENAI_API_KEYentry indeploy.yaml'sconfig-env-varsis harmless and left alone; happy to strip it in a follow-up.Validation
defang compose config --provider awsexits 0 with no warnings, and the rendered output above is taken from it. Not deployable from here — merging deploys it. I'll verify the live site once it lands.Refs #115, #129