Skip to content

Repository files navigation

pi-model-fallback

Join dotfield.xyz on Discord

CI Publish npm version npm downloads License: MIT Pi package Buy Me A Coffee

Pi extension that switches to a fallback model after provider failures such as 429 rate limits.

What this does

pi-model-fallback watches provider failures and automatically moves Pi to a safer fallback model when a matching rule fires.

Current default:

  • source provider: zai
  • matching statuses: 429, 500, 502, 503, 504
  • fallback model: deepseek/deepseek-v4-flash

When a failure matches, the extension also stores persistent fallback state so future sessions can preselect the fallback model until the cooldown expires.

Install

Install from npm:

pi install npm:pi-model-fallback

Install into the current project only:

pi install npm:pi-model-fallback -l

Or install from GitHub:

pi install git:github.com/eiei114/pi-model-fallback

Try it without permanently installing:

pi -e npm:pi-model-fallback

Commands

/model-fallback:status
/model-fallback:reset
  • status: shows whether fallback is enabled, active persistent entries, and current paths
  • reset: clears persistent fallback state and switches back to the remembered original model when possible

Configuration

The extension exposes the model_fallback_config tool for reading, validating, and saving config JSON.

Default config shape:

{
  "version": 1,
  "enabled": true,
  "autoRetry": true,
  "rules": [
    {
      "name": "zai-to-deepseek-flash",
      "matchProviders": ["zai"],
      "statuses": [429, 500, 502, 503, 504],
      "fallback": {
        "provider": "deepseek",
        "model": "deepseek-v4-flash"
      }
    }
  ]
}

Rule fields:

  • matchProviders: match all models from a provider
  • matchModels: match specific provider + model pairs
  • statuses: optional; defaults to 429, 500, 502, 503, 504
  • reasons: optional list of failure reason codes (for example context_length_exceeded); when set, the rule only fires if the parsed failure reason is listed. Reasons come from quoted machine codes in the error message (for example "provider_error_code":"context_length_exceeded") or prose such as exceeds this model's context length. Useful for status codes like 400 that should only trigger a fallback for specific causes, such as context-length overflow.
  • cooldownMs: optional persistent fallback window
  • fallback: target model Pi should switch to

Top-level fields:

  • enabled: toggle the extension without deleting rules
  • autoRetry: when a fallback fires from a failed turn, automatically re-queue the failed user prompt so the turn retries on the fallback model (defaults to true; set to false to only switch models). The failed prompt is queued once per fallback transition, so a fallback chain may replay it more than once.

Rules use first-match order: the first rule whose provider/model and status match wins. Put specific matchModels rules before broad matchProviders rules when they should take priority. model_fallback_config validate, save, read, and status report warning-only diagnostics when a later rule or model entry is completely shadowed by an earlier rule; /model-fallback:status also includes a concise warning summary for the current config. Warning-bearing config remains valid and can still be saved.

Cooldowns

When a rule does not set cooldownMs, the extension uses these defaults:

  • 429 → 72 hours
  • 5xx → 10 minutes

When the provider response includes Retry-After or x-ratelimit-reset* headers, those values override cooldownMs and the defaults for the persisted fallback window.

State and paths

The extension stores:

  • config: model-fallback/config.json
  • state: model-fallback/state.json

If the package is installed project-locally and the current project references it from .pi/settings.json, those files live under the project .pi/ directory. Otherwise they live under the user agent directory.

Behavior notes

  • Successful responses do nothing.
  • Matching failures from after_provider_response can trigger fallback immediately.
  • Assistant error messages parsed at turn_end can also persist fallback state for SDK/provider failures that do not emit the normal response hook. Status extraction looks for HTTP-style tokens (for example status 429, HTTP 503, or rate limit) rather than any bare 3-digit number in the message.
  • With the default autoRetry: true, the failed prompt is automatically queued once per fallback transition on the fallback model. A fallback chain may therefore replay the prompt more than once. Set autoRetry: false to switch models without replaying it.

Rules never fall back to the failing model itself; a rule whose fallback target equals the failing model is skipped so the next matching rule (or no rule) applies.

Fallback chains cascade: when the active fallback model itself fails, another matching rule may move the session further (for example free model → kimi-k2.6 → glm-5.3-flash). A cascade never revisits a model already used in the current run (the original or an earlier fallback), so circular rule configurations cannot ping-pong; when no unvisited target remains, the failure surfaces normally.

Development

npm install
npm run ci

Run locally in Pi:

pi -e .

Links

License

MIT

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages