Model routing & fallback
When a request names a model, ModelRouter turns that into a concrete provider adapter and an ordered execution plan. This is where policy, auto selection, budgets and fallback live.
#Resolution order
- Explicit model — the alias you asked for, if it exists and is enabled.
- Team default — otherwise the team's / gateway's
Gateway:DefaultModel. - 404 — if nothing resolves,
model_not_found.
Before executing, the router also enforces the per-key model allow-list (403 model_not_allowed) and the team budget hard-stop (402).
#Dynamic aliases
Gateway:ModelAliases:{requested} remaps a requested name to another at runtime — handy for redirecting gpt-4 → gpt-4o org-wide without touching apps or the catalog.
#auto strategies
Pass one of these as the model to let Relay choose among enabled models using rolling metrics:
| Model value | Picks |
|---|---|
auto / auto:balanced | A balance of cost, latency and success. |
auto:cheap | The cheapest capable model. |
auto:fast | The lowest-latency model (by rolling EWMA). |
Selection is driven by ModelMetricsRegistry (per-model EWMA latency + success ratio, updated from live outcomes) and the catalog's pricing.
#Resilience
GatewayChatService wraps every call:
- Retry — transient upstream failures (
ProviderException.IsTransient) are retried with exponential backoff (Gateway:MaxRetries,Gateway:RetryBaseDelayMs). - Circuit breaking —
ProviderHealthRegistrytrips a provider to degraded after consecutive transient failures, so the router avoids it. - Fallback — retry against another model when this one fails transiently. See below.
- Active probing —
HealthProbeServiceperiodically checks each provider's reachability/latency; results surface atGET /v1/providers/health.
#Fallback
A model can name another model to be retried against when it fails. Set it per model in Models → Edit → Fallback model, which stores fallback_model_id on the row — no redeploy, no config change.
#Precedence
The router picks a fallback chain in this order, first match wins:
- The model's own
Fallback model, followed transitively —gpt-4o→gpt-4o-mini→claude-sonnet. - The
autorunner-up, when the request asked forauto/auto:cheap/auto:fast/auto:balancedand the chosen model declares no fallback of its own. Second place in the cost/latency ranking becomes the fallback. Gateway:FallbackModel— the global default, used only when the primary declares no fallback. This is the original behaviour and still works; a per-model link simply overrides it.
A chain is capped at 3 targets including the primary. Each hop costs a full retry cycle before it is abandoned, so an unbounded chain would let one slow upstream stretch a single request across every model's retries.
#When it applies
| Path | Falls back | Why |
|---|---|---|
POST /v1/chat/completions | Yes | |
Agents (/v1/agents/{id}/run) | Yes | Runs through GatewayChatService. |
Workflow llm.* nodes | Yes | Same pipeline. |
| Batch jobs | Yes | Applied per row, so a rate limit that clears mid-job lets later rows return to the requested model. |
| Evals / A-B runs | No | An eval measures a named model. A substituted answer would attribute a score to a model that never produced it, and an A/B comparison would be against a mixture. A failed case is honest; a substituted one is not. |
| Embeddings and RAG | No | Two embedding models produce vectors in different spaces. Answering with a different one returns results that are wrong rather than degraded. |
#When it triggers
Fallback is not a general retry. It engages only when all of these hold:
- The failure is transient — upstream 5xx, 429, timeout, or connection error. A 400, 401 or 404 fails immediately, because a malformed or unauthorised request fails identically on any provider and retrying elsewhere only adds latency.
- The primary's own retries are exhausted first (
Gateway:MaxRetries, default 2). - For streaming, no token has reached the client yet. Once SSE output has begun it cannot be taken back, so a mid-stream failure ends the response.
- The next target's provider is not degraded. Degraded providers are skipped as fallbacks; the explicitly requested primary is always attempted regardless.
#Edge cases
- Loops are safe.
A → B → Ais a legitimate "each covers the other" setup and is accepted. The router tracks which models it has already tried and ends the chain rather than re-attempting one. - A model cannot be its own fallback — rejected on save, since it would be a silent no-op.
- Dead links are skipped, not fatal. If the fallback is deleted, disabled, or its provider is disabled, the chain simply ends — a broken backup never stops a healthy primary from serving. The panel shows such a link as
(deleted model). - Cost and telemetry are recorded against whichever model actually served the request. There is currently no flag distinguishing a fallback-served request from a direct one, so a fallback shows up as ordinary traffic to the fallback model.
#Catalog cache
The model + provider catalog is cached in-process for Gateway:CatalogCacheSeconds (default 10) and shared across requests; edits in the panel invalidate it. A fallback link you just saved therefore takes effect within that window.