Relay docs

Model routing & fallback

When a request names a model, ModelRouter turns that into a concrete provider adapter and an ordered execution plan. This is where policy, auto selection, budgets and fallback live.

#Resolution order

  1. Explicit model — the alias you asked for, if it exists and is enabled.
  2. Team default — otherwise the team's / gateway's Gateway:DefaultModel.
  3. 404 — if nothing resolves, model_not_found.

Before executing, the router also enforces the per-key model allow-list (403 model_not_allowed) and the team budget hard-stop (402).

#Dynamic aliases

Gateway:ModelAliases:{requested} remaps a requested name to another at runtime — handy for redirecting gpt-4 → gpt-4o org-wide without touching apps or the catalog.

#auto strategies

Pass one of these as the model to let Relay choose among enabled models using rolling metrics:

Model valuePicks
auto / auto:balancedA balance of cost, latency and success.
auto:cheapThe cheapest capable model.
auto:fastThe lowest-latency model (by rolling EWMA).

Selection is driven by ModelMetricsRegistry (per-model EWMA latency + success ratio, updated from live outcomes) and the catalog's pricing.

#Resilience

GatewayChatService wraps every call:

  • Retry — transient upstream failures (ProviderException.IsTransient) are retried with exponential backoff (Gateway:MaxRetries, Gateway:RetryBaseDelayMs).
  • Circuit breaking — ProviderHealthRegistry trips a provider to degraded after consecutive transient failures, so the router avoids it.
  • Fallback — retry against another model when this one fails transiently. See below.
  • Active probing — HealthProbeService periodically checks each provider's reachability/latency; results surface at GET /v1/providers/health.

#Fallback

A model can name another model to be retried against when it fails. Set it per model in Models → Edit → Fallback model, which stores fallback_model_id on the row — no redeploy, no config change.

#Precedence

The router picks a fallback chain in this order, first match wins:

  1. The model's own Fallback model, followed transitively — gpt-4o → gpt-4o-mini → claude-sonnet.
  2. The auto runner-up, when the request asked for auto / auto:cheap / auto:fast / auto:balanced and the chosen model declares no fallback of its own. Second place in the cost/latency ranking becomes the fallback.
  3. Gateway:FallbackModel — the global default, used only when the primary declares no fallback. This is the original behaviour and still works; a per-model link simply overrides it.

A chain is capped at 3 targets including the primary. Each hop costs a full retry cycle before it is abandoned, so an unbounded chain would let one slow upstream stretch a single request across every model's retries.

#When it applies

PathFalls backWhy
POST /v1/chat/completionsYes
Agents (/v1/agents/{id}/run)YesRuns through GatewayChatService.
Workflow llm.* nodesYesSame pipeline.
Batch jobsYesApplied per row, so a rate limit that clears mid-job lets later rows return to the requested model.
Evals / A-B runsNoAn eval measures a named model. A substituted answer would attribute a score to a model that never produced it, and an A/B comparison would be against a mixture. A failed case is honest; a substituted one is not.
Embeddings and RAGNoTwo embedding models produce vectors in different spaces. Answering with a different one returns results that are wrong rather than degraded.

#When it triggers

Fallback is not a general retry. It engages only when all of these hold:

  • The failure is transient — upstream 5xx, 429, timeout, or connection error. A 400, 401 or 404 fails immediately, because a malformed or unauthorised request fails identically on any provider and retrying elsewhere only adds latency.
  • The primary's own retries are exhausted first (Gateway:MaxRetries, default 2).
  • For streaming, no token has reached the client yet. Once SSE output has begun it cannot be taken back, so a mid-stream failure ends the response.
  • The next target's provider is not degraded. Degraded providers are skipped as fallbacks; the explicitly requested primary is always attempted regardless.

#Edge cases

  • Loops are safe. A → B → A is a legitimate "each covers the other" setup and is accepted. The router tracks which models it has already tried and ends the chain rather than re-attempting one.
  • A model cannot be its own fallback — rejected on save, since it would be a silent no-op.
  • Dead links are skipped, not fatal. If the fallback is deleted, disabled, or its provider is disabled, the chain simply ends — a broken backup never stops a healthy primary from serving. The panel shows such a link as (deleted model).
  • Cost and telemetry are recorded against whichever model actually served the request. There is currently no flag distinguishing a fallback-served request from a direct one, so a fallback shows up as ordinary traffic to the fallback model.

#Catalog cache

The model + provider catalog is cached in-process for Gateway:CatalogCacheSeconds (default 10) and shared across requests; edits in the panel invalidate it. A fallback link you just saved therefore takes effect within that window.