Relay docs

Caching & idempotency

#Response cache

Relay can cache successful responses to cut cost and latency for repeated calls.

  • Enable with Gateway:Cache:TtlSeconds (0 = off). Size is capped by Gateway:Cache:MaxEntries (default 5000).
  • Scope — the cache key is derived from the operation, model and normalised inputs, so identical requests hit the cache.
  • Per-request bypass — send X-Relay-Cache: no (or false) to force a fresh call.
  • Signal — the X-Relay-Cache response header reports HIT or MISS.

Both chat completions and embeddings honour the cache. Cached hits are still recorded to telemetry (so usage stays accurate).

#Idempotency

For safe retries of non-idempotent calls, send an Idempotency-Key header. Relay associates the result with that key so a retry with the same key returns the original result instead of running the request again. Use a fresh unique key per logical operation (e.g. a UUID per user action).

#When to use which

  • Cache — read-heavy, repeatable prompts (classification of the same text, FAQ answers, embeddings of stable content).
  • Idempotency — write-like operations you might retry after a network blip, where running twice would be wrong.

The cache is in-process. In a multi-instance deployment each instance has its own cache; treat it as a best-effort optimisation, not a shared store.