Caching & idempotency
#Response cache
Relay can cache successful responses to cut cost and latency for repeated calls.
- Enable with
Gateway:Cache:TtlSeconds(0 = off). Size is capped byGateway:Cache:MaxEntries(default 5000). - Scope — the cache key is derived from the operation, model and normalised inputs, so identical requests hit the cache.
- Per-request bypass — send
X-Relay-Cache: no(orfalse) to force a fresh call. - Signal — the
X-Relay-Cacheresponse header reportsHITorMISS.
Both chat completions and embeddings honour the cache. Cached hits are still recorded to telemetry (so usage stays accurate).
#Idempotency
For safe retries of non-idempotent calls, send an Idempotency-Key header. Relay associates the result with that key so a retry with the same key returns the original result instead of running the request again. Use a fresh unique key per logical operation (e.g. a UUID per user action).
#When to use which
- Cache — read-heavy, repeatable prompts (classification of the same text, FAQ answers, embeddings of stable content).
- Idempotency — write-like operations you might retry after a network blip, where running twice would be wrong.
The cache is in-process. In a multi-instance deployment each instance has its own cache; treat it as a best-effort optimisation, not a shared store.