Relay docs

Chat completions

POST /v1/chat/completions — the primary endpoint. OpenAI-compatible, so most SDKs work by changing only the base URL and key.

#Request

{
  "model": "gpt-4o",
  "messages": [
    { "role": "system", "content": "You are concise." },
    { "role": "user", "content": "Summarize the benefits of an API gateway." }
  ],
  "temperature": 0.3,
  "max_tokens": 400,
  "stream": false
}
FieldTypeNotes
modelstringA public model alias, or an auto strategy (auto, auto:cheap, auto:fast, auto:balanced).
messagesarray{ role, content } with roles system / user / assistant / tool.
temperature, top_p, max_tokens, stop—Standard generation params (optional).
streambooltrue streams SSE chunks.
tools, response_format, …—Unknown fields pass through to the provider.

#Response

{
  "id": "chatcmpl-…",
  "choices": [{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }],
  "usage": { "prompt_tokens": 42, "completion_tokens": 88, "total_tokens": 130 }
}

Response headers include X-Relay-Request-Id and X-Relay-Cost-Usd.

#Streaming

curl -N http://localhost:5300/v1/chat/completions \
  -H "Authorization: Bearer $RELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "gpt-4o", "stream": true, "messages": [{ "role": "user", "content": "Write a haiku." }] }'

You receive data: {chunk} lines (each an OpenAI-shaped delta) ending with data: [DONE].

#What the gateway does for you

Behind this one call, Relay applies model routing, per-key model allow-list checks, transient retry, provider-health circuit breaking, single-hop fallback, optional response caching, optional moderation, and records telemetry. See Model routing & fallback.

#Caching & idempotency

Enable the response cache with Gateway:Cache:TtlSeconds. Bypass per request with X-Relay-Cache: no. Provide an Idempotency-Key header to make retried writes safe. See Caching & idempotency.