Chat completions
POST /v1/chat/completions — the primary endpoint. OpenAI-compatible, so most SDKs work by changing only the base URL and key.
#Request
{
"model": "gpt-4o",
"messages": [
{ "role": "system", "content": "You are concise." },
{ "role": "user", "content": "Summarize the benefits of an API gateway." }
],
"temperature": 0.3,
"max_tokens": 400,
"stream": false
}
| Field | Type | Notes |
|---|---|---|
model | string | A public model alias, or an auto strategy (auto, auto:cheap, auto:fast, auto:balanced). |
messages | array | { role, content } with roles system / user / assistant / tool. |
temperature, top_p, max_tokens, stop | — | Standard generation params (optional). |
stream | bool | true streams SSE chunks. |
tools, response_format, … | — | Unknown fields pass through to the provider. |
#Response
{
"id": "chatcmpl-…",
"choices": [{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }],
"usage": { "prompt_tokens": 42, "completion_tokens": 88, "total_tokens": 130 }
}
Response headers include X-Relay-Request-Id and X-Relay-Cost-Usd.
#Streaming
curl -N http://localhost:5300/v1/chat/completions \
-H "Authorization: Bearer $RELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "gpt-4o", "stream": true, "messages": [{ "role": "user", "content": "Write a haiku." }] }'
You receive data: {chunk} lines (each an OpenAI-shaped delta) ending with data: [DONE].
#What the gateway does for you
Behind this one call, Relay applies model routing, per-key model allow-list checks, transient retry, provider-health circuit breaking, single-hop fallback, optional response caching, optional moderation, and records telemetry. See Model routing & fallback.
#Caching & idempotency
Enable the response cache with Gateway:Cache:TtlSeconds. Bypass per request with X-Relay-Cache: no. Provide an Idempotency-Key header to make retried writes safe. See Caching & idempotency.