API overview
The gateway API is OpenAI-compatible. Point your existing OpenAI client at the Relay base URL and use a Relay API key.
- Base URL (local):
http://localhost:5300 - Auth:
Authorization: Bearer <relay-api-key>on every request (exceptGET /health). - Content type:
application/json.
#Endpoint map
| Method | Path | Purpose |
|---|---|---|
POST | /v1/chat/completions | Chat completions (streaming or not). |
POST | /v1/embeddings | Text embeddings. |
GET | /v1/models | List available models. |
POST | /v1/responses | Minimal OpenAI Responses shim. |
GET | /v1/prompts · /v1/prompts/{name} | List / fetch a prompt version. |
POST | /v1/prompts/{name}/compile | Compile a prompt with variables. |
GET | /v1/tools | Function-tool registry. |
POST | /v1/realtime/sessions | Mint a token for a realtime voice session. |
GET | /v1/realtime | WebSocket: bidirectional audio relay. |
POST | /v1/agents/{id}/realtime/sessions | Mint a realtime session backed by an agent. |
GET | /v1/agents | List the agents this key can run. |
POST | /v1/agents/{id}/chat/completions | Run a stored agent. |
POST | /v1/knowledge/{id}/documents · /files · /url · /search | RAG ingest & search. |
POST | /v1/batches · GET /v1/batches/{id} · /results | Batch jobs. |
GET | /v1/pricing · POST /v1/cost/simulate | Pricing & cost estimation. |
GET | /health · /v1/providers/health · /v1/models/metrics | Health & metrics. |
#Errors
Errors follow the OpenAI shape:
{ "error": { "message": "…", "type": "invalid_request_error", "code": "model_not_allowed" } }
Common cases: 401 (missing/invalid key), 403 (model_not_allowed — key isn't allowed that model), 402 (budget hard-stop), 404 (model_not_found / agent_not_found), 429 (rate/quota, retryable), 5xx (provider error).
Two more apply to agents, and both are about the acting end user rather than the key: 400 user_header_required when the agent needs X-Relay-User-Id and it is missing or malformed (a dataset agent, or one that is shared with named users), and 403 agent_not_shared when that user is not on the agent's share list.
#Useful response headers
X-Relay-Request-Id— the id for this request (also in telemetry).X-Relay-Cost-Usd— computed cost of the call.X-Relay-Cache—HIT/MISSwhen the response cache is enabled.X-Relay-Thread-Id— set by the agent endpoint when using?thread=.
#Streaming
Set "stream": true on chat requests to receive Server-Sent Events (data: {chunk} lines, terminated by data: [DONE]), exactly like OpenAI. See Chat completions.