Relay docs

Realtime API

Hold a spoken conversation with a model, or with an agent, over a WebSocket. Audio flows both ways: you stream microphone audio up, the model streams speech back, and the provider handles turn detection and interruption.

This is a different wire protocol from /v1/chat/completions, not an option on it. A model is one or the other, which is why the catalogue carries a Realtime flag — see Models.

#Two steps, and why

A browser cannot set request headers on a WebSocket. So a realtime connection is made in two moves:

  1. Mint a session over normal authenticated HTTPS. You get a short-lived token.
  2. Open the socket, presenting that token in the query string.

The alternatives are worse, and it is worth being explicit about why: putting your app_live_… key in a WebSocket query string writes a long-lived credential into every proxy, load balancer and access log along the path; handing the key to browser JavaScript exposes it to anyone who opens dev tools. The minted token is single-use and expires in 60 seconds, so by the time it appears in a log it is already spent or dead. Your provider key never leaves the gateway at all.

#1. Mint a session

POST /v1/realtime/sessions

curl -X POST http://localhost:5300/v1/realtime/sessions \
  -H "Authorization: Bearer $RELAY_API_KEY" -H "Content-Type: application/json" \
  -d '{ "model": "gpt-realtime", "voice": "alloy" }'
{
  "object": "realtime.session",
  "model": "gpt-realtime",
  "agent_id": null,
  "client_secret": { "value": "u7Xh…", "expires_at": 1772539200 },
  "url": "/v1/realtime?session=u7Xh…"
}
FieldMeaning
modelMust be a catalogue model marked realtime-capable, and permitted by the key's model allow-list.
voicePassed through to the provider, e.g. alloy, echo, shimmer.
instructionsExtra system instructions for this session.
sessionProvider-shaped session fields merged into the config — turn detection, audio formats, modalities. Anything the gateway does not model itself. Preview-era key names are rewritten to their GA locations; see the session object.

url is returned as an absolute path, not a full URL. The gateway does not guess its own public origin: behind a reverse proxy that is how you end up telling a browser to open ws:// from an https:// page. Prefix it with your gateway origin and swap the scheme — http → ws, https → wss. Concatenating the origin without swapping produces something that is not a URL at all, and the WebSocket constructor rejects it.

Mixed content applies: a page served over HTTPS cannot open a plain ws:// socket to a remote host. Loopback is exempt — browsers treat localhost, 127.0.0.1 and [::1] as potentially trustworthy — so the usual dev setup (panel on https://localhost:7236, gateway on http://localhost:5300) works. In production, serve the gateway over HTTPS and use wss://.

Errors are the ordinary ones, up front rather than as a socket failure:

StatusWhen
400 unsupported_capabilityThe model exists but is not marked realtime-capable.
403 model_not_allowedThe key's model allow-list excludes it.
404 model_not_foundNo such enabled model. The message names the realtime aliases that do exist, and detects the common mistake of passing the provider model id (an Azure deployment name) where the public alias belongs.

model is the public alias from the catalogue, not the upstream deployment id. With Azure Foundry the two are usually different — the alias is what you named the row in the panel, the deployment id is what Azure calls it — and passing the latter is the most common cause of a 404 here.

Model aliases are resolved globally by name, the same way /v1/chat/completions, /v1/models, embeddings and RAG resolve them. Per-key restriction is the key's model allow-list, which realtime enforces exactly as chat does.

#2. Open the socket

GET /v1/realtime?session=<token>

const mint = await fetch("https://gateway.example.com/v1/realtime/sessions", {
  method: "POST",
  headers: { Authorization: `Bearer ${relayKey}`, "Content-Type": "application/json" },
  body: JSON.stringify({ model: "gpt-realtime", voice: "alloy" })
}).then(r => r.json());

const ws = new WebSocket("wss://gateway.example.com" + mint.url);

From here you are speaking the provider's realtime event protocol directly. The gateway relays frames rather than translating them: the realtime vocabulary is large and moves fast, and your client already speaks it. What the gateway owns is the parts that must not be client-controlled — the upstream credential, which model is reached, and the session configuration.

Client events you will use most:

EventPurpose
input_audio_buffer.appendSend a chunk of microphone audio (base64 PCM16).
response.createAsk for a reply — needed only when server turn detection is off.
conversation.item.createInject a text turn into the conversation.
conversation.item.truncateTell the provider how much of a reply the user actually heard before interrupting.

Server events worth handling:

EventPurpose
response.output_audio.deltaA chunk of speech to play (base64 PCM16).
response.output_audio_transcript.delta / .doneThe assistant's words as text.
response.output_text.delta / .doneA text-only reply, when audio output is off.
conversation.item.audio_transcription.completedWhat the user was heard to say.
input_audio_buffer.speech_startedThe user started speaking — see Handling an interruption below.
errorThe provider rejected something.

The GA event model prefixes response output events with output_ — response.audio.delta became response.output_audio.delta, and conversation.item.input_audio_transcription.completed lost its input_. A client written against the preview names does not error on GA: it connects, streams microphone audio, and plays nothing, because the audio deltas simply never match. The voice console listens for both spellings for exactly that reason, and so should your client.

#Handling an interruption

When the user talks over a reply, server turn detection fires input_audio_buffer.speech_started and the provider stops generating. Three things are then your client's job, and skipping any of them is audible.

Stop the audio you have already scheduled. This is the one that bites. Audio deltas arrive faster than real time, so a client that queues each chunk for smooth playback is holding seconds of speech that has not been played yet. Discarding the queue position is not enough — with the Web Audio API, a buffer passed to start(when) will play at that time whether or not you have moved your cursor, so the old reply keeps talking while the new one begins and the user hears two voices at once. Keep a handle on every buffer you schedule and stop them individually.

Ignore the deltas still in flight. The provider stops generating when it detects speech, but audio produced before that moment is already on the wire and will arrive after the interruption. Drop it by response_id rather than with a flag, so a delta belonging to the new reply is never discarded with it.

Tell the provider what was heard, with conversation.item.truncate naming the interrupted item_id, its content_index, and audio_end_ms — how many milliseconds the user actually heard. Without it the model's history contains the whole sentence it never finished saying, so it answers the next question as though the cut-off half had landed and will not repeat what the user genuinely missed. Clamp audio_end_ms to the audio the provider actually sent; a point beyond it is refused.

#Knowledge bases in a voice session

If the agent has a knowledge base, the gateway declares a search_knowledge_base function on the session and executes the call itself — your client does nothing. When the model decides it needs to look something up, the gateway embeds the query, searches the knowledge base, and returns the passages as a function_call_output, then sends response.create so the model carries on speaking.

This is deliberately not how the chat path works. There, the last user message is embedded and the top-k chunks are prepended to the system prompt. A realtime session has no user turn at the moment it is configured, and the turns that follow are speech the provider transcribes on its own side — so there is nothing to retrieve against up front. Pushing the whole knowledge base into the instructions instead does not scale past a handful of documents. As a tool, the model reaches for the reference when it needs it.

What this means in practice:

  • The agent's retrieval_top_k applies, as does its PII redaction setting — retrieved text is redacted before the model sees it, exactly as on the chat path. A passage read aloud is no less a disclosure than one written down.
  • Results are capped at 1,200 characters per passage and 6,000 in total. A voice model's context is small and every retrieved passage competes with the conversation for it.
  • A failed search still answers the call. If the vector store is unreachable the model receives an error result and is told to say it could not search, because an unanswered function call leaves the session waiting forever while the user talks into silence.
  • A gateway event is sent to your client for visibility: {"type": "relay.knowledge_search", "call_id": …, "query": …, "results": n}. It is not provider protocol — the relay. prefix marks it as gateway-originated, and clients that do not know it can ignore it. The voice console shows it in the transcript.
  • Only this tool is intercepted. Every other function call is forwarded to your client to execute, exactly as before. If the agent already declares a registry tool named search_knowledge_base, the gateway leaves it alone and skips retrieval rather than shadowing your client's tool — rename that tool to enable retrieval.

#The session object is nested in GA

Audio settings moved out of the session root and under session.audio. The gateway emits the GA shape:

{
  "type": "session.update",
  "session": {
    "type": "realtime",
    "instructions": "You are a helpful assistant.",
    "output_modalities": ["audio"],
    "audio": {
      "input":  { "format": { "type": "audio/pcm", "rate": 24000 }, "turn_detection": { "type": "server_vad" } },
      "output": { "voice": "alloy" }
    }
  }
}

Requesting a voice also declares output_modalities: ["audio"], so speech is not contingent on a server-side default — pass your own output_modalities through the session passthrough to change that.

If you send preview-era keys through the session passthrough — voice, modalities, turn_detection, input_audio_transcription, input_audio_format, output_audio_format, max_response_output_tokens — the gateway rewrites each to its GA location rather than letting the provider refuse the whole frame with Unknown parameter: 'session.voice'. An explicit GA value always wins over its legacy equivalent. One limit worth knowing: only the format string pcm16 is translated to a GA format object; any other codec name is relocated untranslated, so the provider reports a precise error instead of the gateway guessing a spelling it could not confirm.

#Audio format

Mono PCM16 at 24 kHz, base64-encoded inside the JSON event. getUserMedia gives you Float32 at the device's rate — usually 48 kHz — so capture has to be downsampled and converted before sending. Playback deltas come back in the same format and must be scheduled back-to-back on an explicit timeline; playing each chunk as it arrives makes network jitter audible as clicks.

#Session refused

A rejected connection is reported as a WebSocket close with a reason, not an HTTP error. A browser cannot read the body of a rejected upgrade, so a 401 would be invisible to your client. Read event.reason on close:

ReasonCause
Session token is invalid, expired, or already used.Older than 60 s, already redeemed, or wrong.
Provider '…' does not support realtime.The model's provider has no realtime capability.

#Provider endpoints

The gateway derives the realtime socket URL from the provider's Base URL, swapping the scheme.

ProviderDerived URL
OpenAI / OpenAI-compatible<base>/realtime?model=<model>
Azure (Foundry / Azure OpenAI)wss://<resource>.openai.azure.com/openai/v1/realtime?model=<deployment>

#Azure specifics

Azure's realtime endpoint differs from its chat endpoint in three ways, all of which produce a 404 on the WebSocket handshake if you get them wrong:

  1. The path is /openai/v1/realtime at the resource root — not under the inference base path. A Foundry base URL of …/models (what chat uses) would derive /models/realtime, which does not exist. The gateway replaces the base path rather than appending to it.
  2. No api-version. Microsoft's guidance is explicit: "use the GA endpoint with /openai/v1 in the URL. Don't use date-based API versions or the api-version query parameter." The gateway strips api-version from the base URL's query for realtime, keeping any other parameters.
  3. The target is model, and its value is the deployment name. The GA endpoint is the OpenAI-compatible surface, so it is ?model=, not ?deployment=. Set the catalogue model's provider model id to your Azure deployment name.

So a provider Base URL of https://<resource>.openai.azure.com/models?api-version=2024-05-01-preview and a provider model id of gpt-realtime-2.1 yields:

wss://<resource>.openai.azure.com/openai/v1/realtime?model=gpt-realtime-2.1

Authentication uses the api-key header, which the gateway already applies for Azure providers.

#Overriding the endpoint

Normally you do not need this. The derivation above is correct for OpenAI and for Azure's GA endpoint. Set Realtime endpoint on the provider (Providers → Edit) only when it is wrong for your deployment.

It accepts wss:///ws:// or https:///http:// (the scheme is swapped) and wins outright:

  • include {model} to place the deployment name yourself — for a resource still on the older preview shape:
  wss://YOUR-RESOURCE.openai.azure.com/openai/realtime?api-version=2024-10-01-preview&deployment={model}

Substitute your own hostname. Pasting an example's placeholder host produces a syntactically valid URL that dials a machine which does not exist — the gateway rejects obvious placeholders on save, but it cannot know your resource name.

  • omit {model} and the target is appended as model=.

Do not hardcode a model or deployment here. The endpoint belongs to the provider; the model is chosen per request — by the API caller, by an agent's configuration, or in the voice console. A fixed value would override all of them, so a session opened for one model would silently run on another. Saving such a URL is rejected, and any already stored is ignored in favour of the per-request model.

#Authentication against Azure

Azure serves realtime on two surfaces that authenticate differently, and a stored endpoint does not reliably say which one it is. The older preview path reads the resource key from an api-key header; the GA /openai/v1 surface is OpenAI-compatible and reads the same key as Authorization: Bearer — Microsoft's own GA sample hands an Azure key to the OpenAI SDK, which sends it as a bearer token.

The gateway therefore attempts both, in the order the endpoint's path implies (/openai/v1 → Bearer first, otherwise api-key first), and stops at whichever completes the handshake. Nothing needs configuring for this. A non-Azure provider gets a single Bearer attempt, since there is no second style to fall back to.

#When the handshake fails

A WebSocket handshake that is refused returns only a status code, and ClientWebSocket discards the response body — which is exactly where a provider explains itself. The gateway replays the failed upgrade once as a plain HTTP request purely to read that body, and includes it in the error as Provider said: …. If a refused upgrade carried no body at all, the same request is repeated without the upgrade headers, which usually does explain itself. The gateway then asks the resource which model ids it actually serves and appends them, because a model that is not a deployment name is the most common cause of a 400 on a correct route. All of it is best-effort and none of it is silent: a probe that fails says why, and the original handshake error is never replaced by one about the diagnostic.

A 400 on an otherwise correct URL is most often the model id. On Azure, model= must be the deployment name from the portal, which is not necessarily the name of the model you deployed — check the model's provider model id in the panel against the deployment list.
A failed handshake reports the exact URL that was attempted, which is the fastest way to tell a wrong path from a wrong deployment name. The credential travels as a header, so it never appears in that message.

#Session type

The GA event model selects the session pattern with session.type: realtime for voice-agent (speech-to-speech), transcription for speech-to-text. The gateway sets realtime unless you supply your own type through the session passthrough.

#Limits

Sessions last a maximum of 60 minutes. Audio is PCM16, mono, 24 kHz, and Microsoft recommends sending it in ~100 ms chunks. Realtime has its own quotas, separate from chat completions.

#Agent-backed sessions

POST /v1/agents/{id}/realtime/sessions

curl -X POST http://localhost:5300/v1/agents/AGENT_ID/realtime/sessions \
  -H "Authorization: Bearer $RELAY_API_KEY" -H "Content-Type: application/json" \
  -d '{ "voice": "sage" }'

The agent supplies the model and the instructions; you do not name a model. The gateway projects into the session:

An agent is visible to an API key when the agent has no workspace, or the same workspace as the key. This is the same rule GET /v1/agents and the agent run path apply, and it is the most common reason a valid-looking id fails: an agent created inside one workspace cannot be opened by a key belonging to another, even though the id is real and the agent is enabled. The error names which of the four causes applies — nonexistent, deleted, disabled, or a different workspace — so use a key created in the agent's workspace, or clear the agent's workspace to make it global.

A non-public agent adds a second rule on top of that one: it can only be opened by the users it has been shared with, so send the acting end user in X-Relay-User-Id and the mint refuses with 400 user_header_required without it, or 403 agent_not_shared for a user who is not on the list. It is checked at the mint deliberately, because that is the last point at which it can be — the socket that follows carries only the token, so the acting user is no longer in the picture. A caller that refuses cleanly here has not yet opened a microphone.

The voice console lists agents by calling GET /v1/agents with the key you paste, not from the workspace selected in the panel. Those are two different scopes, so listing from the panel would offer agents that minting then refuses. It names no acting user, so it can only offer public agents — a shared agent is reachable over the API but not from that console.

A session backed by an agent uses:

  • the agent's system prompt, with its skills folded in first — the same order the chat path uses, so a voice conversation behaves like a text one;
  • its registry tools, as function declarations.

The agent's own model must be realtime-capable. That is the point of attaching realtime to an agent: an operator picks the model once in the panel, and callers just name the agent.

Your instructions are appended, never substituted. An agent's instructions are the operator's policy, and letting whoever holds a session token replace them would make the agent meaningless.

#Two agent features that do not carry over

Both are stated plainly because silently ignoring them would be worse:

  • Retrieval is a tool call, not pre-loaded context. See knowledge bases in a voice session — the model chooses when to search, so a question it decides not to look up is answered from the model's own knowledge. The instructions tell it to search before answering on the knowledge base's subject matter, but that is guidance, not a guarantee.
  • MCP tools are not executed. An MCP tool needs the gateway to run the call and feed the result back, which the relay does not do — it forwards frames. Advertising tools nobody executes would strand the model mid-turn, so MCP servers are omitted from the session. Registry tool declarations are included, and your client executes them, exactly as it does for chat.

#Fallback, cost and limits

  • No model fallback. A realtime session is a stateful conversation; switching models mid-call would discard the audio context and change the voice. A failure ends the call instead. See routing & fallback.
  • The key's model allow-list applies, or realtime would be a way around it.
  • Cost is not yet recorded per session. The gateway relays frames and does not parse usage out of them, so a realtime call does not currently appear in telemetry or count towards a budget. Bill against your provider's own dashboard until this lands.
  • Sessions are single-instance. The token is held in memory on the instance that minted it, so mint and connect must reach the same process. Behind a load balancer that means sticky sessions.

#In the panel

Realtime voice (/realtime) is a working console: pick a realtime model or an agent, choose a voice, connect, and talk. It shows the live transcript and the raw protocol frames, which is the fastest way to see what your own client should be doing.