Agents API
POST /v1/agents/{id}/chat/completions — run a stored agent. The request body is a normal chat completion payload, but the agent supplies the model, system prompt (free-text or a library prompt), skills, tools/MCP and optional knowledge-base retrieval. The model field in your body is ignored.
See Agents for how to compose one.
An agent whose model is realtime-capable can also be spoken to — see Realtime.
An agent with a dataset pinned to it behaves differently in three ways: it requires the acting end user in a request header, it cannot stream, and its response carries a relay_dataset object with the SQL that ran and the rows it returned. See Dataset agents.
An agent that is not public also requires the acting end user, for a different reason: the gateway checks that user against the agent's creator and its share list. It is the same header either way, so an agent that is both private and dataset-backed needs one header, not two. The value is the user's Entra object id — see Agents for how the audience is set and why that id and not another.
#List agents
GET /v1/agents — the agents this key can run, so an app can discover them instead of hard-coding ids.
curl http://localhost:5300/v1/agents \
-H "Authorization: Bearer $RELAY_API_KEY" \
-H "X-Relay-User-Id: $END_USER_ID"
The user header is optional here even though a private agent's run requires it: listing without one returns the public agents, which is the useful answer for a client that has not yet picked a user.
{ "object": "list", "data": [
{
"id": "a1b2c3…",
"name": "support-assistant",
"description": "Answers returns and delivery questions",
"model": "gpt-4o",
"model_allowed": true,
"realtime": false,
"published_version": 3,
"grounded": true,
"retrieval_top_k": 4,
"skills": 2,
"tools": 1,
"mcp_servers": 0,
"dataset": false,
"public": true,
"requires_user_header": false,
"temperature": 0.3,
"max_tokens": 800,
"pii_redaction": false
}
] }
Only enabled agents appear, and only those that are global (no team) or owned by the key's team — the same visibility rule as the prompt library. Non-public agents appear only when X-Relay-User-Id names a user they are shared with.
A few fields are worth knowing:
| Field | Meaning |
|---|---|
model_allowed | false when the key's model allow-list excludes the agent's model, so a run would return 403 model_not_allowed. Reported rather than filtered on — an agent hidden for this reason looks deleted, while an unexplained 403 looks like a bug. |
realtime | Whether the agent's model speaks the bidirectional-audio protocol, so the agent can be opened as a voice session. Reported because it is otherwise unknowable: GET /v1/models does not expose the flag, so an app offering a "talk to an agent" picker could only discover the answer from a 400 unsupported_capability after the user had already chosen. Filter on this, unlike model_allowed — an agent that cannot be spoken to is not a permissions problem to explain, it simply has no voice option. |
published_version | The published version that will actually run. Null means the live draft is in effect, because a run prefers a published snapshot whenever one exists. |
grounded | Whether a knowledge base is attached. retrieval_top_k is present only when it is — otherwise it is a stored default that never takes effect. |
public | false when the agent is shared with named users rather than open to every caller. Because the listing already omits the ones this user may not run, a false here means "you can see this because the user you named created it or was shared it", not "you might be refused". |
requires_user_header | Whether a run needs X-Relay-User-Id. True for a dataset agent (to resolve per-user data access) or a non-public one (to check the share list); either reason is enough, and user_header gives the header name. |
skills / tools / mcp_servers | Counts, not ids. |
This is a directory, not a definition: the system prompt, the prompt id and the skill / tool / MCP ids are deliberately not returned. They are authoring detail, the prompt can carry sensitive instructions, and the ids name resources that may belong to another team. Nothing but id is needed to invoke an agent. Compose and inspect agents in the panel under Agents.
#Request
curl -X POST http://localhost:5300/v1/agents/AGENT_ID/chat/completions \
-H "Authorization: Bearer $RELAY_API_KEY" -H "Content-Type: application/json" \
-d '{
"messages": [{ "role": "user", "content": "How do I return an item?" }],
"stream": false
}'
- The message is optional. If the agent's prompt already contains the user turn (a chat-type prompt), or the system prompt is self-contained, you can omit it.
- Variables. If the agent's prompt has
{{variables}}, pass them as a top-levelvariablesobject; the gateway compiles the system prompt before running (the variables are not forwarded upstream):
{ "messages": [{ "role": "user", "content": "…" }],
"variables": { "company": "Spinneys", "language": "English" } }
#Conversation memory
Add ?thread=<id> to give the agent multi-turn memory. Prior turns for that thread are prepended, and this turn is persisted afterwards. The response includes X-Relay-Thread-Id.
POST /v1/agents/{id}/chat/completions?thread=my-thread-123
Threads are visible in the panel under Conversations.
#Tools & MCP
If the agent has MCP servers attached, the gateway runs a tool loop: it calls the model, executes any requested MCP tool, feeds the result back, and repeats up to Gateway:ToolMaxIterations (default 5). Registry tools (without an MCP backend) are advertised to the model, and their tool_calls are returned to you to execute. See Skills, tools & MCP.
#Streaming vs. tool loop
- Agents without MCP tools (and without
?thread=) stream like a normal chat call. - Agents with MCP tools (or using a thread) run non-streamed and return a single JSON completion.
#Response
Standard chat completion shape, plus X-Relay-Request-Id, X-Relay-Cost-Usd, and (for threads) X-Relay-Thread-Id.