Agents
An agent is a reusable, callable configuration: a model + a system prompt + generation defaults, optionally grounded on a knowledge base and equipped with skills, tools and MCP servers. Apps invoke it by id and get the full behaviour without re-assembling prompts each time.
#What an agent bundles
| Part | Purpose |
|---|---|
| Model | The public alias the agent runs on. |
| System prompt | Either free-text, or a reference to a library prompt (resolved production → latest at run time). |
| Skills | Reusable instruction blocks folded into the system prompt. |
| Tools / MCP servers | Advertised to the model; MCP-backed tools are executed by Relay. |
| Knowledge base | Optional RAG grounding — retrieved chunks are injected with [S#] citations. |
| Dataset | Optional structured-data grounding — the agent writes SQL and Relay runs it. See Dataset agents. |
| Generation defaults | Temperature, max tokens, retrieval top-K. |
| Guardrails | Optional PII redaction before text goes upstream. |
| Audience | Public (the default), or shared with named users — plus its creator, always. See below. |
#Who can run an agent
Two rules apply in order, and both must pass.
First, workspace: an agent is reachable by an API key when the agent is global (no workspace) or belongs to the key's workspace. That rule is unchanged and applies to every agent.
Second, audience. An agent is either public — any caller whose key can see it may run it, which is how every agent behaves by default — or shared with a specific set of users. A shared agent is runnable by its creator and by the users named on it, and the caller has to say which user it is acting for by sending the X-Relay-User-Id header. Without that header the run is refused with 400 user_header_required; with a user who is neither the creator nor on the list, 403 agent_not_shared. GET /v1/agents simply omits the agents the named user may not run, so a client can list and offer only what will actually work.
The creator always keeps access. Whoever created an agent can run it whatever its sharing says, and does not appear in — or need to be added to — the share list. This is why a private agent with nobody on it is still useful: it is a draft only its author can run. The owner is stamped once, server-side, from the signed-in panel user; it is never read from a request, or granting yourself permanent access to someone else's agent would be a matter of posting the right field.
Set this in the panel under Agents: open the agent's editor, switch Access to Specific users, and pick the people from the list.
#Which id a user is
The id an agent is shared by is the user's Entra object id (oid) — the same value the caller puts in X-Relay-User-Id, and the same one Relay forwards to the data app for a dataset agent's row-level security. That is what makes one header enough for an agent that is both private and dataset-backed.
It is deliberately not the panel's own application_user.id, which is a separate id ASP.NET Identity generates; the two are unrelated values for the same person. The panel's picker resolves the object id for you, so this only matters if you are writing share rows by some other route — in which case note the failure mode, because it is a quiet one: a share keyed by the wrong id saves successfully and then matches nobody, and the agent simply appears to be shared with no one.
A user who has never signed in with Microsoft has no object id at all, so there is nothing to match them by. The picker lists them greyed out rather than hiding them, so the reason is visible.
Two things worth knowing:
- The header is caller-asserted. Relay trusts the API key holder to name its end user, exactly as it already does for dataset agents. This controls which agent a trusted application may run on someone's behalf; it does not authenticate that person, and the application is still responsible for signing its own users in.
- Sharing is live, not versioned. A run resolves the audience from the current agent, before any published snapshot is applied. Making an agent private takes effect immediately, and publishing an older version never brings back an older audience.
An agent that is private with nobody on it is runnable by its creator alone. The panel labels that owner only rather than treating it as an error — it is exactly the state you want while building one. The exception is an agent created before owners were recorded: it has no creator, so private and unshared means nobody at all can run it, and the panel labels it nobody.
#Prompt source & variables
When creating an agent you either write a system prompt or select a library prompt. If the prompt has {{variables}}, callers supply them per run and Relay compiles the prompt before executing — pass a top-level variables object (see the Agents API). A chat-type prompt contributes its user/assistant turns to the conversation, so the caller's message can even be omitted.
#Composing the request
At run time the agent builds:
system = skills + prompt system text + retrieved KB context
messages = prompt seed turns + thread history + the caller's message
If there's no user turn anywhere, a self-contained system prompt is promoted to the user message so every provider runs.
#Versions & deployment
Agents are versioned like prompts. Snapshot the current definition, publish a version to deploy it, or restore an old version into the draft. The run endpoint prefers the published snapshot over live draft edits — so you can edit safely and deploy deliberately.
In the panel this is the agent's Versions tab: the draft and every snapshot sit in a timeline on the left (with its note, age and author), and the selected one shows its configuration as readable lines, a diff against the published version or any other, and the raw snapshot JSON. Access is left out of versions on purpose, since it is always the live setting.
#Running an agent
- API:
GET /v1/agentsto discover them,POST /v1/agents/{id}/chat/completionsto run one — see Agents API. - Voice: point an agent at a realtime model and it can be spoken to over a WebSocket, with its prompt, skills and tools applied — see Realtime.
- Panel: each agent has its own page. Overview shows its instructions (a library prompt is shown as the version it will actually serve), tools, knowledge, configuration, who can run it, and what a caller must send. Run is a multi-turn console that calls the agent through a picked API key (the panel never sees the secret) with its variables and acting user, and shows any SQL a dataset agent ran. Traces lists its recent runs and Use agent has cURL, Python and JavaScript snippets. Open in Playground opens the Playground with the agent already selected.
- Editing: creating, editing and duplicating happen on one full page with a section for each concern — basics, instructions, knowledge & data, tools, generation & safety, and access. Saving changes the draft only, so when a version is published callers keep getting it until you publish again.
- Memory: add
?thread=<id>for multi-turn conversations.