Relay docs

The panel, page by page

The control panel is the Blazor admin app. Admins sign in with Microsoft Entra. Here's every page and what it's for.

#Operate

  • Dashboard (/) — usage for the active workspace over 24h / 7d / 30d, optionally compared with the previous period. Eight KPI tiles, a time series switchable between requests, tokens, cost and errors, top models, spend by provider and by app, a prompt-vs-completion token split, and a cost attribution table. Filter by model, provider, app or status. Exports the series as PNG or CSV.
  • Realtime voice (/realtime) — a spoken conversation with a realtime model or an agent. Pick a voice, connect, and talk; shows the live transcript and the raw protocol frames. Only realtime-capable models and agents pointed at one are offered. See Realtime.
  • Playground (/playground) — an interactive chat console. Pick a model or an agent, drop in an API key, and stream responses. Selecting an agent routes through the agent endpoint (prompt, skills, tools, RAG all apply) and shows variable inputs when the prompt needs them. Output renders as Markdown; you can optionally save the chat to Conversations.
  • Telemetry (/telemetry) — the detailed request log with cost/usage attribution across the gateway.

#Catalog

  • Providers (/providers) — upstream LLM endpoints. Secrets live in config (referenced by name), never here.
  • Models (/models) — the catalog callers request. A public alias maps to a provider model, with pricing, limits and an optional fallback model retried against on transient upstream failure (routing & fallback). Clicking a model opens a detail view with everything the catalog holds: identity and owning workspace, upstream mapping, pricing (with a worked per-request example), context and output limits, the resolved fallback chain, which other models fall back to it, 30-day usage pulled from telemetry, and the created/modified audit trail.

#Access & cost

  • API Keys (/api-keys) — issue keys with scopes, a model allow-list, a per-minute rate limit and an expiry date; edit any of those later without changing the secret. The full key is shown once at creation; after that only its prefix, because only a hash is stored. Revoke or delete anytime. See Authentication & keys.
  • Teams (/teams) — workspaces that own API keys, budgets and usage attribution.
  • Budgets (/budgets) — one workspace spend ceiling plus an optional allocation per member, where the member allocations can never sum past the ceiling. Hard-stop budgets block requests with a 402 once spend hits the cap; alert-only budgets just track it. See Budgets & spend limits.
  • Cost simulator (/cost) — estimate a call's cost from token counts or sample text, using catalog pricing.

#Build

  • Prompts (/prompts) — the versioned prompt library, grouped into folders by / in the name. Each prompt has its own page (/prompts/{name}) with a version timeline, labels, a live compile, a diff and fetch snippets; new prompts and versions are written on a full-page editor.
  • Agents (/agents) — compose an agent from a model, a prompt (written or from the library), skills, tools, MCP and an optional knowledge base. Each agent has its own page (/agents/{id}) with an overview, a run console, versions (snapshot, publish, restore, diff), its traces and API snippets; it is created and edited on a full-page editor.
  • Skills (/skills) — reusable instruction blocks attached to agents.
  • Tools (/tools) — the function-tool registry (OpenAI shape).
  • MCP servers (/mcp) — register MCP servers, discover their tools, and attach them to agents.
  • Knowledge (/knowledge) — per-team knowledge bases (Qdrant collections); ingest docs and run semantic search. The Add document dialog takes multiple files at once (Markdown, text, CSV, JSON, logs, PDF, DOCX) — click the drop area or drag files onto it — and shows a row per file: queued, uploading with a percentage, indexing, then the chunk count on success or the reason on failure, with an overall progress bar across the batch. Rows stay on screen after the run, so a failure part-way through is still readable.

#Run workloads

  • Batch (/batch) — queue large workloads; the gateway processes each row asynchronously. Accepts multiple .csv/.jsonl files at once, queueing one job per file (named after the file) so a failure is traceable to its source, with the same per-file progress and outcome display as the knowledge import.
  • Evals & A-B (/evals) — run test cases against a model (+ optional variant), score outputs, compare pass rates.
  • Conversations (/conversations) — persisted agent threads; call an agent with ?thread=<id> for multi-turn memory.

#Govern

  • Audit log (/audit) — every control-plane change: who created/updated/revoked/deleted what, and when.
  • Settings (/settings) — gateway connection health, provider status, and a reference of runtime config keys. Configuration portability exports one JSON bundle and imports one or more at a time, in the order shown, reporting each bundle's result separately.