Observability (traces, scores, evals)
Relay records every request as a trace: a tree of observations covering the model calls, the agent run around them, each tool call, each retrieval, workflow steps, and any spans your app sends itself. Traces carry who they were for (end user, session), how they were tagged, what they cost, and how they were scored. Everything lives in ClickHouse; the panel reads it back under Observability.
Because Relay sits in front of every provider, an app gets full traces without adding any tracing code — it only adds a few headers to say who the request is for.
#What is recorded automatically
| Observation type | Recorded when |
|---|---|
generation | Every chat completion, streamed or not — including agent tool-loop iterations, workflow LLM nodes, batch rows, experiment cases and judge calls. |
embedding | Every /v1/embeddings call. |
agent | An agent run (/v1/agents/{id}/chat/completions); its generations, tools and retrieval nest inside it. |
tool | Each tool the gateway executes for an agent: MCP, Microsoft 365 connector, dataset SQL, workflow-as-tool. |
retriever | Knowledge-base search for an agent (query in, sources out). |
workflow / span | A workflow run, and each node it executes. A sub-workflow nests inside its caller. |
guardrail / event | Moderation verdicts, PII redaction, attachments. |
Each generation records the requested and resolved model, provider, prompt / completion / cached / reasoning tokens, cost, latency, time to first token, retry attempts, whether a fallback served it (and from which provider), whether it was a cache hit or an idempotent replay (both cost $0), the generation parameters, and the prompt name and version that produced it.
Realtime voice sessions are traced too: one span per call and one generation per completed response, with the provider's text and audio token usage.
#Attributing traces: users, sessions, tags
Send these headers (any OpenAI SDK can, through its extra-headers option). Every observation in the request inherits them.
| Header | Meaning |
|---|---|
X-Relay-User-Id | The end user of your app. Drives the End users page, per-user spend limits and erasure. |
X-Relay-Session-Id | Groups traces into a conversation (Sessions page). An agent called with ?thread= uses the thread id automatically. |
X-Relay-Tags | Comma-separated tags. |
X-Relay-Trace-Name | A readable name for the trace. |
X-Relay-Environment / X-Relay-Release | Deployment attribution (a key can carry a default environment). |
X-Relay-Metadata | A JSON object stored with the trace. |
X-Relay-Prompt | name@version of the library prompt you compiled, so its production metrics show on the Prompts page. |
X-Relay-Trace-Id / traceparent | Join an existing trace (yours, or one Relay returned). |
X-Relay-Parent-Span-Id | Nest this call under a specific observation. |
The OpenAI body fields work as well: user becomes the end user, and keys in metadata named session_id, trace_name, tags, environment, release, prompt_name, prompt_version (optionally prefixed relay_ or langfuse_) are read; other keys become metadata. Relay removes metadata before forwarding, because OpenAI rejects it on chat completions.
Every response returns X-Relay-Trace-Id and X-Relay-Observation-Id (plus the existing X-Relay-Request-Id), so your app can score the exact generation later.
#Capture: metrics only, or full bodies
By default Relay records metadata only: cost, tokens, latency, models, users, sessions, tags and the trace structure, but no message text. Turning capture to full also stores inputs and outputs, which enables the transcript view, "add to dataset" and LLM-judge evaluators.
| Level | Where it is set |
|---|---|
| Workspace default | Observability settings page. |
| Per API key | Key editor → Trace body capture. Overrides the workspace. |
| Per agent | Agent editor → Trace body capture. Overrides the key for that agent's runs. |
| Per request | X-Relay-Capture: metadata or none — a caller can lower capture, never raise it. |
Captured text is redacted for PII (emails, phone numbers, cards, IBANs) when the workspace setting is on, truncated to the workspace's character cap, and removed after the body retention period while the metrics stay. Captured text is other people's conversations: keep the default at metadata and enable full per key or per agent where you need it.
#Scores
A score is a quality signal attached to a trace, an observation or a session. Every score has a source:
| Source | Written by |
|---|---|
feedback / api | Your app, via POST /v1/scores (thumbs up/down, ratings, app-side checks). |
annotation | A person in the panel (trace view or annotation queue). |
judge | An LLM-as-a-judge evaluator. |
eval | An experiment run. |
Score configs (Scores → Score configs) define a score's type and allowed values — numeric with a range, boolean, or categorical with labels — so annotations and judges stay consistent. Scores appear on traces, sessions, users and prompt versions, and can be filtered on in the trace explorer.
#Annotation queues
Collect traces into a queue — from the trace list (select, Add to annotation queue) or a single trace — and pick which score configs reviewers fill in. Review walks through pending items one at a time, showing the input and output, with Complete & next and Skip.
#Evaluators (LLM-as-a-judge)
An evaluator scores live traffic continuously. It samples new generations that match its filter (agent, model, environment, tags, prompt, name) and have captured bodies, renders its judge prompt with {{input}}, {{output}} and {{metadata}}, asks the judge model for a verdict, and writes a judge score on the generation. Sampling is deterministic per generation, a per-run cap bounds cost, each generation is judged once, and judge and experiment traffic is never sampled. Judge calls are traced separately (evaluator: <name>) so production traces keep their real cost and latency.
#Datasets and experiments
A dataset is a curated set of test items — usually real traces worth keeping (Add to dataset on a trace or a selection), or imported JSONL/CSV. An experiment runs every active item through a model — optionally a variant model for A/B, and optionally a library prompt version — and scores each output by rule (contains, equals, regex) or with an LLM judge. Each case becomes its own trace with its score, so results can be opened, annotated and compared side by side on the dataset's Compare tab.
#Alerts
Workspace alert rules watch spend, request volume, error rate, p95 latency, p95 time to first token, a score's average, or error types not seen in the previous week — optionally narrowed to one key, model or environment. A rule notifies once when it breaches and once when it recovers, through the operator channel and the rule's own Slack/Teams-compatible webhook (https only, sent through the gateway's SSRF-guarded client).
#End-user spend limits
On a user's page, set a limit per day, week or month. With a hard stop, that user's requests are refused with 402 end_user_budget_exceeded once their spend (refreshed from telemetry every Gateway:BudgetRefreshMinutes) reaches it. These are keyed by your app's end-user id and are separate from panel member budgets.
#Retention, export and erasure
Retention is enforced by ClickHouse TTLs: rows can be deleted after N days, and captured text blanked after M days while metrics and scores stay (Observability settings). Export any window as JSON Lines or CSV from the trace list, or through GET /v1/telemetry/export. Erase data on a user's page deletes every observation and score for that end user (right to erasure); a trace can also be deleted from its page. Both are written to the audit log.
#Panel pages
| Page | What it shows |
|---|---|
| Traces | Every trace, with full-text search and filters (user, session, tags, environment, model, status, operation, prompt, score, cost, latency). Switch to generations or all observations. Live mode streams new traces. |
| Trace | The observation tree with a timing waterfall; for each observation its input/output (chat-rendered), metadata, parameters, usage, cost and errors; scores and annotation; add to dataset or queue; open in Playground; export; delete. |
| Sessions | Conversations with turn count, duration, cost and scores; a transcript view per session. |
| End users | Per-user traces, sessions, spend, errors and scores; per-user charts, models, spend limit and erasure. |
| Scores | Score analytics (averages, distributions, trends), every score, and score configs. |
| Annotation queues, Evaluators, Datasets, Alerts | As described above. |
| Observability settings | Capture default, PII redaction, body size cap and retention. |
The Dashboard adds traces, end users, sessions, cache-hit rate, p95 time to first token, top end users, score averages and spend by agent; the Prompts page has a per-version Metrics tab.
#Storage
| Store | Contents |
|---|---|
ClickHouse relay_observation | One row per observation. Created and evolved automatically at startup (missing columns are added). |
ClickHouse relay_score | Scores (ReplacingMergeTree: re-sending a score id updates it). |
SQL (relay_gateway) | Score configs, annotation queues, evaluators, datasets, alert rules, observability settings, end-user budgets — deployed from the DACPAC. |
The same spans are exported over OTLP when Gateway:Otlp:Enabled is on, with real trace, span and parent ids, so HyperDX shows the same tree. See Observability API for the endpoints.