Relay docs

Observability (traces, scores, evals)

Relay records every request as a trace: a tree of observations covering the model calls, the agent run around them, each tool call, each retrieval, workflow steps, and any spans your app sends itself. Traces carry who they were for (end user, session), how they were tagged, what they cost, and how they were scored. Everything lives in ClickHouse; the panel reads it back under Observability.

Because Relay sits in front of every provider, an app gets full traces without adding any tracing code — it only adds a few headers to say who the request is for.

#What is recorded automatically

Observation typeRecorded when
generationEvery chat completion, streamed or not — including agent tool-loop iterations, workflow LLM nodes, batch rows, experiment cases and judge calls.
embeddingEvery /v1/embeddings call.
agentAn agent run (/v1/agents/{id}/chat/completions); its generations, tools and retrieval nest inside it.
toolEach tool the gateway executes for an agent: MCP, Microsoft 365 connector, dataset SQL, workflow-as-tool.
retrieverKnowledge-base search for an agent (query in, sources out).
workflow / spanA workflow run, and each node it executes. A sub-workflow nests inside its caller.
guardrail / eventModeration verdicts, PII redaction, attachments.

Each generation records the requested and resolved model, provider, prompt / completion / cached / reasoning tokens, cost, latency, time to first token, retry attempts, whether a fallback served it (and from which provider), whether it was a cache hit or an idempotent replay (both cost $0), the generation parameters, and the prompt name and version that produced it.

Realtime voice sessions are traced too: one span per call and one generation per completed response, with the provider's text and audio token usage.

#Attributing traces: users, sessions, tags

Send these headers (any OpenAI SDK can, through its extra-headers option). Every observation in the request inherits them.

HeaderMeaning
X-Relay-User-IdThe end user of your app. Drives the End users page, per-user spend limits and erasure.
X-Relay-Session-IdGroups traces into a conversation (Sessions page). An agent called with ?thread= uses the thread id automatically.
X-Relay-TagsComma-separated tags.
X-Relay-Trace-NameA readable name for the trace.
X-Relay-Environment / X-Relay-ReleaseDeployment attribution (a key can carry a default environment).
X-Relay-MetadataA JSON object stored with the trace.
X-Relay-Promptname@version of the library prompt you compiled, so its production metrics show on the Prompts page.
X-Relay-Trace-Id / traceparentJoin an existing trace (yours, or one Relay returned).
X-Relay-Parent-Span-IdNest this call under a specific observation.

The OpenAI body fields work as well: user becomes the end user, and keys in metadata named session_id, trace_name, tags, environment, release, prompt_name, prompt_version (optionally prefixed relay_ or langfuse_) are read; other keys become metadata. Relay removes metadata before forwarding, because OpenAI rejects it on chat completions.

Every response returns X-Relay-Trace-Id and X-Relay-Observation-Id (plus the existing X-Relay-Request-Id), so your app can score the exact generation later.

#Capture: metrics only, or full bodies

By default Relay records metadata only: cost, tokens, latency, models, users, sessions, tags and the trace structure, but no message text. Turning capture to full also stores inputs and outputs, which enables the transcript view, "add to dataset" and LLM-judge evaluators.

LevelWhere it is set
Workspace defaultObservability settings page.
Per API keyKey editor → Trace body capture. Overrides the workspace.
Per agentAgent editor → Trace body capture. Overrides the key for that agent's runs.
Per requestX-Relay-Capture: metadata or none — a caller can lower capture, never raise it.

Captured text is redacted for PII (emails, phone numbers, cards, IBANs) when the workspace setting is on, truncated to the workspace's character cap, and removed after the body retention period while the metrics stay. Captured text is other people's conversations: keep the default at metadata and enable full per key or per agent where you need it.

#Scores

A score is a quality signal attached to a trace, an observation or a session. Every score has a source:

SourceWritten by
feedback / apiYour app, via POST /v1/scores (thumbs up/down, ratings, app-side checks).
annotationA person in the panel (trace view or annotation queue).
judgeAn LLM-as-a-judge evaluator.
evalAn experiment run.

Score configs (Scores → Score configs) define a score's type and allowed values — numeric with a range, boolean, or categorical with labels — so annotations and judges stay consistent. Scores appear on traces, sessions, users and prompt versions, and can be filtered on in the trace explorer.

#Annotation queues

Collect traces into a queue — from the trace list (select, Add to annotation queue) or a single trace — and pick which score configs reviewers fill in. Review walks through pending items one at a time, showing the input and output, with Complete & next and Skip.

#Evaluators (LLM-as-a-judge)

An evaluator scores live traffic continuously. It samples new generations that match its filter (agent, model, environment, tags, prompt, name) and have captured bodies, renders its judge prompt with {{input}}, {{output}} and {{metadata}}, asks the judge model for a verdict, and writes a judge score on the generation. Sampling is deterministic per generation, a per-run cap bounds cost, each generation is judged once, and judge and experiment traffic is never sampled. Judge calls are traced separately (evaluator: <name>) so production traces keep their real cost and latency.

#Datasets and experiments

A dataset is a curated set of test items — usually real traces worth keeping (Add to dataset on a trace or a selection), or imported JSONL/CSV. An experiment runs every active item through a model — optionally a variant model for A/B, and optionally a library prompt version — and scores each output by rule (contains, equals, regex) or with an LLM judge. Each case becomes its own trace with its score, so results can be opened, annotated and compared side by side on the dataset's Compare tab.

#Alerts

Workspace alert rules watch spend, request volume, error rate, p95 latency, p95 time to first token, a score's average, or error types not seen in the previous week — optionally narrowed to one key, model or environment. A rule notifies once when it breaches and once when it recovers, through the operator channel and the rule's own Slack/Teams-compatible webhook (https only, sent through the gateway's SSRF-guarded client).

#End-user spend limits

On a user's page, set a limit per day, week or month. With a hard stop, that user's requests are refused with 402 end_user_budget_exceeded once their spend (refreshed from telemetry every Gateway:BudgetRefreshMinutes) reaches it. These are keyed by your app's end-user id and are separate from panel member budgets.

#Retention, export and erasure

Retention is enforced by ClickHouse TTLs: rows can be deleted after N days, and captured text blanked after M days while metrics and scores stay (Observability settings). Export any window as JSON Lines or CSV from the trace list, or through GET /v1/telemetry/export. Erase data on a user's page deletes every observation and score for that end user (right to erasure); a trace can also be deleted from its page. Both are written to the audit log.

#Panel pages

PageWhat it shows
TracesEvery trace, with full-text search and filters (user, session, tags, environment, model, status, operation, prompt, score, cost, latency). Switch to generations or all observations. Live mode streams new traces.
TraceThe observation tree with a timing waterfall; for each observation its input/output (chat-rendered), metadata, parameters, usage, cost and errors; scores and annotation; add to dataset or queue; open in Playground; export; delete.
SessionsConversations with turn count, duration, cost and scores; a transcript view per session.
End usersPer-user traces, sessions, spend, errors and scores; per-user charts, models, spend limit and erasure.
ScoresScore analytics (averages, distributions, trends), every score, and score configs.
Annotation queues, Evaluators, Datasets, AlertsAs described above.
Observability settingsCapture default, PII redaction, body size cap and retention.

The Dashboard adds traces, end users, sessions, cache-hit rate, p95 time to first token, top end users, score averages and spend by agent; the Prompts page has a per-version Metrics tab.

#Storage

StoreContents
ClickHouse relay_observationOne row per observation. Created and evolved automatically at startup (missing columns are added).
ClickHouse relay_scoreScores (ReplacingMergeTree: re-sending a score id updates it).
SQL (relay_gateway)Score configs, annotation queues, evaluators, datasets, alert rules, observability settings, end-user budgets — deployed from the DACPAC.

The same spans are exported over OTLP when Gateway:Otlp:Enabled is on, with real trace, span and parent ids, so HyperDX shows the same tree. See Observability API for the endpoints.