Relay docs

Telemetry & dashboards

Every request through Relay is recorded: model, provider, tokens, cost, latency, status, team and API key. That data powers the in-app dashboards, cost attribution and alerting.

#Storage

Relay writes one row per observation — every model call, agent run, tool call, retrieval and workflow step — to relay_observation in ClickHouse, and scores to relay_score, whenever ClickHouse:ConnectionString is set. Both tables are created, and new columns added, automatically at startup. The dashboard, the Telemetry page and every Observability page read from them; see Observability for traces, sessions, users, scores, evaluators and datasets.

OTLP export to a ClickStack collector is independent (Gateway:Otlp:Enabled), so HyperDX can show the same traces. To keep reading the usage dashboard from the collector's otel_traces table instead, set Telemetry:DashboardSource = otel. The old Telemetry:Backend switch no longer stops Relay writing its own store, because traces, captured bodies and scores cannot live in otel_traces.

Builds before the observability store wrote a relay_usage table; Application.Database.Log/schema/05_backfill_relay_usage.sql copies that history into relay_observation.

#OpenTelemetry export

OtlpExporter is a BCL-only OTLP/HTTP JSON exporter (no OTel SDK). It's gated by Gateway:Otlp:Enabled and sends to Gateway:Otlp:Endpoint (traces), deriving the logs/metrics endpoints. It emits:

  • Traces — one span per observation, with the same trace, span and parent ids as the panel's trace tree, and gen_ai.* / relay.* attributes (model, tokens incl. cached and reasoning, cost, latency, status, team, key, user, session, environment, prompt, cache hit, attempts, fallback).
  • Logs — one record per request, correlated to its span.
  • Metrics — cumulative counters (relay.requests, relay.tokens, relay.errors, relay.cost.usd) and a latency gauge, every Gateway:Otlp:MetricsIntervalSeconds.

ClickStack/HyperDX authenticates OTLP ingestion — set Gateway:Otlp:ApiKey to the HyperDX Ingestion API Key, or the collector returns 401: missing or empty authorization header.

#In-app dashboards

The panel reads telemetry via api/Telemetry (configured, summary, timeseries, top-models, recent, attribution):

  • Dashboard — usage for the active workspace: KPI tiles (requests, cost, tokens, average cost per request, p50/p95 latency, error rate and count), a time series switchable between requests / tokens / cost / errors, top models, spend by provider, spend by app, prompt-vs-completion token split, and a cost attribution table groupable by model, provider or app.
  • Telemetry — a detailed request log with cost/usage attribution.

#Filtering

The dashboard filters by model, provider, app (API key) and status. Filter values come from the catalogue rather than from observed traffic, so a model that exists but has not been called yet is still selectable — otherwise "no traffic" and "not an option" would look identical.

End-user attribution comes from the X-Relay-User-Id header (or the OpenAI user field) that apps send. The End users page, the dashboard's top end users and the attribution tables group by it; requests from apps that send nothing are attributed to their API key only.

#Workspace scoping

Telemetry queries are hard-scoped to one workspace. The active workspace is taken from the X-Workspace-ID header the panel sends, it overwrites any teamId in the request body, and a request with no workspace returns an empty result rather than an unscoped one.

This matters more here than for the SQL entities: ClickHouse holds no row-level tenancy and there is no query filter behind these endpoints, so that scoping is the entire boundary between one workspace's usage and another's.

#Budgets & alerting

  • Budgets — BudgetRefreshService recomputes each workspace's spend from telemetry (Gateway:BudgetRefreshMinutes) for hard-stop enforcement. Only workspace-level rows are refreshed: the query filters by workspace, so running it over a per-member budget would write the whole workspace's spend into that member's row. See Budgets & spend limits.
  • Alerting — AlertingService watches error rate (and optional p95 latency) over a window and fires via AlertNotifier (logs + optional Slack-compatible webhook + optional SMTP). Configure under Gateway:Alerts:* and Gateway:HealthProbe:*.