Observability API
Endpoints for apps that call Relay: attribute requests, send feedback, add their own spans, and read back their workspace's traces and metrics. Everything is scoped to the calling key's workspace. See Observability for the concepts.
#Scopes
| Scope | Grants |
|---|---|
scores.write | POST /v1/scores, POST /v1/scores/batch, DELETE /v1/scores/{id} |
telemetry.read | GET /v1/scores, /v1/telemetry/*, POST /v1/metrics — metadata only |
telemetry.bodies | Also return captured prompts and completions (and allow full-text search) |
traces.write | POST /v1/traces/ingest, POST /v1/otlp/v1/traces |
As with every scope, a key with an empty scope list is unrestricted.
#Attribution on any request
Add headers to chat, agent, embeddings or responses calls. Every response carries X-Relay-Trace-Id and X-Relay-Observation-Id.
from openai import OpenAI
client = OpenAI(base_url="https://relay.example.com/v1", api_key="app_live_…")
resp = client.chat.completions.with_raw_response.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Where is my order?"}],
extra_headers={
"X-Relay-User-Id": "customer-8841",
"X-Relay-Session-Id": "chat-2026-09-27-001",
"X-Relay-Tags": "support,web",
"X-Relay-Environment": "production",
"X-Relay-Prompt": "support-answer@7",
},
)
trace_id = resp.headers["X-Relay-Trace-Id"]
observation_id = resp.headers["X-Relay-Observation-Id"]
completion = resp.parse()
const res = await client.chat.completions
.create({ model: "gpt-4o", messages }, { headers: { "X-Relay-User-Id": userId, "X-Relay-Session-Id": sessionId } })
.withResponse();
const traceId = res.response.headers.get("x-relay-trace-id");
Body alternative: "user": "customer-8841" and "metadata": {"session_id": "...", "tags": ["support"], "customer_tier": "gold"}. X-Relay-Capture: metadata (or none) lowers body capture for one request.
#Scores
POST /v1/scores — attach a score to a trace, an observation, a request id, or a session.
{
"trace_id": "4f1c…",
"observation_id": "9a2b…",
"name": "helpful",
"value": true,
"source": "feedback",
"comment": "Solved my problem"
}
| Field | Notes |
|---|---|
trace_id / observation_id / request_id / session_id | At least one. An observation or request id is resolved to its trace, user and session (a just-finished request may take a couple of seconds to become resolvable; trace_id never needs to wait). |
name | 1–200 characters. With config_id, the config's name is used. |
value | A number, a boolean (true/false, 1/0, "thumbs_up"), or a label for categorical scores. |
data_type | NUMERIC, BOOLEAN or CATEGORICAL; inferred from the value when omitted. |
config_id | Validate against a score config (range, allowed labels). |
source | feedback or api (default). Other sources are reserved for Relay. |
id | Optional idempotency id: re-sending it updates the score instead of adding another. |
POST /v1/scores/batch accepts up to 500 scores and reports per-item errors. GET /v1/scores?trace_id=…&session_id=…&user_id=…&name=…&source=…&from=…&to=…&page=0&limit=50 lists them. DELETE /v1/scores/{id} removes one.
#Reading telemetry
| Endpoint | Returns |
|---|---|
GET /v1/telemetry/summary?from&to&model&user_id&environment | Headline usage (requests, tokens, cost, latency percentiles, error rate, cache hits, traces, users, sessions) and a time series. |
POST /v1/telemetry/traces | Traces matching a filter (below), newest first, paged. |
GET /v1/telemetry/traces/{traceId} | Every observation in the trace, plus its scores. |
POST /v1/telemetry/observations | Individual observations (e.g. all failed generations). |
POST /v1/telemetry/sessions, GET /v1/telemetry/sessions/{id} | Sessions, and a session's turns. |
POST /v1/telemetry/users, GET /v1/telemetry/users/{id} | End users, and one user's usage. |
GET /v1/telemetry/export?from&to&type&user_id&session_id | Observations as JSON Lines (streamed). |
The filter body (all fields optional): from, to, search, traceId, name, userId, sessionId, tags (all must match), environment, release, model, provider, status (ok/error), operation, type, apiKeyId, agentId, promptName, promptVersion, level, minCost, maxCost, minLatencyMs, maxLatencyMs, scoreName, minScore, maxScore, hasBody, orderBy (time, cost, latency, tokens), desc, page, pageSize.
#Metrics
POST /v1/metrics runs a whitelisted aggregate: measures by dimensions, optionally bucketed in time.
{
"view": "observations",
"measures": ["count", "cost", "latency_p95", "error_rate"],
"dimensions": ["model"],
"granularity": "day",
"filters": { "from": "2026-09-01T00:00:00Z", "type": "generation", "environment": "production" },
"limit": 500
}
| View | Measures | Dimensions |
|---|---|---|
observations | count, traces, users, sessions, cost, tokens, prompt_tokens, completion_tokens, cached_tokens, latency_avg, latency_p50, latency_p95, latency_p99, ttft_p95, error_rate, cache_hit_rate | model, provider, user, session, environment, release, name, type, operation, status, api_key, agent, prompt, prompt_version, tag, trace, level, error_type, cache_hit |
scores | count, score_avg, score_min, score_max, traces, users | score_name, source, data_type, user, session, environment, config, value |
granularity is none, minute, hour, day, week or month. The response is { "columns": [...], "rows": [ { "time": …, "model": …, "count": … } ] }. Unknown names are rejected with the list of valid ones.
#Sending your own spans
POST /v1/traces/ingest adds observations from work your app does outside Relay — its own retrieval, tool calls or orchestration. Use the trace id Relay returned (or your own, sent on both calls) to put them in the same trace.
{
"batch": [
{
"trace_id": "4f1c…",
"id": "retrieval-1",
"parent_id": "9a2b…",
"type": "retriever",
"name": "vector search",
"start_time": "2026-09-27T10:00:00.120Z",
"end_time": "2026-09-27T10:00:00.340Z",
"input": { "query": "refund policy" },
"output": ["doc-17", "doc-42"],
"metadata": { "index": "policies-v3" },
"user_id": "customer-8841"
}
]
}
Types: generation, embedding, span, agent, tool, retriever, event, guardrail. Generations may carry model, provider, prompt_tokens, completion_tokens, cost_usd, prompt_name, prompt_version. Up to 1,000 per call; the response lists the accepted count and trace ids. Inputs and outputs follow the workspace capture policy.
#OpenTelemetry (OTLP)
Point an OpenTelemetry SDK's OTLP/HTTP JSON trace exporter at https://relay.example.com/v1/otlp/v1/traces with Authorization: Bearer app_live_…. Relay maps the GenAI semantic conventions (gen_ai.request.model, gen_ai.usage.input_tokens …), OpenInference (openinference.span.kind, input.value, llm.token_count.prompt …) and Langfuse (langfuse.session.id, langfuse.user.id, langfuse.trace.tags …) attributes onto observations; user.id, session.id, deployment.environment and service.version become attribution. Protobuf encoding is not accepted.