Relay docs

Observability API

Endpoints for apps that call Relay: attribute requests, send feedback, add their own spans, and read back their workspace's traces and metrics. Everything is scoped to the calling key's workspace. See Observability for the concepts.

#Scopes

ScopeGrants
scores.writePOST /v1/scores, POST /v1/scores/batch, DELETE /v1/scores/{id}
telemetry.readGET /v1/scores, /v1/telemetry/*, POST /v1/metrics — metadata only
telemetry.bodiesAlso return captured prompts and completions (and allow full-text search)
traces.writePOST /v1/traces/ingest, POST /v1/otlp/v1/traces

As with every scope, a key with an empty scope list is unrestricted.

#Attribution on any request

Add headers to chat, agent, embeddings or responses calls. Every response carries X-Relay-Trace-Id and X-Relay-Observation-Id.

from openai import OpenAI

client = OpenAI(base_url="https://relay.example.com/v1", api_key="app_live_…")
resp = client.chat.completions.with_raw_response.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Where is my order?"}],
    extra_headers={
        "X-Relay-User-Id": "customer-8841",
        "X-Relay-Session-Id": "chat-2026-09-27-001",
        "X-Relay-Tags": "support,web",
        "X-Relay-Environment": "production",
        "X-Relay-Prompt": "support-answer@7",
    },
)
trace_id = resp.headers["X-Relay-Trace-Id"]
observation_id = resp.headers["X-Relay-Observation-Id"]
completion = resp.parse()
const res = await client.chat.completions
  .create({ model: "gpt-4o", messages }, { headers: { "X-Relay-User-Id": userId, "X-Relay-Session-Id": sessionId } })
  .withResponse();
const traceId = res.response.headers.get("x-relay-trace-id");

Body alternative: "user": "customer-8841" and "metadata": {"session_id": "...", "tags": ["support"], "customer_tier": "gold"}. X-Relay-Capture: metadata (or none) lowers body capture for one request.

#Scores

POST /v1/scores — attach a score to a trace, an observation, a request id, or a session.

{
  "trace_id": "4f1c…",
  "observation_id": "9a2b…",
  "name": "helpful",
  "value": true,
  "source": "feedback",
  "comment": "Solved my problem"
}
FieldNotes
trace_id / observation_id / request_id / session_idAt least one. An observation or request id is resolved to its trace, user and session (a just-finished request may take a couple of seconds to become resolvable; trace_id never needs to wait).
name1–200 characters. With config_id, the config's name is used.
valueA number, a boolean (true/false, 1/0, "thumbs_up"), or a label for categorical scores.
data_typeNUMERIC, BOOLEAN or CATEGORICAL; inferred from the value when omitted.
config_idValidate against a score config (range, allowed labels).
sourcefeedback or api (default). Other sources are reserved for Relay.
idOptional idempotency id: re-sending it updates the score instead of adding another.

POST /v1/scores/batch accepts up to 500 scores and reports per-item errors. GET /v1/scores?trace_id=…&session_id=…&user_id=…&name=…&source=…&from=…&to=…&page=0&limit=50 lists them. DELETE /v1/scores/{id} removes one.

#Reading telemetry

EndpointReturns
GET /v1/telemetry/summary?from&to&model&user_id&environmentHeadline usage (requests, tokens, cost, latency percentiles, error rate, cache hits, traces, users, sessions) and a time series.
POST /v1/telemetry/tracesTraces matching a filter (below), newest first, paged.
GET /v1/telemetry/traces/{traceId}Every observation in the trace, plus its scores.
POST /v1/telemetry/observationsIndividual observations (e.g. all failed generations).
POST /v1/telemetry/sessions, GET /v1/telemetry/sessions/{id}Sessions, and a session's turns.
POST /v1/telemetry/users, GET /v1/telemetry/users/{id}End users, and one user's usage.
GET /v1/telemetry/export?from&to&type&user_id&session_idObservations as JSON Lines (streamed).

The filter body (all fields optional): from, to, search, traceId, name, userId, sessionId, tags (all must match), environment, release, model, provider, status (ok/error), operation, type, apiKeyId, agentId, promptName, promptVersion, level, minCost, maxCost, minLatencyMs, maxLatencyMs, scoreName, minScore, maxScore, hasBody, orderBy (time, cost, latency, tokens), desc, page, pageSize.

#Metrics

POST /v1/metrics runs a whitelisted aggregate: measures by dimensions, optionally bucketed in time.

{
  "view": "observations",
  "measures": ["count", "cost", "latency_p95", "error_rate"],
  "dimensions": ["model"],
  "granularity": "day",
  "filters": { "from": "2026-09-01T00:00:00Z", "type": "generation", "environment": "production" },
  "limit": 500
}
ViewMeasuresDimensions
observationscount, traces, users, sessions, cost, tokens, prompt_tokens, completion_tokens, cached_tokens, latency_avg, latency_p50, latency_p95, latency_p99, ttft_p95, error_rate, cache_hit_ratemodel, provider, user, session, environment, release, name, type, operation, status, api_key, agent, prompt, prompt_version, tag, trace, level, error_type, cache_hit
scorescount, score_avg, score_min, score_max, traces, usersscore_name, source, data_type, user, session, environment, config, value

granularity is none, minute, hour, day, week or month. The response is { "columns": [...], "rows": [ { "time": …, "model": …, "count": … } ] }. Unknown names are rejected with the list of valid ones.

#Sending your own spans

POST /v1/traces/ingest adds observations from work your app does outside Relay — its own retrieval, tool calls or orchestration. Use the trace id Relay returned (or your own, sent on both calls) to put them in the same trace.

{
  "batch": [
    {
      "trace_id": "4f1c…",
      "id": "retrieval-1",
      "parent_id": "9a2b…",
      "type": "retriever",
      "name": "vector search",
      "start_time": "2026-09-27T10:00:00.120Z",
      "end_time": "2026-09-27T10:00:00.340Z",
      "input": { "query": "refund policy" },
      "output": ["doc-17", "doc-42"],
      "metadata": { "index": "policies-v3" },
      "user_id": "customer-8841"
    }
  ]
}

Types: generation, embedding, span, agent, tool, retriever, event, guardrail. Generations may carry model, provider, prompt_tokens, completion_tokens, cost_usd, prompt_name, prompt_version. Up to 1,000 per call; the response lists the accepted count and trace ids. Inputs and outputs follow the workspace capture policy.

#OpenTelemetry (OTLP)

Point an OpenTelemetry SDK's OTLP/HTTP JSON trace exporter at https://relay.example.com/v1/otlp/v1/traces with Authorization: Bearer app_live_…. Relay maps the GenAI semantic conventions (gen_ai.request.model, gen_ai.usage.input_tokens …), OpenInference (openinference.span.kind, input.value, llm.token_count.prompt …) and Langfuse (langfuse.session.id, langfuse.user.id, langfuse.trace.tags …) attributes onto observations; user.id, session.id, deployment.environment and service.version become attribution. Protobuf encoding is not accepted.