Relay docs

Cost

Estimate spend before you make calls, and read the catalog's pricing.

#Pricing catalog

GET /v1/pricing — per-model rates from the catalog: input/output cost per 1M tokens, context window, and max output tokens.

{ "data": [
  { "model": "gpt-4o", "input_per_1m": 5.00, "output_per_1m": 15.00, "context_window": 128000, "max_output": 16384 }
] }

#Simulate a cost

POST /v1/cost/simulate — estimate USD from token counts, or from sample text (a ~4-characters-per-token heuristic).

{ "model": "gpt-4o", "prompt_tokens": 1200, "completion_tokens": 400 }

or

{ "model": "gpt-4o", "prompt_text": "…a long prompt…", "expected_output_tokens": 500 }

Response:

{ "model": "gpt-4o", "prompt_tokens": 1200, "completion_tokens": 400, "input_cost": 0.006, "output_cost": 0.006, "total_cost": 0.012 }

#Where cost shows up

  • Every chat/embeddings response carries an X-Relay-Cost-Usd header.
  • Actual spend is recorded in telemetry and attributed per team/app/model.
  • The panel's Cost simulator page wraps /v1/cost/simulate; the Dashboard and Telemetry pages show real spend; Budgets enforce caps. See Telemetry & dashboards.