Cost
Estimate spend before you make calls, and read the catalog's pricing.
#Pricing catalog
GET /v1/pricing — per-model rates from the catalog: input/output cost per 1M tokens, context window, and max output tokens.
{ "data": [
{ "model": "gpt-4o", "input_per_1m": 5.00, "output_per_1m": 15.00, "context_window": 128000, "max_output": 16384 }
] }
#Simulate a cost
POST /v1/cost/simulate — estimate USD from token counts, or from sample text (a ~4-characters-per-token heuristic).
{ "model": "gpt-4o", "prompt_tokens": 1200, "completion_tokens": 400 }
or
{ "model": "gpt-4o", "prompt_text": "…a long prompt…", "expected_output_tokens": 500 }
Response:
{ "model": "gpt-4o", "prompt_tokens": 1200, "completion_tokens": 400, "input_cost": 0.006, "output_cost": 0.006, "total_cost": 0.012 }
#Where cost shows up
- Every chat/embeddings response carries an
X-Relay-Cost-Usdheader. - Actual spend is recorded in telemetry and attributed per team/app/model.
- The panel's Cost simulator page wraps
/v1/cost/simulate; the Dashboard and Telemetry pages show real spend; Budgets enforce caps. See Telemetry & dashboards.