Models
GET /v1/models — OpenAI-compatible listing of the models a caller can request. Only enabled models are returned, by their public alias name.
#Request
curl http://localhost:5300/v1/models -H "Authorization: Bearer $RELAY_API_KEY"
#Response
{
"object": "list",
"data": [
{ "id": "gpt-4o", "object": "model", "owned_by": "relay" },
{ "id": "claude-sonnet", "object": "model", "owned_by": "relay" }
]
}
The id is the public alias — the value you pass as model on chat/embeddings requests. Aliases decouple your apps from the underlying provider model id, so a model can be repointed (or upgraded) centrally without app changes.
#Related
- Manage the catalog (create aliases, map to provider model ids, set pricing, choose a fallback model) in the panel under Models — see The panel, page by page.
- A model marked Realtime speaks the bidirectional-audio protocol and is reachable through
/v1/realtimeinstead of/v1/chat/completions— the two are different protocols, so the flag decides which endpoint a model belongs to. - Health & rolling metrics:
GET /v1/providers/healthandGET /v1/models/metrics. - Dynamic remaps,
autoselection and per-model fallback: Model routing & fallback.
#Setting a model up from the provider
The model form can ask the provider what it serves, rather than making you remember a deployment name, and can fill the cost fields from Azure's published prices. Both run on the panel host, because the first needs the provider's credentials — nothing about them reaches the browser.
#Listing what a provider serves
List from provider on the model form calls the provider and offers the result as a picker. What comes back depends on how the provider is configured:
| Provider | Source | What the ids are |
|---|---|---|
| Azure Foundry, with the discovery fields filled | Azure management plane (ARM) | Your real deployment names, each labelled with the base model it runs and its provisioning state |
| Azure Foundry, without them | GET /openai/v1/models | The models the resource can serve — not deployment names |
| OpenAI, Ollama, other OpenAI-compatible | GET {base}/models | Model ids, which are what you call |
The Azure distinction matters and the panel states it every time, because it is invisible in the list itself. Azure has no data-plane endpoint that lists deployments: GET /openai/deployments was retired, and /openai/v1/models returns base and fine-tuned models. An inference call takes the deployment name, so an id from the fallback list may produce a 400 that looks like a broken deployment.
To get real deployment names, fill in the optional Azure deployment discovery fields on the provider: subscription id, resource group, account name, tenant id, client id, and a client secret. The secret is encrypted at rest with the same protector as the provider API key and is never returned to the browser. The service principal needs only read access to the account (Microsoft.CognitiveServices/accounts/deployments/read) — grant it Reader on the resource; nothing here writes.
The picker is a convenience, not a constraint: Provider model id stays editable, so an id the provider does not advertise can still be typed.
#Filling in the cost
Look up cost queries the public Azure retail price list for the model's token meters and offers what it finds. Choosing a candidate fills Input $/1M and Output $/1M and shows the exact meters the figures came from.
It offers candidates rather than one number because one number does not exist. A single model carries many meters:
- input and output are separate meters, and their naming is not consistent between model families —
Inp/Outpfor the gpt-4o family,inp/optfor gpt-realtime; - deployment tiers price differently: global, data zone, regional;
- realtime and multimodal models price each modality separately.
gpt-realtime-2.1ineastus2is $4.00/$24.00 per 1M for text but $32.00/$64.00 for audio — so a voice session costs eight times what the text price suggests. Pick the modality you will actually use.
Meters priced per 1K are normalised to per 1M. Cached-input, hourly and commitment meters are excluded: they are not the rate a request is billed at. The region comes from the resource itself when ARM discovery supplied it, since prices vary by region.
These are retail list prices, not your bill — they ignore any enterprise agreement, commitment or discount. The meter names are shown precisely so a figure can be checked before you rely on it, and both fields stay editable.