Embeddings
POST /v1/embeddings — OpenAI-compatible text embeddings. The request is routed to a model whose provider implements the embeddings capability.
#Request
{
"model": "text-embedding-3-small",
"input": "The quick brown fox",
"encoding_format": "float"
}
inputaccepts a string or an array of strings.dimensionsandencoding_formatare optional and forwarded to the provider.
#Response
{
"object": "list",
"data": [{ "object": "embedding", "index": 0, "embedding": [0.0012, -0.034, ...] }],
"model": "text-embedding-3-small",
"usage": { "prompt_tokens": 4, "total_tokens": 4 }
}
#Notes
- If the routed model's provider doesn't support embeddings, the gateway returns
400 unsupported_capability. - The per-key model allow-list applies here too (
403 model_not_allowed). - Responses can be cached (opt-in via
Gateway:Cache:TtlSeconds); bypass withX-Relay-Cache: no. TheX-Relay-Cacheresponse header reportsHIT/MISS. - Usage is recorded to telemetry like any other call.
Embeddings are the backbone of knowledge bases / RAG — the same routing is used internally when ingesting and searching documents.
One routing feature does not apply here: a model's fallback is ignored for embeddings and RAG. Two embedding models produce vectors in different spaces, so silently answering with a different one returns results that are wrong rather than degraded. An embedding request fails instead of substituting.