Relay docs

Embeddings

POST /v1/embeddings — OpenAI-compatible text embeddings. The request is routed to a model whose provider implements the embeddings capability.

#Request

{
  "model": "text-embedding-3-small",
  "input": "The quick brown fox",
  "encoding_format": "float"
}
  • input accepts a string or an array of strings.
  • dimensions and encoding_format are optional and forwarded to the provider.

#Response

{
  "object": "list",
  "data": [{ "object": "embedding", "index": 0, "embedding": [0.0012, -0.034, ...] }],
  "model": "text-embedding-3-small",
  "usage": { "prompt_tokens": 4, "total_tokens": 4 }
}

#Notes

  • If the routed model's provider doesn't support embeddings, the gateway returns 400 unsupported_capability.
  • The per-key model allow-list applies here too (403 model_not_allowed).
  • Responses can be cached (opt-in via Gateway:Cache:TtlSeconds); bypass with X-Relay-Cache: no. The X-Relay-Cache response header reports HIT/MISS.
  • Usage is recorded to telemetry like any other call.

Embeddings are the backbone of knowledge bases / RAG — the same routing is used internally when ingesting and searching documents.

One routing feature does not apply here: a model's fallback is ignored for embeddings and RAG. Two embedding models produce vectors in different spaces, so silently answering with a different one returns results that are wrong rather than degraded. An embedding request fails instead of substituting.