Relay docs

Relay AI Gateway

Relay is GMRL's internal AI gateway: a single, OpenAI-compatible endpoint that sits in front of every large-language-model provider your teams use. Apps talk to Relay; Relay handles routing, keys, cost, caching, guardrails, retrieval, agents and observability.

Think of it as the control plane for AI in the organisation — one place to manage who can call which models, how much they spend, and what the models are allowed to do.

#Why a gateway

Without a gateway, every app wires directly to OpenAI/Anthropic/Gemini with its own key, its own retry logic, and no shared visibility. Relay centralises all of that:

  • One API, many providers. Apps use the OpenAI wire format; Relay translates to OpenAI, Anthropic, Gemini, Azure Foundry or Ollama behind the scenes.
  • Governance. Scoped API keys, per-team budgets with hard stops, per-key model allow-lists, and a full audit trail.
  • Resilience. Transient-retry, provider health circuit-breaking, and single-hop fallback to a backup model.
  • Cost control. Live cost attribution per team/app/model, a cost simulator, and budget enforcement.
  • Higher-level building blocks. A versioned prompt library, reusable agents, skills, tools, MCP servers, knowledge bases (RAG), batch jobs and evaluations.
  • Observability. Full request telemetry with an in-app dashboard, plus standard OpenTelemetry export to ClickStack/HyperDX.

#The three surfaces

Relay is delivered as three things:

SurfaceWhat it isWho uses it
Gateway APIThe public /v1/* HTTP API (OpenAI-compatible).Your applications, authenticated with an API key.
Control panelA Blazor web app for administration.Platform admins and team owners, signed in with Entra.
This documentationThe site you're reading.Everyone.

#Where to go next

Relay speaks the OpenAI API. In most SDKs you only need to change the base URL to your Relay host and use a Relay API key — no other code changes.