Relay docs

Evaluations

Evals let you measure prompt/model quality with repeatable test cases, and A/B compare a candidate against a baseline before you ship a change.

#Model

An EvalRun holds a set of test cases. Each case has an input and an expected result plus a scoring rule. You run the cases against a model — and optionally a variant model for A/B comparison — and Relay scores each output.

#Scoring rules

EvalProcessorService (a background service, polling Gateway:EvalPollSeconds, concurrency Gateway:EvalConcurrency) runs queued eval runs and scores each output with:

  • contains — the output contains the expected substring.
  • equals — the output exactly equals the expected value.
  • regex — the output matches an expected pattern.

It aggregates a pass rate per model, so an A/B run shows baseline vs. variant side by side.

#Using evals

  • Panel — the Evals & A-B page: define cases, pick a model (+ optional variant), run, and compare pass rates.
  • Control-plane API — api/Evals: POST search, GET {id}, POST (create), POST {id}/cancel, DELETE {id}.

#A typical workflow

  1. Curate 20–50 representative cases for a task (with expected substrings / patterns).
  2. Run against your current production model → record the pass rate.
  3. Change a prompt version or try a cheaper model as the variant.
  4. Compare pass rates; promote the change (e.g. label a new prompt version production) only if it holds up.

Pair this with the prompt library's versioning and the cost simulator to trade off quality against spend deliberately.

#Evals never fall back

A model's fallback is deliberately not applied to eval runs. An eval measures a named model, and an A/B run compares two of them — if a transient failure quietly routed a case to the fallback, the score would be attributed to a model that never answered it, and the comparison would be against a mixture. A failed case is honest; a substituted one is not.