Evaluations
Evals let you measure prompt/model quality with repeatable test cases, and A/B compare a candidate against a baseline before you ship a change.
#Model
An EvalRun holds a set of test cases. Each case has an input and an expected result plus a scoring rule. You run the cases against a model — and optionally a variant model for A/B comparison — and Relay scores each output.
#Scoring rules
EvalProcessorService (a background service, polling Gateway:EvalPollSeconds, concurrency Gateway:EvalConcurrency) runs queued eval runs and scores each output with:
- contains — the output contains the expected substring.
- equals — the output exactly equals the expected value.
- regex — the output matches an expected pattern.
It aggregates a pass rate per model, so an A/B run shows baseline vs. variant side by side.
#Using evals
- Panel — the Evals & A-B page: define cases, pick a model (+ optional variant), run, and compare pass rates.
- Control-plane API —
api/Evals:POST search,GET {id},POST(create),POST {id}/cancel,DELETE {id}.
#A typical workflow
- Curate 20–50 representative cases for a task (with expected substrings / patterns).
- Run against your current production model → record the pass rate.
- Change a prompt version or try a cheaper model as the variant.
- Compare pass rates; promote the change (e.g. label a new prompt version
production) only if it holds up.
Pair this with the prompt library's versioning and the cost simulator to trade off quality against spend deliberately.
#Evals never fall back
A model's fallback is deliberately not applied to eval runs. An eval measures a named model, and an A/B run compares two of them — if a transient failure quietly routed a case to the fallback, the score would be attributed to a model that never answered it, and the comparison would be against a mixture. A failed case is honest; a substituted one is not.