Batch processing
Batches run large workloads asynchronously through the same routing pipeline as live calls. Submit rows once; Relay works through them in the background.
See the Batches API for the request shape.
#The processor
BatchProcessorService is a background service that, on each poll (Gateway:BatchPollSeconds, default 10):
- Spawns due recurring templates — any batch with a
cronschedule that's due creates a fresh run. - Picks the highest-priority queued job and processes its rows with bounded concurrency (
Gateway:BatchConcurrency, default 4, clamped 1–32).
#Reliability features
- Priority — higher-priority jobs run first.
- Checkpoint / resume — progress is checkpointed, so a restart continues where it left off rather than re-running completed rows.
- Dead-letter — rows that fail terminally are recorded to a dead-letter record instead of failing the whole job.
- Webhooks — supply a
webhookUrland Relay POSTs a completion notification when the job finishes. - Recurring (cron) — a template with a cron expression produces runs on schedule.
#Creating batches
- API —
POST /v1/batcheswith inline rows. - Panel — the Batch page accepts a CSV (one prompt per line) or JSONL upload, and lets you cancel jobs.
#Monitoring
GET /v1/batches/{id} returns status and counts (total / completed / failed); GET /v1/batches/{id}/results returns the outputs. The panel shows the same, live.
#Fallback
A row that fails transiently is retried against the model's fallback, if one is configured. This is evaluated per row, not per job — so a rate limit that clears part-way through lets later rows go back to the model you asked for, rather than committing the whole job to the fallback after the first failure.