Relay docs

Knowledge bases (RAG)

A knowledge base grounds answers in your own documents. Each KB is a Qdrant collection scoped to a team. You ingest documents, and Relay retrieves relevant chunks at query time.

#Ingestion

RagService chunks each document, embeds the chunks (via the model router's embeddings), and upserts them into the vector store. You can ingest:

  • Raw content — POST /v1/knowledge/{id}/documents.
  • Files — POST /v1/knowledge/{id}/files (.md, .markdown, .mdx, .txt, .csv, .json, .log, .pdf, .docx; ≤ 25 MB). Text is extracted server-side by BCL-only extractors (Markdown syntax stripping, PDF content-stream walking, DOCX document.xml reading, HTML tag stripping). In the panel you can select several files at once and watch each one's progress and outcome.

#Ownership and visibility

A document belongs to the workspace that owns its knowledge base — not to the workspace of the API key that uploaded it. A key issued to one team ingesting into another team's knowledge base files the document under the knowledge base's owner, which is the only coherent answer when the two differ.

This matters because ingestion is a gateway call while the document list is a panel call, and the two see tenancy differently: the panel stamps rows with the workspace you have selected, whereas the gateway has no active workspace at all. Deriving the document's owner from its knowledge base is what keeps an ingested document visible in the knowledge base it was added to.

If a document was ingested before this behaviour was corrected, it may carry the uploading key's team instead. The panel reads a knowledge base's documents through the knowledge base itself rather than through each document's own team, so those rows still appear, and are realigned as they are touched. A telltale of the older behaviour is a knowledge base whose document count is above zero while the document table looks empty.

#Markdown

Markdown files are run through a dedicated extractor rather than indexed verbatim, because the markup itself would otherwise be embedded — [label](https://a/very/long/url) contributes URL tokens that carry no meaning, |---|---:| table rules are pure noise, and ### dilutes the heading text it decorates. Removed: link and image targets, link-definition lines, emphasis and inline-code markers, list bullets, blockquote markers, table delimiter rows, thematic breaks, HTML comments, and YAML/TOML front matter. Kept: heading text (usually the best summary of the section beneath it) and fenced code bodies (in technical docs the code is often the answer) — only the fence and its language tag are dropped.

  • URLs — POST /v1/knowledge/{id}/url (fetches and strips the page to text).

Manage the KBs and their document metadata in the panel under Knowledge.

#Retrieval

POST /v1/knowledge/{id}/search runs hybrid retrieval:

  1. Dense vector search returns candidates.
  2. BM25 lexical scoring is computed over the same candidates.
  3. Scores are min-max normalised and blended (Gateway:Rag:HybridAlpha, default 0.5 — dense vs. lexical weight).
  4. MMR re-ranks for relevance and diversity (Gateway:Rag:MmrLambda, default 0.7).

This BCL-only reranker (HybridReranker) improves recall over pure vector search without any extra service.

#Using RAG in an agent

Attach a KB to an agent and set a retrieval top-K. On each run the agent uses the effective last user message as the query, retrieves chunks, and folds them into the system prompt inside a --- CONTEXT --- block, instructing the model to cite sources inline as [S1], [S2], … So RAG "just works" for callers — no separate search call.

#PII

If the agent has PII redaction enabled, retrieved context (and user input) is passed through Guardrails.RedactPii before going upstream. See Guardrails & moderation.