Enhanced RAG Pipeline
Hrida.ai's RAG pipeline goes beyond simple embedding search. It stacks multiple complementary retrieval and ranking techniques so answers are grounded in the most relevant chunks — not just the most similar ones.
Query
│
├─► BM25 keyword retrieval ─┐
│ ├─► RRF merge ─► Reranker ─► Score filter ─► LLM
└─► Vector semantic search ─┘ │
Metadata filter (optional)
Every stage is independently configurable. You can run pure semantic search, enable only hybrid retrieval, or stack the full pipeline including reranking and metadata filtering.
Stage 1 — Chunking
Documents are split into chunks before they are embedded and stored. The text splitter controls how splits happen; the chunk size and overlap control how much text each chunk contains.
Text splitters
TEXT_SPLITTER | How it splits | Best for |
|---|---|---|
character (default) | Recursive character splitting — paragraphs, then sentences, then words, until chunks fit CHUNK_SIZE | General prose, PDFs, HTML |
token | By token count (tiktoken) | Precise context-window budgeting |
When ENABLE_MARKDOWN_HEADER_TEXT_SPLITTER is on (the default), Markdown content is first split on H1–H6 headings, then each section is split with the selected splitter. This keeps a section's heading with its text, which helps both BM25 and semantic retrieval.
Individual knowledge bases can override these settings — see Per-knowledge base chunking configuration.
Configuration
| Variable | Default | Description |
|---|---|---|
TEXT_SPLITTER | character | Splitter: character or token |
ENABLE_MARKDOWN_HEADER_TEXT_SPLITTER | True | Split Markdown on headings before applying the splitter |
CHUNK_SIZE | 1000 | Target size per chunk (characters, or tokens with token) |
CHUNK_OVERLAP | 100 | Amount shared between adjacent chunks (prevents boundary gaps) |
A chunk overlap of 10–25% of CHUNK_SIZE prevents answers that span a chunk boundary from being lost. For a CHUNK_SIZE of 1000, an overlap of 150–250 is a good starting point.
Stage 2 — Embedding and Indexing
After chunking, each chunk is converted to a vector using the configured embedding model and stored in the vector database alongside the chunk text, file metadata, and page number.
Batch embedding
Chunks are embedded in batches, not one at a time. The batch size controls throughput during ingestion.
| Variable | Default | Description |
|---|---|---|
RAG_EMBEDDING_BATCH_SIZE | 32 | Chunks submitted per embedding API call |
A batch size of 32 is approximately 30× faster than the original single-chunk-per-call approach for large knowledge bases. For very large documents (10,000+ chunks), increase to 64 or 128 if your embedding endpoint supports it.
Stage 3 — Retrieval (Hybrid Search)
At query time, Hrida.ai can run one or two retrievers in parallel.
Pure semantic search (default)
The query is embedded and the nearest neighbors are found in the vector store using cosine similarity. Returns the top RAG_TOP_K results by embedding similarity.
Limitation: Pure semantic search can miss exact identifiers, version numbers, error codes, and uncommon terms that are not well-represented in the embedding space.
Hybrid BM25 + semantic search
Enable with ENABLE_RAG_HYBRID_SEARCH=True.
Both retrievers run in parallel and their ranked result lists are merged using Reciprocal Rank Fusion (RRF):
RRF score = Σ 1 / (k + rank_in_list)
A chunk that ranks highly in both lists gets a significantly higher fused score than one that ranks highly in only one. This rewards chunks that are both semantically similar and keyword-matched.
| Variable | Default | Description |
|---|---|---|
ENABLE_RAG_HYBRID_SEARCH | False | Enable BM25 + semantic ensemble |
RAG_HYBRID_BM25_WEIGHT | 0.5 | BM25 contribution weight in RRF (0 = pure semantic, 1 = pure BM25) |
ENABLE_RAG_HYBRID_SEARCH_ENRICHED_TEXTS | False | Enrich BM25 index with surrounding context before indexing |
BM25 weight tuning
| Scenario | Recommended weight |
|---|---|
| General Q&A over prose documentation | 0.3–0.5 |
| Technical docs with many exact identifiers (env vars, API paths, error codes) | 0.5–0.7 |
| Multi-language or semantically rich knowledge bases | 0.2–0.3 |
| Keyword-heavy compliance/policy documents | 0.6–0.8 |
Enriched texts
When ENABLE_RAG_HYBRID_SEARCH_ENRICHED_TEXTS=True, each chunk is indexed in BM25 together with a small window of surrounding text from the same document. This improves recall for chunks that contain mostly stopwords or short phrases by adding the terminological context of their neighbors.
Hybrid search reliably improves retrieval accuracy — especially for technical documentation, error messages, and version-specific queries — with minimal latency overhead. Enable it as a baseline for any production knowledge base.
Stage 4 — Reranking
Reranking is a second-pass scoring step that re-scores the top-K candidates from retrieval using a dedicated cross-encoder model.
Why reranking works
Embedding similarity is a fast approximation: the query and chunk are scored independently and compared in vector space. A cross-encoder reads the query and chunk together in a single forward pass, which lets it detect subtle relevance signals that embedding similarity misses.
Reranking runs as part of hybrid search: enable ENABLE_RAG_HYBRID_SEARCH=True and set RAG_RERANKING_MODEL.
Reranking flow
Top RAG_TOP_K candidates from retrieval
│
└─► Cross-encoder model scores (query, chunk₁), (query, chunk₂), … (query, chunkₙ)
│
└─► Sort by reranker score ─► Keep top RAG_TOP_K_RERANKER ─► Drop below threshold
Reranker models
Set RAG_RERANKING_MODEL to any ColBERT-style or cross-encoder model available on your embedding server.
Recommended models:
| Model | Notes |
|---|---|
cross-encoder/ms-marco-MiniLM-L-6-v2 | Fast, small, good general-purpose reranker |
cross-encoder/ms-marco-MiniLM-L-12-v2 | Slightly better accuracy, similar speed |
BAAI/bge-reranker-base | Strong multilingual support |
BAAI/bge-reranker-large | Highest accuracy; requires more memory |
Configuration
| Variable | Default | Description |
|---|---|---|
RAG_RERANKING_MODEL | — | Model name/path for the reranker |
RAG_TOP_K | 3 | Candidates fetched from retrieval before reranking |
RAG_TOP_K_RERANKER | 3 | Chunks kept after reranking |
RAG_RELEVANCE_THRESHOLD | 0.0 | Minimum reranker score; chunks scoring below this are dropped |
Reranking in agent tool calls
Reranking is also active in the agent knowledge retrieval path. When an agent calls query_knowledge_files via a tool server, the full hybrid + rerank pipeline runs — not just a simple vector lookup.
Fetch more candidates at retrieval time (RAG_TOP_K = 10–20) so the reranker has enough material to find the best chunks. Then trim aggressively (RAG_TOP_K_RERANKER = 3–5) to keep the context window focused. A wide funnel into a narrow reranker output consistently outperforms fetching fewer candidates without reranking.
Stage 5 — Metadata Filtering
All retrieval paths support an optional metadata_filter dict that narrows results to chunks whose stored metadata matches specific key-value conditions.
Supported filter fields
Metadata stored per chunk typically includes:
source— the source file path or URLpage— page number (PDF/multi-page documents)file_id— the internal file UUIDcreated_at— ingestion timestamp
Filter mechanics
| Retrieval mode | Filter applied |
|---|---|
| Pure vector search | Pushed down to the vector DB query (fast, pre-fetch) |
| Hybrid BM25 + vector | Applied after RRF merge and reranking (BM25 cannot pre-filter) |
Programmatic usage (Python)
Metadata filtering is available to server-side code (for example, inside a Function or Tool) through query_collection. The REST endpoint POST /api/v1/retrieval/query/doc does not accept a filter.
from hrida_ai_studio.retrieval.utils import query_collection
results = await query_collection(
request,
collection_names=["hr-kb"],
queries=["paternity leave"],
embedding_function=request.app.state.EMBEDDING_FUNCTION,
k=5,
_metadata_filter={"source": "/docs/leave-policy.md"},
)Stage 6 — Citation Formatting
In regular chat, retrieved chunks are wrapped in <source id="N"> tags and the default RAG template asks the model to cite them inline as [1], [2], and so on; the chat UI turns those markers into clickable citations.
When knowledge is injected through a skill, each chunk is formatted with a citation header instead:
[hr-handbook.pdf, p.12] (score: 0.91)
Annual leave entitlement starts at 20 days per year for all full-time employees…
[leave-policy.md, p.3] (score: 0.83)
Requests for annual leave exceeding 10 consecutive days require line manager approval…
That block ends with an instruction to the model:
Cite sources inline as
[filename, p.N]when using information from the context above.
Deduplication
If the same chunk text appears in more than one knowledge base (e.g., a shared policy document indexed in both HR and Legal KBs), it is included only once.
Score threshold
Chunks scoring below RAG_RELEVANCE_THRESHOLD (default 0.0, so nothing is dropped) are suppressed before injection. Raise it to keep low-quality matches out of the context.
Performance Tuning Guide
Recommended production configuration
# Retrieval
ENABLE_RAG_HYBRID_SEARCH=True
RAG_HYBRID_BM25_WEIGHT=0.5
RAG_TOP_K=10
# Reranking (requires hybrid search)
RAG_RERANKING_MODEL=cross-encoder/ms-marco-MiniLM-L-6-v2
RAG_TOP_K_RERANKER=4
RAG_RELEVANCE_THRESHOLD=0.1
# Chunking
TEXT_SPLITTER=character
CHUNK_SIZE=1000
CHUNK_OVERLAP=200
# Ingestion speed
RAG_EMBEDDING_BATCH_SIZE=32
# Prompt caching (critical for follow-up latency)
RAG_SYSTEM_CONTEXT=TrueKV cache optimization
By default, RAG context is injected into the user message. As the conversation grows, the context position shifts, invalidating the LLM's KV prefix cache. Set RAG_SYSTEM_CONTEXT=True to inject context into the system message instead — its position never shifts, so the provider can cache the context after the first question.
Impact: For a 10-message conversation with 5 context chunks, RAG_SYSTEM_CONTEXT=True reduces average response latency by 50–80% from the second message onward (depending on provider caching support).
Supported providers: Ollama, OpenAI, Vertex AI, Anthropic, any provider with system-prompt prefix caching.
Latency breakdown
| Stage | Typical latency | Tuneable via |
|---|---|---|
| Embedding the query | 20–80 ms | Provider + model choice |
| Vector search | 5–50 ms | DB size, index type |
| BM25 search (hybrid) | 2–20 ms | Collection size |
| RRF merge | < 1 ms | — |
| Reranking (MiniLM-L6) | 30–100 ms | RAG_TOP_K (fewer candidates = faster) |
| Context injection | < 1 ms | — |
| Total RAG overhead | ~60–250 ms | See above |
Full Configuration Reference
| Variable | Default | Description |
|---|---|---|
TEXT_SPLITTER | character | Text splitter: character or token |
ENABLE_MARKDOWN_HEADER_TEXT_SPLITTER | True | Split Markdown on headings first |
CHUNK_SIZE | 1000 | Target size per chunk |
CHUNK_OVERLAP | 100 | Overlap between adjacent chunks |
RAG_EMBEDDING_BATCH_SIZE | 32 | Chunks embedded per batch call |
ENABLE_RAG_HYBRID_SEARCH | False | Enable BM25 + semantic ensemble |
RAG_HYBRID_BM25_WEIGHT | 0.5 | BM25 weight in RRF merge (0–1) |
ENABLE_RAG_HYBRID_SEARCH_ENRICHED_TEXTS | False | Enrich BM25 index with neighbor text |
RAG_TOP_K | 3 | Candidates fetched before reranking |
RAG_RERANKING_MODEL | — | Reranker model name or path |
RAG_TOP_K_RERANKER | 3 | Chunks kept after reranking |
RAG_RELEVANCE_THRESHOLD | 0.0 | Min reranker score; lower chunks are dropped |
RAG_SYSTEM_CONTEXT | False | Inject context into system message (enables KV caching) |
Troubleshooting
Retrieval returns irrelevant chunks
→ Enable hybrid search (ENABLE_RAG_HYBRID_SEARCH=True). Pure semantic search underperforms on exact identifiers and technical terms. If already enabled, lower RAG_HYBRID_BM25_WEIGHT (more semantic weight) for prose-heavy knowledge bases.
Reranker drops all chunks
→ RAG_RELEVANCE_THRESHOLD is too high. Lower it toward 0.0, or increase RAG_TOP_K so more candidates reach the reranker. Check that RAG_RERANKING_MODEL is correctly spelled and available on your embedding server.
Ingestion is slow for large knowledge bases
→ Increase RAG_EMBEDDING_BATCH_SIZE to 64 or 128. Also check that the embedding server is not CPU-bound; GPU-accelerated embedding is significantly faster for batch sizes above 32.
Follow-up questions are slow
→ Set RAG_SYSTEM_CONTEXT=True. The first question will be the same speed; every follow-up will be faster because the RAG context is cached by the LLM provider.
Chunks miss answers that span a page boundary
→ Increase CHUNK_OVERLAP. A value of 300–400 with a CHUNK_SIZE of 1000 catches most cross-boundary answers.
Agents don't use knowledge base in tool calls
→ Check that the knowledge base ID is registered in the agent model's knowledge_bases config. Reranking and hybrid search apply automatically to query_knowledge_files tool calls; no extra config needed.
Related
- RAG Overview — document sources, web search, configuration
- Knowledge Base — creating and managing knowledge bases
- Digital Employees — agents with per-agent knowledge bases
- Document Extraction — OCR and advanced PDF parsing