← All chapters

CHAPTER 07 / Knowledge

Make retrieval
observable

Advanced RAG begins when you inspect evidence systematically. The aim is not an answer that sounds confident; it is a repeatable system that retrieves the right support.

Enter the chapter
Original cover of chapter 07: Advanced RAG

CHAPTER 07

Advanced RAG

Tune retrieval as an engineering problem.

On this page

Faithful English web edition · Original chapter, structure and illustrations from the learning guide.

LEARNING OBJECTIVES

Tune retrieval with evidence, not intuition

By the end, you should be able to separate extraction, chunking, retrieval, reranking, context and model failures; change one variable at a time; measure whether retrieval improved; and select a sensible architecture from a personal homelab to a multi-worker deployment.

01. The advanced-RAG mindset

The question is no longer “Does RAG work?” It is “Why did this query retrieve these chunks rather than the ones that contain the answer?” Two failures can look identical in the final response:

FailureEvidenceCorrect response
Retrieval problemThe right passage never appears in Sources.Inspect extraction, chunks, query, index and ranking.
Model/prompt problemThe right passage arrived, but the answer misunderstood or ignored it.Inspect context trimming, instructions and model capability.
Debug in this order: stored text → retrieved Sources → tool calls → final answer.

Replacing an 8B model with a 70B model cannot repair a passage the retriever never supplied. It may only produce a more fluent answer from the same wrong evidence.

02. Chunking strategy is the foundation

Chunking strategy comparison showing good chunks, fragments that are too small, chunks that are too large and overlapEnlarge illustration ↗
Figure 1 · Chunk size changes what retrieval can understand and how much context it consumes.

Chunking decides which information should be retrievable together. A good chunk retains a meaningful unit: a paragraph with its heading, a definition with its explanation, a table row with its column labels or several procedure steps that depend on each other.

Too small

If one chunk contains “30 days” while another contains the heading “Notice period,” the number may be retrieved without enough meaning to answer safely.

Too large

If a two-line answer shares a chunk with three pages of unrelated text, retrieval may technically succeed while attention and context are wasted.

Overlap and document structure

Overlap protects ideas that cross a boundary, but also duplicates text and produces near-identical candidates. Markdown Header Splitting often gives technical documentation better semantic boundaries. A minimum chunk target can merge tiny header sections without exceeding the configured maximum.

Open WebUI Admin Documents configuration screenEnlarge illustration ↗
Admin → Documents exposes extraction, structural splitting, chunking, embeddings and retrieval settings in one place.

03. Hybrid Search: meaning plus exact words

Hybrid search and reranking combining semantic and BM25 candidate setsEnlarge illustration ↗
Figure 2 · Vector and BM25 search widen recall; a reranker narrows the result to useful evidence.

Vector search finds semantic similarity. BM25 finds lexical matches. A real corpus needs both: “What does the vacation policy say?” is semantic, while “Where is 0x80070005?” is exact.

ENABLE_RAG_HYBRID_SEARCH=true
RAG_HYBRID_BM25_WEIGHT=0.5

The documented BM25 weight is a balanced starting point, not a universal optimum. ENABLE_RAG_HYBRID_SEARCH_ENRICHED_TEXTS can also add metadata such as filename, title, section and snippets to lexical search, which helps when people ask using document names or headings.

04. Reranking, thresholds and Top K

Hybrid search casts a wider net. A cross-encoder reranker reads the query and each candidate together, then scores the pair more precisely. Because that is more expensive than first-stage search, it normally operates on a limited candidate pool.

retrieve broadly → rerank carefully → keep the best evidence

RAG_TOP_K=3
RAG_TOP_K_RERANKER=3
RAG_RELEVANCE_THRESHOLD=0.0

These documented defaults are conservative values for modest hardware, not quality targets. Top K is a context budget: too low loses supporting passages; too high invites noise. A relevance threshold can remove weak candidates, but an aggressive threshold can also discard the only useful passage.

05. The context budget

system prompt + chat history + retrieved chunks
+ tool results + answer budget ≤ context window

Raising Top K can appear to do nothing if the model’s effective context cannot hold the additional chunks. More available context also does not guarantee better reasoning: 128K of mostly irrelevant material can perform worse than a precise 8K evidence set.

Full Context remains valuable for a short, always-relevant document because it eliminates selection risk. For a large library, selective retrieval is mandatory. Choose on document size and question behaviour, not on the assumption that one mode is inherently more advanced.

06. Embeddings, reindexing and vector consistency

The embedding model defines the mathematical space occupied by every stored chunk. After changing that model, query vectors and old stored vectors are no longer comparable in a meaningful way.

Reindex does not reopen the original PDF or repeat OCR. If a new extractor is meant to repair bad text, re-upload the file. Standalone chat attachments also are not rebuilt by the Knowledge Base reindex operation and must be uploaded again.

07. Agentic retrieval, grep and kb_exec

Open WebUI Knowledge workspace used by agentic retrieval toolsEnlarge illustration ↗
Knowledge becomes an inspectable library when a capable model can choose semantic search, literal search and file reads.

Native function calling lets the model decide what to search, retry with another query and open specific files. Use query_knowledge_files for meaning and grep_knowledge_files for identifiers, versions or exact strings.

With experimental ENABLE_KB_EXEC=true, the model can use a filesystem-like interface with ls, tree, grep, cat, head, tail, find, wc and pipes.

Agentic retrieval is not automatically more accurate. The model can formulate a poor query, choose the wrong tool or misread an empty result. Full Context also differs from indexed Knowledge: injected full text does not behave like content stored for query_knowledge_files.

08. Evaluate RAG like an engineer

RAG evaluation framework separating retrieval quality, answer quality, latency and costEnlarge illustration ↗
Figure 3 · Measure retrieval and answer quality independently; change one variable at a time.

“It seems better” is not a test. Build a golden set of 10–50 real questions. For each, record the expected source document and passage. Include semantic questions, exact identifiers, ambiguous and multi-document cases, plus questions whose answer is absent.

Retrieval Hit@K
Does at least one correct chunk appear among the first K results?
Source correctness
Does the retrieved text actually support the answer rather than merely share words?
Groundedness
Does the response stay inside the supplied evidence?
Completeness
Does it include every necessary condition or supporting passage?
Citation quality
Do citations point to the chunks that support each claim?
Latency and cost
How much time and compute do embeddings, reranking, retrieval and generation add?

Change one variable at a time. Chunk size, overlap, embedding model, Hybrid Search, reranker, threshold and Top K interact; changing five at once destroys causal evidence.

Open WebUI Arena comparing two model responsesEnlarge illustration ↗
Arena can compare model answers, but a model leaderboard does not replace retrieval-specific evaluation.

09. Production architecture

RAG production architectures from personal use through small team, multi-worker and high throughputEnlarge illustration ↗
Figure 4 · Extraction, embeddings and vector storage should scale as separate workloads.
ScaleReasonable starting architecture
PersonalSingle Open WebUI, local ChromaDB, local embeddings and Ollama.
Small teamPlan RAM, backups and simultaneous ingestion; PostgreSQL + PGVector and external embeddings can simplify operations.
Multi-worker / replicasUse client-server vector storage such as PGVector, Milvus, Qdrant or MariaDB Vector, or a separate Chroma HTTP service.
High throughputSeparate extraction and embedding workloads from interactive chat; use shared durable storage and observed queues.

Local SQLite-backed Chroma is not safe as a shared store for multiple Uvicorn workers or replicas. Extraction also consumes memory and CPU, and a local SentenceTransformer may be loaded once per worker. Externalising embeddings avoids duplicating that model across every process.

10. External Knowledge Sources

Open WebUI can query a vector database you already operate instead of reingesting its documents. The documented experimental providers include Qdrant, Milvus and pgvector. Your infrastructure remains responsible for documents, embeddings and indexing; Open WebUI performs live retrieval.

The query embedding model and vector dimensions must match the ones used to build the external index. Otherwise similarity scores have no useful meaning. Map content, title, URL, document ID, page, metadata and score fields deliberately, then use the interface’s real test query before saving.

11. A practical tuning playbook

  1. Open stored text. If extraction is wrong, change the extractor and re-upload.
  2. Ask one known-answer question.
  3. Open Sources and locate the expected passage.
  4. If the correct file is missing, raise Top K modestly and repeat the same question.
  5. If exact codes or headings matter, enable Hybrid Search.
  6. If strong evidence is buried, add a reranker or calibrate the threshold.
  7. If an answer crosses boundaries, adjust chunk size, overlap or structural splitting.
  8. If Top K changes nothing, inspect the effective context window.
  9. Reindex after chunk or embedding changes; re-upload after extractor changes.
  10. Only then change the chat model if correct evidence arrived but the answer remained poor.

12. Troubleshooting by layer

SymptomLayer and action
File is empty or mangledExtraction: use OCR/Tika/Docling and re-upload.
Sources show the wrong sectionChunking/retrieval: inspect boundaries, Top K, embeddings, Hybrid Search and reranking.
Sources are right; answer is wrongModel/prompt: test clearer instructions or a stronger model.
Exact codes are missedLexical search: use BM25, grep or kb_exec.
Top K changes nothingContext budget: confirm extra chunks are actually included.
Retrieval broke after embedding changeVector mismatch: reindex Knowledge and re-upload chat files.
Worker crashes during ingestionArchitecture: do not share local SQLite-backed Chroma across workers.

PRACTICAL CHECKPOINT

Can you improve quality without guessing?

  • What makes a chunk retrievable but still useless?
  • What does BM25 add to vector search?
  • Why can a high Top K reduce quality?
  • How would you measure retrieval independently from the answer?
  • Why is local Chroma unsafe across multiple workers?
Check your answers

A chunk may lack its heading or contain too much noise. BM25 recovers exact terms. High Top K spends context on weak evidence. Hit@K and source correctness evaluate retrieval before generation. SQLite-backed local Chroma is not a multi-process shared vector service.

13. Essential vocabulary

Hybrid Search
Combined vector retrieval and BM25 keyword search.
Cross-encoder
A reranker that evaluates the query and candidate together.
Relevance threshold
The minimum score accepted for a candidate.
Hit@K
Whether at least one relevant item appears in the first K results.
Groundedness
How fully the answer is supported by retrieved evidence.
Reindex
Rebuild chunks and vectors from stored extracted text.
External Knowledge Source
An external vector store queried without reingesting its corpus.
kb_exec
Experimental filesystem-style Knowledge navigation and search.

Source and further reading

This edition preserves the chapter's teaching sequence and examples. Screenshots reflect the source edition; controls may move between releases.

Open the original chapter ↗Knowledge ↗RAG troubleshooting ↗Scaling ↗Evaluation ↗

Find your next step

Search chapter titles and section headings