Advanced RAG: chunking, hybrid search, reranking
Basic RAG (split → embed → retrieve → answer) works most of the time. It fails in three places: badly cut text, missed exact terms and weak ranking. In this lesson we fix all three — with Qdrant on your own machine and models that run locally.
modifier=IDF; Qdrant/bm25 stems words by English rules — for Bulgarian the stemmer is turned off; the model BAAI/bge-reranker-v2-m3 is not supported by fastembed — the old code stopped with an error, it now goes through sentence-transformers; rr'…' was a syntax error; breakpoint_threshold_amount=0.85 in percentile mode should be around 95; recreate_collection is deprecated; Ollama's /api/embeddings is superseded by /api/embed; the Qdrant image has no curl, so the old healthcheck always failed. The langchain-experimental package (with SemanticChunker) is sunset since May 2026 — we give a short own implementation. The "Art. 45 SIC" and "clinical pathway No. 45" examples were inaccurate — replaced with verified articles (Arts. 40–41 of the Social Insurance Code, Art. 110 of the Obligations and Contracts Act). The figures "+31% recall", "+74% accuracy", "23% → 97%", "48 ms / 312 ms" and "10M documents in 8 GB" had no source — removed or marked as illustrative. A wrong arXiv link is fixed.
01What you'll learn
- Why bad chunking is the most common reason RAG "doesn't find it", and how to pick a strategy by document type.
- How to split a law by article, text by meaning, and a long document into "parent and child".
- How to combine semantic (dense) and exact (BM25) search in Qdrant and fuse them with RRF.
- How to restrict results by course, law or date with filters — without anything leaking between users.
- How a reranking model (cross-encoder) orders 20 candidates so only the best 5 reach the model.
02Before you start
- You have done Blocks 0–3 and have a working basic RAG: you know what an embedding and a vector database are.
- Python 3.10 or newer, a virtual environment and Docker.
- Ollama 🔒 local with
nomic-embed-text(for embeddings) and a model for answers, e.g.llama3.1:8b. - A few GB of free space for the reranking model — it downloads from Hugging Face 🌐 global on first run.
python3 -m venv .venv && source .venv/bin/activate
pip install -U qdrant-client fastembed langchain-text-splitters langchain-ollama sentence-transformers numpy requests
# checked with: qdrant-client 1.19.1 · fastembed 0.8.1 · langchain-text-splitters 1.1.2 · langchain-ollama 1.1.0
ollama pull nomic-embed-text
ollama pull llama3.1:8b03Steps
-
Why basic chunking fails
Fixed-size chunking (say, every 500 characters) doesn't know where a thought ends. A law article can land half at the end of one chunk and half at the start of the next — then no chunk contains the whole rule, retrieval returns a fragment and the model answers incompletely or wrongly. Why does it matter? Because no better model can fix information that never reached it.
example · one article cut in two (teaching example, Bulgarian text)Чънк 1: …Чл. 40. (5) Работодателят изплаща първите два работни дни от временната неработоспособност в размер 70 на сто от среднодневното Чънк 2: брутно възнаграждение… Чл. 41. (1) Паричното обезщетение е 80 на сто от среднодневното брутно възнаграждение… Question: "How much does the employer pay for the first days of sick leave?" → Chunk 1 lacks the base ("gross pay"), chunk 2 starts with a different article.💡The facts in the example (checked with the NSSI)Since 2024 the employer pays the first two working days at 70%, and the National Social Security Institute pays a benefit of 80% of the average daily gross pay (Arts. 40–41 of the Social Insurance Code). The text above is shortened for the lesson, not a verbatim quote. -
Five chunking strategies
There is no single right strategy — it depends on the document. Start with recursive splitting and move to something more complex only when the check (the "Check" section and lesson 01-04b) shows it isn't enough.
Strategy How it works Good for Fixed size + overlap N characters, K overlap. Fast and simple. General prose, news. Bad for laws. Recursive ★ Splits on blank line → newline → ". " — keeps paragraphs. Most documents. A good start. Semantic Compares neighbouring sentences and cuts where the meaning "jumps". Mixed texts without clear structure: reports, opinions. Parent–child Small chunks for search, a big "parent" as context for the model. Long technical documents. Structure-aware ★ Recognises headings (##), articles ("Чл. N."), tables. Laws, regulations, standards. Python · advanced_chunking.pyimport re import numpy as np import requests from langchain_text_splitters import RecursiveCharacterTextSplitter, MarkdownHeaderTextSplitter OLLAMA = "http://localhost:11434" # ── 1. Recursive: a good default for most Bulgarian texts ── def recursive_chunks(text: str, chunk_size: int = 700, overlap: int = 120) -> list[str]: """Splits on \n\n → \n → ". " → space. Size is in characters, not tokens.""" splitter = RecursiveCharacterTextSplitter( chunk_size=chunk_size, chunk_overlap=overlap, separators=["\n\n", "\n", ". ", " ", ""], ) return splitter.split_text(text) # ── 2. Structure-aware: a law → one article = one chunk ── ARTICLE = re.compile(r"^(Чл\.\s*(\d+[а-я]?)\..*?)(?=^Чл\.\s*\d+[а-я]?\.|\Z)", re.S | re.M) def legal_law_chunks(law_text: str) -> list[dict]: """Every article starting a line ("Чл. 110.") becomes its own chunk with the number in metadata.""" return [{"text": m.group(1).strip(), "article": m.group(2)} for m in ARTICLE.finditer(law_text) if len(m.group(1).strip()) > 20] # ── 3. Semantic: cut where the meaning "jumps" ── def embed(texts: list[str], prefix: str = "search_document: ") -> np.ndarray: """nomic-embed-text via Ollama /api/embed (batch).""" r = requests.post(f"{OLLAMA}/api/embed", timeout=120, json={"model": "nomic-embed-text", "input": [prefix + t for t in texts]}) r.raise_for_status() return np.array(r.json()["embeddings"]) def semantic_chunks(text: str, percentile: float = 95, embed_fn=embed) -> list[str]: """Cut between two sentences when their distance is above the given percentile. Higher percentile = fewer cuts = bigger chunks.""" sentences = [s for s in re.split(r"(?<=[.!?])\s+", text) if s.strip()] if len(sentences) < 3: return [text] v = embed_fn(sentences) v = v / np.linalg.norm(v, axis=1, keepdims=True) dist = 1 - (v[:-1] * v[1:]).sum(axis=1) # cosine distance between neighbouring sentences cut = np.percentile(dist, percentile) chunks, current = [], [sentences[0]] for sentence, d in zip(sentences[1:], dist): if d > cut: chunks.append(" ".join(current)) current = [sentence] else: current.append(sentence) chunks.append(" ".join(current)) return chunks # ── 4. Parent–child: search by the small one, pass the big one ── def hierarchical_chunks(text: str) -> list[dict]: parent_splitter = RecursiveCharacterTextSplitter(chunk_size=1500, chunk_overlap=200) child_splitter = RecursiveCharacterTextSplitter(chunk_size=300, chunk_overlap=50) result = [] for i, parent in enumerate(parent_splitter.split_text(text)): for child in child_splitter.split_text(parent): result.append({"child_text": child, # small → for embedding and search "parent_text": parent, # large → for the model "parent_id": i}) return result # ── 5. Markdown: split by headings and keep them as metadata ── def markdown_chunks(md_text: str) -> list[dict]: splitter = MarkdownHeaderTextSplitter( headers_to_split_on=[("#", "h1"), ("##", "h2"), ("###", "h3")], strip_headers=False, ) return [{"text": d.page_content, **d.metadata} for d in splitter.split_text(md_text)]⚠️Characters, not tokens — and a short contextchunk_sizeinRecursiveCharacterTextSplittercounts characters. Cyrillic usually breaks into more tokens than English, so don't size chunks by eye. In the Ollama librarynomic-embed-texthas a 2K-token window (the model itself supports up to 8192) — keep chunks small, or their ends get lost.⚠️Old code on the internetMany examples useSemanticChunkerfromlangchain-experimental. The package has been sunset since May 2026 and is not maintained — so semantic chunking above is our own 20-line function. In percentile mode the threshold is a number from 0 to 100 (default 95), not 0.85.⚖️Illustrative scenario: a law by articlePicture a knowledge base with the Bulgarian Obligations and Contracts Act cut into 500-character pieces. Asked "What is the general limitation period?", retrieval returns the end of Art. 109 and the start of Art. 110 — without the rule that after a five-year limitation period all claims for which the law sets no other period are extinguished. With article-level chunking Art. 110 is whole and carriesarticle="110"in its metadata, so the answer can cite it. The old lesson claimed "23% → 97% accuracy" — we could not verify the source, so the figures are removed. Measure on your own documents. -
Hybrid search: meaning + exact words
Semantic search (dense vectors) is great for questions like "who pays when I'm sick", but misses exact identifiers: "Art. 41", a law's abbreviation, an order number. BM25 (sparse) finds exactly those words but doesn't understand paraphrase. Hybrid search runs both and fuses the results.
Query: "sick-leave benefit Art. 41 SIC" What it finds What it misses Dense (nomic-embed-text) "temporary incapacity", "cash benefit" — related in meaning The exact "Art. 41" Sparse (BM25) Exactly "Art. 41", "SIC" Texts with the same meaning but other words Hybrid (RRF) Documents ranked high in both lists — ✅How RRF fuses the two listsReciprocal Rank Fusion looks only at positions, not scores: each document gets Σ 1/(k + position) from each list. So it doesn't matter that dense and BM25 scores live on different scales. In Qdrant positions start at 0 and by default k = 2; since v1.16 you can set k (the classic paper uses 60), and since v1.17 also a weight per list.Run Qdrant in Docker. The ports are bound to the local interface only, and the key comes from
.env.YAML · docker-compose.yml# docker-compose.yml — Qdrant on your own machine services: qdrant: image: qdrant/qdrant:v1.19.1 # pinned version (latest as of 01.10.2026) restart: unless-stopped ports: - "127.0.0.1:6333:6333" # REST — local only - "127.0.0.1:6334:6334" # gRPC volumes: - ./qdrant_storage:/qdrant/storage environment: QDRANT__SERVICE__API_KEY: ${QDRANT_API_KEY} # from .env healthcheck: # the image has no curl or wget test: ["CMD", "bash", "-c", "exec 3<>/dev/tcp/127.0.0.1/6333 && printf 'GET /readyz HTTP/1.0\\r\\n\\r\\n' >&3 && grep -q ' 200 ' <&3"] interval: 15s timeout: 5s retries: 5⚠️The Qdrant image has no curlThe old checkcurl -f …/healthzalways fails, because curl and wget were removed from the image. The check above usesbashand/readyz, which returns 200 only once Qdrant can serve requests. We have not run it on a live machine — confirm withdocker psthat the status ishealthy.Python · qdrant_hybrid.pyimport os import requests from fastembed import SparseTextEmbedding from qdrant_client import QdrantClient, models OLLAMA = "http://localhost:11434" COLLECTION = "academy_bg" client = QdrantClient(url=os.getenv("QDRANT_URL", "http://localhost:6333"), api_key=os.getenv("QDRANT_API_KEY")) # the key comes from .env, not from code # ── 1. A collection with two kinds of vectors ── if not client.collection_exists(COLLECTION): client.create_collection( collection_name=COLLECTION, vectors_config={ "dense": models.VectorParams(size=768, distance=models.Distance.COSINE, on_disk=True), # nomic-embed-text = 768 }, sparse_vectors_config={ "sparse": models.SparseVectorParams( index=models.SparseIndexParams(on_disk=True), modifier=models.Modifier.IDF, # required for BM25 ), }, ) # index the fields you filter on — otherwise filtering is slow client.create_payload_index(COLLECTION, "course_id", field_schema=models.KeywordIndexParams(type="keyword", is_tenant=True)) # ── 2. Embeddings ── def get_dense(text: str, kind: str = "document") -> list[float]: """nomic-embed-text needs a prefix: search_document: when indexing, search_query: when searching.""" prefix = "search_query: " if kind == "query" else "search_document: " r = requests.post(f"{OLLAMA}/api/embed", timeout=60, json={"model": "nomic-embed-text", "input": prefix + text}) r.raise_for_status() return r.json()["embeddings"][0] # Qdrant/bm25 has no Bulgarian stemmer → turn it off (otherwise it stems by English rules) bm25 = SparseTextEmbedding(model_name="Qdrant/bm25", disable_stemmer=True) def get_sparse(text: str, kind: str = "document") -> models.SparseVector: emb = next(bm25.query_embed(text) if kind == "query" else bm25.embed([text])) return models.SparseVector(indices=emb.indices.tolist(), values=emb.values.tolist()) # ── 3. Indexing ── def index_document(doc_id: int, text: str, metadata: dict): client.upsert(COLLECTION, points=[models.PointStruct( id=doc_id, vector={"dense": get_dense(text), "sparse": get_sparse(text)}, payload={"text": text, **metadata}, )]) # ── 4. Hybrid search: two prefetches + RRF inside Qdrant ── def hybrid_search(query: str, top_k: int = 5, filters: dict | None = None) -> list[dict]: flt = None if filters: flt = models.Filter(must=[models.FieldCondition(key=k, match=models.MatchValue(value=v)) for k, v in filters.items()]) res = client.query_points( collection_name=COLLECTION, prefetch=[ models.Prefetch(query=get_dense(query, "query"), using="dense", limit=20, filter=flt), models.Prefetch(query=get_sparse(query, "query"), using="sparse", limit=20, filter=flt), ], query=models.FusionQuery(fusion=models.Fusion.RRF), # or models.RrfQuery(rrf=models.Rrf(k=60)) limit=top_k, with_payload=True, ) return [{"text": p.payload.get("text", ""), "score": p.score, "metadata": {k: v for k, v in p.payload.items() if k != "text"}} for p in res.points] if __name__ == "__main__": for r in hybrid_search("обезщетение за болничен чл. 41 КСО", top_k=5, filters={"course_id": "42"}): print(f"{r['score']:.3f} | {r['text'][:100]}")⚠️Three traps in hybrid search1. Withoutmodifier=IDFon the collection, BM25 ignores how rare a word is and loses its point. 2.Qdrant/bm25stems with an English stemmer and drops English stop words by default; there is no Bulgarian stemmer, hencedisable_stemmer=True. Usequery_embedfor the query andembedfor documents. 3.nomic-embed-textneeds the prefixsearch_document:when indexing andsearch_query:when searching — without them precision drops.💾Memory:on_disk=True10 million vectors × 768 dimensions × 4 bytes is about 31 GB for the raw vectors alone. Withon_disk=TrueQdrant keeps them in memory-mapped files, and mostly the frequently searched ones stay in RAM. How much RAM you really need depends on the load — measure before buying hardware. The old "8 GB instead of 30 GB" had no source. -
Metadata filters: who can see what
In an academy with many courses, a student of course A must not see course B's materials. A Qdrant filter is applied during the search, not after it — so the 20 returned candidates are already from the right course only. Put the filter in every
Prefetchand index the fields you filter on (for the tenant key —is_tenant=True).Python · metadata_filters.pyimport time from qdrant_client import models # 1. One course only, paid materials only paid_course = models.Filter(must=[ models.FieldCondition(key="course_id", match=models.MatchValue(value="42")), models.FieldCondition(key="price_tier", match=models.MatchValue(value="paid")), ]) # 2. A few laws only, excluding repealed ones laws_in_force = models.Filter( must=[models.FieldCondition(key="domain", match=models.MatchAny(any=["ЗЗД", "ТЗ", "КТ", "КСО"]))], must_not=[models.FieldCondition(key="status", match=models.MatchValue(value="отменен"))], ) # 3. Updated in the last 30 days (Unix time in the updated_ts field) recent = models.Filter(must=[ models.FieldCondition(key="updated_ts", range=models.Range(gte=int(time.time()) - 30 * 24 * 3600)), ]) # Usage: client.query_points(..., query_filter=paid_course) or inside each Prefetch(filter=...)⛔The filter is security, not convenienceIf you filter after the search (in Python), a bug in the code shows someone else's material. If you forget the filter in one of the twoPrefetchblocks, fusion lets foreign documents in. Test with a user who must not see a given document. -
Reranking: 20 candidates → the best 5
Vector search encodes the query and the document separately and compares the vectors. A cross-encoder reads the query and the document together and judges how well the document answers this exact question. It is more precise, but much slower — so you run it only on the 20 candidates from hybrid search.
Document (illustrative values) Position after hybrid Cross-encoder score Position after reranking Art. 40 SIC — the employer pays the first days 2 0.94 🥇 1 Art. 41 SIC — the size of the benefit 1 0.71 2 General overview of social insurance law 3 0.18 5 ↓ For "Who pays the first days of sick leave?" the general overview is close in meaning but doesn't answer — reranking pushes it down.
Python · reranker_pipeline.pyimport time from sentence_transformers import CrossEncoder from qdrant_hybrid import hybrid_search # bge-reranker-v2-m3: multilingual, Apache-2.0 licence, ~570M parameters; the first run downloads the model reranker = CrossEncoder("BAAI/bge-reranker-v2-m3", max_length=512) def rerank(query: str, candidates: list[dict], top_k: int = 5) -> list[dict]: """The cross-encoder reads the query and each document TOGETHER and gives a new score.""" if not candidates: return [] ranked = reranker.rank(query, [c["text"] for c in candidates], top_k=top_k) return [{**candidates[r["corpus_id"]], "rerank_score": float(r["score"])} for r in ranked] def advanced_retrieve(query: str, course_id: str | None = None, top_k: int = 5, candidate_k: int = 20) -> list[dict]: """Hybrid search → 20 candidates → rerank → 5 for the model.""" filters = {"course_id": course_id} if course_id else None return rerank(query, hybrid_search(query, top_k=candidate_k, filters=filters), top_k=top_k) if __name__ == "__main__": q = "кой плаща първите дни от болничния" t0 = time.perf_counter(); plain = hybrid_search(q, top_k=5) t1 = time.perf_counter(); best = advanced_retrieve(q, top_k=5) t2 = time.perf_counter() print(f"without reranker: {(t1 - t0) * 1000:.0f} ms · with reranker: {(t2 - t1) * 1000:.0f} ms") for d in best: print(f"{d['rerank_score']:.3f} | {d['text'][:90]}")⚠️Which model, under which licenceBAAI/bge-reranker-v2-m3is multilingual and Apache-2.0 — fine in a commercial product. The fastembed library does not support it (the old code stopped with an error). The multilingual model fastembed does have (jina-reranker-v2-base-multilingual) is CC-BY-NC-4.0 — non-commercial use only. The Englishms-marcomodels are fast but not built for Cyrillic.⏱️What it costs in timeThe old lesson promised "48 ms → 312 ms". That depends entirely on the machine, text length and whether you have a GPU. The script above measures both paths — run it on your machine and decide whether the delay is worth it. -
All together: from question to cited answer
The last step assembles the chain: hybrid search with a filter → reranking → the best five passages go to the model with a source label. The system prompt tells the model to answer only from the context and to say honestly when the answer is missing. The prompt stays in Bulgarian because the documents and users are Bulgarian.
Python · full_rag_pipeline.pyfrom langchain_core.prompts import ChatPromptTemplate from langchain_ollama import ChatOllama from reranker_pipeline import advanced_retrieve llm = ChatOllama(model="llama3.1:8b", temperature=0.1) # a bigger model = better Bulgarian PROMPT = ChatPromptTemplate.from_messages([ ("system", "Ти си асистент на учебна AI академия. Отговаряй САМО от дадения контекст. " "Ако отговорът не е там, кажи: „Тази информация не е в наличните материали.“ " "Отговаряй на български. При правни въпроси цитирай члена в квадратни скоби."), ("human", "Контекст:\n{context}\n\nВъпрос: {question}"), ]) def answer(question: str, course_id: str | None = None) -> dict: docs = advanced_retrieve(question, course_id=course_id, top_k=5) context = "\n\n".join( f"[{d['metadata'].get('source', 'Документ')}" f"{' · чл. ' + d['metadata']['article'] if d['metadata'].get('article') else ''}]: {d['text']}" for d in docs) reply = (PROMPT | llm).invoke({"context": context, "question": question}) return {"answer": reply.content, "sources": [d["metadata"] for d in docs]} if __name__ == "__main__": r = answer("Кой плаща първите два дни от болничния и колко?") print(r["answer"]) print("Sources:", [s.get("source", "?") for s in r["sources"]])⛔Legal answers are a draftThe system cites the text it found — it doesn't check whether the law has changed. Keep only versions in force in the base (thestatusfield) and present answers as guidance, not legal advice.
04Check
Checklist
- The chunking strategy is chosen per document type; laws are split by article with the number in metadata.
- Embeddings use the
search_document:andsearch_query:prefixes. - The collection has
denseandsparsevectors; sparse hasmodifier=IDF. - BM25 runs without the English stemmer.
- The hybrid query finds both an exact article number and a paraphrase without it.
- The filter is in every
Prefetch; the filtered fields are indexed. - Reranking runs on ~20 candidates; the delay is measured.
- The model answers only from context and cites the article.
- The Qdrant key is in the environment; the ports are not open to the outside.
Quiz
1. Hybrid search is on, but context recall is low (0.42). What is the MOST LIKELY cause?
2. A document has a dense score of 0.92 and a BM25 score of 0.15. How does RRF compute its final score?
3. A cross-encoder is more precise than vector search. Why don't we search with it directly?
4. Why do we set disable_stemmer=True on Qdrant/bm25 for Bulgarian?
05What's next
06Sources
- Qdrant: hybrid queries — prefetch, RRF, DBSF, the k parameter and weights (and from which versions).
- Qdrant: filtering —
FieldCondition,MatchValue,MatchAny,Range. - Qdrant: indexing — sparse index and the IDF modifier, payload indexes.
- Qdrant: releases and Docker healthcheck — why the image has no curl.
- FastEmbed — BM25 and the list of supported models.
- LangChain: text splitters — recursive and Markdown splitting.
- Sunsetting langchain-experimental — the May 2026 announcement.
- Ollama: /api/embed, nomic-embed-text in Ollama and the model card — the
search_document/search_queryprefixes. - BAAI/bge-reranker-v2-m3 🌐 global — the model card and licence; BGE M3-Embedding (2024) — the paper behind its base.
- Sentence Transformers: CrossEncoder —
rank()andpredict(). - Reciprocal Rank Fusion (Cormack et al., 2009) — the original paper.
- NSSI: temporary incapacity benefit (in Bulgarian) — Arts. 40–41 of the Social Insurance Code, the facts in the examples.