The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewVERIFIED 2026-10-01 · UPDATED 2026-10-01
IDENTITY
module
01-04a · Advanced RAG: chunking, hybrid search, reranking
series
Blocks 0–10 · Block 4 — RAG and knowledge systems · Part 1/2
level
Advanced
duration
4–6 h
prerequisites
Blocks 0–3 (LLM basics, inference, prompting, agents); a basic RAG pipeline; Python ≥ 3.10; Docker
trust_label
VERIFIED 2026-10-01 (package versions on PyPI, Qdrant release, Qdrant Query API / RRF / IDF docs, fastembed model list, reranker licences, all source links; code executed with in-memory Qdrant, real BM25 and mocked dense embeddings) · UPDATED 2026-10-01 · NOT end-to-end tested with a real LLM, real nomic-embed-text or bge-reranker-v2-m3; docker healthcheck not run
versions
Qdrant server v1.19.x · qdrant-client 1.19.x · fastembed 0.8.x · langchain-text-splitters 1.1.x · langchain-ollama 1.1.x · sentence-transformers 6.x
language
this page: en · bulgarian edition: /academy/blokove/moduli/01-04a_Блок_4_Част_1_Advanced_RAG.html
next
01-04b_Блок_4_Част_2_GraphRAG_RAGAS.html · GraphRAG + RAGAS evaluation
PURPOSE

Fix the three failure points of basic RAG (split → embed → retrieve → generate): bad chunks, missed exact terms, weak ranking. Chunk by document structure (one law article = one chunk), combine dense and sparse (BM25) retrieval in Qdrant with Reciprocal Rank Fusion, restrict results with payload filters, then rerank the top 20 candidates with a multilingual cross-encoder and pass the top 5 to the LLM.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

01-04b_Блок_4_Част_2_GraphRAG_RAGAS.html · GraphRAG, a production RAG system and RAGAS evaluation · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
ragchunkinghybrid-searchbm25qdrantrrfrerankingcross-encodermetadata-filteringollama
VERIFIED · 01.10.2026 UPDATED · 01.10.2026

Advanced RAG: chunking, hybrid search, reranking

Basic RAG (split → embed → retrieve → answer) works most of the time. It fails in three places: badly cut text, missed exact terms and weak ranking. In this lesson we fix all three — with Qdrant on your own machine and models that run locally.

⏱ 4–6 h Advanced Block 4 · Part 1/2 RAG · search · precision
Qdrant (vector database)🔒 local Ollama · nomic-embed-text🔒 local fastembed (BM25)🔒 local bge-reranker-v2-m3🔒 local Hugging Face (model downloads)🌐 global
🔄
UPDATED · 01.10.2026 — what changed
The code is rewritten and run against qdrant-client 1.19, fastembed 0.8 and langchain-text-splitters 1.1 (checked on PyPI) — with in-memory Qdrant, real BM25 and mocked dense vectors; no real LLM or reranker. The Qdrant image moves from v1.9.0 to v1.19.1. Bugs fixed from the old lesson: the Query API with fusion (RRF) arrived in Qdrant v1.10, not v1.9; BM25 was missing modifier=IDF; Qdrant/bm25 stems words by English rules — for Bulgarian the stemmer is turned off; the model BAAI/bge-reranker-v2-m3 is not supported by fastembed — the old code stopped with an error, it now goes through sentence-transformers; rr'…' was a syntax error; breakpoint_threshold_amount=0.85 in percentile mode should be around 95; recreate_collection is deprecated; Ollama's /api/embeddings is superseded by /api/embed; the Qdrant image has no curl, so the old healthcheck always failed. The langchain-experimental package (with SemanticChunker) is sunset since May 2026 — we give a short own implementation. The "Art. 45 SIC" and "clinical pathway No. 45" examples were inaccurate — replaced with verified articles (Arts. 40–41 of the Social Insurance Code, Art. 110 of the Obligations and Contracts Act). The figures "+31% recall", "+74% accuracy", "23% → 97%", "48 ms / 312 ms" and "10M documents in 8 GB" had no source — removed or marked as illustrative. A wrong arXiv link is fixed.

01What you'll learn

02Before you start

bash · install (versions as of 01.10.2026)
python3 -m venv .venv && source .venv/bin/activate
pip install -U qdrant-client fastembed langchain-text-splitters langchain-ollama sentence-transformers numpy requests
# checked with: qdrant-client 1.19.1 · fastembed 0.8.1 · langchain-text-splitters 1.1.2 · langchain-ollama 1.1.0
ollama pull nomic-embed-text
ollama pull llama3.1:8b

03Steps

  1. Why basic chunking fails

    Fixed-size chunking (say, every 500 characters) doesn't know where a thought ends. A law article can land half at the end of one chunk and half at the start of the next — then no chunk contains the whole rule, retrieval returns a fragment and the model answers incompletely or wrongly. Why does it matter? Because no better model can fix information that never reached it.

    example · one article cut in two (teaching example, Bulgarian text)
    Чънк 1: …Чл. 40. (5) Работодателят изплаща първите два работни дни от
            временната неработоспособност в размер 70 на сто от среднодневното
    Чънк 2: брутно възнаграждение… Чл. 41. (1) Паричното обезщетение е 80 на сто
            от среднодневното брутно възнаграждение…
    
    Question: "How much does the employer pay for the first days of sick leave?"
    → Chunk 1 lacks the base ("gross pay"), chunk 2 starts with a different article.
    💡
    The facts in the example (checked with the NSSI)
    Since 2024 the employer pays the first two working days at 70%, and the National Social Security Institute pays a benefit of 80% of the average daily gross pay (Arts. 40–41 of the Social Insurance Code). The text above is shortened for the lesson, not a verbatim quote.
  2. Five chunking strategies

    There is no single right strategy — it depends on the document. Start with recursive splitting and move to something more complex only when the check (the "Check" section and lesson 01-04b) shows it isn't enough.

    StrategyHow it worksGood for
    Fixed size + overlapN characters, K overlap. Fast and simple.General prose, news. Bad for laws.
    Recursive ★Splits on blank line → newline → ". " — keeps paragraphs.Most documents. A good start.
    SemanticCompares neighbouring sentences and cuts where the meaning "jumps".Mixed texts without clear structure: reports, opinions.
    Parent–childSmall chunks for search, a big "parent" as context for the model.Long technical documents.
    Structure-aware ★Recognises headings (##), articles ("Чл. N."), tables.Laws, regulations, standards.
    Python · advanced_chunking.py
    import re
    import numpy as np
    import requests
    from langchain_text_splitters import RecursiveCharacterTextSplitter, MarkdownHeaderTextSplitter
    
    OLLAMA = "http://localhost:11434"
    
    # ── 1. Recursive: a good default for most Bulgarian texts ──
    def recursive_chunks(text: str, chunk_size: int = 700, overlap: int = 120) -> list[str]:
        """Splits on \n\n → \n → ". " → space. Size is in characters, not tokens."""
        splitter = RecursiveCharacterTextSplitter(
            chunk_size=chunk_size,
            chunk_overlap=overlap,
            separators=["\n\n", "\n", ". ", " ", ""],
        )
        return splitter.split_text(text)
    
    # ── 2. Structure-aware: a law → one article = one chunk ──
    ARTICLE = re.compile(r"^(Чл\.\s*(\d+[а-я]?)\..*?)(?=^Чл\.\s*\d+[а-я]?\.|\Z)", re.S | re.M)
    
    def legal_law_chunks(law_text: str) -> list[dict]:
        """Every article starting a line ("Чл. 110.") becomes its own chunk with the number in metadata."""
        return [{"text": m.group(1).strip(), "article": m.group(2)}
                for m in ARTICLE.finditer(law_text) if len(m.group(1).strip()) > 20]
    
    # ── 3. Semantic: cut where the meaning "jumps" ──
    def embed(texts: list[str], prefix: str = "search_document: ") -> np.ndarray:
        """nomic-embed-text via Ollama /api/embed (batch)."""
        r = requests.post(f"{OLLAMA}/api/embed", timeout=120,
                          json={"model": "nomic-embed-text", "input": [prefix + t for t in texts]})
        r.raise_for_status()
        return np.array(r.json()["embeddings"])
    
    def semantic_chunks(text: str, percentile: float = 95, embed_fn=embed) -> list[str]:
        """Cut between two sentences when their distance is above the given percentile.
        Higher percentile = fewer cuts = bigger chunks."""
        sentences = [s for s in re.split(r"(?<=[.!?])\s+", text) if s.strip()]
        if len(sentences) < 3:
            return [text]
        v = embed_fn(sentences)
        v = v / np.linalg.norm(v, axis=1, keepdims=True)
        dist = 1 - (v[:-1] * v[1:]).sum(axis=1)          # cosine distance between neighbouring sentences
        cut = np.percentile(dist, percentile)
        chunks, current = [], [sentences[0]]
        for sentence, d in zip(sentences[1:], dist):
            if d > cut:
                chunks.append(" ".join(current))
                current = [sentence]
            else:
                current.append(sentence)
        chunks.append(" ".join(current))
        return chunks
    
    # ── 4. Parent–child: search by the small one, pass the big one ──
    def hierarchical_chunks(text: str) -> list[dict]:
        parent_splitter = RecursiveCharacterTextSplitter(chunk_size=1500, chunk_overlap=200)
        child_splitter = RecursiveCharacterTextSplitter(chunk_size=300, chunk_overlap=50)
        result = []
        for i, parent in enumerate(parent_splitter.split_text(text)):
            for child in child_splitter.split_text(parent):
                result.append({"child_text": child,     # small → for embedding and search
                               "parent_text": parent,   # large → for the model
                               "parent_id": i})
        return result
    
    # ── 5. Markdown: split by headings and keep them as metadata ──
    def markdown_chunks(md_text: str) -> list[dict]:
        splitter = MarkdownHeaderTextSplitter(
            headers_to_split_on=[("#", "h1"), ("##", "h2"), ("###", "h3")],
            strip_headers=False,
        )
        return [{"text": d.page_content, **d.metadata} for d in splitter.split_text(md_text)]
    ⚠️
    Characters, not tokens — and a short context
    chunk_size in RecursiveCharacterTextSplitter counts characters. Cyrillic usually breaks into more tokens than English, so don't size chunks by eye. In the Ollama library nomic-embed-text has a 2K-token window (the model itself supports up to 8192) — keep chunks small, or their ends get lost.
    ⚠️
    Old code on the internet
    Many examples use SemanticChunker from langchain-experimental. The package has been sunset since May 2026 and is not maintained — so semantic chunking above is our own 20-line function. In percentile mode the threshold is a number from 0 to 100 (default 95), not 0.85.
    ⚖️
    Illustrative scenario: a law by article
    Picture a knowledge base with the Bulgarian Obligations and Contracts Act cut into 500-character pieces. Asked "What is the general limitation period?", retrieval returns the end of Art. 109 and the start of Art. 110 — without the rule that after a five-year limitation period all claims for which the law sets no other period are extinguished. With article-level chunking Art. 110 is whole and carries article="110" in its metadata, so the answer can cite it. The old lesson claimed "23% → 97% accuracy" — we could not verify the source, so the figures are removed. Measure on your own documents.
  3. Hybrid search: meaning + exact words

    Semantic search (dense vectors) is great for questions like "who pays when I'm sick", but misses exact identifiers: "Art. 41", a law's abbreviation, an order number. BM25 (sparse) finds exactly those words but doesn't understand paraphrase. Hybrid search runs both and fuses the results.

    Query: "sick-leave benefit Art. 41 SIC"What it findsWhat it misses
    Dense (nomic-embed-text)"temporary incapacity", "cash benefit" — related in meaningThe exact "Art. 41"
    Sparse (BM25)Exactly "Art. 41", "SIC"Texts with the same meaning but other words
    Hybrid (RRF)Documents ranked high in both lists—
    ✅
    How RRF fuses the two lists
    Reciprocal Rank Fusion looks only at positions, not scores: each document gets Σ 1/(k + position) from each list. So it doesn't matter that dense and BM25 scores live on different scales. In Qdrant positions start at 0 and by default k = 2; since v1.16 you can set k (the classic paper uses 60), and since v1.17 also a weight per list.

    Run Qdrant in Docker. The ports are bound to the local interface only, and the key comes from .env.

    YAML · docker-compose.yml
    # docker-compose.yml — Qdrant on your own machine
    services:
      qdrant:
        image: qdrant/qdrant:v1.19.1          # pinned version (latest as of 01.10.2026)
        restart: unless-stopped
        ports:
          - "127.0.0.1:6333:6333"             # REST — local only
          - "127.0.0.1:6334:6334"             # gRPC
        volumes:
          - ./qdrant_storage:/qdrant/storage
        environment:
          QDRANT__SERVICE__API_KEY: ${QDRANT_API_KEY}   # from .env
        healthcheck:                          # the image has no curl or wget
          test: ["CMD", "bash", "-c", "exec 3<>/dev/tcp/127.0.0.1/6333 && printf 'GET /readyz HTTP/1.0\\r\\n\\r\\n' >&3 && grep -q ' 200 ' <&3"]
          interval: 15s
          timeout: 5s
          retries: 5
    ⚠️
    The Qdrant image has no curl
    The old check curl -f …/healthz always fails, because curl and wget were removed from the image. The check above uses bash and /readyz, which returns 200 only once Qdrant can serve requests. We have not run it on a live machine — confirm with docker ps that the status is healthy.
    Python · qdrant_hybrid.py
    import os
    import requests
    from fastembed import SparseTextEmbedding
    from qdrant_client import QdrantClient, models
    
    OLLAMA = "http://localhost:11434"
    COLLECTION = "academy_bg"
    
    client = QdrantClient(url=os.getenv("QDRANT_URL", "http://localhost:6333"),
                          api_key=os.getenv("QDRANT_API_KEY"))    # the key comes from .env, not from code
    
    # ── 1. A collection with two kinds of vectors ──
    if not client.collection_exists(COLLECTION):
        client.create_collection(
            collection_name=COLLECTION,
            vectors_config={
                "dense": models.VectorParams(size=768, distance=models.Distance.COSINE,
                                             on_disk=True),          # nomic-embed-text = 768
            },
            sparse_vectors_config={
                "sparse": models.SparseVectorParams(
                    index=models.SparseIndexParams(on_disk=True),
                    modifier=models.Modifier.IDF,                    # required for BM25
                ),
            },
        )
        # index the fields you filter on — otherwise filtering is slow
        client.create_payload_index(COLLECTION, "course_id",
                                    field_schema=models.KeywordIndexParams(type="keyword", is_tenant=True))
    
    # ── 2. Embeddings ──
    def get_dense(text: str, kind: str = "document") -> list[float]:
        """nomic-embed-text needs a prefix: search_document: when indexing, search_query: when searching."""
        prefix = "search_query: " if kind == "query" else "search_document: "
        r = requests.post(f"{OLLAMA}/api/embed", timeout=60,
                          json={"model": "nomic-embed-text", "input": prefix + text})
        r.raise_for_status()
        return r.json()["embeddings"][0]
    
    # Qdrant/bm25 has no Bulgarian stemmer → turn it off (otherwise it stems by English rules)
    bm25 = SparseTextEmbedding(model_name="Qdrant/bm25", disable_stemmer=True)
    
    def get_sparse(text: str, kind: str = "document") -> models.SparseVector:
        emb = next(bm25.query_embed(text) if kind == "query" else bm25.embed([text]))
        return models.SparseVector(indices=emb.indices.tolist(), values=emb.values.tolist())
    
    # ── 3. Indexing ──
    def index_document(doc_id: int, text: str, metadata: dict):
        client.upsert(COLLECTION, points=[models.PointStruct(
            id=doc_id,
            vector={"dense": get_dense(text), "sparse": get_sparse(text)},
            payload={"text": text, **metadata},
        )])
    
    # ── 4. Hybrid search: two prefetches + RRF inside Qdrant ──
    def hybrid_search(query: str, top_k: int = 5, filters: dict | None = None) -> list[dict]:
        flt = None
        if filters:
            flt = models.Filter(must=[models.FieldCondition(key=k, match=models.MatchValue(value=v))
                                      for k, v in filters.items()])
        res = client.query_points(
            collection_name=COLLECTION,
            prefetch=[
                models.Prefetch(query=get_dense(query, "query"), using="dense", limit=20, filter=flt),
                models.Prefetch(query=get_sparse(query, "query"), using="sparse", limit=20, filter=flt),
            ],
            query=models.FusionQuery(fusion=models.Fusion.RRF),   # or models.RrfQuery(rrf=models.Rrf(k=60))
            limit=top_k,
            with_payload=True,
        )
        return [{"text": p.payload.get("text", ""), "score": p.score,
                 "metadata": {k: v for k, v in p.payload.items() if k != "text"}}
                for p in res.points]
    
    if __name__ == "__main__":
        for r in hybrid_search("обезщетение за болничен чл. 41 КСО", top_k=5, filters={"course_id": "42"}):
            print(f"{r['score']:.3f} | {r['text'][:100]}")
    ⚠️
    Three traps in hybrid search
    1. Without modifier=IDF on the collection, BM25 ignores how rare a word is and loses its point. 2. Qdrant/bm25 stems with an English stemmer and drops English stop words by default; there is no Bulgarian stemmer, hence disable_stemmer=True. Use query_embed for the query and embed for documents. 3. nomic-embed-text needs the prefix search_document: when indexing and search_query: when searching — without them precision drops.
    💾
    Memory: on_disk=True
    10 million vectors × 768 dimensions × 4 bytes is about 31 GB for the raw vectors alone. With on_disk=True Qdrant keeps them in memory-mapped files, and mostly the frequently searched ones stay in RAM. How much RAM you really need depends on the load — measure before buying hardware. The old "8 GB instead of 30 GB" had no source.
  4. Metadata filters: who can see what

    In an academy with many courses, a student of course A must not see course B's materials. A Qdrant filter is applied during the search, not after it — so the 20 returned candidates are already from the right course only. Put the filter in every Prefetch and index the fields you filter on (for the tenant key — is_tenant=True).

    Python · metadata_filters.py
    import time
    from qdrant_client import models
    
    # 1. One course only, paid materials only
    paid_course = models.Filter(must=[
        models.FieldCondition(key="course_id", match=models.MatchValue(value="42")),
        models.FieldCondition(key="price_tier", match=models.MatchValue(value="paid")),
    ])
    
    # 2. A few laws only, excluding repealed ones
    laws_in_force = models.Filter(
        must=[models.FieldCondition(key="domain", match=models.MatchAny(any=["ЗЗД", "ТЗ", "КТ", "КСО"]))],
        must_not=[models.FieldCondition(key="status", match=models.MatchValue(value="отменен"))],
    )
    
    # 3. Updated in the last 30 days (Unix time in the updated_ts field)
    recent = models.Filter(must=[
        models.FieldCondition(key="updated_ts", range=models.Range(gte=int(time.time()) - 30 * 24 * 3600)),
    ])
    
    # Usage: client.query_points(..., query_filter=paid_course)  or inside each Prefetch(filter=...)
    ⛔
    The filter is security, not convenience
    If you filter after the search (in Python), a bug in the code shows someone else's material. If you forget the filter in one of the two Prefetch blocks, fusion lets foreign documents in. Test with a user who must not see a given document.
  5. Reranking: 20 candidates → the best 5

    Vector search encodes the query and the document separately and compares the vectors. A cross-encoder reads the query and the document together and judges how well the document answers this exact question. It is more precise, but much slower — so you run it only on the 20 candidates from hybrid search.

    Document (illustrative values)Position after hybridCross-encoder scorePosition after reranking
    Art. 40 SIC — the employer pays the first days20.94🥇 1
    Art. 41 SIC — the size of the benefit10.712
    General overview of social insurance law30.185 ↓

    For "Who pays the first days of sick leave?" the general overview is close in meaning but doesn't answer — reranking pushes it down.

    Python · reranker_pipeline.py
    import time
    from sentence_transformers import CrossEncoder
    from qdrant_hybrid import hybrid_search
    
    # bge-reranker-v2-m3: multilingual, Apache-2.0 licence, ~570M parameters; the first run downloads the model
    reranker = CrossEncoder("BAAI/bge-reranker-v2-m3", max_length=512)
    
    def rerank(query: str, candidates: list[dict], top_k: int = 5) -> list[dict]:
        """The cross-encoder reads the query and each document TOGETHER and gives a new score."""
        if not candidates:
            return []
        ranked = reranker.rank(query, [c["text"] for c in candidates], top_k=top_k)
        return [{**candidates[r["corpus_id"]], "rerank_score": float(r["score"])} for r in ranked]
    
    def advanced_retrieve(query: str, course_id: str | None = None,
                          top_k: int = 5, candidate_k: int = 20) -> list[dict]:
        """Hybrid search → 20 candidates → rerank → 5 for the model."""
        filters = {"course_id": course_id} if course_id else None
        return rerank(query, hybrid_search(query, top_k=candidate_k, filters=filters), top_k=top_k)
    
    if __name__ == "__main__":
        q = "кой плаща първите дни от болничния"
        t0 = time.perf_counter(); plain = hybrid_search(q, top_k=5)
        t1 = time.perf_counter(); best = advanced_retrieve(q, top_k=5)
        t2 = time.perf_counter()
        print(f"without reranker: {(t1 - t0) * 1000:.0f} ms · with reranker: {(t2 - t1) * 1000:.0f} ms")
        for d in best:
            print(f"{d['rerank_score']:.3f} | {d['text'][:90]}")
    ⚠️
    Which model, under which licence
    BAAI/bge-reranker-v2-m3 is multilingual and Apache-2.0 — fine in a commercial product. The fastembed library does not support it (the old code stopped with an error). The multilingual model fastembed does have (jina-reranker-v2-base-multilingual) is CC-BY-NC-4.0 — non-commercial use only. The English ms-marco models are fast but not built for Cyrillic.
    ⏱️
    What it costs in time
    The old lesson promised "48 ms → 312 ms". That depends entirely on the machine, text length and whether you have a GPU. The script above measures both paths — run it on your machine and decide whether the delay is worth it.
  6. All together: from question to cited answer

    The last step assembles the chain: hybrid search with a filter → reranking → the best five passages go to the model with a source label. The system prompt tells the model to answer only from the context and to say honestly when the answer is missing. The prompt stays in Bulgarian because the documents and users are Bulgarian.

    Python · full_rag_pipeline.py
    from langchain_core.prompts import ChatPromptTemplate
    from langchain_ollama import ChatOllama
    from reranker_pipeline import advanced_retrieve
    
    llm = ChatOllama(model="llama3.1:8b", temperature=0.1)   # a bigger model = better Bulgarian
    
    PROMPT = ChatPromptTemplate.from_messages([
        ("system", "Ти си асистент на учебна AI академия. Отговаряй САМО от дадения контекст. "
                   "Ако отговорът не е там, кажи: „Тази информация не е в наличните материали.“ "
                   "Отговаряй на български. При правни въпроси цитирай члена в квадратни скоби."),
        ("human", "Контекст:\n{context}\n\nВъпрос: {question}"),
    ])
    
    def answer(question: str, course_id: str | None = None) -> dict:
        docs = advanced_retrieve(question, course_id=course_id, top_k=5)
        context = "\n\n".join(
            f"[{d['metadata'].get('source', 'Документ')}"
            f"{' · чл. ' + d['metadata']['article'] if d['metadata'].get('article') else ''}]: {d['text']}"
            for d in docs)
        reply = (PROMPT | llm).invoke({"context": context, "question": question})
        return {"answer": reply.content, "sources": [d["metadata"] for d in docs]}
    
    if __name__ == "__main__":
        r = answer("Кой плаща първите два дни от болничния и колко?")
        print(r["answer"])
        print("Sources:", [s.get("source", "?") for s in r["sources"]])
    ⛔
    Legal answers are a draft
    The system cites the text it found — it doesn't check whether the law has changed. Keep only versions in force in the base (the status field) and present answers as guidance, not legal advice.

04Check

Checklist

Quiz

1. Hybrid search is on, but context recall is low (0.42). What is the MOST LIKELY cause?

2. A document has a dense score of 0.92 and a BM25 score of 0.15. How does RRF compute its final score?

3. A cross-encoder is more precise than vector search. Why don't we search with it directly?

4. Why do we set disable_stemmer=True on Qdrant/bm25 for Bulgarian?

05What's next

06Sources

  1. Qdrant: hybrid queries — prefetch, RRF, DBSF, the k parameter and weights (and from which versions).
  2. Qdrant: filtering — FieldCondition, MatchValue, MatchAny, Range.
  3. Qdrant: indexing — sparse index and the IDF modifier, payload indexes.
  4. Qdrant: releases and Docker healthcheck — why the image has no curl.
  5. FastEmbed — BM25 and the list of supported models.
  6. LangChain: text splitters — recursive and Markdown splitting.
  7. Sunsetting langchain-experimental — the May 2026 announcement.
  8. Ollama: /api/embed, nomic-embed-text in Ollama and the model card — the search_document / search_query prefixes.
  9. BAAI/bge-reranker-v2-m3 🌐 global — the model card and licence; BGE M3-Embedding (2024) — the paper behind its base.
  10. Sentence Transformers: CrossEncoder — rank() and predict().
  11. Reciprocal Rank Fusion (Cormack et al., 2009) — the original paper.
  12. NSSI: temporary incapacity benefit (in Bulgarian) — Arts. 40–41 of the Social Insurance Code, the facts in the examples.