The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewUPDATED 2026-10-03
IDENTITY
module
GX10-04-156 · Local venue assistant (RAG + tool calling) for a cultural and sports complex
series
GX10 (local AI server class: NVIDIA GB10, e.g. ASUS Ascent GX10 / DGX Spark)
level
Advanced
duration
2-3 h
prerequisites
A GB10-class machine with Ollama, Python 3.10+, basic RAG knowledge; 04-153 helps
trust_label
UPDATED 2026-10-03 (Ollama model pages, Ollama tool-calling docs, PyPI version read on 2026-10-03) · NOT TESTED on a GB10 machine · no VERIFIED label · all venue data is invented
versions
llama3.1:8b = 4.9 GB · bge-m3 = 1.2 GB (ollama.com, 2026-10-03) · ollama (Python) 0.6.3 (PyPI, 2026-10-03)
language
human view: bg · english edition: /en/academy/gx10/ (same file name)
previous / next
04-153_Event_AI_Assistant.html · 04-158_Event_Demand_Forecast.html
PURPOSE

Build a local assistant for hall questions: markdown knowledge files split by section and embedded with bge-m3 into Chroma, availability from a SQLite bookings table through an Ollama tool call, a FastAPI chat bound to loopback. A person confirms every availability and price; the lesson data is invented.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

04-158 · Event demand forecast (04-158_Event_Demand_Forecast.html) · related: 04-153 Event assistant, 04-84 Enterprise RAG pipeline · series index: kagami.bg/academy/gx10/ · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
gx10nvidia-gb10ragchromadbtool-callingollamafastapivenue
UPDATED · 03.10.2026

Venue RAG Assistant on a Local GX10

A small helper that knows the halls of a cultural and sports complex: it answers about seats, equipment and prices from your own files and checks availability in a calendar through a function. The models run on your machine. All halls, seat counts and prices in the lesson are invented.

⏱ 2–3 h Advanced GX10 RAG · tool calling · FastAPI
Ollama · llama3.1:8b (the answers)🔒 local bge-m3 (search by meaning)🔒 local ChromaDB (hall knowledge)🔒 local SQLite (bookings)🔒 local FastAPI🔒 local
🔄
UPDATED · 03.10.2026 — what
We generalised the example. The lesson is no longer about a specific building: halls, seat counts and prices are invented (prices are in EUR, the old version used the former national currency), and we removed the tenant count, the car park and equipment brands. We fixed the code: the old one passed the function description in a format Ollama does not expect and never handled the function-call reply — so availability was never checked. Now there is a loop following the Ollama documentation. The old server listened on 0.0.0.0 with no login — now it is 127.0.0.1; the client can no longer send system messages; dates and the hall identifier are validated. We simplified: chunks are cut by section instead of word count, and vectors come from bge-m3 through Ollama (we dropped sentence-transformers). We removed: the claim of "no transfer between CPU and GPU", the Russian dialogue and the claim that local vectors are "practically mandatory".
⚠️
What we have not run ourselves
The code was not run on a GB10-class machine and was not connected to a real Ollama and Chroma, so there is no "TESTED" label and no "VERIFIED". None of the commands above was run ("not run"). We did not measure the speed or the quality of the answers. Versions and sizes were checked against the documentation on 03.10.2026; how the small model behaves with function calls you must check yourself.

01What you will learn

02Before you start

Will it fit in memory?

WhatSize (per ollama.com, 03.10.2026)
llama3.1:8b4,9 GB
bge-m31,2 GB
Togetherabout 6.1 GB
Machine memory128 GB, shared by the CPU and the GPU
💡
A calculation, not a promise
The sizes are those of the model files. The conversation context and the index take extra memory, and we did not measure the speed of your machine. A larger model (for example llama3.1:70b, 43 GB) answers better but is slower and heavier on memory — see 04-153.

03Steps

  1. The models and packages

    Why two models: bge-m3 turns text into numbers used for searching (per its Ollama page it supports over 100 languages). llama3.1:8b writes the answer and is marked in Ollama as a model with tool support.

    bash · on the machine · not run
    ollama pull llama3.1:8b
    ollama pull bge-m3
    python3 -m venv .venv && . .venv/bin/activate
    pip install chromadb ollama fastapi "uvicorn[standard]"
  2. The knowledge: small files per hall

    Why this way: each hall is one file and each section ("Basic parameters", "Equipment", "Prices") is a separate chunk. That way search returns the exact section and the answer can show where it came from. Every file carries a "Valid from" date — prices and capacities go stale and someone has to maintain them.

    project tree
    venue-rag/
    ├── knowledge/
    │   ├── halls/
    │   │   ├── main.md
    │   │   ├── conf_a.md
    │   │   ├── conf_b.md
    │   │   └── workshop.md
    │   └── rules/
    │       └── rental_rules.md
    ├── db/                 # Chroma and SQLite are created automatically
    ├── kb.py
    ├── calendar_db.py
    ├── rag.py
    └── main.py
    markdown · knowledge/halls/main.md
    # Main hall (invented)
    
    ## Basic parameters
    - Seats, theatre layout: 1200
    - Seats, banquet layout: 700
    - Area: 1100 sq m
    
    ## Equipment
    - Stage, LED screen, sound system, interpreter booths
    - Rear loading access, dressing rooms: 4
    
    ## Prices (invented, EUR)
    - Full day: 3500 EUR; half day: 2000 EUR
    - Technical staff for concerts: mandatory, per tariff
    - Valid from: 2026-01-01
  3. Indexing

    Why by sections, not by word count: a chunk that starts mid-sentence is searched poorly. Here we cut on the "## " headings and put the hall name back on every chunk. The identifier is file#number, so running it again updates the chunks instead of duplicating them.

    python · kb.py · not run
    # kb.py · hall knowledge (all data is invented)
    from pathlib import Path
    
    import chromadb
    from chromadb.utils.embedding_functions.ollama_embedding_function import (
        OllamaEmbeddingFunction,
    )
    
    OLLAMA = "http://localhost:11434"
    KB_DIR = Path("./knowledge")
    
    chroma = chromadb.PersistentClient(path="./db/chroma")
    embed_fn = OllamaEmbeddingFunction(url=OLLAMA, model_name="bge-m3")
    col = chroma.get_or_create_collection("venue_knowledge", embedding_function=embed_fn)
    
    
    def split_sections(text: str) -> list[str]:
        """Split markdown on "## " headings: one section = one chunk."""
        parts, current = [], []
        for line in text.splitlines():
            if line.startswith("## ") and current:
                parts.append("\n".join(current).strip())
                current = []
            current.append(line)
        if current:
            parts.append("\n".join(current).strip())
        return [p for p in parts if p]
    
    
    def index_kb() -> int:
        ids, docs, metas = [], [], []
        for path in sorted(KB_DIR.rglob("*.md")):
            text = path.read_text(encoding="utf-8")
            title = text.splitlines()[0]  # "# Hall name"
            rel = str(path.relative_to(KB_DIR))
            for i, section in enumerate(split_sections(text)):
                # every chunk carries the hall name so it can be found by meaning
                doc = section if i == 0 else title + "\n" + section
                ids.append(f"{rel}#{i}")
                docs.append(doc)
                metas.append({"source": rel})
        if ids:
            col.upsert(ids=ids, documents=docs, metadatas=metas)
        return len(ids)
    
    
    def search(query: str, n: int = 4) -> list[dict]:
        if col.count() == 0:
            return []
        res = col.query(query_texts=[query], n_results=min(n, col.count()))
        return [{"text": d, "source": m["source"]}
                for d, m in zip(res["documents"][0], res["metadatas"][0])]

    If you delete or rename a file, its old chunks stay in the base. For this lesson it is enough to delete the db/chroma folder and index again.

  4. Calendar and the availability function

    Why a function: bookings change every day and must not live in text the model reads. The function validates the dates and the hall and returns "booked" or "no_booking_in_records". The second value only means "there is no entry in the table" — not "free". That is why a person gives the final confirmation.

    python · calendar_db.py · not run
    # calendar_db.py · bookings in SQLite (invented)
    import sqlite3
    from contextlib import closing
    from datetime import date
    
    DB_PATH = "./db/bookings.db"
    HALLS = {"main": 1200, "conf_a": 400, "conf_b": 200, "workshop": 80}  # seats (invented)
    
    
    def init_db() -> None:
        with closing(sqlite3.connect(DB_PATH)) as conn, conn:
            conn.execute("""
                CREATE TABLE IF NOT EXISTS bookings (
                    id INTEGER PRIMARY KEY AUTOINCREMENT,
                    hall_id   TEXT NOT NULL,
                    title     TEXT,
                    date_from TEXT NOT NULL,          -- YYYY-MM-DD
                    date_to   TEXT NOT NULL,
                    status    TEXT NOT NULL DEFAULT 'confirmed'  -- confirmed / tentative / cancelled
                )""")
    
    
    def check_availability(hall_id: str, date_from: str, date_to: str) -> dict:
        """Returns "booked" or "no_booking_in_records". The second value is NOT a confirmation."""
        if hall_id not in HALLS:
            return {"status": "error", "detail": "unknown hall_id", "known": list(HALLS)}
        try:
            d1, d2 = date.fromisoformat(date_from), date.fromisoformat(date_to)
        except ValueError:
            return {"status": "error", "detail": "dates must be YYYY-MM-DD"}
        if d2 < d1:
            return {"status": "error", "detail": "date_to is before date_from"}
        with closing(sqlite3.connect(DB_PATH)) as conn:
            rows = conn.execute(
                "SELECT title, date_from, date_to, status FROM bookings "
                "WHERE hall_id = ? AND status != 'cancelled' "
                "AND date_from <= ? AND date_to >= ?",
                (hall_id, d2.isoformat(), d1.isoformat()),
            ).fetchall()
        return {
            "hall_id": hall_id,
            "status": "booked" if rows else "no_booking_in_records",
            "conflicts": [{"title": r[0], "from": r[1], "to": r[2], "state": r[3]} for r in rows],
        }
  5. Retrieval + tool calling

    How it works: the model receives the chunks as context and a description of the function. If it decides to call it, we run it and return the result as a message with the role "tool"; then the model writes the answer. That is the pattern in the Ollama tool-calling documentation. The loop is limited to 3 rounds.

    python · rag.py · not run
    # rag.py · retrieval + tool calling
    import json
    from datetime import date
    
    from ollama import chat
    
    from calendar_db import check_availability
    from kb import search
    
    MODEL = "llama3.1:8b"
    
    TOOLS = [{
        "type": "function",
        "function": {
            "name": "check_availability",
            "description": "Check whether a hall has a booking in a date range, from the booking records only",
            "parameters": {
                "type": "object",
                "required": ["hall_id", "date_from", "date_to"],
                "properties": {
                    "hall_id": {"type": "string", "description": "main, conf_a, conf_b or workshop"},
                    "date_from": {"type": "string", "description": "YYYY-MM-DD"},
                    "date_to": {"type": "string", "description": "YYYY-MM-DD"},
                },
            },
        },
    }]
    FUNCS = {"check_availability": check_availability}
    
    SYSTEM = (
        "Ти си помощник за зали в културен и спортен комплекс. Отговаряй на езика на въпроса "
        "(български или английски). За размери, оборудване и цени използвай САМО КОНТЕКСТА по-долу; "
        "ако нещо липсва, кажи го. За наличност на дата ВИНАГИ извикай check_availability, не гадай. "
        "Наличността е по записите в базата, не е потвърждение. Не обещавай нищо на клиент. "
        "Днес е {today}."
    )
    
    
    def ask(question: str, history: list[dict]) -> dict:
        chunks = search(question)
        context = "\n\n".join(f"[{c['source']}]\n{c['text']}" for c in chunks)
        system = SYSTEM.format(today=date.today().isoformat()) + "\n\nКОНТЕКСТ:\n" + context
        clean = [m for m in history[-6:] if m.get("role") in ("user", "assistant")]
        messages = [{"role": "system", "content": system}, *clean,
                    {"role": "user", "content": question}]
        for _ in range(3):  # at most 3 rounds of "model → function → model"
            resp = chat(model=MODEL, messages=messages, tools=TOOLS)
            messages.append(resp.message)
            if not resp.message.tool_calls:
                return {"answer": resp.message.content,
                        "sources": sorted({c["source"] for c in chunks})}
            for call in resp.message.tool_calls:
                fn = FUNCS.get(call.function.name)
                if fn is None:
                    result = {"status": "error", "detail": "unknown tool"}
                else:
                    try:
                        result = fn(**call.function.arguments)
                    except TypeError:
                        result = {"status": "error", "detail": "bad arguments"}
                messages.append({"role": "tool", "tool_name": call.function.name,
                                 "content": json.dumps(result, ensure_ascii=False)})
        return {"answer": "Не стигнах до отговор. Опитай с по-конкретен въпрос.", "sources": []}
    ⚠️
    The small model makes mistakes
    A model with 8 billion parameters sometimes does not call the function, passes a wrong hall identifier or a wrong date format. The code returns an error instead of crashing, but it cannot force the model to ask. So: try at least ten questions of your own about dates, write down the misses and never show an availability answer to a client without a person checking it.
  6. API and chat page

    Why this way: the server listens on 127.0.0.1 only and has no login — do not publish it to the network. The page inserts text with textContent, not as HTML. History from the client is accepted only with the roles "user" and "assistant", so nobody can slip in their own system instructions. The port is arbitrary.

    python · main.py · not run
    # main.py · API and a small chat page
    from contextlib import asynccontextmanager
    
    from fastapi import FastAPI
    from fastapi.responses import HTMLResponse
    from pydantic import BaseModel, Field
    
    from calendar_db import check_availability, init_db
    from kb import index_kb
    from rag import ask
    
    
    @asynccontextmanager
    async def lifespan(_: FastAPI):
        init_db()
        print("indexed sections:", index_kb())
        yield
    
    
    app = FastAPI(title="Venue RAG", lifespan=lifespan)
    
    
    class ChatRequest(BaseModel):
        message: str = Field(max_length=2000)
        history: list[dict] = Field(default_factory=list, max_length=12)
    
    
    PAGE = """<!doctype html><meta charset="utf-8"><title>Venue RAG</title>
    <div id="log"></div>
    <p><input id="m" size="60" placeholder="E.g. Which hall is free on 2026-11-14 for 300 people?"> <button id="b">Send</button></p>
    <script>
    let history = [];
    const log = document.getElementById("log");
    function add(who, text) { const p = document.createElement("p"); p.textContent = who + ": " + text; log.appendChild(p); }
    document.getElementById("b").onclick = async () => {
      const m = document.getElementById("m"); const message = m.value; m.value = "";
      add("You", message);
      const r = await fetch("/chat", {method: "POST", headers: {"Content-Type": "application/json"},
        body: JSON.stringify({message, history})});
      const data = await r.json();
      add("Assistant", data.answer + (data.sources && data.sources.length ? "  [" + data.sources.join(", ") + "]" : ""));
      history.push({role: "user", content: message}, {role: "assistant", content: data.answer});
      history = history.slice(-12);
    };
    </script>"""
    
    
    @app.get("/")
    def index() -> HTMLResponse:
        return HTMLResponse(PAGE)
    
    
    @app.post("/chat")
    def chat_endpoint(req: ChatRequest) -> dict:
        return ask(req.message, req.history)
    
    
    @app.get("/availability/{hall_id}")
    def availability(hall_id: str, date_from: str, date_to: str) -> dict:
        return check_availability(hall_id, date_from, date_to)
    bash · on the machine · not run
    uvicorn main:app --host 127.0.0.1 --port 8156

    From your own computer you reach the page through an SSH tunnel:

    bash · on your computer · not run
    ssh -L 8156:localhost:8156 <user>@<server-address>

    Then open http://localhost:8156.

  7. Try it with invented questions

    Add a few invented bookings to the table and ask the questions below. The last two test "I do not know": if the assistant makes up an answer, tighten the instructions in SYSTEM and try again.

    Question (invented)What to expect
    How many seats does the Main hall have?an answer from the base with source halls/main.md, no function
    Is Conference hall A free on 2026-11-14?a function call; the answer "no entry" or "booked"
    How much does a hall for 5000 people cost?"there is no such hall in the base" — not an invented number
    What is the refund policy?"the base has no information about this"
  8. Limits and safety

    • Every availability and price is checked by a person before it reaches a client.
    • The base holds no personal data. If you later add tenant or contact data, that is personal data processing and has to be agreed with data protection in your organisation.
    • The chat has no login — keep it behind the tunnel.
    • The base of chunks is maintained: a "Valid from" date, an owner, a review when prices change.

04Check

Quiz

1. Why is availability taken with a function and not from text in the base?

2. What does "no_booking_in_records" mean?

3. How much memory do the two models in the lesson take (per ollama.com)?

4. What do you do when the small model does not call the function for a date?

05What next

06Sources

  1. Ollama: llama3.1 · Ollama: bge-m3 🔒 local — file sizes, the "tools" tag, multilingual support (read on 03.10.2026).
  2. Ollama: tool calling — tool schema, the "tool" role message, the loop (read on 03.10.2026).
  3. ollama (Python) 0.6.3 — version on PyPI as of 03.10.2026.
  4. Chroma: embedding function through Ollama — OllamaEmbeddingFunction(url, model_name).
  5. FastAPI: lifespan events — asynccontextmanager.