Venue RAG Assistant on a Local GX10
A small helper that knows the halls of a cultural and sports complex: it answers about seats, equipment and prices from your own files and checks availability in a calendar through a function. The models run on your machine. All halls, seat counts and prices in the lesson are invented.
0.0.0.0 with no login — now it is 127.0.0.1; the client can no longer send system messages; dates and the hall identifier are validated. We simplified: chunks are cut by section instead of word count, and vectors come from bge-m3 through Ollama (we dropped sentence-transformers). We removed: the claim of "no transfer between CPU and GPU", the Russian dialogue and the claim that local vectors are "practically mandatory".
01What you will learn
- How to organise hall knowledge as small markdown files and turn it into a searchable base on your own machine.
- How to let the model call a function for availability instead of guessing from text that may be out of date.
- How to make the assistant name the source of an answer and say "I do not know" when the base has no answer.
- How to keep the chat closed to the network and what a person must check before promising anything to a client.
- Where this approach breaks: a small model may skip the function call or pass wrong arguments.
02Before you start
- A machine of the NVIDIA GB10 class (for example ASUS Ascent GX10 or DGX Spark) with Ollama and Python 3.10 or newer.
- RAG basics: documents → vectors → search → an answer based on what was found.
- It helps to have read 04-153: it uses the same invented halls, with availability from a CalDAV calendar. Here we take it from a simple SQLite table.
- Your own hall data. Everything in the lesson is invented; do not put personal data of visitors or tenants in the base.
Will it fit in memory?
| What | Size (per ollama.com, 03.10.2026) |
|---|---|
llama3.1:8b | 4,9 GB |
bge-m3 | 1,2 GB |
| Together | about 6.1 GB |
| Machine memory | 128 GB, shared by the CPU and the GPU |
llama3.1:70b, 43 GB) answers better but is slower and heavier on memory — see 04-153.03Steps
-
The models and packages
Why two models:
bge-m3turns text into numbers used for searching (per its Ollama page it supports over 100 languages).llama3.1:8bwrites the answer and is marked in Ollama as a model with tool support.bash · on the machine · not runollama pull llama3.1:8b ollama pull bge-m3 python3 -m venv .venv && . .venv/bin/activate pip install chromadb ollama fastapi "uvicorn[standard]" -
The knowledge: small files per hall
Why this way: each hall is one file and each section ("Basic parameters", "Equipment", "Prices") is a separate chunk. That way search returns the exact section and the answer can show where it came from. Every file carries a "Valid from" date — prices and capacities go stale and someone has to maintain them.
project treevenue-rag/ ├── knowledge/ │ ├── halls/ │ │ ├── main.md │ │ ├── conf_a.md │ │ ├── conf_b.md │ │ └── workshop.md │ └── rules/ │ └── rental_rules.md ├── db/ # Chroma and SQLite are created automatically ├── kb.py ├── calendar_db.py ├── rag.py └── main.pymarkdown · knowledge/halls/main.md# Main hall (invented) ## Basic parameters - Seats, theatre layout: 1200 - Seats, banquet layout: 700 - Area: 1100 sq m ## Equipment - Stage, LED screen, sound system, interpreter booths - Rear loading access, dressing rooms: 4 ## Prices (invented, EUR) - Full day: 3500 EUR; half day: 2000 EUR - Technical staff for concerts: mandatory, per tariff - Valid from: 2026-01-01 -
Indexing
Why by sections, not by word count: a chunk that starts mid-sentence is searched poorly. Here we cut on the "## " headings and put the hall name back on every chunk. The identifier is
file#number, so running it again updates the chunks instead of duplicating them.python · kb.py · not run# kb.py · hall knowledge (all data is invented) from pathlib import Path import chromadb from chromadb.utils.embedding_functions.ollama_embedding_function import ( OllamaEmbeddingFunction, ) OLLAMA = "http://localhost:11434" KB_DIR = Path("./knowledge") chroma = chromadb.PersistentClient(path="./db/chroma") embed_fn = OllamaEmbeddingFunction(url=OLLAMA, model_name="bge-m3") col = chroma.get_or_create_collection("venue_knowledge", embedding_function=embed_fn) def split_sections(text: str) -> list[str]: """Split markdown on "## " headings: one section = one chunk.""" parts, current = [], [] for line in text.splitlines(): if line.startswith("## ") and current: parts.append("\n".join(current).strip()) current = [] current.append(line) if current: parts.append("\n".join(current).strip()) return [p for p in parts if p] def index_kb() -> int: ids, docs, metas = [], [], [] for path in sorted(KB_DIR.rglob("*.md")): text = path.read_text(encoding="utf-8") title = text.splitlines()[0] # "# Hall name" rel = str(path.relative_to(KB_DIR)) for i, section in enumerate(split_sections(text)): # every chunk carries the hall name so it can be found by meaning doc = section if i == 0 else title + "\n" + section ids.append(f"{rel}#{i}") docs.append(doc) metas.append({"source": rel}) if ids: col.upsert(ids=ids, documents=docs, metadatas=metas) return len(ids) def search(query: str, n: int = 4) -> list[dict]: if col.count() == 0: return [] res = col.query(query_texts=[query], n_results=min(n, col.count())) return [{"text": d, "source": m["source"]} for d, m in zip(res["documents"][0], res["metadatas"][0])]If you delete or rename a file, its old chunks stay in the base. For this lesson it is enough to delete the
db/chromafolder and index again. -
Calendar and the availability function
Why a function: bookings change every day and must not live in text the model reads. The function validates the dates and the hall and returns "booked" or "no_booking_in_records". The second value only means "there is no entry in the table" — not "free". That is why a person gives the final confirmation.
python · calendar_db.py · not run# calendar_db.py · bookings in SQLite (invented) import sqlite3 from contextlib import closing from datetime import date DB_PATH = "./db/bookings.db" HALLS = {"main": 1200, "conf_a": 400, "conf_b": 200, "workshop": 80} # seats (invented) def init_db() -> None: with closing(sqlite3.connect(DB_PATH)) as conn, conn: conn.execute(""" CREATE TABLE IF NOT EXISTS bookings ( id INTEGER PRIMARY KEY AUTOINCREMENT, hall_id TEXT NOT NULL, title TEXT, date_from TEXT NOT NULL, -- YYYY-MM-DD date_to TEXT NOT NULL, status TEXT NOT NULL DEFAULT 'confirmed' -- confirmed / tentative / cancelled )""") def check_availability(hall_id: str, date_from: str, date_to: str) -> dict: """Returns "booked" or "no_booking_in_records". The second value is NOT a confirmation.""" if hall_id not in HALLS: return {"status": "error", "detail": "unknown hall_id", "known": list(HALLS)} try: d1, d2 = date.fromisoformat(date_from), date.fromisoformat(date_to) except ValueError: return {"status": "error", "detail": "dates must be YYYY-MM-DD"} if d2 < d1: return {"status": "error", "detail": "date_to is before date_from"} with closing(sqlite3.connect(DB_PATH)) as conn: rows = conn.execute( "SELECT title, date_from, date_to, status FROM bookings " "WHERE hall_id = ? AND status != 'cancelled' " "AND date_from <= ? AND date_to >= ?", (hall_id, d2.isoformat(), d1.isoformat()), ).fetchall() return { "hall_id": hall_id, "status": "booked" if rows else "no_booking_in_records", "conflicts": [{"title": r[0], "from": r[1], "to": r[2], "state": r[3]} for r in rows], } -
Retrieval + tool calling
How it works: the model receives the chunks as context and a description of the function. If it decides to call it, we run it and return the result as a message with the role "tool"; then the model writes the answer. That is the pattern in the Ollama tool-calling documentation. The loop is limited to 3 rounds.
python · rag.py · not run# rag.py · retrieval + tool calling import json from datetime import date from ollama import chat from calendar_db import check_availability from kb import search MODEL = "llama3.1:8b" TOOLS = [{ "type": "function", "function": { "name": "check_availability", "description": "Check whether a hall has a booking in a date range, from the booking records only", "parameters": { "type": "object", "required": ["hall_id", "date_from", "date_to"], "properties": { "hall_id": {"type": "string", "description": "main, conf_a, conf_b or workshop"}, "date_from": {"type": "string", "description": "YYYY-MM-DD"}, "date_to": {"type": "string", "description": "YYYY-MM-DD"}, }, }, }, }] FUNCS = {"check_availability": check_availability} SYSTEM = ( "Ти си помощник за зали в културен и спортен комплекс. Отговаряй на езика на въпроса " "(български или английски). За размери, оборудване и цени използвай САМО КОНТЕКСТА по-долу; " "ако нещо липсва, кажи го. За наличност на дата ВИНАГИ извикай check_availability, не гадай. " "Наличността е по записите в базата, не е потвърждение. Не обещавай нищо на клиент. " "Днес е {today}." ) def ask(question: str, history: list[dict]) -> dict: chunks = search(question) context = "\n\n".join(f"[{c['source']}]\n{c['text']}" for c in chunks) system = SYSTEM.format(today=date.today().isoformat()) + "\n\nКОНТЕКСТ:\n" + context clean = [m for m in history[-6:] if m.get("role") in ("user", "assistant")] messages = [{"role": "system", "content": system}, *clean, {"role": "user", "content": question}] for _ in range(3): # at most 3 rounds of "model → function → model" resp = chat(model=MODEL, messages=messages, tools=TOOLS) messages.append(resp.message) if not resp.message.tool_calls: return {"answer": resp.message.content, "sources": sorted({c["source"] for c in chunks})} for call in resp.message.tool_calls: fn = FUNCS.get(call.function.name) if fn is None: result = {"status": "error", "detail": "unknown tool"} else: try: result = fn(**call.function.arguments) except TypeError: result = {"status": "error", "detail": "bad arguments"} messages.append({"role": "tool", "tool_name": call.function.name, "content": json.dumps(result, ensure_ascii=False)}) return {"answer": "Не стигнах до отговор. Опитай с по-конкретен въпрос.", "sources": []}⚠️The small model makes mistakesA model with 8 billion parameters sometimes does not call the function, passes a wrong hall identifier or a wrong date format. The code returns an error instead of crashing, but it cannot force the model to ask. So: try at least ten questions of your own about dates, write down the misses and never show an availability answer to a client without a person checking it. -
API and chat page
Why this way: the server listens on
127.0.0.1only and has no login — do not publish it to the network. The page inserts text withtextContent, not as HTML. History from the client is accepted only with the roles "user" and "assistant", so nobody can slip in their own system instructions. The port is arbitrary.python · main.py · not run# main.py · API and a small chat page from contextlib import asynccontextmanager from fastapi import FastAPI from fastapi.responses import HTMLResponse from pydantic import BaseModel, Field from calendar_db import check_availability, init_db from kb import index_kb from rag import ask @asynccontextmanager async def lifespan(_: FastAPI): init_db() print("indexed sections:", index_kb()) yield app = FastAPI(title="Venue RAG", lifespan=lifespan) class ChatRequest(BaseModel): message: str = Field(max_length=2000) history: list[dict] = Field(default_factory=list, max_length=12) PAGE = """<!doctype html><meta charset="utf-8"><title>Venue RAG</title> <div id="log"></div> <p><input id="m" size="60" placeholder="E.g. Which hall is free on 2026-11-14 for 300 people?"> <button id="b">Send</button></p> <script> let history = []; const log = document.getElementById("log"); function add(who, text) { const p = document.createElement("p"); p.textContent = who + ": " + text; log.appendChild(p); } document.getElementById("b").onclick = async () => { const m = document.getElementById("m"); const message = m.value; m.value = ""; add("You", message); const r = await fetch("/chat", {method: "POST", headers: {"Content-Type": "application/json"}, body: JSON.stringify({message, history})}); const data = await r.json(); add("Assistant", data.answer + (data.sources && data.sources.length ? " [" + data.sources.join(", ") + "]" : "")); history.push({role: "user", content: message}, {role: "assistant", content: data.answer}); history = history.slice(-12); }; </script>""" @app.get("/") def index() -> HTMLResponse: return HTMLResponse(PAGE) @app.post("/chat") def chat_endpoint(req: ChatRequest) -> dict: return ask(req.message, req.history) @app.get("/availability/{hall_id}") def availability(hall_id: str, date_from: str, date_to: str) -> dict: return check_availability(hall_id, date_from, date_to)bash · on the machine · not runuvicorn main:app --host 127.0.0.1 --port 8156From your own computer you reach the page through an SSH tunnel:
bash · on your computer · not runssh -L 8156:localhost:8156 <user>@<server-address>Then open
http://localhost:8156. -
Try it with invented questions
Add a few invented bookings to the table and ask the questions below. The last two test "I do not know": if the assistant makes up an answer, tighten the instructions in
SYSTEMand try again.Question (invented) What to expect How many seats does the Main hall have? an answer from the base with source halls/main.md, no functionIs Conference hall A free on 2026-11-14? a function call; the answer "no entry" or "booked" How much does a hall for 5000 people cost? "there is no such hall in the base" — not an invented number What is the refund policy? "the base has no information about this" -
Limits and safety
- Every availability and price is checked by a person before it reaches a client.
- The base holds no personal data. If you later add tenant or contact data, that is personal data processing and has to be agreed with data protection in your organisation.
- The chat has no login — keep it behind the tunnel.
- The base of chunks is maintained: a "Valid from" date, an owner, a review when prices change.
04Check
ollama listshowsllama3.1:8bandbge-m3.- After start the print shows the number of indexed sections (at least one per hall).
- A capacity question returns a
halls/…source; a question outside the base gets "no information". - A wrong date or an unknown hall returns an error, not a crash.
- The server listens on
127.0.0.1only.
Quiz
1. Why is availability taken with a function and not from text in the base?
2. What does "no_booking_in_records" mean?
3. How much memory do the two models in the lesson take (per ollama.com)?
4. What do you do when the small model does not call the function for a date?
05What next
06Sources
- Ollama: llama3.1 · Ollama: bge-m3 🔒 local — file sizes, the "tools" tag, multilingual support (read on 03.10.2026).
- Ollama: tool calling — tool schema, the "tool" role message, the loop (read on 03.10.2026).
- ollama (Python) 0.6.3 — version on PyPI as of 03.10.2026.
- Chroma: embedding function through Ollama —
OllamaEmbeddingFunction(url, model_name). - FastAPI: lifespan events —
asynccontextmanager.