The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewVERIFIED 2026-10-01 · UPDATED 2026-10-01
IDENTITY
module
01-00c · Lab: first local LLM, live
series
KAGAMI Academy · Blocks 0–10 · Block 0 — AI Fundamentals · part 3 of 3 (series colour: mint)
level
Beginner
duration
2–4 h hands-on
trust_label
VERIFIED 2026-10-01 (model sizes on ollama.com, published DGX Spark decode speeds, PyTorch/CUDA wheel index, BgGPT 3.0 GGUF, LangChain packages, all source links) · UPDATED 2026-10-01 (speed figures, 70B memory math, install commands, BgGPT 3.0, fictional demo document, EUR)
prerequisites
01-00a (tokenisation) · 01-00b (AI stack)
language
human view: en · bulgarian edition: /academy/blokove/moduli/01-00c_Блок_0_Част_3_Лаборатория.html
next
01-01a · Block 1 part 1 · vLLM inference
PURPOSE

Install a local AI environment (GB10-class desktop such as DGX Spark / ASUS Ascent GX10, Ubuntu + NVIDIA GPU, or Apple Silicon), run a model with Ollama, measure decode speed honestly, build a mini RAG that answers "I don't know" outside its documents, call the model from Python three ways, and A/B-test BgGPT 3.0 on Bulgarian tasks.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

01-01a · Block 1 part 1 · vLLM inference (local deployment and optimisation) · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
local-llmollamabenchmarkraglangchainbggptdgx-sparkgb10python-api
VERIFIED · 01.10.2026 UPDATED · 01.10.2026

Lab: your first local LLM, live

Theory is done. Now we install the environment, run a real language model, measure its speed honestly and solve a first business scenario — fully local, no API key, no cloud. This is the final part of Block 0.

⏱ 2–4 h hands-on Beginner Block 0 · Part 3 of 3 Ollama · RAG · Python API
Ollama🔒 local Python + LangChain + Chroma🔒 local BgGPT 3.0🔒 local Hugging Face (download)🌐 global
🔄
UPDATED · 01.10.2026 — what changed
Speed figures now come from Ollama's published DGX Spark measurements (a 70B model decodes at ~4.4 tok/s, not ~140). The memory maths is fixed: 70B at full precision (~141 GB) does not fit in 128 GB — Ollama pulls Q4 (~43 GB) by default. Installation now uses a virtual environment (Ubuntu 24.04 blocks system-wide pip) and PyTorch's CUDA 13 wheels. BgGPT 3.0 now ships official GGUF files — no manual conversion. The RAG demo document is a fictional rulebook (the old one contained inaccurate legal text). Amounts are in euro.

01What you will learn

02Before you start

03Steps

  1. Set up the environment (~20 min)

    The procedure differs slightly per platform. GB10 machines ship with DGX OS (Ubuntu with NVIDIA drivers and CUDA) — there you only check and top up. Why a virtual environment: since Ubuntu 24.04 the system Python is "externally managed" (PEP 668) and pip install outside a venv fails. The lesson's packages live in their own folder and don't interfere with the system.

    bash · GB10 machine (DGX OS) — check and top up
    # 1. System update (NVIDIA also recommends it for firmware)
    sudo apt update && sudo apt dist-upgrade -y
    
    # 2. Is the chip visible?
    nvidia-smi              # you should see NVIDIA GB10
    
    # 3. Ollama — check; install if missing
    ollama --version || curl -fsSL https://ollama.com/install.sh | sh
    
    # 4. Working folder and virtual environment for the lesson
    mkdir -p ~/ai-lab/{models,datasets,outputs} && cd ~/ai-lab
    python3 -m venv .venv && source .venv/bin/activate
    pip install --upgrade pip
    
    # 5. PyTorch with CUDA 13 (wheels exist for ARM64/aarch64 too)
    pip install torch --index-url https://download.pytorch.org/whl/cu130
    python -c "import torch;print(torch.__version__, torch.cuda.is_available(), torch.cuda.get_device_name(0))"
    
    # 6. The lesson's packages
    pip install ollama openai rich requests langchain langchain-community \
                langchain-ollama langchain-chroma langchain-text-splitters
    echo "✅ Ready"
    bash · Ubuntu 24.04 with an NVIDIA card — from scratch
    # 1. Driver (CUDA 13 wheels need a 580-series driver or newer)
    sudo apt update && sudo apt install -y ubuntu-drivers-common python3-venv
    sudo ubuntu-drivers install && sudo reboot
    nvidia-smi              # after reboot: you see the card and driver version
    
    # 2. A separate CUDA Toolkit is NOT needed — PyTorch and Ollama bundle their libraries.
    #    You only need it to compile CUDA code (nvcc).
    
    # 3. Virtual environment + PyTorch
    mkdir -p ~/ai-lab && cd ~/ai-lab
    python3 -m venv .venv && source .venv/bin/activate
    pip install torch --index-url https://download.pytorch.org/whl/cu130
    
    # 4. Ollama
    curl -fsSL https://ollama.com/install.sh | sh
    
    # 5. The lesson's packages
    pip install ollama openai rich requests langchain langchain-community \
                langchain-ollama langchain-chroma langchain-text-splitters
    bash · macOS with Apple Silicon
    # Ollama: the app from ollama.com or via Homebrew
    brew install ollama && ollama serve &
    
    mkdir -p ~/ai-lab && cd ~/ai-lab
    python3 -m venv .venv && source .venv/bin/activate
    pip install torch ollama openai rich requests langchain langchain-community \
                langchain-ollama langchain-chroma langchain-text-splitters
    
    # On a Mac acceleration is MPS (Metal), not CUDA
    python -c "import torch;print('MPS:', torch.backends.mps.is_available())"
    ✅
    Rule: one environment per project
    Every lab gets its own folder and its own .venv. When something breaks, you delete the environment, not the system.
  2. Ollama — your first model (~15 min)

    Ollama 🔒 local works like docker pull for language models: one command, wait for the download, the model is ready. The choice depends on memory and on the task. Sizes below are the default tags in the Ollama library (checked 01.10.2026).

    ModelSizeGood forGB10 (128 GB)24 GB card
    llama3.2:3b2.0 GBquick tests, prototypes✅✅
    llama3.1:8b4.9 GBeveryday work, RAG✅✅
    mistral:7b4.4 GBcode, structured output✅✅
    llama3.1:70b (Q4)43 GBhigher quality, slow✅❌ doesn't fit
    nomic-embed-text274 MBembeddings for RAG✅✅
    ⚠️
    Trap: "70B at full precision"
    70 billion parameters × 2 bytes (FP16/BF16) ≈ 141 GB — that does not fit in 128 GB of unified memory. The llama3.1:70b tag in Ollama is quantised to Q4 (~43 GB); Q8 is ~75 GB. That is the normal choice for a desktop machine.
    bash · download and first chat
    # Any machine: 8B is the right starting point
    ollama pull llama3.1:8b
    ollama pull nomic-embed-text      # for RAG in step 5
    
    # Only with ~50 GB of free memory (e.g. GB10): 70B for comparison
    ollama pull llama3.1:70b
    
    ollama list
    ollama run llama3.1:8b "Explain in 3 sentences what a RAG system is."
    ollama run llama3.1:8b            # interactive mode; exit with /bye
  3. Modelfile — an assistant with character

    A Modelfile sets the system prompt, temperature and context length. It is the first step toward a specialised assistant for a given sector — without training the model at all.

    bash · Bulgarian business assistant
    cat > Modelfile <<'EOF'
    FROM llama3.1:8b
    PARAMETER temperature 0.2
    PARAMETER num_ctx 8192
    PARAMETER top_p 0.9
    SYSTEM """You are an AI assistant for Bulgarian businesses.
    Answer ONLY in Bulgarian.
    Give concrete, practical answers.
    If you don't know something, say so directly.
    Never invent facts or figures.
    For legal and medical questions, refer to a specialist."""
    EOF
    
    ollama create bg-assistant -f Modelfile
    ollama run bg-assistant "What is the difference between an EOOD and an AD?"
  4. Measuring performance (~20 min)

    An architect doesn't just run the model — they measure it. Writing speed (tokens/s), time to first token and memory used are the numbers you use to justify hardware and model choices to a client. Why this way: Ollama reports prompt-reading time and writing time separately; if you divide tokens by the total time, you mix the two and the number lies.

    Python · benchmark.py
    import requests
    from rich.console import Console
    from rich.table import Table
    
    console = Console()
    URL = "http://localhost:11434/api/generate"
    
    def bench(model, prompt):
        d = requests.post(URL, json={"model": model, "prompt": prompt,
                                     "stream": False}, timeout=600).json()
        ns = 1e9
        out_tok = d.get("eval_count", 0)
        decode_s = d.get("eval_duration", 0) / ns
        return {
            "tokens": out_tok,
            "tok_s": round(out_tok / decode_s, 1) if decode_s else 0,   # writing speed
            "first_s": round((d.get("load_duration", 0) + d.get("prompt_eval_duration", 0)) / ns, 2),
            "total_s": round(d.get("total_duration", 0) / ns, 2),
        }
    
    PROMPTS = [
        "Say 'hello' in Bulgarian.",
        "Explain inflation in 5 sentences.",
        "Write a short business plan for an open-air café in Varna.",
        "What are the differences between the GDPR and the EU AI Act?",
    ]
    MODEL = "llama3.1:8b"   # then repeat with llama3.1:70b if you have it
    
    bench(MODEL, "hello")   # warm-up: loads the model into memory
    t = Table(title=f"Benchmark — {MODEL}")
    for c in ("Prompt", "Tokens", "tok/s", "1st token (s)", "Total (s)"):
        t.add_column(c)
    rows = [(p, bench(MODEL, p)) for p in PROMPTS]
    for p, r in rows:
        t.add_row(p[:38], str(r["tokens"]), str(r["tok_s"]), str(r["first_s"]), str(r["total_s"]))
    console.print(t)
    console.print(f"Average: {sum(r['tok_s'] for _, r in rows)/len(rows):.1f} tok/s")

    Reference — what to expect. Ollama's published measurements on DGX Spark (same GB10 chip; October 2025, 500 output tokens):

    ModelQuantisationWriting (tok/s)200-token answer
    llama3.1 8BQ4_K_M~38~5 s
    gemma3 27BQ4_K_M~11~18 s
    llama3.1 70BQ4_K_M~4.4~45 s
    gpt-oss 20B (MoE)MXFP4~58~3.5 s
    gpt-oss 120B (MoE)MXFP4~41~5 s
    💡
    Why 70B is slow and 120B MoE is fast
    For every new token a dense model reads all of its weights from memory. At ~273 GB/s of bandwidth and ~43 GB of weights the ceiling is about 6 tok/s — that's why 70B gives ~4.4. Mixture-of-experts (MoE) models activate only a small share of their weights per token, so 120B writes faster than 70B. The takeaway for the client: for real-time chat on such a machine choose 8B–27B or MoE; 70B is for overnight batch jobs.
    ✅
    Business value: calculate, don't guess
    A local model has no per-request price — you pay for the hardware and electricity, and the data never leaves the building (important under the GDPR). A cloud API charges per million tokens. Take the current price from the provider's page 🌐 global and calculate: requests per day × tokens per request × 365 × price per token. Compare with the machine's cost spread over 3 years. The payback point is the number the client wants to see.
  5. Real scenario: RAG instead of fine-tuning

    The theory from Part 1: a language model hallucinates when the question lies outside the data it was trained on. Regulations and internal rules change often — no model knows last month's version. Without RAG: confidently wrong answers. With RAG: accurate answers with a citation.

    ⛔
    A typical trap (a real-world case, anonymised)
    An accounting firm spent ~€7,700 (BGN 15,000) fine-tuning a cloud model to automate queries on health-insurance rules. The model confidently returned wrong amounts: the rules change often and the data was already stale at launch. The right approach — RAG over an up-to-date base of official documents + a local model + automatic refresh — costs several times less (on the order of ~€1,000). Conclusion: fine-tuning for frequently changing information is almost always the wrong choice.

    A mini RAG in 15 minutes. The document below is a fictional internal rulebook of a sample company — deliberately, so we don't feed the model (or you) inaccurate legal text. Replace it with a real document of your own.

    Python · mini_rag.py
    from langchain_community.document_loaders import TextLoader
    from langchain_text_splitters import RecursiveCharacterTextSplitter
    from langchain_ollama import OllamaEmbeddings, ChatOllama
    from langchain_chroma import Chroma
    from langchain_core.prompts import ChatPromptTemplate
    from langchain_core.runnables import RunnablePassthrough
    from langchain_core.output_parsers import StrOutputParser
    
    # 1. Test document — a FICTIONAL rulebook of a sample company
    with open("rulebook.txt", "w", encoding="utf-8") as f:
        f.write("""INTERNAL RULEBOOK OF "EXAMPLE" LTD. (2026 edition)
    Art. 3. Working hours are 9:00 to 17:30, with a 30-minute lunch break.
    Art. 7. Travel expenses are reported within 5 working days of return.
    Art. 9. Employees may work from home 2 days a week,
    subject to agreement with their line manager.
    Art. 12. Leave requests are filed at least 10 working days in advance.
    Art. 15. Company laptops are returned within 3 days of the contract ending.""")
    
    # 2. Load and split
    docs = TextLoader("rulebook.txt", encoding="utf-8").load()
    chunks = RecursiveCharacterTextSplitter(chunk_size=300, chunk_overlap=50).split_documents(docs)
    print(f"✅ Chunks: {len(chunks)}")
    
    # 3. Embeddings and vector store (in memory)
    store = Chroma.from_documents(chunks, OllamaEmbeddings(model="nomic-embed-text"))
    retriever = store.as_retriever(search_kwargs={"k": 3})
    
    # 4. Model and a strict prompt
    llm = ChatOllama(model="llama3.1:8b", temperature=0.1)
    prompt = ChatPromptTemplate.from_template("""Answer ONLY from the context.
    If the answer is not in the context, say "I don't know from the documents."
    Cite the article.
    Context: {context}
    Question: {question}""")
    
    chain = ({"context": retriever, "question": RunnablePassthrough()}
             | prompt | llm | StrOutputParser())
    
    for q in ["How many work-from-home days are allowed?",
              "When must a leave request be filed?",
              "When is the company laptop returned?",
              "What is the VAT rate in Bulgaria?"]:   # not in the document!
        print(f"\n❓ {q}\n💬 {chain.invoke(q)}")
    Expected output (shortened; wording varies)
    ❓ How many work-from-home days are allowed?
    💬 Under Art. 9 — 2 days a week, subject to agreement with the line manager.
    
    ❓ What is the VAT rate in Bulgaria?
    💬 I don't know from the documents.
    ✅
    The key observation
    The last question shows the most important property of a good RAG: saying "I don't know" instead of inventing. Without RAG the model would answer confidently — sometimes right, sometimes wrong, and you can't tell which. Clients trust systems that know their limits.

    Take it to your sector — only the document changes:

    SectorDocumentsTypical questionExtra measures
    🏥 Healthcareprotocols, internal rules"What is the admission procedure?"guardrails, audit log
    ⚖️ Lawcontracts, statutes"Find the risks in this clause"article citations
    💼 Accountingregulations, VAT Act, Corporate Income Tax Act"How do I book this invoice?"refresh on every change
    🤝 Social servicesstatute, internal procedures"Which documents are needed for…"BgGPT 3.0 for Bulgarian terms
    🌱 ESG / green transitionEU Taxonomy, climate regulations"Does this activity meet a green criterion?"GraphRAG for relations
  6. Python API — three ways (~30 min)

    The command line is for trying things out. Real systems call the model from code — through the ollama library or through Ollama's OpenAI-compatible interface 🔒 local. The second is valuable: code written for a cloud API moves to a local model by changing two lines.

    Python · api_three_ways.py
    import json, ollama
    from openai import OpenAI
    
    MODEL = "llama3.1:8b"
    
    # 1. The ollama library
    r = ollama.generate(model=MODEL, prompt="Explain tokenisation in 2 sentences.",
                        options={"temperature": 0.2, "num_ctx": 4096})
    print(r["response"])
    
    # 2. Streaming (for a chat interface)
    for part in ollama.generate(model=MODEL, prompt="A short story about AI in 100 words.", stream=True):
        print(part["response"], end="", flush=True)
    
    # 3. OpenAI-compatible interface — the key is just a placeholder
    client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
    resp = client.chat.completions.create(model=MODEL, temperature=0.2, messages=[
        {"role": "system", "content": "You are an HR assistant."},
        {"role": "user", "content": "How is severance pay calculated?"}])
    print(resp.choices[0].message.content)
    
    # Bonus: structured JSON from an invoice — the format is a JSON schema
    schema = {"type": "object", "required": ["vendor", "amount", "currency", "date", "vat_included"],
              "properties": {"vendor": {"type": "string"}, "amount": {"type": "number"},
                             "currency": {"type": "string"}, "date": {"type": "string"},
                             "vat_included": {"type": "boolean"}}}
    r = ollama.generate(model=MODEL, format=schema, options={"temperature": 0},
        prompt="Extract the data. Invoice: Vendor: Technologii Ltd., Amount: 1,230 EUR incl. VAT, Date: 15.03.2026")
    try:
        print(json.loads(r["response"]))
    except json.JSONDecodeError:
        print("⚠️ Invalid JSON — retry the request or send it to a human")
    # {'vendor': 'Technologii Ltd.', 'amount': 1230, 'currency': 'EUR', 'date': '2026-03-15', 'vat_included': True}
    ⚠️
    JSON mode is not a guarantee
    format="json" or a JSON schema sharply reduce invalid output, but not to zero — and valid JSON can still carry a wrong value. In production: try/except, schema validation (e.g. Pydantic), a retry, and human review where money or personal data are involved.
  7. Bonus: BgGPT 3.0 vs Llama in Bulgarian (~15 min)

    BgGPT 3.0 by INSAIT is a series of models adapted to Bulgarian, built on Gemma 3 (4B, 12B and 27B). Official GGUF files exist — Ollama 🔒 local pulls them straight from Hugging Face 🌐 global, no manual conversion. The model then runs fully locally.

    bash · download BgGPT 3.0
    # 27B, Q4_K_M (~17 GB) — for GB10 or a 24 GB card
    ollama pull hf.co/INSAIT-Institute/BgGPT-Gemma-3-27B-IT-GGUF:Q4_K_M
    # Smaller machine: BgGPT-Gemma-3-12B-IT-GGUF or BgGPT-Gemma-3-4B-IT-GGUF
    Python · bggpt_ab.py — A/B test
    import ollama
    
    # Tasks stay in Bulgarian — that is what we are testing
    TASKS = {
        "Terminology": "Обясни разликата между 'социален работник' и 'социален педагог'.",
        "Invoice": "Какви реквизити трябва да има една фактура според ЗДДС?",
        "Formal letter": "Напиши кратко официално писмо до общината за достъп до обществена информация.",
    }
    MODELS = {
        "BgGPT 3.0 27B": "hf.co/INSAIT-Institute/BgGPT-Gemma-3-27B-IT-GGUF:Q4_K_M",
        "Llama 3.1 8B": "llama3.1:8b",
    }
    for name, prompt in TASKS.items():
        print(f"\n{'='*60}\n📋 {name}: {prompt}")
        for label, mid in MODELS.items():
            try:
                r = ollama.generate(model=mid, prompt=prompt, options={"temperature": 0.1})
                print(f"\n[{label}]\n{r['response'][:400]}…")
            except Exception as e:
                print(f"\n[{label}] unavailable: {e}")
    👤
    Who judges
    The comparison is not a number but a review by someone who knows the subject: grammar, terminology, factual accuracy. Check the facts in the answers (invoice requirements, deadlines) against the primary source — both models can be wrong.

04Check

Practical challenge — before Block 1:

1. Your RAG assistant answers 8 of 10 questions correctly, but for "What is the electricity price?" it confidently gives a wrong price. The most likely reason?

2. A client wants a 70B model on a machine with 128 GB of unified memory. Which is true?

3. How do you measure writing speed correctly?

4. Does Ollama's JSON mode guarantee 100% valid and correct output?

05What's next

That completes Block 0: you have a local environment, measured numbers, a RAG that knows its limits, and three ways to call the model from code.

06Sources

  1. Ollama library — models, tags and sizes.
  2. Ollama: NVIDIA DGX Spark performance — the measurements in step 4.
  3. Ollama API · OpenAI compatibility · structured outputs · Modelfile.
  4. ollama-python — the official Python library.
  5. LangChain + Ollama · Chroma.
  6. BgGPT-Gemma-3-27B-IT-GGUF · BgGPT 3.0 announcement — INSAIT.
  7. Hugging Face: models in Ollama — pulling with hf.co/….
  8. PyTorch: local install · PEP 668 — why a virtual environment.
  9. NVIDIA DGX Spark documentation · vLLM — for Block 1.