Lab: your first local LLM, live
Theory is done. Now we install the environment, run a real language model, measure its speed honestly and solve a first business scenario — fully local, no API key, no cloud. This is the final part of Block 0.
01What you will learn
- Set up a working local AI environment on a GB10 machine, on Ubuntu with an NVIDIA card, or on a Mac.
- Run a model with Ollama 🔒 local and give it a "character" with a Modelfile.
- Measure speed correctly — and understand why memory, not "power", decides how fast a model writes.
- Build a mini RAG that says "I don't know" instead of making things up.
- Call the model from Python in three ways and extract structured JSON.
- Compare BgGPT 3.0 and Llama on real Bulgarian tasks.
02Before you start
- Block 0 parts 1 and 2 done: tokenisation and the AI stack.
- One of three machines: a GB10 desktop (NVIDIA DGX Spark, ASUS Ascent GX10 and similar, 128 GB unified memory), Linux with an NVIDIA card, or a Mac with Apple Silicon (16 GB or more).
- Internet only to download the models — after that everything runs offline.
- Free disk space: ~10 GB for the small models; ~50 GB if you want to try 70B.
- Terminal basics and Python 3.10+.
03Steps
-
Set up the environment (~20 min)
The procedure differs slightly per platform. GB10 machines ship with DGX OS (Ubuntu with NVIDIA drivers and CUDA) — there you only check and top up. Why a virtual environment: since Ubuntu 24.04 the system Python is "externally managed" (PEP 668) and
pip installoutside a venv fails. The lesson's packages live in their own folder and don't interfere with the system.bash · GB10 machine (DGX OS) — check and top up# 1. System update (NVIDIA also recommends it for firmware) sudo apt update && sudo apt dist-upgrade -y # 2. Is the chip visible? nvidia-smi # you should see NVIDIA GB10 # 3. Ollama — check; install if missing ollama --version || curl -fsSL https://ollama.com/install.sh | sh # 4. Working folder and virtual environment for the lesson mkdir -p ~/ai-lab/{models,datasets,outputs} && cd ~/ai-lab python3 -m venv .venv && source .venv/bin/activate pip install --upgrade pip # 5. PyTorch with CUDA 13 (wheels exist for ARM64/aarch64 too) pip install torch --index-url https://download.pytorch.org/whl/cu130 python -c "import torch;print(torch.__version__, torch.cuda.is_available(), torch.cuda.get_device_name(0))" # 6. The lesson's packages pip install ollama openai rich requests langchain langchain-community \ langchain-ollama langchain-chroma langchain-text-splitters echo "✅ Ready"bash · Ubuntu 24.04 with an NVIDIA card — from scratch# 1. Driver (CUDA 13 wheels need a 580-series driver or newer) sudo apt update && sudo apt install -y ubuntu-drivers-common python3-venv sudo ubuntu-drivers install && sudo reboot nvidia-smi # after reboot: you see the card and driver version # 2. A separate CUDA Toolkit is NOT needed — PyTorch and Ollama bundle their libraries. # You only need it to compile CUDA code (nvcc). # 3. Virtual environment + PyTorch mkdir -p ~/ai-lab && cd ~/ai-lab python3 -m venv .venv && source .venv/bin/activate pip install torch --index-url https://download.pytorch.org/whl/cu130 # 4. Ollama curl -fsSL https://ollama.com/install.sh | sh # 5. The lesson's packages pip install ollama openai rich requests langchain langchain-community \ langchain-ollama langchain-chroma langchain-text-splittersbash · macOS with Apple Silicon# Ollama: the app from ollama.com or via Homebrew brew install ollama && ollama serve & mkdir -p ~/ai-lab && cd ~/ai-lab python3 -m venv .venv && source .venv/bin/activate pip install torch ollama openai rich requests langchain langchain-community \ langchain-ollama langchain-chroma langchain-text-splitters # On a Mac acceleration is MPS (Metal), not CUDA python -c "import torch;print('MPS:', torch.backends.mps.is_available())"✅Rule: one environment per projectEvery lab gets its own folder and its own.venv. When something breaks, you delete the environment, not the system. -
Ollama — your first model (~15 min)
Ollama 🔒 local works like
docker pullfor language models: one command, wait for the download, the model is ready. The choice depends on memory and on the task. Sizes below are the default tags in the Ollama library (checked 01.10.2026).Model Size Good for GB10 (128 GB) 24 GB card llama3.2:3b2.0 GB quick tests, prototypes ✅ ✅ llama3.1:8b4.9 GB everyday work, RAG ✅ ✅ mistral:7b4.4 GB code, structured output ✅ ✅ llama3.1:70b(Q4)43 GB higher quality, slow ✅ ❌ doesn't fit nomic-embed-text274 MB embeddings for RAG ✅ ✅ ⚠️Trap: "70B at full precision"70 billion parameters × 2 bytes (FP16/BF16) ≈ 141 GB — that does not fit in 128 GB of unified memory. Thellama3.1:70btag in Ollama is quantised to Q4 (~43 GB); Q8 is ~75 GB. That is the normal choice for a desktop machine.bash · download and first chat# Any machine: 8B is the right starting point ollama pull llama3.1:8b ollama pull nomic-embed-text # for RAG in step 5 # Only with ~50 GB of free memory (e.g. GB10): 70B for comparison ollama pull llama3.1:70b ollama list ollama run llama3.1:8b "Explain in 3 sentences what a RAG system is." ollama run llama3.1:8b # interactive mode; exit with /bye -
Modelfile — an assistant with character
A Modelfile sets the system prompt, temperature and context length. It is the first step toward a specialised assistant for a given sector — without training the model at all.
bash · Bulgarian business assistantcat > Modelfile <<'EOF' FROM llama3.1:8b PARAMETER temperature 0.2 PARAMETER num_ctx 8192 PARAMETER top_p 0.9 SYSTEM """You are an AI assistant for Bulgarian businesses. Answer ONLY in Bulgarian. Give concrete, practical answers. If you don't know something, say so directly. Never invent facts or figures. For legal and medical questions, refer to a specialist.""" EOF ollama create bg-assistant -f Modelfile ollama run bg-assistant "What is the difference between an EOOD and an AD?" -
Measuring performance (~20 min)
An architect doesn't just run the model — they measure it. Writing speed (tokens/s), time to first token and memory used are the numbers you use to justify hardware and model choices to a client. Why this way: Ollama reports prompt-reading time and writing time separately; if you divide tokens by the total time, you mix the two and the number lies.
Python · benchmark.pyimport requests from rich.console import Console from rich.table import Table console = Console() URL = "http://localhost:11434/api/generate" def bench(model, prompt): d = requests.post(URL, json={"model": model, "prompt": prompt, "stream": False}, timeout=600).json() ns = 1e9 out_tok = d.get("eval_count", 0) decode_s = d.get("eval_duration", 0) / ns return { "tokens": out_tok, "tok_s": round(out_tok / decode_s, 1) if decode_s else 0, # writing speed "first_s": round((d.get("load_duration", 0) + d.get("prompt_eval_duration", 0)) / ns, 2), "total_s": round(d.get("total_duration", 0) / ns, 2), } PROMPTS = [ "Say 'hello' in Bulgarian.", "Explain inflation in 5 sentences.", "Write a short business plan for an open-air café in Varna.", "What are the differences between the GDPR and the EU AI Act?", ] MODEL = "llama3.1:8b" # then repeat with llama3.1:70b if you have it bench(MODEL, "hello") # warm-up: loads the model into memory t = Table(title=f"Benchmark — {MODEL}") for c in ("Prompt", "Tokens", "tok/s", "1st token (s)", "Total (s)"): t.add_column(c) rows = [(p, bench(MODEL, p)) for p in PROMPTS] for p, r in rows: t.add_row(p[:38], str(r["tokens"]), str(r["tok_s"]), str(r["first_s"]), str(r["total_s"])) console.print(t) console.print(f"Average: {sum(r['tok_s'] for _, r in rows)/len(rows):.1f} tok/s")Reference — what to expect. Ollama's published measurements on DGX Spark (same GB10 chip; October 2025, 500 output tokens):
Model Quantisation Writing (tok/s) 200-token answer llama3.1 8B Q4_K_M ~38 ~5 s gemma3 27B Q4_K_M ~11 ~18 s llama3.1 70B Q4_K_M ~4.4 ~45 s gpt-oss 20B (MoE) MXFP4 ~58 ~3.5 s gpt-oss 120B (MoE) MXFP4 ~41 ~5 s 💡Why 70B is slow and 120B MoE is fastFor every new token a dense model reads all of its weights from memory. At ~273 GB/s of bandwidth and ~43 GB of weights the ceiling is about 6 tok/s — that's why 70B gives ~4.4. Mixture-of-experts (MoE) models activate only a small share of their weights per token, so 120B writes faster than 70B. The takeaway for the client: for real-time chat on such a machine choose 8B–27B or MoE; 70B is for overnight batch jobs.✅Business value: calculate, don't guessA local model has no per-request price — you pay for the hardware and electricity, and the data never leaves the building (important under the GDPR). A cloud API charges per million tokens. Take the current price from the provider's page 🌐 global and calculate: requests per day × tokens per request × 365 × price per token. Compare with the machine's cost spread over 3 years. The payback point is the number the client wants to see. -
Real scenario: RAG instead of fine-tuning
The theory from Part 1: a language model hallucinates when the question lies outside the data it was trained on. Regulations and internal rules change often — no model knows last month's version. Without RAG: confidently wrong answers. With RAG: accurate answers with a citation.
⛔A typical trap (a real-world case, anonymised)An accounting firm spent ~€7,700 (BGN 15,000) fine-tuning a cloud model to automate queries on health-insurance rules. The model confidently returned wrong amounts: the rules change often and the data was already stale at launch. The right approach — RAG over an up-to-date base of official documents + a local model + automatic refresh — costs several times less (on the order of ~€1,000). Conclusion: fine-tuning for frequently changing information is almost always the wrong choice.A mini RAG in 15 minutes. The document below is a fictional internal rulebook of a sample company — deliberately, so we don't feed the model (or you) inaccurate legal text. Replace it with a real document of your own.
Python · mini_rag.pyfrom langchain_community.document_loaders import TextLoader from langchain_text_splitters import RecursiveCharacterTextSplitter from langchain_ollama import OllamaEmbeddings, ChatOllama from langchain_chroma import Chroma from langchain_core.prompts import ChatPromptTemplate from langchain_core.runnables import RunnablePassthrough from langchain_core.output_parsers import StrOutputParser # 1. Test document — a FICTIONAL rulebook of a sample company with open("rulebook.txt", "w", encoding="utf-8") as f: f.write("""INTERNAL RULEBOOK OF "EXAMPLE" LTD. (2026 edition) Art. 3. Working hours are 9:00 to 17:30, with a 30-minute lunch break. Art. 7. Travel expenses are reported within 5 working days of return. Art. 9. Employees may work from home 2 days a week, subject to agreement with their line manager. Art. 12. Leave requests are filed at least 10 working days in advance. Art. 15. Company laptops are returned within 3 days of the contract ending.""") # 2. Load and split docs = TextLoader("rulebook.txt", encoding="utf-8").load() chunks = RecursiveCharacterTextSplitter(chunk_size=300, chunk_overlap=50).split_documents(docs) print(f"✅ Chunks: {len(chunks)}") # 3. Embeddings and vector store (in memory) store = Chroma.from_documents(chunks, OllamaEmbeddings(model="nomic-embed-text")) retriever = store.as_retriever(search_kwargs={"k": 3}) # 4. Model and a strict prompt llm = ChatOllama(model="llama3.1:8b", temperature=0.1) prompt = ChatPromptTemplate.from_template("""Answer ONLY from the context. If the answer is not in the context, say "I don't know from the documents." Cite the article. Context: {context} Question: {question}""") chain = ({"context": retriever, "question": RunnablePassthrough()} | prompt | llm | StrOutputParser()) for q in ["How many work-from-home days are allowed?", "When must a leave request be filed?", "When is the company laptop returned?", "What is the VAT rate in Bulgaria?"]: # not in the document! print(f"\n❓ {q}\n💬 {chain.invoke(q)}")Expected output (shortened; wording varies)❓ How many work-from-home days are allowed? 💬 Under Art. 9 — 2 days a week, subject to agreement with the line manager. ❓ What is the VAT rate in Bulgaria? 💬 I don't know from the documents.✅The key observationThe last question shows the most important property of a good RAG: saying "I don't know" instead of inventing. Without RAG the model would answer confidently — sometimes right, sometimes wrong, and you can't tell which. Clients trust systems that know their limits.Take it to your sector — only the document changes:
Sector Documents Typical question Extra measures 🏥 Healthcare protocols, internal rules "What is the admission procedure?" guardrails, audit log ⚖️ Law contracts, statutes "Find the risks in this clause" article citations 💼 Accounting regulations, VAT Act, Corporate Income Tax Act "How do I book this invoice?" refresh on every change 🤝 Social services statute, internal procedures "Which documents are needed for…" BgGPT 3.0 for Bulgarian terms 🌱 ESG / green transition EU Taxonomy, climate regulations "Does this activity meet a green criterion?" GraphRAG for relations -
Python API — three ways (~30 min)
The command line is for trying things out. Real systems call the model from code — through the
ollamalibrary or through Ollama's OpenAI-compatible interface 🔒 local. The second is valuable: code written for a cloud API moves to a local model by changing two lines.Python · api_three_ways.pyimport json, ollama from openai import OpenAI MODEL = "llama3.1:8b" # 1. The ollama library r = ollama.generate(model=MODEL, prompt="Explain tokenisation in 2 sentences.", options={"temperature": 0.2, "num_ctx": 4096}) print(r["response"]) # 2. Streaming (for a chat interface) for part in ollama.generate(model=MODEL, prompt="A short story about AI in 100 words.", stream=True): print(part["response"], end="", flush=True) # 3. OpenAI-compatible interface — the key is just a placeholder client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama") resp = client.chat.completions.create(model=MODEL, temperature=0.2, messages=[ {"role": "system", "content": "You are an HR assistant."}, {"role": "user", "content": "How is severance pay calculated?"}]) print(resp.choices[0].message.content) # Bonus: structured JSON from an invoice — the format is a JSON schema schema = {"type": "object", "required": ["vendor", "amount", "currency", "date", "vat_included"], "properties": {"vendor": {"type": "string"}, "amount": {"type": "number"}, "currency": {"type": "string"}, "date": {"type": "string"}, "vat_included": {"type": "boolean"}}} r = ollama.generate(model=MODEL, format=schema, options={"temperature": 0}, prompt="Extract the data. Invoice: Vendor: Technologii Ltd., Amount: 1,230 EUR incl. VAT, Date: 15.03.2026") try: print(json.loads(r["response"])) except json.JSONDecodeError: print("⚠️ Invalid JSON — retry the request or send it to a human") # {'vendor': 'Technologii Ltd.', 'amount': 1230, 'currency': 'EUR', 'date': '2026-03-15', 'vat_included': True}⚠️JSON mode is not a guaranteeformat="json"or a JSON schema sharply reduce invalid output, but not to zero — and valid JSON can still carry a wrong value. In production:try/except, schema validation (e.g. Pydantic), a retry, and human review where money or personal data are involved. -
Bonus: BgGPT 3.0 vs Llama in Bulgarian (~15 min)
BgGPT 3.0 by INSAIT is a series of models adapted to Bulgarian, built on Gemma 3 (4B, 12B and 27B). Official GGUF files exist — Ollama 🔒 local pulls them straight from Hugging Face 🌐 global, no manual conversion. The model then runs fully locally.
bash · download BgGPT 3.0# 27B, Q4_K_M (~17 GB) — for GB10 or a 24 GB card ollama pull hf.co/INSAIT-Institute/BgGPT-Gemma-3-27B-IT-GGUF:Q4_K_M # Smaller machine: BgGPT-Gemma-3-12B-IT-GGUF or BgGPT-Gemma-3-4B-IT-GGUFPython · bggpt_ab.py — A/B testimport ollama # Tasks stay in Bulgarian — that is what we are testing TASKS = { "Terminology": "Обясни разликата между 'социален работник' и 'социален педагог'.", "Invoice": "Какви реквизити трябва да има една фактура според ЗДДС?", "Formal letter": "Напиши кратко официално писмо до общината за достъп до обществена информация.", } MODELS = { "BgGPT 3.0 27B": "hf.co/INSAIT-Institute/BgGPT-Gemma-3-27B-IT-GGUF:Q4_K_M", "Llama 3.1 8B": "llama3.1:8b", } for name, prompt in TASKS.items(): print(f"\n{'='*60}\n📋 {name}: {prompt}") for label, mid in MODELS.items(): try: r = ollama.generate(model=mid, prompt=prompt, options={"temperature": 0.1}) print(f"\n[{label}]\n{r['response'][:400]}…") except Exception as e: print(f"\n[{label}] unavailable: {e}")👤Who judgesThe comparison is not a number but a review by someone who knows the subject: grammar, terminology, factual accuracy. Check the facts in the answers (invoice requirements, deadlines) against the primary source — both models can be wrong.
04Check
Practical challenge — before Block 1:
- A (easy): run a model and measure tok/s with
benchmark.py. Write the number down — you'll compare it after the optimisations in Block 1. - B (medium): take a real document from your field (a regulation, a contract, technical documentation), run the mini RAG from step 5 on it and ask 5 questions. How many answers are correct?
- C (advanced): measure the tokens of your typical prompt and calculate for 1 year: a cloud API at the current price vs a local model. When does the machine pay for itself?
1. Your RAG assistant answers 8 of 10 questions correctly, but for "What is the electricity price?" it confidently gives a wrong price. The most likely reason?
2. A client wants a 70B model on a machine with 128 GB of unified memory. Which is true?
3. How do you measure writing speed correctly?
4. Does Ollama's JSON mode guarantee 100% valid and correct output?
05What's next
That completes Block 0: you have a local environment, measured numbers, a RAG that knows its limits, and three ways to call the model from code.
06Sources
- Ollama library — models, tags and sizes.
- Ollama: NVIDIA DGX Spark performance — the measurements in step 4.
- Ollama API · OpenAI compatibility · structured outputs · Modelfile.
- ollama-python — the official Python library.
- LangChain + Ollama · Chroma.
- BgGPT-Gemma-3-27B-IT-GGUF · BgGPT 3.0 announcement — INSAIT.
- Hugging Face: models in Ollama — pulling with
hf.co/…. - PyTorch: local install · PEP 668 — why a virtual environment.
- NVIDIA DGX Spark documentation · vLLM — for Block 1.