The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewVERIFIED 2026-10-01 · UPDATED 2026-10-01
IDENTITY
module
01-00b · The AI stack, from the GPU to the application
series
KAGAMI Academy · Blocks · Block 0 — AI fundamentals · part 2 of 3
level
Beginner → Intermediate
duration
~3 h
trust_label
VERIFIED 2026-10-01 (sources, BgGPT 3.0 facts, commands vs official docs) · UPDATED 2026-10-01 (BgGPT 3.0, CUDA 13 / ROCm 7, TGI maintenance, NVLink, DGX Spark row) · hardware speeds and prices NOT re-measured
language
human view: en · bulgarian edition: /academy/blokove/moduli/01-00b_Блок_0_Част_2_AI_Стек.html
prev
01-00a · LLMs and tokenisation
next
01-00c · Lab: full environment + first live LLM
PURPOSE

Give a mental model of every AI system as six dependent layers (hardware → driver/CUDA → ML runtime → inference server → orchestration → application), so faults are diagnosed by layer; size hardware via quantization; know the Bulgarian open model BgGPT 3.0; decide between prompt, RAG, agent and fine-tuning — and when not to use AI at all.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

01-00c · Lab: full environment + first live LLM · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
ai-stackgpucudaquantizationollamavllmbggptragagentsfine-tuning
VERIFIED · 01.10.2026 UPDATED · 01.10.2026

The AI stack: from the GPU to the application

Every AI system — from a local chatbot to a large document-search system — passes through the same six layers. Whoever understands them does not "install and pray" but diagnoses, optimises and scales. In this lesson: the layers, hardware and quantization, the Bulgarian model BgGPT 3.0, and the three questions that separate the architect from the technician.

⏱ ~3 hours Beginner → Intermediate Block 0 · AI fundamentals hardware · quantization · decisions
Ollama🔒 local vLLM · llama.cpp🔒 local BgGPT 3.0🔒 local Cloud APIs (GPT, Gemini…)🌐 global model

01What you will learn

02Before you start

03Steps

  1. Six layers — a bottom-up pyramid

    Each layer depends on the one below it. That is why hardware experience is a big advantage in AI deployments: most "AI problems" in a running system are really infrastructure problems.

    LayerWhat it doesExamples
    6 · ApplicationThe end product — chatbot, document search, agent, automated workflow. This is where the business value is.Open WebUI, AnythingLLM, n8n, custom FastAPI, Flowise
    5 · OrchestrationThe logic: how the AI decides, calls tools, remembers the conversation; a vector database for documents.LangChain, LangGraph, CrewAI, LlamaIndex, Haystack · Qdrant / Chroma · RAGAS for evaluation
    4 · Inference serverLoads the model and answers requests. Tokens/s, latency and capacity are decided here.Ollama, vLLM, llama.cpp, LM Studio, SGLang, LiteLLM (gateway) · formats GGUF, AWQ/GPTQ, SafeTensors
    3 · ML framework / runtimeRuns the tensor operations on the GPU.PyTorch 2.x, ONNX Runtime, TensorRT, ROCm (AMD), Metal (Apple) · Flash Attention, Triton
    2 · Driver / CUDAThe bridge between the operating system and the GPU. The version must match the one PyTorch was built for.CUDA 13.x, cuDNN 9.x, NCCL, ROCm 7.x · nvidia-smi, nvcc --version
    1 · HardwareThe physical base: GPU memory (VRAM), RAM, NVMe, PCIe. Wrong hardware makes good software slow.NVIDIA RTX / RTX PRO / H-series, AMD Radeon / Instinct, Apple M-series, NVIDIA Jetson, NVIDIA GB10 (DGX Spark)
    💡
    Quick diagnosis by layer
    Tokens/s are low while the GPU is busy → memory bandwidth is the bottleneck. The GPU sits idle → the CPU can't keep up with data preparation. The model won't start at all → check bottom-up: driver, CUDA, PyTorch.
    🔄
    UPDATED · 01.10.2026 — versions and tools
    The old text said CUDA 12.x and ROCm 6.x; the current lines are CUDA 13.x and ROCm 7.x. Hugging Face TGI has been in maintenance mode since December 2025 — Hugging Face itself recommends vLLM, SGLang or llama.cpp — so TGI was removed from the examples.
  2. Hardware: memory, speed, power

    Three numbers decide almost everything: how much memory the GPU has (how big a model fits), how fast that memory is (tokens per second) and how much power it draws (the monthly bill).

    PlatformMemoryMax model (Q4)PowerGood for
    RTX 4060 Ti 16 GB16 GB13B165 WStarter, small office
    RTX 409024 GB34B450 WWorkstation, SME
    RTX 509032 GB34B with more context575 WHigh-end workstation
    RTX 6000 Ada48 GB70B300 WOffice server
    NVIDIA GB10 (DGX Spark and similar)128 GB unified70B+ (up to ~200B per NVIDIA)~240 WQuiet local AI station
    Apple M4 Maxup to 128 GB unified70BlowQuiet office, privacy
    NVIDIA H100 SXM80 GB70B+ at full precision700 WEnterprise cluster
    AMD MI300X192 GBvery large models750 WH100 alternative
    NVIDIA Jetson AGX Orin64 GB13Bup to 60 WEdge, industrial
    2× RTX 40902×24 GB70B (split)~900 WMulti-GPU starter
    ⚠️
    Indicative — not measured by us
    Memory and power are manufacturer figures. The speeds (tokens/s) and prices from the old lesson (Q1 2026) were not re-measured, so they are left out — check a current EUR price with a supplier before quoting. The RTX 4090 is no longer produced and is hard to find.
    ⛔
    UPDATED · 01.10.2026 — correction: the RTX 4090 has no NVLink
    The old lesson said "2× RTX 4090 NVLink". RTX 40-series cards have no NVLink. Two cards work, but the model is split between them over PCIe — slower than one card with enough memory. "RTX A6000 Ada" was also corrected to "RTX 6000 Ada" (two different products), and an NVIDIA GB10 row was added.
    🎯
    Starter configuration for an SME
    One card with 24–32 GB (RTX 4090 or 5090), 64 GB RAM, 2 TB NVMe, Ubuntu 24.04 — covers most local projects: models up to 34B in Q4 and a few concurrent users. If you want silence and nothing leaving the office — a unified-memory machine (Apple M-series or NVIDIA GB10).
    ⚡
    PCIe — the hidden bottleneck
    A large model (70B+) loads along NVMe → RAM → GPU. PCIe 4.0 x16 ≈ 32 GB/s, PCIe 3.0 x4 ≈ 4 GB/s — 8 times slower. Check with lspci -vvv | grep LnkSta. For models over 40 GB — NVMe in a PCIe 4.0 x4 slot, not a SATA SSD.
  3. Quantization — memory versus quality

    Quantization lowers the precision of the model weights (from 16 bits down to 4), so memory drops about 4× with a small quality loss. That is how a 24 GB card runs a 34B model instead of only a 7B.

    TypeBits7B13B34B70BFormat
    fp16 (baseline)1614 GB26 GB68 GB140 GBSafeTensors
    Q8_08~7.2 GB~13 GB~34 GB~72 GBGGUF
    Q4_K_M ★ balanced4~4.1 GB~7.9 GB~19 GB~41 GBGGUF
    Q3_K_M3~3.3 GB~6.3 GB~15 GB~33 GBGGUF
    AWQ4~4 GB~7.5 GB~18 GB~39 GBAWQ (vLLM)

    This is the size of the weights alone. Add headroom for the context (KV cache) — the longer the conversation, the more. Q8 is practically lossless, Q4_K_M usually loses little, at Q3 the loss becomes noticeable.

    bash · how much GPU memory is in use
    # Current state
    nvidia-smi
    
    # Live utilisation, refreshed every second
    nvidia-smi dmon -s u -d 1
    
    # Name, memory, temperature, utilisation — every 2 seconds
    nvidia-smi --query-gpu=name,memory.total,memory.used,temperature.gpu,utilization.gpu \
               --format=csv -l 2
    
    # Does PyTorch see the GPU?
    python3 -c "
    import torch
    print(f'CUDA: {torch.cuda.is_available()}')
    print(f'GPU: {torch.cuda.get_device_name(0)}')
    print(f'VRAM: {torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB')
    "
    
    # Which models Ollama has loaded and how much memory they hold
    ollama ps
  4. BgGPT 3.0 — the open model for Bulgarian

    BgGPT 3.0 comes from INSAIT (Sofia) and was announced on 24 March 2026. It is three open models — 4B, 12B and 27B, built on Google's Gemma 3: they understand text and images, hold a long context (~131k tokens) and follow instructions in Bulgarian and English markedly better than version 2.0. The licence is the Gemma terms.

    Why it matters: Cyrillic, legal and administrative terminology, tax and health-insurance documents — exactly where models trained mostly on English slip. And because it runs locally 🔒 local, the data never leaves the machine.

    🔄
    UPDATED · 01.10.2026 — what was corrected about BgGPT
    The old lesson's ollama pull bggpt3 does not exist; the official route is the GGUF file on Hugging Face (below). Voice (speech recognition and synthesis) and web search are features of the bggpt.ai chat app, not of the open models. The "GDPR compliant" and "AI Act ready" labels were removed — compliance depends on how you deploy the model, not on the model itself.
    bash · BgGPT 3.0 (12B, Q4_K_M ≈ 7.3 GB) with Ollama
    # Downloads and runs INSAIT's official GGUF from Hugging Face
    ollama run hf.co/INSAIT-Institute/BgGPT-Gemma-3-12B-IT-GGUF:Q4_K_M
    
    # Ollama's native API (localhost = your own machine)
    curl http://localhost:11434/api/generate -d '{
      "model": "hf.co/INSAIT-Institute/BgGPT-Gemma-3-12B-IT-GGUF:Q4_K_M",
      "prompt": "Обобщи в три изречения: [текст]",
      "stream": false
    }'
    
    # The OpenAI-compatible endpoint is separate: http://localhost:11434/v1/chat/completions
    Task in BulgarianCloud model 🌐Generic small local model 🔒BgGPT 3.0 🔒
    Legal and administrative textgood, with guidancefrequent terminology errorsa strength
    Confidential dataleaves the organisationstays localstays local
    Cost of usepay per tokenpower and hardware onlypower and hardware only
    ⚠️
    Unverified comparison
    This table is a qualitative assessment carried over from the old lesson, not a test we measured. Before choosing a model for a client — trial it on that client's documents.
  5. Three questions before any project

    The technician knows how to install Ollama. The architect knows when Ollama is the wrong choice. Before any code — three questions:

    1 · Is AI needed at all?2 · Which approach?3 · Local or cloud?
    Is the problem described with examples?RAG — there are organisational documentsSensitive data → local
    Is there enough data?Agent — actions and tools are neededRegulated sector (health, law) → local
    Wouldn't a plain algorithm be more reliable?Fine-tuning — a specific style or domainPrototype, exploration → cloud API
    Does the cost pay off?Prompt — clear, simple classificationUnpredictable peaks → cloud with a fallback
    Is the client ready for AI mistakes?Combination — for complex systemsNo internet → local only
  6. RAG, agent, fine-tuning — when to use which

    The most common mistake: fine-tuning for a problem RAG solves, or an agent where a good prompt is enough.

    ✍️ Prompt

    The task is clear and the base model copes. The fastest solution — every task passes here first.

    📚 RAG

    You have documents, rules, procedures and want answers only from them, with citations. Knowledge updates without retraining.

    🤖 Agent

    Actions are needed: search, email, a request to a system, filling a form. A multi-step task.

    🎛️ Fine-tuning

    A specific style, tone or terminology a prompt can't reach. Needs hundreds of good examples (rule of thumb: 500+).

    🔗 RAG + agent

    The agent calls RAG when it needs knowledge and acts otherwise. The most common real-world architecture.

    ❌ Not AI

    Simple create/read, deterministic logic (tax calculations!), critical safety without human oversight, under ~50 examples. Sometimes a regular expression is the better answer.

    🔑
    The frame in action — typical cases
    Social service: RAG (case files) + agent (fills in forms) + BgGPT 3.0 locally.
    ESG consultant: RAG (regulations and measurements) + agent (generates a PDF report).
    Accounting: agent (extracts data from PDF invoices) + prompt (ledger coding).
    Medical practice: RAG (protocols) + fine-tuning (Bulgarian medical terminology) + guardrails.
  7. Practice: check the stack layer by layer

    Before any deployment, confirm that every layer works. Save the script as check_stack.sh, make it executable (chmod +x check_stack.sh) and run it. We will use it again in the lab (Part 3).

    bash · check_stack.sh
    #!/bin/bash
    # AI Stack Diagnostic · KAGAMI Academy
    # Checks the layers bottom-up
    MODEL="${MODEL:-llama3.2}"   # model for the end-to-end test; change to one you have
    
    status() { [ "$2" == "OK" ] && echo "✅ $1" || echo "❌ $1 — $2"; }
    
    echo "== LAYER 1: HARDWARE =="
    if nvidia-smi &>/dev/null; then
      GPU=$(nvidia-smi --query-gpu=name --format=csv,noheader | head -1)
      VRAM=$(nvidia-smi --query-gpu=memory.total --format=csv,noheader | head -1)
      status "GPU: $GPU | VRAM: $VRAM" "OK"
    else
      status "NVIDIA GPU" "not found — running on CPU"
    fi
    
    echo "== LAYER 2: CUDA =="
    CUDA_VER=$(nvcc --version 2>/dev/null | grep "release" | awk '{print $6}')
    [ -n "$CUDA_VER" ] && status "CUDA: $CUDA_VER" "OK" || status "CUDA toolkit" "not installed"
    
    echo "== LAYER 3: RUNTIME =="
    python3 -c "import torch; print(f'PyTorch {torch.__version__} | CUDA: {torch.cuda.is_available()}')" 2>/dev/null \
      && status "PyTorch" "OK" || status "PyTorch" "not installed"
    
    echo "== LAYER 4: INFERENCE SERVER =="
    if curl -s http://localhost:11434/api/tags &>/dev/null; then
      N=$(ollama list 2>/dev/null | tail -n +2 | wc -l)
      status "Ollama running | models: $N" "OK"
    else
      status "Ollama" "not running — start it: ollama serve"
    fi
    
    echo "== LAYER 5: ORCHESTRATION =="
    python3 -c "import langchain; print(f'LangChain {langchain.__version__}')" 2>/dev/null \
      && status "LangChain" "OK" || status "LangChain" "pip install langchain"
    
    echo "== END-TO-END TEST =="
    R=$(curl -s http://localhost:11434/api/generate \
      -d "{\"model\":\"$MODEL\",\"prompt\":\"Answer in one sentence: what is AI?\",\"stream\":false}" \
      | python3 -c "import sys,json; print(json.load(sys.stdin).get('response',''))" 2>/dev/null)
    [ -n "$R" ] && status "Answer: $R" "OK" || status "End-to-end test" "failed — check Ollama and the model"
    👤
    Why bottom-up
    If layer 2 is red, there is no point chasing an error in layer 5. The script shows the first broken layer — that's where you start. Change the test model with MODEL=<your-model> ./check_stack.sh.

04Check

1. A medical practice wants AI that fills in documents from the doctor's dictation. Which stack is right?

2. A client wants an email to go out automatically when an invoice arrives. Which approach fits best?

3. A 24 GB card. What is the largest Q4_K_M model you can run comfortably?

4. torch.cuda.is_available() returns False while nvidia-smi sees the card. Where do you look?

05What's next

06Sources

  1. BgGPT 3.0 — announcement (INSAIT, 24.03.2026) — sizes, Gemma 3, images, context.
  2. BgGPT-Gemma-3-12B-IT-GGUF — files, sizes and the Ollama command.
  3. INSAIT on Hugging Face — all BgGPT versions.
  4. Hugging Face TGI — the maintenance-mode notice.
  5. NVIDIA CUDA Toolkit — installer; always match it to your PyTorch version.
  6. PyTorch — local install — the exact command for your CUDA.
  7. vLLM — quickstart — OpenAI-compatible server.
  8. llama.cpp — the GGUF engine (repository moved to ggml-org).
  9. Ollama — model library · OpenAI-compatible endpoint.
  10. Open WebUI — a web interface for local models.
  11. NVIDIA DGX Spark (GB10) — 128 GB unified memory.
  12. TechPowerUp GPU Database — GPU memory, bandwidth and power.