The AI stack: from the GPU to the application
Every AI system — from a local chatbot to a large document-search system — passes through the same six layers. Whoever understands them does not "install and pray" but diagnoses, optimises and scales. In this lesson: the layers, hardware and quantization, the Bulgarian model BgGPT 3.0, and the three questions that separate the architect from the technician.
01What you will learn
- The six layers of the AI stack and why a problem below shows up everywhere above.
- How to judge hardware: GPU memory, bandwidth, power.
- What quantization is and how much memory a 7, 13, 34 and 70 billion-parameter model needs.
- What BgGPT 3.0 is and how to run it locally.
- Three questions before any project, and when to use RAG, an agent, fine-tuning — or no AI at all.
- A script that checks the stack layer by layer.
02Before you start
- You have done Part 1 of Block 0 — what a language model and a token are.
- A Linux machine (or WSL2) with a terminal. An NVIDIA GPU helps but is not required — the script will tell you if you are CPU-only.
- Optional: Ollama 🔒 local installed with one small model — for the end-to-end test.
03Steps
-
Six layers — a bottom-up pyramid
Each layer depends on the one below it. That is why hardware experience is a big advantage in AI deployments: most "AI problems" in a running system are really infrastructure problems.
Layer What it does Examples 6 · Application The end product — chatbot, document search, agent, automated workflow. This is where the business value is. Open WebUI, AnythingLLM, n8n, custom FastAPI, Flowise 5 · Orchestration The logic: how the AI decides, calls tools, remembers the conversation; a vector database for documents. LangChain, LangGraph, CrewAI, LlamaIndex, Haystack · Qdrant / Chroma · RAGAS for evaluation 4 · Inference server Loads the model and answers requests. Tokens/s, latency and capacity are decided here. Ollama, vLLM, llama.cpp, LM Studio, SGLang, LiteLLM (gateway) · formats GGUF, AWQ/GPTQ, SafeTensors 3 · ML framework / runtime Runs the tensor operations on the GPU. PyTorch 2.x, ONNX Runtime, TensorRT, ROCm (AMD), Metal (Apple) · Flash Attention, Triton 2 · Driver / CUDA The bridge between the operating system and the GPU. The version must match the one PyTorch was built for. CUDA 13.x, cuDNN 9.x, NCCL, ROCm 7.x · nvidia-smi,nvcc --version1 · Hardware The physical base: GPU memory (VRAM), RAM, NVMe, PCIe. Wrong hardware makes good software slow. NVIDIA RTX / RTX PRO / H-series, AMD Radeon / Instinct, Apple M-series, NVIDIA Jetson, NVIDIA GB10 (DGX Spark) 💡Quick diagnosis by layerTokens/s are low while the GPU is busy → memory bandwidth is the bottleneck. The GPU sits idle → the CPU can't keep up with data preparation. The model won't start at all → check bottom-up: driver, CUDA, PyTorch.🔄UPDATED · 01.10.2026 — versions and toolsThe old text said CUDA 12.x and ROCm 6.x; the current lines are CUDA 13.x and ROCm 7.x. Hugging Face TGI has been in maintenance mode since December 2025 — Hugging Face itself recommends vLLM, SGLang or llama.cpp — so TGI was removed from the examples. -
Hardware: memory, speed, power
Three numbers decide almost everything: how much memory the GPU has (how big a model fits), how fast that memory is (tokens per second) and how much power it draws (the monthly bill).
Platform Memory Max model (Q4) Power Good for RTX 4060 Ti 16 GB 16 GB 13B 165 W Starter, small office RTX 4090 24 GB 34B 450 W Workstation, SME RTX 5090 32 GB 34B with more context 575 W High-end workstation RTX 6000 Ada 48 GB 70B 300 W Office server NVIDIA GB10 (DGX Spark and similar) 128 GB unified 70B+ (up to ~200B per NVIDIA) ~240 W Quiet local AI station Apple M4 Max up to 128 GB unified 70B low Quiet office, privacy NVIDIA H100 SXM 80 GB 70B+ at full precision 700 W Enterprise cluster AMD MI300X 192 GB very large models 750 W H100 alternative NVIDIA Jetson AGX Orin 64 GB 13B up to 60 W Edge, industrial 2× RTX 4090 2×24 GB 70B (split) ~900 W Multi-GPU starter ⚠️Indicative — not measured by usMemory and power are manufacturer figures. The speeds (tokens/s) and prices from the old lesson (Q1 2026) were not re-measured, so they are left out — check a current EUR price with a supplier before quoting. The RTX 4090 is no longer produced and is hard to find.⛔UPDATED · 01.10.2026 — correction: the RTX 4090 has no NVLinkThe old lesson said "2× RTX 4090 NVLink". RTX 40-series cards have no NVLink. Two cards work, but the model is split between them over PCIe — slower than one card with enough memory. "RTX A6000 Ada" was also corrected to "RTX 6000 Ada" (two different products), and an NVIDIA GB10 row was added.🎯Starter configuration for an SMEOne card with 24–32 GB (RTX 4090 or 5090), 64 GB RAM, 2 TB NVMe, Ubuntu 24.04 — covers most local projects: models up to 34B in Q4 and a few concurrent users. If you want silence and nothing leaving the office — a unified-memory machine (Apple M-series or NVIDIA GB10).⚡PCIe — the hidden bottleneckA large model (70B+) loads along NVMe → RAM → GPU. PCIe 4.0 x16 ≈ 32 GB/s, PCIe 3.0 x4 ≈ 4 GB/s — 8 times slower. Check withlspci -vvv | grep LnkSta. For models over 40 GB — NVMe in a PCIe 4.0 x4 slot, not a SATA SSD. -
Quantization — memory versus quality
Quantization lowers the precision of the model weights (from 16 bits down to 4), so memory drops about 4× with a small quality loss. That is how a 24 GB card runs a 34B model instead of only a 7B.
Type Bits 7B 13B 34B 70B Format fp16 (baseline) 16 14 GB 26 GB 68 GB 140 GB SafeTensors Q8_0 8 ~7.2 GB ~13 GB ~34 GB ~72 GB GGUF Q4_K_M ★ balanced 4 ~4.1 GB ~7.9 GB ~19 GB ~41 GB GGUF Q3_K_M 3 ~3.3 GB ~6.3 GB ~15 GB ~33 GB GGUF AWQ 4 ~4 GB ~7.5 GB ~18 GB ~39 GB AWQ (vLLM) This is the size of the weights alone. Add headroom for the context (KV cache) — the longer the conversation, the more. Q8 is practically lossless, Q4_K_M usually loses little, at Q3 the loss becomes noticeable.
bash · how much GPU memory is in use# Current state nvidia-smi # Live utilisation, refreshed every second nvidia-smi dmon -s u -d 1 # Name, memory, temperature, utilisation — every 2 seconds nvidia-smi --query-gpu=name,memory.total,memory.used,temperature.gpu,utilization.gpu \ --format=csv -l 2 # Does PyTorch see the GPU? python3 -c " import torch print(f'CUDA: {torch.cuda.is_available()}') print(f'GPU: {torch.cuda.get_device_name(0)}') print(f'VRAM: {torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB') " # Which models Ollama has loaded and how much memory they hold ollama ps -
BgGPT 3.0 — the open model for Bulgarian
BgGPT 3.0 comes from INSAIT (Sofia) and was announced on 24 March 2026. It is three open models — 4B, 12B and 27B, built on Google's Gemma 3: they understand text and images, hold a long context (~131k tokens) and follow instructions in Bulgarian and English markedly better than version 2.0. The licence is the Gemma terms.
Why it matters: Cyrillic, legal and administrative terminology, tax and health-insurance documents — exactly where models trained mostly on English slip. And because it runs locally 🔒 local, the data never leaves the machine.
🔄UPDATED · 01.10.2026 — what was corrected about BgGPTThe old lesson'sollama pull bggpt3does not exist; the official route is the GGUF file on Hugging Face (below). Voice (speech recognition and synthesis) and web search are features of the bggpt.ai chat app, not of the open models. The "GDPR compliant" and "AI Act ready" labels were removed — compliance depends on how you deploy the model, not on the model itself.bash · BgGPT 3.0 (12B, Q4_K_M ≈ 7.3 GB) with Ollama# Downloads and runs INSAIT's official GGUF from Hugging Face ollama run hf.co/INSAIT-Institute/BgGPT-Gemma-3-12B-IT-GGUF:Q4_K_M # Ollama's native API (localhost = your own machine) curl http://localhost:11434/api/generate -d '{ "model": "hf.co/INSAIT-Institute/BgGPT-Gemma-3-12B-IT-GGUF:Q4_K_M", "prompt": "Обобщи в три изречения: [текст]", "stream": false }' # The OpenAI-compatible endpoint is separate: http://localhost:11434/v1/chat/completionsTask in Bulgarian Cloud model 🌐 Generic small local model 🔒 BgGPT 3.0 🔒 Legal and administrative text good, with guidance frequent terminology errors a strength Confidential data leaves the organisation stays local stays local Cost of use pay per token power and hardware only power and hardware only ⚠️Unverified comparisonThis table is a qualitative assessment carried over from the old lesson, not a test we measured. Before choosing a model for a client — trial it on that client's documents. -
Three questions before any project
The technician knows how to install Ollama. The architect knows when Ollama is the wrong choice. Before any code — three questions:
1 · Is AI needed at all? 2 · Which approach? 3 · Local or cloud? Is the problem described with examples? RAG — there are organisational documents Sensitive data → local Is there enough data? Agent — actions and tools are needed Regulated sector (health, law) → local Wouldn't a plain algorithm be more reliable? Fine-tuning — a specific style or domain Prototype, exploration → cloud API Does the cost pay off? Prompt — clear, simple classification Unpredictable peaks → cloud with a fallback Is the client ready for AI mistakes? Combination — for complex systems No internet → local only -
RAG, agent, fine-tuning — when to use which
The most common mistake: fine-tuning for a problem RAG solves, or an agent where a good prompt is enough.
✍️ Prompt
The task is clear and the base model copes. The fastest solution — every task passes here first.
📚 RAG
You have documents, rules, procedures and want answers only from them, with citations. Knowledge updates without retraining.
🤖 Agent
Actions are needed: search, email, a request to a system, filling a form. A multi-step task.
🎛️ Fine-tuning
A specific style, tone or terminology a prompt can't reach. Needs hundreds of good examples (rule of thumb: 500+).
🔗 RAG + agent
The agent calls RAG when it needs knowledge and acts otherwise. The most common real-world architecture.
❌ Not AI
Simple create/read, deterministic logic (tax calculations!), critical safety without human oversight, under ~50 examples. Sometimes a regular expression is the better answer.
🔑The frame in action — typical casesSocial service: RAG (case files) + agent (fills in forms) + BgGPT 3.0 locally.
ESG consultant: RAG (regulations and measurements) + agent (generates a PDF report).
Accounting: agent (extracts data from PDF invoices) + prompt (ledger coding).
Medical practice: RAG (protocols) + fine-tuning (Bulgarian medical terminology) + guardrails. -
Practice: check the stack layer by layer
Before any deployment, confirm that every layer works. Save the script as
check_stack.sh, make it executable (chmod +x check_stack.sh) and run it. We will use it again in the lab (Part 3).bash · check_stack.sh#!/bin/bash # AI Stack Diagnostic · KAGAMI Academy # Checks the layers bottom-up MODEL="${MODEL:-llama3.2}" # model for the end-to-end test; change to one you have status() { [ "$2" == "OK" ] && echo "✅ $1" || echo "❌ $1 — $2"; } echo "== LAYER 1: HARDWARE ==" if nvidia-smi &>/dev/null; then GPU=$(nvidia-smi --query-gpu=name --format=csv,noheader | head -1) VRAM=$(nvidia-smi --query-gpu=memory.total --format=csv,noheader | head -1) status "GPU: $GPU | VRAM: $VRAM" "OK" else status "NVIDIA GPU" "not found — running on CPU" fi echo "== LAYER 2: CUDA ==" CUDA_VER=$(nvcc --version 2>/dev/null | grep "release" | awk '{print $6}') [ -n "$CUDA_VER" ] && status "CUDA: $CUDA_VER" "OK" || status "CUDA toolkit" "not installed" echo "== LAYER 3: RUNTIME ==" python3 -c "import torch; print(f'PyTorch {torch.__version__} | CUDA: {torch.cuda.is_available()}')" 2>/dev/null \ && status "PyTorch" "OK" || status "PyTorch" "not installed" echo "== LAYER 4: INFERENCE SERVER ==" if curl -s http://localhost:11434/api/tags &>/dev/null; then N=$(ollama list 2>/dev/null | tail -n +2 | wc -l) status "Ollama running | models: $N" "OK" else status "Ollama" "not running — start it: ollama serve" fi echo "== LAYER 5: ORCHESTRATION ==" python3 -c "import langchain; print(f'LangChain {langchain.__version__}')" 2>/dev/null \ && status "LangChain" "OK" || status "LangChain" "pip install langchain" echo "== END-TO-END TEST ==" R=$(curl -s http://localhost:11434/api/generate \ -d "{\"model\":\"$MODEL\",\"prompt\":\"Answer in one sentence: what is AI?\",\"stream\":false}" \ | python3 -c "import sys,json; print(json.load(sys.stdin).get('response',''))" 2>/dev/null) [ -n "$R" ] && status "Answer: $R" "OK" || status "End-to-end test" "failed — check Ollama and the model"👤Why bottom-upIf layer 2 is red, there is no point chasing an error in layer 5. The script shows the first broken layer — that's where you start. Change the test model withMODEL=<your-model> ./check_stack.sh.
04Check
1. A medical practice wants AI that fills in documents from the doctor's dictation. Which stack is right?
2. A client wants an email to go out automatically when an invoice arrives. Which approach fits best?
3. A 24 GB card. What is the largest Q4_K_M model you can run comfortably?
4. torch.cuda.is_available() returns False while nvidia-smi sees the card. Where do you look?
05What's next
06Sources
- BgGPT 3.0 — announcement (INSAIT, 24.03.2026) — sizes, Gemma 3, images, context.
- BgGPT-Gemma-3-12B-IT-GGUF — files, sizes and the Ollama command.
- INSAIT on Hugging Face — all BgGPT versions.
- Hugging Face TGI — the maintenance-mode notice.
- NVIDIA CUDA Toolkit — installer; always match it to your PyTorch version.
- PyTorch — local install — the exact command for your CUDA.
- vLLM — quickstart — OpenAI-compatible server.
- llama.cpp — the GGUF engine (repository moved to ggml-org).
- Ollama — model library · OpenAI-compatible endpoint.
- Open WebUI — a web interface for local models.
- NVIDIA DGX Spark (GB10) — 128 GB unified memory.
- TechPowerUp GPU Database — GPU memory, bandwidth and power.