Blocks 0–10 · programme
The block programme: from language models to production, security and go-to-market.
What is an LLM and how it thinks in tokens
Transformers and self-attention without heavy maths, tokens and BPE, why Bulgarian costs more tokens, the context window and embeddings for RAG.
VERIFIED · 01.10.2026UPDATED · 01.10.2026The AI stack: from the GPU to the application
The six layers of an AI system, hardware and quantization, BgGPT 3.0 for Bulgarian, and three questions before any AI project — RAG, agent or fine-tuning.
VERIFIED · 01.10.2026UPDATED · 01.10.2026Lab: your first local LLM, live
Set up a local AI environment, run a model with Ollama, measure speed honestly, build a mini RAG that says “I don't know”, and test BgGPT 3.0 in Bulgarian.
VERIFIED · 01.10.2026UPDATED · 01.10.2026vLLM: a local inference server for many users
When Ollama is enough and when you need vLLM: PagedAttention, continuous batching, install on x86 and ARM64 (GB10), OpenAI-compatible server, AWQ.
VERIFIED · 01.10.2026UPDATED · 01.10.2026Docker Compose production stack for AI
vLLM, Nginx, Redis cache, Prometheus and Grafana in one docker-compose.yml: health checks, secrets out of git, rate limiting, alerts and an ARM64 Ollama variant.
VERIFIED · 01.10.2026UPDATED · 01.10.2026Lab: a full AI stack in one Compose file
Lab: Ollama, Qdrant, Redis, PostgreSQL, n8n, Prometheus, Grafana and Nginx with TLS in one Compose file — health checks, load testing and recovery.
VERIFIED · 01.10.2026UPDATED · 01.10.2026Prompt engineering: from templates to chain-of-thought
Zero-shot, few-shot, chain-of-thought and tree-of-thought; a 6-part system prompt, adversarial tests, token budgets and prefix caching in vLLM.
VERIFIED · 01.10.2026UPDATED · 01.10.2026Structured outputs and prompt evaluation
JSON schema, Pydantic and Instructor for reliable LLM output; A/B tests, LLM-as-a-judge, RAGAS and a regression gate to improve prompts with data.
VERIFIED · 01.10.2026UPDATED · 01.10.2026AI agents with LangChain and LangGraph
The ReAct loop, custom tools, create_agent and StateGraph in LangGraph 1.x: typed state, conditional edges and human approval with interrupt().
VERIFIED · 01.10.2026UPDATED · 01.10.2026CrewAI and n8n: teams of AI agents
Multi-agent teams with CrewAI 1.x on local Ollama: roles, tasks, sequential and hierarchical process, cloud-free memory and async calls from n8n.
VERIFIED · 01.10.2026UPDATED · 01.10.2026Human approval in production: LangGraph + FastAPI
An agent that pauses for a human: interrupt() in LangGraph 1.x, a FastAPI server with /invoke and /resume, pauses kept in PostgreSQL, approval via n8n.
VERIFIED · 01.10.2026UPDATED · 01.10.2026Advanced RAG: chunking, hybrid search, reranking
Structure-aware chunking, dense + BM25 hybrid search fused with RRF in Qdrant, metadata filters and reranking with a multilingual cross-encoder.
VERIFIED · 01.10.2026UPDATED · 01.10.2026GraphRAG, production RAG and evaluation with RAGAS
A knowledge graph from legal text, hybrid graph + vector retrieval, a RAG API with a Redis cache, and nightly RAGAS 0.4 evaluation on a local model.
VERIFIED · 01.10.2026UPDATED · 01.10.2026Fine-tuning with LoRA and QLoRA
When fine-tuning pays off, how LoRA and QLoRA work, how much memory they need, running Unsloth on BgGPT 3 and preparing clean Bulgarian training data.
VERIFIED · 01.10.2026UPDATED · 01.10.2026From training to Ollama: a fine-tuned model in use
Train with SFTTrainer, evaluate honestly in Bulgarian, merge LoRA, export GGUF with llama.cpp and serve it in Ollama with the right chat template.
VERIFIED · 01.10.2026UPDATED · 01.10.2026From AI tools to operational systems
How AI becomes a production unit: input, processing, output and control, five system layers, a value calculation and three cases with human approval.
VERIFIED · 01.10.2026UPDATED · 01.10.2026Multimodal AI: images, voice and documents
Local vision models in Ollama, Whisper and WhisperX for Bulgarian, text-to-speech, and extraction from PDF, Word and scanned documents — with checked code.
VERIFIED · 01.10.2026UPDATED · 01.10.2026Multimodal pipeline: image, audio and document in one answer
Image, audio and document in one answer: modality detection, Whisper, vision models, document parsers, synthesis and a FastAPI service on Ollama.
VERIFIED · 01.10.2026UPDATED · 01.10.2026AI security and guardrails
Five layers of defense for LLM apps: prompt injection, NeMo Guardrails, personal data, a GDPR audit log, the OWASP LLM Top 10 2026 and AI Act dates.
VERIFIED · 01.10.2026UPDATED · 01.10.2026Observability for AI in production
Trace every call in LangSmith or Langfuse, metrics in Prometheus, Grafana dashboards, Slack alerts and a runbook for production LLM systems.
VERIFIED · 01.10.2026UPDATED · 01.10.2026Media pool: an AI content pipeline
A draft from a local AI, checks, human approval, JSON-LD, a sitemap and banner rotation with n8n — and where Google counts automation as spam.
VERIFIED · 01.10.2026UPDATED · 01.10.2026How to sell an AI system: pricing and ROI
Positioning, value-based pricing, an ROI calculator, pilot playbooks and honest answers to objections when selling local AI systems to clients.