The KAGAMI mark КАГАМИ
kagami.bg/en/academy/blokove/ · series index · machine-readable viewUPDATED 2026-10-03 · 22 lessons
IDENTITY
series
Blocks 0–10 · programme
publisher
KAGAMI Ltd. (КАГАМИ ЕООД), Varna
access
free, no registration
languages
bg + en
LESSONS
AGENT INSTRUCTIONS

Each lesson carries its own dated trust labels (VERIFIED / UPDATED / TESTED). Quote the lesson page, not this index.

Blocks 0–10 · programme

The block programme: from language models to production, security and go-to-market.

22 lessonsfreeBG and EN
VERIFIED · 01.10.2026UPDATED · 01.10.2026

What is an LLM and how it thinks in tokens

Transformers and self-attention without heavy maths, tokens and BPE, why Bulgarian costs more tokens, the context window and embeddings for RAG.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

The AI stack: from the GPU to the application

The six layers of an AI system, hardware and quantization, BgGPT 3.0 for Bulgarian, and three questions before any AI project — RAG, agent or fine-tuning.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

Lab: your first local LLM, live

Set up a local AI environment, run a model with Ollama, measure speed honestly, build a mini RAG that says “I don't know”, and test BgGPT 3.0 in Bulgarian.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

vLLM: a local inference server for many users

When Ollama is enough and when you need vLLM: PagedAttention, continuous batching, install on x86 and ARM64 (GB10), OpenAI-compatible server, AWQ.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

Docker Compose production stack for AI

vLLM, Nginx, Redis cache, Prometheus and Grafana in one docker-compose.yml: health checks, secrets out of git, rate limiting, alerts and an ARM64 Ollama variant.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

Lab: a full AI stack in one Compose file

Lab: Ollama, Qdrant, Redis, PostgreSQL, n8n, Prometheus, Grafana and Nginx with TLS in one Compose file — health checks, load testing and recovery.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

Prompt engineering: from templates to chain-of-thought

Zero-shot, few-shot, chain-of-thought and tree-of-thought; a 6-part system prompt, adversarial tests, token budgets and prefix caching in vLLM.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

Structured outputs and prompt evaluation

JSON schema, Pydantic and Instructor for reliable LLM output; A/B tests, LLM-as-a-judge, RAGAS and a regression gate to improve prompts with data.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

AI agents with LangChain and LangGraph

The ReAct loop, custom tools, create_agent and StateGraph in LangGraph 1.x: typed state, conditional edges and human approval with interrupt().

VERIFIED · 01.10.2026UPDATED · 01.10.2026

CrewAI and n8n: teams of AI agents

Multi-agent teams with CrewAI 1.x on local Ollama: roles, tasks, sequential and hierarchical process, cloud-free memory and async calls from n8n.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

Human approval in production: LangGraph + FastAPI

An agent that pauses for a human: interrupt() in LangGraph 1.x, a FastAPI server with /invoke and /resume, pauses kept in PostgreSQL, approval via n8n.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

Advanced RAG: chunking, hybrid search, reranking

Structure-aware chunking, dense + BM25 hybrid search fused with RRF in Qdrant, metadata filters and reranking with a multilingual cross-encoder.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

GraphRAG, production RAG and evaluation with RAGAS

A knowledge graph from legal text, hybrid graph + vector retrieval, a RAG API with a Redis cache, and nightly RAGAS 0.4 evaluation on a local model.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

Fine-tuning with LoRA and QLoRA

When fine-tuning pays off, how LoRA and QLoRA work, how much memory they need, running Unsloth on BgGPT 3 and preparing clean Bulgarian training data.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

From training to Ollama: a fine-tuned model in use

Train with SFTTrainer, evaluate honestly in Bulgarian, merge LoRA, export GGUF with llama.cpp and serve it in Ollama with the right chat template.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

From AI tools to operational systems

How AI becomes a production unit: input, processing, output and control, five system layers, a value calculation and three cases with human approval.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

Multimodal AI: images, voice and documents

Local vision models in Ollama, Whisper and WhisperX for Bulgarian, text-to-speech, and extraction from PDF, Word and scanned documents — with checked code.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

Multimodal pipeline: image, audio and document in one answer

Image, audio and document in one answer: modality detection, Whisper, vision models, document parsers, synthesis and a FastAPI service on Ollama.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

AI security and guardrails

Five layers of defense for LLM apps: prompt injection, NeMo Guardrails, personal data, a GDPR audit log, the OWASP LLM Top 10 2026 and AI Act dates.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

Observability for AI in production

Trace every call in LangSmith or Langfuse, metrics in Prometheus, Grafana dashboards, Slack alerts and a runbook for production LLM systems.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

Media pool: an AI content pipeline

A draft from a local AI, checks, human approval, JSON-LD, a sitemap and banner rotation with n8n — and where Google counts automation as spam.

VERIFIED · 01.10.2026UPDATED · 01.10.2026

How to sell an AI system: pricing and ROI

Positioning, value-based pricing, an ROI calculator, pilot playbooks and honest answers to objections when selling local AI systems to clients.