The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewVERIFIED 2026-10-01 · UPDATED 2026-10-01
IDENTITY
module
01-00a · What is an LLM and how it thinks in tokens
series
KAGAMI Academy · Blocks 0–10 · Block 0 — AI Fundamentals · part 1/3 (chapters 0.1–0.2)
level
Beginner
duration
~3–4 h (reading + practice)
prerequisites
basic Python; intuition for vectors
trust_label
VERIFIED 2026-10-01 (token counts measured with tiktoken 0.14; prices, context windows, embedding models, links) · UPDATED 2026-10-01
language
this page: en · bulgarian edition: /academy/blokove/moduli/01-00a_Блок_0_Част_1_LLM_Токенизация.html
next
01-00b · Block 0 part 2/3 · The AI stack: CUDA → runtime → framework → application
PURPOSE

Give practitioners a working mental model of LLMs without heavy maths: next-token prediction via self-attention; pre-training vs fine-tuning vs inference; parameters × quantization → memory; tokens (BPE) as the unit of cost; the context window as working memory; embeddings as semantic coordinates that power RAG.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

01-00b_Блок_0_Част_2_AI_Стек.html · The AI stack top-down · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
llmtransformerself-attentiontokenizationbpecontext-windowembeddingsragbulgarian
VERIFIED · 01.10.2026 UPDATED · 01.10.2026

What is an LLM and how it thinks in tokens

A solid foundation without unnecessary maths. You will understand how the transformer processes text, why the context window is your most important resource, why Bulgarian text costs more tokens, and how embeddings make meaning computable.

⏱ ~3–4 hours Beginner Block 0 · AI fundamentals Chapters 0.1 and 0.2
tiktoken🔒 local Ollama + embedding model🔒 local OpenAI API / Tokenizer🌐 global model

01What you will learn

02Before you start

03Steps

  1. What is an LLM — the transformer without deep maths

    A Large Language Model (LLM) is a neural network trained to predict the next token (a piece of a word) from all previous tokens. Simple next-word prediction, repeated billions of times, produces an apparent understanding of language.

    Before the transformer (2017, Google, "Attention Is All You Need") networks read text sequentially, word by word — and the end of a sentence "forgot" the beginning. The transformer solves this with self-attention: every word can "look at" every other word at once.

    StageWhat happens
    1. Inputtext is cut into tokens: "The cat" · "eats" · "fish"
    2. Vectorisationeach token becomes a vector (embedding) + its position in the sentence
    3. Processingself-attention + a feed-forward network, repeated over many layers (dozens)
    4. Probabilitieslogits — a probability for every possible next token
    5. Choicesampling with settings such as temperature and top_p

    Self-attention is simpler than it sounds: for every word the mechanism asks "which other words matter for understanding this one?". In "The cat eats fish because it is hungry", when processing "it", attention goes to "the cat", not the fish. This ability to link distant words is what makes transformers powerful.

    💡
    Analogy: attention while reading
    When you read, you don't give every word equal attention. In "He returned the book to the library because he had finished it" — "it" links to "the book", not "the library". Self-attention computes such links automatically.
  2. Pre-training, fine-tuning, inference — when to use which

    The three phases have fundamentally different needs for resources, data and purpose. Mixing them up is a common mistake when planning an AI project.

    PhaseWhat is doneGPUDataWhen
    Pre-trainingtrain the model from scratch on a huge corpusthousands of GPUs · monthstrillions of tokensonly large labs (OpenAI, Meta, Mistral…)
    Fine-tuningfurther train an existing model for a specific task1–8 GPUs · hours/dayshundreds to thousands of examplesa domain (medicine, law, house style)
    Inferenceuse a trained model to answer1 GPU (or CPU) · secondsthe input prompt onlyevery day, in every application
    ✅
    The rule for 95% of projects
    You will work almost only with inference. Fine-tuning is needed when prompt engineering and RAG are not enough. Pre-training — practically never, unless you have a compute budget in the millions.
  3. Model size — parameters and memory

    "7B", "13B", "70B" is the number of parameters (weights). More parameters = better understanding, but more memory and slower answers. Roughly: memory ≈ parameters × bytes per weight. Quantization (e.g. Q4 ≈ half a byte per weight) shrinks the size several times with little quality loss.

    Sizefp16 (no quantization)Q4In practice
    7B~14 GB~4 GBfast, good for simple tasks; runs on an ordinary card
    13B~26 GB~8 GBbalanced quality · a 16–24 GB class card
    34B~68 GB~20 GBQ4 fits a single 24 GB card; fp16 — server class
    70B~140 GB~40 GBserver hardware or a machine with large unified memory
    ⚠️
    Hallucination — the main risk
    An LLM generates statistically plausible text, not verified facts. It can invent names, dates and citations with full confidence. That is why RAG (Retrieval-Augmented Generation) matters so much — it "grounds" the model in real documents.
  4. What is a token — and why Bulgarian costs more

    A token is not a word. BPE (Byte Pair Encoding), the most common algorithm, cuts text by the frequency of character combinations: a common word = 1 token, a rare word = several, a language less frequent in training = more tokens.

    TextWordsTokens · GPT-4 (cl100k)Tokens · GPT-4o (o200k)
    "Изкуственият интелект промени света" (Bulgarian)41610
    "Artificial intelligence changed the world"565
    "невроендокринен" (rare Bulgarian medical word)1106
    📏
    UPDATED · 01.10.2026 — measured, not guessed
    The old lesson gave 11 / 5 / 6 tokens "for GPT-4" and "~2.5× more for Bulgarian". We counted with tiktoken 0.14. Over the full text of the lesson (the same content in both languages): Bulgarian is ~3.1 tokens per word with GPT-4 and ~2.2 with GPT-4o, English ~1.4. So Bulgarian costs ~2× more tokens with GPT-4 and ~1.5× with GPT-4o. Newer tokenizers are kinder to Cyrillic — but it stays more expensive.
    💰
    Tokens have a direct price
    With an API you pay per token. GPT-4o 🌐 global: $2.50 per 1M input tokens (OpenAI list price, in USD, as of 01.10.2026). 10,000 documents a month × 500 words × ~2.2 tokens = ~11M tokens ≈ $27.50 for input alone. Prompt optimisation is not aesthetics — it is a financial decision.
    UPDATED · 01.10.2026 — the old lesson used the launch price $0.005 / 1K ($5 / 1M) and ~700 tokens for 500 Bulgarian words; the correct figure is ~1,100.
    Python · count the tokens yourself 🔒 local
    import tiktoken
    enc = tiktoken.encoding_for_model("gpt-4o")   # o200k_base; for GPT-4: get_encoding("cl100k_base")
    for s in ["Изкуственият интелект промени света",
              "Artificial intelligence changed the world"]:
        print(len(enc.encode(s)), s)
  5. The context window — the model's working memory

    The context window is the maximum number of tokens the model "sees" at once — system prompt, history and the current query together. Everything outside it is invisible.

    ModelWindowNote
    GPT-3.5 Turbo 🌐16Kold model — for comparison; ~12,000 English words
    GPT-4o 🌐128K~200 pages of English text
    Claude Sonnet 5.5 🌐1MClaude Haiku 4.5 — 200K
    Gemini 3.1 Pro (preview) 🌐1M1,048,576 input tokens
    Llama 3.2 🔒 local128K1B and 3B; typical for open models
    🔄
    UPDATED · 01.10.2026 — windows keep growing
    The old lesson listed "GPT-4o / Claude Sonnet — 128K" and "Gemini 1.5 Pro — 1M". Claude Sonnet now has 1M, and Gemini 1.5 has been shut down — replaced here with Gemini 3.1 Pro (Gemini 2.5 Pro is now open only to existing users; checked 01.10.2026). These numbers change often: check the vendor's documentation before a project. And remember that a page of Bulgarian costs more tokens.

    A typical split in a long conversation: system prompt ~15%, history ~45%, current query ~20%, free buffer ~20%. With a 128K window, 57K tokens of history × $2.50 / 1M ≈ $0.14 for every single query — which is why history management is critical for cost.

    Three main strategies when history fills the window:

    Python · three context-management strategies
    # Strategy 1: sliding window — keep the system prompt + the last N messages
    def sliding_window(messages, max_messages=10):
        system = [m for m in messages if m["role"] == "system"]
        history = [m for m in messages if m["role"] != "system"]
        return system + history[-max_messages:]
    
    # Strategy 2: summarisation — shrink old history to a few sentences
    async def summarize_history(old_messages, llm_client):
        prompt = f"""Summarise this conversation history in 3–5 sentences,
    keeping the key facts and decisions:
    {old_messages}"""
        summary = await llm_client.complete(prompt)   # your client for the model
        return [{"role": "system", "content": f"Summary: {summary}"}]
    
    # Strategy 3: token budget — count and drop the oldest
    import tiktoken
    def count_tokens(messages, model="gpt-4o"):
        enc = tiktoken.encoding_for_model(model)
        return sum(len(enc.encode(m["content"])) + 4 for m in messages)  # +4 ≈ overhead tokens
    
    def trim_to_budget(messages, max_tokens=100_000):
        while count_tokens(messages) > max_tokens:
            for i, m in enumerate(messages):
                if m["role"] != "system":
                    messages.pop(i)
                    break
            else:
                break   # only system messages left — stop
        return messages
    ⚠️
    Trap: tiktoken counts for OpenAI only
    tiktoken is the tokenizer of OpenAI's models. Claude, Gemini and open models use their own tokenizers — counts will differ. For them use the vendor's token-counting tool or the model's own tokenizer. (UPDATED · 01.10.2026 — the old lesson claimed tiktoken was compatible with Claude.)
  6. Embeddings — words as coordinates

    An embedding is a dense vector of numbers that represents the meaning of a text. The key property: texts close in meaning have close vectors. This is the basis of RAG, semantic search and classification.

    🔑
    Distance = closeness of meaning
    vector("king") − vector("man") + vector("woman") ≈ vector("queen")
    The classic example shows that embeddings encode analogies, synonyms and opposites as geometry in a many-dimensional space.

    Picture a map: "cat", "dog", "fish", "bird" sit in one corner; "car", "train", "plane", "ship" in another; "freedom", "democracy", "rights" in a third. In reality vectors have hundreds to thousands of dimensions (e.g. 384, 768, 1024, 1536), not two — but the principle is the same. The standard measure is cosine similarity: it looks at the angle between vectors, not their length.

    Python · embeddings with a local model via Ollama 🔒 local
    import ollama
    import numpy as np
    
    # Local: no internet, no per-token fee. First: ollama pull bge-m3
    def embed(text: str) -> list[float]:
        r = ollama.embed(model="bge-m3", input=text)   # multilingual, 1024 dimensions
        return r.embeddings[0]
    
    def cosine_similarity(v1, v2) -> float:
        a, b = np.array(v1), np.array(v2)
        return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))
    
    for w1, w2 in [("котка", "коте"), ("котка", "куче"), ("котка", "автомобил")]:
        print(f"{w1} ↔ {w2}: {cosine_similarity(embed(w1), embed(w2)):.3f}")
    # cat–kitten, cat–dog, cat–car. Expect cat–kitten closest, cat–car farthest.
    # Exact numbers depend on the model — run it and see.
    ModelDimensionsLanguagesSizeNote
    bge-m3 🔒1024multilingual (100+)~1.2 GBgood start for Bulgarian · in Ollama · 8K window
    nomic-embed-text 🔒768mainly English~274 MBfast, 2K window · in Ollama
    mxbai-embed-large 🔒1024mainly English~670 MBstrong for English · 512 window · in Ollama
    all-MiniLM-L6-v2 🔒384English~90 MBvery fast · for English-only systems
    text-embedding-3-small 🌐1536multilingualAPI onlyOpenAI · $0.02 / 1M tokens · not local
    🎯
    UPDATED · 01.10.2026 — the Bulgarian recommendation changed
    The old lesson recommended nomic-embed-text and mxbai-embed-large as multilingual. Checked: both are trained mainly on English (their Hugging Face cards are tagged "en"). For RAG in Bulgarian start with a multilingual model such as bge-m3, and always test with real documents from your domain — quality differences are large.

04Check

1. Which factor MOST determines the memory needed for inference?

2. A medical company wants an AI assistant for clinical protocols. What is the MOST suitable first step?

3. A Bulgarian text of 500 words — roughly how many tokens (GPT-4 / GPT-4o)?

4. Why is cosine similarity preferred over Euclidean distance for embeddings?

Summary

05What's next

06Sources

  1. "Attention Is All You Need" (2017) — the original paper; section 3 describes the layers.
  2. The Illustrated Transformer — Jay Alammar — the best visual explanation.
  3. 3Blue1Brown — Transformers, the tech behind LLMs — video, ~27 min (formerly titled "But what is a GPT?").
  4. OpenAI Tokenizer — type text and see the tokens.
  5. tiktoken — OpenAI's token-counting library.
  6. OpenAI — API pricing — GPT-4o and text-embedding-3-small.
  7. Anthropic — Claude models overview — context windows.
  8. Google — Gemini models — token limits.
  9. bge-m3, nomic-embed-text, mxbai-embed-large — the embedding models in Ollama.
  10. MTEB Leaderboard — ranking of embedding models by task.
  11. KAGAMI's own measurement, 01.10.2026: tiktoken 0.14 on the text of this lesson in both languages.