What is an LLM and how it thinks in tokens
A solid foundation without unnecessary maths. You will understand how the transformer processes text, why the context window is your most important resource, why Bulgarian text costs more tokens, and how embeddings make meaning computable.
01What you will learn
- What an LLM is and how self-attention "connects" the words of a sentence.
- When you need inference, when fine-tuning, and why almost never pre-training.
- How model size and quantization determine the memory you need.
- What a token is, why Bulgarian costs more tokens and how that hits the budget.
- What the context window is and three ways to manage it.
- What embeddings are, how they are compared and which model to choose for RAG in Bulgarian.
02Before you start
- Python basics — you will read and run short examples.
- An intuition for a vector as "a list of numbers" — no linear-algebra exam.
- Optional: Python 3 with
pip install tiktoken🔒 local to count tokens yourself. - Optional: Ollama 🔒 local with an embedding model — for the cosine-similarity example.
03Steps
-
What is an LLM — the transformer without deep maths
A Large Language Model (LLM) is a neural network trained to predict the next token (a piece of a word) from all previous tokens. Simple next-word prediction, repeated billions of times, produces an apparent understanding of language.
Before the transformer (2017, Google, "Attention Is All You Need") networks read text sequentially, word by word — and the end of a sentence "forgot" the beginning. The transformer solves this with self-attention: every word can "look at" every other word at once.
Stage What happens 1. Input text is cut into tokens: "The cat" · "eats" · "fish" 2. Vectorisation each token becomes a vector (embedding) + its position in the sentence 3. Processing self-attention + a feed-forward network, repeated over many layers (dozens) 4. Probabilities logits — a probability for every possible next token 5. Choice sampling with settings such as temperature and top_p Self-attention is simpler than it sounds: for every word the mechanism asks "which other words matter for understanding this one?". In "The cat eats fish because it is hungry", when processing "it", attention goes to "the cat", not the fish. This ability to link distant words is what makes transformers powerful.
💡Analogy: attention while readingWhen you read, you don't give every word equal attention. In "He returned the book to the library because he had finished it" — "it" links to "the book", not "the library". Self-attention computes such links automatically. -
Pre-training, fine-tuning, inference — when to use which
The three phases have fundamentally different needs for resources, data and purpose. Mixing them up is a common mistake when planning an AI project.
Phase What is done GPU Data When Pre-training train the model from scratch on a huge corpus thousands of GPUs · months trillions of tokens only large labs (OpenAI, Meta, Mistral…) Fine-tuning further train an existing model for a specific task 1–8 GPUs · hours/days hundreds to thousands of examples a domain (medicine, law, house style) Inference use a trained model to answer 1 GPU (or CPU) · seconds the input prompt only every day, in every application ✅The rule for 95% of projectsYou will work almost only with inference. Fine-tuning is needed when prompt engineering and RAG are not enough. Pre-training — practically never, unless you have a compute budget in the millions. -
Model size — parameters and memory
"7B", "13B", "70B" is the number of parameters (weights). More parameters = better understanding, but more memory and slower answers. Roughly: memory ≈ parameters × bytes per weight. Quantization (e.g. Q4 ≈ half a byte per weight) shrinks the size several times with little quality loss.
Size fp16 (no quantization) Q4 In practice 7B ~14 GB ~4 GB fast, good for simple tasks; runs on an ordinary card 13B ~26 GB ~8 GB balanced quality · a 16–24 GB class card 34B ~68 GB ~20 GB Q4 fits a single 24 GB card; fp16 — server class 70B ~140 GB ~40 GB server hardware or a machine with large unified memory ⚠️Hallucination — the main riskAn LLM generates statistically plausible text, not verified facts. It can invent names, dates and citations with full confidence. That is why RAG (Retrieval-Augmented Generation) matters so much — it "grounds" the model in real documents. -
What is a token — and why Bulgarian costs more
A token is not a word. BPE (Byte Pair Encoding), the most common algorithm, cuts text by the frequency of character combinations: a common word = 1 token, a rare word = several, a language less frequent in training = more tokens.
Text Words Tokens · GPT-4 (cl100k) Tokens · GPT-4o (o200k) "Изкуственият интелект промени света" (Bulgarian) 4 16 10 "Artificial intelligence changed the world" 5 6 5 "невроендокринен" (rare Bulgarian medical word) 1 10 6 📏UPDATED · 01.10.2026 — measured, not guessedThe old lesson gave 11 / 5 / 6 tokens "for GPT-4" and "~2.5× more for Bulgarian". We counted with tiktoken 0.14. Over the full text of the lesson (the same content in both languages): Bulgarian is ~3.1 tokens per word with GPT-4 and ~2.2 with GPT-4o, English ~1.4. So Bulgarian costs ~2× more tokens with GPT-4 and ~1.5× with GPT-4o. Newer tokenizers are kinder to Cyrillic — but it stays more expensive.💰Tokens have a direct priceWith an API you pay per token. GPT-4o 🌐 global: $2.50 per 1M input tokens (OpenAI list price, in USD, as of 01.10.2026). 10,000 documents a month × 500 words × ~2.2 tokens = ~11M tokens ≈ $27.50 for input alone. Prompt optimisation is not aesthetics — it is a financial decision.
UPDATED · 01.10.2026 — the old lesson used the launch price $0.005 / 1K ($5 / 1M) and ~700 tokens for 500 Bulgarian words; the correct figure is ~1,100.Python · count the tokens yourself 🔒 localimport tiktoken enc = tiktoken.encoding_for_model("gpt-4o") # o200k_base; for GPT-4: get_encoding("cl100k_base") for s in ["Изкуственият интелект промени света", "Artificial intelligence changed the world"]: print(len(enc.encode(s)), s) -
The context window — the model's working memory
The context window is the maximum number of tokens the model "sees" at once — system prompt, history and the current query together. Everything outside it is invisible.
Model Window Note GPT-3.5 Turbo 🌐 16K old model — for comparison; ~12,000 English words GPT-4o 🌐 128K ~200 pages of English text Claude Sonnet 5.5 🌐 1M Claude Haiku 4.5 — 200K Gemini 3.1 Pro (preview) 🌐 1M 1,048,576 input tokens Llama 3.2 🔒 local 128K 1B and 3B; typical for open models 🔄UPDATED · 01.10.2026 — windows keep growingThe old lesson listed "GPT-4o / Claude Sonnet — 128K" and "Gemini 1.5 Pro — 1M". Claude Sonnet now has 1M, and Gemini 1.5 has been shut down — replaced here with Gemini 3.1 Pro (Gemini 2.5 Pro is now open only to existing users; checked 01.10.2026). These numbers change often: check the vendor's documentation before a project. And remember that a page of Bulgarian costs more tokens.A typical split in a long conversation: system prompt ~15%, history ~45%, current query ~20%, free buffer ~20%. With a 128K window, 57K tokens of history × $2.50 / 1M ≈ $0.14 for every single query — which is why history management is critical for cost.
Three main strategies when history fills the window:
Python · three context-management strategies# Strategy 1: sliding window — keep the system prompt + the last N messages def sliding_window(messages, max_messages=10): system = [m for m in messages if m["role"] == "system"] history = [m for m in messages if m["role"] != "system"] return system + history[-max_messages:] # Strategy 2: summarisation — shrink old history to a few sentences async def summarize_history(old_messages, llm_client): prompt = f"""Summarise this conversation history in 3–5 sentences, keeping the key facts and decisions: {old_messages}""" summary = await llm_client.complete(prompt) # your client for the model return [{"role": "system", "content": f"Summary: {summary}"}] # Strategy 3: token budget — count and drop the oldest import tiktoken def count_tokens(messages, model="gpt-4o"): enc = tiktoken.encoding_for_model(model) return sum(len(enc.encode(m["content"])) + 4 for m in messages) # +4 ≈ overhead tokens def trim_to_budget(messages, max_tokens=100_000): while count_tokens(messages) > max_tokens: for i, m in enumerate(messages): if m["role"] != "system": messages.pop(i) break else: break # only system messages left — stop return messages⚠️Trap: tiktoken counts for OpenAI onlytiktoken is the tokenizer of OpenAI's models. Claude, Gemini and open models use their own tokenizers — counts will differ. For them use the vendor's token-counting tool or the model's own tokenizer. (UPDATED · 01.10.2026 — the old lesson claimed tiktoken was compatible with Claude.) -
Embeddings — words as coordinates
An embedding is a dense vector of numbers that represents the meaning of a text. The key property: texts close in meaning have close vectors. This is the basis of RAG, semantic search and classification.
🔑Distance = closeness of meaningvector("king") − vector("man") + vector("woman") ≈ vector("queen")
The classic example shows that embeddings encode analogies, synonyms and opposites as geometry in a many-dimensional space.Picture a map: "cat", "dog", "fish", "bird" sit in one corner; "car", "train", "plane", "ship" in another; "freedom", "democracy", "rights" in a third. In reality vectors have hundreds to thousands of dimensions (e.g. 384, 768, 1024, 1536), not two — but the principle is the same. The standard measure is cosine similarity: it looks at the angle between vectors, not their length.
Python · embeddings with a local model via Ollama 🔒 localimport ollama import numpy as np # Local: no internet, no per-token fee. First: ollama pull bge-m3 def embed(text: str) -> list[float]: r = ollama.embed(model="bge-m3", input=text) # multilingual, 1024 dimensions return r.embeddings[0] def cosine_similarity(v1, v2) -> float: a, b = np.array(v1), np.array(v2) return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))) for w1, w2 in [("котка", "коте"), ("котка", "куче"), ("котка", "автомобил")]: print(f"{w1} ↔ {w2}: {cosine_similarity(embed(w1), embed(w2)):.3f}") # cat–kitten, cat–dog, cat–car. Expect cat–kitten closest, cat–car farthest. # Exact numbers depend on the model — run it and see.Model Dimensions Languages Size Note bge-m3 🔒 1024 multilingual (100+) ~1.2 GB good start for Bulgarian · in Ollama · 8K window nomic-embed-text 🔒 768 mainly English ~274 MB fast, 2K window · in Ollama mxbai-embed-large 🔒 1024 mainly English ~670 MB strong for English · 512 window · in Ollama all-MiniLM-L6-v2 🔒 384 English ~90 MB very fast · for English-only systems text-embedding-3-small 🌐 1536 multilingual API only OpenAI · $0.02 / 1M tokens · not local 🎯UPDATED · 01.10.2026 — the Bulgarian recommendation changedThe old lesson recommended nomic-embed-text and mxbai-embed-large as multilingual. Checked: both are trained mainly on English (their Hugging Face cards are tagged "en"). For RAG in Bulgarian start with a multilingual model such as bge-m3, and always test with real documents from your domain — quality differences are large.
04Check
1. Which factor MOST determines the memory needed for inference?
2. A medical company wants an AI assistant for clinical protocols. What is the MOST suitable first step?
3. A Bulgarian text of 500 words — roughly how many tokens (GPT-4 / GPT-4o)?
4. Why is cosine similarity preferred over Euclidean distance for embeddings?
Summary
- An LLM predicts the next token via self-attention — it doesn't "think", it applies statistical patterns.
- Pre-training ≠ fine-tuning ≠ inference — mixing them up is a costly mistake.
- Parameters × quantization = memory: 7B Q4 ≈ 4 GB · 70B Q4 ≈ 40 GB.
- Token ≠ word. Bulgarian costs ~1.5–2× more tokens than English for the same content.
- The context window is working memory — everything outside it is invisible.
- Embeddings = coordinates of meaning; compared with cosine similarity.
- For RAG in Bulgarian — a multilingual embedding model, tested on your own documents.
05What's next
06Sources
- "Attention Is All You Need" (2017) — the original paper; section 3 describes the layers.
- The Illustrated Transformer — Jay Alammar — the best visual explanation.
- 3Blue1Brown — Transformers, the tech behind LLMs — video, ~27 min (formerly titled "But what is a GPT?").
- OpenAI Tokenizer — type text and see the tokens.
- tiktoken — OpenAI's token-counting library.
- OpenAI — API pricing — GPT-4o and text-embedding-3-small.
- Anthropic — Claude models overview — context windows.
- Google — Gemini models — token limits.
- bge-m3, nomic-embed-text, mxbai-embed-large — the embedding models in Ollama.
- MTEB Leaderboard — ranking of embedding models by task.
- KAGAMI's own measurement, 01.10.2026: tiktoken 0.14 on the text of this lesson in both languages.