The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewVERIFIED 2026-10-01 · UPDATED 2026-10-01
IDENTITY
module
01-01a · vLLM: a local inference server for many users
series
KAGAMI Academy · Blocks (AI engineer) · Block 1 — Production inference · part 1/2
level
Intermediate → Advanced
duration
~3–4 h incl. lab
trust_label
VERIFIED 2026-10-01 (versions, commands, links checked against official sources) · UPDATED 2026-10-01 · NOT TESTED by KAGAMI on own hardware
language
human view: en · bulgarian edition: /academy/blokove/moduli/01-01a_Блок_1_Част_1_vLLM_Inference.html
prerequisites
Block 0 (LLM basics, AI stack, lab)
next
01-01b · Docker Compose production stack
PURPOSE

Decide between Ollama (developer tool, easy start) and vLLM (serving engine for many concurrent requests), understand why vLLM is fast (PagedAttention + continuous batching), install it on x86 or ARM64 (Grace/GB10), serve an OpenAI-compatible API, and use pre-quantized AWQ checkpoints.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

01-01b · Docker Compose production stack (vLLM + Nginx + Redis + Prometheus + Grafana) · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
vllmollamapaged-attentioncontinuous-batchingopenai-apiawqarm64gb10
VERIFIED · 01.10.2026 UPDATED · 01.10.2026

vLLM: a local inference server for many users

Ollama is for development and small teams. When dozens of people query the model at the same time, you need vLLM. Here we unpack the two mechanisms that make it fast — PagedAttention and continuous batching — and bring up an OpenAI-compatible server on an x86 and on an ARM64 machine.

⏱ ~3–4 h incl. lab Intermediate → advanced Block 1 · Production inference vLLM · PagedAttention · AWQ
vLLM🔒 local Ollama🔒 local Docker🔒 local Hugging Face Hub🌐 global (model downloads only)
🔄
UPDATED · 01.10.2026 — what changed
ARM64: vLLM now ships official aarch64 packages on PyPI, and the vllm/vllm-openai image covers amd64 and arm64; NVIDIA publishes its own image for GB10 (DGX Spark / ASUS Ascent GX10). The old "limited support" no longer applies.
Commands: vllm serve instead of python -m vllm.entrypoints.openai.api_server; Python 3.10–3.13 (was 3.9–3.12); install with uv.
AWQ: the --quantization awq flag needs an AWQ model — the examples now use one; the casperhansen/… repository is no longer accessible; AutoAWQ is deprecated in favour of llm-compressor.
Facts: a PagedAttention block is a fixed number of tokens (not "~16 KB"); "24×" is versus HF Transformers, while the paper reports 2–4× versus FasterTransformer and Orca; Ollama is not pure FIFO — it has OLLAMA_NUM_PARALLEL; the NGC link is fixed.
Speed figures in the examples are illustrative, not measured by us — ⚠️ measure on your own hardware.

01What you will learn

02Before you start

03Steps

  1. Ollama or vLLM — which one when

    The two tools solve different problems. Ollama is a developer tool — easy start, good model management. vLLM is a serving engine — built to hold many requests at once. An Olympic sprinter versus a jeep: each is best on its own terrain. A common architectural mistake is using the wrong one for the situation.

    CriterionOllamavLLMWhich when
    Installation1 commandpip/uv or DockerOllama to start
    Parallel requestslimited — OLLAMA_NUM_PARALLEL (default 1 per model)continuous batching, hundreds of sequencesvLLM for many concurrent users
    Context memoryreserves NUM_PARALLEL × contextPagedAttention — dynamicvLLM for many sessions
    OpenAI API✓ compatible✓ compatible (more complete)both work with LangChain
    QuantizationGGUF (llama.cpp)AWQ, GPTQ, FP8 and morevLLM for AWQ/FP8
    ARM64 (GB10)✓✓ (aarch64 packages and images — see step 3)both — updated 01.10.2026
    Monitoringminimalbuilt-in Prometheus metricsvLLM for production
    ✅
    Rule
    Development and a small team → Ollama. A production API or RAG with many concurrent users → vLLM. Decide by the real load (how many people at once, how long the context), not "on principle".
    🏥
    Example scenario (illustration, not a case we measured)
    A hospital runs RAG over national health-insurance protocols on Ollama with default settings. When five doctors ask at once, the requests wait one after another and answers take tens of seconds. Same hardware and model, but on vLLM, the requests are processed together and the wait drops several times. The lesson: Ollama is not "at fault" — it was built for something else.
  2. PagedAttention and continuous batching — why vLLM is fast

    While generating, the model keeps a KV cache (keys and values) for every token in the context. The naive approach reserves space for the whole possible context of each request up front — even if the answer is short. Much of the memory stays reserved but empty, and new requests do not fit.

    PagedAttention (UC Berkeley, SOSP 2023) borrows the idea of virtual memory from operating systems: the KV cache is split into blocks holding a fixed number of tokens, allocated on demand and not necessarily contiguous. Space is lost only in the last block of each request — under 4% according to the vLLM team. The freed memory goes to more concurrent requests.

    Naive allocationPagedAttention
    How it allocatesthe maximum context — up frontblock by block as the request grows
    Memory wastelarge (reserved but unused)only in the last block (<4%)
    Resultfew requests fitmany more requests on the same hardware

    Continuous batching is the second mechanism. With a static batch everyone waits for the longest request: if one generates 500 tokens and another 10, the second sits idle. With continuous batching a finished request leaves and a new one takes its place at every iteration — the GPU does not sit idle.

    📄
    What the sources actually claim
    The vLLM blog (2023): up to 24× higher throughput than HuggingFace Transformers and up to 3.5× than TGI. The paper (Kwon et al., SOSP 2023): 2–4× versus FasterTransformer and Orca at the same latency. The real gain depends on the load — measure it yourself.
  3. Installation: x86, ARM64 (GB10) and Docker

    x86 + NVIDIA. The official path today uses uv — it picks the right PyTorch for your driver by itself.

    bash · x86 or ARM64 Linux · install
    # Checks before installing
    python3 --version     # 3.10–3.13
    nvidia-smi            # the GPU and driver are visible
    
    # Environment and install (recommended path in the vLLM docs)
    uv venv --python 3.12 --seed
    source .venv/bin/activate
    uv pip install vllm --torch-backend=auto
    
    python -c "import vllm; print(vllm.__version__)"

    ARM64 with GB10 (NVIDIA DGX Spark, ASUS Ascent GX10 and similar). Until recently this was the weak spot. As of 01.10.2026 PyPI has aarch64 vLLM packages, and NVIDIA maintains a ready image for these machines — the safest path, because it is built for the exact architecture.

    bash · GB10 · NVIDIA's image
    # The tag changes — take the current one from NGC or NVIDIA's DGX Spark playbook
    export VLLM_TAG=26.05.post1-py3   # the playbook's tag as of 01.10.2026
    docker pull nvcr.io/nvidia/vllm:${VLLM_TAG}
    
    docker run -d --gpus all -p 8000:8000 --name vllm-server \
      nvcr.io/nvidia/vllm:${VLLM_TAG} \
      vllm serve Qwen/Qwen2.5-7B-Instruct-AWQ

    Docker, the general image. vllm/vllm-openai:latest now covers both amd64 and arm64.

    bash · Docker · the official vLLM image
    docker run --gpus all --ipc=host -p 8000:8000 \
      -v ~/.cache/huggingface:/root/.cache/huggingface \
      --name vllm-server \
      vllm/vllm-openai:latest \
      --model Qwen/Qwen2.5-7B-Instruct-AWQ \
      --max-model-len 16384 \
      --served-model-name qwen-7b
    
    # Test right after start-up
    curl http://localhost:8000/v1/models
    ⚠️
    Trap: gated models
    Meta's Llama models on Hugging Face are gated: you must accept the licence and pass a token (HF_TOKEN). Without it the download stops with a 401 error. Never put the token in scripts you share.
  4. First server — three ways of working

    Way 1 — an OpenAI-compatible HTTP server. This is the production way: any application that talks to OpenAI talks to it too.

    bash · vllm serve with the main settings
    # The model is AWQ — vLLM detects the quantization from its config
    vllm serve Qwen/Qwen2.5-7B-Instruct-AWQ \
      --port 8000 \
      --max-model-len 16384 \
      --max-num-seqs 128 \
      --gpu-memory-utilization 0.90 \
      --enable-prefix-caching \
      --served-model-name qwen-7b
    
    # --max-num-seqs           how many sequences per iteration
    # --enable-prefix-caching  reuses a shared leading text (the system prompt)
    # --served-model-name      the name clients use to call the model
    # --trust-remote-code      only for models with custom code, and only if you trust them
    
    # Ask with curl
    curl http://localhost:8000/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d '{
        "model": "qwen-7b",
        "messages": [
          {"role": "system", "content": "You are the AI assistant of a small business."},
          {"role": "user",   "content": "Explain PagedAttention in 3 sentences."}
        ],
        "temperature": 0.2,
        "max_tokens": 300
      }'
    
    # Streaming (for a chat UI): add "stream": true

    Way 2 — Python for batch jobs. When you have a list of questions and no need for a server: hand them all over at once and vLLM schedules them itself.

    python · batch_inference.py
    from vllm import LLM, SamplingParams
    import time
    
    llm = LLM(
        model="Qwen/Qwen2.5-7B-Instruct-AWQ",
        max_model_len=8192,
        enable_prefix_caching=True,
        gpu_memory_utilization=0.90,   # 10% headroom
    )
    params = SamplingParams(temperature=0.2, top_p=0.9, max_tokens=500,
                            repetition_penalty=1.05)
    
    PROMPTS = [
        "Explain the difference between LoRA and full fine-tuning.",
        "Write an example RAG chain with LangChain.",
        "What is PagedAttention and why does it matter?",
        "How is cosine similarity computed?",
    ]
    
    t0 = time.time()
    outputs = llm.generate(PROMPTS, params)
    elapsed = time.time() - t0
    tokens = sum(len(o.outputs[0].token_ids) for o in outputs)
    print(f"{len(PROMPTS)} requests · {tokens} tokens · {elapsed:.1f} s · {tokens/elapsed:.0f} tok/s")
    for p, o in zip(PROMPTS, outputs):
        print("\n?", p[:60]); print(o.outputs[0].text[:200])

    Way 3 — swapping Ollama out of LangChain. vLLM speaks the OpenAI API, so in the RAG code you change only the model's address. Embeddings can stay on Ollama.

    python · LangChain: Ollama → vLLM
    from langchain_openai import ChatOpenAI
    
    # BEFORE: llm = ChatOllama(model="qwen2.5:7b", temperature=0.2)
    llm = ChatOpenAI(
        base_url="http://localhost:8000/v1",
        api_key="not-needed",        # vLLM needs no key unless you set one
        model="qwen-7b",
        temperature=0.2,
        max_tokens=500,
    )
    # chain = ({"context": retriever, "question": RunnablePassthrough()}
    #          | prompt | llm | StrOutputParser())   ← the only change
    
    # Quick comparison: the same 10 requests to both servers
    import time, requests
    def bench(url, model, n=10):
        t0 = time.time()
        for _ in range(n):
            requests.post(url, json={"model": model, "max_tokens": 80,
                "messages": [{"role": "user", "content": "Explain RAG in 50 words."}]})
        return time.time() - t0
    print("Ollama:", bench("http://localhost:11434/v1/chat/completions", "qwen2.5:7b"))
    print("vLLM:  ", bench("http://localhost:8000/v1/chat/completions",  "qwen-7b"))
    ⚠️
    These are not our measurements
    The old version of this lesson showed "42 s vs 14 s" for 10 requests — an illustration, not a result of our test. Sequential requests barely show vLLM's advantage; it shows with parallel requests. Run the comparison with several threads and record your own numbers.
  5. Quantization for production — AWQ, GPTQ, FP8

    AWQ (Lin et al., MIT, 2023) is 4-bit quantization that finds the most important weights (by the real activations) and keeps them more precise. That is why it loses less quality than plain 4-bit quantization at the same memory.

    FormatBitsMemory for a 7B modelWhen
    BF16 (no quantization)16~14 GB for weights alonemachines with plenty of memory (e.g. 128 GB unified)
    AWQ 4-bit4~4–5 GBGPUs with little memory; good quality
    GPTQ 4-bit4~4–5 GBalternative to AWQ
    FP88~7–8 GBnewer GPUs with hardware FP8 support
    GGUF Q4_K_M4~4–5 GBOllama / llama.cpp only
    ⚠️
    No blanket quality percentages
    We removed the old version's "quality vs fp16" percentages and "1.2–1.5× faster" — we found no source confirming them across models. Memory in the table is approximate. Quality depends on the model and the task; compare on your own data.
    python · download a ready AWQ model and start
    # Look for ready AWQ versions: huggingface.co/models?search=awq
    # Official AWQ versions are published e.g. by Qwen; for Llama 3.1 — hugging-quants
    from huggingface_hub import snapshot_download
    path = snapshot_download("Qwen/Qwen2.5-7B-Instruct-AWQ",
                             local_dir="./models/qwen2.5-7b-awq")
    
    from vllm import LLM
    llm = LLM(model=path,
              dtype="float16",            # AWQ kernels run in fp16
              max_model_len=32768,
              gpu_memory_utilization=0.85,
              enable_chunked_prefill=True)  # steadier latency on long prompts
    
    # The same from the command line:
    # vllm serve ./models/qwen2.5-7b-awq --max-model-len 32768
    🔧
    Quantize it yourself?
    The AutoAWQ library is deprecated; the vLLM team has taken the functionality into llm-compressor. If there is no ready AWQ version of your model, quantize with llm-compressor — but look for a ready one first.
  6. GB10 with 128 GB unified memory — strategy

    Machines such as DGX Spark and ASUS Ascent GX10 have 128 GB of memory shared by the CPU and the GPU. That allows large models without quantization, but the memory is slower than that of large data-centre GPUs. For a small team or a pilot course a well-tuned Ollama is often enough; vLLM (from NVIDIA's image) is the next step as concurrent users grow.

    bash · Ollama for a few parallel requests (systemd)
    sudo mkdir -p /etc/systemd/system/ollama.service.d/
    sudo tee /etc/systemd/system/ollama.service.d/override.conf >/dev/null <<'EOF'
    [Service]
    Environment="OLLAMA_NUM_PARALLEL=4"
    Environment="OLLAMA_MAX_LOADED_MODELS=2"
    Environment="OLLAMA_KEEP_ALIVE=30m"
    Environment="OLLAMA_FLASH_ATTENTION=1"
    EOF
    sudo systemctl daemon-reload && sudo systemctl restart ollama
    ollama ps   # which models are loaded
    
    # OLLAMA_NUM_PARALLEL     parallel requests per model (default 1)
    #                         memory grows with NUM_PARALLEL × context length
    # OLLAMA_FLASH_ATTENTION  Ollama enables it automatically where it can; =1 forces it
    python · check: 8 parallel requests
    import requests, concurrent.futures, time
    URL = "http://localhost:11434/v1/chat/completions"   # or :8000 for vLLM
    MODEL = "qwen2.5:7b"
    def one(i):
        t0 = time.time()
        requests.post(URL, json={"model": MODEL, "max_tokens": 50,
            "messages": [{"role": "user", "content": f"Question {i}: what is AI?"}]})
        return time.time() - t0
    t0 = time.time()
    with concurrent.futures.ThreadPoolExecutor(max_workers=8) as ex:
        r = list(ex.map(one, range(8)))
    print(f"total {time.time()-t0:.1f}s · avg {sum(r)/len(r):.1f}s · sequential would be ~{sum(r):.1f}s")
    ⛔
    Never put the model on the internet bare
    Neither Ollama nor vLLM asks for a password by default. Put a reverse proxy (e.g. nginx) in front with rate limiting and authentication — or keep them on the internal network only. How to assemble that is in the next part.
    🎯
    When to move from Ollama to vLLM
    Stay on Ollama: development, testing, a small team, a few concurrent requests.
    Move to vLLM: many concurrent users, need for FP8/AWQ and metrics, agreed response times. ⚠️ Measure the exact threshold on your own hardware — the old version claimed "8 users on a 70B model", which we have not verified.

04Check

1. A RAG system with 30 concurrent users answers slowly on Ollama with default settings. Which vLLM mechanism helps most?

2. Why does AWQ lose less quality than plain 4-bit quantization?

3. What does PagedAttention split the KV cache into?

4. You want vLLM on an ARM64 machine with GB10 (as of 01.10.2026). What is the safest path?

05What's next

06Sources

  1. vLLM — Quickstart — installing with uv, vllm serve, the OpenAI-compatible server on port 8000.
  2. vLLM — GPU installation — requirements (Linux, Python 3.10–3.13, compute capability ≥ 7.5), aarch64.
  3. vllm on PyPI — version 0.30.0 as of 01.10.2026, packages for x86_64 and aarch64.
  4. NVIDIA DGX Spark playbooks — vLLM — the nvcr.io/nvidia/vllm image for GB10.
  5. NVIDIA NGC — vLLM container — current tags.
  6. Kwon et al. — Efficient Memory Management for LLM Serving with PagedAttention (SOSP 2023) — sections 3 and 4 are the core.
  7. The vLLM blog (2023) — 24× vs HF Transformers, waste under 4%.
  8. Lin et al. — AWQ: Activation-aware Weight Quantization.
  9. vLLM — quantization (AWQ, FP8 and more) · llm-compressor.
  10. Qwen2.5-7B-Instruct-AWQ · Llama 3.1 8B AWQ (hugging-quants) — ready AWQ models.
  11. Ollama FAQ — OLLAMA_NUM_PARALLEL, OLLAMA_MAX_LOADED_MODELS, Flash Attention.