vLLM: a local inference server for many users
Ollama is for development and small teams. When dozens of people query the model at the same time, you need vLLM. Here we unpack the two mechanisms that make it fast — PagedAttention and continuous batching — and bring up an OpenAI-compatible server on an x86 and on an ARM64 machine.
vllm/vllm-openai image covers amd64 and arm64; NVIDIA publishes its own image for GB10 (DGX Spark / ASUS Ascent GX10). The old "limited support" no longer applies.Commands:
vllm serve instead of python -m vllm.entrypoints.openai.api_server; Python 3.10–3.13 (was 3.9–3.12); install with uv.AWQ: the
--quantization awq flag needs an AWQ model — the examples now use one; the casperhansen/… repository is no longer accessible; AutoAWQ is deprecated in favour of llm-compressor.Facts: a PagedAttention block is a fixed number of tokens (not "~16 KB"); "24×" is versus HF Transformers, while the paper reports 2–4× versus FasterTransformer and Orca; Ollama is not pure FIFO — it has
OLLAMA_NUM_PARALLEL; the NGC link is fixed.Speed figures in the examples are illustrative, not measured by us — ⚠️ measure on your own hardware.
01What you will learn
- When Ollama is enough and when it is time for vLLM — by the number of concurrent users, not by fashion.
- How PagedAttention and continuous batching free memory and the GPU for more requests.
- How to install vLLM on x86, on ARM64 (GB10) and in Docker.
- Three ways to work: an OpenAI-compatible server, Python for batch jobs, swapping Ollama out of LangChain with one line.
- When to quantize (AWQ, FP8) and what to check when picking a model.
- How to tune Ollama for a few parallel requests while you don't need vLLM yet.
02Before you start
- Block 0 completed (how an LLM works, the AI stack, the lab).
- A Linux machine with an NVIDIA GPU (compute capability 7.5 or newer) — or an ARM64 machine with GB10 and 128 GB of unified memory. vLLM does not run natively on Windows (only via WSL).
- Python 3.10–3.13;
uvrecommended for the environment. - A Hugging Face account 🌐 global if you use gated models such as Llama (accept the licence and use a token). The examples here use Qwen, which is not gated.
- Ollama 🔒 local — for the comparison and for step 6.
03Steps
-
Ollama or vLLM — which one when
The two tools solve different problems. Ollama is a developer tool — easy start, good model management. vLLM is a serving engine — built to hold many requests at once. An Olympic sprinter versus a jeep: each is best on its own terrain. A common architectural mistake is using the wrong one for the situation.
Criterion Ollama vLLM Which when Installation 1 command pip/uv or Docker Ollama to start Parallel requests limited — OLLAMA_NUM_PARALLEL(default 1 per model)continuous batching, hundreds of sequences vLLM for many concurrent users Context memory reserves NUM_PARALLEL × contextPagedAttention — dynamic vLLM for many sessions OpenAI API ✓ compatible ✓ compatible (more complete) both work with LangChain Quantization GGUF (llama.cpp) AWQ, GPTQ, FP8 and more vLLM for AWQ/FP8 ARM64 (GB10) ✓ ✓ (aarch64 packages and images — see step 3) both — updated 01.10.2026 Monitoring minimal built-in Prometheus metrics vLLM for production ✅RuleDevelopment and a small team → Ollama. A production API or RAG with many concurrent users → vLLM. Decide by the real load (how many people at once, how long the context), not "on principle".🏥Example scenario (illustration, not a case we measured)A hospital runs RAG over national health-insurance protocols on Ollama with default settings. When five doctors ask at once, the requests wait one after another and answers take tens of seconds. Same hardware and model, but on vLLM, the requests are processed together and the wait drops several times. The lesson: Ollama is not "at fault" — it was built for something else. -
PagedAttention and continuous batching — why vLLM is fast
While generating, the model keeps a KV cache (keys and values) for every token in the context. The naive approach reserves space for the whole possible context of each request up front — even if the answer is short. Much of the memory stays reserved but empty, and new requests do not fit.
PagedAttention (UC Berkeley, SOSP 2023) borrows the idea of virtual memory from operating systems: the KV cache is split into blocks holding a fixed number of tokens, allocated on demand and not necessarily contiguous. Space is lost only in the last block of each request — under 4% according to the vLLM team. The freed memory goes to more concurrent requests.
Naive allocation PagedAttention How it allocates the maximum context — up front block by block as the request grows Memory waste large (reserved but unused) only in the last block (<4%) Result few requests fit many more requests on the same hardware Continuous batching is the second mechanism. With a static batch everyone waits for the longest request: if one generates 500 tokens and another 10, the second sits idle. With continuous batching a finished request leaves and a new one takes its place at every iteration — the GPU does not sit idle.
📄What the sources actually claimThe vLLM blog (2023): up to 24× higher throughput than HuggingFace Transformers and up to 3.5× than TGI. The paper (Kwon et al., SOSP 2023): 2–4× versus FasterTransformer and Orca at the same latency. The real gain depends on the load — measure it yourself. -
Installation: x86, ARM64 (GB10) and Docker
x86 + NVIDIA. The official path today uses
uv— it picks the right PyTorch for your driver by itself.bash · x86 or ARM64 Linux · install# Checks before installing python3 --version # 3.10–3.13 nvidia-smi # the GPU and driver are visible # Environment and install (recommended path in the vLLM docs) uv venv --python 3.12 --seed source .venv/bin/activate uv pip install vllm --torch-backend=auto python -c "import vllm; print(vllm.__version__)"ARM64 with GB10 (NVIDIA DGX Spark, ASUS Ascent GX10 and similar). Until recently this was the weak spot. As of 01.10.2026 PyPI has aarch64 vLLM packages, and NVIDIA maintains a ready image for these machines — the safest path, because it is built for the exact architecture.
bash · GB10 · NVIDIA's image# The tag changes — take the current one from NGC or NVIDIA's DGX Spark playbook export VLLM_TAG=26.05.post1-py3 # the playbook's tag as of 01.10.2026 docker pull nvcr.io/nvidia/vllm:${VLLM_TAG} docker run -d --gpus all -p 8000:8000 --name vllm-server \ nvcr.io/nvidia/vllm:${VLLM_TAG} \ vllm serve Qwen/Qwen2.5-7B-Instruct-AWQDocker, the general image.
vllm/vllm-openai:latestnow covers both amd64 and arm64.bash · Docker · the official vLLM imagedocker run --gpus all --ipc=host -p 8000:8000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --name vllm-server \ vllm/vllm-openai:latest \ --model Qwen/Qwen2.5-7B-Instruct-AWQ \ --max-model-len 16384 \ --served-model-name qwen-7b # Test right after start-up curl http://localhost:8000/v1/models⚠️Trap: gated modelsMeta's Llama models on Hugging Face are gated: you must accept the licence and pass a token (HF_TOKEN). Without it the download stops with a 401 error. Never put the token in scripts you share. -
First server — three ways of working
Way 1 — an OpenAI-compatible HTTP server. This is the production way: any application that talks to OpenAI talks to it too.
bash · vllm serve with the main settings# The model is AWQ — vLLM detects the quantization from its config vllm serve Qwen/Qwen2.5-7B-Instruct-AWQ \ --port 8000 \ --max-model-len 16384 \ --max-num-seqs 128 \ --gpu-memory-utilization 0.90 \ --enable-prefix-caching \ --served-model-name qwen-7b # --max-num-seqs how many sequences per iteration # --enable-prefix-caching reuses a shared leading text (the system prompt) # --served-model-name the name clients use to call the model # --trust-remote-code only for models with custom code, and only if you trust them # Ask with curl curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "qwen-7b", "messages": [ {"role": "system", "content": "You are the AI assistant of a small business."}, {"role": "user", "content": "Explain PagedAttention in 3 sentences."} ], "temperature": 0.2, "max_tokens": 300 }' # Streaming (for a chat UI): add "stream": trueWay 2 — Python for batch jobs. When you have a list of questions and no need for a server: hand them all over at once and vLLM schedules them itself.
python · batch_inference.pyfrom vllm import LLM, SamplingParams import time llm = LLM( model="Qwen/Qwen2.5-7B-Instruct-AWQ", max_model_len=8192, enable_prefix_caching=True, gpu_memory_utilization=0.90, # 10% headroom ) params = SamplingParams(temperature=0.2, top_p=0.9, max_tokens=500, repetition_penalty=1.05) PROMPTS = [ "Explain the difference between LoRA and full fine-tuning.", "Write an example RAG chain with LangChain.", "What is PagedAttention and why does it matter?", "How is cosine similarity computed?", ] t0 = time.time() outputs = llm.generate(PROMPTS, params) elapsed = time.time() - t0 tokens = sum(len(o.outputs[0].token_ids) for o in outputs) print(f"{len(PROMPTS)} requests · {tokens} tokens · {elapsed:.1f} s · {tokens/elapsed:.0f} tok/s") for p, o in zip(PROMPTS, outputs): print("\n?", p[:60]); print(o.outputs[0].text[:200])Way 3 — swapping Ollama out of LangChain. vLLM speaks the OpenAI API, so in the RAG code you change only the model's address. Embeddings can stay on Ollama.
python · LangChain: Ollama → vLLMfrom langchain_openai import ChatOpenAI # BEFORE: llm = ChatOllama(model="qwen2.5:7b", temperature=0.2) llm = ChatOpenAI( base_url="http://localhost:8000/v1", api_key="not-needed", # vLLM needs no key unless you set one model="qwen-7b", temperature=0.2, max_tokens=500, ) # chain = ({"context": retriever, "question": RunnablePassthrough()} # | prompt | llm | StrOutputParser()) ← the only change # Quick comparison: the same 10 requests to both servers import time, requests def bench(url, model, n=10): t0 = time.time() for _ in range(n): requests.post(url, json={"model": model, "max_tokens": 80, "messages": [{"role": "user", "content": "Explain RAG in 50 words."}]}) return time.time() - t0 print("Ollama:", bench("http://localhost:11434/v1/chat/completions", "qwen2.5:7b")) print("vLLM: ", bench("http://localhost:8000/v1/chat/completions", "qwen-7b"))⚠️These are not our measurementsThe old version of this lesson showed "42 s vs 14 s" for 10 requests — an illustration, not a result of our test. Sequential requests barely show vLLM's advantage; it shows with parallel requests. Run the comparison with several threads and record your own numbers. -
Quantization for production — AWQ, GPTQ, FP8
AWQ (Lin et al., MIT, 2023) is 4-bit quantization that finds the most important weights (by the real activations) and keeps them more precise. That is why it loses less quality than plain 4-bit quantization at the same memory.
Format Bits Memory for a 7B model When BF16 (no quantization) 16 ~14 GB for weights alone machines with plenty of memory (e.g. 128 GB unified) AWQ 4-bit 4 ~4–5 GB GPUs with little memory; good quality GPTQ 4-bit 4 ~4–5 GB alternative to AWQ FP8 8 ~7–8 GB newer GPUs with hardware FP8 support GGUF Q4_K_M 4 ~4–5 GB Ollama / llama.cpp only ⚠️No blanket quality percentagesWe removed the old version's "quality vs fp16" percentages and "1.2–1.5× faster" — we found no source confirming them across models. Memory in the table is approximate. Quality depends on the model and the task; compare on your own data.python · download a ready AWQ model and start# Look for ready AWQ versions: huggingface.co/models?search=awq # Official AWQ versions are published e.g. by Qwen; for Llama 3.1 — hugging-quants from huggingface_hub import snapshot_download path = snapshot_download("Qwen/Qwen2.5-7B-Instruct-AWQ", local_dir="./models/qwen2.5-7b-awq") from vllm import LLM llm = LLM(model=path, dtype="float16", # AWQ kernels run in fp16 max_model_len=32768, gpu_memory_utilization=0.85, enable_chunked_prefill=True) # steadier latency on long prompts # The same from the command line: # vllm serve ./models/qwen2.5-7b-awq --max-model-len 32768🔧Quantize it yourself?The AutoAWQ library is deprecated; the vLLM team has taken the functionality into llm-compressor. If there is no ready AWQ version of your model, quantize with llm-compressor — but look for a ready one first. -
GB10 with 128 GB unified memory — strategy
Machines such as DGX Spark and ASUS Ascent GX10 have 128 GB of memory shared by the CPU and the GPU. That allows large models without quantization, but the memory is slower than that of large data-centre GPUs. For a small team or a pilot course a well-tuned Ollama is often enough; vLLM (from NVIDIA's image) is the next step as concurrent users grow.
bash · Ollama for a few parallel requests (systemd)sudo mkdir -p /etc/systemd/system/ollama.service.d/ sudo tee /etc/systemd/system/ollama.service.d/override.conf >/dev/null <<'EOF' [Service] Environment="OLLAMA_NUM_PARALLEL=4" Environment="OLLAMA_MAX_LOADED_MODELS=2" Environment="OLLAMA_KEEP_ALIVE=30m" Environment="OLLAMA_FLASH_ATTENTION=1" EOF sudo systemctl daemon-reload && sudo systemctl restart ollama ollama ps # which models are loaded # OLLAMA_NUM_PARALLEL parallel requests per model (default 1) # memory grows with NUM_PARALLEL × context length # OLLAMA_FLASH_ATTENTION Ollama enables it automatically where it can; =1 forces itpython · check: 8 parallel requestsimport requests, concurrent.futures, time URL = "http://localhost:11434/v1/chat/completions" # or :8000 for vLLM MODEL = "qwen2.5:7b" def one(i): t0 = time.time() requests.post(URL, json={"model": MODEL, "max_tokens": 50, "messages": [{"role": "user", "content": f"Question {i}: what is AI?"}]}) return time.time() - t0 t0 = time.time() with concurrent.futures.ThreadPoolExecutor(max_workers=8) as ex: r = list(ex.map(one, range(8))) print(f"total {time.time()-t0:.1f}s · avg {sum(r)/len(r):.1f}s · sequential would be ~{sum(r):.1f}s")⛔Never put the model on the internet bareNeither Ollama nor vLLM asks for a password by default. Put a reverse proxy (e.g. nginx) in front with rate limiting and authentication — or keep them on the internal network only. How to assemble that is in the next part.🎯When to move from Ollama to vLLMStay on Ollama: development, testing, a small team, a few concurrent requests.
Move to vLLM: many concurrent users, need for FP8/AWQ and metrics, agreed response times. ⚠️ Measure the exact threshold on your own hardware — the old version claimed "8 users on a 70B model", which we have not verified.
04Check
1. A RAG system with 30 concurrent users answers slowly on Ollama with default settings. Which vLLM mechanism helps most?
2. Why does AWQ lose less quality than plain 4-bit quantization?
3. What does PagedAttention split the KV cache into?
4. You want vLLM on an ARM64 machine with GB10 (as of 01.10.2026). What is the safest path?
05What's next
06Sources
- vLLM — Quickstart — installing with uv,
vllm serve, the OpenAI-compatible server on port 8000. - vLLM — GPU installation — requirements (Linux, Python 3.10–3.13, compute capability ≥ 7.5), aarch64.
- vllm on PyPI — version 0.30.0 as of 01.10.2026, packages for x86_64 and aarch64.
- NVIDIA DGX Spark playbooks — vLLM — the
nvcr.io/nvidia/vllmimage for GB10. - NVIDIA NGC — vLLM container — current tags.
- Kwon et al. — Efficient Memory Management for LLM Serving with PagedAttention (SOSP 2023) — sections 3 and 4 are the core.
- The vLLM blog (2023) — 24× vs HF Transformers, waste under 4%.
- Lin et al. — AWQ: Activation-aware Weight Quantization.
- vLLM — quantization (AWQ, FP8 and more) · llm-compressor.
- Qwen2.5-7B-Instruct-AWQ · Llama 3.1 8B AWQ (hugging-quants) — ready AWQ models.
- Ollama FAQ —
OLLAMA_NUM_PARALLEL,OLLAMA_MAX_LOADED_MODELS, Flash Attention.