How to Measure the Speed of an AI Model on a GX10 Yourself
Speed numbers for models are easy to share and hard to check. This lesson gives you no borrowed numbers — it gives you a method: how to measure, on your own GX10, how many tokens per second a model writes, how to repeat the measurement, how to separate speed from quality, and how to record the conditions so the measurement can be repeated.
--verbose and llama-bench may differ by version.01What you'll learn
- Why a number without conditions and without evidence is not a fact.
- Which quantities are measured: loading, prompt processing and generation (tokens per second).
- How to take the numbers from the Ollama reply and turn them into a speed.
- How to repeat the measurement, discard the first run and take the median.
- How to separate speed from quality and build a small test with your own questions.
- How to write a measurement card so that somebody else can repeat it.
02Before you start
- A working Ollama — see the lessons Open WebUI and Ollama and n8n on GX10.
- At least one downloaded model. As an example we use
llama3.1:8b(about 4.9 GB per the model page as of 03.10.2026) — see "Which AI Models Fit on a GX10". - Python 3 for the script; check with
python3 --version. - A calm machine: do not run other heavy jobs while you measure.
| Quantity | Where it comes from in the Ollama reply | What it shows |
|---|---|---|
| Loading | load_duration (nanoseconds) | How long loading the model into memory takes — present on a "cold" start. |
| Prompt processing | prompt_eval_count and prompt_eval_duration | How fast the model "reads" the input (tokens per second). |
| Generation | eval_count and eval_duration | How fast it writes the answer (tokens per second). This is usually what people call "speed". |
The names follow the Ollama API documentation, read on 03.10.2026.
03Steps
-
Why there is no table of numbers here
The previous version of this lesson showed a table of speeds for a dozen models. We removed it: the numbers were measurements from one machine without evidence — no command, runtime version, quantization, context length or date. A number like that looks like a fact but cannot be checked or repeated. So this lesson now gives you the method: you measure on your own machine and write the conditions next to the number.
-
Write down the conditions before you measure
A number means something only together with its conditions. Before the first run, record:
What Why it matters Machine and system (class, memory, system version) Speed depends on the hardware and drivers. Runtime and version (for example Ollama) A new version can be faster or slower. The exact model tag and quantization llama3.1:8band another quantization of the same model are different weights. Seeollama show <model>— not run by us.Context length and prompt length A longer input and a larger context slow things down. Options: temperature, seed, token limit Different options give different answers and different lengths. Other load on the machine Parallel jobs share the unified memory and its bandwidth. Date The number is true for the day it was measured. -
Quickest:
--verboseThe simplest measurement is to run the model with the
--verboseflag. After the answer Ollama prints statistics. The format and row names depend on the version — read them from your own output.bash · not rundocker compose exec ollama ollama run llama3.1:8b --verbose "Explain in three sentences what quantization is."This is a quick check, not a measurement: you have one run, and the first often includes loading the model. For a number you will write down, use the next step.
-
Measure through the API with a script
In the reply of
/api/generateOllama returns times in nanoseconds:load_duration,prompt_eval_countandprompt_eval_duration,eval_countandeval_duration. The Ollama documentation gives the formula: generation speed =eval_count / eval_duration × 109tokens per second. The script below sends the request 6 times without streaming ("stream": false), with fixed options, and computes the median.python · bench.py · not runimport json, os, statistics, sys, urllib.request URL = os.environ.get("OLLAMA_URL", "http://localhost:11434") MODEL = sys.argv[1] PROMPT = "Explain in three sentences what quantization of a language model is." RUNS = 6 def one_run(): body = json.dumps({ "model": MODEL, "prompt": PROMPT, "stream": False, "options": {"temperature": 0, "seed": 1, "num_predict": 200}, }).encode() req = urllib.request.Request(URL + "/api/generate", body, {"Content-Type": "application/json"}) with urllib.request.urlopen(req, timeout=900) as r: d = json.load(r) gen = d["eval_count"] / d["eval_duration"] * 1e9 pp = None if d.get("prompt_eval_duration"): pp = d["prompt_eval_count"] / d["prompt_eval_duration"] * 1e9 return gen, pp, d["load_duration"] / 1e9 results = [one_run() for _ in range(RUNS)] print("run 1 (warm-up, not counted): load %.1f s" % results[0][2]) gens = [g for g, _, _ in results[1:]] print("generation tokens/s: median %.1f, min %.1f, max %.1f (n=%d)" % (statistics.median(gens), min(gens), max(gens), len(gens))) pps = [p for _, p, _ in results[1:] if p] if pps: print("prompt-processing tokens/s: median %.1f (n=%d)" % (statistics.median(pps), len(pps)))Run it like this (the model name is the argument):
bash · not runpython3 bench.py llama3.1:8bThe output has three lines: the load time of the first (discarded) run, the median with minimum and maximum for generation and — if present — the median for prompt processing. The numbers will be yours; there are deliberately no sample values here.
⚠️How the script reaches OllamaThe script reads the address from theOLLAMA_URLvariable (defaulthttp://localhost:11434). An Ollama installed directly on the system listens on loopback and is reachable from there. In the stack from lesson 04-01 Ollama has no published port: for the measurement either run the script inside the containers' network, or publish the port temporarily to127.0.0.1only ("127.0.0.1:11434:11434") — never to0.0.0.0— and revert the setting after the test. -
Repeats, discarding the first run, the median
The first run is "cold": the model is loaded into memory (see
load_duration) and the number is not typical. That is why the script does not count it. From the rest take the median — a single random slow request does not move it — and write down the minimum and maximum. If the spread is large, something else is running on the machine or it has heated up: pause, let it cool and repeat. NVIDIA gives an ideal ambient operating temperature of 5 to 30 °C (as of 03.10.2026), so measure at normal room temperature with the vents unblocked. -
What else changes the number
Factor How it matters Model size and quantization A bigger file means more data to read per token. In generation, memory bandwidth (273 GB/s for GB10 per NVIDIA) is usually the limit — which is why large models are slower. Dense versus mixture-of-experts (MoE) In MoE models only part of the parameters is used per token, so at the same file size the speed can differ. Check the model card. Context length and prompt length A longer context means more work for every new token and slower input processing. "Thinking" models Part of the tokens is a chain of thought. Check whether and how they are counted in your version, and compare only similar models and tasks. Other load and temperature Parallel jobs and overheating reduce speed. Runtime version Every new version can change the result — which is why you record it. -
Speed is not quality
A fast model can answer badly. So quality is measured separately, with your own questions:
- Prepare 10 questions from your own work with a short known correct answer (a number, a date, a category).
- Run each question 3 times on each model and count the correct answers.
- For text in Bulgarian: show two answers to a colleague without saying which model gave them and ask which is better (a "blind" comparison).
- Also note how often the model "invents" — answers confidently and wrongly.
A small test does not prove much, but it is more honest than somebody else's table — because it runs on your documents and questions.
-
A second tool:
llama-benchIf you work with GGUF models through llama.cpp, there is a ready-made tool. It tests prompt processing (
-p) and generation (-n) separately, repeats each test-rtimes (default 5) and gives an average with deviation. The documentation notes that the measurement does not include tokenization and sampling, so its numbers are not directly comparable with Ollama's.bash · not runllama-bench -m <path-to-model.gguf> -p 512 -n 128 -r 5 -o md -
The measurement card and the publishing rule
Every number you record or share travels with a card. Here is a template:
text · templateMeasurement card date: <yyyy-mm-dd> machine: <class, e.g. GB10 / 128 GB> system: <version> runtime: <name and version> model tag: <e.g. llama3.1:8b> quantization: <as shown by "ollama show"> context: <num_ctx> prompt: <text or a link to it> options: temperature 0, seed 1, num_predict 200 method: <command or script name> runs: 6 (first discarded) result: generation <median> tok/s (min <x> - max <y>) other load: <nothing else running / what ran> notes: <anything unusual>✅RuleA number without a card is not published as a fact. With a card it is a measurement under these conditions — and call it that. Do not carry it over to another machine, version or quantization.
04Check
- You can explain why your number is worthless without its conditions.
- You have a script that sends the request several times and computes a median.
- The first run is discarded and the minimum and maximum are recorded.
- You have a separate small quality test with your own questions.
- Every number has a measurement card.
Quiz
1. Why does the script not count the first run?
2. How is generation speed calculated from the Ollama reply?
3. When can a speed number be published as a measurement?
4. Model A writes 80 tokens per second and model B 40. What follows about their quality?
05What's next
All lessons of the series are in the GX10 index (in the path above).
06Sources
- Ollama: API documentation — the fields
load_duration,prompt_eval_*,eval_*and the tokens-per-second formula (read 03.10.2026). - llama.cpp: llama-bench — parameters
-p,-n,-r,-oand the note on tokenization (03.10.2026). - NVIDIA DGX Spark: hardware overview — 273 GB/s memory bandwidth, ideal temperature 5–30 °C (03.10.2026).
- llama3.1 🔒 local — the size of the model used as an example.