The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewUPDATED 2026-10-03
IDENTITY
module
GX10 · How to measure the speed of an AI model yourself
series
GX10 (local AI server class: NVIDIA GB10, e.g. ASUS Ascent GX10 / DGX Spark)
level
Intermediate
duration
~45 min
prerequisites
A working Ollama with at least one model (lessons 04-07 and 04-01), Python 3
trust_label
UPDATED 2026-10-03 (field names and formulas read from the Ollama API documentation, llama-bench options from the llama.cpp README, hardware figures from NVIDIA docs, all on that date) · NOT TESTED (no GB10 machine available; the script and commands were not run) · no speed values are given
versions
none pinned: record the runtime version yourself; example model llama3.1:8b (about 4.9 GB on its Ollama page)
language
human view: bg · english edition: /en/academy/gx10/ (same file name)
previous / next
ai_модели_сравнение_GX10.html / GX10 series index
PURPOSE

Teach a repeatable method for measuring the generation and prompt-processing speed of a local language model on a GB10-class machine, and for recording the conditions so the result can be reproduced. A previous table of speeds without evidence was removed on purpose; this module publishes no speed values.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

ai_модели_сравнение_GX10.html · Which AI models fit on a GX10 · 04-115_Model_Selection_Trade_offs.html · 04-219_Kakvo_Tezhi_na_GX10.html · series index: kagami.bg/academy/gx10/ · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
gx10nvidia-gb10benchmarkingtokens-per-secondollamallama-benchmethod
UPDATED · 03.10.2026

How to Measure the Speed of an AI Model on a GX10 Yourself

Speed numbers for models are easy to share and hard to check. This lesson gives you no borrowed numbers — it gives you a method: how to measure, on your own GX10, how many tokens per second a model writes, how to repeat the measurement, how to separate speed from quality, and how to record the conditions so the measurement can be repeated.

⏱ ~45 min Intermediate GX10 NVIDIA GB10 · 128 GB unified memory Measuring · Ollama · Python
Ollama (the model and API)🔒 local Python 3 (the script)🔒 local llama.cpp · llama-bench🔒 local
🔄
UPDATED · 03.10.2026 — what changed
The lesson has been rewritten. The old version was a table of test results for 12 models (speed, size and quality). We removed it: they were measurements from one machine without evidence (no command, versions, quantization, context), and such numbers depend on those conditions. We added: a method for measuring yourself, a script over the Ollama API with repeats and a median, a separate quality test with your own questions, a second tool (llama-bench) and a measurement-card template. The lesson is consistent with "Which AI Models Fit on a GX10", which also has no speeds.
⚠️
What we have not run ourselves
We had no GB10-class machine at hand. The script and commands are written from the Ollama and llama.cpp documentation, but we have not run them — which is why there is no "TESTED" or "VERIFIED" label. We give no speed value at all, because we have not measured one with this method. The output format of --verbose and llama-bench may differ by version.

01What you'll learn

02Before you start

QuantityWhere it comes from in the Ollama replyWhat it shows
Loadingload_duration (nanoseconds)How long loading the model into memory takes — present on a "cold" start.
Prompt processingprompt_eval_count and prompt_eval_durationHow fast the model "reads" the input (tokens per second).
Generationeval_count and eval_durationHow fast it writes the answer (tokens per second). This is usually what people call "speed".

The names follow the Ollama API documentation, read on 03.10.2026.

03Steps

  1. Why there is no table of numbers here

    The previous version of this lesson showed a table of speeds for a dozen models. We removed it: the numbers were measurements from one machine without evidence — no command, runtime version, quantization, context length or date. A number like that looks like a fact but cannot be checked or repeated. So this lesson now gives you the method: you measure on your own machine and write the conditions next to the number.

  2. Write down the conditions before you measure

    A number means something only together with its conditions. Before the first run, record:

    WhatWhy it matters
    Machine and system (class, memory, system version)Speed depends on the hardware and drivers.
    Runtime and version (for example Ollama)A new version can be faster or slower.
    The exact model tag and quantizationllama3.1:8b and another quantization of the same model are different weights. See ollama show <model> — not run by us.
    Context length and prompt lengthA longer input and a larger context slow things down.
    Options: temperature, seed, token limitDifferent options give different answers and different lengths.
    Other load on the machineParallel jobs share the unified memory and its bandwidth.
    DateThe number is true for the day it was measured.
  3. Quickest: --verbose

    The simplest measurement is to run the model with the --verbose flag. After the answer Ollama prints statistics. The format and row names depend on the version — read them from your own output.

    bash · not run
    docker compose exec ollama ollama run llama3.1:8b --verbose "Explain in three sentences what quantization is."

    This is a quick check, not a measurement: you have one run, and the first often includes loading the model. For a number you will write down, use the next step.

  4. Measure through the API with a script

    In the reply of /api/generate Ollama returns times in nanoseconds: load_duration, prompt_eval_count and prompt_eval_duration, eval_count and eval_duration. The Ollama documentation gives the formula: generation speed = eval_count / eval_duration × 109 tokens per second. The script below sends the request 6 times without streaming ("stream": false), with fixed options, and computes the median.

    python · bench.py · not run
    import json, os, statistics, sys, urllib.request
    
    URL = os.environ.get("OLLAMA_URL", "http://localhost:11434")
    MODEL = sys.argv[1]
    PROMPT = "Explain in three sentences what quantization of a language model is."
    RUNS = 6
    
    
    def one_run():
        body = json.dumps({
            "model": MODEL,
            "prompt": PROMPT,
            "stream": False,
            "options": {"temperature": 0, "seed": 1, "num_predict": 200},
        }).encode()
        req = urllib.request.Request(URL + "/api/generate", body,
                                     {"Content-Type": "application/json"})
        with urllib.request.urlopen(req, timeout=900) as r:
            d = json.load(r)
        gen = d["eval_count"] / d["eval_duration"] * 1e9
        pp = None
        if d.get("prompt_eval_duration"):
            pp = d["prompt_eval_count"] / d["prompt_eval_duration"] * 1e9
        return gen, pp, d["load_duration"] / 1e9
    
    
    results = [one_run() for _ in range(RUNS)]
    print("run 1 (warm-up, not counted): load %.1f s" % results[0][2])
    gens = [g for g, _, _ in results[1:]]
    print("generation tokens/s: median %.1f, min %.1f, max %.1f (n=%d)"
          % (statistics.median(gens), min(gens), max(gens), len(gens)))
    pps = [p for _, p, _ in results[1:] if p]
    if pps:
        print("prompt-processing tokens/s: median %.1f (n=%d)" % (statistics.median(pps), len(pps)))
    

    Run it like this (the model name is the argument):

    bash · not run
    python3 bench.py llama3.1:8b

    The output has three lines: the load time of the first (discarded) run, the median with minimum and maximum for generation and — if present — the median for prompt processing. The numbers will be yours; there are deliberately no sample values here.

    ⚠️
    How the script reaches Ollama
    The script reads the address from the OLLAMA_URL variable (default http://localhost:11434). An Ollama installed directly on the system listens on loopback and is reachable from there. In the stack from lesson 04-01 Ollama has no published port: for the measurement either run the script inside the containers' network, or publish the port temporarily to 127.0.0.1 only ("127.0.0.1:11434:11434") — never to 0.0.0.0 — and revert the setting after the test.
  5. Repeats, discarding the first run, the median

    The first run is "cold": the model is loaded into memory (see load_duration) and the number is not typical. That is why the script does not count it. From the rest take the median — a single random slow request does not move it — and write down the minimum and maximum. If the spread is large, something else is running on the machine or it has heated up: pause, let it cool and repeat. NVIDIA gives an ideal ambient operating temperature of 5 to 30 °C (as of 03.10.2026), so measure at normal room temperature with the vents unblocked.

  6. What else changes the number

    FactorHow it matters
    Model size and quantizationA bigger file means more data to read per token. In generation, memory bandwidth (273 GB/s for GB10 per NVIDIA) is usually the limit — which is why large models are slower.
    Dense versus mixture-of-experts (MoE)In MoE models only part of the parameters is used per token, so at the same file size the speed can differ. Check the model card.
    Context length and prompt lengthA longer context means more work for every new token and slower input processing.
    "Thinking" modelsPart of the tokens is a chain of thought. Check whether and how they are counted in your version, and compare only similar models and tasks.
    Other load and temperatureParallel jobs and overheating reduce speed.
    Runtime versionEvery new version can change the result — which is why you record it.
  7. Speed is not quality

    A fast model can answer badly. So quality is measured separately, with your own questions:

    • Prepare 10 questions from your own work with a short known correct answer (a number, a date, a category).
    • Run each question 3 times on each model and count the correct answers.
    • For text in Bulgarian: show two answers to a colleague without saying which model gave them and ask which is better (a "blind" comparison).
    • Also note how often the model "invents" — answers confidently and wrongly.

    A small test does not prove much, but it is more honest than somebody else's table — because it runs on your documents and questions.

  8. A second tool: llama-bench

    If you work with GGUF models through llama.cpp, there is a ready-made tool. It tests prompt processing (-p) and generation (-n) separately, repeats each test -r times (default 5) and gives an average with deviation. The documentation notes that the measurement does not include tokenization and sampling, so its numbers are not directly comparable with Ollama's.

    bash · not run
    llama-bench -m <path-to-model.gguf> -p 512 -n 128 -r 5 -o md
  9. The measurement card and the publishing rule

    Every number you record or share travels with a card. Here is a template:

    text · template
    Measurement card
    date:            <yyyy-mm-dd>
    machine:         <class, e.g. GB10 / 128 GB>   system: <version>
    runtime:         <name and version>            model tag: <e.g. llama3.1:8b>
    quantization:    <as shown by "ollama show">   context: <num_ctx>
    prompt:          <text or a link to it>        options: temperature 0, seed 1, num_predict 200
    method:          <command or script name>      runs: 6 (first discarded)
    result:          generation <median> tok/s (min <x> - max <y>)
    other load:      <nothing else running / what ran>
    notes:           <anything unusual>
    ✅
    Rule
    A number without a card is not published as a fact. With a card it is a measurement under these conditions — and call it that. Do not carry it over to another machine, version or quantization.

04Check

Quiz

1. Why does the script not count the first run?

2. How is generation speed calculated from the Ollama reply?

3. When can a speed number be published as a measurement?

4. Model A writes 80 tokens per second and model B 40. What follows about their quality?

05What's next

All lessons of the series are in the GX10 index (in the path above).

06Sources

  1. Ollama: API documentation — the fields load_duration, prompt_eval_*, eval_* and the tokens-per-second formula (read 03.10.2026).
  2. llama.cpp: llama-bench — parameters -p, -n, -r, -o and the note on tokenization (03.10.2026).
  3. NVIDIA DGX Spark: hardware overview — 273 GB/s memory bandwidth, ideal temperature 5–30 °C (03.10.2026).
  4. llama3.1 🔒 local — the size of the model used as an example.