The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewVERIFIED 2026-10-01 · UPDATED 2026-10-01
IDENTITY
module
OpenClaw-10 · Automated model benchmarking (Ollama)
series
OpenClaw · lesson 10 (last of the series)
level
Intermediate
duration
1–2 h
prerequisites
Ollama running locally with several chat models pulled; Python 3 with pip
trust_label
VERIFIED 2026-10-01 (Ollama API docs, Ollama model library tags and capability badges, PyPI and npm versions) · UPDATED 2026-10-01 · NOT TESTED with a real model: the script was run only against a mock server that returns Ollama-shaped JSON
versions
requests 2.34 · tabulate 0.10 · lm-evaluation-harness 0.4.13 (Python ≥ 3.10) · promptfoo 0.123 · Ollama native API /api/generate
language
human view: bg · english edition: /en/academy/openclaw/Обучение 10 · Автоматизирано тестване на модели (benchmark).html
previous / next
Обучение 9 · Мониторинг и alerting с Prometheus + Grafana / series index
PURPOSE

Measure local Ollama models on your own hardware and your own questions: generation speed in tokens per second taken from the API timing fields (not from wall-clock time), cold-load time, and answer quality graded by a separate judge model plus a manual spot check. Use the result to choose a model for an OpenClaw agent. No ranking is published in this lesson: any figure depends on the machine, quantisation and prompts.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

OpenClaw series index (kagami.bg/academy/openclaw/) · quick experiment: kagami.bg/stalbata/

SOURCES
TAGS
ollamabenchmarkllm-as-a-judgetokens-per-secondopenclawlocal-llmpython
VERIFIED · 01.10.2026 UPDATED · 01.10.2026

Which model is fastest and most accurate? Benchmarking local models

Do not trust other people's leaderboards — measure it yourself, on your machine, with your questions. You will write a small script that sends the same 10 questions to several models in Ollama, measures real speed from Ollama's own data and collects a quality grade. The result is a table you can use to choose a model for your OpenClaw agent.

⏱ 1–2 h Intermediate OpenClaw · Lesson 10 Ollama · Python · LLM-as-a-judge
Ollama (the models and the API)🔒 local Python script🔒 local Judge model (a second local model)🔒 local lm-evaluation-harness · promptfoo (optional)🔒 local
🔄
UPDATED · 01.10.2026 — what changed
We removed every number from the old table. It showed “sample results based on a real test” (time, tokens per second, grades) with no source you could reproduce, so it is not honest to keep them as fact. Instead you get a script that produces your own numbers. The script was reworked: tokens per second now come from Ollama's eval_count and eval_duration fields, not from the wall-clock time of the request (which includes model loading); a warm-up call was added; the judge now grades all 10 answers (the old one graded only the first, and its prompt hard-coded “the capital of Bulgaria” for every question); a failed grade is reported as “not graded” instead of silently becoming a 3; the median replaces the mean; all answers are saved to a file for manual review; the Ollama address defaults to localhost and is changed with a variable instead of a hard-coded IP. Fact correction: the old text claimed DeepSeek-R1 does not support tool calling — today its page in the Ollama library shows “tools”. You check that yourself in step 7. Name correction: “Open WebUI” is a separate project and is not OpenClaw. We also added a pointer to ready-made tools (lm-evaluation-harness, promptfoo).
⚠️
What we have not run ourselves
The script was run only against a mock server that returns JSON shaped like Ollama's — so there is no “TESTED” label and the lesson contains no model results. The response format and the tokens-per-second formula were checked against the Ollama documentation; your own run will give the real numbers.

01What you'll learn

02Before you start

⛔
The judge is not one of the tested models
If a model grades its own answer, it will flatter itself. So the judge is a separate model that is not in the test list. Even then the grades are a guide, not a verdict — see step 6.

03Steps

  1. What exactly we measure

    For a request to /api/generate with "stream": false, Ollama returns a bill for its work at the end. These are the fields we need (durations are in nanoseconds):

    FieldMeaning
    eval_counthow many tokens the model produced in the answer
    eval_durationtime spent on generation only
    load_durationtime to load the model into memory
    prompt_eval_count / prompt_eval_durationtokens and time for processing the question
    total_durationtotal time of the request

    Tokens per second = eval_count ÷ eval_duration × 10⁹ — the formula from the Ollama documentation. Why not use a stopwatch? Because the total time includes loading the model and processing the question, and on the first request it is mostly loading. You would be comparing “how fast it loads”, not “how fast it thinks”.

    ⚠️
    Models that “think”
    For reasoning models (for example DeepSeek-R1), reasoning tokens are counted in eval_count too. They produce many more tokens for the same question — tokens per second can look fine while the answer arrives slowly. So we look at total time as well. The request has a think field that turns the reasoning trace on or off; when it is on, the trace comes back in a separate thinking field, not in response.
  2. Pull the models

    The sample list comes from the Ollama library, and the tags were checked on 01.10.2026. Small models are for speed, bigger ones for quality:

    bash
    ollama pull llama3.2:3b
    ollama pull phi4-mini:3.8b
    ollama pull gemma3:4b
    ollama pull mistral:7b
    ollama pull qwen2.5-coder:7b
    ollama pull deepseek-r1:8b
    ollama pull qwen2.5:7b      # the judge - not in the test list
    
    ollama list

    The tags llama3.2:latest and llama3.2:3b are the same model; we write the explicit size so it is clear what was tested. Also note the Ollama version itself (ollama --version) — results depend on it.

  3. The script

    Save it as benchmark_models.py. Before the code, the decisions in it, each with its reason:

    • Warm-up — one “Hello” request to each model before measuring starts. Loading is then not counted in any question. keep_alive keeps the model in memory between requests.
    • temperature: 0 and a fixed seed — so repeated runs give similar answers.
    • Median — one unusually long answer must not spoil the comparison.
    • A grade for every question, with a reference answer where a correct one exists (capital, arithmetic, translation, author).
    • None instead of “3” when the judge returns no digit — we do not invent a grade.
    • Everything in a file — benchmark_results.json keeps every answer so you can read it.
    python · benchmark_models.py
    #!/usr/bin/env python3
    """Benchmark of local Ollama models: speed (from API fields) + quality (LLM-as-a-judge)."""
    import json, os, re, statistics, requests
    from tabulate import tabulate
    
    BASE = os.environ.get("OLLAMA_URL", "http://localhost:11434")
    MODELS = os.environ.get(
        "BENCH_MODELS",
        "llama3.2:3b,phi4-mini:3.8b,gemma3:4b,mistral:7b,qwen2.5-coder:7b,deepseek-r1:8b",
    ).split(",")
    JUDGE = os.environ.get("BENCH_JUDGE", "qwen2.5:7b")  # the judge is NOT one of the tested models
    
    # (question, reference answer or None)
    QUESTIONS = [
        ("What is the capital of Bulgaria?", "Sofia"),
        ("Explain what an AI agent is in plain language.", None),
        ("Write a Python function that checks whether a number is prime.", None),
        ("Calculate 15 * 38 + 127.", "697"),
        ("What are the benefits of running an LLM locally?", None),
        ("Translate to Bulgarian: 'The weather is wonderful today.'", "Днес времето е прекрасно."),
        ("Who is the author of 'War and Peace'?", "Leo Tolstoy"),
        ("Give three ideas for automating an accounting office with n8n.", None),
        ("Explain the difference between SQL and NoSQL.", None),
        ("Write a short poem about AI.", None),
    ]
    
    def generate(model, prompt, timeout=300):
        payload = {"model": model, "prompt": prompt, "stream": False,
                   "keep_alive": "10m", "options": {"temperature": 0, "seed": 42}}
        r = requests.post(f"{BASE}/api/generate", json=payload, timeout=timeout)
        r.raise_for_status()
        d = r.json()
        gen_s = d.get("eval_duration", 0) / 1e9          # nanoseconds -> seconds
        tokens = d.get("eval_count", 0)
        return {
            "response": d.get("response", ""),
            "tokens": tokens,
            "load_s": d.get("load_duration", 0) / 1e9,
            "total_s": d.get("total_duration", 0) / 1e9,
            "tps": tokens / gen_s if gen_s > 0 else 0.0,  # tokens/s as in the Ollama docs
        }
    
    def judge(question, candidate, reference):
        ref = f"Reference answer: {reference}\n" if reference else ""
        prompt = (
            "You are a strict grader. Rate the answer from 1 to 5 "
            "(5 = correct, complete and clear; 1 = wrong or meaningless). Reply with a single digit.\n"
            f"Question: {question}\n{ref}Answer to grade: {candidate}\nScore:"
        )
        try:
            text = generate(JUDGE, prompt, timeout=120)["response"]
        except requests.RequestException:
            return None
        m = re.search(r"[1-5]", text)
        return int(m.group()) if m else None  # None = not graded (never invent a 3)
    
    def run():
        rows, answers = [], {}
        for model in MODELS:
            print(f"Testing {model} ...")
            try:
                generate(model, "Hello")  # warm-up: loads the model into memory
            except requests.RequestException as e:
                print(f"  skipped ({e})")
                continue
            tps, times, tokens, scores = [], [], 0, []
            answers[model] = []
            for q, ref in QUESTIONS:
                try:
                    r = generate(model, q)
                except requests.RequestException as e:
                    print(f"  error: {e}")
                    continue
                tps.append(r["tps"]); times.append(r["total_s"]); tokens += r["tokens"]
                s = judge(q, r["response"], ref)
                if s is not None:
                    scores.append(s)
                answers[model].append({"q": q, "a": r["response"], "score": s})
            if not tps:
                continue
            rows.append([model, f"{statistics.median(times):.2f}", f"{statistics.median(tps):.1f}",
                         tokens, f"{statistics.mean(scores):.2f}" if scores else "-",
                         f"{len(scores)}/{len(QUESTIONS)}"])
        headers = ["Model", "Median time (s)", "Median tok/s", "Tokens", "Mean score (1-5)", "Graded"]
        print("\n" + tabulate(rows, headers=headers, tablefmt="github"))
        with open("benchmark_results.json", "w", encoding="utf-8") as f:
            json.dump({"headers": headers, "rows": rows, "answers": answers}, f, ensure_ascii=False, indent=2)
        print("\nAnswers and scores are in benchmark_results.json - review them by hand.")
    
    if __name__ == "__main__":
        run()
    ✅
    The Ollama address
    By default Ollama listens only on your own machine, on port 11434. If you run the script from another environment (for example a container or another machine), change the address with the OLLAMA_URL variable and allow access on Ollama itself with OLLAMA_HOST — see the Ollama FAQ. Do not expose Ollama to the network unless you need to: it has no authentication.
  4. Run it and read the result

    bash
    pip install requests tabulate
    python3 benchmark_models.py
    
    # only two models and a different judge:
    BENCH_MODELS="llama3.2:3b,phi4-mini:3.8b" BENCH_JUDGE="qwen2.5:7b" python3 benchmark_models.py

    The result is a table like this (shape only — no numbers, the numbers are yours):

    shape example · no figures
    | Model         | Median time (s) | Median tok/s | Tokens | Mean score (1-5) | Graded  |
    |---------------|-----------------|--------------|--------|------------------|---------|
    | <model A>     | …               | …            | …      | …                | …/10    |
    | <model B>     | …               | …            | …      | …                | …/10    |

    How to read it: Median tok/s is “how fast it writes”; Median time is “how long you wait for an answer”; Mean score is the judge's opinion; Graded shows for how many questions the judge returned a digit at all. If it is not 10/10, do not trust the mean score.

  5. The ten questions and what each tests

    #QuestionTests
    1What is the capital of Bulgaria?Factual knowledge
    2Explain what an AI agent is in plain language.Comprehension and explanation
    3Write a Python function that checks whether a number is prime.Code, logic
    4Calculate 15 * 38 + 127.Mathematics, accuracy
    5What are the benefits of running an LLM locally?Analytical skills
    6Translate to Bulgarian: “The weather is wonderful today.”Language ability
    7Who is the author of “War and Peace”?Cultural knowledge
    8Give three ideas for automating an accounting office with n8n.Creativity, practicality
    9Explain the difference between SQL and NoSQL.Technical explanation
    10Write a short poem about AI.Creativity, style

    Ten questions are a probe, not an exam: enough to see big differences, not small ones. Replace them with questions from your own work — then the benchmark answers “which model is best for me”. The Bulgarian edition of this lesson tests in Bulgarian; the English script here uses English questions, so its numbers are not comparable with it.

  6. The judge: what to trust

    A “judge model” (LLM-as-a-judge) is a second model that receives the question, the answer and — when there is one — the correct answer, and returns a digit from 1 to 5. It is convenient because it is automatic, but it errs in known ways: it likes its own style, prefers longer answers and can be swayed by order. This is described in the study by Zheng et al. (2023) — see “Sources”.

    • The judge is larger than or at least as capable as the tested models, and is not among them.
    • Where a correct answer exists (questions 1, 4, 6, 7), pass it as the reference — the script does.
    • Read by hand at least the code (question 3) and the creative answers (8 and 10) from benchmark_results.json. Code gets run — that is the only way to know it works.
    • Do not compare grades between runs that used different judges.
    ⚠️
    The number is an opinion
    A difference of 0.3 points in the mean grade is noise. Look at big differences and at the examples you read by hand.
  7. Choose a model for OpenClaw

    OpenClaw is a personal AI assistant with a Gateway; when you give it tools, the model must be able to call them (tool calling). So speed and quality are not enough — check this too:

    • Tool support. Each model's page in the Ollama library has a capability badge. As of 01.10.2026: llama3.2, phi4-mini, mistral, qwen2.5-coder and deepseek-r1 show “tools”; gemma3 shows only “vision”. Badges change — check again.
    • The right address. OpenClaw talks to Ollama through its native /api/chat interface, not the OpenAI-compatible /v1 — the documentation warns that /v1 breaks tool calling.
    • Cold model. The first request to a model that is not loaded can time out — exactly why the benchmark has a warm-up. See the troubleshooting section of the OpenClaw documentation.
    • What OpenClaw sees. openclaw models list --provider ollama lists the models it discovers on your Ollama (the command comes from the documentation; we have not run it).
    ✅
    A rule for choosing
    For quick chat and simple automations — the fastest model that keeps quality on your questions. For complex work — the slower one with better answers you read by hand. If actions have a real effect (files, messages, payments), add human approval regardless of the model.
  8. When the script is not enough

    • lm-evaluation-harness — a framework for standard academic tests. Install with pip install "lm_eval[api]" (current version 0.4.13, needs Python 3.10+); it works with models over an API. Connecting it to Ollama through its OpenAI-compatible address has not been verified by us ⚠️ — see its documentation.
    • promptfoo — a tool for comparing prompts and models with test cases and assertions; it has a ready Ollama provider (ollama:chat:<model>). Current version 0.123 on npm. Good when you want tests to repeat automatically after every change.

    An idea for later: run the same script periodically and keep the results with the date, hardware and Ollama version — that way you notice regressions when you switch a model or update.

04Check

Checklist

Quiz

1. How are tokens per second calculated according to the Ollama documentation?

2. Why do we do a warm-up before measuring?

3. Who can be the judge model?

4. Before giving a model to OpenClaw with tools, what else do you check besides speed and quality?

05What's next

06Sources

  1. Ollama: API documentation 🔒 local — /api/generate, the timing fields and the tokens-per-second formula; also: thinking models and the FAQ (OLLAMA_HOST).
  2. Ollama library — llama3.2, phi4-mini, gemma3, mistral, qwen2.5-coder, deepseek-r1, qwen2.5 — tags and capability badges.
  3. OpenClaw: Ollama provider — native interface, the /v1 warning, troubleshooting.
  4. Zheng et al.: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — the biases of a judge model.
  5. lm-evaluation-harness · promptfoo: Ollama provider · PyPI · npm.