Which model is fastest and most accurate? Benchmarking local models
Do not trust other people's leaderboards — measure it yourself, on your machine, with your questions. You will write a small script that sends the same 10 questions to several models in Ollama, measures real speed from Ollama's own data and collects a quality grade. The result is a table you can use to choose a model for your OpenClaw agent.
eval_count and eval_duration fields, not from the wall-clock time of the request (which includes model loading); a warm-up call was added; the judge now grades all 10 answers (the old one graded only the first, and its prompt hard-coded “the capital of Bulgaria” for every question); a failed grade is reported as “not graded” instead of silently becoming a 3; the median replaces the mean; all answers are saved to a file for manual review; the Ollama address defaults to localhost and is changed with a variable instead of a hard-coded IP. Fact correction: the old text claimed DeepSeek-R1 does not support tool calling — today its page in the Ollama library shows “tools”. You check that yourself in step 7. Name correction: “Open WebUI” is a separate project and is not OpenClaw. We also added a pointer to ready-made tools (lm-evaluation-harness, promptfoo).
01What you'll learn
- How to measure a model's speed accurately: from the data Ollama returns, not with an outside clock.
- Why the first request is always slower and how one warm-up call fixes the comparison.
- How to send the same 10 questions to several models and keep every answer.
- How a “judge model” (LLM-as-a-judge) works, where it goes wrong and how to check it.
- How to choose a model for OpenClaw — and what to verify before giving it tools.
- When to move on to ready-made benchmark tools.
02Before you start
- Ollama 🔒 local is running and you have done the local-model lessons of the series (pulling a model,
ollama list). - Python 3 and
pip. The script uses two packages:requestsandtabulate. - Room and time to download 6 models plus one judge: each model is several gigabytes. Start with 2–3 if disk or bandwidth is tight.
- The machine should not be busy with anything else during the test — otherwise the numbers measure that too.
- Optional: OpenClaw already installed, so you can see which models it “sees” on your Ollama.
03Steps
-
What exactly we measure
For a request to
/api/generatewith"stream": false, Ollama returns a bill for its work at the end. These are the fields we need (durations are in nanoseconds):Field Meaning eval_counthow many tokens the model produced in the answer eval_durationtime spent on generation only load_durationtime to load the model into memory prompt_eval_count/prompt_eval_durationtokens and time for processing the question total_durationtotal time of the request Tokens per second =
eval_count÷eval_duration× 10⁹ — the formula from the Ollama documentation. Why not use a stopwatch? Because the total time includes loading the model and processing the question, and on the first request it is mostly loading. You would be comparing “how fast it loads”, not “how fast it thinks”.⚠️Models that “think”For reasoning models (for example DeepSeek-R1), reasoning tokens are counted ineval_counttoo. They produce many more tokens for the same question — tokens per second can look fine while the answer arrives slowly. So we look at total time as well. The request has athinkfield that turns the reasoning trace on or off; when it is on, the trace comes back in a separatethinkingfield, not inresponse. -
Pull the models
The sample list comes from the Ollama library, and the tags were checked on 01.10.2026. Small models are for speed, bigger ones for quality:
bashollama pull llama3.2:3b ollama pull phi4-mini:3.8b ollama pull gemma3:4b ollama pull mistral:7b ollama pull qwen2.5-coder:7b ollama pull deepseek-r1:8b ollama pull qwen2.5:7b # the judge - not in the test list ollama listThe tags
llama3.2:latestandllama3.2:3bare the same model; we write the explicit size so it is clear what was tested. Also note the Ollama version itself (ollama --version) — results depend on it. -
The script
Save it as
benchmark_models.py. Before the code, the decisions in it, each with its reason:- Warm-up — one “Hello” request to each model before measuring starts. Loading is then not counted in any question.
keep_alivekeeps the model in memory between requests. temperature: 0and a fixedseed— so repeated runs give similar answers.- Median — one unusually long answer must not spoil the comparison.
- A grade for every question, with a reference answer where a correct one exists (capital, arithmetic, translation, author).
Noneinstead of “3” when the judge returns no digit — we do not invent a grade.- Everything in a file —
benchmark_results.jsonkeeps every answer so you can read it.
python · benchmark_models.py#!/usr/bin/env python3 """Benchmark of local Ollama models: speed (from API fields) + quality (LLM-as-a-judge).""" import json, os, re, statistics, requests from tabulate import tabulate BASE = os.environ.get("OLLAMA_URL", "http://localhost:11434") MODELS = os.environ.get( "BENCH_MODELS", "llama3.2:3b,phi4-mini:3.8b,gemma3:4b,mistral:7b,qwen2.5-coder:7b,deepseek-r1:8b", ).split(",") JUDGE = os.environ.get("BENCH_JUDGE", "qwen2.5:7b") # the judge is NOT one of the tested models # (question, reference answer or None) QUESTIONS = [ ("What is the capital of Bulgaria?", "Sofia"), ("Explain what an AI agent is in plain language.", None), ("Write a Python function that checks whether a number is prime.", None), ("Calculate 15 * 38 + 127.", "697"), ("What are the benefits of running an LLM locally?", None), ("Translate to Bulgarian: 'The weather is wonderful today.'", "Днес времето е прекрасно."), ("Who is the author of 'War and Peace'?", "Leo Tolstoy"), ("Give three ideas for automating an accounting office with n8n.", None), ("Explain the difference between SQL and NoSQL.", None), ("Write a short poem about AI.", None), ] def generate(model, prompt, timeout=300): payload = {"model": model, "prompt": prompt, "stream": False, "keep_alive": "10m", "options": {"temperature": 0, "seed": 42}} r = requests.post(f"{BASE}/api/generate", json=payload, timeout=timeout) r.raise_for_status() d = r.json() gen_s = d.get("eval_duration", 0) / 1e9 # nanoseconds -> seconds tokens = d.get("eval_count", 0) return { "response": d.get("response", ""), "tokens": tokens, "load_s": d.get("load_duration", 0) / 1e9, "total_s": d.get("total_duration", 0) / 1e9, "tps": tokens / gen_s if gen_s > 0 else 0.0, # tokens/s as in the Ollama docs } def judge(question, candidate, reference): ref = f"Reference answer: {reference}\n" if reference else "" prompt = ( "You are a strict grader. Rate the answer from 1 to 5 " "(5 = correct, complete and clear; 1 = wrong or meaningless). Reply with a single digit.\n" f"Question: {question}\n{ref}Answer to grade: {candidate}\nScore:" ) try: text = generate(JUDGE, prompt, timeout=120)["response"] except requests.RequestException: return None m = re.search(r"[1-5]", text) return int(m.group()) if m else None # None = not graded (never invent a 3) def run(): rows, answers = [], {} for model in MODELS: print(f"Testing {model} ...") try: generate(model, "Hello") # warm-up: loads the model into memory except requests.RequestException as e: print(f" skipped ({e})") continue tps, times, tokens, scores = [], [], 0, [] answers[model] = [] for q, ref in QUESTIONS: try: r = generate(model, q) except requests.RequestException as e: print(f" error: {e}") continue tps.append(r["tps"]); times.append(r["total_s"]); tokens += r["tokens"] s = judge(q, r["response"], ref) if s is not None: scores.append(s) answers[model].append({"q": q, "a": r["response"], "score": s}) if not tps: continue rows.append([model, f"{statistics.median(times):.2f}", f"{statistics.median(tps):.1f}", tokens, f"{statistics.mean(scores):.2f}" if scores else "-", f"{len(scores)}/{len(QUESTIONS)}"]) headers = ["Model", "Median time (s)", "Median tok/s", "Tokens", "Mean score (1-5)", "Graded"] print("\n" + tabulate(rows, headers=headers, tablefmt="github")) with open("benchmark_results.json", "w", encoding="utf-8") as f: json.dump({"headers": headers, "rows": rows, "answers": answers}, f, ensure_ascii=False, indent=2) print("\nAnswers and scores are in benchmark_results.json - review them by hand.") if __name__ == "__main__": run()✅The Ollama addressBy default Ollama listens only on your own machine, on port 11434. If you run the script from another environment (for example a container or another machine), change the address with theOLLAMA_URLvariable and allow access on Ollama itself withOLLAMA_HOST— see the Ollama FAQ. Do not expose Ollama to the network unless you need to: it has no authentication. - Warm-up — one “Hello” request to each model before measuring starts. Loading is then not counted in any question.
-
Run it and read the result
bashpip install requests tabulate python3 benchmark_models.py # only two models and a different judge: BENCH_MODELS="llama3.2:3b,phi4-mini:3.8b" BENCH_JUDGE="qwen2.5:7b" python3 benchmark_models.pyThe result is a table like this (shape only — no numbers, the numbers are yours):
shape example · no figures| Model | Median time (s) | Median tok/s | Tokens | Mean score (1-5) | Graded | |---------------|-----------------|--------------|--------|------------------|---------| | <model A> | … | … | … | … | …/10 | | <model B> | … | … | … | … | …/10 |How to read it: Median tok/s is “how fast it writes”; Median time is “how long you wait for an answer”; Mean score is the judge's opinion; Graded shows for how many questions the judge returned a digit at all. If it is not 10/10, do not trust the mean score.
-
The ten questions and what each tests
# Question Tests 1 What is the capital of Bulgaria? Factual knowledge 2 Explain what an AI agent is in plain language. Comprehension and explanation 3 Write a Python function that checks whether a number is prime. Code, logic 4 Calculate 15 * 38 + 127. Mathematics, accuracy 5 What are the benefits of running an LLM locally? Analytical skills 6 Translate to Bulgarian: “The weather is wonderful today.” Language ability 7 Who is the author of “War and Peace”? Cultural knowledge 8 Give three ideas for automating an accounting office with n8n. Creativity, practicality 9 Explain the difference between SQL and NoSQL. Technical explanation 10 Write a short poem about AI. Creativity, style Ten questions are a probe, not an exam: enough to see big differences, not small ones. Replace them with questions from your own work — then the benchmark answers “which model is best for me”. The Bulgarian edition of this lesson tests in Bulgarian; the English script here uses English questions, so its numbers are not comparable with it.
-
The judge: what to trust
A “judge model” (LLM-as-a-judge) is a second model that receives the question, the answer and — when there is one — the correct answer, and returns a digit from 1 to 5. It is convenient because it is automatic, but it errs in known ways: it likes its own style, prefers longer answers and can be swayed by order. This is described in the study by Zheng et al. (2023) — see “Sources”.
- The judge is larger than or at least as capable as the tested models, and is not among them.
- Where a correct answer exists (questions 1, 4, 6, 7), pass it as the reference — the script does.
- Read by hand at least the code (question 3) and the creative answers (8 and 10) from
benchmark_results.json. Code gets run — that is the only way to know it works. - Do not compare grades between runs that used different judges.
⚠️The number is an opinionA difference of 0.3 points in the mean grade is noise. Look at big differences and at the examples you read by hand. -
Choose a model for OpenClaw
OpenClaw is a personal AI assistant with a Gateway; when you give it tools, the model must be able to call them (tool calling). So speed and quality are not enough — check this too:
- Tool support. Each model's page in the Ollama library has a capability badge. As of 01.10.2026:
llama3.2,phi4-mini,mistral,qwen2.5-coderanddeepseek-r1show “tools”;gemma3shows only “vision”. Badges change — check again. - The right address. OpenClaw talks to Ollama through its native
/api/chatinterface, not the OpenAI-compatible/v1— the documentation warns that/v1breaks tool calling. - Cold model. The first request to a model that is not loaded can time out — exactly why the benchmark has a warm-up. See the troubleshooting section of the OpenClaw documentation.
- What OpenClaw sees.
openclaw models list --provider ollamalists the models it discovers on your Ollama (the command comes from the documentation; we have not run it).
✅A rule for choosingFor quick chat and simple automations — the fastest model that keeps quality on your questions. For complex work — the slower one with better answers you read by hand. If actions have a real effect (files, messages, payments), add human approval regardless of the model. - Tool support. Each model's page in the Ollama library has a capability badge. As of 01.10.2026:
-
When the script is not enough
- lm-evaluation-harness — a framework for standard academic tests. Install with
pip install "lm_eval[api]"(current version 0.4.13, needs Python 3.10+); it works with models over an API. Connecting it to Ollama through its OpenAI-compatible address has not been verified by us ⚠️ — see its documentation. - promptfoo — a tool for comparing prompts and models with test cases and assertions; it has a ready Ollama provider (
ollama:chat:<model>). Current version 0.123 on npm. Good when you want tests to repeat automatically after every change.
An idea for later: run the same script periodically and keep the results with the date, hardware and Ollama version — that way you notice regressions when you switch a model or update.
- lm-evaluation-harness — a framework for standard academic tests. Install with
04Check
Checklist
ollama listshows all tested models and the judge; the judge is not in the test list.- The machine does nothing else heavy during the test.
- Speed is in tokens/s from
eval_countandeval_duration; every model has a warm-up. benchmark_results.jsonexists and you have read at least the code and the creative answers.- The “Graded” column is 10/10, or you know why not.
- The model chosen for OpenClaw supports tools (the badge in the library).
- You recorded the date, hardware, Ollama version and the model tags.
Quiz
1. How are tokens per second calculated according to the Ollama documentation?
2. Why do we do a warm-up before measuring?
3. Who can be the judge model?
4. Before giving a model to OpenClaw with tools, what else do you check besides speed and quality?
05What's next
06Sources
- Ollama: API documentation 🔒 local —
/api/generate, the timing fields and the tokens-per-second formula; also: thinking models and the FAQ (OLLAMA_HOST). - Ollama library — llama3.2, phi4-mini, gemma3, mistral, qwen2.5-coder, deepseek-r1, qwen2.5 — tags and capability badges.
- OpenClaw: Ollama provider — native interface, the
/v1warning, troubleshooting. - Zheng et al.: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — the biases of a judge model.
- lm-evaluation-harness · promptfoo: Ollama provider · PyPI · npm.