The same model, another machine: why the result differs
The same model name, the same question — and two machines give a different answer or a different speed. It is rarely a "defect". Most often it is a difference in the model tag, in the default context, in where the model is loaded, or in the settings. That is why a conclusion has to carry the name of the machine.
ollama ps. The three measured cases from our trials are kept as our observations; the GPU architecture is left as the cause we reached by elimination, not as a proven one. The specific machine names were removed — they are now "machine A" and "machine B".⚠️ Unverified: we did not run the code on live hardware in this revision — so the label is VERIFIED, not TESTED.
01What you will learn
- The documented reasons one model can behave differently on two machines.
- Why "qwen2.5:7b" with nothing added is not one unambiguous file — and what the tag says.
- How the default context changes with the VRAM itself — and how to pin it.
- The model-card rule: how to record a conclusion so the next person does not deploy it where it does not hold.
- How to check in about an hour whether a module ports.
02Before you start
- Two machines running Ollama 🔒 local — call them "machine A" and "machine B" — with different VRAM, or a CPU against a GPU.
- The same model pulled on both (see Lesson 2 for installation and picking a model).
- Command-line and API access on each machine (you know the addresses — here they are placeholders).
jqandcurlfor the comparison.
03Steps
-
Three cases from our trials
Two machines: an older consumer GPU with 12 GB (machine A) and a newer node with unified memory (machine B). The same models, the same material, the trials run separately on each. Three results diverged — each in a different way.
Case Machine A (older GPU) Machine B (newer node) embedding model for search 2 of 20 invalid vectors 0 of 20 vision model, 8B crashes the runner on every image does not crash (but the text is unusable) vision model, 2B — meter reading correct 8 times out of 8 wrong reading and wrong serial ℹ️Note the directionThe first two cases break on the older GPU, the third on the newer one. It is not "newer is better" and not "older is more stable" — the two are simply different machines.What was eliminated. The first hypothesis was the runtime version. The check refuted it: the machine with the defect ran the newer version, and the failing inputs were then re-run on the newer version — the defect stayed. Also eliminated: a different weights file (checked by digest), a different input (the same strings, verbatim) and different settings (temperature zero, identical options). What remained is the difference in GPU architecture — two generations apart, different kernels for the same operations. That is a cause we reached by elimination; we have not proven it directly.
⚠️The honest way to write it upThe first wording was "the model is defective, replace it". It fell as soon as the second machine's numbers arrived. The correct one is: "on this machine this model returns invalid vectors". The first is a claim about the model; the second is about the model + machine pair, and only the second is what was measured. The replacement choice, however, stayed the same, for a different reason: not because the first is "broken", but because the second gives a wider margin between the first and second search result. The grounds changed, the choice did not. That is normal and is written down exactly that way. -
"The same model" — but is it the same file
A model name without a tag is a reference, not one unambiguous file. In the Ollama library the same model is offered in several tags with different quantization (the tags page shows variants such as
q4_0,q8_0andfp16). Tighter quantization = a smaller file and slightly different numbers inside. If one machine pulls the "usual" tag and on the other you chose a different one, you are not comparing the same file.✅RuleRecord and compare name:tag, not just the name.ollama show <name:tag>shows details about the model (including the quantization — see also its library page). -
The default context depends on VRAM
This is the most underrated difference. According to the Ollama documentation the default context is 4k tokens under 24 GiB of VRAM, 32k between 24 and 48 GiB, and 256k above 48 GiB. So the same long input can be cut off on a machine with a smaller card, and fit whole on a larger one. The results look "different" while the model is the same.
Check what is loaded — the CONTEXT column of
ollama ps— and set the context explicitly on both machines before you compare.bash · where and with what context the model is loaded# after the first request; columns: PROCESSOR (GPU/CPU) and CONTEXT ollama ps # pin the context for the whole server (example) OLLAMA_CONTEXT_LENGTH=8192 ollama serve⚠️Trap: more context = more memoryMore context eats VRAM. Set as much as you need, and check withollama psthat the model still fits on the card. -
Where it is loaded: GPU, CPU or split
ollama psshows a PROCESSOR column:100% GPU,100% CPUor split (for example48%/52% CPU/GPU). When the model does not fit on the card, part of it goes to system memory — the answer arrives more slowly. Speed is a property of the pair model + machine + context.Ollama also supports different hardware and compute paths (NVIDIA, AMD, Apple, Vulkan — see the hardware document). Different paths may give slightly different numbers; ⚠️ we have not checked this ourselves, so take it as a possible cause, not a proven one.
-
Seed and temperature
The default temperature is 0.8 — the answer is "slightly random". A set seed (
seed) makes the same input give the same text under the same setup (so says the Modelfile reference). To compare two machines settemperature: 0and a fixedseed. Even then there is no guarantee of a match across different machines — if a difference remains, that is a finding, not an error in the test. -
Versions: server, driver, runtime
The Ollama version, the GPU driver and the compute path differ between machines. Do not assume the newer version is the cause (or the older one) — re-run the failing input on the other version and see whether the difference stays. Versions go into the card.
-
The model-card rule
If a conclusion holds for the pair model + machine, the record about the model must carry the machine. Otherwise the next person (or next session) will read "this model reads meters best" and deploy it where it does not.
ℹ️Not the same as the card on ollama.comThe model pages in the Ollama library are also called a "model card". Here the "model card" is your own record of what you checked and where.json · the minimum beside every result{ "machine": "<machine A or B>", "gpu": "<model and VRAM>", "server_version": "<Ollama version>", "model": "<name:tag>", "num_ctx": "<context you set>", "temperature": 0, "seed": "<integer>", "date": "<date>", "repeats": "<how many times>" }Three consequences: (1) results are not merged into one table — one file per machine; merging hides exactly the difference; (2) porting a module means repeating the check, not copying a file; (3) word conclusions as "on this machine, with these settings, this model returned…", not as "the model is defective".
In our observations (the three cases above) embedding and reading from an image diverged between machines more easily than plain text generation — ⚠️ this is our conclusion from few cases, not checked against the literature; so check those first.
-
How to check in about an hour
You need three things: the same input verbatim, the same settings, the two addresses. The addresses are placeholders — put in your own.
bash · one input, two machines, a comparison# request.json — the same file for both machines # {"model":"<name:tag>","prompt":"<the same text>","stream":false, # "options":{"num_ctx":4096,"temperature":0,"seed":42}} for M in A B; do case $M in A) ADDR="<machine-A-address>";; B) ADDR="<machine-B-address>";; esac curl -s "http://$ADDR:11434/api/generate" -d @request.json \ | jq -r '.response' > "out_$M.txt" curl -s "http://$ADDR:11434/api/version" # the version goes into the card done diff out_A.txt out_B.txtThree inputs are worth it: one document to extract from, one short text to embed, one image — if the module looks at images. If the output is identical, record that too — with the date and the versions; next time you repeat only what has changed.
04Check
1. What is Ollama's default context under 24 GiB of VRAM?
2. Why might "qwen2.5:7b" on two machines not be the same file?
3. How should the conclusion be worded?
4. Which of these does NOT belong in the model card?
05What's next
06Sources
- Ollama: Context length — the default context by VRAM (4k / 32k / 256k),
ollama ps. - Ollama: FAQ —
OLLAMA_CONTEXT_LENGTH,num_ctx, how to see GPU/CPU, concurrency. - Ollama: Modelfile —
temperature(default 0.8) andseed. - Ollama: Hardware support — supported hardware and compute paths.
- ollama.com library: the tags of one model 🌐 global — an example of different quantizations.
- Ollama v0.35.0 on GitHub — the current stable version, released on 28.09.2026 (checked 01.10.2026).
- KAGAMI's own earlier trials on two machines — the three cases in step 1; not repeated for this revision.