The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewVERIFIED 2026-10-01 · UPDATED 2026-10-01
IDENTITY
module
KW I25 · The same model, another machine
series
KAGAMI Way · Track I — Infrastructure
level
Intermediate
duration
~15 min
trust_label
VERIFIED 2026-10-01 (against official Ollama documentation) · UPDATED 2026-10-01 · not re-run on live hardware in this revision
language
human view: en · bulgarian edition: /academy/moduli/KW_I25_Same_Model_Other_Machine.html
previous
KW_I24_Numbers_By_Code.html
next
KW_I26_Model_Verdict_Matrix.html · Which model for what: a reference matrix
PURPOSE

Explain why the same model name can give a different result or speed on two machines, list the documented causes, and give an operational rule: a conclusion about a model is a conclusion about the pair (model + machine), so the record of the conclusion carries the machine and the settings.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

KW_I26_Model_Verdict_Matrix.html · Which model for what: a reference matrix · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
ollamaportabilityquantizationcontext-lengthreproducibilitymodel-cardgpu
VERIFIED · 01.10.2026 UPDATED · 01.10.2026

The same model, another machine: why the result differs

The same model name, the same question — and two machines give a different answer or a different speed. It is rarely a "defect". Most often it is a difference in the model tag, in the default context, in where the model is loaded, or in the settings. That is why a conclusion has to carry the name of the machine.

⏱ ~15 min Intermediate Local AI portability · reproducibility · model card
Ollama (the server)🔒 local Downloaded models🔒 local The ollama.com library (download only)🌐 global
🔄
UPDATED · 01.10.2026 — what changed
The lesson was checked against the official Ollama documentation (default context, settings, supported hardware, library tags) and with the release notes on GitHub (current stable 0.35.0, released on 28.09.2026). Added: the default context depends on VRAM (under 24 GiB — 4k), the quantization tag, the check with ollama ps. The three measured cases from our trials are kept as our observations; the GPU architecture is left as the cause we reached by elimination, not as a proven one. The specific machine names were removed — they are now "machine A" and "machine B".
⚠️ Unverified: we did not run the code on live hardware in this revision — so the label is VERIFIED, not TESTED.

01What you will learn

02Before you start

⚠️
What is NOT proven here
The three cases in step 1 come from our earlier trials and were not repeated for this revision. They show that there is a divergence, but do not prove why. The Ollama documentation does not promise bit-identical output on different hardware; we have not checked that ourselves.

03Steps

  1. Three cases from our trials

    Two machines: an older consumer GPU with 12 GB (machine A) and a newer node with unified memory (machine B). The same models, the same material, the trials run separately on each. Three results diverged — each in a different way.

    CaseMachine A (older GPU)Machine B (newer node)
    embedding model for search2 of 20 invalid vectors0 of 20
    vision model, 8Bcrashes the runner on every imagedoes not crash (but the text is unusable)
    vision model, 2B — meter readingcorrect 8 times out of 8wrong reading and wrong serial
    ℹ️
    Note the direction
    The first two cases break on the older GPU, the third on the newer one. It is not "newer is better" and not "older is more stable" — the two are simply different machines.

    What was eliminated. The first hypothesis was the runtime version. The check refuted it: the machine with the defect ran the newer version, and the failing inputs were then re-run on the newer version — the defect stayed. Also eliminated: a different weights file (checked by digest), a different input (the same strings, verbatim) and different settings (temperature zero, identical options). What remained is the difference in GPU architecture — two generations apart, different kernels for the same operations. That is a cause we reached by elimination; we have not proven it directly.

    ⚠️
    The honest way to write it up
    The first wording was "the model is defective, replace it". It fell as soon as the second machine's numbers arrived. The correct one is: "on this machine this model returns invalid vectors". The first is a claim about the model; the second is about the model + machine pair, and only the second is what was measured. The replacement choice, however, stayed the same, for a different reason: not because the first is "broken", but because the second gives a wider margin between the first and second search result. The grounds changed, the choice did not. That is normal and is written down exactly that way.
  2. "The same model" — but is it the same file

    A model name without a tag is a reference, not one unambiguous file. In the Ollama library the same model is offered in several tags with different quantization (the tags page shows variants such as q4_0, q8_0 and fp16). Tighter quantization = a smaller file and slightly different numbers inside. If one machine pulls the "usual" tag and on the other you chose a different one, you are not comparing the same file.

    ✅
    Rule
    Record and compare name:tag, not just the name. ollama show <name:tag> shows details about the model (including the quantization — see also its library page).
  3. The default context depends on VRAM

    This is the most underrated difference. According to the Ollama documentation the default context is 4k tokens under 24 GiB of VRAM, 32k between 24 and 48 GiB, and 256k above 48 GiB. So the same long input can be cut off on a machine with a smaller card, and fit whole on a larger one. The results look "different" while the model is the same.

    Check what is loaded — the CONTEXT column of ollama ps — and set the context explicitly on both machines before you compare.

    bash · where and with what context the model is loaded
    # after the first request; columns: PROCESSOR (GPU/CPU) and CONTEXT
    ollama ps
    
    # pin the context for the whole server (example)
    OLLAMA_CONTEXT_LENGTH=8192 ollama serve
    ⚠️
    Trap: more context = more memory
    More context eats VRAM. Set as much as you need, and check with ollama ps that the model still fits on the card.
  4. Where it is loaded: GPU, CPU or split

    ollama ps shows a PROCESSOR column: 100% GPU, 100% CPU or split (for example 48%/52% CPU/GPU). When the model does not fit on the card, part of it goes to system memory — the answer arrives more slowly. Speed is a property of the pair model + machine + context.

    Ollama also supports different hardware and compute paths (NVIDIA, AMD, Apple, Vulkan — see the hardware document). Different paths may give slightly different numbers; ⚠️ we have not checked this ourselves, so take it as a possible cause, not a proven one.

  5. Seed and temperature

    The default temperature is 0.8 — the answer is "slightly random". A set seed (seed) makes the same input give the same text under the same setup (so says the Modelfile reference). To compare two machines set temperature: 0 and a fixed seed. Even then there is no guarantee of a match across different machines — if a difference remains, that is a finding, not an error in the test.

  6. Versions: server, driver, runtime

    The Ollama version, the GPU driver and the compute path differ between machines. Do not assume the newer version is the cause (or the older one) — re-run the failing input on the other version and see whether the difference stays. Versions go into the card.

  7. The model-card rule

    If a conclusion holds for the pair model + machine, the record about the model must carry the machine. Otherwise the next person (or next session) will read "this model reads meters best" and deploy it where it does not.

    ℹ️
    Not the same as the card on ollama.com
    The model pages in the Ollama library are also called a "model card". Here the "model card" is your own record of what you checked and where.
    json · the minimum beside every result
    {
      "machine": "<machine A or B>",
      "gpu": "<model and VRAM>",
      "server_version": "<Ollama version>",
      "model": "<name:tag>",
      "num_ctx": "<context you set>",
      "temperature": 0,
      "seed": "<integer>",
      "date": "<date>",
      "repeats": "<how many times>"
    }

    Three consequences: (1) results are not merged into one table — one file per machine; merging hides exactly the difference; (2) porting a module means repeating the check, not copying a file; (3) word conclusions as "on this machine, with these settings, this model returned…", not as "the model is defective".

    In our observations (the three cases above) embedding and reading from an image diverged between machines more easily than plain text generation — ⚠️ this is our conclusion from few cases, not checked against the literature; so check those first.

  8. How to check in about an hour

    You need three things: the same input verbatim, the same settings, the two addresses. The addresses are placeholders — put in your own.

    bash · one input, two machines, a comparison
    # request.json — the same file for both machines
    # {"model":"<name:tag>","prompt":"<the same text>","stream":false,
    #  "options":{"num_ctx":4096,"temperature":0,"seed":42}}
    for M in A B; do
      case $M in A) ADDR="<machine-A-address>";; B) ADDR="<machine-B-address>";; esac
      curl -s "http://$ADDR:11434/api/generate" -d @request.json \
        | jq -r '.response' > "out_$M.txt"
      curl -s "http://$ADDR:11434/api/version"   # the version goes into the card
    done
    diff out_A.txt out_B.txt

    Three inputs are worth it: one document to extract from, one short text to embed, one image — if the module looks at images. If the output is identical, record that too — with the date and the versions; next time you repeat only what has changed.

04Check

1. What is Ollama's default context under 24 GiB of VRAM?

2. Why might "qwen2.5:7b" on two machines not be the same file?

3. How should the conclusion be worded?

4. Which of these does NOT belong in the model card?

05What's next

06Sources

  1. Ollama: Context length — the default context by VRAM (4k / 32k / 256k), ollama ps.
  2. Ollama: FAQ — OLLAMA_CONTEXT_LENGTH, num_ctx, how to see GPU/CPU, concurrency.
  3. Ollama: Modelfile — temperature (default 0.8) and seed.
  4. Ollama: Hardware support — supported hardware and compute paths.
  5. ollama.com library: the tags of one model 🌐 global — an example of different quantizations.
  6. Ollama v0.35.0 on GitHub — the current stable version, released on 28.09.2026 (checked 01.10.2026).
  7. KAGAMI's own earlier trials on two machines — the three cases in step 1; not repeated for this revision.