The KAGAMI mark КАГАМИ
kagami.bg/en/blog/2026-08-kakvo-tezhi-na-128gb.html · article · machine-readable viewPUBLISHED 2026-08-16 · VERIFIED 2026-10-03 · UPDATED 2026-10-03
IDENTITY
title
What actually fits in 128 GB — the arithmetic before you download a model
author
KAGAMI Ltd. (КАГАМИ ЕООД), Varna · organisation byline, no individual author named
language
English edition (human view and this agent view) · Bulgarian original: https://kagami.bg/blog/2026-08-kakvo-tezhi-na-128gb.html
topic
AI hardware · memory sizing of local LLMs on 128 GB unified memory (ASUS Ascent GX10 / NVIDIA DGX Spark, GB10)
audience
Small companies and technical staff in Bulgaria deciding which open models can run locally
status
General information. Memory estimates only; no latency or throughput measurements are published on this page.
SUMMARY

Whether a model fits in the 128 GB unified memory of a GX10 / DGX Spark is arithmetic: weights (file size at the chosen quantization) + KV cache (grows linearly with context and depends strongly on architecture) + roughly 15% overhead, against a practical working ceiling of about 110 GB (KAGAMI's own rule of thumb, not a vendor figure). Mixture-of-Experts models reduce compute per token, not storage: all experts must be resident, so memory is counted on total parameters and speed on active parameters. Checked against primary sources on 2026-10-03: Llama 3.3 70B is 43 GB at Q4_K_M and 75 GB at Q8_0; gpt-oss-120b is 65 GB; Nemotron 3.5 Lightning (30B total, 3B active) is 25 GB at Q4_K_M; Muse Glimmer (30B dense) is 18 GB at Q4_K_M; DeepSeek-V4-Flash (284B total, 13B active) is about 160 GB of weights, DeepSeek-V4.1-Flash about 510 GB, Kimi K2.6 about 595 GB and Kimi K3 about 1.56 TB, so none of these fit. The original (2026-08-16) KV-cache figures (about 15 GB at 128K) were not model-independent: for Llama 3.3 70B with an FP16 cache the same context needs about 43 GB (own calculation).

KEY CLAIMS (dated)
WHAT WAS UPDATED (2026-10-03)
SOURCES
AGENT INSTRUCTIONS
TAGS
gx10nvidia-dgx-sparkunified-memoryquantizationkv-cachemixture-of-expertslocal-llmollama
VERIFIED · 03.10.2026 UPDATED · 03.10.2026

AI · Hardware

What really fits in 128 GB: the sum to do before you download the model

The question asked most often about machines of the DGX Spark class is one: "will this model fit?" The answer is thirty seconds of arithmetic, as long as you know which numbers to add up and where to get them. Here is the sum and a table of real models, checked against the primary sources on 3 October 2026.

AI 128 GB unified memory · quantisation · KV cache · MoE Technical
UPDATED · 03.10.2026 · WHAT WAS UPDATED

01 · WEIGHTSWeights: the main part

Every parameter takes up space, and exactly how much depends on the precision in which you store it. The memory of the GX10 is 128 GB unified: the processor and the graphics chip share the same 128 GB. [1][2] A rule of thumb (bytes per parameter):

PrecisionBytes / parameter30B model70B model
BF16 / FP16260 GB140 GB
8 bits (FP8, Q8_0)about 1about 30 GBabout 70 GB
4 bits (Q4, NVFP4)about 0.5about 15 GBabout 35 GB
⚠
A guide, not a file size
This is a rough calculation. Real files are larger, because some of the layers stay in higher precision. Llama 3.3 70B at Q4_K_M is not 35 but 43 GB; at Q8_0 it is 75 GB, and at FP16 141 GB. [3] Before you download, look at the size of the specific file.

The first conclusion stands: a 70B model in full precision does not fit in 128 GB (141 GB). At 8 bits it fits, but little room is left. At 4 bits it fits comfortably, and you pay in quality.

02 · KV CACHEThe KV cache: the term everyone forgets

The weights are only one part. While running, the model keeps keys and values (the KV cache) for every token of the context. The cache grows linearly with the length of the context and with a long context it can rival the weights themselves in size.

Exactly how much it grows depends on the architecture of the model; the first version of the article gave one general coefficient, and it is not right for all models. The formula for a model with classic attention is: 2 (keys and values) × number of layers × number of KV heads × head size × bytes per value. For Llama 3.3 70B the configuration is 80 layers, 8 KV heads, head size 128 [23]. With the cache in FP16 (2 bytes) this gives:

ContextKV cache of Llama 3.3 70BCalculation
8K (8,192 tokens)about 2.7 GB327,680 bytes × 8,192
32K (32,768 tokens)about 10.7 GB327,680 bytes × 32,768
128K (131,072 tokens)about 43 GB327,680 bytes × 131,072

The table is our calculation from the published configuration, not a measurement. 327,680 = 2 × 80 × 8 × 128 × 2 bytes per token.

At a 128K context the cache of this model is as large as its own weights at Q4_K_M (43 GB). This is where most planning fails. A model that loads without a problem collapses at the third long document, because the sum was done only for the weights.

The new architectures cut this cost sharply. For example, DeepSeek-V4.1-Flash states a global KV cache of about 890 bytes per token, about a quarter of the previous V4-Flash. [13][14] So do not carry the numbers over from one model to another: take the cache parameters from the model card.

03 · OVERHEADOverhead and the working ceiling

Activations, memory fragmentation, the runtime framework itself. Allow about 15% on top. Not because it is exactly that much, but because filling memory to the last byte causes failures that you then hunt for hours.

§
Our practical rule: a ceiling of about 110 GB, not 128
The unified memory is shared with the operating system and everything else on the machine. We use the 110 GB ceiling and the 15% overhead as our own rule; they are not manufacturer data. The 18 GB difference is not a loss: it is the reason the system does not fall over.

04 · THE TABLEWhat weighs what, and what fits

The sizes are those of the weights file according to the Ollama library, unless stated otherwise; "active" means the parameters that work for a single token. The "Fits?" column follows our rule (weights plus 15% overhead, a ceiling of about 110 GB, without the KV cache unless stated).

ModelParameters (total / active)VariantWeightsFits?
Nemotron 3.5 Lightning30B MoE / 3BQ4_K_M · Q8_0 · BF1625 · 35 · 66 GB [6][7]yes, comfortably; there is also an official NVFP4 of about 21.6 GB [8]
Muse Glimmer (Meta)about 29.6B, denseQ4_K_M · Q8_0 · BF1618 · 31 · 57 GB [9][10]yes, comfortably
Llama 3.3 70B70B, denseQ4_K_M43 GB [3]yes; with a 128K context about 99 GB (our calculation)
Llama 3.3 70B70B, denseQ8_075 GB [3]on the edge: up to a 32K context yes (about 99 GB), with 128K no
gpt-oss 120B117B MoE / 5.1BMXFP4 (original)65 GB [4][5]yes
DeepSeek V4 Flash284B MoE / 13BFP4 + FP8 mixedabout 160 GB [11]no
Kimi K2.61.04T MoE4-bit weightsabout 595 GB [15][16]no
Kimi K3about 2.8T MoEMXFP4about 1.56 TB [17][18]no: a cluster is needed

The sizes in GB are as Ollama shows them; for Hugging Face they are our sum of the sizes of the weight files in the repository (decimal GB), so they may differ by a few percent. According to Ollama, Kimi K2.6 and K3 are available only as cloud models (":cloud"); there is no local file to download there. [15][17]

ℹ
What changed compared with the first version
The biggest correction is for DeepSeek: V4 Flash has 13 billion active parameters, not "about 22". [11] Its entry in Ollama was withdrawn on 27.08.2026 [12], and the newer V4.1 Flash has 552 billion parameters in its backbone and weighs about 510 GB, even further from 128 GB. [13][14]

05 · THE NEW ONESThe new families that do fit

As of 03.10.2026 the Ollama library lists several newer families, which we selected because they fit in 128 GB. [24] The sizes are for the Q4_K_M variant unless stated otherwise:

FamilyVariantWeightsFits?
Qwen3.5 122B-A10B (10B active)Q4_K_M, 256K context81 GB [19]on the edge: the weights are 81 GB, little is left for the cache
Nemotron 3 Super 120B-A12B (12B active)Q4_K_M · Q8_087 · 132 GB [20]Q4_K_M: on the edge; Q8_0: no
Gemma 4 31BQ4_K_M · Q8_020 · 34 GB [21]yes, comfortably
Qwen3.8 27BQ4_K_M · Q8_018 · 30 GB [22]yes, comfortably

Models of around 120 billion parameters at 4 bits are the largest that 128 GB can hold sensibly. Above them the territory of clusters begins.

06 · MoEThe MoE trap

"Mixture-of-Experts" (MoE) models activate only part of their parameters for each token. DeepSeek V4 Flash has 284 billion in total but activates 13 billion per token. [11] Hence the temptation to reckon that "13B fits".

It does not fit. All 284 billion weights must be in memory, because it is not known in advance which experts will be needed for the next token. It is the computation that is cheap, not the storage.

§
The rule
With MoE, count the total parameters for memory and the active ones only for speed.

That is why gpt-oss 120B (5.1 billion active) and Nemotron 3.5 Lightning (3 billion active) are faster than a dense model with the same file size, but take up as much memory as their file shows. [5][7]

07 · THE FUNCTIONThe function

We rewrote it: it no longer uses a general coefficient for the cache, but takes the bytes of cache per token (from the model's card or configuration) and the size of the weights file.

python · fits_on_128gb
def fits_on_128gb(weights_gb, kv_bytes_per_token, context_tokens):
    kv_cache = kv_bytes_per_token * context_tokens / 1e9
    overhead = (weights_gb + kv_cache) * 0.15
    total    = weights_gb + kv_cache + overhead
    return {"total_gb": round(total, 1), "fits": total < 110}

# Llama 3.3 70B, Q4_K_M: 43 GB, 327,680 bytes of KV per token (FP16 cache)
fits_on_128gb(43, 327_680, 32_768)     # → 61.8 GB · True
fits_on_128gb(43, 327_680, 131_072)    # → 98.8 GB · True
fits_on_128gb(75, 327_680, 131_072)    # → 135.6 GB · False

# DeepSeek V4 Flash: about 160 GB of weights (without the cache) — does not fit
fits_on_128gb(159.6, 0, 0)             # → 183.5 GB · False

The 110 GB ceiling and the 15% overhead are our assumptions (section 03). The examples are calculations, not measurements.

08 · MEASURINGAn estimate does not replace a measurement

The formula tells you whether it makes sense to download the model. It does not tell you whether it will be fast enough. The speed of a dense model is limited by the memory bandwidth, 273 GB/s on the GX10. [1] A rough upper bound is the bandwidth divided by the size of the weights: for Llama 3.3 70B at Q4_K_M (43 GB) that is about 6 tokens per second, and for the Q8_0 variant (75 GB) about 3.6 (our calculation, not a measurement; see also our article on Ollama on the GX10). A model that fits on the edge usually gives single tokens per second: technically it works, practically it is unusable for interactive work.

Always measure after loading, under two different loads: a single request and a batch of eight. MoE models behave differently in the two cases, which is exactly why they exist.

09 · SOURCESSources

  1. NVIDIA, "NVIDIA DGX Spark" (128 GB unified memory, 273 GB/s) — nvidia.com/…/dgx-spark
  2. ASUS, "ASUS Ascent GX10 — Tech Specs" — asus.com/…/asus-ascent-gx10/techspec
  3. Ollama, library "llama3.3", all tags (Q4_K_M 43 GB, Q8_0 75 GB, FP16 141 GB) — ollama.com/library/llama3.3/tags
  4. Ollama, library "gpt-oss", tags (120b: 65 GB) — ollama.com/library/gpt-oss/tags
  5. OpenAI, model card gpt-oss-120b (117B parameters, 5.1B active, MXFP4) — huggingface.co/openai/gpt-oss-120b
  6. Ollama, library "nemotron-3.5-lightning", tags (Q4_K_M 25 GB, Q8_0 35 GB, BF16 66 GB) — ollama.com/library/nemotron-3.5-lightning/tags
  7. NVIDIA, model card NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 (30B total, 3B active) — huggingface.co/nvidia/…-30B-A3B-BF16
  8. NVIDIA, model card NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 (recipe for DGX Spark / GB10) — huggingface.co/nvidia/…-30B-A3B-NVFP4
  9. Ollama, library "muse-glimmer", tags (Q4_K_M 18 GB, Q8_0 31 GB, BF16 57 GB) — ollama.com/library/muse-glimmer/tags
  10. Meta, model card Muse-Glimmer-30B (dense, about 29.6B, Apache 2.0) — huggingface.co/meta-models/Muse-Glimmer-30B
  11. DeepSeek, model card DeepSeek-V4-Flash (284B total, 13B active, 1M context) — huggingface.co/deepseek-ai/DeepSeek-V4-Flash
  12. Ollama, library "deepseek-v4-flash" (marked as withdrawn on 27.08.2026) — ollama.com/library/deepseek-v4-flash
  13. DeepSeek, model card DeepSeek-V4.1-Flash (552B backbone, KV cache 890 bytes per token) — huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
  14. Ollama, library "deepseek-v4.1-flash" — ollama.com/library/deepseek-v4.1-flash
  15. Ollama, library "kimi-k2.6" (1.04T parameters, cloud model only) — ollama.com/library/kimi-k2.6
  16. Moonshot AI, Kimi-K2.6 on Hugging Face (weight files of about 595 GB) — huggingface.co/moonshotai/Kimi-K2.6
  17. Ollama, library "kimi-k3" (2.81T parameters, cloud model only) — ollama.com/library/kimi-k3
  18. Moonshot AI, Kimi-K3 on Hugging Face (weight files of about 1.56 TB) — huggingface.co/moonshotai/Kimi-K3
  19. Ollama, library "qwen3.5", tags (122B-A10B: 81 GB) — ollama.com/library/qwen3.5/tags
  20. Ollama, library "nemotron-3-super", tags (120B-A12B: 87 GB and 132 GB) — ollama.com/library/nemotron-3-super/tags
  21. Ollama, library "gemma4", tags (31B: 20 GB and 34 GB) — ollama.com/library/gemma4/tags
  22. Ollama, library "qwen3.8", tags (27B: 18 GB and 30 GB) — ollama.com/library/qwen3.8/tags
  23. Configuration of Llama 3.3 70B Instruct (80 layers, 8 KV heads, head size 128), a copy by unsloth, a secondary source — huggingface.co/unsloth/Llama-3.3-70B-Instruct
  24. Ollama, list of models in the library — ollama.com/library

Checked on 03.10.2026. The calculations for the KV cache, the 110 GB ceiling, the 15% overhead and the speed bound are ours, not measurements. Model catalogues change quickly; see the primary sources before you quote a size.

10 · RELATEDContinue from here