AI · Hardware
What really fits in 128 GB: the sum to do before you download the model
The question asked most often about machines of the DGX Spark class is one: "will this model fit?" The answer is thirty seconds of arithmetic, as long as you know which numbers to add up and where to get them. Here is the sum and a table of real models, checked against the primary sources on 3 October 2026.
- The model sizes were checked against the Ollama library and the Hugging Face cards.
Llama 3.3 70B, Q4: ~35 GB→ 43 GB;"FP8": ~70 GB→ Q8_0: 75 GB. That is also what our article on Ollama on the GX10 says. [3] gpt-oss 120B, Q4: ~60 GB→ 65 GB: the model was released directly in the 4-bit format MXFP4, not "converted" to Q4. [4][5]Nemotron 3.5 Lightning, FP8: ~30 GB→ Q4_K_M 25 GB, Q8_0 35 GB, BF16 66 GB; the active parameters are 3 billion. [6][7]DeepSeek V4 Flash: "about 22 billion active"→ 13 billion active out of 284 billion in total; the weights are about 160 GB. Its entry in Ollama was withdrawn on 27.08.2026, and its successor V4.1 Flash weighs about 510 GB. [11][12][13]Kimi K2.6: ">400 GB"→ about 595 GB;Kimi K3: ">1 TB"→ about 1.56 TB. In Ollama both are available only as cloud models. [15][16][17][18]KV cache: "128K context ≈ 15 GB"→ depends on the architecture; for Llama 3.3 70B the same context needs about 43 GB (our calculation). The fixed rule "0.12 GB per 1K tokens" has been removed, and the function rewritten. [23]- Added: a section on the newer families that do fit, sources and clearly labelled calculations of our own. The "working ceiling of 110 GB" is now marked as our own practical rule, not as manufacturer data.
01 · WEIGHTSWeights: the main part
Every parameter takes up space, and exactly how much depends on the precision in which you store it. The memory of the GX10 is 128 GB unified: the processor and the graphics chip share the same 128 GB. [1][2] A rule of thumb (bytes per parameter):
| Precision | Bytes / parameter | 30B model | 70B model |
|---|---|---|---|
| BF16 / FP16 | 2 | 60 GB | 140 GB |
| 8 bits (FP8, Q8_0) | about 1 | about 30 GB | about 70 GB |
| 4 bits (Q4, NVFP4) | about 0.5 | about 15 GB | about 35 GB |
The first conclusion stands: a 70B model in full precision does not fit in 128 GB (141 GB). At 8 bits it fits, but little room is left. At 4 bits it fits comfortably, and you pay in quality.
02 · KV CACHEThe KV cache: the term everyone forgets
The weights are only one part. While running, the model keeps keys and values (the KV cache) for every token of the context. The cache grows linearly with the length of the context and with a long context it can rival the weights themselves in size.
Exactly how much it grows depends on the architecture of the model; the first version of the article gave one general coefficient, and it is not right for all models. The formula for a model with classic attention is: 2 (keys and values) × number of layers × number of KV heads × head size × bytes per value. For Llama 3.3 70B the configuration is 80 layers, 8 KV heads, head size 128 [23]. With the cache in FP16 (2 bytes) this gives:
| Context | KV cache of Llama 3.3 70B | Calculation |
|---|---|---|
| 8K (8,192 tokens) | about 2.7 GB | 327,680 bytes × 8,192 |
| 32K (32,768 tokens) | about 10.7 GB | 327,680 bytes × 32,768 |
| 128K (131,072 tokens) | about 43 GB | 327,680 bytes × 131,072 |
The table is our calculation from the published configuration, not a measurement. 327,680 = 2 × 80 × 8 × 128 × 2 bytes per token.
At a 128K context the cache of this model is as large as its own weights at Q4_K_M (43 GB). This is where most planning fails. A model that loads without a problem collapses at the third long document, because the sum was done only for the weights.
The new architectures cut this cost sharply. For example, DeepSeek-V4.1-Flash states a global KV cache of about 890 bytes per token, about a quarter of the previous V4-Flash. [13][14] So do not carry the numbers over from one model to another: take the cache parameters from the model card.
03 · OVERHEADOverhead and the working ceiling
Activations, memory fragmentation, the runtime framework itself. Allow about 15% on top. Not because it is exactly that much, but because filling memory to the last byte causes failures that you then hunt for hours.
04 · THE TABLEWhat weighs what, and what fits
The sizes are those of the weights file according to the Ollama library, unless stated otherwise; "active" means the parameters that work for a single token. The "Fits?" column follows our rule (weights plus 15% overhead, a ceiling of about 110 GB, without the KV cache unless stated).
| Model | Parameters (total / active) | Variant | Weights | Fits? |
|---|---|---|---|---|
| Nemotron 3.5 Lightning | 30B MoE / 3B | Q4_K_M · Q8_0 · BF16 | 25 · 35 · 66 GB [6][7] | yes, comfortably; there is also an official NVFP4 of about 21.6 GB [8] |
| Muse Glimmer (Meta) | about 29.6B, dense | Q4_K_M · Q8_0 · BF16 | 18 · 31 · 57 GB [9][10] | yes, comfortably |
| Llama 3.3 70B | 70B, dense | Q4_K_M | 43 GB [3] | yes; with a 128K context about 99 GB (our calculation) |
| Llama 3.3 70B | 70B, dense | Q8_0 | 75 GB [3] | on the edge: up to a 32K context yes (about 99 GB), with 128K no |
| gpt-oss 120B | 117B MoE / 5.1B | MXFP4 (original) | 65 GB [4][5] | yes |
| DeepSeek V4 Flash | 284B MoE / 13B | FP4 + FP8 mixed | about 160 GB [11] | no |
| Kimi K2.6 | 1.04T MoE | 4-bit weights | about 595 GB [15][16] | no |
| Kimi K3 | about 2.8T MoE | MXFP4 | about 1.56 TB [17][18] | no: a cluster is needed |
The sizes in GB are as Ollama shows them; for Hugging Face they are our sum of the sizes of the weight files in the repository (decimal GB), so they may differ by a few percent. According to Ollama, Kimi K2.6 and K3 are available only as cloud models (":cloud"); there is no local file to download there. [15][17]
05 · THE NEW ONESThe new families that do fit
As of 03.10.2026 the Ollama library lists several newer families, which we selected because they fit in 128 GB. [24] The sizes are for the Q4_K_M variant unless stated otherwise:
| Family | Variant | Weights | Fits? |
|---|---|---|---|
| Qwen3.5 122B-A10B (10B active) | Q4_K_M, 256K context | 81 GB [19] | on the edge: the weights are 81 GB, little is left for the cache |
| Nemotron 3 Super 120B-A12B (12B active) | Q4_K_M · Q8_0 | 87 · 132 GB [20] | Q4_K_M: on the edge; Q8_0: no |
| Gemma 4 31B | Q4_K_M · Q8_0 | 20 · 34 GB [21] | yes, comfortably |
| Qwen3.8 27B | Q4_K_M · Q8_0 | 18 · 30 GB [22] | yes, comfortably |
Models of around 120 billion parameters at 4 bits are the largest that 128 GB can hold sensibly. Above them the territory of clusters begins.
06 · MoEThe MoE trap
"Mixture-of-Experts" (MoE) models activate only part of their parameters for each token. DeepSeek V4 Flash has 284 billion in total but activates 13 billion per token. [11] Hence the temptation to reckon that "13B fits".
It does not fit. All 284 billion weights must be in memory, because it is not known in advance which experts will be needed for the next token. It is the computation that is cheap, not the storage.
That is why gpt-oss 120B (5.1 billion active) and Nemotron 3.5 Lightning (3 billion active) are faster than a dense model with the same file size, but take up as much memory as their file shows. [5][7]
07 · THE FUNCTIONThe function
We rewrote it: it no longer uses a general coefficient for the cache, but takes the bytes of cache per token (from the model's card or configuration) and the size of the weights file.
def fits_on_128gb(weights_gb, kv_bytes_per_token, context_tokens):
kv_cache = kv_bytes_per_token * context_tokens / 1e9
overhead = (weights_gb + kv_cache) * 0.15
total = weights_gb + kv_cache + overhead
return {"total_gb": round(total, 1), "fits": total < 110}
# Llama 3.3 70B, Q4_K_M: 43 GB, 327,680 bytes of KV per token (FP16 cache)
fits_on_128gb(43, 327_680, 32_768) # → 61.8 GB · True
fits_on_128gb(43, 327_680, 131_072) # → 98.8 GB · True
fits_on_128gb(75, 327_680, 131_072) # → 135.6 GB · False
# DeepSeek V4 Flash: about 160 GB of weights (without the cache) — does not fit
fits_on_128gb(159.6, 0, 0) # → 183.5 GB · FalseThe 110 GB ceiling and the 15% overhead are our assumptions (section 03). The examples are calculations, not measurements.
08 · MEASURINGAn estimate does not replace a measurement
The formula tells you whether it makes sense to download the model. It does not tell you whether it will be fast enough. The speed of a dense model is limited by the memory bandwidth, 273 GB/s on the GX10. [1] A rough upper bound is the bandwidth divided by the size of the weights: for Llama 3.3 70B at Q4_K_M (43 GB) that is about 6 tokens per second, and for the Q8_0 variant (75 GB) about 3.6 (our calculation, not a measurement; see also our article on Ollama on the GX10). A model that fits on the edge usually gives single tokens per second: technically it works, practically it is unusable for interactive work.
Always measure after loading, under two different loads: a single request and a batch of eight. MoE models behave differently in the two cases, which is exactly why they exist.
09 · SOURCESSources
- NVIDIA, "NVIDIA DGX Spark" (128 GB unified memory, 273 GB/s) — nvidia.com/…/dgx-spark
- ASUS, "ASUS Ascent GX10 — Tech Specs" — asus.com/…/asus-ascent-gx10/techspec
- Ollama, library "llama3.3", all tags (Q4_K_M 43 GB, Q8_0 75 GB, FP16 141 GB) — ollama.com/library/llama3.3/tags
- Ollama, library "gpt-oss", tags (120b: 65 GB) — ollama.com/library/gpt-oss/tags
- OpenAI, model card gpt-oss-120b (117B parameters, 5.1B active, MXFP4) — huggingface.co/openai/gpt-oss-120b
- Ollama, library "nemotron-3.5-lightning", tags (Q4_K_M 25 GB, Q8_0 35 GB, BF16 66 GB) — ollama.com/library/nemotron-3.5-lightning/tags
- NVIDIA, model card NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 (30B total, 3B active) — huggingface.co/nvidia/…-30B-A3B-BF16
- NVIDIA, model card NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 (recipe for DGX Spark / GB10) — huggingface.co/nvidia/…-30B-A3B-NVFP4
- Ollama, library "muse-glimmer", tags (Q4_K_M 18 GB, Q8_0 31 GB, BF16 57 GB) — ollama.com/library/muse-glimmer/tags
- Meta, model card Muse-Glimmer-30B (dense, about 29.6B, Apache 2.0) — huggingface.co/meta-models/Muse-Glimmer-30B
- DeepSeek, model card DeepSeek-V4-Flash (284B total, 13B active, 1M context) — huggingface.co/deepseek-ai/DeepSeek-V4-Flash
- Ollama, library "deepseek-v4-flash" (marked as withdrawn on 27.08.2026) — ollama.com/library/deepseek-v4-flash
- DeepSeek, model card DeepSeek-V4.1-Flash (552B backbone, KV cache 890 bytes per token) — huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- Ollama, library "deepseek-v4.1-flash" — ollama.com/library/deepseek-v4.1-flash
- Ollama, library "kimi-k2.6" (1.04T parameters, cloud model only) — ollama.com/library/kimi-k2.6
- Moonshot AI, Kimi-K2.6 on Hugging Face (weight files of about 595 GB) — huggingface.co/moonshotai/Kimi-K2.6
- Ollama, library "kimi-k3" (2.81T parameters, cloud model only) — ollama.com/library/kimi-k3
- Moonshot AI, Kimi-K3 on Hugging Face (weight files of about 1.56 TB) — huggingface.co/moonshotai/Kimi-K3
- Ollama, library "qwen3.5", tags (122B-A10B: 81 GB) — ollama.com/library/qwen3.5/tags
- Ollama, library "nemotron-3-super", tags (120B-A12B: 87 GB and 132 GB) — ollama.com/library/nemotron-3-super/tags
- Ollama, library "gemma4", tags (31B: 20 GB and 34 GB) — ollama.com/library/gemma4/tags
- Ollama, library "qwen3.8", tags (27B: 18 GB and 30 GB) — ollama.com/library/qwen3.8/tags
- Configuration of Llama 3.3 70B Instruct (80 layers, 8 KV heads, head size 128), a copy by unsloth, a secondary source — huggingface.co/unsloth/Llama-3.3-70B-Instruct
- Ollama, list of models in the library — ollama.com/library
Checked on 03.10.2026. The calculations for the KV cache, the 110 GB ceiling, the 15% overhead and the speed bound are ours, not measurements. Model catalogues change quickly; see the primary sources before you quote a size.
10 · RELATEDContinue from here
Choosing a Model: Memory, Speed, Quality
How to choose a model according to memory and the task.
Article · BlogOllama on the GX10: 70B models locally, what is true
How much Llama 3.3 70B and Qwen 2.5 72B weigh and how to run Ollama.
Step · The LadderSession B1 · 45 min · €99
We discuss whether a local AI makes sense for your company and what machine you need.