AI · Hardware
Ollama on the GX10: 70B models locally, what is true
The first version of this article promised "70B models in under 8 GB of VRAM". That is not true: a 70B model in its 4-bit form weighs about 43 GB. But on a machine with 128 GB of unified memory, such as the ASUS Ascent GX10, it fits, and that is the real story.
- The headline of the original (06.05.2026) claimed that 70B models run locally in under 8 GB of VRAM.
"70B models locally in under 8GB VRAM"→ Llama 3.3 70B (Q4_K_M) is 43 GB and Qwen 2.5 72B (Q4_K_M) is 47 GB, about five times more than 8 GB; they fit in the 128 GB of unified memory of the GX10. [3][4][1][2] - The original promised "real latency measurements on DGX Spark" but contained none.
"real latency measurements"→ there are no measurements here; we give only a calculated upper bound based on memory bandwidth, clearly labelled as a calculation. [1][3] - Added: GX10 specifications from the official ASUS and NVIDIA pages, installation steps, the current Ollama version (v0.35.0 as of 2 October 2026), a note on the Qwen licence and sources. The original was a stub of about 70 words. [2][6][7][4]
01 · THE MACHINEWhat the GX10 is
The ASUS Ascent GX10 is a compact desktop AI computer on the NVIDIA GB10 platform, the same one the NVIDIA DGX Spark is built on. The key point is the memory: it is unified. The processor and the graphics chip share the same 128 GB, instead of the graphics chip having a small video memory of its own. [1][2]
| Parameter | Value | Source |
|---|---|---|
| Chip | NVIDIA GB10 Grace Blackwell; 20-core Arm processor | [1][2] |
| Memory | 128 GB LPDDR5x, unified (coherent) | [1][2] |
| Memory bandwidth | 273 GB/s | [1] |
| Compute | up to 1 PFLOP at FP4 | [1][2] |
| Networking | 10 GbE and ConnectX-7 (200 Gbps according to NVIDIA) | [1][2] |
| Power | 240 W adapter | [1][2] |
| Size and weight (ASUS) | 150 × 150 × 51 mm · 1.48 kg | [2] |
| What NVIDIA promises | inference of models up to 200 billion parameters; fine-tuning up to 70 billion (with 128 GB) | [1] |
On 13 October 2025 Ollama announced that it had worked with NVIDIA so that the DGX Spark runs fast "out of the box". [5]
02 · THE MEMORYHow much memory 70B models need
Both models from the original article are available in the Ollama library in the 4-bit quantisation Q4_K_M. The sizes are those of the weights file; the context cache and the system come on top. [3][4]
| Model | Parameters | Size (Q4_K_M) | Context | Licence |
|---|---|---|---|---|
| llama3.3:70b | 70.6 billion | 43 GB | 128K | Llama 3.3 Community License |
| qwen2.5:72b | 72.7 billion | 47 GB | 32K (Ollama default) | Qwen License |
For business use, pay attention to the licence: Qwen 2.5 72B is under the Qwen licence, not Apache 2.0 like the smaller models of the family. [4] The Llama 3.3 licence is Meta's "Community License". [3] Read them before building a model into a product.
03 · THE SPEEDHow fast? Honest about speed
The original promised latency measurements. There are none here: we do not publish a number we have not measured for this exact version of Ollama, this model and this context size.
There is, however, a simple calculation for the upper bound. With a dense model, the whole set of weights is read from memory for every new token. With a memory bandwidth of 273 GB/s [1] and a size of 43 GB [3]:
| Model | Calculation | Upper bound |
|---|---|---|
| Llama 3.3 70B (43 GB) | 273 ÷ 43 | about 6 tokens per second |
| Qwen 2.5 72B (47 GB) | 273 ÷ 47 | about 5.8 tokens per second |
The practical conclusion: a 70B model on the GX10 is convenient for tasks where you are not waiting for a real-time answer, such as document analysis, drafts and batch processing. For fast chat, smaller models are a better fit.
04 · INSTALLATIONHow to run Ollama
The GX10 runs Linux on Arm (NVIDIA DGX OS). The official Ollama documentation for Linux gives a one-line installation and an ARM64 package. [6][2]
Install
Run the official script.
bash · installing Ollamacurl -fsSL https://ollama.com/install.sh | shCheck the version
As of 2 October 2026 the latest stable version is v0.35.0 (28 September 2026); the "rc" versions are pre-releases. [7]
bash · versionollama -vRun a model
The first run downloads the file: 43 GB and 47 GB respectively. Allow disk space and time. [3][4]
bash · Llama 3.3 70B and Qwen 2.5 72Bollama run llama3.3:70b ollama run qwen2.5:72bKeep it updated
To upgrade, run the same script again. [6]
05 · CURRENTThe 2024 models and today
Llama 3.3 was released on 6 December 2024. [3] Qwen 2.5 dates from September 2024. [4] In the Ollama library they are marked as updated one and two years ago respectively. [3][4] As of 2 October 2026 the Ollama project description already lists newer families: Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma [7], and the notes for the latest versions mention Qwen 3.8 and Gemma 4. [7]
So treat the two models as an example of size: the rule "as much memory as the model weighs, plus a margin" holds for every new family. Before choosing a model, check the current list in the Ollama library and the size of the specific variant.
06 · SOURCESSources
- NVIDIA, "NVIDIA DGX Spark" (specifications, 128 GB, 273 GB/s, 1 PFLOP FP4) — nvidia.com/…/dgx-spark
- ASUS, "ASUS Ascent GX10 — Tech Specs" — asus.com/…/asus-ascent-gx10/techspec
- Ollama, library "llama3.3" (tag 70b: Q4_K_M, 43 GB, 128K) — ollama.com/library/llama3.3:70b
- Ollama, library "qwen2.5" (tag 72b: Q4_K_M, 47 GB, licence) — ollama.com/library/qwen2.5:72b
- Ollama, "NVIDIA DGX Spark", 13 October 2025 — ollama.com/blog/nvidia-spark
- Ollama, documentation "Linux" (installation, ARM64, updating) — docs.ollama.com/linux
- Ollama, releases on GitHub (v0.35.0 — 28 September 2026) — github.com/ollama/ollama/releases
Checked on 2 October 2026. The speed calculation in section 03 is ours, not a measurement. Sizes and versions change; see the primary sources.