The KAGAMI mark КАГАМИ
kagami.bg/en/blog/2026-05-ollama-gx10.html · article · machine-readable viewPUBLISHED 2026-05-06 · VERIFIED 2026-10-02 · UPDATED 2026-10-02
IDENTITY
title
Ollama on the GX10: 70B models locally — what is actually true
author
KAGAMI Ltd. (КАГАМИ ЕООД), Varna · organisation byline, no individual author named
language
English edition (human view and this agent view) · Bulgarian original: https://kagami.bg/blog/2026-05-ollama-gx10.html
topic
AI hardware · running 70B-class open models with Ollama on the ASUS Ascent GX10 (NVIDIA GB10, same platform as NVIDIA DGX Spark)
audience
Small companies and technical staff in Bulgaria considering local LLM inference
status
General information. No latency or throughput benchmarks are published on this page.
SUMMARY

The original (2026-05-06) headline claimed 70B models run locally "under 8 GB VRAM" with "real latency measurements"; the original contained no measurements. Corrected: in the Ollama library the 4-bit (Q4_K_M) builds of Llama 3.3 70B and Qwen 2.5 72B are 43 GB and 47 GB, so they cannot fit in 8 GB of VRAM. They do fit in the 128 GB coherent unified memory of the ASUS Ascent GX10 / NVIDIA DGX Spark (GB10), whose memory bandwidth is 273 GB/s. Because token generation for a dense model reads all weights per token, a rough upper bound is about 6 tokens/s for the 43 GB model (273 / 43) — a back-of-the-envelope calculation, not a measurement. Ollama installs on Linux ARM64 with a one-line script; the latest stable release on 2026-10-02 is v0.35.0.

KEY CLAIMS (dated)
WHAT WAS UPDATED (2026-10-02)
SOURCES
AGENT INSTRUCTIONS
TAGS
ollamaasus-ascent-gx10nvidia-dgx-sparkllama-3.3-70bqwen-2.5-72blocal-llmquantization
VERIFIED · 02.10.2026 UPDATED · 02.10.2026

AI · Hardware

Ollama on the GX10: 70B models locally, what is true

The first version of this article promised "70B models in under 8 GB of VRAM". That is not true: a 70B model in its 4-bit form weighs about 43 GB. But on a machine with 128 GB of unified memory, such as the ASUS Ascent GX10, it fits, and that is the real story.

AI Ollama · Llama 3.3 70B · Qwen 2.5 72B Technical
UPDATED · 02.10.2026 · WHAT WAS UPDATED

01 · THE MACHINEWhat the GX10 is

The ASUS Ascent GX10 is a compact desktop AI computer on the NVIDIA GB10 platform, the same one the NVIDIA DGX Spark is built on. The key point is the memory: it is unified. The processor and the graphics chip share the same 128 GB, instead of the graphics chip having a small video memory of its own. [1][2]

ParameterValueSource
ChipNVIDIA GB10 Grace Blackwell; 20-core Arm processor[1][2]
Memory128 GB LPDDR5x, unified (coherent)[1][2]
Memory bandwidth273 GB/s[1]
Computeup to 1 PFLOP at FP4[1][2]
Networking10 GbE and ConnectX-7 (200 Gbps according to NVIDIA)[1][2]
Power240 W adapter[1][2]
Size and weight (ASUS)150 × 150 × 51 mm · 1.48 kg[2]
What NVIDIA promisesinference of models up to 200 billion parameters; fine-tuning up to 70 billion (with 128 GB)[1]

On 13 October 2025 Ollama announced that it had worked with NVIDIA so that the DGX Spark runs fast "out of the box". [5]

02 · THE MEMORYHow much memory 70B models need

Both models from the original article are available in the Ollama library in the 4-bit quantisation Q4_K_M. The sizes are those of the weights file; the context cache and the system come on top. [3][4]

ModelParametersSize (Q4_K_M)ContextLicence
llama3.3:70b70.6 billion43 GB128KLlama 3.3 Community License
qwen2.5:72b72.7 billion47 GB32K (Ollama default)Qwen License
⚠
It does not fit in 8 GB
43 GB and 47 GB are about five times more than 8 GB. The claim "70B in under 8 GB of VRAM" is wrong. The model fits with 128 GB of unified memory, not because it is "squeezed" any further, but because there is room for it.

For business use, pay attention to the licence: Qwen 2.5 72B is under the Qwen licence, not Apache 2.0 like the smaller models of the family. [4] The Llama 3.3 licence is Meta's "Community License". [3] Read them before building a model into a product.

03 · THE SPEEDHow fast? Honest about speed

The original promised latency measurements. There are none here: we do not publish a number we have not measured for this exact version of Ollama, this model and this context size.

There is, however, a simple calculation for the upper bound. With a dense model, the whole set of weights is read from memory for every new token. With a memory bandwidth of 273 GB/s [1] and a size of 43 GB [3]:

ModelCalculationUpper bound
Llama 3.3 70B (43 GB)273 ÷ 43about 6 tokens per second
Qwen 2.5 72B (47 GB)273 ÷ 47about 5.8 tokens per second
ℹ
This is a calculation, not a measurement
The real speed is lower than this bound and depends on the Ollama version, the context length and the load. If speed is decisive for your case, measure it on your own machine with your own prompt.

The practical conclusion: a 70B model on the GX10 is convenient for tasks where you are not waiting for a real-time answer, such as document analysis, drafts and batch processing. For fast chat, smaller models are a better fit.

04 · INSTALLATIONHow to run Ollama

The GX10 runs Linux on Arm (NVIDIA DGX OS). The official Ollama documentation for Linux gives a one-line installation and an ARM64 package. [6][2]

  1. Install

    Run the official script.

    bash · installing Ollama
    curl -fsSL https://ollama.com/install.sh | sh
  2. Check the version

    As of 2 October 2026 the latest stable version is v0.35.0 (28 September 2026); the "rc" versions are pre-releases. [7]

    bash · version
    ollama -v
  3. Run a model

    The first run downloads the file: 43 GB and 47 GB respectively. Allow disk space and time. [3][4]

    bash · Llama 3.3 70B and Qwen 2.5 72B
    ollama run llama3.3:70b
    ollama run qwen2.5:72b
  4. Keep it updated

    To upgrade, run the same script again. [6]

05 · CURRENTThe 2024 models and today

Llama 3.3 was released on 6 December 2024. [3] Qwen 2.5 dates from September 2024. [4] In the Ollama library they are marked as updated one and two years ago respectively. [3][4] As of 2 October 2026 the Ollama project description already lists newer families: Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma [7], and the notes for the latest versions mention Qwen 3.8 and Gemma 4. [7]

So treat the two models as an example of size: the rule "as much memory as the model weighs, plus a margin" holds for every new family. Before choosing a model, check the current list in the Ollama library and the size of the specific variant.

§
The rule of thumb
The size of the model file is the minimum. Add room for the context cache and for the system, and choose a machine whose memory is clearly larger than the file.

06 · SOURCESSources

  1. NVIDIA, "NVIDIA DGX Spark" (specifications, 128 GB, 273 GB/s, 1 PFLOP FP4) — nvidia.com/…/dgx-spark
  2. ASUS, "ASUS Ascent GX10 — Tech Specs" — asus.com/…/asus-ascent-gx10/techspec
  3. Ollama, library "llama3.3" (tag 70b: Q4_K_M, 43 GB, 128K) — ollama.com/library/llama3.3:70b
  4. Ollama, library "qwen2.5" (tag 72b: Q4_K_M, 47 GB, licence) — ollama.com/library/qwen2.5:72b
  5. Ollama, "NVIDIA DGX Spark", 13 October 2025 — ollama.com/blog/nvidia-spark
  6. Ollama, documentation "Linux" (installation, ARM64, updating) — docs.ollama.com/linux
  7. Ollama, releases on GitHub (v0.35.0 — 28 September 2026) — github.com/ollama/ollama/releases

Checked on 2 October 2026. The speed calculation in section 03 is ours, not a measurement. Sizes and versions change; see the primary sources.

07 · RELATEDContinue from here