The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewUPDATED 2026-10-03
IDENTITY
module
GX10-04-24 · CLI coding agent with a local model (Aider, Claude Code via Ollama, a small custom agent)
series
GX10 (local AI server class: NVIDIA GB10, e.g. ASUS Ascent GX10 / DGX Spark)
level
Intermediate
duration
40–60 min
prerequisites
A GB10-class machine with Ollama reachable from the host on loopback, a model that supports tool calling, git, Python 3.8–3.13 (for Aider), a terminal
trust_label
UPDATED 2026-10-03 (checked against the official Aider, Ollama and Ollama-library documentation) · NOT TESTED on a GB10 machine · commands not run by the author · the custom agent code was not run
versions
no version is pinned in this lesson; model sizes and context lengths are taken from ollama.com/library on 2026-10-03
language
human view: bg · english edition: /en/academy/gx10/ (same file name)
previous / next
GX10 series index / GX10 series index
PURPOSE

Run a coding agent in the terminal against a local model: it reads files, edits code, runs commands, looks at the result and repeats. Use Aider (git-integrated pair programming) or Claude Code pointed at Ollama's Anthropic-compatible API, and understand the loop by writing a small tool-calling agent with a human approval step before every shell command. Nothing leaves the machine.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

GX10 series index: kagami.bg/en/academy/gx10/ · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
gx10nvidia-gb10ollamaaidercoding-agenttool-callingclilocal-llm
UPDATED · 03.10.2026

CLI Coding Agent on GX10: Aider and Ollama

A programming helper that lives in the terminal: it reads your files, writes and fixes code, runs commands, looks at the result and tries again. The model is local — your code never leaves the machine. You learn to get it running, to keep it within bounds and to understand how it works inside.

⏱ 40–60 min Intermediate GX10 NVIDIA GB10 · 128 GB unified memory Aider · Claude Code · Ollama · Python · Git
Ollama (the model)🔒 local Aider🔒 local Claude Code through Ollama🔒 local Your own agent (Python)🔒 local
🔄
UPDATED · 03.10.2026 — what changed
The lesson was rebuilt against the current documentation. We removed what was outdated or wrong: the model "qwen3:72b" (the Ollama library has no such size of Qwen3 — the sizes are 0.6 / 1.7 / 4 / 8 / 14 / 30 / 32 / 235 billion), the ollama/ prefix (Aider's documentation recommends ollama_chat/), the api-base line in the configuration (the address is set with OLLAMA_API_BASE), the pip … --break-system-packages installation (now: aider-install), the claim that Aider does not make automatic Git commits by default (it does — which is why /undo exists), the invented results in the demos (such as "8 passed in 0.3s"), and the description of Continue.dev, which we could not verify. We added: choosing a model with tool support (qwen3-coder:30b), the context window size (at least 64,000 tokens per the Ollama documentation), how to connect to the Ollama from the 04-01 stack (its port is closed and must be opened only to 127.0.0.1), Claude Code through Ollama's compatible API, permissions, safe habits and a quiz. We reworked the small agent: it has a step limit, works in one folder and asks the human before every command.
⚠️
What we have not run ourselves
We had no GB10-class machine at hand during the check. The commands and settings were checked against the official documentation as of 03.10.2026, but we have not run them ourselves — which is why there is no "TESTED" label. Especially unchecked: the speed of qwen3-coder:30b on GB10, how Claude Code behaves with a local model, and the code of the small agent (it has not been executed). How well a model writes code depends on the task — verify everything it suggests.

01What you'll learn

02Before you start

💡
Why the model must "know tools"
An agent does not only write text — it asks the program to carry out an action (read a file, run a command). This is called tool calling and only some models support it. In the Ollama library such models carry the "tools" label. qwen3-coder has 30 billion parameters, about 19 GB and a context window of 256 thousand tokens (ollama.com, 03.10.2026).

03Steps

  1. What a coding agent is

    An ordinary chat answers and stops. An agent works in a loop: it thinks about the task, takes an action (reads a file, writes code, runs a command), sees the result and decides what comes next — until the task is done or it cannot continue.

    🔄
    The agent loop
    plan → action → observation → correction → again, until it is done.

    An example (invented) run for the task "make a small web interface for a to-do list":

    1. the agent looks around the folder and reads the existing files;
    2. it creates the files with the code;
    3. it runs the program — and gets an error about a missing library;
    4. it installs the library and runs again;
    5. it checks the response with one request and reports what it did.

    The important point: from the second and third steps the agent is already changing things and running programs on your behalf. That is why this lesson has so much about limits and approval.

  2. Prepare Ollama and the model

    Download the model and check that it is loaded:

    bash · on the machine
    ollama pull qwen3-coder:30b
    ollama run qwen3-coder:30b "Say 'ready' in one word"
    ollama ps

    ollama ps shows whether the model is loaded, how much memory it takes and what context it runs with. The memory is shared by the CPU and the GPU (128 GB), so do not keep several large models loaded at once.

    The context must be large

    Context is how much text the model "keeps in its head". The Ollama documentation recommends at least 64,000 tokens for agents and coding tools. Ollama picks a default based on available memory (per its documentation: under 24 GiB — 4k, 24–48 GiB — 32k, from 48 GiB — 256k). How this applies to unified memory such as GB10's we have not checked, so verify with ollama ps and set the value yourself if needed. For Ollama installed directly on the system:

    bash
    OLLAMA_CONTEXT_LENGTH=64000 ollama serve

    A larger context needs more memory. Aider (step 4) raises the window itself for each request, so this matters less with it, but it is important for Claude Code and your own agent.

    Ollama from the lesson 04-01 stack

    In the stack from lesson 04-01 Ollama is an internal service with no published port — only the other containers can reach it. A tool running on the machine itself will not get to it. The safest option is to publish the port to loopback only, in the ollama service section:

    yaml · compose.yaml · the ollama service
        ports:
          - "127.0.0.1:11434:11434"
        environment:
          OLLAMA_CONTEXT_LENGTH: "64000"

    After the change: docker compose up -d. This way Ollama is visible from the machine but not from the network.

    ⚠️
    Never 0.0.0.0
    Do not publish the port as "11434:11434" and do not set OLLAMA_HOST=0.0.0.0 — that lets anyone on your network use your model. We have not run this Compose change.
  3. Install Aider

    Aider is a terminal program for "pair programming" with a model. It works on your Git repository and records every change as a separate commit, so changes can be reverted. The official way:

    bash
    python -m pip install aider-install
    aider-install
    aider --version

    aider-install puts it in its own environment and, if needed, brings a separate Python 3.12 — so you do not touch the system Python. (If pip refuses to install into the system, that is the distribution's protection, not an error — do not bypass it; use aider-install.)

  4. Connect Aider to Ollama

    First tell Aider where Ollama is, then start it in the project folder with the right model prefix. Its documentation recommends ollama_chat/ over ollama/:

    bash
    export OLLAMA_API_BASE=http://127.0.0.1:11434
    cd ~/my-project          # a Git repository
    aider --model ollama_chat/qwen3-coder:30b

    So you do not type the model every time, put it in a configuration file .aider.conf.yml (in your home folder or in the project):

    yaml · .aider.conf.yml
    model: ollama_chat/qwen3-coder:30b

    The Ollama address stays in the OLLAMA_API_BASE variable (put it in your shell profile to keep it). Aider may show a warning about a model it does not know well — that is normal for newer models.

    💡
    Git is the safety net
    By default Aider records each of its changes in Git with a description. If you do not like one — /undo reverts the last. Do not turn this off (--no-auto-commits, --no-git) until you are used to it; with --no-git you have no way back except your own copies.
  5. Working with Aider: commands and a first task

    Commands in Aider's chat start with /:

    CommandWhat it does
    /add fileAdds a file Aider may edit
    /read-only fileAdds a file for reading only (for example a description or rules)
    /lsShows which files are in the conversation
    /run commandRuns a shell command and optionally adds its output to the conversation
    /test commandRuns a command and adds the output if it fails
    /diffShows the changes since the last message
    /undoReverts Aider's last commit in Git
    /commitCommits changes made outside the conversation
    /askQuestions about the code without editing files
    /exitExit

    A sample first task in a practice project:

    in Aider's chat
    /add app.py
    Add a function that reads a CSV file and returns the average of the "price" column. Handle a missing file.
    /diff
    /run python app.py

    Read /diff before you continue. If something is wrong — /undo. Start with small, clear tasks.

    ✅
    Where agents are strongest
    Refactoring existing code · adding tests · rewriting a function to a clear specification · hunting a specific bug · documentation · small migrations. Weak spot: vague tasks with no way to check the result.
  6. Optional: Claude Code pointed at Ollama

    Ollama has an Anthropic-compatible API through which Claude Code can work with a local model. The shortest path per the Ollama documentation:

    bash
    ollama launch claude --model qwen3-coder:30b

    The manual variant (after Claude Code has been installed following its official instructions):

    bash
    export ANTHROPIC_AUTH_TOKEN=ollama
    export ANTHROPIC_API_KEY=""
    export ANTHROPIC_BASE_URL=http://127.0.0.1:11434
    claude --model qwen3-coder:30b

    Here the "token" is only a placeholder — Ollama ignores it; no real key goes in. The Ollama documentation recommends a context of 64k tokens or more.

    ⚠️
    We have not tried it
    We have not run Claude Code with a local model on GB10. Quality and speed depend on the model — a small local model is not equal to the large cloud ones. Read Ollama's documentation for this option as well.
  7. Your own small agent (to understand how it works)

    All coding agents are variations of one loop. Here is a teaching example that reads files, writes files and runs commands through the local model. Three important differences from a naive version: it has a step limit, it works only in the current folder and it asks you before every command.

    bash · preparation
    mkdir -p ~/agent-demo && cd ~/agent-demo
    git init
    python3 -m venv .venv
    . .venv/bin/activate
    pip install openai
    python · mini_agent.py
    #!/usr/bin/env python3
    # mini_agent.py — a teaching coding agent with a local model through Ollama
    import json
    import subprocess
    from pathlib import Path
    from openai import OpenAI
    
    # Ollama speaks the OpenAI-compatible protocol; the key is required but ignored
    client = OpenAI(base_url="http://127.0.0.1:11434/v1/", api_key="ollama")
    MODEL = "qwen3-coder:30b"
    MAX_STEPS = 15
    WORKDIR = Path.cwd().resolve()
    
    TOOLS = [
        {"type": "function", "function": {
            "name": "run_shell",
            "description": "Run a shell command in the project folder and return its output",
            "parameters": {"type": "object",
                           "properties": {"cmd": {"type": "string"}},
                           "required": ["cmd"]}}},
        {"type": "function", "function": {
            "name": "read_file",
            "description": "Read a text file from the project folder",
            "parameters": {"type": "object",
                           "properties": {"path": {"type": "string"}},
                           "required": ["path"]}}},
        {"type": "function", "function": {
            "name": "write_file",
            "description": "Write text to a file in the project folder",
            "parameters": {"type": "object",
                           "properties": {"path": {"type": "string"},
                                          "content": {"type": "string"}},
                           "required": ["path", "content"]}}},
    ]
    
    
    def safe_path(rel):
        """Allows only paths inside the working folder."""
        p = (WORKDIR / rel).resolve()
        if p != WORKDIR and WORKDIR not in p.parents:
            raise ValueError("The path is outside the working folder")
        return p
    
    
    def run_tool(name, args):
        try:
            if name == "run_shell":
                print(f"\nCommand: {args['cmd']}")
                if input("Run it? [y/N] ").strip().lower() != "y":
                    return "Refused by the human."
                r = subprocess.run(args["cmd"], shell=True, cwd=WORKDIR,
                                   capture_output=True, text=True, timeout=30)
                return (r.stdout + r.stderr)[-4000:]
            if name == "read_file":
                return safe_path(args["path"]).read_text()[:8000]
            if name == "write_file":
                safe_path(args["path"]).write_text(args["content"])
                return "Written: " + args["path"]
            return "Unknown tool"
        except Exception as e:  # hand the error back to the model so it can fix it
            return "Error: " + str(e)
    
    
    def agent(task):
        messages = [{"role": "user", "content": task}]
        for _ in range(MAX_STEPS):
            resp = client.chat.completions.create(
                model=MODEL, messages=messages, tools=TOOLS)
            msg = resp.choices[0].message
            if not msg.tool_calls:
                print("\nDone:", msg.content)
                return
            messages.append(msg.model_dump(exclude_none=True))
            for tc in msg.tool_calls:
                args = json.loads(tc.function.arguments)
                result = run_tool(tc.function.name, args)
                print("tool:", tc.function.name)
                messages.append({"role": "tool", "tool_call_id": tc.id,
                                 "content": result})
        print("Step limit reached. Stopped.")
    
    
    if __name__ == "__main__":
        agent(input("Task: "))

    Run it with a harmless task:

    bash
    python mini_agent.py
    # Task: Create prices.csv with 5 rows (name,price) and a script avg.py that prints the average price

    Watch how every step is "tool → result → back to the model". That is the whole secret; ready-made agents add more tools, memory of the repository and better rules.

    ⚠️
    The code has not been run
    We have not executed this example. It is deliberately simple: the path guard and the "Run it?" question are a minimum, not complete security. A command that looks harmless to you can still delete files. Do not point it at valuable folders.
  8. Safe habits

    • Work in a separate branch (git switch -c trial) or in a practice folder — easy to throw away.
    • Review every change (/diff) before you accept it, and run the tests.
    • Approve commands by hand until you know the model's behaviour. Do not turn on "no asking" modes on valuable data.
    • Keep no secrets (passwords, keys) in the project the agent reads — it sees them and can print them.
    • Run unfamiliar code in a container or an isolated folder.
    • The Ollama port is on 127.0.0.1 only.
    • The model errs confidently — the human is responsible for the code.

04Check

Quiz

1. Which model prefix does Aider's documentation recommend for Ollama?

2. Why is a large context (at least 64,000 tokens) recommended for a coding agent?

3. Ollama from the lesson 04-01 stack has no published port, and you want Aider on the machine to use it. What is right?

4. What does the small agent in the lesson do before it runs a shell command?

05What's next

06Sources

  1. Aider: installation · Aider: Ollama 🔒 local — aider-install, ollama_chat/, OLLAMA_API_BASE, context.
  2. Aider: in-chat commands · Aider: Git integration — /add, /read-only, /undo, automatic commits.
  3. Ollama: context length — defaults, the 64k-token recommendation, ollama ps.
  4. Ollama: OpenAI compatibility — the /v1/ address, key "required but ignored", tools.
  5. Ollama: Claude Code — ollama launch claude, environment variables.
  6. qwen3-coder · qwen3 🔒 local — sizes and context windows.
  7. NVIDIA DGX Spark: hardware — unified memory of the GB10 class.