CLI Coding Agent on GX10: Aider and Ollama
A programming helper that lives in the terminal: it reads your files, writes and fixes code, runs commands, looks at the result and tries again. The model is local — your code never leaves the machine. You learn to get it running, to keep it within bounds and to understand how it works inside.
ollama/ prefix (Aider's documentation recommends ollama_chat/), the api-base line in the configuration (the address is set with OLLAMA_API_BASE), the pip … --break-system-packages installation (now: aider-install), the claim that Aider does not make automatic Git commits by default (it does — which is why /undo exists), the invented results in the demos (such as "8 passed in 0.3s"), and the description of Continue.dev, which we could not verify. We added: choosing a model with tool support (qwen3-coder:30b), the context window size (at least 64,000 tokens per the Ollama documentation), how to connect to the Ollama from the 04-01 stack (its port is closed and must be opened only to 127.0.0.1), Claude Code through Ollama's compatible API, permissions, safe habits and a quiz. We reworked the small agent: it has a step limit, works in one folder and asks the human before every command.
qwen3-coder:30b on GB10, how Claude Code behaves with a local model, and the code of the small agent (it has not been executed). How well a model writes code depends on the task — verify everything it suggests.01What you'll learn
- What a coding agent is, and why the "plan → action → observation → correction" loop sets it apart from an ordinary chat.
- How to choose a model that can call tools, and give it enough context.
- How to run Aider against a local model through Ollama and work with its commands.
- How to point Claude Code at Ollama and what to expect.
- How to write a small agent of your own, and why a human approves every command in it.
- Which tasks suit an agent and which do not.
02Before you start
- A machine of the NVIDIA GB10 class (for example ASUS Ascent GX10 or DGX Spark) with a working Ollama — from lesson 04-01 or installed directly on the system.
- Ollama must be reachable from the machine itself at
127.0.0.1:11434. If you run it as an internal service with no published port (as in lesson 04-01), tools on the machine will not see it — see step 2. - Git and Python (for Aider — versions 3.8 to 3.13 per its documentation).
- Free space: the model in this lesson is about 19 GB.
- A small practice project in a Git repository — do not work on something valuable the first time.
qwen3-coder has 30 billion parameters, about 19 GB and a context window of 256 thousand tokens (ollama.com, 03.10.2026).03Steps
-
What a coding agent is
An ordinary chat answers and stops. An agent works in a loop: it thinks about the task, takes an action (reads a file, writes code, runs a command), sees the result and decides what comes next — until the task is done or it cannot continue.
🔄The agent loopplan → action → observation → correction → again, until it is done.An example (invented) run for the task "make a small web interface for a to-do list":
- the agent looks around the folder and reads the existing files;
- it creates the files with the code;
- it runs the program — and gets an error about a missing library;
- it installs the library and runs again;
- it checks the response with one request and reports what it did.
The important point: from the second and third steps the agent is already changing things and running programs on your behalf. That is why this lesson has so much about limits and approval.
-
Prepare Ollama and the model
Download the model and check that it is loaded:
bash · on the machineollama pull qwen3-coder:30b ollama run qwen3-coder:30b "Say 'ready' in one word" ollama psollama psshows whether the model is loaded, how much memory it takes and what context it runs with. The memory is shared by the CPU and the GPU (128 GB), so do not keep several large models loaded at once.The context must be large
Context is how much text the model "keeps in its head". The Ollama documentation recommends at least 64,000 tokens for agents and coding tools. Ollama picks a default based on available memory (per its documentation: under 24 GiB — 4k, 24–48 GiB — 32k, from 48 GiB — 256k). How this applies to unified memory such as GB10's we have not checked, so verify with
ollama psand set the value yourself if needed. For Ollama installed directly on the system:bashOLLAMA_CONTEXT_LENGTH=64000 ollama serveA larger context needs more memory. Aider (step 4) raises the window itself for each request, so this matters less with it, but it is important for Claude Code and your own agent.
Ollama from the lesson 04-01 stack
In the stack from lesson 04-01 Ollama is an internal service with no published port — only the other containers can reach it. A tool running on the machine itself will not get to it. The safest option is to publish the port to loopback only, in the
ollamaservice section:yaml · compose.yaml · the ollama serviceports: - "127.0.0.1:11434:11434" environment: OLLAMA_CONTEXT_LENGTH: "64000"After the change:
docker compose up -d. This way Ollama is visible from the machine but not from the network.⚠️Never 0.0.0.0Do not publish the port as"11434:11434"and do not setOLLAMA_HOST=0.0.0.0— that lets anyone on your network use your model. We have not run this Compose change. -
Install Aider
Aider is a terminal program for "pair programming" with a model. It works on your Git repository and records every change as a separate commit, so changes can be reverted. The official way:
bashpython -m pip install aider-install aider-install aider --versionaider-installputs it in its own environment and, if needed, brings a separate Python 3.12 — so you do not touch the system Python. (Ifpiprefuses to install into the system, that is the distribution's protection, not an error — do not bypass it; useaider-install.) -
Connect Aider to Ollama
First tell Aider where Ollama is, then start it in the project folder with the right model prefix. Its documentation recommends
ollama_chat/overollama/:bashexport OLLAMA_API_BASE=http://127.0.0.1:11434 cd ~/my-project # a Git repository aider --model ollama_chat/qwen3-coder:30bSo you do not type the model every time, put it in a configuration file
.aider.conf.yml(in your home folder or in the project):yaml · .aider.conf.ymlmodel: ollama_chat/qwen3-coder:30bThe Ollama address stays in the
OLLAMA_API_BASEvariable (put it in your shell profile to keep it). Aider may show a warning about a model it does not know well — that is normal for newer models.💡Git is the safety netBy default Aider records each of its changes in Git with a description. If you do not like one —/undoreverts the last. Do not turn this off (--no-auto-commits,--no-git) until you are used to it; with--no-gityou have no way back except your own copies. -
Working with Aider: commands and a first task
Commands in Aider's chat start with
/:Command What it does /add fileAdds a file Aider may edit /read-only fileAdds a file for reading only (for example a description or rules) /lsShows which files are in the conversation /run commandRuns a shell command and optionally adds its output to the conversation /test commandRuns a command and adds the output if it fails /diffShows the changes since the last message /undoReverts Aider's last commit in Git /commitCommits changes made outside the conversation /askQuestions about the code without editing files /exitExit A sample first task in a practice project:
in Aider's chat/add app.py Add a function that reads a CSV file and returns the average of the "price" column. Handle a missing file. /diff /run python app.pyRead
/diffbefore you continue. If something is wrong —/undo. Start with small, clear tasks.✅Where agents are strongestRefactoring existing code · adding tests · rewriting a function to a clear specification · hunting a specific bug · documentation · small migrations. Weak spot: vague tasks with no way to check the result. -
Optional: Claude Code pointed at Ollama
Ollama has an Anthropic-compatible API through which Claude Code can work with a local model. The shortest path per the Ollama documentation:
bashollama launch claude --model qwen3-coder:30bThe manual variant (after Claude Code has been installed following its official instructions):
bashexport ANTHROPIC_AUTH_TOKEN=ollama export ANTHROPIC_API_KEY="" export ANTHROPIC_BASE_URL=http://127.0.0.1:11434 claude --model qwen3-coder:30bHere the "token" is only a placeholder — Ollama ignores it; no real key goes in. The Ollama documentation recommends a context of 64k tokens or more.
⚠️We have not tried itWe have not run Claude Code with a local model on GB10. Quality and speed depend on the model — a small local model is not equal to the large cloud ones. Read Ollama's documentation for this option as well. -
Your own small agent (to understand how it works)
All coding agents are variations of one loop. Here is a teaching example that reads files, writes files and runs commands through the local model. Three important differences from a naive version: it has a step limit, it works only in the current folder and it asks you before every command.
bash · preparationmkdir -p ~/agent-demo && cd ~/agent-demo git init python3 -m venv .venv . .venv/bin/activate pip install openaipython · mini_agent.py#!/usr/bin/env python3 # mini_agent.py — a teaching coding agent with a local model through Ollama import json import subprocess from pathlib import Path from openai import OpenAI # Ollama speaks the OpenAI-compatible protocol; the key is required but ignored client = OpenAI(base_url="http://127.0.0.1:11434/v1/", api_key="ollama") MODEL = "qwen3-coder:30b" MAX_STEPS = 15 WORKDIR = Path.cwd().resolve() TOOLS = [ {"type": "function", "function": { "name": "run_shell", "description": "Run a shell command in the project folder and return its output", "parameters": {"type": "object", "properties": {"cmd": {"type": "string"}}, "required": ["cmd"]}}}, {"type": "function", "function": { "name": "read_file", "description": "Read a text file from the project folder", "parameters": {"type": "object", "properties": {"path": {"type": "string"}}, "required": ["path"]}}}, {"type": "function", "function": { "name": "write_file", "description": "Write text to a file in the project folder", "parameters": {"type": "object", "properties": {"path": {"type": "string"}, "content": {"type": "string"}}, "required": ["path", "content"]}}}, ] def safe_path(rel): """Allows only paths inside the working folder.""" p = (WORKDIR / rel).resolve() if p != WORKDIR and WORKDIR not in p.parents: raise ValueError("The path is outside the working folder") return p def run_tool(name, args): try: if name == "run_shell": print(f"\nCommand: {args['cmd']}") if input("Run it? [y/N] ").strip().lower() != "y": return "Refused by the human." r = subprocess.run(args["cmd"], shell=True, cwd=WORKDIR, capture_output=True, text=True, timeout=30) return (r.stdout + r.stderr)[-4000:] if name == "read_file": return safe_path(args["path"]).read_text()[:8000] if name == "write_file": safe_path(args["path"]).write_text(args["content"]) return "Written: " + args["path"] return "Unknown tool" except Exception as e: # hand the error back to the model so it can fix it return "Error: " + str(e) def agent(task): messages = [{"role": "user", "content": task}] for _ in range(MAX_STEPS): resp = client.chat.completions.create( model=MODEL, messages=messages, tools=TOOLS) msg = resp.choices[0].message if not msg.tool_calls: print("\nDone:", msg.content) return messages.append(msg.model_dump(exclude_none=True)) for tc in msg.tool_calls: args = json.loads(tc.function.arguments) result = run_tool(tc.function.name, args) print("tool:", tc.function.name) messages.append({"role": "tool", "tool_call_id": tc.id, "content": result}) print("Step limit reached. Stopped.") if __name__ == "__main__": agent(input("Task: "))Run it with a harmless task:
bashpython mini_agent.py # Task: Create prices.csv with 5 rows (name,price) and a script avg.py that prints the average priceWatch how every step is "tool → result → back to the model". That is the whole secret; ready-made agents add more tools, memory of the repository and better rules.
⚠️The code has not been runWe have not executed this example. It is deliberately simple: the path guard and the "Run it?" question are a minimum, not complete security. A command that looks harmless to you can still delete files. Do not point it at valuable folders. -
Safe habits
- Work in a separate branch (
git switch -c trial) or in a practice folder — easy to throw away. - Review every change (
/diff) before you accept it, and run the tests. - Approve commands by hand until you know the model's behaviour. Do not turn on "no asking" modes on valuable data.
- Keep no secrets (passwords, keys) in the project the agent reads — it sees them and can print them.
- Run unfamiliar code in a container or an isolated folder.
- The Ollama port is on
127.0.0.1only. - The model errs confidently — the human is responsible for the code.
- Work in a separate branch (
04Check
ollama psshows a loaded model with tool support and a context of at least 64k tokens.- Ollama is reachable from the machine at
127.0.0.1:11434and is not published to the network. aider --versionanswers; Aider is started withollama_chat/….- You did one small task in a practice repository, reviewed
/diffand tried/undo. - The small agent does a harmless task and asks before every command.
- You know which tasks suit an agent and which do not.
Quiz
1. Which model prefix does Aider's documentation recommend for Ollama?
2. Why is a large context (at least 64,000 tokens) recommended for a coding agent?
3. Ollama from the lesson 04-01 stack has no published port, and you want Aider on the machine to use it. What is right?
4. What does the small agent in the lesson do before it runs a shell command?
05What's next
06Sources
- Aider: installation · Aider: Ollama 🔒 local —
aider-install,ollama_chat/,OLLAMA_API_BASE, context. - Aider: in-chat commands · Aider: Git integration —
/add,/read-only,/undo, automatic commits. - Ollama: context length — defaults, the 64k-token recommendation,
ollama ps. - Ollama: OpenAI compatibility — the
/v1/address, key "required but ignored", tools. - Ollama: Claude Code —
ollama launch claude, environment variables. - qwen3-coder · qwen3 🔒 local — sizes and context windows.
- NVIDIA DGX Spark: hardware — unified memory of the GB10 class.