Multi-Agent Chatbot with a Router on GX10
One model is rarely the best programmer, mathematician and conversation partner at the same time. Here we build a chatbot made of several agents: a small, fast model (the router) reads the question and hands it to the right specialist. Everything runs locally through Ollama, and the data never leaves the machine.
qwen3:7b as the router and qwen3:72b for math and research. Neither is in the Ollama catalogue (the qwen3 sizes are 0.6b, 1.7b, 4b, 8b, 14b, 30b, 32b and 235b). Now: router and chat — llama3.2:3b, code — qwen2.5-coder:7b, math and research — qwen3:32b. We fixed the interface code. In Gradio 6 the tuple format of the chatbot was removed; the history is a list of messages with a role and text — so the interface was rewritten. We fixed the history: "the last 6 messages" means 6 messages (3 exchanges), not 6 pairs. We added: a JSON mode for the router and a safe fallback to the chat agent, a split between the core and the interface, a virtual environment, a closed port (the interface listens only on 127.0.0.1) with access through an SSH tunnel, and a memory budget. We removed the bilingual interface text and the elaborate branded Gradio look.
01What you'll learn
- Why several specialised agents often do a better job than one all-purpose model.
- How the "router → specialist" pattern works and how the router returns a precise answer as JSON.
- How to pass the earlier part of the conversation to the specialist.
- How to give each agent its own model and its own instruction (system prompt).
- How to build a simple web interface with Gradio and open it safely from your own computer.
- How to check whether the router routes correctly, and what to do when it errs.
02Before you start
- A machine of the NVIDIA GB10 class (for example ASUS Ascent GX10 or DGX Spark) with DGX OS and access to its terminal — directly or over SSH.
- Ollama, installed and running on the machine. If you do not have it yet, start with the lesson Open WebUI and Ollama. Check with
ollama --version. - Python 3 with the virtual-environment module (
venv). If it is missing on Ubuntu, install thepython3-venvpackage. - Internet access to download the models — about 27 GB in total (see the table in step 3) — and free disk space (check with
df -h).
qwen3:32b with qwen3:8b (5.2 GB).03Steps
-
How one question flows
The system has one router and four specialists. Why this way? The router does only one cheap job — it decides "whose question is this". The expensive large model wakes up only when it really has to.
Agent What it is for Model Router Classifies the question and returns JSON llama3.2:3bcode Programming, scripts, debugging qwen2.5-coder:7bmath Equations, calculations, step by step qwen3:32bresearch Explanations, comparisons, structured answers qwen3:32bchat Conversation, greetings, small tasks llama3.2:3bThe path is: your message → router → chosen specialist → answer. If the router returns something unintelligible, the question goes to the chat agent — so the system does not break.
-
Prepare a folder and a virtual environment
A virtual environment keeps the project's packages apart from the system Python — they do not break each other.
bash · on the machinemkdir -p ~/multi-agent && cd ~/multi-agent python3 -m venv .venv source .venv/bin/activate pip install "gradio>=6,<7" openaiWe use the
openaipackage only as a client: Ollama has an interface compatible with it athttp://localhost:11434/v1. Gradio is pinned to version 6 (6.29.1 as of 03.10.2026) because the code below is written for it. ⚠️ We have not run these commands ourselves. -
Download the models
bashollama pull llama3.2:3b ollama pull qwen2.5-coder:7b ollama pull qwen3:32b ollama listModel Size (per ollama.com) Role llama3.2:3b2.0 GB router and chat qwen2.5-coder:7b4.7 GB code qwen3:32b20 GB math and research ✅Why the router is not qwen3Qwen3 is a "thinking" model: before the answer it may write a long block of reasoning. For the router that is a needless delay — it only has to say one word. So we use the smallllama3.2:3b, which answers immediately. If you want another router, change it in one variable. -
The core: the file multi_agent.py
All the logic is here — without an interface. That way you can try it in the terminal first and put whatever screen you like on top of it afterwards.
python · ~/multi-agent/multi_agent.pyimport json from openai import OpenAI # Ollama ignores the key, but the client requires some value client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama") ROUTER_MODEL = "llama3.2:3b" AGENTS = { "code": { "model": "qwen2.5-coder:7b", "system": "You are an experienced programmer. You write clean code with short comments. " "You explain what the code does, show a usage example " "and mention edge cases. You prefer Python and Bash.", }, "math": { "model": "qwen3:32b", "system": "You are a mathematician. You solve problems step by step, " "explain what you do and why, and check the result at the end.", }, "research": { "model": "qwen3:32b", "system": "You are a research assistant. You give accurate, complete answers " "organised with headings and bullet points. You separate facts from opinions " "and say when you are not sure.", }, "chat": { "model": "llama3.2:3b", "system": "You are a friendly assistant. You answer briefly, clearly " "and helpfully, in the language the user writes in.", }, } ROUTER_SYSTEM = """Classify the question into EXACTLY ONE of these types: - "code": programming, scripts, debugging - "math": mathematics, equations, calculations, statistics - "research": explanations, concepts, history, comparisons, analysis - "chat": conversation, greetings, personal opinion, small tasks Reply ONLY with JSON: {"agent": "code|math|research|chat"}""" def route(message: str) -> str: """Returns the name of the agent; in any doubt - 'chat'.""" r = client.chat.completions.create( model=ROUTER_MODEL, temperature=0.1, response_format={"type": "json_object"}, messages=[ {"role": "system", "content": ROUTER_SYSTEM}, {"role": "user", "content": message}, ], ) try: agent = json.loads(r.choices[0].message.content).get("agent", "chat") except (json.JSONDecodeError, TypeError, AttributeError): agent = "chat" return agent if agent in AGENTS else "chat" def ask_agent(name: str, message: str, history: list) -> str: """history is a list of {"role": ..., "content": ...}; we take the last 6.""" agent = AGENTS[name] messages = [{"role": "system", "content": agent["system"]}] messages += history[-6:] messages.append({"role": "user", "content": message}) r = client.chat.completions.create( model=agent["model"], temperature=0.7, messages=messages, ) return r.choices[0].message.content def respond(message: str, history: list): """Returns (agent name, answer).""" name = route(message) return name, ask_agent(name, message, history) if __name__ == "__main__": history = [] while True: msg = input("You: ").strip() if msg.lower() in ("exit", "quit"): break if not msg: continue name, answer = respond(msg, history) history += [ {"role": "user", "content": msg}, {"role": "assistant", "content": answer}, ] print(f"\n[{name.upper()}] {answer}\n")💡What happens in the codeJSON mode (response_format) is supported by Ollama's compatible interface and makes the model return valid JSON. Theagent in AGENTScheck catches the case where the router invents an unknown category. The history is a plain list; "the last 6" are 6 messages, that is 3 exchanges. The low temperature of the router (0.1) makes classification stable, while 0.7 for the specialists leaves room for expression. -
Try it in the terminal
bashpython multi_agent.pyType "Hello, how are you?" and then "Write a Python script that reads a CSV file". In front of every answer you should see the agent's name. The first answer from each model is slower — the model is being loaded into memory. Leave with
exit. -
The interface: the file app.py
Gradio makes a web page in a few lines. The important part here is the history: in Gradio 6 it is a list of dicts with
roleandcontent. The content can be text or a list of parts, so we reduce it to plain text before returning it to a model. We show the agent-name tag on screen but strip it from the history so it does not confuse the specialists.python · ~/multi-agent/app.pyimport re import gradio as gr from multi_agent import respond TAG = re.compile(r"^\*\*\[[A-Z]+\]\*\*\n\n") def as_text(content) -> str: """Gradio may give text or a list of parts; we return plain text.""" if isinstance(content, str): return content if isinstance(content, list): return "".join( p.get("text", "") if isinstance(p, dict) else str(p) for p in content ) return str(content) def chat(message, history): if not message.strip(): return "", history plain = [ {"role": h["role"], "content": TAG.sub("", as_text(h["content"]))} for h in history ] name, answer = respond(message, plain) history = history + [ {"role": "user", "content": message}, {"role": "assistant", "content": f"**[{name.upper()}]**\n\n{answer}"}, ] return "", history with gr.Blocks(title="Multi-agent chatbot") as demo: gr.Markdown("# Multi-agent chatbot\nThe router sends the question to: **code** · **math** · **research** · **chat**") chatbot = gr.Chatbot(height=500, label="Conversation") msg = gr.Textbox(placeholder="Ask a question...", label="Message") with gr.Row(): send_btn = gr.Button("Send", variant="primary") clear_btn = gr.Button("Clear") send_btn.click(chat, [msg, chatbot], [msg, chatbot]) msg.submit(chat, [msg, chatbot], [msg, chatbot]) clear_btn.click(lambda: [], None, chatbot) demo.launch(server_name="127.0.0.1", server_port=7860)bashpython app.pyThe interface listens only on
127.0.0.1:7860— others on the network cannot see it. From your own computer you reach it through an SSH tunnel:bash · on your computerssh -L 7860:localhost:7860 <user>@<server-address>Leave the connection open and open
http://localhost:7860in your browser. If you work directly on the machine, just open the same address.⚠️Do not use 0.0.0.0The old version of this lesson started the interface on0.0.0.0, that is, open to the whole network — with no password. We do not do that: anyone on the network could then talk to your models. -
Check the routing
Ask these four questions and see which agent answers:
Question Expected agent Model "Write a Python script that reads a CSV file" CODE qwen2.5-coder:7b"Solve: 3x² + 5x − 2 = 0" MATH qwen3:32b"Explain what the transformer architecture is" RESEARCH qwen3:32b"Hello, how are you?" CHAT llama3.2:3bFor the equation you expect the roots
x = 1/3andx = −2— that way you can check whether the mathematician is right. The small router can err, especially on mixed questions ("write code that solves an equation"). The fix then is inROUTER_SYSTEM: add one short example for each category. -
Traps and improvements
- Slow at first. Loading a 20 GB model takes time. Later answers are quick while the model stays in memory (5 minutes after the last request). If you want to free memory right away, use
ollama stop <model>. - History costs. Every old message is sent to the model again — more history means a slower answer. That is why we keep only the last 6.
- The router is a suggestion. It does not know everything. If it matters, add an "other agent" button or show the agent's name (as we do) so it is clear who answered.
- A human approves actions. This chatbot only answers. Once you give it permission to send emails or write to systems, every action with a real effect must go through a person.
- Slow at first. Loading a 20 GB model takes time. Later answers are quick while the model stays in memory (5 minutes after the last request). If you want to free memory right away, use
04Check
ollama listshowsllama3.2:3b,qwen2.5-coder:7bandqwen3:32b.python multi_agent.pyanswers a greeting through the chat agent.- The four test questions go to CODE, MATH, RESEARCH and CHAT.
- The mathematician solves the equation step by step and reaches
x = 1/3andx = −2. - The interface opens through the tunnel at
http://localhost:7860. - A follow-up question ("can you explain that more simply?") uses the earlier messages.
- Nothing listens on
0.0.0.0.
Quiz
1. Why is the router a small model rather than the largest one?
2. What does the code do if the router returns JSON with an unknown category?
3. Why is Gradio started with server_name="127.0.0.1"?
4. In what form is the conversation history in the Gradio 6 chatbot?
05What's next
06Sources
- Ollama: OpenAI compatibility 🔒 local — the
/v1address,response_format, the key is ignored. - Ollama: FAQ — 5 minutes in memory,
ollama stop, several models at once. - llama3.2 · qwen2.5-coder · qwen3 (tags) 🔒 local — available sizes: 2.0 GB, 4.7 GB, 20 GB.
- Gradio on PyPI · Gradio: changelog — version 6.29.1; the tuple format of the chatbot was removed.
- openai on PyPI — the client we use to talk to Ollama.
- NVIDIA DGX Spark: hardware — 128 GB of shared memory.