The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewUPDATED 2026-10-03
IDENTITY
module
GX10-04-27 · Multi-agent chatbot with a router and specialist agents
series
GX10 (local AI server class: NVIDIA GB10, e.g. ASUS Ascent GX10 / DGX Spark)
level
Intermediate
duration
about 1 h, plus model downloads (about 27 GB in total)
prerequisites
A GB10-class machine with DGX OS (Ubuntu 24.04, Arm64), shell access (local or SSH), Ollama installed and running on the machine (lesson 04-07), Python 3 with venv, internet for the first downloads
trust_label
UPDATED 2026-10-03 (model names, sizes and package versions checked against ollama.com, PyPI and the Ollama docs) · NOT TESTED on a GB10 machine · code not run end to end · not marked VERIFIED
versions
Gradio 6.29.1 (PyPI, 2026-10-03; chatbot history is a list of role/content messages, the tuple format was removed in Gradio 6) · openai Python package (current release, used only as an HTTP client for Ollama) · Ollama latest release 0.35.1 on GitHub (2026-10-03)
models
llama3.2:3b (2.0 GB) as router and chat agent · qwen2.5-coder:7b (4.7 GB) as code agent · qwen3:32b (20 GB) as math and research agent. Note: qwen3:7b and qwen3:72b do not exist on ollama.com; the qwen3 sizes are 0.6b, 1.7b, 4b, 8b, 14b, 30b, 32b, 235b
language
human view: en · bulgarian edition: /academy/gx10/ (same file name)
previous / next
GX10 series index / GX10 series index
PURPOSE

Build a chatbot made of one router and four specialist agents on a local AI server. The router (a small, fast model at temperature 0.1) classifies each user message into code, math, research or chat and answers with JSON; the matching specialist answers with its own system prompt and model. The last 6 messages of the conversation are passed to the specialist. The web UI is Gradio, bound to 127.0.0.1 and reached through an SSH tunnel. Everything runs locally; no external API.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

GX10 series index: kagami.bg/en/academy/gx10/ · related lessons: 04-07 Open WebUI and Ollama (04-07_Open_WebUI_Ollama.html), 04-01 n8n on GX10 (04-01_n8n_Training_GX10.html) · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
gx10nvidia-gb10multi-agentrouterollamagradiolocal-llmchatbot
UPDATED · 03.10.2026

Multi-Agent Chatbot with a Router on GX10

One model is rarely the best programmer, mathematician and conversation partner at the same time. Here we build a chatbot made of several agents: a small, fast model (the router) reads the question and hands it to the right specialist. Everything runs locally through Ollama, and the data never leaves the machine.

⏱ ~1 h Intermediate GX10 Router · specialists · history Ollama · Python · Gradio
Ollama (the models)🔒 local Python · Gradio🔒 local
🔄
UPDATED · 03.10.2026 — what changed
We fixed models that do not exist. The previous version used qwen3:7b as the router and qwen3:72b for math and research. Neither is in the Ollama catalogue (the qwen3 sizes are 0.6b, 1.7b, 4b, 8b, 14b, 30b, 32b and 235b). Now: router and chat — llama3.2:3b, code — qwen2.5-coder:7b, math and research — qwen3:32b. We fixed the interface code. In Gradio 6 the tuple format of the chatbot was removed; the history is a list of messages with a role and text — so the interface was rewritten. We fixed the history: "the last 6 messages" means 6 messages (3 exchanges), not 6 pairs. We added: a JSON mode for the router and a safe fallback to the chat agent, a split between the core and the interface, a virtual environment, a closed port (the interface listens only on 127.0.0.1) with access through an SSH tunnel, and a memory budget. We removed the bilingual interface text and the elaborate branded Gradio look.
⚠️
What we have not run ourselves
We had no GB10-class machine and did not run the code end to end — which is why there are no "TESTED" and "VERIFIED" labels. Model names, their sizes and package versions were checked online (03.10.2026). Not checked: the exact behaviour of the Gradio 6 history with more complex content (which is why the code reduces it to plain text), model speed on your machine, and how often the small router errs on Bulgarian. Try it and tell us how it went.

01What you'll learn

02Before you start

💡
Why memory is not a problem here
The machine has 128 GB of memory, shared by the CPU and the GPU. By default Ollama keeps a model in memory for 5 minutes after the last request and can hold up to three at once if they fit (per the Ollama documentation). Our three models total about 27 GB, so they fit easily. On a smaller machine, replace qwen3:32b with qwen3:8b (5.2 GB).

03Steps

  1. How one question flows

    The system has one router and four specialists. Why this way? The router does only one cheap job — it decides "whose question is this". The expensive large model wakes up only when it really has to.

    AgentWhat it is forModel
    RouterClassifies the question and returns JSONllama3.2:3b
    codeProgramming, scripts, debuggingqwen2.5-coder:7b
    mathEquations, calculations, step by stepqwen3:32b
    researchExplanations, comparisons, structured answersqwen3:32b
    chatConversation, greetings, small tasksllama3.2:3b

    The path is: your message → router → chosen specialist → answer. If the router returns something unintelligible, the question goes to the chat agent — so the system does not break.

  2. Prepare a folder and a virtual environment

    A virtual environment keeps the project's packages apart from the system Python — they do not break each other.

    bash · on the machine
    mkdir -p ~/multi-agent && cd ~/multi-agent
    python3 -m venv .venv
    source .venv/bin/activate
    pip install "gradio>=6,<7" openai

    We use the openai package only as a client: Ollama has an interface compatible with it at http://localhost:11434/v1. Gradio is pinned to version 6 (6.29.1 as of 03.10.2026) because the code below is written for it. ⚠️ We have not run these commands ourselves.

  3. Download the models

    bash
    ollama pull llama3.2:3b
    ollama pull qwen2.5-coder:7b
    ollama pull qwen3:32b
    ollama list
    ModelSize (per ollama.com)Role
    llama3.2:3b2.0 GBrouter and chat
    qwen2.5-coder:7b4.7 GBcode
    qwen3:32b20 GBmath and research
    ✅
    Why the router is not qwen3
    Qwen3 is a "thinking" model: before the answer it may write a long block of reasoning. For the router that is a needless delay — it only has to say one word. So we use the small llama3.2:3b, which answers immediately. If you want another router, change it in one variable.
  4. The core: the file multi_agent.py

    All the logic is here — without an interface. That way you can try it in the terminal first and put whatever screen you like on top of it afterwards.

    python · ~/multi-agent/multi_agent.py
    import json
    from openai import OpenAI
    
    # Ollama ignores the key, but the client requires some value
    client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
    
    ROUTER_MODEL = "llama3.2:3b"
    
    AGENTS = {
        "code": {
            "model": "qwen2.5-coder:7b",
            "system": "You are an experienced programmer. You write clean code with short comments. "
                      "You explain what the code does, show a usage example "
                      "and mention edge cases. You prefer Python and Bash.",
        },
        "math": {
            "model": "qwen3:32b",
            "system": "You are a mathematician. You solve problems step by step, "
                      "explain what you do and why, and check the result at the end.",
        },
        "research": {
            "model": "qwen3:32b",
            "system": "You are a research assistant. You give accurate, complete answers "
                      "organised with headings and bullet points. You separate facts from opinions "
                      "and say when you are not sure.",
        },
        "chat": {
            "model": "llama3.2:3b",
            "system": "You are a friendly assistant. You answer briefly, clearly "
                      "and helpfully, in the language the user writes in.",
        },
    }
    
    ROUTER_SYSTEM = """Classify the question into EXACTLY ONE of these types:
    - "code": programming, scripts, debugging
    - "math": mathematics, equations, calculations, statistics
    - "research": explanations, concepts, history, comparisons, analysis
    - "chat": conversation, greetings, personal opinion, small tasks
    
    Reply ONLY with JSON: {"agent": "code|math|research|chat"}"""
    
    
    def route(message: str) -> str:
        """Returns the name of the agent; in any doubt - 'chat'."""
        r = client.chat.completions.create(
            model=ROUTER_MODEL,
            temperature=0.1,
            response_format={"type": "json_object"},
            messages=[
                {"role": "system", "content": ROUTER_SYSTEM},
                {"role": "user", "content": message},
            ],
        )
        try:
            agent = json.loads(r.choices[0].message.content).get("agent", "chat")
        except (json.JSONDecodeError, TypeError, AttributeError):
            agent = "chat"
        return agent if agent in AGENTS else "chat"
    
    
    def ask_agent(name: str, message: str, history: list) -> str:
        """history is a list of {"role": ..., "content": ...}; we take the last 6."""
        agent = AGENTS[name]
        messages = [{"role": "system", "content": agent["system"]}]
        messages += history[-6:]
        messages.append({"role": "user", "content": message})
        r = client.chat.completions.create(
            model=agent["model"],
            temperature=0.7,
            messages=messages,
        )
        return r.choices[0].message.content
    
    
    def respond(message: str, history: list):
        """Returns (agent name, answer)."""
        name = route(message)
        return name, ask_agent(name, message, history)
    
    
    if __name__ == "__main__":
        history = []
        while True:
            msg = input("You: ").strip()
            if msg.lower() in ("exit", "quit"):
                break
            if not msg:
                continue
            name, answer = respond(msg, history)
            history += [
                {"role": "user", "content": msg},
                {"role": "assistant", "content": answer},
            ]
            print(f"\n[{name.upper()}] {answer}\n")
    💡
    What happens in the code
    JSON mode (response_format) is supported by Ollama's compatible interface and makes the model return valid JSON. The agent in AGENTS check catches the case where the router invents an unknown category. The history is a plain list; "the last 6" are 6 messages, that is 3 exchanges. The low temperature of the router (0.1) makes classification stable, while 0.7 for the specialists leaves room for expression.
  5. Try it in the terminal

    bash
    python multi_agent.py

    Type "Hello, how are you?" and then "Write a Python script that reads a CSV file". In front of every answer you should see the agent's name. The first answer from each model is slower — the model is being loaded into memory. Leave with exit.

  6. The interface: the file app.py

    Gradio makes a web page in a few lines. The important part here is the history: in Gradio 6 it is a list of dicts with role and content. The content can be text or a list of parts, so we reduce it to plain text before returning it to a model. We show the agent-name tag on screen but strip it from the history so it does not confuse the specialists.

    python · ~/multi-agent/app.py
    import re
    import gradio as gr
    from multi_agent import respond
    
    TAG = re.compile(r"^\*\*\[[A-Z]+\]\*\*\n\n")
    
    
    def as_text(content) -> str:
        """Gradio may give text or a list of parts; we return plain text."""
        if isinstance(content, str):
            return content
        if isinstance(content, list):
            return "".join(
                p.get("text", "") if isinstance(p, dict) else str(p) for p in content
            )
        return str(content)
    
    
    def chat(message, history):
        if not message.strip():
            return "", history
        plain = [
            {"role": h["role"], "content": TAG.sub("", as_text(h["content"]))}
            for h in history
        ]
        name, answer = respond(message, plain)
        history = history + [
            {"role": "user", "content": message},
            {"role": "assistant", "content": f"**[{name.upper()}]**\n\n{answer}"},
        ]
        return "", history
    
    
    with gr.Blocks(title="Multi-agent chatbot") as demo:
        gr.Markdown("# Multi-agent chatbot\nThe router sends the question to: **code** · **math** · **research** · **chat**")
        chatbot = gr.Chatbot(height=500, label="Conversation")
        msg = gr.Textbox(placeholder="Ask a question...", label="Message")
        with gr.Row():
            send_btn = gr.Button("Send", variant="primary")
            clear_btn = gr.Button("Clear")
        send_btn.click(chat, [msg, chatbot], [msg, chatbot])
        msg.submit(chat, [msg, chatbot], [msg, chatbot])
        clear_btn.click(lambda: [], None, chatbot)
    
    demo.launch(server_name="127.0.0.1", server_port=7860)
    bash
    python app.py

    The interface listens only on 127.0.0.1:7860 — others on the network cannot see it. From your own computer you reach it through an SSH tunnel:

    bash · on your computer
    ssh -L 7860:localhost:7860 <user>@<server-address>

    Leave the connection open and open http://localhost:7860 in your browser. If you work directly on the machine, just open the same address.

    ⚠️
    Do not use 0.0.0.0
    The old version of this lesson started the interface on 0.0.0.0, that is, open to the whole network — with no password. We do not do that: anyone on the network could then talk to your models.
  7. Check the routing

    Ask these four questions and see which agent answers:

    QuestionExpected agentModel
    "Write a Python script that reads a CSV file"CODEqwen2.5-coder:7b
    "Solve: 3x² + 5x − 2 = 0"MATHqwen3:32b
    "Explain what the transformer architecture is"RESEARCHqwen3:32b
    "Hello, how are you?"CHATllama3.2:3b

    For the equation you expect the roots x = 1/3 and x = −2 — that way you can check whether the mathematician is right. The small router can err, especially on mixed questions ("write code that solves an equation"). The fix then is in ROUTER_SYSTEM: add one short example for each category.

  8. Traps and improvements

    • Slow at first. Loading a 20 GB model takes time. Later answers are quick while the model stays in memory (5 minutes after the last request). If you want to free memory right away, use ollama stop <model>.
    • History costs. Every old message is sent to the model again — more history means a slower answer. That is why we keep only the last 6.
    • The router is a suggestion. It does not know everything. If it matters, add an "other agent" button or show the agent's name (as we do) so it is clear who answered.
    • A human approves actions. This chatbot only answers. Once you give it permission to send emails or write to systems, every action with a real effect must go through a person.

04Check

Quiz

1. Why is the router a small model rather than the largest one?

2. What does the code do if the router returns JSON with an unknown category?

3. Why is Gradio started with server_name="127.0.0.1"?

4. In what form is the conversation history in the Gradio 6 chatbot?

05What's next

06Sources

  1. Ollama: OpenAI compatibility 🔒 local — the /v1 address, response_format, the key is ignored.
  2. Ollama: FAQ — 5 minutes in memory, ollama stop, several models at once.
  3. llama3.2 · qwen2.5-coder · qwen3 (tags) 🔒 local — available sizes: 2.0 GB, 4.7 GB, 20 GB.
  4. Gradio on PyPI · Gradio: changelog — version 6.29.1; the tuple format of the chatbot was removed.
  5. openai on PyPI — the client we use to talk to Ollama.
  6. NVIDIA DGX Spark: hardware — 128 GB of shared memory.