The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewTESTED · VERIFIED 2026-09-30
IDENTITY
module
KW I22 · Two-node pattern: worker + brain
series
KAGAMI Way · Track I — Infrastructure
level
Intermediate
duration
~12 min
trust_label
TESTED (two live internal pairs, no public date) · VERIFIED 2026-09-30 (/api/version endpoint, OLLAMA_HOST binding, Docker host network against official docs)
language
human view: en · bulgarian edition: /academy/moduli/KW_I22_Two_Node_Pattern.html
prev
KW_I21_Model_Sync.html
next
../moduli_index.html (all lessons)
PURPOSE

A reusable two-node pattern for local AI: one node runs the OCR/worker container (Docker + Tesseract), a second node runs the model server with the GPU. They talk over a private mesh network. Any site becomes a working "reader" without duplicating the GPU node.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

../moduli_index.html (all lessons) · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
two-nodedockerollamaprivate-meshocrlocal-llmreusable
TESTED VERIFIED · 30.09.2026

Two nodes: worker + brain

A reusable pattern: one node runs the worker (OCR in Docker), the other runs the model; they talk over a private network. That way every site becomes a working "reader" without duplicating the expensive GPU node.

⏱ ~12 min Intermediate Local AI private network · Docker · two nodes
Docker + Tesseract🔒 local Model server (Ollama)🔒 local

01What you will learn

02Before you start

03Steps

  1. The idea: separate the worker from the brain

    A GPU is expensive and scarce. But reading a document doesn't need a GPU — the OCR and the orchestration run on a cheap node with an ordinary processor. Only the thinking (the model) needs a GPU. So we split: the worker (Docker + OCR) sits where the documents are; the brain (the model server) sits on the GPU machine. One brain serves many workers.

    diagram
    [ worker: Docker + OCR ]  --private network-->  [ brain: model server :11434 + GPU ]
      cheap node, watched folder                      GPU machine, text model
  2. The two roles

    RoleWhat it runsRequires
    WorkerOCR container, watched folderDocker
    Brainmodel server, text modelGPU + the model
    🔗
    One shape, different machines
    In our setup the pattern works on two different pairs of machines. The worker node doesn't even need a file-sync platform — a watched folder and a container are enough.
  3. Connecting over the private network

    The worker reaches the brain by the brain's private address in the mesh network — nothing is exposed publicly. The container starts with --network host so it uses the machine's network and can reach the model's port. The brain must listen on a reachable address, not only on localhost — this is set with OLLAMA_HOST in the service configuration.

    bash · on the brain: listen on the private-network address
    sudo systemctl edit ollama
    # in the file that opens:
    # [Service]
    # Environment="OLLAMA_HOST=<brain-mesh-address>:11434"
    sudo systemctl restart ollama
    bash · on the worker: check and run
    # check that the brain is visible
    curl -s http://<brain-mesh-address>:11434/api/version
    # run the container, pointing at the model on the other node
    docker run --rm --network host \
      -e OLLAMA_URL=http://<brain-mesh-address>:11434 -e MODEL=<text-model> \
      -v /path/input:/data doc-ocr /data
    ⛔
    Don't open the model to everyone
    The model's local API doesn't ask for a login — it answers anyone who can reach it over the network. Set the address in the private network, not "all interfaces", and don't forward the port to the internet.
    ⚠️
    Only the worker needs Docker
    The brain doesn't need Docker — only the model server has to listen. Don't overcomplicate it: if the site has no file sync, a watched folder and a container are enough.
  4. Why it is a pattern, not a one-off setup

    Once described, the pattern travels: change the brain's address and the folder — and you have a new "reader" in a new place. One GPU node serves many workers, with no new graphics cards.

    ✅
    Ownership ≠ account
    You may use hardware that is yours but sits in someone else's mesh network (for example a partner's). Then you use the node, but you don't touch the network account. The distinction "my hardware ≠ my account" keeps relationships clean.

04Check

1. Which node needs a GPU?

2. How does the worker reach the brain?

3. Which node must have Docker?

4. Why is this a pattern and not a one-off setup?

05What's next

06Sources

  1. Ollama API — GET /api/version to check that the server is visible.
  2. Ollama FAQ — OLLAMA_HOST and configuring the systemd service.
  3. Docker: host network — what --network host does.
  4. Tailscale — one example of a private mesh network.
  5. KAGAMI's own experience: the pattern works on two pairs of machines.