Two nodes: worker + brain
A reusable pattern: one node runs the worker (OCR in Docker), the other runs the model; they talk over a private network. That way every site becomes a working "reader" without duplicating the expensive GPU node.
01What you will learn
- Why reading a document doesn't need a GPU but thinking does — and how that splits the system into two roles.
- How the worker finds the brain over a private network, with nothing exposed publicly.
- What each of the two nodes must have — and what it must not.
- Why this is a pattern: changing the address and the folder gives you a new reader in a new place.
02Before you start
- The worker from the lesson I19 · OCR→LLM — the container that reads documents.
- A model server with the text model on the GPU machine (how a model is moved — in I21).
- Both machines in one private mesh network (for example Tailscale or your own WireGuard).
- Docker on the worker machine.
03Steps
-
The idea: separate the worker from the brain
A GPU is expensive and scarce. But reading a document doesn't need a GPU — the OCR and the orchestration run on a cheap node with an ordinary processor. Only the thinking (the model) needs a GPU. So we split: the worker (Docker + OCR) sits where the documents are; the brain (the model server) sits on the GPU machine. One brain serves many workers.
diagram[ worker: Docker + OCR ] --private network--> [ brain: model server :11434 + GPU ] cheap node, watched folder GPU machine, text model -
The two roles
Role What it runs Requires Worker OCR container, watched folder Docker Brain model server, text model GPU + the model 🔗One shape, different machinesIn our setup the pattern works on two different pairs of machines. The worker node doesn't even need a file-sync platform — a watched folder and a container are enough. -
Connecting over the private network
The worker reaches the brain by the brain's private address in the mesh network — nothing is exposed publicly. The container starts with
--network hostso it uses the machine's network and can reach the model's port. The brain must listen on a reachable address, not only on localhost — this is set withOLLAMA_HOSTin the service configuration.bash · on the brain: listen on the private-network addresssudo systemctl edit ollama # in the file that opens: # [Service] # Environment="OLLAMA_HOST=<brain-mesh-address>:11434" sudo systemctl restart ollamabash · on the worker: check and run# check that the brain is visible curl -s http://<brain-mesh-address>:11434/api/version # run the container, pointing at the model on the other node docker run --rm --network host \ -e OLLAMA_URL=http://<brain-mesh-address>:11434 -e MODEL=<text-model> \ -v /path/input:/data doc-ocr /data⛔Don't open the model to everyoneThe model's local API doesn't ask for a login — it answers anyone who can reach it over the network. Set the address in the private network, not "all interfaces", and don't forward the port to the internet.⚠️Only the worker needs DockerThe brain doesn't need Docker — only the model server has to listen. Don't overcomplicate it: if the site has no file sync, a watched folder and a container are enough. -
Why it is a pattern, not a one-off setup
Once described, the pattern travels: change the brain's address and the folder — and you have a new "reader" in a new place. One GPU node serves many workers, with no new graphics cards.
✅Ownership ≠ accountYou may use hardware that is yours but sits in someone else's mesh network (for example a partner's). Then you use the node, but you don't touch the network account. The distinction "my hardware ≠ my account" keeps relationships clean.
04Check
1. Which node needs a GPU?
2. How does the worker reach the brain?
3. Which node must have Docker?
4. Why is this a pattern and not a one-off setup?
05What's next
06Sources
- Ollama API —
GET /api/versionto check that the server is visible. - Ollama FAQ —
OLLAMA_HOSTand configuring the systemd service. - Docker: host network — what
--network hostdoes. - Tailscale — one example of a private mesh network.
- KAGAMI's own experience: the pattern works on two pairs of machines.