Installing Docker, Ollama, and Hermes, and the 64K Context Trap
Assembling Docker, Ollama, and Hermes on the hardened VPS, and discovering that a model 'compatible with Ollama' isn't necessarily usable as-is by an agentic orchestrator. Third article in the series.
By Nicolas Cousin — Published on September 29, 2026
Installing Docker, Ollama, and Hermes, and the 64K Context Trap
TL;DR
On the hardened VPS from the previous article, I installed Docker, created a private network so containers can talk to each other without exposing anything publicly, then launched Ollama and Hermes on it. First trap I hit: a model Ollama lists as compatible isn't necessarily compatible with an orchestrator like Hermes, which can demand more context than the default. Extending that context is possible, but it isn't free.
Table of contents
- A private Docker network, no public access
- Ollama: no port exposed
- Hermes in Docker
- The trap: what a model advertises isn't what you can use
A private Docker network, no public access
With Docker and Compose installed, the first thing I do before launching a single container is create a dedicated network:
docker network create agentic
Every container in the agentic stack (Ollama, Hermes, and whatever comes next) joins this network. They can see each other by container name, without any port published to the outside. The hardened VPS from the previous article (key-only SSH, UFW, Tailscale) stays the only way in; Docker doesn't accidentally open a second one.
Ollama: no port exposed
docker run -d --name ollama \
--network agentic \
-v ollama-data:/root/.ollama \
ollama/ollama
No -p in this command, deliberately. Ollama is only reachable from other
containers on the agentic network, at http://ollama:11434, resolved by
Docker's internal DNS. No port is open on the VPS's public interface. The
ollama-data volume persists downloaded models across container restarts.
Hermes in Docker
docker run -d --name hermes \
--network agentic \
-v ~/.hermes:/opt/data \
hermes-local
Same logic: Hermes joins the agentic network and exposes nothing
publicly. The ~/.hermes:/opt/data mount persists its configuration and
history on the host side, independent of the container's lifecycle.
Communication follows this path:
Hermes → agentic → Ollama:11434
Hermes calls Ollama by its container name on the private network, never by a public IP.
The trap: what a model advertises isn't what you can use
With Ollama and Hermes in place, first model tested, first snag: the
context was simply too short for what Hermes was trying to fit into it,
once you add up the system prompt, tool definitions, and conversation
history. A model like qwen2.5-coder:7b has a native context of 32768
tokens on Ollama's side, and some Hermes configurations demand more than
that to work correctly.
"Compatible with Ollama" doesn't mean "directly usable by the orchestrator calling it." The context actually available depends on how the model is loaded, not just on what the model advertises it can handle.
The fix exists on Ollama's side: create a model variant with an extended
context via a Modelfile:
FROM qwen2.5-coder:7b
PARAMETER num_ctx 65536
ollama create qwen2.5-coder:7b-64k -f Modelfile
Technically, it works: the new variant accepts a context of 65536 tokens.
But it's not a free setting. A larger context consumes more memory, and on
a CPU-only VPS like mine, every extra token in the context costs compute
time. Bumping num_ctx to 64K says nothing about how fast that context
will actually be processed, or about result quality once you approach that
limit: "model available in Ollama," "context configured," and "model
actually usable agentically" remain three different things, and confusing
the first two with the third cost me time I didn't need to lose.
Docker, Ollama, and Hermes are now running on the private agentic
network, with a context configured to hold the load. What's left to find
out is whether what I gained in context, I didn't lose in speed: on to the
first real benchmark, measurements included.