This article is published in English.
Practical notes: Run Hermes Agent Locally (and securely) with Ollama and
Operable walkthrough of Practical notes: Run Hermes Agent Locally (and securely) with Ollama and: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “Run Hermes Agent Locally (and securely) with Ollama and Rootless Podman on Arch Linux”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
1. Install rootless Podman
For the 1 Install rootless Podman stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
sudo pacman -Syu podman
1) crun
2) krun
3) runc
1
podman --version
podman info --format '{{.Host.OCIRuntime.Name}}'
crun
grep "^$USER:" /etc/subuid
grep "^$USER:" /etc/subgid
2. Fix rootless OverlayFS if necessary
For the 2 Fix rootless OverlayFS stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
kernel does not support overlay fs:
'overlay' is not supported over extfs
sudo pacman -S fuse-overlayfs
which fuse-overlayfs
/usr/bin/fuse-overlayfs
~/.config/containers/storage.conf
[storage]
driver = "overlay"
[storage.options.overlay]
mount_program = "/usr/bin/fuse-overlayfs"
podman info --debug | grep -Ei 'graphDriverName|mount_program|overlay'
3. Optional: move Podman’s storage to a larger drive
For the 3 Optional move Podman stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine. For the 3 Optional move Podman stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
write /var/tmp/container_images_storage...:
no space left on device
$HOME/.local/share/containers/storage
/var/tmp
/path/to/large-drive/podman/
├── storage/
└── tmp/
sudo mkdir -p /path/to/large-drive/podman/storage
sudo mkdir -p /path/to/large-drive/podman/tmp
sudo chown -R "$USER:$USER" /path/to/large-drive/podman
mkdir -p ~/.config/containers
nano ~/.config/containers/storage.conf
[storage]
driver = "overlay"
graphroot = "/path/to/large-drive/podman/storage"
[storage.options.overlay]
mount_program = "/usr/bin/fuse-overlayfs"
export TMPDIR=/path/to/large-drive/podman/tmp
echo 'export TMPDIR=/path/to/large-drive/podman/tmp' >> ~/.bashrc
source ~/.bashrc
podman info --format 'GraphRoot: {{.Store.GraphRoot}}'
podman info --debug | grep -Ei 'graphRoot|imageCopyTmpDir|mount_program'
~/.local/share/containers/storage
4. Create the only host folder Hermes will be allowed to access
When working through the 4 Create the only stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
export HERMES_WORKSPACE="$HOME/path/to/Hermes-Workspace"
mkdir -p "$HERMES_WORKSPACE"
Obsidian Vault/
├── Personal/
├── Work/
├── Research/
└── Agent Workspace/ ← only this folder is exposed
podman volume create hermes-data
podman volume create ollama-models
5. Configure AMD GPU access
When working through the 5 Configure AMD GPU stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
/dev/kfd
/dev/dri
ls -l /dev/kfd
ls -l /dev/dri/render*
groups
sudo usermod -aG video,render "$USER"
groups
groups
video render
6. Check /dev/net/tun
When working through the 6 Check dev net stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs. When working through the 6 Check dev net stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
pasta failed with exit code 1:
Failed to open() /dev/net/tun: No such device
ls -l /dev/net/tun
sudo modprobe tun
uname -r
ls /usr/lib/modules/
7. Choose a local model
The 7 Choose a local stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
gemma4:12b
export HERMES_MODEL="gemma4:12b"
8. Pull the container images
The 8 Pull the container stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
podman pull docker.io/ollama/ollama:latest
podman pull docker.io/ollama/ollama:rocm
podman pull docker.io/nousresearch/hermes-agent:latest
9. Download the model into a persistent volume
The 9 Download the model stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. The 9 Download the model stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
podman rm -f ollama-bootstrap 2>/dev/null
podman run -d \
--name ollama-bootstrap \
-v ollama-models:/root/.ollama \
docker.io/ollama/ollama:latest
podman rm -f ollama-bootstrap 2>/dev/null
podman run -d \
--network host \
--name ollama-bootstrap \
-v ollama-models:/root/.ollama \
docker.io/ollama/ollama:latest
podman exec -it ollama-bootstrap \
ollama pull "$HERMES_MODEL"
podman exec ollama-bootstrap ollama list
gemma4:12b
podman rm -f ollama-bootstrap
container name "ollama-bootstrap" is already in use
podman rm -f ollama-bootstrap
10. Create an internal-only network
For the 10 Create an internal-only stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
--network none
podman network create \
--ignore \
--internal \
hermes-internal
podman pod create \
--name hermes-local \
--network hermes-internal \
--userns=keep-id:uid=10000,gid=10000
11. Start Ollama with AMD ROCm
For the 11 Start Ollama with stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
podman run -d \
--name ollama \
--pod hermes-local \
--device /dev/kfd \
--device /dev/dri \
--group-add keep-groups \
-e HOME=/root \
-e OLLAMA_MODELS=/root/.ollama/models \
-e OLLAMA_CONTEXT_LENGTH=64000 \
-v ollama-models:/root/.ollama \
docker.io/ollama/ollama:rocm
--device /dev/kfd
--device /dev/dri
--group-add keep-groups
ollama/ollama:rocm
HOME=/root
OLLAMA_MODELS=/root/.ollama/models
podman exec ollama ollama list
podman logs ollama
HOME=/
OLLAMA_MODELS=/.ollama/models
/root/.ollama/models
HOME=/root
OLLAMA_MODELS=/root/.ollama/models
podman exec ollama ollama list
12. Verify ROCm is actually using the GPU
For the 12 Verify ROCm is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine. For the 12 Verify ROCm is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
podman logs ollama 2>&1 | grep -Ei 'gpu|rocm|amd|gfx'
podman exec -it ollama \
ollama run gemma4:12b \
"Reply with exactly: AMD GPU test successful"
podman exec ollama ollama ps
PROCESSOR
CONTEXT
64000
13. Start Hermes with exactly one host-folder mount
When working through the 13 Start Hermes with stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
podman run -d \
--name hermes \
--pod hermes-local \
--security-opt=no-new-privileges \
--pids-limit 512 \
-v hermes-data:/opt/data \
-v "$HERMES_WORKSPACE:/opt/data/workspace:rw,nodev,nosuid" \
-w /opt/data/workspace \
docker.io/nousresearch/hermes-agent:latest \
sleep infinity
/:/host
/home:/home
~/.ssh
~/.config
Docker socket
Podman socket
$HERMES_WORKSPACE
↓
/opt/data/workspace
hermes-data
↓
/opt/data
14. Verify the mount boundary
When working through the 14 Verify the mount stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
podman inspect hermes \
--format '{{range .Mounts}}{{println .Type .Source "->" .Destination}}{{end}}'
Podman volume -> /opt/data
your allowed folder -> /opt/data/workspace
podman exec hermes sh -c \
'test ! -S /var/run/docker.sock && echo "No Docker socket exposed"'
podman exec hermes sh -lc \
'test -e "$HOME/.ssh" && echo "SSH directory visible" || echo "Host SSH directory not visible"'
15. Verify Hermes can reach Ollama
When working through the 15 Verify Hermes can stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs. When working through the 15 Verify Hermes can stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
podman exec hermes python -c \
'import urllib.request; print(urllib.request.urlopen("http://127.0.0.1:11434/v1/models").read().decode())'
podman exec hermes python - <<'PY'
...
PY
podman exec -i hermes python - <<'PY'
import urllib.request
print(
urllib.request.urlopen(
"http://127.0.0.1:11434/v1/models"
).read().decode()
)
PY
16. Verify that external network access is blocked
The 16 Verify that external stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
podman exec hermes python -c \
'import urllib.request; print(urllib.request.urlopen("https://example.com", timeout=5).read())'
podman ps
11434/tcp
0.0.0.0:11434->11434/tcp
17. Configure Hermes to use local Ollama
The 17 Configure Hermes to stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
podman exec -it \
--user 10000:10000 \
-w /opt/data/workspace \
hermes \
hermes model
Ollama Cloud
Custom endpoint (enter URL manually)
http://127.0.0.1:11434/v1
gemma4:12b
64000
18. Use Hermes’ local terminal backend — inside the container
The 18 Use Hermes local stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. The 18 Use Hermes local stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
terminal:
backend: local
Hermes "local"
↓
Hermes container
Hermes "local"
↓
your Arch workstation
podman exec --user 10000:10000 hermes \
hermes config set terminal.backend local
podman exec --user 10000:10000 hermes \
hermes config set terminal.cwd /opt/data/workspace
podman exec --user 10000:10000 hermes \
hermes config set terminal.home_mode profile
19. Launch Hermes
For the 19 Launch Hermes stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
podman exec -it \
--user 10000:10000 \
-w /opt/data/workspace \
hermes \
hermes
20. Start the environment automatically at boot
For the 20 Start the environment stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
mkdir -p ~/.config/systemd/user
nano ~/.config/systemd/user/hermes-local.service
[Unit]
Description=Hermes Local AI Pod
After=default.target
[Service]
Type=oneshot
RemainAfterExit=yesExecStart=/usr/bin/podman pod start hermes-local
ExecStop=/usr/bin/podman pod stop -t 30 hermes-localTimeoutStartSec=120
TimeoutStopSec=60[Install]
WantedBy=default.target
systemctl --user daemon-reload
systemctl --user enable hermes-local.service
systemctl --user start hermes-local.service
systemctl --user status hermes-local.service
Active: active (exited)
sudo loginctl enable-linger "$USER"
loginctl show-user "$USER" -p Linger
Linger=yes
podman ps
podman exec -it \
--user 10000:10000 \
-w /opt/data/workspace \
hermes \
hermes
Troubleshooting the things that actually went wrong
For the Troubleshooting the things that stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine. For the Troubleshooting the things that stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Problem:
kernel does not support overlay fs
Fix:
Install fuse-overlayfs and configure it as the overlay mount program.
Problem:
no space left on device under /var/tmp
Fix:
Move Podman's graphroot if needed AND set TMPDIR.
Changing Podman's --tmpdir is not the same thing.
Problem:
ollama-bootstrap name already in use
Fix:
podman rm -f ollama-bootstrap
Problem:
pasta cannot open /dev/net/tun
Fix:
sudo modprobe tun
If the installed modules don't match the running kernel, reboot.
Problem:
Installed nvidia-container-toolkit on an AMD machine
Fix:
Don't.
Use ollama/ollama:rocm with /dev/kfd and /dev/dri.
Problem:
Added myself to video/render but `groups` still didn't show them
Fix:
Log out and back in.
The existing login session retains its original supplementary groups.
Problem:
Ollama model files and manifest exist, but `ollama list` is empty
Fix:
Check:
podman logs ollamaIf Ollama is using:
OLLAMA_MODELS=/.ollama/modelswhile the volume is mounted under:
/root/.ollamaset explicitly:
HOME=/root
OLLAMA_MODELS=/root/.ollama/models
Problem:
Hermes → Ollama Python test silently prints nothing
Fix:
If using `python -` with a heredoc, add `podman exec -i`.
Or simply use `python -c`.
Problem:
The Hermes provider menu doesn't contain "local Ollama"
Fix:
Choose:
Custom endpoint (enter URL manually)Then:
http://127.0.0.1:11434/v1
Problem:
Everything works, but Hermes reports an inadequate context window
Fix:
Set OLLAMA_CONTEXT_LENGTH=64000 server-side and configure Hermes for the same value.
Verify the real allocation with:
ollama ps
Why you prefer this over simply trusting the agent
When working through the Why you prefer this stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
Prompt restrictions
↓
Hermes file write protections
↓
Hermes container filesystem
↓
Rootless Podman user namespace
↓
Host filesystem permissions
A note about fast-moving local AI tooling
When working through the A note about fast-moving stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
Does Podman see the GPU?
↓
Does Ollama see the model?
↓
Does Ollama actually use the GPU?
↓
Is the context really 64K?
↓
Can Hermes reach /v1/models?
↓
Can Hermes perform an actual file operation?
↓
Can Hermes reach anything it shouldn't?
Final result
When working through the Final result stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs. When working through the Final result stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Local inference YES
AMD GPU acceleration YES
Hermes persistent memory YES
One writable host workspace YES
Cloud LLM required NO
Host filesystem exposed NO
Podman/Docker socket exposed NO
Normal Internet egress NO
Entire Obsidian vault exposed NO
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for bab1ff410bd9: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.