A conductor in a black tailcoat, seen from behind, standing on a wooden podium in a dark concert hall with baton raised. Three music stands face the podium, each holding a small grey box lit by a soft blue glow. Three overlapping rings of pale light hang in the air above the stage.

The stupidest possible way to build a K3s cluster

ai · · 8 min read

Hermes, running inside Herdr, drove three remote Claude Code agents to tear down and rebuild a K3s cluster. It was a daft way to do it. That was the point.

by Colin Domoney

I have a perfectly good Ansible playbook for K3s. I have shell scripts. The K3s installer itself is a one-liner that has never once let me down. If you asked me the correct way to stand up a three-node cluster on a Wednesday evening, I would tell you to run the script and go and make a cup of tea.

So, naturally, I did not do that. I gave the job to an AI agent sitting on my Mac, which then went and found three other AI agents sitting on the three servers, and asked them to do it instead.

This is, by any sensible measure, the stupidest possible way to install K3s. The point was never the cluster. The point was to find out whether two tools I have been slowly falling for could handle something with real consequences: root on three machines, a datastore, a network overlay, and a live chance of destroying something I cared about.

They could. And the way they did it has changed my sense of what I can get done in a week.

The two things I have caught religion about

The first is Herdr. The shortest description is tmux rebuilt around coding agents, and like most short descriptions it undersells it. Herdr is a single Rust binary running as a persistent server, with real terminals living inside it. Close the laptop lid, drop the network, walk away for a weekend; the agents in those panes keep working, and you attach again from wherever you happen to be. I have picked up a running Claude Code session from my phone over Tailscale and carried on as if nothing had happened.

It also knows what is in each pane: “Claude Code, working” or “Codex, blocked, waiting for you”. When an agent stops and needs a human, Herdr tells you. That sounds trivial until you have had five tabs open and no idea which one has been sitting on a permission prompt for twenty minutes. Since 0.9 one client can show local and remote machines together, so a box in the loft and a box in a datacentre sit in the same sidebar.

And the whole thing is scriptable through a CLI and a socket API. That is the part that matters for this story. A pane stops being a terminal you look at and becomes a terminal another program can prompt, wait on and read back.

The second is Hermes Agent from Nous Research, which has been my orchestration layer for a while. Hermes is the thing I actually talk to. It has memory, it runs tools, it has a view of my whole environment, and I can point it at whichever model suits the job. It is not a coding agent; it is the thing that decides which coding agent does what, and then checks their homework.

Put the two together and you get a control plane for a small herd of remote operators.

The set-up

Three small x86-64 boxes, ser5a, ser5b and ser5c. Each one running its own Herdr server with a Claude Code session live in a pane. On my Mac, Hermes, also inside Herdr. Four AI sessions in total, one of them in charge.

The brief was deliberately dull. Build a clean, reproducible three-node K3s cluster. All three nodes servers, all three voting members of embedded etcd, all three schedulable. Bundled containerd, Flannel, CoreDNS, Traefik, ServiceLB, local-path-provisioner. No Docker, no Calico, no Cilium, no Longhorn, no Rancher, no dashboard, no Argo. Nothing decorative.

And one instruction that turned out to matter more than all the rest: before changing anything, inspect all three machines and report.

Hermes as operator

The first thing Hermes did was work out that the three Claude agents were not where it expected. Herdr scopes agent names to a server, so its local session saw only itself. It reasoned through the topology on its own: SSH to each host, drive that host’s Herdr CLI against its local session, and address the pane that way. Then it renamed each remote agent after its host, so three parallel command streams were not all called w1:p1.

Before touching anything privileged it sent a harmless read-only prompt to one box and read the answer back, proving the whole chain end to end. Only then did the real work start.

The pre-flight inspections ran in parallel. Herdr’s prompt-wait timed out a couple of times while the remote Claude was still busy. Hermes treated the timeout as “the client stopped waiting” rather than “the prompt failed”, checked the agent’s state, waited again, and read the transcript. No duplicate commands. That one distinction is the difference between an orchestrator and a script that retries until something breaks.

The bit where it stopped

The reports came back, and the “clean hosts” were not clean at all.

All three machines were already K3s servers, forming a healthy etcd cluster. The wider cluster had eight nodes, including four Raspberry Pi workers. There were real workloads on it: Prometheus, node exporters, a local-volume PVC, scheduled etcd snapshots. The old install had also left the cluster token in world-readable systemd command lines, which is the sort of thing I would rather find myself than read about later.

An eager automation loop would have followed my original instruction and erased the lot. Hermes did not. It laid out what it had found, offered two paths (preserve and modernise, or destroy and rebuild), and recommended preservation because there was data. It would not proceed until I typed, in capitals, that nothing needed preserving.

I have spent thirty years telling people that the dangerous part of any privileged system is the gap between what the operator intended and what the operator actually said. Watching a language model refuse to close that gap on my behalf was the moment this stopped being a toy.

Tear-down, build, prove

Cleanup ran in parallel, each agent starting with K3s’s own uninstall script and then hunting residue. They left the LVM-backed /var/lib/rancher filesystem mounted and empty for the new install, and did not touch the Pi workers, which were outside scope. Hermes then verified the cleanup itself over SSH rather than believing the agents’ prose.

The build went the other way: parallel where safe, serial where etcd demands it. Hermes pinned the resolved stable version, bootstrapped ser5a, and only when that node was healthy let the other two join as servers. The token went into a root-only file on each host, moved without ever being printed, and never appeared in a transcript.

Twice during this phase an approval expired because I had wandered off. Herdr flagged the blocked agent; Hermes treated the expired approval as a command that had not run, stopped before anything dependent, and waited for me. Nothing was papered over.

Then, rather than declaring victory because k3s.service was active, it wrote a kubeconfig on my workstation and checked from outside the cluster: three Ready nodes, three API endpoints, three etcd voters, every default component healthy. It handed a throwaway workload to one agent that scheduled across all three nodes, pushed traffic across node boundaries, resolved through CoreDNS and served HTTP through Traefik, then cleaned up after itself.

The claim at the end was not “the installer exited zero”. It was that packets and workloads had crossed the system.

What it’s actually about

I could have had this cluster in fifteen minutes with a script. What I could not have had in fifteen minutes is the confidence I now have in it. Every step was inspected before it was taken and verified after, and the write-up ended with an honest list of what had not been checked.

That is the productivity shift, and it is not the one people usually mean. I am not going to build clusters this way every day. But the pattern, one coordinator holding the constraints and treating a set of capable remote operators as fallible witnesses whose claims need checking, is one I can point at almost anything. Herdr makes the operators addressable from anywhere I have SSH, keeps them running when I am not looking, and rings a bell when one of them needs me. Hermes makes them a team. Claude Code makes each one competent at the console it sits at.

The flashy part is that one agent drove three others across three machines. The part I keep coming back to is that it stopped when it should have.

Stay in the loop

#_

Writing worth reading

I write about security, AI, and occasionally cycling. No spam, no pitches — just things I find interesting, when I find them interesting.

Related posts

I Didn't Want to Give My AI Agent SSH

code · · 7 min read

I Didn't Want to Give My AI Agent SSH

Built to destroy itself

security · · 9 min read

Built to destroy itself

Your coding agent, but as an auditor

security · · 6 min read

Your coding agent, but as an auditor

Cooked to within an inch of its life

code · · 10 min read

Cooked to within an inch of its life