code · · 7 min read
by Colin Domoney
This post was written by a pipeline I built last week. Five Hermes Agent profiles, one job each, coordinated through Hermes’ own Kanban board. I did not type a sentence of it. I did tap a phone four times, and those four taps are the only part of the system I actually want to talk about.
The cast, briefly. desk takes a drop from my phone, files it as a brief and parks the gates. scout researches and refuses any claim it cannot attach a URL to. pen drives a writing session with no tools and never types a sentence itself. sub reviews on a different model vendor and is allowed exactly two words: pass or fail. press stages the post, builds the site, opens the PR, makes the cover and posts to Buffer. Three posts went through in the first week. This is the fourth, and it is about the pipeline.
There are two obvious ways to put a human in a system like this, and I did not want either.
The first is the notification. A worker finishes a stage, fires a message into Slack or Telegram, and carries on. Somebody reads it, or does not. The system has no idea which. The “human in the loop” is a hope expressed as a chat message, and the loop closes whether or not the human turned up.
The second is to skip the human altogether. Let it publish, review afterwards, fix in a follow-up commit. This is what “fully autonomous” means in most of the demos, and for a scratch repo I have no objection. It is my name at the top of this blog. I can live with an agent drafting under it, researching under it, even arguing with me under it. Publishing under it without asking is the one thing I drew a line around before I wrote a single prompt, and the line has not moved since.
What Scribe does instead is dull, and that is the point. Every human decision is a task on the board. Not a notification about a task. A task.
The Hermes Kanban moves a task through triage | todo | ready | running | blocked | review | done | archived, and it gives workers a kanban_block tool. So the gate is a task that runs, immediately blocks itself, and sits there. Four of them per post: angle, approve, cover, merge. The board records parent-to-child links, and a child only promotes to ready when every parent is done. Nothing downstream of a gate can start until I unblock it. The pipeline is not asking my permission. It is structurally incapable of proceeding without it.
A few choices that made this hold together:
I did not invent the pattern, which is reassuring. LangGraph calls it an interrupt: interrupt() “pauses graph execution at specific points and waits for external input before continuing,” and when it fires, “LangGraph saves the graph state using its persistence layer and waits indefinitely until you resume execution.” Their canonical example is a node that asks “Do you approve this action?” and resumes only when the caller answers. That is the approve gate, in a framework I am not using.
The Hermes docs draw the same line from the other direction. They contrast the Kanban with plain subagent delegation: delegate_task is a single fire-and-forget RPC with no resumability and no human in the loop, whereas the board is “a durable message queue + state machine,” resumable across crashes, where a human can comment or unblock at any point. I have been banging on about this shape for a while, so it was pleasant to find the vendor saying it in their own documentation.
The first version did not hold the gates on this board at all. I tried to run the coordination through Linear, because it was already there and had a nice app. The gates became tickets in someone else’s SaaS, the workers were polling something they did not own, and it was the wrong shape within days. Then the first native install broke on runtime drift and went back into a pinned container, which is where it still runs. And a small model on the front desk still loops when a message arrives without its header, which I have not fixed and am not pretending I have.
None of that is the gates failing. The gates are the bit that has not moved.
The interesting engineering in an agent pipeline is not the automation. Anyone can chain five prompts. The interesting part is deciding where you refuse to automate, and then making that refusal a first-class state in the system rather than a message hoping someone reads it. Anthropic’s own guidance ends with the line that “human review remains crucial for ensuring solutions align with broader system requirements,” even when the automated checks pass. I agree, with one amendment: the review has to be somewhere the system cannot route around.
Four taps per post, and no fewer on purpose: the angle, the approval, the cover, and the merge. Each is a decision I want to own; everything between them is not.
The board runs itself right up to the moment a judgement is needed. Then it stops, and it stays stopped until I say otherwise. That is not a limitation I am working around. It is the design.