Two abstract geometric figures face each other over a grid of small blank squares, the darker angular one striking through several while a faint six-node arc passes behind them, in muted teal, amber and cream tones.

Your coding agent, but as an auditor

security · · 6 min read

Cloudflare's security-audit skill turns Claude Code or Codex into a six-phase auditor with a coverage ledger and a second agent whose only job is to kill the first one's findings. This is what a run looks like.

by Colin Domoney

I have typed “find security issues in this repo” into a coding agent more times than I would like to admit. What comes back is always the same: a page of hedged maybes, every one of them “possibly” or “could in theory”, and a warm feeling that a review has taken place. Cloudflare has the line for it: “Ask a model to find bugs, and it will find them, whether the code has any or not.”

So when they open-sourced the security-audit skill that seeded their internal vulnerability harness, I was less interested in the prompt and more interested in what it turned out to be: a job description. Six phases, a set of files the agent has to produce, and a rule that no agent gets to mark its own homework. It installs into the same skills directory Claude Code and Codex already read from, and you invoke it with one sentence:

Security audit this codebase. Full audit mode. All six phases.

The “full audit mode” matters. The skill is guidance by default; loading it does not authorise the whole workflow or any file creation. Say the words and it goes to work.

What a run looks like

Recon. Parallel research agents read the repo and write architecture.md, then coverage-ledger.json. This is the map. Nothing hunts until the system has been described, which is the opposite of what a bare agent does when it grabs the first suspicious-looking function and fills its context window with a sliver of the codebase.

Coverage-led hunting. Isolated hunters are assigned units from the ledger, each working an attack class rather than a vibe. The packs shipped with the skill cover web and auth, memory safety, AI and LLM issues, client-side, supply chain, cloud and IAM, RPC and messaging, resource exhaustion, tenant isolation, and desktop, mobile and local IPC. A hunter has to state its threat model before it files anything. Coverage critics run alongside looking for what the hunters missed.

Adversarial validation. Every candidate is handed to a fresh verifier whose only job is to disprove it. The README puts it plainly: “The agent that checks a finding is never the agent that found it.” Source inspection is read-only, and target code only runs inside an OS-enforced sandbox with no network, an empty allowlisted environment and hard resource limits. If that sandbox is not available, the skill does not improvise; the lead stays as needs_validation with a safe validation plan attached.

Structured output. Survivors land in findings.json as confirmed, needs_validation or rejected, and the file is mechanically checked against report-schema.json with a zero-dependency Node validator. The coverage claim gets the same treatment.

Independent re-check. Fresh agents re-verify the source claims in every record. Not the hunters, not the first verifiers. New eyes.

Report. REPORT.md, FINDINGS-DETAIL.md and NEEDS-VALIDATION.md are derived from the findings, and everything is written to ~/security-audit-skill/<repo-name>/run-1 by default. Outside the repo, unless you explicitly point it at an ignored directory. Small thing, but I have seen enough tooling that litters a working tree to appreciate it.

Verdicts that mean something

This is where the skill earns its keep. A confirmed finding carries a complete source trace, a bounded observed result, a working proof of concept against untouched code and a proposed patch. A needs_validation record names the exact fact that is unresolved and carries no severity at all. A rejected record says what was tried and why it failed.

Two rules sit underneath that. “Severity requires impact.” And “Defense-in-depth gaps are not vulnerabilities.” A missing header or a permissive default is a hardening note, not a CVE, and the report treats it as one. Anyone who has triaged a scanner dump will feel the relief in that sentence. I have written before about wanting triage that fits inside the pull request; this is the same instinct applied to an agent that would otherwise happily invent work for you.

Run it twice

Cloudflare volunteers a number most vendors would bury: in their test runs a single pass found roughly half of the vulnerabilities that repeated runs found in total, and the half it found skewed toward the simpler and less subtle. Rather than pretend otherwise, the skill is built around it. Run two reads the ledger and findings from run one and goes for the gaps instead of re-finding the same easy hits. Runs are additive. Treat the first pass as a strong first cut and the ledger as the thing you keep pushing until it stops moving.

They are equally blunt about what they will not claim. There is no labelled set of every real bug in a codebase, so any recall figure is speculative and they decline to give one. The metric they care about is keeping the number of unconfirmed findings put in front of a human as close to zero as possible.

Where it came from

The skill started as about 450 lines that Cloudflare ran on one repository, adjusting the prompts until it surfaced real bugs. About six weeks later each phase had been lifted almost one-to-one into its own agent, with a database behind it and an orchestrator in front, scanning 128 repos. That harness is not what has been released; the public skill is the single-repo seed it grew from, and the blog post says the harness itself will “hopefully” follow. Their advice for anyone tempted to build the big thing first: start with a skill, get the prompts working, and only add the next stage when its absence is the specific thing slowing you down.

What it’s actually about

The sub-agents are plumbing. What Cloudflare actually shipped is a refusal: no finding without a trust-boundary break, no finding without a source trace, no finding that has not had a second agent try to kill it. Generic prompts produce a faster way to make junk. This produces a report you can hand to someone and stand behind.

Pick a repo you know well. Run it this week. Then run it again.

Installing it

npx skills add https://github.com/cloudflare/security-audit-skill --skill security-audit --global

You need an agent whose model supports tool use and parallel sub-agents, Node.js for the validators, and ideally an OS-enforced sandbox. Without the sandbox it still runs; it just keeps its hands off the target’s code. MIT licensed, at cloudflare/security-audit-skill, with the provenance story in Build your own vulnerability harness.

Stay in the loop

#_

Writing worth reading

I write about security, AI, and occasionally cycling. No spam, no pitches — just things I find interesting, when I find them interesting.

Related posts

I Didn't Want to Give My AI Agent SSH

code · · 7 min read

I Didn't Want to Give My AI Agent SSH

Built to destroy itself

security · · 9 min read

Built to destroy itself

Your coding agent, but as an auditor

security · · 6 min read

Your coding agent, but as an auditor

Cooked to within an inch of its life

code · · 10 min read

Cooked to within an inch of its life