code · · 7 min read
Seven hundred and thirty-one Semgrep findings, five questions each, classified in fifteen seconds for five cents. I have been turning that number over since Saturday, because it changes where triage can live rather than just what it costs.
The model is Jev, from TypeSafe, which launched on Thursday. It is not a chat model. You hand it a piece of structured state and a set of typed questions, each with a closed set of answers you define, and it returns a probability for every answer. No prose, no reasoning to read, no JSON to parse. I wrote about the idea on Friday in Type-safe does not mean correct, and the caveat in that title still holds. This post is about the economics, which turned out to be the more interesting part.
Launch pricing is 4.2 cents per million input tokens and output is free. That is a vendor figure, and one the vendor itself hedges; I would not build a budget on it surviving contact with demand. TypeSafe quotes 70 to 500 milliseconds end to end, measured from the US West Coast where the service runs.
The latency figures below are mine, measured from London. A single call took a median of 309 milliseconds, of which the server reported about 130; the rest was the Atlantic and TLS. Asking eight questions instead of one cost 24 percent more time, not eight times more, because every question in a request is evaluated in parallel against the same state. Twenty calls at a concurrency of eight ran 4.6 times faster than the same twenty in sequence, with per-call latency unchanged, and nothing rate-limited me. The published limits during early access are 250,000 tokens per second and 1,200 requests per minute, which is a long way above anything a single repository’s Semgrep output will produce.
The default, at every company I have worked with, is that SAST output goes somewhere else. A nightly scan, a dashboard, a queue that somebody is nominally responsible for. The developer who wrote the code sees the finding days later, if at all, stripped of the context they had when they wrote it. The queue grows, the team learns which rules are noise, and eventually everyone learns to ignore the tool. That is not a detection failure. It is a dismissal failure, and it happens because dismissing a false positive costs an engineer-minute and there are hundreds of them.
Put the latency and the price together and that constraint goes away. Triage fits inside the pull request check, after Semgrep and before a human, at a cost that does not appear on an invoice. The finding arrives with a probability attached while the author still remembers what the code does.
The corpus was six deliberately vulnerable applications: Juice Shop, WebGoat, crAPI, NodeGoat, VAmPI and DVWA, scanned with fourteen registry rulesets, 731 unique findings after deduplication. Each finding went to Jev as an object: the rule, its message, the matched line, fifteen lines of code either side, and a handful of flags.
Of the 731, 265 cleared the threshold for reporting, 127 were suppressed at high confidence, and 339 went to a review queue with the full distribution attached. Two patterns show what the middle bucket is for.
The 37 sample JWT tokens sitting in documentation and tests were suppressed at a mean probability of real of 0.03. Nobody needs to look at those.
The 90 hits from a Django template rule firing on WebGoat and NodeGoat HTML, neither of which is Django, landed in review at 0.56. Not confident enough to report, not confident enough to bin. That is the right answer. It takes a human thirty seconds each to see the pattern, or one glance at the queue to see all ninety and kill the rule for that repo, rather than an engineer discovering it on their own three months in.
Coverage. Semgrep found 22 of the 113 vulnerable locations I had labelled in DVWA, and 9 of 33 in Juice Shop. Everything it missed was logic: missing authorisation checks, absent rate limits, a CAPTCHA that is never re-verified, a chatbot that takes instructions from users. That is consistent with the wider literature, where static analysis misses somewhere between 47 and 80 percent of vulnerabilities under test conditions and struggles most with anything that depends on developer intent. No classifier can rank a finding that was never reported. Cheap triage of a thin list is still a thin list, and a very clean queue makes it easier to forget how much never entered it.
The other thing it does not change is the need for a labelled sample of your own. The thresholds I set came from one PHP application with known-safe controls. I would not carry them to a Java codebase without checking, and the model version is pinned next to them because a new model moves every probability. On TypeSafe’s own benchmark Jev sits at about 68 percent accuracy, tied with a mid-tier LLM and five or six points behind the top models, and there is no independent reproduction yet. A calibrated probability is a well-formed answer, not a correct one. It is just an answer you can measure.
If you already pay for Semgrep, you already have the data. Take a year of historical findings that carry a human disposition, run them through with your own questions, and see whether the ranking holds against what your people decided. At these prices that costs about the same as the coffee you drink while it runs, and it tells you whether the thresholds are yours or mine.
The engineer-minute spent dismissing a false positive has been the accepted overhead of static analysis for as long as I have worked with it. When that minute becomes a quarter of a second and a rounding error, the overhead stops being a cost and starts being a choice.