39ms, Open Source, and It Never Chats: Cloudflare's Clef and the Rise of Decision Models
October 3, 2026 · 5 min read
While everyone was racing to build chattier AI, Cloudflare shipped a model that refuses to chat. Clef doesn't generate a single word of text — it only makes decisions: yes or no, A or B, each with a probability attached. And Cloudflare says it runs 2.5 times faster than the startup that invented this playbook.
On October 1, Cloudflare released Clef and Clef-flash, the first models trained by its own Workers AI team, and they walk the "decision model" line exactly: instead of taking a prompt and writing back a paragraph, Clef reads an input state plus a set of typed questions and returns the probability of every allowed answer. The agent doesn't parse anything — it just acts: route the ticket, block the request, escalate to a human. The weights are open source under Apache 2.0, on Hugging Face and hosted on Workers AI itself. Cloudflare claims a median latency of 209 milliseconds for Clef, about 2.5x faster than TypeSafe's Jev, and the flash version drops to 39 milliseconds.
Why the agent era needs models that only judge
Think about what an agent actually does all day. Most of its "thinking" isn't creative writing — it's thousands of tiny judgments: is this request an attack, is this customer angry enough to need a human, is this transaction fraud? Forcing a giant chatbot to answer each one is like hiring a poet to stamp parking tickets. It's expensive, it's slow, and worse, it's inconsistent: ask the same question twice and the wording drifts.
A decision model strips all of that away. No prose, no pleasantries — just calibrated probabilities you can plug straight into code. When the stakes are "block this request or let it through," you don't want eloquence. You want a number, in 39 milliseconds, every single time.
Jev: the startup that started a stampede
Two weeks earlier, a startup called TypeSafe released Jev — a model that doesn't chat and doesn't generate text, only making calibrated decisions: yes/no, scores, pick-between-two. On October 3, the Wall Street Journal reported that Jev is setting off an "LLM replacement" debate in Silicon Valley, and the copycats are already showing up. Founder Diogo Almeida, a former OpenAI employee, claims 25% of Fortune 500 companies are already using it, processing roughly one trillion tokens a day.
The Information reports the company is now in talks for a new funding round above $1 billion, at a potential valuation over $10 billion — for a company that is, essentially, selling judgment without conversation. When Cloudflare drops a free, open-source answer two weeks later, you can read the calendar as the argument: the incumbents are not going to let one startup own this lane.
Open vs. closed: the old fight, a new battlefield
Here's the twist people keep missing: this is the first AI race where the open option arrived almost immediately. Clef is Apache 2.0 — download it, self-host it, fine-tune it, and never send your data to anyone. For banks, hospitals, and anyone whose data can't leave the building, that's not a philosophical preference; it's the whole game. Jev's pitch is performance and polish behind an API. Clef's pitch is: it's yours, it's free, and it's fast enough to run at the edge.
What it means for you
As a user, you'll never notice decision models — that's the point. The chatbot you talk to stays chatty. But behind it, more of the invisible machinery — fraud checks, moderation, routing, triage — will quietly switch from giant chat models to small, fast judges. If you're a developer, the question to start asking is brutal and simple: for each AI call in my product, do I need prose — or just a decision? Every "just a decision" you find is money and latency you get back.