Clef and Clef-flash: Cloudflare releases its first in-house models — decision models for agents, not chatbots
🔎 A CDN training its own models: the announcement you shouldn't skim over
On October 1, 2026, right in the middle of Birthday Week, Cloudflare crossed a line that no infrastructure player had crossed before: publishing its first models trained by its own Workers AI team. They're called Clef and Clef-flash. And no, they're not just more chatbots.
They're "decision models": models that generate no text. You give them a state (text, JSON, images, or video) and a schema of typed questions — yes/no, choice from a list, score — and they return a probability for each allowed option. Their sole job: deciding how a software agent acts, at the millisecond level.
The timing is no accident. The wave of System 1 decision models — Laya, Kev, DeepOpen — has been gaining momentum for a few months, driven by small open source teams. With Clef, it's an infrastructure giant, present in millions of projects, that's bringing its firepower to bear. According to the briefing from AI News Online, it's one of the most strategic moves of the season. For anyone building agents, this is an architectural component that just changed status: it's moving from an under-the-radar GitHub project to the catalog of a global provider.
Key takeaways
- Clef (27 billion parameters, Qwen3.8-27B backbone) and Clef-flash (9 billion, Qwen3.5-9B): the first in-house trained models from Cloudflare, released on October 1, 2026 during Birthday Week.
- Decision models, not LLMs: zero text generation. A state as input, typed questions, one probability per option as output.
- Open weights under Apache 2.0 on Hugging Face, inference on Cloudflare's edge GPUs.
- Pricing: $0.24 per million input tokens for Clef, $0.09 for Clef-flash — output tokens not billed (October 2026, check on cloudflare.com).
- Figures announced by Cloudflare: macro-F1 of 94.20 on BANKING77 versus 79.74 for Jev (Typesafe AI); median latency of 38.8 ms for Clef-flash versus 524.1 ms for Jev.
- Jev API compatibility: a migration path planned for the existing ecosystem.
- A self-serve RL fine-tuning service accompanies the release, built on AI Gateway, Containers, and a new Trainer component.
Recommended Tools
| Tool | Primary use | Price (October 2026) | Ideal for |
|---|---|---|---|
| Clef on Workers AI | Multimodal decisions (text, JSON, images, video), 64k context | $0.24/M input tokens, output not billed (verify on cloudflare.com) | Complex agents with rich state |
| Clef-flash on Workers AI | Ultra-fast decisions, 9B model | $0.09/M input tokens (verify on cloudflare.com) | High-frequency arbitration, latency-critical paths |
| Clef weights on Hugging Face | Self-hosting, auditing, in-house fine-tuning | Free (Apache 2.0) | Teams that want to keep control |
| Jev (Typesafe AI) | The market's reference decision model | n/a (see vendor's website) | Comparison and migration |
| Hostinger | Hosting for the orchestrator and agentic stack | VPS from ~$5/month (Oct. 2026, verify on hostinger.com) | Deploy your agents outside the hyperscalers |
What exactly is a "decision model"?
A decision model doesn't write anything: it decides. Where an LLM like GPT-5.5 or Claude Opus 4.7 produces text token by token, Clef scores options. You give it a state of the world and a list of typed questions; it returns, for each question, a probability per allowed option. Nothing else — and that's precisely what makes it valuable.
Concretely, the input can be text, JSON, images, or video. The question schema is typed: boolean (escalate this ticket?), single choice (refund, request information, close), score (confidence level between 0 and 1). Output: probabilities that regular code can work with — thresholds, rules, logs, audit. No more prose to parse, no more risk of a "yes" turning into an ambiguous paragraph.
The technical trick: prefill-only, not generation
Under the hood, Clef is not autoregressive. The so-called "prefill-only" approach scores options instead of generating them: no sequential decoding, so no latency accumulating token after token. This explains both the speed and the bill — output tokens aren't billed, quite simply because there aren't any.
The name itself sums up the philosophy. Cloudflare explains it without detours: "A decision model is analogous to a music clef because it helps define the domain of the context and the subsequent notes (actions) that follow it". The clef defines the context; the notes — the agent's actions — follow.
A concrete example
Let's take a customer support agent. At each step, it queries Clef with the conversation history and a screenshot, along with three typed questions: "escalate to a human?" (yes/no), "priority?" (low/medium/high), "next action?" (refund/request a part/close). The orchestrator reads the probabilities, applies a threshold, logs the decision. No hallucination possible in the control loop — exactly what you'd expect from an agent in production.
The announced numbers — and how to read them
They're impressive on paper — and they come from the vendor itself. Cloudflare claims that Clef beats gpt-oss-120b and Typesafe AI's Jev model on accuracy and latency across dozens of benchmarks. The most telling public figures cover two dimensions.
| Metric | Clef (27B) | Clef-flash (9B) | Jev (Typesafe) |
|---|---|---|---|
| Backbone | Qwen3.8-27B | Qwen3.5-9B | n/a |
| Context window | 64k tokens | n/a | n/a |
| Macro-F1 (BANKING77) | 94.20 | n/a | 79.74 |
| Median latency | n/a | 38.8 ms | 524.1 ms |
| p95 latency | n/a | 122.4 ms | 536.0 ms |
| Input price (/M tokens) | $0.24 | $0.09 | n/a |
Two readings stand out. On accuracy, the gap on BANKING77 — a banking intent classification dataset, close to the intended use case — is massive: 94.20 vs. 79.74 macro-F1 for Clef. On latency, Clef-flash is claimed to be roughly 13× faster than Jev at the median (38.8 ms vs. 524.1 ms), and the gap holds even at p95 (122.4 ms vs. 536.0 ms), according to AI Weekly. Clef-flash is also explicitly positioned for latency-critical paths.
A necessary step back
Keep a cool head: these numbers are published by Cloudflare, on its own hardware. BANKING77 is an excellent proxy for typed questions, but it is not an end-to-end agentic benchmark. And "dozens of benchmarks" without a detailed publication is something that needs verifying. Best practice: wait for independent replication, and above all, test against your own agent traces. Cryptobriefing notes that inference runs on Cloudflare's edge GPUs — a hardware context that is hard to compare directly with that of other players in the market.
Why is a CDN getting into training models?
Because latency is its business — and decision models are the perfect workload for the edge. A 27 or 9 billion parameter model fits on the GPUs Cloudflare has already deployed in its datacenters for Workers AI. No need for massive clusters: what's needed is proximity to the user.
The economic argument is tangible. An agent making hundreds of decisions per minute can't afford a 500 ms round trip to a hyperscaler on every call. With a median latency of 38.8 ms at the network edge, Clef-flash makes viable a pattern that no general-purpose LLM allows: querying the decision model at every step, systematically, without a second thought.
There's also an obvious platform logic. The launch comes with a self-serve reinforcement fine-tuning service, built on AI Gateway, Containers, and a new Trainer component: you start from your traces, you train, and the tuned model is redeployed on Workers AI. The data → fine-tune → inference loop stays entirely within Cloudflare. It's lock-in, but useful lock-in — and it's exactly the kind of loop that keeps teams coming back.
Cloudflare used to sell the web's bandwidth, then that of applications; with Clef, it's selling agents' decisions. Whoever hosts the decisions hosts the agents. This isn't a lab experiment: it's a deliberate business positioning, and Birthday Week merely served as the showcase.
The System 1 wave gains an infrastructure heavyweight
Clef isn't arriving in a vacuum. Over the past few months, a family of System 1 decision models has been taking shape: Laya, an open source 421M-parameter decision model, Kev, the open source System 1 decision model family released by Jared Palmer, as well as DeepOpen. Small, fast models, specialized in deciding — not in writing.
hwchase17's reaction, relayed by daily.dev, sums up the mood: "decision model season." When the creator of LangChain talks about a season, it's because he's seeing a pattern spread at high speed.
What Cloudflare brings to this movement is the channel. Laya and Kev have the ideas and the weights; Cloudflare has millions of developers already on Workers, global edge infrastructure, listed pricing, and managed fine-tuning. Historically, it's exactly these kinds of distribution channels that turn a niche into an industry standard.
My take: the generation/decision decoupling is going to become the default architecture pattern for agents in 2026. The big LLM (GPT-5.5, Gemini 3.1 Pro, Claude Opus 4.7) writes and reasons; the decision model (Clef, Laya, Kev) arbitrates. Two models, two budgets, two latency profiles. Anyone still designing their agents as a single giant LLM call will have to rethink their stack — and will probably come out ahead on both reliability and cost.
Open weights, multimodal, and a price war in the background
Clef is available for download. The weights are open under Apache 2.0 on Hugging Face, which allows self-hosting, auditing, and in-house fine-tuning. A choice that stands in stark contrast to the current U.S. regulatory climate — the White House is now pushing to verify AI models before their release — and that says a lot about where the market is headed: openness is winning over small specialized models first, not frontier models.
On the capabilities side, Clef reads state as text, JSON, images, or video, with a vision encoder limited to 4 images per request and a 64k-token context window. Enough to feed the model with an agent's actual state — screenshots, structured logs, compressed history — without building a heavy preprocessing pipeline.
And the pricing fits squarely into the ongoing battle. $0.24 and $0.09 per million input tokens is flash-model territory, the same ground where DeepSeek V4 now comes in Pro and Flash variants and where GLM-5.3-Flash and Qwen3.8-Flash-Next were released on the very same day. Delicious irony: Clef runs on a Qwen3.8 backbone — the very family of models leading this price war.
A detail that matters for budgets: since Clef generates nothing, output tokens are never billed. The bill for an agent making 500 decisions per minute becomes predictable, bounded by input volume. That's a luxury no generative LLM can offer, and one more argument for taking decisions out of the generative loop.
How to test Clef today
Two main paths, depending on how much control you want — and a third option for demanding teams.
The fastest: Workers AI. The official documentation lists Clef and Clef-flash among the available models. You enable Workers AI, call the model with your state and your typed question schema, and get probabilities back. If you're coming from the Typesafe ecosystem, Jev API compatibility is planned: an endpoint change, not a code rewrite.
The most sovereign: the weights. On Hugging Face, under Apache 2.0. Count on a serious GPU for Clef (27B); Clef-flash (9B) runs on more modest setups. Your orchestrator, meanwhile, only needs a VPS — Hostinger does the job just fine for the stack that surrounds the model (check prices on hostinger.com). And if you run your agents locally, our guide on AI agents with Ollama will show you how to plug a self-hosted instance of Clef-flash into your decision chain.
The most advanced: RL fine-tuning. Cloudflare's self-serve service (AI Gateway + Containers + Trainer) takes your traces, trains via reinforcement, then redeploys the adjusted model on Workers AI. It's the royal road if your decisions have business-specific nuances that the base model doesn't capture — and it's the real product behind the announcement.
An architecture reminder before you dive in: Clef decides, it doesn't generate. Choosing the model that writes and reasons remains a separate problem — our comparison of the best LLMs for AI agents will help you settle it.
❌ Common Mistakes
Mistake 1: Using Clef as a small LLM
Clef doesn't summarize, doesn't write, doesn't answer open-ended questions. Using it to generate text destroys its only advantage: speed and price. The solution: separate the roles. Generation and reasoning go to the big models (GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro); typed decisions go to Clef.
Mistake 2: Swallowing the vendor's benchmarks
The numbers (94.20 macro-F1, 38.8 ms latency) are reported by Cloudflare itself, on its own GPUs. BANKING77 is a classification proxy, not an end-to-end agentic test. The solution: evaluate on your own agent traces before migrating anything, and keep an eye out for the independent replications that are sure to follow.
Mistake 3: Confusing model latency with pipeline latency
38.8 ms is the model's inference at the edge of Cloudflare's network. Your agent, on the other hand, adds up the network, orchestration, calls to the main LLM, and your tools. The solution: measure the p95 of your full pipeline, not the isolated model's — it's the only number your users actually feel.
Mistake 4: Exceeding input limits
Vision encoder limited to 4 images per request, 64k token context: you don't throw in three hours of raw history and twenty screenshots. The solution: compress and structure the state before the call. That's precisely what Clef is designed to read efficiently — it's up to you to feed it clean, not raw.
❓ Frequently Asked Questions
Is Clef free?
The weights are free and open source (Apache 2.0) on Hugging Face, so self-hosting only costs you the infrastructure. On Cloudflare's infrastructure, usage is billed: $0.24 per million input tokens for Clef, $0.09 for Clef-flash, with output tokens not billed (October 2026, check on cloudflare.com).
Can it replace GPT-5.5 or Claude?
No, and that's not its purpose. Clef answers typed questions with probabilities; it doesn't generate text and doesn't reason in natural language. The winning architecture combines both: a large model for writing and reasoning, Clef for fast, cheap arbitration at every step of the agent.
What's the difference with Jev from Typesafe AI?
Same principle — decision models that score options — and API compatibility between the two. Cloudflare claims higher accuracy (macro-F1 of 94.20 vs 79.74 on BANKING77) and much lower latency (38.8 ms vs 524.1 ms median for Clef-flash). These figures remain vendor-side: compare them against your own tests.
Can you fine-tune it on your own data?
Yes, via two paths. Cloudflare's self-serve service trains the model with reinforcement learning from your traces (via AI Gateway, Containers, and the Trainer component), then redeploys it on Workers AI. Or you can download the Apache 2.0 weights and build your own fine-tuning pipeline, with full independence.
Why the name "Clef"?
Cloudflare embraces the musical metaphor: "a decision model is analogous to a music clef because it helps define the domain of the context and the subsequent notes (actions) that follow it". The clef sets the context; the notes — the agent's actions — follow from it. Clef-flash, the fast 9B version, follows exactly the same logic.
✅ Conclusion
With Clef and Clef-flash, Cloudflare isn't trying to compete with GPT-5.5: it wants to become the place where agents make their decisions — fast, cheap, at the edge of the network, and with open weights. If you're building agents, try out the decision/generation pattern this week, and browse our selection of the best autonomous AI agents to see what's already being done on the tools front.