Cloudflare's Clef Gives Agents a Decision Model That Skips Text

Cloudflare's Clef is an open-weight decision model that returns typed probabilities instead of text - 38.8 ms median latency, Apache 2.0 weights on Hugging Face. Here's what the category actually does, and where Clef-flash's accuracy falls apart.

Cover art for Cloudflare's Clef Gives Agents a Decision Model That Skips Text

An agent on your support pipeline triage asks a general-purpose LLM whether a ticket is a billing issue, a bug, or an account question. The model generates three sentences explaining its thinking, then buries the answer in the last clause. Your code parses the text, strips whitespace, and hopes. Cloudflare released a different kind of model on October 1 that skips all of that: you pass it a state and a schema of typed questions, and it returns a probability for each answer option with no generated text at all.

That model is Clef - and understanding what it actually does, and where it quietly breaks down, is worth an hour of your time if you build agentic workflows.

What an open-weight decision model actually is

A decision model takes structured questions instead of open-ended prompts. A request carries a state - text, JSON, or images - plus typed questions: yes/no (noul), multiple-choice (choice), and ordered rating (score). Clef returns a probability for every allowed option of every question rather than free-form text. The model never enters a decoding loop; it scores options in a single prefill pass. That is where the speed comes from.

Clef is a family of two open-weight decision models Cloudflare released on October 1, 2026: Clef, built on Qwen 3.8-27B with rank-256 low-rank adapters, and the faster Clef-flash, built on Qwen 3.5-9B.

Both run on Workers AI, and the weights are on Hugging Face under the Apache 2.0 license, so you can download them and run them on your own hardware.

This is also Cloudflare's first internally trained model. The company has trained its own AI models for the first time. The Qwen backbone is not novel; the post-training and schema head on top of it are what Cloudflare built.

38.8 msClef-flash median latencyvs. 524 ms for Jev on the same benchmark
209 msClef (27B) median latencystill faster than Jev at 95th percentile
97.4Clef macro-F1 on CLINC150+OOSversus 66.8 for Clef-flash on the same eval

The accuracy cliff nobody is leading with

Coverage of Clef has focused on the latency win. The number that matters more for teams choosing between the two sizes is this: on multi-class intent classification (CLINC150+OOS, 151 categories), Clef scored 97.43 macro-F1 while Clef-flash scored 66.77 - and in this evaluation, which classifies input into many categories, the larger Clef is far ahead.

That is a 30-point gap on the task decision models are most often pitched for. On binary and simple function-call tasks the gap shrinks considerably: on BFCL's per-case exact-match rate, Clef-flash scored 98.76 and Clef 98.47, so the lighter model was slightly higher.

The practical reading: Clef-flash is genuinely fast and accurate for yes/no routing and two-to-four option classification. Point it at a 150-category taxonomy and it degrades significantly. Clef (27B) holds up across both.

The full results on the model card show that F1 on RAGTruth - which detects errors in generated content - was 79.4 for Clef and 35.6 for Clef-flash. The difference is 43.8 points. If you are using a decision model to validate agent outputs rather than just classify inputs, that gap matters more than the latency win.

How decision models fit an agentic workflow - and what they replace

The honest comparison is not decision model vs. general LLM. It is decision model vs. a structured call to a small classifier or a function-calling prompt on a cheap model tier.

Function calling overhead is real: tool definitions and function schemas consume input tokens on every call. A complex agent with 20 or more tools can spend 2,000-5,000 tokens just on tool definitions. A decision model priced on input tokens only - no output tokens billed - removes that overhead entirely.

All of the recent decision models make the same argument: an agent that asks a 27B model to write JSON is paying for a capability it does not use. That is fair. Where it gets complicated is the API-compatibility play.

Cloudflare built Clef to be fully compatible with the API of TypeSafe AI's Jev System One model, so the request format is the same.

Cloudflare claims better results than Jev in three out of four domains of the Jev Decision Index benchmark. It only loses out on agent trace analysis. Clef-flash responds in under 40 milliseconds while Jev takes more than half a second.

The API compatibility is strategically interesting: teams evaluating Jev can benchmark Clef without rewriting their integration. It is also how Cloudflare is positioning against a closed model from a startup - same interface, open weights, edge distribution.

Beagle in action#customer-support, 8:47am
The ask
180 new tickets overnight, triage bot is blocked waiting for LLM call to classify priority
Beagle drafts
a teammate like Beagle drafts the routing summary once classification returns, so the approval step is human judgment on routing logic, not waiting for the model
You approve
triage completes in one pass; a decision model handles the classification, the human approves the downstream action
Do this in your workspace →

What is genuinely new and what is incremental

The category is not new. Embedding-based classifiers and BERT-size models have done structured classification for years. What is new:

  • Open weights on a 27B model trained specifically for typed decisions, with a schema head that scores options in a single forward pass rather than decoding tokens.
  • Jev-compatible API, which means the decision model category now has a de facto interface standard, the same way OpenAI's function-calling spec became the implicit standard for tool use.

Cloudflare's angle is distribution rather than architecture. Clef runs on the same edge network your Workers already sit on, which is why 38.8 milliseconds is plausible in production and not only on a benchmark rig.

- The company also announced a reinforcement learning service for fine-tuning Clef on your own traffic, built from AI Gateway, Workers AI, and Containers.

What is incremental: the Qwen backbone underneath is unchanged. The architecture story - prefill-only scoring, no decoding loop - is the same one Jev introduced two weeks earlier. Cloudflare is bringing edge deployment and open weights to an idea TypeSafe AI prototyped first.

The pricing criticism that surfaced immediately is worth understanding. The figures in circulation - $0.24 per million input tokens for Clef and $0.09 for Clef-flash - come from secondary coverage rather than from Cloudflare directly. For a model sold on cost per decision, that is the number you would expect first. Cloudflare's own changelog links to a pricing page rather than stating the figure in the announcement. At $0.24/M input, Clef is priced between a budget general-purpose model and a mid-tier one - reasonable if you are replacing function-calling calls to a $2/M model, tight if you are replacing a purpose-built BERT classifier at near-zero marginal cost.

Classifying 1,000 support tickets for priority routing
Without Beagle
each ticket hits a general LLM with a 20-tool schema; 3,000 tokens average per call including tool definitions; output tokens billed; ~$0.006 per ticket at mid-tier rates
With Beagle
each ticket hits Clef-flash; input only, no output tokens; 38.8 ms median; ~$0.0001 per ticket at listed Workers AI rates - if you can live with the 66.8 macro-F1 on fine-grained taxonomies

The cost math looks convincing until you recall that many teams are already running a small fine-tuned classifier at effectively zero marginal cost for exactly this use case. Decision models make the most sense at the sweet spot where: (a) you need more than a four-class classifier but fewer than 20 free-text output classes, (b) latency under 50ms is a hard requirement, and (c) you don't want to maintain a fine-tuning pipeline.

The category is real. The hype that any routing decision should now go through a decision model instead of a cheaper purpose-built classifier is not.

Cloudflare Clef and the open-weight decision model: common questions

What is a decision model, and how does it differ from an LLM?

A decision model takes typed questions with predefined answer options and returns probabilities for each option in a single pass - no text is generated. An LLM generates tokens until it produces a full output. Decision models are faster and cheaper for classification tasks but cannot answer open-ended questions or explain their reasoning.

Can I run Clef locally, off Cloudflare's network?

Yes. Clef is listed on Workers AI at $0.24 per million input tokens. The weights for Clef and Clef-flash are also free to download from Hugging Face under the Apache-2.0 license, so teams can run the models on their own GPUs instead of paying for hosted inference.

When should I use Clef-flash versus Clef?

Use Clef-flash for binary routing, yes/no decisions, and schemas with four or fewer options where 38.8ms latency matters. Use Clef (27B) for multi-class taxonomies - the macro-F1 gap of roughly 30 points on 150-category classification makes the flash model unsuitable for fine-grained routing without re-evaluation on your own data.

Is Clef compatible with the Jev System One API?

Cloudflare built Clef to be fully compatible with the API of TypeSafe AI's Jev, so the request format is the same. Teams evaluating Jev can benchmark Clef without rewriting their integration code.

What tasks should still go to a general LLM instead of a decision model?

Any task where the output cannot be predefined as a schema of options: writing a draft, summarizing an unknown document, reasoning through an edge case, or explaining a decision to a user. Decision models classify; they do not compose. For agentic systems, the practical split is decision model for routing and triage steps, general LLM for any step that produces content a human reads.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle