If a decision in your agent pipeline can be written down as a fixed set of criteria with a typed answer (a category, a severity, a yes or no), you are probably overpaying for it with a full LLM call. On 1 October 2026, Cloudflare shipped two open-weight decision models, Clef and Clef-flash, under Apache 2.0 on Hugging Face and hosted on Workers AI, aimed squarely at those bounded gates. The practical consequence: checks that used to cost a multi-second LLM round trip can now run in tens of milliseconds, which changes not just your latency budget but how many guardrails you can afford to scatter through a pipeline.
The honest answer to “should I swap” is narrower than the announcement suggests. Swap where the decision fits a fixed criteria schema with typed output: classification, triage, severity routing, batched adjudication. Keep a general LLM for open-ended judgment, because an independent paired evaluation of system-1 decision models found neither model it tested beat chance on zero-shot model routing. And every Clef performance figure in this article is Cloudflare-reported; kaer.ai notes the scores have yet to be reproduced for ranking on the official Decision Index.
What shipped on 1 October: Clef, Clef-flash, and the Jev drop-in claim
Cloudflare released two tiers. Clef is the precision model; Clef-flash is the latency model. The vendor’s Workers AI changelog reports a median of 209.3 ms and a p95 of 238.6 ms for Clef, against 38.8 ms median and 122.4 ms p95 for Clef-flash, and 524.1 ms median and 536.0 ms p95 for Jev in the same table.
Two things make this more than a model drop. First, the API story: Cloudflare says Clef follows the System One API, so an existing Jev integration switches by changing the endpoint and model. Jev is Typesafe’s decision model, the current reference point for this model class, so the claim targets teams already paying for decision-model infrastructure rather than asking them to adopt a new abstraction. Note that this compatibility is a vendor claim, not a tested migration.
Second, the capability delta Cloudflare asserts over Jev: a vision encoder and a 64k context window, versus Jev’s text-only classification and 32k limit. That comparison is Cloudflare’s characterization as relayed in news coverage, not Typesafe documentation, and the same caution applies to the benchmark and pricing deltas below.
The economics of a 38.8 ms gate: why decision latency compounds in agent loops
A single gate’s latency looks trivial. Agent loops multiply it. A study of decision models in a penetration-testing harness (arXiv:2609.28940, author-reported) quantifies the compounding: “An agent that takes 3 seconds to decide whether to try the next payload completes 20 confirmation attempts in 60 seconds; one that takes 33 milliseconds completes 1,800.” (arXiv:2609.28940) That is a 90x difference in attempts per minute in the paper’s arithmetic, decided entirely by the per-gate cost of asking “should I try this next.”
The same paper describes a second lever: batching. With System One batching, 13 findings can be adjudicated per API call in large-surface triage, instead of one LLM call per finding. Thirteen calls collapse into one, and each call costs milliseconds instead of seconds. That 13 is the batch limit in the paper’s setup; Clef’s changelog allows up to 64 questions per request, so do not import the lower ceiling when swapping endpoints.
Cloudflare’s own worked example points the same direction. In a Threat Intelligence test, Clef fetched, rendered, and classified a domain in 2.2 seconds, returning categories like 95% fashion, 85% ecommerce, and under 1% phishing, while gpt-oss-120b, Cloudflare’s fastest general model, took 4.7 seconds and returned only two classifications. This is a vendor-run demo on one workflow, so read it as an existence proof, not a benchmark. What it shows is the shape of the win: on a bounded classification, the decision model was both faster and returned richer typed output.
The strategic consequence is the one worth sitting with. When a gate costs 3 seconds and a real token bill, engineers ration gates. When a gate costs 39 ms, per Cloudflare’s reported Clef-flash median, you can put a severity check after every tool call and a routing check before every branch. The constraint stops being tokens and becomes something less familiar: the quality of your criteria schemas. Cheap gates invite gate sprawl, and every new gate is a schema someone has to design, test, and maintain.
Clef vs Clef-flash: choosing precision or latency per gate
Cloudflare’s reported numbers, all vendor-run and unreplicated, frame the split:
| Axis | Clef (precision) | Clef-flash (latency) |
|---|---|---|
| Median / p95 latency (vendor changelog) | 209.3 ms / 238.6 ms | 38.8 ms / 122.4 ms |
| Self-hosting VRAM floor, single concurrency, 64k context (kaer.ai, attributed to Michelle Chen) | 85 GB | 41 GB |
| Intended role | Accuracy-critical gates | High-frequency, latency-critical gates |
On accuracy, Cloudflare’s changelog reports Clef at 98.47 BFCL case-exact and 94.20 BANKING77 macro-F1, versus 95.75 and 79.74 for Jev. The same table gives Clef-flash 98.76 on BFCL, slightly above Clef, and 90.93 on BANKING77, not far behind. On those two benchmarks the 9B model gives up little, but the same table shows where the tier split is real: Clef-flash scores 66.77 on CLINC150+OOS macro-F1 against Clef’s 97.43, below even Jev’s 89.27, and the launch blog reports a When2Call accuracy gap of 72.37 to 65.58. On out-of-scope-heavy intent tasks the drop is large, so the precision-versus-latency choice is benchmark-dependent, not settled by the tier names. The BANKING77 gap to Jev, if it holds up, is the kind of margin that matters for intent classification. The Jev figures are Cloudflare’s characterization of a competitor, and independent replication should be treated as pending.
Pricing cuts the other way. Clef costs $0.24 per million tokens, nearly six times Jev’s $0.042, per news-reported pricing. At decision-model volumes, where each call is small, the absolute numbers stay modest, but a team already running Jev at high gate frequency should do the arithmetic before assuming a swap saves money. The savings case against a general LLM call is much clearer than the case against Jev.
Keep-or-swap: mapping each workflow step to a model class
Here is the framework the evidence supports. I would classify every decision point in the pipeline before touching any endpoints:
- Swap candidates (bounded, schema-defined): domain or intent classification, severity triage, allow/deny checks against a written criteria set, batched adjudication of many similar findings, routing among a small fixed set of well-described options. These are the gates where the reported latency and batching economics apply, and where typed output is a feature rather than a constraint.
- Keep the LLM (open-ended judgment): zero-shot routing to a model or tool the decision model has not been tuned for, relevance judgments on unfamiliar retrieval corpora, anything where “the right answer” is not enumerable in advance. The independent evidence below shows system-1 models at chance on exactly this class of task.
- Choose Clef over Clef-flash when a wrong gate is expensive (a misrouted severity, a missed phishing classification) and a vendor-reported ~210 ms median, ~240 ms p95 fits the budget. Choose Clef-flash when the gate fires inside a tight loop and its vendor-reported 38.8 ms median (p95 122.4 ms) is the point; on out-of-scope-heavy intent tasks the 9B model’s accuracy drop is large, 66.77 versus Clef’s 97.43 on CLINC150+OOS, so check the benchmark nearest your gate before choosing flash. If self-hosting, the 41 GB versus 85 GB VRAM floors may make this decision for you.
- Stay on Jev, for now, if your gates already work and the roughly 6x token price gap matters more than Cloudflare’s claimed accuracy delta, at least until someone outside Cloudflare replicates those numbers.
That last row deserves emphasis. The swap case is strongest against general LLM calls, weaker against an existing decision-model integration.
Migration checklist
Built from the vendor’s own examples and the independent failure data, in the order I would work through it:
- Write the criteria schema first. Every gate you plan to swap needs its decision criteria made explicit and its output typed. This is where the real migration cost lives, and it is work you would have to do for any decision model, Clef or otherwise.
- The endpoint swap is the easy part, on paper. Cloudflare says switching a Jev integration means changing the endpoint and model. Treat that as a claim to verify in staging, not a guarantee.
- Test order stability. In the independent paired evaluation, the open-weight Laya model changed 30% of its answers when the order of options was reversed. Whether Clef shares this failure mode is untested, so run your evaluation set with option order shuffled both ways before you trust a single accuracy number.
- Test on your ambiguous inputs. The documented failure modes of this model class concentrate exactly where inputs are underspecified. Your staging set should over-represent the cases your current LLM gate handles by reasoning rather than pattern.
- Confirm the data-handling terms. Cloudflare states it does not read, store, or train on requests or responses for the hosted models unless you opt into its fine-tuning product. For gates that see customer data, that guarantee may be the deciding factor over self-hosting.
- Decide on the fine-tuning path deliberately. Opting into fine-tuning changes the data-handling posture above. It is also the likely escape hatch if a gate sits near the boundary of what the base model handles, so treat it as a roadmap item with a privacy cost, not a free upgrade.
- Price it honestly. Compare Clef’s $0.24 per million tokens against both your current LLM gate spend and any existing Jev spend, at your actual gate frequency.
What independent evidence shows, and where system-1 models fail
The third-party evaluations cited here tested Jev and the open-weight Laya, not Clef, so they bound the category rather than grade Cloudflare’s specific models. Two findings from the paired evaluation across 11 agent decision points (arXiv:2610.02267) matter for a keep-or-swap decision.
The positive finding: Jev was significantly more accurate than Laya on 9 of 11 decision points, with gains of +10.8 to +46.0 percentage points, and on tool selection with similar distractors, Laya fell to 31.3% at 50 nearest-neighbour options, while Jev scored 98.9% on items with a unique correct tool, 84.5% overall. A well-built decision model can be genuinely strong on structured gates.
The limiting finding: “Neither model beats chance on zero-shot model routing, and they tie on RAG relevance gating.” Chance performance means a coin flip would do as well. If your gate asks the model to pick among options it has never been conditioned on, or to judge relevance in an unfamiliar retrieval setup, a decision model is not a smaller LLM; it is the wrong tool. This is why the framework above draws the swap line at enumerable answers rather than at task size.
What still needs replication, and the verdict
Everything Clef-specific in the record is one vendor’s account of its own models: the 209.3/38.8 ms medians, the BFCL and BANKING77 scores, the 2.2-second classification demo, the Jev compatibility claim, and the Jev comparison figures, which Typesafe has not corroborated. None of that makes the numbers wrong; it makes them provisional.
My read: if you are routing bounded gates through general LLM calls today, the throughput math alone justifies trialing Clef-flash on your highest-frequency gate, with the order-stability and ambiguity tests from the checklist as the gate on rollout. If a decision cannot be written as a fixed schema with a typed answer, keep the LLM and pay its latency, because the independent evidence says the cheap model will not reliably make that call at all. The migration is real work, but it is schema work, and unlike a token bill, a good criteria schema is an asset you keep.

Join the discussion
Share a useful perspective or ask a question about this article.