BlogAi

Clef (Cloudflare) explained: open decision models you can run yourself

Cloudflare's Clef is an Apache-2.0 decision model with Jev's API. What the benchmarks really say, what the packaging gets wrong, and how to choose.

Nicolás Torres

The open decision model field got its first big cloud platform three weeks after the category was defined. Cloudflare released Clef and Clef-flash on 1 October 2026: open-weight models that read a state and a set of typed questions and return a probability for every allowed answer, with no text to parse. If you read my explainer on Jev and the System One idea, you already know the shape of the category. This is the same category from a company with a GPU network, which is why the models matter more than the announcement does.

Clef is not the music symbol, not a file format, and not an LLM you chat with. It is a 27B decision model built on Qwen3.8-27B, and Clef-flash is a 9B version built on Qwen3.5-9B. Both are Apache-2.0, both are on Workers AI, and both are, in Cloudflare's words, "fully Jev-API compatible", which means the swap is one string in code you already wrote.

A word on the field around it, because "Cloudflare invented this" is not the story. Nimble (Bespoke Labs) and Tev1 (Together AI) were open-weight decision models on Ollama from 29 September, and the community Decision Index already had 70 entrants by then, including Decider, Kev, Rune, Winnow and Perplexity's pplx-decider. Liquid AI's d1 and Fastino's GLiDE are decision models behind hosted APIs. What Cloudflare adds is the first credible open-weight pair from a platform big enough to host it for you as well as let you download it, plus vision.

What Cloudflare shipped

Everything in this table comes from Cloudflare's launch post, its Hugging Face model cards and its Workers AI model page, read on 6 October 2026.

FactClefClef-flashJev 1.13 (reference)
MakerCloudflareCloudflareTypeSafe AI
Released1 October 20261 October 202615 September 2026
Size27B, built on Qwen3.8-27B9B, built on Qwen3.5-9BNot published
WeightsApache-2.0, on Hugging FaceApache-2.0, on Hugging FaceClosed, API only
InputText, JSON, images, videoText, JSON, images, videoText and JSON
Question typesChoice, Score, NoulSameChoice, Score, Noul
Published context64k64k64k per request, 32k for state plus the longest question
Hosted ID@cf/cloudflare/clef@cf/cloudflare/clef-flashtypesafe/jev, jev-1.13.0
Price per million input tokens$0.24$0.09$0.042
Output tokensFreeFreeFree
Runs on your own hardwareYesYesNo

The architectural difference that matters is the joint schema head. Cloudflare froze the Qwen backbone, added a small transformer head that reads the backbone's final hidden states and scores every option of every question in one forward pass, and trained it with rank-256 low-rank adapters. Nothing is sampled, no text is generated, no JSON is parsed. A separate 128M-parameter head is where the decision actually happens, which becomes important when you try to run this locally.

That non-autoregressive step is also why Clef behaved differently from Jev under repetition in the one independent test that measured it, as we will get to. Cloudflare also shipped a vision encoder, so a support ticket that arrives as a screenshot is a decision you can make now, and a reinforcement learning fine-tuning service for customers who want their own labels trained in.

Read the scoreboard like a claims adjuster

Cloudflare's launch post says Clef "is currently the leader" on the Decision Index, the community leaderboard for this model class. Its own run of the suite gives Clef 61.21 and Clef-flash 57.07, against Jev's 57.91 on the board itself. If those numbers hold up, Clef would be first and Clef-flash fourth.

Four things about that table are worth knowing before you quote it.

The two Cloudflare rows are flagged as self-reported. In the leaderboard's own data file, the Clef and Clef-flash entries carry a self_reported flag and a coverage field of 36 benchmarks, against 38 for every community row. The two that are missing are HLE and iSarcasmEval, and both are scored as zero in Cloudflare's total. Nobody outside Cloudflare has reproduced the Clef column on the board's hardware. Creative AI News re-implemented the index formula and reproduced the arithmetic; that says nothing about whether the underlying scores are right.

Clef would not be first even taking everyone's word for it. Fastino published GLiDE on 30 September with a self-reported 64.81 on the same 0.2.1 scorer, 6.9 skill points above Jev and about 3.6 above Clef, including an 11.5-point lead in Knowledge and Reasoning. Fastino ran its own numbers too, and nobody has reproduced those either. Two vendor tables, two leaders, neither checked.

Winning rows is the wrong way to read a vendor's table. I counted them anyway, against Jev's community numbers: Clef wins 25 of the 36 published index benchmarks and loses 11, with no ties. The losses have a shape. Clef is far weaker where the answer needs stored knowledge:

BenchmarkClefClef-flashJev
GPQA Diamond (graduate science)48.051.078.3
BBH (hard reasoning)73.768.992.9
MMLU-Pro65.965.382.7
When2Call (should the agent act?)72.465.681.0
HoVer and NLI4CT (grounding)65.2 and 82.961.2 and 78.672.9 and 84.1

The counterweight matters as much, and it is why I would not call Clef the weaker reasoner. It wins MuSR (multi-step reasoning) 83.5 against 66.1, CLadder (causal questions) 94.0 against 72.6, CRUXEval (code execution) 86.7 against 73.0 and GSM8K 80.8 against 79.9. On Cloudflare's own area breakdown, its Knowledge and Reasoning skill is 0.512 against Jev's 0.514, a tie, while Fastino claims GLiDE leads that area by 11.5 points. The honest reading is narrower than better or worse. Clef is strong at routing and labelling (BANKING77 intent 94.2 against 79.7, CLINC150 out-of-scope 97.4 against 89.3, tool-call selection 98.5 against 95.8), strong on the step-by-step reasoning tests, and weak on benchmarks that need facts it does not have. When2Call is the row agent builders should read twice, because it asks whether a model knows to call a tool, ask a question, or decline.

Clef-flash is not a smaller Clef. On most rows it tracks the 27B model closely, and on two it collapses, both of them checks people put in production. RAGTruth, which measures hallucination detection, is 35.6 for Clef-flash against 79.4 for Clef and 76.5 for Jev, below that benchmark's 51.8 chance line, so it earns zero skill there. CLINC150 with out-of-scope requests is 66.8 against 97.4. If your pipeline uses a decision model to say "this does not belong to any of my categories", the fast model is the wrong one.

One column in Cloudflare's table mixes two different entrants. The DiffusionGemma column's accuracy figures (BFCL 96.52, BANKING77 74.28, CLINC150 83.49) match the leaderboard entry JoshuaSP diffusiongemma (open-jev), index 49.47. Its latency figures (median 84.4 ms, p95 211.2 ms) match a different entry, djev, index 40.28. I checked both value by value against the leaderboard JSON. Two models, one column.

On latency, be specific about what was measured

Cloudflare reports 209.3 ms median for Clef, 38.8 ms for Clef-flash and 524.1 ms for Jev. Those three numbers do not measure the same thing. The leaderboard labels Jev's figure a hosted API round trip from its lab, "not comparable to the on-card single-process figures", and Cloudflare runs Clef on its own infrastructure, which it says is the point: models at the edge, low network latency, decisions in the hot path.

Independent measurements are messier. Construct called all three models from a laptop through Cloudflare's REST API and found a floor of 340 to 384 ms for every model, because the network path dominates; at the median, Clef was the slowest of the three and both Clef models had longer tails. Knowviq, running inside their own Worker on large multi-question requests, measured Clef at roughly 1.1 to 1.4 seconds against a 209 ms claim, and recorded a fifteen-minute window where one-question Clef-flash calls took a 13.7 second median, with one call at 108 seconds, while the same model was stable under a batch run. Their requests are bigger than the benchmark's, so this is not a straight contradiction, but it is the number a production caller is more likely to see.

Two independent tests, both much smaller than the launch table

The most useful published test is Construct's drop-in study, and it deserves the caveats it states itself. It took five decisions from a real agent, 117 labelled cases, three calls per model, and ran them on all three models using thresholds that had been tuned for Jev. Its own summary is honest about the size: Clef 92.3%, Jev 88.6%, Clef-flash 82.1%, with a 95% interval on the lead of minus 1.1 to plus 8.8 points, which is not statistically significant. Jev won email triage, made no wrong merges, and was perfect on the memory gate.

Three findings from that work matter more than the headline accuracy.

  1. Clef gave the same answer every time; Jev did not. Clef and Clef-flash returned the same probability to four decimal places on all 351 repeat calls each, including with cache-busting variations. Jev's action probability moved by up to 0.10 between identical calls, and two cases changed verdict. TypeSafe promises consistent answers for similar inputs, not determinism, so this is consistent with its documentation, but a decision layer acts when a probability crosses a line, so any wobble means a case near the line is decided by which call you happened to make.
  2. Jev's scores moved under the same version string. Construct calibrated three probes against Jev in September at 0.71 to 0.82, set a 0.70 threshold in the gap, and re-asked in October. None of 60 repeat calls reached 0.70, on Workers AI or on TypeSafe's own API with jev-1.13.0 pinned. Three probes are three probes, and they are not claiming Jev got worse in general, but the lesson generalizes in both directions: a version string is not a guarantee, and the hosted Clef alias Construct tested carries no version number either, so pin your own weights if you need the model frozen.
  3. Thresholds do not transfer between models. For a tool-call risk decision, the configured threshold was 0.80, the best threshold was 0.49 for Jev, 0.22 for Clef and 0.31 for Clef-flash. Clef-flash scored 82.1% being read with Jev's ruler and 90.6% with its own, level with Jev's 90.9%. The 9B model was never bad; it was being measured against a different model's probabilities.

Knowviq's write-up, from 938 checks of its own production traffic, shows both sides of the same picture. Clef's sentence-support scoring trails Jev badly (AUC 0.786 against 0.978, on a random sample of 50 sentences), because it often marks a sentence unsupported even when the evidence says the same thing word for word. But at the answerable gate, Clef hides far fewer good answers than Jev does (7.3% against 22.2% at a 0.75 threshold), and it ties Jev on ranking results (AUC 0.894 against 0.899). The threshold problem shows up again: Knowviq's badge rule, tuned to Jev's confidence values at 0.7, hid 82% of relevant results under Clef and 87% under Clef-flash, against 41% under Jev. Three teams, three times, the same conclusion. The model is not the only thing you migrate.

What to know before you run it locally

This is the part I checked by hand, because it is where the open weights stop being a marketing claim and start being files on your machine. It is also the part I got wrong first time round, so the method note is worth stating: checking whether a repo contains a file called joint_head.safetensors tells you nothing about whether the head is in the model, because in a GGUF conversion the head tensors live inside the quantised file.

llama.cpp is now a first-class path. PR 29831, merged on 3 October, added a native clef architecture, and the server exposes /v1/systemone for models whose output_modalities include decisions. Two commands start a working endpoint:

llama serve -hf ggml-org/Clef-GGUF
llama serve -hf ggml-org/Clef-Flash-GGUF

ggml-org's conversion logs list 72 decision.* tensors, and the model metadata reports architecture clef at 27.02B parameters against 26.90B for a plain Qwen build, the difference being the head. bartowski's quants keep the head at Q8_0 in every file. One warning from the server docs: the probabilities it returns are scaled with the temperatures stored in the model file, and "are not guaranteed to be calibrated for your data".

Two small repos are the ones that cannot decide anything. prithivMLmods/clef-GGUF and OlyMahmud/clef-flash-Q4_K_M-GGUF report architecture qwen35 with no decision tensors, so they load as text models. Between them they have 790 downloads. Check the architecture field before you download anything.

Ollama is still the shortest path. Four decision models are in the library: clef (18GB, 27B, text and images), clef-flash (11 to 12GB, 9B, text and images), nimble (9.5GB, Bespoke Labs, text) and tev1 (4.5GB for 4B, 800MB for 0.8B, Together AI, text). They all answer on /v1/systemone, the endpoint Ollama added on 29 September, and Clef needs Ollama 0.35.1 or later. The model pages carry curl examples and ollama.systemone(...) examples in Python and JavaScript, and the client libraries added it on the same day the models landed: ollama-python 0.6.3, whose release notes list "add System One API support", and ollama-js 0.6.4, both published on 29 September. There is no per-decision API cost and no network hop.

The context number in the tag is the architecture maximum, not your working limit. Every one of those four models advertises a "256K context window" on its Ollama tag. The GGUF metadata carries a configured context_length of 262144, which is where that number comes from. The makers document something smaller, on the same page:

Model tagTag metadataDocumented limit on the same page
clef, clef-flash256K"64K context window"
nimble256K"that prompt has to fit in Nimble's 8,192-token context"
tev1256K"runs with a context of about 2,000 tokens"

The model card for Nimble is stricter still: 8,192 tokens, and oversized inputs are rejected rather than truncated. For Clef, Cloudflare serves 64k and the Hugging Face loader defaults to 16,384. Plan for the smallest number you can find. If you send a 20,000-token state because the tag said 256K, you will find out at the worst moment.

The hardware, honestly. Clef in bf16 is 55GB of weights before any cache, so an 80GB-class GPU or a multi-GPU box. Clef-flash is 19.1GB and fits a 24GB card, and Livesport runs a llama.cpp build of it in production on a Quadro RTX 4000 with 8GB. Clef-flash in MLX 4-bit is 6.2GB, which fits a 16GB Mac. Perplexity's pplx-decider-v1-27b wants about 49 GiB and a CUDA GPU. Kev has 0.8B, 4B, 9B and 27B checkpoints, and the smaller ones run on Apple Silicon.

Two traps that will not show up as errors. Bespoke's Nimble requires a temperature of 2.179 when you turn its logits into probabilities, and its card warns that loading only the adapter and applying a plain softmax silently gives you an uncalibrated distribution; nothing on the Ollama page mentions it. And Tev1 is not a Jev-style runtime at all: it keeps Qwen's next-token head and returns an option letter, its card states no license, and it is best treated as a cheap classifier rather than a model with a type guarantee.

Images have rules, and they differ by host. On Ollama, images must be raw base64 in the request, and URLs and data URLs are not supported. On Workers AI, up to four embedded images, PNG, JPEG or WebP, 4 MiB and 16 megapixels each, remote URLs rejected. llama.cpp takes base64 data URLs.

The rest of the field, in one table

All index scores below are the community Decision Index 0.2.1 (28 September 2026), measured on one RTX PRO 6000, so those latencies are on-card and comparable to each other, and not to Jev's hosted 524 ms.

ModelIndexSizeMedian latencyLicenseWhat it is for
Surogate Rune 26B-A4B v357.4425.8B MoE120.5 msApache-2.0High accuracy on one card, GGUF available
Decider chat · Gemma-4-31B57.3332.7B108.5 msTechnique, Apache-2.0 codePrompting method, no new weights
pplx-decider-v1-27b56.4027.8B101.4 msApache-2.0Best calibration in the top group (ECE 0.018), needs about 49 GiB
Jebadiah 27B54.6727.8B110.4 msApache-2.0Best calibration among the 27B group (ECE 0.014)
Winnow-12B50.0212.0B72.5 msApache-2.0Mid-size compromise, Q8 GGUF
Decider 4B40.704.7B12.6 msApache-2.0Best size-to-quality at the low end, one forward pass
Bespoke Nimble 9B v239.579.7B77.3 msApache-2.0Honest card, 255 choices, on Ollama
Kev 9B / 4B / 0.8B38.48 / 34.64 / 14.609.7B / 4.7B / 0.9B51.4 / 52.1 / 41.6 msApache-2.0Full training pipeline, System One server
Tev1-4B29.244.7B35.8 msNot stated on the cardCheap routing, autoregressive, not type-safe
Decider 2B28.972.3B8.1 msApache-2.0Cheapest useful gate, one forward pass
Laya6.040.42B5.8 msApache-2.0Needs fine-tuning on your labels; near-zero zero-shot
GLiNER2.5-Decide11.210.49B23.3 msApache-2.0CPU-friendly classifier, strong once fine-tuned

Two reading notes. First, the index is chance-corrected accuracy: each benchmark is mapped to (score minus chance) divided by (1 minus chance) before averaging, with unanswered requests counted as wrong. Speed, cost and calibration are not in it, so a 4B model outranking a 27B model is a real accuracy result on this panel rather than a composite trick; what it cannot tell you is whether the panel looks like your traffic. The composite that does blend speed and cost is JevBench, which has a different board and a sealed test set. Second, several of the highest rows are techniques rather than weights, which means you can apply them to a model you already trust instead of adopting a new one.

How to choose

The decision rules that fall out of all of the above:

  • Closed label set, high volume, cheap mistakes (routing, tagging, extraction, model selection): Clef-flash locally or hosted, or a 4B like Decider. Retune thresholds on your own labels first. This is the narrow judgment call that loop engineering keeps in code rather than in a prompt.
  • Anything that needs judgment rather than lookup (should this tool call wait, is this answer grounded, does this request belong to any of my categories): the 27B Clef, or Jev, or a frontier model with a schema. Not Clef-flash, whose RAGTruth and out-of-scope numbers are below the incumbent.
  • Knowledge-heavy decisions, or reasoning you have not decomposed: none of the open models has a checked lead here. Clef trails badly on GPQA, MMLU-Pro and BBH, Jev leads on the community board, and GLiDE claims a lead nobody has reproduced. Measure it yourself, or keep the frontier model.
  • Vision: Clef or Clef-flash, because Jev cannot read an image today. Test it on your own assets, since no image benchmark has been published.
  • You need a version you can pin, data that cannot leave your network, or a fine-tune of your own: the open weights are the whole point. That is the argument Apache-2.0 makes, and it is a better one than any leaderboard row.
  • Whatever you pick: swap in shadow first, log the probabilities, and set each threshold from your own data. The uncomfortable lesson of this week is that a threshold is a property of a model and a prompt together, and the models move underneath you. If reading calibration and thresholds is new to you, my guide to LLM evals covers how to measure all of this on your own cases.

Where this leaves Jev

Of the scores anyone has actually checked, Jev is still the best on knowledge and the cheapest hosted decision per token, by about 2.1x against Clef-flash and 5.7x against Clef. What it lost in the last two weeks is exclusivity. There is now an open-weight pair from a large platform: downloadable, pinnable, fine-tunable and auditable, with a vision encoder and, in the one test that measured it, identical answers on every repeat call. That matters more than three points of accuracy if you have ever debugged an agent that did something different the second time.

The category also has a new failure mode, and it is not on any leaderboard. A vendor table can be eight points better than the incumbent on the rows the vendor chose, and still be third on self-reports once two other vendors publish their own. Three claims, three leaders, none reproduced. The number that matters is the one you measure on your own labelled cases, which is the argument from my guide to LLM evals, now with an industry attached to it.

Everything above is from primary sources read on 6 October 2026: TypeSafe's docs and jaggedness page, Cloudflare's launch post, model cards and Workers AI pages, the Decision Index data, the leaderboard file behind Cloudflare's numbers, the Ollama library pages, Hugging Face file listings, GGUF metadata and conversion logs, the llama.cpp pull request and server docs, plus three independently published tests. Vendor numbers are labelled as vendor numbers. This category is three weeks old, so prices, limits and leaderboards will all move.

Frequently asked questions

What is Clef?

Clef is a decision model from Cloudflare, released on 1 October 2026. It reads a state (text, JSON, images or video) and a set of typed questions, and returns a probability for every allowed answer instead of generating text. There are two sizes: Clef at 27B parameters, built on Qwen3.8-27B, and Clef-flash at 9B, built on Qwen3.5-9B. Both are Apache-2.0 open weights on Hugging Face, both are hosted on Workers AI, and both speak the same request format as TypeSafe's Jev.

Is Clef better than Jev?

On Cloudflare's own numbers, Clef would top the community Decision Index at 61.21 against Jev's 57.91. That run is self-reported, covers 36 of the index's 38 benchmarks, and has not been reproduced independently, so treat it as a vendor claim. Two other vendors make larger claims of their own: Fastino reports 64.81 for GLiDE, which would put Clef third once everyone's own numbers are taken at face value. Reading Cloudflare's table row by row, Clef wins 25 of the 36 benchmarks it published and loses the knowledge-heavy ones by wide margins: GPQA Diamond 48.0 against 78.3, BBH 73.7 against 92.9, MMLU-Pro 65.9 against 82.7, When2Call 72.4 against 81.0. It wins most of the step-by-step reasoning rows instead.

Can I run Clef locally in LM Studio or llama.cpp?

Yes. Clef's probabilities come from a separate joint schema head, and the good GGUF conversions include it inside the quantised file. bartowski's repos store the head at Q8_0 in every quant, ggml-org's conversions carry the decision tensors with llama.cpp's native clef architecture, and llama serve -hf ggml-org/Clef-GGUF answers on /v1/systemone. Cloudflare's own repositories with their Python loader, the MLX 4-bit builds and Ollama also work. The exceptions are two small repos, prithivMLmods/clef-GGUF and OlyMahmud/clef-flash-Q4_K_M-GGUF, which are plain Qwen builds with no head: those load as text models and produce meaningless text. Check the architecture field before downloading.

How much does Clef cost?

On Workers AI, Clef is $0.24 and Clef-flash is $0.09 per million input tokens, with no charge for output because a decision model generates no text. Jev is $0.042 per million input tokens, so hosted Clef is about 5.7 times Jev's price per token and Clef-flash about 2.1 times. In absolute terms the amounts are small: Construct measured $93 per million Clef decisions against $23 for Jev, and its entire 1,053-call benchmark cost about five cents. The weights are free to download and self-host under Apache 2.0.

Does Clef read images?

Yes, and that is the clearest capability difference from Jev, which is text and JSON only. Cloudflare's hosted endpoint accepts up to four embedded PNG, JPEG or WebP images per request, 4 MiB and 16 megapixels each, with remote URLs rejected. llama.cpp accepts base64 data URLs for the same models. The open weights also accept video frames. Two caveats: the vision encoder is unchanged from stock Qwen, and Cloudflare published no image benchmark, so measure it on your own assets before trusting it.

Should I use Clef or Clef-flash?

Use Clef-flash for high-volume gates that can only skip or downgrade work, and only with thresholds measured for it. Use the 27B Clef for anything that merges, deletes, sends or passes a policy, and never use either flash model as a hallucination or out-of-scope guardrail: on the Decision Index, Clef-flash scores 35.6 on RAGTruth against 79.4 for Clef and 76.5 for Jev, below that benchmark's 51.8 chance line, and 66.8 on CLINC150 with out-of-scope requests against 97.4 for Clef.