AIOctober 2, 202611 min read
What a decision model is, and why Clef is not the cheap one
Cloudflare open sourced Clef on 1 October and charges $0.240 per million input tokens for it. TypeSafe's Jev, the model Cloudflare measures itself against, publishes $0.042.

A decision model is an AI model that never writes a sentence. You give it a situation and a list of multiple choice questions, and it returns a probability for every answer you allowed, which your own code reads directly instead of pulling values out of a paragraph.
Cloudflare released 2 of them on 1 October, Clef and a smaller, faster one called Clef flash, with the weights free to download under a license that lets anyone use them commercially. Calling Clef on Cloudflare's own network costs $0.240 per million tokens of input, and nothing at all for the answer, because no text comes back to charge for. TypeSafe's Jev, the model Cloudflare measures itself against all the way through its announcement, publishes $0.042 for that same million tokens.
What is a decision model, and why does it never write a sentence?
A decision model reads a situation and answers a fixed list of questions with probabilities, and it has no ability to produce prose at all. The definition in Cloudflare's own announcement, written by Michelle Chen, Alex Reneau and Kevin Flansburg, is that "a decision model makes classifications to help agents decide how to act, based on certain probabilities".

The example Cloudflare uses is a customer support message. You send the text of the complaint, then ask whether a service is down and which team should handle it, and back comes a named team, a yes or no answer with the probability of yes, and a severity level on the scale you defined. Your code reads those values and routes the ticket. Nobody has to read a reply and work out what the model meant by it.
That is the difference from asking ChatGPT or Claude to make the same judgment. A chat model can do the job, so you ask it to answer in JSON, and then you write code to handle the day it answers in a slightly different shape, or wraps the whole thing in an apology, or invents a team name that does not exist in your system. Clef cannot do any of that, because the list of allowed answers travels inside the request and the model only ever scores the options you supplied.
Clef's model card says the model does no free form text generation and no output parsing. The clearest evidence sits in Cloudflare's own price list, where every other model carries 2 rates, one for what you send and one for what it writes back. Clef and Clef flash each carry a single rate, for the input, and the cell where an output rate would go is empty.
The uses are the unglamorous machinery underneath software rather than anything a person ever sees. Routing a ticket to a queue, deciding whether a document is complete enough to file, scoring whether a retrieved passage deserves to reach an expensive model, checking whether a message breaks a rule before it reaches a customer. Cloudflare says its own threat intelligence team has been running Clef to classify website domains, and that the model returned more categories in about half the time its fastest general model needed for the same job.
What does Clef cost on Workers AI, and what runs for nothing?
Clef's rate on Cloudflare's Workers AI buys more than most readers expect, because every account gets a free allowance of Neurons each day, which at the rate Cloudflare publishes for Clef works out at roughly 450,000 tokens of input before anything is billed. Neurons are Cloudflare's own billing unit, and the daily counter resets at midnight UTC.

Clef flash costs $0.090 for the same million tokens of input, so the free daily allowance stretches well past a million tokens on the small model. A million tokens is a great deal of work in this setting, because what you send a decision model is a ticket, a log line, an invoice or a short document rather than a novel, so that budget is thousands of separate decisions rather than one long read.
Clef is a post trained version of Alibaba's Qwen3.8, and Cloudflare sells that same Qwen model on the same platform, where feeding it costs almost twice what feeding Clef costs, and its written answers are then billed on top at more than a dozen times Clef's input rate. Teaching a model to answer in probabilities instead of prose made it cheaper to run in both directions.
With a chat model, the expensive half of the bill is usually the writing, as the rates in our comparison of what the big model APIs charge show. With a decision model there is no writing, so the only meter that turns is the one measuring what you sent.
Is Clef cheaper than TypeSafe's Jev?
No, Clef is the more expensive of the 2 on published rates. TypeSafe's own documentation prices Jev well under Cloudflare's rate for Clef, and both companies charge nothing for the answer itself. Cloudflare's announcement calls Clef smarter and faster than Jev and never puts the 2 price lines side by side.

| Clef | Clef-flash | Jev 1.13 | |
|---|---|---|---|
| Made by | Cloudflare | Cloudflare | TypeSafe AI |
| Price per million input tokens | $0.240 | $0.090 | $0.042 |
| Price for the answer | nothing | nothing | nothing |
| What it can read | text, JSON, images, video | text, JSON, images, video | text only |
| Context its maker publishes | 64k tokens | 64k tokens | 64k per request, 32k for the state plus the longest question |
| Weights you can download | yes, 55 GB | yes, 19 GB | no, API only |
| Licence | Apache 2.0 | Apache 2.0 | not published |
| Median latency on Cloudflare's run | 209.3 ms | 38.8 ms | 524.1 ms |
What the extra money buys is real enough. Clef has a vision encoder, so it reads images and video frames as well as text, and TypeSafe's documentation says Jev takes text only and tells you to turn anything else into text before you send it. Clef's weights are also yours to download, where Jev is reached only through TypeSafe's own service, and the company says access is currently shaped by rate limits it can change without notice while it serves a very large volume of demand.
Cloudflare's post says its model has a bigger context window than Jev, comparing its own figure against a smaller one, and TypeSafe's models page publishes that same larger figure for a whole request, with the smaller number applying only to the situation plus the single longest question in it. The 2 figures being set against each other are not measuring the same thing, and a reader checking both pages this morning finds the same headline number on each.
Neither company sells a model you can meaningfully compare on a price list alone, which is why the licence matters as much as the rate. Clef ships under the Apache license, and so do both Qwen models underneath it, which is rarer than it sounds once you start reading the actual files, as we found when we ranked the open models by what their licences allow.
Does Clef really beat Jev on Cloudflare's own tests?
Cloudflare calls Clef "currently the leader when evaluated against the Jev Decision Index", and the table published on Clef's own model card shows Jev still ahead of Clef in 15 of the 41 scored rows. Clef takes the rest, so the record is a clear win overall and nothing like a sweep.

We counted those rows by hand on the morning of 2 October, going down the Decision Index table on the model card and leaving out the 2 latency rows at the bottom, which measure speed rather than quality. Jev holds its lead in the broad general knowledge tests and wins the graduate level science set by a wide margin, which makes sense for a model asked to judge written material. Clef wins most of the rows about sorting support messages, reading contracts and picking the right tool, which is the work these models are actually sold for.
Clef flash, the small cheap model, is ahead of the bigger Clef in 18 of those same rows while answering several times faster, which is the count that changes what you would do tonight. The expensive model is not automatically the right call, and on plenty of ordinary classification work the small one is both better and cheaper. Cloudflare says as much itself, noting that its flash model performs exceptionally well given how much faster it is.
None of this has been checked by anyone outside Cloudflare. Cloudflare ran the Decision Index itself, and it also ran TypeSafe's own workflow tests itself rather than waiting for TypeSafe to run them. TypeSafe's evaluation site lists TypeSafe, OpenAI, Anthropic and Fireworks among the providers it scores, with no Cloudflare entry on it when we checked on 2 October. A benchmark table published by the company that wins it is a starting point for your own test, not a result.
Can you run Clef on your own machine tonight?
Not the big one, on most machines. Clef's weights come to 55 GB spread across a dozen files on Hugging Face, and Cloudflare says it tested the model on a single H200, which is a data centre card rather than anything that goes into a laptop. Clef flash is far smaller and lands within reach of a well specified desktop.

What you download is mostly somebody else's model. Cloudflare left the Qwen weights frozen and trained a small extra piece that sits on top of them, reading the model's internal state and scoring every allowed option against the situation in a single pass. That design explains the licence, because the base was already free to use commercially, and it explains the size, because the Qwen model underneath is a full one rather than a trimmed down classifier.
Size is not the only obstacle. Clef is not a single file you point a local app at, because the decision head that turns the model's internal state into scored options ships as its own file, with a Python module beside it that does the encoding and the scoring. Cloudflare's published example loads the model through that module, and anyone wanting to serve Clef locally has to carry that code along with the weights.
Ollama's library does not carry either model, so the usual one line pull that works for most open models does not work here, and the download numbers tell the same story. Cloudflare's Clef page had collected hundreds of likes and 18 downloads when we checked on 2 October, which is the signature of a release people are bookmarking rather than running. Community conversions into the smaller local formats did appear within a day, every one of them still showing no downloads, and nobody has yet reported whether those files carry the decision head or only the Qwen model underneath it.
The realistic way to try Clef tonight is the hosted one, inside the free daily allowance on Workers AI, which costs nothing and needs no card. If what you want is a decision model running entirely on your own hardware, the practical choice today is one of the small models already in Ollama's library, which install the way everything else in it does. Our guide to picking a local model that fits your memory covers the arithmetic for working out what your machine can hold.
Who else shipped a decision model this week?
TypeSafe AI started this category with Jev, and in the days around Cloudflare's launch, Amazon, Bespoke Labs and Together AI all put out a decision model of their own. Cloudflare's own announcement acknowledges it, saying the market is getting increasingly saturated with decision models.

The clearest sign that this is now a category rather than one company's idea is that Ollama has given it a shelf. Its model library carries a decision filter beside the ones for vision and embeddings, and the 2 entries on it are nimble from Bespoke Labs and tev1 from Together AI, both of which install with a single command and have collected thousands of pulls each in their first days. TechCrunch reported on 1 October that Amazon released Strands Decider 2B, open sourced and small enough to run locally, and Latent Space's roundup of OpenAI's DevDay lists a Decisions API among the announcements there.
The spread in what these models charge is wider than the spread in what they do. They all do a similar job, they all refuse to write prose, and the published rates for the same million tokens of input already differ by several times over. Nobody has settled on what a decision is worth, which is the normal condition of a market a few weeks old and a good reason to test 2 of them rather than committing to one.
There is a second reason the category appeared so suddenly. Agents only became useful once software could act on a model's judgment without a person reading it first, and that needs an answer with a known shape and a number attached. Every agent platform has been working around the absence of one, including the hosted agent runtimes we priced recently, where the model still writes text that something downstream has to interpret.
What do we still not know about Clef?
No independent evaluator has published a Clef score. Every figure in circulation on 2 October comes from Cloudflare's own run, including its run of TypeSafe's workflow suite, and the questions below are open rather than answered.
- No outside party has measured Clef against Jev, so the row by row record stands entirely on the company that published it.
- The fine tuning product is not self serve yet. Cloudflare says it starts as a hands on engagement with its forward deployed engineering team, and no price for that work appears anywhere.
- The context figure in the announcement does not match the one in the configuration file shipped with the weights, and Cloudflare has not explained which ceiling applies to a request on Workers AI.
- Nobody has reported whether the community conversions of Clef keep the decision head, which is what separates this model from the Qwen model it was built on.
2 other lines in the announcement matter before you send anything real through this. Cloudflare says it does not read, store or train on your requests or responses on Workers AI, with one exception, the fine tuning product, which exists to learn from your traffic. That product is where the company wants this to go, and the plan it describes puts its own gateway in front of your requests to capture them, a sandbox to run the training, and the tuned model redeployed on the same network.
A model that answers in probabilities can still be confidently wrong, and that is the objection worth holding onto, because a wrong answer carrying a high confidence score is harder to catch than a wrong paragraph, because a paragraph reads oddly to a human and a number does not. TypeSafe says as much in its own documentation, where confidence is treated as a second signal to route on rather than as proof. Anything you wire a decision model into needs a threshold below which it asks a person, and you only find the right threshold by running your own data through it.
If you want to spend an evening on this, the useful test is small. Take a few dozen support messages, or invoices, or whatever your software currently asks a chat model to classify, run them through Clef inside the free daily allowance, and compare both the answers and the bill against what you pay today. Keep the cases where the model was unsure and read those yourself, because that is where your threshold gets set. The weights behind all of this are open in a way worth appreciating, since both Clef and the Qwen base underneath it carry the same permissive licence, which we unpacked in our explainer on what open weights actually let you do.
Questions people ask
What is a decision model in AI?
A decision model is a model that returns typed answers with probabilities instead of text. You send it a situation and a list of questions with their allowed answers, and it scores every option you defined. Cloudflare describes it as making classifications that help agents decide how to act, based on probabilities.
What does Cloudflare's Clef cost?
Clef costs $0.240 per million tokens of input on Cloudflare's Workers AI, and the smaller Clef flash costs a good deal less again. Neither model charges anything for the answer it returns, because it produces no text to bill for.
Is Clef free to use?
The weights are free to download from Hugging Face under the Apache license, and the hosted version is free up to Workers AI's daily allowance, which covers roughly 450,000 tokens of input a day on Clef. Past that allowance you pay Cloudflare's published rate.
What is the difference between Clef and Jev?
Jev came first and is made by TypeSafe AI, costs less per million tokens of input, and reads text only. Clef is made by Cloudflare, costs more, reads images and video as well as text, and you can download its weights and run them yourself.
Can I run Clef on my laptop?
The larger Clef is unlikely to fit, since its weights run to dozens of gigabytes and Cloudflare tested it on a data centre graphics card. Clef flash is far smaller and within reach of a strong desktop, though neither model installs through Ollama today and both need the Python module Cloudflare ships beside the weights.
Do decision models hallucinate?
A decision model cannot invent an answer outside the list you gave it, which removes the broken shapes and invented values you get when a chat model is asked for structured output. It can still pick the wrong option with high confidence, so a threshold that sends uncertain cases to a person is still needed.
Why does Clef have no output price?
Clef never generates text, so there are no output tokens to meter. It scores the options supplied in the request and returns a probability for each, which is why Cloudflare's pricing table shows an input rate for it and an empty output cell.
