Cloudflare releases open-weight decision models for AI agents at the edge
The Clef and Clef-Flash models, announced by Cloudflare during its Birthday Week, trade text generation for typed probabilities, with open weights and latency designed to run inside the critical path of AI agents.
During its recent Birthday Week, Cloudflare announced Clef, a family of open-weight models designed to choose among predefined options instead of generating free-form text. The company released versions with 9 billion and 27 billion parameters, along with a platform for adapting the models to specific decision tasks, according to a report by InfoQ.
The launch targets developers who build AI agents directly: the idea is to separate the "decide what to do" step from the "generate a response" step, something that today is usually handled by forcing a general-purpose LLM to return JSON and hoping the parser doesn't break.
What a decision model actually is
Clef is a 27B-parameter multimodal model that receives a state (text, JSON, image, or video) and a schema of typed questions, and returns a probability for each allowed option, for each question, in a single forward pass. There is no free-form text generation or output parsing: the agent receives numbers directly, which it can use to decide how to act, whether routing a support ticket, escalating to a human, or deferring the decision.
In practice, this eliminates an entire class of bugs common in agent pipelines: the LLM that returns malformed JSON, the field that comes back as a string when it should be an enum, the silent retry because the parser failed. The decision model trades generation for classification, and classification has typed output by definition.
Two sizes, two latencies
Besides the 27B Clef, Cloudflare released Clef-Flash, a 9B-parameter version designed for latency-sensitive decisions. The numbers disclosed by the company, in its own benchmarks, show the difference between the two:

| Model | Parameters | Median latency (Cloudflare) |
|---|---|---|
| Clef | 27B | 209.3 ms |
| Clef-Flash | 9B | 38.8 ms |
The more than 5x difference in latency is Cloudflare's central argument for placing Clef-Flash in the middle of a real-time decision flow, where each additional network call costs user experience.
The edge argument: decisions near the user, action with a separate LLM
Because they are hosted on Cloudflare's own infrastructure, both models run on the GPUs the company already distributes across the edge, and that is the point the company uses to justify the architecture design it is proposing. Michelle Chen (group product manager), Alex Reneau (principal machine learning engineer), and Kevin Flansburg (senior engineering manager), all from Cloudflare, write:
Because they are hosted on Cloudflare's infrastructure, we're able to take advantage of our GPUs at the edge, leading to low network latency and faster decisions. This means that you could put Clef into the hot path for agents to make decisions and combine that with one of our LLMs on Workers AI to take action.
Michelle Chen, Alex Reneau and Kevin Flansburg, Cloudflare
The architectural idea, then, is explicit: Clef stays in the hot path, deciding quickly, and a generative LLM steps in afterward, only when the action requires producing text (a reply to a customer, a summary, an email). For those who currently use a single general-purpose model to do both things, this is a new separation of responsibilities within the agent's own pipeline, with implications for cost (a smaller model running most of the time) and reliability (typed output instead of text to be parsed).
Compatibility with Jev System One, from Typesafe AI
Clef's API is compatible with Jev System One, Typesafe AI's decision model that already had a presence in the market before Cloudflare's announcement. Chen, Reneau, and Flansburg detail the technical differences between the two:
First, it has a vision encoder so it's able to take in images and classify visual content. This is different from Jev, which only does text classification today. Secondly, our model has a 64k context window (compared to Jev's 32k), which allows users to squeeze more input state for the model to classify against.
Michelle Chen, Alex Reneau and Kevin Flansburg, Cloudflare
API compatibility matters for those who have already integrated Jev: migrating to Clef, in theory, doesn't require rewriting the integration layer, just swapping the endpoint and testing whether the typed question schema behaves the same way.
Community skepticism about benchmarks and the "new paradigm"
The announcement sparked discussion on both Hacker News and Reddit, with part of the community questioning whether decision models are truly a new category or a repackaging of supervised classification. Jacek Złydach summed up the skepticism:
It's not a 'new paradigm', it's a low-hanging fruit that's been lying around for years; Typesafe were the first to bother to stop and pick it up, and market the shit out of it.
Jacek Złydach, comment on Hacker News
The benchmarks released by Cloudflare were also questioned. On Hacker News, user SebastianSosa warned about the risk of optimizing for the wrong metric:
Public benchmarks are easy to cheat, if I am Typesafe, I would also release a public benchmark to distract otherwise competent people from overfitting to a benchmark instead of making something actually useful.
SebastianSosa, comment on Hacker News
Another user, bugra_sa, proposed a more practical test than just looking at the published accuracy:
I'd care more about whether Clef knows when to punt than its raw accuracy score. Test it on cases where a false positive is much more expensive than a miss, then change the data enough to see when its confidence falls apart. If it stays confident through that, the benchmark number doesn't mean much.
bugra_sa, comment on Reddit
Bugra_sa's point matters for anyone evaluating Clef before putting it into production: aggregate accuracy on a public benchmark says little about behavior in tail cases, which are precisely where an autonomous AI agent tends to cause the most expensive damage.
What's still missing: self-service fine-tuning and real use cases
Cloudflare said it plans to fine-tune Clef for specific tasks, such as support triage and bot classification, using the company's own historical labeled data to improve accuracy and speed. It also announced a fine-tuning service that lets customers adapt Clef to their own workloads, initially with support from Cloudflare engineers; a self-service platform is planned for a future release, with no firm date announced.
The models are already available via Workers AI and as downloadable weights on Hugging Face, and Cloudflare invites developers to try them out and give feedback. For those evaluating adoption now, this means the production path (human-assisted fine-tuning) doesn't yet have parity with the self-service path the company promises to deliver later, which weighs on any architecture decision that depends on custom model tuning.
Translated from the Brazilian Portuguese original · Read the original
Snapdragon X2 Plus debuts in Lenovo's Yoga Slim 7x with 80 TOPS NPU for local AI
Qualcomm's chip arrives in Lenovo's Yoga Slim 7x with an 80 TOPS NPU dedicated to local AI. For software developers, the advance reopens the question of ARM64 compatibility on Windows.