Cloudflare Clef Decision Models: A Fixed-Choice Answer in Under 40 Milliseconds

Cloudflare Launches Clef and Clef-flash: Open-Source Models That Pick a Decision Instead of Writing Text
The short answer: On October 1, 2026, Cloudflare announced Clef and Clef-flash, two open-source "decision models" released under the Apache 2.0 license. The model does not write a free-form reply. It reads the input, chooses from a list of answers you define in advance, and returns a probability for each option. The lighter Clef-flash is built on a 9-billion-parameter Qwen 3.5 model and has a median response time of 38.8 milliseconds. The more precise Clef is built on a 27-billion-parameter Qwen 3.8 model, with a median of 209.3 milliseconds. Both are available on Workers AI, and their weights are published on Hugging Face for anyone who wants to run them on their own servers.
Much of the AI in customer service is really classification: complaint or order? Urgent or not? Many teams use a general chat model for this and get slower answers that can change between runs.
What a Decision Model Actually Is
Cloudflare describes a decision model as one that produces bounded structured outputs "cheaply, quickly and consistently". It contrasts this with large language models, which it describes as largely non-deterministic.
According to the model card on Hugging Face, you send a "state" describing the situation, as text or JSON, plus images or video if needed, then a list of questions, each with allowed options. The model returns a probability for each option and generates no free-form text. Cloudflare says your code uses these probabilities to route the ticket, trigger an escalation, or defer to a human.
The Numbers as Cloudflare Published Them
- Capabilities: a 64k-token context window according to the announcement (the model card sets a default input limit of 16,384 tokens, which can be changed), and both versions can read and classify images.
- Median latency: 38.8 milliseconds for Clef-flash and 209.3 milliseconds for Clef, against 524.1 milliseconds for Jev, a competing model from Typesafe AI.
- Internal test: on Cloudflare's threat intelligence team, fetching, rendering and classifying a website took 2.2 seconds with Clef, versus 4.7 seconds with gpt-oss-120b, its fastest general LLM, which returned only two classifications.
- Slowest requests: 95% of requests finish within 122.4 milliseconds with Clef-flash and within 238.6 milliseconds with Clef.
On CLINC150, an intent classification test that also checks for out-of-scope requests, Clef scored 97.43 macro-F1 against 66.77 for Clef-flash. On BANKING77, which covers banking customer queries, the two were close: 94.20 versus 90.93, and on API-Bank Clef-flash came out ahead with 93.11 accuracy versus 91.93. That is why Cloudflare calls Clef its precision model and Clef-flash the better fit for latency-critical decisions. The biggest gap shows up when out-of-scope requests must be caught.
Where This Helps Saudi Businesses
Use cases Cloudflare names include routing support tickets by urgency and team, classifying website domains for its threat intelligence team, and telling good bots from bad ones, while invoice processing and security incidents appear among its tests. In a Saudi business, that looks like this:
- WhatsApp triage: every message is instantly classified as an order, a complaint, a price inquiry or spam, then sent to the right employee.
- Support tickets: set the urgency and the responsible department before any employee opens the ticket.
- Documents in the ERP: identify the type of incoming document, whether a supplier invoice, a purchase order or a credit note, and send it down the right path.
- Protecting forms and stores: spot automated sign-ups and orders on contact forms and e-commerce stores.
Running It on Your Own Servers Keeps the Data in the Kingdom
The most important feature for the Saudi market is not speed. It is that the weights are published and can be run inside the Kingdom. Customer messages and supplier invoices contain personal and commercial data, and the Personal Data Protection Law places controls on transferring personal data outside the Kingdom. A model running in a local data center avoids the cross-border transfer question for classification tasks, while the law's other obligations still apply. The model card says it was tested on a single H200 GPU, so plan your hardware early.
Limits to Know Before You Start
- Arabic is not mentioned: neither the announcement nor the model cards list supported languages. Test it on real Arabic messages in Saudi dialect.
- No price in the announcement: Cloudflare did not publish the cost of using the two models on Workers AI in its blog post, so check the pricing page before estimating costs.
- Customization is not self-serve yet: tuning the model on your data with reinforcement learning currently goes through a team of Cloudflare engineers, with a self-serve platform coming later.
How to Build With It, Step by Step
- Inventory your decisions: find every question in your systems that has limited answers.
- Collect a real sample: past messages and tickets your staff classified by hand.
- Compare three options: Clef-flash, Clef and your current solution, on accuracy and response time.
- Set a confidence threshold: if the top option's probability is high, the system acts automatically; if it is low, the case goes to an employee.
- Split the work: the decision model picks the path, and the large language model writes the reply only when needed.
We apply this pattern at Origami when building customer service and process automation systems: a small model decides, a large model writes, and an employee reviews unclear cases. The Clef models make the first part easier, and it can stay inside the Kingdom.
Sources
Frequently asked questions
How is Clef different from a chat model like ChatGPT?+
Clef does not write text. It picks from options you define and returns a probability for each one, which makes it faster and more consistent for classification and routing. The chat model stays in charge of writing and drafting replies.
Does Clef support Arabic?+
Cloudflare did not list supported languages in the announcement or in the model cards on Hugging Face. Test it on a sample of your Arabic customer messages, including Saudi dialect, before relying on it.
Can I run Clef on my own servers inside Saudi Arabia?+
Yes. The weights are published on Hugging Face under the Apache 2.0 license, which permits commercial use. You will need a server with a GPU, and the model card says it was tested on a single H200.
Which should I choose for my system, Clef or Clef-flash?+
Clef-flash is much faster, with a median of 38.8 milliseconds, comes close to Clef on tests like BANKING77 and even beats it on API-Bank. Clef is clearly more accurate when out-of-scope requests must be caught, so compare both on your real data.
Follow Origami in Google
Pin Origami as a preferred source and our articles will surface first for you in Google Search and Top Stories.

Related articles
- Artificial IntelligenceGoogle Hands Gemini 4 Argon to Cyber Defenders First, With a 1M-Token Output LimitGoogle's Gemini 4 Argon finds and patches vulnerabilities on its own and outputs up to 1M tokens at $2/$10 per million. Who gets it first, and how to prepare.
- Artificial IntelligenceClaude Marketplace Puts 2,000 Connectors One Click Away From Your Company DataAnthropic's Claude Marketplace launched with 2,000+ connectors, plugins and agents. What to check before linking Claude to your systems, and when to build your own.
- Artificial IntelligenceOpenAI's Dots Work While You Sleep: What Do They Do Without Asking You?OpenAI's dots are always-on AI agents with their own cloud computer and 4,000+ app connections. What they do alone, what needs your approval, and how to set them up.
- Artificial IntelligenceEleven v4 Speaks Arabic in 100 Milliseconds. The Hard Part Is What It Tells Your CustomerElevenLabs' Eleven v4 and v4 Turbo bring ~100 ms voice generation and 90+ languages including Arabic. Prices, limits, and what to fix before a voice agent answers customers.
- Artificial IntelligenceXiaomi Opens MiMo-V2.6 Under MIT, and Its Light Model Reads a Million Tokens for 14 CentsXiaomi released MiMo-V2.6 open source under MIT: one model for text, audio, images and video, with a 1M-token context and prices from $0.14 per million tokens.
- Artificial IntelligenceAn Open Agent Built for Days of Work, Not Minutes: Atria Dawn and Its Real Running BillShanghai AI Lab released Atria Dawn under MIT: 744 billion parameters aimed at long multi-step tasks. What it is actually good for, and what running it in-house costs.
Have a project in mind?
We build custom systems, apps and websites for your business. Tell us your idea and we will give you a straight answer on it.
