Over the final few weeks, there has been many of buzz about decision models specified as Typesafe AI’s Jev System One model. While classifier models have been about for several time, Jev introduces a new decision example idea into the earth of AI — a example that produces bounded organized outputs cheaply, quickly and consistently that can be added into a workflow whenever a decision is required. These models are capable adequate to activity complete any set of inputs without continually retraining the example to merge new classification categories. This contrasts alongside the earth of Large Language Models (LLMs), which are mostly non-deterministic, but are open-ended adequate to logic and create content and tool calls for agentic workloads.
Today, we’re releasing two Cloudflare-trained decision models, Clef and Clef-flash, hosted on Workers AI. Clef is currently the chief whenever evaluated against the Jev Decision Index, you can perspective complete results on the live benchmark demo site. These models are smarter, faster, and completely Jev-API compatible, so you can test alongside these hosted models easily. We’re completely open-sourcing these models on Hugging Face under an Apache 2.0 licence for you to run locally and test alongside yourselves.

Lastly, we’re enthusiastic to debut our new reinforcement learning (RL) product, which allows customers to fine-tune Clef to lawsuit their use cases as well.
What is a decision model?
A decision example makes classifications to assistance agents decide how to act, according to certain probabilities. For example, you can continue in a client assistance communication (inputs) and ask if it is urgent and which squad should grip it. A decision example volition come back typed answers alongside probabilities (outputs), which your code can use to path the ticket, trigger an escalation, or defer to a human. This method that a individual does not necessarily need to be in the iteration for agentic decisions anymore — agents can programmatically collect context, create decisions, and obtain actions on tasks, or defer to a individual whenever needed.

Specifically at Cloudflare, we’ve been evaluation our new Clef example on our Threat Intelligence squad to assistance us categorize website domains. By giving a domain to Clef (with Browser Run) it can quickly acknowledge categories that the domain falls under — for example, it power categorize a domain alongside a 95% chance it is a manner website, 85% ecommerce, <1% phishing, etc. This classification took our Clef example 2.2s to fetch, render, and categorize the website. In contrast, our fastest broad LLM gpt-oss-120b took 4.7s in the identical workflow, and lone returned two classifications. As a user, you can ideate how a 2x funds in latency and results can assistance us enhance our danger intellect workflows and be faster in identifying malicious or lawful domains. Generalize this to any use case anywhere you need to create quick programmatic decisions, and you unlock mighty agentic workflows that are capable to autonomously decide, reason, and execute.
In music theory, a clef is a sign placed at the commencement of a musical personnel that assigns particular throw names to the lines and spaces. A decision example is akin to a music clef since it helps define the domain of the environment and the consequent notes (actions) that prosecute it. We chose Clef as the name of our family of decision models, as it serves akin purposes, and the CF hearkens to Cloudflare.
How is Clef distinct from another decision models?
Although the market is getting increasingly soaked alongside decision models, Clef has several distinctive properties that create us enthusiastic to publish it to the public. First, it has a imagination encoder so it’s capable to obtain in images and categorize ocular content. This is distinct from Jev, which lone does content classification today. Secondly, our example has a 64k environment opening (compared to Jev’s 32k), which allows users to compression additional input province for the example to categorize against.
Third, our example is exact and powerful, scoring competitively against another decision models on the market throughout assorted norm benchmarks. We shortlisted several evaluations below that are crucial for decision-making as defined by the Jev Decision Index and scored several of the additional famous models on the market for it. Check out the array below for benchmarks, or perspective the scores on our live decision indicator demo site:
Benchmark | DiffusionGemma Jev | |||||
BFCL · case exact | 98.47 | 98.76 | 95.75 | 96.52 | 94.51 | 38.13 |
ToolRet · nDCG@10 | 69.19 | 66.43 | 65.28 | 61.21 | 64.26 | 12.69 |
API-Bank · accuracy | 91.93 | 93.11 | 88.19 | 83.66 | 56.30 | 11.41 |
Home appliances · case exact | 82.95 | 97.73 | 52.27 | 42.05 | 25.00 | 0.00 |
When2Call · accuracy | 72.37 | 65.58 | 80.97 | 75.44 | 49.62 | 11.94 |
BANKING77 · macro-F1 | 94.20 | 90.93 | 79.74 | 74.28 | 84.83 | 14.29 |
CLINC150+OOS · macro-F1 | 97.43 | 66.77 | 89.27 | 83.49 | 79.03 | 3.19 |
BRIGHT · nDCG@10 | 45.91 | 39.26 | 47.52 | 42.94 | 38.53 | 19.90 |
Amazon ESCI · macro-F1 | 57.48 | 57.39 | 55.21 | 53.37 | 49.22 | 24.40 |
PhishNChips · accuracy | 79.60 | 75.05 | 62.55 | 85.35 | 50.75 | 50.15 |
We additionally ran benchmarks throughout Typesafe’s own eval suite and our Clef models fared well, beating Jev in 3 out of 4 areas. Notably, our Clef-flash performs exceptionally well, stated how much faster it is.
Workflow | |||
Invoice processing | 64.7 | 57.1 | 61.8 |
Customer service | 76.3 | 77 | 76.0 |
Security incidents | 62.9 | 61.7 | 61.7 |
Agent trace observability | 68.5 | 69.8 | 71.6 |
Across the 43 eval benchmarks that we ran, our Clef models attack the decision models on latency (except for Laya which is extremely accelerated but trades off norm in the benchmarks above):
Benchmark | DiffusionGemma Jev | |||||
Median latency · ms | 209.3 | 38.8 | 524.1 | 84.4 | 51.4 | 5.8 |
p95 latency · ms | 238.6 | 122.4 | 536.0 | 211.2 | 187.9 | 222.5 |
On top of the latency benefits from the example itself, our Clef models are hosted on Workers AI. Because they are hosted on Cloudflare’s infrastructure, we’re capable to obtain advantage of our GPUs at the edge, foremost to low network latency and faster decisions. This method that you could put Clef into the hot way for agents to create decisions and merge that alongside one of our LLMs on Workers AI to obtain action.
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \\ -X POST \\ -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \\ -d '{ "model": "clef", "state": "Checkout has been failing for all client for the final hour.", "questions": { "urgent": { "type": "noul", "instructions": "Is this assistance petition urgent?" }, "team": { "type": "choice", "instructions": "Which squad should grip this request?", "criteria": { "billing": "Payments, invoices, and refunds", "technical": "Outages, errors, and configuration", "sales": "Plans and upgrades" } }, "severity": { "type": "score", "instructions": "How serious is the client impact?", "criteria": ["No impact", "Minor", "Major", "Critical"] } } }'
Clef additionally produces strictly typed outputs akin to Jev and is completely API-compatible, so you can create the toggle extremely easily. The larger Clef example is your additional mighty precision model, during the Clef-Flash example is awesome for latency-critical decisions. The models are enterprise-ready alongside our justify that we don’t read, store, or train on your requests or responses (unless you desire to use our fine-tuning product, which we go into below). You can get started alongside the Clef models today, starting alongside our developer documentation or perform about alongside the open-source example on the Hugging Face repo.
If you’d akin assistance tuning Clef for a particular workload, we are additionally offering fine-tuning services — archetypal as a hands-on partner alongside our forward-deployed engineer (FDE) team, and afterward afterward as a self-serve fine-tuning phase for customers to train and redeploy the example onto Cloudflare.
How we trained Clef
In the identical week that Jev came out, we posted concerning several experiments we had alongside our own homegrown decision model. Our demo goes into how we adapted the DiffusionGemma example to output deterministic probabilities by exposing the logprobs that are generated by a ample tongue model. Our first method built upon autonomous investigation by Matt Mastracci, who has been energetic in the device learning (ML) community alongside sharing new ideas and pull requests to vLLM conclusion motor to create DiffusionGemma assistance stronger.
Clef builds upon this concept, but uses a distinct basis example as the backbone. We currently use Qwen as the basis example and post-trained it to lawsuit decision example use cases. During inference, Clef uses Qwen for a prefill-only pass, afterward scores the valid schema choices in parallel. The decision stage is non-autoregressive, so there’s no intermediate content to create token by token, making Clef considerably faster than autoregressive LLMs. Rather than generating intermediate content to create organized answers, Clef and Clef-flash get schema choices immediately from inner backbone representations. This method relies on a specialized two-stage notice routing process: all valid choice extracts environment applicable to the prompt, allowing idiosyncratic site parameters to cross-attend alongside another sectors and rear to the first payload previous to scoring. By leveraging a lexical prior, the example preserves semantic intent throughout options. Ultimately, the architecture unites option-specific evidence routing, shared cross-field attention, and schema-bound scoring.
By freezing Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, we jointly optimized the routing caput alongside rank-256 low-rank adapters. Our post-training utilizes label-smoothed cross-entropy for valid schema outputs paired alongside a Brier defeat to refine probability calibration. This training leverages our own inner synthetic datasets permutating site orders, prompts, and schema structures. We additionally developed Reinforcement Learning for Calibrated Decisions (RLCD) to assist as a secondary optimization target, granting partial credit to neighboring ordinal choices, rewarding completely exact document outputs, and applying a citation penalty to forestall allocation shift, giving us improved accuracy and generalization.
This method that we were capable to accomplish a few novel things alongside Clef: we improved accuracy of the example in classification, constrained it to output lone probabilities alternatively of content generation, and made it faster than Jev and the basis Qwen models.
How fine-tuning can broaden the capabilities of Clef
We heard a lot of inner use cases that required fine-tuning our Clef example to be built into our agentic workflows at Cloudflare. For example, inner teams desire a classifier example to be capable to measure Trust & Safety submissions, assistance us triage Cloudflare Support requests, or equal to be built-in to our Bot products to decide if a crawler is a fine bot or bad bot.
These use cases are incredibly particular and we have had many years of labelled decisions that we could use to train a particular classifier. When you fine-tune a model, you may provision up several broad intent achievement in toggle for higher accuracy in a particular domain.. Because Cloudflare has additional than 15 years of network data throughout distinct domains, we can fine-tune a example to fit these particular use cases which is additional exact and faster than our generic Clef model. We’re operating alongside inner teams already to fig out how we can post-train Clef to create mighty ML models that boost our effect and enhance workflows throughout Cloudflare. These inner teams and use cases are the next remit of our new FDE fine-tuning squad and basis for our reinforcement learning (RL) product.
Our new RL service
We are offering a assistance to assistance customers fine-tune Clef to lawsuit their workloads alongside our hands-on FDE team. From that, we’ll study from our hands-on experiences to build a self-serve phase that customers can use to grasp data, fine-tune, and redeploy the model, all on Cloudflare.
This has really been a lengthy period coming — we’ve been construction our AI phase to have the correct primitives anywhere we could be construction a tradition RL product. The involvement in Jev shows the need for a fast, small, specific, classifier model, and we chose this to be our niche to commencement experimenting alongside RL environments.
To do this, we leverage the primitives that we already have built on our Cloudflare platform:
- Cloudflare AI Gateway – continue all your AI traffic through AI Gateway and automatically create a dataset of requests for your use case
- Cloudflare Workers AI – create rollouts against the basis Clef model
- Cloudflare Containers – RL sandbox for scoring and replaying delegate actions
- [NEW] Trainer – update weights of fine-tuned Clef model
- Cloudflare Workers AI + BYO Model – redeploy the fine-tuned example on Workers AI
This combines a few work-in-progress pieces of the AI Platform that we’ve been operating on, including AI Gateway that captures your AI traffic so you can leverage your own request/response data, Containers for RL Sandboxes, and Workers AI’s Bring Your Own Model (Cog) activity that has been progressing since our acquisition of Replicate.

Try it out today
We’re enthusiastic to initiate our archetypal Cloudflare-trained ML example from the Workers AI squad today. We’re motionless first current and have a lot additional improvements in store, but it is a fantastic archetypal showcase of the difficult activity we’ve been doing on the AI Platform team. We accept that Clef has the capability to disrupt the way we use agents, which fits naturally into Cloudflare’s goal of being the delegate cloud.
If you have particular use cases and are already customers of these products — we’d affection to conversation alongside you and be scheme partners as we test in this space.
Try out the Clef models hosted on Workers AI, download the weights on Hugging Face if you’d akin to examine for yourself, and attain out if you have fine-tuning use cases you’d akin us to assistance with.
Our ML squad has been expanding in impact, from example optimizations to example training research. If you’re curious in joining our mission, check out our open roles.