AI Affairs, home

Saturday 3 October 2026

Technology

Cloudflare releases Clef decision models and a fine-tuning product

The Apache 2.0-licensed models run on Workers AI. Cloudflare reports faster responses than Jev in its evaluations and is offering hands-on help to customise Clef.

Rows of colourful lava lamps line shelves at Cloudflare.
Photo: F ASTILY, CC BY-SA 4.0, via Wikimedia Commons (cropped)

Cloudflare released Clef and Clef-flash, two decision models hosted on Workers AI, and introduced a reinforcement learning product for fine-tuning Clef on 1 October 2026, according to its announcement. Both models are available on Hugging Face under an Apache 2.0 licence. With the release, Cloudflare is offering its own models alongside the infrastructure on which customers can run them.

Key points

  • Clef and Clef-flash produce structured answers for classification tasks and are compatible with the Jev API.
  • In Cloudflare’s evaluations, both models scored above Jev on BFCL case exact, while Jev led on some other tasks.
  • Cloudflare is offering hands-on fine-tuning assistance for customers adapting Clef to their own workloads.

Clef returns typed answers with probabilities

A decision model takes information and answers a set of defined questions about it. For a support message, those questions might ask whether the request is urgent and which team should receive it. Clef returns answers in specified types, together with probabilities, so software can use them to choose a route or defer the decision to a human. A general language model can produce an open-ended reply instead, which gives it more room to respond but is a different shape of output for a task requiring set categories.

Sorting a customer support message could use an urgency answer and a suggested destination to route the request, escalate it or leave it for a human. The probabilities would accompany those bounded answers rather than a free-form reply.

Cloudflare says Clef has a vision encoder, allowing it to classify images as well as text. Typesafe AI’s Jev currently classifies text, according to Cloudflare. Clef also accepts a 64k context window, against the 32k window Cloudflare gives for Jev. That window governs how much material can be supplied for a classification, rather than how many categories the model can return.

Cloudflare says its Threat Intelligence team tested Clef on website domains using Browser Run. Fetching and rendering a site and then classifying it took 2.2 seconds with Clef, compared with 4.7 seconds for gpt-oss-120b in the same workflow. Cloudflare says the general language model returned only two classifications in that test.

BFCL scores favour Clef, but not every test does

Cloudflare evaluated the models on tasks it selected from the Jev Decision Index. On BFCL case exact, Clef scored 98.47 and Clef-flash scored 98.76, against 95.75 for Jev. These are Cloudflare’s measurements on named evaluation tasks, rather than timings from the domain-classification workflow.

The ordering changes across the tests. On When2Call accuracy, Cloudflare reports 80.97 for Jev, 72.37 for Clef and 65.58 for Clef-flash. Jev also scores above both Clef models on BRIGHT, which the table measures using nDCG@10. Cloudflare says it ran a separate set of evaluations drawn from Typesafe AI’s own suite, where its models beat Jev in three of four areas.

For latency, Cloudflare reports results across 43 evaluation benchmarks. Its median figures are 209.3 milliseconds for Clef, 38.8 milliseconds for Clef-flash and 524.1 milliseconds for Jev. At the 95th percentile, the reported figures are 238.6 milliseconds, 122.4 milliseconds and 536.0 milliseconds respectively. Those timings concern the benchmark runs; Cloudflare’s website-domain test also included fetching and rendering the site.

Speed is relevant when a classification sits inside a longer sequence of actions, because the next action must wait for an answer. Cloudflare proposes putting Clef in that decision-making path alongside a language model on Workers AI. Its domain test gives one example of the complete sequence, while the evaluation table measures the models on a broader set of tasks with differing scores.

Workers AI hosts Clef and Clef-flash

Cloudflare hosts both models on Workers AI and says its edge GPUs help reduce network latency. The models are also fully compatible with the Jev API, giving developers an existing interface through which to try them. The Apache 2.0 release on Hugging Face provides another route: running the models locally rather than using Cloudflare’s hosted service.

The two versions serve different priorities in Cloudflare’s account of the release. It presents Clef as the option for tasks needing greater precision and Clef-flash as the option for decisions where latency matters most. The benchmark results put that choice in more concrete terms: Clef-flash has the lower reported median latency, although neither model leads on every evaluation task.

Cloudflare says it does not read, store or train on requests and responses sent to the hosted models unless a customer uses its fine-tuning product. Fine-tuning is intended to adapt Clef to a particular workload through reinforcement learning. Cloudflare says its fine-tuning service begins with hands-on assistance from its forward-deployed engineer team.

Topics: Agents, Foundation models, Open source