Cloudflare trains its first AI models: Clef and Clef-flash, open decision models for agents
Cloudflare's Clef models skip text generation and return calibrated probabilities over predefined answers, with open Apache 2.0 weights on Hugging Face.
Cloudflare released its first models trained in-house on October 1: Clef and Clef-flash, decision models that skip text generation entirely and return calibrated probabilities over predefined answers. Send the model an input state and a set of typed questions, and it answers with numbers: how likely the answer is yes, which of your options fits best, or where the input sits on a rubric you defined. No free-form output to parse, no reasoning tokens to wait for.
Clef is the 27-billion-parameter precision model, post-trained on Qwen3.8-27B; Clef-flash is the 9-billion-parameter sibling on Qwen3.5-9B, aimed at latency-critical paths. Both have a 64K context window and a vision encoder that accepts up to four images alongside the state, which sets them apart from TypeSafe's Jev, the text-only decision model that kicked off the category on September 15.
The weights are open-sourced on Hugging Face under Apache 2.0, and both models run on Workers AI, where they can be called through the standard binding or REST API. They speak TypeSafe's System One API, so a team with an existing Jev integration can switch by changing the endpoint and model name. Cloudflare also launched a reinforcement learning fine-tuning service, delivered first through its forward-deployed engineer team and later as a self-serve platform built on AI Gateway, Workers AI and Containers.
Cloudflare's benchmark numbers are the headline: across 43 of its own runs, median latency was 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash, against 524.1 milliseconds for Jev, or 2.5x and 13x faster at the median. On BANKING77, Clef scored 94.20 macro-F1 against 79.74 for Jev; on BFCL case-exact, 98.47 against 95.75. Cloudflare says a Clef model scores highest on 7 of 10 decision benchmarks. It also published a workflow test: classifying a website domain by fetching, rendering and categorizing it took Clef 2.2 seconds with Browser Run, against 4.7 seconds for gpt-oss-120b in the same workflow.
The timing is pointed. OpenAI mentioned a Decisions API in limited preview at DevDay on September 29, and Amazon shipped its own open decision model, Strands Decider 2B, the same day as Clef. Cloudflare is the only one of the three giving away the weights. The launch post even explains the name: a decision model is like a musical clef, defining the domain of the context and the actions that follow, with CF nodding to Cloudflare.
As with any vendor-run benchmark, treat the comparisons as Cloudflare's case, not an independent verdict. The weights being on Hugging Face means developers can run their own evaluations.