The Apache 2.0 release covers routing, triage, and schema selection: small classifiers meant to replace the slow, expensive LLM call a typical agent makes for every routine choice.
When an agent has to decide whether a support ticket is urgent or whether a domain is a phishing page, the choice is not really reasoning. It is classification. Today that classification usually costs an LLM call: slow, expensive, and inconsistent.
Cloudflare wants to swap that call for a smaller class of model it calls a "decision model": a compact, bounded-output classifier that returns a label and a probability, trained to slot into an agent workflow the way a router hands a request to the right backend. The company has open-sourced two such models, Clef and Clef-flash, on Hugging Face under Apache 2.0, and published a blog post describing how they are built and where Cloudflare itself is using them.
The release is the first time a major infrastructure vendor has open-sourced this category rather than gating it behind a paid API. The combination also makes custom local classifiers feasible for a small team: a builder can download the weights, run them on a single H200, and fine-tune a custom model without paying per call to a frontier system.
Clef and Clef-flash share a 27-billion-parameter Qwen3.8 backbone with a vision encoder and a joint schema head. At inference the model runs a prefill pass on the prompt once, then scores multiple candidate answers in parallel using cross-field attention. Cloudflare says the routing head and rank-256 adapters were trained with a mix of classification, calibration, and reinforcement-learning objectives, and the request schema accepts one to 64 typed questions plus up to four embedded images.
The clearest production example comes from Cloudflare's own Threat Intelligence team. In the company's description, the team uses Clef to classify a domain as "benign," "suspicious," or "malicious" and return calibrated probabilities, with Cloudflare framing the output shape as roughly 95 percent confidence on clear-cut benign traffic, 85 percent on harder cases, and under 1 percent on the long tail where the model abstains and a human takes over.
Cloudflare's own Decision Index 0.2.1 run, published on the model card, reports a median request latency of 209.3 ms for Clef and 38.8 ms for Clef-flash. The same internal run lists GPQA Diamond accuracy of 48.0 for Clef, 51.0 for Clef-flash, and 78.3 for a rival system, Typesafe AI's Jev System One. RAGTruth F1 scores land at 79.4, 35.6, and 76.5. None of these numbers are independently reproduced, and Jev outperforms Clef on the reasoning benchmark, but the table is also a market signal: Jev wins on reasoning, Clef-flash wins on latency.
The latency story is harder to read. The blog cites a 2.2-second end-to-end run for one internal domain-classification workflow with Clef versus 4.7 seconds with gpt-oss-120b, but that figure is an illustrative vendor observation, not a controlled study. The same caveats that apply to any LLM latency number apply here: batching, network, hardware, and input shape.
Two service limits are worth pinning down. The hosted endpoint advertises a 65,536-token context window and $0.24 per million input tokens, with the 1-to-64 question schema above. The model card, by contrast, defaults to a 16,384-token encoder in its local usage instructions and notes video support, which the hosted request schema does not currently advertise. A builder who plans to run Clef on a single GPU should size their context to the local default, not the hosted number.
The fine-tuning story is the access beat. Cloudflare is selling a managed service today where its Forward Deployed Engineers help customers adapt Clef to their own decision categories, with a self-service RL fine-tuning platform described as future work. Once that self-service path lands, the combination of Apache-2.0 weights and a hosted fine-tuning job puts custom local classifiers within reach of a team that does not have a research lab.
There are real limits. Decision models are bounded by their training categories, and Cloudflare keeps a human in the loop whenever confidence is low or the stakes are high. The benchmarks are vendor-run, the production numbers are illustrative, and the rival on the reasoning test is not the rival on the latency test. For the first time, the weights, the code, and the path to a custom fine-tune are all in the open.