Cloudflare Releases Clef and Clef-flash, Open-Weight Decision Models That Return Typed Probabilities
Cloudflare has released Clef and Clef-flash, the Workers AI team's first trained models: Apache 2.0 decision models that return probabilities for typed questions instead of generated text and are compatible with TypeSafe AI's Jev API.
The models implement the same API as Jev, the "System One" model that TypeSafe AI launched on Sept. 15, 2026, and open alternatives such as Kev-9B and Laya followed with the same interface. Switching from Jev requires changing the endpoint and the model name. Clef supports three question types: noul, which returns the probability of yes; choice, which picks one named option and returns per-option probabilities plus a confidence value; and score, which rates an input against an ordered rubric and returns a probability-weighted score. A single request on Workers AI can carry up to 64 questions and up to four images.
Clef is post-trained from Qwen3.8-27B and Clef-flash from Qwen3.5-9B, and both keep the backbone's vision encoder. Inference runs in two stages: the backbone performs a single prefill-only pass over the state and questions, then a small transformer called the joint schema head reads the final hidden states, routes evidence to each question, lets fields cross-attend and scores all options jointly. A per-question softmax converts logits into probabilities. Training froze both backbones and jointly optimized the routing head with rank-256 low-rank adapters; the loss pairs label-smoothed cross-entropy with a Brier loss for calibration, and a secondary objective, Reinforcement Learning for Calibrated Decisions, gives partial credit to adjacent ordinal choices.
On Cloudflare's 10-benchmark shortlist from the Decision Index 0.2.1 suite, a Clef model scored highest on seven. On BANKING77, Clef reached 94.20 macro-F1 against Jev's 79.74; on CLINC150+OOS it scored 97.43 against 89.27; and on the home appliances set Clef-flash reached 97.73 case-exact against Jev's 52.27. The model card shows Jev ahead on the knowledge-heavy tests: GPQA Diamond at 78.3 against 48.0, MMLU-Pro at 82.7 against 65.9 and BBH at 92.9 against 73.7.
On TypeSafe's own workflow evaluations, Clef beat Jev in three of four areas by small margins: invoice processing at 64.7 against 61.8, customer service at 76.3 against 76.0 and security incidents at 62.9 against 61.7. Jev led agent trace observability at 71.6 against 68.5. In Cloudflare's threat intelligence workflow, Clef classified a domain in 2.2 seconds, the company reported, while gpt-oss-120b took 4.7 seconds. All figures are vendor-reported and have not yet been independently replicated.
Cloudflare reports median latencies of 209.3 ms for Clef, 38.8 ms for Clef-flash and 524.1 ms for Jev, and hosted input prices of $0.24, $0.09 and $0.042 per million tokens respectively. Clef and Clef-flash accept image input with a 65,536-token context window; Jev's listing shows a 32K context per Cloudflare and no image input. Among the other open implementations of the System One API, Jared Palmer's Kev-9B is a 9B model with a 45.4M LoRA and a median latency of 51.4 ms, and Convai Innovations' Laya is a 421M model at 5.8 ms.
Both models can be called through the Workers AI binding, the REST API or AI Gateway, and the model cards list testing on a single H200 with BF16 weights. Cloudflare also announced a reinforcement learning service for tuning Clef on private data. It starts with the company's forward-deployed engineers, with a self-serve platform planned later; the pipeline combines AI Gateway, Workers AI, Containers and a new Trainer component, and teams can apply through a design partner form.