Skip to content
System One

Jev vs Clef: TypeSafe's decision model against Cloudflare's

Clef is Cloudflare's decision model, launched on 1 October 2026 in two sizes, 27B and 9B, with open Apache 2.0 weights and a hosted version on Workers AI. It takes the same state and typed questions as Jev, adds image input and a 65,536-token context, and costs $0.24 per million input tokens against Jev's $0.042. Every benchmark comparing the two is Cloudflare's own, and Jev wins some of them.

Updated

The short answer

Clef is the closest thing to a drop-in neighbour that Jev has. Cloudflare trained it, hosts it on Workers AI, and says it is “fully Jev-API compatible”, so a request with a state and typed Noul, Choice and Score questions runs on either. Clef also reads images, has twice Jev’s context window as Cloudflare states it, and ships open weights under Apache 2.0.

Jev is cheaper per input token and has a longer public record. Every comparison between the two comes from Cloudflare’s launch post, which Cloudflare ran itself, so read the numbers below as one side’s results. Jev wins several of them.

Jev vs Clef at a glance

Jev Clef Clef-flash
Maker TypeSafe AI Cloudflare Cloudflare
Size and base Not published 27B, frozen Qwen3.8-27B backbone 9B, frozen Qwen3.5-9B backbone
Weights Closed, API only Apache 2.0 on Hugging Face Apache 2.0 on Hugging Face
Hosted API TypeSafe, OpenRouter, Cloudflare Workers AI as @cf/cloudflare/clef Workers AI as @cf/cloudflare/clef-flash
Question types Choice, Score, Noul Choice, Score, Noul Choice, Score, Noul
Input Text Text, JSON, images, video Text, JSON, images, video
Context window 32,000 tokens per Cloudflare’s post 65,536 tokens 65,536 tokens
Questions per request Not checked on this page 1 to 64 1 to 64
Input price per million tokens $0.042, output free $0.24 $0.09
Median latency (Cloudflare’s run) 524.1ms 209.3ms 38.8ms
p95 latency (Cloudflare’s run) 536.0ms 238.6ms 122.4ms

Sources: TypeSafe’s models page, Cloudflare’s launch post and the Workers AI model pages, checked 2 October 2026. The latency figures are the medians and 95th percentiles from Cloudflare’s 43-eval run. TypeSafe’s own claim for Jev is 70ms to 500ms end to end.

What each model does

Both models read one state and answer a set of typed questions with probabilities. Neither writes text. Jev’s three question types are covered in Choice, Score and Noul, and Clef uses the same three with the same field names.

Clef adds an optional images array of up to four PNG, JPEG or WebP images, which the Workers AI docs call a Clef extension to the System One API. It also requires a model field set to clef or clef-flash. A client written for Jev needs those two changes plus a Cloudflare account and token.

How Clef was built

Cloudflare’s post says each model keeps a frozen Qwen backbone and trains a routing head and rank-256 low-rank adapters on top. At inference it makes one prefill pass and then scores every valid answer in parallel, so no text is generated token by token. Training used label-smoothed cross-entropy plus a Brier loss for calibration, synthetic data, and a second stage Cloudflare calls Reinforcement Learning for Calibrated Decisions. Cloudflare does not say that stage matches TypeSafe’s RLCD, which is still unpublished.

Jev’s architecture is not published, so there is no like-for-like description to set beside this. How Jev is built covers what is known.

What Cloudflare’s benchmarks show

All of these are vendor-run: Cloudflare chose the tasks, ran the models and published the tables. Nobody else has reproduced them as of 2 October 2026.

On 10 tasks it picked from the Decision Index:

Task and metric Clef Clef-flash Jev
BFCL, case exact 98.47 98.76 95.75
ToolRet, nDCG@10 69.19 66.43 65.28
API-Bank, accuracy 91.93 93.11 88.19
Home appliances, case exact 82.95 97.73 52.27
When2Call, accuracy 72.37 65.58 80.97
BANKING77, macro-F1 94.20 90.93 79.74
CLINC150+OOS, macro-F1 97.43 66.77 89.27
BRIGHT, nDCG@10 45.91 39.26 47.52
Amazon ESCI, macro-F1 57.48 57.39 55.21
PhishNChips, accuracy 79.60 75.05 62.55

Clef scores higher than Jev on eight of the ten. Jev is ahead on When2Call and on BRIGHT. Clef-flash is below Jev on three tasks, and by the widest margin on CLINC150+OOS.

On TypeSafe’s own workflow evals:

Workflow Clef Clef-flash Jev
Invoice processing 64.7 57.1 61.8
Customer service 76.3 77 76.0
Security incidents 62.9 61.7 61.7
Agent trace observability 68.5 69.8 71.6

Clef is ahead of Jev on three of the four, by 0.3 of a point on customer service, and behind on agent trace observability. Cloudflare also says Clef is currently the leader when evaluated against the Jev Decision Index, and publishes a live leaderboard for it. Clef did not appear on the public Decision Index page as of 2 October 2026.

Speed, cost and hosting

Cloudflare reports a median of 209.3ms for Clef and 38.8ms for Clef-flash against 524.1ms for Jev, across the 43 evals it ran. The post does not say how Jev was called or from where. The Jev median is above TypeSafe’s own upper figure of 500ms, so the gap may partly reflect the network path. Measure both from your own region.

Price runs the other way. Clef is $0.24 per million input tokens and Clef-flash $0.09, against $0.042 for Jev. The Workers AI pages list no output price. Jev’s output is free. Because the weights are Apache 2.0, you can also run either Clef model on your own GPU, which Jev does not allow.

Cloudflare says it does not read, store or train on requests or responses unless you opt into its fine-tuning service. Jev speed and pricing covers Jev’s published figures.

When to use which

  • You need to classify images. Clef takes up to four per request. Jev reads text only.
  • You send long states. Clef’s window is 65,536 tokens against 32,000 for Jev.
  • Cost per call is the main limit. Jev is $0.042 per million input tokens, under half of Clef-flash’s $0.09 and under a fifth of Clef’s $0.24.
  • Latency is the main limit. Cloudflare’s numbers favour Clef-flash by a wide margin, but they are its own. Test in your region.
  • Your tasks look like When2Call or agent traces. Cloudflare’s tables put Jev ahead there. Test on your own data.
  • You want weights you control. Only Clef offers them. See Jev alternatives for other open options.

FAQ

Is Clef a Jev alternative?

Yes, for the request format. Cloudflare says Clef is fully Jev-API compatible, so the same state and Noul, Choice and Score questions run on both. Moving means changing the endpoint, adding a model field and rerunning your own evals. It is not identical. Clef takes images and a longer context, and its prices differ.

Is Clef better than Jev?

Cloudflare’s own benchmarks say Clef is higher on most of the tasks it chose, and lower on When2Call, BRIGHT and agent trace observability. Cloudflare ran them and nobody else has reproduced them. Which is better for you depends on your data, so run both on a sample of it.

Is Clef free or open source?

The weights are free under Apache 2.0 on Hugging Face, for both Clef and Clef-flash. The hosted versions on Workers AI are not free. Clef costs $0.24 per million input tokens and Clef-flash $0.09. Running the weights yourself costs the GPU time, and the 27B model needs a large GPU.

Does Clef read images?

Yes. The Workers AI docs say Clef reads the state as text, JSON, images or video, and accept up to four PNG, JPEG or WebP images per request through an images field. Cloudflare calls that field a Clef extension to the System One API. Jev reads text only, according to Cloudflare’s launch post.

Is Clef faster than Jev?

In Cloudflare’s own run, yes: a median of 209.3ms for Clef and 38.8ms for Clef-flash against 524.1ms for Jev. Cloudflare does not say how it called Jev, and TypeSafe claims 70ms to 500ms end to end. Treat both sets of numbers as vendor claims and time them yourself.

Examples

Introducing Clef: our open-source decision models, and new RL fine-tuning platform

Cloudflare's launch post for Clef and Clef-flash, two decision models on Workers AI that accept the Jev request format and open weights under Apache 2.0. It covers the architecture, Cloudflare's own benchmark tables against Jev, Kev and Laya, and a new reinforcement learning service.

Related guides