
Post on X
- Views
- 41,477
- Likes
- 377
- Reposts
- 36
Read from X on . Counts change daily.
Cloudflare's live leaderboard for decision models on the Decision Index suite, with Clef, Clef-flash, Jev, Kev-9B and Laya plotted by score and latency. Cloudflare built and ran it, so treat the ranking as vendor-run.

Read from X on . Counts change daily.

Cloudflare's launch post for Clef and Clef-flash, two decision models on Workers AI that accept the Jev request format and open weights under Apache 2.0. It covers the architecture, Cloudflare's own benchmark tables against Jev, Kev and Laya, and a new reinforcement learning service.

AI News daily roundup leading with the Jev launch, summarising RLCD and TypeSafe's 20-200x speed and 40-400x cost claims. Its recap of reactions notes posters using Jev as a structured classifier, judge and routing policy where text generation is unnecessary.
Argues Jev's speed comes mainly from constrained single-token generation and aggressive prefill, techniques reproducible with existing LLMs, rather than a new architecture. Still treats very fast structured output as a new primitive, while doubting Jev matches frontier models on raw intelligence.
An independent, evidence-based map of where Jev's calibrated-decision claim holds up and where it breaks down, built from real API-call receipts rather than a leaderboard.
An adversarial evaluation of Jev with nine experiments and 28 predictions fixed before any data was collected. One run covered 123,805 requests over 138 minutes for $12.69, with five failures, plus a 13-rule prompting guide.

A 49-task, 8-site browser-agent benchmark. Jev paired with a small model for argument generation solved 100% of tasks at far lower cost than screenshot-based computer use, and adding WebMCP tool exposure roughly doubled Jev's own solve rate from 25 to 49 tasks.
239k viewsSolve rate 25/49 to 49/49Cost vs Astra 112x lower