Skip to content
System One

Microsoft-Decision-1 vs Clef: Microsoft's decision model against Cloudflare's

Microsoft-Decision-1 and Cloudflare's Clef-flash are decision models built on the same base, Alibaba's Qwen3.5-9B. Microsoft-Decision-1 is closed, text only, has a 32,768-token context and costs $0.042 per million input tokens on Foundry and OpenRouter. Clef and Clef-flash have open Apache 2.0 weights, read images and take Jev's request format, at $0.24 and $0.038 on Workers AI. Neither vendor tested the other's model, and only Clef is on the independent Decision Index, so no head-to-head result exists.

Updated

The short answer

Microsoft-Decision-1 and Clef are decision models from two large cloud companies, and the smaller Clef shares its base with Microsoft’s model. Microsoft post-trained Alibaba’s open-weight Qwen3.5-9B. Cloudflare’s Clef-flash keeps a frozen Qwen3.5-9B backbone and trains a routing head and adapters on top. The full-size Clef uses Qwen3.8-27B.

The two companies made different choices with that base. Clef’s weights are open under Apache 2.0, it reads images, and it takes the same request body as Jev. Microsoft-Decision-1’s weights are closed, it reads text only, and Microsoft has not published its request format. On price they are close: $0.042 per million input tokens for Microsoft’s model, $0.038 for Clef-flash and $0.24 for Clef on Workers AI, with output free on all three.

Do not expect a verdict on quality. Clef is not in Microsoft’s launch charts, and Microsoft-Decision-1 is not in Cloudflare’s. Each vendor ran its own tests on its own suite. The one independent board that lists Clef, the Decision Index, does not list Microsoft-Decision-1.

Microsoft-Decision-1 vs Clef at a glance

Microsoft-Decision-1 Clef Clef-flash
Maker Microsoft Cloudflare Cloudflare
Base model Qwen3.5-9B, post-trained; dense, 5B to 15B parameters per the model card Frozen Qwen3.8-27B backbone, about 27.4B parameters Frozen Qwen3.5-9B backbone, about 9.4B parameters
Weights Closed, API only Apache 2.0 on Hugging Face Apache 2.0 on Hugging Face
Hosted API Microsoft Foundry; OpenRouter as microsoft/microsoft-decision-1, served by Azure Workers AI as @cf/cloudflare/clef; OpenRouter as cloudflare/clef Workers AI as @cf/cloudflare/clef-flash; OpenRouter as cloudflare/clef-flash
Request format Not published Jev’s state and questions, plus a model field and optional images Same as Clef
Question types Yes/no, multiple choice, rating Noul, Choice, Score Noul, Choice, Score
Input Text only Text, JSON, up to 4 images Text, JSON, up to 4 images
Context window 32,768 tokens 65,536 tokens; 16,384 through PrimeIntellect 24,576 tokens on Workers AI; 16,384 through PrimeIntellect
Questions per request Not published 1 to 64 1 to 64
Input price per million tokens $0.042, output free $0.24, output free $0.038, output free
Median latency, vendor-run 85ms (Microsoft, through Foundry) 209.3ms (Cloudflare) 38.8ms (Cloudflare)
p95 latency, vendor-run 125ms 238.6ms 122.4ms
Decision Index 0.3 Full score Not listed 53.08 47.61
Launched 9 October 2026 1 October 2026 1 October 2026

Sources: Microsoft’s launch post and Foundry model card, Cloudflare’s launch post, and the Workers AI model pages, all checked on 10 October 2026. OpenRouter ids, providers and prices come from OpenRouter’s endpoints API on the same day. Each latency figure comes from its own vendor’s harness, so the two sets cannot be compared. Decision Index figures are from edition 0.3, generated on 7 October 2026.

What they share

Both models read some context and a closed question and return a calibrated probability for each allowed answer. Neither writes text. Microsoft lists yes/no, multiple-choice and rating questions, which match Noul, Choice and Score, the three types Clef uses by name. Both vendors aim them at the same jobs: intent and model routing, confidence-gated actions and agent routing and skill selection.

The shared base is Qwen3.5-9B, which the Foundry model card and Cloudflare’s post each name. Microsoft says it will later rebase its model on other models, including MAI and OpenAI’s, so the overlap may not last.

Where they differ

The request. Cloudflare calls Clef “fully Jev-API compatible”. It takes Jev’s state and typed questions, plus a required model field and an optional images array, so a client written for Jev needs small changes. Microsoft has not published the request format, the endpoint path, or limits on questions and options. OpenRouter’s listing of Microsoft-Decision-1 does not describe the request body either. So on 10 October 2026 nobody outside Microsoft can say whether it accepts a Jev request.

Input. Clef and Clef-flash read text, JSON and up to four PNG, JPEG or WebP images per request. Microsoft-Decision-1 reads text only, and its model card says it accepts no image, audio or video.

Context. Microsoft-Decision-1 takes 32,768 tokens. Clef takes 65,536 on Workers AI. The Clef-flash page on Workers AI now gives 24,576 tokens, smaller than Microsoft’s model, while OpenRouter still lists 65,536 for Cloudflare’s Clef-flash endpoint. Through PrimeIntellect on OpenRouter, both Clef models take 16,384.

Weights. You can download either Clef model and run it on your own GPU. Microsoft-Decision-1 runs only as a hosted API.

Scope. Microsoft says its model is tuned for English and may do worse in other languages and in medical, legal and financial work. It also says the model should not be the only basis for decisions about people, such as credit, hiring or housing.

What each vendor’s benchmarks show

Every figure in this section is vendor-run. Microsoft and Cloudflare each picked the tasks, ran the models and published the results.

Microsoft reports 83.5% average accuracy across 36 benchmarks and 147,137 questions. That is against 81.9% for Quyet-1.0-Large, 79.4% for GPT-6 Luna Decisions and 77.2% for H2O-Lightning-4B. Its calibration score is 92.2 out of 100, second to Quyet’s 93.1. On requests altered in eight ways, it changes its answer 1.3% of the time on average. Clef does not appear in any of these charts.

Cloudflare compared Clef with Jev, Kev, Laya and its own earlier DiffusionGemma experiment. On 10 tasks from the Decision Index, Clef scored above Jev on eight. Microsoft-Decision-1 launched eight days after Cloudflare’s post and is not in it.

The two sets of results use different tasks, different metrics and different harnesses. You cannot line Microsoft’s 83.5% up against any Clef number. Latency has the same problem. Microsoft’s 85ms median and Cloudflare’s 38.8ms median for Clef-flash come from different test setups, and Cloudflare’s post does not say what hardware or network path it used.

What independent boards show

The public Decision Index, a third-party board that runs one suite for every entry, added both Clef models in edition 0.3, generated on 7 October 2026. Its Full score takes 20% from public benchmarks, 50% from private tests of the same skills and 30% from private tasks in new domains. Clef scores 53.08 and Clef-flash 47.61. On the public benchmarks alone they score 61.71 and 56.15. Their expected calibration errors are 0.040 and 0.025, where lower is better.

Microsoft-Decision-1 is not on the Decision Index. Benchmark Heaven’s JevBench lists Clef and Clef-flash but not Microsoft’s model, checked on 10 October 2026. Microsoft quotes JevBench latency medians for other models in its launch chart, but the board has no row for Microsoft-Decision-1 itself. Until one of these boards runs it, the only quality figures for it are Microsoft’s.

Price, speed and hosting

Microsoft charges $0.042 per million input tokens with free output, from the launch post. The Foundry model card shows no price. OpenRouter lists the same $0.042 through Azure, with a 32,768-token context.

On Workers AI, Clef costs $0.24 and Clef-flash $0.038 per million input tokens, and the docs say Clef models do not charge for output. Clef-flash was $0.09 when this site first checked on 2 October. OpenRouter now lists both Cloudflare and PrimeIntellect at those same prices, checked on 10 October 2026. So Clef-flash is now slightly cheaper than Microsoft-Decision-1, and Clef costs nearly six times as much.

Cloudflare launched a third model on 9 October 2026, Clef-omni. It is post-trained from Qwen3-Omni-30B-A3B, a mixture-of-experts model with about 3B parameters active per token, and reads audio and video as well as text and images. Workers AI and OpenRouter list it at $0.15 per million input tokens. No outside board has scored it yet.

How both compare with Jev

Jev costs $0.042 per million input tokens with free output, the same as Microsoft-Decision-1. It reads text only and its weights are closed, like Microsoft’s model. Its request format is public, and Clef accepts it, which is why Jev vs Clef calls Clef the closest drop-in neighbour Jev has.

Jev appears in both vendors’ launch material, though not on equal terms. Cloudflare scored Jev on its accuracy tables and its latency chart. Microsoft’s post shows Jev only in its latency chart, at a 240ms JevBench median for Jev 1.13.0, with no accuracy or calibration score. On the Decision Index, Jev’s Full score of 60.11 is above both Clef models. On the public benchmarks alone, the full-size Clef leads Jev, 61.71 to 57.96. Jev vs Microsoft-Decision-1 covers the Microsoft side in detail.

When to use which

  • You need to classify images. Only Clef reads them, up to four per request.
  • You want weights you control. Only Clef offers them, under Apache 2.0.
  • You already send Jev requests. Clef takes the same body with a model field added. Microsoft’s format is unpublished.
  • You build on Azure. Microsoft-Decision-1 is a Direct from Azure model in Foundry.
  • Cost per call is the main limit. Clef-flash at $0.038 and Microsoft-Decision-1 at $0.042 are close. Clef costs $0.24.
  • You send long states. Clef on Workers AI takes 65,536 tokens. Microsoft-Decision-1 takes 32,768, and Clef-flash on Workers AI 24,576.
  • You need quality numbers you can check. Only Clef has independent scores. Run both on a sample of your own data before you choose.

FAQ

Are Microsoft-Decision-1 and Clef the same model?

No. They share a base model, Qwen3.5-9B, in the case of Clef-flash. Microsoft post-trained it and keeps the weights closed. Cloudflare froze it, trained a routing head and adapters on top, and released the weights under Apache 2.0. The full-size Clef uses a different base, Qwen3.8-27B. Microsoft has not published how it trained its model.

Which is more accurate, Microsoft-Decision-1 or Clef?

Nobody has tested them side by side. Clef is not in Microsoft’s launch charts, and Microsoft-Decision-1 is not in Cloudflare’s. Each vendor ran its own suite. The Decision Index scores Clef at 53.08 and Clef-flash at 47.61 but has no row for Microsoft’s model, checked on 10 October 2026. Test both on your own data.

Which is cheaper?

On input price, Clef-flash, at $0.038 per million input tokens on Workers AI against $0.042 for Microsoft-Decision-1. The full Clef costs $0.24. All three charge nothing for output. You can also run either Clef model on your own hardware, since the weights are open, which Microsoft does not allow.

Can I send a Jev request to Microsoft-Decision-1?

Unknown. Microsoft had not published the request format or endpoint path on 10 October 2026, and OpenRouter’s listing does not describe one. Clef does take Jev’s request body, with a model field added and an optional images array. If you need to swap between models without rewriting your client, Clef is the safer choice today.

Is Microsoft-Decision-1 on OpenRouter?

Yes. OpenRouter lists it as microsoft/microsoft-decision-1, served by Azure at $0.042 per million input tokens with a 32,768-token context, checked on 10 October 2026. Clef and Clef-flash are also listed, served by Cloudflare and PrimeIntellect. Cloudflare’s newer Clef Omni appeared on OpenRouter on 9 October 2026.

Examples

Introducing Clef: our open-source decision models, and new RL fine-tuning platform

Cloudflare's launch post for Clef and Clef-flash, two decision models on Workers AI that accept the Jev request format and open weights under Apache 2.0. It covers the architecture, Cloudflare's own benchmark tables against Jev, Kev and Laya, and a new reinforcement learning service.

Microsoft-Decision-1: Our model for fast decision-making

Microsoft's launch post for Microsoft-Decision-1, a Qwen3.5-9B decision model on Foundry at $0.042 per million input tokens. It reports Microsoft's own accuracy, calibration and latency charts against Quyet-1.0-Large, GPT-6 Luna Decisions and Jev, plus internal use at Xbox Research and Copilot.

Related guides