Skip to content
System One

Jev vs Microsoft-Decision-1: same price, different evidence

Microsoft-Decision-1 is Microsoft's decision model, announced on 9 October 2026 and post-trained from Qwen3.5-9B. Like Jev, it costs $0.042 per million input tokens with free output, has closed weights and answers yes/no, multiple-choice and rating questions with probabilities. It takes 32,768 tokens of text per request against Jev's 64,000. Its accuracy and calibration figures come from Microsoft's own tests, which give Jev no accuracy score. Neither the Decision Index nor JevBench lists it as of 10 October 2026.

Updated

The short answer

Microsoft-Decision-1 and Jev cost the same: $0.042 per million input tokens, with output free. Both are closed-weight, hosted-only decision models. You send text and a question with fixed answer options, and each returns a probability per option instead of writing text. Both handle yes/no, multiple-choice and rating questions, which Jev calls Noul, Choice and Score.

The difference is in what you can check. Jev has a published endpoint, two SDKs, a documented error and rate-limit policy, and results on two independent boards. Microsoft announced its model on 9 October 2026 and has not published its Foundry request format. Its accuracy and calibration numbers come from Microsoft’s own tests, and those tests give Jev no accuracy score at all. Its headline speed figure is the strongest claim in its favour, but it was measured in a different way from the Jev figure it sits next to.

Jev vs Microsoft-Decision-1 at a glance

Jev Microsoft-Decision-1
Maker TypeSafe AI Microsoft
Released 15 September 2026 Announced 9 October 2026; the model card gives 8 October
Model and base jev-1.13.0, base not published Post-trained from Qwen3.5-9B
Weights Closed, API only Closed, API only
Where to call it TypeSafe API, OpenRouter, Cloudflare Workers AI Microsoft Foundry; OpenRouter as microsoft/microsoft-decision-1, served by Azure
Endpoint POST https://api.typesafe.ai/v1/systemone Foundry format not published as of 10 October 2026
Question types Choice, Score, Noul Yes/no, multiple-choice, rating
Input Text Text only
Context 64,000 tokens per request; state plus the longest question must fit in 32,000 32,768 tokens
Options per Choice 255 Not published
Levels per Score 10 Not published
Rate limit 40 requests and 100,000 tokens per second, as of 30 September 2026 Not published as of 10 October 2026
SDKs Python and TypeScript None named by Microsoft
Input price per million tokens $0.042, output free $0.042, output free
Latency as published 70ms to 500ms end to end, TypeSafe’s claim 85ms median, 125ms p95 through Foundry, Microsoft’s run
Independent boards Decision Index 0.3 and JevBench v1.6.1 Not listed on either as of 10 October 2026

Sources: TypeSafe’s models page, API reference and SDK docs; Microsoft’s launch post and Foundry model card; OpenRouter’s models and endpoints API, all checked 10 October 2026. The latency rows are each vendor’s own figures and do not compare with each other, as the speed section explains.

What each model does

Both models read a block of text, which Jev calls the state, and answer questions about it with calibrated probabilities. Neither writes a reply or explains its answer. Microsoft’s model card also lists grading AI responses and agent actions against a rubric. Jev does the same job through a Score or a Noul with your own criteria.

The two vendors rule out the same work. Microsoft’s card excludes text generation, open questions, conversation, translation and summarization. TypeSafe’s jaggedness page for Jev 1.13 lists nine weak spots, among them arithmetic, date comparison and generation. Microsoft says its model is tuned for English and may do worse in other languages and in medical, legal and financial work. TypeSafe also calls English Jev’s primary language.

API and drop-in compatibility

Jev’s request is documented: one POST to /v1/systemone with a state and named questions. TypeSafe’s Python and TypeScript SDKs wrap it, and the API reference lists the error codes and retry rules. Several other vendors accept the same format, which Jev alternatives lists.

On Foundry, Microsoft has not published Microsoft-Decision-1’s request format, endpoint path, option limits or question limits as of 10 October 2026. The launch post mentions only a simple structured API call. So it is unknown whether a Jev request or TypeSafe’s SDKs work against it on Foundry.

OpenRouter is the exception. It listed the model on 9 October 2026, with Azure as the only provider, at the same price and with a 32,768-token context. There it runs through OpenRouter’s Decisions API at POST https://openrouter.ai/api/alpha/decisions, the same endpoint OpenRouter uses for Jev. OpenRouter’s code sample for it sends a state and Noul, Choice and Score questions, the same fields as a Jev request. If you already call Jev through OpenRouter, trying Microsoft-Decision-1 may be a change of model id. Run your own evals before you rely on that, because OpenRouter’s alpha path means the endpoint is still unfinished.

One more difference matters for repeatable results. Jev has a pinned version, jev-1.13.0, and TypeSafe ties its weakness list to that version. OpenRouter’s listing for Microsoft-Decision-1 says its weights are updated continually while the API shape stays the same. If you need answers to stay stable over time, ask Microsoft how versions are pinned.

What Microsoft’s benchmarks show

Every accuracy and calibration figure for Microsoft-Decision-1 comes from Microsoft’s launch post. Microsoft chose the tasks, ran the models and drew the charts. No one else has reproduced them.

Microsoft’s chart Average accuracy, 36 benchmarks Calibration, 100 is perfect
Microsoft-Decision-1 83.5% 92.2
Quyet-1.0-Large 81.9% 93.1
GPT-6 Luna Decisions 79.4% 89.9
H2O-Lightning-4B v1.1 77.2% 91.8
Jev 1.13.0 Not ranked Not scored

The accuracy figure is a mean over 147,137 questions. Jev appears in the same chart, but only with a latency figure. Microsoft gives it no accuracy or calibration score. So Microsoft’s data gives no basis for saying Microsoft-Decision-1 is more accurate than Jev.

Microsoft also tested whether answers hold when the wording changes. It altered each request in eight ways, and the model changed its answer on 1.3% of the altered requests on average. It never changed its answer when options were paraphrased, reversed or shuffled. TypeSafe publishes no figure like this for Jev, though its jaggedness page shows a Noul and a Choice giving different answers to the same idea. Xbox Research sorted more than 10,000 pieces of feedback with it and found it competitive with GPT-6 Sol at over 14 times the speed. That is also Microsoft’s own report.

What the independent boards show

Neither independent board lists Microsoft-Decision-1 as of 10 October 2026.

The Decision Index is a third-party board with one suite and one set of rules for every entry. Edition 0.3 was generated on 7 October 2026, two days before Microsoft’s launch. Jev has a Full score of 60.11 and shares second place of 115 rows. Its score on public benchmarks alone is 57.96, and its calibration error is 0.074, where lower is better.

Benchmark Heaven’s JevBench v1.6.1 keeps Jev as the reference row on its main board and also scores it on a separate board for hosted APIs. Microsoft-Decision-1 appears on neither board, nor in the newer v1.6.2 release.

Until one of these boards scores it, the only quality numbers for Microsoft-Decision-1 are Microsoft’s. Jev has third-party results on two boards and has been public since 15 September 2026.

Speed, cost and hosting

Microsoft measured a median of 85ms and a 95th percentile of 125ms per request, through Foundry in the same region. Its chart sets that next to 240ms for Jev 1.13.0. The 240ms is not Microsoft’s measurement. It is JevBench v1.6.1’s adjusted median, which Microsoft checked on 7 October 2026. JevBench’s own data gives Jev a median of 239ms and a 95th percentile of 296ms.

These numbers are not like-for-like. Microsoft timed its own model through Foundry, and Benchmark Heaven timed Jev through TypeSafe’s hosted API with its own harness. Different people ran them on different setups. The harness alone moves Jev’s number a lot: the Decision Index measured a 524.1ms median for Jev as a hosted round trip from its lab. TypeSafe claims 70ms to 500ms end to end. Microsoft-Decision-1 may well be faster, but the chart does not show by how much. Time both from your own region on your own requests.

Cost is a tie on paper. Both charge $0.042 per million input tokens with free output, and OpenRouter charges the same for each. Jev’s whole request can hold 64,000 tokens against 32,768 for Microsoft-Decision-1. Jev caps the state plus the longest question at 32,000 of that, so for a single long document the two limits are close.

Hosting is where Microsoft’s case is strongest for some buyers. Foundry sells it as a Direct from Azure model. Microsoft describes these as bought and managed through Azure under one licence, with unified billing and governance, pay-as-you-go or reserved capacity. If your company already buys through Azure, that may be easier than adding TypeSafe as a new vendor. Jev reaches more places today: TypeSafe’s own API, OpenRouter and Cloudflare Workers AI.

When to use which

  • You need a documented API and SDKs today. Use Jev. Its endpoint, limits, errors and Python and TypeScript clients are published.
  • You want independent results before you commit. Use Jev. It is scored on the Decision Index and JevBench. Microsoft-Decision-1 is on neither yet.
  • You send long requests with many questions. Jev’s 64,000-token request is larger, though its state limit is close to Microsoft’s 32,768.
  • You already buy through Azure. Microsoft-Decision-1 fits your existing contract and billing in Foundry.
  • Latency is your main limit. Microsoft’s 85ms median is its own measurement on its own platform. Test both in your region before you choose on speed.
  • You care about answer stability when options are reworded. Microsoft publishes how often its answers flip when a request is reworded, and TypeSafe does not. Test it on your own prompts.
  • You use OpenRouter. Both are listed there at the same price, through the same Decisions API, so you can compare them on your own data cheaply.

FAQ

Is Microsoft-Decision-1 a Jev alternative?

Yes, for the job. Both answer yes/no, multiple-choice and rating questions about a block of text with probabilities, at the same $0.042 per million input tokens. Microsoft has not published its Foundry request format as of 10 October 2026. On OpenRouter both run through the same Decisions API with the same fields, so trying it there may only need a new model id.

Is Microsoft-Decision-1 better than Jev?

No one knows yet. Microsoft’s own tests put it first on accuracy among the models it scored, but they give Jev no accuracy or calibration score. Neither the Decision Index nor JevBench lists Microsoft-Decision-1 as of 10 October 2026. Jev has independent scores on both. Test the two on a sample of your own data.

Is Microsoft-Decision-1 faster than Jev?

Microsoft reports an 85ms median through Foundry against 240ms for Jev. The Jev figure comes from JevBench, a different harness with a different network path, so the gap is not like-for-like. Jev’s measured median also varies by harness, from 239ms on JevBench to 524.1ms on the Decision Index. Time both from your own region.

Can I use TypeSafe’s SDKs with Microsoft-Decision-1?

Unknown as of 10 October 2026. TypeSafe’s SDKs call /v1/systemone, and Microsoft has not published the request format or endpoint path for the model on Foundry. On OpenRouter, it uses the same Decisions API as Jev, with state and questions fields. Test a call before you build on either route.

Does Microsoft-Decision-1 have open weights?

No. Microsoft post-trained it from Alibaba’s open-weight Qwen3.5-9B, but it does not release the result. You can only call it as a hosted API, through Microsoft Foundry or OpenRouter. Jev is also closed and hosted only, so neither model can be run on your own hardware.

Examples

Microsoft-Decision-1: Our model for fast decision-making

Microsoft's launch post for Microsoft-Decision-1, a Qwen3.5-9B decision model on Foundry at $0.042 per million input tokens. It reports Microsoft's own accuracy, calibration and latency charts against Quyet-1.0-Large, GPT-6 Luna Decisions and Jev, plus internal use at Xbox Research and Copilot.

JevBench by Benchmark Heaven: Jev-class decision model leaderboard

Benchmark Heaven's own leaderboard for Jev-class decision models, unrelated to the dhruvmehra/jevbench repo, ranking 106 of 112 systems on 1,624 choice, score and noul decisions each in release v1.5.4. Jev 1.13.0 leads on capability at 80.0, while on the four-axis composite that adds speed and cost, Cygnet and Winnow-12B Q8 tie first and Jev is third.

Article

Related guides