Skip to content
System One

Microsoft-Decision-1 vs pplx-decider: Microsoft's decision model against Perplexity's

Microsoft-Decision-1 launched on 9 October 2026 as a closed, text-only decision model on Microsoft Foundry at $0.042 per million input tokens. Perplexity's pplx-decider-v1.1-27b costs $0.02, reads images, takes 262,144 input tokens and ships Apache 2.0 weights. pplx-decider tops Decision Index 0.3 with a Full score of 62.75. Microsoft-Decision-1 is on no independent board yet, and Microsoft's launch charts do not include any Perplexity model, so no head-to-head result exists.

Updated

The short answer

Microsoft-Decision-1 and pplx-decider do the same job. Each reads some context and a closed question with fixed answers, then returns a probability for each answer instead of writing text. Both cover yes/no, multiple-choice and rating questions, the Noul, Choice and Score types.

On paper, pplx-decider offers more. It costs $0.02 per million input tokens against Microsoft’s $0.042. It reads images, takes 262,144 input tokens against 32,768, and you can download its weights under Apache 2.0. Perplexity also publishes its request format and limits, which Microsoft had not done as of 10 October 2026.

Microsoft-Decision-1 launched on 9 October 2026 and is closed and hosted only. Microsoft’s own tests say it is fast, at a median of 85 ms per request, but no independent board has measured it yet. pplx-decider tops Decision Index 0.3 with a Full score of 62.75. The two have never been tested side by side. Microsoft’s launch charts include no Perplexity model, and the Decision Index does not list Microsoft-Decision-1.

Microsoft-Decision-1 vs pplx-decider at a glance

Microsoft-Decision-1 pplx-decider
Maker Microsoft Perplexity
Model and base Microsoft-Decision-1, post-trained from Qwen3.5-9B, 5B to 15B parameters pplx-decider-v1.1-27b, from Qwen3.8-27B; pplx-decider-v1-27b still accepted
Weights Closed, hosted only Apache 2.0 on Hugging Face, not gated
Where to call it Microsoft Foundry; OpenRouter as microsoft/microsoft-decision-1 Perplexity’s Decisions API; OpenRouter as perplexity/pplx-decider-v1.1-27b; your own GPU
Endpoint and request format Not published as of 10 October 2026 POST https://api.perplexity.ai/v1/decisions
Question types Yes/no, multiple choice, rating Noul, Choice, Score
Input Text only Text, JSON and base64 images
Context 32,768 input tokens 262,144 input tokens, 32 MiB body
Options per Choice Not published 1 to 255
Levels per Score Not published Up to 10
Questions per call Not published 128
Input price per million tokens $0.042, output free $0.02, output free
Rate limit Not published 10 requests per second per organization
SDKs None named None for decisions, plain httpx or fetch
Latency as published Median 85 ms, 95th percentile 125 ms through Foundry, Microsoft’s figures Under 2 seconds for a few hundred tokens, about 23 seconds near the limit, Perplexity’s figures
Independent boards None as of 10 October 2026 Decision Index 0.3, first of 115 entries; JevBench v1.6.1, self-hosted weights
Status Generally available on Foundry, announced 9 October 2026 Generally available since 1 October 2026; v1.1 current since early October

Sources: Microsoft’s launch post and Foundry model card, Perplexity’s quickstart and changelog, and OpenRouter’s model API, all checked 10 October 2026. The Decision Index row is from edition 0.3, generated 7 October 2026.

What each model does

Both models return probabilities over a fixed set of answers and never generate text. See Choice, Score and Noul for how the three question types work.

Perplexity documents its request in full. You send model, a state and up to 128 named questions to https://api.perplexity.ai/v1/decisions, with the key in an Authorization: Bearer header. The state can be a string, an object or an array. The array can hold base64 PNG, JPEG or WebP images, each up to 2,048 tiles of 32 by 32 pixels. A Noul returns the probability of yes, a Choice returns one probability per option, and a Score returns a probability-weighted average level.

Microsoft describes its request only in outline: a context or state, a question and a fixed set of answer options. It has not published the field names, the endpoint path, or how many options, levels or questions a request can hold. The Foundry card says input is text only, with no images, audio or video. It also lists grading AI responses and agent actions against a rubric. Microsoft says the model is tuned for English and may do worse in other languages.

OpenRouter now lists Microsoft-Decision-1, served by Azure at $0.042 per million input tokens with a 32,768-token context. Its listing went up late on 9 October 2026. That gives a second way to call it, but Microsoft’s own request format is still unpublished.

What the benchmarks show

No test covers both models under the same conditions. There are three kinds of numbers, and none of them sets one model against the other.

Microsoft’s launch charts. These are vendor-run. Microsoft reports 83.5% average accuracy across 36 benchmarks and 147,137 questions. It compares that with Quyet-1.0-Large at 81.9%, GPT-6 Luna Decisions at 79.4% and H2O-Lightning-4B at 77.2%. Its calibration score is 92.2 out of 100, behind Quyet’s 93.1. The post’s appendix also names Surogate Rune 26B-A4B, deck-31B and Strands-Decider 2B. No Perplexity model appears anywhere in the post or the charts. Strands-Decider 2B is a separate project, not Perplexity’s Decider. On requests altered in eight ways, Microsoft says the model changes its answer 1.3% of the time on average.

Perplexity’s model cards. These are vendor-run too. The v1 card reports 85.71% across 11 benchmarks, and Perplexity measured its own model through its API. The v1.1 card reports a Decision Index score of 61.56. Microsoft used 36 benchmarks and Perplexity 11, with different scoring, so 83.5% and 85.71% do not compare.

Independent boards. The Decision Index runs every entry through one suite. Edition 0.3, generated 7 October 2026, ranks pplx-decider v1.1 first of 115 entries with a Full score of 62.75. That score mixes public benchmarks with private tests. On the public benchmarks alone it scores 62.25. Its expected calibration error, the average gap between stated confidence and how often it is right, is 0.071. Lower is better.

JevBench v1.6.1 by Benchmark Heaven added self-hosted pplx-decider v1.1 weights on 7 October 2026. It gives a Capability Score of 82.8 and a median of 0.39 seconds on its own setup. It places the model outside its Jev-class list because the estimated self-hosting cost, $0.22 per 1,000 decisions, is 6.7 times its Jev reference cost. That figure is JevBench’s estimate for its own hardware, not Perplexity’s API price.

Microsoft-Decision-1 is on neither board as of 10 October 2026. Until someone runs both models through the same suite, there is no way to say which is more accurate.

Speed, cost and hosting

Microsoft measured a median of 85 ms and a 95th percentile of 125 ms per request through Foundry. Its latency chart sets those against JevBench medians, which come from a different harness. Perplexity publishes no median. Its own tests from 30 September 2026 give under 2 seconds for a few hundred input tokens and about 23 seconds near the input limit. The two vendors measured different things in different ways, so the figures cannot be compared. Time both from your own region with your own inputs.

Perplexity is cheaper on price per token. Microsoft’s $0.042 per million input tokens is 2.1 times Perplexity’s $0.02, and both give output away free. A billion input tokens a month costs $42 on Microsoft-Decision-1 and $20 on pplx-decider. OpenRouter lists both at the same prices as their makers.

Limits favour Perplexity as well, mostly because Microsoft has published few. Perplexity caps requests at 10 per second per organization and states the 128-question, 255-option and 10-level limits. Microsoft gives no rate limit and no option or question limits.

Only pplx-decider can run on your own hardware. Its card asks for Python 3.12 or newer and a CUDA GPU with about 49 GiB for the weights plus working memory. Microsoft does not distribute its weights.

How both compare with Jev

Jev costs $0.042 per million input tokens, the same as Microsoft-Decision-1 and 2.1 times pplx-decider’s price. It reads text only, like Microsoft’s model, and accepts 64,000 tokens per request.

Jev has independent results that Microsoft-Decision-1 lacks. On Decision Index 0.3 it has a Full score of 60.11, behind pplx-decider’s 62.75. Microsoft’s latency chart includes Jev 1.13.0 at a 240 ms JevBench median, but Jev gets no accuracy or calibration score in Microsoft’s charts.

Format matters if you already use Jev. pplx-decider has its own /v1/decisions format, so TypeSafe’s SDKs do not work with it. Microsoft has not published its format, so it is unknown whether Microsoft-Decision-1 accepts Jev’s /v1/systemone request. The two Jev pages cover each pairing in detail: Jev vs pplx-decider and Jev vs Microsoft-Decision-1.

When to use which

  • You classify images: pplx-decider reads them. Microsoft-Decision-1 is text only.
  • You send long inputs: pplx-decider takes 262,144 input tokens, eight times Microsoft’s 32,768.
  • You want the lowest price per token: pplx-decider, at $0.02 against $0.042.
  • You want weights you control: only pplx-decider has them, under Apache 2.0.
  • You need independent results now: pplx-decider is on the Decision Index and JevBench. Microsoft-Decision-1 is on neither yet.
  • You already run on Microsoft Foundry: Microsoft-Decision-1 is sold there as a Direct from Azure model. Wait for Microsoft’s request format before you write code against it.
  • Latency is your main limit: Microsoft reports an 85 ms median, but that is its own figure. Measure both before you choose.

See Jev alternatives for the other decision models.

FAQ

Which is better, Microsoft-Decision-1 or pplx-decider?

Nobody has tested them side by side. Microsoft’s launch charts leave out every Perplexity model, and no independent board lists Microsoft-Decision-1 as of 10 October 2026. pplx-decider leads Decision Index 0.3 with a Full score of 62.75. Microsoft reports 83.5% accuracy on its own 36-benchmark suite. Run both on a sample of your own data.

Is pplx-decider in Microsoft’s launch benchmarks?

No. Microsoft’s charts compare Microsoft-Decision-1 with Quyet-1.0-Large, GPT-6 Luna Decisions, H2O-Lightning-4B and GPT-6 Sol, with Jev in the latency chart only. The appendix adds Surogate Rune 26B-A4B, deck-31B and Strands-Decider 2B, which is not Perplexity’s model. No head-to-head result between the two exists.

Which is cheaper?

pplx-decider. It costs $0.02 per million input tokens on Perplexity’s API and on OpenRouter. Microsoft-Decision-1 costs $0.042, the same as Jev. Output is free on both. You can also run pplx-decider on your own GPU under Apache 2.0, which Microsoft does not allow for its model.

Can either model read images?

Only pplx-decider. It takes base64 PNG, JPEG or WebP images in the state array, up to 2,048 tiles of 32 by 32 pixels each, and never fetches image URLs. The Microsoft-Decision-1 model card says it is text only and accepts no images, audio or video.

Can I swap one for the other in my code?

Not yet with any certainty. pplx-decider uses Perplexity’s documented /v1/decisions format. Microsoft had not published Microsoft-Decision-1’s request format or endpoint path as of 10 October 2026, so nobody can say how much code a switch needs. Check Microsoft’s docs before you plan a move.

Examples

Microsoft-Decision-1: Our model for fast decision-making

Microsoft's launch post for Microsoft-Decision-1, a Qwen3.5-9B decision model on Foundry at $0.042 per million input tokens. It reports Microsoft's own accuracy, calibration and latency charts against Quyet-1.0-Large, GPT-6 Luna Decisions and Jev, plus internal use at Xbox Research and Copilot.

JevBench by Benchmark Heaven: Jev-class decision model leaderboard

Benchmark Heaven's own leaderboard for Jev-class decision models, unrelated to the dhruvmehra/jevbench repo, ranking 106 of 112 systems on 1,624 choice, score and noul decisions each in release v1.5.4. Jev 1.13.0 leads on capability at 80.0, while on the four-axis composite that adds speed and cost, Cygnet and Winnow-12B Q8 tie first and Jev is third.

Article

Related guides