The short answer
Jev has no drop-in replacement. What you pick depends on which part of Jev you need.
If you want someone else to run it, there are eight other hosted options. Tev1 on Together AI and Kev-4B on OpenRouter cost the same $0.042 per million input tokens as Jev. Tev1 only answers Choice questions. Kev-4B answers Choice, Score and Noul and takes the same request as Jev.
The six newest are commercial models. Solar Decide from Upstage answers all three question types on the same /v1/systemone request schema as Jev. It is in beta at $0.10 per million input tokens, or $0.05 on OpenRouter for a limited time. Span-01 from Respan scores agent traces for behaviors you define, at $0.02 per million input tokens, and has a free Lite tier. d1 from Liquid AI answers all three types too, with no published paid price. Decider 1 from meraGPT also answers all three on the same schema, at $0.03 per million input tokens, but a request can hold only 4,096 tokens. Clef from Cloudflare answers all three on the same request schema and also reads images. It runs on Workers AI at $0.24 per million input tokens, or $0.09 for the smaller Clef-flash, and every quality figure for it is Cloudflare’s own. pplx-decider from Perplexity launched the same day and answers Choice, Score and Noul questions about text or images. It costs $0.04 per million input tokens, with a cut planned, but it uses Perplexity’s own /v1/decisions request format instead of /v1/systemone. The remaining hosted route is an ordinary LLM with structured outputs, which is slower and costs more per decision.
If you want weights you can run and keep, the serious candidates are Laya, Kev, Bespoke Nimble, CLM, Von and Fastino Labs’ GLiNER2.5-Decide, released 24 September 2026. Each one ships trained weights under Apache 2.0, as do Cloudflare’s Clef and Clef-flash, which are also hosted on Workers AI, and Perplexity’s pplx-decider, which is also hosted on Perplexity’s API. Behind them sit community clones such as SemIf, OpenJev and NanoJev. Those read decisions off open models, and most are research code a week old.
If your labels are fixed and you have training data, skip the whole category. A fine-tuned encoder classifier beat every open System One option in the one public benchmark that included one.
What “alternative” means here
People searching for a Jev alternative usually want one of three things:
- The same API somewhere else: you have code that calls
POST /v1/systemoneand you want it off TypeSafe. Many open projects speak that wire format, and so do the hosted Solar Decide, d1 and Decider 1, so the change is one base URL. Clef takes the same fields but at a Workers AI URL, with an extramodelfield. Perplexity’s pplx-decider does not: it uses its own/v1/decisionsformat. - A model you own: weights on your disk, no text leaving your network, no rate limit. That is the “jev local” and “jev open source” question.
- The same job done another way: a classification model, an NLI model or an LLM with a JSON schema. None of these is a System One model, and all of them can sort text into labels.
No alternative copies Jev’s training. TypeSafe has not published RLCD, its method for training calibrated probabilities, so every project here copies the interface and not the method. RLCD explained covers what is known.
Jev alternatives compared
Stars were read from the GitHub API on 30 September 2026 and change daily. “Own” in the evidence column means the authors measured it themselves.
| Option | How you use it | Base model and size | Licence | Stars | Eval evidence |
|---|---|---|---|---|---|
| Jev 1.13 | TypeSafe API, $0.042 per million input tokens | Not published | Closed | n/a | TypeSafe’s own figures; outside tests below |
| Tev1-4B-experimental | Together AI serverless API, $0.042 per million input tokens | Qwen3.5-4B, LoRA fine-tune | Code MIT; the weights card declares no licence | 173 | 880 of 1,000 on a reused dev set (own) |
| Solar Decide | Upstage API in beta at $0.10 per million input tokens, or OpenRouter at $0.05 for a limited time | Solar Mini 4, 35B mixture-of-experts with 3B active | No downloadable weights | n/a | 95.8% on a 48-question pool against Jev’s 93.8%, from a third-party pull request; no Upstage figures |
| Span-01 | Respan API at /api/v1/scores or OpenRouter, $0.02 per million input tokens; Span-01 Lite is free with a daily cap |
Not listed | Not listed | n/a | 0.806 F1 against Jev’s 0.716 on Respan’s own behavior benchmark (own); 85.4% against Jev’s 93.8% on the third-party 48-question pool |
| d1 | Liquid AI API as d1:free, no paid price published |
Not published | No downloadable weights | n/a | 58.9 against Jev’s 57.9 on Liquid’s own reproduction of Decision Index 0.2.1 (own); not on the public index |
| Decider 1 | meraGPT API at $0.03 per million input tokens, up to 4,096 tokens per request | Not published | No downloadable weights | n/a | 0.768 accuracy against Jev’s 0.727 on typed-decisions, a benchmark meraGPT built (own); no outside test |
| Clef | Cloudflare Workers AI at $0.24 per million input tokens (Clef) or $0.09 (Clef-flash); weights also self-hostable | Qwen3.8-27B (Clef); Qwen3.5-9B (Clef-flash) | Apache 2.0 | n/a | Above Jev on 8 of 10 Decision Index tasks and 3 of 4 TypeSafe workflow evals, below it on the rest (own, Cloudflare’s run) |
| pplx-decider | Perplexity Decisions API at $0.04 per million input tokens, with a cut planned; weights also self-hostable | Qwen3.8-27B | Apache 2.0 | n/a | 85.71% against Jev’s 84.51% across 11 benchmarks, measured by Perplexity (own); 56.40 against Jev’s 57.91 on the Decision Index |
| Laya | Self-hosted, GPU or CPU | ModernBERT-large, 421M; mmBERT-base, 322M | Apache 2.0 | 28,759 | TREC 86.6%, BANKING77 37.0% (SOTAAZ); Berkeley exam 31.2% |
| Kev | Self-hosted, CUDA, ROCm or Apple MLX; Kev-4B also on OpenRouter at $0.042 per million input tokens | Qwen3.5 at 0.8B, 4B and 9B; Qwen3.8-27B for the 27B | Apache 2.0 | 7,948 | Kev-27B 0.848 and Kev-9B 0.822 against Jev’s 0.857 on unseen datasets (own) |
| GLiNER2.5-Decide | Self-hosted, CPU or GPU; Fastino also hosts inference at agent.fastino.ai | DeBERTa-v3-large fine-tune, 340M per the model card (Hugging Face’s own weight metadata reports 486M) | Apache 2.0 | 2,237 | 60.1% on Fastino’s own 17-dataset suite, ahead of JevK5 at 57.5% |
| Bespoke Nimble | Self-hosted, Apple Silicon or NVIDIA | Qwen3.5-9B, LoRA fine-tune | Apache 2.0 on the model card | 1,915 | 90.12% against Jev’s 93.21% on 324 held-out examples (own) |
| CLM | Self-hosted, Linux and NVIDIA | Qwen3-8B plus a 75MB head | Apache 2.0 | 2,540 | Claims parity with Jev at up to 9x the speed (own, unreproduced) |
| Von | Self-hosted, GPU or CPU | ModernBERT, 395M | Apache 2.0 | 776 | Von 1.1 72.0% against Jev’s 96.6% on a 49-task suite (own) |
| Decider | Self-hosted, GPU | Qwen3.5-2B-Base, fine-tuned | Apache 2.0 | 979 | None published |
| SemIf (formerly openjev) | Self-hosted, RTX 3090, Mac or CPU | Frozen Qwen3.5-4B | MIT code | 4,583 | Berkeley exam 61.6%; 0.845 agreement on a TypeSafe subset where Jev gets 0.883 (own) |
| OpenJev (razorback16) | Self-hosted, or free on Codiv | Frozen DiffusionGemma 26B-A4B | Apache 2.0 | 540 | BANKING77 66.9%, TREC 66.0%, AG News 85.0% (SOTAAZ) |
| NanoJev | Self-hosted, CUDA | Qwen3-0.6B | MIT | 2,440 | TREC 18.4%, BANKING77 20.1% (SOTAAZ); trained on games |
| jevlike | Self-hosted, CPU or GPU | Own small byte-level model | MIT | 1,336 | None published |
| Rizzo Flow | Self-hosted through llama.cpp | Frozen Spark-X2.5, 4B or 1.7B | Apache 2.0 | 758 | Tied with SemIf on SemIf’s own test set (own) |
Sources: each project’s README or model card, the model pages on this site, and the SOTAAZ and Berkeley exam write-ups linked below, all checked 24 September 2026, except that star counts, licences and the Von and Kev rows were rechecked on 30 September 2026. The Solar Decide, Span-01, d1 and Decider 1 rows come from their model pages, checked 30 September 2026, and the Clef and pplx-decider rows from their model pages, checked 2 October 2026.
The Decision Index
The broadest independent test so far is the Decision Index 0.2.1 by multimodalart, updated 28 September 2026. It sends the same 119,898 scored questions to Jev and 70 open models and clones, runs every open one on a single RTX PRO 6000, and publishes its code. The headline is a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas.
| Entry | Index score | Calibration error |
|---|---|---|
| Jev 1.13 | 57.91 | 0.074 |
| Surogate Rune 26B-A4B v3 | 57.44 | 0.120 |
| Decider chat with Gemma-4-31B | 57.33 | 0.047 |
| pplx-decider-v1-27b | 56.40 | 0.018 |
| Decider 35B-A3B | 47.11 | 0.023 |
| Bespoke Nimble 9B v2 | 39.57 | 0.024 |
| Kev-9B | 38.48 | 0.138 |
| Tev1-4B-experimental | 29.24 | 0.104 |
| GLiNER2.5-Decide | 11.21 | 0.088 |
| CLM-v0.1-8B | 7.40 | 0.323 |
| Laya | 6.04 | 0.140 |
Calibration error is the average gap between a model’s stated confidence and how often it is right. d1 is not in this table. Liquid’s own reproduction of Decision Index 0.2.1 gives d1 58.9 and Jev 57.9, and the public data, last generated on 28 September 2026, had no d1 entry on 30 September. meraGPT’s Decider 1 is not in it either, and neither is Cloudflare’s Clef, which does not appear on the public index page as of 2 October 2026. Perplexity’s pplx-decider is in it, listed under the engine id autojev-27b and earlier shown here as AutoJev-27B; the two Decider rows are the unrelated open Decider. Jev leads, but three open entries come within 1.5 points: Surogate Rune 26B-A4B v3, a full fine-tune of Gemma-4-26B-A4B-it; Decider chat with Gemma-4-31B, which runs Decider on a stock Gemma model; and Perplexity’s pplx-decider-v1-27b, a full fine-tune of Qwen3.8-27B. The small encoders that do well on narrow label sets score near random on this broad suite. It is one person’s suite, and it changed between editions, most recently on 28 September when 0.2.1 reweighted the areas and swapped in newer entrants, so check the current page.
Hosted options
Tev1 on Together AI
Tev1 is a System One style model you can call as a paid service on Together AI. Together fine-tuned Qwen3.5-4B to read a state, a question and 2 to 24 options, then return one option letter. It costs $0.042 per million input tokens with free output, the same as Jev.
It is a narrower product. It has no Score and no Noul. The model card says it “retains Qwen’s standard next-token language-model head,” so it is a normal language model trained to answer with one letter, not a model that skips text generation. The repo says its logprobs are not calibrated confidence. Its one published result, 880 of 1,000, comes from a development set the team reused while building the recipe. Together also published the recipe and says a training run costs about $17. That makes Tev1 as much a template for training your own classifier as a product.
Solar Decide on Upstage
Solar Decide is Upstage’s System One model, in beta since 22 September 2026. It runs on Solar Mini 4, a 35B mixture-of-experts model with 3B parameters active per token, and reads a state of up to 512K tokens. It answers Choice, Score and Noul questions and uses the same /v1/systemone request schema as Jev. Upstage lists $0.10 per million input tokens with free output. OpenRouter added it on 28 September at $0.05, a discount its launch post calls limited-time.
A Choice takes 2 to 26 options, and Upstage warns that the schema may change before general availability. Upstage publishes no accuracy figures. The only measurements are third-party. In a pull request to zero-shot-ie-bench, its author reports 95.8% on a 48-question sentiment and topic pool against 93.8% for Jev. That is a small sample from one person’s test set.
Span-01 on Respan
Span-01 does a different job from Jev. You send a conversation span and behaviors written in plain language, such as “the user expresses frustration,” and it returns the probability that each is present, absent or not observable. Noul is the closest System One type, but the output does not match it exactly. Respan’s API takes POST /api/v1/scores, and its docs do not mention /v1/systemone, so do not assume a Jev client works against it.
It costs $0.02 per million input tokens with free output, on Respan’s API and on OpenRouter. Span-01 Lite is free with a daily cap, and it is the default model on Respan’s API. On Respan’s own behavior benchmark Span-01 scores 0.806 F1 against 0.716 for Jev. That is Respan’s number on its own dataset. The one outside test, on a 48-question sentiment and topic pool, has Span-01 at 85.4% and Lite at 79.2% against 93.8% for Jev. Respan’s launch post of 24 September 2026 says Span-01 is public, but its docs pages still describe a waitlist.
d1 from Liquid AI
d1 is Liquid AI’s first decision model, announced on 29 September 2026. It answers Choice, Score and Noul questions and writes no text. Liquid’s docs call POST https://api.liquid.ai/decisions/v1/systemone with TypeSafe’s Python and TypeScript SDKs, using the model id d1:free. Liquid publishes no paid price, and it has not published the base model or the parameter count.
Its launch post says d1 beats Jev on multilingual evals, prompt injection resistance and longer inputs, and gives no numbers for any of the three. The one chart is Liquid’s own reproduction of the Decision Index, where d1 leads Jev overall but trails it on Knowledge, 43.3 against 51.3. Nobody outside Liquid has reproduced it. d1 was not on OpenRouter on 30 September 2026. Vercel’s AI Gateway lists it as liquid/d1.
Decider 1 from meraGPT
Decider 1 is meraGPT’s decision model, announced on 22 September 2026. It answers Choice, Score and Noul questions and writes no text. It serves the same /v1/systemone request schema as Jev at https://meragpt.com, so TypeSafe’s SDKs work once you change the base URL. It costs $0.03 per million input tokens with free output. A Choice takes up to 10 labels, a request up to 64 questions, and the state and questions together must fit in 4,096 tokens, so long documents are out. It is a hosted API only, with no weights. It is not related to the open Decider covered below.
Every accuracy figure comes from meraGPT’s own run on typed-decisions, a synthetic benchmark that meraGPT told this site it built and that it published under the LocalLLaMA organisation on Hugging Face. There Decider 1 scores 0.768 against Jev’s 0.727. The benchmark’s own card says scores well above 0.735 mean a model is learning the quirks of the teacher model that wrote the answers. Nobody outside meraGPT had measured Decider 1 by 30 September 2026, and it is not on Benchmark Heaven’s JevBench.
Clef from Cloudflare
Clef is Cloudflare’s own decision model, launched on 1 October 2026 in two sizes: Clef at 27B on Qwen3.8-27B and Clef-flash at 9B on Qwen3.5-9B. Both run on Workers AI and take Jev’s state and questions request with Noul, Choice and Score questions. Clef adds an optional images array and a 65,536-token context. Clef costs $0.24 per million input tokens and Clef-flash $0.09. The weights are on Hugging Face under Apache 2.0.
Every quality number is Cloudflare’s own run. On 10 tasks it chose from the Decision Index, Clef beats Jev on eight and loses on When2Call (72.37 against 80.97) and BRIGHT (45.91 against 47.52). Clef-flash scores 66.77 on CLINC150+OOS against Jev’s 89.27. On TypeSafe’s workflow evals Clef wins three of four and loses agent trace observability, 68.5 to 71.6. Cloudflare reports a median latency of 209.3ms for Clef and 38.8ms for Clef-flash, against 524.1ms for Jev. It calls its second training stage RLCD, which it describes itself and does not say matches TypeSafe’s method. No outside run of Clef had appeared by 2 October 2026.
pplx-decider from Perplexity
pplx-decider, full name pplx-decider-v1-27b, is Perplexity’s decision model. It launched on 1 October 2026, the same day as Clef, on a Qwen3.8-27B base with Apache 2.0 weights and image input. Both are on Hugging Face. That is where the overlap ends. The Decisions API takes up to 128 named questions per call at POST /v1/decisions, not Jev’s /v1/systemone, and answers Noul, Choice and Score questions. A request can hold 262,144 input tokens, four times Clef’s limit, and the rate limit is 10 requests per second per organization.
It costs $0.04 per million input tokens as of 2 October 2026, with output free, and Perplexity says it plans to lower the price. The model card reports 85.71% across 11 benchmarks against 84.51% for Jev, with Jev ahead on 6 of the 11. Perplexity ran these itself and does not say how it ran Jev, so they are vendor-run. On the independent Decision Index, checked 2 October 2026, it scores 56.40 against Jev’s 57.91, with a calibration error of 0.018 against Jev’s 0.074. Its Python and TypeScript SDKs have no decisions method, so you call the endpoint with httpx or fetch.
OpenAI Decisions API
The OpenAI Decisions API runs on GPT-6 Luna and was announced on 29 September 2026. You define questions with fixed answers and send text or images, and it returns a selection. It is in limited preview for selected API customers, with no published endpoint, price or request format as of 30 September 2026. Every, which had preview access, found it ahead of Jev on some early tests and behind on its broader evals.
An LLM with structured outputs
The oldest alternative is the one most teams already have. Ask GPT, Claude or an open LLM for JSON that matches a schema, and let the provider enforce the schema. You get any label set you like and no new vendor. You also get LLM latency and LLM prices. There is no probability to threshold on, only token log-probabilities where the provider exposes them.
SOTAAZ ran one LLM on BANKING77, a set of 77 banking intents. GPT-5.6 Terra got 83.8% at a median of 1,351ms per message. That was better than every open System One option in the same test, and about 50 times slower than Laya. Jev vs GPT and Jev vs Claude work through cost per decision. Libraries like Outlines and Instructor give you the same guarantee on models you host yourself.
Open-source System One models
These ship trained weights you download and run. Kev-4B is the one exception with a paid hosted API, on OpenRouter since 25 September 2026.
Laya
Laya, from Convai Innovations, is an encoder with a decision head: ModernBERT-large for English at 421M parameters, and mmBERT-base for other languages at 322M. It answers Choice, Score and Noul. Convai measures 32.8ms to 39.5ms for one question on a Tesla T4 and gives 193ms to 464ms on CPU. It is the most-starred project in the space by a long way.
Accuracy depends heavily on the task. Convai’s own README calls Laya “a fast base to specialise, not a zero-shot decision engine.” SOTAAZ found it strong with few labels: 86.6% on TREC’s six question types in 23ms, and 94.5% on AG News. It fell to 37.0% on BANKING77 because all 77 labels share one token budget. Giving it 20 candidates shortlisted by a small embedding model raised that to 59.1%. On the Berkeley exam test it scored 31.2%, against about 30.1% for random guessing. Its calibration error is 0.466 as shipped and 0.081 after fitting a temperature, so fit one on your own data.
Kev
Kev is Jared Palmer’s family of LoRA adapters on Qwen3.5 at 0.8B, 4B and 9B, plus a 27B on Qwen3.8. It serves TypeSafe’s /v1/systemone shape, so the official typesafe-sdk Python client works against it. It runs on CUDA, ROCm or Apple Silicon through MLX. The author reports Kev-4B at 41.5ms to 145ms of model time on an L40S for a new state, and Kev-27B at 46.5ms to 178ms on a B200.
Kev publishes a direct comparison with Jev and says it is not a controlled one. On datasets Kev was not trained on, Kev-9B scores 0.822 on the development set against Jev’s 0.857, and Kev-27B scores 0.848. Kev-27B needs an 80 GB GPU and starts from Qwen’s post-trained release, whose training data is unknown. On knowledge questions the gap is wider: 0.74 on MMLU for Kev-9B and 0.84 for Kev-27B, against Jev’s 0.90. Training covered at most 384 state tokens, so long documents fall outside what it has seen.
Bespoke Nimble
Bespoke Labs’ Nimble fine-tunes Qwen3.5-9B with LoRA and answers choice and true-or-false fields. The team published its 2,676 training examples and 324 held-out examples. On those 324, Nimble matched 90.12% of reference labels, Jev 1.13 matched 93.21%, and untuned Qwen3.5-9B matched 66.36%. The README flags the limits: the labels are synthetic and come from six source families. The 9B weights need about 18GB before quantisation. Median latency was 106ms on an H100 and 444ms on an M5 Pro Mac, against 246.7ms for Jev through its API in the same table.
CLM
CLM takes a different route. It trains separate encoders for states and actions and scores them against each other, on a frozen Qwen3-8B with a small trained head. Because action embeddings can be cached, it is built for agent loops that see the same options again and again. The authors claim parity with Jev and up to 9x the speed on computer-use, gaming and tool-calling tasks. Nobody has reproduced that yet. It needs Linux and an NVIDIA GPU.
GLiNER2.5-Decide
GLiNER2.5-Decide is Fastino Labs’ entry, released 24 September 2026: a 340M-parameter encoder (Hugging Face’s own weight metadata reports 486M, an unresolved discrepancy) fine-tuned from Fastino’s gliner2-large-v1, a DeBERTa-v3-large base. It evaluates typed questions and rules against a passage in one pass and, beyond a typed answer, can also return spans, relations and structured records in the same call. Fastino’s own benchmark, an internal 17-dataset suite called “Fast Decisions,” gives it 60.1% average accuracy against 57.5% for what Fastino calls JevK5, 56.4% for SemIf and 46.6% for Laya, leading on 9 of 17 datasets. Latency, also from Fastino, is a p50 of 167ms on a 48-vCPU CPU and 38 to 47ms on GPU for short requests. None of these figures are independently reproduced. It does not serve TypeSafe’s /v1/systemone wire format; it loads through its own gliner2 Python package. Jev vs GLiNER covers it against Jev and against the original span-extracting GLiNER models in the same family.
Von and Decider
Von is a 395M ModernBERT decision model that serves the same API and installs as von-sdk from pip or npm. Its own table puts Von 1.1 well behind Jev, at 72.0% on a 49-task suite where Jev scores 96.6%. Decider fine-tunes Qwen3.5-2B-Base, serves TypeSafe’s wire format and claims calibrated probabilities, but publishes no measurement.
Community clones and experiments
Most of the “jev open source alternative” results are here. These projects take an existing open model, frozen, and read the answer off its output scores. They train nothing new, or only a little. They show that the interface is easy to copy, and they show little about quality.
- SemIf, formerly openjev, by TheoLeeCJ. It reads option probabilities from a frozen Qwen3.5-4B on one RTX 3090, with MLX and CPU paths too. On 21 yes-or-no questions about one state, it answered in 1.02s against 5.33s for the same model writing a JSON array. It scored 61.6% on the Berkeley exam test, well above Laya and well below Jev’s 83.7%.
- OpenJev by razorback16 (repo). It turns DiffusionGemma 26B-A4B, a model that fills in all its output tokens at once, into a Jev-compatible server. It was the best open option on BANKING77 in the SOTAAZ test at 66.9%. Codiv hosts it free with 100M input tokens per account, billed as a public experiment, not a paid service. The Gemma team’s djev-run deploys a related DiffusionGemma server to Google Cloud Run, and open-jev by JoshuaSP is an earlier experiment on the same idea.
- NanoJev, a 0.6B replica with weights, dataset and a full training pipeline. It is a good codebase to learn from. Its released checkpoint was trained on game decisions like Snake and mazes. It scored 18.4% on TREC and 20.1% on BANKING77 in the SOTAAZ test, so do not use it for text classification.
- jevlike, a tiny byte-level model trained from scratch to pick among changing options, with Doom and chess demos. It is a research toy, not a classifier you deploy.
- Rizzo Flow (repo) runs frozen Spark-X2.5 models through llama.cpp. That means Metal, CUDA, Vulkan or CPU from one install. It reports about 49ms per decision on an RTX 5060 Ti and says it is tied with SemIf on quality. It claims no lead over either project.
- Laptop and Mac experiments: jevmlx, mini-jev and jev-on-a-laptop. Each one reads options off a small model on Apple Silicon. openjev-sglang is the opposite end: a 35B model on a B200.
The full list, with filters, is in the Jev alternatives directory.
Two corrections to older roundups. Latent.Space’s 19 September list describes Kev as a 0.5B adapter on Qwen2.5. The current family is 0.8B to 27B, on Qwen3.5 and Qwen3.8. The same list says Laya was trained with PPO. Laya’s model card says REINFORCE with a group-mean baseline.
Classifiers that were never System One models
If your label set is stable, these often beat everything above on cost and accuracy.
In the SOTAAZ test, a logistic regression on MiniLM embeddings scored 90.3% on BANKING77 and 90.4% on TREC, at under 7ms on a CPU. No open System One option came within 20 points on BANKING77. Jev itself scored 76.3% on BANKING77 in a separate test by nibzard that SOTAAZ cites. We have not checked that run ourselves. The catch is that the classifier needed labelled training data, and a new label means retraining.
- Fine-tuned encoders: ModernBERT, DeBERTa and BERT. Jev vs BERT covers the tradeoff in detail.
- SetFit: a usable classifier from a few dozen labelled examples.
- Zero-shot NLI models: labels at request time with no training, but one forward pass per label, so cost grows with the label count.
- GLiNER: the original GLiNER pulls entity spans out of text, not decisions about it. Fastino’s separate GLiNER2.5-Decide model, covered above, does answer typed decisions. See Jev vs GLiNER.
How to choose
- You need a hosted API and only pick one option per question. Try Tev1 on Together. Test it on your data first, since its only published number is from a reused dev set.
- You need a hosted API with Score and Noul. Stay on Jev, or try Solar Decide, d1, Decider 1, Clef, pplx-decider or Kev-4B on OpenRouter. An LLM with structured outputs also works if you accept the latency. Solar Decide is in beta, d1 has no published paid price, and Decider 1 takes at most 4,096 tokens per request. Kev-4B’s accuracy figures are its author’s own, d1’s are Liquid’s own reproduction, Decider 1’s come from a benchmark meraGPT built, Clef’s are Cloudflare’s own, pplx-decider’s are Perplexity’s own, and Solar Decide’s come from one third-party test on 48 questions.
- Your code calls Jev and you want another hosted API. Solar Decide, d1 and Decider 1 use the
/v1/systemonerequest shape, Clef takes the same fields through Workers AI, and Kev-4B takes the same request as Jev. Span-01 and pplx-decider do not. - You need to classify images as well as text. Clef takes up to four images per request, and pplx-decider takes images in the state array. Both vendors’ numbers are vendor-run, so test them on your own images.
- You need a long input in one request. pplx-decider accepts up to 262,144 input tokens. Its quickstart gives vendor-run times of about 5 seconds for roughly 90,000 tokens and 23 seconds near the limit.
- You want to score behaviors in agent traces. Respan aims Span-01 at that job. It returns present, absent or not observable for each behavior, and its accuracy figures are Respan’s own.
- Your code already calls Jev and must move to your own hardware. Kev, Decider, Von, CLM and OpenJev all serve the same endpoint, so you change the base URL. Rerun your own evals, because every public test shows a quality drop.
- You want the closest quality to Jev on your own GPU. Kev-27B, Nimble and Kev-9B publish the smallest gaps to Jev, all on their own test sets. Kev-27B needs an 80 GB GPU. Check them on yours.
- You need CPU only or very low latency with a few labels. Laya, or Von. Shortlist candidates first if you have dozens of labels.
- You have labelled data and fixed labels. Train an encoder or a logistic regression on embeddings. It is cheaper, faster and, on the public evidence, more accurate.
- You want to learn how these models work. NanoJev, jevlike and SemIf publish their full pipelines and results.
Measure calibration as well as accuracy. Confidence gating, acting only when the probability clears a threshold, is the main reason to use a System One model at all. Most open projects either disclaim calibration or fit a temperature afterwards, so a threshold tuned on Jev will not transfer.
FAQ
Is there an open-source version of Jev?
No. TypeSafe has not released Jev’s weights, architecture or training method. The open projects copy its API shape on other models. Laya, Kev, Nimble, CLM, Von and GLiNER2.5-Decide ship trained weights under Apache 2.0. SemIf, OpenJev and most other clones read answers off frozen open models. None reproduces RLCD, TypeSafe’s unpublished training method for calibrated probabilities.
Can I run Jev locally?
No. Jev is a hosted API only, with no weight download or self-hosted build. To run a System One model locally you run a different one. Laya and Von run on CPU. Kev, Nimble and SemIf run on Apple Silicon or one consumer GPU. CLM and OpenJev need a larger NVIDIA card, or you can try OpenJev free on Codiv.
What is the best Jev alternative?
It depends on the job. For a hosted API, Solar Decide, d1, Decider 1, Clef, pplx-decider or Kev-4B on OpenRouter, or Tev1 on Together for Choice questions only. For the smallest published gap to Jev on your own GPU, Kev-27B on 80 GB, or Kev-9B or Nimble. For a few labels on CPU, Laya. For fixed labels with training data, a fine-tuned encoder beat every open option on BANKING77.
Are the open alternatives as accurate as Jev?
Some come close, most do not. The broadest independent test, the Decision Index, puts Jev first at 57.91 on a 0 to 100 scale, with three open models within 1.5 points, led by Surogate Rune at 57.44. Kev-9B scores 38.48 and GLiNER2.5-Decide 11.21. On the Berkeley exam set, Jev scored 83.7% and SemIf 61.6%. Fastino’s own benchmark puts GLiNER2.5-Decide ahead of JevK5.
Which alternatives work with the TypeSafe SDK?
Kev, Decider, Von, CLM, OpenJev, openjev-sglang and Rizzo Flow serve TypeSafe’s POST /v1/systemone request and response shape, and so do Solar Decide and meraGPT’s Decider 1, unrelated to the open Decider. Liquid’s docs use the SDKs with d1 at https://api.liquid.ai. Change the SDK’s base URL and existing code runs unchanged. Tev1 and Span-01 do not. Tev1 uses Together’s own API and returns one letter, and Span-01 uses /api/v1/scores.