The short answer
TypeSafe publishes one set of numbers for Jev: its own workflow evals and the speed and price multiples in the launch post. Independent builders have published at least twenty more benchmarks since launch on 15 September 2026. Most come with code and raw outputs you can rerun.
Speed is the result the independent tests agree on most. Measured medians for single calls land between 154 ms and 422 ms, and every LLM they were compared with was slower.
Accuracy depends on the task. Jev matched or beat the alternatives on topic classification, intent routing, reranking and agent-failure attribution. It lost clearly to Claude Haiku 4.5 on phishing, to OpenAI structured outputs on medical-note extraction, and to plain classifiers trained on labelled data. It also gets worse outside English.
What the evidence doesn’t give you: a peer-reviewed study, a large test repeated by several independent groups, or any test run on your data. Most numbers below come from one run by one author on jev-1.13.0.
TypeSafe’s own numbers
These are vendor claims. They come from TypeSafe’s own harness and have not been reproduced.
The launch post gives 70ms to 500ms end to end for Jev, against 3 to 329 seconds for frontier models, and “40x-200x faster” for System One shaped queries. The post says the headline “193.6x faster, 444.6x cheaper” figures come from its workflow evals, and that TypeSafe expects “these are on the higher end of real world gains.” The speed and pricing guide covers what those ranges leave out.
The workflow evals site runs four tasks: security incidents, agent-trace observability, invoice processing and customer service. Every model answers the same narrow questions, and code turns the answers into an action. Averaged across the four, read on 24 September 2026:
| Model, workflow mode | Accuracy | Cost per case | Seconds per case |
|---|---|---|---|
| Jev | 67.8% | $0.0004 | 0.4 |
| terra | 67.9% | $0.0304 | 10.1 |
| sonnet 5 | 67.8% | $0.1174 | 78.1 |
| opus 5 | 73.1% | $0.1761 | 37.8 |
| sol | 74.1% | $0.0836 | 23.3 |
| haiku 4.5 | 53.6% | $0.0195 | 12.5 |
Model names are as the site labels them. The reference answers are “an average of the responses of GPT-6 Astra and Claude Fable 5.1, both at high thinking,” so accuracy means agreement with two other models, not with ground truth. Per task the picture moves: Jev scored 61.8% on invoice processing against 78.4% for Opus 5, and 76.0% on customer service against 72.4%.
Independent benchmarks at a glance
Each number below was read at the linked source on 24 September 2026, or appears with a quote and link in the claims field of the directory entry.
Classification and extraction
| Benchmark | Who ran it | Task | Compared with | Headline result | Sample and method |
|---|---|---|---|---|---|
| OpenRouter Ori Eval | OpenRouter | Label requests as one of 30 task types | GPT-5.6 Luna, DeepSeek V4.1 Flash, Qwen3.8 Flash, GLM 5.3 Flash | Jev median 154 ms against 860 ms for the next fastest. OpenRouter says accuracy was within a handful of cases across all five | 200 synthetic cases, sequential. Posted on X, raw data not published |
| Jev vs Claude | Ben Greenberg, dev.to | Classify evidence as satisfied, not satisfied or insufficient | Claude Sonnet 5, high reasoning | Jev 100.0% at 378ms median. Sonnet 5 99.0% at 3,554ms median | 102 submissions run three times, 306 decisions. One bounded task |
| jevbench | dhruvmehra | SST-2, AG News, Banking77 | Claude Sonnet 5, GPT-5 mini, Laya, fine-tuned DistilBERT, BART-MNLI | Jev 95.4% on SST-2 (Sonnet 5 95.6%), 84.3% on AG News (Sonnet 5 89.6%), 76.4% on Banking77 (Sonnet 5 77.4%, DistilBERT 88.0%). Jev p50 376ms to 389ms | 500 examples per dataset. Author notes about ±2.5 points of noise at that size |
| sysone-bench | instax-dutta | Nine suites, same bytes to both models | Laya, Qwen2.5-1.5B | Jev best or tied on six of nine suites, including moderation 0.989 against Laya 0.833. Laya best on AG News (0.940 vs 0.910) and MNLI (0.983 vs 0.867) | 751 states. Several suites are curated by the author, not public |
| jev-vs-open-decision-models | elcronos | Zero-shot topic and emotion | Laya, PrismNLI-0.4B, logistic regression trained on labels | Jev 0.793 on tweet_topic against 0.633 and 0.632. On dair-ai/emotion Jev and Laya tie at 0.587. Logistic regression with labels beats every zero-shot system on all four datasets | Frozen protocol, 2,000 rows on the main emotion run. Jev’s ECE ranged 0.063 to 0.281 across datasets |
| Phishing bench | anisselbd | Should an email agent click the link? | Claude Haiku 4.5 | Haiku 81.3%, Jev 62.6%. Jev p50 239 ms against 687 ms. $0.038 against $0.462 per 1,000 emails | 2,000 emails from PhishNChips v5.2. McNemar test, intervals on every figure. A two-line regex scored 91.6% |
| Smoking extraction | vclic | Ten typed fields from outpatient notes | OpenAI chat-latest structured outputs |
All ten fields right: OpenAI 98.7%, Jev 92.4%. Jev $0.18 against $24.54 per 1,000 notes, median 0.375 s against 1.782 s | 1,000 synthetic notes. Author calls these pipeline results on synthetic data |
| Spam eval | bitnovus | Ham, spam or phishing, zero-shot | TF-IDF logistic regression trained on labels | Jev 98.64%, regression 98.87% | 5,733-email test. Adding link targets and Reply-To raised Jev from 93.62% to 97.98% |
| Jev vs fine-tuned Laya | Alexander Ollman | Banking77 plus five public workflows | Laya before and after fine-tuning | Banking77: Jev 80.0%, fine-tuned Laya 79.4%. Fine-tuned Laya beat Jev on four of five other tasks | All 3,080 Banking77 test questions. Laya trained on 1,001 examples |
| Classifier eval | Malte Ubl | An existing internal classifier eval | Gemini 2.5 Flash Lite | “saturated the eval” and ran 6x faster | One post on X. Eval not published |
Ranking and retrieval
| Benchmark | Who ran it | Task | Compared with | Headline result | Sample and method |
|---|---|---|---|---|---|
| jev-rerank-bench | anessbelbati | Rerank 30 BM25 results | Cohere Rerank 4 Pro and Fast, zerank-2, DeepSeek V4.1 Flash | nDCG@10: Jev rubric 0.692, Cohere Pro 0.691. Author says no winner. Jev 422 ms against 844 ms | 8 English datasets, 1,617 scored questions. Weighting each query equally puts Cohere ahead |
| Finnish legal rerank | laguagu | Rerank two Finnish legal corpora | Voyage rerank-2.5, GPT-4.1 mini, GPT-5.6 Luna, Laya | MuPLeR-fi recall@1: Jev 97.0%, Voyage 95.5%. On the harder private corpus Voyage led, 58.3% against 56.9% | 200 and 84 queries. The second corpus is not public |
| Search rerank eval | zhuyansen | Rerank an agent-skills catalog | bge-m3, text-embedding-3-small, BM25 | Jev alone did not beat bge-m3 (+0.012 nDCG@10, interval crosses zero). Jev fused with bge-m3 was best, +0.090 | 164 queries, 9,831 labelled pairs. Measures and removes judge bias |
Security and agents
| Benchmark | Who ran it | Task | Compared with | Headline result | Sample and method |
|---|---|---|---|---|---|
| jev-sec-bench | Gaurav Gosain | Prompt injection, vulnerable code | No other model | Injection: 96.5% accuracy, ROC-AUC 0.9927, p50 325ms. Vulnerable half ranked above its secure twin in 178 of 200 pairs | 662 labelled messages, 200 code pairs. Raw per-sample output committed |
| Agent failure attribution | TokenTrim | Who&When Pro: which agent broke a run, and why | Paper figures for GPT-5.4, Claude Sonnet 4.6, GLM-5, Qwen3.5-122B | Error type F1: Jev 23.7, next best 22.2. Jev also led on agent and step | 6,257 traces, official scorer. Jev picked from listed options while the LLMs generated, so only the error-type column is like for like |
| Qwen on Cerebras vs Jev | iammrduncan | Seven synthetic app workloads | Qwen 3.8 27B on Cerebras | p50 176 ms for Jev against 215 ms. Estimated cost $0.011919 against $0.310581 | About 480 requests each. Author says it is not a controlled speed ranking |
Other tasks
| Benchmark | Who ran it | Task | Compared with | Headline result | Sample and method |
|---|---|---|---|---|---|
| Headline A/B bench | Gaurav Gosain | Pick the winning Upworthy headline | Rules of thumb | 64.5% on 10,984 tests, 74.7% where the gap was decisive, 53.1% where there was no real gap | Design fixed on one split, then run once on the other. Found and removed a position bias |
| jev-acento | marcosmartinez | Same tasks in English and Spanish | Jev against itself | Spanish text cost 3.0 to 6.4 points of accuracy on all four datasets. Spanish instructions did not help | 19,200 calls, pre-registered, model version pinned |
| Adversarial evaluation | Will Kelly | Nine experiments, 28 predictions fixed in advance | Jev against its own claims | ECE 0.075 on ticket routing. On random 3-SAT it answered satisfiable for every formula | 123,805 requests for $12.69 |
| LLM Chess | Maxim Saplin | Pick a legal move each turn | The LLM Chess leaderboard | Rank 59, Elo 242.9±117.5, next to o4-mini-medium at 240.3 | 80 games at about $0.0015 each |
Games and head-to-head demos
These are single matches or short runs. They show speed and behaviour, not general quality.
- Pokemon Showdown against Opus 5: Sid Arya reported that Jev won one battle, spending $0.0029 and 37s against $2.35 and 6m 29s for Opus 5. One game.
- Snake, local Laya against cloud Jev: 86.5 decisions per second for Laya on a laptop against 3.2 for Jev, whose calls went over the cloud API.
- Tetris, Jev against Laya-MLX: the author reports Laya at about 84ms per decision and Jev winning all three rounds.
- Pong on Show HN (item 49754516): the submitter reported Jev at 227ms average against 3.5s for GPT-5.6 Sol. The Jev vs GPT comparison has the details.
What the evidence shows
Speed holds up. The single-call medians above are all under half a second, measured from outside TypeSafe’s network, which fits the vendor’s 70ms to 500ms range. Bigger requests take longer: sysone-bench measured 925 ms to 1,068 ms for a five-question call. The multiples vary a lot because they depend on which LLM was chosen and how much it was allowed to reason.
Accuracy is task-shaped. Jev does well when the answer is a fixed label about text it can read directly: topics, intents, relevance, injection attempts. It does worse when the task needs extraction of exact values, counting or multi-step logic. The 3-SAT result and the invoice workflow are examples.
Calibration varies by task too. Reported ECE runs from 0.026 on SST-2 to 0.281 on an emotion dataset. The headline-bench and sec-bench authors found that higher confidence meant more correct answers. The phishing bench found Haiku better calibrated than Jev. Expected calibration error (ECE) is the average gap between stated probability and actual hit rate; the glossary entry explains how it is computed.
Labels still win. Where authors trained a plain classifier on a few hundred to a few thousand labelled examples, it usually matched or beat zero-shot Jev. Jev’s advantage is that you can start without labels.
Context and wording move the score. Telling Jev what the assistant was for raised injection accuracy from 89.7% to 96.5%. Adding link targets raised spam accuracy by more than four points. Naming each subject instead of numbering it took one of Will Kelly’s tests from 0.420 to 1.000.
Jev vs Opus
Two public data points compare the two directly. On TypeSafe’s workflow evals, Opus 5 averaged 73.1% against Jev’s 67.8%, at $0.1761 and 37.8 seconds per case against $0.0004 and 0.4 seconds. In the Pokemon Showdown post, Jev won the one battle played. The Jev vs Claude comparison covers the Sonnet 5 test and when to use each.
Jev 1.13 jaggedness
TypeSafe publishes a jaggedness page for each version: a list of what the model does badly. For jev-1.13 it lists literal interpretation, unreliable counting, weak arithmetic and date ordering, multi-hop indirection, distraction from irrelevant state, and no defence against prompt injection inside the state. It also says a yes/no question and its negation are not guaranteed to sum to 1.
The independent results line up with that list. The 3-SAT failure is multi-hop logic. The “decision already taken” injection that moved 147 of 200 tickets in Will Kelly’s run is injection inside the state. The glossary entry explains the term.
What has not been measured
We found no peer-reviewed study of Jev as of 24 September 2026. Most runs above cover one model version and one run, and several use synthetic or author-curated data.
Two sources are left out of the tables. Near Here’s event-validation test against Mistral and Gemini is listed in the directory, but its page refused our requests, so its figures are not repeated here. A 23 September post on sotaaz.com often turns up in searches for Jev benchmarks. It tests laya, openjev and NanoJev on BANKING77, TREC and AG News, but its author says they had no Jev API access, so it holds no Jev numbers.
Cloudflare’s 1 October 2026 launch post for Clef compares Jev with its own models on 10 Decision Index tasks and four of TypeSafe’s workflow evals. Jev scores higher on When2Call, BRIGHT and agent trace observability, and lower on the rest. Cloudflare ran those numbers itself and is a competitor, so they are not in the tables above. Jev vs Clef lays them out.
How to benchmark Jev on your own data
A public benchmark tells you how Jev did on someone else’s task. To decide whether to use it, you need your own numbers.
- Collect a few hundred real inputs with known answers. Several authors above used 500, and the Janus tool’s author suggests starting with a few hundred.
- Send the state and questions you would use in production, at production length. Keep the instructions in English.
- Run the model you would otherwise deploy on the same inputs. For classification that is often a small LLM or a fine-tuned encoder, not a frontier reasoning model.
- Record accuracy, the full latency distribution under your real concurrency, and cost.
- Check calibration. Bucket answers by probability and compare each bucket’s average probability to how often it was right. The RLCD guide walks through this under “How to check calibration on your own data.”
- Pick a confidence threshold from the results and route everything below it to a person or a larger model. That is confidence gating.
- If your options appear in a list, ask each question twice with the order swapped. The headline bench found a position bias of about 7 points.
Tools in the directory can do some of this for you. abhixhek/jevcal fits thresholds to your labelled data and fails CI when a model update breaks them.
FAQ
Is there an official Jev benchmark?
TypeSafe’s workflow evals site is the only one. It covers four tasks and grades against the averaged answers of GPT-6 Astra and Claude Fable 5.1. Averaged across the four tasks, Jev scored 67.8% at $0.0004 and 0.4 seconds per case. Everything else on this page comes from independent authors.
How fast is Jev?
Independent medians range from 154 ms in OpenRouter’s test to 422 ms in the reranking bench, all measured over the network. TypeSafe’s own range is 70ms to 500ms end to end. Latency grows with the size of the state you send, so time it on your own inputs.
Is Jev more accurate than Claude Opus?
Not on TypeSafe’s own evals, where Opus 5 averaged 73.1% against Jev’s 67.8%. Jev took 0.4 seconds per case on those tasks against 37.8, and cost $0.0004 against $0.1761. On single narrow tasks, independent tests have Jev matching or beating Claude models. On phishing, Claude Haiku 4.5 beat it by almost 19 points.
Is JevBench a real benchmark?
jevbench is a GitHub project by dhruvmehra that runs Jev, Laya, two LLMs and two BERT models over SST-2, AG News and Banking77, with 500 examples each. It is reproducible with one OpenRouter key. It is a community project, not a standard maintained by TypeSafe.
Does Jev’s confidence mean anything?
In several tests, yes: higher confidence meant more correct answers, and the headline bench found 78.5% accuracy above 0.80 confidence against 51.7% below 0.60. Calibration still varies by task and language, and the phishing bench found it worse than Haiku’s. Measure it on your data before you set a threshold.
Where does Jev do badly?
Counting, arithmetic, date ordering, multi-step logic and long states full of irrelevant text, per TypeSafe’s jaggedness page. Independent tests add non-English input, exact value extraction and injection text that claims a decision was already made. See the Jev model page for the full profile.