Skip to content
System One

Jev benchmarks: every public evaluation in one place

Jev has been measured by TypeSafe itself and by at least twenty independent builders. The independent runs mostly agree on speed: single-question medians between 154 ms and 422 ms per call, faster than every LLM they were compared with. They disagree on accuracy. Jev wins on topic classification, routing, reranking and agent-failure attribution, loses to Claude Haiku 4.5 on phishing and to OpenAI on structured extraction, and trails models trained on labelled data. None of the tests is peer reviewed, and most are single runs by one author.

Updated

The short answer

TypeSafe publishes one set of numbers for Jev: its own workflow evals and the speed and price multiples in the launch post. Independent builders have published at least twenty more benchmarks since launch on 15 September 2026. Most come with code and raw outputs you can rerun.

Speed is the result the independent tests agree on most. Measured medians for single calls land between 154 ms and 422 ms, and every LLM they were compared with was slower.

Accuracy depends on the task. Jev matched or beat the alternatives on topic classification, intent routing, reranking and agent-failure attribution. It lost clearly to Claude Haiku 4.5 on phishing, to OpenAI structured outputs on medical-note extraction, and to plain classifiers trained on labelled data. It also gets worse outside English.

What the evidence doesn’t give you: a peer-reviewed study, a large test repeated by several independent groups, or any test run on your data. Most numbers below come from one run by one author on jev-1.13.0.

TypeSafe’s own numbers

These are vendor claims. They come from TypeSafe’s own harness and have not been reproduced.

The launch post gives 70ms to 500ms end to end for Jev, against 3 to 329 seconds for frontier models, and “40x-200x faster” for System One shaped queries. The post says the headline “193.6x faster, 444.6x cheaper” figures come from its workflow evals, and that TypeSafe expects “these are on the higher end of real world gains.” The speed and pricing guide covers what those ranges leave out.

The workflow evals site runs four tasks: security incidents, agent-trace observability, invoice processing and customer service. Every model answers the same narrow questions, and code turns the answers into an action. Averaged across the four, read on 24 September 2026:

Model, workflow mode Accuracy Cost per case Seconds per case
Jev 67.8% $0.0004 0.4
terra 67.9% $0.0304 10.1
sonnet 5 67.8% $0.1174 78.1
opus 5 73.1% $0.1761 37.8
sol 74.1% $0.0836 23.3
haiku 4.5 53.6% $0.0195 12.5

Model names are as the site labels them. The reference answers are “an average of the responses of GPT-6 Astra and Claude Fable 5.1, both at high thinking,” so accuracy means agreement with two other models, not with ground truth. Per task the picture moves: Jev scored 61.8% on invoice processing against 78.4% for Opus 5, and 76.0% on customer service against 72.4%.

Independent benchmarks at a glance

Each number below was read at the linked source on 24 September 2026, or appears with a quote and link in the claims field of the directory entry.

Classification and extraction

Benchmark Who ran it Task Compared with Headline result Sample and method
OpenRouter Ori Eval OpenRouter Label requests as one of 30 task types GPT-5.6 Luna, DeepSeek V4.1 Flash, Qwen3.8 Flash, GLM 5.3 Flash Jev median 154 ms against 860 ms for the next fastest. OpenRouter says accuracy was within a handful of cases across all five 200 synthetic cases, sequential. Posted on X, raw data not published
Jev vs Claude Ben Greenberg, dev.to Classify evidence as satisfied, not satisfied or insufficient Claude Sonnet 5, high reasoning Jev 100.0% at 378ms median. Sonnet 5 99.0% at 3,554ms median 102 submissions run three times, 306 decisions. One bounded task
jevbench dhruvmehra SST-2, AG News, Banking77 Claude Sonnet 5, GPT-5 mini, Laya, fine-tuned DistilBERT, BART-MNLI Jev 95.4% on SST-2 (Sonnet 5 95.6%), 84.3% on AG News (Sonnet 5 89.6%), 76.4% on Banking77 (Sonnet 5 77.4%, DistilBERT 88.0%). Jev p50 376ms to 389ms 500 examples per dataset. Author notes about ±2.5 points of noise at that size
sysone-bench instax-dutta Nine suites, same bytes to both models Laya, Qwen2.5-1.5B Jev best or tied on six of nine suites, including moderation 0.989 against Laya 0.833. Laya best on AG News (0.940 vs 0.910) and MNLI (0.983 vs 0.867) 751 states. Several suites are curated by the author, not public
jev-vs-open-decision-models elcronos Zero-shot topic and emotion Laya, PrismNLI-0.4B, logistic regression trained on labels Jev 0.793 on tweet_topic against 0.633 and 0.632. On dair-ai/emotion Jev and Laya tie at 0.587. Logistic regression with labels beats every zero-shot system on all four datasets Frozen protocol, 2,000 rows on the main emotion run. Jev’s ECE ranged 0.063 to 0.281 across datasets
Phishing bench anisselbd Should an email agent click the link? Claude Haiku 4.5 Haiku 81.3%, Jev 62.6%. Jev p50 239 ms against 687 ms. $0.038 against $0.462 per 1,000 emails 2,000 emails from PhishNChips v5.2. McNemar test, intervals on every figure. A two-line regex scored 91.6%
Smoking extraction vclic Ten typed fields from outpatient notes OpenAI chat-latest structured outputs All ten fields right: OpenAI 98.7%, Jev 92.4%. Jev $0.18 against $24.54 per 1,000 notes, median 0.375 s against 1.782 s 1,000 synthetic notes. Author calls these pipeline results on synthetic data
Spam eval bitnovus Ham, spam or phishing, zero-shot TF-IDF logistic regression trained on labels Jev 98.64%, regression 98.87% 5,733-email test. Adding link targets and Reply-To raised Jev from 93.62% to 97.98%
Jev vs fine-tuned Laya Alexander Ollman Banking77 plus five public workflows Laya before and after fine-tuning Banking77: Jev 80.0%, fine-tuned Laya 79.4%. Fine-tuned Laya beat Jev on four of five other tasks All 3,080 Banking77 test questions. Laya trained on 1,001 examples
Classifier eval Malte Ubl An existing internal classifier eval Gemini 2.5 Flash Lite “saturated the eval” and ran 6x faster One post on X. Eval not published

Ranking and retrieval

Benchmark Who ran it Task Compared with Headline result Sample and method
jev-rerank-bench anessbelbati Rerank 30 BM25 results Cohere Rerank 4 Pro and Fast, zerank-2, DeepSeek V4.1 Flash nDCG@10: Jev rubric 0.692, Cohere Pro 0.691. Author says no winner. Jev 422 ms against 844 ms 8 English datasets, 1,617 scored questions. Weighting each query equally puts Cohere ahead
Finnish legal rerank laguagu Rerank two Finnish legal corpora Voyage rerank-2.5, GPT-4.1 mini, GPT-5.6 Luna, Laya MuPLeR-fi recall@1: Jev 97.0%, Voyage 95.5%. On the harder private corpus Voyage led, 58.3% against 56.9% 200 and 84 queries. The second corpus is not public
Search rerank eval zhuyansen Rerank an agent-skills catalog bge-m3, text-embedding-3-small, BM25 Jev alone did not beat bge-m3 (+0.012 nDCG@10, interval crosses zero). Jev fused with bge-m3 was best, +0.090 164 queries, 9,831 labelled pairs. Measures and removes judge bias

Security and agents

Benchmark Who ran it Task Compared with Headline result Sample and method
jev-sec-bench Gaurav Gosain Prompt injection, vulnerable code No other model Injection: 96.5% accuracy, ROC-AUC 0.9927, p50 325ms. Vulnerable half ranked above its secure twin in 178 of 200 pairs 662 labelled messages, 200 code pairs. Raw per-sample output committed
Agent failure attribution TokenTrim Who&When Pro: which agent broke a run, and why Paper figures for GPT-5.4, Claude Sonnet 4.6, GLM-5, Qwen3.5-122B Error type F1: Jev 23.7, next best 22.2. Jev also led on agent and step 6,257 traces, official scorer. Jev picked from listed options while the LLMs generated, so only the error-type column is like for like
Qwen on Cerebras vs Jev iammrduncan Seven synthetic app workloads Qwen 3.8 27B on Cerebras p50 176 ms for Jev against 215 ms. Estimated cost $0.011919 against $0.310581 About 480 requests each. Author says it is not a controlled speed ranking

Other tasks

Benchmark Who ran it Task Compared with Headline result Sample and method
Headline A/B bench Gaurav Gosain Pick the winning Upworthy headline Rules of thumb 64.5% on 10,984 tests, 74.7% where the gap was decisive, 53.1% where there was no real gap Design fixed on one split, then run once on the other. Found and removed a position bias
jev-acento marcosmartinez Same tasks in English and Spanish Jev against itself Spanish text cost 3.0 to 6.4 points of accuracy on all four datasets. Spanish instructions did not help 19,200 calls, pre-registered, model version pinned
Adversarial evaluation Will Kelly Nine experiments, 28 predictions fixed in advance Jev against its own claims ECE 0.075 on ticket routing. On random 3-SAT it answered satisfiable for every formula 123,805 requests for $12.69
LLM Chess Maxim Saplin Pick a legal move each turn The LLM Chess leaderboard Rank 59, Elo 242.9±117.5, next to o4-mini-medium at 240.3 80 games at about $0.0015 each

Games and head-to-head demos

These are single matches or short runs. They show speed and behaviour, not general quality.

What the evidence shows

Speed holds up. The single-call medians above are all under half a second, measured from outside TypeSafe’s network, which fits the vendor’s 70ms to 500ms range. Bigger requests take longer: sysone-bench measured 925 ms to 1,068 ms for a five-question call. The multiples vary a lot because they depend on which LLM was chosen and how much it was allowed to reason.

Accuracy is task-shaped. Jev does well when the answer is a fixed label about text it can read directly: topics, intents, relevance, injection attempts. It does worse when the task needs extraction of exact values, counting or multi-step logic. The 3-SAT result and the invoice workflow are examples.

Calibration varies by task too. Reported ECE runs from 0.026 on SST-2 to 0.281 on an emotion dataset. The headline-bench and sec-bench authors found that higher confidence meant more correct answers. The phishing bench found Haiku better calibrated than Jev. Expected calibration error (ECE) is the average gap between stated probability and actual hit rate; the glossary entry explains how it is computed.

Labels still win. Where authors trained a plain classifier on a few hundred to a few thousand labelled examples, it usually matched or beat zero-shot Jev. Jev’s advantage is that you can start without labels.

Context and wording move the score. Telling Jev what the assistant was for raised injection accuracy from 89.7% to 96.5%. Adding link targets raised spam accuracy by more than four points. Naming each subject instead of numbering it took one of Will Kelly’s tests from 0.420 to 1.000.

Jev vs Opus

Two public data points compare the two directly. On TypeSafe’s workflow evals, Opus 5 averaged 73.1% against Jev’s 67.8%, at $0.1761 and 37.8 seconds per case against $0.0004 and 0.4 seconds. In the Pokemon Showdown post, Jev won the one battle played. The Jev vs Claude comparison covers the Sonnet 5 test and when to use each.

Jev 1.13 jaggedness

TypeSafe publishes a jaggedness page for each version: a list of what the model does badly. For jev-1.13 it lists literal interpretation, unreliable counting, weak arithmetic and date ordering, multi-hop indirection, distraction from irrelevant state, and no defence against prompt injection inside the state. It also says a yes/no question and its negation are not guaranteed to sum to 1.

The independent results line up with that list. The 3-SAT failure is multi-hop logic. The “decision already taken” injection that moved 147 of 200 tickets in Will Kelly’s run is injection inside the state. The glossary entry explains the term.

What has not been measured

We found no peer-reviewed study of Jev as of 24 September 2026. Most runs above cover one model version and one run, and several use synthetic or author-curated data.

Two sources are left out of the tables. Near Here’s event-validation test against Mistral and Gemini is listed in the directory, but its page refused our requests, so its figures are not repeated here. A 23 September post on sotaaz.com often turns up in searches for Jev benchmarks. It tests laya, openjev and NanoJev on BANKING77, TREC and AG News, but its author says they had no Jev API access, so it holds no Jev numbers.

Cloudflare’s 1 October 2026 launch post for Clef compares Jev with its own models on 10 Decision Index tasks and four of TypeSafe’s workflow evals. Jev scores higher on When2Call, BRIGHT and agent trace observability, and lower on the rest. Cloudflare ran those numbers itself and is a competitor, so they are not in the tables above. Jev vs Clef lays them out.

How to benchmark Jev on your own data

A public benchmark tells you how Jev did on someone else’s task. To decide whether to use it, you need your own numbers.

  1. Collect a few hundred real inputs with known answers. Several authors above used 500, and the Janus tool’s author suggests starting with a few hundred.
  2. Send the state and questions you would use in production, at production length. Keep the instructions in English.
  3. Run the model you would otherwise deploy on the same inputs. For classification that is often a small LLM or a fine-tuned encoder, not a frontier reasoning model.
  4. Record accuracy, the full latency distribution under your real concurrency, and cost.
  5. Check calibration. Bucket answers by probability and compare each bucket’s average probability to how often it was right. The RLCD guide walks through this under “How to check calibration on your own data.”
  6. Pick a confidence threshold from the results and route everything below it to a person or a larger model. That is confidence gating.
  7. If your options appear in a list, ask each question twice with the order swapped. The headline bench found a position bias of about 7 points.

Tools in the directory can do some of this for you. abhixhek/jevcal fits thresholds to your labelled data and fails CI when a model update breaks them.

FAQ

Is there an official Jev benchmark?

TypeSafe’s workflow evals site is the only one. It covers four tasks and grades against the averaged answers of GPT-6 Astra and Claude Fable 5.1. Averaged across the four tasks, Jev scored 67.8% at $0.0004 and 0.4 seconds per case. Everything else on this page comes from independent authors.

How fast is Jev?

Independent medians range from 154 ms in OpenRouter’s test to 422 ms in the reranking bench, all measured over the network. TypeSafe’s own range is 70ms to 500ms end to end. Latency grows with the size of the state you send, so time it on your own inputs.

Is Jev more accurate than Claude Opus?

Not on TypeSafe’s own evals, where Opus 5 averaged 73.1% against Jev’s 67.8%. Jev took 0.4 seconds per case on those tasks against 37.8, and cost $0.0004 against $0.1761. On single narrow tasks, independent tests have Jev matching or beating Claude models. On phishing, Claude Haiku 4.5 beat it by almost 19 points.

Is JevBench a real benchmark?

jevbench is a GitHub project by dhruvmehra that runs Jev, Laya, two LLMs and two BERT models over SST-2, AG News and Banking77, with 500 examples each. It is reproducible with one OpenRouter key. It is a community project, not a standard maintained by TypeSafe.

Does Jev’s confidence mean anything?

In several tests, yes: higher confidence meant more correct answers, and the headline bench found 78.5% accuracy above 0.80 confidence against 51.7% below 0.60. Calibration still varies by task and language, and the phishing bench found it worse than Haiku’s. Measure it on your data before you set a threshold.

Where does Jev do badly?

Counting, arithmetic, date ordering, multi-step logic and long states full of irrelevant text, per TypeSafe’s jaggedness page. Independent tests add non-English input, exact value extraction and injection text that claims a decision was already made. See the Jev model page for the full profile.

Related guides