What Kai is
Kai is Hanzo AI’s decision model, and a System One model. Hanzo announced Kai 1 on 28 September 2026. You give it a state and listed questions, and it returns a typed answer with a probability for every option. It writes no text, so usage.output_tokens is 0. Hanzo says the weights are closed and Kai is served by API only.
Per Hanzo’s docs, it is a mmBERT-base encoder initialised from Laya’s multilingual weights, with a cross-attention head, written in Rust. Each option is scored on its own, so listing the options in another order gives the same answer by design. Probabilities are temperature-scaled on held-out calibration rows.
What it returns
The request is POST https://api.hanzo.ai/v1/decisions with model set to kai. The same endpoint also serves Jev. A Choice returns a probability per label, a Score returns a rating against levels you define, and a Noul returns the probability of yes. A request holds 1 to 100 questions, and a question plus its state can take 1,024 tokens at most.
What it is good at
Hanzo lists model and tool selection, deciding when to stop or ask a human, lead scoring, intent detection and escalation, which fit intent and model routing, typed tool dispatch and confidence-gated actions. Input costs $0.021 per million tokens, half of Jev’s price.
Hanzo built and ran its own 12-suite harness (hanzoai/benchmarks, commit 2c74e2d), recorded on 3 October 2026 on CUDA in BF16. Mean accuracy is 0.857 for Kai against 0.780 for Jev 1.13. Kai is ahead on 11 of 12 suites and behind on jailbreak (0.9175 against 0.9400), and its calibration error averages 0.093 against 0.139. On MASSIVE, a 51-language set, it scores 0.9096 against 0.8902. These are vendor-run.
Hanzo states the caveats. Kai trained on these suites’ train splits and is scored on their test splits. Held-out rows have not been run on the served build, and production runs on CPU where the harness has not been run. Hanzo’s other pages give different totals: its Kai page says 12 of 12 suites and 86.2% against 77.2%. This page uses the docs figures. Kai is not listed on Benchmark Heaven’s JevBench or the Decision Index as of 9 October 2026. The Decision Index’s Decision-2.0-Kai-0.6B is a different model, from vllm-sr.
What it is not for
Hanzo says Kai is not yet competitive on dependent multi-question programs, on negated yes/no questions, or on calibration for typed decisions and jailbreak. Wide choices are weak: it scores 0.9075 at 150 options against Jev’s 0.940, and 0.005 at 1,000. Hanzo claims no multimodal accuracy and publishes no latency. The weights cannot be self-hosted.
Access
Kai is live now. Create a key in Hanzo’s console, then call the endpoint. Hanzo’s generated SDKs, among them hanzoai on npm and PyPI, are built from the same API document that lists /v1/decisions.
Specifications
| Question types | ChoiceScoreNoul |
| Max Choice options | Not documented |
| Score levels | Not documented |
| Questions per call | 100 |
| Total context | 1,024 tokens |
| State budget | Not documented |
| Rate limit | A key that sends requests faster than its per-minute rate gets HTTP 429 with a Retry-After header, and so does an organization whose plan usage window is full. The numeric rate is not published as of 9 October 2026. |
| Endpoint | POST https://api.hanzo.ai/v1/decisions |
| SDKs | TypeScript: hanzoaiPython: hanzoai |
A question and its state may take at most 1,024 tokens together, and each option up to 512 tokens. A request holds 1 to 100 questions. The docs give no hard cap on options per Choice. A Choice above 160 options is first shortlisted by a retrieval step that passes 160 options to the scorer. Hanzo's cardinality run measured 0.9675, 0.935, 0.9075 and 0.9075 accuracy at 4, 16, 77 and 150 options, against 0.905, 0.8525, 0.8425 and 0.940 for Jev, and 0.005 at 1,000 options and 0 at 10,000 and 100,000. Jev refuses from 1,000 options. Questions in one request are answered independently.
Versions
- kai-1, 28 Sep 2026, The served checkpoint, requested as model kai. Hanzo's docs describe it as stage a7's weights reading options up to 512 tokens, plus one routed capability that answers questions keyed helpfulness, correctness, coherence, complexity or verbosity. Every answer's routing names the checkpoint, the SHA-256 of its weights and the hash of its calibration tables. The weights SHA-256 begins 0834a74f. Release notes
Use cases
What people use Kai for, one page per pattern.
Intent and model routing with System One models
One Jev call reads an incoming request and returns its intent as a label plus a difficulty rating on a scale you wrote. Your router reads both numbers and picks the handler: deterministic code, a cheap model, an expensive one, or a human queue.
Typed tool dispatch with System One models
Ask Jev one Choice question over your tool names and one per closed-set argument, so every value that comes back is a value the function already accepts. Nouls decide whether an optional argument was mentioned at all. Your code assembles the call and runs it.
Confidence-gated actions with System One models
Jev returns a confidence value from 0 to 1 alongside every Choice and Score answer. Your code treats it as a separate axis: act automatically when it's high, confirm or flag when it's middling, hand the decision to a person when it's low. Riskier actions get higher bars.