Open source, self-hosted. There is no hosted API. You download the weights and run them on your own hardware. How open-source models are listed.
What CLM is
CLM, short for Contrastive Language Model, is an open-weights System One model from a research group led by Jacky Kwok, whose X bio lists a Stanford CS PhD and Berkeley EECS background, with six co-authors credited in the repo’s citation. It reads a state and a set of candidate actions and scores each, rather than generating text. clm-serve exposes this over a POST /v1/systemone endpoint shaped like TypeSafe’s API, so it accepts Choice, Score and Noul questions the same way Jev does, plus a lower-level /v1/rank endpoint that scores any list of free-form candidates against a state.
How it is built
CLM trains two encoders, one for states and one for actions, with a contrastive objective (InfoNCE) that pulls a state’s embedding toward the action actually taken and away from the others. At inference, a typed question becomes a state plus a set of candidate action texts, and a softmax over the similarity scores is the answer distribution. The served model pairs a frozen Qwen3-8B backbone with a roughly 75MB trainable projection head, so most of the cost is one embedding per fresh text.
The authors describe three training stages: pre-training on about 60 million Nemotron question-answer pairs, mid-training on about 30 million synthetic hard negatives, and post-training on about 1 million agentic trajectories from public agent-trace datasets. The repo also documents scaling-law fits relating test contrastive loss to compute, model size and data size, published in full on the linked Notion blog.
What it’s good at, according to the authors
The authors report CLM-8B performs on par with Jev across computer-use, gaming and tool-calling tasks while running up to 9 times faster, with the largest gains when many candidate actions can be embedded once and reused. After lightweight fine-tuning, they report state-of-the-art verifier results on two agentic coding benchmarks: 87.6% on Terminal-Bench 2.1 and 81.6% on DeepSWE, both evaluated on held-out task sets, where they say Jev does not serve as an effective verifier. These are the authors’ own claims, run on their own benchmark harness, and have not been independently reproduced.
The disaggregated state and action encoders mean a fixed action set’s embeddings can be cached and reused as the state keeps changing, which the README frames as the main latency win for agent loops that revisit the same options.
The Decision Index 0.2.1 by multimodalart, updated 28 September 2026, does not support the parity claim on its suite. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, CLM-v0.1-8B scores 7.40 against Jev’s 57.91. The index covers general decision tasks, not only the agent loops CLM targets.
What it’s not for
CLM’s own state and action encoders make it a two-embedding classifier, not a generative model, and its typed-question support depends on clm-serve translating Choice, Score and Noul requests into that shape rather than TypeSafe having defined the wire format. The repo requires a Linux host with an NVIDIA GPU to serve the reference encoder, so there is no documented CPU or Apple Silicon path.
Access today
Code and the CLM-8B weights are Apache 2.0 and hosted on GitHub and Hugging Face, with no waitlist. Running it means standing up a vllm serve pooling endpoint for the Qwen3-8B encoder plus clm-serve for the API and playground.
Specifications
| Question types | ChoiceScoreNoul |
| Max Choice options | Not documented |
| Score levels | Not documented |
| Questions per call | Not documented |
| Total context | 2,048 tokens |
| State budget | Not documented |
| Rate limit | Not documented |
| Endpoint | POST /v1/systemone on your own server |
| SDKs | Python: clm |
The README's quickstart serves the Qwen3-8B encoder with --max-model-len 2048; the repo does not otherwise document a fixed cap on Choice options, Score levels, or questions per call. Score criteria must be an ordered list of at least 2 levels, matching TypeSafe's Score primitive. State can be a string, an object rendered as key: value text, or an array rendered as list lines, never JSON.
Versions
- Contrastive-LM/CLM-v0.1-8B, 23 Sep 2026, The reference projection head served as clm-latest, about 75MB, trained against a Qwen3-8B backbone with last-token pooling. Apache 2.0 on Hugging Face. Release notes
Use cases
What people use CLM for, one page per pattern.
Support inbox triage with System One models
Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket.
Agent routing and skill selection with System One models
An agent choosing from a long skill roster reads one truncated line per entry and often loads the wrong thing. Jev ranks every entry in one request and separately answers whether any skill applies at all, so the agent gets a short hint instead of a guess.
Examples built with CLM
The most-starred and most-viewed entries in the directory. Browse all examples.

CLM: Contrastive Language Models
Open System One model that scores states against actions with a contrastive objective, serving CLM-8B behind a TypeSafe-compatible API. The authors report Jev-level results on computer-use, gaming and tool-calling with up to 9x lower latency, plus scaling laws and a fine-tuning guide.
215 starsvs Jev latency up to 9x fasterDeepSWE (tuned) 81.6%

Decision Index: Jev against 50+ open decision models
A leaderboard that runs Jev and 54 open System One models and clones through the same 120,000-question suite on one RTX PRO 6000, scoring accuracy and calibration. Jev leads the 0.2 edition at 51.67 on a chance-corrected scale, with AutoJev-27B close behind at 50.94.

JevBench by Benchmark Heaven: Jev-class decision model leaderboard
Benchmark Heaven's own leaderboard for Jev-class decision models, unrelated to the dhruvmehra/jevbench repo, ranking 106 of 112 systems on 1,624 choice, score and noul decisions each in release v1.5.4. Jev 1.13.0 leads on capability at 80.0, while on the four-axis composite that adds speed and cost, Cygnet and Winnow-12B Q8 tie first and Jev is third.