Skip to content
System One

CLM: Contrastive Language Models

Open System One model that scores states against actions with a contrastive objective, serving CLM-8B behind a TypeSafe-compatible API. The authors report Jev-level results on computer-use, gaming and tool-calling with up to 9x lower latency, plus scaling laws and a fine-tuning guide.

The CLM playground: a state on the left with three typed questions, their answer distributions on the right
Image from github.com

Contrastive-LM/CLM on GitHub

Stars
215
Forks
13
Language
Python
License
Apache-2.0
Last push
24 Sep 2026

Post on X

Views
208,246
Likes
2,163
Reposts
227

Read from GitHub and X on . Counts change daily.

Reported by the author

vs Jev latency
up to 9x fasterup to 9× faster inference than Jevx.com
DeepSWE (tuned)
81.6%DeepSWE (81.6%)x.com
Terminal-Bench 2.1
87.6%Terminal-Bench 2.1 (87.6%)x.com

CLM (Contrastive Language Models) is an open alternative to Jev rather than a wrapper around it. CLM-8B is trained with a contrastive objective that connects states and actions: its core operation scores a candidate against a state. The repo serves it behind a TypeSafe-compatible API, and questions can be written as Noul, Choice and Score objects or as plain wire-format dicts, so a request written for TypeSafe replays against CLM unchanged.

The authors pre-trained it on 60 million Nemotron question-and-answer pairs, mid-trained on 30 million synthetic hard negatives, and post-trained on 1 million agentic trajectories. They report performance on par with Jev across computer-use, gaming and tool-calling tasks, with up to 9 times lower latency. States and actions are embedded separately, so both can be cached and reused, which the authors say cuts latency most when the state keeps changing while the action set stays fixed. With lightweight fine-tuning, they report new state-of-the-art verifier scores on two agentic coding benchmarks, 81.6% on DeepSWE and 87.6% on Terminal-Bench 2.1, and say Jev does not work as an effective verifier on these longer tasks. They also publish scaling laws showing the contrastive loss falls predictably as a power law in compute, model size and data size, and ship a local playground and a fine-tuning guide. Serving it needs Linux and an NVIDIA GPU.

How it uses Jev

from clm import CLMClient, Choice, Noul, Score

client = CLMClient()                       # CLM_BASE_URL (default http://127.0.0.1:8700), CLM_API_KEY
r = client.system_one(
    state="Customer: my invoice was charged twice and nobody answers the phone!",
    questions={
        "urgency": Noul(instructions="Is this urgent?"),
        "department": Choice(instructions="Which team should handle this?",
                             criteria={"billing": "Charges, invoices, refunds",
                                       "technical": "Bugs and outages"}),
        "frustration": Score(instructions="How frustrated is the customer?",
                             criteria=["Calm", "Frustrated", "Very angry"]),
    },
)
r.answers["urgency"].noul                  # 0.0–1.0

View in src/clm/client.py

Open on GitHub

PrimitivesChoiceScoreNoul

More like this

Keep browsing