
Contrastive-LM/CLM on GitHub
- Stars
- 215
- Forks
- 13
- Language
- Python
- License
- Apache-2.0
- Last push
- 24 Sep 2026
Post on X
- Views
- 208,246
- Likes
- 2,163
- Reposts
- 227
Read from GitHub and X on . Counts change daily.
CLM (Contrastive Language Models) is an open alternative to Jev rather than a wrapper around it. CLM-8B is trained with a contrastive objective that connects states and actions: its core operation scores a candidate against a state. The repo serves it behind a TypeSafe-compatible API, and questions can be written as Noul, Choice and Score objects or as plain wire-format dicts, so a request written for TypeSafe replays against CLM unchanged.
The authors pre-trained it on 60 million Nemotron question-and-answer pairs, mid-trained on 30 million synthetic hard negatives, and post-trained on 1 million agentic trajectories. They report performance on par with Jev across computer-use, gaming and tool-calling tasks, with up to 9 times lower latency. States and actions are embedded separately, so both can be cached and reused, which the authors say cuts latency most when the state keeps changing while the action set stays fixed. With lightweight fine-tuning, they report new state-of-the-art verifier scores on two agentic coding benchmarks, 81.6% on DeepSWE and 87.6% on Terminal-Bench 2.1, and say Jev does not work as an effective verifier on these longer tasks. They also publish scaling laws showing the contrastive loss falls predictably as a power law in compute, model size and data size, and ship a local playground and a fine-tuning guide. Serving it needs Linux and an NVIDIA GPU.
How it uses Jev
from clm import CLMClient, Choice, Noul, Score
client = CLMClient() # CLM_BASE_URL (default http://127.0.0.1:8700), CLM_API_KEY
r = client.system_one(
state="Customer: my invoice was charged twice and nobody answers the phone!",
questions={
"urgency": Noul(instructions="Is this urgent?"),
"department": Choice(instructions="Which team should handle this?",
criteria={"billing": "Charges, invoices, refunds",
"technical": "Bugs and outages"}),
"frustration": Score(instructions="How frustrated is the customer?",
criteria=["Calm", "Frustrated", "Very angry"]),
},
)
r.answers["urgency"].noul # 0.0–1.0

