Skip to content
System One

abhixhek/jevcal

Fits a per-question confidence threshold to your accuracy target on your own labeled data, checks it on a held-out split, reports how much traffic still needs an LLM, and fails CI when a model update breaks the locked thresholds. Publishes no Jev numbers on purpose.

Open on GitHub

More like this

Use cases this is tagged with