Skip to content
System One

Zaious/jev-capability-atlas

An independent, evidence-based map of where Jev's calibrated-decision claim holds up and where it breaks down, built from real API-call receipts rather than a leaderboard.

Zaious/jev-capability-atlas on GitHub

Stars
26
Forks
6
Language
Python
License
MIT
Last push
23 Sep 2026

Read from GitHub on . Counts change daily.

How it uses Jev

def call(client, state, questions, gold_keys):
    t0 = time.perf_counter()
    resp = client.system_one(state=state, model="jev-latest", questions=questions)
    ms = (time.perf_counter() - t0) * 1000
    return {"probs": answer_vectors(resp, questions, gold_keys), "model": resp.model,
            "ms": round(ms, 1), "input_tokens": resp.usage.input_tokens}


def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("--workers", type=int, default=4)
    ap.add_argument("--limit", type=int, default=0)
    a = ap.parse_args()
    client = get_client()

View in suites/laya-head-to-head/run_jev.py

Open on GitHub

More like this

Keep browsing