Skip to content
System One

willkelly/jev-evaluation

An adversarial evaluation of Jev with nine experiments and 28 predictions fixed before any data was collected. One run covered 123,805 requests over 138 minutes for $12.69, with five failures, plus a 13-rule prompting guide.

willkelly/jev-evaluation on GitHub

Stars
0
Forks
1
Language
Python
License
MIT
Last push
22 Sep 2026

Read from GitHub on . Counts change daily.

How it uses Jev

# Endpoint and model
# --------------------------------------------------------------------------

API_URL = os.environ.get("TYPESAFE_API_URL", "https://api.typesafe.ai/v1/systemone")

# The alias we request. The *versioned* string the server returns (e.g.
# "jev-1.13.0") is recorded on every logged call and reproduced in the report
# header; results are only meaningful against a pinned version.
MODEL_ALIAS = os.environ.get("JEV_MODEL", "jev-latest")

# $42 per 10^9 input tokens, output free. Used for cost-per-decision.
USD_PER_INPUT_TOKEN = 42.0 / 1e9
USD_PER_OUTPUT_TOKEN = 0.0

# --------------------------------------------------------------------------

View in jeveval/config.py

Open on GitHub

More like this

Keep browsing