Skip to content
System One

TokenTrim/jev-agent-failure-benchmark

Benchmarks Jev on the 6,257 text traces of the Who and When Pro agent-failure-attribution set, predicting the responsible agent, the decisive step and the error type. Scored with the official whowhen_eval scorer against the paper's LLM baselines, at about $1.28 for the full run.

Open on GitHub

PrimitivesChoice

More like this

Keep browsing