Skip to content
System One

anisselbd/jev-phishing-bench

Reproducible comparison of Jev and Claude Haiku 4.5 on 2,000 phishing emails, asking whether an email agent should click the link. Haiku wins on accuracy, 81.3% against 62.6%, while Jev is faster and cheaper, and a logistic regression over Jev's five signal questions reaches 95.1%.

Open on GitHub

More like this

Use cases this is tagged with