Skip to content
System One

johnhughes3/LegalForecastBench

Benchmark that tests whether models can forecast federal motion-to-dismiss rulings from the judge's written record, scored with claim-defendant micro-Brier metrics and clustered intervals. Ships a Jev integration under integrations/jev alongside the frontier-model workflows.

Open on GitHub

More like this

Keep browsing