Skip to content
System One

Jev: one judge call, or twelve dimension scores? I measured both on three tasks

Measures one direct Jev call per row against 12 to 14 Jev-scored dimensions fed into a locally trained linear model, across three classification tasks. Decomposition lifted Japanese NLI from 0.837 to 0.908 but was about 25 times worse on false positives against hard benign input, and the whole run cost $1.43 over 5,477 test rows.

Open on Blog

More like this

Use cases this is tagged with