Skip to content
System One

vclic/smoking-extraction-benchmark

Paired comparison of Jev and OpenAI structured outputs on 1,000 synthetic outpatient notes, both given the same note, policy, candidate values and ten typed questions. OpenAI got 98.7% of records fully correct against Jev's 92.4%, while Jev cost about 133 times less with 5.2 times lower mean latency.

Open on GitHub

More like this

Use cases this is tagged with