Skip to content
System One

jev-rerank-bench

Measures whether Jev reranks two Finnish legal corpora better than Voyage, GPT-based rerankers and open-weights Laya over identical shortlists and questions. Batched Jev reached 97 percent recall at 1, matching the shortlist ceiling on the MuPLeR-fi corpus.

Forty near-identical stacks of legal documents on a concrete floor, one lit in amber
Image from github.com

laguagu/jev-rerank-bench on GitHub

Stars
0
Forks
0
Language
TypeScript
License
MIT
Last push
23 Sep 2026

Read from GitHub on . Counts change daily.

How it uses Jev

  const res = await client.systemOne({
    state: { question, passage: passageState(c) },
    questions: { answers_the_question: noul(RELEVANCE_INSTRUCTIONS, RELEVANCE_CRITERIA) },
  });
  scores.set(c.chunkId, res.answers.answers_the_question.noul);
  inputTokens += res.usage.input_tokens;
  outputTokens += res.usage.output_tokens;
} catch {
  failed++;

View in src/rerank/jev.ts

Open on GitHub

PrimitivesNoulScore

More like this

Use cases this is tagged with