Skip to content
System One

zhuyansen/jev-search-rerank-eval

Graded relevance eval of a Jev score rerank against BM25, bge-m3 and other rankers over the Agent Skills Hub catalog, with 164 queries, 9,831 labelled pairs and the judge-circularity bias measured. Jev alone does not beat a good embedding ranker, but fusing the two wins even with Jev removed from the judging.

Open on GitHub

PrimitivesScore

More like this

Use cases this is tagged with