Skip to content
System One

Jev vs fine-tuned Laya

Compares Jev with Laya before and after task-specific fine-tuning across Banking77 and five other public workflows, plus a chat-moderation study. Fine-tuning raised Laya's Banking77 accuracy from 51.3 to 79.4 percent, close to Jev's 80.0 percent.

Alexander-Ollman/laya-ft on GitHub

Stars
0
Forks
0
Language
Python
Last push
21 Sep 2026

Read from GitHub on . Counts change daily.

How it uses Jev

DATA = ROOT / 'train/moderation-study'
RESULTS = ROOT / 'bench/results/moderation-study'
MODEL = 'typesafe/jev-1.13'
LOCAL = threading.local()


def request(case, rec, split='heldout'):
    if not hasattr(LOCAL, 'target'):
        LOCAL.target = make_target('jev:' + MODEL)
    start = time.perf_counter()
    answers, usage, resolved, error, prediction, probs = {}, {}, None, None, None, None
    try:
        answers, usage, resolved = LOCAL.target.ask(case)
        prediction, probs = parse_answer(answers, ['safe', 'unsafe'])
    except Exception as exc:

View in bench/evaluate_moderation_jev.py

Open on GitHub

More like this

Keep browsing