Skip to content
System One

Tev1 4B-experimental

Tev1-4B-experimental is Together AI's Jev-like classifier, a Qwen3.5-4B fine-tune that reads a state and a set of 2 to 24 options and returns one answer letter. Together released it alongside the training recipe and a tutorial on 23 September 2026.

Updated

What Tev1 is

Tev1 is Together AI’s fine-tune of Qwen3.5-4B into a Jev-style classifier: a System One model that reads a state and a closed set of options and returns one answer, rather than generated prose. Together published it as together/Tev1-4B-experimental on its serverless API and, alongside it, the full data recipe and a tutorial for training your own version.

Unlike Jev’s typed Choice, Score and Noul questions, Tev1 has one output shape: given a state, a question and 2 to 24 labeled options, it returns the letter of the option it picked. The repo’s example script, decide.py, sends a support ticket and four possible intents and gets back a JSON object with the chosen label and key. There is no separate Score or Noul response format documented.

How it was built

Together fine-tuned Qwen/Qwen3.5-4B with LoRA (rank 8, one epoch, learning rate 5e-5, a 2,048-token sequence limit) on 37,840 to 38,340 training examples, depending on whether you read the GitHub README or the tutorial article; the two give slightly different totals for the same dataset. The mix draws from MultiNLI, BoolQ, Banking77, AG News and SST-5, plus generated policy, routing and research-classification examples. Together says the run cost about $17 and took roughly 25 minutes on their fine-tuning service.

The README reports 880 of 1,000 correct on its main development set and 300 of 300 on a policy-transfer set. It is explicit that these are reused development benchmarks the team checked its own recipe against, not a held-out final test, and that endpoint access depends on your own Together account to reproduce.

What it’s good at

Short classification calls with a small, fixed option list: routing a support ticket to an intent, a yes/no/unsure read of a policy question, or picking a sentiment bucket, the same shape as the training datasets. Output tokens are free, so cost scales with the size of the state you send in.

The Decision Index 0.2.1 by multimodalart, updated 28 September 2026, gives an independent number. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, Tev1-4B-experimental scores 29.24 and Tev1-0.8B-experimental 12.85, against Jev’s 57.91.

What it’s not for

Tev1 does not generate text, and it has no documented Score or Noul primitive, so anything needing a calibrated probability or a numeric level rather than a picked letter falls outside what’s published. The GitHub README calls its own benchmark numbers reused development results rather than an independent test, so treat any accuracy comparison with Jev as unverified.

Access today

The fine-tuning repo is MIT licensed and the weights are on Hugging Face under togethercomputer/Tev1-4B-experimental. The repo’s training and deployment scripts use Together’s fine-tuning and dedicated-endpoint services. The hosted together/Tev1-4B-experimental model is live on Together’s serverless API now, billed per the pricing above, with no waitlist mentioned in the launch materials. It is also listed on OpenRouter as togethercomputer/tev1-4b-experimental since 30 September 2026, served by Together at the same price.

Specifications

Question typesChoice
Max Choice options24
Score levelsNot documented
Questions per callNot documented
Total context2,048 tokens
State budgetNot documented
Rate limitNot documented
SDKs

The repo's README says the model takes a state, a question, and 2 to 24 options and returns one answer letter. The saved training recipe used a 2,048-token sequence limit; the maximum number of questions per call and any separate state-token budget are not documented. There is no distinct Score or Noul output: every answer is a selected option letter. OpenRouter accepts a 32,768-token context for the model, well above the 2,048-token training sequence limit, and sets a maximum of 8 completion tokens. Inputs past what training covered fall outside the tested range.

Versions

  • together/Tev1-4B-experimental, 23 Sep 2026, LoRA fine-tune of Qwen/Qwen3.5-4B, rank 8, one epoch, learning rate 5e-5. The repo's saved dev-set results: 880 of 1,000 on the main decision set and 300 of 300 on a policy-transfer set. The README says these are reused development benchmarks, not untouched final tests. Release notes

Use cases

What people use Tev1 for, one page per pattern.

Workflow controlStarter

Support inbox triage with System One models

Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket.

ChoiceScoreNoul

Examples built with Tev1

The most-starred and most-viewed entries in the directory. Browse all examples.