Skip to content
System One

Span-01

Span-01 is Respan's behavior classifier for agent traces. It reads a conversation span and behaviors you define in plain language, and returns the probability that each is present, absent or not observable. It costs $0.02 per million input tokens, and a free Lite tier exists.

Updated

What Span-01 is

Span-01 is Respan’s classification model for agent traces. Like a System One model, it returns probabilities in one forward pass and writes no text. You send a conversation span, meaning the earlier messages plus the turn to judge, and a list of behaviors written in plain language, such as “the user expresses frustration.” Respan says the behaviors are not a fixed list. It trained the model for general classification reasoning with RLAIF, reinforcement learning from AI feedback, then specialised it for behavior detection.

What it returns

For each behavior you get three probabilities: p_present, p_absent and p_not_observable. They sum to about 1. Not observable means the trace cannot answer the question, for example when it ends before the customer replies. Respan says to read that as unknown, not absent.

Every behavior in a request is scored in one forward pass, and output tokens are free. Respan’s own API takes POST /api/v1/scores with span.input, span.output and a behaviors list of ids and definitions. Its docs do not mention the /v1/systemone schema. The three-way output does not match Choice, Score or Noul exactly, and Noul is the closest. A third-party benchmark repo calls it through OpenRouter’s System One endpoint.

What it is good at

Respan aims it at evaluation, guardrails and monitoring of LLM and agent output. Its example scores escalation, user frustration and prompt injection in tool output, then uses thresholds in code to hand off, block or alert. That fits LLM guardrails and confidence-gated actions.

On Respan’s own production behavior benchmark, Span-01 scores 0.806 overall F1, against 0.716 for Jev, 0.719 for Sonnet 5 and 0.885 for GPT-6 Sol. Its best domains are privacy and secrets at 1.000 and agent and tool reliability at 0.845. These are Respan’s numbers on its own dataset.

The one outside test found is not about behaviors. On a 48-question sentiment and topic pool, the zero-shot-ie-bench author measured 85.4% for Span-01 and 79.2% for Lite, against 93.8% for Jev.

Span-01 Lite

Lite is the free, lighter tier of Span-01. On Respan’s overall behavior chart it scores 0.761 F1 against 0.843 for Span-01 and 0.715 for Jev. It is the default model on Respan’s API.

What it is not for

It does not write text or explain a score, and it reads text only. It is not a general classifier over a fixed option list. Respan publishes no context length or latency, so test long traces yourself.

Access

Respan’s launch post, dated 24 September 2026, says Span-01 is public. Its docs pages still describe an early access waitlist, with 403 errors until your organization is enabled. OpenRouter has listed it since 26 September, from Respan as the only provider.

Specifications

Question typesNoul
Max Choice optionsNot documented
Score levelsNot documented
Questions per callNot documented
Total contextNot documented
State budgetNot documented
Rate limitSpan-01 Lite has a daily cap that resets at 00:00 UTC. Respan does not publish the size of the cap or any rate limit for Span-01.
EndpointPOST https://api.respan.ai/api/v1/scores
SDKs

Respan does not publish a context length, and OpenRouter lists 0. Respan says there is no per-request cap on the number of behaviors, and that Span-01 scores text only, with each message's content as a string.

Versions

  • span-01-pro, 24 Sep 2026, Span-01 on Respan's API. OpenRouter lists it as respan/span-01, snapshot span-01-20260925, at $0.02 per million input tokens. Release notes
  • span-01-free, 24 Sep 2026, Span-01 Lite, the free lighter tier and the default model on Respan's API. OpenRouter lists it as respan/span-01-lite and respan/span-01-lite:free. On Respan's overall behavior benchmark it scores 0.761 F1 against 0.843 for Span-01. Release notes

Use cases

What people use Span-01 for, one page per pattern.

Safety and qualityIntermediate

LLM guardrails with System One models

Put one Jev request in front of an LLM and one behind it. Yes/no questions return the probability that each hazard holds, a Score rates how much harm complying would do, and your thresholds turn those numbers into pass, review, block, or a crisis path.

NoulScore
Workflow controlIntermediate

Confidence-gated actions with System One models

Jev returns a confidence value from 0 to 1 alongside every Choice and Score answer. Your code treats it as a separate axis: act automatically when it's high, confirm or flag when it's middling, hand the decision to a person when it's low. Riskier actions get higher bars.

ChoiceScore