Skip to content
System One

Clef-omni

Clef-omni is Cloudflare's multimodal decision model. It reads audio and video as well as text and images, answers the same typed questions as Jev, and runs on Workers AI at $0.15 per million input tokens. It is a 30B mixture-of-experts model with 3B parameters active, and the weights are open under Apache 2.0. Cloudflare launched it on 9 October 2026.

Updated

What Clef-omni is

Clef-omni is the third model in Cloudflare’s Clef family and a System One model. Cloudflare launched it on 9 October 2026, eight days after Clef and Clef-flash. You send a state and typed questions, and it returns a probability for every allowed option, with no generated text. What sets it apart is input: it reads audio (WAV or MP3) and video (MP4 or WebM) as well as text, JSON and images, all in one call. Clef and Clef-flash take text and images only.

It is also built differently. Clef and Clef-flash sit on dense Qwen backbones. Clef-omni is post-trained from Qwen3-Omni-30B-A3B-Instruct, a mixture-of-experts model that uses 3B of its 30B parameters per token. Cloudflare keeps the backbone frozen and trains low-rank adapters and a scoring head on top, with the same loss as Clef: label-smoothed cross-entropy plus a Brier term for calibration. The base model’s speech-output parts are left unused.

What it returns

The request is Jev’s state and questions plus a required model field, and the Hugging Face card calls the API “fully compatible with Jev and SystemOne”. It answers Choice, Score and Noul questions, 1 to 64 per request. Optional images, audio and videos arrays carry the media, embedded as data, never as remote URLs.

What Cloudflare reports

Every figure here is vendor-run. Cloudflare ran its own copy of the Decision Index 0.2.1 suite and TypeSafe’s workflow evals; it did not build either benchmark. On the 10 tasks in its launch post, Clef-omni beats Jev on eight and trails on When2Call (63.3 against 80.97) and BRIGHT (42.0 against 47.52). It beats Clef on four of the 10. On the five workflow scores it trails Jev on four and ties on security incidents, and it trails Clef on all five. Cloudflare gives a median of about 130 ms for text and 150 ms with images.

It is not on Decision Index 0.3, generated before the launch, or on Benchmark Heaven’s JevBench as of 10 October 2026.

What it’s good at

Decisions that need sound or moving pictures: whether a recording has breaking glass in it, whether a dashcam clip shows a collision, or whether a machine sounds normal. One call replaces a speech-to-text or captioning step before the decision.

What it’s not for

For text-only work, Cloudflare’s own tables favour Clef, and Clef-flash costs a quarter of the price. It writes no text. Self-hosting needs about 64 GB of GPU memory in bfloat16.

Access today

Clef-omni is live on Workers AI as @cf/cloudflare/clef-omni at $0.15 per million input tokens, output free, with media billed as input tokens. OpenRouter lists it as cloudflare/clef-omni at the same price. The weights are on Hugging Face under Apache 2.0.

Specifications

Question typesChoiceScoreNoul
Max Choice optionsNot documented
Score levelsNot documented
Questions per call64
Total context64,000 tokens
State budgetNot documented
Rate limitNot published on the Workers AI model page as of 10 October 2026.
EndpointPOST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef-omni, or env.AI.run("@cf/cloudflare/clef-omni", {...}) from a Worker
SDKs

Workers AI lists a 64,000-token context window; OpenRouter lists 65,536 tokens with a 58,982-token maximum completion, checked 2026-10-10. Media tokens count toward the window together with the questions; a request whose media exceeds it fails, and otherwise the text state is truncated to fit. A request takes 1 to 64 questions. It accepts up to 4 images (PNG, JPEG or WebP, 4 MiB and 16 megapixels each, 8 MiB in total), up to 4 audio clips (8 MiB and 300 seconds each) and up to 2 videos (16 MiB and 60 seconds each, sampled at 2 frames per second). Audio and video together may total 16 MiB. Media must be embedded; remote URLs are not accepted. The docs publish no cap on options per Choice or levels per Score.

Versions

  • @cf/cloudflare/clef-omni, 9 Oct 2026, Post-trained from Qwen3-Omni-30B-A3B-Instruct, a mixture-of-experts model with 30B total and 3B active parameters. Hugging Face counts 35.3B parameters (35,259,818,545) because the repository keeps the base model's unused speech-output weights. Apache 2.0 weights as Cloudflare/clef-omni. Release notes

Use cases

What people use Clef-omni for, one page per pattern.

Workflow controlStarter

Support inbox triage with System One models

Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket.

ChoiceScoreNoul
Workflow controlIntermediate

Confidence-gated actions with System One models

Jev returns a confidence value from 0 to 1 alongside every Choice and Score answer. Your code treats it as a separate axis: act automatically when it's high, confirm or flag when it's middling, hand the decision to a person when it's low. Riskier actions get higher bars.

ChoiceScore

Examples built with Clef-omni

The most-starred and most-viewed entries in the directory. Browse all examples.

Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash

Cloudflare's launch post for Clef-omni, a 30B-A3B mixture-of-experts decision model that reads audio, video, images and text, at $0.15 per million input tokens. It also cuts Clef-flash to $0.038 with a 24K hosted context, and gives Cloudflare's own benchmark and latency tables.