What Clef-omni is
Clef-omni is the third model in Cloudflare’s Clef family and a System One model. Cloudflare launched it on 9 October 2026, eight days after Clef and Clef-flash. You send a state and typed questions, and it returns a probability for every allowed option, with no generated text. What sets it apart is input: it reads audio (WAV or MP3) and video (MP4 or WebM) as well as text, JSON and images, all in one call. Clef and Clef-flash take text and images only.
It is also built differently. Clef and Clef-flash sit on dense Qwen backbones. Clef-omni is post-trained from Qwen3-Omni-30B-A3B-Instruct, a mixture-of-experts model that uses 3B of its 30B parameters per token. Cloudflare keeps the backbone frozen and trains low-rank adapters and a scoring head on top, with the same loss as Clef: label-smoothed cross-entropy plus a Brier term for calibration. The base model’s speech-output parts are left unused.
What it returns
The request is Jev’s state and questions plus a required model field, and the Hugging Face card calls the API “fully compatible with Jev and SystemOne”. It answers Choice, Score and Noul questions, 1 to 64 per request. Optional images, audio and videos arrays carry the media, embedded as data, never as remote URLs.
What Cloudflare reports
Every figure here is vendor-run. Cloudflare ran its own copy of the Decision Index 0.2.1 suite and TypeSafe’s workflow evals; it did not build either benchmark. On the 10 tasks in its launch post, Clef-omni beats Jev on eight and trails on When2Call (63.3 against 80.97) and BRIGHT (42.0 against 47.52). It beats Clef on four of the 10. On the five workflow scores it trails Jev on four and ties on security incidents, and it trails Clef on all five. Cloudflare gives a median of about 130 ms for text and 150 ms with images.
It is not on Decision Index 0.3, generated before the launch, or on Benchmark Heaven’s JevBench as of 10 October 2026.
What it’s good at
Decisions that need sound or moving pictures: whether a recording has breaking glass in it, whether a dashcam clip shows a collision, or whether a machine sounds normal. One call replaces a speech-to-text or captioning step before the decision.
What it’s not for
For text-only work, Cloudflare’s own tables favour Clef, and Clef-flash costs a quarter of the price. It writes no text. Self-hosting needs about 64 GB of GPU memory in bfloat16.
Access today
Clef-omni is live on Workers AI as @cf/cloudflare/clef-omni at $0.15 per million input tokens, output free, with media billed as input tokens. OpenRouter lists it as cloudflare/clef-omni at the same price. The weights are on Hugging Face under Apache 2.0.
Specifications
| Question types | ChoiceScoreNoul |
| Max Choice options | Not documented |
| Score levels | Not documented |
| Questions per call | 64 |
| Total context | 64,000 tokens |
| State budget | Not documented |
| Rate limit | Not published on the Workers AI model page as of 10 October 2026. |
| Endpoint | POST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef-omni, or env.AI.run("@cf/cloudflare/clef-omni", {...}) from a Worker |
| SDKs |
Workers AI lists a 64,000-token context window; OpenRouter lists 65,536 tokens with a 58,982-token maximum completion, checked 2026-10-10. Media tokens count toward the window together with the questions; a request whose media exceeds it fails, and otherwise the text state is truncated to fit. A request takes 1 to 64 questions. It accepts up to 4 images (PNG, JPEG or WebP, 4 MiB and 16 megapixels each, 8 MiB in total), up to 4 audio clips (8 MiB and 300 seconds each) and up to 2 videos (16 MiB and 60 seconds each, sampled at 2 frames per second). Audio and video together may total 16 MiB. Media must be embedded; remote URLs are not accepted. The docs publish no cap on options per Choice or levels per Score.
Versions
- @cf/cloudflare/clef-omni, 9 Oct 2026, Post-trained from Qwen3-Omni-30B-A3B-Instruct, a mixture-of-experts model with 30B total and 3B active parameters. Hugging Face counts 35.3B parameters (35,259,818,545) because the repository keeps the base model's unused speech-output weights. Apache 2.0 weights as Cloudflare/clef-omni. Release notes
Use cases
What people use Clef-omni for, one page per pattern.
Support inbox triage with System One models
Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket.
Confidence-gated actions with System One models
Jev returns a confidence value from 0 to 1 alongside every Choice and Score answer. Your code treats it as a separate axis: act automatically when it's high, confirm or flag when it's middling, hand the decision to a person when it's low. Riskier actions get higher bars.
Examples built with Clef-omni
The most-starred and most-viewed entries in the directory. Browse all examples.

Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
Cloudflare's launch post for Clef-omni, a 30B-A3B mixture-of-experts decision model that reads audio, video, images and text, at $0.15 per million input tokens. It also cuts Clef-flash to $0.038 with a 24K hosted context, and gives Cloudflare's own benchmark and latency tables.