Slop TVNewsLatest
News

Cloudflare ships Clef, the first models it trained itself

The two open-weight models return a probability instead of a sentence, accept images and video, and are free to download under Apache 2.0.

Illustration: Cloudflare ships Clef, the first models it trained itself
Illustration: SLOP TV News. Logos are trademarks of their owners.

Key takeaways

  • Cloudflare introduced Clef and Clef-flash on October 1, 2026, the first models it has trained in-house, live on its Workers AI platform.
  • Clef is 27 billion parameters with a 64,000-token context and accepts text, JSON, images and video; Clef-flash is 9 billion and built for the fastest calls.
  • Both are published on Hugging Face under Apache 2.0, and hosted use costs $0.24 per million input tokens for Clef and $0.09 for Clef-flash, with output tokens not billed.
  • Cloudflare's own benchmarks put Clef far ahead of TypeSafe's Jev, but the leaderboard marks those results self-reported, and an independent local test of Clef-flash found it trailing Jev on accuracy.

Cloudflare introduced Clef and Clef-flash on October 1, 2026, a pair of open-weight decision models that answer with a probability instead of a sentence, and the first models the company has trained in-house.

A decision model skips the writing step. Where a chat model generates an answer token by token, Clef reads a state and a schema once and scores every allowed option, returning a yes-or-no probability, a choice among named alternatives or a ranking. Nothing generated means no invented option and no malformed JSON, which is why the pair run fast.

Clef is the larger model, 27 billion parameters built on Qwen3.8-27B, with a 64,000-token context window and inputs covering text, JSON, images and video. Clef-flash is 9 billion parameters on Qwen3.5-9B and targets latency-critical paths. Both carry a vision encoder that handles up to four images per request. Cloudflare published the weights on Hugging Face under the Apache 2.0 licence, so they can run on a developer's own hardware, and priced hosted use on Workers AI at $0.24 per million input tokens for Clef and $0.09 for Clef-flash, with output tokens not billed at all.

The performance claims are the part to read carefully. Cloudflare says Clef scored a macro-F1 of 94.20 on the BANKING77 classification benchmark against 79.74 for TypeSafe's Jev, and that Clef-flash answered in a median 38.8 milliseconds against Jev's 524.1 milliseconds. Those are the vendor's numbers, and the leaderboard Cloudflare points to flags its results as self-reported and not yet reproduced by the board. One independent test, a 42-case run of Clef-flash on a single 24 GB RTX 3090, found it agreed with Jev's API on 83% of calls but trailed on nuanced triage, scoring 66.7% against Jev's 71.4%. The full Clef needs roughly 54 GB of memory in bf16, while the 9B model fits on one consumer card.

For anyone building a video pipeline with agents, routing, moderation and shot selection are the jobs that burn the most tokens for the least creativity. A model that returns a bounded answer cheaply is the point, and Clef accepting video and image input is what makes it applicable to a footage workflow rather than only to text.

The weights are on Hugging Face now, and Cloudflare is pairing the launch with a reinforcement-learning fine-tuning product.

Sources

  1. blog.cloudflare.com - Cloudflare's launch post and its own benchmarks
  2. cryptobriefing.com - launch date, parameters, pricing
  3. particle.news - the independent local test and its caveats
  4. flaviocopes.com - how a decision model returns a probability instead of prose
  5. dev.to - the independent 42-case Clef-flash benchmark on a single RTX 3090