Shopify's Canvas Just Cut Merchant Time-to-Store From Two Weeks to Twenty Minutes. Here's What That Means for Every E-Commerce Activation Playbook.
On October 1, 2026, Cloudflare released Clef and Clef-flash on Workers AI: open-weight models that take a state plus typed questions and return calibrated probabilities instead of generating tokens. At $0.24/$0.09 per million input tokens with no output billing and 38.8 ms median latency, they reframe how enterprises budget agentic AI routing and guardrail costs.
On October 1, 2026, Cloudflare released Clef and Clef-flash on its Workers AI platform — the company's first internally trained open-weight models. They are decision models: instead of generating text, they take a state and a set of typed questions and return calibrated probability distributions over predefined answers in a single forward pass. Output tokens are not generated. Output tokens are not billed. The median latency for Clef-flash is 38.8 milliseconds.
The announcement is technically narrow. It is strategically broad. The agentic AI cost curve has a long tail that most enterprise teams have not yet fully mapped, and decision-layer compute is one of the largest line items hiding inside it.
What a Decision Model Is and Why It Matters
A standard large language model generates text. That is what it does: given an input, it samples from a probability distribution over its vocabulary to produce an output sequence. Even when you use a language model for a classification task — routing an intent, evaluating whether an output is safe, deciding which tool to call next — the model is still generating tokens. Those output tokens cost money. They also add latency: autoregressive generation requires sequential sampling, and even a short classification response takes time to produce.
A decision model does not generate text. Clef takes a state (any combination of text, structured JSON, or images) plus a schema of typed questions with defined answer sets, and produces a probability vector over the allowed answers for each question in a single forward pass. There is no output token generation. The response is a structured probability distribution that is directly machine-parseable without parsing logic.
The Decisions API that Clef uses looks like this in practice: you define a schema with questions like "Is this request a refund inquiry? [yes, no, unsure]" and "What is the urgency level? [low, medium, high, critical]", you send the current state as input, and you receive calibrated probabilities for each answer to each question simultaneously. One forward pass answers all questions at once.
This architecture has two operational consequences that matter for enterprise deployment:
Output tokens are never billed. Cloudflare prices Clef at $0.24 per million input tokens for the 27 billion parameter model and $0.09 per million input tokens for Clef-flash. There is no output token charge. For a typical classification call at 500 input tokens, Clef-flash costs $0.000045 per call.
Latency is determined only by the forward pass. Clef-flash posts 38.8 ms median latency on Workers AI. TypeSafe's Jev, the incumbent decision model that Clef is designed to replace, posts 524.1 ms. The difference is not model size — it is architecture. Sequential token generation is structurally slower than a single forward pass for a fixed-answer classification task.
The Agentic Cost Stack That Enterprises Are Not Fully Accounting For
The token cost conversation in enterprise AI has been dominated by generation costs: the per-token rates for Claude, GPT-6, and Gemini that appear on API invoices and get into budget spreadsheets during procurement. Those rates have fallen dramatically — tokens are roughly 98% cheaper than they were in 2024 — and the falling cost of generation has obscured a different category of AI compute cost that is growing as agentic deployments mature.
Agentic AI workflows do not just generate content. They make decisions — hundreds or thousands of decisions per task, depending on the complexity of the workflow. Which tool to call next. Whether an input is safe to process. What intent the user has. Whether a retrieved document is relevant to the current query. Whether to continue, branch, or terminate a step. Whether an output meets quality standards before returning it.
In the current dominant architecture, these decisions are made by the same frontier model handling generation. Every routing decision, every guardrail check, every relevance judgment is billed at frontier generation rates — including the output tokens for the classification label and any rationale the model generates. At scale, the decision layer becomes a significant fraction of total model spend.
The math compounds quickly. Consider an enterprise customer support agent handling 1 million interactions per month. Each interaction requires: an intent classification (is this a refund, a complaint, a technical issue, or something else?), a safety evaluation (does the input contain anything that should trigger escalation or refusal?), a relevance check per retrieved document (is this context actually useful for answering this question?), and a quality evaluation before sending the response. If there are four retrieved documents per interaction and all four checks run at frontier model rates, that is six decision calls per interaction, or 6 million decision calls per month — before a single word of the response is generated.
At $5–10 per million input tokens (frontier pricing) with typical decision call volumes, that is $30,000–$60,000 per month in decision-layer costs alone for a single customer support deployment. Clef-flash at $0.09 per million input tokens brings the same 6 million calls to $540 per month. The savings — roughly 98% — are not marginal. They are structural.
The Open-Weight Architecture and Its Implications
Both Clef and Clef-flash are released under Apache 2.0 and available on Hugging Face. The 27B parameter Clef model is a post-training fine-tune of Qwen3.8-27B with a vision encoder and 64k context window added. Clef-flash is built on Qwen3.5-9B with the same additions.
The open-weight release is strategically important for several reasons.
First, it removes Workers AI lock-in as a concern for enterprise adoption. Enterprises evaluating Clef do not have to commit to Cloudflare's inference infrastructure to use the model. The weights are downloadable; the model can run on private infrastructure with the same API interface. For enterprises with existing on-premises AI infrastructure or sovereign cloud requirements, the open weights make Clef a viable choice without routing decision traffic through Cloudflare's network.
Second, open weights enable fine-tuning on proprietary decision schemas. Enterprise decision tasks are often domain-specific: a legal document routing model needs to recognize legal concepts; a financial transaction classifier needs to understand payment terminology; a medical record triage system needs domain-specific scoring criteria. The open weights allow enterprises to fine-tune Clef on their proprietary decision datasets without sharing that data with Cloudflare or any external service. The fine-tuned model can then be deployed on private infrastructure or run on Workers AI depending on the enterprise's infrastructure preference.
Third, the Apache 2.0 license removes the commercial use restrictions that have complicated enterprise adoption of some open-weight models. Nvidia's acquisition of Hugging Face created questions about the long-term openness of the open-weight model ecosystem; a Cloudflare model under Apache 2.0 provides a credible alternative source for open-weight decision models that enterprises can rely on without concern about licensing changes.
Clef Against the Incumbent: TypeSafe Jev
TypeSafe's Jev is the incumbent decision model that Clef directly targets. Jev has been the primary commercial offering in the decision model category, built on TypeSafe's System One API — the standardized interface for submitting states and question schemas and receiving probability distributions in return.
Cloudflare made Clef fully compatible with the System One API. Existing integrations built on Jev can switch to Clef by changing the model endpoint. The calling code does not change.
The performance comparison on TypeSafe's own published benchmarks, as reported by Developers Digest:
| Metric | Clef-flash | Jev System One |
|---|---|---|
| Median latency | 38.8 ms | 524.1 ms |
| Input token pricing | $0.09/M | $0.30/M (estimated) |
| Output token pricing | None | Billed |
| Model weights | Open (Apache 2.0) | Closed |
| Parameters | 9B | Undisclosed |
| Vision input | ✓ | Limited |
| Context window | 64k | 32k |
The latency difference — 38.8 ms versus 524.1 ms — is not a marginal improvement. It is a 13x difference in response time for a decision call. For agentic workflows where decisions gate subsequent steps, decision latency compounds: a 10-step workflow with a 500ms decision at each step adds 5 seconds of decision overhead before accounting for generation latency. At 38.8 ms per decision, the same workflow adds 388 ms.
The pricing difference depends on workload characteristics. If existing Jev deployments send short inputs, the per-token difference is small. If inputs are long (full documents, long conversation histories, detailed state objects), the per-token differential becomes the dominant cost factor.
The Infrastructure Design Choice Clef Forces
The emergence of purpose-built decision models does not make frontier models less important. It does force an architectural decision that most enterprise AI teams have been deferring: should every layer of an agentic workflow run on the same model, or should different layers use different models optimized for their specific task type?
The current dominant architecture — sometimes called the "one model, all tasks" pattern — routes generation, routing, guardrailing, relevance scoring, and tool selection through a single frontier model. The advantages are simplicity: one vendor relationship, one API, one billing account, one performance optimization focus. The disadvantages are cost and latency for decision-layer operations that do not require frontier-level capability.
The disaggregated architecture that Clef enables:
1. Generation layer (frontier models): Complex content generation, multi-step reasoning, code synthesis, document analysis. Claude Opus 5.5, GPT-6 Astra, Gemini 4 Argon. High capability, higher cost per call.
2. Decision layer (purpose-built decision models): Routing, classification, guardrails, relevance scoring, control flow. Clef, Clef-flash. Low latency, minimal cost, structured output.
3. Retrieval layer (embedding models): Semantic search, document retrieval, context selection. Embedding-specific models optimized for similarity. Low cost, high throughput.
4. Code execution layer (specialized code models): Code generation, test execution, debugging. Purpose-built coding models where relevant.
Building this architecture adds integration complexity — you maintain relationships with multiple model vendors, manage multiple API contracts, and build routing logic that directs tasks to the appropriate layer. The cost savings at enterprise scale can be substantial enough to justify the complexity. The enterprise AI cost optimization conversation has been dominated by frontier model pricing negotiations; disaggregation is an architectural alternative that reduces costs through workload matching rather than price negotiation.
What the "No Output Token Billing" Structure Changes
The zero output token charge on Clef is not just a pricing feature. It is a design constraint that encodes the product's value proposition.
Decision models produce no text output. They produce probability distributions over typed answer sets. There is nothing to bill for on the output side because there is no output in the traditional sense — only an internal computation result expressed as a probability vector. The architecture and the billing model are the same design decision: a model that cannot generate text cannot be billed for text generation.
This matters for enterprise budget modeling. Current AI cost models treat every model call as having an input cost and an output cost, with the ratio determined by the task. Classification calls are output-cheap (short responses) but still incur output charges. Clef eliminates the output cost category entirely for the decision layer, which simplifies cost projection significantly: decision layer costs are purely a function of input volume and average input length. There is no output volume to estimate, no token generation to account for, and no variation in output cost based on how verbose the model decides to be.
For enterprises building cost models for new agentic deployments, this simplification is operationally valuable. The variable that most frequently causes AI cost forecasts to miss budget is output volume — models that are slightly more verbose than projected in testing can double or triple output token costs at production scale. Clef removes that variable from the decision layer entirely.
The Workers AI Distribution Strategy
Cloudflare built Kitesurf, an agent-first browser for AI agent workloads, in August 2026. Now it has built its own model. The pattern is consistent: Cloudflare is constructing an AI infrastructure stack for agent-specific workloads, layer by layer, using its global network distribution as the common foundation.
Workers AI's competitive position is latency and distribution. Cloudflare operates more than 300 data centers globally and routes inference requests to the nearest node with available capacity. For latency-sensitive applications — and decision models, at the heart of an agentic loop, are among the most latency-sensitive AI workloads — edge distribution produces meaningful performance advantages over centralized cloud inference.
The Clef launch ties the open-weight model strategy to the Workers AI distribution advantage. Enterprises that deploy Clef on Workers AI get the model and the distribution infrastructure. Enterprises that deploy the Clef weights on their own hardware get the model without Cloudflare's network advantage but with full infrastructure control. Both paths serve Cloudflare's long-term goal: establishing Clef as the standard decision model interface, with Workers AI as the preferred deployment target for teams that do not have private AI infrastructure.
The Enterprise Procurement Playbook for Decision Models
For enterprise teams evaluating whether to add a dedicated decision layer to their agentic architecture, the evaluation framework has five components:
1. Audit current decision call volume and cost. Before evaluating models, measure what you are currently spending on decision-type calls — classification, routing, guardrails, relevance scoring — in your existing agentic deployments. Enterprise teams frequently discover that decision calls constitute 40–60% of total model call volume once they segment their API logs by call type. If your decision call volume is below 1 million per month, the complexity of adding a separate decision model may not justify the cost savings. Above 5 million calls per month, the economics strongly favor a dedicated decision layer.
2. Benchmark Clef and Clef-flash on your specific decision schemas. General benchmarks do not predict accuracy on domain-specific classification tasks. Build a representative sample of your decision inputs — 200–500 examples per decision type — and measure Clef's accuracy against your existing frontier model baseline. For well-defined, bounded decision categories (intent routing, binary safety classification), Clef's performance on its purpose-built training distribution should match or exceed frontier model performance. For ambiguous, nuanced decisions that require broader world knowledge, frontier models may retain an accuracy advantage that outweighs the cost savings.
3. Evaluate latency requirements at the workflow level, not the call level. A 38.8 ms decision call versus a 500 ms frontier model call is a 461 ms difference per call. In a 10-step agentic workflow with one decision per step, that is 4.6 seconds of latency reduction — significant for interactive applications, negligible for background processing. Map your latency-sensitive workflows and prioritize decision model adoption there first.
4. Model the infrastructure integration cost. Adding a decision model layer requires building routing logic that directs different call types to different models, managing a second vendor API relationship, and maintaining a second set of latency and accuracy monitoring metrics. Estimate this integration cost before projecting ROI. For teams with mature agent infrastructure, the integration is typically a few days of engineering work. For teams building their first agentic deployment, adding the decision model architecture adds complexity that may slow initial shipping.
5. Plan the fine-tuning roadmap. The open-weight Clef model enables enterprise fine-tuning on proprietary decision datasets. If your decision schemas are highly domain-specific, plan a fine-tuning experiment at 3–6 months post-deployment to measure whether a fine-tuned Clef variant outperforms the base model on your specific tasks. The Apache 2.0 license and open weights make this practically achievable without Cloudflare involvement.
What the October 2026 Decision Model Launch Signals
Cloudflare's Clef is the first decision model from a major infrastructure provider. TypeSafe built the category from a startup position. Cloudflare's entry signals that the decision model layer has graduated from a niche optimization to an infrastructure component that cloud providers consider worth building and distributing.
The pattern here is recognizable: embedding models followed a similar path, from specialized startup offerings (Cohere, VoyageAI) to infrastructure-level offerings from every major cloud provider. Decision models appear to be following the same trajectory, roughly 18–24 months behind. If the pattern holds, the major hyperscalers — AWS, Azure, Google Cloud — will have their own decision model offerings within the next 12–18 months. Cloudflare has moved ahead of that wave with open weights, API compatibility, and a clear performance advantage over the existing commercial incumbent.
For enterprise teams building agentic AI infrastructure today, the decision model category is immature enough that the early adoption cost — integration work, evaluation time, monitoring overhead — is real. The long-term cost savings at scale are substantial enough that teams building high-volume agentic deployments should include the decision layer in their architecture from the start rather than retrofitting it later.
Takeaway: Cloudflare Clef is not a better version of an existing model category — it is a new component category for a workload that enterprise agentic deployments already have at scale. Decision calls — routing, classification, guardrails, relevance scoring — represent a large fraction of total model call volume in mature agent deployments, but they are often hidden inside frontier model billing because the current dominant architecture routes everything through the same model. At $0.09 per million input tokens with zero output billing and 38.8 ms latency, Clef-flash brings decision-layer costs down by 96–98% compared to frontier model classification. For any enterprise running more than 5 million decision calls per month, the ROI of adding a dedicated decision layer is straightforward. The open weights under Apache 2.0 remove vendor lock-in as a blocking concern. The System One API compatibility means Jev integrations migrate by changing an endpoint. The remaining question is not whether to adopt purpose-built decision models — it is how quickly enterprise infrastructure teams build the routing layer that sends the right calls to the right model tier.
Frequently Asked Questions
What is Cloudflare Clef and how is it different from a regular language model?
Cloudflare Clef is a decision model, not a text generation model. A standard large language model takes a prompt and generates a sequence of tokens — words, sentences, structured output — through an autoregressive sampling process. Clef takes a state (a support ticket, a web page, a JSON object, an agent trace) plus a schema of typed questions and returns a calibrated probability distribution over predefined answers for each question, all in a single forward pass without generating any output tokens at all. The practical difference is significant: a text generation model used for classification or routing must still generate text — even if only a single word — incurring output token costs and adding latency from the sampling loop. Clef produces no output tokens. It processes the input, computes probabilities internally, and returns the probability vector directly. This means output tokens are not billed, latency is determined only by the forward pass, and the response is immediately machine-parseable without a parsing step. Clef-flash achieves 38.8 milliseconds median latency on Workers AI. The 27 billion parameter Clef model optimizes for accuracy over speed. Both are released under Apache 2.0 and available on Hugging Face, so enterprises can run them on their own infrastructure as well as through Workers AI.
What use cases is Cloudflare Clef built for?
Cloudflare designed Clef specifically for the decision and routing layer of agentic workflows — the part of an agent pipeline that determines what to do next, which tool to call, or how to classify an input, rather than the part that generates content. Core use cases are: intent classification (routing a user message to the correct sub-agent or function), guardrail evaluation (determining whether an input or output is safe, relevant, or policy-compliant), relevance scoring (ranking retrieved documents or tool outputs by relevance before passing them to a more expensive generation model), routing in multi-agent systems (deciding which specialized agent should handle a task based on its characteristics), and agentic control flow (deciding at each step of a workflow whether to continue, branch, or terminate). The Decisions API exposes a standardized interface: you send a state object plus a schema of questions with typed answer sets, and receive probabilities back. The schema approach means Clef can answer multiple questions about a single input in one forward pass — for example, simultaneously classifying intent, evaluating safety, and scoring urgency — without multiple model calls.
How does Cloudflare Clef pricing compare to using a frontier model for classification tasks?
The cost difference between Clef and using a frontier model for classification is substantial, and it compounds with scale. A frontier model like GPT-6 Sol or Claude Opus 5.5 charges for both input and output tokens. A classification call that sends 500 tokens and generates 10 tokens (the class label and maybe a brief rationale) costs roughly $0.003–$0.005 at current frontier pricing for that token volume. At 10 million classification calls per day — a reasonable volume for an enterprise guardrail or routing layer — that is $30,000–$50,000 per day, or $10–18 million per year for a single decision layer. Clef charges $0.24 per million input tokens for the 27B model and $0.09 per million input tokens for Clef-flash, with zero output token charges. At the same 500-token input volume and 10 million calls per day, that is $1,200–$450 per day for Clef and Clef-flash respectively, versus $30,000+ for a frontier model. The savings are approximately 96–98%. The trade-off is capability: Clef is not a general-purpose reasoner and cannot replace a frontier model for generation tasks. But for pure classification and routing decisions — which represent a large fraction of total model calls in a mature agentic deployment — the cost differential justifies purpose-built decision models as a separate layer in the architecture.
Is Cloudflare Clef compatible with existing decision model APIs like TypeSafe Jev?
Yes. Cloudflare explicitly engineered Clef to be compatible with the TypeSafe System One API, which is the interface used by Jev — the incumbent decision model that Clef is positioned against. This means existing integrations built on the Jev API can switch to Clef by changing the model endpoint without rewriting the calling code. The API compatibility is strategic: TypeSafe has established a developer audience around Jev and the System One interface, and Cloudflare's Clef adoption path is designed to be as frictionless as possible for that existing audience. The System One interface defines how states and question schemas are submitted and how probability responses are structured. Clef implements this interface, adds a vision encoder (enabling image and document inputs alongside text), and runs on Workers AI's global network infrastructure. The latency advantage Clef-flash demonstrates against Jev — 38.8 ms versus 524.1 ms — is large enough that the API compatibility may be the deciding factor for latency-sensitive workloads regardless of capability differences. Enterprises evaluating Clef should run their specific workload on both models to measure accuracy on their own decision schemas, since general benchmarks may not reflect performance on domain-specific classification tasks.
What does Cloudflare Clef mean for enterprise AI infrastructure architecture?
Clef represents a design philosophy for agentic AI infrastructure that disaggregates the model layer by task type. The current dominant architecture routes all agent tasks — intent classification, content generation, tool selection, output validation — through the same frontier model, which simplifies the implementation but uses the most expensive compute for every task regardless of whether that compute is warranted. A disaggregated architecture routes generation tasks to frontier models, decision and routing tasks to purpose-built decision models like Clef, retrieval tasks to embedding models, and code execution to specialized code models. Each tier uses the minimum capable model for its task type. For enterprise teams managing AI infrastructure costs at scale, this architecture can reduce total model spend by 40–70% while maintaining or improving latency for decision-heavy parts of the pipeline. Cloudflare's decision to open-weight Clef under Apache 2.0 means enterprises can embed the model in their own infrastructure without Workers AI dependency. The commercial advantage of Workers AI is latency and global distribution, not model exclusivity. For enterprises already running private AI infrastructure, the open weights enable the same cost disaggregation benefit on-premises.
How does Cloudflare's Clef relate to its broader AI infrastructure strategy?
Cloudflare has been methodically building an AI infrastructure stack that competes with hyperscalers on the specific characteristics that hyperscalers struggle with: latency, global distribution, and developer simplicity. Workers AI launched in 2023 as a serverless inference platform running on Cloudflare's global network of 300+ data centers — the premise being that inference at the edge is fundamentally faster than inference at a centralized cloud region. The AI Gateway followed as a model routing and observability layer. Clef represents the third leg of this strategy: Cloudflare is now training and distributing its own models, not just running other companies' models on its infrastructure. The decision to start with a decision model rather than a general-purpose language model is deliberate: decision models are a defined, bounded problem space where Cloudflare can credibly compete on performance and cost without attempting to match the frontier labs on general capability. The Kitesurf agent browser, released in August 2026, showed Cloudflare's willingness to build specialized infrastructure for agent-specific workloads rather than general-purpose tools. Clef is the model-layer equivalent of that strategy — a purpose-built component for a specific workload class, rather than a general model trying to do everything.