Skip to main content

Cloudflare Clef Provider Guide

A provider of decide — the same typed boolean / choice / score answers as TypeSafe's Jev, Laya, XOR and Perplexity, from Cloudflare's Clef models on Workers AI, and it also reads images. It emits no text at all.

This is not the Cloudflare Workers AI text provider (cloudflare, which runs chat models). The two are separate providers that read the same CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID; see One token, two providers.

Read Limits before you rely on it. Workers AI reads only about the first 2,048 tokens of the state, far less than the 64K context Cloudflare documents, and ignores text past that point without an error (hosted service or model: unknown). NeuroLink refuses a state it estimates as longer than that rather than let a decision be made on text the model never saw.

Overview​

Clef is a decision model: it reads a state and a set of typed questions and returns a probability for every allowed answer, with no free-form text and no reasoning tokens to wait for. Cloudflare publishes two sizes, clef (27B) and clef-flash (9B), and says it is releasing their weights under the Apache 2.0 license. Both answered in about a second from a developer machine. Cloudflare's pages call them available, and do not say whether they are generally available or in beta.

Key Facts​

Provider idcloudflare-clef
Modelsclef (27B, the default) and clef-flash (9B). @cf/cloudflare/clef and @cf/cloudflare/clef-flash are accepted too
Question typesboolean (sent as noul), choice (2 to 255 options), score (2 to 10 levels)
Questions per request64
Imagesup to 4 per request: PNG, JPEG or WebP. No video
StateWorkers AI reads about the first 2,048 tokens; NeuroLink refuses a state it estimates at more than 1,500
Request size256,000 bytes in NeuroLink, base64 image data included
Endpointhttps://api.cloudflare.com/client/v4/accounts/<account id>/ai/run/@cf/cloudflare/<model>
Credentialsan API token with Workers AI permission, and the account id
Price$0.24 per million input tokens for clef, $0.09 for clef-flash; no output price is listed (Cloudflare's Workers AI pricing page)
Measured latency2026-10-03: 0.3 to 1.0 s for a small request; 64 questions: 1.1 s (clef-flash), 1.3 s (clef); 2026-10-04: 1.5 s / 2.3 s
Default decide providerlast, after TypeSafe, Laya, XOR and Perplexity
Verified live2026-10-03 and 2026-10-04, with a real account (see Limits)

Quick Start​

1. Get a token and your account id​

Create an API token with the Workers AI: Read + Write permission at dash.cloudflare.com/profile/api-tokens, and copy your account id from the dashboard URL or the "Account ID" panel.

2. Configure​

export CLOUDFLARE_API_KEY=your-workers-ai-token
export CLOUDFLARE_ACCOUNT_ID=your-account-id

# Optional
export CLOUDFLARE_CLEF_MODEL=clef-flash # default: clef
export CLOUDFLARE_CLEF_BASE_URL=https://api.cloudflare.com/client/v4 # default

Or pass them in code, which wins over the environment:

import { NeuroLink } from "@juspay/neurolink";

const neurolink = new NeuroLink({
credentials: {
cloudflareClef: { apiKey: "…", accountId: "…" }, // baseURL is optional
},
});

Both the token and the account id are needed: the account id is part of the route, so a token alone does not count as configured.

3. Use it​

const result = await neurolink.decide({
provider: "cloudflare-clef",
state: "Checkout has been failing for every customer for the last hour.",
questions: {
urgent: {
type: "boolean",
instructions: "Is this support request urgent?",
},
team: {
type: "choice",
instructions: "Which team should handle this request?",
criteria: {
billing: "Payments, invoices, and refunds",
technical: "Outages, errors, and configuration",
sales: "Plans and upgrades",
},
},
severity: {
type: "score",
instructions: "How severe is the customer impact?",
criteria: ["No impact", "Minor", "Major", "Critical"],
},
},
});

result.answers.urgent; // { type: "boolean", probability: 0.99 }
result.answers.team; // { type: "choice", choice: "technical", probabilities: {…}, confidence: 0.53 }
result.answers.severity; // { type: "score", score: 2.96, legend: {…}, probabilities: {…}, confidence: 0.92 }

The figures in the comments are what a live call to clef returned for this example on 2026-10-03, rounded; clef-flash answered the same example with 0.96, 0.82 and 0.50, 2.72. Pick the faster model for one call with model: "clef-flash".

From the CLI​

neurolink decide "Checkout has been failing for every customer for the last hour." \
--provider cloudflare-clef \
--questions '{"urgent":{"type":"boolean","instructions":"Is this support request urgent?"}}'

Images​

Pass up to four images beside the state, as a Buffer, a file path or a data:image/…;base64, URL; an http(s) URL is refused. PNG, JPEG and WebP are read; image/jpg is accepted as a JPEG alias and sent as image/jpeg. Other formats are refused before they are sent.

await neurolink.decide({
provider: "cloudflare-clef",
state: "Look at the attached image.",
images: ["./photo.png"],
questions: {
color: {
type: "choice",
instructions: "What colour is the image?",
criteria: { red: "red", blue: "blue" },
},
},
});
  • They travel in their own images array, not inside the state. Cloudflare places them before the state.
  • The Workers AI API accepts no video. Its model page says it does, but the API refuses a video and has no field for one, so NeuroLink refuses a video too.
  • Pixels: Cloudflare documents 16 megapixels an image. 16.00 megapixels were accepted (4096 × 3900 at 15.97 MP, 8000 × 2000 and 16000 × 1000 at 16.00 MP); 16.38 megapixels (4096 × 4000) were refused with a 422. The limit is on the pixel count, not on either side.
  • Cost: an image costs at most about 1,000 input tokens however large it is. A test request with one 512 × 512 image counted 389 input tokens in all, one with a 1,024 × 1,024 image counted 1,157, and larger images up to the pixel limit counted between 1,125 and 1,157.
  • How many: exactly four tiny PNG/JPEG/WebP images succeeded on both models on 2026-10-04; a fifth was refused with a 422. Four is also NeuroLink's limit.
  • NeuroLink's size limit is 256,000 bytes for the encoded request body, including base64 image data. A 195 KB PNG encodes to about 260,000 characters and is refused locally. Resize or compress a photo to well under 190 KB first. The 2026-10-03 API probe accepted a 195 KB PNG and refused a 202 KB one. On 2026-10-04 the text-only request ceiling moved (see below); the current image-byte ceiling was not remeasured, so those image figures are historical.

Built-in features that call decide() pick the first decision provider that is configured, in this order: TypeSafe, Laya, XOR, Perplexity, then Clef. A host with none of the first four, but with CLOUDFLARE_API_KEY and CLOUDFLARE_ACCOUNT_ID set for the Workers AI text provider, therefore has Clef as its default decision provider. Naming provider: "cloudflare-clef" reaches it whatever else is configured.

Because of the state window below, a built-in feature whose state is longer than about 1,500 estimated tokens gets a refusal from Clef, which each of them treats as "carry on as before". Clef suits short decisions: routing a request, a guardrail on one action, classifying one message or one picture.

What is sent to Cloudflare​

POST https://api.cloudflare.com/client/v4/accounts/<account id>/ai/run/@cf/cloudflare/clef
Authorization: Bearer <token>

{
"model": "clef",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {
"urgent": { "type": "noul", "instructions": "Is this support request urgent?" }
},
"images": ["data:image/png;base64,…"]
}

boolean is the SDK's name for the wire's noul. model is sent although the path already names it: Cloudflare's schema marks it required, and a body whose model differs from the path is refused. A request with no model was accepted live and answered by clef-flash, so sending it matters only for following the schema. Cloudflare wraps every answer in its own envelope, and NeuroLink reads the answers out of it:

{
"result": {
"model": "clef",
"answers": { "urgent": { "type": "noul", "noul": 0.9551 } },
"usage": { "input_tokens": 346, "output_tokens": 0 }
},
"success": true,
"errors": [],
"messages": []
}

The request id is the cf-ai-req-id response header; where that is absent (a rejected token never reaches the model) it is the edge's cf-ray.

One token, two providers​

cloudflare-clef and cloudflare read the same two variables, but each has its own credentials slice, so the two cannot be mixed up:

cloudflare (text)cloudflare-clef (decide)
TokenCLOUDFLARE_API_KEYCLOUDFLARE_API_KEY
Account idCLOUDFLARE_ACCOUNT_IDCLOUDFLARE_ACCOUNT_ID
SDK credentialscredentials.cloudflarecredentials.cloudflareClef
Servesgenerate() and stream()decide() only

credentials.cloudflare does not configure decide. Setting the two variables for the text provider also lets decide() use Clef when no other decision provider is configured, and no switch turns that off while the two variables are set in the environment. A host that wants the Workers AI text provider without Clef can give the text provider its token and account id only through credentials.cloudflare, and leave the environment variables unset.

Limits​

Measured on a real account, October 2026​

The Workers AI endpoint ignores text past about 2,048 tokens, far below Cloudflare's documented 64K context (hosted service or model: unknown). Request-size refusals also changed between two measurement dates. The original measurements were on 2026-10-03; the bounded follow-up campaign used 133 probe calls on 2026-10-04, with one request at a time.

Which model: on both clef-flash and clef, the follow-up still read the fact at the 2026-10-03 clef-flash lower bound for logs, number lists, digit arrays and compact JSON, and did not read it about 2.5% further on. English prose and random CJK had already matched on both models. It also checked the question, id, option, score-level, PNG/JPEG/WebP, pixel and four/five-image boundaries on clef, and the four-image and image/jpg cases on both models. Natural-script prose was measured on clef; a reworked many-key object probe matched on both. TypeScript, minified JSON and synthetic Devanagari/emoji cuts remain 9B-only measurements. Historical latency, billing and burst figures retain their dates.

LimitDocumentedMeasured
Statea 65,536-token context window; the schema adds that long text "is truncated to fit the model's token limit"text past about 2,048 tokens is ignored without an error; hosted service or model: unknown
Request size4 MiB an image, 8 MiB in all, 13 MiB for the body2026-10-04, both models: 520,000 text characters accepted, 525,000 refused with 413/code 5021; NeuroLink retains its 256,000-byte cap
Questions1 to 6464 sent, the 65th refused with a 422
Question idsletters, digits, _, ., -, up to 100the same; 101 characters and a space refused
Options and levels2 to 255 options, 2 to 10 levelsthe same; 256 and 11 refused
Imagesup to 4, 16 megapixels each4 tiny images accepted, 5 refused; PNG/JPEG/WebP and image/jpg accepted; 16.00 MP accepted, 16.38 MP refused on both models
Videothe model page says it reads videorefused: no field, and not accepted in images
Ratenot documented on the model page; the pricing page links a limits page and a free allowance of 10,000 neurons a day60 requests sent at once: 56 answered, 4 refused with a 429

The state window. A fact placed past about 2,048 tokens of state was ignored (a 75-character sentence repeated for 60,000 characters: the fact was used at 10,500 characters and ignored from 12,000), a question about the end of a long text was answered from its beginning, and the reported input tokens stayed at about 2,190 for any longer text, which is also what is billed. The cut falls at a different character count for each kind of text. The first five rows are real text; the last three are synthetic, built from cycled code points, so natural text in those scripts may tokenise differently:

TextCharacters before the cutPer token
English prose (Cloudflare's launch post)7,5173.7 characters
TypeScript (this repository's providers)8,7684.3 characters
Minified JSON (a documentation index)9,9004.8 characters
Log lines (generated)2,8631.4 characters
A list of numbers (generated)2,0401.0 characters (each digit is a token)
A JSON array of single digits (generated)2,0431.0 characters (commas count as well)
Four-digit numbers, space separated2,0391.0 characters
Records as an object array or compact JSON5,141 (97 records)2.5 characters
The same records, pretty-printed JSON text4,055 (49 records)2.0 characters
Random CJK ideographs (synthetic)1,1701.75 tokens a character
Cycled Devanagari (synthetic)1,7911.14 tokens a character
Cycled emoji (synthetic)7072.9 tokens a character

An array of records passed as state was cut at the same place as compact JSON text. On both models, the follow-up still read the fact at the 2026-10-03 clef-flash lower bound and did not read it about 2.5% further on: 2,863–2,935 characters for logs, 2,040–2,091 for number lists, 2,043–2,095 for digit arrays, and 97–100 records for compact JSON. These two-point checks do not establish exact tokenizer boundaries.

The many-key object control failed three times, then passed after an explicit field-name question and zero-padded keys were introduced together. On both models the hidden field stopped being read between 128 and 134 preceding keys. The lower bound is 4,883 characters of compact JSON and 2,299 tokens from the provider's estimator, above the local 1,500-token cap. This verifies the tested object shape; it does not establish that every structured state is serialized identically by Workers AI.

Natural samples were composed as about 3,000 code points of ordinary prose and repeated to 12,000 to find a cut. The intervals below are on clef (27B); the estimate is evaluated at the lower, still-visible bound.

SampleCharacters before cut, lower–upperNeuroLink estimate at lower bound
Chinese prose3,749–4,4995,584
Japanese prose3,749–4,4995,588
Korean prose4,499–5,2495,309
Hindi prose2,999–3,7493,751
Emoji-rich English conversation5,249–5,9992,396

All exceed 1,500 estimated tokens before the observed cut, so no estimator rate changed. The mixed emoji conversation is not a measurement of every emoji or multi-code-point sequence.

The request size changed. On 2026-10-03, clef-flash accepted 262,000 text characters and refused 270,000; a 195 KB PNG was accepted and a 202 KB one refused. The 250,000-character request was first tested on 2026-10-04. On that date both models accepted 270,000, 500,000, 510,000 and 520,000 text characters, then refused 525,000, 530,000 and 1,000,000 with 413/code 5021.

The refusal estimate agrees with the encoded body characters / 4, rounded up. What changed was the refusal threshold: on 2026-10-03, clef-flash accepted an estimate of 65,527 and refused 67,527; on 2026-10-04, both models accepted 130,026 (clef) / 130,027 (clef-flash) and refused 131,276 / 131,277. The error still prints a 65,536-token context. Inference: the new interval contains 131,072, twice the printed figure; the reason is unknown. The observed text ceiling is between 520,000 and 525,000 characters, not a stable API guarantee. The image-byte boundary was not repeated on this date.

NeuroLink keeps the 256,000-byte encoded-body cap. It remains below the accepted request sizes and preserves the conservative image-size behavior. The 256,000-byte cap with images has mocked coverage but is not live-tested. Live canary 19.11 checks the text-only service ceiling on clef-flash.

RefusedLimitError kind
A state estimated over 1,500 tokensdigits 1 token each, punctuation 0.75, emoji 3, other non-ASCII text 1.5, letters about 4 characters a tokenmax_tokens_exceeded
More than 64 questions in one decide()64max_tokens_exceeded
A request body over 256,000 bytes256,000invalid_request
More than 4 images, a video, or a format other than PNG, JPEG or WebPinvalid_request
A model name that is not a plain name, such as clef/../xinvalid_request
A missing token or account id, or a base URL that cannot workauthentication or invalid_request

NeuroLink.tryDecide(), which every built-in feature calls, splits a question map at 64 and runs up to four batches at once, so only a direct decide() meets the 64-question refusal.

Why 1,500 and not 2,048. The limit is an estimate in NeuroLink's own tokens, not the model's, and it was chosen so that for every kind of text in the table above the estimate reaches 1,500 before the endpoint ignores the remaining state. A rate for digits alone was not enough: a JSON array of single digits, whose commas are tokens too, passed with 15% of its length ignored without an error, which is why punctuation is charged 0.75 and emoji 3. The price is a state of ordinary text uses only about 60% of the window before it is refused: English prose is refused at about 1,500 estimated tokens, which is about 1,200 of the model's (the estimate runs a few percent low for prose, and about 12% high for code), and compact JSON at half the window. The natural samples above also reached the local cap before their observed cuts. The estimates remain conservative measurements for those samples, rather than a guarantee for every script or emoji sequence.

What is not verified​

  • The status Cloudflare returns for a token without Workers AI permission. No such token was available. A 403 is treated as an ordinary invalid_request, not authentication, so that if it does mean a missing permission, fixing it in the dashboard works at once, without restarting the process.
  • Any 5xx. Cloudflare answered none during the probe, so the retry and classification of 500, 502, 503 and 504 rest on ordinary HTTP conventions.
  • What happens to an account that has run out of credit.
  • The 27B TypeScript, minified-JSON and synthetic Devanagari/emoji cuts, and the image-byte ceiling after the observed request-size change; see the model coverage above.
  • The rate limit. Only a burst of 60 was tried.
  • Whether Cloudflare's service is generally available or in beta.
  • Whether the 2,048-token state window and the changing request-size limit belong to the hosted service or to the model itself. The open weights on Hugging Face were not run.
  • Whether either limit changes. The live decide suite carries three canary cases (19.10, 19.10b and 19.11) that fail with an instruction to re-measure if it does; 19.11 checks both sides of the text-only request ceiling on clef-flash.

Latency and the timeout​

On 2026-10-03, a small request answered in 0.3 to 1.0 s from a developer machine; 64 questions took 1.1 s on clef-flash and 1.3 s on clef; the slowest of 60 requests sent at once took 2.2 s. On 2026-10-04, a 64-question request with the same questions and a shorter state took 1.5 s on clef-flash and 2.3 s on clef. The default timeout is 5 s. Pass timeoutMs to change it for one call.

Errors​

Cloudflare answers with its own envelope, { "success": false, "errors": [{ "code", "message" }] }, and a refusal from the model nests a second envelope, with a trailing request id, inside message. NeuroLink flattens both to one line and takes the request id out of the text.

StatusCodeSeen forKindRetried
40110000a token Cloudflare does not knowauthenticationno; trips the breaker
403never seen; a token without Workers AI permission may draw itinvalid_requestno; does not trip the breaker
4005006no questions; an unknown question typeinvalid_requestno
4007000a model path that does not exist ("No route for that URI")invalid_requestno
4006003a body that is not JSONinvalid_requestno
4225012a validation failure: 65 questions, a fifth image, an unknown field, an empty instructioninvalid_requestno
4135021a request past the estimate abovemax_tokens_exceededno
4293040"Capacity temporarily exceeded"; no Retry-Afterrate_limityes, after the default backoff
5xxnever seen from Cloudflareserver, or overloaded for a 503yes, once

Only a 401 is authentication. After one, that provider instance refuses further calls without sending anything; construct a new one with a working token. "Yes" means one retry, after about 250 to 500 ms; the base class retries once.

Troubleshooting​

  • "requires the account id" — set CLOUDFLARE_ACCOUNT_ID or credentials.cloudflareClef.accountId.
  • "The state is ~N tokens; Cloudflare's … model reads at most 1500" — shorten the state, or send only the part the question is about. Cloudflare would have cut the rest without telling you.
  • "No route for that URI" — the model name or the base URL is wrong. Clef is served as clef and clef-flash, and the base URL must end in /client/v4.
  • A 429 "Capacity temporarily exceeded" — retried for you; if it keeps happening, send fewer requests at once.
  • An image is refused as "The request is N bytes; Cloudflare accepts at most 256000" — Cloudflare counts the base64 text of the image, so NeuroLink refuses it locally under the retained conservative cap. Resize or compress it to well under 190 KB; a 195 KB PNG, which Cloudflare accepted, is about 260,000 characters in base64 and is refused here.
  • decide() uses Clef although you never chose it — you have the Workers AI token and account id set and no other decision provider; see When NeuroLink uses it.

See also​