Cloudflare AI Search fits a text-first collection when files stay within their format caps and instance limits. Images need an image-capable embedding model for direct visual matching; scanned PDFs need OCR. Before billing starts November 1, 2026, estimate the final indexed tokens, average stored index size and monthly query mix for your account.
Cloudflare announced general availability on October 1. AI Search provides built-in file storage and indexing, with optional website or R2 sources. For a collection made mostly of Markdown, code, notes, or small text PDFs, that can remove the need to run a separate parser, vector index, and retrieval service. A mixed collection still needs a file-by-file check before you upload everything. Cloudflare’s GA announcement and its built-in storage documentation describe the managed path. If your application also needs custom runtime logic, our Python Workers compatibility guide covers that layer while AI Search handles managed indexing and retrieval.
Start with extensions and file sizes
The size limit depends on the format and whether OCR is enabled. This matters because the launch post describes larger text files broadly, while Cloudflare’s detailed format table and release note make narrower distinctions. For a go/no-go decision, use the detailed per-format limits and test files close to the boundary.
| Collection item | Documented limit | Fit check |
|---|---|---|
| Plain text, Markdown, code, JSON and the other listed plain-text formats | 10 MiB per file | Direct fit if your extension is in the plain-text list. |
| PDF without OCR | 4 MiB per file | Suitable for smaller PDFs with extractable text. |
| PDF with OCR enabled | 10 MiB per file | Use when scanned pages need searchable text; account for OCR processing and reindexing. |
| HTML, CSV, images and other supported formats | 4 MiB per file in the “other formats” limit group | Check the extension and test larger files. Images still need an image-capable model for direct visual matching. |
The distinction comes from Cloudflare’s supported-format table and its October 1 release note: plain-text and code files, plus OCR-enabled PDFs, can be 10 MiB; PDFs without OCR and other supported formats are limited to 4 MiB. The announcement also names HTML and CSV in its broader 10 MiB sentence. Since the detailed table classifies HTML and CSV as formats converted to Markdown, treat 4 MiB as the safer operating limit for them until Cloudflare clarifies the difference.
For example, the documented caps put an 8 MiB Markdown file below its 10 MiB limit, a 6 MiB scanned PDF below the OCR-enabled PDF cap, and a 6 MiB CSV above the 4 MiB cap. This is a source-based file-fit check, not a live upload result. New instances default to hybrid search. Cloudflare lists 100,000 files per instance on Workers Free, and up to 1 million on Paid plans or 500,000 for hybrid search. Check the applicable plan and file limits and search-mode setting against your collection size.
Files above a limit are not indexed and appear in the error logs. Converting a rich file to Markdown may bring it into the 10 MiB plain-text path, but that changes what enters the index. Inspect the converted text and keep a copy of the source if layout or metadata matters. The docs confirm conversion to Markdown; they do not promise that every source layout survives conversion unchanged. If you plan a separate extraction stage before indexing, our Cohere Parse guide explains that document-parsing workflow.
Images need a model choice; scanned PDFs need OCR
AI Search supports two image paths. With a text-only embedding model, it creates captions for images and image queries. With a multimodal embedding model, it embeds image content directly, so retrieval can use visual details that a caption might omit. That difference matters for product texture, screenshot states, diagrams, charts, or similar-image search. Cloudflare’s supported-model table lists the default @cf/qwen/qwen3-embedding-0.6b as text-only and @cf/qwen/qwen3-vl-embedding-2b as image-capable. It also lists Google’s gemini-embedding-2 with image support.
If visual matching is central, choose an image-capable embedding model when creating the instance. Cloudflare’s model settings fix that embedding choice at creation, so changing it later means creating a new instance and reindexing the collection. Third-party embeddings can be connected through AI Gateway, but provider charges are separate from the included Workers AI embedding and reranking usage. For a separate comparison of embedding options and direct inference costs, see our Cohere Embed 5 comparison.
Image queries are available through the REST API. This unexecuted request follows Cloudflare’s documented image-query shape. First create and index an instance, then replace the account, instance, and token placeholders with your values. Supply an HTTPS image URL the API can reach, or use the documented inline image-data form. The REST API requires an API token with AI Search:Edit and AI Search:Run permissions.
curl -X POST \
"https://api.cloudflare.com/client/v4/accounts/<ACCOUNT_ID>/ai-search/namespaces/default/instances/<INSTANCE_NAME>/search" \
-H "Authorization: Bearer <API_TOKEN>" \
-H "Content-Type: application/json" \
-d '{
"messages": [{
"role": "user",
"content": [{
"type": "image_url",
"image_url": { "url": "https://example.com/reference-product.jpg" }
}]
}],
"ai_search_options": {
"query_rewrite": { "enabled": false }
}
}'
With a text-only model, the image is captioned before search. With an image-capable model, retrieval can use visual content directly. The REST Search API returns ranked content chunks with source-item details, not a generated answer; use chat completions when you want a generated response.
For existing AutoRAG users, Cloudflare says legacy routes continue to work, while new features use the AI Search REST routes. The Workers binding documentation is inconsistent: a notice says message content must be a string, while its parameter table permits arrays. Use REST or the public endpoint for image and file parts until that notice is clarified; see the Workers binding guide.
OCR solves a different problem. It extracts searchable text from scanned PDFs and images; it is disabled by default. Enable indexing_options.use_ocr when creating or updating the instance. Changing the setting triggers a full reindex, so enable it before indexing a collection if scanned pages are part of the intended search experience. Cloudflare says OCR is available on every account and meters its extracted text as image-processing tokens. The data-source docs list the supported formats and OCR behavior.
Price the monthly meters separately
Cloudflare’s pricing page, checked October 4, 2026, gives every account monthly included usage. The allowances are account-wide, not one fresh pool for each instance. Semantic, vector, and hybrid calls share one query allowance; full-text calls have a separate one.
| Meter | Included each month | Overage rate |
|---|---|---|
| Base ingestion | 5 million tokens | $0.75 per million tokens |
| Image processing add-on | Uses the shared 5 million ingestion-token allowance | Additional $0.50 per million image-processing tokens |
| Stored index data | 10 GB-month | $2 per GB-month |
| Semantic, vector, and hybrid queries | 1,000 queries | $0.75 per 1,000 queries |
| Full-text queries | 1,000 queries | $0.10 per 1,000 queries |
AI Search counts ingestion on the final chunks after parsing and chunking. Each chunk uses the cl100k_base tokenizer, and overlapping text counts again wherever it appears. Estimate from the text that will actually be indexed, not the original file size; the chunking guide and pricing page document these billing units.
For images and scanned files using OCR, extracted text counts as base ingestion and as image-processing tokens. Cloudflare says image-processing tokens share the 5 million token ingestion allowance, so do not budget a separate free 5 million for images. The accessed pricing page does not give a sample allocation formula for every mixed workload. Keep those lines separate in a projection and verify the account meter before making a budget commitment. Once the shared allowance is used, the published rates add to $1.25 per million OCR text tokens: $0.75 base plus the $0.50 image-processing add-on.
Storage is billed by GB-month. Cloudflare’s Stats API reports indexing status, payload and metadata bytes, and vector counts, but the accessed Stats API and pricing docs do not define a formula converting those fields to billed GB. After billing data is available, the billable usage dashboard reports total and billable usage plus cost; Cloudflare says completed-period totals match the invoice. The dashboard requires a Pay-as-you-go account.
Managed Workers AI embedding and reranking are included in AI Search pricing. Answer generation, query rewriting and external model providers are billed through their applicable service; the pricing documentation separates these charges from the AI Search meters.
Worked text-only billing month
Here is one hypothetical billing month after November 1. The numbers are account-wide totals across all instances, with the listed allowances fully available and no other usage consuming them. It assumes indexing in that billing month. October activity is not included, and this example makes no inference about how Cloudflare treats pre-November activity.
- 8 million account-wide final indexed text tokens, including chunk overlap.
- 12 GB-month of average billable stored index data across the account.
- 1,800 account-wide semantic or hybrid queries and 800 account-wide full-text queries.
- No image or OCR ingestion or third-party provider use; query rewriting is disabled with
ai_search_options.query_rewrite.enabled=false, and no answers are generated.
| Line item | Arithmetic | Cost |
|---|---|---|
| Base ingestion | (8M – 5M included) x $0.75 per million | $2.25 |
| Storage | (12 – 10 included) GB-month x $2 | $4.00 |
| Semantic or hybrid queries | (1,800 – 1,000 included) x $0.75 per 1,000 | $0.60 |
| Full-text queries | 800, within its separate 1,000-query allowance | $0.00 |
| Total for the assumed month | AI Search meters only | $6.85 |
This is a worksheet, not an account quote. Replace the account-wide token, average billable storage, and query totals with your own readings. The example applies each included allowance once to the whole account; it leaves separate answer-generation and third-party provider charges out of the total.
What to capture in an October pilot
Use files and search questions from the collection you plan to index, including examples near its format caps, representative images and scanned pages if present. Record which items index, any parse errors, whether the top returned chunks help answer the intended search, final indexed-token estimates and the expected query-mode split. After billing data is available, compare actual usage with the billable usage dashboard. Recheck Cloudflare’s pricing page before budgeting a later billing month.
AI disclosure: AI tools supported the research and drafting of this article. It does not report a hands-on AI Search test.