AI
What it costs to run a retrieval-augmented AI system on Cloudflare
The honest cost anatomy of a RAG system on Cloudflare — embeddings, Vectorize, model inference, storage, and the request path — plus how the bill scales and where it surprises people. A method to estimate yours, not a fake total.
“How much does it cost to run?” is the question that decides whether a retrieval-augmented AI feature ships or stays a demo. It is also the question most vendors dodge, because the honest answer is “it depends on your traffic and your choices” — which is true, and useless without a way to reason about it. So here is the way to reason about it.
A retrieval-augmented generation system — RAG — answers questions by pulling relevant chunks from your own content and handing them to a language model as context. On Cloudflare, the whole pipeline can run on Workers, Vectorize, and either Workers AI or an external model. This piece breaks the running cost into its parts, shows which part dominates at which scale, and names the places the bill tends to surprise people. We will not hand you a made-up monthly total; we will hand you the arithmetic so you can build your own. If you want the code that produces these costs, we wrote that separately in production RAG on Cloudflare without LangChain.
The two cost phases: ingest and query
The first clarifying move is to separate the two things a RAG system does, because they bill completely differently.
Ingest happens when you load or update your content: chunk each document, embed every chunk, and store the vectors. It is roughly a one-time cost per document, paid again only when that document changes. A corpus of ten thousand documents is a large ingest bill once and almost nothing thereafter.
Query happens every time a user asks a question: embed the query, search the vector store, and call a model to generate the answer. This scales with usage, forever. It is the line that grows with success.
Most people worry about the wrong one. Ingest feels big because you do it all at once, but query is what determines whether the feature is affordable at ten thousand users a month. Model your query cost first.
The five line items
Every RAG bill is some combination of five things. Specific prices move, so treat any figure below as a shape to verify against current Cloudflare and model-provider pricing before you plan around it — not a quote.
1. Embeddings. Turning text into vectors. You pay for it twice conceptually: once per chunk at ingest, and once per query at ask-time. On Cloudflare, Workers AI offers embedding models billed by usage, with a free daily allowance that covers low volume; external providers like OpenAI charge per token. Embeddings are usually cheap per call. The place they add up is a large ingest — embedding an entire corpus — or re-embedding everything when you change models.
2. Vector store (Vectorize). Cloudflare’s vector database charges along two axes: storage of the vectors you keep, and the queries you run against them. Storage is billed by the number of vectors and their dimensions; querying is billed by how many vectors are scanned. For a small-to-mid corpus this is typically one of the smaller lines. It grows if your corpus is very large, if you use high-dimension embeddings, or if every user query fans out into many sub-queries.
3. Model inference. The language model that writes the answer. This is almost always the line that dominates a query-heavy system, and it is where your architectural choices show up most. Workers AI runs open models billed by usage with a free daily allowance; external models like Claude or GPT are billed per input and output token, and are more capable but cost more per call. The two dials that move this number are which model you choose and how much context you stuff into each call — a bloated prompt with fifteen retrieved chunks costs multiples of a tight one with four.
4. Storage of source content. Vectorize holds vectors, not your actual text. The chunk text and source documents live in D1 (structured) or R2 (files and large blobs). These are inexpensive at the scale most RAG systems run, and R2’s zero-egress model means serving that content back does not add a transfer bill. Storage is rarely the line that hurts.
5. The request path — Workers. The compute that runs the pipeline: chunking, orchestration, hydrating text, assembling the prompt. Workers are billed by requests and CPU time, and for the light glue work of a RAG pipeline this is typically the smallest line on the bill. The expensive work is the model call the Worker makes, not the Worker itself.
How the costs scale
Put the pieces on a timeline and the picture clarifies:
- At low volume, most of the pipeline fits inside free or near-free allowances. Your real cost is a handful of external model calls, if you use an external model at all. A pilot can genuinely cost very little to run.
- As queries grow, model inference pulls away from the pack and becomes the number that matters. Everything else stays in the noise. This is the regime where “which model, how much context” decides your bill.
- As the corpus grows, ingest embedding cost and Vectorize storage rise, but these are gentle and mostly one-time. A big corpus with modest traffic is cheap to run; a small corpus with huge traffic is not.
The single most useful sentence: your monthly cost is approximately your query volume times your per-query model cost. Estimate those two numbers honestly and you have your answer to within the margin that matters. Everything else is a rounding adjustment on top.
A method to estimate yours
You can get a defensible figure in four steps.
- Estimate queries per month. Be realistic, then double it for the month the feature actually gets used.
- Price one query. Take a representative question, run it through your chosen model with a realistic amount of retrieved context, and note the input and output tokens. Multiply by the model’s per-token rate. This one measured number is worth more than any calculator.
- Multiply. Queries per month times cost per query is the line that dominates.
- Add the rest. Estimate ingest embeddings for your corpus size (mostly one-time), Vectorize storage and query, and a small allowance for Workers and content storage. These are the correction terms, not the headline.
Do this before you build, and again with real numbers once you have a week of traffic. The gap between the two is usually smaller than people fear, provided you measured step two honestly.
Where the bill surprises people
- Over-stuffed context. Retrieving fifteen chunks “to be safe” when four would answer the question multiplies your input tokens on every single query. This is the most common source of a bigger-than-expected inference bill, and the easiest to fix.
- Query fan-out. Techniques that rewrite one user question into several retrieval queries improve answer quality and multiply your embedding and Vectorize costs per ask. Worth it, often — but budget for it.
- Re-embedding the whole corpus. Changing your embedding model means re-embedding everything, which is a fresh ingest bill. Choose deliberately so you do it rarely.
- Retries and failures. A model call that times out and retries is billed work with no answer to show for it. At scale this is a real, if small, tax.
- Assuming external and local models are interchangeable on price. They are not. Moving inference from an open model on Workers AI to an external frontier model can change the query cost by an order of magnitude in either direction. It is a deliberate quality-versus-cost decision, not a detail.
The reassuring counterpoint: none of these are hidden once you know to look. A RAG system’s cost is legible if you built it with the five line items visible, which is one more reason we keep the pipeline plain rather than wrapping it in a framework that obscures where the money goes. The broader case for building on Cloudflare — one platform, one bill, one observability surface — is on our Cloudflare page.
If you are weighing a retrieval-augmented feature and want a straight, written estimate for your corpus and your expected traffic before you commit engineering time, that is exactly the scoping we do on our RAG development work. Send us two paragraphs about what you want it to answer and roughly how often, and we will reply within one business day.