
How Much Does a RAG Chatbot Cost?
A demo costs a weekend. A version you can put in front of customers costs considerably more, and almost none of the difference is the model.

People estimating an AI build usually price the part they can see: calling a model. That is the cheapest line on the invoice, and it falls in price every year.
What costs money is everything required to make the answers trustworthy enough to show a customer. A demo takes a weekend. A product takes the rest.
| Line | Share of a typical AI build | Why |
|---|---|---|
| Getting the data usable | Often the largest | Documents, permissions, formats nobody standardised |
| Retrieval and context | Large | Deciding what the model sees is most of answer quality |
| Evaluation | Underestimated by nearly everyone | Without it you cannot tell whether a change helped |
| The product around it | Normal software cost | Auth, UI, billing, admin — unchanged by AI |
| The model call itself | Smallest | A commodity, and getting cheaper |
The first row is the one that surprises people, and it is where projects stall. If the answers must come from your own documents and nobody has ever organised those documents, the first phase of the project is data work with no visible AI in it at all.
Ranges across the market for a first production version, in USD. These describe what buyers are quoted generally rather than any one firm's prices:
Normal software costs roughly the same to run whether a customer uses it once a month or hourly. Inference does not: it scales with usage, so one enthusiastic customer can cost more than a hundred ordinary ones, and your gross margin becomes a function of behaviour you do not control.
Current rates are published by the vendors — OpenAI and Anthropic both list them — and they fall regularly, so build the estimate from your own expected token counts rather than from a figure you read once. What it costs to run a SaaS covers the rest of the monthly bill.
Two controls belong in the build rather than being retrofitted: a hard per-account cap so no single user can run up an unbounded bill, and caching for anything asked repeatedly. Both are cheap while the system is being designed and awkward afterwards.
We scope and quote each AI integration and RAG build individually, because the five questions above move the number far more than the feature description does. You get a written scope, a timeline and a fixed price before any work starts.
On the first call we will tell you if your problem is a data problem, a scoping problem, or not an AI problem — all three are common, and hearing it before you commit a budget is worth more than a quote. Adding AI to an existing SaaS covers choosing the job worth giving a model at all.
Across the market, roughly $10k–$30k for a single AI feature inside an existing product, $25k–$65k for an assistant answering from your own content with permissions and citations, and $60k–$150k or more for an AI-native product built from nothing — because that includes a full MVP as well as the AI engineering. Regulated domains start higher.
It usually is not the AI that costs more. The expensive parts are getting your data into a usable state, building retrieval that returns the right context, and creating an evaluation set so you can tell whether a change improved answers. The model call itself is the cheapest line and gets cheaper every year.
Inference is a variable cost that scales with usage rather than with user count, so one heavy customer can cost more than a hundred ordinary ones. Build the estimate from your own expected token counts using the vendors' published rates. Two controls belong in the build itself: a hard per-account cap, and caching for anything asked repeatedly.
Most often data that was not as ready as assumed. Other common causes are discovering that different users may see different documents, which moves permission filtering into the retrieval query, and deciding late that the model should be able to act rather than only answer — which adds confirmation steps, permission scoping and audit trails.