ChatGPT API cost: what an AI feature costs to run
What the ChatGPT, Claude and Gemini APIs cost to run in a real feature: how token pricing works, a worked monthly example, and how to keep the bill in check.
By Ali HassanSenior engineer for SaaS and AI products
The ChatGPT API cost of a feature is the price of the tokens it moves, not a flat monthly fee: you pay per token sent in and per token generated out, quoted per million tokens. The same billing model applies to Claude and Gemini, so the sums here carry across all three. A small support feature might cost a few hundred dollars a month to run; the same feature on a top-tier model, or with a bloated prompt, can cost several times that.
How ChatGPT API cost is charged
Every request has two prices: the tokens you send (input) and the tokens the model writes back (output). A token is roughly three-quarters of a word. Output almost always costs several times more than input, so a model that returns long answers costs more than the input figures alone suggest.
The same structure holds for Anthropic's Claude and Google's Gemini. Only the per-million rates differ. Here are the current standard-tier rates, per 1 million tokens:
| Model | Input | Output |
|---|---|---|
| OpenAI GPT-6 Astra | $10.00 | $50.00 |
| OpenAI GPT-6.1 Sol | $2.00 | $10.00 |
| OpenAI GPT-6 Luna | $0.10 | $0.50 |
| Anthropic Claude Opus 5.5 | $4.00 | $20.00 |
| Anthropic Claude Sonnet 5.5 | $2.00 | $10.00 |
| Anthropic Claude Haiku 5.5 | $0.10 | $0.50 |
| Google Gemini 3.1 Pro (preview) | $2.00 | $12.00 |
| Google Gemini 3.8 Flash | $0.75 | $3.75 |
Prices as of 8 October 2026, from each vendor's pricing page; they change, so check before you budget.
A few caveats on those numbers. The OpenAI prices are for prompts up to 272K tokens, Gemini 3.1 Pro's are for prompts up to 200K tokens, and Claude Haiku 5.5's are for prompts up to 100K tokens; go over those sizes and the rate steps up. Gemini 3.8 Flash rises to $1.50 / $7.50 per million from 1 January 2027.
What actually drives the bill
The headline rate is the smallest part of the story. What you pay per request depends on how much you feed the model and how many times you call it.
- The prompt and system prompt. Every instruction, example and formatting rule you prepend is charged on every single request, forever.
- Retrieved chunks. In a RAG feature, the passages you pull from your knowledge base are added to the prompt each time. Ten chunks of 300 words each is a few thousand input tokens before the user has typed anything.
- Conversation history. A chat feature usually resends the whole thread on each turn, so a long conversation gets dearer with every message.
- Output length. Output tokens cost the most, so a feature that writes essays costs far more than one that returns a short answer or a structured result.
- Retries. A failed or malformed call that you retry is billed twice.
- Agent loops. An agent that calls the model to plan, act and check can run ten or more calls for one user task, each carrying the full context again.
That last one catches people out. A single user action can be dozens of model calls under the surface, and the bill is the sum of all of them.
A worked example
Take a support feature that sends 2,000 input tokens (system prompt, a few retrieved chunks, the user's question) and gets back 400 output tokens, across 100,000 requests a month.
That is 200 million input tokens (2,000 × 100,000) and 40 million output tokens (400 × 100,000) a month.
On a mid-tier $2 / $10 model, such as GPT-6.1 Sol or Claude Sonnet 5.5:
- Input: 200M × $2 per million = $400
- Output: 40M × $10 per million = $400
- Total: $800 a month
Same traffic on a small model, GPT-6 Luna or Claude Haiku 5.5 at $0.10 / $0.50:
- Input: 200M × $0.10 per million = $20
- Output: 40M × $0.50 per million = $20
- Total: $40 a month
Same traffic on a large model, GPT-6 Astra at $10 / $50:
- Input: 200M × $10 per million = $2,000
- Output: 40M × $50 per million = $2,000
- Total: $4,000 a month
Notice that input is five times the token volume of output here, yet costs the same, because output is priced five times higher per token. That is why trimming the prompt and capping the output both matter: on these numbers, neither half of the bill is small enough to ignore.
Same feature, same traffic: $40 to $4,000 depending only on which model answers. That spread is why model choice is the first lever, not the last. And it scales with volume, so a feature that is cheap in a demo can turn into a real monthly line once it has users. You can run your own numbers in the LLM cost calculator.
Keeping the bill down
A handful of levers move the number, roughly in order of effort.
- Prompt caching. If your system prompt and retrieved context are stable across requests, caching them means you are not paying full price to resend the same tokens every time. All three vendors offer it.
- A smaller model for simple steps. Not every step needs the top model. Sending classification, extraction and routine replies to a cheaper model, and keeping the expensive one for the hard cases, is often the biggest saving.
- Trim the context. Fewer, better retrieved chunks beat ten mediocre ones. Shorter system prompts, and summarised history instead of the full thread.
- Cap the output. Set a maximum on output tokens and ask for short, structured responses. Output is the dear half of the bill.
- Meter usage per customer. Track tokens per customer or per workspace, so you know who is expensive and can price accordingly.
- Set spend caps and alerts. A monthly ceiling, and a warning as you approach it, turns a runaway bill into an email rather than a surprise.
I lean on the last two in my own work. On Evoriqa, my multi-tenant support platform, every AI action (chat, voice, actions and knowledge ingestion) draws from one credit balance, so a workspace can see exactly what its AI use costs. A workspace can set a monthly spend cap, and owners are emailed as usage crosses thresholds.
Measure before you optimise
Every one of those levers is a guess until you have measured. You cannot tell whether caching or a smaller model is worth the work until you know what a single request costs today and which part of it is expensive.
On an AI audit reviewer I built for a CPA firm, input and output tokens were logged on every model call, which is how the cost landed at a measured $0.24–0.28 per review on sample documents rather than an estimate. Once you have a real per-request number, you know which lever is worth pulling and which would save pennies.
When to bring in help
If your model bill is climbing faster than your usage, or you cannot say what a single request costs, that is worth measuring before you start cutting. An LLM cost and RAG quality audit does exactly that: a measured cost per request, and the changes that move it most. Otherwise the levers above and the calculator will get you a long way on your own.


