AI Agent 月度成本怎么算:token 账单公式与六款模型实测
AI Agent Cost per Month: The Token Maths, Worked Out
作者给出 AI Agent 月度运行成本公式,并以小型企业客服机器人为例算出 3,000 次对话在六款模型上花费 $17–$342/月(未含缓存)。核心机制是每轮重发全部历史,历史成本随轮数平方增长,示例中占模型账单 44%;开启 prompt caching 后 Claude Sonnet 5.5 账单降 59%。
A four-message support chat with a tool-using AI agent can send the model more than 45,000 input tokens. That's about twelve times Anthropic's own estimate of roughly 3,700 tokens for a support conversation (Anthropic pricing). Nothing is broken: each turn simply resends everything that came before it.
That's why most guides to AI agent cost per month give you wide ranges instead of a number. This one gives you a formula, worked through a small-business example and priced against six current models using each provider's official pricing page.
TL;DR
- Separate build cost from run cost. This post covers only run cost: model tokens, tools, storage, monitoring and the people who handle escalations.
- Conversation history is the biggest cost. In our example, resending earlier turns accounts for 44% of the model bill. Because each turn resends all the earlier ones, history cost grows with the square of the number of turns.
- 3,000 support conversations a month cost $17 to $342 in model fees across six current models, before caching. Turning on prompt caching cut the Sonnet 5.5 bill by 59%.
Why "AI agent cost" articles never add up
Most articles on this topic mix two different bills:
- Build cost: design, integration, prompt work, testing. You pay it once, and it's usually much larger.
- Run cost: what you pay every month after launch. It grows with usage and never goes away.
One top-ranking article says plainly that "the published per-token price tells you almost nothing about what an agent will actually spend" (CloudZero). True, but behaviour can be measured — and anything measurable belongs in a formula.
Not sure you need an agent at all? Read Zapier, n8n or an AI agent — often one AI step inside a normal workflow is enough.
The real cost drivers
Four things set your monthly bill:
- Conversations per month. This is your volume.
- Model calls per conversation. A tool-using agent usually makes at least two calls per customer message: one to pick a tool, one to write the answer.
- Input tokens per call. Every call resends the system prompt, the tool definitions, the conversation so far and any retrieved documents. Tools add hidden tokens too. Anthropic adds a tool-use system prompt (496 tokens on Haiku 4.5, 286 on Sonnet 5.5) on top of your tool schemas (Anthropic pricing).
- Loops and retries. These are failed tool calls, re-asked questions and agents that check their own work. Each one is another full-price call.
One more catch: the same text produces different token counts on different models. Anthropic says the tokenizer in its newer models "produces approximately 30% more tokens for the same text" (Anthropic pricing). Count tokens on the model you'll actually use.
The formula
Monthly run cost is conversations times per-conversation token cost, plus retry overhead, plus fixed costs. Here it is in full:
Monthly run cost = N × (I × p_in + O × p_out) / 1,000,000 × (1 + r) + fixed costs
N = conversations per month
I = input tokens per conversation (summed over every model call)
O = output tokens per conversation
p_in = price per million input tokens; p_out = price per million output tokens
r = retry/loop overhead (0.10 = 10% extra calls)
fixed = vector DB, monitoring, hosting, search fees, human fallback
The hard part is calculating I correctly. Each call's input is fixed prefix + history so far + new content. With T turns, k calls per turn and h tokens added to the history each turn:
History tokens per conversation = k × h × T(T − 1) / 2
The T(T − 1) term is the important part. Going from 4 turns to 8 turns multiplies history tokens by 4.7 (from 6 to 28 in the T(T − 1)/2 term), not by 2.
Here's the same calculation as a small pricing calculator you can edit:
def agent_cost(convs, turns, prefix, user, tool_call, tool_result, answer,
price_in, price_out, overhead=0.10):
history, tokens_in, tokens_out = 0, 0, 0
for _ in range(turns):
tokens_in += prefix + history + user # call 1: pick a tool
tokens_in += prefix + history + user + tool_call + tool_result # call 2: answer
tokens_out += tool_call + answer
history += user + tool_call + tool_result + answer
per_conv = (tokens_in * price_in + tokens_out * price_out) / 1e6
return per_conv * (1 + overhead) * convs
print(agent_cost(3000, 4, 2000, 100, 50, 1500, 250, 1.00, 5.00)) # ≈ 170.94
Worked example: AI agent cost per month for a small-business support bot
This example prices a small-business support bot at $170.94 a month on Claude Haiku 4.5, with every step shown so you can swap in your own numbers.
The setup (these are assumptions, so swap in your own): a support agent that searches a knowledge base and answers customers.
| Assumption | Value |
|---|---|
| Conversations per month (N) | 3,000 (about 100 a day) |
| Customer messages per conversation (T) | 4 |
| Model calls per turn (k) | 2 (choose tool → answer) |
| System prompt + tool definitions | 2,000 tokens |
| Customer message / tool call / retrieved docs / answer | 100 / 50 / 1,500 / 250 tokens |
| Retry overhead (r) | 10% |
Step 1: tokens added to history per turn. 100 + 50 + 1,500 + 250 = 1,900. History at the start of turns 1–4 is 0, 1,900, 3,800 and 5,700 tokens.
Step 2: input per turn.
- Call 1 = 2,000 + history + 100 = 2,100 + history
- Call 2 = 2,000 + history + 100 + 50 + 1,500 = 3,650 + history
- Per turn = 5,750 + 2 × history
Step 3: input per conversation. 4 × 5,750 = 23,000. History adds 2 × (0 + 1,900 + 3,800 + 5,700) = 22,800. That matches the formula: 2 × 1,900 × 6 = 22,800. Total I = 45,800 tokens.
Step 4: output per conversation. (50 + 250) × 4 = O = 1,200 tokens.
Step 5: monthly tokens, with 10% overhead.
- Input: 45,800 × 1.1 × 3,000 = 151.14M
- Output: 1,200 × 1.1 × 3,000 = 3.96M
Step 6: price it. On Claude Haiku 4.5 ($1 input / $5 output per million tokens): 151.14 × $1 + 3.96 × $5 = $151.14 + $19.80 = $170.94 a month, or about $0.057 per conversation. The Python above gives the same figure.
Notice the ratio: the agent reads 38 tokens for every token it writes. For tool-using agents, the input price matters far more than the output price.
Provider price comparison
Across six current models, the same worked example costs $17 to $342 a month in model fees alone. Prices as of 3 October 2026, taken from each provider's official pricing page. Re-check them before relying on this; they change often. Monthly cost uses the worked example (151.14M input, 3.96M output) with no caching and no batching.
| Model | Input / Output per 1M tokens | Cached input per 1M | Example monthly cost |
|---|---|---|---|
| OpenAI gpt-6-luna | $0.10 / $0.50 | $0.01 | $17.09 |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | $0.03 + storage | $55.24 |
| Claude Haiku 4.5 | $1 / $5 | $0.10 | $170.94 |
| Gemini 3.5 Flash | $1.50 / $9 | $0.15 + storage | $262.35 |
| Claude Sonnet 5.5 | $2 / $10 | $0.20 | $341.88 |
| OpenAI gpt-6.1-sol | $2 / $10 | $0.10 | $341.88 |
Sources: Anthropic, OpenAI (short-context rates, up to 272K input tokens), Google (Gemini caching also charges $1 per million tokens per hour of storage).
Prompt caching is the biggest discount you can control. On Anthropic, a cache read costs 0.1× the base input price and a 5-minute cache write costs 1.25× (on Sonnet 5.5, that's $0.20 and $2.50 per million). With automatic caching, each new request reads everything up to the previous message from cache (Anthropic prompt caching). In our example, the eight calls add up to 45,800 input tokens. Of these, 36,450 are cache reads and 9,350 are writes:
- Input: 36,450 × $0.20 + 9,350 × $2.50 = $0.0307 per conversation (down from $0.0916)
- Add output ($0.012), the 10% overhead and the 3,000 conversations: $140.79 a month instead of $341.88, a 59% cut.
Two catches: the default cache lasts 5 minutes (a 1-hour cache costs 2× to write), so a quiet customer breaks the chain; and Haiku 4.5 only caches prompts of at least 4,096 tokens, so the early calls here wouldn't be cached (Anthropic prompt caching).
Batch APIs take 50% off at Anthropic, OpenAI and Google, but only for jobs that can wait, like nightly ticket summaries. They're no use for live chat.
Watch for introductory prices. Gemini 3.8 Flash is $0.75 / $3.75 until 31 December 2026, then $1.50 / $7.50 from 1 January 2027 (Google). If you budget at the introductory price, your bill will double when it ends.
Hidden costs beyond tokens
The model fee is only part of the cost of running an AI chatbot. The rest:
- Vector database. Pinecone's free Starter plan covers up to 2 GB of storage, 2M write units and 1M read units a month. Its Standard plan has a $50/month minimum, then $0.33/GB/month for storage and $16–$18 per million read units (Pinecone). OpenAI's hosted file search charges $0.10/GB per day after the first free GB (OpenAI).
- Monitoring and logging. You need traces to find loops. Langfuse Cloud is free for 50k units a month, and its Core plan is $29/month for 100k units plus $8 per extra 100k. You can also self-host it for free (Langfuse). Our example makes about 24,000 model calls a month before retries, so check how your tool counts units.
- Search and hosted runtimes. Web search costs $10 per 1,000 searches on Anthropic and $10 per 1,000 calls on OpenAI. Google includes 5,000 free searches a month across Gemini 3.x, then charges $14 per 1,000. Anthropic's Managed Agents add $0.08 per session-hour of runtime.
- Human fallback and app hosting. Escalated conversations cost your team's time, far more than the tokens. The server running the agent loop is usually a small cost, but never zero.
For comparison, Intercom's Fin charges $0.99 per resolved conversation (Intercom) — not like-for-like, since that price includes the product itself, but a useful ceiling.
When self-hosting an open model beats per-token pricing
Running an open model on a GPU you rent turns a per-token bill into a per-hour one. On Lambda, one NVIDIA A10 (24 GB) is $1.29/hour and one H100 PCIe is $3.29/hour (Lambda). Running all month (about 730 hours), that's about $942 and $2,402.
Our whole example costs $17–$342 a month through an API. At small-business volume, self-hosting costs more, often several times more. It starts to make sense when:
- the GPU is busy most of the day, because your volume is high enough to use the capacity you pay for
- your data can't leave your own infrastructure
- you need a fine-tuned or specialised open model
To find your break-even point, benchmark how many tokens per second the GPU handles on your workload, then add the engineering time for patching, scaling and uptime.
Quick self-check: is your spend normal or bloated?
Pull a week of logs and check these:
- Input-to-output ratio. A high ratio is normal for tool-using agents (ours is 38:1). A sudden rise usually means history or retrieval is growing unchecked.
- Calls per turn. If you designed for 2 and see 5, the agent is looping. Set a hard cap on steps.
- Cost by turn count. Long conversations cost far more than short ones. Summarise or trim history after a few turns.
-
Cache hit rate. The API's usage fields (for example,
cache_read_input_tokens) show whether caching is working. - Model fit. Are you using a top-tier model for questions a small model can answer? Send easy questions to a cheaper model.
- Most expensive 5% of conversations. Read them. Bugs and abuse usually show up there.
Want this run against your own numbers?
If you want this formula applied to your real logs, or an agent built with budget caps and caching from the start, that's part of our AI integration & agents work. Send us your volumes and we'll tell you what it should cost to run — or if an agent isn't worth it at all.
FAQ
How much does an AI agent cost per month for a small business?
In our worked example, 3,000 support conversations cost $17–$342 a month in model fees across six current models, before caching. On top of that come vector storage, monitoring, search and your team's time on escalations.
Why is my AI chatbot's running cost higher than the per-token price suggests?
Every call resends the system prompt, tool definitions and the conversation so far. History cost grows with the square of the number of turns, and retries add more full-price calls.
Does prompt caching really reduce LLM API costs?
Usually. In our example it cut the Sonnet 5.5 bill by 59%. Check your model's minimum cacheable prompt length and how long the cache lasts.
Is self-hosting an open-source model cheaper?
Rarely at small-business volume. One rented A10 GPU running all month costs about $942, which is more than our example's highest API bill.
What does your agent's input-to-output ratio look like, and which cost driver surprised you most? Tell us in the comments.
Originally published on banxal.com.
来源:Google AI:DEV 作者专属(RSS) · dev.to


