Token cost benchmark for an autonomous Accounts Payable agent, across 13 models. Prices as of 14 Jun 2026.
| Model | $/1M in | $/1M out | Cost / outcome | Cost / month* |
|---|---|---|---|---|
| GPT-4o mini | $0.15 | $0.60 | $0.0602 | $602 |
| Llama 4 Maverick | $0.27 | $0.85 | $0.105 | $1,054 |
| Gemini 2.5 Flash | $0.30 | $2.50 | $0.137 | $1,367 |
| GPT-4.1 mini | $0.40 | $1.60 | $0.160 | $1,605 |
| DeepSeek V4 | $0.44 | $0.87 | $0.165 | $1,653 |
| Claude Haiku 4.5 | $1.00 | $5.00 | $0.414 | $4,138 |
| Gemini 2.5 Pro | $1.25 | $10.00 | $0.565 | $5,645 |
| Mistral Large 3 | $2.00 | $6.00 | $0.777 | $7,772 |
| GPT-4.1 | $2.00 | $8.00 | $0.802 | $8,024 |
| GPT-4o | $2.50 | $10.00 | $1.00 | $10,030 |
| Claude Sonnet 4.6 | $3.00 | $15.00 | $1.24 | $12,414 |
| Claude Opus 4.8 | $5.00 | $25.00 | $2.07 | $20,690 |
| Claude Fable 5 | $10.00 | $50.00 | $4.14 | $41,380 |
*At 10,000 outcomes per month. Cheapest model highlighted.
The clean-path steps this benchmark prices:
This path runs 17 steps: 4 tool calls, 3 reasoning steps, 10 decision points and 0 human checkpoints. Tool steps make two model calls each, and the agent re-reads its growing context on every call. That compounding is why one Accounts Payable outcome costs about 53x a single chat message ($1.24 on Claude Sonnet 4.6), not the price of one message.
On the clean path with default assumptions, an agent for Accounts Payable costs about $0.0602 to $4.14 per outcome depending on the model, or roughly $602 to $41,380 per month at 10,000 outcomes. The cheapest model here is GPT-4o mini at $0.0602; the most expensive is Claude Fable 5 at $4.14.
An agent does not make one model call. It plans, calls tools, retrieves context and re-reads its growing working context on every step. For Accounts Payable that adds up to about 53x the cost of a single chat message.
Across the 13 models benchmarked, GPT-4o mini is cheapest at $0.0602 per outcome and Claude Fable 5 is the most expensive at $4.14. A cheaper model is not always the right choice, but it sets the floor for this workflow.
The biggest levers are prompt caching on the base context, fewer planning loops, smaller tool results, less retrieval, and choosing a cheaper model where quality allows. You can test each lever in the live estimator.