Token cost benchmark for an autonomous Offboarding agent, across 26 models. Prices as of 26 Jul 2026.
| Model | $/1M in | $/1M out | Cost / outcome | Cost / month* |
|---|---|---|---|---|
| GPT-4o mini | $0.15 | $0.60 | $0.0291 | $291 |
| Llama 4 Maverick | $0.27 | $0.85 | $0.0507 | $507 |
| Gemini 2.5 Flash | $0.30 | $2.50 | $0.0676 | $676 |
| GPT-4.1 mini | $0.40 | $1.60 | $0.0776 | $776 |
| DeepSeek V4 | $0.43 | $0.87 | $0.0781 | $781 |
| Mistral Large 3 | $0.50 | $1.50 | $0.0934 | $934 |
| Qwen3.5 397B | $0.60 | $3.60 | $0.125 | $1,250 |
| Kimi K2.6 | $0.95 | $4.00 | $0.186 | $1,857 |
| Claude Haiku 4.5 | $1.00 | $5.00 | $0.201 | $2,012 |
| Grok 4.3 | $1.25 | $2.50 | $0.224 | $2,245 |
| Qwen3.7 Max | $1.25 | $3.75 | $0.233 | $2,335 |
| GLM-5.2 | $1.40 | $4.40 | $0.263 | $2,630 |
| Gemini 2.5 Pro | $1.25 | $10.00 | $0.279 | $2,785 |
| Mistral Medium 3.5 | $1.50 | $7.50 | $0.302 | $3,018 |
| Gemini 3.5 Flash | $1.50 | $9.00 | $0.313 | $3,126 |
| GPT-4.1 | $2.00 | $8.00 | $0.388 | $3,880 |
| Claude Sonnet 5 | $2.00 | $10.00 | $0.402 | $4,024 |
| GPT-4o | $2.50 | $10.00 | $0.485 | $4,850 |
| GPT-5.4 | $2.50 | $15.00 | $0.521 | $5,210 |
| GPT-5.6 Terra | $2.50 | $15.00 | $0.521 | $5,210 |
| Claude Sonnet 4.6 | $3.00 | $15.00 | $0.604 | $6,036 |
| Kimi K3 | $3.00 | $15.00 | $0.604 | $6,036 |
| Claude Opus 4.8 | $5.00 | $25.00 | $1.01 | $10,060 |
| GPT-5.5 | $5.00 | $30.00 | $1.04 | $10,420 |
| GPT-5.6 Sol | $5.00 | $30.00 | $1.04 | $10,420 |
| Claude Fable 5 | $10.00 | $50.00 | $2.01 | $20,120 |
*At 10,000 outcomes per month. Cheapest model highlighted.
The clean-path steps this benchmark prices:
This path runs 8 steps: 4 tool calls, 1 reasoning step and 3 decision points. Tool steps make two model calls each, and the agent re-reads its growing context on every call. That compounding is why one Offboarding outcome costs about 26x a single chat message ($0.604 on Claude Sonnet 4.6), not the price of one message.
These figures are modeled estimates, not metered bills. We price a generic, representative Offboarding workflow across 26 models using the same cost engine as the live estimator, at each model’s published list price (checked 26 Jul 2026), under documented default assumptions for planning loops, tool calls, memory retrieval, sub-agents and context size. Your own process will differ, so use these as starting points, tune the assumptions in the estimator, and validate against your real usage. Illustrative estimates, not financial advice.
On the clean path with default assumptions, an agent for Offboarding costs about $0.0291 to $2.01 per outcome depending on the model, or roughly $291 to $20,120 per month at 10,000 outcomes. The cheapest model here is GPT-4o mini at $0.0291; the most expensive is Claude Fable 5 at $2.01.
An agent does not make one model call. It plans, calls tools, retrieves context and re-reads its growing working context on every step. For Offboarding that adds up to about 26x the cost of a single chat message.
Across the 26 models benchmarked, GPT-4o mini is cheapest at $0.0291 per outcome and Claude Fable 5 is the most expensive at $2.01. A cheaper model is not always the right choice, but it sets the floor for this workflow.
The biggest levers are prompt caching on the base context, fewer planning loops, smaller tool results, less retrieval, and choosing a cheaper model where quality allows. You can test each lever in the live estimator.
These are modeled estimates, not metered bills. Each figure prices a generic, representative Offboarding workflow across 26 models with the same cost engine as the live estimator, at each model's published list price (checked 26 Jul 2026), under documented default assumptions for planning loops, tool calls, memory retrieval, sub-agents and context size. Your own process will differ, so treat these as starting points, tune them in the estimator, and validate against your own usage.
Cost is one axis. Bolting AI onto an unstable process just makes the mess run faster, and regulators expect provable human oversight. See a governed, human-in-the-loop version of this process, where the AI recommends, a human decides at every control point, and closure produces attestation evidence. See the governed Process Animation →