Know your cost per resolution before you scale the deployment.
Built for CIOs, platform owners and AI leaders being asked what production AI costs to run. Token volume, model mix, caching and human escalation in one unit-economic model.
Your inputs
Benchmarks show typical enterprise ranges — override every field with your own numbers.
Conversations, tickets or tasks that hit a model in production
Benchmark: 50k–1M for an enterprise deployment at scale
Includes retrieval reranking, tool selection, generation and evaluation calls
Benchmark: 2–6 for a RAG or agentic pattern
Retrieved context usually dominates — count it
Benchmark: 2k–8k with enterprise RAG context
Weighted across your current model mix, before optimisation
Benchmark: $2–$12 depending on frontier model share
Prompt and semantic caching on repeated enterprise traffic
Benchmark: 20–45% on repetitive workloads
Simple intents that do not need the frontier model
Benchmark: 30–60% of enterprise traffic
Benchmark: 70–92% cheaper per token
Escalation cost belongs in cost per resolution
Benchmark: 15–35% in year one
Benchmark: $3–$12 in US enterprise service operations
Saving from caching and model routing at your volume, holding escalation cost constant.
Directional estimate. Assumes a $180k engineering investment to implement caching, routing and cost telemetry, and that escalation volume is unchanged by optimisation.
Three-scenario view
Finance reviewers expect a range. These scenarios flex adoption and implementation cost around the model you entered.
Slower adoption, higher integration effort
Your inputs as entered
Strong sponsorship, clean data, phased scale-up
How enterprise leaders use this model
- Why model AI run cost separately from build cost?
- Build cost is a one-time line item finance already understands. Run cost scales with volume, and it is what erodes the business case in year two. Cost per resolution — inference, retrieval, evaluation and the human escalation behind it — is the unit economic that decides whether a deployment stays funded.
- What is a healthy cost per resolution?
- It has to be judged against the human alternative, not against a token price. Most enterprise deployments land between $0.05 and $0.60 per fully-resolved interaction, versus $3–$12 for a human-handled one. If your blended AI cost exceeds roughly a third of the human cost, optimisation should come before scaling.
- How much does caching and model routing actually save?
- Prompt and semantic caching typically remove 20–45% of token spend on repetitive enterprise traffic. Routing simple intents to a small model and reserving the frontier model for complex work usually saves another 25–50% of blended cost with no measurable quality loss when it is evaluated properly.
- Does this replace our FinOps tooling?
- No. It sizes the prize before you instrument. Use it to set a cost-per-resolution target and to justify the optimisation work, then hold the deployment to that target with real telemetry.