- 01
Falling token prices do not lower enterprise AI cost, because consumption per completed task rises faster than unit price falls.
- 02
By 2029 the dominant cost lines are integration, evaluation and supervision — not inference.
- 03
Cost per completed outcome is the only measure that survives a model change; cost per token does not.
The thesis
Every year since 2023 the price of a unit of model capability has fallen sharply, and every year enterprise AI spend has risen. That is not a contradiction and it is not waste — it is the standard pattern when a technology becomes cheap enough to use for things nobody attempted before. Our view is that this continues through 2029, and that leaders who build budgets around falling inference prices will be wrong in a specific and expensive way. Consumption grows faster than price falls, because agentic patterns replace single-shot ones: a workflow that reasons, calls tools, checks its own work and retries consumes an order of magnitude more capability per business event than a chat completion did. Meanwhile the costs that never appear in a model provider's price list — integration, evaluation, supervision, migration — grow with the number of workflows in production and do not fall at all. Planning for that shape is the difference between an AI portfolio that compounds and one that gets frozen after a budget review.
Where the money actually goes
Across enterprise programmes the mature cost structure has five lines and inference is rarely the largest. Integration: building and maintaining the governed interfaces into systems of record, including identity, entitlements and the unglamorous work of making a thirty-year-old core system respond predictably. Evaluation: golden sets, test infrastructure, and the human time to curate them, which grows with every workflow and every regulation. Supervision: the standing team that watches production behaviour and handles escalations. Inference and platform: the visible bill everyone forecasts. Migration: the periodic, unavoidable cost of moving between models, frameworks or vendors. In the first year of a programme, inference can look like most of the spend because the other lines are still being capitalised into a project. By the third year, in the estates we see, it is typically the minority. Budgets that model only the visible line will be reset mid-year, which is how AI portfolios lose executive confidence even when the workflows are working.
Why agentic patterns reset the baseline
A single-turn assistant answers once. An agentic workflow plans, retrieves, calls several tools, evaluates its own output, sometimes retries, and produces a trace for audit. The same business event therefore consumes many multiples of the tokens, plus retrieval and tool-call costs that scale with the number of systems involved. Two consequences follow for planning. First, any per-task cost benchmark drawn from an assistant pilot is invalid as soon as the workflow becomes agentic, and re-baselining is not optional. Second, cost control moves upstream into design: how many tool calls the workflow needs, whether cheap models can handle routing and extraction with an expensive model reserved for the hard step, whether results are cached, and whether the workflow knows when to stop. These are architecture decisions with a direct and large financial effect, which is why unit economics belongs in the design review rather than in the monthly finance report.
What gets cheaper, and what does not
Assume three things fall through 2029: the price of a given level of model capability, the cost of standard retrieval and vector infrastructure, and the cost of the second and third workflow on a platform you have already built. Assume three things do not fall: integration with systems of record, because that cost is set by your estate rather than by the market; evaluation, because the bar rises as deployments become consequential; and supervision, because it scales with the number of production workflows. The practical implication is that marginal cost is where the leverage sits. An enterprise where workflow four costs a fifth of workflow one has built a platform. An enterprise where workflow four costs the same as workflow one has built four projects, and its cost curve will be linear at exactly the moment the board expects it to bend. That ratio — the cost of the latest workflow against the first — is the single most diagnostic number a CFO can ask for.
Concentration risk and the cost of moving
Every enterprise should assume at least one forced migration before 2029: a model retired, a vendor repriced, a framework abandoned, a regulator requiring a change in where processing happens. The cost of that migration is decided years earlier by architecture. If prompts, tool definitions and evaluation suites are portable, migration is weeks of work per workflow. If business logic lives inside a vendor's orchestration canvas and quality is defined by demonstration rather than by a test suite, migration is a rebuild. This is a financial argument as much as a technical one, and it belongs in the business case as a line item rather than as a risk register entry. The practical hedge is cheap: keep the evaluation suite outside the vendor, keep tool definitions in your own layer, and run at least one workflow against a second model each quarter so the ability to switch is exercised rather than assumed.
Implications by role
CFO: replace productivity percentages with cost per completed outcome, and require the marginal-cost ratio between the newest and first workflow at every portfolio review. CIO: fund the reusable layers centrally as infrastructure, because charging them to the first workflow makes that workflow look uneconomic and makes every later one look artificially cheap. Chief AI Officer: make unit economics a gate in the design review, since most of the cost is fixed by architecture before any code runs. Enterprise Architecture: keep evaluation and tool definitions vendor-external and prove portability quarterly. COO: understand that supervision cost is permanent and grows with the estate — a plan that shows it decaying to zero is not a plan. Procurement: negotiate for price protection and exit terms rather than the lowest headline rate, because the headline rate is the line most likely to fall on its own.
How to build a 2027–2029 budget that holds
Five practices produce budgets that survive contact with reality. Separate infrastructure from workflows: one envelope for tool layer, retrieval, evaluation and observability, and per-workflow funding for everything else. Publish a cost-per-outcome number for every production workflow, monthly, next to its quality number, so nobody optimises one blind to the other. Assume consumption growth of a different order than price decline, and stress-test the plan at several times current volume rather than at last quarter's. Fund supervision as headcount rather than as contingency. And book a migration reserve — a defined percentage of the AI run budget set aside for the forced move you cannot yet name. Enterprises that adopt these five look conservative in 2026 and turn out to be the ones still expanding in 2029, because their numbers were believed. The goal is not to spend less on AI; it is to be able to explain, at any point, what the next unit of spend buys.
- Falling token prices do not lower enterprise AI cost, because consumption per completed task rises faster than unit price falls.
- By 2029 the dominant cost lines are integration, evaluation and supervision — not inference.
- Cost per completed outcome is the only measure that survives a model change; cost per token does not.
- Agentic workflows consume five to twenty times the tokens of a single-turn assistant for the same business event, and plans must assume that step change.
- The cheapest enterprises are not the ones using the cheapest models but the ones that stopped re-solving the same integration problem for every workflow.
Questions leaders ask us
- If model prices keep falling, why plan for rising cost?
- Because consumption per business event rises faster. Agentic workflows use many multiples of the capability a single-turn assistant used, and the non-inference lines — integration, evaluation, supervision — do not fall at all.
- What single number should we track?
- Cost per completed outcome, reported next to quality for that workflow. It survives model changes, comparisons across workflows and conversations with finance.
- How do we know whether we built a platform or a set of projects?
- Compare the cost of your most recent workflow with your first. A large drop means the reusable layers are real; a flat line means each team is re-solving the same integration and evaluation problems.
- Is a migration reserve really necessary?
- Assume at least one forced model, framework or vendor move before 2029. A reserve turns that from a mid-year budget crisis into a planned piece of work.
Sources
- [1] The large majority of enterprise generative AI pilots never produce a measurable production outcome. The GenAI Divide: State of AI in Business — MIT NANDA / Project NANDA, 2025