- Product squads (Outcomes): Own the KPI, prioritize agents by ROI, run experiments
- AI platform (Capability): Tools, memory, orchestration, evaluation harness
- Safety function (Policy): Guardrails, red-team, audit trail, model risk
- CoE / Chief AI Officer (Governance): Portfolio, funding model, cross-BU standards
Why AI programs stall after the pilot
The pattern repeats across every enterprise we assess. Twenty to forty pilots run, three or four demonstrate real value, and none reach production at scale. The blocking constraints are almost never model quality. They are: no owner accountable for the business metric, data that cannot be retrieved with the right permissions, no path for an AI system to write into a system of record, and no evaluation harness — so nobody can prove the thing is safe to scale. Consulting that does not attack those four constraints is producing slides, not transformation.
What AI transformation consulting should actually deliver
Six deliverables, in order: a readiness diagnostic grounded in your data and estate, a prioritised workload portfolio with modelled economics, a target operating model naming who owns what, a reference architecture that survives vendor change, a governance and evaluation framework, and at least one workload in production during the engagement. If the statement of work ends at the roadmap, you have bought a research report at consulting prices.
Engagement models and how to price them
Three shapes work. A fixed-scope readiness and roadmap sprint (four to six weeks) when leadership needs an evidence-based plan. A build-and-transfer engagement (one to two quarters) where the advisor ships the first workloads and hands the operating model to your team. A managed AI operations model for enterprises without a standing platform team. Avoid open-ended time and materials advisory with no production deliverable — it is the single most common source of AI spend with no attributable outcome.
Choosing an advisor: five questions that separate them
Ask to see an evaluation harness they built and run in CI. Ask which system of record their last engagement wrote to, and who approved the entitlement. Ask for a workload where they recommended against building. Ask how inference cost was modelled at peak, not pilot, volume. Ask which of their people will still be on the account in month nine. Firms that answer all five concretely are delivery organisations; firms that pivot to a maturity model are not.
The 12-month transformation sequence
Quarter one: readiness diagnostic, data and permission remediation on one domain, one workload in production. Quarter two: operating model stood up, governance and evaluation in CI, second and third workloads shipped. Quarter three: platform consolidation and cost controls as inference volume grows. Quarter four: portfolio review, decommission the losers, and move the CoE from build to enablement. Each quarter ends with a measured business number, not a maturity score.
Cost of ownership nobody models at the start
Beyond licences and consulting fees: inference at production volume, the evaluation and QA harness, integration maintenance as source systems change, model and prompt version management, and the standing run team. In our experience run cost overtakes build cost in year two. Programs that never modelled it end up quietly rationing usage — which is how a successful pilot becomes a shelved capability.
Measuring transformation honestly
Track four families: business outcome per workload (cost, cycle time, revenue, risk), adoption depth (share of eligible transactions actually handled), quality and safety (evaluation pass rate, incident count), and unit economics (fully loaded cost per successful task). Maturity scores are useful for internal narrative and useless as a decision instrument. Fund the next wave on the first three.
What enterprises are actually buying
The phrase covers three very different engagements. Strategy work produces a portfolio, a value case and an operating model — useful, and worthless on its own. Delivery work builds and ships specific workflows into production. Capability work leaves behind a platform, standards and people who can continue without the consultant. Enterprises get poor value when they buy the first and expect the third. The engagement model that works pairs a small strategy phase with a delivery phase that proves the strategy on real workflows, and writes capability transfer into the contract as a deliverable with acceptance criteria rather than a closing sentiment.
Diagnostic before roadmap
A credible transformation diagnostic examines six things: the data estate and whether entitled, real-time access exists; the integration surface of the systems that matter; the current portfolio of pilots and why they have not scaled; the operating model, including who can approve a production change; the talent distribution between strategy, platform and delivery; and the financial model, including how AI spend is attributed today. Roadmaps written without this evidence tend to be sequenced by ambition, which is why they collide with reality in the second quarter.
Portfolio design and the value case
Structure the portfolio as three horizons running concurrently rather than sequentially: proven workflows that ship this quarter and fund the program, platform investments that make the next ten workflows cheaper, and exploratory bets with an explicit kill date. Each item carries a named business owner, a baseline metric, a target and a review date. Value cases should quantify cost to serve, cycle time, capacity released and revenue protected, and state the assumptions plainly so the CFO can challenge them. Portfolios where every item is a strategic bet, or where every item is a cost cut, both fail — for opposite reasons.
Choosing and managing a partner
Assess partners on production evidence rather than practice size: workflows live in production at comparable scale, the engineers who will actually staff your work, the reference architecture they bring, their evaluation and quality methodology, their position on model and vendor neutrality, and their willingness to put fees at risk against your metric. Ask who owns the intellectual property, what happens to the platform when the engagement ends, and how knowledge transfer is evidenced. A partner who cannot describe how they will make themselves unnecessary is selling a dependency.
How to know the engagement is working
Ninety-day markers are more honest than milestone plans: a workflow running in shadow against real traffic, a golden evaluation set in CI, a named business owner reviewing a scorecard, an entitlement and data access path solved once and reused, and at least one internal engineer able to ship a change without the partner. If those five are absent at day ninety, additional slideware will not recover the program — reset the scope to a single workflow and rebuild credibility with a shipped result.
Engagement shapes that work
Three shapes recur in successful programs. A four-to-six week diagnostic that produces a portfolio, a value case and a delivery plan grounded in evidence. A ninety-day proving engagement that ships one workflow to production while building the reusable platform pieces it needs. And an embedded capability engagement where partner engineers pair with internal teams on a rolling backlog with explicit transfer milestones. What fails is the open-ended advisory retainer with no production deliverable — it produces artefacts everyone agrees with and nothing anyone runs.
Internal capability versus partner leverage
Decide deliberately what you will never outsource: the intent and workflow knowledge, the outcome metrics, the data and entitlement model, and the evaluation sets. These are the assets that keep the program yours. Partners add most value in reference architecture, platform engineering acceleration, hard integration work, evaluation infrastructure and the delivery discipline of having done it before. Enterprises that outsource the first list buy speed now and dependency later; those that keep it can change partners without losing momentum.
Working with the CFO from the start
Bring finance into the program early rather than at approval. Agree the cost model — platform, inference, delivery, run — the attribution method, the baseline for each metric, and the review cadence at which workflows are continued or retired. Finance partners who helped design the measurement approach defend the program during budget pressure; finance partners presented with results after the fact interrogate them. This is a relationship decision more than an analytical one, and it changes program survival rates markedly.
Managing stakeholder expectations
Executive expectations are shaped by consumer AI and vendor marketing, which set an unrealistic pace for enterprise delivery where entitlement, integration and governance dominate the timeline. Reset expectations explicitly at the start with a realistic ninety-day picture, then over-communicate progress against it. Show working software early and often, even in shadow mode. Programs that go quiet for a quarter while building foundations lose sponsorship precisely when they are closest to delivering.
Exit criteria and what good looks like at twelve months
Define what the partner leaves behind: workflows in production with named business owners, a shared tool layer with documentation, evaluation infrastructure running in CI, governance evidence produced automatically, internal engineers shipping changes independently, and a portfolio process the enterprise runs itself. Write these as acceptance criteria in the contract. A transformation engagement without exit criteria tends to end when budget ends rather than when capability exists.
- AI programs stall on ownership, data permissions, write access and evaluation — not model quality
- A credible engagement puts at least one workload in production, not just a roadmap
- Five questions separate delivery firms from advisory theatre — start with the evaluation harness
- Run cost overtakes build cost in year two; model inference at peak volume from day one
- Buy strategy, delivery and capability transfer as one engagement, with transfer as a contractual deliverable.
- Diagnose data access, integration surface, operating model, talent and cost attribution before writing a roadmap.
- Run three portfolio horizons concurrently: fund-the-program workflows, platform investment and bets with kill dates.
- At day ninety, look for shadow traffic, a golden set in CI, a business owner scorecard and an internal engineer shipping changes.
Questions leaders ask us
- What is AI transformation consulting?
- AI transformation consulting is advisory and delivery work that takes an enterprise from scattered AI pilots to production workloads with a named owner, a governance framework, a target operating model and measurable business outcomes. Credible engagements ship at least one workload into production, not just a roadmap.
- How much does AI transformation consulting cost?
- Fixed-scope readiness and roadmap sprints typically run four to six weeks; build-and-transfer engagements run one to two quarters. The larger cost to model is ownership: inference at production volume, evaluation tooling, integration maintenance and the standing run team, which usually overtake build cost in year two.
- How do we choose an AI consulting partner?
- Ask to see an evaluation harness they run in CI, which system of record their last engagement wrote to, a workload they recommended against, how they modelled inference at peak volume, and which named people stay on the account in month nine.
- Why do most enterprise AI pilots never scale?
- Four repeatable constraints: no owner accountable for the business metric, data that cannot be retrieved with correct permissions, no write path into a system of record, and no evaluation harness to prove the workload is safe to scale.
- How long does an AI transformation take?
- A realistic first year is: readiness and one production workload in quarter one, operating model and governance plus two more workloads in quarter two, platform and cost consolidation in quarter three, and portfolio rationalisation in quarter four.
- What should an AI transformation engagement deliver in the first quarter?
- A workflow running in shadow against real traffic, an evaluation set in CI, a reusable data access and entitlement path, and a business owner reviewing a scorecard — not a roadmap alone.
- How do we evaluate AI consulting partners?
- On production evidence at comparable scale, the named engineers who will staff the work, their evaluation methodology, vendor neutrality, IP terms, and willingness to put fees at risk against your metric.
- Should strategy and delivery be separate vendors?
- Separating them is a common and expensive mistake: the strategy is written without delivery constraints, and the delivery partner spends the first quarter renegotiating it. Keep them accountable to one outcome.
- How do we avoid dependency on a consulting partner?
- Write capability transfer into the contract with acceptance criteria, keep IP and platform ownership, pair internal engineers on every workstream, and measure how many changes ship without the partner.