- 01
Measure the baseline before anything changes; a baseline reconstructed after launch is never believed at review.
- 02
Separate cash benefit from capacity benefit and label each honestly — conflating them is why cases fail their first audit.
- 03
Retained cost is the number most cases omit: the exceptions left behind are the expensive ones.
Why automation cases fail review
Most back-office automation cases are not rejected for weak returns; they are discredited at the first review because the numbers cannot be reconciled with the general ledger. Four causes recur. The baseline was estimated rather than measured, so the improvement is unverifiable. Benefit was stated as hours saved and implicitly treated as cash, when no headcount, contractor spend or overtime actually changed. Retained cost was omitted, so the exceptions that still require human handling — and which are systematically harder than the average case the baseline measured — appear nowhere. And run cost was modelled at pilot volume, missing the inference, platform and operating cost that scales with production throughput. The remedy is not conservatism for its own sake but construction discipline. A case built with a measured baseline, benefits separated by type, retained cost included and run cost modelled at scale will typically show a smaller headline number and a far higher chance of surviving. That trade is worth making, because a programme whose first review is credible receives funding for the second and third workflow, which is where the platform economics actually pay back. A programme whose first review collapses rarely gets a second.
Measuring the baseline properly
The baseline needs four components per process. Volume: annual instances with seasonal distribution, taken from system counts rather than team estimates. Fully loaded cost per instance: not just the processing team's salaries but supervision, quality assurance, training, facilities, systems licensing and an allocation of the support functions that exist because the volume exists. Cycle time: end to end from receipt to completion, distinguishing the straight-through path from the exception path, because the averages hide everything that matters. Quality: error and rework rates, downstream correction cost, and any penalty, interest or service credit incurred because of processing failure. Gather these over a representative period — a single month distorted by a system outage or a seasonal peak will be challenged. Split the population into first-pass and exception cases and cost each separately; in most back-office processes the exception minority consumes the majority of the effort, and a case that reports only the blended average will misprice both the benefit and what remains. Have finance validate the baseline before the programme starts. The cost of that validation is a few days; the cost of skipping it is an unwinnable argument at review about whether anything improved.
Categorising benefit without overstating it
Use four categories and label every line. Cash cost reduction is benefit that will visibly change a cost line: reduced contractor or outsourced volume, avoided recruitment against an approved plan, eliminated overtime, retired licences. This is the only category finance will treat as hard. Cost avoidance is growth absorbed without added cost — real and defensible when volume growth is documented and the hiring plan it displaces was approved. Working capital is the ledger-visible effect of faster processing: reduced days sales outstanding, captured early-payment discounts, reduced interest on late settlement. Capacity release is time returned to existing staff; it is genuine but is not cash until something changes, so state what the capacity will be used for and who has committed to redeploying it. Quality and risk benefits — fewer errors, faster audit response, reduced penalty exposure — should be quantified where a historical loss record exists and described qualitatively where it does not. Fabricating a number for an unquantifiable benefit damages the whole case. Reviewers discount every line once they find one they distrust, so the discipline of honest labelling protects the credible lines rather than weakening the case.
Retained cost, run cost and the true unit economics
Two cost categories decide whether a case is real. Retained cost is the human work that remains: exception review, quality assurance sampling, escalation handling, and the supervision the residual team still needs. Because automation absorbs the easy instances first, the remaining population is harder than the baseline average — cost per retained case rises even as total volume falls, and cases that apply the baseline average to the residual overstate benefit substantially. Estimate retained cost from the exception population specifically, and re-measure after the first quarter in production. Run cost is what the automation consumes: inference by model and step at production volume, document processing, platform and storage, monitoring, and the permanent team that maintains evaluation sets, investigates drift, manages releases and handles incidents. That team is not optional and does not disappear after go-live; omitting it is the second most common flaw after missing retained cost. Model unit economics at three volumes — pilot, first-year production and steady state — because inference and platform costs scale roughly with volume while engineering cost does not, and the shape of that curve determines whether the case improves or deteriorates as the programme succeeds.
Sensitivity analysis and the numbers to stress
Present the case as a range with the assumptions that drive it exposed. Stress four variables. Straight-through rate: what happens at half the assumed rate, which is a realistic first-year outcome for a process with messier data than expected. Retained cost per exception: what happens if it is a third higher than modelled. Inference cost: what happens if volume mix shifts toward cases requiring more model calls, and conversely what the case looks like after the optimisation most teams achieve in the first two quarters. Timeline: what happens if production ramp slips a quarter, which delays benefit while run cost begins accruing. Show the payback period at the pessimistic case, not only at the expected case, and state the point at which the programme should be stopped — a case that cannot describe its own failure condition is not a case. Reviewers respond well to this framing because it demonstrates the numbers were built rather than reverse-engineered from a target, and it converts the review conversation from advocacy into a shared assessment of which assumptions to monitor.
The ninety-day review and what to bring to it
Schedule the review before the programme starts and agree its contents in the funding paper. Bring: actual versus modelled straight-through rate with the reasons for variance; measured cost per case against baseline, with retained cost separated; actual run cost by component at current volume; reviewer minutes per exception; error and rework rates against baseline quality; the working-capital effects visible in the ledger; and a short list of what was learned and what will change as a result. Where a benefit has not materialised, say so plainly and explain whether it is delayed or wrong — programmes that report honestly at ninety days retain credibility to continue, while those that present only favourable metrics lose the sponsor's trust at the second review regardless of underlying performance. Use the review to decide the next workflow and to publish the deployment cost curve, since the argument for continued platform investment rests on the second deployment costing meaningfully less than the first. A programme that can show that curve has moved from a project conversation to a capability conversation, and capability conversations attract multi-year funding.
Portfolio economics across multiple workflows
A single-workflow case rarely justifies the platform investment honestly, and stretching it to do so is how programmes acquire numbers they cannot defend. Present the economics as a portfolio instead. Separate the investment into platform — ingestion, retrieval, tool layer, evaluation harness, observability, audit evidence — and workflow-specific build. Amortise the platform across the workflows the roadmap commits to, and show the marginal cost and payback of each successive deployment. This produces a curve rather than a point: the first workflow may pay back slowly or barely, the second considerably faster, and the fourth and fifth quickly enough to be uncontroversial. Sponsors understand this shape because it matches every other platform investment they have approved. It also creates the right governance question at the ninety-day review — not whether the first workflow succeeded in isolation, but whether the marginal cost of the next one fell as predicted. Programmes that cannot demonstrate a falling curve after three deployments are running projects on shared branding, and the honest response is to fix the platform reuse before funding a fourth.
Common measurement traps
Five traps recur often enough to check for explicitly. Comparing automated cases against the blended baseline rather than against comparable cases, which flatters the result because automation takes the easy population first. Counting benefit from volume that would have declined anyway due to upstream fixes or product changes, which is attribution error and will be spotted. Ignoring the productivity dip during transition, which is real for a quarter and shows up in the numbers whether or not the case acknowledged it. Measuring quality only on the automated path while error rates rise on the human path because the residual work got harder and staffing did not adjust. And declaring success at a point chosen for favourable numbers rather than at the pre-agreed review date. Guard against all five by fixing the measurement definitions and the review date in the funding paper, having finance own the calculation rather than the programme, and reporting the unfavourable metrics alongside the favourable ones. The credibility this buys is what funds the next three workflows.
- Measure the baseline before anything changes; a baseline reconstructed after launch is never believed at review.
- Separate cash benefit from capacity benefit and label each honestly — conflating them is why cases fail their first audit.
- Retained cost is the number most cases omit: the exceptions left behind are the expensive ones.
- Model run cost per unit at production volume, not at pilot volume, and include the permanent operating team.
- Commit to a ninety-day review with pre-agreed metrics; cases without a scheduled review become unfalsifiable.
Questions leaders ask us
- How long should the baseline measurement period be?
- Long enough to be representative — typically one to three months covering a normal operating cycle, excluding periods distorted by outages, system changes or extreme seasonality. Take volume and cycle time from system data rather than estimates, and have finance validate the fully loaded cost per instance before the programme begins.
- Can we count hours saved as financial benefit?
- Only when something downstream changes: contractor spend reduced, approved hiring displaced, overtime eliminated, outsourced volume cut. Otherwise label it as capacity release and state what that capacity will be redeployed to and who has committed to it. Presenting capacity as cash is the fastest way to lose credibility at review.
- What payback period is realistic for a first workflow?
- For a well-chosen first workflow with existing system access, twelve to eighteen months is a defensible expectation once retained cost and the permanent run team are included. Subsequent workflows reusing the same platform layers usually pay back considerably faster, which is why the programme case should be presented across a portfolio rather than on the first deployment alone.
- How do we account for inference costs that keep changing?
- Model at today's prices with your measured token and call profile, then show sensitivity in both directions. Instrument cost per case from the first production day so the model is corrected by observation within weeks. Teams that measure typically find meaningful reduction available through model routing, caching and context trimming during the first two quarters.
- Who should own the business case after approval?
- The business owner of the process, not the technology team. They accept the outcome metric, report at the ninety-day review, and hold the authority to pause the rollout if quality degrades. Cases owned by the delivery team drift toward reporting delivery milestones rather than business results.
Sources
- [1] A majority of organisations report difficulty quantifying realised value from AI initiatives against a measured baseline. The state of AI: Global survey — McKinsey & Company, 2024
- [2] Most generative AI projects that fail to scale do so for reasons of cost, data readiness and unclear business value rather than model performance. Gartner predictions on generative AI project abandonment — Gartner, 2024