AI money ends up spread thin across small pilots for a reason that has little to do with caution. Nobody in the room can defend the big number. A case assembled from someone else's results holds a single multiplication and no answer to the first question a finance director asks, so it gets quietly withdrawn rather than formally rejected, and what survives is a set of experiments small enough to approve without anybody making a decision.

The case for spending first

BCG made the argument plainly in July. Its piece on the cost of caution says that "for many, the real risk in AI isn't spending too much; often, it's spending too little," that leaders "concentrate resources" rather than "distributing small budgets broadly," and that "excessive precision can become a strategic constraint in fast-moving environments." The direction is right, and most of it is worth taking seriously.

Two things about that piece are worth naming before anyone leans on it. It carries no author, no methodology, no sample size and no definition of the groups it compares, which is unremarkable for a three-minute article and a problem for a load-bearing citation. And the two figures inside it, 1.7% of revenue against 0.8%, appear in BCG's own AI Radar in January as something else entirely: a year-over-year forecast across all companies, from 2,360 executives in sixteen markets, with AI spend rising from roughly 0.8% of revenue in 2025 to about 1.7% in 2026. Same digits, different claim. The forecast is the one that can be checked.

The same article also says "this does not mean scaling blindly," which is the sentence the headline buries. So the disagreement worth having is narrower than the headline suggests, and it sits on one line: the companies pulling ahead are "willing to invest before fully mapping out all returns."

The forecasts that did not arrive

There is a reason executives get told to stop demanding precision. The precision on offer has been bad.

Last year McKinsey asked respondents at AI-using organisations how AI would change their headcount over the coming year, and 32% expected a decrease. Asking again this year, 14% report that AI contributed to an actual decline. In every function McKinsey asked about, the share reporting a real reduction was smaller than the share that had predicted one a year earlier. The share expecting cuts over the coming year went up anyway, to 39%.

That comparison deserves its caveat, in a piece about caveats. It sets two survey waves side by side rather than one fixed panel, though McKinsey reports that the 552 respondents who answered in both years showed the same pattern. What it measures is what people attribute to AI, not what happened to headcount. It is still the cleanest published read on the distance between what companies expect from AI and what they report getting.

The pattern turns up again where the measurement is harder to argue with. Researchers at Chicago Booth and the University of Copenhagen linked adoption surveys to Danish administrative records across seven thousand workplaces in the occupations most exposed to chatbots, and found what they describe as "precise null effects on earnings and recorded hours at both the worker and workplace levels, ruling out effects larger than 2%." Scope that honestly. It covers chatbots in the first two years after ChatGPT, before agents and before reasoning models, and it is a working paper rather than a verdict on what AI can do now. The authors also do not read their own null as inaction: their finding is that employers absorbed the reported time saving by reorganising tasks rather than by cutting labour input. Which is the point. The saving was real to the people reporting it, and it never reached pay or recorded hours.

The most useful evidence sits inside a single firm. Seven hundred and fifty-eight consultants at BCG ran the same pre-registered experiment on the same frozen version of GPT-4, with a randomly assigned third of them working without it. On tasks inside what the researchers called the jagged frontier, the ones with the tool completed more work and completed it faster. On a task placed just outside that frontier they were 19 percentage points less likely to reach the correct answer than the consultants working without it. Same firm, same tool, two randomised halves of one consultant pool. The sign of the effect changed with the fit between the tool and the task.

Which is the finding underneath all three. The shape of an AI result travels between companies. The size of it does not.

What a borrowed case is made of

A borrowed case is built forward from a claim. Somebody reads that AI handles 40% of first-line tickets, or cuts document review by 70%. The percentage is carried across whole, as a single figure rather than a range. It gets applied to the coarsest base available without having to ask anyone, which usually means a line on the P&L: total support cost, total finance salaries. One multiplication produces a saving. Cost is the vendor quote. Payback is the saving divided by the quote.

A case built from inside starts somewhere else, at the level a person actually does the work. Not the support function but one inbound ticket of a given type. Not finance but one monthly reconciliation of one entity. Not sales but one quote for a project under fifty thousand euros.

Exhibit 1
The same move, modelled twice.
Borrowed
Built inside
Unit
The department
One quote under €50K
Volume
Estimated from memory
System export, filter stated, twelve months
Time per unit
Whole task, average
Only the steps the change touches
Cost of that time
Average salary
Loaded rate for whoever does it
Share AI handles
The vendor's percentage
A named range, from your own case mix
Run cost
not in the model
Per unit, scaling with the same volume
Change cost
not in the model
Rework, training, parallel run, expert hours
Conversion to cash
not in the model
A hire not made, contractor spend removed, or booked as capacity
Baseline
not in the model
Value, source, date, owner
Output
One number
A range, with the swing driver named
Both columns model the same move. The four cells marked "not in the model" are usually not errors in the borrowed case, they are terms it never contained, which is why it cannot answer a question about any of them.

The difference that matters is not effort. The borrowed case is denominated in percentages and the inside case is denominated in units.

A percentage travels between companies effortlessly and means nothing on arrival. A unit count does not travel at all, which is exactly what makes it worth something.

A borrowed case is also one multiplication over an aggregate, so it has no joints. Challenge any term and the whole thing goes with it, which is why the sponsor stops defending it after the second question. A case built from inside has a source behind each term, so challenging one number moves the answer instead of ending the meeting.

Why every term belongs to one company

Take the calculation apart and the reason turns mechanical rather than philosophical. Annual value is the volume of the work, times the time it takes per unit, times the loaded cost of that time, times the share the change can realistically handle, and then net of what it costs to run and what it costs to change.

Volume is definitionally yours. There is no industry figure for how many quotes your company writes. Sector ratios exist, and they average across business models that differ on the things driving volume most, which are customer count, product complexity and how much your customers already do themselves. When volume gets guessed it comes in high and at the wrong granularity, because the person estimating is remembering the pain rather than counting the work.

Time per unit is a function of your process design, your handoffs, your approval layers and your rework rate, which is exactly what a benchmark strips out. Two firms writing the same quote can differ by a multiple, not a margin, because one has a configured price list and one rebuilds pricing every time. The common failure here is arithmetic rather than optimism: elapsed time gets substituted for touch time, and whole-task time gets credited to a change that touches one step of it. If drafting is 40% of a ninety-minute quote, a perfect drafting capability saves thirty-six minutes, not ninety.

Loaded cost is the term where a benchmark is least wrong in size and most wrong in relevance. Salary surveys give gross pay for a job title. The calculation needs the fully loaded hourly cost of whoever actually does this task, divided by productive hours rather than contracted ones. The valuable discovery is usually not the number but the name attached to it, and the sentence that moves a case most is some version of "a partner does this."

The share AI can handle decides the answer, and it is the one term that cannot be sourced from anywhere at all. Published rates are measured against another company's case mix and another company's definition of success. Ask what counted as handled, and whether a ticket that reopened the next day was counted once. What can be established internally is the mix: take a hundred of your own real cases and sort them into fully handled, partly handled, and untouchable. That bounds the question without answering it, which is why the share belongs in a case as a named range rather than a figure buried inside a total.

Then the two terms a borrowed case usually does not contain at all. Run cost scales with the same volume that drives the value, so a case can be positive at pilot size and negative at full size, and the human review layer is part of it. Change cost is rework, training, the parallel run, and the hours your own experts spend during the build.

And one line that decides whether any of it is money. What do the freed hours become? A hire that was in the plan and now does not get made. Throughput absorbed at the same headcount, if there is demand waiting to absorb it. Contractor or overtime spend that comes off a named ledger line. If none of those applies, the value is capacity rather than cash, and a case that says so plainly is stronger than one that pretends otherwise.

The formula is not a secret. Any competent analyst would write the same shape. What makes the number hard to produce is that three of those four value terms and both cost terms live with different people at different levels of the business, and none of them holds more than a piece. Volume sits in a system nobody queries for this purpose. Step time sits with the person doing the work, and it usually contradicts what their manager believes. Loaded cost sits with finance and has rarely been asked for at task level. That is why the number cannot be looked up, and it is also why a tool with a proven track record somewhere else still arrives without one.

The pilot is not a decision

None of this is something a leadership team consciously decides. It is what happens when the paper cannot survive a meeting.

A case with a soft term at its centre does not get argued down. It gets withdrawn, or shrunk until it is small enough that nobody has to sign for it. That is the honest reason companies end up running a shelf of pilots instead of one committed change. A pilot is what a proposal decays into when nobody can defend the number behind it, and it costs so little to approve that no one has to.

The cost of that is real and nobody books it. A year of small experiments leaves no baseline, so next year's case rests on the same absence of evidence as this year's, and the argument for the bigger move never gets stronger. The expectation keeps climbing regardless. The same survey that recorded a forecast missing by more than half also recorded the next forecast going up.

Worth noticing who is saying what here. Even the company with the most to gain from you committing capital immediately does not tell you to skip the analysis. OpenAI's own description of its deployment work says a typical engagement "will begin with a focused diagnostic of where AI can create the most value." The order in that sentence is the whole argument.

The question you can answer this week

Take the largest AI commitment your company has made this year and ask whoever sponsored it for four numbers. The volume of the work it touches, and where that count came from. The loaded cost of the person who does that work. The share of it the change is assumed to handle, and what that share is based on. Then ask what the number was before you started.

If those come back from a system export, a payroll line and a sample of your own cases, you have a business case. If they come back from a vendor deck or a conference talk, you have a forecast about a company that is not yours.