Capability stopped being the constraint. A new leading AI model landed every five and a half days over the past year, from eighteen different labs, and what stands between that supply and the work inside most companies has nothing to do with access, price or technical skill. Inside most of the companies we work with, deciding what a new model means here is nobody's job, so the releases pile up outside the business and the company keeps running on whatever it happened to pick the last time somebody looked.
What a year of releases actually looked like
Counting model releases only means something if someone defines what counts, so we printed the rule on the chart: one mark per generation of a model line that has, at least once this year, been its maker's most capable publicly available model, dated at first public availability. A cheaper or smaller variant of something already counted gets no mark of its own, unless no general model shipped in that generation. The rule kept 65 releases and cut 40, and anyone who wants the forty can ask for them.
What survives is a year with almost no quiet stretches in it. Eighteen labs, sixty-five releases, one every five and a half days on average, and a company that settled on a provider last September has watched fifty-seven land since.
The case for ignoring all of it
There is a serious argument that none of this belongs near a leadership agenda, and it deserves to be made properly before it gets answered. A survey of 150 technical decision-makers in the summer of 2025 found that across the whole sample, 66% had upgraded within the provider they already had, 23% had made no model change at all, and 11% had moved to a different vendor.
Two other findings point the same way. When one team was forced off a model that had been retired underneath them, none of three migration methods beat the original prompt by a statistically significant margin, because that prompt travelled between models far better than anyone expected. And only 16% of enterprise deployments qualify as true agents, meaning the model plans, acts, watches what happens and adapts. Most of the rest run a fixed sequence around a single call.
Read together, model churn looks like what it usually is in practice, a maintenance chore that lands on an engineer's backlog and never reaches a board. For a great deal of the work sitting inside companies today, that reading is simply correct, and any argument that skips past it is selling something.
The gap between what shipped and what got used
What it fails to explain is the direction of travel. The most recent read on the buying end runs only three months into that release year, and it already points the other way. Three providers account for 88% of enterprise AI spending, and the enterprise share going to open models fell from 19% to 11% over the twelve months to that date. Choice multiplied at the supply end while the buying end narrowed.
Worth naming the exposure in that number before someone else does. It is a venture firm's own modelled estimate, built from self-reported workload shares, and the firm is an investor in the provider it ranks first. The conflict touches the ranking rather than the total, and the total is what matters here, because 88% in three hands is the same finding whichever of the three leads. Nobody has published a comparable read on the nine months since, which is its own piece of evidence: the supply side gets counted weekly, the absorption side once a year.
Inertia explains part of that, and inertia is real. The more useful explanation is cost, because finding out whether a new model does your work better is expensive in a way that buying it is not. One research group costed a single honest evaluation pass at roughly $40,000, and nearer $320,000 once the runs are repeated enough times to trust the answer. That is research scale rather than company scale, but it tells you the shape of the problem: the model is cheap and the knowledge about the model is expensive. So the question gets deferred, and because it belongs to nobody in particular, deferring it costs nothing visible until the day it does.
The decision that arrives whether or not you make one
Changing nothing works right up to the point where the choice gets taken away, and the AI providers publish the dates on which that happens. Anthropic is the transparent end of the range: every active model carries a date written as a floor rather than a plan, guaranteeing at least twelve months from release, with sixty days' notice to customers running it in production. Two of its models retired this year exactly 365 days after release.
The other two are less generous. OpenAI deprecated six models in a single notice in June, with the shutdown set for December. Google publishes shutdown dates while stating plainly that they are the earliest possible dates a model might be retired, which is a notice period of zero.
That is the part that belongs on a leadership agenda rather than in a backlog, because a retirement is contractual, dated and somebody else's decision, while an upgrade stays optional and yours, and only one of the two puts a deadline in your calendar.
The bill for the involuntary version is not fixed either. Researchers running a nine-by-nine matrix of model switches found task performance moving anywhere from eight points worse to thirteen points better, with roughly a quarter of the switches moving by a statistically significant margin. Most swaps are uneventful, a minority are not, and nothing tells you in advance which one you are holding. That rarely surfaces on simple work. On the agentic work companies are now building it is the entire result.
None of that argues for chasing releases. It argues that the question of what a new capability changes here should belong to somebody by name, on a standing basis, and should be answerable before a retirement notice sets the timing for you.
The uncomfortable question
Sixty-five leading models landed in the last year. Name the person whose job it was to look at them and decide what, if anything, any of them changed here, and say what they decided. If that question never reached anyone, then it did not go unanswered, it was answered by default, and it will keep being answered that way until a date somebody else published turns it into a deadline.