Eighty-eight percent of organizations now use AI in at least one business function. That number describes adoption, not results. According to McKinsey's "The State of AI in 2025" survey of 1,993 respondents across 105 countries, fielded June 25 to July 29, 2025, only about 6 percent of organizations clear the bar McKinsey sets for an "AI high performer": attributing more than 5 percent of earnings before interest and taxes (EBIT) to AI. Adoption is nearly universal. Profit is not.
How widespread is AI adoption today?
McKinsey's survey puts adoption at 88 percent of organizations using AI in at least one business function. Scaling tells a narrower story: only about 23 percent of organizations are scaling an agentic AI system anywhere in the business, and within any single business function, no more than roughly 10 percent report agents running at scale. Most organizations have started. Few have moved past a pilot in any one function, and fewer still are running at scale across the business.
Why do so few organizations see AI hit the P&L?
Adoption measures whether a tool is in use somewhere. It does not measure whether that use changes the bottom line. McKinsey's 6 percent figure applies a stricter bar: organizations that can point to more than 5 percent of EBIT coming from AI. The distance between 88 percent and 6 percent is the value gap. Most organizations that use AI cannot show that it moved a number that matters to the business.
Scaling and profit are also different questions. An organization can roll a workflow out to every team in the business and still not see it move EBIT, if no one measured whether the workflow's output changed a number that mattered in the first place. Scale without a measured baseline is adoption spread wider, not value proven.
Is this a McKinsey-only finding?
No two survey methodologies agree on an exact figure, but the direction is consistent across independent research. Gartner's own data shows a similar pattern: of 782 infrastructure-and-operations leaders surveyed, only 28 percent of their AI use cases met return-on-investment expectations., April 2026. Two different surveys, run by two different firms, land on the same conclusion: a small minority of AI deployments produce a result that leadership can defend in a budget review.
Why does adoption not turn into value?
The most common reason has little to do with the model. Most organizations simply never establish a baseline in the first place. An organization that deploys AI into a workflow without first measuring that workflow's cost, speed, or error rate has no number to compare against afterward. When someone later asks whether the tool is working, the honest answer is that no one recorded what "before" looked like. Impact cannot be proven, even when the tool genuinely works, because there is nothing to measure it against.
A second reason compounds the first. Without a target set in advance for what success means, a deployment stays in pilot indefinitely. It is not scaled, because no one agreed beforehand on the number that would justify scaling it, and it is not killed, because no one agreed on the number that would justify stopping it. It sits in the gap between adoption and value, consuming budget without a decision either way.
Is the problem the model or the process?
Model capability has advanced faster than most organizations' ability to measure what the model changed. A more capable model can still be deployed into a workflow with no recorded starting point, and the result is the same regardless of how good the model is: no one can say, with a number, whether it worked. The value gap does not show that the technology underperforms. Most deployments were simply never set up to be measured in the first place.
What counts as a measured baseline?
A measured baseline is a number recorded before an agent touches the workflow, not after. At minimum, it covers three things: how long the workflow currently takes, how often it produces an error or requires rework, and what it costs per unit of output: per case, per quote, per ticket, per document. The baseline is signed off by the team running the workflow and by finance, so it cannot be redefined after deployment to make a middling result look successful. Without that sign-off, a post-deployment number is a claim, not a measurement.
What does skipping the baseline cost a business?
The direct cost is budget spent on a pilot that cannot be defended in a P&L review. The larger cost is organizational. A leadership team that cannot point to a single AI deployment with a measured result loses the internal case for the next one. Every new pilot then starts from the same skepticism the last one ended in, and the business spends its time re-litigating whether AI works at all instead of scaling the workflows that already do. Over several pilot cycles, this pattern compounds: budget keeps flowing to new experiments, while the organization accumulates no evidence, favorable or unfavorable, about any of them.
What should leadership ask before the next AI deployment?
Three questions separate a measured deployment from another unproven pilot. What does this workflow cost today, in time and money, before AI touches it? What number, tied to an outcome the business already tracks, would justify scaling this beyond the pilot? Who has agreed, in writing, to those numbers before deployment starts? An organization that cannot answer all three is not yet ready to deploy. It is ready to guess.
What is the fix?
Eighty-eight percent of organizations have already adopted AI somewhere. What is missing is measurement, applied before deployment rather than after.
Mirai360 AI's approach to a new deployment follows a fixed sequence:
- Pick one workflow, narrow enough to measure, not a company-wide rollout.
- Record a measured baseline for that workflow before any agent touches it: current cost, speed, and error rate, signed off by the team and by finance.
- Deploy against a fixed budget and time box, with the target outcome tied to a business number, such as cost saved, hours freed, or EBIT impact, set before deployment rather than after.
- Scale only against the pre-set numbers. If the deployment clears the bar, scale it. If it does not, fix it or kill it. Either decision is made against a number agreed on in advance, not an opinion offered after the fact.
This sequence does not guarantee that every workflow clears the bar, but it does guarantee that the business knows, with a number, whether it did, which is the one thing most organizations currently cannot show.
Where does this leave a business evaluating AI today?
Being part of the 88 percent that use AI is no longer a meaningful distinction. Being part of the roughly 6 percent that can show AI moved a number that matters is. The difference between the two groups has nothing to do with the tool. It comes down to whether someone signed a baseline before the tool went live.
Mirai360 AI works with businesses to pick one workflow, measure it before deployment, and scale only against numbers set in advance. Book a 30-minute call to identify the workflow worth measuring first.
Sources
- McKinsey & Company, "The State of AI in 2025": survey of 1,993 respondents across 105 countries, fielded June 25 to July 29, 2025, published November 2025.
- Gartner: survey of 782 infrastructure-and-operations leaders; 28% of their AI use cases met ROI expectations., April 2026
FAQ
- Why do so few organizations see AI hit the P&L?
- Adoption measures whether a tool is in use somewhere, not whether it changes the bottom line. McKinsey finds only about 6% of organizations attribute more than 5% of EBIT to AI, and the gap between the two is the value gap.
- Why does adoption not turn into value?
- Most deployments skip a measured baseline. Without a recorded 'before' state, impact cannot be proven even when the tool works, and there is no pre-set number to decide whether to scale, fix, or kill the pilot.
- What counts as a measured baseline?
- A number recorded before deployment covering current cost, speed, and error rate for one workflow, signed off by the team running it and by finance, so it cannot be redefined after the fact.