MIT's Project NANDA put a hard number on what a lot of HR and ops leaders already suspected: ninety-five percent of enterprise generative AI pilots are not delivering measurable ROI. Not “almost there.” Not “just needs more time.” Ninety-five percent are producing nothing you could confidently slide across the CFO's desk without breaking into a cold sweat. The report, “The GenAI Divide: State of AI in Business 2025,” came out last summer, and I still watch its exact findings play out live every time I sit down with a company running a pilot.
Ask that company how the pilot is going, and you will get an answer, usually delivered with the confidence of a kid showing off a perfect attendance certificate. “Adoption is up 40% month over month.” “We had 600 prompts logged last week.” “Engagement with the tool is at an all-time high.” Someone built a dashboard. Someone is patting themselves on the back.
Here is the uncomfortable truth nobody wants to say out loud: none of those numbers actually tell you if AI is doing anything useful. They tell you people are using it. That is a different question entirely, and frankly, it is the wrong one.
What makes an AI metric a vanity metric?
You know a vanity metric when you see one, even if you have never called it that. It is the number that climbs up and to the right, making everyone in the meeting feel like they just won a prize, while the business quietly wonders if anything actually improved. Social media taught us this lesson years ago. Remember when “impressions” were all the rage, right before everyone realized impressions do not pay the bills? AI is running the same playbook, just with fancier graphics.
“AI adoption rate.” “Number of prompts per employee.” “Percentage of team trained on the tool.” Not bad things to track, exactly. But these are activity metrics wearing impact-metric costumes, and that difference matters a lot when someone is about to sign off on a six or seven-figure AI contract based on them.
Activity metrics vs. metrics that actually mean something
|
What people usually track |
What it actually tells you |
What to check instead |
|
AI adoption rate |
People opened the tool |
Cost per hire, time to fill |
|
Prompts per employee |
People are typing into it |
Error rate on payroll or benefits processing |
|
Percentage of team trained |
People sat through a session |
Turnover in the first 90 days |
|
Engagement score for the tool |
People did not immediately abandon it |
Revenue per employee, cycle time on a process |
So why does this keep happening? Partly because AI vendors love handing you a dashboard overflowing with usage stats, because usage stats always go up. Give people a shiny new tool and tell them to use it, and surprise, the numbers climb. A chart that only goes up is easy to sell. The other reason is that measuring real impact is harder and slower, and it requires you to have done your homework before you ever flipped the AI switch. Which, let's be honest, almost nobody actually did.
What were you measuring before AI showed up?
Here is a question worth sitting with. If you turned off every AI tool in your company tomorrow, would any of your actual business metrics move? Cost per hire. Time to fill. Error rates on payroll processing. Employee turnover in your first 90 days. Revenue per employee. Cycle time on a performance review. NPS from your internal HR service desk.
If you do not know the answer, do not beat yourself up. Almost nobody does, because almost nobody bothered to write down what those numbers looked like before the AI parade rolled into town. Without that baseline, you are not measuring AI's impact. You are describing whatever your metrics look like today and crossing your fingers that it is somehow related to the tool you bought in March.
This is the part that should keep people up at night even more than the MIT stat. The 95 percent failure number is ugly, sure. But a good chunk of that 95 percent might not even be a real failure. It could just be a company that never built the baseline to know if it succeeded, and defaulted to “no visible ROI” because nobody could prove otherwise. You cannot fail a test you never took.
The fix is not to invent a shiny new AI-flavored metric. It is to go back to the metrics you already had, the ones your finance team trusts, the ones sitting in your HRIS, your ATS, and your engagement survey data, and ask a much less glamorous question: did this number move, and can we actually tie that move to the tool?
What can HR learn from how engineering teams got AI metrics wrong first?
If you want a preview of where this is headed, look at what has happened in software engineering, because that team has been living this exact argument for two years already.
When AI coding tools first showed up, everyone reached for metrics like lines of code generated, commits per week, or tickets closed. Big numbers, great for the slide. Except a bunch of engineering leaders started noticing something uncomfortable: teams were shipping more code and somehow drowning in more problems. Pull requests got reverted more often. Code review queues backed up. Deployments that used to be routine started breaking things downstream.
It turned out “how much code did AI help us write” was the wrong question the entire time. The right question was closer to: is our change failure rate going up or down, is lead time for changes actually shrinking, are we creating more rework than we are eliminating. Those are not new AI metrics. They are the same operational metrics good engineering teams have tracked for a decade. AI did not require a new scorecard. It required someone to actually read the old one.
HR and people teams are about eighteen months behind on learning this same lesson, and I would really prefer you skip the expensive tuition.
What should you actually do about this?
Not build a new dashboard. That is the temptation, and it is the wrong move. The move is smaller and less exciting, which is exactly why almost nobody is doing it: find the metrics your business already runs on. The ones finance already believes. The ones that predate the word AI showing up in your vendor conversations. Ask, deal by deal and workflow by workflow, whether those numbers have actually shifted since AI entered the picture, and whether you can defend that connection to someone whose job is to be skeptical of you.
Some of that work is straightforward. A lot of it is not, because it requires you to have documented a “before” picture, which is something almost nobody bothers to do until it is too late. Connecting a lagging metric like margin contribution to a specific AI workflow six months later takes more than a hunch and a good story. There is a real difference between an operational metric moving because of AI and an operational metric moving because your best recruiter finally cracked LinkedIn, and most companies have no way to tell those two apart.
That is the real hard part. Not picking metrics, but building the discipline and the baseline to know, with some confidence, which of your numbers AI actually touched. In the AI-readiness audits I run at livingHR, this is the single most common gap: a company with a real AI budget and zero record of what its numbers looked like the day before that budget got approved.
Is this really a story about AI failing?
The 95 percent statistic is not really a story about AI failure. It is a story about measurement failure, with AI serving as the scapegoat. Companies that skipped establishing a baseline and relied on adoption metrics in place of real business outcomes will struggle to justify their AI investment at renewal time, regardless of whether the tool delivered anything real.
You do not need another AI-specific number. You need an honest look at the numbers you already tracked, and a clear-eyed read on whether they changed, and why. That “why” is where most people get stuck, and it is worth a real conversation rather than another dashboard. If you are staring at your own metrics wondering whether AI is actually moving the needle or just making noise, that conversation is the one to have, and it is the kind of readiness work livingHR's AI-People Solutions team does.
FAQ
What did MIT's AI ROI report actually find?
MIT Project NANDA's “The GenAI Divide: State of AI in Business 2025,” published last summer, found that 95 percent of enterprise generative AI pilots showed no measurable impact on profit and loss, while about 5 percent of integrated pilots produced real, significant value.
Why do AI adoption metrics fail to show real business impact?
Adoption metrics like usage rate, prompts per employee, or training completion measure activity, not outcomes. They tend to climb no matter what because using a new tool is easy; they do not tell you whether a real business metric, like cost per hire or error rate, actually moved.
What should companies measure instead of AI adoption rates?
The business metrics they already track and trust: cost per hire, time to fill, turnover, error rates, revenue per employee, and similar operational numbers, compared against a documented baseline from before the AI tool was introduced.
Why is a “before AI” baseline so important?
Without a baseline, there is no way to tell whether a metric changed because of the AI tool or for an unrelated reason, like a strong hire or a demand shift. Most companies skip this step and then have no way to defend their AI investment later.
Does a high AI adoption rate mean the tool is working?
Not by itself. High adoption shows people are using the tool, which is a different question from whether the tool moved a real business outcome.