Thought Leadership

AI Readiness Assessment: What It Should Actually Measure

Written by Anthony Onesto | Sep 10, 2026, 7:12:51 PM

A real AI readiness assessment measures four things: proficiency distribution by function and role, governance maturity, manager coaching capability, and workflow readiness. License counts and seat activation are not on that list, and a score built on them will not tell you anything you can act on Monday morning. 

Here is the finding that should reorder how you think about this. HiBob's 2026 AI skills research, built on 1,200 AI decision-makers across six market groupings, asked which everyday AI behaviors matter most to the business. The top two came back tied at 52 percent: proactively reviewing output quality, and documenting workflow decisions. Prompt fluency did not lead. Output speed did not lead. The behaviors organizations value most cluster around accuracy, documentation, and reliability. 

That inverts the premise almost every readiness assessment is built on. We have been measuring how fast people can produce with AI. The market is saying it cares more about whether people can be trusted with it. Those are different capabilities, they live in different parts of your org chart, and only one of them shows up in a usage dashboard. 

I raised this at the CPO Collective roundtable livingHR ran with HiBob in July, “AI Skills Defined: A Practical Framework for CPOs Building Their AI-Ready Workforce,” with Dr. Ken Matos. The chat kept returning to the same problem from different directions. People had numbers. Nobody had a decision. 

Why do most AI readiness assessments measure the wrong thing? 

Because seat data is easy to pull and behavior data is not.

Your HRIS or your vendor console can tell you that 94 percent of the commercial team activated a license and 61 percent logged in last week. It cannot tell you whether anyone checked an output before sending it to a client. So the assessment gets built around what the system already knows, and the resulting score describes your procurement history rather than your workforce. 

The second failure is aggregation. A company-level readiness number averages a function running at 80 percent capability against one running at 15 percent, and reports something in the fifties. That number is technically accurate and operationally useless. You cannot fund a fifty. You can fund a gap. 

McKinsey's State of AI research puts a number on the cost of measuring the wrong thing: adoption is nearly universal, but only 39 percent of organizations report any enterprise-level EBIT impact. The separator among the ones seeing returns is workflow redesign, which is exactly the variable a license-count assessment cannot see. 

What is the difference between a surface-level and a decision-grade AI readiness assessment?

Dimension

Surface-level assessment

Decision-grade assessment

Adoption

Seats activated, weekly logins

Observable behaviors demonstrated in real work

Proficiency

Self-reported confidence, 1 to 5

Evidence of output verification, data judgment, documentation

Unit of analysis

Company-wide average

Distribution by function and role

Governance

Whether an AI use policy exists

Whether decision rights, logging, and escalation actually operate

Managers

Whether managers completed training

Whether managers can evaluate a skill they may not personally hold

Work design

Which tools are approved

Which workflows have been rebuilt around the tool

Output

A single readiness score

A ranked list of gaps with owners and a sequence

What it lets you decide

Renew or cancel licenses

Where to invest, what to redesign, what to govern first

The right-hand column is harder to build. It is also the only column that survives contact with a CFO. 

How do you measure AI proficiency distribution by function and role?

Stop asking people how proficient they feel. Confidence and capability correlate weakly, and the gap runs in both directions. Quiet, competent people underrate themselves, and enthusiastic novices do the opposite.

Measure behaviors instead, in the work people already do. Four levels are enough. 

Level

What it looks like in practice

Unaware

No meaningful use, or use that is hidden from the manager

Assisted

Uses AI for drafting and summarizing, accepts output largely as written

Verified

Checks output against a source before acting, knows which data never goes into an external tool

Redesigning

Restructures how the work is sequenced, documents AI-assisted decisions, teaches the pattern to others

Good, at the organizational level, is not everyone at Redesigning. Good is knowing where each function actually sits and whether that matches what the function needs. Your legal team living at Verified may be exactly right. Your marketing team living at Assisted is a growth problem. Your finance team sitting at Assisted while touching material numbers is a risk problem. Same score, three different responses. 

Run this at role level, not department level. Two people with the same title and different scopes will land in different places, and role is the unit you eventually hire, promote, and pay against. 

How do you measure AI governance maturity? 

Ask one question: what happens next?

Someone pastes customer data into an unapproved tool. What happens next? A model produces a confidently wrong number that reaches a board deck. What happens next? An employee wants to use a tool that is not on the approved list. What happens next?

If the answer to all three is “we have a policy,” you have a document, not a capability. Governance maturity is measurable across five components, and each one is either operating or it is not: 

  • Decision rights. Named owners for tool approval, data classification, and exception handling. 

  • Logging. A record of AI-assisted decisions that someone could actually retrieve six months later.

  • Human-in-the-loop thresholds. Written rules for which decisions require review, by risk tier rather than by department preference.

  • Data classification. People know which categories of information never leave the building, and can say so without opening a document.

  • Escalation. A path that a mid-level employee would use without fear of being made an example of. 

Score each as absent, documented, or operating. The distance between documented and operating is where most organizations actually sit, and it is the distance that matters. Gartner's 2026 CHRO priorities make the point that most organizations and vendors are still experimenting, and that CHROs need an HR-specific AI strategy rather than a borrowed enterprise one. Governance is where that distinction becomes concrete. 

How do you measure manager coaching capability? 

This is the one most assessments skip entirely, and it is the one most likely to stall you.

The same HiBob research found that direct managers are the group organizations most expect to build AI capability across their teams, and that only 36 percent of respondents see those managers as highly prepared to do it. Read that as a sequencing instruction. You can run every upskilling program in the catalog, and if the manager cannot reinforce it in a one-to-one, it decays inside a quarter. 

Three things to measure, none of which require a survey about confidence: 

Can the manager evaluate the behavior?

Give a manager two AI-assisted work samples, one where the output was verified and one where it was not, and see whether they can tell the difference and articulate why. Managers who cannot distinguish good AI-assisted work from plausible AI-assisted work cannot coach it, no matter how much they support the initiative.

Does the manager have shared vocabulary?

If your organization has four names for the same proficiency level, coaching conversations become negotiations. Measure whether managers use the same words for the same behaviors, because a taxonomy nobody speaks is a spreadsheet.

Is it safe for the manager to be learning too?

People conceal AI use when they think it makes them look replaceable, and that includes managers. Hidden use is unmeasurable use, which quietly corrupts every other number in your assessment. Measure whether managers admit uncertainty in front of their teams. It is a soft signal with hard consequences. 

How do you measure workflow readiness?

Proficiency without workflow change produces faster versions of processes that were already wrong.

Workflow readiness asks whether the work itself has been rebuilt, and it has three observable markers. First, has any end-to-end process been redesigned with AI as the default path rather than an optional accelerator? Second, do the handoffs still assume a human step that no longer adds anything? Third, does the output of the AI-assisted step feed a system, or does someone retype it into one?

Pick your two or three highest-value workflows and assess those. An organization-wide workflow audit is a consulting engagement, not an assessment, and the returns concentrate in a small number of processes anyway. 

What should the output of an AI readiness assessment let a CPO decide?

If the assessment produces a number and nothing else, it failed. A decision-grade output lets you answer four questions in a room with your CEO:

  • Where do we invest first? The function with the widest gap between current proficiency and required proficiency, weighted by that function's contribution to the business. 

  • What do we govern before we scale? The risk tier with the largest volume of AI-assisted decisions and the weakest logging. 

  • Which managers get support and which get accountability? Different interventions, and the assessment should tell you which people need which. 

  • What do we stop? Tools, pilots, and training tracks that are producing activity without capability movement.

That last one is where the money is, and it is the one an adoption dashboard will never surface for you.

livingHR's AI-People Readiness Index is built to produce that four-part output rather than a score, which is why it breaks results down by function rather than reporting a company-level number. 

What will an AI readiness assessment not tell you?

It will not tell you what a gap is costing you. It will not tell you whether closing it is worth the investment relative to everything else on your plate. And it will not rank your gaps against each other, because ranking requires business context the instrument does not have.

I say this because the failure mode I see most often is a CPO with an accurate diagnostic and no path from it to a budget conversation. The assessment tells you where you are. Deciding what to do about it is still a human judgment, informed by data, made by someone who understands the business. That combination, humans and machines together rather than one substituting for the other, is the whole argument for doing this properly in the first place.

For the structural work of connecting a readiness gap to org design, see HR as the Architecture Behind AI-Ready Organizations and Workforce Planning in an AI-First Environment. 

Frequently asked questions

What is an AI readiness assessment?

An AI readiness assessment measures an organization's capability to use AI productively and safely, across proficiency, governance, manager capability, and workflow design. It is a capability diagnostic, not a technology audit.

What is the difference between AI readiness and AI proficiency?

AI proficiency describes an individual's demonstrated skill with AI tools. AI readiness describes whether the organization around that individual, including its governance, managers, and workflows, can turn that skill into business value.

How long should an AI readiness assessment take?

A behavior-based assessment should take an individual under ten minutes. Organization-level analysis and interpretation typically runs two to four weeks depending on how many functions you cover.

Who should own the AI readiness assessment?

HR should own it, with IT and legal supplying governance and risk input. HR is the function that already holds role architecture, performance, and development, which is where readiness data has to land to be useful.

How often should you reassess?

Every six months during active AI adoption. Proficiency distribution moves faster than most workforce metrics, and a year-old readiness score is a historical document.

Where to start

Pick two functions. Assess the four dimensions in those two. Compare the results against what each function actually needs rather than against each other.

The AI-People Readiness Index gives you that breakdown in about five minutes per participant, scored by function rather than rolled into a single company number. Take the Index and see where your distribution actually sits.

Anthony Onesto is Principal of AI-People Solutions at livingHR. This piece draws on research presented at the CPO Collective roundtable with HiBob, July 2026.