ProcessStrategic Technology Consulting
How an AI readiness assessment works, step by step
A useful AI readiness assessment does not end in one maturity score for the whole organisation. It ends in a ranked list of specific use cases, each rated on the data, platform, governance and skills it depends on, with the work each needs before it can start. Here are the working steps, the evidence behind each rating and how to act on the result.
On this page
- The decision an AI readiness assessment should inform
- From candidate use cases to a funded sequence
- Six working steps and what each one produces
- Rating each use case on five dimensions
- A hypothetical insurer works through three candidates
- What to do with each rating
- Evidence to gather before the first interview
- Five ways readiness assessments mislead
- Questions and answers
- Sources
The decision an AI readiness assessment should inform
Most organisations commission a readiness assessment because someone has to approve money. A board wants to know whether next year's AI budget will produce working systems. The assessment makes that approval defensible: which use cases to fund first, which to defer and what must be fixed before any of them can start.
That framing rules out two common outputs. One is a maturity score for the whole company, which averages away the differences that matter: finance may hold clean, well-governed ledgers while customer service runs on free-text tickets nobody owns. The other is a long list of ideas with no view on feasibility. Neither tells a sponsor what to fund.
In ColdAI's consulting practice, a readiness assessment usually runs as a sprint assessment of two to four weeks and rates data availability and quality, infrastructure, governance and skills gaps for each use case and business unit1. The method below shows how those ratings are reached.
From candidate use cases to a funded sequence
- Frame the decision
Agree what will be funded, by whom and by when, so every later rating answers that question.
- Inventory candidates
Collect proposed use cases from business units and vendors in one consistent format.
- Evidence the data
Inspect samples, lineage and access rights for the data each candidate needs.
- Test the platform path
Check where models would run, which systems they must reach and who operates them.
- Map governance duties
Identify the obligations and internal policies each candidate would trigger.
- Score and sequence
Rate each candidate on every dimension and order them by readiness and value.
Six working steps and what each one produces
Frame the decision
Start with the approval the assessment must support: the budget line, the deadline and who signs. Write down what makes a use case fundable, such as a named process owner and a measurable change in cycle time or error rate.
Inventory the candidates
Gather every proposed use case, including vendor pitches and shadow pilots, and describe each the same way: the task it changes, the people affected, the systems involved and how the work is done today.
Evidence the data
For each candidate, pull a real sample of the records it would use. Check whether the fields exist, how often they are empty or contradictory, where they come from and whether you have the right to use them for this purpose.
Test the platform path
Trace how a working system would be built and run: where inference happens, which systems it must read from or write to, how access is enforced and which team is on call when it fails.
Map governance duties
Establish which external obligations and internal policies apply to each candidate before scoring, because they change the build. A system placed on the EU market may fall under the EU AI Act (Regulation (EU) 2024/1689)2; many organisations also align their risk process with the NIST AI Risk Management Framework3 or a management system under ISO/IEC 420014.
Score and sequence
Rate every candidate on each dimension, take the lowest rating as its overall readiness, and order the list by readiness and expected value. For each deferred candidate, record the specific work that would move it up.
Rating each use case on five dimensions
| Dimension | Not ready | Ready to pilot | Ready to scale |
|---|---|---|---|
| Data availability | Key fields missing or held in documents nobody can export | Fields exist and a representative sample can be extracted | Data flows continuously through a documented pipeline |
| Data quality and rights | Unknown lineage, or consent and licence terms unclear | Known gaps with a workaround; rights confirmed for a pilot | Quality monitored and rights confirmed for production use |
| Infrastructure and integration | No agreed place to run models or reach source systems | A sandbox with read access to the systems involved | Approved runtime, write access controls and an on-call owner |
| Governance and accountability | No owner for the decision the system would influence | Owner named and obligations mapped for a bounded pilot | Controls, monitoring and escalation agreed and tested |
| Skills and process ownership | No one to maintain it after the vendor or project leaves | A team that can run a pilot with outside support | A team that can operate, evaluate and change it alone |
Most first projects should aim for the middle column on every row; scaling comes after the pilot has produced its own evidence.
A hypothetical insurer works through three candidates
What to do with each rating
- If
A use case is ready to pilot on every dimension.
ThenFund a bounded pilot with an evaluation set and success measure agreed before any build starts.
Agreeing the test in advance stops a demonstration deciding the outcome.
- If
One dimension is not ready and the fix is known.
ThenFund the prerequisite as its own small project with an owner and a date, and re-rate afterwards.
Building around a known gap usually costs more than closing it first.
- If
Governance is the weakest dimension.
ThenBring risk and legal into scoping now and narrow the use case until its obligations are clear.
Obligations discovered late tend to force a redesign rather than a sign-off.
- If
Several candidates fail on the same platform gap.
ThenTreat the shared gap as a platform investment and evaluate it on the value of everything it unblocks.
Charging it to one use case makes that use case look worse than it is.
Evidence to gather before the first interview
Five ways readiness assessments mislead
A single maturity score for the whole organisation
Early signalThe headline result is a number or a level with no named use cases behind it.
MitigationRate each candidate separately; describe overall maturity only as a pattern across those ratings.
Data judged from documentation
Early signalRatings cite data dictionaries and catalogue entries but no extracted samples.
MitigationRequire a real sample behind every data rating and record what it showed.
Governance left to the end
Early signalRisk and legal first see the recommendations at the final presentation.
MitigationMap obligations before scoring so they shape which use cases are proposed and how.
A vendor demonstration counted as platform evidence
Early signalThe platform rating rests on a demo run on the vendor's own data and infrastructure.
MitigationTest the path on your own systems and identity controls, even if only in a sandbox.
Contract constraints discovered after scoring
Early signalA top-ranked use case turns out to conflict with a cloud commitment, a data-processing agreement or a vendor exclusivity clause.
MitigationRead the contracts that constrain platform and data choices before rating infrastructure and governance.
Questions and answers
How long does an AI readiness assessment take?
At ColdAI it usually runs as a sprint assessment of two to four weeks from the signed brief, enough to rate a focused list of candidates properly. Its length depends more on how quickly data owners provide samples and access than on the number of interviews.
Do we need clean data before we start?
No. Finding out where the data is incomplete, contradictory or legally restricted is one of the main purposes of the assessment. What you do need is people who can extract representative samples and explain where the records come from. A candidate held back by data problems still gets a clear list of what to fix and who owns each fix.
Is the result a maturity score?
Not as the headline. The main output is a ranked list of specific use cases, each rated on data, platform, governance and skills, with the prerequisite work for every deferred candidate. A pattern across those ratings can describe overall maturity, but a score alone cannot tell a sponsor what to fund first.
Can our own team run the assessment instead?
Often, yes, if your architects and data owners have time and no stake in a particular answer; the method on this page is designed to be repeatable. Outside help is most useful when internal teams disagree, when vendors are competing for the same budget, or when the board wants a view from someone without a prior position.
What happens after the assessment?
The sponsor approves, amends or defers the sequence and names an owner for each funded use case. Prerequisite work, such as digitising documents or agreeing an operating owner, gets dates. The first pilot then starts against an evaluation set agreed in advance, and its results feed the next round of ratings.
Sources
- Strategic Technology Consulting: engagement models and deliverables — ColdAI
- Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) — EUR-Lex · checked 10 October 2026
- AI Risk Management Framework — National Institute of Standards and Technology · checked 10 October 2026
- ISO/IEC 42001:2023 — AI management systems — International Organization for Standardization · checked 10 October 2026