AI Readiness Assessment: Evidence to Gather Before a Pilot

An AI readiness assessment examines whether an organization can test or operate a particular AI use case with the people, data, systems and controls it has available. A useful assessment produces a documented decision, the evidence behind it and a list of gaps with owners. The scope matters: readiness to test a drafting assistant is different from readiness to let software act on customer accounts.

This guide provides an evidence checklist you can adapt to that decision. It does not assign an organization a maturity score, predict a return on investment or certify it as ready for every application.

NIST’s AI Risk Management Framework provides a public reference for considering context, accountability, measurement and risk treatment. Its four functions are Govern, Map, Measure and Manage. NIST describes the framework as voluntary and adaptable; its Playbook is not a mandatory checklist. The assessment procedure below is editorial guidance, not a NIST scoring instrument. NIST AI RMF 1.0, NIST Playbook.

Public sources checked on 14 September 2026. Examples below are illustrations, not reported client results.

Define the decision before assessing readiness

Write a short description of the proposed use before collecting evidence. Specify:

  • The task, people affected and current way of doing the work.
  • What the AI would read, generate, recommend or change.
  • The users, locations and systems included in the proposal.
  • Who would review outputs or authorize actions.
  • The result you want to improve and how you would measure it.
  • What would make the proposal unacceptable.

For example, “Help support agents draft replies using approved product documentation” defines a smaller decision than “Automate customer support.” The second description leaves unanswered whether the system can send messages, change subscriptions or issue refunds. Record those permissions explicitly.

This focus on purpose and operating context follows NIST’s Map guidance. It also makes the assessment reviewable: another person can check whether the evidence actually covers the proposed use. NIST Playbook: Map.

AI readiness checklist: what evidence to gather

The following groups organize evidence for review. They are not weighted dimensions, maturity levels or universal technology requirements. Add checks that follow from the specific task and remove irrelevant ones with a written reason.

AreaEvidence to inspectDecision it informs
Workflow and baselineProcess description, sample cases, current quality and time records, exception handlingIs the problem defined well enough to test an improvement?
Data and permissionSample inputs, source owners, quality checks, access approvals and permitted usesCan the proposed system use the required information in the intended setting?
System integrationProposed data flow, permissions, interface tests and vendor configurationCan the system perform the bounded task without unintended access or actions?
People and accountabilityNamed process owner, users’ task knowledge, review responsibilities and escalation routeCan people operate the system, challenge outputs and make the release decision?
EvaluationRepresentative cases, expected outcomes, failure categories, acceptance criteria and test recordsCan the team distinguish an acceptable result from a convincing but incorrect one?
Operation and exitMonitoring plan, support ownership, running-cost estimate, fallback and shutdown procedureCan the organization sustain the proposed use and stop it when necessary?

For each item, distinguish evidence inspected from information reported in an interview. A policy document can establish a rule; a test record can show whether a particular configuration followed it. Mark an untested control as untested, even when its owner expects it to work.

Examine the parts that could block the use case

Data quality and access

Use examples from the workflow: incomplete records, older documents, different formats and cases that require escalation. Record which inputs are permitted, who owns them and how changes will reach the system. A successful test on a clean sample leaves the messy cases unresolved.

Assess the data path the application will actually use. For a purchased tool, inspect its configuration and the information users will submit. For a system that retrieves internal documents, check document permissions and update handling. A centralized data platform or a custom model-training environment is not a requirement of this checklist; justify any such investment through the proposed application.

Human review and responsibility

Name the person who owns the business outcome and the people who can approve access, accept risk or stop a release. Ask users to demonstrate how they would check an output and report a problem. If review is part of the design, include its time and workload in the assessment.

NIST’s Govern guidance supports clear responsibilities, appropriate training and the ability to question system decisions. It does not specify that every company needs the same committee or team structure. NIST Playbook: Govern.

Evaluation that matches the work

Write acceptance criteria before a pilot. Include failures that matter to the people using or affected by the system, then decide what evidence would be sufficient to permit the next step. Retain failures and ambiguous results rather than selecting only successful demonstrations.

For generative AI, check the factual basis of outputs and the limits of the tested conditions. NIST’s Generative AI Profile cautions against generalizing capability from narrow or anecdotal tests and recommends checking generated sources and citations. A polished answer still needs an appropriate accuracy check. NIST Generative AI Profile, actions MS-2.5-001 and MS-2.5-003.

Set measures for the use case instead of importing an expected productivity percentage. The published Generative AI at Work study examined a support assistant at one business-software company and found effects that varied with workers’ experience and skill. That study informs what to examine; it does not predict the outcome for a different workforce or task. Brynjolfsson, Li and Raymond, 2025.

How to conduct the assessment

  1. Agree on scope and ownership. Identify the decision being made, who makes it and who needs to contribute evidence. Include people doing the work, not only sponsors.
  2. Inspect available records. Use current documents, configurations, sample inputs and operating data. Record dates and versions so later changes are visible.
  3. Resolve contradictions. If an interview says access is restricted but the test account can retrieve excluded records, document that discrepancy and its effect on the proposal.
  4. List missing evidence. For each gap, specify the test, document or decision needed, its owner and any dependency. A lack of evidence is not a passing result.
  5. Make the bounded decision. Approve the next test, hold a proposed release, revise the design or stop the proposal. Record conditions and the event that triggers another review.

There is no need to fill every box with a number. NIST’s Measure guidance includes documenting risks that cannot yet be measured. Use that distinction to keep unknowns visible instead of hiding them inside an average. NIST Playbook: Measure.

Example: assess a support drafting assistant

The following fictional example shows how findings can lead to different decisions for different scopes. It contains no measured performance, customer data or delivery estimate.

Proposed capabilityEvidence in the illustrationRemaining issuePossible assessment decision
Draft from approved product documentationA document owner has identified the permitted source setEvaluation cases still need expected answersPrepare an offline test
Use customer account informationRequired fields are knownAccess permissions and data handling remain unverifiedHold this part of the proposal
Send a reply without human reviewDrafting is the only evaluated behaviorNo evidence covers autonomous sending or escalationKeep sending outside the test scope
Expand to another product lineThe current source set covers one productNew documentation and exception cases have not been checkedAssess the additional scope separately

The useful output is the reason for each decision. An overall readiness score would not resolve the specific access and evaluation gaps in this example.

Turn findings into an action plan

Create a short decision record containing the use case, approved scope, evidence reviewed, unresolved issues, accountable owner and next review trigger. Put supporting records behind it so reviewers can inspect the basis for the decision.

For every action, write what completion looks like. “Improve data readiness” is difficult to verify. “Check the evaluation set against current product documentation and have the document owner resolve conflicting answers” describes a task that can be reviewed. Sequence the work by dependency: an evaluation that requires restricted data cannot begin until its access question is resolved.

Set dates from the work involved, access lead times and available people. You can use a planning window to organize the backlog, but the end of that window does not by itself justify a pilot or release. The AI adoption roadmap explains how to carry these decisions into implementation.

For an assessment covering several use cases, keep a shared list of reusable capabilities and separate decisions for each application. A common identity system may support several projects; approval for one use of data does not establish permission for all of them. Use the AI governance guide to organize oversight and the AI ROI guide to examine cost and benefit assumptions.

Questions or factual corrections can be submitted through the contact page.

Frequently Asked Questions

What does an AI readiness assessment measure?

It examines evidence relevant to a defined use case: the task, data, system access, people, evaluation and operation. State the intended decision first, then select checks that help determine whether that next step is justified.

What is a good AI readiness score?

This guide uses no composite score. A decision should identify which conditions are met, which remain unverified and what scope the evidence supports. A number cannot substitute for resolving a missing access approval or an untested failure condition.

How long does an AI readiness assessment take?

Estimate the duration after defining the scope and evidence needed. Document availability, system access, stakeholder capacity and testing requirements determine the work. Record those dependencies instead of assuming a standard number of weeks.

Who should take part?

Include the process owner, people who do the work, the relevant data and technical owners, and whoever can evaluate the proposal’s risks. Involve additional specialists when the application requires their expertise. Assign decision authority explicitly.

Does readiness require a centralized data platform or an AI team?

Assess what the proposed application needs. Identify the necessary data, skills and operating responsibilities, then check whether existing arrangements cover them. Make any platform purchase or staffing change part of a specific, justified gap action.

Can we assess readiness internally?

An internal team can gather evidence and document decisions. Consider independent review when important expertise is missing or the people evaluating a proposal also have strong incentives to approve it. Record the reviewer, evidence and reasoning either way; an external assessor is not automatically more accurate.

When should we reassess?

Revisit the decision before expanding its scope and when relevant conditions change, such as data sources, models, permissions, users or observed failures. Set a review cadence suited to the application as well as these event-based checks. NIST’s Manage guidance treats continued use as a decision informed by ongoing evidence. NIST Playbook: Manage.


Related reading


From strategy to systems in production

Book a briefing with a Transformation Lead — we confirm scope and recommend the right place to start.

Book a briefing