TL;DR
Before an AI pilot starts, write down what work the tool may do, what improvement would justify using it and which failures would stop deployment.
Measure the work after people check and correct the output. A fast first draft can still leave the team with more work.
If the test changes, keep the original result. Run another test under the new conditions before calling the pilot a success.
If you can’t show me what would have made your AI pilot fail, I don’t know what “passed” means.
A demo can show something useful. People can like the tool. The team can learn enough to justify another experiment. None of that tells the executive approving deployment what the tool has actually earned permission to do.
The trouble starts when we treat those different decisions as one. Someone asks whether the technology looks promising. A few weeks later, the answer has become permission to put it into a workflow people depend on. The pilot never tested that decision.
On August 7, 2026, the National Institute of Standards and Technology, NIST, announced a draft framework for evaluating AI systems and invited comments through October 6, 2026. It starts evaluation with the organization’s goals and the people who need the results. That’s a useful place to start. The harder management job is agreeing on what result would justify deployment before the team knows whether the tool can produce it.
The demo can choose the test for you
Consider a hypothetical pilot for an AI assistant that drafts incident handoff summaries. A responder provides the incident records. The assistant produces a summary for the next person taking over.
The demo looks good because the output is readable and arrives quickly. Someone suggests measuring how long the assistant takes to produce it. That’s easy to count, so it becomes the pilot’s success measure.
But the receiving responder needs an accurate handoff. They need to know which systems were affected and what actions were actually taken. If someone has to reopen the records, correct invented details and rewrite the summary, the assistant’s speed tells us very little about the time saved.
I covered that measurement problem in The 51-Point Gap: Your People Got Faster. The Work Didn’t.. Here, there’s an earlier decision to make. Who gets to define “done” before the pilot produces an attractive number?
It shouldn’t be decided solely by the person running the demo or the executive who wants the purchase to work out. The person responsible for the workflow has to help define the result. Otherwise, the team can pass a test that the people receiving its work would never have chosen.
Write the deployment decision before the pilot
In The Problem Statement Problem, I wrote about executing against a brief whose assumptions nobody challenged. An AI pilot needs that same scrutiny, followed by something more specific: written conditions for the decision it will support.
For the handoff example, I’d start with a narrow proposal: use the assistant to draft summaries that a responder checks before sharing. The assistant doesn’t send the summary or decide which response action to take.
Then I’d ask that person to agree to a short decision record with the sponsor. It could look like this:
What would we authorize after a pass? Drafting handoff summaries in the named workflow, with a responder checking each one.
What improvement would justify using it? A stated reduction in time to a checked, usable handoff compared with the current process, without sacrificing accuracy. Count checking and corrections. Record both elapsed time and the time people spend doing the work.
Which failures would stop deployment? Name them: claiming a responder isolated a system when they didn’t, identifying the wrong affected system or disclosing information to someone who shouldn’t receive it. Specify what would block deployment even if the summary were later corrected.
What will the test include? Cases chosen separately from the demo, including incomplete records and conflicting information. Use the data access planned for deployment.
Who makes the decision? Name the person responsible for the workflow and the sponsor. Record who can approve a change to the test or accept a remaining risk.
There isn’t a percentage in this example because I don’t know the team’s workload or the consequences of a bad handoff. Those facts should determine the threshold. A round number chosen without that context wouldn’t tell the executive whether the tool was ready to use.
Write down the comparison process too. Have experienced responders agree on what each test case’s records establish and what remains uncertain. Judge both the current process and the AI-assisted process against that standard. Where practical, hide how each summary was produced from the responder judging it. If today’s handoffs are incomplete, beating them on speed alone wouldn’t establish that the new process produces usable handoffs.
This is my recommendation for running the pilot. NIST’s draft doesn’t prescribe this decision record or require a particular approval process.
Give the test a real chance to fail
A test built from the cases that made the demo look good can confirm very little. Include the conditions the tool will encounter when people are busy and the records are imperfect. Keep some cases out of development so the team has to show that its changes work on more than the examples it used to make them.
NIST’s September ARIA Evaluation Planning Manual describes combining three kinds of testing. Model testing uses prepared inputs to check the system’s responses. Red teaming asks people to try to make it behave badly. User testing examines how it performs when people use it for its intended task.
For our handoff pilot, test whether the responder spots a wrong system name and checks a confident statement against the record. Measure how much time that checking takes. A requirement for human review only helps if the review can catch the errors that matter.
Keep the individual failures visible. A high overall score can hide the one kind of error that makes the workflow unacceptable. In Evaluating Agents You Can’t Trust Yet, I examined different tests for different agent roles and looked beyond the aggregate score. That article also described deploying after an automated HOLD. Here, I’d make the record explicit: approving an exception doesn’t change the failed test result.
A small test that finds no serious errors also doesn’t prove that serious errors won’t occur. Record how many cases you tested and what kinds of cases they were. If a consequential failure would be rare, the owner needs to know that the pilot may not have had much opportunity to encounter it.
That can justify more testing or a smaller initial deployment. It can’t justify quietly replacing “we didn’t see it” with “it doesn’t happen.”
Keep the original result when you change the test
Early experiments should help a team learn. They may reveal that the proposed workflow is wrong, the comparison is weak or the tool needs a different task. Change the plan when the evidence warrants it.
But keep a record of what changed and why. An experiment used to develop the tool serves a different purpose from a test used to approve it.
Suppose the assistant fails on incidents with conflicting records. The team decides it will only draft summaries when the records agree. That might be a sensible restriction. It also changes what the tool would be allowed to do.
The report should retain the first result, describe the narrower use and show the result of a new test. Include how often the restriction would exclude real incidents and who would handle those handoffs. Count that work when estimating the benefit. Otherwise, the executive may approve a tool expecting help across the workload when the evidence supports only a small part of it.
The same applies if the team changes the model, its instructions or the review process. Those changes may fix the problem. Test that claim on fresh cases rather than treating the corrected development examples as proof.
If the sponsor wants to proceed despite a failed condition, record the exception. Name who had authority to accept it, what limits apply and what would end the exception. Calling that a pass would erase information the next owner needs.
A pass should name what it authorizes
In March, NIST and the General Services Administration announced work on AI evaluation for federal procurement, including evaluation before deployment and monitoring afterward. That work gives government buyers another reason to ask what a test establishes before a purchase. For our pilot, monitoring during deployment should help the owner catch problems the test missed. It doesn’t supply evidence missing from the current decision.
A successful handoff pilot could support using the assistant in the tested workflow with the tested review process. It wouldn’t establish that the assistant can write customer notifications, operate with broader access or send summaries without review. Each change creates a decision the original pilot may not have tested.
Before approving your next AI pilot for deployment, ask for the written conditions and the results against them. Read the failures and any changes to the test. The person responsible for the workflow should be able to explain what the evidence supports without relying on the enthusiasm of the project team.
If the conditions were never written, the pilot may still have taught you something useful. Record what you learned, then agree on the test needed to make the deployment decision.
Resources
NIST, The TEVV-Athlon Framework for Evaluating AI Systems. Initial public draft announced August 7, 2026; comment deadline October 6, 2026. Full draft, NIST AI 200-2. The framework starts with organizational objectives and develops an evaluation suited to the intended use.
NIST, ARIA Evaluation Planning Manual: Elements of ARIA-style AI Evaluations, September 18, 2026. Describes evaluations combining model testing, red teaming and user testing.
NIST, CAISI signs MOU with GSA to boost AI evaluation science for federal procurement, March 18, 2026. Announces work on evaluations before deployment and monitoring after deployment.


