A 90-day AI pilot in wholesale: scope, metrics and exit criteria
A practical plan for a time-boxed AI pilot in a wholesale business, with a defined scope, measurable targets and agreed conditions for stopping or scaling.

KEY TAKEAWAYS Summary by the editors
- A 90-day wholesale AI pilot should address one process, with a baseline, a small group of users, agreed metrics and written exit criteria for scaling, changing or stopping.
- McKinsey's 2026 State of AI survey found that 44 percent of respondents say AI is scaling across their enterprise and 37 percent attribute at least some EBIT impact to AI, so many organisations still lack measurable financial returns.
- The best first use cases in wholesale are narrow and data-ready, such as order status questions, product data checks or visit briefings.
- Data readiness, user adoption and the cost of review often decide the outcome more than model quality.
- A pilot that ends without a decision has failed, regardless of the technical result.
A good 90-day AI pilot in wholesale picks one narrow process, measures it against a baseline, involves a small group of real users and sets down in advance what result would lead to scaling, adjusting or stopping. The aim is a decision, not a demonstration.
Why run a time-boxed pilot?
McKinsey's 2026 State of AI survey (1,719 respondents in 97 countries, fielded May to June 2026) reports that 44 percent say AI is scaling across their enterprise, up from 38 percent a year earlier, and that 37 percent attribute at least some EBIT impact to AI, essentially unchanged. A year earlier, as summarised by Le New Black, the survey found that the majority of organisations were still experimenting or piloting. The picture suggests that many pilots do not reach measurable financial results. A fixed time frame forces a decision and limits cost.
A time box also protects the team. Open-ended experiments consume the attention of the people who know the business best, and they tend to be the first to be pulled back to daily work. Ninety days is long enough to see behaviour change and short enough that a sponsor can reasonably protect the time.
What should the scope be?
Choose one process with a clear owner, available data and a measurable output. Avoid broad aims such as improving the customer experience. Good candidates in wholesale include answering routine order and delivery questions, checking price lists and product data before release, preparing briefings for sales reps, and drafting follow-up messages after market-week appointments.
| Pilot | Main prerequisite | Typical metric |
|---|---|---|
| Order status assistant for customer service | Reliable order and shipment data | Share of queries answered correctly, handling time |
| Price list and product data checks | Documented rules, clean master data | Errors caught before release, review time |
| Rep visit briefings | Complete account and order history | Prep time saved, rep rating of usefulness |
| Re-order reminders for small accounts | Order history, consent to contact | Re-order rate versus control group |
Check data readiness early. Run a one-week review of the data the pilot needs: whether it exists, who owns it, how current it is and whether the AI tool is permitted to use it. Many pilots lose weeks to discovering late that the order history is spread over two systems or that a customer file includes restricted data. A short data review before day 15 avoids this.

How should the 90 days be structured?
- Days 1 to 15, define: choose the process, owner, users, baseline measures and success and exit criteria.
- Days 16 to 30, prepare: check data, set access rights and decide how errors will be logged.
- Days 31 to 60, run: use the tool with a small group, collect logs and feedback weekly and fix issues.
- Days 61 to 80, measure: compare with the baseline and, where possible, a control group.
- Days 81 to 90, decide: hold a review against the exit criteria and record the decision.
Document the baseline carefully. If customer service currently takes an average time to answer an order query, record that with the same definition you will use at the end. Without a baseline, any improvement is an impression, and an impression is not enough to justify scaling a tool across a sales organisation.
Which metrics matter?
Use a small set covering three areas. Outcome metrics reflect business results, such as time to answer a customer, errors reaching buyers or re-order rate. Quality metrics reflect accuracy, such as the share of answers verified as correct, and the number of serious errors. Adoption and cost metrics show whether people use the tool and what it costs, including subscription, usage fees and the time spent on checking its work.
The 2026 McKinsey survey notes that about 20 percent of respondents say operating costs, including tokens, have constrained their AI use. Include running costs in the pilot, not just set-up effort.
Include a measure of the effort needed to check the tool's output. A system that is 90 percent accurate sounds good until the team realises it must still read every answer to find the other 10 percent. Review time can wipe out the gain, so log it from the first week.
What are sensible exit criteria?
- Scale: the tool meets the quality threshold, shows a measurable benefit over the baseline and users keep using it voluntarily.
- Adjust: the benefit is plausible but data, process or training issues are clear and fixable within a defined period.
- Stop: quality is below the threshold, benefit is not measurable or cost of oversight exceeds the gain.
- Hard stops: a serious error, such as exposing confidential data or sending wrong prices to buyers, triggers an immediate review.
Write the thresholds as numbers before the pilot starts, for example the minimum share of verified correct answers and the maximum acceptable review time per answer. Setting them afterwards invites the common temptation to define success by whatever the pilot happened to achieve. Agree also who has the authority to stop the pilot early if a hard stop is triggered.

What commonly goes wrong?
The most frequent failures are unclear ownership, no baseline, data problems found late, and users who were never asked. Another is a pilot that quietly continues because nobody wants to take the decision. Name a sponsor who will make the call on day 90. Finally, be honest about build or buy: the McKinsey survey reports that 32 percent of respondents decided against buying at least one software product because agentic coding tools let them build it in house, but this adds maintenance and risk that should be costed.
Treat staff concerns as part of the pilot. Reps and customer service teams may fear that the tool is intended to replace them, which affects how they use it. The McKinsey survey reports that 39 percent of respondents expect AI to reduce headcount in the coming year, so the concern is not irrational. Be clear about the intent of the pilot, what will change in roles if it scales, and how feedback will be used.
Finally, plan for the day after. If the decision is to scale, the cost, support and data work needed for a wider rollout are usually larger than the pilot, and they should be estimated before the review. If the decision is to stop, record what was learned so that the next pilot does not repeat the same mistakes.
Frequently asked questions
How long should an AI pilot in wholesale last?
Around 90 days is enough to define, run and evaluate a narrow use case. Longer pilots tend to lose focus unless they have explicit milestones.
Which wholesale AI use case is best for a first pilot?
One with clean data, a clear owner and measurable results, such as order status questions, price list and product data checks or briefings for sales reps.
What are exit criteria for an AI pilot?
Pre-agreed conditions for scaling, adjusting or stopping, based on quality, measured benefit, user adoption and cost, plus hard stops for serious errors.
Why do many AI pilots fail to scale?
Common reasons include unclear ownership, missing baselines, poor data, costs of review and a lack of a decision at the end. McKinsey's 2026 survey shows only a minority attribute significant EBIT impact to AI.
One edition every weekday morning. Read in five minutes. Free for industry professionals.




