How to measure the ROI of AI in fashion
Most companies struggle to prove that AI pays off. A practical framework for fashion: baselines, control groups, the right metrics per use case and the full cost picture.
KEY TAKEAWAYS Summary by the editors
- Measuring the ROI of AI in fashion requires a baseline before the project starts, a control group during the pilot and a full cost count that includes integration, data work and human review time.
- MIT's NANDA initiative reported in 2025 that about 95 per cent of the generative AI pilots it studied had little to no measurable impact on profit and loss.
- In McKinsey's 2026 State of AI survey, 37 per cent of respondents attributed at least some EBIT impact to AI, unchanged from 2025, while only 6 per cent attributed 5 per cent or more.
- Gartner predicts that over 40 per cent of agentic AI projects will be cancelled by the end of 2027 because of escalating costs, unclear business value or inadequate risk controls.
- Leading indicators such as clicks or time saved are useful early, but the business case should be judged on outcome metrics such as margin, stock levels, conversion or cost per order.
To measure the ROI of AI in fashion, set a baseline before the project starts, compare the results of AI-supported work with a control group working the old way, and count all costs, including integration, data preparation, licences and the time people spend checking output. ROI is the measured improvement in a business metric, translated into money, divided by that full cost.
That sounds obvious, yet most companies do not do it. The result is a large number of pilots with enthusiastic anecdotes and no defensible numbers.
Why do so many AI projects fail to show a return?
The evidence on AI returns is sobering. MIT's NANDA initiative, in its 2025 report The GenAI Divide, found that about 95 per cent of the generative AI pilots it studied delivered little to no measurable impact on profit and loss, according to Fortune. The report also observed that more than half of generative AI budgets went to sales and marketing tools, while the largest returns were found in back-office automation.
Survey data points the same way. In McKinsey's 2026 State of AI survey, 37 per cent of respondents attributed at least some EBIT impact to AI, unchanged from the previous year, and only 6 per cent attributed 5 per cent or more of EBIT to AI. Gartner, meanwhile, predicts that over 40 per cent of agentic AI projects will be cancelled by the end of 2027 because of escalating costs, unclear business value or inadequate risk controls.
The common causes are familiar: no baseline, no control group, benefits counted as time saved but never converted into lower cost or higher output, and costs that grow after the pilot as usage scales.
Which metrics matter for fashion AI use cases?
Each use case needs one primary outcome metric that matters to the business, supported by a few leading indicators that show early whether things are moving in the right direction.
| Use case | Primary outcome metric | Leading indicators | How value is realised |
|---|---|---|---|
| Demand forecasting and allocation | Gross margin, excess stock at season end | Forecast error by style and size, stock-outs | Fewer markdowns, fewer lost sales |
| Product content generation | Cost and time per published product | Error rate, edits per text | Faster launches, less agency spend |
| Search and recommendations | Conversion rate, revenue per visitor | Click-through, zero-result searches | Higher sales from the same traffic |
| Customer service assistance | Cost per contact, customer satisfaction | Handling time, escalation rate | Lower service cost or higher capacity |
| Wholesale order processing | Cost per order, order error rate | Manual minutes per order | Capacity for more orders without new hires |
| Markdown and pricing | Full-price sell-through, margin | Discount depth, weeks of cover | Higher margin on the same volume |
Leading indicators are useful but should not be confused with returns. Zalando's assistant, built with OpenAI, recorded a 23 per cent increase in product clicks and more than 40 per cent more products added to wishlists after an upgrade, compared with the previous version, according to OpenAI. Those are meaningful engagement gains; whether they translate into revenue and margin is a separate question that each company must measure for itself.
How do you set a baseline and attribute impact?
- Measure the primary metric for the scope of the pilot over a representative period before the project starts, adjusting for seasonality where possible.
- Choose a comparable control group: similar stores, categories, markets or customer segments that continue with the current process.
- Run the pilot long enough to cover normal variation, which in fashion often means at least part of a season for planning use cases.
- Use controlled experiments (A/B tests) for digital use cases such as search, recommendations and content, where traffic can be split.
- Record what else changed during the pilot, such as promotions, price changes or new collections, so that effects are not wrongly attributed to AI.
Time saved deserves special scrutiny. If an AI tool saves each copywriter an hour a day, that only becomes a return if the team publishes more products, reduces agency spend or takes on other work. Business cases should state which of these will happen.
What costs belong in an AI business case?
- Licences, subscriptions or usage fees, including how they grow as volume increases.
- Integration with ERP, product information, e-commerce and other systems.
- Data preparation and ongoing data quality work.
- Human review of AI output, especially for customer-facing content.
- Training, change management and support for users.
- Monitoring, security, legal review and compliance work.
- The internal time of business experts who define, test and improve the solution.
When should a fashion company stop an AI project?
Stopping rules should be agreed before a pilot begins. Reasonable triggers include no measurable improvement in the primary metric against the control group, costs at full scale exceeding the expected benefit, error rates that require so much review that savings disappear, or low adoption by the intended users after training. Ending a project on these grounds is evidence of good management, not failure.
How should leaders report AI ROI?
A simple portfolio view works best: each use case with its baseline, current result, full cost to date, expected cost at scale and status (pilot, scaling, in production or stopped). Reviewing this quarterly lets leaders shift investment toward use cases that work. It also helps to separate one-off costs, such as integration and data clean-up, from recurring costs, such as usage fees and review time, because a use case that looks expensive in its first year can be attractive over three years, and the reverse is also possible. Over time, the share of AI spending that sits in proven, measured use cases is itself a useful indicator of whether the programme is creating value.
Frequently asked questions
How do you calculate the ROI of AI?
Measure a business metric before and during the AI project, compare the result with a control group, convert the improvement into money and divide by the full cost. Full cost includes licences, integration, data work, training and the time people spend reviewing AI output. Without a baseline and a control group, ROI figures are guesses.
Why do most AI projects not deliver ROI?
Common reasons are a missing baseline, tools that are not integrated into real workflows, benefits counted as time saved without changing cost or output, and costs that rise at scale. MIT's NANDA initiative found in 2025 that about 95 per cent of the generative AI pilots it studied had little to no measurable impact on profit and loss.
What is a good KPI for AI in fashion?
A good KPI is the business metric the use case is meant to change, such as margin and season-end stock for forecasting, conversion for search, or cost per order for wholesale order processing. Leading indicators such as clicks or forecast error help track progress early. They should not replace the outcome metric.
How long does it take to see ROI from AI?
Digital use cases such as search or content can show measurable results within weeks through A/B tests. Planning and forecasting use cases usually need at least part of a season to show effects on stock and margin. Business cases should state the expected timeline in advance.
One edition every weekday morning. Read in five minutes. Free for industry professionals.