Do AI stylists sell? The metrics to judge a conversational shopping assistant
Chat volume says little about whether an AI stylist earns its keep. Which metrics show real commercial impact, how to test them fairly and what published data does and does not prove.

KEY TAKEAWAYS Summary by the editors
- Whether an AI stylist sells can only be judged with a controlled comparison, because shoppers who choose to chat are usually more motivated than average.
- Adobe reported that in the 2025 US holiday season, visits referred by generative AI tools converted 31% better than other traffic and spent 45% longer on retail sites.
- AI referral data from external assistants is not the same as the performance of a retailer's own on-site stylist, and should not be used as a business case without own testing.
- Zalando reported six million users of its assistant in March 2026, four times the previous year, but did not publish conversion or revenue effects.
- A balanced scorecard for fashion should include incremental conversion, order value, return rate, contact rate, answer accuracy and cost per conversation.
AI stylists can sell, but published evidence is still thin and much of it measures traffic from external assistants rather than the impact of a retailer's own conversational stylist. The only reliable way for a fashion retailer to know is a controlled test that compares incremental conversion, order value and returns against a holdout group, alongside accuracy and cost metrics.
What does the published data say about AI and shopping conversion?
The most widely cited figures come from Adobe Analytics. For the 2025 US holiday season, Adobe reported that generative AI traffic to retail sites rose 693% year on year, that AI referrals converted 31% better than non-AI sources, and that AI-referred visitors spent 45% longer on sites and viewed 13% more pages. On Thanksgiving, AI referrals converted 54% better and on Black Friday 38% better.
These numbers are encouraging but describe a different thing from an on-site stylist. They measure shoppers who arrive from tools such as ChatGPT or Perplexity after doing research there. Those visitors may already be further along in their decision. Adobe's figures also cover all retail, not fashion specifically, and come from a vendor with an analytics product to sell.
Retailers' own disclosures tend to report adoption, not effect. Zalando said in March 2026 that its assistant had six million users, four times the number a year earlier, without publishing conversion or revenue impact.
Why are AI stylist results often overstated?
- Self-selection: shoppers who open a chat are typically more engaged, so comparing chatters with non-chatters exaggerates the effect.
- Attribution overlap: if the stylist recommends an item the shopper would have bought anyway, the order is counted but the value is not incremental.
- Ignoring returns: in fashion, a sale that comes back is not a sale. Higher conversion with higher returns can destroy margin.
- Short test windows: novelty effects can inflate early usage and results.
- Missing cost side: model and integration costs per conversation are often left out of reported results.

Which metrics should a fashion retailer track?
No single number captures the value of a conversational assistant. The table below combines commercial, quality and cost measures so that a rise in one metric cannot hide a fall in another. In fashion, net revenue after returns should carry the most weight, because conversion gains that are later returned do not create value and add logistics cost.
| Metric | What it shows | How to measure | Watch-out |
|---|---|---|---|
| Incremental conversion | Extra orders caused by the assistant | A/B test with holdout group that never sees the assistant | Compare by assignment, not by usage |
| Average order value | Basket building and cross-sell | Same test design | Discount-driven increases may hurt margin |
| Net revenue after returns | Real value in fashion | Track returns per test group over 30 to 60 days | Needs patience before reporting |
| Return rate and reasons | Fit and expectation quality | Return reason codes per group | Size-related returns are the key signal |
| Contact rate | Service deflection or extra workload | Tickets per order per group | Errors can create new contacts |
| Answer accuracy | Trust and legal risk | Weekly sampled review against catalogue | Set a minimum threshold |
| Cost per conversation | Economics | Model, hosting and staff costs divided by sessions | Rises with longer conversations |
How should an AI stylist test be designed?
- Randomly assign visitors (or logged-in customers) to a test group that can use the stylist and a control group that cannot.
- Analyse results by group assignment, including shoppers in the test group who never used the stylist, so that self-selection does not bias the result.
- Run the test long enough to cover at least one full returns window and avoid major sale periods distorting the outcome.
- Pre-define success thresholds for net revenue, return rate and accuracy before looking at results.
- Segment results by category and customer type, because a stylist may help with occasionwear but not with basics.
Do AI stylists reduce returns?
This is plausible but under-evidenced. In Adobe's holiday 2025 survey of US consumers, 68% of AI shopping assistant users said they were less likely to return a product bought with AI help. That is a stated intention in a survey, not measured return behaviour, and it concerns AI tools in general. Retailers should test the effect on their own return rates, ideally separating size-related reasons from others.
When does an AI stylist make commercial sense?
The strongest cases tend to combine a wide assortment, where shoppers need help navigating, with complex decisions such as fit, occasion or outfit building, and good product data to ground answers. McKinsey's agentic commerce report also notes that consumers may not trust an agent simply because a brand deploys it, so quality matters more than presence. For narrow assortments of basics, better filters and size guidance may deliver more value at lower cost.

What should executives ask before scaling?
Ask for incremental net revenue after returns versus a holdout, the accuracy rate from sampled reviews, the full cost per conversation, and evidence that results held after the novelty period. If those four answers are positive, scaling is justified; if only usage numbers are available, the case is not yet made.
It is also worth being clear about what the assistant is meant to achieve. A stylist designed to deflect service contacts should be judged mainly on resolution rate and customer satisfaction, while one designed to build outfits should be judged on order value and incremental revenue. Mixing goals makes results hard to interpret and often leads to vague success claims. Agree the primary goal, one or two secondary metrics and the cost ceiling with finance before launch, and report against those consistently over at least two seasons, because fashion demand patterns and assortment mix differ strongly between spring and autumn collections.
Frequently asked questions
Do AI shopping assistants increase conversion?
Adobe reported that in the 2025 US holiday season, traffic referred by generative AI tools converted 31% better than other traffic. That concerns external assistants; the effect of a retailer's own stylist needs to be measured with a controlled test.
How do you measure ROI of an AI stylist?
Compare a randomly assigned test group with a holdout group on incremental net revenue after returns, order value and contact rate. Subtract model, integration and staff costs per conversation, and check answer accuracy through sampled reviews.
Do AI stylists reduce fashion returns?
Evidence is limited. In an Adobe survey, 68% of AI shopping assistant users said they were less likely to return items bought with AI help, but that is self-reported. Retailers should measure actual return rates by test group.
How many people use Zalando's AI assistant?
Zalando reported in March 2026 that its assistant had six million users, four times the number a year earlier. The company did not publish conversion or revenue effects for the assistant.
One edition every weekday morning. Read in five minutes. Free for industry professionals.



