How to measure virtual try-on: a test design and the numbers to track
Try-on users are not a random sample, so naive comparisons overstate results. A test design with holdouts, the metrics that matter and the traps to avoid.

KEY TAKEAWAYS Summary by the editors
- Reported results are promising but largely company claims: Zalando's director of applied science told Business of Fashion that returns fell 40% after an April 2023 test, according to eMarketer.
- People who choose to use a try-on feature differ from those who do not, so comparing users with non-users overstates the effect; randomise who is offered the feature instead.
- The primary metric should be net margin per exposed visitor, with conversion, return rate by reason, average order value and repeat purchase as supporting metrics.
- A difference-in-differences design with a control group, as used in a 2021 Journal of Operations Management study of free returns, is a practical option when randomising visitors is not possible.
- Zalando's own account separates pilots from scale: it reports return rates down by up to 40% in Virtual Fitting Room pilots, and a 10% reduction in size related returns for its measurement tool, which shows effects depend on category and design.
To measure virtual try-on properly, randomly offer the feature to part of your traffic, keep a holdout group without it, and compare net margin per visitor rather than usage or conversion alone. Users who opt in are not representative, so a simple comparison of users and non-users will flatter the technology. Define metrics, duration and stopping rules before launch.
What results have retailers reported for virtual try-on?
eMarketer reported in April 2026 that Levi Strauss was developing a generative AI feature called Imagine on Me, that Zara had introduced an AI try-on feature for app users earlier in 2026, and that Zalando planned to roll its virtual try-on technology out to all customers in 2026. It cited a Business of Fashion interview in which Zalando's director of applied science said returns fell 40% after an April 2023 test. A start-up, Catches, expects a 10% increase in conversions and a 20X to 30X return on investment, which is an expectation from a vendor rather than a measured outcome. eMarketer also noted that companies are mostly in early testing and that most shoppers are not interested in using AI try-on tools.
Zalando's June 2026 account adds context. It says return rates fell by up to 40% in Virtual Fitting Room pilots, and that the feature uses body measurements to create a 3D avatar for comparing sizes of the same item. An earlier tool for size recommendations from two photos was reported by Just Style in July 2023 to have reduced size related returns by 10% in women's tops and dresses in Germany, Austria and Switzerland. The gap between 10% and up to 40% shows how much results depend on the tool, the category and the comparison.
How do you design a valid virtual try-on test?
- Choose a narrow scope: one or two categories where fit uncertainty is high and return reasons are well recorded, such as jeans or dresses.
- Randomise exposure at visitor or customer level so that the group offered try-on and the holdout group are alike. Persist assignment so that a returning visitor sees the same experience.
- Record whether each exposed visitor used the feature, so you can estimate both the effect of being offered it (intention to treat) and the effect among users, while keeping the first as the headline number.
- Fix the metric set, the minimum run time and the decision rule in writing before launch. Use a power calculation to set sample size rather than stopping when a number looks good.
- Follow orders until returns have materialised, which in fashion means waiting at least through your returns window plus processing time.
- Report results by category, device and new versus returning customers, and keep the holdout in place until the final analysis.

Which numbers should you track?
| Metric | Why it matters | Watch for |
|---|---|---|
| Net margin per exposed visitor | Combines conversion, basket, returns cost and feature cost | Needs reliable return cost allocation |
| Conversion rate | Shows whether try-on removes hesitation | Can rise while returns also rise |
| Return rate by reason (size, fit, look) | Shows whether the feature acts on the intended problem | Reason codes are often inconsistent |
| Average order value and items per order | Detects bracketing, where several sizes are ordered | A lower basket can be a good outcome if returns fall |
| Feature usage and completion rate | Shows whether shoppers can use it | Usage is not impact |
| Repeat purchase within a defined period | Tests the loyalty claim | Needs longer follow-up |
| Cost per session (compute, licensing) | Needed for ROI | Unit costs change with volume |
What are the common mistakes in try-on measurement?
- Comparing users with non-users. People who try on may already be more decided or more engaged, so the gap overstates the feature.
- Measuring too early. Returns lag orders. A conversion gain with no returns data tells you little.
- Counting the pilot as the rollout. A pilot in one category with enthusiastic customers rarely transfers unchanged.
- Ignoring image quality and avatar realism. If the render misleads on fit or colour, returns can rise for new reasons. Track returns coded as not as pictured.
- Skipping the cost side. ROI claims need compute, integration, support and licensing costs alongside gains.
What if you cannot randomise visitors?
A quasi-experimental design can work. A 2021 study of free returns at a Swedish online fashion retailer, published in the Journal of Operations Management, used a market where the change was introduced (Denmark, from 1 November 2017) and other European markets as a control group, comparing before and after with difference-in-differences. The authors tested robustness with parallel trends, event studies and placebo tests. The same logic applies if you launch try-on in one market first, though you must account for seasonal differences and promotions that affect only one market.
Plan for the data you will need in advance. Order, return and session records must be joined at customer level, with consistent return reason codes and a flag for exposure. If your return reasons are free text, spend time cleaning them before the test, because the whole analysis of whether try-on acts on size and fit depends on them. Agree with finance how return handling costs are allocated, so that the margin figure survives scrutiny when results are presented.

How should you read the result?
Look first at net margin per exposed visitor and its confidence interval. If it is positive and the return rate fell for the intended reasons, scale to similar categories. If conversion rose but returns did not fall, the feature may be a marketing tool rather than a returns tool, which is a different business case. If usage is low, improve placement and speed before judging the technology. Finally, document what you tested and for how long, so that the next category can reuse the design.
Frequently asked questions
Does virtual try-on reduce fashion returns?
Some retailers report that it does. Zalando says return rates fell by up to 40% in Virtual Fitting Room pilots and, per eMarketer, that returns fell 40% after an April 2023 test. These are company figures from pilots, so test the effect on your own assortment before assuming similar results.
What is the best metric for a virtual try-on test?
Net margin per exposed visitor, because it includes conversion, basket size, returns cost and the cost of running the feature. Track return rate by reason and conversion as supporting metrics.
How long should a virtual try-on test run?
Long enough to reach the sample size from your power calculation and to observe returns, which means at least your returns window plus processing time. Avoid stopping early because a result looks favourable.
Why not just compare try-on users with other shoppers?
Because users self-select. They may already be more engaged or more certain, so the comparison overstates the feature's effect. Randomly offering the feature, and analysing everyone offered it, removes that bias.
One edition every weekday morning. Read in five minutes. Free for industry professionals.
SOURCES
- eMarketer: Retailers rely on virtual try-on to curb returns, boost conversions
- Zalando Corporate: How Zalando uses technology to help customers find the right size
- Just Style: Zalando reduces size-related returns with new tool
- Hanken School of Economics: The impact of free returns on online purchase behavior, evidence from an intervention at an online retailer




