How Vinted uses AI in search and recommendations for second-hand fashion
Vinted's engineering team has published detailed accounts of its search, autocomplete and recommendation systems. This case study sets out what is documented, what the results were and what other marketplaces can learn.

KEY TAKEAWAYS Summary by the editors
- Vinted's engineers report that its search engine held one billion searchable items as of 11 November 2024, with mean data-layer query latency under 20 milliseconds.
- Vinted uses a two-tower model, with a frozen multilingual CLIP model on the query side and item metadata plus primary-image embeddings on the item side, to retrieve items that keyword search misses.
- According to Vinted, autocomplete now starts more than 20 percent of search sessions, and a learning-to-rank model with 63 features re-ranks suggestions on each keystroke.
- Vinted rejected one autocomplete feature, scoped suggestions, because it raised click-through but made users about 1.3 percent less likely to buy in the same session.
- Vinted's published engineering material covers search, autocomplete and recommendations in detail but does not describe price suggestion or trust and safety models, so those areas cannot be assessed from public sources.
Vinted, the second-hand fashion marketplace, uses machine learning mainly to help buyers find one-off items in an enormous, constantly changing catalogue. Its engineers have published how dense (embedding-based) retrieval, personalised autocomplete and a two-tower recommender work in production. The published evidence is strong on search and recommendation and silent on price guidance and trust, which this article therefore does not cover.
Why is search a hard AI problem for a second-hand marketplace?
A conventional retailer sells repeated stock keeping units with clean attributes. A peer-to-peer marketplace lists items that exist once, are described by sellers in many languages, and are photographed in very different ways. Keyword search often returns nothing for such catalogues, which Vinted's engineers described as a missed business opportunity in their account of dense retrieval.
Scale adds to the difficulty. In January 2025 Vinted reported that its index passed one billion searchable documents on 11 November 2024, up from about 100 million in late 2019, with mean data-layer latency under 20 milliseconds. It had earlier moved its search engine from Elasticsearch to Vespa, a platform that combines search, vector retrieval and machine learning ranking.
How does Vinted use dense retrieval to find items?
Dense retrieval began as hackathon experiments in 2022. It was first used to fill sessions with few results, and from spring 2024 it was rolled out to all sessions after about 50 A/B tests. The model maps queries and items into a shared 256-dimension space. The query tower is a frozen multilingual CLIP model with a trained projection head. The item tower combines categorical embeddings, such as brand and category, with the embedding of the primary product image.
The team reported that scaling training data tenfold, to more than 100 million positive pairs, improved results, and that computing image embeddings for all items took about one month. It also reports an error rate below 0.02 percent. The engineers are open about the cost: keeping result counts, filters and sort orders consistent added substantial complexity that they hope to reduce later.
| System | What it does | Published detail |
|---|---|---|
| Dense retrieval | Finds items that keyword matching misses, using text and image embeddings | Two-tower model, 256-dimension embeddings, about 50 A/B tests before full rollout |
| Personalised autocomplete | Suggests queries as the user types | More than 20 percent of search sessions start with it; LightGBM LambdaRank re-ranks with 63 features |
| Homepage recommendations | Retrieves candidate listings for each user | Two-tower model with user interaction sequences; first-stage P99 latency about 50 ms |
| Search platform | Serves queries across the catalogue | One billion documents as of 11 November 2024; mean data-layer latency under 20 ms |

How does personalised autocomplete work?
In April 2026 Vinted described a pipeline that generates about 125 million candidate queries across 24 languages. Candidates come from product metadata and search logs, and are regenerated twice a week. They are scored on measures such as sell-through rate, sales volume and suggestion click-through, with a diversity filter so that newer and niche suggestions stay visible.
A LightGBM ranking model then re-ranks the top 20 exact-prefix matches using 63 features covering the query, popularity, user behaviour and context. The service handles 4,700 queries per second at 31 milliseconds P99. Vinted reports that the full suggestion system raised suggestion click-through by about 49 percent and search transactions by about 0.8 percent in short-term experiments, and that the personalisation layer added about 8 percent more click-through on top.
How does Vinted generate recommendations?
Vinted's homepage recommender has three stages. The first retrieves candidates with approximate nearest neighbour search. An in-house two-tower model encodes listings (brand, price, size and photos) on one side and a user's sequence of clicks, favourites and purchases on the other, and the distance between the two embeddings serves as a relevance score.
An experiment is worth noting. Approximate search had recall of about 60 to 70 percent against exact search, yet users did not notice. When the team tested exact search on half of users, latency rose about 40 percent and satisfaction did not improve enough to justify it, so approximate search stayed.
What did Vinted learn, and what remains unproven?
Three lessons stand out. First, the team measured business outcomes such as transactions and buyer gross merchandise value, not only click-through, and used that discipline to reject scoped suggestions: click-through rose about 2.4 percent, but users were about 1.3 percent less likely to buy in the same session. Second, latency budgets shaped model choices, since native prefix queries at about 220 milliseconds P99 were too slow for autocomplete. Third, much of the work is infrastructure, not modelling.
- Results are reported by Vinted, from its own experiments, and are mostly short-term.
- The published material does not cover price suggestion, fraud or trust and safety models.
- Large language model suggestion generation is described as a next step, constrained by a 100 millisecond latency budget.
How should a marketplace measure whether search AI works?
Vinted's posts offer a useful template. The team ran about 50 A/B tests before dense retrieval was fully enabled, reported results with significance levels, and examined side effects such as latency, result consistency and in-session purchases. A feature that raised click-through but lowered buying was not shipped. Teams that judge search changes only on clicks risk optimising for curiosity instead of commerce.
- Define the business outcome first, for example transactions per active user or buyer gross merchandise value.
- Run controlled experiments on a share of users before rollout.
- Track latency at the tail (P99), since slow suggestions are ignored.
- Check consistency effects, such as result counts changing when a filter is applied.
- Review short-term gains against longer-term behaviour before declaring success.
The same posts also show the limits of self-reported results. Most figures are from short-term experiments run by Vinted, they are not independently audited, and they describe one platform with its own catalogue and users. A fashion retailer with a standard catalogue would see different effects.

What can other fashion marketplaces take from this?
Marketplaces with unique inventory benefit from embeddings because exact keyword matching fails when listings are sparse and multilingual. The prerequisites are clean categorical data, good photographs, search logs and an A/B testing culture that can reject features that look good on click-through. Smaller platforms should note that Vinted's gains came from years of incremental engineering on a platform built for this purpose, not from a single model.
Vinted has also worked with an outside advertising technology provider to show personalised post-purchase offers, according to trade reports from August 2024, which shows that AI on a marketplace extends beyond search into monetisation. Public detail on that is limited.
Frequently asked questions
Does Vinted use AI?
Yes. Vinted's engineering blog describes machine learning in dense retrieval for search, personalised autocomplete ranking and a two-tower recommendation model. Its published posts do not describe price suggestion or trust and safety models.
What search engine does Vinted use?
Vinted migrated search from Elasticsearch to Vespa. In January 2025 it reported one billion searchable documents in the index, with mean data-layer query latency under 20 milliseconds.
How does Vinted personalise autocomplete?
A LightGBM LambdaRank model re-ranks the top exact-prefix suggestions using 63 features covering the query, popularity, user behaviour and context. Candidates themselves are generated twice a week from product metadata and search logs.
What is dense retrieval in e-commerce search?
Dense retrieval converts queries and items into numerical embeddings and finds the nearest matches in that space. It can return relevant items even when no keywords match, which helps in multilingual, image-led catalogues.
One edition every weekday morning. Read in five minutes. Free for industry professionals.
SOURCES
- Vinted Engineering: Dense Retrieval
- Vinted Engineering: How Vinted Serves Personalised Search Autocomplete
- Vinted Engineering: Vinted Search Scaling Chapter 9, Billion-Scale Search
- Vinted Engineering: Adopting the Vespa search engine for serving personalized second-hand fashion recommendations
- Decision Marketing: Vinted signs deal with AI retail marketing specialist




