What is RAG? How AI answers questions from your product and order data
Retrieval-augmented generation lets an AI assistant look up your catalogue, price lists and orders before it answers. How it works, what data it needs, and where it goes wrong.
KEY TAKEAWAYS Summary by the editors
- Retrieval-augmented generation (RAG) is a technique in which an AI system first retrieves relevant information from a company's own sources and then passes it to a large language model as context for its answer.
- The term was introduced in a 2020 research paper by Patrick Lewis and colleagues, which combined a pretrained language model with a searchable index of documents.
- RAG lets a fashion business ground answers in current product, price and order data without retraining a model, and allows answers to cite the records they used.
- RAG answers are only as good as the underlying data: incomplete attributes, outdated price lists or duplicate article records lead directly to wrong answers.
- The OWASP Top 10 for LLM Applications 2025 lists prompt injection, sensitive information disclosure and vector and embedding weaknesses among the main security risks for such systems.
Retrieval-augmented generation (RAG) is a way of making an AI assistant answer from your own data instead of from its general training. When someone asks a question, the system first searches your product catalogue, price lists, order history or policy documents, then hands the relevant extracts to a large language model, which writes an answer based on them. For fashion wholesale and retail teams, it is the most common architecture behind assistants that answer questions like "Which styles from the summer line are still available in size 38?"
What is retrieval-augmented generation?
The term comes from a 2020 research paper by Patrick Lewis and colleagues, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks". The authors observed that large pretrained language models store knowledge in their parameters but struggle to access it precisely, to show where an answer came from, or to update what they know. Their solution combined the model's internal "parametric" memory with an external, searchable index of documents. The paper reported that this produced more specific and factual answers than a language model working alone.
AWS summarises the idea in business terms: RAG makes a language model consult an authoritative knowledge base outside its training data before it responds, which extends the model to an organisation's internal knowledge without retraining it.
How does RAG work, step by step?
- Prepare the data: product records, line sheets, terms and conditions, order and delivery data are collected from their source systems. Text is split into chunks and converted into numerical vectors (embeddings), usually stored in a vector database; structured data can also be queried directly.
- Retrieve: the user's question is converted in the same way and matched against the store to find the most relevant records, for example the article master data for a style and its current stock.
- Augment: the retrieved records are inserted into the prompt as context, together with instructions such as "answer only from these records and cite them".
- Generate: the language model writes the answer, ideally with references to the records it used.
- Refresh: as AWS notes, the documents and embeddings must be updated in real time or in regular batches, otherwise the assistant answers from yesterday's data.

What can RAG do with product and order data in fashion?
RAG is useful wherever people spend time looking things up across several systems. Typical examples in fashion wholesale and retail include:
| Use case | Typical question | Data the system must retrieve |
|---|---|---|
| Sales rep support | What are the fabric and delivery window for this style? | Product master data, collection calendar |
| Wholesale customer service | Where is my order and what is still open? | Order, shipment and invoice data per account |
| Re-order assistance | Which of my best sellers can I still re-order this season? | Sell-out or order history, available-to-promise stock |
| Product content | Draft a description for this new article | Attributes, care instructions, brand tone guide |
| Internal policies | What are our return conditions for key accounts? | Contracts, terms, policy documents |
| Buying and merchandising | Summarise last season's performance in knitwear | Sales reports, assortment plans |
The common thread is that the answer must be correct for one specific company, customer and moment. A general model cannot know that; a retrieval step can supply it.
Not every question should be answered by searching text. Questions about stock levels, open order quantities or prices are best answered by querying the source system directly with exact filters, while questions about care instructions, delivery terms or collection stories suit document search. Many practical assistants combine both: a structured query for the numbers and document retrieval for the explanation, with the language model only phrasing the final answer.
Why is RAG better than training a model on your data?
Fine-tuning a model bakes information into its weights, which is costly to repeat and still does not guarantee correct recall. Product and order data change daily, so retraining would never keep up. AWS describes RAG as a cost-effective way to improve model output compared with retraining foundation models, and highlights two further advantages: answers can draw on frequently updated sources, and they can include citations that users can check. For a sales team, that last point is decisive: a rep needs to see which price list or order record an answer came from before quoting it to a customer.
What goes wrong with RAG in practice?
- Poor master data: if attributes are missing or inconsistent across systems, retrieval returns incomplete or contradictory records, and the model fills gaps with guesses.
- Wrong retrieval: similar style names, colour codes or seasons can cause the system to fetch the wrong article. Structured filters (season, market, customer account) reduce this.
- Stale indexes: stock and order status change constantly; an index refreshed weekly will give confident but outdated answers.
- Access control: a wholesale customer must never see another account's prices or orders. Permissions have to be applied at retrieval, not only in the user interface.
- Security: the OWASP Top 10 for LLM Applications 2025 lists prompt injection, sensitive information disclosure, and vector and embedding weaknesses among the main risks for these systems.

How should a fashion company prepare for RAG?
The groundwork is mostly data work. Companies need one reliable source for each kind of record (articles, prices, stock, orders, customers), consistent identifiers that link them, and clear ownership for keeping them current. Documents such as terms, line sheets and care guides should be in machine-readable form and versioned. Start with a narrow, high-volume question type, such as order status for wholesale accounts or product facts for sales reps, require the assistant to cite its sources, and keep a human escalation path. Only when answer quality is measured and stable is it sensible to widen the scope or let the system trigger actions such as drafting a re-order.
Frequently asked questions
What does RAG stand for in AI?
RAG stands for retrieval-augmented generation. It describes AI systems that retrieve relevant information from external sources, such as a company's product or order data, and give it to a large language model as context before the model generates an answer.
Is RAG the same as fine-tuning?
No. Fine-tuning changes a model's internal weights through additional training, while RAG leaves the model unchanged and supplies relevant information at the moment of the question. RAG is better suited to data that changes often, such as stock, prices and orders, and makes it easier to show sources.
Can RAG still hallucinate?
Yes. If retrieval returns the wrong or incomplete records, or the model ignores them, answers can still be wrong. Requiring citations, limiting answers to retrieved data and testing against known questions reduce the risk but do not eliminate it.
What data do I need for a RAG assistant in fashion wholesale?
At minimum, clean product master data, current price lists, stock or availability data and order records per customer account, linked by consistent identifiers. Documents such as terms and line sheets should be digital and versioned, and access rights must be defined per user and account.
One edition every weekday morning. Read in five minutes. Free for industry professionals.


