AI quality checks for PIM data: catching missing, wrong and inconsistent attributes
A practical checklist for using rules and AI to find gaps, errors and contradictions in fashion product data before they reach webshops, marketplaces and customers.
KEY TAKEAWAYS Summary by the editors
- Product data quality problems fall into three groups: missing values, wrong values and inconsistent values across attributes, images, channels or systems, and each needs a different kind of check.
- Deterministic rules should catch what can be defined exactly (mandatory fields, GTIN per size and colour, composition adding to 100 percent); AI is best used for what rules cannot see, such as text that contradicts attributes or images that contradict the product record.
- In an Akeneo consumer survey reported by Just Style in April 2026, 43 percent of consumers said they had returned a product in the past year because of inaccurate pre-purchase information.
- Zalando reported that AI-supported size recommendations reduced size related returns by 8 percent in 2025, which illustrates that fit information is among the most valuable attributes to get right.
- AI checks need thresholds, human review queues and feedback loops; a model that flags too many false positives is ignored, and one that auto-corrects without review can spread errors to every channel.
AI quality checks for PIM data combine fixed validation rules with machine learning models that look for problems rules cannot express: product descriptions that contradict attributes, images that show a different colour or neckline than the record, and implausible measurements. The aim is to catch missing, wrong and inconsistent data before it reaches customers, because inaccurate product information is a recognised driver of returns.
Why does product data quality matter for returns?
Returns are one of the largest cost lines in online fashion, and product information is one of the few levers a brand fully controls. In a consumer survey by PIM vendor Akeneo, reported by Just Style in April 2026, 43 percent of consumers said they had returned a product in the past year because of inaccurate pre-purchase information, and the survey attributed 58 percent of apparel returns to sizing concerns. A separate survey of 1,882 online shoppers, commissioned by Overnight Glasses in June 2026 and reported by Retail Insider, found clothing to be the largest return category, with incorrect sizing or poor fit as the main reason.
Survey figures vary with method and sponsor, but the direction is consistent. Zalando also reported that its AI-supported size recommendations reduced size related returns by 8 percent in 2025, according to FashionUnited, which shows how much value lies in accurate fit and size data.
What kinds of data errors should checks find?
| Error type | Example | Best detected by | Typical impact |
|---|---|---|---|
| Missing | No size chart, no material for lining, no care text | Completeness rules per category and channel | Listing rejected, customer uncertainty |
| Wrong format | Composition as free text, GTIN with wrong length | Format and checksum validation | Channel rejection, EDI errors |
| Implausible | Inseam of 120 cm on a regular size, 140 percent cotton | Range rules plus statistical outlier detection | Wrong size advice, returns |
| Inconsistent across fields | Description says "wool blend", composition says 100 percent polyester | Language models comparing text and attributes | Customer complaints, legal exposure |
| Inconsistent with images | Record says navy, image shows black; long sleeve shown as short | Image recognition comparing image and attributes | Returns as "not as described" |
| Inconsistent across channels | Different composition on webshop and marketplace | Comparison of exported data per channel | Loss of trust, partner queries |
Which checks should be rules and which should be AI?
A common mistake is to use AI where a rule would be cheaper and more reliable. Anything that can be defined precisely belongs in rules: mandatory fields, formats, value lists and identifiers. GS1 states, for example, that each size, each colour and each combination needs its own GTIN, which a simple rule can verify. Marketplaces also enforce rules: Zalando requires category specific attributes and a mandatory image set and reports validation results in a product status report.
AI adds value where judgement is needed: reading free text, interpreting images, spotting outliers relative to similar products and suggesting likely correct values. The two layers work best together, with rules running first and AI reviewing what passes. In practice this means a record can be technically complete, with every mandatory field filled, and still be wrong: a correctly formatted composition that belongs to last season's fabric, or a colour value that does not match the photo. Those are the cases where an AI layer earns its cost.
What belongs on a PIM data quality checklist?
Completeness and format (rules)
- Every sellable size and colour has a valid, unique GTIN.
- All mandatory attributes per category and per channel are filled.
- Composition is structured per component and each component adds up to 100 percent.
- Care information, country of origin and size chart are present.
- Values come from controlled lists, not free text, where a list exists.
Plausibility and consistency (AI assisted)
- Measurements are within expected ranges for category, fit and size, compared with similar styles.
- Product titles and descriptions do not contradict structured attributes (material, fit, length, closure).
- Images match recorded colour, pattern, sleeve length and neckline.
- Generated descriptions contain no features that are absent from the attribute record.
- Translations preserve regulated terms such as fibre names and care instructions.
Process and governance
- Each flagged issue goes to a named owner with a due date.
- AI suggestions are marked as suggestions until a person approves them.
- Corrections are made in the master system (for example PLM for composition), not patched in a channel.
- False positive rates are monitored and thresholds adjusted.
- Returns reasons are fed back to identify attributes that drive "not as described" returns.
How do you roll out AI data quality checks?
- Measure the baseline. Score current completeness and error rates per category and channel.
- Fix the rules layer first. It is cheaper and removes most noise before AI is applied.
- Pilot AI checks on one category with a known error set, and measure precision (how many flags are real) and recall (how many real errors are found).
- Set confidence thresholds that keep review volumes manageable for the team.
- Link to returns data to prioritise the attributes that cost most when wrong.
What are the limits of AI checks?
AI checks are probabilistic. Image models can misread colours under different lighting, and language models can miss contradictions phrased in unusual ways. They also inherit gaps in the master data: a model cannot know a lining composition that was never recorded anywhere. Costs include review time and the effort of maintaining reference data. Used as a filter that directs human attention, rather than as an automatic fixer, they make a well run PIM more reliable without hiding responsibility for the data.
Frequently asked questions
How can AI improve product data quality?
AI can compare text, attributes and images to find contradictions, detect implausible measurements relative to similar products and suggest missing values. It works best on top of fixed validation rules and with a human review step for suggested corrections.
Does better product information reduce returns?
Surveys suggest it does. In an Akeneo consumer survey reported by Just Style, 43 percent of consumers said they had returned a product because of inaccurate pre-purchase information. Fit and size information is particularly important for apparel.
What is the most common product data error in fashion?
Missing or inconsistent attributes are common, especially composition held as free text, incomplete size charts and colours that differ between record and images. Which errors dominate varies by brand, so a baseline measurement is the first step.
Should AI correct product data automatically?
Generally not for customer facing or regulated attributes. AI should flag issues and propose values, and a responsible person should approve changes in the master system so that corrections reach all channels.
One edition every weekday morning. Read in five minutes. Free for industry professionals.
SOURCES
- Just Style: Returns surge driven by poor product information
- Retail Insider: Online clothing leads e-commerce returns, with sizing driving most send-backs: Overnight Glasses
- FashionUnited: 'Faster than ever before': Why Zalando is betting on AI
- Zalando Developers: Products Onboarding Overview
- GS1: How many GS1 GTINs do I need when I have a product with many sizes and colours?