GuidesMaterial master data

How to find duplicate material codes in your SAP material master

Why duplicate materials pile up in SAP and ERP systems, why text search can't find them, and a step-by-step method — classify, extract specifications, match on a canonical key — that does.

Enterprise intelligence · 6 min read · 7 October 2026

Every plant that has run SAP for more than a few years has the same part under several codes. A taper roller bearing lives once as "BRG TPR RLLR DBL ROW 510X950", again as "BEARING, TAPER ROLLER, ID 510MM", and a third time under the manufacturer's part number with no description at all. Each one has its own stock, its own reorder point and its own purchase history.

The cost is rarely visible in one place. It shows up as stock bought twice, as a part that was "out of stock" while forty of them sat under another code, and as a material master nobody trusts enough to search before raising a new one. This guide explains why duplicates happen, why the obvious fix doesn't work, and a method that does.

01Why duplicates happen

Material descriptions in SAP are free text. They were typed by different people, at different plants, over different decades, under different abbreviation habits. A mass upload during a merger adds a whole second vocabulary. A new code is faster to raise than an old one is to find, so people raise new ones.

  • Free-text descriptions with no enforced structure
  • Abbreviations that differ by plant, person and era (BRG, BEARNG, BEARING)
  • Specifications buried in the text in mixed units (mm and inches, kW and HP)
  • Manufacturer names written many ways (SKF, SKF AG, S.K.F.)
  • No search-first step before a new code is created
  • Mergers, acquisitions and system migrations that bring their own material lists

02Why text search can't find them

The intuitive fix is a fuzzy text search: find descriptions that look alike. It catches the easy cases and misses the ones that matter. Two descriptions of the same bearing can share no words at all. Two descriptions that are nearly identical can be different parts — the same text with one digit changed in the bore diameter.

Text similarity measures how a description was written, not what the part is. To find duplicates you have to compare what the parts are.

03A method that works

The approach that holds up is to turn each description into a structured record first, then compare the records. In practice that is a pipeline:

  1. Normalise the text. Expand abbreviations, standardise units (convert everything to millimetres, kilowatts, bar), and separate the manufacturer from the description.
  2. Classify each record into a material class — bearing, motor, valve, pump, gearbox, seal and so on. Classification decides which specifications matter: a bearing is defined by bore, outside diameter, width and type; a motor by power, speed, voltage and frame.
  3. Extract the critical specifications for that class into named fields, each with a source and a confidence. A value that cannot be read with confidence is left blank and flagged, not guessed.
  4. Build a canonical key from the critical specifications only. Two records with the same key are exact duplicates, however differently they were written.
  5. Find near-duplicates by comparing specifications within engineering tolerances — the same bearing can be recorded as 510 mm or 509.9 mm — and by checking that the critical fields agree.
  6. Review with evidence. Show a person the two records side by side, the specifications that matched, and the rule that matched them. People approve merges; the system proposes them.
  7. Stop new duplicates at the door. Before a new code can be created, the request is checked against the cleaned master by specification, not by wording, and a match has to be justified in writing.

04Three kinds of duplicate, and why the difference matters

  • Exact duplicates: the same part under different codes. These can be merged into one master record with the other codes mapped to it.
  • Functional duplicates: different parts that do the same job and could replace each other. These are not merged; they are linked, so a buyer sees the alternative before ordering.
  • Interchangeable with conditions: a part that could serve if a stated condition holds — a higher insulation class, a different seal material. These are shown with the condition, never silently substituted.

Treating all three the same way is how clean-ups cause damage: merging two parts that only looked alike, or hiding a real alternative because it was written differently.

05The traps

  • Specification overlap is not identity. An actuator's drive motor can share every listed spec with a general-purpose motor and still be a different piece of equipment. Equipment kind has to be checked before specifications are compared.
  • Parts "for" a machine are not the machine. "BRUSH HOLDER F/ MOTOR" is a holder, not a motor. Descriptions in SAP are usually head-noun first; the words after "for" say what it belongs to.
  • Units lie quietly. 120 HP and 90 kW are the same motor; 300 PSI and 20 bar are the same valve rating. Compare in one unit system or you will miss half the matches.
  • Bigger is not a safe substitute. A larger motor is not an upgrade when the overload relay, contactor and cable were sized to the original. A rating that other equipment was sized against is never a free improvement.
  • Confidence has to be measured. A classifier that is "90% sure" should be right nine times in ten on a hand-checked sample. If nobody has checked, the number means nothing.

06What you have at the end

A material master where each physical part has one master record, every other code maps to it, and every value carries the evidence of where it came from. A creation process that checks before it creates. And a measured figure — precision and recall on a labelled sample — that tells you how far to trust the result, rather than a demo that looked good.

In short

  • Duplicates come from free text, mixed units and the absence of a check before creation.
  • Text similarity finds how things were written; duplicate detection has to compare what the parts are.
  • Classify, extract specifications, match on a canonical key, review with evidence, then gate new codes.
  • Exact, functional and conditional duplicates are different things and need different handling.
  • Measure precision and recall on a labelled sample before trusting any automated clean-up.

Questions

Can SAP find duplicate materials on its own?

SAP can check for similar descriptions and, with Master Data Governance, enforce some rules at creation. What it does not do natively is read free text into specifications and match on them, which is where most duplicates hide.

Does our material data have to leave our network to do this?

No. The whole pipeline — normalisation, classification, extraction, matching — can run on your own infrastructure with local models, which is how we build it.

How many duplicates should we expect?

It varies enormously by industry, age of the system and how creation has been governed. Anyone who quotes you a percentage before looking at your data is guessing. A sample of a few thousand records is enough to measure yours.

What should we do first?

Pick one material class with high spend — motors or bearings are typical — clean it end to end, measure the result, and put the creation gate in place for that class before widening.

Sounds like your problem?

Tell us about it. We'll say honestly whether the swarm can help, and what it would take.

More guides