A specialty retailer that has not started with AI, and whose system reorders store stock automatically, should spend its first AI budget on the accuracy of store inventory records, because that system triggers replenishment from the stock level it records.
ECR Retail Loss, a retailer and manufacturer working group, has measured how often those records are wrong. In Measuring the Sales Impact of Improving Inventory Records, published July 6, 2020, three university professors worked with seven unnamed retailers, all among Europe's largest, in four countries, some in grocery and general merchandise and some in fashion and apparel. At stocktake the quantity on hand differed from the system's figure by one unit or more for 63.39 percent of audited SKUs at the grocery and general merchandise retailers and 54.08 percent at the fashion and apparel ones. Test stores then counted and corrected their records while matched control stores did not, and over the next twelve weeks the test stores sold more at every retailer in that experiment, by between 3.83 and 8.38 percent. The authors write that because fashion items often carry a higher average unit value, fixing the records matters as much to sales in fashion and apparel as in grocery. That fashion and apparel group is the part of the study nearest a specialty chain.
The same three authors published Smart Inventory Record Inaccuracy (IRI) Prediction and Management on June 24, 2026, funded by grants from companies including RGIS, NCR Voyix and Retail Insight, which sells AI software that corrects store inventory records. It says that when a record shows more stock than the store holds, replenishment is not triggered even though the shelf is empty, and it tests a machine learning model built on more than 1.3 million stock audits at six grocery retailers, collected from May 2018 to April 2022 and scored only on later periods the model had not seen. Each morning the model ranks the items each store should count. At one retailer, 87.5 percent of the five items it picked per store per day were inaccurate, and it found about 19 percent more inaccuracies than ranking items by sales velocity, the best method without a model in that comparison. It runs on data retail systems already generate, product and store master data among them, though its main investment is data quality. All six retailers are grocers, the report has no measured sales result for the model, and the authors call a disciplined pilot the next step before scale deployment.
The report lays out a possible 90-day path. In weeks one to six the data feeds are aligned, the model is built and its daily list is connected to the store tasking tool. Department and category leads set tolerances in weeks seven to ten while associates are trained, and in weeks eleven to thirteen matched test and control stores are compared on availability, removal of stock the system shows and the shelf lacks, audit productivity and sales. In the test stores the store manager reviews the list each day, inventory associates do the checks, and replenishment and planning adjust reorder settings where the same errors keep returning. If the test stores show no gain in availability or sales over their controls after the four to eight weeks of comparison the authors advise, the retailer stops before any rollout.
At two of the fashion and apparel retailers in the 2020 study the discrepancies where the store held more stock than its records showed exceeded those where it held less. The authors suggest as a possible cause the product returns that are very frequent in fashion.