Detecting Silent Data Failures in Machine Learning Pipelines for Retail Replenishment
DOI:
https://doi.org/10.64137/31079458/IJCSEI-V2I3P105Keywords:
Data Quality, Retail Forecasting, Anomaly Detection, Replenishment, Reconciliation, Synthetic FailuresAbstract
Schema-valid errors in sales feeds can change replenishment decisions without interrupting a machine learning pipeline. This study evaluates six batch monitors and two automatic repair rules using historical M5 retail sales. Eight synthetic failure types are injected into two 112-day evaluation windows covering 30,490 item-store series. Detector fitting, threshold calibration and evaluation are chronologically separated. Twenty paired injection replicates measure detection performance; ten assess downstream loss from a fixed weekly quantile forecaster. With an independent auxiliary total containing approximately 1% multiplicative noise, reconciliation achieves 82.1% corrupted-batch recall at a 1.12% false-positive rate. A maximum-score combination of thirteen statistics achieves 63.0% recall at a 0.79% false-positive rate. Reconciliation recall falls to 1.1% when the auxiliary source shares the injected quantity errors. Total-preserving identifier permutations remain difficult to detect, and more alarms do not necessarily improve decisions. Unmonitored corruption increases mean ordering loss by 3.21%. Reconciliation with rescale-only repair reduces loss by 0.0932 normalized units per item-week relative to the corrupted feed, but does not restore clean-history performance. These results support evaluating source independence, threshold calibration and the repair action jointly. They demonstrate conditional performance on designed faults and historical recorded sales, rather than incident prevalence or realized inventory savings.
References
[1] S. Shankar, R. R. Garcia, J. M. Hellerstein, and A. Parameswaran, “Operationalizing Machine Learning: An Interview Study,” arXiv (Cornell University), Sept. 2022, doi: 10.48550/arxiv.2209.09125
[2] E. Breck, N. Polyzotis, S. Roy, S. E. Whang and M. Zinkevich, "Data validation for machine learning," in Proceedings of Machine Learning and Systems (MLSys), vol. 1, pp. 334–347, 2019. [Online]. Available: https://research.google/pubs/data-validation-for-machine-learning/
[3] N. Sambasivan, S. Kapania, H. Highfill, D. Akrong, P. Paritosh, and L. Aroyo, “‘Everyone wants to do the model work, not the data work’: Data Cascades in High-Stakes AI,” CHI ’21: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 2021, doi: 10.1145/3411764.3445518.
[4] S. Schelter, D. Lange, P. Schmidt, M. Celikel, F. Biessmann, and A. Grafberger, “Automating large-scale data quality verification,” Proceedings of the VLDB Endowment, vol. 11, no. 12, pp. 1781–1794, Aug. 2018, doi: 10.14778/3229863.3229867.
[5] S. Redyuk, Z. Kaoudi, V. Markl and S. Schelter, "Automating data quality validation for dynamic data ingestion," in Proceedings of the 24th International Conference on Extending Database Technology (EDBT), pp. 61–72, 2021. doi: 10.5441/002/edbt.2021.07.
[6] D. C. Montgomery, Introduction to Statistical Quality Control, 8th ed. Hoboken, NJ, USA: Wiley, 2019.
[7] S. Rabanser, S. Günnemann and Z. C. Lipton, "Failing loudly: An empirical study of methods for detecting dataset shift," Proceedings of the 33rd International Conference on Neural Information Processing Systems, pp. 1396 - 1408, Dec 2019.
[8] D. Sculley et al., "Hidden technical debt in machine learning systems," NIPS'15: Proceedings of the 29th International Conference on Neural Information Processing Systems, vol. 2, pp. 2503–2511, 2015.
[9] E. Breck, S. Cai, E. Nielsen, M. Salib, and D. Sculley, “The ML test score: A rubric for ML production readiness and technical debt reduction,” 2017 IEEE International Conference on Big Data (Big Data), Dec. 2017, doi: 10.1109/bigdata.2017.8258038.
[10] J. Gama, I. Žliobaitė, A. Bifet, M. Pechenizkiy, and A. Bouchachia, “A survey on concept drift adaptation,” ACM Computing Surveys, vol. 46, no. 4, pp. 1–37, Mar. 2014, doi: 10.1145/2523813.
[11] J. Lu, A. Liu, F. Dong, F. Gu, J. Gama, and G. Zhang, “Learning under Concept Drift: A Review,” IEEE Transactions on Knowledge and Data Engineering, vol. 31, no. 12, pp. 1–1, 2018, doi: 10.1109/tkde.2018.2876857.
[12] Bifet and R. Gavaldà, “Learning from Time-Changing Data with Adaptive Windowing,” Proceedings of the 2007 SIAM International Conference on Data Mining, Apr. 2007, doi: 10.1137/1.9781611972771.42.
[13] V. Chandola, A. Banerjee, and V. Kumar, “Anomaly Detection: A Survey,” ACM Computing Surveys, vol. 41, no. 3, pp. 1–58, July 2009, doi: 10.1145/1541880.1541882.
[14] F. T. Liu, K. M. Ting, and Z.-H. Zhou, “Isolation Forest,” 2008 Eighth IEEE International Conference on Data Mining, pp. 413–422, Dec. 2008, doi: 10.1109/icdm.2008.17
[15] F. J. Massey, “The Kolmogorov-Smirnov Test for Goodness of Fit,” Journal of the American Statistical Association, vol. 46, no. 253, pp. 68–78, Mar. 1951, doi: 10.1080/01621459.1951.10500769.
[16] N. DeHoratius and A. Raman, “Inventory Record Inaccuracy: An Empirical Analysis,” Management Science, vol. 54, no. 4, pp. 627–641, Apr. 2008, doi: 10.1287/mnsc.1070.0789.
[17] F. Petropoulos, “Forecasting: Theory and practice,” International Journal of Forecasting, vol. 38, no. 3, Jan. 2022, doi: 10.1016/j.ijforecast.2021.11.001.
[18] N. Polyzotis, S. Roy, S. E. Whang, and M. Zinkevich, “Data Lifecycle Challenges in Production Machine Learning,” ACM SIGMOD Record, vol. 47, no. 2, pp. 17–28, Dec. 2018, doi: 10.1145/3299887.3299891.
[19] T. Schröder and M. Schulz, “Monitoring machine learning models: A categorization of challenges and methods,” Data Science and Management, vol. 5, no. 3, Aug. 2022, doi: 10.1016/j.dsm.2022.07.004.
[20] S. Shankar and A. G. Parameswaran, “Towards Observability for Production Machine Learning Pipelines,” Proceedings of the VLDB Endowment, vol. 15, no. 13, pp. 4015–4022, Sept. 2022, doi: 10.14778/3565838.3565853.
[21] S. Makridakis, E. Spiliotis, and V. Assimakopoulos, “M5 accuracy competition: Results, findings, and conclusions,” International Journal of Forecasting, vol. 38, no. 4, Jan. 2022, doi: 10.1016/j.ijforecast.2021.11.013.
[22] P. J. Rousseeuw and C. Croux, “Alternatives to the Median Absolute Deviation,” Journal of the American Statistical Association, vol. 88, no. 424, pp. 1273–1283, Dec. 1993, doi: 10.1080/01621459.1993.10476408.
[23] G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye and T.-Y. Liu, "LightGBM: A highly efficient gradient boosting decision tree," 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, vol. 30, pp. 1-9, 2017.
[24] E. A. Silver, D. F. Pyke, and D. J. Thomas, Inventory and production management in supply chains. Boca Raton: Crc Press, 2017.


