Tabular In-Context Learning for Low-Resource Rare Industrial Fault Detection: a Cost-Sensitive Comparison on Scania APS and SECOM
DOI:
https://doi.org/10.70393/6a6374616d.343334ARK:
https://n2t.net/ark:/40704/JCTAM.v3n4a01Disciplines:
Software SystemsSubjects:
Software EngineeringReferences:
18Keywords:
Tabular Foundation Model, In-context Learning, Industrial Fault Detection, Class Imbalance, Cost-sensitive LearningAbstract
Rare industrial fault detection is complicated by class imbalance, missing sensor values, asymmetric errors, and limited training resources. We compared TabICLv2 with TabM, XGBoost, and logistic regression on Scania APS Failure and SECOM. For Scania, nested stratified samples of 5,000 and 10,000 training records were drawn under five prespecified seeds. Shared three-fold splits supported out-of-fold (OOF) evaluation; tunable models were selected by OOF AUPRC, and decision thresholds were locked by minimizing OOF cost, defined as C = 10 × FP + 500 × FN. The official test set remained locked until all selections were complete. SECOM used three repeats of five-fold outer cross-validation with three-fold inner selection. Paired class-stratified bootstrap analyses used 2,000 replicates. At the prespecified 10,000-record Scania budget, TabICLv2 achieved an AUPRC of 0.8879, an official cost of 13,002, an MCC of 0.6514, and a Brier score of 0.00685; XGBoost achieved 0.8549, 16,456, 0.6125, and 0.00840, respectively. The TabICLv2-minus-XGBoost difference was 0.03294 for AUPRC (95% CI 0.02089–0.04570) and −3,454 for cost (95% CI −5,370.60 to −1,711.95). Across 15 SECOM outer folds, mean AUPRC was 0.2083 for TabICLv2 and 0.1705 for XGBoost. After averaging repeated OOF probabilities per record, the paired AUPRC difference was 0.05110 (95% CI −0.00693 to 0.11358). TabICLv2 required 992.2 s, approximately 413 times the XGBoost time. TabICLv2 improved the prespecified Scania outcomes but incurred substantial task-side computational cost; SECOM provided suggestive rather than definitive support.
References
[1] He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on knowledge and data engineering, 21(9), 1263-1284.
[2] Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PloS one, 10(3), e0118432.
[3] Chicco, D., & Jurman, G. (2020). The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC genomics, 21(1), 6.
[4] Glenn, W. B. (1950). Verification of forecasts expressed in terms of probability. Monthly weather review, 78(1), 1-3.
[5] Fouladvand, S., Noshad, M., Goldstein, M. K., Periyakoil, V. J., & Chen, J. H. (2023). Mild cognitive impairment: data-driven prediction, risk factors, and workup. AMIA Summits on Translational Science Proceedings, 2023, 167.
[6] Grinsztajn, L., Oyallon, E., & Varoquaux, G. (2022). Why do tree-based models still outperform deep learning on typical tabular data?. Advances in neural information processing systems, 35, 507-520.
[7] Ye, H. J., Liu, S. Y., Cai, H. R., Zhou, Q. L., & Zhan, D. C. (2024). A closer look at deep learning methods on tabular datasets. arXiv preprint arXiv:2407.00956.
[8] Hollmann, N., Müller, S., Purucker, L., Krishnakumar, A., Körfer, M., Hoo, S. B., ... & Hutter, F. (2025). Accurate predictions on small data with a tabular foundation model. Nature, 637(8045), 319-326.
[9] Qu, J., HolzmÞller, D., Varoquaux, G., & Morvan, M. L. (2025). Tabicl: A tabular foundation model for in-context learning on large data. arXiv preprint arXiv:2502.05564.
[10] Qu, J., HolzmÞller, D., Varoquaux, G., & Morvan, M. L. (2026). TabICLv2: A better, faster, scalable, and open tabular foundation model. arXiv preprint arXiv:2602.11139.
[11] Gorishniy, Y., Kotelnikov, A., & Babenko, A. (2025, May). Tabm: Advancing tabular deep learning with parameter-efficient ensembling. In International Conference on Learning Representations (Vol. 2025, pp. 77899-77935).
[12] Niculescu-Mizil, A., & Caruana, R. (2005, August). Predicting good probabilities with supervised learning. In Proceedings of the 22nd international conference on Machine learning (pp. 625-632).
[13] Varma, S., & Simon, R. (2006). Bias in error estimation when using cross-validation for model selection. BMC bioinformatics, 7(1), 91.
[14] Cawley, G. C., & Talbot, N. L. (2010). On over-fitting in model selection and subsequent selection bias in performance evaluation. The Journal of Machine Learning Research, 11, 2079-2107.
[15] Scania CV AB. (2017). APS failure at Scania Trucks [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C51S51
[16] McCann, M., & Johnston, A. (2008). SECOM [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C54305
[17] Effron, B., & Tibshirani, R. J. (1993). An introduction to the bootstrap. Monographs on statistics and applied probability, 57, 436.
[18] Loza, M., Chushig-Muzo, D., Milara, E., Bote-Curiel, L., Estrada-Petrocelli, L., & Grijalva, F. (2026). Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models. arXiv preprint arXiv:2607.26000.
Downloads
Published
How to Cite
Issue
Section
ARK
License
Copyright (c) 2026 The author retains copyright and grants the journal the right of first publication.

This work is licensed under a Creative Commons Attribution 4.0 International License.









