TY - JOUR
T1 - Revisiting weakly supervised tabular anomaly detection from a cell-level perspective
AU - Wang, Jiahui
AU - Peng, Zhen
AU - Jia, Xujing
AU - Lin, Qika
AU - Ma, Lan
AU - Shi, Bin
N1 - Publisher Copyright:
© 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2026/11
Y1 - 2026/11
N2 - Thanks to the full utilization of limited annotations along with abundant unlabeled data in reality, weakly supervised tabular anomaly detection has achieved improved performance over the unsupervised paradigm in recent years and has attracted widespread attention. However, almost all existing efforts follow the classic sample-level problem definition, which aims to provide a binary judgment of whether a sample is abnormal or not, without further indicating the detailed abnormalities within the samples. Actually, tabular data form a collection of numerous cells, inherently possessing the capability to reveal anomalies at the granularity of individual cells. Based on this intuition, we shift the studied problem to a more fundamental cell level by reformulating coarse-grained tabular anomaly detection into a fine-grained topic termed cell anomaly detection. Note that readily available labels are limited and remain at the sample level without precisely indicating which table cells are problematic, undoubtedly increasing the challenge. To this end, we propose a novel weakly supervised framework named Trail, which directly quantifies the abnormality of each table cell under the mixture of incomplete and inexact supervision, rather than producing a single anomaly score for the entire sample. Specifically, Trail derives discriminative representations that help reveal unusual cell patterns by adaptively learning a sparse topology between features and a self-supervised task called masked tabular modeling. With a few imprecise sample-level labels available, we draw inspiration from the idea of multiple instance learning (MIL) to assign higher scores to deviant cells in abnormal samples. The spotted cell anomalies not only serve as further indicators for determining suspicious samples, but also intuitively reflect the location and severity of abnormalities within the sample. Experiments on ten benchmark datasets corroborate the effectiveness of Trail w.r.t. AUC-ROC and AUC-PR. Case studies demonstrate that Trail provides an easily understandable way for humans to comprehend tabular anomalies from the granularity of cells22We are willing to release the source code after the review. We hope that our work will advance future research on the fine-grained task of cell anomaly detection, particularly in the development of annotated datasets for cell anomalies.
AB - Thanks to the full utilization of limited annotations along with abundant unlabeled data in reality, weakly supervised tabular anomaly detection has achieved improved performance over the unsupervised paradigm in recent years and has attracted widespread attention. However, almost all existing efforts follow the classic sample-level problem definition, which aims to provide a binary judgment of whether a sample is abnormal or not, without further indicating the detailed abnormalities within the samples. Actually, tabular data form a collection of numerous cells, inherently possessing the capability to reveal anomalies at the granularity of individual cells. Based on this intuition, we shift the studied problem to a more fundamental cell level by reformulating coarse-grained tabular anomaly detection into a fine-grained topic termed cell anomaly detection. Note that readily available labels are limited and remain at the sample level without precisely indicating which table cells are problematic, undoubtedly increasing the challenge. To this end, we propose a novel weakly supervised framework named Trail, which directly quantifies the abnormality of each table cell under the mixture of incomplete and inexact supervision, rather than producing a single anomaly score for the entire sample. Specifically, Trail derives discriminative representations that help reveal unusual cell patterns by adaptively learning a sparse topology between features and a self-supervised task called masked tabular modeling. With a few imprecise sample-level labels available, we draw inspiration from the idea of multiple instance learning (MIL) to assign higher scores to deviant cells in abnormal samples. The spotted cell anomalies not only serve as further indicators for determining suspicious samples, but also intuitively reflect the location and severity of abnormalities within the sample. Experiments on ten benchmark datasets corroborate the effectiveness of Trail w.r.t. AUC-ROC and AUC-PR. Case studies demonstrate that Trail provides an easily understandable way for humans to comprehend tabular anomalies from the granularity of cells22We are willing to release the source code after the review. We hope that our work will advance future research on the fine-grained task of cell anomaly detection, particularly in the development of annotated datasets for cell anomalies.
KW - Tabular anomaly detection
KW - Weakly supervised learning
UR - https://www.scopus.com/pages/publications/105040749202
U2 - 10.1016/j.neunet.2026.109202
DO - 10.1016/j.neunet.2026.109202
M3 - 文章
AN - SCOPUS:105040749202
SN - 0893-6080
VL - 203
JO - Neural Networks
JF - Neural Networks
M1 - 109202
ER -