-
摘要:
针对无监督异常行为检测任务中预测器易衍生出异常泛化及被检目标存在尺度差异等问题,提出一种基于Mamba模型改进的U型异常行为检测网络。网络从全局特征和局部特征2个方面改进,约束预测不良泛化能力。在编码器部分引入状态空间模型加强提取全局特征能力;设计多尺度空间通道融合(M-SCF)策略,融合不同感受野下特征信息,降低尺度差异对局部特征干扰;解码器采用跳跃连接丰富浅层特征信息,增强捕捉上下文信息的能力。所提方法在UCSD Ped2、Avenue和ShanghaiTech等公开数据集进行验证,识别准确率分别达到98.1%、89.8%和78.5%,与近年诸多先进算法相比,具有较高的准确度。结果表明,Mamba可有效提升异常行为检测精度。
Abstract:A U-shaped network is suggested for anomaly behavior recognition based on Mamba in order to overcome difficulties in unsupervised anomaly behavior detection, such as the predictor’s propensity for abnormal generalization and the target objects' scale disparities. The network improves both global and local features to constrain undesired generalization ability in predictions. A state space model is introduced in the encoder to strengthen the extraction of global features. A multi-scale spatial channel fusion (M-SCF) strategy is designed to integrate feature information from different receptive fields, thereby reducing the interference of scale differences on local features. Skip connections are used in the decoder to enrich shallow feature information and enhance the ability to capture contextual information. The proposed method has been extensively validated on the UCSD Ped2, Avenue, and Shanghai Tech datasets, with respective recognition accuracies of 98.1%, 89.8%, and 78.5%. Mamba can successfully increase the accuracy of abnormal behavior identification, as evidenced by the findings, which demonstrate superior accuracy when compared to several sophisticated algorithms in recent years.
-
Key words:
- anomaly behavior detection /
- unsupervised learning /
- Mamba model /
- multi-scale /
- feature fusion
-
表 1 数据集详细信息
Table 1. Details of datasets
表 2 不同数据集上的AUC结果对比
Table 2. Comparison of AUC on different datasets
类别 方法 年份 AUC/% UCSD Ped2
数据集[27]Avenue
数据集[28]ShanghaiTech
数据集[29]Re Conv-AE[12] 2016 85.0 80.0 60.9 MemAE[13] 2019 94.1 83.3 71.2 MNAD[6] 2020 90.2 82.8 69.8 sRNN-AE[14] 2021 92.2 83.5 69.6 Huang等[15] 2023 95.5 87.6 76.6 Pr FFP[16] 2018 95.4 84.9 72.8 MNAD[6] 2020 97.0 88.5 70.5 Multisapce[17] 2021 95.4 86.8 73.6 AMMC-Net[18] 2021 96.9 86.6 73.7 STC-Net[30] 2022 96.7 87.8 73.1 LGN-Net[8] 2022 97.1 89.3 73.0 AMUM-AE[31] 2023 96.6 86.2 Le等[32] 2023 97.4 86.7 73.6 Wang 等[33] 2023 97.7 90.9 71.3 Wang等[34] 2023 96.5 92.2 75.4 USTN-DSC[35] 2023 98.1 89.9 74.7 TransMem[36] 2024 98.1 88.5 72.5 本文 2024 98.1 89.8 78.5 注:加粗数值表示性能最优。 表 3 UCSD Ped2数据集上消融实验
Table 3. Ablation study on UCSD Ped2 dataset
$ {L}_{\mathrm{df}} $ Mamba Skip M-SCF AUC/% 95.9 √ 96.2 √ √ 96.5 √ √ √ 97.3 √ √ √ √ 98.1 表 4 M-SCF策略消融实验
Table 4. Ablation study of M-SCF
方法 UCSD Ped2 AUC/% MNAD* 96.3 MNAD* + M-SCF 97.0 LGN-Net* 95.4 LGN-Net* + M-SCF 96.3 注:“*”表示模型未调整参数的实际复现结果。 -
[1] Yu J, Lee Y, Yow K C, et al. Abnormal event detection and localization via adversarial event prediction[J]. IEEE Transactions on Neural Networks and Learning Systems, 2022, 33(8): 3572-3586. [2] Chen D Y, Wang P T, Yue L Y, et al. Anomaly detection in surveillance video based on bidirectional prediction[J]. Image and Vision Computing, 2020, 98: 103915. [3] Dong F, Zhang Y, Nie X S. Dual discriminator generative adversarial network for video anomaly detection[J]. IEEE Access, 2020, 8: 88170-88176. [4] Lee S, Kim H G, Ro Y M. BMAN: bidirectional multi-scale aggregation networks for abnormal event detection[J]. IEEE Transactions on Image Processing, 2020, 29: 2395-2408. [5] Pang G S, Shen C H, Cao L B, et al. Deep learning for anomaly detection: a review[J]. ACM Computing Surveys (CSUR), 2021, 54(2): 1-38. [6] Park H, Noh J, Ham B. Learning memory-guided normality for anomaly detection[C]//Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2020: 14360-14369. [7] V H, Chen C, Cui Z, et al. Learning normal dynamics in videos with meta prototype network[C]//Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2021: 15420-15429. [8] Zhao M Y, Zeng X H, Liu Y, et al. LGN-Net: local-global normality network for video anomaly detection[EB/OL]. (2023-01-08)[2024-02-03]. https://doi.org/10.48550/arXiv.07454. [9] Gong M G, Zeng H M, Xie Y, et al. Local distinguishability aggrandizing network for human anomaly detection[J]. Neural Networks, 2020, 122(C): 364-373. [10] Liu W R, Chang H, Ma B P, et al. Diversity-measurable anomaly detection[C]//Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2023: 12147-12156. [11] Gu A, Dao T. Mamba: linear-time sequence modeling with selective state spaces[EB/OL]. (2023-12-01)[2024-02-04]. https://doi.org/10.48550/arXiv.2312.00752. [12] Hasan M, Choi J, Neumann J, et al. Learning temporal regularity in video sequences[C]//Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2016: 733-742. [13] Gong D, Liu L Q, Le V, et al. Memorizing normality to detect anomaly: memory-augmented deep autoencoder for unsupervised anomaly detection[C]//Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE Press, 2020: 1705-1714. [14] Luo W X, Liu W, Lian D Z, et al. Video anomaly detection with sparse coding inspired deep neural networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43(3): 1070-1084. [15] Huang X Y, Zhao C D, Gao C X, et al. Synthetic pseudo anomalies for unsupervised video anomaly detection: a simple yet efficient framework based on masked autoencoder[C]//Proceedings of the ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing. Piscataway: IEEE Press, 2023: 1-5. [16] Liu W, Luo W X, Lian D Z, et al. Future frame prediction for anomaly detection-a new baseline[C]//Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2018: 6536-6545. [17] Hang Y, Nie X S, He R D, et al. Normality learning in multispace for video anomaly detection[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 31(9): 3694-3706. [18] Cai R C, Zhang H, Liu W, et al. Appearance-motion memory consistency network for video anomaly detection[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2021, 35(2): 938-946. [19] Wang J, Chen J, Chen D, et al. Large window-based mamba unet for medical image segmentation: beyond convolution and self-attention[EB/OL]. arXiv preprint arXiv: 2403.07332, 2024. [20] Ruan J C, Li J C, Xiang S C. VM-UNet: vision mamba UNet for medical image segmentation[EB/OL]. (2024-02-04)[2024-03-02]. https://doi.org/10.48550/arXiv.2402.02491. [21] Ma J, Li F F, Wang B. U-mamba: enhancing long-range dependency for biomedical image segmentation[EB/OL]. (2024-01-09)[2024-03-02]. https://doi.org/10.48550/arXiv.2401.04722. [22] Liu Y, Tian Y J, Zhao Y Z, et al. VMamba: visual state space model[EB/OL]. (2024-01-18)[2024-03-03]. https://doi.org/10.48550/arXiv.2401.10166. [23] Liu J R, Yang H, Zhou H Y, et al. Swin-UMamba: mamba-based UNet with ImageNet-based pretraining[EB/OL]. (2024-02-05)[2024-03-05]. https://doi.org/10.48550/arXiv.2402.03302. [24] Gu A, Goel K, Ré C. Efficiently modeling long sequences with structured state spaces[EB/OL]. (2022-08-05)[2024-03-06]. https://doi.org/10.48550/arXiv.2111.00396. [25] Ronneberger O, Fischer P, Brox T. U-Net: convolutional networks for biomedical image segmentation[C]//Proceedings of the Medical Image Computing and Computer-Assisted Intervention-MICCAI 2015. Berlin: Springer, 2015: 234-241. [26] Dai J F, Qi H Z, Xiong Y W, et al. Deformable convolutional networks[C]//Proceedings of the 2017 IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2017: 764-773. [27] Wang S, Miao Z J. Anomaly detection in crowd scene[C]//Proceedings of the IEEE 10th International Conference on Signal Processing Proceedings. Piscataway: IEEE Press, 2010: 1220-1223. [28] Lu C W, Shi J P, Jia J Y. Abnormal event detection at 150 FPS in MATLAB[C]//Proceedings of the 2013 IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2014: 2720-2727. [29] Luo W X, Liu W, Gao S H. A revisit of sparse coding based anomaly detection in stacked RNN framework[C]//Proceedings of the 2017 IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2017: 341-349. [30] Zhao M Y, Liu Y, Liu J, et al. Exploiting spatial-temporal correlations for video anomaly detection[C]//Proceedings of the 2022 26th International Conference on Pattern Recognition. Piscataway: IEEE Press, 2022: 1727-1733. [31] Jin Z H, Zhao Y X. Appearance-motion united memory autoencoder for video anomaly detection[C]//Proceedings of the 2023 International Conference on Cyber-Enabled Distributed Computing and Knowledge Discovery. Piscataway: IEEE Press, 2024: 179-187. [32] Le V T, Kim Y G. Attention-based residual autoencoder for video anomaly detection[J]. Applied Intelligence, 2023, 53(3): 3240-3254. [33] Wang L, Tian J W, Zhou S P, et al. Memory-augmented appearance-motion network for video anomaly detection[J]. Pattern Recognition, 2023, 138: 109335. [34] Wang Z Q, Gu X J, Hu J Y, et al. Ensemble anomaly score for video anomaly detection using denoise diffusion model and motion filters[J]. Neurocomputing, 2023, 553: 126589. [35] Yang Z W, Liu J, Wu Z Y, et al. Video event restoration based on keyframes for video anomaly detection[C]//Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2023: 14592-14601. [36] Wang Z Q, Gu X J, Gu X S, et al. Enhancing video anomaly detection with learnable memory network: a new approach to memory-based auto-encoders[J]. Computer Vision and Image Understanding, 2024, 241: 103946. -


下载: