留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

基于Mamba-CNN的多尺度异常行为检测方法

石洋宇 谢承杰 郑棣文 卢树华

石洋宇,谢承杰,郑棣文,等. 基于Mamba-CNN的多尺度异常行为检测方法[J]. 北京航空航天大学学报,2026,52(7):2672-2680
引用本文: 石洋宇,谢承杰,郑棣文,等. 基于Mamba-CNN的多尺度异常行为检测方法[J]. 北京航空航天大学学报,2026,52(7):2672-2680
Shi Y Y,Xie C J,Zheng D W,et al. Multi-scale anomaly behavior detection method based on Mamba-CNN[J]. Journal of Beijing University of Aeronautics and Astronautics,2026,52(7):2672-2680 (in Chinese)
Citation: Shi Y Y,Xie C J,Zheng D W,et al. Multi-scale anomaly behavior detection method based on Mamba-CNN[J]. Journal of Beijing University of Aeronautics and Astronautics,2026,52(7):2672-2680 (in Chinese)

基于Mamba-CNN的多尺度异常行为检测方法

doi: 10.13700/j.bh.1001-5965.2024.0416
基金项目: 

中国人民公安大学双一流创新研究项目(2024yjsky036)

详细信息
    通讯作者:

    E-mail:lushuhua@ppsuc.edu.cn

  • 中图分类号: TP391.4

Multi-scale anomaly behavior detection method based on Mamba-CNN

Funds: 

Double First-Class Innovation Research Project for People’s Public Security University of China (2024yjsky036)

More Information
  • 摘要:

    针对无监督异常行为检测任务中预测器易衍生出异常泛化及被检目标存在尺度差异等问题,提出一种基于Mamba模型改进的U型异常行为检测网络。网络从全局特征和局部特征2个方面改进,约束预测不良泛化能力。在编码器部分引入状态空间模型加强提取全局特征能力;设计多尺度空间通道融合(M-SCF)策略,融合不同感受野下特征信息,降低尺度差异对局部特征干扰;解码器采用跳跃连接丰富浅层特征信息,增强捕捉上下文信息的能力。所提方法在UCSD Ped2、Avenue和ShanghaiTech等公开数据集进行验证,识别准确率分别达到98.1%、89.8%和78.5%,与近年诸多先进算法相比,具有较高的准确度。结果表明,Mamba可有效提升异常行为检测精度。

     

  • 图 1  异常行为检测网络结构

    Figure 1.  Network structure of anomaly behavior detection

    图 2  编码器网络结构

    Figure 2.  Network structure of encoder

    图 3  SS2D扫描机制

    Figure 3.  SS2D scan mechanism

    图 4  M-SCF策略机制

    Figure 4.  M-SCF strategy mechanism

    图 5  解码器网络结构

    Figure 5.  Network structure of decoder

    图 6  数据集异常行为展示

    Figure 6.  Some samples of abnormal events of datasets

    图 7  检测结果异常分数曲线

    Figure 7.  Anomaly score curve of detection results

    图 8  消融实验结果对比

    Figure 8.  Comparison of the ablation study results

    表  1  数据集详细信息

    Table  1.   Details of datasets

    数据集 总计/帧 训练集/帧 测试集/帧 尺寸/像素×像素
    UCSD Ped2[27] 4 560 2 550 2 010 $ 360\times 240 $
    Avenue[28] 30 652 15 328 15 324 $ 640\times 360 $
    ShanghaiTech[29] 317 398 274 515 42 883 $ 856\times 480 $
    下载: 导出CSV

    表  2  不同数据集上的AUC结果对比

    Table  2.   Comparison of AUC on different datasets

    类别 方法 年份 AUC/%
    UCSD Ped2
    数据集[27]
    Avenue
    数据集[28]
    ShanghaiTech
    数据集[29]
    Re Conv-AE[12] 2016 85.0 80.0 60.9
    MemAE[13] 2019 94.1 83.3 71.2
    MNAD[6] 2020 90.2 82.8 69.8
    sRNN-AE[14] 2021 92.2 83.5 69.6
    Huang等[15] 2023 95.5 87.6 76.6
    Pr FFP[16] 2018 95.4 84.9 72.8
    MNAD[6] 2020 97.0 88.5 70.5
    Multisapce[17] 2021 95.4 86.8 73.6
    AMMC-Net[18] 2021 96.9 86.6 73.7
    STC-Net[30] 2022 96.7 87.8 73.1
    LGN-Net[8] 2022 97.1 89.3 73.0
    AMUM-AE[31] 2023 96.6 86.2
    Le等[32] 2023 97.4 86.7 73.6
    Wang 等[33] 2023 97.7 90.9 71.3
    Wang等[34] 2023 96.5 92.2 75.4
    USTN-DSC[35] 2023 98.1 89.9 74.7
    TransMem[36] 2024 98.1 88.5 72.5
    本文 2024 98.1 89.8 78.5
     注:加粗数值表示性能最优。
    下载: 导出CSV

    表  3  UCSD Ped2数据集上消融实验

    Table  3.   Ablation study on UCSD Ped2 dataset

    $ {L}_{\mathrm{df}} $ Mamba Skip M-SCF AUC/%
    95.9
    96.2
    96.5
    97.3
    98.1
    下载: 导出CSV

    表  4  M-SCF策略消融实验

    Table  4.   Ablation study of M-SCF

    方法 UCSD Ped2 AUC/%
    MNAD* 96.3
    MNAD* + M-SCF 97.0
    LGN-Net* 95.4
    LGN-Net* + M-SCF 96.3
     注:“*”表示模型未调整参数的实际复现结果。
    下载: 导出CSV
  • [1] Yu J, Lee Y, Yow K C, et al. Abnormal event detection and localization via adversarial event prediction[J]. IEEE Transactions on Neural Networks and Learning Systems, 2022, 33(8): 3572-3586.
    [2] Chen D Y, Wang P T, Yue L Y, et al. Anomaly detection in surveillance video based on bidirectional prediction[J]. Image and Vision Computing, 2020, 98: 103915.
    [3] Dong F, Zhang Y, Nie X S. Dual discriminator generative adversarial network for video anomaly detection[J]. IEEE Access, 2020, 8: 88170-88176.
    [4] Lee S, Kim H G, Ro Y M. BMAN: bidirectional multi-scale aggregation networks for abnormal event detection[J]. IEEE Transactions on Image Processing, 2020, 29: 2395-2408.
    [5] Pang G S, Shen C H, Cao L B, et al. Deep learning for anomaly detection: a review[J]. ACM Computing Surveys (CSUR), 2021, 54(2): 1-38.
    [6] Park H, Noh J, Ham B. Learning memory-guided normality for anomaly detection[C]//Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2020: 14360-14369.
    [7] V H, Chen C, Cui Z, et al. Learning normal dynamics in videos with meta prototype network[C]//Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2021: 15420-15429.
    [8] Zhao M Y, Zeng X H, Liu Y, et al. LGN-Net: local-global normality network for video anomaly detection[EB/OL]. (2023-01-08)[2024-02-03]. https://doi.org/10.48550/arXiv.07454.
    [9] Gong M G, Zeng H M, Xie Y, et al. Local distinguishability aggrandizing network for human anomaly detection[J]. Neural Networks, 2020, 122(C): 364-373.
    [10] Liu W R, Chang H, Ma B P, et al. Diversity-measurable anomaly detection[C]//Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2023: 12147-12156.
    [11] Gu A, Dao T. Mamba: linear-time sequence modeling with selective state spaces[EB/OL]. (2023-12-01)[2024-02-04]. https://doi.org/10.48550/arXiv.2312.00752.
    [12] Hasan M, Choi J, Neumann J, et al. Learning temporal regularity in video sequences[C]//Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2016: 733-742.
    [13] Gong D, Liu L Q, Le V, et al. Memorizing normality to detect anomaly: memory-augmented deep autoencoder for unsupervised anomaly detection[C]//Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE Press, 2020: 1705-1714.
    [14] Luo W X, Liu W, Lian D Z, et al. Video anomaly detection with sparse coding inspired deep neural networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43(3): 1070-1084.
    [15] Huang X Y, Zhao C D, Gao C X, et al. Synthetic pseudo anomalies for unsupervised video anomaly detection: a simple yet efficient framework based on masked autoencoder[C]//Proceedings of the ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing. Piscataway: IEEE Press, 2023: 1-5.
    [16] Liu W, Luo W X, Lian D Z, et al. Future frame prediction for anomaly detection-a new baseline[C]//Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2018: 6536-6545.
    [17] Hang Y, Nie X S, He R D, et al. Normality learning in multispace for video anomaly detection[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 31(9): 3694-3706.
    [18] Cai R C, Zhang H, Liu W, et al. Appearance-motion memory consistency network for video anomaly detection[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2021, 35(2): 938-946.
    [19] Wang J, Chen J, Chen D, et al. Large window-based mamba unet for medical image segmentation: beyond convolution and self-attention[EB/OL]. arXiv preprint arXiv: 2403.07332, 2024.
    [20] Ruan J C, Li J C, Xiang S C. VM-UNet: vision mamba UNet for medical image segmentation[EB/OL]. (2024-02-04)[2024-03-02]. https://doi.org/10.48550/arXiv.2402.02491.
    [21] Ma J, Li F F, Wang B. U-mamba: enhancing long-range dependency for biomedical image segmentation[EB/OL]. (2024-01-09)[2024-03-02]. https://doi.org/10.48550/arXiv.2401.04722.
    [22] Liu Y, Tian Y J, Zhao Y Z, et al. VMamba: visual state space model[EB/OL]. (2024-01-18)[2024-03-03]. https://doi.org/10.48550/arXiv.2401.10166.
    [23] Liu J R, Yang H, Zhou H Y, et al. Swin-UMamba: mamba-based UNet with ImageNet-based pretraining[EB/OL]. (2024-02-05)[2024-03-05]. https://doi.org/10.48550/arXiv.2402.03302.
    [24] Gu A, Goel K, Ré C. Efficiently modeling long sequences with structured state spaces[EB/OL]. (2022-08-05)[2024-03-06]. https://doi.org/10.48550/arXiv.2111.00396.
    [25] Ronneberger O, Fischer P, Brox T. U-Net: convolutional networks for biomedical image segmentation[C]//Proceedings of the Medical Image Computing and Computer-Assisted Intervention-MICCAI 2015. Berlin: Springer, 2015: 234-241.
    [26] Dai J F, Qi H Z, Xiong Y W, et al. Deformable convolutional networks[C]//Proceedings of the 2017 IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2017: 764-773.
    [27] Wang S, Miao Z J. Anomaly detection in crowd scene[C]//Proceedings of the IEEE 10th International Conference on Signal Processing Proceedings. Piscataway: IEEE Press, 2010: 1220-1223.
    [28] Lu C W, Shi J P, Jia J Y. Abnormal event detection at 150 FPS in MATLAB[C]//Proceedings of the 2013 IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2014: 2720-2727.
    [29] Luo W X, Liu W, Gao S H. A revisit of sparse coding based anomaly detection in stacked RNN framework[C]//Proceedings of the 2017 IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2017: 341-349.
    [30] Zhao M Y, Liu Y, Liu J, et al. Exploiting spatial-temporal correlations for video anomaly detection[C]//Proceedings of the 2022 26th International Conference on Pattern Recognition. Piscataway: IEEE Press, 2022: 1727-1733.
    [31] Jin Z H, Zhao Y X. Appearance-motion united memory autoencoder for video anomaly detection[C]//Proceedings of the 2023 International Conference on Cyber-Enabled Distributed Computing and Knowledge Discovery. Piscataway: IEEE Press, 2024: 179-187.
    [32] Le V T, Kim Y G. Attention-based residual autoencoder for video anomaly detection[J]. Applied Intelligence, 2023, 53(3): 3240-3254.
    [33] Wang L, Tian J W, Zhou S P, et al. Memory-augmented appearance-motion network for video anomaly detection[J]. Pattern Recognition, 2023, 138: 109335.
    [34] Wang Z Q, Gu X J, Hu J Y, et al. Ensemble anomaly score for video anomaly detection using denoise diffusion model and motion filters[J]. Neurocomputing, 2023, 553: 126589.
    [35] Yang Z W, Liu J, Wu Z Y, et al. Video event restoration based on keyframes for video anomaly detection[C]//Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2023: 14592-14601.
    [36] Wang Z Q, Gu X J, Gu X S, et al. Enhancing video anomaly detection with learnable memory network: a new approach to memory-based auto-encoders[J]. Computer Vision and Image Understanding, 2024, 241: 103946.
  • 加载中
图(8) / 表(4)
计量
  • 文章访问数:  623
  • HTML全文浏览量:  178
  • PDF下载量:  47
  • 被引次数: 0
出版历程
  • 收稿日期:  2024-06-07
  • 录用日期:  2024-10-04
  • 网络出版日期:  2024-11-15
  • 整期出版日期:  2026-07-31

目录

    /

    返回文章
    返回
    常见问答