Cross-layer high-efficiency phase-aware Transformer for colorectal polyp image segmentation algorithm
-
摘要:
针对结直肠息肉图像病变区域形状不规则、边缘轮廓模糊及与正常区域相似性高,导致细节信息丢失和病变区域误分割等问题,提出一种跨层高效相位感知Transformer的结直肠息肉图像分割算法。利用金字塔视觉Transformer分段编码器逐层提取输入特征图的全局语义信息和空间细节,多尺度解析结直肠息肉病变特征;通过极化自注意力模块对病变特征进行回归预测,加深特征语义信息关联性;设计高效相位感知模块,同时提取全局与局部信息,精准定位病变位置;采用跨层融合与传播模块,整合边缘细节,提升高级特征复用率。在CVC-ClinicDB、Kvasir-SEG、ETIS-LaribPolypDB、CVC-ConlonDB和CVC-T这5个数据集上进行实验,Dice指数分别为0.940、0.923、0.801、0.810和0.896,分割性能优于现有CaraNet、MSRAFormer等网络。结果表明,对于空间结构复杂、边缘模糊的结直肠息肉图像,所提算法均有较高的分割精度。
-
关键词:
- 结直肠息肉 /
- Transformer /
- 极化自注意力模块 /
- 高效相位感知模块 /
- 跨层融合与传播模块
Abstract:This paper proposes a cross-layer efficient phase-aware Transformer segmentation algorithm for colorectal polyps in order to address the issues of irregular shape of the lesion region, fuzzy edge contour, and high similarity with normal region, which result in the loss of detail information and mis-segmentation of the lesion region. Firstly, the pyramid vision Transformer encoder is used to extract the global semantic information and spatial details of the input feature map layer by layer, and to analyze the colorectal polyp lesion features at multiple scales; secondly, the polarized self-attention module is used to regressively predict the lesion features, and to deepen the correlation of the semantic information of the features;thirdly, the high-efficiency phase-aware module is designed to extract the global and local information to precisely. The final one is the cross-layer fusion and propagation module, which enhances the rate of advanced feature reuse by integrating the edge details. Experiments were conducted on five datasets: CVC-ClinicDB, Kvasir-SEG, ETIS-LaribPolypDB, CVC-ConlonDB, and CVC-T, achieving Dice coefficients of 0.940, 0.923, 0.801, 0.810, and 0.896, respectively. This demonstrates superior segmentation performance over existing networks such as CaraNet and MSRAFormer. Both colorectal polyp images with fuzzy edges and complicated spatial organization exhibit great segmentation accuracy, according to the evaluation results.
-
表 1 数据集划分及尺寸
Table 1. Dataset partitioning and sizing
数据集 训练集 测试集 图片尺寸 CVC-ClinicDB 550 62 352×352 Kvasir-SEG 900 100 352×352 ETIS-LaribPolyDB 0 196 352×352 CVC-ConlonDB 0 380 352×352 CVC-T 0 60 352×352 表 2 实验环境
Table 2. Experimental environment
实验环境 设置 系统 Windows11 处理器 Inter Core i5-13600 CPU 显卡 NVIDIA GeForce RTX 4070Ti GPU CUDA CUDA 12.1 深度学习框架 Pytorch 3.9 表 3 不同算法在CVC-ClinicDB和Kvasir-SEG数据集上的对比
Table 3. Comparison of different algorithms on CVC-ClinicDB and Kvasir-SEG datasets
数据集 算法 Dice MIoU R PC F2 MAE U-Net[15] 0.824 0.755 0.834 0.839 0.827 0.019 PraNet 0.901 0.851 0.910 0.907 0.901 0.009 CaraNet[31] 0.936 0.887 0.951 0.928 0.938 0.007 CVC-ClinicDB SSFormer-s[33] 0.918 0.874 0.904 0.939 0.909 0.007 MSRAFormer[34] 0.934 0.884 0.953 0.924 0.944 0.008 PolypPVT[32] 0.934 0.879 0.937 0.936 0.927 0.012 本文 0.940 0.894 0.968 0.948 0.951 0.005 U-Net[15] 0.820 0.746 0.856 0.857 0.827 0.054 PraNet 0.900 0.842 0.911 0.915 0.901 0.029 CaraNet[31] 0.918 0.867 0.912 0.938 0.914 0.023 Kvasir-SEG SSFormer-s[33] 0.926 0.873 0.917 0.941 0.920 0.017 MSRAFormer[34] 0.919 0.870 0.921 0.938 0.918 0.020 PolypPVT[32] 0.920 0.868 0.913 0.907 0.910 0.023 本文 0.923 0.876 0.925 0.954 0.923 0.023 表 4 不同算法在ETIS-LaribPolypDB、CVC-ConlonDB和CVC-T数据集上的对比
Table 4. Comparison of different algorithms on ETIS-LaribPolypDB、CVC-ConlonDB and CVC-T datasets
数据集 算法 Dice MIoU R PC F2 MAE U-Net[15] 0.511 0.437 0.523 0.620 0.510 0.058 PraNet 0.716 0.642 0.739 0.755 0.717 0.043 CaraNet[31] 0.747 0.672 0.811 0.731 0.777 0.017 ETIS-LaribPolypDB SSFormer-s[33] 0.769 0.694 0.856 0.743 0.800 0.016 MSRAFormer[34] 0.747 0.673 0.823 0.786 0.781 0.012 PolypPVT[32] 0.789 0.703 0.869 0.748 0.816 0.020 本文 0.801 0.728 0.864 0.795 0.819 0.011 U-Net[15] 0.405 0.333 0.482 0.439 0.428 0.036 PraNet 0.630 0.566 0.686 0.628 0.648 0.031 CaraNet[31] 0.773 0.689 0.837 0.753 0.796 0.042 CVC-ColonDB SSFormer-s[33] 0.774 0.698 0.777 0.836 0.766 0.035 MSRAFormer[34] 0.782 0.707 0.803 0.872 0.787 0.028 PolypPVT[32] 0.804 0.725 0.835 0.822 0.812 0.031 本文 0.810 0.730 0.844 0.834 0.815 0.026 U-Net[15] 0.716 0.627 0.754 0.766 0.735 0.022 PraNet 0.872 0.797 0.841 0.940 0.904 0.009 CaraNet[31] 0.886 0.823 0.851 0.944 0.922 0.011 CVC-T SSFormer-s[33] 0.889 0.823 0.862 0.946 0.915 0.009 MSRAFormer[34] 0.894 0.828 0.863 0.954 0.927 0.006 PolypPVT[32] 0.886 0.817 0.848 0.942 0.912 0.008 本文 0.896 0.828 0.857 0.961 0.929 0.008 表 5 不同算法的参数量、浮点运算速度和单轮训练时长
Table 5. Number of parameters, floating-point operations per second, and single-round training duration for different algorithms
表 6 高效相位感知模块在CVC-ClinicDB和CVC- ColonDB数据集上的对比结果
Table 6. Comparison results of efficient phase-aware module on CVC-ClinicDB and CVC- ColonDB datasets
数据集 算法 Dice MIoU R PC F2 MAE 文献[24] 0.934 0.886 0.940 0.940 0.934 0.008 CVC-ClinicDB 文献[25] 0.928 0.882 0.925 0.945 0.932 0.012 本文 0.940 0.894 0.968 0.948 0.951 0.005 文献[24] 0.788 0.705 0.831 0.802 0.786 0.035 CVC-ColonDB 文献[25] 0.786 0.705 0.799 0.849 0.801 0.042 本文 0.810 0.730 0.844 0.834 0.815 0.026 表 7 各模块在CVC-ClinicDB数据集上的消融结果
Table 7. Ablation results for each module on CVC-ClinicDB dataset
模块 解码器 Dice MIoU R PC F2 极化 高效 混合 M1 √ √ 0.933 0.888 0.928 0.934 0.933 M2 √ √ 0.923 0.874 0.922 0.946 0.920 M3 √ √ 0.908 0.855 0.901 0.943 0.903 M4 √ √ √ 0.940 0.894 0.938 0.948 0.951 表 8 各模块在ETIS-LaribPolypDB数据集上的消融结果
Table 8. Ablation results for each module on ETIS-LaribPolypDB dataset
模块 解码器 Dice MIoU R PC F2 极化 高效 混合 M1 √ √ 0.786 0.712 0.830 0.830 0.804 M2 √ √ 0.764 0.693 0.793 0.801 0.773 M3 √ √ 0.703 0.619 0.668 0.731 0.748 M4 √ √ √ 0.801 0.728 0.864 0.795 0.819 -
[1] Liang H, Cheng Z M, Zhong H Q, et al. A region-based convolutional network for nuclei detection and segmentation in microscopy images[J]. Biomedical Signal Processing and Control, 2022, 71: 103276. [2] Banik D, Roy K, Bhattacharjee D, et al. Polyp-Net: a multimodel fusion network for polyp segmentation[J]. IEEE Transactions on Instrumentation and Measurement, 2021, 70: 4000512. [3] Jha D, Smedsrud P H, Johansen D, et al. A comprehensive study on colorectal polyp segmentation with ResUNet++, conditional random field and test-time augmentation[J]. IEEE Journal of Biomedical and Health Informatics, 2021, 25(6): 2029-2040. [4] 关少亚, 张诚, 孟偲, 等. 基于相位对称性的血管超声图像分割算法[J]. 北京航空航天大学学报, 2023, 49(10): 2645-2650.Guan S Y, Zhang C, Meng S, et al. Vascular ultrasound image segmentation algorithm based on phase symmetry[J]. Journal of Beijing University of Aeronautics and Astronautics, 2023, 49(10): 2645-2650(in Chinese). [5] Zhou L S, Liang L M, Sheng X Q. GA-Net: ghost convolution adaptive fusion skin lesion segmentation network[J]. Computers in Biology and Medicine, 2023, 164: 107273. [6] Shao D G, Xu C R, Xiang Y, et al. Ultrasound image segmentation with multilevel threshold based on differential search algorithm[J]. IET Image Processing, 2019, 13(6): 998-1005. [7] Eriyanti N A, Sigit R, Harsono T. Classification of colon polyp on endoscopic image using support vector machine[C]//Proceedings of the International Electronics Symposium. Piscataway: IEEE Press, 2021: 244-250. [8] Senthilkumaran N, Vaithegi S. Image segmentation by using thresholding techniques for medical images[J]. Computer Science & Engineering, 2016, 6(1): 1-13. [9] Yue G H, Han W W, Jiang B, et al. Boundary constraint network with cross layer feature integration for polyp segmentation[J]. IEEE Journal of Biomedical and Health Informatics, 2022, 26(8): 4090-4099. [10] Näppi J, Yoshida H. Automated detection of polyps with CT colonography evaluation of volumetric features for reduction of false-positive findings[J]. Academic Radiology, 2002, 9(4): 386-397. [11] Jerebko A K, Franaszek M, Summers R M. Radon transform based polyp segmentation method for CT colonography computer aided diagnosis[C]//Proceedings of the Society of Photo-Optical Instrumentation Engineers Conference on Medical Imaging: Physiology and Function-Methods, Systems, and Applications. Bellingham: SPIE, 2002, 225: 257-258. [12] Yao J H, Miller M, Franaszek M, et al. Colonic polyp segmentation in CT colonography-based on fuzzy clustering and deformable models[J]. IEEE Transactions on Medical Imaging, 2004, 23(11): 1344-1352. [13] Wu H S, Zhao Z B, Zhong J F, et al. PolypSeg: a lightweight context-aware network for real-time polyp segmentation[J]. IEEE Transactions on Cybernetics, 2023, 53(4): 2610-2621. [14] Long J, Shelhamer E, Darrell T. Fully convolutional networks for semantic segmentation[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2015: 3431-3440. [15] Ronneberger O, Fischer P, Brox T. U-Net: convolutional networks for biomedical image segmentation[C]//Proceedings of the 18th International Conference on Medical Image Computing and Computer-Assisted Intervention. Berlin: Springer, 2015: 234-241. [16] Ali S M F, Khan M T, Haider S U, et al. Depth-wise separable atrous convolution for polyps segmentation in gastro[C]//Proceedings of the MediaEval. [S.l.:s.n.]: 2020. [17] Khan M A, Khan M A, Ahmed F, et al. Gastrointestinal diseases segmentation and classification based on duo-deep architectures[J]. Pattern Recognition Letters, 2020, 131: 193-204. [18] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. New York: ACM, 2017: 6000-6010. [19] Peng Z L, Huang W, Gu S Z, et al. Conformer: local features coupling global representations for visual recognition[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE Press, 2022: 357-366. [20] Wang J F, Huang Q M, Tang F L, et al. Stepwise feature fusion: local guides global[C]//Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention. Berlin: Springer, 2022: 110-120. [21] Wang W H, Xie E Z, Li X, et al. PVTv2: improved baselines with pyramid vision Transformer[J]. Computational Visual Media, 2022, 8(3): 415-424. [22] Liu H J, Liu F Q, FAN X Y, et al. Polarized self-attention: towards high-quality pixel-wise regression[EB/OL]. (2021-07-08)[2024-05-01]. https://arxiv.org/abs/2107.00782. [23] Woo S, Park J, Lee J Y, et al. CBAM: convolutional block attention module[C]//Proceedings of the European Conference on Computer Vision. Berlin: Springer, 2018: 3-19. [24] Wang Q L, Wu B G, Zhu P F, et al. ECA-Net: efficient channel attention for deep convolutional neural networks[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2020: 11531-11539. [25] Tang Y H, Han K, Guo J Y, et al. An image patch is a wave: phase-aware vision MLP[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2022: 10925-10934. [26] Zhou T, Zhou Y, Gong C, et al. Feature aggregation and propagation network for camouflaged object detection[J]. IEEE Transactions on Image Processing, 2022, 31: 7036-7047. [27] Bernal J, Sánchez F J, Fernández-esparrach G, et al. WM-DOVA maps for accurate polyp highlighting in colonoscopy: validation vs. saliency maps from physicians[J]. Computerized Medical Imaging and Graphics, 2015, 43: 99-111. [28] Dutta A, Zisserman A. The VIA annotation software for images, audio and video[C]//Proceedings of the 27th ACM International Conference on Multimedia. New York: ACM, 2019: 2276-2279. [29] Tajbakhsh N, Gurudu S R, Liang J M. Automated polyp detection in colonoscopy videos using shape and context information[J]. IEEE Transactions on Medical Imaging, 2016, 35(2): 630-644. [30] Silva J, Histace A, Romain O, et al. Toward embedded detection of polyps in WCE images for early diagnosis of colorectal cancer[J]. International Journal of Computer Assisted Radiology and Surgery, 2014, 9(2): 283-293. [31] Lou A G, Guan S Y, Ko H, et al. CaraNet: context axial reverse attention network for segmentation of small medical objects[C]// Proceedings of the Conference on Medical Imaging-Image Processing. Bellingham: SPIE, 2022, 12032: 81-92. [32] Dong B, Wang W H, Fan D P, et al. Polyp-PVT: polyp segmentation with pyramid vision transformers[EB/OL]. (2024-02-19)[2024-05-01]. https://arxiv.org/abs/2108.06932. [33] Shi W T, Xu J, Gao P. SSformer: a lightweight Transformer for semantic segmentation[C]//Proceedings of the IEEE 24th International Workshop on Multimedia Signal Processing. Piscataway: IEEE Press, 2022: 1-5. [34] Wu C, Long C, Li S J, et al. MSRAformer: multiscale spatial reverse attention network for polyp segmentation[J]. Computers in Biology and Medicine, 2022, 151: 106274. -


下载: