Aerial-ground person re-identification method via position-aware and local pre-interaction
-
摘要:
针对现有空-地行人重识别方法中存在的行人特征判别性较差、模型泛化能力不足等问题,提出一种位置感知与局部预交互的空-地行人重识别方法。提出位置偏置自注意力机制,通过将位置信息融入注意力分数,引导模型聚焦输入序列中关键位置的令牌,增强模型对图块之间空间位置关系的感知能力,提升行人特征的判别性与鲁棒性;设计提示-局部预交互模块,通过一次较低成本的预交互,建立提示语义与局部特征之间的细粒度关联,增强提示向量对当前视角图像局部具体细节的感知能力,提升模型的局部细节重建能力;引入标签平滑正则化训练策略,通过将原始One-Hot编码的硬标签转换为软标签,使模型学习更平滑、泛化性更强的特征表示,缓解模型过拟合问题,提升模型的性能与泛化能力。在空-地行人重识别公开数据集LAGPeR上进行大量实验,结果表明,所提方法能够提升重识别性能,具有一定有效性。
-
关键词:
- 空-地行人重识别 /
- Transformer /
- 自注意力机制 /
- 局部特征预交互 /
- 标签平滑正则化
Abstract:In response to the issues of poor discriminative power in pedestrian features and insufficient model generalization in existing aerial-ground person re-identification methods, this paper proposes an aerial-ground person re-identification method via position-aware and local pre-interaction. A positionally biased self-attention mechanism is proposed, which incorporates positional information into attention scores to guide the model in focusing on tokens at key positions within the input sequence. This enhances the model's perception of spatial relationships between image patches and improves the discriminability and robustness of pedestrian features. A prompt-local pre-interaction module is created that uses a single, inexpensive pre-interaction to create fine-grained linkages between prompt semantics and local attributes. This strengthens the prompt vector's ability to perceive specific local details in the current view and enhances the model's capability for fine-grained detail reconstruction. The label smoothing regularization training strategy is introduced, converting original one-hot encoded hard labels into soft labels. This encourages the model to learn smoother and more generalizable feature representations, mitigates overfitting, and improves overall model performance and generalization. The validity of the proposed strategy is confirmed by extensive trials on the public aerial-ground person re-identification dataset LAGPeR, which show that it effectively enhances Re-Identification performance.
-
表 1 LAGPeR数据集情况介绍
Table 1. Introduction to LAGPeR dataset
实验设置 子集 视角 摄像头个数 行人个数 图像张数 Train A+G 12 2708 40770 A→G Query A 3 1523 3046 Gallery G 6 1523 15533 G→A Query G 6 1523 3046 Gallery A 3 1523 7717 G→A+G Query G 6 1523 3046 Gallery A+G 9 1523 20204 表 2 在LAGPeR数据集上的对比实验结果
Table 2. Comparative experimental results on LAGPeR dataset
% 方法 主干网络 Rank-1 mAP mINP A→G G→A G→A+G A→G G→A G→A+G A→G G→A G→A+G ViT[13] ViT 38.67 32.04 18.88 27.25 30.69 15.31 TransReID[22] ViT 38.80 33.00 22.90 28.80 32.10 18.80 CLIP-ReID[23] CLIP 24.40 21.30 12.30 17.60 20.80 10.20 MIP[24] ViT 39.30 33.90 21.00 29.30 32.60 17.30 AG-ReID[9] ViT 40.48 32.96 22.03 28.89 31.91 17.89 VDT[15] ViT 40.15 33.55 19.50 28.97 31.98 16.45 SeCap[16] ViT 39.89 33.68 21.04 28.67 32.00 16.80 11.16 23.04 6.59 本文方法 ViT 41.99 34.93 22.23 30.52 33.23 18.17 12.32 24.02 7.24 表 3 在LAGPeR数据集上的消融实验结果
Table 3. Results of ablation study on LAGPeR dataset
% PBSAM PLPIM LSR Rank-1 mAP mINP A→G G→A G→A+G A→G G→A G→A+G A→G G→A G→A+G 39.89 33.68 21.04 28.67 32.00 16.80 11.16 23.04 6.59 √ 40.28 33.29 21.57 29.20 32.12 17.42 11.44 23.18 6.79 √ 41.27 35.06 21.96 29.72 32.69 17.29 11.88 23.35 6.62 √ 39.79 34.24 21.31 28.96 32.34 17.22 11.08 23.39 7.02 √ √ 41.92 34.50 21.50 29.74 33.03 17.40 11.48 23.93 6.80 √ √ √ 41.99 34.93 22.23 30.52 33.23 18.17 12.32 24.02 7.24 -
[1] 孙义博, 张文靖, 王蓉, 等. 基于通道注意力机制的行人重识别方法[J]. 北京航空航天大学学报, 2022, 48(5): 881-889.Sun Y B, Zhang W J, Wang R, et al. Pedestrian re-identification method based on channel attention mechanism[J]. Journal of Beijing University of Aeronautics and Astronautics, 2022, 48(5): 881-889(in Chinese). [2] 宋晓勇, 孙学宏, 刘丽萍, 等. 基于多细粒度双流网络的行人重识别[J]. 电子测量与仪器学报, 2025, 39(8): 250-257.Song X Y, Sun X H, Liu L P, et al. Pedestrian re-identification based on a multi-granularity dual-stream network[J]. Journal of Electronic Measurement and Instrumentation, 2025, 39(8): 250-257(in Chinese). [3] Huang Y K, Zha Z J, Fu X Y, et al. Illumination-invariant person re-identification[C]//Proceedings of the 27th ACM International Conference on Multimedia. New York: ACM, 2019: 365-373. [4] Hou R B, Ma B P, Chang H, et al. VRSTC: occlusion-free video person re-identification[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2020: 7176-7185. [5] Huang H J, Li D W, Zhang Z, et al. Adversarially occluded samples for person re-identification[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2018: 5098-5107. [6] Li X, Zheng W S, Wang X J, et al. Multi-scale learning for low-resolution person re-identification[C]//Proceedings of the IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2016: 3765-3773. [7] Wang Y, Wang L Q, You Y R, et al. Resource aware person re-identification across multiple resolutions[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2018: 8042-8051. [8] Schumann A, Metzler J. Person re-identification across aerial and ground-based cameras by deep feature fusion[J]. Automatic Target Recognition XXVII, 2017, 10202: 102020A. [9] Nguyen H, Nguyen K, Sridharan S, et al. Aerial-ground person re-ID[C]//Proceedings of the IEEE International Conference on Multimedia and Expo. Piscataway: IEEE Press, 2023: 2585-2590. [10] 贝俊仁, 张权, 赖剑煌. 基于隐式解码对齐的空地行人重识别方法[J]. 自动化学报, 2025, 51(9): 1988-2000.Bei J R, Zhang Q, Lai J H. Implicit decoder alignment for aerial-ground person re-identification[J]. Acta Automatica Sinica, 2025, 51(9): 1988-2000(in Chinese). [11] Wang Y H, Hu X, Wang L X, et al. SD-ReID: view-aware stable diffusion for aerial-ground person re-identification[EB/OL]. (2026-05-14)[2026-05-29]. https://arxiv.org/abs/2504.09549. [12] Mei L, Cheng Y W, Chen H X, et al. Unsupervised aerial-ground re-identification from pedestrian to group for UAV-based surveillance[J]. Drones, 2025, 9(4): 244. [13] Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: Transformers for image recognition at scale[EB/OL]. (2021-06-03)[2025-12-02]. https://arxiv.org/abs/2010.11929. [14] Szegedy C, Vanhoucke V, Ioffe S, et al. Rethinking the inception architecture for computer vision[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2016: 2818-2826. [15] Zhang Q, Wang L, Patel V M, et al. View-decoupled transformer for person re-identification under aerial-ground camera network[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2024: 22000-22009. [16] Wang S N, Wang Y L, Wu R Q, et al. SeCap: self-calibrating and adaptive prompts for cross-view person re-identification in aerial-ground networks[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2025: 22119-22128. [17] Wang X G, Doretto G, Sebastian T, et al. Shape and appearance context modeling[C]//Proceedings of the IEEE 11th International Conference on Computer Vision. Piscataway: IEEE Press, 2007: 1-8. [18] Zheng L, Shen L Y, Tian L, et al. Scalable person re-identification: a benchmark[C]//Proceedings of the IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2016: 1116-1124. [19] Ye M, Shen J B, Lin G J, et al. Deep learning for person re-identification: a survey and outlook[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(6): 2872-2893. [20] Deng J, Dong W, Socher R, et al. ImageNet: a large-scale hierarchical image database[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2009: 248-255. [21] Kingma D P, BA J. Adam: a method for stochastic optimization[EB/OL]. (2017-01-30)[2025-12-02]. https://arxiv.org/abs/1412.6980. [22] He S T, Luo H, Wang P C, et al. TransReID: Transformer-based object re-identification[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE Press, 2022: 14993-15002. [23] Li S Y, Sun L, Li Q L. CLIP-ReID: exploiting vision-language model for image re-identification without concrete text labels[EB/OL]. (2023-01-01)[2025-12-02]. https://arxiv.org/abs/2211.13977. [24] Wu R Q, Jiao B L, Wang W X, et al. Enhancing visible-infrared person re-identification with modality- and instance-aware visual prompt learning[EB/OL]. (2024-06-18)[2025-12-02]. https://arxiv.org/abs/2406.12316. -


下载: