Volume 52 Issue 9
Sep.  2026
Turn off MathJax
Article Contents
Liu S Y,He L,Zeng J H. Fine-grained semantic-enhanced cross-modal image-text retrieval method for civil aviation[J]. Journal of Beijing University of Aeronautics and Astronautics,2026,52(9):3075-3088 (in Chinese)
Citation: Liu S Y,He L,Zeng J H. Fine-grained semantic-enhanced cross-modal image-text retrieval method for civil aviation[J]. Journal of Beijing University of Aeronautics and Astronautics,2026,52(9):3075-3088 (in Chinese)

Fine-grained semantic-enhanced cross-modal image-text retrieval method for civil aviation

doi: 10.13700/j.bh.1001-5965.2025.0549
Funds:

Civil Aircraft Special Research Project of the Ministry of Industry and Information Technology (MJZ2-3N21)

More Information
  • Corresponding author: E-mail:heliu1219@126.com
  • Received Date: 06 Aug 2025
  • Accepted Date: 29 Sep 2025
  • Available Online: 13 Oct 2025
  • Publish Date: 09 Oct 2025
  • In the fields of low-altitude economy and civil aviation, security assurance tasks heavily rely on the efficient correlation of cross-modal information such as images and texts. However, while mainstream cross-modal retrieval models perform well on general datasets, they underperform in these areas, which require high levels of fine-grained semantic understanding. Based on existing civil aviation datasets, a cross-modal retrieval approach with fine-grained semantic augmentation is suggested as a solution to this problem, creating a whole pipeline that includes data processing, model development, and training. First, text descriptions are optimized and enhanced based on a large multimodal model to construct a cross-modal retrieval dataset containing rich semantic information. Second, a model is created using a popular cross-modal retrieval framework. To improve the model's ability to express fine-grained semantic features, techniques such class supervision, key semantic information masking, and a fine-grained feature extraction module are introduced. Experimental results on two datasets verify the effectiveness of the proposed method, providing a reference technical path for cross-modal retrieval in low-altitude economy scenarios.

     

  • loading
  • [1]
    郭辰阳, 敖万忠, 吕宜宏. 充分把握发展机遇, 加快推进低空经济高质量发展[J]. 财经界, 2022(25): 36-38.

    Guo C Y, Ao W Z, Lyu Y H. Fully grasp the development opportunities and accelerate the high-quality development of low-altitude economy[J]. Money China, 2022(25): 36-38(in Chinese).
    [2]
    王淑鹤. 人工智能时代低空经济高质量发展的挑战和应对策略[J]. 产业创新研究, 2025(14): 39-41.

    Wang S H. Challenges and countermeasures of high-quality development of low-altitude economy in the era of artificial intelligence[J]. Industrial Innovation, 2025(14): 39-41(in Chinese).
    [3]
    张若愚, 聂婕, 宋宁, 等. 基于布局化-语义联合表征遥感图文检索方法[J]. 北京航空航天大学学报, 2024, 50(2): 671-683.

    Zhang R Y, Nie J, Song N, et al. Remote sensing image-text retrieval based on layout semantic joint representation[J]. Journal of Beijing University of Aeronautics and Astronautics, 2024, 50(2): 671-683(in Chinese).
    [4]
    王丹, 张峰, 张辉, 等. 基于信息互补与交叉注意力的跨模态检索方法[J]. 计算机应用研究, 2025, 42(7): 2032-2038.

    Wang D, Zhang F, Zhang H, et al. Information complementarity and cross-attention for cross-modal retrieval[J]. Application Research of Computers, 2025, 42(7): 2032-2038(in Chinese).
    [5]
    Vinyals O, Toshev A, Bengio S, et al. Show and tell: lessons learned from the 2015 MSCOCO image captioning challenge[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(4): 652-663.
    [6]
    Plummer B A, Wang L W, Cervantes C M, et al. Flickr30k entities: collecting region-to-phrase correspondences for richer image-to-sentence models[J]. International Journal of Computer Vision, 2017, 123(1): 74-93.
    [7]
    Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C]//Proceedings of the International Conference on Machine Learning. Cambridge: JMLR, 2021: 8748-8763.
    [8]
    Li J, Selvaraju R, Gotmare A, et al. Align before fuse: vi-sion and language representation learning with momentum distillation[J]. Advances in Neural Information Processing Systems, 2021, 34: 9694-9705.
    [9]
    Li Y H, Fan H Q, Hu R H, et al. Scaling language-image pre-training via masking[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2023: 23390-23400.
    [10]
    刘光才, 黄利萍, 李章萍, 等. 通用航空视角下我国低空经济的发展研判及欧美发展启示[J]. 中国物价, 2025(6): 42-48.

    Liu G C, Huang L P, Li Z P, et al. Analysis and prospects of China’s low-altitude economy development from the general aviation perspective and enlightenment from European and American development[J]. China Price Journal, 2025(6): 42-48(in Chinese).
    [11]
    Li D G, Dimitrova N, Li M K, et al. Multimedia content processing through cross-modal association[C]//Proceedings of the Eleventh ACM International Conference on Multimedia. New York: ACM, 2003: 604-611.
    [12]
    Zhang H, Liu Y, Ma Z G. Fusing inherent and external knowledge with nonlinear learning for cross-media retrieval[J]. Neurocomputing, 2013, 119: 10-16.
    [13]
    Frome A, Corrado G S, Shlens J, et al. Devise: a deep visual-semantic embedding model[C]//Proceedings of the 27th International Conference on Neural Information Processing Systems. New York: ACM, 2013: 2121-2129.
    [14]
    Mikolov T, Chen K, Corrado G, et al. Efficient estimation of word representations in vector space[EB/OL]. (2023-01-16)[2025-07-20]. https://arxiv.org/abs/1301.3781.
    [15]
    Lee K H, Chen X, Hua G, et al. Stacked cross attention for image-text matching[C]//Proceedings of the Computer Vision-ECCV. Berlin: Springer, 2018: 212-228.
    [16]
    Wehrmann J, Kolling C, C Barros R. Adaptive cross-modal embeddings for image-text alignment[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34(7): 12313-12320.
    [17]
    Qu L G, Liu M, Cao D, et al. Context-aware multi-view summarization network for image-text matching[C]//Proceedings of the 28th ACM International Conference on Multimedia. New York: ACM, 2020: 1047-1055.
    [18]
    Chen Y C, Li L J, Yu L C, et al. UNITER: universal image-text representation learning[C]//Proceedings of the Computer Vision-ECCV. Berlin: Springer, 2020: 104-120.
    [19]
    Li X J, Yin X, Li C Y, et al. Oscar: object-semantics aligned pre-training for vision-language tasks[C]//Proceedings of the Computer Vision-ECCV. Berlin: Springer, 2020: 121-137.
    [20]
    Li J N, Li D X, Xiong C M, et al. BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation[C]//Proceedings of the International Conference on Machine Learning. Cambridge: JMLR, 2022: 12888-12900.
    [21]
    Yu J H, Wang Z R, Vasudevan V, et al. CoCa: contrastive captioners are image-text foundation models[EB/OL]. (2022-06-14)[2025-07-30]. https://arxiv.org/abs/2205.01917.
    [22]
    Xiao R, Kim S, Georgescu M I, et al. FLAIR: VLM with fine-grained language-informed image representations[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2025: 24884-24894.
    [23]
    Tschannen M, Gritsenko A, Wang X, et al. SigLIP 2: multilingual vision-language encoders with improved semantic understanding, localization, and dense features[EB/OL]. (2025-02-20) [2025-08-01]. https://arxiv.org/abs/2502.14786.
    [24]
    Csizmadia D, Codreanu A, Sim V, et al. Distill CLIP (DCLIP): enhancing image-text retrieval via cross-modal transformer distillation[EB/OL]. (2025-06-16) [2025-08-01]. https://arxiv.org/abs/2505.21549.
    [25]
    Zhan G Q, Liu Y P, Han K, et al. ELIP: enhanced visual-language foundation models for image retrieval[EB/OL]. (2025-05-27) [2025-08-01]. https://arxiv.org/abs/2502.15682.
    [26]
    Beaumont R, Cherti M, Coombes T, et al. LAION-5B: an open large-scale dataset for training next generation image-text models[C]//Proceedings of the Advances in Neural Information Processing Systems 35. New York: ACM, 2022: 25278-25294.
    [27]
    中国民航大学. 世界民航事故调查跟踪系列报告: 2010-2021[R]. 天津: 中国民航大学, 2010-2021.

    Civil Aviation University of China. World civil aviation accident investigation tracking report series: 2010-2021[R]. Tianjin: Civil Aviation University of China, 2010-2021(in Chinese).
    [28]
    Maji S, Rahtu E, Kannala J, et al. Fine-grained visual classification of aircraft[EB/OL]. (2013-06-21) [2025-08-01]. https://arxiv.org/abs/1306.5151.
    [29]
    Openai, Achiam J, Adler S, et al. GPT-4 technical report[EB/OL]. (2024-03-04)[2025-08-01]. https://arxiv.org/abs/2303.08774.
    [30]
    Fu S H, Yang Q Z, Mo Q J, et al. LLMDet: learning strong open-vocabulary object detectors under the supervision of large language models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2025: 14987-14997.
    [31]
    Zheng L M, Chiang W L, Sheng Y, et al. Judging LLM-as-a-judge with MT-bench and chatbot arena[C]//Proceedings of the 37th International Conference on Neural Information Processing Systems. New York: ACM, 2023: 46595-46623.
    [32]
    Reimers N, Gurevych I. Sentence-BERT: sentence embeddings using Siamese BERT-networks[EB/OL]. (2019-08-27) [2025-08-01]. https://arxiv.org/abs/1908.10084.
    [33]
    He K M, Zhang X Y, Ren S Q, et al. Deep residual learning for image recognition[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2016: 770-778.
    [34]
    Hu J, Shen L, Sun G. Squeeze-and-excitation networks[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2018: 7132-7141.
    [35]
    Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: transformers for image recognition at scale[EB/OL]. (2021-06-03)[2025-08-01]. https://arxiv.org/abs/2010.11929.
    [36]
    Devlin J, Chang M W, Lee K, et al. BERT: pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Kerrville: Association for Computational Linguistics 2019: 4171-4186.
    [37]
    张振兴, 王亚雄. 图文跨模态检索研究综述[J]. 北京交通大学学报, 2024, 48(2): 23-36.

    Zhang Z X, Wang Y X. A survey on image-text cross-modal retrieval[J]. Journal of Beijing Jiaotong University, 2024, 48(2): 23-36(in Chinese).
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(8)  / Tables(6)

    Article Metrics

    Article views(264) PDF downloads(29) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return