Unsupervised Cross-Modal Hashing Algorithms for Web Multimedia Retrieval

Authors

  • Yang-hao Li China Southern Power Grid Artificial Intelligence Technology Co. Ltd., Guangzhou 510700, Guangdong, China
  • Zhao-jie Dong China Southern Power Grid Artificial Intelligence Technology Co. Ltd., Guangzhou 510700, Guangdong, China
  • Shi-song Wu China Southern Power Grid Artificial Intelligence Technology Co. Ltd., Guangzhou 510700, Guangdong, China
  • Xuan-ang Li China Southern Power Grid Artificial Intelligence Technology Co. Ltd., Guangzhou 510700, Guangdong, China
  • Lian-yu Sha China Southern Power Grid Artificial Intelligence Technology Co. Ltd., Guangzhou 510700, Guangdong, China

DOI:

https://doi.org/10.13052/jwe1540-9589.2563

Keywords:

Unsupervised learning, cross-modal hashing, web multimedia retrieval, semantic transfer, deep neural networks

Abstract

With the Web witnessing a rapid surge in multimodal data, it’s becoming increasingly vital to develop efficient and budget-friendly cross-modal (CM) retrieval techniques to elevate the user experience in Web applications. Traditional hashing methods, however, often neglect the valuable semantic information hidden within the text descriptions that come with Web images. Moreover, they tend to lean heavily on supervised learning, which poses a challenge when it comes to adapting to real-world Web scenarios where annotations are often in short supply. To tackle this, this research introduces an unsupervised CM hashing algorithm for Web multimedia retrieval. By mining the semantic structure of text associated with Web images and utilizing a deep network to achieve semantic transfer from text to vision, a unified and efficient hashing learning framework is constructed. Experiments indicate that the introduced approach achieves mAP values of 0.3370 and 0.6990 with 16-bit hash codes (HCs). When the HC length is increased to 128 bits, the mAP increases to 0.3632 and 0.7575, representing an absolute improvement of 13.0% and 4.04% compared to the best performing baseline method. Further analysis shows that the semantic transfer mechanism significantly improves the semantic representation ability of the HCs. Even in a semi-supervised setting using only 20% of labeled data, the retrieval mAP can still reach 0.8871. The method requires only a portion of image-text pairs during the training phase and supports pure image queries during the retrieval phase, achieving millisecond-level response times under Hamming distance calculation. This provides an efficient and practical solution for Web-scale multimedia retrieval.

Downloads

Download data is not yet available.

Author Biographies

Yang-hao Li, China Southern Power Grid Artificial Intelligence Technology Co. Ltd., Guangzhou 510700, Guangdong, China

Yang-hao Li (June 1996), male, holds a master’s degree in Business Administration from Guangdong University of Technology, China. Currently working as an Assistant Engineer at China Southern Power Grid Artificial Intelligence Technology Co. Ltd. His research and work focus on large models for the power industry, power informatization and power marketing.

Zhao-jie Dong, China Southern Power Grid Artificial Intelligence Technology Co. Ltd., Guangzhou 510700, Guangdong, China

Zhao-jie Dong (September 1985), male, holds a master’s degree in Computer Technology from Sun Yat-sen University, China. Currently, he works as a Senior Engineer at China Southern Power Grid Artificial Intelligence Technology Co. Ltd. His work and research focus on artificial intelligence, large language models and intelligent customer service.

Shi-song Wu, China Southern Power Grid Artificial Intelligence Technology Co. Ltd., Guangzhou 510700, Guangdong, China

Shi-song Wu (April 1986), male, holds a master’s degree in Software Engineering from Tsinghua University, China. Currently, he works as a Senior Engineer at China Southern Power Grid Artificial Intelligence Technology Co. Ltd. His research focuses on artificial intelligence, large language models and intelligent customer service.

Xuan-ang Li, China Southern Power Grid Artificial Intelligence Technology Co. Ltd., Guangzhou 510700, Guangdong, China

Xuan-ang Li (June 1993), male, holds a master’s degree in Computer Technology from Guangxi University, China. Currently, he works as an Engineer at China Southern Power Grid Artificial Intelligence Technology Co. Ltd. His research focuses on artificial intelligence and power informatization.

Lian-yu Sha, China Southern Power Grid Artificial Intelligence Technology Co. Ltd., Guangzhou 510700, Guangdong, China

Lian-yu Sha (November 1990), male, obtained his bachelor’s degree in Computer Technology from Harbin Normal University, China. Currently, he works as an Engineer at China Southern Power Grid Artificial Intelligence Technology Co. Ltd. His research and work focus on artificial intelligence, data analysis and power marketing.

References

Zhu Y, Wu Y, Sebe N, Yan Y. Vision+ x: A survey on multimodal learning in the light of data. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(12): 9102–9122. DOI:10.1109/TPAMI.2024.3420239.

Han Z, Azman A B, Mustaffa M R B, Khalid, F B. Cross-modal retrieval: A review of methodologies, datasets, and future perspectives. IEEE Access, 2024, 12(1): 115716–115741. DOI:10.1109/ACCESS.2024.3444817.

Li T, Kong L, Yang X, Wang B, Xu J. Bridging modalities: A survey of cross-modal image-text retrieval. Chinese Journal of Information Fusion, 2024, 1(1): 79–92. DOI:10.62762/CJIF.2024.361895.

Ma X, Yang M, Li Y, Hu P, Lv J, Peng X. Cross-modal retrieval with noisy correspondence via consistency refining and mining. IEEE Transactions on Image Processing, 2024, 33(1): 2587–2598. DOI:10.1109/tip.2024.3374221.

Wang Z, Xu X, Wei J, Xie N, Yang Y, Shen H T. Semantics disentangling for cross-modal retrieval. IEEE Transactions on Image Processing, 2024, 33(1): 2226–2237. DOI:10.1109/tip.2024.3374111.

Hu Z, Cheung Y M, Li M, Lan W. Cross-modal hashing method with properties of hamming space: A new perspective. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(12): 7636–7650. DOI:10.1109/tpami.2024.3392763.

Zhu L, Zheng C, Guan W, Li J, Yang Y, Shen H T. Multi-modal hashing for efficient multimedia retrieval: A survey. IEEE Transactions on Knowledge and Data Engineering, 2023, 36(1): 239–260. DOI:10.1109/tkde.2023.3282921.

Liang M, Du J, Liang Z, Xing Y, Huang W, Xue Z. Self-supervised multi-modal knowledge graph contrastive hashing for cross-modal search. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38(12): 13744–13753. DOI:10.1609/aaai.v38i12.29280.

Sun Y, Wang M, Ma Y. Semantic-alignment Transformer and adversary hashing for cross-modal retrieval. Applied Intelligence, 2024, 54(17): 7581–7602. DOI:10.1007/s10489-024-05501-2.

Han K, Liu Y, Wei R, Zhou K, Xu J, Long K. Supervised hierarchical online hashing for cross-modal retrieval. ACM Transactions on Multimedia Computing, Communications and Applications, 2024, 20(4): 1–23. DOI:10.1145/3632527.

Li F, Wang B, Zhu L, Li J, Zhang Z, Chang X. Cross-domain transfer hashing for efficient cross-modal retrieval. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(10): 9664–9677. DOI:10.1109/tcsvt.2024.3374791.

Liu X, Li J, Nie X, Zhang X, Wang S, Yin Y. Scalable unsupervised hashing via exploiting robust cross-modal consistency. IEEE Transactions on Big Data, 2024, 10(4): 514–527. DOI:10.1109/tbdata.2024.3350541.

Song G, Huang K, Su H, Song F, Yang M. Deep ranking distribution preserving hashing for robust multi-label cross-modal retrieval. IEEE Transactions on Multimedia, 2024, 26(1): 7027–7042. DOI:10.1109/tmm.2024.3358995.

Wang J, Zeng Z, Chen B, Wang Y, Liao D, Li G, Xia S T. Hugs bring double benefits: Unsupervised cross-modal hashing with multi-granularity aligned Transformers. International Journal of Computer Vision, 2024, 132(8): 2765–2797. DOI:10.1007/s11263-024-02009-7.

Li M, Ge M. Enhanced-similarity attention fusion for unsupervised cross-modal hashing retrieval. Data Science and Engineering, 2025, 10(2): 258–276. DOI:10.1007/s41019-024-00274-7.

Sun Y, Dai J, Ren Z, Chen Y, Peng D, Hu P. Dual self-paced cross-modal hashing. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38(14): 15184–15192. DOI:10.1609/aaai.v38i14.29441.

Zuo R, Zheng C, Li F, Zhu L, Zhang Z. Privacy-enhanced prototype-based federated cross-modal hashing for cross-modal retrieval. ACM Transactions on Multimedia Computing, Communications and Applications, 2024, 20(9): 1–19. DOI:10.1145/3674507.

Tu J, Liu X, Hao Y, Hong R. A unified generative hashing for cross-modal retrieval. ACM Transactions on Multimedia Computing, Communications and Applications, 2025, 21(12): 1–15. DOI:10.1145/3744567.

Adilakshmi K, Srinivas M, Kodali A, Srilakshmi V. Optimized RMDL with transfer learning for sentiment classification in the MapReduce framework. Journal of Web Engineering, 2023, 22(8): 1101–1132. DOI:10.13052/jwe1540-9589.2282.

Medhat S, Abdel-Galil H, Aboutabl A E, Saleh, H. Iterative magnitude pruning-based light-version of AlexNet for skin cancer classification. Neural Computing and Applications, 2024, 36(3): 1413–1428. DOI:10.1007/s00521-023-09111-w.

Li X. Design of a Web content personalized recommendation system based on collaborative filtering improved by combining k-means and LightGBM. Journal of Web Engineering, 2025, 24(2): 267–290. DOI:10.13052/jwe1540-9589.2425.

Hu P, Zhu H, Lin J, Peng D, Zhao Y P, Peng X. Unsupervised contrastive cross-modal hashing. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(3): 3877–3889. DOI:10.1109/tpami.2022.3177356.

Zhu L, Wu X, Li J, Zhang Z, Guan W, Shen H T. Work together: Correlation-identity reconstruction hashing for unsupervised cross-modal retrieval. IEEE Transactions on Knowledge & Data Engineering, 2023, 35(09): 8838–8851. DOI:10.1109/tkde.2022.3218656.

Wu Q, Zhang Z, Liu Y, Zhang J, Nie L. Contrastive multi-bit collaborative learning for deep cross-modal hashing. IEEE Transactions on Knowledge and Data Engineering, 2024, 36(11): 5835–5848. DOI:10.1109/TKDE.2024.3419577.

Downloads

Published

2026-08-22

How to Cite

Li, Y.- hao ., Dong, Z.- jie ., Wu, S.- song ., Li, X.- ang ., & Sha, . L.- yu . (2026). Unsupervised Cross-Modal Hashing Algorithms for Web Multimedia Retrieval. Journal of Web Engineering, 25(06), 1085–1110. https://doi.org/10.13052/jwe1540-9589.2563

Issue

Section

Advanced Practice in Web Engineering in Asia