Unsupervised Cross-Modal Hashing Algorithms for Web Multimedia Retrieval
DOI:
https://doi.org/10.13052/jwe1540-9589.2563Keywords:
Unsupervised learning, cross-modal hashing, web multimedia retrieval, semantic transfer, deep neural networksAbstract
With the Web witnessing a rapid surge in multimodal data, it’s becoming increasingly vital to develop efficient and budget-friendly cross-modal (CM) retrieval techniques to elevate the user experience in Web applications. Traditional hashing methods, however, often neglect the valuable semantic information hidden within the text descriptions that come with Web images. Moreover, they tend to lean heavily on supervised learning, which poses a challenge when it comes to adapting to real-world Web scenarios where annotations are often in short supply. To tackle this, this research introduces an unsupervised CM hashing algorithm for Web multimedia retrieval. By mining the semantic structure of text associated with Web images and utilizing a deep network to achieve semantic transfer from text to vision, a unified and efficient hashing learning framework is constructed. Experiments indicate that the introduced approach achieves mAP values of 0.3370 and 0.6990 with 16-bit hash codes (HCs). When the HC length is increased to 128 bits, the mAP increases to 0.3632 and 0.7575, representing an absolute improvement of 13.0% and 4.04% compared to the best performing baseline method. Further analysis shows that the semantic transfer mechanism significantly improves the semantic representation ability of the HCs. Even in a semi-supervised setting using only 20% of labeled data, the retrieval mAP can still reach 0.8871. The method requires only a portion of image-text pairs during the training phase and supports pure image queries during the retrieval phase, achieving millisecond-level response times under Hamming distance calculation. This provides an efficient and practical solution for Web-scale multimedia retrieval.
Downloads
References
Zhu Y, Wu Y, Sebe N, Yan Y. Vision+ x: A survey on multimodal learning in the light of data. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(12): 9102–9122. DOI:10.1109/TPAMI.2024.3420239.
Han Z, Azman A B, Mustaffa M R B, Khalid, F B. Cross-modal retrieval: A review of methodologies, datasets, and future perspectives. IEEE Access, 2024, 12(1): 115716–115741. DOI:10.1109/ACCESS.2024.3444817.
Li T, Kong L, Yang X, Wang B, Xu J. Bridging modalities: A survey of cross-modal image-text retrieval. Chinese Journal of Information Fusion, 2024, 1(1): 79–92. DOI:10.62762/CJIF.2024.361895.
Ma X, Yang M, Li Y, Hu P, Lv J, Peng X. Cross-modal retrieval with noisy correspondence via consistency refining and mining. IEEE Transactions on Image Processing, 2024, 33(1): 2587–2598. DOI:10.1109/tip.2024.3374221.
Wang Z, Xu X, Wei J, Xie N, Yang Y, Shen H T. Semantics disentangling for cross-modal retrieval. IEEE Transactions on Image Processing, 2024, 33(1): 2226–2237. DOI:10.1109/tip.2024.3374111.
Hu Z, Cheung Y M, Li M, Lan W. Cross-modal hashing method with properties of hamming space: A new perspective. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(12): 7636–7650. DOI:10.1109/tpami.2024.3392763.
Zhu L, Zheng C, Guan W, Li J, Yang Y, Shen H T. Multi-modal hashing for efficient multimedia retrieval: A survey. IEEE Transactions on Knowledge and Data Engineering, 2023, 36(1): 239–260. DOI:10.1109/tkde.2023.3282921.
Liang M, Du J, Liang Z, Xing Y, Huang W, Xue Z. Self-supervised multi-modal knowledge graph contrastive hashing for cross-modal search. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38(12): 13744–13753. DOI:10.1609/aaai.v38i12.29280.
Sun Y, Wang M, Ma Y. Semantic-alignment Transformer and adversary hashing for cross-modal retrieval. Applied Intelligence, 2024, 54(17): 7581–7602. DOI:10.1007/s10489-024-05501-2.
Han K, Liu Y, Wei R, Zhou K, Xu J, Long K. Supervised hierarchical online hashing for cross-modal retrieval. ACM Transactions on Multimedia Computing, Communications and Applications, 2024, 20(4): 1–23. DOI:10.1145/3632527.
Li F, Wang B, Zhu L, Li J, Zhang Z, Chang X. Cross-domain transfer hashing for efficient cross-modal retrieval. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(10): 9664–9677. DOI:10.1109/tcsvt.2024.3374791.
Liu X, Li J, Nie X, Zhang X, Wang S, Yin Y. Scalable unsupervised hashing via exploiting robust cross-modal consistency. IEEE Transactions on Big Data, 2024, 10(4): 514–527. DOI:10.1109/tbdata.2024.3350541.
Song G, Huang K, Su H, Song F, Yang M. Deep ranking distribution preserving hashing for robust multi-label cross-modal retrieval. IEEE Transactions on Multimedia, 2024, 26(1): 7027–7042. DOI:10.1109/tmm.2024.3358995.
Wang J, Zeng Z, Chen B, Wang Y, Liao D, Li G, Xia S T. Hugs bring double benefits: Unsupervised cross-modal hashing with multi-granularity aligned Transformers. International Journal of Computer Vision, 2024, 132(8): 2765–2797. DOI:10.1007/s11263-024-02009-7.
Li M, Ge M. Enhanced-similarity attention fusion for unsupervised cross-modal hashing retrieval. Data Science and Engineering, 2025, 10(2): 258–276. DOI:10.1007/s41019-024-00274-7.
Sun Y, Dai J, Ren Z, Chen Y, Peng D, Hu P. Dual self-paced cross-modal hashing. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38(14): 15184–15192. DOI:10.1609/aaai.v38i14.29441.
Zuo R, Zheng C, Li F, Zhu L, Zhang Z. Privacy-enhanced prototype-based federated cross-modal hashing for cross-modal retrieval. ACM Transactions on Multimedia Computing, Communications and Applications, 2024, 20(9): 1–19. DOI:10.1145/3674507.
Tu J, Liu X, Hao Y, Hong R. A unified generative hashing for cross-modal retrieval. ACM Transactions on Multimedia Computing, Communications and Applications, 2025, 21(12): 1–15. DOI:10.1145/3744567.
Adilakshmi K, Srinivas M, Kodali A, Srilakshmi V. Optimized RMDL with transfer learning for sentiment classification in the MapReduce framework. Journal of Web Engineering, 2023, 22(8): 1101–1132. DOI:10.13052/jwe1540-9589.2282.
Medhat S, Abdel-Galil H, Aboutabl A E, Saleh, H. Iterative magnitude pruning-based light-version of AlexNet for skin cancer classification. Neural Computing and Applications, 2024, 36(3): 1413–1428. DOI:10.1007/s00521-023-09111-w.
Li X. Design of a Web content personalized recommendation system based on collaborative filtering improved by combining k-means and LightGBM. Journal of Web Engineering, 2025, 24(2): 267–290. DOI:10.13052/jwe1540-9589.2425.
Hu P, Zhu H, Lin J, Peng D, Zhao Y P, Peng X. Unsupervised contrastive cross-modal hashing. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(3): 3877–3889. DOI:10.1109/tpami.2022.3177356.
Zhu L, Wu X, Li J, Zhang Z, Guan W, Shen H T. Work together: Correlation-identity reconstruction hashing for unsupervised cross-modal retrieval. IEEE Transactions on Knowledge & Data Engineering, 2023, 35(09): 8838–8851. DOI:10.1109/tkde.2022.3218656.
Wu Q, Zhang Z, Liu Y, Zhang J, Nie L. Contrastive multi-bit collaborative learning for deep cross-modal hashing. IEEE Transactions on Knowledge and Data Engineering, 2024, 36(11): 5835–5848. DOI:10.1109/TKDE.2024.3419577.

