A Robust Audio DNA-Based Method for OTT Content Recognition in Noisy Web Streaming Environments

Authors

  • Byeongchan Park Department of Computer Science & Engineering, Soongsil University, Korea https://orcid.org/0000-0002-8060-0561
  • Sun-Jib Kim Department of Convergence Security, Hansei University, Korea
  • Seok-Yoon Kim Department of Computer Science & Engineering, Soongsil University, Korea https://orcid.org/0000-0001-8029-515X
  • Youngmo Kim Department of Computer Science & Engineering, Soongsil University, Korea https://orcid.org/0000-0003-1415-3908
  • Yoontaek Sung Media & Advertising Research Institute, Korea Broadcasting Advertising Corporation (KOBACO), Korea

DOI:

https://doi.org/10.13052/jwe1540-9589.2562

Keywords:

Web Streaming, Audio Fingerprinting, OTT Content Recognition, Noise Robustness, Computational Efficiency

Abstract

With the proliferation of Over-The-Top (OTT) platforms via web browsers and mobile web apps, the demand for real-time copyright protection in web streaming environments has surged. However, unstructured noise generated when users consume web content in public environments (e.g., subways, cafes) increases the false positive rate of existing clean-audio-based identification systems and heavily burdens web server computations. This paper proposes a robust, audio DNA-based content recognition method capable of fast and accurate retrieval from large-scale media databases even in noisy web streaming environments. The proposed method extracts dual-stage (Coarse-Fine) features based on Mel-spectrograms at the client side and performs a highly efficient three-stage matching pipeline (Coarse Matching, Fine Matching, and Post-verification) at the web server side. Experimental results on 649 noisy audio samples demonstrated that applying web-optimized parameters (FFT length of 4096, Hop length of 1470) achieved a precision of 0.9965, a recall of 0.8814, and an F1-score of 0.9354. The proposed method significantly accelerates retrieval speed for large-scale web streaming data through binary hash matching in the Coarse stage while maintaining high accuracy, proving its effectiveness for real-time OTT copyright protection and web media monitoring systems.

Downloads

Download data is not yet available.

Author Biographies

Byeongchan Park, Department of Computer Science & Engineering, Soongsil University, Korea

Byeongchan Park received the bachelor’s degree in computer engineering through the Academic Credit Bank System in 2015, and the master’s and Doctor of Philosophy degrees in computer science and engineering from Soongsil University, Korea, in 2018 and 2023, respectively. He is currently a Visiting Professor at the Department of Computer Science, Soongsil University. His research areas include copyright technology and the promotion of its utilization.

Sun-Jib Kim, Department of Convergence Security, Hansei University, Korea

Sun-Jib Kim is currently a Professor in the School of IT, Hansei University, Korea. His research areas include information security, the Internet of Things (IoT), cloud computing, and AI system authentication.

Seok-Yoon Kim, Department of Computer Science & Engineering, Soongsil University, Korea

Seok-Yoon Kim received the B.S. degree in electrical and electronic engineering from Seoul National University, Korea, in 1980, and the M.S. and Ph.D. degrees in electrical and computer engineering from the University of Texas at Austin, USA, in 1990 and 1993, respectively. From 1982 to 1987, he was a Researcher with the Electronics and Telecommunications Research Institute (ETRI). From 1993 to 1995, he worked as a Senior Researcher at Motorola. Since 1995, he has been a Professor at Soongsil University. His main research interests include copyright protection and the promotion of its utilization.

Youngmo Kim, Department of Computer Science & Engineering, Soongsil University, Korea

Youngmo Kim received the B.S., M.S., and Ph.D. degrees in computer engineering from Daejeon University, Korea, in 2003, 2005, and 2011, respectively. Since 2012, he has been a Professor at Soongsil University. His main research interests include copyright protection and the promotion of its utilization.

Yoontaek Sung, Media & Advertising Research Institute, Korea Broadcasting Advertising Corporation (KOBACO), Korea

Yoontaek Sung received his Ph.D. in Communication from Sungkyunkwan University, Korea. He is currently a Principal Research Fellow at the Media Advertising Research Institute of the Korea Broadcasting Advertising Corporation (KOBACO). His research interests include the application of technology and data-driven approaches to advancing the media and advertising industries, on which he continues to conduct related projects and studies.

References

Korea Information Society Development Institute, Digital Content Industry Trends Report, KISDI, 2021.

A. Wang, “An Industrial-Strength Audio Search Algorithm,” in Proceedings of the 4th International Conference on Music Information Retrieval, pp. 7–13, 2003.

J.-S. Seo, M. Jin, S. Lee, D. Jang, S. Lee, and C. D. Yoo, “Audio Fingerprinting Based on Normalized Spectral Subband Moments,” IEEE Signal Processing Letters, vol. 13, no. 4, pp. 209–212, 2006, doi:10.1109/LSP.2005.863678.

V. Chandrasekhar, M. Sharifi, and D. A. Ross, “Survey and Evaluation of Audio Fingerprinting Schemes for Mobile Query-by-Example Applications,” in Proceedings of the 12th International Society for Music Information Retrieval Conference, pp. 801–806, 2011.

Q. Xiao, M. Suzuki, and K. Kita, “Fast Hamming Space Search for Audio Fingerprinting Systems,” in Proceedings of the 12th International Society for Music Information Retrieval Conference, pp. 133–138, 2011.

A. Baez-Suarez, N. Shah, J. A. Nolazco-Flores, S.-H. S. Huang, O. Gnawali, and W. Shi, “SAMAF: Sequence-to-Sequence Autoencoder Model for Audio Fingerprinting,” ACM Transactions on Multimedia Computing, Communications, and Applications, vol. 16, no. 2, pp. 1–23, 2020, doi:10.1145/3380828.

S. Chang, D. Lee, J. Park, H. Lim, K. Lee, K. Ko, and Y. Han, “Neural Audio Fingerprint for High-Specific Audio Retrieval Based on Contrastive Learning,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, pp. 3025–3029, 2021, doi:10.1109/ICASSP39728.2021.9414337.

J. Six, “Olaf: A Lightweight, Portable Audio Search System,” Journal of Open Source Software, vol. 8, no. 87, article 5459, 2023, doi:10.21105/joss.05459.

A. Marafioti, N. Holighaus, and P. Majdak, “Time-Frequency Phase Retrieval for Audio – The Effect of Transform Parameters,” IEEE Transactions on Signal Processing, vol. 69, pp. 3585–3596, 2021, doi:10.1109/TSP.2021.3088581.

Y. Zhou, X. Li, C. Xiong, H. Yao, and C. Qin, “A Survey of Perceptual Hashing for Multimedia,” ACM Transactions on Multimedia Computing, Communications, and Applications, 2025, doi:10.1145/3727880.

Downloads

Published

2026-08-22

How to Cite

Park, B. ., Kim, S.-J. ., Kim, S.-Y. ., Kim, Y. ., & Sung, Y. . (2026). A Robust Audio DNA-Based Method for OTT Content Recognition in Noisy Web Streaming Environments. Journal of Web Engineering, 25(06), 1067–1084. https://doi.org/10.13052/jwe1540-9589.2562

Issue

Section

ECTI