A Robust Audio DNA-Based Method for OTT Content Recognition in Noisy Web Streaming Environments
DOI:
https://doi.org/10.13052/jwe1540-9589.2562Keywords:
Web Streaming, Audio Fingerprinting, OTT Content Recognition, Noise Robustness, Computational EfficiencyAbstract
With the proliferation of Over-The-Top (OTT) platforms via web browsers and mobile web apps, the demand for real-time copyright protection in web streaming environments has surged. However, unstructured noise generated when users consume web content in public environments (e.g., subways, cafes) increases the false positive rate of existing clean-audio-based identification systems and heavily burdens web server computations. This paper proposes a robust, audio DNA-based content recognition method capable of fast and accurate retrieval from large-scale media databases even in noisy web streaming environments. The proposed method extracts dual-stage (Coarse-Fine) features based on Mel-spectrograms at the client side and performs a highly efficient three-stage matching pipeline (Coarse Matching, Fine Matching, and Post-verification) at the web server side. Experimental results on 649 noisy audio samples demonstrated that applying web-optimized parameters (FFT length of 4096, Hop length of 1470) achieved a precision of 0.9965, a recall of 0.8814, and an F1-score of 0.9354. The proposed method significantly accelerates retrieval speed for large-scale web streaming data through binary hash matching in the Coarse stage while maintaining high accuracy, proving its effectiveness for real-time OTT copyright protection and web media monitoring systems.
Downloads
References
Korea Information Society Development Institute, Digital Content Industry Trends Report, KISDI, 2021.
A. Wang, “An Industrial-Strength Audio Search Algorithm,” in Proceedings of the 4th International Conference on Music Information Retrieval, pp. 7–13, 2003.
J.-S. Seo, M. Jin, S. Lee, D. Jang, S. Lee, and C. D. Yoo, “Audio Fingerprinting Based on Normalized Spectral Subband Moments,” IEEE Signal Processing Letters, vol. 13, no. 4, pp. 209–212, 2006, doi:10.1109/LSP.2005.863678.
V. Chandrasekhar, M. Sharifi, and D. A. Ross, “Survey and Evaluation of Audio Fingerprinting Schemes for Mobile Query-by-Example Applications,” in Proceedings of the 12th International Society for Music Information Retrieval Conference, pp. 801–806, 2011.
Q. Xiao, M. Suzuki, and K. Kita, “Fast Hamming Space Search for Audio Fingerprinting Systems,” in Proceedings of the 12th International Society for Music Information Retrieval Conference, pp. 133–138, 2011.
A. Baez-Suarez, N. Shah, J. A. Nolazco-Flores, S.-H. S. Huang, O. Gnawali, and W. Shi, “SAMAF: Sequence-to-Sequence Autoencoder Model for Audio Fingerprinting,” ACM Transactions on Multimedia Computing, Communications, and Applications, vol. 16, no. 2, pp. 1–23, 2020, doi:10.1145/3380828.
S. Chang, D. Lee, J. Park, H. Lim, K. Lee, K. Ko, and Y. Han, “Neural Audio Fingerprint for High-Specific Audio Retrieval Based on Contrastive Learning,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, pp. 3025–3029, 2021, doi:10.1109/ICASSP39728.2021.9414337.
J. Six, “Olaf: A Lightweight, Portable Audio Search System,” Journal of Open Source Software, vol. 8, no. 87, article 5459, 2023, doi:10.21105/joss.05459.
A. Marafioti, N. Holighaus, and P. Majdak, “Time-Frequency Phase Retrieval for Audio – The Effect of Transform Parameters,” IEEE Transactions on Signal Processing, vol. 69, pp. 3585–3596, 2021, doi:10.1109/TSP.2021.3088581.
Y. Zhou, X. Li, C. Xiong, H. Yao, and C. Qin, “A Survey of Perceptual Hashing for Multimedia,” ACM Transactions on Multimedia Computing, Communications, and Applications, 2025, doi:10.1145/3727880.

