A Web-Based Streaming Video Identification Method Using Self-Supervised Structural Embedding and Vector Similarity Search
DOI:
https://doi.org/10.13052/jwe1540-9589.2567Keywords:
Streaming Video Identification, Self-Supervised Learning, Structural Embedding, Vector Similarity Search, Vector DatabaseAbstract
Streaming video services have become a major channel for content distribution owing to the rapid growth of over-the-top (OTT) platforms and web-based media services. As a result, accurate video identification is increasingly required for copyright protection, audience measurement, content management, and search and recommendation services. However, conventional feature point-based methods often show limited robustness against transformations frequently observed in streaming environments, including compression, resolution change, re-encoding, frame-rate reduction, aspect-ratio conversion, overlays, rotation, and flipping. They also require the storage and comparison of many local descriptors, which can limit retrieval efficiency in large-scale content databases. In this paper, we propose a web-based streaming video identification method using self-supervised structural embedding and vector similarity search. The proposed method extracts representative frames from original and query videos, applies preprocessing, generates 512-dimensional structural embeddings using a self-supervised copy detection model, and performs cosine similarity search in a vector database. To improve robustness against geometric transformations, transformed reference frames are additionally registered. For query videos, multiple representative frames are used, and the final video-level identification result is determined by majority voting over frame-level retrieval results. Each frame is represented by a fixed-length 512-dimensional vector, requiring 2048 bytes per frame and 10,240 bytes for a five-frame query. Experiments on 4000 query videos covering 14 transformation categories and 40 detailed transformation settings show that the proposed method achieves a recognition rate of 99.23%, a missed recognition rate of 0.38%, and a false recognition rate of 0.40%. These results demonstrate the feasibility of the proposed method for identifying transformed copies of the same source video in web-based streaming environments.
Downloads
References
I. Yoo, Y. Kim, B. Park, S. Jang, and S. Kim, “A web-based identification method for illegal streaming videos using low-frequency components of the fast Fourier transform,” Journal of Korea Multimedia Society, 2024.
Sandvine, Global Internet Phenomena Report 2024. Sandvine Inc., 2024. [Online]. Available: https://www.sandvine.com/global-internet-phenomena-report.
Nielsen, “Time spent streaming surges to 40.3% of total TV usage in June 2024,” The Gauge, 2024.
Nielsen, Streaming Measurement. Nielsen, 2024.
Nielsen, “What’s the difference between OTT, CTV and streaming?” Nielsen Insights, 2024.
A. González-Neira and C. Quintas-Froufe, “Television audience measurement: The challenge posed by the rise of streaming and OTT platforms,” Comunicación y Sociedad, 2020.
G. Anselmi et al., “A first look at automatic content recognition tracking in smart TVs,” arXiv preprint arXiv:2409.06203, 2024.
D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International Journal of Computer Vision, vol. 60, no. 2, pp. 91–110, 2004.
H. Bay, T. Tuytelaars, and L. Van Gool, “SURF: Speeded up robust features,” Computer Vision and Image Understanding, vol. 110, no. 3, pp. 346–359, 2008.
E. Rublee, V. Rabaud, K. Konolige, and G. Bradski, “ORB: An efficient alternative to SIFT or SURF,” in Proc. IEEE Int. Conf. Computer Vision (ICCV), 2011.
P. F. Alcantarilla, J. Nuevo, and A. Bartoli, “Fast explicit diffusion for accelerated features in nonlinear scale spaces,” in Proc. British Machine Vision Conf. (BMVC), 2013.
M. Douze, M. Furon, and H. Jégou, “Self-supervised copy detection for large-scale image retrieval,” in Proc. IEEE/CVF Int. Conf. Computer Vision (ICCV), 2021.
J. Wang et al., “Milvus: A purpose-built vector data management system,” in Proc. ACM SIGMOD Int. Conf. Management of Data, 2021.
Y. Aumüller, E. Bernhardsson, and A. Faithfull, “ANN-benchmarks: A benchmarking tool for approximate nearest neighbor algorithms,” Information Systems, vol. 87, 2020.

