Design of a Multi-Tenant Real-Time Inference Framework Based on OpenStack and SR-IOV GPU Virtualization

Authors

  • Ma Rui Earthquake Agency of Xinjiang Uygur Autonomous Region, Urumqi 830011, Xinjiang, China
  • Shi Bingfeng Earthquake Agency of Xinjiang Uygur Autonomous Region, Urumqi 830011, Xinjiang, China
  • Ma Xin Earthquake Agency of Xinjiang Uygur Autonomous Region, Urumqi 830011, Xinjiang, China
  • Wei Bin Earthquake Agency of Xinjiang Uygur Autonomous Region, Urumqi 830011, Xinjiang, China

DOI:

https://doi.org/10.13052/jicts2245-800X.1434

Keywords:

Real-time inference, Multi-tenant cloud computing, GPU virtualization, SR-IOV, OpenStack, Tail-latency, Hardware-assisted isolation, ICT standards

Abstract

The deployment of real-time artificial intelligence inference services on shared cloud infrastructure poses significant challenges due to resource contention, latency variability, and tail-latency amplification. While cloud platforms offer scalability and flexibility, conventional accelerator sharing mechanisms often fail to provide the determinism required by latency-sensitive inference workloads. This paper presents a standards-based, multi-tenant cloud inference framework that integrates OpenStack orchestration with Single Root I/O Virtualization (SR-IOV)-enabled graphics processing unit (GPU) partitioning to achieve predictable and isolated real-time inference execution. In the proposed architecture, each tenant is assigned an exclusive GPU virtual function, enabling hardware-level isolation while remaining fully compatible with native OpenStack scheduling and resource management mechanisms. A comprehensive experimental evaluation is conducted on a private OpenStack cloud to assess inference latency distribution, tail behavior, scalability, robustness to background network and control-plane activity, and throughput-latency trade-offs. Experimental results show that median inference latency remains stable across single-tenant and multi-tenant configurations, while P95 and P99 tail latencies exhibit no measurable amplification under concurrent execution. The system scales linearly with the number of available GPU virtual functions, maintaining consistent latency behavior until hardware capacity is reached. Additional experiments demonstrate that background network traffic and control-plane operations introduce negligible impact on inference latency. Throughput analysis reveals a well-defined saturation knee, enabling clear identification of safe operating regions for real-time inference services. By leveraging mature ICT standards and open-source cloud infrastructure, this work provides a reusable reference architecture for deploying latency-sensitive inference services in private and hybrid clouds. The results highlight the effectiveness of hardware-assisted accelerator isolation in balancing performance determinism, scalability, and operational simplicity, and offer practical guidance for future system design and standardization efforts.

Downloads

Download data is not yet available.

Author Biographies

Ma Rui, Earthquake Agency of Xinjiang Uygur Autonomous Region, Urumqi 830011, Xinjiang, China

Ma Rui received his B.S. in Communication Engineering from Beijing University of Posts and Telecommunications, China, in July 2008, and his M.S. in Geological Resources and Geological Engineering from Xinjiang University in July 2015. His research interests include the construction and maintenance of information network systems. He has authored several academic papers and was awarded the title of Engineer in September 2018. He currently works at the Monitoring and Information Center of the Xinjiang Earthquake Agency.

Shi Bingfeng, Earthquake Agency of Xinjiang Uygur Autonomous Region, Urumqi 830011, Xinjiang, China

Shi Bingfeng, assistant engineer, graduated with a bachelor’s degree in Mechanical Design, Manufacturing and Automation from Henan University of Technology, China, in July 2021. Currently working at the Monitoring and Information Center of Xinjiang Seismological Bureau, he specializes in the construction and operation & maintenance of information network systems, undertaking tasks such as virtualization platform management, server operation & maintenance, and hardware troubleshooting. With comprehensive technical capabilities, he ensures the stable operation of seismological monitoring-related systems.

Ma Xin, Earthquake Agency of Xinjiang Uygur Autonomous Region, Urumqi 830011, Xinjiang, China

Ma Xin received his B.Eng. in Electrical Engineering and Automation from Bohai University, China, in 2019. He is currently an Assistant Engineer at the Monitoring and Information Center of the Xinjiang Earthquake Agency. His research focuses on the operation and maintenance of seismic monitoring networks and earthquake early warning systems. He has published several academic papers in these fields.

Wei Bin, Earthquake Agency of Xinjiang Uygur Autonomous Region, Urumqi 830011, Xinjiang, China

Wei Bin received his B.S. in theoretical physics from Kashi Normal University, China, in July 1991. His research focuses on the construction of seismic monitoring systems, data quality evaluation, and the development of intelligent early warning technologies. He has published several academic papers and was promoted to Senior Engineer (full professor level) in January 2024. He is currently Director of the Monitoring and Information Center at the Xinjiang Earthquake Agency, dedicated to enhancing the seismic monitoring capabilities in Xinjiang.

References

Aslani, A., & Ghobaei-Arani, M. (2025). Machine learning inference serving models in serverless computing: A survey. Computing, 107(1), 47–78. https://doi.org/10.1007/s00607-024-01243-1

Dean, J., & Barroso, L. A. (2013). The tail at scale. Communications of the ACM, 56(2), 74–80.

Nigade, V. V. (2023). Latency-critical inference serving for deep learning. ACM Queue, 21(4), 44–67. https://doi.org/10.1145/3617306

Zhou, Z., Ning, X., Hong, K., Fu, T., Xu, J., Li, S., et al. (2024). A survey on efficient inference for large language models. arXiv preprint arXiv:2404.14294

Hong, C. H., Spence, I., & Nikolopoulos, D. S. (2017). GPU virtualization and scheduling methods: A comprehensive survey. ACM Computing Surveys (CSUR), 50(3), 1–37.

Prekas, G., Primorac, M., Belay, A., Kozyrakis, C., & Bugnion, E. (2015, August). Energy proportionality and workload consolidation for latency-critical applications. In Proceedings of the Sixth ACM symposium on cloud computing (pp. 342–355).

Crankshaw, D., et al. (2017). Clipper: A low-latency online prediction serving system. USENIX NSDI, 613–627.

Olston, C., Fiedel, N., Gorovoy, K., Harmsen, J., Lao, L., Li, F., Rajashekhar, V., Ramesh, S., & Soyke, J. (2017). TensorFlow-Serving: Flexible, high-performance ML serving. In NIPS 2017 Workshop on Machine Learning Systems. arXiv:1712.06139. https://doi.org/10.48550/arXiv.1712.06139

Rzadca, K., Findeisen, P., Swiderski, J., Zych, P., Broniek, P., Kusmierek, J., … & Wilkes, J. (2020, April). Autopilot: workload autoscaling at Google. In Proceedings of the Fifteenth European Conference on Computer Systems (pp. 1–16).

Luu, H., Pumperla, M., & Zhang, Z. (2024). Model serving infrastructure. In MLOps with Ray. Berkeley, CA: Apress. https://doi.org/10.1007/979-8-8688-0376-5

Moritz, P., et al. (2018). Ray: A distributed framework for emerging AI applications. USENIX OSDI, 561–577.

Lin, X., Yang, C., Wang, W., Li, Y., Du, C., Feng, F., … & Chua, T. S. (2024). Efficient inference for large language model-based generative recommendation. In The Thirteenth International Conference on Learning Representations (ICLR 2025). arXiv preprint arXiv:2410.05165. https://doi.org/10.48550/arXiv.2410.05165

Ratul, I. J., Zhou, Y., & Yang, K. (2025). Accelerating deep learning inference: A comparative analysis of modern acceleration frameworks. Electronics, 14(15), 2977.

Shi, L., et al. (2011). vCUDA: GPU-accelerated HPC in virtual machines. IEEE Transactions on Computers, 61(6), 804–816.

NVIDIA Corporation. (2022). NVIDIA virtual GPU architecture. White paper.

Fischer, M., & Wiedner, F. (2021). Survey on SR-IOV performance. Network, 43, 45–52.

Avasalcai, C., Tsigkanos, C., & Dustdar, S. (2021). Resource management for latency-sensitive IoT applications with satisfiability. IEEE Transactions on Services Computing, 15(5), 2982–2993.

PCI-SIG. (2021). Single Root I/O Virtualization (SR-IOV) specification.

Dong, Y., Yang, X., Li, J., Liao, G., Tian, K., & Guan, H. (2012). High performance network virtualization with SR-IOV. Journal of Parallel and Distributed Computing, 72(11), 1471–1480.

de Oliveira Filho, A. T., Freitas, E., do Carmo, P. R., Souto, E., Kelner, J., & Sadok, D. F. (2024). Analysis of SR-IOV in Docker containers using RTT measurements. Computer Communications, 228, 107961.

Zhang, Z., Wang, L., Liu, R., & Fan, J. (2022). Development of cloud computing platform based on neural network. Mathematical Problems in Engineering, 2022(1), 1513081.

Chen, X., Ying, R., Ma, H., Wang, Y., Meng, X., Xie, G., … & Wu, F. (2024, May). An SR-IOV SSD Optimized for QoS-Sensitive IaaS Cloud Storage. In 2024 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW) (pp. 1161–1163).

Maene, P., Götzfried, J., De Clercq, R., Müller, T., Freiling, F., & Verbauwhede, I. (2017). Hardware-based trusted computing architectures for isolation and attestation. IEEE Transactions on Computers, 67(3), 361–374.

OpenStack Foundation. (2018). Placement API specification.

Pepple, K. (2011). Deploying OpenStack. O’Reilly Media.

Scharf, M., et al. (2015). Network-aware instance scheduling in OpenStack. ICCCN, 1–6.

Lai, W. K., Wang, Y. C., & Wei, S. C. (2023). Delay-aware container scheduling in Kubernetes. IEEE Internet of Things Journal, 10(13), 11813–11824.

Sun, Q., Yi, L., Yang, H., Li, M., Luan, Z., & Qian, D. (2022). QoS-aware dynamic resource allocation with improved utilization and energy efficiency on GPU. Parallel Computing, 113, 102958.

Zhou, M., Mu, X., & Liang, Y. (2025). SOE: A multi-objective traffic scheduling engine for DDoS mitigation with isolation-aware optimization. Mathematics, 13(11), 1853.

Yu, F., Wang, D., Shangguan, L., Zhang, M., Liu, C., & Chen, X. (2022). A survey of multi-tenant deep learning inference on GPU. In MLSys 2022 Workshop on Cloud Intelligence/AIOps. arXiv:2203.09040. https://doi.org/10.48550/arXiv.2203.09040

Rosa, L., Foschini, L., & Corradi, A. (2024). Empowering cloud computing with network acceleration: A survey. IEEE Communications Surveys & Tutorials, 26(4), 2729–2768.

Downloads

Published

2026-08-09

How to Cite

Rui, M. ., Bingfeng, S. ., Xin, M. ., & Bin, W. . (2026). Design of a Multi-Tenant Real-Time Inference Framework Based on OpenStack and SR-IOV GPU Virtualization. Journal of ICT Standardization, 14(03), 357–390. https://doi.org/10.13052/jicts2245-800X.1434

Issue

Section

Intelligent System Concepts, architecture, standards, tools and applications