Immersive Web Virtual Environment for Language Learning: Multimodal Interaction and Error Correction Based on Speech, Gesture and Eye Tracking

Authors

  • Shuwen Yu Information Technology Center, Wenzhou University, Wenzhou 325035, China

DOI:

https://doi.org/10.13052/jwe1540-9589.2565

Keywords:

Immersive Web Virtual Environment, Language Learning, Multimodal Interaction, Speech Error Correction, Eye Tracking, WebXR

Abstract

In order to solve the problems of existing Web language learning tools, such as single interaction dimension, lagging error correction feedback and insufficient attention perception, this paper designs and verifies a multimodal immersive Web virtual environment optimization scheme which integrates voice, gesture and eye tracking. The core breakthroughs are as follows: (1) a dual-mechanism adaptive fusion algorithm of “attention weight+dynamic threshold” is proposed, which solves the problem of error correction failure of the traditional fixed threshold scheme in high-fluctuation interactive scenarios (such as voice discontinuity, gesture occlusion, attention drift); (2) a four-stage closed-loop control mechanism of perception-decision-feedback-adjustment” is constructed, which realizes the end-side lightweight deployment. Based on WebXR, MediaPipe and other native technologies, the cross-end compatible architecture is constructed, and the comparative experiments on three data sets and 10 typical working conditions show that, compared with the traditional Web platform, the collaborative error correction accuracy of the scheme is 93. 2% (improved by 41%), and the response time is compressed to 0.15 seconds (shortened by 62%). End-side resource occupancy is reduced to 16.8%, and learning efficiency is 4.8 times per minute. The research confirms that the scheme effectively balances the immersion, error correction accuracy and end-to-end adaptability, and provides technical support for the large-scale application of immersive Web language education.

Downloads

Download data is not yet available.

Author Biography

Shuwen Yu, Information Technology Center, Wenzhou University, Wenzhou 325035, China

Shuwen Yu was born in Lanzhou, Gansu, P.R. China, in 1977. He received the bachelor‘s degree from Northwest Normal University. He works at the Information Technology Center of Wenzhou University. His research interests include artificial intelligence in education, algorithm fusion and teaching big data analysis.

References

Slocum C, Jingwen Huang, Jiasi Chen. VIA: Visibility-aware Web-based Virtual Reality. Proceedings of the 26th International Conference on 3D Web Technology (Web3D ’21). Association for Computing Machinery, New York, NY, USA, Article 7, 1–9, 2021. DOI:10.1145/3485444.3487641.

Choi J M, Park K H. Performance Analysis of GLTF/GLB to Improve 3D Content Rendering Performance. Journal of Platform Technology 11(4):13–18, 2023.

Petropoulos J, Kamarianakis M, Protopsaltis A, et al. pyGANDALF – An open-source Geometric, ANimation, Directed, Algorithmic, Learning Framework for Computer Graphics. SIGGRAPH Asia 2024 Educator’s Forum 1–9, 2024. DOI:10.1145/3680533.3697057.

Ma Xiangchun, Xu Da, Zhong Yongjiang. Design of Situational English Learning System for Primary Schools Based on WebXR and AI Agents. China Information Technology Education 18:86–89, 2024.

Chojnowski O, Kirschbaum A, Neef C, Jeschke S, Richert A. Exploring the Role of Co-Speech Gestures: An In-the-Wild Study with a Virtual Agent in a Museum. Proceedings of the 13th International Conference on Human-Agent Interaction, HAI, November 10–13, 2025, Yokohama, Kanagawa-Pref., Japan.

Yaseen, Kwon O J, Kim J, et al. Next-Gen Dynamic Hand Gesture Recognition: MediaPipe, Inception-v3 and LSTM-Based Enhanced Deep Learning Model. Electronics, 13(16):3233, 2024. DOI:10.3390/electronics13163233.

Guo H, Fan W, Wei B, et al. AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding. IEEE Transactions on Circuits and Systems for Video Technology, 35(10):10238–10249, 2024. DOI:10.1109/TCSVT.2025.3569731.

Chen X, Zhang J, Zhao Y, Chen Q, Chen B, Xu N, Jin E, Shen Y, Tian Y, Shen M, Gao Z. PATE Model: A 30-Year Review and Analysis of Gestural Interaction Research. Human Factors 68:42–77, 2025.

Wu Lan, Yang Pan, Li Binquan, Wang Han. Multimodal audio-visual speech recognition in noisy environment with large vocabulary. Guangxi Science 30(1): 52–60, 2023.

Khaustova V, Pyshkin E, Khaustov V, Blake J, Bogach N. CAPTuring Accents: An Approach to Personalize Pronunciation Training for Learners with Different L1 Backgrounds. In SPECOM 2023. Lecture Notes in Computer Science 14339, 2023. Springer, Cham. DOI:10.1007/978-3-031-48312-7_5.

Li Wentao, Wang Xijun, Chen Li. Design of multi-modal cooperative reasoning system for edge computation. China Mobile 49(03):72–77, 2025.

Zhao Q, Sun G, Zhang C, Xu M, Zheng T F. Enhancing Quantised End-to-End ASR Models Via Personalisation. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Republic of Korea, 12426–12430, 2024. DOI:10.1109/ICASSP48485.2024.10448012.

Yang X, Krajbich I. Webcam-based online eye-tracking for behavioral research. Judgment and Decision Making 16(6):1485–1505, 2021. DOI:10.1017/S1930297500008512.

Hangbin Zheng, Tianyuan Liu, Jiayu Liu, Jinsong Bao. Visual analytics for digital twins: a conceptual framework and case study. Journal of Intelligent Manufacturing 35(4):1671–1686, 2024.

Du P, Guo W. Cheng S. Using eye-tracking for real-time translation: a new approach to improving reading experience. CCF Trans. Pervasive Comp. Interact. 6:150–164, 2024. DOI:10.1007/s42486-024-00150-3.

Fan Z, Li J, Wumaier A, Kadeer Z, Abdurahman A. A Multifaceted Approach to Oral Assessment Based on the Conformer Architecture. IEEE Access 11:28318–28329, 2023. DOI:10.1109/ACCESS.2023.3255986.

Zhang Q. Imamiya A. Go K. Mao X. Resolving ambiguities of a gaze and speech interface. Proceedings of the Eye Tracking Research & Application Symposium, ETRA, San Antonio, Texas, USA, March 22–24, 85–95, 2004.

Vidal J, Wigham C R. Multimodal strategies allowing corrective feedback to be softened during Web conferencing-supported interactions. Paper presented at the Conference on Telecollaboration in Higher Education, Dublin, Ireland, April 21–23, 139–146, 2016. DOI:10.14705/rpnet.2016.telecollab2016.500.

Lei W, Ge Y, Yi K, et al. ViT-Lens: Towards Omni-modal Representations. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 26637–26647, 2024. DOI:10.1109/CVPR52733.2024.02516.

Downloads

Published

2026-08-22

How to Cite

Yu, S. . (2026). Immersive Web Virtual Environment for Language Learning: Multimodal Interaction and Error Correction Based on Speech, Gesture and Eye Tracking. Journal of Web Engineering, 25(06), 1133–1160. https://doi.org/10.13052/jwe1540-9589.2565

Issue

Section

Advanced Practice in Web Engineering in Asia