VLDB 2026 Research / reviewers in the wild / expert
Boyu Gao 0003
dblp:161/3491
· DBLP profile ↗
44ranked-venue papers
8as first author
37since 2021 · last 2026
0000-0001-8523-2828ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 17 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning from Scoring Disagreements: Contrastive Error Mining for Efficient and Robust LLM-based AssessmentabstractAutomated grading of student responses still faces numerous challenges, particularly when dealing with complex and ambiguous answers. In particular, large models are prone to scoring bias when handling uncertain responses, and few-shot reasoning methods often lack stability, which limits their applicability in real educational scenarios. To tackle these challenges, we propose the Contrastive Error Mining and FineTuning (CEM-FT) framework, which automatically identifies high-value hard samples by analyzing scoring disagreements between a full fine-tuned model and a few-shot model. A lightweight LoRA adapter is then trained on these samples to refine model performance with minimal computational overhead. Experiments on the SciEntsbank, Beetle, and Mohler datasets show that CEM-FT can improve QWK by up to 3.9% compared to the fine-tuned Qwen model on SciEntsbank datasets, which is a significant improvement over the few-shot baseline. The proposed framework substantially enhances both scoring accuracy and consistency, providing a practical, robust solution for reliable automated assessment with large language models. Lei Chen 0088, Tengteng Cheng 0002, Boyu Gao 0003, Zitao Liu 0001, Weiqi Luo 0002 |
AAAI | 3 |
| 2026 | Exploring the Impact of Sensory Conflict on State Mindfulness in VR Meditation Training
Shuting Chang, Tianren Luo, Pengxiang Wang 0006, Xiehaoxuan Tang, Gaozhang Chen, Boyu Gao 0003, Qi Wang 0192, Teng Han, Yachun Fan, Feng Tian 0001 |
CHI | 7 |
| 2026 | Two-Handed Click and Tap: Expanding Input Vocabulary of Controllers for Virtual Reality InteractionabstractThis study explores the design space of two-handed input (i.e., clicking or tapping with the thumb) on the touchpads of controllers for virtual reality (VR) interaction. Four experiments were conducted to fulfill this purpose. Experiment 1 investigated how users employed two VR controllers to perform four representative interaction tasks in VR and identified 14 potentially usable two-handed operations that involved tapping or clicking. Experiments 2 and 3 analyzed user performance of the 14 operations, providing insights into their interaction characteristics in terms of completion time, accuracy, and subjective feedback. In Experiment 4, we designed a command-input technique based on the proposed operations. We verified its effectiveness compared to context menus and marking menus in a VR text entry scenario. Our technique generally had shorter times and similar accuracy to the two menu types. Our work contributes to the design of VR interactions using two-handed controllers. Huawei Tu, Boyu Gao 0003, Yujun Lu, Weiqiang Xin, Hui Cui 0002, Weiqi Luo 0002, Jian Weng 0001, Henry Been-Lirn Duh |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2026 | SemanticAction: A Semantic-Driven and Behavior-Aware Password Framework for Adaptive VR Authentication Under Observation AttacksabstractAs immersive Virtual Reality (VR) applications become increasingly widespread, ensuring secure and usable authentication is critical. Traditional knowledge-based methods (e.g., passwords, PINs) suffer from memorability issues and are highly vulnerable to observation attacks such as Man-in-the-Room (MITR). Meanwhile, biometric and behavioral approaches raise concerns regarding practicality, privacy, and cross-platform deployment. We present SemanticAction, a semantic-driven, behavior-augmented authentication framework that integrates knowledge-based passwords with gesture-based behavioral biometrics. Passwords are encoded as scene-anchored directions executed through intuitive hand gestures, enabling semantic meaning to guide user interactions. To counter observation attacks, SemanticAction employs randomized scene prompts and decoy scenes, while a dual-constraint verification mechanism adapts to both cold-start/few-shot conditions. Two user studies provide initial evidence that SemanticAction can mitigate MITR attacks while alleviating memorability challenges, maintaining favorable usability and security performance even with limited behavioral data. This work offers early insights and practical design considerations for behavior-aware authentication in VR. Tingjie Wan, Yalin Deng, Zixuan Guo 0003, Xubo Yang, Boyu Gao 0003 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2026 | Understanding the Effect of Latency on User Performance of Target Selection in Virtual RealityabstractHigh latency is often introduced due to limited computational capabilities and high hardware demands. It has proven to significantly impair user performance in target selection, a fundamental interaction task. Existing research has established that latency negatively impacts selection times and success rates in 2D interactive systems; however, the underlying behavioral mechanisms remain unclear. This article investigates the effects of latency on selection times, success rates, and endpoint distributions in Virtual Reality (VR) with controller-based raycasting and bare-hand direct touch-the two most common selection methods. Our results from a user study (N = 31) revealed distinct patterns between the two methods, leading to two novel mathematical models that account for latency, target width, and movement amplitude. These two models were validated via a new dataset collected from a second user study (N = 16) and were demonstrated to outperform the existing models. Our findings provide actionable recommendations to mitigate the negative impacts of latency and improve user experience in VR interface design. Yushi Wei, Rongkai Shi, Kemu Xu, Jialin Wang 0002, Boyu Gao 0003, Pan Hui 0001, Lingyun Yu 0001, Hai-Ning Liang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | StarPicker: A Technique for Selecting Dense Small Targets in AR-Based Data Visualization EnvironmentsabstractIn the era of Big Data, augmented reality (AR)-based 3D visualization technology is gradually becoming an essential tool of effective information dissemination. However, 3D data visualization inevitably faces a series of challenges, particularly when visualizing large datasets within limited physical spaces. This often results in high spatial density, characterized by smaller target objects and occlusion of data objects, which greatly complicate object selection and subsequent interaction. To address these issues, this paper proposes a multimodal progressive target selection technique-StarPicker, which integrates the SpotLight metaphor and Clock metaphor. By leveraging wrist rotation and gesture interactions, StarPicker facilitates a coarse-to-fine semantic disambiguation and precise selection of target objects. To validate the effectiveness of StarPicker, a comparative user study was conducted against the state-of-the-art techniques (GridWall and FlowerCone) under high-density (up to 360 objects) and small-target (1cm) conditions. Experimental results demonstrate that StarPicker significantly outperforms the baselines in terms of target selection accuracy and user satisfaction, while achieving comparable or better completion times, especially in denser scenes. This work offers a novel approach to target selection and interaction technology in the field of AR-based 3D Big Data visualization. Huiyue Wu, Xinle Wang, Zhitong Ma, Boyu Gao 0003, Huawei Tu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | FootEyePorting: Design and Evaluation of Foot-Eye Teleportation Techniques in Virtual RealityabstractVarious locomotion techniques, such as teleportation, walking in place, redirected walking, and walking, have been proposed. However, conventional methods overlooked the human ability to coordinate multiple modalities (e.g., eyes and feet) for natural virtual locomotion within a small physical space. Inspired by the natural coordination of the eyes and feet in human walking, as well as insights from prior work, this study investigates how eye-foot coordination can be leveraged more effectively for VR teleportation. We present four novel teleportation techniques based on eye-foot coordination, using users' gaze behavior and 3D foot positions as input modalities. A user study with 20 participants compared our techniques with a state-of-the-art foot-based locomotion method. Results demonstrate the superiority of our approaches: task completion times were significantly reduced, NASA-TLX workload scores and SUS usability scores were markedly improved, and participants expressed a clear preference for our techniques over the baseline. These findings provide a strong foundation for the design and implementation of future eye-foot coordinated teleportation methods in VR. Tingjie Wan, Boyu Gao 0003, Huawei Tu, Henry Been-Lirn Duh |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | Effects of Postures on Identifying Users for Selection-Based Behavioral Authentication in Virtual RealityabstractBehavioral authentication has become increasingly popular as a natural method for authentication in Virtual Reality (VR). However, existing studies often overlook the fact that users may perform behavioral authentication in different postures (i.e., sitting, standing, reclining) during VR use. Therefore, understanding how posture variations affect classification accuracy is crucial for designing posture-robust systems. In this study, we conducted a controlled experiment (N = 30) to investigate the impact of posture on classification accuracy during a target-selection task. We collected behavioral trajectory data and analyzed it using multivariate time series classification algorithms, addressing authentication performance under three different postures. In a within-posture authentication, reclining took longer but achieved the highest classification accuracy, with an interaction effect between posture and target vertical layout. In cross-posture authentication, transfers from sitting to standing/reclining were more effective than direct transfers between standing and reclining, with vertical layout crucial for classification accuracy. In mixed-posture training, the cross-posture classification accuracy increased, particularly when standing and reclining data were combined to help the model indirectly learn features of sitting posture. These findings provide valuable insights for designing tasks and data collection strategies that support the development of robust cross-posture authentication systems. GuanYu Ye, Tingjie Wan, Huawei Tu, Jian Weng 0001, Boyu Gao 0003 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Modeling Locomotion with Body Angular Movements in Virtual Reality
Zijun Mai, Boyu Gao 0003, Huawei Tu, Dasheng Li, HyungSeok Kim 0001, Weiqi Luo 0002 |
CHI | 2 |
| 2025 | Exploring Plausible Preference of Body-Centric Locomotion with Reinforcement Learning in Virtual RealityabstractInvestigating users' plausible preferences for body-centric locomotion can help researchers better understand the impact of different factors on this locomotion method and optimize the locomotion configuration for providing a plausible locomotion experience. In this paper, we propose to evaluate users' plausible preferences for body-centric locomotion using a reinforcement learning method. This method can intelligently infer users' plausible preferences for different factor levels involved in the virtual locomotion by proposing possible modifications to the factor levels and asking users to accept or reject the modifications after experiencing the locomotion. We conducted a within-subject experiment to examine the impact of different factors (i.e., body parts used for virtual locomotion, the point of view, auditory feedback, the transfer function, and the coefficients of the transfer function) on users' plausible preferences for body-centric locomotion in sitting and standing postures. The results mainly indicated that (1) users preferred using arm swinging in standing posture, whereas they preferred using head tilting in sitting posture; (2) The point of view was identified as the most important factor in standing posture, while it was less important in sitting posture; (3) Participants showed consistent plausible preferences for auditory feedback, transfer function, and coefficient of the transfer function in both standing and sitting postures. Our research findings can guide the design of VR applications to enhance a plausible walking experience in different postures. Zijun Mai, Boyu Gao 0003, Huawei Tu, Haojun Zheng, Henry Been-Lirn Duh |
ISMAR | 2 |
| 2025 | A Dual-Stick Controller for Enhancing Raycasting Interactions with Virtual ObjectsabstractThis work presents Dual-Stick, a novel controller with two sticks connected at the end that innovates a Dual-Ray interaction paradigm to enrich raycasting input in Virtual Reality (VR). Dual-Stick leverages the inherent human dexterity in using everyday tools such as clamps and tweezers to adjust the relative angle between two sticks. This design supports Dual-Ray interactions that provide with a heuristics-based enhanced mechanism. It also offers more flexible manipulation by taking advantages of additional degrees of freedom provided by clamping angle. We conducted two studies to evaluate the effectiveness of Dual-Ray in target selection and manipulation tasks. The results indicated that Dual-Ray significantly improved efficiency in target selection compared to single-ray input but did not outperform the enhanced single-ray technique. In terms of manipulation, Dual-Ray effectively reduced completion time and mode switching compared to single-ray input. Nianlong Li, Zhenxuan He, Luyao Shen, Tianren Luo, Teng Han, Boyu Gao 0003, Yu Zhang 0199, Liuxin Zhang, Feng Tian 0001, Qianying Wang 0002 |
VR | 7 |
| 2025 | Effects of Arm Posture and Speed Estimation Methods on the Performance and User Preference of Virtual Locomotion Using Arm SwingabstractArm swing is a viable approach for locomotion in virtual reality (VR). However, previous studies have not investigated the effects of arm posture (straight arms and bent arms) and speed estimation methods on speed estimation error, locomotion performance, and user preference. We proposed two data-driven approaches for virtual locomotion based on a one-dimensional convolutional neural network (1-D CNN) and the support vector regression (SVR), respectively, and compared their speed estimation errors using treadmill walking data for both bent arms and straight arms. We then compared the locomotion performance and user preference of the two methods against two existing virtual locomotion methods through virtual locomotion tasks. Experimental results suggest that it is necessary to consider the influence of arm posture and speed estimation methods when designing virtual locomotion methods based on arm swing. The results of this work provide insights into the research and the design of arm-swing-based locomotion interfaces. Mingjun Shao, Boyu Gao 0003 |
Int. J. Hum. Comput. Interact. | 4 |
| 2025 | AudioGest: Gesture-Based Interaction for Virtual Reality Using Audio DevicesabstractCurrent virtual reality (VR) system takes gesture interaction based on camera, handle and touch screen as one of the mainstream interaction methods, which can provide accurate gesture input for it. However, limited by application forms and the volume of devices, these methods cannot extend the interaction area to such surfaces as walls and tables. To address the above challenge, we propose AudioGest, a portable, plug-and-play system that detects the audio signal generated by finger tapping and sliding on the surface through a set of microphone devices without extensive calibration. First, an audio synthesis-recognition pipeline based on micro-contact dynamics simulation is constructed to generate modal audio synthesis from different materials and physical properties. Then the accuracy and effectiveness of the synthetic audio are verified by mixing the synthetic audio with real audio proportionally as the training sets. Finally, a series of desktop office applications are developed to demonstrate the application potential of AudioGest's scalability and versatility in VR scenarios. Yi Xiao 0009, Mingwei Hu, Hao Sha 0004, Shining Ma, Boyu Gao 0003, Shihui Guo, Yue Liu 0005 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | SummonBrush: Enhancing Touch Interaction on Large XR User Interfaces by Augmenting Users' Hands with Virtual BrushesabstractTouch interaction is one of the fundamental interaction paradigms in XR, as users have become very familiar with touch interactions on physical touchscreens. However, users typically need to perform extensive arm movements for engaging with XR user interfaces much larger than mobile device touchscreens. We propose the SummonBrush technique to facilitate easy access to hidden windows while interacting with large XR user interfaces, requiring minimal arm movements. The SummonBrush technique adds a virtual brush to the index fingertip of a user's hand. Upon making contact with a virtual user interface, the brush bends and diverges and ink starts to diffuse in it. The more the brush bends and diverges, the more the ink diffuses. The user can summon hidden windows or background applications in situ, which is achieved by firstly pressing the brush against the user interface to make ink fully fill the brush and then perform swipe gestures. Also, the user can press the brush against the thumbtails of background applications in situ to quickly cycle them through. Ecological studies showed that SummonBrush significantly reduced the arm movement time by 39% and 34% in summoning hidden windows and activating/closing background applications, respectively, leading to a significant decrease in reported physical demand. Yang Tian 0008, Zhao Su, Tianren Luo, Teng Han, Shengdong Zhao 0001, Boyu Gao 0003, Dangxiao Wang |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2025 | Casual-VRAuth: A Design Framework Bridging Focused and Casual Interactions for Behavioral Authentication in Virtual RealityabstractCurrent behavioral authentication systems in Virtual Reality (VR) require sustained focused interaction during task execution - an assumption frequently incompatible with real-world constraints across two factors: (1) physical limitations (e.g., restricted hand/eye mobility), and (2) psychological barriers (e.g., task-switching fatigue or break-in-presence). To address this attentional gap, we propose a design framework bridging focused and casual interactions in behavior-based VR authentication (Casual-VRAuth). Based on this framework, we designed an authentication prototype using a modified ball-and-tunnel task (propelling a ball along a circular path), supporting three interaction modes: baseline Touch, and two eyes-free options (Hover/Tapping). Experimental results demonstrate that our framework effectively guides the design of authentication systems with varying interaction engagement levels (Touch > Hover > Tapping) to accommodate scenarios requiring casual interaction (e.g., multitasking or eyes-free operation). Furthermore, we revealed that reducing interaction engagement enhances resistance to mimicry attacks while decreasing cognitive workload and error rates in multitasking or eyes-free environments. However, this approach compromises the average classification accuracy of Interaction behavior under different algorithms (including InceptionTime, FCN, ResNet, CNN, MLP, and MCDCNN). Notably, moderate reduction of interaction engagement enhances authentication speed, while excessive reduction may conversely slow it down. Overall, our work establishes a novel design paradigm for VR authentication that supports casual interactions and offers valuable insights into balancing usability and security. GuanYu Ye, Boyu Gao 0003, Huawei Tu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Knowledge Tracing as Language Processing: A Large-Scale Autoregressive Paradigm
Bojun Zhan, Teng Guo 0002, Xueyi Li 0005, Mingliang Hou, Qianru Liang, Boyu Gao 0003, Weiqi Luo 0002, Zitao Liu 0001 |
AIED (1) | 6 |
| 2024 | Extending Context Window of Attention Based Knowledge Tracing Models via Length ExtrapolationabstractKnowledge tracing (KT) is a prediction task that aims to predict students’ future performance based on their past learning data. The rapid progress in attention mechanisms has led to the emergence of various high-performing attention based KT models. However, in online or personalized education settings, students’ varying learning paths result in different lengths of student interaction sequences, which poses a significant challenge for attention based KT models as their context window sizes are fixed during both training and prediction stages. We refer to this as the length extrapolation of KT model. In this paper, we propose extraKT to facilitate better extrapolation that learn from student interactions with a short context window and continue to perform well across various longer context window sizes at prediction stage. Specifically, we negatively bias attention scores with linearly decreasing penalties that are proportional to query-key distance, which efficiently represents short-term forgetting characteristics of student knowledge states. We conduct comprehensive and rigorous experiments on three real-world educational datasets. The results show that our extraKT model exhibits robust length extrapolation capability and outperforms state-of-the-art baseline models in terms of AUC and accuracy. To encourage reproducible research, we merge our data and code to the publicly available pyKT benchmark at https://github.com/pykt-team/pykt-toolkit. Xueyi Li 0005, Youheng Bai, Teng Guo 0002, Ying Zheng 0010, Mingliang Hou, Bojun Zhan, Yaying Huang, Zitao Liu 0001, Boyu Gao 0003, Weiqi Luo 0002 |
ECAI | 9 |
| 2024 | SATPose: Improving Monocular 3D Pose Estimation with Spatial-aware Ground TactilityabstractEstimating 3D human poses from monocular images is an important research area with many practical applications. However, the depth ambiguity of 2D solutions limits their accuracy in actions where occlusion exits or where slight centroid shifts can result in significant 3D pose variations. In this paper, we introduce a novel multimodal approach to mitigate the depth ambiguity inherent in monocular solutions by integrating spatial-aware pressure information. We first establish a data collection system with a pressure mat and a monocular camera, and construct a large-scale multimodal human activity dataset comprising over 600,000 frames of motion data. Utilizing this dataset, we propose a pressure image reconstruction network to extract pressure priors from monocular images. Subsequently, we introduce a Transformer-based multimodal pose estimation network to combine pressure priors with monocular images, achieving a world mean per joint position error of 51.6mm, outperforming state-of-the-art methods. Extensive experiments demonstrate the effectiveness of our multimodal 3D human pose estimation method across various actions and joints, highlighting the significance of spatial-aware pressure in improving the accuracy of monocular-vision-based methods. Our dataset is available at: https://github.com/LishuangZhan/SATPose. Lishuang Zhan, Enting Ying, Jiabao Gan, Shihui Guo, Boyu Gao 0003, Yipeng Qin |
ACM Multimedia | 5 |
| 2024 | Evaluating Plausible Preference of Body-Centric Locomotion using Subjective Matching in Virtual RealityabstractBody-centric locomotion in Virtual Reality (VR) involves multiple factors, including the point of view, avatar representations, tracked body parts for locomotion control and transfer functions that map body movement to the displacement of the virtual viewpoint. Understanding the role of these factors in evoking a plausible walking experience using within- or between-subject experimental designs based on questionnaires and/or objective measurements can be time-consuming and challenging due to the interrelated effects of these factors. This study employed the subjective matching method to evaluate the sense of plausible walking experience during body-centric locomotion in VR. Five relevant factors that may affect locomotion experience were identified by analyzing existing studies, i.e., point of view, the avatar appearance, body parts for locomotion control, transfer functions and the coefficients of transfer functions. A virtual locomotion experiment with these five factors based on subjective matching was conducted. Results showed that participants regarded the point of view as the most critical factor for walking experience enhancement, followed by body parts, transfer functions, the coefficients of transfer functions and finally the avatar appearance. Additionally, participants’ preferences for different body parts and the coefficients of transfer functions affected the choice of transfer functions. These results could serve as the guidelines for virtual locomotion experience design that involves combinations of multiple factors and can help achieve a plausible walking experience in VR. Boyu Gao 0003, Haojun Zheng, Huawei Tu, HyungSeok Kim 0001, Henry Been-Lirn Duh |
VR | 1 |
| 2024 | Exploring Controller-based Techniques for Precise and Rapid Text Selection in Virtual RealityabstractText selection is a common task in interactive systems. Often, it can be difficult because the letters and words are too small and clustered together to allow precise selection. Compared to traditional 2D interfaces, text selection is more challenging in virtual reality (VR) head-mounted displays (HMDs) because users interact with the immersive 3D space via mid-air interaction, which has higher degrees of freedom but becomes more imprecise and involves a higher workload due to the lack of support from a fixed structure like a desk. There has been limited exploration of techniques that support precise and rapid text selection at the character, word, sentence, or paragraph levels in VR HMDs. To fill this gap, we propose three controller-based text selection methods: Joystick Movement, Depth Movement, and Wrist Orientation. They are evaluated against a baseline selection method via a user study with 24 participants. Results show that the three proposed techniques significantly improved the performance and user experience over the baseline, especially for the selection beyond the character level. Jianbin Song, Rongkai Shi, Yue Li 0023, Boyu Gao 0003, Hai-Ning Liang |
VR | 4 |
| 2024 | Climbing Keyboard: A Tilt-Based Selection Keyboard Entry for Virtual RealityabstractText input is one of the common interaction tasks in virtual environments. However, current inputting methods (e.g., laser-input: aim-and-shoot technique) have many limitations, such as inefficiency, lack of precision, and fatigue of long-text inputting. We propose a Climbing Keyboard method to allow easier, faster, and more accurate text input, and use tilt instead of precise aiming. The selected target changes from a specific letter to a group of letters with such a tilt-based interaction based on the QWERTY layout, which aims to reduce the learning cost, especially for novice users. Meanwhile, expert users can focus on the screen without looking at the keyboard. We designed three user studies to evaluate the performance of proposed method, including the verification of the usability of the tilt interaction method in the first study, optimization of the tilt angle range in the second study, and evaluation of the learning curve of Climbing keyboard in the last study. Our results showed that participants can reach 16.48 words per minute after an hours of training. Junfeng Huang, Minghui Sun 0001, Boyu Gao 0003, Guihe Qin |
Int. J. Hum. Comput. Interact. | 4 |
| 2024 | Secure and Memorable Authentication Using Dynamic Combinations of 3D Objects in Virtual RealityabstractAs Virtual Reality (VR) applications gain popularity, the need for a secure, usable, and memorable user authentication method becomes crucial. However, security and privacy in such VR applications are often ignored. Current methods are insufficient in preventing man-in-the-room (MITR) attacks, which allow attackers to observe user interactions in VR while remaining invisible, and inputted passwords can easily be stolen. In this study, we propose a dynamic combination of multi-attribute authentication methods for VR, where various 3D objects and their attributes can be created and displayed. Users must select combinations of 3D objects and their attributes provided by our designed principles for identity authentication. We explore the impact of method parameters on security and provide three specific parameter schemes to deploy the practical authentication system. We designed three user studies to evaluate the usability, security, and memorability of our authentication system. The results show that the proposed scheme can effectively resist both shoulder surfing and MITR attacks with unsuccessful attack rates of 100% and 95.83%, respectively. Furthermore, this research provides suggestions to secure VR applications while maintaining usability and enhancing the memorability of the authentication method. Boyu Gao 0003, Huawei Tu, Hai-Ning Liang, Zitao Liu 0001, Weiqi Luo 0002, Jian Weng 0001 |
Int. J. Hum. Comput. Interact. | 2 |
| 2024 | Exploring Bimanual Haptic Feedback for Spatial Search in Virtual RealityabstractSpatial search tasks are common and crucial in many Virtual Reality (VR) applications. Traditional methods to enhance the performance of spatial search often employ sensory cues such as visual, auditory, or haptic feedback. However, the design and use of bimanual haptic feedback with two VR controllers for spatial search in VR remains largely unexplored. In this work, we explored bimanual haptic feedback with various combinations of haptic properties, where four types of bimanual haptic feedback were designed, for spatial search tasks in VR. Two experiments were designed to evaluate the effectiveness of bimanual haptic feedback on spatial direction guidance and search in VR. The results from the first experiment reveal that our proposed bimanual haptic schemes significantly enhanced the recognition of spatial directions in terms of accuracy and speed compared to spatial audio feedback. The second experiment's findings suggest that the performance of bimanual haptic feedback was comparable to or even better than the visual arrow, especially in reducing the angle of head movement and enhancing searching targets behind the participants, which was supported by subjective feedback as well. Based on these findings, we have derived a set of design recommendations for spatial search using bimanual haptic feedback in VR. Boyu Gao 0003, Tong Shao, Huawei Tu, Qizi Ma, Zitao Liu 0001, Teng Han |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Analysis and Design of Efficient Authentication Techniques for Password Entry with the Qwerty Keyboard for VR EnvironmentsabstractAuthentication in digital security relies heavily on text-based passwords, even with other available methods like biometrics and graphical passwords. While virtual reality (VR) keyboards are typically invisible to onlookers, the presence of inconspicuous sensors, including accelerometers, gyroscopes, and barometers, poses a potential risk of unauthorized observation and recording. Traditional defense shoulder-surfing attack methods typically involve breaking apart the Qwerty layout, which destroys the user's inherent familiarity with the layout. This research addresses the need for secure password entry in VR environments while retaining the Qwerty layout. We explore three keyboard-related position alteration strategies to ensure security while mitigating the decline in user experience. These strategies involve moving the entire keyboard, cursor, and keys. Our theoretical study assesses the effectiveness of these strategies against shoulder-surfing attacks. Two user studies, employing ray-based and position-based text entry methods, respectively, evaluate the practical effectiveness of the three strategies in resisting shoulder-surfing attacks, as well as their impact on typing performance and user experience. Our findings demonstrate that the three strategies achieve shoulder-surfing attack resistance comparable to a random layout keyboard. Moreover, compared to a random layout, the two strategies involving the movement of the entire keyboard and the repositioning of keys support faster entry rates and enhanced user experience. Tingjie Wan, Liangyuting Zhang, Yunxin Xu, Zixuan Guo 0003, Boyu Gao 0003, Hai-Ning Liang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | An Interactive Framework of Balancing Evaluation Cost and Prediction Accuracy for Knowledge TracingabstractThe development of online intelligent educational systems has revolutionized personalized learning, presenting an opportunity for the integration of knowledge tracing (KT). KT is an essential task that leverages students’ historical interactions to model their knowledge states, enabling accurate predictions of their future performance. The application of deep learning models on KT tasks, also known as deep learning based knowledge tracing (DLKT) models, has accelerated the process of KT tasks in recent years. Although DLKT models have achieved promising results, there is a huge challenge to avoid models generating wrong estimations that yield bad guidance in educational contexts. Hence, in this work, we propose a simulation framework to explore the feasibility of minimizing the evaluation cost while guaranteeing prediction performance. In the framework, we initially design a simple yet efficient DLKT model that learns the actual process of students’ knowledge acquisition. We then select the reliable predictions generated by the proposed model and assign the unreliable ones to teaching professionals based on confidence estimations. We present the results of a proof-of-concept experiment on three real-world publicly available datasets to demonstrate that our framework can obtain the balance of human cost and automatic evaluation accuracy, which can be flexibly deployed to real-world educational contexts in the future. To encourage reproducible research, we make our code publicly available at https://github.com/gwbnwnwh/human-in-the-loop. Weili Sun, Boyu Gao 0003, Jiahao Chen 0006, Shuyan Huang, Weiqi Luo 0002 |
IEEE Big Data | 2 |
| 2023 | Assessing Student Performance with Multi-granularity Attention from Online Classroom DialogueabstractAccurately judging students' ongoing performance is very crucial for real-world educational scenarios. In this work, we focus on the task of automatically predicting students' levels of mastery of math questions from teacher-student classroom dialogue data in the online learning environment. We propose a novel neural network armed with a multi-granularity attention mechanism to capture the personalized pedagogical instructions from the very noisy teacher-student dialogue transcriptions. We conduct experiments on a real-world educational dataset and the results demonstrate the superiority and availability of our model in terms of various evaluation metrics. Jiahao Chen 0006, Zitao Liu 0001, Shuyan Huang, Yaying Huang, Xiangyu Zhao 0001, Boyu Gao 0003, Weiqi Luo 0002 |
CIKM | 6 |
| 2023 | Enhancing Deep Knowledge Tracing with Auxiliary TasksabstractKnowledge tracing (KT) is the problem of predicting students’ future performance based on their historical interactions with intelligent tutoring systems. Recent studies have applied multiple types of deep neural networks to solve the KT problem. However, there are two important factors in real-world educational data that are not well represented. First, most existing works augment input representations with the co-occurrence matrix of questions and knowledge components1 (KCs) but fail to explicitly integrate such intrinsic relations into the final response prediction task. Second, the individualized historical performance of students has not been well captured. In this paper, we proposed AT-DKT to improve the prediction performance of the original deep knowledge tracing model with two auxiliary learning tasks, i.e., question tagging (QT) prediction task and individualized prior knowledge (IK) prediction task. Specifically, the QT task helps learn better question representations by predicting whether questions contain specific KCs. The IK task captures students’ global historical performance by progressively predicting student-level prior knowledge that is hidden in students’ historical learning interactions. We conduct comprehensive experiments on three real-world educational datasets and compare the proposed approach to both deep sequential KT models and non-sequential models. Experimental results show that AT-DKT outperforms all sequential models with more than 0.9% improvements of AUC for all datasets, and is almost the second best compared to non-sequential models. Furthermore, we conduct both ablation studies and quantitative analysis to show the effectiveness of auxiliary tasks and the superior prediction outcomes of AT-DKT. To encourage reproducible research, we make our data and code publicly available at https://github.com/pykt-team/pykt-toolkit 2. Zitao Liu 0001, Qiongqiong Liu, Jiahao Chen 0006, Shuyan Huang, Boyu Gao 0003, Weiqi Luo 0002, Jian Weng 0001 |
WWW | 5 |
| 2023 | Text Pin: Improving text selection with mode-augmented handles on touchscreen mobile devices
Huawei Tu, Boyu Gao 0003, Huiyue Wu, Fei Lyu 0001 |
Int. J. Hum. Comput. Stud. | 2 |
| 2023 | Effects of Transfer Functions and Body Parts on Body-Centric Locomotion in Virtual RealityabstractBody-centric locomotion allows users to control both movement speed and direction with body parts (e.g., head tilt, arm swing or torso lean) to navigate in virtual reality (VR). However, there is little research to systematically investigate the effects of body parts for speed and direction control on virtual locomotion by taking in account different transfer functions(L: linear function, P: power function, and CL: piecewise function with constant and linear function). Therefore, we conducted an experiment to evaluate the combinational effects of the three factors (body parts for direction control, body parts for speed control, and transfer functions) on virtual locomotion. Results showed that (1) the head outperformed the torso for movement direction control in task completion time and environmental collisions; (2) Arm-based speed control led to shorter traveled distances than both head and knee. Head-based speed control had fewer environmental collisions than knee; (3) Body-centric locomotion with CL function was faster but less accurate than both L and P functions. Task time significantly decreased from P, L to CL functions, while traveled distance and overshoot significantly increased from P, L to CL functions. L function was rated with the highest score of USE-S, -pragmatic and -hedonic; (4) Transfer function had a significant main effect on motion sickness: the participants felt more headache and nausea when performing locomotion with CL function. Our results provide implications for body-centric locomotion design in VR applications. Boyu Gao 0003, Zijun Mai, Huawei Tu, Henry Been-Lirn Duh |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | LDGC-SR: Integrating long-range dependencies and global context information for session-based recommendation
Nan Qiu, Boyu Gao 0003, Huawei Tu, Feiran Huang, Quanlong Guan, Weiqi Luo 0002 |
Knowl. Based Syst. | 2 |
| 2021 | Effects of Different Proximity-Based Feedback on Virtual Hand Pointing in Virtual Reality
Yujun Lu, Boyu Gao 0003, Huawei Tu, Weiqi Luo 0002, HyungSeok Kim 0001 |
CGI | 2 |
| 2021 | Incorporating Global Context into Multi-task Learning for Session-Based Recommendation
Nan Qiu, Boyu Gao 0003, Feiran Huang, Huawei Tu, Weiqi Luo 0002 |
KSEM | 2 |
| 2021 | Evaluation of Body-centric Locomotion with Different Transfer Functions in Virtual RealityabstractBody-centric locomotion allows users to navigate virtual environments with body parts (e.g. head tilt, arm swing or torso lean). Transfer functions are an important determinant of the locus of such a locomotion method. However, there is little known about the effects of transfer functions on virtual locomotion with different body parts. In this work, we selected four typical transfer functions (linear function: L, power function: P, a piecewise function with constant and linear functions: CL, and a piecewise function with constant and power functions: CP) and four body parts (head, arm, torso, and knee) from existing works, and conducted an experiment to evaluate their effects on virtual locomotion under three distances (5, 10, and 15 m) in Virtual Reality (VR). Results show that (1) CP function generally led to the longest task time with a low rate of failed trials, while CL function had the shortest task time with a high rate of failed trials; (2) body parts significantly affected the rate of failed trials, but not task time and final position offset. Head and torso resulted in the lowest and highest rate of failed trials respectively; (3) body parts did not differ in User Experience Questionnaire-Short (UEQ-S), UEQ-S Pragmatic and UEQ-S Hedonic. L was rated as the highest score for UEQ-S, UEQ-S Pragmatic and UEQ-S Hedonic, but CP had the lowest score. According to the results, we provide implications of designing body-centric locomotion with different transfer functions in VR. Boyu Gao 0003, Ziiun Mai, Huawei Tu, Henry Been-Lirn Duh |
VR | 1 |
| 2021 | The effects of audiovisual landmarks on spatial learning and recalling for image browsing interface in virtual environments
Boyu Gao 0003, Huawei Tu, Feiran Huang |
J. Syst. Archit. | 1 |
| 2021 | A Cognitive Joint Angle Compensation System Based on Self-Feedback Fuzzy Neural Network With Incremental LearningabstractJoint angle error of robotic arm has great impacts on the accuracy of the end-effector, which is critical in industrial applications. Therefore, in this article, an online cognitive joint angle error compensation method based on incremental learning is proposed to reduce joint angle error. The proposed method consists of a joint angle error solver and a compensation module, which ensure that the robot can obtain effective joint angle compensation in various situations. The joint angle error solver is used to solve joint angle error online. It uses the redundant constraint method for multilink position measurement so as to calculate the position error of the robot accurately later. The compensation module uses the self-feedback incremental fuzzy neural network (SFIFN) to predict and update the compensation in real time. SFIFN is a variant of the fuzzy neural network (FNN), which uses long short-term memory to introduce a feedback mechanism based on FNN. The incremental learning capability of SFIFN reduces the time for solving error and makes the module runs in real time. Specifically, two inertial measurement units mounted at the ends of links are used to measure pose changes of the ends of corresponding links. Both the simulated and the real experiments show that the proposed method yields good compensations to joint angle error and its potentials for smart manipulation. Guanglong Du, Yinhao Liang, Boyu Gao 0003, Sattam Al Otaibi, Di Li 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | A Convolution Bidirectional Long Short-Term Memory Neural Network for Driver Emotion RecognitionabstractReal-time recognition of driver emotions can greatly improve traffic safety. With the rapid development of communication technology, it becomes possible to process large amounts of video data and identify the driver's emotions in real time. To effectively recognize driver's emotions, this paper proposes a new deep learning framework called Convolution Bidirectional Long Short-term Memory Neural Network (CBLNN). This method predicts the driver's emotion based on the geometric features extracted from facial skin information and the heart rate extracted from changes in RGB components. The facial geometry features obtained by using Convolutional Neural Network (CNN) are intermediate variables for the heart rate analysis of Bidirectional Long Short Term Memory (Bi-LSTM). Subsequently, the output of Bi-LSTM is used as input to the CNN module to extract the hear rate features. CBLNN uses Multi-modal factorized bilinear pooling (MFB) to fuse the extracted information and classifies it into five common emotions: happiness, anger, sadness, fear and neutrality. Our emotion recognition method was tested, proving that it can be used to quickly and steadily recognize emotions in real time. Guanglong Du, Zhiyao Wang, Boyu Gao 0003, Shahid Mumtaz, Khamael M. Abualnaja, Cuifeng Du |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Deep Attentive Multimodal Network Representation Learning for Social Media ImagesabstractThe analysis for social networks, such as the socially connected Internet of Things, has shown a deep influence of intelligent information processing technology on industrial systems for Smart Cities. The goal of social media representation learning is to learn dense, low-dimensional, and continuous representations for multimodal data within social networks, facilitating many real-world applications. Since social media images are usually accompanied by rich metadata (e.g., textual descriptions, tags, groups, and submitted users), simply modeling the image is not effective to learn the comprehensive information from social media images. In this work, we treat the image and its textual description as multimodal content, and transform other metainformation into the links between contents (such as two images marked by the same tag or submitted by the same user). Based on the multimodal content and social links, we propose a Deep Attentive Multimodal Graph Embedding model named DAMGE for more effective social image representation learning. We introduce both small- and large-scale datasets to conduct extensive experiments, of which the results confirm the superiority of the proposal on the tasks of social image classification and link prediction. Feiran Huang, Chaozhuo Li, Boyu Gao 0003, Yun Liu 0017, Sattam Al Otaibi, Hao Chen 0062 |
ACM Trans. Internet Techn. | 3 |
| 2020 | Transfer of Coordination Skill to the Unpracticed Hand in Immersive EnvironmentsabstractPhysical practice with one hand results in performance gains of the other (un-practiced) hand in a unilateral motor task. Yet how it induces performance gains of interlimb coordination in the bimanual movements between trained limb and the opposite, untrained limb is unclear. The present study designed a game-like interactive system for physical practice, in which an avatar’s hands could be controlled itself or by the subject during a bimanual movement task in an immersive virtual reality environment. Participants practiced with the bimanual task by simultaneously drawing non-symmetric three-sided squares (e.g., U and C) to learn limb coordination with the following training strategies: (1) performing and seeing a bimanual task (BH-BH); (2) performing a unimanual task with right hand and seeing a bimanual action (RH-BH); (3) not performing a task but seeing a bimanual action (noH-BH); (4) performing and seeing a unimanual task (RH-RH). We found that the learning performance was better after BH-BH and RH-BH compared with other training strategies. In addition, we examined the effects of virtual hand representations on the learning performance after RH-BH. We found that the performance after training was increased with the realism level of virtual hands. These findings suggest that the proposed approach of RH-BH with realistic virtual hand would result in transfer of coordination skill to the unpracticed hand, which puts forward a new approach for learning and rehabilitation of coordination skill in patients with unilateral motor deficit in immersive environments. Shan Xiao, Xupeng Ye, Yaqiu Guo, Boyu Gao 0003, Jinyi Long |
VR | 4 |
| 2020 | Effects of holding postures on user-defined touch gestures for tablet interaction
Huawei Tu, Qihan Huang, Yanchao Zhao, Boyu Gao 0003 |
Int. J. Hum. Comput. Stud. | 4 |
| 2020 | Natural Human-Machine Interface With Gesture Tracking and Cartesian Platform for Contactless Electromagnetic Force FeedbackabstractIn this article, a novel human-machine interface, in which two Leap Motion (LM) controllers and a coil are attached to a Cartesian platform to provide contactless electromagnetic force feedback for enhancing the accuracy and efficiency of human-robot manipulation tasks is presented. To implement such an interface, an interval Kalman filter, an improved particle filter, and a mean filter are integrated to estimate accurately the position and orientation of the hand gesture tracked by the two LM controllers, and to smoothen the movement of the Cartesian platform. The back propagation neural network is employed to regulate the electric currents of the coil attached to the Cartesian platform for accurate force feedback. A series of comparative experiments are performed, and the results show that the presented interface greatly improved the efficiency and accuracy of human-robot manipulation tasks in comparison with existing methods, indicating its great potentials for many industry scenarios. Guanglong Du, Chunquan Li 0001, Boyu Gao 0003, Peter Xiaoping Liu |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | Amphitheater Layout with Egocentric Distance-Based Item Sizing and Landmarks for Browsing in Virtual RealityabstractTo allow more rapid and accurate accesses in the virtual reality browsing interfaces, an amphitheater layout with varying egocentric distance-based item sizing (EDIS) is presented, and items at different egocentric distances with sizes proportional to the distances are placed. The effects of EDIS variations on the retrieval and recall performance depend on the distance configurations. Experiments 1, 2, and 3 showed that small and medium EDIS variations gave efficient trial completions in near-field distance perception for small and medium item sets. For large-item sets, the further items in the amphitheater layout may lie beyond the near-field distance perception. We therefore adopt additional 3D visual landmarks in the amphitheater layout (Experiment 4); while both the location-fixed and user-defined 3D pins greatly improved completion and accuracy performance, they showed no significant difference between them, indicating that location-fixed landmarks would be suffice. These findings suggest the use of spatial and visual landmarks for effective VR browsing interfaces. Boyu Gao 0003, Jee-In Kim, HyungSeok Kim 0001 |
Int. J. Hum. Comput. Interact. | 1 |
| 2018 | Sensory and Perceptual Consistency for Believable Response in Action Feedback LoopabstractMost Virtual Reality1(VR) applications are dedicated to providing a realistic response to users. Prior works have argued that the importance of credibility or believability to evoke realistic responses. However, there are few empirical works focused on the concept of believability itself. This work further explores the evoke-condition of believability from the views of sensory and perceptual aspects. Sensory believability refers to the actions supported by the Virtual Environment (VE) match with common sense, whereas perceptual believability refers to the behaviors supported by the VE are consistent with the expectation formed during interactions or from prior experiences. Towards the end, an experiment that examines the sensory (rock appearance, environmental visual scene, environmental sound) and perceptual (dynamic behavior) factors of consistency on believable response in a virtual rock climbing environment was conducted. The results showed that the consistency of dynamics features contributed the most to perceptual believability, followed by a higher level of rock appearance along with climbing actions, even with the medium level of the environmental visual scene. In contrast, the medium and higher level of environmental visual scene contributed the most to sensory believability. This study suggests the sensory and perceptual consistency together could evoke the believable response in action feedback loop for VR. Boyu Gao 0003, Jee-In Kim, HyungSeok Kim 0001 |
CGI | 1 |
| 2018 | Effects of Continuous Auditory Feedback on Drawing Trajectory-Based Finger GesturesabstractThe well-known “fat finger” issue limits the interaction performance of trajectory-based finger gestures. To alleviate this issue, this work focuses on the possibility of using additional continuous auditory feedback to assist trajectory-based finger gestures. First, the experiment validated that, with the visual feedback only, the bare fingertip led to more errors in drawing of intersectional points, endpoints of closed gestures, and gestural length and shape variability compared to when the finger-attached pen was used. Then, we designed different types of auditory feedback (discrete beep, static, gradual) to provide additional information on the spatial relationship between finger-contact point and the endpoints or intersections of predefined gestures. An experiment that evaluates the effects of individual or combination of designed auditory feedback on trajectory-based finger gestures was conducted. These results show a few differences between them. However, a combination of gradual (amplitude and frequency) continuous sound and beep reached the highest drawing accuracy for trajectory-based finger gestures, which is similar to that of a finger-attached pen. This research offers insights and implications for the future design of continuous auditory feedback on small touchscreens. Boyu Gao 0003, HyungSeok Kim 0001, Hasup Lee 0001, Jooyoung Lee 0003, Jee-In Kim |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2015 | Multiple devices as windows for virtual environmentabstractWe introduce a method for using multiple devices as windows for interacting with 3-D virtual environment. Motivation of our work has come from generating collaborative workspace with multiple devices which can be found in our daily lives, like desktop PC and mobile devices. Provided with life size virtual environment, each device shows a scene of 3-D virtual space on its position and direction, and users would be able to perceive virtual space in more immersive way with it. By adopting mobile device to our system, users not only see outer space of stationary screen by relocating their mobile device, but also have personalized view in working space. To acquiring knowledge of device's pose and orientation, we adopt vision-based approaches. For the last, we introduce an implementation of a system for managing multiple device and letting them have synchronized performance. Jooyoung Lee 0003, Hasup Lee 0001, Boyu Gao 0003, HyungSeok Kim 0001, Jee-In Kim |
VR | 3 |