Yue Liu 0005

dblp:74/1932-5 · DBLP profile ↗
← Back
93ranked-venue papers
0as first author
39since 2021 · last 2026
0000-0002-6784-8802ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 73 · 28 since 2021Human-computer interaction and ubiquitous computing · 25 · 7 since 2021Artificial intelligence and machine learning · 11 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 since 2021Systems, architecture and hardware · 2Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Fully utilizing cross modal features to achieve precise segmentation of brain gliomas
Weiye Cao, Kaiyan Zhu, Jianhao Xu, Yue Liu 0005
Multim. Syst.5
2026 Focus-guided feature fusion network for lightweight image super-resolution
Tongtai Cao, Songyue Yang, Yue Liu 0005
Neural Networks5
2026 Super-Resolution Reconstruction of OCTA via Multi-Field-of-View Representation Learning
abstract
High-resolution Optical Coherence Tomography Angiography (OCTA) images are essential for morphological analysis and biomarker measurement of the retinal vasculature. They can also provide underlying biomarkers for the accurate analysis of eye-related diseases. The trade-off between the high resolution (HR) and large scanning field-of-view (FOV) is a long-standing problem for OCTA image instrument. A large FOV image provides more retinal information with shorter acquisition time but often suffers from low resolution (LR), high scatter noise, and poor vascular contrast. In order to obtain HR OCTA images with larger FOV, we propose a novel self-similar dynamic domain adaptation network based on cross-field-of-view representation learning. The network enables LR images (i.e., $6\times \text{6}\,\text{mm}^{2}$) to learn HR image (i.e., $3\times \text{3}\,\text{mm}^{2}$) feature representations specialized for OCTA by constructing feature mapping relations for cross-field-of-view OCTA scans. To be specific, a multiple random degradation model is proposed on HR images to generate various synthetic LR images. Further, we propose a dynamic domain adaptation framework that prompts feature dynamic alignment of the LR image reconstruction results with those of synthetic LR images. Finally, a novel self-similar supervision loss is proposed to optimize the reconstruction results from LR to HR by exploiting the similarity between vessels in different regions. Experimental results on three OCTA datasets show that the proposed method surpasses existing state-of-the-art ones, significantly enhancing retinal structure segmentation and disease classification. Our OCTA dataset (the first dataset in this research area with paired $3\times 3$ and $6\times \text{6}\,\text{mm}^{2}$ OCTA images) and code are publicly available.
Huaying Hao, Shaoyi Leng, Yanda Meng, Yonghuai Liu, Yalin Zheng, Huazhu Fu, Jiong Zhang 0004, Quanyong Yi, Yue Liu 0005, Jingfeng Zhang, Yitian Zhao
IEEE J. Biomed. Health Informatics9
2026 PortInput: Enabling Always-Available Micro-Gesture Input With Pressure Array Sensor
abstract
Micro-gestures provide a natural, efficient, and privacy-preserving input modality; however, existing techniques often depend on environmental conditions, limiting their robustness and applicability in real-world settings. In this work, we present three portable prototypes- - finger-cot, finger-worn, and surface-based-that integrate compact pressure array sensors to support environment-independent micro-gesture interaction. We further propose a deep learning-based recognition model that accurately classifies 14 micro-gestures by analyzing temporal pressure patterns. Building upon these components, we introduce PortInput, a real-time interactive system that enables robust micro-gesture tracking and detection. We conducted two user studies with augmented reality (AR) head-mounted displays (HMDs). The first study evaluates input performance under both sitting and walking conditions, while the second compares PortInput with a commercial pressure-based ring device. The results show that PortInput improves usability and user experience, achieves comparable accuracy, and enables faster input with lower perceived workload. Overall, PortInput have potential to offer efficient, robust, and comfortable input across diverse application scenarios-ranging from AR/Virtual Reality (VR) headsets to smart homes and in-car systems-even in noisy or cluttered environments. This work provides a foundation for integrating pressure array sensors into ring-based or other portable devices, advancing always-available micro-gesture interaction for ubiquitous computing environments.
Henry Been-Lirn Duh, Mingwei Hu, Yue Liu 0005, Yongtian Wang
IEEE Trans. Vis. Comput. Graph.6
2026 Investigating the Effects of Egocentric FPV vs. Exocentric TPV Perspectives and Collaboration Modes on Sense of Agency in VR Teleoperation
abstract
With the advancement of virtual reality (VR), human-robot interaction within remote collaboration contexts has emerged as a pivotal research area. However, the nuanced interplay between the user's perspective and the robot's level of automation, and its subsequent impact on the operator's sense of agency (SoA), remains underexplored. This study investigates how perspective and collaboration mode-differentiated by the robot's level of automation-affect the SoA in VR-based human-robot collaboration. We implemented a 2 (perspective: first-person view[FPV] vs. third-person view[TPV]) × 3 (collaboration mode: Direct Remote Control, Supervisory Cooperation, Spectator Mode) experimental design. SoA was assessed via a multimodal approach, measuring both the explicit judgement of agency (JoA) and the implicit feeling of agency (FoA). The results revealed a dissociation between these two components: explicit JoA was dictated by the collaboration mode, with active modes yielding a higher explicit JoA than the passive Spectator Mode. In contrast, implicit FoA exhibited a significant interaction effect. Under FPV, the influence of collaboration mode was pronounced, where Direct Remote Control elicited the strongest FoA; however, under TPV, the differences among collaboration modes disappeared. The relationship between cybersickness and SoA was also explored. These findings suggest that SoA is not a unitary concept but a multifaceted experience, co-shaped by the user's perspective and their role as either a decision-maker or an executor. Furthermore, we investigate the relationship between cybersickness and SoA. Our study offers novel insights for the design of VR-based human-robot systems.
Zhonghao Zhu, Kang Yueh, Songyue Yang, Fanlu Zeng, Yue Liu 0005
IEEE Trans. Vis. Comput. Graph.5
2026 Efficient shallow feature extraction for lightweight super-resolution via transformer
Tongtai Cao, Pengjie Zhao, Yue Liu 0005
Vis. Comput.5
2025 Lightweight Image Super-Resolution Using Fine-Grained Feature Distillation in a Dense Residual U-Net
Tongtai Cao, Huaying Hao, Yue Liu 0005
CGI (3)5
2025 Cross-Media Color Appearance Reproduction in Optical See-Through Augmented Reality
abstract
In optical see-through (OST) augmented reality (AR), displayed colors blend with the real-world scene, affecting perceived color. Studies showed that AR's color appearance depends not only on the additive chromaticity of the display and the real scene but also on ambient illumination. However, these studies often overlook changes in observer's adaptation state under varying illumination. To address this, a series of color-matching experiments between AR and display devices was conducted in an immersive lighting environment. The first experiment under a D65 illuminant found a slightly higher correlated color temperature (CCT) level of the internal white point in OST AR than in the display. Considering media's white point differences, this study proposes a three-step chromatic adaptation transform (CAT) framework to improve color appearance reproduction accuracy in AR. The second experiment used three Planckian radiators varying in CCT and two offPlanckian colorful ones to validate the proposed CAT under varying illumination, indicating that AR reached a more complete adaptation state than the display, especially under high luminance. A third validation experiment had observers rate color differences between reference and reproduced stimuli in AR. Results showed the effectiveness of our three-step CAT for cross-media color reproduction of OST AR under diverse illumination conditions.
Jiahong Luo, Shining Ma, Yue Liu 0005, Yongtian Wang
ISMAR3
2025 Improving Pointing Accuracy for 3D Target Selection in Virtual Reality Through Depth Perception Biases Correction
abstract
Accurate 3D target selection in virtual reality (VR) is fundamentally impeded by pointing uncertainty along the depth axis, a challenge that existing 2D pointing models fail to address due to the complexities of depth perception. Near-eye interactions in VR are influenced by binocular depth cues and vergence-accommodation conflicts (VAC), which introduce significant depth perception biases that impair predictive performance. To address this issue, we first investigate these factors and derive a Gaussian distribution to model near-field depth biases within a 2.5m range. Second, to analyze pointing performance across this extended depth range, we classify 3D target motions into three distinct types: motion-indepth, motion-in-plane, and combined motion. Our analysis identifies that interaction depth and motion amplitude are the two most critical factors influencing pointing accuracy. Accordingly, by incorporating these factors alongside our perceptual bias Gaussian into the Ternary-Gaussian framework, we demonstrate significantly improved predictive performance across diverse 3D motion scenarios. These findings enhance the understanding of user perception in virtual environments and support the development of precise, context-aware interaction cues. Future research can extend these models to design real-time adaptive interfaces, thereby elevating user experiences in VR.
Songyue Yang, Kang Yue, Haolin Gao, Yiyi Yang, Mei Guo, Yu Liu 0081, Zhonghao Zhu, Yue Liu 0005
ISMAR8
2025 Lossless Intrinsic Image Decomposition via Learning Shading Feature Filtering
abstract
Intrinsic image decomposition decomposes an image into reflectance and shading. It has been applied in image editing, augmented reality, and geometry estimation. However, the complete decoupling between reflectance and shading, as well as the consistency of the reconstructed image with the original image, have become the main challenges in the application of intrinsic image decomposition. To improve the performance of the intrinsic image decomposition algorithm for these two challenges, we propose a novel deep learning framework that works separately to learn features unique to different intrinsic images. Based on this framework, we developed more effective loss functions to strengthen the decoupling of reflectance and shading and to maintain the decomposition without losing as much information of the original image as possible. We trained the network on a mixture of synthetic and real datasets and evaluated the results of the experiments on real datasets. The results show that our proposed method not only outperformed existing state-of-the-art methods in qualitative and quantitative comparisons in terms of reflectance but was also competitive in terms of reconstructed consistency and shading. Finally, we implemented several realistic image-editing applications, and the results were visually superior to other results.
Hao Sha 0004, Yu Han 0011, Yi Xiao 0009, Yue Liu 0005
Comput. Vis. Media5
2025 Lightweight Local-Global Dual-Path Feature Fusion Network for Infrared Small Target Image Super-Resolution and Enhancement
abstract
Infrared imaging is widely used in remote sensing and military target recognition due to its strong resistance to interference in complex environments. However, imaging mechanisms and hardware limitations cause infrared images to have low-resolution, sparse textures, and significant background noise, which severely restrict the detection of small targets such as low-altitude drones and weak thermal emitters. To overcome these limitations, we propose a lightweight Local–Global Dual-Path Feature Fusion Network (LDFF-Net) that enhances the resolution and quality of infrared images, providing high-quality inputs for subsequent detection tasks. The network includes a Small Target Feature Recognition Module (STFRM) composed of three key components. The Enhanced High Frequency Perception Module (EHFPM) strengthens high-frequency details of small targets while suppressing noise, enabling robust local feature extraction. The State-Space Model (SSM) captures long-range dependencies with linear complexity and models semantic relationships between targets and background to compensate for the limited receptive field of local features. The Adaptive Feature Fusion Unit (AFFU) combines local and global features adaptively to improve the saliency of small targets. During training, we introduce a realistic degradation process based on visible-light images to generate training samples that include complex degradation patterns and noise, which enhances the model’s robustness and generalization. Evaluation on the ARCHIVE and SIRST datasets demonstrates that LDFF-Net outperforms existing state-of-the-art methods across eight widely used full-reference and no-reference metrics, including PSNR, LPIPS, FID, and NIQE. This result confirms the model’s effectiveness in enhancing both the super-resolution and detection performance of infrared small target images. The code and pretrained model weights are publicly available at https://github.com/98Hao/LDFF-Net.
Songyue Yang, Tongtai Cao, Yue Liu 0005
IEEE Trans. Geosci. Remote. Sens.5
2025 Cross-Optical Property Image Translation for Face Anti-Spoofing: From Visible to Polarization
abstract
Despite the development of spectral sensors and spectral data-driven learning methods which have led to significant advances in face anti-spoofing (FAS), the singular dimensionality of spectral information often results in poor robustness and weak generalization. Polarization, another fundamental property of light, can reveal intrinsic differences between genuine and fake faces with advantaged performance in precision, robustness, and generalizability. In this paper, we propose a facial image translation method from visible light (VIS) to polarization (VPT), capable of generating valuable polarimetric optical characteristics for facial presentation attack detection using VIS spectrum information input only. Specifically, the VPT method adopts a multi-stream network structure, comprising a main network and two branch networks, to translate VIS images into degree of polarization (DoP) images and Stokes polarization parameters${S}_{1}$and${S}_{2}$. To further improve image translation quality, we introduce a frequency-domain consistency loss as a complement to the existing spatial losses to narrow the gap in the frequency domain. The physical mapping relations for the DoP and Stokes parameters are employed, and the Stokes loss is designed to ensure that the generated polarization modalities conform to objective physical laws. Extensive experiments on the CASIA-Polar and CASIA-SURF datasets demonstrate the superiority of VPT over other baseline methods in terms of polarization image quality and its remarkable performance in the FAS task. This work leverages the inherent physical advantages of polarization information in material discrimination tasks while addressing hardware limitations in polarization image collection, proposing a novel solution for face recognition system security control.
Yu Tian 0017, Kunbo Zhang, Yalin Huang, Leyuan Wang, Yue Liu 0005, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.5
2025 Uncertainty Quantification for Incomplete Multi-View Data Using Divergence Measures
abstract
Existing multi-view classification and clustering methods typically improve task accuracy by leveraging and fusing information from different views. However, ensuring the reliability of multi-view integration and final decisions is crucial, particularly when dealing with noisy or corrupted data. Current methods often rely on Kullback-Leibler (KL) divergence to estimate uncertainty of network predictions, ignoring domain gaps between different modalities. To address this issue, KPHD-Net, based on Hölder divergence, is proposed for multi-view classification and clustering tasks. Generally, our KPHD-Net employs a variational Dirichlet distribution to represent class probability distributions, models evidences from different views, and then integrates it with Dempster-Shafer evidence theory (DST) to improve uncertainty estimation effects. Our theoretical analysis demonstrates that Proper Hölder divergence offers a more effective measure of distribution discrepancies, ensuring enhanced performance in multi-view learning. Moreover, Dempster-Shafer evidence theory, recognized for its superior performance in multi-view fusion tasks, is introduced and combined with the Kalman filter to provide future state estimations. This integration further enhances the reliability of the final fusion results. Extensive experiments show that the proposed KPHD-Net outperforms the current state-of-the-art methods in both classification and clustering tasks regarding accuracy, robustness, and reliability, with theoretical guarantees.
Zhipeng Xue 0001, Yan Zhang 0119, Ming Li 0073, Yue Liu 0005, F. Richard Yu
IEEE Trans. Image Process.5
2025 Utilizing Gaze-Contingent Rendering to Maintain Visual Attention in Educational VR
abstract
In educational Virtual Reality (VR) environments, objects irrelevant to learning can lead to students' inattention, which adversely affects learning. However, removing these objects from virtual scenes is not feasible, as they are crucial for creating a realistic and immersive experience. Balancing the need to maintain students' attention while preserving the integrity of scenarios is a challenging task. In this paper, we introduce a gaze-contingent rendering (GCR) technique to address such an issue, which is independent of specific elements or configurations in virtual scenes and adaptable across various contexts. Specifically, we utilize gaze-aware rendering adjustments to adaptively reduce the visibility of objects irrelevant to learning while highlighting relevant ones. We develop three GCR strategies (i.e., blur, pixelation, and underexposure) and investigate how these strategies affect students' visual attention, academic achievement, and perceptions of the learning activity across different scenarios. Our findings indicate that the proposed rendering strategies effectively achieve the goals of sustaining visual attention and improving academic achievement without significantly impacting immersion or engagement. As an initial exploration of GCR for maintaining attention within educational VR, this study may inspire new directions in future research on GCR and visual attention maintenance in immersive VR.
Yu Han 0011, Hao Sha 0004, Yi Xiao 0009, Yue Liu 0005
IEEE Trans. Vis. Comput. Graph.6
2025 Errata to "Depth Perception in Optical See-Through Augmented Reality: Investigating the Impact of Texture Density, Luminance Contrast, and Color Contrast"
abstract
In This paper, the information regarding the corresponding authors is missing. The corresponding authors of the paper should be Shining Ma and Weitao Song.
Chaochao Liu, Shining Ma, Yue Liu 0005, Yongtian Wang
IEEE Trans. Vis. Comput. Graph.3
2025 AudioGest: Gesture-Based Interaction for Virtual Reality Using Audio Devices
abstract
Current virtual reality (VR) system takes gesture interaction based on camera, handle and touch screen as one of the mainstream interaction methods, which can provide accurate gesture input for it. However, limited by application forms and the volume of devices, these methods cannot extend the interaction area to such surfaces as walls and tables. To address the above challenge, we propose AudioGest, a portable, plug-and-play system that detects the audio signal generated by finger tapping and sliding on the surface through a set of microphone devices without extensive calibration. First, an audio synthesis-recognition pipeline based on micro-contact dynamics simulation is constructed to generate modal audio synthesis from different materials and physical properties. Then the accuracy and effectiveness of the synthetic audio are verified by mixing the synthetic audio with real audio proportionally as the training sets. Finally, a series of desktop office applications are developed to demonstrate the application potential of AudioGest's scalability and versatility in VR scenarios.
Yi Xiao 0009, Mingwei Hu, Hao Sha 0004, Shining Ma, Boyu Gao 0003, Shihui Guo, Yue Liu 0005
IEEE Trans. Vis. Comput. Graph.8
2025 Influence of Object Height, Shadow and Adapting Luminance on Outdoor Depth Perception in Augmented Reality
abstract
Augmented reality (AR) technology has great potential in the applications of training, exhibition, and visual guidance, all of which demand precise virtual-real registration in perceived depth. Many AR applications such as navigation and tourism guidance are usually implemented in outdoor environments. However, prior research on depth perception in AR predominantly focused on the indoor environment, characterized by a lower illumination level and more confined space compared to outdoor settings. To address this gap, this paper presented a systematic investigation into the depth perception in outdoor environments. Two experiments were conducted in this study: the first one aimed to explore how to eliminate the bias induced by the floating object and how the knowledge of object height influences the perceived depth. The second experiment examined how ambient luminance affects depth estimation in AR. Our findings revealed an overestimation of perceived depth when participants were unaware of the actual height of the floating object, but an underestimation when they were informed of this information prior to the experiment. Additionally, shadows effectively reduced depth errors regardless of whether participants were informed of the object's height. The second experiment further indicated that, in outdoor environments, reducing ambient luminance significantly improves the accuracy of depth perception in AR.
Shining Ma, Chaochao Liu, Yue Liu 0005, Yongtian Wang
IEEE Trans. Vis. Comput. Graph.4
2025 Intrinsic Decomposition With Robustly Separating and Restoring Colored Illumination
abstract
Intrinsic decomposition separates an image into reflectance and shading, which contributes to image editing, augmented reality, etc. Despite recent efforts dedicated to this field, effectively separating colored illumination from reflectance and correctly restoring it into shading remains an challenge. We propose a deep intrinsic decomposition method to address this issue. Specifically, by transforming intrinsic decomposition process in RGB image domains into the combination of intensity and chromaticity domains, we propose a novel macro intrinsic decomposition network framework. This framework enables the generation of finer intrinsic components through more relevant features propagation and more detailed sub-constraints guidance. In order to expand the macro network, we integrate multiple attention mechanism modules in key positions of encoders, which enhances the extraction of distinct features. We also propose a skip connection module based on specific deep features guidance, which can filter out features that are physically irrelevant to each intrinsic component. Our method not only outperforms state-of-the-art methods across multiple datasets, but also robustly separates illumination from reflectance and restores it into shading in various types of images. By leveraging our intrinsic images, we achieve visually superior image editing effects compared to other methods, while also being able to manipulate the inherent lighting of the original scene.
Hao Sha 0004, Shining Ma, Tongtai Cao, Yu Han 0011, Yu Liu 0081, Yue Liu 0005
IEEE Trans. Vis. Comput. Graph.6
2025 Dynamic Changes of Latency Perception Threshold in Virtual Reality: Behavioral and EEG Evidence
abstract
Virtual Reality (VR) technologies in fields such as telehealth, teleconferencing, and virtual education are significantly affected by end-to-end latency, which notably impacts users' interactive experience and performance. Previous research suggests that a perceptual threshold may exist-once latency is reduced below a certain level, users no longer perceive it, and their interactive performance remains largely unaffected. However, there is no consensus on the exact value of this absolute latency perception threshold. In this study, we employed an experimental design based on Fitts' law to investigate whether interaction strategies and task difficulty can alter the latency perception threshold (LPT), and how variations in this threshold influence users' interactive performance. The results show that the LPT is approximately 130-170 ms, and that when interaction strategies prioritize speed or when tasks become more challenging, users exhibit heightened sensitivity to latency. Due to the presence of the LPT, the effect of latency on interactive performance follows a nonlinear pattern, and building on this finding, we refined a Fitts' law model to incorporate the influence of latency. Notably, electroencephalogram (EEG) signals can still capture users' perception of latency when they are unaware of minor latency, demonstrating a level of sensitivity that exceeds conscious awareness. Our findings provide insights into latency effects on performance and perception, guiding the design of more responsive VR interaction systems.
Songyue Yang, Kang Yue, Haolin Gao, Mei Guo, Yu Liu 0081, Dan Zhang 0014, Yue Liu 0005
IEEE Trans. Vis. Comput. Graph.7
2025 Learning intrinsic decomposition with semantic information fusion based on transformer
abstract
Intrinsic decomposition, the process of decomposing an image into reflectance and shading, is widely used in virtual and augmented reality tasks. Reflectance and shading often exhibit large gradients at the object edges, and the intrinsic properties on the same object tend to be similar. This spatial coherence is closely related to semantic consistency because objects within the same semantic category often exhibit similar intrinsic properties. Therefore, incorporating semantic segmentation into a deep intrinsic decomposition framework helps the network distinguish between different object instances and understand high-level scene structures. To this end, we design an intrinsic decomposition network jointly trained with a dedicated semantic segmentation module, allowing semantic cues to enhance the decomposition of reflectance and shading. The semantic module provides guidance during training but is removed during inference, improving performance without increasing the inference cost. Additionally, to capture the global contextual dependencies critical for intrinsic decomposition, we adopt a Transformer-based backbone. The proposed backbone enables the model to associate distant regions with similar material properties, thereby maintaining consistency in reflectance and learning smooth illumination patterns across a scene. A convolutional decoder is also designed to output predictions with improved details. Experiments demonstrate that our approach achieves state-of-the-art performance in the quantitative evaluations on the Intrinsic Images in the Wild (IIW) and Shading Annotations in the wild (SAW) datasets.
Pengjie Zhao, Hao Sha 0004, Yongtian Wang, Yue Liu 0005
Virtual Real. Intell. Hardw.4
2024 Exploring Depth-based Perception Conflicts in Virtual Reality through Error-Related Potentials
abstract
Virtual Reality (VR) offers a valuable platform for real-life skills training. However, previous research has indicated that human’s perception of depth in VR differs from that of the real world. Such perceptual conflicts can impact immersion and the learning of skills, thus attracting widespread attention. Various methods have been proposed to enhance users’ depth perception, yet the underlying mechanisms of depth perception conflicts still require further research. In this paper, we used Error-Related Potentials (ErrPs) from electroencephalography (EEG) data to investigate the differences in participants’ perceptions at varying depths within the near-field. We designed a within-subjects experiment to successfully introduce depth perception conflicts. From participants exposed to three distinct depths, we collected questionnaire results, performance data, and EEG data. Our findings showed that EEG can effectively detect depth perception conflicts and, following each conflict, participants’ behavioral patterns showed significant changes. In situations with shallower depths, participants exhibited stronger responses to the designed conflicts. This increased sensitivity correlates with their accuracy in depth estimation. This study represents a novel approach to depth perception in VR using ErrPs, setting the stage for further use of physiological signals to measure the granularity of depth perception in VR/AR environments.
Haolin Gao, Kang Yue, Songyue Yang, Yu Liu 0081, Mei Guo, Yue Liu 0005
VR6
2024 Realtime Recognition of Dynamic Hand Gestures in Practical Applications
abstract
Dynamic hand gesture acting as a semaphoric gesture is a practical and intuitive mid-air gesture interface. Nowadays benefiting from the development of deep convolutional networks, the gesture recognition has already achieved a high accuracy, however, when performing a dynamic hand gesture such as gestures of direction commands, some unintentional actions are easily misrecognized due to the similarity of the hand poses. This hinders the application of dynamic hand gestures and cannot be solved by just improving the accuracy of the applied algorithm on public datasets, thus it is necessary to study such problems from the perspective of human-computer interaction. In this article, two methods are proposed to avoid misrecognition by introducing activation delay and using asymmetric gesture design. First the temporal process of a dynamic hand gesture is decomposed and redefined, then a realtime dynamic hand gesture recognition system is built through a two-dimensional convolutional neural network. In order to investigate the influence of activation delay and asymmetric gesture design on system performance, a user study is conducted and experimental results show that the two proposed methods can effectively avoid misrecognition. The two methods proposed in this article can provide valuable guidance for researchers when designing realtime recognition system in practical applications.
Yi Xiao 0009, Yu Han 0011, Yue Liu 0005, Yongtian Wang
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Depth Perception in Optical See-Through Augmented Reality: Investigating the Impact of Texture Density, Luminance Contrast, and Color Contrast
abstract
The immersive augmented reality (AR) system necessitates precise depth registration between virtual objects and the real scene. Prior studies have emphasized the efficacy of surface texture in providing depth cues to enhance depth perception across various media, including the real scene, virtual reality, and AR. However, these studies predominantly focus on black-and-white textures, leaving a gap in understanding the effectiveness of colored textures. To address this gap and further explore texture-related factors in AR, a series of experiments were conducted to investigate the effects of different texture cues on depth perception using the perceptual matching method. Findings indicate that the absolute depth error increases with decreasing contrast under black-and-white texture. Moreover, textures with higher color contrast also contribute to enhanced accuracy of depth judgments in AR. However, no significant effect of texture density on depth perception was observed. The findings serve as a theoretical reference for texture design in AR, aiding in the optimization of virtual-real registration processes.
Chaochao Liu, Shining Ma, Yue Liu 0005, Yongtian Wang
IEEE Trans. Vis. Comput. Graph.3
2024 Effects of virtual agents on interaction efficiency and environmental immersion in MR environments
abstract
Physical entity interactions in mixed reality (MR) environments aim to harness human capabilities in manipulating physical objects, thereby enhancing virtual environment (VEs) functionality. In MR, a common strategy is to use virtual agents as substitutes for physical entities, balancing interaction efficiency with environmental immersion. However, the impact of virtual agent size and form on interaction performance remains unclear. Two experiments were conducted to explore how virtual agent size and form affect interaction performance, immersion, and preference in MR environments. The first experiment assessed five virtual agent sizes (25%, 50%, 75%, 100%, and 125% of physical size). The second experiment tested four types of frames (no frame, consistent frame, half frame, and surrounding frame) across all agent sizes. Participants, utilizing a head-mounted display, performed tasks involving moving cups, typing words, and using a mouse. They completed questionnaires assessing aspects such as the virtual environment effects, interaction effects, collision concerns, and preferences. Results from the first experiment revealed that agents matching physical object size produced the best overall performance. The second experiment demonstrated that consistent framing notably enhances interaction accuracy and speed but reduces immersion. To balance efficiency and immersion, frameless agents matching physical object sizes were deemed optimal. Virtual agents matching physical entity sizes enhance user experience and interaction performance. Conversely, familiar frames from 2D interfaces detrimentally affect interaction and immersion in virtual spaces. This study provides valuable insights for the future development of MR systems.
Yihua Bao, Jie Guo 0004, Dongdong Weng, Yue Liu 0005, Zeyu Tian
Virtual Real. Intell. Hardw.4
2023 Skeleton Based Dynamic Hand Gesture Recognition using Short Term Sampling Neural Networks (STSNN)
Aamrah Ikram, Yue Liu 0005
ICIG (1)2
2023 Polarized Image Translation From Nonpolarized Cameras for Multimodal Face Anti-Spoofing
abstract
In face antispoofing, it is desirable to have multimodal images to demonstrate liveness cues from various perspectives. However, in most face recognition scenarios, only a single modality, namely visible lighting (VIS) facial images is available. This paper first investigates the possibility of generating polarized (Polar) images from VIS cameras without changing the existing recognition devices to improve the accuracy and robustness of Presentation Attack Detection (PAD) in face biometrics. A novel multimodal face antispoofing framework is proposed based on the machine-learning relationship between VIS and Polar images of genuine faces. Specifically, a dual-modal central differential convolutional network (CDCN) is developed to capture the inherent spoofing features between the VIS and the generated Polar modalities. Quantitative and qualitative experimental results show that our proposed framework not only generates realistic Polar face images but also improves the state-of-the-art face anti-spoofing results on the VIS modal database (i.e. CASIA-SURF). Moreover, a polar face database, CASIA-Polar, has been constructed and will be shared with the public at http://biometrics.idealtest.org to inspire future applications within the biometric anti-spoofing field.
Yu Tian 0017, Yalin Huang, Kunbo Zhang, Yue Liu 0005, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.4
2022 Multi-scale Vertical Cross-layer Feature Aggregation and Attention Fusion Network for Object Detection
Wenting Gao, Xiaojuan Li 0005, Yu Han 0011, Yue Liu 0005
ICANN (4)4
2022 NailRing: An Intelligent Ring for Recognizing Micro-gestures in Mixed Reality
abstract
Gesture interaction is currently a main interaction technology in the field of mixed reality. However, long-term and large-scale gesture in mid-air will lead to muscle fatigue and privacy problems, which cannot meet the comfort requirements of continuous interaction and inevitably hinder the development of mixed reality systems. To solve this problem, we propose NailRing, an intelligent ring to recognize fingertip micro-gestures using a micro-close-focus camera on a fingertip bracket. Such fingertip physiological characteristics as the changes in fingertip color distribution and muscle shape changes caused by fingertip pressure have been studied. According to the recognition principle, ten types of micro-gestures have been designed and used for contact interaction and one-hand interaction respectively. The accuracy of gesture recognition (cross-session $ F_{Macro}=98.3\%$; cross-person $ F_{Macro}=86.4\%$) in user studies verifies the performances of NailRing under different interaction conditions. Finally, the capability of NailRing in a series of potential application scenarios has also been discussed and analyzed.
Yue Liu 0005, Shining Ma, Mingwei Hu
ISMAR2
2022 GAN-based image-to-friction generation for tactile simulation of fabric material
Shaoyu Cai, Yuki Ban, Takuji Narumi, Yue Liu 0005, Kening Zhu
Comput. Graph.5
2022 Professional and High-Level Gamers: Differences in Performance, Muscle Activity, and Hand Kinematics for Different Mice
abstract
Computer mouse design can impact user comfort and performance. The effect of mouse design on gamers, who use a mouse for long hours and apply higher velocity movements than office workers, is uncertain. Professional (N = 29) and high-level (N = 19) gamers participated in this laboratory study and performed Fitts’ and gaming tasks (OverwatchTM) with different mice, and this analysis compared results from a light-weight (87 g) wireless mouse and a very light-weight (80 g) wireless mouse. There was little difference between the mice on muscle activity, hand motion, usability, or fatigue; however, professional gamers preferred the 80 g mouse. Professional gamers achieved higher levels of throughput and lower levels of path deviation than high-level gamers and their hand movement velocities and accelerations were significantly higher (p<0.05). The observed hand movement patterns used by professionals may be useful for instructing other gamers on techniques to improve performance.
Guangchuan Li, Mengcheng Wang, Federico Arippa, Alan Barr, David Rempel, Yue Liu 0005, Carisa Harris-Adamson
Int. J. Hum. Comput. Interact.6
2022 Sparse-Based Domain Adaptation Network for OCTA Image Super-Resolution Reconstruction
abstract
Retinal Optical Coherence Tomography Angiography (OCTA) with high-resolution is important for the quantification and analysis of retinal vasculature. However, the resolution of OCTA images is inversely proportional to the field of view at the same sampling frequency, which is not conducive to clinicians for analyzing larger vascular areas. In this paper, we propose a novel Sparse-based domain Adaptation Super-Resolution network (SASR) for the reconstruction of realistic [Formula: see text]/low-resolution (LR) OCTA images to high-resolution (HR) representations. To be more specific, we first perform a simple degradation of the [Formula: see text]/high-resolution (HR) image to obtain the synthetic LR image. An efficient registration method is then employed to register the synthetic LR with its corresponding [Formula: see text] image region within the [Formula: see text] image to obtain the cropped realistic LR image. We then propose a multi-level super-resolution model for the fully-supervised reconstruction of the synthetic data, guiding the reconstruction of the realistic LR images through a generative-adversarial strategy that allows the synthetic and realistic LR images to be unified in the feature domain. Finally, a novel sparse edge-aware loss is designed to dynamically optimize the vessel edge structure. Extensive experiments on two OCTA sets have shown that our method performs better than state-of-the-art super-resolution reconstruction methods. In addition, we have investigated the performance of the reconstruction results on retina structure segmentations, which further validate the effectiveness of our approach.
Huaying Hao, Dan Zhang 0026, Qifeng Yan, Jiong Zhang 0004, Yue Liu 0005, Yitian Zhao
IEEE J. Biomed. Health Informatics6
2022 Investigate the Neuro Mechanisms of Stereoscopic Visual Fatigue
abstract
Stereoscopic visual fatigue (SVF) due to prolonged immersion in the virtual environment can lead to negative user experience, thus hindering the development of virtual reality (VR) industry. Previous studies have focused on investigating the evaluation indicators associated with SVF, while few studies have been conducted to reveal the underlying neural mechanism, especially in VR applications. In this paper, a modified Go/NoGo paradigm was adopted to induce SVF in VR environment with Go trials for maintaining participants' attention and NoGo trials for investigating the neural effects under SVF. Random dot stereograms (RDSs) with 11 disparities were presented to evoke the depth-related visual evoked potentials (DVEPs) during 64-channel EEG recordings. EEG datasets collected from 15 participants in NoGo trials were selected to conduct individual processing and group analysis, in which the characteristics of the DVEPs components for various fatigue degrees were compared and independent components were clustered to explore the original cortex areas related to SVF. Point-by-point permutation statistics revealed that DVEPs sample points from 230 ms to 280 ms (component P2) in most brain areas changed significantly when SVF increased. Additionally, independent component analysis (ICA) identified that component P2 which originated from posterior cingulate cortex and precuneus, was associated statistically with SVF. We believe that SVF is rather a conscious status concerning the changes of self-awareness or self-location awareness than the performance reduction of retinal image processing. Moreover, we suggest that indicators representing higher conscious state may be a better indicator for SVF evaluation in VR environments.
Kang Yue, Mei Guo, Yue Liu 0005, Haochen Hu, Danli Wang
IEEE J. Biomed. Health Informatics3
2022 Navigation in virtual and real environment using brain computer interface: a progress report
abstract
A brain-computer interface (BCI) facilitates bypassing the peripheral nervous system and directly communicating with surrounding devices. Navigation technology using BCI has developed—from exploring the prototype paradigm in the virtual environment (VE) to accurately completing the locomotion intention of the operator in the form of a powered wheelchair or mobile robot in a real environment. This paper summarizes BCI navigation applications that have been used in both real and VEs in the past 20 years. Horizontal comparisons were conducted between various paradigms applied to BCI and their unique signal-processing methods. Owing to the shift in the control mode from synchronous to asynchronous, the development trend of navigation applications in the VE was also reviewed. The contrast between highlevel commands and low-level commands is introduced as the main line to review the two major applications of BCI navigation in real environments: mobile robots and unmanned aerial vehicles (UAVs). Finally, applications of BCI navigation to scenarios outside the laboratory; research challenges, including human factors in navigation application interaction design; and the feasibility of hybrid BCI for BCI navigation are discussed in detail.
Haochen Hu, Yue Liu 0005, Kang Yue, Yongtian Wang
Virtual Real. Intell. Hardw.2
2021 Investigating the Factors that Influence Technology Acceptance of an Educational Game Integrating Mixed Reality and Concept Maps
abstract
With the rapid development of mobile learning technologies as well as relevant hardware and software platforms, there is a bright prospect for mixed reality (MR) applying in the field of education. However, appropriate knowledge navigation and concept arrangement methods are urgent to diminish students' cognitive load in the virtual learning environment. In this preliminary study, concept maps were selected as scaffolding tools to help students navigate through the MR learning space, and an educational game prototype named MMRCM integrating MR and concept maps was developed to investigate students' technology acceptance about it. Considering the novelty of this new type of MR application, an extended version of the Technology Acceptance Model (TAM) was constructed with the MR game design elements and the concept map usefulness as external factors. A middle school physics experiment using MMRCM was conducted to help students learn the abstract concepts of friction. The evaluation results showed that the model's external factors have significant correlations with both perceived ease of use and perceived usefulness. The findings indicated explicit intention to use MMRCM, which implied that MR gaming and concept maps could be integrated as an effective instructional tool in science education.
Yu Liu 0081, Yue Liu 0005, Kang Yue
ICALT2
2021 A New Dataset and Recognition for Egocentric Microgesture Designed by Ergonomists
Guangchuan Li, Yue Liu 0005, Yongtian Wang
ICIG (2)2
2021 Single Scene Image Editing Based on Deep Intrinsic Decomposition
Hao Sha 0004, Yue Liu 0005, Chenguang Lu, Hengrun Chen, Yongtian Wang
ICIG (3)2
2021 A Paradigm to Enhance Motor Imagery through Immersive Virtual Reality with Visuo-Tactile Stimulus
abstract
Motor imagery brain-computer interfaces have a wide range of promising applications in medical, entertainment and home applications, however, such problems as few motor imagery brain-computer paradigms and weak EEG features need to be addressed. In order to explore good experimental paradigms to induce users to perform motor imagery tasks more effectively, provide more effective EEG features and save training time and increase system robustness, this paper designs a novel motor imagination brain-computer interaction paradigm, which combines virtual reality and haptic stimulation to provide synchronized visual-haptic feedback to the system and enhance the user's motor imagination ability. The experimental results show that there are significant differences between the traditional paradigms and the proposed paradigm in terms of event-related spectral perturbations, which are more obvious in the 10-20Hz and 25-30Hz in the proposed, and the embodiment scores in the proposed paradigm are also higher. In summary, the proposed paradigm can improve users' motor imagery ability by enhancing their embodiments.
Kang Yue, Haochen Hu, Yue Liu 0005
SMC4
2021 Tactile Perceptual Thresholds of Electrovibration in VR
abstract
Haptic sensation plays an important role in providing physical information to users in both real environments and virtual environments. To produce high-fidelity haptic feedback, various haptic devices and tactile rendering methods have been explored in myriad scenarios, and perception deviation between a virtual environment and a real environment has been investigated. However, the tactile sensitivity for touch perception in a virtual environment has not been fully studied; thus, the necessary guidance to design haptic feedback quantitatively for virtual reality systems is lacking. This paper aims to investigate users' tactile sensitivity and explore the perceptual thresholds when users are immersed in a virtual environment by utilizing electrovibration tactile feedback and by generating tactile stimuli with different waveform, frequency and amplitude characteristics. Hence, two psychophysical experiments were designed, and the experimental results were analyzed. We believe that the significance and potential of our study on tactile perceptual thresholds can promote future research that focuses on creating a favorable haptic experience for VR applications.
Yue Liu 0005
IEEE Trans. Vis. Comput. Graph.2
2021 Analysis of teenagers' preferences and concerns regarding HMDs in education
abstract
Virtual reality (VR) has become a powerful and promising tool for education, and numerous studies have investigated the application and effectiveness of VR education. However, few studies have focused on the expectations and concerns of teenagers regarding head-mounted displays (HMDs), which are used for this purpose. In this paper, we aim to explore the current problems and necessary advancements required in VR education based on a survey of 163 senior high school students who experience VR educational content for 1h. The usability and comfort of the HMD system, the physical and psychological effects on the students, and their preferences and concerns are investigated. The results show that HMDs increase students' interest, concentration, and enthusiasm for learning. However, isolated virtual environments make students feel nervous and afraid. The immersive environment also makes them worry about VR addiction and confusing the physical world with the virtual one. VR has great potential in the field of education, but the issue of safety needs to be considered in the future.
Jie Guo 0004, Dongdong Weng, Yue Liu 0005, Qiyong Chen, Yongtian Wang
Virtual Real. Intell. Hardw.3
2020 Exploring the Differences of Visual Discomfort Caused by Long-term Immersion between Virtual Environments and Physical Environments
abstract
To investigate the effects of visual discomfort caused by long-term immersing in virtual environments (VEs), we conducted a comparative study to evaluate users’ visual discomfort in an eight-hour working rhythm and compared the differences between the VEs and the physical environments. Twenty-seven participants performed four different visual tasks with a head-mounted display (HMD) for the VE condition and with a monitor for the physical condition. Their subjective visual discomfort and objective oculomotor indicators were measured to evaluate their visual performances. The results show that the subjective visual fatigue symptoms, the objective pupil size, and the relative accommodation response vary across time for the two conditions, in which VEs affects visual fatigue the most compared to the physical environments. The results also show that pupil size is negatively related to subjective visual fatigue, and the long-term work based on displays only influences the maximum accommodation response of participants. This work is a supplement to the necessary but insufficient-researched field of visual fatigue in long-term immersing in VEs, which should be valuable to researchers involved in the evaluation of visual fatigue using HMDs.
Jie Guo 0004, Dongdong Weng, Zhenliang Zhang 0002, Jiamin Ping, Yue Liu 0005, Yongtian Wang
VR6
2020 Implementation and Evaluation of Touch-based Interaction Using Electrovibration Haptic Feedback in Virtual Environments
abstract
Presentation of more haptic information to improve the users interactive capability is an important and challenging task in the virtual environment. As an emerging haptic feedback technology, electrovibration is an underlying solution to enhance the systems interactivity and improve user experience. However, such a solution is rarely studied in a virtual environment. In this work, we explore a new VR interaction method based on electrovibration technology with a touch screen and conduct evaluations about it. The key idea is to incorporate a set of manipulation gestures and three types of electrovibration in the VR interaction to help users acquire different kinds of tactile perception in the virtual manipulation. We present the evaluation in which we compare user performance measured first in a Fitts law task to evaluate different electrovibration types and then in a virtual office application to assess the interactive user interface. Our results show that the precision of interactions is significantly improved with the electrovibration haptic feedback. To the best of our knowledge, this is the first work to introduce electrovibration haptic feedback into the VR human-computer interaction and our work enlightens the potential of the electrovibration touchscreen-based interaction in virtual environments.
Yue Liu 0005, Dejiang Ye, Zhuoluo Ma
VR2
2020 Analysis on Mitigation of Visually Induced Motion Sickness by Applying Dynamical Blurring on a User's Retina
abstract
Visually induced motion sickness (MS) experienced in a 3D immersive virtual environment (VE) limits the widespread use of virtual reality (VR). This paper studies the effects of a saliency detection-based approach on the reduction of MS when the display on a user's retina is dynamic blurred. In the experiment, forty participants were exposed to a VR experience under a control condition without applying dynamic blurring, and an experimental condition applying dynamic blurring. The experimental results show that the participants under the experimental condition report a statistically significant reduction in the severity of MS symptoms on average during the VR experience compared to those under the control condition, which demonstrates that the proposed approach may alleviate visually induced MS in VR and enable users to remain in a VE for a longer period of time.
Guang-Yu Nie, Henry Been-Lirn Duh, Yue Liu 0005, Yongtian Wang
IEEE Trans. Vis. Comput. Graph.3
2019 Multi-Level Context Ultra-Aggregation for Stereo Matching
abstract
Exploiting multi-level context information to cost volume can improve the performance of learning-based stereo matching methods. In recent years, 3-D Convolution Neural Networks (3-D CNNs) show the advantages in regularizing cost volume but are limited by unary features learning in matching cost computation. However, existing methods only use features from plain convolution layers or a simple aggregation of multi-level features to calculate cost volume, which is insufficient because stereo matching requires discriminative features to identify corresponding pixels in rectified stereo image pairs. In this paper, we propose a unary features descriptor using multi-level context ultra-aggregation (MCUA), which encapsulates all convolutional features into a more discriminative representation by intra- and inter-level features combination. Specifically, a child module that takes low-resolution images as input captures larger context information; the larger context information from each layer is densely connected to the main branch of the network. MCUA makes good usage of multi-level features with richer context and performs the image-to-image prediction holistically. We introduce our MCUA scheme for cost volume calculation and test it on PSM-Net. We also evaluate our method on Scene Flow and KITTI 2012/2015 stereo datasets. Experimental results show that our method outperforms state-of-the-art methods by a notable margin and effectively improves the accuracy of stereo matching.
Guang-Yu Nie, Ming-Ming Cheng, Yun Liu 0011, Zhengfa Liang, Deng-Ping Fan, Yue Liu 0005, Yongtian Wang
CVPR6
2019 Implementation and Evaluation of Touch and Gesture Interaction Modalities for In-vehicle Infotainment Systems
Yue Liu 0005
ICIG (3)3
2019 High-Fidelity Grasping in Virtual Reality using a Glove-based System
abstract
This paper presents a design that jointly provides hand pose sensing, hand localization, and haptic feedback to facilitate real-time stable grasps in Virtual Reality (VR). The design is based on an easy-to-replicate glove-based system that can reliably perform (i) a high-fidelity hand pose sensing in real time through a network of 15 IMUs, and (ii) the hand localization using a Vive Tracker. The supported physics-based simulation in VR is capable of detecting collisions and contact points for virtual object manipulation, which drives the collision event to trigger the physical vibration motors on the glove to signal the user, providing a better realism inside virtual environments. A caging-based approach using collision geometry is integrated to determine whether a grasp is stable. In the experiment, we showcase successful grasps of virtual objects with large geometry variations. Comparing to the popular LeapMotion sensor, we demonstrate the proposed glove-based design yields a higher success rate in various tasks in VR. We hope such a glove-based system can simplify the data collection of human manipulations with VR.
Hangxin Liu, Zhenliang Zhang 0002, Xu Xie 0001, Yixin Zhu 0001, Yue Liu 0005, Yongtian Wang, Song-Chun Zhu
ICRA5
2019 Toward an Efficient Hybrid Interaction Paradigm for Object Manipulation in Optical See-Through Mixed Reality
abstract
Human-computer interaction (HCI) plays an important role in the near-field mixed reality, in which the hand-based interaction is one of the most widely-used interaction modes, especially in the applications based on optical see-through head-mounted displays (OST-HMDs). In this paper, such interaction modes as gesture-based interaction (GBI) and physics-based interaction (PBI) are developed to construct a mixed reality system to evaluate the advantages and disadvantages of different interaction modes. The ultimate goal is to find an efficient hybrid paradigm for mixed reality applications based on OST-HMDs to deal with the situations that a single interaction mode cannot handle. The results of the experiment, which compares GBI and PBI, show that PBI leads to a better performance of users regarding their work efficiency in the proposed two tasks. Some statistical tests, including T-test and one-way ANOVA, have also been adopted to prove that the difference regarding the efficiency between different interaction modes is significant. Experiments for combining both interaction modes are put forward in order to seek a good experience for manipulation, which proves that the partially-overlapping style would help to improve work efficiency for manipulation tasks. The experimental results of the proposed two hand-based interaction modes and their hybrid forms can provide some practical suggestions for the development of mixed reality systems based on OST-HMDs.
Zhenliang Zhang 0002, Dongdong Weng, Jie Guo 0004, Yue Liu 0005, Yongtian Wang
IROS4
2019 Mixed Reality Office System Based on Maslow's Hierarchy of Needs: Towards the Long-Term Immersion in Virtual Environments
abstract
In a mixed reality (MR) environment that combines the physical objects with the virtual environments, users' feelings are immersed in the virtual world, while their bodies remain in the physical world. Compared to the purely physical environments, such characteristic has led to some special needs for users' long-term immersion. However, the deficiency needs that we have to face for long-term immersion still need further research. In this paper, we apply the theory of Maslow's Hierarchy of Needs (MHN) to guide the design of MR systems for long-term immersion. Taking the normal biological rhythm of human beings as the basic unit (24 hours), we propose the fundamental needs for long-term immersion in VEs through combining the theory of MHN with the special needs of virtual reality (VR). In order to verify whether those needs can satisfy users' long-term immersion, we design an MR office system for basic operations based on the theory of MHN. A long-term exposure experiment (duration of 8 hours) is designed to evaluate those needs by comparing the results with a physical work environment after a short-term preliminary study. The physiological and psychological effects are tested in both two environments and the deficiency needs for short-term immersion and long-term immersion are also compared. The results showed that the design based on the theory of MHN can support users' long-term immersion, which means that it can be a guideline for long-term use of MR systems.
Jie Guo 0004, Dongdong Weng, Zhenliang Zhang 0002, Yue Liu 0005, Yongtian Wang, Henry Been-Lirn Duh
ISMAR5
2019 Evaluation of Maslows Hierarchy of Needs on Long-Term Use of HMDs - A Case Study of Office Environment
abstract
Long-term exposure to VR will become more and more important, but what we need for long term immersion to meet users fundamental needs is still under-researched. In this paper, we apply the theory of Maslows Hierarchy of Needs to guide the design of VR for longterm immersion based on the normal biological rhythm of human beings (24 hours). An office environment is designed to verify those needs. The efficiency, the physical and the psychological effects of this VR office system are tested. The results show that the VR office environment is as comfortable as the physical environment at short-term immersion and it can support users basic immersion. It means that the Maslows Hierarchy of Needs can be a guideline for long-term immersion.
Jie Guo 0004, Dongdong Weng, Zhenliang Zhang 0002, Yue Liu 0005, Yongtian Wang
VR4
2019 Real Time 3D Magnetic Field Visualization Based on Augmented Reality
abstract
In physics teaching, electromagnetism is one of the most difficult concepts for students to understand. This paper proposes a real time visualization method for 3-D magnetic field based on the augmented reality technology, which can not only visualize magnetic flux lines in real time, but also simulates the approximate sparse distribution of magnetic flux lines in space. An application utilizing this method is also presented. It permits leaners to freely and interactively move the magnets in 3-D space and to observe the magnetic flux lines in real time. As a result, the proposed method visualizes the invisible factors in 3-D magnetic field, with which students will have real-life reference when studying electromagnetic.
Yue Liu 0005, Yongtian Wang
VR2
2019 Exploring Stereovision-Based 3-D Scene Reconstruction for Augmented Reality
abstract
Three-dimensional (3-D) scene reconstruction is one of the key techniques in Augmented Reality (AR), which is related to the integration of image processing and display systems of complex information. Stereo matching is a computer vision based approach for 3-D scene reconstruction. In this paper, we explore an improved stereo matching network, SLED-Net, in which a Single Long Encoder-Decoder is proposed to replace the stacked hourglass network in PSM-Net for better contextual information learning. We compare SLED-Net to state-of-the-art methods recently published, and demonstrate its superior performance on Scene Flow and KITTI2015 test sets.
Guang-Yu Nie, Yun Liu 0011, Yongtian Wang, Yue Liu 0005
VR5
2019 Building AR-based Optical Experiment Applications in a VR Course
abstract
The demand for VR courses in universities is growing since VR technology is widely used in industry, entertainment, and education today. Traditional lectures and exercises have problems in motivating and engaging students, especially non-CS-majors. We design a VR course with a teamwork project assignment for Opto-Electronic Engineering undergraduates, providing them with project-based learning (PBL) experience. The task requires three students to group a team to build an AR-based optical experiment application over the course, aiming to develop students' practical engineering ability. The process of the project consists of three main stages: preparation, designing and implementing. We also evaluate students' work from different aspects and survey to analyze the students' attitude toward the project.
Huan Wei, Yue Liu 0005, Yongtian Wang
VR2
2019 Symmetrical Reality: Toward a Unified Framework for Physical and Virtual Reality
abstract
In this paper, we review the background of physical reality, virtual reality, and some traditional mixed forms of them. Based on the current knowledge, we propose a new unified concept called symmetrical reality to describe the physical and virtual world in a unified perspective. Under the framework of symmetrical reality, the traditional virtual reality, augmented reality, inverse virtual reality, and inverse augmented reality can be interpreted using a unified presentation. We analyze the characteristics of symmetrical reality from two different observation locations (i.e., from the physical world and from the virtual world), where all other forms of physical and virtual reality can be treated as special cases of symmetrical reality.
Zhenliang Zhang 0002, Dongdong Weng, Yue Liu 0005, Yongtian Wang
VR4
2019 Analyzing the Usability of Gesture Interaction in Virtual Driving System
abstract
In this study, an experiment is presented aiming at verifying the applicability of gesture interaction in the virtual driving environment. 30 participants are recruited to perform the secondary tasks with gesture and touch interaction. The task completion rate and reaction time of two interaction modalities under different road conditions are adopted as evaluation indexes. In addition, visual attention, NASA-TLX, and subjective questionnaire are collected as evaluation factors for fuzzy comprehensive evaluation based on entropy to evaluate the usability gestures. The research results show that gesture interaction not only shows excellence in safety, but also favors more than 90% of users.
Yue Liu 0005, Yongtian Wang
VR2
2019 Effect of haptic feedback on a virtual lab about friction
abstract
With the increase in recent years of the utilization of multimedia devices in education, new haptic devices for education have been gradually adopted and developed. As compared with visual and auditory channels, the development of applications with a haptic channel is still in the initial stages. For example, it is unclear how force feedback influences an instructional effect of an educational application and the subjective feeling of users. In this study, we designed an educational application with a haptic device (Haply) to explore the effects of force feedback on selflearning.Subjects in an experiment group used a designed application to study friction by themselves using force feedback, whereas subjects in a control group studied the same knowledge without force feedback. A post-test and questionnaire were designed to assess the learning outcomes. The experimental result indicates that force feedback is beneficial to an educational application, and using a haptic device can improve the effect of the application and motivate students.
Zhuoluo Ma, Yue Liu 0005
Virtual Real. Intell. Hardw.2
2018 A Mobile Augmented Reality System For Illumination Consistency With Fisheye Camera Based On Client-Server Model
abstract
This paper proposes a client-server based mobile augmented reality system for illumination consistency between the virtual objects and the real world. The server system adopts a fisheye camera to capture real world illumination and transmits the light source information to the client system, while the client system uses the received information to render the virtual objects in the augmented reality application. The proposed method possesses the virtue of reducing the calculation cost of the clients, which is especially important for mobile devices. An evaluation method of AR image's illumination consistency has also been proposed based on analyzing the distance of the histogram of virtual and real parts of AR image with different channels.
Wankui Liu, Yue Liu 0005, Hua Huang 0001
CASA2
2018 Stereo Generation from a Single Image Using Deep Residual Network
abstract
In this paper, we propose a framework to generate stereoscopic content from a single image using the relative depth label predicted from deep residual network. Specifically, our framework first obtains a coarse relative depth label from the network and refines it to painting depth by sampling and interpolation, then an unsupervised clustering algorithm is employed to separate pixels of different depths into different layers to generate stereoscopic images. Experimental results with good visual effects demonstrate that the proposed method can be generally applied in both outdoor and indoor scenes. Meanwhile the quantitative results on relative depth estimation from a single image are comparable to state-of-the-art. Further experiments show the application possibility of our method in VR and panorama.
Tianteng Bi, Yue Liu 0005, Yongtian Wang
ICIP3
2018 Inverse Virtual Reality: Intelligence-Driven Mutually Mirrored World
abstract
Since artificial intelligence has been integrated into virtual reality, a new branch of virtual reality, which is called inverse virtual reality (IVR), is created. A typical IVR system contains both the intelligence-driven virtual reality and the physical reality, thus constructing an intelligence-driven mutually mirrored world. We propose the concept of IVR, and describe the details about the definition, structure and implementation of a typical IVR system. The parallel living environment is proposed as a typical application of IVR, which reveals that IVR has a significant potential to extend the human living environment.
Zhenliang Zhang 0002, Benyang Cao, Jie Guo 0004, Dongdong Weng, Yue Liu 0005, Yongtian Wang
VR5
2018 Evaluation of Hand-Based Interaction for Near-Field Mixed Reality with Optical See-Through Head-Mounted Displays
abstract
Hand-based interaction is one of the most widely-used interaction modes in the applications based on optical see-through head-mounted displays (OST-HMDs). In this paper, such interaction modes as gesture-based interaction (GBI) and physics-based interaction (PBI) are developed to construct a mixed reality system to evaluate the advantages and disadvantages of different interaction modes for near-field mixed reality. The experimental results show that PBI leads to a better performance of users regarding their work efficiency in the proposed tasks. The statistical analysis of T-test has been adopted to prove that the difference of efficiency between different interaction modes is significant.
Zhenliang Zhang 0002, Benyang Cao, Dongdong Weng, Yue Liu 0005, Yongtian Wang, Hua Huang 0001
VR4
2018 Physics-Inspired Input Method for Near-Field Mixed Reality Applications Using Latent Active Correction
abstract
Calibration accuracy is one of the most important factors to affect the user experience in mixed reality applications. For a typical mixed reality system built with the optical see-through head-mounted display (OST-HMD), a key problem is how to guarantee the accuracy of hand-eye coordination by decreasing the instability of the eye and the HMD in long-term use. In this paper, we propose a real-time latent active correction (LAC) algorithm to decrease hand-eye calibration errors accumulated over time. Experimental results show that we can successfully use the LAC algorithm to physics-inspired virtual input methods.
Zhenliang Zhang 0002, Dongdong Weng, Yue Liu 0005, Yongtian Wang
VR4
2018 Simultaneous Trajectory Association and Clustering for Motion Segmentation
abstract
Trajectory association and clustering are two key problems in motion analysis. While association links the points of interest to form trajectories, clustering discovers motion patterns of these trajectories and group them into clusters. Despite mutually related, the two problems have been typically studied separately in the literature. In this letter, we formulate them as a unified optimization problem and take the advantage of high-order information to capture the interrelations for Motion Segmentation by Trajectory Association and Clustering (MSTAC). To solve this unified problem, we propose an alternating optimization strategy to improve the association and clustering in each iteration. Specifically, a tensor-based multidimensional assignment method with high-order motion context information is proposed for trajectory association; and a minimum cost multicut-based trajectory clustering method is introduced for trajectory clustering. While the association process provides incomplete trajectories to clustering, the clustering method presents high-order context information to improve the performance of association; and thus they benefit from each other. Experiments on the Hopkins 155 dataset and a realistic airport sequence demonstrate that the proposed MSTAC framework obtains high accuracy on both trajectory association and clustering.
Yuxi Wang 0002, Yue Liu 0005, Erik Blasch, Haibin Ling
IEEE Signal Process. Lett.2
2017 Effects of using HMDs on visual fatigue in virtual environments
abstract
There are few negative effects to make people discomfort using virtual reality systems. In this paper, we investigated the effects of visual fatigue when wearing head-mounted displays (HMD) and compared the results with those from the smartphones. Forty subjects were recruited and divided into two different groups. The visual fatigue scale was measured to assess the subjects' performance. The results indicated that visual fatigue caused by the conflict of focal distance and vergence distance was less severe than visual fatigue caused by long-term focus without accommodation.
Jie Guo 0004, Dongdong Weng, Henry Been-Lirn Duh, Yue Liu 0005, Yongtian Wang
VR4
2017 Evaluation of labelling layout methods in augmented reality
abstract
View management techniques are commonly used for labelling of objects in augmented reality environments. Combining with image analysis, search space and adaptive representations, they can be utilized to achieve desired labelling tasks. However, the evaluation of different search space methods on labelling are still an open problem. In this paper, we propose an image analysis based view management method, which first adopts the image processing to superimpose 2D labels to the specific object. We then conduct three search space methods to an augmented reality scenario. Without the requirements of setting rules and constraints for occlusion among the labels, the results of three search space methods are evaluated by using objective analysis of related parameters. The evaluation results indicate that different search space methods could generate different time costs and occlusion, thereby affecting the final labelling effects.
Yue Liu 0005, Yongtian Wang
VR2
2017 RIDE: Region-induced data enhancement method for dynamic calibration of optical see-through head-mounted displays
abstract
The most commonly used single point active alignment method (SPAAM) is based on a static pinhole camera model, in which it is assumed that both the eye and the HMD are fixed. This leads to a limitation for calibration precision. In this work, we propose a dynamic pinhole camera model according to the fact that the human eye would experience an obvious displacement over the whole calibration process. Based on such a camera model, we propose a new calibration data acquisition method called the region-induced data enhancement (RIDE) to revise the calibration data. The experimental results prove that the proposed dynamic model performs better than the traditional static model in actual calibration.
Zhenliang Zhang 0002, Dongdong Weng, Yue Liu 0005, Yongtian Wang, Xinjun Zhao
VR3
2017 Fast hand posture classification using depth features extracted from random line segments
Weizhi Nai, Yue Liu 0005, David Rempel, Yongtian Wang
Pattern Recognit.2
2016 Visual tracking via sparsity pattern learning
abstract
Recently sparse representation has been applied to visual tracking by modeling the target appearance using a sparse approximation over the template set. However, this approach is limited by the high computational cost of the ℓ1-norm minimization involved, which also impacts on the amount of particle samples that we can have. This paper introduces a basic constraint on the self-representation of the target set. The sparsity pattern in the self-representation allows us to recover the “sparse coefficients” of the candidate samples by some small-scale ℓ2-norm minimization; this results in a fast tracking algorithm. It also leads to a principled dictionary update mechanism which is crucial for good performance. Experiments on a recently released benchmark with 50 challenging video sequences show significant runtime efficiency and tracking accuracy achieved by the proposed algorithm.
Yuxi Wang 0002, Yue Liu 0005, Zhuwen Li, Loong Fah Cheong, Haibin Ling
ICPR2
2016 A tour guiding system of historical relics based on augmented reality
abstract
Yuanmingyuan is a relic park and only few cultural relics are left due to the looting and burning down in history, which makes that most of the scenic spots of the park look boring. To address such issue, a game-based guidance system for Yuanmingyuan and a time travel game called MAGIC-EYES has been proposed with Augmented Reality technology. Six interactive modes are designed in the proposed system to guide tourists to visit the specified place. The evaluation results of a pilot study shows that the proposed guidance system has significantly improved the tourist experiences.
Dongdong Weng, Yue Liu 0005, Yongtian Wang
VR3
2015 Deformable 3D Fusion: From Partial Dynamic 3D Observations to Complete 4D Models
abstract
Capturing the 3D motion of dynamic, non-rigid objects has attracted significant attention in computer vision. Existing methods typically require either complete 3D volumetric observations, or a shape template. In this paper, we introduce a template-less 4D reconstruction method that incrementally fuses highly-incomplete 3D observations of a deforming object, and generates a complete, temporally-coherent shape representation of the object. To this end, we design an online algorithm that alternatively registers new observations to the current model estimate and updates the model. We demonstrate the effectiveness of our approach at reconstructing non-rigidly moving objects from highly-incomplete measurements on both sequences of partial 3D point clouds and Kinect videos.
WeiPeng Xu, Mathieu Salzmann, Yongtian Wang, Yue Liu 0005
ICCV4
2015 Context-Aware Based Mobile Augmented Reality Browser and its Optimization Design
Yue Liu 0005, Yongtian Wang
ICIG (2)2
2015 Design of a Simulated Michelson Interferometer for Education Based on Virtual Reality
Hongling Sun, Yue Liu 0005
ICIG (2)3
2015 Omnidirectional-view three-dimensional displays using multiple mini-projectors
abstract
We have developed omnidirectional-view 3D displays using multiple mini-projectors. Three types of synchronization structure are developed to ensure the accurate synchronization between the projectors and the rotating screen. Therefore, low-cost and low-speed display devices can be used to realize natural-looking three-dimensional (3D) scene with full color and high resolution.
Qiudong Zhu, Dongdong Weng, Yue Liu 0005, Yongtian Wang
VCIP4
2014 Nonrigid Surface Registration and Completion from RGBD Images
WeiPeng Xu, Mathieu Salzmann, Yongtian Wang, Yue Liu 0005
ECCV (2)4
2013 Real-time keystone correction for hand-held projectors with an RGBD camera
abstract
This paper introduces a novel and simple approach to realtime continuous keystone correction for hand-held projectors. An RGBD camera is attached to the projector to form a projector-RGBD-camera system. The system is first calibrated in an offline stage. At run-time, we then estimate the relative pose between the projector and the screen using the RGBD camera, which lets us correct the keystone distortion by warping the projected image accordingly. Experimental results show that our method outperforms existing techniques in terms of both accuracy and efficiency.
WeiPeng Xu, Yongtian Wang, Yue Liu 0005, Dongdong Weng, Mengwen Tan, Mathieu Salzmann
ICIP3
2013 Outdoor scenes identification on mobile device by integrating vision and inertial sensors
abstract
This paper addresses the identification of large scale outdoor scenes on smart phone by fusing outputs of inertial sensors and computer vision techniques. The main contributions can be summarized as follows: Firstly, we propose an overlap region divide (ORD) method to plot image position area, which is fast enough to find the nearest visiting area and can also reduce the search range compared with the traditional approaches. Secondly, the vocabulary tree based approach is improved by introducing fast geometric consistency constraints (FGCC). Our method involves no operation in the high-dimensional feature space and does not assume a global transform between a pair of images. Thus, it substantially reduces the computational complexity and memory usage, which makes the city scale image recognition feasible on the smartphone. Experiments on a collected database including 0.16 million images show that the proposed method demonstrates excellent identification performance, while maintaining the average identification time of less than 1s.
Zhenwen Gui, Yongtian Wang, Yue Liu 0005, Jing Chen 0018
IWCMC3
2013 Robust structure from motion with affine camera via low-rank matrix recovery
abstract
Abstract We present a novel approach to structure from motion that can deal with missing data and outliers with an affine camera. We model the corruptions as sparse error. Therefore the structure from motion problem is reduced to the problem of recovering a low-rank matrix from corrupted observations. We first decompose the matrix of trajectories of features into low-rank and sparse components by nuclear-norm and ℓ 1-norm minimization, and then obtain the motion and structure from the low-rank components by the classical factorization method. Unlike pervious methods, which have some drawbacks such as depending on the initial value selection and being sensitive to the large magnitude errors, our method uses a convex optimization technique that is guaranteed to recover the low-rank matrix from highly corrupted and incomplete observations. Experimental results demonstrate that the proposed approach is more efficient and robust to large-scale outliers.
Lun Wu, Yongtian Wang, Yue Liu 0005, Yuxi Wang 0002
Sci. China Inf. Sci.3
2013 Panoramic Gaussian Mixture Model and large-scale range background substraction method for PTZ camera-based surveillance systems
Kang Xue, Yue Liu 0005, Gbolabo Ogunmakin, Jing Chen 0018, Jiangen Zhang
Mach. Vis. Appl.2
2012 A modified KLT multiple objects tracking framework based on global segmentation and adaptive template
Kang Xue, Patricio A. Vela, Yue Liu 0005, Yongtian Wang
ICPR3
2011 Application of Pen-Based Planar Haptic Interface in Physics Education
abstract
This paper proposes a pen-based interaction system with 2D co-located haptic and visual feedback for physics education. The system combines a 3DOF pen-based planar haptic device and simulated 2D physical world based on physical engine Box 2D and 2D game engine HGE. The inputs of haptic device are translation and rotation of pen tip and the outputs are translational and rotational force displayed at pen tip. In addition, visual display and haptic display are coincident and well integrated. With this system, user can create and design simulated physical world by drawing, selecting, moving or rotating objects on screen and setting their physical properties in a natural and intuitive way. This system provides an entertaining and cartoony tool for designing and carrying out physics experiments, and will greatly promote students' learning interest and creativity.
Liping Lin, Yongtian Wang, Yue Liu 0005, Makoto Sato
CAD/Graphics3
2011 Segmentation Based on Routing Image Algorithms
abstract
This paper presents a novel image segmentation method in which energy function is based on global region information while not only on edge information. Image segmentation can be viewed as a routing problem. In order to obtain the optimal segmentation, the Shortest Path Faster Algorithm (SPFA) is used to optimize the discrete grid energy function. As the commonly used Live-Wire algorithm is easy to obtain mistake segmentation when the strong edges and the weak edges are close to each other, the interactive segmentation method is proposed for the precise boundaries estimation. The developed method has been tested on both clinical medical images and natural scene images. It can be seen that the developed method is very fast and effective, and can obtain good segmentation results.
Hongzhe Yang, Jian Yang 0009, Yongtian Wang, Yue Liu 0005
ICIG4
2011 PTZ camera-based adaptive panoramic and multi-layered background model
abstract
In this paper, we present a novel approach for constructing an adaptive panoramic and multi-layered background model for Pan-tilt-zoom (PTZ) camera that provides fast registration of the observed frame and localizes the foreground targets with arbitrary camera position and scale (optical zoom). Our method consists of two stages. (1) An adaptive panoramic background mixture model is generated off-line for foreground detection. (2) A layered correspondence is generated off-line from frames captured at different optical zoom values of the camera, and a correspondence propagation method is used to register the observed frame with the panoramic background online. We demonstrate the advantages of the proposed adaptive panoramic and multi-layered background model within wide field of view (FOV) and over large scale range.
Kang Xue, Gbolabo Ogunmakin, Yue Liu 0005, Patricio A. Vela, Yongtian Wang
ICIP3
2011 "Soul Hunter": A novel augmented reality application in theme parks
abstract
This paper introduces a novel augmented reality shooting game named “Soul Hunter”, which has been successfully operating in a theme park in China. Soul Hunter adopts an innovative infrared marker scheme to build a mobile augmented reality application in a wide area. It is an extension of the traditional first person game, in which a player is able to fight with virtual ghost through a gunlike device in real environment. This paper describes the challenges of applying augmented reality in theme parks and shares some experiences in solving the problems encountered in practical applications.
Dongdong Weng, WeiPeng Xu, Dong Li 0013, Yongtian Wang, Yue Liu 0005
ISMAR5
2011 Sensor fusion based head pose tracking for lightweight flight cockpit systems
Yongtian Wang, Yue Liu 0005
Multim. Tools Appl.3
2010 Augmented reality registration algorithm based on nature feature recognition
Jing Chen 0018, Yongtian Wang, Junwei Guo, Jingdun Lin, Kang Xue, Yue Liu 0005
Sci. China Inf. Sci.7
2009 Key Issues of Wide-Area Tracking System for Multi-user Augmented Reality Adventure Game
abstract
Augmented reality (AR) Adventure Game is a wide-area indoor AR application for multiple users in Guangdong science center. Tracking the pose of users’ head in wide area is crucial for alignment between virtual and real scene. This paper studies the key issues of wide-area indoor tracking system. Different from previous inside-looking-out vision-based tracking, coded infrared (IR) markers are installed both on walls and ceiling in the proposed system. An automatic method is designed to calibrate the transformation between tracking camera and scene camera. The linear algorithm used for pose estimation is presented and problems of registration in rendering engine are also discussed. Experiments are conducted and applications in the AR Adventure Game prove that the proposed tracking system can provide precise and stable registration in actual systems.
Yetao Huang, Dongdong Weng, Yue Liu 0005, Yongtian Wang
ICIG3
2009 A Remote Control System Based on Real-Time Image Processing
abstract
A novel human-computer interaction (HCI) system based on real-time image processing is proposed in this paper. With the help of infrared tracking technology, the proposed system achieves real-time processing and stable operation on an ADSP-BF533 hardware platform. Compared with the conventional remote control methods, the proposed system enables a user to control the cursor on the screen of a TV by targeting it, and provides users with new experiences of remote control. To realize such a new remote control system, cross-ratio invariant, which is an important characteristic of the projective transformation, is also studied. Experimental results show the potential of the proposed system in TV remote control.
Yongtian Wang, Yue Liu 0005, Dongdong Weng, Xiaoming Hu 0001
ICIG3
2009 Sensor Fusion for Vision-Based Indoor Head Pose Tracking
abstract
Accurate head pose tracking is a key issue for indoor augmented reality systems. This paper proposes a novel approach to track head pose of indoor users using sensor fusion. The proposed approach utilizes a track-to-track fusion framework composed of extended Kalman filters and fusion filter to fuse the poses from the two complementary tracking modes of inside-out tracking (IOT) and outside-in tracking (OIT). A vision-based head tracker is constructed to verify our approach. Primary experimental results show that the tracker is capable of achieving more accurate and stable pose than the single tracking mode of IOT or OIT, which validates the usefulness of the proposed sensor fusion approach.
Yongtian Wang, Yue Liu 0005
ICIG3
2009 GPU Based Real-time Correction for Optical Distortions in Head-Mounted Displays
abstract
This paper presents a GPU-based real-time method to correct optical distortions in head-mounted displays (HMDs). The HMD to be corrected is a lightweight and wide field-of-view HMD system with free-form-surface (FFS) prism, in which the image distortion is not rectilinear and centrosymmetric. A special predistortion model is constructed to correct the distortion of the HMD. Although the distortion correction can be performed with an extensional optics system, the system will be too expensive and additional weight will be imposed to the HMD. With the method presented in this paper, each pixel in the original image is remapped to a new position with GPU and forms a predistortion image. The remapping process is based on a prior formatted distortion map, which is similar to the normal map used in the bump mapping process. The distortion map is an RGBA image that corresponds to the X and Y coordinates of a pixel offset from the original image. The remapping process accomplished by GPU is very fast. The performance of the proposed method is analyzed and validated via a demonstration system.
Dongdong Weng, Yongtian Wang, Yue Liu 0005
ICIG3
2009 Adaptive Real-Time Labeling and Recognition of Multiple Infrared Markers Using FPGA
abstract
In this paper, we propose a real-time and adaptive method for labeling and recognition of multiple infrared markers. A single-frame based iteration process is developed to obtain the suitable threshold. A sliding window is proposed to propagate the preliminary labels and to reduce the capacity of equivalent tables handling. The labeling and recognition are realized by merging the results of domains with label collisions. Multi-stage pipelines are developed to implement all the operations including smoothing filter, adaptive threshold, preliminary labeling and recognition. Experimental results show that the proposed method can label and recognize multiple infrared target markers with a latency of 259 ns and with an accuracy of sub-pixel. The proposed method can be applied in applications that require real-time performance such as surgery navigation, intelligent control and visual measurement.
Yue Liu 0005, Yongtian Wang
ICIG2
2009 Marker-less registration based on template tracking for augmented reality
Yongtian Wang, Yue Liu 0005, Caiming Xiong
Multim. Tools Appl.3
2009 Novel Approach for 3-D Reconstruction of Coronary Arteries From Two Uncalibrated Angiographic Images
abstract
Three-dimensional reconstruction of vessels from digital X-ray angiographic images is a powerful technique that compensates for limitations in angiography. It can provide physicians with the ability to accurately inspect the complex arterial network and to quantitatively assess disease induced vascular alterations in three dimensions. In this paper, both the projection principle of single view angiography and mathematical modeling of two view angiographies are studied in detail. The movement of the table, which commonly occurs during clinical practice, complicates the reconstruction process. On the basis of the pinhole camera model and existing optimization methods, an algorithm is developed for 3-D reconstruction of coronary arteries from two uncalibrated monoplane angiographic images. A simple and effective perspective projection model is proposed for the 3-D reconstruction of coronary arteries. A nonlinear optimization method is employed for refinement of the 3-D structure of the vessel skeletons, which takes the influence of table movement into consideration. An accurate model is suggested for the calculation of contour points of the vascular surface, which fully utilizes the information in the two projections. In our experiments with phantom and patient angiograms, the vessel centerlines are reconstructed in 3-D space with a mean positional accuracy of 0.665 mm and with a mean back projection error of 0.259 mm. This shows that the algorithm put forward in this paper is very effective and robust.
Jian Yang 0009, Yongtian Wang, Yue Liu 0005, Songyuan Tang, Wufan Chen
IEEE Trans. Image Process.3
2007 Study on an Indoor Tracking System Based on Primary and Assistant Infrared Markers
abstract
An indoor tracking system based on primary and assistant infrared markers is presented in this paper. The system can track the user's head in a large area with high stability and accuracy. And the price of the system is very low. The novel assistant infrared markers with particular spatial characteristic and the primary infrared markers with particular spatio-temporal characteristic are proposed, which can avoid the synchronization between infrared markers and user system. Various numbers of users can be supported by the proposed system and experimental result shows the effectiveness and robustness of the system.
Dongdong Weng, Yue Liu 0005, Yongtian Wang, Lun Wu
CAD/Graphics2
2007 Fiducial Marker Based on Projective Invariant for Augmented Reality
Yongtian Wang, Yue Liu 0005
J. Comput. Sci. Technol.3
2005 An Improved Colored-Marker Based Registration Method for AR Applications
Xiaowei Li 0004, Yue Liu 0005, Yongtian Wang, Dayuan Yan, Dongdong Weng
ICCSA (3)2
2005 Autocalibration of an Electronic Compass for Augmented Reality
abstract
Electronic compass is often used to provide the absolute heading reference for tracking the user's head and hands in virtual reality (VR) and augmented reality (AR), especially for outdoor AR applications. However, compass is vulnerable to environment magnetism disturbance. Existing compass calibration methods require complex steps and true heading reference which is often impossible to be obtained in outdoor AR applications, and is useful only when compass is in horizontal plane. An autocalibration method without the need of heading reference and redundant sensors is proposed in This work. First the compass error model based on physical principle is presented, then the algorithm to calculate the compensation coefficients with a set of sample measurements of the sensors in the compass is described. Because the influence of the environmental disturbance has been effectively compensated, the calibrated compass can provided accurate heading even when it is under large tilt attitude.
Xiaoming Hu 0001, Yue Liu 0005, Yongtian Wang, Yanling Hu, Dayuan Yan
ISMAR2