Yongtian Wang

dblp:00/852 · DBLP profile ↗
← Back
144ranked-venue papers
9as first author
41since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 73 · 1 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 45 · 8 first-author · 20 since 2021Artificial intelligence and machine learning · 32 · 3 since 2021Human-computer interaction and ubiquitous computing · 20 · 1 since 2021Systems, architecture and hardware · 4Computer networks · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Artificial intelligence for virtual reality: a review
Lili Wang 0006, Yebin Liu, Miao Wang 0004, Xubo Yang, Lan Xu 0003, Zhangyao Tan, Runze Fan, Hongwen Zhang 0001, Yijian Wen, Haozhong Yang, Jian Wu 0033, Jiahui Fan, Hui Wang 0045, Qixuan Zhang, Yongtian Wang, Qinping Zhao
Sci. China Inf. Sci.20
2026 DiffuST: A Latent Diffusion Model for Spatial Transcriptomics Denoising
abstract
Spatial transcriptomics technologies have enabled comprehensive measurements of gene expression profiles while retaining spatial information, with most platforms also providing matched pathology images. However, noise resulting from low RNA capture efficiency and experimental steps needed to keep spatial information may corrupt the biological signals and obstruct analyses. Here, we develop a latent diffusion model DiffuST to denoise spatial transcriptomics. DiffuST employs a graph autoencoder and a pre-trained model to extract different-scale features from spatial information and pathology images. Then, a latent diffusion model is leveraged to map different scales of features to the same space for denoising. The evaluation based on various spatial transcriptomics datasets showed the superiority of DiffuST over existing denoising methods. Furthermore, the results demonstrated that DiffuST can enhance downstream analysis of spatial transcriptomics and yield significant biological insights.
Shaoqing Jiao, Dazhi Lu, Tao Wang 0082, Yongtian Wang, Yunwei Dong, Jiajie Peng
IEEE Trans. Comput. Biol. Bioinform.5
2026 Double-Decomposition Motion Tracking of Intraoperative 3D Structures via Cross-Spatio-Temporal Semantics Alignment
abstract
3D motion tracking in X-ray image-guided operations using pre- and intra-operative image registration has recently gained attention. However, due to pre- and intra-operative acquisitions exist spatio-temporal misalignment (i.e., limited 3D prior versus continuous 2D images) and distinct respiratory phase difference, recent methods still struggle to accurately estimate 3D dynamic structures from X-ray images. To overcome these issues, we propose a novel double-decomposition tracking (DD-Track) framework that aligns with multi-organ motion characteristics via two alignment pipes: 1) Temporal alignment aims to compensate in-plane respiratory phases difference between the projection of static 3D prior and continuous X-ray images. A dual-excitation mechanism in the image and frequency domains is proposed to extract discriminate motion features while suppressing irrelevant background information. 2) Spatial alignment subsequently integrates the extracted 2D motion features into the cross-modal registration process to accurately warp the 3D prior. Further, we decompose the motion tracking into the common trajectory and organ-specific deformation to align with the multi-organ motion nature, avoiding excessive organ stretching for sliding compensation. Comprehensive quantitative and qualitative experiments on simulated and clinical multi-organ datasets demonstrate that DD-Track outperforms state-of-the-art methods, and we also validate its generalization for tracking intra-organ lesions on simulated data.
Haixiao Geng, Jingfan Fan, Danni Ai, Deqiang Xiao, Tianyu Fu 0003, Hong Song 0003, Feng Duan 0001, Yongtian Wang, Jian Yang 0009
IEEE Trans. Medical Imaging10
2026 PortInput: Enabling Always-Available Micro-Gesture Input With Pressure Array Sensor
abstract
Micro-gestures provide a natural, efficient, and privacy-preserving input modality; however, existing techniques often depend on environmental conditions, limiting their robustness and applicability in real-world settings. In this work, we present three portable prototypes- - finger-cot, finger-worn, and surface-based-that integrate compact pressure array sensors to support environment-independent micro-gesture interaction. We further propose a deep learning-based recognition model that accurately classifies 14 micro-gestures by analyzing temporal pressure patterns. Building upon these components, we introduce PortInput, a real-time interactive system that enables robust micro-gesture tracking and detection. We conducted two user studies with augmented reality (AR) head-mounted displays (HMDs). The first study evaluates input performance under both sitting and walking conditions, while the second compares PortInput with a commercial pressure-based ring device. The results show that PortInput improves usability and user experience, achieves comparable accuracy, and enables faster input with lower perceived workload. Overall, PortInput have potential to offer efficient, robust, and comfortable input across diverse application scenarios-ranging from AR/Virtual Reality (VR) headsets to smart homes and in-car systems-even in noisy or cluttered environments. This work provides a foundation for integrating pressure array sensors into ring-based or other portable devices, advancing always-available micro-gesture interaction for ubiquitous computing environments.
Henry Been-Lirn Duh, Mingwei Hu, Yue Liu 0005, Yongtian Wang
IEEE Trans. Vis. Comput. Graph.7
2026 Augmented reality surgical navigation: Clinical applications, key technologies, and future directions
abstract
Surgical navigation has evolved significantly through advances in augmented reality, virtual reality, and mixed reality, improving precision and safety across many clinical applications, including neurosurgery, maxillofacial, spinal, and arthroplasty procedures. By integrating preoperative imaging with real-time intraoperative data, these systems provide dynamic guidance, reduce radiation exposure, and minimize tissue damage. Key challenges persist, including intraoperative registration accuracy, flexible tissue deformation, respiratory compensation, and real-time imaging quality. Emerging solutions include artificial intelligence-driven segmentation, deformation-field modeling, and hybrid registration techniques. Future developments will include lightweight, portable systems, improved non-rigid registration algorithms, and greater clinical adoption. Despite advances in rigid-tissue applications, soft-tissue navigation requires additional innovation to address motion variability and registration reliability, ultimately advancing minimally invasive surgery and precision medicine.
Jingfan Fan, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Long Shao, Tao Chen 0022, Hong Song 0003, Yongtian Wang, Jian Yang 0009
Virtual Real. Intell. Hardw.11
2025 Cross-Media Color Appearance Reproduction in Optical See-Through Augmented Reality
abstract
In optical see-through (OST) augmented reality (AR), displayed colors blend with the real-world scene, affecting perceived color. Studies showed that AR's color appearance depends not only on the additive chromaticity of the display and the real scene but also on ambient illumination. However, these studies often overlook changes in observer's adaptation state under varying illumination. To address this, a series of color-matching experiments between AR and display devices was conducted in an immersive lighting environment. The first experiment under a D65 illuminant found a slightly higher correlated color temperature (CCT) level of the internal white point in OST AR than in the display. Considering media's white point differences, this study proposes a three-step chromatic adaptation transform (CAT) framework to improve color appearance reproduction accuracy in AR. The second experiment used three Planckian radiators varying in CCT and two offPlanckian colorful ones to validate the proposed CAT under varying illumination, indicating that AR reached a more complete adaptation state than the display, especially under high luminance. A third validation experiment had observers rate color differences between reference and reproduced stimuli in AR. Results showed the effectiveness of our three-step CAT for cross-media color reproduction of OST AR under diverse illumination conditions.
Jiahong Luo, Shining Ma, Yue Liu 0005, Yongtian Wang
ISMAR4
2025 BioRAGent: natural language biomedical querying with retrieval-augmented multiagent systems
abstract
Understanding the roles of genes, phenotypes, and diseases is crucial for advancing biomedical research. However, efficient and accessible retrieval of biomedical knowledge remains a challenge due to the complexity of the relevant data. We introduce BioRAGent, an intelligent biomedical assistant that combines Tool-augmented retrieval-augmented generation (RAG) with a multiagent system. Leveraging the ability of large language models, BioRAGent facilitates natural language queries about genes, phenotypes, diseases, and their interrelationships. BioRAGent employs three specialized agents: Guide (query optimization), Retriever (data retrieval), and Reviewer (answer validation) to access authoritative biomedical databases and to generate accurate responses. We evaluate the performance of BioRAGent on a benchmark of eleven single-hop and three multi-hop tasks, demonstrating superior results compared with state-of-the-art models. User evaluations highlight the practicality and robust user experience of BioRAGent, particularly in handling complex multi-hop queries. Moreover, ablation experiments validate the contribution of each agent in improving retrieval accuracy.
Manlian Bi, Zhijie Bao, Dongna Xie, Xiaohan Xie, Changxiao Yang, Tao Wang 0082, Yongtian Wang, Jiajie Peng
Briefings Bioinform.7
2025 Inference of gene coexpression networks from single-cell transcriptome data based on variance decomposition analysis
abstract
Gene regulation varies across different cell types and developmental stages, leading to distinct cellular roles across cellular populations. Investigating cell type-specific gene coexpression is therefore crucial for understanding gene functions and disease pathology. However, reconstructing gene coexpression networks from single-cell transcriptome data is challenging due to artifacts, noise, and data sparsity. Here, we present an efficient method for inference of gene coexpression networks via variance decomposition analysis (GCNVDA) to explore the underlying gene regulatory mechanisms from single-cell transcriptome data. Our model incorporates multiple sources of variability, including a random effect term $G$ to capture gene-level variance and a random effect term $E$ to account for residual errors. We applied GCNVDA to three real-world single-cell datasets, demonstrating that our method outperforms existing state-of-the-art algorithms in both sensitivity and specificity for identifying tissue- or state-specific gene regulations. Furthermore, GCNVDA facilitates the discovery of functional modules that play critical roles in key biological processes such as embryonic development. These findings provide new insights into cell-specific regulatory mechanisms and have the potential to significantly advance research in developmental biology and disease pathology.
Bin Lian, Haohui Zhang, Tao Wang 0082, Yongtian Wang, Xuequn Shang 0001, N. Ahmad Aziz, Jialu Hu
Briefings Bioinform.4
2025 cfMethylPre: deep transfer learning enhances cancer detection based on circulating cell-free DNA methylation profiling
abstract
Cancer remains a significant global health burden, underscoring the need for innovative diagnostic tools to enable early detection and improve patient outcomes. While circulating cell-free DNA (cfDNA) methylation has emerged as a promising biomarker for noninvasive cancer diagnostics, existing methods often face limitations in handling the high-dimensionality of methylation data, small sample sizes, and a lack of biological interpretability. To address these challenges, we propose cfMethylPre, a novel deep transfer learning framework tailored for cancer detection using cfDNA methylation data. cfMethylPre leverages large language model pretrained embeddings from DNA sequence information and integrates them with methylation profiles to enhance feature representation. The deep transfer learning process involves pretraining on bulk DNA methylation data encompassing 2801 samples across 82 cancer types and normal controls, followed by fine-tuning with cfDNA methylation data. This approach ensures robust adaptation to cfDNA's unique characteristics while improving predictive accuracy. Our model achieved superior predictive accuracy compared with state-of-the-art methods, with a weighted Matthews Correlation Coefficient of 0.926 and a weighted F1-score of 0.942. Through model interpretation and biological experimental validation, we identified three novel breast cancer genes-PCDHA10, PRICKLE2, and PRTG-demonstrating their inhibitory effects on cell proliferation and migration in breast cancer cell lines. These findings establish cfMethylPre as a powerful and interpretable tool for cancer diagnostics and biological discovery, paving the way for its application in precision oncology.
Xuchao Zhang, Yongtian Wang, Jialu Hu, Jiajie Peng, Xuequn Shang 0001, Yanpu Wang, Tao Wang 0082
Briefings Bioinform.3
2025 MAEST: accurately spatial domain detection in spatial transcriptomics with graph masked autoencoder
abstract
Spatial transcriptomics (ST) technology provides gene expression profiles with spatial context, offering critical insights into cellular interactions and tissue architecture. A core task in ST is spatial domain identification, which involves detecting coherent regions with similar spatial expression patterns. However, existing methods often fail to fully exploit spatial information, leading to limited representational capacity and suboptimal clustering accuracy. Here, we introduce MAEST, a novel graph neural network model designed to address these limitations in ST data. MAEST leverages graph masked autoencoders to denoise and refine representations while incorporating graph contrastive learning to prevent feature collapse and enhance model robustness. By integrating one-hop and multi-hop representations, MAEST effectively captures both local and global spatial relationships, improving clustering precision. Extensive experiments across diverse datasets, including the human brain, mouse hippocampus, olfactory bulb, brain, and embryo, demonstrate that MAEST outperforms seven state-of-the-art methods in spatial domain identification. Furthermore, MAEST showcases its ability to integrate multi-slice data, identifying joint domains across horizontal tissue sections with high accuracy. These results highlight MAEST's versatility and effectiveness in unraveling the spatial organization of complex tissues. The source code of MAEST can be obtained at https://github.com/clearlove2333/MAEST.
Han Shu, Yongtian Wang, Jialu Hu, Jiajie Peng, Xuequn Shang 0001, Zhen Tian 0004, Tao Wang 0082
Briefings Bioinform.3
2025 Incremental energy-based recurrent transformer-KAN for time series deformation simulation of soft tissue
Jiaxi Jiang, Tianyu Fu 0003, Jingfan Fan, Hong Song 0003, Danni Ai, Deqiang Xiao, Yongtian Wang, Jian Yang 0009
Expert Syst. Appl.8
2025 Collective Migration-Inspired Large-Deformation Compensation for Nonrigid Image Registration
Dingkun Liu, Danni Ai, Hong Song 0003, Jingfan Fan, Tianyu Fu 0003, Deqiang Xiao, Yongtian Wang, Jian Yang 0009
Int. J. Comput. Vis.8
2025 Integrative Graph-Based Framework for Predicting circRNA Drug Resistance Using Disease Contextualization and Deep Learning
abstract
Circular RNAs (circRNAs) play a crucial role in gene regulation and have been implicated in the development of drug resistance in cancer, representing a significant challenge in oncological therapeutics. Despite advancements in computational models predicting RNA-drug interactions, existing frameworks often overlook the complex interplay between circRNAs, drug mechanisms, and disease contexts. This study aims to bridge this gap by introducing a novel computational model, circRDRP, that enhances prediction accuracy by integrating disease-specific contexts into the analysis of circRNA-drug interactions. It employs a hybrid graph neural network that combines features from Graph Attention Networks (GAT) and Graph Convolutional Networks (GCN) in a two-layer structure, with further enhancement through convolutional neural networks. This approach allows for sophisticated feature extraction from integrated networks of circRNAs, drugs, and diseases. Our results demonstrate that the circRDRP model outperforms existing models in predicting drug resistance, showing significant improvements in accuracy, precision, and recall. Specifically, the model shows robust predictive capability in case studies involving major anticancer drugs such as Cisplatin and Methotrexate, indicating its potential utility in precision medicine. In conclusion, circRDRP offers a powerful tool for understanding and predicting drug resistance mediated by circRNAs, with implications for designing more effective cancer therapies.
Yongtian Wang, Wenkai Shen, Yewei Shen, Shang Feng, Tao Wang 0082, Xuequn Shang 0001, Jiajie Peng
IEEE J. Biomed. Health Informatics1
2025 Errata to "Depth Perception in Optical See-Through Augmented Reality: Investigating the Impact of Texture Density, Luminance Contrast, and Color Contrast"
abstract
In This paper, the information regarding the corresponding authors is missing. The corresponding authors of the paper should be Shining Ma and Weitao Song.
Chaochao Liu, Shining Ma, Yue Liu 0005, Yongtian Wang
IEEE Trans. Vis. Comput. Graph.4
2025 Influence of Object Height, Shadow and Adapting Luminance on Outdoor Depth Perception in Augmented Reality
abstract
Augmented reality (AR) technology has great potential in the applications of training, exhibition, and visual guidance, all of which demand precise virtual-real registration in perceived depth. Many AR applications such as navigation and tourism guidance are usually implemented in outdoor environments. However, prior research on depth perception in AR predominantly focused on the indoor environment, characterized by a lower illumination level and more confined space compared to outdoor settings. To address this gap, this paper presented a systematic investigation into the depth perception in outdoor environments. Two experiments were conducted in this study: the first one aimed to explore how to eliminate the bias induced by the floating object and how the knowledge of object height influences the perceived depth. The second experiment examined how ambient luminance affects depth estimation in AR. Our findings revealed an overestimation of perceived depth when participants were unaware of the actual height of the floating object, but an underestimation when they were informed of this information prior to the experiment. Additionally, shadows effectively reduced depth errors regardless of whether participants were informed of the object's height. The second experiment further indicated that, in outdoor environments, reducing ambient luminance significantly improves the accuracy of depth perception in AR.
Shining Ma, Chaochao Liu, Yue Liu 0005, Yongtian Wang
IEEE Trans. Vis. Comput. Graph.5
2025 Optimizing locomotion techniques in virtual reality: a comparative analysis of posture, interaction freedom, and methods
Qiuxin Du, Dongdong Weng, Yongtian Wang
Vis. Comput.8
2025 Learning intrinsic decomposition with semantic information fusion based on transformer
abstract
Intrinsic decomposition, the process of decomposing an image into reflectance and shading, is widely used in virtual and augmented reality tasks. Reflectance and shading often exhibit large gradients at the object edges, and the intrinsic properties on the same object tend to be similar. This spatial coherence is closely related to semantic consistency because objects within the same semantic category often exhibit similar intrinsic properties. Therefore, incorporating semantic segmentation into a deep intrinsic decomposition framework helps the network distinguish between different object instances and understand high-level scene structures. To this end, we design an intrinsic decomposition network jointly trained with a dedicated semantic segmentation module, allowing semantic cues to enhance the decomposition of reflectance and shading. The semantic module provides guidance during training but is removed during inference, improving performance without increasing the inference cost. Additionally, to capture the global contextual dependencies critical for intrinsic decomposition, we adopt a Transformer-based backbone. The proposed backbone enables the model to associate distant regions with similar material properties, thereby maintaining consistency in reflectance and learning smooth illumination patterns across a scene. A convolutional decoder is also designed to output predictions with improved details. Experiments demonstrate that our approach achieves state-of-the-art performance in the quantitative evaluations on the Intrinsic Images in the Wild (IIW) and Shading Annotations in the wild (SAW) datasets.
Pengjie Zhao, Hao Sha 0004, Yongtian Wang, Yue Liu 0005
Virtual Real. Intell. Hardw.3
2024 Accurately deciphering spatial domains for spatially resolved transcriptomics with stCluster
abstract
Spatial transcriptomics provides valuable insights into gene expression within the native tissue context, effectively merging molecular data with spatial information to uncover intricate cellular relationships and tissue organizations. In this context, deciphering cellular spatial domains becomes essential for revealing complex cellular dynamics and tissue structures. However, current methods encounter challenges in seamlessly integrating gene expression data with spatial information, resulting in less informative representations of spots and suboptimal accuracy in spatial domain identification. We introduce stCluster, a novel method that integrates graph contrastive learning with multi-task learning to refine informative representations for spatial transcriptomic data, consequently improving spatial domain identification. stCluster first leverages graph contrastive learning technology to obtain discriminative representations capable of recognizing spatially coherent patterns. Through jointly optimizing multiple tasks, stCluster further fine-tunes the representations to be able to capture complex relationships between gene expression and spatial organization. Benchmarked against six state-of-the-art methods, the experimental results reveal its proficiency in accurately identifying complex spatial domains across various datasets and platforms, spanning tissue, organ, and embryo levels. Moreover, stCluster can effectively denoise the spatial gene expression patterns and enhance the spatial trajectory inference. The source code of stCluster is freely available at https://github.com/hannshu/stCluster.
Tao Wang 0082, Han Shu, Jialu Hu, Yongtian Wang, Jin Chen 0004, Jiajie Peng, Xuequn Shang 0001
Briefings Bioinform.4
2024 Multi-scale convolutional neural networks and saliency weight maps for infrared and visible image fusion
Chenxuan Yang, Yunan He, Bingkun Chen, Jie Cao 0004, Yongtian Wang, Qun Hao
J. Vis. Commun. Image Represent.6
2024 Feature decomposition-based gaze estimation with auxiliary head pose regression
Ke Ni, Yongtian Wang
Pattern Recognit. Lett.6
2024 DSC-Recon: Dual-Stage Complementary 4-D Organ Reconstruction From X-Ray Image Sequence for Intraoperative Fusion
abstract
Accurately reconstructing 4D critical organs contributes to the visual guidance in X-ray image-guided interventional operation. Current methods estimate intraoperative dynamic meshes by refining a static initial organ mesh from the semantic information in the single-frame X-ray images. However, these methods fall short of reconstructing an accurate and smooth organ sequence due to the distinct respiratory patterns between the initial mesh and X-ray image. To overcome this limitation, we propose a novel dual-stage complementary 4D organ reconstruction (DSC-Recon) model for recovering dynamic organ meshes by utilizing the preoperative and intraoperative data with different respiratory patterns. DSC-Recon is structured as a dual-stage framework: 1) The first stage focuses on addressing a flexible interpolation network applicable to multiple respiratory patterns, which could generate dynamic shape sequences between any pair of preoperative 3D meshes segmented from CT scans. 2) In the second stage, we present a deformation network to take the generated dynamic shape sequence as the initial prior and explore the discriminate feature (i.e., target organ areas and meaningful motion information) in the intraoperative X-ray images, predicting the deformed mesh by introducing a designed feature mapping pipeline integrated into the initialized shape refinement process. Experiments on simulated and clinical datasets demonstrate the superiority of our method over state-of-the-art methods in both quantitative and qualitative aspects.
Haixiao Geng, Jingfan Fan, Sigeng Chen, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Hong Song 0003, Feng Duan 0001, Yongtian Wang, Jian Yang 0009
IEEE Trans. Medical Imaging11
2024 Realtime Recognition of Dynamic Hand Gestures in Practical Applications
abstract
Dynamic hand gesture acting as a semaphoric gesture is a practical and intuitive mid-air gesture interface. Nowadays benefiting from the development of deep convolutional networks, the gesture recognition has already achieved a high accuracy, however, when performing a dynamic hand gesture such as gestures of direction commands, some unintentional actions are easily misrecognized due to the similarity of the hand poses. This hinders the application of dynamic hand gestures and cannot be solved by just improving the accuracy of the applied algorithm on public datasets, thus it is necessary to study such problems from the perspective of human-computer interaction. In this article, two methods are proposed to avoid misrecognition by introducing activation delay and using asymmetric gesture design. First the temporal process of a dynamic hand gesture is decomposed and redefined, then a realtime dynamic hand gesture recognition system is built through a two-dimensional convolutional neural network. In order to investigate the influence of activation delay and asymmetric gesture design on system performance, a user study is conducted and experimental results show that the two proposed methods can effectively avoid misrecognition. The two methods proposed in this article can provide valuable guidance for researchers when designing realtime recognition system in practical applications.
Yi Xiao 0009, Yu Han 0011, Yue Liu 0005, Yongtian Wang
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Depth Perception in Optical See-Through Augmented Reality: Investigating the Impact of Texture Density, Luminance Contrast, and Color Contrast
abstract
The immersive augmented reality (AR) system necessitates precise depth registration between virtual objects and the real scene. Prior studies have emphasized the efficacy of surface texture in providing depth cues to enhance depth perception across various media, including the real scene, virtual reality, and AR. However, these studies predominantly focus on black-and-white textures, leaving a gap in understanding the effectiveness of colored textures. To address this gap and further explore texture-related factors in AR, a series of experiments were conducted to investigate the effects of different texture cues on depth perception using the perceptual matching method. Findings indicate that the absolute depth error increases with decreasing contrast under black-and-white texture. Moreover, textures with higher color contrast also contribute to enhanced accuracy of depth judgments in AR. However, no significant effect of texture density on depth perception was observed. The findings serve as a theoretical reference for texture design in AR, aiding in the optimization of virtual-real registration processes.
Chaochao Liu, Shining Ma, Yue Liu 0005, Yongtian Wang
IEEE Trans. Vis. Comput. Graph.4
2023 Enhanced RNA Sequence Representation through Sequence Masking and Subsequence Consistency Optimization
abstract
In the burgeoning field of RNA research, accurate and efficient RNA sequence representation remains a pivotal challenge, exacerbated by the complexity and diversity of RNA sequences. Addressing the critical need for enhanced sequence representation and the issues of sequence context and structural alignment, this study introduces a novel, comprehensive approach. The proposed model seamlessly integrates sequence masking and subsequence consistency optimization, offering a robust solution to the intricate problem of RNA sequence representation. Utilizing the filtered RNAStralign dataset, encompassing 20,923 sequences, the model's performance is rigorously evaluated employing a Support Vector Machine (SVM) for subsequent RNA family classification tasks. Despite the inherent imbalance in RNA family sequence distribution, the model demonstrates exemplary performance, achieving high classification accuracy and AUPRC values across diverse RNA sequence groups. This balanced and unbiased assessment, ensured by the use of AUPRC as an evaluation metric, highlights the model's practical utility for comprehensive RNA sequence analysis and classification. In essence, this research presents a method for enhanced RNA sequence representation and laying a robust foundation for future advancements in the nuanced field of RNA sequence analysis.
Yewei Shen, Zongyu Li, Xinmeng Liu, Xuequn Shang 0001, Yongtian Wang
BIBM6
2023 Deep Learning Integration with Phenotypic Similarities and Heterogeneous Networks for Drug-Target Interaction Prediction
abstract
In the field of drug discovery, the accurate prediction of drug-target interactions (DTIs) is a critical yet challenging task, hindered by the intricate dynamics of biological systems and molecular interplay. To address this, we propose the DTI-VGAE model, a novel deep learning framework that integrates variational graph autoencoders (VGAE) with a multi-layer perceptron (MLP) for robust DTI prediction. Our approach focuses on three key aspects: learning distinct representations of drugs and proteins from heterogeneous networks, constructing Drug-Protein Pair (DPP) networks to capture the complex interactions, and employing MLP for the final prediction of DTIs. This comprehensive methodology not only enhances the accuracy of DTI predictions but also ensures greater reliability and stability. Validated through extensive 5-fold cross-validation, the DTI-VGAE model consistently outperforms existing methods, achieving superior average AUROC, AUPR scores, and accuracy. The DTI-VGAE model's innovative integration of VGAE and MLP offers a significant advancement in the computational approach to drug discovery, paving the way for more efficient and precise drug development processes.
Yongtian Wang, Yewei Shen, Xuequn Shang 0001
BIBM1
2023 Collaborative deep learning improves disease-related circRNA prediction based on multi-source functional information
abstract
Emerging studies have shown that circular RNAs (circRNAs) are involved in a variety of biological processes and play a key role in disease diagnosing, treating and inferring. Although many methods, including traditional machine learning and deep learning, have been developed to predict associations between circRNAs and diseases, the biological function of circRNAs has not been fully exploited. Some methods have explored disease-related circRNAs based on different views, but how to efficiently use the multi-view data about circRNA is still not well studied. Therefore, we propose a computational model to predict potential circRNA-disease associations based on collaborative learning with circRNA multi-view functional annotations. First, we extract circRNA multi-view functional annotations and build circRNA association networks, respectively, to enable effective network fusion. Then, a collaborative deep learning framework for multi-view information is designed to get circRNA multi-source information features, which can make full use of the internal relationship among circRNA multi-view information. We build a network consisting of circRNAs and diseases by their functional similarity and extract the consistency description information of circRNAs and diseases. Last, we predict potential associations between circRNAs and diseases based on graph auto encoder. Our computational model has better performance in predicting candidate disease-related circRNAs than the existing ones. Furthermore, it shows the high practicability of the method that we use several common diseases as case studies to find some unknown circRNAs related to them. The experiments show that CLCDA can efficiently predict disease-related circRNAs and are helpful for the diagnosis and treatment of human disease.
Yongtian Wang, Xinmeng Liu, Yewei Shen, Xuerui Song, Tao Wang 0082, Xuequn Shang 0001, Jiajie Peng
Briefings Bioinform.1
2023 scMultiGAN: cell-specific imputation for single-cell transcriptomes with multiple deep generative adversarial networks
abstract
The emergence of single-cell RNA sequencing (scRNA-seq) technology has revolutionized the identification of cell types and the study of cellular states at a single-cell level. Despite its significant potential, scRNA-seq data analysis is plagued by the issue of missing values. Many existing imputation methods rely on simplistic data distribution assumptions while ignoring the intrinsic gene expression distribution specific to cells. This work presents a novel deep-learning model, named scMultiGAN, for scRNA-seq imputation, which utilizes multiple collaborative generative adversarial networks (GAN). Unlike traditional GAN-based imputation methods that generate missing values based on random noises, scMultiGAN employs a two-stage training process and utilizes multiple GANs to achieve cell-specific imputation. Experimental results show the efficacy of scMultiGAN in imputation accuracy, cell clustering, differential gene expression analysis and trajectory analysis, significantly outperforming existing state-of-the-art techniques. Additionally, scMultiGAN is scalable to large scRNA-seq datasets and consistently performs well across sequencing platforms. The scMultiGAN code is freely available at https://github.com/Galaxy8172/scMultiGAN.
Tao Wang 0082, Yungang Xu, Yongtian Wang, Xuequn Shang 0001, Jiajie Peng, Bing Xiao 0001
Briefings Bioinform.4
2023 DFinder: a novel end-to-end graph embedding-based method to identify drug-food interactions
abstract
MOTIVATION: Drug-food interactions (DFIs) occur when some constituents of food affect the bioaccessibility or efficacy of the drug by involving in drug pharmacodynamic and/or pharmacokinetic processes. Many computational methods have achieved remarkable results in link prediction tasks between biological entities, which show the potential of computational methods in discovering novel DFIs. However, there are few computational approaches that pay attention to DFI identification. This is mainly due to the lack of DFI data. In addition, food is generally made up of a variety of chemical substances. The complexity of food makes it difficult to generate accurate feature representations for food. Therefore, it is urgent to develop effective computational approaches for learning the food feature representation and predicting DFIs. RESULTS: In this article, we first collect DFI data from DrugBank and PubMed, respectively, to construct two datasets, named DrugBank-DFI and PubMed-DFI. Based on these two datasets, two DFI networks are constructed. Then, we propose a novel end-to-end graph embedding-based method named DFinder to identify DFIs. DFinder combines node attribute features and topological structure features to learn the representations of drugs and food constituents. In topology space, we adopt a simplified graph convolution network-based method to learn the topological structure features. In feature space, we use a deep neural network to extract attribute features from the original node attributes. The evaluation results indicate that DFinder performs better than other baseline methods. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/23AIBox/23AIBox-DFinder. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tao Wang 0082, Jinjin Yang, Yifu Xiao, Yuxian Wang, Yongtian Wang, Jiajie Peng
Bioinform.7
2023 Geometry-guided multilevel RGBD fusion for surface normal estimation
Yanfeng Tong, Jing Chen 0018, Yongtian Wang
Comput. Commun.3
2023 f2-ToF: A feature-alignment and frequency-division time-of-flight data denoise network
Yanfeng Tong, Jing Chen 0018, Zhen Leng, Yongtian Wang
Comput. Commun.5
2022 CircRNA-Disease Association Prediction based on Heterogeneous Graph Representation
abstract
Circular RNAs (circRNAs) have important effects on various biological processes, and their dysfunction is closely related to the emergence and development of diseases. Identifying the associations between circRNAs and diseases is helpful in analyzing the pathogenesis of diseases. Therefore, it is necessary to develop effective computational methods for predicting circRNA-disease associations. Here, we present a computational model called HRCDA to predict associations between circRNA and disease based on heterogeneous graph representation. Firstly, an integrated network of circRNA functional similarity is built by Random Walk with Restart in the view of biological functions of circRNA. Then, a heterogeneous graph of circRNAs and diseases is constructed with known circRNA-disease associations. Finally, we design a heterogeneous graph representation learn model based on Graph Auto-Encoder (GAE) to predict circRNA-disease associations. Experiments have shown that the proposed method perform better than existing state-of-the-art methods and can be an effective tool to predict potential disease-related circRNAs.
Xinmeng Liu, Yewei Shen, Xuequn Shang 0001, Yongtian Wang
BIBM5
2022 Discovering eQTL Regulatory Patterns Through eQTLMotif
abstract
The expression quantitative trait loci (eQTL) analysis has become important for understanding the regulatory function of genomic variants on gene expression in a tissuespecific manner and has been widely applied across species from microbes to mammals. Current eQTL studies mainly focus on the simple one-to-one regulation between variant and gene. Recent research have demonstrated there are also more complex regulatory patterns between eQTLs and genes. However, there is a lack of studies and relevant methods to systematically discover the regulatory patterns between multiple eQTLs and multiple genes. In this regard, this study has proposed a novel computational framework, called eQTLMotif, to discover regulation patterns of eQTLs in a many-to-many manner. This framework mainly consists of two steps: (1) construct a novel eQTL regulatory network by integrating bipartite eQTL network, eQTL mediation effects, and gene regulatory network; (2) perform motif mining through exactly enumerating frequently appeared eQTL regulatory structures. Based on this framework, we for the first time systematically investigated the eQTL regulatory patterns in the human frontal cortex based on a large cohort of postmortem human brains. Experiments have demonstrated that our framework can effectively reveal novel eQTL regulatory patterns. And some are in similar structure to the existing gene regulation patterns, such as feed-forward loop (FFL)-like motif, single input module (SIM)-like motif, and dense overlapping regulons (DOR)- like motif. Our method and findings will further enhance the understanding of regulatory mechanisms of eQTLs in multiple tissues and species.
Tao Wang 0082, Yifu Xiao, Hanzi Yang, Xipeng Yin, Yongtian Wang, Bing Xiao 0001, Xuequn Shang 0001, Jiajie Peng
BIBM6
2022 A network-based method for brain disease gene prediction by integrating brain connectome and molecular network
abstract
Brain disease gene identification is critical for revealing the biological mechanism and developing drugs for brain diseases. To enhance the identification of brain disease genes, similarity-based computational methods, especially network-based methods, have been adopted for narrowing down the searching space. However, these network-based methods only use molecular networks, ignoring brain connectome data, which have been widely used in many brain-related studies. In our study, we propose a novel framework, named brainMI, for integrating brain connectome data and molecular-based gene association networks to predict brain disease genes. For the consistent representation of molecular-based network data and brain connectome data, brainMI first constructs a novel gene network, called brain functional connectivity (BFC)-based gene network, based on resting-state functional magnetic resonance imaging data and brain region-specific gene expression data. Then, a multiple network integration method is proposed to learn low-dimensional features of genes by integrating the BFC-based gene network and existing protein-protein interaction networks. Finally, these features are utilized to predict brain disease genes based on a support vector machine-based model. We evaluate brainMI on four brain diseases, including Alzheimer's disease, Parkinson's disease, major depressive disorder and autism. brainMI achieves of 0.761, 0.729, 0.728 and 0.744 using the BFC-based gene network alone and enhances the molecular network-based performance by 6.3% on average. In addition, the results show that brainMI achieves higher performance in predicting brain disease genes compared to the existing three state-of-the-art methods.
Ruijiang Han, Menghan Zhang, Yuxian Wang, Tao Wang 0082, Yongtian Wang, Xuequn Shang 0001, Jiajie Peng
Briefings Bioinform.6
2022 Enhancing discoveries of molecular QTL studies with small sample size using summary statistic imputation
abstract
Quantitative trait locus (QTL) analyses of multiomic molecular traits, such as gene transcription (eQTL), DNA methylation (mQTL) and histone modification (haQTL), have been widely used to infer the functional effects of genome variants. However, the QTL discovery is largely restricted by the limited study sample size, which demands higher threshold of minor allele frequency and then causes heavy missing molecular trait-variant associations. This happens prominently in single-cell level molecular QTL studies because of sample availability and cost. It is urgent to propose a method to solve this problem in order to enhance discoveries of current molecular QTL studies with small sample size. In this study, we presented an efficient computational framework called xQTLImp to impute missing molecular QTL associations. In the local-region imputation, xQTLImp uses multivariate Gaussian model to impute the missing associations by leveraging known association statistics of variants and the linkage disequilibrium (LD) around. In the genome-wide imputation, novel procedures are implemented to improve efficiency, including dynamically constructing a reused LD buffer, adopting multiple heuristic strategies and parallel computing. Experiments on various multiomic bulk and single-cell sequencing-based QTL datasets have demonstrated high imputation accuracy and novel QTL discovery ability of xQTLImp. Finally, a C++ software package is freely available at https://github.com/stormlovetao/QTLIMP.
Tao Wang 0082, Yongzhuang Liu, Quanwei Yin, Jiaquan Geng, Jin Chen 0004, Xipeng Yin, Yongtian Wang, Xuequn Shang 0001, Chunwei Tian, Yadong Wang 0001, Jiajie Peng
Briefings Bioinform.7
2022 Correction to: Enhancing discoveries of molecular QTL studies with small sample size using summary statistic imputation
abstract
In the originally published version of this manuscript, there was an error in the Funding section; ‘National Natural Science Foundation of China (6210071334, 62072376)’ has now been corrected to ‘National Natural Science Foundation of China (62102319, 62072376)’.
Tao Wang 0082, Yongzhuang Liu, Quanwei Yin, Jiaquan Geng, Jin Chen 0004, Xipeng Yin, Yongtian Wang, Xuequn Shang 0001, Chunwei Tian, Yadong Wang 0001, Jiajie Peng
Briefings Bioinform.7
2022 Navigation in virtual and real environment using brain computer interface: a progress report
abstract
A brain-computer interface (BCI) facilitates bypassing the peripheral nervous system and directly communicating with surrounding devices. Navigation technology using BCI has developed—from exploring the prototype paradigm in the virtual environment (VE) to accurately completing the locomotion intention of the operator in the form of a powered wheelchair or mobile robot in a real environment. This paper summarizes BCI navigation applications that have been used in both real and VEs in the past 20 years. Horizontal comparisons were conducted between various paradigms applied to BCI and their unique signal-processing methods. Owing to the shift in the control mode from synchronous to asynchronous, the development trend of navigation applications in the VE was also reviewed. The contrast between highlevel commands and low-level commands is introduced as the main line to review the two major applications of BCI navigation in real environments: mobile robots and unmanned aerial vehicles (UAVs). Finally, applications of BCI navigation to scenarios outside the laboratory; research challenges, including human factors in navigation application interaction design; and the feasibility of hybrid BCI for BCI navigation are discussed in detail.
Haochen Hu, Yue Liu 0005, Kang Yue, Yongtian Wang
Virtual Real. Intell. Hardw.4
2021 Cross-Domain Transfer Learning for Vessel Segmentation in Computed Tomographic Coronary Angiographic Images
Ruirui An, Danni Ai, Yongtian Wang, Jian Yang 0009
ICIG (2)5
2021 Novel Augmented Reality System for Oral and Maxillofacial Surgery
Lele Ding, Long Shao, Zehua Zhao, Tao Zhang 0152, Danni Ai, Jian Yang 0009, Yongtian Wang
ICIG (2)7
2021 A New Dataset and Recognition for Egocentric Microgesture Designed by Ergonomists
Guangchuan Li, Yue Liu 0005, Yongtian Wang
ICIG (2)5
2021 Single Scene Image Editing Based on Deep Intrinsic Decomposition
Hao Sha 0004, Yue Liu 0005, Chenguang Lu, Hengrun Chen, Yongtian Wang
ICIG (3)6
2021 Analysis of teenagers' preferences and concerns regarding HMDs in education
abstract
Virtual reality (VR) has become a powerful and promising tool for education, and numerous studies have investigated the application and effectiveness of VR education. However, few studies have focused on the expectations and concerns of teenagers regarding head-mounted displays (HMDs), which are used for this purpose. In this paper, we aim to explore the current problems and necessary advancements required in VR education based on a survey of 163 senior high school students who experience VR educational content for 1h. The usability and comfort of the HMD system, the physical and psychological effects on the students, and their preferences and concerns are investigated. The results show that HMDs increase students' interest, concentration, and enthusiasm for learning. However, isolated virtual environments make students feel nervous and afraid. The immersive environment also makes them worry about VR addiction and confusing the physical world with the virtual one. VR has great potential in the field of education, but the issue of safety needs to be considered in the future.
Jie Guo 0004, Dongdong Weng, Yue Liu 0005, Qiyong Chen, Yongtian Wang
Virtual Real. Intell. Hardw.5
2020 Learning to Deblur Face Images via Sketch Synthesis
abstract
The success of existing face deblurring methods based on deep neural networks is mainly due to the large model capacity. Few algorithms have been specially designed according to the domain knowledge of face images and the physical properties of the deblurring process. In this paper, we propose an effective face deblurring algorithm based on deep convolutional neural networks (CNNs). Motivated by the conventional deblurring process which usually involves the motion blur estimation and the latent clear image restoration, the proposed algorithm first estimates motion blur by a deep CNN and then restores latent clear images with the estimated motion blur. However, estimating motion blur from blurry face images is difficult as the textures of the blurry face images are scarce. As most face images share some common global structures which can be modeled well by sketch information, we propose to learn face sketches by a deep CNN so that the sketches can help the motion blur estimation. With the estimated motion blur, we then develop an effective latent image restoration algorithm based on a deep CNN. Although involving the several components, the proposed algorithm is trained in an end-to-end fashion. We analyze the effectiveness of each component on face image deblurring and show that the proposed algorithm is able to deblur face images with favorable performance against state-of-the-art methods.
Songnan Lin, Jiawei Zhang 0002, Jinshan Pan, Yicun Liu, Yongtian Wang, Jing S. J. Chen, Jimmy S. J. Ren
AAAI5
2020 A General Endoscopic Image Enhancement Method Based on Pre-trained Generative Adversarial Networks
abstract
Endoscopic images frequently have image quality problems due to the limitations of surgical instruments and the impact of surgical operations, such as uneven illumination, smogginess and color deviation. For deep learning based on enhancement methods, independent training lacks sufficient defect images and generalization capability, and combined training with mixture of data cannot identify diverse specific tasks. To address these issues, we propose a general method based on pre-trained generative adversarial network with a specified transfer learning strategy to obtain high-quality images. Initially, we independently train a standard network based on a universal task, e.g., uneven illumination, where a pre-trained model is extracted as a backbone with partially shared generator. Then, we transfer the backbone to more potential image enhancement tasks. Experiments on uneven illumination, smogginess, and color deviation indicate that the model successfully shares common features of high-quality images and responds specifically to different defects as well.
Jingfan Fan, Danni Ai, Hong Song 0003, Yongtian Wang, Jian Yang 0009
BIBM5
2020 Learning Event-Driven Video Deblurring and Interpolation
Songnan Lin, Jiawei Zhang 0002, Jinshan Pan, Dongqing Zou, Yongtian Wang, Jing Chen 0018, Jimmy S. J. Ren
ECCV (8)6
2020 Exploring the Differences of Visual Discomfort Caused by Long-term Immersion between Virtual Environments and Physical Environments
abstract
To investigate the effects of visual discomfort caused by long-term immersing in virtual environments (VEs), we conducted a comparative study to evaluate users’ visual discomfort in an eight-hour working rhythm and compared the differences between the VEs and the physical environments. Twenty-seven participants performed four different visual tasks with a head-mounted display (HMD) for the VE condition and with a monitor for the physical condition. Their subjective visual discomfort and objective oculomotor indicators were measured to evaluate their visual performances. The results show that the subjective visual fatigue symptoms, the objective pupil size, and the relative accommodation response vary across time for the two conditions, in which VEs affects visual fatigue the most compared to the physical environments. The results also show that pupil size is negatively related to subjective visual fatigue, and the long-term work based on displays only influences the maximum accommodation response of participants. This work is a supplement to the necessary but insufficient-researched field of visual fatigue in long-term immersing in VEs, which should be valuable to researchers involved in the evaluation of visual fatigue using HMDs.
Jie Guo 0004, Dongdong Weng, Zhenliang Zhang 0002, Jiamin Ping, Yue Liu 0005, Yongtian Wang
VR7
2020 Evaluating individual genome similarity with a topic model
abstract
MOTIVATION: Evaluating genome similarity among individuals is an essential step in data analysis. Advanced sequencing technology detects more and rarer variants for massive individual genomes, thus enabling individual-level genome similarity evaluation. However, the current methodologies, such as the principal component analysis (PCA), lack the capability to fully leverage rare variants and are also difficult to interpret in terms of population genetics. RESULTS: Here, we introduce a probabilistic topic model, latent Dirichlet allocation, to evaluate individual genome similarity. A total of 2535 individuals from the 1000 Genomes Project (KGP) were used to demonstrate our method. Various aspects of variant choice and model parameter selection were studied. We found that relatively rare (0.001 20 000 bp) variants are more efficient for genome similarity evaluation. At least 100 000 such variants are necessary. In our results, the populations show significantly less mixed and more cohesive visualization than the PCA results. The global similarities among the KGP genomes are consistent with known geographical, historical and cultural factors. AVAILABILITY AND IMPLEMENTATION: The source code and data access are available at: https://github.com/lrjuan/LDA_genome. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Liran Juan, Yongtian Wang, Jingyi Jiang, Guohua Wang 0001, Yadong Wang 0001
Bioinform.2
2020 Cross-spectral stereo matching for facial disparity estimation in the dark
Songnan Lin, Jiawei Zhang 0002, Jing Chen 0018, Yongtian Wang, Yicun Liu, Jimmy S. J. Ren
Comput. Vis. Image Underst.4
2020 Prior information constrained alternating direction method of multipliers for longitudinal compressive sensing MR imaging
Ruirui Kang, Danni Ai, Gangrong Qu, Qingbo Li, Yurong Jiang, Yong Huang 0002, Hong Song 0003, Yongtian Wang, Jian Yang 0009
Neurocomputing9
2020 Groupwise registration with global-local graph shrinkage in atlas construction
Tianyu Fu 0003, Jian Yang 0009, Danni Ai, Hong Song 0003, Yurong Jiang, Yongtian Wang, Alejandro F. Frangi
Medical Image Anal.7
2020 Topology Optimization Using Multiple-Possibility Fusion for Vasculature Extraction
abstract
Vascular centerline extraction from angiography images plays an important role in computer-aided diagnosis of vascular disease. To solve the common problems related to noise and inconsistent vasculatures from uneven perfusion, this paper proposes an automatic framework for accurate vascular centerline extraction from angiograms that uses multi-probability fusion-based topology optimization. In this framework, vascular region is first segmented using a learning-based method. Then, initial centerlines are obtained by applying iterative filtering operation and multi-direction indexed non-maximum suppression. Topology optimization is achieved by gap filling. A connection probability map is constructed utilizing the information of initial centerlines, texture, and orientation of vasculatures. Shortest path tracking is employed to search for optimal connections around gaps in the initial centerlines. The proposed framework is evaluated using simulative and clinical coronary angiographies. The experimental results demonstrate that the proposed method can extract centerlines with F1 score of 97.28% ± 1.2% for vasculatures in 12 clinical angiographic images. It is evident that the proposed method can extract complete and accurate vascular centerlines from angiograms and can be used to repair gaps in other filamentary structures, such as roads and retinal blood vessels. This endows our method a great potential in the analysis of filamentary structures.
Huihui Fang, Danni Ai, Weijian Cong, Yong Huang 0002, Hong Song 0003, Yongtian Wang, Jian Yang 0009
IEEE Trans. Circuits Syst. Video Technol.8
2020 Greedy Soft Matching for Vascular Tracking of Coronary Angiographic Image Sequences
abstract
Vascular tracking of coronary angiographic image sequences is one of the most clinically important tasks in the diagnostic assessment and interventional guidance of cardiac disease. It is difficult to automate this application because the vascular structure is complex; moreover, unsatisfactory angiography image quality may exacerbate the difficulty of vasculature extraction. This paper converts vascular tracking into branch matching and proposes a novel and automatic greedy soft match algorithm. Our method is based on a graph framework. A graph model building module is proposed to represent the vascular structure. Then, a greedy branch searching method is adopted to acquire all possible paths in the graph that may match the reference vessel. Finally, a soft batch matching method that combines branch descriptor and dynamic time warping is presented to select the best matching branch. The solution to the problem takes advantage of both spatial and temporal continuity between successive frames. The experimental results demonstrate that the proposed algorithm is effective and robust for vascular tracking. The F1 score of a single branch dataset, which contains 12 angiographic image sequences with 77 angiograms of contrast agent-filled vessels, is 0.89 ± 0.06 and of a vessel tree dataset which contains nine sequences with 58 angiograms is 0.88 ± 0.05. Extensive experimental results well demonstrate the superior performance of the algorithm. In addition, it provides a universal solution to address the problem of filamentary structure tracking.
Huihui Fang, Danni Ai, Yong Huang 0002, Yurong Jiang, Hong Song 0003, Yongtian Wang, Jian Yang 0009
IEEE Trans. Circuits Syst. Video Technol.7
2020 Spatio-Temporal Constrained Online Layer Separation for Vascular Enhancement in X-Ray Angiographic Image Sequence
abstract
Automatic vascular enhancement is crucial to vascular structure identification in X-ray angiographic (XRA) image sequences. In this work, we propose a novel spatio-temporal constrained online layer separation (STOLS) method to achieve vascular enhancement in XRA image sequences. The proposed method integrates the motion consistency of structures into the temporal-constrained online robust principal component analysis (ORPCA) to remove quasi-static structures (e.g., bones) from the enhanced vascular images. Furthermore, smoothing technique is integrated into the spatial-constrained ORPCA to reduce motion artifacts and the noise introduced by non-uniform illumination. To make the proposed method more adaptive to various vascular structures, the spatial-constrained ORPCA is adjusted by an adaptive weight using the proportion of the vessel region in the previous frame. The performance of the proposed method is compared with five state-of-the-art subtraction methods with respect to local and global revised contrast-to-noise ratios (rCNRs) and reconstruction errors. For the proposed method, the local and global rCNRs of the final vessel layer reached 2.54 and 1.24, respectively, while the error between the original and reconstructed images from the respiratory, background, and vessel layer reached 0.0354. The proposed STOLS can enhance the angiograms in a real-time and online manner without fine-tuning parameters, and can thus be used for intra-operation diagnosis and interventional procedures of coronary artery diseases.
Shuang Song 0005, Chenbing Du, Danni Ai, Yong Huang 0002, Hong Song 0003, Yongtian Wang, Jian Yang 0009
IEEE Trans. Circuits Syst. Video Technol.6
2020 Analysis on Mitigation of Visually Induced Motion Sickness by Applying Dynamical Blurring on a User's Retina
abstract
Visually induced motion sickness (MS) experienced in a 3D immersive virtual environment (VE) limits the widespread use of virtual reality (VR). This paper studies the effects of a saliency detection-based approach on the reduction of MS when the display on a user's retina is dynamic blurred. In the experiment, forty participants were exposed to a VR experience under a control condition without applying dynamic blurring, and an experimental condition applying dynamic blurring. The experimental results show that the participants under the experimental condition report a statistically significant reduction in the severity of MS symptoms on average during the VR experience compared to those under the control condition, which demonstrates that the proposed approach may alleviate visually induced MS in VR and enable users to remain in a VE for a longer period of time.
Guang-Yu Nie, Henry Been-Lirn Duh, Yue Liu 0005, Yongtian Wang
IEEE Trans. Vis. Comput. Graph.4
2019 Multi-Level Context Ultra-Aggregation for Stereo Matching
abstract
Exploiting multi-level context information to cost volume can improve the performance of learning-based stereo matching methods. In recent years, 3-D Convolution Neural Networks (3-D CNNs) show the advantages in regularizing cost volume but are limited by unary features learning in matching cost computation. However, existing methods only use features from plain convolution layers or a simple aggregation of multi-level features to calculate cost volume, which is insufficient because stereo matching requires discriminative features to identify corresponding pixels in rectified stereo image pairs. In this paper, we propose a unary features descriptor using multi-level context ultra-aggregation (MCUA), which encapsulates all convolutional features into a more discriminative representation by intra- and inter-level features combination. Specifically, a child module that takes low-resolution images as input captures larger context information; the larger context information from each layer is densely connected to the main branch of the network. MCUA makes good usage of multi-level features with richer context and performs the image-to-image prediction holistically. We introduce our MCUA scheme for cost volume calculation and test it on PSM-Net. We also evaluate our method on Scene Flow and KITTI 2012/2015 stereo datasets. Experimental results show that our method outperforms state-of-the-art methods by a notable margin and effectively improves the accuracy of stereo matching.
Guang-Yu Nie, Ming-Ming Cheng, Yun Liu 0011, Zhengfa Liang, Deng-Ping Fan, Yue Liu 0005, Yongtian Wang
CVPR7
2019 Deep Surface Normal Estimation With Hierarchical RGB-D Fusion
abstract
The growing availability of commodity RGB-D cameras has boosted the applications in the field of scene understanding. However, as a fundamental scene understanding task, surface normal estimation from RGB-D data lacks thorough investigation. In this paper, a hierarchical fusion network with adaptive feature re-weighting is proposed for surface normal estimation from a single RGB-D image. Specifically, the features from color image and depth are successively integrated at multiple scales to ensure global surface smoothness while preserving visually salient details. Meanwhile, the depth features are re-weighted with a confidence map estimated from depth before merging into the color branch to avoid artifacts caused by input depth corruption. Additionally, a hybrid multi-scale loss function is designed to learn accurate normal estimation given noisy ground-truth dataset. Extensive experimental results validate the effectiveness of the fusion strategy and the loss design, outperforming state-of-the-art normal estimation schemes.
Jin Zeng 0004, Yanfeng Tong, Yunmu Huang, Qiong Yan, Wenxiu Sun, Jing Chen 0018, Yongtian Wang
CVPR7
2019 Spatial Probabilistic Distribution Map Based 3D FCN for Visual Pathway Segmentation
Zhiqi Zhao, Danni Ai, Jingfan Fan, Hong Song 0003, Yongtian Wang, Jian Yang 0009
ICIG (2)6
2019 High-Fidelity Grasping in Virtual Reality using a Glove-based System
abstract
This paper presents a design that jointly provides hand pose sensing, hand localization, and haptic feedback to facilitate real-time stable grasps in Virtual Reality (VR). The design is based on an easy-to-replicate glove-based system that can reliably perform (i) a high-fidelity hand pose sensing in real time through a network of 15 IMUs, and (ii) the hand localization using a Vive Tracker. The supported physics-based simulation in VR is capable of detecting collisions and contact points for virtual object manipulation, which drives the collision event to trigger the physical vibration motors on the glove to signal the user, providing a better realism inside virtual environments. A caging-based approach using collision geometry is integrated to determine whether a grasp is stable. In the experiment, we showcase successful grasps of virtual objects with large geometry variations. Comparing to the popular LeapMotion sensor, we demonstrate the proposed glove-based design yields a higher success rate in various tasks in VR. We hope such a glove-based system can simplify the data collection of human manipulations with VR.
Hangxin Liu, Zhenliang Zhang 0002, Xu Xie 0001, Yixin Zhu 0001, Yue Liu 0005, Yongtian Wang, Song-Chun Zhu
ICRA6
2019 Toward an Efficient Hybrid Interaction Paradigm for Object Manipulation in Optical See-Through Mixed Reality
abstract
Human-computer interaction (HCI) plays an important role in the near-field mixed reality, in which the hand-based interaction is one of the most widely-used interaction modes, especially in the applications based on optical see-through head-mounted displays (OST-HMDs). In this paper, such interaction modes as gesture-based interaction (GBI) and physics-based interaction (PBI) are developed to construct a mixed reality system to evaluate the advantages and disadvantages of different interaction modes. The ultimate goal is to find an efficient hybrid paradigm for mixed reality applications based on OST-HMDs to deal with the situations that a single interaction mode cannot handle. The results of the experiment, which compares GBI and PBI, show that PBI leads to a better performance of users regarding their work efficiency in the proposed two tasks. Some statistical tests, including T-test and one-way ANOVA, have also been adopted to prove that the difference regarding the efficiency between different interaction modes is significant. Experiments for combining both interaction modes are put forward in order to seek a good experience for manipulation, which proves that the partially-overlapping style would help to improve work efficiency for manipulation tasks. The experimental results of the proposed two hand-based interaction modes and their hybrid forms can provide some practical suggestions for the development of mixed reality systems based on OST-HMDs.
Zhenliang Zhang 0002, Dongdong Weng, Jie Guo 0004, Yue Liu 0005, Yongtian Wang
IROS5
2019 Mixed Reality Office System Based on Maslow's Hierarchy of Needs: Towards the Long-Term Immersion in Virtual Environments
abstract
In a mixed reality (MR) environment that combines the physical objects with the virtual environments, users' feelings are immersed in the virtual world, while their bodies remain in the physical world. Compared to the purely physical environments, such characteristic has led to some special needs for users' long-term immersion. However, the deficiency needs that we have to face for long-term immersion still need further research. In this paper, we apply the theory of Maslow's Hierarchy of Needs (MHN) to guide the design of MR systems for long-term immersion. Taking the normal biological rhythm of human beings as the basic unit (24 hours), we propose the fundamental needs for long-term immersion in VEs through combining the theory of MHN with the special needs of virtual reality (VR). In order to verify whether those needs can satisfy users' long-term immersion, we design an MR office system for basic operations based on the theory of MHN. A long-term exposure experiment (duration of 8 hours) is designed to evaluate those needs by comparing the results with a physical work environment after a short-term preliminary study. The physiological and psychological effects are tested in both two environments and the deficiency needs for short-term immersion and long-term immersion are also compared. The results showed that the design based on the theory of MHN can support users' long-term immersion, which means that it can be a guideline for long-term use of MR systems.
Jie Guo 0004, Dongdong Weng, Zhenliang Zhang 0002, Yue Liu 0005, Yongtian Wang, Henry Been-Lirn Duh
ISMAR6
2019 Evaluation of Maslows Hierarchy of Needs on Long-Term Use of HMDs - A Case Study of Office Environment
abstract
Long-term exposure to VR will become more and more important, but what we need for long term immersion to meet users fundamental needs is still under-researched. In this paper, we apply the theory of Maslows Hierarchy of Needs to guide the design of VR for longterm immersion based on the normal biological rhythm of human beings (24 hours). An office environment is designed to verify those needs. The efficiency, the physical and the psychological effects of this VR office system are tested. The results show that the VR office environment is as comfortable as the physical environment at short-term immersion and it can support users basic immersion. It means that the Maslows Hierarchy of Needs can be a guideline for long-term immersion.
Jie Guo 0004, Dongdong Weng, Zhenliang Zhang 0002, Yue Liu 0005, Yongtian Wang
VR5
2019 A Neural Motion Deblurring Approach to Restore Rich Textures for Visual SLAM
abstract
In this paper, we present a sequential video deblurring method based on a spatio-temporal recurrent network for visual SLAM. The method can be applied to any SLAM systems to make sure continuous localization even with blurred images. The quality of the deblurring method is evaluated on real-world problems: feature points extraction and SLAM, which prove the method can significantly improve the performance of tracking accuracy especially in some severe cases containing strong camera shake or fast motion.
Guojing Jin, Jing Chen 0018, Yongtian Wang
VR4
2019 Real Time 3D Magnetic Field Visualization Based on Augmented Reality
abstract
In physics teaching, electromagnetism is one of the most difficult concepts for students to understand. This paper proposes a real time visualization method for 3-D magnetic field based on the augmented reality technology, which can not only visualize magnetic flux lines in real time, but also simulates the approximate sparse distribution of magnetic flux lines in space. An application utilizing this method is also presented. It permits leaners to freely and interactively move the magnets in 3-D space and to observe the magnetic flux lines in real time. As a result, the proposed method visualizes the invisible factors in 3-D magnetic field, with which students will have real-life reference when studying electromagnetic.
Yue Liu 0005, Yongtian Wang
VR3
2019 Exploring Stereovision-Based 3-D Scene Reconstruction for Augmented Reality
abstract
Three-dimensional (3-D) scene reconstruction is one of the key techniques in Augmented Reality (AR), which is related to the integration of image processing and display systems of complex information. Stereo matching is a computer vision based approach for 3-D scene reconstruction. In this paper, we explore an improved stereo matching network, SLED-Net, in which a Single Long Encoder-Decoder is proposed to replace the stacked hourglass network in PSM-Net for better contextual information learning. We compare SLED-Net to state-of-the-art methods recently published, and demonstrate its superior performance on Scene Flow and KITTI2015 test sets.
Guang-Yu Nie, Yun Liu 0011, Yongtian Wang, Yue Liu 0005
VR4
2019 Building AR-based Optical Experiment Applications in a VR Course
abstract
The demand for VR courses in universities is growing since VR technology is widely used in industry, entertainment, and education today. Traditional lectures and exercises have problems in motivating and engaging students, especially non-CS-majors. We design a VR course with a teamwork project assignment for Opto-Electronic Engineering undergraduates, providing them with project-based learning (PBL) experience. The task requires three students to group a team to build an AR-based optical experiment application over the course, aiming to develop students' practical engineering ability. The process of the project consists of three main stages: preparation, designing and implementing. We also evaluate students' work from different aspects and survey to analyze the students' attitude toward the project.
Huan Wei, Yue Liu 0005, Yongtian Wang
VR3
2019 Symmetrical Reality: Toward a Unified Framework for Physical and Virtual Reality
abstract
In this paper, we review the background of physical reality, virtual reality, and some traditional mixed forms of them. Based on the current knowledge, we propose a new unified concept called symmetrical reality to describe the physical and virtual world in a unified perspective. Under the framework of symmetrical reality, the traditional virtual reality, augmented reality, inverse virtual reality, and inverse augmented reality can be interpreted using a unified presentation. We analyze the characteristics of symmetrical reality from two different observation locations (i.e., from the physical world and from the virtual world), where all other forms of physical and virtual reality can be treated as special cases of symmetrical reality.
Zhenliang Zhang 0002, Dongdong Weng, Yue Liu 0005, Yongtian Wang
VR5
2019 Analyzing the Usability of Gesture Interaction in Virtual Driving System
abstract
In this study, an experiment is presented aiming at verifying the applicability of gesture interaction in the virtual driving environment. 30 participants are recruited to perform the secondary tasks with gesture and touch interaction. The task completion rate and reaction time of two interaction modalities under different road conditions are adopted as evaluation indexes. In addition, visual attention, NASA-TLX, and subjective questionnaire are collected as evaluation factors for fuzzy comprehensive evaluation based on entropy to evaluate the usability gestures. The research results show that gesture interaction not only shows excellence in safety, but also favors more than 90% of users.
Yue Liu 0005, Yongtian Wang
VR3
2019 Prioritizing candidate diseases-related metabolites based on literature and functional similarity
abstract
BACKGROUND: As the terminal products of cellular regulatory process, functional related metabolites have a close relationship with complex diseases, and are often associated with the same or similar diseases. Therefore, identification of disease related metabolites play a critical role in understanding comprehensively pathogenesis of disease, aiming at improving the clinical medicine. Considering that a large number of metabolic markers of diseases need to be explored, we propose a computational model to identify potential disease-related metabolites based on functional relationships and scores of referred literatures between metabolites. First, obtaining associations between metabolites and diseases from the Human Metabolome database, we calculate the similarities of metabolites based on modified recommendation strategy of collaborative filtering utilizing the similarities between diseases. Next, a disease-associated metabolite network (DMN) is built with similarities between metabolites as weight. To improve the ability of identifying disease-related metabolites, we introduce scores of text mining from the existing database of chemicals and proteins into DMN and build a new disease-associated metabolite network (FLDMN) by fusing functional associations and scores of literatures. Finally, we utilize random walking with restart (RWR) in this network to predict candidate metabolites related to diseases. RESULTS: We construct the disease-associated metabolite network and its improved network (FLDMN) with 245 diseases, 587 metabolites and 28,715 disease-metabolite associations. Subsequently, we extract training sets and testing sets from two different versions of the Human Metabolome database and assess the performance of DMN and FLDMN on 19 diseases, respectively. As a result, the average AUC (area under the receiver operating characteristic curve) of DMN is 64.35%. As a further improved network, FLDMN is proven to be successful in predicting potential metabolic signatures for 19 diseases with an average AUC value of 76.03%. CONCLUSION: In this paper, a computational model is proposed for exploring metabolite-disease pairs and has good performance in predicting potential metabolites related to diseases through adequate validation. This result suggests that integrating literature and functional associations can be an effective way to construct disease associated metabolite network for prioritizing candidate diseases-related metabolites.
Yongtian Wang, Liran Juan, Jiajie Peng, Tianyi Zang, Yadong Wang 0001
BMC Bioinform.1
2019 LncDisAP: a computation model for LncRNA-disease association prediction based on multiple biological datasets
abstract
BACKGROUND: Over the past decades, a large number of long non-coding RNAs (lncRNAs) have been identified. Growing evidence has indicated that the mutation and dysregulation of lncRNAs play a critical role in the development of many complex human diseases. Consequently, identifying potential disease-related lncRNAs is an effective means to improve the quality of disease diagnostics and treatment, which is the motivation of this work. Here, we propose a computational model (LncDisAP) for potential disease-related lncRNA identification based on multiple biological datasets. First, the associations between lncRNA and different data sources are collected from different databases. With these data sources as dimensions, we calculate the functional associations between lncRNAs by the recommendation strategy of collaborative filtering. Subsequently, a disease-associated lncRNA functional network is built with functional similarities between lncRNAs as the weight. Ultimately, potential disease-related lncRNAs can be identified based on ranked scores derived by random walking with restart (RWR). Then, training sets and testing sets are extracted from two different versions of a disease-lncRNA dataset to assess the performance of LncDisAP on 54 diseases. RESULTS: A lncRNA functional network is built based on the proposed computational model, and it contains 66,060 associations among 364 lncRNAs associated with 182 diseases in total. We extract 218 known disease-lncRNA pairs associated with 54 diseases to assess the network. As a result, the average AUC (area under the receiver operating characteristic curve) of LncDisAP is 78.08%. CONCLUSION: In this article, a computational model integrating multiple lncRNA-related biological datasets is proposed for identifying potential disease-related lncRNAs. The result shows that LncDisAP is successful in predicting novel disease-related lncRNA signatures. In addition, with several common cancers taken as case studies, we found some unknown lncRNAs that could be associated with these diseases through our network. These results suggest that this method can be helpful in improving the quality for disease diagnostics and treatment.
Yongtian Wang, Liran Juan, Jiajie Peng, Tianyi Zang, Yadong Wang 0001
BMC Bioinform.1
2019 A mobilized automatic human body measure system using neural network
Likun Xia, Jian Yang 0009, Huiming Xu, Yitian Zhao, Yongtian Wang
Multim. Tools Appl.7
2019 Patch-Based Adaptive Background Subtraction for Vascular Enhancement in X-Ray Cineangiograms
abstract
OBJECTIVE: Automatic vascular enhancement in X-ray cineangiography is of crucial interest, for instance, for better visualizing and quantifying coronary arteries in diagnostic and interventional procedures. METHODS: A novel patch-based adaptive background subtraction method (PABSM) is proposed automatically enhancing vessels in coronary X-ray cineangiography. First, pixels in the cineangiogram are described by the vesselness and Gabor features. Second, a classifier is utilized to separate the cineangiogram into the rough vascular and non-vascular region. Dilation is applied to the classified binary image to include more vascular region. Third, a patch-based background synthesis is utilized to fill the removed vascular region. RESULTS: A database containing 320 cineangiograms of 175 patients was collected, and then an interventional cardiologist annotated all vascular structures. The performance of PABSM is compared with six state-of-the-art vascular enhancement methods regarding the precision-recall curve and C-value. The area under the precision-recall curve is 0.7133, and the C-value is 0.9659. CONCLUSION: PABSM can automatically enhance the coronary artery in the cineangiograms. It preserves the integrity of vascular topological structures, particularly in complex vascular regions, and removes noise caused by the non-uniform gray-level distribution in the cineangiogram. SIGNIFICANCE: PABSM can avoid the motion artifacts and it eases the subsequent vascular segmentation, which is crucial for the diagnosis and interventional procedures of coronary artery diseases.
Shuang Song 0005, Alejandro F. Frangi, Jian Yang 0009, Danni Ai, Chenbing Du, Yong Huang 0002, Hong Song 0003, Luosha Zhang, Yechen Han, Yongtian Wang
IEEE J. Biomed. Health Informatics10
2019 On the launch of Virtual Reality & Intelligent Hardware
Yongtian Wang
Virtual Real. Intell. Hardw.1
2018 Inter/Intra-Constraints Optimization for Fast Vessel Enhancement in X-ray Angiographic Image Sequence
Chenbing Du, Shuang Song 0005, Danni Ai, Hong Song 0003, Yong Huang 0002, Yongtian Wang, Jian Yang 0009
BIBM6
2018 DeepDNA: a hybrid convolutional and recurrent neural network for compressing human mitochondrial genomes
Yan-Shuo Chu, Yongtian Wang, Mingrui Sun, Junyi Li 0004, Tianyi Zang, Yadong Wang 0001
BIBM5
2018 Identifying Candidate Diseases-related Metabolites Based on Disease Similarity
Yongtian Wang, Liran Juan, Chunpu Liu, Tianyi Zang
BIBM1
2018 Predicting candidate disease-related lncRNAs based on network random walk
Yongtian Wang, Liran Juan, Jiajie Peng, Tianyi Zang, Yadong Wang 0001
BIBM1
2018 Stereo Generation from a Single Image Using Deep Residual Network
abstract
In this paper, we propose a framework to generate stereoscopic content from a single image using the relative depth label predicted from deep residual network. Specifically, our framework first obtains a coarse relative depth label from the network and refines it to painting depth by sampling and interpolation, then an unsupervised clustering algorithm is employed to separate pixels of different depths into different layers to generate stereoscopic images. Experimental results with good visual effects demonstrate that the proposed method can be generally applied in both outdoor and indoor scenes. Meanwhile the quantitative results on relative depth estimation from a single image are comparable to state-of-the-art. Further experiments show the application possibility of our method in VR and panorama.
Tianteng Bi, Yue Liu 0005, Yongtian Wang
ICIP4
2018 Inverse Virtual Reality: Intelligence-Driven Mutually Mirrored World
abstract
Since artificial intelligence has been integrated into virtual reality, a new branch of virtual reality, which is called inverse virtual reality (IVR), is created. A typical IVR system contains both the intelligence-driven virtual reality and the physical reality, thus constructing an intelligence-driven mutually mirrored world. We propose the concept of IVR, and describe the details about the definition, structure and implementation of a typical IVR system. The parallel living environment is proposed as a typical application of IVR, which reveals that IVR has a significant potential to extend the human living environment.
Zhenliang Zhang 0002, Benyang Cao, Jie Guo 0004, Dongdong Weng, Yue Liu 0005, Yongtian Wang
VR6
2018 Evaluation of Hand-Based Interaction for Near-Field Mixed Reality with Optical See-Through Head-Mounted Displays
abstract
Hand-based interaction is one of the most widely-used interaction modes in the applications based on optical see-through head-mounted displays (OST-HMDs). In this paper, such interaction modes as gesture-based interaction (GBI) and physics-based interaction (PBI) are developed to construct a mixed reality system to evaluate the advantages and disadvantages of different interaction modes for near-field mixed reality. The experimental results show that PBI leads to a better performance of users regarding their work efficiency in the proposed tasks. The statistical analysis of T-test has been adopted to prove that the difference of efficiency between different interaction modes is significant.
Zhenliang Zhang 0002, Benyang Cao, Dongdong Weng, Yue Liu 0005, Yongtian Wang, Hua Huang 0001
VR5
2018 Physics-Inspired Input Method for Near-Field Mixed Reality Applications Using Latent Active Correction
abstract
Calibration accuracy is one of the most important factors to affect the user experience in mixed reality applications. For a typical mixed reality system built with the optical see-through head-mounted display (OST-HMD), a key problem is how to guarantee the accuracy of hand-eye coordination by decreasing the instability of the eye and the HMD in long-term use. In this paper, we propose a real-time latent active correction (LAC) algorithm to decrease hand-eye calibration errors accumulated over time. Experimental results show that we can successfully use the LAC algorithm to physics-inspired virtual input methods.
Zhenliang Zhang 0002, Dongdong Weng, Yue Liu 0005, Yongtian Wang
VR5
2018 Local statistical deformation models for deformable image registration
Songyuan Tang, Weijian Cong, Jian Yang 0009, Tianyu Fu 0003, Hong Song 0003, Danni Ai, Yongtian Wang
Neurocomputing7
2018 Automatic 2-D/3-D Vessel Enhancement in Multiple Modality Images Using a Weighted Symmetry Filter
abstract
Automated detection of vascular structures is of great importance in understanding the mechanism, diagnosis, and treatment of many vascular pathologies. However, automatic vascular detection continues to be an open issue because of difficulties posed by multiple factors, such as poor contrast, inhomogeneous backgrounds, anatomical variations, and the presence of noise during image acquisition. In this paper, we propose a novel 2-D/3-D symmetry filter to tackle these challenging issues for enhancing vessels from different imaging modalities. The proposed filter not only considers local phase features by using a quadrature filter to distinguish between lines and edges, but also uses the weighted geometric mean of the blurred and shifted responses of the quadrature filter, which allows more tolerance of vessels with irregular appearance. As a result, this filter shows a strong response to the vascular features under typical imaging conditions. Results based on eight publicly available datasets (six 2-D data sets, one 3-D data set, and one 3-D synthetic data set) demonstrate its superior performance to other state-of-the-art methods.
Yitian Zhao, Yalin Zheng, Yonghuai Liu, Yifan Zhao 0001, Lingling Luo, Tong Na, Yongtian Wang, Jiang Liu 0001
IEEE Trans. Medical Imaging8
2017 Pre-SCNAClonal: Efficient GC bias correction for SCNA based tumor subclonal populations inferring
abstract
Somatic copy number alternations (SCNAs) can be utilized to infer tumor subclonal populations in whole genome seuqncing studies, where usually their read count ratios between tumor-normal paired samples serve as the inferring proxy. We found that, in a GC study, the GC contents and read count ratios on SCNA segments present a Log linear biased pattern. However, currently no subclonal inferring tools take into account this information. We provide Pre-SCNAClonal, a comprehensive GC bias correction tool for inferring tumor subclonal populations based on SCNAs. Results show that Pre-SCNAClonal could effectively and robustly correct the GC bias and improve the performance of the SCNAs based tumor subclonal population inferring tools. Pre-SCNAClonal could be strung together with the SCNAs based subclonal population inferring tool as a pipeline or run individually as needed.
Yan-Shuo Chu, Mingxiang Teng, Yongtian Wang, Yadong Wang 0001
BIBM4
2017 A framework for analyzing DNA methylation data from Illumina Infinium HumanMethylation450 BeadChip
abstract
DNA methylation has been identified to be widely associated to complex diseases. Among biological platforms to profile DNA methylation in human, the Illumina Infinium HumanMethylation450 BeadChip (450K) has been accepted as one of the most efficient technologies. However, challenges exist in analysis of DNA methylation data generated by this technology due to widespread biases. Here we proposed a generalized framework for evaluating data analysis methods for Illumina 450K array. This framework considers the following steps towards a successful analysis: importing data, quality control, within-array normalization, correcting type bias, detecting differentially methylated probes or regions and biological interpretation. We evaluated five methods using three real datasets, and proposed outperform methods for the Illumina 450K array data analysis.
Yan-Shuo Chu, Yongtian Wang
BIBM3
2017 FNSemSim: An improved disease similarity method based on network fusion
abstract
Discovering similar diseases is very helpful for revealing the pathogenesis of diseases and making direction in drug use. And related diseases are often triggered by disease-related genes. Therefore, function interaction networks structured by disease-related genes are suitable for measurement of disease similarity, and some methods have utilized the advantage of function interaction of disease-related genes. However, all of them were developed by using a single gene functional network, some of them ignoring the effect of non-neighbour nodes in a functional interaction network. In this study, we propose a new method, FNSemSim, for computing relatedness between diseases by fusing two protein networks, which could be utilized fully based on random walk with restart (RWR). And a benchmark set of similar disease pairs are used to assess the performance of FNSemSim. As a result, FNSemSim achieved a very good performance with a high AUC (area under the receiver operating characteristic curve) reached 98.7%. Furthermore, we further studied the impact of different data sources, including function interaction networks and disease-related genes databases. It was found that the quality of the data sources has a greater impact on the performance of disease similarity calculation than the size of the data source, and utilizing function interaction networks and gene-disease association data could improve the performance of FNSemSim.
Yongtian Wang, Liran Juan, Yan-Shuo Chu, Tianyi Zang
BIBM1
2017 Effects of using HMDs on visual fatigue in virtual environments
abstract
There are few negative effects to make people discomfort using virtual reality systems. In this paper, we investigated the effects of visual fatigue when wearing head-mounted displays (HMD) and compared the results with those from the smartphones. Forty subjects were recruited and divided into two different groups. The visual fatigue scale was measured to assess the subjects' performance. The results indicated that visual fatigue caused by the conflict of focal distance and vergence distance was less severe than visual fatigue caused by long-term focus without accommodation.
Jie Guo 0004, Dongdong Weng, Henry Been-Lirn Duh, Yue Liu 0005, Yongtian Wang
VR5
2017 Evaluation of labelling layout methods in augmented reality
abstract
View management techniques are commonly used for labelling of objects in augmented reality environments. Combining with image analysis, search space and adaptive representations, they can be utilized to achieve desired labelling tasks. However, the evaluation of different search space methods on labelling are still an open problem. In this paper, we propose an image analysis based view management method, which first adopts the image processing to superimpose 2D labels to the specific object. We then conduct three search space methods to an augmented reality scenario. Without the requirements of setting rules and constraints for occlusion among the labels, the results of three search space methods are evaluated by using objective analysis of related parameters. The evaluation results indicate that different search space methods could generate different time costs and occlusion, thereby affecting the final labelling effects.
Yue Liu 0005, Yongtian Wang
VR3
2017 RIDE: Region-induced data enhancement method for dynamic calibration of optical see-through head-mounted displays
abstract
The most commonly used single point active alignment method (SPAAM) is based on a static pinhole camera model, in which it is assumed that both the eye and the HMD are fixed. This leads to a limitation for calibration precision. In this work, we propose a dynamic pinhole camera model according to the fact that the human eye would experience an obvious displacement over the whole calibration process. Based on such a camera model, we propose a new calibration data acquisition method called the region-induced data enhancement (RIDE) to revise the calibration data. The experimental results prove that the proposed dynamic model performs better than the traditional static model in actual calibration.
Zhenliang Zhang 0002, Dongdong Weng, Yue Liu 0005, Yongtian Wang, Xinjun Zhao
VR4
2017 Saliency driven vasculature segmentation with infinite perimeter active contour model
Yitian Zhao, Jingliang Zhao, Jian Yang 0009, Yonghuai Liu, Yifan Zhao 0001, Yalin Zheng, Likun Xia, Yongtian Wang
Neurocomputing8
2017 Registration and fusion quantification of augmented reality based nasal endoscopic surgery
Yakui Chu, Jian Yang 0009, Shaodong Ma, Danni Ai, Hong Song 0003, Duanduan Chen, Lei Chen 0073, Yongtian Wang
Medical Image Anal.10
2017 Fast hand posture classification using depth features extracted from random line segments
Weizhi Nai, Yue Liu 0005, David Rempel, Yongtian Wang
Pattern Recognit.4
2017 Intensity and Compactness Enabled Saliency Estimation for Leakage Detection in Diabetic and Malarial Retinopathy
abstract
Leakage in retinal angiography currently is a key feature for confirming the activities of lesions in the management of a wide range of retinal diseases, such as diabetic maculopathy and paediatric malarial retinopathy. This paper proposes a new saliency-based method for the detection of leakage in fluorescein angiography. A superpixel approach is firstly employed to divide the image into meaningful patches (or superpixels) at different levels. Two saliency cues, intensity and compactness, are then proposed for the estimation of the saliency map of each individual superpixel at each level. The saliency maps at different levels over the same cues are fused using an averaging operator. The two saliency maps over different cues are fused using a pixel-wise multiplication operator. Leaking regions are finally detected by thresholding the saliency map followed by a graph-cut segmentation. The proposed method has been validated using the only two publicly available datasets: one for malarial retinopathy and the other for diabetic retinopathy. The experimental results show that it outperforms one of the latest competitors and performs as well as a human expert for leakage detection and outperforms several state-of-the-art methods for saliency detection.
Yitian Zhao, Yalin Zheng, Yonghuai Liu, Jian Yang 0009, Yifan Zhao 0001, Duanduan Chen, Yongtian Wang
IEEE Trans. Medical Imaging7
2017 Convex Hull Aided Registration Method (CHARM)
abstract
Non-rigid registration finds many applications such as photogrammetry, motion tracking, model retrieval, and object recognition. In this paper we propose a novel convex hull aided registration method (CHARM) to match two point sets subject to a non-rigid transformation. First, two convex hulls are extracted from the source and target respectively. Then, all points of the point sets are projected onto the reference plane through each triangular facet of the hulls. From these projections, invariant features are extracted and matched optimally. The matched feature point pairs are mapped back onto the triangular facets of the convex hulls to remove outliers that are outside any relevant triangular facet. The rigid transformation from the source to the target is robustly estimated by the random sample consensus (RANSAC) scheme through minimizing the distance between the matched feature point pairs. Finally, these feature points are utilized as the control points to achieve non-rigid deformation in the form of thin-plate spline of the entire source point set towards the target one. The experimental results based on both synthetic and real data show that the proposed algorithm outperforms several state-of-the-art ones with respect to sampling, rotational angle, and data noise. In addition, the proposed CHARM algorithm also shows higher computational efficiency compared to these methods.
Jingfan Fan, Jian Yang 0009, Yitian Zhao, Danni Ai, Yonghuai Liu, Ge Wang 0001, Yongtian Wang
IEEE Trans. Vis. Comput. Graph.7
2016 Wavefront modulation by complex amplitude in holographic display
abstract
Complex amplitude modulation is an important technology to control the wavefront arbitrary, which is widely used in many optics field. It is reviewed in this paper, and the basic principles are introduced and the features of the methods are discussed. Though there are still some disadvantages in these modulation methods, because of offering the more flexibilities, it is believed that complex amplitude modulation plays more significant role in optics.
Yongtian Wang
INDIN3
2016 A tour guiding system of historical relics based on augmented reality
abstract
Yuanmingyuan is a relic park and only few cultural relics are left due to the looting and burning down in history, which makes that most of the scenic spots of the park look boring. To address such issue, a game-based guidance system for Yuanmingyuan and a time travel game called MAGIC-EYES has been proposed with Augmented Reality technology. Six interactive modes are designed in the proposed system to guide tourists to visit the specified place. The evaluation results of a pilot study shows that the proposed guidance system has significantly improved the tourist experiences.
Dongdong Weng, Yue Liu 0005, Yongtian Wang
VR4
2016 3-Points Convex Hull Matching (3PCHM) for fast and robust point set registration
Jingfan Fan, Jian Yang 0009, Feng Lu 0005, Danni Ai, Yitian Zhao, Yongtian Wang
Neurocomputing6
2016 Shape context and projection geometry constrained vasculature matching for 3D reconstruction of coronary artery
Ruoxiu Xiao, Jian Yang 0009, Jingfan Fan, Danni Ai, Guangzhi Wang, Yongtian Wang
Neurocomputing6
2016 Local statistics and non-local mean filter for speckle noise reduction in medical ultrasound image
Jian Yang 0009, Jingfan Fan, Danni Ai, Xuehu Wang, Yongchang Zheng, Songyuan Tang, Yongtian Wang
Neurocomputing7
2016 Region-based saliency estimation for 3D shape analysis and understanding
Yitian Zhao, Yonghuai Liu, Baogang Wei, Jian Yang 0009, Yifan Zhao 0001, Yongtian Wang
Neurocomputing7
2016 Convex hull indexed Gaussian mixture model (CH-GMM) for 3D point set registration
Jingfan Fan, Jian Yang 0009, Danni Ai, Likun Xia, Yitian Zhao, Yongtian Wang
Pattern Recognit.7
2016 A Review of Dynamic Holographic Three-Dimensional Display: Algorithms, Devices, and Systems
abstract
Dynamic holographic three-dimensional (3-D) display, which features reconstructing 3-D images with full depth cues has great potential in various fields such as medical, and military industries obtained a broad attention in the last decades. Combing parallel computation techniques, hologram synthesis algorithms are capable of calculating diffraction and interference pattern dynamically or even in real-time manner. In addition, the development of various types of optical modulators including liquid crystal (LC)-based, opto-acoustic-typed, digital micromirror devices (DMDs), and opto-opto materials made it possible that device can load and display such holographic patterns dynamically. In accordance with the progress of algorithms and modulators, we can expect that dynamic holographic 3-D display will be able to reconstruct 3-D images with real-time refresh rate, full color, wide viewing angle, and large image size, and the system will be commercialized in a near future.
Yijie Pan, Yongtian Wang
IEEE Trans. Ind. Informatics4
2016 Recognizing Focal Liver Lesions in CEUS With Dynamically Trained Latent Structured Models
abstract
This work investigates how to automatically classify Focal Liver Lesions (FLLs) into three specific benign or malignant types in Contrast-Enhanced Ultrasound (CEUS) videos, and aims at providing a computational framework to assist clinicians in FLL diagnosis. The main challenge for this task is that FLLs in CEUS videos often show diverse enhancement patterns at different temporal phases. To handle these diverse patterns, we propose a novel structured model, which detects a number of discriminative Regions of Interest (ROIs) for the FLL and recognize the FLL based on these ROIs. Our model incorporates an ensemble of local classifiers in the attempt to identify different enhancement patterns of ROIs, and in particular, we make the model reconfigurable by introducing switch variables to adaptively select appropriate classifiers during inference. We formulate the model learning as a non-convex optimization problem, and present a principled optimization method to solve it in a dynamic manner: the latent structures (e.g. the selections of local classifiers, and the sizes and locations of ROIs) are iteratively determined along with the parameter learning. Given the updated model parameters in each step, the data-driven inference is also proposed to efficiently determine the latent structures by using the sequential pruning and dynamic programming method. In the experiments, we demonstrate superior performances over the state-of-the-art approaches. We also release hundreds of CEUS FLLs videos used to quantitatively evaluate this work, which to the best of our knowledge forms the largest dataset in the literature. Please find more information at "http://vision.sysu.edu.cn/projects/fllrecog/".
Xiaodan Liang, Liang Lin 0004, Qingxing Cao, Rui Huang 0001, Yongtian Wang
IEEE Trans. Medical Imaging5
2015 Deformable 3D Fusion: From Partial Dynamic 3D Observations to Complete 4D Models
abstract
Capturing the 3D motion of dynamic, non-rigid objects has attracted significant attention in computer vision. Existing methods typically require either complete 3D volumetric observations, or a shape template. In this paper, we introduce a template-less 4D reconstruction method that incrementally fuses highly-incomplete 3D observations of a deforming object, and generates a complete, temporally-coherent shape representation of the object. To this end, we design an online algorithm that alternatively registers new observations to the current model estimate and updates the model. We demonstrate the effectiveness of our approach at reconstructing non-rigidly moving objects from highly-incomplete measurements on both sequences of partial 3D point clouds and Kinect videos.
WeiPeng Xu, Mathieu Salzmann, Yongtian Wang, Yue Liu 0005
ICCV3
2015 Context-Aware Based Mobile Augmented Reality Browser and its Optimization Design
Yue Liu 0005, Yongtian Wang
ICIG (2)3
2015 Recent progress on holographic 3D display at Beijing Institute of Technology
abstract
Our recent progress in dynamic holographic 3D display is introduced, including compressive holographic 3D display, full-color holographic 3D display, medium for holographic 3D applications, occlusion culling algorithm and improved full analytical polygon-based method.
Gaolei Xue, Yongtian Wang
INDIN3
2015 Omnidirectional-view three-dimensional displays using multiple mini-projectors
abstract
We have developed omnidirectional-view 3D displays using multiple mini-projectors. Three types of synchronization structure are developed to ensure the accurate synchronization between the projectors and the rotating screen. Therefore, low-cost and low-speed display devices can be used to realize natural-looking three-dimensional (3D) scene with full color and high resolution.
Qiudong Zhu, Dongdong Weng, Yue Liu 0005, Yongtian Wang
VCIP5
2015 Family genome browser: visualizing genomes with pedigree information
abstract
MOTIVATION: Families with inherited diseases are widely used in Mendelian/complex disease studies. Owing to the advances in high-throughput sequencing technologies, family genome sequencing becomes more and more prevalent. Visualizing family genomes can greatly facilitate human genetics studies and personalized medicine. However, due to the complex genetic relationships and high similarities among genomes of consanguineous family members, family genomes are difficult to be visualized in traditional genome visualization framework. How to visualize the family genome variants and their functions with integrated pedigree information remains a critical challenge. RESULTS: We developed the Family Genome Browser (FGB) to provide comprehensive analysis and visualization for family genomes. The FGB can visualize family genomes in both individual level and variant level effectively, through integrating genome data with pedigree information. Family genome analysis, including determination of parental origin of the variants, detection of de novo mutations, identification of potential recombination events and identical-by-decent segments, etc., can be performed flexibly. Diverse annotations for the family genome variants, such as dbSNP memberships, linkage disequilibriums, genes, variant effects, potential phenotypes, etc., are illustrated as well. Moreover, the FGB can automatically search de novo mutations and compound heterozygous variants for a selected individual, and guide investigators to find high-risk genes with flexible navigation options. These features enable users to investigate and understand family genomes intuitively and systematically. AVAILABILITY AND IMPLEMENTATION: The FGB is available at http://mlg.hit.edu.cn/FGB/.
Liran Juan, Yongzhuang Liu, Yongtian Wang, Mingxiang Teng, Tianyi Zang, Yadong Wang 0001
Bioinform.3
2014 Nonrigid Surface Registration and Completion from RGBD Images
WeiPeng Xu, Mathieu Salzmann, Yongtian Wang, Yue Liu 0005
ECCV (2)3
2013 Rigid registration of 3-D medical image using convex hull matching
abstract
In this paper, a robust approach called convex hull matching (CHM) technique is proposed for registration of medical images that differ from each other with Euclidean transformation. Firstly, point sets on the surface of the medical image are extracted, and then the 3-D convex hull is constructed from the point sets and triangle patches on the surface of convex hulls are specified by predefining their normal vectors. Secondly, each edge of the referenced triangle is compared with all the edges of the triangle in other point set to find the congruent pair set and also to obtain the scaling factor. Thereafter, the transformation parameters of each triangle pairs including rotation and translation are optimized by minimizing the Euclidian distance between the corresponding vertex pairs. Hence, rigid transformation of the two point sets is obtained by iteratively enumerating and evaluating similarity measures of the triangle patches chosen. Global optimization is achieved through RANSAC optimization by removing the correspondence pairs that may lead to large matching errors of the whole point sets. The experiments evaluate the performance of the proposed algorithm on simulated data with the presence of outliers and noise. The results show the efficiency of CHM by quantitative analysis and comparative study with existing approaches like EM-ICP, LM-ICP and 4PCS. Finally, the real clinical data experiments confirm the proposed algorithm is a strong performer in medical image registration.
Jingfan Fan, Jian Yang 0009, Mahima Goyal, Yongtian Wang
BIBM4
2013 Real-time keystone correction for hand-held projectors with an RGBD camera
abstract
This paper introduces a novel and simple approach to realtime continuous keystone correction for hand-held projectors. An RGBD camera is attached to the projector to form a projector-RGBD-camera system. The system is first calibrated in an offline stage. At run-time, we then estimate the relative pose between the projector and the screen using the RGBD camera, which lets us correct the keystone distortion by warping the projected image accordingly. Experimental results show that our method outperforms existing techniques in terms of both accuracy and efficiency.
WeiPeng Xu, Yongtian Wang, Yue Liu 0005, Dongdong Weng, Mengwen Tan, Mathieu Salzmann
ICIP2
2013 Outdoor scenes identification on mobile device by integrating vision and inertial sensors
abstract
This paper addresses the identification of large scale outdoor scenes on smart phone by fusing outputs of inertial sensors and computer vision techniques. The main contributions can be summarized as follows: Firstly, we propose an overlap region divide (ORD) method to plot image position area, which is fast enough to find the nearest visiting area and can also reduce the search range compared with the traditional approaches. Secondly, the vocabulary tree based approach is improved by introducing fast geometric consistency constraints (FGCC). Our method involves no operation in the high-dimensional feature space and does not assume a global transform between a pair of images. Thus, it substantially reduces the computational complexity and memory usage, which makes the city scale image recognition feasible on the smartphone. Experiments on a collected database including 0.16 million images show that the proposed method demonstrates excellent identification performance, while maintaining the average identification time of less than 1s.
Zhenwen Gui, Yongtian Wang, Yue Liu 0005, Jing Chen 0018
IWCMC2
2013 Robust structure from motion with affine camera via low-rank matrix recovery
abstract
Abstract We present a novel approach to structure from motion that can deal with missing data and outliers with an affine camera. We model the corruptions as sparse error. Therefore the structure from motion problem is reduced to the problem of recovering a low-rank matrix from corrupted observations. We first decompose the matrix of trajectories of features into low-rank and sparse components by nuclear-norm and ℓ 1-norm minimization, and then obtain the motion and structure from the low-rank components by the classical factorization method. Unlike pervious methods, which have some drawbacks such as depending on the initial value selection and being sensitive to the large magnitude errors, our method uses a convex optimization technique that is guaranteed to recover the low-rank matrix from highly corrupted and incomplete observations. Experimental results demonstrate that the proposed approach is more efficient and robust to large-scale outliers.
Lun Wu, Yongtian Wang, Yue Liu 0005, Yuxi Wang 0002
Sci. China Inf. Sci.2
2013 Auto learning temporal atomic actions for activity classification
Jiangen Zhang, Benjamin Z. Yao, Yongtian Wang
Pattern Recognit.3
2012 Modelling Atomic Actions for Activity Classification
abstract
In this paper, we present a model for learning atomic actions for complex activities classification. A video sequence is first represented by a collection of visual interest points. The model automatically clusters visual words into atomic actions based on their co-occurrence and temporal proximity using an extension of Hierarchical Dirichlet Process (HDP) mixture model. Our approach is robust to noisy interest points caused by various conditions because HDP is a generative model. Based on the atomic actions learned from our model, we use both a Naive Bayesian and a linear SVM classifier for activity classification. We first use a synthetic example to demonstrate the intermediate result, then we apply on the complex Olympic Sport 16-class dataset and show that our model outperforms other state-of-art methods.
Jiangen Zhang, Benjamin Z. Yao, Yongtian Wang
ICME3
2012 A modified KLT multiple objects tracking framework based on global segmentation and adaptive template
Kang Xue, Patricio A. Vela, Yue Liu 0005, Yongtian Wang
ICPR4
2012 Reconfigurable templates for robust vehicle detection and classification
abstract
In this paper, we learn a reconfigurable template for detecting vehicles and classifying their types. We adopt a popular design for the part based model that has one coarse template covering entire object window and several small high-resolution templates representing parts. The reconfigurable template can learn part configurations that capture the spatial correlation of features for a deformable part based model. The features of templates are Histograms of Gradients (HoG). In order to better describe the actual dimensions and locations of “parts” (i.e. features with strong spatial correlations), we design a dictionary of rectangular primitives of various sizes, aspect-ratios and positions. A configuration is defined as a subset of non-overlapping primitives from this dictionary. To learn the optimal configuration using SVM amounts, we need to find the subset of parts that minimize the regularized hinge loss, which leads to a non-convex optimization problem. We solve this problem by replacing the hinge loss with a negative sigmoid loss that can be approximately decomposed into losses (or negative sigmoid scores) of individual parts. In the experiment, we compare our method empirically with group lasso and a state of the art method [7] and demonstrate that models learned with our method outperform others on two computer vision applications: vehicle localization and vehicle model recognition.
Benjamin Z. Yao, Yongtian Wang, Song-Chun Zhu
WACV3
2012 Application of ICA to X-ray coronary digital subtraction angiography
Songyuan Tang, Yongtian Wang, Yen-Wei Chen 0001
Neurocomputing2
2012 A completely affine invariant image-matching method based on perspective projection
Yongtian Wang, Jing Chen 0018, Junwei Guo
Mach. Vis. Appl.2
2012 Object categorization with sketch representation and generalized samples
Liang Lin 0004, Xiaobai Liu, Shaowu Peng, Hongyang Chao, Yongtian Wang, Bo Jiang 0002
Pattern Recognit.5
2011 Application of Pen-Based Planar Haptic Interface in Physics Education
abstract
This paper proposes a pen-based interaction system with 2D co-located haptic and visual feedback for physics education. The system combines a 3DOF pen-based planar haptic device and simulated 2D physical world based on physical engine Box 2D and 2D game engine HGE. The inputs of haptic device are translation and rotation of pen tip and the outputs are translational and rotational force displayed at pen tip. In addition, visual display and haptic display are coincident and well integrated. With this system, user can create and design simulated physical world by drawing, selecting, moving or rotating objects on screen and setting their physical properties in a natural and intuitive way. This system provides an entertaining and cartoony tool for designing and carrying out physics experiments, and will greatly promote students' learning interest and creativity.
Liping Lin, Yongtian Wang, Yue Liu 0005, Makoto Sato
CAD/Graphics2
2011 Segmentation Based on Routing Image Algorithms
abstract
This paper presents a novel image segmentation method in which energy function is based on global region information while not only on edge information. Image segmentation can be viewed as a routing problem. In order to obtain the optimal segmentation, the Shortest Path Faster Algorithm (SPFA) is used to optimize the discrete grid energy function. As the commonly used Live-Wire algorithm is easy to obtain mistake segmentation when the strong edges and the weak edges are close to each other, the interactive segmentation method is proposed for the precise boundaries estimation. The developed method has been tested on both clinical medical images and natural scene images. It can be seen that the developed method is very fast and effective, and can obtain good segmentation results.
Hongzhe Yang, Jian Yang 0009, Yongtian Wang, Yue Liu 0005
ICIG3
2011 PTZ camera-based adaptive panoramic and multi-layered background model
abstract
In this paper, we present a novel approach for constructing an adaptive panoramic and multi-layered background model for Pan-tilt-zoom (PTZ) camera that provides fast registration of the observed frame and localizes the foreground targets with arbitrary camera position and scale (optical zoom). Our method consists of two stages. (1) An adaptive panoramic background mixture model is generated off-line for foreground detection. (2) A layered correspondence is generated off-line from frames captured at different optical zoom values of the camera, and a correspondence propagation method is used to register the observed frame with the panoramic background online. We demonstrate the advantages of the proposed adaptive panoramic and multi-layered background model within wide field of view (FOV) and over large scale range.
Kang Xue, Gbolabo Ogunmakin, Yue Liu 0005, Patricio A. Vela, Yongtian Wang
ICIP5
2011 "Soul Hunter": A novel augmented reality application in theme parks
abstract
This paper introduces a novel augmented reality shooting game named “Soul Hunter”, which has been successfully operating in a theme park in China. Soul Hunter adopts an innovative infrared marker scheme to build a mobile augmented reality application in a wide area. It is an extension of the traditional first person game, in which a player is able to fight with virtual ghost through a gunlike device in real environment. This paper describes the challenges of applying augmented reality in theme parks and shares some experiences in solving the problems encountered in practical applications.
Dongdong Weng, WeiPeng Xu, Dong Li 0013, Yongtian Wang, Yue Liu 0005
ISMAR4
2011 Sensor fusion based head pose tracking for lightweight flight cockpit systems
Yongtian Wang, Yue Liu 0005
Multim. Tools Appl.2
2010 Robust Photometric Stereo via Low-Rank Matrix Completion and Recovery
Lun Wu, Arvind Ganesh, Boxin Shi, Yasuyuki Matsushita, Yongtian Wang, Yi Ma 0001
ACCV (3)5
2010 Augmented reality registration algorithm based on nature feature recognition
Jing Chen 0018, Yongtian Wang, Junwei Guo, Jingdun Lin, Kang Xue, Yue Liu 0005
Sci. China Inf. Sci.2
2009 Markerless tracking for augmented reality applied in reconstruction of Yuanmingyuan archaeological site
abstract
This paper presents an algorithm based on the method of supervised machine learning and multi-keyframes to achieve markerless augmented reality (AR) application when there is a locally planar object in the scene. The main goal is to solve the problem of AR tracking in outdoor environment by only using vision and natural features. Instead of tracking fiducial markers, we track natural keypoints, during which the point correspondences are established from the classification perspective. The tracking range is able to be extended by employing many reference images. These results in a promising algorithm that successfully tested on a touring guide system, which provides views of virtual original appearance superimposed to ruins of Yuanmingyuan archaeological site of China. Comparisons are also made between ARToolkit and the proposed algorithm in indoor environment. Experimental results demonstrated that our algorithm is characterized by fairly robustness and high time efficiency in both indoor and outdoor application.
Junwei Guo, Yongtian Wang, Jing Chen 0018, Jingdun Lin, Lun Wu, Kang Xue, Jiangen Zhang
CAD/Graphics2
2009 Key Issues of Wide-Area Tracking System for Multi-user Augmented Reality Adventure Game
abstract
Augmented reality (AR) Adventure Game is a wide-area indoor AR application for multiple users in Guangdong science center. Tracking the pose of users’ head in wide area is crucial for alignment between virtual and real scene. This paper studies the key issues of wide-area indoor tracking system. Different from previous inside-looking-out vision-based tracking, coded infrared (IR) markers are installed both on walls and ceiling in the proposed system. An automatic method is designed to calibrate the transformation between tracking camera and scene camera. The linear algorithm used for pose estimation is presented and problems of registration in rendering engine are also discussed. Experiments are conducted and applications in the AR Adventure Game prove that the proposed tracking system can provide precise and stable registration in actual systems.
Yetao Huang, Dongdong Weng, Yue Liu 0005, Yongtian Wang
ICIG4
2009 A Remote Control System Based on Real-Time Image Processing
abstract
A novel human-computer interaction (HCI) system based on real-time image processing is proposed in this paper. With the help of infrared tracking technology, the proposed system achieves real-time processing and stable operation on an ADSP-BF533 hardware platform. Compared with the conventional remote control methods, the proposed system enables a user to control the cursor on the screen of a TV by targeting it, and provides users with new experiences of remote control. To realize such a new remote control system, cross-ratio invariant, which is an important characteristic of the projective transformation, is also studied. Experimental results show the potential of the proposed system in TV remote control.
Yongtian Wang, Yue Liu 0005, Dongdong Weng, Xiaoming Hu 0001
ICIG2
2009 Sensor Fusion for Vision-Based Indoor Head Pose Tracking
abstract
Accurate head pose tracking is a key issue for indoor augmented reality systems. This paper proposes a novel approach to track head pose of indoor users using sensor fusion. The proposed approach utilizes a track-to-track fusion framework composed of extended Kalman filters and fusion filter to fuse the poses from the two complementary tracking modes of inside-out tracking (IOT) and outside-in tracking (OIT). A vision-based head tracker is constructed to verify our approach. Primary experimental results show that the tracker is capable of achieving more accurate and stable pose than the single tracking mode of IOT or OIT, which validates the usefulness of the proposed sensor fusion approach.
Yongtian Wang, Yue Liu 0005
ICIG2
2009 GPU Based Real-time Correction for Optical Distortions in Head-Mounted Displays
abstract
This paper presents a GPU-based real-time method to correct optical distortions in head-mounted displays (HMDs). The HMD to be corrected is a lightweight and wide field-of-view HMD system with free-form-surface (FFS) prism, in which the image distortion is not rectilinear and centrosymmetric. A special predistortion model is constructed to correct the distortion of the HMD. Although the distortion correction can be performed with an extensional optics system, the system will be too expensive and additional weight will be imposed to the HMD. With the method presented in this paper, each pixel in the original image is remapped to a new position with GPU and forms a predistortion image. The remapping process is based on a prior formatted distortion map, which is similar to the normal map used in the bump mapping process. The distortion map is an RGBA image that corresponds to the X and Y coordinates of a pixel offset from the original image. The remapping process accomplished by GPU is very fast. The performance of the proposed method is analyzed and validated via a demonstration system.
Dongdong Weng, Yongtian Wang, Yue Liu 0005
ICIG2
2009 Adaptive Real-Time Labeling and Recognition of Multiple Infrared Markers Using FPGA
abstract
In this paper, we propose a real-time and adaptive method for labeling and recognition of multiple infrared markers. A single-frame based iteration process is developed to obtain the suitable threshold. A sliding window is proposed to propagate the preliminary labels and to reduce the capacity of equivalent tables handling. The labeling and recognition are realized by merging the results of domains with label collisions. Multi-stage pipelines are developed to implement all the operations including smoothing filter, adaptive threshold, preliminary labeling and recognition. Experimental results show that the proposed method can label and recognize multiple infrared target markers with a latency of 259 ns and with an accuracy of sub-pixel. The proposed method can be applied in applications that require real-time performance such as surgery navigation, intelligent control and visual measurement.
Yue Liu 0005, Yongtian Wang
ICIG3
2009 Marker-less registration based on template tracking for augmented reality
Yongtian Wang, Yue Liu 0005, Caiming Xiong
Multim. Tools Appl.2
2009 Novel Approach for 3-D Reconstruction of Coronary Arteries From Two Uncalibrated Angiographic Images
abstract
Three-dimensional reconstruction of vessels from digital X-ray angiographic images is a powerful technique that compensates for limitations in angiography. It can provide physicians with the ability to accurately inspect the complex arterial network and to quantitatively assess disease induced vascular alterations in three dimensions. In this paper, both the projection principle of single view angiography and mathematical modeling of two view angiographies are studied in detail. The movement of the table, which commonly occurs during clinical practice, complicates the reconstruction process. On the basis of the pinhole camera model and existing optimization methods, an algorithm is developed for 3-D reconstruction of coronary arteries from two uncalibrated monoplane angiographic images. A simple and effective perspective projection model is proposed for the 3-D reconstruction of coronary arteries. A nonlinear optimization method is employed for refinement of the 3-D structure of the vessel skeletons, which takes the influence of table movement into consideration. An accurate model is suggested for the calculation of contour points of the vascular surface, which fully utilizes the information in the two projections. In our experiments with phantom and patient angiograms, the vessel centerlines are reconstructed in 3-D space with a mean positional accuracy of 0.665 mm and with a mean back projection error of 0.259 mm. This shows that the algorithm put forward in this paper is very effective and robust.
Jian Yang 0009, Yongtian Wang, Yue Liu 0005, Songyuan Tang, Wufan Chen
IEEE Trans. Image Process.2
2008 An integrated background model for video surveillance based on primal sketch and 3D scene geometry
abstract
This paper presents a novel integrated background model for video surveillance. Our model uses a primal sketch representation for image appearance and 3D scene geometry to capture the ground plane and major surfaces in the scene. The primal sketch model divides the background image into three types of regions - flat, sketchable and textured. The three types of regions are modeled respectively by mixture of Gaussians, image primitives and LBP histograms. We calibrate the camera and recover important planes such as ground, horizontal surfaces, walls, stairs in the 3D scene, and use geometric information to predict the sizes and locations of foreground blobs to further reduce false alarms. Compared with the state-of-the-art background modeling methods, our approach is more effective, especially for indoor scenes where shadows, highlights and reflections of moving objects and camera exposure adjusting usually cause problems. Experiment results demonstrate that our approach improves the performance of background/foreground separation at pixel level, and the integrated video surveillance system at the object and trajectory level.
Wenze Hu, Haifeng Gong, Song-Chun Zhu, Yongtian Wang
CVPR4
2008 An interactive scene annotation tool for video surveillance
abstract
An interactive scene annotation tool for video surveillance is presented in this paper. The annotation process is divided into three stages. (1) camera rough calibration;(2) calibration refinement; (3) major surfaces annotation. Inputs are then rendered in a 3D environment, which again help users check calibration accuracy and annotation correctness. Experiments show that this tool is easy to use and attains acceptable annotation accuracy. The interactive procedure helps users without knowledge in computer vision to complete camera calibration as well as surface annotation.
Wenze Hu, Jianting Wen, Haifeng Gong, Yongtian Wang
ICPR4
2008 New approach to the automatic segmentation of coronary artery in X-ray angiograms
Shoujun Zhou, Wufan Chen, Yongtian Wang
Sci. China Ser. F Inf. Sci.4
2007 Study on an Indoor Tracking System Based on Primary and Assistant Infrared Markers
abstract
An indoor tracking system based on primary and assistant infrared markers is presented in this paper. The system can track the user's head in a large area with high stability and accuracy. And the price of the system is very low. The novel assistant infrared markers with particular spatial characteristic and the primary infrared markers with particular spatio-temporal characteristic are proposed, which can avoid the synchronization between infrared markers and user system. Various numbers of users can be supported by the proposed system and experimental result shows the effectiveness and robustness of the system.
Dongdong Weng, Yue Liu 0005, Yongtian Wang, Lun Wu
CAD/Graphics3
2007 Layered Graph Match with Graph Editing
abstract
Many vision tasks are posed as either graph partitioning (coloring) or graph matching (correspondence) problems. The former include segmentation and grouping, and the latter include wide baseline stereo, large motion, object tracking and recognition. In this paper, we present an integrated solution for both graph matching and graph partition using an effective sampling algorithm in a Bayesian framework. Given two images for matching, we extract two graphs using a primal sketch algorithm [4]. The graph nodes are linelets and primitives (junctions). Both graphs are automatically partitioned into an unknown number of K + 1 layers of subgraphs so that K pairs of subgraphs are matched and the remaining layer contains unmatched backgrounds. Each matched pair represent a "moving object" with a TPS (thin-plate-spline) transform to account for its deformations and a set of graph operators to edit the pair of subgraphs to achieve perfect structural match. The matching energy between two subgraphs includes geometric deformations, appearance dissimilarities, and the cost of graph editing operators. We demonstrate its application on two tasks: (i) large motion with occlusion, and (ii) automatic detection and recognition of common objects in a pair of images.
Liang Lin 0004, Song-Chun Zhu, Yongtian Wang
CVPR3
2007 An Empirical Study of Object Category Recognition: Sequential Testing with Generalized Samples
abstract
In this paper we present an empirical study of object category recognition using generalized samples and a set of sequential tests. We study 33 categories, each consisting of a small data set of 30 instances. To increase the amount of training data we have, we use a compositional object model to learn a representation for each category from which we select 30 additional templates with varied appearance from the training set. These samples better span the appearance space and form an augmented training set ΩTof 1980 (60×33) training templates. To perform recognition on a testing image, we use a set of sequential tests to project ΩTinto different representation spaces to narrow the number of candidate matches in ΩT. We use"graphlets"(structural elements), as our local features and model OmegaTat each stage using histograms of graphlets over categories, histograms of graphlets over object instances, histograms of pairs of graphlets over objects, shape context. Each test is increasingly computationally expensive, and by the end of the cascade we have a small candidate set remaining to use with our most powerful test, a top-down graph matching algorithm. We achieve an 81.4 % classification rate on classifying 800 testing images in 33 categories, 15.2% more accurate than a method without generalized samples.
Liang Lin 0004, Shaowu Peng, Jake Porway, Song-Chun Zhu, Yongtian Wang
ICCV5
2007 The Application of ICA to the X-Ray Digital Subtraction Angiography
Songyuan Tang, Yongtian Wang, Yen-Wei Chen 0001
ISNN (2)2
2007 Fiducial Marker Based on Projective Invariant for Augmented Reality
Yongtian Wang, Yue Liu 0005
J. Comput. Sci. Technol.2
2005 An Improved Colored-Marker Based Registration Method for AR Applications
Xiaowei Li 0004, Yue Liu 0005, Yongtian Wang, Dayuan Yan, Dongdong Weng
ICCSA (3)3
2005 Autocalibration of an Electronic Compass for Augmented Reality
abstract
Electronic compass is often used to provide the absolute heading reference for tracking the user's head and hands in virtual reality (VR) and augmented reality (AR), especially for outdoor AR applications. However, compass is vulnerable to environment magnetism disturbance. Existing compass calibration methods require complex steps and true heading reference which is often impossible to be obtained in outdoor AR applications, and is useful only when compass is in horizontal plane. An autocalibration method without the need of heading reference and redundant sensors is proposed in This work. First the compass error model based on physical principle is presented, then the algorithm to calculate the compensation coefficients with a set of sample measurements of the sensors in the compass is described. Because the influence of the environmental disturbance has been effectively compensated, the calibrated compass can provided accurate heading even when it is under large tilt attitude.
Xiaoming Hu 0001, Yue Liu 0005, Yongtian Wang, Yanling Hu, Dayuan Yan
ISMAR3
2000 Vision-Based Registration Using 3-D Fiducial for Augmented Reality
Yongtian Wang, Dayuan Yan
ICMI2