VLDB 2026 Research / reviewers in the wild / expert
Xiaoqing Yu
dblp:19/5905
· DBLP profile ↗
34ranked-venue papers
7as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An explainable eye-tracking-based framework for enhanced level-specific situational awareness recognition in air traffic control
Xing Yao, Chun-Hsien Chen, Bufan Liu, Guorui Ma, Xiaoqing Yu |
Adv. Eng. Informatics | 5 |
| 2026 | Graph structure learning with joint node and structural feature representation for node classification
Junsheng Wu, Weigang Li 0005, Xiaoqing Yu |
Neurocomputing | 4 |
| 2026 | Hybrid-Aligned Domain Adaptation for Driver Distraction RecognitionabstractDriver distraction recognition is a critical component of human-machine collaborative driving systems. Accurately identifying driver distraction behaviors is of great significance for improving road traffic safety. However, previous studies often focus on single experimental settings, neglecting the significant variations in data distribution caused by factors such as camera angles, lighting conditions, and experimental subjects across different environments. This leads to significant challenges in model generalization across domains and in diverse and uncertain real-world scenarios. To address these issues, this paper proposes a novel hybrid unsupervised domain adaptation framework. The proposed method achieves hybrid-aligned domain adaptation by minimizing subdomain-level feature distribution discrepancies between the source and target domains, while simultaneously reducing the divergence between the logits of classifiers that exhibit prediction discrepancies on the target domain. This enhances the accuracy of driver distraction behavior recognition in the target domain, improving the model’s generalization ability and robustness. Extensive experiments on four distracted driving datasets demonstrate that our proposed strategy outperforms previous methods, especially when dealing with datasets with imbalanced class distributions, which are more representative of real-world scenarios. Ximing Zhou, Xiaoqing Yu, Haohan Yang |
IEEE Internet Things J. | 3 |
| 2025 | ScPanKD: Distilling Pan-Cancer Knowledge for Enhanced T Cell Subtypes Annotation in Single-Cell Transcriptomics DataabstractSingle-cell RNA sequencing (scRNA-seq) enables high-resolution characterization of cellular heterogeneity, and annotating major cell types has become a standard practice in scRNA -seq analysis pipelines. However, accurately identifying fine-grained subtypes within major cell types remains challenging, particularly in heterogeneous tissues such as cancer samples. Here, we present ScPanKD, a computational frame-work for accurate and robust classification of fine-grained T cell subtypes in cancer samples. Unlike existing methods that suffer from cell type mismatches between reference and query datasets due to cancer heterogeneity, ScPanKD leverages knowledge distillation (KD) to accurately identify T cell subtypes even when the reference dataset contains more subtype diversity than the query. ScPanKD learns a cancer-invariant feature space and employs a two-step strategy, anchor cell selection followed by KD, to mitigate distribution shifts between reference and query datasets. Across extensive experiments using pan-cancer level CD4+ or CD8+ T cell atlases as references, we demonstrate that ScPanKD outperforms conventional annotation methods and single-cell foundation models, achieving more accurate and robust T cell subtype classification. ScPanKD and all reproducible scripts are available at https://github.com/marvinquiet/ScPanKD. Wenjing Ma, Xiaoqing Yu, Jiaying Lu 0001 |
BIBM | 2 |
| 2025 | Mamba in Mamba: Offline Reinforcement Learning via Sequence Modeling with Inner and Outer Selective State Spaces
Kaixin Jin, Xiaoqing Yu |
ICIC (20) | 6 |
| 2025 | PTsFusion: Multimodal Medical Image Fusion Based on Parallel Transformers with Global-Local Feature InteractionabstractMultimodal medical image fusion aim is to integrate complementary information from different modal images. In response to the existing multimodal medical image fusion methods, transformers struggles to extract local features, and loss of detailed information in the feature extraction process. Our proposed solution is parallel transformers with global-local feature interaction (PTsFusion). The PTsFusion mostly composed of three modules — Encoder, Fusion, and Decoder. Within the Encoder module, we use parallel transformers with global-local feature interaction, which consist of global transformer blocks, local transformer blocks, and global-local feature interaction blocks. The global transformer block focuses on extracting global deep features using the global attention mechanism. The local transformer block handles local deep features through the local attention mechanism. The global-local feature interaction block capture correlations between the global deep features and local deep features. In the Fusion module, the global deep features and local deep features from various modal images are fused by the -norm-based image-sequence matrix fusion rule. Ultimately, the Decoder module remodels the fused image by deconvolution. Experimental findings illustrate that the generated fused image exhibits clear textures edges and local information, surpassing other comparative methods in qualitative and quantitative analyses. Kaixin Jin, Xiaoqing Yu |
IJCNN | 6 |
| 2025 | Cognitive workload quantification for air traffic controllers: An ensemble semi-supervised learning approach
Xiaoqing Yu, Chun-Hsien Chen, Haohan Yang |
Adv. Eng. Informatics | 1 |
| 2025 | ROICellTrack: a deep learning framework for integrating cellular imaging modalities in subcellular spatial transcriptomic profiling of tumor tissuesabstractMOTIVATION: Spatial transcriptomic (ST) technologies, such as GeoMx Digital Spatial Profiler, are increasingly utilized to investigate the role of diverse tumor microenvironment components, particularly in relation to cancer progression, treatment response, and therapeutic resistance. However, in many ST studies, the spatial information obtained from immunofluorescence imaging is primarily used for identifying regions of interest (ROIs) rather than as an integral part of downstream transcriptomic data analysis and interpretation. RESULTS: We developed ROICellTrack, a deep learning-based framework that better integrates cellular imaging with spatial transcriptomic profiling. By analyzing 56 ROIs from urothelial carcinoma of the bladder and upper tract urothelial carcinoma, ROICellTrack identified distinct cancer-immune cell mixtures, characterized by specific transcriptomic and morphological signatures and receptor-ligand interactions linked to tumor content and immune infiltrations. Our findings demonstrate the value of integrating imaging with transcriptomics to analyze spatial omics data, improving our understanding of tumor heterogeneity and its relevance to personalized and targeted therapies. AVAILABILITY AND IMPLEMENTATION: ROICellTrack is publicly available at https://github.com/wanglab1/ROICellTrack. Xiaofei Song, Xiaoqing Yu, Carlos Moran Segura, Hongzhi Xu, Tingyi Li, Joshua T. Davis, Aram Vosoughi, G. Daniel Grass, Roger Li |
Bioinform. | 2 |
| 2025 | Human operators' cognitive workload recognition with a dual attention-enabled multimodal fusion framework
Xiaoqing Yu, Haohan Yang, Chun-Hsien Chen |
Expert Syst. Appl. | 1 |
| 2025 | Contrastive Graph Semantic Learning via prototype for recommendation
Mi Wen, Weiwei Li 0007, Zizhu Fan, Xiaoqing Yu |
Inf. Sci. | 5 |
| 2024 | A Novel Approach to Assessing Air Traffic Controllers' Situation Awareness with Deep Learning and Eye TrackingabstractAir traffic control (ATC) serves a critical role in the aviation industry and is responsible for the safety and efficiency of aircraft movement. High situation awareness (SA) involves continuously monitoring, understanding, and anticipating the state of the air traffic environment to ensure safe and efficient flight operations, which is critical for air traffic controllers (ATCOs) to prevent accidents. It is a challenge to accurately and promptly identify ATCOs’ situation awareness in a non-intrusive manner. This study aims to integrate learning-based methods with eye-tracking technology to assess the ATCOs’ amount of SA. Eye movement data from 26 participants were collected as they monitored aircraft on a simulated ATC radar screen and responded to freeze-probe queries targeting different levels of SA. Several conventional machine learning and deep learning models are trained using eye movement data and the models’ performance is extensively evaluated. Notably, a hybrid model of a convolutional neural network (CNN) and a long shortterm memory network (LSTM) has optimal performance. The CNN-LSTM model can learn useful features from the temporal sequential eye movement data and achieves an accuracy and F1 score of $\mathbf{9 2. 7 \%}$ and $\mathbf{9 0. 2 \%}$, respectively. The effectiveness and empirical implications of this learning-based method can contribute to future works of developing real-time SA monitoring systems to detect loss of SA in ATCOs, reducing human errors and improving aviation safety. Xiaoqing Yu, Xing Yao, Chun-Hsien Chen |
CW | 1 |
| 2024 | BatchFLEX: feature-level equalization of X-batchabstractMOTIVATION: Integrative analysis of heterogeneous expression data remains challenging due to variations in platform, RNA quality, sample processing, and other unknown technical effects. Selecting the approach for removing unwanted batch effects can be a time-consuming and tedious process, especially for more biologically focused investigators. RESULTS: Here, we present BatchFLEX, a Shiny app that can facilitate visualization and correction of batch effects using several established methods. BatchFLEX can visualize the variance contribution of a factor before and after correction. As an example, we have analyzed ImmGen microarray data and enhanced its expression signals that distinguishes each immune cell type. Moreover, our analysis revealed the impact of the batch correction in altering the gene expression rank and single-sample GSEA pathway scores in immune cell types, highlighting the importance of real-time assessment of the batch correction for optimal downstream analysis. AVAILABILITY AND IMPLEMENTATION: Our tool is available through Github https://github.com/shawlab-moffitt/BATCH-FLEX-ShinyApp with an online example on Shiny.io https://shawlab-moffitt.shinyapps.io/batch_flex/. Joshua T. Davis, Alyssa N. Obermayer, Alex C. Soupir, Rebecca S. Hesterberg, Thac Duong, Ching-Yao Yang, Ken Phong Dao, Brandon J. Manley, G. Daniel Grass, Dorina Avram, Paulo C. Rodriguez, Brooke L. Fridley, Xiaoqing Yu, Mingxiang Teng, Timothy I. Shaw |
Bioinform. | 13 |
| 2024 | A robust operators' cognitive workload recognition method based on denoising masked autoencoder
Xiaoqing Yu, Chun-Hsien Chen |
Knowl. Based Syst. | 1 |
| 2023 | Personalized federated adaptive regularization for heterogeneous medical image classification: COVID-19 CT resultsabstractDuring outbreaks of large infectious diseases like COVID-19, there is a strain on healthcare resources worldwide. To alleviate the burden on healthcare workers during the initial stages of the outbreak, there is an urgent need for the development of automated tools. Federated learning offers a privacypreserving solution to the challenge that limited annotated data within single healthcare facilities. However, existing federated aggregation strategies cannot adapt to the real-world medical image data problem of mixed heterogeneous. This paper introduces Personalized Federated Adaptive Regularization (pFedAR), an adaptive framework for federated learning that effectively utilizes multi-site COVID-19 CT datasets with mixed heterogeneous of distribution discrepancies. To solve the mixed distribution of label skew and feature shift, we propose a two-stage adaptive regularization. In the client training stage, we use the balance loss term to balance the COVID-19 clients with missing label. In the aggregation stage, we correct the client gradient conflict. Our method is developed and evaluated using eight real-world COVID-19 diagnosis datasets composed of CT images. Extensive experiments demonstrate the consistent improvement achieved by our method across the datasets. Yuan Liu 0038, Xiaoqing Yu, Jianxin Wang 0001 |
BIBM | 5 |
| 2023 | Air traffic controllers' mental fatigue recognition: A multi-sensor information fusion-based deep learning approach
Xiaoqing Yu, Chun-Hsien Chen, Haohan Yang |
Adv. Eng. Informatics | 1 |
| 2023 | Fast all versus all genotype comparison using DNA/RNA sequencing data: method and workflowabstractBACKGROUND: Massively parallel sequencing includes many liquid handling steps which introduce the possibility of sample swaps, mixing, and duplication. The unique profile of inherited variants in human genomes allows for comparison of sample identity using sequence data. A comparison of all samples vs. each other (all vs. all) provides both identification of mismatched samples and the possibility of resolving swapped samples. However, all vs. all comparison complexity grows as the square of the number of samples, so efficiency becomes essential. RESULTS: We have developed a tool for fast all vs. all genotype comparison using low level bitwise operations built into the Perl programming language. Importantly, we have also developed a complete workflow allowing users to start with either raw FASTQ sequence files, aligned BAM files, or genotype VCF files and automatically generate comparison metrics and summary plots. The tool is freely available at https://github.com/teerjk/TimeAttackGenComp/ . CONCLUSIONS: A fast and easy to use method for genotype comparison as described here is an important tool to ensure high quality and robust results in sequencing studies. Steven Eschrich, Xiaoqing Yu, Jamie K. Teer |
BMC Bioinform. | 2 |
| 2023 | Intra-tumor heterogeneity, turnover rate and karyotype space shape susceptibility to missegregation-induced extinctionabstractThe phenotypic efficacy of somatic copy number alterations (SCNAs) stems from their incidence per base pair of the genome, which is orders of magnitudes greater than that of point mutations. One mitotic event stands out in its potential to significantly change a cell's SCNA burden-a chromosome missegregation. A stochastic model of chromosome mis-segregations has been previously developed to describe the evolution of SCNAs of a single chromosome type. Building upon this work, we derive a general deterministic framework for modeling missegregations of multiple chromosome types. The framework offers flexibility to model intra-tumor heterogeneity in the SCNAs of all chromosomes, as well as in missegregation- and turnover rates. The model can be used to test how selection acts upon coexisting karyotypes over hundreds of generations. We use the model to calculate missegregation-induced population extinction (MIE) curves, that separate viable from non-viable populations as a function of their turnover- and missegregation rates. Turnover- and missegregation rates estimated from scRNA-seq data are then compared to theoretical predictions. We find convergence of theoretical and empirical results in both the location of MIE curves and the necessary conditions for MIE. When a dependency of missegregation rate on karyotype is introduced, karyotypes associated with low missegregation rates act as a stabilizing refuge, rendering MIE impossible unless turnover rates are exceedingly high. Intra-tumor heterogeneity, including heterogeneity in missegregation rates, increases as tumors progress, rendering MIE unlikely. Gregory Kimmel, Richard J. Beck, Xiaoqing Yu, Thomas Veith, Samuel Bakhoum, Philipp M. Altrock, Noemi Andor |
PLoS Comput. Biol. | 3 |
| 2022 | Deep MRI glioma segmentation via multiple guidances and hybrid enhanced-gradient cross-entropy loss
Jinjing Zhang, Lijun Zhao 0002, Jianchao Zeng 0001, Pinle Qin, Xiaoqing Yu |
Expert Syst. Appl. | 6 |
| 2020 | A novel augmented reality framework based on monocular semi-dense simultaneous localization and mappingabstractSummary Markerless tracking has been a trend in augmented reality (AR) applications nowadays, but it no longer satisfies users who want virtual characters to interact with the real world such as collision. Some sparse or dense simultaneous localization and mapping (SLAM) methods are proposed aiming to solve this problem. However, sparse methods only extract a plane from the sparse map, which cannot allow virtual characters to move realistically. Meanwhile, dense methods usually require powerful graphics processing unit (GPU) for dense mapping. In this paper, we present a real‐time AR framework based on a semi‐dense method with central processing unit (CPU). Specifically, the semi‐dense method searches pixels with high gradients in each keyframe and estimates accurate depths by fusing matching pixels in other keyframes. We propose an outlier removal method that excludes three‐dimensional points outside the camera trajectory. By integrating this method, our framework preserves clean edges of the real environment. The experimental results on the dataset show that our proposed framework has better surface reconstruction accuracy than other methods and our tracking thread runs in an acceptable speed when the semi‐dense mapping thread runs backend. With the benefit of the robust camera tracking and the aligned surface, virtual characters of our AR application enable realistic movement and collision. Lianyao Wu, Xiaoqing Yu, Chunkai Ye, A. A. M. Muzahid |
Comput. Animat. Virtual Worlds | 3 |
| 2020 | Efficient Mining Multi-Mers in a Variety of Biological SequencesabstractCounting the occurrence frequency of each $k$k-mer in a biological sequence is a preliminary yet important step in many bioinformatics applications. However, most $k$k-mer counting algorithms rely on a given $k$k to produce single-length $k$k-mers, which is inefficient for sequence analysis for different $k$k. Moreover, existing $k$k-mer counters focus more on DNA and RNA sequences and less on protein ones. In practice, the analysis of $k$k-mers in protein sequences can provide substantial biological insights in structure, function, and evolution. To this end, an efficient algorithm, called MulMer (Multiple-Mer mining), is proposed to mine $k$k-mers of various lengths termed multi-mers via inverted-index technique, which is orders of magnitude faster than the conventional forward-index methods. Moreover, to the best of our knowledge, MulMer is the first able to mine multi-mers in a variety of sequences, including DNA, RNA, and protein sequences. Jingsong Zhang, Jianmei Guo, Xiangtian Yu, Xiaoqing Yu, Weifeng Guo, Tao Zeng 0003, Luonan Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2019 | Multiple-kernel learning for genomic data mining and predictionabstractBACKGROUND: Advances in medical technology have allowed for customized prognosis, diagnosis, and treatment regimens that utilize multiple heterogeneous data sources. Multiple kernel learning (MKL) is well suited for the integration of multiple high throughput data sources. MKL remains to be under-utilized by genomic researchers partly due to the lack of unified guidelines for its use, and benchmark genomic datasets. RESULTS: We provide three implementations of MKL in R. These methods are applied to simulated data to illustrate that MKL can select appropriate models. We also apply MKL to combine clinical information with miRNA gene expression data of ovarian cancer study into a single analysis. Lastly, we show that MKL can identify gene sets that are known to play a role in the prognostic prediction of 15 cancer types using gene expression data from The Cancer Genome Atlas, as well as, identify new gene sets for the future research. CONCLUSION: Multiple kernel learning coupled with modern optimization techniques provides a promising learning tool for building predictive models based on multi-source genomic data. MKL also provides an automated scheme for kernel prioritization and parameter tuning. The methods used in the paper are implemented as an R package called RMKL package, which is freely available for download through CRAN at https://CRAN.R-project.org/package=RMKL . Kaiqiao Li, Xiaoqing Yu, Pei Fen Kuan |
BMC Bioinform. | 3 |
| 2018 | Eye landmarks detection via two-level cascaded CNNs with multi-task learning
Bin Huang 0024, Renwen Chen, Qinbang Zhou, Xiaoqing Yu |
Signal Process. Image Commun. | 4 |
| 2017 | Mining K-mers of Various Lengths in Biological Sequences
Jingsong Zhang, Jianmei Guo, Xiaoqing Yu, Xiangtian Yu, Weifeng Guo, Tao Zeng 0003, Luonan Chen |
ISBRA | 3 |
| 2016 | Global copy number profiling of cancer genomesabstractUNLABELLED: In this article, we introduce a robust and efficient strategy for deriving global and allele-specific copy number alternations (CNA) from cancer whole exome sequencing data based on Log R ratios and B-allele frequencies. Applying the approach to the analysis of over 200 skin cancer samples, we demonstrate its utility for discovering distinct CNA events and for deriving ancillary information such as tumor purity. AVAILABILITY AND IMPLEMENTATION: https://github.com/xfwang/CLOSE CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mengjie Chen, Xiaoqing Yu, Natapol Pornputtapong, Hao Chen 0064, Nancy Ruonan Zhang, R. Scott Powers, Michael Krauthammer |
Bioinform. | 3 |
| 2015 | A trimming-and-retrieving alignment scheme for reduced representation bisulfite sequencingabstractAbstract Summary: Currently available bisulfite sequencing tools frequently suffer from low mapping rates and low methylation calls, especially for data generated from the Illumina sequencer, NextSeq. Here, we introduce a sequential trimming-and-retrieving alignment approach for investigating DNA methylation patterns, which significantly improves the number of mapped reads and covered CpG sites. The method is implemented in an automated analysis toolkit for processing bisulfite sequencing reads. Availability and implementation: http://mysbfiles.stonybrook.edu/~xuefenwang/software.html and https://github.com/xfwang/BStools. Contact: [email protected] Supplementary information: Supplementary materials are available at Bioinformatics online. Xiaoqing Yu, Wei Zhu 0008, W. Richard McCombie, Eric Antoniou, R. Scott Powers, Nicholas O. Davidson, Ellen Li, Jennie Williams |
Bioinform. | 2 |
| 2014 | A Robust and Fast Reconstruction Framework for Noisy and Large Point Cloud DataabstractIn this paper we present a robust reconstruction framework on noisy and large point cloud data. Though Poisson reconstruction performs well in recovering the surface from noisy point cloud data, it's problematic to reconstruct underlying surface from large cloud data, especially on a general processor. An inaccurate estimation of point normal for noisy and large dataset would result in local distortion on the reconstructed mesh. We adopt a systematical combination of Poisson-disk sampling, normal estimation and Poisson reconstruction to avoid the inaccuracy of normal calculated from k-nearest neighbors. With the fewer dataset obtained by sampling on original points, the normal estimated is more reliable for subsequent Poisson reconstruction and the time spent in normal estimation and reconstruction is much less. We demonstrate the effectiveness of the framework in recovering topology and geometry information when dealing with point cloud data from real world. The experiment results indicate that the framework is superior to Poisson reconstruction directly on raw point dataset in the aspects of time consumption and visual fidelity. Xiang Feng 0007, Xiaoqing Yu, Fabien Pfaender, J. Alfredo Sánchez 0001 |
CCGRID | 2 |
| 2013 | Parallel Simulation of Large-Scale Universal Particle Systems Using CUDAabstractParticle systems' greatest advantage is well suited for modeling complex fuzzy phenomena, such as explosions, fountain, tornado and fireworks, etc. in 3D graphics. With the increasing requirements on the number of particles and particle-particle interactions, the computational complexity of simulation in particle systems has increased rapidly. Particle systems are traditionally implemented on a general-purpose CPU, and the computational complexity of particle systems limits the number of particles that can be computed at interactive rates. This paper focuses on real-time simulation of large-scale particle systems. We discuss optional integration algorithms based on CUDA (Compute Unified Device Architecture) for both graphic and scientific simulation. The speed of particle systems has been greatly improved, with parallel-core GPUs working in tandem with multi-core CPUs. In order to provide a scalable and portable API library, the object-oriented programming method is adopted to encapsulate the functions of parallel particle system. Results show that our proposed APIs are user-friendly and the parallel implementations are significantly efficient. Xiangfei Li, Xuzhi Wang, Xiaoqiang Zhu, Xiaoqing Yu |
DASC | 5 |
| 2013 | MethyQA: a pipeline for bisulfite-treated methylation sequencing quality assessmentabstractBACKGROUND: DNA methylation is an epigenetic event that adds a methyl-group to the 5' cytosine. This epigenetic modification can significantly affect gene expression in both normal and diseased cells. Hence, it is important to study methylation signals at the single cytosine site level, which is now possible utilizing bisulfite conversion technique (i.e., converting unmethylated Cs to Us and then to Ts after PCR amplification) and next generation sequencing (NGS) technologies. Despite the advances of NGS technologies, certain quality issues remain. Some of the more prevalent quality issues involve low per-base sequencing quality at the 3' end, PCR amplification bias, and bisulfite conversion rates. Therefore, it is important to conduct quality assessment before downstream analysis. To the best of our knowledge, no existing software packages can generally assess the quality of methylation sequencing data generated based on different bisulfite-treated protocols. RESULTS: To conduct the quality assessment of bisulfite methylation sequencing data, we have developed a pipeline named MethyQA. MethyQA combines currently available open-source software packages with our own custom programs written in Perl and R. The pipeline can provide quality assessment results for tens of millions of reads in under an hour. The novelty of our pipeline lies in its examination of bisulfite conversion rates and of the DNA sequence structure of regions that have different conversion rates or coverage. CONCLUSIONS: MethyQA is a new software package that provides users with a unique insight into the methylation sequencing data they are researching. It allows the users to determine the quality of their data and better prepares them to address the research questions that lie ahead. Due to the speed and efficiency at which MethyQA operates, it will become an important tool for studies dealing with bisulfite methylation sequencing data. Shuying Sun, Aaron Noviski, Xiaoqing Yu |
BMC Bioinform. | 3 |
| 2013 | Comparing a few SNP calling algorithms using low-coverage sequencing dataabstractBACKGROUND: Many Single Nucleotide Polymorphism (SNP) calling programs have been developed to identify Single Nucleotide Variations (SNVs) in next-generation sequencing (NGS) data. However, low sequencing coverage presents challenges to accurate SNV identification, especially in single-sample data. Moreover, commonly used SNP calling programs usually include several metrics in their output files for each potential SNP. These metrics are highly correlated in complex patterns, making it extremely difficult to select SNPs for further experimental validations. RESULTS: To explore solutions to the above challenges, we compare the performance of four SNP calling algorithm, SOAPsnp, Atlas-SNP2, SAMtools, and GATK, in a low-coverage single-sample sequencing dataset. Without any post-output filtering, SOAPsnp calls more SNVs than the other programs since it has fewer internal filtering criteria. Atlas-SNP2 has stringent internal filtering criteria; thus it reports the least number of SNVs. The numbers of SNVs called by GATK and SAMtools fall between SOAPsnp and Atlas-SNP2. Moreover, we explore the values of key metrics related to SNVs' quality in each algorithm and use them as post-output filtering criteria to filter out low quality SNVs. Under different coverage cutoff values, we compare four algorithms and calculate the empirical positive calling rate and sensitivity. Our results show that: 1) the overall agreement of the four calling algorithms is low, especially in non-dbSNPs; 2) the agreement of the four algorithms is similar when using different coverage cutoffs, except that the non-dbSNPs agreement level tends to increase slightly with increasing coverage; 3) SOAPsnp, SAMtools, and GATK have a higher empirical calling rate for dbSNPs compared to non-dbSNPs; and 4) overall, GATK and Atlas-SNP2 have a relatively higher positive calling rate and sensitivity, but GATK calls more SNVs. CONCLUSIONS: Our results show that the agreement between different calling algorithms is relatively low. Thus, more caution should be used in choosing algorithms, setting filtering parameters, and designing validation studies. For reliable SNV calling results, we recommend that users employ more than one algorithm and use metrics related to calling quality and coverage as filtering criteria. Xiaoqing Yu, Shuying Sun |
BMC Bioinform. | 1 |
| 2012 | PEAQ Compatible Audio Quality Estimation Using Computational Auditory Model
Xiaoqing Yu |
ICONIP (4) | 4 |
| 2012 | Intrinsic Bayesian model for high-dimensional unsupervised reduction
Longcun Jin, Yongliang Wu, Bin Cui 0006, Xiaoqing Yu |
Neurocomputing | 5 |
| 2011 | Audio Quality Assessment Improvement via Circular and Flexible OverlapabstractThis paper proposed an improved audio quality metric via circular and flexible overlap. Based on the Power Spectrum Estimation via circular overlap, we use a novel circular overlap sub-frame to assess highly impaired audio. The relationship between the fraction of overlap and the metric of audio quality is also examined, and it has been proven that the accuracy of audio quality assessment increased with the fraction of overlap. By integrating circular and flexible overlap into ITU-R BS.1387, which is also known as Perceptual Evaluation of Audio Quality, our method can be applied to the quality assessment of highly impaired audio. Xiaoqing Yu |
ISM | 3 |
| 2010 | Packet Loss Concealment for compressed audio stream using sinusoidal frequency estimationabstractIn this paper we propose a Packet Loss Concealment (PLC) method for audio compressed in the Modified Discrete Cosine Transform (MDCT) domain. Based on sinusoidal and noise model in MDCT domain, a novel sinusoidal frequency estimation method is adopted to reconstruct the sign of MDCT coefficients in a lost frame. Further optimization is achieved by investigation of the cosine property. Both subjective and objective assessments are examined, and the results show that our proposed algorithm achieves almost first-class quality, while a small amount of calculation and memory are consumed. Xiaoqing Yu |
ICME | 3 |
| 2001 | Auditory model based speech recognition in noisy environmentabstractThe main purpose of this paper is to present how to raise the speech recognition performance in noisy environment. So far the most popularly used speech feature in speech recognition is probably the so-called MFCC. The recognition rate of speech recognition algorithm using MFCC and CDHMM is known to be very high in clean speech environment, but it deteriorates greatly in noisy environment, especially in the white noisy environment. In this paper, we propose a new speech feature, the ASBF speech feature based on the mathematical model of inner ear of human auditory system. This new speech feature is extracted using both mathematical model of inner ear and primary auditory nerve processing model of human auditory system, and it can track the speech formants effectively. In the experiment, the performance of MFCC and the ASBF are compared in both clean and noisy environments when using left-to-right CDHMM with 6 states and 5 Gaussian mixtures. The experimental result shows that the ASBF is much more robust to noise than MFCC. When only 5 dimension is used in ASBF vector, the recognition rate is approximately 38.6% higher than the traditional MFCC with 39 dimension in the condition of S/N=10dB with white noise. Xiaoqing Yu, Daniel Pak-Kong Lun |
INTERSPEECH | 1 |