Xiaoquan Su

dblp:68/10762 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-2144-1991ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 MEFE: microbiome signature identification based on elastic feature extraction
Guosen Hou, Yuxiao Fei, Xiaoquan Su
Frontiers Comput. Sci.4
2025 Concisemizer: An Asymmetric Redundancy-Removal Algorithm for Efficient Seed Sampling in Large-Scale Sequence Alignment
abstract
Modern sequence alignment tools rely on seed-based anchoring to efficiently align large-scale sequencing data. Minimizers are widely used for seed extraction but often generate excessive and redundant seeds that each base is typically sampled twice on average, thereby limiting alignment throughput. To address seed redundancy, we propose Concisemizer, an asymmetric, simultaneous de-redundancy algorithm. Concisemizer introduces positional filtering steps applied to sliding seed triples that simultaneously extract and filter seeds from the reference sequence, resulting in a more efficient seed set. By employing an asymmetric strategy that removes redundancy only from the reference sequences while retaining minimizers in the query sequences, Concisemizer reduces seed count while preserving alignment sensitivity. Comparative evaluations on six model species demonstrated Concisemizer's superior performance in seed density and repetition. Replacing the minimizer module in minimap2 with Concisemizer led to a$9.6 \times$speedup in human genome alignment, with minimal impact on accuracy. Concisemizer substantially reduces redundancy in seed extraction while maintaining accuracy, thereby improving the efficiency and scalability of large-scale sequence alignment. It holds strong potential for widespread adoption in genomic analysis pipelines.
Xiaoquan Su
BIBM3
2025 TheraPepNet: Enhancing Therapeutic Peptide Recognition via Temporal Sequence Prediction Networks and Preprocessed Features
abstract
Therapeutic peptides are gaining attention in drug development due to their specificity and biocompatibility. Unlike small-molecule drugs, peptide therapeutics offer better target specificity, efficacy, and fewer side effects. However, current ther-apeutic peptide predictors focus mainly on model architecture, neglecting the integration of both sequential and physicochemical properties of peptides, making accurate and interpretable iden-tification a challenge. To address this, we propose TheraPepN et, a prediction model that combines multiple embedding strategies with the temporal learning network to identify therapeutic pep-tides. TheraPepNet integrates three types of embed dings-ami no acid spectral features from Complex Fast Fourier Transform, pretrained feature vectors from ProtBERT, and improved biolog-ical encodings from the AAindex database-effectively combining amino acid properties with sequential information to enhance peptide prediction. Cross-validation and independent testing on benchmark datasets demonstrate TheraPepNet's strong per-formance and generalization ability, showing its potential in bioengineering and clinical applications. Compared to existing methods, TheraPepN et performs exceptionally well in terms of accuracy, F1-score, and other metrics. TheraPepNet is presented through a user-friendly web server at http://www.qdubiolab.top.
Zhenming Wu, Xiaoquan Su
BIBM3
2025 CAM-net: a context-aware network for identifying reliable microbial relations via propagation cliques
abstract
Abstract Motivation The human gut microbiome is a complex ecosystem essential to host health, yet its interspecies interactions remain incompletely understood. Conventional correlation analyses typically rely on isolated pairwise measures (e.g., Spearman correlation or SparCC), which overlook the broader ecological context. This narrow perspective often produces spurious associations in large-scale datasets, obscuring genuine microbial interactions. Methods Here we present CAM-Net, a context-aware microbial correlation framework that fundamentally departs from pairwise approaches. Instead of treating associations in isolation, CAM-Net dynamically constructs a ‘local environment’ around a target microbe and infers relational propagation cliques that reflect ecologically meaningful interactions. By explicitly controlling for community background, CAM-Net reduces false positives and captures context-dependent ecological dynamics such as facilitation and competition. Results We applied CAM-Net to whole-genome sequencing (WGS) and 16S rRNA gene sequencing datasets encompassing over 30,000 human gut microbiomes, focusing on two microbes: Akkermansia muciniphila and Lactobacillus acidophilus. For A. muciniphila, an autochthonous bacterium, CAM-Net identified a relational propagation clique that was reproducible across independent datasets, underscoring its symbiotic ecological role. In contrast, for L. acidophilus, an exogenous species of the human gut, CAM-Net recovered only a weakly correlated clique, consistent with prior knowledge of its transient colonization. Notably, neither species produced stable associations in 16S datasets, reflecting limitations of amplicon sequencing—including data heterogeneity, batch effects, and restricted taxonomic resolution. Conclusion These findings show CAM-Net as a robust, context-aware alternative to conventional correlations, moving beyond sparse pairwise comparisons to infer microbial interactions and decode microbial ecology.
Wenfei Xu, Jieqi Xing, Shi Huang, Xiaoquan Su
Briefings Bioinform.6
2025 TaxaCal: enhancing species-level profiling accuracy of 16S amplicon data
abstract
BACKGROUND: 16S rRNA amplicon sequencing is a widely used method for microbiome composition analysis due to its cost-effectiveness and lower data requirements compared to metagenomic whole-genome sequencing (WGS). However, inherent limitations in 16S-based approach often lead to profiling discrepancies, particularly at the species level, compromising the accuracy and reliability of findings. RESULTS: To address this issue, we present TaxaCal (Taxonomic Calibrator), a machine learning algorithm designed to calibrate species-level taxonomy profiles in 16S amplicon data using a two-tier correction strategy. Validation on in-house produced and public datasets shows that TaxaCal effectively reduces biases in amplicon sequencing, mitigating discrepancies between microbial profiles derived from 16S and WGS. Moreover, TaxaCal enables seamless cross-platform comparisons between these two sequencing approaches, significantly improving disease detection in 16S-based microbiome data. CONCLUSIONS: Therefore, TaxaCal offers a cost-effective solution for generating high-resolution microbiome species profiles that closely align with WGS results, enhancing the utility of 16S-based profiling in microbiome research. As microbiome-based diagnostics continue to evolve, TaxaCal has the potential to be a crucial tool in advancing the utility of 16S sequencing in clinical and research settings.
Qingrong Shen, Xiaoqian Fan, Xiaoquan Su
BMC Bioinform.5
2024 A deep learning framework for predicting disease-gene associations with functional modules and graph augmentation
abstract
BACKGROUND: The exploration of gene-disease associations is crucial for understanding the mechanisms underlying disease onset and progression, with significant implications for prevention and treatment strategies. Advances in high-throughput biotechnology have generated a wealth of data linking diseases to specific genes. While graph representation learning has recently introduced groundbreaking approaches for predicting novel associations, existing studies always overlooked the cumulative impact of functional modules such as protein complexes and the incompletion of some important data such as protein interactions, which limits the detection performance. RESULTS: Addressing these limitations, here we introduce a deep learning framework called ModulePred for predicting disease-gene associations. ModulePred performs graph augmentation on the protein interaction network using L3 link prediction algorithms. It builds a heterogeneous module network by integrating disease-gene associations, protein complexes and augmented protein interactions, and develops a novel graph embedding for the heterogeneous module network. Subsequently, a graph neural network is constructed to learn node representations by collectively aggregating information from topological structure, and gene prioritization is carried out by the disease and gene embeddings obtained from the graph neural network. Experimental results underscore the superiority of ModulePred, showcasing the effectiveness of incorporating functional modules and graph augmentation in predicting disease-gene associations. This research introduces innovative ideas and directions, enhancing the understanding and prediction of gene-disease relationships.
Xianghu Jia, Weiwen Luo, Jieqi Xing, Hongjie Sun, Shunyao Wu, Xiaoquan Su
BMC Bioinform.7
2023 Flex Meta-Storms elucidates the microbiome local beta-diversity under specific phenotypes
abstract
MOTIVATION: Beta-diversity quantitatively measures the difference among microbial communities thus enlightening the association between microbiome composition and environment properties or host phenotypes. The beta-diversity analysis mainly relies on distances among microbiomes that are calculated by all microbial features. However, in some cases, only a small fraction of members in a community plays crucial roles. Such a tiny proportion is insufficient to alter the overall distance, which is always missed by end-to-end comparison. On the other hand, beta-diversity pattern can also be interfered due to the data sparsity when only focusing on nonabundant microbes. RESULTS: Here, we develop Flex Meta-Storms (FMS) distance algorithm that implements the "local alignment" of microbiomes for the first time. Using a flexible extraction that considers the weighted phylogenetic and functional relations of microbes, FMS produces a normalized phylogenetic distance among members of interest for microbiome pairs. We demonstrated the advantage of FMS in detecting the subtle variations of microbiomes among different states using artificial and real datasets, which were neglected by regular distance metrics. Therefore, FMS effectively discriminates microbiomes with higher sensitivity and flexibility, thus contributing to in-depth comprehension of microbe-host interactions, as well as promoting the utilization of microbiome data such as disease screening and prediction. AVAILABILITY AND IMPLEMENTATION: FMS is implemented in C++, and the source code is released at https://github.com/qdu-bioinfo/flex-meta-storms.
Mingqian Zhang, Wenke Zhang, Shunyao Wu, Xiaoquan Su
Bioinform.6
2020 Dynamic Meta-Storms enables comprehensive taxonomic and phylogenetic comparison of shotgun metagenomes at the species level
abstract
MOTIVATION: An accurate and reliable distance (or dissimilarity) among shotgun metagenomes is fundamental to deducing the beta-diversity of microbiomes. To compute the distance at the species level, current methods either ignore the evolutionary relationship among species or fail to account for unclassified organisms that cannot be mapped to definite tip nodes in the phylogenic tree, thus can produce erroneous beta-diversity pattern. RESULTS: To solve these problems, we propose the Dynamic Meta-Storms (DMS) algorithm to enable the comprehensive comparison of metagenomes on the species level with both taxonomy and phylogeny profiles. It compares the identified species of metagenomes with phylogeny, and then dynamically places the unclassified species to the virtual nodes of the phylogeny tree via their higher-level taxonomy information. Its high speed and low memory consumption enable pairwise comparison of 100 000 metagenomes (synthesized from 3688 bacteria) within 6.4 h on a single computing node. AVAILABILITY AND IMPLEMENTATION: An optimized implementation of DMS is available on GitHub (https://github.com/qibebt-bioinfo/dynamic-meta-storms) under a GNU GPL license. It takes the species-level profiles of metagenomes as input, and generates their pairwise distance matrix. The bacterial species-level phylogeny tree and taxonomy information of MetaPhlAn2 have been integrated into this implementation, while customized tree and taxonomy are also supported. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gongchao Jing, Jian Xu 0020, Xiaoquan Su
Bioinform.6
2016 MultiGeMS: detection of SNVs from multiple samples using model selection on high-throughput sequencing data
abstract
MOTIVATION: Single nucleotide variant (SNV) detection procedures are being utilized as never before to analyze the recent abundance of high-throughput DNA sequencing data, both on single and multiple sample datasets. Building on previously published work with the single sample SNV caller genotype model selection (GeMS), a multiple sample version of GeMS (MultiGeMS) is introduced. Unlike other popular multiple sample SNV callers, the MultiGeMS statistical model accounts for enzymatic substitution sequencing errors. It also addresses the multiple testing problem endemic to multiple sample SNV calling and utilizes high performance computing (HPC) techniques. RESULTS: A simulation study demonstrates that MultiGeMS ranks highest in precision among a selection of popular multiple sample SNV callers, while showing exceptional recall in calling common SNVs. Further, both simulation studies and real data analyses indicate that MultiGeMS is robust to low-quality data. We also demonstrate that accounting for enzymatic substitution sequencing errors not only improves SNV call precision at low mapping quality regions, but also improves recall at reference allele-dominated sites with high mapping quality. AVAILABILITY AND IMPLEMENTATION: The MultiGeMS package can be downloaded from https://github.com/cui-lab/multigems CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gabriel H. Murillo, Na You, Xiaoquan Su, Muredach P. Reilly, Mingyao Li, Kang Ning 0001, Xinping Cui
Bioinform.3
2015 Condensing Raman spectrum for single-cell phenotype analysis
abstract
BACKGROUND: In recent years, high throughput and non-invasive Raman spectrometry technique has matured as an effective approach to identification of individual cells by species, even in complex, mixed populations. Raman profiling is an appealing optical microscopic method to achieve this. To fully utilize Raman proling for single-cell analysis, an extensive understanding of Raman spectra is necessary to answer questions such as which filtering methodologies are effective for pre-processing of Raman spectra, what strains can be distinguished by Raman spectra, and what features serve best as Raman-based biomarkers for single-cells, etc. RESULTS: In this work, we have proposed an approach called rDisc to discretize the original Raman spectrum into only a few (usually less than 20) representative peaks (Raman shifts). The approach has advantages in removing noises, and condensing the original spectrum. In particular, effective signal processing procedures were designed to eliminate noise, utilising wavelet transform denoising, baseline correction, and signal normalization. In the discretizing process, representative peaks were selected to signicantly decrease the Raman data size. More importantly, the selected peaks are chosen as suitable to serve as key biological markers to differentiate species and other cellular features. Additionally, the classication performance of discretized spectra was found to be comparable to full spectrum having more than 1000 Raman shifts. Overall, the discretized spectrum needs about 5storage space of a full spectrum and the processing speed is considerably faster. This makes rDisc clearly superior to other methods for single-cell classication.
Shiwei Sun, Xuetao Wang, Xin Gao 0001, Lihui Ren, Xiaoquan Su, Dongbo Bu, Kang Ning 0001
BMC Bioinform.5
2014 GPU-Meta-Storms: computing the structure similarities among massive amount of microbial community samples using GPU
abstract
MOTIVATION: The number of microbial community samples is increasing with exponential speed. Data-mining among microbial community samples could facilitate the discovery of valuable biological information that is still hidden in the massive data. However, current methods for the comparison among microbial communities are limited by their ability to process large amount of samples each with complex community structure. SUMMARY: We have developed an optimized GPU-based software, GPU-Meta-Storms, to efficiently measure the quantitative phylogenetic similarity among massive amount of microbial community samples. Our results have shown that GPU-Meta-Storms would be able to compute the pair-wise similarity scores for 10 240 samples within 20 min, which gained a speed-up of >17 000 times compared with single-core CPU, and >2600 times compared with 16-core CPU. Therefore, the high-performance of GPU-Meta-Storms could facilitate in-depth data mining among massive microbial community samples, and make the real-time analysis and monitoring of temporal or conditional changes for microbial communities possible. AVAILABILITY AND IMPLEMENTATION: GPU-Meta-Storms is implemented by CUDA (Compute Unified Device Architecture) and C++. Source code is available at http://www.computationalbioenergy.org/meta-storms.html.
Xiaoquan Su, Xuetao Wang, Gongchao Jing, Kang Ning 0001
Bioinform.1
2012 Meta-Storms: efficient search for similar microbial communities based on a novel indexing scheme and similarity score for metagenomic data
abstract
BACKGROUND: It has long been intriguing scientists to effectively compare different microbial communities (also referred as 'metagenomic samples' here) in a large scale: given a set of unknown samples, find similar metagenomic samples from a large repository and examine how similar these samples are. With the current metagenomic samples accumulated, it is possible to build a database of metagenomic samples of interests. Any metagenomic samples could then be searched against this database to find the most similar metagenomic sample(s). However, on one hand, current databases with a large number of metagenomic samples mostly serve as data repositories that offer few functionalities for analysis; and on the other hand, methods to measure the similarity of metagenomic data work well only for small set of samples by pairwise comparison. It is not yet clear, how to efficiently search for metagenomic samples against a large metagenomic database. RESULTS: In this study, we have proposed a novel method, Meta-Storms, that could systematically and efficiently organize and search metagenomic data. It includes the following components: (i) creating a database of metagenomic samples based on their taxonomical annotations, (ii) efficient indexing of samples in the database based on a hierarchical taxonomy indexing strategy, (iii) searching for a metagenomic sample against the database by a fast scoring function based on quantitative phylogeny and (iv) managing database by index export, index import, data insertion, data deletion and database merging. We have collected more than 1300 metagenomic data from the public domain and in-house facilities, and tested the Meta-Storms method on these datasets. Our experimental results show that Meta-Storms is capable of database creation and effective searching for a large number of metagenomic samples, and it could achieve similar accuracies compared with the current popular significance testing-based methods. CONCLUSION: Meta-Storms method would serve as a suitable database management and search system to quickly identify similar metagenomic samples from a large pool of samples. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaoquan Su, Jian Xu 0020, Kang Ning 0001
Bioinform.1
2012 SNP calling using genotype model selection on high-throughput sequencing data
abstract
MOTIVATION: A review of the available single nucleotide polymorphism (SNP) calling procedures for Illumina high-throughput sequencing (HTS) platform data reveals that most rely mainly on base-calling and mapping qualities as sources of error when calling SNPs. Thus, errors not involved in base-calling or alignment, such as those in genomic sample preparation, are not accounted for. RESULTS: A novel method of consensus and SNP calling, Genotype Model Selection (GeMS), is given which accounts for the errors that occur during the preparation of the genomic sample. Simulations and real data analyses indicate that GeMS has the best performance balance of sensitivity and positive predictive value among the tested SNP callers. AVAILABILITY: The GeMS package can be downloaded from https://sites.google.com/a/bioinformatics.ucr.edu/xinping-cui/home/software or http://computationalbioenergy.org/software.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Na You, Gabriel H. Murillo, Xiaoquan Su, Xiaowei Zeng, Jian Xu 0020, Kang Ning 0001, Shoudong Zhang, Jiankang Zhu, Xinping Cui
Bioinform.3
2011 An Open-source Collaboration Environment for Metagenomics Research
abstract
By analyzing metagenomic data from microbial communities, the taxonomical and functional component of hundreds of previously unknown microbial communities have been elucidated in the past few years. However, metagenomic data analyses are both data- and computation-intensive, which require extensive computational power. Most of the current metagenomic data analysis software were designed to be used on a single PC (Personal Computer), which could not match with the fast increasing number of large metagenomic projects' computational requirements. Therefore, advanced computational environment has to be developed to cope with such needs. In this paper, we proposed an open-source collaboration environment for metagenomic data analysis, which enabled the parallel analysis of multiple metagenomic datasets at the same time. By using this collaboration environment, researchers from different locations could submit their data, collaboratively configure the analysis pipeline, and perform data analysis efficiently. As of now, more than 30 metagenomic data analysis projects have already been conducted based on this environment.
Xiaoquan Su, Yongzheng Ma, Xingzhi Chang, Kai Nan, Jian Xu 0020, Kang Ning 0001
eScience1