EDBT 2026 Demo / reviewers in the wild / expert
Xiaotao Shen
dblp:236/1596
· DBLP profile ↗
9ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-9608-9964ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Kun-peng enables scalable and accurate pan-domain metagenomic classificationabstractComprehensive pan-domain metagenomic classification is increasingly constrained by the memory and runtime costs of building and querying the rapidly expanding reference genome space. We introduce Kun-peng, a taxonomic classifier powered by an intelligent block-partitioned database structure and optimized search strategies, enabling ultra-scalable, memory-efficient pan-domain profiling. Using the Critical Assessment of Metagenome Interpretation II benchmark, Kun-peng substantially reduces the memory usage of database-building and querying by up to 24-fold, and accelerates sample classification by up to 4.73-fold compared with Kraken2. Kun-peng achieves competitive accuracy with fewer false positives than Kraken2, Centrifuger, and even KrakenUniq, while maintaining consistently high sensitivity across diverse datasets. In a real-world evaluation of 586 metagenomic samples spanning air, water, soil, and human-associated environments, we performed classification using a 4.3 TB pan-domain database comprising 204,477 genomes, which was built by Kun-peng with only 4.1 GB peak memory. Kun-peng processed each sample in 0.2-11.2 min with 4.0-35.4 GB peak memory, corresponding to a 54-473-fold reduction in memory usage relative to Kraken2. Compared with Sylph, Kun-peng achieved up to a 46-fold speedup while requiring 21-fold less memory. Kun-peng classified 69.8%-94.3% of reads, improving coverage by 20%-60% over the standard Kraken2 database with 62,026 genomes. This improvement reflects expanded reference coverage, although a small fraction of false positives is inherent to k-mer-based methods. Overall, Kun-peng effectively eliminates the long-standing memory bottleneck in pan-domain database building and classification, enabling rapid and scalable pan-domain taxonomic analysis of complex environmental, ecological, and exposomic sequencing datasets. Boliang Zhang, Xiaotao Shen |
Briefings Bioinform. | 6 |
| 2026 | MetaNet: a scalable and integrated tool for reproducible omics network analysisabstractMOTIVATION: Network analysis has become a central strategy for dissecting complex biological and environmental systems, particularly as modern omics technologies generate increasingly large and heterogeneous datasets. However, current tools often lack the scalability, flexibility, and native multi-omics support required for high-dimensional data analysis. We developed MetaNet, a high-performance R package that unifies network construction, visualization, and analysis across diverse omics layers. RESULTS: MetaNet enables fast and scalable correlation-based network construction for datasets with more than 10 000 features, providing over 40 layout algorithms, rich annotation utilities, and visualization options compatible with both static and interactive platforms. It further offers comprehensive topological and stability metrics for in-depth network characterization. Benchmarking shows that MetaNet delivers up to a 100-fold improvement in computation time and a 50-fold reduction in memory usage compared to existing R packages. We demonstrate its utility through two representative applications: (1) longitudinal microbial co-occurrence networks revealing airborne microbiome dynamics, and (2) an integrative exposome-transcriptome network of over 40 000 features, uncovering distinct regulatory impacts of biological and chemical exposures. By offering a robust, reproducible, and biologically informed framework, MetaNet advances multi-omics network analysis across biological, ecological, and environmental domains. AVAILABILITY: MetaNet package is freely available at https://github.com/Asa12138/MetaNet. Liuyiqi Jiang, Zinuo Huang, Xiaotao Shen |
Bioinform. | 8 |
| 2025 | Longitudinal urine metabolic profiling and gestational age prediction in human pregnancyabstractPregnancy is a vital period affecting both maternal and fetal health, with impacts on maternal metabolism, fetal growth, and long-term development. While the maternal metabolome undergoes significant changes during pregnancy, longitudinal shifts in maternal urine have been largely unexplored. In this study, we applied liquid chromatography-mass spectrometry-based untargeted metabolomics to analyze 346 maternal urine samples collected throughout pregnancy from 36 women with diverse backgrounds and clinical profiles. Key metabolite changes included glucocorticoids, lipids, and amino acid derivatives, indicating systematic pathway alterations. We also developed a machine learning model to accurately predict gestational age using urine metabolites, offering a non-invasive pregnancy dating method. Additionally, we demonstrated the ability of the urine metabolome to predict time-to-delivery, providing a complementary tool for prenatal care and delivery planning. This study highlights the clinical potential of urine untargeted metabolomics in obstetric care. Xiaotao Shen, Songjie Chen, Monika Avina, Hanyah Zackriah, Laura Jelliffe-Pawlowski, Larry Rand, Michael Snyder 0001 |
Briefings Bioinform. | 1 |
| 2024 | Generalized reporter score-based enrichment analysis for omics dataabstractEnrichment analysis contextualizes biological features in pathways to facilitate a systematic understanding of high-dimensional data and is widely used in biomedical research. The emerging reporter score-based analysis (RSA) method shows more promising sensitivity, as it relies on P-values instead of raw values of features. However, RSA cannot be directly applied to multi-group and longitudinal experimental designs and is often misused due to the lack of a proper tool. Here, we propose the Generalized Reporter Score-based Analysis (GRSA) method for multi-group and longitudinal omics data. A comparison with other popular enrichment analysis methods demonstrated that GRSA had increased sensitivity across multiple benchmark datasets. We applied GRSA to microbiome, transcriptome and metabolome data and discovered new biological insights in omics studies. Finally, we demonstrated the application of GRSA beyond functional enrichment using a taxonomy database. We implemented GRSA in an R package, ReporterScore, integrating with a powerful visualization module and updatable pathway databases, which is available on the Comprehensive R Archive Network (https://cran.r-project.org/web/packages/ReporterScore). We believe that the ReporterScore package will be a valuable asset for broad biomedical research fields. Shangjin Tan, Xiaotao Shen |
Briefings Bioinform. | 4 |
| 2022 | Deep learning-based pseudo-mass spectrometry imaging analysis for precision medicineabstractLiquid chromatography-mass spectrometry (LC-MS)-based untargeted metabolomics provides systematic profiling of metabolic. Yet, its applications in precision medicine (disease diagnosis) have been limited by several challenges, including metabolite identification, information loss and low reproducibility. Here, we present the deep-learning-based Pseudo-Mass Spectrometry Imaging (deepPseudoMSI) project (https://www.deeppseudomsi.org/), which converts LC-MS raw data to pseudo-MS images and then processes them by deep learning for precision medicine, such as disease diagnosis. Extensive tests based on real data demonstrated the superiority of deepPseudoMSI over traditional approaches and the capacity of our method to achieve an accurate individualized diagnosis. Our framework lays the foundation for future metabolic-based precision medicine. Xiaotao Shen, Wei Shao 0008, Chuchu Wang, Songjie Chen, Mirabela Rusu, Michael Snyder 0001 |
Briefings Bioinform. | 1 |
| 2022 | metID: an R package for automatable compound annotation for LC-MS-based dataabstractSUMMARY: Accurate and efficient compound annotation is a long-standing challenge for LC-MS-based data (e.g. untargeted metabolomics and exposomics). Substantial efforts have been devoted to overcoming this obstacle, whereas current tools are limited by the sources of spectral information used (in-house and public databases) and are not automated and streamlined. Therefore, we developed metID, an R package that combines information from all major databases for comprehensive and streamlined compound annotation. metID is a flexible, simple and powerful tool that can be installed on all platforms, allowing the compound annotation process to be fully automatic and reproducible. A detailed tutorial and a case study are provided in Supplementary Materials. AVAILABILITY AND IMPLEMENTATION: https://jaspershen.github.io/metID. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaotao Shen, Songjie Chen, Kévin Contrepois, Zheng-Jiang Zhu, Michael Snyder 0001 |
Bioinform. | 1 |
| 2022 | massDatabase: utilities for the operation of the public compound and pathway databaseabstractSUMMARY: One of the major challenges in liquid chromatography coupled to mass spectrometry data is converting many metabolic feature entries to biological function information, such as metabolite annotation and pathway enrichment, which are based on the compound and pathway databases. Multiple online databases have been developed. However, no tool has been developed for operating all these databases for biological analysis. Therefore, we developed massDatabase, an R package that operates the online public databases and combines with other tools for streamlined compound annotation and pathway enrichment. massDatabase is a flexible, simple and powerful tool that can be installed on all platforms, allowing the users to leverage all the online public databases for biological function mining. A detailed tutorial and a case study are provided in the Supplementary Material. AVAILABILITY AND IMPLEMENTATION: https://massdatabase.tidymass.org/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaotao Shen, Chuchu Wang, Michael Snyder 0001 |
Bioinform. | 1 |
| 2019 | MetFlow: an interactive and integrated workflow for metabolomics data cleaning and differential metabolite discoveryabstractSUMMARY: Mass spectrometry-based metabolomics aims to profile the metabolic changes in biological systems and identify differential metabolites related to physiological phenotypes and aberrant activities. However, many confounding factors during data acquisition complicate metabolomics data, which is characterized by high dimensionality, uncertain degrees of missing and zero values, nonlinearity, unwanted variations and non-normality. Therefore, prior to differential metabolite discovery analysis, various types of data cleaning such as batch alignment, missing value imputation, data normalization and scaling are essentially required for data post-processing. Here, we developed an interactive web server, namely, MetFlow, to provide an integrated and comprehensive workflow for metabolomics data cleaning and differential metabolite discovery. AVAILABILITY AND IMPLEMENTATION: The MetFlow is freely available on http://metflow.zhulab.cn/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaotao Shen, Zheng-Jiang Zhu |
Bioinform. | 1 |
| 2019 | LipidIMMS Analyzer: integrating multi-dimensional information to support lipid identification in ion mobility - mass spectrometry based lipidomicsabstractSUMMARY: Ion mobility-mass spectrometry (IM-MS) has showed great application potential for lipidomics. However, IM-MS based lipidomics is significantly restricted by the available software for lipid structural identification. Here, we developed a software tool, namely, LipidIMMS Analyzer, to support the accurate identification of lipids in IM-MS. For the first time, the software incorporates a large-scale database covering over 260 000 lipids and four-dimensional structural information for each lipid [i.e. m/z, retention time (RT), collision cross-section (CCS) and MS/MS spectra]. Therefore, multi-dimensional information can be readily integrated to support lipid identifications, and significantly improve the coverage and confidence of identification. Currently, the software supports different IM-MS instruments and data acquisition approaches. AVAILABILITY AND IMPLEMENTATION: The software is freely available at: http://imms.zhulab.cn/LipidIMMS/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaotao Shen, Jia Tu, Zheng-Jiang Zhu |
Bioinform. | 2 |