EDBT 2026 Demo / reviewers in the wild / expert
Mingxuan Gao
dblp:291/1191
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2022
0000-0002-6619-6507ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Multi-level attention graph neural network based on co-expression gene modules for disease diagnosis and prognosisabstractMOTIVATION: Advanced deep learning techniques have been widely applied in disease diagnosis and prognosis with clinical omics, especially gene expression data. In the regulation of biological processes and disease progression, genes often work interactively rather than individually. Therefore, investigating gene association information and co-functional gene modules can facilitate disease state prediction. RESULTS: To explore the gene modules and inter-gene relational information contained in the omics data, we propose a novel multi-level attention graph neural network (MLA-GNN) for disease diagnosis and prognosis. Specifically, we format omics data into co-expression graphs via weighted correlation network analysis, and then construct multi-level graph features, finally fuse them through a well-designed multi-level graph feature fully fusion module to conduct predictions. For model interpretation, a novel full-gradient graph saliency mechanism is developed to identify the disease-relevant genes. MLA-GNN achieves state-of-the-art performance on transcriptomic data from TCGA-LGG/TCGA-GBM and proteomic data from coronavirus disease 2019 (COVID-19)/non-COVID-19 patient sera. More importantly, the relevant genes selected by our model are interpretable and are consistent with the clinical understanding. AVAILABILITYAND IMPLEMENTATION: The codes are available at https://github.com/TencentAILabHealthcare/MLA-GNN. Xiaohan Xing, Fan Yang 0081, Jun Zhang 0018, Yu Zhao 0009, Mingxuan Gao, Junzhou Huang, Jianhua Yao 0001 |
Bioinform. | 6 |
| 2021 | scSparkXMBD: High-Performance scRNA-seq Data Processing with SparkabstractHigh-throughput single-cell RNA sequencing (scRNA-seq) data processing pipelines integrate multiple modules to transform raw scRNA-seq data to gene expression matrices, including barcode processing, sequence quality control, genome alignment and transcript quantification. With the rapid growth in data volume, the speed of scRNA-seq data processing pipeline has become a major bottleneck to large-scale scRNA-seq studies. We present scSparkXMBD1(denoted as scSpark), a cloud computing based scRNA-seq data processing pipeline. By leveraging the in-memory computing capability of Apache Spark, scSpark significantly improves the processing speed of scRNA-seq data, and achieves around 5-20 times faster than the state-of-the-art processing pipelines under the same CPU core consumption. In addition, thanks to the inherent scalability of Spark in a cloud computing environment, scSpark can further reduce the processing time for a typical scRNA-seq dataset (e.g., 640 million reads) from hours to minutes when multiple computer nodes (e.g., 16) are used. Biological evaluation also confirmed that the results generated by scSpark are highly consistent with existing scRNA-seq data processing pipelines.1XMBD refers to Xiamen Big Data, which is a biomedical open software initiative in the National Institute for Data Science in Health and Medicine, Xiamen University, China Mingxuan Gao, Lixuan Tan, Hongjin Liu, Yating Lin, Rongshan Yu |
BIBM | 2 |
| 2021 | An Interpretable Multi-Level Enhanced Graph Attention Network for Disease Diagnosis with Gene Expression DataabstractClinical omics, especially gene expression data, have been widely studied and successfully applied for disease diagnosis using machine learning techniques. As genes often work interactively rather than individually, investigating co-functional gene modules can improve our understanding of disease mechanisms and facilitate disease state prediction. To this end, we in this paper propose a novel Multi-Level Enhanced Graph ATtention (MLE-GAT) network to explore the gene modules and intergene relational information contained in the omics data. In specific, we first format the omics data of each patient into co-expression graphs using weighted correlation network analysis (WGCNA) and then feed them to a well-designed multi-level graph feature fully fusion (MGFFF) module for disease diagnosis. For model interpretation, we develop a novel full-gradient graph saliency (FGS) mechanism to identify the disease-relevant genes. Comprehensive experiments show that our proposed MLE-GAT achieves state-of-the-art performance on transcriptomics data from TCGA-LGG/TCGA-GBM and proteomics data from COVID-19/non-COVID-19 patient sera. Xiaohan Xing, Fan Yang 0081, Jun Zhang 0018, Yu Zhao 0009, Mingxuan Gao, Junzhou Huang, Jianhua Yao 0001 |
BIBM | 6 |
| 2021 | DT-MIL: Deformable Transformer for Multi-instance Learning on Histopathological Image
Fan Yang 0081, Yu Zhao 0009, Xiaohan Xing, Jun Zhang 0018, Mingxuan Gao, Junzhou Huang, Liansheng Wang 0002, Jianhua Yao 0001 |
MICCAI (8) | 6 |
| 2021 | Comparison of high-throughput single-cell RNA sequencing data processing pipelinesabstractWith the development of single-cell RNA sequencing (scRNA-seq) technology, it has become possible to perform large-scale transcript profiling for tens of thousands of cells in a single experiment. Many analysis pipelines have been developed for data generated from different high-throughput scRNA-seq platforms, bringing a new challenge to users to choose a proper workflow that is efficient, robust and reliable for a specific sequencing platform. Moreover, as the amount of public scRNA-seq data has increased rapidly, integrated analysis of scRNA-seq data from different sources has become increasingly popular. However, it remains unclear whether such integrated analysis would be biassed if the data were processed by different upstream pipelines. In this study, we encapsulated seven existing high-throughput scRNA-seq data processing pipelines with Nextflow, a general integrative workflow management framework, and evaluated their performance in terms of running time, computational resource consumption and data analysis consistency using eight public datasets generated from five different high-throughput scRNA-seq platforms. Our work provides a useful guideline for the selection of scRNA-seq data processing pipelines based on their performance on different real datasets. In addition, these guidelines can serve as a performance evaluation framework for future developments in high-throughput scRNA-seq data processing. Mingxuan Gao, Mingyi Ling, Xinwei Tang, Rongshan Yu |
Briefings Bioinform. | 1 |
| 2021 | Diamond: a multi-modal DIA mass spectrometry data processing pipelineabstractSUMMARY: Currently, various software tools are used to support two mainstream workflows for data-independent acquisition (DIA) mass spectrometry (MS) data processing, namely, spectrum-centric scoring (SCS) and peptide-centric scoring (PCS). However, a fully automatic, easily reproducible and freely accessible pipeline that simultaneously integrates SCS and PCS strategies and supports both library-free and library-based modes is absent. We developed Diamond, a Nextflow-based, containerized, multi-modal DIA-MS data processing pipeline for peptide identification and quantification. Diamond integrated two mainstream workflows for DIA data analysis, namely, SCS and PCS, for use cases both with and without assay libraries. This multi-modal pipeline serves as a versatile, easy-to-use and easily extendable toolbox for large-scale DIA data processing. AVAILABILITY: Diamond is hosted on GitHub (https://github.com/xmuyulab/Diamond) and is released under the highly permissive MIT license to encourage further customization and modification. The Docker image for Diamond is freely accessible at https://hub.docker.com/r/zeroli/diamond. Chenxin Li, Mingxuan Gao, Chuanqi Zhong, Rongshan Yu |
Bioinform. | 2 |