EDBT 2026 Demo / reviewers in the wild / expert
Xuan Guo 0004
dblp:82/2905-4
· DBLP profile ↗
18ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-2777-4482ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Frequency Regulated Channel-Spatial Attention module for improved image classification
Chengyuan Zhuang, Xiaohui Yuan 0001, Lichuan Gu, Zhenchun Wei, Yuqi Fan 0001, Xuan Guo 0004 |
Expert Syst. Appl. | 6 |
| 2025 | A Multiple Attention Layer-shareable Method for Link Prediction in Multilayer NetworksabstractLink prediction in multilayer networks aims to predict missing links at the target layer by incorporating structural information from both auxiliary layers and the target layer. Existing methods tend to learn layer-specific knowledge to maximize the link prediction performance on a specific network layer. However, they have difficulty incorporating multilayer structural information to improve the link prediction performance. Therefore, we propose a Multiple Attention Layer-shareable Method (MALM) for link prediction in multilayer networks, which consists of a feature encoder, a knowledge learner, and a fusion predictor. The feature encoder introduces multiple attention mechanisms to encode the feature representations of links by differentiating the importance of structural information for each link. In cooperation with the feature encoder, the knowledge learner splits the link prediction tasks into different layers and employs meta-learning to learn layer-shareable knowledge from these link prediction tasks. Finally, the fusion predictor combines the learned layer-shareable knowledge with the layer-specific knowledge at the target layer for link prediction. Experiments on real-world datasets demonstrate that the proposed MALM outperforms existing state-of-the-art baselines in link prediction in multilayer networks. Huan Wang 0005, Yu Teng, Lingsong Qin, Xuan Guo 0004, Po Hu 0001 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2024 | SEMQuant: Extending Sipros-Ensemble with Match-Between-Runs for Comprehensive Quantitative Metaproteomics
Bailu Zhang, Shichao Feng, Manushi Parajuli, Chongle Pan, Xuan Guo 0004 |
ISBRA (3) | 6 |
| 2023 | Transformer-Based De Novo Peptide Sequencing for Data-Independent Acquisition Mass SpectrometryabstractTandem mass spectrometry (MS/MS) stands as the predominant high-throughput technique for comprehensively analyzing protein content within biological samples. This methodology is a cornerstone driving the advancement of proteomics. In recent years, substantial strides have been made in Data-Independent Acquisition (DIA) strategies, facilitating impartial and non-targeted fragmentation of precursor ions. The DIA-generated MS/MS spectra present a formidable obstacle due to their inherent high multiplexing nature. Each spectrum encapsulates fragmented product ions originating from multiple precursor peptides. This intricacy poses a particularly acute challenge in de novo peptide/protein sequencing, where current methods are ill-equipped to address the multiplexing conundrum. In this paper, we introduce Casanovo-DIA, a deep-learning model based on transformer architecture. It deciphers peptide sequences from DIA mass spectrometry data. Our results show significant improvements over existing STOA methods, including DeepNovo-DIA and PepNet. Casanovo-DIA enhances precision by 15.14% to 34.8%, recall by 11.62% to 31.94% at the amino acid level, and boosts precision by 59% to 81.36% at the peptide level. Integrating DIA data and our Casanovo-DIA model holds considerable promise to uncover novel peptides and more comprehensive profiling of biological samples. Casanovo-DIA is freely available under the GNU GPL license at https://github.com/Biocomputing-Research-Group/Casanovo-DIA. Shiva Ebrahimi, Xuan Guo 0004 |
BIBE | 2 |
| 2023 | Meta-learning adaptation network for few-shot link prediction in heterogeneous social networks
Huan Wang 0005, Jiaxin Mi, Xuan Guo 0004, Po Hu 0001 |
Inf. Process. Manag. | 3 |
| 2022 | Deep Learning Based MS2 Feature Detection for Data-Independent Shotgun ProteomicsabstractAccuracy of peptide identification in LC-MS analysis is crucial for information regarding the aspects of proteins that aid in biomarker discovery and the profiling of complex proteomes. The detection of peptide fragment ions in tandem mass spectrometry is still challenging given that current tools were not created or tested for the low-abundance, low-peak fragments of peptides found in MS2 data. Feature detection, a crucial pre-processing step in the LC-MS analysis pipeline that quantifies peptides by their mass-to-charge ratio, retention time, and intensity, is particularly challenging due to the overlapping nature of peptides and weak signals that are often indistinguishable from noises, thus creating a reliance on rigid mathematical structures and heuristics. In this study, we developed a deep-learning-based model with an innovative sliding window process that enables high-resolution processing of quantitative MS/MS data to conduct MS2 feature detection. Experimental results show that our model can produce more accurate values and identifications than existing feature detection tools, as well as a high rate of true positive features quantified. Therefore, we believe that our model illustrates the advantages of deep learning techniques applied towards computational proteomics. Jonathan He, Olivia Liu, Xuan Guo 0004 |
BIBM | 3 |
| 2022 | IDIA: An Integrative Signal Extractor for Data-Independent Acquisition ProteomicsabstractIn proteomics, data-independent acquisition (DIA) has been shown to provide less biased and more reproducible results than data-dependent acquisition. Recently, many researchers have developed a series of methods to identify peptides and proteins by using spectrum libraries for DIA data. However, spectrum libraries are not always available for novel organisms or microbial communities. To detect peptides and proteins without a spectrum library, we developed IDIA, a library-free method using DIA data to generate pseudo-spectra that can be searched using conventional sequence database searching software. IDIA integrates two isotopic trace detection strategies and employs B-spline and Gaussian filters to help extract high-quality pseudo-spectra from the complex DIA data. The experimental results on human and yeast data demonstrated that our approach remarkably produced more peptide and protein identifications than the two state-of-the-art library-free methods, i.e., DIA-Umpire and Group-DIA. IDIA is freely available under the GNU GPL license at https://github.com/Biocomputing-Research-Group/IDIA. Jiancheng Li, Chongle Pan, Xuan Guo 0004 |
BIBM | 3 |
| 2022 | FineFDR: Fine-grained Taxonomy-specific False Discovery Rates Control in MetaproteomicsabstractMicrobial community proteomics, also termed metaproteomics, investigates all proteins expressed by a microbiota. Tandem mass spectrometry (MS/MS) is the typical method for identifying proteins in metaproteomics, which involves searching the mass spectra against a protein sequence database. A major post-analysis step is controlling the false discovery rate (FDR), i.e., the ratio of false positives to the total number of annotations. The current popular target-decoy FDR estimation method treats all the peptides and proteins equally and overlooks that they could have varied probabilities of being identified. In this study, we report FineFDR, a framework for FDR assessment at fine-grained levels with taxonomy information considered. FineFDR groups the identified peptide-spectrum matches, peptides, and proteins from different taxonomic units and estimates the FDR in each group separately. Empirical experiments on the simulated and real-world data sets demonstrate that our FineFDR achieved higher precision and more peptide and protein identifications when compared to the state-of-the-art methods, such as Comet, Percolator, TIDD, and Tailor. FineFDR is freely available under the GNU GPL license at https://github.com/Biocomputing-Research-Group/FDR. Shengze Wang 0004, Shichao Feng, Chongle Pan, Xuan Guo 0004 |
BIBM | 4 |
| 2022 | MetaLP: An integrative linear programming method for protein inference in metaproteomicsabstractMetaproteomics based on high-throughput tandem mass spectrometry (MS/MS) plays a crucial role in characterizing microbiome functions. The acquired MS/MS data is searched against a protein sequence database to identify peptides, which are then used to infer a list of proteins present in a metaproteome sample. While the problem of protein inference has been well-studied for proteomics of single organisms, it remains a major challenge for metaproteomics of complex microbial communities because of the large number of degenerate peptides shared among homologous proteins in different organisms. This challenge calls for improved discrimination of true protein identifications from false protein identifications given a set of unique and degenerate peptides identified in metaproteomics. MetaLP was developed here for protein inference in metaproteomics using an integrative linear programming method. Taxonomic abundance information extracted from metagenomics shotgun sequencing or 16s rRNA gene amplicon sequencing, was incorporated as prior information in MetaLP. Benchmarking with mock, human gut, soil, and marine microbial communities demonstrated significantly higher numbers of protein identifications by MetaLP than ProteinLP, PeptideProphet, DeepPep, PIPQ, and Sipros Ensemble. In conclusion, MetaLP could substantially improve protein inference for complex metaproteomes by incorporating taxonomic abundance information in a linear programming model. Shichao Feng, Hong-Long Ji, Huan Wang 0005, Bailu Zhang, Ryan Sterzenbach, Chongle Pan, Xuan Guo 0004 |
PLoS Comput. Biol. | 7 |
| 2022 | EditorialabstractThis special section gives the opportunity to know recent advances in the application of intelligent optimization algorithms in genomics and precision medicine. Precision medicine is designed to optimize the pathway for diagnosis, therapeutic intervention, and prognosis by using multidimensional biological datasets that capture individual variability in genes, function, and environment. Recent advances in -omics technologies provide substantial novel opportunities to study and/or identify biomarkers of chronic diseases by interpreting multi-omics data, including transcriptomics, epigenomics, genomics, and proteomics, that, together may improve understanding of precision medicine. Precision medicine is drugs or treatments designed for small groups, rather than large populations, based on characteristics, such as medical history, genetic makeup, and data recorded by wearable devices. The use of genomic data can support precision medicine to enable clinicians to predict the most appropriate course of action quickly, efficiently, and accurately for a patient. This offers clinicians the opportunity to tailor early interventions to each patient more carefully. Xiuzhen Huang, Yu Zhang 0150, Xuan Guo 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Identifying and Evaluating Anomalous Structural Change-based Nodes in Generalized Dynamic Social NetworksabstractRecently, dynamic social network research has attracted a great amount of attention, especially in the area of anomaly analysis that analyzes the anomalous change in the evolution of dynamic social networks. However, most of the current research focused on anomaly analysis of the macro representation of dynamic social networks and failed to analyze the nodes that have anomalous structural changes at a micro level. To identify and evaluate anomalous structural change-based nodes in generalized dynamic social networks that only have limited structural information, this research considers undirected and unweighted graphs and develops a multiple-neighbor superposition similarity method ( ), which mainly consists of a multiple-neighbor range algorithm ( ) and a superposition similarity fluctuation algorithm ( ). introduces observation nodes, characterizes the structural similarities of nodes within multiple-neighbor ranges, and proposes a new multiple-neighbor similarity index on the basis of extensional similarity indices. Subsequently, maximally reflects the structural change of each node, using a new superposition similarity fluctuation index from the perspective of diverse multiple-neighbor similarities. As a result, based on and , not only identifies anomalous structural change-based nodes by detecting the anomalous structural changes of nodes but also evaluates their anomalous degrees by quantifying these changes. Results obtained by comparing with state-of-the-art methods via extensive experiments show that can accurately identify anomalous structural change-based nodes and evaluate their anomalous degrees well. Huan Wang 0005, Chunming Qiao, Xuan Guo 0004, Lei Fang 0001, Ying Sha, Zhiguo Gong |
ACM Trans. Web | 3 |
| 2018 | Sipros Ensemble improves database searching and filtering for complex metaproteomicsabstractMotivation: Complex microbial communities can be characterized by metagenomics and metaproteomics. However, metagenome assemblies often generate enormous, and yet incomplete, protein databases, which undermines the identification of peptides and proteins in metaproteomics. This challenge calls for increased discrimination of true identifications from false identifications by database searching and filtering algorithms in metaproteomics. Results: Sipros Ensemble was developed here for metaproteomics using an ensemble approach. Three diverse scoring functions from MyriMatch, Comet and the original Sipros were incorporated within a single database searching engine. Supervised classification with logistic regression was used to filter database searching results. Benchmarking with soil and marine microbial communities demonstrated a higher number of peptide and protein identifications by Sipros Ensemble than MyriMatch/Percolator, Comet/Percolator, MS-GF+/Percolator, Comet & MyriMatch/iProphet and Comet & MyriMatch & MS-GF+/iProphet. Sipros Ensemble was computationally efficient and scalable on supercomputers. Availability and implementation: Freely available under the GNU GPL license at http://sipros.omicsbio.org. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Xuan Guo 0004, Qiuming Yao, Ryan S. Mueller, Jimmy K. Eng, David L. Tabb, IV William Judson Hervey, Chongle Pan |
Bioinform. | 1 |
| 2017 | Searching Genome-Wide Multi-Locus Associations for Multiple Diseases Based on Bayesian InferenceabstractTaking the advantage of high-throughput single nucleotide polymorphism (SNP) genotyping technology, large genome-wide association studies (GWASs) have been considered to hold promise for unraveling complex relationships between genotypes and phenotypes. Current multi-locus-based methods are insufficient to detect interactions with diverse genetic effects on multifarious diseases. Also, statistic tests for high-order epistasis ( ≥ 2 SNPs) raise huge computational and analytical challenges because the computation increases exponentially as the growth of the cardinality of SNPs combinations. In this paper, we provide a simple, fast and powerful method, named DAM, using Bayesian inference to detect genome-wide multi-locus epistatic interactions in multiple diseases. Experimental results on simulated data demonstrate that our method is powerful and efficient. We also apply DAM on two GWAS datasets from WTCCC, i.e., Rheumatoid Arthritis and Type 1 Diabetes, and identify some novel findings. Therefore, we believe that our method is suitable and efficient for the full-scale analysis of multi-disease-related interactions in GWASs. Xuan Guo 0004, Jing Zhang 0010, Zhipeng Cai 0001, Ding-Zhu Du, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2015 | DAM: A Bayesian Method for Detecting Genome-wide Associations on Multiple Diseases
Xuan Guo 0004, Jing Zhang 0010, Zhipeng Cai 0001, Ding-Zhu Du, Yi Pan 0001 |
ISBRA | 1 |
| 2015 | DNA AS X: An Information-Coding-Based Model to Improve the Sensitivity in Comparative Gene Analysis
Ning Yu 0004, Xuan Guo 0004, Feng Gu 0001, Yi Pan 0001 |
ISBRA | 2 |
| 2015 | Searching High-Order SNP Combinations for Complex Diseases Based on Energy Distribution DifferenceabstractSingle nucleotide polymorphisms, a dominant type of genetic variants, have been used successfully to identify defective genes causing human single gene diseases. However, most common human diseases are complex diseases and caused by gene-gene and gene-environment interactions. Many SNP-SNP interaction analysis methods have been introduced but they are not powerful enough to discover interactions more than three SNPs. The paper proposes a novel method that analyzes all SNPs simultaneously. Different from existing methods, the method regards an individual's genotype data on a list of SNPs as a point with a unit of energy in a multi-dimensional space, and tries to find a new coordinate system where the energy distribution difference between cases and controls reaches the maximum. The method will find different multiple SNPs combinatorial patterns between cases and controls based on the new coordinate system. The experiment on simulated data shows that the method is efficient. The tests on the real data of age-related macular degeneration (AMD) disease show that it can find out more significant multi-SNP combinatorial patterns than existing methods. Jianxin Wang 0001, Alex Zelikovsky, Xuan Guo 0004, Minzhu Xie, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2014 | Cloud computing for detecting high-order genome-wide epistatic interaction via dynamic clusteringabstractBACKGROUND: Taking the advantage of high-throughput single nucleotide polymorphism (SNP) genotyping technology, large genome-wide association studies (GWASs) have been considered to hold promise for unravelling complex relationships between genotype and phenotype. At present, traditional single-locus-based methods are insufficient to detect interactions consisting of multiple-locus, which are broadly existing in complex traits. In addition, statistic tests for high order epistatic interactions with more than 2 SNPs propose computational and analytical challenges because the computation increases exponentially as the cardinality of SNPs combinations gets larger. RESULTS: In this paper, we provide a simple, fast and powerful method using dynamic clustering and cloud computing to detect genome-wide multi-locus epistatic interactions. We have constructed systematic experiments to compare powers performance against some recently proposed algorithms, including TEAM, SNPRuler, EDCF and BOOST. Furthermore, we have applied our method on two real GWAS datasets, Age-related macular degeneration (AMD) and Rheumatoid arthritis (RA) datasets, where we find some novel potential disease-related genetic factors which are not shown up in detections of 2-loci epistatic interactions. CONCLUSIONS: Experimental results on simulated data demonstrate that our method is more powerful than some recently proposed methods on both two- and three-locus disease models. Our method has discovered many novel high-order associations that are significantly enriched in cases from two real GWAS datasets. Moreover, the running time of the cloud implementation for our method on AMD dataset and RA dataset are roughly 2 hours and 50 hours on a cluster with forty small virtual machines for detecting two-locus interactions, respectively. Therefore, we believe that our method is suitable and effective for the full-scale analysis of multiple-locus epistatic interactions in GWAS. Xuan Guo 0004, Meng Yu 0001, Ning Yu 0004, Yi Pan 0001 |
BMC Bioinform. | 1 |
| 2013 | Cloud Computing for De Novo Metagenomic Sequence Assembly
Xuan Guo 0004, Meng Yu 0001, Yi Pan 0001 |
ISBRA | 1 |