Ximei Luo

dblp:213/6693 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-2956-6799ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021
YearPublicationVenuePosition
2026 AutoGERN: single-cell RNA-seq gene regulatory network inference via explicit link modeling and adaptive architectures
abstract
MOTIVATION: Single-cell RNA sequencing (scRNA-seq) enables transcriptome-wide profiling at single-cell resolution, revealing heterogeneous regulatory programs and making gene regulatory network (GRN) inference both central and challenging. Recent graph neural network (GNN)-based approaches for GRN inference typically model edges only implicitly (for example, via concatenated node embeddings), which limits their ability to capture complex regulatory dependencies. In addition, distributional shifts across scRNA-seq datasets make a single fixed GNN architecture poorly suited for broad generalization. RESULTS: We present AutoGERN, a GNN framework tailored for GRN inference from scRNA-seq data. AutoGERN explicitly models regulatory information in the message-passing space and learns expressive link (edge) embeddings, which a lightweight multilayer perceptron uses to score gene-gene regulatory associations. To enhance flexibility and representational power, AutoGERN employs dual message-passing spaces (within-layer and cross-layer) and integrates a robust AutoGNN-based architecture search to adapt the network design to differing dataset distributions. Extensive experiments on multiple real scRNA-seq datasets demonstrate that AutoGERN consistently achieves superior performance and robustness compared with state-of-the-art baselines. AVAILABILITY AND IMPLEMENTATION: The code and data of AutoGERN are available on GitHub at https://github.com/JChander/AutoGERN and on Zenodo at https://doi.org/10.5281/zenodo.18659807.
Jiacheng Wang 0009, Yaojia Chen, Quan Zou 0001, Ximei Luo
Bioinform.4
2025 Assessment and applications of joint profiling of single-cell chromatin accessibility and transcriptome
abstract
Joint profiling technologies combining single-cell chromatin accessibility (CA) and transcriptome sequencing enable cellular heterogeneity analysis from both gene and cis-regulatory element perspectives, greatly advancing molecular biology at a cellular resolution. These techniques have been used to construct gene regulatory networks across diverse cell types and biological tissues, contributing significantly to the mapping of cell developmental trajectories. In this review, we summarize existing single-cell joint profiling methods for CA and transcriptomics and systematically evaluate the data quality of each modality using consistent criteria: the median number of genes detected per cell (RNA) and the median number of accessible peaks per cell (ATAC). Furthermore, we examine relevant bioinformatics tools and highlight their applications in various omics research contexts. Finally, we discuss the current limitations of joint profiling technologies, prospects for future improvement, the extensibility of computational tools, and the potential for co-assaying with additional omics data.
Jiechen Wang, Quan Zou 0001, Ximei Luo
Briefings Bioinform.5
2025 BridgeSyn: a bridging fusion framework for drug combination synergy prediction
abstract
Drug combination is a promising therapeutic strategy for complex diseases. However, only a small fraction of potential drug combinations exhibit true synergistic effects, making the prediction of drug synergy a critical yet challenging task. In this study, we propose BridgeSyn, a novel bridge fusion framework for drug synergy prediction. BridgeSyn leverages the knowledge from pretrained biological language models to enrich both drug compound and cell line representations. We introduce a bridging fusion mechanism that employs a set of shared latent tokens derived from global features, serving as a semantic interface to effectively fuse the representations of drug pairs and cell lines. By combining biological prior knowledge with this fusion strategy, BridgeSyn can capture complex biological interactions and achieve superior prediction results. Extensive experiments on two public datasets demonstrate that BridgeSyn consistently outperforms existing computation methods.
Suwan Mao, Quan Zou 0001, Xi Su, Junjie Wang 0005, Ximei Luo
Briefings Bioinform.9
2025 ReAlign-P: a vertical iterative realignment method for protein multiple sequence alignment
abstract
MOTIVATION: Reliable protein multiple sequence alignment (MSA) is essential for downstream biomedical research and directly impacts the accuracy of analytical results. However, protein sequences often exhibit low similarity and complex alignment patterns, and existing general alignment tools frequently fall short in terms of accuracy. Many current realignment methods are outdated, suffering from issues such as code obsolescence and inadequate precision. As a result, there is a pressing need for realignment methods that can better address these challenges. RESULTS: This study introduces ReAlign-P, a realignment tool designed specifically for protein MSA. ReAlign-P first divides the initial alignment into three regions and applies a novel vertical iterative realignment strategy to optimize the more conserved middle region. This method is by default compatible with MUSCLE5 for realignment, leading to a significant improvement in accuracy. We evaluated initial alignments generated using 10 different MSA parameter configurations across four protein benchmark datasets. The results demonstrate that ReAlign-P consistently outperforms or matches the quality of the initial alignments in all cases. In contrast, RASCAL-the only other currently functional protein realignment tool-sometimes even reduces alignment quality. ReAlign-P not only delivers more substantial improvements but also exhibits greater stability, effectively addressing the gap in available protein realignment tools. AVAILABILITY AND IMPLEMENTATION: The source code and test data for ReAlign-P are available on GitHub (https://github.com/malabz/ReAlign-P).
Yixiao Zhai, Pinglu Zhang, Quan Zou 0001, Ximei Luo
Bioinform.4
2025 Identifying the DNA methylation preference of transcription factors using ProtBERT and SVM
abstract
Transcription factors (TFs) can affect gene expression by binding to certain specific DNA sequences. This binding process of TFs may be modulated by DNA methylation. A subset of TFs that serve as methylation readers preferentially binds to certain methylated DNA and is defined as TFPM. The identification of TFPMs enhances our understanding of DNA methylation's role in gene regulation. However, their experimental identification is resource-demanding. In this study, we propose a novel two-step computational approach to classify TFs and TFPMs. First, we employed a fine-tuned ProtBERT model to differentiate between the classes of TFs and non-TFs. Second, we combined the Reduced Amino Acid Category (RAAC) with K-mer and SVM to predict the potential of TFs to bind to methylated DNA. Comparative experiments demonstrate that our proposed methods outperform all existing approaches and emphasize the efficiency of our computational framework in classifying TFs and TFPMs. Cross-species validation on an independent mouse dataset further demonstrates the generalizability of our proposed framework In addition, we conducted predictions on all human transcription factors and found that most of the top 20 proteins belong to the Krueppel C2H2-type Zinc-finger family. So far, some studies have demonstrated a partial correlation between this family and DNA methylation and confirmed the preference of some of its members, thereby showing the robustness of our approach.
Quan Zou 0001, Antony Stalin, Ximei Luo
PLoS Comput. Biol.5
2023 Predicting active enhancers with DNA methylation and histone modification
abstract
BACKGROUND: Enhancers play a crucial role in gene regulation, and some active enhancers produce noncoding RNAs known as enhancer RNAs (eRNAs) bi-directionally. The most commonly used method for detecting eRNAs is CAGE-seq, but the instability of eRNAs in vivo leads to data noise in sequencing results. Unfortunately, there is currently a lack of research focused on the noise inherent in CAGE-seq data, and few approaches have been developed for predicting eRNAs. Bridging this gap and developing widely applicable eRNA prediction models is of utmost importance. RESULTS: In this study, we proposed a method to reduce false positives in the identification of eRNAs by adjusting the statistical distribution of expression levels. We also developed eRNA prediction models using joint gene expressions, DNA methylation, and histone modification. These models achieved impressive performance with an AUC value of approximately 0.95 for intra-cell prediction and 0.9 for cross-cell prediction. CONCLUSIONS: Our method effectively attenuates the noise generated by stochastic RNA production, resulting in more accurate detection of eRNAs. Furthermore, our eRNA prediction model exhibited significant accuracy in both intra-cell and cross-cell validation, highlighting its robustness and potential application in various cellular contexts.
Ximei Luo, Yan Liu 0085, Quan Zou 0001, Ying Zhang 0060, Lei Xu 0047
BMC Bioinform.1
2023 Recall DNA methylation levels at low coverage sites using a CNN model in WGBS
abstract
DNA methylation is an important regulator of gene transcription. WGBS is the gold-standard approach for base-pair resolution quantitative of DNA methylation. It requires high sequencing depth. Many CpG sites with insufficient coverage in the WGBS data, resulting in inaccurate DNA methylation levels of individual sites. Many state-of-arts computation methods were proposed to predict the missing value. However, many methods required either other omics datasets or other cross-sample data. And most of them only predicted the state of DNA methylation. In this study, we proposed the RcWGBS, which can impute the missing (or low coverage) values from the DNA methylation levels on the adjacent sides. Deep learning techniques were employed for the accurate prediction. The WGBS datasets of H1-hESC and GM12878 were down-sampled. The average difference between the DNA methylation level at 12× depth predicted by RcWGBS and that at >50× depth in the H1-hESC and GM2878 cells are less than 0.03 and 0.01, respectively. RcWGBS performed better than METHimpute even though the sequencing depth was as low as 12×. Our work would help to process methylation data of low sequencing depth. It is beneficial for researchers to save sequencing costs and improve data utilization through computational methods.
Ximei Luo, Yansu Wang, Quan Zou 0001, Lei Xu 0047
PLoS Comput. Biol.1
2022 scESI: evolutionary sparse imputation for single-cell transcriptomes from nearest neighbor cells
abstract
The ubiquitous dropout problem in single-cell RNA sequencing technology causes a large amount of data noise in the gene expression profile. For this reason, we propose an evolutionary sparse imputation (ESI) algorithm for single-cell transcriptomes, which constructs a sparse representation model based on gene regulation relationships between cells. To solve this model, we design an optimization framework based on nondominated sorting genetics. This framework takes into account the topological relationship between cells and the variety of gene expression to iteratively search the global optimal solution, thereby learning the Pareto optimal cell-cell affinity matrix. Finally, we use the learned sparse relationship model between cells to improve data quality and reduce data noise. In simulated datasets, scESI performed significantly better than benchmark methods with various metrics. By applying scESI to real scRNA-seq datasets, we discovered scESI can not only further classify the cell types and separate cells in visualization successfully but also improve the performance in reconstructing trajectories differentiation and identifying differentially expressed genes. In addition, scESI successfully recovered the expression trends of marker genes in stem cell differentiation and can discover new cell types and putative pathways regulating biological processes.
Qiaoming Liu, Ximei Luo, Jie Li 0055, Guohua Wang 0001
Briefings Bioinform.2
2022 Effector-GAN: prediction of fungal effector proteins based on pretrained deep representation learning methods and generative adversarial networks
abstract
MOTIVATION: Phytopathogenic fungi secrete effector proteins to subvert host defenses and facilitate infection. Systematic analysis and prediction of candidate fungal effector proteins are crucial for experimental validation and biological control of plant disease. However, two problems are still considered intractable to be solved in fungal effector prediction: one is the high-level diversity in effector sequences that increases the difficulty of protein feature learning, and the other is the class imbalance between effector and non-effector samples in the training dataset. RESULTS: In our study, pretrained deep representation learning methods are presented to represent multiple characteristics of sequences for predicting fungal effectors and generative adversarial networks are adapted to create synthetic feature samples to address the data imbalance problem. Compared with the state-of-the-art fungal effector prediction methods, Effector-GAN shows an overall improvement in accuracy in the independent test set. AVAILABILITY AND IMPLEMENTATION: Effector-GAN offers a user-friendly interface to inspect potential fungal effector proteins (http://lab.malab.cn/~wys/webserver/Effector-GAN). The Python script can be downloaded from http://lab.malab.cn/~wys/gitlab/effector-gan. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yansu Wang, Ximei Luo, Quan Zou 0001
Bioinform.2