Xiaofei Nan

dblp:37/7094 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0001-9145-1111ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 PGDiff: A Physics-Guided Equivariant Diffusion Model for Structure-Based Drug Design
abstract
Structure-based drug design (SBDD) is a critical approach in the development of new drugs, aiming to generate ligand molecules that can effectively bind to specific target proteins. Recent progress in diffusion-based generative models has enabled target-conditioned 3D molecular generation. However, many existing methods overlook the fundamental physical interactions that are essential for accurate protein-ligand binding. As a result, they often generate molecules with unrealistic geometries or weak binding affinities, limiting their applicability in real-world drug design. In this work, we introduce PGDiff, a physics-informed generative framework that integrates both structural and energetic priors into a 3D denoising diffusion process. PGDiff explicitly models key non-covalent interactions, including van der Waals forces, hydrogen bonding, and hydrophobic effects, to promote the generation of ligands with realistic geometries and favorable binding energetics. To maintain molecular symmetry, the model utilizes an SE(3)-equivariant graph neural network that preserves rotational and translational equivariance during the entire generation process. Experimental results on the CrossDocked2020 benchmark demonstrate that PGDiff achieves competitive or superior performance compared to state-of-the-art baselines across multiple key metrics These findings emphasize the critical role of physical constraints in guiding molecular generation for structure-based drug discovery.
Xiaofei Nan, Xuezhen Liu, Xing You, Yongsheng Du, Chengxiang Ji, Jinshuai Song
BIBM1
2025 Enhanced Corneal Endothelial Cell Segmentation via Frequency-Selected Residual Fourier Diffusion Models
abstract
Segmenting corneal endothelial cells in conditions like Fuchs endothelial dystrophy (FED) is challenging due to guttae obscuring cell details and complicating imaging. This is further compounded by labor-intensive manual annotations and a lack of large annotated datasets. To address these issues, we introduce a novel two-stage framework using Denoising Diffusion Probabilistic Models (DDPMs) for generating training pairs of corneal endothelial cell images. In the first stage, we generate synthetic endothelial labels, which are then used to guide the production of high-resolution corneal images in the second stage. We also present the Fourier Residual Block with Frequency Selection (FRB-FS), which enhances important high-frequency details for clearer textures and edges, while suppressing irrelevant low-frequency components. This is the first application of diffusion models to corneal endothelial cell segmentation. Extensive experiments and ablation studies on two benchmark datasets demonstrate the effectiveness of our framework.
Xiaofei Nan, Yunze Wang, Zhenkai Gao, Jingxin Liu 0005
ICASSP2
2025 CGLDM: A Conditional Geometric Latent Diffusion Model for 3D Molecular Generation
Xuezhen Liu, Chuanghui Wang, Xing You, Chengxiang Ji, Xiaofei Nan
ICIC (25)5
2025 DefGAN-Im: Adversarial Industrial Defect Synthesis with Conditional Feature Disentanglement on Imbalanced Datasets
Xiaofei Nan, Jinyu Fan, Wenyang Li, Xiaoheng Jiang
ICIC (11)1
2025 HSSPPI: hierarchical and spatial-sequential modeling for PPIs prediction
abstract
MOTIVATION: Protein-protein interactions play a fundamental role in biological systems. Accurate detection of protein-protein interaction sites (PPIs) remains a challenge. And, the methods of PPIs prediction based on biological experiments are expensive. Recently, a lot of computation-based methods have been developed and made great progress. However, current computational methods only focus on one form of protein, using only protein spatial conformation or primary sequence. And, the protein's natural hierarchical structure is ignored. RESULTS: In this study, we propose a novel network architecture, HSSPPI, through hierarchical and spatial-sequential modeling of protein for PPIs prediction. In this network, we represent protein as a hierarchical graph, in which a node in the protein is a residue (residue-level graph) and a node in the residue is an atom (atom-level graph). Moreover, we design a spatial-sequential block for capturing complex interaction relationships from spatial and sequential forms of protein. We evaluate HSSPPI on public benchmark datasets and the predicting results outperform the comparative models. This indicates the effectiveness of hierarchical protein modeling and also illustrates that HSSPPI has a strong feature extraction ability by considering spatial and sequential information simultaneously. AVAILABILITY AND IMPLEMENTATION: The code of HSSPPI is available at https://github.com/biolushuai/Hierarchical-Spatial-Sequential-Modeling-of-Protein.
Yuguang Li, Zhen Tian 0004, Xiaofei Nan, Shoutao Zhang, Qinglei Zhou
Briefings Bioinform.3
2025 A model-based infrared and visible image fusion network with cooperative optimization
Tianqing Hu, Xiaofei Nan, Qinglei Zhou, Renhao Lin
Expert Syst. Appl.2
2023 A color image decomposition model for image enhancement
Tianqing Hu, Qinglei Zhou, Xiaofei Nan, Renhao Lin
Neurocomputing3
2023 Protein-Protein Interaction Site Prediction Based on Attention Mechanism and Convolutional Neural Networks
abstract
Proteins usually perform their cellular functions by interacting with other proteins. Accurate identification of protein-protein interaction sites (PPIs) from sequence is import for designing new drugs and developing novel therapeutics. A lot of computational models for PPIs prediction have been developed because experimental methods are slow and expensive. Most models employ a sliding window approach in which local neighbors are concatenated to present a target residue. However, those neighbors are not distinguished by pairwise information between a neighbor and the target. In this study, we propose a novel PPIs prediction model AttCNNPPISP, which combines attention mechanism and convolutional neural networks (CNNs). The attention mechanism dynamically captures the pairwise correlation of each neighbor-target pair within a sliding window, and therefore makes a better understanding of the local environment of target residue. And then, CNNs take the local representation as input to make prediction. Experiments are employed on several public benchmark datasets. Compared with the state-of-the-art models, AttCNNPPISP improves the prediction performance. Also, the experimental results demonstrate that the attention mechanism is effective in terms of constructing comprehensive context information of target residue.
Yuguang Li, Xiaofei Nan, Shoutao Zhang
IEEE ACM Trans. Comput. Biol. Bioinform.4
2022 Robustness evaluation for deep neural networks via mutation decision boundaries analysis
Renhao Lin, Qinglei Zhou, Xiaofei Nan
Inf. Sci.4
2022 Leveraging Sequential and Spatial Neighbors Information by Using CNNs Linked With GCNs for Paratope Prediction
abstract
Antibodies consisting of variable and constant regions, are a special type of proteins playing a vital role in immune system of the vertebrate. They have the remarkable ability to bind a large range of diverse antigens with extraordinary affinity and specificity. This malleability of binding makes antibodies an important class of biological drugs and biomarkers. In this article, we propose a method to identify which amino acid residues of an antibody directly interact with its associated antigen based on the features from sequence and structure. Our algorithm uses convolution neural networks (CNNs) linked with graph convolution networks (GCNs) to make use of information from both sequential and spatial neighbors to understand more about the local environment of target amino acid residue. Furthermore, we process the antigen partner of an antibody by employing an attention layer. Our method improves on the state-of-the-art methodology.
Yuguang Li, Fei Wang 0125, Xiaofei Nan, Shoutao Zhang
IEEE ACM Trans. Comput. Biol. Bioinform.4
2021 Attention-based Convolutional Neural Networks for Protein-Protein Interaction Site Prediction
abstract
Protein-protein interactions are of great importance in the life cycles of living cells. Accurate prediction of the proteinprotein interaction site (PPIs) from protein sequence improves our understanding of protein-protein interaction, contributes to the protein-protein docking and is crucial for drug design. However, practical experimental methods are costly and time-consuming so that many sequence-based computational methods have been developed. Most of those methods employ a sliding window approach, which utilize local neighbor information within a window size. However, they don’t distinguish and use the effect of each individual neighboring residue at different position. We propose a novel sequence-based deep learning method consisting of convolutional neural networks (CNNs) and attention mechanism to improve the performance of PPIs prediction. Our attention-based CNNs captures the different effect of each neighboring residue within a sliding window, and therefore making a better understanding of the local environment of target residue. We employ experiments on several public benchmark datasets. The experimental results demonstrate that our proposed method significantly outperforms the state-of-the-art techniques. The source code can be obtained from https://github.com/biolushuai/attention-based-CNNsfor-PPIs-prediction.
Yuguang Li, Xiaofei Nan, Shoutao Zhang
BIBM3
2021 A Sequence-Based Antibody Paratope Prediction Model Through Combing Local-Global Information and Partner Features
Yuguang Li, Xiaofei Nan, Shoutao Zhang
ISBRA3
2020 Predicting the results of molecular specific hybridization using boosted tree algorithm
abstract
Summary In the field of bioinformatics and DNA computing, simulated hybridization experiments can replace real molecular hybridization experiments to some extent, avoiding some disadvantages of the actual experimental design. However, the core techniques, which are employed by the popular DNA simulation software, are limited to the exponential computational complexity of the combinatorial problems. As a result, it is impossible to decide whether a specific hybridization among complex DNA molecules is effective or not within acceptable time. To address this common problem, we hereby introduce a new method based on the machine learning technique. First, a sample set is employed to train the boosted tree algorithm, which resulted in a corresponding machine learning model. Second, this model is applied to predict the classification results of molecular hybridization for a given group of DNA molecular coding. The experiment results showed that the new method had an average accuracy level of 94.2% and an average efficiency level 90 839 times higher than that of the existing representative approaches. Especially for the case study in this paper, the efficiency of the new method is 235 000, 250 000, and 990 000 times higher than that of the three existing methods, respectively. These experimental results indicate that our new approach can quickly and accurately determine the biological effectiveness of molecular hybridization for a given DNA design.
Weijun Zhu, Yingjie Han, Huanmei Wu, Yang Liu 0050, Xiaofei Nan, Qinglei Zhou
Concurr. Comput. Pract. Exp.5
2016 DTSP-V: A trend-based Top Scoring Pairs method for classification of time series gene expression data
abstract
Time series gene expression is a genetic data collected at different time points in the process of biological growth. As it contains a large amount of biological information related to a specific time period, the study of classification of time series gene data is a very significant task. With low sample size and high dimensionality, gene expression data classified by traditional machine learning methods suffers not only the curse of dimensionality but also lack of the interpretability of the complex model derived from the data. Inspired by the idea of top gene pair comparison, we propose the algorithm of Dynamic Top Scoring Pairs on Variance (DTSP-V), and introduce the concept of trend to regularize the variances between time stamps. In order to make full use of the advantages of the Top Scoring Pairs algorithm, we use the idea of trend to process time series gene. DTSP-V calculation can classify time series data and select one or more pairs of features with the most classification ability simultaneously. Experiments show that in comparison with traditional machine learning algorithms, DTSP-V algorithm, achieves high classification accuracy, and more importantly, provides explanatory computational model that valuable biological information could be extracted.
Kaimin Wu, Xiaofei Nan, Yumei Chai
BIBM2
2012 Implementation of multiple-instance learning in drug activity prediction
abstract
BACKGROUND: In the context of drug discovery and development, much effort has been exerted to determine which conformers of a given molecule are responsible for the observed biological activity. In this work we aimed to predict bioactive conformers using a variant of supervised learning, named multiple-instance learning. A single molecule, treated as a bag of conformers, is biologically active if and only if at least one of its conformers, treated as an instance, is responsible for the observed bioactivity; and a molecule is inactive if none of its conformers is responsible for the observed bioactivity. The implementation requires instance-based embedding, and joint feature selection and classification. The goal of the present project is to implement multiple-instance learning in drug activity prediction, and subsequently to identify the bioactive conformers for each molecule. METHODS: We encoded the 3-dimensional structures using pharmacophore fingerprints which are binary strings, and accomplished instance-based embedding using calculated dissimilarity distances. Four dissimilarity measures were employed and their performances were compared. 1-norm SVM was used for joint feature selection and classification. The approach was applied to four data sets, and the best proposed model for each data set was determined by using the dissimilarity measure yielding the smallest number of selected features. RESULTS: The predictive abilities of the proposed approach were compared with three classical predictive models without instance-based embedding. The proposed approach produced the best predictive models for one data set and second best predictive models for the rest of the data sets, based on the external validations. To validate the ability of the proposed approach to find bioactive conformers, 12 small molecules with co-crystallized structures were seeded in one data set. 10 out of 12 co-crystallized structures were indeed identified as significant conformers using the proposed approach. CONCLUSIONS: The proposed approach was proven not to suffer from overfitting and to be highly competitive with classical predictive models, so it is very powerful for drug activity prediction. The approach was also validated as a useful method for pursuit of bioactive conformers.
Xiaofei Nan, Haining Liu, Ronak Y. Patel, Pankaj R. Daga, Yixin Chen 0002, Dawn Wilkins, Robert J. Doerksen
BMC Bioinform.2
2012 Biomarker discovery using 1-norm regularization for multiclass earthworm microarray gene expression data
Xiaofei Nan, Nan Wang 0002, Ping Gong 0001, Yixin Chen 0002, Dawn Wilkins
Neurocomputing1
2011 Leveraging domain information to restructure biological prediction
abstract
BACKGROUND: It is commonly believed that including domain knowledge in a prediction model is desirable. However, representing and incorporating domain information in the learning process is, in general, a challenging problem. In this research, we consider domain information encoded by discrete or categorical attributes. A discrete or categorical attribute provides a natural partition of the problem domain, and hence divides the original problem into several non-overlapping sub-problems. In this sense, the domain information is useful if the partition simplifies the learning task. The goal of this research is to develop an algorithm to identify discrete or categorical attributes that maximally simplify the learning task. RESULTS: We consider restructuring a supervised learning problem via a partition of the problem space using a discrete or categorical attribute. A naive approach exhaustively searches all the possible restructured problems. It is computationally prohibitive when the number of discrete or categorical attributes is large. We propose a metric to rank attributes according to their potential to reduce the uncertainty of a classification task. It is quantified as a conditional entropy achieved using a set of optimal classifiers, each of which is built for a sub-problem defined by the attribute under consideration. To avoid high computational cost, we approximate the solution by the expected minimum conditional entropy with respect to random projections. This approach is tested on three artificial data sets, three cheminformatics data sets, and two leukemia gene expression data sets. Empirical results demonstrate that our method is capable of selecting a proper discrete or categorical attribute to simplify the problem, i.e., the performance of the classifier built for the restructured problem always beats that of the original problem. CONCLUSIONS: The proposed conditional entropy based metric is effective in identifying good partitions of a classification problem, hence enhancing the prediction performance.
Xiaofei Nan, Zhengdong Zhao, Ronak Y. Patel, Haining Liu, Pankaj R. Daga, Robert J. Doerksen, Xin Dang, Yixin Chen 0002, Dawn Wilkins
BMC Bioinform.1
2010 Gene selection using 1-norm regularization for multi-class microarray data
abstract
Explosive compounds such as TNT and RDX have various toxicological effects on the natural environment. The goal of the earthworm microarray experiment is to unearth the biomarker for toxicity evaluation. We propose a novel recursive gene selection method which can handle the multi-class setting effectively and efficiently. The selection is performed iteratively. In each iteration, a linear multi-class classifier is trained using 1-norm regularization, which leads to sparse weight vectors, i.e., many feature weights are exactly zero. Those zero-weight features are eliminated in the next iteration. The empirical results demonstrate that the selected features (genes) have very competitive discriminative power. In addition, the selection process has fast rate of convergence.
Xiaofei Nan, Nan Wang 0002, Ping Gong 0001, Yixin Chen 0002, Dawn Wilkins
BIBM1
2009 Clustering of Defect Reports Using Graph Partitioning Algorithms
Vasile Rus, Xiaofei Nan, Sajjan G. Shiva
SEKE2