EDBT 2026 Demo / reviewers in the wild / expert
Xuefeng Cui
dblp:96/3307
· DBLP profile ↗
48ranked-venue papers
7as first author
35since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 42 · 6 first-author · 33 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Longest common subsequence including at most k segments
Haitao Jiang 0005, Xuefeng Cui, Haodi Feng, Daming Zhu, Lusheng Wang 0001 |
Inf. Process. Lett. | 3 |
| 2026 | Dual-Axis Message Passing on Molecular Graphs: Layerwise Multi-hop Convolution Meets Depthwise Dense Mixing
Xuefeng Cui, Shiwei Sun |
ISBRA (1) | 1 |
| 2026 | HCPDA: Identification of PiRNA-Disease Association Based on Heterogeneous Graph Attention and Contrastive Learning
Yulian Ding, Xuefeng Cui |
ISBRA (1) | 5 |
| 2025 | CoDiST: Combining Unimodal Contrastive Denoising and Crossmodal Disentanglement for Spatial TranscriptomicsabstractIntegrating spatial transcriptomics (ST) with histology enables precise delineation of spatial domains in complex tissues. However, effective multimodal integration faces two primary challenges: (1) High sparsity and frequent dropout events in transcriptomic data and artifacts in histological images together obscure true biological signals, resulting in intra-modal noise. (2) Excessive emphasis on modality alignment in current fusion methods allows redundant information to dominate over modality -specific features, leading to inter-modal redundancy. To address these challenges, we propose CoDiST, a novel multi-modal representation learning framework. Specifically, CoDiST uses histological information to strengthen spatial neighborhood relationships and employs unimodal contrastive learning to enhance robustness against technical noise in both modalities. Furthermore, to overcome inter-modal redundancy, CoDiST introduces the Gene Image Disentanglement Network (GIDNet). It disentangles representations into a shared subspace and specific subspaces, achieving crossmodal semantic alignment while preserving complementary features unique to each modality. In benchmarking on multiple ST datasets, CoDiST delivers clear gains over existing methods. On the Human Breast Cancer dataset, CoDiST achieves a 0.66 Adjusted Rand Index (ARI), representing a 3.6% improvement over state-of-the-art methods. Kai Hu 0002, Xuefeng Cui, Fa Zhang 0001 |
BIBM | 3 |
| 2025 | IMPDI: Integrating Multiple Dynamic Information for Protein Representation LearningabstractProtein Representation Learning (PRL) holds signif-icant value in various fields. However, existing methods primarily focus on static amino acid sequences or structures of proteins, paying less attention to their dynamic behaviors, which limits their ability to capture the intrinsic properties of proteins. In this paper, a novel multimodal protein representation learning method named IMPDI is proposed, which integrates amino acid sequences, structural information, and multiple dynamic information to construct a unified protein representation learning framework. To validate the effectiveness of the IMPDI model, we applied it to three typical protein-related downstream tasks: Protein-Protein Interaction (PPI) Prediction, Protein Secondary Structure Prediction, and Protein Thermostability Prediction. The results indicate that IMPDI outperforms baseline methods in various metrics, demonstrating strong generalizability and application potential. This study not only offers a new perspective on protein representation methods, but also provides robust technical support for related downstream tasks. Lingtao Su, Yanlong Gong, Zhenyu Cui, Gonglei Zhang, Xuefeng Cui |
BIBM | 6 |
| 2025 | ICM-Prog: An Interpretable Framework for Breast Cancer Prognosis Prediction by Integrating Clinical and Multi-Omics DataabstractExisting breast cancer prognosis prediction models often rely on single-type data such as clinical or mRNA expression data, neglecting the prognostic value of gene mutations linked to tumor aggressiveness. Multi-omics integration remains difficult due to the sparsity of mutation data and the high dimensionality of expression data. To address these limitations, we propose ICMProg, an interpretable framework for breast cancer prognosis prediction by integrating clinical and multi-omics data. ICMProg uses a Frequency-aware Sparse Variational Autoencoder (FS-VAE) to extract informative features from sparse mutation data, and a Dynamic Pathway-Guided Variational Autoencoder (DPG-VAE) to capture biologically meaningful features from high-dimensional expression data. For model interpretability, we develop an interpretable method based on Grad-SHAP, systematically identifying critical biomarkers in clinical, mutational, gene, and pathway data. Experimental results on two breast cancer datasets show that ICM-Prog outperforms state-of-the-art methods, and it can significantly separate high-risk and low-risk patient groups. These results suggest that ICM-Prog can not only improve prognostic accuracy, but also support the discovery of potential therapeutic targets. Lingtao Su, Gonglei Zhang, Zhenyu Cui, Yanlong Gong, Xuefeng Cui |
BIBM | 6 |
| 2025 | Enough Consecutive Matches in k-Tuple Common Substrings
Siqi Jiang, Haitao Jiang 0005, Lianrong Pu, Haodi Feng, Xuefeng Cui, Li-Zhen Cui 0001, Daming Zhu |
ICIC (26) | 6 |
| 2025 | MambaST: Hexagonal State Space Modeling for Spatial Domain Identification
Kai Hu 0002, Xuefeng Cui, Fa Zhang 0001 |
ISBRA (1) | 3 |
| 2025 | MetaGIN: a lightweight framework for molecular property prediction
Haitao Jiang 0005, Xuefeng Cui |
Frontiers Comput. Sci. | 6 |
| 2025 | Central Feature Network Enables Accurate Detection of Both Small and Large Particles in Cryo-Electron Tomography
Yaoyu Wang, Fa Zhang 0001, Xuefeng Cui |
J. Comput. Sci. Technol. | 5 |
| 2024 | Revolutionizing Enzyme Turnover Predictions with Non-Binary Reaction Fingerprints and AIabstractThe turnover number (kcat) is a crucial measure for evaluating enzyme catalytic efficiency, significantly influencing studies of cellular metabolism and resource allocation. Due to the high costs associated with determining kcat, predictive machine learning models such as TurNuP have been developed to predict the turnover numbers of kinetically uncharacterized enzymes. Existing methods rely on binary Differential Reaction Fingerprints (DRFPs) to represent reactions. However, these binary reaction fingerprints have two main limitations. First, they only show changes between substrates and products, without specifying whether the change occurs in the substrates or products. Second, they omit information about similarities between substrates and products, missing crucial reaction details. To overcome these issues, we developed Quaternary Reaction Fingerprints (QRFPs) by incorporating quaternary structural features. Thus, QRFPs extend DRFPs, and DRFPs can be derived from QRFPs. We proposed a deep learning method based on QRFPs and pretrained protein language model ESM2, called TurNuP4. Experimental results show that TurNuP4outperforms TurNuP by 18% in R2to predict kcatvalues for natural reactions of wild-type enzymes. Code is available at https://github.com/xfcui/TurNuP4. Jibin Cheng, Guishan Cui, Haitao Jiang 0005, Xuefeng Cui |
BIBM | 5 |
| 2024 | Enhanced Protein-Ligand Affinity Prediction with Conditional Updating and Proximity EmbeddingabstractThe protein-ligand affinity prediction task aims to predict the binding strength of small molecule ligands to specific proteins, which is crucial in the fields of drug design and molecular biology, and can accelerate the drug discovery. The structure complementarity between protein and ligand plays a critical role in determining binding strength , but most of current deep learning-based affinity prediction models usually extracted the features of protein and ligand by these two detached modules. which limits the exchange of information for capturing interactions and struggles to capture proteins’ important residues. To address these limitations. we introduce CIP. which takes the combination of GNN, Conditional Updating and Proximity Embedding for the first time. Compared to existing models, CIP has several significant advantages. First, Conditional updating modifies the ligand’s local features based on the protein’s global features, and vice versa , enhancing structural complementarity to capture intricate interactions. Second, encoding the relative distances between proximal residue-atom pairs highlights critical residues. Additionally, our model integrates covalent and noncovalent interactions to obtain more comprehensive graph representations. Experiments on the PDB-bind 2016 benchmark demonstrate that CIP outperforms the original method with improvements of 2.3%, 3.2%, 2.8%, 3.8%, and 3.2% across five baselines. Furthermore, visualization results reveal that CIP effectively captures intricate interactions and crucial residues.The implemented code and dataset are available online at https://github.com/xfcui/CIP. Shuai Cui, Zizheng Nie, Haitao Jiang 0005, Guishan Cui, Xuefeng Cui |
BIBM | 6 |
| 2024 | Advancing Template-Based Flexible Docking of P450-Heme Complexes via JAX MDabstractAlphaFold2 has advanced protein structure prediction but often overlooks complexes with essential ligands, crucial for understanding protein-ligand interactions relevant to drug design. AlphaFill was developed to integrate ligands and ions from experimental structures into predicted models. However, many resulting complexes lack precise spatial constraints, leading to clashes between filled heme atoms and the P450 structure, as well as discrepancies in the key Fe-S bond.So, we propose FlexFill, a rapid flexible docking algorithm based on JAX MD. Specifically, FlexFill takes the predicted P450 structure as input to generate the P450-heme complex structure. During the docking process, we leverage JAX MD to account for the structural flexibility of P450. Unlike traditional molecular dynamics methods, we developed a customized energy function that focuses on the atoms near the pocket, significantly reducing runtime while satisfying spatial constraints.The results indicate that AlphaFill experienced almost 50% structural crashes, while FlexFill exhibited a substantially lower rate at less than 5%. Furthermore, FlexFill is 25% faster than AlphaFill, with further speed improvements by removing redundant entries from the template library. These observations suggest that FlexFill is capable of generating high-fidelity P450heme complexes and has the ability to flexibly dock a diverse array of protein-ligand complexes. Code and pre-trained model are released at https://github.com/xfcui/FlexFill. Rongxu Guo, Jiaheng Sang, Guishan Cui, Fa Zhang 0001, Xuefeng Cui, Shengying Li |
BIBM | 6 |
| 2024 | Iterative Glycan Structure Generation by Component Localization and Classification from Mass SpectraabstractAnalyzing intact glycopeptides via mass spectrometry data elucidates the glycan composition and attachment sites, which are pivotal for investigating protein folding, stability, and function. GlycanFinder, an advanced technique, builds the glycan structure starting from the peptide (i.e., the root) and sequentially adds monosaccharide components (i.e., the leaves). However, its major limitation is the dependence on predefined rules with strong assumptions for leaf addition locations. To address this issue, we introduce MS2Glycan, a deep learning method for iteratively generating novel glycan structures from mass spectrometry data. It models component generation similarly to an object detection task, incorporating both localization and classification sub-tasks, and constructs the glycan structure component by component. MS2Glycan offers a key improvement over GlycanFinder by using a data-driven approach to predict component locations, rather than relying on predefined rules. By analyzing intact N-linked glycopeptides collected from five different mouse tissues, we assessed MS2Glycan through five-fold cross-validation, observing a notable enhancement in structural accuracy, which rose from 32% to 74%. Additionally, our findings indicate that our model more effectively mimics the principles of glycan structure construction and offers valuable insights into fully leveraging mass spectrometry data. The Python implementation can be accessed via the following link: https://github.com/xfcui/MS2Glycan. Defeng Li, Zizheng Nie, Yanmin Liu, Xiaojun Cai, Xuefeng Cui |
BIBM | 6 |
| 2024 | Deep Representation Learning for Electron Ionization Mass Spectra RetrievalabstractThe task of retrieving and analyzing mass spectra is indispensable for the identification of compounds in mass spectrometry (MS). This methodology is of critical importance as it enables researchers to correlate observed spectra with established databases, thereby precisely determining the chemical composition of samples. The primary challenges to its efficacy lie in optimizing the balance between retrieval accuracy and processing speed. Empirical studies have demonstrated that by converting mass spectra into embeddings via deep learning, it is possible to achieve both high accuracy and speed in retrieval. Nevertheless, there are complex challenges associated with employing deep learning for spectral embedding, particularly within the domain of electron ionization mass spectrometry (EI-MS). In this paper, we introduce a novel representation learning technique termed EI-MS2VEC for EI-MS retrieval. Our spectrum retrieval methodology surpasses current state-of-the-art techniques such as FastEI. For the in-silico library, we attain hit rate@1 and hit rate@10 of 43.6% and 84.5%, respectively, compared to FastEI’s 36.7% and 80.4%. Moreover, our retrieval approach operates with an order of magnitude greater speed than FastEI. The source code is available on Github (https://github.com/xfcui/EI-MS2VEC). Anlei Jiao, Shiwei Sun, Longyang Dian, Xuefeng Cui |
BIBM | 6 |
| 2024 | How to Integrate Non-Sequence Features for Peptide Collision Cross Section Prediction with Deep Learning?abstractThe collision cross section (CCS) is a crucial parameter that reflects a molecule’s size and shape, offering valuable insights into its structural conformation. Accurate CCS predictions can significantly enhance the identification and separation of compounds in mass spectrometry, proving particularly useful in fields such as proteomics and drug discovery. The CCS of a peptide is a reproducible physicochemical property influenced by its mass, charge, and gas-phase structure. Despite its importance, predicting CCS using both sequence data and physical attributes remains challenging. To address this, we introduce PEP2CCS, a deep learning model designed to predict CCS. This model integrates enhanced physical features of peptides, thereby improving the accuracy of CCS predictions. The analysis revealed that PEP2CCS had a Median Absolute Percent Error of 1.139% in predicting the synthetic proteomic peptide cross-section, which is 18.6% lower than state-of-the-art method. Furthermore, the results suggest that the integration of the charge state into the model is a significant contributor to the accurate prediction of CCS values. Code and pre-trained model are released at https://github.com/xfcui/PEP2CCS. Zhimeng Tian, Zizheng Nie, Daming Zhu, Xuefeng Cui |
BIBM | 5 |
| 2024 | How to Train Your Neural Network for Molecular Structure Generation from Mass Spectra?abstractMass spectrometry serves as a pivotal tool for the analysis of small molecules through an examination of their mass-to-charge ratios. Recent advancements in deep learning have markedly enhanced the analysis of mass spectrometric data, facilitating the prediction of novel small molecule structures without the necessity of extensive databases. Nonetheless, the paucity of annotated datasets impedes the efficacious training of molecular generation models predicated on MS2spectra. To mitigate this limitation, we introduce ctMSNovelist, an avant-garde method that amalgamates pre-training, fine-tuning, and co-training techniques to construct a more precise model for the generation of molecular structures from tandem mass spectrometry data. This novel approach augments both the training regimen and the predictive accuracy of the MSNovelist model, thereby surmounting the obstacle of limited data. The methodology commences with the pretraining of a Variational Autoencoder (VAE) to generate molecular fingerprints derived from SMILES strings. Subsequently, it undergoes fine-tuning to emulate noisy fingerprints originating from mass spectrometry (MS) data. Concurrently, MSNovelist is co-trained utilizing these simulated fingerprints, inclusive of the highly noisy variants produced in the early stages of VAE training. The incorporation of a substantial volume of noisy data serves to enhance model accuracy and avert overfitting. We evaluated ctMSNovelist using the GNPS dataset and attained a SMILES prediction accuracy of 48.8%, representing a 4.1% enhancement over MSNovelist. It is pertinent to note that the sole distinction between ctMSNovelist and MSNovelist in this experiment was the training process. The code and models are publicly available at https://github.com/xfcui/ctMSNovelist. Yanmin Liu, Longyang Dian, Shiwei Sun, Xuefeng Cui |
BIBM | 5 |
| 2024 | DTMIReID: Person Re-identification Based on Deformable Transformer to Incorporate Mutual Information Between Images
Haodi Feng, Xuefeng Cui |
ICPR (14) | 3 |
| 2024 | Central Feature Network Enables Accurate Detection of Both Small and Large Particles in Cryo-Electron Tomography
Yaoyu Wang, Fa Zhang 0001, Xuefeng Cui |
ISBRA (1) | 5 |
| 2023 | Using Deep Learning to Classify Full-Length Transcriptome SequencesabstractThe emergence of Third-Generation Sequencing (TGS) has revolutionized transcriptome sequencing, allowing the production of long reads that span multiple kilobases. This breakthrough has enabled the sequencing of entire transcript sequences. However, a major challenge is posed by the high error rates associated with TGS, making it difficult to accurately classify transcript sequences against reference sequences using traditional algorithms. Fortunately, deep learning-based embedding methods can be trained to overcome these errors. In this pioneering study, we introduce trxCNN, a deep learning model that exhibits remarkable accuracy in classifying erroneous transcript sequences compared to reference sequences. Specifically, evaluations of simulated data have revealed that trxCNN has an impressive classification accuracy of 87.1%. This accuracy exceeds that of the Minimap2 and magicBLAST aligners, both designed for TGS data, by 10.7% and 9.0%, respectively. Furthermore, we provide evidence that trxCNN is capable of accurately estimating the abundance of transcripts. These findings strongly suggest that deep learning methods have great potential to effectively process errors-affected sequencing data. Junchi Ma, Cuiyuan Li, Xuefeng Cui |
BIBM | 5 |
| 2023 | Variational Clustering and Denoising of Spatial TranscriptomicsabstractSpatial transcriptomics data provides a unique opportunity to investigate both gene expression and spatial structure in tissues at the same time. However, incorporating spatial information to accurately identify spatial domains is difficult due to factors such as high-dimensionality, sparsity, noise, and dropout events. To address these issues, we introduce vGraphST, a novel graph-based deep learning approach tailored for spatial transcriptomics data. Our method combines auto-encoder and contrastive learning techniques to process high-dimensional data and generate meaningful low-dimensional embeddings. Additionally, we use continuous distributions instead of discrete values in both the latent space and the denoised gene expression space. Specifically, Gaussian distributions are used to model the latent space, while zero-inflated Poisson distributions are used to model the denoised gene expression space. Experimental results demonstrate the effectiveness of vGraphST in accurately representing and analyzing spatial transcriptomics data. When compared to other methods using the DLPFC dataset, vGraphST achieves an average Adjusted Rand Index (ARI) of 0.58, demonstrating its superiority in segmenting spatial domains and recognizing biologically relevant spatiotemporal patterns. Cuiyuan Li, Fa Zhang 0001, Kai Hu 0002, Xuefeng Cui |
BIBM | 4 |
| 2023 | De Novo Molecular Structure Generation from Mass SpectraabstractMass spectrometry is a key technology for the identification of small molecules. However, traditional methods that rely on database comparisons have difficulty with newly discovered molecules that are not in the database. Recent advances in deep learning allow for direct analysis of mass spectra, which makes it possible to predict chemical structures without using a database. We have found that the accurate prediction of hydrogen atoms is a major challenge for the prediction of chemical structures, especially since they are not explicitly represented in SMILES. To address this challenge, we introduce MS2SMILES, a novel approach that treats hydrogen atoms as implicitly linked to heavy atoms. This method enables the model to predict both heavy atoms and hydrogen atoms accurately (instead of just focusing on heavy atoms) during the training phase. Additionally, MS2SMILES incorporates the SMILES grammatical rules when predicting chemical structures, increasing the reliability of the generated SMILES representations. We tested MS2SMILES using the GNPS and CASMI 2016 datasets, and it achieved SMILES prediction accuracies of 53.6% and 63.8%, respectively. These results demonstrate a significant improvement of 19.9% and 10.9% compared to the current leading method. Yanmin Liu, Daming Zhu, Xuefeng Cui |
BIBM | 5 |
| 2023 | Learned Fingerprint Embedding for Large-Scale Peptide Mass Spectra RetrievalabstractTandem mass spectrometry (MS/MS) is a widely used technique for protein identification, post-translational modifications, immunotherapy, and other applications. As the amount of MS/MS spectra data increases, new computational methods are needed to efficiently search through these databases. This study introduces MS2VEC, a novel fingerprint embedding model designed to facilitate large-scale retrieval of peptide mass spectra. MS2VEC captures the relationships between distant peaks and incorporates position-aware fingerprint features from all peaks. To do this, dilated convolutions are used to capture remote relationships, and a novel position-aware multi-head attention pooling mechanism is used to abstract fingerprint features. The results demonstrate that MS2VEC achieves a top-1 retrieval accuracy of 0.810, outperforming existing methods by 5.1%. Interestingly, the precursor charge is not essential for the retrieval task, as the spectra itself contains enough information to accurately predict the charge. Additionally, the results suggest that weight-balanced fragment ions and water losses are important contributors to fingerprint features. Yongshuai Wang, Xiaojun Cai, Defeng Li, Shiwei Sun, Xuefeng Cui |
BIBM | 6 |
| 2023 | ContactLib-ATT: A Structure-Based Search Engine for Homologous ProteinsabstractGeneral-purpose protein structure embedding can be used for many important protein biology tasks, such as protein design, drug design and binding affinity prediction. Recent researches have shown that attention-based encoder layers are more suitable to learn high-level features. Based on this key observation, we propose a two-level general-purpose protein structure embedding neural network, called ContactLib-ATT. On local embedding level, a biologically more meaningful contact context is introduced. On global embedding level, attention-based encoder layers are employed for better global representation learning. Our general-purpose protein structure embedding framework is trained and tested on the SCOP40 2.07 dataset. As a result, ContactLib-ATT achieves a SCOP superfamily classification accuracy of 82.4% (i.e., 6.7% higher than state-of-the-art method). On the same dataset, ContactLib-ATT is used to simulate a structure-based search engine for remote homologous proteins, and our top-10 candidate list contains at least one remote homolog with a probability of 91.9%. Cheng Chen 0031, Yuguo Zha, Daming Zhu, Kang Ning 0001, Xuefeng Cui |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | Hexagonal Convolutional Neural Network for Spatial Transcriptomics ClassificationabstractRecent advances in spatial transcriptomics have enabled the comprehensive measurement of transcriptional profiles while retaining the spatial contextual information. Identifying spatial domains is a critical step in the analysis of spatially resolved transcriptomics. Existing unsupervised methods perform poorly on this task owing to the large amount of noise and dropout events in the transcriptomic profiles. To address this problem, we first extend an unsupervised algorithm to a supervised learning method that can identify useful features and reduce noise hindrance. Second, inspired by the classical convolution in convolutional neural networks (CNNs), we designed a regular hexagonal convolution to compensate for the missing gene expression patterns from adjacent nodes. Compared with the graph convolution in graph neural networks (GNNs), our hexagonal convolution can preserve the relative spatial location information of different nodes in graph-structured data. Third, based on the hexagonal convolution, a novel hexagonal Convolutional Neural Network (hexCNN) is proposed for spatial transcriptomics classification. Finally, we compared the proposed hexCNN with existing methods on the DLPFC dataset. The results show that hexCNN achieves a classification accuracy of 87.2% and an average Rand index (ARI) of 78.2% (1.9% and 3.3% higher than those of GNNs). Fa Zhang 0001, Kai Hu 0002, Xuefeng Cui |
BIBM | 4 |
| 2022 | Longest k-tuple Common Sub-StringsabstractWe focus on a new problem that is formulated to find a longest k-tuple of common sub-strings (abbr. k-CSSs) of two or more strings. We present a suffix tree based algorithm for this problem, which can find a longest k-CSS of m strings in $O(kmn^{k})$ time and $O(kmn)$ space where n is the length sum of the m strings. This algorithm can be used to approximate the longest k-CSS problem to a performance ratio $\frac{1}{\epsilon}$ in $O(kmn^{\lceil\epsilon k\rceil})$ time for $\epsilon\in(0,1]$. Since the algorithm has the space complexity in linear order of n, it will show advantage in comparing particularly long strings. This algorithm proves that the problem that asks to find a longest gapped pattern of non-constant number of strings is polynomial time solvable if the gap number is restricted constant, although the problem without any restriction on the gap number was proved NP-Hard. Using a C++ tool that is reliant on the algorithm, we performed experiments of finding longest 2-CSSs, 3-CSSs and 5-CSSs of 2 ~ 14 COVID-19 S-proteins. Under the help of longest 2-CSSs and 3-CSSs of COVID-19 S-proteins, we identified the mutation sites in the S-proteins of two COVID-19 variants Delta and Omicron. The algorithm based tool is available for downloading at https://github.com/lytt0/k-CSS. Daming Zhu, Haitao Jiang 0005, Haodi Feng, Xuefeng Cui |
BIBM | 5 |
| 2022 | Pretraining Transformers for TCR-pMHC Binding PredictionabstractThe knowledge concerning antigen presentation by the main histocompatibility complex (MHC) to T-cell receptor (TCR) and TCR binding specificity can facilitate the application of T-cell immunity in modern medicine, such as tumor immunotherapy and drug and vaccine design cases. With the development of high-throughput sequencing technology and artificial intelligence, data-driven approaches can be employed to help understand the rules of TCR-pMHC binding. Simulating the biological binding process of TCRs and pMHCs, we propose a novel pipeline, pMTattn, using transfer learning based on an attention mechanism for TCR-pMHC binding prediction. During the pretraining stage, partner-specific training strategies can capture useful local binding features. In the fine-tuning stage, an attention block is employed to aggregate the TCR encoding and pMHC encoding information, forming a better global TCR-pMHC representation. Visualization experiments indicate that the pMTattn model focuses more on the voxels near the binding sites of pMHCs and TCRs. This key observation effectively supports our hypothesis that attention is critical for TCR-pMHC binding prediction. In addition, on an independent test set, the area under the precision-recall curve (AUPR) and the area under the receiver operating characteristic curve (AUC) are improved from 0.533 to 0.583 and from 0.830 to 0.866, respectively, by pMTattn compared to those of the state-of-the-art model. Simultaneously, we also explore the influences of different sequence lengths and dataset differences on the model effect, and pMTattn exhibits better robustness than other models. These results suggest that pMTattn has the ability to be used as an adjunct tool for screening and discovering neoantigens. Jinsheng Shang, Qihong Jiao, Daming Zhu, Xuefeng Cui |
BIBM | 5 |
| 2022 | Boosting Deep Learning-based Docking with Cross-attention and Centrality EmbeddingabstractDocking is a classic computational biology problem that is widely used to predict binding conformations and to virtually screen binding molecules. Recently, a deep learning-based method, DeepDock was proposed to address the docking problem. The method shows great performance on conformation prediction. One major limitation of the method is that the interaction of ligands and targets is too simple. Here, we introduce caDeepDock, which is a geometric deep learning model for protein-ligand binding conformation prediction. Inspired by DeepDock, caDeepDock has two major advantages over DeepDock. First, cross-attention is employed to enable communications between the molecule and the protein binding pocket. Second, a positional embedding based on node degrees is used to incorporate both data-dependent and position-dependent communications. Experiments on the CASF-2016 benchmark have shown that the potential function learned by caDeepDock is able to pick a near-native conformation (with RMSD$\leqslant$2Å) from a set of decoy conformations with a success rate of 91.6%. This result outperforms not only DeepDock by 4.6% but also AutoDock by 1.4%. Consequently, the potential function learned by caDeepDock yields 5.96% more near-native conformations in conformation optimization experiments. Zongzhao Qiu, Zhenghe Yang, Xuefeng Cui |
BIBM | 6 |
| 2022 | Graph-based Reaction Classification by Contrasting between Precursors and ProductsabstractThe classification of organic reactions is a complex and tedious process, and it often requires certain domain knowledge to understand the classification rules. To reduce such requirements on domain knowledge, BERT-based deep learning methods have been applied to classify reactions based on SMILES strings. However, the same reaction can be represented by different but equivalent SMILES strings, and it can be observed that BERT-based methods are highly sensitive to the choice of SMILES strings. Logically, GNN-based methods are robust to equivalent SMILES strings. Here, we propose a graph isomorphism network (GIN) with contrastive learning, called ContraGIN, to encourage feature fusion learning between precursors and products of the reaction. Indeed, experiments have shown that the new method focuses on learning the atoms near the bond changes before and after the reaction, and consequently achieves a classification accuracy of 99.30% on the USPTO 1k TPL dataset. Additional experiments have also shown that ContraGIN is more robust and much faster for lager reactions (with more atoms) and complicated reactions (with rings). Yuxiao Wang 0002, Zhaoxu Meng, Zhenghe Yang, Xuefeng Cui |
BIBM | 7 |
| 2022 | CoAtGIN: Marrying Convolution and Attention for Graph-based Molecule Property PredictionabstractMolecule property prediction based on computational strategies plays a key role in the process of drug discovery and design processes, such as DFT. However, these traditional methods are time-consuming, labor-intensive, and cannot satisfy the need for biomedicine. Owning to the development of deep learning, there are many variants of Graph Neural Networks (GNN) for molecular representation learning. However, the existing well-performing graph-based methods that have a number of parameters or light models cannot achieve good grades on various tasks. To manage the trade-off between efficiency and performance, we propose a novel model architecture, CoAtGIN, using both Convolution and Attention. At the local level, k-hop convolution is designed to capture long-range neighbor information. At the global level, in addition to using the virtual node to pass identical messages, we utilize linear attention to the aggregate global graph representation according to the importance of each node and edge. In the recent Open Graph Benchmark (OGB) Large-Scale Benchmark, CoAtGIN achieves the 0.0901 Mean Absolute Error (MAE) on the large-scale dataset PCQM4Mv2 with only 6.4 M model parameters. Moreover, using the linear attention block improves the performance, which helps to capture the global representation. Zhaoxu Meng, Zhenghe Yang, Haitao Jiang 0005, Xuefeng Cui |
BIBM | 6 |
| 2021 | Jointly Learning to Align and Aggregate with Cross Attention Pooling for Peptide-MHC Class I Binding PredictionabstractPredicting binding affinities of peptide antigens presented on major histocompatibility complex (MHC) is of great importance in T-cell immune response research. Accurate prediction of peptide-MHC binding affinities is essential for vaccine design and disease treatment. Recent deep learning-based prediction methods have shown that effective sequence embedding is critical to accurately predict binding affinities. One common neural network layer shared by these methods is the global average pooling layer that aggregates features. However, can we design a better global pooling layer? Here, we introduce a novel cross attention pooling (caPool) layer to aggregate features. As our initial application of caPool, a novel end-to-end transformer model, called capTransformer, is proposed for peptide-MHC class I binding prediction. In our model, caPool jointly aligns peptide-MHC residual pairs and aggregates residual features. Thus, instead of treating all residues equally and independently, caPool focuses more on correlated residue pairs that are potentially contact pairs contributing major forces to stabilize the complex structure. Using a five-fold cross-validation experiment, we found that caPool achieved the highest PCC value of 0.845, which was 0.139 higher than a global average pooling. Here, the global pooling layer was the only difference between the two tested models, and this observation indicated that global average pooling was not always the best choice. Importantly, our capTransformer model achieved a SRCC value of 0.614 (i.e., 6.4% higher than the best-performing method) when applied to the IEDB dataset. Cheng Chen 0051, Zongzhao Qiu, Zhenghe Yang, Bin Yu 0007, Xuefeng Cui |
BIBM | 5 |
| 2021 | Hydrogen bonds meet self-attention: all you need for protein structure embeddingabstractGeneral-purpose protein structure embedding can be used for many important protein biology tasks, such as protein design, drug design and binding affinity prediction. Recent researches have shown that attention-based encoder layers are more suitable to learn high-level features. Based on this key observation, we treat low-level representation learning and high-level representation learning separately, and propose a two-level general-purpose protein structure embedding neural network, called ContactLib-ATT. On the local embedding level, a simple yet meaningful hydrogen-bond representation is learned. On the global embedding level, attention-based encoder layers are employed for global representation learning. In our experiments, ContactLib-ATT achieves a SCOP superfamily classification accuracy of 82.4% (i.e., 6.7% higher than state-of-the-art method) on the SCOP 40 2.07 dataset. Moreover, ContactLib-ATT is demonstrated to successfully simulate a structure-based search engine for remote homologous proteins, and our top-10 candidate list contains at least one remote homolog with a probability of 91.9%. Cheng Chen 0031, Yuguo Zha, Daming Zhu, Kang Ning 0001, Xuefeng Cui |
BIBM | 5 |
| 2021 | Edge-Gated Graph Neural Network for Predicting Protein-Ligand Binding AffinitiesabstractPredicting Protein-ligand binding affinities using Deep Learning can significantly shorten the drug development cycle. Recently, Graph neural network models have been developed, and are successfully used to accelerate the development of potential drugs. One major limitation of these GNN models is that they focus on node features (i.e., atom features), as these nodes carry the most important information of molecules. However, atoms are connected via different bonds in molecules, and we argue that such chemical bonds carry critical information for assessing how atomic features should be aggregated. To overcome the lack of bond-related information in earlier models, we here proposed a novel edge-gated graph neural network (egGNN) that predict the binding affinities between proteins and ligands. Specifically, our model treats chemical bonds as gates that control how information is extracted between atoms, this modification enables our model to extract more accurate information for different bonds. We tested our model using the CASF-2016 dataset, and found that the Pearson’s correlation coefficient (R) of egGNN is three percent higher than that of the best tested method (i.e., 0.86 vs 0.83), and the Root Mean Square Error (RMSE) is significantly lower than that of the best tested method (i.e., 1.12 vs 1.23) when compared to state-of-the-art-approaches. In ablation experiments, we demonstrate that our edge-gated feature extraction (EGFE) consoderably improves the performance of GNNs. These results indicate that egGNN represents a promising tool applicable for virtual screening, and should greatly assist in accelerating drug development. Qihong Jiao, Zongzhao Qiu, Yuxiao Wang 0002, Zhenghe Yang, Xuefeng Cui |
BIBM | 6 |
| 2021 | rzMLP-DTA: gMLP network with ReZero for sequence-based drug-target affinity predictionabstractComputational algorithms are being successfully used to speed up drug development processes, primarily by way of turning biochemical problems into data problems. Recently, with increasing amounts of available biological data generated by biochemical methods measured affinity, some computational algorithms based deep learning for predicting drug-target affinity (DTA) become promising directions for accelerating the process of drug development. These deep learning models first attempts focus on representation learning of individual amino acids or atoms. Next global average pooling layers in those models are used to to combine such individual features to global features, and finally simple feed forward networks are adopted to yield affinity predictions. Notably, research which has been undertaken on global feature aggregations (e.g., the global pooling and the feed forward layers) for the drug-target affinity problem is still lacked currently. To address this issue, we propose a new rzMLP block featured newly designed global feature aggregations. This rzMLP block is based on two recent technologies in deep learning research: the gMLP model and the ReZero layer. We use gMLP model to aggregate input features with a constant size, while the ReZero layer is used to smooth the training process of this block. Our rzMLP is capable of learning complicated global features while overcoming the problems caused by the model being too deep. Importantly, when we compared a model contained rzMLP block to others, the mean squared error(MSE) decreases by 33%. Comparing to state-of-the-art methods for predicting affinity, rzMLP-DTA achieves the lowest MSE and highest CI on two benchmarks, Davis and KIBA datasets, respectively. Zongzhao Qiu, Qihong Jiao, Yuxiao Wang 0002, Daming Zhu, Xuefeng Cui |
BIBM | 6 |
| 2021 | Structure-Based Protein-Drug Affinity Prediction with Spatial Attention MechanismsabstractDiscovery of new drugs heavily relies on predicting the binding affinities of drug molecules to suitable drug targets. To accelerate this process of identifying accurate affinities, computer-aided methods need to be applied in drug discovery pipeline. While various computational methods have been developed in the past ten years, the most successful methods to date use 3D convolutional neural networks (3D-CNNs). These 3D-CNN networks are based on deep learning models, and are both faster and more accurate than machine learning methods. However, currently used CNN is difficult to learn global and spatial features, while we hypothesis that spacial features should be critical for structure-based binding affinity predictions. Here we propose an end-to-end 3D-CNN with spatial attention mechanisms, called saCNN, to encourage spatial feature learning. When visualizing the learned spacial attentions in our experiments, it can be observed that saCNN model focuses more on the voxels near interaction centers. This key observation well supports our hypothesis that spacial features are critical for binding affinity predictions. In additions, we show that our model improves the Root Mean Square Error (RMSE) of the predicted binding affinities by 11.5% (with an absolute value of 1.117) and the Pearson Correlation Coefficient (R) by 3.2% (with an absolute value of 0.865) compared to currently used models on the PDBbind v.2016 core set. Importantly, the generalization abilities of our model is further demonstrated on CASF-2013 and CASF-2007 datasets. Yuxiao Wang 0002, Zongzhao Qiu, Qihong Jiao, Zhaoxu Meng, Xuefeng Cui |
BIBM | 6 |
| 2020 | Dilated-DenseNet for Macromolecule Classification in Cryo-electron Tomography
Renmin Han, Xuefeng Cui, Zhiyong Liu 0002, Min Xu 0009, Fa Zhang 0001 |
ISBRA | 4 |
| 2020 | SVLR: Genome Structure Variant Detection Using Long Read Sequencing Data
Wenyan Gu, Aizhong Zhou, Lusheng Wang 0001, Shiwei Sun, Xuefeng Cui, Daming Zhu |
ISBRA | 5 |
| 2020 | SHREC 2020: Classification in cryo-electron tomograms
Ilja Gubins, Marten L. Chaillet, Gijs van der Schot, Remco C. Veltkamp, Friedrich Förster, Xiaohua Wan 0001, Xuefeng Cui, Fa Zhang 0001, Emmanuel Moebel, Xiao Wang 0004, Daisuke Kihara, Min Xu 0009, Nguyen P. Nguyen, Tommi A. White, Filiz Bunyak |
Comput. Graph. | 8 |
| 2019 | When sparse coding meets ranking: a joint framework for learning sparse codes and ranking scores
Jim Jing-Yan Wang, Xuefeng Cui, Xin Gao 0001 |
Neural Comput. Appl. | 2 |
| 2018 | Learning Protein Structural Fingerprints under the Label-Free Supervision of Domain Knowledge
Yaosen Min, Chenyao Lou, Xuefeng Cui |
BIBM | 4 |
| 2018 | Mmalloc: A Dynamic Memory Management on Many-core Coprocessor for the Acceleration of Storage-intensive Bioinformatics Application
Mingzhe Zhang 0005, Jingrong Zhang, Rui Yan 0009, Zhiyong Liu 0002, Fa Zhang 0001, Xuefeng Cui |
BIBM | 8 |
| 2017 | K-nearest uphill clustering in the protein structure space
Xuefeng Cui, Xin Gao 0001 |
Neurocomputing | 1 |
| 2017 | Multi-instance dictionary learning via multivariate performance measure optimization
Jim Jing-Yan Wang, Ivor W. Tsang, Xuefeng Cui, Zhiwu Lu 0001, Xin Gao 0001 |
Pattern Recognit. | 3 |
| 2016 | CMsearch: simultaneous exploration of protein sequence space and structure space improves not only protein homology detection but also protein structure predictionabstractMOTIVATION: Protein homology detection, a fundamental problem in computational biology, is an indispensable step toward predicting protein structures and understanding protein functions. Despite the advances in recent decades on sequence alignment, threading and alignment-free methods, protein homology detection remains a challenging open problem. Recently, network methods that try to find transitive paths in the protein structure space demonstrate the importance of incorporating network information of the structure space. Yet, current methods merge the sequence space and the structure space into a single space, and thus introduce inconsistency in combining different sources of information. METHOD: We present a novel network-based protein homology detection method, CMsearch, based on cross-modal learning. Instead of exploring a single network built from the mixture of sequence and structure space information, CMsearch builds two separate networks to represent the sequence space and the structure space. It then learns sequence-structure correlation by simultaneously taking sequence information, structure information, sequence space information and structure space information into consideration. RESULTS: We tested CMsearch on two challenging tasks, protein homology detection and protein structure prediction, by querying all 8332 PDB40 proteins. Our results demonstrate that CMsearch is insensitive to the similarity metrics used to define the sequence and the structure spaces. By using HMM-HMM alignment as the sequence similarity metric, CMsearch clearly outperforms state-of-the-art homology detection methods and the CASP-winning template-based protein structure prediction methods. AVAILABILITY AND IMPLEMENTATION: Our program is freely available for download from http://sfb.kaust.edu.sa/Pages/Software.aspx CONTACT: : [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xuefeng Cui, Zhiwu Lu 0001, Sheng Wang 0001, Jim Jing-Yan Wang, Xin Gao 0001 |
Bioinform. | 1 |
| 2015 | Finding optimal interaction interface alignments between biological complexesabstractMOTIVATION: Biological molecules perform their functions through interactions with other molecules. Structure alignment of interaction interfaces between biological complexes is an indispensable step in detecting their structural similarities, which are key S: to understanding their evolutionary histories and functions. Although various structure alignment methods have been developed to successfully access the similarities of protein structures or certain types of interaction interfaces, existing alignment tools cannot directly align arbitrary types of interfaces formed by protein, DNA or RNA molecules. Specifically, they require a ': blackbox preprocessing ': to standardize interface types and chain identifiers. Yet their performance is limited and sometimes unsatisfactory. RESULTS: Here we introduce a novel method, PROSTA-inter, that automatically determines and aligns interaction interfaces between two arbitrary types of complex structures. Our method uses sequentially remote fragments to search for the optimal superimposition. The optimal residue matching problem is then formulated as a maximum weighted bipartite matching problem to detect the optimal sequence order-independent alignment. Benchmark evaluation on all non-redundant protein -: DNA complexes in PDB shows significant performance improvement of our method over TM-align and iAlign (with the ': blackbox preprocessing ': ). Two case studies where our method discovers, for the first time, structural similarities between two pairs of functionally related protein -: DNA complexes are presented. We further demonstrate the power of our method on detecting structural similarities between a protein -: protein complex and a protein -: RNA complex, which is biologically known as a protein -: RNA mimicry case. AVAILABILITY AND IMPLEMENTATION: The PROSTA-inter web-server is publicly available at http://www.cbrc.kaust.edu.sa/prosta/. Xuefeng Cui, Hammad Naveed, Xin Gao 0001 |
Bioinform. | 1 |
| 2014 | Fingerprinting protein structures effectively and efficientlyabstractMOTIVATION: One common task in structural biology is to assess the similarities and differences among protein structures. A variety of structure alignment algorithms and programs has been designed and implemented for this purpose. A major drawback with existing structure alignment programs is that they require a large amount of computational time, rendering them infeasible for pairwise alignments on large collections of structures. To overcome this drawback, a fragment alphabet learned from known structures has been introduced. The method, however, considers local similarity only, and therefore occasionally assigns high scores to structures that are similar only in local fragments. METHOD: We propose a novel approach that eliminates false positives, through the comparison of both local and remote similarity, with little compromise in speed. Two kinds of contact libraries (ContactLib) are introduced to fingerprint protein structures effectively and efficiently. Each contact group of the contact library consists of one local or two remote fragments and is represented by a concise vector. These vectors are then indexed and used to calculate a new combined hit-rate score to identify similar protein structures effectively and efficiently. RESULTS: We tested our method on the high-quality protein structure subset of SCOP30 containing 3297 protein structures. For each protein structure of the subset, we retrieved its neighbor protein structures from the rest of the subset. The best area under the Receiver-Operating Characteristic curve, archived by ContactLib, is as high as 0.960. This is a significant improvement compared with 0.747, the best result achieved by FragBag. We also demonstrated that incorporating remote contact information is critical to consistently retrieve accurate neighbor protein structures for all- query protein structures. AVAILABILITY AND IMPLEMENTATION: https://cs.uwaterloo.ca/∼xfcui/contactlib/. Xuefeng Cui, Shuaicheng Li 0001, Lin He 0002, Ming Li 0001 |
Bioinform. | 1 |
| 2013 | Towards Reliable Automatic Protein Structure Alignment
Xuefeng Cui, Shuaicheng Li 0001, Dongbo Bu, Ming Li 0001 |
WABI | 1 |
| 2012 | How Accurately Can We Model Protein Structures with Dihedral Angles?
Xuefeng Cui, Shuaicheng Li 0001, Dongbo Bu, Babak Alipanahi, Ming Li 0001 |
WABI | 1 |