EDBT 2026 Demo / reviewers in the wild / expert
Xinqi Gong
dblp:200/7127
· DBLP profile ↗
13ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0003-2802-6176ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PCIPG2.0: multi-omics fusion and structure-aware graph autoencoding for protein complex identificationabstractMOTIVATION: Protein complexes execute cellular functions, yet identifying them from protein-protein interaction (PPI) networks remains challenging because interactomes are incomplete and purely topology-driven clustering often lacks mechanistic interpretability. Here we present PCIPG 2.0, an unsupervised framework that explicitly addresses two major bottlenecks in PPI-based complex discovery: missing interactions and limited mechanistic specificity. PCIPG 2.0 first fuses multiple omics views to prioritize high-confidence candidate protein associations and enhance the observed interactome, and then learns structure-aware node representations by aggregating residue embeddings on residue graphs, coupled with a PPI-level graph autoencoder to infer latent complex memberships. RESULTS: Across five yeast benchmarks, PCIPG 2.0 consistently improves complex recovery over representative baselines and yields predicted complexes with significantly higher Gene Ontology semantic coherence than size-matched random sets. Literature-supported case studies and AlphaFold3-based assembly analyses further suggest that representative predictions are consistent with functionally coherent and structurally plausible protein assemblies. Together, these results suggest that combining multi-omics-driven interactome completion with residue-informed representation learning provides a useful and mechanistically informed framework for protein complex identification under incomplete interactome measurements. AVAILABILITY: The source code is available at GitHub: https://github.com/hyx-1/PCIPG2.0. The complete reproducibility package, including the code, processed data, configuration files and materials required to reproduce the experiments reported in this manuscript, has been archived on Zenodo with an archival DOI: https://doi.org/10.5281/zenodo.20133228. The GitHub repository also provides the README-based usage instructions. Jiudong Wang, Xinqi Gong |
Bioinform. | 4 |
| 2025 | TopoQA: a topological deep learning-based approach for protein complex structure interface quality assessmentabstractEven with the significant advances of AlphaFold-Multimer (AF-Multimer) and AlphaFold3 (AF3) in protein complex structure prediction, their accuracy is still not comparable with monomer structure prediction. Efficient and effective quality assessment (QA) or estimation of model accuracy models that can evaluate the quality of the predicted protein-complexes without knowing their native structures are of key importance for protein structure generation and model selection. In this paper, we leverage persistent homology (PH) to capture the atomic-level topological information around residues and design a topological deep learning-based QA method, TopoQA, to assess the accuracy of protein complex interfaces. We integrate PH from topological data analysis into graph neural networks (GNNs) to characterize complex higher-order structures that GNNs might overlook, enhancing the learning of the relationship between the topological structure of complex interfaces and quality scores. Our TopoQA model is extensively validated based on the two most-widely used benchmark datasets, Docking Benchmark5.5 AF2 (DBM55-AF2) and Heterodimer-AF2 (HAF2), along with our newly constructed ABAG-AF3 dataset to facilitate comparisons with AF3. For all three datasets, TopoQA outperforms AF-Multimer-based AF2Rank and shows an advantage over AF3 in nearly half of the targets. In particular, in the DBM55-AF2 dataset, a ranking loss of 73.6% lower than AF-Multimer-based AF2Rank is obtained. Further, other than AF-Multimer and AF3, we have also extensively compared with nearly-all the state-of-the-art models (as far as we know), it has been found that our TopoQA can achieve the highest Top 10 Hit-rate on the DBM55-AF2 dataset and the lowest ranking loss on the HAF2 dataset. Ablation experiments show that our topological features significantly improve the model's performance. At the same time, our method also provides a new paradigm for protein structure representation learning. Bingqing Han, Xinqi Gong, Kelin Xia |
Briefings Bioinform. | 4 |
| 2025 | $\mathcal{S}$ able: bridging the gap in protein structure understanding with an empowering and versatile pre-training paradigmabstractProtein pre-training has emerged as a transformative approach for solving diverse biological tasks. While many contemporary methods focus on sequence-based language models, recent findings highlight that protein sequences alone are insufficient to capture the extensive information inherent in protein structures. Recognizing the crucial role of protein structure in defining function and interactions, we introduce $\mathcal{S}$able, a versatile pre-training model designed to comprehensively understand protein structures. $\mathcal{S}$able incorporates a novel structural encoding mechanism that enhances inter-atomic information exchange and spatial awareness, combined with robust pre-training strategies and lightweight decoders optimized for specific downstream tasks. This approach enables $\mathcal{S}$able to consistently outperform existing methods in tasks such as generation, classification, and regression, demonstrating its superior capability in protein structure representation. The code and models can be accessed via GitHub repository at https://github.com/baaihealth/Sable. Jiashan Li, Mingliang Zeng, Jingcheng Yu, Xinqi Gong, Qiwei Ye |
Briefings Bioinform. | 6 |
| 2025 | Harnessing pre-trained models for accurate prediction of protein-ligand binding affinityabstractBACKGROUND: The binding between proteins and ligands plays a crucial role in the field of drug discovery. However, this area currently faces numerous challenges. On one hand, existing methods are constrained by the limited availability of labeled data, often performing inadequately when addressing complex protein-ligand interactions. On the other hand, many models struggle to effectively capture the flexible variations and relative spatial relationships between proteins and ligands. These issues not only significantly hinder the advancement of protein-ligand binding research but also adversely affect the accuracy and efficiency of drug discovery. Therefore, in response to these challenges, our study aims to enhance predictive capabilities through innovative approaches, providing more reliable support for drug discovery efforts. METHODS: This study leverages a pre-trained model with spatial awareness to enhance the prediction of protein-ligand binding affinity. By perturbing the structures of small molecules in a manner consistent with physical constraints and employing self-supervised tasks, we improve the representation of small molecule structures, allowing for better adaptation to affinity predictions. Meanwhile, our approach enables the identification of potential binding sites on proteins. RESULTS: Our model demonstrates a significantly higher correlation coefficient in binding affinity predictions. Extensive evaluation on the PDBBind v2019 refined set, CASF, and Merck FEP benchmarks confirms the model's robustness and strong generalization across diverse datasets. Additionally, the model achieves over 95% in classification ROC for binding site identification, underscoring its high accuracy in pinpointing protein-ligand interaction regions. CONCLUSION: This research presents a novel approach that not only enhances the accuracy of binding affinity predictions but also facilitates the identification of binding sites, showcasing the potential of pre-trained models in computational drug design. Data and code are available at https://github.com/MIALAB-RUC/SableBind . Jiashan Li, Xinqi Gong |
BMC Bioinform. | 2 |
| 2024 | ABAG-docking benchmark: a non-redundant structure benchmark dataset for antibody-antigen computational dockingabstractAccurate prediction of antibody-antigen complex structures is pivotal in drug discovery, vaccine design and disease treatment and can facilitate the development of more effective therapies and diagnostics. In this work, we first review the antibody-antigen docking (ABAG-docking) datasets. Then, we present the creation and characterization of a comprehensive benchmark dataset of antibody-antigen complexes. We categorize the dataset based on docking difficulty, interface properties and structural characteristics, to provide a diverse set of cases for rigorous evaluation. Compared with Docking Benchmark 5.5, we have added 112 cases, including 14 single-domain antibody (sdAb) cases and 98 monoclonal antibody (mAb) cases, and also increased the proportion of Difficult cases. Our dataset contains diverse cases, including human/humanized antibodies, sdAbs, rodent antibodies and other types, opening the door to better algorithm development. Furthermore, we provide details on the process of building the benchmark dataset and introduce a pipeline for periodic updates to keep it up to date. We also utilize multiple complex prediction methods including ZDOCK, ClusPro, HDOCK and AlphaFold-Multimer for testing and analyzing this dataset. This benchmark serves as a valuable resource for evaluating and advancing docking computational methods in the analysis of antibody-antigen interaction, enabling researchers to develop more accurate and effective tools for predicting and designing antibody-antigen complexes. The non-redundant ABAG-docking structure benchmark dataset is available at https://github.com/Zhaonan99/Antibody-antigen-complex-structure-benchmark-dataset. Bingqing Han, Cuicui Zhao, Jinbo Xu, Xinqi Gong |
Briefings Bioinform. | 5 |
| 2023 | DSR: Dynamical Surface Representation as Implicit Neural Networks for ProteinabstractWe propose a novel neural network-based approach to modeling protein dynamics using an implicit representation of a protein’s surface in 3D and time. Our method utilizes the zero-level set of signed distance functions (SDFs) to represent protein surfaces, enabling temporally and spatially continuous representations of protein dynamics. Our experimental results demonstrate that our model accurately captures protein dynamic trajectories and can interpolate and extrapolate in 3D and time. Importantly, this is the first study to introduce this method and successfully model large-scale protein dynamics. This approach offers a promising alternative to current methods, overcoming the limitations of first-principles-based and deep learning methods, and provides a more scalable and efficient approach to modeling protein dynamics. Additionally, our surface representation approach simplifies calculations and allows identifying movement trends and amplitudes of protein domains, making it a useful tool for protein dynamics research. Codes are available at https://github.com/Sundw-818/DSR, and we have a project webpage that shows some video results, https://sundw-818.github.io/DSR/. Daiwen Sun, Xinqi Gong, Qiwei Ye |
NeurIPS | 4 |
| 2022 | Inter-chain contact map prediction for protein complex based on graph attention network and triangular multiplication updateabstractResidue-residue interactions between individual subunits of protein complexes are critical for predicting complex structures and can serve as distance constraints to guide complex structure modeling. Some recent studies have made some progress in predicting protein inter-chain contact maps based on multiple sequence alignments and deep learning models. Here we develop a new model based on graph attention network and triangular multiplication update to predict interchain contact maps for homologous protein complexes, named PGT (P is Protein, G is Graph attention network and T is Triangular multiplication update). Different from other methods which need to perform multiple sequence alignment processes and extract complicated manual features, PGT extracts embeddings of residues through the protein language model. Besides, we introduce structural information through the graph attention network to learn the spatial information of subunits from the complex structure and utilize the triangular multiplication module to capture triangular constraints between residues. To demonstrate the effectiveness of our method, we compare PGT with previous works such as DeepHomo, DRCon and Glinter on two independent test sets. The results show that PGT substantially outperforms these methods. Furthermore, we also perform two ablation experiments to demonstrate the necessity of introducing graph attention network and triangular multiplication update. In all, our framework presents new modules to accurately predict inter-chain contact maps in homologous protein complexes and it’s also useful to analyze interactions in other type of protein complexes. Jiashan Li, Wenda Wang 0004, Xinqi Gong |
BIBM | 5 |
| 2022 | An adaptive variational model for multireference alignment with mixed noiseabstractThe multireference alignment (MRA) problem is to estimate an underlying signal from a large number of noisy circularly-shifted observations. The existing methods are under the hypothesis of a single Gaussian noise. However, the hypothesis of a single-type noise is inefficient for solving practical problems like single particle cryo-EM. In this paper, we derive an adaptive variational model by combining maximum a posteriori (MAP) estimation and the soft-max method under the assumption of Gaussian mixture noise. There are two adaptive weights for detecting cyclical shifts and types of noise separately. The existence of a minimizer is mathematically proved. We design a novel algorithm for the proposed model using the alternating direction iterative method and the augmented Lagrange method. There are some convergence analyses for the proposed algorithm under some conditions. The numerical results show that the proposed model performs better than the existing methods in that the level of one Gaussian noise is high and the other is low. Cuicui Zhao, Jun Liu 0029, Xinqi Gong |
BIBM | 3 |
| 2021 | Inter-protein contact map generated only from intra-monomer by image inpaintingabstractIntra-protein contact or distance prediction has seen many progresses in recent years with the development of deep learning algorithms combined with direct coupling analysis. Some works try to extend these state-of-art intra-protein methods to inter-protein. However, it’s difficult to build high-quality joint multiple sequence alignments (MSAs) for protein dimers, especially for heterodimers. Here we propose a generative adversarial network (GAN) based method, Protein Dimer Image Inpainting (PDII), to predict inter-protein contact map (CM) solely from monomer structures without any co-evolution information. PDII can learn intrinsic contact patterns with local coherence and global consistency for several cases. It’s robust to monomer structure quality as our method is only based on joint CMs and doesn’t need precise structure information. Furthermore, PDII not only works well on bound monomers but also on unbound ones. When evaluating on 3Dcomplex datasets, our method works better than DNCON-inter and RaptorX-ComplexContact. Tested on homodimer proteins with C2 symmetry type, PDII can make good predictions if homologous sequences are relatively small. Besides, it presents good patch accuracy on the MaSif-PPI dataset. In all, PDII is an effective way to predict contacts in dimer protein complexes and provide new understandings for interactions in protein complexes. Chengshi Zeng, Xinqi Gong |
BIBM | 3 |
| 2020 | Heterogeneous multiple kernel learning for breast cancer outcome evaluationabstractBACKGROUND: Breast cancer is one of the common kinds of cancer among women, and it ranks second among all cancers in terms of incidence, after lung cancer. Therefore, it is of great necessity to study the detection methods of breast cancer. Recent research has focused on using gene expression data to predict outcomes, and kernel methods have received a lot of attention regarding the cancer outcome evaluation. However, selecting the appropriate kernels and their parameters still needs further investigation. RESULTS: We utilized heterogeneous kernels from a specific kernel set including the Hadamard, RBF and linear kernels. The mixed coefficients of the heterogeneous kernel were computed by solving the standard convex quadratic programming problem of the quadratic constraints. The algorithm is named the heterogeneous multiple kernel learning (HMKL). Using the particle swarm optimization (PSO) in HMKL, we selected the kernel parameters, then we employed HMKL to perform the breast cancer outcome evaluation. By testing real-world microarray datasets, the HMKL method outperforms the methods of the random forest, decision tree, GA with Rotation Forest, BFA + RF, SVM and MKL. CONCLUSIONS: On one hand, HMKL is effective for the breast cancer evaluation and can be utilized by physicians to better understand the patient's condition. On the other hand, HMKL can choose the function and parameters of the kernel. At the same time, this study proves that the Hadamard kernel is effective in HMKL. We hope that HMKL could be applied as a new method to more actual problems. Xingheng Yu, Xinqi Gong |
BMC Bioinform. | 2 |
| 2019 | Attention mechanism enhanced LSTM with residual architecture and its application for protein-protein interaction residue pairs predictionabstractBACKGROUND: Recurrent neural network(RNN) is a good way to process sequential data, but the capability of RNN to compute long sequence data is inefficient. As a variant of RNN, long short term memory(LSTM) solved the problem in some extent. Here we improved LSTM for big data application in protein-protein interaction interface residue pairs prediction based on the following two reasons. On the one hand, there are some deficiencies in LSTM, such as shallow layers, gradient explosion or vanishing, etc. With a dramatic data increasing, the imbalance between algorithm innovation and big data processing has been more serious and urgent. On the other hand, protein-protein interaction interface residue pairs prediction is an important problem in biology, but the low prediction accuracy compels us to propose new computational methods. RESULTS: In order to surmount aforementioned problems of LSTM, we adopt the residual architecture and add attention mechanism to LSTM. In detail, we redefine the block, and add a connection from front to back in every two layers and attention mechanism to strengthen the capability of mining information. Then we use it to predict protein-protein interaction interface residue pairs, and acquire a quite good accuracy over 72%. What's more, we compare our method with random experiments, PPiPP, standard LSTM, and some other machine learning methods. Our method shows better performance than the methods mentioned above. CONCLUSION: We present an attention mechanism enhanced LSTM with residual architecture, and make deeper network without gradient vanishing or explosion to a certain extent. Then we apply it to a significant problem- protein-protein interaction interface residue pairs prediction and obtain a better accuracy than other methods. Our method provides a new approach for protein-protein interaction computation, which will be helpful for related biomedical researches. Xinqi Gong |
BMC Bioinform. | 2 |
| 2019 | Protein-Protein Interaction Interface Residue Pair Prediction Based on Deep Learning ArchitectureabstractMOTIVATION: Proteins usually fulfill their biological functions by interacting with other proteins. Although some methods have been developed to predict the binding sites of a monomer protein, these are not sufficient for prediction of the interaction between two monomer proteins. The correct prediction of interface residue pairs from two monomer proteins is still an open question and has great significance for practical experimental applications in the life sciences. We hope to build a method for the prediction of interface residue pairs that is suitable for those applications. RESULTS: Here, we developed a novel deep network architecture called the multi-layered Long-Short Term Memory networks (LSTMs) approach for the prediction of protein interface residue pairs. First, we created three new descriptions and used other six worked characterizations to describe an amino acid, then we employed these features to discriminate between interface residue pairs and non-interface residue pairs. Second, we used two thresholds to select residue pairs that are more likely to be interface residue pairs. Furthermore, this step increases the proportion of interface residue pairs and reduces the influence of imbalanced data. Third, we built deep network architectures based on Long-Short Term Memory networks algorithm to organize and refine the prediction of interface residue pairs by employing features mentioned above. We trained the deep networks on dimers in the unbound state in the international Protein-protein Docking Benchmark version 3.0. The updated data sets in the versions 4.0 and 5.0 were used as the validation set and test set respectively. For our best model, the accuracy rate was over 62 percent when we chose the top 0.2 percent pairs of every dimer in the test set as predictions, which will be very helpful for the understanding of protein-protein interaction mechanisms and for guidance in biological experiments. Zhenni Zhao, Xinqi Gong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | Understanding Protein-Protein Interface Formation Mechanism in a New Probability Way at Amino Acid Level
Yongxiao Yang, Xinqi Gong |
ISBRA | 2 |