EDBT 2026 Demo / reviewers in the wild / expert
Xue Li 0019
dblp:181/2710-19
· DBLP profile ↗
16ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-1489-3095ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Structurally Stabilized Representations for Lossless DNA StorageabstractThis paper presents Reed-Solomon coded single-stranded representation learning (RSRL), a novel end-to-end model for learning representations for lossless DNA data storage. In contrast to existing learning-based methods, RSRL is inspired by both error-correction codec and structural biology. Specifically, RSRL first learns the representations for the subsequent storage from the binary data transformed by the Reed-Solomon codec (RS code). Then, the representations are masked by an RS-code-informed mask to focus on correcting the burst errors occurring in the learning process. The synergy of RS masks and graph attention enables active error localization, breaking through the limitations of traditional passive error correction. With the decoded representations with error corrections, a novel biologically stabilized loss is formulated to regularize the data representations to possess stable single-stranded structures. By incorporating these novel strategies, RSRL can learn highly durable, dense, and lossless representations for subsequent storage tasks in DNA sequences. The proposed RSRL has been compared with a number of baselines in real-world tasks of multi-type data storage. The experimental results obtained demonstrate that RSRL can store diverse types of data with much higher information density and durability, but much lower error rates. Ben Cao, Xue Li 0019, Tiantian He 0001, Bin Wang 0005, Shihua Zhou, Qiang Zhang 0008 |
AAAI | 2 |
| 2026 | Predicting protein-protein interaction sites based on dynamic perception mechanism within a hierarchical E(n)-equivariant graphabstractAccurate prediction of protein-protein interaction sites is crucial to understanding biological processes, elucidating disease mechanisms, and accelerating drug discovery. Although graph neural network methods have shown potential in this field, but existing methods are limited by the static integrate multi-group features and insufficient perception of hierarchical 3D spatial geometric information, leading to insufficient predictive ability of orphan sites. To address these issues, this paper proposes a Dperception mechanism within a Hierarchical E(n)-equivariant Graph architecture (DHEG). DHEG introduces a dynamic feature importance perception mechanism that adaptively perceives the contextual inter-dependencies of features and assigns weights to feature groups based on their relevance to the interaction relationship. And a hierarchical gated architecture based on E(n)-equivariant graph neural networks that effectively captures protein 3D spatial structures while mitigating over-smoothing problems. The results show that DHEG achieves improvements in 11 of 13 key metrics, with an enhancement 8% in Matthews correlation coefficient, indicating that DHEG not only predicts more interaction sites but also does so with greater reliability. Furthermore, case studies and visualization analyzes show that DHEG aligns better with the biological mechanism and has excellent predictive capabilities for both orphan sites and continuous regions, demonstrating interpretability, and application potential. Xue Li 0019, Suheng Qiao, Shihua Zhou, Jianmin Wang 0016, Bin Wang 0005, Tao Song 0001, Ben Cao |
Briefings Bioinform. | 1 |
| 2026 | An end-to-end DNA storage coding method based on a low-complexity multiple biological constraints loss and RL-inspired differentiable solver
Yanfen Zheng, Xue Li 0019, Bin Wang 0005, Shihua Zhou, Ben Cao, Pan Zheng 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Integrating histology and spatial transcriptomics via multimodal transformers and contrastive representation learning for accurate gene expression prediction
Liuming Shi, Xue Li 0019, Bin Wang 0005, Shihua Zhou, Ben Cao, Pan Zheng 0001 |
J. Biomed. Informatics | 3 |
| 2026 | PTPPI: A Study on Protein Inhibitor Prediction Methods Using Multimodal Feature Fusion and Attention MechanismabstractProtein-protein interactions (PPIs) are fundamental to many biological processes, including cell signaling, gene expression regulation, immune responses, and protein complex formation. Small molecule inhibitors targeting specific PPIs are expected to treat diseases such as cancer and viral infections by modulating pathophysiological processes. Despite their clinical importance, the development of PPI inhibitors is challenging due to limited experimental validation data, which complicates the accurate prediction of novel inhibitors. Therefore, there is an urgent need for advanced computational methods that can effectively integrate multiple data types and improve prediction accuracy. In this study, we proposed a new framework, PTPPI, to efficiently predict protein-protein interaction inhibitors (PPIIs). PTPPI integrates multiple molecular features, including extended connectivity fingerprints (ECFPs) for structural representation and deep semantic embeddings of SMILES sequences generated by the ChemBERTa pre-trained model. These features are processed by independent encoders and fused using an interactive attention mechanism, which enhances the molecular representation. In addition, PTPPI adopts a multi-task learning approach, enabling the model to both reconstruct input features and accurately predict inhibition scores. Experimental results on eight PPI target families, focusing on inhibitor identification and potency prediction, demonstrate that PTPPI outperforms existing methods. It not only integrates multiple molecular features effectively but also achieves superior prediction performance. This makes PTPPI a valuable and reliable tool for discovering new PPI inhibitors, thus opening up new possibilities for drug discovery and disease treatment. Zhao Yang Dong, Peifu Han, Xue Li 0019, Tao Song 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | TriPercept: A Unified Fine-Grained Structural Perception Model for Molecular Property PredictionabstractAccurate prediction of molecular properties is a fundamental component of Artificial Intelligence-driven Drug Design (AIDD). In molecular property prediction, even slight structural variations, such as atomic arrangement or bond length, could impact molecular properties. Existing methods primarily focus on coarse-grained molecular modeling on SMILES, 2D, or 3D graphs. However, this paradigm struggles to adequately capture complex internal molecular interactions and overlooks finer-grained structural details, such as atom-bond-spatial positions. In this paper, we propose a molecular property prediction model, TriPercept, based on unified fine-grained representation learning. Specifically, TriPercept integrates atomic, topological, and geometric features, fully exploiting their intrinsic complementarity. We introduce three dedicated encoders to separately learn features at the atomic, bond, and spatial distance levels, thus addressing the neglect of complex structural details. Additionally, we design a structure-aware graph neural network to integrate these features and enhance structural consistency modeling with a contrastive and self-supervised based pretraining strategy. To evaluate the performance of TriPercept, we conduct systematic experiments on six classification datasets and six regression datasets. The results demonstrate that TriPercept performs excellently on most datasets. Visualization experiments further verify that TriPercept not only confirms the effectiveness of multi-level structural integration but also clearly reveals its interpretability in focusing on chemically relevant features and capturing complex structural relationships. The data and code are available at https://github.com/TiAW-Go/TriPercept. Xue Li 0019, Peifu Han, Kexin Jin, Tao Song 0001 |
BIBM | 2 |
| 2025 | MVSO-PPIS: a structured objective learning model for protein-protein interaction sites prediction via multi-view graph information integrationabstractMOTIVATION: Predicting protein-protein interaction (PPI) sites is essential for advancing our understanding of protein interactions, as accurate predictions can significantly reduce experimental costs and time. While considerable progress has been made in identifying binding sites at the level of individual amino acid residues, the prediction accuracy for residue subsequences at transitional boundaries-such as those represented by patterns like singular structures (mutation characteristics of contiguous interacting-residue segments) or edge structures (boundary transitions between interacting/non-interacting residue segments) still requires improvement. RESULTS: we propose a novel PPI site prediction method named MVSO-PPIS. This method integrates two complementary feature extraction modules, a subgraph-based module and an enhanced graph attention module. The extracted features are fused using an attention-based fusion mechanism, producing a composite representation that captures both local protein substructures and global contextual dependencies. MVSO-PPIS is trained to jointly optimize three objectives: overall PPI site prediction accuracy, edge structural consistency, and recognition of unique structural patterns in PPI site sequences. Experimental results on benchmark datasets demonstrate that MVSO-PPIS outperforms existing baseline models in both accuracy and structural interpretability. AVAILABILITY AND IMPLEMENTATION: The datasets, source codes, and models of MVSO-PPIS are all available at https://github.com/Edwardblue282/MVSO-PPIS. Tianle Ma, Kaiyu Dong, Peifu Han, Xue Li 0019, Junteng Ma, Tao Song 0001 |
Bioinform. | 5 |
| 2025 | Multi-scale low-frequency enhanced spectral neural operator for reducing low-frequency error in partial differential equations solving
Fengrui Jing, Chuchu Zhai, Peizhi Zhao, Xue Li 0019, Peifu Han, Hongzhen Ding, Yunlong Dong, Long Hao, Tao Song 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | TGF-M: Topology-augmented geometric features enhance molecular property predictionabstractAccurate prediction of molecular properties is a key component of Artificial Intelligence-driven Drug Design (AIDD). Despite significant progress in improving these predictive models, balancing accuracy with computational complexity remains a challenge. Molecular topological and geometric features provide rich spatial information, crucial for improving prediction accuracy, but their extraction typically increases model complexity. To address this, we propose TGF-M (Topology-augmented Geometric Features for Molecular Property Prediction), a novel predictive model that optimizes feature extraction to enhance information capture and improve model accuracy, and reduces model complexity to lower computational cost. This approach enhances the model's ability to leverage both topological and geometric features without unnecessary complexity. On the re-segmented PCQM4Mv2 dataset, TGF-M performs remarkably, achieving a low mean absolute error (MAE) of 0.0647 in the HOMO-LUMO gap prediction task with only 6.4M parameters. Compared to two recent state-of-the-art models evaluated within a unified validation framework, TGF-M demonstrates comparable performance with less than one-tenth of the parameters. We conducted an in-depth analysis of TGF-M's chemical interpretability. The results further validate the method's effectiveness in leveraging complex molecular topology and geometry during model learning, underscoring its potential and advantages. The trained models and source code of TGF-M are publicly available at https://github.com/TiAW-Go/TGF-M. Xue Li 0019, Peifu Han, Tao Song 0001 |
PLoS Comput. Biol. | 3 |
| 2025 | Gene-MOE: A Sparsely Gated Cancer Diagnosis and Prognosis Framework Exploiting Pan-Cancer Genomic InformationabstractImproved cancer genomic diagnosis and prognosis are vital to accurate medical therapy. Deep learning methods offered an end-to-end solution to enhance the precision of analysis. With the fast pace of pre-trained Transformer models, it remains uncertain whether some novel approaches such as the sparsely gated mixture of expert (MOE) and self-attention mechanisms can further improve the precision of cancer prognosis and classification. In this paper, we introduce a novel sparsely gated cancer diagnosis and prognosis framework called Gene-MOE exploiting the potential of the MOE layers and the proposed mixture of attention expert (MOAE) layers to enhance the analysis accuracy. Additionally, we address overfitting challenges by integrating pan-cancer information from 33 distinct cancer types through pre-training. For survival analysis, Gene-MOE achieves the best Concordance Index compared with state-of-the-art models on 12 of 14 cancer types. For cancer classification, the total accuracy of the classification model for 33 cancer classifications reached 95.8%, representing the best performance compared to state-of-the-art models. For cancer subtyping, Gene-MOE achieves the best result on at least one metric of the log10 P-values and the number of significant clinical on seven of nine cancers. These results indicate that Gene-MOE holds strong potential for these downstream tasks. Xiangyu Meng 0005, Xue Li 0019, Huanhuan Dai, Lian Qiao, Hongzhen Ding, Long Hao, Xun Wang 0010 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | Predicting Mutation-Disease Associations Through Protein Interactions Via Deep LearningabstractDisease is one of the primary factors affecting life activities, with complex etiologies often influenced by gene expression and mutation. Currently, wet lab experiments have analyzed the mechanisms of mutations, but these are usually limited by the costs of wet experiments and constraints in sample types and scales. Therefore, this paper constructs a real-world mutation-induced disease dataset and proposes Capsule and Graph topology networks with Multi-head attention (CGM) to predict the mutation-disease associations. CGM can accurately predict protein mutation-disease associations, and to further elucidate the pathogenicity of protein mutations, we also verified that protein mutations lead to protein structural alterations by the model, which suggests that mutation-induced conformational changes may be an important pathogenic factor. Limited by the size of the mutated protein dataset, we also performed experiments on benchmark and imbalanced datasets, where CGM mined 22 unknown protein interaction pairs from the benchmark dataset, better illustrating the potential of CGM in predicting mutation-disease associations. In summary, this paper curates a real dataset. It proposes that CGM predicts protein mutations and disease associations, providing a novel tool for further understanding of biomolecular pathways and disease mechanisms. Xue Li 0019, Ben Cao, Jianmin Wang 0016, Xiangyu Meng 0005, Yu Huang 0004, Enrico Petretto, Tao Song 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Semi-Template Retrosynthesis Prediction with 3D Spatial Structures and Reaction Center Disconnection RulesabstractRetrosynthesis constructs rational synthetic pathways by predicting reactants from the target product. Previous research utilizing Graph Neural Networks often relies on two-dimensional molecular representations. The two-dimensional representations of functional groups of the same type are identical, and ignoring differences in their three-dimensional(3D) structures can cause confusion in identifying reaction centers. Additionally, they disconnect the reaction centers, hindering the accurate identification of chemical transformation rules. To address these issues, we propose the Retro3D model based on a semi-template-based method, integrating 3D structural information to enhance molecular representation. Due to the discrepancies in bond lengths and bond angles of the same type of functional groups in different 3D spatial structures, our method improves the accuracy of identifying reaction centers in the target product. Furthermore, our method determines whether reaction centers are disconnected, differentiating between types of reaction centers to obtain two types of chemical transformation rules. Given the known reaction class, experiments demonstrate that our model achieves superior accuracy in predicting reaction centers and final reactants compared to previous benchmark models. In summary, our methodology aligns with established chemical principles, enhancing its interpretability and broad applicability across diverse synthetic challenges. Xin Li 0244, Zishuai Wei, Peifu Han, Xue Li 0019, Tao Song 0001 |
BIBM | 5 |
| 2024 | KGM4DTI: A Depression Knowledge Graph-Driven Multi-semantic Multi-view Computational Framework for Drug-Target Interaction PredictionabstractDepression is a prevalent global mental health issue that significantly impacts both individuals and societies. Current methods often overlook the heterogeneity of depressive symptoms and fail to fully leverage available data, underscoring the need for more advanced computational methods DTI prediction. This paper introduces KGM4DTI, a multi-semantic, multi-view, and multi-level meta-path Transformer framework that integrates Transformer and GNN. KGM4DTI effectively captures both global semantic and local structural information between interaction pairs within a depression knowledge graph (DepKG), which comprises 2,258 drugs, 3,360 proteins, 6 subtypes of depression, 196 phenotypes, and 238 pathways. Comparative experiments have demonstrated that KGM4DTI significantly outperforms existing state-of-the-art methods, achieving an AUROC of 99.72 and an AUPR of 99.76. The robustness of framework is further confirmed through extensive ablation studies, which highlight the crucial role of integrating both global and local information. These results underscore the potential of KGM4DTI in accurately predicting DTI. This approach has promising implications for drug discovery, particularly in developing effective treatments for depression, and could extend its applicability to broader mental health disorders. The data and code are available at https://github.com/Tianxyuu/KGM4DTI/. Peifu Han, Xue Li 0019, Hongzhen Ding, Fengrui Jing, Xun Wang 0010, Tao Song 0001 |
BIBM | 3 |
| 2024 | MEG-PPIS: a fast protein-protein interaction site prediction method based on multi-scale graph information and equivariant graph neural networkabstractMOTIVATION: Protein-protein interaction sites (PPIS) are crucial for deciphering protein action mechanisms and related medical research, which is the key issue in protein action research. Recent studies have shown that graph neural networks have achieved outstanding performance in predicting PPIS. However, these studies often neglect the modeling of information at different scales in the graph and the symmetry of protein molecules within three-dimensional space. RESULTS: In response to this gap, this article proposes the MEG-PPIS approach, a PPIS prediction method based on multi-scale graph information and E(n) equivariant graph neural network (EGNN). There are two channels in MEG-PPIS: the original graph and the subgraph obtained by graph pooling. The model can iteratively update the features of the original graph and subgraph through the weight-sharing EGNN. Subsequently, the max-pooling operation aggregates the updated features of the original graph and subgraph. Ultimately, the model feeds node features into the prediction layer to obtain prediction results. Comparative assessments against other methods on benchmark datasets reveal that MEG-PPIS achieves optimal performance across all evaluation metrics and gets the fastest runtime. Furthermore, specific case studies demonstrate that our method can predict more true positive and true negative sites than the current best method, proving that our model achieves better performance in the PPIS prediction task. AVAILABILITY AND IMPLEMENTATION: The data and code are available at https://github.com/dhz234/MEG-PPIS.git. Hongzhen Ding, Xue Li 0019, Peifu Han, Fengrui Jing, Tao Song 0001, Hanjiao Fu, Na Kang |
Bioinform. | 2 |
| 2024 | MIPPIS: protein-protein interaction site prediction network with multi-information fusionabstractBACKGROUND: The prediction of protein-protein interaction sites plays a crucial role in biochemical processes. Investigating the interaction between viruses and receptor proteins through biological techniques aids in understanding disease mechanisms and guides the development of corresponding drugs. While various methods have been proposed in the past, they often suffer from drawbacks such as long processing times, high costs, and low accuracy. RESULTS: Addressing these challenges, we propose a novel protein-protein interaction site prediction network based on multi-information fusion. In our approach, the initial amino acid features are depicted by the position-specific scoring matrix, hidden Markov model, dictionary of protein secondary structure, and one-hot encoding. Simultaneously, we adopt a multi-channel approach to extract deep-level amino acids features from different perspectives. The graph convolutional network channel effectively extracts spatial structural information. The bidirectional long short-term memory channel treats the amino acid sequence as natural language, capturing the protein's primary structure information. The ProtT5 protein large language model channel outputs a more comprehensive amino acid embedding representation, providing a robust complement to the two aforementioned channels. Finally, the obtained amino acid features are fed into the prediction layer for the final prediction. CONCLUSION: , Matthews correlation coefficient, and area under the precision recall curve, which demonstrates the superiority of our model. Kaiyu Dong, Dingming Liang, Yunjing Zhang, Xue Li 0019, Tao Song 0001 |
BMC Bioinform. | 5 |
| 2023 | MARPPI: boosting prediction of protein-protein interactions with multi-scale architecture residual networkabstractProtein-protein interactions (PPIs) are a major component of the cellular biochemical reaction network. Rich sequence information and machine learning techniques reduce the dependence of exploring PPIs on wet experiments, which are costly and time-consuming. This paper proposes a PPI prediction model, multi-scale architecture residual network for PPIs (MARPPI), based on dual-channel and multi-feature. Multi-feature leverages Res2vec to obtain the association information between residues, and utilizes pseudo amino acid composition, autocorrelation descriptors and multivariate mutual information to achieve the amino acid composition and order information, physicochemical properties and information entropy, respectively. Dual channel utilizes multi-scale architecture improved ResNet network which extracts protein sequence features to reduce protein feature loss. Compared with other advanced methods, MARPPI achieves 96.03%, 99.01% and 91.80% accuracy in the intraspecific datasets of Saccharomyces cerevisiae, Human and Helicobacter pylori, respectively. The accuracy on the two interspecific datasets of Human-Bacillus anthracis and Human-Yersinia pestis is 97.29%, and 95.30%, respectively. In addition, results on specific datasets of disease (neurodegenerative and metabolic disorders) demonstrate the ability to detect hidden interactions. To better illustrate the performance of MARPPI, evaluations on independent datasets and PPIs network suggest that MARPPI can be used to predict cross-species interactions. The above shows that MARPPI can be regarded as a concise, efficient and accurate tool for PPI datasets. Xue Li 0019, Peifu Han, Changnan Gao, Tao Song 0001, Muyuan Niu, Alfonso Rodríguez-Patón |
Briefings Bioinform. | 1 |