Wenjie Du 0003

dblp:96/5778-3 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 2 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MSAnchor: De Novo Molecular Generation from Mass Spectrometry Data with Anchor-Extended Molecular Scaffolds
abstract
Tandem mass spectrometry (MS/MS) is a critical tool for identifying molecular structures. By efficiently separating molecular fragments based on their mass-to-charge (m/z) ratios, it facilitates molecular generation and subsequent scientific discoveries. However, de novo molecular generation from MS/MS spectra remains fundamentally constrained by two paramount challenges: the vast chemical space requires effective structural constraints, and the absence of fine-grained substructural generation weakens the correspondences between spectral features and molecular structures. In this work, we propose MSAnchor, a novel two-stage framework for MS/MS-based molecular structure generation. We mitigate the search space challenge through the introduction of Anchor-Extended Molecular Scaffold (AEMS) representation that explicitly encodes side-chain anchoring points, thereby dramatically reducing combinatorial complexity. Leveraging the explicit attachment sites provided by AEMS, we develop anchor-specific priors that establish effective alignments between spectral features and molecular substructures. This fine-grained substructural correspondence is further enhanced by a modified Conditional Information Bottleneck (CIB) module that extracts the most informative spectral components in a structure-aware manner. These innovations enable MSAnchor to generate molecular structures that closely reflect spectral characteristics while constraining combinatorial complexity. Extensive experiments on the CANOPUS and MassSpecGym datasets demonstrate that MSAnchor achieves state-of-the-art performance in molecular structure prediction from MS/MS spectra, with performance improvements that are particularly more pronounced for molecules with higher complexity.
Xiaohan Qin, Zhengyang Zhou, Linjiang Chen, Wenjie Du 0003, Yang Wang 0015
AAAI5
2025 ReNovo: Retrieval-Based \emph{De Novo} Mass Spectrometry Peptide Sequencing
abstract
Proteomics is the large-scale study of proteins. Tandem mass spectrometry, as the only high-throughput technique for protein sequence identification, plays a pivotal role in proteomics research. One of the long-standing challenges in this field is peptide identification, which entails determining the specific peptide (sequence of amino acids) that corresponds to each observed mass spectrum. The conventional approach involves database searching, wherein the observed mass spectrum is scored against a pre-constructed peptide database. However, the reliance on pre-existing databases limits applicability in scenarios where the peptide is absent from existing databases. Such circumstances necessitate \emph{de novo} peptide sequencing, which derives peptide sequence solely from input mass spectrum, independent of any peptide database. Despite ongoing advancements in \emph{de novo} peptide sequencing, its performance still has considerable room for improvement, which limits its application in large-scale experiments. In this study, we introduce a novel \textbf{Re}trieval-based \emph{De \textbf{Novo}} peptide sequencing methodology, termed \textbf{ReNovo}, which draws inspiration from database search methods. Specifically, by constructing a datastore from training data, ReNovo can retrieve information from the datastore during the inference stage to conduct retrieval-based inference, thereby achieving improved performance. This innovative approach enables ReNovo to effectively combine the strengths of both methods: utilizing the assistance of the datastore while also being capable of predicting novel peptides that are not present in pre-existing databases. A series of experiments have confirmed that ReNovo outperforms state-of-the-art models across multiple widely-used datasets, incurring only minor storage and time consumption, representing a significant advancement in proteomics. Supplementary materials include the code.
Shaorong Chen, Jun Xia 0001, Lecheng Zhang, Zhangyang Gao, Bozhen Hu, Cheng Tan 0012, Wenjie Du 0003, Stan Z. Li
ICLR8
2025 Iterative Substructure Extraction for Molecular Relational Learning with Interactive Graph Information Bottleneck
abstract
Molecular relational learning (MRL) seeks to understand the interaction behaviors between molecules, a pivotal task in domains such as drug discovery and materials science. Recently, extracting core substructures and modeling their interactions have emerged as mainstream approaches within machine learning-assisted methods. However, these methods still exhibit some limitations, such as insufficient consideration of molecular interactions or capturing substructures that include excessive noise, which hampers precise core substructure extraction. To address these challenges, we present an integrated dynamic framework called Iterative Substructure Extraction (ISE). ISE employs the Expectation-Maximization (EM) algorithm for MRL tasks, where the core substructures of interacting molecules are treated as latent variables and model parameters, respectively. Through iterative refinement, ISE gradually narrows the interactions from the entire molecular structures to just the core substructures. Moreover, to ensure the extracted substructures are concise and compact, we propose the Interactive Graph Information Bottleneck (IGIB) theory, which focuses on capturing the most influential yet minimal interactive substructures. In summary, our approach, guided by the IGIB theory, achieves precise substructure extraction within the ISE framework and is encapsulated in the IGIB-ISE} Extensive experiments validate the superiority of our model over state-of-the-art baselines across various tasks in terms of accuracy, generalizability, and interpretability.
Junfeng Fang, Xuqiang Li, Hongxin Xiang, Alan Xia, Wenjie Du 0003, Yang Wang 0015
ICLR7
2025 Enhancing Graph Invariant Learning from a Negative Inference Perspective
abstract
The out-of-distribution (OOD) generalization challenge is a longstanding problem in graph learning. Through studying the fundamental cause of data distribution shift, i.e., the changes of environments, significant progress has been achieved in addressing this issue. However, we observe that existing works still fail to effectively address complex environment shifts. Existing practices place excessive attention on extracting causal subgraphs, inevitably treating spurious subgraphs as environment variables. While spurious subgraphs are controlled by environments, the space of environment changes encompass more than the scale of spurious subgraphs. Therefore, existing efforts have a limited inference space for environments, leading to failure under severe environment changes. To tackle this issue, we propose a negative inference graph OOD framework (NeGo) to broaden the inference space for environment factors. Inspired by the successful practice of prompt learning in capturing underlying semantics and causal associations in large language models, we design a negative prompt environment inference to extract underlying environment information. We further introduce the environment-enhanced invariant subgraph learning to effectively exploit inferred environment embedding, ensuring the robust extraction of causal subgraph in the environment shifts. Lastly, we conduct a comprehensive evaluation of NeGo on real-world datasets and synthetic datasets across domains. NeGo outperforms baselines on nearly all datasets, which verify the effectiveness of our framework.
Kuo Yang 0002, Zhengyang Zhou, Qihe Huang, Wenjie Du 0003, Wu Jiang, Yang Wang 0015
ICML4
2025 DO-CoLM: Dynamic 3D Conformation Relationships Capture with Self-Adaptive Ordering Molecular Relational Modeling in Language Models
abstract
Molecular Relational Learning (MRL) aims to understand interactions between molecular pairs, playing a critical role in advancing biochemical research. Recently, Large Language Models (LLMs), with their extensive knowledge bases and advanced reasoning capabilities, have emerged as powerful tools for MRL. However, existing LLMs, which primarily rely on SMILES strings and molecular graphs, face two major challenges. They struggle to capture molecular stereochemistry and dynamics, as molecules possess multiple 3D conformations with varying reactivity and dynamic transformation relationships that are essential for accurately predicting molecular interactions but cannot be effectively represented by 1D SMILES or 2D molecular graphs. Additionally, these models do not consider the autoregressive nature of LLMs, overlooking the impact of input order on model performance. To address these issues, we propose DO-CoLM: a Dynamic relationship capture and self-adaptive Ordering 3D molecular Conformation LM for MRL. By introducing modules to dynamically model intra-molecular and inter-molecular conformational relationships and adaptively adjust the molecular modality input order, DO-CoLM achieves superior performance, as demonstrated by experimental results on 12 cross-domain datasets.
Hongxin Xiang, Jianmin Wang 0016, Wenjie Du 0003, Yang Wang 0015
IJCAI6
2025 MTGIB-UNet: A Multi-Task Graph Information Bottleneck and Uncertainty Weighted Network for ADMET Prediction
abstract
Accurate prediction of ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) properties is crucial in drug development, as these properties directly impact a drug's efficacy and safety. However, existing multi-task learning models often face challenges related to noise interference and task conflicts when dealing with complex molecular structures. To address these issues, we propose a novel multi-task Graph Neural Network (GNN) model, \textbf{MTGIB-UNet}. The model begins by encoding molecular graphs to capture intricate molecular structure information. Subsequently, based on the Graph Information Bottleneck (GIB) principle, the model compresses the information flow by extracting subgraphs, retaining task-relevant features while removing noise for each task. These embeddings are then fused through a gated network that dynamically adjusts the contribution weights of auxiliary tasks to the primary task. Specifically, an uncertainty weighting (UW) strategy is applied, with additional emphasis placed on the primary task, allowing dynamic adjustment of task weights while strengthening the influence of the primary task on model training. Experiments on standard ADMET datasets demonstrate that our model outperforms existing methods. Additionally, the model shows good interpretability by identifying key molecular substructures related to specific ADMET endpoints.
Xuqiang Li, Wenjie Du 0003, Jun Xia 0001, Jianmin Wang 0016, Yang Wang 0015
IJCAI2
2025 Electron Density-enhanced Molecular Geometry Learning
abstract
Electron density (ED), which describes the probability distribution of electrons in space, is crucial for accurately understanding the energy and force distribution in molecular force fields (MFF). Existing machine learning force fields (MLFF) focus on mining appropriate physical quantities from the atom-level conformation to enhance the molecular geometry representation while ignoring the unique information from microscopic electrons. In this work, we propose an efficient Electronic Density representation framework to enhance molecular Geometric learning (called EDG), which leverages images rendered from ED to boost molecular geometric representations in MLFF. Specifically, we construct a novel image-based ED representation, which consists of 2 million 6-view images with RGB-D channels, and design an ED representation learning model, called ImageED, to learn ED-related knowledge from these images. We further propose an efficient ED-aware teacher and introduce a cross-modal distillation strategy to transfer knowledge from the image-based teacher to the geometry-based students. Extensive experiments on QM9 and rMD17 demonstrate that EDG can be directly integrated into existing geometry-based models and significantly improves the capabilities of these models (e.g., SchNet, EGNN, SphereNet, ViSNet) for geometry representation learning in MLFF with a maximum average performance increase of 33.7%. Code and appendix are available at https://github.com/HongxinXiang/EDG
Hongxin Xiang, Jun Xia 0001, Xin Jin 0014, Wenjie Du 0003, Xiangxiang Zeng
IJCAI4
2025 ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models
abstract
Molecular Relational Learning (MRL) aims to understand interactions between molecular pairs, playing a critical role in advancing biochemical research. With the recent development of large language models (LLMs), a growing number of studies have explored the integration of MRL with LLMs and achieved promising results. However, the increasing availability of diverse LLMs and molecular structure encoders has significantly expanded the model space, presenting major challenges for benchmarking. Currently, there is no LLM framework that supports both flexible molecular input formats and dynamic architectural switching. To address these challenges, reduce redundant coding, and ensure fair model comparison, we propose ModuLM, a framework designed to support flexible LLM-based model construction and diverse molecular representations. ModuLM provides a rich suite of modular components, including 8 types of 2D molecular graph encoders, 11 types of 3D molecular conformation encoders, 7 types of interaction layers, and 7 mainstream LLM backbones. Owing to its highly flexible model assembly mechanism, ModuLM enables the dynamic construction of over 50,000 distinct model configurations. In addition, we provide comprehensive benchmark results to demonstrate the effectiveness of ModuLM in supporting LLM-based MRL tasks.
Yizhen Zheng, Huan Yee Koh, Hongxin Xiang, Linjiang Chen, Wenjie Du 0003, Yang Wang 0015
NeurIPS6
2025 Integrating Drug Substructures and Longitudinal Electronic Health Records for Personalized Drug Recommendation
abstract
Drug recommendation systems aim to identify optimal drug combinations for patient care, balancing therapeutic efficacy and safety. Advances in large-scale longitudinal EHRs have enabled learning-based approaches that leverage patient histories such as diagnoses, procedures, and previously prescribed drugs, to model complex patient-drug relationships. Yet, many existing solutions overlook standard clinical practices that favor certain drugs for specific conditions and fail to fully integrate the influence of molecular substructures on drug efficacy and safety. In response, we propose \textbf{SubRec}, a unified framework that integrates representation learning across both patient and drug spaces. Specifically, SubRec introduces a conditional information bottleneck to extract core drug substructures most relevant to patient conditions, thereby enhancing interpretability and clinical alignment. Meanwhile, an adaptive vector quantization mechanism is designed to generate patient–drug interaction patterns into a condition-aware codebook which reuses clinically meaningful patterns, reduces training overhead, and provides a controllable latent space for recommendation. Crucially, the synergy between condition-specific substructure learning and discrete patient prototypes allows SubRec to make accurate and personalized drug recommendations. Experimental results on the real-world MIMIC III and IV demonstrate our model's advantages. The source code is available at \href{https://anonymous.4open.science/r/DrugRecommendation-5173}{https://anonymous.4open.science/}.
Wenjie Du 0003, Xuqiang Li, Jinke Feng, Yang Wang 0015
NeurIPS1
2025 Bridging the Gap Between Cross-Domain Theory and Practical Application: A Case Study on Molecular Dissolution
abstract
Artificial intelligence (AI) has played a transformative role in chemical research, greatly facilitating the prediction of small molecule properties, simulation of catalytic processes, and material design. These advances are driven by increases in computing power, open source machine learning frameworks, and extensive chemical datasets. However, a persistent challenge is the limited amount of high-quality real-world data, while models calculated based on large amounts of theoretical data are often costly and difficult to deploy, which hinders the applicability of AI models in real-world scenarios. In this study, we enhance the prediction of solute-solvent properties by proposing a novel sample selection method: the iterative core subset extraction (CSIE) framework. CSIE iteratively updates the core sample subset based on information gain to remove redundant features in theoretical data and optimize the performance of the model on real chemical datasets. Furthermore, we introduce an asymmetric molecular interaction graph neural network (AMGNN) that combines positional information and bidirectional edge connections to simulate real-world chemical reaction scenarios to better capture solute-solvent interactions. Experimental results show that our method can accurately extract the core subset and improve the prediction accuracy.
Wenjie Du 0003, Yang Wang 0015
NeurIPS2
2025 EDBench: Large-Scale Electron Density Data for Molecular Modeling
abstract
Existing molecular machine learning force fields (MLFFs) generally focus on the learning of atoms, molecules, and simple quantum chemical properties (such as energy and force), but ignore the importance of electron density (ED) $\rho(r)$ in accurately understanding molecular force fields (MFFs). ED describes the probability of finding electrons at specific locations around atoms or molecules, which uniquely determines all ground state properties (such as energy, molecular structure, etc.) of interactive multi-particle systems according to the Hohenberg-Kohn theorem. However, the calculation of ED relies on the time-consuming first-principles density functional theory (DFT), which leads to the lack of large-scale ED data and limits its application in MLFFs. In this paper, we introduce EDBench, a large-scale, high-quality dataset of ED designed to advance learning-based research at the electronic scale. Built upon the PCQM4Mv2, EDBench provides accurate ED data, covering 3.3 million molecules. To comprehensively evaluate the ability of models to understand and utilize electronic information, we design a suite of ED-centric benchmark tasks spanning prediction, retrieval, and generation. Our evaluation of several state-of-the-art methods demonstrates that learning from EDBench is not only feasible but also achieves high accuracy. Moreover, we show that learning-based methods can efficiently calculate ED with comparable precision while significantly reducing the computational cost relative to traditional DFT calculations. All data and benchmarks from EDBench will be freely available, laying a robust foundation for ED-driven drug discovery and materials science.
Hongxin Xiang, Mingquan Liu, Zhixiang Cheng, Wenjie Du 0003, Jun Xia 0001, Xin Jin 0014, Xiangxiang Zeng
NeurIPS6
2025 Dynamic and Chemical Constraints to Enhance the Molecular Masked Graph Autoencoders
abstract
Masked Graph Autoencoders (MGAEs) have gained significant attention recently. Their proxy tasks typically involve random corruption of input graphs followed by reconstruction. However, in the molecular domain, two main issues arise: the predetermined mask ratio and reconstruction objectives can lead to suboptimal performance or negative transfer due to overly simplified or complex tasks, and these tasks may deviate from chemical priors. To tackle these challenges, we propose Dynamic and Chemical Constraints (DyCC) for MGAEs. This includes a masking strategy called GIBMS, which preserves essential semantic information during graph masking while adaptively adjusting the mask ratio and content for each molecule. Additionally, we introduce a Soft Label Generator (SLG) that reconstructs masked tokens as learnable prototypes (soft labels) rather than hard labels. These components adhere to chemical constraints and allow dynamic variation of proxy tasks during training. We integrate the model-agnostic DyCC into various MGAEs and conduct comprehensive experiments, demonstrating significant performance improvements. Our code is available at \url{https://github.com/forever-ly/DyCC}.
Wenjie Du 0003, Yang Wang 0015
NeurIPS2
2025 Enhancing the Maximum Effective Window for Long-Term Time Series Forecasting
abstract
Long-term time series forecasting (LTSF) aims to predict future trends based on historical data. While longer lookback windows theoretically offer more comprehensive insights, Transformer-based models often struggle with them. On one hand, longer windows introduce more noise and redundancy, hindering the model's learning process. On the other hand, Transformers suffer from attention dispersion and are prone to overfitting to noise, especially when processing long sequences. In this paper, we introduce the Maximum Effective Window (MEW) metric to assess a model's ability to effectively utilize the lookback window. We also propose two model-agnostic modules to enhance MEW, enabling models to better leverage historical data for improved performance. Specifically, to reduce redundancy and noise, we introduce the Information Bottleneck Filter (IBF), which employs information bottleneck theory to extract the most essential subsequences from the input. Additionally, we propose the Hybrid-Transformer-Mamba (HTM), which incorporates the Mamba mechanism for selective forgetting of long sequences while harnessing the Transformer's strong modeling capabilities for shorter sequences. We integrate these two modules into various Transformer-based models, and experimental results show that they effectively enhance MEW, leading to improved overall performance. Our code is available at \url{https://github.com/forever-ly/PIH}.
Zhengyang Zhou, Wenjie Du 0003, Yang Wang 0015
NeurIPS3
2025 An image-based protein-ligand binding representation learning framework via multi-level flexible dynamics trajectory pre-training
abstract
MOTIVATION: Accurate prediction of protein-ligand binding (PLB) relationships plays a crucial role in drug discovery, which helps identify drugs that modulate the activity of specific targets. Traditional biological assays for measuring PLB relationships are time consuming and costly. In addition, models for predicting PLB relationships have been developed and widely used in drug discovery tasks. However, learning more accurate PLB representations is essential to meet the stringent standards required for drug discovery. RESULTS: We propose an image-based PLB representation learning framework, called ImagePLB, which equips ligand representation learner (LRL) and protein representation learner (PRL) to accept 3D multi-view ligand images and protein graphs as input, respectively, and learns rich interaction information between ligand and protein through a binding representation learner (BRL). Considering the scarcity of protein-ligand pairs, we further propose a multi-level next trajectory prediction (MLNTP) task to pre-train ImagePLB on the 4D flexible dynamics trajectory of 16 972 complexes, including ligand level, protein level, and complex level, to learn information related to trajectories. Besides, by introducing trajectory regularization (TR), we effectively alleviate the problem of high (even almost identical) feature similarity caused by adjacent trajectories. Compared with the current state-of-the-art methods, ImagePLB has achieved competitive improvements on PLB-related prediction tasks, including protein-ligand affinity and efficacy prediction tasks. This study opens the door to the image-based PLB learning paradigm. AVAILABILITY AND IMPLEMENTATION: All data and implementation details of code can be obtained from https://github.com/HongxinXiang/ImagePLB.
Hongxin Xiang, Mingquan Liu, Linlin Hou, Shuting Jin, Jianmin Wang 0016, Jun Xia 0001, Wenjie Du 0003, Sisi Yuan, Xiangzheng Fu, Lei Xu 0047
Bioinform.7
2025 Soft causal learning for generalized molecule property prediction: An environment modeling perspective
Zhengyang Zhou, Kuo Yang 0002, Wenjie Du 0003, Pengkun Wang 0001, Yang Wang 0015
Knowl. Inf. Syst.4
2025 TRGH-PPI: Effective and Generalized Prediction of Protein-Protein Interactions Through Transformer and Graph
abstract
Protein-protein interaction (PPI) is a fundamental means of function and signaling in biological systems. The significant increase in demand and cost associated with experimental PPI research requires computational tools to automatically predict and understand PPI. However, existing methods either heavily rely on protein sequences for PPI prediction, or focus on protein structure based on the belief that structure is the key determinant of interactions. But in fact, both have a significant impact on the function of proteins. Therefore, we propose an integrated framework TRGH-PPI based on Transformer and GNN, including a protein feature extraction module and a PPI prediction module, which can simultaneously model both types of protein information and predict PPI. In the protein feature extraction module, we repeatedly use Transformer and GCN to iteratively update the sequence representation and structural features of proteins. The combination of Transformer and GCN enables them to leverage their respective advantages, promote model innovation, and improve the efficiency of graph data processing. In the PPI prediction module, we propose dpdGAT, which uses dot product operations more suitable for PPI prediction and has a dynamic attention mechanism. Numerous experiments have shown that TRGH-PPI is superior to current advanced methods and has shown high accuracy and generalization ability in predicting PPI. By integrating the synergistic modeling of sequence and structural features, the model has achieved an average improvement of 2% -5% in key indicators compared to the SOTA on three datasets.
Xuqiang Li, Jianmin Wang 0016, Wenjie Du 0003, Yang Wang 0015
IEEE Trans. Comput. Biol. Bioinform.6
2025 IIB-DDI: Invariant Information Bottleneck Theory for Out-of-Distribution Drug-Drug Interaction Prediction
abstract
Rapid and accurate identification of drug-drug interactions (DDIs) among multiple medications is crucial for various medical treatments, and clinical therapies. Currently, the increasing significance of molecular substructure interactions in DDI prediction has become a consensus. However, due to the uneven distribution of substructures in molecules and most methods do not consider the differences in substructure distributions, models may rely on spurious substructure relationships, thereby weakening their out-of-distribution (OOD) generalization ability. To address this challenge, we propose an OOD-DDI framework called IIB-DDI. Specifically, information bottleneck theory is initially utilized to extract core subgraphs of drug pairs. Then, considering the diversity and unknown nature of environments, we introduce the vector quantization to design an environment codebook, where the potential environments in the dataset are clustered into a specified number of categories. Subsequently, we position the extracted core subgraphs under various latent environmental factors to attain invariant core substructures (rationales). Additionally, the learned environmental distribution also could be acted as noise injection for optimizing mutual information, achieving a smoother and more stable training curve, thereby leading to lower loss. Extensive experiments conducted on real-world DDI datasets demonstrate the superiority of our model over state-of-the-art baselines.
Xuqiang Li, Di Wu 0057, Wenjie Du 0003, Yang Wang 0015
IEEE Trans. Comput. Biol. Bioinform.7
2024 MolCLW: Molecular Contrastive Learning With Learnable Weighted Substructures
abstract
Predicting molecular properties is crucial across scientific and industrial domains such as drug discovery and material science. Contrastive learning has gained prominence as an effective method. However, routine strategies often overlook substructure information and run the risk of guiding the model to learn incorrect knowledge due to molecular property changes during augmentation. To address these issues, we propose a novel molecular contrastive learning model named Molecular Contrastive Learning with Learnable Weighted Substructures (MolCLW). Our model aims to enhance the learning of molecular substructures by comparing molecular representations generated at the molecular-level and substructure-level. Meanwhile, a transformer encoder module is dedicated to learning the weight scores of substructures, enabling the model to adjust the similarity level between positive pairs based on the varying masked substructures. In comparison to baselines, MolCLW demonstrates an averaged 2.2% improvement in ROC-AUC on 6 classification benchmarks and an averaged 4.3% decrease in error on 5 regression benchmarks. Furthermore, during the finetuning process, the model can attend to the substructure relevant to the downstream task by adjusting the weights of different substructures. The visualization results establish a link between downstream targets and substructure function, providing valuable insights and strong support for drug discovery, chemical reactions, and research.
Jiahe Li 0011, Wenjie Du 0003, Yang Wang 0015
BIBM2
2024 EMoNet: An environment causal learning for molecule OOD generalization
abstract
Data-driven molecular computing has become an increasingly popular topic in AI for molecular and bioinformatic science. Current molecular modeling exploits the Graph Neural Networks (GNNs) to achieve the representation but they mostly fail to generalize to out-of-distribution (OOD) samples. Even though recent advances on OOD-oriented graph learning discovered the invariant rationale on graphs, they still ignore two important issues, i.e., 1) the increasing number and types of molecules expand patterns of environments on graphs, resulting in failures of invariant rationale based models, 2) the associations between discovered molecular subgraphs and corresponding properties are complex where causal substructures cannot fully interpret the labels. To this end, we propose an environment causal learning framework, EMoNet, to tackle the unresolved OOD challenge in molecular science. Specifically, we model the graph environments via bypassing invariant subgraphs. We first incorporate chemistry principle into our graph growth generator to imitate environment growth, and then devise an environment-GIB to squash out environment and finally introduce a cross-attention causal aggregation, allowing dynamic interactions between environments and invariances. We perform experiments on seven datasets and extensive experiments demonstrate strong generalization ability of EMoNet.
Kuo Yang 0002, Wenjie Du 0003, Zhongchao Yi, Zhengyang Zhou, Yang Wang 0015
BIBM3
2024 MMGNN: A Molecular Merged Graph Neural Network for Explainable Solvation Free Energy Prediction
Wenjie Du 0003, Di Wu 0057, Jun Xia 0001, Ziyuan Zhao, Junfeng Fang, Yang Wang 0015
IJCAI1
2024 AdaNovo: Towards Robust \emph{De Novo} Peptide Sequencing in Proteomics against Data Biases
abstract
Tandem mass spectrometry has played a pivotal role in advancing proteomics, enabling the high-throughput analysis of protein composition in biological tissues. Despite the development of several deep learning methods for predicting amino acid sequences (peptides) responsible for generating the observed mass spectra, training data biases hinder further advancements of \emph{de novo} peptide sequencing. Firstly, prior methods struggle to identify amino acids with Post-Translational Modifications (PTMs) due to their lower frequency in training data compared to canonical amino acids, further resulting in unsatisfactory peptide sequencing performance. Secondly, various noise and missing peaks in mass spectra reduce the reliability of training data (Peptide-Spectrum Matches, PSMs). To address these challenges, we propose AdaNovo, a novel and domain knowledge-inspired framework that calculates Conditional Mutual Information (CMI) between the mass spectra and amino acids or peptides, using CMI for robust training against above biases. Extensive experiments indicate that AdaNovo outperforms previous competitors on the widely-used 9-species benchmark, meanwhile yielding 3.6\% - 9.4\% improvements in PTMs identification. The supplements contain the code.
Jun Xia 0001, Shaorong Chen, Xiaojun Shan, Wenjie Du 0003, Zhangyang Gao, Cheng Tan 0012, Bozhen Hu, Jiangbin Zheng 0002, Stan Z. Li
NeurIPS5
2024 FlexMol: A Flexible Toolkit for Benchmarking Molecular Relational Learning
abstract
Molecular relational learning (MRL) is crucial for understanding the interaction behaviors between molecular pairs, a critical aspect of drug discovery and development. However, the large feasible model space of MRL poses significant challenges to benchmarking, and existing MRL frameworks face limitations in flexibility and scope. To address these challenges, avoid repetitive coding efforts, and ensure fair comparison of models, we introduce FlexMol, a comprehensive toolkit designed to facilitate the construction and evaluation of diverse model architectures across various datasets and performance metrics. FlexMol offers a robust suite of preset model components, including 16 drug encoders, 13 protein sequence encoders, 9 protein structure encoders, and 7 interaction layers. With its easy-to-use API and flexibility, FlexMol supports the dynamic construction of over 70, 000 distinct combinations of model architectures. Additionally, we provide detailed benchmark results and code examples to demonstrate FlexMol’s effectiveness in simplifying and standardizing MRL model development and comparison. FlexMol is open-sourced and available at https://github.com/Steven51516/FlexMol.
Sizhe Liu, Jun Xia 0001, Lecheng Zhang, Yue Liu 0008, Wenjie Du 0003, Zhangyang Gao, Bozhen Hu, Cheng Tan 0012, Hongxin Xiang, Stan Z. Li
NeurIPS6
2024 NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics
abstract
Tandem mass spectrometry has played a pivotal role in advancing proteomics, enabling the analysis of protein composition in biological tissues. Many deep learning methods have been developed for \emph{de novo} peptide sequencing task, i.e., predicting the peptide sequence for the observed mass spectrum. However, two key challenges seriously hinder the further research of this important task. Firstly, since there is no consensus for the evaluation datasets, the empirical results in different research papers are often not comparable, leading to unfair comparison. Secondly, the current methods are usually limited to amino acid-level or peptide-level precision and recall metrics. In this work, we present the first unified benchmark NovoBench for \emph{de novo} peptide sequencing, which comprises diverse mass spectrum data, integrated models, and comprehensive evaluation metrics. Recent impressive methods, including DeepNovo, PointNovo, Casanovo, InstaNovo, AdaNovo and $\pi$-HelixNovo are integrated into our framework. In addition to amino acid-level and peptide-level precision and recall, we also evaluate the models' performance in terms of identifying post-tranlational modifications (PTMs), efficiency and robustness to peptide length, noise peaks and missing fragment ratio, which are important influencing factors while seldom be considered. Leveraging this benchmark, we conduct a large-scale study of current methods, report many insightful findings that open up new possibilities for future development. The benchmark is open-sourced to facilitate future research and application. The code is available at \url{https://github.com/Westlake-OmicsAI/NovoBench}.
Shaorong Chen, Jun Xia 0001, Sizhe Liu, Tianze Ling, Wenjie Du 0003, Yue Liu 0008, Jianwei Yin, Stan Z. Li
NeurIPS6
2023 Improving efficiency in rationale discovery for Out-of-Distribution molecular representations
abstract
Graph Neural Networks (GNNs), as the dominant approach in Molecular Representation Learning (MRL), have exhibited remarkable efficacy in diverse tasks, including molecular property prediction and drug discovery. Considering the dynamic and diverse nature of test molecules in real-world contexts, the previous independently and identically distributed (i.i.d.) assumption for training and test molecules is not align with the requirements of practical applications in chemistry. Rationalization has been proposed to enhance the Out-of-Distribution (OOD) generalization of GNN, but a critical concern remains unexplored: low efficiency in rationale discovery. We consider this bottleneck is due to large search space and overly flexible modeling. Here, we introduce a framework called Molecule Stratifier and Invariant Rational Improver (MSIRI) to overcome the challenge. Specifically, MSIRI adopt vector quantization to obtain invariant rationales and the search space is narrowed by substructure and two assumptions based on subgraph matching. Experimental results on ten datasets demonstrate that our approach both achieves significant improvements across various GNN backbones and outperforms other six Out-of-Distribution method, clearly demonstrating its effectiveness in addressing the OOD challenges in MRL.
Jiahe Li 0011, Wenjie Du 0003, Di Wu 0057, Yang Wang 0015
BIBM3
2023 Fusing 2D and 3D molecular graphs as unambiguous molecular descriptors for conformational and chiral stereoisomers
abstract
The rapid progress of machine learning (ML) in predicting molecular properties enables high-precision predictions being routinely achieved. However, many ML models, such as conventional molecular graph, cannot differentiate stereoisomers of certain types, particularly conformational and chiral ones that share the same bonding connectivity but differ in spatial arrangement. Here, we designed a hybrid molecular graph network, Chemical Feature Fusion Network (CFFN), to address the issue by integrating planar and stereo information of molecules in an interweaved fashion. The three-dimensional (3D, i.e., stereo) modality guarantees precision and completeness by providing unabridged information, while the two-dimensional (2D, i.e., planar) modality brings in chemical intuitions as prior knowledge for guidance. The zipper-like arrangement of 2D and 3D information processing promotes cooperativity between them, and their synergy is the key to our model's success. Experiments on various molecules or conformational datasets including a special newly created chiral molecule dataset comprised of various configurations and conformations demonstrate the superior performance of CFFN. The advantage of CFFN is even more significant in datasets made of small samples. Ablation experiments confirm that fusing 2D and 3D molecular graphs as unambiguous molecular descriptors can not only effectively distinguish molecules and their conformations, but also achieve more accurate and robust prediction of quantum chemical properties.
Wenjie Du 0003, Xiaoting Yang, Di Wu 0057, Fenfen Ma, Baicheng Zhang, Chaochao Bao, Yaoyuan Huo, Yang Wang 0015
Briefings Bioinform.1