VLDB 2026 Research / reviewers in the wild / expert
Jiancong Xie
dblp:46/2144
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0001-9030-0709ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | De Novo Molecular Generation from Mass Spectra via Many-Body Enhanced DiffusionabstractMolecular structure generation from mass spectrometry is fundamental for understanding cellular metabolism and discovering novel compounds. Although tandem mass spectrometry (MS/MS) enables the high-throughput acquisition of fragment fingerprints, these spectra often reflect higher-order interactions involving the concerted cleavage of multiple atoms and bonds-crucial for resolving complex isomers and non-local fragmentation mechanisms. However, most existing methods adopt atom-centric and pairwise interaction modeling, overlooking higher-order edge interactions and lacking the capacity to systematically capture essential many-body characteristics for structure generation. To overcome these limitations, we present MBGen, a Many-Body enhanced diffusion framework for de novo molecular structure Generation from mass spectra. By integrating a many-body attention mechanism and higher-order edge modeling, MBGen comprehensively leverages the rich structural information encoded in MS/MS spectra, enabling accurate de novo generation and isomer differentiation for novel molecules. Experimental results on the NPLIB1 and MassSpecGym benchmarks demonstrate that MBGen achieves superior performance, with improvements of up to 230% over state-of-the-art methods, highlighting the scientific value and practical utility of many-body modeling for mass spectrometry-based molecular generation. Further analysis and ablation studies show that our approach effectively captures higher-order interactions and exhibits enhanced sensitivity to complex isomeric and non-local fragmentation information. Xichen Sun, Jiahua Rao, Jiancong Xie, Yuedong Yang |
AAAI | 4 |
| 2026 | Informative Subgraph Extraction with Deep Reinforcement Learning for Drug-Drug Interaction PredictionabstractDrug-drug interaction (DDI) prediction is pivotal for drug safety and clinical decision-making. Recently, subgraph-based methods utilizing knowledge graphs (KGs) and domain information have achieved promising results by extracting informative subgraphs for DDI prediction. However, existing subgraph extraction methods are typically coarse-grained and nonspecific, facing two key limitations: First, they are constrained by the vast and noisy nature of real-world KGs, making it challenging to identify the most informative substructures from the massive space of candidate subgraphs. Second, current methods often fail to exploit the molecular structural specificity of drugs to selectively extract relevant subgraphs, lacking effective integration of molecular structure information with knowledge graph context. To address these challenges, we propose RISE-DDI, a novel framework for Reinforced-based Informative Subgraph Extraction approach for drug-drug interaction prediction. Specifically, RISE-DDI formulates the subgraph extraction as a Markov Decision Process (MDP) and leverages a deep reinforcement learning (RL) agent to dynamically and adaptively extract the most informative and context-specific subgraphs for each drug pair. The agent is guided by a learnable structure-aware reward model that considers both the topological context from the knowledge graph and the molecular features of the drug pairs, thereby encouraging the selection of subgraphs that are both structurally relevant and biologically informative. Extensive experiments on DDI benchmark datasets demonstrate that our method outperforms state-of-the-art baselines in both transductive and inductive scenarios, achieving improvements of up to 20%. Furthermore, visualization analyses of the extracted subgraphs highlight the interpretability of our model, providing insights into the underlying mechanisms of drug interactions. Jiancong Xie, Jiahua Rao, Yuedong Yang |
AAAI | 1 |
| 2025 | Interpretably Predicting Chemical Perturbation via Biologically Informed Visible Neural NetworkabstractRecent advances in single-cell transcriptomics have enabled high-resolution characterization of cellular responses to drug perturbations, a critical capability for precision medicine and drug discovery. Deep learning has emerged as a powerful tool for accurate and scalable prediction of single-cell responses to drug perturbations. However, most existing approaches fail to consider gene-level structural priors and biological knowledge. Furthermore, they often operate as black boxes with limited interpretability. To offer reliable drug response prediction in real-world applications, there is an urgent need to develop a model that combines high predictive accuracy with strong interpretability. In this work, we propose a novel framework based on the Visible Neural Network (VNN) to predict cellular responses to drug perturbations in single-cell transcriptomic profiles. In contrast to traditional black-box models, VNN incorporates prior biological knowledge into its architecture. This integration enables interpretable predictions while maintaining biological coherence. We further integrate deep learning models with a knowledge graph of gene-gene relationships to improve the understanding and generalization of transcriptomic responses to unseen drug perturbations by leveraging prior biological knowledge of gene interactions. Our analysis reveals that the hierarchical structure of the VNN aligns well with known biological processes and supports robust feature attribution. In general, our findings underscore the potential of visible architectures to bridge deep learning and systems biology for a scalable and interpretable modeling of cellular drug responses. Jiancong Xie, Yuedong Yang |
BIBM | 2 |
| 2025 | Multi-modal Contrastive Learning with Negative Sampling Calibration for Phenotypic Drug DiscoveryabstractPhenotypic drug discovery presents a promising strategy for identifying first-in-class drugs by bypassing the need for specific drug targets. Recent advances in cell-based phenotypic screening tools, including Cell Painting and the LINCS L1000, provide essential cellular data that capture biological responses to compounds. While the integration of the multi-modal data enhances the use of contrastive learning (CL) methods for molecular phenotypic representation, these approaches treat all negative pairs equally, failing to discriminate molecules with similar phenotypes. To address these challenges, we introduce a foundational framework MINER that dynamically estimates the likelihoods of sample pairs as negative pairs based on uni-modal disentangled representations. In addition, our approach incorporates a mixture fusion strategy to effectively integrate multimodal data, even in cases where certain modalities are missing. Extensive experiments demonstrate that our method enhances both molecular property prediction and molecule-phenotype retrieval accuracy. Moreover, it successfully recommends drug candidates from phenotype for complex diseases documented in the literature. These findings underscore MINER’s potential to advance drug discovery by enabling deeper insights into disease mechanisms and improving drug candidate recommendations. Jiahua Rao, Hanjing Lin, Leyu Chen, Jiancong Xie, Shuangjia Zheng, Yuedong Yang |
CVPR | 4 |
| 2025 | Incorporating Retrieval-based Causal Learning with Information Bottlenecks for Interpretable Molecular Graph LearningabstractGraph Neural Networks (GNNs) have gained considerable traction for modeling molecular structures and predicting properties, but their interpretability remains a significant challenge in understanding chemical behaviors. Current interpretation methods often rely on post-hoc explanations, which aim to provide transparency in GNN decisions. However, these approaches struggle with interpreting complex subgraphs and fail to leverage explanations to enhance predictive capabilities. While transparent methods can enhance GNN predictions, they typically compromise on explanation precision. This limitation underscores the need for a new strategy that effectively integrates GNN explanations and predictions. In this study, we have developed a novel interpretable causal GNN framework that combines retrieval-based causal learning with Graph Information Bottleneck (GIB) theory. Our framework semi-parametrically identifies crucial subgraphs through GIB and compresses explanatory subgraphs using a causal module. The framework consistently outperformed state-of-the-art methods, achieving a 32.72% increase in precision for scientific explanation tasks involving diverse substructures. More importantly, the learned explanations were also shown to be able to improve GNN prediction performance. This advancement is particularly vital for molecular graph learning, as it addresses the critical need to interpret how molecular structures influence predicted properties, thereby aiding drug discovery and materials science by providing insights into chemical mechanisms. Jiahua Rao, Hanjing Lin, Jiancong Xie, Zhen Wang 0036, Shuangjia Zheng, Yuedong Yang |
KDD (2) | 3 |
| 2025 | Reinforced Active Learning for Large-Scale Virtual Screening with Learnable Policy ModelabstractVirtual Screening (VS) is vital for drug discovery but struggles with low hit rates and high computational costs. While Active Learning (AL) has shown promise in improving the efficiency of VS, traditional methods rely on inflexible and handcrafted heuristics, limiting adaptability in complex chemical spaces, particularly in balancing molecular diversity and selection accuracy.
To overcome these challenges, we propose GLARE, a reinforced active learning framework that reformulates VS as a Markov Decision Process (MDP). Using Group Relative Policy Optimization (GRPO), GLARE dynamically balances chemical diversity, biological relevance, and computational constraints, eliminating the need for inflexible heuristics.
Experiments show GLARE outperforms state-of-the-art AL methods, with a 64.8% average improvement in Enrichment Factors (EF). Additionally, GLARE enhances the performance of VS foundation models like DrugCLIP, achieving up to an 8-fold improvement in EF$_{0.5\\%}$ with as few as 15 active molecules. These results highlight the transformative potential of GLARE for adaptive and efficient drug discovery. Yicong Chen, Jiahua Rao, Jiancong Xie, Dahao Xu, Zhen Wang 0004, Yuedong Yang |
NeurIPS | 3 |
| 2025 | A 3D pocket-aware lead optimization model with knowledge guidance and its application for discovery of new glutaminyl cyclase inhibitorsabstractLead optimization, aimed at improving binding affinity or other properties of hit compounds, is a crucial task in drug discovery. Though deep learning-based 3D generative models showed promise in enhancing the efficiency of de novo drug design recently, less research and attention has garnered for structure-based lead optimization. Herein, we propose a 3D pocket-aware diffusion model named Diffleop, which explicitly incorporates the knowledge of protein-ligand binding affinity and information on covalent bonds to guide the denoising sampling process for lead optimization with enhanced binding affinity and rational properties. Specifically, the bond constraint is achieved through diffusion on fully connected molecular graphs, and the determination of atom positions, atom and bond types in each sampling step is guided by the gradient of the binding affinity that is predicted through fitting with an E(3)-equivariant expert network. The comprehensive evaluations indicated that Diffleop outperforms baseline models on lead optimization with higher affinity and more binding interactions, and can generate more drug-like molecules with more rational structures. Diffleop was further applied to optimize 5-methyl-1H-imidazole, our newly discovered lead compound targeting human glutaminyl cyclases (QCs). Three synthesized compounds exhibit substantially improved inhibitory activities against QCs, with the most effective one showing an IC50 value of 8 nM and 3.5-fold better than clinical candidate PQ912. Anjie Qiao, Weifeng Huang, Hao Zhang 0200, Qirui Deng, Jiahua Rao, Ji Deng, Zhen Wang 0004, Mingyuan Xu, Hongming Chen 0001, Jiancong Xie, Shuangjia Zheng, Yuedong Yang, Guo-Bo Li, Jinping Lei |
Briefings Bioinform. | 13 |
| 2024 | Interpretable Drug Response Prediction through Molecule Structure-aware and Knowledge-Guided Visible Neural NetworkabstractPrecise prediction of anti-cancer drug responses has become a crucial obstruction in anti-cancer drug design and clinical applications. In recent years, various deep learning methods have been applied to drug response prediction and become more accurate. However, they are still criticized as being non-transparent. To offer reliable drug response prediction in real-world applications, there is still a pressing demand to develop a model with high predictive performance as well as interpretability. In this study, we propose DrugVNN, an end-to-end interpretable drug response prediction framework, which extracts gene features of cell lines through a knowledge-guided visible neural network (VNN) and learns drug representation through a node-edge communicative message passing network (CMPNN). Additionally, between these two networks, a novel drug-aware gene attention gate is designed to direct the drug representation to VNN to simulate the effects of drugs. By evaluating on the GDSC dataset, DrugVNN achieved state-of-the-art performance. Moreover, DrugVNN can identify active genes and relevant signaling pathways for specific drug-cell line pairs with supporting evidence in the literature, implying the interpretability of our model. Jiancong Xie, Youyou Li, Jiahua Rao, Yuedong Yang |
BIBM | 1 |
| 2024 | From intuition to AI: evolution of small molecule representations in drug discoveryabstractWithin drug discovery, the goal of AI scientists and cheminformaticians is to help identify molecular starting points that will develop into safe and efficacious drugs while reducing costs, time and failure rates. To achieve this goal, it is crucial to represent molecules in a digital format that makes them machine-readable and facilitates the accurate prediction of properties that drive decision-making. Over the years, molecular representations have evolved from intuitive and human-readable formats to bespoke numerical descriptors and fingerprints, and now to learned representations that capture patterns and salient features across vast chemical spaces. Among these, sequence-based and graph-based representations of small molecules have become highly popular. However, each approach has strengths and weaknesses across dimensions such as generality, computational cost, inversibility for generative applications and interpretability, which can be critical in informing practitioners' decisions. As the drug discovery landscape evolves, opportunities for innovation continue to emerge. These include the creation of molecular representations for high-value, low-data regimes, the distillation of broader biological and chemical knowledge into novel learned representations and the modeling of up-and-coming therapeutic modalities. Miles McGibbon, Steven R. Shave, Yumiao Gao, Douglas R. Houston, Jiancong Xie, Yuedong Yang, Philippe Schwaller, Vincent Blay |
Briefings Bioinform. | 6 |
| 2023 | Fast and accurate protein function prediction from sequence through pretrained language model and homology-based label diffusionabstractProtein function prediction is an essential task in bioinformatics which benefits disease mechanism elucidation and drug target discovery. Due to the explosive growth of proteins in sequence databases and the diversity of their functions, it remains challenging to fast and accurately predict protein functions from sequences alone. Although many methods have integrated protein structures, biological networks or literature information to improve performance, these extra features are often unavailable for most proteins. Here, we propose SPROF-GO, a Sequence-based alignment-free PROtein Function predictor, which leverages a pretrained language model to efficiently extract informative sequence embeddings and employs self-attention pooling to focus on important residues. The prediction is further advanced by exploiting the homology information and accounting for the overlapping communities of proteins with related functions through the label diffusion algorithm. SPROF-GO was shown to surpass state-of-the-art sequence-based and even network-based approaches by more than 14.5, 27.3 and 10.1% in area under the precision-recall curve on the three sub-ontology test sets, respectively. Our method was also demonstrated to generalize well on non-homologous proteins and unseen species. Finally, visualization based on the attention mechanism indicated that SPROF-GO is able to capture sequence domains useful for function prediction. The datasets, source codes and trained models of SPROF-GO are available at https://github.com/biomed-AI/SPROF-GO. The SPROF-GO web server is freely available at http://bio-web1.nscc-gz.cn/app/sprof-go. Qianmu Yuan, Jiancong Xie, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 3 |