EDBT 2026 Demo / reviewers in the wild / expert
Jun Xia 0001
dblp:22/3650-1
· DBLP profile ↗
62ranked-venue papers
12as first author
54since 2021 · last 2026
0000-0002-7993-0803ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 7 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Regressor-guided Diffusion Model for De Novo Peptide Sequencing with Explicit Mass ControlabstractThe discovery of novel proteins relies on sensitive protein identification, for which de novo peptide sequencing (DNPS) from mass spectra is a crucial approach. While deep learning has advanced DNPS, existing models inadequately enforce the fundamental mass consistency constraint—that a predicted peptide's mass must match the experimental measured precursor mass. Previous DNPS methods often treat this critical information as a simple input feature or use it in post-processing, leading to numerous implausible predictions that do not adhere to this fundamental physical property. To address this limitation, we introduce DiffuNovo, a novel regressor-guided diffusion model for de novo peptide sequencing that provides explicit peptide-level mass control. Our approach integrates the mass constraint at two critical stages: during training, a novel peptide-level mass loss guides model optimization, while at inference, regressor-based guidance from gradient-based updates in the latent space steers the generation to compel the predicted peptide adheres to the mass constraint. Comprehensive evaluations on established benchmarks demonstrate that DiffuNovo surpasses state-of-the-art methods in DNPS accuracy. Additionally, as the first DNPS model to employ a diffusion model as its core backbone, DiffuNovo leverages the powerful controllability of diffusion architecture and achieves a significant reduction in mass error, thereby producing much more physically plausible peptides. These innovations represent a substantial advancement toward robust and broadly applicable DNPS. The source code is available in the supplementary material. Shaorong Chen, Jun Xia 0001 |
AAAI | 3 |
| 2026 | Departures: Distributional Transport for Single-Cell Perturbation Prediction with Neural Schrödinger BridgesabstractPredicting single-cell perturbation outcomes directly advances gene function analysis and facilitates drug candidate selection, making it a key driver of both basic and translational biomedical research. However, a major bottleneck in this task is the unpaired nature of single-cell data, as the same cell cannot be observed both before and after perturbation due to the destructive nature of sequencing. Although some neural generative transport models attempt to tackle unpaired single-cell perturbation data, they either lack explicit conditioning or depend on prior spaces for indirect distribution alignment, limiting precise perturbation modeling. In this work, we approximate Schrödinger Bridge (SB), which defines stochastic dynamic mappings recovering the entropy-regularized optimal transport (OT), to directly align the distributions of control and perturbed single-cell populations across different perturbation conditions. Unlike prior SB approximations that rely on bidirectional modeling to infer optimal source-target sample coupling, we leverage Minibatch-OT based pairing to avoid such bidirectional inference and the associated ill-posedness of defining the reverse process. This pairing directly guides bridge learning, yielding a scalable approximation to the SB. We approximate two SB models, one modeling discrete gene activation states and the other continuous expression distributions. Joint training enables accurate perturbation modeling and captures single-cell heterogeneity. Experiments on public genetic and drug perturbation datasets show that our model effectively captures heterogeneous single-cell responses and achieves state-of-the-art performance. Changxi Chi, Yufei Huang 0002, Jun Xia 0001, Jiangbin Zheng 0002, Yunfan Liu 0002, Zelin Zang, Stan Z. Li |
AAAI | 3 |
| 2026 | VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token AttentionabstractGraph Transformer has demonstrated impressive capabilities in the field of graph representation learning. However, existing approaches face two critical challenges: (1) most models suffer from exponentially increasing computational complexity, making it difficult to scale to large graphs; (2) attention mechanisms based on node-level operations limit the flexibility of the model and result in poor generalization performance in out-of-distribution (OOD) scenarios. To address these issues, we propose VecFormer (the Vec tor Quantized Graph Transformer ), an efficient and highly generalizable model for node classification, particularly under OOD settings. VecFormer adopts a two-stage training paradigm. In the first stage, two codebooks are used to reconstruct the node features and the graph structure, aiming to learn the rich semantic Graph Codes. In the second stage, attention mechanisms are performed at the Graph Token level based on the transformed cross codebook, reducing computational complexity while enhancing the model's generalization capability. Extensive experiments on datasets of various sizes demonstrate that VecFormer outperforms the existing Graph Transformer in both performance and speed. Jun Xia 0001, Siyuan Li 0002, Yunfan Liu 0002, Yufei Huang 0002, Changxi Chi, Mutian Hong, Zhuoli Ouyang, Chang Yu 0001, Stan Z. Li |
WWW | 2 |
| 2026 | A Survey of Deep Graph Clustering: Taxonomy, Challenge, Application, and Open ResourceabstractGraph clustering, which aims to divide nodes in the graph into several distinct clusters, is a fundamental yet challenging task. Benefiting from the powerful representation capability of deep learning, deep graph clustering methods have achieved great success in recent years. However, the corresponding survey paper is relatively scarce, and it is imminent to make a summary of this field. From this motivation, we conduct a comprehensive survey of deep graph clustering. Firstly, we introduce formulaic definition, evaluation, and development in this field. Secondly, the taxonomy of deep graph clustering methods is presented based on four different criteria, including graph type, network architecture, learning paradigm, and clustering method. Thirdly, we carefully analyze the existing methods via extensive experiments and summarize the challenges and opportunities from five perspectives, including graph data quality, stability, scalability, discriminative capability, and unknown cluster number. Besides, the applications of deep graph clustering methods in six domains, including computer vision, natural language processing, recommendation systems, social network analyses, bioinformatics, and medical science, are presented. Last but not least, this paper provides open resource supports, including 1) a collection (https://github.com/yueliu1999/Awesome-Deep-Graph-Clustering) of state-of-the-art deep graph clustering methods (papers, codes, and datasets) and 2) a flexible and extensible Python library (https://github.com/Marigoldwu/PyDGC) for deep graph clustering. We hope this work can serve as a quick guide and help researchers overcome challenges in this vibrant field. Yue Liu 0008, Jun Xia 0001, Benyu Wu, Sihang Zhou 0001, Xihong Yang, Ke Liang 0006, Guoxian Yu, Stan Z. Li, Xinwang Liu 0002, Kunlun He |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2026 | SKIP: A Prototype-Based Scalable Knowledge Graph Representation Learning MethodabstractThe field of knowledge graph representation learning (KGRL) has been rapidly expanding. To effectively apply KGRL models to large real-world knowledge graphs (KGs), anchor-based methods have been proposed. These methods aim to reduce computational costs and parameter requirements by encoding entities using a small set of entity anchors. However, existing anchor selection approaches are often rudimentary and sometimes yield suboptimal results. In this article, we propose a scalable anchor-based KGRL method called SKIP. By leveraging prototype information, our method selects representative entities as anchors. The SKIP method consists of two main steps. First, pretraining models are employed to encode entities by utilizing the topological structure and textual information in KGs. Second, the prototype learning module (PLM) extracts entity prototypes, which are then used to sample entity anchors that contain valuable prototype information. These settings enable SKIP to identify representative and reasonable entity anchors, leading to improved performance while requiring fewer computational resources. Extensive experiments conducted on various downstream tasks using KGs of different scales demonstrate the superiority and effectiveness of SKIP. Particularly, on the large OGB WikiKG 2 dataset, our method achieves comparable performance while reducing running time by approximately 21.28% and requiring 21.43% fewer model parameters compared to the baseline. This indicates the superior scalability of SKIP. Yue Liu 0008, Ke Liang 0006, Jun Xia 0001, Meng Liu 0014, Xihong Yang, Xinwang Liu 0002, Sihang Zhou 0001, Stan Z. Li |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | MeToken: Uniform Micro-environment Token Boosts Post-Translational Modification PredictionabstractPost-translational modifications (PTMs) profoundly expand the complexity and functionality of the proteome, regulating protein attributes and interactions that are crucial for biological processes. Accurately predicting PTM sites and their specific types is therefore essential for elucidating protein function and understanding disease mechanisms. Existing computational approaches predominantly focus on protein sequences to predict PTM sites, driven by the recognition of sequence-dependent motifs. However, these approaches often overlook protein structural contexts. In this work, we first compile a large-scale sequence-structure PTM dataset, which serves as the foundation for fair comparison. We introduce the MeToken model, which tokenizes the micro-environment of each amino acid, integrating both sequence and structural information into unified discrete tokens. This model not only captures the typical sequence motifs associated with PTMs but also leverages the spatial arrangements dictated by protein tertiary structures, thus providing a holistic view of the factors influencing PTM sites. Designed to address the long-tail distribution of PTM types, MeToken employs uniform sub-codebooks that ensure even the rarest PTMs are adequately represented and distinguished. We validate the effectiveness and generalizability of MeToken across multiple datasets, demonstrating its superior performance in accurately identifying PTM types. The results underscore the importance of incorporating structural data and highlight MeToken's potential in facilitating accurate and comprehensive PTM predictions, which could significantly impact proteomics research. Cheng Tan 0012, Zhenxiao Cao, Zhangyang Gao, Lirong Wu, Siyuan Li 0002, Yufei Huang 0002, Jun Xia 0001, Bozhen Hu, Stan Z. Li |
ICLR | 7 |
| 2025 | ReNovo: Retrieval-Based \emph{De Novo} Mass Spectrometry Peptide SequencingabstractProteomics is the large-scale study of proteins. Tandem mass spectrometry, as the only high-throughput technique for protein sequence identification, plays a pivotal role in proteomics research. One of the long-standing challenges in this field is peptide identification, which entails determining the specific peptide (sequence of amino acids) that corresponds to each observed mass spectrum. The conventional approach involves database searching, wherein the observed mass spectrum is scored against a pre-constructed peptide database. However, the reliance on pre-existing databases limits applicability in scenarios where the peptide is absent from existing databases. Such circumstances necessitate \emph{de novo} peptide sequencing, which derives peptide sequence solely from input mass spectrum, independent of any peptide database. Despite ongoing advancements in \emph{de novo} peptide sequencing, its performance still has considerable room for improvement, which limits its application in large-scale experiments. In this study, we introduce a novel \textbf{Re}trieval-based \emph{De \textbf{Novo}} peptide sequencing methodology, termed \textbf{ReNovo}, which draws inspiration from database search methods. Specifically, by constructing a datastore from training data, ReNovo can retrieve information from the datastore during the inference stage to conduct retrieval-based inference, thereby achieving improved performance. This innovative approach enables ReNovo to effectively combine the strengths of both methods: utilizing the assistance of the datastore while also being capable of predicting novel peptides that are not present in pre-existing databases. A series of experiments have confirmed that ReNovo outperforms state-of-the-art models across multiple widely-used datasets, incurring only minor storage and time consumption, representing a significant advancement in proteomics. Supplementary materials include the code. Shaorong Chen, Jun Xia 0001, Lecheng Zhang, Zhangyang Gao, Bozhen Hu, Cheng Tan 0012, Wenjie Du 0003, Stan Z. Li |
ICLR | 2 |
| 2025 | Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo
Jun Xia 0001, Sizhe Liu, Shaorong Chen, Hongxin Xiang, Zicheng Liu 0006, Yue Liu 0008, Stan Z. Li |
ICLR | 1 |
| 2025 | GRAPE: Heterogeneous Graph Representation Learning for Genetic Perturbation with Coding and Non-Coding BiotypeabstractPredicting genetic perturbations enables the identification of potentially crucial genes prior to wet-lab experiments, significantly improving overall experimental efficiency. Since genes are the foundation of cellular life, building gene regulatory networks (GRN) is essential to understand and predict the effects of genetic perturbations. However, current methods fail to fully leverage gene-related information, and solely rely on simple evaluation metrics to construct coarse-grained GRN. More importantly, they ignore functional differences between biotypes, limiting the ability to capture potential gene interactions. In this work, we leverage pre-trained large language model and DNA sequence model to extract features from gene descriptions and DNA sequence data, respectively, which serve as the initialization for gene representations. Additionally, we introduce gene biotype information for the first time in genetic perturbation, simulating the distinct roles of genes with different biotypes in regulating cellular processes, while capturing implicit gene relationships through graph structure learning (GSL). We propose GRAPE, a heterogeneous graph neural network (HGNN) that leverages gene representations initialized with features from descriptions and sequences, models the distinct roles of genes with different biotypes, and dynamically refines the GRN through GSL. The results on publicly available datasets show that our method achieves state-of-the-art performance. The code for reproducing the results can be seen at the link: https://github.com/ChangxiChi/GRAPE. Changxi Chi, Jun Xia 0001, Jiabei Cheng, Chang Yu 0001, Stan Z. Li |
IJCAI | 2 |
| 2025 | MTGIB-UNet: A Multi-Task Graph Information Bottleneck and Uncertainty Weighted Network for ADMET PredictionabstractAccurate prediction of ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) properties is crucial in drug development, as these properties directly impact a drug's efficacy and safety. However, existing multi-task learning models often face challenges related to noise interference and task conflicts when dealing with complex molecular structures. To address these issues, we propose a novel multi-task Graph Neural Network (GNN) model, \textbf{MTGIB-UNet}. The model begins by encoding molecular graphs to capture intricate molecular structure information. Subsequently, based on the Graph Information Bottleneck (GIB) principle, the model compresses the information flow by extracting subgraphs, retaining task-relevant features while removing noise for each task. These embeddings are then fused through a gated network that dynamically adjusts the contribution weights of auxiliary tasks to the primary task. Specifically, an uncertainty weighting (UW) strategy is applied, with additional emphasis placed on the primary task, allowing dynamic adjustment of task weights while strengthening the influence of the primary task on model training. Experiments on standard ADMET datasets demonstrate that our model outperforms existing methods. Additionally, the model shows good interpretability by identifying key molecular substructures related to specific ADMET endpoints. Xuqiang Li, Wenjie Du 0003, Jun Xia 0001, Jianmin Wang 0016, Yang Wang 0015 |
IJCAI | 3 |
| 2025 | A Comprehensive and Systematic Review for Deep Learning-Based De Novo Peptide SequencingabstractTandem mass spectrometry (MS/MS) has revolutionized the field of proteomics, enabling the high-throughput identification of proteins. However, one of the central challenges in mass spectrometry-based proteomics remains peptide identification, especially in the absence of a comprehensive peptide database. While traditional database search methods compare observed mass spectra to pre-existing protein databases, they are limited by the availability and completeness of these databases. \emph{De novo} peptide sequencing, which derives peptide sequences directly from mass spectra, has emerged as a crucial approach in such cases. In recent years, deep learning has made significant strides in this domain. These methods train deep neural networks for translating mass spectra into peptide sequences without relying on any pre-constructed databases. Despite significant progress, this field still lacks a comprehensive and systematic review. In this paper, we provide the first review of deep learning-based \emph{de novo} peptide sequencing techniques from the perspectives of data types, model architectures, decoding strategies, applications and evaluation metrics. We also identify key challenges and highlight promising avenues for future research, providing a valuable resource for the AI and scientific communities. Jun Xia 0001, Shaorong Chen, Tianze Ling, Stan Z. Li |
IJCAI | 1 |
| 2025 | Electron Density-enhanced Molecular Geometry LearningabstractElectron density (ED), which describes the probability distribution of electrons in space, is crucial for accurately understanding the energy and force distribution in molecular force fields (MFF). Existing machine learning force fields (MLFF) focus on mining appropriate physical quantities from the atom-level conformation to enhance the molecular geometry representation while ignoring the unique information from microscopic electrons. In this work, we propose an efficient Electronic Density representation framework to enhance molecular Geometric learning (called EDG), which leverages images rendered from ED to boost molecular geometric representations in MLFF. Specifically, we construct a novel image-based ED representation, which consists of 2 million 6-view images with RGB-D channels, and design an ED representation learning model, called ImageED, to learn ED-related knowledge from these images. We further propose an efficient ED-aware teacher and introduce a cross-modal distillation strategy to transfer knowledge from the image-based teacher to the geometry-based students. Extensive experiments on QM9 and rMD17 demonstrate that EDG can be directly integrated into existing geometry-based models and significantly improves the capabilities of these models (e.g., SchNet, EGNN, SphereNet, ViSNet) for geometry representation learning in MLFF with a maximum average performance increase of 33.7%. Code and appendix are available at https://github.com/HongxinXiang/EDG Hongxin Xiang, Jun Xia 0001, Xin Jin 0014, Wenjie Du 0003, Xiangxiang Zeng |
IJCAI | 2 |
| 2025 | PRESCRIBE: Predicting Single-Cell Responses with Bayesian EstimationabstractIn single-cell perturbation prediction, a central task is to forecast the effects of perturbing a gene unseen in the training data. The efficacy of such predictions depends on two factors: (1) the similarity of the target gene to those covered in the training data, which informs model (epistemic) uncertainty, and (2) the quality of the corresponding training data, which reflects data (aleatoric) uncertainty. Both factors are critical for determining the reliability of a prediction, particularly as gene perturbation is an inherently stochastic biochemical process. In this paper, we propose PRESCRIBE (PREdicting Single-Cell Response wIth Bayesian Estimation), a multivariate deep evidential regression framework designed to measure both sources of uncertainty jointly. Our analysis demonstrates that PRESCRIBE effectively estimates a confidence score for each prediction, which strongly correlates with its empirical accuracy. This capability enables the filtering of untrustworthy results, and in our experiments, it achieves steady accuracy improvements of over 3% compared to comparable baselines. Jiabei Cheng, Changxi Chi, Hongyi Xin, Jun Xia 0001 |
NeurIPS | 5 |
| 2025 | EDBench: Large-Scale Electron Density Data for Molecular ModelingabstractExisting molecular machine learning force fields (MLFFs) generally focus on the learning of atoms, molecules, and simple quantum chemical properties (such as energy and force), but ignore the importance of electron density (ED) $\rho(r)$ in accurately understanding molecular force fields (MFFs). ED describes the probability of finding electrons at specific locations around atoms or molecules, which uniquely determines all ground state properties (such as energy, molecular structure, etc.) of interactive multi-particle systems according to the Hohenberg-Kohn theorem. However, the calculation of ED relies on the time-consuming first-principles density functional theory (DFT), which leads to the lack of large-scale ED data and limits its application in MLFFs. In this paper, we introduce EDBench, a large-scale, high-quality dataset of ED designed to advance learning-based research at the electronic scale. Built upon the PCQM4Mv2, EDBench provides accurate ED data, covering 3.3 million molecules. To comprehensively evaluate the ability of models to understand and utilize electronic information, we design a suite of ED-centric benchmark tasks spanning prediction, retrieval, and generation. Our evaluation of several state-of-the-art methods demonstrates that learning from EDBench is not only feasible but also achieves high accuracy. Moreover, we show that learning-based methods can efficiently calculate ED with comparable precision while significantly reducing the computational cost relative to traditional DFT calculations. All data and benchmarks from EDBench will be freely available, laying a robust foundation for ED-driven drug discovery and materials science. Hongxin Xiang, Mingquan Liu, Zhixiang Cheng, Wenjie Du 0003, Jun Xia 0001, Xin Jin 0014, Xiangxiang Zeng |
NeurIPS | 7 |
| 2025 | Complex hierarchical structures analysis in single-cell data with Poincaré deep manifold transformationabstractSingle-cell RNA sequencing (scRNA-seq) offers remarkable insights into cellular development and differentiation by capturing the gene expression profiles of individual cells. The role of dimensionality reduction and visualization in the interpretation of scRNA-seq data has gained widely acceptance. However, current methods face several challenges, including incomplete structure-preserving strategies and high distortion in embeddings, which fail to effectively model complex cell trajectories with multiple branches. To address these issues, we propose the Poincaré deep manifold transformation (PoincaréDMT) method, which maps high-dimensional scRNA-seq data to a hyperbolic Poincaré disk. This approach preserves global structure from a graph Laplacian matrix while achieving local structure correction through a structure module combined with data augmentation. Additionally, PoincaréDMT alleviates batch effects by integrating a batch graph that accounts for batch labels into the low-dimensional embeddings during network training. Furthermore, PoincaréDMT introduces the Shapley additive explanations method based on trained model to identify the important marker genes in specific clusters and cell differentiation process. Therefore, PoincaréDMT provides a unified framework for multiple key tasks essential for scRNA-seq analysis, including trajectory inference, pseudotime inference, batch correction, and marker gene selection. We validate PoincaréDMT through extensive evaluations on both simulated and real scRNA-seq datasets, demonstrating its superior performance in preserving global and local data structures compared to existing methods. Yongjie Xu 0001, Zelin Zang, Bozhen Hu, Cheng Tan 0012, Jun Xia 0001, Stan Z. Li |
Briefings Bioinform. | 6 |
| 2025 | SP-DTI: subpocket-informed transformer for drug-target interaction predictionabstractMOTIVATION: Drug-target interaction (DTI) prediction is crucial for drug discovery, significantly reducing costs and time in experimental searches across vast drug compound spaces. While deep learning has advanced DTI prediction accuracy, challenges remain: (i) existing methods often lack generalizability, with performance dropping significantly on unseen proteins and cross-domain settings; and (ii) current molecular relational learning often overlooks subpocket-level interactions, which are vital for a detailed understanding of binding sites. RESULTS: We introduce SP-DTI, a subpocket-informed transformer model designed to address these challenges through: (i) detailed subpocket analysis using the Cavity Identification and Analysis Routine for interaction modeling at both global and local levels, and (ii) integration of pre-trained language models into graph neural networks to encode drugs and proteins, enhancing generalizability to unlabeled data. Benchmark evaluations show that SP-DTI consistently outperforms state-of-the-art models, achieving an area under the receiver operating characteristic curve of 0.873 in unseen protein settings, an 11% improvement over the best baseline. AVAILABILITY AND IMPLEMENTATION: The model scripts are available at https://github.com/Steven51516/SP-DTI. Sizhe Liu, Haofeng Xu, Jun Xia 0001, Stan Z. Li |
Bioinform. | 4 |
| 2025 | An image-based protein-ligand binding representation learning framework via multi-level flexible dynamics trajectory pre-trainingabstractMOTIVATION: Accurate prediction of protein-ligand binding (PLB) relationships plays a crucial role in drug discovery, which helps identify drugs that modulate the activity of specific targets. Traditional biological assays for measuring PLB relationships are time consuming and costly. In addition, models for predicting PLB relationships have been developed and widely used in drug discovery tasks. However, learning more accurate PLB representations is essential to meet the stringent standards required for drug discovery. RESULTS: We propose an image-based PLB representation learning framework, called ImagePLB, which equips ligand representation learner (LRL) and protein representation learner (PRL) to accept 3D multi-view ligand images and protein graphs as input, respectively, and learns rich interaction information between ligand and protein through a binding representation learner (BRL). Considering the scarcity of protein-ligand pairs, we further propose a multi-level next trajectory prediction (MLNTP) task to pre-train ImagePLB on the 4D flexible dynamics trajectory of 16 972 complexes, including ligand level, protein level, and complex level, to learn information related to trajectories. Besides, by introducing trajectory regularization (TR), we effectively alleviate the problem of high (even almost identical) feature similarity caused by adjacent trajectories. Compared with the current state-of-the-art methods, ImagePLB has achieved competitive improvements on PLB-related prediction tasks, including protein-ligand affinity and efficacy prediction tasks. This study opens the door to the image-based PLB learning paradigm. AVAILABILITY AND IMPLEMENTATION: All data and implementation details of code can be obtained from https://github.com/HongxinXiang/ImagePLB. Hongxin Xiang, Mingquan Liu, Linlin Hou, Shuting Jin, Jianmin Wang 0016, Jun Xia 0001, Wenjie Du 0003, Sisi Yuan, Xiangzheng Fu, Lei Xu 0047 |
Bioinform. | 6 |
| 2024 | Cross-Gate MLP with Protein Complex Invariant Embedding Is a One-Shot Antibody DesignerabstractAntibodies are crucial proteins produced by the immune system in response to foreign substances or antigens. The specificity of an antibody is determined by its complementarity-determining regions (CDRs), which are located in the variable domains of the antibody chains and form the antigen-binding site. Previous studies have utilized complex techniques to generate CDRs, but they suffer from inadequate geometric modeling. Moreover, the common iterative refinement strategies lead to an inefficient inference. In this paper, we propose a simple yet effective model that can co-design 1D sequences and 3D structures of CDRs in a one-shot manner. To achieve this, we decouple the antibody CDR design problem into two stages: (i) geometric modeling of protein complex structures and (ii) sequence-structure co-learning. We develop a novel macromolecular structure invariant embedding, typically for protein complexes, that captures both intra- and inter-component interactions among the backbone atoms, including Calpha, N, C, and O atoms, to achieve comprehensive geometric modeling. Then, we introduce a simple cross-gate MLP for sequence-structure co-learning, allowing sequence and structure representations to implicitly refine each other. This enables our model to design desired sequences and structures in a one-shot manner. Extensive experiments are conducted to evaluate our results at both the sequence and structure level, which demonstrate that our model achieves superior performance compared to the state-of-the-art antibody CDR design methods. Cheng Tan 0012, Zhangyang Gao, Lirong Wu, Jun Xia 0001, Jiangbin Zheng 0002, Xihong Yang, Yue Liu 0008, Bozhen Hu, Stan Z. Li |
AAAI | 4 |
| 2024 | DiscoGNN: A Sample-Efficient Framework for Self-Supervised Graph Representation LearningabstractSelf-supervised graph representation learning has received increasing research interest recently, with generative and contrastive modeling being two dominant ways. Typically, generative learning first masks parts of each graph and then recovers the masked parts based on the encoding results of the corrupted graph. However, these methods only mask fixed parts of each graph and fail to train on all the nodes and edges, which hinders them from getting the most out of each graph. As a remedy, we propose a novel self-supervised strategy, dubbed DetCor, where we first randomly replace some nodes and edges with alternative ones and then pre-train GNNs to detect and correct the replaced ones from all the nodes and edges. Additionally, for graph-level learning, the vanilla contrastive framework cannot reflect the distinction between the in-batch negatives. To alleviate this issue, we propose RankGCL, which enables the contrastive framework to capture the similarity ranking information between graphs and shows special superiority in graph similarity-based practical tasks. DetCor and RankGCL together constitute a unified self-supervised framework, DiscoGNN, which matches or outperforms state-of-the-art strategies on multiple datasets from various domains. Also, DiscoGNN is a sample-efficient framework that can achieve better performance than competitive methods with much less pre-training data. We release the codes at: https://github.com/junxia97/DiscoGNN-ICDE. Jun Xia 0001, Shaorong Chen, Yue Liu 0008, Zhangyang Gao, Jiangbin Zheng 0002, Xihong Yang, Stan Z. Li |
ICDE | 1 |
| 2024 | KW-Design: Pushing the Limit of Protein Design via Knowledge RefinementabstractRecent studies have shown competitive performance in protein inverse folding, while most of them disregard the importance of predictive confidence, fail to cover the vast protein space, and do not incorporate common protein knowledge. Given the great success of pretrained models on diverse protein-related tasks and the fact that recovery is highly correlated with confidence, we wonder whether this knowledge can push the limits of protein design further. As a solution, we propose a knowledge-aware module that refines low-quality residues. We also introduce a memory-retrieval mechanism to save more than 50\% of the training time. We extensively evaluate our proposed method on the CATH, TS50, TS500, and PDB datasets and our results show that our KW-Design method outperforms the previous PiFold method by approximately 9\% on the CATH dataset. KW-Design is the first method that achieves 60+\% recovery on all these benchmarks. We also provide additional analysis to demonstrate the effectiveness of our proposed method. The code is publicly available via \href{https://github.com/A4Bio/ProteinInvBench}{GitHub}. Zhangyang Gao, Cheng Tan 0012, Xingran Chen, Jun Xia 0001, Siyuan Li 0002, Stan Z. Li |
ICLR | 5 |
| 2024 | Deciphering RNA Secondary Structure Prediction: A Probabilistic K-Rook Matching PerspectiveabstractThe secondary structure of ribonucleic acid (RNA) is more stable and accessible in the cell than its tertiary structure, making it essential for functional prediction. Although deep learning has shown promising results in this field, current methods suffer from poor generalization and high complexity. In this work, we reformulate the RNA secondary structure prediction as a K-Rook problem, thereby simplifying the prediction process into probabilistic matching within a finite solution space. Building on this innovative perspective, we introduce RFold, a simple yet effective method that learns to predict the most matching K-Rook solution from the given sequence. RFold employs a bi-dimensional optimization strategy that decomposes the probabilistic matching problem into row-wise and column-wise components to reduce the matching complexity, simplifying the solving process while guaranteeing the validity of the output. Extensive experiments demonstrate that RFold achieves competitive performance and about eight times faster inference efficiency than the state-of-the-art approaches. The code is available at https://github.com/A4Bio/RFold. Cheng Tan 0012, Zhangyang Gao, Hanqun Cao, Xingran Chen, Lirong Wu, Jun Xia 0001, Jiangbin Zheng 0002, Stan Z. Li |
ICML | 7 |
| 2024 | A Graph is Worth K Words: Euclideanizing Graph using Pure TransformerabstractCan we model Non-Euclidean graphs as pure language or even Euclidean vectors while retaining their inherent information? The Non-Euclidean property have posed a long term challenge in graph modeling. Despite recent graph neural networks and graph transformers efforts encoding graphs as Euclidean vectors, recovering the original graph from vectors remains a challenge. In this paper, we introduce GraphsGPT, featuring an Graph2Seq encoder that transforms Non-Euclidean graphs into learnable Graph Words in the Euclidean space, along with a GraphGPT decoder that reconstructs the original graph from Graph Words to ensure information equivalence. We pretrain GraphsGPT on $100$M molecules and yield some interesting findings: (1) The pretrained Graph2Seq excels in graph representation learning, achieving state-of-the-art results on $8/9$ graph classification and regression tasks. (2) The pretrained GraphGPT serves as a strong graph generator, demonstrated by its strong ability to perform both few-shot and conditional graph generation. (3) Graph2Seq+GraphGPT enables effective graph mixup in the Euclidean space, overcoming previously known Non-Euclidean challenges. (4) The edge-centric pretraining framework GraphsGPT demonstrates its efficacy in graph domain tasks, excelling in both representation and generation. Code is available at https://github.com/A4Bio/GraphsGPT. Zhangyang Gao, Daize Dong, Cheng Tan 0012, Jun Xia 0001, Bozhen Hu, Stan Z. Li |
ICML | 4 |
| 2024 | MMGNN: A Molecular Merged Graph Neural Network for Explainable Solvation Free Energy Prediction
Wenjie Du 0003, Di Wu 0057, Jun Xia 0001, Ziyuan Zhao, Junfeng Fang, Yang Wang 0015 |
IJCAI | 4 |
| 2024 | An Image-enhanced Molecular Graph Representation Learning Framework
Hongxin Xiang, Shuting Jin, Jun Xia 0001, Jianmin Wang 0016, Xiangxiang Zeng |
IJCAI | 3 |
| 2024 | AdaNovo: Towards Robust \emph{De Novo} Peptide Sequencing in Proteomics against Data BiasesabstractTandem mass spectrometry has played a pivotal role in advancing proteomics, enabling the high-throughput analysis of protein composition in biological tissues. Despite the development of several deep learning methods for predicting amino acid sequences (peptides) responsible for generating the observed mass spectra, training data biases hinder further advancements of \emph{de novo} peptide sequencing. Firstly, prior methods struggle to identify amino acids with Post-Translational Modifications (PTMs) due to their lower frequency in training data compared to canonical amino acids, further resulting in unsatisfactory peptide sequencing performance. Secondly, various noise and missing peaks in mass spectra reduce the reliability of training data (Peptide-Spectrum Matches, PSMs). To address these challenges, we propose AdaNovo, a novel and domain knowledge-inspired framework that calculates Conditional Mutual Information (CMI) between the mass spectra and amino acids or peptides, using CMI for robust training against above biases. Extensive experiments indicate that AdaNovo outperforms previous competitors on the widely-used 9-species benchmark, meanwhile yielding 3.6\% - 9.4\% improvements in PTMs identification. The supplements contain the code. Jun Xia 0001, Shaorong Chen, Xiaojun Shan, Wenjie Du 0003, Zhangyang Gao, Cheng Tan 0012, Bozhen Hu, Jiangbin Zheng 0002, Stan Z. Li |
NeurIPS | 1 |
| 2024 | Learning Complete Protein Representation by Dynamically Coupling of Sequence and StructureabstractLearning effective representations is imperative for comprehending proteins and deciphering their biological functions. Recent strides in language models and graph neural networks have empowered protein models to harness primary or tertiary structure information for representation learning. Nevertheless, the absence of practical methodologies to appropriately model intricate inter-dependencies between protein sequences and structures has resulted in embeddings that exhibit low performance on tasks such as protein function prediction. In this study, we introduce CoupleNet, a novel framework designed to interlink protein sequences and structures to derive informative protein representations. CoupleNet integrates multiple levels and scales of features in proteins, encompassing residue identities and positions for sequences, as well as geometric representations for tertiary structures from both local and global perspectives. A two-type dynamic graph is constructed to capture adjacent and distant sequential features and structural geometries, achieving completeness at the amino acid and backbone levels. Additionally, convolutions are executed on nodes and edges simultaneously to generate comprehensive protein embeddings. Experimental results on benchmark datasets showcase that CoupleNet outperforms state-of-the-art methods, exhibiting particularly superior performance in low-sequence similarities scenarios, adeptly identifying infrequently encountered functions and effectively capturing remote homology relationships in proteins. Bozhen Hu, Cheng Tan 0012, Jun Xia 0001, Yue Liu 0008, Lirong Wu, Jiangbin Zheng 0002, Yongjie Xu 0001, Yufei Huang 0002, Stan Z. Li |
NeurIPS | 3 |
| 2024 | ProtGO: Function-Guided Protein Modeling for Unified Representation LearningabstractProtein representation learning is indispensable for various downstream applications of artificial intelligence for bio-medicine research, such as drug design and function prediction. However, achieving effective representation learning for proteins poses challenges due to the diversity of data modalities involved, including sequence, structure, and function annotations. Despite the impressive capabilities of large language models in biomedical text modelling, there remains a pressing need for a framework that seamlessly integrates these diverse modalities, particularly focusing on the three critical aspects of protein information: sequence, structure, and function. Moreover, addressing the inherent data scale differences among these modalities is essential. To tackle these challenges, we introduce ProtGO, a unified model that harnesses a teacher network equipped with a customized graph neural network (GNN) and a Gene Ontology (GO) encoder to learn hybrid embeddings. Notably, our approach eliminates the need for additional functions as input for the student network, which shares the same GNN module. Importantly, we utilize a domain adaptation method to facilitate distribution approximation for guiding the training of the teacher-student framework. This approach leverages distributions learned from latent representations to avoid the alignment of individual samples. Benchmark experiments highlight that ProtGO significantly outperforms state-of-the-art baselines, clearly demonstrating the advantages of the proposed unified framework. Bozhen Hu, Cheng Tan 0012, Yongjie Xu 0001, Zhangyang Gao, Jun Xia 0001, Lirong Wu, Stan Z. Li |
NeurIPS | 5 |
| 2024 | FlexMol: A Flexible Toolkit for Benchmarking Molecular Relational LearningabstractMolecular relational learning (MRL) is crucial for understanding the interaction behaviors between molecular pairs, a critical aspect of drug discovery and development. However, the large feasible model space of MRL poses significant challenges to benchmarking, and existing MRL frameworks face limitations in flexibility and scope. To address these challenges, avoid repetitive coding efforts, and ensure fair comparison of models, we introduce FlexMol, a comprehensive toolkit designed to facilitate the construction and evaluation of diverse model architectures across various datasets and performance metrics. FlexMol offers a robust suite of preset model components, including 16 drug encoders, 13 protein sequence encoders, 9 protein structure encoders, and 7 interaction layers. With its easy-to-use API and flexibility, FlexMol supports the dynamic construction of over 70, 000 distinct combinations of model architectures. Additionally, we provide detailed benchmark results and code examples to demonstrate FlexMol’s effectiveness in simplifying and standardizing MRL model development and comparison. FlexMol is open-sourced and available at https://github.com/Steven51516/FlexMol. Sizhe Liu, Jun Xia 0001, Lecheng Zhang, Yue Liu 0008, Wenjie Du 0003, Zhangyang Gao, Bozhen Hu, Cheng Tan 0012, Hongxin Xiang, Stan Z. Li |
NeurIPS | 2 |
| 2024 | End-to-end Learnable Clustering for Intent Learning in RecommendationabstractIntent learning, which aims to learn users' intents for user understanding and item recommendation, has become a hot research spot in recent years. However, existing methods suffer from complex and cumbersome alternating optimization, limiting performance and scalability. To this end, we propose a novel intent learning method termed \underline{ELCRec}, by unifying behavior representation learning into an \underline{E}nd-to-end \underline{L}earnable \underline{C}lustering framework, for effective and efficient \underline{Rec}ommendation. Concretely, we encode user behavior sequences and initialize the cluster centers (latent intents) as learnable neurons. Then, we design a novel learnable clustering module to separate different cluster centers, thus decoupling users' complex intents. Meanwhile, it guides the network to learn intents from behaviors by forcing behavior embeddings close to cluster centers. This allows simultaneous optimization of recommendation and clustering via mini-batch data. Moreover, we propose intent-assisted contrastive learning by using cluster centers as self-supervision signals, further enhancing mutual promotion. Both experimental results and theoretical analyses demonstrate the superiority of ELCRec from six perspectives. Compared to the runner-up, ELCRec improves NDCG@5 by 8.9\% and reduces computational costs by 22.5\% on the Beauty dataset. Furthermore, due to the scalability and universal applicability, we deploy this method on the industrial recommendation system with 130 million page views and achieve promising results. The codes are available on GitHub\footnote{https://github.com/yueliu1999/ELCRec}. A collection (papers, codes, datasets) of deep group recommendation/intent learning methods is available on GitHub\footnote{https://github.com/yueliu1999/Awesome-Deep-Group-Recommendation}. Yue Liu 0008, Jun Xia 0001, Yingwei Ma, Xinwang Liu 0002, Shengju Yu, Leon Wenliang Zhong |
NeurIPS | 3 |
| 2024 | NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in ProteomicsabstractTandem mass spectrometry has played a pivotal role in advancing proteomics, enabling the analysis of protein composition in biological tissues. Many deep learning methods have been developed for \emph{de novo} peptide sequencing task, i.e., predicting the peptide sequence for the observed mass spectrum. However, two key challenges seriously hinder the further research of this important task. Firstly, since there is no consensus for the evaluation datasets, the empirical results in different research papers are often not comparable, leading to unfair comparison. Secondly, the current methods are usually limited to amino acid-level or peptide-level precision and recall metrics. In this work, we present the first unified benchmark NovoBench for \emph{de novo} peptide sequencing, which comprises diverse mass spectrum data, integrated models, and comprehensive evaluation metrics. Recent impressive methods, including DeepNovo, PointNovo, Casanovo, InstaNovo, AdaNovo and $\pi$-HelixNovo are integrated into our framework. In addition to amino acid-level and peptide-level precision and recall, we also evaluate the models' performance in terms of identifying post-tranlational modifications (PTMs), efficiency and robustness to peptide length, noise peaks and missing fragment ratio, which are important influencing factors while seldom be considered. Leveraging this benchmark, we conduct a large-scale study of current methods, report many insightful findings that open up new possibilities for future development. The benchmark is open-sourced to facilitate future research and application. The code is available at \url{https://github.com/Westlake-OmicsAI/NovoBench}. Shaorong Chen, Jun Xia 0001, Sizhe Liu, Tianze Ling, Wenjie Du 0003, Yue Liu 0008, Jianwei Yin, Stan Z. Li |
NeurIPS | 3 |
| 2024 | Deep Graph Neural Networks via Posteriori-Sampling-based Node-Adaptative Residual ModuleabstractGraph Neural Networks (GNNs), a type of neural network that can learn from graph-structured data through neighborhood information aggregation, have shown superior performance in various downstream tasks. However, as the number of layers increases, node representations becomes indistinguishable, which is known as over-smoothing. To address this issue, many residual methods have emerged. In this paper, we focus on the over-smoothing issue and related residual methods. Firstly, we revisit over-smoothing from the perspective of overlapping neighborhood subgraphs, and based on this, we explain how residual methods can alleviate over-smoothing by integrating multiple orders neighborhood subgraphs to avoid the indistinguishability of the single high-order neighborhood subgraphs. Additionally, we reveal the drawbacks of previous residual methods, such as the lack of node adaptability and severe loss of high-order neighborhood subgraph information, and propose a \textbf{Posterior-Sampling-based, Node-Adaptive Residual module (PSNR)}. We theoretically demonstrate that PSNR can alleviate the drawbacks of previous residual methods. Furthermore, extensive experiments verify the superiority of the PSNR module in fully observed node classification and missing feature scenarios. Our code
is available at \href{https://github.com/jingbo02/PSNR-GNN}{https://github.com/jingbo02/PSNR-GNN}. Ruqiong Zhang, Jun Xia 0001, Zhizhi Yu, Zelin Zang, Di Jin 0001, Carl Yang 0001, Stan Z. Li |
NeurIPS | 4 |
| 2024 | GNN Cleaner: Label Cleaner for Graph Structured DataabstractGraph Neural Network (GNN) has emerged as a predominant tool for graph data analysis. Despite their proliferation, the low-quality labels of many real-world graphs will undermine their performance dramatically. Existing studies on learning neural networks with noisy labels mainly focus on independent data and thus cannot fully exploit the structural information of graph data. Currently, there are few studies of robustness to noisy labels for graph-structured data even if this problem is commonly seen in real-world settings. To remedy this deficiency, we proposeGNN Cleanerwhich utilizes structural information of graph data to combat noisy labels. More specifically, a pseudo label is computed from the neighboring labels for each node in the training set via a modified version of label propagation. Additionally, a novel method is developed to learn to correct the labels adaptively and dynamically. Extensive experiments show that GNN Cleaner can train GNNs robustly and correct both the synthetic and real-world noisy labels even if the noise is severe. Moreover, GNN Cleaner is model-agnostic and can be combined with various GNNs to improve their robustness against label noise. Jun Xia 0001, Yongjie Xu 0001, Cheng Tan 0012, Lirong Wu, Siyuan Li 0002, Stan Z. Li |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Temporal Attention Unit: Towards Efficient Spatiotemporal Predictive LearningabstractSpatiotemporal predictive learning aims to generate future frames by learning from historical frames. In this paper, we investigate existing methods and present a general framework of spatiotemporal predictive learning, in which the spatial encoder and decoder capture intra-frame features and the middle temporal module catches inter-frame correlations. While the mainstream methods employ recurrent units to capture long-term temporal dependencies, they suffer from low computational efficiency due to their unparallelizable architectures. To parallelize the temporal module, we propose the Temporal Attention Unit (TAU), which decomposes temporal attention into intra-frame statical attention and inter-frame dynamical attention. Moreover, while the mean squared error loss focuses on intra-frame errors, we introduce a novel differential divergence regularization to take inter-frame variations into account. Extensive experiments demonstrate that the proposed method enables the derived model to achieve competitive performance on various spatiotemporal prediction benchmarks. Cheng Tan 0012, Zhangyang Gao, Lirong Wu, Yongjie Xu 0001, Jun Xia 0001, Siyuan Li 0002, Stan Z. Li |
CVPR | 5 |
| 2023 | CVT-SLR: Contrastive Visual-Textual Transformation for Sign Language Recognition with Variational AlignmentabstractSign language recognition (SLR) is a weakly supervised task that annotates sign videos as textual glosses. Recent studies show that insufficient training caused by the lack of large-scale available sign datasets becomes the main bottleneck for SLR. Most SLR works thereby adopt pretrained visual modules and develop two mainstream solutions. The multi-stream architectures extend multi-cue visual features, yielding the current SOTA performances but requiring complex designs and might introduce potential noise. Alternatively, the advanced single-cue SLR frameworks using explicit cross-modal alignment between visual and textual modalities are simple and effective, potentially competitive with the multi-cue framework. In this work, we propose a novel contrastive visual-textual transformation for SLR, CVT-SLR, to fully explore the pretrained knowledge of both the visual and language modalities. Based on the single-cue cross-modal alignment framework, we propose a variational autoencoder (VAE) for pretrained contextual knowledge while introducing the complete pretrained language module. The VAE implicitly aligns visual and textual modalities while benefiting from pretrained contextual knowledge as the traditional contextual module. Meanwhile, a contrastive cross-modal alignment algorithm is designed to explicitly enhance the consistency constraints. Extensive experiments on public datasets (PHOENIX-2014 and PHOENIX-2014T) demonstrate that our proposed CVT-SLR consistently outperforms existing single-cue methods and even outperforms SOTA multi-cue methods. The source codes and models are available at https://github.com/binbinjiang/CVT-SLR. Jiangbin Zheng 0002, Yile Wang 0001, Cheng Tan 0012, Siyuan Li 0002, Jun Xia 0001, Yidong Chen 0001, Stan Z. Li |
CVPR | 6 |
| 2023 | Deep Manifold Graph Auto-Encoder For Attributed Graph EmbeddingabstractRepresenting graph data in a low-dimensional space for subsequent tasks is the purpose of attributed graph embedding. Most existing neural network approaches learn latent representations by minimizing reconstruction errors. Rare work considers the data distribution and the topological structure of latent codes simultaneously, which often results in inferior embeddings in real-world graph data. This paper proposes a novel Deep Manifold (Variational) Graph Auto-Encoder (DMVGAE/DMGAE) method for attributed graph data to improve the stability and quality of learned representations to tackle the crowding problem. The node-to-node geodesic similarity is preserved between the original and latent space under a pre-defined distribution. The proposed method surpasses state-of-the-art baseline algorithms by a significant margin on different downstream tasks across popular datasets, which validates our solutions. We promise to release the code after acceptance. Bozhen Hu, Zelin Zang, Jun Xia 0001, Lirong Wu, Cheng Tan 0012, Stan Z. Li |
ICASSP | 3 |
| 2023 | Global-Context Aware Generative Protein DesignabstractThe linear sequence of amino acids determines protein structure and function. Protein design, known as the inverse of protein structure prediction, aims to obtain a novel protein sequence that will fold into the defined structure. Recent works on computational protein design have studied designing sequences for the desired backbone structure with local positional information and achieved competitive performance. However, similar local environments in different backbone structures may result in different amino acids, which indicates the global context of protein structure matters. Thus, we propose the Global-Context Aware generative de novo protein design method (GCA), consisting of local modules and global modules. While local modules focus on relationships between neighbor amino acids, global modules explicitly capture non-local contexts. Experimental results demonstrate that the proposed GCA method achieves state-of-the-art performance on structure-based protein design. Our code and pretrained model have been released on Github1. Cheng Tan 0012, Zhangyang Gao, Jun Xia 0001, Bozhen Hu, Stan Z. Li |
ICASSP | 3 |
| 2023 | Wordreg: Mitigating the Gap between Training and Inference with Worst-Case Drop RegularizationabstractDropout has emerged as one of the most frequently used techniques for training deep neural networks (DNNs). Although effective, the sampled sub-model by random dropout during training is inconsistent with the full model (without dropout) during inference. To mitigate this undesirable gap, we propose WordReg, a simple yet effective regularization built on dropout that enforces the consistency between the outputs of different sub-models sampled by dropout. Specifically, WordReg first obtains the worst-case dropout by maximizing the divergence between the outputs with two sub-models with different random dropouts. And then, it encourages the agreements between the outputs of the two sub-models with worstcase divergence. Extensive experiments on diverse DNNs and tasks reveal that WordReg can achieve notable and consistent improvements over non-regularized models and yields some state-of-the-art results. Theoretically, we verify that WordReg can reduce the gap between training and inference. Jun Xia 0001, Bozhen Hu, Cheng Tan 0012, Jiangbin Zheng 0002, Yongjie Xu 0001, Stan Z. Li |
ICASSP | 1 |
| 2023 | Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules
Jun Xia 0001, Chengshuai Zhao, Bozhen Hu, Zhangyang Gao, Cheng Tan 0012, Yue Liu 0008, Siyuan Li 0002, Stan Z. Li |
ICLR | 1 |
| 2023 | Dink-Net: Neural Clustering on Large GraphsabstractDeep graph clustering, which aims to group the nodes of a graph into disjoint clusters with deep neural networks, has achieved promising progress in recent years. However, the existing methods fail to scale to the large graph with million nodes. To solve this problem, a scalable deep graph clustering method (Dink-Net) is proposed with the idea of dilation and shrink. Firstly, by discriminating nodes, whether being corrupted by augmentations, representations are learned in a self-supervised manner. Meanwhile, the cluster centers are initialized as learnable neural parameters. Subsequently, the clustering distribution is optimized by minimizing the proposed cluster dilation loss and cluster shrink loss in an adversarial manner. By these settings, we unify the two-step clustering, i.e., representation learning and clustering optimization, into an end-to-end framework, guiding the network to learn clustering-friendly features. Besides, Dink-Net scales well to large graphs since the designed loss functions adopt the mini-batch data to optimize the clustering distribution even without performance drops. Both experimental results and theoretical analyses demonstrate the superiority of our method. Compared to the runner-up, Dink-Net achieves $9.62%$ NMI improvement on the ogbn-papers100M dataset with 111 million nodes and 1.6 billion edges. The source code is released: https://github.com/yueliu1999/Dink-Net. Besides, a collection (papers, codes, and datasets) of deep graph clustering is shared on GitHub https://github.com/yueliu1999/Awesome-Deep-Graph-Clustering. Yue Liu 0008, Ke Liang 0006, Jun Xia 0001, Sihang Zhou 0001, Xihong Yang, Xinwang Liu 0002, Stan Z. Li |
ICML | 3 |
| 2023 | A Systematic Survey of Chemical Pre-trained ModelsabstractDeep learning has achieved remarkable success in learning representations for molecules, which is crucial for various biochemical applications, ranging from property prediction to drug design. However, training Deep Neural Networks (DNNs) from scratch often requires abundant labeled molecules, which are expensive to acquire in the real world. To alleviate this issue, tremendous efforts have been devoted to Chemical Pre-trained Models (CPMs), where DNNs are pre-trained using large-scale unlabeled molecular databases and then fine-tuned over specific downstream tasks. Despite the prosperity, there lacks a systematic review of this fast-growing field. In this paper, we present the first survey that summarizes the current progress of CPMs. We first highlight the limitations of training molecular representation models from scratch to motivate CPM studies. Next, we systematically review recent advances on this topic from several key perspectives, including molecular descriptors, encoder architectures, pre-training strategies, and applications. We also highlight the challenges and promising avenues for future research, providing a useful resource for both machine learning and scientific communities. Jun Xia 0001, Yanqiao Zhu 0001, Yuanqi Du, Stan Z. Li |
IJCAI | 1 |
| 2023 | Reinforcement Graph Clustering with Unknown Cluster NumberabstractDeep graph clustering, which aims to group nodes into disjoint clusters by neural networks in an unsupervised manner, has attracted great attention in recent years. Although the performance has been largely improved, the excellent performance of the existing methods heavily relies on an accurately predefined cluster number, which is not always available in the real-world scenario. To enable the deep graph clustering algorithms to work without the guidance of the predefined cluster number, we propose a new deep graph clustering method termed Reinforcement Graph Clustering (RGC). In our proposed method, cluster number determination and unsupervised representation learning are unified into a uniform framework by the reinforcement learning mechanism. Concretely, the discriminative node representations are first learned with the contrastive pretext task. Then, to capture the clustering state accurately with both local and global information in the graph, both node and cluster states are considered. Subsequently, at each state, the qualities of different cluster numbers are evaluated by the quality network, and the greedy action is executed to determine the cluster number. In order to conduct feedback actions, the clustering-oriented reward function is proposed to enhance the cohesion of the same clusters and separate the different clusters. Extensive experiments demonstrate the effectiveness and efficiency of our proposed method. The source code of RGC is shared at https://github.com/yueliu1999/RGC and a collection (papers, codes and, datasets) of deep graph clustering is shared at https://github.com/yueliu1999/Awesome-Deep-Graph-Clustering on Github. Yue Liu 0008, Ke Liang 0006, Jun Xia 0001, Xihong Yang, Sihang Zhou 0001, Meng Liu 0014, Xinwang Liu 0002, Stan Z. Li |
ACM Multimedia | 3 |
| 2023 | CONVERT: Contrastive Graph Clustering with Reliable AugmentationabstractContrastive graph node clustering via learnable data augmentation is a hot research spot in the field of unsupervised graph learning. The existing methods learn the sampling distribution of a pre-defined augmentation to generate data-driven augmentations automatically. Although promising clustering performance has been achieved, we observe that these strategies still rely on pre-defined augmentations, the semantics of the augmented graph can easily drift. The reliability of the augmented view semantics for contrastive learning can not be guaranteed, thus limiting the model performance. To address these problems, we propose a novel CONtrastiVe Graph ClustEring network with Reliable AugmenTation (COVERT). Specifically, in our method, the data augmentations are processed by the proposed reversible perturb-recover network. It distills reliable semantic information by recovering the perturbed latent embeddings. Moreover, to further guarantee the reliability of semantics, a novel semantic loss is presented to constrain the network via quantifying the perturbation and recovery. Lastly, a label-matching mechanism is designed to guide the model by clustering information through aligning the semantic labels and the selected high-confidence clustering pseudo labels. Extensive experimental results on seven datasets demonstrate the effectiveness of the proposed method. We release the code and appendix of CONVERT at https://github.com/xihongyang1999/CONVERT on GitHub. Xihong Yang, Cheng Tan 0012, Yue Liu 0008, Ke Liang 0006, Siwei Wang 0001, Sihang Zhou 0001, Jun Xia 0001, Stan Z. Li, Xinwang Liu 0002, En Zhu |
ACM Multimedia | 7 |
| 2023 | Understanding the Limitations of Deep Models for Molecular property prediction: Insights and SolutionsabstractMolecular Property Prediction (MPP) is a crucial task in the AI-driven Drug Discovery (AIDD) pipeline, which has recently gained considerable attention thanks to advancements in deep learning. However, recent research has revealed that deep models struggle to beat traditional non-deep ones on MPP. In this study, we benchmark 12 representative models (3 non-deep models and 9 deep models) on 15 molecule datasets. Through the most comprehensive study to date, we make the following key observations: \textbf{(\romannumeral 1)} Deep models are generally unable to outperform non-deep ones; \textbf{(\romannumeral 2)} The failure of deep models on MPP cannot be solely attributed to the small size of molecular datasets; \textbf{(\romannumeral 3)} In particular, some traditional models including XGB and RF that use molecular fingerprints as inputs tend to perform better than other competitors. Furthermore, we conduct extensive empirical investigations into the unique patterns of molecule data and inductive biases of various models underlying these phenomena. These findings stimulate us to develop a simple-yet-effective feature mapping method for molecule data prior to feeding them into deep models. Empirically, deep models equipped with this mapping method can beat non-deep ones in most MoleculeNet datasets. Notably, the effectiveness is further corroborated by extensive experiments on cutting-edge dataset related to COVID-19 and activity cliff datasets. Jun Xia 0001, Lecheng Zhang, Yue Liu 0008, Zhangyang Gao, Bozhen Hu, Cheng Tan 0012, Jiangbin Zheng 0002, Siyuan Li 0002, Stan Z. Li |
NeurIPS | 1 |
| 2023 | Co-supervised Pre-training of Pocket and Ligand
Zhangyang Gao, Cheng Tan 0012, Jun Xia 0001, Stan Z. Li |
ECML/PKDD (1) | 3 |
| 2022 | Using Context-to-Vector with Graph Retrofitting to Improve Word EmbeddingsabstractJiangbin Zheng, Yile Wang, Ge Wang, Jun Xia, Yufei Huang, Guojiang Zhao, Yue Zhang, Stan Li. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Jiangbin Zheng 0002, Yile Wang 0001, Jun Xia 0001, Yufei Huang 0002, Guojiang Zhao, Yue Zhang 0004, Stan Z. Li |
ACL (1) | 4 |
| 2022 | OT Cleaner: Label Correction as Optimal TransportabstractDatasets with noisy labels present challenges for training Deep Neural Networks (DNNs) with high generalization ability. An direct idea is to correct the noisy labels for robust learning. However, existing label correction methods can not handle with heavy noise or datasets with samples of many categories so well. We explain the reasons and introduce a global label distribution regularization to remedy these deficiencies. With this regularization, we convert the label correction to the Optimal Transport (OT) formulation and propose to utilize a fast version of the Sinkhorn-Knopp algorithm for finding an approximate solution efficiently at scale. Experiments on benchmark datasets with both synthetic and real-world label noise show that the superiority of our OT Cleaner in terms of both training efficiency and classification accuracy. The code is available at: https://github.com/junxia97/OT-Cleaner. Jun Xia 0001, Cheng Tan 0012, Lirong Wu, Yongjie Xu 0001, Stan Z. Li |
ICASSP | 1 |
| 2022 | ProGCL: Rethinking Hard Negative Mining in Graph Contrastive LearningabstractContrastive Learning (CL) has emerged as a dominant technique for unsupervised representation learning which embeds augmented versions of the anchor close to each other (positive samples) and pushes the embeddings of other samples (negatives) apart. As revealed in recent studies, CL can benefit from hard negatives (negatives that are most similar to the anchor). However, we observe limited benefits when we adopt existing hard negative mining techniques of other domains in Graph Contrastive Learning (GCL). We perform both experimental and theoretical analysis on this phenomenon and find it can be attributed to the message passing of Graph Neural Networks (GNNs). Unlike CL in other domains, most hard negatives are potentially false negatives (negatives that share the same class with the anchor) if they are selected merely according to the similarities between anchor and themselves, which will undesirably push away the samples of the same class. To remedy this deficiency, we propose an effective method, dubbed \textbf{ProGCL}, to estimate the probability of a negative being true one, which constitutes a more suitable measure for negatives’ hardness together with similarity. Additionally, we devise two schemes (i.e., \textbf{ProGCL-weight} and \textbf{ProGCL-mix}) to boost the performance of GCL. Extensive experiments demonstrate that ProGCL brings notable and consistent improvements over base GCL methods and yields multiple state-of-the-art results on several unsupervised benchmarks or even exceeds the performance of supervised ones. Also, ProGCL is readily pluggable into various negatives-based GCL methods for performance improvement. We release the code at \textcolor{magenta}\url{https://github.com/junxia97/ProGCL}. Jun Xia 0001, Lirong Wu, Jintao Chen 0001, Stan Z. Li |
ICML | 1 |
| 2022 | GraphMixup: Improving Class-Imbalanced Node Classification by Reinforcement Mixup and Self-supervised Context Prediction
Lirong Wu, Jun Xia 0001, Zhangyang Gao, Cheng Tan 0012, Stan Z. Li |
ECML/PKDD (4) | 2 |
| 2022 | Generalized Clustering and Multi-Manifold Learning with Geometric Structure PreservationabstractThough manifold-based clustering has become a popular research topic, we observe that one important factor has been omitted by these works, namely that the defined clustering loss may corrupt the local and global structure of the latent space. In this paper, we propose a novel Generalized Clustering and Multi-manifold Learning (GCML) framework with geometric structure preservation for generalized data, i.e., not limited to 2-D image data and has a wide range of applications in speech, text, and biology domains. In the proposed framework, manifold clustering is done in the latent space guided by a clustering loss. To overcome the problem that the clustering-oriented loss may deteriorate the geometric structure of the latent space, an isometric loss is proposed for preserving intra-manifold structure locally and a ranking loss for inter-manifold structure globally. Extensive experimental results have shown that GCML exhibits superior performance to counterparts in terms of qualitative visualizations and quantitative metrics, which demonstrates the effectiveness of preserving geometric structure. Code has been made available at: https://github.com/LirongWu/GCML. Lirong Wu, Zicheng Liu 0006, Jun Xia 0001, Zelin Zang, Siyuan Li 0002, Stan Z. Li |
WACV | 3 |
| 2022 | SimGRACE: A Simple Framework for Graph Contrastive Learning without Data AugmentationabstractGraph contrastive learning (GCL) has emerged as a dominant technique for graph representation learning which maximizes the mutual information between paired graph augmentations that share the same semantics. Unfortunately, it is difficult to preserve semantics well during augmentations in view of the diverse nature of graph data. Currently, data augmentations in GCL broadly fall into three unsatisfactory ways. First, the augmentations can be manually picked per dataset by trial-and-errors. Second, the augmentations can be selected via cumbersome search. Third, the augmentations can be obtained with expensive domain knowledge as guidance. All of these limit the efficiency and more general applicability of existing GCL methods. To circumvent these crucial issues, we propose a Simple framework for GRAph Contrastive lEarning, SimGRACE for brevity, which does not require data augmentations. Specifically, we take original graph as input and GNN model with its perturbed version as two encoders to obtain two correlated views for contrast. SimGRACE is inspired by the observation that graph data can preserve their semantics well during encoder perturbations while not requiring manual trial-and-errors, cumbersome search or expensive domain knowledge for augmentations selection. Also, we explain why SimGRACE can succeed. Furthermore, we devise adversarial training scheme, dubbed AT-SimGRACE, to enhance the robustness of graph contrastive learning and theoretically explain the reasons. Albeit simple, we show that SimGRACE can yield competitive or better performance compared with state-of-the-art methods in terms of generalizability, transferability and robustness, while enjoying unprecedented degree of flexibility and efficiency. The code is available at: https://github.com/junxia97/SimGRACE. Jun Xia 0001, Lirong Wu, Jintao Chen 0001, Bozhen Hu, Stan Z. Li |
WWW | 1 |
| 2022 | Multi-level disentanglement graph neural network
Lirong Wu, Jun Xia 0001, Cheng Tan 0012, Stan Z. Li |
Neural Comput. Appl. | 3 |
| 2022 | BA_EnCaps: Dense Capsule Architecture for Thermal ScrutinyabstractRemote sensing integrated with deep learning (DL) improves wildfire assessment. The research has been done to scrutinize the areas affected by disastrous wildfires in Yunnan using DL. Wildfire identification and demarcation of the affected area have been limited to primitive thresholding and outdated machine learning classification techniques. Therefore, the research work incorporated DL in the wildfire scrutiny, and several of the most important considerations are investigated. The proposed research objective is to exploit the recent advancement of capsule-based DL together with the wildfire domain. The proposed dense structure provides highly efficient detection and segmentation of the burned area (BA). The BA dense capsule network (BA_EnCaps) is employed to extract and localize the burned zone with an overall accuracy of 98%. The model is evaluated quantitatively using accuracy, binary_cross-entropy, dice_loss, and mean square error (mse). The research aims to utilize the segmentation model to estimate the BA with great results. BA-EnCaps shows excellent accuracy in discriminating the spectral indices for the burned zone. The proposed method surpasses other segmentation benchmark techniques (U-Net, U-Net3p, SegCaps, Deep U-Net, and U-Net+) by substantially lessening the computing power. Finally, BA_EnCaps is compared with standard segmentation techniques and shows that DL-based models can assess wildfire better than conventional algorithms. Qurratulain Safder, Fangrong Zhou, Zezhong Zheng, Jun Xia 0001, Mingcang Zhu, Yong He 0007, Jiang Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Co-learning: Learning from Noisy Labels with Self-supervisionabstractNoisy labels, resulting from mistakes in manual labeling or webly data collecting for supervised learning, can cause neural networks to overfit the misleading information and degrade the generalization performance. Self-supervised learning works in the absence of labels and thus eliminates the negative impact of noisy labels. Motivated by co-training with both supervised learning view and self-supervised learning view, we propose a simple yet effective method called Co-learning for learning with noisy labels. Co-learning performs supervised learning and self-supervised learning in a cooperative way. The constraints of intrinsic similarity with the self-supervised module and the structural similarity with the noisily-supervised module are imposed on a shared common feature encoder to regularize the network to maximize the agreement between the two constraints. Co-learning is compared with peer methods on corrupted data from benchmark datasets fairly, and extensive results are provided which demonstrate that Co-learning is superior to many state-of-the-art approaches. Cheng Tan 0012, Jun Xia 0001, Lirong Wu, Stan Z. Li |
ACM Multimedia | 2 |
| 2021 | Invertible Manifold Learning for Dimension Reduction
Siyuan Li 0002, Zelin Zang, Lirong Wu, Jun Xia 0001, Stan Z. Li |
ECML/PKDD (3) | 5 |
| 2020 | Change of Impervious Surface of Chengdu City, ChinaabstractImpervious surfaces have become the most intuitive indicator in the process of urbanization. Timely and accurate information on impervious surfaces from remote sensing images is essential. It not only helps us understand the process of land use/cover change, but also the influences on human society and the environment. In this study, convolutional neural network (CNN) was used to extract the impervious surface in Chengdu city, Sichuan province, China. The overall accuracy in 2009 and 2017 were 98.75% and 99.76% respectively. From the results for 2009 and 2017, the impervious surface increased by 51.24 km2, Growth rate is 13.8%. During the process of urban expansion, suburban farmland was replaced by impervious surfaces and the area of impervious surface gradually increased. Jibao Shi, Jun Xia 0001, Tao Weng, Zezhong Zheng |
IGARSS | 4 |
| 2020 | Land Use and Land Cover Change of GhanaabstractLand use and cover change (LUCC) is a central component in current strategies for managing natural resources and monitoring environmental change. In this paper, we used maximum likelihood classification algorithm to obtain the supervised land use and cover classification. Four major land use and cover classes are identified and mapped from 2000 to 2015. The changes of land use and cover using Landsat images of the study area were analyzed. The results showed that: From 2005 to 2015, closed forest has increased and the annual rate of change was (+)3.3%. Open forest has an annual rate of change of (+)1.21%. Water bodies had an annual rate of change of (+)0.81%. While the settlements and bare lands had a decrease of 52.93 km2and the annual rate of change was (-)5.3%. Ankai Hou, Abrado Blankson Samuel, Mujie Li, Zezhong Zheng, Jun Xia 0001, Xiang Zhang 0002, Guoqing Zhou 0001 |
IGARSS | 5 |
| 2020 | Drought Monitoring in Sub-Sahara AfricaabstractDrought is one of the main natural hazards affecting the environment and economy of countries all over the world. Fusing weather data with satellite images therefore becomes a superior method of identifying and monitoring drought in a given region. We established the relationship between land surface temperature (LST), the normalized differential vegetation index (NDVI) and rainfall data to derive areas of drought. Then, we obtained the indexes from the rainfall anomaly and NDVI anomaly as indicators which confirm the drought indicative claims of the maps produced. Our further examination of the NDVI, LST and rainfall maps indicate that the western, central and Volta Regions of the study area are the least prone to drought, with Axim (one of the most southern towns) in Ghana recording the highest rainfall in the country each year. Fan Mou, Twum-Antwi Akwasi, Mujie Li, Mingcang Zhu, Yong He 0007, Zhanyong He, Juan Ren, Jun Xia 0001, Xiang Zhang 0002, Zezhong Zheng, Guoqing Zhou 0001 |
IGARSS | 9 |
| 2019 | Land Price Assesment Based on Deep Neural NetworkabstractThe land resource is becoming scarcer and scarcer for a rapidly developing city. Thus, the land price assessment is important for the government to auction the land appropriately. In the paper, we introduced the deep neural network to evaluate the land price, taking the Shenzhen city in China as a case. Firstly, twenty influencing factors and land price data were gathered. Then, Shenzhen city was segmented into many grids with a size of 300 × 300 m. Secondly, the land price of each grid was derived with Kriging approach based upon the samples of land price. And the twenty influencing factors was quantified. Thirdly, the land price data and influencing factors were partitioned into training and testing datasets with the ratio of 8:1, and the training data were utilized to train the deep neural network based on regression analysis and classification with different hidden layers. Finally, the results were analyzed, and the deep neural network with the highest accuracy was selected as the optimum model. Therefore, our proposed method is an efficient approach to evaluate the land price with deep neural network. Ankai Hou, Guoqing Zhou 0001, Hongsheng Zhang 0001, Jiang Li 0001, Yuxuan Tao, Shaobin Jiang, Kai Li 0011, Zezhong Zheng, Jun Xia 0001, Yong He 0007, Mingcang Zhu |
IGARSS | 10 |
| 2019 | Urban Functional Regions Discovering Based on Deep LearningabstractIn recent years, the big data industry chain has become more mature. Analyzing and managing cities by utilizing various big data in cities has become a hot research topic. Urban functional regions discovering is one of the important applications. The mainstream in urban functional regions discovering are probabilistic topic models, such as latent Dirichlet allocation (LDA) based topic model, which seeing the regions as documents and their functions are their topics. These methods require feature engineering by hand, which will construct features of limited expressiveness. To overcome these methods' shortcomings, we introduced a deep learning topic model called document neural autoregressive distribution estimation (DocNADE) into urban functional regions mining. And we did an experiment to test its effect. The experimental result shows that this DocNADE framework has achieved a considerable result in urban function inference compared with Dirichlet Multinomial Regression (DMR) based topic model which is a state of the art of urban functional regions discovering. Fan Mou, Zhigang Liu 0013, Ankai Hou, Shengli Wang, Jiang Li 0001, Kai Li 0011, Zezhong Zheng, Jun Xia 0001, Yong He 0007, Mingcang Zhu, Guoqing Zhou 0001, Hongsheng Zhang 0001 |
IGARSS | 10 |
| 2018 | Monitoring of Drought Change in the Middle Reach of Yangtze RiverabstractDrought is a weather phenomenon widespread worldwide due to the water shortage or unbalance of supply and demand, and it's also one of the most serious natural disasters for human life and agricultural production. The middle reach of Yangtze river, one of China's most important grain producer, subjected to the sub-tropical monsoon climate, is prone to have droughts. This paper has practical implications as it build a model by depending on the normalized difference vegetation index (NDVI) and land surface temperature (LST) of moderate resolution imaging spectroradiometer (MODIS) between 2005 and 2009. Firstly, the 8-day LST and 16-day NDVI data, 8-day LST and 30-day NDVI data were utilized to construct the LST/NDVI feature space. Secondly, the temperature vegetation dryness index (TVDI) images of the reach were derived respectively. Thirdly, the temporal evolution and spatial variation of drought was analyzed. Finally, the results of two different years were compared to analyze the drought in the reach. Our study showed the drought in May was more severe than that in other months. Therefore, a severe drought event is more likely to happen in May in the middle reach of Yangtze river and more measures should be taken to alleviate the loss for the governments. Pingchuan Zhang, Zezhong Zheng, Jun Xia 0001, Xiang Zhang 0002, Mingcang Zhu, Guoqing Zhou 0001, Jiang Li 0001 |
IGARSS | 4 |
| 2015 | The application of ant colony algorithm in emergency rescue with GISabstractUnder the indoor building environment, when the fires and other accidents occur, how to effectively organize the masses evacuation and fire rescue, is closely related to the safety of people's lives and property and has become a critical problem of public concern. This paper presents an improved ant colony algorithm (ACO) to solve the problem of how to optimize the evacuation route and rescue route when an accident occurs. According to the key factors affecting people emergency evacuation, such as indoor building environment, fire and its combustion products, problem of path's optimal selection, etc., we propose an emergency evacuation model, based on the model it can give an optimal evacuation route for the mass and an optimal rescue route for the firefighters. We also analyzes the search results, it shows that the search results is robust and reasonable. Yufeng Lu, Yong He 0007, Jun Xia 0001, Zezhong Zheng, Huan Wei, Yalan Liu, Xiang Zhang 0002, Guoqing Zhou 0001, Zhanmang Liao, Guiyun Zhou, Hongsheng Zhang 0001, Jiang Li 0001 |
IGARSS | 3 |
| 2015 | Drought monitoring and warning in the middle reach of Yangtze River with MODISabstractIn China, drought is one of the major environmental disasters, which bring great harm to the people. The middle reach of Yangtze River is the most important base to produce grains in China. Influenced by the summer monsoon, the drought occurs frequently. In our paper, the NDVI and LST from MODIS data were utilized to calculate the TVDI (Temperature Vegetation Dryness Index), which were used to monitor the drought of the study area. Meteorological drought indices were calculated from 10-day precipitation, temperature and evaporation data of 94 meteorological stations, including precipitation standardized variables, dryness and relative moisture index were used to analyze the degree of drought and the area of drought. The results showed that TVDI is significantly related to soil moisture. Lanying Yuan, Mingcang Zhu, Zezhong Zheng, Jun Xia 0001, Xiang Zhang 0002, Yong He 0007, Guoqing Zhou 0001, Xiaowen Li 0001, Guiyun Zhou, Yufeng Lu, Shi Qiu 0003, Hongsheng Zhang 0001, Jiang Li 0001 |
IGARSS | 4 |