Weitao Du

dblp:17/10015 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0001-7643-4671ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Generative modeling · 28% Graph learning · 23% Deep learning architectures and training · 13%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 56% Computational science and engineering · 44%

Topics — the 27 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.522025
GDiffRetro: Retrosynthesis Prediction with Dual Graph Enhanced Molecular Representation and Diffusion Generation · AAAI 2025
A Flexible Diffusion Model · ICML 2023
Machine learning › Graph learning › graph neural network › geometric graph neural network
equivariant graph neural network
1.222023
A new perspective on building efficient and expressive 3D equivariant graph neural networks · NeurIPS 2023
SE(3) Equivariant Graph Neural Networks with Complete Local Frames · ICML 2022
Machine learning › Graph learning
graph neural network
1.122022
SE(3) Equivariant Graph Neural Networks with Complete Local Frames · ICML 2022
Towards Deepening Graph Neural Networks: A GNTK-based Optimization Perspective · ICLR 2022
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
1.122022
Towards Deepening Graph Neural Networks: A GNTK-based Optimization Perspective · ICLR 2022
On the Neural Tangent Kernel of Deep Networks with Orthogonal Initialization · IJCAI 2021
Machine learning › Generative modeling
molecular generation
0.912025
GDiffRetro: Retrosynthesis Prediction with Dual Graph Enhanced Molecular Representation and Diffusion Generation · AAAI 2025
Bioinformatics and computational biology › molecular informatics
cheminformatics
0.912025
GDiffRetro: Retrosynthesis Prediction with Dual Graph Enhanced Molecular Representation and Diffusion Generation · AAAI 2025
Computational science and engineering › computational chemistry
retrosynthesis prediction
0.912025
GDiffRetro: Retrosynthesis Prediction with Dual Graph Enhanced Molecular Representation and Diffusion Generation · AAAI 2025
Computer vision › 3D vision › object modeling › geometric modeling
geometric representation
0.712023
A new perspective on building efficient and expressive 3D equivariant graph neural networks · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning
geometric representation learning
0.712023
Symmetry-Informed Geometric Representation for Molecules, Proteins, and Crystalline Materials · NeurIPS 2023
Machine learning › Representation and self-supervised learning › pre-training
molecular pre-training
0.712023
A Group Symmetric Stochastic Differential Equation Model for Molecule Multi-modal Pretraining · ICML 2023
Machine learning › Generative modeling › diffusion model
score-based generative model
0.712023
A Flexible Diffusion Model · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › continuous-time model
stochastic differential equations
0.712023
A Flexible Diffusion Model · ICML 2023
Machine learning › Generative modeling › generative model › continuous-time generative model
stochastic differential equation models
0.712023
A Group Symmetric Stochastic Differential Equation Model for Molecule Multi-modal Pretraining · ICML 2023
Bioinformatics and computational biology › molecular informatics
molecular representation learning
0.712023
Molecule Joint Auto-Encoding: Trajectory Pretraining with 2D and 3D Diffusion · NeurIPS 2023
Machine learning › Graph learning › graph neural network
deep graph neural network
0.612022
Towards Deepening Graph Neural Networks: A GNTK-based Optimization Perspective · ICLR 2022
Machine learning › Deep learning architectures and training
equivariant neural network
0.612022
SE(3) Equivariant Graph Neural Networks with Complete Local Frames · ICML 2022
Computer vision › 3D vision
geometric deep learning
0.612022
SE(3) Equivariant Graph Neural Networks with Complete Local Frames · ICML 2022
Machine learning › Graph learning › graph kernel
graph neural tangent kernel
0.612022
Towards Deepening Graph Neural Networks: A GNTK-based Optimization Perspective · ICLR 2022
Computer vision › 3D vision › geometric deep learning
SE(3) equivariance
0.612022
SE(3) Equivariant Graph Neural Networks with Complete Local Frames · ICML 2022
Machine learning › Deep learning architectures and training › training dynamics
lazy training regime
0.512021
On the Neural Tangent Kernel of Deep Networks with Orthogonal Initialization · IJCAI 2021
Machine learning › Deep learning architectures and training › weight initialization
orthogonal initialization
0.512021
On the Neural Tangent Kernel of Deep Networks with Orthogonal Initialization · IJCAI 2021
Machine learning › Deep learning architectures and training
training dynamics
0.512021
On the Neural Tangent Kernel of Deep Networks with Orthogonal Initialization · IJCAI 2021
Machine learning › Representation and self-supervised learning › pre-training
geometric pretraining
0.212023
Symmetry-Informed Geometric Representation for Molecules, Proteins, and Crystalline Materials · NeurIPS 2023
Machine learning › Representation and self-supervised learning
pre-training
0.212023
Symmetry-Informed Geometric Representation for Molecules, Proteins, and Crystalline Materials · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.212023
Molecule Joint Auto-Encoding: Trajectory Pretraining with 2D and 3D Diffusion · NeurIPS 2023
Bioinformatics and computational biology
drug discovery
0.212023
A Group Symmetric Stochastic Differential Equation Model for Molecule Multi-modal Pretraining · ICML 2023
Bioinformatics and computational biology
molecular property prediction
0.212023
A new perspective on building efficient and expressive 3D equivariant graph neural networks · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

dual graph neural networks · 1.7conditional diffusion model · 1.7mutual information maximization · 1.3joint auto-encoding · 1.3group symmetric stochastic differential equation · 1.32d and 3d diffusion · 1.3variational optimization · 0.7symplectic geometry · 0.7riemannian geometry · 0.7local substructure encoding · 0.7frame transition encoding · 0.7equivariance · 0.7SE(3)-equivariance · 0.7SE(3) equivariance · 0.7
YearPublicationVenuePosition
2025 GDiffRetro: Retrosynthesis Prediction with Dual Graph Enhanced Molecular Representation and Diffusion Generation
abstract
Retrosynthesis prediction focuses on identifying reactants capable of synthesizing a target product. Typically, the retrosynthesis prediction involves two phases: Reaction Center Identification and Reactant Generation. However, we argue that most existing methods suffer from two limitations in the two phases: (i) Existing models do not adequately capture the ``face'' information in molecular graphs for the reaction center identification. (ii) Current approaches for the reactant generation predominantly use sequence generation in a 2D space, which lacks versatility in generating reasonable distributions for completed reactive groups and overlooks molecules' inherent 3D properties. To overcome the above limitations, we propose GDiffRetro. For the reaction center identification, GDiffRetro uniquely integrates the original graph with its corresponding dual graph to represent molecular structures, which helps guide the model to focus more on the faces in the graph. For the reactant generation, GDiffRetro employs a conditional diffusion model in 3D to further transform the obtained synthon into a complete reactant. Our experimental findings reveal that GDiffRetro outperforms state-of-the-art semi-template models across various evaluative metrics.
Shengyin Sun, Wenhao Yu 0014, Yuxiang Ren, Weitao Du, Xuecang Zhang, Chen Ma 0001
AAAI4
2025 Sculpting molecules in text-3D space: a flexible substructure aware framework for text-oriented molecular optimization
abstract
The integration of deep learning, particularly AI-Generated Content, with high-quality data derived from ab initio calculations has emerged as a promising avenue for transforming the landscape of scientific research. However, the challenge of designing molecular drugs or materials that incorporate multi-modality prior knowledge remains a critical and complex undertaking. Specifically, achieving a practical molecular design necessitates not only meeting the diversity requirements but also addressing structural and textural constraints with various symmetries outlined by domain experts. In this article, we present an innovative approach to tackle this inverse design problem by formulating it as a multi-modality guidance optimization task. Our proposed solution involves a textural-structure alignment symmetric diffusion framework for the implementation of molecular optimization tasks, namely 3DToMolo. 3DToMolo aims to harmonize diverse modalities including textual description features and graph structural features, aligning them seamlessly to produce molecular structures adhere to specified symmetric structural and textural constraints by experts in the field. Experimental trials across three guidance optimization settings have shown a superior hit optimization performance compared to state-of-the-art methodologies. Moreover, 3DToMolo demonstrates the capability to discover potential novel molecules, incorporating specified target substructures, without the need for prior knowledge. This work not only holds general significance for the advancement of deep learning methodologies but also paves the way for a transformative shift in molecular design strategies. 3DToMolo creates opportunities for a more nuanced and effective exploration of the vast chemical space, opening new frontiers in the development of molecular entities with tailored properties and functionalities.
Kaiwei Zhang, Yange Lin, Guangcheng Wu, Yuxiang Ren, Xuecang Zhang, Weitao Du
BMC Bioinform.8
2024 CGCL: Collaborative Graph Contrastive Learning Without Handcrafted Graph Data Augmentations
Yuxiang Ren, Wenzheng Feng, Weitao Du, Xuecang Zhang
DASFAA (6)4
2023 A Flexible Diffusion Model
abstract
Denoising diffusion (score-based) generative models have become a popular choice for modeling complex data. Recently, a deep connection between forward-backward stochastic differential equations (SDEs) and diffusion-based models has been established, leading to the development of new SDE variants such as sub-VP and critically-damped Langevin. Despite the empirical success of some hand-crafted forward SDEs, many potentially promising forward SDEs remain unexplored. In this work, we propose a general framework for parameterizing diffusion models, particularly the spatial part of forward SDEs, by leveraging the symplectic and Riemannian geometry of the data manifold. We introduce a systematic formalism with theoretical guarantees and connect it with previous diffusion models. Finally, we demonstrate the theoretical advantages of our method from a variational optimization perspective. We present numerical experiments on synthetic datasets, MNIST and CIFAR10 to validate the effectiveness of our framework.
Weitao Du, Yuanqi Du
ICML1
2023 A Group Symmetric Stochastic Differential Equation Model for Molecule Multi-modal Pretraining
abstract
Molecule pretraining has quickly become the go-to schema to boost the performance of AI-based drug discovery. Naturally, molecules can be represented as 2D topological graphs or 3D geometric point clouds. Although most existing pertaining methods focus on merely the single modality, recent research has shown that maximizing the mutual information (MI) between such two modalities enhances the molecule representation ability. Meanwhile, existing molecule multi-modal pretraining approaches approximate MI based on the representation space encoded from the topology and geometry, thus resulting in the loss of critical structural information of molecules. To address this issue, we propose MoleculeSDE. MoleculeSDE leverages group symmetric (e.g., SE(3)-equivariant and reflection-antisymmetric) stochastic differential equation models to generate the 3D geometries from 2D topologies, and vice versa, directly in the input space. It not only obtains tighter MI bound but also enables prosperous downstream tasks than the previous work. By comparing with 17 pretraining baselines, we empirically verify that MoleculeSDE can learn an expressive representation with state-of-the-art performance on 26 out of 32 downstream tasks.
Shengchao Liu, Weitao Du, Zhiming Ma, Jian Tang 0005
ICML2
2023 Molecule Joint Auto-Encoding: Trajectory Pretraining with 2D and 3D Diffusion
abstract
Recently, artificial intelligence for drug discovery has raised increasing interest in both machine learning and chemistry domains. The fundamental building block for drug discovery is molecule geometry and thus, the molecule's geometrical representation is the main bottleneck to better utilize machine learning techniques for drug discovery. In this work, we propose a pretraining method for molecule joint auto-encoding (MoleculeJAE). MoleculeJAE can learn both the 2D bond (topology) and 3D conformation (geometry) information, and a diffusion process model is applied to mimic the augmented trajectories of such two modalities, based on which, MoleculeJAE will learn the inherent chemical structure in a self-supervised manner. Thus, the pretrained geometrical representation in MoleculeJAE is expected to benefit downstream geometry-related tasks. Empirically, MoleculeJAE proves its effectiveness by reaching state-of-the-art performance on 15 out of 20 tasks by comparing it with 12 competitive baselines.
Weitao Du, Jiujiu Chen 0001, Xuecang Zhang, Zhiming Ma, Shengchao Liu
NeurIPS1
2023 A new perspective on building efficient and expressive 3D equivariant graph neural networks
abstract
Geometric deep learning enables the encoding of physical symmetries in modeling 3D objects. Despite rapid progress in encoding 3D symmetries into Graph Neural Networks (GNNs), a comprehensive evaluation of the expressiveness of these network architectures through a local-to-global analysis lacks today. In this paper, we propose a local hierarchy of 3D isomorphism to evaluate the expressive power of equivariant GNNs and investigate the process of representing global geometric information from local patches. Our work leads to two crucial modules for designing expressive and efficient geometric GNNs; namely local substructure encoding (\textbf{LSE}) and frame transition encoding (\textbf{FTE}). To demonstrate the applicability of our theory, we propose LEFTNet which effectively implements these modules and achieves state-of-the-art performance on both scalar-valued and vector-valued molecular property prediction tasks. We further point out future design space for 3D equivariant graph neural networks. Our codes are available at \url{https://github.com/yuanqidu/LeftNet}.
Weitao Du, Yuanqi Du, Limei Wang, Dieqiao Feng, Shuiwang Ji, Carla P. Gomes, Zhiming Ma
NeurIPS1
2023 Symmetry-Informed Geometric Representation for Molecules, Proteins, and Crystalline Materials
abstract
Artificial intelligence for scientific discovery has recently generated significant interest within the machine learning and scientific communities, particularly in the domains of chemistry, biology, and material discovery. For these scientific problems, molecules serve as the fundamental building blocks, and machine learning has emerged as a highly effective and powerful tool for modeling their geometric structures. Nevertheless, due to the rapidly evolving process of the field and the knowledge gap between science ({\eg}, physics, chemistry, & biology) and machine learning communities, a benchmarking study on geometrical representation for such data has not been conducted. To address such an issue, in this paper, we first provide a unified view of the current symmetry-informed geometric methods, classifying them into three main categories: invariance, equivariance with spherical frame basis, and equivariance with vector frame basis. Then we propose a platform, coined Geom3D, which enables benchmarking the effectiveness of geometric strategies. Geom3D contains 16 advanced symmetry-informed geometric representation models and 14 geometric pretraining methods over 52 diverse tasks, including small molecules, proteins, and crystalline materials. We hope that Geom3D can, on the one hand, eliminate barriers for machine learning researchers interested in exploring scientific problems; and, on the other hand, provide valuable guidance for researchers in computational chemistry, structural biology, and materials science, aiding in the informed selection of representation techniques for specific applications. The source code is available on \href{https://github.com/chao1224/Geom3D}{the GitHub repository}.
Shengchao Liu, Weitao Du, Yanjing Li, Zhuoxinran Li, Zhiling Zheng, Chenru Duan, Zhiming Ma, Omar Yaghi, Anima Anandkumar, Christian Borgs, Jennifer T. Chayes, Jian Tang 0005
NeurIPS2
2022 Towards Deepening Graph Neural Networks: A GNTK-based Optimization Perspective
Wei Huang 0034, Yayong Li, Weitao Du, Jie Yin 0001, Ling Chen 0006, Miao Zhang 0022
ICLR3
2022 SE(3) Equivariant Graph Neural Networks with Complete Local Frames
abstract
Group equivariance (e.g. SE(3) equivariance) is a critical physical symmetry in science, from classical and quantum physics to computational biology. It enables robust and accurate prediction under arbitrary reference transformations. In light of this, great efforts have been put on encoding this symmetry into deep neural networks, which has been shown to improve the generalization performance and data efficiency for downstream tasks. Constructing an equivariant neural network generally brings high computational costs to ensure expressiveness. Therefore, how to better trade-off the expressiveness and computational efficiency plays a core role in the design of the equivariant deep learning models. In this paper, we propose a framework to construct SE(3) equivariant graph neural networks that can approximate the geometric quantities efficiently. Inspired by differential geometry and physics, we introduce equivariant local complete frames to graph neural networks, such that tensor information at given orders can be projected onto the frames. The local frame is constructed to form an orthonormal basis that avoids direction degeneration and ensure completeness. Since the frames are built only by cross product operations, our method is computationally efficient. We evaluate our method on two tasks: Newton mechanics modeling and equilibrium molecule conformation generation. Extensive experimental results demonstrate that our model achieves the best or competitive performance in two types of datasets.
Weitao Du, Yuanqi Du, Wei Chen 0034, Nanning Zheng 0001, Bin Shao 0002, Tie-Yan Liu
ICML1
2021 On the Neural Tangent Kernel of Deep Networks with Orthogonal Initialization
abstract
The prevailing thinking is that orthogonal weights are crucial to enforcing dynamical isometry and speeding up training. The increase in learning speed that results from orthogonal initialization in linear networks has been well-proven. However, while the same is believed to also hold for nonlinear networks when the dynamical isometry condition is satisfied, the training dynamics behind this contention have not been thoroughly explored. In this work, we study the dynamics of ultra-wide networks across a range of architectures, including Fully Connected Networks (FCNs) and Convolutional Neural Networks (CNNs) with orthogonal initialization via neural tangent kernel (NTK). Through a series of propositions and lemmas, we prove that two NTKs, one corresponding to Gaussian weights and one to orthogonal weights, are equal when the network width is infinite. Further, during training, the NTK of an orthogonally-initialized infinite-width network should theoretically remain constant. This suggests that the orthogonal initialization cannot speed up training in the NTK (lazy training) regime, contrary to the prevailing thoughts. In order to explore under what circumstances can orthogonality accelerate training, we conduct a thorough empirical investigation outside the NTK regime. We find that when the hyper-parameters are set to achieve a linear regime in nonlinear activation, orthogonal initialization can improve the learning speed with a large learning rate or large depth.
Wei Huang 0034, Weitao Du
IJCAI2
2020 Mean Field Theory for Deep Dropout Networks: Digging up Gradient Backpropagation Deeply
abstract
In recent years, the mean field theory has been applied to the study of neural networks and has achieved a great deal of success. The theory has been applied to various neural network structures, including CNNs, RNNs, Residual networks, and Batch normalization. Inevitably, recent work has also covered the use of dropout. The mean field theory shows that the existence of depth scales that limit the maximum depth of signal propagation and gradient backpropagation. However, the gradient backpropagation is derived under the gradient independence assumption that weights used during feed forward are drawn independently from the ones used in backpropagation. This is not how neural networks are trained in a real setting. Instead, the same weights used in a feed-forward step needs to be carried over to its corresponding backpropagation. Using this realistic condition, we perform theoretical computation on linear dropout networks and a series of experiments on dropout networks. Our empirical results show an interesting phenomenon that the length gradients can backpropagate for a single input and a pair of inputs are governed by the same depth scale. Besides, we study the relationship between variance and mean of statistical metrics of the gradient and shown an emergence of universality. Finally, we investigate the maximum trainable length for deep dropout networks through a series of experiments using MNIST and CIFAR10 and provide a more precise empirical formula that describes the trainable length than original work.
Wei Huang 0034, Weitao Du, Yutian Zeng, Yunce Zhao
ECAI3