VLDB 2026 Research / reviewers in the wild / expert
Weinan E
dblp:06/9390
· DBLP profile ↗
21ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0003-0272-9500ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 1 first-author · 10 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PaSa: An LLM Agent for Comprehensive Academic Paper SearchabstractYichen He, Guanhua Huang, Peiyuan Feng, Yuan Lin, Yuchen Zhang, Hang Li, Weinan E. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Guanhua Huang, Peiyuan Feng, Weinan E |
ACL (1) | 7 |
| 2025 | The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-TrainingabstractTransformers have become the cornerstone of modern AI. Unlike traditional architectures, transformers exhibit a distinctive characteristic: diverse types of building blocks, such as embedding layers, normalization layers, self-attention mechanisms, and point-wise feed-forward networks, work collaboratively. Understanding the disparities and interactions among these blocks is therefore important.
In this paper, we uncover a clear **sharpness disparity** across these blocks, which intriguingly emerges early in training and persists throughout the training process.
Building on this insight, we propose a novel **Blockwise Learning Rate (LR)** strategy to accelerate large language model (LLM) pre-training. Specifically, by integrating Blockwise LR into AdamW, we consistently achieve lower terminal loss and nearly $2\times$ speedup compared to vanilla AdamW. This improvement is demonstrated across GPT-2 and LLaMA models, with model sizes ranging from 0.12B to 1.1B and datasets including OpenWebText and MiniPile.
Finally, we incorporate Blockwise LR into Adam-mini (Zhang et al., 2024), a recently proposed memory-efficient variant of Adam, achieving a combined $2\times$ speedup and $2\times$ memory savings. These results underscore the potential of leveraging the sharpness disparity principle to improve LLM training. Jinbo Wang 0003, Zhanpeng Zhou, Junchi Yan, Weinan E |
ICML | 5 |
| 2025 | On the Expressive Power of Mixture-of-Experts for Structured Complex TasksabstractMixture-of-experts networks (MoEs) have demonstrated remarkable efficiency in modern deep learning. Despite their empirical success, the theoretical foundations underlying their ability to model complex tasks remain poorly understood.
In this work, we conduct a systematic study of the expressive power of MoEs in modeling complex tasks with two common structural priors: low-dimensionality and sparsity.
For shallow MoEs, we prove that they can efficiently approximate functions supported on low-dimensional manifolds, overcoming the curse of dimensionality.
For deep MoEs, we show that $\mathcal{O}(L)$-layer MoEs with $E$ experts per layer can approximate piecewise functions comprising $E^L$ pieces with compositional sparsity, i.e., they can exhibit an exponential number of structured tasks.
Our analysis reveals the roles of critical architectural components and hyperparameters in MoEs, including the gating mechanism, expert networks, the number of experts, and the number of layers, and offers natural suggestions for MoE variants. Weinan E |
NeurIPS | 2 |
| 2024 | Exploring Molecular Pretraining Model at ScaleabstractIn recent years, pretraining models have made significant advancements in the fields of natural language processing (NLP), computer vision (CV), and life sciences. The significant advancements in NLP and CV are predominantly driven by the expansion of model parameters and data size, a phenomenon now recognized as the scaling laws. However, research exploring scaling law in molecular pretraining model remains unexplored. In this work, we present an innovative molecular pretraining model that leverages a two-track transformer to effectively integrate features at the atomic level, graph level, and geometry structure level. Along with this, we systematically investigate the scaling law within molecular pretraining models, examining the power-law correlations between validation loss and model size, dataset size, and computational resources. Consequently, we successfully scale the model to 1.1 billion parameters through pretraining on 800 million conformations, making it the largest molecular pretraining model to date. Extensive experiments show the consistent improvement on the downstream tasks as the model size grows up. The model with 1.1 billion parameters also outperform over existing methods, achieving an average 27\% improvement on the QM9 and 14\% on COMPAS-1D dataset. Xiaohong Ji, Zhifeng Gao, Linfeng Zhang 0002, Guolin Ke, Weinan E |
NeurIPS | 7 |
| 2024 | Improving Generalization and Convergence by Enhancing Implicit RegularizationabstractIn this work, we propose an Implicit Regularization Enhancement (IRE) framework to accelerate the discovery of flat solutions in deep learning, thereby improving generalization and convergence.
Specifically, IRE decouples the dynamics of flat and sharp directions, which boosts the sharpness reduction along flat directions while maintaining the training stability in sharp directions. We show that IRE can be practically incorporated with *generic base optimizers* without introducing significant computational overload. Experiments show that IRE consistently improves the generalization performance for image classification tasks across a variety of benchmark datasets (CIFAR-10/100, ImageNet) and models (ResNets and ViTs).
Surprisingly, IRE also achieves a $2\times$ *speed-up* compared to AdamW in the pre-training of Llama models (of sizes ranging from 60M to 229M) on datasets including Wikitext-103, Minipile, and Openwebtext. Moreover, we provide theoretical guarantees, showing that IRE can substantially accelerate the convergence towards flat minima in Sharpness-aware Minimization (SAM). Jinbo Wang 0003, Haotian He, Guanhua Huang, Feiyu Xiong, Weinan E |
NeurIPS | 8 |
| 2024 | Understanding the Expressive Power and Mechanisms of Transformer for Sequence ModelingabstractWe conduct a systematic study of the approximation properties of Transformer for sequence modeling with long, sparse and complicated memory.
We investigate the mechanisms through which different components of Transformer, such as the dot-product self-attention, positional encoding and feed-forward layer, affect its expressive power, and we study their combined effects through establishing explicit approximation rates.
Our study reveals the roles of critical parameters in the Transformer, such as the number of layers and the number of attention heads.
These theoretical insights are validated experimentally and offer natural suggestions for alternative architectures. Weinan E |
NeurIPS | 2 |
| 2023 | An Iteratively Parallel Generation Method with the Pre-Filling Strategy for Document-level Event ExtractionabstractIn document-level event extraction (DEE) tasks, a document typically contains many event records with multiple event roles.Therefore, accurately extracting all event records is a big challenge since the number of event records is not given.Previous works present the entitybased directed acyclic graph (EDAG) generation methods to autoregressively generate event roles, which requires a given generation order.Meanwhile, parallel methods are proposed to generate all event roles simultaneously, but suffer from the inadequate training which manifests zero accuracies on some event roles.In this paper, we propose an Iteratively Parallel Generation method with the Pre-Filling strategy (IPGPF).Event roles in an event record are generated in parallel to avoid order selection, and the event records are iteratively generated to utilize historical results.Experiments on two public datasets show our IPGPF improves 11.7 F1 than previous parallel models and up to 5.1 F1 than auto-regressive models under the control variable settings.Moreover, our enhanced IPGPF outperforms other entityenhanced models and achieves new state-ofthe-art performance 1 .* Work was done when Guanhua was an intern at ByteDance AI Lab.† Corresponding author. 1 Our code is available at https://github.com/ CarlanLark/IPGPF [S6] …, Jinggong Group increased its holdings of the company's stock by 182,038 shares through the secondary market on Dec 15, 2011,… [S7] …, the shares held by Jinggong Group in the company increased from 90,880,020 shares to 91,062,058 shares, … [S9] on Dec 16, 2011, Jinggong Group reduced its holdings of ... 35,000 shares, with an average price of 19.88.[S14] As of the date of this announcement, Jinggong Group holds 91,027,058 shares of the company, … EquityOverweight EquityHolder Jinggong Group Guanhua Huang, Runxin Xu, Jiaze Chen, Zhouwang Yang, Weinan E |
EMNLP | 6 |
| 2022 | Approximation and Optimization Theory for Linear Continuous-Time Recurrent Neural NetworksabstractWe perform a systematic study of the approximation properties and optimization dynamics of recurrent neural networks (RNNs) when applied to learn input-output relationships in temporal data. We consider the simple but representative setting of using continuous-time linear RNNs to learn from data generated by linear relationships. On the approximation side, we prove a direct and an inverse approximation theorem of linear functionals using RNNs, which reveal the intricate connections between memory structures in the target and the corresponding approximation efficiency. In particular, we show that temporal relationships can be effectively approximated by RNNs if and only if the former possesses sufficient memory decay. On the optimization front, we perform detailed analysis of the optimization dynamics, including a precise understanding of the difficulty that may arise in learning relationships with long-term memory. The term “curse of memory” is coined to describe the uncovered phenomena, akin to the “curse of dimension” that plagues high-dimensional function approximation. These results form a relatively complete picture of the interaction of memory and recurrent structures in the linear dynamical setting. Zhong Li 0004, Jiequn Han, Weinan E, Qianxiao Li |
J. Mach. Learn. Res. | 3 |
| 2022 | A Mathematical Model for Universal SemanticsabstractWe characterize the meaning of words with language-independent numerical fingerprints, through a mathematical analysis of recurring patterns in texts. Approximating texts by Markov processes on a long-range time scale, we are able to extract topics, discover synonyms, and sketch semantic fields from a particular document of moderate length, without consulting external knowledge-base or thesaurus. Our Markov semantic model allows us to represent each topical concept by a low-dimensional vector, interpretable as algebraic invariants in succinct statistical operations on the document, targeting local environments of individual words. These language-independent semantic representations enable a robot reader to both understand short texts in a given language (automated question-answering) and match medium-length texts across different languages (automated word translation). Our semantic fingerprints quantify local meaning of words in 14 representative languages across five major language families, suggesting a universal and cost-effective mechanism by which human languages are processed at the semantic level. Our protocols and source codes are publicly available on https://github.com/yajun-zhou/linguae-naturalis-principia-mathematica. Weinan E, Yajun Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
Zhong Li 0004, Jiequn Han, Weinan E, Qianxiao Li |
ICLR | 3 |
| 2020 | A Priori Estimates of the Generalization Error for AutoencodersabstractAutoencoder is a machine learning model which aims for dimensionality reduction, by reconstructing its input through a bottleneck with lower dimension than the input. It is among the most popular models used in unsupervised learning and semi-supervised learning. In this paper, we build theoretical understanding about autoencoders. Specifically, assuming the existence of the underlying groundtruth encoder and decoder, we establish a priori estimates of the generalization error for autoencoders when an appropriately chosen regularization term is applied. The estimate is a priori in the sense that it only depend on some norms of the groundtruth encoder and decoder, but not the model parameters. The bound acheives nearly optimal rates with respect to the number of data and parameters. To our knowledge, this is the first try to build a priori estimates to unsupervised learning models. Numerical experiments show the tightness of the bounds. Zehao Don, Weinan E, Chao Ma 0012 |
ICASSP | 2 |
| 2020 | Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningabstractIt is not clear yet why ADAM-alike adaptive gradient algorithms suffer from worse generalization performance than SGD despite their faster training speed. This work aims to provide understandings on this generalization gap by analyzing their local convergence behaviors. Specifically, we observe the heavy tails of gradient noise in these algorithms. This motivates us to analyze these algorithms through their Levy-driven stochastic differential equations (SDEs) because of the similar convergence behaviors of an algorithm and its SDE. Then we establish the escaping time of these SDEs from a local basin. The result shows that (1) the escaping time of both SGD and ADAM~depends on the Radon measure of the basin positively and the heaviness of gradient noise negatively; (2) for the same basin, SGD enjoys smaller escaping time than ADAM, mainly because (a) the geometry adaptation in ADAM~via adaptively scaling each gradient coordinate well diminishes the anisotropic structure in gradient noise and results in larger Radon measure of a basin; (b) the exponential gradient average in ADAM~smooths its gradient and leads to lighter gradient noise tails than SGD. So SGD is more locally unstable than ADAM~at sharp minima defined as the minima whose local basins have small Radon measure, and can better escape from them to flatter ones with larger Radon measure. As flat minima here which often refer to the minima at flat or asymmetric basins/valleys often generalize better than sharp ones~\cite{keskar2016large,he2019asymmetric}, our result explains the better generalization performance of SGD over ADAM. Finally, experimental results confirm our heavy-tailed gradient noise assumption and theoretical affirmation. Pan Zhou 0002, Jiashi Feng, Chao Ma 0012, Caiming Xiong, Steven C. H. Hoi, Weinan E |
NeurIPS | 6 |
| 2020 | Pushing the limit of molecular dynamics with ab initio accuracy to 100 million atoms with machine learningabstractFor 35 years, ab initio molecular dynamics (AIMD) has been the method of choice for modeling complex atomistic phenomena from first principles. However, most AIMD applications are limited by computational cost to systems with thousands of atoms at most. We report that a machine learning based simulation protocol (Deep Potential Molecular Dynamics), while retaining ab initio accuracy, can simulate more than 1 nanosecond-long trajectory of over 100 million atoms per day, using a highly optimized code (GPU DeePMD-kit) on the Summit supercomputer. Our code can efficiently scale up to the entire Summit supercomputer, attaining 91 PFLOPS in double precision (45.5% of the peak) and 162/275 PFLOPS in mixed-single/half precision. The great accomplishment of this work is that it opens the door to simulating unprecedented size and time scales with ab initio accuracy. It also poses new challenges to the next-generation supercomputer for a better integration of machine learning and physical modeling. Weile Jia, Han Wang 0006, Mohan Chen 0002, Denghui Lu, Lin Lin 0001, Roberto Car, Weinan E, Linfeng Zhang 0002 |
SC | 7 |
| 2019 | Stochastic Modified Equations and Dynamics of Stochastic Gradient Algorithms I: Mathematical FoundationsabstractWe develop the mathematical foundations of the stochastic modified equations (SME) framework for analyzing the dynamics of stochastic gradient algorithms, where the latter is approximated by a class of stochastic differential equations with small noise parameters. We prove that this approximation can be understood mathematically as an weak approximation, which leads to a number of precise and useful results on the approximations of stochastic gradient descent (SGD), momentum SGD and stochastic Nesterov's accelerated gradient method in the general setting of stochastic objectives. We also demonstrate through explicit calculations that this continuous-time approach can uncover important analytical insights into the stochastic gradient algorithms under consideration that may not be easy to obtain in a purely discrete-time setting. Qianxiao Li, Cheng Tai, Weinan E |
J. Mach. Learn. Res. | 3 |
| 2018 | How SGD Selects the Global Minima in Over-parameterized Learning: A Dynamical Stability PerspectiveabstractThe question of which global minima are accessible by a stochastic gradient decent (SGD) algorithm with specific learning rate and batch size is studied from the perspective of dynamical stability. The concept of non-uniformity is introduced, which, together with sharpness, characterizes the stability property of a global minimum and hence the accessibility of a particular SGD algorithm to that global minimum. In particular, this analysis shows that learning rate and batch size play different roles in minima selection. Extensive empirical results seem to correlate well with the theoretical findings and provide further support to these claims. Chao Ma 0012, Weinan E |
NeurIPS | 3 |
| 2018 | End-to-end Symmetry Preserving Inter-atomic Potential Energy Model for Finite and Extended SystemsabstractMachine learning models are changing the paradigm of molecular modeling, which is a fundamental tool for material science, chemistry, and computational biology. Of particular interest is the inter-atomic potential energy surface (PES). Here we develop Deep Potential - Smooth Edition (DeepPot-SE), an end-to-end machine learning-based PES model, which is able to efficiently represent the PES for a wide variety of systems with the accuracy of ab initio quantum mechanics models. By construction, DeepPot-SE is extensive and continuously differentiable, scales linearly with system size, and preserves all the natural symmetries of the system. Further, we show that DeepPot-SE describes finite and extended systems including organic molecules, metals, semiconductors, and insulators with high fidelity. Linfeng Zhang 0002, Jiequn Han, Han Wang 0006, Wissam Saidi, Roberto Car, Weinan E |
NeurIPS | 6 |
| 2017 | Stochastic Modified Equations and Adaptive Stochastic Gradient AlgorithmsabstractWe develop the method of stochastic modified equations (SME), in which stochastic gradient algorithms are approximated in the weak sense by continuous-time stochastic differential equations. We exploit the continuous formulation together with optimal control theory to derive novel adaptive hyper-parameter adjustment policies. Our algorithms have competitive performance with the added benefit of being robust to varying models and datasets. This provides a general methodology for the analysis and design of stochastic gradient algorithms. Qianxiao Li, Cheng Tai, Weinan E |
ICML | 3 |
| 2017 | Joint Learning of Response Ranking and Next Utterance Suggestion in Human-Computer Conversation SystemabstractConversation systems are of growing importance since they enable an easy interaction interface between humans and computers: using natural languages. To build a conversation system with adequate intelligence is challenging, and requires abundant resources including an acquisition of big data and interdisciplinary techniques, such as information retrieval and natural language processing. Along with the prosperity of Web 2.0, the massive data available greatly facilitate data-driven methods such as deep learning for human-computer conversation systems. Owing to the diversity of Web resources, a retrieval-based conversation system will come up with at least some results from the immense repository for any user inputs. Given a human issued message, i.e., query, a traditional conversation system would provide a response after adequate training and learning of how to respond. In this paper, we propose a new task for conversation systems: joint learning of response ranking featured with next utterance suggestion. We assume that the new conversation mode is more proactive and keeps user engaging. We examine the assumption in experiments. Besides, to address the joint learning task, we propose a novel Dual-LSTM Chain Model to couple response ranking and next utterance suggestion simultaneously. From the experimental results, we demonstrate the usefulness of the proposed task and the effectiveness of the proposed model. Rui Yan 0001, Dongyan Zhao 0001, Weinan E |
SIGIR | 3 |
| 2017 | Maximum Principle Based Algorithms for Deep Learning
Qianxiao Li, Cheng Tai, Weinan E |
J. Mach. Learn. Res. | 4 |
| 2016 | Multiscale Adaptive Representation of Signals: I. The Basic FrameworkabstractWe introduce a framework for designing multi-scale, adaptive, shift-invariant frames and bi-frames for representing signals. The new framework, called AdaFrame, improves over dictionary learning-based techniques in terms of computational efficiency at inference time. It improves classical multi-scale basis such as wavelet frames in terms of coding efficiency. It provides an attractive alternative to dictionary learning-based techniques for low level signal processing tasks, such as compression and denoising, as well as high level tasks, such as feature extraction for object recognition. Connections with deep convolutional networks are also discussed. In particular, the proposed framework reveals a drawback in the commonly used approach for visualizing the activations of the intermediate layers in convolutional networks, and suggests a natural alternative. Cheng Tai, Weinan E |
J. Mach. Learn. Res. | 2 |
| 2011 | SelInv - An Algorithm for Selected Inversion of a Sparse Symmetric MatrixabstractWe describe an efficient implementation of an algorithm for computing selected elements of a general sparse symmetric matrix A that can be decomposed as A = LDLT , where L is lower triangular and D is diagonal. Our implementation, which is called SelInv , is built on top of an efficient supernodal left-looking LDLT factorization of A . We discuss how computational efficiency can be gained by making use of a relative index array to handle indirect addressing. We report the performance of SelInv on a collection of sparse matrices of various sizes and nonzero structures. We also demonstrate how SelInv can be used in electronic structure calculations. Lin Lin 0001, Chao Yang 0001, Juan C. Meza, Jianfeng Lu 0001, Lexing Ying, Weinan E |
ACM Trans. Math. Softw. | 6 |