VLDB 2026 Research / reviewers in the wild / expert
Yeming Wen
dblp:217/1501
· DBLP profile ↗
10ranked-venue papers
6as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Trustworthy machine learning · 26% Efficient and distributed learning · 23% Kernel, tree and ensemble methods · 14% | |
| Software engineering, system software, and programming languages
2 papers |
Program synthesis and code generation · 54% Compilers and program optimization · 23% Program analysis · 23% |
Topics — the 25 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
uncertainty estimation |
1.1 | 2 | 2023 | A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023 Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors · ICML 2020 |
Natural language and speech › Language models and text generation
code generation |
1.0 | 2 | 2024 | Synthesize, Partition, then Adapt: Eliciting Diverse Samples from Foundation Models · NeurIPS 2024 Batched Low-Rank Adaptation of Foundation Models · ICLR 2024 |
Machine learning › Kernel, tree and ensemble methods
model ensemble |
0.9 | 2 | 2021 | Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021 BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong Learning · ICLR 2020 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
diverse response generation |
0.8 | 1 | 2024 | Synthesize, Partition, then Adapt: Eliciting Diverse Samples from Foundation Models · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation |
0.8 | 1 | 2024 | Batched Low-Rank Adaptation of Foundation Models · ICLR 2024 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.8 | 1 | 2024 | Batched Low-Rank Adaptation of Foundation Models · ICLR 2024 |
Machine learning › Trustworthy machine learning › uncertainty estimation › predictive uncertainty
distance-aware uncertainty |
0.7 | 1 | 2023 | A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023 |
Program synthesis and code generation
code generation from natural language |
0.7 | 1 | 2023 | Natural Language to Code Generation in Interactive Data Science Notebooks · ACL (1) 2023 |
Machine learning › Trustworthy machine learning
calibration |
0.5 | 1 | 2021 | Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
ensemble calibration |
0.5 | 1 | 2021 | Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021 |
Compilers and program optimization
code generation |
0.5 | 1 | 2021 | Neural Program Generation Modulo Static Analysis · NeurIPS 2021 |
Program synthesis and code generation
neural program synthesis |
0.5 | 1 | 2021 | Neural Program Generation Modulo Static Analysis · NeurIPS 2021 |
Program analysis
static analysis |
0.5 | 1 | 2021 | Neural Program Generation Modulo Static Analysis · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks |
0.4 | 1 | 2020 | Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors · ICML 2020 |
Machine learning › Learning paradigms
lifelong learning |
0.4 | 1 | 2020 | BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong Learning · ICLR 2020 |
Machine learning › Deep learning architectures and training
weight perturbation |
0.3 | 1 | 2018 | Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches · ICLR (Poster) 2018 |
Machine learning › Transfer learning and domain adaptation
model adaptation |
0.2 | 1 | 2024 | Synthesize, Partition, then Adapt: Eliciting Diverse Samples from Foundation Models · NeurIPS 2024 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
multilingual speech recognition |
0.2 | 1 | 2024 | Batched Low-Rank Adaptation of Foundation Models · ICLR 2024 |
Machine learning › Trustworthy machine learning › calibration
neural network calibration |
0.2 | 1 | 2023 | A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.2 | 1 | 2023 | A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023 |
Machine learning › Deep learning architectures and training
data augmentation |
0.1 | 1 | 2021 | Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021 |
Machine learning › Deep learning architectures and training
transformer |
0.1 | 1 | 2021 | Neural Program Generation Modulo Static Analysis · NeurIPS 2021 |
Machine learning › Efficient and distributed learning
parameter-efficient model |
0.1 | 1 | 2020 | Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors · ICML 2020 |
Machine learning › Graph learning › graph neural network training
mini-batch training |
0.1 | 1 | 2018 | Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches · ICLR (Poster) 2018 |
Machine learning › Optimization for machine learning
stochastic optimization |
0.1 | 1 | 2018 | Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches · ICLR (Poster) 2018 |
Methods — techniques the papers use, named apart from their topics
variational inference · 0.8synthetic data · 0.8low-rank adaptation · 0.8influence functions · 0.8data attribution · 0.8batching · 0.8spectral normalization · 0.7program synthesis · 0.7minimax learning · 0.7large language model · 0.7gaussian process · 0.7static program analysis · 0.5neurosymbolic learning · 0.5deep generative model · 0.5data augmentation · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Batched Low-Rank Adaptation of Foundation ModelsabstractLow-Rank Adaptation (LoRA) has recently gained attention for fine-tuning foundation models by incorporating trainable low-rank matrices, thereby reducing the number of trainable parameters. While \lora/ offers numerous advantages, its applicability for real-time serving to a diverse and global user base
is constrained by its incapability to handle multiple task-specific adapters efficiently. This imposes a performance bottleneck in scenarios requiring personalized, task-specific adaptations for each incoming request.
To address this, we introduce FLoRA (Fast LoRA), a framework in which each input example in a minibatch can be associated with its unique low-rank adaptation weights, allowing for efficient batching of heterogeneous requests. We empirically demonstrate that \flora/ retains the performance merits of \lora/, showcasing competitive results on the MultiPL-E code generation benchmark spanning over 8 languages and a multilingual speech recognition task across 6 languages. Yeming Wen, Swarat Chaudhuri |
ICLR | 1 |
| 2024 | Synthesize, Partition, then Adapt: Eliciting Diverse Samples from Foundation ModelsabstractPresenting users with diverse responses from foundation models is crucial for enhancing user experience and accommodating varying preferences.
However, generating multiple high-quality and diverse responses without sacrificing accuracy remains a challenge, especially when using greedy sampling.
In this work, we propose a novel framework, Synthesize-Partition-Adapt (SPA), that leverages the abundant synthetic data available in many domains to elicit diverse responses from foundation models.
By leveraging signal provided by data attribution methods such as influence functions, SPA partitions data into subsets, each targeting unique aspects of the data, and trains multiple model adaptations optimized for these subsets.
Experimental results demonstrate the effectiveness of our approach in diversifying foundation model responses while maintaining high quality, showcased through the HumanEval and MBPP tasks in the code generation domain and several tasks in the natural language understanding domain, highlighting its potential to enrich user experience across various applications. Yeming Wen, Swarat Chaudhuri |
NeurIPS | 1 |
| 2023 | Natural Language to Code Generation in Interactive Data Science NotebooksabstractPengcheng Yin, Wen-Ding Li, Kefan Xiao, Abhishek Rao, Yeming Wen, Kensen Shi, Joshua Howland, Paige Bailey, Michele Catasta, Henryk Michalewski, Oleksandr Polozov, Charles Sutton. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Wen-Ding Li, Kefan Xiao, Abhishek Rao, Yeming Wen, Kensen Shi, Joshua Howland, Paige Bailey, Michele Catasta, Henryk Michalewski, Oleksandr Polozov, Charles Sutton |
ACL (1) | 5 |
| 2023 | A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-AwarenessabstractAccurate uncertainty quantification is a major challenge in deep learning, as neural networks can make overconfident errors and assign high confidence predictions to out-of-distribution (OOD) inputs. The most popular approaches to estimate predictive uncertainty in deep learning are methods that combine predictions from multiple neural networks, such as Bayesian neural networks (BNNs) and deep ensembles. However their practicality in real-time, industrial-scale applications are limited due to the high memory and computational cost. Furthermore, ensembles and BNNs do not necessarily fix all the issues with the underlying member networks. In this work, we study principled approaches to improve the uncertainty property of a single network, based on a single, deterministic representation. By formalizing the uncertainty quantification as a minimax learning problem, we first identify distance awareness, i.e., the model's ability to quantify the distance of a testing example from the training data, as a necessary condition for a DNN to achieve high-quality (i.e., minimax optimal) uncertainty estimation. We then propose Spectral-normalized Neural Gaussian Process (SNGP), a simple method that improves the distance-awareness ability of modern DNNs with two simple changes: (1) applying spectral normalization to hidden weights to enforce bi-Lipschitz smoothness in representations and (2) replacing the last output layer with a Gaussian process layer. On a suite of vision and language understanding benchmarks and on modern architectures (Wide-ResNet and BERT), SNGP consistently outperforms other single-model approaches in prediction, calibration and out-of-domain detection. Furthermore, SNGP provides complementary benefits to popular techniques such as deep ensembles and data augmentation, making it a simple and scalable building block for probabilistic deep learning. Jeremiah Z. Liu, Shreyas Padhy, Jie Ren 0006, Zi Lin, Yeming Wen, Ghassen Jerfel, Zachary Nado, Jasper Snoek, Dustin Tran, Balaji Lakshminarayanan |
J. Mach. Learn. Res. | 5 |
| 2021 | Combining Ensembles and Data Augmentation Can Harm Your Calibration
Yeming Wen, Ghassen Jerfel, Rafael Muller, Michael Dusenberry, Jasper Snoek, Balaji Lakshminarayanan, Dustin Tran |
ICLR | 1 |
| 2021 | Neural Program Generation Modulo Static AnalysisabstractState-of-the-art neural models of source code tend to be evaluated on the generation of individual expressions and lines of code, and commonly fail on long-horizon tasks such as the generation of entire method bodies. We propose to address this deficiency using weak supervision from a static program analyzer. Our neurosymbolic method allows a deep generative model to symbolically compute, using calls to a static analysis tool, long-distance semantic relationships in the code that it has already generated. During training, the model observes these relationships and learns to generate programs conditioned on them. We apply our approach to the problem of generating entire Java methods given the remainder of the class that contains the method. Our experiments show that the approach substantially outperforms a state-of-the-art transformer and a model that explicitly tries to learn program semantics on this task, both in terms of producing programs free of basic semantic errors and in terms of syntactically matching the ground truth. Rohan Mukherjee 0001, Yeming Wen, Dipak Chaudhari, Thomas W. Reps, Swarat Chaudhuri, Chris Jermaine |
NeurIPS | 2 |
| 2020 | An Empirical Study of Stochastic Gradient Descent with Structured Covariance NoiseabstractThe choice of batch-size in a stochastic optimization algorithm plays a substantial role for both optimization and generalization. Increasing the batch-size used typically improves optimization but degrades generalization. To address the problem of improving generalization while maintaining optimal convergence in large-batch training, we propose to add covariance noise to the gradients. We demonstrate that the learning performance of our method is more accurately captured by the structure of the covariance matrix of the noise rather than by the variance of gradients. Moreover, over the convex-quadratic, we prove in theory that it can be characterized by the Frobenius norm of the noise matrix. Our empirical studies with standard deep learning model-architectures and datasets shows that our method not only improves generalization performance in large-batch training, but furthermore, does so in a way where the optimization performance remains desirable and the training duration is not elongated. Yeming Wen, Kevin Luk, Maxime Gazeau, Guodong Zhang 0006, Harris Chan, Jimmy Ba |
AISTATS | 1 |
| 2020 | BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong Learning
Yeming Wen, Dustin Tran, Jimmy Ba |
ICLR | 1 |
| 2020 | Efficient and Scalable Bayesian Neural Nets with Rank-1 FactorsabstractBayesian neural networks (BNNs) demonstrate promising success in improving the robustness and uncertainty quantification of modern deep learning. However, they generally struggle with underfitting at scale and parameter efficiency. On the other hand, deep ensembles have emerged as alternatives for uncertainty quantification that, while outperforming BNNs on certain problems, also suffer from efficiency issues. It remains unclear how to combine the strengths of these two approaches and remediate their common issues. To tackle this challenge, we propose a rank-1 parameterization of BNNs, where each weight matrix involves only a distribution on a rank-1 subspace. We also revisit the use of mixture approximate posteriors to capture multiple modes, where unlike typical mixtures, this approach admits a significantly smaller memory increase (e.g., only a 0.4% increase for a ResNet-50 mixture of size 10). We perform a systematic empirical study on the choices of prior, variational posterior, and methods to improve training. For ResNet-50 on ImageNet, Wide ResNet 28-10 on CIFAR-10/100, and an RNN on MIMIC-III, rank-1 BNNs achieve state-of-the-art performance across log-likelihood, accuracy, and calibration on the test sets and out-of-distribution variants. Michael Dusenberry, Ghassen Jerfel, Yeming Wen, Yi-An Ma, Jasper Snoek, Katherine A. Heller, Balaji Lakshminarayanan, Dustin Tran |
ICML | 3 |
| 2018 | Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches
Yeming Wen, Paul Vicol, Jimmy Ba, Dustin Tran, Roger B. Grosse |
ICLR (Poster) | 1 |