Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yeming Wen

dblp:217/1501 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Trustworthy machine learning · 26% Efficient and distributed learning · 23% Kernel, tree and ensemble methods · 14%
Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 54% Compilers and program optimization · 23% Program analysis · 23%

Topics — the 25 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
uncertainty estimation
1.122023
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023
Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors · ICML 2020
Natural language and speech › Language models and text generation
code generation
1.022024
Synthesize, Partition, then Adapt: Eliciting Diverse Samples from Foundation Models · NeurIPS 2024
Batched Low-Rank Adaptation of Foundation Models · ICLR 2024
Machine learning › Kernel, tree and ensemble methods
model ensemble
0.922021
Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021
BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong Learning · ICLR 2020
Natural language and speech › Question answering and dialogue systems › dialogue generation
diverse response generation
0.812024
Synthesize, Partition, then Adapt: Eliciting Diverse Samples from Foundation Models · NeurIPS 2024
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.812024
Batched Low-Rank Adaptation of Foundation Models · ICLR 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.812024
Batched Low-Rank Adaptation of Foundation Models · ICLR 2024
Machine learning › Trustworthy machine learning › uncertainty estimation › predictive uncertainty
distance-aware uncertainty
0.712023
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023
Program synthesis and code generation
code generation from natural language
0.712023
Natural Language to Code Generation in Interactive Data Science Notebooks · ACL (1) 2023
Machine learning › Trustworthy machine learning
calibration
0.512021
Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021
Machine learning › Kernel, tree and ensemble methods › ensemble learning
ensemble calibration
0.512021
Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021
Compilers and program optimization
code generation
0.512021
Neural Program Generation Modulo Static Analysis · NeurIPS 2021
Program synthesis and code generation
neural program synthesis
0.512021
Neural Program Generation Modulo Static Analysis · NeurIPS 2021
Program analysis
static analysis
0.512021
Neural Program Generation Modulo Static Analysis · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.412020
Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors · ICML 2020
Machine learning › Learning paradigms
lifelong learning
0.412020
BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong Learning · ICLR 2020
Machine learning › Deep learning architectures and training
weight perturbation
0.312018
Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches · ICLR (Poster) 2018
Machine learning › Transfer learning and domain adaptation
model adaptation
0.212024
Synthesize, Partition, then Adapt: Eliciting Diverse Samples from Foundation Models · NeurIPS 2024
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
multilingual speech recognition
0.212024
Batched Low-Rank Adaptation of Foundation Models · ICLR 2024
Machine learning › Trustworthy machine learning › calibration
neural network calibration
0.212023
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.212023
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023
Machine learning › Deep learning architectures and training
data augmentation
0.112021
Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021
Machine learning › Deep learning architectures and training
transformer
0.112021
Neural Program Generation Modulo Static Analysis · NeurIPS 2021
Machine learning › Efficient and distributed learning
parameter-efficient model
0.112020
Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors · ICML 2020
Machine learning › Graph learning › graph neural network training
mini-batch training
0.112018
Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches · ICLR (Poster) 2018
Machine learning › Optimization for machine learning
stochastic optimization
0.112018
Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches · ICLR (Poster) 2018

Methods — techniques the papers use, named apart from their topics

variational inference · 0.8synthetic data · 0.8low-rank adaptation · 0.8influence functions · 0.8data attribution · 0.8batching · 0.8spectral normalization · 0.7program synthesis · 0.7minimax learning · 0.7large language model · 0.7gaussian process · 0.7static program analysis · 0.5neurosymbolic learning · 0.5deep generative model · 0.5data augmentation · 0.5
YearPublicationVenuePosition
2024 Batched Low-Rank Adaptation of Foundation Models
abstract
Low-Rank Adaptation (LoRA) has recently gained attention for fine-tuning foundation models by incorporating trainable low-rank matrices, thereby reducing the number of trainable parameters. While \lora/ offers numerous advantages, its applicability for real-time serving to a diverse and global user base is constrained by its incapability to handle multiple task-specific adapters efficiently. This imposes a performance bottleneck in scenarios requiring personalized, task-specific adaptations for each incoming request. To address this, we introduce FLoRA (Fast LoRA), a framework in which each input example in a minibatch can be associated with its unique low-rank adaptation weights, allowing for efficient batching of heterogeneous requests. We empirically demonstrate that \flora/ retains the performance merits of \lora/, showcasing competitive results on the MultiPL-E code generation benchmark spanning over 8 languages and a multilingual speech recognition task across 6 languages.
Yeming Wen, Swarat Chaudhuri
ICLR1
2024 Synthesize, Partition, then Adapt: Eliciting Diverse Samples from Foundation Models
abstract
Presenting users with diverse responses from foundation models is crucial for enhancing user experience and accommodating varying preferences. However, generating multiple high-quality and diverse responses without sacrificing accuracy remains a challenge, especially when using greedy sampling. In this work, we propose a novel framework, Synthesize-Partition-Adapt (SPA), that leverages the abundant synthetic data available in many domains to elicit diverse responses from foundation models. By leveraging signal provided by data attribution methods such as influence functions, SPA partitions data into subsets, each targeting unique aspects of the data, and trains multiple model adaptations optimized for these subsets. Experimental results demonstrate the effectiveness of our approach in diversifying foundation model responses while maintaining high quality, showcased through the HumanEval and MBPP tasks in the code generation domain and several tasks in the natural language understanding domain, highlighting its potential to enrich user experience across various applications.
Yeming Wen, Swarat Chaudhuri
NeurIPS1
2023 Natural Language to Code Generation in Interactive Data Science Notebooks
abstract
Pengcheng Yin, Wen-Ding Li, Kefan Xiao, Abhishek Rao, Yeming Wen, Kensen Shi, Joshua Howland, Paige Bailey, Michele Catasta, Henryk Michalewski, Oleksandr Polozov, Charles Sutton. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Wen-Ding Li, Kefan Xiao, Abhishek Rao, Yeming Wen, Kensen Shi, Joshua Howland, Paige Bailey, Michele Catasta, Henryk Michalewski, Oleksandr Polozov, Charles Sutton
ACL (1)5
2023 A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness
abstract
Accurate uncertainty quantification is a major challenge in deep learning, as neural networks can make overconfident errors and assign high confidence predictions to out-of-distribution (OOD) inputs. The most popular approaches to estimate predictive uncertainty in deep learning are methods that combine predictions from multiple neural networks, such as Bayesian neural networks (BNNs) and deep ensembles. However their practicality in real-time, industrial-scale applications are limited due to the high memory and computational cost. Furthermore, ensembles and BNNs do not necessarily fix all the issues with the underlying member networks. In this work, we study principled approaches to improve the uncertainty property of a single network, based on a single, deterministic representation. By formalizing the uncertainty quantification as a minimax learning problem, we first identify distance awareness, i.e., the model's ability to quantify the distance of a testing example from the training data, as a necessary condition for a DNN to achieve high-quality (i.e., minimax optimal) uncertainty estimation. We then propose Spectral-normalized Neural Gaussian Process (SNGP), a simple method that improves the distance-awareness ability of modern DNNs with two simple changes: (1) applying spectral normalization to hidden weights to enforce bi-Lipschitz smoothness in representations and (2) replacing the last output layer with a Gaussian process layer. On a suite of vision and language understanding benchmarks and on modern architectures (Wide-ResNet and BERT), SNGP consistently outperforms other single-model approaches in prediction, calibration and out-of-domain detection. Furthermore, SNGP provides complementary benefits to popular techniques such as deep ensembles and data augmentation, making it a simple and scalable building block for probabilistic deep learning.
Jeremiah Z. Liu, Shreyas Padhy, Jie Ren 0006, Zi Lin, Yeming Wen, Ghassen Jerfel, Zachary Nado, Jasper Snoek, Dustin Tran, Balaji Lakshminarayanan
J. Mach. Learn. Res.5
2021 Combining Ensembles and Data Augmentation Can Harm Your Calibration
Yeming Wen, Ghassen Jerfel, Rafael Muller, Michael Dusenberry, Jasper Snoek, Balaji Lakshminarayanan, Dustin Tran
ICLR1
2021 Neural Program Generation Modulo Static Analysis
abstract
State-of-the-art neural models of source code tend to be evaluated on the generation of individual expressions and lines of code, and commonly fail on long-horizon tasks such as the generation of entire method bodies. We propose to address this deficiency using weak supervision from a static program analyzer. Our neurosymbolic method allows a deep generative model to symbolically compute, using calls to a static analysis tool, long-distance semantic relationships in the code that it has already generated. During training, the model observes these relationships and learns to generate programs conditioned on them. We apply our approach to the problem of generating entire Java methods given the remainder of the class that contains the method. Our experiments show that the approach substantially outperforms a state-of-the-art transformer and a model that explicitly tries to learn program semantics on this task, both in terms of producing programs free of basic semantic errors and in terms of syntactically matching the ground truth.
Rohan Mukherjee 0001, Yeming Wen, Dipak Chaudhari, Thomas W. Reps, Swarat Chaudhuri, Chris Jermaine
NeurIPS2
2020 An Empirical Study of Stochastic Gradient Descent with Structured Covariance Noise
abstract
The choice of batch-size in a stochastic optimization algorithm plays a substantial role for both optimization and generalization. Increasing the batch-size used typically improves optimization but degrades generalization. To address the problem of improving generalization while maintaining optimal convergence in large-batch training, we propose to add covariance noise to the gradients. We demonstrate that the learning performance of our method is more accurately captured by the structure of the covariance matrix of the noise rather than by the variance of gradients. Moreover, over the convex-quadratic, we prove in theory that it can be characterized by the Frobenius norm of the noise matrix. Our empirical studies with standard deep learning model-architectures and datasets shows that our method not only improves generalization performance in large-batch training, but furthermore, does so in a way where the optimization performance remains desirable and the training duration is not elongated.
Yeming Wen, Kevin Luk, Maxime Gazeau, Guodong Zhang 0006, Harris Chan, Jimmy Ba
AISTATS1
2020 BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong Learning
Yeming Wen, Dustin Tran, Jimmy Ba
ICLR1
2020 Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors
abstract
Bayesian neural networks (BNNs) demonstrate promising success in improving the robustness and uncertainty quantification of modern deep learning. However, they generally struggle with underfitting at scale and parameter efficiency. On the other hand, deep ensembles have emerged as alternatives for uncertainty quantification that, while outperforming BNNs on certain problems, also suffer from efficiency issues. It remains unclear how to combine the strengths of these two approaches and remediate their common issues. To tackle this challenge, we propose a rank-1 parameterization of BNNs, where each weight matrix involves only a distribution on a rank-1 subspace. We also revisit the use of mixture approximate posteriors to capture multiple modes, where unlike typical mixtures, this approach admits a significantly smaller memory increase (e.g., only a 0.4% increase for a ResNet-50 mixture of size 10). We perform a systematic empirical study on the choices of prior, variational posterior, and methods to improve training. For ResNet-50 on ImageNet, Wide ResNet 28-10 on CIFAR-10/100, and an RNN on MIMIC-III, rank-1 BNNs achieve state-of-the-art performance across log-likelihood, accuracy, and calibration on the test sets and out-of-distribution variants.
Michael Dusenberry, Ghassen Jerfel, Yeming Wen, Yi-An Ma, Jasper Snoek, Katherine A. Heller, Balaji Lakshminarayanan, Dustin Tran
ICML3
2018 Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches
Yeming Wen, Paul Vicol, Jimmy Ba, Dustin Tran, Roger B. Grosse
ICLR (Poster)1