VLDB 2026 Research / reviewers in the wild / expert
Alex Bie
dblp:254/0822
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Learning theory · 71% Generative modeling · 14% Trustworthy machine learning · 12% | |
| Network and information security
4 papers |
Privacy and data protection · 100% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Privacy and data protection
differential privacy |
1.9 | 4 | 2023 | Private Distribution Learning with Public Data: The View from Sample Compression · NeurIPS 2023 Private Estimation with Public Data · NeurIPS 2022 Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn Divergence · NeurIPS 2021 |
Machine learning › Learning theory › online learning
adaptive adversaries |
0.9 | 1 | 2025 | On the Learnability of Distribution Classes with Adaptive Adversaries · ICML 2025 |
Machine learning › Learning theory › computational learning theory
learnability |
0.9 | 1 | 2025 | On the Learnability of Distribution Classes with Adaptive Adversaries · ICML 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | On the Learnability of Distribution Classes with Adaptive Adversaries · ICML 2025 |
Machine learning › Learning theory › PAC learning
agnostic learning |
0.7 | 1 | 2023 | Distribution Learnability and Robustness · NeurIPS 2023 |
Machine learning › Learning theory
distribution learning |
0.7 | 1 | 2023 | Distribution Learnability and Robustness · NeurIPS 2023 |
Machine learning › Learning theory
PAC learning |
0.7 | 1 | 2023 | Distribution Learnability and Robustness · NeurIPS 2023 |
Machine learning › Learning theory › computational learning theory
robust learnability |
0.7 | 1 | 2023 | Distribution Learnability and Robustness · NeurIPS 2023 |
Machine learning › Learning theory › computational learning theory
sample compression |
0.7 | 1 | 2023 | Private Distribution Learning with Public Data: The View from Sample Compression · NeurIPS 2023 |
Privacy and data protection › differential privacy › private statistical estimation
private distribution learning |
0.7 | 1 | 2023 | Private Distribution Learning with Public Data: The View from Sample Compression · NeurIPS 2023 |
Privacy and data protection › differential privacy
private statistical estimation |
0.6 | 1 | 2022 | Private Estimation with Public Data · NeurIPS 2022 |
Mathematical optimization
statistical estimation |
0.6 | 1 | 2022 | Private Estimation with Public Data · NeurIPS 2022 |
Machine learning › Generative modeling › generative model
differentially private generative model |
0.5 | 1 | 2021 | Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn Divergence · NeurIPS 2021 |
Machine learning › Generative modeling
optimal transport based generative models |
0.5 | 1 | 2021 | Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn Divergence · NeurIPS 2021 |
Privacy and data protection › differential privacy › synthetic data generation
differentially private generative model |
0.5 | 1 | 2021 | Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn Divergence · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture model |
0.2 | 1 | 2023 | Private Distribution Learning with Public Data: The View from Sample Compression · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
differential privacy · 2.5sample compression scheme · 1.3sample compression · 1.3list learning · 1.3huber contamination · 1.3public data augmentation · 1.1sinkhorn divergence · 1.0optimal transport · 1.0gradient estimation · 1.0PAC learning framework · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How to DP-Fy Your Data: A Practical Guide to Generating Synthetic Data With Differential PrivacyabstractHigh quality data is of vital importance for unlocking the full potential of AI for end users. Villalobos et al. stated in 2024 that finding new sources of such data is getting harder as most publicly-available human generated data will soon have been used. Additionally, publicly available data often is not representative of users of a particular system — for example, a research speech dataset of contractors interacting with an AI assistant will likely be more homogeneous, well articulated and self-censored that real world commands that end users will issue. Therefore unlocking high-quality data grounded in real user interactions is of vital interest to both system creators and end users themselves. However, the direct use of user data comes with significant privacy risks, which must be addressed before the data can be used. Differential Privacy (DP) is a well established framework for reasoning about and limiting information leakage, and is a gold standard for protecting user privacy. The focus of this work, Differentially Private Synthetic data, refers to synthetic data that preserves the overall trends of source data (often user-generated), while providing strong privacy guarantees to individuals that contributed to the source dataset. DP synthetic data can unlock the value of datasets that have previously been inaccessible due to privacy concerns. Additionally, DP synthetic data can replace the use of sensitive datasets that previously have only had rudimentary protections like ad-hoc rule-based anonymization. In this survey we explore the full suite of techniques surrounding DP synthetic data, the types of privacy protections different generation approaches can offer, and the state-of-the-art for various modalities including image, tabular, text and federated (decentralized) data. We outline all the components needed in a system that generates DP synthetic data, from sensitive data handling and preparation, to tracking the use of synthetic data and empirical privacy testing. We hope that work will result in increased adoption of DP synthetic data, spur additional research in still underexplored domains, and additionally increase trust in DP synthetic data approaches. Natalia Ponomareva 0001, Zheng Xu 0002, H. Brendan McMahan, Peter Kairouz, Lucas Rosenblatt, Vincent Cohen-Addad, Cristóbal Guzmán, Ryan McKenna, Galen Andrew, Alex Bie, Alexey Kurakin, Morteza Zadimoghaddam, Sergei Vassilvitskii, Andreas Terzis |
J. Artif. Intell. Res. | 10 |
| 2025 | On the Learnability of Distribution Classes with Adaptive AdversariesabstractWe consider the question of learnability of distribution classes in the presence of adaptive adversaries – that is, adversaries capable of intercepting the samples requested by a learner and applying manipulations with full knowledge of the samples before passing it on to the learner. This stands in contrast to oblivious adversaries, who can only modify the underlying distribution the samples come from but not their i.i.d. nature. We formulate a general notion of learnability with respect to adaptive adversaries, taking into account the budget of the adversary. We show that learnability with respect to additive adaptive adversaries is a strictly stronger condition than learnability with respect to additive oblivious adversaries. Tosca Lechner, Alex Bie, Gautam Kamath 0001 |
ICML | 2 |
| 2025 | Escaping Collapse: The Strength of Weak Data for Large Language Model TrainingabstractSynthetically-generated data plays an increasingly larger role in training large language models. However, while synthetic data has been found to be useful, studies have also shown that without proper curation it can cause LLM performance to plateau, or even "collapse", after many training iterations. In this paper, we formalize this question and develop a theoretical framework to investigate how much curation is needed in order to ensure that LLM performance continually improves. Our analysis is inspired by boosting, a classic machine learning technique that leverages a very weak learning algorithm to produce an arbitrarily good classifier. The approach we analyze subsumes many recently proposed methods for training LLMs on synthetic data, and thus our analysis sheds light on why they are successful, and also suggests opportunities for future improvement. We present experiments that validate our theory, and show that dynamically focusing labeling resources on the most challenging examples --- in much the same way that boosting focuses the efforts of the weak learner --- leads to improved performance. Kareem Amin 0002, Sara Babakniya, Alex Bie, Umar Syed, Sergei Vassilvitskii |
NeurIPS | 3 |
| 2023 | Distribution Learnability and RobustnessabstractWe examine the relationship between learnability and robust learnability for the problem of distribution learning.
We show that learnability implies robust learnability if the adversary can only perform additive contamination (and consequently, under Huber contamination), but not if the adversary is allowed to perform subtractive contamination.
Thus, contrary to other learning settings (e.g., PAC learning of function classes), realizable learnability does not imply agnostic learnability.
We also explore related implications in the context of compression schemes and differentially private learnability. Shai Ben-David, Alex Bie, Gautam Kamath 0001, Tosca Lechner |
NeurIPS | 2 |
| 2023 | Private Distribution Learning with Public Data: The View from Sample CompressionabstractWe study the problem of private distribution learning with access to public data. In this setup, which we refer to as *public-private learning*, the learner is given public and private samples drawn from an unknown distribution $p$ belonging to a class $\mathcal Q$, with the goal of outputting an estimate of $p$ while adhering to privacy constraints (here, pure differential privacy) only with respect to the private samples.
We show that the public-private learnability of a class $\mathcal Q$ is connected to the existence of a sample compression scheme for $\mathcal Q$, as well as to an intermediate notion we refer to as \emph{list learning}. Leveraging this connection: (1) approximately recovers previous results on Gaussians over $\mathbb R^d$; and (2) leads to new ones, including sample complexity upper bounds for arbitrary $k$-mixtures of Gaussians over $\mathbb R^d$, results for agnostic and distribution-shift resistant learners, as well as closure properties for public-private learnability under taking mixtures and products of distributions. Finally, via the connection to list learning, we show that for Gaussians in $\mathbb R^d$, at least $d$ public samples are necessary for private learnability, which is close to the known upper bound of $d+1$ public samples. Shai Ben-David, Alex Bie, Clément L. Canonne, Gautam Kamath 0001, Vikrant Singhal |
NeurIPS | 2 |
| 2022 | Private Estimation with Public DataabstractWe initiate the study of differentially private (DP) estimation with access to a small amount of public data. For private estimation of $d$-dimensional Gaussians, we assume that the public data comes from a Gaussian that may have vanishing similarity in total variation distance with the underlying Gaussian of the private data. We show that under the constraints of pure or concentrated DP, $d+1$ public data samples are sufficient to remove any dependence on the range parameters of the private data distribution from the private sample complexity, which is known to be otherwise necessary without public data. For separated Gaussian mixtures, we assume that the underlying public and private distributions are the same, and we consider two settings: (1) when given a dimension-independent amount of public data, the private sample complexity can be improved polynomially in terms of the number of mixture components, and any dependence on the range parameters of the distribution can be removed in the approximate DP case; (2) when given an amount of public data linear in the dimension, the private sample complexity can be made independent of range parameters even under concentrated DP, and additional improvements can be made to the overall sample complexity. Alex Bie, Gautam Kamath 0001, Vikrant Singhal |
NeurIPS | 1 |
| 2021 | Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn DivergenceabstractAlthough machine learning models trained on massive data have led to breakthroughs in several areas, their deployment in privacy-sensitive domains remains limited due to restricted access to data. Generative models trained with privacy constraints on private data can sidestep this challenge, providing indirect access to private data instead. We propose DP-Sinkhorn, a novel optimal transport-based generative method for learning data distributions from private data with differential privacy. DP-Sinkhorn minimizes the Sinkhorn divergence, a computationally efficient approximation to the exact optimal transport distance, between the model and data in a differentially private manner and uses a novel technique for controlling the bias-variance trade-off of gradient estimates. Unlike existing approaches for training differentially private generative models, which are mostly based on generative adversarial networks, we do not rely on adversarial objectives, which are notoriously difficult to optimize, especially in the presence of noise imposed by privacy constraints. Hence, DP-Sinkhorn is easy to train and deploy. Experimentally, we improve upon the state-of-the-art on multiple image modeling benchmarks and show differentially private synthesis of informative RGB images. Tianshi Cao, Alex Bie, Arash Vahdat, Sanja Fidler, Karsten Kreis |
NeurIPS | 2 |