Alex Bie

dblp:254/0822 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Learning theory · 71% Generative modeling · 14% Trustworthy machine learning · 12%
Network and information security
4 papers
Privacy and data protection · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Privacy and data protection
differential privacy
1.942023
Private Distribution Learning with Public Data: The View from Sample Compression · NeurIPS 2023
Private Estimation with Public Data · NeurIPS 2022
Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn Divergence · NeurIPS 2021
Machine learning › Learning theory › online learning
adaptive adversaries
0.912025
On the Learnability of Distribution Classes with Adaptive Adversaries · ICML 2025
Machine learning › Learning theory › computational learning theory
learnability
0.912025
On the Learnability of Distribution Classes with Adaptive Adversaries · ICML 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
On the Learnability of Distribution Classes with Adaptive Adversaries · ICML 2025
Machine learning › Learning theory › PAC learning
agnostic learning
0.712023
Distribution Learnability and Robustness · NeurIPS 2023
Machine learning › Learning theory
distribution learning
0.712023
Distribution Learnability and Robustness · NeurIPS 2023
Machine learning › Learning theory
PAC learning
0.712023
Distribution Learnability and Robustness · NeurIPS 2023
Machine learning › Learning theory › computational learning theory
robust learnability
0.712023
Distribution Learnability and Robustness · NeurIPS 2023
Machine learning › Learning theory › computational learning theory
sample compression
0.712023
Private Distribution Learning with Public Data: The View from Sample Compression · NeurIPS 2023
Privacy and data protection › differential privacy › private statistical estimation
private distribution learning
0.712023
Private Distribution Learning with Public Data: The View from Sample Compression · NeurIPS 2023
Privacy and data protection › differential privacy
private statistical estimation
0.612022
Private Estimation with Public Data · NeurIPS 2022
Mathematical optimization
statistical estimation
0.612022
Private Estimation with Public Data · NeurIPS 2022
Machine learning › Generative modeling › generative model
differentially private generative model
0.512021
Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn Divergence · NeurIPS 2021
Machine learning › Generative modeling
optimal transport based generative models
0.512021
Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn Divergence · NeurIPS 2021
Privacy and data protection › differential privacy › synthetic data generation
differentially private generative model
0.512021
Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn Divergence · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture model
0.212023
Private Distribution Learning with Public Data: The View from Sample Compression · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

differential privacy · 2.5sample compression scheme · 1.3sample compression · 1.3list learning · 1.3huber contamination · 1.3public data augmentation · 1.1sinkhorn divergence · 1.0optimal transport · 1.0gradient estimation · 1.0PAC learning framework · 0.9
YearPublicationVenuePosition
2026 How to DP-Fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy
abstract
High quality data is of vital importance for unlocking the full potential of AI for end users. Villalobos et al. stated in 2024 that finding new sources of such data is getting harder as most publicly-available human generated data will soon have been used. Additionally, publicly available data often is not representative of users of a particular system — for example, a research speech dataset of contractors interacting with an AI assistant will likely be more homogeneous, well articulated and self-censored that real world commands that end users will issue. Therefore unlocking high-quality data grounded in real user interactions is of vital interest to both system creators and end users themselves. However, the direct use of user data comes with significant privacy risks, which must be addressed before the data can be used. Differential Privacy (DP) is a well established framework for reasoning about and limiting information leakage, and is a gold standard for protecting user privacy. The focus of this work, Differentially Private Synthetic data, refers to synthetic data that preserves the overall trends of source data (often user-generated), while providing strong privacy guarantees to individuals that contributed to the source dataset. DP synthetic data can unlock the value of datasets that have previously been inaccessible due to privacy concerns. Additionally, DP synthetic data can replace the use of sensitive datasets that previously have only had rudimentary protections like ad-hoc rule-based anonymization. In this survey we explore the full suite of techniques surrounding DP synthetic data, the types of privacy protections different generation approaches can offer, and the state-of-the-art for various modalities including image, tabular, text and federated (decentralized) data. We outline all the components needed in a system that generates DP synthetic data, from sensitive data handling and preparation, to tracking the use of synthetic data and empirical privacy testing. We hope that work will result in increased adoption of DP synthetic data, spur additional research in still underexplored domains, and additionally increase trust in DP synthetic data approaches.
Natalia Ponomareva 0001, Zheng Xu 0002, H. Brendan McMahan, Peter Kairouz, Lucas Rosenblatt, Vincent Cohen-Addad, Cristóbal Guzmán, Ryan McKenna, Galen Andrew, Alex Bie, Alexey Kurakin, Morteza Zadimoghaddam, Sergei Vassilvitskii, Andreas Terzis
J. Artif. Intell. Res.10
2025 On the Learnability of Distribution Classes with Adaptive Adversaries
abstract
We consider the question of learnability of distribution classes in the presence of adaptive adversaries – that is, adversaries capable of intercepting the samples requested by a learner and applying manipulations with full knowledge of the samples before passing it on to the learner. This stands in contrast to oblivious adversaries, who can only modify the underlying distribution the samples come from but not their i.i.d. nature. We formulate a general notion of learnability with respect to adaptive adversaries, taking into account the budget of the adversary. We show that learnability with respect to additive adaptive adversaries is a strictly stronger condition than learnability with respect to additive oblivious adversaries.
Tosca Lechner, Alex Bie, Gautam Kamath 0001
ICML2
2025 Escaping Collapse: The Strength of Weak Data for Large Language Model Training
abstract
Synthetically-generated data plays an increasingly larger role in training large language models. However, while synthetic data has been found to be useful, studies have also shown that without proper curation it can cause LLM performance to plateau, or even "collapse", after many training iterations. In this paper, we formalize this question and develop a theoretical framework to investigate how much curation is needed in order to ensure that LLM performance continually improves. Our analysis is inspired by boosting, a classic machine learning technique that leverages a very weak learning algorithm to produce an arbitrarily good classifier. The approach we analyze subsumes many recently proposed methods for training LLMs on synthetic data, and thus our analysis sheds light on why they are successful, and also suggests opportunities for future improvement. We present experiments that validate our theory, and show that dynamically focusing labeling resources on the most challenging examples --- in much the same way that boosting focuses the efforts of the weak learner --- leads to improved performance.
Kareem Amin 0002, Sara Babakniya, Alex Bie, Umar Syed, Sergei Vassilvitskii
NeurIPS3
2023 Distribution Learnability and Robustness
abstract
We examine the relationship between learnability and robust learnability for the problem of distribution learning. We show that learnability implies robust learnability if the adversary can only perform additive contamination (and consequently, under Huber contamination), but not if the adversary is allowed to perform subtractive contamination. Thus, contrary to other learning settings (e.g., PAC learning of function classes), realizable learnability does not imply agnostic learnability. We also explore related implications in the context of compression schemes and differentially private learnability.
Shai Ben-David, Alex Bie, Gautam Kamath 0001, Tosca Lechner
NeurIPS2
2023 Private Distribution Learning with Public Data: The View from Sample Compression
abstract
We study the problem of private distribution learning with access to public data. In this setup, which we refer to as *public-private learning*, the learner is given public and private samples drawn from an unknown distribution $p$ belonging to a class $\mathcal Q$, with the goal of outputting an estimate of $p$ while adhering to privacy constraints (here, pure differential privacy) only with respect to the private samples. We show that the public-private learnability of a class $\mathcal Q$ is connected to the existence of a sample compression scheme for $\mathcal Q$, as well as to an intermediate notion we refer to as \emph{list learning}. Leveraging this connection: (1) approximately recovers previous results on Gaussians over $\mathbb R^d$; and (2) leads to new ones, including sample complexity upper bounds for arbitrary $k$-mixtures of Gaussians over $\mathbb R^d$, results for agnostic and distribution-shift resistant learners, as well as closure properties for public-private learnability under taking mixtures and products of distributions. Finally, via the connection to list learning, we show that for Gaussians in $\mathbb R^d$, at least $d$ public samples are necessary for private learnability, which is close to the known upper bound of $d+1$ public samples.
Shai Ben-David, Alex Bie, Clément L. Canonne, Gautam Kamath 0001, Vikrant Singhal
NeurIPS2
2022 Private Estimation with Public Data
abstract
We initiate the study of differentially private (DP) estimation with access to a small amount of public data. For private estimation of $d$-dimensional Gaussians, we assume that the public data comes from a Gaussian that may have vanishing similarity in total variation distance with the underlying Gaussian of the private data. We show that under the constraints of pure or concentrated DP, $d+1$ public data samples are sufficient to remove any dependence on the range parameters of the private data distribution from the private sample complexity, which is known to be otherwise necessary without public data. For separated Gaussian mixtures, we assume that the underlying public and private distributions are the same, and we consider two settings: (1) when given a dimension-independent amount of public data, the private sample complexity can be improved polynomially in terms of the number of mixture components, and any dependence on the range parameters of the distribution can be removed in the approximate DP case; (2) when given an amount of public data linear in the dimension, the private sample complexity can be made independent of range parameters even under concentrated DP, and additional improvements can be made to the overall sample complexity.
Alex Bie, Gautam Kamath 0001, Vikrant Singhal
NeurIPS1
2021 Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn Divergence
abstract
Although machine learning models trained on massive data have led to breakthroughs in several areas, their deployment in privacy-sensitive domains remains limited due to restricted access to data. Generative models trained with privacy constraints on private data can sidestep this challenge, providing indirect access to private data instead. We propose DP-Sinkhorn, a novel optimal transport-based generative method for learning data distributions from private data with differential privacy. DP-Sinkhorn minimizes the Sinkhorn divergence, a computationally efficient approximation to the exact optimal transport distance, between the model and data in a differentially private manner and uses a novel technique for controlling the bias-variance trade-off of gradient estimates. Unlike existing approaches for training differentially private generative models, which are mostly based on generative adversarial networks, we do not rely on adversarial objectives, which are notoriously difficult to optimize, especially in the presence of noise imposed by privacy constraints. Hence, DP-Sinkhorn is easy to train and deploy. Experimentally, we improve upon the state-of-the-art on multiple image modeling benchmarks and show differentially private synthesis of informative RGB images.
Tianshi Cao, Alex Bie, Arash Vahdat, Sanja Fidler, Karsten Kreis
NeurIPS2