EDBT 2026 Demo / reviewers in the wild / expert
Adji B. Dieng
dblp:188/6478 · also Adji Bousso Dieng
· DBLP profile ↗
12ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Probabilistic and Bayesian machine learning · 31% Optimization for machine learning · 14% Trustworthy machine learning · 13% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
1.2 | 3 | 2022 | Markov Chain Score Ascent: A Unifying Framework of Variational Inference with Markovian Gradients · NeurIPS 2022 Augment and Reduce: Stochastic Inference for Large Categorical Distributions · ICML 2018 Variational Inference via \chi Upper Bound Minimization · NIPS 2017 |
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
0.8 | 1 | 2024 | Quality-Weighted Vendi Scores And Their Application To Diverse Experimental Design · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning
experimental design |
0.8 | 1 | 2024 | Quality-Weighted Vendi Scores And Their Application To Diverse Experimental Design · ICML 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | Beyond Aesthetics: Cultural Competence in Text-to-Image Models · NeurIPS 2024 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.6 | 1 | 2022 | Markov Chain Score Ascent: A Unifying Framework of Variational Inference with Markovian Gradients · NeurIPS 2022 |
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization |
0.5 | 1 | 2021 | Consistency Regularization for Variational Auto-Encoders · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.5 | 1 | 2021 | Consistency Regularization for Variational Auto-Encoders · NeurIPS 2021 |
Machine learning › Generative modeling
variational autoencoder |
0.5 | 1 | 2021 | Consistency Regularization for Variational Auto-Encoders · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › exponential family
categorical distribution |
0.3 | 1 | 2018 | Augment and Reduce: Stochastic Inference for Large Categorical Distributions · ICML 2018 |
Machine learning › Probabilistic and Bayesian machine learning
marginal likelihood |
0.3 | 1 | 2018 | Noisin: Unbiased Regularization for Recurrent Neural Networks · ICML 2018 |
Machine learning › Deep learning architectures and training › regularization
noise injection |
0.3 | 1 | 2018 | Noisin: Unbiased Regularization for Recurrent Neural Networks · ICML 2018 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2018 | Noisin: Unbiased Regularization for Recurrent Neural Networks · ICML 2018 |
Machine learning › Deep learning architectures and training
regularization |
0.3 | 1 | 2018 | Noisin: Unbiased Regularization for Recurrent Neural Networks · ICML 2018 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
stochastic variational inference |
0.3 | 1 | 2018 | Augment and Reduce: Stochastic Inference for Large Categorical Distributions · ICML 2018 |
Machine learning › Trustworthy machine learning › uncertainty estimation
bayesian uncertainty quantification |
0.3 | 1 | 2017 | Variational Inference via \chi Upper Bound Minimization · NIPS 2017 |
Natural language and speech › Language models and text generation › neural language model
recurrent neural network language model |
0.3 | 1 | 2017 | TopicRNN: A Recurrent Neural Network with Long-Range Semantic Dependency · ICLR (Poster) 2017 |
Natural language and speech › Information extraction and text analysis
topic model |
0.3 | 1 | 2017 | TopicRNN: A Recurrent Neural Network with Long-Range Semantic Dependency · ICLR (Poster) 2017 |
Computer vision › Vision and language › vision-language model
vision-language model evaluation |
0.2 | 1 | 2024 | Beyond Aesthetics: Cultural Competence in Text-to-Image Models · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
vendi scores · 1.5structured knowledge base · 0.8quality-weighted diversity · 0.8large language model · 0.8markov chain gradient descent · 0.6kullback-leibler divergence · 0.6variational inference · 0.5KL divergence regularization · 0.5marginal likelihood maximization · 0.3dropout · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Behavior-Inspired Neural Networks for Relational InferenceabstractFrom pedestrians to Kuramoto oscillators, interactions between agents govern how dynamical systems evolve in space and time. Discovering how these agents relate to each other has the potential to improve our understanding of the often complex dynamics that underlie these systems. Recent works learn to categorize relationships between agents based on observations of their physical behavior. These approaches model relationship categories as outcomes of a categorical distribution which is limiting and contrary to real-world systems, where relationship categories often intermingle and interact. In this work, we introduce a level of abstraction between the observable behavior of agents and the latent categories that determine their behavior. To do this, we learn a mapping from agent observations to agent preferences for a set of latent categories. The learned preferences and inter-agent proximity are integrated in a nonlinear opinion dynamics model, which allows us to naturally identify mutually exclusive categories, predict an agent’s evolution in time, and control an agent’s behavior. Through extensive experiments, we demonstrate the utility of our model for learning interpretable categories, and the efficacy of our model for long-horizon trajectory prediction. Yulong Yang 0003, Bowen Feng, Keqin Wang, Naomi Ehrich Leonard, Adji B. Dieng, Christine Allen-Blanchette |
AISTATS | 5 |
| 2024 | Cousins Of The Vendi Score: A Family Of Similarity-Based Diversity Metrics For Science And Machine Learning
Amey P. Pasarkar, Adji B. Dieng |
AISTATS | 2 |
| 2024 | Quality-Weighted Vendi Scores And Their Application To Diverse Experimental DesignabstractExperimental design techniques such as active search and Bayesian optimization are widely used in the natural sciences for data collection and discovery. However, existing techniques tend to favor exploitation over exploration of the search space, which causes them to get stuck in local optima. This collapse problem prevents experimental design algorithms from yielding diverse high-quality data. In this paper, we extend the Vendi scores—a family of interpretable similarity-based diversity metrics—to account for quality. We then leverage these quality-weighted Vendi scores to tackle experimental design problems across various applications, including drug discovery, materials discovery, and reinforcement learning. We found that quality-weighted Vendi scores allow us to construct policies for experimental design that flexibly balance quality and diversity, and ultimately assemble rich and diverse sets of high-performing data points. Our algorithms led to a 70%–170% increase in the number of effective discoveries compared to baselines. Adji B. Dieng |
ICML | 2 |
| 2024 | Beyond Aesthetics: Cultural Competence in Text-to-Image ModelsabstractText-to-Image (T2I) models are being increasingly adopted in diverse global communities where they create visual representations of their unique cultures. Current T2I benchmarks primarily focus on faithfulness, aesthetics, and realism of generated images, overlooking the critical dimension of cultural competence. In this work, we introduce a framework to evaluate cultural competence of T2I models along two crucial dimensions: cultural awareness and cultural diversity, and present a scalable approach using a combination of structured knowledge bases and large language models to build a large dataset of cultural artifacts to enable this evaluation. In particular, we apply this approach to build CUBE (CUltural BEnchmark for Text-to-Image models), a first-of-its-kind benchmark to evaluate cultural competence of T2I models. CUBE covers cultural artifacts associated with 8 countries across different geo-cultural regions and along 3 concepts: cuisine, landmarks, and art. CUBE consists of 1) CUBE-1K, a set of high-quality prompts that enable the evaluation of cultural awareness, and 2) CUBE-CSpace, a larger dataset of cultural artifacts that serves as grounding to evaluate cultural diversity. We also introduce cultural diversity as a novel T2I evaluation component, leveraging quality-weighted Vendi score. Our evaluations reveal significant gaps in the cultural awareness of existing models across countries and provide valuable insights into the cultural diversity of T2I outputs for underspecified prompts. Our methodology is extendable to other cultural regions and concepts and can facilitate the development of T2I models that better cater to the global population. Nithish Kannen, Arif Ahmad, Marco Andreetto, Vinodkumar Prabhakaran, Utsav Prabhu, Adji B. Dieng, Pushpak Bhattacharyya, Shachi Dave |
NeurIPS | 6 |
| 2022 | Markov Chain Score Ascent: A Unifying Framework of Variational Inference with Markovian GradientsabstractMinimizing the inclusive Kullback-Leibler (KL) divergence with stochastic gradient descent (SGD) is challenging since its gradient is defined as an integral over the posterior. Recently, multiple methods have been proposed to run SGD with biased gradient estimates obtained from a Markov chain. This paper provides the first non-asymptotic convergence analysis of these methods by establishing their mixing rate and gradient variance. To do this, we demonstrate that these methods—which we collectively refer to as Markov chain score ascent (MCSA) methods—can be cast as special cases of the Markov chain gradient descent framework. Furthermore, by leveraging this new understanding, we develop a novel MCSA scheme, parallel MCSA (pMCSA), that achieves a tighter bound on the gradient variance. We demonstrate that this improved theoretical result translates to superior empirical performance. Kyurae Kim, Jisu Oh, Jacob R. Gardner, Adji B. Dieng, Hongseok Kim |
NeurIPS | 4 |
| 2021 | Consistency Regularization for Variational Auto-EncodersabstractVariational Auto-Encoders (VAEs) are a powerful approach to unsupervised learning. They enable scalable approximate posterior inference in latent-variable models using variational inference. A VAE posits a variational family parameterized by a deep neural network---called an encoder---that takes data as input. This encoder is shared across all the observations, which amortizes the cost of inference. However the encoder of a VAE has the undesirable property that it maps a given observation and a semantics-preserving transformation of it to different latent representations. This "inconsistency" of the encoder lowers the quality of the learned representations, especially for downstream tasks, and also negatively affects generalization. In this paper, we propose a regularization method to enforce consistency in VAEs. The idea is to minimize the Kullback-Leibler (KL) divergence between the variational distribution when conditioning on the observation and the variational distribution when conditioning on a random semantics-preserving transformation of this observation. This regularization is applicable to any VAE. In our experiments we apply it to four different VAE variants on several benchmark datasets and found it always improves the quality of the learned representations but also leads to better generalization. In particular, when applied to the Nouveau VAE (NVAE), our regularization method yields state-of-the-art performance on MNIST, CIFAR-10, and CELEBA. We also applied our method to 3D data and found it learns representations of superior quality as measured by accuracy on a downstream classification task. Finally, we show our method can even outperform the triplet loss, an advanced and popular contrastive learning-based method for representation learning. Samarth Sinha, Adji B. Dieng |
NeurIPS | 2 |
| 2020 | Topic Modeling in Embedding SpacesabstractTopic modeling analyzes documents to learn meaningful patterns of words. However, existing topic models fail to learn interpretable topics when working with large and heavy-tailed vocabularies. To this end, we develop the embedded topic model (etm), a generative model of documents that marries traditional topic models with word embeddings. More specifically, the etm models each word with a categorical distribution whose natural parameter is the inner product between the word’s embedding and an embedding of its assigned topic. To fit the etm, we develop an efficient amortized variational inference algorithm. The etm discovers interpretable topics even with large vocabularies that include rare words and stop words. It outperforms existing document models, such as latent Dirichlet allocation, in terms of both topic quality and predictive performance. Adji B. Dieng, Francisco J. R. Ruiz, David M. Blei |
Trans. Assoc. Comput. Linguistics | 1 |
| 2019 | Avoiding Latent Variable Collapse with Generative Skip ModelsabstractVariational autoencoders (VAEs) learn distributions of high-dimensional data. They model data with a deep latent-variable model and then fit the model by maximizing a lower bound of the log marginal likelihood. VAEs can capture complex distributions, but they can also suffer from an issue known as "latent variable collapse," especially if the likelihood model is powerful. Specifically, the lower bound involves an approximate posterior of the latent variables; this posterior "collapses" when it is set equal to the prior, i.e., when the approximate posterior is independent of the data. While VAEs learn good generative models, latent variable collapse prevents them from learning useful representations. In this paper, we propose a simple new way to avoid latent variable collapse by including skip connections in our generative model; these connections enforce strong links between the latent variables and the likelihood function. We study generative skip models both theoretically and empirically. Theoretically, we prove that skip models increase the mutual information between the observations and the inferred latent variables. Empirically, we study images (MNIST and Omniglot) and text (Yahoo). Compared to existing VAE architectures, we show that generative skip models maintain similar predictive performance but lead to less collapse and provide more meaningful representations of the data. Adji B. Dieng, Alexander M. Rush, David M. Blei |
AISTATS | 1 |
| 2018 | Noisin: Unbiased Regularization for Recurrent Neural NetworksabstractRecurrent neural networks (RNNs) are powerful models of sequential data. They have been successfully used in domains such as text and speech. However, RNNs are susceptible to overfitting; regularization is important. In this paper we develop Noisin, a new method for regularizing RNNs. Noisin injects random noise into the hidden states of the RNN and then maximizes the corresponding marginal likelihood of the data. We show how Noisin applies to any RNN and we study many different types of noise. Noisin is unbiased–it preserves the underlying RNN on average. We characterize how Noisin regularizes its RNN both theoretically and empirically. On language modeling benchmarks, Noisin improves over dropout by as much as 12.2% on the Penn Treebank and 9.4% on the Wikitext-2 dataset. We also compared the state-of-the-art language model of Yang et al. 2017, both with and without Noisin. On the Penn Treebank, the method with Noisin more quickly reaches state-of-the-art performance. Adji B. Dieng, Rajesh Ranganath, Jaan Altosaar, David M. Blei |
ICML | 1 |
| 2018 | Augment and Reduce: Stochastic Inference for Large Categorical DistributionsabstractCategorical distributions are ubiquitous in machine learning, e.g., in classification, language models, and recommendation systems. However, when the number of possible outcomes is very large, using categorical distributions becomes computationally expensive, as the complexity scales linearly with the number of outcomes. To address this problem, we propose augment and reduce (A&R), a method to alleviate the computational complexity. A&R uses two ideas: latent variable augmentation and stochastic variational inference. It maximizes a lower bound on the marginal likelihood of the data. Unlike existing methods which are specific to softmax, A&R is more general and is amenable to other categorical models, such as multinomial probit. On several large-scale classification problems, we show that A&R provides a tighter bound on the marginal likelihood and has better predictive performance than existing approaches. Francisco J. R. Ruiz, Michalis K. Titsias, Adji B. Dieng, David M. Blei |
ICML | 3 |
| 2017 | TopicRNN: A Recurrent Neural Network with Long-Range Semantic Dependency
Adji B. Dieng, Chong Wang 0002, Jianfeng Gao 0001, John W. Paisley |
ICLR (Poster) | 1 |
| 2017 | Variational Inference via \chi Upper Bound MinimizationabstractVariational inference (VI) is widely used as an efficient alternative to Markov chain Monte Carlo. It posits a family of approximating distributions $q$ and finds the closest member to the exact posterior $p$. Closeness is usually measured via a divergence $D(q || p)$ from $q$ to $p$. While successful, this approach also has problems. Notably, it typically leads to underestimation of the posterior variance. In this paper we propose CHIVI, a black-box variational inference algorithm that minimizes $D_{\chi}(p || q)$, the $\chi$-divergence from $p$ to $q$. CHIVI minimizes an upper bound of the model evidence, which we term the $\chi$ upper bound (CUBO). Minimizing the CUBO leads to improved posterior uncertainty, and it can also be used with the classical VI lower bound (ELBO) to provide a sandwich estimate of the model evidence. We study CHIVI on three models: probit regression, Gaussian process classification, and a Cox process model of basketball plays. When compared to expectation propagation and classical VI, CHIVI produces better error rates and more accurate estimates of posterior variance. Adji B. Dieng, Dustin Tran, Rajesh Ranganath, John W. Paisley, David M. Blei |
NIPS | 1 |