VLDB 2026 Research / reviewers in the wild / expert
Meihua Dang
dblp:270/9145
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0001-8241-1943ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Probabilistic and Bayesian machine learning · 29% Generative modeling · 24% Efficient and distributed learning · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › tractable probabilistic model
probabilistic circuit |
2.4 | 4 | 2025 | Scaling Probabilistic Circuits via Monarch Matrices · ICML 2025 Sparse Probabilistic Circuits via Pruning and Growing · NeurIPS 2022 Juice: A Julia Package for Logic and Probabilistic Circuits · AAAI 2021 |
Machine learning › Generative modeling
diffusion model |
1.6 | 2 | 2025 | Personalized Preference Fine-tuning of Diffusion Models · CVPR 2025 Diffusion Model Alignment Using Direct Preference Optimization · CVPR 2024 |
Machine learning › Efficient and distributed learning
model compression |
1.4 | 2 | 2025 | Scaling Probabilistic Circuits via Monarch Matrices · ICML 2025 Sparse Probabilistic Circuits via Pruning and Growing · NeurIPS 2022 |
Machine learning › Probabilistic and Bayesian machine learning
tractable probabilistic model |
1.2 | 2 | 2023 | Tractable Control for Autoregressive Language Generation · ICML 2023 Group Fairness by Probabilistic Modeling with Latent Fair Decisions · AAAI 2021 |
Machine learning › Generative modeling › diffusion model › diffusion model adaptation
personalized diffusion model |
0.9 | 1 | 2025 | Personalized Preference Fine-tuning of Diffusion Models · CVPR 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | Personalized Preference Fine-tuning of Diffusion Models · CVPR 2025 |
Machine learning › Efficient and distributed learning › model compression › sparsity
sparse parameterization |
0.9 | 1 | 2025 | Scaling Probabilistic Circuits via Monarch Matrices · ICML 2025 |
Machine learning › Generative modeling › diffusion model › diffusion model adaptation
diffusion model alignment |
0.8 | 1 | 2024 | Diffusion Model Alignment Using Direct Preference Optimization · CVPR 2024 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.8 | 1 | 2024 | Diffusion Model Alignment Using Direct Preference Optimization · CVPR 2024 |
Natural language and speech › Language models and text generation › text generation
constrained text generation |
0.7 | 1 | 2023 | Tractable Control for Autoregressive Language Generation · ICML 2023 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
hidden markov model |
0.7 | 1 | 2023 | Tractable Control for Autoregressive Language Generation · ICML 2023 |
Natural language and speech › Machine translation › constrained machine translation
lexical constraints |
0.7 | 1 | 2023 | Tractable Control for Autoregressive Language Generation · ICML 2023 |
Machine learning › Reinforcement learning
multi-objective reinforcement learning |
0.7 | 1 | 2023 | Scaling Pareto-Efficient Decision Making via Offline Multi-Objective RL · ICLR 2023 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.6 | 1 | 2022 | Sparse Probabilistic Circuits via Pruning and Growing · NeurIPS 2022 |
Bioinformatics and computational biology
genetic variation |
0.6 | 1 | 2022 | Tractable and Expressive Generative Models of Genetic Variation Data · RECOMB 2022 |
Bioinformatics and computational biology
genomics |
0.6 | 1 | 2022 | Tractable and Expressive Generative Models of Genetic Variation Data · RECOMB 2022 |
Machine learning › Trustworthy machine learning
fairness |
0.5 | 1 | 2021 | Group Fairness by Probabilistic Modeling with Latent Fair Decisions · AAAI 2021 |
Machine learning › Trustworthy machine learning › fairness
group fairness |
0.5 | 1 | 2021 | Group Fairness by Probabilistic Modeling with Latent Fair Decisions · AAAI 2021 |
Machine learning › Probabilistic and Bayesian machine learning
probabilistic inference |
0.5 | 1 | 2021 | Juice: A Julia Package for Logic and Probabilistic Circuits · AAAI 2021 |
Computer vision › Vision and language
vision-language model |
0.3 | 1 | 2025 | Personalized Preference Fine-tuning of Diffusion Models · CVPR 2025 |
Machine learning › Generative modeling
variational autoencoder |
0.2 | 1 | 2022 | Sparse Probabilistic Circuits via Pruning and Growing · NeurIPS 2022 |
Privacy and data protection
privacy-preserving data analysis |
0.1 | 1 | 2021 | Group Fairness by Probabilistic Modeling with Latent Fair Decisions · AAAI 2021 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 0.9tensorized operations · 0.9monarch matrices · 0.9cross-attention · 0.9RLHF · 0.9DPO · 0.9reinforcement learning from human feedback · 0.8direct preference optimization · 0.8offline reinforcement learning · 0.7distilled hidden markov models · 0.7generative modeling · 0.6probabilistic circuits · 0.5latent variable modeling · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Personalized Preference Fine-tuning of Diffusion ModelsabstractRLHF techniques like DPO can significantly improve the generation quality of text-to-image diffusion models. However, these methods optimize for a single reward that aligns model generation with population-level preferences, neglecting the nuances of individual users’ beliefs or values. This lack of personalization limits the efficacy of these models. To bridge this gap, we introduce PPD, a multi-reward optimization objective that aligns diffusion models with personalized preferences. With PPD, a diffusion model learns the individual preferences of a population of users in a few-shot way, enabling generalization to unseen users. Specifically, our approach (1) leverages a vision-language model (VLM) to extract personal preference embeddings from a small set of pairwise preference examples, and then (2) incorporates the embeddings into diffusion models through cross attention. Conditioning on user embeddings, the text-to-image models are fine-tuned with the DPO objective, simultaneously optimizing for alignment with the preferences of multiple users. Empirical results demonstrate that our method effectively optimizes for multiple reward functions and can interpolate between them during inference. In real-world user scenarios, with as few as four preference examples from a new user, our approach achieves an average win rate of 76% over Stable Cascade, generating images that more accurately reflect specific user preferences. Meihua Dang, Anikait Singh, Linqi Zhou, Stefano Ermon, Jiaming Song |
CVPR | 1 |
| 2025 | Scaling Probabilistic Circuits via Monarch MatricesabstractProbabilistic Circuits (PCs) are tractable representations of probability distributions allowing for exact and efficient computation of likelihoods and marginals. Recent advancements have improved the scalability of PCs either by leveraging their sparse properties or through the use of tensorized operations for better hardware utilization. However, no existing method fully exploits both aspects simultaneously. In this paper, we propose a novel sparse and structured parameterization for the sum blocks in PCs. By replacing dense matrices with sparse Monarch matrices, we significantly reduce the memory and computation costs, enabling unprecedented scaling of PCs. From a theory perspective, our construction arises naturally from circuit multiplication; from a practical perspective, compared to previous efforts on scaling up tractable probabilistic models, our approach not only achieves state-of-the-art generative modeling performance on challenging benchmarks like Text8, LM1B and ImageNet, but also demonstrates superior scaling behavior, achieving the same performance with substantially less compute as measured by the number of floating-point operations (FLOPs) during training. Honghua Zhang, Meihua Dang, Benjie Wang 0001, Stefano Ermon, Nanyun Peng 0001, Guy Van den Broeck |
ICML | 2 |
| 2024 | Diffusion Model Alignment Using Direct Preference OptimizationabstractLarge language models (LLMs) are fine-tuned using human comparison data with Reinforcement Learning from Human Feedback (RLHF) methods to make them better aligned with users' preferences. In contrast to LLMs, human preference learning has not been widely explored in text-to-image diffusion models; the best existing approach is to fine-tune a pretrained model using carefully curated high quality images and captions to improve visual appeal and text alignment. We propose Diffusion-DPO, a method to align diffusion models to human preferences by directly optimizing on human comparison data. Diffusion-DPO is adapted from the recently developed Direct Preference Optimization (DPO) [36], a simpler alternative to RLHF which directly optimizes a policy that best satisfies human preferences under a classification objective. We re-formulate DPO to account for a diffusion model notion of likelihood, utilizing the evidence lower bound to derive a differentiable objective. Using the Pick-a-Pic dataset of 851K crowdsourced pairwise preferences, we fine-tune the base model of the state-of-the-art Stable Diffusion XL (SDXL)-1.0 model with Diffusion-DPO. Our fine-tuned base model significantly outperforms both base SDXL-1.0 and the larger SDXL-1.0 model consisting of an additional refinement model in human evaluation, improving visual appeal and prompt alignment. We also develop a variant that uses AI feedback and has comparable performance to training on human preferences, opening the door for scaling of diffusion model alignment methods. Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq R. Joty |
CVPR | 2 |
| 2023 | Scaling Pareto-Efficient Decision Making via Offline Multi-Objective RL
Baiting Zhu, Meihua Dang, Aditya Grover |
ICLR | 2 |
| 2023 | Tractable Control for Autoregressive Language GenerationabstractDespite the success of autoregressive large language models in text generation, it remains a major challenge to generate text that satisfies complex constraints: sampling from the conditional distribution ${\Pr}(\text{text} | \alpha)$ is intractable for even the simplest lexical constraints $\alpha$. To overcome this challenge, we propose to use tractable probabilistic models (TPMs) to impose lexical constraints in autoregressive text generation models, which we refer to as GeLaTo (Generating Language with Tractable Constraints). To demonstrate the effectiveness of this framework, we use distilled hidden Markov models, where we can efficiently compute ${\Pr}(\text{text} | \alpha)$, to guide autoregressive generation from GPT2. GeLaTo achieves state-of-the-art performance on challenging benchmarks for constrained text generation (e.g., CommonGen), beating various strong baselines by a large margin. Our work not only opens up new avenues for controlling large language models but also motivates the development of more expressive TPMs. Honghua Zhang, Meihua Dang, Nanyun Peng 0001, Guy Van den Broeck |
ICML | 2 |
| 2022 | Sparse Probabilistic Circuits via Pruning and GrowingabstractProbabilistic circuits (PCs) are a tractable representation of probability distributions allowing for exact and efficient computation of likelihoods and marginals. There has been significant recent progress on improving the scale and expressiveness of PCs. However, PC training performance plateaus as model size increases. We discover that most capacity in existing large PC structures is wasted: fully-connected parameter layers are only sparsely used. We propose two operations: pruning and growing, that exploit the sparsity of PC structures. Specifically, the pruning operation removes unimportant sub-networks of the PC for model compression and comes with theoretical guarantees. The growing operation increases model capacity by increasing the dimensions of latent states. By alternatingly applying pruning and growing, we increase the capacity that is meaningfully used, allowing us to significantly scale up PC learning. Empirically, our learner achieves state-of-the-art likelihoods on MNIST-family image datasets and an Penn Tree Bank language data compared to other PC learners and less tractable deep generative models such as flow-based models and variational autoencoders (VAEs). Meihua Dang, Anji Liu, Guy Van den Broeck |
NeurIPS | 1 |
| 2022 | Tractable and Expressive Generative Models of Genetic Variation Data
Meihua Dang, Anji Liu, Xinzhu Wei, Sriram Sankararaman, Guy Van den Broeck |
RECOMB | 1 |
| 2022 | Strudel: A fast and accurate learner of structured-decomposable probabilistic circuits
Meihua Dang, Antonio Vergari, Guy Van den Broeck |
Int. J. Approx. Reason. | 1 |
| 2021 | Group Fairness by Probabilistic Modeling with Latent Fair DecisionsabstractMachine learning systems are increasingly being used to make impactful decisions such as loan applications and criminal justice risk assessments, and as such, ensuring fairness of these systems is critical. This is often challenging as the labels in the data are biased. This paper studies learning fair probability distributions from biased data by explicitly modeling a latent variable that represents a hidden, unbiased label. In particular, we aim to achieve demographic parity by enforcing certain independencies in the learned model. We also show that group fairness guarantees are meaningful only if the distribution used to provide those guarantees indeed captures the real-world data. In order to closely model the data distribution, we employ probabilistic circuits, an expressive and tractable probabilistic model, and propose an algorithm to learn them from incomplete data. We show on real-world datasets that our approach not only is a better model of how the data was generated than existing methods but also achieves competitive accuracy. Moreover, we also evaluate our approach on a synthetic dataset in which observed labels indeed come from fair labels but with added bias, and demonstrate that the fair labels are successfully retrieved. YooJung Choi 0001, Meihua Dang, Guy Van den Broeck |
AAAI | 2 |
| 2021 | Juice: A Julia Package for Logic and Probabilistic CircuitsabstractJuice is an open-source Julia package providing tools for logic and probabilistic reasoning and learning based on logic circuits (LCs) and probabilistic circuits (PCs). It provides a range of efficient algorithms for probabilistic inference queries, such as computing marginal probabilities (MAR), as well as many more advanced queries. Certain structural circuit properties are needed to achieve this tractability, which Juice helps validate. Additionally, it supports several parameter and structure learning algorithms proposed in the recent literature. By leveraging parallelism (on both CPU and GPU), Juice provides a fast implementation of circuit-based algorithms, which makes it suitable for tackling large-scale datasets and models. Meihua Dang, Pasha Khosravi, Yitao Liang, Antonio Vergari, Guy Van den Broeck |
AAAI | 1 |