VLDB 2026 Research / reviewers in the wild / expert
Henry Li
dblp:31/6498
· DBLP profile ↗
14ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Generative modeling · 52% Vision and language · 17% Probabilistic and Bayesian machine learning · 12% | |
| Computer graphics and multimedia
2 papers |
Image and video coding · 77% Image and video processing · 23% |
Topics — the 20 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
3.6 | 5 | 2025 | Dual Diffusion for Unified Image Generation and Understanding · CVPR 2025 Solving Inverse Problems via Diffusion Optimal Control · NeurIPS 2024 Boosting Alignment for Post-Unlearning Text-to-Image Generative Models · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
multimodal diffusion model |
0.9 | 1 | 2025 | Dual Diffusion for Unified Image Generation and Understanding · CVPR 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal understanding and generation |
0.9 | 1 | 2025 | Dual Diffusion for Unified Image Generation and Understanding · CVPR 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
unified image understanding and generation |
0.9 | 1 | 2025 | Dual Diffusion for Unified Image Generation and Understanding · CVPR 2025 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Boosting Alignment for Post-Unlearning Text-to-Image Generative Models · NeurIPS 2024 |
Computer vision › Vision and language › cross-modal alignment
image-text alignment |
0.8 | 1 | 2024 | Boosting Alignment for Post-Unlearning Text-to-Image Generative Models · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
inverse problem solving |
0.8 | 1 | 2024 | Solving Inverse Problems via Diffusion Optimal Control · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
machine unlearning |
0.8 | 1 | 2024 | Boosting Alignment for Post-Unlearning Text-to-Image Generative Models · NeurIPS 2024 |
Robotics › Motion planning and robot control › robot control
optimal control |
0.8 | 1 | 2024 | Solving Inverse Problems via Diffusion Optimal Control · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | Boosting Alignment for Post-Unlearning Text-to-Image Generative Models · NeurIPS 2024 |
Image and video coding
lossless compression |
0.8 | 1 | 2024 | Likelihood Training of Cascaded Diffusion Models via Hierarchical Volume-preserving Maps · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation |
0.6 | 1 | 2022 | Neural Inverse Transform Sampler · ICML 2022 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
neural density estimation |
0.6 | 1 | 2022 | Neural Inverse Transform Sampler · ICML 2022 |
Machine learning › Probabilistic and Bayesian machine learning
sampling |
0.6 | 1 | 2022 | Neural Inverse Transform Sampler · ICML 2022 |
Machine learning › Generative modeling
variational autoencoder |
0.4 | 1 | 2020 | Variational Diffusion Autoencoders with Random Walk Sampling · ECCV (23) 2020 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
deep clustering |
0.3 | 1 | 2018 | SpectralNet: Spectral Clustering using Deep Neural Networks · ICLR (Poster) 2018 |
Machine learning › Generative modeling › diffusion model › discrete diffusion model
discrete diffusion language model |
0.3 | 1 | 2025 | Dual Diffusion for Unified Image Generation and Understanding · CVPR 2025 |
Machine learning › Generative modeling › diffusion model › latent diffusion model
stable diffusion |
0.2 | 1 | 2024 | Boosting Alignment for Post-Unlearning Text-to-Image Generative Models · NeurIPS 2024 |
Image and video processing
image reconstruction |
0.2 | 1 | 2024 | Solving Inverse Problems via Diffusion Optimal Control · NeurIPS 2024 |
Machine learning › Generative modeling › generative model
likelihood-based generative model |
0.2 | 1 | 2022 | Neural Inverse Transform Sampler · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
wavelet transform · 1.5score matching · 1.5optimal control · 1.5laplacian pyramid · 1.5iterative linear quadratic regulator · 1.5diffusion transformer · 0.9cross-modal maximum likelihood estimation · 0.9monotonic model update · 0.8machine unlearning · 0.8dataset diversification · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dual Diffusion for Unified Image Generation and UnderstandingabstractDiffusion models have gained tremendous success in text-to-image generation, yet still struggle with visual understanding tasks, an area dominated by autoregressive vision-language models. We propose a large-scale and fully end-to-end diffusion model for multi-modal understanding and generation that significantly improves on existing diffusion-based multimodal models, and is the first of its kind to support the full suite of vision-language modeling capabilities. Inspired by the multimodal diffusion transformer (MM-DiT) and recent advances in discrete diffusion language modeling, we leverage a cross-modal maximum likelihood estimation framework that simultaneously trains the conditional likelihoods of both images and text jointly under a single loss function, which is back-propagated through both branches of the diffusion transformer. The resulting model is highly flexible and capable of a wide range of tasks including image generation, captioning, and visual question answering. Our model attained competitive performance compared to recent unified image understanding and generation models, demonstrating the potential of multimodal diffusion modeling as a promising alternative to autoregressive next-token prediction models. Henry Li, Yichun Shi, Amir Barati Farimani, Yuval Kluger |
CVPR | 2 |
| 2025 | Supervised fine-tuning of pre-trained antibody language models improves antigen specificity predictionabstractAntibodies play a crucial role in the adaptive immune response, with their specificity to antigens being a fundamental determinant of immune function. Accurate prediction of antibody-antigen specificity is vital for understanding immune responses, guiding vaccine design, and developing antibody-based therapeutics. In this study, we present a method of supervised fine-tuning for antibody language models, which improves on pre-trained antibody language model embeddings in binding specificity prediction to SARS-CoV-2 spike protein and influenza hemagglutinin. We perform supervised fine-tuning on four pre-trained antibody language models to predict specificity to these antigens and demonstrate that fine-tuned language model classifiers exhibit enhanced predictive accuracy compared to classifiers trained on pre-trained model embeddings. Additionally, we investigate the change of model attention activations after supervised fine-tuning to gain insights into the molecular basis of antigen recognition by antibodies. Furthermore, we apply the supervised fine-tuned models to BCR repertoire data related to influenza and SARS-CoV-2 vaccination, demonstrating their ability to capture changes in repertoire following vaccination. Overall, our study highlights the effect of supervised fine-tuning on pre-trained antibody language models as valuable tools to improve antigen specificity prediction. Jonathan Patsenker, Henry Li, Yuval Kluger, Steven H. Kleinstein |
PLoS Comput. Biol. | 3 |
| 2024 | Likelihood Training of Cascaded Diffusion Models via Hierarchical Volume-preserving MapsabstractCascaded models are multi-scale generative models with a marked capacity for producing perceptually impressive samples at high resolutions. In this work, we show that they can also be excellent likelihood models, so long as we overcome a fundamental difficulty with probabilistic multi-scale models: the intractability of the likelihood function. Chiefly, in cascaded models each intermediary scale introduces extraneous variables that cannot be tractably marginalized out for likelihood evaluation. This issue vanishes by modeling the diffusion process on latent spaces induced by a class of transformations we call hierarchical volume-preserving maps, which decompose spatially structured data in a hierarchical fashion without introducing local distortions in the latent space. We demonstrate that two such maps are well-known in the literature for multiscale modeling: Laplacian pyramids and wavelet transforms. Not only do such reparameterizations allow the likelihood function to be directly expressed as a joint likelihood over the scales, we show that the Laplacian pyramid and wavelet transform also produces significant improvements to the state-of-the-art on a selection of benchmarks in likelihood modeling, including density estimation, lossless compression, and out-of-distribution detection. Investigating the theoretical basis of our empirical gains we uncover deep connections to score matching under the Earth Mover's Distance (EMD), which is a well-known surrogate for perceptual similarity. Henry Li, Ronen Basri, Yuval Kluger |
ICLR | 1 |
| 2024 | Boosting Alignment for Post-Unlearning Text-to-Image Generative ModelsabstractLarge-scale generative models have shown impressive image-generation capabilities, propelled by massive data. However, this often inadvertently leads to the generation of harmful or inappropriate content and raises copyright concerns. Driven by these concerns, machine unlearning has become crucial to effectively purge undesirable knowledge from models. While existing literature has studied various unlearning techniques, these often suffer from either poor unlearning quality or degradation in text-image alignment after unlearning, due to the competitive nature of these objectives. To address these challenges, we propose a framework that seeks an optimal model update at each unlearning iteration, ensuring monotonic improvement on both objectives. We further derive the characterization of such an update.
In addition, we design procedures to strategically diversify the unlearning and remaining datasets to boost performance improvement. Our evaluation demonstrates that our method effectively removes target classes from recent diffusion-based generative models and concepts from stable diffusion models while maintaining close alignment with the models' original trained states, thus outperforming state-of-the-art baselines. Myeongseob Ko, Henry Li, Zhun Wang, Jonathan Patsenker, Jiachen T. Wang, Qinbin Li, Ming Jin 0002, Dawn Song, Ruoxi Jia 0001 |
NeurIPS | 2 |
| 2024 | Solving Inverse Problems via Diffusion Optimal ControlabstractExisting approaches to diffusion-based inverse problem solvers frame the signal recovery task as a probabilistic sampling episode, where the solution is drawn from the desired posterior distribution. This framework suffers from several critical drawbacks, including the intractability of the conditional likelihood function, strict dependence on the score network approximation, and poor $\mathbf{x}_0$ prediction quality. We demonstrate that these limitations can be sidestepped by reframing the generative process as a discrete optimal control episode. We derive a diffusion-based optimal controller inspired by the iterative Linear Quadratic Regulator (iLQR) algorithm. This framework is fully general and able to handle any differentiable forward measurement operator, including super-resolution, inpainting, Gaussian deblurring, nonlinear deblurring, and even highly nonlinear neural classifiers. Furthermore, we show that the idealized posterior sampling equation can be recovered as a special case of our algorithm. We then evaluate our method against a selection of neural inverse problem solvers, and establish a new baseline in image reconstruction with inverse problems. Henry Li, Marcus Pereira |
NeurIPS | 1 |
| 2024 | Anomaly Detection with Variance Stabilized Density EstimationabstractWe propose a modified density estimation problem that is highly effective for detecting anomalies in tabular data. Our approach assumes that the density function is relatively stable (with lower variance) around normal samples. We have verified this hypothesis empirically using a wide range of real-world data. Then, we present a variance-stabilized density estimation problem for maximizing the likelihood of the observed samples while minimizing the variance of the density around normal samples. To obtain a reliable anomaly detector, we introduce a spectral ensemble of autoregressive models for learning the variance-stabilized distribution. We have conducted an extensive benchmark with 52 datasets, demonstrating that our method leads to state-of-the-art results while alleviating the need for data-specific hyperparameter tuning. Finally, we have used an ablation study to demonstrate the importance of each of the proposed components, followed by a stability analysis evaluating the robustness of our model. Amit Rozner, Barak Battash, Henry Li, Lior Wolf, Ofir Lindenbaum |
UAI | 3 |
| 2023 | Joint Energy-Based Model for Robust Speech Classification System Against Dirty-Label Backdoor Poisoning AttacksabstractOur novel technique utilizes a Joint Energy-based Model (JEM) that integrates both discriminative and generative approaches to increase resistance against dirty-label backdoor attacks. Our approach is especially effective when the trigger is short or hardly perceivable. We simulate the attack on the Speech Commands Dataset consisting of 1s audio clips. During training, we use JEM to model a view of the input implemented by a randomly selected 610ms window. During inference, we combine all (40) possible views utilizing a generative part of JEM. The resulting system has slightly decreased accuracy but significantly increased resistance shown in multiple scenarios. Interestingly, replacing JEM with a standard discriminative model (Disc) provides increased resistance with a lesser effect compared to JEM but maintains accuracy. We introduce an extension motivated by semi-supervised training that further improves JEM but not Disc. JEM can also benefit from Gaussian noise during evaluation. Martin Sustek, Sonal Joshi, Henry Li, Thomas Thebaud, Jesús Villalba 0001, Sanjeev Khudanpur, Najim Dehak |
ASRU | 3 |
| 2023 | Clustering Unsupervised Representations as Defense Against Poisoning Attacks on Speech Commands Classification SystemabstractPoisoning attacks entail attackers intentionally tampering with training data. In this paper, we consider a dirty-label poisoning attack scenario on a speech commands classification system. The threat model assumes that certain utterances from one of the classes (source class) are poisoned by superimposing a trigger on it, and its label is changed to another class selected by the attacker (target class). We propose a filtering defense against such an attack. First, we use DIstillation with NO labels (DINO) to learn unsupervised representations for all the training examples. Next, we use K-means and LDA to cluster these representations. Finally, we keep the utterances with the most repeated label in their cluster for training and discard the rest. For a 10% poisoned source class, we demonstrate a drop in attack success rate from 99.75% to 0.25%. We test our defense against a variety of threat models, including different target and source classes, as well as trigger variations. Thomas Thebaud, Sonal Joshi, Henry Li, Martin Sustek, Jesús Villalba 0001, Sanjeev Khudanpur, Najim Dehak |
ASRU | 3 |
| 2023 | Support recovery with Projected Stochastic Gates: Theory and application for linear models
Soham Jana, Henry Li, Yutaro Yamada, Ofir Lindenbaum |
Signal Process. | 2 |
| 2022 | Neural Inverse Transform SamplerabstractAny explicit functional representation $f$ of a density is hampered by two main obstacles when we wish to use it as a generative model: designing $f$ so that sampling is fast, and estimating $Z = \int f$ so that $Z^{-1}f$ integrates to 1. This becomes increasingly complicated as $f$ itself becomes complicated. In this paper, we show that when modeling one-dimensional conditional densities with a neural network, $Z$ can be exactly and efficiently computed by letting the network represent the cumulative distribution function of a target density, and applying a generalized fundamental theorem of calculus. We also derive a fast algorithm for sampling from the resulting representation by the inverse transform method. By extending these principles to higher dimensions, we introduce the \textbf{Neural Inverse Transform Sampler (NITS)}, a novel deep learning framework for modeling and sampling from general, multidimensional, compactly-supported probability densities. NITS is a highly expressive density estimator that boasts end-to-end differentiability, fast sampling, and exact and cheap likelihood evaluation. We demonstrate the applicability of NITS by applying it to realistic, high-dimensional density estimation tasks: likelihood-based generative modeling on the CIFAR-10 dataset, and density estimation on the UCI suite of benchmark datasets, where NITS produces compelling results rivaling or surpassing the state of the art. Henry Li, Yuval Kluger |
ICML | 1 |
| 2021 | MIMIC: an optimization method to identify cell type-specific marker panel for cell sortingabstractMulti-omics data allow us to select a small set of informative markers for the discrimination of specific cell types and study of cellular heterogeneity. However, it is often challenging to choose an optimal marker panel from the high-dimensional molecular profiles for a large amount of cell types. Here, we propose a method called Mixed Integer programming Model to Identify Cell type-specific marker panel (MIMIC). MIMIC maintains the hierarchical topology among different cell types and simultaneously maximizes the specificity of a fixed number of selected markers. MIMIC was benchmarked on the mouse ENCODE RNA-seq dataset, with 29 diverse tissues, for 43 surface markers (SMs) and 1345 transcription factors (TFs). MIMIC could select biologically meaningful markers and is robust for different accuracy criteria. It shows advantages over the standard single gene-based approaches and widely used dimensional reduction methods, such as multidimensional scaling and t-SNE, both in accuracy and in biological interpretation. Furthermore, the combination of SMs and TFs achieves better specificity than SMs or TFs alone. Applying MIMIC to a large collection of 641 RNA-seq samples covering 231 cell types identifies a panel of TFs and SMs that reveal the modularity of cell type association networks. Finally, the scalability of MIMIC is demonstrated by selecting enhancer markers from mouse ENCODE data. MIMIC is freely available at https://github.com/MengZou1/MIMIC. Meng Zou, Zhana Duren, Qiuyue Yuan, Henry Li, Andrew Paul Hutchins, Wing Hung Wong, Yong Wang 0001 |
Briefings Bioinform. | 4 |
| 2020 | Variational Diffusion Autoencoders with Random Walk Sampling
Henry Li, Ofir Lindenbaum, Xiuyuan Cheng, Alexander Cloninger |
ECCV (23) | 1 |
| 2018 | SpectralNet: Spectral Clustering using Deep Neural Networks
Uri Shaham 0001, Kelly P. Stanton, Henry Li, Ronen Basri, Boaz Nadler, Yuval Kluger |
ICLR (Poster) | 3 |
| 2008 | ArrayWiki: an enabling technology for sharing public microarray data repositories and meta-analysesabstractBACKGROUND: A survey of microarray databases reveals that most of the repository contents and data models are heterogeneous (i.e., data obtained from different chip manufacturers), and that the repositories provide only basic biological keywords linking to PubMed. As a result, it is difficult to find datasets using research context or analysis parameters information beyond a few keywords. For example, to reduce the "curse-of-dimension" problem in microarray analysis, the number of samples is often increased by merging array data from different datasets. Knowing chip data parameters such as pre-processing steps (e.g., normalization, artefact removal, etc), and knowing any previous biological validation of the dataset is essential due to the heterogeneity of the data. However, most of the microarray repositories do not have meta-data information in the first place, and do not have a a mechanism to add or insert this information. Thus, there is a critical need to create "intelligent" microarray repositories that (1) enable update of meta-data with the raw array data, and (2) provide standardized archiving protocols to minimize bias from the raw data sources. RESULTS: To address the problems discussed, we have developed a community maintained system called ArrayWiki that unites disparate meta-data of microarray meta-experiments from multiple primary sources with four key features. First, ArrayWiki provides a user-friendly knowledge management interface in addition to a programmable interface using standards developed by Wikipedia. Second, ArrayWiki includes automated quality control processes (caCORRECT) and novel visualization methods (BioPNG, Gel Plots), which provide extra information about data quality unavailable in other microarray repositories. Third, it provides a user-curation capability through the familiar Wiki interface. Fourth, ArrayWiki provides users with simple text-based searches across all experiment meta-data, and exposes data to search engine crawlers (Semantic Agents) such as Google to further enhance data discovery. CONCLUSIONS: Microarray data and meta information in ArrayWiki are distributed and visualized using a novel and compact data storage format, BioPNG. Also, they are open to the research community for curation, modification, and contribution. By making a small investment of time to learn the syntax and structure common to all sites running MediaWiki software, domain scientists and practioners can all contribute to make better use of microarray technologies in research and medical practices. ArrayWiki is available at http://www.bio-miblab.org/arraywiki. Todd H. Stokes, J. T. Torrance, Henry Li, May D. Wang |
BMC Bioinform. | 3 |