VLDB 2026 Research / reviewers in the wild / expert
AnhDung Dinh
dblp:344/3730 · also Anh-Dung Dinh
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Generative modeling · 75% Video understanding and tracking · 14% Representation and self-supervised learning · 7% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
4.7 | 7 | 2025 | DiffAct++: Diffusion Action Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Representative Guidance: Diffusion Model Sampling with Coherence · ICLR 2025 Boosting Diffusion Models with an Adaptive Momentum Sampler · IJCAI 2024 |
Machine learning › Generative modeling › diffusion model
diffusion sampling |
1.6 | 2 | 2025 | Representative Guidance: Diffusion Model Sampling with Coherence · ICLR 2025 Boosting Diffusion Models with an Adaptive Momentum Sampler · IJCAI 2024 |
Computer vision › Video understanding and tracking
action segmentation |
1.5 | 2 | 2025 | DiffAct++: Diffusion Action Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Diffusion Action Segmentation · ICCV 2023 |
Machine learning › Generative modeling › diffusion model › score-based generative model
denoising diffusion |
0.9 | 1 | 2025 | DiffAct++: Diffusion Action Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.9 | 1 | 2025 | Representative Guidance: Diffusion Model Sampling with Coherence · ICLR 2025 |
Machine learning › Generative modeling › diffusion model › guided diffusion
classifier guidance |
0.7 | 1 | 2023 | Rethinking Conditional Diffusion Sampling with Progressive Guidance · NeurIPS 2023 |
Machine learning › Generative modeling › diffusion model
conditional sampling |
0.7 | 1 | 2023 | Rethinking Conditional Diffusion Sampling with Progressive Guidance · NeurIPS 2023 |
Machine learning › Generative modeling › diffusion model
guided diffusion |
0.7 | 1 | 2023 | PixelAsParam: A Gradient View on Diffusion Sampling with Guidance · ICML 2023 |
Machine learning › Generative modeling
image generation |
0.7 | 1 | 2023 | PixelAsParam: A Gradient View on Diffusion Sampling with Guidance · ICML 2023 |
Machine learning › Reinforcement learning
policy learning |
0.7 | 1 | 2023 | Learning to Schedule in Diffusion Probabilistic Models · KDD 2023 |
Computer vision › Video understanding and tracking
long video understanding |
0.3 | 1 | 2025 | DiffAct++: Diffusion Action Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Methods — techniques the papers use, named apart from their topics
denoising diffusion · 1.5self-supervised learning · 0.9masking strategy · 0.9consistency gradient guidance · 0.9classifier-free guidance · 0.9momentum sampler · 0.8projection method · 0.7langevin dynamics · 0.7iterative refinement · 0.7gradient-based optimization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Effects of Tempo and Tonality on Listener Enjoyment of Automated Pop MashupsabstractPop mashup is a form of audio remixing prevalent among musicians and automatic mashup systems due to the inherent compatibility of songs within this genre. Recent advancements in machine learning models have facilitated tools for source separation, harmonic analysis, and matching, enabling the generation of high-quality mashups. However, selecting the optimal musical features, specifically tempo and tonality, to maximize listener enjoyment remains a challenge. To address this question, we conducted a series of listening experiments featuring mashups composed of two pop songs, systematically varying these musical features and evaluating the overall preferences of survey participants. Our findings contribute to an understanding of how adjustments to mashup components influence audience perception, offering valuable insights for the development of automated mashup algorithms aimed at enhancing output quality. AnhDung Dinh, Xinyang Wu 0005, Andrew Brian Horner |
IEEE Big Data | 1 |
| 2025 | Representative Guidance: Diffusion Model Sampling with CoherenceabstractThe diffusion sampling process faces a persistent challenge stemming from its incoherence, attributable to varying noise directions across different timesteps.
Our Representative Guidance (RepG) offers a new perspective to address this issue by reformulating the sampling process with a coherent direction toward a representative target.
From this perspective, classic classifier guidance reveals its drawback in lacking meaningful representative information, as the features it relies on are optimized for discrimination and tend to highlight only a narrow set of class-specific cues. This focus often sacrifices diversity and increases the risk of adversarial generation.
In contrast, we leverage self-supervised representations as the coherent target and treat sampling as a downstream task—one that focuses on refining image details and correcting generation errors, rather than settling for oversimplified outputs.
Our Representative Guidance achieves superior performance and demonstrates the potential of pre-trained self-supervised models in guiding diffusion sampling. Our findings show that RepG not only significantly improves vanilla diffusion sampling, but also surpasses state-of-the-art benchmarks when combined with classifier-free guidance. AnhDung Dinh, Daochang Liu, Chang Xu 0002 |
ICLR | 1 |
| 2025 | DiffAct++: Diffusion Action SegmentationabstractUnderstanding long-form videos requires precise temporal action segmentation. While existing studies typically employ multi-stage models that follow an iterative refinement process, we present a novel framework based on the denoising diffusion model that retains this core iterative principle. Within this framework, the model iteratively produces action predictions starting with random noise, conditioned on the features of the input video. To effectively capture three key characteristics of human actions, namely the position prior, the boundary ambiguity, and the relational dependency, we propose a cohesive masking strategy for the conditioning features. Moreover, a consistency gradient guidance technique is proposed, which maximizes the similarity between outputs with or without the masking, thereby enriching conditional information during the inference process. Extensive experiments are performed on four datasets, i.e., GTEA, 50Salads, Breakfast, and Assembly101. The results indicate that our proposed method outperforms or is on par with existing state-of-the-art techniques, underscoring the potential of generative approaches for action segmentation. Daochang Liu, Qiyue Li 0002, AnhDung Dinh, Tingting Jiang 0001, Mubarak Shah, Chang Xu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Boosting Diffusion Models with an Adaptive Momentum Sampler
AnhDung Dinh, Daochang Liu, Chang Xu 0002 |
IJCAI | 2 |
| 2023 | Diffusion Action SegmentationabstractTemporal action segmentation is crucial for understanding long-form videos. Previous works on this task commonly adopt an iterative refinement paradigm by using multi-stage models. We propose a novel framework via denoising diffusion models, which nonetheless shares the same inherent spirit of such iterative refinement. In this framework, action predictions are iteratively generated from random noise with input video features as conditions. To enhance the modeling of three striking characteristics of human actions, including the position prior, the boundary ambiguity, and the relational dependency, we devise a unified masking strategy for the conditioning inputs in our framework. Extensive experiments on three benchmark datasets, i.e., GTEA, 50Salads, and Breakfast, are performed and the proposed method achieves superior or comparable results to state-of-the-art methods, showing the effectiveness of a generative approach for action segmentation. Code is at tinyurl.com/DiffAct. Daochang Liu, Qiyue Li 0002, AnhDung Dinh, Tingting Jiang 0001, Mubarak Shah, Chang Xu 0002 |
ICCV | 3 |
| 2023 | PixelAsParam: A Gradient View on Diffusion Sampling with GuidanceabstractDiffusion models recently achieved state-of-the-art in image generation. They mainly utilize the denoising framework, which leverages the Langevin dynamics process for image sampling. Recently, the guidance method has modified this process to add conditional information to achieve a controllable generator. However, the current guidance on denoising processes suffers from the trade-off between diversity, image quality, and conditional information. In this work, we propose to view this guidance sampling process from a gradient view, where image pixels are treated as parameters being optimized, and each mathematical term in the sampling process represents one update direction. This perspective reveals more insights into the conflict problems between updated directions on the pixels, which cause the trade-off as mentioned previously. We investigate the conflict problems and propose to solve them by a simple projection method. The experimental results evidently improve over different baselines on datasets with various resolutions. AnhDung Dinh, Daochang Liu, Chang Xu 0002 |
ICML | 1 |
| 2023 | Learning to Schedule in Diffusion Probabilistic ModelsabstractRecently, the field of generative models has seen a significant advancement with the introduction of Diffusion Probabilistic Models (DPMs). The Denoising Diffusion Implicit Model (DDIM) was designed to reduce computational time by skipping a number of steps in the inference process of DPMs. However, the hand-crafted sampling schedule in DDIM, which relies on human expertise, has its limitations in considering all relevant factors in the sampling process. Additionally, the assumption that all instances should have the same schedule is not always valid. To address these problems, this paper proposes a method that leverages reinforcement learning to automatically search for an optimal sampling schedule for DPMs. This is achieved by a policy network that predicts the next step to visit based on the current state of the noisy image. The optimization of the policy network is accomplished using an episodic actor-critic framework, which incorporates reinforcement learning. Empirical results demonstrate the superiority of our approach over various datasets with different timesteps. We also observe that the trained sampling schedule has a strong generalization ability across different DPM baselines. Yunke Wang, AnhDung Dinh, Bo Du 0001, Chang Xu 0002 |
KDD | 3 |
| 2023 | Rethinking Conditional Diffusion Sampling with Progressive GuidanceabstractThis paper tackles two critical challenges encountered in classifier guidance for diffusion generative models, i.e., the lack of diversity and the presence of adversarial effects. These issues often result in a scarcity of diverse samples or the generation of non-robust features. The underlying cause lies in the mechanism of classifier guidance, where discriminative gradients push samples to be recognized as conditions aggressively. This inadvertently suppresses information with common features among relevant classes, resulting in a limited pool of features with less diversity or the absence of robust features for image construction. We propose a generalized classifier guidance method called Progressive Guidance, which mitigates the problems by allowing relevant classes' gradients to contribute to shared information construction when the image is noisy in early sampling steps. In the later sampling stage, we progressively enhance gradients to refine the details in the image toward the primary condition. This helps to attain a high level of diversity and robustness compared to the vanilla classifier guidance. Experimental results demonstrate that our proposed method further improves the image quality while offering a significant level of diversity as well as robust features. AnhDung Dinh, Daochang Liu, Chang Xu 0002 |
NeurIPS | 1 |