VLDB 2026 Research / reviewers in the wild / expert
Duy M. H. Nguyen
dblp:199/8349 · also Duy Minh Ho Nguyen
· DBLP profile ↗
22ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0002-7826-2444ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric PerspectiveabstractAs embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually similar objects. In non-Markovian settings, critical decision cues lie in object histories rather than the current scene. Without persistent memory of prior interactions (what was used, where it was placed, or how it changed), visuomotor policies may fail, repeat past actions, or overlook completed ones. To surface this challenge, we introduce LIBERO-Mem, a non-Markovian task suite for stress-testing robotic manipulation under object-level partial observability. It combines short- and long-horizon object tracking with temporally sequenced subgoals, requiring reasoning beyond the current frame. However, vision-language-action (VLA) models often struggle in such settings, with token scaling quickly becoming intractable even for tasks spanning just a few hundred frames. We propose Embodied-SlotSSM, a slot-centric VLA framework built for temporal scalability. It maintains spatio-temporally consistent slot identities and leverages them through two mechanisms: (1) slot-state-space modeling for reconstructing short-term history, and (2) a relational encoder to align the input tokens with action decoding. Together, these components enable temporally grounded, context-aware action prediction. Experiments show Embodied-SlotSSM's baseline performance on LIBERO-Mem and general tasks, offering a scalable solution for non-Markovian reasoning in object-centric policies. Nhat Chung, Taisei Hanyu, Toan Nguyen 0004, Huy Le 0001, Frederick Bumgarner, Duy M. H. Nguyen, Viet-Khoa Vo-Ho, Kashu Yamazaki, Chase Rainwater, Tung Kieu, Anh Nguyen 0003, T. Hoang Ngan Le |
AAAI | 6 |
| 2026 | Reinforce Trustworthiness in Multimodal Emotional Support SystemabstractIn today’s world, emotional support is increasingly essential, yet it remains challenging for both those seeking help and those offering it. Multimodal approaches to emotional support show great promise by integrating diverse data sources to provide empathetic, contextually relevant responses, fostering more effective interactions. However, current methods have notable limitations, often relying solely on text or converting other data types into text, or providing emotion recognition only, thus overlooking the full potential of multimodal inputs. Moreover, many studies prioritize response generation without accurately identifying critical emotional support elements or ensuring the reliability of outputs. To overcome these issues, we introduce MULTIMOOD, a new framework that (i) leverages multimodal embeddings from video, audio, and text to predict emotional components and to produce responses responses aligned with professional therapeutic standards. To improve trustworthiness, we (ii) incorporate novel psychological criteria and apply Reinforcement Learning (RL) to optimize large language models (LLMs) for consistent adherence to these standards. We also (iii) analyze several advanced LLMs to assess their multimodal emotional support capabilities. Experimental results show that MultiMood achieves state-of-the-art on MESC and DFEW datasets while RL-driven trustworthiness improvements are validated through human and LLM evaluations, demonstrating its superior capability in applying a multimodal framework in this domain. Huy M. Le, Tien Dat Nguyen, Ngan T. T. Vo, Tuan D. Q. Nguyen, Nguyen Binh Le, Duy M. H. Nguyen, Daniel Sonntag, Lizi Liao, Binh T. Nguyen 0001 |
AAAI | 6 |
| 2025 | On Zero-Initialized Attention: Optimal Prompt and Gating Factor EstimationabstractLLaMA-Adapter has recently emerged as an efficient fine-tuning technique for LLaMA models, leveraging zero-initialized attention to stabilize training and enhance performance. However, despite its empirical success, the theoretical foundations of zero-initialized attention remain largely unexplored. In this paper, we provide a rigorous theoretical analysis, establishing a connection between zero-initialized attention and mixture-of-expert models. We prove that both linear and non-linear prompts, along with gating functions, can be optimally estimated, with non-linear prompts offering greater flexibility for future applications. Empirically, we validate our findings on the open LLM benchmarks, demonstrating that non-linear prompts outperform linear ones. Notably, even with limited training data, both prompt types consistently surpass vanilla attention, highlighting the robustness and adaptability of zero-initialized attention. Nghiem Tuong Diep, Minh Le, Duy M. H. Nguyen, Daniel Sonntag, Mathias Niepert, Nhat Ho |
ICML | 5 |
| 2025 | ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language ModelsabstractState-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLaVA-Med and BioMedGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach can lead to weak vision-language alignment, making these models overly dependent on costly instruction-following data. To address this, we introduce ExGra-Med, a novel multi-graph alignment framework that jointly aligns images, instruction responses, and extended captions in the latent space, advancing semantic grounding and cross-modal coherence. To scale to large LLMs (e.g., LLaMa-7B), we develop an efficient end-to-end training scheme using black-box gradient estimation, enabling fast and scalable optimization. Empirically, ExGra-Med matches LLaVA-Med’s performance using just 10\% of pre-training data, achieving a 20.13\% gain on VQA-RAD and approaching full-data performance. It also outperforms strong baselines like BioMedGPT and RadFM on visual chatbot and zero-shot classification tasks, demonstrating its promise for efficient, high-quality vision-language integration in medical AI. Duy M. H. Nguyen, Nghiem Tuong Diep, Hoang Bao Le, Tai D. Nguyen, Anh-Tien Nguyen, TrungTin Nguyen, Nhat Ho, Pengtao Xie, Roger Wattenhofer, Daniel Sonntag, James Zou 0001, Mathias Niepert |
NeurIPS | 1 |
| 2025 | Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance SamplingabstractRecently, Direct Alignment Algorithms (DAAs) such as Direct Preference Optimization (DPO) have emerged as alternatives to the standard Reinforcement Learning from Human Feedback (RLHF) for aligning large language models (LLMs) with human values.
Surprisingly, while DAAs do not use a separate proxy reward model as in RLHF, their performance can still deteriorate over the course of training -- an over-optimization phenomenon found in RLHF where the learning policy exploits the overfitting to inaccuracies of the reward model to achieve high rewards.
One attributed source of over-optimization in DAAs is the under-constrained nature of their offline optimization, which can gradually shift probability mass toward non-preferred responses not presented in the preference dataset. This paper proposes a novel importance-sampling approach to mitigate the distribution shift problem of offline DAAs.
This approach, called (IS-DAAs), multiplies the DAA objective with an importance ratio that accounts for the reference policy distribution. IS-DAAs additionally avoid the high variance issue associated with importance sampling by clipping the importance ratio to a maximum value. Our extensive experiments demonstrate that IS-DAAs can effectively mitigate over-optimization, especially under low regularization strength, and achieve better performance than other methods designed to address this problem. Phuc Minh Nguyen, Ngoc-Hieu Nguyen, Duy M. H. Nguyen, Anji Liu, An Mai, Binh T. Nguyen 0001, Daniel Sonntag, Khoa D. Doan |
NeurIPS | 3 |
| 2025 | How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?abstractRecent advances in 3D point cloud transformers have led to state-of-the-art results in tasks such as semantic segmentation and reconstruction. However, these models typically rely on dense token representations, incurring high computational and memory costs during training and inference. In this work, we present the finding that tokens are remarkably redundant, leading to substantial inefficiency. We introduce \textbf{GitMerge3D}, a \textbf{g}lobally \textbf{i}nformed graph \textbf{t}oken \textbf{merging} method that can reduce the token count by up to 90–95\% while maintaining competitive performance. This finding challenges the prevailing assumption that more tokens inherently yield better performance and highlights that many current models are over-tokenized and under-optimized for scalability. We validate our method across multiple 3D vision tasks and show consistent improvements in computational efficiency. This work is the first to assess redundancy in large-scale 3D transformer models, providing insights into the development of more efficient 3D foundation architectures. Our code and checkpoints are publicly available at \href{https://gitmerge3d.github.io/}{https://gitmerge3d.github.io}. Duy M. H. Nguyen, Hoai-Chau Tran, Michael Barz, Khoa D. Doan, Roger Wattenhofer, Ngo Anh Vien, Mathias Niepert, Daniel Sonntag, Paul Swoboda |
NeurIPS | 2 |
| 2024 | Dude: Dual Distribution-Aware Context Prompt Learning For Large Vision-Language Model
Duy M. H. Nguyen, An T. Le 0001, Trung Quoc Nguyen, Nghiem Tuong Diep, Tai Nguyen 0008, Duy Duong-Tran, Jan Peters 0001, Li Shen 0001, Mathias Niepert, Daniel Sonntag |
ACML | 1 |
| 2024 | Structure-Aware E(3)-Invariant Molecular Conformer Aggregation NetworksabstractA molecule’s 2D representation consists of its atoms, their attributes, and the molecule’s covalent bonds. A 3D (geometric) representation of a molecule is called a conformer and consists of its atom types and Cartesian coordinates. Every conformer has a potential energy, and the lower this energy, the more likely it occurs in nature. Most existing machine learning methods for molecular property prediction consider either 2D molecular graphs or 3D conformer structure representations in isolation. Inspired by recent work on using ensembles of conformers in conjunction with 2D graph representations, we propose E(3)-invariant molecular conformer aggregation networks. The method integrates a molecule’s 2D representation with that of multiple of its conformers. Contrary to prior work, we propose a novel 2D–3D aggregation mechanism based on a differentiable solver for the Fused Gromov-Wasserstein Barycenter problem and the use of an efficient conformer generation method based on distance geometry. We show that the proposed aggregation mechanism is E(3) invariant and propose an efficient GPU implementation. Moreover, we demonstrate that the aggregation mechanism helps to significantly outperform state-of-the-art molecule property prediction methods on established datasets. Duy M. H. Nguyen, Nina Lukashina, Tai Nguyen 0008, An T. Le 0001, TrungTin Nguyen, Nhat Ho, Jan Peters 0001, Daniel Sonntag, Viktor Zaverkin, Mathias Niepert |
ICML | 1 |
| 2024 | Accelerating Transformers with Spectrum-Preserving Token MergingabstractIncreasing the throughput of the Transformer architecture, a foundational component used in numerous state-of-the-art models for vision and language tasks (e.g., GPT, LLaVa), is an important problem in machine learning. One recent and effective strategy is to merge token representations within Transformer models, aiming to reduce computational and memory requirements while maintaining accuracy. Prior work has proposed algorithms based on Bipartite Soft Matching (BSM), which divides tokens into distinct sets and merges the top $k$ similar tokens. However, these methods have significant drawbacks, such as sensitivity to token-splitting strategies and damage to informative tokens in later layers. This paper presents a novel paradigm called PiToMe, which prioritizes the preservation of informative tokens using an additional metric termed the \textit{energy score}. This score identifies large clusters of similar tokens as high-energy, indicating potential candidates for merging, while smaller (unique and isolated) clusters are considered as low-energy and preserved. Experimental findings demonstrate that PiToMe saved from 40-60\% FLOPs of the base models while exhibiting superior off-the-shelf performance on image classification (0.5\% average performance drop of ViT-MAEH compared to 2.6\% as baselines), image-text retrieval (0.3\% average performance drop of Clip on Flick30k compared to 4.5\% as others), and analogously in visual questions answering with LLaVa-7B. Furthermore, PiToMe is theoretically shown to preserve intrinsic spectral properties to the original token space under mild conditions. Chau Tran, Duy M. H. Nguyen, Duy Nguyen 0003, TrungTin Nguyen, T. Hoang Ngan Le, Pengtao Xie, Daniel Sonntag, James Zou 0001, Mathias Niepert |
NeurIPS | 2 |
| 2023 | Joint Self-Supervised Image-Volume Representation Learning with Intra-inter Contrastive ClusteringabstractCollecting large-scale medical datasets with fully annotated samples for training of deep networks is prohibitively expensive, especially for 3D volume data. Recent breakthroughs in self-supervised learning (SSL) offer the ability to overcome the lack of labeled training samples by learning feature representations from unlabeled data. However, most current SSL techniques in the medical field have been designed for either 2D images or 3D volumes. In practice, this restricts the capability to fully leverage unlabeled data from numerous sources, which may include both 2D and 3D data. Additionally, the use of these pre-trained networks is constrained to downstream tasks with compatible data dimensions. In this paper, we propose a novel framework for unsupervised joint learning on 2D and 3D data modalities. Given a set of 2D images or 2D slices extracted from 3D volumes, we construct an SSL task based on a 2D contrastive clustering problem for distinct classes. The 3D volumes are exploited by computing vectored embedding at each slice and then assembling a holistic feature through deformable self-attention mechanisms in Transformer, allowing incorporating long-range dependencies between slices inside 3D volumes. These holistic features are further utilized to define a novel 3D clustering agreement-based SSL task and masking embedding prediction inspired by pre-trained language models. Experiments on downstream tasks, such as 3D brain segmentation, lung nodule detection, 3D heart structures segmentation, and abnormal chest X-ray detection, demonstrate the effectiveness of our joint 2D and 3D SSL approach. We improve plain 2D Deep-ClusterV2 and SwAV by a significant margin and also surpass various modern 2D and 3D SSL approaches. Duy M. H. Nguyen, Truong Thanh Nhat Mai, Tri Cao, Binh T. Nguyen 0001, Nhat Ho, Paul Swoboda, Shadi Albarqouni, Pengtao Xie, Daniel Sonntag |
AAAI | 1 |
| 2023 | LVM-Med: Learning Large-Scale Self-Supervised Vision Models for Medical Imaging via Second-order Graph MatchingabstractObtaining large pre-trained models that can be fine-tuned to new tasks with limited annotated samples has remained an open challenge for medical imaging data. While pre-trained networks on ImageNet and vision-language foundation models trained on web-scale data are the prevailing approaches, their effectiveness on medical tasks is limited due to the significant domain shift between natural and medical images. To bridge this gap, we introduce LVM-Med, the first family of deep networks trained on large-scale medical datasets. We have collected approximately 1.3 million medical images from 55 publicly available datasets, covering a large number of organs and modalities such as CT, MRI, X-ray, and Ultrasound. We benchmark several state-of-the-art self-supervised algorithms on this dataset and propose a novel self-supervised contrastive learning algorithm using a graph-matching formulation. The proposed approach makes three contributions: (i) it integrates prior pair-wise image similarity metrics based on local and global information; (ii) it captures the structural constraints of feature embeddings through a loss function constructed through a combinatorial graph-matching objective, and (iii) it can be trained efficiently end-to-end using modern gradient-estimation techniques for black-box solvers. We thoroughly evaluate the proposed LVM-Med on 15 downstream medical tasks ranging from segmentation and classification to object detection, and both for the in and out-of-distribution settings. LVM-Med empirically outperforms a number of state-of-the-art supervised, self-supervised, and foundation models. For challenging tasks such as Brain Tumor Classification or Diabetic Retinopathy Grading, LVM-Med improves previous vision-language models trained on 1 billion masks by 6-7% while using only a ResNet-50. Duy M. H. Nguyen, Nghiem Tuong Diep, Tan Ngoc Pham, Tri Cao, Binh T. Nguyen 0001, Paul Swoboda, Nhat Ho, Shadi Albarqouni, Pengtao Xie, Daniel Sonntag, Mathias Niepert |
NeurIPS | 1 |
| 2022 | LMGP: Lifted Multicut Meets Geometry Projections for Multi-Camera Multi-Object TrackingabstractMulti-Camera Multi-Object Tracking is currently drawing attention in the computer vision field due to its superior performance in real-world applications such as video surveillance with crowded scenes or in wide spaces. In this work, we propose a mathematically elegant multi-camera multiple object tracking approach based on a spatial-temporal lifted multicut formulation. Our model utilizes state-of-the-art tracklets produced by single-camera trackers as proposals. As these tracklets may contain ID-Switch errors, we refine them through a novel pre-clustering obtained from 3D geometry projections. As a result, we derive a better tracking graph without ID switches and more precise affinity costs for the data association phase. Tracklets are then matched to multi-camera trajectories by solving a global lifted multicut formulation that incorporates short and long-range temporal interactions on tracklets located in the same camera as well as inter-camera ones. Experimental results on the WildTrack dataset yield near-perfect performance, outperforming state-of-the-art trackers on Campus while being on par on the PETS-09 dataset. We will release our implementations at this link https://github.com/nhmduy/LMGP. Duy M. H. Nguyen, Roberto Henschel, Bodo Rosenhahn, Daniel Sonntag, Paul Swoboda |
CVPR | 1 |
| 2022 | ASMCNN: An efficient brain extraction using active shape model and convolutional neural networks
Duy M. H. Nguyen, Duy M. Nguyen, Truong Thanh Nhat Mai, Thu Nguyen 0001, Khanh T. Tran, Anh Triet Nguyen, Bao T. Pham, Binh T. Nguyen 0001 |
Inf. Sci. | 1 |
| 2022 | TATL: Task agnostic transfer learning for skin attributes detection
Duy M. H. Nguyen, Thu T. Nguyen, Huong Vu, Quang Pham, Duy Nguyen 0003, Binh T. Nguyen 0001, Daniel Sonntag |
Medical Image Anal. | 1 |
| 2021 | EPEM: Efficient Parameter Estimation for Multiple Class Monotone Missing Data
Thu Nguyen 0001, Duy M. H. Nguyen, Binh T. Nguyen 0001, Bruce A. Wade |
Inf. Sci. | 2 |
| 2020 | Posterior concentration and fast convergence rates for generalized Bayesian learning
Lam Si Tung Ho, Binh T. Nguyen 0001, Vu C. Dinh, Duy M. H. Nguyen |
Inf. Sci. | 4 |
| 2019 | Automatically Generate Hymns Using Variational Attention Models
Han K. Cao, Duyen T. Ly, Duy M. H. Nguyen, Binh T. Nguyen 0001 |
ISNN (2) | 3 |
| 2019 | An active learning framework for set inversion
Binh T. Nguyen 0001, Duy M. H. Nguyen, Lam Si Tung Ho, Vu C. Dinh |
Knowl. Based Syst. | 2 |
| 2018 | OASIS: An Active Framework for Set InversionabstractIn this work, we introduce a novel method for solving the set inversion problem by formulating it as a binary classification problem. Aiming to develop a fast algorithm that can work effectively with high-dimensional and computationally expensive nonlinear models, we focus on active learning, a family of new and powerful techniques which can achieve the same level of accuracy with fewer data points compared to traditional learning methods. Specifically, we propose OASIS, an active learning framework using Support Vector Machine algorithms for solving the problem of set inversion. Our method works well in high dimensions and its computational cost is relatively robust to the increase of dimension. We illustrate the performance of OASIS by several simulation studies and show that our algorithm outperforms VISIA, the state-of-the-art method. Binh T. Nguyen 0001, Duy M. H. Nguyen, Lam Si Tung Ho, Vu C. Dinh |
SoMeT | 2 |
| 2017 | 3D-Brain Segmentation Using Deep Neural Network and Gaussian Mixture ModelabstractAutomatic segmentation of major brain tissues from high-resolution magnetic resonance images (MRIs) plays an important role in clinical diagnostics and neuroscience research. In this paper, we present a novel approach to extract brain tissues including gray matter, white matter and cerebrospinal fluid by using Gaussian mixture models (GMMs), Convolution neural networks (CNNs) and Deep neural networks (DNNs). GMMs are applied to classify voxels which have distinct intensity information and are easy to recognize while DNNs and CNNs are treating voxels which are similar in appearance and usually recognized insufficiently by traditional approaches. The empirical results on IBSR 18 dataset show that the proposed method outperforms 13 state-of-the-art algorithms, surpassing all other methods by a significant margin. Duy M. H. Nguyen, Huy T. Vu, Quang Huy Ung, Binh T. Nguyen 0001 |
WACV | 1 |
| 2016 | Fast learning rates with heavy-tailed lossesabstractWe study fast learning rates when the losses are not necessarily bounded and may have a distribution with heavy tails. To enable such analyses, we introduce two new conditions: (i) the envelope function $\sup_{f \in \mathcal{F}}|\ell \circ f|$, where $\ell$ is the loss function and $\mathcal{F}$ is the hypothesis class, exists and is $L^r$-integrable, and (ii) $\ell$ satisfies the multi-scale Bernstein's condition on $\mathcal{F}$. Under these assumptions, we prove that learning rate faster than $O(n^{-1/2})$ can be obtained and, depending on $r$ and the multi-scale Bernstein's powers, can be arbitrarily close to $O(n^{-1})$. We then verify these assumptions and derive fast learning rates for the problem of vector quantization by $k$-means clustering with heavy-tailed distributions. The analyses enable us to obtain novel learning rates that extend and complement existing results in the literature from both theoretical and practical viewpoints. Vu C. Dinh, Lam Si Tung Ho, Binh T. Nguyen 0001, Duy M. H. Nguyen |
NIPS | 4 |
| 2015 | Learning from Non-iid Data: Fast Rates for the One-vs-All Multiclass Plug-in Classifiers
Vu C. Dinh, Lam Si Tung Ho, Viet Cuong Nguyen, Duy M. H. Nguyen, Binh T. Nguyen 0001 |
TAMC | 4 |