EDBT 2026 Demo / reviewers in the wild / expert
An Vuong
dblp:342/1527
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction
An Vuong, Minh-Hao Van, Chen Zhao 0010, Xintao Wu |
IEEE Big Data | 1 |
| 2025 | Perception-based multiplicative noise removal with Diffusion modelsabstractWe present a novel approach to perform multiplicative noise removal, utilizing the recent developments of diffusion models. We show that multiplicative noise, which commonly appears in images produced by synthetic aperture radar (SAR), laser, or optical lenses, can be well-modeled by a Geometric Brownian process in the logarithmic domain. This process admits a time-reversal stochastic differential equation (SDE), which is utilized to perform noise removal. We conduct extensive experiments to compare our approach with classical methods as well as state-of-the-art Deep Learning-based approaches. Our models significantly outperform others in terms of perception-based metrics such as LPIPS and FID, while remaining competitive in traditional pixel-based metrics like PSNR and SSIM. An Vuong, Thinh Nguyen |
ICMLA | 1 |
| 2025 | GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature LearningabstractGrasp detection is a fundamental robotic task critical to the success of many industrial applications. However, current language-driven models for this task often struggle with cluttered images, lengthy textual descriptions, or slow inference speed. We introduce GraspMamba, a new language-driven grasp detection method that employs hierarchical feature fusion with Mamba vision to tackle these challenges. By leveraging rich visual features of the Mamba-based backbone alongside textual information, our approach effectively enhances the fusion of multimodal features. GraspMamba represents the first Mamba-based grasp detection model to extract vision and language features at multiple scales, delivering robust performance and rapid inference time. Intensive experiments show that GraspMamba outperforms recent methods by a clear margin. We validate our approach through real-world robotic experiments, highlighting its fast inference speed. An Vuong, Anh Nguyen 0003, Ian D. Reid 0001, Minh Nhat Vu |
IROS | 2 |
| 2024 | Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
Toan Nguyen 0004, Minh Nhat Vu, Baoru Huang, An Vuong, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
ECCV (19) | 4 |
| 2024 | Lightweight Language-driven Grasp Detection using Conditional Consistency ModelabstractLanguage-driven grasp detection is a fundamental yet challenging task in robotics with various industrial applications. This work presents a new approach for language-driven grasp detection that leverages lightweight diffusion models to achieve fast inference time. By integrating diffusion processes with grasping prompts in natural language, our method can effectively encode visual and textual information, enabling more accurate and versatile grasp positioning that aligns well with the text query. To overcome the long inference time problem in diffusion models, we leverage the image and text features as the condition in the consistency model to reduce the number of denoising timesteps during inference. The intensive experimental results show that our method outperforms other recent grasp detection methods and lightweight diffusion models by a clear margin. We further validate our method in real-world robotic experiments to demonstrate its fast inference time capability. Minh Nhat Vu, Baoru Huang, An Vuong, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
IROS | 4 |
| 2024 | Language-driven Grasp Detection with Mask-guided AttentionabstractGrasp detection is an essential task in robotics with various industrial applications. However, traditional methods often struggle with occlusions and do not utilize language for grasping. Incorporating natural language into grasp detection remains a challenging task and largely unexplored. To address this gap, we propose a new method for language-driven grasp detection with mask-guided attention by utilizing the transformer attention mechanism with semantic segmentation features. Our approach integrates visual data, segmentation mask features, and natural language instructions, significantly improving grasp detection accuracy. Our work introduces a new framework for language-driven grasp detection, paving the way for language-driven robotic applications. Intensive experiments show that our method outperforms other recent baselines by a clear margin, with a 10.0% success score improvement. We further validate our method in real-world robotic experiments, confirming the effectiveness of our approach. Tuan Van Vo, Minh Nhat Vu, Baoru Huang, An Vuong, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
IROS | 4 |
| 2024 | HabiCrowd: A High Performance Simulator for Crowd-Aware Visual NavigationabstractVisual navigation, a foundational aspect of Embodied AI (E-AI) and robotics has been extensively studied in the past few years. While many 3D simulators have been introduced for the visual navigation tasks, scarcely works have combined human dynamics, creating the gap between simulation and real-world applications. Furthermore, current 3D simulators incorporating human dynamics have several limitations, particularly in terms of computational efficiency, which is a promise of modern simulators. To overcome these issues, we introduce HabiCrowd, the new standard benchmark for crowd-aware visual navigation that includes a crowd dynamics model with diverse human settings into photorealistic environments. Empirical evaluations demonstrate that our proposed human dynamics model achieves state-of-the-art performance in collision avoidance while exhibiting superior computational efficiency compared to its counterparts. We leverage HabiCrowd to conduct several comprehensive studies on crowd-aware visual navigation tasks and human-robot interactions. The source code and data can be found at https://habicrowd.github.io/. An Vuong, Toan Nguyen 0004, Minh Nhat Vu, Baoru Huang, Huynh Thi Thanh Binh, Thieu Vo, Anh Nguyen 0003 |
IROS | 1 |
| 2024 | On Minimizing Symbol Error Probability for Antipodal Beamforming in MIMO Gaussian Wiretap ChannelsabstractThis paper investigates a beamforming scheme designed to minimize the symbol error probability (SEP) for a legitimate user while guaranteeing that the likelihood of an eavesdropper correctly recovering symbols remains below a predefined threshold. The focus is on finding an optimal beamforming vector for binary antipodal signal detection in multiple-input multiple-output (MIMO) Gaussian wiretap channels. Finding the optimal beamforming vector is a non-convex problem, and thus conventional computationally efficient algorithms for convex problems cannot be applied in this context. To that end, our proposed algorithm relies on Karush–Kuhn–Tucker (KKT) conditions and the generalized eigen-decomposition method to find an exact solution. The numerical results are presented to assess the performance of the proposed method for various scenarios. Nam Nguyen 0004, An Vuong, Thuan Nguyen 0001, Thinh Nguyen |
VTC Fall | 2 |
| 2023 | Open-Vocabulary Affordance Detection in 3D Point CloudsabstractAffordance detection is a challenging problem with a wide variety of robotic applications. Traditional affordance detection methods are limited to a predefined set of affordance labels, hence potentially restricting the adaptability of intelligent robots in complex and dynamic environments. In this paper, we present the Open-Vocabulary Affordance Detection (OpenAD) method, which is capable of detecting an unbounded number of affordances in 3D point clouds. By simultaneously learning the affordance text and the point feature, OpenAD successfully exploits the semantic relationships between affordances. Therefore, our proposed method enables zero-shot detection and can be able to detect previously unseen affordances without a single annotation example. Intensive experimental results show that OpenAD works effectively on a wide range of affordance detection setups and outperforms other baselines by a large margin. Additionally, we demonstrate the practicality of the proposed OpenAD in real-world robotic applications with a fast inference speed. Our project is available at https://openad2023.github.io. Toan Nguyen 0004, Minh Nhat Vu, An Vuong, Dzung Nguyen, Thieu Vo, T. Hoang Ngan Le, Anh Nguyen 0003 |
IROS | 3 |
| 2023 | Capacity achieving quantizer design for multiple-input multiple-output thresholding channelsabstractWe consider a communication channel whose input is modeled as a discrete random variable X with distribution pX. X is transmitted over a noisy channel and distorted by a continuous-valued noise to result in a continuous-valued output signal U at the receiver. A thresholding quantizer Q is applied to reconstruct a discrete signal V = Q(U) from the continuous-valued U. Our goal is to jointly design both the input distribution pXand the thresholding quantizer Q to maximize the mutual information I(X; V) between the input X and V since the accuracy of any decoding algorithm that estimates X from V fundamentally depends on I(X; V). In this paper, an alternating maximization algorithm is proposed that guarantees to achieve a locally optimal solution. In addition, we numerically show that by randomly selecting a set of initial starting points, the proposed algorithm is capable of achieving the globally optimal solution. Both the theoretical and numerical results are provided to justify our approach. An Vuong, Thuan Nguyen 0001, Thinh Nguyen |
VTC2023-Spring | 1 |