Hai-Dang Nguyen

dblp:165/3880 · DBLP profile ↗
← Back
24ranked-venue papers
0as first author
21since 2021 · last 2026
0000-0003-0888-8908ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 16 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DiffCAS: Inference-time CT-free diffusion model for physics-aware multi-slice attenuation correction in cardiac SPECT
abstract
Attenuation artifacts remain a critical challenge in cardiac Myocardial Perfusion Imaging (MPI) using Single-Photon Emission Computed Tomography (SPECT), often degrading diagnostic accuracy and clinical interpretability. While hybrid SPECT and Computed Tomography (CT) systems mitigate these artifacts using CT-derived attenuation maps, their high cost, radiation exposure, and limited accessibility restrict widespread clinical use. To address these challenges, we propose DiffCAS, an inference-time CT-free diffusion model for physics-aware multi-slice attenuation correction in cardiac SPECT. DiffCAS integrates a Brownian Bridge diffusion process with physics-guided supervision, enabling the generation of attenuation-corrected (AC) images directly from non-attenuation-corrected (NAC) inputs. Specifically, a physics-aware reconstruction module predicts voxel-wise attenuation coefficients and path lengths, then combines them via the Beer-Lambert law into an attenuation correction factor applied at each diffusion step, keeping the AC images physically consistent. The model introduces two key innovations that jointly enhance structural understanding and physics consistency. The first is multi-slice contextual learning, which captures cross-slice anatomical dependencies and improves spatial coherence in reconstructed images. The second is the 3D Computed Tomography Vision Transformer that models long-range volumetric structures and provides physics-consistent attenuation priors to guide the diffusion process. To enable CT-free attenuation correction, DiffCAS introduces the teacher-student distillation framework that transfers physics-informed knowledge from CT-conditioned training to a student network that requires no CT input at inference time, ensuring stability and interpretability. Evaluations on the CardiAC dataset, which comprises 424 patient studies with paired NAC and AC, and CT-based attenuation maps, demonstrate the strong performance of DiffCAS, evaluated using global pixel-level metrics and myocardium-specific clinical metrics. The proposed method surpasses state-of-the-art image generative methods, achieving superior reconstruction accuracy, structural consistency, and diagnostic reliability. These results highlight the proposed DiffCAS as a clinically promising, inference-time CT-free solution for attenuation correction in cardiac SPECT imaging.
Hoang Minh Vu, Trung-Kien Pham, Thi Ha Chi Nguyen, Hai-Dang Nguyen, Dac Thai Nguyen, Hong Son Mai, Thanh Trung Nguyen, Trung Thanh Nguyen 0006, Phi-Le Nguyen
Artif. Intell. Medicine4
2026 FUGC: Benchmarking Semi-Supervised Learning Methods for Cervical Segmentation
abstract
Accurate segmentation of cervical structures in transvaginal ultrasound (TVS) is critical for assessing the risk of spontaneous preterm birth (PTB), yet the scarcity of labeled data limits the performance of supervised learning approaches. This paper introduces the Fetal Ultrasound Grand Challenge (FUGC), the first benchmark for semi-supervised learning in cervical segmentation, hosted at ISBI 2025. FUGC provides a dataset of 890 TVS images, including 500 training images, 90 validation images, and 300 test images. Methods were evaluated using the Dice Similarity Coefficient (DSC), Hausdorff Distance (HD), and runtime (RT), with a weighted combination of 0.4/0.4/0.2. The challenge attracted 10 teams with 82 participants submitting innovative solutions. The best-performing methods for each individual metric achieved 90.26% mDSC, 38.88 mHD, and 32.85 ms RT, respectively. FUGC establishes a standardized benchmark for cervical segmentation, demonstrates the efficacy of semi-supervised methods with limited labeled data, and provides a foundation for AI-assisted clinical PTB risk assessment.
Jieyun Bai, Yitong Tang, Mahdi Islam, Musarrat Tabassum, Enrique Almar-Munoz, Nianjiang Lv, Yu Chen 0099, Zilun Peng, Yusong Xiao, Li Xiao 0002, Nam-Khanh Tran, Dac-Phu Phan-Le, Hai-Dang Nguyen, Xiao Liu 0037, Jiale Hu, Mingxu Huang, Jitao Liang, Chaolu Feng, Xuezhi Zhang, Lyuyang Tong, Bo Du 0001, Ha-Hieu Pham, Thanh-Huy Nguyen, Min Xu 0009, Juntao Jiang, Jiangning Zhang, Yong Liu 0007, Md. Kamrul Hasan 0002, Zhuonan Liang, Tom Weidong Cai, Gongning Luo, Mohammad Yaqub, Karim Lekadir
IEEE Trans. Medical Imaging17
2025 Can Frequency Filtering Approximate CNNs for Enhancing Segment Anything?
Xuan-Tung Nguyen 0002, Hoang-Minh Nguyen, Trong-Hieu Nguyen Mau, Tuong-Vy Truong-Thuy, Ngoc-Linh Nguyen-Ha, Minh-Triet Tran, Hai-Dang Nguyen
ICCCI (2)7
2024 Rethinking Sampling for Music-Driven Long-Term Dance Generation
Tuong-Vy Truong-Thuy, Gia-Cat Bui-Le, Hai-Dang Nguyen, Trung-Nghia Le
ACCV (5)3
2024 PGDS: Pose-Guidance Deep Supervision for Mitigating Clothes-Changing in Person Re-Identification
abstract
Person Re-Identification (Re-ID) task seeks to enhance the tracking of multiple individuals by surveillance cameras. It supports multimodal tasks, including text-based person retrieval and human matching. One of the most significant challenges faced in Re-ID is clothes-changing, where the same person may appear in different outfits. While previous methods have made notable progress in maintaining clothing data consistency and handling clothing change data, they still rely excessively on clothing information, which can limit performance due to the dynamic nature of human appearances. To mitigate this challenge, we propose the Pose-Guidance Deep Supervision (PGDS), an effective framework for learning pose guidance within the Re-ID task. It consists of three modules: a human encoder, a pose encoder, and a Pose-to-Human Projection module(PHP). Our framework guides the human encoder, i.e., the main re-identification model, with pose information from the pose encoder through multiple layers via the knowledge transfer mechanism from the PHP module, helping the human encoder learn body parts information without increasing computation resources in the inference stage. Through extensive experiments, our method surpasses the performance of current state-of-the-art methods, demonstrating its robustness and effectiveness for real-world applications. Our code is available at https://github.com/huyquoctrinh/PGDS.
Quoc-Huy Trinh, Nhat-Tan Bui, Dinh-Hieu Hoang, Phuoc-Thao Vo Thi, Hai-Dang Nguyen, Debesh Jha, Ulas Bagci, T. Hoang Ngan Le, Minh-Triet Tran
AVSS5
2024 SAM-EG: Segment Anything Model with Egde Guidance framework for efficient Polyp Segmentation
Quoc-Huy Trinh, Hai-Dang Nguyen, Bao-Tram Nguyen Ngoc, Debesh Jha, Ulas Bagci, Minh-Triet Tran
BMVC2
2024 Pose-to-Human (P2H): A pose guidance framework via Gram matrix for Occluded Person Re-identification
abstract
The person identification task aims to generate robust human representation embeddings. Traditional methods have achieved competitive results, but they face occlusion challenges, making it difficult to distinguish between humans occluded by objects or other humans. This paper proposes Pose-to-Human (P2H), a pose guidance framework via Gram matrix and knowledge transfer mechanism to tackle occlusion problems in Person Re-identification to address this issue. Our framework comprises four modules: Human Encoder, Pose Encoder, Compact Modules (CPM), and the Gram Matrix Guiding (GMG). The Pose Encoder generates pose information used to guide the human embedding models, facilitating learning of other parts of the human anatomy, which enables the model to focus on the unoccluded parts of the body, thus mitigating the reliance on the data for part-to-part matching, which is the first limitation of previous works. Additionally, it can reduce the demand for general appearance information from the human encoder model. Our method achieves competitive results through extensive experiments on five datasets compared to state-of-the-art methods, making it a promising framework for addressing the Re-Identification task.
Quoc-Huy Trinh, Phuoc-Thao Vo Thi, Minh-Triet Tran, Hai-Dang Nguyen
KES4
2023 Graph for Transformer Feature: A New Approach for Face Anti-Spoofing
abstract
Face recognition is popular nowadays, however, Face antispoofing (FAS) poses a significant challenge for recognition systems due to the threat of external attacks.While many deep learning methods have been proposed to address this issue, they often face challenges in industry settings.Experiments found that patch extraction modules, such as the Vision Transformer and Swin Transformer, are effective for FAS in single images and perform well in industrial environments.From this point, we propose a model that leverages Transformer features and Graph Neural Networks to learn global information and identify correlations between patch features, which are critical for FAS.
Quoc-Huy Trinh, Xuan-Mao Nguyen, Hai-Dang Nguyen
ESANN5
2023 PEFNet: Positional Embedding Feature for Polyp Segmentation
Trong-Hieu Nguyen Mau, Quoc-Huy Trinh, Nhat-Tan Bui, Phuoc-Thao Vo Thi, Minh-Van Nguyen, Xuan-Nam Cao, Minh-Triet Tran, Hai-Dang Nguyen
MMM (2)8
2023 TextANIMAR: Text-based 3D animal fine-grained retrieval
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Quan Le, Viet-Tham Huynh, Trong-Le Do, Khanh-Duy Le, Mai-Khiem Tran, Nhat Hoang-Xuan, Thang-Long Nguyen-Ho, Vinh-Tiep Nguyen, Tuong-Nghiem Diep, Khanh-Duy Ho, Xuan-Hieu Nguyen, Thien-Phuc Tran, Tuan-Anh Yang, Kim-Phat Tran, Nhu-Vinh Hoang, Minh-Quang Nguyen, E-Ro Nguyen, Minh-Khoi Nguyen-Nhat, Tuan-An To, Trung-Truc Huynh-Le, Nham-Tan Nguyen, Hoang-Chau Luong, Truong Hoai Phong, Nhat-Quynh Le-Pham, Huu-Phuc Pham, Trong-Vu Hoang, Quang-Binh Nguyen, Hai-Dang Nguyen, Akihiro Sugimoto, Minh-Triet Tran
Comput. Graph.31
2023 SketchANIMAR: Sketch-based 3D animal fine-grained retrieval
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Quan Le, Viet-Tham Huynh, Trong-Le Do, Khanh-Duy Le, Mai-Khiem Tran, Nhat Hoang-Xuan, Thang-Long Nguyen-Ho, Vinh-Tiep Nguyen, Nhat-Quynh Le-Pham, Huu-Phuc Pham, Trong-Vu Hoang, Quang-Binh Nguyen, Trong-Hieu Nguyen Mau, Tuan-Luc Huynh, Thanh-Danh Le, Ngoc-Linh Nguyen-Ha, Tuong-Vy Truong-Thuy, Truong Hoai Phong, Tuong-Nghiem Diep, Khanh-Duy Ho, Xuan-Hieu Nguyen, Thien-Phuc Tran, Tuan-Anh Yang, Kim-Phat Tran, Nhu-Vinh Hoang, Minh-Quang Nguyen, Hoai-Danh Vo, Minh-Hoa Doan, Hai-Dang Nguyen, Akihiro Sugimoto, Minh-Triet Tran
Comput. Graph.32
2022 Efficient loss functions for GAN-based style transfer
abstract
Style transfer aims to render a new artistic image based on a content image and given artwork style. Recent style transfer techniques often suffer structure distortion and artifact problems that abate the quality of stylized images. Motivated by these observations and the previous works, we introduce a novel GAN framework to enhance the aesthetics, faithfulness and flexibility in the style transfer process. The key factor of our model is the Laplacian Pyramid loss that naturally forces the content preservation and the ResidualStyle discriminator block to capture the artwork’s painting style better. In contrast to existing methods that calculate the Euclidean distance between the features of generated image and content image, our Laplacian Pyramid loss better captures the content representation by different frequency bands of the content image. As evaluated by experimental results, our framework surmounts the unrealistic artifacts to synthesize the photorealistic artworks in real-time, hence attaining striking visual effects.
Nhat-Tan Bui, Hai-Dang Nguyen, Trung-Nam Bui Huynh, Ngoc-Thao Nguyen, Xuan-Nam Cao
ICMV2
2022 Tiny convolution contextual neural network: a lightweight model for skin lesion detection
abstract
Skin Lesion is a controversial disease all over the world, particularly Melanoma which is a kind of skin cancer. In recent years, there are several methods of using the Convolutional Neural Network and Vision Transformer model have been proposed for the detection and classification of skin images and have achieved competitive results. In this paper, we introduce and demonstrate the efficiency of the Tiny Convolution Contextual Neural Network (TCC Neural Network) which is a tiny model with light-weight architecture than the previous architecture, and have fewer parameter than the popular model for classification of nine lesions from skin images. Our proposal achieves 0.75 on accuracy and 0.55 on F1 score with 5.6 million parameters in the Skin Lesion classification task.
Quoc-Huy Trinh, Trong-Hieu Nguyen Mau, Phuoc-Thao Vo Thi, Hai-Dang Nguyen
ICMV4
2022 SHREC 2022 track on online detection of heterogeneous gestures
Marco Emporio, Ariel Caputo, Andrea Giachetti 0001, Marco Cristani, Guido Borghi, Andrea D'Eusanio, Minh-Quan Le, Hai-Dang Nguyen, Minh-Triet Tran, Felix Ambellan, Martin Hanik, Esfandiar Nava-Yazdani, Christoph von Tycowicz
Comput. Graph.8
2022 SHREC'22 track: Open-Set 3D Object Retrieval
Yifan Feng 0001, Yue Gao 0002, Xibin Zhao, Yandong Guo, Nihar Bagewadi, Nhat-Tan Bui, Hieu Dao, Shankar Gangisetty, Ripeng Guan, Xie Han 0001, Cong Hua, Chidambar Hunakunti, Yu Jiang 0006, Shichao Jiao, Yuqi Ke, Liqun Kuang, Anan Liu, Dinh-Huan Nguyen, Hai-Dang Nguyen, Weizhi Nie, Bang-Dang Pham, Karthik Raikar, Qingmei Tang, Minh-Triet Tran, Jialong Wan, Chenggang Yan 0001, Haoxuan You, Difei Zhu
Comput. Graph.19
2022 SHREC'22 track: Sketch-based 3D shape retrieval in the wild
Jie Qin 0004, Shuaihang Yuan, Jiaxin Chen 0002, Boulbaba Ben Amor, Yi Fang 0006, Nhat Hoang-Xuan, Chi-Bien Chu, Khoi-Nguyen Nguyen-Ngoc, Thien-Tri Cao, Nhat-Khang Ngô, Tuan-Luc Huynh, Hai-Dang Nguyen, Minh-Triet Tran, Haoyang Luo, Jianning Wang, Zheng Zhang 0006, Zihao Xin, Yang Wang 0023, Haiqin Chen, Qunying Zhou
Comput. Graph.12
2022 SHREC 2022: Fitting and recognition of simple geometric primitives on point clouds
Chiara Romanengo, Andrea Raffo, Silvia Biasotti, Bianca Falcidieno, Vlassis Fotis, Ioannis Romanelis, Eleftheria Psatha, Konstantinos Moustakas, Ivan Sipiran, Chi-Bien Chu, Khoi-Nguyen Nguyen-Ngoc, Dinh-Khoi Vo, Tuan-An To, Nham-Tan Nguyen, Nhat-Quynh Le-Pham, Hai-Dang Nguyen, Minh-Triet Tran, Yifan Qie, Nabil Anwer
Comput. Graph.17
2022 SHREC 2022: Pothole and crack detection in the road pavement using images and RGB-D data
Elia Moscoso Thompson, Andrea Ranieri, Silvia Biasotti, Miguel Chicchón, Ivan Sipiran, Minh-Khoi Pham, Thang-Long Nguyen-Ho, Hai-Dang Nguyen, Minh-Triet Tran
Comput. Graph.8
2021 Effectiveness of Detection-based and Regression-based Approaches for Estimating Mask-Wearing Ratio
abstract
Estimating the mask-wearing ratio in public places is important as it enables health authorities to promptly analyze and implement policies. Methods for estimating the mask-wearing ratio on the basis of image analysis have been reported. However, there is still a lack of comprehensive research on both methodologies and datasets. Most recent reports straightforwardly propose estimating the ratio by applying conventional object detection and classification methods. It is feasible to use regression-based approaches to estimate the number of people wearing masks, especially for congested scenes with tiny and occluded faces, but this has not been well studied. A large-scale and well-annotated dataset is still in demand. In this paper, we proposed two different methods for ratio estimation that are leveraged by either detection-based or regression-based approaches. For the detection-based approach, we improved a state-of-the-art face detector, RetinaFace, for the ratio estimation. For the regression-based approach, we utilized a baseline network, CSRNet, and finetuned it to estimate the density maps for masked and unmasked faces. We also proposed the first large-scale dataset, the “NFM,” which contains 581,108 face annotations extracted from 18,088 video frames in 17 street-view videos11The annotations (bounding boxes and labels), and pre-trained models will be released with the publication of our paper. , Through experiments, the RetinaFace-based method achieves better accuracy under different situations, while the CSRNet-based method is superior in terms of operation time thanks to its compactness.
Khanh-Duy Nguyen, Hai-Dang Nguyen, Trung-Nghia Le, Junichi Yamagishi, Isao Echizen
FG2
2021 SHREC 2021: Skeleton-based hand gesture recognition in the wild
Ariel Caputo, Andrea Giachetti 0001, Simone Soso, Deborah Pintani, Andrea D'Eusanio, Stefano Pini, Guido Borghi, Alessandro Simoni, Roberto Vezzani, Rita Cucchiara, Andrea Ranieri, Franca Giannini, Katia Lupinetti, Marina Monti, Mehran Maghoumi, Joseph J. LaViola Jr., Minh-Quan Le, Hai-Dang Nguyen, Minh-Triet Tran
Comput. Graph.18
2021 SHREC 2021: Retrieval and classification of protein surfaces equipped with physical and chemical properties
Andrea Raffo, Ulderico Fugacci, Silvia Biasotti, Walter Rocchia, Yonghuai Liu, Ekpo Otu, Reyer Zwiggelaar, David Hunter, Evangelia I. Zacharaki, Eleftheria Psatha, Dimitrios Laskos, Gerasimos Arvanitis, Konstantinos Moustakas, Tunde Aderinwale, Charles Christoffer, Woong-Hee Shin, Daisuke Kihara, Andrea Giachetti 0001, Huu-Nghia Nguyen, Tuan-Duy Nguyen, Vinh-Thuyen Nguyen-Truong, Danh Le-Thanh, Hai-Dang Nguyen, Minh-Triet Tran
Comput. Graph.23
2020 SHREC 2020: Retrieval of digital surfaces with similar geometric reliefs
Elia Moscoso Thompson, Silvia Biasotti, Andrea Giachetti 0001, Claudio Tortorici, Naoufel Werghi, Ahmad Obeid 0001, Stefano Berretti, Hoang-Phuc Nguyen-Dinh, Minh-Quan Le, Hai-Dang Nguyen, Minh-Triet Tran, Leonardo Gigli, Santiago Velasco-Forero, Beatriz Marcotegui, Ivan Sipiran, Benjamin Bustos, Ioannis Romanelis, Vlassis Fotis, Ramamoorthy Luxman
Comput. Graph.10
2019 Enhancing Endoscopic Image Classification with Symptom Localization and Data Augmentation
abstract
Inspired by recent advances in computer vision and deep learning, we propose new enhancements to tackle problems appearing in endoscopic image analysis, especially abnormality finding and anatomical landmark detection. In details, a combination of Residual Neural Network and Faster R-CNN are jointly applied in order to take all of their advantages and improve the overall performance. Nevertheless, novel data augmentation is designed and adapted to corresponding domains. Our approaches prove their competitive results in term of not only the accuracy but also the inference time in Medico: The 2018 Multimedia for Medicine Task and The Biomedia ACM MM Grand Challenge 2019. These results show the great potential of the collaborating between deep learning models and data augmentation in medical image analysis applications. Especially, more than 4900 bounding boxes localizing the symptom of some classes from KVASIR dataset that we annotated and used in this project are shared online for future research.
Trung-Hieu Hoang, Hai-Dang Nguyen, Thanh-An Nguyen, Vinh-Tiep Nguyen, Minh-Triet Tran
ACM Multimedia2
2014 Apply lightweight recognition algorithms in optical music recognition
abstract
The problems of digitalization and transformation of musical scores into machine-readable format are necessary to be solved since they help people to enjoy music, to learn music, to conserve music sheets, and even to assist music composers. However, the results of existing methods still require improvements for higher accuracy. Therefore, the authors propose lightweight algorithms for Optical Music Recognition to help people to recognize and automatically play musical scores. In our proposal, after removing staff lines and extracting symbols, each music symbol is represented as a grid of identical M ∗ N cells, and the features are extracted and classified with multiple lightweight SVM classifiers. Through experiments, the authors find that the size of 10 ∗ 12 cells yields the highest precision value. Experimental results on the dataset consisting of 4929 music symbols taken from 18 modern music sheets in the Synthetic Score Database show that our proposed method is able to classify printed musical scores with accuracy up to 99.56%.
Viet-Khoi Pham, Hai-Dang Nguyen, Tung-Anh Nguyen-Khac, Minh-Triet Tran
ICMV2