EDBT 2026 Demo / reviewers in the wild / expert
Hieu H. Pham 0001
dblp:254/2861-1 · also Hieu Huy Pham 0001, Huy Hieu Pham 0001
· DBLP profile ↗
25ranked-venue papers
4as first author
22since 2021 · last 2025
0000-0003-4851-2518ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ConstStyle: Robust Domain Generalization with Unified Style TransformationabstractDeep neural networks often suffer performance drops when test data distribution differs from training data. Domain Generalization (DG) aims to address this by focusing on domain-invariant features or augmenting data for greater diversity. However, these methods often struggle with limited training domains or significant gaps between seen (training) and unseen (test) domains. To enhance DG robustness, we hypothesize that it is essential for the model to be trained on data from domains that closely resemble unseen test domains-an inherently difficult task due to the absence of prior knowledge about the unseen domains. Accordingly, we propose ConstStyle, a novel approach that leverages a unified domain to capture domain-invariant features and bridge the domain gap with theoretical analysis. During training, all samples are mapped onto this unified domain, optimized for seen domains. During testing, unseen domain samples are projected similarly before predictions. By aligning both training and testing data within this unified domain, ConstStyle effectively reduces the impact of domain shifts, even with large domain gaps or few seen domains. Extensive experiments demonstrate that ConstStyle consistently outperforms existing methods across diverse scenarios. Notably, when only a limited number of seen domains are available, ConstStyle can boost accuracy up to 19.82\% compared to the next best approach. Nam Duong Tran, Nam Nguyen Phuong, Hieu H. Pham 0001, Phi-Le Nguyen, My T. Thai |
ICCV | 3 |
| 2025 | NeurFlow: Interpreting Neural Networks through Neuron Groups and Functional InteractionsabstractUnderstanding the inner workings of neural networks is essential for enhancing model performance and interpretability. Current research predominantly focuses on examining the connection between individual neurons and the model's final predictions, which suffers from challenges in interpreting the internal workings of the model, particularly when neurons encode multiple unrelated features. In this paper, we propose a novel framework that transitions the focus from analyzing individual neurons to investigating groups of neurons, shifting the emphasis from neuron-output relationships to the functional interactions between neurons. Our automated framework, NeurFlow, first identifies core neurons and clusters them into groups based on shared functional relationships, enabling a more coherent and interpretable view of the network’s internal processes. This approach facilitates the construction of a hierarchical circuit representing neuron interactions across layers, thus improving interpretability while reducing computational costs. Our extensive empirical studies validate the fidelity of our proposed NeurFlow. Additionally, we showcase its utility in practical applications such as image debugging and automatic concept labeling, thereby highlighting its potential to advance the field of neural network explainability. Tue Minh Cao, Nhat Hoang-Xuan, Hieu H. Pham 0001, Phi-Le Nguyen, My T. Thai |
ICLR | 3 |
| 2025 | Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report GenerationabstractVision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains, applying these models to medical imaging remains challenging due to the limited availability of diverse imaging modalities and multilingual clinical data. Most existing medical VLMs are trained on a subset of imaging modalities and focus primarily on high-resource languages, thus limiting their generalizability and clinical utility. To address these limitations, we introduce a novel Vietnamese-language multimodal medical dataset consisting of 2,757 whole-body PET/CT volumes from independent patients and their corresponding full-length clinical reports. This dataset is designed to fill two pressing gaps in medical AI development: (1) the lack of PET/CT imaging data in existing VLMs training corpora, which hinders the development of models capable of handling functional imaging tasks; and (2) the underrepresentation of low-resource languages, particularly the Vietnamese language, in medical vision-language research. To the best of our knowledge, this is the first dataset to provide comprehensive PET/CT-report pairs in Vietnamese. We further introduce a training framework to enhance VLMs' learning, including data augmentation and expert-validated test sets. We conduct comprehensive experiments benchmarking state-of-the-art VLMs on downstream tasks, including medical report generation and visual question answering. The experimental results show that incorporating our dataset significantly improves the performance of existing VLMs. However, despite these advancements, the models still underperform on clinically critical criteria, particularly the diagnosis of lung cancer, indicating substantial room for future improvement. We believe this dataset and benchmark will serve as a pivotal step in advancing the development of more robust VLMs for medical imaging, particularly in low-resource languages, and improving their clinical relevance in Vietnamese healthcare. Huu Tien Nguyen, Dac Thai Nguyen, The Minh Duc Nguyen, Trung Thanh Nguyen 0006, Truong Thao Nguyen, Hieu H. Pham 0001, Johan Barthelemy, Minh Quan Tran, Nguyen Quoc Viet Hung, Thanh Tam Nguyen, Hong Son Mai, Quynh Anh Chau, Thanh Hong Nguyen, Phi-Le Nguyen |
NeurIPS | 6 |
| 2025 | CT to PET Translation: A Large-Scale Dataset and Domain-Knowledge-Guided Diffusion ApproachabstractPositron Emission Tomography (PET) and Computed Tomography (CT) are essential for diagnosing, staging, and monitoring various diseases, particularly cancer. Despite their importance, the use of PET/CT systems is limited by the necessity for radioactive materials, the scarcity of PET scanners, and the high cost associated with PET imaging. In contrast, CT scanners are more widely available and significantly less expensive. In response to these challenges, our study addresses the issue of generating PET images from CT images, aiming to reduce both the medical examination cost and the associated health risks for patients. Our contributions are twofold: First, we introduce a conditional diffusion model named CPDM, which, to our knowledge, is one of the initial attempts to employ a diffusion model for translating from CT to PET images. Second, we provide the largest CT-PET dataset to date, comprising 2,028,628 paired CT-PET images, which facilitates the training and evaluation of CT-to-PET translation models. For the CPDM model, we incorporate domain knowledge to develop two conditional maps: the Attention map and the Attenuation map. The former helps the diffusion process focus on areas of interest, while the latter improves PET data correction and ensures accurate diagnostic information. Experimental evaluations across various benchmarks demonstrate that CPDM surpasses existing methods in generating high-quality PET images in terms of multiple metrics. The source code and data samples are available at https://github.com/thanhhff/CPDM. Dac Thai Nguyen, Trung Thanh Nguyen 0006, Huu Tien Nguyen, Thanh Trung Nguyen, Hieu H. Pham 0001, Thanh-Hung Nguyen, Truong Thao Nguyen, Phi-Le Nguyen |
WACV | 5 |
| 2025 | Noisy data-based attack: A new type of untargeted attack in Federated Learning and its countermeasures
Manh Cuong Dao, Phi-Le Nguyen, Hieu H. Pham 0001, Thanh-Hung Nguyen, Peng Chen 0035, Mohamed Wahib, Truong Thao Nguyen |
Future Gener. Comput. Syst. | 3 |
| 2025 | SAFA: Handling Sparse and Scarce Data in Federated Learning With Accumulative LearningabstractFederated Learning (FL) has emerged as an effective paradigm allowing multiple parties to collaboratively train a global model while protecting their private data. However, it is observed that the performance of FL approaches tends to degrade significantly when data are sparsely distributed across clients with small datasets. This is referred to as the sparse-and-scarce challenge, where data held by each client is both sparse (does not contain examples to all classes) and scarce (small dataset). Sparse-and-scarce data diminishes the generalizability of clients’ data, leading to intensive over-fitting and massive domain shifts in the local models and, ultimately, decreasing the aggregated model's performance. Interestingly, while this scenario is a specific manifestation of the well-known non-IID11This refers to the generic situation where local data distributions are not identical and independently distributed.challenge in FL, it has not been distinctly addressed. Our empirical investigation highlights that generic approaches to the non-IID challenge often prove inadequate in mitigating the sparse-and-scarce issue. To bridge this gap, we develop SAFA, a novel FL algorithm that specifically addresses the sparse-and-scarce challenge via a novel continual model iteration procedure. SAFA maximally exposes local models to the inter-client diversity of data with minimal effects of catastrophic forgetting. Our experiments show that SAFA outperforms existing FL solutions, up to 17.86%, compared to the prominent baseline. The code is accessible viahttps://github.com/HungNguyen20/SAFA. Nang Hung Nguyen, Truong Thao Nguyen, Trong Nghia Hoang, Hieu H. Pham 0001, Thanh-Hung Nguyen, Phi-Le Nguyen |
IEEE Trans. Computers | 4 |
| 2024 | D-SarcNet: A Dual-stream Deep Learning Framework for Automatic Analysis of Sarcomere Structures in Fluorescently Labeled hiPSC-CMsabstractHuman-induced pluripotent stem cell-derived cardiomyocytes (hiPSC-CMs) are a powerful tool in advancing cardiovascular research and clinical applications. The maturation of sarcomere organization in hiPSC-CMs is crucial, as it supports the contractile function and structural integrity of these cells. Traditional methods for assessing this maturation like manual annotation and feature extraction are labor-intensive, time-consuming, and unsuitable for high-throughput analysis. To address this, we propose D-SarcNet, a dual-stream deep learning framework that takes fluorescent hiPSC-CM single-cell images as input and outputs the stage of the sarcomere structural organization on a scale from 1.0 to 5.0. The framework also integrates Fast Fourier Transform (FFT), deep learning-generated local patterns, and gradient magnitude to capture detailed structural information at both global and local levels. Experiments on a publicly available dataset from the Allen Institute for Cell Science show that the proposed approach not only achieves a Spearman correlation of 0.868—marking a 3.7% improvement over the previous state-of-the-art—but also significantly enhances other key performance metrics, including MSE, MAE, and R2score. Beyond establishing a new state-of-the-art in sarcomere structure assessment from hiPSC-CM images, our ablation studies highlight the significance of integrating global and local information to enhance deep learning networks’ ability to discern and learn vital visual features of sarcomere structure. Huyen Le, Khiet Dang, Nhung Nguyen, Mai Tran, Hieu H. Pham 0001 |
BIBM | 5 |
| 2024 | FedBlock: A Blockchain Approach to Federated Learning against Backdoor AttacksabstractFederated Learning (FL) is a machine learning method for training with private data locally stored in distributed machines without gathering them into one place for central learning. Despite its promises, FL is prone to critical security risks. First, because FL depends on a central server to aggregate local training models, this is a single point of failure. The server might function maliciously. Second, due to its distributed nature, FL might encounter backdoor attacks by participating clients. They can poison the local model before submitting to the server. Either type of attack, on the server or the client side, would severely degrade learning accuracy. We propose FedBlock, a novel blockchain-based FL framework that addresses both of these security risks. FedBlock is uniquely desirable in that it involves only smart contract programming, thus deployable atop any blockchain network. Our framework is substantiated with a comprehensive evaluation study using real-world datasets. Its robustness against backdoor attacks is competitive with the literature of FL backdoor defense. The latter, however, does not address the server risk as we do. Duong H. Nguyen, Phi L. Nguyen, Truong T. Nguyen, Hieu H. Pham 0001, Duc A. Tran |
IEEE Big Data | 4 |
| 2024 | Improving Time Series Encoding with Noise-Aware Self-Supervised Learning and an Efficient EncoderabstractIn this work, we investigate the time series representation learning problem using self-supervised techniques. Contrastive learning is well-known in this area as it is a powerful method for extracting information from the series and generating task-appropriate representations. Despite its proficiency in capturing time series characteristics, these techniques often overlook a critical factor - the inherent noise in this type of data, a consideration usually emphasized in general time series analysis. Moreover, there is a notable absence of attention to developing efficient yet lightweight encoder architectures, with an undue focus on delivering contrastive losses. Our work address these gaps by proposing an innovative training strategy that promotes consistent representation learning, accounting for the presence of noise-prone signals in natural time series. Furthermore, we propose an encoder architecture that incorporates dilated convolution within the Inception block, resulting in a scalable and robust network with a wide receptive field. Experimental findings underscore the effectiveness of our method, consistently outperforming state-of-the-art approaches across various tasks, including forecasting, classification, and abnormality detection. Notably, our method attains the top rank in over two-thirds of the classification UCR datasets, utilizing only 40% of the parameters compared to the second-best approach. Duy A. Nguyen, Trang H. Tran, Hieu H. Pham 0001, Phi-Le Nguyen, Lam M. Nguyen |
ICDM | 3 |
| 2024 | Training-Free Condition Video Diffusion Models for Single Frame Spatial-Semantic Echocardiogram Synthesis
Nguyen Van Phi, Tri Nhan Luong Ha, Hieu H. Pham 0001, Quoc Long Tran |
MICCAI (6) | 3 |
| 2024 | FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive RegularizationabstractFederated Learning (FL) is a method for training machine learning models using distributed data sources. It ensures privacy by allowing clients to collaboratively learn a shared global model while storing their data locally. However, a significant challenge arises when dealing with missing modalities in clients’ datasets, where certain features or modalities are unavailable or incomplete, leading to heterogeneous data distribution. While previous studies have addressed the issue of complete-modality missing1, they fail to tackle partial-modality missing2on account of severe heterogeneity among clients at an instance level, where the pattern of missing data can vary significantly from one sample to another. To tackle this challenge, this study proposes a novel framework named FedMAC, designed to address multimodality missing under conditions of partial-modality missing in FL. Additionally, to avoid trivial aggregation of multi-modal features, we introduce contrastive-based regularization to impose additional constraints on the latent representation space. The experimental results demonstrate the effectiveness of FedMAC across various client configurations with statistical heterogeneity, outperforming baseline methods by up to 26% in severe missing scenarios, highlighting its potential as a solution for the challenge of partially missing modalities in federated systems.1Complete missing is when one or more modalities are absent in server and clients’ data.2Partial missing is when only parts of one or more modalities are absent in server and clients’ data Manh Duong Nguyen, Trung Thanh Nguyen 0006, Hieu H. Pham 0001, Trong Nghia Hoang, Phi-Le Nguyen |
NCA | 3 |
| 2024 | Backdoor attacks and defenses in federated learning: Survey, challenges and future research directions
Thuy Dung Nguyen, Phi-Le Nguyen, Hieu H. Pham 0001, Khoa D. Doan, Kok-Seng Wong |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | GAMMA: A universal model for calibrating sensory data of multiple low-cost air monitoring devicesabstractDue to global air pollution , there is a growing demand for accurate and large-scale air quality monitoring systems. Consequently, low-cost air monitoring devices have emerged as a potential alternative to expensive conventional ones. However, the low-cost devices’ major drawback is their insufficient level of accuracy. This work investigates the problem of calibrating the sensory data, especially PM2.5 concentration, collected by low-cost sensor-based air quality monitoring devices. Recently, deep learning has emerged as a potential solution for data calibration instead of using traditional methods, whose accuracy is relatively low. Nevertheless, it generally incurs significant costs. Moreover, it is necessary to employ a dedicated calibration model for each device to increase precision, resulting in additional expenditures. To address the issue, this study provides a novel approach named GAMMA, which entails the development of a deep learning-based model capable of accurately calibrating data for multiple devices simultaneously. The proposed method leverages the multitask learning paradigm to solve the challenge of concurrently processing several devices’ data. This involves capturing common features across all devices’ data while also distinguishing the device-specific characteristics. Furthermore, GAMMA also employs the adversarial training approach to augment the accuracy of predictions. This method has been implemented and integrated into an air quality monitoring system in Hanoi, Vietnam. Comprehensive experiments are conducted on real-world data to demonstrate the superiority of GAMMA against the comparison benchmarks in terms of various metrics. Notably, GAMMA reduces MAE from 60.19% to 74.09% compared to the best comparison baseline. The source code is available at https://github.com/anhduy0911/FimiCalibIdea/tree/multi_attention . Anh Duy Nguyen, Thu Hang Phung, Thuy Dung Nguyen, Hieu H. Pham 0001, Kien Nguyen 0002, Phi-Le Nguyen |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | MPCNN: A Novel Matrix Profile Approach for CNN-based Single Lead Sleep Apnea in Classification ProblemabstractSleep apnea (SA) is a significant respiratory condition that poses a major global health challenge. Deep Learning (DL) has emerged as an efficient tool for the classification problem in electrocardiogram (ECG)-based SA diagnoses. Despite these advancements, most common conventional feature extractions derived from ECG signals in DL, such as R-peaks and RR intervals, may fail to capture crucial information encompassed within the complete ECG segments. In this study, we propose an innovative approach to address this diagnostic gap by delving deeper into the comprehensive segments of the ECG signal. The proposed methodology draws inspiration from Matrix Profile algorithms, which generate an Euclidean distance profile from fixed-length signal subsequences. From this, we derived the Min Distance Profile (MinDP), Max Distance Profile (MaxDP), and Mean Distance Profile (MeanDP) based on the minimum, maximum, and mean of the profile distances, respectively. To validate the effectiveness of our approach, we use the modified LeNet-5 architecture as the primary CNN model, along with two existing lightweight models, BAFNet and SE-MSCNN. Our experiment results on the PhysioNet Apnea-ECG dataset (70 overnight recordings), and the UCDDB dataset (25 overnight recordings) revealed that our new feature extraction method achieved per-segment accuracies of up to 92.11% and 81.25%, respectively. Moreover, using the PhysioNet data, we achieved a per-recording accuracy of 100% and yielded the highest correlation of 0.989 compared to state-of-the-art methods. By introducing a new feature extraction method based on distance relationships, we enhanced the performance of certain lightweight models in DL, showing potential for home sleep apnea test (HSAT) and SA detection in IoT devices. The source code for this work is made publicly available in GitHub: https://github.com/vinuni-vishc/MPCNN-Sleep-Apnea. Hieu X. Nguyen, Duong V. Nguyen, Hieu H. Pham 0001, Cuong Do 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | FedDCT: Federated Learning of Large Convolutional Neural Networks on Resource-Constrained Devices Using Divide and Collaborative TrainingabstractIn Federated Learning (FL), the size of local models matters. On the one hand, it is logical to use large-capacity neural networks in pursuit of high performance. On the other hand, deep convolutional neural networks (CNNs) are exceedingly parameter-hungry, which makes memory a significant bottleneck when training large-scale CNNs on hardware-constrained devices such as smartphones or wearables sensors. Current state-of-the-art (SOTA) FL approaches either only test their convergence properties on tiny CNNs with inferior accuracy or assume clients have the adequate processing power to train large models, which remains a formidable obstacle in actual practice. To overcome these issues, we introduce FedDCT, a novel distributed learning paradigm that enables the usage of large, high-performance CNNs on resource-limited edge devices. As opposed to traditional FL approaches, which require each client to train the full-size neural network independently during each training round, the proposed FedDCT allows a cluster of several clients to collaboratively train a large deep learning model by dividing it into an ensemble of several small sub-models and train them on multiple devices in parallel while maintaining privacy. In this collaborative training process, clients from the same cluster can also learn from each other, further improving their ensemble performance. In the aggregation stage, the server takes a weighted average of all the ensemble models trained by all the clusters. FedDCT reduces the memory requirements and allows low-end devices to participate in FL. We empirically conduct extensive experiments on standardized datasets, including CIFAR-10, CIFAR-100, and two real-world medical datasets HAM10000 and VAIPE. Experimental results show that FedDCT outperforms a set of current SOTA FL methods with interesting convergence behaviors. Furthermore, compared to other existing approaches, FedDCT achieves higher accuracy and substantially reduces the number of communication rounds (with 4-8 times fewer memory requirements) to achieve the desired accuracy on the testing dataset without incurring any extra training cost on the server side. Hieu H. Pham 0001, Kok-Seng Wong, Phi-Le Nguyen, Truong Thao Nguyen, Minh N. Do |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2023 | CADIS: Handling Cluster-skewed Non-IID Data in Federated Learning with Clustered Aggregation and Knowledge DIStilled RegularizationabstractFederated learning enables edge devices to train a global model collaboratively without exposing their data. Despite achieving outstanding advantages in computing efficiency and privacy protection, federated learning faces a significant challenge when dealing with non-IID data, i.e., data generated by clients that are typically not independent and identically distributed. In this paper, we tackle a new type of Non-IID data, called cluster-skewed non-IID, discovered in actual data sets. The cluster-skewed non-IID is a phenomenon in which clients can be grouped into clusters with similar data distributions. By performing an in-depth analysis of the behavior of a classification model's penultimate layer, we introduce a metric that quantifies the similarity between two clients' data distributions without violating their privacy. We then propose an aggregation scheme that guarantees equality between clusters. In addition, we offer a novel local training regularization based on the knowledge-distillation technique that reduces the overfitting problem at clients and dramatically boosts the training scheme's performance. We theoretically prove the superiority of the proposed aggregation over the benchmark FedAvg. Extensive experimental results on both standard public datasets and our in-house real-world dataset demonstrate that the proposed approach improves accuracy by up to 16% compared to the FedAvg algorithm. Nang Hung Nguyen, Duc Long Nguyen, Trong Bang Nguyen, Thanh-Hung Nguyen, Hieu H. Pham 0001, Truong Thao Nguyen, Phi-Le Nguyen |
CCGrid | 5 |
| 2023 | FedGrad: Mitigating Backdoor Attacks in Federated Learning Through Local Ultimate Gradients InspectionabstractFederated learning (FL) enables multiple clients to train a model without compromising sensitive data. The decentralized nature of FL makes it susceptible to adversarial attacks, especially backdoor insertion during training. Recently, the edge-case backdoor attack employing the tail of the data distribution has been proposed as a powerful one, raising questions about the shortfall in current defenses' robustness guarantees. Specifically, most existing defenses cannot eliminate edge-case backdoor attacks or suffer from a trade-off between backdoor-defending effectiveness and overall performance on the primary task. To tackle this challenge, we propose FedGrad, a novel backdoor-resistant defense for FL that is resistant to cutting-edge backdoor attacks, including the edge-case attack, and performs effectively under heterogeneous client data and a large number of compromised clients. FedGrad is designed as a two-layer filtering mechanism that thoroughly analyzes the ultimate layer's gradient to identify suspicious local updates and remove them from the aggregation process. We evaluate FedGrad under different attack scenarios and show that it significantly outperforms state-of-the-art defense mechanisms. Notably, FedGrad can almost 100% correctly detect the malicious participants, thus providing a significant reduction in the backdoor effect (e.g., backdoor accuracy is less than 8%) while not reducing main accuracy on the primary task. Thuy Dung Nguyen, Anh Duy Nguyen, Thanh-Hung Nguyen, Kok-Seng Wong, Hieu H. Pham 0001, Truong Thao Nguyen, Phi-Le Nguyen |
IJCNN | 5 |
| 2022 | Multi-stream Fusion for Class Incremental Learning in Pill Image Classification
Trong-Tung Nguyen, Hieu H. Pham 0001, Phi-Le Nguyen, Thanh-Hung Nguyen, Minh Do |
ACCV (2) | 2 |
| 2022 | FedDRL: Deep Reinforcement Learning-based Adaptive Aggregation for Non-IID Data in Federated LearningabstractThe uneven distribution of local data across different edge devices (clients) results in slow model training and accuracy reduction in federated learning. Naive federated learning (FL) strategy and most alternative solutions attempted to achieve more fairness by weighted aggregating deep learning models across clients. This work introduces a novel non-IID type encountered in real-world datasets, namely cluster-skew, in which groups of clients have local data with similar distributions, causing the global model to converge to an over-fitted solution. To deal with non-IID data, particularly the cluster-skewed data, we propose FedDRL, a novel FL model that employs deep reinforcement learning to adaptively determine each client’s impact factor (which will be used as the weights in the aggregation process). Extensive experiments on a suite of federated datasets confirm that the proposed FedDRL improves favorably against FedAvg and FedProx methods, e.g., up to 4.05% and 2.17% on average for the CIFAR-100 dataset, respectively. Nang Hung Nguyen, Phi-Le Nguyen, Thuy Dung Nguyen, Trung Thanh Nguyen 0006, Duc Long Nguyen, Thanh-Hung Nguyen, Hieu H. Pham 0001, Truong Thao Nguyen |
ICPP | 7 |
| 2022 | A Novel Approach for Pill-Prescription Matching with GNN Assistance and Contrastive Learning
Trung Thanh Nguyen 0006, Hoang Dang Nguyen, Thanh-Hung Nguyen, Hieu H. Pham 0001, Ichiro Ide, Phi-Le Nguyen |
PRICAI (1) | 4 |
| 2021 | VinDr-SpineXR: A Deep Learning Framework for Spinal Lesions Detection and Classification from Radiographs
Hieu T. Nguyen 0003, Hieu H. Pham 0001, Nghia T. Nguyen, Ha Q. Nguyen 0001, Thang Q. Huynh, Minh Dao, Van H. Vu |
MICCAI (5) | 2 |
| 2021 | Interpreting chest X-rays via CNNs that exploit hierarchical disease dependencies and uncertainty labels
Hieu H. Pham 0001, Tung T. Le, Dat Q. Tran, Dat T. Ngo 0001, Ha Q. Nguyen 0001 |
Neurocomputing | 1 |
| 2019 | Learning to recognise 3D human action from a new skeleton-based representation using deep convolutional neural networksabstractRecognising human actions in untrimmed videos is an important challenging task. An effective three‐dimensional (3D) motion representation and a powerful learning model are two key factors influencing recognition performance. In this study, the authors introduce a new skeleton‐based representation for 3D action recognition in videos. The key idea of the proposed representation is to transform 3D joint coordinates of the human body carried in skeleton sequences into RGB images via a colour encoding process. By normalising the 3D joint coordinates and dividing each skeleton frame into five parts, where the joints are concatenated according to the order of their physical connections, the colour‐coded representation is able to represent spatio‐temporal evolutions of complex 3D motions, independently of the length of each sequence. They then design and train different deep convolutional neural networks based on the residual network architecture on the obtained image‐based representations to learn 3D motion features and classify them into classes. Their proposed method is evaluated on two widely used action recognition benchmarks: MSR Action3D and NTU‐RGB+D, a very large‐scale dataset for 3D human action recognition. The experimental results demonstrate that the proposed method outperforms previous state‐of‐the‐art approaches while requiring less computation for training and prediction. Hieu H. Pham 0001, Louahdi Khoudour, Alain Crouzil, Pablo Zegers, Sergio A. Velastin |
IET Comput. Vis. | 1 |
| 2018 | Skeletal Movement to Color Map: A Novel Representation for 3D Action Recognition with Inception Residual NetworksabstractWe propose a novel skeleton-based representation for 3D action recognition in videos using Deep Convolutional Neural Networks (D-CNNs). Two key issues have been addressed: First, how to construct a robust representation that easily captures the spatial-temporal evolutions of motions from skeleton sequences. Second, how to design D-CNNs capable of learning discriminative features from the new representation in a effective manner. To address these tasks, a skeleton-based representation, namely, SPMF (Skeleton Pose-Motion Feature) is proposed. The SPMFs are built from two of the most important properties of a human action: postures and their motions. Therefore, they are able to effectively represent complex actions. For learning and recognition tasks, we design and optimize new D-CNNs based on the idea of Inception Residual networks to predict actions from SPMFs. Our method is evaluated on two challenging datasets including MSR Action3D and NTU-RGB+D. Experimental results indicated that the proposed method surpasses state-of-the-art methods whilst requiring less computation. Hieu H. Pham 0001, Louahdi Khoudour, Alain Crouzil, Pablo Zegers, Sergio A. Velastin |
ICIP | 1 |
| 2018 | Exploiting deep residual networks for human action recognition from skeletal data
Hieu H. Pham 0001, Louahdi Khoudour, Alain Crouzil, Pablo Zegers, Sergio A. Velastin |
Comput. Vis. Image Underst. | 1 |