VLDB 2026 Research / reviewers in the wild / expert
Truong Thao Nguyen
dblp:233/1462
· DBLP profile ↗
34ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0003-3641-374XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Computer networks · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLaTEC: An efficient federated learning scheme across the Thing-Edge-Cloud environment
Van An Le, Jason H. Haga, Yusuke Tanimura, Truong Thao Nguyen |
Future Gener. Comput. Syst. | 4 |
| 2025 | ROOT: Rethinking Offline Optimization as Distributional Translation via Probabilistic BridgeabstractThis paper studies the black-box optimization task which aims to find the maxima of a black-box function using a static set of its observed input-output pairs. This is often achieved via learning and optimizing a surrogate function with that offline data. Alternatively, it can also be framed as an inverse modeling task that maps a desired performance to potential input candidates that achieve it. Both approaches are constrained by the limited amount of offline data. To mitigate this limitation, we introduce a new perspective that casts offline optimization as a distributional translation task. This is formulated as learning a probabilistic bridge transforming an implicit distribution of low-value inputs (i.e., offline data) into another distribution of high-value inputs (i.e., solution candidates). Such probabilistic bridge can be learned using low- and high-value inputs sampled from synthetic functions that resemble the target function. These synthetic functions are constructed as the mean posterior of multiple Gaussian processes fitted with different parameterizations on the offline data, alleviating the data bottleneck. The proposed approach is evaluated on an extensive benchmark comprising most recent methods, demonstrating significant improvement and establishing a new state-of-the-art performance. Our code is publicly available at https://github.com/cuong-dm/ROOT. Cuong Dao, The Hung Tran, Phi-Le Nguyen, Truong Thao Nguyen, Nghia Hoang |
NeurIPS | 4 |
| 2025 | Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report GenerationabstractVision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains, applying these models to medical imaging remains challenging due to the limited availability of diverse imaging modalities and multilingual clinical data. Most existing medical VLMs are trained on a subset of imaging modalities and focus primarily on high-resource languages, thus limiting their generalizability and clinical utility. To address these limitations, we introduce a novel Vietnamese-language multimodal medical dataset consisting of 2,757 whole-body PET/CT volumes from independent patients and their corresponding full-length clinical reports. This dataset is designed to fill two pressing gaps in medical AI development: (1) the lack of PET/CT imaging data in existing VLMs training corpora, which hinders the development of models capable of handling functional imaging tasks; and (2) the underrepresentation of low-resource languages, particularly the Vietnamese language, in medical vision-language research. To the best of our knowledge, this is the first dataset to provide comprehensive PET/CT-report pairs in Vietnamese. We further introduce a training framework to enhance VLMs' learning, including data augmentation and expert-validated test sets. We conduct comprehensive experiments benchmarking state-of-the-art VLMs on downstream tasks, including medical report generation and visual question answering. The experimental results show that incorporating our dataset significantly improves the performance of existing VLMs. However, despite these advancements, the models still underperform on clinically critical criteria, particularly the diagnosis of lung cancer, indicating substantial room for future improvement. We believe this dataset and benchmark will serve as a pivotal step in advancing the development of more robust VLMs for medical imaging, particularly in low-resource languages, and improving their clinical relevance in Vietnamese healthcare. Huu Tien Nguyen, Dac Thai Nguyen, The Minh Duc Nguyen, Trung Thanh Nguyen 0006, Truong Thao Nguyen, Hieu H. Pham 0001, Johan Barthelemy, Minh Quan Tran, Nguyen Quoc Viet Hung, Thanh Tam Nguyen, Hong Son Mai, Quynh Anh Chau, Thanh Hong Nguyen, Phi-Le Nguyen |
NeurIPS | 5 |
| 2025 | CT to PET Translation: A Large-Scale Dataset and Domain-Knowledge-Guided Diffusion ApproachabstractPositron Emission Tomography (PET) and Computed Tomography (CT) are essential for diagnosing, staging, and monitoring various diseases, particularly cancer. Despite their importance, the use of PET/CT systems is limited by the necessity for radioactive materials, the scarcity of PET scanners, and the high cost associated with PET imaging. In contrast, CT scanners are more widely available and significantly less expensive. In response to these challenges, our study addresses the issue of generating PET images from CT images, aiming to reduce both the medical examination cost and the associated health risks for patients. Our contributions are twofold: First, we introduce a conditional diffusion model named CPDM, which, to our knowledge, is one of the initial attempts to employ a diffusion model for translating from CT to PET images. Second, we provide the largest CT-PET dataset to date, comprising 2,028,628 paired CT-PET images, which facilitates the training and evaluation of CT-to-PET translation models. For the CPDM model, we incorporate domain knowledge to develop two conditional maps: the Attention map and the Attenuation map. The former helps the diffusion process focus on areas of interest, while the latter improves PET data correction and ensures accurate diagnostic information. Experimental evaluations across various benchmarks demonstrate that CPDM surpasses existing methods in generating high-quality PET images in terms of multiple metrics. The source code and data samples are available at https://github.com/thanhhff/CPDM. Dac Thai Nguyen, Trung Thanh Nguyen 0006, Huu Tien Nguyen, Thanh Trung Nguyen, Hieu H. Pham 0001, Thanh-Hung Nguyen, Truong Thao Nguyen, Phi-Le Nguyen |
WACV | 7 |
| 2025 | Noisy data-based attack: A new type of untargeted attack in Federated Learning and its countermeasures
Manh Cuong Dao, Phi-Le Nguyen, Hieu H. Pham 0001, Thanh-Hung Nguyen, Peng Chen 0035, Mohamed Wahib, Truong Thao Nguyen |
Future Gener. Comput. Syst. | 7 |
| 2025 | SAFA: Handling Sparse and Scarce Data in Federated Learning With Accumulative LearningabstractFederated Learning (FL) has emerged as an effective paradigm allowing multiple parties to collaboratively train a global model while protecting their private data. However, it is observed that the performance of FL approaches tends to degrade significantly when data are sparsely distributed across clients with small datasets. This is referred to as the sparse-and-scarce challenge, where data held by each client is both sparse (does not contain examples to all classes) and scarce (small dataset). Sparse-and-scarce data diminishes the generalizability of clients’ data, leading to intensive over-fitting and massive domain shifts in the local models and, ultimately, decreasing the aggregated model's performance. Interestingly, while this scenario is a specific manifestation of the well-known non-IID11This refers to the generic situation where local data distributions are not identical and independently distributed.challenge in FL, it has not been distinctly addressed. Our empirical investigation highlights that generic approaches to the non-IID challenge often prove inadequate in mitigating the sparse-and-scarce issue. To bridge this gap, we develop SAFA, a novel FL algorithm that specifically addresses the sparse-and-scarce challenge via a novel continual model iteration procedure. SAFA maximally exposes local models to the inter-client diversity of data with minimal effects of catastrophic forgetting. Our experiments show that SAFA outperforms existing FL solutions, up to 17.86%, compared to the prominent baseline. The code is accessible viahttps://github.com/HungNguyen20/SAFA. Nang Hung Nguyen, Truong Thao Nguyen, Trong Nghia Hoang, Hieu H. Pham 0001, Thanh-Hung Nguyen, Phi-Le Nguyen |
IEEE Trans. Computers | 2 |
| 2024 | A Bandwidth-Optimal All-to-All Communication in Two-Dimensional Fully Connected NetworkabstractHigh-radix direct interconnection networks enable many simultaneous, non-blocking point-to-point communications at every compute node of high-performance computing systems. This feature provides a novel opportunity to improve the performance of collective communications. In this study, we focus on two-dimensional fully connected network topologies (2dfc) which are basic blocks of cutting-edge high-radix network topologies such as Dragonfly and HyperX. We aim to design a bandwidth-optimal algorithm for the All-to-All collective communication in the 2dfc network. By using parallel point-to-point communications, the proposed algorithm enables parallel non-blocking intra- and inter-group communications with striping data placement. Simulation results using SimGrid illustrated that the proposed algorithm outperforms the conventional algorithms significantly. Compared with topology-aware algorithms, the proposed approach reduces 1.66× and 1.58× communication time of an existing and a topology-aware collective algorithm in a 16 × 16 2dfc network, respectively. Kien Trung Pham, Truong Thao Nguyen, Michihiro Koibuchi |
CCGrid | 2 |
| 2024 | SFETEC: Split-FEderated Learning Scheme Optimized for Thing-Edge-Cloud EnvironmentabstractThis paper introduces SFETEC, an innovative federated learning framework addressing the limitations of traditional methods like FedAvg. SFETEC splits training into base models and core models, reducing communication overhead, mitigating non-IID data issues, and enhancing training speed. Base models are trained at client devices, while core models are trained at edge servers and aggregated at the cloud. Preliminary results show that SFETEC significantly reduces communication overhead and training duration compared to the state-of-the-art baselines while enhancing privacy. Van An Le, Jason H. Haga, Yusuke Tanimura, Truong Thao Nguyen |
e-Science | 4 |
| 2024 | Boosting Offline Optimizers with Surrogate SensitivityabstractOffline optimization is an important task in numerous material engineering domains where online experimentation to collect data is too expensive and needs to be replaced by an in silico maximization of a surrogate of the black-box function. Although such a surrogate can be learned from offline data, its prediction might not be reliable outside the offline data regime, which happens when the surrogate has narrow prediction margin and is (therefore) sensitive to small perturbations of its parameterization. This raises the following questions: (1) how to regulate the sensitivity of a surrogate model; and (2) whether conditioning an offline optimizer with such less sensitive surrogate will lead to better optimization performance. To address these questions, we develop an optimizable sensitivity measurement for the surrogate model, which then inspires a sensitivity-informed regularizer that is applicable to a wide range of offline optimizers. This development is both orthogonal and synergistic to prior research on offline optimization, which is demonstrated in our extensive experiment benchmark. Manh Cuong Dao, Phi-Le Nguyen, Truong Thao Nguyen, Trong Nghia Hoang |
ICML | 3 |
| 2024 | Enhancing the Generalization of Personalized Federated Learning with Multi-head Model and Ensemble VotingabstractFederated Learning has emerged as a transformative paradigm in the realm of collaborative machine learning, enabling the training of global models across decentralized devices without the need for centralizing data. While Federated Learning has shown remarkable promise, a critical limitation lies in its ability to personalize models to individual clients. Current approaches predominantly emphasize improving the accuracy of trained clients, inadvertently sidelining the significance of accommodating unseen clients. Furthermore, most of the existing personalized federated learning approaches require new clients to provide labeled data and undergo extensive retraining, posing a substantial barrier and hindering the broader adoption and engagement of potential users within these systems.In this paper, we introduce a novel and comprehensive solution to address these challenges: a generalized method for Personalized Federated Learning. Our approach transcends the limitations of conventional Federated Learning techniques by not only optimizing the accuracy of trained clients but also ensuring exceptional performance among unseen clients, even in diverse settings. Throughout extensive experiments, our method demonstrates significant improvements concerning the performance of seen and unseen clients, respectively, while eliminating the need for labeled data and model re-training among unseen clients. Van An Le, Nam Duong Tran, Phuong Nam Nguyen, Thanh-Hung Nguyen, Phi-Le Nguyen, Truong Thao Nguyen, Yusheng Ji |
IPDPS | 6 |
| 2024 | FedCert: Federated Accuracy Certification
Minh Hieu Nguyen 0003, Huu Tien Nguyen, Trung Thanh Nguyen 0006, Manh Duong Nguyen, Trong Nghia Hoang, Truong Thao Nguyen, Phi-Le Nguyen |
NCA | 6 |
| 2024 | Incorporating Surrogate Gradient Norm to Improve Offline Optimization TechniquesabstractOffline optimization has recently emerged as an increasingly popular approach to mitigate the prohibitively expensive cost of online experimentation. The key idea is to learn a surrogate of the black-box function that underlines the target experiment using a static (offline) dataset of its previous input-output queries. Such an approach is, however, fraught with an out-of-distribution issue where the learned surrogate becomes inaccurate outside the offline data regimes. To mitigate this, existing offline optimizers have proposed numerous conditioning techniques to prevent the learned surrogate from being too erratic. Nonetheless, such conditioning strategies are often specific to particular surrogate or search models, which might not generalize to a different model choice. This motivates us to develop a model-agnostic approach instead, which incorporates a notion of model sharpness into the training loss of the surrogate as a regularizer. Our approach is supported by a new theoretical analysis demonstrating that reducing surrogate sharpness on the offline dataset provably reduces its generalized sharpness on unseen data. Our analysis extends existing theories from bounding generalized prediction loss (on unseen data) with loss sharpness to bounding the worst-case generalized surrogate sharpness with its empirical estimate on training data, providing a new perspective on sharpness regularization. Our extensive experimentation on a diverse range of optimization tasks also shows that reducing surrogate sharpness often leads to significant improvement, marking (up to) a noticeable 9.6% performance boost. Our code is publicly available at https://github.com/cuong-dm/IGNITE. Cuong Dao, Phi-Le Nguyen, Truong Thao Nguyen, Nghia Hoang |
NeurIPS | 3 |
| 2024 | Combating Quality Distortion in Federated Learning with Collaborative Data Selection
Duc Long Nguyen, Phi-Le Nguyen, Truong Thao Nguyen |
PAKDD (3) | 3 |
| 2024 | A data-driven approach for high accurate spatiotemporal precipitation estimation
Pham Minh Khiem, Phi-Le Nguyen, Viet Hung Vu, Truong Thao Nguyen, Hoa Vo-Van, Thanh Ngo-Duc |
Neural Comput. Appl. | 4 |
| 2024 | FedDCT: Federated Learning of Large Convolutional Neural Networks on Resource-Constrained Devices Using Divide and Collaborative TrainingabstractIn Federated Learning (FL), the size of local models matters. On the one hand, it is logical to use large-capacity neural networks in pursuit of high performance. On the other hand, deep convolutional neural networks (CNNs) are exceedingly parameter-hungry, which makes memory a significant bottleneck when training large-scale CNNs on hardware-constrained devices such as smartphones or wearables sensors. Current state-of-the-art (SOTA) FL approaches either only test their convergence properties on tiny CNNs with inferior accuracy or assume clients have the adequate processing power to train large models, which remains a formidable obstacle in actual practice. To overcome these issues, we introduce FedDCT, a novel distributed learning paradigm that enables the usage of large, high-performance CNNs on resource-limited edge devices. As opposed to traditional FL approaches, which require each client to train the full-size neural network independently during each training round, the proposed FedDCT allows a cluster of several clients to collaboratively train a large deep learning model by dividing it into an ensemble of several small sub-models and train them on multiple devices in parallel while maintaining privacy. In this collaborative training process, clients from the same cluster can also learn from each other, further improving their ensemble performance. In the aggregation stage, the server takes a weighted average of all the ensemble models trained by all the clusters. FedDCT reduces the memory requirements and allows low-end devices to participate in FL. We empirically conduct extensive experiments on standardized datasets, including CIFAR-10, CIFAR-100, and two real-world medical datasets HAM10000 and VAIPE. Experimental results show that FedDCT outperforms a set of current SOTA FL methods with interesting convergence behaviors. Furthermore, compared to other existing approaches, FedDCT achieves higher accuracy and substantially reduces the number of communication rounds (with 4-8 times fewer memory requirements) to achieve the desired accuracy on the testing dataset without incurring any extra training cost on the server side. Hieu H. Pham 0001, Kok-Seng Wong, Phi-Le Nguyen, Truong Thao Nguyen, Minh N. Do |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2023 | CADIS: Handling Cluster-skewed Non-IID Data in Federated Learning with Clustered Aggregation and Knowledge DIStilled RegularizationabstractFederated learning enables edge devices to train a global model collaboratively without exposing their data. Despite achieving outstanding advantages in computing efficiency and privacy protection, federated learning faces a significant challenge when dealing with non-IID data, i.e., data generated by clients that are typically not independent and identically distributed. In this paper, we tackle a new type of Non-IID data, called cluster-skewed non-IID, discovered in actual data sets. The cluster-skewed non-IID is a phenomenon in which clients can be grouped into clusters with similar data distributions. By performing an in-depth analysis of the behavior of a classification model's penultimate layer, we introduce a metric that quantifies the similarity between two clients' data distributions without violating their privacy. We then propose an aggregation scheme that guarantees equality between clusters. In addition, we offer a novel local training regularization based on the knowledge-distillation technique that reduces the overfitting problem at clients and dramatically boosts the training scheme's performance. We theoretically prove the superiority of the proposed aggregation over the benchmark FedAvg. Extensive experimental results on both standard public datasets and our in-house real-world dataset demonstrate that the proposed approach improves accuracy by up to 16% compared to the FedAvg algorithm. Nang Hung Nguyen, Duc Long Nguyen, Trong Bang Nguyen, Thanh-Hung Nguyen, Hieu H. Pham 0001, Truong Thao Nguyen, Phi-Le Nguyen |
CCGrid | 6 |
| 2023 | FedGrad: Mitigating Backdoor Attacks in Federated Learning Through Local Ultimate Gradients InspectionabstractFederated learning (FL) enables multiple clients to train a model without compromising sensitive data. The decentralized nature of FL makes it susceptible to adversarial attacks, especially backdoor insertion during training. Recently, the edge-case backdoor attack employing the tail of the data distribution has been proposed as a powerful one, raising questions about the shortfall in current defenses' robustness guarantees. Specifically, most existing defenses cannot eliminate edge-case backdoor attacks or suffer from a trade-off between backdoor-defending effectiveness and overall performance on the primary task. To tackle this challenge, we propose FedGrad, a novel backdoor-resistant defense for FL that is resistant to cutting-edge backdoor attacks, including the edge-case attack, and performs effectively under heterogeneous client data and a large number of compromised clients. FedGrad is designed as a two-layer filtering mechanism that thoroughly analyzes the ultimate layer's gradient to identify suspicious local updates and remove them from the aggregation process. We evaluate FedGrad under different attack scenarios and show that it significantly outperforms state-of-the-art defense mechanisms. Notably, FedGrad can almost 100% correctly detect the malicious participants, thus providing a significant reduction in the backdoor effect (e.g., backdoor accuracy is less than 8%) while not reducing main accuracy on the primary task. Thuy Dung Nguyen, Anh Duy Nguyen, Thanh-Hung Nguyen, Kok-Seng Wong, Hieu H. Pham 0001, Truong Thao Nguyen, Phi-Le Nguyen |
IJCNN | 6 |
| 2023 | KAKURENBO: Adaptively Hiding Samples in Deep Neural Network TrainingabstractThis paper proposes a method for hiding the least-important samples during the training of deep neural networks to increase efficiency, i.e., to reduce the cost of training. Using information about the loss and prediction confidence during training, we adaptively find samples to exclude in a given epoch based on their contribution to the overall learning process, without significantly degrading accuracy. We explore the converge properties when accounting for the reduction in the number of SGD updates. Empirical results on various large-scale datasets and models used directly in image classification and segmentation show that while the with-replacement importance sampling algorithm performs poorly on large datasets, our method can reduce total training time by up to 22\% impacting accuracy only by 0.4\% compared to the baseline. Truong Thao Nguyen, Balazs Gerofi, Edgar Josafat Martinez-Noriega, François Trahay, Mohamed Wahib |
NeurIPS | 1 |
| 2023 | Effective switchless inter-FPGA memory networks
Truong Thao Nguyen, Kien Trung Pham, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi |
J. Parallel Distributed Comput. | 1 |
| 2023 | Simeuro: A Hybrid CPU-GPU Parallel Simulator for Neuromorphic Computing ChipsabstractWith the success of deep learning, there have been numerous efforts to build hardware for it. One approach that is gaining momentum is neuromorphic computing with spiking neural networks (SNNs), which are multiplication-free and open the possibility of using analog computing via novel technologies. However, to design effective and efficient hardware for such architectures, a fast and accurate software simulator is key. This article presents Simeuro, a fast and scalable system-level simulator for SNN models used in neuromorphic accelerators. The simulator uses spike-level details and configurable architectural constraints that are independent of the underlying hardware implementation. Simeuro supports a wide range of features including analog computing, novel memory (currently, RRAM is supported), and a full network-on-chip. The simulator can provide detailed simulation results such as routing statistics, energy consumption, delay, and accuracy of arbitrarily defined SNN architectures. Our simulator leverages a CPU-GPU hybrid environment to expedite the simulation by scaling out to multi-nodes equipped with multi-GPUs. We are able to conduct core simulations for a system-scale SNN chip of 20,000 neuromorphic cores on up to 512 A100 GPUs in a few minutes. Huaipeng Zhang, Nhut-Minh Ho, Dogukan Yigit Polat, Peng Chen 0035, Mohamed Wahib, Truong Thao Nguyen, Jintao Meng 0001, Rick Siow Mong Goh, Satoshi Matsuoka, Tao Luo 0014, Weng-Fai Wong |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | FedDRL: Deep Reinforcement Learning-based Adaptive Aggregation for Non-IID Data in Federated LearningabstractThe uneven distribution of local data across different edge devices (clients) results in slow model training and accuracy reduction in federated learning. Naive federated learning (FL) strategy and most alternative solutions attempted to achieve more fairness by weighted aggregating deep learning models across clients. This work introduces a novel non-IID type encountered in real-world datasets, namely cluster-skew, in which groups of clients have local data with similar distributions, causing the global model to converge to an over-fitted solution. To deal with non-IID data, particularly the cluster-skewed data, we propose FedDRL, a novel FL model that employs deep reinforcement learning to adaptively determine each client’s impact factor (which will be used as the weights in the aggregation process). Extensive experiments on a suite of federated datasets confirm that the proposed FedDRL improves favorably against FedAvg and FedProx methods, e.g., up to 4.05% and 2.17% on average for the CIFAR-100 dataset, respectively. Nang Hung Nguyen, Phi-Le Nguyen, Thuy Dung Nguyen, Trung Thanh Nguyen 0006, Duc Long Nguyen, Thanh-Hung Nguyen, Hieu H. Pham 0001, Truong Thao Nguyen |
ICPP | 8 |
| 2022 | Why Globally Re-shuffle? Revisiting Data Shuffling in Large Scale Deep LearningabstractStochastic gradient descent (SGD) is the most prevalent algorithm for training Deep Neural Networks (DNN). SGD iterates the input data set in each training epoch processing data samples in a random access fashion. Because this puts enormous pressure on the I/O subsystem, the most common approach to distributed SGD in HPC environments is to replicate the entire dataset to node local SSDs. However, due to rapidly growing data set sizes this approach has become increasingly infeasible. Surprisingly, the questions of why and to what extent random access is required have not received a lot of attention in the literature from an empirical standpoint. In this paper, we revisit data shuffling in DL workloads to investigate the viability of partitioning the dataset among workers and performing only a partial distributed exchange of samples in each training epoch. Through extensive experiments on up to 2,048 GPUs of ABCI and 4,096 compute nodes of Fugaku, we demonstrate that in practice validation accuracy of global shuffling can be maintained when carefully tuning the partial distributed exchange. We provide a solution implemented in PyTorch that enables users to control the proposed data exchange scheme. Truong Thao Nguyen, François Trahay, Jens Domke, Aleksandr Drozd, Emil Vatai, Jianwei Liao 0001, Mohamed Wahib, Balazs Gerofi |
IPDPS | 1 |
| 2022 | Scalable Low-Latency Inter-FPGA NetworksabstractA cutting-edge FPGA card can be equipped with many high-bandwidth I/Os by means of high-density optical integration, e.g., onboard Si-photonics transceivers, to provide high network bandwidth for memory-to-memory inter-FPGA communication. This study presents its scalable switchless net-work architecture by exploiting an indirect path, consisting of two one-hop paths, for enabling a diameter-2 network topology. It then takes a Kautz network topology with a diameter of two for connecting d(d + 1) FPGAs with a degree of$d$, which is close to the theoretical upper bound. The Kautz network topologies have bi-directional links and uni-directional links which form triangles. Uni-directional links introduce difficulty in avoiding channel buffer overflow because the existing link-level flow control assumes a bi-directional link. This study presents an indirect flow control along a uni-directional triangle embedded in the Kautz network topology. It then develops a combination of unicasts that forms multi-port collective communications to mitigate the influence of the startup latency on the execution time. Since a high-degree FPGA card introduces difficulty in storing many I/O ports at the panel of a 1- U compute server, we propose using WDM (Wavelength Division Multiplexing) as an alternative and present its efficient mapping onto arrayed waveguide grating (AWG). The required number of wavelengths becomes d on d+ 1 AWG equipments. Based on our experimental results with OPTWEB of custom Stratix10 FPGA cards, SimGrid simulation results show that our collective communication is 7 × faster than that of Dragonfly with 272 FPGAs. Kien Trung Pham, Truong Thao Nguyen, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi |
IPDPS | 2 |
| 2022 | Deep Reinforcement Learning-based Charging Algorithm for Target Coverage and Connectivity in WRSNsabstractTarget coverage and connectivity are two of the most crucial issues in handling wireless sensor networks. However, maintaining these two factors is challenging due to the energy constraint of sensors. To this end, wireless charging has emerged as a promising solution to prolong the sensor's lifetime. In a wireless charging sensor network, a mobile charger moves around the network, stops at several charging locations and charges the sensor via electromagnetic waves. In this study, we investigate the problem of optimizing the charging location and charging time of the mobile charger to ensure the target coverage and connectivity of the network. Our main idea is to leverage the Deep Reinforcement Learning approach. Specifically, the mobile charger will act as an agent, which receives a state including the energy information of the sensors. The mobile charger then decides the following charging location and charging time using the state information and the knowledge learned in the past. Experimental results have shown that our algorithm can extend the network lifetime (i.e., the time until the network coverage and connectivity are not guaranteed) up to 245.9 times compared to the existing algorithms. Hung Cuong Nguyen, Manh Cuong Dao, Thanh Trung Nguyen, Ngoc Khanh Doan, Thanh-Hung Nguyen, Truong Thao Nguyen, Phi-Le Nguyen |
PIMRC | 6 |
| 2022 | Spatial-temporal Coverage Maximization in Vehicle-based Mobile Crowdsensing for Air Quality MonitoringabstractIn this paper, we address vehicle-based mobile crowdsensing for air quality monitoring applications. We tackle a novel issue that asks to determine monitoring frequencies for maximizing spatial-temporal coverage while reducing the monitoring costs and balancing load across the vehicles. We begin by theoretically formulating the problem and proposing an objective function that considers the three goals. We then leverage the evolutionary approach to develop an algorithm for determining the optimal monitoring frequency. We conduct comprehensive experiments to evaluate the performance of the proposed approach and compare it to the other methods. The results indicate that our approach can enhance the objective function by a factor of 1.33 to 4 compared to the others. Tuan Anh Nguyen Dinh, Anh Duy Nguyen, Truong Thao Nguyen, Thanh-Hung Nguyen, Phi-Le Nguyen |
WCNC | 3 |
| 2022 | Deep Reinforcement Learning-based Offloading for Latency Minimization in 3-tier V2X NetworksabstractMulti-access edge computing (MEC) is seen as an effective technique for decreasing service latency in a V2X network by offloading computational activities. With MEC, a three-tier offloading architecture can be developed, where a vehicle can offload computational tasks to a cloud by communicating with a base station (gNB) and Road Side Units (RSUs). In this paper, we focus on three-tier V2X networks which rely on three offloading paths: Vehicle-to-Infrastructure (i.e., vehicle to RSU), Vehicle-to-Cloud (i.e., vehicle to gNB), and Infrastructure-to-Cloud (i.e., RSU to gNB). We propose an offloading strategy based on deep reinforcement learning with the goal of reducing the average latency of tasks. To be more specific, we leverage the Deep Q Network to estimate the goodness of action-state value to determine the offloading decision. We also propose a novel exploration scheme and a new model training strategy. The experimental findings indicate that our proposed offloading method outperforms the state-of-the-art, particularly in critical circumstances characterized by a high rate of vehicle arrival or packet generation. Hieu Dinh, Nang Hung Nguyen, Trung Thanh Nguyen 0006, Thanh-Hung Nguyen, Truong Thao Nguyen, Phi-Le Nguyen |
WCNC | 5 |
| 2022 | Fuzzy Q-Learning-Based Opportunistic Communication for MEC-Enhanced Vehicular CrowdsensingabstractThis study focuses on MEC-enhanced, vehicle-based crowdsensing systems that rely on devices installed on automobiles. We investigate an opportunistic communication paradigm in which devices can transmit measured data directly to a crowdsensing server over a 4G communication channel or to nearby devices or so-called Road Side Units positioned along the road via Wi-Fi. We tackle a new problem that is how to reduce the cost of 4G while preserving the latency. We propose an offloading strategy that combines a reinforcement learning technique known as Q-learning with Fuzzy logic to accomplish the purpose. Q-learning assists devices in learning to decide the communication channel. Meanwhile, Fuzzy logic is used to optimize the reward function in Q-learning. The experiment results show that our offloading method significantly cuts down around 30-40% of the 4G communication cost while keeping the latency of 99% packets below the required threshold. Trung Thanh Nguyen 0006, Truong Thao Nguyen, Thanh-Hung Nguyen, Phi-Le Nguyen |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2021 | An Allreduce Algorithm and Network Co-design for Large-Scale Training of Distributed Deep LearningabstractDistributed training of Deep Neural Networks (DNNs) on High-Performance Computing (HPC) systems is becoming increasingly common. HPC systems dedicated entirely or mainly to Deep Learning (DL) workloads are becoming a reality. The collective communication overhead for calculating the average of weight gradients, e.g., an Allreduce operations, is one of the main factors limiting the scaling of data parallelism training. Several active efforts across different layers of the training stack including the training algorithms, parallelism strategy, communication algorithms, and system design have been proposed to cope with this communication challenge when scaling distributed training of DNNs. However, even with those methods, communication still becomes a bottleneck with the steady increase in model sizes, e.g.,100-10,000s MB, and the number of compute nodes, e.g., 1,000-10,000s of GPUs. In this work, we investigate the benefits of co-design of Allreduce algorithms and the network system. We propose to replace the Fat-tree network topology with a variant of Distributed Loop Network topology that guarantees a fixed routing paths length between any pairs of computing nodes for the communication pattern of halving-doubling Allreduce algorithm. We also propose a technique to eliminate/mitigate the network contention. Truong Thao Nguyen, Mohamed Wahib |
CCGRID | 1 |
| 2021 | An Oracle for Guiding Large-Scale Model/Hybrid Parallel Training of Convolutional Neural NetworksabstractDeep Neural Network (DNN) frameworks use distributed training to enable faster time to convergence and alleviate memory capacity limitations when training large models and/or using high dimension inputs. With the steady increase in datasets and model sizes, model/hybrid parallelism is deemed to have an important role in the future of distributed training of DNNs. We analyze the compute, communication, and memory requirements of Convolutional Neural Networks (CNNs) to understand the trade-offs between different parallelism approaches on performance and scalability. We leverage our model-driven analysis to be the basis for an oracle utility which can help in detecting the limitations and bottlenecks of different parallelism approaches at scale. We evaluate the oracle on six parallelization strategies, with four CNN models and multiple datasets (2D and 3D), on up to 1024 GPUs. The results demonstrate that the oracle has an average accuracy of about 86.74% when compared to empirical results, and as high as 97.57% for data parallelism. Albert Kahira, Truong Thao Nguyen, Leonardo Arturo Bautista-Gomez, Ryousei Takano, Rosa M. Badia, Mohamed Wahib |
HPDC | 2 |
| 2021 | Q-learning-based Opportunistic Communication for Real-time Mobile Air Quality Monitoring SystemsabstractWe focus on real-time air quality monitoring systems that rely on devices installed on automobiles in this research. We investigate an opportunistic communication model in which devices can send the measured data directly to the air quality server through a 4G communication channel or via Wi-Fi to adjacent devices or the so-called Road Side Units deployed along the road. We aim to reduce 4G costs while assuring data latency, where the data latency is defined as the amount of time it takes for data to reach the server. We propose an offloading scheme that leverages Q-learning to accomplish the purpose. The experiment results show that our offloading method significantly cuts down around 40-50% of the 4G communication cost while keeping the latency of 99.5% packets smaller than the required threshold. Trung Thanh Nguyen 0006, Truong Thao Nguyen, Tuan Anh Nguyen Dinh, Thanh-Hung Nguyen, Phi-Le Nguyen |
IPCCC | 2 |
| 2021 | Efficient MPI-AllReduce for large-scale deep learning on GPU-clustersabstractSummary Training models on large‐scale GPUs‐accelerated clusters are becoming a commonplace due to the increase in complexity and size in deep learning models. One of the main challenges for distributed training is the collective communication overhead for large message sizes: up to hundreds of MB. In this paper, we propose two hierarchical distributed memory multileader AllReduce algorithms optimized for GPU‐accelerated clusters (named lr_lr and lr_rab ), in which GPUs inside a computing node perform an intra‐node communication phase to gather and store results of local reduced values to designated GPUs (known as node leaders). Node leaders then keep a role as an inter‐node communicator. Each leader exchanges one part of reduced values to the leaders of the other nodes in parallel. Hence, we are capable of significantly reducing the time for injecting data into the inter‐node network. We also overlap the inter‐node and intra‐node communication by implementing our proposal in a pipelined manner. We evaluate those algorithms on the discrete‐event simulation Simgrid. We show that our algorithms, lr_lr and lr_rab , can cut down the execution time of an AllReduce microbenchmark that uses the logical ring algorithm ( lr ) by up to 45% and 51%, respectively. With the pipelined implementation, our lr_lr_pipe achieves 15% performance improvement when compared with lr_lr . In addition, the simulation result also projects power savings for the network devices of up to 23% and 32%. Truong Thao Nguyen, Mohamed Wahib, Ryousei Takano |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | Scaling distributed deep learning workloads beyond the memory capacity with KARMAabstractThe dedicated memory of hardware accelerators can be insufficient to store all weights and/or intermediate states of large deep learning models. Although model parallelism is a viable approach to reduce the memory pressure issue, significant modification of the source code and considerations for algorithms are required. An alternative solution is to use out-of-core methods instead of, or in addition to, data parallelism. We propose a performance model based on the concurrency analysis of out-of-core training behavior, and derive a strategy that combines layer swapping and redundant recomputing. We achieve an average of 1. 52x speedup in six different models over the state-of-the-art out-of-core methods. We also introduce the first method to solve the challenging problem of out-of-core multi-node training by carefully pipelining gradient exchanges and performing the parameter updates on the host. Our data parallel out-of-core solution can outperform complex hybrid model parallelism in training large models, e.g. Megatron-LM and Turning-NLG. Mohamed Wahib, Truong Thao Nguyen, Aleksandr Drozd, Jens Domke, Lingqi Zhang 0001, Ryousei Takano, Satoshi Matsuoka |
SC | 3 |
| 2019 | Topology-aware Sparse Allreduce for Large-scale Deep LearningabstractData parallelism is the dominant method used to scale-up deep learning (DL) training across multiple compute nodes. Collective communication of the local gradients between nodes is a critical bottleneck due to the significant increase in complexity and size of DL models. Researchers cope with this problem by one of the following solutions: a) optimizing the collective communication algorithm to account for the underlying network topology (topology-aware), and b) reducing the amount of transferred data. In the latter approach, sparse communication techniques communicate only the essential data, which helps to significantly cut down the communication cost. However, the diversity of message sizes, and unknown overlap of indicessets among compute nodes (i.e. irregular size), restricts their communication into the many-to-many scheme, which is ineffective in large-scale implementations. In this paper, we present allreduce algorithms that can exploit both sparse-communication and topology-aware techniques by mixing heterogeneous data formats (dense and sparse) with a trivial cost of computation. Truong Thao Nguyen, Mohamed Wahib, Ryousei Takano |
IPCCC | 1 |
| 2018 | Low-Reliable Low-Latency Networks Optimized for HPC Parallel ApplicationsabstractHigh-end network standards, such as 400GbE, have been introduced Forwarding Error Correction (FEC) for maintaining the same bit error rate (BER) as that in traditional low-bandwidth interconnection networks. However, FEC operation latency overhead surprisingly becomes higher than the sum of all the other switch operation overheads, e.g., routing computation and switch allocation. FEC operation latency overhead significantly degrades the performance of parallel applications in HPC systems. Instead, in this study, we exploit the low-latency network design using a Hamming code that does not provide rigid error-free communication. Since it is consistent with existing frame format based on standard Reed-Solomon RS(544,514) with DC(64b/66b) direct linecode and TC(256b/257b) transcode, respectively, the influences upon the other network layer design are limited. Interestingly, a large number of parallel applications can accept the BER in such a Hamming code. Since lowering such a BER improves switch operation latency, the proposed network using the Hamming code improves the execution time of NAS Parallel Benchmarks by 56% on average when compared to the counterpart RS-FEC networks. Truong Thao Nguyen, Hiroki Matsutani, Michihiro Koibuchi |
NCA | 1 |