VLDB 2026 Research / reviewers in the wild / expert
Trung Thanh Nguyen 0006
dblp:18/1411-6
· DBLP profile ↗
20ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0001-8976-2922ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Computer networks · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | View-aware Cross-modal Distillation for Multi-view Action RecognitionabstractThe widespread use of multi-sensor systems has increased research in multi-view action recognition. While existing approaches in multi-view setups with fully overlapping sensors benefit from consistent view coverage, partially overlapping settings where actions are visible in only a subset of views remain underexplored. This challenge becomes more severe in real-world scenarios, as many systems provide only limited input modalities and rely on sequence-level annotations instead of dense frame-level labels. In this study, we propose View-aware Cross-modal Knowledge Distillation (ViCoKD), a framework that distills knowledge from a fully supervised multi-modal teacher to a modality- and annotation-limited student. ViCoKD employs a cross-modal adapter with cross-modal attention, allowing the student to exploit multi-modal correlations while operating with incomplete modalities. Moreover, we propose a View-aware Consistency module to address view misalignment, where the same action may appear differently or only partially across viewpoints. It enforces prediction alignment when the action is co-visible across views, guided by human-detection masks and confidence-weighted Jensen–Shannon divergence between their predicted class distributions. Experiments on the real-world MultiSensor-Home dataset show that ViCoKD consistently outperforms competitive distillation methods across multiple backbones and environments, delivering significant gains and surpassing the teacher model under limited conditions. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide |
WACV | 1 |
| 2026 | PADM: A Physics-aware Diffusion Model for Attenuation CorrectionabstractAttenuation artifacts remain a significant challenge in cardiac Myocardial Perfusion Imaging (MPI) using Single-Photon Emission Computed Tomography (SPECT), often compromising diagnostic accuracy and reducing clinical interpretability. While hybrid SPECT/CT systems mitigate these artifacts through CT-derived attenuation maps, their high cost, limited accessibility, and added radiation exposure hinder widespread clinical adoption. In this study, we propose a novel CT-free solution to attenuation correction in cardiac SPECT. Specifically, we introduce Physics-aware Attenuation Correction Diffusion Model (PADM), a diffusion-based generative method that incorporates explicit physics priors via a teacher–student distillation mechanism. This approach enables attenuation artifact correction using only Non-Attenuation-Corrected (NAC) input, while still benefiting from physics-informed supervision during training. To support this work, we also introduce CardiAC, a comprehensive dataset comprising 424 patient studies with paired NAC and Attenuation-Corrected (AC) reconstructions, alongside high-resolution CT-based attenuation maps. Extensive experiments demonstrate that PADM outperforms state-of-the-art generative models, delivering superior reconstruction fidelity across both quantitative metrics and visual assessment. Trung-Kien Pham, Hoang Minh Vu, Anh Duc Chu, Dac Thai Nguyen, Trung Thanh Nguyen 0006, Thao Nguyen Truong, Hong Son Mai, Phi-Le Nguyen |
WACV | 5 |
| 2026 | DiffCAS: Inference-time CT-free diffusion model for physics-aware multi-slice attenuation correction in cardiac SPECTabstractAttenuation artifacts remain a critical challenge in cardiac Myocardial Perfusion Imaging (MPI) using Single-Photon Emission Computed Tomography (SPECT), often degrading diagnostic accuracy and clinical interpretability. While hybrid SPECT and Computed Tomography (CT) systems mitigate these artifacts using CT-derived attenuation maps, their high cost, radiation exposure, and limited accessibility restrict widespread clinical use. To address these challenges, we propose DiffCAS, an inference-time CT-free diffusion model for physics-aware multi-slice attenuation correction in cardiac SPECT. DiffCAS integrates a Brownian Bridge diffusion process with physics-guided supervision, enabling the generation of attenuation-corrected (AC) images directly from non-attenuation-corrected (NAC) inputs. Specifically, a physics-aware reconstruction module predicts voxel-wise attenuation coefficients and path lengths, then combines them via the Beer-Lambert law into an attenuation correction factor applied at each diffusion step, keeping the AC images physically consistent. The model introduces two key innovations that jointly enhance structural understanding and physics consistency. The first is multi-slice contextual learning, which captures cross-slice anatomical dependencies and improves spatial coherence in reconstructed images. The second is the 3D Computed Tomography Vision Transformer that models long-range volumetric structures and provides physics-consistent attenuation priors to guide the diffusion process. To enable CT-free attenuation correction, DiffCAS introduces the teacher-student distillation framework that transfers physics-informed knowledge from CT-conditioned training to a student network that requires no CT input at inference time, ensuring stability and interpretability. Evaluations on the CardiAC dataset, which comprises 424 patient studies with paired NAC and AC, and CT-based attenuation maps, demonstrate the strong performance of DiffCAS, evaluated using global pixel-level metrics and myocardium-specific clinical metrics. The proposed method surpasses state-of-the-art image generative methods, achieving superior reconstruction accuracy, structural consistency, and diagnostic reliability. These results highlight the proposed DiffCAS as a clinically promising, inference-time CT-free solution for attenuation correction in cardiac SPECT imaging. Hoang Minh Vu, Trung-Kien Pham, Thi Ha Chi Nguyen, Hai-Dang Nguyen, Dac Thai Nguyen, Hong Son Mai, Thanh Trung Nguyen, Trung Thanh Nguyen 0006, Phi-Le Nguyen |
Artif. Intell. Medicine | 8 |
| 2026 | MultiSensor-Home: Multi-modal multi-view dataset and benchmarks for action recognition in home environmentsabstractMulti-modal multi-view action recognition is a rapidly growing area in computer vision, with important applications in surveillance, smart homes, and assistive robotics. However, existing datasets often fail to capture real-world challenges such as distributed sensor layouts, asynchronous data streams, and limited frame-level annotations. To address these limitations, we introduce MultiSensor-Home, a novel multi-modal multi-view dataset specifically designed for realistic residential environments. It comprises 5,250 untrimmed videos recorded in two distinct residential environments, Home-1 and Home-2, using five distributed RGB and audio sensor units, where each video contains multiple sequential actions along with background segments. Each frame in the recorded sequences is manually annotated with fine-grained frame-level action labels, making it, to the best of our knowledge, the first densely annotated multi-view dataset for home activity recognition. To benchmark this dataset, we further propose Act ion Selection Learning-guided Transformer-based Sensor Fusion (ActFusion), a unified method that jointly models temporal dynamics and cross-view correspondence. It dynamically models cross-view relationships and selects informative frames, enabling robust training under both frame-level supervision, where the start and end timings of each action are labeled, and sequence-level supervision, where only action labels are provided. To support reproducible evaluation, we establish a comprehensive benchmark with standardized training and testing protocols. Extensive experiments on MultiSensor-Home and the existing MM-Office datasets show that ActFusion consistently outperforms baseline methods across diverse scenarios. By capturing the challenges in a realistic setting, MultiSensor-Home sets a new benchmark and encourages future research on robust and generalizable action recognition methods. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide |
Pattern Recognit. | 1 |
| 2026 | Hierarchical Global-Local Fusion for One-stage Open-vocabulary Temporal Action DetectionabstractOpen-vocabulary Temporal Action Detection (Open-vocab TAD) extends the detection scope of Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) to unseen action classes specified by vocabularies not included in the training data, within untrimmed video. Typical Open-vocab TAD methods adopt a two-stage approach that first proposes candidate action intervals and then identifies those actions. However, errors in the first stage can affect the subsequent stage and the final detection results. Moreover, conventional methods for temporal context analyses tend to focus solely on either global or local context. Focusing solely on the global context can lead to lack of momentary detail, making it difficult to distinguish one action from another. Conversely, focusing only on the local context makes it challenging to determine the start and end timings of action intervals. To address these challenges, we introduce a one-stage approach named Hierarchical Open-vocab TAD (HOTAD), consisting of two branches: Temporal Context Analysis (TCA) and Video–Text Alignment (VTA). The former utilizes Hierarchical Encoder (HE) to fuse global and local temporal features, enabling a comprehensive capture of temporal actions, while the latter branch exploits the synergy between visual and textual modalities for precisely detecting unseen actions in the Open-vocab setting. Experiments and in-depth analysis using the widely recognized datasets THUMOS14 and ActivityNet-1.3 are performed to show the effectiveness of HOTAD. The results highlight remarkable accuracy in detecting a wide range of unseen actions. Furthermore, HOTAD significantly reduces wrong labels and localizes action instances with high precision, showcasing its robustness in complex and dynamic video settings. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor FusionabstractMulti-modal multi-view action recognition is a rapidly growing field in computer vision, offering significant potential for applications in surveillance. However, current datasets often fail to address real-world challenges such as widearea distributed settings, asynchronous data streams, and the lack of frame-level annotations. Furthermore, existing methods face difficulties in effectively modeling inter-view relationships and enhancing spatial feature learning. In this paper, we introduce the MultiSensor-Home dataset, a novel benchmark designed for comprehensive action recognition in home environments, and also propose the Multi-modal Multi-view Transformer-based Sensor Fusion (MultiTSF) method. The proposed MultiSensor-Home dataset features untrimmed videos captured by distributed sensors, providing high-resolution RGB and audio data along with detailed multi-view frame-level action labels. The proposed MultiTSF method leverages a Transformer-based fusion mechanism to dynamically model inter-view relationships. Furthermore, the proposed method integrates a human detection module to enhance spatial feature learning, guiding the model to prioritize frames with human activity to enhance action the recognition accuracy. Experiments on the proposed MultiSensor-Home and the existing MM-Office datasets demonstrate the superiority of MultiTSF over the state-of-the-art methods. Quantitative and qualitative results highlight the effectiveness of the proposed method in advancing real-world multi-modal multi-view action recognition. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide |
FG | 1 |
| 2025 | IntentVC 2025: The ACM Multimedia Grand Challenge on Intention-Oriented Controllable Video CaptioningabstractThe IntentVC Challenge, held in conjunction with ACM Multimedia 2025, introduces a novel benchmark for intention-oriented controllable video captioning. Unlike conventional captioning methods that generate generic, scene-level summaries, IntentVC focuses on intention-specific generation. Participants are required to produce captions explicitly conditioned on user-defined intentions, such as emphasizing a specific object tracked within a video. To support this task, the challenge provides an extended version of the LaSOT dataset annotated with intention-focused captions across 70 object categories. A standardized evaluation protocol and public leaderboard enable fair and reproducible comparison among submitted methods. By advancing research in personalized and adaptive video understanding, IntentVC offers a platform for exploring controllable vision-language modeling with practical relevance for accessibility, retrieval, and human-AI interaction. As a result, a total of 23 teams and 58 active participants have participated, and a total of 1,443 entries have been submitted. More information and resources are available at https://sites.google.com/view/intentvc/. Takahiro Komamizu, Marc A. Kastner 0001, Yasutomo Kawanishi, Trung Thanh Nguyen 0006, Junan Chen 0004 |
ACM Multimedia | 4 |
| 2025 | Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video CaptioningabstractRecent advances in video captioning are driven by large-scale pretrained models, which follow the standard “pre-training followed by fine-tuning” paradigm, where the full model is fine-tuned for downstream tasks. Although effective, this approach becomes computationally prohibitive as the model size increases. The Parameter-Efficient Fine-Tuning (PEFT) approach offers a promising alternative, but primarily focuses on the language components of Multimodal Large Language Models (MLLMs). Despite recent progress, PEFT remains underexplored in multimodal tasks and lacks sufficient understanding of visual information during fine-tuning the model. To bridge this gap, we propose Query-Adapter (Q-Adapter), a lightweight visual adapter module designed to enhance MLLMs by enabling efficient fine-tuning for the video captioning task. Q-Adapter introduces learnable query tokens and a gating layer into Vision Encoder, enabling effective extraction of sparse, caption-relevant features without relying on external textual supervision. We evaluate Q-Adapter on two well-known video captioning datasets, MSR-VTT and MSVD, where it achieves state-of-the-art performance among the methods that take the PEFT approach across BLEU@4, METEOR, ROUGE-L, and CIDEr metrics. Q-Adapter also achieves competitive performance compared to methods that take the full fine-tuning approach while requiring only 1.4% of the parameters. We further analyze the impact of key hyperparameters and design choices on fine-tuning effectiveness, providing insights into optimization strategies for adapter-based learning. These results highlight the strong potential of Q-Adapter in balancing caption quality and parameter efficiency, demonstrating its scalability for video–language modeling. Junan Chen 0004, Trung Thanh Nguyen 0006, Takahiro Komamizu, Ichiro Ide |
MMAsia | 2 |
| 2025 | MultiSensor-Home: Benchmark for Multi-modal Multi-view Action Recognition in Home EnvironmentsabstractWe present MultiSensor-Home, a benchmark for multi-modal multi-view action recognition in home environments. It consists of a dataset with 5,250 videos captured from five synchronized RGB and Audio sensors across two home environments, covering 16 distinct household activities. Each frame is manually annotated with fine-grained action labels, resulting in a densely labeled multi-view dataset for home activity recognition. To support fair and reproducible evaluation, we establish a comprehensive benchmark with standardized training and testing protocols. Evaluations are conducted under joint-training and environment-specific settings to assess model generalization across diverse home layouts. We benchmark state-of-the-art methods on the proposed dataset, including our previously proposed MultiASL, which uses action selection learning for view-aware fusion under weak supervision, and MultiTSF, which applies sensor fusion and human-centric attention for multi-view action recognition. The results highlight significant performance gaps across domains, especially in challenging scenes, providing valuable insights into the strengths and limitations of current approaches. This underscores the challenges and opportunities in multi-modal multi-view home action understanding and demonstrates the need for further advancements in the field. Trung Thanh Nguyen 0006 |
MMAsia | 1 |
| 2025 | Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report GenerationabstractVision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains, applying these models to medical imaging remains challenging due to the limited availability of diverse imaging modalities and multilingual clinical data. Most existing medical VLMs are trained on a subset of imaging modalities and focus primarily on high-resource languages, thus limiting their generalizability and clinical utility. To address these limitations, we introduce a novel Vietnamese-language multimodal medical dataset consisting of 2,757 whole-body PET/CT volumes from independent patients and their corresponding full-length clinical reports. This dataset is designed to fill two pressing gaps in medical AI development: (1) the lack of PET/CT imaging data in existing VLMs training corpora, which hinders the development of models capable of handling functional imaging tasks; and (2) the underrepresentation of low-resource languages, particularly the Vietnamese language, in medical vision-language research. To the best of our knowledge, this is the first dataset to provide comprehensive PET/CT-report pairs in Vietnamese. We further introduce a training framework to enhance VLMs' learning, including data augmentation and expert-validated test sets. We conduct comprehensive experiments benchmarking state-of-the-art VLMs on downstream tasks, including medical report generation and visual question answering. The experimental results show that incorporating our dataset significantly improves the performance of existing VLMs. However, despite these advancements, the models still underperform on clinically critical criteria, particularly the diagnosis of lung cancer, indicating substantial room for future improvement. We believe this dataset and benchmark will serve as a pivotal step in advancing the development of more robust VLMs for medical imaging, particularly in low-resource languages, and improving their clinical relevance in Vietnamese healthcare. Huu Tien Nguyen, Dac Thai Nguyen, The Minh Duc Nguyen, Trung Thanh Nguyen 0006, Truong Thao Nguyen, Hieu H. Pham 0001, Johan Barthelemy, Minh Quan Tran, Nguyen Quoc Viet Hung, Thanh Tam Nguyen, Hong Son Mai, Quynh Anh Chau, Thanh Hong Nguyen, Phi-Le Nguyen |
NeurIPS | 4 |
| 2025 | CT to PET Translation: A Large-Scale Dataset and Domain-Knowledge-Guided Diffusion ApproachabstractPositron Emission Tomography (PET) and Computed Tomography (CT) are essential for diagnosing, staging, and monitoring various diseases, particularly cancer. Despite their importance, the use of PET/CT systems is limited by the necessity for radioactive materials, the scarcity of PET scanners, and the high cost associated with PET imaging. In contrast, CT scanners are more widely available and significantly less expensive. In response to these challenges, our study addresses the issue of generating PET images from CT images, aiming to reduce both the medical examination cost and the associated health risks for patients. Our contributions are twofold: First, we introduce a conditional diffusion model named CPDM, which, to our knowledge, is one of the initial attempts to employ a diffusion model for translating from CT to PET images. Second, we provide the largest CT-PET dataset to date, comprising 2,028,628 paired CT-PET images, which facilitates the training and evaluation of CT-to-PET translation models. For the CPDM model, we incorporate domain knowledge to develop two conditional maps: the Attention map and the Attenuation map. The former helps the diffusion process focus on areas of interest, while the latter improves PET data correction and ensures accurate diagnostic information. Experimental evaluations across various benchmarks demonstrate that CPDM surpasses existing methods in generating high-quality PET images in terms of multiple metrics. The source code and data samples are available at https://github.com/thanhhff/CPDM. Dac Thai Nguyen, Trung Thanh Nguyen 0006, Huu Tien Nguyen, Thanh Trung Nguyen, Hieu H. Pham 0001, Thanh-Hung Nguyen, Truong Thao Nguyen, Phi-Le Nguyen |
WACV | 2 |
| 2024 | One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-Scale and Action Label FeaturesabstractOpen-vocabulary Temporal Action Detection (Open-vocab TAD) is an advanced video analysis approach that expands Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) capabilities. Closed-vocab TAD is typically confined to localizing and classifying actions based on a predefined set of categories. In contrast, Open-vocab TAD goes further and is not limited to these predefined categories. This is particularly useful in real-world scenarios where the variety of actions in videos can be vast and not always predictable. The prevalent methods in Open-vocab TAD typically employ a 2-stage approach, which involves generating action proposals and then identifying those actions. However, errors made during the first stage can adversely affect the subsequent action identification accuracy. Additionally, existing studies face challenges in handling actions of different durations owing to the use of fixed temporal processing methods. Therefore, we propose a L-stage approach consisting of two primary modules: Multi-scale Video Analysis (MVA) and Video-Text Alignment (VTA). The MVA module captures actions at varying temporal resolutions, overcoming the challenge of detecting actions with diverse durations. The VTA module leverages the synergy between visual and textual modalities to precisely align video segments with corresponding action labels, a critical step for accurate action identification in Open-vocab scenarios. Evaluations on widely recognized datasets THUMOSl4 and ActivityNet-I.3, showed that the proposed method achieved superior results compared to the other methods in both Open-vocab and Closed-vocab settings. This serves as a strong demonstration of the effectiveness of the proposed method in the TAD task. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
FG | 1 |
| 2024 | Action Selection Learning for Multi-label Multi-view Action RecognitionabstractMulti-label multi-view action recognition aims to recognize multiple concurrent or sequential actions from untrimmed videos captured by multiple cameras. Existing work has focused on multi-view action recognition in a narrow area with strong labels available, where the onset and offset of each action are labeled at the frame-level. This study focuses on real-world scenarios where cameras are distributed to capture a wide-range area with only weak labels available at the video-level. We propose the method named Multi-view Action Selection Learning (MultiASL), which leverages action selection learning to enhance view fusion by selecting the most useful information from different viewpoints. The proposed method includes a Multi-view Spatial-Temporal Transformer video encoder to extract spatial and temporal features from multi-viewpoint videos. Action Selection Learning is employed at the frame-level, using pseudo ground-truth obtained from weak labels at the video-level, to identify the most relevant frames for action recognition. Experiments in a real-world office environment using the MM-Office dataset demonstrate the superior performance of the proposed method compared to existing methods. The source code is available at https://github.com/thanhhff/MultiASL/. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
MMAsia | 1 |
| 2024 | FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive RegularizationabstractFederated Learning (FL) is a method for training machine learning models using distributed data sources. It ensures privacy by allowing clients to collaboratively learn a shared global model while storing their data locally. However, a significant challenge arises when dealing with missing modalities in clients’ datasets, where certain features or modalities are unavailable or incomplete, leading to heterogeneous data distribution. While previous studies have addressed the issue of complete-modality missing1, they fail to tackle partial-modality missing2on account of severe heterogeneity among clients at an instance level, where the pattern of missing data can vary significantly from one sample to another. To tackle this challenge, this study proposes a novel framework named FedMAC, designed to address multimodality missing under conditions of partial-modality missing in FL. Additionally, to avoid trivial aggregation of multi-modal features, we introduce contrastive-based regularization to impose additional constraints on the latent representation space. The experimental results demonstrate the effectiveness of FedMAC across various client configurations with statistical heterogeneity, outperforming baseline methods by up to 26% in severe missing scenarios, highlighting its potential as a solution for the challenge of partially missing modalities in federated systems.1Complete missing is when one or more modalities are absent in server and clients’ data.2Partial missing is when only parts of one or more modalities are absent in server and clients’ data Manh Duong Nguyen, Trung Thanh Nguyen 0006, Hieu H. Pham 0001, Trong Nghia Hoang, Phi-Le Nguyen |
NCA | 2 |
| 2024 | FedCert: Federated Accuracy Certification
Minh Hieu Nguyen 0003, Huu Tien Nguyen, Trung Thanh Nguyen 0006, Manh Duong Nguyen, Trong Nghia Hoang, Truong Thao Nguyen, Phi-Le Nguyen |
NCA | 3 |
| 2022 | FedDRL: Deep Reinforcement Learning-based Adaptive Aggregation for Non-IID Data in Federated LearningabstractThe uneven distribution of local data across different edge devices (clients) results in slow model training and accuracy reduction in federated learning. Naive federated learning (FL) strategy and most alternative solutions attempted to achieve more fairness by weighted aggregating deep learning models across clients. This work introduces a novel non-IID type encountered in real-world datasets, namely cluster-skew, in which groups of clients have local data with similar distributions, causing the global model to converge to an over-fitted solution. To deal with non-IID data, particularly the cluster-skewed data, we propose FedDRL, a novel FL model that employs deep reinforcement learning to adaptively determine each client’s impact factor (which will be used as the weights in the aggregation process). Extensive experiments on a suite of federated datasets confirm that the proposed FedDRL improves favorably against FedAvg and FedProx methods, e.g., up to 4.05% and 2.17% on average for the CIFAR-100 dataset, respectively. Nang Hung Nguyen, Phi-Le Nguyen, Thuy Dung Nguyen, Trung Thanh Nguyen 0006, Duc Long Nguyen, Thanh-Hung Nguyen, Hieu H. Pham 0001, Truong Thao Nguyen |
ICPP | 4 |
| 2022 | A Novel Approach for Pill-Prescription Matching with GNN Assistance and Contrastive Learning
Trung Thanh Nguyen 0006, Hoang Dang Nguyen, Thanh-Hung Nguyen, Hieu H. Pham 0001, Ichiro Ide, Phi-Le Nguyen |
PRICAI (1) | 1 |
| 2022 | Deep Reinforcement Learning-based Offloading for Latency Minimization in 3-tier V2X NetworksabstractMulti-access edge computing (MEC) is seen as an effective technique for decreasing service latency in a V2X network by offloading computational activities. With MEC, a three-tier offloading architecture can be developed, where a vehicle can offload computational tasks to a cloud by communicating with a base station (gNB) and Road Side Units (RSUs). In this paper, we focus on three-tier V2X networks which rely on three offloading paths: Vehicle-to-Infrastructure (i.e., vehicle to RSU), Vehicle-to-Cloud (i.e., vehicle to gNB), and Infrastructure-to-Cloud (i.e., RSU to gNB). We propose an offloading strategy based on deep reinforcement learning with the goal of reducing the average latency of tasks. To be more specific, we leverage the Deep Q Network to estimate the goodness of action-state value to determine the offloading decision. We also propose a novel exploration scheme and a new model training strategy. The experimental findings indicate that our proposed offloading method outperforms the state-of-the-art, particularly in critical circumstances characterized by a high rate of vehicle arrival or packet generation. Hieu Dinh, Nang Hung Nguyen, Trung Thanh Nguyen 0006, Thanh-Hung Nguyen, Truong Thao Nguyen, Phi-Le Nguyen |
WCNC | 3 |
| 2022 | Fuzzy Q-Learning-Based Opportunistic Communication for MEC-Enhanced Vehicular CrowdsensingabstractThis study focuses on MEC-enhanced, vehicle-based crowdsensing systems that rely on devices installed on automobiles. We investigate an opportunistic communication paradigm in which devices can transmit measured data directly to a crowdsensing server over a 4G communication channel or to nearby devices or so-called Road Side Units positioned along the road via Wi-Fi. We tackle a new problem that is how to reduce the cost of 4G while preserving the latency. We propose an offloading strategy that combines a reinforcement learning technique known as Q-learning with Fuzzy logic to accomplish the purpose. Q-learning assists devices in learning to decide the communication channel. Meanwhile, Fuzzy logic is used to optimize the reward function in Q-learning. The experiment results show that our offloading method significantly cuts down around 30-40% of the 4G communication cost while keeping the latency of 99% packets below the required threshold. Trung Thanh Nguyen 0006, Truong Thao Nguyen, Thanh-Hung Nguyen, Phi-Le Nguyen |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2021 | Q-learning-based Opportunistic Communication for Real-time Mobile Air Quality Monitoring SystemsabstractWe focus on real-time air quality monitoring systems that rely on devices installed on automobiles in this research. We investigate an opportunistic communication model in which devices can send the measured data directly to the air quality server through a 4G communication channel or via Wi-Fi to adjacent devices or the so-called Road Side Units deployed along the road. We aim to reduce 4G costs while assuring data latency, where the data latency is defined as the amount of time it takes for data to reach the server. We propose an offloading scheme that leverages Q-learning to accomplish the purpose. The experiment results show that our offloading method significantly cuts down around 40-50% of the 4G communication cost while keeping the latency of 99.5% packets smaller than the required threshold. Trung Thanh Nguyen 0006, Truong Thao Nguyen, Tuan Anh Nguyen Dinh, Thanh-Hung Nguyen, Phi-Le Nguyen |
IPCCC | 1 |