Nguyen Anh Tu

dblp:157/7974 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0002-0650-8169ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 3 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Federated Aerial Video Captioning With Effective Temporal Adaptation
abstract
Aerial video captioning (VC) facilitates the automatic interpretation of dynamic scenes in remote sensing (RS), supporting critical applications such as disaster response, traffic monitoring, and environmental surveillance. However, challenges like extreme angles and continuous camera motion require adaptive modeling of complex temporal relationships. To tackle these challenges, we leverage an image-language model as the vision encoder and introduce a temporal adaptation module that combines convolution with self-attention layers to both capture local semantics across neighboring frames and model global temporal dependencies. This design allows our model to exploit the multimodal knowledge of the vision encoder while effectively reasoning over the spatiotemporal dynamics. In addition, privacy concerns often restrict access to annotated aerial datasets, posing further challenges for model training. To address this, we develop a federated learning (FL) framework that enables collaborative model training across decentralized clients. Within this framework, we establish a unified benchmark for systematic comparison of temporal adapters, text decoders, and FL strategies, hence filling a gap in existing literature. Extensive experiments validate the robustness of our approach and its potential for advancing aerial VC.
Nguyen Anh Tu, Nursultan Makhanov, Kenzhebek Taniyev, Ton Duc Do
IEEE Geosci. Remote. Sens. Lett.1
2025 Blockchain-Based Assessment System Proposal for Personalized Learning Pathways
Pham-Duc Tho, Nguyen Anh Tu, Vu-Hong Son, Do-Anh Tuan, Nguyen-Nhu Tung
ICALT2
2025 Improving Vision-Language Models With Attention Mechanisms for Aerial Video Classification
abstract
Vision-language models (VLMs), particularly contrastive language-image pretraining (CLIP), have recently demonstrated great success across various vision tasks. However, their potential in aerial video understanding, an increasingly active area of remote sensing (RS), remains underexplored. This is due to challenges posed by aerial data, such as UAV movement, extreme camera angles, and complex spatiotemporal dependencies. To tackle these challenges, we propose an effective method called CLIP-AVC, which adapts CLIP to classify aerial videos into predefined classes. Specifically, we leverage CLIP’s multimodal transferability by utilizing its encoders to extract robust visual and textual features. We then employ a temporal transformer to capture the interactions among the visual features. To address the lack of inductive bias in the CLIP’s visual encoder, we integrate the temporal transformer’s outputs with 3-D features using a cross-transformer, thereby allowing the spatiotemporal locality of aerial videos. In addition, existing methods often fail to explore the semantic alignment between classes and video features. To further overcome these limitations, we propose a context-enriched transformer that employs self-attention mechanisms to adaptively refine visual and textual representations. Experimental results on two benchmark datasets validate the robustness of CLIP-AVC, demonstrating its potential to significantly advance VLMs for aerial scene understanding.
Nguyen Anh Tu, Nartay Aikyn
IEEE Geosci. Remote. Sens. Lett.1
2025 Towards good practice for convolution and attention with PANs in federated medical image classification
Nursultan Makhanov, Nhan Duc Ho, Kok-Seng Wong, Nguyen Anh Tu
J. Supercomput.4
2024 Multi-Stream GCN and CNN for Skeleton-Based Action Recognition
abstract
This paper presents two novel approaches for improving skeleton-based action recognition using Graph Convolutional Networks (GCN) and Convolutional Neural Networks (CNN). In the first approach, we combine GCN and CNN streams that process position and velocity features to improve classification accuracy. In the second approach, we use GCN as an embedding layer for support network CNN to extract features from skeleton data, which significantly improves recognition accuracy. Our experiments on the JHMDB dataset demonstrate that our approaches outperform state-of-the-art methods while using significantly fewer parameters. Additionally, we extended our evaluation to the Kinetics-400 dataset, where our methods showed comparable results with considerably lower model complexity. Our work contributes to the development of more efficient and robust action recognition models.
Kenzhebek Taniyev, Tomiris Zhaksylyk, Nguyen Anh Tu
SMC3
2024 Efficient facial expression recognition framework based on edge computing
Nartay Aikyn, Ardan Zhanegizov, Temirlan Aidarov, Dinh-Mao Bui, Nguyen Anh Tu
J. Supercomput.5
2023 Joint Multiple Intent Detection and Slot Filling with Supervised Contrastive Learning and Self-Distillation
abstract
Multiple intent detection and slot filling are two fundamental and crucial tasks in spoken language understanding. Motivated by the fact that the two tasks are closely related, joint models that can detect intents and extract slots simultaneously are preferred to individual models that perform each task independently. The accuracy of a joint model depends heavily on the ability of the model to transfer information between the two tasks so that the result of one task can correct the result of the other. In addition, since a joint model has multiple outputs, how to train the model effectively is also challenging. In this paper, we present a method for multiple intent detection and slot filling by addressing these challenges. First, we propose a bidirectional joint model that explicitly employs intent information to recognize slots and slot features to detect intents. Second, we introduce a novel method for training the proposed joint model using supervised contrastive learning and self-distillation. Experimental results on two benchmark datasets MixATIS and MixSNIPS show that our method outperforms state-of-the-art models in both tasks. The results also demonstrate the contributions of both bidirectional design and the training method to the accuracy improvement. Our source code is available at https://github.com/anhtunguyen98/BiSLU.
Nguyen Anh Tu, Hoang Thi Thu Uyen, Tu Minh Phuong, Ngo Xuan Bach
ECAI1
2023 A Bidirectional Joint Model for Spoken Language Understanding
abstract
Intent detection and slot filling are two fundamental and important tasks in spoken language understanding (SLU). Motivated by the fact that the intent and slots in a user utterance have a strong relationship, joint models that deal with both tasks in a single framework have become a predominant choice in SLU research. Most existing joint models build two different decoders on top of a shared weight encoder or exploit intent information to detect slots. Some joint models transfer information between two tasks implicitly. In this paper, we propose a bidirectional joint model for SLU that explicitly incorporates intent information into slot filling and slot information into intent detection. Specifically, we first predict a soft intent signal, which is fed into a biaffine classifier to recognize slots. Slot features are then employed along with the utterance representation to predict the final intent. We also introduce a loss function that takes into account three types of losses: soft intent detection, final intent detection, and slot filling. Experimental results on three benchmark datasets ATIS, Snips, and PhoATIS show that our model outperforms previous state-of-the-art models in both tasks with relative error reductions ranging from 6% to 22%.
Nguyen Anh Tu, Duong Xuan Hieu, Tu Minh Phuong, Ngo Xuan Bach
ICASSP1
2022 Vietnamese Capitalization and Punctuation Recovery Models
Hoang Thi Thu Uyen, Nguyen Anh Tu, Ta Duc Huy
INTERSPEECH2
2022 Meta Pseudo Labels for Chest X-ray Image Classification
abstract
Deep Learning methods are getting more and more extensively applied to medical imaging tasks. Nevertheless, very frequently medical images appear unlabelled making it difficult for AI algorithms to utilize the features of the images for classification purposes. Thus, such limitations make it almost impossible to develop robust and accurate algorithm for medical image classification. In this study, we have used a semi-supervised learning method Meta Pseudo Labels which allowed us to train models with a limited amount of labelled data extracted from chest X-ray images. The approach has demonstrated promising results achieving 92.5% of accuracy on the data labelled only for 16%. Additionally, we have also implemented the Transfer Learning approach to obtain higher accuracy on data labelled for only 0.5%. The approach involved initializing the model with the weights obtained from training it on a dataset with higher portion of labelled data. The approach has been proven to be successful averagely increasing the model accuracy on 0.5% of labeled data by 26 percent.
Assanali Abu, Yerkin Abdukarimov, Nguyen Anh Tu, Min-Ho Lee
SMC3
2022 2D Skeleton-based Action Recognition Using Action-Snippets and Sequential Deep Learning
abstract
Human action recognition (HAR) is an active and crucial field of computer vision due to its various applications, such as smart surveillance and human-computer interaction. Recently, the human skeleton, which is compact and intuitive for representing actions and body movements, has been widely used in numerous HAR frameworks. Despite the great success of the skeleton-based HAR, several challenges remain, such as intra-class variability and inter-class similarity. In this paper, we address this task by first proposing a discriminative representation of the action-snippet (i.e., the very short sequence) that captures meaningful characteristics of human pose and body transition. We then employ adequate deep sequential neural networks (DSNNs) to thoroughly learn the temporal relation of action-snippets in a whole sequence. In experiments, the results show that the proposed approach achieves high recognition rates on benchmark datasets while maintaining good computational efficiency (i.e., lightweight networks and high recognition speed).
Aizada Askar, Min-Ho Lee, Thien Huynh-The, Nguyen Anh Tu
SMC4
2021 ViMQ: A Vietnamese Medical Question Dataset for Healthcare Dialogue System Development
Ta Duc Huy, Nguyen Anh Tu, Tran Hoang Vu, Nguyen Phuc Minh, Nguyen Phan, Trung H. Bui, Steven Quoc Hung Truong
ICONIP (6)2
2021 Analyzing Vietnamese Legal Questions Using Deep Neural Networks with Biaffine Classifiers
Nguyen Anh Tu, Hoang Thi Thu Uyen, Tu Minh Phuong, Ngo Xuan Bach
ICONIP (2)1
2021 Physical Activity Recognition With Statistical-Deep Fusion Model Using Multiple Sensory Data for Smart Health
abstract
Nowadays, enhancing the living standard with smart healthcare via the Internet of Things is one of the most critical goals of smart cities, in which artificial intelligence plays as the core technology. Many smart services, deployed according to wearable sensor-based physical activity recognition, have been able to early detect unhealthy daily behaviors and further medical risks. Numerous approaches have studied shallow handcrafted features coupled with traditional machine learning (ML) techniques, which find it difficult to model real-world activities. In this work, by revealing deep features from deep convolutional neural networks (DCNNs) in fusion with conventional handcrafted features, we learn an intermediate fusion framework of human activity recognition (HAR). According to transforming the raw signal value to pixel intensity value, segmentation data acquired from a multisensor system are encoded to an activity image for deep model learning. Formulated by several novel residual triple convolutional blocks, the proposed DCNN allows extracting multiscale spatiotemporal signal-level and sensor-level correlations simultaneously from the activity image. In the fusion model, the hybrid feature merged from the handcrafted and deep features is learned by a multiclass support vector machine (SVM) classifier. Based on several experiments of performance evaluation, our fusion approach for activity recognition has achieved the accuracy over 96.0% on three public benchmark data sets, including Daily and Sport Activities, Daily Life Activities, and RealWorld. Furthermore, the method outperforms several state-of-the-art HAR approaches and demonstrates the superiority of the proposed intermediate fusion model in multisensor systems.
Thien Huynh-The, Cam-Hao Hua, Nguyen Anh Tu, Dong-Seong Kim 0002
IEEE Internet Things J.3
2021 Energy efficiency in cloud computing based on mixture power spectral density prediction
Dinh-Mao Bui, Nguyen Anh Tu, Eui-nam Huh
J. Supercomput.2
2021 Toward efficient and intelligent video analytics with visual privacy protection for large-scale surveillance
Nguyen Anh Tu, Thien Huynh-The, Kok-Seng Wong, M. Fatih Demirci, Young-Koo Lee
J. Supercomput.1
2020 Learning Geometric Features with Dual-stream CNN for 3D Action Recognition
abstract
Recently, regarding several beneficial properties of depth camera, numerous 3D action recognition frameworks have studied high-level features by exploiting deep learning techniques, but nevertheless they cannot seize the meaningful characteristics of static human pose and dynamic action motion of a whole sequence. This paper introduces a deep network configured by two parallel streams of convolutional stacks for fully learning the deep intra-frame joint associations and inter-frame joint correlations, wherein the structure of each stream is learned from Inception-v3. In experiments, besides the compatibility verification with various backbone networks, the proposed approach achieves the state-of-theart performance in battle with several deep learning-based methods on the updated NTU RGB+D 120 dataset..
Thien Huynh-The, Cam-Hao Hua, Nguyen Anh Tu, Dong-Seong Kim 0002
ICASSP3
2020 Learning 3D spatiotemporal gait feature by convolutional network for person identification
Thien Huynh-The, Cam-Hao Hua, Nguyen Anh Tu, Dong-Seong Kim 0002
Neurocomputing3
2019 ML-HDP: A Hierarchical Bayesian Nonparametric Model for Recognizing Human Actions in Video
abstract
Action recognition from videos is an important area of computer vision research due to its various applications, ranging from visual surveillance to human-computer interaction. To address action recognition problems, this paper presents a framework that jointly models multiple complex actions and motion units at different hierarchical levels. We achieve this by proposing a generative topic model, namely, multi-label hierarchical Dirichlet process (ML-HDP). The ML-HDP model formulates the co-occurrence relationship of actions and motion units, and enables highly accurate recognition. In particular, our topic model possesses the three-level representation in action understanding, where low-level local features are connected to high-level actions via mid-level atomic actions. This allows the recognition model to work discriminatively. In our ML-HDP, atomic actions are treated as latent topics and automatically discovered from data. In addition, we incorporate the notion of class labels into our model in a semi-supervised fashion to effectively learn and infer multi-labeled videos. Using discovered topics and inferred labels, which are jointly assigned to local features, we present the straightforward methods to perform three recognition tasks including action classification, joint classification and segmentation of continuous actions, and spatiotemporal action localization. In experiments, we explore the use of three different features and demonstrate the effectiveness of our proposed approach for these tasks on four public datasets: KTH, MSR-II, Hollywood2, and UCF101.
Nguyen Anh Tu, Thien Huynh-The, Kifayat-Ullah Khan, Young-Koo Lee
IEEE Trans. Circuits Syst. Video Technol.1
2018 Selective bit embedding scheme for robust blind color image watermarking
Thien Huynh-The, Cam-Hao Hua, Nguyen Anh Tu, Tae Ho Hur, Jae Hun Bang, Dohyeong Kim, Muhammad Bilal Amin, Byeong Ho Kang 0001, Hyonwoo Seung, Sungyoung Lee 0001
Inf. Sci.3
2018 Hierarchical topic modeling with pose-transition feature for action recognition using 3D skeleton data
Thien Huynh-The, Cam-Hao Hua, Nguyen Anh Tu, Tae Ho Hur, Jae Hun Bang, Dohyeong Kim, Muhammad Bilal Amin, Byeong Ho Kang 0001, Hyonwoo Seung, Soo-Yong Shin, Eun-Soo Kim, Sungyoung Lee 0001
Inf. Sci.3
2017 Featured correspondence topic model for semantic search on social image collections
Nguyen Anh Tu, Kifayat-Ullah Khan, Young-Koo Lee
Expert Syst. Appl.1
2017 Faster compression methods for a weighted graph using locality sensitive hashing
Kifayat-Ullah Khan, Batjargal Dolgorsuren, Nguyen Anh Tu, Waqas Nawaz, Young-Koo Lee
Inf. Sci.3
2016 iTri: Index-based triangle listing in massive graphs
Mostofa Kamal Rasel, Yongkoo Han, Jinseung Kim, Kisung Park 0001, Nguyen Anh Tu, Young-Koo Lee
Inf. Sci.5
2016 Topic modeling and improvement of image representation for large-scale image retrieval
Nguyen Anh Tu, Dong-Luong Dinh, Mostofa Kamal Rasel, Young-Koo Lee
Inf. Sci.1