Duy-Linh Nguyen

dblp:197/3328 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
22since 2021 · last 2025
0000-0001-6184-4133ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Artificial Behavior Intelligence: Technology, Challenges, and Future Directions
abstract
Understanding and predicting human behavior has emerged as a core capability in various AI application domains such as autonomous driving, smart healthcare, surveillance systems, and social robotics. This paper defines the technical frame-work of Artificial Behavior Intelligence (ABI), which comprehensively analyzes and interprets human posture, facial expressions, emotions, behavioral sequences, and contextual cues. It details the essential components of ABI, including pose estimation, face and emotion recognition, sequential behavior analysis, and context-aware modeling. Furthermore, we highlight the transformative potential of recent advances in large-scale pretrained models, such as large language models (LLMs), vision foundation models, and multimodal integration models, in significantly improving the accuracy and interpretability of behavior recognition. Our research team has a strong interest in the ABI domain and is actively conducting research, particularly focusing on the development of intelligent lightweight models capable of efficiently inferring complex human behaviors. This paper identifies several technical challenges that must be addressed to deploy ABI in real-world applications including learning behavioral intelligence from limited data, quantifying uncertainty in complex behavior prediction, and optimizing model structures for low-power, real-time inference. To tackle these challenges, our team is exploring various optimization strategies including lightweight transformers, graph-based recognition architectures, energy-aware loss functions, and multimodal knowledge distillation, while validating their applicability in real-time environments.
Kang-Hyun Jo, Jehwan Choi, Kwanho Kim, Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Tien-Dat Tran
HSI5
2025 Efficient Human Behavior Detector for Vision-based Emergency Evacuation Systems
abstract
The emergency evacuation systems are often installed in crowded places such as airports, train stations, and shopping malls to evacuate and protect people when incidents occur quickly. With the development of surveillance cameras, vision-based emergency evacuation systems have demonstrated their ability to observe and promptly warn flexibly. This paper proposes a human behavior detector by fine-tuning the YOLOv11n detection network with the Global Attention Mechanism (GAM) to enhance the individual human action recognition. Extensive experiments are trained and evaluated on the Human Behavior Detection Dataset (HBDset) using a NVIDIA Tesla V100 32GB GPU. The proposed network achieves 62.2% of mAP and an inference speed of 1.3 milliseconds (ms), and outperforms other networks of the same scale.
Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Jehwan Choi, Kang-Hyun Jo
HSI1
2025 A Hybrid Vision Transformer and Convolutional Neural Network Architecture for Banana Leaf Disease Classification
abstract
The classification of banana leaf disease plays a crucial role in early disease detection and in preventing the condition from worsening. To handle this task, this study proposes a hybrid Vision Transformer (ViT) architecture that leverages the strengths of both the convolutional and self-attention layers. By leveraging convolutional layers in the earlier stages and self-attention layers in the later stages, the proposed architecture aims to balance effective feature learning and computational cost while achieving better efficiency. Experimental results show that this model achieves an outstanding accuracy of up to 97.65% while maintaining a moderate tradeoff in computational complexity.
Thi-Kim-Anh Pham, Duy-Linh Nguyen, Xuan-Thuy Vo, Kang-Hyun Jo
HSI2
2025 Efficiency-Accuracy Trade-Off of Facial Attribute Classifier Supporting Human-Robot Interaction
abstract
The advancement of robotics has been driven by the integration of artificial intelligence, machine learning, and sophisticated sensing technologies, enabling more seamless Human-Robot Interaction (HRI). Facial Attribute Classifier (FAC) plays a crucial role in HRI by helping robots understand human emotions, intentions, and social cues, fostering personalized and intuitive interactions. However, while existing methods achieve high accuracy, their computational complexity limits real-time applications on low-cost or CPU-based devices, highlighting the need for lightweight models that balance accuracy and efficiency. This work proposes an Efficient Network (ENet) designed to achieve an optimal trade-off between efficiency and accuracy of FAC. ENet introduces an Enhanced Sequential Efficient Attention Module (ESEAM) to improve the quality of feature maps while maintaining high efficiency. Accordingly, ENet demonstrates a compromise between efficiency and accuracy on the CelebA and LFWA datasets. The proposed ENet is computationally efficient, generating a few parameters, making it well-suited for CPU-based applications. When combined with a face detector, the optimized FAC achieves a processing speed of 25.88 frames per second (FPS) on an Intel Core i7-9750H CPU, demonstrating its suitability for real-time use.
Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Ge Cao, Jehwan Choi, Kang-Hyun Jo
HSI2
2025 Efficient Multi-Scale Spatial Interactions for Visual Recognition Tasks
abstract
Convolution operation has local connectivity and translation equivalence while self-attention operation captures long-range spatial dependencies. Adopting the merits of convolution and self-attention operations in hierarchical networks can result in better visual representation and generalization performance. However, integrating self-attention layers into earlier stages is inefficient because self-attention operation has quadratic complexity with token lengths. In this work, we tackle this issue and propose an Efficient Multi-scale Spatial interaction Network (EMSNet) that takes advantage of hybrid networks. The EMSNet has key insights: (1) Each stage efficiently models both short-range and long-range spatial interactions via the design of the multi-scale tokens; (2) The novel convolution-based multi-head self-attention (C-MHSA) operation is introduced to learn spatial interactions inside local regions; (3) The efficient combination of the depthwise convolution, coordinate depthwise convolution, C-MHSA, and global multi-head self-attention (G-MHSA) are performed via channel splitting strategy, extracting wide ranges of frequencies and multi-order interactions. Extensive experiments on ImageNet-1K image classification, MS-COCO object detection, and segmentation tasks verify the effectiveness and generalization ability of the EMSNet. For instance, the EMS Net-XTiny gets 77.1% Top-1 accuracy on ImageNet-1K which is much greater than PVTvl-Tiny by 2% with only 22% parameters and 37% GFLOPs.
Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Jehwan Choi, Kang-Hyun Jo
HSI2
2025 A High-Accuracy and Faster Face Recognizer Supporting Biometric Continuous Authentication for Smart Factory Workers
abstract
Smart factories require secure and sustainable worker authentication for safe operations. Biometric continuous authentication based on facial recognition is one of the most convenient mechanisms. This method applies a face recognition task to verify the captured face as an authorized user. However, existing methods that employ large networks for high-accuracy face recognition incur high computational costs and slow down the process, rendering them unsuitable for continuous operation. This work proposes an efficient and rapid face recognizer with high accuracy. It offers a faster face residual network, containing efficient FasterFace blocks and efficient channel spatial attention for improved feature extraction. As a result, the proposed network achieves 97.08% based on average accuracy, outperforming the other networks on five benchmark datasets. It performs faster at 19.91 frames per second in real time on CPU-based hardware when integrated with a face detector, showcasing its capacity to support real-time biometric continuous authentication for smart factory workers.
Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Muhamad Dwisnanto Putro, Ge Cao, Kang-Hyun Jo
IEEE Trans. Ind. Informatics2
2025 Local Self-Attention With Mixing Abstract Tokens for Urban Autonomous Driving
abstract
Although local self-attentions exhibit translation equivariance and locality similar to convolution, the model has limited receptive fields and weak modeling ability. The main reason is that self-attention is computed within nonoverlapped windows. To overcome this issue, common methods need further operations to communicate the information across windows, such as window shifting, and sliding. These operations are memory unfriendly, not well supported, and optimized by modern deep-learning frameworks. Alternatively, this article exchanges information across nonoverlapped windows via efficiently mixing abstract tokens (MAT). The MAT block includes the following steps. First, the image tokens are partitioned into windows and each window is merged with an abstract token. Second, in each window, interactions of image tokens and the abstract token to image tokens are performed. Third, because the abstract token learns abstract information from each corresponding window, mixing all abstract tokens via transformer encoder helps to exchange information between local windows and result in global context modeling. Fourth, the global information of the mixed tokens is propagated back to the image tokens through transformer decoder. The MAT block is efficient and easy to implement, only containing matrix multiplications. In addition, this article also proposes a bilinear patch embedding that samples relevant regions of the input tokens based on learned offsets. Extensive experiments are conducted and evaluated with various tasks such as image classification, object detection, and segmentation. As a result, our method achieves promising performances across tasks. For example, MAT-2 accomplishes79.0%top-1 accuracy on ImageNet-1 K with0.7GFLOPs and outperforms the baseline Swin-0.7 G by4.6%while reducing15.2 mson CPU and0.53 mson GPU devices. The MAT-4 surpasses Swin-T by1.8%mIoU with only70%GFLOPs.
Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Ge Cao, Jehwan Choi, Kang-Hyun Jo
IEEE Trans. Ind. Informatics2
2024 Efficient Vision Transformers with Partial Attention
Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Kang-Hyun Jo
ECCV (83)2
2024 EMPCNet: Facial Attribute Recognition Using Efficient Multi - Perspective Convolution for Human-Robot Interaction
abstract
Human-robot interaction has evolved into a significant field in robotics. In this domain, facial attributes are essential as they enable robots to understand human emotions, intentions, and preferences. In robot applications, which typically involve low-cost devices, efficient recognition technology is crucial for promising real-time operation by robots. This work proposes EMPCNet to perform facial attribute recognition, consisting of an Efficient Multi-Perspective Convolution (EMPC) block used to efficiently extract and capture various information from multiple perspectives using different kernel sizes and shapes of convolutional operations. The proposed network, which only utilizes a few parameters and low computational operations, achieves competitive performance on the CelebA and LFWA datasets. Additionally, when integrated with face detection, the proposed EMPCNet operates efficiently in real-time on a CPU with Intel Core i7-9750H, achieving a frame rate of 21.27 frames per second (FPS) with an image input size of$224\times 224$consisting of a face area.
Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, RussoMohammadAshraf Uddin, Kang-Hyun Jo
HSI2
2024 Simple Human Fall Surveillance System Based on Person Detection
abstract
Human fall is a common problem that often occurs with the elderly, disabled people, and people with bone diseases and neurological diseases. Sometimes, it also comes from human carelessness. Detecting and warning of human falls can minimize the unfortunate risks. Therefore, human fall detection has been widely applied in medical care and surveillance systems. This paper proposes a simple human fall surveillance system based on a person detection network. This system utilizes the pre-trained YOLOv8 network architecture with a related person body dataset. The proposed system reduces the computational complexity and simplifies the use of available datasets for building a surveillance system. As a result, the proposed system achieves the best speed at 206 Frames per second (FPS) when testing on a GeForce GTX 1080Ti 11GB GPU.
Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Duc-Vuong Nguyen, Thi-Le-Hang Nguyen, Kang-Hyun Jo
IECON1
2024 Wider Neighborhood-Aware Attention in Improving YOLOv8n for One-Stage Human Fall Detection
abstract
Human fall detection has become a crucial technology in bolstering intelligent surveillance systems. A one-stage human fall detection model based on the YOLO network emerges as an ideal solution for implementation in limited resource environments, supporting real-time operation with faster speed. This work introduces a Wider Neighborhood-Aware Attention (WN2A) module to enhance YOLOv8n performance for one-stage human fall detection on a CPU device. WN2A enables the YOLOv8n network to focus on crucial information within the feature map based on the channel while considering a wider neighborhood area from a spatial point of view. As a result, the proposed WN2A applied on the YOLOv8n network outperforms the other methods based on the mean Average Precision (mAP) of two benchmark datasets. Moreover, the improved YOLOv8n network enables operating at 27.38 frames per second on an Intel Core i7-9750H CPU while providing higher mAP.
Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Jehwan Choi, Kang-Hyun Jo
IECON2
2024 Lightweight CNN-Based Driver Eye Status Surveillance for Smart Vehicles
abstract
Traffic accidents are the leading death rate among accident categories. One of the major causes of road traffic accidents is driver drowsiness. Many studies have paid attention to this issue and developed driver assistance tools to reduce the risk. These methods mainly analyze driver behavior, vehicle behavior, and driver physiology. This article proposes a driver eye status surveillance system based on lightweight convolutional neural networks (CNNs). The overall system consists of the following three stages: Face detection, eye detection, and eye classification. In the first stage, the system utilizes a small real-time face detector, named nano YOLO5Face. The second stage focuses on exploiting the compact CNN network architecture combined with the inception network, and triplet attention mechanism. Finally, the system uses a simple classification network architecture to classify open or closed eye status. Additionally, this work also provides the datasets for the eye detection task comprised of 10 659 images and 21 318 labels. As a result, the real-time testing reached 33.12 frames per second (FPS) and 25.11 FPS on an Intel Core I7-4770 CPU @ 3.40 GHz [personal computer (PC)] and a 128-core Nvidia Maxwell GPU (Jetson Nano device), respectively.
Duy-Linh Nguyen, Muhamad Dwisnanto Putro, Kang-Hyun Jo
IEEE Trans. Ind. Informatics1
2023 Vehicle Detector Based on Improved YOLOv5 Architecture for Traffic Management and Control Systems
abstract
Vehicle detection is an important module in traffic management and control systems. These systems require compactness, mobility, and high accuracy when deployed in a real-time context. Based on the YOLOv5 network architecture, this paper proposes several improvements to increase the performance and speed of the network when applied to vehicle detection. The research aims to redesign the backbone and neck modules with lightweight convolutional network architectures such as EfficientNet, PP-LCNet, and MobileNet. In addition, the Squeeze-and-Excitation (SE) attention architecture is also used inside the above-mentioned architectures to help the network focus on salient information during feature extraction. The network is trained and evaluated on a modified and normalized dataset of the UA-DETRAC dataset. As a result, the proposed network achieves 58.1% of [email protected] and 40.1% of [email protected]:0.95 with just over ten million network parameters. This result outperforms other methods and is comparable to the lightweight architectures of the YOLOv5 family.
Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Kang-Hyun Jo
IECON1
2023 Facial Attribute Recognition Using Lightweight Multi-Label CNN-Transformer Architecture for Intelligent Advertising
abstract
In modern cities, intelligent advertising platforms have been widely engaged in public areas. A facial attribute recognition technique is essential to assist these platforms in delivering suitable adverts for each audience. These platforms also require a recognition technology that can operate at least suitably on a CPU device to reduce implementation costs. This work proposed a lightweight multi-label CNN-Transformer architecture with an efficient inception block (EIB) and squeeze channel transformer encoder (SCTE) to perform facial attribute recognition efficiently. EIB is used to extract face features in multi-scale and levels supported by SCTE in improving its feature map's quality. The proposed architecture produces fewer parameters with low operations and gains competitive accuracy on the CelebA and LWFA datasets consisting of images with multi-label. Moreover, the proposed architecture integrated with face detection can perform sufficiently on a CPU configuration in real-time with 21 frames per second (FPS) using 224 × 224 input size of face area image.
Adri Priadana, Muhamad Dwisnanto Putro, Jinsu An, Duy-Linh Nguyen, Xuan-Thuy Vo, Kang-Hyun Jo
IECON4
2022 Multi-level Feature Reweighting and Fusion for Instance Segmentation
abstract
Accurate instance segmentation requires high-resolution features for performing a dense pixel-wise prediction task. However, using high-resolution feature maps results in highly expensive model complexity and ineffective receptive fields. To overcome the problems of high-resolution features, conventional methods explore multi-level feature fusion that exchanges the information between low-level features at earlier layers and high-level features at top layers. Both low and high information is extracted by the hierarchical backbone network where high-level features contain more semantic cues and low-level features encompass more specific patterns. Thus, adopting these features to the training segmentation model is necessary, and designing a more efficient multi-level feature fusion is crucial. Existing methods balance such information by using top-down and bottom-up pathway connections with more inefficient convolution layers to produce richer multi-scale features. In this work, we contribute two folds: (1) a simple but effective multilevel feature reweighting layer is proposed to strengthen deep high-level features based on channel reweighting generated from multiple features of the backbone, and (2) an efficient fusion block is proposed to process low-resolution features in a depth-to-spatial manner and combine enhanced multi-level features together. These designs enable the segmentation models to predict instance kernels for mask generation on high-level feature maps. To verify the effectiveness of the proposed method, we conduct experiments on the challenging benchmark dataset MS-COCO. Surprisingly, our simple network outperforms the baseline in both accuracy and inference speed. More specifically, we achieve 35.4% APmaskat 19.5 FPS on a GPU device, becoming a state-of-the-art instance segmentation method.
Xuan-Thuy Vo, Tien-Dat Tran, Duy-Linh Nguyen, Kang-Hyun Jo
INDIN3
2022 A Fast CPU Real-Time Facial Expression Detector Using Sequential Attention Network for Human-Robot Interaction
abstract
Facial expression detection is a method to predict human facial emotions. This work is a trending research topic that can be implemented for human-robot interaction. More recently, deep convolutional neural network provides a robust extractor features but tends to be slow in real-time implementations and often requires a large memory and graphics processing units for fast execution. In this article, an efficient CPU-based facial expression detector is proposed using a sequential attention network to improve the baseline performance. The proposed attention network consists of three modules, global representation to capture the global features, channel representation, and dimension representation, which are focused on the channel and using spatial attention to discriminate local features. The efficient partial transfer module is also presented as a light backbone to extract facial features from an image. The entire module is trained and tested on several benchmarks to classify seven facial expressions. As a result, the proposed model reaches an accuracy of 98.18%, 98.75%, 95.63%, and 74.17% on CK+, JAFFE, KDEF, and FER-2013, respectively. It achieves competitive performance when compared to state-of-the-art methods. Lastly, it is integrated with a face detector and runs in real-time without a constraint at 69 frames per second on a CPU.
Muhamad Dwisnanto Putro, Duy-Linh Nguyen, Kang-Hyun Jo
IEEE Trans. Ind. Informatics2
2021 Eye State Recognizer Using Light-Weight Architecture for Drowsiness Warning
Duy-Linh Nguyen, Muhamad Dwisnanto Putro, Kang-Hyun Jo
ACIIDS1
2021 Real-Time Multi-view Face Mask Detector on Edge Device for Supporting Service Robots in the COVID-19 Pandemic
Muhamad Dwisnanto Putro, Duy-Linh Nguyen, Kang-Hyun Jo
ACIIDS2
2021 Efficient Face Detector Using Spatial Attention Module in Real-Time Application on an Edge Device
Muhamad Dwisnanto Putro, Duy-Linh Nguyen, Kang-Hyun Jo
ICIC (1)2
2021 Regression-Aware Classification Feature for Pedestrian Detection and Tracking in Video Surveillance Systems
Xuan-Thuy Vo, Tien-Dat Tran, Duy-Linh Nguyen, Kang-Hyun Jo
ICIC (1)3
2021 Light-weight Convolutional Neural Network for Distracted Driver Classification
abstract
Driving is an activity that requires the coordination of many senses with complex manipulations. However, the driver can be affected by a several factors such as using a mobile phone, adjusting audio equipment, smoking, drinking, eating, talking to a passenger or drowsy. Therefore, the development of assistant applications to warn distracted driver is very necessary. Because of the limited space and mobility, the equipment also requires compact, energy-saving and efficient. This paper proposes a lightweight Convolutional Neural Network for a distracted driver warning system. The method is built based on a combination of standard convolution and Depthwise Separable Convolution operation to optimize the network parameters but still ensure the important information and speed. The network was trained and evaluated on two datasets, AUC (the American University in Cairo) and StateFarm dataset from Kaggle’s competition. As a result, the evaluation accuracy reached 95.36% and 99.95%, respectively.
Duy-Linh Nguyen, Muhamad Dwisnanto Putro, Xuan-Thuy Vo, Kang-Hyun Jo
IECON1
2021 Dynamic Multi-Loss Weighting for Multiple People Tracking in Video Surveillance Systems
abstract
Multiple people tracking is a fundamental yet challenging task in the computer vision field, which served as a primary process for high-level tasks such as human behaviors, action recognition, pose estimation. Person tracking is decomposed into detection and re-identification (re-ID) sub-tasks. Conventionally, the detection learns classification and regression objectives simultaneously; and the re-ID sub-task is treated as a classification task. Therefore, person tracking is multiple task learning corresponding to multiple loss functions (multiple objectives) with one bounding box regression and two classifications. The difference between various tasks is as follows: the ranges of each objective are inconsistent, the contribution of each task to the overall gradient is altered, and the learning pace of each task is different (level of difficulty). It leads to an objective imbalance in multi-task learning. Previous methods proposed weighting factors as new hyper-parameters to balance the ranges of each task. The dimension of search space for manually tuning these hyper-parameters is high, which depends on the number of tasks. Accordingly, selecting reasonable weighting factors is difficult and complicated. This paper introduces dynamic multi-loss weighting (DMW) with simple but effective in which the weighting factors are dynamically changed during training without introducing any hyper-parameters. The dynamic weights are optimized to balance regression and classification objectives, which depend on the difficulty level of each task and the correlation between each task. Additionally, the general convolution operations are spatially invariant to some degree, which hinders the network’s performance. Hence, this work employs the position-sensitive operation improving feature extraction. The proposed method is conducted on the MOT17 challenging benchmark, which outperforms the online multiple people trackers without using additional data.
Xuan-Thuy Vo, Tien-Dat Tran, Duy-Linh Nguyen, Kang-Hyun Jo
INDIN3
2020 Lightweight Convolutional Neural Network for Real-Time Face Detector on CPU Supporting Interaction of Service Robot
abstract
Face detection plays an essential role in the success of the interaction between service robots and consumers. This method is the initial stage for face-related applications. Practical applications require face detection to work in real-time and can be implemented on low-cost devices such as CPU. Traditional methods have problems when the face is not frontal, blocked, and partially covered, but real-time speed is not an obstacle. On the other hand, deep learning has succeeded in accurately distinguishing facial features and backgrounds. Face sizes that tend to be medium and large when robot interaction with consumers so it can employ Convolutional Neural Networks (CNN) with light weights. In this paper, a real-time face detector is built that can work on the CPU. This detector will be implemented explicitly in service robots to support interactions with consumers. It can overcome the occlusion and not-frontal face. Detector architecture consists of the backbone as rapidly features extractor, transition module as a transformer of prediction map, and the dual-detection layer is head of a network prediction based on scale assignment. As a result, the detector can work at speeds of 301 frames per second on CPU without ignoring the accuracy.
Muhamad Dwisnanto Putro, Duy-Linh Nguyen, Kang-Hyun Jo
HSI2
2020 Eyes Status Detector Based on Light-weight Convolutional Neural Networks supporting for Drowsiness Detection System
abstract
The drowsiness is the leading cause of many accidents on the road. These causes can be reduced by using the drowsiness alarm or drowsiness detection system. These systems monitor drivers while driving and alarm when they don't focus or have some abnormal signs in the driver's body. Currently, most methodologies use the analysis of human behaviors, vehicle behaviors, and human physiological conditions. This paper regards eyes status analysis based on deep learning method using proposed Convolutional Neural Networks (CNN) with two stages are face detection and eyes classification. The face detector employs a single detector module and shallow layer, then the eyes classifier using simple CNN without ignoring the accuracy. As a result, the average speed was tested in real-time by 50.03 fps (frames per second) on Intel Core I7-4770 CPU @ 3.40 GHz.
Duy-Linh Nguyen, Muhamad Dwisnanto Putro, Kang-Hyun Jo
IECON1
2020 A Dual Attention Module for Real-time Facial Expression Recognition
abstract
In this paper, a real-time face expression based on a Convolutional Neural Network with a Dual Attention Module is presented for classifying various facial emotions. The system contains two main components. Firstly, local convolutional features of faces are extracted by the VGG13 baseline. Secondly, dual attention masks are automatically employed to enhance the backbone end based on the global probability of features. It represents the position and channel of the feature map. The local features from baseline are combined with the attention to infer the emotional label module. A single network is trained in an end-to-end scheme with five million parameters. Experiments on benchmark datasets show the attention module gives increased accuracy. Besides, this module provides lightweight and runs 60.20 frames per second when working in real-time on CPU devices.
Muhamad Dwisnanto Putro, Duy-Linh Nguyen, Kang-Hyun Jo
IECON2