EDBT 2026 Demo / reviewers in the wild / expert
Xuan-Thuy Vo
dblp:270/5527
· DBLP profile ↗
23ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0002-7411-0697ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dilated multi-Layer perceptron mixer for faster neural networks
Van-Dung Hoang, Xuan-Thuy Vo, Kang-Hyun Jo |
Neural Networks | 2 |
| 2026 | Optimal Proxy Mining Contrastive Network for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) performance enhancement hinges on extracting the most informative features from unlabeled person datasets. In recent approaches, proxy-based contrastive learning with awareness of camera labels has been adopted for model training, thereby achieving highly promising results. However, inappropriate selections of contrastive pairs can significantly degrade the performance of these models. To address this issue, we propose the Optimal Proxy Mining Contrastive Network (OPMCN), a novel framework designed to strategically optimize the selection of proxies for positive and negative pair formation, thus enhancing the efficacy of contrastive training. The OPMCN framework proposes two specific contrastive losses: Hardest Camera Proxy Mining (HCPM) and False Negative Proxies Mining (FNPM), each essential for enhancing model performance in unsupervised settings. The HCPM loss targets proxies from the most challenging cameras to maximize semantic differences between pairs while ensuring minimal background shifts. In contrast, the FNPM loss counters noise in pseudo labels by prioritizing similarity rankings over clustering results to effectively identify and correct false negatives among proxies. Moreover, we have developed the Pyramid Kernel Global Context (PKGC) block, which employs an attention mechanism that focuses on identity-invariant semantic cues in instances. This module utilizes optimally sized convolutional kernels to enhance identity recognition consistency across camera-based variations, thereby improving the precision of feature extraction. Experimental results on several popular datasets prove that our work surpasses existing unsupervised person Re-ID approaches to a remarkable extent. Ge Cao, Qing Tang 0004, Xuan-Thuy Vo, Adri Priadana, Kang-Hyun Jo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Artificial Behavior Intelligence: Technology, Challenges, and Future DirectionsabstractUnderstanding and predicting human behavior has emerged as a core capability in various AI application domains such as autonomous driving, smart healthcare, surveillance systems, and social robotics. This paper defines the technical frame-work of Artificial Behavior Intelligence (ABI), which comprehensively analyzes and interprets human posture, facial expressions, emotions, behavioral sequences, and contextual cues. It details the essential components of ABI, including pose estimation, face and emotion recognition, sequential behavior analysis, and context-aware modeling. Furthermore, we highlight the transformative potential of recent advances in large-scale pretrained models, such as large language models (LLMs), vision foundation models, and multimodal integration models, in significantly improving the accuracy and interpretability of behavior recognition. Our research team has a strong interest in the ABI domain and is actively conducting research, particularly focusing on the development of intelligent lightweight models capable of efficiently inferring complex human behaviors. This paper identifies several technical challenges that must be addressed to deploy ABI in real-world applications including learning behavioral intelligence from limited data, quantifying uncertainty in complex behavior prediction, and optimizing model structures for low-power, real-time inference. To tackle these challenges, our team is exploring various optimization strategies including lightweight transformers, graph-based recognition architectures, energy-aware loss functions, and multimodal knowledge distillation, while validating their applicability in real-time environments. Kang-Hyun Jo, Jehwan Choi, Kwanho Kim, Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Tien-Dat Tran |
HSI | 6 |
| 2025 | Efficient Human Behavior Detector for Vision-based Emergency Evacuation SystemsabstractThe emergency evacuation systems are often installed in crowded places such as airports, train stations, and shopping malls to evacuate and protect people when incidents occur quickly. With the development of surveillance cameras, vision-based emergency evacuation systems have demonstrated their ability to observe and promptly warn flexibly. This paper proposes a human behavior detector by fine-tuning the YOLOv11n detection network with the Global Attention Mechanism (GAM) to enhance the individual human action recognition. Extensive experiments are trained and evaluated on the Human Behavior Detection Dataset (HBDset) using a NVIDIA Tesla V100 32GB GPU. The proposed network achieves 62.2% of mAP and an inference speed of 1.3 milliseconds (ms), and outperforms other networks of the same scale. Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Jehwan Choi, Kang-Hyun Jo |
HSI | 2 |
| 2025 | A Hybrid Vision Transformer and Convolutional Neural Network Architecture for Banana Leaf Disease ClassificationabstractThe classification of banana leaf disease plays a crucial role in early disease detection and in preventing the condition from worsening. To handle this task, this study proposes a hybrid Vision Transformer (ViT) architecture that leverages the strengths of both the convolutional and self-attention layers. By leveraging convolutional layers in the earlier stages and self-attention layers in the later stages, the proposed architecture aims to balance effective feature learning and computational cost while achieving better efficiency. Experimental results show that this model achieves an outstanding accuracy of up to 97.65% while maintaining a moderate tradeoff in computational complexity. Thi-Kim-Anh Pham, Duy-Linh Nguyen, Xuan-Thuy Vo, Kang-Hyun Jo |
HSI | 3 |
| 2025 | Efficiency-Accuracy Trade-Off of Facial Attribute Classifier Supporting Human-Robot InteractionabstractThe advancement of robotics has been driven by the integration of artificial intelligence, machine learning, and sophisticated sensing technologies, enabling more seamless Human-Robot Interaction (HRI). Facial Attribute Classifier (FAC) plays a crucial role in HRI by helping robots understand human emotions, intentions, and social cues, fostering personalized and intuitive interactions. However, while existing methods achieve high accuracy, their computational complexity limits real-time applications on low-cost or CPU-based devices, highlighting the need for lightweight models that balance accuracy and efficiency. This work proposes an Efficient Network (ENet) designed to achieve an optimal trade-off between efficiency and accuracy of FAC. ENet introduces an Enhanced Sequential Efficient Attention Module (ESEAM) to improve the quality of feature maps while maintaining high efficiency. Accordingly, ENet demonstrates a compromise between efficiency and accuracy on the CelebA and LFWA datasets. The proposed ENet is computationally efficient, generating a few parameters, making it well-suited for CPU-based applications. When combined with a face detector, the optimized FAC achieves a processing speed of 25.88 frames per second (FPS) on an Intel Core i7-9750H CPU, demonstrating its suitability for real-time use. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Ge Cao, Jehwan Choi, Kang-Hyun Jo |
HSI | 3 |
| 2025 | Efficient Multi-Scale Spatial Interactions for Visual Recognition TasksabstractConvolution operation has local connectivity and translation equivalence while self-attention operation captures long-range spatial dependencies. Adopting the merits of convolution and self-attention operations in hierarchical networks can result in better visual representation and generalization performance. However, integrating self-attention layers into earlier stages is inefficient because self-attention operation has quadratic complexity with token lengths. In this work, we tackle this issue and propose an Efficient Multi-scale Spatial interaction Network (EMSNet) that takes advantage of hybrid networks. The EMSNet has key insights: (1) Each stage efficiently models both short-range and long-range spatial interactions via the design of the multi-scale tokens; (2) The novel convolution-based multi-head self-attention (C-MHSA) operation is introduced to learn spatial interactions inside local regions; (3) The efficient combination of the depthwise convolution, coordinate depthwise convolution, C-MHSA, and global multi-head self-attention (G-MHSA) are performed via channel splitting strategy, extracting wide ranges of frequencies and multi-order interactions. Extensive experiments on ImageNet-1K image classification, MS-COCO object detection, and segmentation tasks verify the effectiveness and generalization ability of the EMSNet. For instance, the EMS Net-XTiny gets 77.1% Top-1 accuracy on ImageNet-1K which is much greater than PVTvl-Tiny by 2% with only 22% parameters and 37% GFLOPs. Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Jehwan Choi, Kang-Hyun Jo |
HSI | 1 |
| 2025 | A High-Accuracy and Faster Face Recognizer Supporting Biometric Continuous Authentication for Smart Factory WorkersabstractSmart factories require secure and sustainable worker authentication for safe operations. Biometric continuous authentication based on facial recognition is one of the most convenient mechanisms. This method applies a face recognition task to verify the captured face as an authorized user. However, existing methods that employ large networks for high-accuracy face recognition incur high computational costs and slow down the process, rendering them unsuitable for continuous operation. This work proposes an efficient and rapid face recognizer with high accuracy. It offers a faster face residual network, containing efficient FasterFace blocks and efficient channel spatial attention for improved feature extraction. As a result, the proposed network achieves 97.08% based on average accuracy, outperforming the other networks on five benchmark datasets. It performs faster at 19.91 frames per second in real time on CPU-based hardware when integrated with a face detector, showcasing its capacity to support real-time biometric continuous authentication for smart factory workers. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Muhamad Dwisnanto Putro, Ge Cao, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Local Self-Attention With Mixing Abstract Tokens for Urban Autonomous DrivingabstractAlthough local self-attentions exhibit translation equivariance and locality similar to convolution, the model has limited receptive fields and weak modeling ability. The main reason is that self-attention is computed within nonoverlapped windows. To overcome this issue, common methods need further operations to communicate the information across windows, such as window shifting, and sliding. These operations are memory unfriendly, not well supported, and optimized by modern deep-learning frameworks. Alternatively, this article exchanges information across nonoverlapped windows via efficiently mixing abstract tokens (MAT). The MAT block includes the following steps. First, the image tokens are partitioned into windows and each window is merged with an abstract token. Second, in each window, interactions of image tokens and the abstract token to image tokens are performed. Third, because the abstract token learns abstract information from each corresponding window, mixing all abstract tokens via transformer encoder helps to exchange information between local windows and result in global context modeling. Fourth, the global information of the mixed tokens is propagated back to the image tokens through transformer decoder. The MAT block is efficient and easy to implement, only containing matrix multiplications. In addition, this article also proposes a bilinear patch embedding that samples relevant regions of the input tokens based on learned offsets. Extensive experiments are conducted and evaluated with various tasks such as image classification, object detection, and segmentation. As a result, our method achieves promising performances across tasks. For example, MAT-2 accomplishes79.0%top-1 accuracy on ImageNet-1 K with0.7GFLOPs and outperforms the baseline Swin-0.7 G by4.6%while reducing15.2 mson CPU and0.53 mson GPU devices. The MAT-4 surpasses Swin-T by1.8%mIoU with only70%GFLOPs. Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Ge Cao, Jehwan Choi, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Efficient Vision Transformers with Partial Attention
Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Kang-Hyun Jo |
ECCV (83) | 1 |
| 2024 | EMPCNet: Facial Attribute Recognition Using Efficient Multi - Perspective Convolution for Human-Robot InteractionabstractHuman-robot interaction has evolved into a significant field in robotics. In this domain, facial attributes are essential as they enable robots to understand human emotions, intentions, and preferences. In robot applications, which typically involve low-cost devices, efficient recognition technology is crucial for promising real-time operation by robots. This work proposes EMPCNet to perform facial attribute recognition, consisting of an Efficient Multi-Perspective Convolution (EMPC) block used to efficiently extract and capture various information from multiple perspectives using different kernel sizes and shapes of convolutional operations. The proposed network, which only utilizes a few parameters and low computational operations, achieves competitive performance on the CelebA and LFWA datasets. Additionally, when integrated with face detection, the proposed EMPCNet operates efficiently in real-time on a CPU with Intel Core i7-9750H, achieving a frame rate of 21.27 frames per second (FPS) with an image input size of$224\times 224$consisting of a face area. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, RussoMohammadAshraf Uddin, Kang-Hyun Jo |
HSI | 3 |
| 2024 | Simple Human Fall Surveillance System Based on Person DetectionabstractHuman fall is a common problem that often occurs with the elderly, disabled people, and people with bone diseases and neurological diseases. Sometimes, it also comes from human carelessness. Detecting and warning of human falls can minimize the unfortunate risks. Therefore, human fall detection has been widely applied in medical care and surveillance systems. This paper proposes a simple human fall surveillance system based on a person detection network. This system utilizes the pre-trained YOLOv8 network architecture with a related person body dataset. The proposed system reduces the computational complexity and simplifies the use of available datasets for building a surveillance system. As a result, the proposed system achieves the best speed at 206 Frames per second (FPS) when testing on a GeForce GTX 1080Ti 11GB GPU. Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Duc-Vuong Nguyen, Thi-Le-Hang Nguyen, Kang-Hyun Jo |
IECON | 2 |
| 2024 | Wider Neighborhood-Aware Attention in Improving YOLOv8n for One-Stage Human Fall DetectionabstractHuman fall detection has become a crucial technology in bolstering intelligent surveillance systems. A one-stage human fall detection model based on the YOLO network emerges as an ideal solution for implementation in limited resource environments, supporting real-time operation with faster speed. This work introduces a Wider Neighborhood-Aware Attention (WN2A) module to enhance YOLOv8n performance for one-stage human fall detection on a CPU device. WN2A enables the YOLOv8n network to focus on crucial information within the feature map based on the channel while considering a wider neighborhood area from a spatial point of view. As a result, the proposed WN2A applied on the YOLOv8n network outperforms the other methods based on the mean Average Precision (mAP) of two benchmark datasets. Moreover, the improved YOLOv8n network enables operating at 27.38 frames per second on an Intel Core i7-9750H CPU while providing higher mAP. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Jehwan Choi, Kang-Hyun Jo |
IECON | 3 |
| 2023 | Vehicle Detector Based on Improved YOLOv5 Architecture for Traffic Management and Control SystemsabstractVehicle detection is an important module in traffic management and control systems. These systems require compactness, mobility, and high accuracy when deployed in a real-time context. Based on the YOLOv5 network architecture, this paper proposes several improvements to increase the performance and speed of the network when applied to vehicle detection. The research aims to redesign the backbone and neck modules with lightweight convolutional network architectures such as EfficientNet, PP-LCNet, and MobileNet. In addition, the Squeeze-and-Excitation (SE) attention architecture is also used inside the above-mentioned architectures to help the network focus on salient information during feature extraction. The network is trained and evaluated on a modified and normalized dataset of the UA-DETRAC dataset. As a result, the proposed network achieves 58.1% of [email protected] and 40.1% of [email protected]:0.95 with just over ten million network parameters. This result outperforms other methods and is comparable to the lightweight architectures of the YOLOv5 family. Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Kang-Hyun Jo |
IECON | 2 |
| 2023 | Facial Attribute Recognition Using Lightweight Multi-Label CNN-Transformer Architecture for Intelligent AdvertisingabstractIn modern cities, intelligent advertising platforms have been widely engaged in public areas. A facial attribute recognition technique is essential to assist these platforms in delivering suitable adverts for each audience. These platforms also require a recognition technology that can operate at least suitably on a CPU device to reduce implementation costs. This work proposed a lightweight multi-label CNN-Transformer architecture with an efficient inception block (EIB) and squeeze channel transformer encoder (SCTE) to perform facial attribute recognition efficiently. EIB is used to extract face features in multi-scale and levels supported by SCTE in improving its feature map's quality. The proposed architecture produces fewer parameters with low operations and gains competitive accuracy on the CelebA and LWFA datasets consisting of images with multi-label. Moreover, the proposed architecture integrated with face detection can perform sufficiently on a CPU configuration in real-time with 21 frames per second (FPS) using 224 × 224 input size of face area image. Adri Priadana, Muhamad Dwisnanto Putro, Jinsu An, Duy-Linh Nguyen, Xuan-Thuy Vo, Kang-Hyun Jo |
IECON | 5 |
| 2022 | Multi-level Feature Reweighting and Fusion for Instance SegmentationabstractAccurate instance segmentation requires high-resolution features for performing a dense pixel-wise prediction task. However, using high-resolution feature maps results in highly expensive model complexity and ineffective receptive fields. To overcome the problems of high-resolution features, conventional methods explore multi-level feature fusion that exchanges the information between low-level features at earlier layers and high-level features at top layers. Both low and high information is extracted by the hierarchical backbone network where high-level features contain more semantic cues and low-level features encompass more specific patterns. Thus, adopting these features to the training segmentation model is necessary, and designing a more efficient multi-level feature fusion is crucial. Existing methods balance such information by using top-down and bottom-up pathway connections with more inefficient convolution layers to produce richer multi-scale features. In this work, we contribute two folds: (1) a simple but effective multilevel feature reweighting layer is proposed to strengthen deep high-level features based on channel reweighting generated from multiple features of the backbone, and (2) an efficient fusion block is proposed to process low-resolution features in a depth-to-spatial manner and combine enhanced multi-level features together. These designs enable the segmentation models to predict instance kernels for mask generation on high-level feature maps. To verify the effectiveness of the proposed method, we conduct experiments on the challenging benchmark dataset MS-COCO. Surprisingly, our simple network outperforms the baseline in both accuracy and inference speed. More specifically, we achieve 35.4% APmaskat 19.5 FPS on a GPU device, becoming a state-of-the-art instance segmentation method. Xuan-Thuy Vo, Tien-Dat Tran, Duy-Linh Nguyen, Kang-Hyun Jo |
INDIN | 1 |
| 2022 | A review on anchor assignment and sampling heuristics in deep learning-based object detection
Xuan-Thuy Vo, Kang-Hyun Jo |
Neurocomputing | 1 |
| 2022 | Accurate Bounding Box Prediction for Single-Shot Object DetectionabstractAccurate single-shot object detection is an extremely challenging task in real environments because of complex scenes, occlusion, ambiguities, blur, and shadow, i.e., these factors are called uncertainty problem. It leads to unreliable labeling of bounding box annotation and makes detectors arduous to learn bounding box localization. Previous methods viewed the ground truth box coordinates as a rigid distribution omitting localization uncertainty in real datasets. This article proposes a novel bounding box encoding algorithm integrated into the single-shot detector (BBENet) to consider the flexible distribution of bounding box localization. First, discretized ground truth labels are generated by decomposing each object’s boundary into multiple boundaries. The new representation of ground truth boxes is more arbitrary and flexible to cover any case of complex scenes. During training, the detector directly learns discretized box locations instead of continuous domain. Second, the bounding box encoding algorithm reorganizes bounding box predictions to be more accurate. Furthermore, another problem in existing methods is inconsistency in estimating detection quality. The single-shot detection consists of classification and localization tasks, but the popular detectors consider the classification score as the final detection quality. Thus, it lacks localization quality and hinders the overall performance because both tasks have a positive correlation. To overcome this problem, BBENet introduces detection quality by combining the localization and classification quality to rank detection during nonmaximum suppression. The localization quality is computed based on how uncertain the predicted boxes are, which is a new perspective in detection literature. The proposed BBENet is evaluated on three benchmark datasets, i.e., MS-COCO, Pascal VOC, and CrowdHuman. Without bells and whistles, BBENet outperforms the existing methods by a large margin with comparable speed, achieving the state-of-the-art single-shot detector. Xuan-Thuy Vo, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Regression-Aware Classification Feature for Pedestrian Detection and Tracking in Video Surveillance Systems
Xuan-Thuy Vo, Tien-Dat Tran, Duy-Linh Nguyen, Kang-Hyun Jo |
ICIC (1) | 1 |
| 2021 | Light-weight Convolutional Neural Network for Distracted Driver ClassificationabstractDriving is an activity that requires the coordination of many senses with complex manipulations. However, the driver can be affected by a several factors such as using a mobile phone, adjusting audio equipment, smoking, drinking, eating, talking to a passenger or drowsy. Therefore, the development of assistant applications to warn distracted driver is very necessary. Because of the limited space and mobility, the equipment also requires compact, energy-saving and efficient. This paper proposes a lightweight Convolutional Neural Network for a distracted driver warning system. The method is built based on a combination of standard convolution and Depthwise Separable Convolution operation to optimize the network parameters but still ensure the important information and speed. The network was trained and evaluated on two datasets, AUC (the American University in Cairo) and StateFarm dataset from Kaggle’s competition. As a result, the evaluation accuracy reached 95.36% and 99.95%, respectively. Duy-Linh Nguyen, Muhamad Dwisnanto Putro, Xuan-Thuy Vo, Kang-Hyun Jo |
IECON | 3 |
| 2021 | Dynamic Multi-Loss Weighting for Multiple People Tracking in Video Surveillance SystemsabstractMultiple people tracking is a fundamental yet challenging task in the computer vision field, which served as a primary process for high-level tasks such as human behaviors, action recognition, pose estimation. Person tracking is decomposed into detection and re-identification (re-ID) sub-tasks. Conventionally, the detection learns classification and regression objectives simultaneously; and the re-ID sub-task is treated as a classification task. Therefore, person tracking is multiple task learning corresponding to multiple loss functions (multiple objectives) with one bounding box regression and two classifications. The difference between various tasks is as follows: the ranges of each objective are inconsistent, the contribution of each task to the overall gradient is altered, and the learning pace of each task is different (level of difficulty). It leads to an objective imbalance in multi-task learning. Previous methods proposed weighting factors as new hyper-parameters to balance the ranges of each task. The dimension of search space for manually tuning these hyper-parameters is high, which depends on the number of tasks. Accordingly, selecting reasonable weighting factors is difficult and complicated. This paper introduces dynamic multi-loss weighting (DMW) with simple but effective in which the weighting factors are dynamically changed during training without introducing any hyper-parameters. The dynamic weights are optimized to balance regression and classification objectives, which depend on the difficulty level of each task and the correlation between each task. Additionally, the general convolution operations are spatially invariant to some degree, which hinders the network’s performance. Hence, this work employs the position-sensitive operation improving feature extraction. The proposed method is conducted on the MOT17 challenging benchmark, which outperforms the online multiple people trackers without using additional data. Xuan-Thuy Vo, Tien-Dat Tran, Duy-Linh Nguyen, Kang-Hyun Jo |
INDIN | 1 |
| 2020 | Enhanced Feature Pyramid Networks by Feature Aggregation Module and Refinement ModuleabstractFeature pyramids executing refinements on the raw feature maps produced by the backbone (e.g., ResNet, VGG) are universally employed in object detection tasks (e.g., Faster R-CNN, Mask R-CNN, YOLO, SSD, RetinaNet) to mitigate scale variation problem. Although these object detections with feature pyramids accomplish a boost in accuracy without compromising speed, they have some limitations since that they only naturally design the feature pyramid with consecutive scales, the pyramidal architecture of the backbone, which are initially constructed for the classification task. This problem leads to the feature imbalance between high-level features and low-level features in object detection. In this work, the proposed method introduces Feature Aggregation Module (FAM) and Refinement Module (RM) to obtain more powerful feature pyramids for predicting objects of different scales. First, the multi-level feature maps (i.e., multiple layers) extracted by the backbone network are aggregated as the basic feature. Second, the basic feature is enhanced by a refinement module exploiting long-range dependency. Three, to construct a feature pyramid for object detection, the proposed FAM is used by converting the basic feature (after utilizing a refinement module) into multi-level features. Finally, refined multi-level features and raw features generated by the backbone could be enhanced through shortcut connections to capture more representative. To perform the efficiency, the proposed method integrates the FAM and the RAM into the architecture of Faster R-CNN called EFPN Faster R-CNN. Especially on the MS-COCO dataset, EFPN Faster R-CNN achieves 2.2 points higher Average Precision (AP) than FPN Faster R-CNN without bells and whistles. Xuan-Thuy Vo, Kang-Hyun Jo |
HSI | 1 |
| 2020 | Bidirectional Non-local Networks for Object Detection
Xuan-Thuy Vo, Li-Hua Wen, Tien-Dat Tran, Kang-Hyun Jo |
ICCCI | 1 |