VLDB 2026 Research / reviewers in the wild / expert
Adri Priadana
dblp:325/0036
· DBLP profile ↗
17ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-1553-7631ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimal Proxy Mining Contrastive Network for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) performance enhancement hinges on extracting the most informative features from unlabeled person datasets. In recent approaches, proxy-based contrastive learning with awareness of camera labels has been adopted for model training, thereby achieving highly promising results. However, inappropriate selections of contrastive pairs can significantly degrade the performance of these models. To address this issue, we propose the Optimal Proxy Mining Contrastive Network (OPMCN), a novel framework designed to strategically optimize the selection of proxies for positive and negative pair formation, thus enhancing the efficacy of contrastive training. The OPMCN framework proposes two specific contrastive losses: Hardest Camera Proxy Mining (HCPM) and False Negative Proxies Mining (FNPM), each essential for enhancing model performance in unsupervised settings. The HCPM loss targets proxies from the most challenging cameras to maximize semantic differences between pairs while ensuring minimal background shifts. In contrast, the FNPM loss counters noise in pseudo labels by prioritizing similarity rankings over clustering results to effectively identify and correct false negatives among proxies. Moreover, we have developed the Pyramid Kernel Global Context (PKGC) block, which employs an attention mechanism that focuses on identity-invariant semantic cues in instances. This module utilizes optimally sized convolutional kernels to enhance identity recognition consistency across camera-based variations, thereby improving the precision of feature extraction. Experimental results on several popular datasets prove that our work surpasses existing unsupervised person Re-ID approaches to a remarkable extent. Ge Cao, Qing Tang 0004, Xuan-Thuy Vo, Adri Priadana, Kang-Hyun Jo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Large Vision-Language Models with PEFT for Generating Descriptive Annotations in Person Re-IdentificationabstractGenerating descriptive annotations for person re-identification (Re-ID) images is essential for bridging vision and language domain, improving both interpretability and cross-modal retrieval performance. However, large vision-language models (LVLMs), which trained on broad web-scale corpora, often struggle to generate accurate, context-relevant descriptions for ReID samples due to inherent challenges such as occlusions, low resolution, varying illumination, and diverse viewpoints. In this paper, we propose to apply Parameter-Efficient Fine-Tuning (PEFT) via Low-Rank Adaptation (LoRA) to tune Qwen2-VL for ReID-specific captioning tasks. Leveraging existing Re-ID datasets with paired image-text annotations, our fine-tuned model generates domain-aligned and discriminative captions. Experiments show significant improvements in caption relevance and identity descriptiveness, highlighting the potential of PEFT-tuned LVLMs for real-world ReID applications. Ge Cao, Qing Tang 0004, Adri Priadana, Tran Tien Dat, Ashraf Uddin Russo, Kang-Hyun Jo |
HSI | 3 |
| 2025 | Artificial Behavior Intelligence: Technology, Challenges, and Future DirectionsabstractUnderstanding and predicting human behavior has emerged as a core capability in various AI application domains such as autonomous driving, smart healthcare, surveillance systems, and social robotics. This paper defines the technical frame-work of Artificial Behavior Intelligence (ABI), which comprehensively analyzes and interprets human posture, facial expressions, emotions, behavioral sequences, and contextual cues. It details the essential components of ABI, including pose estimation, face and emotion recognition, sequential behavior analysis, and context-aware modeling. Furthermore, we highlight the transformative potential of recent advances in large-scale pretrained models, such as large language models (LLMs), vision foundation models, and multimodal integration models, in significantly improving the accuracy and interpretability of behavior recognition. Our research team has a strong interest in the ABI domain and is actively conducting research, particularly focusing on the development of intelligent lightweight models capable of efficiently inferring complex human behaviors. This paper identifies several technical challenges that must be addressed to deploy ABI in real-world applications including learning behavioral intelligence from limited data, quantifying uncertainty in complex behavior prediction, and optimizing model structures for low-power, real-time inference. To tackle these challenges, our team is exploring various optimization strategies including lightweight transformers, graph-based recognition architectures, energy-aware loss functions, and multimodal knowledge distillation, while validating their applicability in real-time environments. Kang-Hyun Jo, Jehwan Choi, Kwanho Kim, Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Tien-Dat Tran |
HSI | 7 |
| 2025 | Efficient Human Behavior Detector for Vision-based Emergency Evacuation SystemsabstractThe emergency evacuation systems are often installed in crowded places such as airports, train stations, and shopping malls to evacuate and protect people when incidents occur quickly. With the development of surveillance cameras, vision-based emergency evacuation systems have demonstrated their ability to observe and promptly warn flexibly. This paper proposes a human behavior detector by fine-tuning the YOLOv11n detection network with the Global Attention Mechanism (GAM) to enhance the individual human action recognition. Extensive experiments are trained and evaluated on the Human Behavior Detection Dataset (HBDset) using a NVIDIA Tesla V100 32GB GPU. The proposed network achieves 62.2% of mAP and an inference speed of 1.3 milliseconds (ms), and outperforms other networks of the same scale. Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Jehwan Choi, Kang-Hyun Jo |
HSI | 3 |
| 2025 | A Lightweight CNN-Based Framework for Infant Emotional State Recognition in Real-Time ScenariosabstractRecognizing infant emotion is challenging, mainly due to the lack of facial expression data specifically focused on infants. Most publicly available datasets are developed for the general public or adults, which may not accurately capture the distinct facial characteristics of infants. This paper presents the initial stages of developing a deep learning-based framework for infant emotion recognition, utilizing a state-of-the-art and lightweight CNN backbone. To construct a suitable dataset, this paper filters infant faces from an existing dataset using age-related characteristics to create a more focused training set. The curation dataset is then used to refine the current CNN architecture selection. During testing, facial regions are detected in real-time from input videos to localize areas of interest before classification. This approach aims to improve the reliability of infant emotion recognition while maintaining efficiency suitable for real-time applications. This paper describes the dataset preparation process, model evaluation, and the design of a real-time workflow. Future work will explore improvements in data quality, model performance, and comprehensive application scenarios. Rahmatullah Arrizal Pranatadesta, Adri Priadana, Kang-Hyun Jo |
HSI | 2 |
| 2025 | Efficiency-Accuracy Trade-Off of Facial Attribute Classifier Supporting Human-Robot InteractionabstractThe advancement of robotics has been driven by the integration of artificial intelligence, machine learning, and sophisticated sensing technologies, enabling more seamless Human-Robot Interaction (HRI). Facial Attribute Classifier (FAC) plays a crucial role in HRI by helping robots understand human emotions, intentions, and social cues, fostering personalized and intuitive interactions. However, while existing methods achieve high accuracy, their computational complexity limits real-time applications on low-cost or CPU-based devices, highlighting the need for lightweight models that balance accuracy and efficiency. This work proposes an Efficient Network (ENet) designed to achieve an optimal trade-off between efficiency and accuracy of FAC. ENet introduces an Enhanced Sequential Efficient Attention Module (ESEAM) to improve the quality of feature maps while maintaining high efficiency. Accordingly, ENet demonstrates a compromise between efficiency and accuracy on the CelebA and LFWA datasets. The proposed ENet is computationally efficient, generating a few parameters, making it well-suited for CPU-based applications. When combined with a face detector, the optimized FAC achieves a processing speed of 25.88 frames per second (FPS) on an Intel Core i7-9750H CPU, demonstrating its suitability for real-time use. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Ge Cao, Jehwan Choi, Kang-Hyun Jo |
HSI | 1 |
| 2025 | Efficient Multi-Scale Spatial Interactions for Visual Recognition TasksabstractConvolution operation has local connectivity and translation equivalence while self-attention operation captures long-range spatial dependencies. Adopting the merits of convolution and self-attention operations in hierarchical networks can result in better visual representation and generalization performance. However, integrating self-attention layers into earlier stages is inefficient because self-attention operation has quadratic complexity with token lengths. In this work, we tackle this issue and propose an Efficient Multi-scale Spatial interaction Network (EMSNet) that takes advantage of hybrid networks. The EMSNet has key insights: (1) Each stage efficiently models both short-range and long-range spatial interactions via the design of the multi-scale tokens; (2) The novel convolution-based multi-head self-attention (C-MHSA) operation is introduced to learn spatial interactions inside local regions; (3) The efficient combination of the depthwise convolution, coordinate depthwise convolution, C-MHSA, and global multi-head self-attention (G-MHSA) are performed via channel splitting strategy, extracting wide ranges of frequencies and multi-order interactions. Extensive experiments on ImageNet-1K image classification, MS-COCO object detection, and segmentation tasks verify the effectiveness and generalization ability of the EMSNet. For instance, the EMS Net-XTiny gets 77.1% Top-1 accuracy on ImageNet-1K which is much greater than PVTvl-Tiny by 2% with only 22% parameters and 37% GFLOPs. Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Jehwan Choi, Kang-Hyun Jo |
HSI | 3 |
| 2025 | A High-Accuracy and Faster Face Recognizer Supporting Biometric Continuous Authentication for Smart Factory WorkersabstractSmart factories require secure and sustainable worker authentication for safe operations. Biometric continuous authentication based on facial recognition is one of the most convenient mechanisms. This method applies a face recognition task to verify the captured face as an authorized user. However, existing methods that employ large networks for high-accuracy face recognition incur high computational costs and slow down the process, rendering them unsuitable for continuous operation. This work proposes an efficient and rapid face recognizer with high accuracy. It offers a faster face residual network, containing efficient FasterFace blocks and efficient channel spatial attention for improved feature extraction. As a result, the proposed network achieves 97.08% based on average accuracy, outperforming the other networks on five benchmark datasets. It performs faster at 19.91 frames per second in real time on CPU-based hardware when integrated with a face detector, showcasing its capacity to support real-time biometric continuous authentication for smart factory workers. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Muhamad Dwisnanto Putro, Ge Cao, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Local Self-Attention With Mixing Abstract Tokens for Urban Autonomous DrivingabstractAlthough local self-attentions exhibit translation equivariance and locality similar to convolution, the model has limited receptive fields and weak modeling ability. The main reason is that self-attention is computed within nonoverlapped windows. To overcome this issue, common methods need further operations to communicate the information across windows, such as window shifting, and sliding. These operations are memory unfriendly, not well supported, and optimized by modern deep-learning frameworks. Alternatively, this article exchanges information across nonoverlapped windows via efficiently mixing abstract tokens (MAT). The MAT block includes the following steps. First, the image tokens are partitioned into windows and each window is merged with an abstract token. Second, in each window, interactions of image tokens and the abstract token to image tokens are performed. Third, because the abstract token learns abstract information from each corresponding window, mixing all abstract tokens via transformer encoder helps to exchange information between local windows and result in global context modeling. Fourth, the global information of the mixed tokens is propagated back to the image tokens through transformer decoder. The MAT block is efficient and easy to implement, only containing matrix multiplications. In addition, this article also proposes a bilinear patch embedding that samples relevant regions of the input tokens based on learned offsets. Extensive experiments are conducted and evaluated with various tasks such as image classification, object detection, and segmentation. As a result, our method achieves promising performances across tasks. For example, MAT-2 accomplishes79.0%top-1 accuracy on ImageNet-1 K with0.7GFLOPs and outperforms the baseline Swin-0.7 G by4.6%while reducing15.2 mson CPU and0.53 mson GPU devices. The MAT-4 surpasses Swin-T by1.8%mIoU with only70%GFLOPs. Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Ge Cao, Jehwan Choi, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Efficient Vision Transformers with Partial Attention
Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Kang-Hyun Jo |
ECCV (83) | 3 |
| 2024 | Inverted Residual Bottlenecks with Large Kernel Attention for Remote Scene ClassificationabstractRemote sensing image classification plays a piv-otal role in environmental monitoring and urban planning, yet it faces the challenge of accurately interpreting complex and high-resolution images with fast inference speed for real time applications. To address this, we introduce the Mobile Large Kernel Attention Network (MLKANet), which integrates MobileNetV2's inverted residual structures with the large kernel attention mechanism from the Visual Attention Network (VAN). Our proposed MLKANet achieves a compelling balance of computational efficiency and sophisticated feature extraction, while maintaining the speed from the MobileNetV2 baseline. This study evaluates MLKANet's performance against state-of-the-art models using the Aerial Image Dataset (AID), demonstrating superior accuracy and efficiency. The architecture's effectiveness is further evidenced through an ablation study highlighting the scalability of our approach and class-wise performance analysis that showcases MLKANet's proficiency across various scene types. RussoMohammadAshraf Uddin, Adri Priadana, Ge Cao, Kang-Hyun Jo |
HSI | 2 |
| 2024 | EMPCNet: Facial Attribute Recognition Using Efficient Multi - Perspective Convolution for Human-Robot InteractionabstractHuman-robot interaction has evolved into a significant field in robotics. In this domain, facial attributes are essential as they enable robots to understand human emotions, intentions, and preferences. In robot applications, which typically involve low-cost devices, efficient recognition technology is crucial for promising real-time operation by robots. This work proposes EMPCNet to perform facial attribute recognition, consisting of an Efficient Multi-Perspective Convolution (EMPC) block used to efficiently extract and capture various information from multiple perspectives using different kernel sizes and shapes of convolutional operations. The proposed network, which only utilizes a few parameters and low computational operations, achieves competitive performance on the CelebA and LFWA datasets. Additionally, when integrated with face detection, the proposed EMPCNet operates efficiently in real-time on a CPU with Intel Core i7-9750H, achieving a frame rate of 21.27 frames per second (FPS) with an image input size of$224\times 224$consisting of a face area. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, RussoMohammadAshraf Uddin, Kang-Hyun Jo |
HSI | 1 |
| 2024 | Simple Human Fall Surveillance System Based on Person DetectionabstractHuman fall is a common problem that often occurs with the elderly, disabled people, and people with bone diseases and neurological diseases. Sometimes, it also comes from human carelessness. Detecting and warning of human falls can minimize the unfortunate risks. Therefore, human fall detection has been widely applied in medical care and surveillance systems. This paper proposes a simple human fall surveillance system based on a person detection network. This system utilizes the pre-trained YOLOv8 network architecture with a related person body dataset. The proposed system reduces the computational complexity and simplifies the use of available datasets for building a surveillance system. As a result, the proposed system achieves the best speed at 206 Frames per second (FPS) when testing on a GeForce GTX 1080Ti 11GB GPU. Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Duc-Vuong Nguyen, Thi-Le-Hang Nguyen, Kang-Hyun Jo |
IECON | 3 |
| 2024 | Wider Neighborhood-Aware Attention in Improving YOLOv8n for One-Stage Human Fall DetectionabstractHuman fall detection has become a crucial technology in bolstering intelligent surveillance systems. A one-stage human fall detection model based on the YOLO network emerges as an ideal solution for implementation in limited resource environments, supporting real-time operation with faster speed. This work introduces a Wider Neighborhood-Aware Attention (WN2A) module to enhance YOLOv8n performance for one-stage human fall detection on a CPU device. WN2A enables the YOLOv8n network to focus on crucial information within the feature map based on the channel while considering a wider neighborhood area from a spatial point of view. As a result, the proposed WN2A applied on the YOLOv8n network outperforms the other methods based on the mean Average Precision (mAP) of two benchmark datasets. Moreover, the improved YOLOv8n network enables operating at 27.38 frames per second on an Intel Core i7-9750H CPU while providing higher mAP. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Jehwan Choi, Kang-Hyun Jo |
IECON | 1 |
| 2023 | YOLOv5 with Combination of Coordinate Attention and CBAM for Object Detection on DroneabstractObject detection is an important study in computer vision to discriminate the position and class of an object in an image. Object detection in drone images is a technology that automatically detects and classifies objects using deep learning algorithms in flight images taken by drones. Object detection using drone images can rescue human life in disaster situations, grasp the situation at the disaster site, and identify the growth status of crops or pests in agriculture. In addition, it can be used in various fields such as infrastructure management, roads and railways, and city planning. A quick calculation is required. Although rapid computation is possible due to recent hardware development, there are many difficulties in using GPUs in industrial settings. In order to utilize drones in industrial sites, an object detection algorithm capable of real-time operation in a low-cost device is required. In this paper, we propose YOLOv5 with the combination of Coordinate Attention and CBAM for Object Detection on Drone for an algorithm capable of real-time operation in a low-cost device. The proposed architecture makes the model lighter by reducing the number of parameters and improves the object detection rate of the model through Coordinate Attention and CBAM. The model is trained using the VisDrone dataset, and the object detection rate, mAP, increased by about 10% to 22.2mAP, and the number of parameters decreased by about 70% to 2,147,589. Jinsu An, Muhamad Dwisnanto Putro, Adri Priadana, Youlkyeong Lee, Junmyeong Kim, Kang-Hyun Jo |
IECON | 3 |
| 2023 | Vehicle Detector Based on Improved YOLOv5 Architecture for Traffic Management and Control SystemsabstractVehicle detection is an important module in traffic management and control systems. These systems require compactness, mobility, and high accuracy when deployed in a real-time context. Based on the YOLOv5 network architecture, this paper proposes several improvements to increase the performance and speed of the network when applied to vehicle detection. The research aims to redesign the backbone and neck modules with lightweight convolutional network architectures such as EfficientNet, PP-LCNet, and MobileNet. In addition, the Squeeze-and-Excitation (SE) attention architecture is also used inside the above-mentioned architectures to help the network focus on salient information during feature extraction. The network is trained and evaluated on a modified and normalized dataset of the UA-DETRAC dataset. As a result, the proposed network achieves 58.1% of [email protected] and 40.1% of [email protected]:0.95 with just over ten million network parameters. This result outperforms other methods and is comparable to the lightweight architectures of the YOLOv5 family. Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Kang-Hyun Jo |
IECON | 3 |
| 2023 | Facial Attribute Recognition Using Lightweight Multi-Label CNN-Transformer Architecture for Intelligent AdvertisingabstractIn modern cities, intelligent advertising platforms have been widely engaged in public areas. A facial attribute recognition technique is essential to assist these platforms in delivering suitable adverts for each audience. These platforms also require a recognition technology that can operate at least suitably on a CPU device to reduce implementation costs. This work proposed a lightweight multi-label CNN-Transformer architecture with an efficient inception block (EIB) and squeeze channel transformer encoder (SCTE) to perform facial attribute recognition efficiently. EIB is used to extract face features in multi-scale and levels supported by SCTE in improving its feature map's quality. The proposed architecture produces fewer parameters with low operations and gains competitive accuracy on the CelebA and LWFA datasets consisting of images with multi-label. Moreover, the proposed architecture integrated with face detection can perform sufficiently on a CPU configuration in real-time with 21 frames per second (FPS) using 224 × 224 input size of face area image. Adri Priadana, Muhamad Dwisnanto Putro, Jinsu An, Duy-Linh Nguyen, Xuan-Thuy Vo, Kang-Hyun Jo |
IECON | 1 |