VLDB 2026 Research / reviewers in the wild / expert
Jehwan Choi
dblp:306/3535
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2025
0009-0005-8494-2170ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Waste Detection on Low-Light EnvironmentabstractOn waste-sorting conveyor lines, illumination often fluctuates sharply or drops to very low levels, causing conventional object detectors to fail. To achieve lighting-robust performance without extra lamps, vision rooms, or site-specific retraining, we propose an Illumination-Invariant Convolution (IIC) block that can be inserted into any backbone network. Working in log-intensity space under a Lambertian model, IIC applies learnable zero-mean cross-channel filters that mute lighting artefacts and boost material cues, and then merges the resulting maps with base features to produce lighting-robust representations. We integrate IIC into the “nano” versions of YOLOv5/8/10/11 and lightweight RT-DETR, training on roughly 200 k conveyor-belt images (29 classes) from the AI-Hub waste dataset. IIC raises YOLO mAP50 by 2.5-4.5 pp and mAP50-95 by up to 4.6 pp; even the data-hungry, transformer-based RT-DETR gains up to 1.5 pp. In low-light video tests, the IIC-augmented models successfully detected objects the original networks missed, raising recall. This demonstrates that inserting the IIC module at a network's input provides a straightforward path to illumination-robust, field-ready waste-sorting systems. Jehwan Choi, Minseung Kim, Kang-Hyun Jo |
HSI | 1 |
| 2025 | Artificial Behavior Intelligence: Technology, Challenges, and Future DirectionsabstractUnderstanding and predicting human behavior has emerged as a core capability in various AI application domains such as autonomous driving, smart healthcare, surveillance systems, and social robotics. This paper defines the technical frame-work of Artificial Behavior Intelligence (ABI), which comprehensively analyzes and interprets human posture, facial expressions, emotions, behavioral sequences, and contextual cues. It details the essential components of ABI, including pose estimation, face and emotion recognition, sequential behavior analysis, and context-aware modeling. Furthermore, we highlight the transformative potential of recent advances in large-scale pretrained models, such as large language models (LLMs), vision foundation models, and multimodal integration models, in significantly improving the accuracy and interpretability of behavior recognition. Our research team has a strong interest in the ABI domain and is actively conducting research, particularly focusing on the development of intelligent lightweight models capable of efficiently inferring complex human behaviors. This paper identifies several technical challenges that must be addressed to deploy ABI in real-world applications including learning behavioral intelligence from limited data, quantifying uncertainty in complex behavior prediction, and optimizing model structures for low-power, real-time inference. To tackle these challenges, our team is exploring various optimization strategies including lightweight transformers, graph-based recognition architectures, energy-aware loss functions, and multimodal knowledge distillation, while validating their applicability in real-time environments. Kang-Hyun Jo, Jehwan Choi, Kwanho Kim, Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Tien-Dat Tran |
HSI | 2 |
| 2025 | Waste Object Detection Using Bright to Dark Feature AlignmentabstractFactory environments vary significantly in lighting and camera conditions, necessitating models that are robust to illumination changes. Although prior studies have addressed this issue by constructing separate low-light datasets, such approaches face scalability challenges due to the high cost of data collection. To overcome this issue, the proposed approach leverages DARK-ISP(Low-light Image Synthesis Pipeline) from the prior work DAI-Net to construct a synthetic low-light image dataset. This approach eliminates the need to collect real dark data. Retraining a model from scratch to handle low-light conditions is not cost-effective. A more practical approach is to leverage high-performance models pretrained in bright industrial environments. While fine-tuning such models is a possible solution, it often suffers from performance degradation due to domain shift. This study adopts a teacher-student framework to perform back-bone feature-level domain alignment between a teacher model trained on well-lit images and a student model trained on low-light images. The alignment is achieved using MMD(Maximum Mean Discrepancy) loss, which effectively mitigates the domain shift problem and reduces the representational gap between the two models. The detection model is based on RT-DETRv2, a lightweight ViT-based architecture that achieves both real-time performance and high object detection accuracy. Building upon the aligned features, a knowledge distillation method based on KD-DETR was applied. This method, specifically tailored for DETR architectures, further enhanced detection performance under low-light conditions. Experiments conducted on ROBOne recyclable waste dataset show that the proposed method achieves a 6.4% higher mAP compared to DAI-Net and a 2.2% improvement over simple fine-tuning on the target domain. Minseung Kim, Jehwan Choi, Jongchae Lee, Kyubin Hwang, Kang-Hyun Jo |
HSI | 2 |
| 2025 | Efficient Human Behavior Detector for Vision-based Emergency Evacuation SystemsabstractThe emergency evacuation systems are often installed in crowded places such as airports, train stations, and shopping malls to evacuate and protect people when incidents occur quickly. With the development of surveillance cameras, vision-based emergency evacuation systems have demonstrated their ability to observe and promptly warn flexibly. This paper proposes a human behavior detector by fine-tuning the YOLOv11n detection network with the Global Attention Mechanism (GAM) to enhance the individual human action recognition. Extensive experiments are trained and evaluated on the Human Behavior Detection Dataset (HBDset) using a NVIDIA Tesla V100 32GB GPU. The proposed network achieves 62.2% of mAP and an inference speed of 1.3 milliseconds (ms), and outperforms other networks of the same scale. Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Jehwan Choi, Kang-Hyun Jo |
HSI | 4 |
| 2025 | Efficiency-Accuracy Trade-Off of Facial Attribute Classifier Supporting Human-Robot InteractionabstractThe advancement of robotics has been driven by the integration of artificial intelligence, machine learning, and sophisticated sensing technologies, enabling more seamless Human-Robot Interaction (HRI). Facial Attribute Classifier (FAC) plays a crucial role in HRI by helping robots understand human emotions, intentions, and social cues, fostering personalized and intuitive interactions. However, while existing methods achieve high accuracy, their computational complexity limits real-time applications on low-cost or CPU-based devices, highlighting the need for lightweight models that balance accuracy and efficiency. This work proposes an Efficient Network (ENet) designed to achieve an optimal trade-off between efficiency and accuracy of FAC. ENet introduces an Enhanced Sequential Efficient Attention Module (ESEAM) to improve the quality of feature maps while maintaining high efficiency. Accordingly, ENet demonstrates a compromise between efficiency and accuracy on the CelebA and LFWA datasets. The proposed ENet is computationally efficient, generating a few parameters, making it well-suited for CPU-based applications. When combined with a face detector, the optimized FAC achieves a processing speed of 25.88 frames per second (FPS) on an Intel Core i7-9750H CPU, demonstrating its suitability for real-time use. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Ge Cao, Jehwan Choi, Kang-Hyun Jo |
HSI | 5 |
| 2025 | Efficient Multi-Scale Spatial Interactions for Visual Recognition TasksabstractConvolution operation has local connectivity and translation equivalence while self-attention operation captures long-range spatial dependencies. Adopting the merits of convolution and self-attention operations in hierarchical networks can result in better visual representation and generalization performance. However, integrating self-attention layers into earlier stages is inefficient because self-attention operation has quadratic complexity with token lengths. In this work, we tackle this issue and propose an Efficient Multi-scale Spatial interaction Network (EMSNet) that takes advantage of hybrid networks. The EMSNet has key insights: (1) Each stage efficiently models both short-range and long-range spatial interactions via the design of the multi-scale tokens; (2) The novel convolution-based multi-head self-attention (C-MHSA) operation is introduced to learn spatial interactions inside local regions; (3) The efficient combination of the depthwise convolution, coordinate depthwise convolution, C-MHSA, and global multi-head self-attention (G-MHSA) are performed via channel splitting strategy, extracting wide ranges of frequencies and multi-order interactions. Extensive experiments on ImageNet-1K image classification, MS-COCO object detection, and segmentation tasks verify the effectiveness and generalization ability of the EMSNet. For instance, the EMS Net-XTiny gets 77.1% Top-1 accuracy on ImageNet-1K which is much greater than PVTvl-Tiny by 2% with only 22% parameters and 37% GFLOPs. Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Jehwan Choi, Kang-Hyun Jo |
HSI | 4 |
| 2025 | Local Self-Attention With Mixing Abstract Tokens for Urban Autonomous DrivingabstractAlthough local self-attentions exhibit translation equivariance and locality similar to convolution, the model has limited receptive fields and weak modeling ability. The main reason is that self-attention is computed within nonoverlapped windows. To overcome this issue, common methods need further operations to communicate the information across windows, such as window shifting, and sliding. These operations are memory unfriendly, not well supported, and optimized by modern deep-learning frameworks. Alternatively, this article exchanges information across nonoverlapped windows via efficiently mixing abstract tokens (MAT). The MAT block includes the following steps. First, the image tokens are partitioned into windows and each window is merged with an abstract token. Second, in each window, interactions of image tokens and the abstract token to image tokens are performed. Third, because the abstract token learns abstract information from each corresponding window, mixing all abstract tokens via transformer encoder helps to exchange information between local windows and result in global context modeling. Fourth, the global information of the mixed tokens is propagated back to the image tokens through transformer decoder. The MAT block is efficient and easy to implement, only containing matrix multiplications. In addition, this article also proposes a bilinear patch embedding that samples relevant regions of the input tokens based on learned offsets. Extensive experiments are conducted and evaluated with various tasks such as image classification, object detection, and segmentation. As a result, our method achieves promising performances across tasks. For example, MAT-2 accomplishes79.0%top-1 accuracy on ImageNet-1 K with0.7GFLOPs and outperforms the baseline Swin-0.7 G by4.6%while reducing15.2 mson CPU and0.53 mson GPU devices. The MAT-4 surpasses Swin-T by1.8%mIoU with only70%GFLOPs. Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Ge Cao, Jehwan Choi, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Vehicle Movement Status Network on Drone-Perspective View with Adaptive Adversarial LearningabstractThe rapidly developing autonomous driving field now needs a more secure transportation system through information between multiple mobility. Deep learning that can judge traffic conditions by convergence of various sensor data and in particular, research on the convolutional neural network using computer vision are being actively conducted. In addition, recognizing many objects at once in a large area through drone images and understanding the movement of the object is used as safe traffic assistance information. In this study, an image classification study is conducted to determine the status of the vehicle on the road through drone flight image data. The goal is to build a new image classification model robust to the proposed image classification network by applying the weighted adversarial learning method. Weight adversarial learning is a method of securing robust performance in image classification of various statuses while disturbing the model by forcibly reflecting the slope value in reverse when updating the network through the reverse gradient layer. In the experiment, model performance is evaluated through the collected drone flight data set. Youlkyeong Lee, Jehwan Choi, Kang-Hyun Jo |
IECON | 2 |
| 2024 | Wider Neighborhood-Aware Attention in Improving YOLOv8n for One-Stage Human Fall DetectionabstractHuman fall detection has become a crucial technology in bolstering intelligent surveillance systems. A one-stage human fall detection model based on the YOLO network emerges as an ideal solution for implementation in limited resource environments, supporting real-time operation with faster speed. This work introduces a Wider Neighborhood-Aware Attention (WN2A) module to enhance YOLOv8n performance for one-stage human fall detection on a CPU device. WN2A enables the YOLOv8n network to focus on crucial information within the feature map based on the channel while considering a wider neighborhood area from a spatial point of view. As a result, the proposed WN2A applied on the YOLOv8n network outperforms the other methods based on the mean Average Precision (mAP) of two benchmark datasets. Moreover, the improved YOLOv8n network enables operating at 27.38 frames per second on an Intel Core i7-9750H CPU while providing higher mAP. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Jehwan Choi, Kang-Hyun Jo |
IECON | 4 |
| 2023 | VSNet: Vehicle State Classification for Drone Image with Mosaic Augmentation and Soft-Label Assignment
Youlkyeong Lee, Jehwan Choi, Kang-Hyun Jo |
ACIIDS (1) | 2 |
| 2023 | CSA: Channel-Wise Similarity Attention for Vehicle State ClassificationabstractDeveloped for specific missions, CNNs have gradually improved the performance of object classification networks by using various architectures. The weight of the convolutional layer is a crucial factor in feature extraction. However, as the number of layers increases, performance degradation can occur due to problems such as the vanishing gradient. To overcome this problem, networks have evolved to continuously incorporate information from previous feature maps using various attention mechanisms. In this study, a Channel-wise Similarity Attention (CSA) method is proposed to measure the similarity of feature maps between channels and enhance positive information by highlighting it. Additionally, a deformable convolutional kernel is embedded to apply a flexible receptive field around the object area in the image, replacing the fixed receptive field of the conventional CNN layer. The network is trained end-to-end to classify the condition of vehicles on the road using collected drone flight images. The proposed model achieves an accuracy of 86.13% and 302 frames per second with a number of parameters of 1,273,504. Youlkyeong Lee, Jehwan Choi, Jinsu An, Kang-Hyun Jo |
IECON | 2 |
| 2022 | Low Computational Vehicle Re-Identification for Unlabeled Drone Flight ImagesabstractRecently advanced vehicle re-identification frameworks are mainly based on convolutional neural networks (CNN) and labeled information. Previous frameworks face two difficulties. First CNN includes complicated architectures, which require expensive GPU devices to perform computation. The second difficulty is that annotating vehicle identities for every frame is expensive and time-consuming. To tackle these two difficulties, this study proposes a simple but effective method to perform re-ID without CNN and labeled identities. The proposed method has two streams of vehicle re-identification. The object detector takes charge of detecting vehicles on the road. With the position of vehicles in the image, the condition module extracts the vehicle movement information and sets the condition to match the same vehicle between current and subsequent frames. To train the object detector and test the proposed algorithm, a set of drone flight images collect and annotate for studying the traffic road. It contains 9,776 train images and 2,200 test images for object detection. In the experiments, three different traffic video clips were applied for testing the proposed method. Youlkyeong Lee, Qing Tang 0004, Jehwan Choi, Kang-Hyun Jo |
IECON | 3 |
| 2021 | Attention based Object Classification for Drone ImageryabstractThis paper shows how to make the drone imagery for surveillance or tracking the object in the ground. To detect or classify objects on the ground, convolutional neural networks was adopted and compared with some existed methods and the proposed attention blocks in it. The objects on the ground from the drone images are relatively very small and diversity of the appearance from its perspective projections. This is mainly due to the arbitrary viewpoints from the bird eye views. Furthermore, the distance from its viewpoint in the sky is quite much changeable so that the image of the object is too diverse in appearance and its size. However, the drone is so useful to see widely while navigating in the sky. It is much more attentive to use for real application. Here, some proposed target objects are mainly located in the ground, like static and dynamic objects such as street lamps or trees, vehicles, trucks and pedestrians. These works were done for the national projects to establish the general AI services in Korea recently. For the experiments such as buildup the ground truth of target objects after taken in regulated distance and viewing angles and performed to detect exactly objects in an arbitrary image. For experiments of detection and classification of five categories of objects, attention based CNN architecture was adopted and compared comprehensively with the existed networks like MobileNet, VGG16, SqueezeNet, and ResNet. The experimental results outperformed for the archived drone image dataset with 87.12% in precision. The architecture shows almost 3 times faster with respect to VGG16 or 2 times faster than MobileNet in the speed but a half slimer and twice thicker respectively in the number of parameters. Thus, the Attention Block is useful while a drone navigates through a certain route according to the ground location regardless of the appearance and size of the target region in image. Jehwan Choi, Kang-Hyun Jo |
IECON | 1 |