Kwang-Ju Kim

dblp:225/0417 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0001-8458-4506ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ZeRA: Zero-Reindex Multimodal RAG via Heterogeneous Embedding Alignment for Lightweight Query Encoding
Dasom Ahn, Hye Rim Kim, Sangwon Kim 0004, Kwang-Ju Kim, ByoungChul Ko
ICPR (11)5
2025 Pixel-Level Fire Origin Localization via Digital Twin Mapping for Wildfire Surveillance Framework
abstract
Wildfire monitoring systems play a critical role in minimizing environmental and societal damage. Recent advances in computer vision, particularly deep learning-based fire detection, have enabled more accurate and scalable solutions. However, conventional fire detection methods often struggle with wildfire scenarios due to wide spatial extent, the demand for precise localization, and the urgency of early response. To overcome these challenges, we propose a wildfire monitoring framework capable of pixel-level fire origin localization mapped onto a GPS-calibrated digital twin of mountainous terrain. Our system integrates visual fire detection with terrain-aware 3D projection, enabling accurate mapping of fire origins to real-world coordinates. Experimental results on wildfire datasets demonstrate that our method achieves high accuracy in both early fire detection and precise localization, offering a practical and scalable solution for real-world wildfire monitoring.
Dongyoung Kim, In-Su Jang, Kwang-Ju Kim, Kyoungoh Lee
AVSS4
2025 LAttE: A label-free and multimodal framework for context-aware person re-identification
Dasom Ahn, Sangwon Kim 0004, Kwang-Ju Kim, ByoungChul Ko
Neurocomputing3
2024 EQ-CBM: A Probabilistic Concept Bottleneck with Energy-Based Models and Quantized Vectors
Sangwon Kim 0004, Dasom Ahn, ByoungChul Ko, In-Su Jang, Kwang-Ju Kim
ACCV (7)5
2024 MOVES: Motion-Oriented VidEo Sampling for Natural Language-Based Vehicle Retrieval
abstract
Retrieving the target vehicle through natural language descriptions plays a crucial role in intelligent transportation systems. Existing methods tackle this task by employing models that leverage the correlation between textual and visual representations, such as CLIP. However, these models struggle to capture the temporal characteristics of video data, and researchers enhance temporal understanding performance through various data augmentation and video encoders. Yet, conventional approaches in previous studies often overlook the detailed temporal characteristics of vehicles. To overcome this limitation, we introduce a MOVES: Motion-Oriented VidEo Sampling method to effectively utilize the motion information of the target vehicle. Furthermore, we construct a robust model by implementing a re-ranking algorithm to address a variety of vehicle attributes. As a result, our proposed model achieves state-of-the-art performance on the public vehicle retrieval dataset.
Dongyoung Kim, Kyoungoh Lee, In-Su Jang, Kwang-Ju Kim, Pyong-Kun Kim, Jaejun Yoo 0001
AVSS4
2024 TRET: Two Stream-Based Regionally Enhanced Transformers for Person Re-Identification
abstract
Person Re-IDentification (ReID) is a pivotal method for pedestrian tracking and retrieval. This research is inherently challenged by large changes in intra-class or small changes in inter-class. To address this challenge, many researchers have recently introduced transformer-based models, which have shown excellent results. The primary objective of these models is to generate robust features that effectively distinguish between classes and enable generalization. However, existing methods still suffer from class discrimination due to unnecessary noise, including the background. To overcome this limitation, we propose a novel approach called Two stream-based Regionally Enhanced Transformers (TRET) that focuses on the target to be identified. To concentrate on the target region, the TRET utilizes a structure that leverages the pedestrian mask. Furthermore, the proposed model generalizes well by utilizing Contrastive Language-Image Pretraining as the backbone. Finally, our proposed model achieves state-of-the-art performance on the public datasets.
Kyoungoh Lee, Kwang-Ju Kim, Pyong-Kun Kim, In-Su Jang
ICASSP2
2024 Corrigendum to 'Collaborative multi-modal deep learning and radiomic features for classification of strokes within 6 h' [Expert Systems Appl. (2023), 228, 120473]
Chiho Yoon, Sampa Misra, Kwang-Ju Kim, Chulhong Kim, Bum Joon Kim
Expert Syst. Appl.3
2024 SurgT challenge: Benchmark of soft-tissue trackers for robotic surgery
João Cartucho, Alistair Weld, Samyakh Tukra, Haozheng Xu, Hiroki Matsuzaki, Taiyo Ishikawa, Minjun Kwon, Yongeun Jang, Kwang-Ju Kim, Gwang Lee, Bizhe Bai, Lüder A. Kahrs, Lars Boecking, Simeon Allmendinger, Leopold Müller, Yueming Jin, Sophia Bano, Francisco Vasconcelos 0001, Wolfgang Reiter, Jonas Hajek, Estevão Lima, João L. Vilaça, Sandro F. Queiros, Stamatia Giannarou
Medical Image Anal.9
2023 Collaborative multi-modal deep learning and radiomic features for classification of strokes within 6 h
Chiho Yoon, Sampa Misra, Kwang-Ju Kim, Chulhong Kim, Bum Joon Kim
Expert Syst. Appl.3
2022 REET: Region-Enhanced Transformer for Person Re-Identification
abstract
Person re-identification (ReID) plays a significant role in intelligent surveillance systems. However, it is challenging due to large variations in the intra-class, where the same person is captured in different scenes or cameras. The current person ReID research focuses on creating robust features for class distinction and generalizing neural networks for covering various target domains to address the issue. Recently, after the achievement of vision transformers, the application of transformers has also begun to person ReID studies. The transformer-based methods have improved quantitative performance of person ReID; however, they still suffer from class distinction. Therefore, this paper proposes a novel region-enhanced transformer (REET) to create robust ReID features. Unlike conventional transformer-based approaches, the REET emphasizes the tokens generated by region-level. Our method achieves state-of-the-art results on three public datasets; Market1501, DukeMTMC, and CUHK-03.
Kyoungoh Lee, In-Su Jang, Kwang-Ju Kim, Pyong-Kun Kim
AVSS3
2022 GAITTAKE: Gait Recognition by Temporal Attention and Keypoint-Guided Embedding
abstract
Gait recognition, which refers to the recognition or identification of a person based on their body shape and walking styles, derived from video data captured from a distance, is widely used in crime prevention, forensic identification, and social security. However, to the best of our knowledge, most of the existing methods use appearance, posture and temporal feautures without considering a learned temporal attention mechanism for global and local information fusion. In this paper, we propose a novel gait recognition framework, called Temporal Attention and Keypoint-guided Embedding (GaitTAKE), which effectively fuses temporal-attention-based global and local appearance feature and temporal aggregated human pose feature. Experimental results show that our proposed method achieves a new SOTA in gait recognition with rank-1 accuracy of 98.0% (normal), 97.5% (bag) and 92.2% (coat) on the CASIA-B gait dataset; 90.4% accuracy on the OU-MVLP gait dataset.
Hung-Min Hsu, Yizhou Wang 0005, Cheng-Yen Yang, Jenq-Neng Hwang, Le Uyen Thuc Hoang, Kwang-Ju Kim
ICIP6
2021 ROD2021 Challenge: A Summary for Radar Object Detection Challenge for Autonomous Driving Applications
abstract
The Radar Object Detection 2021 (ROD2021) Challenge, held in the ACM International Conference on Multimedia Retrieval (ICMR) 2021, has been introduced to detect and classify objects purely using an FMCW radar for autonomous driving applications. As a robust sensor to all-weather conditions, radar has rich information hidden in the radio frequencies, which can potentially achieve object detection and classification. This insight will provide a new object perception solution for an autonomous vehicle even in adverse driving scenarios. The ROD2021 Challenge is the first public benchmark focusing on this topic, which attracts great attention and participation. There are more than 260 participants among 37 teams from more than 10 countries with different academic and industrial affiliations, contributing about 300 submissions in the first phase and 400 submissions in the second phase. The final performance is evaluated by average precision (AP). Results add strong value and a better understanding of the radar object detection task for the autonomous vehicle community.
Yizhou Wang 0005, Jenq-Neng Hwang, Gaoang Wang, Hui Liu 0011, Kwang-Ju Kim, Hung-Min Hsu, Jiarui Cai, Haotian Zhang 0005, Zhongyu Jiang, Renshu Gu
ICMR5
2021 Multi-Target Multi-Camera Tracking of Vehicles Using Metadata-Aided Re-ID and Trajectory-Based Camera Link Model
abstract
In this paper, we propose a novel framework for multi-target multi-camera tracking (MTMCT) of vehicles based on metadata-aided re-identification (MA-ReID) and the trajectory-based camera link model (TCLM). Given a video sequence and the corresponding frame-by-frame vehicle detections, we first address the isolated tracklets issue from single camera tracking (SCT) by the proposed traffic-aware single-camera tracking (TSCT). Then, after automatically constructing the TCLM, we solve MTMCT by the MA-ReID. The TCLM is generated from camera topological configuration to obtain the spatial and temporal information to improve the performance of MTMCT by reducing the candidate search of ReID. We also use the temporal attention model to create more discriminative embeddings of trajectories from each camera to achieve robust distance measures for vehicle ReID. Moreover, we train a metadata classifier for MTMCT to obtain the metadata feature, which is concatenated with the temporal attention based embeddings. Finally, the TCLM and hierarchical clustering are jointly applied for global ID assignment. The proposed method is evaluated on the CityFlow dataset, achieving IDF1 76.77%, which outperforms the state-of-the-art MTMCT methods.
Hung-Min Hsu, Jiarui Cai, Yizhou Wang 0005, Jenq-Neng Hwang, Kwang-Ju Kim
IEEE Trans. Image Process.5
2020 Colorectal Cancer Image Segmentation and Classification with Deep Neural Network Based on Information Theory
abstract
Colorectal cancer (CRC) is the development of cancer from the colon or rectum. Microsatellite instability (MSI) status can be considered as an indicator to predict the prognosis of CRC. We employ MSI prediction of CRC image by designing a neural network model of which base network is DeepLabv3+ with OctaveResNet. Additionally, we add a channel sort module to divide a feature map along with channel intensity. Then each feature map goes through distinct convolution paths. Each convolution path is designed based on information theory: the most important feature goes through the lightest convolution path, vice versa. By dividing feature map and applying different amount of convolutional operation, the model can extract features efficiently. In the experiment, total model weight is reduced but accuracy increases.
Hwa-Rang Kim, Kwang-Ju Kim, Kil-Taek Lim, Doo-Hyun Choi
BIBM2
2018 Performance Enhancement of YOLOv3 by Adding Prediction Layers with Spatial Pyramid Pooling for Vehicle Detection
abstract
In recent years, vision-based object detection methods using convolutional neural network (CNN) have been very successful. However, the object detection method using the CNN feature has a disadvantage that lots of feature maps should be generated in order to be robust against the scale change and the occlusion of the object. Also, simply raising a large number of feature maps does not improve performance. We propose a multi-scale vehicle detection with spatial pyramid pooling method which is robust to the scale change of the vehicle and the occlusion by improving the conventional YOLOv3 algorithm. The proposed method was evaluated through the UA-DETRAC benchmark and obtain the state-of-the-art mAP, which is better than those of the DPM, ACF, R-CNN, CompACT, NANO, SA-FRCNN, and Faster-RCNN2.
Kwang-Ju Kim, Pyong-Kun Kim, Yun-Su Chung, Doo-Hyun Choi
AVSS1
2018 UA-DETRAC 2018: Report of AVSS2018 & IWT4S Challenge on Advanced Traffic Monitoring
abstract
A desirable smart traffic-monitoring and street-safety system can elicit and support the intervention of law enforcement agencies or medical staff. Recently, there has been a dramatically higher demand for such smart systems. To this end, the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S) was organized in conjunction with the 15th IEEE International Conference on Advanced Video and Signal-based Surveillance (AVSS 2018). Our goal is to advance the state-of-the-art detection and tracking algorithms and provide a comprehensive performance evaluation for them. We evaluate 5 submitted detection and 7 submitted tracking methods on the large-scale UA-DETRAC benchmark, and the results are shared publicly on the website http://detrac-db. rit.albany.edu. We expect this challenge to advance the research and development of new detection and tracking methods for transportation applications.
Siwei Lyu, Ming-Ching Chang, Dawei Du, Wenbo Li 0001, Yi Wei 0006, Marco Del Coco, Pierluigi Carcagnì, Arne Schumann, Bharti Munjal, Dinh-Quoc-Trung Dang, Doo-Hyun Choi, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Guna Seetharaman, Jang-Woon Baek, Jong Taek Lee, Kannappan Palaniappan, Kil-Taek Lim, Kiyoung Moon, Kwang-Ju Kim, Lars Wilko Sommer, Meltem Brandlmaier, Minsung Kang, Moongu Jeon, Noor Al-Shakarji, Oliver Acatay, Pyong-Kun Kim, Sikandar Amin, Thomas Sikora, Tien Ba Dinh, Tobias Senst, Vu-Gia-Hy Che, Young-Chul Lim, Yun-Su Chung
AVSS21