Kyoungoh Lee

dblp:203/7566 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-5273-0131ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Structure and sensitivity in 3D human pose similarity quantification and estimation
Kyoungoh Lee, Jungwoo Huh, Jiwoo Kang 0001, Sanghoon Lee 0001
Pattern Recognit.1
2025 Pixel-Level Fire Origin Localization via Digital Twin Mapping for Wildfire Surveillance Framework
abstract
Wildfire monitoring systems play a critical role in minimizing environmental and societal damage. Recent advances in computer vision, particularly deep learning-based fire detection, have enabled more accurate and scalable solutions. However, conventional fire detection methods often struggle with wildfire scenarios due to wide spatial extent, the demand for precise localization, and the urgency of early response. To overcome these challenges, we propose a wildfire monitoring framework capable of pixel-level fire origin localization mapped onto a GPS-calibrated digital twin of mountainous terrain. Our system integrates visual fire detection with terrain-aware 3D projection, enabling accurate mapping of fire origins to real-world coordinates. Experimental results on wildfire datasets demonstrate that our method achieves high accuracy in both early fire detection and precise localization, offering a practical and scalable solution for real-world wildfire monitoring.
Dongyoung Kim, In-Su Jang, Kwang-Ju Kim, Kyoungoh Lee
AVSS5
2025 PETS2025: Multi-Authority Multi-Sensor Maritime Surveillance Challenge and Evaluation
abstract
This paper presents the outcomes of the PETS2025 challenge, held in conjunction with AVSS 2025 and sponsored by the EU-funded EURMARS project. The challenge introduces a novel maritime surveillance dataset comprising image sequences captured by diverse multi-altitude, multimodal sensors, reflecting the real-world multi-authority environment. The key tasks include: (1) object detection using various sensors across different platforms (ground-based and low-altitude aerial) and spectral ranges (visible, thermal, ultraviolet (UV), and short-wave infrared (SWIR)); (2) long-term tracking of targets in maritime environments spanning both sea and land; and (3) approximating target geolocations by using sensor imagery and telemetry data. Performance evaluations of results submitted by 12 international participants are discussed. The results show the effectiveness of these submissions and highlight ongoing challenges posed by heterogeneous sensors and complex environments. These challenges emphasise the need to further improve detection, tracking, and geolocation approximation for maritime and coastal surveillance.
Thanet Markchom, Jonathan N. Boyle, Lulu Chen, James M. Ferryman, Matteo Marturini, Stephan Veigl, Andreas Opitz, Andreas Kriechbaum-Zabini, Romaios Bratskas, Anastasios Gkamaris, Dimitris Papachristos, George Leventakis, Wenjun Fan, Hsiang-Wei Huang, Jeng-Neng Hwang, Pyong-Kun Kim, Kwangju Kim, Chung-I Huang, Kenta Saito, Shunta Kaneko, Kyoko Sudo, Nguyen Thanh Thien, Meng-Yu Kao, Jun-Wei Hsieh, Teepakorn Lilek, Tossapol Pomsuwan, Jinjie Gu, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler, Stephanie Stacy, Alfredo Gabaldon, Peter Tu, Dongyoung Kim, Kyoungoh Lee
AVSS37
2024 MOVES: Motion-Oriented VidEo Sampling for Natural Language-Based Vehicle Retrieval
abstract
Retrieving the target vehicle through natural language descriptions plays a crucial role in intelligent transportation systems. Existing methods tackle this task by employing models that leverage the correlation between textual and visual representations, such as CLIP. However, these models struggle to capture the temporal characteristics of video data, and researchers enhance temporal understanding performance through various data augmentation and video encoders. Yet, conventional approaches in previous studies often overlook the detailed temporal characteristics of vehicles. To overcome this limitation, we introduce a MOVES: Motion-Oriented VidEo Sampling method to effectively utilize the motion information of the target vehicle. Furthermore, we construct a robust model by implementing a re-ranking algorithm to address a variety of vehicle attributes. As a result, our proposed model achieves state-of-the-art performance on the public vehicle retrieval dataset.
Dongyoung Kim, Kyoungoh Lee, In-Su Jang, Kwang-Ju Kim, Pyong-Kun Kim, Jaejun Yoo 0001
AVSS2
2024 TRET: Two Stream-Based Regionally Enhanced Transformers for Person Re-Identification
abstract
Person Re-IDentification (ReID) is a pivotal method for pedestrian tracking and retrieval. This research is inherently challenged by large changes in intra-class or small changes in inter-class. To address this challenge, many researchers have recently introduced transformer-based models, which have shown excellent results. The primary objective of these models is to generate robust features that effectively distinguish between classes and enable generalization. However, existing methods still suffer from class discrimination due to unnecessary noise, including the background. To overcome this limitation, we propose a novel approach called Two stream-based Regionally Enhanced Transformers (TRET) that focuses on the target to be identified. To concentrate on the target region, the TRET utilizes a structure that leverages the pedestrian mask. Furthermore, the proposed model generalizes well by utilizing Contrastive Language-Image Pretraining as the backbone. Finally, our proposed model achieves state-of-the-art performance on the public datasets.
Kyoungoh Lee, Kwang-Ju Kim, Pyong-Kun Kim, In-Su Jang
ICASSP1
2023 Unlocking Potential of 3D-aware GAN for More Expressive Face Generation
abstract
As style-based image generators have achieved disentanglement in features by converting latent vector space to style vector space, numerous efforts have been made to enhance the controllability of the latent. However, existing methods for controllable models have limitations in precisely creating high-resolution faces with large expressions. The degradation is due to the dependence on the training dataset, as the high-resolution face datasets do not have sufficient expressive images. To tackle this challenge, we propose a robust training framework for 3D-aware generative adversarial networks to learn the high-quality generation of more expressive faces through a signed distance field. First, we propose a novel 3D enforcement loss to generate more expressive images in an unsupervised manner. Second, we introduce a partial training method to fine-tune the network on multiple datasets without loss of image resolution. Finally, we propose a ray-scaling scheme for the volume renderer to represent a face at arbitrary scales. Through the proposed framework, the network learns 3D face priors, such as expressional shapes of the parametric facial model, to generate detailed faces. The experimental results outperform the methods of the state of the art, showing strong benefits in the generation of high-resolution facial expressions.
Juheon Hwang, Jiwoo Kang 0001, Kyoungoh Lee, Sanghoon Lee 0001
ICMR3
2023 From Human Pose Similarity Metric to 3D Human Pose Estimator: Temporal Propagating LSTM Networks
abstract
Predicting a 3D pose directly from a monocular image is a challenging problem. Most pose estimation methods proposed in recent years have shown 'quantitatively' good results (below ∼ 50mm). However, these methods remain 'perceptually' flawed because their performance is only measured via a simple distance metric. Although this fact is well understood, the reliance on 'quantitative' information implies that the development of 3D pose estimation methods has been slowed down. To address this issue, we first propose a perceptual Pose SIMilarity (PSIM) metric, by assuming that human perception (HP) is highly adapted to extracting structural information from a given signal. Second, we present a perceptually robust 3D pose estimation framework: Temporal Propagating Long Short-Term Memory networks (TP-LSTMs). Toward this, we analyze the information-theory-based spatio-temporal posture correlations, including joint interdependency, temporal consistency, and HP. The experimental results clearly show that the proposed PSIM metric achieves a superior correlation with users' subjective opinions than conventional pose metrics. Furthermore, we demonstrate the significant quantitative and perceptual performance improvements of TP-LSTMs compared to existing state-of-the-art methods.
Kyoungoh Lee, Woojae Kim, Sanghoon Lee 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 REET: Region-Enhanced Transformer for Person Re-Identification
abstract
Person re-identification (ReID) plays a significant role in intelligent surveillance systems. However, it is challenging due to large variations in the intra-class, where the same person is captured in different scenes or cameras. The current person ReID research focuses on creating robust features for class distinction and generalizing neural networks for covering various target domains to address the issue. Recently, after the achievement of vision transformers, the application of transformers has also begun to person ReID studies. The transformer-based methods have improved quantitative performance of person ReID; however, they still suffer from class distinction. Therefore, this paper proposes a novel region-enhanced transformer (REET) to create robust ReID features. Unlike conventional transformer-based approaches, the REET emphasizes the tokens generated by region-level. Our method achieves state-of-the-art results on three public datasets; Market1501, DukeMTMC, and CUHK-03.
Kyoungoh Lee, In-Su Jang, Kwang-Ju Kim, Pyong-Kun Kim
AVSS1
2022 Self-Updatable Database System Based on Human Motion Assessment Framework
abstract
Recently, human motion-centric videos have been attracting attention in the field of computer vision. Observing and detecting human motion in intelligent surveillance camera systems is essential for understanding the intentions of target subjects. However, these videos have vast amounts of disparate and complex information, and hence they are difficult to process and label automatically. As a result, building and maintaining a database using motion-centric videos requires considerable labor in trimming and classifying the videos. Therefore, we propose a self-updatable motion database system based on a human motion assessment framework for evaluating complex human movements. The framework quantifies three primitive motion properties: stability, liveliness, and attention. This assessment highlights the semantics of human motion in the input video. The semantic motion sequence obtained after the motion assessment is compared with a similarity motion database to determine whether the database needs to be updated; for efficient comparison, we introduce a sequential autoencoder model with a long short-term memory neural network. The proposed system maintains the database within a surveillance camera system using a motion update algorithm; unseen motions in the database are updated using a camera-based surveillance system. In addition, this framework combines state-of-art action recognition methods to improve performance by up to 11% via the self-update of motion.
Kyoungoh Lee, Yeseung Park, Jungwoo Huh, Jiwoo Kang 0001, Sanghoon Lee 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Optimal Camera Point Selection Toward the Most Preferable View of 3-D Human Pose
abstract
Answering the question “what is the most preferable view of a three-dimensional (3-D) human model?” is a challenge in computer vision, computer graphics, and cinematography applications because the appearance of a human, for a given pose, relies on the viewpoint of the user. Currently, to the best of the authors’ knowledge, solid research on the most preferable viewing angle for obtaining numerical subjective evaluation scores has not been conducted. In this study, we investigate a metric that can be used to quantify the view of a 3-D human model, whose value is maximized at the most favorable camera angle in accordance with subjective assessments done by users. For an objective assessment in a numerical form, in this study, we define three view selection metrics: the 1)normalized limb length sum; 2)normalized area of a two-dimensional bounding box; and 3)normalized visible area of a 3-D bounding box. Finally, we formulate a viewpoint optimization problem whose objective function is the sum of the metrics. However, the objective function is nonconcave, and the solution set of the constraint is nonconvex. To overcome this difficulty, we employ decomposition and penalty methods. From the simulation results, it is verified that the average of the viewpoint selection error between the ground truth viewpoint and the optimal viewpoint obtained by the proposed algorithm is very close to the lower bound of the viewpoint selection error.
Beom Kwon, Jungwoo Huh, Kyoungoh Lee, Sanghoon Lee 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2018 Propagating LSTM: 3D Pose Estimation Based on Joint Interdependency
Kyoungoh Lee, Inwoong Lee, Sanghoon Lee 0001
ECCV (7)1