VLDB 2026 Research / reviewers in the wild / expert
Kimin Yun
dblp:89/10771
· DBLP profile ↗
24ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0002-4493-9437ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Augmentation Strategy Selection for Incremental Object Detection
Yujeong Oh, Kimin Yun |
ICPR (14) | 4 |
| 2025 | Distributional Uncertainty for Out-of-Distribution DetectionabstractEstimating uncertainty from deep neural networks is a widely used approach for detecting out-of-distribution (OoD) samples, which typically exhibit high predictive uncertainty. However, conventional methods such as Monte Carlo (MC) Dropout often focus solely on either model or data uncertainty, failing to align with the semantic objective of OoD detection. To address this, we propose the Free-Energy Posterior Network, a novel framework that jointly models distributional uncertainty and identifying OoD and misclassified regions using free energy. Our method introduces two key contributions: (1) a free-energy-based density estimator parameterized by a Beta distribution, which enables fine-grained uncertainty estimation near ambiguous or unseen regions; and (2) a loss integrated within a posterior network, allowing direct uncertainty estimation from learned parameters without requiring stochastic sampling. By integrating our approach with the residual prediction branch (RPL) framework, the proposed method goes beyond post-hoc energy thresholding and enables the network to learn OoD regions by leveraging the variance of the Beta distribution, resulting in a semantically meaningful and computationally efficient solution for uncertainty-aware segmentation. We validate the effectiveness of our method on challenging real-world benchmarks, including Fishyscapes, RoadAnomaly, and Segment-Me-If-You-Can. Kimin Yun, Jeonghyo Song, Young Joon Yoo |
AVSS | 3 |
| 2025 | CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought ReasoningabstractEffective Out-of-Distribution (OOD) detection is critical for ensuring the reliability of semantic segmentation models, particularly in complex road environments where safety and accuracy are paramount. Despite recent advancements in large language models (LLMs), notably GPT-4, which significantly enhanced multimodal reasoning through Chain-of-Thought (CoT) prompting, the application of CoT-based visual reasoning for OOD semantic segmentation remains largely unexplored. In this paper, through extensive analyses of the road scene anomalies, we identify three challenging scenarios where current state-of-the-art OOD segmentation methods consistently struggle: (1) densely packed and overlapping objects, (2) distant scenes with small objects, and (3) large foreground-dominant objects. To address the presented challenges, we propose a novel CoT-based framework targeting OOD detection in road anomaly scenes. Our method leverages the extensive knowledge and reasoning capabilities of foundation models, such as GPT-4, to enhance OOD detection through improved image understanding and prompt-based reasoning aligned with observed problematic scene attributes. Extensive experiments show that our framework consistently outperforms state-of-the-art methods on both standard benchmarks and our newly defined challenging subset of the RoadAnomaly dataset, offering a robust and interpretable solution for OOD semantic segmentation in complex driving environments. Jeonghyo Song, Kimin Yun, Young Joon Yoo |
AVSS | 2 |
| 2025 | Task-Adaptive Open-Set Detection with Prompt-Tuned AdaptorsabstractThis paper presents a task-adaptive open-set detection framework that preserves zero-shot performance while incorporating task-specific adaptations for enhanced visual understanding. Our method integrates a frozen zero-shot detector with a learnable, task-specific adaptor module, and employs a token-level conditional inference mechanism using prompt-based feature masking. This approach selectively combines features from the pre-trained zero-shot model and the adapted module within a single forward pass, allowing both general and task-specific representations to contribute effectively. Unlike conventional full fine-tuning that transforms an open-set detector into a closed-set detector, our design maintains the inherent open-set capabilities, thereby mitigating overfitting to task-specific biases. Experimental results on the IHP and VFP290K datasets demonstrate that our method outperforms existing techniques in fallen person detection, underscoring its robustness and practical applicability. Kimin Yun, Kangmin Bae, Yu-Seok Bae |
AVSS | 1 |
| 2022 | SWAG-Net: Semantic Word-Aware Graph Network for Temporal Video GroundingabstractIn this paper, to effectively capture non-sequential dependencies among semantic words for temporal video grounding, we propose a novel framework called Semantic Word-Aware Graph Network (SWAG-Net), which adopts graph-guided semantic word embedding in an end-to-end manner. Specifically, we define semantic word features as node features of semantic word-aware graphs and word-to-word correlations as three edge types (i.e., intrinsic, extrinsic, and relative edges) for diverse graph structures. We then apply Semantic Word-aware Graph Convolutional Networks (SW-GCNs) to the graphs for semantic word embedding. For modality fusion and context modeling, the embedded features and video segment features are merged into bi-modal features, and the bi-modal features are aggregated by incorporating local and global contextual information. Leveraging the aggregated features, the proposed method effectively finds a temporal boundary semantically corresponding to a sentence query in an untrimmed video. We verify that our SWAG-Net outperforms state-of-the-art methods on Charades-STA and ActivityNet Captions datasets. Sunoh Kim, Taegil Ha, Kimin Yun, Jin Young Choi 0002 |
CIKM | 3 |
| 2021 | The Dataset and Baseline Models to Detect Human Postural States Robustly against Irregular PosturesabstractIn many visual applications, we often encounter people with irregular postures, such as lying down. Many approaches adopted two-step methods to handle a person with irregular postures: 1) person detection and 2) posture prediction based on the detected person. However, it is challenging to detect irregular postures because the existing detectors were trained with datasets consisting of most upright postures. Therefore, we propose a new Irregular Human Posture (IHP) dataset to handle various postures captured from real-world surveillance cameras. The IHP dataset provides sufficient annotations to understand the posture of person, including segmentation, keypoints, and postural states. This paper also provides two baseline net-works for postural state estimation of the people trained on the IHP dataset. Moreover, we show that our baseline networks effectively detect the people with irregular postures that may be in an urgent situation in a surveillance environment. Kangmin Bae, Kimin Yun, Jungchan Cho, Yu-Seok Bae |
AVSS | 2 |
| 2021 | Position-aware Location Regression Network for Temporal Video GroundingabstractThe key to successful grounding for video surveillance is to understand a semantic phrase corresponding to important actors and objects. Conventional methods ignore comprehensive contexts for the phrase or require heavy computation for multiple phrases. To understand comprehensive contexts with only one semantic phrase, we propose Position-aware Location Regression Network (PLRN) which exploits position-aware features of a query and a video. Specifically, PLRN first encodes both the video and query using positional information of words and video segments. Then, a semantic phrase feature is extracted from an encoded query with attention. The semantic phrase feature and encoded video are merged and made into a context-aware feature by reflecting local and global contexts. Finally, PLRN predicts start, end, center, and width values of a grounding boundary. Our experiments show that PLRN achieves competitive performance over existing methods with less computation time and memory. Sunoh Kim, Kimin Yun, Jin Young Choi 0002 |
AVSS | 2 |
| 2020 | Anti-Litter Surveillance based on Person Understanding via Multi-Task Learning
Kangmin Bae, Kimin Yun, Hyungil Kim, Youngwan Lee, Jongyoul Park |
BMVC | 2 |
| 2020 | Unsupervised Moving Object Detection through Background Models for PTZ CameraabstractMoving object detection in a video plays an important role in many vision applications. Recently, moving object detection using appearance modeling based on a convolutional neural network has been actively developed. However, the CNN-based methods usually require the user's supervision of the first frame so that it becomes highly dependent on the training dataset. In contrast, the method of finding a foreground, which models a background occupying a large proportion in an image, can detect a moving object efficiently in an unsupervised manner. However, existing methods based on background modeling in a pan-tilt-zoom (PTZ) camera suffer many false positives or loss of moving objects due to the estimation error of camera motion. To overcome the aforementioned limitations, we propose a moving object detection method for a PTZ camera through two background models. In an unsupervised way, our method builds the two background models that have different roles: 1) a coarse background model for detecting large changes, and 2) a fine background model for detecting small changes. In more detail, the coarse background model builds a block-based Gaussian model, and the fine model builds a sample consensus model. Both models are adaptively updated according to the estimated camera motion in the video recorded by a PTZ camera. Then, each foreground result from two background models is incorporated to fill the moving object region. Through experiments, the proposed method achieves better performance than the state-of-the-art methods and operates in real-time without parallel processing. In addition, we showed the effectiveness of the proposed model through improved results of moving object detection through combination with the latest supervised method. Kimin Yun, Hyungil Kim, Kangmin Bae, Jongyoul Park |
ICPR | 1 |
| 2019 | Skeleton-Based Action Recognition of People Handling ObjectsabstractIn visual surveillance systems, it is necessary to recognize the behavior of people handling objects such as a phone, a cup, or a plastic bag. In this paper, to address this problem, we propose a new framework for recognizing object-related human actions by graph convolutional networks using human and object poses. In this framework, we construct skeletal graphs of reliable human poses by selectively sampling the informative frames in a video, which include human joints with high confidence scores obtained in pose estimation. The skeletal graphs generated from the sampled frames represent human poses related to the object position in both the spatial and temporal domains, and these graphs are used as inputs to the graph convolutional networks. Through experiments over an open benchmark and our own data sets, we verify the validity of our framework in that our method outperforms the state-of-the-art method for skeleton-based action recognition. Sunoh Kim, Kimin Yun, Jongyoul Park, Jin Young Choi 0002 |
WACV | 2 |
| 2018 | Action-Driven Visual Object Tracking With Deep Reinforcement LearningabstractIn this paper, we propose an efficient visual tracker, which directly captures a bounding box containing the target object in a video by means of sequential actions learned using deep neural networks. The proposed deep neural network to control tracking actions is pretrained using various training video sequences and fine-tuned during actual tracking for online adaptation to a change of target and background. The pretraining is done by utilizing deep reinforcement learning (RL) as well as supervised learning. The use of RL enables even partially labeled data to be successfully utilized for semisupervised learning. Through the evaluation of the object tracking benchmark data set, the proposed tracker is validated to achieve a competitive performance at three times the speed of existing deep network-based trackers. The fast version of the proposed method, which operates in real time on graphics processing unit, outperforms the state-of-the-art real-time trackers with an accuracy improvement of more than 8%. Sangdoo Yun, Jongwon Choi 0002, Young Joon Yoo, Kimin Yun, Jin Young Choi 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Action-Decision Networks for Visual Tracking with Deep Reinforcement LearningabstractThis paper proposes a novel tracker which is controlled by sequentially pursuing actions learned by deep reinforcement learning. In contrast to the existing trackers using deep networks, the proposed tracker is designed to achieve a light computation as well as satisfactory tracking accuracy in both location and scale. The deep network to control actions is pre-trained using various training sequences and fine-tuned during tracking for online adaptation to target and background changes. The pre-training is done by utilizing deep reinforcement learning as well as supervised learning. The use of reinforcement learning enables even partially labeled data to be successfully utilized for semi-supervised learning. Through evaluation of the OTB dataset, the proposed tracker is validated to achieve a competitive performance that is three times faster than state-of-the-art, deep network-based trackers. The fast version of the proposed method, which operates in real-time on GPU, outperforms the state-of-the-art real-time trackers. Sangdoo Yun, Jongwon Choi 0002, Young Joon Yoo, Kimin Yun, Jin Young Choi 0002 |
CVPR | 4 |
| 2017 | Appearance and motion based deep learning architecture for moving object detection in moving cameraabstractBackground subtraction from the given image is a widely used method for moving object detection. However, this method is vulnerable to dynamic background in a moving camera video. In this paper, we propose a novel moving object detection approach using deep learning to achieve a robust performance even in a dynamic background. The proposed approach considers appearance features as well as motion features. To this end, we design a deep learning architecture composed of two networks: an appearance network and a motion network. The two networks are combined to detect moving object robustly to the background motion by utilizing the appearance of the target object in addition to the motion difference. In the experiment, it is shown that the proposed method achieves 50 fps speed in GPU and outperforms state-of-the-art methods for various moving camera videos. Byeongho Heo, Kimin Yun, Jin Young Choi 0002 |
ICIP | 2 |
| 2017 | Motion interaction field for detection of abnormal interactions
Kimin Yun, Young Joon Yoo, Jin Young Choi 0002 |
Mach. Vis. Appl. | 1 |
| 2017 | Scene conditional background update for moving object detection in a moving camera
Kimin Yun, Jongin Lim 0002, Jin Young Choi 0002 |
Pattern Recognit. Lett. | 1 |
| 2016 | Visual Path Prediction in Complex Scenes with Crowded Moving ObjectsabstractThis paper proposes a novel path prediction algorithm for progressing one step further than the existing works focusing on single target path prediction. In this paper, we consider moving dynamics of co-occurring objects for path prediction in a scene that includes crowded moving objects. To solve this problem, we first suggest a two-layered probabilistic model to find major movement patterns and their cooccurrence tendency. By utilizing the unsupervised learning results from the model, we present an algorithm to find the future location of any target object. Through extensive qualitative/quantitative experiments, we show that our algorithm can find a plausible future path in complex scenes with a large number of moving objects. Young Joon Yoo, Kimin Yun, Sangdoo Yun, Jonghee Hong, Hawook Jeong, Jin Young Choi 0002 |
CVPR | 2 |
| 2016 | Attention-inspired moving object detection in monocular dashcam videosabstractThis paper proposes a moving object detection algorithm for a monocular dashcam mounted on a vehicle. To deal with dynamic changes of the scene from the dashcam, we propose a new scheme inspired by human-attention inclination for change detection. Humans do not build a detailed visual representation and perceive a change of the scene based on the structure of an interesting region. In this perspective, our method focuses on a sky and road region of the scene and builds an abstracted background model, which is updated with a spatially adaptive learning rate according to the center-focused tendency of the human gaze. To improve the robustness of detection, the final detection map is refined by combining the results from twin processes applied to the original image and the median-filtered image, respectively. In experiments, we have found that our method outperforms state-of-the-art methods qualitatively and quantitatively on a realistic dashcam video. Kimin Yun, Jongin Lim 0002, Sangdoo Yun, Soo Wan Kim, Jin Young Choi 0002 |
ICPR | 1 |
| 2015 | Robust pan-tilt-zoom tracking via optimization combining motion features and appearance correlationsabstractThis paper proposes a new pan-tilt-zoom (PTZ) tracking method to improve the robustness against occlusions and appearance changes by using motion likelihood map and scale change estimation as well as appearance correlation filter. For this purpose, we introduce a motion likelihood map constructed from motion detection result in addition to the correlation filter. The motion likelihood map is generated by blurring the motion detection result, which shows high probability in the center of target. To combine the correlation filter and the motion likelihood map, we formulate an optimization problem. In addition, to handle the scale change of target, we repeat the combining process for various scale of bounding box. The experiments show that the proposed method outperforms the state-of-the-art methods. Byeongju Lee, Kimin Yun, Jongwon Choi 0002, Jin Young Choi 0002 |
AVSS | 2 |
| 2015 | Gradient preserving RGB-to-gray conversion using random forestabstractThis paper proposes a new algorithm for color-to-gray conversion preserving the gradient information in input color image. To preserve the gradient in a color image, we construct a random forest representing the relation between color intensity and gradient in an input image. The leaf nodes of random trees indicate the gray colors (single channel colors) corresponding to the input RGB colored pixels. From these initial gray colors obtained by the random forest, we determine the final gray scale by keeping the balance between intensity and luminance channels. In our experiments, we show that the proposed method outperforms the state-of-the-arts in view of color constrast preserving ratio and mean squared error versus luminance. Byeongju Lee, Jongwon Choi 0002, Kimin Yun, Jin Young Choi 0002 |
ICIP | 3 |
| 2015 | Robust and fast moving object detection in a non-stationary camera via foreground probability based samplingabstractThis paper proposes a robust and fast scheme to detect moving objects in a non-stationary camera. The state-of-the art methods still do not give a satisfactory performance due to drastic frame changes in a non-stationary camera. To improve the robustness in performance, we additionally use the spatio-temporal properties of moving objects. We build the foreground probability map which reflects the spatio-temporal properties, then we selectively apply the detection procedure and update the background model only to the selected pixels using the foreground probability. The foreground probability is also used to refine the initial detection results to obtain a clear foreground region. We compare our scheme quantitatively and qualitatively to the state-of-the-art methods in the detection quality and speed. The experimental results show that our scheme outperforms all other compared methods. Kimin Yun, Jin Young Choi 0002 |
ICIP | 1 |
| 2014 | Visual surveillance briefing system: Event-based video retrieval and summarizationabstractThis paper presents a visual surveillance briefing (VSB) system which provides event-based retrieval and briefing functions. Traditional event-based video retrieval systems usually aim to analyze the appearance of objects rather than the motion information (e.g. trajectory) of objects. The VSB system adopts the video summarization technique which temporally abstracts the retrieved events to understand the motion patterns of objects. We propose various event features including object's appearances and motion patterns for the purpose of event retrieval and design the energy function to abstract the retrieved events in real-time. To avoid the occlusion problem in the briefed events, we propose an animated displaying method that separately presents the global motion and the local motion of moving objects. Effectiveness of the implemented VSB system is evaluated through several surveillance videos. Sangdoo Yun, Kimin Yun, Soo Wan Kim, Young Joon Yoo, Jiyeoup Jeong |
AVSS | 2 |
| 2014 | Motion Interaction Field for Accident Detection in Traffic Surveillance VideoabstractThis paper presents a novel method for modeling of interaction among multiple moving objects to detect traffic accidents. The proposed method to model object interactions is motivated by the motion of water waves responding to moving objects on water surface. The shape of the water surface is modeled in a field form using Gaussian kernels, which is referred to as the Motion Interaction Field (MIF). By utilizing the symmetric properties of the MIF, we detect and localize traffic accidents without solving complex vehicle tracking problems. Experimental results show that our method outperforms the existing works in detecting and localizing traffic accidents. Kimin Yun, Hawook Jeong, Kwang Moo Yi, Soo Wan Kim, Jin Young Choi 0002 |
ICPR | 1 |
| 2014 | Spatio-temporal weighting in local patches for direct estimation of camera motion in video stabilization
Soo Wan Kim, Shimin Yin, Kimin Yun, Jin Young Choi 0002 |
Comput. Vis. Image Underst. | 3 |
| 2013 | Detection of moving objects with a moving camera using non-panoramic background model
Soo Wan Kim, Kimin Yun, Kwang Moo Yi, Sun Jung Kim, Jin Young Choi 0002 |
Mach. Vis. Appl. | 2 |