Ue-Hwan Kim

dblp:189/7930 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-2201-2988ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Frequency-aligned supervision for few-shot neural rendering
Suji Jang, Ue-Hwan Kim
Pattern Recognit.2
2025 Towards Generalizable Scene Change Detection
abstract
While current state-of-the-art Scene Change Detection (SCD) approaches achieve impressive results in well-trained research data, they become unreliable under unseen environments and different temporal conditions; in-domain performance drops from 77.6% to 8.0% in a previously unseen environment and to 4.6% under a different temporal condition—calling for generalizable SCD and benchmark. In this work, we propose the Generalizable Scene Change Detection Framework (GeSCF), which addresses unseen domain performance and temporal consistency—to meet the growing demand for anything SCD. Our method leverages the pre-trained Segment Anything Model (SAM) in a zero-shot manner. For this, we design Initial Pseudo-mask Generation and Geometric-Semantic Mask Matching—seamlessly turning user-guided prompt and single-image based segmentation into scene change detection for a pair of inputs without guidance. Furthermore, we define the Generalizable Scene Change Detection (GeSCD) benchmark along with novel metrics and an evaluation protocol to facilitate SCD research in generalizability. In the process, we introduce the ChangeVPR dataset, a collection of challenging image pairs with diverse environmental scenarios—including urban, suburban, and rural settings. Extensive experiments across various datasets demonstrate that GeSCF achieves an average performance gain of 19.2% on existing SCD datasets and 30.0% on the ChangeVPR dataset, nearly doubling the prior art performance. We believe our work can lay a solid foundation for robust and generalizable SCD research.
Jae-Woo Kim, Ue-Hwan Kim
CVPR2
2024 FourierAugment: Frequency-based image encoding for resource-constrained vision tasks
Jiae Yoon, Myeongjin Lee, Ue-Hwan Kim
Knowl. Based Syst.3
2023 SPU-BERT: Faster human multi-trajectory prediction from socio-physical understanding of BERT
abstract
Accurately predicting pedestrian trajectories requires a human-like socio-physical understanding of movement, nearby pedestrians, and obstacles. However, traditional methods struggle to generate multiple trajectories in the same situation based on socio-physical understanding and are computationally intensive, making real-time application difficult. To overcome these limitations, we propose SPU-BERT, a fast multi-trajectory prediction model that incorporates two non-recursive BERTs for multi-goal prediction (MGP) and trajectory-to-goal prediction (TGP). First, MGP predicts multiple goals through generative models, followed by TGP generating trajectories that approach the predicted goals. SPU-BERT can simultaneously understand movement, social interaction, and scene context from trajectories and semantic maps using a single Transformer encoder, providing explainable results as evidence of socio-physical understanding. In experiments, SPU-BERT accurately predicted future trajectories (with 0.19 m and 7.54 pixels of ADE20 for the ETH/UCY datasets and SDD) with over 100 times faster computation (0.132 s) than the state-of-the-art method. The code is available at https://github.com/kina4147/SPUBERT.
Ki-In Na, Ue-Hwan Kim, Jong-Hwan Kim 0001
Knowl. Based Syst.2
2022 Dual Task Learning by Leveraging Both Dense Correspondence and Mis-Correspondence for Robust Change Detection With Imperfect Matches
abstract
Accurate change detection enables a wide range of tasks in visual surveillance, anomaly detection and mobile robotics. However, contemporary change detection approaches assume an ideal matching between the current and stored scenes, whereas only coarse matching is possible in real-world scenarios. Thus, contemporary approaches fail to show the reported performance in real-world settings. To overcome this limitation, we propose SimSaC. SimSaC concurrently conducts scene flow estimation and change detection and is able to detect changes with imperfect matches. To train SimSaC without additional manual labeling, we propose a training scheme with random geometric transformations and the cut-paste method. Moreover, we design an evaluation protocol which reflects performance in realworld settings. In designing the protocol, we collect a test benchmark dataset, which we claim as another contribution. Our comprehensive experiments verify that SimSaC displays robust performance even given imperfect matches and the performance margin compared to contemporary approaches is huge.
Jin-Man Park, Ue-Hwan Kim, Seon-Hoon Lee, Jong-Hwan Kim 0001
CVPR2
2022 SimVODIS: Simultaneous Visual Odometry, Object Detection, and Instance Segmentation
abstract
Intelligent agents need to understand the surrounding environment to provide meaningful services to or interact intelligently with humans. The agents should perceive geometric features as well as semantic entities inherent in the environment. Contemporary methods in general provide one type of information regarding the environment at a time, making it difficult to conduct high-level tasks. Moreover, running two types of methods and associating two resultant information requires a lot of computation and complicates the software architecture. To overcome these limitations, we propose a neural architecture that simultaneously performs both geometric and semantic tasks in a single thread: simultaneous visual odometry, object detection, and instance segmentation (SimVODIS). SimVODIS is built on top of Mask-RCNN which is trained in a supervised manner. Training the pose and depth branches of SimVODIS requires unlabeled video sequences and the photometric consistency between input image frames generates self-supervision signals. The performance of SimVODIS outperforms or matches the state-of-the-art performance in pose estimation, depth map prediction, object detection, and instance segmentation tasks while completing all the tasks in a single thread. We expect SimVODIS would enhance the autonomy of intelligent agents and let the agents provide effective services to humans.
Ue-Hwan Kim, Se-Ho Kim, Jong-Hwan Kim 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Convolutional Recurrent Reconstructive Network for Spatiotemporal Anomaly Detection in Solder Paste Inspection
abstract
Surface mount technology (SMT) is a process for producing printed-circuit boards. The solder paste printer (SPP), package mounter, and solder reflow oven are used for SMT. The board on which the solder paste is deposited from the SPP is monitored by the solder paste inspector (SPI). If SPP malfunctions due to the printer defects, the SPP produces defective products, and then abnormal patterns are detected by SPI. In this article, we propose a convolutional recurrent reconstructive network (CRRN), which decomposes the anomaly patterns generated by the printer defects, from SPI data. CRRN learns only normal data and detects the anomaly pattern through the reconstruction error. CRRN consists of a spatial encoder (S-Encoder), a spatiotemporal encoder and decoder (ST-Encoder-Decoder), and a spatial decoder (S-Decoder). The ST-Encoder-Decoder consists of multiple convolutional spatiotemporal memories (CSTMs) with a spatiotemporal attention (ST-Attention) mechanism. CSTM is developed to extract spatiotemporal patterns efficiently. In addition, an ST-Attention mechanism is designed to facilitate transmitting information from the spatiotemporal encoder to the spatiotemporal decoder, which can solve the long-term dependency problem. We demonstrate that the proposed CRRN outperforms the other conventional models in anomaly detection. Moreover, we show the discriminative power of the anomaly map decomposed by the proposed CRRN through the printer defect classification.
Yong-Ho Yoo, Ue-Hwan Kim, Jong-Hwan Kim 0001
IEEE Trans. Cybern.2
2021 Type Anywhere You Want: An Introduction to Invisible Mobile Keyboard
abstract
Contemporary soft keyboards possess limitations: the lack of physical feedback results in an increase of typos, and the interface of soft keyboards degrades the utility of the screen. To overcome these limitations, we propose an Invisible Mobile Keyboard (IMK), which lets users freely type on the desired area without any constraints. To facilitate a data-driven IMK decoding task, we have collected the most extensive text-entry dataset (approximately 2M pairs of typing positions and the corresponding characters). Additionally, we propose our baseline decoder along with a semantic typo correction mechanism based on self-attention, which decodes such unconstrained inputs with high accuracy (96.0%). Moreover, the user study reveals that the users could type faster and feel convenience and satisfaction to IMK with our decoder. Lastly, we make the source code and the dataset public to contribute to the research community.
Sahng-Min Yoo, Ue-Hwan Kim, Yewon Hwang, Jong-Hwan Kim 0001
IJCAI2
2021 ChangeSim: Towards End-to-End Online Scene Change Detection in Industrial Indoor Environments
abstract
We present a challenging dataset, ChangeSim, aimed at online scene change detection (SCD) and more. The data is collected in photo-realistic simulation environments with the presence of environmental non-targeted variations, such as air turbidity and light condition changes, as well as targeted object changes in industrial indoor environments. By collecting data in simulations, multi-modal sensor data and precise ground truth labels are obtainable such as the RGB image, depth image, semantic segmentation, change segmentation, camera poses, and 3D reconstructions. While the previous online SCD datasets evaluate models given well-aligned image pairs, ChangeSim also provides raw unpaired sequences that present an opportunity to develop an online SCD model in an end-to-end manner, considering both pairing and detection. Experiments show that even the latest pair-based SCD models suffer from the bottleneck of the pairing process, and it gets worse when the environment contains the non-targeted variations. Our dataset is available at https://sammica.github.io/ChangeSim/.
Jin-Man Park, Jae-Hyuk Jang, Sahng-Min Yoo, Sun-Kyung Lee 0001, Ue-Hwan Kim, Jong-Hwan Kim 0001
IROS5
2021 I-Keyboard: Fully Imaginary Keyboard on Touch Devices Empowered by Deep Neural Decoder
abstract
Text entry aims to provide an effective and efficient pathway for humans to deliver their messages to computers. With the advent of mobile computing, the recent focus of text-entry research has moved from physical keyboards to soft keyboards. Current soft keyboards, however, increase the typo rate due to a lack of tactile feedback and degrade the usability of mobile devices due to their large portion on screens. To tackle these limitations, we propose a fully imaginary keyboard (I-Keyboard) with a deep neural decoder (DND). The invisibility of I-Keyboard maximizes the usability of mobile devices and DND empowered by a deep neural architecture allows users to start typing from any position on the touch screens at any angle. To the best of our knowledge, the eyes-free ten-finger typing scenario of I-Keyboard which does not necessitate both a calibration step and a predefined region for typing is first explored in this article. For the purpose of training DND, we collected the largest user data in the process of developing I-Keyboard. We verified the performance of the proposed I-Keyboard and DND by conducting a series of comprehensive simulations and experiments under various conditions. I-Keyboard showed 18.95% and 4.06% increases in typing speed (45.57 words per minute) and accuracy (95.84%), respectively, over the baseline.
Ue-Hwan Kim, Sahng-Min Yoo, Jong-Hwan Kim 0001
IEEE Trans. Cybern.1
2021 Recurrent Reconstructive Network for Sequential Anomaly Detection
abstract
Anomaly detection identifies anomaly samples that deviate significantly from normal patterns. Usually, the number of anomaly samples is extremely small compared to the normal samples. To handle such imbalanced sample distribution, one-class classification has been widely used in identifying the anomaly by modeling the features of normal data using only normal data. Recently, recurrent autoencoder (RAE) has shown outstanding performance in the sequential anomaly detection compared to the other conventional methods. However, RAE, which has a long-term dependency problem, is optimized only to handle the fixed-length inputs. To overcome the limitations of RAE, we propose recurrent reconstructive network (RRN) as a novel RAE, with three functionalities for anomaly detection of streaming data: 1) a self-attention mechanism; 2) hidden state forcing; and 3) skip transition. The designed self-attention mechanism and the hidden state forcing between the encoder and decoder effectively manage the input sequences of varying length. The skip transition with the attention gate improves the reconstruction performance. We conduct a series of comprehensive experiments on four datasets and verify the superior performance of the proposed RRN in the sequential anomaly detection tasks.
Yong-Ho Yoo, Ue-Hwan Kim, Jong-Hwan Kim 0001
IEEE Trans. Cybern.2
2020 A Stabilized Feedback Episodic Memory (SF-EM) and Home Service Provision Framework for Robot and IoT Collaboration
abstract
The automated home referred to as Smart Home is expected to offer fully customized services to its residents, reducing the amount of home labor, thus improving human beings' welfare. Service robots and Internet of Things (IoT) play the key roles in the development of Smart Home. The service provision with these two main components in a Smart Home environment requires: 1) learning and reasoning algorithms and 2) the integration of robot and IoT systems. Conventional computational intelligence-based learning and reasoning algorithms do not successfully manage dynamic changes in the Smart Home data, and the simple integrations fail to fully draw the synergies from the collaboration of the two systems. To tackle these limitations, we propose: 1) a stabilized memory network with a feedback mechanism which can learn user behaviors in an incremental manner and 2) a robot-IoT service provision framework for a Smart Home which utilizes the proposed memory architecture as a learning and reasoning module and exploits synergies between the robot and IoT systems. We conduct a set of comprehensive experiments under various conditions to verify the performance of the proposed memory architecture and the service provision framework and analyze the experiment results.
Ue-Hwan Kim, Jong-Hwan Kim 0001
IEEE Trans. Cybern.1
2020 3-D Scene Graph: A Sparse and Semantic Representation of Physical Environments for Intelligent Agents
abstract
Intelligent agents gather information and perceive semantics within the environments before taking on given tasks. The agents store the collected information in the form of environment models that compactly represent the surrounding environments. The agents, however, can only conduct limited tasks without an efficient and effective environment model. Thus, such an environment model takes a crucial role for the autonomy systems of intelligent agents. We claim the following characteristics for a versatile environment model: accuracy, applicability, usability, and scalability. Although a number of researchers have attempted to develop such models that represent environments precisely to a certain degree, they lack broad applicability, intuitive usability, and satisfactory scalability. To tackle these limitations, we propose 3-D scene graph as an environment model and the 3-D scene graph construction framework. The concise and widely used graph structure readily guarantees usability as well as scalability for 3-D scene graph. We demonstrate the accuracy and applicability of the 3-D scene graph by exhibiting the deployment of the 3-D scene graph in practical applications. Moreover, we verify the performance of the proposed 3-D scene graph and the framework by conducting a series of comprehensive experiments under various conditions.
Ue-Hwan Kim, Jin-Man Park, Taek-Jin Song, Jong-Hwan Kim 0001
IEEE Trans. Cybern.1
2016 A fuzzy expert system for designing customized workout programs
abstract
Due to the change in life style and diet, modern people suffer from obesity, diabetes, and other types of diseases. Regular practice of exercise can alleviate the negative effects from the diseases and even cure the diseases in certain cases. In addition, regular practice of exercise improves the quality of life. These facts have drawn much attention and people nowadays recognize the importance of exercise. As a result, more and more people hope to start exercising but they lack the knowledge of how and what to exercise. Professional counseling costs relatively expensive and thus it is difficult for ordinary people to access a counselor. To tackle these issues we propose a fuzzy expert system that designs a workout program. The system receives user's body information, preference on exercise style, and available time. Then, the system generates a customized workout program based on fuzzy reasoning. We conduct experiments to verify the performance of the proposed system. The participants enter their body condition, preference and available time and receive customized workout programs from the system. The experiments verifies the applicability of the system. The future research includes the extension of the system to meet various user demands and to reflect a number of expert knowledge sources.
Ue-Hwan Kim, Jong-Hwan Kim 0001
FUZZ-IEEE1