VLDB 2026 Research / reviewers in the wild / expert
Hyungil Kim
dblp:77/6143 · also Hyung-Il Kim, Hyung-il Kim
· DBLP profile ↗
45ranked-venue papers
14as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 12 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Task Prototype-Based Knowledge Retrieval for Multi-Task Learning from Partially Annotated DataabstractMulti-task learning (MTL) is critical in real-world applications such as autonomous driving and robotics, enabling simultaneous handling of diverse tasks. However, obtaining fully annotated data for all tasks is impractical due to labeling costs. Existing methods for partially labeled MTL typically rely on predictions from unlabeled tasks, making it difficult to establish reliable task associations and potentially leading to negative transfer and suboptimal performance. To address these issues, we propose a prototype-based knowledge retrieval framework that achieves robust MTL instead of relying on predictions from unlabeled tasks. Our framework consists of two key components: (1) a task prototype embedding task-specific characteristics and quantifying task associations, and (2) a knowledge retrieval transformer that adaptively refines feature representations based on these associations. To achieve this, we introduce an association knowledge generating (AKG) loss to ensure the task prototype consistently captures task-specific characteristics. Extensive experiments demonstrate the effectiveness of our framework, highlighting its potential for robust multi-task learning, even when only a subset of tasks is annotated. Youngmin Oh 0003, Hyungil Kim, Jung Uk Kim |
AAAI | 2 |
| 2026 | Foreword to the special section on Recent Advances in Industrial eXtended Reality (XR)
Bernardo Marques, Samuel S. Silva, Hyungil Kim |
Comput. Graph. | 3 |
| 2026 | Adaptive integration of textual context and visual embeddings for underrepresented vision classification
Seongyeop Kim, Hyungil Kim, Yong Man Ro |
Pattern Recognit. | 2 |
| 2025 | Foreword to the special section on eXtended Reality for Industrial and Occupational Supports (XRIOS)
Isaac Cho, Heejin Jeong, Kangsoo Kim, Hyungil Kim, Myounghoon Jeon 0001 |
Comput. Graph. | 4 |
| 2025 | Prompt Tuning of Deep Neural Networks for Speaker-Adaptive Visual Speech RecognitionabstractVisual Speech Recognition (VSR) aims to infer speech into text depending on lip movements alone. As it focuses on visual information to model the speech, its performance is inherently sensitive to personal lip appearances and movements, and this makes the VSR models show degraded performance when they are applied to unseen speakers. In this paper, to remedy the performance degradation of the VSR model on unseen speakers, we propose prompt tuning methods of Deep Neural Networks (DNNs) for speaker-adaptive VSR. Specifically, motivated by recent advances in Natural Language Processing (NLP), we finetune prompts on adaptation data of target speakers instead of modifying the pre-trained model parameters. Different from the previous prompt tuning methods mainly limited to Transformer variant architecture, we explore different types of prompts, the addition, the padding, and the concatenation form prompts that can be applied to the VSR model which is composed of CNN and Transformer in general. With the proposed prompt tuning, we show that the performance of the pre-trained VSR model on unseen speakers can be largely improved by using a small amount of adaptation data (e.g., less than 5 minutes), even if the pre-trained model is already developed with large speaker variations. Moreover, by analyzing the performance and parameters of different types of prompts, we investigate when the prompt tuning is preferred over the finetuning methods. The effectiveness of the proposed method is evaluated on both word- and sentence-level VSR databases, LRW-ID and GRID. Minsu Kim 0001, Hyungil Kim, Yong Man Ro |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Improving Open Set Recognition via Visual Prompts Distilled from Common-Sense KnowledgeabstractOpen Set Recognition (OSR) poses significant challenges in distinguishing known from unknown classes. In OSR, the overconfidence problem has become a persistent obstacle, where visual recognition models often misclassify unknown objects as known objects with high confidence. This issue stems from the fact that visual recognition models often lack the integration of common-sense knowledge, a feature that is naturally present in language-based models but lacking in visual recognition systems. In this paper, we propose a novel approach to enhance OSR performance by distilling common-sense knowledge into visual prompts. Utilizing text prompts that embody common-sense knowledge about known classes, the proposed visual prompt is learned by extracting semantic common-sense features and aligning them with image features from visual recognition models. The unique aspect of this work is the training of individual visual prompts for each class to encapsulate this common-sense knowledge. Our methodology is model-agnostic, capable of enhancing OSR across various visual recognition models, and computationally light as it focuses solely on training the visual prompts. This research introduces a method for addressing OSR, aiming at a more systematic integration of visual recognition systems with common-sense knowledge. The obtained results indicate an enhancement in recognition accuracy, suggesting the applicability of this approach in practical settings. Seongyeop Kim, Hyungil Kim, Yong Man Ro |
AAAI | 2 |
| 2024 | MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
Youngmin Oh 0003, Hyungil Kim, Seong Tae Kim 0001, Jung Uk Kim |
ECCV (10) | 2 |
| 2024 | Whirling Interface: Hand-based Motion Matching Selection for Small Target on XR DisplaysabstractWe introduce “Whirling Interface,” a selection method for XR displays using bare-hand motion matching gestures as an input technique. We extend the motion matching input method, by introducing different input states to provide visual feedback and guidance to the users. Using the wrist joint as the primary input modality, our technique reduces user fatigue and improves performance while selecting small and distant targets. In a study with 16 participants, we compared the whirling interface with a standard ray casting method using hand gestures. The results demonstrate that the Whirling Interface consistently achieves high success rates, especially for distant targets, averaging 95.58% with a completion time of 5.58 seconds. Notably, it requires a smaller camera sensing field of view of only 21.45° horizontally and 24.7° vertically. Participants reported lower workloads on distant conditions and expressed a higher preference for the Whirling Interface in general. These findings suggest that the Whirling Interface could be a useful alternative input method for XR displays with a small camera sensing FOV or when interacting with small targets. Seoyoung Oh, Minju Baeck, Hui-Shyong Yeo, Hyungil Kim, Thad Starner, Woontack Woo |
ISMAR | 5 |
| 2024 | Kernel adaptive memory network for blind video super-resolution
Jun-Seok Yun, Minhyuk Kim, Hyungil Kim, Seok Bong Yoo |
Expert Syst. Appl. | 3 |
| 2024 | Text-guided distillation learning to diversify video embeddings for text-video retrieval
Sangmin Lee 0001, Hyungil Kim, Yong Man Ro |
Pattern Recognit. | 2 |
| 2023 | OmniSense: Exploring Novel Input Sensing and Interaction Techniques on Mobile Device with an Omni-Directional CameraabstractAn omni-directional (360°) camera captures the entire viewing sphere surrounding its optical center. Such cameras are growing in use to create highly immersive content and viewing experiences. When such a camera is held by a user, the view includes the user’s hand grip, finger, body pose, face, and the surrounding environment, providing a complete understanding of the visual world and context around it. This capability opens up numerous possibilities for rich mobile input sensing. In OmniSense, we explore the broad input design space for mobile devices with a built-in omni-directional camera and broadly categorize them into three sensing pillars: i) near device ii) around device and iii) surrounding device. In addition we explore potential use cases and applications that leverage these sensing capabilities to solve user needs. Following this, we develop a working system to put these concepts into action, by leveraging these sensing capabilities to enable potential use cases and applications. We studied the system in a technical evaluation and a preliminary user study to gain initial feedback and insights. Collectively these techniques illustrate how a single, omni-purpose sensor on a mobile device affords many compelling ways to enable expressive input, while also affording a broad range of novel applications that improve user experience during mobile interaction. Hui-Shyong Yeo, Erwin Wu, Daehwa Kim, Hyungil Kim, Seoyoung Oh, Luna Takagi, Woontack Woo, Hideki Koike, Aaron J. Quigley |
CHI | 5 |
| 2023 | Exploiting recollection effects for memory-based video object segmentation
Enki Cho, Minkuk Kim, Hyungil Kim, Jinyoung Moon, Seong Tae Kim 0001 |
Image Vis. Comput. | 3 |
| 2023 | Stereoscopic Vision Recalling Memory for Monocular 3D Object DetectionabstractMonocular 3D object detection has drawn increasing attention in various human-related applications, such as autonomous vehicles, due to its cost-effective property. On the other hand, a monocular image alone inherently contains insufficient information to infer the 3D information. In this paper, we propose a new monocular 3D object detector that can recall the stereoscopic visual information about an object, given a left-view monocular image. Here, we devise a location embedding module to handle each object by being aware of its location. Next, given the object appearance of the left-view monocular image, we devise Monocular-to-Stereoscopic (M2S) memory that can recall the object appearance of the right-view and depth information. For this purpose, we introduce a stereoscopic vision memorizing loss that guides the M2S memory to store the stereoscopic visual information. Furthermore, we propose a binocular vision association loss to guide the M2S memory that can associate the information of the left-right view about the object when estimating the depth. As a result, our monocular 3D object detector with the M2S memory can effectively exploit the recalled stereoscopic visual information in the inference phase. The comprehensive experimental results on two public datasets, KITTI 3D Object Detection Benchmark and Waymo Open Dataset, demonstrate the effectiveness of the proposed method. We claim that our method is a step-forward method that follows the behaviors of humans that can recall the stereoscopic visual information even when one eye is closed. Jung Uk Kim, Hyungil Kim, Yong Man Ro |
IEEE Trans. Image Process. | 2 |
| 2023 | Visualizing Hand Force with Wearable Muscle Sensing for Enhanced Mixed Reality Remote CollaborationabstractIn this paper, we present a prototype system for sharing a user's hand force in mixed reality (MR) remote collaboration on physical tasks, where hand force is estimated using wearable surface electromyography (sEMG) sensor. In a remote collaboration between a worker and an expert, hand activity plays a crucial role. However, the force exerted by the worker's hand has not been extensively investigated. Our sEMG-based system reliably captures the worker's hand force during physical tasks and conveys this information to the expert through hand force visualization, overlaid on the worker's view or on the worker's avatar. A user study was conducted to evaluate the impact of visualizing a worker's hand force on collaboration, employing three distinct visualization methods across two view modes. Our findings demonstrate that sensing and sharing hand force in MR remote collaboration improves the expert's awareness of the worker's task, significantly enhances the expert's perception of the collaborator's hand force and the weight of the interacting object, and promotes a heightened sense of social presence for the expert. Based on the findings, we provide design implications for future mixed reality remote collaboration systems that incorporate hand force sensing and visualization. Hyungil Kim, Boram Yoon, Seoyoung Oh, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Effects of Avatar Transparency on Social Presence in Task-Centric Mixed Reality Remote CollaborationabstractDespite the importance of avatar representation on user experience for Mixed Reality (MR) remote collaboration involving various device environments and large amounts of task-related information, studies on how controlling visual parameters for avatars can benefit users in such situations have been scarce. Thus, we conducted a user study comparing the effects of three avatars with different transparency levels (Nontransparent, Semi-transparent, and Near-transparent) on social presence for users in Augmented Reality (AR) and Virtual Reality (VR) during task-centric MR remote collaboration. Results show that avatars with a strong visual presence are not required in situations where accomplishing the collaborative task is prioritized over social interaction. However, AR users preferred more vivid avatars than VR users. Based on our findings, we suggest guidelines on how different levels of avatar transparency should be applied based on the context of the task and device type for MR remote collaboration. Boram Yoon, Jae-eun Shin, Hyungil Kim, Seoyoung Oh, Dooyoung Kim 0001, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | HAZE-Net: High-Frequency Attentive Super-Resolved Gaze Estimation in Low-Resolution Face Images
Jun-Seok Yun, Youngju Na, Hee Hyeon Kim, Hyungil Kim, Seok Bong Yoo |
ACCV (5) | 4 |
| 2022 | Weakly Paired Associative Learning for Sound and Image Representations via Bimodal Associative MemoryabstractData representation learning without labels has attracted increasing attention due to its nature that does not require human annotation. Recently, representation learning has been extended to bimodal data, especially sound and image which are closely related to basic human senses. Existing sound and image representation learning methods necessarily require a large number of sound and image with corresponding pairs. Therefore, it is difficult to ensure the effectiveness of the methods in the weakly paired condition, which lacks paired bimodal data. In fact, according to human cognitive studies, the cognitive functions in the human brain for a certain modality can be enhanced by receiving other modalities, even not directly paired ones. Based on the observation, we propose a new problem to deal with the weakly paired condition: How to boost a certain modal representation even by using other unpaired modal data. To address the issue, we introduce a novel bimodal associative memory (BMA-Memory) with key-value switching. It enables to build sound-image association with small paired bimodal data and to boost the built association with the eas-ily obtainable large amount of unpaired data. Through the proposed associative learning, it is possible to reinforce the representation of a certain modality (e.g., sound) even by using other unpaired modal data (e.g., images). Sangmin Lee 0001, Hyungil Kim, Yong Man Ro |
CVPR | 2 |
| 2022 | The Effects of Device and Spatial Layout on Social Presence During a Dynamic Remote Collaboration Task in Mixed RealityabstractThis paper evaluates factors of social presence during a dynamic remote collaboration task in a technologically asymmetric Mixed Reality (MR) setting for two spatial layouts. While active movement during MR remote collaboration is afforded by how the shared 3D space is mediated and configured, studies investigating the impact of these conditions on user experience have been scarce. In a between-group study $(\mathrm{n}=48)$, a host user in Augmented Reality (AR) and a remote user in Virtual Reality (VR), both wearing Head Mounted Displays (HMDs), simultaneously moved around the shared space to find and assemble parts of a Mars exploration rover together, one group in a Peripheral layout and the other in a Scattered layout with disparate levels of spatial affordance. Results show that while VR facilitates higher co-presence and spatial presence than AR through HMDs, the Peripheral layout enables users to pay more attention to one another than the Scattered. We analyze the results and derive implications aimed at bridging the AR-VR gap in social presence for dynamic MR remote collaboration through the adaptive placement of virtual content in shared spaces. Jae-eun Shin, Boram Yoon, Dooyoung Kim 0001, Hyungil Kim, Woontack Woo |
ISMAR | 4 |
| 2021 | Video Prediction Recalling Long-Term Motion Context via Memory Alignment LearningabstractOur work addresses long-term motion context issues for predicting future frames. To predict the future precisely, it is required to capture which long-term motion context (e.g., walking or running) the input motion (e.g., leg movement) belongs to. The bottlenecks arising when dealing with the long-term motion context are: (i) how to predict the long-term motion context naturally matching input sequences with limited dynamics, (ii) how to predict the long-term motion context with high-dimensionality (e.g., complex motion). To address the issues, we propose novel motion context-aware video prediction. To solve the bottle-neck (i), we introduce a long-term motion context memory (LMC-Memory) with memory alignment learning. The pro-posed memory alignment learning enables to store long-term motion contexts into the memory and to match them with sequences including limited dynamics. As a result, the long-term context can be recalled from the limited in-put sequence. In addition, to resolve the bottleneck (ii), we propose memory query decomposition to store local motion context (i.e., low-dimensional dynamics) and recall the suitable local context for each local part of the input individually. It enables to boost the alignment effects of the memory. Experimental results show that the proposed method outperforms other sophisticated RNN-based methods, especially in long-term condition. Further, we validate the effectiveness of the proposed network designs by conducting ablation studies and memory feature analysis. The source code of this work is available†. Sangmin Lee 0001, Hak Gu Kim, Dae Hwi Choi, Hyungil Kim, Yong Man Ro |
CVPR | 4 |
| 2020 | Anti-Litter Surveillance based on Person Understanding via Multi-Task Learning
Kangmin Bae, Kimin Yun, Hyungil Kim, Youngwan Lee, Jongyoul Park |
BMVC | 3 |
| 2020 | Unsupervised Moving Object Detection through Background Models for PTZ CameraabstractMoving object detection in a video plays an important role in many vision applications. Recently, moving object detection using appearance modeling based on a convolutional neural network has been actively developed. However, the CNN-based methods usually require the user's supervision of the first frame so that it becomes highly dependent on the training dataset. In contrast, the method of finding a foreground, which models a background occupying a large proportion in an image, can detect a moving object efficiently in an unsupervised manner. However, existing methods based on background modeling in a pan-tilt-zoom (PTZ) camera suffer many false positives or loss of moving objects due to the estimation error of camera motion. To overcome the aforementioned limitations, we propose a moving object detection method for a PTZ camera through two background models. In an unsupervised way, our method builds the two background models that have different roles: 1) a coarse background model for detecting large changes, and 2) a fine background model for detecting small changes. In more detail, the coarse background model builds a block-based Gaussian model, and the fine model builds a sample consensus model. Both models are adaptively updated according to the estimated camera motion in the video recorded by a PTZ camera. Then, each foreground result from two background models is incorporated to fill the moving object region. Through experiments, the proposed method achieves better performance than the state-of-the-art methods and operates in real-time without parallel processing. In addition, we showed the effectiveness of the proposed model through improved results of moving object detection through combination with the latest supervised method. Kimin Yun, Hyungil Kim, Kangmin Bae, Jongyoul Park |
ICPR | 2 |
| 2020 | Evaluating Remote Virtual Hands Models on Social Presence in Hand-based 3D Remote CollaborationabstractThis study investigates the effects of a virtual hand representation on the user experience including social presence during hand-based 3D remote collaboration. Although a remote hand appearance is a critical parts of a hand-based telepresence, it has been rarely studied in comparison to studies on the self-embodiment of virtual hands in a 3D environment. Thus, we conducted a user study comparing the three virtual hands models (Skeleton, Low Polygon and Realistic) while performing a remote collaborative task based on the American Sign Language (ASL) using both Augmented Reality (AR) and Virtual Reality (VR) environments. We found that the realistic type was perceived as the most sense of being together, human-like, and trustable representation. The low polygon model could also convey a clear sign and moderate level of social presence. Although the system was configured asymmetrically in AR and VR, little difference in perception was found except for the participant's mental load and message understanding. We then discuss the results and suggest design implications for future hand-based 3D telepresence systems. Boram Yoon, Hyungil Kim, Seoyoung Oh, Woontack Woo |
ISMAR | 2 |
| 2020 | Toward Real-Time Estimation of Driver Situation Awareness: An Eye-tracking Approach based on Moving Objects of InterestabstractEye-tracking techniques have the potential for estimating driver awareness of road hazards. However, traditional eye-movement measures based on static areas of interest may not capture the unique characteristics of driver eyeglance behavior and challenge the real-time application of the technology on the road. This article proposes a novel method to operationalize driver eye-movement data analysis based on moving objects of interest. A human-subject experiment conducted in a driving simulator demonstrated the potential of the proposed method. Correlation and regression analyses between indirect (i.e., eye-tracking) and direct measures of driver awareness identified some promising variables that feature both spatial and temporal aspects of driver eye-glance behavior relative to objects of interest. Results also suggest that eye-glance behavior might be a promising but insufficient predictor of driver awareness. This work is a preliminary step toward real-time, on-road estimation of driver awareness of road hazards. The proposed method could be further combined with computer-vision techniques such as object recognition to fully automate eye-movement data processing as well as machine learning approaches to improve the accuracy of driver awareness estimation. Hyungil Kim, Sujitha Martin, Ashish Tawari, Teruhisa Misu, Joseph L. Gabbard |
IV | 1 |
| 2019 | Is Any Room Really OK? The Effect of Room Size and Furniture on Presence, Narrative Engagement, and Usability During a Space-Adaptive Augmented Reality GameabstractOne of the main challenges in creating narrative-driven Augmented Reality (AR) content for Head Mounted Displays (HMDs) is to make them equally accessible and enjoyable in different types of indoor environments. However, little has been studied in regards to whether such content can indeed provide similar, if not the same, levels of experience across different spaces. To gain more understanding towards this issue, we examine the effect of room size and furniture on the player experience of Fragments, a space-adaptive, indoor AR crime-solving game created for the Microsoft HoloLens. The study compares factors of player experience in four types of spatial conditions: (1) Large Room - Fully Furnished; (2) Large Room - Scarcely Furnished; (3) Small Room - Fully Furnished; and (4) Small Room - Scarcely Furnished. Our results show that while large spaces facilitate a higher sense of presence and narrative engagement, fully-furnished rooms raise perceived workload. Based on our findings, we propose design suggestions that can support narrative-driven, space-adaptive indoor HMD-based AR content in delivering optimal experiences for various types of rooms. Jae-eun Shin, Hayun Kim, Callum Parker, Hyungil Kim, Seoyoung Oh, Woontack Woo |
ISMAR | 4 |
| 2019 | WRIST: Watch-Ring Interaction and Sensing Technique for Wrist Gestures and Macro-Micro PointingabstractTo better explore the incorporation of pointing and gesturing into ubiquitous computing, we introduce WRIST, an interaction and sensing technique that leverages the dexterity of human wrist motion. WRIST employs a sensor fusion approach which combines inertial measurement unit (IMU) data from a smartwatch and a smart ring. The relative orientation difference of the two devices is measured as the wrist rotation that is independent from arm rotation, which is also position and orientation invariant. Employing our test hardware, we demonstrate that WRIST affords and enables a number of novel yet simplistic interaction techniques, such as (i) macro-micro pointing without explicit mode switching and (ii) wrist gesture recognition when the hand is held in different orientations (e.g., raised or lowered). We report on two studies to evaluate the proposed techniques and we present a set of applications that demonstrate the benefits of WRIST. We conclude with a discussion of the limitations and highlight possible future pathways for research in pointing and gesturing with wearable devices. Hui-Shyong Yeo, Hyungil Kim, Aakar Gupta, Andrea Bianchi, Daniel Vogel 0001, Hideki Koike, Woontack Woo, Aaron J. Quigley |
MobileHCI | 3 |
| 2019 | The Effect of Avatar Appearance on Social Presence in an Augmented Reality Remote CollaborationabstractThis paper investigates the effect of avatar appearance on Social Presence and users' perception in an Augmented Reality (AR) telep-resence system. Despite the development of various commercial 3D telepresence systems, there has been little evaluation and discussions about the appearance of the collaborator's avatars. We conducted two user studies comparing the effect of avatar appearances with three levels of body part visibility (head & hands, upper body, and whole body) and two different character styles (realistic and cartoon-like) on Social Presence while performing two different remote collaboration tasks. We found that a realistic whole body avatar was perceived as being the best for remote collaboration, but an upper body or cartoon style could be considered as a substitute depending on the collaboration context. We discuss these results and suggest guidelines for designing future avatar-mediated AR remote collaboration systems. Boram Yoon, Hyungil Kim, Gun A. Lee, Mark Billinghurst, Woontack Woo |
VR | 2 |
| 2018 | Effect of Volumetric Displays on Depth Perception in Augmented RealityabstractAugmented reality (AR) head-up displays (HUD) have previously been explored as a potential information delivery system for drivers. In driving scenarios, correct perception of virtual object distance assists with effective use of AR HUDs in safety-critical applications (e.g., collision warnings). AR volumetric displays purportedly offer increased accuracy of distance perception through consistent presentation of oculomotor cues such as vergence and accommodation over traditional displays. For this paper, we investigated volumetric AR displays as a mean of enhancing perception of virtual objects registered to the real world, specifically in terms of distance perception as a result of binocular cues. We designed and ran an experiment where participants controlled and placed virtual objects next to real-world counterparts at ranges between 7-12 meters. We found that the volumetric AR display outperformed a traditional fixed focal plane AR HUD when considering distances 5 meters greater than the traditional display's focal depth. Lee Lisle, Kyle Tanous, Hyungil Kim, Joseph L. Gabbard, Doug A. Bowman |
AutomotiveUI | 3 |
| 2018 | Driver Behavior and Performance with Augmented Reality Pedestrian Collision Warning: An Outdoor User StudyabstractThis article investigates the effects of visual warning presentation methods on human performance in augmented reality (AR) driving. An experimental user study was conducted in a parking lot where participants drove a test vehicle while braking for any cross traffic with assistance from AR visual warnings presented on a monoscopic and volumetric head-up display (HUD). Results showed that monoscopic displays can be as effective as volumetric displays for human performance in AR braking tasks. The experiment also demonstrated the benefits of conformal graphics, which are tightly integrated into the real world, such as their ability to guide drivers' attention and their positive consequences on driver behavior and performance. These findings suggest that conformal graphics presented via monoscopic HUDs can enhance driver performance by leveraging the effectiveness of monocular depth cues. The proposed approaches and methods can be used and further developed by future researchers and practitioners to better understand driver performance in AR as well as inform usability evaluation of future automotive AR applications. Hyungil Kim, Joseph L. Gabbard, Alexandre Miranda Añon, Teruhisa Misu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2018 | Augmented Reality Interface Design Approaches for Goal-directed and Stimulus-driven Driving TasksabstractThe automotive industry is rapidly developing new in-vehicle technologies that can provide drivers with information to aid awareness and promote quicker response times. Particularly, vehicles with augmented reality (AR) graphics delivered via head-up displays (HUDs) are nearing mainstream commercial feasibility and will be widely implemented over the next decade. Though AR graphics have been shown to provide tangible benefits to drivers in scenarios like forward collision warnings and navigation, they also create many new perceptual and sensory issues for drivers. For some time now, designers have focused on increasing the realism and quality of virtual graphics delivered via HUDs, and recently have begun testing more advanced 3D HUD systems that deliver volumetric spatial information to drivers. However, the realization of volumetric graphics adds further complexity to the design and delivery of AR cues, and moreover, parameters in this new design space must be clearly and operationally defined and explored. In this work, we present two user studies that examine how driver performance and visual attention are affected when using fixed and animated AR HUD interface design approaches in driving scenarios that require top-down and bottom-up cognitive processing. Results demonstrate that animated design approaches can produce some driving gains (e.g., in goal-directed navigation tasks) but often come at the cost of response time and distance. Our discussion yields AR HUD design recommendations and challenges some of the existing assumptions of world-fixed conformal graphic approaches to design. Coleman Merenda, Hyungil Kim, Kyle Tanous, Joseph L. Gabbard, Blake Feichtl, Teruhisa Misu, Chihiro Suga |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | Did You See Me?: Assessing Perceptual vs. Real Driving Gains Across Multi-Modal Pedestrian Alert SystemsabstractIn-vehicle support systems have the potential to reduce the risk of pedestrian collisions and promote gains in braking performance and visual attention when scanning for threats on the road. This study investigated changes in driver behavior in pedestrian collision scenarios with increasing urgency while using varying levels of pedestrian alert system (PAS) support in a medium fidelity driving simulator. During pedestrian collision scenarios, we assessed drivers' eye gaze behavior, braking performance, and acceptance ratings across three levels of PAS and four levels of increasing urgency, defined as time to collision (TTC). Results suggest that both audio- and visually-based PAS do not produce gains in the localization of pedestrians, but can nevertheless improve drivers' braking performance in events where pedestrians may pose a threat. Our results further suggest that drivers exhibit both innate and direct confidence in visually-based PAS support, despite no concurrent gains in visual scanning performance. Coleman Merenda, Hyungil Kim, Joseph L. Gabbard, Samantha Leong, David R. Large, Gary E. Burnett |
AutomotiveUI | 2 |
| 2017 | Color channel-wise recurrent learning for facial expression recognitionabstractFacial expression recognition is increasingly gaining importance in emerging affective computing applications. In practice, achieving accurate facial expression recognition is still challenging due to environmental variations. In this paper, we propose a color channel-wise recurrent facial feature learning. The proposed method adopts recurrent neural network to learn expression features sequentially along color channels. The proposed network preserves discriminative expression feature through a long short-term memory for the sequence of color spatial features. Comprehensive experiments have been conducted on the publically available CMU Multi-PIE dataset under illumination variations. Experimental results showed that the proposed method achieved higher recognition rates compared to the state-of-the-art methods. Jinhyeok Jang, Dae Hoe Kim, Hyungil Kim, Yong Man Ro |
ICASSP | 3 |
| 2017 | Effective and efficient human action recognition using dynamic frame skipping and trajectory rejection
Jeong-Jik Seo, Hyungil Kim, Wesley De Neve, Yong Man Ro |
Image Vis. Comput. | 2 |
| 2016 | Collaborative facial color feature learning of multiple color spaces for face recognitionabstractFacial color is known as playing an important role in face recognition. Color face recognition has been investigated in the last decade. Recently, deep learning has attracted considerable attention due to their high performance in face recognition. The importance of the color in a deep learning framework is not fully investigated yet. In this paper, we have conducted experiments to investigate the effectiveness of facial color in face recognition with deep learning. Through experimental results, we have demonstrated that facial color is helpful for enhancing the recognition performance in deep learning and color space selection is crucial to achieve high performance. Moreover, by fusing features from multiple color spaces, the face recognition accuracy has been considerably improved. Hyungil Kim, Yong Man Ro |
ICIP | 1 |
| 2016 | Look at Me: Augmented Reality Pedestrian Warning System Using an In-Vehicle Volumetric Head Up DisplayabstractCurrent pedestrian collision warning systems use either auditory alarms or visual symbols to inform drivers. These traditional approaches cannot tell the driver where the detected pedestrians are located, which is critical for the driver to respond appropriately. To address this problem, we introduce a new driver interface taking advantage of a volumetric head-up display (HUD). In our experimental user study, sixteen participants drove a test vehicle in a parking lot while braking for crossing pedestrians using different interface designs on the HUD. Our results showed that spatial information provided by conformal graphics on the HUD resulted in not only better driver performance but also smoother braking behavior as compared to the baseline. Hyungil Kim, Alexandre Miranda Añon, Teruhisa Misu, Nanxiang Li, Ashish Tawari, Kikuo Fujimura |
IUI | 1 |
| 2016 | Casting shadows: Ecological interface design for augmented reality pedestrian collision warningabstractEcological interface design (EID) has the opportunity to complement current approaches for augmented reality (AR) interface design by considering human-environment interaction and leveraging the inherent benefit of AR interfaces: conformal graphics. This work applies EID to design a novel interface for pedestrian collision warning for an automotive AR head-up display (HUD). Our initial usability evaluation shows potential benefits of incorporating EID into AR interface design. Hyungil Kim, Jessica D. Isleib, Joseph L. Gabbard |
VR | 1 |
| 2016 | Feature scalability for a low complexity face recognition with unconstrained spatial resolution
Hyungil Kim, Seung-Ho Lee, Yong Man Ro |
Multim. Tools Appl. | 1 |
| 2015 | Face image assessment learned with objective and relative face image qualities for improved face recognitionabstractConsiderable research efforts have been made for face recognition in various real-world applications. However, degraded face images, acquired in the real-world, make face recognition difficult. In this paper, we propose a new face image quality assessment that aims to realize a robust and reliable face recognition system. The proposed method considers two factors for face image quality, i.e., visual quality and mismatch between training and test face images. A face image quality assessor is learned based on the two factors to discriminate useful faces from unuseful ones. The proposed face image quality assessment model is robust and adaptive to face recognition systems by employing a learned assessment. Our experimental results on a challenging database show significant improvement in face recognition accuracy by the proposed method. Hyungil Kim, Seung-Ho Lee, Yong Man Ro |
ICIP | 1 |
| 2015 | Multispectral Texture Features from Visible and Near-Infrared Synthetic Face Images for Face RecognitionabstractRecently, high-performance face recognition has attracted research attention in real-world scenarios. Thanks to the advances in sensor technology, face recognition system equipped with multiple sensors has been widely researched. Among them, face recognition system with near-infrared imagery has been one important research topic. In this paper, complementary effect resided in face images captured by nearinfrared and visible rays is exploited by combining two distinct spectral images (i.e., face images captured by near-infrared and visible rays). We propose a new texture feature (i.e., multispectral texture feature) extraction method with synthesized face images to achieve high-performance face recognition with illumination-invariant property. The experimental results show that the proposed method enhances the discriminative power of features thanks the complementary effect. Hyungil Kim, Seung-Ho Lee, Yong Man Ro |
ISM | 1 |
| 2015 | Pose-Robust and Discriminative Feature Representation by Multi-task Deep Learning for Multi-view Face RecognitionabstractAutomatic face recognition (FR) under uncontrolled environments has attracted considerable research attention. In the uncontrolled environments, pose variation is known as one of the crucial factors that influences FR performance. In this paper, we propose a discriminative and pose-robust feature representation using the multi-task learning in deep convolutional neural networks (ConvNet). We introduce four tasks (i.e., maximizing inter-class variation, minimizing intraclass variation, minimizing intra-pose variation, and preserving pose continuity) to learn the ConvNet. Moreover, two-stage learning strategy is proposed to minimize the error functions in learning the deep ConvNet. The extensive experimental results (with the challenging CMU MultiPIE dataset containing pose variations) show that the proposed method outperform stateof-the-art in terms of FR accuracy. Furthermore, the proposed method shows significant improvement even for the face images whose poses are not included in training set. Jeong-Jik Seo, Hyungil Kim, Yong Man Ro |
ISM | 2 |
| 2014 | Adaptive feature extraction for blurred face images in facial expression recognitionabstractIn real world facial expression recognition, blurred face images could hamper achieving high performance due to the lack of distinct edges and textures. In this paper, we propose a new feature extraction method that is robust to blurred face images for facial expression recognition. In the proposed method, the facial feature is extracted adaptively depending on the image sharpness, aiming to achieve robustness against blurred face images. Experimental results on blurred face images demonstrate that the proposed method outperforms the exiting feature extraction method. Hyungil Kim, Seung-Ho Lee, Yong Man Ro |
ICIP | 1 |
| 2014 | Investigating Cascaded Face Quality Assessment for Practical Face Recognition SystemabstractRecently, the development of practical face recognition (FR) system has received much attention. Despite of its extensive study, the FR performance could be severely degraded in real-life scenario (e.g., CCTV surveillance), due to uncontrolled face image conditions of pose/alignment, blur, and brightness. This paper proposes new automated face quality assessment (FQA) framework built-in to a practical FR system. In the proposed framework, three quality factors in face images are rapidly evaluated owing to a cascaded classification. Only face images that have been verified by the FQA are used in recognition phase. Our experiment shows that the cascaded FQA can successfully discard face images that could negatively affect FR. Hyungil Kim, Seung-Ho Lee, Yong Man Ro |
ISM | 1 |
| 2014 | Behind the Glass: Driver Challenges and Opportunities for AR Automotive ApplicationsabstractAs the automotive industry moves toward the car of the future, technology companies are developing cutting-edge systems, in vehicle and out, that aim to make driving safer, more pleasant, and more convenient. While we are already seeing some successful video-based augmented reality (AR) auxiliary displays (e.g., center-mounted backup aid systems), the application opportunities of optical see-through AR as presented on a drivers' windshield are yet to be fully tapped; nor are the visual perceptual and attention challenges fully understood. As we race to field AR applications in transportation, we should first consider the perceptual and distraction issues that are known in both the AR and transportation communities, with a focus on the unique and intersecting aspects for driving applications. This paper describes the some opportunities and driver challenges associated with AR applications in the automotive domain. We first present a basic research space to assist in these inquiries, which delineates head-mounted from heads-up and center-mounted displays; video from optical see-through displays; and world-fixed from screen-fixed AR graphics. We then address benefits of AR related to primary, secondary, and tertiary driver tasks as well as driver perception and cognition challenges inherent in automotive AR systems. Joseph L. Gabbard, Gregory M. Fitch, Hyungil Kim |
Proc. IEEE | 3 |
| 2013 | Exploring head-up augmented reality interfaces for crash warning systemsabstractCrash warning systems are designed to help avoid vehicle accidents by notifying drivers of potential hazards. In typical crash warning systems, primary warning information is provided through visual, audible and/or haptic cues. In general, the use of crash warning systems results in safer driving. However, driver vehicle interfaces that employ visual warning elements, such as text messages appearing on the center console display and blindspot detection icons in side view mirrors, may take drivers' eyes off the road momentarily, and may lead to divided attention, distracted driving and increased crash risks. To address this, we propose an augmented reality (AR) head-up display interface for crash warning systems that displays visual cues on the drivers' view of the road, with the ultimate goal of increasing driver awareness and safety. In this paper, we describe a simulator-based comparative user study to begin understanding the effect of AR interface design features on driver performance, mental workload, and preferences. Our results support the hypothesis that head-up AR display interfaces for crash warning systems have potential safety benefits and a high likelihood of driver acceptance. Hyungil Kim, Xuefang Wu, Joseph L. Gabbard, Nicholas F. Polys |
AutomotiveUI | 1 |
| 2004 | Feature-Based Prediction of Unknown Preferences for Nearest-Neighbor Collaborative FilteringabstractRecommendation systems analyze user preferences and recommend items to a user by predicting the user's preference for those items. Among various kinds of recommendation methods, collaborative filtering (CF) has been widely used and successfully applied to practical applications. However, collaborative filtering has two inherent problems: data sparseness and the cold-start problems. In this paper, we propose a method of integrating additional feature information of users and items into CF to overcome the difficulties caused by sparseness and improve the accuracy of recommendation. Several experimental results that show the effectiveness of the proposed method are also presented. Hyungil Kim, Juntae Kim, Jon Herlocker |
ICDM | 1 |
| 2004 | Integrating Feature Information for Improving Accuracy of Collaborative Filtering
Hyungil Kim, Juntae Kim, Jon Herlocker |
PRICAI | 1 |