EDBT 2026 Demo / reviewers in the wild / expert
Ig-Jae Kim
dblp:27/4612
· DBLP profile ↗
66ranked-venue papers
5as first author
39since 2021 · last 2026
0000-0002-2741-7047ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 4 first-author · 22 since 2021Artificial intelligence and machine learning · 27 · 18 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The AI Genie Phenomenon and Three Types of AI Chatbot Addiction: Escapist Roleplays, Pseudosocial Companions, and Epistemic Rabbit Holes
M. Karen Shen, Jessica Huang, Olivia Liang, Ig-Jae Kim, Dongwook Yoon |
CHI | 4 |
| 2026 | Cloning the Self for Mental Well-Being: A Framework for Designing Safe and Therapeutic Self-Clone ChatbotsabstractAs digital tools increasingly mediate mental health care, self-clone chatbots can offer a uniquely novel approach to intra-personal exploration and self-derived support. Trained to replicate users’ conversational patterns, self-clones allow users to talk to themselves through their digital replicas. Despite the promises, these systems may carry risks around identity confusion, negative reinforcement, and blurred user agency. Through interviews with 16 mental health professionals and 6 general users, we aim to uncover tensions and design opportunities in this emerging space to guide responsible self-clone design. Our analysis produces a design framework organized around three priorities: (1) defining goals and grounding the approach in existing therapeutic models, (2) design dimensions including the self-clone persona and user-clone relationship dynamics, and (3) considerations for minimizing potential emotional and ethical harms. This framework contributes an interdisciplinary foundation for designing self-clone chatbots as AI-mediated self-interaction tools that are emotionally and ethically attuned in mental health contexts. Mehrnoosh Sadat Shirvani, Jackie Crowley, Cher Peng, Jackie Liu, Thomas Chao, Suky Martinez, Laura Brandt, Ig-Jae Kim, Dongwook Yoon |
CHI | 8 |
| 2026 | PASTA: A Scalable Framework for Multi-Policy AI Compliance EvaluationabstractAI compliance is becoming increasingly critical as AI systems grow more powerful and pervasive. Yet the rapid expansion of AI policies creates substantial burdens for resource-constrained practitioners lacking policy expertise. Existing approaches typically address one policy at a time, making multi-policy compliance costly. We present PASTA, a scalable compliance tool integrating four innovations: (1) a comprehensive model-card format supporting descriptive inputs across development stages; (2) a policy normalization scheme; (3) an efficient LLM-powered pairwise evaluation engine with cost-saving strategies; and (4) an interface delivering interpretable evaluations via compliance heatmaps and actionable recommendations. Expert evaluation shows PASTA’s judgments closely align with human experts (ρ ≥.626). The system evaluates five major policies in under two minutes at approximately $3. A user study (N = 12) confirms practitioners found outputs easy-to-understand and actionable, introducing a novel framework for scalable automated AI governance. Ig-Jae Kim, Dongwook Yoon |
CHI | 2 |
| 2026 | VAST-ReID: A Low-Light Benchmark Dataset for Person Re-Identification with Visual and Attribute-Rich Semantic TrackingabstractPerson Re-Identification (ReID) task is important for designing intelligent surveillance systems. ReID can be highly challenging in low-light and low resolution scenarios. Existing ReID datasets predominantly feature cropped pedestrian images captured in well-lit environments, often lacking semantic richness, frame-level temporal continuity, and robustness to adverse conditions. To address these limitations, we introduce VAST-ReID, a new benchmark dataset specifically designed for the low-light person ReID task in real-world surveillance contexts. VAST-ReID consists of 1,441 surveillance videos collected at 24 different locations, capturing 256 distinct pedestrians of various age groups. The dataset emphasizes naturally low-light and visually degraded scenarios. Each identity is annotated with dense bounding boxes and enriched with auxiliary semantic labels, including pedestrian attributes and LLM-generated descriptions. While these annotations are not used during supervised training, they provide valuable semantic context for advancing research in language-guided retrieval and attribute-aware modeling. Additionally, we release identity-aligned image crops under the BoxTrack-ReID subset, which has over 18.7K frames sampled at 1fps from the raw videos, with standard training, gallery, and query splits compatible with the Market-1501 evaluation protocol, enabling straightforward benchmarking. The dataset has been benchmarked against SOTA methods, and experiments reveal that there is huge scope for improvement in ReID research. VAST-ReID is available at: https://github.com/Byte0wl/VAST-ReID Hammad Khan, Rakesh Kumar Giri, Thakare Kamalakar Vijay, Heeseung Choi, Hyungjoo Jung, Debi Prosad Dogra, Ig-Jae Kim |
WACV | 7 |
| 2026 | IMPACT: Interpretable Most Important Person Analysis and Classification using Transformer-based ModelsabstractIdentifying the Most Important Person (MIP) in complex social and sports events remains a challenging problem due to the dynamic nature of group interactions, subtle visual cues, and context-dependent semantics. Traditional methods often struggle to accurately capture the interplay between individuals and the overarching activity, especially in unstructured real-world environments. In addition, the lack of strong supervision and the need for a deeper contextual understanding further complicate the task. In this work, we propose IMPACT, a novel multi-modal framework that leverages recent advances in vision language models to bridge the gap between visual perception and semantic reasoning. Our approach integrates structured scene understanding, natural language generation, and cross-modal learning to jointly model activity recognition and MIP localization. The method integrates language, vision, and spatial reasoning to improve scene interpretability as well as accuracy in group activity recognition tasks. By incorporating language-based representations, the proposed method enables interpretable and robust performance in sports-centric group activity scenarios. Comprehensive experiments on C-Sports and NCAA datasets demonstrate that the framework significantly enhances the localization of key individuals as well as the accuracy of activity prediction, laying the groundwork for a holistic scene understanding in human-centric video and image analysis. Our proposed method achieves an accuracy of 81.6% when compared with human annotator markings and an increase in mAP scores by ∼ 5% for MIP identification. Akshat Rampuria, Kamakshya Prasad Nayak, Thakare Kamalakar Vijay, Tushar Joshi, Aditya Dhananjay Singh, Haesol Park, Heeseung Choi, Hyungjoo Jung, Debi Prosad Dogra, Ig-Jae Kim |
WACV | 10 |
| 2026 | Social Media Clones: Exploring the Impact of Social Delegation with AI Clones through a Design Workbook Study CSCW036abstractSocial media clones are AI-powered social delegates of ourselves created using our personal data. As our identities and online personas intertwine, these technologies have the potential to greatly enhance our social media experience. If mismanaged however, these clones may also pose new risks to our social reputation and online relationships. To set the foundation for a productive and responsible integration, we set out to understand how social media clones will impact our online behavior and interactions. We conducted a series of semi-structured interviews introducing eight speculative clone concepts to 32 social media users through a design workbook. Applying existing work in AI-mediated communication in the context of social media, we found that although clones can offer convenience and comfort, they can also threaten the user’s authenticity and increase distrust within the online community. As a result users tend to behave more like their clones to mitigate discrepancies and interaction breakdowns. These findings are discussed through the lens of past literature in identity and impression management to highlight challenges in the adoption of social media clones by the general public, and propose design considerations for their successful integration into social media platforms. Jackie Liu, Mehrnoosh Sadat Shirvani, Hwajung Hong, Ig-Jae Kim, Dongwook Yoon |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2025 | Navigating Label Ambiguity for Facial Expression Recognition in the WildabstractFacial expression recognition (FER) remains a challenging task due to label ambiguity caused by the subjective nature of facial expressions and noisy samples. Additionally, class imbalance, which is common in real-world datasets, further complicates FER. Although many studies have shown impressive improvements, they typically address only one of these issues, leading to suboptimal results. To tackle both challenges simultaneously, we propose a novel framework called Navigating Label Ambiguity (NLA), which is robust under real-world conditions. The motivation behind NLA is that dynamically estimating and emphasizing ambiguous samples at each iteration helps mitigate noise and class imbalance by reducing the model's bias toward majority classes. To achieve this, NLA consists of two main components: Noise-aware Adaptive Weighting (NAW) and consistency regularization. Specifically, NAW adaptively assigns higher importance to ambiguous samples and lower importance to noisy ones, based on the correlation between the intermediate prediction scores for the ground truth and the nearest negative. Moreover, we incorporate a regularization term to ensure consistent latent distributions. Consequently, NLA enables the model to progressively focus on more challenging ambiguous samples, which primarily belong to the minority class, in the later stages of training. Extensive experiments demonstrate that NLA outperforms existing methods in both overall and mean accuracy, confirming its robustness against noise and class imbalance. To the best of our knowledge, this is the first framework to address both problems simultaneously. JunGyu Lee 0003, Yeji Choi, Haksub Kim, Ig-Jae Kim, Gi Pyo Nam |
AAAI | 4 |
| 2025 | AvatARoid: A Motion-Mapped AR Overlay to Bridge the Embodiment Gap Between Robots and Teleoperators in Robot-Mediated Telepresence
Amit Ghimire, Anova Hou, Ig-Jae Kim, Dongwook Yoon |
CHI | 3 |
| 2025 | Mirror to Companion: Exploring Roles, Values, and Risks of AI Self-Clones through Story Completion
Jessica Huang, Ig-Jae Kim, Dongwook Yoon |
CHI | 2 |
| 2025 | Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor ScenesabstractWe propose a diffusion-based inverse rendering framework that decomposes a single RGB image into geometry, material, and lighting. Inverse rendering is inherently ill-posed, making it difficult to predict a single accurate solution. To address this challenge, recent generative model-based methods aim to present a range of possible solutions. However, finding a single accurate solution and generating diverse solutions can be conflicting. In this paper, we propose a channel-wise noise scheduling approach that allows a single diffusion model architecture to achieve two conflicting objectives. The resulting two diffusion models, trained with different channel-wise noise schedules, can predict a single highly accurate solution and present multiple possible solutions. The experimental results demonstrate the superiority of our two models in terms of both diversity and accuracy, which translates to enhanced performance in downstream applications such as object insertion and material editing. Junyong Choi, Min-Cheol Sagong, SeokYeong Lee, Seung-Won Jung, Ig-Jae Kim, Junghyun Cho |
CVPR | 5 |
| 2025 | Effective SAM Combination for Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation aims to assign pixel-level labels to images across an unlimited range of classes. Traditional methods address this by sequentially connecting a powerful mask proposal generator, such as the Segment Anything Model (SAM), with a pre-trained vision-language model like CLIP. But these two-stage approaches often suffer from high computational costs, memory inefficiencies. In this paper, we propose ESC-Net, a novel one-stage open-vocabulary segmentation model that leverages the SAM decoder blocks for class-agnostic segmentation within an efficient inference framework. By embedding pseudo prompts generated from image-text correlations into SAM’s promptable segmentation framework, ESC-Net achieves refined spatial aggregation for accurate mask predictions. Additionally, a Vision-Language Fusion (VLF) module enhances the final mask prediction through image and text guidance. ESC-Net and PASCAL-Context, outperforming prior methods in both efficiency and accuracy. Comprehensive ablation studies further demonstrate its robustness across challenging conditions. Minhyeok Lee, Suhwan Cho, Sunghun Yang, Heeseung Choi, Ig-Jae Kim, Sangyoun Lee |
CVPR | 6 |
| 2025 | VIGFace: Virtual Identity Generation for Privacy-Free Face Recognition Dataset
Min-Cheol Sagong, Gi Pyo Nam, Junghyun Cho, Ig-Jae Kim |
ICCV | 5 |
| 2025 | CLIPping Imbalances: A Novel Evaluation Baseline and PEARL Dataset for Pedestrian Attribute RecognitionabstractPedestrian Attribute Recognition (PAR) serves as a fun-damental task in computer vision and is crucial for upgradign security systems. It helps in precisely identifying and characterizing various attributes of pedestrians. However, current PAR datasets have certain issues in representing a wide range of attributes correctly, which makes the ex-isting PAR methods less effective in real-world scenarios. Addressing this limitation, this paper introduces PEARL, a comprehensive dataset comprising of diverse pedestrian images annotated with 146 attributes. These samples have been sourced from surveillance videos across twelve coun-tries. This paper also formulates an image-based PAR using language-image fusion strategy and utilizes CLIP as a new evaluation baseline. Specifically, we leverage textual infor-mation by transforming sets of attributes into meaningful sentences. Addressing the inherent data imbalance in PAR, we provide three types of prompt settings to optimize the training of the CLIP model. Our evaluation encompasses a thorough assessment of the proposed baseline model across various datasets, including PEARL dataset as well as estab-lished PAR benchmarks such as PA100K, RAP, and PETA. Thakare Kamalakar Vijay, Lalit Lohani, Kamakshya Prasad Nayak, Debi Prosad Dogra, Heeseung Choi, Hyungjoo Jung, Ig-Jae Kim |
WACV | 7 |
| 2025 | MAIR++: Improving Multi-View Attention Inverse Rendering With Implicit Lighting RepresentationabstractIn this paper, we propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, SVBRDF, and 3D spatially-varying lighting. While multi-view images have been widely used for object-level inverse rendering, scene-level inverse rendering has primarily been studied using single-view images due to the lack of a dataset containing high dynamic range multi-view images with ground-truth geometry, material, and spatially-varying lighting. To improve the quality of scene-level inverse rendering, a novel framework called Multi-view Attention Inverse Rendering (MAIR) was recently introduced. MAIR performs scene-level multi-view inverse rendering by expanding the OpenRooms dataset, designing efficient pipelines to handle multi-view images, and splitting spatially-varying lighting. Although MAIR showed impressive results, its lighting representation is fixed to spherical Gaussians, which limits its ability to render images realistically. Consequently, MAIR cannot be directly used in applications such as material editing. Moreover, its multi-view aggregation networks have difficulties extracting rich features because they only focus on the mean and variance between multi-view features. In this paper, we propose its extended version, called MAIR++. MAIR++ addresses the aforementioned limitations by introducing an implicit lighting representation that accurately captures the lighting conditions of an image while facilitating realistic rendering. Furthermore, we design a directional attention-based multi-view aggregation network to infer more intricate relationships between views. Experimental results show that MAIR++ not only outperforms MAIR and single-view-based methods but also demonstrates robust performance on unseen real-world scenes. Junyong Choi, SeokYeong Lee, Haesol Park, Seung-Won Jung, Ig-Jae Kim, Junghyun Cho |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Few-Shot Neural Radiance Fields under Unconstrained IlluminationabstractIn this paper, we introduce a new challenge for synthesizing novel view images in practical environments with limited input multi-view images and varying lighting conditions. Neural radiance fields (NeRF), one of the pioneering works for this task, demand an extensive set of multi-view images taken under constrained illumination, which is often unattainable in real-world settings. While some previous works have managed to synthesize novel views given images with different illumination, their performance still relies on a substantial number of input multi-view images. To address this problem, we suggest ExtremeNeRF, which utilizes multi-view albedo consistency, supported by geometric alignment. Specifically, we extract intrinsic image components that should be illumination-invariant across different views, enabling direct appearance comparison between the input and novel view under unconstrained illumination. We offer thorough experimental results for task evaluation, employing the newly created NeRF Extreme benchmark—the first in-the-wild benchmark for novel view synthesis under multiple viewing directions and varying illuminations. SeokYeong Lee, Junyong Choi, Seungryong Kim, Ig-Jae Kim, Junghyun Cho |
AAAI | 4 |
| 2024 | Dual Prototype Attention for Unsupervised Video Object SegmentationabstractUnsupervised video object segmentation (VOS) aims to detect and segment the most salient object in videos. The primary techniques used in unsupervised VOS are 1) the collaboration of appearance and motion information; and 2) temporal fusion between different frames. This paper proposes two novel prototype-based attention mechanisms, inter-modality attention (IMA) and inter-frame attention (IFA), to incorporate these techniques via dense propagation across different modalities and frames. IMA densely in-tegrates context information from different modalities based on a mutual refinement. IFA injects global context of a video to the query frame, enabling a full utilization of useful prop-erties from multiple frames. Experimental results on public benchmark datasets demonstrate that our proposed approach outperforms all existing methods by a substantial margin. The proposed two components are also thoroughly validated via ablative study. Code and models are available at https://github.com/Hydragon516/DPA. Suhwan Cho, Minhyeok Lee, Seunghoon Lee 0008, Dogyoon Lee, Heeseung Choi, Ig-Jae Kim, Sangyoun Lee |
CVPR | 6 |
| 2024 | Pedestrian Attribute Recognition Using Hierarchical Transformers
Lalit Lohani, Thakare Kamalakar Vijay, Kamakshya Prasad Nayak, Debi Prosad Dogra, Heeseung Choi, Hyungjoo Jung, Ig-Jae Kim |
ICPR (16) | 7 |
| 2024 | Let's Observe Them Over Time: An Improved Pedestrian Attribute Recognition ApproachabstractDespite poor image quality, occlusions, and small training datasets, recent pedestrian attribute recognition (PAR) methods have achieved considerable performance. However, leveraging only spatial information of different attributes limits their reliability and generalizability. This paper introduces a multi-perspective approach to reduce over-dependence on spatial clues of a single perspective and exploits other aspects available in multiple perspectives. In order to tackle image quality and occlusions, we exploit different spatial clues present across images and handpick the best attribute-specific features to classify. Precisely, we extract the class-activation energy of each attribute and correlate it with the corresponding energy present across other images using the proposed Self-Attentive Cross Relation Module. In the next stage, we fuse this correlation information with similar clues accumulated from the other images. Lastly, we train a classification neural network using combined correlation information with two different losses. We have validated our method on four widely used PAR datasets, namely Market1501, PETA, PA-100k, and Duke. Our method achieves superior performance over most existing methods, demonstrating the effectiveness of a multi-perspective approach in PAR. Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, Ig-Jae Kim |
WACV | 5 |
| 2023 | MAIR: Multi-View Attention Inverse Rendering with 3D Spatially-Varying Lighting EstimationabstractWe propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, a SVBRDF, and 3D spatially-varying lighting. Because multi-view images provide a variety of information about the scene, multi-view images in object-level inverse rendering have been taken for granted. However, owing to the absence of multi-view HDR synthetic dataset, scene-level inverse rendering has mainly been studied using single-view image. We were able to successfully perform scene-level inverse rendering using multi-view images by expanding OpenRooms dataset and designing efficient pipelines to handle multi-view images, and splitting spatially-varying lighting. Our experiments show that the proposed method not only achieves better performance than single-view-based methods, but also achieves robust performance on unseen real-world scene. Also, our sophisticated 3D spatially-varying lighting volume allows for photorealistic object insertion in any 3D location. Junyong Choi, SeokYeong Lee, Haesol Park, Seung-Won Jung, Ig-Jae Kim, Junghyun Cho |
CVPR | 5 |
| 2023 | Face Photo-Sketch Synthesis Via Domain-Invariant Feature EmbeddingabstractFace photo-sketch synthesis involves transforming photos into sketches and vice versa. A well-transformed image should preserve its original identity characteristics and naturalness. However, identity preservation remains a challenge because of the large discrepancy between the photo and sketch domains. To this end, we propose a novel face photo-sketch synthesis framework that uses domain-invariant feature embedding (DIFE). The DIFE framework generates images assuming the domain-invariant feature of an image pair for the same person to be the identity information. A joint feature embedding module considers latent features from two different domains as input and transfers them into the domain-invariant latent space. Subsequently, a semantic-aware decoder completes the desired image guided by multiscale facial parsing masks. Experimental results demonstrate that the DIFE method outperforms state-of-the-art approaches visually and perceptually. Yeji Choi, Kwanghoon Sohn, Ig-Jae Kim |
ICIP | 3 |
| 2023 | DyAnNet: A Scene Dynamicity Guided Self-Trained Video Anomaly Detection NetworkabstractUnsupervised approaches for video anomaly detection may not perform as good as supervised approaches. However, learning unknown types of anomalies using an unsupervised approach is more practical than a supervised approach as annotation is an extra burden. In this paper, we use isolation tree-based unsupervised clustering to partition the deep feature space of the video segments. The RGB-stream generates a pseudo anomaly score and the flow stream generates a pseudo dynamicity score of a video segment. These scores are then fused using a majority voting scheme to generate preliminary bags of positive and negative segments. However, these bags may not be accurate as the scores are generated only using the current segment which does not represent the global behavior of a typical anomalous event. We then use a refinement strategy based on a cross-branch feed-forward network designed using a popular I3D network to refine both scores. The bags are then refined through a segment re-mapping strategy. The intuition of adding the dynamicity score of a segment with the anomaly score is to enhance the quality of the evidence. The method has been evaluated on three popular video anomaly datasets, i.e., UCF-Crime, CCTV-Fights, and UBI-Fights. Experimental results reveal that the proposed framework achieves competitive accuracy as compared to the state-of-the-art video anomaly detection methods. Thakare Kamalakar Vijay, Yash Raghuwanshi, Debi Prosad Dogra, Heeseung Choi, Ig-Jae Kim |
WACV | 5 |
| 2023 | Speculating on Risks of AI Clones to Selfhood and Relationships: Doppelganger-phobia, Identity Fragmentation, and Living MemoriesabstractDigitally replicating the appearance and behaviour of individuals is becoming feasible with recent advancements in deep-learning technologies such as interactive deepfake applications, voice conversion, and virtual actors. Interactive applications of such agents, termed AI clones, pose risks related to impression management, identity abuse, and unhealthy dependencies. Identifying concerns AI clones will generate is a prerequisite to establishing the basis of discourse around how this technology will impact a source individual's selfhood and interpersonal relationships. We presented 20 participants of diverse ages and backgrounds with 8 speculative scenarios to explore their perception towards the concept of AI clones. We found that (1. doppelganger-phobia) the abusive potential of AI clones to exploit and displace the identity of an individual elicits negative emotional reactions; (2. identity fragmentation) creating replicas of a living individual threatens their cohesive self-perception and unique individuality; and (3. living memories) interacting with a clone of someone with whom the user has an existing relationship poses risks of misrepresenting the individual or developing over-attachment to the clone. These findings provide an avenue to discuss preliminary ethical implications, respect for identity and authenticity, and design recommendations for creating AI clones. Patrick Yung Kang Lee, Ning F. Ma, Ig-Jae Kim, Dongwook Yoon |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2023 | RareAnom: A Benchmark Video Dataset for Rare Type AnomaliesabstractExisting video anomaly detection methods and datasets suffer from restricted anomaly categories containing single-source (CCTV) videos recorded in controlled environment, inadequate annotations, and lack of adequate supervision. To mitigate these problems, we introduce a new dataset ( RareAnom ) containing 17 rare types of real-world anomalies (2200 videos) recorded using multiple sources (e.g., CCTV , handheld cameras, dash-cams, and mobile phones) with rich temporal annotations. A new fully unsupervised anomaly detection and classification method has been proposed. It has three stages: training of a 3D Convolution Autoencoder using pseudo-labelled video segments, anomaly detection using latent features, and classification. Unlike the existing datasets, we have benchmarked RareAnom using three levels of supervision: fully, weakly, and unsupervised. It has been compared with UCF-Crime and XD-Violence datasets. The proposed anomaly detection and classification method beats the latest unsupervised methods by 4.49%, 8.66%, and 6.77% on RareAnom, UCF-Crime, and XD-violence datasets, respectively. Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, Ig-Jae Kim |
Pattern Recognit. | 5 |
| 2023 | Detection of Road Accidents Using Synthetically Generated Multi-Perspective Accident VideosabstractRoad accidents are often caused by short abnormal events, including traffic violations, abrupt change in vehicular motion, driver fatigue, etc. Observing an accident event from the right camera perspective plays a crucial role while detecting accidents. However, it may not be possible to capture such abnormal events from a limited camera perspective. We present a deep learning framework to analyze the accident events recorded from multiple perspectives. First, we estimate feature similarity in videos recorded from multiple perspectives. We then divided the video samples into high and low feature similarity groups. Next, we extract spatio-temporal features from each group using two-branch DCNNs and fuse them using a rank-based weighted average pooling strategy followed by classification. We present a new road accident video dataset (MP-RAD), where each accident event is synthetically generated and captured from five independent camera perspectives using a computer gaming platform. Most of the existing road accident datasets use egocentric views or they are captured in fixed camera setups. However, our dataset is large and multi-perspective that can be used to validate ITS-related tasks such as accident detection, accident localization, traffic monitoring, etc. The dataset contains 400 accident events with a total of 2000 videos. We provide temporal annotations of all videos. The proposed framework and the dataset have been cross-validated with latest accident detection baselines trained on real-world road accident videos and vice-versa. The sub-optimal detection accuracy obtained using the baselines indicates that the proposed framework and the dataset can be useful for ITS related research. Code and dataset is available at: https://github.com/draxler1/MP-RAD-Dataset-ITS- Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Gi Pyo Nam, Ig-Jae Kim |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | PPL: Pairwise Prototype Learning for Masked Face Recognition
Gi Pyo Nam, Yu-Jin Hong, Ig-Jae Kim |
BMVC | 4 |
| 2022 | Probabilistic Representations for Video Contrastive LearningabstractThis paper presents Probabilistic Video Contrastive Learning, a self-supervised representation learning method that bridges contrastive learning with probabilistic representation. We hypothesize that the clips composing the video have different distributions in short-term duration, but can represent the complicated and sophisticated video distribution through combination in a common embedding space. Thus, the proposed method represents video clips as normal distributions and combines them into a Mixture of Gaussians to model the whole video distribution. By sampling embeddings from the whole video distribution, we can circumvent the careful sampling strategy or transformations to generate augmented views of the clips, unlike previous deterministic methods that have mainly focused on such sample generation strategies for contrastive learning. We further propose a stochastic contrastive loss to learn proper video distributions and handle the inherent uncertainty from the nature of the raw video. Experimental results verify that our probabilistic embedding stands as a state-of-the-art video representation learning for action recognition and video retrieval on the most popular benchmarks, including UCF101 and HMDB51. Jungin Park, Jiyoung Lee 0005, Ig-Jae Kim, Kwanghoon Sohn |
CVPR | 3 |
| 2022 | A Log-Structured Merge Tree-aware Message Authentication Scheme for Persistent Key-Value Stores
Ig-Jae Kim, J. Hyun Kim, Minu Chung, Hyungon Moon, Sam H. Noh |
FAST | 1 |
| 2022 | Person re-identification in indoor videos by information fusion using Graph Convolutional Networks
Komal Soni, Debi Prosad Dogra, Arif Ahmed 0002, Samarjit Kar, Heeseung Choi, Ig-Jae Kim |
Expert Syst. Appl. | 6 |
| 2022 | A multi-stream deep neural network with late fuzzy fusion for real-world anomaly detection
Thakare Kamalakar Vijay, Nitin Sharma 0004, Debi Prosad Dogra, Heeseung Choi, Ig-Jae Kim |
Expert Syst. Appl. | 5 |
| 2022 | Object Interaction-Based Localization and Description of Road Accident Events Using Deep LearningabstractDetection and localization of road accidents in real-time is an integral part of the Intelligent Transportation System (ITS). Even though the existing road accident detection methods show promising results, the process suffers from some drawbacks. For example, existing methods require a large number of sample videos for feature learning. Moreover, features such as temporal gradients or flow fields are time-consuming. To address these issues, we introduce a new method that uses objects and their positions to detect accidents in real-time. Apart from localization of the accident events in videos, we perform a high-level post processing to describe the severity and context of an accident. Firstly, we divide an input video into pre-accident, accident and post-accident stages to extract object interactions. These interaction proposals are then filtered using a refinement algorithm. We then adopt an iterative training procedure to classify normal and accident interactions. We also highlight the damaged zone using heat maps. Finally, we generate high-level textual descriptions to quantify the context and severity of an accident. We have trained the proposed model using offline setups. However, it can be deployed online to detect road accident events in real-time by taking the video inputs directly from the CCTV camera. Moreover, with a minimal supervision, the model can be retrained for online surveillance. Extensive experiments carried out on UCF Crime and CADP datasets reveal that the proposed framework achieves state-of-the-art performance when compared with the recently proposed accident event detection methods in terms of AUC (UCF Crime: 69.70% and CADP: 72.59%) and FAR (UCF Crime: 0.8 and CADP: 2.2). The high-level description of the accident is an added advantage that will certainly help the traffic police to react in a timely manner. Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, Ig-Jae Kim |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Cross-Domain Grouping and Alignment for Domain Adaptive Semantic SegmentationabstractExisting techniques to adapt semantic segmentation networks across source and target domains within deep convolutional neural networks (CNNs) deal with all the samples from the two domains in a global or category-aware manner. They do not consider an inter-class variation within the target domain itself or estimated category, providing the limitation to encode the domains having a multi-modal data distribution. To overcome this limitation, we introduce a learnable clustering module, and a novel domain adaptation framework, called cross-domain grouping and alignment. To cluster the samples across domains with an aim to maximize the domain alignment without forgetting precise segmentation ability on the source domain, we present two loss functions, in particular, for encouraging semantic consistency and orthogonality among the clusters. We also present a loss so as to solve a class imbalance problem, which is the other limitation of the previous methods. Our experiments show that our method consistently boosts the adaptation performance in semantic segmentation, outperforming the state-of-the-arts on various domain adaptation settings. Sunghun Joung, Seungryong Kim, Jungin Park, Ig-Jae Kim, Kwanghoon Sohn |
AAAI | 5 |
| 2021 | Prototype-Guided Saliency Feature Learning for Person SearchabstractExisting person search methods integrate person detection and re-identification (re-ID) module into a unified system. Though promising results have been achieved, the misalignment problem, which commonly occurs in person search, limits the discriminative feature representation for re-ID. To overcome this limitation, we introduce a novel framework to learn the discriminative representation by utilizing prototype in OIM loss. Unlike conventional methods using prototype as a representation of person identity, we utilize it as guidance to allow the attention network to consistently highlight multiple instances across different poses. Moreover, we propose a new prototype update scheme with adaptive momentum to increase the discriminative ability across different instances. Extensive ablation experiments demonstrate that our method can significantly enhance the feature discriminative power, outperforming the state-of-the-art results on two person search benchmarks including CUHK-SYSU and PRW. Hanjae Kim, Sunghun Joung, Ig-Jae Kim, Kwanghoon Sohn |
CVPR | 3 |
| 2021 | Learning Canonical 3D Object Representation for Fine-Grained RecognitionabstractWe propose a novel framework for fine-grained object recognition that learns to recover object variation in 3D space from a single image, trained on an image collection without using any ground-truth 3D annotation. We accomplish this by representing an object as a composition of 3D shape and its appearance, while eliminating the effect of camera viewpoint, in a canonical configuration. Unlike conventional methods modeling spatial variation in 2D images only, our method is capable of reconfiguring the appearance feature in a canonical 3D space, thus enabling the subsequent object classifier to be invariant under 3D geometric variation. Our representation also allows us to go beyond existing methods, by incorporating 3D shape variation as an additional cue for object recognition. To learn the model without ground-truth 3D annotation, we deploy a differentiable renderer in an analysis-by-synthesis frame- work. By incorporating 3D shape and appearance jointly in a deep representation, our method learns the discriminative representation of the object and achieves competitive performance on fine-grained image recognition and vehicle re-identification. We also demonstrate that the performance of 3D shape reconstruction is improved by learning fine-grained shape deformation in a boosting manner. Sunghun Joung, Seungryong Kim, Ig-Jae Kim, Kwanghoon Sohn |
ICCV | 4 |
| 2021 | A 3d Model-Based Approach For Fitting Masks To Faces In The WildabstractFace recognition now requires a large number of labelled masked face images in the era of this unprecedented COVID19 pandemic. Unfortunately, the rapid spread of the virus has left us little time to prepare for such dataset in the wild. To circumvent this issue, we present a 3D model-based approach called WearMask3D for augmenting face images of various poses to the masked face counterparts. Our method proceeds by first fitting a 3D morphable model on the input image, second overlaying the mask surface onto the face model and warping the respective mask texture, and last projecting the 3D mask back to 2D. The mask texture is adapted based on the brightness and resolution of the input image. By working in 3D, our method can produce more natural masked faces of diverse poses from a single mask texture. To compare precisely between different augmentation approaches, we have constructed a dataset comprising masked and unmasked faces with labels called MFW-mini. Experimental results demonstrate WearMask3D1produces more realistic masked faces, and utilizing these images for training leads to state-of-the-art recognition accuracy for masked faces. Je Hyeong Hong, Hanjo Kim, Gi Pyo Nam, Junghyun Cho, Hyeong-Seok Ko, Ig-Jae Kim |
ICIP | 7 |
| 2021 | Robot Facial Expression Framework for Enhancing Empathy in Human-Robot InteractionabstractA social robot interacts with humans based on social intelligence, for which related applications are being developed across diverse fields to be increasingly integrated in modern society. In this regard, social intelligence and interaction are the keywords of a social robot. Social intelligence refers to the ability to control interactions or thoughts and feelings of relationships with other people; primal empathy, which is the ability to empathize by perceiving emotional signals, among the components of social intelligence was applied to the robot in this study. We proposed that the empathic ability of a social robot can be improved if the social robot can create facial expressions based on the emotional state of a user. Moreover, we suggested a framework of facial expressions for robots. These facial expressions can be repeatedly used in various social robot platforms to achieve such a strategy. Ung Park, Youngeun Jang, GiJae Lee, KangGeon Kim, Ig-Jae Kim, Jongsuk Choi |
RO-MAN | 6 |
| 2021 | Digital Social Interaction in Older Adults During the COVID-19 PandemicabstractThroughout the COVID-19 pandemic, older adults have been encouraged to stay indoors and isolated, leading to potential disruptions in their social activities and interpersonal relationships. This interview study ($N=24$) provides a close examination of older adults' communication technology adoption and usage in light of the pandemic. Our interviews revealed that the pandemic motivated many older adults to learn new technology and become more tech-savvy in an effort to stay connected with others. However, older adults also reported challenges related to the pandemic that were major impediments to technology adoption. These were: (1) lack of access to in-person technology support under physical distancing mandates, (2) lack of opportunities for online participation due to negative age stereotypes and assumptions, and (3) increased apprehension to seek help from family members and friends who were suffering from pandemic-related stresses. This study extends technology adoption literature and contributes an up-to-date examination of the "grey digital divide" (the gap between older adults who use technology and those who do not). Our findings demonstrate that despite the rapidly increasing number of tech-savvy seniors, a digital divide not only persists, but has been exacerbated by the transition to virtual-only offerings. We reveal the challenges and coping strategies of older adults who remain separated from technology and propose actionable solutions to increase digital access during the COVID-19 pandemic and beyond. Frances Jihae Sin, Sophie Berger, Ig-Jae Kim, Dongwook Yoon |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2021 | Relational Deep Feature Learning for Heterogeneous Face RecognitionabstractHeterogeneous Face Recognition (HFR) is a task that matches faces across two different domains such as visible light (VIS), near-infrared (NIR), or the sketch domain. Due to the lack of databases, HFR methods usually exploit the pre-trained features on a large-scale visual database that contain general facial information. However, these pre-trained features cause performance degradation due to the texture discrepancy with the visual domain. With this motivation, we propose a graph-structured module called Relational Graph Module (RGM) that extracts global relational information in addition to general facial features. Because each identity's relational information between intra-facial parts is similar in any modality, the modeling relationship between features can help cross-domain matching. Through the RGM, relation propagation diminishes texture dependency without losing its advantages from the pre-trained features. Furthermore, the RGM captures global facial geometrics from locally correlated convolutional features to identify long-range relationships. In addition, we propose a Node Attention Unit (NAU) that performs node-wise recalibration to concentrate on the more informative nodes arising from relation-based propagation. Furthermore, we suggest a novel conditional-margin loss function ($C$ -softmax) for the efficient projection learning of the embedding vector in HFR. The proposed method outperforms other state-of-the-art methods on five HFR databases. Furthermore, we demonstrate performance improvement on three backbones because our module can be plugged into any pre-trained face recognition backbone to overcome the limitations of a small HFR database. MyeongAh Cho, Taeoh Kim, Ig-Jae Kim, Kyungjae Lee 0003, Sangyoun Lee |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Scene Adaptive Online Surveillance Video Synopsis via Dynamic Tube Rearrangement Using OctreeabstractVisual surveillance produces a significant amount of raw video data that can be time consuming to browse and analyze. In this work, we present a video synopsis methodology called "scene adaptive online video synopsis via dynamic tube rearrangement using octree (SSOcT)" that can effectively condense input surveillance videos. Our method entailed summarizing the input video by analyzing scene characteristics and determining an effective spatio-temporal 3D structure for video synopsis. For this purpose, we first analyzed the attributes of each extracted tube with respect to scene geometry and complexity. Then, we adaptively grouped the tubes using an online grouping algorithm that exploits these scene characteristics. Finally, the tube groups were dynamically rearranged using the proposed octree-based algorithm that efficiently inserted and refined tubes containing high spatio-temporal movements in real time. Extensive video synopsis experimental results are provided, demonstrating the effectiveness and efficiency of our method in summarizing real-world surveillance videos with diverse scene characteristics. Yoonsik Yang, Haksub Kim, Heeseung Choi, Seungho Chae, Ig-Jae Kim |
IEEE Trans. Image Process. | 5 |
| 2021 | Planar Abstraction and Inverse Rendering of 3D Indoor EnvironmentsabstractScanning and acquiring a 3D indoor environment suffers from complex occlusions and misalignment errors. The reconstruction obtained from an RGB-D scanner contains holes in geometry and ghosting in texture. These are easily noticeable and cannot be considered as visually compelling VR content without further processing. On the other hand, the well-known Manhattan World priors successfully recreate relatively simple structures. In this article, we would like to push the limit of planar representation in indoor environments. Given an initial 3D reconstruction captured by an RGB-D sensor, we use planes not only to represent the environment geometrically but also to solve an inverse rendering problem considering texture and light. The complex process of shape inference and intrinsic imaging is greatly simplified with the help of detected planes and yet produces a realistic 3D indoor environment. The generated content can adequately represent the spatial arrangements for various AR/VR applications and can be readily composited with virtual objects possessing plausible lighting and texture. Young Min Kim 0001, Sangwoo Ryu, Ig-Jae Kim |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | Cylindrical Convolutional Networks for Joint Object Detection and Viewpoint EstimationabstractExisting techniques to encode spatial invariance within deep convolutional neural networks only model 2D transformation fields. This does not account for the fact that objects in a 2D space are a projection of 3D ones, and thus they have limited ability to severe object viewpoint changes. To overcome this limitation, we introduce a learnable module, cylindrical convolutional networks (CCNs), that exploit cylindrical representation of a convolutional kernel defined in the 3D space. CCNs extract a view-specific feature through a view-specific convolutional kernel to predict object category scores at each viewpoint. With the view-specific feature, we simultaneously determine objective category and viewpoints using the proposed sinusoidal soft-argmax module. Our experiments demonstrate the effectiveness of the cylindrical convolutional networks on joint object detection and viewpoint estimation. Sunghun Joung, Seungryong Kim, Hanjae Kim, Ig-Jae Kim, Junghyun Cho, Kwanghoon Sohn |
CVPR | 5 |
| 2020 | SumGraph: Video Summarization via Recursive Graph Modeling
Jungin Park, Jiyoung Lee 0005, Ig-Jae Kim, Kwanghoon Sohn |
ECCV (25) | 3 |
| 2020 | Shape-Adaptive Kernel Network for Dense Object DetectionabstractDense object detectors that are applied over a regular, dense grid have advanced and drawn their attention in recent days. Their fully convolutional nature greatly advances the computational efficiency of object detectors compared to the two-stage detectors. However, the lack of the ability to adjust shape variation on a regular grid is still limited. In this paper we introduce a new framework, shape-adaptive kernel network, to handle spatial manipulation of input data in convolutional kernel space. At the heart of out approach is to align the original kernel space recovering shape variation of each input feature on regular grid. To this end, we propose a shape-adaptive kernel sampler to adjust dynamic convolutional kernel conditioned on input. To increase the flexibility of geometric transformation, a cascade refinement module is designed, which first estimates the global transformation grid and then estimates local offset in convolutional kernel space. Our experiments demonstrate the effectiveness of the shape-adaptive kernel network for dense object detection on various benchmarks. Hanjae Kim, Sunghun Joung, Ig-Jae Kim, Kwanghoon Sohn |
ICIP | 3 |
| 2020 | Person Re-identification in Videos by Analyzing Spatio-temporal TubesabstractAbstract Typical person re-identification frameworks search for k best matches in a gallery of images that are often collected in varying conditions. The gallery usually contains image sequences for video re-identification applications. However, such a process is time consuming as video re-identification involves carrying out the matching process multiple times. In this paper, we propose a new method that extracts spatio-temporal frame sequences or tubes of moving persons and performs the re-identification in quick time. Initially, we apply a binary classifier to remove noisy images from the input query tube. In the next step, we use a key-pose detection-based query minimization technique. Finally, a hierarchical re-identification framework is proposed and used to rank the output tubes. Experiments with publicly available video re-identification datasets reveal that our framework is better than existing methods. It ranks the tubes with an average increase in the CMC accuracy of 6-8% across multiple datasets. Also, our method significantly reduces the number of false positives. A new video re-identification dataset, named Tube-based Re-identification Video Dataset (TRiViD), has been prepared with an aim to help the re-identification research community. Arif Ahmed 0002, Debi Prosad Dogra, Heeseung Choi, Seungho Chae, Ig-Jae Kim |
Multim. Tools Appl. | 5 |
| 2020 | Memetic algorithm for multivariate time-series segmentation
Hyunki Lim, Heeseung Choi, Yeji Choi, Ig-Jae Kim |
Pattern Recognit. Lett. | 4 |
| 2020 | Query-Based Video Synopsis for Intelligent Traffic Monitoring ApplicationsabstractSynopsis of a long-duration video has many applications in intelligent transportation systems. It can help to monitor traffic with lesser manpower. However, generating meaningful synopsis of a long-duration video recording can be challenging. Often summarized outputs include redundant contents or activities that may not be helpful to the observer. Moving object trajectories are possible sources of information that can be used to generate the synopsis of long-duration videos. The synopsis generation faces challenges due to object tracking, grouping of the trajectories with respect to activity type, object category, and contextual information, and generating smooth synopsis according to a query. In this paper, we propose a method to generate meaningful and smooth synopsis of long-duration videos according to the users' query. We have tracked moving objects and adopted deep learning to classify the objects into known categories (e.g., car, bike, and pedestrians). We then identify regions in the surveillance scene with the help of unsupervised clustering. Each tube (spatiotemporal object trajectory) is represented by the source and the destination. In the final stage, we take a query from the user and generate the synopsis video by smoothly blending the appropriate tubes over the background frame through energy minimization. The proposed method has been evaluated on two publicly available datasets and our own surveillance datasets. We have compared the method with popular state-of-the-art techniques. The experiments reveal that the proposed method is superior to the existing techniques and it produces visually seamless video synopsis. Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Renuka Patnaik, Seung-Cheol Lee, Heeseung Choi, Gi Pyo Nam, Ig-Jae Kim |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2019 | Cancelable fingerprint template design with randomized non-negative least squares
Jun Beom Kho, Jaihie Kim, Ig-Jae Kim, Andrew Beng Jin Teoh |
Pattern Recognit. | 3 |
| 2017 | Age face simulation using aging functions on global and local features with residual images
Sung Eun Choi, Jaeik Jo, Sanghak Lee, Heeseung Choi, Ig-Jae Kim, Jaihie Kim |
Expert Syst. Appl. | 5 |
| 2017 | Face alignment using a deep neural network with local feature learning and recurrent regression
Byung-Hwa Park, Se-Young Oh, Ig-Jae Kim |
Expert Syst. Appl. | 3 |
| 2015 | Single-view-based 3D facial reconstruction method robust against pose variations
Jaeik Jo, Heeseung Choi, Ig-Jae Kim, Jaihie Kim |
Pattern Recognit. | 3 |
| 2014 | Detecting driver drowsiness using feature-level fusion and user-specific classification
Jaeik Jo, Sung Joo Lee, Kang Ryoung Park, Ig-Jae Kim, Jaihie Kim |
Expert Syst. Appl. | 4 |
| 2012 | Web image-based super-resolution
Jongho Lee 0004, Sang Chul Ahn, Hwasup Lim, Ig-Jae Kim, Hyoung-Gon Kim |
ICPR | 4 |
| 2011 | Back Talk: An auditory environment for sociable television viewingabstractVideo content is being consumed in a host of new ways-viewers are no longer restricted to same-time or same-place viewing. However, the experience of watching with a group is inherently social and often desirable despite the physical distribution of group members. This paper introduces Back Talk, a system designed to create a sociable television watching experience. We enhance TV viewing with an auditory environment around a listener. We have explored and leveraged the richness of audio to convey presence of remote viewers. We have developed a novel framework for capturing and translating engagement of an individual into a set of audio cues that are played spatially around a listener. Such auditory enhancements can augment video content consumption in the future. Andrea Colaco, Ig-Jae Kim, Chris Schmandt |
CCNC | 2 |
| 2011 | Indoor location sensing using geo-magnetismabstractWe present an indoor positioning system that measures location using disturbances of the Earth's magnetic field caused by structural steel elements in a building. The presence of these large steel members warps the geomagnetic field in a way that is spatially varying but temporally stable. To localize, we measure the magnetic field using an array of e-compasses and compare the measurement with a previously obtained magnetic map. We demonstrate accuracy within 1 meter 88% of the time in experiments in two buildings and across multiple floors within the buildings. We discuss several constraint techniques that can maintain accuracy as the sample space increases. Jaewoo Chung, Matt Donahoe, Chris Schmandt, Ig-Jae Kim, Pedram Razavai, Micaela Wiseman |
MobiSys | 4 |
| 2011 | Highlighted depth-of-field photography: Shining light on focusabstractWe present a photographic method to enhance intensity differences between objects at varying distances from the focal plane. By combining a unique capture procedure with simple image processing techniques, the detected brightness of an object is decreased proportional to its degree of defocus. A camera-projector system casts distinct grid patterns onto a scene to generate a spatial distribution of point reflections. These point reflections relay a relative measure of defocus that is utilized in postprocessing to generate a highlighted DOF photograph. Trade-offs between three different projector-processing pairs are analyzed, and a model is developed to help describe a new intensity-dependent depth of field that is controlled by the pattern of illumination. Results are presented for a primary single snapshot design as well as a scanning method and a comparison method. As an application, automatic matting results are presented. Roarke Horstmeyer, Ig-Jae Kim, Ramesh Raskar |
ACM Trans. Graph. | 3 |
| 2010 | Introduction to augmented reality and its applicationsabstractThis course introduces and defines augmented reality and user interfaces that can apply AR to enhance users' perceptions of reality. Ig-Jae Kim |
SIGGRAPH ASIA (Courses) | 1 |
| 2009 | MART-MAF: Media File Format for AR Tour Guide ServiceabstractWe are currently developing a new tour guide service on mobile phone using augmented reality. Among the components for the service, in-situ annotation is one of the most important components for capturing of lots of personal experiences in the form of digital multimedia during his/her trip. Our in-situ annotation system provides several pieces of functionality that include recording MART (Mobile Augmented Reality based Tour) media format, transmitting the recorded media and retrieving the related experiences of others. MART media data should be easy for users to store, interchange and manage. Additionally, it can integrate various types of media, such as video, audio, GPS data, motion, 2D/3D graphics and annotated texts. For this, we define a new format for MART media, called MART-MAF, which can compose MPEG standards and non-MPEG standards for various kinds of media. Ig-Jae Kim, Injun Song, Jane Hwang, Sang Chul Ahn, Hyoung-Gon Kim, Heedong Ko |
ISM | 1 |
| 2008 | Adaptive Modeling of a User's Daily Life with a Wearable Sensor NetworkabstractIn an environment where the contexts of users are complex and the degree of freedom of user activity is very high, such as in daily life, several factors need to be considered for constructing user models. Such a model should include changes in the meanings of activities that reflect the user's situation both temporally and individually. In this paper we propose a novel approach for personalizing the user model and adapting it to individual circumstances with a wearable sensor network. We also describe the process for determining the repetitive activities of a user by using incremental clustering and Bayesian network. We show experimental results for an adaptive user model based on a real wearable sensor platform. Multimedia data of user experience are acquired from the multimodal sensors, and processed to metadata that have meanings. Hyoungnyoun Kim, Ig-Jae Kim, Hyoung-Gon Kim, Ji-Hyung Park |
ISM | 2 |
| 2007 | Design and Implementation of A Mobile and Portable Lifelog Media System
Baud Haryo Prananto, Ig-Jae Kim, Hyoung-Gon Kim |
MoMM | 2 |
| 2007 | 3D Lip-Synch Generation with Data-Faithful Machine LearningabstractAbstract This paper proposes a new technique for generating three‐dimensional speech animation. The proposed technique takes advantage of both data‐driven and machine learning approaches. It seeks to utilize the most relevant part of the captured utterances for the synthesis of input phoneme sequences. If highly relevant data are missing or lacking, then it utilizes less relevant (but more abundant) data and relies more heavily on machine learning for the lip‐synch generation. This hybrid approach produces results that are more faithful to real data than conventional machine learning approaches, while being better able to handle incompleteness or redundancy in the database than conventional data‐driven approaches. Experimental results, obtained by applying the proposed technique to the utterance of various words and phrases, show that (1) the proposed technique generates lip‐synchs of different qualities depending on the availability of the data, and (2) the new technique produces more realistic results than conventional machine learning approaches. Ig-Jae Kim, Hyeong-Seok Ko |
Comput. Graph. Forum | 1 |
| 2006 | Video Surveillance using Dynamic Configuration of Mutiple Active CamerasabstractIn this paper, we present a coordinated video surveillance system that can minimize the spatial limitation and can precisely extract the 3D position of objects. To do this, our system used an agent based system and also tracked the normalized object using active wide-baseline stereo method. The system is composed of two parts: multiple camera agents (CAs) and a support module (SM). Each CA treats image processing and camera controlling. A SM performs a role that manages communication between CAs. Our proposed system extracts object positions independent of environment via the collaboration of CAs and a SM. Finally, through experimental results we show that the proposed system successfully tracks an object on real-time. Nyoun Kim, Ig-Jae Kim, Hyoung-Gon Kim |
ICIP | 2 |
| 2002 | Stereo vision based 3D input deviceabstractThis paper concerns extracting 3D motion information from a 3D input device in real time focused to enabling effective human-computer. interaction. In particular, we develop a novel algorithm for extracting 6 degrees-of-freedom motion information from a 3D input device by employing an epipolar geometry of stereo camera, color, motion, and structure information, free from requiring the aid of camera calibration object. To extract 3D motion, we first determine the epipolar geometry of stereo camera by computing the perspective projection matrix and perspective distortion matrix. We then incorporate the proposed “Motion Adaptive Weighted Unmatched Pixel Count” algorithm performing color transformation, unmatched pixel counting, discrete Kalman filtering, and principal component analysis. The extracted 3D motion information can be applied to controlling virtual objects or aiding the navigation device that controls the viewpoint of a user in virtual reality setting. Since the stereo vision-based 3D input device is wireless, it provides users with a means for more natural and efficient interface, thus effectively realizing a feeling of immersion. SangMin Yoon, Ig-Jae Kim, Sang Chul Ahn, Heedong Ko, Hyoung-Gon Kim |
ICASSP | 2 |
| 2002 | The Making of Kyongju VR TheatreabstractRecently we have built the largest Virtual Reality (VR) theatre in the world for the Kyongju World Culture EXPO 2000. Unlike single user VR systems, the VR theatre is characterized by a single shared screen and controlled by a kind of tightly coupled user inputs from several hundreds of people in the audience. The large computer-generated stereo images by the huge cylindrical screen provide the immersive feeling augmenting the physical audience space with of 3D virtual space. In addition to the visual immersion, the theatre provides 3D audio, vibration and olfactory display as well as keypads for the audience in their seats interactively controlling the virtual environment. This paper introduces the issues raised and addressed during the design of making such a versatile VR theatre, production and presentation of the virtual heritage at Kyongju, one thousand years ago. ChangHoon Park, Heedong Ko, Ig-Jae Kim, Sang Chul Ahn, Yong-Moo Kwon, Hyoung-Gon Kim |
VR | 3 |
| 2001 | 3D tracking of multi-objects using color and stereo for HCIabstractWe present a 3D tracking method of multi-objects by color and stereo. The results are applied to navigation and manipulation in a virtual environment. We choose a human face and hands as tracking targets for HCI. To extract the area of face and hands in a complex background, we transform an input color image using the GSCD (generalized skin color distribution). Based on the transformed image, we detect the area of face and hands by WUPC. Furthermore, the face and hands candidate region are estimated by a Kalman filter in the following frame. Then we can extract the depth information about the only corresponding region, respectively, by stereo matching. Finally, we can navigate and manipulate objects based on the extracted 3D information of face and hands naturally in the virtual environment without any additional device. Ig-Jae Kim, Shwan Lee, Sang Chul Ahn, Yong-Moo Kwon, Hyoung-Gon Kim |
ICIP (3) | 1 |
| 2001 | Audience interaction for virtual reality theater and its implementationabstractRecently we have built a VR(Virtual Reality) theater in Kyongju, Korea. It combines the advantages of VR and IMAX theater. The VR theater can be characterized by a single shared screen and by multiple inputs from several hundreds of people. In this case, multi-user interaction is different from that of networked VR systems and must be reconsidered. This paper defines the multi-user interaction in such a VR theater as Audience Interaction, and discusses key issues for the implementation of the Audience Interaction. This paper also presents a real implementation example in the Kyongju VR theater. Sang Chul Ahn, Ig-Jae Kim, Hyoung-Gon Kim, Yong-Moo Kwon, Heedong Ko |
VRST | 2 |
| 2000 | Automatic FDP/FAP generation from an image sequenceabstractThis paper presents an automatic FDP (Facial Definition Parameters) and FAP (Facial Animation Parameters) generation method from an image sequence that captures a frontal face. The proposed method is based on facial feature tracking without markers on a face. We present an efficient method to extract 2D facial features and to generate the FDP by applying 2D features to a generic face model. We also propose a template matching based FAP generation method. The advantage of this approach is that it can be easily applied to single camera MPEG-4 SNHC encoding systems. Munjae Song, Ig-Jae Kim, Yong-Moo Kwon, Hyoung-Gon Kim, Sang Chul Ahn |
ISCAS | 3 |
| 2000 | Web-based 3D media information systemabstractThis paper introduces web-based 3D media information system. We first address two promising 3D modeling techniques, i.e., image-based 3D modeling and laser scanning based 3D modeling. Especially, we present two approaches of the image-based 3D modeling. One is an off-line approach using multiview images which is captured with single camera and a robot arm. The another one is an on-line approach that extends a commercial triclops camera system. We also utilize a 3D modeling scheme based on a laser scanner and a 3D reverse modeler. Using our 3D modeling environments, we construct several kinds of 3D models. We also implement web-based 3D media information management and retrieval system using XML data server, which provides services of 3D models. Our web-based 3D media information system has a goal of services of various types of 3D models and contents through WWW, which is currently focused on the development and management of 3D models of Korea cultural heritage. Yong-Moo Kwon, Ig-Jae Kim, Sang Chul Ahn, Hyoung-Gon Kim |
VRST | 2 |