EDBT 2026 Demo / reviewers in the wild / expert
Sangmin Lee 0001
dblp:68/311-1
· DBLP profile ↗
36ranked-venue papers
11as first author
22since 2021 · last 2026
0000-0002-8606-2549ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 10 first-author · 18 since 2021Artificial intelligence and machine learning · 18 · 6 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unsupervised domain adaptation for medical image segmentation using adaptogen-perturbationabstractDomains shift originated from differences in devices or patients in the medical field, poses a significant challenge when applying pre-trained models to clinical applications. To tackle this challenge, domain adaptation methods have been explored. However, most existing methods are designed for a single target domain adaptation or require sharing all target domain data for adaptation, which is infeasible in the medical field due to privacy issues. In this paper, we propose a novel unsupervised multi-target domain adaptation method without requiring data sharing. To this end, we introduce an additional signal, termed Adaptogen-Perturbation (AP) optimized to bridge the gap between the source and target domains. The optimized AP is injected into the latent feature and facilitates the adaptation of the pre-trained model to the target domain. Moreover, we propose a Spectral/Geometric Consistency learning framework to optimize the AP in an unsupervised manner. This promotes consistent predictions across two types of transformations: geometric and frequency-space spectral transformations, enhancing robustness to both variations. Extensive experiments with multiple medical segmentation datasets demonstrate the effectiveness of APs. Hong Joo Lee 0001, Yuan Bi, Sangmin Lee 0001, Gyeong-Moon Park, Jung Uk Kim, Seong Tae Kim 0001, Zhongliang Jiang, Nassir Navab |
Medical Image Anal. | 3 |
| 2025 | Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight DetectionabstractThe goal of video moment retrieval and highlight detection is to identify specific segments and highlights based on a given text query. With the rapid growth of video content and the overlap between these tasks, recent works have addressed both simultaneously. However, they still struggle to fully capture the overall video context, making it challenging to determine which words are most relevant. In this paper, we present a novel Video Context-aware Keyword Attention module that overcomes this limitation by capturing keyword variation within the context of the entire video. To achieve this, we introduce a video context clustering module that provides concise representations of the overall video context, thereby enhancing the understanding of keyword dynamics. Furthermore, we propose a keyword weight detection module with keyword-aware contrastive learning that incorporates keyword information to enhance fine-grained alignment between visual and textual features. Extensive experiments on the QVHighlights, TVSum, and Charades-STA benchmarks demonstrate that our proposed method significantly improves performance in moment retrieval and highlight detection tasks compared to existing approaches. Sung Jin Um, Sangmin Lee 0001, Jung Uk Kim |
AAAI | 3 |
| 2025 | SocialGesture: Delving into Multi-person Gesture UnderstandingabstractPrevious research in human gesture recognition has largely overlooked multi-person interactions, which are crucial for understanding the social context of naturally occurring gestures. This limitation in existing datasets presents a significant challenge in aligning human gestures with other modalities like language and speech. To address this issue, we introduce SocialGesture, the first large-scale dataset specifically designed for multi-person gesture analysis. SocialGesture features a diverse range of natural scenarios and supports multiple gesture analysis tasks, including video-based recognition and temporal localization, providing a valuable resource for advancing the study of gesture during complex social interactions. Furthermore, we propose a novel visual question answering (VQA) task to benchmark vision language models’ (VLMs) performance on social gesture understanding. Our findings highlight several limitations of current gesture recognition models, offering insights into future directions for improvement in this field. SocialGesture is available at hugging-face.co/datasets/IrohXu/SocialGesture. Pranav Virupaksha, Wenqi Jia 0001, Bolin Lai, Fiona Ryan, Sangmin Lee 0001, James M. Rehg |
CVPR | 6 |
| 2025 | Question-Aware Gaussian Experts for Audio-Visual Question AnsweringabstractAudio-Visual Question Answering (AVQA) requires not only question-based multimodal reasoning but also precise temporal grounding to capture subtle dynamics for accurate prediction. However, existing methods mainly use question information implicitly, limiting focus on question-specific details. Furthermore, most studies rely on uniform frame sampling, which can miss key question-relevant frames. Although recent Top-K frame selection methods aim to address this, their discrete nature still overlooks fine-grained temporal details. This paper proposes QA-TIGER, a novel framework that explicitly incorporates question information and models continuous temporal dynamics. Our key idea is to use Gaussian-based modeling to adaptively focus on both consecutive and non-consecutive frames based on the question, while explicitly injecting question information and applying progressive refinement. We leverage a Mixture of Experts (MoE) to flexibly implement multiple Gaussian models, activating temporal experts specifically tailored to the question. Extensive experiments on multiple AVQA benchmarks show that QA-TIGER consistently achieves state-of-the-art performance. Code is available at https://aim-skku.github.io/QA-TIGER/ Hongyeob Kim, Inyoung Jung, Dayoon Suh, Youjia Zhang, Sangmin Lee 0001, Sungeun Hong |
CVPR | 5 |
| 2025 | Unleashing In-context Learning of Autoregressive Models for Few-shot Image ManipulationabstractText-guided image manipulation has experienced notable advancement in recent years. In order to mitigate linguistic ambiguity, few-shot learning with visual examples has been applied for instructions that are underrepresented in the training set, or difficult to describe purely in language. However, learning from visual prompts requires strong reasoning capability, which diffusion models are struggling with. To address this issue, we introduce a novel multi-modal autoregressive model, dubbed InstaManip, that can instantly learn a new image manipulation operation from textual and visual guidance via in-context learning, and apply it to new query images. Specifically, we propose an innovative group self-attention mechanism to break down the in-context learning process into two separate stages – learning and applying, which simplifies the complex problem into two easier tasks. We also introduce a relation regularization method to further disentangle image transformation features from irrelevant contents in exemplar images. Extensive experiments suggest that our method surpasses previous few-shot image manipulation models by a notable margin (≥19% in human evaluation). We also find our model can be further boosted by increasing the number or diversity of exemplar images. Please check out our project page (https://bolinlai.github.io/projects/InstaManip/). Bolin Lai, Felix Juefei-Xu, Miao Liu 0007, Xiaoliang Dai, Nikhil Mehta 0002, Zeyi Huang, James M. Rehg, Sangmin Lee 0001, Tong Xiao 0003 |
CVPR | 9 |
| 2025 | Gaze-LLE: Gaze Target Estimation via Large-Scale Learned EncodersabstractWe address the problem of gaze target estimation, which aims to predict where a person is looking in a scene. Predicting a person’s gaze target requires reasoning both about the person’s appearance and the contents of the scene. Prior works have developed increasingly complex, handcrafted pipelines for gaze target estimation that carefully fuse features from separate scene encoders, head encoders, and auxiliary models for signals like depth and pose. Motivated by the success of general-purpose feature extractors on a variety of visual tasks, we propose Gaze-LLE, a novel transformer framework that streamlines gaze target estimation by leveraging features from a frozen DINOv2 encoder. We extract a single feature representation for the scene, and apply a person-specific positional prompt to decode gaze with a lightweight module. We demonstrate state-of-the-art performance across several gaze benchmarks and provide extensive analysis to validate our design choices. Our code and models are available at: http://github.com/fkryan/gazelle. Fiona Ryan, Ajay Bati, Sangmin Lee 0001, Daniel Bolya, Judy Hoffman, James M. Rehg |
CVPR | 3 |
| 2025 | Object-aware Sound Source Localization via Audio-Visual Scene UnderstandingabstractAudio-visual sound source localization task aims to spatially localize sound-making objects within visual scenes by integrating visual and audio cues. However, existing methods struggle with accurately localizing sound-making objects in complex scenes, particularly when visually similar silent objects coexist. This limitation arises primarily from their reliance on simple audio-visual correspondence, which does not capture fine-grained semantic differences between sound-making and silent objects. To address these challenges, we propose a novel sound source localization framework leveraging Multimodal Large Language Models (MLLMs) to generate detailed contextual information that explicitly distinguishes between sound-making foreground objects and silent background objects. To effectively integrate this detailed information, we introduce two novel loss functions: Object-aware Contrastive Alignment (OCA) loss and Object Region Isolation (ORI) loss. Extensive experimental results on MUSIC and VGGSound datasets demonstrate the effectiveness of our approach, significantly outperforming existing methods in both single-source and multi-source localization scenarios. Code and generated detailed contextual information are available at: https://github.com/VisualAIKHU/OA-SSL. Sung Jin Um, Sangmin Lee 0001, Jung Uk Kim |
CVPR | 3 |
| 2025 | Toward Human Deictic Gesture Target EstimationabstractHumans have a remarkable ability to use co-speech deictic gestures, such as pointing and showing, to enrich verbal communication and support social interaction. These gestures are so fundamental that infants begin to use them even before they acquire spoken language, which highlights their central role in human communication. Understanding the intended targets of another individual's deictic gestures enables inference of their intentions, comprehension of their current actions, and prediction of upcoming behaviors. Despite its significance, gesture target estimation remains an underexplored task within the computer vision community. In this paper, we introduce GestureTarget, a novel task designed specifically for comprehensive evaluation of social deictic gesture semantic target estimation. To address this task, we propose TransGesture, a set of Transformer-based gesture target prediction models. Given an input image and the spatial location of a person, our models predict the intended target of their gesture within the scene. Critically, our gaze-aware joint cross attention fusion model demonstrates how incorporating gaze-following cues significantly improves gesture target mask prediction IoU by 6% and gesture existence prediction accuracy by 10%. Our results underscore the complexity and importance of integrating gaze cues into deictic gesture intention understanding, advocating for increased research attention to this emerging area. All data, code will be made publicly available upon acceptance. Code of TransGesture is available at GitHub.com/IrohXu/TransGesture. Pranav Virupaksha, Sangmin Lee 0001, Bolin Lai, Wenqi Jia 0001, Jintai Chen, James M. Rehg |
NeurIPS | 3 |
| 2024 | Learning to Visually Localize Sound Sources from Mixtures without Prior Source KnowledgeabstractThe goal of the multi-sound source localization task is to localize sound sources from the mixture individu-ally. While recent multi-sound source localization meth-ods have shown improved performance, they face chal-lenges due to their reliance on prior information about the number of objects to be separated. In this paper, to overcome this limitation, we present a novel multi-sound source localization method that can perform localization without prior knowledge of the number of sound sources. To achieve this goal, we propose an iterative object iden-tification (101) module, which can recognize sound-making objects in an iterative manner. After finding the regions of sound-making objects, we devise object similarity-aware clustering (OSC) loss to guide the 101 module to effectively combine regions of the same object but also dis-tinguish between different objects and backgrounds. It enables our method to perform accurate localization of sound-making objects without any prior knowledge. Exten-sive experimental results on the MUSIC and VGGSound benchmarks show the significant performance improve-ments of the proposed method over the existing methods for both single and multi-source. Our code is available at: https://github.comNisuaIAIKHUINoPrior_MultiSSL. Sung Jin Um, Sangmin Lee 0001, Jung Uk Kim |
CVPR | 3 |
| 2024 | Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned RepresentationsabstractUnderstanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However, most prior works on multimodal social cues focus predominantly on single-person behav-iors or rely on holistic visual representations that are not aligned to utterances in multi-party environments. Conse-quently, they are limited in modeling the intricate dynam-ics of multi-party interactions. In this paper, we introduce three new challenging tasks to model the fine-grained dy-namics between multiple people: speaking target identification, pronoun coreference resolution, and mentioned player prediction. We contribute extensive data annotations to cu-rate these new challenges in social deduction game settings. Furthermore, we propose a novel multimodal baseline that leverages densely aligned language-visual representations by synchronizing visual features with their corresponding utterances. This facilitates concurrently capturing verbal and non-verbal cues pertinent to social reasoning. Exper-iments demonstrate the effectiveness of the proposed approach with densely aligned multimodal representations in modeling fine-grained social interactions. Project website: https://sangmin-git.github.iolprojectslMMSI. Sangmin Lee 0001, Bolin Lai, Fiona Ryan, Bikram Boote, James M. Rehg |
CVPR | 1 |
| 2024 | Analyzing Visible Articulatory Movements in Speech Production For Speech-Driven 3D Facial AnimationabstractSpeech-driven 3D facial animation aims to generate realistic facial meshes based on input speech signals. However, due to a lack of understanding of visible articulatory movements, current state-of-the-art methods result in inaccurate lip and jaw movements. Traditional evaluation metrics, such as lip vertex error (LVE), often fail to represent the quality of visual results. Based on our observation, we reveal the problems with existing evaluation metrics and raise the necessity for separate evaluation approaches for 3D axes. Comprehensive analysis shows that most recent methods struggle to precisely predict lip and jaw movements in 3D space. Hyung Kyu Kim, Sangmin Lee 0001, Hak Gu Kim |
ICIP | 2 |
| 2024 | Text-guided distillation learning to diversify video embeddings for text-video retrieval
Sangmin Lee 0001, Hyungil Kim, Yong Man Ro |
Pattern Recognit. | 1 |
| 2022 | Weakly Paired Associative Learning for Sound and Image Representations via Bimodal Associative MemoryabstractData representation learning without labels has attracted increasing attention due to its nature that does not require human annotation. Recently, representation learning has been extended to bimodal data, especially sound and image which are closely related to basic human senses. Existing sound and image representation learning methods necessarily require a large number of sound and image with corresponding pairs. Therefore, it is difficult to ensure the effectiveness of the methods in the weakly paired condition, which lacks paired bimodal data. In fact, according to human cognitive studies, the cognitive functions in the human brain for a certain modality can be enhanced by receiving other modalities, even not directly paired ones. Based on the observation, we propose a new problem to deal with the weakly paired condition: How to boost a certain modal representation even by using other unpaired modal data. To address the issue, we introduce a novel bimodal associative memory (BMA-Memory) with key-value switching. It enables to build sound-image association with small paired bimodal data and to boost the built association with the eas-ily obtainable large amount of unpaired data. Through the proposed associative learning, it is possible to reinforce the representation of a certain modality (e.g., sound) even by using other unpaired modal data (e.g., images). Sangmin Lee 0001, Hyungil Kim, Yong Man Ro |
CVPR | 1 |
| 2022 | Audio-Visual Mismatch-Aware Video Retrieval via Association and Adjustment
Sangmin Lee 0001, Sungjune Park, Yong Man Ro |
ECCV (14) | 1 |
| 2022 | IVIST: Interactive Video Search Tool in VBS 2022
Sangmin Lee 0001, Sungjune Park, Yong Man Ro |
MMM (2) | 1 |
| 2022 | On-the-Fly Facial Expression Prediction Using LSTM Encoded Appearance-Suppressed DynamicsabstractEncoding the facial expression dynamics is efficient in classifying and recognizing facial expressions. Most facial dynamics-based methods assume that a sequence is temporally segmented before prediction. This requires the prediction to wait until a full sequence is available, resulting in prediction delay. To reduce the prediction delay and enable prediction “on-the-fly” (as frames are fed to the system), we propose new dynamics feature learning method that allows prediction with partial (incomplete) sequences. The proposed method utilizes the readiness of recurrent neural networks (RNNs) for on-the-fly prediction, and introduces novel learning constraints to induce early prediction with partial sequences. We further show that a delay in accurate prediction using RNNs could originate from the effect that the subject appearance has on the spatio-temporal features encoded by the RNN. We refer to that effect as “appearance bias”. We propose the appearance suppressed dynamics feature, which utilizes a static sequence to suppress the appearance bias. Experimental results have shown that the proposed method achieved higher recognition rates compared to the state-of-the-art methods on publicly available datasets. The results also verified that the proposed method improved on-the-fly prediction at subtle expression frames early in the sequence, using partial sequence inputs. Wissam J. Baddar, Sangmin Lee 0001, Yong Man Ro |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Assessing Individual VR Sickness Through Deep Feature Fusion of VR Video and Physiological ResponseabstractRecently, VR sickness assessment for VR videos is highly demanded in industry and research fields to address VR viewing safety issues. Especially, it is difficult to evaluate VR sickness of individuals due to individual differences. To achieve the challenging goal, we focus on deep feature fusion of sickness-related information. In this paper, we propose a novel deep learning-based assessment framework which estimates VR sickness of individual viewers with VR videos and corresponding physiological responses. We design the content stimulus guider imitating the phenomenon that humans feel VR sickness. The content stimulus guider extracts a deep stimulus feature from a VR video to reflect VR sickness caused by VR videos. In addition, we devise the physiological response guider to encode physiological responses that are acquired while humans experience VR videos. Each physiology sickness feature extractor (EEG, ECG, and GSR) in the physiological response guider is designed to suit their physiological characteristics. Extracted physiology sickness features are then fused into a deep physiology feature that comprehensively reflects individual deviations of VR sickness. Finally, the VR sickness predictor assesses individual VR sickness effectively with the fusion of the deep stimulus feature and the deep physiology feature. To validate the proposed method extensively, we built two benchmark datasets which contain 360-degree VR videos with physiological responses (EEG, ECG, and GSR) and SSQ scores. Experimental results show that the proposed method achieves meaningful correlations with human SSQ scores. Further, we validate the effectiveness of the proposed network designs by conducting analysis on feature fusion and visualization. Sangmin Lee 0001, Seongyeop Kim, Hak Gu Kim, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Towards a Better Understanding of VR Sickness: Physical Symptom Prediction for VR ContentsabstractWe address the black-box issue of VR sickness assessment (VRSA) by evaluating the level of physical symptoms of VR sickness. For the VR contents inducing the similar VR sickness level, the physical symptoms can vary depending on the characteristics of the contents. Most of existing VRSA methods focused on assessing the overall VR sickness score. To make better understanding of VR sickness, it is required to predict and provide the level of major symptoms of VR sickness rather than overall degree of VR sickness. In this paper, we predict the degrees of main physical symptoms affecting the overall degree of VR sickness, which are disorientation, nausea, and oculomotor. In addition, we introduce a new large-scale dataset for VRSA including 360 videos with various frame rates, physiological signals, and subjective scores. On VRSA benchmark and our newly collected dataset, our approach shows a potential to not only achieve the highest correlation with subjective scores, but also to better understand which symptoms are the main causes of VR sickness. Hak Gu Kim, Sangmin Lee 0001, Seongyeop Kim, Heoun-taek Lim, Yong Man Ro |
AAAI | 2 |
| 2021 | Visual Comfort Aware-Reinforcement Learning for Depth Adjustment of Stereoscopic 3D ImagesabstractDepth adjustment aims to enhance the visual experience of stereoscopic 3D (S3D) images, which accompanied with improving visual comfort and depth perception. For a human expert, the depth adjustment procedure is a sequence of iterative decision making. The human expert iteratively adjusted the depth until he is satisfied with the both levels of visual comfort and the perceived depth. In this work, we present a novel deep reinforcement learning (DRL)-based approach for depth adjustment named VCA-RL (Visual Comfort Aware Reinforcement Learning) to explicitly model human sequential decision making in depth editing operations. We formulate the depth adjustment process as a Markov decision process where actions are defined as camera movement operations to control the distance between the left and right cameras. Our agent is trained based on the guidance of an objective visual comfort assessment metric to learn the optimal sequence of camera movement actions in terms of perceptual aspects in stereoscopic viewing. With extensive experiments and user studies, we show the effectiveness of our VCA-RL model on three different S3D databases. Hak Gu Kim, Minho Park 0002, Sangmin Lee 0001, Seongyeop Kim, Yong Man Ro |
AAAI | 3 |
| 2021 | Video Prediction Recalling Long-Term Motion Context via Memory Alignment LearningabstractOur work addresses long-term motion context issues for predicting future frames. To predict the future precisely, it is required to capture which long-term motion context (e.g., walking or running) the input motion (e.g., leg movement) belongs to. The bottlenecks arising when dealing with the long-term motion context are: (i) how to predict the long-term motion context naturally matching input sequences with limited dynamics, (ii) how to predict the long-term motion context with high-dimensionality (e.g., complex motion). To address the issues, we propose novel motion context-aware video prediction. To solve the bottle-neck (i), we introduce a long-term motion context memory (LMC-Memory) with memory alignment learning. The pro-posed memory alignment learning enables to store long-term motion contexts into the memory and to match them with sequences including limited dynamics. As a result, the long-term context can be recalled from the limited in-put sequence. In addition, to resolve the bottleneck (ii), we propose memory query decomposition to store local motion context (i.e., low-dimensional dynamics) and recall the suitable local context for each local part of the input individually. It enables to boost the alignment effects of the memory. Experimental results show that the proposed method outperforms other sophisticated RNN-based methods, especially in long-term condition. Further, we validate the effectiveness of the proposed network designs by conducting ablation studies and memory feature analysis. The source code of this work is available†. Sangmin Lee 0001, Hak Gu Kim, Dae Hwi Choi, Hyungil Kim, Yong Man Ro |
CVPR | 1 |
| 2021 | CUA Loss: Class Uncertainty-Aware Gradient Modulation for Robust Object DetectionabstractRecently, a wide range of research on object detection has shown breakthrough performance. However, in a challenging environment, such as occlusion and small object cases, object detectors still produce inaccurate or erroneous predictions. To effectively cope with such conditions, most of the existing methods have suggested loss functions to guide the object detectors by modulating the magnitude of their loss. However, when modulating the loss function, they are highly dependent on the classification score of the object detector. It is a known fact that deep neural networks tend to be overconfident in their predictions. In this article, to alleviate the problem of the object detectors which heavily rely on the prediction in the training phase, we devise a novel loss function called class uncertainty-aware (CUA) loss. CUA loss considers the predictive ambiguity as well as the predictions on classification score when modulating loss function. In addition to the classification score, CUA loss further modulates the loss gradient in an increasing way when the object detectors output an uncertain prediction. Therefore, object detectors with CUA loss effectively cope with challenging environments where prediction results are uncertain. With comprehensive experiments on three public datasets (i.e. PASCAL VOC, MS COCO, and Berkeley DeepDrive), we verified that our CUA loss enhanced the accuracy of the object detectors and outperformed previous state-of-the-art loss functions. Jung Uk Kim, Seong Tae Kim 0001, Hong Joo Lee 0001, Sangmin Lee 0001, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Robust Video Frame Interpolation With Exceptional Motion MapabstractVideo frame interpolation has increasingly attracted attention in computer vision and video processing fields. When motion patterns in a video are complex, large and non-linear (exceptional motion), the generated intermediate frame is blurred and likely to have large artifacts. In this paper, we propose a novel video frame interpolation considering the exceptional motion patterns. The proposed video frame interpolation takes into account an exceptional motion map that contains the location and intensity of the exceptional motion. The proposed method consists of three parts, which are optical flow based frame interpolation, exceptional motion detection, and frame refinement. The optical flow based frame interpolation predicts an optical flow which is used to synthesize the pre-generated intermediate frame. The exceptional motion detection detects the position and intensity of complex and large motion with the current frame and the previous frame sequence. The frame refinement focuses on the exceptional motion region of the pre-generated intermediate frame by using the exceptional motion map. The proposed video frame interpolation can be robust against the exceptional motion including complex and large motion. Experimental results showed that the proposed video frame interpolation achieved high performance on various public video datasets and especially on videos with exceptional motion patterns. Minho Park 0002, Hak Gu Kim, Sangmin Lee 0001, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Structure Boundary Preserving Segmentation for Medical Image With Ambiguous BoundaryabstractIn this paper, we propose a novel image segmentation method to tackle two critical problems of medical image, which are (i) ambiguity of structure boundary in the medical image domain and (ii) uncertainty of the segmented region without specialized domain knowledge. To solve those two problems in automatic medical segmentation, we propose a novel structure boundary preserving segmentation framework. To this end, the boundary key point selection algorithm is proposed. In the proposed algorithm, the key points on the structural boundary of the target object are estimated. Then, a boundary preserving block (BPB) with the boundary key point map is applied for predicting the structure boundary of the target object. Further, for embedding experts' knowledge in the fully automatic segmentation, we propose a novel shape boundary-aware evaluator (SBE) with the ground-truth structure information indicated by experts. The proposed SBE could give feedback to the segmentation network based on the structure boundary key point. The proposed method is general and flexible enough to be built on top of any deep learning-based segmentation network. We demonstrate that the proposed method could surpass the state-of-the-art segmentation network and improve the accuracy of three different segmentation network models on different types of medical image datasets. Hong Joo Lee 0001, Jung Uk Kim, Sangmin Lee 0001, Hak Gu Kim, Yong Man Ro |
CVPR | 3 |
| 2020 | SACA Net: Cybersickness Assessment of Individual Viewers for VR Content via Graph-Based Symptom Relation Embedding
Sangmin Lee 0001, Jung Uk Kim, Hak Gu Kim, Seongyeop Kim, Yong Man Ro |
ECCV (23) | 1 |
| 2020 | Video Frame Interpolation Via Exceptional Motion-Aware SynthesisabstractIn this paper, we propose a novel video frame interpolation method via exceptional motion-aware synthesis, in which accurate optical flow could be estimated even with exceptional motion patterns. Specifically, we devise two deep learning modules: exceptional motion detection and frame interpolation with refined flow. The motion detection module detects the position and intensity of exceptional motion patterns in current frame given the past frame sequence. The flow refinement module refines the pre-estimated optical flow for synthesizing the intermediate frame using the information of exceptional motion. The proposed modules improve the quality of the synthesized intermediate frame by making the optical flow robust against exceptional case of motion. Experimental results showed that the proposed method outperforms the state-of-the-art methods qualitatively and quantitatively. Minho Park 0002, Sangmin Lee 0001, Yong Man Ro |
ICASSP | 2 |
| 2020 | Fake Video Detection With Certainty-Based Attention NetworkabstractDeepFake synthesizes realistic fake videos that could be used maliciously such as manipulation and harassment. In order to prevent such malicious usages, detecting fake videos is immediately needed. In this paper, we propose a novel fake video detection method by adopting predictive uncertainty in detection. We devise the certainty-based attention network which guides to focus certainty-key frames in detecting fake videos. In addition, certainty-based attention is proposed for refining the features with consideration for frame-level certainty. Experiments are performed to validate the effectiveness of the proposed method by comparing the existing methods on Celeb-DF, the latest DeepFake dataset. Dae Hwi Choi, Hong Joo Lee 0001, Sangmin Lee 0001, Jung Uk Kim, Yong Man Ro |
ICIP | 3 |
| 2020 | Comprehensive Facial Expression Synthesis Using Human-Interpretable LanguageabstractRecent advances in facial expression synthesis have shown promising results using diverse expression representations including facial action units. Facial action units for an elaborate facial expression synthesis need to be intuitively represented for human comprehension, not a numeric categorization of facial action units. To address this issue, we utilize human-friendly approach: use of natural language where language helps human grasp conceptual contexts. In this paper, therefore, we propose a new facial expression synthesis model from language-based facial expression description. Our method can synthesize the facial image with detailed expressions. In addition, effectively embedding language features on facial features, our method can control individual word to handle each part of facial movement. Extensive qualitative and quantitative evaluations were conducted to verify the effectiveness of the natural language. Joanna Hong, Jung Uk Kim, Sangmin Lee 0001, Yong Man Ro |
ICIP | 3 |
| 2020 | Learning Style Correlation for Elaborate Few-Shot ClassificationabstractFew-shot classification is defined as a task where the network aims to classify unseen classes given only a few samples. Recent approaches, especially metric-based methods, have great progress in few-shot classification. However, the existing metric-based methods have a limitation in deploying discriminative features for elaborate comparison. They usually extract features from the embedding network without direct consideration of the relationship between support and query sets. To address the relationship, we propose a novel architecture, Style Correlated Module (SCM) to learn style correlation between support and query sets for few-shot classification. The proposed module leads support and query feature maps to focus on significant style correlated features and encourage the metric network to conduct an elaborate comparison. Furthermore, the proposed module can be generally applied to the existing metric-based approaches by adding the SCM behind the embedding network. We evaluate our proposed method with comprehensive experiments on two publicly available datasets and demonstrate its effectiveness with comparable results. Minsu Kim 0001, Jung Uk Kim, Hong Joo Lee 0001, Sangmin Lee 0001, Joanna Hong, Yong Man Ro |
ICIP | 5 |
| 2020 | Class Incremental Learning With Task-SelectionabstractDespite the success of the deep neural networks (DNNs), in case of incremental learning, DNNs are known to suffer from catastrophic forgetting problems which are the phenomenon of entirely forgetting previously learned task information upon learning current task information. To alleviate this problem, we propose a novel knowledge distillation-based class incremental learning method with a task-selective autoencoder (TsAE). By learning the TsAE to reconstruct the feature map of each task, the proposed method effectively memorizes not only the classes of the current task but also the classes of previously learned tasks. Since the proposed TsAE has a simple but powerful architecture, it can be easily generalized to other knowledge distillation-based class incremental learning methods. Our experimental results on various datasets, including iCIFAR-100 and iILSVRC-small, demonstrated that the proposed method achieves higher classification accuracy and less forgetting compared to the stateof-the-art methods. Eun Sung Kim, Jung Uk Kim, Sangmin Lee 0001, Sang-Keun Moon, Yong Man Ro |
ICIP | 3 |
| 2020 | Robust Video Facial Authentication With Unsupervised Mode DisentanglementabstractDeep learning-based video facial authentication has limitations when it comes to real-world applications, due to large mode variations such as illumination, pose, and eyeglasses variations in real-life situations. Many of existing mode-invariant facial authentication methods need labels of each mode. However, the label information could not be always available in practice. To alleviate this problem, we develop an unsupervised mode disentangling method for video facial authentication. By matching both disentangled identity features and dynamic features of two facial videos, our proposed method shows significant face verification and identification performances on three publicly available datasets, KAIST-MPMI, UVA-NEMO, and YTF. Minsu Kim 0001, Hong Joo Lee 0001, Sangmin Lee 0001, Yong Man Ro |
ICIP | 3 |
| 2020 | Estimating VR Sickness Caused By Camera Shake in VR VideographyabstractRecent development of Virtual Reality (VR) technology provides more realistic experience for viewers with a variety of contents. While the viewing safety of the viewers is one of the important issues in VR industry, the necessity of VR sickness estimation has been drawing attentions. Inspired by the observations that camera shake in VR videography is one of the major causes of VR sickness, we propose a novel deep network that predicts VR sickness level of individuals caused by camera shake. The proposed method is designed to comprehensively identify changes in direction and speed of the VR video scenes with camera shake. Sparse selection of optical flow maps with different intervals allows the proposed network to efficiently extract stimulus features with a variety of camera shake patterns. We built a new benchmark database for the evaluation of the proposed method that consists of 360-degree videos including various camera shake movements, physiological signals, and Simulation Sickness Questionnaires (SSQ) scores of the experimental participants. Experimental results of the sickness prediction show the effectiveness of the proposed method on the built benchmark database. Seongyeop Kim, Sangmin Lee 0001, Yong Man Ro |
ICIP | 2 |
| 2020 | BMAN: Bidirectional Multi-Scale Aggregation Networks for Abnormal Event DetectionabstractAbnormal event detection is an important task in video surveillance systems. In this paper, we propose a novel bidirectional multi-scale aggregation networks (BMAN) for abnormal event detection. The proposed BMAN learns spatiotemporal patterns of normal events to detect deviations from the learned normal patterns as abnormalities. The BMAN consists of two main parts: an inter-frame predictor and an appearancemotion joint detector. The inter-frame predictor is devised to encode normal patterns, which generates an inter-frame using bidirectional multi-scale aggregation based on attention. With the feature aggregation, robustness for object scale variations and complex motions is achieved in normal pattern encoding. Based on the encoded normal patterns, abnormal events are detected by the appearance-motion joint detector in which both appearance and motion characteristics of scenes are considered. Comprehensive experiments are performed, and the results show that the proposed method outperforms the existing state-of-the-art methods. The resulting abnormal event detection is interpretable on the visual basis of where the detected events occur. Further, we validate the effectiveness of the proposed network designs by conducting ablation study and feature visualization. Sangmin Lee 0001, Hak Gu Kim, Yong Man Ro |
IEEE Trans. Image Process. | 1 |
| 2019 | Deep Objective Assessment Model Based on Spatio-Temporal Perception of 360-Degree Video for VR Sickness PredictionabstractIn virtual reality (VR) environment, viewing safety is one of increasing concerns because of physical symptoms induced by VR sickness. Distortion of VR video is one of main causes. In this paper, we investigate the degradation of spatial resolution as distortion causing VR sickness. We propose a novel deep learning-based VR sickness assessment framework for predicting VR sickness caused by degradation of spatial resolution. The proposed method takes into account visual perception of 360-degree videos in spatio-temporal domain for assessing VR sickness. In cooperating visual quality and the temporal flickering with deep latent feature in training stage, the proposed network could effectively learn the spatio-temporal characteristics causing VR sickness. To evaluate the performance of the proposed method, we built a new dataset consisting of 360-degree videos and ground truths (physiological signals and SSQ scores). The dataset will be open publicly. Experimental results demonstrated that the proposed VR sickness assessment had a high correlation with human subjective scores. Ki Hyun Kim, Sangmin Lee 0001, Hak Gu Kim, Minho Park 0002, Yong Man Ro |
ICIP | 2 |
| 2019 | Physiological Fusion Net: Quantifying Individual VR Sickness with Content Stimulus and Physiological ResponseabstractQuantifying Virtual Reality (VR) sickness is demanded in industry to address viewing safety issue. In this paper, we develop a new method to quantify VR sickness. We propose a novel physiological fusion deep network which estimates individual VR sickness with content stimulus and physiological response. In the proposed framework, content stimulus guider and physiological response guider are devised to effectively represent feature related with VR sickness. Deep stimulus feature from the content stimulus guiders reflects the content sickness tendency while deep physiology feature from the physiological response guider reflects the individual sickness characteristics. By combining those features, VR sickness predictor quantifies individual Simulation Sickness Questionnaires (SSQ) scores. To evaluate the performance of the proposed method, we built a new dataset that consists of 360-degree videos with physiological signals and SSQ scores. Experimental results show that the proposed method achieved meaningful correlation with human subjective scores. Sangmin Lee 0001, Seongyeop Kim, Hak Gu Kim, Min Seob Kim, Seokho Yun, Bumseok Jeong, Yong Man Ro |
ICIP | 1 |
| 2019 | VRSA Net: VR Sickness Assessment Considering Exceptional Motion for 360° VR VideoabstractThe viewing safety is one of the main issues in viewing virtual reality (VR) content. In particular, VR sickness could occur when watching immersive VR content. To deal with the viewing safety for VR content, objective assessment of VR sickness is of great importance. In this paper, we propose a novel objective VR sickness assessment (VRSA) network based on deep generative model for automatically predicting the VR sickness score. The proposed method takes into account motion patterns of VR videos in which an exceptional motion is a critical factor inducing excessive VR sickness in human motion perception. The proposed VRSA network consists of two parts, which are VR video generator and VR sickness score predictor. By training the VR video generator with common videos with non-exceptional motion, the generator learns the tolerance of VR sickness in human motion perception. As a result, the difference between the original and the generated videos by the VR video generator could represent exceptional motion of VR video causing VR sickness. In the VR sickness score predictor, the VR sickness score is predicted by projecting the difference between the original and the generated videos onto the subjective score space. For the evaluation of VR sickness assessment, we built a new dataset which consists of 360° videos (stimuli), corresponding physiological signals, and subjective questionnaires from subjective assessment experiments. Experimental results demonstrated that the proposed VRSA network achieved a high correlation with human perceptual score for VR sickness. Hak Gu Kim, Heoun-taek Lim, Sangmin Lee 0001, Yong Man Ro |
IEEE Trans. Image Process. | 3 |
| 2018 | Stan: Spatio- Temporal Adversarial Networks for Abnormal Event DetectionabstractIn this paper, we propose a novel abnormal event detection method with spatio-temporal adversarial networks (STAN). We devise a spatio-temporal generator which synthesizes an inter- frame by considering spatio-temporal characteristics with bidirectional ConvLSTM. A proposed spatio-temporal discriminator determines whether an input sequence is real-normal or not with 3D convolutional layers. These two networks are trained in an adversarial way to effectively encode spatio-temporal features of normal patterns. After the learning, the generator and the discriminator can be independently used as detectors, and deviations from the learned normal patterns are detected as abnormalities. Experimental results show that the proposed method achieved competitive performance compared to the state-of-the-art methods. Further, for the interpretation, we visualize the location of abnormal events detected by the proposed networks using a generator loss and discriminator gradients. Sangmin Lee 0001, Hak Gu Kim, Yong Man Ro |
ICASSP | 1 |