EDBT 2026 Demo / reviewers in the wild / expert
Hak Gu Kim
dblp:143/2074
· DBLP profile ↗
39ranked-venue papers
11as first author
15since 2021 · last 2026
0000-0003-2137-934XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 11 first-author · 14 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unsharp-Inspired Adversarial Point Cloud Perturbation via Low-Rank ApproximationabstractDeep neural networks (DNNs) for 3D point cloud recognition have achieved remarkable performance, but remain highly vulnerable to adversarial perturbations. Existing adversarial methods often suffer from a trade-off between imperceptibility and attack strength, either producing noticeable outliers or requiring excessive point displacements. In this paper, we propose a novel unsharp-inspired adversarial perturbation that leverages low-rank approximation to balance global structure preservation and fine-detail manipulation. Specifically, the point cloud is decomposed via eigenvalue decomposition (EVD) into global and residual components, where the low-rank reconstruction captures the overall shape and the residual highlights salient details. Perturbations are then selectively applied to these fine-detail regions and smoothed according to local magnitude and orientation, ensuring geometric consistency. Extensive experiments on ModelNet40 and ShapeNet-Part using PointNet and DGCNN demonstrate that our method achieves 100% attack success rate while significantly reducing geometric distortion compared to state-of-the-art methods. These results highlight the effectiveness of spectral-domain low-rank modeling for generating adversarial point clouds that are both strong and imperceptible. Kyo Seok Lee, Han-nyoung Lee, Hak Gu Kim |
IEEE Signal Process. Lett. | 3 |
| 2026 | Deformable 3-D Point Cloud Perturbations Using Cage-Based Deformation for Semantic ConsistencyabstractDeep neural networks for 3D point cloud analysis are widely used in applications such as autonomous driving and robotics, yet they remain highly vulnerable to adversarial attacks. Existing methods typically minimize point-wise distances to preserve geometry, which constrains perturbations and leads to a trade-off between imperceptibility and attack strength. To address this limitation, we propose a cage-based adversarial deformation framework that generates semantically consistent perturbations aligned with natural intra-class variations. Our method refines a source cage, predicts adversarial cage displacements by fusing source–target features, and computes smooth point-wise offsets using solid-angle– and distance-aware weights. This enables globally coherent deformations that appear natural to humans while effectively misleading classifiers. Extensive experiments on ModelNet40, ShapeNet-Part, and ScanObjectNN show that our approach achieves consistently high attack success rates while simultaneously improving point uniformity and reducing local geometric distortions. Furthermore, the perturbations remain effective against various defense methods. Kyo Seok Lee, Hak Gu Kim |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
Hyung Kyu Kim, Hak Gu Kim |
ICCV | 3 |
| 2025 | Enhancing 3D Scene Representation with Structural Dissimilarity-Aware LearningabstractNovel view synthesis aims to generate high-quality unseen views from images at different viewpoints. However, existing methods often struggle to preserve fine details, leading to structural distortions in complex regions. In this paper, we introduce a simple yet effective structure-aware objective function designed to enhance structural information in novel view synthesis. By leveraging the Structural Similarity Index (SSIM), our method attends to regions exhibiting significant structural distortions. We incorporate structural dissimilarity-based attention to highlight discrepancies in challenging regions between predicted and ground-truth images. It enables recent 3D scene representation models to achieve improved structural preservation, leading to more coherent representations. Experiments on synthetic and real-world datasets demonstrate that our method enhances structural consistency, particularly in challenging regions. Ho Jun Kim, Hak Gu Kim |
ICIP | 3 |
| 2025 | Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial Animation
Hyung Kyu Kim, Hak Gu Kim |
INTERSPEECH | 2 |
| 2025 | MSCoTDet: Language-Driven Multi-Modal Fusion for Improved Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection is attractive for around-the-clock applications due to the complementary information between RGB and thermal modalities. However, current models often fail to detect pedestrians in certain cases (e.g., thermal-obscured pedestrians), particularly due to the modality bias learned from statistically biased datasets. In this paper, we investigate how to mitigate modality bias in multispectral pedestrian detection using a Large Language Model (LLM). Accordingly, we design a Multispectral Chain-of-Thought (MSCoT) prompting strategy, which prompts the LLM to perform multispectral pedestrian detection. Moreover, we propose a novel Multispectral Chain-of-Thought Detection (MSCoTDet) framework that integrates MSCoT prompting into multispectral pedestrian detection. To this end, we design a Language-driven Multi-modal Fusion (LMF) strategy that enables fusing the outputs of MSCoT prompting with the detection results of vision-based multispectral pedestrian detection models. Extensive experiments validate that MSCoTDet effectively mitigates modality biases and improves multispectral pedestrian detection. Taeheon Kim, Sangyun Chung, Damin Yeom, Youngjoon Yu, Hak Gu Kim, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Causal Mode Multiplexer: A Novel Framework for Unbiased Multispectral Pedestrian DetectionabstractRGBT multispectral pedestrian detection has emerged as a promising solution for safety-critical applications that require day/night operations. However, the modality bias problem remains unsolved as multispectral pedestrian detectors learn the statistical bias in datasets. Specifically, datasets in multispectral pedestrian detection mainly distribute between ROTO11R⋆T⋆ refers to the visibility (O/X) in each modality. Generally, ROTO refers to daytime images, and RXTO refers to nighttime images. ROTX refers to daytime images in obscured situations. (day) and RXTO (night) data; the majority of the pedestrian labels statistically co-occur with their thermal features. As a result, multispectral pedestrian detectors show poor generalization ability on examples beyond this statistical correlation, such as ROTX data. To address this problem, we propose a novel Causal Mode Multiplexer (CMM) framework that effectively learns the causalities between multispectral inputs and predictions. Moreover, we construct a new dataset (ROTX-MP) to evaluate modality bias in multispectral pedestrian detection. ROTX-MP mainly includes ROTX examples not presented in previous datasets. Extensive experiments demonstrate that our proposed CMM framework generalizes well on existing datasets (KAIST, CVC-14, FLIR) and the new ROTX-MP. Our code and dataset are available at: https://github.com/ssbin0914/Causal-Mode-Multiplexer.git. Taeheon Kim, Sebin Shin, Youngjoon Yu, Hak Gu Kim, Yong Man Ro |
CVPR | 4 |
| 2024 | Analyzing Visible Articulatory Movements in Speech Production For Speech-Driven 3D Facial AnimationabstractSpeech-driven 3D facial animation aims to generate realistic facial meshes based on input speech signals. However, due to a lack of understanding of visible articulatory movements, current state-of-the-art methods result in inaccurate lip and jaw movements. Traditional evaluation metrics, such as lip vertex error (LVE), often fail to represent the quality of visual results. Based on our observation, we reveal the problems with existing evaluation metrics and raise the necessity for separate evaluation approaches for 3D axes. Comprehensive analysis shows that most recent methods struggle to precisely predict lip and jaw movements in 3D space. Hyung Kyu Kim, Sangmin Lee 0001, Hak Gu Kim |
ICIP | 3 |
| 2024 | Super-Resolution Neural Radiance Field via Learning High Frequency Details for High-Fidelity Novel View SynthesisabstractWhile neural rendering approaches facilitate photorealistic rendering in novel view synthesis tasks, the challenge of high-resolution rendering persists due to the substantial costs associated with acquiring and training data. Recently, several studies have been proposed that render high-resolution scenes by either super-sampling points or using reference images, aiming to restore details in low-resolution (LR) images. However, supersampling is computationally expensive, and methods with reference images require high-resolution (HR) images for inference. In this paper, we propose a novel super-resolution (SR) neural radiance field (NeRF) framework for high-fidelity novel view synthesis. To address the representation of high-fidelity HR images from the captured LR images, we learn a mapping function that maps LR rendering images to the Fourier space to restore insufficient high frequency details and render HR images at higher resolution. Experiments demonstrate that our results are quantitatively and qualitatively better than those of the existing SR methods in novel view synthesis. By visualizing the estimated dominant frequency components, we provide visual interpretations of the performance improvement. Han-nyoung Lee, Hak Gu Kim |
IEEE Signal Process. Lett. | 2 |
| 2022 | Natural-Looking Adversarial Examples from Freehand SketchesabstractDeep neural networks (DNNs) have achieved great success in image classification and recognition compared to previous methods. However, recent works have reported that DNNs are very vulnerable to adversarial examples that are intentionally generated to mislead the predictions of the DNNs. Here, we present a novel freehand sketch-based natural-looking adversarial example generator that we call SketchAdv. To generate a natural-looking adversarial example from a sketch, we force the encoded edge information (i.e., the visual attributes) to be close to the latent random vector fed to the edge generator and adversarial example generator. This preserves the spatial consistency of the adversarial example generated from the random vector with the edge information. In addition, by employing a sketch-edge encoder with a novel sketch-edge matching loss, we reduce the gap between edges and sketches. We evaluate the proposed method on several dominant classes of SketchyCOCO, the benchmark dataset for sketch to image translation. Our experiments show that our SketchAdv produces visually plausible adversarial examples while remaining competitive with other adversarial attack methods. Hak Gu Kim, Davide Nanni, Sabine Süsstrunk |
ICASSP | 1 |
| 2022 | Assessing Individual VR Sickness Through Deep Feature Fusion of VR Video and Physiological ResponseabstractRecently, VR sickness assessment for VR videos is highly demanded in industry and research fields to address VR viewing safety issues. Especially, it is difficult to evaluate VR sickness of individuals due to individual differences. To achieve the challenging goal, we focus on deep feature fusion of sickness-related information. In this paper, we propose a novel deep learning-based assessment framework which estimates VR sickness of individual viewers with VR videos and corresponding physiological responses. We design the content stimulus guider imitating the phenomenon that humans feel VR sickness. The content stimulus guider extracts a deep stimulus feature from a VR video to reflect VR sickness caused by VR videos. In addition, we devise the physiological response guider to encode physiological responses that are acquired while humans experience VR videos. Each physiology sickness feature extractor (EEG, ECG, and GSR) in the physiological response guider is designed to suit their physiological characteristics. Extracted physiology sickness features are then fused into a deep physiology feature that comprehensively reflects individual deviations of VR sickness. Finally, the VR sickness predictor assesses individual VR sickness effectively with the fusion of the deep stimulus feature and the deep physiology feature. To validate the proposed method extensively, we built two benchmark datasets which contain 360-degree VR videos with physiological responses (EEG, ECG, and GSR) and SSQ scores. Experimental results show that the proposed method achieves meaningful correlations with human SSQ scores. Further, we validate the effectiveness of the proposed network designs by conducting analysis on feature fusion and visualization. Sangmin Lee 0001, Seongyeop Kim, Hak Gu Kim, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Towards a Better Understanding of VR Sickness: Physical Symptom Prediction for VR ContentsabstractWe address the black-box issue of VR sickness assessment (VRSA) by evaluating the level of physical symptoms of VR sickness. For the VR contents inducing the similar VR sickness level, the physical symptoms can vary depending on the characteristics of the contents. Most of existing VRSA methods focused on assessing the overall VR sickness score. To make better understanding of VR sickness, it is required to predict and provide the level of major symptoms of VR sickness rather than overall degree of VR sickness. In this paper, we predict the degrees of main physical symptoms affecting the overall degree of VR sickness, which are disorientation, nausea, and oculomotor. In addition, we introduce a new large-scale dataset for VRSA including 360 videos with various frame rates, physiological signals, and subjective scores. On VRSA benchmark and our newly collected dataset, our approach shows a potential to not only achieve the highest correlation with subjective scores, but also to better understand which symptoms are the main causes of VR sickness. Hak Gu Kim, Sangmin Lee 0001, Seongyeop Kim, Heoun-taek Lim, Yong Man Ro |
AAAI | 1 |
| 2021 | Visual Comfort Aware-Reinforcement Learning for Depth Adjustment of Stereoscopic 3D ImagesabstractDepth adjustment aims to enhance the visual experience of stereoscopic 3D (S3D) images, which accompanied with improving visual comfort and depth perception. For a human expert, the depth adjustment procedure is a sequence of iterative decision making. The human expert iteratively adjusted the depth until he is satisfied with the both levels of visual comfort and the perceived depth. In this work, we present a novel deep reinforcement learning (DRL)-based approach for depth adjustment named VCA-RL (Visual Comfort Aware Reinforcement Learning) to explicitly model human sequential decision making in depth editing operations. We formulate the depth adjustment process as a Markov decision process where actions are defined as camera movement operations to control the distance between the left and right cameras. Our agent is trained based on the guidance of an objective visual comfort assessment metric to learn the optimal sequence of camera movement actions in terms of perceptual aspects in stereoscopic viewing. With extensive experiments and user studies, we show the effectiveness of our VCA-RL model on three different S3D databases. Hak Gu Kim, Minho Park 0002, Sangmin Lee 0001, Seongyeop Kim, Yong Man Ro |
AAAI | 1 |
| 2021 | Video Prediction Recalling Long-Term Motion Context via Memory Alignment LearningabstractOur work addresses long-term motion context issues for predicting future frames. To predict the future precisely, it is required to capture which long-term motion context (e.g., walking or running) the input motion (e.g., leg movement) belongs to. The bottlenecks arising when dealing with the long-term motion context are: (i) how to predict the long-term motion context naturally matching input sequences with limited dynamics, (ii) how to predict the long-term motion context with high-dimensionality (e.g., complex motion). To address the issues, we propose novel motion context-aware video prediction. To solve the bottle-neck (i), we introduce a long-term motion context memory (LMC-Memory) with memory alignment learning. The pro-posed memory alignment learning enables to store long-term motion contexts into the memory and to match them with sequences including limited dynamics. As a result, the long-term context can be recalled from the limited in-put sequence. In addition, to resolve the bottleneck (ii), we propose memory query decomposition to store local motion context (i.e., low-dimensional dynamics) and recall the suitable local context for each local part of the input individually. It enables to boost the alignment effects of the memory. Experimental results show that the proposed method outperforms other sophisticated RNN-based methods, especially in long-term condition. Further, we validate the effectiveness of the proposed network designs by conducting ablation studies and memory feature analysis. The source code of this work is available†. Sangmin Lee 0001, Hak Gu Kim, Dae Hwi Choi, Hyungil Kim, Yong Man Ro |
CVPR | 2 |
| 2021 | Robust Video Frame Interpolation With Exceptional Motion MapabstractVideo frame interpolation has increasingly attracted attention in computer vision and video processing fields. When motion patterns in a video are complex, large and non-linear (exceptional motion), the generated intermediate frame is blurred and likely to have large artifacts. In this paper, we propose a novel video frame interpolation considering the exceptional motion patterns. The proposed video frame interpolation takes into account an exceptional motion map that contains the location and intensity of the exceptional motion. The proposed method consists of three parts, which are optical flow based frame interpolation, exceptional motion detection, and frame refinement. The optical flow based frame interpolation predicts an optical flow which is used to synthesize the pre-generated intermediate frame. The exceptional motion detection detects the position and intensity of complex and large motion with the current frame and the previous frame sequence. The frame refinement focuses on the exceptional motion region of the pre-generated intermediate frame by using the exceptional motion map. The proposed video frame interpolation can be robust against the exceptional motion including complex and large motion. Experimental results showed that the proposed video frame interpolation achieved high performance on various public video datasets and especially on videos with exceptional motion patterns. Minho Park 0002, Hak Gu Kim, Sangmin Lee 0001, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Structure Boundary Preserving Segmentation for Medical Image With Ambiguous BoundaryabstractIn this paper, we propose a novel image segmentation method to tackle two critical problems of medical image, which are (i) ambiguity of structure boundary in the medical image domain and (ii) uncertainty of the segmented region without specialized domain knowledge. To solve those two problems in automatic medical segmentation, we propose a novel structure boundary preserving segmentation framework. To this end, the boundary key point selection algorithm is proposed. In the proposed algorithm, the key points on the structural boundary of the target object are estimated. Then, a boundary preserving block (BPB) with the boundary key point map is applied for predicting the structure boundary of the target object. Further, for embedding experts' knowledge in the fully automatic segmentation, we propose a novel shape boundary-aware evaluator (SBE) with the ground-truth structure information indicated by experts. The proposed SBE could give feedback to the segmentation network based on the structure boundary key point. The proposed method is general and flexible enough to be built on top of any deep learning-based segmentation network. We demonstrate that the proposed method could surpass the state-of-the-art segmentation network and improve the accuracy of three different segmentation network models on different types of medical image datasets. Hong Joo Lee 0001, Jung Uk Kim, Sangmin Lee 0001, Hak Gu Kim, Yong Man Ro |
CVPR | 4 |
| 2020 | SACA Net: Cybersickness Assessment of Individual Viewers for VR Content via Graph-Based Symptom Relation Embedding
Sangmin Lee 0001, Jung Uk Kim, Hak Gu Kim, Seongyeop Kim, Yong Man Ro |
ECCV (23) | 3 |
| 2020 | BBC Net: Bounding-Box Critic Network for Occlusion-Robust Object DetectionabstractObject detection has received significant interest in the research field of computer vision and is widely used in human-centric applications. The occlusion problem is a frequent obstacle that degrades detection quality. In this paper, we propose a novel object detection framework targeting robust object detection in occlusion. The proposed deep learning-based network consists mainly of two parts: 1) object detection framework, which classifies the object categories and localizes the object location and 2) plug-in bounding-box (BB) estimator, which estimates the object and occlusion region from the feature map of the backbone network and the corresponding critic network for evaluating the predicted BB map. The BB estimator and the critic network are the plug-in modules added to the object detection framework and learned competitively with adversarial manner. As the plug-in BB estimator is learned to estimate the BB map containing the object and occlusion pattern information, the backbone network can embed this information to enable robust detection under occlusion in the test phase. The comprehensive experimental results on the PASCAL VOC, MS COCO, and KITTI dataset showed that the performance is improved with the plug-in BB-Critic network by predicting and criticizing object and occlusion in general generic object detection framework. Jung Uk Kim, Jungsu Kwon, Hak Gu Kim, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Deep Virtual Reality Image Quality Assessment With Human Perception Guider for Omnidirectional ImageabstractIn this paper, we propose a novel deep learning-based virtual reality image quality assessment method that automatically predicts the visual quality of an omnidirectional image. In order to assess the visual quality in viewing the omnidirectional image, we propose deep networks consisting of virtual reality (VR) quality score predictor and human perception guider. The proposed VR quality score predictor learns the positional and visual characteristics of the omnidirectional image by encoding the positional feature and visual feature of a patch on the omnidirectional image. With the encoded positional feature and visual feature, patch weight and patch quality score are estimated. Then, by aggregating all weights and scores of the patches, the image quality score is predicted. The proposed human perception guider evaluates the predicted quality score by referring to the human subjective score (i.e., ground-truth obtained by subjects) using an adversarial learning. With adversarial learning, the VR quality score predictor is trained to accurately predict the quality score in order to deceive the guider, while the proposed human perception guider is trained to precisely distinguish between the predictor score and the ground-truth subjective score. To verify the performance of the proposed method, we conducted comprehensive subjective experiments and evaluated the performance of the proposed method. The experimental results show that the proposed method outperforms the existing two-dimentional image quality models and the state-of-the-art image quality models for omnidirectional images. Hak Gu Kim, Heoun-taek Lim, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | MCSIP Net: Multichannel Satellite Image Prediction via Deep Neural NetworkabstractSatellite image prediction is important in weather nowcasting. In this article, we propose a novel multichannel satellite image prediction network (MCSIP Net) for predicting satellite images. The proposed MCSIP Net consists of three parts such as the satellite image predictor, the spatio-temporal 3-D discriminators, and the domain knowledge critic networks. The satellite image predictor takes a multichannel satellite image as an input and predicts a multichannel satellite image by learning spatio-temporal characteristics of each input channel. The spatio-temporal 3-D discriminators are trained to distinguish whether the input satellite image consists of a real satellite image or predicted image. By learning the spatio-temporal 3-D discriminator to distinguish and the satellite image predictor to deceive, the satellite image predictor can generate satellite image more similar to real satellite image distribution. The domain knowledge critic networks take the satellite image and the corresponding analysis data (which is obtained from a meteorological model) as an input and learn to distinguish whether the input satellite image is real or predicted on the basis of the analysis data. By utilizing the analysis data, the proposed MCSIP Net could take the meteorological knowledge into account efficiently. For the purpose of verification of the proposed method, ablation study and qualitative evaluation were conducted. Experimental results demonstrated that the proposed MCSIP Net could be learned efficiently and predict a multichannel satellite image with remarkable quality. Jae-Hyeok Lee 0001, Sangmin S. Lee, Hak Gu Kim, Sa-Kwang Song, Seongchan Kim, Yong Man Ro |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | BMAN: Bidirectional Multi-Scale Aggregation Networks for Abnormal Event DetectionabstractAbnormal event detection is an important task in video surveillance systems. In this paper, we propose a novel bidirectional multi-scale aggregation networks (BMAN) for abnormal event detection. The proposed BMAN learns spatiotemporal patterns of normal events to detect deviations from the learned normal patterns as abnormalities. The BMAN consists of two main parts: an inter-frame predictor and an appearancemotion joint detector. The inter-frame predictor is devised to encode normal patterns, which generates an inter-frame using bidirectional multi-scale aggregation based on attention. With the feature aggregation, robustness for object scale variations and complex motions is achieved in normal pattern encoding. Based on the encoded normal patterns, abnormal events are detected by the appearance-motion joint detector in which both appearance and motion characteristics of scenes are considered. Comprehensive experiments are performed, and the results show that the proposed method outperforms the existing state-of-the-art methods. The resulting abnormal event detection is interpretable on the visual basis of where the detected events occur. Further, we validate the effectiveness of the proposed network designs by conducting ablation study and feature visualization. Sangmin Lee 0001, Hak Gu Kim, Yong Man Ro |
IEEE Trans. Image Process. | 2 |
| 2019 | Deep Objective Assessment Model Based on Spatio-Temporal Perception of 360-Degree Video for VR Sickness PredictionabstractIn virtual reality (VR) environment, viewing safety is one of increasing concerns because of physical symptoms induced by VR sickness. Distortion of VR video is one of main causes. In this paper, we investigate the degradation of spatial resolution as distortion causing VR sickness. We propose a novel deep learning-based VR sickness assessment framework for predicting VR sickness caused by degradation of spatial resolution. The proposed method takes into account visual perception of 360-degree videos in spatio-temporal domain for assessing VR sickness. In cooperating visual quality and the temporal flickering with deep latent feature in training stage, the proposed network could effectively learn the spatio-temporal characteristics causing VR sickness. To evaluate the performance of the proposed method, we built a new dataset consisting of 360-degree videos and ground truths (physiological signals and SSQ scores). The dataset will be open publicly. Experimental results demonstrated that the proposed VR sickness assessment had a high correlation with human subjective scores. Ki Hyun Kim, Sangmin Lee 0001, Hak Gu Kim, Minho Park 0002, Yong Man Ro |
ICIP | 3 |
| 2019 | Physiological Fusion Net: Quantifying Individual VR Sickness with Content Stimulus and Physiological ResponseabstractQuantifying Virtual Reality (VR) sickness is demanded in industry to address viewing safety issue. In this paper, we develop a new method to quantify VR sickness. We propose a novel physiological fusion deep network which estimates individual VR sickness with content stimulus and physiological response. In the proposed framework, content stimulus guider and physiological response guider are devised to effectively represent feature related with VR sickness. Deep stimulus feature from the content stimulus guiders reflects the content sickness tendency while deep physiology feature from the physiological response guider reflects the individual sickness characteristics. By combining those features, VR sickness predictor quantifies individual Simulation Sickness Questionnaires (SSQ) scores. To evaluate the performance of the proposed method, we built a new dataset that consists of 360-degree videos with physiological signals and SSQ scores. Experimental results show that the proposed method achieved meaningful correlation with human subjective scores. Sangmin Lee 0001, Seongyeop Kim, Hak Gu Kim, Min Seob Kim, Seokho Yun, Bumseok Jeong, Yong Man Ro |
ICIP | 3 |
| 2019 | Generative Guiding Block: Synthesizing Realistic Looking Variants Capable of Even Large Change DemandsabstractRealistic image synthesis is to generate an image that is perceptually indistinguishable from an actual image. Generating realistic looking images with large variations (e.g., large spatial deformations and large pose change), however, is very challenging. Handing large variations as well as preserving appearance needs to be taken into account in the realistic looking image generation. In this paper, we propose a novel realistic looking image synthesis method, especially in large change demands. To do that, we devise generative guiding blocks. The proposed generative guiding block includes realistic appearance preserving discriminator and naturalistic variation transforming discriminator. By taking the proposed generative guiding blocks into generative model, the latent features at the layer of generative model are enhanced to synthesize both realistic looking- and target variation- image. With qualitative and quantitative evaluation in experiments, we demonstrated the effectiveness of the proposed generative guiding blocks, compared to the state-of-the-arts. Minho Park 0002, Hak Gu Kim, Yong Man Ro |
ICIP | 2 |
| 2019 | Photo-Realistic Facial Emotion Synthesis Using Multi-level Critic Networks with Multi-level Generative Model
Minho Park 0002, Hak Gu Kim, Yong Man Ro |
MMM (2) | 2 |
| 2019 | Binocular Fusion Net: Deep Learning Visual Comfort Assessment for Stereoscopic 3DabstractIn this paper, we propose a novel deep learning-based visual comfort assessment (VCA) for stereoscopic images. To assess the overall degree of visual discomfort in stereoscopic viewing, we devise a binocular fusion deep network (BFN) learning binocular characteristics between stereoscopic images. The proposed BFN learns the latent binocular feature representations for the visual comfort score prediction. In the BFN, the binocular feature is encoded by fusing the spatial features extracted from left and right views. Finally, the visual comfort score is predicted by projecting the binocular feature onto the subjective score space. In addition, we devise a disparity regularization network (DRN) for improving the prediction results. The proposed DRN takes the binocular feature from the BFN and estimates disparity maps from the feature in order to embed disparity relations between left and right views into the deep network. The proposed deep network with BFN and DRN is end-to-end trained in a unified framework in which the DRN acts as disparity regularization. We evaluated the prediction performance of the proposed deep network for VCA by the comparison of existing objective VCA metrics. Further, we demonstrated that the proposed BFN showed various factors causing visual discomfort by using network visualization. Hak Gu Kim, Hyunwook Jeong, Heoun-taek Lim, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | VRSA Net: VR Sickness Assessment Considering Exceptional Motion for 360° VR VideoabstractThe viewing safety is one of the main issues in viewing virtual reality (VR) content. In particular, VR sickness could occur when watching immersive VR content. To deal with the viewing safety for VR content, objective assessment of VR sickness is of great importance. In this paper, we propose a novel objective VR sickness assessment (VRSA) network based on deep generative model for automatically predicting the VR sickness score. The proposed method takes into account motion patterns of VR videos in which an exceptional motion is a critical factor inducing excessive VR sickness in human motion perception. The proposed VRSA network consists of two parts, which are VR video generator and VR sickness score predictor. By training the VR video generator with common videos with non-exceptional motion, the generator learns the tolerance of VR sickness in human motion perception. As a result, the difference between the original and the generated videos by the VR video generator could represent exceptional motion of VR video causing VR sickness. In the VR sickness score predictor, the VR sickness score is predicted by projecting the difference between the original and the generated videos onto the subjective score space. For the evaluation of VR sickness assessment, we built a new dataset which consists of 360° videos (stimuli), corresponding physiological signals, and subjective questionnaires from subjective assessment experiments. Experimental results demonstrated that the proposed VRSA network achieved a high correlation with human perceptual score for VR sickness. Hak Gu Kim, Heoun-taek Lim, Sangmin Lee 0001, Yong Man Ro |
IEEE Trans. Image Process. | 1 |
| 2018 | Stan: Spatio- Temporal Adversarial Networks for Abnormal Event DetectionabstractIn this paper, we propose a novel abnormal event detection method with spatio-temporal adversarial networks (STAN). We devise a spatio-temporal generator which synthesizes an inter- frame by considering spatio-temporal characteristics with bidirectional ConvLSTM. A proposed spatio-temporal discriminator determines whether an input sequence is real-normal or not with 3D convolutional layers. These two networks are trained in an adversarial way to effectively encode spatio-temporal features of normal patterns. After the learning, the generator and the discriminator can be independently used as detectors, and deviations from the learned normal patterns are detected as abnormalities. Experimental results show that the proposed method achieved competitive performance compared to the state-of-the-art methods. Further, for the interpretation, we visualize the location of abnormal events detected by the proposed networks using a generator loss and discriminator gradients. Sangmin Lee 0001, Hak Gu Kim, Yong Man Ro |
ICASSP | 2 |
| 2018 | VR IQA NET: Deep Virtual Reality Image Quality Assessment Using Adversarial LearningabstractIn this paper, we propose a novel virtual reality image quality assessment (VR IQA) with adversarial learning for omnidirectional images. To take into account the characteristics of the omnidirectional image, we devise deep networks including novel quality score predictor and human perception guider. The proposed quality score predictor automatically predicts the quality score of distorted image using the latent spatial and position feature. The proposed human perception guider criticizes the predicted quality score of the predictor with the human perceptual score using adversarial learning. For evaluation, we conducted extensive subjective experiments with omnidirectional image dataset. Experimental results show that the proposed VR IQA metric outperforms the 2-D IQA and the state-of-the-arts VR IQA. Heoun-taek Lim, Hak Gu Kim, Yang Man Ra |
ICASSP | 2 |
| 2018 | Object Bounding Box-Critic Networks for Occlusion-Robust Object Detection in Road SceneabstractObject detection in a road scene has received a significant attention from research fields of developing autonomous vehicle and automatic road monitoring systems. However, object occlusion problems frequently occur in generic road scenes. Due to such occlusion problems, previous object detection methods have limitations of not being able to detect objects accurately. In this paper, we propose a novel object detection network which is robust in occlusions. For effective object detection even with occlusion, the proposed network mainly consists of two parts; 1) Object detection framework, 2) Multiple object bounding box (OBB)-Critic network for predicting a BB map which estimates both object region and occlusion region. Comprehensive experimental results on a KITTI Vision Benchmark Suite dataset showed that the proposed object detection network outperformed the state-of-the-art methods. Jung Uk Kim, Jungsu Kwon, Hak Gu Kim, Haesung Lee, Yong Man Ro |
ICIP | 3 |
| 2018 | Adversarial Spatial Frequency Domain Critic Learning for Age and Gender ClassificationabstractThis paper proposes a novel deep learning framework for age and gender classification with the adversarial spatial frequency domain critic. In the proposed framework, the encoder-generator synthesizes realistic facial images with real images and corresponding age and gender label. An adversarial critic is devised to make generated images more proper for age and gender classification. In particular, we analyze the characteristic of age and gender attributes in the spatial frequency domain. Based on our investigation, we devise the spatial frequency domain critic network for considering the specific frequency bands which are dominant on age and gender attributes. Our discriminator is designed to simultaneously perform age and gender classification tasks. For this purpose, alternating learning is performed for multi-task classification. Experimental results showed that the proposed method outperformed other state-of-the-art methods in age and gender classifications. Sangmin S. Lee, Hak Gu Kim, Ki Hyun Kim, Yong Man Ro |
ICIP | 2 |
| 2018 | Teacher and Student Joint Learning for Compact Facial Landmark Detection Network
Hong Joo Lee 0001, Wissam J. Baddar, Hak Gu Kim, Seong Tae Kim 0001, Yong Man Ro |
MMM (1) | 3 |
| 2017 | Visual comfort assessment of stereoscopic images using deep visual and disparity features based on human attentionabstractThis paper proposes a novel visual comfort assessment (VCA) for stereoscopic images using deep learning. To predict visual discomfort of human visual system in stereoscopic viewing, we devise VCA deep networks to latently encode perceptual cues, which are visual differences between stereoscopic images and human attention-based disparity magnitude and gradient information. To extract the visual difference features from left and right views, a Siamese network is employed. In addition, human attention region-based disparity magnitude and gradient maps are fed to two individual deep convolutional neural networks (DCNNs) for disparity-related features based on human visual system (HVS). Finally, by aggregating these perceptual features, the proposed method directly predicts the final visual comfort score. Extensive and comparative experiments have been conducted on IEEE-SA dataset. Experimental results show that the proposed method can yield excellent correlation performance compared to existing methods. Hyunwook Jeong, Hak Gu Kim, Yong Man Ro |
ICIP | 2 |
| 2017 | Measurement of exceptional motion in VR video contents for VR sickness assessment using deep convolutional autoencoderabstractThis paper proposes a new objective metric of exceptional motion in VR video contents for VR sickness assessment. In VR environment, VR sickness can be caused by several factors which are mismatched motion, field of view, motion parallax, viewing angle, etc. Similar to motion sickness, VR sickness can induce a lot of physical symptoms such as general discomfort, headache, stomach awareness, nausea, vomiting, fatigue, and disorientation. To address the viewing safety issues in virtual environment, it is of great importance to develop an objective VR sickness assessment method that predicts and analyses the degree of VR sickness induced by the VR content. The proposed method takes into account motion information that is one of the most important factors in determining the overall degree of VR sickness. In this paper, we detect the exceptional motion that is likely to induce VR sickness. Spatio-temporal features of the exceptional motion in the VR video content are encoded using a convolutional autoencoder. For objectively assessing the VR sickness, the level of exceptional motion in VR video content is measured by using the convolutional autoencoder as well. The effectiveness of the proposed method has been successfully evaluated by subjective assessment experiment using simulator sickness questionnaires (SSQ) in VR environment. Hak Gu Kim, Wissam J. Baddar, Heoun-taek Lim, Hyunwook Jeong, Yong Man Ro |
VRST | 1 |
| 2017 | Multiview Stereoscopic Video Hole Filling Considering Spatiotemporal Consistency and Binocular Symmetry for Synthesized 3D VideoabstractThis paper proposes a new hole-filling method with spatiotemporal consistency and binocular symmetry for synthesized 3D videos in view extrapolation. Disocclusion regions in the synthesized views at virtual viewpoints result in regions with missing content. These regions will be referred to as hole regions. To provide the high-quality synthesized 3D videos via 3D display, the hole regions need to be filled considering the characteristics of human visual perception. From the perceptual point of view, binocular asymmetry between synthesized left- and right-eye videos (i.e., stereo pair video) is one of the most important factors that induce visual discomfort in stereoscopic viewing. In addition, binocular symmetry without temporal consistency between temporally neighboring frames could cause visual discomfort by annoying flickering artifacts. In this paper, to maintain the spatiotemporal consistency and binocular symmetry in synthesized 3D videos at multiple virtual viewpoints, we propose a global optimization-based hole-filling method using the information from the already filled adjacent view and previous frame. Furthermore, to reduce the computational cost of the global optimization, we propose a label propagation method, which propagates reliable labels used in the adjacent view and previous frame to the target image to be filled. The performance of the proposed method has been evaluated by objective assessments of 3D image quality, temporal consistency, and computational efficiency. In addition, subjective assessment is conducted for measuring the visual comfort and overall quality. The experimental results proved that the proposed method provides hole filling results with spatiotemporal consistency and binocular symmetry. Hak Gu Kim, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Measurement of critical temporal inconsistency for quality assessment of synthesized videoabstractThis paper proposes a new temporal consistency measure for quality assessment of synthesized video. Disocclusion regions appear hole regions of the synthesized video at virtual viewpoints. Filling hole regions could be problematic when the synthesized video is perceived through multi-view displays. In particular, the temporal inconsistency caused by hole filling process in view synthesis could affect the perceptual quality of the synthesized video. In the proposed method, we extract excessive flicker regions between consecutive frames and quantify the perceptual effects of the temporal inconsistency on them by measuring the structural similarity. We have demonstrated the validity of the proposed quality measure by comparisons of subjective ratings and existing objective metrics. Experimental results have shown that the proposed temporal inconsistency measure is highly correlated with the overall quality of the synthesized video. Hak Gu Kim, Yong Man Ro |
ICIP | 1 |
| 2016 | Critical Binocular Asymmetry Measure for the Perceptual Quality Assessment of Synthesized Stereo 3D Images in View SynthesisabstractIn human vision, excessive binocular asymmetry between the left- and right-eye images can be very problematic to perceive single binocular vision (causing visual discomfort) in the viewing of stereoscopic images. In this paper, we propose a critical binocular asymmetry (CBA) measure for objectively assessing the perceptual quality of the synthesized stereo 3D images generated through a depth-image-based rendering (DIBR) process. The proposed method detects critical regions that are likely to induce excessive binocular asymmetry. In particular, this paper considers view extrapolation since it introduces much more artifacts due to the lack of data. We measure structural similarity on the critical regions between the left and right images to quantify the perceptual effects of binocular asymmetry. The effectiveness of the proposed quality measure has been successfully evaluated by subjective assessment experiments using various types of synthesized stereoscopic images generated by four different DIBR-based view synthesis algorithms. We demonstrate the validity of the proposed quality measure by the comparison of subjective ratings and existing objective methods. Experimental results show that the combined use of the proposed binocular asymmetry measure and existing quality measures substantially improves the performance of the quality measures by explicitly considering the perceptual effects of CBA. Yong Ju Jung, Hak Gu Kim, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Temporally consistent hole filling method based on global optimization with label propagation for 3D videoabstractThis paper presents a new temporally consistent hole filling method based on global optimization for a synthesized 3D vid-eo.1For the temporal consistency, the proposed method adaptively utilizes the already filled region in a previous frame under the guidance of motion vectors to fill a hole region in a current frame (i.e., target frame to-be-filled). In addition, when filling the hole region in the target frame, reliable labels are stored and propagated to a next target frame in order to reduce the computational cost of the global optimization. Experimental results show that the proposed method can achieve the temporal consistency and a high computational gain than existing hole filling methods. Hak Gu Kim, Soo Sung Yoon, Yong Man Ro |
ICIP | 1 |
| 2015 | Multi-frame de-raining algorithm using a motion-compensated non-local mean filter for rainy video sequences
Hak Gu Kim, Seung Ji Seo, Byung Cheol Song |
J. Vis. Commun. Image Represent. | 1 |