EDBT 2026 Demo / reviewers in the wild / expert
Sanghoon Lee 0001
dblp:58/6214-1
· DBLP profile ↗
158ranked-venue papers
12as first author
40since 2021 · last 2026
0000-0001-9895-5347ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 104 · 11 first-author · 27 since 2021Computer networks · 28 · 2 since 2021Artificial intelligence and machine learning · 24 · 16 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DipGuava: Disentangling Personalized Gaussian Features for 3D Head Avatars from Monocular VideoabstractWhile recent 3D head avatar creation methods attempt to animate facial dynamics, they often fail to capture personalized details, limiting realism and expressiveness. To fill this gap, we present DipGuava (Disentangled and Personalized Gaussian UV Avatar), a novel 3D Gaussian head avatar creation method that successfully generates avatars with personalized attributes from monocular video. DipGuava is the first method to explicitly disentangle facial appearance into two complementary components, trained in a structured two-stage pipeline that significantly reduces learning ambiguity and enhances reconstruction fidelity. In the first stage, we learn a stable geometry-driven base appearance that captures global facial structure and coarse expression-dependent variations. In the second stage, the personalized residual details not captured in the first stage are predicted, including high-frequency components and nonlinearly varying features such as wrinkles and subtle skin deformations. These components are fused via dynamic appearance fusion that integrates residual details after deformation, ensuring spatial and semantic alignment. This disentangled design enables DipGuava to generate photorealistic, identity-preserving avatars, consistently outperforming prior methods in both visual quality and quantitative performance, as demonstrated in extensive experiments. Jeonghaeng Lee, Seokkeun Choi, Weisi Lin, Sanghoon Lee 0001 |
AAAI | 5 |
| 2026 | DMGNet: Discriminative multi-view geometry learning with hybrid-domain enhancement for light field occlusion removal
Jieyu Chen, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Ce Zhu, Sanghoon Lee 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Structure and sensitivity in 3D human pose similarity quantification and estimation
Kyoungoh Lee, Jungwoo Huh, Jiwoo Kang 0001, Sanghoon Lee 0001 |
Pattern Recognit. | 4 |
| 2026 | SeCo: Semantic-Guided Multimodal Color Splash EffectsabstractColor splash is a widely used image editing effect that highlights selected regions by retaining color while rendering the rest of the image in grayscale. However, existing tools often struggle with achieving high precision, efficiency, and user flexibility in controlling the effect. In this article, we propose Semantic-Guided Multimodal Color Splash Effects (SeCo), a novel framework for generating stylized and customizable color splash effects from natural language instructions and color palettes. SeCo decomposes the task into two key components: Semantic-Guided Object Isolation (SGOI) and Palette-Driven Color Adjustment (PDCA). SGOI accurately identifies and isolates user-referred objects with fine-grained transparency, while the PDCA module recolors the isolated regions under user-specified palette guidance. Our approach supports arbitrary object selection, handles transparency, and enables diverse stylization patterns. Experimental results on both synthetic and real-world datasets demonstrate that SeCo outperforms existing methods in precision and controllability, offering a practical and expressive solution for visual editing and content creation. Jing-Xuan Chen, Ling Lo, Si-Yu Lu, Wen-Huang Cheng, Jungwoo Huh, Sanghoon Lee 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2026 | Learning a Domain-Specialized Network for Light Field Spatial-Angular Super-ResolutionabstractLight field (LF) imaging is inherently constrained by the trade-off between spatial resolution and angular sampling density. To overcome this obstacle, spatial-angular super-resolution (SR) methods have been developed to achieve concurrent enhancement in both dimensions. Traditional spatial-angular SR methods treat spatial and angular SR as separate tasks, resulting in parameter redundancy and error accumulation. While recent end-to-end approaches attempt joint processing, their uniform treatment of these distinct problems overlooks critical domain-specific requirements. To address these challenges, we propose a domain-specialized framework that deploys stage-tailored strategies to satisfy domain-specific demands. Specifically, in the angular SR stage, we introduce a cross-view consistency modulation module that enhances inter-view coherence through long-range dependency modeling of angular features. In the spatial SR stage, we propose a detail-aware state space model to reconstruct fine-grained detail. Finally, we develop a cross-domain integration module that explores spatial-angular correlations by fusing multi-representational features from both domains to foster synergistic optimization. Experimental results on public LF datasets demonstrate substantial improvements over state-of-the-art methods in both qualitative and quantitative comparisons, with approximately 50% fewer model parameters compared to competing methods. Xinpeng Huang, Deyang Liu, Ping An 0001, Sanghoon Lee 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | SDAS: Semantic Data Acquisition System for Minimizing Redundancy and Maximizing DiversityabstractIn this paper, we propose SDAS, a new motion data assessment and storage system designed to acquire new motion data with reduced redundancy and maximizing diversity. SDAS collects data in the field, retrieves the most similar data from the database in real-time, and provides visualization tools that allow for the comparison of differences between the capture data and the stored data. Through this system, researchers can efficiently build and manage a database. The demonstration video is available at https://youtu.be/vqW0uMDnZTw. Yeseung Park, Hyunse Yoon, Jungwoo Huh, Jungsu Kim, Jeongwook Choi, Sanghoon Lee 0001 |
AAAI | 6 |
| 2025 | Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURAabstractAmodal segmentation aims to infer the complete shape of occluded objects, even when the occluded region's appearance is unavailable. However, current amodal segmentation methods lack the capability to interact with users through text input and struggle to understand or reason about implicit and complex purposes. While methods like LISA integrate multi-modal large language models (LLMs) with segmentation for reasoning tasks, they are limited to predicting only visible object regions and face challenges in handling complex occlusion scenarios. To address these limitations, we propose a novel task named amodal reasoning segmentation, aiming to predict the complete amodal shape of occluded objects while providing answers with elaborations based on user text input. We develop a generalizable dataset generation pipeline and introduce a new dataset focusing on daily life scenarios, encompassing diverse real-world occlusions. Furthermore, we present AURA (Amodal Understanding and Reasoning Assistant), a novel model with advanced global and spatial-level designs specifically tailored to handle complex occlusions. Extensive experiments validate AURA's effectiveness on the proposed dataset. Hyunse Yoon, Sanghoon Lee 0001, Weisi Lin |
ICCV | 3 |
| 2025 | Relightable and Dynamic Gaussian Avatar Reconstruction from Monocular VideoabstractModeling relightable and animatable human avatars from monocular video is a long-standing and challenging task. Recently, Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) methods have been employed to reconstruct the avatars. However, they often produce unsatisfactory photo-realistic results because of insufficient geometrical details related to body motion, such as clothing wrinkles. In this paper, we propose a 3DGS-based human avatar modeling framework, termed as Relightable and Dynamic Gaussian Avatar (RnD-Avatar), that presents accurate pose-variant deformation for high-fidelity geometrical details. To achieve this, we introduce dynamic skinning weights that define the human avatar's articulation based on pose while also learning additional deformations induced by body motion. We also introduce a novel regularization to capture fine geometric details under sparse visual cues. Furthermore, we present a new multi-view dataset with varied lighting conditions to evaluate relight. Our framework enables realistic rendering of novel poses and views while supporting photo-realistic lighting effects under arbitrary lighting conditions. Our method achieves state-of-the-art performance in novel view synthesis, novel pose rendering, and relighting. Seonghwa Choi, Moonkyeong Choi, Mingyu Jang, Jaekyung Kim, Jianfei Cai 0001, Wen-Huang Cheng, Sanghoon Lee 0001 |
ACM Multimedia | 7 |
| 2025 | Permission to Dance: An End-to-End Dance Enhancement System from Dance Capture to AnalysisabstractIn this demonstration, we present Permission to Dance, an end-to-end dance enhancement system designed to capture, enhance, and analyze user's dance performance. Our system consists of a dance capture module, a dance enhancement module, and a dance feedback module. Using the system, users can acquire their dance data in an enhanced version, followed by textual feedback on how to achieve better dance performance. The demonstration video is available at https://youtu.be/lFw7Xic48KU Jungsu Kim, Jungwoo Huh, Yeseung Park, Seongjean Kim, Jeongwook Choi, Sanghoon Lee 0001 |
ACM Multimedia | 6 |
| 2025 | Domain Crossover Non-Rigid Registration for 3D Human MeshesabstractNon-rigid registration is essential for reconstructing dynamic and incomplete 3D human meshes, yet traditional methods often fail to achieve robust alignment in the sequence of high-motion deformations and missing geometry. We propose a domain crossover non-rigid registration (DCNRR) framework that addresses these challenges by effectively transferring informative features from 2D image space into the 3D mesh domain of three key stages: multi-view projection, hierarchical non-rigid registration, and topology-consistent completion. In the first stage, multi-view projections are used to extract 2D joint locations and deep features, which guide deformation in the 3D space. In the second stage, hierarchical joint priors and deep features collaboratively guide mesh alignment, enabling more accurate deformation in distal regions and complex poses. In the final stage, we apply a diffusion-based completion process in UV coordinates to reconstruct incomplete surface normals and refine missing mesh areas with topological consistency. Our approach achieves highly detailed and perceptually accurate mesh deformation. To validate our approach, we evaluate performance on a newly constructed dynamic human motion (DHM) dataset, as well as public datasets. Our method demonstrates state-of-the-art results in both geometric accuracy and stability, showing particular robustness in dynamic and incomplete mesh sequences. Kyungjune Lee, Seongjean Kim, Hoseok Tong, Hyucksang Lee, Seongmin Lee 0002, Weisi Lin, Ping An 0001, Sanghoon Lee 0001 |
ACM Multimedia | 8 |
| 2025 | Visibility-Aware Multi-View Stereo by Surface Normal Weighting for Occlusion RobustnessabstractRecent learning-based multi-view stereo (MVS) still exhibits insufficient accuracy in large occlusion cases, such as environments with significant inter-camera distance or when capturing objects with complex shapes. This is because incorrect image features extracted from occluded areas serve as significant noise in the cost volume construction. To address this, we propose a visibility-aware MVS using surface normal weighting (SnowMVSNet) based on explicit 3D geometry. It selectively suppresses mismatched features in the cost volume construction by computing inter-view visibility. Additionally, we present a geometry-guided cost volume regularization that enhances true depth among depth hypotheses using a surface normal prior. We also propose intra-view visibility that distinguishes geometrically more visible pixels within a reference view. Using intra-view visibility, we introduce the visibility-weighted training and depth estimation methods. These methods enable the network to achieve accurate 3D point cloud reconstruction by focusing on visible regions. Based on simple inter-view and intra-view visibility computations, SnowMVSNet accomplishes substantial performance improvements relative to computational complexity, particularly in terms of occlusion robustness. To evaluate occlusion robustness, we constructed a multi-view human (MVHuman) dataset containing general human body shapes prone to self-occlusion. Extensive experiments demonstrated that SnowMVSNet significantly outperformed state-of-the-art methods in both low- and high-occlusion scenarios. Hyucksang Lee, Seongmin Lee 0002, Sanghoon Lee 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Perceptually-Guided VR Style TransferabstractVirtual reality (VR) makes it possible to provide immersive multimedia content composed of omnidirectional videos (ODVs). Towards enabling more immersive and satisfying VR content, methods are needed to manipulate VR scenes, taking into account perceptual factors related to viewers' quality of experience (QoE). For example, style transfer methods can be applied to VR content, allowing users to create artistic or surreal effects in their immersive environments. Here, we study perceptual factors that affect the sensation of stylized immersiveness, including color dynamics and spatio-temporal consistency. To do this, we introduce an immersiveness sensitivity model of luminance and color perception, and use it to measure the color dynamics and spatio-temporal consistency of stylized VR contents. We subsequently use this model to construct a perceptually-guided VR style transfer model called VR Style Transfer GAN (VRST-GAN). VRST-GAN learns to transfer a desired style into VR to enhance immersiveness by considering color dynamics while preserving spatio-temporal consistency. We demonstrate the effectiveness of VRST-GAN via qualitative and quantitative experiments. We also develop a VR Immersiveness Predictor (VR-IP) that is able to predict the sensation of immersiveness using the perceptual model. In our experiments, VR-IP predicts immersiveness with an accuracy of 91%. Seonghwa Choi, Jungwoo Huh, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2025 | A Novel Intelligent Video Surveillance System Using Low-Traffic Scene-Preserving Video AnonymizationabstractWith the development of computer vision technology, intelligent video surveillance systems have been developed for automatic monitoring. However, the problem of personal information protection has also emerged. Existing systems attempted to solve this problem by anonymizing a video by, for example, sending only low-dimensional abstract information such as a person’s 2D pose or blurring a person’s face in the video before sending it to the central cloud server. However, these approaches failed to balance scene-preservation and traffic efficiency, because abstract information is too limited for preserving the entire scene, and video modification generates massive traffic. This article proposes a novel intelligent video surveillance system to overcome such limitations that preserves the scene information and generates minimal traffic through video anonymization. The proposed system reconstructs 3D human models and estimates segmentation masks to preserve a scene captured by a surveillance camera in its entirety. Parametric models represent 3D human models with several sets of parameters, and dictionary coding compresses the segmentation mask with a high compression ratio. The system follows the edge-cloud architecture, where the edge node extracts and transmits the scene information and the central cloud server generates the final anonymized video. We demonstrate the effectiveness of the proposed system by conducting experiments on processing time, scene preservation, and traffic efficiency. Our proposed system runs in real-time ( \(>\) 25fps) in a typical hardware setting and has a data compression ratio of more than 5,000 compared with raw data transfer while maintaining over 85% scene-preservation correlation with the original video. Jungwoo Huh, Jiwoo Kang 0001, Jongwook Woo, Sanghoon Lee 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2025 | DMESH: A Structure-Preserving Diffusion Model for 3-D Mesh DenoisingabstractDenoising diffusion models have shown a powerful capacity for generating high-quality image samples by progressively removing noise. Inspired by this, we present a diffusion-based mesh denoiser that progressively removes noise from mesh. In general, the iterative algorithm of diffusion models attempts to manipulate the overall structure and fine details of target meshes simultaneously. For this reason, it is difficult to apply the diffusion process to a mesh denoising task that removes artifacts while maintaining a structure. To address this, we formulate a structure-preserving diffusion process. Instead of diffusing the mesh vertices to be distributed as zero-centered isotopic Gaussian distribution, we diffuse each vertex into a specific noise distribution, in which the entire structure can be preserved. In addition, we propose a topology-agnostic mesh diffusion model by projecting the vertex into multiple 2-D viewpoints to efficiently learn the diffusion using a deep network. This enables the proposed method to learn the diffusion of arbitrary meshes that have an irregular topology. Finally, the denoised mesh can be obtained via refinement based on 2-D projections obtained from reverse diffusion. Through extensive experiments, we demonstrate that our method outperforms the state-of-the-art mesh denoising methods in both quantitative and qualitative evaluations. Seongmin Lee 0002, Suwoong Heo, Sanghoon Lee 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | 3D Facial Shape Similarity with Deep Perceptual RepresentationsabstractComparing different 3D shapes is challenging due to their irregularities. Motivated by the human visual system mechanism, where the entire 3D geometry is clearly perceived as a series of multiple projections, we propose a novel facial shape similarity measurement using multiview deep perceptual representations. We introduce a multiview disentangling scheme that accurately represents a facial mesh in multiple coordinates and the training strategy with view specificity and regional consistency to reliably train the network with multiple projections. View specificity pertains to the human visual perception to better recognize facial similarity. Regional consistency mitigates regional redundancy among views. Hence, robust perceptual features with respect to views are embedded and accurate similarity can be measured. Consequently, the view-specific integration scheme incorporates the similarities of all views, allowing for highly consistent measurement. The experiments demonstrate that the proposed similarity outperforms state-of-the-arts and significantly improves the details in terms of geometry and human perception. Seongmin Lee 0002, Jiwoo Kang 0001, Sanghoon Lee 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Towards 360 VR Sickness Mitigation: From Virtual Reality Eye-Tracking to Visual CommunicationabstractMost 360 virtual reality (VR) contents have been developed without considering that users could be affected by VR sickness. Accordingly, users' viewing safety has been steadily highlighted as a critical problem in the VR market. In this study, we investigate a novel VR sickness mitigation framework based on human visual characteristics for the rendered VR content. First, we build a large-scale 360 VR content database termed VRSP360 (VR Sickness and Presence 360) dedicated to the analysis of VR sickness and thoroughly conduct eye-tracking experiments to measure human perception. In the experiment, we observe that the users' gaze distribution is highly center-biased when they experience excessive VR sickness. From this observation, we design a foveated filtering framework that limits high-frequency textures in the peripheral view to mitigate VR sickness. Particularly, given the human visual system's (HVS) non-uniform resolution with respect to the fovea, we also adopt the foveation-based filtering method using the trade-off between sickness mitigation and presence conservation, which reduces any loss in perceptual quality despite the filtering. We further demonstrate that our framework can effectively compress visual information by applying foveated compression. In addition, we develop two metrics (visual texture index and perceptual information index) to measure the effective preservation of user-perceived information despite the filtration of peripheral vision textures by our proposed mitigation method. Through rigorous subjective evaluation on both original content and its VR-sickness-mitigated version, we demonstrate that the proposed framework successfully mitigates VR sickness with a reduction rate of $\sim$∼19% on the proposed dataset. Jeonghaeng Lee, Woojae Kim, Chao Yang 0021, Ping An 0001, Sanghoon Lee 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Speech-Driven Emotional 3d Talking Face Animation Using Emotional EmbeddingsabstractExisting emotional talking 3D facial animation primarily focus on animating emotional faces using a specific emotion condition. However, in real-world situations, no one consistently speaks with just one emotion. Thus, previous emotion-based approaches have very limited applicability in real-world applications. To address this issue, we propose SDETalk, a novel learning framework that animates the emotional talking faces by leveraging the emotional source from a speech. Unlike previous studies, which use static one-hot emotion conditions, the proposed network regresses complex emotional states from speech. It enables the network to animate natural facial animation from an emotional speech without using a specific emotional condition. Furthermore, we design the proposed method to produce head motions because head motion is an important factor to enhance the naturalness of talking face animation. By doing this, our approach simultaneously achieves accurate lip motion, natural expressions, and rhythmical head motions from emotional speech. Through extensive experiments in both qualitative and quantitative manners, it is demonstrated that our method outperforms other state-of-the-art methods by animating realistic and expressive 3D faces. Seongmin Lee 0002, Jeonghaeng Lee, Hyewon Song, Sanghoon Lee 0001 |
ICASSP | 4 |
| 2024 | InViTe: Individual Virtual Transfer for Personalized 3D Face Generation System
Mingyu Jang, Kyungjune Lee, Seongmin Lee 0002, Hoseok Tong, Juwan Chung, Yusung Ro, Sanghoon Lee 0001 |
IJCAI | 7 |
| 2024 | AVIN-Chat: An Audio-Visual Interactive Chatbot System with Emotional State Tuning
Chanhyuk Park, Jungbin Cho, Junwan Kim, Seongmin Lee 0002, Jungsu Kim, Sanghoon Lee 0001 |
IJCAI | 6 |
| 2024 | DanceMimic: Awaken Your Dancing Instinct through a Real-time Dance Imitation Capture System
Seongjean Kim, Jungwoo Huh, Yeseung Park, Jungsu Kim, Sanghoon Lee 0001 |
ACM Multimedia | 5 |
| 2024 | Double reverse diffusion for realistic garment reconstruction from images
Jeonghaeng Lee, Jongyoo Kim, Jiwoo Kang 0001, Sanghoon Lee 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | 3D-PSSIM: Projective Structural Similarity for 3D Mesh Quality Assessment Robust to Topological IrregularitiesabstractDespite acceleration in the use of 3D meshes, it is difficult to find effective mesh quality assessment algorithms that can produce predictions highly correlated with human subjective opinions. Defining mesh quality features is challenging due to the irregular topology of meshes, which are defined on vertices and triangles. To address this, we propose a novel 3D projective structural similarity index ( 3D- PSSIM) for meshes that is robust to differences in mesh topology. We address topological differences between meshes by introducing multi-view and multi-layer projections that can densely represent the mesh textures and geometrical shapes irrespective of mesh topology. It also addresses occlusion problems that occur during projection. We propose visual sensitivity weights that capture the perceptual sensitivity to the degree of mesh surface curvature. 3D- PSSIM computes perceptual quality predictions by aggregating quality-aware features that are computed in multiple projective spaces onto the mesh domain, rather than on 2D spaces. This allows 3D- PSSIM to determine which parts of a mesh surface are distorted by geometric or color impairments. Experimental results show that 3D- PSSIM can predict mesh quality with high correlation against human subjective judgments, across the presence of noise, even when there are large topological differences, outperforming existing mesh quality assessment models. Seongmin Lee 0002, Jiwoo Kang 0001, Sanghoon Lee 0001, Weisi Lin, Alan C. Bovik |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Kinematic Diversity and Rhythmic Alignment in Choreographic Quality Transformers for Dance Quality AssessmentabstractIn recent years, the dance entertainment industry has experienced significant growth, driven by the desire of consumers to learn and improve their dancing skills. To effectively improve their skills, dancers require evaluation and feedback, which traditionally relies heavily on professional dancers. To address this challenge, researchers have proposed objective assessment methods for dance performance via kinematic data captured by sensors. However, these existing methods primarily focus on assessing the rhythmic accuracy of movements synchronized to music. In this paper, we propose Dance Quality Assessment (DanceQA) Framework to evaluate dance performance, considering choreographic factors that are important criteria in subjective DanceQA. We find that kinematic diversity and rhythmic alignment are significant choreographic factors from human perception perspective. Based on these factors, we design two metrics: kinematic information entropy (KIE) and kinematic-music beat similarity (BSIM). Our study demonstrates that these metrics are closely related to specific body parts in each choreography. To validate the effectiveness of our metrics, we capture dance performance by OptiTrack system providing precise three-dimensional data at very high sampling rate. We then label their dance quality via subjective test. The metrics give strong correlation with subjective opinion, but it is difficult to tell which body part is the most correlated. To comprehensively understand the dance quality, we propose choreographic quality transformers (CQTs), which learn the aforementioned choreographic factors by embedding KIE and BSIM into attention matrices. In numerous experiments, the CQTs outperforms previous methods, graph convolutional networks and multimodal transformers, at least by up to 0.146 in correlation coefficient. Taewan Kim 0002, Inwoong Lee, Sanghoon Lee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Single-Image 3-D Reconstruction: Rethinking Point Cloud DeformationabstractSingle-image 3-D reconstruction has long been a challenging problem. Recent deep learning approaches have been introduced to this 3-D area, but the ability to generate point clouds still remains limited due to inefficient and expensive 3-D representations, the dependency between the output and the number of model parameters, or the lack of a suitable computing operation. In this article, we present a novel deep-learning-based method to reconstruct a point cloud of an object from a single still image. The proposed method can be decomposed into two steps: feature fusion and deformation. The first step extracts both global and point-specific shape features from a 2-D object image, and then injects them into a randomly generated point cloud. In the second step, which is deformation, we introduce a new layer termed as GraphX that considers the interrelationship between points like common graph convolutions but operates on unordered sets. The framework can be applicable to realistic image data with background as we optionally learn a mask branch to segment objects from input images. To complement the quality of point clouds, we further propose an objective function to control the point uniformity. In addition, we introduce different variants of GraphX that cover from best performance to best memory budget. Moreover, the proposed model can generate an arbitrary-sized point cloud, which is the first deep method to do so. Extensive experiments demonstrate that we outperform the existing models and set a new height for different performance metrics in single-image 3-D reconstruction. Seonghwa Choi, Woojae Kim, Jongyoo Kim, Heeseok Oh, Jiwoo Kang 0001, Sanghoon Lee 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Fusing Explicit and Implicit Flow for Optical Flow EstimationabstractEstimating optical flow for large movement remains a challenging issue due to inconsistency in features between frames. To resolve this challenge, we propose a novel sequence-based deep learning network that jointly trains explicit flow and implicit flow to accurately estimate optical flow. To do this, we implemented three submodules: explicit flow embedder, implicit flow embedder, and flow fusion network. Explicit flow embedder learns the pair-wise correlation between visible pixels based on the spatial attention made per image. Implicit flow embedder learns implicit flow based on the temporal context of motion from all frames in the sequence. To effectively learn the implicit flow, we give a longer sequence of frames as input. Flow fusion network fuses features from explicit and implicit embedder to output the final optical flow. Through extensive experiments, our model demonstrates its robustness against the large motion while providing accurate flow estimation for pixels without pairs in the next frame. Hyunse Yoon, Seongmin Lee 0002, Sanghoon Lee 0001 |
ICIP | 3 |
| 2023 | Unlocking Potential of 3D-aware GAN for More Expressive Face GenerationabstractAs style-based image generators have achieved disentanglement in features by converting latent vector space to style vector space, numerous efforts have been made to enhance the controllability of the latent. However, existing methods for controllable models have limitations in precisely creating high-resolution faces with large expressions. The degradation is due to the dependence on the training dataset, as the high-resolution face datasets do not have sufficient expressive images. To tackle this challenge, we propose a robust training framework for 3D-aware generative adversarial networks to learn the high-quality generation of more expressive faces through a signed distance field. First, we propose a novel 3D enforcement loss to generate more expressive images in an unsupervised manner. Second, we introduce a partial training method to fine-tune the network on multiple datasets without loss of image resolution. Finally, we propose a ray-scaling scheme for the volume renderer to represent a face at arbitrary scales. Through the proposed framework, the network learns 3D face priors, such as expressional shapes of the parametric facial model, to generate detailed faces. The experimental results outperform the methods of the state of the art, showing strong benefits in the generation of high-resolution facial expressions. Juheon Hwang, Jiwoo Kang 0001, Kyoungoh Lee, Sanghoon Lee 0001 |
ICMR | 4 |
| 2023 | Video-Based Stabilized 3D Face Alignment Using Temporal Multi-DiscriminationabstractExisting 3D face alignment primarily aim to achieve accurate face alignment result for a static facial image. While these methods have strong alignment performance under large poses, occlusion, and extreme lighting conditions, they often result in trembling artifacts in video-based sequential 3D face alignment. Reducing temporal misalignment remains a challenging task because a single misaligned frame can propagate errors to other frames along the temporal axis. To address this issue, we propose a novel temporal discriminating scheme that learns the distribution gap between the face alignment results and ground truth face animation. By leveraging the discrimination results as a guide, the proposed method can effectively align the 3D faces to the input video by reducing temporal trembling artifacts. To effectively learn the distribution gap, we introduce a multi-discriminating scheme that separately discriminates facial animation based on identity and expression changes. It enables the proposed method to produce a stabilized alignment result, especially in dynamic and fast movement. Through extensive experiments in both qualitative and quantitative evaluations, it is confirmed that our method outperforms state-of-the-art 3D face alignment methods by animating stabilized results in the video. Seongmin Lee 0002, Hyunse Yoon, Jiwoo Kang 0001, Jungsu Kim, Jiwan Son, Jungwoo Huh, Sanghoon Lee 0001 |
MMSP | 7 |
| 2023 | MNET++: Music-Driven Pluralistic Dancing Toward Multiple Dance Genre SynthesisabstractNumerous task-specific variants of autoregressive networks have been developed for dance generation. Nonetheless, a severe limitation remains in that all existing algorithms can return repeated patterns for a given initial pose, which may be inferior. We examine and analyze several key challenges of previous works, and propose variations in both model architecture (namely MNET++) and training methods to address these. In particular, we devise the beat synchronizer and dance synthesizer. First, generated dance should be locally and globally consistent with given music beats, circumvent repetitive patterns, and look realistic. To achieve this, the beat synchronizer implicitly catches the rhythm enabling it to stay in sync with the music as it dances. Then, the dance synthesizer infers the dance motions in a seamless patch-by-patch manner conditioned by music. Second, to generate diverse dance lines, adversarial learning is performed by leveraging the transformer architecture. Furthermore, MNET++ learns a dance genre-aware latent representation that is scalable for multiple domains to provide fine-grained user control according to the dance genre. Compared with the state-of-the-art methods, our method synthesizes plausible and diverse outputs according to multiple dance genres as well as generates remarkable dance sequences qualitatively and quantitatively. Jinwoo Kim 0005, Beom Kwon, Jongyoo Kim, Sanghoon Lee 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | From Human Pose Similarity Metric to 3D Human Pose Estimator: Temporal Propagating LSTM NetworksabstractPredicting a 3D pose directly from a monocular image is a challenging problem. Most pose estimation methods proposed in recent years have shown 'quantitatively' good results (below ∼ 50mm). However, these methods remain 'perceptually' flawed because their performance is only measured via a simple distance metric. Although this fact is well understood, the reliance on 'quantitative' information implies that the development of 3D pose estimation methods has been slowed down. To address this issue, we first propose a perceptual Pose SIMilarity (PSIM) metric, by assuming that human perception (HP) is highly adapted to extracting structural information from a given signal. Second, we present a perceptually robust 3D pose estimation framework: Temporal Propagating Long Short-Term Memory networks (TP-LSTMs). Toward this, we analyze the information-theory-based spatio-temporal posture correlations, including joint interdependency, temporal consistency, and HP. The experimental results clearly show that the proposed PSIM metric achieves a superior correlation with users' subjective opinions than conventional pose metrics. Furthermore, we demonstrate the significant quantitative and perceptual performance improvements of TP-LSTMs compared to existing state-of-the-art methods. Kyoungoh Lee, Woojae Kim, Sanghoon Lee 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Progressive Contextual Aggregation Empowered by Pixel-Wise Confidence Scoring for Image InpaintingabstractImage inpainting methods leverage the similarity of adjacent pixels to create alternative content. However, as the invisible region becomes larger, the pixels completed in the deeper hole are difficult to infer from the surrounding pixel signal, which is more prone to visual artifacts. To help fill this void, we adopt an alternative progressive hole-filling scheme that hierarchically fills the corrupted region in the feature and image spaces. This technique allows us to utilize reliable contextual information of the surrounding pixels, even for large hole samples, and then gradually complete the details as the resolution increases. For a more realistic representation of the completed region, we devise a pixel-wise dense detector. By distinguishing each pixel as whether it is a masked region or not, and passing the gradient to all resolutions, the generator further enhances the potential quality of the compositing. Furthermore, the completed images at different resolutions are then merged using a proposed structure transfer module (STM) that incorporates fine-grained local and coarse-grained global interactions. In this new mechanism, each completed image at the different resolutions attends its closest composition at fine granularity adjacent image and thus can capture the global continuity by interacting both short- and long-range dependencies. By comparing our solutions qualitatively and quantitatively with state-of-the-art methods, we conclude that our model exhibits a significantly improved visual quality, even in the case of large holes. Jinwoo Kim 0005, Woojae Kim, Heeseok Oh, Sanghoon Lee 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | A Brand New Dance Partner: Music-Conditioned Pluralistic Dancing Controlled by Multiple Dance GenresabstractWhen coming up with phrases of movement, choreographers all have their habits as they are used to their skilled dance genres. Therefore, they tend to return certain patterns of the dance genres that they are familiar with. What if artificial intelligence could be used to help choreographers blend dance genres by suggesting various dances, and one that matches their choreographic style? Numerous task-specific variants of autoregressive networks have been developed for dance generation. Yet, a serious limitation remains that all existing algorithms can return repeated patterns for a given initial pose sequence, which may be inferior. To mitigate this issue, we propose MNET, a novel and scalable approach that can perform music-conditioned pluralistic dance generation synthesized by multiple dance genres using only a single model. Here, we learn a dancegenre aware latent representation by training a conditional generative adversarial network leveraging Transformer architecture. We conduct extensive experiments on AIST++ along with user studies. Compared to the state-of-the-art methods, our method synthesizes plausible and diverse outputs according to multiple dance genres as well as generates outperforming dance sequences qualitatively and quantitatively. Jinwoo Kim 0005, Heeseok Oh, Seongjean Kim, Hoseok Tong, Sanghoon Lee 0001 |
CVPR | 5 |
| 2022 | Gradient Flow Evolution for 3D Fusion From a Single Depth SensorabstractWe present a novel real-time framework for non-rigid 3D reconstruction that is robust to noise, camera poses, and large deformation from a single depth camera. KinectFusion has achieved high-quality 3D object reconstructions in real-time by implicitly representing an object’s surface with a signed distance field (SDF) representation from a single depth camera. Many studies for incremental reconstruction have been presented since then, with the surface estimation improving over time. Previous works primarily focused on improving conventional SDF matching and deformation schemes. In contrast to these works, the proposed framework tackles the problem of temporal inconsistency caused by SDF approximation and fusion to manipulate SDFs and reconstruct a target more accurately over time. In our reconstruction pipeline, we introduce a refinement evolution method, where an erroneous SDF from a depth sensor is recovered more accurately in a few iterations by propagating erroneous SDF values from the surface. Reliable gradients of refined SDFs enable more accurate non-rigid tracking of a target object. Furthermore, we propose a level-set evolution for SDF fusion, enabling SDFs to be manipulated stably in the reconstruction pipeline over time. The proposed methods are fully parallelizable and can be executed in real-time. Qualitative and quantitative evaluations show that incorporating the refinement and fusion methods into the reconstruction pipeline improves 3D reconstruction accuracy and temporal reliability by avoiding cumulative errors over time. Evaluation results show that our pipeline results in more accurate reconstruction that is robust to noise and large motions, as well as outperforms previous state-of-the-art reconstruction methods. Jiwoo Kang 0001, Seongmin Lee 0002, Mingyu Jang, Sanghoon Lee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Self-Updatable Database System Based on Human Motion Assessment FrameworkabstractRecently, human motion-centric videos have been attracting attention in the field of computer vision. Observing and detecting human motion in intelligent surveillance camera systems is essential for understanding the intentions of target subjects. However, these videos have vast amounts of disparate and complex information, and hence they are difficult to process and label automatically. As a result, building and maintaining a database using motion-centric videos requires considerable labor in trimming and classifying the videos. Therefore, we propose a self-updatable motion database system based on a human motion assessment framework for evaluating complex human movements. The framework quantifies three primitive motion properties: stability, liveliness, and attention. This assessment highlights the semantics of human motion in the input video. The semantic motion sequence obtained after the motion assessment is compared with a similarity motion database to determine whether the database needs to be updated; for efficient comparison, we introduce a sequential autoencoder model with a long short-term memory neural network. The proposed system maintains the database within a surveillance camera system using a motion update algorithm; unseen motions in the database are updated using a camera-based surveillance system. In addition, this framework combines state-of-art action recognition methods to improve performance by up to 11% via the self-update of motion. Kyoungoh Lee, Yeseung Park, Jungwoo Huh, Jiwoo Kang 0001, Sanghoon Lee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | A Deep Motion Sickness Predictor Induced by Visual Stimuli in Virtual RealityabstractIn a virtual reality (VR) environment, where visual stimuli predominate over other stimuli, the user experiences cybersickness because the balance of the body collapses due to self-motion. Accordingly, the VR experience is accompanied by unavoidable sickness referred to as visually induced motion sickness (VIMS). In this article, our primary purpose is to simultaneously estimate the VIMS score by referring to the content and calculate the temporally induced VIMS sensitivity. To seek our goals, we propose a novel architecture composed of two consecutive networks: 1) neurological representation and 2) spatiotemporal representation. In the first stage, the network imitates and learns the neurological mechanism of motion sickness. In the second stage, the significant feature of the spatial and temporal domains is expressed over the generated frames. After the training procedure, our model can calculate VIMS sensitivity for each frame of the VR content by using the weakly supervised approach for unannotated temporal VIMS scores. Furthermore, we release a massive VR content database. In the experiments, the proposed framework demonstrates excellent performance for VIMS score prediction compared with existing methods, including feature engineering and deep learning-based approaches. Furthermore, we propose a way to visualize the cognitive response to visual stimuli and demonstrate that the induced sickness tends to be activated in a similar tendency, as done in clinical studies. Jinwoo Kim 0005, Heeseok Oh, Woojae Kim, Seonghwa Choi, Wookho Son, Sanghoon Lee 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Competitive Learning of Facial Fitting and Synthesis Using UV EnergyabstractThe three-dimensional morphable model (3DMM) is the most widely used representative model for obtaining a three-dimensional (3-D) face from a target on an image. Although 3DMMs have demonstrated the powerful capability to represent various facial shapes on natural images, they are limited to capturing texture variations of in-the-wild human faces. Based on the fact that fitting a 3-D facial model to an image determines the corresponding UV map, we propose a novel method for facial fitting and synthesis by competitively training two deep learning networks for facial alignment and UV texture completion. When the completion network is trained using well-aligned UV maps, it can model facial textures precisely and, consequently, fill the missing regions more completely. Accordingly, we use a UV completion network, denoted as a UV energy-based generative adversarial network (UV EB-GAN), to discriminate whether a UV map from the alignment network is well aligned by defining the generative loss of the completion network as the energy. Competitive learning facilitates training the completion network without ground-truth facial UV maps and training the alignment network without hard constraints and regularization terms. The proposed network can be trained in an end-to-end manner. The facial texture, albedo, lighting parameters, and 3-D facial shape can be obtained through this network. The results of the experiments on 2-D alignment, 3-D reconstruction, texture synthesis, and illumination estimation verified that the proposed method achieves remarkable improvements over the state-of-the-art methods. Jiwoo Kang 0001, Seongmin Lee 0002, Sanghoon Lee 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Optimal Camera Point Selection Toward the Most Preferable View of 3-D Human PoseabstractAnswering the question “what is the most preferable view of a three-dimensional (3-D) human model?” is a challenge in computer vision, computer graphics, and cinematography applications because the appearance of a human, for a given pose, relies on the viewpoint of the user. Currently, to the best of the authors’ knowledge, solid research on the most preferable viewing angle for obtaining numerical subjective evaluation scores has not been conducted. In this study, we investigate a metric that can be used to quantify the view of a 3-D human model, whose value is maximized at the most favorable camera angle in accordance with subjective assessments done by users. For an objective assessment in a numerical form, in this study, we define three view selection metrics: the 1)normalized limb length sum; 2)normalized area of a two-dimensional bounding box; and 3)normalized visible area of a 3-D bounding box. Finally, we formulate a viewpoint optimization problem whose objective function is the sum of the metrics. However, the objective function is nonconcave, and the solution set of the constraint is nonconvex. To overcome this difficulty, we employ decomposition and penalty methods. From the simulation results, it is verified that the average of the viewpoint selection error between the ground truth viewpoint and the optimal viewpoint obtained by the proposed algorithm is very close to the lower bound of the viewpoint selection error. Beom Kwon, Jungwoo Huh, Kyoungoh Lee, Sanghoon Lee 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | WarpingFusion: Accurate Multi-View TSDF Fusion with Local Perspective WarpabstractIn this paper, we propose the novel 3D reconstruction framework, where the surface of a target object is reconstructed accurately and robustly from multi-view depth maps. A depth map of a moving object tends to have the spatially-varying perspective warps due to motion blur and rolling shutter artifacts. Incorporating those misaligned points from the views into the world coordinate leads to significant artifacts in the reconstructed shape. We address the mismatches by the patch-based depth-to-surface alignment using implicit surface-based distance measurement. The patch-based minimization finds spatial warps on the depth map fast and accurately with the global transformation preserved. The proposed framework efficiently optimizes the local alignments against depth occlusions and local variants thanks to the point to surface distance based on an implicit representation. The proposed method shows significant improvements over the other reconstruction methods, demonstrating efficiency and benefits of our method in the multi-view reconstruction. Jiwoo Kang 0001, Seongmin Lee 0002, Mingyu Jang, Hyunse Yoon, Sanghoon Lee 0001 |
ICIP | 5 |
| 2021 | Deep Chessboard Corner Detection Using Multi-task LearningabstractCamera calibration is an indispensable step in the fields of robotics and computer vision, which includes augmented reality, 3D reconstruction, and camera motion estimation. Before camera calibration, detecting matching correspondence is necessary to understand the structure of the world from multiple images. For an accurate result, a calibration object, such as a chessboard, is used. Existing handcrafted feature methods precisely detect chessboard corners but are weak against blurs, noises, and severe lens distortion. Conversely, neural network-based methods can detect corners regardless of noises in the image. Both methods do not utilize the information of camera priors, which are lens distortion and intrinsic parameters, affecting the location of chessboard corners. Learning of lens distortion and intrinsic parameters enables the proposed network to understand the alignment of corners more precisely. Therefore, in this paper, we propose a novel multi-task learning framework to detect chessboard corners and simultaneously estimate lens distortion and intrinsic parameters. In order to train these three tasks, synthetic images of the chessboard are generated with ground-truth labels corresponding to each task. Hence, by learning the camera priors, the proposed network can more precisely locate the corners than other state-of-the-art corner detection methods while robust to noises, blurs, and distortion. Hyunse Yoon, Seongmin Lee 0002, Jiwoo Kang 0001, Sanghoon Lee 0001 |
MMSP | 4 |
| 2021 | VR Sickness Versus VR Presence: A Statistical Prediction ModelabstractAlthough it is well-known that the negative effects of VR sickness, and the desirable sense of presence are important determinants of a user's immersive VR experience, there remains a lack of definitive research outcomes to enable the creation of methods to predict and/or optimize the trade-offs between them. Most VR sickness assessment (VRSA) and VR presence assessment (VRPA) studies reported to date have utilized simple image patterns as probes, hence their results are difficult to apply to the highly diverse contents encountered in general, real-world VR environments. To help fill this void, we have constructed a large, dedicated VR sickness/presence (VR-SP) database, which contains 100 VR videos with associated human subjective ratings. Using this new resource, we developed a statistical model of spatio-temporal and rotational frame difference maps to predict VR sickness. We also designed an exceptional motion feature, which is expressed as the correlation between an instantaneous change feature and averaged temporal features. By adding additional features (visual activity, content features) to capture the sense of presence, we use the new data resource to explore the relationship between VRSA and VRPA. We also show the aggregate VR-SP model is able to predict VR sickness with an accuracy of 90% and VR presence with an accuracy of 75% using the new VR-SP dataset. Woojae Kim, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2021 | 3-D Human Behavior Understanding Using Generalized TS-LSTM NetworksabstractThis paper addresses the problems of skeleton feature representation and the modeling of temporal dynamics to recognize human actions consisting of poses. In contrast to traditional methods which generally used relative coordinate systems dependent on some joints, or modeled only the long-term dependency, we attempt to understand 3D human behavior with observation by taking temporally different windows. Instead of taking raw skeletons as the input, we transform the skeletons into another coordinate system to obtain the robustness to scale, rotation and translation, and extract motion features between adjacent skeletons, which finally constructs an efficient hybrid-stream combining both pose and motion streams. We propose novel generalized Temporal Sliding Long Short-term Memory (TS-LSTM) networks. The proposed networks are composed of multiple TS-LSTM networks with various hyper-parameters, which can capture various temporal dynamics of actions. We also propose a novel hyper-parameter searching method, which finds decent hyper-parameters of generalized TS-LSTM to handle temporal dynamics of actions. In the experiment, we evaluate the proposed networks to verify the effectiveness of the proposed methods, and compare them with the other methods on three challenging datasets. Additionally, we analyze a relation between the recognized actions and the hyper-parameters, and visualize the layers of the proposed models. Inwoong Lee, Sanghoon Lee 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | Statistical Convolution On Unordered Point SetabstractIn this paper, we propose a new convolutional layer for neural networks on unordered and irregular point set. Most research advanced to date usually face multiple problem related to point cloud density and may require ad-hoc neural network architectures, which overlooks the huge treasure of architectures from computer vision or language processing. To mitigate these shortcomings, we process a point set at its distribution level by introducing statistical convolution (StatsConv). The spotlight feature of StatsConv is that it extracts various statistics to characterize the distribution of the input point set, which makes it highly scalable compared to existing point convolution operators. StatsConv is fundamentally simple, and can be used as a drop-in in any contemporary neural network architecture with negligible changes. Thorough experiments on point cloud classification and segmentation demonstrate the competence of StatsConv compared to the state of the art. Seonghwa Choi, Woojae Kim, Sanghoon Lee 0001, Weisi Lin |
ICIP | 4 |
| 2020 | PatchMatch based Multiview Stereo with Local Quadric WindowabstractAlthough various stereo matching methods are studied in many years, the accurate 3D reconstruction from multiview stereos in high-fidelity is still challenging due to the surface inconsistency caused by various factors such as specular illumination. In this paper, we propose an accurate PatchMatch based multiview stereo matching method with a quadric support window that efficiently captures the surface of a complex structured object. Our method takes three novel contributions. Firstly, delicate surface configurations are used for representing the complex structure of an object. By using a general 3D quadric function, the structured object surfaces can be estimated more accurately. In addition, an illumination robust framework is proposed, where the patch dissimilarities are precisely measured with disentangled representation. The matching cost is defined based on disentangled measurements of the object photometric and geometric properties, balancing the pixel intensities between images robust to illumination. Lastly, a multiview propagation method is proposed to confirm shape consistency among views. Through the disparity refinement to unify plane parameters of the views, the object surface is estimated from a global perspective. Consequently, the dense and smooth 3D shape of the object is reconstructed accurately. We evaluate our proposed method on the Middlebury stereo set and conduct comprehensive experiments on facial images. Both quantitative and qualitative results demonstrate that the proposed method shows significant improvements over state-of-the-art methods. Hyewon Song, Jaeseong Park, Suwoong Heo, Jiwoo Kang 0001, Sanghoon Lee 0001 |
ACM Multimedia | 5 |
| 2020 | Dynamic Receptive Field Generation for Full-Reference Image Quality AssessmentabstractMost full-reference image quality assessment (FR-IQA) methods advanced to date have been holistically designed without regard to the type of distortion impairing the image. However, the perception of distortion depends nonlinearly on the distortion type. Here we propose a novel FR-IQA framework that dynamically generates receptive fields responsive to distortion type. Our proposed method-dynamic receptive field generation based image quality assessor (DRF-IQA)-separates the process of FR-IQA into two streams: 1) dynamic error representation and 2) visual sensitivity-based quality pooling. The first stream generates dynamic receptive fields on the input distorted image, implemented by a trained convolutional neural network (CNN), then the generated receptive field profiles are convolved with the distorted and reference images, and differenced to produce spatial error maps. In the second stream, a visual sensitivity map is generated. The visual sensitivity map is used to weight the spatial error map. The experimental results show that the proposed model achieves state-of-the-art prediction accuracy on various open IQA databases. Woojae Kim, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2019 | A Simple Way of Multimodal and Arbitrary Style TransferabstractWe re-define multimodality and introduce a simple approach to multimodal and arbitrary style transfer. Conventionally, style transfer methods are limited to synthesizing a deterministic output based on a single style, and there has been no work that can generate multiple images of various details, or multimodality, given a single style. In this work, we explore a way to achieve multimodal and arbitrary style transfer by injecting noise to a unimodal method. This novel approach does not require any trainable parameters, and can be readily applied to any unimodal style transfer methods with separate style encoding sub-network in literature. Experimental results show that while being able to transfer an image to multiple domains in various ways, the image quality is highly competitive with contemporary models in style transfer. Seonghwa Choi, Woojae Kim, Sanghoon Lee 0001 |
ICASSP | 4 |
| 2019 | A Deep Cybersickness Predictor Based on Brain Signal Analysis for Virtual Reality ContentsabstractWhat if we could interpret the cognitive state of a user while experiencing a virtual reality (VR) and estimate the cognitive state from a visual stimulus? In this paper, we address the above question by developing an electroencephalography (EEG) driven VR cybersickness prediction model. The EEG data has been widely utilized to learn the cognitive representation of brain activity. In the first stage, to fully exploit the advantages of the EEG data, it is transformed into the multi-channel spectrogram which enables to account for the correlation of spectral and temporal coefficient. Then, a convolutional neural network (CNN) is applied to encode the cognitive representation of the EEG spectrogram. In the second stage, we train a cybersickness prediction model on the VR video sequence by designing a Recurrent Neural Network (RNN). Here, the encoded cognitive representation is transferred to the model to train the visual and cognitive features for cybersickness prediction. Through the proposed framework, it is possible to predict the cybersickness level that reflects brain activity automatically. We use 8-channels EEG data to record brain activity while more than 200 subjects experience 44 different VR contents. After rigorous training, we demonstrate that the proposed framework reliably estimates cognitive states without the EEG data. Furthermore, it achieves state-of-the-art performance comparing to existing VR cybersickness prediction models. Jinwoo Kim 0005, Woojae Kim, Heeseok Oh, Seongmin Lee 0002, Sanghoon Lee 0001 |
ICCV | 5 |
| 2019 | GraphX-Convolution for Point Cloud Deformation in 2D-to-3D ConversionabstractIn this paper, we present a novel deep method to reconstruct a point cloud of an object from a single still image. Prior arts in the field struggle to reconstruct an accurate and scalable 3D model due to either the inefficient and expensive 3D representations, the dependency between the output and number of model parameters or the lack of a suitable computing operation. We propose to overcome these by deforming a random point cloud to the object shape through two steps: feature blending and deformation. In the first step, the global and point-specific shape features extracted from a 2D object image are blended with the encoded feature of a randomly generated point cloud, and then this mixture is sent to the deformation step to produce the final representative point set of the object. In the deformation process, we introduce a new layer termed as GraphX that considers the inter-relationship between points like common graph convolutions but operates on unordered sets. Moreover, with a simple trick, the proposed model can generate an arbitrary-sized point cloud, which is the first deep method to do so. Extensive experiments verify that we outperform existing models and halve the state-of-the-art distance score in single image 3D reconstruction. Seonghwa Choi, Woojae Kim, Sanghoon Lee 0001 |
ICCV | 4 |
| 2019 | Point Cloud Deformation for Single Image 3d ReconstructionabstractWe propose an approach to reconstruct a precise and dense 3d point cloud from a single image. Previous works employed reconstruction to complexity 3D shape or directly regression location from image. However, while the former requires overhead construction of 3D shape or is inefficient because of high computing cost, the latter does not scale well as the number of trainable parameters depends on the number of output points. In this paper, we explore a method to infer a point cloud representation given an input image. We extract shape information from an input image, and then we embed the two kinds of shape information into the point cloud: point-specific and global shape features. After that, we deform a randomly generated point cloud to the final representation based on the embedded point cloud feature. Our method does not require overhead construction, and is efficient and scalable because the number of trainable parameters is independent of the point cloud size, which is the first work to be able to do so according to our knowledge. Thorough experimental results suggest that our proposed method outperforms with other state-of-the-art methods in dense and precise point cloud generation. Seonghwa Choi, Jinwoo Kim 0005, Sewoong Ahn, Sanghoon Lee 0001 |
ICIP | 5 |
| 2019 | CNN-Based Blind Quality Prediction On Stereoscopic Images Via Patch To Image Feature PoolingabstractIn previous quality assessment studies on stereoscopic 3D (S3D) images, researchers have concentrated on deriving manually extracted features which represent the quality of images. These features are based on the human visual system or natural scene statistics, but they have not been revealed as a deterministic function, preventing to guarantee the robustness of features. To solve this problem, we introduce a deep learning method for predicting the quality of S3D images without a reference. A convolutional neural network (CNN) model is trained through two-step learning. First, to overcome the lack of training data, patch-based CNNs are introduced. And then, automatically extracted patch features are pooled into image features. Finally, the trained CNN model parameters are updated iteratively using holistic image labeling, i.e., mean opinion score (MOS). The proposed method represents a significant improvement compared to other no-reference (NR) S3D image quality assessment (IQA) algorithms. Jinwoo Kim 0005, Sewoong Ahn, Heeseok Oh, Sanghoon Lee 0001 |
ICIP | 4 |
| 2019 | Distribution Padding in Convolutional Neural NetworksabstractEven though zero padding is usually a staple in convolutional neural networks to maintain the output size, it is highly suspicious because it significantly alters the input distribution around border region. To mitigate this problem, in this paper, we propose a new padding technique termed as distribution padding. The goal of the method is to approximately maintain the statistics of the input border regions. We introduce two different ways to achieve our goal. In both approaches, the padded values are derived from the means of the border patches, but those values are handled in a different way in each variant. Through extensive experiments on image classification and style transfer using different architectures, we demonstrate that the proposed padding technique consistently outperforms the default zero padding, and hence can be a potential candidate for its replacement. Seonghwa Choi, Woojae Kim, Sewoong Ahn, Jinwoo Kim 0005, Sanghoon Lee 0001 |
ICIP | 6 |
| 2019 | Deep Visual Saliency on Stereoscopic ImagesabstractVisual saliency on stereoscopic 3D (S3D) images has been shown to be heavily influenced by image quality. Hence, this dependency is an important factor in image quality prediction, image restoration and discomfort reduction, but it is still very difficult to predict such a nonlinear relation in images. In addition, most algorithms specialized in detecting visual saliency on pristine images may unsurprisingly fail when facing distorted images. In this paper, we investigate a deep learning scheme named Deep Visual Saliency (DeepVS) to achieve a more accurate and reliable saliency predictor even in the presence of distortions. Since visual saliency is influenced by low-level features (contrast, luminance and depth information) from a psychophysical point of view, we propose seven low-level features derived from S3D image pairs and utilize them in the context of deep learning to detect visual attention adaptively to human perception. During analysis, it turns out that the low-level features play a role to extract distortion and saliency information. To construct saliency predictors, we weight and model the human visual saliency through two different network architectures, a regression and a fully convolutional neural networks (CNNs). Our results from thorough experiments confirm that the predicted saliency maps are up to 70 % correlated with human gaze patterns, which emphasize the need for the hand-crafted features as input to deep neural networks in S3D saliency detection. Jongyoo Kim, Heeseok Oh, Haksub Kim, Weisi Lin, Sanghoon Lee 0001 |
IEEE Trans. Image Process. | 6 |
| 2019 | Deep CNN-Based Blind Image Quality PredictorabstractImage recognition based on convolutional neural networks (CNNs) has recently been shown to deliver the state-of-the-art performance in various areas of computer vision and image processing. Nevertheless, applying a deep CNN to no-reference image quality assessment (NR-IQA) remains a challenging task due to critical obstacles, i.e., the lack of a training database. In this paper, we propose a CNN-based NR-IQA framework that can effectively solve this problem. The proposed method-deep image quality assessor (DIQA)-separates the training of NR-IQA into two stages: 1) an objective distortion part and 2) a human visual system-related part. In the first stage, the CNN learns to predict the objective error map, and then the model learns to predict subjective score in the second stage. To complement the inaccuracy of the objective error map prediction on the homogeneous region, we also propose a reliability map. Two simple handcrafted features were additionally employed to further enhance the accuracy. In addition, we propose a way to visualize perceptual error maps to analyze what was learned by the deep CNN model. In the experiments, the DIQA yielded the state-of-the-art accuracy on the various databases. Jongyoo Kim, Sanghoon Lee 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Deep Video Quality Assessor: From Spatio-Temporal Visual Sensitivity to a Convolutional Neural Aggregation Network
Woojae Kim, Jongyoo Kim, Sewoong Ahn, Jinwoo Kim 0005, Sanghoon Lee 0001 |
ECCV (1) | 5 |
| 2018 | Propagating LSTM: 3D Pose Estimation Based on Joint Interdependency
Kyoungoh Lee, Inwoong Lee, Sanghoon Lee 0001 |
ECCV (7) | 3 |
| 2018 | Deep Blind Image Quality Assessment by Learning Sensitivity MapabstractApplying a deep convolutional neural network CNN to no-reference image quality assessment (NR-IQA) is a challenging task due to the lack of a training database. In this paper, we propose a CNN-based NR-IQA framework that can effectively solve this problem. The proposed method-the Deep Blind image Quality Assessment predictor (DeepBQA)-adopts two step training stages to avoid overfitting. In the first stage, a ground-truth objective error map is generated and used as a proxy training target. Then, in the second stage, subjective score is predicted by learning a sensitivity map, which weights each pixel in the predicted objective error map. To compensate the inaccurate prediction of the objective error on the homogeneous regions, we additionally suggest a reliability map. Experiments showed that DeepBQA yields a state-of-the-art correlation with human opinions. Jongyoo Kim, Woojae Kim, Sanghoon Lee 0001 |
ICASSP | 3 |
| 2018 | Visual Preference Prediction for Enhanced Images on Ultra-High-Definition DisplayabstractDue to the evolution of ultra-high-definition (UHD) technologies, viewers can enjoy more realistic contents. Furthermore, in order to maximize visual attraction, post-processing is conducted in commercial devices. In this paper, we propose a new terminology called visual preference to quantify viewer's preferences for sharpness- and contrast- enhanced UHD images in a particular viewing geometry. Visual preferences depend on the spatial characteristics and are affected by the viewing geometry of display like resolution, display size, and viewing distance. Therefore, we propose a method called visual preference assessment model that accounts for content enhancement features and diverse viewing geometry. By rigorous experiments, our proposed model outperforms other state-of-the-art models. Sewoong Ahn, Woojae Kim, Jinwoo Kim 0005, Jaekyung Kim, Sanghoon Lee 0001 |
ICIP | 5 |
| 2018 | Deep Blind Video Quality Assessment Based on Temporal Human PerceptionabstractThe high performance video quality assessment (VQA) algorithm is a necessary skill to provide high quality video to viewers. However, since the nonlinear perception function between the distortion level of the video and the subjective quality score is not precisely defined, there are many limitations in accurately predicting the quality of the video. In this paper, we propose a deep learning scheme named Deep Blind Video Quality Assessment (DeepBVQA) to achieve a more accurate and reliable video quality predictor by considering various spatial and temporal cues which have not been considered before. We used CNN to extract the spatial cues of each video in VQA and proposed new hand-crafted features for temporal cues. Performance experiments show that performance is better than other state-of-the-art no-reference (NR) VQA models and the introduction of hand-crafted temporal features is very efficient in VQA. Sewoong Ahn, Sanghoon Lee 0001 |
ICIP | 2 |
| 2018 | Fitting Facial Models to Spatial Points: Blendshape Approaches and BenchmarkabstractBlendshape is one of the most common facial representation used for 3D animation, 3D game and virtual reality. In this paper, four representative blendshape approaches are benchmarked: global, delta, mean-delta, and SVD-based blend-shapes. When fitting the blendshape models to sparse facial points, the obtained facial shape highly depends on fitting approach due to the lack of the fitted points. Therefore, it is important to set up appropriate criteria for comparing and verifying the performance of the approaches. In this paper, we use four kinds of metrics that are utilized to measure the performance of the approaches: fitting, landmark, and vertex errors and coefficient sparsity. Through the experimental results, it is verified that the benchmarks are very effective to measure the subjective quality of blendshape. Taelim Choi, Jiwoo Kang 0001, Hyewon Song, Sanghoon Lee 0001 |
ICIP | 4 |
| 2018 | Multiple Level Feature-Based Universal Blind Image Quality Assessment ModelabstractThe direct use of a deep convolutional neural network (CNN) in no-reference image quality assessment (NR-IQA) usually struggles for a good performance due to a lack of training data, which can be alleviated by transfer learning. However, depending on the similarity between the source and target tasks, the final performance differs vastly. In particular, various kinds of distortion types exist in IQA, which requires different kinds of features to predict visual quality. In this paper, to make the transferred model robust to various distortion types, we propose a Multiple-level Feature-based Image Quality Assessor (MFIQA) which considers multiple levels of features simultaneously. Through rigorous experiments, we prove that MFIQA consistently yields state-of-the-art performance regardless of the distortion types including synthetic and authentic corruption. Jongyoo Kim, Sewoong Ahn, Chong Luo 0001, Sanghoon Lee 0001 |
ICIP | 5 |
| 2018 | Robust Facial Pose Estimation Using Landmark Selection Method for Binocular Stereo VisionabstractIn this paper, we present a robust framework for facial pose estimation from binocular stereoscopic vision. Unlike prior work on the facial pose estimation that employs the whole landmarks even located in the wrong position, we propose a landmark selection method to remove the erroneous landmarks for better performance, especially in the large facial pose case. For this purpose, we train a convolutional neural network (CNN) in order to measure the confidence of each facial landmark detected by using a well-known landmark detection algorithm. Also, by fitting selected landmarks to 3D space, our framework becomes more robust even when a small number of landmarks are selected. Due to the absence of public dataset for the binocular stereo facial pose, we construct facial pose data sets using a motion sensor for performance validation. In our experiments, our method achieves the higher accuracy of the pose estimation than the previous method, especially for large facial pose cases. Jaeseong Park, Suwoong Heo, Kyungjune Lee, Hyewon Song, Sanghoon Lee 0001 |
ICIP | 5 |
| 2018 | ConcatNet: A Deep Architecture of Concatenation-Assisted Network for Dense Facial Landmark AlignmentabstractFacial landmark is one of the most basic elements for obtaining facial information such as facial expression and emotion. However, detecting dense landmarks on an image is challenging due to various facial poses. In this paper, a deep architecture for dense facial landmark detection, called ConcatNet, is proposed. In our architecture, we propose a CNN-based dense landmark detector on part regions of a face, which extends a given set of sparse landmarks to more accurate and dense landmarks. By introducing interface layers for coordinate normalization and part region localization, we concatenate a network for sparse landmark detection to ConcatNet in a global-to-local manner and the whole network to operate in an end-to-end manner. The experimental results on LFW and 300W datasets show that ConcatNet not only expands the number of the sparse landmarks but also increases the accuracy of the landmark positions remarkably. Also, ConcatNet shows high accuracy in detecting the dense landmarks with a smaller dataset and without additional data on an image such as 3D position annotations when compared to 3D model-based detection method. Hyewon Song, Jiwoo Kang 0001, Sanghoon Lee 0001 |
ICIP | 3 |
| 2018 | A Comparative Quality Assessment Study for Gaming and Non-Gaming VideosabstractRecent years have seen a tremendous increase in video traffic with the rise of Over The Top (OTT) services. Along with traditional Video on demand (VoD) streaming services (e.g., Netflix, YouTube), live video services (e.g., Twitch. tv, YouTubeGaming, Facebook Live) have also resulted in a tremendous share of Internet traffic. Among the live streaming services, gaming video streaming has a major share, with Twitch.tv alone currently responsible for the fourth highest peak Internet traffic in the US. As a consequence of this, and due to the fact that gaming videos are artificial and synthetic, it is worth investigating the specificity of gaming videos in relation to compression and the consequent end user QoE. In this paper, we present an objective and subjective quality comparison study for regular videos and gaming videos, with 30 video sequences (15 per type), encoded using the state of the art encoder HEVC. We discuss the similarity and dissimilarity between the two video types and also discuss how these observations can be used to improve the end user QoE. Nabajeet Barman, Maria G. Martini, Saman Zad Tootaghaj, Sebastian Möller 0001, Sanghoon Lee 0001 |
QoMEX | 5 |
| 2018 | Virtual Reality Sickness Predictor: Analysis of visual-vestibular conflict and VR contentsabstractPredicting the degree of sickness is an imperative goal to guarantee viewing safety when watching virtual reality (VR) contents. Ideally, such predictive models should be explained in terms of the human visual system (HVS). When viewing VR contents using a head mounted display (HMD), there is a conflict between user's actual motion and visually perceived motion. This results in an unnatural visual-vestibular sensory mismatch that causes side effects such as onset of nausea, oculomotor, disorientation, asthenopia (eyestrain). In this paper, we propose a framework called VR sickness predictor (VRSP) using the interaction model between user's motion and the vestibular system. VRSP extracts two types of features: a) perceptual motion feature through a visual-vestibular interaction model, and b) statistical content feature that affects user motion perception. Furthermore, we build a VR sickness database including 36 virtual scenes to evaluate the performance of VRSP. Through rigorous experiments, we demonstrate that the correlation between the proposed model and the subjective sickness score yields ~72 %. Jaekyung Kim, Woojae Kim, Sewoong Ahn, Jinwoo Kim 0005, Sanghoon Lee 0001 |
QoMEX | 5 |
| 2018 | 3D Active Vessel Tracking Using an Elliptical PriorabstractIn this paper, we propose a novel vessel tracking method, called active vessel tracking (AVT). The proposed method retains the major advantages that most 2D segmentation methods have demonstrated for 3D tracking while overcoming the drawbacks of previous 3D vessel tracking methods. Under the assumption that the vessel is cylindrical, thereby making its cross-section elliptical, the AVT finds a plane perpendicular to the vessel axis while tracking the vessel along its length. Also, We propose a method for vessel branch detection to automatically track complete vascular networks from a single starting point, whereas the previously proposed solutions have usually been limited in handling vessel bifurcations precisely on 3D or have required considerable user interaction. Our results show that the method is robust and accurate in both synthetic and clinical cases. In an experiment on synthetic data sets, the proposed method achieved a tracking accuracy of 96.1±0.5, detecting 99.1% of the branches. In an experiment on abdominal CTA data sets, it achieved a tracking accuracy of 98.4±0.5 for six target vessels, detecting 98.3% of the branches. These results show that the proposed method can outperform previous methods for vessel tracking. Jiwoo Kang 0001, Suwoong Heo, Woo Jin Hyung, Joon Seok Lim, Sanghoon Lee 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | Deep Visual Discomfort Predictor for Stereoscopic 3D ImagesabstractMost prior approaches to the problem of stereoscopic 3D (S3D) visual discomfort prediction (VDP) have focused on the extraction of perceptually meaningful handcrafted features based on models of visual perception and of natural depth statistics. Towards advancing performance on this problem, we have developed a deep learning based VDP model named Deep Visual Discomfort Predictor (DeepVDP). DeepVDP uses a convolutional neural network (CNN) to learn features that are highly predictive of experienced visual discomfort. Since a large amount of reference data is needed to train a CNN, we develop a systematic way of dividing S3D image into local regions defined as patches, and model a patch-based CNN using two sequential training steps. Since it is very difficult to obtain human opinions on each patch, instead a proxy ground-truth label that is generated by an existing S3D visual discomfort prediction algorithm called 3D-VDP is assigned to each patch. These proxy ground-truth labels are used to conduct the first stage of training the CNN. In the second stage, the automatically learned local abstractions are aggregated into global features via a feature aggregation layer. The learned features are iteratively updated via supervised learning on subjective 3D discomfort scores, which serve as ground-truth labels on each S3D image. The patchbased CNN model that has been pretrained on proxy groundtruth labels is subsequently retrained on true global subjective scores. The global S3D visual discomfort scores predicted by the trained DeepVDP model achieve state-of-the-art performance as compared to previous VDP algorithms. Heeseok Oh, Sewoong Ahn, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2018 | Cross-antenna interference cancellation and channel estimation for MISO-FBMC/QAM-based eMBMS
Beom Kwon, Sanghoon Lee 0001 |
Wirel. Networks | 2 |
| 2017 | Deep Learning of Human Visual Sensitivity in Image Quality Assessment FrameworkabstractSince human observers are the ultimate receivers of digital images, image quality metrics should be designed from a human-oriented perspective. Conventionally, a number of full-reference image quality assessment (FR-IQA) methods adopted various computational models of the human visual system (HVS) from psychological vision science research. In this paper, we propose a novel convolutional neural networks (CNN) based FR-IQA model, named Deep Image Quality Assessment (DeepQA), where the behavior of the HVS is learned from the underlying data distribution of IQA databases. Different from previous studies, our model seeks the optimal visual weight based on understanding of database information itself without any prior knowledge of the HVS. Through the experiments, we show that the predicted visual sensitivity maps agree with the human subjective opinions. In addition, DeepQA achieves the state-of-the-art prediction accuracy among FR-IQA models. Jongyoo Kim, Sanghoon Lee 0001 |
CVPR | 2 |
| 2017 | Ensemble Deep Learning for Skeleton-Based Action Recognition Using Temporal Sliding LSTM NetworksabstractThis paper addresses the problems of feature representation of skeleton joints and the modeling of temporal dynamics to recognize human actions. Traditional methods generally use relative coordinate systems dependent on some joints, and model only the long-term dependency, while excluding short-term and medium term dependencies. Instead of taking raw skeletons as the input, we transform the skeletons into another coordinate system to obtain the robustness to scale, rotation and translation, and then extract salient motion features from them. Considering that Long Shortterm Memory (LSTM) networks with various time-step sizes can model various attributes well, we propose novel ensemble Temporal Sliding LSTM (TS-LSTM) networks for skeleton-based action recognition. The proposed network is composed of multiple parts containing short-term, mediumterm and long-term TS-LSTM networks, respectively. In our network, we utilize an average ensemble among multiple parts as a final feature to capture various temporal dependencies. We evaluate the proposed networks and the additional other architectures to verify the effectiveness of the proposed networks, and also compare them with several other methods on five challenging datasets. The experimental results demonstrate that our network models achieve the state-of-the-art performance through various temporal features. Additionally, we analyze a relation between the recognized actions and the multi-term TS-LSTM features by visualizing the softmax features of multiple parts. Inwoong Lee, Seoungyoon Kang, Sanghoon Lee 0001 |
ICCV | 4 |
| 2017 | Visual entropy: A new framework for quantifying visual information based on human perceptionabstractIn recent years, how to quantify visualizations of an object and surface displayed in 3D space is now more prominent with a rapid increase in the demand for three-dimensional (3D) content. In order to measure the content information in terms of human visual perception, it is necessary to quantify the visual information in accordance with the human visual system. In this paper, we propose a framework for expressing visual information in bits termed visual entropy based on information theory. The visual entropy of 2D content (2DVE) is composed of texture entropy on the 2D surface and depth entropy based on the monocular cue. In addition to 2DVE, the visual entropy of 3D content (3DVE) includes the depth entropy based on the binocular cue. A series of simulations are conducted to demonstrate the effectiveness of visual entropy, including a performance trade-off between 2D and 3D visualizations measured according to the bitrate. Sewoong Ahn, Kwanghyun Lee, Sanghoon Lee 0001 |
ICIP | 3 |
| 2017 | Deep blind image quality assessment by employing FR-IQAabstractIn this paper, we propose a convolutional neural network (CNN)-based no-reference image quality assessment (NR-IQA). Though deep learning has yielded superior performance in a number of computer vision studies, applying the deep CNN to the NR-IQA framework is not straightforward, since we face a few critical problems: 1) lack of training data; 2) absence of local ground truth targets. To alleviate these problems, we employ the full-reference image quality assessment (FR-IQA) metrics as intermediate training targets of the CNN. In addition, we incorporate the pooling stage in the training stage, so that the whole parameters of the model can be optimized in an end-to-end framework. The proposed model, named as a blind image evaluator based on a convolutional neural network (BIECON), achieves state-of-the-art prediction accuracy that is comparable with that of FR-IQA methods. Jongyoo Kim, Sanghoon Lee 0001 |
ICIP | 2 |
| 2017 | A restoration method for distorted comics to improve comic contents identification
Sanghoon Lee 0005, Sagar Jadhav, Sanghoon Lee 0001 |
Int. J. Document Anal. Recognit. | 4 |
| 2017 | An identification framework for print-scan books in a large database
Sanghoon Lee 0005, Jongyoo Kim, Sanghoon Lee 0001 |
Inf. Sci. | 3 |
| 2017 | Scattered Reference Symbol-Based Channel Estimation and Equalization for FBMC-QAM SystemsabstractIn this paper, we derive the closed forms of residual interference that occurs over a multipath channel of a filter bank multicarrier-quadrature amplitude modulation (FBMC-QAM) system, in which the system adopts two different prototype filters to guarantee that the orthogonality condition is met for both even and odd subcarriers without using a guard interval. The filter coefficients of the two filters are designed differently to meet the orthogonality condition, and thus, their degrees of robustness are different over the multipath channel. In order to clarify the influence of the multipath channel on the performance of the FBMC-QAM, we both derive an effective channel frequency response that includes the filter response and analyze the residual interference over the multipath channel. Based on this analysis, we propose a unique scheme for both even-biased reference symbol mapping and symbol/signal-level channel estimation and equalization that is suitable for the FBMC-QAM system transceiver structure. The simulation results show that the performance of the proposed scheme is comparable to that of a conventional orthogonal frequency division multiplexing-quadrature amplitude modulation system in terms of the mean squared error, bit error rate, and effective data rate. Beom Kwon, Seonghyun Kim, Sanghoon Lee 0001 |
IEEE Trans. Commun. | 3 |
| 2017 | Perceptual Crosstalk Prediction on Autostereoscopic 3D DisplayabstractPerceptual crosstalk prediction for autostereoscopic 3D displays is of fundamental importance in determining the level of quality perceived by humans in terms of the display performance and the 3D viewing experience. However, no robust framework exists to quantify perceptual crosstalk while taking into account the hardware structure of a display as well as its content characteristics via content analysis. In this paper, we present a 3D perceptual crosstalk predictor (3D-PCP) that can be used to predict crosstalk in a unique way when viewing autostereoscopic 3D displays. 3D-PCP captures hardware features using an optical Fourier transform-light measurement device and content features through content analysis based on information theory. By deriving the disparity, luminance, color, and texture maps, this approach defines the visual entropy, mutual information, and relative entropy in order to investigate the influences of the 3D scene characteristics on perceptual crosstalk. The experimental results demonstrate that the 3D-PCP output is highly correlated with subjective scores. Taewan Kim 0002, Jongyoo Kim, Sanghoon Lee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | Blind Sharpness Prediction for Ultrahigh-Definition Video Based on Human Visual ResolutionabstractWe explore a no-reference sharpness assessment model for predicting the perceptual sharpness of ultrahigh-definition (UHD) videos through analysis of visual resolution variation in terms of viewing geometry and scene characteristics. The quality and sharpness of UHD videos are influenced by viewer perception of the spatial resolution afforded by the UHD display, which depends on viewing geometry parameters including display resolution, display size, and viewing distance. In addition, viewers may perceive different degrees of quality and sharpness according to the statistical behavior of the visual signals, such as the motion, texture, and edge, which vary over both spatial and temporal domains. The model also accounts for the resolution variation associated with fixation and foveal regions, which is another important factor affecting the sharpness prediction of UHD video over the spatial domain and which is caused by the nonuniform distribution of the photoreceptors. We calculate the transition of the visually salient statistical characteristics resulting from changing the display's screen size and resolution. Moreover, we calculated the temporal variation in sharpness over consecutive frames in order to evaluate the temporal sharpness perception of UHD video. We verify that the proposed model outperforms other sharpness models in both spatial and temporal sharpness assessments. Haksub Kim, Jongyoo Kim, Taegeun Oh, Sanghoon Lee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Quality Assessment of Perceptual Crosstalk on Two-View Auto-Stereoscopic DisplaysabstractCrosstalk is one of the most severe factors affecting the perceived quality of stereoscopic 3D images. It arises from a leakage of light intensity between multiple views, as in auto-stereoscopic displays. Well-known determinants of crosstalk include the co-location contrast and disparity of the left and right images, which have been dealt with in prior studies. However, when a natural stereo image that contains complex naturalistic spatial characteristics is viewed on an auto-stereoscopic display, other factors may also play an important role in the perception of crosstalk. Here, we describe a new way of predicting the perceived severity of crosstalk, which we call the Binocular Perceptual Crosstalk Predictor (BPCP). BPCP uses measurements of three complementary 3D image properties (texture, structural duplication, and binocular summation) in combination with two well-known factors (co-location contrast and disparity) to make predictions of crosstalk on two-view auto-stereoscopic displays. The new BPCP model includes two masking algorithms and a binocular pooling method. We explore a new masking phenomenon that we call duplicated structure masking, which arises from structural correlations between the original and distorted objects. We also utilize an advanced binocular summation model to develop a binocular pooling algorithm. Our experimental results indicate that BPCP achieves high correlations against subjective test results, improving upon those delivered by previous crosstalk prediction models. Jongyoo Kim, Taewan Kim 0002, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2017 | Blind Deep S3D Image Quality Evaluation via Local to Global Feature AggregationabstractPreviously, no-reference (NR) stereoscopic 3D (S3D) image quality assessment (IQA) algorithms have been limited to the extraction of reliable hand-crafted features based on an understanding of the insufficiently revealed human visual system or natural scene statistics. Furthermore, compared with full-reference (FR) S3D IQA metrics, it is difficult to achieve competitive quality score predictions using the extracted features, which are not optimized with respect to human opinion. To cope with this limitation of the conventional approach, we introduce a novel deep learning scheme for NR S3D IQA in terms of local to global feature aggregation. A deep convolutional neural network (CNN) model is trained in a supervised manner through two-step regression. First, to overcome the lack of training data, local patch-based CNNs are modeled, and the FR S3D IQA metric is used to approximate a reference ground-truth for training the CNNs. The automatically extracted local abstractions are aggregated into global features by inserting an aggregation layer in the deep structure. The locally trained model parameters are then updated iteratively using supervised global labeling, i.e., subjective mean opinion score (MOS). In particular, the proposed deep NR S3D image quality evaluator does not estimate the depth from a pair of S3D images. The S3D image quality scores predicted by the proposed method represent a significant improvement over those of previous NR S3D IQA algorithms. Indeed, the accuracy of the proposed method is competitive with FR S3D IQA metrics, having ~ 91% correlation in terms of MOS. Heeseok Oh, Sewoong Ahn, Jongyoo Kim, Sanghoon Lee 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Enhancement of Visual Comfort and Sense of Presence on Stereoscopic 3D ImagesabstractConventional stereoscopic 3D (S3D) displays do not provide accommodation depth cues of the 3D image or video contents being viewed. The sense of content depths is thus limited to cues supplied by motion parallax (for 3D video), stereoscopic vergence cues created by presenting left and right views to the respective eyes, and other contextual and perspective depth cues. The absence of accommodation cues can induce two kinds of accommodation vergence mismatches (AVM) at the fixation and peripheral points, which can result in severe visual discomfort. With the aim of alleviating discomfort arising from AVM, we propose a new visual comfort enhancement approach for processing S3D visual signals to deliver a more comfortable 3D viewing experience at the display. This is accomplished via an optimization process whereby a predictive indicator of visual discomfort is minimized, while still aiming to maintain the viewer's sense of 3D presence by performing a suitable parallax shift, and by directed blurring of the signal. Our processing framework is defined on 3D visual coordinates that reflect the nonuniform resolution of retinal sensors and that uses a measure of 3D saliency strength. An appropriate level of blur that corresponds to the degree of parallax shift is found, making it possible to produce synthetic accommodation cues implemented using a perceptively relevant filter. By this method, AVM, the primary contributor to the discomfort felt when viewing S3D images, is reduced. We show via a series of subjective experiments that the proposed approach improves visual comfort while preserving the sense of 3D presence. Heeseok Oh, Jongyoo Kim, Jinwoo Kim 0005, Taewan Kim 0002, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2016 | No-reference perceptual sharpness assessment for ultra-high-definition imagesabstractSince ultra-high-definition (UHD) display has larger resolution and various display size, it is necessary to measure image sharpness considering variation in visual resolution caused by diverse viewing geometry. In this paper, we propose a no-reference perceptual sharpness assessment model of UHD images. The proposed model analyzes viewing geometry in terms of display resolution and viewing environment. Then, we measure the local adaptive sharpness score in accordance with the textural motion blur, texture, and edge. In addition, we propose a spatial pooling method associated with foveal regions, which is caused by nonuniform distribution of the photoreceptors on a human retina. Through the rigorous experiments, we demonstrate that the proposed model can measure the sharpness of UHD images more accurately than other image sharpness assessment methods. Woojae Kim, Haksub Kim, Heeseok Oh, Jongyoo Kim, Sanghoon Lee 0001 |
ICIP | 5 |
| 2016 | Visual attention analysis on stereoscopic images for subjective discomfort evaluationabstractBy analyzing the statistical behaviors on human visual attention, we discover a clue that the fixation behaviors are highly correlated with how much the viewers feel visual discomfort on stereoscopic images differently from conventional subjective assessments. In order to quantify the correlation between visual attention and discomfort, we explore a novel methodology termed transition of visual attention (ToVA) according to various disparities, which accounts depth attributes of 3D images by eye-tracker experiments. Moreover, the saliency entropy is defined to quantify the distribution of fixations for 3D images. Then, we measure ToVA in terms of the relative saliency entropy using Kullback-Leibler divergence. In order to evaluate the effectiveness of ToVA, a successful example application is also provided, whereby ToVA is applied to obtaining subjective results of measuring discomfort experienced when viewing 3D displays rather than relying on the conventional subjective test by using scoring system. Sewoong Ahn, Haksub Kim, Sanghoon Lee 0001 |
ICME | 4 |
| 2016 | A New Framework for Measuring 2D and 3D Visual Information in Terms of EntropyabstractIn recent years, the problem of how to more perceptually quantify visualizations of an object and a surface displayed in 3D space between the human eye and a display has been more prominent with rapid increase in the demand for 3D/ultrahigh-definition content. In order to quantify the content information in terms of human visual perception, it is necessary to measure the visual information of 2D and 3D videos accurately in accordance with the human visual system. In this paper, we investigate a new framework for expressing visual information in bits termed visual entropy, based on information theory. 2D visual entropy (2DVE) is composed of two major components: 1) texture entropy on the 2D surface and 2) depth entropy based on the monocular cue. In contrast, 3D visual entropy (3DVE) includes the depth entropy based on the binocular cue in addition to the 2DVE. A series of simulations is conducted to demonstrate the effectiveness of visual entropy, including the degree of accuracy needed to represent perceptual information by means of the texture and monocular and binocular depth entropies. In the simulation results, two successful example applications are also provided, whereby visual entropy is applied to the problems of predicting the visual discomfort experienced when viewing 3D displays and of analyzing performance tradeoffs between 2D and 3D contents. Kwanghyun Lee, Sanghoon Lee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Visual Presence: Viewing Geometry Visual Information of UHD S3D EntertainmentabstractTo maximize the presence experienced by humans, visual content has evolved to achieve a higher visual presence in a series of high definition (HD), ultra HD (UHD), 8K UHD, and 8K stereoscopic 3D (S3D). Several studies have introduced visual presence delivered from content when viewing UHD S3D from a content analysis perspective. Nevertheless, no clear definition has been presented for visual presence, and only a subjective evaluation has been relied upon. The main reason for this is that there is a limitation to defining visual presence via the use of content information itself. In this paper, we define the visual presence for each viewing environment, and investigate a novel methodology to measure the experienced visual presence when viewing both 2D and 3D via the definition of a new metric termed volume of visual information by quantifying the influence of the viewing geometry between the display and viewer. To achieve this goal, the viewing geometry and display parameters for both flat and atypical displays are analyzed in terms of human perception by introducing a novel concept of pixel-wise geometry. In addition, perceptual weighting through analysis of content information is performed in accordance with monocular and binocular vision characteristics. In the experimental results, it is shown that the constructed model based on the viewing geometry, content, and perceptual characteristics has a high correlation of about 84% with subjective evaluations. Heeseok Oh, Sanghoon Lee 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Stereoscopic 3D Visual Discomfort Prediction: A Dynamic Accommodation and Vergence Interaction ModelabstractThe human visual system perceives 3D depth following sensing via its binocular optical system, a series of massively parallel processing units, and a feedback system that controls the mechanical dynamics of eye movements and the crystalline lens. The process of accommodation (focusing of the crystalline lens) and binocular vergence is controlled simultaneously and symbiotically via cross-coupled communication between the two critical depth computation modalities. The output responses of these two subsystems, which are induced by oculomotor control, are used in the computation of a clear and stable cyclopean 3D image from the input stimuli. These subsystems operate in smooth synchronicity when one is viewing the natural world; however, conflicting responses can occur when viewing stereoscopic 3D (S3D) content on fixed displays, causing physiological discomfort. If such occurrences could be predicted, then they might also be avoided (by modifying the acquisition process) or ameliorated (by changing the relative scene depth). Toward this end, we have developed a dynamic accommodation and vergence interaction (DAVI) model that successfully predicts visual discomfort on S3D images. The DAVI model is based on the phasic and reflex responses of the fast fusional vergence mechanism. Quantitative models of accommodation and vergence mismatches are used to conduct visual discomfort prediction. Other 3D perceptual elements are included in the proposed method, including sharpness limits imposed by the depth of focus and fusion limits implied by Panum's fusional area. The DAVI predictor is created by training a support vector machine on features derived from the proposed model and on recorded subjective assessment results. The experimental results are shown to produce accurate predictions of experienced visual discomfort. Heeseok Oh, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2015 | 3D perception based quality pooling on stereoscopic imageabstractIn this paper, we investigate a unique approach for characterizing the process of human 3D perception by reflecting psychological observation to quality assessment. When a human being perceives the 3D structure, the brain classifies the scene into the binocular- or monocular-vision region depending on the availability of binocular depth perception in the unit of a certain region. Furthermore, we include the 3D perception on the quality assessment, and propose a human 3D perception based stereo image quality pooling model. Haksub Kim, Sanghoon Lee 0001 |
ICIP | 3 |
| 2015 | 3D visual discomfort predictor based on neural activity statisticsabstractVisual discomfort assessment (VDA) on stereoscopic images is of fundamental importance for making decisions regarding visual fatigue caused by unnatural binocular alignment. Nevertheless, no solid framework exists to quantify this discomfort using models of the responses of visual neurons. Binocular vision is realized by means of neural mechanisms that subserve the sensorimotor control of eye movements. We propose a neuronal model-based framework called Neural 3D Visual Discomfort Predictor (N3D-VDP) that automatically predicts the level of visual discomfort experienced when viewing stereoscopic 3D (S3D) images. The N3D-VDP model extracts features derived by estimating the neural activity associated with the processing of binocular disparities. In this regard we deploy a model of disparity processing in the extra-striate middle temporal (MT) region of occipital lobe. We compare the performance of N3D-VDP with other recent VDA algorithms using correlations against reported subjective visual discomfort, and show that N3D-VDP is statistically superior to the other methods. Heeseok Oh, Jongyoo Kim, Sanghoon Lee 0001, Alan C. Bovik |
ICIP | 3 |
| 2015 | Video sharpness prediction based on motion blur analysisabstractFor high bit rate video, it is important to acquire the video contents with high resolution, the quality of which may be degraded due to the motion blur from the movement of an object(s) or the camera. However, conventional sharpness assessments are designed to find focal blur caused either by defocusing or by compression distortion targeted for low bit rates. To overcome this limitation, we present a no-reference framework of a visual sharpness assessment (VSA) for high-resolution video based on the motion and scene classification. In the proposed framework, the accuracy of the sharpness estimation can be improved via pooling weighted by the visual perception from the object and camera movements and by the strong influence from the region with the highest sharpness. Based on the motion blur characteristics, the variance and the contrast over the spectral domain are used to quantify the perceived sharpness. Moreover, for the VSA, we extract the highly influential sharper regions and emphasize them by utilizing the scene adaptive pooling. Jongyoo Kim, Woojae Kim, Jisoo Lee, Sanghoon Lee 0001 |
ICME | 5 |
| 2015 | Low-complexity and robust comic fingerprint method for comic identification
Taegeun Oh, Nakyeon Choi, Sanghoon Lee 0001 |
Signal Process. Image Commun. | 4 |
| 2015 | Transfer Function Model of Physiological Mechanisms Underlying Temporal Visual Discomfort Experienced When Viewing Stereoscopic 3D ImagesabstractWhen viewing 3D images, a sense of visual comfort (or lack of) is developed in the brain over time as a function of binocular disparity and other 3D factors. We have developed a unique temporal visual discomfort model (TVDM) that we use to automatically predict the degree of discomfort felt when viewing stereoscopic 3D (S3D) images. This model is based on physiological mechanisms. In particular, TVDM is defined as a second-order system capturing relevant neuronal elements of the visual pathway from the eyes and through the brain. The experimental results demonstrate that the TVDM transfer function model produces predictions that correlate highly with the subjective visual discomfort scores contained in the large public databases. The transfer function analysis also yields insights into the perceptual processes that yield a stable S3D image. Taewan Kim 0002, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2015 | 3D Visual Discomfort Predictor: Analysis of Disparity and Neural Activity StatisticsabstractBeing able to predict the degree of visual discomfort that is felt when viewing stereoscopic 3D (S3D) images is an important goal toward ameliorating causative factors, such as excessive horizontal disparity, misalignments or mismatches between the left and right views of stereo pairs, or conflicts between different depth cues. Ideally, such a model should account for such factors as capture and viewing geometries, the distribution of disparities, and the responses of visual neurons. When viewing modern 3D displays, visual discomfort is caused primarily by changes in binocular vergence while accommodation in held fixed at the viewing distance to a flat 3D screen. This results in unnatural mismatches between ocular fixations and ocular focus that does not occur in normal direct 3D viewing. This accommodation vergence conflict can cause adverse effects, such as headaches, fatigue, eye strain, and reduced visual ability. Binocular vision is ultimately realized by means of neural mechanisms that subserve the sensorimotor control of eye movements. Realizing that the neuronal responses are directly implicated in both the control and experience of 3D perception, we have developed a model-based neuronal and statistical framework called the 3D visual discomfort predictor (3D-VDP)that automatically predicts the level of visual discomfort that is experienced when viewing S3D images. 3D-VDP extracts two types of features: 1) coarse features derived from the statistics of binocular disparities and 2) fine features derived by estimating the neural activity associated with the processing of horizontal disparities. In particular, we deploy a model of horizontal disparity processing in the extrastriate middle temporal region of occipital lobe. We compare the performance of 3D-VDP with other recent discomfort prediction algorithms with respect to correlation against recorded subjective visual discomfort scores,and show that 3D-VDP is statistically superior to the other methods. Jincheol Park, Heeseok Oh, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2015 | Link Capacity-Energy Aware WDC for Network Lifetime MaximizationabstractWith the increase in flexibility and capabilities of wireless networks, the use of distributed computing over wireless network environments is being researched in order to maximize network sustainability and interoperability among distributed nodes. To this end, a new paradigm is required for optimization of a more generalized environment. This environment would include various nodes of different processing and communication abilities constrained by circuit powers and residual energy over individual dynamic wireless channels. In this paper, we present a novel strategy named link capacity-energy aware wireless distributed computing (LEA-WDC) for maximizing the lifetime of a wireless network. The major advantage of LEA-WDC is its achievement of lifetime maximization by systematically reconciling highly coupled system parameters (tasks, processing power, communication power, and residual energy) in terms of the role of nodes and the layer of each node. To attain an optimal solution, we perform unique interworking optimization via decomposition in accordance with the roles of header and slave nodes. The evaluation results of our simulation verify that the lifetime is further maximized by finding the optimal transmission power of each node according to the Shannon capacity. Seonghyun Kim, Sanghoon Lee 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2015 | Transition of Visual Attention Assessment in Stereoscopic Images With Evaluation of Subjective Visual Quality and DiscomfortabstractThrough statistical analysis of the behaviors of human visual attention, we discovered that fixation behaviors are highly correlated with the degree to which the viewer experiences visual quality/discomfort with stereoscopic images. To quantify the correlation between visual attention and quality/discomfort with respect to various distortions and disparities, we explore a novel methodology called transition of visual attention (ToVA) that accounts for diverse low-level luminance, chrominance, and depth attributes of 3D images in eye-tracker experiments. Moreover, saliency entropy is defined in order to quantify the distribution of fixations for 3D images. Using saliency entropy, we interpret perceptual factors in accordance with the non-uniform resolution of the human eye and the stereoscopic limits caused by foveation and Panum's fusional area. We then measure ToVA in terms of relative saliency entropy using Kullback-Leibler divergence. To evaluate the effectiveness of ToVA, a successful example application is also provided for which ToVA is applied to obtain subjective measurements of the quality and discomfort experienced when viewing 3D displays, rather than relying on conventional subjective tests using a scoring system. Haksub Kim, Sanghoon Lee 0001 |
IEEE Trans. Multim. | 2 |
| 2015 | A downlink power control algorithm for long-term energy efficiency of small cell network
Beom Kwon, Seonghyun Kim, Hojae Lee, Sanghoon Lee 0001 |
Wirel. Networks | 4 |
| 2014 | Quality assessment of perceptual crosstalk in autostereoscopic displayabstractCrosstalk is one of the most annoying problems in an autostereoscopic display causing perceptual quality degradation and visual discomfort. To predict the perceived crosstalk when viewing an autostereoscopic display, it is necessary to consider the characteristics of human perception, displaying mechanism, viewing environment and so on. Therefor, we propose a novel metric for predicting the perceptual crosstalk that is based on human visual system (HVS); non-linear sensitivity of luminance and masking effects. The proposed model adopts the duplicated structure masking, yielding predictive power that is statistically superior to prior models that rely on 2D quality metric. Jongyoo Kim, Taewan Kim 0002, Sanghoon Lee 0001 |
ICIP | 3 |
| 2014 | Optimal phase control for joint transmission and reception with beamformingabstractIn this paper, we propose a joint transmission and reception with phase control for beamforming in a multi-cell environment. For generated transmit weight vectors of multiple base stations (BSs), a mobile station (MS) calculates phases to maximize an achievable rate with low rate feedback. By using the phases, the multiple transmit weight vectors are coordinated to improve the signal to noise ratio (SNR). In order to find optimal phases, we present a phase control method for the effective channel via geometrical approach. Seonghyun Kim, Hojae Lee, Beom Kwon, Inwoong Lee, Sanghoon Lee 0001 |
IPCCC | 5 |
| 2014 | Combinatorial JPT based on orthogonal beamforming for two-cell cooperationabstractIn this paper, we investigate efficient multi-cell cooperation based on CoMP-joint processing and transmission (CoMP-JPT) with orthogonal beamforming. Through the use of a combinatorial optimization algorithm, the optimal user scheduling for joint transmission using multiple transmitters is accomplished. The throughput of the CoMP-JPT can be significantly improved while maintaining fairness among users over a multi-cell environment. Hojae Lee, Beom Kwon, Seonghyun Kim, Inwoong Lee, Sanghoon Lee 0001 |
IPCCC | 5 |
| 2014 | Combinatorial Orthogonal Beamforming for Joint Processing and TransmissionabstractCoordinated multiple point transmission (CoMP) technologies have recently been proposed to improve the performance of cell-edge users who suffer strong inter-cell interference. Nevertheless, as more transmitters get involved in cooperation, the complexity associated with the selection of multi-dimensional parameters increases exponentially. In this work, we investigate efficient multi-cell cooperation based on CoMP-joint processing and transmission (CoMP-JPT) with orthogonal beamforming using limited feedback. Through the utilization of combinatorial optimization, optimal user scheduling for joint transmission via multiple transmitters is accomplished, while the computational complexity is significantly reduced. In particular, a generalized beam assignment problem (GBAP) is formulated and solved using a combinatorial algorithm that is generalized in terms of the number of transmitters o\mathcal{B}o. The performance of the combinatorial orthogonal beamforming (COBF) scheme is mathematically analyzed so as to demonstrate its superiority and capability to maintain fairness among users in a multi-cell environment. In the simulation results, a performance gain of more than 50% for cell-edge users is obtained without a loss in the average throughput for the total number of users. In addition, the COBF method can reduce the complexity by more than 60% when compared to the conventional exhaustive search technique. Hojae Lee, Seonghyun Kim, Sanghoon Lee 0001 |
IEEE Trans. Commun. | 3 |
| 2014 | Saliency Prediction on Stereoscopic VideosabstractWe describe a new 3D saliency prediction model that accounts for diverse low-level luminance, chrominance, motion, and depth attributes of 3D videos as well as high-level classifications of scenes by type. The model also accounts for perceptual factors, such as the nonuniform resolution of the human eye, stereoscopic limits imposed by Panum's fusional area, and the predicted degree of (dis) comfort felt, when viewing the 3D video. The high-level analysis involves classification of each 3D video scene by type with regard to estimated camera motion and the motions of objects in the videos. Decisions regarding the relative saliency of objects or regions are supported by data obtained through a series of eye-tracking experiments. The algorithm developed from the model elements operates by finding and segmenting salient 3D space-time regions in a video, then calculating the saliency strength of each segment using measured attributes of motion, disparity, texture, and the predicted degree of visual discomfort experienced. The saliency energy of both segmented objects and frames are weighted using models of human foveation and Panum's fusional area yielding a single predictor of 3D saliency. Haksub Kim, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2014 | 3D Visual Activity Assessment Based on Natural Scene StatisticsabstractOne of the most challenging ongoing issues in the field of 3D visual research is how to perceptually quantify object and surface visualizations that are displayed within a virtual 3D space between a human eye and 3D display. To seek an effective method of quantification, it is necessary to measure various elements related to the perception of 3D objects at different depths. We propose a new framework for quantifying 3D visual information that we call 3D visual activity (3DVA), which utilizes natural scene statistics measured over 3D visual coordinates. We account for important aspects of 3D perception by carrying out a 3D coordinate transform reflecting the nonuniform sampling resolution of the eye and the process of stereoscopic fusion. The 3DVA utilizes the empirical distortions of wavelet coefficients to a parametric generalized Gaussian probability distribution model and a set of 3D perceptual weights. We conducted a series of simulations that demonstrate the effectiveness of the 3DVA for quantifying the statistical dynamics of visual 3D space with respect to disparity, motion, texture, and color. A successful example application is also provided, whereby 3DVA is applied to the problem of predicting visual fatigue experienced when viewing 3D displays. Kwanghyun Lee, Anush K. Moorthy, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2014 | No-Reference Sharpness Assessment of Camera-Shaken Images by Analysis of Spectral StructureabstractThe tremendous explosion of image-, video-, and audio-enabled mobile devices, such as tablets and smart-phones in recent years, has led to an associated dramatic increase in the volume of captured and distributed multimedia content. In particular, the number of digital photographs being captured annually is approaching 100 billion in just the U.S. These pictures are increasingly being acquired by inexperienced, casual users under highly diverse conditions leading to a plethora of distortions, including blur induced by camera shake. In order to be able to automatically detect, correct, or cull images impaired by shake-induced blur, it is necessary to develop distortion models specific to and suitable for assessing the sharpness of camera-shaken images. Toward this goal, we have developed a no-reference framework for automatically predicting the perceptual quality of camera-shaken images based on their spectral statistics. Two kinds of features are defined that capture blur induced by camera shake. One is a directional feature, which measures the variation of the image spectrum across orientations. The second feature captures the shape, area, and orientation of the spectral contours of camera shaken images. We demonstrate the performance of an algorithm derived from these features on new and existing databases of images distorted by camera shake. Taegeun Oh, Jincheol Park, Kalpana Seshadrinathan, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2014 | Multimodal Interactive Continuous Scoring of Subjective 3D Video Quality of ExperienceabstractPeople experience a variety of 3D visual programs, such as 3D cinema, 3D TV and 3D games, making it necessary to deploy reliable methodologies for predicting each viewer's subjective experience. We propose a new methodology that we call multimodal interactive continuous scoring of quality (MICSQ). MICSQ is composed of a device interaction process between the 3D display and a separate device (PC, tablet, etc.) used as an assessment tool, and a human interaction process between the subject(s) and the separate device. The scoring process is multimodal, using aural and tactile cues to help engage and focus the subject(s) on their tasks by enhancing neuroplasticity. Recorded human responses to 3D visualizations obtained via MICSQ correlate highly with measurements of spatial and temporal activity in the 3D video content. We have also found that 3D quality of experience (QoE) assessment results obtained using MICSQ are more reliable over a wide dynamic range of content than obtained by the conventional single stimulus continuous quality evaluation (SSCQE) protocol. Moreover, the wireless device interaction process makes it possible for multiple subjects to assess 3D QoE simultaneously in a large space such as a movie theater, at different viewing angles and distances. We conducted a series of interesting 3D experiments showing the accuracy and versatility of the new system, while yielding new findings on visual comfort in terms of disparity, motion and an interesting relation between the naturalness and depth of field (DOF) of a stereo camera. Taewan Kim 0002, Jiwoo Kang 0001, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Multim. | 3 |
| 2013 | Visually Weighted Compressive Sensing: Measurement and ReconstructionabstractCompressive sensing (CS) makes it possible to more naturally create compact representations of data with respect to a desired data rate. Through wavelet decomposition, smooth and piecewise smooth signals can be represented as sparse and compressible coefficients. These coefficients can then be effectively compressed via the CS. Since a wavelet transform divides image information into layered blockwise wavelet coefficients over spatial and frequency domains, visual improvement can be attained by an appropriate perceptually weighted CS scheme. We introduce such a method in this paper and compare it with the conventional CS. The resulting visual CS model is shown to deliver improved visual reconstructions. Hyungkeuk Lee, Heeseok Oh, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2013 | Video Quality Pooling Adaptive to Perceptual Distortion SeverityabstractIt is generally recognized that severe video distortions that are transient in space and/or time have a large effect on overall perceived video quality. In order to understand this phenomena, we study the distribution of spatio-temporally local quality scores obtained from several video quality assessment (VQA) algorithms on videos suffering from compression and lossy transmission over communication channels. We propose a content adaptive spatial and temporal pooling strategy based on the observed distribution. Our method adaptively emphasizes "worst" scores along both the spatial and temporal dimensions of a video sequence and also considers the perceptual effect of large-area cohesive motion flow such as egomotion. We demonstrate the efficacy of the method by testing it using three different VQA algorithms on the LIVE Video Quality database and the EPFL-PoliMI video quality database. Jincheol Park, Kalpana Seshadrinathan, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2013 | Effective channel control random beamforming for single user MIMO systems
Hojae Lee, Jongrok Park, Hyukmin Son, Sanghoon Lee 0001 |
Wirel. Networks | 4 |
| 2013 | Hierarchical modulation based cooperative relaying over a multi-cell OFDMA network
Hyukmin Son, Sanghoon Lee 0001 |
Wirel. Networks | 2 |
| 2012 | M2-m2 Beamforming for Virtual MIMO Broadcasting in Multi-Hop Relay NetworksabstractThe emerging demand for multi-hop relay networks requires significant improvement of the end-to-end throughput through the development of a more advanced transmission technology. Research on multiple-input multiple-output (MIMO) has been actively pursued for achieving channel throughput improvements for cellular networks. In a multi-hop ad-hoc network, it is difficult for each sensor node to have multiple transmit antennas, therefore, virtual MIMO has been introduced as an extended version of cellular MIMO. The conventional virtual MIMO scheme requires additional complexity in order to implement the same functionality used for cellular systems. Therefore, we propose a new framework for the broadcast virtual MIMO system (BVMS) by developing an innovative max-min/min-max (M2-m2) beamforming technology optimized for the multi-hop relay network. Compared to the conventional singular value decomposition (SVD)-based or random beamforming technologies, M2-m2beamforming significantly improves the end-to-end channel throughput up to the optimal bound for the BVMS over a multi-hop relay network. Jongrok Park, Sanghoon Lee 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2012 | MIMO Broadcast Channels Based on SINR Feedback using a Non-Orthogonal Beamforming MatrixabstractThe signal to interference plus noise ratio (SINR) feedback has been utilized in random beamforming (RBF) to select users for multiple-input multiple-output (MIMO) systems where a large number of users are required to obtain a multi-user diversity gain. However, if the number of users is not large enough, it may be difficult to obtain the performance gain and be easy to have a performance degradation conversely. To resolve this problem, it is necessary to find orthogonal random beams close to the channel matrix of selected users. In this paper, we present RBF using a non-orthogonal beamforming matrix (RBF NOBM), demonstrating an efficient search of random beams correlated with the channel matrix of selected users over a small number of users. For single user RBF using NOBM (SU-RBF NOBM) with two transmit antennas, the beamforming matrix for user data service is obtained through closed form expressions, and SU-RBF NOBM is then expanded to more than two transmit antennas. We also propose user selection algorithms for SU-RBF NOBM and multi-user RBF using NOBM (MU-RBF NOBM). Via simulation results, we demonstrate that the performances of RBF NOBM represent significant improvements compared to conventional beamforming schemes. Jongrok Park, Hojae Lee, Sanghoon Lee 0001 |
IEEE Trans. Commun. | 3 |
| 2012 | A Multi-User MIMO Downlink Receiver and Quantizer Design Based on SINR OptimizationabstractIn a quantization based multiuser multiple-input/multiple-output (MU-MIMO) broadcast channel, the effective channel gain (i.e., the norm of a channel vector obtained from the receive combiner) needs to be considered in addition to a reduction of quantization errors for maximizing the signal-to-interference plus noise ratio (SINR). In this work, we prove that the effective channel gains in the space limited by an N_R x M_T MIMO channel form a min(M_T,N_R)-dimensional ellipsoid in the channel space. Utilizing the geometric proof, the achievable effective channel gain is derived as a function of its direction for a given MIMO channel. Based on the derivation, we finally propose a quantization vector selection criterion and an optimal receive combiner to maximize the SINR. The proposed maximum SINR based combining (MSC) is proven to be a better solution for maximizing the SINR compared to quantization based combining (QBC) and maximum gain based combining (MGC), each of which attempts to minimize quantization errors and maximize the effective channel gain. Hyukmin Son, Seonghyun Kim, Sanghoon Lee 0001 |
IEEE Trans. Commun. | 3 |
| 2012 | Receiver Design for MIMO Relay Stations in Multi-Cell Downlink SystemabstractIn multi-cell downlink system, quality of service (QoS) of multiuser MIMO is an important issue, in particular, for cell-edge users. For QoS of cell-edge users in the system, the utilization of multiple-input multiple-output (MIMO) relay stations (RSs) is a promising solution that improves the signal-to-interference plus noise ratio (SINR) from neighboring base stations (BSs) to each RS. Since the capacity between the BS and the RS relies heavily on the receive performance, it is necessary to reflect interference channels in the design of the receive weight vector. In this paper, we use a geometric approach to derive the effective channel gain of the RS according to the receive weight vector. The geometric relationship between the desired and interference channels is also described. Using the description, we propose a receiver, called the maximum lower bound of expected SINR receiver (MLESR). Through a performance analysis of the MLESR over the interference-limited regime, we derive closed terms for the performance bounds in terms of the number of RS antennas, the interference channel rank, and the BS transmit power according to the rate gap of the MLESR with respect to an ideal case, i.e., a no interference case. Seonghyun Kim, Hyukmin Son, Sanghoon Lee 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2011 | A cross-layer optimization for energy-efficient MAC protocol with delay and rate constraintsabstractWe propose the energy efficient MAC algorithm in this pa per. In the proposed algorithm, each node sets the contention window size with respect to the residual energy, the harvesting energy and the transmit power. This algorithm makes the sensor nodes consume their energy efficiently. To achieve this goal, we use the game theory and the cross-layer optimization. Introducing the non-cooperative game, we can formulate the utility function easily. In this paper, we can allocate the optimal power by the cross-layer optimization on the PHY and the MAC. Haksub Kim, Hyungkeuk Lee, Sanghoon Lee 0001 |
ICASSP | 3 |
| 2011 | Optimal image transmission over Visual Sensor NetworksabstractIn this paper, we propose a methodology for optimal image transmission over a VSNs (Visual Sensor Networks) via cross-layer optimization. Toward this goal, we control the compression ratio of a captured image and network parameters such as source rate, flow rate and routing path. In particular, since this scheme is based on distributed optimization, we can avoid energy concentration in a specific node such as CH (Cluster Head) which increases the network lifetime. In the simulation, we demonstrate the network adaptation procedure over a randomly deployed VSNs and evaluate the quality of the transmitted image using the SSIM (Structural Similarity) index. Sanghoon Lee 0001, Alan C. Bovik |
ICIP | 2 |
| 2011 | A new block compressive sensingto control the number of measurementsabstractCompressive Sensing (CS) aims to recover a sparse signal from a small number of projections onto random vectors. Because of its great practical possibility, both academia and industries have made efforts to develop the CS's reconstruction performance, but most of existing works remain at the theoretical study. In this paper, we propose a new Block Compres-sive Sensing (nBCS), which has several benefits compared to the general CS methods. In particular, the nBCS can be dynamically adaptive to varying channel capacity because it conveys the good inheritance of the wavelet transform. Hyungkeuk Lee, Heeseok Oh, Sanghoon Lee 0001 |
ICIP | 3 |
| 2011 | Spatio-temporal quality pooling accounting for transient severe impairments and egomotionabstractWith the increasing popularity of video applications, the reliable measurement of perceived video quality has increased in importance. We study methods for pooling video quality scores over space and time. The method accounts for localized severe impairments of the signal which exhibit significant influence on the subjective impression of the overall signal quality. It also accounts for the effect of camera motion (egomotion) on perceived quality. The method arrived at is tested on the LIVE Video Quality Database and is shown to perform quite well. Jincheol Park, Kalpana Seshadrinathan, Sanghoon Lee 0001, Alan C. Bovik |
ICIP | 3 |
| 2011 | Visual Capacity Analysis of Wireless NetworksabstractThis paper models and analyzes an asymptotic visual capacity based on distortions by encoding and wireless transmission. In a limited 2-dimensional area, there are N source and destination node pairs. They are transmitting, receiving and relaying others' traffic. The analysis covers encoding distortion, wireless transmission error and end-to-end delay distribution. Our simulation results show that the visual capacity of the network decreases over the rate of 1/(n log n). Sanghoon Lee 0001, Hyukmin Son, Jongrok Park, Sanghoon Lee 0005 |
VTC Spring | 1 |
| 2011 | Beamforming Matrix Transformation for Random BeamformingabstractThe signal to interference plus noise ratio (SINR) feedback has been utilized in random beamforming (RBF) to select users for the provision of service in multiple-input multiple-output (MIMO) systems. A large number of users are required to obtain the gain of multi-user diversity for a downlink transmission. However, if the number is not large enough, it may be difficult to obtain multi-user diversity, leading to a rapid degradation in performance. To resolve this problem, it is necessary to find beamforming matrix close to the channels of the selected users. In this paper, a beamforming matrix transformation for RBF (BMT-RBF) is presented, demonstrating an efficient search of random beams correlated with user channels as opposed to conventional ones, over a small number of users. Jongrok Park, Hojae Lee, Sanghoon Lee 0001, Sanghoon Lee 0005 |
VTC Spring | 3 |
| 2011 | An Efficient FRS-Cooperative Strategy with No CSIT in OFDM-Based NetworksabstractIn conventional cellular systems with a frequency reuse factor (FRF) of one, quality of service (QoS) may not be guaranteed for users located at the cell boundary due to inter cell interference (ICI). Even if relay nodes are adopted for cooperative communications, ICI and inter-relay interference (IRI) remain significant obstacles to reliable data transmission at the border region of each cell. This paper presents an efficient fixed relay station (FRS)-cooperative strategy termed interference cancelation based FRS-cooperation (ICFC) to improve the QoS. The proposed scheme enables to improve the performances of two users at cell-edge in each cell through cooperation and interference cancelation, simultaneously. In particular, the ICFC scheme is well applicable to practical systems because no CSIT is required. Hyukmin Son, Sanghoon Lee 0005, Hojae Lee, Sanghoon Lee 0001 |
VTC Spring | 4 |
| 2011 | Iterative Best Beam Selection for Random Unitary BeamformingabstractDue to multi-user diversity, the performance of a random unitary beamforming (RUB) can be improved upon in proportion to the number of users at the expense of an increase in feedback. It is essential to identify a mechanism to determine a set of orthonormal beams for improving the channel capacity. This paper proposes a unique scheme, named iterative best beam selection-based RUB (I-RUB), in which the base station selects the best beamforming vectors from unitary matrices generated by the proposed iterative algorithm. Using the selected beamforming vector at each stage, the beam set can ultimately be constructed for high sum-rate capacity of downlink, which leads to a reduction in feedback compared to those of other RUB schemes based on the beam set selection. Hyukmin Son, Sanghoon Lee 0001 |
IEEE Trans. Commun. | 2 |
| 2011 | Perceptually Scalable Extension of H.264abstractWe propose a novel visual scalable video coding (VSVC) framework, named VSVC H.264/AVC. In this approach, the non-uniform sampling characteristic of the human eye is used to modify scalable video coding (SVC) H.264/AVC. We exploit the visibility of video content and the scalability of the video codec to achieve optimal subjective visual quality given limited system resources. To achieve the largest coding gain with controlled perceptual quality degradation, a perceptual weighting scheme is deployed wherein the compressed video is weighted as a function of visual saliency and of the non-uniform distribution of retinal photoreceptors. We develop a resource allocation algorithm emphasizing both efficiency and fairness by controlling the size of the salient region in each quality layer. Efficiency is emphasized on the low quality layer of the SVC. The bits saved by eliminating perceptual redundancy in regions of low interest are allocated to lower block-level distortions in salient regions. Fairness is enforced on the higher quality layers by enlarging the size of the salient regions. The simulation results show that the proposed VSVC framework significantly improves the subjective visual quality of compressed videos. Hojin Ha, Jincheol Park, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Cross-Layer Optimization for Downlink Wavelet Video TransmissionabstractCross-layer optimization for efficient multimedia communications is an important emerging issue towards providing better quality-of-service (QoS) over capacity-limited wireless channels. This paper presents a cross-layer optimization approach that operates between the application and physical layers to achieve high fidelity downlink video transmission by optimizing with respect to a quality criterion termed “visual entropy” using Lagrangian relaxation. By utilizing the natural layered structure of wavelet coding, an optimal level of power allocation is determined, which permits the throughput of visual entropy to be maximized over a multi-cell environment. A theoretical approach to optimization using the Shannon capacity and the Karush-Kuhn-Tucker (KKT) conditions is explored when coupling the application with the physical layers. Simulations show that the throughput gain for cross-layer optimization by visual entropy is increased by nearly 80% at the cell boundary as compared with peak signal-to-noise ratio (PSNR). Hyungkeuk Lee, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Multim. | 2 |
| 2011 | CoMP-CSB for ICI Nulling with User SelectionabstractThe capacity of downlink multiple-input multiple-output (MIMO) cellular networks is significantly limited by inter-cell interference (ICI), particularly at cell boundaries. Recently, two types of coordinated multiple point transmission (CoMP) technologies, joint processing and transmission (JPT) and coordinated scheduling and beamforming (CSB), were proposed. These technologies are intended for the latest cellular communication standard in order to improve the performance of cell-edge users who suffer from significant ICI. In this paper, we propose an ICI cancellation technique based on a user selection algorithm for CoMP-CSB. Under partial channel state information (CSI) and no data sharing condition, each base station (BS) concentrates more on the direction of interference to the adjacent cell's users, during the user selection process. Unlike prior concepts for a single-cell environment, in which each BS generates a precoding matrix for selected users to be served, the proposed technique considers the effects of interference to users located in adjacent cells. Although there are obvious trade-offs between ICI mitigation and the number of simultaneous scheduled users in terms of system capacity, the simulation results demonstrate that our proposed algorithm achieves higher sector throughput and is more robust against ICI if the system is limited by interference. Furthermore, through simulation we are able to obtain the preferred option for the coordination distance (R) and the number of degrees of freedom for ICI nulling (ξ). Uk Jang, Hyukmin Son, Jongrok Park, Sanghoon Lee 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2011 | Definition of a new service class, artPS for video services over WiBro/Mobile WiMAX systems
Hyungkeuk Lee, Sanghoon Lee 0001 |
Wirel. Networks | 2 |
| 2011 | Hierarchical modulation-based cooperation utilizing relay-assistant full-duplex diversity
Hyukmin Son, Jongrok Park, Sanghoon Lee 0001 |
Wirel. Networks | 3 |
| 2010 | Temporal pooling of video quality estimates using perceptual motion modelsabstractEmerging multimedia applications have increased the need for video quality measurement. Motion is critical to this task, but is complicated owing to a variety of object movements and movement of the camera. Here, we categorize the various motion situations and deploy appropriate perceptual models to each category. We use these models to create a new approach to objective video quality assessment. Performance evaluation on the Laboratory for Image and Video Engineering (LIVE) Video Quality Database shows competitive performance compared to the leading contemporary VQA algorithms. Kwanghyun Lee, Jincheol Park, Sanghoon Lee 0001, Alan C. Bovik |
ICIP | 3 |
| 2010 | Perceptually Unequal Packet Loss Protection by Weighting Saliency and Error PropagationabstractWe describe a method for achieving perceptually minimal video distortion over packet-erasure networks using perceptually unequal loss protection (PULP). There are two main ingredients in the algorithm. First, a perceptual weighting scheme is employed wherein the compressed video is weighted as a function of the nonuniform distribution of retinal photoreceptors. Secondly, packets are assigned temporal importance within each group of pictures (GOP), recognizing that the severity of error propagation increases with elapsed time within a GOP. Using both frame-level perceptual importance and GOP-level hierarchical importance, the PULP algorithm seeks efficient forward error correction assignment that balances efficiency and fairness by controlling the size of identified salient region(s) relative to the channel state. PULP demonstrates robust performance and significantly improved subjective and objective visual quality in the face of burst packet losses. Hojin Ha, Jincheol Park, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Node distribution-based localization for large-scale wireless sensor networks
Sangjin Han, Sanghoon Lee 0001, Jongjun Park, Sangjoon Park |
Wirel. Networks | 3 |
| 2010 | Location-based QoS enhanced dynamic carrier allocation over multi-cell environments
Sanghoon Lee 0001 |
Wirel. Networks | 2 |
| 2010 | Optimal transmission methodology for QoS provision of multi-hop cellular network
Sanghoon Lee 0001 |
Wirel. Networks | 2 |
| 2010 | Analysis of resource allocation with macroscopic diversity for uplink OFDMA systems
Sanghoon Lee 0001, Gye-Tae Gil |
Wirel. Networks | 2 |
| 2009 | Optimal power allocation for minimizing visual distortion over MIMO communication systemsabstractA recent dynamic increase in demand for wireless multimedia services has greatly accelerated the research on cross layer optimization techniques for transmitting multimedia data over wireless channel. In this paper, we explore a novel theoretical approach for joint optimization between the rate distortion (RD) of H.264/AVC video and the link-capacity of MIMO parallel subchannels. We obtain the optimal power level of subchannels through an optimization problem to minimize total visual distortion. In the simulation results, compared to the water filling (WF) method, the proposed scheme provides better results in aspects of visual quality in the face of sum rate loss. Jincheol Park, Uk Jang, Taegeun Oh, Sanghoon Lee 0001, Alan C. Bovik |
ICIP | 4 |
| 2009 | Optimal Channel Adaptation of Scalable Video Over a Multicarrier-Based Multicell EnvironmentabstractTo achieve seamless multimedia streaming services over wireless networks, it is important to overcome inter-cell interference (ICI), particularly in cell border regions. In this regard scalable video coding (SVC) has been actively studied due to its advantage of channel adaptation. We explore an optimal solution for maximizing the expected visual entropy over an orthogonal frequency division multiplexing (OFDM)-based broadband network from the perspective of cross-layer optimization. An optimization problem is parameterized by a set of source and channel parameters that are acquired along the user location over a multicell environment. A suboptimal solution is suggested using a greedy algorithm that allocates the radio resources to the scalable bitstreams as a function of their visual importance. The simulation results show that the greedy algorithm effectively resists ICI in the cell border region, while conventional nonscalable coding suffers severely because of ICI. Jincheol Park, Hyungkeuk Lee, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Multim. | 3 |
| 2009 | Throughput and QoS improvement via fixed relay station cooperated beam-formingabstractThroughput and quality-of-service (QoS) over multicell environments are two of the most challenging issues that must be addressed when developing next generation wireless network standards. Currently, multiple-input/multiple-output (MIMO), inter-cell coordination and multi-hop relay technologies are viable options for improving channel capacity or coverage extension. Nevertheless, severe QoS degradation occurs in the outer region of multi-cells due to significant interference from neighboring cells or relay stations, thereby limiting overall performance. This paper describes an effective technique, fixed relay station cooperated beam-forming (FCBF), which combines MIMO, multi-hop relay and multi-cell coordination. Simulated testing of FCBF demonstrates an increase of 10% in the average sum-rate and a decrease of 25% in the outage probability compared with conventional techniques. In particular, throughput at the cell boundary is remarkably increased with FCBF compared with traditional beam-forming technologies. Hyukmin Son, Jongrok Park, Sanghoon Lee 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2009 | Multi-cell communications for OFDM-based asynchronous networks over multi-cell environments
Hyukmin Son, Sanghoon Lee 0001 |
Wirel. Networks | 2 |
| 2008 | Link Capacity Improvement by Utilizing Overlaid Bandwidth and Region DivisionabstractA cell planning and resource allocation scheme for improving channel capacity and for maintaining QoS (quality of service) over a downlink OFDMA (Orthogonal Frequency Division Multiple Access) system is proposed. The frequency overlay is applied to improve link capacity and the sectorization is reduce the influence of carrier collision. This consideration enhances link capacity of system with a proper outage probability. Taegeun Oh, Sanghoon Lee 0001, Gye-Tae Gil |
VTC Spring | 2 |
| 2008 | Dynamic Session Control for Scalable Video Coding over IMSabstractAt the view of Application layer, there are many researches to achieve the cross-layer optimization with Physical layer. Scalable video coding is good example. Although it is necessary to consider the session layer which lies halfway between two layers, the research about that is insufficient. We present a feasible solution of dynamic session control for scalable video coding over IMS. Suyoung Park, Sanghoon Lee 0001 |
VTC Spring | 2 |
| 2008 | The MAI Mitigation Scheme for OFDM-based Asynchronous Networks over Multi-Cell EnvironmentsabstractThis paper investigates the problem of MAI (multiple access interference) occurrence when timing misalignment occurs in the downlink channel of OFDM-based networks. The adaptive guard band allocation scheme is mainly designed to achieve timing synchronization while reducing guard carrier redundancy and maintaining high channel throughput in asynchronous OFDM-based networks. Thus, the proposed scheme reduces the effect of MAI for users receiving multiple transmission signals from surrounding BSs (Base Stations). In a simulation, it is noteworthy that this adaptive carrier allocation is very effective in reducing MAI for asynchronous multi-cell network. Hyukmin Son, Sanghoon Lee 0001, Gye-Tae Gil |
VTC Spring | 2 |
| 2008 | Optimal Carrier Loading Control for the Enhancement of Visual Quality over OFDMA Cellular NetworksabstractA recent dynamic increase in demand for wireless multimedia services has greatly accelerated the research on dynamic channel adaptation of high quality video applications. In this paper, we explore a theoretical approach to cross-layer optimization between multimedia and wireless networks by means of a quality criterion termed ldquovisual throughputrdquo for downlink video transmission using a layered coding algorithm. We obtain the optimal loading ratio of orthogonal frequency division multiple access (OFDMA) subcarriers through an optimization problem balancing the tradeoff relationship between inter-cell interference (ICI) and channel throughput. Through numerical link capacity analysis, we show that the upper bounds of the visual throughput gain at the cell boundary is obtained at about 27%. Uk Jang, Hyungkeuk Lee, Sanghoon Lee 0001 |
IEEE Trans. Multim. | 3 |
| 2007 | Optimal Carrier Loading for Maximizing Visual Entropy Over OFDMA Cellular NetworksabstractExplosively increasing demands for the wireless multimedia data have accelerated to research for adapting higher quality video applications in extensive efforts. In this paper, we explore a theoretical approach to cross-layer optimization between multimedia and wireless network by means of a quality criterion termed "visual entropy" for downlink video transmission, using a layered coding algorithm. We obtain the optimal loading ratio through an optimization problem, which aims at balancing the trade-off relationship between ICI (Inter-Cell Interference) and channel throughput. In the simulation, we show that the throughput gain in terms of visual entropy at the cell boundary is increased by up to about 32%. Uk Jang, Hyungkeuk Lee, Sanghoon Lee 0001 |
ICIP (5) | 3 |
| 2007 | Cross-Layer Optimization for Scalable Video Coding and Transmission Over Broadband Wireless NetworksabstractFor seamless multimedia streaming services, it is very important to overcome severe ICI (intercell interference) over wireless networks, particularly in the cell border region. SVC (scalable video coding) has been actively studied due to its advantage of channel adaptation. In this paper, we study an optimal solution for maximizing the expected visual entropy over an OFDM (orthogonal frequency division multiplexing)-based broadband network from the aspect of cross layer optimization. Based on this approach, an optimal set of source and channel parameters is obtained according to the user location over a multi-cell environment. From the numerical results, we quantify the optimal visual gain attained from the single layer video coding or the SVC according to channel capacity v.s. source data rate. Jincheol Park, Hyungkeuk Lee, Sanghoon Lee 0001 |
ICIP (3) | 3 |
| 2007 | Flexible Channel Allocation with Macroscopic Diversity for Uplink OFDMA SystemsabstractIn this paper, we evaluated effects of ICI (inter-cell interference) on an OFDM (orthogonal frequency division multiplexing) based wireless communication system's reverse link capacity and suggested macroscopic diversity oriented resource allocation scheme to reduce interference and enhance reverse link capacity. For this, the exact amount of ICIs in the system is measured and represented as a closed form. Sungjun Ham, Sanghoon Lee 0001, Gye-Tae Gil |
PIMRC | 2 |
| 2007 | Coexistence Performance Evaluation of IEEE 802.15.4 Under IEEE 802.11B Interference in Fading ChannelsabstractThe IEEE 802.15.4 standard specifies the physical and medium access control layers designed for low-rate wireless personal area networks. Its operational frequency band includes the 2.4 GHz industrial, scientific and medical band, which is also used by other IEEE 802 wireless standards. This paper presents the coexistence model of IEEE 802.15.4 with IEEE 802.1 lb interference in fading channels and proposes two adaptive channel allocation schemes. The first avoids the IEEE 802.15.4 interference only and the second avoids both of the IEEE 802.15.4 and the IEEE 802.11b interferences. Numerical results show that by selecting a channel which gives the maximum signal to noise ratio to the system, the proposed algorithms are effective for avoiding the interferences and for max-imizing the network capacity. Sangjin Han, Sanghoon Lee 0001, Yeonsoo Kim |
PIMRC | 3 |
| 2007 | The Reverse-Link Capacity Analysis of Multihop Cellular Networks over Multi-Cell EnvironmentsabstractIn this paper, a framework of link capacity analysis for the uplink MCN (Multi-hop Cellular Network) is presented, and the goal of which is to increase link capacity while mitigating the effect of interference. An overlaid network architecture is employed as the network topology : the multi-hop (single-hop) network at the outer (inner) region of the cell. In order to verify the improvement in capacity accrued from the inter-network cooperation, inter-network and intra-network interferences are redefined. In a simulation, the MCN exhibits a significant increase of 1.2 ~ 1.8 times in link capacity compared to a cellular network. Sangjin Han, Sungjun Ham, Sanghoon Lee 0001 |
PIMRC | 4 |
| 2007 | OFDM-Based Semi-Soft Handover for High Data Rate ServicesabstractVarious approaches for analyzing handover have been developed to guarantee the QoS (Quality of Service) of multimedia services over mobile communication networks. However, no framework for multicarrier-based broadband systems, such as MC-CDMA (Multi-Carrier Code Division Multiple Access) or OFDMA (Orthogonal Frequency Division Multiple Access) is available, from the perspective of link capacity. This paper presents a handover technique, referred to as semi-soft handover utilizing macro diversity, which permits both hard and soft handover advantages for services over multicarrier-based broadband networks to be retained. A theoretical analysis is then performed to measure the handover gain over the forward link in terms of an outage probability. The simulation data verifies that the semi-soft handover outperforms other traditional handover techniques, in particular, for high data rate services. Hyungkeuk Lee, Hyukmin Son, Sanghoon Lee 0001 |
PIMRC | 3 |
| 2007 | The Cell Planning Scheme for ICI MitigationabstractFRPA(frequency reuse power allocation) technique by employing the frequency reuse notion as a strategy for overcoming the ICI(intercell interference) and maintaining a QoS(quality of service) at the cell boundary is described for broadband cellular networks. In the scheme, the total bandwidth is divided into sub-bands and two different power levels are then allocated to sub-bands based on the frequency reuse for forward-link cell planning. In order to prove the effectiveness of the proposed algorithm, a Monte Carlo simulation was performed. The simulation shows that this technique can achieve high channel throughput while maintaining the required QoS at the cell boundary. Hyukmin Son, Sanghoon Lee 0001 |
PIMRC | 2 |
| 2006 | Performance Analysis of Location-based Dynamic Carrier Allocation in OFDM SystemsabstractIt is well known that high multi-user and multi-carrier diversity gains can be achieved via the use of DCA (dynamic channel allocation). However, there is no framework for analyzing the system performance of DCA algorithms over multi-cell environments. This paper presents a numerical analysis for measuring the performance gain for various DCA algorithms. Based on the analysis, an LQE (location-based QoS enhanced) DCA algorithm is proposed. In the simulation, we demonstrate that the proposed DCA algorithm is very effective from the perspective of fairness and throughput. Sungjun Ham, Sungho Jeon 0001, Younghyun Jeon, Sanghoon Lee 0001 |
GLOBECOM | 4 |
| 2006 | A Cell Search Technique based on Known Postfix for OFDM Cellular SystemsabstractA cell search technique utilizing a known postfix for OFDM (orthogonal frequency division multiplexing) cellular systems is described. The known postfix is generated in the time domain by inserting pilots in the frequency domain and plays the role of the cyclic prefix in general OFDM systems. Since it demonstrates good correlation properties, it can be facilitated to synchronize each symbol with an identified postfix. In this paper, two different known postfixes are allocated to each cell. One is used for cell identification and symbol synchronization, which is designed to be different among neighboring cells. The other is used for frame synchronization and is the same for all cells. In the simulation, the cell search is accomplished with a probability greater than 10-3at -27 dB in a vehicular channel. Even at -30 dB, the cell search probability is greater than 10-2in a pedestrian channel as well as 10-3in the AWGN (additive white Gaussian noise) channel. Younghyun Jeon, Jonghyung Kwun, Sanghoon Lee 0001 |
GLOBECOM | 3 |
| 2006 | Cross-Layer Optimization for Downlink Wavelet Image TransmissionabstractCross-layer optimization for efficient multimedia communications has been an emerging issue, in terms of providing better QoS (quality of service) over a capacity-limited wireless channel. This paper presents a cross-layer optimization approach between multimedia and wireless network layers for downlink image transmission by means of a quality criterion termed "visual entropy" using Lagrangian relaxation. By utilizing the natural layered structure of wavelet coding, an optimal level of power allocation is determined, to permit the throughput of visual entropy for downlink broadband wireless systems to be maximized over a multi-cell environment. In the simulation, the throughput gain of visual entropy is dramatically increased up at the cell boundary, with respect to the case of no considerations. Hyungkeuk Lee, Sungho Jeon 0001, Sanghoon Lee 0001 |
GLOBECOM | 3 |
| 2006 | Visual Entropy Gain for Wavelet Image CodingabstractWavelet image coding exhibits a robust error resilience performance by utilizing a naturally layered bitstream construction over a band-limited channel. In this letter, a new measure that appears to provide a better assessment of visual entropy for comparing and evaluating progressive image coders is defined based on a visual weight over the wavelet domain. This visual weight is characterized by the human visual system (HVS) over the frequency and spatial domains and is then utilized as a criterion for determining the coding order of wavelet coefficients, resulting in improved visual quality. A transmission gain, which is expressed by visual entropy, of up to about 23% can be obtained at a normalized channel throughput of about 0.3. In accordance with the subjective visual quality, a relatively high gain can be obtained at a low channel capacity. Hyungkeuk Lee, Sanghoon Lee 0001 |
IEEE Signal Process. Lett. | 2 |
| 2005 | Visual data rate gain for wavelet foveated image codingabstractWavelet image coding exhibits a robust error resilience performance by utilizing the naturally layered bitstream construction over a band-limited channel. In this paper, a closed form of the visual entropy is defined based on a visual weight over the wavelet domain, which is characterized by the HVS (human vision system) over the frequency and spatial domains. This visual weight is then utilized as a criterion for determining the coding order of wavelet coefficients and the resulting improved visual quality is demonstrated. In terms of visual entropy, a transmission gain of up to 20 % can be obtained at a normalized channel throughput of about 0.3. Hyungkeuk Lee, Sanghoon Lee 0001 |
ICIP (3) | 2 |
| 2005 | High quality, low delay foveated visual communications over mobile channels
Sanghoon Lee 0001, Alan C. Bovik, Young Yong Kim |
J. Vis. Commun. Image Represent. | 1 |
| 2004 | SOVA equalization for multi-code CDMA system with low spreading factorabstractIn a CDMA system, the RAKE receiver is commonly used to attain the diversity gain by taking advantage of the good correlation properties of the spreading codes. However, at low spreading gains the good correlation properties of the spreading codes are lost and the RAKE receiver performance is severely degraded by interpath interference (IPI). In the case of multi-code CDMA system, the multi-code interference (MCI) exists in the system. In order to suppress MPI and MCI, a novel receiver based on soft-output Viterbi algorithm (SOVA) equalization is proposed in this paper. The SOVA equalization is applied to symbol sequences after RAKE combining and MCI cancellation to effectively eliminate the IPI during transmission of high rate data in wideband DS-CDMA systems. Simulation results show that the proposed receiver significantly outperform the traditional RAKE and RAKE-VA receivers. Junhui Zhao 0001, Sanghoon Lee 0001, Dongming Wang 0002, Xiaohu You 0001 |
PIMRC | 2 |
| 2003 | Fast algorithms for foveated video processingabstractThis paper explores the problem of communicating high-quality, foveated video streams in real time. Foveated video exploits the nonuniform resolution of the human visual system by preferentially allocating bits according to the proximity to assumed visual fixation points, thus delivering perceptually high quality at greatly reduced bandwidths. Foveated video streams possess specific data density properties that can be exploited to enhance the efficiency of subsequent video processing. Here, we exploit these properties to construct several efficient foveated video processing algorithms: foveation filtering (local bandwidth reduction), motion estimation, motion compensation, video rate control, and video postprocessing. Our approach leads to enhanced computational efficiency by interpreting nonuniform-density foveated images on the uniform domain and by using a foveation protocol between the encoder and the decoder. Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | Foveated video quality assessmentabstractMost image and video compression algorithms that have been proposed to improve picture quality relative to compression efficiency have either been designed based on objective criteria such as signal-to-noise-ratio (SNR) or have been evaluated, post-design, against competing methods using an objective sample measure. However, existing quantitative design criteria and numerical measurements of image and video quality both fail to adequately capture those attributes deemed important by the human visual system, except, perhaps, at very low error rates. We present a framework for assessing the quality of and determining the efficiency of foveated and compressed images and video streams. Image foveation is a process of nonuniform sampling that accords with the acquisition of visual information at the human retina. Foveated image/video compression algorithms seek to exploit this reduction of sensed information by nonuniformly reducing the resolution of the visual data. We develop unique algorithms for assessing the quality of foveated image/video data using a model of human visual response. We demonstrate these concepts on foveated, compressed video streams using modified (foveated) versions of H.263 that are standard-compliant. We rind that quality vs. compression is enhanced considerably by the foveation approach. Sanghoon Lee 0001, Marios S. Pattichis, Alan C. Bovik |
IEEE Trans. Multim. | 1 |
| 2001 | Foveated video compression with optimal rate controlabstractPreviously, fovcated video compression algorithms have been proposed which, in certain applications, deliver high-quality video at reduced bit rates by seeking to match the nonuniform sampling of the human retina. We describe such a framework here where foveated video is created by a nonuniform filtering scheme that increases the compressibility of the video stream. We maximize a new foveal visual quality metric. the foveal signal-to-noise ratio (FSNR) to determine the best compression and rate control parameters for a given target bit rate. Specifically, we establish a new optimal rate control algorithm for maximizing the FSNR using a Lagrange multiplier method defined on a curvilinear coordinate system. For optimal rate control, we also develop a piecewise R-D (rate-distortion)/R-Q (rate-quantization) model. A fast algorithm for searching for an optimal Lagrange multiplier lambda* is subsequently presented. For the new models, we show how the reconstructed video quality is affected, where the FSNR is maximized, and demonstrate the coding performance for H.263,+,++/MPEG-4 video coding. For H.263/MPEG video coding, a suboptimal rate control algorithm is developed for fast, high-performance applications. In the simulations, we compare the reconstructed pictures obtained using optimal rate control methods for foveated and normal video. We show that foveated video coding using the suboptimal rate control algorithm delivers excellent performance under 64 kb/s. Sanghoon Lee 0001, Marios S. Pattichis, Alan C. Bovik |
IEEE Trans. Image Process. | 1 |
| 2000 | Unequal Error Protection for Foveation-Based Error Resilience over Mobile NetworksabstractIn this paper, we introduce an unequal error protection technique for foveation-based error resilience over highly error-prone mobile networks. For point-to-point visual communications, visual quality can be significantly increased by using foveation-based error resilience where each frame is divided into foveated and background layers according to the gaze direction of the human eye, and two bitstreams are generated. In an effort to increase the source throughput of the foveated layer, we employ unequal delay-constrained ARQ and RCPC (rate compatible punctured convolutional) codes in H.223 Annex C. In the simulation, the visual quality is increased in the range of 0.3 dB to 1 dB over channel SNR 5 dB to 15 dB. Sanghoon Lee 0001, Christine Podilchuk, Vidhya Krishnan, Alan C. Bovik |
ICIP | 1 |
| 1999 | Very low bit rate foveated video coding for H.263abstractRecently, foveated video has been introduced as an important emerging method for very low bit rate multimedia applications. In this paper, we develop several rate control algorithms, and measure the performance of foveated video. We utilize H.263 video, and compare the performance with regular video based on the SNRC (signal-to-noise-ratio in curvilinear coordinates). In order to maximize compression, we use a maximum quantization parameter (QP=31) for the regular video, and code a foveated video sequence at the equivalent bit rate. In simulation, we improve the PSNRC to 3.64 (1.62)dB under 30 (14) kbits/sec for P pictures in CIF "News" ("Akiyo") standard video sequence. Sanghoon Lee 0001, Alan C. Bovik |
ICASSP | 1 |
| 1999 | Motion Estimation and Compensation for Foveated VideoabstractUtilizing the nonuniform resolution property of the human visual system, foveated video can provide high visual quality relative to non-foveated video by allocating more bits to the central foveation area. In this paper, we present a motion estimation and compensation algorithm for foveated video and measure the performance using a new measure of visual fidelity termed foveal mean absolute distortion. The computation redundancy reduction is achieved by subsampling searching area dependent on the local bandwidth in the sense of Nyquist sampling criterion. In addition, we reduce motion compensated errors by increasing temporal correlation when single or multiple foveation points are added or subtracted. Sanghoon Lee 0001, Alan C. Bovik |
ICIP (2) | 1 |
| 1999 | Low Delay Foveated Visual Communications over Wireless ChannelsabstractThe great potential of "foveated imaging" lies in the entropy reduction relative to the original image while minimizing the loss of visual information. Utilizing human foveation combined with video compression, as well as communication and human-machine interface techniques, more efficient multimedia services are expected to be provided in the near future. In this paper, we introduce a prototype for foveated visual communications as one of future human interactive multimedia applications, and demonstrate the benefit of the foveation over fading statistics in the downtown area of Austin, Texas. In order to compare the performance with regular video, we use spatial/temporal resolution and source transmission delay as the evaluation criteria. Sanghoon Lee 0001, Alan C. Bovik, Young Yong Kim |
ICIP (3) | 1 |
| 1998 | Rate Control for Foveated MPEG/H.263 VideoabstractGiven a set of target bits, video rate control algorithms that use Lagrange multipliers have been generally known as an optimal solution for maximizing the picture quality in the uniform spatial domain. Even if the SNR (signal-to-noise) of a picture is maximized by the rate control scheme, the visual quality can be enhanced using a suitable algorithm for the human visual system. We establish a new optimal rate control algorithm for maximizing the SNRC (signal-to-noise ratio in curvilinear coordinates) using the Lagrange multiplier. In addition, a target bit allocation technique for foveated video is introduced for simplified rate control over MEPG/H.263 video standards. Sanghoon Lee 0001, Marios S. Pattichis, Alan C. Bovik |
ICIP (2) | 1 |
| 1998 | Maximally Flat Bandwidth Allocation for Variable Bit Rate VideoabstractWhen a variable bit rate (VBR) video bitstream is transmitted into a network, the channel utilization is dependent on the maximum peak value of the traffic pattern and improved by reducing the peak rate. A lower bound on the transmission rate is derived from the given traffic cells. Based on the transmission rate constraint, the optimal bandwidth that is needed to meet the delay requirement during an allocation interval is also derived. From the delay bound and the constraints, we demonstrate that the traffic has been made as flat as possible. Sanghoon Lee 0001, Alan C. Bovik |
ICIP (2) | 1 |
| 1994 | Dynamic Bandwidth Allocation for Multiple VBR MPEG Video SourcesabstractIn this paper, we present a dynamic bandwidth allocation scheme which allots available transmission rate to each video source in proportion to its traffic. While individual elementary streams are encoded with limited VBR, the bit rate of multiplexed bit streams can be CBR or VBR depending on the transmission networks. The simulation result shows improvement more than 1 dB in average PSNR and smaller variance in PSNR compared to the CBR video sources among multiplexed different video sources. This scheme is able to be applied to satellite or terrestrial broadcasting as well as ATM transmission.> Sanghoon Lee 0001, Seong Hwan Jang, Jeong Su Lee |
ICIP (1) | 1 |