Jiwoo Kang 0001

dblp:141/0692-1 · also Ji-Woo Kang 0001 · DBLP profile ↗
← Back
28ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0001-7622-0817ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UniCross3D: Unified Cross-View and Cross-Domain Diffusion for Consistent Single-Image 3D Generation
abstract
Abstract Reconstructing detailed geometry and realistic appearance from a single RGB image is essential yet fundamentally challenging due to inherent ambiguities such as occlusion, lighting variations, and texture‐geometry entanglement. While recent diffusion‐based generative models have significantly improved novel view synthesis, existing approaches suffer from two critical limitations: lack of cross‐view geometric consistency and insufficient cross‐domain semantic alignment. To address these issues, we introduce U ni C ross 3D , a unified cross‐view and cross‐domain diffusion framework designed explicitly for consistent and physically coherent 3D generation. U ni C ross 3D features two novel contributions: (1) a cross‐view latent regularization that enforces cross‐view geometric consistency across synthesized viewpoints by penalizing latent variance, and (2) a cross‐domain mutual information objective grounded in the physics of image formation, explicitly aligning synthesized color and normal maps. Extensive experiments demonstrate that U ni C ross 3D achieves significantly improved view consistency and semantic alignment over state‐of‐the‐art methods and yields higher‐fidelity reconstructions, particularly under challenging textures and ambiguous viewpoints.
U-Chae Jun, Jaeeun Ko, Jiwoo Kang 0001
Comput. Graph. Forum3
2026 Latent Diffusion-GAN: Adversarial Learning in the Autoencoded Latent Space
abstract
Abstract Diffusion models are powerful generative frameworks for producing high‐quality images by denoising latent variables from random noise. However, training with likelihood‐based objectives, such as denoising score matching, can lead to locally oversmoothed high‐frequency details, including fine textures and sharp edges, thereby limiting perceptual fidelity and structural detail. Adversarial training with GANs enhances sharpness but typically requires additional discriminator networks, increasing computational costs and destabilizing training. To this end, we propose Latent Diffusion Generative Adversarial Networks (LD‐GAN), a novel framework that seamlessly integrates adversarial learning into diffusion models without modifying their original pipeline. LD‐GAN leverages the pretrained variational autoencoder (VAE) in latent diffusion models as an energy‐based discriminator, enabling adversarial training without extra parameters and preserving the structured latent priors learned from large datasets. We also introduce a structural consistency energy that aligns encoder and decoder feature representations, thereby enhancing perceptual quality and compatibility with the pretrained latent space. Extensive experiments demonstrate that LD‐GAN significantly improves sample fidelity, perceptual sharpness, and diversity over state‐of‐the‐art baseline methods across various generation tasks while ensuring efficient training dynamics.
U-Chae Jun, Jaeeun Ko, Jiwoo Kang 0001
Comput. Graph. Forum3
2026 Structure and sensitivity in 3D human pose similarity quantification and estimation
Kyoungoh Lee, Jungwoo Huh, Jiwoo Kang 0001, Sanghoon Lee 0001
Pattern Recognit.3
2025 SinWaveFusion: Learning a single image diffusion model in wavelet domain
Jiwoo Kang 0001, Taewan Kim 0002, Heeseok Oh
Image Vis. Comput.2
2025 Collaborative feature aggregation for face super-resolution and robust re-identification
Juheon Hwang, Taewan Kim 0002, Jiwoo Kang 0001
Multim. Syst.3
2025 Convolutional neural shading for high-quality 3D reconstruction from multi-view images
Juheon Hwang, Taewan Kim 0002, Heeseok Oh, Jiwoo Kang 0001
Multim. Syst.4
2025 Face and voice cross-modal association with learning convex feature embedding
Taewan Kim 0002, Jiwoo Kang 0001
Multim. Syst.2
2025 A Novel Intelligent Video Surveillance System Using Low-Traffic Scene-Preserving Video Anonymization
abstract
With the development of computer vision technology, intelligent video surveillance systems have been developed for automatic monitoring. However, the problem of personal information protection has also emerged. Existing systems attempted to solve this problem by anonymizing a video by, for example, sending only low-dimensional abstract information such as a person’s 2D pose or blurring a person’s face in the video before sending it to the central cloud server. However, these approaches failed to balance scene-preservation and traffic efficiency, because abstract information is too limited for preserving the entire scene, and video modification generates massive traffic. This article proposes a novel intelligent video surveillance system to overcome such limitations that preserves the scene information and generates minimal traffic through video anonymization. The proposed system reconstructs 3D human models and estimates segmentation masks to preserve a scene captured by a surveillance camera in its entirety. Parametric models represent 3D human models with several sets of parameters, and dictionary coding compresses the segmentation mask with a high compression ratio. The system follows the edge-cloud architecture, where the edge node extracts and transmits the scene information and the central cloud server generates the final anonymized video. We demonstrate the effectiveness of the proposed system by conducting experiments on processing time, scene preservation, and traffic efficiency. Our proposed system runs in real-time ( \(>\) 25fps) in a typical hardware setting and has a data compression ratio of more than 5,000 compared with raw data transfer while maintaining over 85% scene-preservation correlation with the original video.
Jungwoo Huh, Jiwoo Kang 0001, Jongwook Woo, Sanghoon Lee 0001
ACM Trans. Intell. Syst. Technol.2
2025 3D Facial Shape Similarity with Deep Perceptual Representations
abstract
Comparing different 3D shapes is challenging due to their irregularities. Motivated by the human visual system mechanism, where the entire 3D geometry is clearly perceived as a series of multiple projections, we propose a novel facial shape similarity measurement using multiview deep perceptual representations. We introduce a multiview disentangling scheme that accurately represents a facial mesh in multiple coordinates and the training strategy with view specificity and regional consistency to reliably train the network with multiple projections. View specificity pertains to the human visual perception to better recognize facial similarity. Regional consistency mitigates regional redundancy among views. Hence, robust perceptual features with respect to views are embedded and accurate similarity can be measured. Consequently, the view-specific integration scheme incorporates the similarities of all views, allowing for highly consistent measurement. The experiments demonstrate that the proposed similarity outperforms state-of-the-arts and significantly improves the details in terms of geometry and human perception.
Seongmin Lee 0002, Jiwoo Kang 0001, Sanghoon Lee 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Real-time Abnormal Behavior Recognition for Patient Monitoring in Hospitals
abstract
Due to a shortage of medical staff, psychiatric nurses often find themselves responsible for as many as 16 or more patients, making it challenging to provide personalized attention to individuals requiring both physical and mental care. For this reason, we propose a real-time abnormal behavior recognition algorithm in hospitals. Our system utilizes real-time video analysis to detect and track the locations of mental patients, enabling the identification of their abnormal behaviors. Specifically, we have defined distinct abnormal behaviors commonly observed in closed wards, such as Self-Harm, Falldown, and Hit. To improve recognition performance, we applied the continual learning method, allowing the system to adapt and enhance its capabilities. In addition, the architecture can create our abnormal behavior dataset. The average abnormal behavior recognition accuracy of the system exceeds 90%. By decreasing the likelihood of encountering dangerous incidents, our proposed method not only improves the wellbeing of patients but also fosters a safer working environment for medical staff.
Hyewon Song, Jiwoo Kang 0001, Taewan Kim 0002
AVSS2
2024 Convex Feature Embedding for Face and Voice Association
Jiwoo Kang 0001, Taewan Kim 0002, Young-Ho Park 0002
SIGIR1
2024 Double reverse diffusion for realistic garment reconstruction from images
Jeonghaeng Lee, Jongyoo Kim, Jiwoo Kang 0001, Sanghoon Lee 0001
Eng. Appl. Artif. Intell.4
2024 EMOVA: Emotion-driven neural volumetric avatar
Juheon Hwang, Byung-Gyu Kim, Taewan Kim 0002, Heeseok Oh, Jiwoo Kang 0001
Image Vis. Comput.5
2024 3D-PSSIM: Projective Structural Similarity for 3D Mesh Quality Assessment Robust to Topological Irregularities
abstract
Despite acceleration in the use of 3D meshes, it is difficult to find effective mesh quality assessment algorithms that can produce predictions highly correlated with human subjective opinions. Defining mesh quality features is challenging due to the irregular topology of meshes, which are defined on vertices and triangles. To address this, we propose a novel 3D projective structural similarity index ( 3D- PSSIM) for meshes that is robust to differences in mesh topology. We address topological differences between meshes by introducing multi-view and multi-layer projections that can densely represent the mesh textures and geometrical shapes irrespective of mesh topology. It also addresses occlusion problems that occur during projection. We propose visual sensitivity weights that capture the perceptual sensitivity to the degree of mesh surface curvature. 3D- PSSIM computes perceptual quality predictions by aggregating quality-aware features that are computed in multiple projective spaces onto the mesh domain, rather than on 2D spaces. This allows 3D- PSSIM to determine which parts of a mesh surface are distorted by geometric or color impairments. Experimental results show that 3D- PSSIM can predict mesh quality with high correlation against human subjective judgments, across the presence of noise, even when there are large topological differences, outperforming existing mesh quality assessment models.
Seongmin Lee 0002, Jiwoo Kang 0001, Sanghoon Lee 0001, Weisi Lin, Alan C. Bovik
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Single-Image 3-D Reconstruction: Rethinking Point Cloud Deformation
abstract
Single-image 3-D reconstruction has long been a challenging problem. Recent deep learning approaches have been introduced to this 3-D area, but the ability to generate point clouds still remains limited due to inefficient and expensive 3-D representations, the dependency between the output and the number of model parameters, or the lack of a suitable computing operation. In this article, we present a novel deep-learning-based method to reconstruct a point cloud of an object from a single still image. The proposed method can be decomposed into two steps: feature fusion and deformation. The first step extracts both global and point-specific shape features from a 2-D object image, and then injects them into a randomly generated point cloud. In the second step, which is deformation, we introduce a new layer termed as GraphX that considers the interrelationship between points like common graph convolutions but operates on unordered sets. The framework can be applicable to realistic image data with background as we optionally learn a mask branch to segment objects from input images. To complement the quality of point clouds, we further propose an objective function to control the point uniformity. In addition, we introduce different variants of GraphX that cover from best performance to best memory budget. Moreover, the proposed model can generate an arbitrary-sized point cloud, which is the first deep method to do so. Extensive experiments demonstrate that we outperform the existing models and set a new height for different performance metrics in single-image 3-D reconstruction.
Seonghwa Choi, Woojae Kim, Jongyoo Kim, Heeseok Oh, Jiwoo Kang 0001, Sanghoon Lee 0001
IEEE Trans. Neural Networks Learn. Syst.6
2023 Audio-visual Neural Face Generation with Emotional Stimuli
abstract
In communication, it is important to generate realistic 3D facial avatars that can express various expressions. However, existing models have limitations in that they cannot show facial expressions in sufficient detail. To this end, we propose an audio-visual volumetric avatar generation method that reconstructs detailed facial expressions using two emotional stimuli. Emotional features extracted from the input and the rendered image induce the model to generate an avatar with the same emotion. At the same time, a speech clip is used as input along with the images, helping the model recognize the emotion of the target that is difficult to understand with the images. We showed qualitatively and qualitatively that our method achieved significantly more realistic and higher-quality avatar generation than current state-of-the-art methods.
Juheon Hwang, Jiwoo Kang 0001
IEEE Big Data2
2023 Unlocking Potential of 3D-aware GAN for More Expressive Face Generation
abstract
As style-based image generators have achieved disentanglement in features by converting latent vector space to style vector space, numerous efforts have been made to enhance the controllability of the latent. However, existing methods for controllable models have limitations in precisely creating high-resolution faces with large expressions. The degradation is due to the dependence on the training dataset, as the high-resolution face datasets do not have sufficient expressive images. To tackle this challenge, we propose a robust training framework for 3D-aware generative adversarial networks to learn the high-quality generation of more expressive faces through a signed distance field. First, we propose a novel 3D enforcement loss to generate more expressive images in an unsupervised manner. Second, we introduce a partial training method to fine-tune the network on multiple datasets without loss of image resolution. Finally, we propose a ray-scaling scheme for the volume renderer to represent a face at arbitrary scales. Through the proposed framework, the network learns 3D face priors, such as expressional shapes of the parametric facial model, to generate detailed faces. The experimental results outperform the methods of the state of the art, showing strong benefits in the generation of high-resolution facial expressions.
Juheon Hwang, Jiwoo Kang 0001, Kyoungoh Lee, Sanghoon Lee 0001
ICMR2
2023 Video-Based Stabilized 3D Face Alignment Using Temporal Multi-Discrimination
abstract
Existing 3D face alignment primarily aim to achieve accurate face alignment result for a static facial image. While these methods have strong alignment performance under large poses, occlusion, and extreme lighting conditions, they often result in trembling artifacts in video-based sequential 3D face alignment. Reducing temporal misalignment remains a challenging task because a single misaligned frame can propagate errors to other frames along the temporal axis. To address this issue, we propose a novel temporal discriminating scheme that learns the distribution gap between the face alignment results and ground truth face animation. By leveraging the discrimination results as a guide, the proposed method can effectively align the 3D faces to the input video by reducing temporal trembling artifacts. To effectively learn the distribution gap, we introduce a multi-discriminating scheme that separately discriminates facial animation based on identity and expression changes. It enables the proposed method to produce a stabilized alignment result, especially in dynamic and fast movement. Through extensive experiments in both qualitative and quantitative evaluations, it is confirmed that our method outperforms state-of-the-art 3D face alignment methods by animating stabilized results in the video.
Seongmin Lee 0002, Hyunse Yoon, Jiwoo Kang 0001, Jungsu Kim, Jiwan Son, Jungwoo Huh, Sanghoon Lee 0001
MMSP3
2022 Gradient Flow Evolution for 3D Fusion From a Single Depth Sensor
abstract
We present a novel real-time framework for non-rigid 3D reconstruction that is robust to noise, camera poses, and large deformation from a single depth camera. KinectFusion has achieved high-quality 3D object reconstructions in real-time by implicitly representing an object’s surface with a signed distance field (SDF) representation from a single depth camera. Many studies for incremental reconstruction have been presented since then, with the surface estimation improving over time. Previous works primarily focused on improving conventional SDF matching and deformation schemes. In contrast to these works, the proposed framework tackles the problem of temporal inconsistency caused by SDF approximation and fusion to manipulate SDFs and reconstruct a target more accurately over time. In our reconstruction pipeline, we introduce a refinement evolution method, where an erroneous SDF from a depth sensor is recovered more accurately in a few iterations by propagating erroneous SDF values from the surface. Reliable gradients of refined SDFs enable more accurate non-rigid tracking of a target object. Furthermore, we propose a level-set evolution for SDF fusion, enabling SDFs to be manipulated stably in the reconstruction pipeline over time. The proposed methods are fully parallelizable and can be executed in real-time. Qualitative and quantitative evaluations show that incorporating the refinement and fusion methods into the reconstruction pipeline improves 3D reconstruction accuracy and temporal reliability by avoiding cumulative errors over time. Evaluation results show that our pipeline results in more accurate reconstruction that is robust to noise and large motions, as well as outperforms previous state-of-the-art reconstruction methods.
Jiwoo Kang 0001, Seongmin Lee 0002, Mingyu Jang, Sanghoon Lee 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Self-Updatable Database System Based on Human Motion Assessment Framework
abstract
Recently, human motion-centric videos have been attracting attention in the field of computer vision. Observing and detecting human motion in intelligent surveillance camera systems is essential for understanding the intentions of target subjects. However, these videos have vast amounts of disparate and complex information, and hence they are difficult to process and label automatically. As a result, building and maintaining a database using motion-centric videos requires considerable labor in trimming and classifying the videos. Therefore, we propose a self-updatable motion database system based on a human motion assessment framework for evaluating complex human movements. The framework quantifies three primitive motion properties: stability, liveliness, and attention. This assessment highlights the semantics of human motion in the input video. The semantic motion sequence obtained after the motion assessment is compared with a similarity motion database to determine whether the database needs to be updated; for efficient comparison, we introduce a sequential autoencoder model with a long short-term memory neural network. The proposed system maintains the database within a surveillance camera system using a motion update algorithm; unseen motions in the database are updated using a camera-based surveillance system. In addition, this framework combines state-of-art action recognition methods to improve performance by up to 11% via the self-update of motion.
Kyoungoh Lee, Yeseung Park, Jungwoo Huh, Jiwoo Kang 0001, Sanghoon Lee 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Competitive Learning of Facial Fitting and Synthesis Using UV Energy
abstract
The three-dimensional morphable model (3DMM) is the most widely used representative model for obtaining a three-dimensional (3-D) face from a target on an image. Although 3DMMs have demonstrated the powerful capability to represent various facial shapes on natural images, they are limited to capturing texture variations of in-the-wild human faces. Based on the fact that fitting a 3-D facial model to an image determines the corresponding UV map, we propose a novel method for facial fitting and synthesis by competitively training two deep learning networks for facial alignment and UV texture completion. When the completion network is trained using well-aligned UV maps, it can model facial textures precisely and, consequently, fill the missing regions more completely. Accordingly, we use a UV completion network, denoted as a UV energy-based generative adversarial network (UV EB-GAN), to discriminate whether a UV map from the alignment network is well aligned by defining the generative loss of the completion network as the energy. Competitive learning facilitates training the completion network without ground-truth facial UV maps and training the alignment network without hard constraints and regularization terms. The proposed network can be trained in an end-to-end manner. The facial texture, albedo, lighting parameters, and 3-D facial shape can be obtained through this network. The results of the experiments on 2-D alignment, 3-D reconstruction, texture synthesis, and illumination estimation verified that the proposed method achieves remarkable improvements over the state-of-the-art methods.
Jiwoo Kang 0001, Seongmin Lee 0002, Sanghoon Lee 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2021 WarpingFusion: Accurate Multi-View TSDF Fusion with Local Perspective Warp
abstract
In this paper, we propose the novel 3D reconstruction framework, where the surface of a target object is reconstructed accurately and robustly from multi-view depth maps. A depth map of a moving object tends to have the spatially-varying perspective warps due to motion blur and rolling shutter artifacts. Incorporating those misaligned points from the views into the world coordinate leads to significant artifacts in the reconstructed shape. We address the mismatches by the patch-based depth-to-surface alignment using implicit surface-based distance measurement. The patch-based minimization finds spatial warps on the depth map fast and accurately with the global transformation preserved. The proposed framework efficiently optimizes the local alignments against depth occlusions and local variants thanks to the point to surface distance based on an implicit representation. The proposed method shows significant improvements over the other reconstruction methods, demonstrating efficiency and benefits of our method in the multi-view reconstruction.
Jiwoo Kang 0001, Seongmin Lee 0002, Mingyu Jang, Hyunse Yoon, Sanghoon Lee 0001
ICIP1
2021 Deep Chessboard Corner Detection Using Multi-task Learning
abstract
Camera calibration is an indispensable step in the fields of robotics and computer vision, which includes augmented reality, 3D reconstruction, and camera motion estimation. Before camera calibration, detecting matching correspondence is necessary to understand the structure of the world from multiple images. For an accurate result, a calibration object, such as a chessboard, is used. Existing handcrafted feature methods precisely detect chessboard corners but are weak against blurs, noises, and severe lens distortion. Conversely, neural network-based methods can detect corners regardless of noises in the image. Both methods do not utilize the information of camera priors, which are lens distortion and intrinsic parameters, affecting the location of chessboard corners. Learning of lens distortion and intrinsic parameters enables the proposed network to understand the alignment of corners more precisely. Therefore, in this paper, we propose a novel multi-task learning framework to detect chessboard corners and simultaneously estimate lens distortion and intrinsic parameters. In order to train these three tasks, synthetic images of the chessboard are generated with ground-truth labels corresponding to each task. Hence, by learning the camera priors, the proposed network can more precisely locate the corners than other state-of-the-art corner detection methods while robust to noises, blurs, and distortion.
Hyunse Yoon, Seongmin Lee 0002, Jiwoo Kang 0001, Sanghoon Lee 0001
MMSP3
2020 PatchMatch based Multiview Stereo with Local Quadric Window
abstract
Although various stereo matching methods are studied in many years, the accurate 3D reconstruction from multiview stereos in high-fidelity is still challenging due to the surface inconsistency caused by various factors such as specular illumination. In this paper, we propose an accurate PatchMatch based multiview stereo matching method with a quadric support window that efficiently captures the surface of a complex structured object. Our method takes three novel contributions. Firstly, delicate surface configurations are used for representing the complex structure of an object. By using a general 3D quadric function, the structured object surfaces can be estimated more accurately. In addition, an illumination robust framework is proposed, where the patch dissimilarities are precisely measured with disentangled representation. The matching cost is defined based on disentangled measurements of the object photometric and geometric properties, balancing the pixel intensities between images robust to illumination. Lastly, a multiview propagation method is proposed to confirm shape consistency among views. Through the disparity refinement to unify plane parameters of the views, the object surface is estimated from a global perspective. Consequently, the dense and smooth 3D shape of the object is reconstructed accurately. We evaluate our proposed method on the Middlebury stereo set and conduct comprehensive experiments on facial images. Both quantitative and qualitative results demonstrate that the proposed method shows significant improvements over state-of-the-art methods.
Hyewon Song, Jaeseong Park, Suwoong Heo, Jiwoo Kang 0001, Sanghoon Lee 0001
ACM Multimedia4
2018 Fitting Facial Models to Spatial Points: Blendshape Approaches and Benchmark
abstract
Blendshape is one of the most common facial representation used for 3D animation, 3D game and virtual reality. In this paper, four representative blendshape approaches are benchmarked: global, delta, mean-delta, and SVD-based blend-shapes. When fitting the blendshape models to sparse facial points, the obtained facial shape highly depends on fitting approach due to the lack of the fitted points. Therefore, it is important to set up appropriate criteria for comparing and verifying the performance of the approaches. In this paper, we use four kinds of metrics that are utilized to measure the performance of the approaches: fitting, landmark, and vertex errors and coefficient sparsity. Through the experimental results, it is verified that the benchmarks are very effective to measure the subjective quality of blendshape.
Taelim Choi, Jiwoo Kang 0001, Hyewon Song, Sanghoon Lee 0001
ICIP2
2018 ConcatNet: A Deep Architecture of Concatenation-Assisted Network for Dense Facial Landmark Alignment
abstract
Facial landmark is one of the most basic elements for obtaining facial information such as facial expression and emotion. However, detecting dense landmarks on an image is challenging due to various facial poses. In this paper, a deep architecture for dense facial landmark detection, called ConcatNet, is proposed. In our architecture, we propose a CNN-based dense landmark detector on part regions of a face, which extends a given set of sparse landmarks to more accurate and dense landmarks. By introducing interface layers for coordinate normalization and part region localization, we concatenate a network for sparse landmark detection to ConcatNet in a global-to-local manner and the whole network to operate in an end-to-end manner. The experimental results on LFW and 300W datasets show that ConcatNet not only expands the number of the sparse landmarks but also increases the accuracy of the landmark positions remarkably. Also, ConcatNet shows high accuracy in detecting the dense landmarks with a smaller dataset and without additional data on an image such as 3D position annotations when compared to 3D model-based detection method.
Hyewon Song, Jiwoo Kang 0001, Sanghoon Lee 0001
ICIP2
2018 3D Active Vessel Tracking Using an Elliptical Prior
abstract
In this paper, we propose a novel vessel tracking method, called active vessel tracking (AVT). The proposed method retains the major advantages that most 2D segmentation methods have demonstrated for 3D tracking while overcoming the drawbacks of previous 3D vessel tracking methods. Under the assumption that the vessel is cylindrical, thereby making its cross-section elliptical, the AVT finds a plane perpendicular to the vessel axis while tracking the vessel along its length. Also, We propose a method for vessel branch detection to automatically track complete vascular networks from a single starting point, whereas the previously proposed solutions have usually been limited in handling vessel bifurcations precisely on 3D or have required considerable user interaction. Our results show that the method is robust and accurate in both synthetic and clinical cases. In an experiment on synthetic data sets, the proposed method achieved a tracking accuracy of 96.1±0.5, detecting 99.1% of the branches. In an experiment on abdominal CTA data sets, it achieved a tracking accuracy of 98.4±0.5 for six target vessels, detecting 98.3% of the branches. These results show that the proposed method can outperform previous methods for vessel tracking.
Jiwoo Kang 0001, Suwoong Heo, Woo Jin Hyung, Joon Seok Lim, Sanghoon Lee 0001
IEEE Trans. Image Process.1
2014 Multimodal Interactive Continuous Scoring of Subjective 3D Video Quality of Experience
abstract
People experience a variety of 3D visual programs, such as 3D cinema, 3D TV and 3D games, making it necessary to deploy reliable methodologies for predicting each viewer's subjective experience. We propose a new methodology that we call multimodal interactive continuous scoring of quality (MICSQ). MICSQ is composed of a device interaction process between the 3D display and a separate device (PC, tablet, etc.) used as an assessment tool, and a human interaction process between the subject(s) and the separate device. The scoring process is multimodal, using aural and tactile cues to help engage and focus the subject(s) on their tasks by enhancing neuroplasticity. Recorded human responses to 3D visualizations obtained via MICSQ correlate highly with measurements of spatial and temporal activity in the 3D video content. We have also found that 3D quality of experience (QoE) assessment results obtained using MICSQ are more reliable over a wide dynamic range of content than obtained by the conventional single stimulus continuous quality evaluation (SSCQE) protocol. Moreover, the wireless device interaction process makes it possible for multiple subjects to assess 3D QoE simultaneously in a large space such as a movie theater, at different viewing angles and distances. We conducted a series of interesting 3D experiments showing the accuracy and versatility of the new system, while yielding new findings on visual comfort in terms of disparity, motion and an interesting relation between the naturalness and depth of field (DOF) of a stereo camera.
Taewan Kim 0002, Jiwoo Kang 0001, Sanghoon Lee 0001, Alan C. Bovik
IEEE Trans. Multim.2