Taewan Kim 0002

dblp:79/2453-2 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0003-3319-7797ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Quality Is Not Comfort: Oculomotor Demand Entropy for Short-Form Videos
Taewan Kim 0002
IEEE Signal Process. Lett.1
2025 SinWaveFusion: Learning a single image diffusion model in wavelet domain
Jiwoo Kang 0001, Taewan Kim 0002, Heeseok Oh
Image Vis. Comput.3
2025 Collaborative feature aggregation for face super-resolution and robust re-identification
Juheon Hwang, Taewan Kim 0002, Jiwoo Kang 0001
Multim. Syst.2
2025 Convolutional neural shading for high-quality 3D reconstruction from multi-view images
Juheon Hwang, Taewan Kim 0002, Heeseok Oh, Jiwoo Kang 0001
Multim. Syst.2
2025 Face and voice cross-modal association with learning convex feature embedding
Taewan Kim 0002, Jiwoo Kang 0001
Multim. Syst.1
2024 Real-time Abnormal Behavior Recognition for Patient Monitoring in Hospitals
abstract
Due to a shortage of medical staff, psychiatric nurses often find themselves responsible for as many as 16 or more patients, making it challenging to provide personalized attention to individuals requiring both physical and mental care. For this reason, we propose a real-time abnormal behavior recognition algorithm in hospitals. Our system utilizes real-time video analysis to detect and track the locations of mental patients, enabling the identification of their abnormal behaviors. Specifically, we have defined distinct abnormal behaviors commonly observed in closed wards, such as Self-Harm, Falldown, and Hit. To improve recognition performance, we applied the continual learning method, allowing the system to adapt and enhance its capabilities. In addition, the architecture can create our abnormal behavior dataset. The average abnormal behavior recognition accuracy of the system exceeds 90%. By decreasing the likelihood of encountering dangerous incidents, our proposed method not only improves the wellbeing of patients but also fosters a safer working environment for medical staff.
Hyewon Song, Jiwoo Kang 0001, Taewan Kim 0002
AVSS3
2024 Convex Feature Embedding for Face and Voice Association
Jiwoo Kang 0001, Taewan Kim 0002, Young-Ho Park 0002
SIGIR2
2024 EMOVA: Emotion-driven neural volumetric avatar
Juheon Hwang, Byung-Gyu Kim, Taewan Kim 0002, Heeseok Oh, Jiwoo Kang 0001
Image Vis. Comput.3
2024 Kinematic Diversity and Rhythmic Alignment in Choreographic Quality Transformers for Dance Quality Assessment
abstract
In recent years, the dance entertainment industry has experienced significant growth, driven by the desire of consumers to learn and improve their dancing skills. To effectively improve their skills, dancers require evaluation and feedback, which traditionally relies heavily on professional dancers. To address this challenge, researchers have proposed objective assessment methods for dance performance via kinematic data captured by sensors. However, these existing methods primarily focus on assessing the rhythmic accuracy of movements synchronized to music. In this paper, we propose Dance Quality Assessment (DanceQA) Framework to evaluate dance performance, considering choreographic factors that are important criteria in subjective DanceQA. We find that kinematic diversity and rhythmic alignment are significant choreographic factors from human perception perspective. Based on these factors, we design two metrics: kinematic information entropy (KIE) and kinematic-music beat similarity (BSIM). Our study demonstrates that these metrics are closely related to specific body parts in each choreography. To validate the effectiveness of our metrics, we capture dance performance by OptiTrack system providing precise three-dimensional data at very high sampling rate. We then label their dance quality via subjective test. The metrics give strong correlation with subjective opinion, but it is difficult to tell which body part is the most correlated. To comprehensively understand the dance quality, we propose choreographic quality transformers (CQTs), which learn the aforementioned choreographic factors by embedding KIE and BSIM into attention matrices. In numerous experiments, the CQTs outperforms previous methods, graph convolutional networks and multimodal transformers, at least by up to 0.146 in correlation coefficient.
Taewan Kim 0002, Inwoong Lee, Sanghoon Lee 0001
IEEE Trans. Circuits Syst. Video Technol.2
2017 Perceptual Crosstalk Prediction on Autostereoscopic 3D Display
abstract
Perceptual crosstalk prediction for autostereoscopic 3D displays is of fundamental importance in determining the level of quality perceived by humans in terms of the display performance and the 3D viewing experience. However, no robust framework exists to quantify perceptual crosstalk while taking into account the hardware structure of a display as well as its content characteristics via content analysis. In this paper, we present a 3D perceptual crosstalk predictor (3D-PCP) that can be used to predict crosstalk in a unique way when viewing autostereoscopic 3D displays. 3D-PCP captures hardware features using an optical Fourier transform-light measurement device and content features through content analysis based on information theory. By deriving the disparity, luminance, color, and texture maps, this approach defines the visual entropy, mutual information, and relative entropy in order to investigate the influences of the 3D scene characteristics on perceptual crosstalk. The experimental results demonstrate that the 3D-PCP output is highly correlated with subjective scores.
Taewan Kim 0002, Jongyoo Kim, Sanghoon Lee 0001
IEEE Trans. Circuits Syst. Video Technol.1
2017 Quality Assessment of Perceptual Crosstalk on Two-View Auto-Stereoscopic Displays
abstract
Crosstalk is one of the most severe factors affecting the perceived quality of stereoscopic 3D images. It arises from a leakage of light intensity between multiple views, as in auto-stereoscopic displays. Well-known determinants of crosstalk include the co-location contrast and disparity of the left and right images, which have been dealt with in prior studies. However, when a natural stereo image that contains complex naturalistic spatial characteristics is viewed on an auto-stereoscopic display, other factors may also play an important role in the perception of crosstalk. Here, we describe a new way of predicting the perceived severity of crosstalk, which we call the Binocular Perceptual Crosstalk Predictor (BPCP). BPCP uses measurements of three complementary 3D image properties (texture, structural duplication, and binocular summation) in combination with two well-known factors (co-location contrast and disparity) to make predictions of crosstalk on two-view auto-stereoscopic displays. The new BPCP model includes two masking algorithms and a binocular pooling method. We explore a new masking phenomenon that we call duplicated structure masking, which arises from structural correlations between the original and distorted objects. We also utilize an advanced binocular summation model to develop a binocular pooling algorithm. Our experimental results indicate that BPCP achieves high correlations against subjective test results, improving upon those delivered by previous crosstalk prediction models.
Jongyoo Kim, Taewan Kim 0002, Sanghoon Lee 0001, Alan C. Bovik
IEEE Trans. Image Process.2
2017 Enhancement of Visual Comfort and Sense of Presence on Stereoscopic 3D Images
abstract
Conventional stereoscopic 3D (S3D) displays do not provide accommodation depth cues of the 3D image or video contents being viewed. The sense of content depths is thus limited to cues supplied by motion parallax (for 3D video), stereoscopic vergence cues created by presenting left and right views to the respective eyes, and other contextual and perspective depth cues. The absence of accommodation cues can induce two kinds of accommodation vergence mismatches (AVM) at the fixation and peripheral points, which can result in severe visual discomfort. With the aim of alleviating discomfort arising from AVM, we propose a new visual comfort enhancement approach for processing S3D visual signals to deliver a more comfortable 3D viewing experience at the display. This is accomplished via an optimization process whereby a predictive indicator of visual discomfort is minimized, while still aiming to maintain the viewer's sense of 3D presence by performing a suitable parallax shift, and by directed blurring of the signal. Our processing framework is defined on 3D visual coordinates that reflect the nonuniform resolution of retinal sensors and that uses a measure of 3D saliency strength. An appropriate level of blur that corresponds to the degree of parallax shift is found, making it possible to produce synthetic accommodation cues implemented using a perceptively relevant filter. By this method, AVM, the primary contributor to the discomfort felt when viewing S3D images, is reduced. We show via a series of subjective experiments that the proposed approach improves visual comfort while preserving the sense of 3D presence.
Heeseok Oh, Jongyoo Kim, Jinwoo Kim 0005, Taewan Kim 0002, Sanghoon Lee 0001, Alan C. Bovik
IEEE Trans. Image Process.4
2015 Transfer Function Model of Physiological Mechanisms Underlying Temporal Visual Discomfort Experienced When Viewing Stereoscopic 3D Images
abstract
When viewing 3D images, a sense of visual comfort (or lack of) is developed in the brain over time as a function of binocular disparity and other 3D factors. We have developed a unique temporal visual discomfort model (TVDM) that we use to automatically predict the degree of discomfort felt when viewing stereoscopic 3D (S3D) images. This model is based on physiological mechanisms. In particular, TVDM is defined as a second-order system capturing relevant neuronal elements of the visual pathway from the eyes and through the brain. The experimental results demonstrate that the TVDM transfer function model produces predictions that correlate highly with the subjective visual discomfort scores contained in the large public databases. The transfer function analysis also yields insights into the perceptual processes that yield a stable S3D image.
Taewan Kim 0002, Sanghoon Lee 0001, Alan C. Bovik
IEEE Trans. Image Process.1
2014 Quality assessment of perceptual crosstalk in autostereoscopic display
abstract
Crosstalk is one of the most annoying problems in an autostereoscopic display causing perceptual quality degradation and visual discomfort. To predict the perceived crosstalk when viewing an autostereoscopic display, it is necessary to consider the characteristics of human perception, displaying mechanism, viewing environment and so on. Therefor, we propose a novel metric for predicting the perceptual crosstalk that is based on human visual system (HVS); non-linear sensitivity of luminance and masking effects. The proposed model adopts the duplicated structure masking, yielding predictive power that is statistically superior to prior models that rely on 2D quality metric.
Jongyoo Kim, Taewan Kim 0002, Sanghoon Lee 0001
ICIP2
2014 Multimodal Interactive Continuous Scoring of Subjective 3D Video Quality of Experience
abstract
People experience a variety of 3D visual programs, such as 3D cinema, 3D TV and 3D games, making it necessary to deploy reliable methodologies for predicting each viewer's subjective experience. We propose a new methodology that we call multimodal interactive continuous scoring of quality (MICSQ). MICSQ is composed of a device interaction process between the 3D display and a separate device (PC, tablet, etc.) used as an assessment tool, and a human interaction process between the subject(s) and the separate device. The scoring process is multimodal, using aural and tactile cues to help engage and focus the subject(s) on their tasks by enhancing neuroplasticity. Recorded human responses to 3D visualizations obtained via MICSQ correlate highly with measurements of spatial and temporal activity in the 3D video content. We have also found that 3D quality of experience (QoE) assessment results obtained using MICSQ are more reliable over a wide dynamic range of content than obtained by the conventional single stimulus continuous quality evaluation (SSCQE) protocol. Moreover, the wireless device interaction process makes it possible for multiple subjects to assess 3D QoE simultaneously in a large space such as a movie theater, at different viewing angles and distances. We conducted a series of interesting 3D experiments showing the accuracy and versatility of the new system, while yielding new findings on visual comfort in terms of disparity, motion and an interesting relation between the naturalness and depth of field (DOF) of a stereo camera.
Taewan Kim 0002, Jiwoo Kang 0001, Sanghoon Lee 0001, Alan C. Bovik
IEEE Trans. Multim.1