Kwanggyoon Seo

dblp:274/1094 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0003-0570-4915ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation
Federico Nocentini, Kwanggyoon Seo, Qingju Liu, Claudio Ferrari, Stefano Berretti, David Ferman, Hyeongwoo Kim, Pablo Garrido 0001, Akin Caliskan
FG2
2026 Emotion Manipulation for Talking-Head Videos via Facial Landmarks
abstract
Manipulating the emotion of a performer in a video is a challenging task. The lip motion needs to be preserved while performing the desired changes in the emotion of the subject; however, simply utilizing existing image-based editing methods sabotages the original lip synchronization. We tackle this problem by utilizing a pretrained StyleGAN paired with a landmark-based editing module that modifies the bias present in the edit direction used in image manipulation. The proposed editing module consists of a latent-based landmark detection network and an editing network that modifies the editing direction to match the original lip synchronization while preserving the desired emotion manipulation results. This is realized by taking the facial landmarks as control points. Both networks operate on the latent space, which enables fast training and inference. We show that the proposed method runs significantly faster and performs better in terms of visual quality than alternative approaches, which was validated through a perceptual study. The proposed method can also be extended to perform face reenactment to generate a talking-head video from a single image and face image manipulation using facial landmarks as control points.
Kwanggyoon Seo, Rene Culaway, Byeong-Uk Lee, Jun-yong Noh
ACM Trans. Graph.1
2025 Neural Face Skinning for Mesh-agnostic Facial Expression Cloning
abstract
Abstract Accurately retargeting facial expressions to a face mesh while enabling manipulation is a key challenge in facial animation retargeting. Recent deep‐learning methods address this by encoding facial expressions into a global latent code, but they often fail to capture fine‐grained details in local regions. While some methods improve local accuracy by transferring deformations locally, this often complicates overall control of the facial expression. To address this, we propose a method that combines the strengths of both global and local deformation models. Our approach enables intuitive control and detailed expression cloning across diverse face meshes, regardless of their underlying structures. The core idea is to localize the influence of the global latent code on the target mesh. Our model learns to predict skinning weights for each vertex of the target face mesh through indirect supervision from predefined segmentation labels. These predicted weights localize the global latent code, enabling precise and region‐specific deformations even for meshes with unseen shapes. We supervise the latent code using Facial Action Coding System (FACS)‐based blendshapes to ensure interpretability and allow straightforward editing of the generated animation. Through extensive experiments, we demonstrate improved performance over state‐of‐the‐art methods in terms of expression fidelity, deformation transfer accuracy, and adaptability across diverse mesh structures.
Sihun Cha, Serin Yoon, Kwanggyoon Seo, Jun-yong Noh
Comput. Graph. Forum3
2025 Speed-Aware Audio-Driven Speech Animation using Adaptive Windows
abstract
We present a novel method that can generate realistic speech animations of a 3D face from audio using multiple adaptive windows. In contrast to previous studies that use a fixed size audio window, our method accepts an adaptive audio window as input, reflecting the audio speaking rate to use consistent phonemic information. Our system consists of three parts. First, the speaking rate is estimated from the input audio using a neural network trained in a self-supervised manner. Second, the appropriate window size that encloses the audio features is predicted adaptively based on the estimated speaking rate. Another key element lies in the use of multiple audio windows of different sizes as input to the animation generator: a small window to concentrate on detailed information and a large window to consider broad phonemic information near the center frame. Finally, the speech animation is generated from the multiple adaptive audio windows. Our method can generate realistic speech animations from in-the-wild audios at any speaking rate, i.e., fast raps, slow songs, as well as normal speech. We demonstrate via extensive quantitative and qualitative evaluations including a user study that our method outperforms state-of-the-art approaches.
Sunjin Jung, Yeongho Seol, Kwanggyoon Seo, Hyeonho Na, Seonghyeon Kim, Vanessa Tan, Jun-yong Noh
ACM Trans. Graph.3
2025 A Deep Learning-based Virtual Oculoplastic Surgery Simulator
abstract
Oculoplastic surgery is a critical treatment for various eye conditions, such as ptosis, which can cause both aesthetic and functional issues. Due to the anxiety about the outcome, patients are often hesitant to undergo the necessary procedures required for the surgery. Virtual oculoplastic surgery simulation technology offers a solution to alleviate these concerns by providing realistic previews of post-surgical results. In this paper, we present a novel deep learning-based virtual oculoplastic surgery simulation system that addresses the limitations of existing methods. The proposed system aims to improve the accuracy of simulations by considering the anatomical structure and characteristics of the eye. Our method utilizes a deformable parametric mesh to enhance the controllability of the image transformation process. Furthermore, the combination of a style-based generator and a neural texture has been implemented to generate high-quality results. The proposed system is expected to facilitate better communication between doctors and patients by providing anatomically inspired high-quality simulation results. The development of this advanced virtual simulation system has the potential to enhance patient experiences and improve satisfaction with outcomes in the field of oculoplastic surgery.
Seonghyeon Kim, Chang Wook Seo, Kwanggyoon Seo, Seung Han Song, Jun-yong Noh
ACM Trans. Graph.3
2024 StyleCineGAN: Landscape Cinemagraph Generation Using a Pre-trained StyleGAN
abstract
We propose a method that can generate cinemagraphs automatically from a still landscape image using a pre-trained StyleGAN. Inspired by the success of recent un-conditional video generation, we leverage a powerful pre-trained image generator to synthesize high-quality cinema-graphs. Unlike previous approaches that mainly utilize the latent space of a pre-trained StyleGAN, our approach utilizes its deep feature space for both GAN inversion and cin-emagraph generation. Specifically, we propose multi-scale deep feature warping (MSDFW), which warps the intermediate features of a pre-trained StyleGAN at different resolutions. by using MSDFW, the generated cinemagraphs are of high resolution and exhibit plausible looping animation. We demonstrate the superiority of our method through user studies and quantitative comparisons with state-of-the-art cinemagraph generation methods and a video generation method that uses a pre-trained StyleGAN.
Kwanggyoon Seo, Amirsaman Ashtari, Jun-yong Noh
CVPR2
2024 LeGO: Leveraging a Surface Deformation Network for Animatable Stylized Face Generation with One Example
abstract
Recent advances in 3D face stylization have made significant strides in few to zero-shot settings. However, the degree of stylization achieved by existing methods is often not sufficient for practical applications because they are mostly based on statistical 3D Morphable Models (3DMM) with limited variations. To this end, we propose a method that can produce a highly stylized 3D face model with desired topology. Our methods train a surface deformation network with 3DMM and translate its domain to the target style using a paired exemplar. The network achieves stylization of the 3D face mesh by mimicking the style of the target using a differentiable renderer and directional CLIP losses. Additionally, during the inference process, we utilize a Mesh Agnostic Encoder (MAGE) that takes deformation target, a mesh of diverse topologies as input to the stylization process and encodes its shape into our latent space. The resulting stylized face model can be animated by commonly used 3DMM blend shapes. A set of quantitative and qualitative evaluations demonstrate that our method can produce highly stylized face meshes according to a given style and output them in a desired topology. We also demonstrate example applications of our method including image-based stylized avatar generation, linear interpolation of geometric styles, and facial animation of stylized avatars.
Soyeon Yoon, Kwan Yun, Kwanggyoon Seo, Sihun Cha, Jung Eun Yoo, Jun-yong Noh
CVPR3
2024 Stylized Face Sketch Extraction via Generative Prior with Limited Data
abstract
Abstract Facial sketches are both a concise way of showing the identity of a person and a means to express artistic intention. While a few techniques have recently emerged that allow sketches to be extracted in different styles, they typically rely on a large amount of data that is difficult to obtain. Here, we propose StyleSketch, a method for extracting high‐resolution stylized sketches from a face image. Using the rich semantics of the deep features from a pretrained StyleGAN, we are able to train a sketch generator with 16 pairs of face and the corresponding sketch images. The sketch generator utilizes part‐based losses with two‐stage learning for fast convergence during training for high‐quality sketch extraction. Through a set of comparisons, we show that StyleSketch outperforms existing state‐of‐the‐art sketch extraction methods and few‐shot image adaptation methods for the task of extracting high‐resolution abstract face sketches. We further demonstrate the versatility of StyleSketch by extending its use to other domains and explore the possibility of semantic editing. The project page can be found in https://kwanyun.github.io/stylesketch_project .
Kwan Yun, Kwanggyoon Seo, Chang Wook Seo, Soyeon Yoon, Soohyun Ji, Amirsaman Ashtari, Jun-yong Noh
Comput. Graph. Forum2
2024 Real-Time CNN Training and Compression for Neural-Enhanced Adaptive Live Streaming
abstract
We propose a real-time convolutional neural network (CNN) training and compression method for delivering high-quality live video even in a poor network environment. The server delivers a low-resolution video segment along with the corresponding CNN for super resolution (SR), after which the client applies the CNN to the segment in order to recover high-resolution video frames. To generate a trained CNN corresponding to a video segment in real-time, our method rapidly increases the training accuracy by promoting the overfitting property of the CNN while also using curriculum-based training. In addition, assuming that the pretrained CNN is already downloaded on the client side, we transfer only residual values between the updated and pretrained CNN parameters. These values can be quantized with low bits in real time while minimizing the amount of loss, as the distribution range is significantly narrower than that of the updated CNN. Quantitatively, our neural-enhanced adaptive live streaming pipeline (NEALS) achieves higher SR accuracy and a lower CNN compression loss rate within a constrained training time compared to the state-of-the-art CNN training and compression method. NEALS achieves 15 to 48% higher quality of the user experience compared to state-of-the-art neural-enhanced live streaming systems.
Seunghwa Jeong, Bumki Kim, Seunghoon Cha, Kwanggyoon Seo, Hayoung Chang, Jungjin Lee, Younghui Kim, Jun-yong Noh
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Generating Texture for 3D Human Avatar from a Single Image using Sampling and Refinement Networks
abstract
Abstract There has been significant progress in generating an animatable 3D human avatar from a single image. However, recovering texture for the 3D human avatar from a single image has been relatively less addressed. Because the generated 3D human avatar reveals the occluded texture of the given image as it moves, it is critical to synthesize the occluded texture pattern that is unseen from the source image. To generate a plausible texture map for 3D human avatars, the occluded texture pattern needs to be synthesized with respect to the visible texture from the given image. Moreover, the generated texture should align with the surface of the target 3D mesh. In this paper, we propose a texture synthesis method for a 3D human avatar that incorporates geometry information. The proposed method consists of two convolutional networks for the sampling and refining process. The sampler network fills in the occluded regions of the source image and aligns the texture with the surface of the target 3D mesh using the geometry information. The sampled texture is further refined and adjusted by the refiner network. To maintain the clear details in the given image, both sampled and refined texture is blended to produce the final texture map. To effectively guide the sampler network to achieve its goal, we designed a curriculum learning scheme that starts from a simple sampling task and gradually progresses to the task where the alignment needs to be considered. We conducted experiments to show that our method outperforms previous methods qualitatively and quantitatively.
Sihun Cha, Kwanggyoon Seo, Amirsaman Ashtari, Jun-yong Noh
Comput. Graph. Forum2
2022 StylePortraitVideo: Editing Portrait Videos with Expression Optimization
abstract
Abstract High‐quality portrait image editing has been made easier by recent advances in GANs (e.g., StyleGAN) and GAN inversion methods that project images onto a pre‐trained GAN's latent space. However, extending the existing image editing methods, it is hard to edit videos to produce temporally coherent and natural‐looking videos. We find challenges in reproducing diverse video frames and preserving the natural motion after editing. In this work, we propose solutions for these challenges. First, we propose a video adaptation method that enables the generator to reconstruct the original input identity, unusual poses, and expressions in the video. Second, we propose an expression dynamics optimization that tweaks the latent codes to maintain the meaningful motion in the original video. Based on these methods, we build a StyleGAN‐based high‐quality portrait video editing system that can edit videos in the wild in a temporally coherent way at up to 4K resolution.
Kwanggyoon Seo, Seoung Wug Oh, Jingwan Lu, Joon-Young Lee, Seonghyeon Kim, Jun-yong Noh
Comput. Graph. Forum1
2021 Virtual Camera Layout Generation using a Reference Video
abstract
We propose a method that generates a virtual camera layout of a 3D animation scene by following the cinematic intention of a reference video. From a reference video, cinematic features such as the start frame, end frame, framing, camera movement, and the visual features of the subjects are extracted automatically. The extracted information is used to generate the virtual camera layout, which resembles the camera layout of the reference video. Our method handles stylized as well as human characters with body proportions different from those of humans. We demonstrate the effectiveness of our approach with various reference videos and 3D animation scenes. The user evaluation results show that the generated layouts are comparable to layouts created by the artist, allowing us to assert that our method can provide effective assistance to both novice and professional users when positioning a virtual camera.
Jung Eun Yoo, Kwanggyoon Seo, Sanghun Park, Jaedong Kim, Dawon Lee, Jun-yong Noh
CHI2
2021 Deep Learning-Based Unsupervised Human Facial Retargeting
abstract
Abstract Traditional approaches to retarget existing facial blendshape animations to other characters rely heavily on manually paired data including corresponding anchors, expressions, or semantic parametrizations to preserve the characteristics of the original performance. In this paper, inspired by recent developments in face swapping and reenactment, we propose a novel unsupervised learning method that reformulates the retargeting of 3D facial blendshape‐based animations in the image domain. The expressions of a source model is transferred to a target model via the rendered images of the source animation. For this purpose, a reenactment network is trained with the rendered images of various expressions created by the source and target models in a shared latent space. The use of shared latent space enable an automatic cross‐mapping obviating the need for manual pairing. Next, a blendshape prediction network is used to extract the blendshape weights from the translated image to complete the retargeting of the animation onto a 3D target model. Our method allows for fully unsupervised retargeting of facial expressions between models of different configurations, and once trained, is suitable for automatic real‐time applications.
Seonghyeon Kim, Sunjin Jung, Kwanggyoon Seo, Roger Blanco Ribera, Jun-yong Noh
Comput. Graph. Forum3
2020 Neural crossbreed: neural based image metamorphosis
abstract
We propose Neural Crossbreed, a feed-forward neural network that can learn a semantic change of input images in a latent space to create the morphing effect. Because the network learns a semantic change, a sequence of meaningful intermediate images can be generated without requiring the user to specify explicit correspondences. In addition, the semantic change learning makes it possible to perform the morphing between the images that contain objects with significantly different poses or camera views. Furthermore, just as in conventional morphing techniques, our morphing network can handle shape and appearance transitions separately by disentangling the content and the style transfer for rich usability. We prepare a training dataset for morphing using a pre-trained BigGAN, which generates an intermediate image by interpolating two latent vectors at an intended morphing value. This is the first attempt to address image morphing using a pre-trained generative model in order to learn semantic transformation. The experiments show that Neural Crossbreed produces high quality morphed images, overcoming various limitations associated with conventional approaches. In addition, Neural Crossbreed can be further extended for diverse applications such as multi-image morphing, appearance transfer, and video frame interpolation.
Sanghun Park, Kwanggyoon Seo, Jun-yong Noh
ACM Trans. Graph.2