VLDB 2026 Research / reviewers in the wild / expert
Seonghyeon Kim
dblp:217/9233
· DBLP profile ↗
15ranked-venue papers
3as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Visualizing and Securing Linguistic Patterns: POS Tag Graphs with Encryption TechniquesabstractDigital Content protection is important Ensuring document integrity in the digital age is a critical challenge. Traditional visible and invisible watermarking methods offer partial protection but have limitations. Visible watermarking can be removed and degrade quality, while invisible watermarking requires specialized tools and may degrade through compression. TThis study presents an innovative document authentication framework that leverages part-of-speech (POS) tag-based graph representations. By transforming textual data into structured graphs, the method captures the unique syntactic signature of each document. This linguistic fingerprint is then encrypted and stored independently, enabling robust verification of document authenticity. The approach not only reinforces security and integrity but also seamlessly integrates with existing natural language processing (NLP) pipelines. Seonghyeon Kim, Lei Chen 0029, Jongyeop Kim, Jongho Seol |
SERA | 1 |
| 2025 | Deep Learning Approaches for Credit Card Fraud Detection: A Data Balancing PerspectiveabstractCredit card fraud detection is a critical challenge in modern financial systems, requiring robust and efficient solutions to identify suspicious transactions accurately. This study addresses the issue by utilizing machine learning techniques to detect potentially fraudulent transactions. A key focus of the research is the handling of imbalanced datasets, where genuine transactions vastly outnumber fraudulent ones. To address this imbalance, we increased the sample size of fraudulent data using oversampling techniques and subsequently applied four distinct machine learning models to assess their performance. Through iterative experimentation, we identified the optimal magnification ratio that enhances the model’s ability to distinguish between legitimate and fraudulent transactions. The results demonstrate that balancing the dataset significantly improves detection accuracy, providing insights into effective model configurations for realworld applications. This research contributes to the development of more reliable and efficient fraud detection systems in the financial sector. Jongyeop Kim, Jongho Seol, Seonghyeon Kim, Lei Chen 0029 |
SERA | 3 |
| 2025 | Deep-Learning-Based Facial Retargeting Using Local PatchesabstractAbstract In the era of digital animation, the quest to produce lifelike facial animations for virtual characters has led to the development of various retargeting methods. While the retargeting facial motion between models of similar shapes has been very successful, challenges arise when the retargeting is performed on stylized or exaggerated 3D characters that deviate significantly from human facial structures. In this scenario, it is important to consider the target character's facial structure and possible range of motion to preserve the semantics assumed by the original facial motions after the retargeting. To achieve this, we propose a local patch‐based retargeting method that transfers facial animations captured in a source performance video to a target stylized 3D character. Our method consists of three modules. The Automatic Patch Extraction Module extracts local patches from the source video frame. These patches are processed through the Reenactment Module to generate correspondingly re‐enacted target local patches. The Weight Estimation Module calculates the animation parameters for the target character at every frame for the creation of a complete facial animation sequence. Extensive experiments demonstrate that our method can successfully transfer the semantic meaning of source facial expressions to stylized characters with considerable variations in facial feature proportion. Yeonsoo Choi, Inyup Lee, Sihun Cha, Seonghyeon Kim, Sunjin Jung, Jun-yong Noh |
Comput. Graph. Forum | 4 |
| 2025 | Speed-Aware Audio-Driven Speech Animation using Adaptive WindowsabstractWe present a novel method that can generate realistic speech animations of a 3D face from audio using multiple adaptive windows. In contrast to previous studies that use a fixed size audio window, our method accepts an adaptive audio window as input, reflecting the audio speaking rate to use consistent phonemic information. Our system consists of three parts. First, the speaking rate is estimated from the input audio using a neural network trained in a self-supervised manner. Second, the appropriate window size that encloses the audio features is predicted adaptively based on the estimated speaking rate. Another key element lies in the use of multiple audio windows of different sizes as input to the animation generator: a small window to concentrate on detailed information and a large window to consider broad phonemic information near the center frame. Finally, the speech animation is generated from the multiple adaptive audio windows. Our method can generate realistic speech animations from in-the-wild audios at any speaking rate, i.e., fast raps, slow songs, as well as normal speech. We demonstrate via extensive quantitative and qualitative evaluations including a user study that our method outperforms state-of-the-art approaches. Sunjin Jung, Yeongho Seol, Kwanggyoon Seo, Hyeonho Na, Seonghyeon Kim, Vanessa Tan, Jun-yong Noh |
ACM Trans. Graph. | 5 |
| 2025 | A Deep Learning-based Virtual Oculoplastic Surgery SimulatorabstractOculoplastic surgery is a critical treatment for various eye conditions, such as ptosis, which can cause both aesthetic and functional issues. Due to the anxiety about the outcome, patients are often hesitant to undergo the necessary procedures required for the surgery. Virtual oculoplastic surgery simulation technology offers a solution to alleviate these concerns by providing realistic previews of post-surgical results. In this paper, we present a novel deep learning-based virtual oculoplastic surgery simulation system that addresses the limitations of existing methods. The proposed system aims to improve the accuracy of simulations by considering the anatomical structure and characteristics of the eye. Our method utilizes a deformable parametric mesh to enhance the controllability of the image transformation process. Furthermore, the combination of a style-based generator and a neural texture has been implemented to generate high-quality results. The proposed system is expected to facilitate better communication between doctors and patients by providing anatomically inspired high-quality simulation results. The development of this advanced virtual simulation system has the potential to enhance patient experiences and improve satisfaction with outcomes in the field of oculoplastic surgery. Seonghyeon Kim, Chang Wook Seo, Kwanggyoon Seo, Seung Han Song, Jun-yong Noh |
ACM Trans. Graph. | 1 |
| 2024 | Learning to Explore for Stochastic Gradient MCMCabstractBayesian Neural Networks(BNNs) with high-dimensional parameters pose a challenge for posterior inference due to the multi-modality of the posterior distributions. Stochastic Gradient Markov Chain Monte Carlo(SGMCMC) with cyclical learning rate scheduling is a promising solution, but it requires a large number of sampling steps to explore high-dimensional multi-modal posteriors, making it computationally expensive. In this paper, we propose a meta-learning strategy to build SGMCMC which can efficiently explore the multi-modal target distributions. Our algorithm allows the learned SGMCMC to quickly explore the high-density region of the posterior landscape. Also, we show that this exploration property is transferrable to various tasks, even for the ones unseen during a meta-training stage. Using popular image classification benchmarks and a variety of downstream tasks, we demonstrate that our method significantly improves the sampling efficiency, achieving better performance than vanilla SGMCMC without incurring significant computational overhead. Seohyeon Jung, Seonghyeon Kim, Juho Lee 0001 |
ICML | 3 |
| 2023 | Towards Unified Scene Text Spotting Based on Sequence GenerationabstractSequence generation models have recently made significant progress in unifying various vision tasks. Although some auto-regressive models have demonstrated promising results in end-to-end text spotting, they use specific detection formats while ignoring various text shapes and are limited in the maximum number of text instances that can be detected. To overcome these limitations, we propose a UNIfied scene Text Spotter, called UNITS. Our model unifies various detection formats, including quadrilaterals and polygons, allowing it to detect text in arbitrary shapes. Additionally, we apply starting-point prompting to enable the model to extract texts from an arbitrary starting point, thereby extracting more texts beyond the number of instances it was trained on. Experimental results demonstrate that our method achieves competitive performance compared to state-of-the-art methods. Further analysis shows that UNITS can extract a larger number of texts than it was trained on. We provide the code for our method at https://github.com/clovaai/units. Taeho Kil, Seonghyeon Kim, Sukmin Seo, Yoonsik Kim |
CVPR | 2 |
| 2023 | TRACE: Table Reconstruction Aligned to Corner and Edges
Youngmin Baek, Daehyun Nam, Jaeheung Surh, Seung Shin, Seonghyeon Kim |
ICDAR (5) | 5 |
| 2022 | Generator Knows What Discriminator Should Learn in Unconditional GANs
Gayoung Lee, Hyunsu Kim, Seonghyeon Kim, Jung-Woo Ha 0001, Yunjey Choi |
ECCV (17) | 4 |
| 2022 | StylePortraitVideo: Editing Portrait Videos with Expression OptimizationabstractAbstract High‐quality portrait image editing has been made easier by recent advances in GANs (e.g., StyleGAN) and GAN inversion methods that project images onto a pre‐trained GAN's latent space. However, extending the existing image editing methods, it is hard to edit videos to produce temporally coherent and natural‐looking videos. We find challenges in reproducing diverse video frames and preserving the natural motion after editing. In this work, we propose solutions for these challenges. First, we propose a video adaptation method that enables the generator to reconstruct the original input identity, unusual poses, and expressions in the video. Second, we propose an expression dynamics optimization that tweaks the latent codes to maintain the meaningful motion in the original video. Based on these methods, we build a StyleGAN‐based high‐quality portrait video editing system that can edit videos in the wild in a temporally coherent way at up to 4K resolution. Kwanggyoon Seo, Seoung Wug Oh, Jingwan Lu, Joon-Young Lee, Seonghyeon Kim, Jun-yong Noh |
Comput. Graph. Forum | 5 |
| 2021 | Deep Learning-Based Unsupervised Human Facial RetargetingabstractAbstract Traditional approaches to retarget existing facial blendshape animations to other characters rely heavily on manually paired data including corresponding anchors, expressions, or semantic parametrizations to preserve the characteristics of the original performance. In this paper, inspired by recent developments in face swapping and reenactment, we propose a novel unsupervised learning method that reformulates the retargeting of 3D facial blendshape‐based animations in the image domain. The expressions of a source model is transferred to a target model via the rendered images of the source animation. For this purpose, a reenactment network is trained with the rendered images of various expressions created by the source and target models in a shared latent space. The use of shared latent space enable an automatic cross‐mapping obviating the need for manual pairing. Next, a blendshape prediction network is used to extract the blendshape weights from the translated image to complete the retargeting of the animation onto a 3D target model. Our method allows for fully unsupervised retargeting of facial expressions between models of different configurations, and once trained, is suitable for automatic real‐time applications. Seonghyeon Kim, Sunjin Jung, Kwanggyoon Seo, Roger Blanco Ribera, Jun-yong Noh |
Comput. Graph. Forum | 1 |
| 2021 | Effect of virtual error sensor location for active sound quality control in a car cabinabstractSummary A virtual error microphone‐based active sound quality control (VEM‐ASQC) algorithm is proposed to make engine sound quality better inside a car cabin in this study. In a single‐input and single‐output system, if a VEM is installed at either left or right ear, the sound pressure level (SPL) at the other ear can be fluctuated, and this leads the excessive SPL difference between the two ears when previous methods were applied. In this paper, an uncomplicated method that can make engine sounds at both left and right ears track the target engine sound levels with less errors and reduces the left‐right tracking error (TE) (in SPL) difference was investigated. The suggested VEM‐ASQC algorithm was implemented in an actual car for real‐time control experiment, and target profiles of nine engine orders were designed to control the sound quality actively. The control experiment showed that the proposed approach reduces the errors to track the target levels and the left‐right TE differences more compared to those by the previous methods. Seokhoon Ryu, Seonghyeon Kim, Young-Sup Lee |
Concurr. Comput. Pract. Exp. | 2 |
| 2020 | Understanding How People Reason about Aesthetic Evaluations of Artificial IntelligenceabstractArtificial intelligence (AI) algorithms are making remarkable achievements even in creative fields such as aesthetics. However, whether those outside the machine learning (ML) community can sufficiently interpret or agree with their results, especially in such highly subjective domains, is being questioned. In this paper, we try to understand how different user communities reason about AI algorithm results in subjective domains. We designed AI Mirror, a research probe that tells users the algorithmically predicted aesthetic scores of photographs. We conducted a user study of the system with 18 participants from three different groups: AI/ML experts, domain experts (photographers), and general public members. They performed tasks consisting of taking photos and reasoning about AI Mirror's prediction algorithm with think-aloud sessions, surveys, and interviews. The results showed the following: (1) Users understood the AI using their own group-specific expertise; (2) Users employed various strategies to close the gap between their judgments and AI predictions overtime; (3) The difference between users' thoughts and AI pre-dictions was negatively related with users' perceptions of the AI's interpretability and reasonability. We also discuss design considerations for AI-infused systems in subjective domains. Changhoon Oh, Seonghyeon Kim, Jinhan Choi, Jinsu Eun, Soomin Kim 0001, Juho Kim 0001, Joonhwan Lee, Bongwon Suh |
Conference on Designing Interactive Systems | 2 |
| 2020 | Few-Shot Compositional Font Generation with Dual Memory
Junbum Cha, Sanghyuk Chun, Gayoung Lee, Bado Lee, Seonghyeon Kim, Hwalsuk Lee |
ECCV (19) | 5 |
| 2018 | I Lead, You Help but Only with Enough Details: Understanding User Experience of Co-Creation with Artificial IntelligenceabstractRecent advances in artificial intelligence (AI) have increased the opportunities for users to interact with the technology. Now, users can even collaborate with AI in creative activities such as art. To understand the user experience in this new user--AI collaboration, we designed a prototype, DuetDraw, an AI interface that allows users and the AI agent to draw pictures collaboratively. We conducted a user study employing both quantitative and qualitative methods. Thirty participants performed a series of drawing tasks with the think-aloud method, followed by post-hoc surveys and interviews. Our findings are as follows: (1) Users were significantly more content with DuetDraw when the tool gave detailed instructions. (2) While users always wanted to lead the task, they also wanted the AI to explain its intentions but only when the users wanted it to do so. (3) Although users rated the AI relatively low in predictability, controllability, and comprehensibility, they enjoyed their interactions with it during the task. Based on these findings, we discuss implications for user interfaces where users can collaborate with AI in creative works. Changhoon Oh, Jungwoo Song, Jinhan Choi, Seonghyeon Kim, Sungwoo Lee, Bongwon Suh |
CHI | 4 |