VLDB 2026 Research / reviewers in the wild / expert
Yoonsik Kim
dblp:194/2556
· DBLP profile ↗
16ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Toward an Autonomous Purple Teaming Framework for Security and Safety in Large Language ModelsabstractLarge Language Models (LLMs) have rapidly advanced in reasoning capability and accessibility, driving their deployment across diverse applications. Yet this progress has also widened the surface for safety and security vulnerabilities. Adversaries can exploit prompt diversity, dialog memory, or multimodal inputs to induce unsafe or confidential outputs, while continual fine-tuning and third-party integration render static assurance infeasible. This paper introduces our ongoing national R&D project on developing the AutoPT Framework-an Autonomous Purple Teaming architecture that extends the collaborative principles of purple teaming toward self-adaptive, continuously verifiable LLM assurance. AutoPT unifies autonomous adversarial exploration and adaptive defensive reinforcement through two co-evolving agents. The red module, AutoPT-Red, employs coverage-guided fuzzing and internal measurement metrics to autonomously uncover vulnerabilities. The blue module, AutoPT-Blue, performs self-healing adaptation by updating guardrails and detecting integrity or confidentiality violations using embedding-based feedback. Preliminary case studies on jailbreak fuzzing and backdoor-poisoning defense validate the feasibility of this closed-loop, self-adapting architecture. As part of a broader national initiative, this work lays the conceptual and technical foundation for transitioning industrial purple teaming into a fully autonomous, scalable, and measurable assurance paradigm for generative AI systems. Leo Hyun Park, Yoonsik Kim, Eunbi Hwang, Sangsoo Han, Hyoungshick Kim, Taekyoung Kwon 0002 |
PRDC | 2 |
| 2025 | Continuous Authentication for Secure and Seamless User-Avatar Integration in Multidevice MetaversesabstractThe metaverse connects the virtual and real worlds, enabling users to interact as avatars across multiple devices, including smartphones, HMDs, and other devices. While multi-device access enhances convenience, it also expands attack surfaces, increasing security risks. Continuous authentication is crucial, but traditional methods like fuzzy extractors struggle with dynamic data, making reliable identification difficult. Moreover, conventional authentication focuses on user-side verification, failing to detect avatar manipulation attacks like avatar hijacking. This paper proposes a continuous authentication system that integrates user and avatar behavior data in multi-device environments. A transformer-based embedding model processes data on edge devices and securely transmits it via JSON Web Tokens (JWT). The authentication model binds user and avatar data in real-time to compute confidence scores and detect avatar manipulation. We implemented a VRSpace-based metaverse on NGINX and conducted simulations using open datasets—HMOG, Liebers, and BOXRR—to evaluate authentication accuracy and continuity. The Smartphone+HMD_6DOF+Avt_act+Window_(20) model achieved an average FAR of 0.0034%, an EER of 0.3386%, and an ADR of up to 97.67% for avatar manipulation detection. Based on our work, industry-driven research is expected to explore real-world applications, further validating our approach in evolving multi-device metaverse ecosystems. Eunbi Hwang, Yoonsik Kim, Taekyoung Kwon 0002 |
IEEE Internet Things J. | 2 |
| 2025 | A Continuous Authentication Framework for Securing Metaverse IdentitiesabstractIn the Metaverse, continuous authentication is essential for verifying the ongoing connection between a user’s physical identity and avatar, ensuring secure access to various services. This process is crucial for confirming identities, maintaining security, and preventing unauthorized activities that could compromise legitimate services. However, traditional biometric-based authentication methods are susceptible to threats such as impersonation, replay attacks, and disguise, primarily due to the difficulty in directly using biometric information to represent the connection between virtual and physical identities. To address these challenges, some studies have proposed using blockchain schemes to mitigate security threats. Despite this, these approaches often encounter issues like insufficient network protection for authentication connections, prolonged data processing times, and latency. To overcome these limitations, we propose a secure continuous authentication framework that leverages standard protocols such as QUIC and JWT to verify user identities efficiently. Our approach employs embedding models on edge devices to generate and transmit biometric data. In contrast, a deep learning-based model on the server validates the user’s credentials, ensuring both high performance and availability. Experimental results show that our QUIC and JWT-based protocol delivers superior security and effectiveness compared to traditional biometric approaches and blockchain-based methods, achieving an AUC of 0.97, an EER of 3.77, and an F1 score of 0.96. Sangsoo Han, Eunbi Hwang, Yoonsik Kim, Taekyoung Kwon 0002 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Towards Unified Scene Text Spotting Based on Sequence GenerationabstractSequence generation models have recently made significant progress in unifying various vision tasks. Although some auto-regressive models have demonstrated promising results in end-to-end text spotting, they use specific detection formats while ignoring various text shapes and are limited in the maximum number of text instances that can be detected. To overcome these limitations, we propose a UNIfied scene Text Spotter, called UNITS. Our model unifies various detection formats, including quadrilaterals and polygons, allowing it to detect text in arbitrary shapes. Additionally, we apply starting-point prompting to enable the model to extract texts from an arbitrary starting point, thereby extracting more texts beyond the number of instances it was trained on. Experimental results demonstrate that our method achieves competitive performance compared to state-of-the-art methods. Further analysis shows that UNITS can extract a larger number of texts than it was trained on. We provide the code for our method at https://github.com/clovaai/units. Taeho Kil, Seonghyeon Kim, Sukmin Seo, Yoonsik Kim |
CVPR | 4 |
| 2023 | Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language ModelsabstractGeewook Kim, Hodong Lee, Daehee Kim, Haeji Jung, Sanghee Park, Yoonsik Kim, Sangdoo Yun, Taeho Kil, Bado Lee, Seunghyun Park. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Geewook Kim, Hodong Lee, Daehee Kim 0003, Haeji Jung, Sanghee Park, Yoonsik Kim, Sangdoo Yun, Taeho Kil, Bado Lee, Seunghyun Park 0001 |
EMNLP | 6 |
| 2023 | SCOB: Universal Text Understanding via Character-wise Supervised Contrastive Learning with Online Text Rendering for Bridging Domain GapabstractInspired by the great success of language model (LM)-based pre-training, recent studies in visual document understanding have explored LM-based pre-training methods for modeling text within document images. Among them, pre-training that reads all text from an image has shown promise, but often exhibits instability and even fails when applied to broader domains, such as those involving both visual documents and scene text images. This is a substantial limitation for real-world scenarios, where the processing of text image inputs in diverse domains is essential. In this paper, we investigate effective pre-training tasks in the broader domains and also propose a novel pre-training method called SCOB that leverages character-wise supervised contrastive learning with online text rendering to effectively pre-train document and scene text domains by bridging the domain gap. Moreover, SCOB enables weakly supervised learning, significantly reducing annotation costs. Extensive benchmarks demonstrate that SCOB generally improves vanilla pre-training methods and achieves comparable performance to state-of-the-art methods. Our findings suggest that SCOB can be served generally and effectively for read-type pre-training methods. The code will be available at https://github.com/naver-ai/scob. Daehee Kim 0003, Yoonsik Kim, Donghyun Kim 0012, Yumin Lim, Geewook Kim, Taeho Kil |
ICCV | 2 |
| 2023 | On Web-based Visual Corpus Construction for Visual Document Understanding
Donghyun Kim 0012, Teakgyu Hong, Moonbin Yim, Yoonsik Kim, Geewook Kim |
ICDAR (3) | 4 |
| 2022 | Multi-modal Text Recognition Networks: Interactive Enhancements Between Visual and Semantic Features
Byeonghu Na, Yoonsik Kim, Sungrae Park |
ECCV (28) | 2 |
| 2021 | SynthTIGER: Synthetic Text Image GEneratoR Towards Better Text Recognition Models
Moonbin Yim, Yoonsik Kim, Hancheol Cho, Sungrae Park |
ICDAR (4) | 2 |
| 2020 | Transfer Learning From Synthetic to Real-Noise Denoising With Adaptive Instance NormalizationabstractReal-noise denoising is a challenging task because the statistics of real-noise do not follow the normal distribution, and they are also spatially and temporally changing. In order to cope with various and complex real-noise, we propose a well-generalized denoising architecture and a transfer learning scheme. Specifically, we adopt an adaptive instance normalization to build a denoiser, which can regularize the feature map and prevent the network from overfitting to the training set. We also introduce a transfer learning scheme that transfers knowledge learned from synthetic-noise data to the real-noise denoiser. From the proposed transfer learning, the synthetic-noise denoiser can learn general features from various synthetic-noise data, and the real-noise denoiser can learn the real-noise characteristics from real data. From the experiments, we find that the proposed denoising method has great generalization ability, such that our network trained with synthetic-noise achieves the best performance for Darmstadt Noise Dataset (DND) among the methods from published papers. We can also see that the proposed transfer learning scheme robustly works for real-noise images through the learning with a very small number of labeled data. Yoonsik Kim, Jae Woong Soh, Gu Yong Park, Nam Ik Cho |
CVPR | 1 |
| 2020 | Dual Path Denoising Network for Real Photographic NoiseabstractThis letter presents a convolutional neural network (CNN) for image denoising, especially for the reduction of real noises. As a network topology, we adopt the dual path network (DPN) that combines the advantages of residual and densely connected networks. Using the DPN as a basic building block, we design a network that connects the DPN in dual path again with an attention mechanism. For efficient denoising of real noise images, we build a training set where noisy images are obtained from a heteroscedastic Gaussian noise model and in-camera pipeline. In addition, we augment the synthetic training set with a relatively small number of real noise data. In the experiments, the proposed method is shown to provide state-of-the-art performance in reducing both synthetic and real noises. Yeong Il Jang, Yoonsik Kim, Nam Ik Cho |
IEEE Signal Process. Lett. | 2 |
| 2020 | A Pseudo-Blind Convolutional Neural Network for the Reduction of Compression ArtifactsabstractThis paper presents methods based on convolutional neural networks (CNNs) for removing compression artifacts. We modify the Inception module for the image restoration problem and use it as a building block for constructing blind and non-blind artifact removal networks. It is known that a CNN trained in a non-blind scenario (known compression quality factor) performs better than the one trained in a blind scenario (unknown factor), and our network is not an exception. However, the blind system is more practical because the compression quality factor is not always available or does not reflect the actual quality when the image is a transcoded or requantized image. Hence, in this paper, we also propose a pseudo-blind system that estimates the quality factor for a given compressed image and then applies a network that is trained with a similar quality factor. For this purpose, we propose a CNN that estimates the compression quality factor and prepare several non-blind artifact removal networks that are trained for some specific compression quality factors. We train the networks and conduct experiments on widely used compression standards, such as JPEG, MPEG-2, H.264, and HEVC. In addition, we conduct experiments for dynamically changing and transcoded videos to demonstrate the effectiveness of the quality estimation method. The experimental results show that the proposed pseudo-blind network performs better than the blind one for the various cases stated above and requires fewer computations. Yoonsik Kim, Jae Woong Soh, Jaewoo Park 0005, Byeongyong Ahn, Hyun-Seung Lee 0001, Young-Su Moon, Nam Ik Cho |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Adaptively Tuning a Convolutional Neural Network by Gating Process for Image DenoisingabstractThis paper presents a new framework that controls feature maps of a convolutional neural network (CNN) according to the noise level such that the network can have different properties to different levels. Unlike the conventional non-blind approach which reloads all the parameters of CNN or switches to other CNNs for different noise levels, we adjust the CNN activation feature maps without changing the parameters at the test phase. For this, we additionally construct a noise level indicator network, which gives appropriate weighting values to the feature maps for the given situation. The noise level indicator network is so simple that it can be implemented as a low-dimensional look-up table at the test phase and thus does not increase the overall complexity. From the experiments on noise reduction, we can observe that the proposed method achieves better performance compared to the baseline network. Yoonsik Kim, Jae Woong Soh, Nam Ik Cho |
ICIP | 1 |
| 2019 | Pupil Center Detection Based on the UNet for the User Interaction in VR and AR EnvironmentsabstractFinding the location of a pupil center is important for the human-computer interaction especially for the user interface in AR/VR devices. In this paper, we propose an indirect use of the convolutional neural network (CNN) for the task, which first segments the pupil region by a CNN, and then finds the center of mass of the region. For this, we create a dataset by labeling the pupil area on 111,581 images from 29 IR video sequences. We also label the pupil region of widely used datasets to test and validate our method on a variety of inputs. Experiments show that the proposed method provides better accuracies than the conventional ones, showing robustness to the noise. Sang Yoon Han, Yoonsik Kim, Sang Hwa Lee, Nam Ik Cho |
VR | 2 |
| 2017 | Skin detection based on multi-seed propagation in a multi-layer graph for regional and color consistencyabstractWe propose a new skin detection method based on multi-seeds propagation in a multi-layer graph representation of an image. Initially, some of nodes in the graph are set to be foreground or background seeds based on a simple Bayesian skin detector, and they are propagated through the graph to find the skin probability in the manner of semi-supervised learning. The graph is designed to consider not only local and global coherence but also to consider the color consistency by constructing a multilayer graph of image and cluster layers. Extensive experiments on several datasets are conducted, which demonstrate that our method outperforms the existing methods in terms of various quantitative measures, such as accuracy, precision, recall and F-measure. Insung Hwang, Yoonsik Kim, Nam Ik Cho |
ICASSP | 2 |
| 2017 | Convolutional neural networks and training strategies for skin detectionabstractThis paper presents two convolutional neural networks (CNN) and their training strategies for skin detection. The first CNN, consisting of 20 convolution layers with 3 × 3 filters, is a kind of VGG network. The second is composed of 20 networkin-network (NiN) layers which can be considered a modification of Inception structure. When training these networks for human skin detection, we consider patch-based and whole-image-based training. The first method focuses on local features such as skin color and texture, and the second on the human-related shape features as well as color and texture. Experiments show that the proposed CNNs yield better performance than the conventional methods and also than the existing deep-learning based method. Also, it is found that the NiN structure generally shows higher accuracy than the VGG-based structure. The experiments also show that the whole-image-based training that learns the shape features yields better accuracy than the patch-based learning that focuses on local color and texture only. Yoonsik Kim, Insung Hwang, Nam Ik Cho |
ICIP | 1 |