VLDB 2026 Research / reviewers in the wild / expert
Kyeongbo Kong
dblp:218/1547
· DBLP profile ↗
23ranked-venue papers
2as first author
22since 2021 · last 2026
0000-0002-1135-7502ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 19 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Group-Wise Layer Aggregation of Self-Supervised Representations for Speech Emotion Recognition
Sunchan Park, Hyung Soon Kim, Kyeongbo Kong |
IEEE Signal Process. Lett. | 3 |
| 2026 | Motion-to-Attention: Enhancing Attention Maps to Improve Performance of Text-Guided Video EditingabstractRecent research in text-guided video editing aims to extend image-based editing models to video domains. A significant challenge in this transition is ensuring temporal consistency across frames. However, existing methods often exhibit limited editing accuracy when processing prompts associated with motion, such as “ floating ” or “ moving .” Our analysis indicates that this limitation arises from inaccurate attention maps corresponding to motion-related prompts. To address this, we introduce the Motion-to-Attention (M2A) module, explicitly integrating motion information for enhanced video editing precision. Specifically, we first convert optical flow extracted from the video into a comprehensive motion map. Optionally, users can specify directional information to refine motion map extraction further. The proposed M2A module incorporates two complementary techniques: “ Attention–Motion Swap ,” which directly substitutes the imprecise attention map of motion prompts with the extracted motion map, and “ Attention–Motion Fusion ,” which adaptively enhances attention maps based on the correlation with the motion map using a carefully selected Fusion metric. Experimental validation demonstrates that incorporating our M2A module into existing text-to-video editing frameworks significantly improves both quantitative performance metrics (CLIP-Acc, Masked PSNR, BRISQUE) and qualitative visual quality. Extensive experiments and comparative studies confirm the superior editability and robustness of our method over current state-of-the-art approaches. Comprehensive results are publicly available at https://currycurry915.github.io/Motion-to-Attention/ . Seong-Hun Jeong, In-Hwan Jin, Haesoo Choo, Hyeonjun Na, Kyeongbo Kong |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Optimizing 4D Gaussians for Dynamic Scene Video from Single Landscape ImagesabstractTo achieve realistic immersion in landscape images, fluids such as water and clouds need to move within the image while revealing new scenes from various camera perspectives. Recently, a field called dynamic scene video has emerged, which combines single image animation with 3D photography. These methods use pseudo 3D space, implicitly represented with Layered Depth Images (LDIs). LDIs separate a single image into depth-based layers, which enables elements like water and clouds to move within the image while revealing new scenes from different camera perspectives. However, as landscapes typically consist of continuous elements, including fluids, the representation of a 3D space separates a landscape image into discrete layers, and it can lead to diminished depth perception and potential distortions depending on camera movement. Furthermore, due to its implicit modeling of 3D space, the output may be limited to videos in the 2D domain, potentially reducing their versatility. In this paper, we propose representing a complete 3D space for dynamic scene video by modeling explicit representations, specifically 4D Gaussians, from a single image. The framework is focused on optimizing 3D Gaussians by generating multi-view images from a single image and creating 3D motion to optimize 4D Gaussians. The most important part of proposed framework is consistent 3D motion estimation, which estimates common motion among multi-view images to bring the motion in 3D space closer to actual motions. As far as we know, this is the first attempt that considers animation while representing a complete 3D space from a single landscape image. Our model demonstrates the ability to provide realistic immersion in various landscape images through diverse experiments and metrics. Extensive experimental results are https://cvsp-lab.github.io/ICLR2025_3D-MOM/. In-Hwan Jin, Haesoo Choo, Seong-Hun Jeong, Park Heemoon, Oh-joon Kwon, Kyeongbo Kong |
ICLR | 7 |
| 2025 | Unsupervised Knowledge Distillation via Local Representations for Vision-Language ModelsabstractRecent vision-language models (VLMs) have adopted prompt learning for downstream adaptation, yet they often suffer from poor generalization and require labeled data. Unsupervised knowledge distillation (UKD) offers a promising alternative, but existing methods primarily distill global predictions, neglecting the potential of local representations for fine-grained recognition. Interestingly, we observe that large VLMs such as CLIP-L/14 yield strong global features but produce weak and noisy local tokens. To address this, we propose a two-stage UKD framework that introduces an assistant model—comparable in size to the student—to refine and transfer reliable local cues under teacher supervision. Extensive experiments on 11 benchmarks demonstrate that our method consistently outperforms global-only distillation approaches in domain generalization and unseen class recognition tasks. Changwoo Baek, Jou Won Song, Kyeongbo Kong |
IEEE Signal Process. Lett. | 3 |
| 2025 | Programmable-Room: Interactive Textured 3D Room Meshes Generation Empowered by Large Language Models
Kyeongbo Kong, Suk-Ju Kang |
IEEE Trans. Multim. | 3 |
| 2024 | Person in Place: Generating Associative Skeleton-Guidance Maps for Human-Object Interaction Image EditingabstractRecently, there were remarkable advances in image editing tasks in various ways. Nevertheless, existing image editing models are not designed for Human-Object Interaction (HOI) image editing. One of these approaches (e.g. ControlNet) employs the skeleton guidance to offer precise representations of human, showing better results in HOI image editing. However, using conventional methods, manually creating HOI skeleton guidance is necessary. This paper proposes the object interactive diffuser with associative attention that considers both the interaction with objects and the joint graph structure, automating the generation of HOI skeleton guidance. Additionally, we propose the HOI loss with novel scaling parameter, demonstrating its effectiveness in generating skeletons that interact better. To evaluate generated object-interactive skeletons, we propose two metrics, top-N accuracy and skeleton probabilistic distance. Our framework integrates object interactive diffuser thatgenerates object-interactive skeletons with previous methods, demonstrating the outstanding results in HOI image editing. Finally, we present potentials of our framework beyond HOI image editing, as applications to human-to-human interaction, skeleton editing, and 3D mesh optimization. The code is available at https://github.com/YangChangHee/CVPR2024_Person-In-Place_RELEASE ChangHee Yang, Chanhee Kang, Kyeongbo Kong, Hanni Oh, Suk-Ju Kang |
CVPR | 3 |
| 2024 | AttentionHand: Text-Driven Controllable Hand Image Generation for 3D Hand Reconstruction in the Wild
Kyeongbo Kong, Suk-Ju Kang |
ECCV (60) | 2 |
| 2024 | Embedding-Free Transformer with Inference Spatial Reduction for Efficient Semantic Segmentation
Hyunwoo Yu, Yubin Cho, Beoungwoo Kang, Seunghun Moon, Kyeongbo Kong, Suk-Ju Kang |
ECCV (42) | 5 |
| 2024 | Human Motion Aware Text-to-Video Generation with Explicit Camera ControlabstractWith the rise in expectations related to generative models, text-to-video (T2V) models are being actively studied. Existing text-to-video models have limitations such as in generating complex movements replicating human motions. These model often generate unintended human motions, and the scale of the subject is incorrect. To overcome these limitations and generate high-quality videos that depict human motion under plausible viewing angles, we propose a two stage framework in this study. In the first stage a text-driven human motion generation network generates three-dimensional (3D) human motion from input text prompts and then motion-to-skeleton projection module projects generated motions onto a two-dimensional (2D) skeleton. In the second stage, the projected skeletons are used to generate a video in which the movements of a subject are well-represented. We demonstrated that the proposed framework quantitatively and qualitatively outperforms the existing T2V models. Previously reported human motion generation models use texts only or texts and human skeletons. However, our framework only uses texts and outputs a video related to human motion. Moreover, our framework benefits from using skeleton as an additional condition in the text-to-human motion generation networks. To the best of our knowledge, our framework is the first of its kind that uses text-driven human motion generation networks to generate high-quality videos related to human motions. The corresponding codes are available at https://github.com/CSJasper/HMTV. Chanhee Kang, JaeHyuk Park, Daun Jeong, ChangHee Yang, Suk-Ju Kang, Kyeongbo Kong |
WACV | 7 |
| 2024 | Image clustering using generated text centroids
Daehyeon Kong, Kyeongbo Kong, Suk-Ju Kang |
Signal Process. Image Commun. | 2 |
| 2024 | MosaicMVS: Mosaic-Based Omnidirectional Multi-View Stereo for Indoor ScenesabstractWe present MosaicMVS, a novel learning-based depth estimation framework for a mosaic-based omnidirectional multi-view stereo (MVS) camera setup. It uses a regular field of view (FOV) MVS network for an omnidirectional imaging setup with explicit consideration of hypothetical voxel-wise FOV overlaps. The resulting depth predictions are accurate and agree on the omnidirectional multi-view geometry. Unlike existing MVS setups, MosaicMVS camera setup can be easily applied to omnidirectional indoor scenes without having to account for constraints such as intricate epipolar constraints and the distortion of omnidirectional cameras. We validate the effectiveness of our framework on a new challenging indoor dataset in terms of depth estimation, reconstruction, and view synthesis. We also present new evaluation metric to check reconstruction performance using post-processed masks for accurate evaluation without any ground truth depth map or laser-scanned reconstructions. Experimental results show that our framework outperforms the state-of-the-art MVS methods in a large margin in all test scenes. Min-Jung Shin, Woojune Park, Minji Cho, Kyeongbo Kong, Hoseong Son, Joonsoo Kim, Kugjin Yun, Gwangsoon Lee, Suk-Ju Kang |
IEEE Trans. Multim. | 4 |
| 2023 | FeedFormer: Revisiting Transformer Decoder for Efficient Semantic SegmentationabstractWith the success of Vision Transformer (ViT) in image classification, its variants have yielded great success in many downstream vision tasks. Among those, the semantic segmentation task has also benefited greatly from the advance of ViT variants. However, most studies of the transformer for semantic segmentation only focus on designing efficient transformer encoders, rarely giving attention to designing the decoder. Several studies make attempts in using the transformer decoder as the segmentation decoder with class-wise learnable query. Instead, we aim to directly use the encoder features as the queries. This paper proposes the Feature Enhancing Decoder transFormer (FeedFormer) that enhances structural information using the transformer decoder. Our goal is to decode the high-level encoder features using the lowest-level encoder feature. We do this by formulating high-level features as queries, and the lowest-level feature as the key and value. This enhances the high-level features by collecting the structural information from the lowest-level feature. Additionally, we use a simple reformation trick of pushing the encoder blocks to take the place of the existing self-attention module of the decoder to improve efficiency. We show the superiority of our decoder with various light-weight transformer-based decoders on popular semantic segmentation datasets. Despite the minute computation, our model has achieved state-of-the-art performance in the performance computation trade-off. Our model FeedFormer-B0 surpasses SegFormer-B0 with 1.8% higher mIoU and 7.1% less computation on ADE20K, and 1.7% higher mIoU and 14.4% less computation on Cityscapes, respectively. Code will be released at: https://github.com/jhshim1995/FeedFormer. Jae-hun Shim, Hyunwoo Yu, Kyeongbo Kong, Suk-Ju Kang |
AAAI | 3 |
| 2023 | SEFD: Learning to Distill Complex Pose and OcclusionabstractThis paper addresses the problem of three-dimensional (3D) human mesh estimation in complex poses and occluded situations. Although many improvements have been made in 3D human mesh estimation using the two-dimensional (2D) pose with occlusion between humans, occlusion from complex poses and other objects remains a consistent problem. Therefore, we propose the novel Skinned Multi-Person Linear (SMPL) Edge Feature Distillation (SEFD) that demonstrates robustness to complex poses and occlusions, without increasing the number of parameters compared to the baseline model. The model generates an SMPL overlapping edge similar to the ground truth that contains target person boundary and occlusion information, performing subsequent feature distillation in a simple edge map. We also perform experiments on various benchmarks and exhibit fidelity both qualitatively and quantitatively. Extensive experiments prove that our method outperforms the state-of-the-art method by 2.8% in MPJPE and 1.9% in MPVPE on a benchmark 3DPW dataset in the presence of domain gap. Also, our method is superior in 3DPW-OCC, 3DPW-PC, RH-Dataset, OCHuman, Crowd-Pose, and LSP dataset in which occlusion, complex pose, and domain gap exist. The code and occlusion & complex pose annotation will be available at https://github.com/YangChangHee/ICCV2023_SEFD_RELEASE/. ChangHee Yang, Kyeongbo Kong, Sung-Jun Min, Dongyoon Wee, Ho-Deok Jang, Geonho Cha, Suk-Ju Kang |
ICCV | 2 |
| 2023 | An Unified Framework for Language Guided Image CompletionabstractImage completion is a research field which aims to generate visual contents for unknown regions of an image. Image outpainting and wide-range image blending, which we refer to as extensive painting, are considered challenging because compared to the large unknown regions, relatively less context is provided. Some recent studies have tried to decrease the complexity of extensive painting by generating image hints for the missing regions. In this paper, we introduce a novel modality of hints, the natural language. Moreover, we propose a Captioning-based Extensive Painting (CEP) module, which combines models for two different multi-modal tasks: image captioning and text-guided image completion. In order to generate appropriate captions for masked images, the image captioning model is optimized using self-critical sequence training (SCST) method with random masks. The biggest benefit of our methodology is the accessibility to well-designed image captioning and text-guided image manipulation models such as OFA and GLIDE without the need for additional architectural changes. In evaluation, our model demonstrates remarkable performance even with complicated image datasets both quantitatively and qualitatively. Seong-Hun Jeong, Kyeongbo Kong, Suk-Ju Kang |
WACV | 3 |
| 2023 | Out-of-Focus Image Deblurring for Mobile Display Vision InspectionabstractIn vision inspection tasks, moiré patterns caused by frequency aliasing can severely degrade image quality. To prevent moiré patterns, we used images that were intentionally out-of-focused, and we performed deblurring to restore details during the acquisition of the images. As existing deblurring methods fail to output satisfactory results for low-contrast Mura images, we applied some simple techniques, minimum-maximum normalization, and edge mask fine-tuning to one of the state-of-the-art non-blind deblurring methods by utilizing parametric generalized Gaussian kernels. Structural image details were preserved through edge mask fine-tuning, and image contrast was improved with minimum-maximum normalization. By parameterizing the blur kernel as a generalized Gaussian kernel, we greatly improved the robustness of the blind image deblurring. We evaluated the effects of each module by conducting thorough experiments. The proposed method showed better performance than existing blind deblurring methods for blur-specific no-reference metrics, the image profile, and frequency domain analysis. Sung-Jun Min, Kyeongbo Kong, Suk-Ju Kang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Subject-Invariant Deep Neural Networks Based on Baseline Correction for EEG Motor Imagery BCIabstractElectroencephalography (EEG)-based brain-computer interface (BCI) systems have been extensively used in various applications, such as communication, control, and rehabilitation. However, individual anatomical and physiological differences cause subject-specific variability of EEG signals for the same task, and BCI systems thus require a calibration procedure that adjusts system parameters to each subject. To overcome this problem, we propose a subject-invariant deep neural network (DNN) using baseline-EEG signals that can be recorded from subjects resting in comfortable states. We first modeled the deep features of EEG signals as a decomposition of subject-invariant and subject-variant features corrupted by anatomical/physiological characteristics. Subject-variant features were then removed from the deep features by learning the network with a baseline correction module (BCM) using the underlying individual information in baseline-EEG signals. The subject-invariant loss forces the BCM to assemble subject-invariant features that have the same class, irrespective of the subject. Using 1-min baseline-EEG signals of the new subject, our algorithm can eliminate subject-variant components from test data without the calibration process. The experimental results show that our subject-invariant DNN framework significantly increases decoding accuracies of the conventional DNN methods for BCI systems. Furthermore, feature visualizations illustrate that the proposed BCM extracts subject-invariant features that are close to each other in the same class. Youngchul Kwak, Kyeongbo Kong, Woo-Jin Song, Seong-Eun Kim |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Human Body-Aware Feature Extractor Using Attachable Feature Corrector for Human Pose EstimationabstractTop-down pose estimation generally employs a person detector and estimates the keypoints of the detected person. This method assumes that only a single person exists within the bounding box cropped by detection. However, this assumption leads to some challenges in practice. First, a loose-fitted bounding box may include certain body parts of a non-target person. Second, spatial interference between several people exists owing to occlusion, so more than a single person can exist in the cropped image. In such scenarios, the pose estimation may falsely predict the keypoints of two or more persons as those of a single person. To tackle these issues, this paper proposes the human body-aware feature extractor based on the global- and local-reasoning features. The global-reasoning feature considers the entire body using transformer's non-local computation property and the local-reasoning feature concentrates on the individual body parts using convolutional neural networks. With those two features, we extract corrected features by filtering unnecessary features and supplementing necessary features using our proposed novel architecture. Hence, the proposed method can focus on the target person's keypoints, thereby mitigating the aforementioned concerns. Our method achieves noticeable improvement when applied to state-of-the-art top-down pose estimation networks. Ginam Kim, Kyeongbo Kong, Jou Won Song, Suk-Ju Kang |
IEEE Trans. Multim. | 3 |
| 2022 | Selective TransHDR: Transformer-Based Selective HDR Imaging Using Ghost Region Mask
Jou Won Song, Ye In Park, Kyeongbo Kong, Jaeho Kwak, Suk-Ju Kang |
ECCV (17) | 3 |
| 2022 | Image-Adaptive Hint Generation via Vision Transformer for OutpaintingabstractImage outpainting has recently received considerable attention because it can be useful in tasks such as image retargeting and panorama image generation. In general, the problem of extending an image beyond its given boundaries is still ill-posed. Conventional methods predominantly attempt image outpainting by using complex network structures. Some recent studies have tried to decrease the problem complexity through the conversion techniques from outpainting to inpainting. Although these methodologies work well in simple cases, their performance reduces considerably for asymmetrical images. This paper proposes a novel hint-based outpainting methodology that can adaptively select the most plausible patches as hints from a given image to reduce the difficulty of outpainting. To estimate high-quality hints, inspired by patch-based image inpainting methods, we utilize Vision Transformer that considers self-attention for each patch. The estimated hints are attached on both boundaries of the input image and the inside missing regions are predicted by using an inpainting network. After finishing the prediction, the output image is obtained by removing the hints. Experiments show that our image-adaptive hint framework, when employed in representative inpainting networks, can consistently improve its performance compared to the other conversion techniques from outpainting to inpainting on SUN and Beach benchmark datasets. Daehyeon Kong, Kyeongbo Kong, Kyunghun Kim, Sung-Jun Min, Suk-Ju Kang |
WACV | 2 |
| 2022 | Penalty based robust learning with noisy labels
Kyeongbo Kong, Junggi Lee, Youngchul Kwak, Young-Rae Cho, Seong-Eun Kim, Woo-Jin Song |
Neurocomputing | 1 |
| 2022 | Dynamic Hand Gesture Recognition Using Improved Spatio-Temporal Graph Convolutional NetworkabstractHand gesture recognition is essential to human-computer interaction as the most natural way of communicating. Furthermore, with the development of 3D hand pose estimation technology and the performance improvement of low-cost depth cameras, skeleton-based dynamic hand gesture recognition has received much attention. This paper proposes a novel multi-stream improved spatio-temporal graph convolutional network (MS-ISTGCN) for skeleton-based dynamic hand gesture recognition. We adopt an adaptive spatial graph convolution that can learn the relationship between distant hand joints and propose an extended temporal graph convolution with multiple dilation rates that can extract informative temporal features from short to long periods. Furthermore, we add a new attention layer consisting of effective spatio-temporal attention and channel attention between the spatial and temporal graph convolution layers to find and focus on key features. Finally, we propose a multi-stream structure that feeds multiple data modalities (i.e., joints, bones, and motions) as inputs to improve performance using the ensemble technique. Each of the three-stream networks is independently trained and fused to predict the final hand gesture. The performance of the proposed method is verified through extensive experiments with two widely used public dynamic hand gesture datasets: SHREC’17 Track and DHG-14/28. Our proposed method achieves the highest recognition accuracy in various gesture categories for both datasets compared with state-of-the-art methods. Jae-Hun Song, Kyeongbo Kong, Suk-Ju Kang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Painting Outside as Inside: Edge Guided Image Outpainting via Bidirectional Rearrangement with Progressive Step LearningabstractImage outpainting is a very intriguing problem as the outside of a given image can be continuously filled by considering as the context of the image. This task has two main challenges. The first is to maintain the spatial consistency in contents of generated regions and the original input. The second is to generate a high-quality large image with a small amount of adjacent information. Conventional image outpainting methods generate inconsistent, blurry, and repeated pixels. To alleviate the difficulty of an outpainting problem, we propose a novel image outpainting method using bidirectional boundary region rearrangement. We rear-range the image to benefit from the image inpainting task by reflecting more directional information. The bidirectional boundary region rearrangement enables the generation of the missing region using bidirectional information similar to that of the image inpainting task, thereby generating the higher quality than the conventional methods using unidirectional information. Moreover, we use the edge map generator that considers images as original input with structural information and hallucinates the edges of unknown regions to generate the image. Our proposed method is compared with other state-of-the-art outpainting and inpainting methods both qualitatively and quantitatively. We further compared and evaluated them using BRISQUE, one of the No-Reference image quality assessment (IQA) metrics, to evaluate the naturalness of the output. The experimental results demonstrate that our method outperforms other methods and generates new images with 360°panoramic characteristics. Kyunghun Kim, Yeohun Yun, Keon-Woo Kang, Kyeongbo Kong, Siyeong Lee, Suk-Ju Kang |
WACV | 4 |
| 2019 | How to Estimate Global Motion Non-Iteratively From a Coarsely Sampled Motion Vector FieldabstractNowadays, as the amount of video content increases, the importance of global motion estimation is increasing more and more. This paper considers motion estimation within a compressed domain and presents a robust non-iterative global motion estimation algorithm that efficiently eliminates outliers. The said algorithm is comprised of two parts. The first part uses the distinct-outlier mask based on statistical analysis to eliminate the distinct outliers and uses the current-object mask based on the reference-object mask to eliminate the object outliers. In the second part, the proposed median error designed to be robust to the outlier ratio has been used to estimate weights with higher values as the motion vector gets closer to global motion. Using these weights, the proposed algorithm can accurately estimate global motion non-iteratively, although the outlier ratio may change. The simulation results demonstrate that the proposed algorithm demonstrates the highest estimation accuracy and processing speed in synthetic motion fields as well as real video sequences. Kyeongbo Kong, Seung-Jun Shin, Junggi Lee, Woo-Jin Song |
IEEE Trans. Circuits Syst. Video Technol. | 1 |