Taeoh Kim

dblp:226/2517 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
16since 2021 · last 2025
0000-0001-7252-5525ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 12 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused Images
abstract
3D Gaussian Splatting (3DGS) has attracted significant attention for its high-quality novel view rendering, inspiring research to address real-world challenges. While conventional methods depend on sharp images for accurate scene reconstruction, real-world scenarios are often affected by defocus blur due to finite depth of field, making it essential to account for realistic 3D scene representation. In this study, we propose CoCoGaussian, a Circle of Confusion-aware Gaussian Splatting that enables precise 3D scene representation using only defocused images. CoCoGaussian addresses the challenge of defocus blur by modeling the Circle of Confusion (CoC) through a physically grounded approach based on the principles of photographic defocus. Exploiting 3D Gaussians, we compute the CoC diameter from depth and learnable aperture information, generating multiple Gaussians to precisely capture the CoC shape. Furthermore, we introduce a learnable scaling factor to enhance robustness and provide more flexibility in handling unreliable depth in scenes with reflective or refractive surfaces. Experiments on both synthetic and real-world datasets demonstrate that CoCoGaussian achieves state-of-the-art performance across multiple benchmarks.
Suhwan Cho, Taeoh Kim, Ho-Deok Jang, Minhyeok Lee, Geonho Cha, Dongyoon Wee, Dogyoon Lee, Sangyoun Lee
CVPR3
2025 Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
abstract
Video large language models (LLMs) achieve strong video understanding by leveraging a large number of spatio-temporal tokens, but suffer from quadratic computational scaling with token count. To address this, we propose a training-free spatio-temporal token merging method, named STTM. Our key insight is to exploit local spatial and temporal redundancy in video data which has been overlooked in prior work. STTM first transforms each frame into multi-granular spatial tokens using a coarse-to-fine search over a quadtree structure, then performs directed pairwise merging across the temporal dimension. This decomposed merging approach outperforms existing token reduction methods across six video QA benchmarks. Notably, STTM achieves a 2$\times$ speed-up with only a 0.5% accuracy drop under a 50% token budget, and a 3$\times$ speed-up with just a 2% drop under a 30% budget. Moreover, STTM is query-agnostic, allowing KV cache reuse across different questions for the same video. The project page is available at https://www.jshyun.me/projects/sttm.
Jeongseok Hyun, Sukjun Hwang, Su Ho Han, Taeoh Kim, Inwoong Lee, Dongyoon Wee, Joon-Young Lee, Seon Joo Kim, Minho Shim
ICCV4
2025 CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Images
abstract
3D Gaussian Splatting (3DGS) has gained significant attention due to its high-quality novel view rendering, motivating research to address real-world challenges. A critical issue is the camera motion blur caused by movement during exposure, which hinders accurate 3D scene reconstruction. In this study, we propose CoMoGaussian, a Continuous Motion-Aware Gaussian Splatting that reconstructs precise 3D scenes from motion-blurred images while maintaining real-time rendering speed. Considering the complex motion patterns inherent in real-world camera movements, we predict continuous camera trajectories using neural ordinary differential equations (ODEs). To ensure accurate modeling, we employ rigid body transformations, preserving the shape and size of the object but rely on the discrete integration of sampled frames. To better approximate the continuous nature of motion blur, we introduce a continuous motion refinement (CMR) transformation that refines rigid transformations by incorporating additional learnable parameters. By revisiting fundamental camera theory and leveraging advanced neural ODE techniques, we achieve precise modeling of continuous camera trajectories, leading to improved reconstruction accuracy. Extensive experiments demonstrate state-of-the-art performance both quantitatively and qualitatively on benchmark datasets, which include a wide range of motion blur scenarios, from moderate to extreme blur.
Donghyeong Kim, Dogyoon Lee, Suhwan Cho, Minhyeok Lee, Wonjoon Lee, Taeoh Kim, Dongyoon Wee, Sangyoun Lee
ICCV7
2025 Prototypes Are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
abstract
In a retrieval system, simultaneously achieving search accuracy and efficiency is inherently challenging. This challenge is particularly pronounced in partially relevant video retrieval (PRVR), where incorporating more diverse context representations at varying temporal scales for each video enhances accuracy but increases computational and memory costs. To address this dichotomy, we propose a prototypical PRVR framework that encodes diverse contexts within a video into a fixed number of prototypes. We then introduce several strategies to enhance text association and video understanding within the prototypes, along with an orthogonal objective to ensure that the prototypes capture a diverse range of content. To keep the prototypes searchable via text queries while accurately encoding video contexts, we implement cross- and uni-modal reconstruction tasks. The cross-modal reconstruction task aligns the prototypes with textual features within a shared space, while the uni-modal reconstruction task preserves all video contexts during encoding. Additionally, we employ a video mixing technique to provide weak guidance to further align prototypes and associated textual representations. Extensive evaluations on TVR, ActivityNet-Captions, and QVHighlights validate the effectiveness of our approach without sacrificing efficiency.
WonJun Moon, Cheol-Ho Cho, Woojin Jun, Taeoh Kim, Inwoong Lee, Dongyoon Wee, Minho Shim, Jae-Pil Heo
ICCV4
2024 Classification Matters: Improving Video Action Detection with Class-Specific Attention
Jinsung Lee, Taeoh Kim, Inwoong Lee, Minho Shim, Dongyoon Wee, Minsu Cho, Suha Kwak
ECCV (20)2
2024 Towards Multi-Domain Learning for Generalizable Video Anomaly Detection
abstract
Most of the existing Video Anomaly Detection (VAD) studies have been conducted within single-domain learning, where training and evaluation are performed on a single dataset. However, the criteria for abnormal events differ across VAD datasets, making it problematic to apply a single-domain model to other domains. In this paper, we propose a new task called Multi-Domain learning forVAD (MDVAD) to explore various real-world abnormal events using multiple datasets for a general model. MDVAD involves training on datasets from multiple domains simultaneously, and we experimentally observe that Abnormal Conflicts between domains hinder learning and generalization. The task aims to address two key objectives: (i) better distinguishing between general normal and abnormal events across multiple domains, and (ii) being aware of ambiguous abnormal conflicts. This paper is the first to tackle abnormal conflict issue and introduces a new benchmark, baselines, and evaluation protocols for MDVAD. As baselines, we propose a framework with Null(Angular)-Multiple Instance Learning and an Abnormal Conflict classifier. Through experiments on a MDVAD benchmark composed of six VAD datasets and using four different evaluation protocols, we reveal abnormal conflicts and demonstrate that the proposed baseline effectively handles these conflicts, showing robustness and adaptability across multiple domains.
MyeongAh Cho, Taeoh Kim, Minho Shim, Dongyoon Wee, Sangyoun Lee
NeurIPS2
2024 A Nonlinear, Regularized, and Data-independent Modulation for Continuously Interactive Image Processing Network
Hyeongmin Lee, Taeoh Kim, Hanbin Son, Sangwook Baek, Minsu Cheon, Sangyoun Lee
Int. J. Comput. Vis.2
2023 Frequency Selective Augmentation for Video Representation Learning
abstract
Recent self-supervised video representation learning methods focus on maximizing the similarity between multiple augmented views from the same video and largely rely on the quality of generated views. However, most existing methods lack a mechanism to prevent representation learning from bias towards static information in the video. In this paper, we propose frequency augmentation (FreqAug), a spatio-temporal data augmentation method in the frequency domain for video representation learning. FreqAug stochastically removes specific frequency components from the video so that learned representation captures essential features more from the remaining information for various downstream tasks. Specifically, FreqAug pushes the model to focus more on dynamic features rather than static features in the video via dropping spatial or temporal low-frequency components. To verify the generality of the proposed method, we experiment with FreqAug on multiple self-supervised learning frameworks along with standard augmentations. Transferring the improved representation to five video action recognition and two temporal action localization downstream tasks shows consistent improvements over baselines.
Jinhyung Kim, Taeoh Kim, Minho Shim, Dongyoon Han, Dongyoon Wee, Junmo Kim 0002
AAAI2
2023 Decomposed Cross-Modal Distillation for RGB-based Temporal Action Detection
abstract
Temporal action detection aims to predict the time intervals and the classes of action instances in the video. Despite the promising performance, existing two-stream models exhibit slow inference speed due to their reliance on computationally expensive optical flow. In this paper, we introduce a decomposed cross-modal distillation framework to build a strong RGB-based detector by transferring knowledge of the motion modality. Specifically, instead of direct distillation, we propose to separately learn RGB and motion representations, which are in turn combined to perform action localization. The dual-branch design and the asymmetric training objectives enable effective motion knowledge transfer while preserving RGB information intact. In addition, we introduce a local attentive fusion to better exploit the multimodal complementarity. It is designed to preserve the local discriminability of the features that is important for action localization. Extensive experiments on the benchmarks verify the effectiveness of the proposed method in enhancing RGB-based action detectors. Notably, our framework is agnostic to backbones and detection heads, bringing consistent gains across different model combinations.
Pilhyeon Lee, Taeoh Kim, Minho Shim, Dongyoon Wee, Hyeran Byun
CVPR2
2023 Exploring Temporally Dynamic Data Augmentation for Video Recognition
Taeoh Kim, Jinhyung Kim, Minho Shim, Sangdoo Yun, Myunggu Kang, Dongyoon Wee, Sangyoun Lee
ICLR1
2022 Unsupervised video anomaly detection via normalizing flows with implicit latent features
MyeongAh Cho, Taeoh Kim, Woo Jin Kim, Suhwan Cho, Sangyoun Lee
Pattern Recognit.2
2022 Geometry-Aware Deep Video Deblurring via Recurrent Feature Refinement
abstract
Blurring in videos is a frequent phenomenon in real-world video data owing to camera shake or object movement at different scene depths. Hence, video deblurring is an ill-posed problem that requires understanding of geometric and temporal information. Traditional model-based optimization methods first define a degradation model and then solve an optimization problem to recover the latent frames with a variational model for additional external information, such as optical flow, segmentation, depth, or camera movement. Recent deep-learning-based approaches learn from numerous training pairs of blurred and clean latent frames, with the powerful representation ability of deep convolutional neural networks. Although deep models have achieved remarkable performances without the explicit model, existing deep methods do not utilize geometrical information as strong priors. Therefore, they cannot handle extreme blurring caused by large camera shake or scene depth variations. In this paper, we propose a geometry-aware deep video deblurring method via a recurrent feature refinement module that exploits optimization-based and deep-learning-based schemes. In addition to the off-the-shelf deep geometry estimation modules, we design an effective fusion module for geometrical information with deep video features. Specifically, similar to model-based optimization, our proposed module recurrently refines video features as well as geometrical information to restore more precise latent frames. To evaluate the effectiveness and generalization of our framework, we perform tests on eight baseline networks whose structures are motivated by the previous research. The experimental results show that our framework offers greater performances than the eight baselines and produces state-of-the-art performance on four video deblurring benchmark datasets.
Taeoh Kim, Sangyoun Lee
IEEE Trans. Image Process.1
2022 Enhanced Standard Compatible Image Compression Framework Based on Auxiliary Codec Networks
abstract
Recent deep neural network-based research to enhance image compression performance can be divided into three categories: learnable codecs, postprocessing networks, and compact representation networks. The learnable codec has been designed for end-to-end learning beyond the conventional compression modules. The postprocessing network increases the quality of decoded images using example-based learning. The compact representation network is learned to reduce the capacity of an input image, reducing the bit rate while maintaining the quality of the decoded image. However, these approaches are not compatible with existing codecs or are not optimal for increasing coding efficiency. Specifically, it is difficult to achieve optimal learning in previous studies using a compact representation network due to the inaccurate consideration of the codecs. In this paper, we propose a novel standard compatible image compression framework based on auxiliary codec networks (ACNs). In addition, ACNs are designed to imitate image degradation operations of the existing codec, which delivers more accurate gradients to the compact representation network. Therefore, compact representation and postprocessing networks can be learned effectively and optimally. We demonstrate that the proposed framework based on the JPEG and High Efficiency Video Coding standard substantially outperforms existing image compression algorithms in a standard compatible manner.
Hanbin Son, Taeoh Kim, Hyeongmin Lee, Sangyoun Lee
IEEE Trans. Image Process.2
2021 Test-Time Adaptation for Out-Of-Distributed Image Inpainting
abstract
Deep-learning-based image inpainting algorithms have shown great performance via powerful learned priors from numerous external natural images. However, they show unpleasant results for test images whose distributions are far from those of the training images because their models are biased toward the training images. In this paper, we propose a simple image inpainting algorithm with test-time adaptation named AdaFill. Given a single out-of-distributed test image, our goal is to complete hole region more naturally than the pre-trained inpainting models. To achieve this goal, we treat the remaining valid regions of the test image as an another training cue because natural images have strong internal similarities. From this test-time adaptation, our network can exploit externally learned image priors from the pre-trained features as well as the internal priors of the test image explicitly. The experimental results show that AdaFill outperforms other models on various out-of-distribution test images. Furthermore, the model named ZeroFill, which is not pre-trained also outperforms the pre-trained models sometimes.
Chajin Shin, Taeoh Kim, Sangyoun Lee
ICIP2
2021 Short-term photovoltaic power generation predicting by input/output structure of weather forecast using deep learning
Dongha Shin, Eungyu Ha, Taeoh Kim, Changbok Kim
Soft Comput.3
2021 Relational Deep Feature Learning for Heterogeneous Face Recognition
abstract
Heterogeneous Face Recognition (HFR) is a task that matches faces across two different domains such as visible light (VIS), near-infrared (NIR), or the sketch domain. Due to the lack of databases, HFR methods usually exploit the pre-trained features on a large-scale visual database that contain general facial information. However, these pre-trained features cause performance degradation due to the texture discrepancy with the visual domain. With this motivation, we propose a graph-structured module called Relational Graph Module (RGM) that extracts global relational information in addition to general facial features. Because each identity's relational information between intra-facial parts is similar in any modality, the modeling relationship between features can help cross-domain matching. Through the RGM, relation propagation diminishes texture dependency without losing its advantages from the pre-trained features. Furthermore, the RGM captures global facial geometrics from locally correlated convolutional features to identify long-range relationships. In addition, we propose a Node Attention Unit (NAU) that performs node-wise recalibration to concentrate on the more informative nodes arising from relation-based propagation. Furthermore, we suggest a novel conditional-margin loss function ($C$ -softmax) for the efficient projection learning of the embedding vector in HFR. The proposed method outperforms other state-of-the-art methods on five HFR databases. Furthermore, we demonstrate performance improvement on three backbones because our module can be plugged into any pre-trained face recognition backbone to overcome the limitations of a small HFR database.
MyeongAh Cho, Taeoh Kim, Ig-Jae Kim, Kyungjae Lee 0003, Sangyoun Lee
IEEE Trans. Inf. Forensics Secur.2
2020 AdaCoF: Adaptive Collaboration of Flows for Video Frame Interpolation
abstract
Video frame interpolation is one of the most challenging tasks in video processing research. Recently, many studies based on deep learning have been suggested. Most of these methods focus on finding locations with useful information to estimate each output pixel using their own frame warping operations. However, many of them have Degrees of Freedom (DoF) limitations and fail to deal with the complex motions found in real world videos. To solve this problem, we propose a new warping module named Adaptive Collaboration of Flows (AdaCoF). Our method estimates both kernel weights and offset vectors for each target pixel to synthesize the output frame. AdaCoF is one of the most generalized warping modules compared to other approaches, and covers most of them as special cases of it. Therefore, it can deal with a significantly wide domain of complex motions. To further improve our framework and synthesize more realistic outputs, we introduce dual-frame adversarial loss which is applicable only to video frame interpolation tasks. The experimental results show that our method outperforms the state-of-the-art methods for both fixed training set environments and the Middlebury benchmark. Our source code is available at https://github.com/HyeongminLEE/AdaCoF-pytorch
Hyeongmin Lee, Taeoh Kim, Tae-Young Chung, Daehyun Pak, Yuseok Ban, Sangyoun Lee
CVPR2
2020 Extrapolative-Interpolative Cycle-Consistency Learning For Video Frame Extrapolation
abstract
Video frame extrapolation is a task to predict future frames when the past frames are given. Unlike previous studies that usually have been focused on the design of modules or construction of networks, we propose a novel ExtrapolativeInterpolative Cycle (EIC) loss using pre-trained frame interpolation module to improve extrapolation performance. Cycle-consistency loss has been used for stable prediction between two function spaces in many visual tasks. We formulate this cycle-consistency using two mapping functions; frame extrapolation and interpolation. Since it is easier to predict intermediate frames than to predict future frames in terms of the object occlusion and motion uncertainty, interpolation module can give guidance signal effectively for training the extrapolation function. EIC loss can be applied to any existing extrapolation algorithms and guarantee consistent prediction in the short future as well as long future frames. Experimental results show that simply adding EIC loss to the existing baseline increases extrapolation performance on both UCF101 [1] and KITTI [2] datasets.
Hyeongmin Lee, Taeoh Kim, Sangyoun Lee
ICIP3
2019 SF-CNN: A Fast Compression Artifacts Removal via Spatial-To-Frequency Convolutional Neural Networks
abstract
In this paper, we propose SF-CNN, a fast convolutional neural network structure for JPEG image compression artifacts removal. Recently, Convolutional Neural Network (CNN)-based image restoration has shown great performance improvement. However, its heavy computational cost makes it difficult to apply to other uses such as high-level vision tasks. Since heavy computation arises from maintaining the spatial resolution of an input image, some works make a structure that is composed of spatial downsampling and upsampling operations. SF-CNN takes Spatial input and predicts residual Frequency using downsampling operations only. Since every 8×8 pixel is grouped and spatially invariant in the JPEG DCT domain, it is possible to down sample the input by a factor of 8 to reduce the computational cost. We show this simple structure is effective for compression artifacts removal. Our scalable baseline networks achieve results comparable to to the reference networks in reduced computations.
Taeoh Kim, Hyeongmin Lee, Hanbin Son, Sangyoun Lee
ICIP1
2018 Collabonet: Collaboration of Generative Models by Unsupervised Classification
abstract
Designing models for learning dataset with complex distributions is one of the main challenges that still remains in machine learning areas. We propose CollaboNet, which can divide a large dataset into sub-datasets, train two generative models separately, and let two models work together to achieve better performance. The proposed algorithm divides a large dataset without label since the capability difference between two generative models in performing tasks on each data is the main criterion for dividing a large dataset. In other words, the classification model can be trained by unsupervised manner. Autoencoder experiments for pure MNIST and the datasets combined artificially from two image sets shows that CollaboNet successfully splits large datasets without labels, improving the performance of generative models.
Hyeongmin Lee, Taeoh Kim, Eungyeol Song, Sangyoun Lee
ICIP2