VLDB 2026 Research / reviewers in the wild / expert
Donghyeong Kim
dblp:329/4347
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-1974-6270ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalizing CLIP prompts for zero-shot anomaly detection
Donghyeong Kim, Suhwan Cho, Hyeonjeong Lim, Sangyoun Lee |
Pattern Recognit. | 1 |
| 2025 | CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Imagesabstract3D Gaussian Splatting (3DGS) has gained significant attention due to its high-quality novel view rendering, motivating research to address real-world challenges. A critical issue is the camera motion blur caused by movement during exposure, which hinders accurate 3D scene reconstruction. In this study, we propose CoMoGaussian, a Continuous Motion-Aware Gaussian Splatting that reconstructs precise 3D scenes from motion-blurred images while maintaining real-time rendering speed. Considering the complex motion patterns inherent in real-world camera movements, we predict continuous camera trajectories using neural ordinary differential equations (ODEs). To ensure accurate modeling, we employ rigid body transformations, preserving the shape and size of the object but rely on the discrete integration of sampled frames. To better approximate the continuous nature of motion blur, we introduce a continuous motion refinement (CMR) transformation that refines rigid transformations by incorporating additional learnable parameters. By revisiting fundamental camera theory and leveraging advanced neural ODE techniques, we achieve precise modeling of continuous camera trajectories, leading to improved reconstruction accuracy. Extensive experiments demonstrate state-of-the-art performance both quantitatively and qualitatively on benchmark datasets, which include a wide range of motion blur scenarios, from moderate to extreme blur. Donghyeong Kim, Dogyoon Lee, Suhwan Cho, Minhyeok Lee, Wonjoon Lee, Taeoh Kim, Dongyoon Wee, Sangyoun Lee |
ICCV | 2 |
| 2025 | CMTM: Cross-Modal Token Modulation for Unsupervised Video Object SegmentationabstractRecent advances in unsupervised video object segmentation have highlighted the potential of two-stream architectures that integrate appearance and motion cues. However, fully leveraging these complementary sources of information requires effectively modeling their interdependencies. In this paper, we introduce cross-modality token modulation, a novel approach designed to strengthen the interaction between appearance and motion cues. Our method establishes dense connections between tokens from each modality, enabling efficient intra-modal and inter-modal information propagation through relation transformer blocks. To improve learning efficiency, we incorporate a token masking strategy that addresses the limitations of relying solely on increased model complexity. Our approach achieves state-of-the-art performance across all public benchmarks, outperforming existing methods. The code is released on https://github.com/InSeokJeon/CMTM Inseok Jeon, Suhwan Cho, Minhyeok Lee, Donghyeong Kim, Sangyoun Lee |
ICIP | 8 |
| 2025 | Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal GroundingabstractVideo Temporal Grounding (VTG) aims to localize temporal segments in long, untrimmed videos that align with a given natural language query. This task typically comprises two subtasks: \textit{Moment Retrieval (MR)} and \textit{Highlight Detection (HD)}. While recent advances have been progressed by powerful pretrained vision-language models such as CLIP and InternVideo2, existing approaches commonly treat all text tokens uniformly during cross-modal attention, disregarding their distinct semantic roles. To validate the limitations of this approach, we conduct controlled experiments demonstrating that VTG models overly rely on [EOS]-driven global semantics while failing to effectively utilize word-level signals, which limits their ability to achieve fine-grained temporal alignment. Motivated by this limitation, we propose DualGround, a dual-branch architecture that explicitly separates global and local semantics by routing the [EOS] token through a sentence-level path and clustering word tokens into phrase-level units for localized grounding. Our method introduces (1) token-role-aware cross modal interaction strategies that align video features with sentence-level and phrase-level semantics in a structurally disentangled manner, and (2) a joint modeling framework that not only improves global sentence-level alignment but also enhances fine-grained temporal grounding by leveraging structured phrase-aware context. This design allows the model to capture both coarse and localized semantics, enabling more expressive and context-aware video grounding. DualGround achieves state-of-the-art performance on both Moment Retrieval and Highlight Detection tasks across QVHighlights and Charades-STA benchmarks, demonstrating the effectiveness of disentangled semantic modeling in video-language alignment. Minhyeok Lee, Donghyeong Kim, Sangyoun Lee |
NeurIPS | 4 |
| 2025 | Sparse-DeRF: Deblurred Neural Radiance Fields From Sparse ViewabstractRecent studies construct deblurred neural radiance fields (DeRF) using dozens of blurry images, which are not practical scenarios if only a limited number of blurry images are available. This paper focuses on constructing DeRF from sparse-view for more pragmatic real-world scenarios. As observed in our experiments, establishing DeRF from sparse views proves to be a more challenging problem due to the inherent complexity arising from the simultaneous optimization of blur kernels and NeRF from sparse view. Sparse-DeRF successfully regularizes the complicated joint optimization, presenting alleviated overfitting artifacts and enhanced quality on radiance fields. The regularization consists of three key components: Surface smoothness, helps the model accurately predict the scene structure utilizing unseen and additional hidden rays derived from the blur kernel based on statistical tendencies of real-world; Modulated gradient scaling, helps the model adjust the amount of the backpropagated gradient according to the arrangements of scene objects; Perceptual distillation improves the perceptual quality by overcoming the ill-posed multi-view inconsistency of image deblurring and distilling the pre-deblurred information, compensating for the lack of clean information in blurry images. We demonstrate the effectiveness of the Sparse-DeRF with extensive quantitative and qualitative experimental results by training DeRF from 2-view, 4-view, and 6-view blurry images. Dogyoon Lee, Donghyeong Kim, Minhyeok Lee, Seunghoon Lee 0008, Sangyoun Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Fast video anomaly detection via context-aware shortcut exploration and abnormal feature distance learning
Donghyeong Kim, MyeongAh Cho, Minjung Kim 0002, Minseok Lee, Seungwook Park, Sangyoun Lee |
Pattern Recognit. | 2 |
| 2024 | FIMP: Future Interaction Modeling for Multi-Agent Motion PredictionabstractMulti-agent motion prediction is a crucial concern in autonomous driving, yet it remains a challenge owing to the ambiguous intentions of dynamic agents and their intricate interactions. Existing studies have attempted to capture interactions between road entities by using the definite data in history timesteps, as future information is not available and involves high uncertainty. However, without sufficient guidance for capturing future states of interacting agents, they frequently produce unrealistic trajectory overlaps. In this work, we propose Future Interaction modeling for Motion Prediction (FIMP), which captures potential future interactions in an end-to-end manner. FIMP adopts a future decoder that implicitly extracts the potential future information in an intermediate feature-level, and identifies the interacting entity pairs through future affinity learning and top-k filtering strategy. Experiments show that our future interaction modeling improves the performance remarkably, leading to superior performance on the Argoverse motion forecasting benchmark. Sungmin Woo, Minjung Kim 0002, Donghyeong Kim, Sungjun Jang, Sangyoun Lee |
ICRA | 3 |
| 2023 | FAPM: Fast Adaptive Patch Memory for Real-Time Industrial Anomaly DetectionabstractFeature embedding-based methods have shown exceptional performance in detecting industrial anomalies by comparing features of target images with normal images. However, some methods do not meet the speed requirements of real-time inference, which is crucial for real-world applications. To address this issue, we propose a new method called Fast Adaptive Patch Memory (FAPM) for real-time industrial anomaly detection. FAPM utilizes patch-wise and layer-wise memory banks that store the embedding features of images at the patch and layer level, respectively, which eliminates unnecessary repetitive computations. We also propose patch-wise adaptive coreset sampling for faster and more accurate detection. FAPM performs well in both accuracy and speed compared to other state-of-the-art methods. Donghyeong Kim, Suhwan Cho, Sangyoun Lee |
ICASSP | 1 |
| 2023 | Two-Stream Decoder Feature Normality Estimating Network for Industrial Anomaly DetectionabstractImage reconstruction-based anomaly detection has recently been in the spotlight because of the difficulty of constructing anomaly datasets. These approaches work by learning to model normal features without seeing abnormal samples during training and then discriminating anomalies at test time based on the reconstructive errors. However, these models have limitations in reconstructing the abnormal samples due to their indiscriminate conveyance of features. Moreover, these approaches are not explicitly optimized for distinguishable anomalies. To address these problems, we propose a two-stream decoder network (TSDN), designed to learn both normal and abnormal features. Additionally, we propose a feature normality estimator (FNE) to eliminate abnormal features and prevent high-quality reconstruction of abnormal regions. Evaluation on a standard benchmark demonstrated performance better than state-of-the-art models. Minhyeok Lee, Suhwan Cho, Donghyeong Kim, Sangyoun Lee |
ICASSP | 4 |
| 2023 | MSV-RGNN: Multiscale Voxel Graph Neural Network for 3D Object DetectionabstractThis paper proposes a two-stage 3D object detection framework, multiscale voxel graph neural network (MSV-RGNN) which aims to fully exploit multiple scale graph features by establishing global and local relationships between voxel features at different 3D convolutional neural network (CNN) layers. In contrast to conventional graph-based methods, our proposed multiscale-voxel-graph region-of-interest (RoI) pooling module constructs graphs across diverse voxel resolutions to obtain geometric structure information on voxel features. Initially, our multiscale-voxel-graph RoI pooling module sample voxel center points with voxel-wise feature vectors and 3D region proposals from backbone network. Subsequently, graphs are constructed at different scales and graph features are aggregated for second-stage refinement. The experimental results demonstrate the potential of using multiscale graphs across different voxel resolutions for 3D object detection, achieving decent experimental results with state-of-the-art methods. Wonjoon Lee, Sungmin Woo, Donghyeong Kim, Sangyoun Lee |
ICIP | 3 |
| 2023 | Treating Motion as Option to Reduce Motion Dependency in Unsupervised Video Object SegmentationabstractUnsupervised video object segmentation (VOS) aims to detect the most salient object in a video sequence at the pixel level. In unsupervised VOS, most state-of-the-art methods leverage motion cues obtained from optical flow maps in addition to appearance cues to exploit the property that salient objects usually have distinctive movements compared to the background. However, as they are overly dependent on motion cues, which may be unreliable in some cases, they cannot achieve stable prediction. To reduce this motion dependency of existing two-stream VOS methods, we propose a novel motion-as-option network that optionally utilizes motion cues. Additionally, to fully exploit the property of the proposed network that motion is not always required, we introduce a collaborative network learning strategy. On all the public benchmark datasets, our proposed network affords state-of-the-art performance with real-time inference speed. Code and models are available at https://github.com/suhwan-cho/TMO. Suhwan Cho, Minhyeok Lee, Seunghoon Lee 0008, Donghyeong Kim, Sangyoun Lee |
WACV | 5 |