VLDB 2026 Research / reviewers in the wild / expert
Hoonhee Cho
dblp:323/9541
· DBLP profile ↗
21ranked-venue papers
13as first author
21since 2021 · last 2026
0000-0003-0896-6793ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 12 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 11 first-author · 18 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temporal Stereo Matching From Event Cameras via Joint Learning With Stereoscopic FlowabstractEvent cameras are dynamic vision sensors inspired by the biological retina, offering high dynamic range, high temporal resolution, and low power consumption. These qualities allow them to perceive 3D environments even in extreme conditions. Event data is continuously recorded over time, capturing pixel movements in detail. To leverage this temporal density, we introduce a temporal event stereo framework that continuously uses past information. The event stereo matching network is jointly trained with stereoscopic flow, which tracks pixel movements from stereo cameras. Instead of relying on optical flow ground truth, our method trains motion flows using disparity maps. The temporal aggregation of information via stereoscopic flow boosts stereo matching performance, achieving state-of-the-art results on MVSEC, DSEC, M3ED, and EVIMO2 datasets. Our method also demonstrates computational efficiency by stacking past data in a cascading manner. Jae-Young Kang, Hoonhee Cho, Kuk-Jin Yoon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Ev-3DOD: Pushing the Temporal Boundaries of 3D Object Detection with Event CamerasabstractDetecting 3D objects in point clouds plays a crucial role in autonomous driving systems. Recently, advanced multimodal methods incorporating camera information have achieved notable performance. For a safe and effective autonomous driving system, algorithms that excel not only in accuracy but also in speed and low latency are essential. However, existing algorithms fail to meet these requirements due to the latency and bandwidth limitations of fixed frame rate sensors, e.g., LiDAR and camera. To address this limitation, we introduce asynchronous event cameras into 3D object detection for the first time. We leverage their high temporal resolution and low bandwidth to enable high-speed 3D object detection. Our method enables detection even during inter-frame intervals when synchronized data is unavailable, by retrieving previous 3D information through the event camera. Furthermore, we introduce the first event-based 3D object detection datasets, Ev-Waymo and DSEC-3DOD, both of which include ground truth 3D bounding boxes at 100 FPS, establishing the first benchmarks for event-based 3D detectors. The code and dataset are available at https://github.com/mickeykang16/Ev3DOD. Hoonhee Cho, Jae-Young Kang, Kuk-Jin Yoon |
CVPR | 1 |
| 2025 | Learning Large Motion Estimation from Intermediate Representations with a High-Resolution Optical Flow Dataset Featuring Long-Range Dynamic Motion
Hoonhee Cho, Yuhwan Jeong, Kuk-Jin Yoon |
ICCV | 1 |
| 2025 | Unleashing the Temporal Potential of Stereo Event Cameras for Continuous-Time 3D Object Detectionabstract3D object detection is essential for autonomous systems, enabling precise localization and dimension estimation. While LiDAR and RGB cameras are widely used, their fixed frame rates create perception gaps in high-speed scenarios. Event cameras, with their asynchronous nature and high temporal resolution, offer a solution by capturing motion continuously. The recent approach, which integrates event cameras with conventional sensors for continuous-time detection, struggles in fast-motion scenarios due to its dependency on synchronized sensors. We propose a novel stereo 3D object detection framework that relies solely on event cameras, eliminating the need for conventional 3D sensors. To compensate for the lack of semantic and geometric information in event data, we introduce a dual filter mechanism that extracts both. Additionally, we enhance regression by aligning bounding boxes with object-centric information. Experiments show that our method outperforms prior approaches in dynamic environments, demonstrating the potential of event cameras for robust, continuous-time 3D perception. The code is available at https://github.com/mickeykang16/Ev-Stereo3D. Jae-Young Kang, Hoonhee Cho, Kuk-Jin Yoon |
ICCV | 2 |
| 2025 | From Sharp to Blur: Unsupervised Domain Adaptation for 2D Human Pose Estimation Under Extreme Motion Blur Using Event Cameras
Hoonhee Cho, Kuk-Jin Yoon |
ICCV | 2 |
| 2025 | VR-Drive: Viewpoint-Robust End-to-End Driving with Feed-Forward 3D Gaussian SplattingabstractEnd-to-end autonomous driving (E2E-AD) has emerged as a promising paradigm that unifies perception, prediction, and planning into a holistic, data-driven framework. However, achieving robustness to varying camera viewpoints, a common real-world challenge due to diverse vehicle configurations, remains an open problem. In this work, we propose VR-Drive, a novel E2E-AD framework that addresses viewpoint generalization by jointly learning 3D scene reconstruction as an auxiliary task to enable planning-aware view synthesis. Unlike prior scene-specific synthesis approaches, VR-Drive adopts a feed-forward inference strategy that supports online training-time augmentation from sparse views without additional annotations. To further improve viewpoint consistency, we introduce a viewpoint-mixed memory bank that facilitates temporal interaction across multiple viewpoints and a viewpoint-consistent distillation strategy that transfers knowledge from original to synthesized views. Trained in a fully end-to-end manner, VR-Drive effectively mitigates synthesis-induced noise and improves planning under viewpoint shifts. In addition, we release a new benchmark dataset to evaluate E2E-AD performance under novel camera viewpoints, enabling comprehensive analysis. Our results demonstrate that VR-Drive is a scalable and robust solution for the real-world deployment of end-to-end autonomous driving systems. Hoonhee Cho, Jae-Young Kang, Giwon Lee, Hyemin Yang, Heejun Park, Seokwoo Jung, Kuk-Jin Yoon |
NeurIPS | 1 |
| 2025 | Unifying Low-Resolution and High-Resolution Alignment by Event Cameras for Space-Time Video Super-ResolutionabstractEvent cameras deliver asynchronous pixel intensity changes, which result in sparse event data that offers the advantages of high temporal resolution. These high temporal characteristics make researchers naturally incorporate event cameras into video frame interpolation (VFI) and video super-resolution (VSR). In this paper, we make the first attempt to solve the space-time video super-resolution (STVSR) task effectively, addressing both VFI and VSR simultaneously, by leveraging temporally dense events. STVSR aims to generate intermediate high-resolution (HR) videos between consecutive low-resolution (LR) frames. To fully exploit the high temporal frequency of events for STVSR, we focus on temporal alignment in two stages, at low-resolution and after up-sampling in high-resolution. In temporal alignment at low-resolution, to upsample spatial dimensions effectively, we leverage high temporal features to preserve spatial context. On the other hand, for temporal alignment at the high-resolution stage, we employ a deformable sampling process from events to achieve accurate alignment with forward and backward directions. In addition, we provide the SuperREST dataset, which features high-frequency details and complex motion in an RGB-Event setup. Experimental results on several datasets demonstrate that our method achieves a significant performance gain on STVSR tasks with low computational cost. Our codes and datasets are available at h t t ps: //github.com/Chohoonhee/ESTNet. Hoonhee Cho, Jae-Young Kang, Taewoo Kim 0003, Yuhwan Jeong, Kuk-Jin Yoon |
WACV | 1 |
| 2024 | Frequency-Aware Event-Based Video Deblurring for Real-World Motion BlurabstractVideo deblurring aims to restore sharp frames from blurred video clips. Despite notable progress in video deblurring works, it is still a challenging problem because of the loss of motion information during the duration of the exposure time. Since event cameras can capture clear motion information asynchronously with high temporal resolution, several works exploit the event camera for deblurring as they can provide abundant motion information. However, despite these approaches, there were few cases of actively exploiting the long-range temporal dependency of videos. To tackle these deficiencies, we present an event-based video deblurring framework by actively utilizing temporal information from videos. To be specific, we first introduce a frequency-based cross-modal feature enhancement module. Second, we propose event-guided video alignment modules by considering the valuable characteristics of the event and videos. In addition, we designed a hybrid camera system to collect the first real-world event-based video deblurring dataset. For the first time, we build a dataset containing synchronized high-resolution real-world blurred videos and corresponding sharp videos and event streams. Experimental results validate that our frameworks significantly outperform the state-of-the-art frame-based and event-based deblurring works in the various datasets. The project pages are available at https://sites.google.com/view/fevd-cvpr2024. Taewoo Kim 0003, Hoonhee Cho, Kuk-Jin Yoon |
CVPR | 2 |
| 2024 | TTA-EVF: Test-Time Adaptation for Event-based Video Frame Interpolation via Reliable Pixel and Sample EstimationabstractVideo Frame Interpolation (VFI), which aims at gener-ating high-frame-rate videos from low-frame-rate inputs, is a highly challenging task. The emergence of bio-inspired sensors known as event cameras, which boast microsecond-level temporal resolution, has ushered in a transformative era for VFI. Nonetheless, the application of event-based VFI techniques in domains with distinct environments from the training data can be problematic. This is mainly because event camera data distribution can undergo substan-tial variations based on camera settings and scene conditions, presenting challenges for effective adaptation. In this paper, we propose a test-time adaptation method for event-based VFI to address the gap between the source and target domains. Our approach enables sequential learning in an online manner on the target domain, which only provides low-frame-rate videos. We present an approach that lever-ages confident pixels as pseudo ground-truths, enabling stable and accurate online learning from low-frame-rate videos. Furthermore, to prevent overfitting during the con-tinuous online process where the same scene is encountered repeatedly, we propose a method of blending historical sam-ples with current scenes. Extensive experiments validate the effectiveness of our method, both in cross-domain and con-tinuous domain shifting setups. The code is available at https://github.com/Chohoonhee/TTA-EVF. Hoonhee Cho, Taewoo Kim 0003, Yuhwan Jeong, Kuk-Jin Yoon |
CVPR | 1 |
| 2024 | Temporal Event Stereo via Joint Learning with Stereoscopic Flow
Hoonhee Cho, Jae-Young Kang, Kuk-Jin Yoon |
ECCV (35) | 1 |
| 2024 | Finding Meaning in Points: Weakly Supervised Semantic Segmentation for Event Cameras
Hoonhee Cho, Sung-Hoon Yoon 0001, Hyeokjun Kweon, Kuk-Jin Yoon |
ECCV (40) | 1 |
| 2024 | Towards Robust Event-Based Networks for Nighttime via Unpaired Day-to-Night Event Translation
Yuhwan Jeong, Hoonhee Cho, Kuk-Jin Yoon |
ECCV (67) | 2 |
| 2024 | CMTA: Cross-Modal Temporal Alignment for Event-Guided Video Deblurring
Taewoo Kim 0003, Hoonhee Cho, Kuk-Jin Yoon |
ECCV (52) | 2 |
| 2024 | Towards Real-World Event-Guided Low-Light Video Enhancement and Deblurring
Taewoo Kim 0003, Jaeseok Jeong 0001, Hoonhee Cho, Yuhwan Jeong, Kuk-Jin Yoon |
ECCV (12) | 3 |
| 2024 | A Benchmark Dataset for Event-Guided Human Pose Estimation and Tracking in Extreme ConditionsabstractMulti-person pose estimation and tracking have been actively researched by the computer vision community due to their practical applicability. However, existing human pose estimation and tracking datasets have only been successful in typical scenarios, such as those without motion blur or with well-lit conditions. These RGB-based datasets are limited to learning under extreme motion blur situations or poor lighting conditions, making them inherently vulnerable to such scenarios.As a promising solution, bio-inspired event cameras exhibit robustness in extreme scenarios due to their high dynamic range and micro-second level temporal resolution. Therefore, in this paper, we introduce a new hybrid dataset encompassing both RGB and event data for human pose estimation and tracking in two extreme scenarios: low-light and motion blur environments. The proposed Event-guided Human Pose Estimation and Tracking in eXtreme Conditions (EHPT-XC) dataset covers cases of motion blur caused by dynamic objects and low-light conditions individually as well as both simultaneously. With EHPT-XC, we aim to inspire researchers to tackle pose estimation and tracking in extreme conditions by leveraging the advantageous of the event camera. Project pages are available at https://github.com/Chohoonhee/EHPT-XC. Hoonhee Cho, Taewoo Kim 0003, Yuhwan Jeong, Kuk-Jin Yoon |
NeurIPS | 1 |
| 2023 | Learning Adaptive Dense Event Stereo from the Image DomainabstractRecently, event-based stereo matching has been studied due to its robustness in poor light conditions. However, existing event-based stereo networks suffer severe performance degradation when domains shift. Unsupervised domain adaptation (UDA) aims at resolving this problem without using the target domain ground-truth. However, traditional UDA still needs the input event data with ground- truth in the source domain, which is more challenging and costly to obtain than image data. To tackle this issue, we propose a novel unsupervised domain Adaptive Dense Event Stereo (ADES), which resolves gaps between the different domains and input modalities. The proposed ADES framework adapts event-based stereo networks from abundant image datasets with ground-truth on the source domain to event datasets without ground-truth on the target domain, which is a more practical setup. First, we propose a self-supervision module that trains the network on the target domain through image reconstruction, while an artifact prediction network trained on the source domain assists in removing intermittent artifacts in the reconstructed image. Secondly, we utilize the feature-level normalization scheme to align the extracted features along the epipolar line. Finally, we present the motion-invariant consistency module to impose the consistent output between the perturbed motion. Our experiments demonstrate that our approach achieves remarkable results in the adaptation ability of event-based stereo matching from the image domain. Hoonhee Cho, Jegyeong Cho, Kuk-Jin Yoon |
CVPR | 1 |
| 2023 | Non-Coaxial Event-guided Motion Deblurring with Spatial AlignmentabstractMotion deblurring from a blurred image is a challenging computer vision problem because frame-based cameras lose information during the blurring process. Several attempts have compensated for the loss of motion information by using event cameras, which are bio-inspired sensors with a high temporal resolution. Even though most studies have assumed that image and event data are pixel-wise aligned, this is only possible with low-quality active-pixel sensor (APS) images and synthetic datasets. In real scenarios, obtaining per-pixel aligned event-RGB data is technically challenging since event and frame cameras have different optical axes. For the application of the event camera, we propose the first Non-coaxial Event-guided Image Deblurring (NEID) approach that utilizes the camera setup composed of a standard frame-based camera with a non-coaxial single event camera. To consider the per-pixel alignment between the image and event without additional devices, we propose the first NEID network that spatially aligns events to images while refining the image features from temporally dense event features. For training and evaluation of our network, we also present the first large-scale dataset, consisting of RGB frames with non-aligned events aimed at a breakthrough in motion deblurring with an event camera. Extensive experiments on various datasets demonstrate that the proposed method achieves significantly better results than the prior works in terms of performance and speed, and it can be applied for practical uses of event cameras. Hoonhee Cho, Yuhwan Jeong, Taewoo Kim 0003, Kuk-Jin Yoon |
ICCV | 1 |
| 2023 | Label-Free Event-based Object Recognition via Joint Learning with Image Reconstruction from EventsabstractRecognizing objects from sparse and noisy events becomes extremely difficult when paired images and category labels do not exist. In this paper, we study label-free event-based object recognition where category labels and paired images are not available. To this end, we propose a joint formulation of object recognition and image reconstruction in a complementary manner. Our method first reconstructs images from events and performs object recognition through Contrastive Language-Image Pretraining (CLIP), enabling better recognition through a rich context of images. Since the category information is essential in reconstructing images, we propose category-guided attraction loss and category-agnostic repulsion loss to bridge the textual features of predicted categories and the visual features of reconstructed images using CLIP. Moreover, we introduce a reliable data sampling strategy and local-global reconstruction consistency to boost joint learning of two tasks. To enhance the accuracy of prediction and quality of reconstruction, we also propose a prototype-based approach using unpaired images. Extensive experiments demonstrate the superiority of our method and its extensibility for zero-shot object recognition. Our project code is available at https://github.com/Chohoonhee/Ev-LaFOR. Hoonhee Cho, Hyeonseong Kim, Yujeong Chae, Kuk-Jin Yoon |
ICCV | 1 |
| 2023 | Efficient Reference-based Video Super-Resolution (ERVSR): Single Reference Image Is All You NeedabstractReference-based video super-resolution (RefVSR) is a promising domain of super-resolution that recovers high-frequency textures of a video using reference video. The multiple cameras with different focal lengths in mobile devices aid recent works in RefVSR, which aim to super-resolve a low-resolution ultra-wide video by utilizing wide-angle videos. Previous works in RefVSR used all reference frames of a Ref video at each time step for the super-resolution of low-resolution videos. However, computation on higher-resolution images increases the runtime and memory consumption, hence hinders the practical application of RefVSR. To solve this problem, we propose an Efficient Reference-based Video Super-Resolution (ERVSR) that exploits a single reference frame to super-resolve whole low-resolution video frames. We introduce an attention-based feature align module and an aggregation upsampling module that attends LR features using the correlation between the reference and LR frames. The proposed ERVSR achieves 12× faster speed, 1/4 memory consumption than previous state-of-the-art RefVSR networks, and competitive performance on the RealMCVSR dataset while using a single reference image. Youngrae Kim 0001, Jinsu Lim, Hoonhee Cho, Dongman Lee, Kuk-Jin Yoon, Ho-Jin Choi |
WACV | 3 |
| 2022 | Event-Image Fusion Stereo Using Cross-Modality Feature PropagationabstractEvent cameras asynchronously output the polarity values of pixel-level log intensity alterations. They are robust against motion blur and can be adopted in challenging light conditions. Owing to these advantages, event cameras have been employed in various vision tasks such as depth estimation, visual odometry, and object detection. In particular, event cameras are effective in stereo depth estimation to find correspondence points between two cameras under challenging illumination conditions and/or fast motion. However, because event cameras provide spatially sparse event stream data, it is difficult to obtain a dense disparity map. Although it is possible to estimate disparity from event data at the edge of a structure where intensity changes are likely to occur, estimating the disparity in a region where event occurs rarely is challenging. In this study, we propose a deep network that combines the features of an image with the features of an event to generate a dense disparity map. The proposed network uses images to obtain spatially dense features that are lacking in events. In addition, we propose a spatial multi-scale correlation between two fused feature maps for an accurate disparity map. To validate our method, we conducted experiments using synthetic and real-world datasets. Hoonhee Cho, Kuk-Jin Yoon |
AAAI | 1 |
| 2022 | Selection and Cross Similarity for Event-Image Deep Stereo
Hoonhee Cho, Kuk-Jin Yoon |
ECCV (32) | 1 |