VLDB 2026 Research / reviewers in the wild / expert
Haesol Park
dblp:98/9446
· DBLP profile ↗
12ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0002-7615-6231ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IMPACT: Interpretable Most Important Person Analysis and Classification using Transformer-based ModelsabstractIdentifying the Most Important Person (MIP) in complex social and sports events remains a challenging problem due to the dynamic nature of group interactions, subtle visual cues, and context-dependent semantics. Traditional methods often struggle to accurately capture the interplay between individuals and the overarching activity, especially in unstructured real-world environments. In addition, the lack of strong supervision and the need for a deeper contextual understanding further complicate the task. In this work, we propose IMPACT, a novel multi-modal framework that leverages recent advances in vision language models to bridge the gap between visual perception and semantic reasoning. Our approach integrates structured scene understanding, natural language generation, and cross-modal learning to jointly model activity recognition and MIP localization. The method integrates language, vision, and spatial reasoning to improve scene interpretability as well as accuracy in group activity recognition tasks. By incorporating language-based representations, the proposed method enables interpretable and robust performance in sports-centric group activity scenarios. Comprehensive experiments on C-Sports and NCAA datasets demonstrate that the framework significantly enhances the localization of key individuals as well as the accuracy of activity prediction, laying the groundwork for a holistic scene understanding in human-centric video and image analysis. Our proposed method achieves an accuracy of 81.6% when compared with human annotator markings and an increase in mAP scores by ∼ 5% for MIP identification. Akshat Rampuria, Kamakshya Prasad Nayak, Thakare Kamalakar Vijay, Tushar Joshi, Aditya Dhananjay Singh, Haesol Park, Heeseung Choi, Hyungjoo Jung, Debi Prosad Dogra, Ig-Jae Kim |
WACV | 6 |
| 2026 | A foundational research framework for real-world abandoned object detection: train-free baseline and a standardized benchmark
Dong-Bum Kim, Deok-Hyun Ahn, Yong-Jin Jo, Haesol Park, Sangyoun Lee, Haksub Kim |
Expert Syst. Appl. | 4 |
| 2025 | GMAR: Gradient-Driven Multi-Head Attention Rollout for Vision Transformer InterpretabilityabstractThe Vision Transformer (ViT) has made significant advancements in computer vision, utilizing self-attention mechanisms to achieve state-of-the-art performance across various tasks, including image classification, object detection, and segmentation. Its architectural flexibility and capabilities have made it a preferred choice among researchers and practitioners. However, the intricate multi-head attention mechanism of ViT presents significant challenges to interpretability, as the underlying prediction process remains opaque. A critical limitation arises from an observation commonly noted in transformer architectures: "Not all attention heads are equally meaningful." Overlooking the relative importance of specific heads highlights the limitations of existing interpretability methods. To address these challenges, we introduce Gradient-Driven Multi-Head Attention Rollout (GMAR), a novel method that quantifies the importance of each attention head using gradient-based scores. These scores are normalized to derive a weighted aggregate attention score, effectively capturing the relative contributions of individual heads. GMAR clarifies the role of each head in the prediction process, enabling more precise interpretability at the head level. Experimental results demonstrate that GMAR consistently outperforms traditional attention rollout techniques. This work provides a practical contribution to transformer-based architectures, establishing a robust framework for enhancing the interpretability of Vision Transformer models. Sehyeong Jo, Gangjae Jang, Haesol Park |
ICIP | 3 |
| 2025 | MAIR++: Improving Multi-View Attention Inverse Rendering With Implicit Lighting RepresentationabstractIn this paper, we propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, SVBRDF, and 3D spatially-varying lighting. While multi-view images have been widely used for object-level inverse rendering, scene-level inverse rendering has primarily been studied using single-view images due to the lack of a dataset containing high dynamic range multi-view images with ground-truth geometry, material, and spatially-varying lighting. To improve the quality of scene-level inverse rendering, a novel framework called Multi-view Attention Inverse Rendering (MAIR) was recently introduced. MAIR performs scene-level multi-view inverse rendering by expanding the OpenRooms dataset, designing efficient pipelines to handle multi-view images, and splitting spatially-varying lighting. Although MAIR showed impressive results, its lighting representation is fixed to spherical Gaussians, which limits its ability to render images realistically. Consequently, MAIR cannot be directly used in applications such as material editing. Moreover, its multi-view aggregation networks have difficulties extracting rich features because they only focus on the mean and variance between multi-view features. In this paper, we propose its extended version, called MAIR++. MAIR++ addresses the aforementioned limitations by introducing an implicit lighting representation that accurately captures the lighting conditions of an image while facilitating realistic rendering. Furthermore, we design a directional attention-based multi-view aggregation network to infer more intricate relationships between views. Experimental results show that MAIR++ not only outperforms MAIR and single-view-based methods but also demonstrates robust performance on unseen real-world scenes. Junyong Choi, SeokYeong Lee, Haesol Park, Seung-Won Jung, Ig-Jae Kim, Junghyun Cho |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Enhancing Multi-view Pedestrian Detection Through Generalized 3D Feature PullingabstractThe main challenge in multi-view pedestrian detection is integrating view-specific features into a unified space for comprehensive end-to-end perception. Prior multi-view detection methods have focused on projecting perspective-view features onto the ground plane, creating a "bird’s eye view" (BEV) representation of the scene. This paper proposes a simple but effective architecture that utilizes a nonparametric 3D feature-pulling strategy. This strategy directly extracts the corresponding 2D features for each valid voxel within the 3D feature volume, addressing the feature loss that may arise in previous methods. The proposed framework introduces three novel modules, each crafted to bolster the generalization capabilities of multi-view detection systems. Through extensive experiments, the efficacy of the proposed model is demonstrated. The results show a new state-of-the-art accuracy, both in conventional scenarios and particularly in the context of scene generalization benchmarks. Sithu Aung, Haesol Park, Hyungjoo Jung, Junghyun Cho |
WACV | 2 |
| 2023 | MAIR: Multi-View Attention Inverse Rendering with 3D Spatially-Varying Lighting EstimationabstractWe propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, a SVBRDF, and 3D spatially-varying lighting. Because multi-view images provide a variety of information about the scene, multi-view images in object-level inverse rendering have been taken for granted. However, owing to the absence of multi-view HDR synthetic dataset, scene-level inverse rendering has mainly been studied using single-view image. We were able to successfully perform scene-level inverse rendering using multi-view images by expanding OpenRooms dataset and designing efficient pipelines to handle multi-view images, and splitting spatially-varying lighting. Our experiments show that the proposed method not only achieves better performance than single-view-based methods, but also achieves robust performance on unseen real-world scene. Also, our sophisticated 3D spatially-varying lighting volume allows for photorealistic object insertion in any 3D location. Junyong Choi, SeokYeong Lee, Haesol Park, Seung-Won Jung, Ig-Jae Kim, Junghyun Cho |
CVPR | 3 |
| 2022 | Weakly supervised Branch Network with Template Mask for Classifying Masses in 3D Automated Breast UltrasoundabstractAutomated breast ultrasound (ABUS) is being rapidly utilized for screening and diagnosing breast cancer. Breast masses, including cancers shown in ABUS scans, often appear as irregular hypoechoic areas that are hard to distinguish from background shadings. We propose a novel branch network architecture incorporating segmentation information of masses in the training process. By providing the spatial attention effect, the branch network boosts the performance of existing neural network classifiers, helping to learn meaningful features around the mass. For the segmentation information, we leverage the existing radiology reports without additional labeling efforts. The reports should include the characteristics of breast masses, such as shape and orientation, and a template mask can be created in a rule-based manner. Experimental results show that the proposed branch network with a template mask significantly improves the performance of existing classifiers. Daekyung Kim, Changmo Nam, Haesol Park, Mijung Jang, Kyong Joon Lee |
WACV | 3 |
| 2018 | Joint Blind Motion Deblurring and Depth Estimation of Light Field
Haesol Park, In Kyu Park, Kyoung Mu Lee |
ECCV (16) | 2 |
| 2017 | Joint Estimation of Camera Pose, Depth, Deblurring, and Super-Resolution from a Blurred Image SequenceabstractThe conventional methods for estimating camera poses and scene structures from severely blurry or low resolution images often result in failure. The off-the-shelf deblurring or super-resolution methods may show visually pleasing results. However, applying each technique independently before matching is generally unprofitable because this naive series of procedures ignores the consistency between images. In this paper, we propose a pioneering unified framework that solves four problems simultaneously, namely, dense depth reconstruction, camera pose estimation, super-resolution, and deblurring. By reflecting a physical imaging process, we formulate a cost minimization problem and solve it using an alternating optimization technique. The experimental results on both synthetic and real videos show high-quality depth maps derived from severely degraded images that contrast the failures of naive multi-view stereo methods. Our proposed method also produces outstanding deblurred and super-resolved images unlike the independent application or combination of conventional video deblurring, super-resolution methods. Haesol Park, Kyoung Mu Lee |
ICCV | 1 |
| 2017 | Look Wider to Match Image Patches With Convolutional Neural NetworksabstractWhen a human matches two images, the viewer has a natural tendency to view the wide area around the target pixel to obtain clues of right correspondence. However, designing a matching cost function that works on a large window in the same way is difficult. The cost function is typically not intelligent enough to discard the information irrelevant to the target pixel, resulting in undesirable artifacts. In this letter, we propose a novel convolutional neural network (CNN) module to learn a stereo matching cost with a large-sized window. Unlike conventional pooling layers with strides, the proposed per-pixel pyramid-pooling layer can cover a large area without a loss of resolution and detail. Therefore, the learned matching cost function can successfully utilize the information from a large area without introducing the fattening effect. The proposed method is robust despite the presence of weak textures, depth discontinuity, illumination, and exposure difference. The proposed method achieves near-peak performance on the Middlebury benchmark. Haesol Park, Kyoung Mu Lee |
IEEE Signal Process. Lett. | 1 |
| 2014 | Stereo reconstruction using high-order likelihoods
Ho Yub Jung, Haesol Park, In Kyu Park, Kyoung Mu Lee, Sang Uk Lee |
Comput. Vis. Image Underst. | 2 |
| 2011 | GPU-friendly multi-view stereo reconstruction using surfel representation and graph cuts
Ju Yong Chang, Haesol Park, In Kyu Park, Kyoung Mu Lee, Sang Uk Lee |
Comput. Vis. Image Underst. | 2 |