VLDB 2026 Research / reviewers in the wild / expert
Sungmin Eum
dblp:42/10859
· DBLP profile ↗
24ranked-venue papers
9as first author
8since 2021 · last 2026
0000-0002-0767-0395ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery Using Gaussian SplattingabstractDespite significant advancements in dynamic neural rendering, existing methods fail to address the unique challenges posed by UAV-captured scenarios, particularly those involving monocular camera setups, top-down perspective, and multiple small, moving humans, which are not adequately represented in existing datasets. In this work, we introduce UAV4D, a framework for enabling photorealistic rendering for dynamic real-world scenes captured by UAVs. Specifically, we address the challenge of reconstructing dynamic scenes with multiple moving pedestrians from monocular video data without the need for additional sensors. We use a combination of a 3D foundation model and a human mesh reconstruction model to reconstruct both the scene background and humans. We propose a novel approach to resolve the scene scale ambiguity and place both humans and the scene in world coordinates by identifying human-scene contact points. Additionally, we exploit the SMPL model and background mesh to initialize Gaussian splats, enabling holistic scene rendering. We evaluated our method on three complex UAV-captured datasets: VisDrone, Manipal-UAV, and Okutama-Action, each with distinct characteristics and 10-50 humans. Our results demonstrate the benefits of our approach over existing methods in novel view synthesis, achieving a 1.5 dB PSNR improvement and superior visual sharpness. Dongki Jung, Christopher Maxey, Sungmin Eum, Yonghan Lee 0001, Dinesh Manocha, Heesung Kwon |
AAAI | 4 |
| 2026 | MoRe: Monocular Geometry Refinement via Graph Optimization for Cross-View ConsistencyabstractMonocular 3D foundation models offer an extensible solution for perception tasks, making them attractive for broader 3D vision applications. In this paper, we propose MoRe, a training-free Monocular Geometry Refinement method designed to improve cross-view consistency and achieve scale alignment. To induce inter-frame relationships, our method employs feature matching between frames to establish correspondences. Rather than applying simple least squares optimization on these matched points, we formulate a graph-based optimization framework that performs local planar approximation using the estimated 3D points and surface normals estimated by monocular foundation models. This formulation addresses the scale ambiguity inherent in monocular geometric priors while preserving the underlying 3D structure. We further demonstrate that MoRe not only enhances 3D reconstruction but also improves novel view synthesis, particularly in sparse-view rendering scenarios. Dongki Jung, Yonghan Lee 0001, Sungmin Eum, Heesung Kwon, Dinesh Manocha |
WACV | 4 |
| 2026 | SynPlay: Large-Scale Synthetic Human Data with Real-World Diversity for Aerial-View PerceptionabstractWe introduce SynPlay, a large-scale synthetic human dataset purpose-built for advancing multi-perspective human localization, with a predominant focus on aerial-view perception. SynPlay departs from traditional synthetic datasets by addressing a critical but underexplored challenge: localizing humans in aerial scenes where subjects often occupy only tens of pixels in the image. In such scenarios, fine-grained details like facial features or textures become irrelevant, shifting the burden of recognition to human motion, behavior, and interactions. To meet this need, SynPlay implements a novel rule-guided motion generation framework that combines real-world motion capture with motion evolution graphs. This design enables human actions to evolve dynamically through high-level game rules rather than predefined scripts, resulting in effectively uncountable motion variations. Unlike existing synthetic datasets—which either focus on static visual traits or reuse a limited set of mocap-driven actions—SynPlay captures a wide spectrum of spontaneous behaviors, including complex interactions that naturally emerge from unscripted gameplay scenarios. SynPlay also introduces an extensive multi-camera setup that spans UAVs at random altitudes, CCTVs, and a freely roaming UGV, achieving true near-to-far perspective coverage in a single dataset. The majority of instances are captured from aerial viewpoints at varying scales, directly supporting the development of models for long-range human analysis—a setting where existing datasets fall short. Our data contains over 73k images and 6.5M human instances, with detailed annotations for detection, segmentation, and keypoint tasks. Extensive experiments demonstrate that training with SynPlay significantly improves human localization performance, especially in few-shot and data-scarce scenarios. Jinsub Yim, Hyungtae Lee, Sungmin Eum, Yi-Ting Shen, Heesung Kwon, Shuvra S. Bhattacharyya |
WACV | 3 |
| 2025 | Autocompose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMsabstractComposed pose retrieval (CPR) enables users to search for human poses by specifying a reference pose and a transition description, but progress in this field is hindered by the scarcity and inconsistency of annotated pose transitions. Existing CPR datasets rely on costly human annotations or heuristic-based rule generation, both of which limit scalability and diversity. In this work, we introduce AutoComPose, the first framework that leverages multimodal large language models (MLLMs) to automatically generate rich and structured pose transition descriptions. Our method enhances annotation quality by structuring transitions into fine-grained body part movements and introducing mirrored/swapped variations, while a cyclic consistency constraint ensures logical coherence between forward and reverse transitions. To advance CPR research, we construct and release two dedicated benchmarks, AIST-CPR and PoseFixCPR, supplementing prior datasets with enhanced attributes. Extensive experiments demonstrate that training retrieval models with AutoComPose yields superior performance over human-annotated and heuristic-based methods, significantly reducing annotation costs while improving retrieval quality. Our work pioneers the automatic annotation of pose transitions, establishing a scalable foundation for future CPR research. Yi-Ting Shen, Sungmin Eum, Doheon Lee, Rohit Shete, Chiao-Yi Wang, Heesung Kwon, Shuvra S. Bhattacharyya |
ICCV | 2 |
| 2024 | Two Teachers Are Better Than One: Leveraging Depth In Training Only For Unsupervised Obstacle SegmentationabstractWe present a novel unsupervised obstacle segmentation architecture that follows a novel Relation Distillation (RD) paradigm. Our architecture design was inspired by a self-supervised teacher-student approach that relies on the Semantic Distillation originally devised for representation learning. While the teacher in the Semantic Distillation considers a single patch at a time, the teacher within RD takes a ‘pair of patches’ instead to transfer the local Semantic Co-occurrence Localization (SCooL) relationship that focuses more on the segmentation-boosting signals. To further improve the proposed architecture, we introduce the utilization of another teacher that leverages the depth information which inherently separates the entities at different physical distances, often tied with the boundaries of the obstacles. As the depth is distilled towards the student network only at the time of training, it adds zero computational/hardware cost at run-time. As no relevant public dataset is available, we have curated the Avoiding Obstacles In unstructured Driving (AvOID) dataset as a new testbed for unsupervised obstacle segmentation. We have validated that both the Relation Distillation and depth contribute to boosting the no-annotation segmentation performance on AvOID and KITTI-Obstacles. Sungmin Eum, Hyungtae Lee, Heesung Kwon, Philip R. Osteen, Andre Harrison |
IROS | 1 |
| 2022 | Negative Samples are at Large: Leveraging Hard-Distance Elastic Loss for Re-identification
Hyungtae Lee, Sungmin Eum, Heesung Kwon |
ECCV (24) | 2 |
| 2022 | MA3: Model-Accuracy Aware Anytime Planning with Simulation Verification for Navigating Complex TerrainsabstractOff-road and unstructured environments often contain complex patches of various types of terrain, rough elevation changes, deformable objects, etc. An autonomous ground vehicle traversing such environments experiences physical interactions that are extremely hard to model at scale and thus very hard to predict. Nevertheless, planning a safely traversable path through such an environment requires the ability to predict the outcomes of these interactions instead of avoiding them. One approach to doing this is to learn the interaction model offline based on collected data. Unfortunately, though, this requires large amounts of data and can often be brittle. Alternatively, models using physics-based simulators can generate large data and provide a reliable prediction. However, they are very slow to query online within the planning loop. This work proposes an algorithmic framework that utilizes the combination of a learned model and a physics-based simulation model for fast planning. Specifically, it uses the learned model as much as possible to accelerate planning while sparsely using the physics-based simulator to verify the feasibility of the planned path. We provide a theoretical analysis of the algorithm and its empirical evaluation showing a significant reduction in planning times. Manash Pratim Das, Damon M. Conover, Sungmin Eum, Heesung Kwon, Maxim Likhachev |
SOCS | 3 |
| 2022 | Exploring Cross-Domain Pretrained Model for Hyperspectral Image ClassificationabstractA pretrain-finetune strategy is widely used to reduce the overfitting that can occur when data are insufficient for convolutional neural network (CNN) training. The first few layers of a CNN pretrained on a large-scale RGB dataset are capable of acquiring general image characteristics, which are remarkably effective in tasks targeted for different RGB datasets. However, when it comes down to the hyperspectral domain where each domain has its unique spectral properties, the pretrain-finetune strategy no longer can be deployed in a conventional way while presenting three major issues: 1) inconsistent spectral characteristics among the domains (e.g., frequency range); 2) inconsistent number of data channels among the domains; and 3) absence of large-scale hyperspectral dataset. We seek to train a universal cross-domain model, which can later be deployed for various spectral domains. To achieve, we physically furnish multiple inlets to the model while having a universal portion, which is designed to handle the inconsistent spectral characteristics among different domains. Note that only the universal portion is used in the finetune process. This approach naturally enables the learning of our model on multiple domains simultaneously, which acts as an effective workaround for the issue of the absence of large-scale dataset. We have carried out a study to extensively compare models that were trained using cross-domain approach with ones trained from scratch. Our approach was found to be superior both in accuracy and training efficiency. In addition, we have verified that our approach effectively reduces the overfitting issue, enabling us to deepen the model up to 13 layers (from 9) without compromising the accuracy. Hyungtae Lee, Sungmin Eum, Heesung Kwon |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | S-DOD-CNN: Doubly Injecting Spatially-Preserved Object Information for Event RecognitionabstractWe present a novel event recognition approach called Spatially-preserved Doubly-injected Object Detection CNN (S-DOD-CNN), which incorporates the spatially preserved object detection information in both a direct and an indirect way. Indirect injection is carried out by simply sharing the weights between the object detection modules and the event recognition module. Meanwhile, our novelty lies in the fact that we have preserved the spatial information for the direct injection. Once multiple regions-of-intereset (RoIs) are acquired, their feature maps are computed and then projected onto a spatially-preserving combined feature map using one of the four Rol Projection approaches we present. In our architecture, combined feature maps are generated for object detection which are directly injected to the event recognition module. Our method provides the state-of-the-art accuracy for malicious event recognition. Hyungtae Lee, Sungmin Eum, Heesung Kwon |
ICASSP | 2 |
| 2020 | Semantics to Space(S2S): Embedding semantics into spatial space for zero-shot verb-object query inferencingabstractWe present a novel deep zero-shot learning (ZSL) model for inferencing human-object-interaction with verb-object (VO) query. While the previous two-stream ZSL approaches only use the semantic/textual information to be fed into the query stream, we seek to incorporate and embed the semantics into the visual representation stream as well. Our approach is powered by Semantics-to-Space (S2S) architecture where semantics derived from the residing objects are embedded into a spatial space of the visual stream. This architecture allows the co-capturing of the semantic attributes of the human and the objects along with their location/size/silhouette information. To validate, we have constructed a new dataset, Verb-Transferability 60 (VT60). VT60 provides 60 different VO pairs with overlapping verbs tailored for testing two-stream ZSL approaches with VO query. Experimental evaluations show that our approach not only outperforms the state-of-the-art, but also shows the capability of consistently improving performance regardless of which ZSL baseline architecture is used. Sungmin Eum, Heesung Kwon |
ICPR | 1 |
| 2020 | ME R-CNN: Multi-Expert R-CNN for Object DetectionabstractWe introduce Multi-Expert Region-based Convolutional Neural Network (ME R-CNN) which is equipped with multiple experts (ME) where each expert is learned to process a certain type of regions of interest (RoIs). This architecture better captures the appearance variations of the RoIs caused by different shapes, poses, and viewing angles. In order to direct each RoI to the appropriate expert, we devise a novel "learnable" network, which we call, expert assignment network (EAN). EAN automatically learns the optimal RoI-expert relationship even without any supervision of expert assignment. As the major components of ME R-CNN, ME and EAN, are mutually affecting each other while tied to a shared network, neither an alternating nor a naive end-to-end optimization is likely to fail. To address this problem, we introduce a practical training strategy which is tailored to optimize ME, EAN, and the shared network in an end-to-end fashion. We show that both of the architectures provide considerable performance increase over the baselines on PASCAL VOC 07, 12, and MS COCO datasets. Hyungtae Lee, Sungmin Eum, Heesung Kwon |
IEEE Trans. Image Process. | 2 |
| 2019 | Object and Text-guided Semantics for CNN-based Activity RecognitionabstractMany previous methods have demonstrated the importance of considering semantically relevant objects for carrying out video-based human activity recognition, yet none of the methods have harvested the power of large text corpora to relate the objects and the activities to be transferred into learning a unified deep convolutional neural network. We present a novel activity recognition CNN which co-learns the object recognition task in an end-to-end multitask learning scheme to improve upon the baseline activity recognition performance. We further improve upon the multitask learning approach by exploiting a text-guided semantic space to select the most relevant objects with respect to the target activities. To the best of our knowledge, we are the first to investigate this approach. Sungmin Eum, Christopher Reale, Heesung Kwon, Claire Bonial, Clare R. Voss |
ICASSP | 1 |
| 2019 | DOD-CNN: Doubly-injecting Object Information for Event RecognitionabstractRecognizing an event in an image can be enhanced by detecting relevant objects in two ways: 1) indirectly utilizing object detection information within the unified architecture or 2) directly making use of the object detection output results. We introduce a novel approach, referred to as Doubly-injected Object Detection CNN (DOD-CNN), exploiting the object information in both ways for the task of event recognition. The structure of this network is inspired by the Integrated Object Detection CNN (IOD-CNN) where object information is indirectly exploited by the event recognition module through the shared portion of the network. In the DOD-CNN architecture, the intermediate object detection outputs are directly injected into the event recognition network while keeping the indirect sharing structure inherited from the IOD-CNN, thus being `doubly-injected'. We also introduce a batch pooling layer which constructs one representative feature map from multiple object hypotheses. We have demonstrated the effectiveness of injecting the object detection information in two different ways in the task of malicious event recognition. Hyungtae Lee, Sungmin Eum, Heesung Kwon |
ICASSP | 2 |
| 2019 | Is Pretraining Necessary for hyperspectral image classification?abstractWe address two questions for training a convolutional neural network (CNN) for hyperspectral image classification: i) is it possible to build a pre-trained network? and ii) is the pre-training effective in furthering the performance? To answer the first question, we have devised an approach that pre-trains a network on multiple source datasets that differ in their hyperspectral characteristics and fine-tunes on a target dataset. This approach effectively resolves the architectural issue that arises when transferring meaningful information between the source and the target networks. To answer the second question, we carried out several ablation experiments. Based on the experimental results, a network trained from scratch performs as good as a network fine-tuned from a pre-trained network. However, we observed that pre-training the network has its own advantage in achieving better performances when deeper networks are required. Hyungtae Lee, Sungmin Eum, Heesung Kwon |
IGARSS | 2 |
| 2019 | A RUGD Dataset for Autonomous Navigation and Visual Perception in Unstructured Outdoor EnvironmentsabstractResearch in autonomous driving has benefited from a number of visual datasets collected from mobile platforms, leading to improved visual perception, greater scene understanding, and ultimately higher intelligence. However, this set of existing data collectively represents only highly structured, urban environments. Operation in unstructured environments, e.g., humanitarian assistance and disaster relief or off-road navigation, bears little resemblance to these existing data. To address this gap, we introduce the Robot Unstructured Ground Driving (RUGD) dataset with video sequences captured from a small, unmanned mobile robot traversing in unstructured environments. Most notably, this data differs from existing autonomous driving benchmark data in that it contains significantly more terrain types, irregular class boundaries, minimal structured markings, and presents challenging visual properties often experienced in off road navigation, e.g., blurred frames. Over 7, 000 frames of pixel-wise annotation are included with this dataset, and we perform an initial benchmark using state-of-the-art semantic segmentation architectures to demonstrate the unique challenges this data introduces as it relates to navigation tasks. Maggie B. Wigness, Sungmin Eum, John G. Rogers III, David K. Han, Heesung Kwon |
IROS | 2 |
| 2019 | Planar content selection in images and videos using frontalness
Sungmin Eum, David S. Doermann |
Pattern Recognit. Lett. | 1 |
| 2018 | Exploitation of Semantic Keywords for Malicious Event ClassificationabstractLearning an event classifier is challenging when the scenes are semantically different but visually similar. However, as humans, we typically handle such tasks painlessly by adding our background semantic knowledge. Motivated by this observation, we aim to provide an empirical study about how additional information such as semantic keywords can boost up the discrimination of such events. To demonstrate the validity of this study, we first construct a novel Malicious Crowd Dataset containing crowd images with two events, benign and malicious, which look visually similar. Note that the primary focus of this paper is not to provide the state-of-the-art performance on this dataset but to show the beneficial aspects of using semantically-driven keyword information. By leveraging crowd-sourcing platforms, such as Amazon Mechanical Turk, we collect semantic keywords associated with images and then subsequently identify a subset of keywords (e.g. police, fire, etc.) unique to specific events. We first show that by using recently introduced attention models, a naive CNN-based event classifier actually learns to primarily focus on local attributes associated with the discriminant semantic keywords identified by the Turks. We further show that incorporating the keyword-driven information into early-and late-fusion approaches can significantly enhance malicious event classification. Hyungtae Lee, Sungmin Eum, Joel Levis, Heesung Kwon, James Michaelis, Michael Kolodny |
ICASSP | 2 |
| 2018 | Cross-Domain CNN for Hyperspectral Image ClassificationabstractIn this paper, we address the dataset scarcity issue with the hyperspectral image classification. As only a few thousands of pixels are available for training, it is difficult to effectively learn high-capacity Convolutional Neural Networks (CNNs). To cope with this problem, we propose a novel cross-domain CNN containing the shared parameters which can co-learn across multiple hyperspectral datasets. the network also contains the non-shared portions designed to handle the dataset-specific spectral characteristics and the associated classification tasks. Our approach is the first attempt to learn a CNN for multiple hyperspectral datasets, in an end-to-end fashion. Moreover, we have experimentally shown that the proposed network trained on three of the widely used datasets outperform all the baseline networks which are trained on single dataset. Hyungtae Lee, Sungmin Eum, Heesung Kwon |
IGARSS | 2 |
| 2017 | IOD-CNN: Integrating object detection networks for event recognitionabstractMany previous methods have showed the importance of considering semantically relevant objects for performing event recognition, yet none of the methods have exploited the power of deep convolutional neural networks to directly integrate relevant object information into a unified network. We present a novel unified deep CNN architecture which integrates architecturally different, yet semantically-related object detection networks to enhance the performance of the event recognition task. Our architecture allows the sharing of the convolutional layers and a fully connected layer which effectively integrates event recognition, rigid object detection and non-rigid object detection. Sungmin Eum, Hyungtae Lee, Heesung Kwon, David S. Doermann |
ICIP | 1 |
| 2016 | Content selection using frontalness evaluation of multiple framesabstractThis paper addresses the problem of selecting instances of a planar object in a video or from a set of images based on an evaluation of its “frontalness”. We introduce the idea of “evaluating the frontalness” by computing how close the object's surface normal aligns with the optical axis of a camera. The unique and novel aspect of our method is that unlike previous planar object pose estimation methods, our method does not require the true frontal image as a reference. The intuition is that a true frontal image can be used to produce other non-frontal images by perspective projection, while the non-frontal images have limited ability to produce other non-frontal images. We show that this intuition of comparing ‘frontal’ and ‘non-frontal’ can be extended to comparing ‘more frontal’ and ‘less frontal’ images. Based on this observation, our method estimates the relative frontalness of an image by exploiting the objective space error. We also propose the usage of K-invariant space to evaluate the frontalness even when the camera intrinsic parameters are unknown (e.g., images/videos from the web). We show that our method outperforms the homography decomposition-based method which also does not require reference images. In addition, a qualitative evaluation is carried out to show that our method can be applied in selecting the most frontal characters from a set of images captured in various viewpoints. Sungmin Eum, David S. Doermann |
ICPR | 1 |
| 2015 | JH2R: Joint Homography Estimation for Highlight RemovalabstractImagine being in an art museum where there are paintings or pictures held inside glass-frames for protection. There are pieces which you wish to capture using a camera, but you experience difficulties avoiding highlights which are generated by indoor lighting reflected off the glossy surfaces. Similar problems occur when capturing contents off of whiteboards, documents printed on glossy surfaces, objects such as books or CDs with plastic covers. In this work, we address the problem of removing unwanted highlight regions in images generated by reflections of light sources on glossy surfaces. Although there have been efforts made to synthetically fill in the missing regions using the neighboring patterns by applying methods like inpainting [3, 4], it is impossible to recover the missing information in completely saturated regions. Therefore, we need to use multiple images where corresponding regions are not covered by the saturated highlights. Unlike other methods, our method uses the relationship between the highlight regions resulting in more robust removal of saturated highlights. Our method Overview Our method was motivated by a widely acknowledged physical phenomenon referred to as the ‘motion parallax’. Without loss of generality, we can similarly view the relationship between the desired content (e.g., a painting) and the highlights. Since the highlights caused by the light source are the result of the reflection on the glossy surface before they reach the camera, the light source can be modeled to virtually exist on the other side of the content. Note that, the distance from the light source is always larger than the distance from the content (D > d, in Figure 1). Sungmin Eum, Hyungtae Lee, David S. Doermann |
BMVC | 1 |
| 2014 | Sharpness-aware document image mosaicing using graphcutsabstractThere are numerous types of documents which are difficult to scan or capture in a single pass due to their physical size or the size of their content. One possible solution that has been proposed is mosaicing multiple overlapping images to capture the complete document. In this paper, we present a novel Graphcut-based document image mosaicing method which seeks to overcome the known limitations of the previous approaches. First, our method does not require any prior knowledge of the content of the given document images, making it more widely applicable and robust. Second, information regarding the geometrical disposition between the overlapping images is exploited to minimize the errors at the boundary regions. Third, our method incorporates a sharpness measure which induces cut generation in a way that results in the mosaic including the sharpest pixels. Our method is shown to outperform previous methods, both quantitatively and qualitatively. Sungmin Eum, David S. Doermann |
ICIP | 1 |
| 2013 | Enhancing Light Blob Detection for Intelligent Headlight Control Using Lane DetectionabstractIn this paper, we propose an enhanced method for detecting light blobs (LBs) for intelligent headlight control (IHC). The main function of the IHC system is to automatically convert high-beam headlights to low beam when vehicles are found in the vicinity. Thus, to implement the IHC, it is necessary to detect preceding or oncoming vehicles. Generally, this process of detecting vehicles is done by detecting LBs in the images. Previous works regarding LB detection can largely be categorized into two approaches by the image type they use: low-exposure (LE) images or autoexposure (AE) images. While they each have their own strengths and weaknesses, the proposed method combines them by integrating the use of the partial region of the AE image confined by the lane detection information and the LE image. Consequently, the proposed method detects headlights at various distances and taillights at close distances using LE images while handling taillights at distant locations by exploiting the confined AE images. This approach enhances the performance of detecting the distant LBs while maintaining low false detections. Sungmin Eum, Ho Gi Jung |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2012 | Recognizability assessment of facial images for automated teller machine applications
Jae Kyu Suhr, Sungmin Eum, Ho Gi Jung, Gen Li 0011, Gahyun Kim, Jaihie Kim |
Pattern Recognit. | 2 |