Hyungtae Lee

dblp:25/9990 · DBLP profile ↗
← Back
34ranked-venue papers
18as first author
13since 2021 · last 2026
0000-0002-0631-9894ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 12 first-author · 8 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 7 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SynPlay: Large-Scale Synthetic Human Data with Real-World Diversity for Aerial-View Perception
abstract
We introduce SynPlay, a large-scale synthetic human dataset purpose-built for advancing multi-perspective human localization, with a predominant focus on aerial-view perception. SynPlay departs from traditional synthetic datasets by addressing a critical but underexplored challenge: localizing humans in aerial scenes where subjects often occupy only tens of pixels in the image. In such scenarios, fine-grained details like facial features or textures become irrelevant, shifting the burden of recognition to human motion, behavior, and interactions. To meet this need, SynPlay implements a novel rule-guided motion generation framework that combines real-world motion capture with motion evolution graphs. This design enables human actions to evolve dynamically through high-level game rules rather than predefined scripts, resulting in effectively uncountable motion variations. Unlike existing synthetic datasets—which either focus on static visual traits or reuse a limited set of mocap-driven actions—SynPlay captures a wide spectrum of spontaneous behaviors, including complex interactions that naturally emerge from unscripted gameplay scenarios. SynPlay also introduces an extensive multi-camera setup that spans UAVs at random altitudes, CCTVs, and a freely roaming UGV, achieving true near-to-far perspective coverage in a single dataset. The majority of instances are captured from aerial viewpoints at varying scales, directly supporting the development of models for long-range human analysis—a setting where existing datasets fall short. Our data contains over 73k images and 6.5M human instances, with detailed annotations for detection, segmentation, and keypoint tasks. Extensive experiments demonstrate that training with SynPlay significantly improves human localization performance, especially in few-shot and data-scarce scenarios.
Jinsub Yim, Hyungtae Lee, Sungmin Eum, Yi-Ting Shen, Heesung Kwon, Shuvra S. Bhattacharyya
WACV2
2025 Diversifying Human Pose In Synthetic Data For Aerial-View Human Detection
abstract
Synthetic data generation has emerged as a promising solution to the data scarcity issue in aerial-view human detection. However, creating datasets that accurately reflect varying real-world human appearances—particularly diverse poses—remains challenging and labor-intensive. To address this, we propose SynPoseDiv, a novel framework that diversifies human poses within existing synthetic datasets. SynPoseDiv tackles two key challenges: generating realistic, diverse 3D human poses using a diffusion-based pose generator, and producing images of virtual characters in novel poses through a source-to-target image translator. The framework incrementally transitions characters into new poses using optimized pose sequences identified via Dijkstra’s algorithm. Experiments demonstrate that SynPoseDiv significantly improves detection accuracy across multiple aerial-view human detection benchmarks, especially in low-shot scenarios, and remains effective regardless of the training approach or dataset size.
Yi-Ting Shen, Hyungtae Lee, Heesung Kwon, Shuvra S. Bhattacharyya
ICIP2
2025 TK-Planes: Tiered K-Planes with High Dimensional Feature Vectors for Dynamic UAV-based Scenes
abstract
In this paper, we present a new approach to improve the neural rendering fidelity of in-the-wild unmanned aerial vehicle (UAV)-based scenes. Our formulation is designed for dynamic scenes, consisting of small moving objects or human actions in particular. We propose an extension of K-Planes Neural Radiance Field (NeRF), wherein our algorithm stores a set of tiered high dimensional feature vectors. The tiered feature vectors are generated to effectively model conceptual information about a scene as well as to be processed by an image decoder that transforms output feature maps into RGB images. Our technique leverages the information among both static and dynamic objects within a scene and is able to capture salient scene attributes of high altitude videos. We evaluate its performance on challenging datasets, including Okutama Action and UG2, and observe considerable improvement in accuracy over state of the art neural rendering methods.
Christopher Maxey, Yonghan Lee 0001, Hyungtae Lee, Dinesh Manocha, Heesung Kwon
IROS4
2024 MeshGS: Adaptive Mesh-Aligned Gaussian Splatting for High-Quality Rendering
Yonghan Lee 0001, Hyungtae Lee, Heesung Kwon, Dinesh Manocha
ACCV (9)3
2024 Exploring the Potential of Synthetic Data to Replace Real Data
abstract
The potential of synthetic data to replace real data creates a huge demand for synthetic data in data-hungry AI. This potential is even greater when synthetic data is used for training along with a small number of real images from domains other than the test domain. We find that this potential varies depending on (i) the number of cross-domain real images and (ii) the test set on which the trained model is evaluated. We introduce two new metrics, the train2test distance and $\mathrm{AP}_{\mathrm{t} 2 \mathrm{t}}$, to evaluate the ability of a cross-domain training set using synthetic data to represent the characteristics of test instances in relation to training performance. Using these metrics, we delve deeper into the factors that influence the potential of synthetic data and uncover some interesting dynamics about how synthetic data impacts training performance. We hope these discoveries will encourage more widespread use of synthetic data.
Hyungtae Lee, Heesung Kwon, Shuvra S. Bhattacharyya
ICIP1
2024 UAV-Sim: NeRF-based Synthetic Data Generation for UAV-based Perception
abstract
Tremendous variations coupled with large degrees of freedom in UAV-based imaging conditions lead to a significant lack of data in adequately learning UAV-based perception models. Using various synthetic renderers in conjunction with perception models is prevalent to create synthetic data to augment the learning in the ground-based imaging domain. However, severe challenges in the austere UAV-based domain require distinctive solutions to image synthesis for data augmentation. In this work, we leverage recent advancements in neural rendering to improve static and dynamic novel-view UAV-based image synthesis, especially from high altitudes, capturing salient scene attributes. Finally, we demonstrate a considerable performance boost is achieved when a state-of-the-art detection model is optimized primarily on hybrid sets of real and synthetic data instead of the real or synthetic data separately.
Christopher Maxey, Hyungtae Lee, Dinesh Manocha, Heesung Kwon
ICRA3
2024 Two Teachers Are Better Than One: Leveraging Depth In Training Only For Unsupervised Obstacle Segmentation
abstract
We present a novel unsupervised obstacle segmentation architecture that follows a novel Relation Distillation (RD) paradigm. Our architecture design was inspired by a self-supervised teacher-student approach that relies on the Semantic Distillation originally devised for representation learning. While the teacher in the Semantic Distillation considers a single patch at a time, the teacher within RD takes a ‘pair of patches’ instead to transfer the local Semantic Co-occurrence Localization (SCooL) relationship that focuses more on the segmentation-boosting signals. To further improve the proposed architecture, we introduce the utilization of another teacher that leverages the depth information which inherently separates the entities at different physical distances, often tied with the boundaries of the obstacles. As the depth is distilled towards the student network only at the time of training, it adds zero computational/hardware cost at run-time. As no relevant public dataset is available, we have curated the Avoiding Obstacles In unstructured Driving (AvOID) dataset as a new testbed for unsupervised obstacle segmentation. We have validated that both the Relation Distillation and depth contribute to boosting the no-annotation segmentation performance on AvOID and KITTI-Obstacles.
Sungmin Eum, Hyungtae Lee, Heesung Kwon, Philip R. Osteen, Andre Harrison
IROS2
2023 Progressive Transformation Learning for Leveraging Virtual Images in Training
abstract
To effectively interrogate UAV-based images for detecting objects of interest, such as humans, it is essential to acquire large-scale UAV-based datasets that include human instances with various poses captured from widely varying viewing angles. As a viable alternative to laborious and costly data curation, we introduce Progressive Transformation Learning (PTL), which gradually augments a training dataset by adding transformed virtual images with enhanced realism. Generally, a virtual2real transformation generator in the conditional GAN framework suffers from quality degradation when a large domain gap exists between real and virtual images. To deal with the domain gap, PTL takes a novel approach that progressively iterates the following three steps: 1) select a subset from a pool of virtual images according to the domain gap, 2) transform the selected virtual images to enhance realism, and 3) add the transformed virtual images to the training set while removing them from the pool. In PTL, accurately quantifying the domain gap is critical. To do that, we theoretically demonstrate that the feature representation space of a given object detector can be modeled as a multivariate Gaussian distribution from which the Mahalanobis distance between a virtual object and the Gaussian distribution of each object category in the representation space can be readily computed. Experiments show that PTL results in a substantial performance increase over the baseline, especially in the small data and the cross-domain regime.
Yi-Ting Shen, Hyungtae Lee, Heesung Kwon, Shuvra S. Bhattacharyya
CVPR2
2023 NEV-NCD: Negative Learning, Entropy, and Variance Regularization Based Novel Action Categories Discovery
abstract
Novel Categories Discovery (NCD) facilitates learning from a partially annotated label space and enables deep learning (DL) models to operate in an open-world setting by identifying and differentiating instances of novel classes based on the labeled data notions. One of the primary assumptions of NCD is that the novel label space is perfectly disjoint and can be equipartitioned, but it is rarely realized by most NCD approaches in practice. To better align with this assumption, we propose a novel single-stage joint optimization-based NCD method, Negative learning, Entropy, and Variance regularization NCD (NEV-NCD). We demonstrate the efficacy of NEV-NCD in previously unexplored NCD applications of video action recognition (VAR) with the public UCF101 dataset and a curated in-house partial action-space annotated multi-view video dataset. Further, we perform a thorough ablation study by varying the composition of final joint loss and associated hyper-parameters. During our experiments with UCF101 and multi-view action dataset, NEV-NCD achieves ≈ 83% classification accuracy in test instances of labeled data. NEV-NCD achieves ≈ 70% clustering accuracy over unlabeled data outperforming both naive baselines and state-of-the-art pseudo-labeling-based approaches by ≈ 40% and ≈ 3.5% over both datasets.
Zahid Hasan 0001, Masud Ahmed, Abu Zaher Md Faridee, Sanjay Purushotham, Heesung Kwon, Hyungtae Lee, Nirmalya Roy
ICIP6
2022 Negative Samples are at Large: Leveraging Hard-Distance Elastic Loss for Re-identification
Hyungtae Lee, Sungmin Eum, Heesung Kwon
ECCV (24)1
2022 Self-Supervised Contrastive Learning for Cross-Domain Hyperspectral Image Representation
abstract
Recently, self-supervised learning has attracted attention due to its remarkable ability to acquire meaningful representations for classification tasks without using semantic labels. This paper introduces a self-supervised learning framework suitable for hyperspectral images that are inherently challenging to annotate. The proposed framework architecture leverages cross-domain CNN [1], allowing for learning representations from different hyperspectral images with varying spectral characteristics and no pixel-level annotation. In the framework, cross-domain representations are learned via contrastive learning where neighboring spectral vectors in the same image are clustered together in a common representation space encompassing multiple hyperspectral images. In contrast, spectral vectors in different hyperspectral images are separated into distinct clusters in the space. To verify that the learned representation through contrastive learning is effectively transferred into a downstream task, we perform a classification task on hyperspectral images. The experimental results demonstrate the advantage of the proposed self-supervised representation over models trained from scratch or other transfer learning methods.
Hyungtae Lee, Heesung Kwon
ICASSP1
2022 Exploring Cross-Domain Pretrained Model for Hyperspectral Image Classification
abstract
A pretrain-finetune strategy is widely used to reduce the overfitting that can occur when data are insufficient for convolutional neural network (CNN) training. The first few layers of a CNN pretrained on a large-scale RGB dataset are capable of acquiring general image characteristics, which are remarkably effective in tasks targeted for different RGB datasets. However, when it comes down to the hyperspectral domain where each domain has its unique spectral properties, the pretrain-finetune strategy no longer can be deployed in a conventional way while presenting three major issues: 1) inconsistent spectral characteristics among the domains (e.g., frequency range); 2) inconsistent number of data channels among the domains; and 3) absence of large-scale hyperspectral dataset. We seek to train a universal cross-domain model, which can later be deployed for various spectral domains. To achieve, we physically furnish multiple inlets to the model while having a universal portion, which is designed to handle the inconsistent spectral characteristics among different domains. Note that only the universal portion is used in the finetune process. This approach naturally enables the learning of our model on multiple domains simultaneously, which acts as an effective workaround for the issue of the absence of large-scale dataset. We have carried out a study to extensively compare models that were trained using cross-domain approach with ones trained from scratch. Our approach was found to be superior both in accuracy and training efficiency. In addition, we have verified that our approach effectively reduces the overfitting issue, enabling us to deepen the model up to 13 layers (from 9) without compromising the accuracy.
Hyungtae Lee, Sungmin Eum, Heesung Kwon
IEEE Trans. Geosci. Remote. Sens.1
2021 DBF: Dynamic Belief Fusion for Combining Multiple Object Detectors
abstract
In this article, we propose a novel and highly practical score-level fusion approach called dynamic belief fusion ( DBF) that directly integrates inference scores of individual detections from multiple object detection methods. To effectively integrate the individual outputs of multiple detectors, the level of ambiguity in each detection score is estimated using a confidence model built on a precision-recall relationship of the corresponding detector. For each detector output, DBF then calculates the probabilities of three hypotheses (target, non-target, and intermediate state (target or non-target)) based on the confidence level of the detection score conditioned on the prior confidence model of individual detectors, which is referred to as basic probability assignment. The probability distributions over three hypotheses of all the detectors are optimally fused via the Dempster's combination rule. Experiments on the ARL, PASCAL VOC 07, and 12 datasets show that the detection accuracy of the DBF is significantly higher than any of the baseline fusion approaches as well as individual detectors used for the fusion.
Hyungtae Lee, Heesung Kwon
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 S-DOD-CNN: Doubly Injecting Spatially-Preserved Object Information for Event Recognition
abstract
We present a novel event recognition approach called Spatially-preserved Doubly-injected Object Detection CNN (S-DOD-CNN), which incorporates the spatially preserved object detection information in both a direct and an indirect way. Indirect injection is carried out by simply sharing the weights between the object detection modules and the event recognition module. Meanwhile, our novelty lies in the fact that we have preserved the spatial information for the direct injection. Once multiple regions-of-intereset (RoIs) are acquired, their feature maps are computed and then projected onto a spatially-preserving combined feature map using one of the four Rol Projection approaches we present. In our architecture, combined feature maps are generated for object detection which are directly injected to the event recognition module. Our method provides the state-of-the-art accuracy for malicious event recognition.
Hyungtae Lee, Sungmin Eum, Heesung Kwon
ICASSP1
2020 ME R-CNN: Multi-Expert R-CNN for Object Detection
abstract
We introduce Multi-Expert Region-based Convolutional Neural Network (ME R-CNN) which is equipped with multiple experts (ME) where each expert is learned to process a certain type of regions of interest (RoIs). This architecture better captures the appearance variations of the RoIs caused by different shapes, poses, and viewing angles. In order to direct each RoI to the appropriate expert, we devise a novel "learnable" network, which we call, expert assignment network (EAN). EAN automatically learns the optimal RoI-expert relationship even without any supervision of expert assignment. As the major components of ME R-CNN, ME and EAN, are mutually affecting each other while tied to a shared network, neither an alternating nor a naive end-to-end optimization is likely to fail. To address this problem, we introduce a practical training strategy which is tailored to optimize ME, EAN, and the shared network in an end-to-end fashion. We show that both of the architectures provide considerable performance increase over the baselines on PASCAL VOC 07, 12, and MS COCO datasets.
Hyungtae Lee, Sungmin Eum, Heesung Kwon
IEEE Trans. Image Process.1
2019 DOD-CNN: Doubly-injecting Object Information for Event Recognition
abstract
Recognizing an event in an image can be enhanced by detecting relevant objects in two ways: 1) indirectly utilizing object detection information within the unified architecture or 2) directly making use of the object detection output results. We introduce a novel approach, referred to as Doubly-injected Object Detection CNN (DOD-CNN), exploiting the object information in both ways for the task of event recognition. The structure of this network is inspired by the Integrated Object Detection CNN (IOD-CNN) where object information is indirectly exploited by the event recognition module through the shared portion of the network. In the DOD-CNN architecture, the intermediate object detection outputs are directly injected into the event recognition network while keeping the indirect sharing structure inherited from the IOD-CNN, thus being `doubly-injected'. We also introduce a batch pooling layer which constructs one representative feature map from multiple object hypotheses. We have demonstrated the effectiveness of injecting the object detection information in two different ways in the task of malicious event recognition.
Hyungtae Lee, Sungmin Eum, Heesung Kwon
ICASSP1
2019 Is Pretraining Necessary for hyperspectral image classification?
abstract
We address two questions for training a convolutional neural network (CNN) for hyperspectral image classification: i) is it possible to build a pre-trained network? and ii) is the pre-training effective in furthering the performance? To answer the first question, we have devised an approach that pre-trains a network on multiple source datasets that differ in their hyperspectral characteristics and fine-tunes on a target dataset. This approach effectively resolves the architectural issue that arises when transferring meaningful information between the source and the target networks. To answer the second question, we carried out several ablation experiments. Based on the experimental results, a network trained from scratch performs as good as a network fine-tuned from a pre-trained network. However, we observed that pre-training the network has its own advantage in achieving better performances when deeper networks are required.
Hyungtae Lee, Sungmin Eum, Heesung Kwon
IGARSS1
2018 Exploitation of Semantic Keywords for Malicious Event Classification
abstract
Learning an event classifier is challenging when the scenes are semantically different but visually similar. However, as humans, we typically handle such tasks painlessly by adding our background semantic knowledge. Motivated by this observation, we aim to provide an empirical study about how additional information such as semantic keywords can boost up the discrimination of such events. To demonstrate the validity of this study, we first construct a novel Malicious Crowd Dataset containing crowd images with two events, benign and malicious, which look visually similar. Note that the primary focus of this paper is not to provide the state-of-the-art performance on this dataset but to show the beneficial aspects of using semantically-driven keyword information. By leveraging crowd-sourcing platforms, such as Amazon Mechanical Turk, we collect semantic keywords associated with images and then subsequently identify a subset of keywords (e.g. police, fire, etc.) unique to specific events. We first show that by using recently introduced attention models, a naive CNN-based event classifier actually learns to primarily focus on local attributes associated with the discriminant semantic keywords identified by the Turks. We further show that incorporating the keyword-driven information into early-and late-fusion approaches can significantly enhance malicious event classification.
Hyungtae Lee, Sungmin Eum, Joel Levis, Heesung Kwon, James Michaelis, Michael Kolodny
ICASSP1
2018 Cross-Domain CNN for Hyperspectral Image Classification
abstract
In this paper, we address the dataset scarcity issue with the hyperspectral image classification. As only a few thousands of pixels are available for training, it is difficult to effectively learn high-capacity Convolutional Neural Networks (CNNs). To cope with this problem, we propose a novel cross-domain CNN containing the shared parameters which can co-learn across multiple hyperspectral datasets. the network also contains the non-shared portions designed to handle the dataset-specific spectral characteristics and the associated classification tasks. Our approach is the first attempt to learn a CNN for multiple hyperspectral datasets, in an end-to-end fashion. Moreover, we have experimentally shown that the proposed network trained on three of the widely used datasets outperform all the baseline networks which are trained on single dataset.
Hyungtae Lee, Sungmin Eum, Heesung Kwon
IGARSS1
2017 Deep Network Shrinkage Applied to Cross-Spectrum Face Recognition
abstract
In recent years, deep learning has emerged as a dominant methodology in virtually all machine learning problems. While it has been shown to produce state-of-the-art results for a variety of applicatons (including face recognition and heterogeneous face recognition), one aspect of deep networks that has not been extensively researched is how to determine the optimal network structure. This problem is generally solved by ad hoc methods. In this work we address a subproblem of this task: determining the breadth (number of nodes) of each layer. We show how to use group-sparsity-inducing regularization to effectively replace these hyper-parameters with a single hyperparameter which can be determined by cross-validation. We demonstrate our method by using it to reduce the size of networks on two commonly used NIR face datasets.
Christopher Reale, Hyungtae Lee, Heesung Kwon, Rama Chellappa
FG2
2017 Enhanced object detection via fusion with prior beliefs from image classification
abstract
In this paper, we introduce a novel fusion method that can enhance object detection performance by fusing decisions from two different types of computer vision tasks: object detection and image classification. In the proposed work, the class label of an image obtained from the image classification task is viewed as prior knowledge about existence or non-existence of certain objects. The prior knowledge is then fused with the decisions of object detection to improve detection accuracy by mitigating false positives of an object detector that are strongly contradicted with the prior knowledge. A recently introduced novel fusion approach called dynamic belief fusion (DBF) is used to fuse the detector output with the classification prior. Experimental results show that the detection performance of all the detection algorithms used in the proposed work is improved on benchmark datasets via the proposed fusion framework.
Yilun Cao, Hyungtae Lee, Heesung Kwon
ICIP2
2017 IOD-CNN: Integrating object detection networks for event recognition
abstract
Many previous methods have showed the importance of considering semantically relevant objects for performing event recognition, yet none of the methods have exploited the power of deep convolutional neural networks to directly integrate relevant object information into a unified network. We present a novel unified deep CNN architecture which integrates architecturally different, yet semantically-related object detection networks to enhance the performance of the event recognition task. Our architecture allows the sharing of the convolutional layers and a fully connected layer which effectively integrates event recognition, rigid object detection and non-rigid object detection.
Sungmin Eum, Hyungtae Lee, Heesung Kwon, David S. Doermann
ICIP2
2017 Going Deeper With Contextual CNN for Hyperspectral Image Classification
abstract
In this paper, we describe a novel deep convolutional neural network (CNN) that is deeper and wider than other existing deep networks for hyperspectral image classification. Unlike current state-of-the-art approaches in CNN-based hyperspectral image classification, the proposed network, called contextual deep CNN, can optimally explore local contextual interactions by jointly exploiting local spatio-spectral relationships of neighboring individual pixel vectors. The joint exploitation of the spatio-spectral information is achieved by a multi-scale convolutional filter bank used as an initial component of the proposed CNN pipeline. The initial spatial and spectral feature maps obtained from the multi-scale filter bank are then combined together to form a joint spatio-spectral feature map. The joint feature map representing rich spectral and spatial properties of the hyperspectral image is then fed through a fully convolutional network that eventually predicts the corresponding label of each pixel vector. The proposed approach is tested on three benchmark data sets: the Indian Pines data set, the Salinas data set, and the University of Pavia data set. Performance comparison shows enhanced classification performance of the proposed approach over the current state-of-the-art on the three data sets.
Hyungtae Lee, Heesung Kwon
IEEE Trans. Image Process.1
2016 Weakly Supervised Localization Using Deep Feature Maps
Archith J. Bency, Heesung Kwon, Hyungtae Lee, S. Karthikeyan 0001, B. S. Manjunath
ECCV (1)3
2016 DTM: Deformable template matching
abstract
A novel template matching algorithm that can incorporate the concept of deformable parts, is presented in this paper. Unlike the deformable part model (DPM) employed in object recognition, the proposed template-matching approach called Deformable Template Matching (DTM) does not require a training step. Instead, deformation is achieved by a set of predefined basic rules (e.g. the left sub-patch cannot pass across the right patch). Experimental evaluation of this new method using the PASCAL VOC 07 dataset demonstrated substantial performance improvement over conventional template matching algorithms. Additionally, to confirm the applicability of DTM, the concept is applied to the generation of a rotation-invariant SIFT descriptor. Experimental evaluation employing deformable matching of SIFT features shows an increased number of matching features compared to a conventional SIFT matching.
Hyungtae Lee, Heesung Kwon, Ryan M. Robinson, William D. Nothwang
ICASSP1
2016 Contextual deep CNN based hyperspectral classification
abstract
In this paper, we describe a novel deep convolutional neural networks (CNN) based approach called contextual deep CNN that can jointly exploit spatial and spectral features for hyperspectral image classification. The contextual deep CNN first concurrently applies multiple 3-dimensional local convolutional filters with different sizes jointly exploiting spatial and spectral features of a hyperspectral image. The initial spatial and spectral feature maps obtained from applying the variable size convolutional filters are then combined together to form a joint spatio-spectral feature map. The joint feature map representing rich spectral and spatial properties of the hyperspectral image is then fed through fully convolutional layers that eventually predict the corresponding label of each pixel vector. The proposed approach is tested on two benchmark datasets: the Indian Pines dataset and the Pavia University scene dataset. Performance comparison shows enhanced classification performance of the proposed approach over the current state of the art on both datasets.
Hyungtae Lee, Heesung Kwon
IGARSS1
2016 Task-conversions for integrating human and machine perception in a unified task
abstract
The different strategies for feature extraction and synthesis employed by humans and computers are often complementary, hence combining the two into an integrated object recognition system may considerably improve performance over either used in isolation. Rapid Serial Visual Presentation (RSVP) is one well-established technique that has shown promise integrating human perception into a machine perception system. In this paper, we apply computer vision techniques to image data filtered through human RSVP. We introduce “task conversions” to integrate the two modalities, applying the precise localization capabilities of computer vision with the detection capabilities of RSVP. We employ naive Bayesian fusion and a novel method, dynamic belief fusion (DBF), in a joint scheme as fusion approaches. Preliminary experiments demonstrate that DBF extracts complementary information from both human and machine sources to improve performance for both target classification and object detection.
Hyungtae Lee, Heesung Kwon, Ryan M. Robinson, Daniel Donavanik, William D. Nothwang, Amar R. Marathe
IROS1
2016 Dynamic belief fusion for object detection
abstract
A novel approach for the fusion of heterogeneous object detection methods is proposed. In order to effectively integrate the outputs of multiple detectors, the level of ambiguity in each individual detection score is estimated using the precision/recall relationship of the corresponding detector. The main contribution of the proposed work is a novel fusion method, called Dynamic Belief Fusion (DBF), which dynamically assigns probabilities to hypotheses (target, non-target, intermediate state (target or non-target)) based on confidence levels in the detection results conditioned on the prior performance of individual detectors. In DBF, a joint basic probability assignment, optimally fusing information from all detectors, is determined by the Dempster's combination rule, and is easily reduced to a single fused detection score. Experiments on ARL and PASCAL VOC 07 datasets demonstrate that the detection accuracy of DBF is considerably greater than conventional fusion approaches as well as individual detectors used for the fusion.
Hyungtae Lee, Heesung Kwon, Ryan M. Robinson, William D. Nothwang, Amar M. Marathe
WACV1
2015 JH2R: Joint Homography Estimation for Highlight Removal
abstract
Imagine being in an art museum where there are paintings or pictures held inside glass-frames for protection. There are pieces which you wish to capture using a camera, but you experience difficulties avoiding highlights which are generated by indoor lighting reflected off the glossy surfaces. Similar problems occur when capturing contents off of whiteboards, documents printed on glossy surfaces, objects such as books or CDs with plastic covers. In this work, we address the problem of removing unwanted highlight regions in images generated by reflections of light sources on glossy surfaces. Although there have been efforts made to synthetically fill in the missing regions using the neighboring patterns by applying methods like inpainting [3, 4], it is impossible to recover the missing information in completely saturated regions. Therefore, we need to use multiple images where corresponding regions are not covered by the saturated highlights. Unlike other methods, our method uses the relationship between the highlight regions resulting in more robust removal of saturated highlights. Our method Overview Our method was motivated by a widely acknowledged physical phenomenon referred to as the ‘motion parallax’. Without loss of generality, we can similarly view the relationship between the desired content (e.g., a painting) and the highlights. Since the highlights caused by the light source are the result of the reflection on the glossy surface before they reach the camera, the light source can be modeled to virtually exist on the other side of the content. Note that, the distance from the light source is always larger than the distance from the content (D > d, in Figure 1).
Sungmin Eum, Hyungtae Lee, David S. Doermann
BMVC2
2015 Human-autonomy sensor fusion for rapid object detection
abstract
Human-autonomy sensor fusion is an emerging technology with a wide range of applications, including object detection/recognition, surveillance, collaborative control, and prosthetics. For object detection, humans and computer-vision-based systems employ different strategies to locate targets, likely providing complementary information. However, little effort has been made in combining the outputs of multiple autonomous detectors and multiple human-generated responses. This paper presents a method for integrating several sources of human- and autonomy-generated information for rapid object detection tasks. Human electroencephalography (EEG) and button-press responses from rapid serial visual presentation (RSVP) experiments are fused with outputs from trained object detection algorithms. Three fusion methods—Bayesian, Dempster-Shafer, and Dynamic Dempster-Shafer—are implemented for comparison. Results demonstrate that fusion of these human classifiers with computer-vision-based detectors improves object detection accuracy over purely computer-vision-based detection (5% relative increase in mean average precision) and the best individual computer vision algorithm (28% relative increase in mean average precision). Computer vision fused with button press response and/or the XDAWN + Bayesian Linear Discriminant Analysis neural classifier provides considerable improvement, while computer vision fused with other neural classifiers provides little or no improvement. Of the three fusion methods, Dynamic Dempster-Shafer Theory (DDST) Fusion exhibits the greatest performance in this application.
Ryan M. Robinson, Hyungtae Lee, Michael J. McCourt, Amar R. Marathe, Heesung Kwon, Chau Ton, William D. Nothwang
IROS2
2015 Clauselets: Leveraging Temporally Related Actions for Video Event Analysis
abstract
We propose clause lets, sets of concurrent actions and their temporal relationships, and explore their application to video event analysis. We train clause lets in two stages. We initially train first level clause let detectors that find a limited set of actions in particular qualitative temporal configurations based on Allen's interval relations. In the second stage, we apply the first level detectors to training videos, and discriminatively learn temporal patterns between activations that involve more actions over longer durations and lead to improved second level clause let models. We demonstrate the utility of clause lets by applying them to the task of "in-the-wild" video event recognition on the TRECVID MED 11 dataset. Not only do clause lets achieve state-of-the-art results on this task, but qualitative results suggest that they may also lead to semantically meaningful descriptions of videos in terms of detected actions and their temporal relationships.
Hyungtae Lee, Vlad I. Morariu, Larry Davis 0001
WACV1
2012 Qualitative Pose Estimation by Discriminative Deformable Part Models
Hyungtae Lee, Vlad I. Morariu, Larry Davis 0001
ACCV (2)1
2011 AVSS 2011 demo session: A large-scale benchmark dataset for event recognition in surveillance video
abstract
Summary form only given. We present a concept for automatic construction site monitoring by taking into account 4D information (3D over time), that is acquired from highly-overlapping digital aerial images. On the one hand today's maturity of flying micro aerial vehicles (MAVs) enables a low-cost and an efficient image acquisition of high-quality data that maps construction sites entirely from many varying viewpoints. On the other hand, due to low-noise sensors and high redundancy in the image data, recent developments in 3D reconstruction workflows have benefited the automatic computation of accurate and dense 3D scene information. Having both an inexpensive high-quality image acquisition and an efficient 3D analysis workflow enables monitoring, documentation and visualization of observed sites over time with short intervals. Relating acquired 4D site observations, composed of color, texture, geometry over time, largely supports automated methods toward full scene understanding, the acquisition of both the change and the construction site's progress.
Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai
AVSS9
2011 A large-scale benchmark dataset for event recognition in surveillance video
abstract
We introduce a new large-scale video dataset designed to assess the performance of diverse visual event recognition algorithms with a focus on continuous visual event recognition (CVER) in outdoor areas with wide coverage. Previous datasets for action recognition are unrealistic for real-world surveillance because they consist of short clips showing one action by one individual [15, 8]. Datasets have been developed for movies [11] and sports [12], but, these actions and scene conditions do not apply effectively to surveillance videos. Our dataset consists of many outdoor scenes with actions occurring naturally by non-actors in continuously captured videos of the real world. The dataset includes large numbers of instances for 23 event types distributed throughout 29 hours of video. This data is accompanied by detailed annotations which include both moving object tracks and event examples, which will provide solid basis for large-scale evaluation. Additionally, we propose different types of evaluation modes for visual recognition tasks and evaluation metrics along with our preliminary experimental results. We believe that this dataset will stimulate diverse aspects of computer vision research and help us to advance the CVER tasks in the years ahead.
Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai
CVPR9