VLDB 2026 Research / reviewers in the wild / expert
Heesung Kwon
dblp:10/5946
· DBLP profile ↗
74ranked-venue papers
18as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 9 first-author · 12 since 2021Artificial intelligence and machine learning · 23 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 4 first-author · 1 since 2021Systems, architecture and hardware · 6 · 3 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery Using Gaussian SplattingabstractDespite significant advancements in dynamic neural rendering, existing methods fail to address the unique challenges posed by UAV-captured scenarios, particularly those involving monocular camera setups, top-down perspective, and multiple small, moving humans, which are not adequately represented in existing datasets. In this work, we introduce UAV4D, a framework for enabling photorealistic rendering for dynamic real-world scenes captured by UAVs. Specifically, we address the challenge of reconstructing dynamic scenes with multiple moving pedestrians from monocular video data without the need for additional sensors. We use a combination of a 3D foundation model and a human mesh reconstruction model to reconstruct both the scene background and humans. We propose a novel approach to resolve the scene scale ambiguity and place both humans and the scene in world coordinates by identifying human-scene contact points. Additionally, we exploit the SMPL model and background mesh to initialize Gaussian splats, enabling holistic scene rendering. We evaluated our method on three complex UAV-captured datasets: VisDrone, Manipal-UAV, and Okutama-Action, each with distinct characteristics and 10-50 humans. Our results demonstrate the benefits of our approach over existing methods in novel view synthesis, achieving a 1.5 dB PSNR improvement and superior visual sharpness. Dongki Jung, Christopher Maxey, Sungmin Eum, Yonghan Lee 0001, Dinesh Manocha, Heesung Kwon |
AAAI | 7 |
| 2026 | Investigating AI-induced Technostress and Coping Strategies of ProfessionalsabstractWhile the rise of AI has benefited professionals, it also induces technostress that threatens their expertise and jobs. To ensure the human-centered advancement of technology, a deep understanding of users technostress and how to cope with it is essential. Despite technostress having long been discussed, the growing integration of AI tools into professionals’ everyday work amplifies these challenges and calls for further exploration. Accordingly, this is a timely moment to examine their real-world experiences and voices. Thus, our study aims to investigate AI-induced technostress experienced by professionals, and the coping strategies they employ. Through focus group interviews with 19 professionals from diverse fields, we identified seven AI-induced technostressors and examined their coping strategies along two dimensions: stress Coping Style (problem-focused and emotion-focused) and Value Orientation (AI-oriented and humanness-oriented). Drawing on professionals’ coping strategies, we suggest practical implications to support users in coping with AI-induced technostress. Heesung Kwon, Jeesun Oh, Suyoun Lee, Sunok Lee |
CHI | 1 |
| 2026 | MoRe: Monocular Geometry Refinement via Graph Optimization for Cross-View ConsistencyabstractMonocular 3D foundation models offer an extensible solution for perception tasks, making them attractive for broader 3D vision applications. In this paper, we propose MoRe, a training-free Monocular Geometry Refinement method designed to improve cross-view consistency and achieve scale alignment. To induce inter-frame relationships, our method employs feature matching between frames to establish correspondences. Rather than applying simple least squares optimization on these matched points, we formulate a graph-based optimization framework that performs local planar approximation using the estimated 3D points and surface normals estimated by monocular foundation models. This formulation addresses the scale ambiguity inherent in monocular geometric priors while preserving the underlying 3D structure. We further demonstrate that MoRe not only enhances 3D reconstruction but also improves novel view synthesis, particularly in sparse-view rendering scenarios. Dongki Jung, Yonghan Lee 0001, Sungmin Eum, Heesung Kwon, Dinesh Manocha |
WACV | 5 |
| 2026 | SynPlay: Large-Scale Synthetic Human Data with Real-World Diversity for Aerial-View PerceptionabstractWe introduce SynPlay, a large-scale synthetic human dataset purpose-built for advancing multi-perspective human localization, with a predominant focus on aerial-view perception. SynPlay departs from traditional synthetic datasets by addressing a critical but underexplored challenge: localizing humans in aerial scenes where subjects often occupy only tens of pixels in the image. In such scenarios, fine-grained details like facial features or textures become irrelevant, shifting the burden of recognition to human motion, behavior, and interactions. To meet this need, SynPlay implements a novel rule-guided motion generation framework that combines real-world motion capture with motion evolution graphs. This design enables human actions to evolve dynamically through high-level game rules rather than predefined scripts, resulting in effectively uncountable motion variations. Unlike existing synthetic datasets—which either focus on static visual traits or reuse a limited set of mocap-driven actions—SynPlay captures a wide spectrum of spontaneous behaviors, including complex interactions that naturally emerge from unscripted gameplay scenarios. SynPlay also introduces an extensive multi-camera setup that spans UAVs at random altitudes, CCTVs, and a freely roaming UGV, achieving true near-to-far perspective coverage in a single dataset. The majority of instances are captured from aerial viewpoints at varying scales, directly supporting the development of models for long-range human analysis—a setting where existing datasets fall short. Our data contains over 73k images and 6.5M human instances, with detailed annotations for detection, segmentation, and keypoint tasks. Extensive experiments demonstrate that training with SynPlay significantly improves human localization performance, especially in few-shot and data-scarce scenarios. Jinsub Yim, Hyungtae Lee, Sungmin Eum, Yi-Ting Shen, Heesung Kwon, Shuvra S. Bhattacharyya |
WACV | 6 |
| 2025 | Autocompose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMsabstractComposed pose retrieval (CPR) enables users to search for human poses by specifying a reference pose and a transition description, but progress in this field is hindered by the scarcity and inconsistency of annotated pose transitions. Existing CPR datasets rely on costly human annotations or heuristic-based rule generation, both of which limit scalability and diversity. In this work, we introduce AutoComPose, the first framework that leverages multimodal large language models (MLLMs) to automatically generate rich and structured pose transition descriptions. Our method enhances annotation quality by structuring transitions into fine-grained body part movements and introducing mirrored/swapped variations, while a cyclic consistency constraint ensures logical coherence between forward and reverse transitions. To advance CPR research, we construct and release two dedicated benchmarks, AIST-CPR and PoseFixCPR, supplementing prior datasets with enhanced attributes. Extensive experiments demonstrate that training retrieval models with AutoComPose yields superior performance over human-annotated and heuristic-based methods, significantly reducing annotation costs while improving retrieval quality. Our work pioneers the automatic annotation of pose transitions, establishing a scalable foundation for future CPR research. Yi-Ting Shen, Sungmin Eum, Doheon Lee, Rohit Shete, Chiao-Yi Wang, Heesung Kwon, Shuvra S. Bhattacharyya |
ICCV | 6 |
| 2025 | Diversifying Human Pose In Synthetic Data For Aerial-View Human DetectionabstractSynthetic data generation has emerged as a promising solution to the data scarcity issue in aerial-view human detection. However, creating datasets that accurately reflect varying real-world human appearances—particularly diverse poses—remains challenging and labor-intensive. To address this, we propose SynPoseDiv, a novel framework that diversifies human poses within existing synthetic datasets. SynPoseDiv tackles two key challenges: generating realistic, diverse 3D human poses using a diffusion-based pose generator, and producing images of virtual characters in novel poses through a source-to-target image translator. The framework incrementally transitions characters into new poses using optimized pose sequences identified via Dijkstra’s algorithm. Experiments demonstrate that SynPoseDiv significantly improves detection accuracy across multiple aerial-view human detection benchmarks, especially in low-shot scenarios, and remains effective regardless of the training approach or dataset size. Yi-Ting Shen, Hyungtae Lee, Heesung Kwon, Shuvra S. Bhattacharyya |
ICIP | 3 |
| 2025 | TK-Planes: Tiered K-Planes with High Dimensional Feature Vectors for Dynamic UAV-based ScenesabstractIn this paper, we present a new approach to improve the neural rendering fidelity of in-the-wild unmanned aerial vehicle (UAV)-based scenes. Our formulation is designed for dynamic scenes, consisting of small moving objects or human actions in particular. We propose an extension of K-Planes Neural Radiance Field (NeRF), wherein our algorithm stores a set of tiered high dimensional feature vectors. The tiered feature vectors are generated to effectively model conceptual information about a scene as well as to be processed by an image decoder that transforms output feature maps into RGB images. Our technique leverages the information among both static and dynamic objects within a scene and is able to capture salient scene attributes of high altitude videos. We evaluate its performance on challenging datasets, including Okutama Action and UG2, and observe considerable improvement in accuracy over state of the art neural rendering methods. Christopher Maxey, Yonghan Lee 0001, Hyungtae Lee, Dinesh Manocha, Heesung Kwon |
IROS | 6 |
| 2024 | MeshGS: Adaptive Mesh-Aligned Gaussian Splatting for High-Quality Rendering
Yonghan Lee 0001, Hyungtae Lee, Heesung Kwon, Dinesh Manocha |
ACCV (9) | 4 |
| 2024 | Exploring the Potential of Synthetic Data to Replace Real DataabstractThe potential of synthetic data to replace real data creates a huge demand for synthetic data in data-hungry AI. This potential is even greater when synthetic data is used for training along with a small number of real images from domains other than the test domain. We find that this potential varies depending on (i) the number of cross-domain real images and (ii) the test set on which the trained model is evaluated. We introduce two new metrics, the train2test distance and $\mathrm{AP}_{\mathrm{t} 2 \mathrm{t}}$, to evaluate the ability of a cross-domain training set using synthetic data to represent the characteristics of test instances in relation to training performance. Using these metrics, we delve deeper into the factors that influence the potential of synthetic data and uncover some interesting dynamics about how synthetic data impacts training performance. We hope these discoveries will encourage more widespread use of synthetic data. Hyungtae Lee, Heesung Kwon, Shuvra S. Bhattacharyya |
ICIP | 3 |
| 2024 | UAV-Sim: NeRF-based Synthetic Data Generation for UAV-based PerceptionabstractTremendous variations coupled with large degrees of freedom in UAV-based imaging conditions lead to a significant lack of data in adequately learning UAV-based perception models. Using various synthetic renderers in conjunction with perception models is prevalent to create synthetic data to augment the learning in the ground-based imaging domain. However, severe challenges in the austere UAV-based domain require distinctive solutions to image synthesis for data augmentation. In this work, we leverage recent advancements in neural rendering to improve static and dynamic novel-view UAV-based image synthesis, especially from high altitudes, capturing salient scene attributes. Finally, we demonstrate a considerable performance boost is achieved when a state-of-the-art detection model is optimized primarily on hybrid sets of real and synthetic data instead of the real or synthetic data separately. Christopher Maxey, Hyungtae Lee, Dinesh Manocha, Heesung Kwon |
ICRA | 5 |
| 2024 | Two Teachers Are Better Than One: Leveraging Depth In Training Only For Unsupervised Obstacle SegmentationabstractWe present a novel unsupervised obstacle segmentation architecture that follows a novel Relation Distillation (RD) paradigm. Our architecture design was inspired by a self-supervised teacher-student approach that relies on the Semantic Distillation originally devised for representation learning. While the teacher in the Semantic Distillation considers a single patch at a time, the teacher within RD takes a ‘pair of patches’ instead to transfer the local Semantic Co-occurrence Localization (SCooL) relationship that focuses more on the segmentation-boosting signals. To further improve the proposed architecture, we introduce the utilization of another teacher that leverages the depth information which inherently separates the entities at different physical distances, often tied with the boundaries of the obstacles. As the depth is distilled towards the student network only at the time of training, it adds zero computational/hardware cost at run-time. As no relevant public dataset is available, we have curated the Avoiding Obstacles In unstructured Driving (AvOID) dataset as a new testbed for unsupervised obstacle segmentation. We have validated that both the Relation Distillation and depth contribute to boosting the no-annotation segmentation performance on AvOID and KITTI-Obstacles. Sungmin Eum, Hyungtae Lee, Heesung Kwon, Philip R. Osteen, Andre Harrison |
IROS | 3 |
| 2024 | EDIR: Efficient Distributed Image Retrieval of Novel Objects in Mobile NetworksabstractCrowdsourcing data collection from a network of mobile devices is useful in various applications. Mobile devices store a large amount of visual data that can aid in different application scenarios. Trained Convolutional Neural Networks (CNNs) can be deployed on mobile devices to be used in searching for objects of interest. Querying for novel objects, for which models have not been trained yet, presents some unique challenges. When novel objects are queried, new models must be trained and distributed to all edge devices. In this paper, we propose an efficient method and a system, called EDIR, which enables answering these queries while taking into account the bandwidth limitations encountered in wireless networks, as well as the limited energy and computational power on mobile devices. Through extensive experimentation, we show that using distance-based classifiers, specifically those relying on the Cosine distance, leads to more efficient utilization of network resources by reducing the number of false positives. We perform analysis that enables the requester to tune the parameters of interest before issuing the query, and validate our theoretical results. EDIR reduces the amount of transferred data by more than 45% compared to other approaches while simultaneously achieving a good F1 score. Noor Felemban, Fidan Mehmeti, Thomas La Porta, Heesung Kwon |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | Progressive Transformation Learning for Leveraging Virtual Images in TrainingabstractTo effectively interrogate UAV-based images for detecting objects of interest, such as humans, it is essential to acquire large-scale UAV-based datasets that include human instances with various poses captured from widely varying viewing angles. As a viable alternative to laborious and costly data curation, we introduce Progressive Transformation Learning (PTL), which gradually augments a training dataset by adding transformed virtual images with enhanced realism. Generally, a virtual2real transformation generator in the conditional GAN framework suffers from quality degradation when a large domain gap exists between real and virtual images. To deal with the domain gap, PTL takes a novel approach that progressively iterates the following three steps: 1) select a subset from a pool of virtual images according to the domain gap, 2) transform the selected virtual images to enhance realism, and 3) add the transformed virtual images to the training set while removing them from the pool. In PTL, accurately quantifying the domain gap is critical. To do that, we theoretically demonstrate that the feature representation space of a given object detector can be modeled as a multivariate Gaussian distribution from which the Mahalanobis distance between a virtual object and the Gaussian distribution of each object category in the representation space can be readily computed. Experiments show that PTL results in a substantial performance increase over the baseline, especially in the small data and the cross-domain regime. Yi-Ting Shen, Hyungtae Lee, Heesung Kwon, Shuvra S. Bhattacharyya |
CVPR | 3 |
| 2023 | NEV-NCD: Negative Learning, Entropy, and Variance Regularization Based Novel Action Categories DiscoveryabstractNovel Categories Discovery (NCD) facilitates learning from a partially annotated label space and enables deep learning (DL) models to operate in an open-world setting by identifying and differentiating instances of novel classes based on the labeled data notions. One of the primary assumptions of NCD is that the novel label space is perfectly disjoint and can be equipartitioned, but it is rarely realized by most NCD approaches in practice. To better align with this assumption, we propose a novel single-stage joint optimization-based NCD method, Negative learning, Entropy, and Variance regularization NCD (NEV-NCD). We demonstrate the efficacy of NEV-NCD in previously unexplored NCD applications of video action recognition (VAR) with the public UCF101 dataset and a curated in-house partial action-space annotated multi-view video dataset. Further, we perform a thorough ablation study by varying the composition of final joint loss and associated hyper-parameters. During our experiments with UCF101 and multi-view action dataset, NEV-NCD achieves ≈ 83% classification accuracy in test instances of labeled data. NEV-NCD achieves ≈ 70% clustering accuracy over unlabeled data outperforming both naive baselines and state-of-the-art pseudo-labeling-based approaches by ≈ 40% and ≈ 3.5% over both datasets. Zahid Hasan 0001, Masud Ahmed, Abu Zaher Md Faridee, Sanjay Purushotham, Heesung Kwon, Hyungtae Lee, Nirmalya Roy |
ICIP | 5 |
| 2023 | A Multi-Purpose Realistic Haze Benchmark With Quantifiable Haze Levels and Ground TruthabstractImagery collected from outdoor visual environments is often degraded due to the presence of dense smoke or haze. A key challenge for research in scene understanding in these degraded visual environments (DVE) is the lack of representative benchmark datasets. These datasets are required to evaluate state-of-the-art object recognition and other computer vision algorithms in degraded settings. In this paper, we address some of these limitations by introducing the first realistic haze image benchmark, from both aerial and ground view, with paired haze-free images, and in-situ haze density measurements. This dataset was produced in a controlled environment with professional smoke generating machines that covered the entire scene, and consists of images captured from the perspective of both an unmanned aerial vehicle (UAV) and an unmanned ground vehicle (UGV). We also evaluate a set of representative state-of-the-art dehazing approaches as well as object detectors on the dataset. The full dataset presented in this paper, including the ground truth object classification bounding boxes and haze density measurements, is provided for the community to evaluate their algorithms at: https://a2i2-archangel.vision. A subset of this dataset has been used for the "Object Detection in Haze" Track of CVPR UG2 2022 challenge at https://cvpr2022.ug2challenge.org/track1.html. Priya Narayanan, Zhenyu Wu 0002, Matthew D. Thielke, John G. Rogers III, Andre Harrison, John A. D'Agostino, James D. Brown, Long Quang, James R. Uplinger, Heesung Kwon, Zhangyang Wang |
IEEE Trans. Image Process. | 11 |
| 2022 | Negative Samples are at Large: Leveraging Hard-Distance Elastic Loss for Re-identification
Hyungtae Lee, Sungmin Eum, Heesung Kwon |
ECCV (24) | 3 |
| 2022 | Self-Supervised Contrastive Learning for Cross-Domain Hyperspectral Image RepresentationabstractRecently, self-supervised learning has attracted attention due to its remarkable ability to acquire meaningful representations for classification tasks without using semantic labels. This paper introduces a self-supervised learning framework suitable for hyperspectral images that are inherently challenging to annotate. The proposed framework architecture leverages cross-domain CNN [1], allowing for learning representations from different hyperspectral images with varying spectral characteristics and no pixel-level annotation. In the framework, cross-domain representations are learned via contrastive learning where neighboring spectral vectors in the same image are clustered together in a common representation space encompassing multiple hyperspectral images. In contrast, spectral vectors in different hyperspectral images are separated into distinct clusters in the space. To verify that the learned representation through contrastive learning is effectively transferred into a downstream task, we perform a classification task on hyperspectral images. The experimental results demonstrate the advantage of the proposed self-supervised representation over models trained from scratch or other transfer learning methods. Hyungtae Lee, Heesung Kwon |
ICASSP | 2 |
| 2022 | MA3: Model-Accuracy Aware Anytime Planning with Simulation Verification for Navigating Complex TerrainsabstractOff-road and unstructured environments often contain complex patches of various types of terrain, rough elevation changes, deformable objects, etc. An autonomous ground vehicle traversing such environments experiences physical interactions that are extremely hard to model at scale and thus very hard to predict. Nevertheless, planning a safely traversable path through such an environment requires the ability to predict the outcomes of these interactions instead of avoiding them. One approach to doing this is to learn the interaction model offline based on collected data. Unfortunately, though, this requires large amounts of data and can often be brittle. Alternatively, models using physics-based simulators can generate large data and provide a reliable prediction. However, they are very slow to query online within the planning loop. This work proposes an algorithmic framework that utilizes the combination of a learned model and a physics-based simulation model for fast planning. Specifically, it uses the learned model as much as possible to accelerate planning while sparsely using the physics-based simulator to verify the feasibility of the planned path. We provide a theoretical analysis of the algorithm and its empirical evaluation showing a significant reduction in planning times. Manash Pratim Das, Damon M. Conover, Sungmin Eum, Heesung Kwon, Maxim Likhachev |
SOCS | 4 |
| 2022 | Exploring Cross-Domain Pretrained Model for Hyperspectral Image ClassificationabstractA pretrain-finetune strategy is widely used to reduce the overfitting that can occur when data are insufficient for convolutional neural network (CNN) training. The first few layers of a CNN pretrained on a large-scale RGB dataset are capable of acquiring general image characteristics, which are remarkably effective in tasks targeted for different RGB datasets. However, when it comes down to the hyperspectral domain where each domain has its unique spectral properties, the pretrain-finetune strategy no longer can be deployed in a conventional way while presenting three major issues: 1) inconsistent spectral characteristics among the domains (e.g., frequency range); 2) inconsistent number of data channels among the domains; and 3) absence of large-scale hyperspectral dataset. We seek to train a universal cross-domain model, which can later be deployed for various spectral domains. To achieve, we physically furnish multiple inlets to the model while having a universal portion, which is designed to handle the inconsistent spectral characteristics among different domains. Note that only the universal portion is used in the finetune process. This approach naturally enables the learning of our model on multiple domains simultaneously, which acts as an effective workaround for the issue of the absence of large-scale dataset. We have carried out a study to extensively compare models that were trained using cross-domain approach with ones trained from scratch. Our approach was found to be superior both in accuracy and training efficiency. In addition, we have verified that our approach effectively reduces the overfitting issue, enabling us to deepen the model up to 13 layers (from 9) without compromising the accuracy. Hyungtae Lee, Sungmin Eum, Heesung Kwon |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | EDIR: Efficient Distributed Image Retrieval of Novel Objects in Mobile NetworksabstractCrowdsourcing data collection from a network of mobile devices is useful in various applications. Mobile devices store a large amount of visual data that aid in different situations. Trained CNNs can be deployed on mobile devices to be used in searching for objects of interest. Querying for novel objects, for which models have not been trained, presents unique challenges. When novel objects are queried, new models must be trained and distributed to all edge devices, which can be cumbersome. In this paper we propose EDIR, an efficient method and a system that enables answering these queries while taking into account the bandwidth limitations in wireless networks, and the limited energy and computational power on mobile devices. Results show that EDIR reduces the amount of data transfer by 45%compared to other approaches while achieving a good F1 score. Noor Felemban, Fidan Mehmeti, Thomas La Porta, Heesung Kwon |
MASS | 4 |
| 2021 | DBF: Dynamic Belief Fusion for Combining Multiple Object DetectorsabstractIn this article, we propose a novel and highly practical score-level fusion approach called dynamic belief fusion ( DBF) that directly integrates inference scores of individual detections from multiple object detection methods. To effectively integrate the individual outputs of multiple detectors, the level of ambiguity in each detection score is estimated using a confidence model built on a precision-recall relationship of the corresponding detector. For each detector output, DBF then calculates the probabilities of three hypotheses (target, non-target, and intermediate state (target or non-target)) based on the confidence level of the detection score conditioned on the prior confidence model of individual detectors, which is referred to as basic probability assignment. The probability distributions over three hypotheses of all the detectors are optimally fused via the Dempster's combination rule. Experiments on the ARL, PASCAL VOC 07, and 12 datasets show that the detection accuracy of the DBF is significantly higher than any of the baseline fusion approaches as well as individual detectors used for the fusion. Hyungtae Lee, Heesung Kwon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | S-DOD-CNN: Doubly Injecting Spatially-Preserved Object Information for Event RecognitionabstractWe present a novel event recognition approach called Spatially-preserved Doubly-injected Object Detection CNN (S-DOD-CNN), which incorporates the spatially preserved object detection information in both a direct and an indirect way. Indirect injection is carried out by simply sharing the weights between the object detection modules and the event recognition module. Meanwhile, our novelty lies in the fact that we have preserved the spatial information for the direct injection. Once multiple regions-of-intereset (RoIs) are acquired, their feature maps are computed and then projected onto a spatially-preserving combined feature map using one of the four Rol Projection approaches we present. In our architecture, combined feature maps are generated for object detection which are directly injected to the event recognition module. Our method provides the state-of-the-art accuracy for malicious event recognition. Hyungtae Lee, Sungmin Eum, Heesung Kwon |
ICASSP | 3 |
| 2020 | Semantics to Space(S2S): Embedding semantics into spatial space for zero-shot verb-object query inferencingabstractWe present a novel deep zero-shot learning (ZSL) model for inferencing human-object-interaction with verb-object (VO) query. While the previous two-stream ZSL approaches only use the semantic/textual information to be fed into the query stream, we seek to incorporate and embed the semantics into the visual representation stream as well. Our approach is powered by Semantics-to-Space (S2S) architecture where semantics derived from the residing objects are embedded into a spatial space of the visual stream. This architecture allows the co-capturing of the semantic attributes of the human and the objects along with their location/size/silhouette information. To validate, we have constructed a new dataset, Verb-Transferability 60 (VT60). VT60 provides 60 different VO pairs with overlapping verbs tailored for testing two-stream ZSL approaches with VO query. Experimental evaluations show that our approach not only outperforms the state-of-the-art, but also shows the capability of consistently improving performance regardless of which ZSL baseline architecture is used. Sungmin Eum, Heesung Kwon |
ICPR | 2 |
| 2020 | CLIO: enabling automatic compilation of deep learning pipelines across IoT and cloudabstractRecent years have seen dramatic advances in low-power neural accelerators that aim to bring deep learning analytics to IoT devices; simultaneously, there have been considerable advances in the design of low-power radios to enable efficient compute offload from IoT devices to the cloud. Neither is a panacea --- deep learning models are often too large for low-power accelerators and bandwidth needs are often too high for low-power radios. While there has been considerable work on deep learning for smartphone-class devices, these methods do not work well for small battery-powered IoT devices that are considerably more resource-constrained. Colin Samplawski, Deepak Ganesan, Benjamin M. Marlin, Heesung Kwon |
MobiCom | 5 |
| 2020 | ME R-CNN: Multi-Expert R-CNN for Object DetectionabstractWe introduce Multi-Expert Region-based Convolutional Neural Network (ME R-CNN) which is equipped with multiple experts (ME) where each expert is learned to process a certain type of regions of interest (RoIs). This architecture better captures the appearance variations of the RoIs caused by different shapes, poses, and viewing angles. In order to direct each RoI to the appropriate expert, we devise a novel "learnable" network, which we call, expert assignment network (EAN). EAN automatically learns the optimal RoI-expert relationship even without any supervision of expert assignment. As the major components of ME R-CNN, ME and EAN, are mutually affecting each other while tied to a shared network, neither an alternating nor a naive end-to-end optimization is likely to fail. To address this problem, we introduce a practical training strategy which is tailored to optimize ME, EAN, and the shared network in an end-to-end fashion. We show that both of the architectures provide considerable performance increase over the baselines on PASCAL VOC 07, 12, and MS COCO datasets. Hyungtae Lee, Sungmin Eum, Heesung Kwon |
IEEE Trans. Image Process. | 3 |
| 2019 | SENSE: Semantically Enhanced Node Sequence EmbeddingabstractEffectively representing graph node sequences in the form of vector embeddings is critical to many applications. We achieve this by (i) first learning vector embeddings of single graph nodes and (ii) then composing them to compactly represent node sequences. Specifically, we propose SENSE-S (Semantically Enhanced Node Sequence Embedding - for Single nodes), a skip-gram based novel embedding mechanism, for single graph nodes that co-learns graph structure as well as their textual descriptions. We demonstrate that SENSE-S vectors increase the accuracy of multi-label classification tasks by up to 50% and link-prediction tasks by up to 78% under a variety of scenarios using real datasets. Based on SENSE-S, we next propose generic SENSE to compute composite vectors that represent a sequence of nodes, where preserving the node order is important. We prove that this approach is efficient in embedding node sequences, and our experiments on real data confirm its high accuracy. Swati Rallapalli, Liang Ma 0002, Mudhakar Srivatsa, Ananthram Swami, Heesung Kwon, Graham A. Bent, Christopher Simpkin |
IEEE BigData | 5 |
| 2019 | Object and Text-guided Semantics for CNN-based Activity RecognitionabstractMany previous methods have demonstrated the importance of considering semantically relevant objects for carrying out video-based human activity recognition, yet none of the methods have harvested the power of large text corpora to relate the objects and the activities to be transferred into learning a unified deep convolutional neural network. We present a novel activity recognition CNN which co-learns the object recognition task in an end-to-end multitask learning scheme to improve upon the baseline activity recognition performance. We further improve upon the multitask learning approach by exploiting a text-guided semantic space to select the most relevant objects with respect to the target activities. To the best of our knowledge, we are the first to investigate this approach. Sungmin Eum, Christopher Reale, Heesung Kwon, Claire Bonial, Clare R. Voss |
ICASSP | 3 |
| 2019 | DOD-CNN: Doubly-injecting Object Information for Event RecognitionabstractRecognizing an event in an image can be enhanced by detecting relevant objects in two ways: 1) indirectly utilizing object detection information within the unified architecture or 2) directly making use of the object detection output results. We introduce a novel approach, referred to as Doubly-injected Object Detection CNN (DOD-CNN), exploiting the object information in both ways for the task of event recognition. The structure of this network is inspired by the Integrated Object Detection CNN (IOD-CNN) where object information is indirectly exploited by the event recognition module through the shared portion of the network. In the DOD-CNN architecture, the intermediate object detection outputs are directly injected into the event recognition network while keeping the indirect sharing structure inherited from the IOD-CNN, thus being `doubly-injected'. We also introduce a batch pooling layer which constructs one representative feature map from multiple object hypotheses. We have demonstrated the effectiveness of injecting the object detection information in two different ways in the task of malicious event recognition. Hyungtae Lee, Sungmin Eum, Heesung Kwon |
ICASSP | 3 |
| 2019 | Delving Into Robust Object Detection From Unmanned Aerial Vehicles: A Deep Nuisance Disentanglement ApproachabstractObject detection from images captured by Unmanned Aerial Vehicles (UAVs) is becoming increasingly useful. Despite the great success of the generic object detection methods trained on ground-to-ground images, a huge performance drop is observed when they are directly applied to images captured by UAVs. The unsatisfactory performance is owing to many UAV-specific nuisances, such as varying flying altitudes, adverse weather conditions, dynamically changing viewing angles, etc. Those nuisances constitute a large number of fine-grained domains, across which the detection model has to stay robust. Fortunately, UAVs will record meta-data that depict those varying attributes, which are either freely available along with the UAV images, or can be easily obtained. We propose to utilize those free meta-data in conjunction with associated UAV images to learn domain-robust features via an adversarial training framework dubbed Nuisance Disentangled Feature Transform (NDFT), for the specific challenging problem of object detection in UAV images, achieving a substantial gain in robustness to those nuisances. We demonstrate the effectiveness of our proposed algorithm, by showing state-of-the- art performance (single model) on two existing UAV-based object detection benchmarks. The code is available at https://github.com/TAMU-VITA/UAV-NDFT. Zhenyu Wu 0002, Karthik Suresh 0001, Priya Narayanan, Hongyu Xu, Heesung Kwon, Zhangyang Wang |
ICCV | 5 |
| 2019 | Is Pretraining Necessary for hyperspectral image classification?abstractWe address two questions for training a convolutional neural network (CNN) for hyperspectral image classification: i) is it possible to build a pre-trained network? and ii) is the pre-training effective in furthering the performance? To answer the first question, we have devised an approach that pre-trains a network on multiple source datasets that differ in their hyperspectral characteristics and fine-tunes on a target dataset. This approach effectively resolves the architectural issue that arises when transferring meaningful information between the source and the target networks. To answer the second question, we carried out several ablation experiments. Based on the experimental results, a network trained from scratch performs as good as a network fine-tuned from a pre-trained network. However, we observed that pre-training the network has its own advantage in achieving better performances when deeper networks are required. Hyungtae Lee, Sungmin Eum, Heesung Kwon |
IGARSS | 3 |
| 2019 | A RUGD Dataset for Autonomous Navigation and Visual Perception in Unstructured Outdoor EnvironmentsabstractResearch in autonomous driving has benefited from a number of visual datasets collected from mobile platforms, leading to improved visual perception, greater scene understanding, and ultimately higher intelligence. However, this set of existing data collectively represents only highly structured, urban environments. Operation in unstructured environments, e.g., humanitarian assistance and disaster relief or off-road navigation, bears little resemblance to these existing data. To address this gap, we introduce the Robot Unstructured Ground Driving (RUGD) dataset with video sequences captured from a small, unmanned mobile robot traversing in unstructured environments. Most notably, this data differs from existing autonomous driving benchmark data in that it contains significantly more terrain types, irregular class boundaries, minimal structured markings, and presents challenging visual properties often experienced in off road navigation, e.g., blurred frames. Over 7, 000 frames of pixel-wise annotation are included with this dataset, and we perform an initial benchmark using state-of-the-art semantic segmentation architectures to demonstrate the unique challenges this data introduces as it relates to navigation tasks. Maggie B. Wigness, Sungmin Eum, John G. Rogers III, David K. Han, Heesung Kwon |
IROS | 5 |
| 2019 | Cooperative Learning for Multi-perspective Image ClassificationabstractData gathered from dense sensor networks is often highly correlated across collocated sensors. For example, in video surveillance networks, multiple cameras can observe the same object from multiple angles. Despite the spatial and temporal dependencies between video frames from different cameras, the deep learning algorithms used in today's video analytics problems treat all frames as independent inputs to image classifiers and object detectors. The outputs of these classifiers and detectors on multiple frames are then fused to extract information about the underlying sensor region. We present a cooperative learning framework that allows sensors to train deep learning systems on their own local data and compressed insights from neighboring sensors' input data. This system fuses sensor data before classification to allow learning agents to more naturally handle correlated inputs and cooperate with neighboring sensors with minimal communication costs. Nick Nordlund, Heesung Kwon, Leandros Tassiulas |
SMARTCOMP | 2 |
| 2019 | Online Distributed Analytics at the Edge with Multiple Service GradesabstractIn this paper, we study the problem of how to allocate bandwidth and computation resources to deliver data analytics services at the edge. The types of services we envision consist of a chain of tasks that must be carried out sequentially, and where the number of tasks executed in the chain determines the grade in which a service is delivered. An example of such type of service is video analytics where different deep-learning algorithms are combined to provide a more accurate description of a scene. The contributions of the paper are to formulate the static resource allocation problem as a linear program, to discuss the challenges of static formulations in dynamic settings, and to propose a control-type formulation that uses approximate system dynamics and time-varying cost functions. The work also highlights the need for policies that can operate the network and learn its characteristics simultaneously. Víctor Valls, Geeth de Mel, Heesung Kwon, Leandros Tassiulas |
SMARTCOMP | 3 |
| 2018 | Exploitation of Semantic Keywords for Malicious Event ClassificationabstractLearning an event classifier is challenging when the scenes are semantically different but visually similar. However, as humans, we typically handle such tasks painlessly by adding our background semantic knowledge. Motivated by this observation, we aim to provide an empirical study about how additional information such as semantic keywords can boost up the discrimination of such events. To demonstrate the validity of this study, we first construct a novel Malicious Crowd Dataset containing crowd images with two events, benign and malicious, which look visually similar. Note that the primary focus of this paper is not to provide the state-of-the-art performance on this dataset but to show the beneficial aspects of using semantically-driven keyword information. By leveraging crowd-sourcing platforms, such as Amazon Mechanical Turk, we collect semantic keywords associated with images and then subsequently identify a subset of keywords (e.g. police, fire, etc.) unique to specific events. We first show that by using recently introduced attention models, a naive CNN-based event classifier actually learns to primarily focus on local attributes associated with the discriminant semantic keywords identified by the Turks. We further show that incorporating the keyword-driven information into early-and late-fusion approaches can significantly enhance malicious event classification. Hyungtae Lee, Sungmin Eum, Joel Levis, Heesung Kwon, James Michaelis, Michael Kolodny |
ICASSP | 4 |
| 2018 | Cross-Domain CNN for Hyperspectral Image ClassificationabstractIn this paper, we address the dataset scarcity issue with the hyperspectral image classification. As only a few thousands of pixels are available for training, it is difficult to effectively learn high-capacity Convolutional Neural Networks (CNNs). To cope with this problem, we propose a novel cross-domain CNN containing the shared parameters which can co-learn across multiple hyperspectral datasets. the network also contains the non-shared portions designed to handle the dataset-specific spectral characteristics and the associated classification tasks. Our approach is the first attempt to learn a CNN for multiple hyperspectral datasets, in an end-to-end fashion. Moreover, we have experimentally shown that the proposed network trained on three of the widely used datasets outperform all the baseline networks which are trained on single dataset. Hyungtae Lee, Sungmin Eum, Heesung Kwon |
IGARSS | 3 |
| 2017 | Deep Network Shrinkage Applied to Cross-Spectrum Face RecognitionabstractIn recent years, deep learning has emerged as a dominant methodology in virtually all machine learning problems. While it has been shown to produce state-of-the-art results for a variety of applicatons (including face recognition and heterogeneous face recognition), one aspect of deep networks that has not been extensively researched is how to determine the optimal network structure. This problem is generally solved by ad hoc methods. In this work we address a subproblem of this task: determining the breadth (number of nodes) of each layer. We show how to use group-sparsity-inducing regularization to effectively replace these hyper-parameters with a single hyperparameter which can be determined by cross-validation. We demonstrate our method by using it to reduce the size of networks on two commonly used NIR face datasets. Christopher Reale, Hyungtae Lee, Heesung Kwon, Rama Chellappa |
FG | 3 |
| 2017 | Enhanced object detection via fusion with prior beliefs from image classificationabstractIn this paper, we introduce a novel fusion method that can enhance object detection performance by fusing decisions from two different types of computer vision tasks: object detection and image classification. In the proposed work, the class label of an image obtained from the image classification task is viewed as prior knowledge about existence or non-existence of certain objects. The prior knowledge is then fused with the decisions of object detection to improve detection accuracy by mitigating false positives of an object detector that are strongly contradicted with the prior knowledge. A recently introduced novel fusion approach called dynamic belief fusion (DBF) is used to fuse the detector output with the classification prior. Experimental results show that the detection performance of all the detection algorithms used in the proposed work is improved on benchmark datasets via the proposed fusion framework. Yilun Cao, Hyungtae Lee, Heesung Kwon |
ICIP | 3 |
| 2017 | IOD-CNN: Integrating object detection networks for event recognitionabstractMany previous methods have showed the importance of considering semantically relevant objects for performing event recognition, yet none of the methods have exploited the power of deep convolutional neural networks to directly integrate relevant object information into a unified network. We present a novel unified deep CNN architecture which integrates architecturally different, yet semantically-related object detection networks to enhance the performance of the event recognition task. Our architecture allows the sharing of the convolutional layers and a fully connected layer which effectively integrates event recognition, rigid object detection and non-rigid object detection. Sungmin Eum, Hyungtae Lee, Heesung Kwon, David S. Doermann |
ICIP | 3 |
| 2017 | Going Deeper With Contextual CNN for Hyperspectral Image ClassificationabstractIn this paper, we describe a novel deep convolutional neural network (CNN) that is deeper and wider than other existing deep networks for hyperspectral image classification. Unlike current state-of-the-art approaches in CNN-based hyperspectral image classification, the proposed network, called contextual deep CNN, can optimally explore local contextual interactions by jointly exploiting local spatio-spectral relationships of neighboring individual pixel vectors. The joint exploitation of the spatio-spectral information is achieved by a multi-scale convolutional filter bank used as an initial component of the proposed CNN pipeline. The initial spatial and spectral feature maps obtained from the multi-scale filter bank are then combined together to form a joint spatio-spectral feature map. The joint feature map representing rich spectral and spatial properties of the hyperspectral image is then fed through a fully convolutional network that eventually predicts the corresponding label of each pixel vector. The proposed approach is tested on three benchmark data sets: the Indian Pines data set, the Salinas data set, and the University of Pavia data set. Performance comparison shows enhanced classification performance of the proposed approach over the current state-of-the-art on the three data sets. Hyungtae Lee, Heesung Kwon |
IEEE Trans. Image Process. | 2 |
| 2016 | Weakly Supervised Localization Using Deep Feature Maps
Archith J. Bency, Heesung Kwon, Hyungtae Lee, S. Karthikeyan 0001, B. S. Manjunath |
ECCV (1) | 2 |
| 2016 | DTM: Deformable template matchingabstractA novel template matching algorithm that can incorporate the concept of deformable parts, is presented in this paper. Unlike the deformable part model (DPM) employed in object recognition, the proposed template-matching approach called Deformable Template Matching (DTM) does not require a training step. Instead, deformation is achieved by a set of predefined basic rules (e.g. the left sub-patch cannot pass across the right patch). Experimental evaluation of this new method using the PASCAL VOC 07 dataset demonstrated substantial performance improvement over conventional template matching algorithms. Additionally, to confirm the applicability of DTM, the concept is applied to the generation of a rotation-invariant SIFT descriptor. Experimental evaluation employing deformable matching of SIFT features shows an increased number of matching features compared to a conventional SIFT matching. Hyungtae Lee, Heesung Kwon, Ryan M. Robinson, William D. Nothwang |
ICASSP | 2 |
| 2016 | Contextual deep CNN based hyperspectral classificationabstractIn this paper, we describe a novel deep convolutional neural networks (CNN) based approach called contextual deep CNN that can jointly exploit spatial and spectral features for hyperspectral image classification. The contextual deep CNN first concurrently applies multiple 3-dimensional local convolutional filters with different sizes jointly exploiting spatial and spectral features of a hyperspectral image. The initial spatial and spectral feature maps obtained from applying the variable size convolutional filters are then combined together to form a joint spatio-spectral feature map. The joint feature map representing rich spectral and spatial properties of the hyperspectral image is then fed through fully convolutional layers that eventually predict the corresponding label of each pixel vector. The proposed approach is tested on two benchmark datasets: the Indian Pines dataset and the Pavia University scene dataset. Performance comparison shows enhanced classification performance of the proposed approach over the current state of the art on both datasets. Hyungtae Lee, Heesung Kwon |
IGARSS | 2 |
| 2016 | Task-conversions for integrating human and machine perception in a unified taskabstractThe different strategies for feature extraction and synthesis employed by humans and computers are often complementary, hence combining the two into an integrated object recognition system may considerably improve performance over either used in isolation. Rapid Serial Visual Presentation (RSVP) is one well-established technique that has shown promise integrating human perception into a machine perception system. In this paper, we apply computer vision techniques to image data filtered through human RSVP. We introduce “task conversions” to integrate the two modalities, applying the precise localization capabilities of computer vision with the detection capabilities of RSVP. We employ naive Bayesian fusion and a novel method, dynamic belief fusion (DBF), in a joint scheme as fusion approaches. Preliminary experiments demonstrate that DBF extracts complementary information from both human and machine sources to improve performance for both target classification and object detection. Hyungtae Lee, Heesung Kwon, Ryan M. Robinson, Daniel Donavanik, William D. Nothwang, Amar R. Marathe |
IROS | 2 |
| 2016 | Dynamic belief fusion for object detectionabstractA novel approach for the fusion of heterogeneous object detection methods is proposed. In order to effectively integrate the outputs of multiple detectors, the level of ambiguity in each individual detection score is estimated using the precision/recall relationship of the corresponding detector. The main contribution of the proposed work is a novel fusion method, called Dynamic Belief Fusion (DBF), which dynamically assigns probabilities to hypotheses (target, non-target, intermediate state (target or non-target)) based on confidence levels in the detection results conditioned on the prior performance of individual detectors. In DBF, a joint basic probability assignment, optimally fusing information from all detectors, is determined by the Dempster's combination rule, and is easily reduced to a single fused detection score. Experiments on ARL and PASCAL VOC 07 datasets demonstrate that the detection accuracy of DBF is considerably greater than conventional fusion approaches as well as individual detectors used for the fusion. Hyungtae Lee, Heesung Kwon, Ryan M. Robinson, William D. Nothwang, Amar M. Marathe |
WACV | 2 |
| 2015 | Shapely value based random subspace selection for hyperspectral image classificationabstractIn this paper, an algorithm to randomly select feature sub-spaces for hyperspectral image classification using the principle of coalition game theory is presented. The feature selection algorithms associated with non-linear kernel based Support Vector Machines (SVM) are either NP-hard or greedy and hence, not very optimal. To deal with this problem, a metric based on the principles of coalition game theory called Shapely value and a sampling approximation is used to determine the contributions of individual features towards the classification task. Feature subsets are randomly drawn from a probability distribution function generated using normalized Shapely values of the individual features. These feature subsets are then used to build kernels corresponding to individual weak classifiers in the Sparse Kernel-based Ensemble Learning (SKEL) framework. By weighting the kernels optimally and sparsely, a small number of useful subsets of features are selected which improve the generalization performance of the ensemble classifier. The algorithm is applied on real hyper-spectral datasets and the results are presented in the paper. Prudhvi Gurram, Heesung Kwon, Charles Davidson |
IGARSS | 2 |
| 2015 | Human-autonomy sensor fusion for rapid object detectionabstractHuman-autonomy sensor fusion is an emerging technology with a wide range of applications, including object detection/recognition, surveillance, collaborative control, and prosthetics. For object detection, humans and computer-vision-based systems employ different strategies to locate targets, likely providing complementary information. However, little effort has been made in combining the outputs of multiple autonomous detectors and multiple human-generated responses. This paper presents a method for integrating several sources of human- and autonomy-generated information for rapid object detection tasks. Human electroencephalography (EEG) and button-press responses from rapid serial visual presentation (RSVP) experiments are fused with outputs from trained object detection algorithms. Three fusion methods—Bayesian, Dempster-Shafer, and Dynamic Dempster-Shafer—are implemented for comparison. Results demonstrate that fusion of these human classifiers with computer-vision-based detectors improves object detection accuracy over purely computer-vision-based detection (5% relative increase in mean average precision) and the best individual computer vision algorithm (28% relative increase in mean average precision). Computer vision fused with button press response and/or the XDAWN + Bayesian Linear Discriminant Analysis neural classifier provides considerable improvement, while computer vision fused with other neural classifiers provides little or no improvement. Of the three fusion methods, Dynamic Dempster-Shafer Theory (DDST) Fusion exhibits the greatest performance in this application. Ryan M. Robinson, Hyungtae Lee, Michael J. McCourt, Amar R. Marathe, Heesung Kwon, Chau Ton, William D. Nothwang |
IROS | 5 |
| 2014 | Coalition game theory based feature subset selection for hyperspectral image classificationabstractIn this paper, an algorithm to select feature subsets for hyper-spectral image classification using the principle of coalition game theory is presented. The feature selection algorithms associated with non-linear kernel based Support Vector Machines (SVM) are either NP-hard or greedy and hence, not very optimal. To deal with this problem, a metric based on the principles of coalition game theory called Shapely value and a sampling approximation is used to determine the contribution of a subset of features towards the classification task. Starting with a few subsets of features, we successively partition each of them into smaller parts if the smaller parts contribute more than a pre-determined threshold compared to the parent subset. The algorithm is terminated when the subsets of features do not change from one iteration to the next. The final subsets of features are then used in multiple kernels and sparse weights of these kernels are optimally learned to build a maximum margin classifier. The algorithm is applied on real hyperspectral datasets and the results are presented in the paper. Prudhvi Gurram, Heesung Kwon |
IGARSS | 2 |
| 2013 | Contextual SVM Using Hilbert Space Embedding for Hyperspectral ClassificationabstractIn this letter, a kernel-based contextual classification approach built on the principle of a newly introduced mapping technique, called Hilbert space embedding, is proposed. The proposed technique, called contextual support vector machine (SVM), is aimed at jointly exploiting both local spectral and spatial information in a reproducing kernel Hilbert space (RKHS) by collectively embedding a set of spectral signatures within a confined local region into a single point in the RKHS that can uniquely represent the corresponding local hyperspectral pixels. Embedding is conducted by calculating the weighted empirical mean of the mapped points in the RKHS to exploit the similarities and variations in the local spectral and spatial information. The weights are adaptively estimated based on the distance between the mapped point in consideration and its neighbors in the RKHS. An SVM separating hyperplane is built to maximize the margin between classes formed by weighted empirical means. The proposed technique showed significant improvement over the composite kernel-based SVM on several hyperspectral images. Prudhvi Gurram, Heesung Kwon |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2013 | Sparse Kernel-Based Ensemble Learning With Fully Optimized Kernel Parameters for Hyperspectral Classification ProblemsabstractRecently, a kernel-based ensemble learning technique for hyperspectral detection/classification problems has been introduced by the authors, to provide robust classification over hyperspectral data with relatively high level of noise and background clutter. The kernel-based ensemble technique first randomly selects spectral feature subspaces from the input data. Each individual classifier, which is in fact a support vector machine (SVM), then independently conducts its own learning within its corresponding spectral feature subspace and hence constitutes a weak classifier. The decisions from these weak classifiers are equally or adaptively combined to generate the final ensemble decision. However, in such ensemble learning, little attempt has been previously made to jointly optimize the weak classifiers and the aggregating process for combining the subdecisions. The main goal of this paper is to achieve an optimal sparse combination of the subdecisions by jointly optimizing the separating hyperplane obtained by optimally combining the kernel matrices of the SVM classifiers and the corresponding weights of the subdecisions required for the aggregation process. Sparsity is induced by applying an$l1$norm constraint on the weighting coefficients. Consequently, the weights of most of the subclassifiers become zero after the optimization, and only a few of the subclassifiers with non-zero weights contribute to the final ensemble decision. Moreover, in this paper, an algorithm to determine the optimal full-diagonal bandwidth parameters of the Gaussian kernels of the individual SVMs is also presented by minimizing the radius-margin bound. The optimized full-diagonal bandwidth Gaussian kernels are used by the sparse SVM ensemble to perform binary classification. The performance of the proposed technique with optimized kernel parameters is compared to that of the one with single-bandwidth parameter obtained using cross-validation by testing them on various data sets. On an average, the proposed sparse kernel-based ensemble learning algorithm with optimized full-diagonal bandwidth parameters shows an improvement of 20$\%$over the existing ensemble learning techniques. Prudhvi Gurram, Heesung Kwon |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2012 | Contextual SVM for hyperspectral classification using Hilbert Space EmbeddingabstractIn this paper, a contextual Support Vector Machine (SVM) technique based on the principle of Hilbert Space Embedding (HSE) of a local hyperspectral data distribution into an Reproducing Kernel Hilbert Space (RKHS) is proposed to optimally exploit the spectral and local spatial information of the hyperspectral image. The idea of embedding is to map hyperspectral pixels in a local neighborhood into a single point in the RKHS that can uniquely represent those pixels collectively. Previously, the authors have employed an HSE called empirical mean map to build the contextual SVM. In this work, a weighted empirical mean map is utilized to exploit the similarities and variation in the local spatial information. For every pixel, a small set of the neighboring pixels in a hyperspectral image are mapped into an RKHS induced by a certain kernel (Eg. Gaussian RBF kernel) and then, the embedded point of these group of pixels is obtained by calculating the weighted empirical mean of these mapped points. The weights are determined based on the distance between the pixel in consideration and its neighbors. An SVM separating hyperplane is built to maximize the margin between classes formed by weighted empirical means. The proposed technique showed significant improvement over the existing contextual and composite kernels on two hyperspectral image data sets. Prudhvi Gurram, Heesung Kwon |
IGARSS | 2 |
| 2012 | Sparse Kernel-Based Hyperspectral Anomaly DetectionabstractIn this letter, a novel ensemble-learning approach for anomaly detection is presented. The proposed technique aims to optimize an ensemble of kernel-based one-class classifiers, such as support vector data description (SVDD) classifiers, by estimating optimal sparse weights of the subclassifiers. In this method, the features of a given multivariate data set representing normalcy are first randomly subsampled into a large number of feature subspaces. An enclosing hypersphere that defines the support of the normalcy data in the reproducing kernel Hilbert space (RKHS) of each respective feature subspace is estimated using standard SVDD. The joint hypersphere in the RKHS of the combined kernel is learned by optimally combining the weighted individual kernels while imposing the$l1$constraint on the combining weights. The joint hypersphere representing the optimal compact support of the multivariate data in the joint RKHS is then used to test a new data point to determine if it belongs to the normalcy data or not. A performance comparison between the proposed algorithm and regular SVDD is reported using hyperspectral image data as well as general multivariate data. Prudhvi Gurram, Heesung Kwon, Timothy Han |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2011 | Generalized optimal Kernel-Based Ensemble Learning for hyperspectral classification problemsabstractIn this paper, a Generalized Kernel-based Ensemble Learning (GKEL) algorithm for hyperspectral classification problems is presented. The proposed algorithm generalizes the Sparse Kernel-based Ensemble Learning (SKEL) technique, developed previously by the authors. SKEL optimally and sparsely weights and aggregates an ensemble of individual SVM classifiers which independently conduct learning within their corresponding randomly selected spectral feature sub-space using a Gaussian kernel. This ensemble decision is fully optimal, if the dimensionality of the randomly selected feature subspaces and the initial number of the sub-classifiers are determined optimally and is sub-optimal, otherwise. This sub optimality issue is addressed by taking a bottom-up approach. Individual sub-classifiers are added one-by-one optimally to the ensemble until the ensemble converges. The feature sub space of each individual classifier is optimally selected. The ensemble is modeled as a Quadratically-Constrained Linear Programming (QCLP) problem and optimized by combining Multiple Kernel Learning (MKL) with a greedy, non-linear integer programming method for non-monotonic sparse feature sub-space selection. Hyperspectral image data as well as multivariate data are used to verify the performance improvement of the proposed GKEL algorithm over SKEL in detecting difficult targets. Prudhvi Gurram, Heesung Kwon |
IGARSS | 2 |
| 2011 | Support-Vector-Based Hyperspectral Anomaly Detection Using Optimized Kernel ParametersabstractIn this letter, a method to optimally determine the kernel bandwidth of the Gaussian radial basis function (RBF) kernel for support vector (SV)-based hyperspectral anomaly detection is presented. In this method, the support of a local background distribution is first nonparametrically learned by a technique called SV data description (SVDD). The SVDD optimally models an enclosing hypersphere around the local background data in a high-dimensional feature space associated with the Gaussian RBF kernel. Any test pixel that lies outside this hypersphere surrounding the local background is considered an anomaly and, hence, a possible target pixel. Considerable improvement in detection performance due to kernel parameter optimization can be seen in the simulation results when the algorithm is applied to hyperspectral images. Prudhvi Gurram, Heesung Kwon |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2010 | A full diagonal bandwidth gaussian kernel SVM based ensemble learning for hyperspectral chemical plume detectionabstractRecently, a sparse kernel-based SVM ensemble learning technique has been introduced by the authors for hyperspectral plume detection/classification. This technique first randomly selects spectral feature subspaces from the input data. Each individual SVM classifier then independently conducts its own learning within its corresponding spectral feature space using a Gaussian kernel with a single bandwidth parameter. Each classifier constitutes a weak classifier. The sub-classifiers are sparsely weighted and aggregated to make an ensemble decision. In this paper, in order to further improve the generalization performance of the ensemble classifier, Gaussian kernel with full diagonal bandwidth parameter matrix is used for each sub-classifier where the parameters are optimally learned by minimizing a bound of the generalization error estimate using a gradient descent algorithm. A performance comparison between the aggregating techniques - sparse kernel-based technique and majority voting with single bandwidth and full diagonal optimized bandwidth parameters as applied to hyperspectral chemical plume detection is presented in the paper. Prudhvi Gurram, Heesung Kwon |
IGARSS | 2 |
| 2010 | Optimal kernel bandwidth estimation for hyperspectral kernel-based anomaly detectionabstractA kernel-based anomaly detection technique called Kernel RX algorithm has been developed earlier by one of the authors, to be used as a prescreening tool that non-linearly detects anomalous pixel spectra in hyperspectral images. Targets of interest are then identified among the prescreened anomalous spectra based on reference spectral information using supervised classification/detection techniques. Kernel RX algorithm uses kernels like the Gaussian radial basis function (RBF) kernel to transform the given data into higher-dimensional (possibly infinite) feature space before detecting the anomalies. The efficiency of the algorithm depends on this transformation which in turn depends on the respective kernel parameters. The Gaussian RBF kernel has a parameter called bandwidth parameter. In this paper, a new method to determine the optimal full diagonal bandwidth parameters of the Gaussian RBF kernel is presented. First, cross-validation technique is used to estimate an optimal single bandwidth parameter. Then, the full diagonal parameters are estimated from this single parameter using the variances of the spectral bands of the hyperspectral image. It will be shown that the optimal full diagonal bandwidth parameters provide a better probability of detection at a given false alarm rate compared to the optimal single bandwidth parameter and other suboptimal bandwidth parameters when tested on hyperspectral imagery for military target detection. Heesung Kwon, Prudhvi Gurram |
IGARSS | 1 |
| 2007 | Kernel Spectral Matched Filter for Hyperspectral Imagery
Heesung Kwon, Nasser M. Nasrabadi |
Int. J. Comput. Vis. | 1 |
| 2007 | Kernel Eigenspace Separation Transform for Subspace Anomaly Detection in Hyperspectral ImageryabstractThis letter proposes a nonlinear version of the eigenspace separation transform (EST) for subspace anomaly detection in hyperspectral imaging. The EST is defined in terms of the eigenvectors of the difference correlation matrix (DCOR) obtained using the data from the two classes. Using ideas found in the machine learning literature (i.e., the kernel trick), a nonlinear version-kernel EST (KEST)-is achieved by expressing the DCOR in terms of dot products in feature space and replacing all dot products with a Mercer kernel function that is defined in terms of input data space. Experimental results indicate that KEST outperforms many other commonly used subspace anomaly detection algorithms. Hirsh R. Goldberg, Heesung Kwon, Nasser M. Nasrabadi |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2006 | Kernel adaptive subspace detector for hyperspectral imageryabstractIn this letter, we present a kernel-based nonlinear version of the adaptive subspace detector (ASD) that implicitly detects signals of interest in a high-dimensional (possibly infinite) feature space associated with a particular nonlinear mapping. In order to address the high dimensionality of the feature space, ASD is first implicitly formulated in the feature space, which is then converted into an expression in terms of kernel functions via the kernel trick property of the Mercer kernels. Experimental results based on simulated data and real hyperspectral imagery show that the proposed kernel-based ASD outperforms the conventional ASD and a nonlinear anomaly detector so called the kernel RX-algorithm. Heesung Kwon, Nasser M. Nasrabadi |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2006 | Kernel Matched Subspace Detectors for Hyperspectral Target DetectionabstractIn this paper, we present a kernel realization of a matched subspace detector (MSD) that is based on a subspace mixture model defined in a high-dimensional feature space associated with a kernel function. The linear subspace mixture model for the MSD is first reformulated in a high-dimensional feature space and then the corresponding expression for the generalized likelihood ratio test (GLRT) is obtained for this model. The subspace mixture model in the feature space and its corresponding GLRT expression are equivalent to a nonlinear subspace mixture model with a corresponding nonlinear GLRT expression in the original input space. In order to address the intractability of the GLRT in the feature space, we kernelize the GLRT expression using the kernel eigenvector representations as well as the kernel trick where dot products in the feature space are implicitly computed by kernels. The proposed kernel-based nonlinear detector, so-called kernel matched subspace detector (KMSD), is applied to several hyperspectral images to detect targets of interest. KMSD showed superior detection performance over the conventional MSD when tested on several synthetic data and real hyperspectral imagery. Heesung Kwon, Nasser M. Nasrabadi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Kernel adaptive subspace detector for hyperspectral target detectionabstractIn this paper, we present a kernel-based nonlinear version of the adaptive subspace detector (ASD) that detects signals of interest in a high dimensional (possibly infinite) feature space associated with a certain nonlinear mapping. In order to address the high dimensionality of the feature space, ASD is first implicitly formulated in the feature space which is then converted into an expression in terms of kernel functions via the kernel trick of the Mercer kernels. The proposed kernel-based ASD (KASD) exploits the nonlinear correlations between the spectral bands that is ignored by the conventional ASD. Experimental results based on the given hyperspectral image show that the proposed KASD outperforms the conventional ASD. Heesung Kwon, Nasser M. Nasrabadi |
ICASSP (4) | 1 |
| 2005 | Kernel spectral matched filter for hyperspectral target detectionabstractIn this paper a kernel-based nonlinear spectral matched filter is introduced for target detection in hyperspectral imagery. The proposed spectral matched filter is defined in a kernel feature space which is equivalent to a nonlinear matched filter in the original input space. This nonlinear spectral matched filter is based on the notion that performing matched filtering in the high dimensional feature space increases the separability of spectral data mainly because it exploits the higher order correlation between the spectral bands. It is also shown that the nonlinear spectral matched filter can easily be implemented in terms of kernel functions using the so called kernel trick property of the Mercer kernels. The kernel version of the nonlinear spectral matched filter is implemented and simulation results on hyperspectral imagery are shown to outperform the linear version. Nasser M. Nasrabadi, Heesung Kwon |
ICASSP (4) | 2 |
| 2005 | Hyperspectral target detection using kernel orthogonal subspace projectionabstractIn this paper, a kernel-based nonlinear version of the orthogonal subspace projection (OSP) classifier is defined in terms of kernel functions. Input data is implicitly mapped into a high dimensional kernel feature space by a nonlinear mapping which is associated with a kernel function. The OSP expression is then derived in the feature space which is kernelized in terms of kernel functions in order to avoid explicit computation in the high dimensional feature space. The resulting kernelized OSP algorithm is equivalent to a nonlinear OSP in the original input space. Experimental results are presented for target detection in hyperspectral imagery and it is shown that the kernel OSP outperforms the conventional OSP classifier. Heesung Kwon, Nasser M. Nasrabadi |
ICIP (2) | 1 |
| 2005 | Kernel RX-algorithm: a nonlinear anomaly detector for hyperspectral imageryabstractWe present a nonlinear version of the well-known anomaly detection method referred to as the RX-algorithm. Extending this algorithm to a feature space associated with the original input space via a certain nonlinear mapping function can provide a nonlinear version of the RX-algorithm. This nonlinear RX-algorithm, referred to as the kernel RX-algorithm, is basically intractable mainly due to the high dimensionality of the feature space produced by the nonlinear mapping function. However, in this paper it is shown that the kernel RX-algorithm can easily be implemented by kernelizing the RX-algorithm in the feature space in terms of kernels that implicitly compute dot products in the feature space. Improved performance of the kernel RX-algorithm over the conventional RX-algorithm is shown by testing several hyperspectral imagery for military target and mine detection. Heesung Kwon, Nasser M. Nasrabadi |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2005 | Kernel orthogonal subspace projection for hyperspectral signal classificationabstractIn this paper, a kernel-based nonlinear version of the orthogonal subspace projection (OSP) operator is defined in terms of kernel functions. Input data are implicitly mapped into a high-dimensional kernel feature space by a nonlinear mapping, which is associated with a kernel function. The OSP expression is then derived in the feature space, which is kernelized in terms of the kernel functions in order to avoid explicit computation in the high-dimensional feature space. The resulting kernelized OSP algorithm is equivalent to a nonlinear OSP in the original input space. Experimental results are presented for detection of roads, roof tops, mines, and targets in hyperspectral imagery, and it is shown that the kernelized OSP method outperforms the conventional OSP approach. Heesung Kwon, Nasser M. Nasrabadi |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2004 | Hyperspectral target detection using kernel matched subspace detectorabstractIn this paper we present a nonlinear realization of a subspace signal detection approach based on the generalized likelihood ratio test (GLRT) - so called matched subspace detectors (MSD). The linear model for MSD is first extended to a high, possibly infinite, dimensional feature space and then the corresponding nonlinear GLRT expression is obtained. In order to address the intractability of the GLRT in the nonlinear feature space we kernelize the nonlinear GLRT using kernel eigenvector representations as well as the kernel trick where dot products in the nonlinear feature space are implicitly computed by kernels. The proposed kernel-based nonlinear detector, so called kernel matched subspace detector (KMSD), is applied to a given hyperspectral imagery - HYDICE (hyperspectral digital imagery collection experiment) images - to detect targets of interest. KMSD showed superior detection performance over MSD for the HYDICE images tested in this paper. Heesung Kwon, Nasser M. Nasrabadi |
ICIP | 1 |
| 2004 | Hyperspectral anomaly detection using kernel rx-algorithmabstractIn this paper we present a nonlinear version of the well-known anomaly detection method, referred to as the RX-algorithm, by extending this algorithm in a feature space associated with the original input space via a certain nonlinear mapping function. An expression for the nonlinear form of the RX-algorithm is derived which is basically intractable mainly due to the high dimensionality of the feature space. We convert the nonlinear RX expression into kernels, which implicitly compute dot products in the nonlinear domain. The proposed kernel RX-algorithm is applied to hyperspectral images for anomaly detection. Improved performance of the kernel RX over the conventional RX is shown for the HYDICE (hyperspectral digital imagery collection experiment) images tested. Heesung Kwon, Nasser M. Nasrabadi |
ICIP | 1 |
| 2004 | Kernel-based subpixel target detection in hyperspectral imagesabstractIn This work we present a nonlinear realization of a signal detection approach that uses the generalized likelihood ratio tests (GLRTs). It is based on converting the linear mixture subspace model, so called matched subspace detector (MSD) into its corresponding nonlinear subspace model. The linear model for the GLRT of MSD is first extended to a high dimensional feature space (equivalent to a non-linear space in the input domain) and then the corresponding nonlinear GLRT expression is obtained. In order to address the intractability of the GLRT in the feature space we kernelize the nonlinear GLRT using kernel eigenvector representations as well as the kernel trick where dot products in the feature space are implicitly computed by kernels. The proposed kernel-based nonlinear detector, so called kernel matched subspace detector (KMSD), is applied to a given hyperspectral imagery - HYDICE (hyperspectral digital imagery collection experiment) images - to detect targets of interest. KMSD showed superior detection performance over MSD for the HYDICE images tested in this paper. Heesung Kwon, Nasser M. Nasrabadi |
IJCNN | 1 |
| 2003 | Projection-based adaptive anomaly detection for hyperspectral imageryabstractAdaptive anomaly detectors that find any materials whose spectral characteristics are out of context with those of the neighboring materials are proposed. We use a dual rectangular window that separates the local area into two regions- the inner window region (IWR) and outer window region (OWR). The statistical differences between the IWR and OWR is exploited by generating projection vectors onto which the IWR and OWR vectors are projected. Anomalies are detected if the projection separation between the IWR and OWR vectors is greater than a predefined threshold. Four different methods are used to produce the projection vectors. The proposed anomaly detectors have been applied to HYDICE (HYper-spectral Digital Imagery Collection Experiment) images and detection performance for each method has been measured. Heesung Kwon, Sandor Z. Der, Nasser M. Nasrabadi |
ICIP (1) | 1 |
| 2001 | An adaptive segmentation algorithm using iterative local feature extraction for hyperspectral imageryabstractWe present an adaptive segmentation algorithm based on the iterative use of a modified minimum-distance classifier. Local adaptivity is achieved by gradually updating each class centroid over a local region whose size is reduced progressively during a segmentation process. The proposed method provides improved segmentation performance over template matching segmentation techniques because it adapts to the local context. The proposed algorithm can be applied to virtually any hyperspectral image regardless of size, dimensionality, and spectral sensitivity. Experimental results on a set of visible to near-infrared hyperspectral images using both the proposed algorithm and a standard template matching technique are presented. Heesung Kwon, Nasser M. Nasrabadi |
ICIP (1) | 1 |
| 2001 | Unsupervised segmentation algorithm based on an iterative spectral dissimilarity measure for hyperspectral imagery
Heesung Kwon, Sandor Z. Der, Nasser M. Nasrabadi |
VCIP | 1 |
| 2000 | An Adaptive Hierarchical Segmentation Algorithm Based on Quadtree Decomposition for Hyperspectral ImageryabstractWe present an adaptive hierarchical segmentation algorithm based on quadtree decomposition and a modified minimum-distance classifier. The proposed algorithm uses quadtree decomposition because this technique can adapt to the local characteristics of the hyperspectral data. A feature vector (i.e., a class centroid) for each material type is recursively estimated and updated, so that increasingly accurate segmentation results are achieved as the decomposition proceeds. The proposed method provides improved segmentation performance over standard template-matching segmentation techniques because it adapts to the local context. It also imposes a spatial smoothness constraint on the pixel classification that provides spatial continuity during the segmentation process. Both the proposed algorithm and a standard template-matching technique were applied to a set of visible to near-infrared hyperspectral images results are presented. Heesung Kwon, Sandor Z. Der, Nasser M. Nasrabadi |
ICIP | 1 |
| 1998 | Very-Low-Bit-Rate Video Coding using Quadtree Decomposition and Cache-based Vector Quantization
Heesung Kwon, Mahesh Venkatraman, Nasser M. Nasrabadi |
ICIP (3) | 1 |
| 1997 | Very Low Bit-Rate Video Coding Using Variable Block-Size Entropy-Constrained Residual Vector QuantizersabstractWe present a practical video coding algorithm for use at very low bit rates. For efficient coding at very low bit rates, it is important to intelligently allocate bits within a frame, and so a powerful variable-rate algorithm is required. We use vector quantization to encode the motion-compensated residue signal in an H.263-like framework. For a given complexity, it is well understood that structured vector quantizers perform better than unstructured and unconstrained vector quantizers. A combination of structured vector quantizers is used in our work to encode the video sequences. The proposed codec is a multistage residual vector quantizer, with transform vector quantizers in the initial stages. The transform-VQ captures the low-frequency information, using only a small portion of the bit budget, while the later stage residual VQ captures the high-frequency information, using the remaining bits. We used a strategy to adaptively refine only areas of high activity, using recursive decomposition and selective refinement in the later stages. An entropy constraint was used to modify the codebooks to allow better entropy coding of the indexes. We evaluate the performance of the proposed codec, and compare this data with the performance of the H.263-based codec. Experimental results show that the proposed codec delivered significantly better perceptual quality along with better quantitative performance. Heesung Kwon, Mahesh Venkatraman, Nasser M. Nasrabadi |
IEEE J. Sel. Areas Commun. | 1 |
| 1996 | Segmentation based wavelet coding of digital imagesabstractIn this paper, we present a segmentation based wavelet coding scheme, in which an image is segmented into two regions: stationary areas (background) and the areas containing edge information (foreground). These regions are then encoded independently using two dedicated encoders that are optimized for each region. A 2-D edge operator is used for segmenting the image. We use the embedded zerotree wavelet (EZW) algorithm for encoding the background due to its good performance on stationary areas. The foreground area is, however, encoded using a predictive residual vector quantizer (PRVQ). Experimental results show that the proposed technique improves the quality of the reconstructed images, both numerically (in terms of mean square error) and perceptually when compared to EZW at the same bit rate. Euee S. Jang, Heesung Kwon, Lin-Cheng Wang, Syed A. Rizvi, Nasser M. Nasrabadi |
ICASSP | 2 |