David K. Han

dblp:12/5692 · DBLP profile ↗
← Back
58ranked-venue papers
0as first author
19since 2021 · last 2024
0000-0001-5055-5408ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 14 since 2021Artificial intelligence and machine learning · 27 · 12 since 2021Systems, architecture and hardware · 6 · 3 since 2021
YearPublicationVenuePosition
2024 Data Augmentation Pipeline for Enhanced UAV Surveillance
Solmaz Arezoomandan, John Klohoker, David K. Han
ICPR (6)3
2024 Enhancing Object Detection by Leveraging Large Language Models for Contextual Knowledge
Amirreza Rouhi, Diego Alberto Patiño Cortes, David K. Han
ICPR (17)3
2024 Towards high-fidelity facial UV map generation in real-world
Yuanming Li, Jeong-gi Kwak, Bonhwa Ku, David K. Han, Hanseok Ko
Pattern Recognit. Lett.4
2023 Cross-Referencing Self-Training Network for Sound Event Detection in Audio Mixtures
abstract
Sound event detection is an important facet of audio tagging that aims to identify sounds of interest and define both the sound category and time boundaries for each sound event in a continuous recording. With advances in deep neural networks, there has been tremendous improvement in the performance of sound event detection systems, although at the expense of costly data collection and labeling efforts. In fact, current state-of-the-art methods employ supervised training methods that leverage large amounts of data samples and corresponding labels in order to facilitate identification of sound category and time stamps of events. As an alternative, the current study proposes a semi-supervised method for generating pseudo-labels from unsupervised data using a student-teacher scheme that balances self-training and cross-training. Additionally, this paper explores post-processing which extracts sound intervals from network prediction, for further improvement in sound event detection performance. The proposed approach is evaluated on sound event detection task for the DCASE2020 challenge. The results of these methods on both "validation" and "public evaluation" sets of DESED database show significant improvement compared to the state-of-the art systems in semi-supervised learning.
Sangwook Park 0002, David K. Han, Mounya Elhilali
IEEE Trans. Multim.2
2022 Injecting 3D Perception of Controllable NeRF-GAN into StyleGAN for Editable Portrait Image Synthesis
Jeong-gi Kwak, Yuanming Li, Dongsik Yoon, David K. Han, Hanseok Ko
ECCV (17)5
2022 3D Human Motion Generation from the Text Via Gesture Action Classification and the Autoregressive Model
abstract
In this paper, a deep learning-based model for 3D human motion generation from the text is proposed via gesture action classification and an autoregressive model. The model focuses on generating special gestures that express human thinking, such as waving and nodding. To achieve the goal, the proposed method predicts expression from the sentences using a text classification model based on a pretrained language model and generates gestures using the gate recurrent unit-based autoregressive model. Especially, we proposed the loss for the embedding space for restoring raw motions and generating intermediate motions well. Moreover, the novel data augmentation method and stop token are proposed to generate variable length motions. To evaluate the text classification model and 3D human motion generation model, a gesture action classification dataset and action-based gesture dataset are collected. With several experiments, the proposed method successfully generates perceptually natural and realistic 3D human motion from the text. Moreover, we verified the effectiveness of the proposed method using a public-available action recognition dataset to evaluate cross-dataset generalization performance.
Gwantae Kim, Youngsuk Ryu, Junyeop Lee, David K. Han, Jeongmin Bae 0006, Hanseok Ko
ICIP4
2022 DIFAI: Diverse Facial Inpainting using StyleGAN Inversion
abstract
Image inpainting is an old problem in computer vision that restores occluded regions and completes damaged images. In the case of facial image inpainting, most of the methods generate only one result for each masked image, even though there are other reasonable possibilities. To prevent any potential biases and unnatural constraints stemming from generating only one image, we propose a novel framework for diverse facial inpainting exploiting the embedding space of StyleGAN. Our framework employs pSp encoder and SeFa algorithm to identify semantic components of the StyleGAN embeddings and feed them into our proposed SPARN decoder that adopts region normalization for plausible inpainting. We demonstrate that our proposed method outperforms several state-of-the-art methods.
Dongsik Yoon, Jeong-gi Kwak, Yuanming Li, David K. Han, Hanseok Ko
ICIP4
2022 Discriminatory and Orthogonal Feature Learning for Noise Robust Keyword Spotting
abstract
Keyword Spotting (KWS) is an essential component in a smart device for alerting the system when a user prompts it with a command. As these devices are typically constrained by computational and energy resources, the KWS model should be designed with a small footprint. In our previous work, we developed lightweight dynamic filters which extract a robust feature map within a noisy environment. The learning variables of the dynamic filter are jointly optimized with KWS weights by using Cross-Entropy (CE) loss. CE loss alone, however, is not sufficient for high performance when the SNR is low. In order to train the network for more robust performance in noisy environments, we introduce the LOw Variant Orthogonal (LOVO) loss. The LOVO loss is composed of a triplet loss applied on the output of the dynamic filter, a spectral norm-based orthogonal loss, and an inner class distance loss applied in the KWS model. These losses are particularly useful in encouraging the network to extract discriminatory features in unseen noise environments.
Kyungdeuk Ko, David K. Han, Hanseok Ko
IEEE Signal Process. Lett.3
2021 CPNet: Cross-Parallel Network for Efficient Anomaly Detection
abstract
Anomaly detection in video streams is a challenging problem because of the scarcity of abnormal events and the difficulty of accurately annotating them. To alleviate these issues, unsupervised learning-based prediction methods have been previously applied. These approaches train the model with only normal events and predict a future frame from a sequence of preceding frames by use of encoder-decoder architectures so that they result in small prediction errors on normal events but large errors on abnormal events. The architecture, however, comes with the computational burden as some anomaly detection tasks require low computational cost without sacrificing performance. In this paper, Cross-Parallel Network (CPNet) for efficient anomaly detection is proposed here to minimize computations without performance drops. It consists of N smaller parallel U-Net, each of which is designed to handle a single input frame, to make the calculations significantly more efficient. Additionally, an inter-network shift module is incorporated to capture temporal relationships among sequential frames to enable more accurate future predictions. The quantitative results show that our model requires less computational cost than the baseline U-Net while delivering equivalent performance in anomaly detection.
Youngsaeng Jin, Jonghwan Hong, David K. Han, Hanseok Ko
AVSS3
2021 Deep Degradation Prior for Real-World Super-Resolution
Kyungdeuk Ko, Bokyeung Lee, Jonghwan Hong, David K. Han, Hanseok Ko
BMVC4
2021 Adaptive Content Feature Enhancement GAN for Multimodal Selfie to Anime Translation
Yuanming Li, Jeong-gi Kwak, Dongsik Yoon, Youngsaeng Jin, David K. Han, Hanseok Ko
BMVC5
2021 Reference Guided Image Inpainting using Facial Attributes
Dongsik Yoon, Youngsaeng Jin, Jeong-gi Kwak, Yuanming Li, David K. Han, Hanseok Ko
BMVC5
2021 Few-Shot Learning for Ct Scan Based Covid-19 Diagnosis
abstract
Coronavirus disease 2019 (COVID-19) is a Public Health Emergency of International Concern infecting more than 40 million people across 188 countries and territories. Chest computed tomography (CT) imaging technique benefits from its high diagnostic accuracy and robustness, it has become an indispensable way for COVID-19 mass testing. Recently, deep learning approaches have become an effective tool for automatic screening of medical images, and it is also being considered for COVID-19 diagnosis. However, the high infection risk involved with COVID-19 leads to relative sparseness of collected labeled data limiting the performance of such methodologies. Moreover, accurately labeling CT images require expertise of radiologists making the process expensive and time-consuming. In order to tackle the above issues, we propose a supervised domain adaption based COVID-19 CT diagnostic method which can perform effectively when only a small samples of labeled CT scans are available. To compensate for the sparseness of labeled data, the proposed method utilizes a large amount of synthetic COVID-19 CT images and adjusts the networks from the source domain (synthetic data) to the target domain (real data) with a cross-domain training mechanism. Experimental results show that the proposed method achieves state-of-the- art performance on few-shot COVID-19 CT imaging based diagnostic tasks.
Yifan Jiang 0002, Hanseok Ko, David K. Han
ICASSP4
2021 Self-Training for Sound Event Detection in Audio Mixtures
abstract
Sound event detection (SED) takes on the task of identifying presence of specific sound events in a complex audio recording. SED has tremendous implications in video analytics, smart speaker algorithms and audio tagging. Recent advances in deep learning have afforded remarkable advances in performance of SED systems; albeit at the cost of extensive labeling efforts to train supervised methods using fully described sound class labels and timestamps. In order to address limitations in availability of training data, this work proposes a self-training technique to leverage unlabeled datasets in supervised learning using pseudo label estimation. This approach proposes a dual-term objective function: a classification loss for the original labels and expectation loss for pseudo labels. The proposed self training technique is applied to sound event detection in the context of the DCASE 2020 challenge, and reports a notable improvement over the baseline system for this task. The self-training approach is particularly effective in extending the labeled database with concurrent sound events.
Sangwook Park 0002, Ashwin Bellur, David K. Han, Mounya Elhilali
ICASSP3
2021 SpecMix : A Mixed Sample Data Augmentation Method for Training with Time-Frequency Domain Features
abstract
A mixed sample data augmentation strategy is proposed to enhance the performance of models on audio scene classification, sound event classification, and speech enhancement tasks.While there have been several augmentation methods shown to be effective in improving image classification performance, their efficacy toward time-frequency domain features of audio is not assured.We propose a novel audio data augmentation approach named "Specmix" specifically designed for dealing with time-frequency domain features.The augmentation method consists of mixing two different data samples by applying time-frequency masks effective in preserving the spectral correlation of each audio sample.Our experiments on acoustic scene classification, sound event classification, and speech enhancement tasks show that the proposed Specmix improves the performance of various neural network architectures by a maximum of 2.7%.
Gwantae Kim, David K. Han, Hanseok Ko
Interspeech2
2021 Consolidating Kinematic Models to Promote Coordinated Mobile Manipulations
abstract
We construct a Virtual Kinematic Chain (VKC) that readily consolidates the kinematics of the mobile base, the arm, and the object to be manipulated in mobile manipulations. Accordingly, a mobile manipulation task is represented by altering the state of the constructed VKC, which can be converted to a motion planning problem, formulated and solved by trajectory optimization. This new VKC perspective of mobile manipulation allows a service robot to (i) produce well-coordinated motions, suitable for complex household environments, and (ii) perform intricate multi-step tasks while interacting with multiple objects without an explicit definition of intermediate goals. In simulated experiments, we validate these advantages by comparing the VKC-based approach with baselines that solely optimize individual components. The results manifest that VKC-based joint modeling and planning promote task success rates and produce more efficient trajectories.
Ziyuan Jiao, Zeyu Zhang 0001, David K. Han, Song-Chun Zhu, Yixin Zhu 0001, Hangxin Liu
IROS4
2021 Efficient Task Planning for Mobile Manipulation: a Virtual Kinematic Chain Perspective
abstract
We present a Virtual Kinematic Chain (VKC) perspective, a simple yet effective method, to improve task planning efficacy for mobile manipulation. By consolidating the kinematics of the mobile base, the arm, and the object being manipulated collectively as a whole, this novel VKC perspective naturally defines abstract actions and eliminates unnecessary predicates in describing intermediate poses. As a result, these advantages simplify the design of the planning domain and significantly reduce the search space and branching factors in solving planning problems. In experiments, we implement a task planner using Planning Domain Definition Language (PDDL) with VKC. Compared with conventional domain definition, our VKC-based domain definition is more efficient in both planning time and memory. In addition, abstract actions perform better in producing feasible motion plans and trajectories. We further scale up the VKC-based task planner in complex mobile manipulation tasks. Taken together, these results demonstrate that task planning using VKC for mobile manipulation is not only natural and effective but also introduces new capabilities.
Ziyuan Jiao, Zeyu Zhang 0001, Weiqi Wang 0004, David K. Han, Song-Chun Zhu, Yixin Zhu 0001, Hangxin Liu
IROS4
2021 Memory-based Semantic Segmentation for Off-road Unstructured Natural Environments
abstract
With the availability of many datasets tailored for autonomous driving in real-world urban scenes, semantic segmentation for urban driving scenes achieves significant progress. However, semantic segmentation for off-road, unstructured environments is not widely studied. Directly applying existing segmentation networks often results in performance degradation as they cannot overcome intrinsic problems in such environments, such as illumination changes. In this paper, a built-in memory module for semantic segmentation is proposed to overcome these problems. The memory module stores significant representations of training images as memory items. In addition to the encoder embedding like items together, the proposed memory module is specifically designed to cluster together instances of the same class even when there are significant variances in embedded features. Therefore, it makes segmentation networks better deal with unexpected illumination changes. A triplet loss is used in training to minimize redundancy in storing discriminative representations of the memory module. The proposed memory module is general so that it can be adopted in a variety of networks. We conduct experiments on the Robot Unstructured Ground Driving (RUGD) dataset and RELLIS dataset, which are collected from off-road, unstructured natural environments. Experimental results show that the proposed memory module improves the performance of existing segmentation networks and contributes to capturing unclear objects over various off-road, unstructured natural scenes with equivalent computational cost and network parameters. As the proposed method can be integrated into compact networks, it presents a viable approach for resource-limited small autonomous platforms.
Youngsaeng Jin, David K. Han, Hanseok Ko
IROS2
2021 TrSeg: Transformer for semantic segmentation
Youngsaeng Jin, David K. Han, Hanseok Ko
Pattern Recognit. Lett.2
2020 CAFE-GAN: Arbitrary Face Attribute Editing with Complementary Attention Feature
Jeong-gi Kwak, David K. Han, Hanseok Ko
ECCV (14)2
2020 Dual Stage Learning Based Dynamic Time-Frequency Mask Generation for Audio Event Classification
Jaihyun Park, David K. Han, Hanseok Ko
INTERSPEECH3
2020 Spectro-Temporal Attention-Based Voice Activity Detection
abstract
Voice Activity Detection (VAD) systems suffer from unexpected and non-stationary background noises at magnitudes sufficiently high to mask the speech signal.Although several methods of increasing the performance of VAD have been proposed, their approaches have yet to mitigate the influence of the background noise itself. This letter proposes an effective noise-robust VAD system approach. The proposed method uses spectral attention and temporal attention through applying a deep learning-based attention mechanism. The proposed method is demonstrated and compared with several other deep learning-based methods in terms of the area under the curve in experiments with either known or unknown noise-added, and real-world noisy data. The results show that the proposed method outperforms the other methods in all the scenarios considered, but moreover generalizes well in environments of unknown or unexpected noise.
Younglo Lee, Jeongki Min, David K. Han, Hanseok Ko
IEEE Signal Process. Lett.3
2020 Amphibian Sounds Generating Network Based on Adversarial Learning
abstract
This letter proposes a generative network based on adversarial learning for synthesizing short-time audio streams and investigates the effectiveness of data augmentation for amphibian call sounds classification. Based on Fourier analysis, the generator is designed by a multi-layer perceptron composed of frequency basis learning layers and an output layer, and a discriminator is constructed by a convolutional neural network. Additionally, regularization on weights is introduced to train the networks with practical data that includes some disturbances. Synthetic audio streams are evaluated by quantitative comparison using inception score, and classification results are compared for real versus synthetic data. In conclusion, the proposed generative network is shown to produce realistic sounds and therefore useful for data augmentation.
Sangwook Park 0002, Mounya Elhilali, David K. Han, Hanseok Ko
IEEE Signal Process. Lett.3
2020 Fusion of Heterogeneous Adversarial Networks for Single Image Dehazing
abstract
In this paper, we propose a novel image dehazing method. Typical deep learning models for dehazing are trained on paired synthetic indoor dataset. Therefore, these models may be effective for indoor image dehazing but less so for outdoor images. We propose a heterogeneous Generative Adversarial Networks (GAN) based method composed of a cycle-consistent Generative Adversarial Networks (CycleGAN) for producing haze-clear images and a conditional Generative Adversarial Networks (cGAN) for preserving textural details. We introduce a novel loss function in the training of the fused network to minimize GAN generated artifacts, to recover fine details, and to preserve color components. These networks are fused via a convolutional neural network (CNN) to generate dehazed image. Extensive experiments demonstrate that the proposed method significantly outperforms the state-of-the-art methods on both synthetic and real-world hazy images.
Jaihyun Park, David K. Han, Hanseok Ko
IEEE Trans. Image Process.2
2019 Self-Subtraction Network for End to End Noise Robust Classification
abstract
Acoustic event classification in surveillance applications typically employs deep learning-based end-to-end learning methods. In real environments, their performance degrades significantly due to noise. While various approaches have been proposed to overcome the noise problem, most of these methodologies rely on supervised learning-based feature representation. Supervised learning system, however, requires a pair of noise free and noisy audio streams. Acquisition of ground truth and noisy acoustic event data requires significant efforts to adequately capture the varieties of noise types for training. This paper proposes a novel supervised learning method for noise robust acoustic event classification in an end-to-end framework named Self Subtraction Network (SSN). SSN extracts noise features from an input audio spectrogram and removes them from the input using LSTMs and an auto-encoder. Our method applied to Urbansound8k dataset with 8 noise types at four different levels demonstrates improved performances compared to the state of the art methods.
David K. Han, Hanseok Ko
AVSS2
2019 A Study of a Cross-Language Perception Based on Cortical Analysis Using Biomimetic STRFs
Sangwook Park 0002, David K. Han, Mounya Elhilali
INTERSPEECH2
2019 A RUGD Dataset for Autonomous Navigation and Visual Perception in Unstructured Outdoor Environments
abstract
Research in autonomous driving has benefited from a number of visual datasets collected from mobile platforms, leading to improved visual perception, greater scene understanding, and ultimately higher intelligence. However, this set of existing data collectively represents only highly structured, urban environments. Operation in unstructured environments, e.g., humanitarian assistance and disaster relief or off-road navigation, bears little resemblance to these existing data. To address this gap, we introduce the Robot Unstructured Ground Driving (RUGD) dataset with video sequences captured from a small, unmanned mobile robot traversing in unstructured environments. Most notably, this data differs from existing autonomous driving benchmark data in that it contains significantly more terrain types, irregular class boundaries, minimal structured markings, and presents challenging visual properties often experienced in off road navigation, e.g., blurred frames. Over 7, 000 frames of pixel-wise annotation are included with this dataset, and we perform an initial benchmark using state-of-the-art semantic segmentation architectures to demonstrate the unique challenges this data introduces as it relates to navigation tasks.
Maggie B. Wigness, Sungmin Eum, John G. Rogers III, David K. Han, Heesung Kwon
IROS4
2019 Relay dueling network for visual tracking with broad field-of-view
abstract
A deep reinforcement‐learning‐based method is presented for visual object tracking tasks. The key objective is to generate a sequence of actions which can move or scale the bounding box in the previous frame to track the target in the current frame. Two intelligent agents are trained to accomplish the above task with a special dueling deep Q‐learning network (Dueling DQN), referred to as a relay dueling network. The proposed model is divided into two agents: the movement agent and the scaling agent. The former performs horizontal or vertical movements and the latter generates scaling actions to change the size of the bounding box. The model has multiple inputs that cover both the bounding box region and the enlarged search region to improve the agents’ perception of the surroundings. The proposed method has a broader field of vision than other similar trackers and its distribution of actions makes it easy to train and improve its tracking performance. The proposed network is tested on popular standard tracker benchmarks and its performance is compared with state‐of‐the‐art trackers. The proposed network is found to be competitive in tracking accuracy and execution effectiveness when compared to conventional methods.
Yifan Jiang 0002, David K. Han, Hanseok Ko
IET Comput. Vis.2
2018 Image fusion and influence function for performance improvement of ATM vandalism action recognition
abstract
Rising rate of vandalism against Automatic Teller Machines (ATMs) is a serious issue within banking industries, prompting needs of a technology to autonomously recognize such events. A vision based fusion method proposed here for classifying these incidents is rooted on visually recognizing heavy or sharp objects potentially used for detecting vandalism actions inferred from optical flow. The recognition performance has been improved chiefly by a novel employment of influence functions in selecting data points of each class useful in learning. We show that the tool recognition performance can be improved when the training data is selected from the ImageNet data set as guided by the influence function.
Jeongseop Yun, Junyeop Lee, Seongkyu Mun, Chul Jin Cho, David K. Han, Hanseok Ko
AVSS5
2018 Hierarchical spatial object detection for ATM vandalism surveillance
abstract
In this paper, a multi-modal classification is proposed for recognizing vandalism against Automatic Teller Machines (ATMs). The visual and textual information base model is developed here to identify external threats on ATMs. The model discriminates threatening behaviors from those that are benign in the image. It provides a level of confidence in the threat recognition by visual object classification coupled with word vector distance measure. To achieve our goal, real-time object detection based on a Region Convolutional Neural Network (R-CNN) first detects objects in the scene and word embedding technique allows to measure distance between the detected object label with predefined tools assumed to be used for vandalizing ATMs. Similarity measure from word embedding not only determines whether the scene may lead to any nefarious activities, but also would provide the level of confidence in occurrence of such incidents. From the experimental evaluation, it is shown that the method is effective and delivers a quantitative measure on decisions it makes.
Junyeop Lee, Chul Jin Cho, David K. Han, Hanseok Ko
AVSS3
2017 Deep Neural Network based learning and transferring mid-level audio features for acoustic scene classification
abstract
Deep Neural Network (DNN) based transfer learning has been shown to be effective in Visual Object Classification (VOC) for complementing the deficit of target domain training samples by adapting classifiers that have been pre-trained for other large-scaled DataBase (DB). Although there exists an abundance of acoustic data, it can also be said that datasets of specific acoustic scenes are sparse for training Acoustic Scene Classification (ASC) models. By exploiting VOC DNN's ability of learning beyond its pre-trained environments, this paper proposes DNN based transfer learning for ASC. Effectiveness of the proposed method is demonstrated on the database of IEEE DCASE Challenge 2016 Task 1 and home surveillance environment via representative experiments. Its improved performance is verified by comparing it to prominent conventional methods.
Seongkyu Mun, Suwon Shon, Wooil Kim, David K. Han, Hanseok Ko
ICASSP4
2017 Subspace projection cepstral coefficients for noise robust acoustic event recognition
abstract
In this paper, a novel feature for noise robust sound event recognition is proposed. The proposed feature is obtained by a two-step procedure. First, a subspace bank is established via target event analysis in complex vector space. Then, by projecting observation vectors onto the subspace bank, noise effect can be reduced while generating discriminant characters originated from differing event subspaces. To demonstrate robustness of the proposed feature, experiments with several classifiers were conducted with varying SNR cases under four noisy environments. According to the experimental results, the proposed method has shown superior performance over prominent conventional methods.
Sangwook Park 0002, Younglo Lee, David K. Han, Hanseok Ko
ICASSP3
2017 Online multi-person tracking with two-stage data association and online appearance model learning
abstract
This study addresses the automatic multi‐person tracking problem in complex scenes from a single, static, uncalibrated camera. In contrast with offline tracking approaches, a novel online multi‐person tracking method is proposed based on a sequential tracking‐by‐detection framework, which can be applied to real‐time applications. A two‐stage data association is first developed to handle the drifting targets stemming from occlusions and people's abrupt motion changes. Subsequently, a novel online appearance learning is developed by using the incremental/decremental support vector machine with an adaptive training sample collection strategy to ensure reliable data association and rapid learning. Experimental results show the effectiveness and robustness of the proposed method while demonstrating its compatibility with real‐time applications.
Jaeyong Ju, Daehun Kim, Bonhwa Ku, David K. Han, Hanseok Ko
IET Comput. Vis.4
2017 A feature descriptor based on the local patch clustering distribution for illumination-robust image matching
Han Wang 0018, Sangmin Yoon, David K. Han, Hanseok Ko
Pattern Recognit. Lett.3
2017 Continuous hand gesture recognition based on trajectory shape information
Cheoljong Yang, David K. Han, Hanseok Ko
Pattern Recognit. Lett.2
2016 Nighttime image dehazing with local atmospheric light and weighted entropy
abstract
In this paper, we propose a novel framework for nighttime image dehazing based on a nighttime haze model which accounts for varying light sources and their glow. First, glow effects are decomposed using relative smoothness. Atmospheric light is then estimated by combining global and local atmospheric lights using a local atmospheric selection map. The transmission is estimated by maximizing an objective function designed with weighted entropy. Finally, haze is removed using two estimated parameters which are atmospheric light and transmission. Experimental results validate the proposed method can achieve haze-free results while alleviating the glow effect.
Dubok Park, David K. Han, Hanseok Ko
ICIP2
2015 Maximum likelihood Linear Dimension Reduction of heteroscedastic feature for robust Speaker Recognition
abstract
This paper analyzes heteroscedasticity in i-vector for robust forensics and surveillance speaker recognition system. Linear Discriminant Analysis (LDA), a widely-used linear dimension reduction technique, assumes that classes are homoscedastic within a same covariance. In this paper it is assumed that general speech utterances contain both homoscedastic and heteroscedastic elements. We show the validity of this assumption by employing several analyses and also demonstrate that dimension reduction using principal components is feasible. To effectively handle the presence of heteroscedastic and homoscedastic elements, we propose a fusion approach of applying both LDA and Heteroscedastic-LDA (HLDA). The experiments are conducted to show its effectiveness and compare to other methods using the telephone database of National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE) 2010 extended.
Suwon Shon, Seongkyu Mun, David K. Han, Hanseok Ko
AVSS3
2015 Acoustic event recognition using dominant spectral basis vectors
Woohyun Choi, Sangwook Park 0002, David K. Han, Hanseok Ko
INTERSPEECH3
2014 Generalized cross-correlation based noise robust abnormal acoustic event localization utilizing non-negative matrix factorization
abstract
In this paper, robust sound source localization for surveillance system is presented. In particular, we propose an algorithm for abnormal acoustic event localization using non-negative matrix factorization based frequency bin weighting. Based on the abnormal acoustic event localization experiments in real acoustic environment, the proposed algorithm's excellent strength is validated in terms of representative performance measures compared to the conventional method.
Sungkyu Moon, Suwon Shon, Wooil Kim, David K. Han
AVSS4
2014 Image enhancement for extremely low light conditions
abstract
In this paper, a novel methodology is proposed for contrast enhancement and noise reduction in very noisy data with low dynamic range on images captured by surveillance camera under extremely low light condition. For the initial noise reduction, a motion adaptive temporal filtering based on the Kalman filter is employed. Then, the denoised image is first inverted and subsequently dehazed as a tone mapping to enhance the visibility based on the observation that the inverted low light image presents quite similar characteristics to hazy image. Finally, the remaining noise is removed using the Non-local means (NLM) denoising step. The overall approach essentially transforms very dark images progressively into more visible form and effectively reduces the high intensity noise generated by the tone mapping process. From the experimental results, effectiveness of the proposed method is validated by comparing with the most recent and leading conventional method.
Dubok Park, Bonhwa Ku, Sangmin Yoon, David K. Han
AVSS5
2014 Single image dehazing with image entropy and information fidelity
abstract
In this paper, we propose a new single image dehazing approach based on information fidelity and image entropy. The global atmospheric light is estimated by quadtree subdivision using transformed hazy images. Then, transmission is estimated by an objective function which is comprised of information fidelity and image entropy at non-overlapped sub-block regions. This is further refined by a Weighted Least Squares (WLS) optimization procedure to alleviate block artifacts. We compared performance of the proposed method with conventional methods to validate its effectiveness in an experiment.
Dubok Park, Hyungjo Park, David K. Han, Hanseok Ko
ICIP3
2014 Single image haze removal using novel estimation of atmospheric light and transmission
abstract
This paper presents a new single image dehaze approach that uses a novel estimation of the atmospheric light and media transmission. Conventional dehaze methods often result in degraded images with low contrast and/or oversaturation of color in some regions. In order to mitigate these problems we use local atmospheric light and estimate the media transmission for each local region by using an objective function represented by modified saturation evaluation metric and intensity difference. Experimental results on a variety of outdoor haze images show that the proposed method achieves excellent restoration in terms of contrast, color fidelity and image visibility.
Hyungjo Park, Dubok Park, David K. Han, Hanseok Ko
ICIP3
2014 Rule-based trajectory segmentation for modeling hand motion trajectory
Jounghoon Beh, David K. Han, Hanseok Ko
Pattern Recognit.2
2014 Hidden Markov Model on a unit hypersphere space for gesture trajectory recognition
Jounghoon Beh, David K. Han, Ramani Durasiwami, Hanseok Ko
Pattern Recognit. Lett.2
2013 Abnormal acoustic event localization based on selective frequency bin in high noise environment for audio surveillance
abstract
In this paper, a method for source localization for surveillance system is presented. In particular, we propose an algorithm for abnormal acoustic event localization based on a novel approach of relevant frequency bin selections by statistical analyses. By means of selective frequency bin, it becomes possible to localize the event more accurately in high noise environment with low computational complexity. The effectiveness is verified through the experimental results in varied noise environments with different levels of Signal to Noise Ratio (SNR).
Suwon Shon, David K. Han, Hanseok Ko
AVSS2
2013 Single image haze removal with WLS-based edge-preserving smoothing filter
abstract
Images captured under hazy conditions have low contrast and poor color. This is primarily due to air-light which degrades image quality according to the transmission map. The approach to enhance these hazy images we introduce here is based on the `Dark-Channel Prior' method with image refinement by the `Weighted Least Square' based edge-preserving smoothing. Local contrast is further enhanced by multi-scale tone manipulation. The proposed method improves the contrast, color and detail for the entire image domain effectively. In the experiment, we compare the proposed method with conventional methods to validate performance.
Dubok Park, David K. Han, Hanseok Ko
ICASSP2
2013 Multimodal image fusion via sparse representation with local patch dictionaries
abstract
Sparse representation is a promising technique for the field of image processing and pattern recognition. It generally exploits over-complete dictionaries which is fixed and known in advance, or learned using training algorithm such as K-SVD. In this paper, we propose a new multimodal image fusion approach based on the sparsity model with local patch dictionaries generated directly from input images. For every location in the image, dictionary is simply constructed with neighboring patches. Experimental results show that the proposed method is efficient and competitive with some existing image fusion methods.
David K. Han, Hanseok Ko
ICIP2
2012 Combining Infrared and Visible Images Using Novel Transform and Statistical Information
abstract
This paper proposes a novel combining method of infrared (IR) and visible images based on a Discrete Wavelet Frame (DWF) approach. In contrast to existing methods, IR image is transformed first using statistical information of the visible image to emphasize relevant information. In a multi-scale domain, we then assign appropriate weights to each pixel of sub-band approximation images through pixel level weighted average for emphasizing relevant information of the IR image while keeping texture information of the visible image. Representative experiments show that the proposed method outperforms exiting methods in image quality.
Bonhwa Ku, David K. Han, Hanseok Ko
AVSS3
2012 Selective Background Adaptation Based Abnormal Acoustic Event Recognition for Audio Surveillance
abstract
In this paper, a method for abnormal acoustic event recognition in an audio surveillance system is presented. We propose a recognition scheme based on a hierarchical structure using a feature combination of Mel-Frequency Cepstral Coefficient (MFCC), timbre, and spectral statistics. A selective background adaptation is proposed for robust abnormal acoustic event recognition in real-world situations. For training, we use a database containing 9 abnormal events (scream, glass breaking, and etc.) and 6 background noise types collected under various surveillance situations. Gaussian Mixture Model (GMM) is considered for classifying the representative abnormal acoustic events and for selecting the background noise for adaptation. Effectiveness of the proposed method is demonstrated via representative experimental results.
Woohyun Choi, Jinsang Rho, David K. Han, Hanseok Ko
AVSS3
2011 Robust background subtraction using data fusion for real elevator scene
abstract
This paper proposes a background subtraction technique robust in elevator environments. Sudden local illumination changes arise frequently in an elevator environment due to opening and closing of the elevator door as well as the inner walls of elevator being made of reflective materials. We present a novel method sequentially fusing a Gaussian mixture model for background subtraction, motion information and a spatial likelihood model based on textured features. Experimental results on real video data demonstrate effectiveness of the proposed approach.
Taeyup Song, David K. Han, Hanseok Ko
AVSS2
2011 Adaptive height-modified histogram equalization and chroma correction in YCbCr color space for fast backlight image compensation
Bonghyup Kang, Changwon Jeon, David K. Han, Hanseok Ko
Image Vis. Comput.3
2010 Robust Dynamic Super Resolution under Inaccurate Motion Estimation
abstract
In image reconstruction, dynamic super resolution image reconstruction algorithms have been investigated to enhance video frames sequentially, where explicit motion estimation is considered as a major factor in the performance. This paper proposes a novel measurement validation method to attain robust image reconstruction results under inaccurate motion estimation. In addition, we present an effective scene change detection method dedicated to the proposed super resolution technique for minimizing erroneous results when abrupt scene changes occur in the video frames. Representative experimental results show excellent performance of the proposed algorithm in terms of the reconstruction quality and processing speed.
Bonhwa Ku, Daesung Chung, Hyunhak Shin, Bonghyup Kang, David K. Han, Hanseok Ko
AVSS6
2010 License Plate Detection Using Local Structure Patterns
abstract
We address the problem of license plate detection in video surveillance systems. The Adaboost based approach, known for relative ease of implementation, makes use of discriminative features such as edges or Haar-like features. In this paper, we propose a novel detection algorithm based on local structure patterns for license plate detection. The proposed algorithm includes post-processing methods to reduce false positive rate using positional and color information of license plates. Experimental results demonstrate effectiveness of the proposed method compared to both the edge and Haar-like feature based methods.
Younghyun Lee, Taeyup Song, Bonhwa Ku, Seoungseon Jeon, David K. Han, Hanseok Ko
AVSS5
2010 Sound source separation by using matched beamforming and time-frequency masking
abstract
This paper proposes a two-stage algorithm to separate two sound sources by using matched beamforming and time-frequency masking techniques. At first, beamforming was used to separate the sound mixtures back to the original sources while preserving the original contents to the maximum extent. The residual interference was then suppressed by the time-frequency masking technique. A sequential least squares method was used in developing a matched beamformer to estimate the relative transfer function (RTF). From experimental results, it has been shown that the proposed method exhibits improved performance in sound source separation compared to conventional methods. Signal enhanced factor (SEF) was improved by an average of 8.39 dB over the baseline.
Jounghoon Beh, Taekjin Lee, David K. Han, Hanseok Ko
IROS3
2008 More powerful discriminants for classifying phylogenetic signals in dinucleotide frequencies
abstract
Microbial DNA fragments are classified according to species using compositional features and “genomic signatures” the oldest of which is the dinucleotide relative abundance profile defined by Karlin et al. More informative features, including higher order signatures, have demonstrated greater species-specificity in comparison to the baseline established by the dinucleotide signature using “delta-distance” to assess dissimilarity; but lack of standard methods has precluded rigorous comparison. We describe a new method for classifier evaluation that reduces any number of pair-wise inter-genomic comparisons to a single performance measure. To illustrate the method, we compare delta-distance to quadratic and linear discriminants prescribed by elementary pattern recognition theory, and find that the quadratic form is significantly more powerful.
Robert H. Baran, Changwon Jeon, David K. Han, Hanseok Ko
ICASSP3
2007 Enabling directional human-robot speech interface via adaptive beamforming and spatial noise reduction
abstract
This paper introduces a home robot application of multi-channel based spatial noise reduction for creating human-robot speech interfaces. A microphone array is employed first to create a speech-only directional conduit, which is realized through adaptive beamforming. Through the directional conduit, the intended speech signal from the desired direction is processed for detection and recognition, while unintended speech-like-sources or undesirable noise from other angles is suppressed. If speech signal is absent among the incoming signals through the conduit, further attenuation of undesirable signals is achieved by using a spatial noise reduction filter. Experimental validation of the technique was conducted using a computer simulation and also an online Samsung AnyBot test. Although the environments exhibited highly non-stationary noise, the method achieved an average speech recognition rate of 87.4% in the case of the computer simulation and 81.6% for the online Samsung AnyBot test. From the cases tested so far, the proposed implementation seems to be effective for practical robot applications in highly non-stationary noise environment.
Jounghoon Beh, Taekjin Lee, Sungjoo Ahn, David K. Han, Hanseok Ko
IROS5
2006 Svm-based Phoneme Classification and Lip Shape Refinement in Real-time Lip-synch System
abstract
In this paper, we present a real time lip-synch system that activates 2-D avatar's lip motion in synch with incoming speech utterance. To achieve the real time operation of the system, the processing time was minimized by "merge and split" procedures resulting in coarse-to-fine phoneme classification. At each stage of phoneme classification, the support vector machine (SVM) method was applied to reduce the computational load while maintaining the desired accuracy. The coarse-to-fine phoneme classification, is accomplished via two_stages of feature extraction: in the first stage, each speech frame is acoustically analyzed for three classes of lip opening using Mel Frequency Cepstral Coefficients (MFCC) as a feature; in the second stage, each frame is further refined for detailed lip shape using formant information. The method was implemented in 2-D lip animation and it was demonstrated that the system was effective in accomplishing real-time lip-synch. This approach was tested on a PC using the Microsoft Visual Studio with an Intel Pentium IV 1.4 Giga Hz CPU and 384 MB RAM. It was observed that the methods of phoneme merging and SVM achieved about twice the speed in recognition than the method employing the Hidden Markov Model (HMM). A typical latency time per a single frame observed using the proposed method was in the order of 18.22 milliseconds while an HMM method under identical conditions resulted about 30.67 milliseconds.
Hanseok Ko, David K. Han
Int. J. Pattern Recognit. Artif. Intell.2
2002 Multiple vehicle tracking based on regional estimation in nighttime CCD images
abstract
In this paper, we develop an image based tracking algorithm of multiple vehicles focused to effective detection and segmentation of moving objects for tracking under poor environmental conditions. In particular, we propose a novel image-tracking algorithm aimed at being robust to occlusion, false alarms, missed detection, and partial or multiple detection of target objects, adverse conditions considered as important issues in Intelligent Transportation System (ITS). Upon applying the Retinex algorithm as preprocessing to reduce the illumination effects at nighttime images, we apply a two-step tracking procedure, performing regional search and track. A regional estimation is first achieved based on a gating using probability data association, to initiate the search. We then invoke an object-oriented multiprocessing for multiple vehicle tracking under poor conditions. Representative experimental results show that the proposed method is effective in CCD images.
Ilkwang Lee, Hanseok Ko, David K. Han
ICASSP3