EDBT 2026 Demo / reviewers in the wild / expert
Hanseok Ko
dblp:67/5314
· DBLP profile ↗
137ranked-venue papers
3as first author
49since 2021 · last 2026
0000-0002-8744-4514ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 89 · 29 since 2021Artificial intelligence and machine learning · 48 · 3 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 12 since 2021Systems, architecture and hardware · 5 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Image dehazing via RGB-FIR multimodal fusion and collaborative learning
Ruolin Du, Han Wang 0018, Wenjie Liu 0004, Guangcheng Wang, Kui Jiang, Hanseok Ko |
Pattern Recognit. | 6 |
| 2025 | EmoBiMamba-TTS: Bidirectional State Space Model for Emotion-Intensity Controllable Text-to-SpeechabstractRecent advances have led to emotion-intensity controllable text-to-speech (TTS) models. However, these systems often suffer from degraded speech quality and unnatural emotional expressions, creating a critical gap between human-like expressiveness and synthesized speech. To address these challenges, we propose EmoBiMamba-TTS, a TTS framework that replaces Transformer architectures with Bidirectional State Space Models (SSMs), enabling linear-time processing and robust context modeling. Our approach incorporates Emotion-Guided Cross Attention (EGCA) for fine-grained emotion intensity control by modeling interactions between emotional cues and acoustic features. Additionally, we introduce dual discriminative learning via a Joint Conditional and Unconditional (JCU) discriminator to enhance speech quality while preserving natural emotional expressions. Experimental results demonstrate that EmoBiMamba-TTS outperforms existing baselines in speech naturalness and computational efficiency. Insung Ham, Bonwha Ku, Hanseok Ko |
ASRU | 3 |
| 2025 | Less is more: Efficient Scene Graph Generation with reparameterizationabstractScene Graph Generation (SGG) aims to identify objects and their relationships in visual scenes but faces two key challenges: high computational overhead, particularly for real-time applications, and the long-tailed distribution of predicates, which biases models toward frequent relationships. To address these challenges, we propose Reparams-SGG, a lightweight and efficient network architecture composed of multi-path residual blocks. This architecture reduces computational overhead by leveraging a reparameterization strategy that minimizes sequential and parallel processing, making it highly efficient during inference. Moreover, we introduce a dynamic focal loss that dynamically adjusts the temperature scale during training to focus learning on rare predicates, promoting progressively unbiased learning. Additionally, we propose a dynamic distribution loss, compensating for learning limitations solely from one-hot distributions under data imbalance conditions. We evaluate our method on the widely-used Visual Genome and the recent PSG dataset. Reparams-SGG achieves competitive performance with significantly fewer parameters than state-of-the-art models, demonstrating its efficiency and suitability for deployment in resource-constrained environments. Jonghwan Hong, Seonghyeok Noh, Bonhwa Ku, Hanseok Ko |
ICASSP | 4 |
| 2025 | Diversity Seeking Techniques for Red-Teaming Large Language ModelsabstractIn this paper, we present new techniques for increasing the diversity of red-teaming prompts generated by automated machine learning-based methods, thereby enabling the discovery of more vulnerabilities in large language models. Using reinforcement learning to train models to output effective prompts for this task results in the models converging deterministically to a single output. Our first technique, which we term Defender, acts by blocking the reward signal for prompts that have already been discovered, thus making what was a stationary problem into a non-stationary problem that compels the reward maximizing algorithm to continually seek new prompts. Our second technique, Teamplay, trains two prompt generation models in tandem and adds the KL divergence between them to the reward in order to make them search in disparate regions of the space of prompts. Our techniques are shown experimentally to increase the effectiveness and diversity of prompts generated by existing reinforcement learning baselines. Seok-Han Lee, Bonhwa Ku, Hanseok Ko |
ICASSP | 3 |
| 2025 | Dropout Connects Transformers and CNNs: Transfer General Knowledge for Knowledge DistillationabstractThanks to their long-range dependencies, transformers obtain state-of-the-art performance in diverse research fields such as computer vision and audio processing. In practical scenarios, convolutional neural networks (CNNs) are used more than Transformers due to their low complexity. So, Transformer-to-CNN knowledge distillation (KD) research, where the Transformer is the teacher and the CNN is the student, is in demand and receiving attention. In Transformer-to-CNN KD training, the capacity gap problem arising from structural differences between the teacher and student networks is the main factor of performance degradation of the student network, unlike homogenous architecture KD. However, previous KD studies transfer all of a teacher's knowledge to the student without consid-ering structural differences. They cannot overcome problems caused by structural differences and show poor performance in Transformer-to-CNN KD. In this paper, we iden-tify general and specific knowledge in feature maps of the teacher and student. General and specific knowledge are the generalized and non-generalized feature representation. We propose a novel KD framework DropKD, which extracts general knowledge from the teacher and student while re-moving specific knowledge and then allows general knowledge of the student network to learn general knowledge of the teacher. Our DropKD empowers the student network to achieve generalization by effectively managing general and specific knowledge. Through extensive experiments on challenging image classification datasets, we demonstrate that the proposed method is superior to existing methods. Bokyeung Lee, Jonghwan Hong, Hyunuk Shin, Bonhwa Ku, Hanseok Ko |
WACV | 5 |
| 2024 | ViVid-1-to-3: Novel View Synthesis with Video Diffusion ModelsabstractGenerating novel views of an object from a single image is a challenging task. It requires an understanding of the underlying 3D structure of the object from an image and ren-dering high-quality, spatially consistent new views. While recent methods for view synthesis based on diffusion have shown great progress, achieving consistency among various view estimates and at the same time abiding by the desired camera pose remains a critical problem yet to be solved. In this work, we demonstrate a strikingly simple method, where we utilize a pre-trained video diffusion model to solve this problem. Our key idea is that synthesizing a novel view could be reformulated as synthesizing a video of a cam-era going around the object of interest-a scanning video-which then allows us to leverage the powerful priors that a video diffusion model would have learned. Thus, to perform novel-view synthesis, we create a smooth camera trajectory to the target view that we wish to render, and denoise using both a view-conditioned diffusion model and a video diffusion model. By doing so, we obtain a highly consistent novel view synthesis, outperforming the state of the art. Jeong-gi Kwak, Erqun Dong, Yuhe Jin, Hanseok Ko, Shweta Mahajan, Kwang Moo Yi |
CVPR | 4 |
| 2024 | Towards Multi-Domain Face Landmark Detection with Synthetic Data from Diffusion ModelabstractRecently, deep learning-based facial landmark detection for in-the-wild faces has achieved significant improvement. However, there are still challenges in face landmark detection in other domains (e.g. cartoon, caricature, etc). This is due to the scarcity of extensively annotated training data. To tackle this concern, we design a two-stage training approach that effectively leverages limited datasets and the pre-trained diffusion model to obtain aligned pairs of landmarks and face in multiple domains. In the first stage, we train a landmark-conditioned face generation model on a large dataset of real faces. In the second stage, we fine-tune the above model on a small dataset of image-landmark pairs with text prompts for controlling the domain. Our new designs enable our method to generate high-quality synthetic paired datasets from multiple domains while preserving the alignment between landmarks and facial features. Finally, we fine-tuned a pre-trained face landmark detection model on the synthetic dataset to achieve multi-domain face landmark detection. Our qualitative and quantitative results demonstrate that our method outperforms existing methods on multi-domain face landmark detection. Yuanming Li, Gwantae Kim, Jeong-gi Kwak, Bonhwa Ku, Hanseok Ko |
ICASSP | 5 |
| 2024 | Prune Channel And Distill: Discriminative Knowledge Distillation For Semantic SegmentationabstractThe goal of knowledge distillation (KD) for semantic segmentation is to transfer discriminative knowledge, enabling the network to distinguish pixels into each class, from a teacher to a student network. Recent KD studies for semantic segmentation fail to convey discriminative knowledge effectively to the student. Consequently, a student network with previous KD cannot generate segmentation maps that effectively distinguish the boundaries of small objects, unlike a teacher network. In this work, we propose a novel KD learning framework, prune channel and distill (PCD), which consists of channel pruning and distillation processes. To transfer the discriminative knowledge of the teacher to the student network, we propose a discriminative score from the perspective of the difference between class responses and student matching distillation, allowing the student to selectively learn channels of pruned feature maps from the teacher. Our PCD directly provides discriminative knowledge from the teacher to the student. In extensive experiments, PCD outperforms state-of-the-art methods on various semantic segmentation datasets. Representative results demonstrate that the proposed method enhances the granularity of the segmentation maps produced by the student network. Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Hanseok Ko |
ICIP | 4 |
| 2024 | Hard Sample-aware Consistency for Low-resolution Facial Expression RecognitionabstractFacial expression recognition (FER) plays a pivotal role in computer vision applications, encompassing video understanding and human-computer interaction. Despite notable advancements in FER, performance still falters when handling low-resolution facial images encountered in real-world scenarios and datasets. While consistency constraint techniques have garnered attention for generating robust convolutional neural network models that accommodate input variations through augmentation, their efficacy is diminished in the realm of low-resolution FER. This decline in performance can be attributed to augmented samples that networks struggle to extract expressive features. In this paper, we identify hard samples that cause an overfitting problem when considering various degrees of resolution and propose novel hard sample-aware consistency (HSAC) loss functions, which include combined attention consistency and label distribution learning. The combined attention consistency aligns an attention map from multi-scale low-resolution images with an appropriate target attention map by combining activation maps from high-resolution and flipped low-resolution images. We measure the classification difficulty for low-resolution face images and adaptively apply label distribution learning by combining the original target and predictions of high-resolution input. Our HSAC empowers the network to achieve generalization by effectively managing hard samples. Extensive experiments on various FER datasets demonstrate the superiority of our proposed method over existing approaches for multiscale low-resolution images. Furthermore, we achieved a new state-of-the-art performance of 90.97% on the original RAF-DB dataset. Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Hanseok Ko |
WACV | 4 |
| 2024 | Noisy label facial expression recognition via face-specific label distribution learning
Hyunuk Shin, Bokyeung Lee, Bonhwa Ku, Hanseok Ko |
Image Vis. Comput. | 4 |
| 2024 | Classification and Magnitude Estimation of Global and Local Seismic Events Using Conformer and Low-Rank Adaptation Fine-TuningabstractClassifying seismic events and estimating their magnitude are crucial topics in the study of seismic waves. Due to the disparities between global and local geologic features, models exclusively trained on global data may exhibit suboptimal performance in local contexts. To solve this problem, this letter proposes a method to evaluate the effectiveness of the low-rank adaptation (LoRA) technique in seismic wave research using the convolution-augmented transformer (Conformer). We simplified and modified the Conformer model, reducing the number of parameters by more than 169-fold, and applied the LoRA technique to this model. Experimental results using the Stanford Earthquake Dataset (STEAD) and the Korean Peninsula Earthquake Dataset (KPED) from 2017 to 2018 showed that fine-tuning the model with a significantly reduced number of parameters using the proposed method is suitable for research on seismological applications. Our approach achieved over 99.99% accuracy in seismic event classification for both datasets. Additionally, our model demonstrated a 7% decrease in mean absolute error (MAE) on the STEAD dataset and a 48% decrease on the KPED dataset compared to the state-of-the-art model. Furthermore, the results also indicate that the Conformer is suitable for seismic event classification and magnitude estimation. The model’s performance in the seismic event classification task decreased by 0.1%, despite reducing the number of retrain parameters by 59 times. Additionally, in the magnitude estimation task, there was an 89-fold decrease in the number of retrain parameters, yet the performance decreased by 1%. Yooseok Jin, Gwantae Kim, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | ConSeisGen: Controllable Synthetic Seismic Waveform GenerationabstractWhile generative adversarial network (GAN) models have shown success in generating synthetic data of acoustic, image, and speech, research on generating seismic waves using GAN is receiving great attention. Although some methods have been successful in generating seismic data, they lack the ability to control the generated seismic waves according to earthquake parameters. This letter proposes a novel approach for controllable seismic wave synthesis using auxiliary classifier GAN (ACGAN). Our method focuses on the generation of synthetic seismic waveforms associated with earthquakes of different epicenteral distances. To incorporate distance information into our model, we introduce a distance regression loss function. In addition, we incorporate a feature-level diversity improvement regularization into our model to enhance the diversity of the generated seismic data. The proposed model was trained on KiK-net datasets, and the quality of the generated data was rigorously validated using various validation methods. Experimental results demonstrate the effectiveness of our proposed model in generating seismic waves by adjusting the earthquake epicenter distance. Yuanming Li, Dongsik Yoon, Bonhwa Ku, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | WaveVC: Speech and Fundamental Frequency Consistent Raw Audio Voice ConversionabstractAbstract Voice conversion (VC) is a task for changing the speech of a source speaker to the target voice while preserving linguistic information of the source speech. The existing VC methods typically use mel-spectrogram as both input and output, so a separate vocoder is required to transform mel-spectrogram into waveform. Therefore, the VC performance varies depending on the vocoder performance, and noisy speech can be generated due to problems such as train-test mismatch. In this paper, we propose a speech and fundamental frequency consistent raw audio voice conversion method called WaveVC. Unlike other methods, WaveVC does not require a separate vocoder and can perform VC directly on raw audio waveform using 1D convolution. This eliminates the issue of performance degradation caused by the train-test mismatch of the vocoder. In the training phase, WaveVC employs speech loss and F0 loss to preserve the content of the source speech and generate F0 consistent speech using the pre-trained networks. WaveVC is capable of converting voices while maintaining consistency in speech and fundamental frequency. In the test phase, the F0 feature of the source speech is concatenated with a content embedding vector to ensure the converted speech follows the fundamental frequency flow of the source speech. WaveVC achieves higher performances than baseline methods in both many-to-many VC and any-to-any VC. The converted samples are available online. Kyungdeuk Ko, Kyungseok Oh, Hanseok Ko |
Neural Process. Lett. | 4 |
| 2024 | Towards high-fidelity facial UV map generation in real-world
Yuanming Li, Jeong-gi Kwak, Bonhwa Ku, David K. Han, Hanseok Ko |
Pattern Recognit. Lett. | 5 |
| 2024 | KFA: Keyword Feature Augmentation for Open Set Keyword SpottingabstractIn recent years, with the advancement of deep learning technology and the emergence of smart devices, there has been a growing interest in keyword spotting (KWS), which is used to activate AI systems with automatic speech recognition and text-to-speech. However, smart devices with KWS often encounter false alarm errors when inputting unexpected words. To address this issue, existing KWS methods typically train non-target words as anunknownclass. Despite these efforts, there is still a possibility that unseen words not trained as part of theunknownclass could be misclassified as one of the target words. To overcome this limitation, we propose a new method named Keyword Feature Augmentation (KFA) for open-set KWS. KFA performs feature augmentation through adversarial learning to increase the loss. The augmented features are constrained within a limited space using label smoothing. Unlike other generative model-based open set recognition (OSR) methods, KFA does not require any additional training parameters or repeated operation for inference. As a result, KFA has achieved a 0.955 AUROC score and 97.34% target class accuracy for Google Speech Commands V1, and a 0.959 AUROC score and 98.17% target class accuracy for Google Speech Commands V2, which is the highest performance when compared to various OSR methods. Kyungdeuk Ko, Bokyeung Lee, Jonghwan Hong, Hanseok Ko |
IEEE Signal Process. Lett. | 4 |
| 2023 | MPE4G : Multimodal Pretrained Encoder for Co-Speech Gesture GenerationabstractWhen virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some input modalities are removed or contain noise, the model may not generate the gestures properly. To acquire robust and generalized encodings, we propose a novel framework with a multimodal pre-trained encoder for co-speech gesture generation. In the proposed method, the multi-head-attention-based encoder is trained with self-supervised learning to contain the information on each modality. Moreover, we collect full-body gestures that consist of 3D joint rotations to improve visualization and apply gestures to the extensible body model. Through the series of experiments and human evaluation, the proposed method renders realistic co-speech gestures not only when all input modalities are given but also when the input modalities are missing or noisy. The project page is available here1 Gwantae Kim, Seonghyeok Noh, Insung Ham, Hanseok Ko |
ICASSP | 4 |
| 2023 | Domain-agnostic single-image super-resolution via a meta-transfer neural architecture search
Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Hanseok Ko |
Neurocomputing | 4 |
| 2023 | Fast Non-Local Attention network for light super-resolution
Jonghwan Hong, Bokyeung Lee, Kyungdeuk Ko, Hanseok Ko |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | Estimation of Magnitude and Epicentral Distance From Seismic Waves Using Deeper CRNNabstractEstimating earthquake parameters is an essential process for an earthquake analysis system. In particular, the magnitude and epicentral distance of an earthquake are the most basic parameters in earthquake analysis. To estimate these, the existing approaches require long waveform data from multiple stations. In this letter, we propose a novel estimation method based on multitasking deep learning and a convolutional recurrent neural network (CRNN) using only a single station. We also use the stream maximum of the input waveform to accurately estimate the earthquake magnitude. Based on the evaluation using the Stanford Earthquake dataset (STEAD) and the Kiban Kyoshin Network (KiK-net) dataset, we verify the high performance of the proposed method. Dongsik Yoon, Yuanming Li, Bonhwa Ku, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Channel Shuffle Neural Architecture Search for Key Word SpottingabstractThe evolution of Network Architecture (NA) allowed Key-Word Spotting (KWS) to exhibit high performance. Generally, NA for KWS is required to have low parameter and computation complexity maintaining high classification performance. Most of the attempts so far have been based on manual approaches, and often the architectures developed from such efforts dwell in the balance of the performance and the network complexity. Then, several KWS models based on Neural Architecture Search (NAS) technique have been proposed. However, these methods do not consider the number of parameters and FLOPs for NA in the search process and manually adjusted the complexity of NA by reducing the number of cells. It may not produce optimized NA with a balance between network complexity and performance. To develop effective network architecture for KWS, network complexity and performance must be considered. In this letter, we propose Channel Shuffle Neural Architecture Search (CSNAS) with channel weights. CSNAS selects whether each channel of the input feature is reflected in the computation or not in the search process and simultaneously controls the number of parameters, FLOPs, and performance. Experiment results show that CSNAS can generate NA that satisfies complexity and performance conditions, and NAs generated by CSNAS outperform state-of-the-art KWS methods. Bokyeung Lee, Gwantae Kim, Hanseok Ko |
IEEE Signal Process. Lett. | 4 |
| 2022 | Injecting 3D Perception of Controllable NeRF-GAN into StyleGAN for Editable Portrait Image Synthesis
Jeong-gi Kwak, Yuanming Li, Dongsik Yoon, David K. Han, Hanseok Ko |
ECCV (17) | 6 |
| 2022 | 3D Human Motion Generation from the Text Via Gesture Action Classification and the Autoregressive ModelabstractIn this paper, a deep learning-based model for 3D human motion generation from the text is proposed via gesture action classification and an autoregressive model. The model focuses on generating special gestures that express human thinking, such as waving and nodding. To achieve the goal, the proposed method predicts expression from the sentences using a text classification model based on a pretrained language model and generates gestures using the gate recurrent unit-based autoregressive model. Especially, we proposed the loss for the embedding space for restoring raw motions and generating intermediate motions well. Moreover, the novel data augmentation method and stop token are proposed to generate variable length motions. To evaluate the text classification model and 3D human motion generation model, a gesture action classification dataset and action-based gesture dataset are collected. With several experiments, the proposed method successfully generates perceptually natural and realistic 3D human motion from the text. Moreover, we verified the effectiveness of the proposed method using a public-available action recognition dataset to evaluate cross-dataset generalization performance. Gwantae Kim, Youngsuk Ryu, Junyeop Lee, David K. Han, Jeongmin Bae 0006, Hanseok Ko |
ICIP | 6 |
| 2022 | DIFAI: Diverse Facial Inpainting using StyleGAN InversionabstractImage inpainting is an old problem in computer vision that restores occluded regions and completes damaged images. In the case of facial image inpainting, most of the methods generate only one result for each masked image, even though there are other reasonable possibilities. To prevent any potential biases and unnatural constraints stemming from generating only one image, we propose a novel framework for diverse facial inpainting exploiting the embedding space of StyleGAN. Our framework employs pSp encoder and SeFa algorithm to identify semantic components of the StyleGAN embeddings and feed them into our proposed SPARN decoder that adopts region normalization for plausible inpainting. We demonstrate that our proposed method outperforms several state-of-the-art methods. Dongsik Yoon, Jeong-gi Kwak, Yuanming Li, David K. Han, Hanseok Ko |
ICIP | 5 |
| 2022 | Efficient Dynamic Filter For Robust and Low Computational Feature ExtractionabstractThe unseen noise signal is difficult to anticipate, and various approaches have been developed to address this issue. In our earlier work, we proposed a lightweight dynamic filter by splitting the filter into kernel and spatial parts. This small footprint model showed robust results in an unseen noisy environment. However, a simple pooling process for dividing the feature would limit the performance. In this paper, we propose an efficient dynamic filter to enhance the performance of the existing dynamic filter. Instead of the simple feature mean, we separate the input features as non-overlapping chunks, and separable convolutions take place for each feature direction. We also propose a dynamic filter based attention pooling method. These methods are applied to the kernel part in our previous work, and experiments are carried out for keyword spotting and speaker verification. We confirm that our proposed method performs better in unseen environments than the recently developed models. Jeong-gi Kwak, Hanseok Ko |
SLT | 3 |
| 2022 | Unsupervised domain adaptation based COVID-19 CT infection segmentation network
Yifan Jiang 0002, Murray H. Loew, Hanseok Ko |
Appl. Intell. | 4 |
| 2022 | Graph Convolution Networks for Seismic Events Classification Using Raw Waveform Data From Multiple StationsabstractThis letter proposes a multiple station-based seismic event classification model using a deep convolution neural network (CNN) and graph convolution network (GCN). To classify various seismic events, such as natural earthquakes, artificial earthquakes, and noise, the proposed model consists of weight-shared convolution layers, graph convolution layers, and fully connected layers. We employed graph convolution layers in order to aggregate features from multiple stations. Representative experimental results with the Korean peninsula earthquake datasets from 2016 to 2019 showed that the proposed model is superior to the single-station based state-of the-art methods. Moreover, the proposed model significantly reduced false alarms when using continuous waveforms of long duration. The code is available at.1 Gwantae Kim, Bonhwa Ku, Jae-Kwang Ahn, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Feature Sparse Coding With CoordConv for Side Scan Sonar Image EnhancementabstractIn this letter, we propose a learning-based compressive sensing (CS) algorithm for denoising side scan sonar (SSS) images. The proposed method is a deep learning-based CS method with enhanced nonlinearity based on an iterative shrinkage and thresholding algorithm (ISTA). Since noise intensity varies depending on the position within SSS images, the proposed method also incorporates CoordConv, which provides coordinate information to the network to help remove nonhomogeneous noise. Through end-to-end training, both the deep learning module and the CS characteristics can be jointly optimized. Representative experimental results show that the proposed method is better than state-of-art methods in terms of both noise removal and memory requirements. Bokyeung Lee, Bonhwa Ku, Wan-Jin Kim, Seungil Kim, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Feedback Network With Curriculum Learning for Earthquake Event ClassificationabstractIn this letter, we propose an earthquake event classification model utilizing a feedback network and curriculum learning (CL). In particular, we propose the CL method with a feature concatenation using gated convolution so that CL can be effectively performed in consideration of the feedback structure. We show that the proposed model is effective through comparison experiments with the existing model using the earthquake dataset for Korean Peninsula and the Stanford earthquake dataset. Jeongki Min, Bonhwa Ku, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Learnable Maximum Amplitude Structure for Earthquake Event ClassificationabstractRecently, most research has been conducted to minimize damage from earthquakes by establishing an early warning system through the analysis of short seismic waves. In particular, deep learning is widely used as it allows to learn complex patterns for earthquake detection from seismic data without complex physical knowledge. In this letter, we propose an improved ConvNetQuake for earthquake event classification by adding learnable features related to the maximum amplitude of the seismic waveform. Since the maximum amplitude is a major factor representing the characteristics of an earthquake, we presented a deep learning structure that can apply this factor in the process of determining whether an earthquake occurs. In the proposed structure, the maximum amplitude is transformed into a feature learned through multi-layer perceptron (MLP) and then concatenates with features extracted through a convolutional neural network (CNN). On the STanford EArthquake Dataset (STEAD) dataset, the proposed method significantly increases the performance for an earthquake event classification than the previous state-of-the-art (SOTA) method by only adding a few parameters. Shou Zhang, Bonhwa Ku, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Discriminatory and Orthogonal Feature Learning for Noise Robust Keyword SpottingabstractKeyword Spotting (KWS) is an essential component in a smart device for alerting the system when a user prompts it with a command. As these devices are typically constrained by computational and energy resources, the KWS model should be designed with a small footprint. In our previous work, we developed lightweight dynamic filters which extract a robust feature map within a noisy environment. The learning variables of the dynamic filter are jointly optimized with KWS weights by using Cross-Entropy (CE) loss. CE loss alone, however, is not sufficient for high performance when the SNR is low. In order to train the network for more robust performance in noisy environments, we introduce the LOw Variant Orthogonal (LOVO) loss. The LOVO loss is composed of a triplet loss applied on the output of the dynamic filter, a spectral norm-based orthogonal loss, and an inner class distance loss applied in the KWS model. These losses are particularly useful in encouraging the network to extract discriminatory features in unseen noise environments. Kyungdeuk Ko, David K. Han, Hanseok Ko |
IEEE Signal Process. Lett. | 4 |
| 2022 | Prototypical Knowledge Distillation for Noise Robust Keyword SpottingabstractKeyword Spotting (KWS) is an essential component in contemporary audio-based deep learning systems and should be of minimal design when the system is working in streaming and on-device environments. We presented a robust feature extraction with a single-layer dynamic convolution model in our previous work. In this letter, we expand our earlier study into multi-layers of operation and propose a robust Knowledge Distillation (KD) learning method. Based on the distribution between class-centroids and embedding vectors, we compute three distinct distance metrics for the KD training and feature extraction processes. The results indicate that our KD method shows similar KWS performance over state-of-the-art models in terms of KWS but with low computational costs. Furthermore, our proposed method results in a more robust performance in noisy environments than conventional KD methods. Gwantae Kim, Bokyeung Lee, Hanseok Ko |
IEEE Signal Process. Lett. | 4 |
| 2022 | Information Bottleneck Measurement for Compressed Sensing Image ReconstructionabstractImage Compressed Sensing (CS) has achieved a lot of performance improvement thanks to advances in deep networks. The CS method is generally composed of a sensing and a decoder. The sensing and decoder networks have a significant impact on the reconstruction performance, and it is obvious that both two networks must be in harmony. However, previous studies have focused on designing the loss function considering only the decoder network. In this paper, we propose a novel training process that can learn sensing and decoder networks simultaneously using Information Bottleneck (IB) theory. By maximizing importance through proposed importance generator, the sensing network is trained to compress important information for image reconstruction of the decoder network. The representative experimental results demonstrate that the proposed method is applied in recently proposed CS algorithms and increases the reconstruction performance with large margin in all CS ratios. Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Bonhwa Ku, Hanseok Ko |
IEEE Signal Process. Lett. | 5 |
| 2021 | Action Recognition with Domain Invariant Features of Skeleton ImageabstractDue to the fast processing-speed and robustness it can achieve, skeleton-based action recognition has recently received the attention of the computer vision community. The recent Convolutional Neural Network (CNN)-based methods have shown commendable performance in learning spatio-temporal representations for skeleton sequence, which use skeleton image as input to a CNN. Since the CNN-based methods mainly encoding the temporal and skeleton joints simply as rows and columns, respectively, the latent correlation related to all joints may be lost caused by the 2D convolution. To solve this problem, we propose a novel CNN-based method with adversarial training for action recognition. We introduce a two-level domain adversarial learning to align the features of skeleton images from different view angles or subjects, respectively, thus further improve the generalization. We evaluated our proposed method on NTU RGB+D. It achieves competitive results compared with state-of-the-art methods and 2.4%, 1.9%accuracy gain than the baseline for cross-subject and cross-view. Yifan Jiang 0002, Hanseok Ko |
AVSS | 3 |
| 2021 | CPNet: Cross-Parallel Network for Efficient Anomaly DetectionabstractAnomaly detection in video streams is a challenging problem because of the scarcity of abnormal events and the difficulty of accurately annotating them. To alleviate these issues, unsupervised learning-based prediction methods have been previously applied. These approaches train the model with only normal events and predict a future frame from a sequence of preceding frames by use of encoder-decoder architectures so that they result in small prediction errors on normal events but large errors on abnormal events. The architecture, however, comes with the computational burden as some anomaly detection tasks require low computational cost without sacrificing performance. In this paper, Cross-Parallel Network (CPNet) for efficient anomaly detection is proposed here to minimize computations without performance drops. It consists of N smaller parallel U-Net, each of which is designed to handle a single input frame, to make the calculations significantly more efficient. Additionally, an inter-network shift module is incorporated to capture temporal relationships among sequential frames to enable more accurate future predictions. The quantitative results show that our model requires less computational cost than the baseline U-Net while delivering equivalent performance in anomaly detection. Youngsaeng Jin, Jonghwan Hong, David K. Han, Hanseok Ko |
AVSS | 4 |
| 2021 | Injecting Sparsity in Anomaly Detection for Efficient InferenceabstractAnomaly detection in the video is a challenging problem in computer vision tasks. Deep networks recently have been successfully applied and achieved competitive performance in anomaly detection. Modern deep networks employ many modules which extract important features. The anomaly detection approaches just developed network architecture and inserted additional networks to improve performance, however, these methods generally require a tremendous amount of computational load and training parameters. Because of limitations in the real world such as field equipment, mobile system, etc., reducing the number of trainable parameters and model capacity is an important issue in anomaly detection. Moreover, the method, which improves the performance of the anomaly detection algorithm, should be developed without additional trainable parameters. In this paper, we propose a sparsity injecting module which reinforces the feature representation of the existing model and presents the abnormality score function using sparsity. In experimental results, our sparsity injecting module improves the performance of state-of-the-art methods without additional trainable parameters. Bokyeung Lee, Hanseok Ko |
AVSS | 2 |
| 2021 | PETS2021: Through-foliage detection and tracking challenge and evaluationabstractThis paper presents the outcomes of the PETS2021 challenge held in conjunction with AVSS2021 and sponsored by the EU FOLDOUT project. The challenge comprises the publication of a novel video surveillance dataset on through-foliage detection, the defined challenges addressing person detection and tracking in fragmented occlusion scenarios, and quantitative and qualitative performance evaluation of challenge results submitted by six worldwide participants. The results show that while several detection and tracking methods achieve overall good results, through-foliage detection and tracking remains a challenging task for surveillance systems especially as it serves as the input to behaviour (threat) recognition. Jose Luis Patino, Jonathan N. Boyle, James M. Ferryman, Jonas Auer, Julian Pegoraro, Roman P. Pflugfelder, Mertcan Cokbas, Janusz Konrad, Prakash Ishwar, Giulia Slavic, Lucio Marcenaro, Yifan Jiang 0002, Youngsaeng Jin, Hanseok Ko, Guangliang Zhao, Guy Ben-Yosef, Jianwei Qiu |
AVSS | 14 |
| 2021 | Deep Degradation Prior for Real-World Super-Resolution
Kyungdeuk Ko, Bokyeung Lee, Jonghwan Hong, David K. Han, Hanseok Ko |
BMVC | 5 |
| 2021 | Adverse Weather Image Translation with Asymmetric and Uncertainty-aware GAN
Jeong-gi Kwak, Youngsaeng Jin, Yuanming Li, Dongsik Yoon, Hanseok Ko |
BMVC | 6 |
| 2021 | Adaptive Content Feature Enhancement GAN for Multimodal Selfie to Anime Translation
Yuanming Li, Jeong-gi Kwak, Dongsik Yoon, Youngsaeng Jin, David K. Han, Hanseok Ko |
BMVC | 6 |
| 2021 | Reference Guided Image Inpainting using Facial Attributes
Dongsik Yoon, Youngsaeng Jin, Jeong-gi Kwak, Yuanming Li, David K. Han, Hanseok Ko |
BMVC | 6 |
| 2021 | Few-Shot Learning for Ct Scan Based Covid-19 DiagnosisabstractCoronavirus disease 2019 (COVID-19) is a Public Health Emergency of International Concern infecting more than 40 million people across 188 countries and territories. Chest computed tomography (CT) imaging technique benefits from its high diagnostic accuracy and robustness, it has become an indispensable way for COVID-19 mass testing. Recently, deep learning approaches have become an effective tool for automatic screening of medical images, and it is also being considered for COVID-19 diagnosis. However, the high infection risk involved with COVID-19 leads to relative sparseness of collected labeled data limiting the performance of such methodologies. Moreover, accurately labeling CT images require expertise of radiologists making the process expensive and time-consuming. In order to tackle the above issues, we propose a supervised domain adaption based COVID-19 CT diagnostic method which can perform effectively when only a small samples of labeled CT scans are available. To compensate for the sparseness of labeled data, the proposed method utilizes a large amount of synthetic COVID-19 CT images and adjusts the networks from the source domain (synthetic data) to the target domain (real data) with a cross-domain training mechanism. Experimental results show that the proposed method achieves state-of-the- art performance on few-shot COVID-19 CT imaging based diagnostic tasks. Yifan Jiang 0002, Hanseok Ko, David K. Han |
ICASSP | 3 |
| 2021 | SpecMix : A Mixed Sample Data Augmentation Method for Training with Time-Frequency Domain FeaturesabstractA mixed sample data augmentation strategy is proposed to enhance the performance of models on audio scene classification, sound event classification, and speech enhancement tasks.While there have been several augmentation methods shown to be effective in improving image classification performance, their efficacy toward time-frequency domain features of audio is not assured.We propose a novel audio data augmentation approach named "Specmix" specifically designed for dealing with time-frequency domain features.The augmentation method consists of mixing two different data samples by applying time-frequency masks effective in preserving the spectral correlation of each audio sample.Our experiments on acoustic scene classification, sound event classification, and speech enhancement tasks show that the proposed Specmix improves the performance of various neural network architectures by a maximum of 2.7%. Gwantae Kim, David K. Han, Hanseok Ko |
Interspeech | 3 |
| 2021 | Memory-based Semantic Segmentation for Off-road Unstructured Natural EnvironmentsabstractWith the availability of many datasets tailored for autonomous driving in real-world urban scenes, semantic segmentation for urban driving scenes achieves significant progress. However, semantic segmentation for off-road, unstructured environments is not widely studied. Directly applying existing segmentation networks often results in performance degradation as they cannot overcome intrinsic problems in such environments, such as illumination changes. In this paper, a built-in memory module for semantic segmentation is proposed to overcome these problems. The memory module stores significant representations of training images as memory items. In addition to the encoder embedding like items together, the proposed memory module is specifically designed to cluster together instances of the same class even when there are significant variances in embedded features. Therefore, it makes segmentation networks better deal with unexpected illumination changes. A triplet loss is used in training to minimize redundancy in storing discriminative representations of the memory module. The proposed memory module is general so that it can be adopted in a variety of networks. We conduct experiments on the Robot Unstructured Ground Driving (RUGD) dataset and RELLIS dataset, which are collected from off-road, unstructured natural environments. Experimental results show that the proposed memory module improves the performance of existing segmentation networks and contributes to capturing unclear objects over various off-road, unstructured natural scenes with equivalent computational cost and network parameters. As the proposed method can be integrated into compact networks, it presents a viable approach for resource-limited small autonomous platforms. Youngsaeng Jin, David K. Han, Hanseok Ko |
IROS | 3 |
| 2021 | Side-Scan Sonar Image Synthesis Based on Generative Adversarial Network for Images in Multiple FrequenciesabstractThe side-scan sonar (SSS) is a critical sensor device used to explore underwater environments in the deep sea. Gathering SSS data, however, is an expensive and time-consuming task because it requires sensor towing and involves complicated field operations. Recently, deep learning has been making advances rapidly in the field of computer vision. Benefiting from this development, generative adversarial networks (GANs) have been demonstrated to produce realistic synthetic data of various types, including images and acoustics signals. In this letter, we propose a GAN-based semantic image synthesis model based on GAN that can generate high-quality SSS images at a low cost in less time. We evaluate the proposed model using both shallow and deep water SSS data sets that include a diverse range of imaging conditions. such as high and low sonar operating frequencies and different landscapes. The experimental results show that the proposed method can effectively generate synthesized SSS data characterized by the shape and style of real data, thereby demonstrating its promising potential for SSS data augmentation in diverse SSS relevant machine learning tasks. Yifan Jiang 0002, Bonhwa Ku, Wan-Jin Kim, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | Multifeature Fusion-Based Earthquake Event Classification Using Transfer LearningabstractThis letter proposes a multifeature fusion model using deep convolution neural networks and transfer learning approach for earthquake event classification. There are several feature representations for seismic analysis, such as the time domain, the frequency domain, and the time–frequency domain. To successfully classify various earthquake events, we propose a novel model that combines these features hierarchically. In addition, we apply a transfer learning to mitigate overfitting problem of deep learning model while achieving high classification performance. To evaluate our approach, we conduct experiments with the Korean peninsula earthquake database from 2016 to 2018 and a large earthquake database on the Circum-Pacific belt in 2019. The experimental results show that the proposed method outperforms over the compared state-of-the-art methods. Gwantae Kim, Bonhwa Ku, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Attention-Based Convolutional Neural Network for Earthquake Event ClassificationabstractThis letter presents a deep convolutional neural network (CNN) with attention module that improves the performance of the classification of various earthquake events. Addressing all possible earthquake events, including not only microearthquakes and artificial-earthquakes but also large-earthquakes, requires both suitable feature expression and a classifier that can effectively discriminate seismic waveforms under adverse conditions. To robustly classify earthquake events, a deep CNN with an attention module was proposed in raw seismic waveforms. Representative experimental results show that the proposed method provides an effective structure for earthquake events classification and, with the Korean peninsula earthquake database from 2016 to 2018, outperforms previous state-of-the-art methods. Bonhwa Ku, Gwantae Kim, Jae-Kwang Ahn, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2021 | Earthquake Event Classification Using Multitasking Deep LearningabstractThis letter proposes an attention-based convolutional neural network architecture for multitasking learning to accurately classify not only the presence of an earthquake but also the event type of the earthquake. In particular, to improve the performance in earthquake-type classification, we develop an attention-based feature aggregation framework embedded in multitask learning architecture. Representative experimental results show that the proposed method provides an effective structure for an earthquake detection and event classification with an earthquake database of the Korean peninsula and the Circum-Pacific belt. Bonhwa Ku, Jeongki Min, Jae-Kwang Ahn, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2021 | TrSeg: Transformer for semantic segmentation
Youngsaeng Jin, David K. Han, Hanseok Ko |
Pattern Recognit. Lett. | 3 |
| 2021 | COVID-19 CT Image Synthesis With a Conditional Generative Adversarial NetworkabstractCoronavirus disease 2019 (COVID-19) is an ongoing global pandemic that has spread rapidly since December 2019. Real-time reverse transcription polymerase chain reaction (rRT-PCR) and chest computed tomography (CT) imaging both play an important role in COVID-19 diagnosis. Chest CT imaging offers the benefits of quick reporting, a low cost, and high sensitivity for the detection of pulmonary infection. Recently, deep-learning-based computer vision methods have demonstrated great promise for use in medical imaging applications, including X-rays, magnetic resonance imaging, and CT imaging. However, training a deep-learning model requires large volumes of data, and medical staff faces a high risk when collecting COVID-19 CT data due to the high infectivity of the disease. Another issue is the lack of experts available for data labeling. In order to meet the data requirements for COVID-19 CT imaging, we propose a CT image synthesis approach based on a conditional generative adversarial network that can effectively generate high-quality and realistic COVID-19 CT images for use in deep-learning-based medical imaging tasks. Experimental results show that the proposed method outperforms other state-of-the-art image synthesis methods with the generated COVID-19 CT images and indicates promising for various machine learning applications including semantic segmentation and classification. Yifan Jiang 0002, Murray H. Loew, Hanseok Ko |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | CAFE-GAN: Arbitrary Face Attribute Editing with Complementary Attention Feature
Jeong-gi Kwak, David K. Han, Hanseok Ko |
ECCV (14) | 3 |
| 2020 | Convolutional Recurrent Neural Networks for Earthquake Epicentral Distance Estimation Using Single-Channel Seismic WaveformabstractThis paper proposes a deep learning method for epicentral distance estimation using a single-channel seismic waveform. The model is based on a convolutional recurrent neural network structure to extract spatial and temporal features. Since the proposed model needs only single-channel data, it can also perform the distance estimation even when some channels of the sensor are adversely disabled. To evaluate our approach, we conduct distance estimation experiments with the Korean peninsula earthquake database from 2016 to 2018, which include microearthquakes and distant earthquakes. The epicentral distance estimation by the proposed method show an absolute mean error of 0.50 km with 9.16km standard deviation of error distribution, which shows the best estimation result among the competing model structures. The promising result indicates that the proposed approach can be deployed for epicentral localization task as part of realizing a robust earthquake monitoring system. Gwantae Kim, Bonhwa Ku, Yuanming Li, Jeongki Min, Hanseok Ko |
IGARSS | 6 |
| 2020 | Seismic Signal Synthesis by Generative Adversarial Network with Gated Convolutional Neural Network StructureabstractDetecting earthquake events from seismic time series signal is a challenging task. Recently, detection methods based on machine learning have been developed to improve the accuracy and efficiency. However, accuracy of those methods rely on sufficient amount of high-quality training data. In many situations, the high-quality data is difficulty to obtain. We address and resolve this issue by using a Generative Adversarial Network (GAN) model for seismic signal synthesis. GAN already shows its powerful capability in generating high quality synthetic samples in multiple domains. In this paper, we propose a GAN model with gated CNN which can excellently capture sequential structure of seismic time series. We demonstrate its effectiveness via earthquake classification performance. The results show the synthetic data generated by our model indeed can improve the classification performance over the one trained with only real samples. Yuanming Li, Bonhwa Ku, Gwantae Kim, Jae-Kwang Ahn, Hanseok Ko |
IGARSS | 5 |
| 2020 | Dual Stage Learning Based Dynamic Time-Frequency Mask Generation for Audio Event Classification
Jaihyun Park, David K. Han, Hanseok Ko |
INTERSPEECH | 4 |
| 2020 | Spectro-Temporal Attention-Based Voice Activity DetectionabstractVoice Activity Detection (VAD) systems suffer from unexpected and non-stationary background noises at magnitudes sufficiently high to mask the speech signal.Although several methods of increasing the performance of VAD have been proposed, their approaches have yet to mitigate the influence of the background noise itself. This letter proposes an effective noise-robust VAD system approach. The proposed method uses spectral attention and temporal attention through applying a deep learning-based attention mechanism. The proposed method is demonstrated and compared with several other deep learning-based methods in terms of the area under the curve in experiments with either known or unknown noise-added, and real-world noisy data. The results show that the proposed method outperforms the other methods in all the scenarios considered, but moreover generalizes well in environments of unknown or unexpected noise. Younglo Lee, Jeongki Min, David K. Han, Hanseok Ko |
IEEE Signal Process. Lett. | 4 |
| 2020 | Amphibian Sounds Generating Network Based on Adversarial LearningabstractThis letter proposes a generative network based on adversarial learning for synthesizing short-time audio streams and investigates the effectiveness of data augmentation for amphibian call sounds classification. Based on Fourier analysis, the generator is designed by a multi-layer perceptron composed of frequency basis learning layers and an output layer, and a discriminator is constructed by a convolutional neural network. Additionally, regularization on weights is introduced to train the networks with practical data that includes some disturbances. Synthetic audio streams are evaluated by quantitative comparison using inception score, and classification results are compared for real versus synthetic data. In conclusion, the proposed generative network is shown to produce realistic sounds and therefore useful for data augmentation. Sangwook Park 0002, Mounya Elhilali, David K. Han, Hanseok Ko |
IEEE Signal Process. Lett. | 4 |
| 2020 | Fusion of Heterogeneous Adversarial Networks for Single Image DehazingabstractIn this paper, we propose a novel image dehazing method. Typical deep learning models for dehazing are trained on paired synthetic indoor dataset. Therefore, these models may be effective for indoor image dehazing but less so for outdoor images. We propose a heterogeneous Generative Adversarial Networks (GAN) based method composed of a cycle-consistent Generative Adversarial Networks (CycleGAN) for producing haze-clear images and a conditional Generative Adversarial Networks (cGAN) for preserving textural details. We introduce a novel loss function in the training of the fused network to minimize GAN generated artifacts, to recover fine details, and to preserve color components. These networks are fused via a convolutional neural network (CNN) to generate dehazed image. Extensive experiments demonstrate that the proposed method significantly outperforms the state-of-the-art methods on both synthetic and real-world hazy images. Jaihyun Park, David K. Han, Hanseok Ko |
IEEE Trans. Image Process. | 3 |
| 2019 | Self-Subtraction Network for End to End Noise Robust ClassificationabstractAcoustic event classification in surveillance applications typically employs deep learning-based end-to-end learning methods. In real environments, their performance degrades significantly due to noise. While various approaches have been proposed to overcome the noise problem, most of these methodologies rely on supervised learning-based feature representation. Supervised learning system, however, requires a pair of noise free and noisy audio streams. Acquisition of ground truth and noisy acoustic event data requires significant efforts to adequately capture the varieties of noise types for training. This paper proposes a novel supervised learning method for noise robust acoustic event classification in an end-to-end framework named Self Subtraction Network (SSN). SSN extracts noise features from an input audio spectrogram and removes them from the input using LSTMs and an auto-encoder. Our method applied to Urbansound8k dataset with 8 noise types at four different levels demonstrates improved performances compared to the state of the art methods. David K. Han, Hanseok Ko |
AVSS | 3 |
| 2019 | Relay dueling network for visual tracking with broad field-of-viewabstractA deep reinforcement‐learning‐based method is presented for visual object tracking tasks. The key objective is to generate a sequence of actions which can move or scale the bounding box in the previous frame to track the target in the current frame. Two intelligent agents are trained to accomplish the above task with a special dueling deep Q‐learning network (Dueling DQN), referred to as a relay dueling network. The proposed model is divided into two agents: the movement agent and the scaling agent. The former performs horizontal or vertical movements and the latter generates scaling actions to change the size of the bounding box. The model has multiple inputs that cover both the bounding box region and the enlarged search region to improve the agents’ perception of the surroundings. The proposed method has a broader field of vision than other similar trackers and its distribution of actions makes it easy to train and improve its tracking performance. The proposed network is tested on popular standard tracker benchmarks and its performance is compared with state‐of‐the‐art trackers. The proposed network is found to be competitive in tracking accuracy and execution effectiveness when compared to conventional methods. Yifan Jiang 0002, David K. Han, Hanseok Ko |
IET Comput. Vis. | 3 |
| 2019 | Nonhomogeneous Noise Removal From Side-Scan Sonar Images Using Structural SparsityabstractThe image quality of side-scan sonar (SSS) is determined by its operating frequency. SSS operating at a low frequency produces low-quality images due to high levels of noise. This noise is randomly generated from a number of different sources, including equipment noise and underwater environmental interference. In addition, to compensate for transmission loss in a received signal, the signal is amplified by time-varied gain correction, and consequently, SSS images contain nonhomogeneous noise, unlike natural images whose noise is assumed to be homogeneous. In this letter, a structural sparsity-based image denoising algorithm is proposed to remove nonhomogeneous noise from SSS images. The algorithm incorporates both local and nonlocal models in the structural features domain in order to guarantee sparsity and enhance nonlocal self-similarity. Using structural features also preserves fine-scale structures, leading to denoised images with natural seabed textures. The patch weights in the nonlocal model are corrected in consideration of the nonhomogeneity of the noise. Experimental results show that the proposed algorithm is qualitatively and quantitatively comparable to conventional algorithms. Youngsaeng Jin, Bonhwa Ku, Jaekyun Ahn, Seongil Kim, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2018 | Image fusion and influence function for performance improvement of ATM vandalism action recognitionabstractRising rate of vandalism against Automatic Teller Machines (ATMs) is a serious issue within banking industries, prompting needs of a technology to autonomously recognize such events. A vision based fusion method proposed here for classifying these incidents is rooted on visually recognizing heavy or sharp objects potentially used for detecting vandalism actions inferred from optical flow. The recognition performance has been improved chiefly by a novel employment of influence functions in selecting data points of each class useful in learning. We show that the tool recognition performance can be improved when the training data is selected from the ImageNet data set as guided by the influence function. Jeongseop Yun, Junyeop Lee, Seongkyu Mun, Chul Jin Cho, David K. Han, Hanseok Ko |
AVSS | 6 |
| 2018 | Hierarchical spatial object detection for ATM vandalism surveillanceabstractIn this paper, a multi-modal classification is proposed for recognizing vandalism against Automatic Teller Machines (ATMs). The visual and textual information base model is developed here to identify external threats on ATMs. The model discriminates threatening behaviors from those that are benign in the image. It provides a level of confidence in the threat recognition by visual object classification coupled with word vector distance measure. To achieve our goal, real-time object detection based on a Region Convolutional Neural Network (R-CNN) first detects objects in the scene and word embedding technique allows to measure distance between the detected object label with predefined tools assumed to be used for vandalizing ATMs. Similarity measure from word embedding not only determines whether the scene may lead to any nefarious activities, but also would provide the level of confidence in occurrence of such incidents. From the experimental evaluation, it is shown that the method is effective and delivers a quantitative measure on decisions it makes. Junyeop Lee, Chul Jin Cho, David K. Han, Hanseok Ko |
AVSS | 4 |
| 2018 | Precise Regression for Bounding Box Correction for Improved Tracking Based on Deep Reinforcement LearningabstractIn this paper, we propose a precise regression approach for correcting imprecise bounding box using deep reinforcement learning. Object tracking task essentially builds trajectory of a moving object based on detection and tracking algorithms and its current state is indicated by having the object encapsulated with a bounding box corresponding to its position and size. However due to the imperfect detection and tracking algorithms operating in complex scene, it is difficult to obtain the precise bounding box as errors frequently occur producing oversized, partial, and false bounding box, respectively. To correct the error, we train an intelligent agent that move the bounding box to the right position and scale it to its correct size matching to that of the true target. The agent is trained by deep Q-Iearning and evaluated on several state-of-the-art multiple object tracking approaches. The experimental results demonstrate that our proposed framework can effectively eliminate the object tracking bounding box error and its robustness is verified by realizing improved tracking performance in complex scene. Yifan Jiang 0002, Hyunhak Shin, Hanseok Ko |
ICASSP | 3 |
| 2018 | Man-Made Radio Frequency Interference Suppression for Compact HF Surface Wave RadarabstractHigh-frequency surface wave radar (HFSWR) suffers from a man-made interference because its amplitude is high enough to mask the Bragg scattering signal. Although several methods have been proposed for resolving this problem, they are inapplicable to compact HFSWR due to their antenna structures. This letter proposes an effective method of suppressing man-made radio frequency interference for compact HFSWR. The proposed method is composed of man-made interference detection and suppression by using regression based on probabilistic signal model. The proposed method is demonstrated in comparison with conventional methods in terms of root-mean-square error in experiments using synthetic and real data. The results show that the proposed method outperforms other methods in both simulated and practical situations. Younglo Lee, Sangwook Park 0002, Chul Jin Cho, Bonhwa Ku, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2017 | Online pedestrian tracking with multi-stage re-identificationabstractNowadays the task of tracking pedestrians is often addressed within a tracking-by-detection framework, which in most cases entails that the position of each target has been detected before tracking begins. However in some cases, a pedestrian who is being tracked may be obscured by other targets or obstacles, and during this period they may change their trajectory or speed (track drift), and sometimes such a target may leave the FOV (Field of View) [10] but appear again later. These temporary disappearances and absence of detections disrupt the work of the detectors to such an extent that there is a significant decline in performance. In this paper, we propose a novel approach to pedestrian tracking based on multi-stage re-identification. To deal with the problems discussed above, the proposed framework is comprised of a two-stage re-identification algorithm dealing with cases of track drift and re-entry into the FOV individually, in order to match the identities of lost and reappeared targets through a comparison of the affinities between their appearance, size and position, and also to update the status of re-identified targets through this assessment. The experimental results demonstrate that this framework can effectively handle complex temporary lost and re-entry situations with robustness, and that its performance is state of the art. Yifan Jiang 0002, Hyunhak Shin, Jaeyong Ju, Hanseok Ko |
AVSS | 4 |
| 2017 | Deep Neural Network based learning and transferring mid-level audio features for acoustic scene classificationabstractDeep Neural Network (DNN) based transfer learning has been shown to be effective in Visual Object Classification (VOC) for complementing the deficit of target domain training samples by adapting classifiers that have been pre-trained for other large-scaled DataBase (DB). Although there exists an abundance of acoustic data, it can also be said that datasets of specific acoustic scenes are sparse for training Acoustic Scene Classification (ASC) models. By exploiting VOC DNN's ability of learning beyond its pre-trained environments, this paper proposes DNN based transfer learning for ASC. Effectiveness of the proposed method is demonstrated on the database of IEEE DCASE Challenge 2016 Task 1 and home surveillance environment via representative experiments. Its improved performance is verified by comparing it to prominent conventional methods. Seongkyu Mun, Suwon Shon, Wooil Kim, David K. Han, Hanseok Ko |
ICASSP | 5 |
| 2017 | Subspace projection cepstral coefficients for noise robust acoustic event recognitionabstractIn this paper, a novel feature for noise robust sound event recognition is proposed. The proposed feature is obtained by a two-step procedure. First, a subspace bank is established via target event analysis in complex vector space. Then, by projecting observation vectors onto the subspace bank, noise effect can be reduced while generating discriminant characters originated from differing event subspaces. To demonstrate robustness of the proposed feature, experiments with several classifiers were conducted with varying SNR cases under four noisy environments. According to the experimental results, the proposed method has shown superior performance over prominent conventional methods. Sangwook Park 0002, Younglo Lee, David K. Han, Hanseok Ko |
ICASSP | 4 |
| 2017 | Recursive Whitening Transformation for Speaker Recognition on Language Mismatched ConditionabstractRecently in speaker recognition, performance degradation due to the channel domain mismatched condition has been actively addressed. However, the mismatches arising from language is yet to be sufficiently addressed. This paper proposes an approach which employs recursive whitening transformation to mitigate the language mismatched condition. The proposed method is based on the multiple whitening transformation, which is intended to remove un-whitened residual components in the dataset associated with i-vector length normalization. The experiments were conducted on the Speaker Recognition Evaluation 2016 trials of which the task is non-English speaker recognition using development dataset consist of both a large scale out-of-domain (English) dataset and an extremely low-quantity in-domain (non-English) dataset. For performance comparison, we develop a state-of- the-art system using deep neural network and bottleneck feature, which is based on a phonetically aware model. From the experimental results, along with other prior studies, effectiveness of the proposed method on language mismatched condition is validated. Suwon Shon, Seongkyu Mun, Hanseok Ko |
INTERSPEECH | 3 |
| 2017 | Autoencoder Based Domain Adaptation for Speaker Recognition Under Insufficient Channel InformationabstractIn real-life conditions, mismatch between development and test domain degrades speaker recognition performance. To solve the issue, many researchers explored domain adaptation approaches using matched in-domain dataset. However, adaptation would be not effective if the dataset is insufficient to estimate channel variability of the domain. In this paper, we explore the problem of performance degradation under such a situation of insufficient channel information. In order to exploit limited in-domain dataset effectively, we propose an unsupervised domain adaptation approach using Autoencoder based Domain Adaptation (AEDA). The proposed approach combines an autoencoder with a denoising autoencoder to adapt resource-rich development dataset to test domain. The proposed technique is evaluated on the Domain Adaptation Challenge 13 experimental protocols that is widely used in speaker recognition for domain mismatched condition. The results show significant improvements over baselines and results from other prior studies. Suwon Shon, Seongkyu Mun, Wooil Kim, Hanseok Ko |
INTERSPEECH | 4 |
| 2017 | Online multi-person tracking with two-stage data association and online appearance model learningabstractThis study addresses the automatic multi‐person tracking problem in complex scenes from a single, static, uncalibrated camera. In contrast with offline tracking approaches, a novel online multi‐person tracking method is proposed based on a sequential tracking‐by‐detection framework, which can be applied to real‐time applications. A two‐stage data association is first developed to handle the drifting targets stemming from occlusions and people's abrupt motion changes. Subsequently, a novel online appearance learning is developed by using the incremental/decremental support vector machine with an adaptive training sample collection strategy to ensure reliable data association and rapid learning. Experimental results show the effectiveness and robustness of the proposed method while demonstrating its compatibility with real‐time applications. Jaeyong Ju, Daehun Kim, Bonhwa Ku, David K. Han, Hanseok Ko |
IET Comput. Vis. | 5 |
| 2017 | Compact HF Surface Wave Radar Data Generating Simulator for Ship Detection and TrackingabstractToward a maritime surveillance objective, many ship detection and tracking algorithms have been investigated but are faced with poor performance in practical ocean environments. Compact high-frequency (HF) radar has also faced critical issues due to its long coherent processing interval and varying response from its orthogonal antenna structure. Hence, a simulator based on compact HF radar is proposed in this letter to provide a guideline for effective assessment of ship detection and tracking algorithms while considering these practical issues. To validate the proposed simulator, the simulator generated data has been compared with real data obtained by the compact HF radar sites. Sangwook Park 0002, Chul Jin Cho, Bonhwa Ku, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2017 | A feature descriptor based on the local patch clustering distribution for illumination-robust image matching
Han Wang 0018, Sangmin Yoon, David K. Han, Hanseok Ko |
Pattern Recognit. Lett. | 4 |
| 2017 | Continuous hand gesture recognition based on trajectory shape information
Cheoljong Yang, David K. Han, Hanseok Ko |
Pattern Recognit. Lett. | 3 |
| 2016 | Single object tracking based on active and passive detection information in distributed heterogeneous sensor networkabstractIn this paper, a single object tracking method based on fusion of detection information collected from a distributed heterogeneous sensor network is proposed. The considered sensor network is composed of one active type source and multiple receivers. It is assumed that the heterogeneous network is capable of acquiring both passive and active information simultaneously. By means of fusion of the acquired heterogeneous data, the proposed method estimates the candidate region of target location. Then, position of the object is estimated by Maximum Likelihood Estimation. In the experimental results, the performance of the proposed method is demonstrated in terms of deployment strategy of the heterogeneous sensor network. Hyunhak Shin, Chul Jin Cho, Hanseok Ko |
AVSS | 3 |
| 2016 | Nighttime image dehazing with local atmospheric light and weighted entropyabstractIn this paper, we propose a novel framework for nighttime image dehazing based on a nighttime haze model which accounts for varying light sources and their glow. First, glow effects are decomposed using relative smoothness. Atmospheric light is then estimated by combining global and local atmospheric lights using a local atmospheric selection map. The transmission is estimated by maximizing an objective function designed with weighted entropy. Finally, haze is removed using two estimated parameters which are atmospheric light and transmission. Experimental results validate the proposed method can achieve haze-free results while alleviating the glow effect. Dubok Park, David K. Han, Hanseok Ko |
ICIP | 3 |
| 2016 | Deep Neural Network Bottleneck Features for Acoustic Event Recognition
Seongkyu Mun, Suwon Shon, Wooil Kim, Hanseok Ko |
INTERSPEECH | 4 |
| 2015 | Maximum likelihood Linear Dimension Reduction of heteroscedastic feature for robust Speaker RecognitionabstractThis paper analyzes heteroscedasticity in i-vector for robust forensics and surveillance speaker recognition system. Linear Discriminant Analysis (LDA), a widely-used linear dimension reduction technique, assumes that classes are homoscedastic within a same covariance. In this paper it is assumed that general speech utterances contain both homoscedastic and heteroscedastic elements. We show the validity of this assumption by employing several analyses and also demonstrate that dimension reduction using principal components is feasible. To effectively handle the presence of heteroscedastic and homoscedastic elements, we propose a fusion approach of applying both LDA and Heteroscedastic-LDA (HLDA). The experiments are conducted to show its effectiveness and compare to other methods using the telephone database of National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE) 2010 extended. Suwon Shon, Seongkyu Mun, David K. Han, Hanseok Ko |
AVSS | 4 |
| 2015 | Acoustic event recognition using dominant spectral basis vectors
Woohyun Choi, Sangwook Park 0002, David K. Han, Hanseok Ko |
INTERSPEECH | 4 |
| 2015 | Video-Based Dynamic Stagger Measurement of Railway Overhead Power Lines Using Rotation-Invariant Feature MatchingabstractIn this paper we propose an effective method of assessing the reliability of railway overhead power lines by measuring the dynamic stagger of contact wires based on a video monitoring technique. Previously developed video monitoring methods may produce severe errors when applied to tilting trains due to changes in position and orientation of the pantograph. In particular, we propose to employ feature-based image matching techniques that are invariant to rotation and robust to changes in camera viewpoint. A pantograph tilting model is first developed from the video data acquired from an actual train based on the motion dynamics of the stagger behavior on moving train platform. We then evaluate the proposed method by comparing it with the conventional template matching in terms of tracking error. The experimental results confirm that the proposed method shows superior performance in all train traveling sequences, particularly over the pantograph tilting train motion segment. Chul Jin Cho, Hanseok Ko |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2014 | Single image dehazing with image entropy and information fidelityabstractIn this paper, we propose a new single image dehazing approach based on information fidelity and image entropy. The global atmospheric light is estimated by quadtree subdivision using transformed hazy images. Then, transmission is estimated by an objective function which is comprised of information fidelity and image entropy at non-overlapped sub-block regions. This is further refined by a Weighted Least Squares (WLS) optimization procedure to alleviate block artifacts. We compared performance of the proposed method with conventional methods to validate its effectiveness in an experiment. Dubok Park, Hyungjo Park, David K. Han, Hanseok Ko |
ICIP | 4 |
| 2014 | Single image haze removal using novel estimation of atmospheric light and transmissionabstractThis paper presents a new single image dehaze approach that uses a novel estimation of the atmospheric light and media transmission. Conventional dehaze methods often result in degraded images with low contrast and/or oversaturation of color in some regions. In order to mitigate these problems we use local atmospheric light and estimate the media transmission for each local region by using an objective function represented by modified saturation evaluation metric and intensity difference. Experimental results on a variety of outdoor haze images show that the proposed method achieves excellent restoration in terms of contrast, color fidelity and image visibility. Hyungjo Park, Dubok Park, David K. Han, Hanseok Ko |
ICIP | 4 |
| 2014 | Rule-based trajectory segmentation for modeling hand motion trajectory
Jounghoon Beh, David K. Han, Hanseok Ko |
Pattern Recognit. | 3 |
| 2014 | Hidden Markov Model on a unit hypersphere space for gesture trajectory recognition
Jounghoon Beh, David K. Han, Ramani Durasiwami, Hanseok Ko |
Pattern Recognit. Lett. | 4 |
| 2013 | Abnormal acoustic event localization based on selective frequency bin in high noise environment for audio surveillanceabstractIn this paper, a method for source localization for surveillance system is presented. In particular, we propose an algorithm for abnormal acoustic event localization based on a novel approach of relevant frequency bin selections by statistical analyses. By means of selective frequency bin, it becomes possible to localize the event more accurately in high noise environment with low computational complexity. The effectiveness is verified through the experimental results in varied noise environments with different levels of Signal to Noise Ratio (SNR). Suwon Shon, David K. Han, Hanseok Ko |
AVSS | 3 |
| 2013 | Robust sound source localization using a Wiener filterabstractRecently, human-robot interaction, or human-robot communication, has been extended to practical application in real environments. One of the important factors in human-robot interaction is sound source localization. Moreover, communication is more complex in sound-corrupted environments than in controlled environments. In the present study, the a priori SNR Wiener Scalart algorithm for noise reduction is utilized and integrated in a sound source localization system. The system is evaluated in the sound-corrupted context of a vacuum cleaning robot performing cleaning operations. Hyungi Cho, Jongsuk Choi, Hanseok Ko |
ETFA | 3 |
| 2013 | Single image haze removal with WLS-based edge-preserving smoothing filterabstractImages captured under hazy conditions have low contrast and poor color. This is primarily due to air-light which degrades image quality according to the transmission map. The approach to enhance these hazy images we introduce here is based on the `Dark-Channel Prior' method with image refinement by the `Weighted Least Square' based edge-preserving smoothing. Local contrast is further enhanced by multi-scale tone manipulation. The proposed method improves the contrast, color and detail for the entire image domain effectively. In the experiment, we compare the proposed method with conventional methods to validate performance. Dubok Park, David K. Han, Hanseok Ko |
ICASSP | 3 |
| 2013 | Multimodal image fusion via sparse representation with local patch dictionariesabstractSparse representation is a promising technique for the field of image processing and pattern recognition. It generally exploits over-complete dictionaries which is fixed and known in advance, or learned using training algorithm such as K-SVD. In this paper, we propose a new multimodal image fusion approach based on the sparsity model with local patch dictionaries generated directly from input images. For every location in the image, dictionary is simply constructed with neighboring patches. Experimental results show that the proposed method is efficient and competitive with some existing image fusion methods. David K. Han, Hanseok Ko |
ICIP | 3 |
| 2012 | Combining Infrared and Visible Images Using Novel Transform and Statistical InformationabstractThis paper proposes a novel combining method of infrared (IR) and visible images based on a Discrete Wavelet Frame (DWF) approach. In contrast to existing methods, IR image is transformed first using statistical information of the visible image to emphasize relevant information. In a multi-scale domain, we then assign appropriate weights to each pixel of sub-band approximation images through pixel level weighted average for emphasizing relevant information of the IR image while keeping texture information of the visible image. Representative experiments show that the proposed method outperforms exiting methods in image quality. Bonhwa Ku, David K. Han, Hanseok Ko |
AVSS | 4 |
| 2012 | Selective Background Adaptation Based Abnormal Acoustic Event Recognition for Audio SurveillanceabstractIn this paper, a method for abnormal acoustic event recognition in an audio surveillance system is presented. We propose a recognition scheme based on a hierarchical structure using a feature combination of Mel-Frequency Cepstral Coefficient (MFCC), timbre, and spectral statistics. A selective background adaptation is proposed for robust abnormal acoustic event recognition in real-world situations. For training, we use a database containing 9 abnormal events (scream, glass breaking, and etc.) and 6 background noise types collected under various surveillance situations. Gaussian Mixture Model (GMM) is considered for classifying the representative abnormal acoustic events and for selecting the background noise for adaptation. Effectiveness of the proposed method is demonstrated via representative experimental results. Woohyun Choi, Jinsang Rho, David K. Han, Hanseok Ko |
AVSS | 4 |
| 2012 | Crowd Density Estimation Using Multi-class AdaboostabstractIn this paper, we propose a crowd density estimation algorithm based on multi-class Adaboost using spectral texture features. Conventional methods based on self-organizing maps have shown unsatisfactory performance in practical scenarios, and in particular, they have exhibited abrupt degradation in performance under special conditions of crowd densities. In order to address these problems, we have developed a new training strategy by incorporating multi-class Adaboost with spectral texture features that represent a global texture pattern. According to the representative experimental results, the proposed method shows an average improvement of about 30% in the correct recognition rate, as compared to existing conventional methods. Daehun Kim, Younghyun Lee, Bonhwa Ku, Hanseok Ko |
AVSS | 4 |
| 2011 | Hierarchical approach for abnormal acoustic event classification in an elevatorabstractIn this paper, we propose a hierarchical method to detect and classify abnormal acoustic events occurring in an elevator environment. The Gaussian Mixture Model (GMM) based event classifier essentially employs two types of acoustic features; Mel Frequency Cepstral Coefficient (MFCC) and Timbre. We explore the effectiveness of various combinations of the two features in terms of classification performance. In addition, we design a hierarchical approach for realizing acoustic event classification and compare it with a single-level approach. It can be verified from an experiment, that the classification performance is improved when the proposed hierarchical approach is applied. In particular, for detection of abnormal situations, we employ a maximum likelihood estimation approach for acoustic event recognition at the 1ststep, and then on the 2ndstep we determine the abnormal contexts by using the ratio of abnormal events to cumulative events during a certain period. For performance evaluation, we employ a database collected in an actual elevator under several scenarios. By experimental results, our proposed method demonstrates 91% correct detection rate and 2.5% error detection rate for abnormal context. Kwangyoun Kim, Hanseok Ko |
AVSS | 2 |
| 2011 | Resolution enhancement of ROI from surveillance video using Bernstein interpolationabstractIn visual surveillance system, a small region-of-interest (ROI) is in most cases set for the target to detect, track, and recognize. In this paper, we propose novel image restoration algorithm using stochastic data regularization (i.e. Bernstein interpolation) in real world surveillance video. We firstly demonstrate the capability of Bernstein function for image enhancement technique such as denoising or deblurring. In addition, a promising approach for Super Resolution algorithm via Bernstein interpolation is also proposed. Representative experimental results prove the effectiveness of the proposed method for ROI from synthetic and real image sequences. Hanseok Ko |
AVSS | 2 |
| 2011 | Robust background subtraction using data fusion for real elevator sceneabstractThis paper proposes a background subtraction technique robust in elevator environments. Sudden local illumination changes arise frequently in an elevator environment due to opening and closing of the elevator door as well as the inner walls of elevator being made of reflective materials. We present a novel method sequentially fusing a Gaussian mixture model for background subtraction, motion information and a spatial likelihood model based on textured features. Experimental results on real video data demonstrate effectiveness of the proposed approach. Taeyup Song, David K. Han, Hanseok Ko |
AVSS | 3 |
| 2011 | Adaptive height-modified histogram equalization and chroma correction in YCbCr color space for fast backlight image compensation
Bonghyup Kang, Changwon Jeon, David K. Han, Hanseok Ko |
Image Vis. Comput. | 4 |
| 2010 | Robust Dynamic Super Resolution under Inaccurate Motion EstimationabstractIn image reconstruction, dynamic super resolution image reconstruction algorithms have been investigated to enhance video frames sequentially, where explicit motion estimation is considered as a major factor in the performance. This paper proposes a novel measurement validation method to attain robust image reconstruction results under inaccurate motion estimation. In addition, we present an effective scene change detection method dedicated to the proposed super resolution technique for minimizing erroneous results when abrupt scene changes occur in the video frames. Representative experimental results show excellent performance of the proposed algorithm in terms of the reconstruction quality and processing speed. Bonhwa Ku, Daesung Chung, Hyunhak Shin, Bonghyup Kang, David K. Han, Hanseok Ko |
AVSS | 7 |
| 2010 | License Plate Detection Using Local Structure PatternsabstractWe address the problem of license plate detection in video surveillance systems. The Adaboost based approach, known for relative ease of implementation, makes use of discriminative features such as edges or Haar-like features. In this paper, we propose a novel detection algorithm based on local structure patterns for license plate detection. The proposed algorithm includes post-processing methods to reduce false positive rate using positional and color information of license plates. Experimental results demonstrate effectiveness of the proposed method compared to both the edge and Haar-like feature based methods. Younghyun Lee, Taeyup Song, Bonhwa Ku, Seoungseon Jeon, David K. Han, Hanseok Ko |
AVSS | 6 |
| 2010 | Reinforced blocking matrix with cross channel projection for speech enhancement
Jongsung Yoon, Hanseok Ko |
INTERSPEECH | 4 |
| 2010 | Sound source separation by using matched beamforming and time-frequency maskingabstractThis paper proposes a two-stage algorithm to separate two sound sources by using matched beamforming and time-frequency masking techniques. At first, beamforming was used to separate the sound mixtures back to the original sources while preserving the original contents to the maximum extent. The residual interference was then suppressed by the time-frequency masking technique. A sequential least squares method was used in developing a matched beamformer to estimate the relative transfer function (RTF). From experimental results, it has been shown that the proposed method exhibits improved performance in sound source separation compared to conventional methods. Signal enhanced factor (SEF) was improved by an average of 8.39 dB over the baseline. Jounghoon Beh, Taekjin Lee, David K. Han, Hanseok Ko |
IROS | 4 |
| 2009 | Extension of two-channel transfer function based generalized sidelobe canceller for dealing with both background and point-source noise
Kihyeon Kim, Robert H. Baran, Hanseok Ko |
Speech Commun. | 3 |
| 2008 | More powerful discriminants for classifying phylogenetic signals in dinucleotide frequenciesabstractMicrobial DNA fragments are classified according to species using compositional features and “genomic signatures” the oldest of which is the dinucleotide relative abundance profile defined by Karlin et al. More informative features, including higher order signatures, have demonstrated greater species-specificity in comparison to the baseline established by the dinucleotide signature using “delta-distance” to assess dissimilarity; but lack of standard methods has precluded rigorous comparison. We describe a new method for classifier evaluation that reduces any number of pair-wise inter-genomic comparisons to a single performance measure. To illustrate the method, we compare delta-distance to quadratic and linear discriminants prescribed by elementary pattern recognition theory, and find that the quadratic form is significantly more powerful. Robert H. Baran, Changwon Jeon, David K. Han, Hanseok Ko |
ICASSP | 4 |
| 2008 | Bregman Divergences and the Self Organising Map
Eunsong Jang, Colin Fyfe, Hanseok Ko |
IDEAL | 3 |
| 2008 | Feature Locations in Images
Hokun Kim, Colin Fyfe, Hanseok Ko |
IDEAL | 3 |
| 2008 | Combining acoustic echo cancellation and adaptive beamforming for achieving robust speech interface in mobile robotabstractThis paper proposes a combined scheme in which adaptive beamforming and acoustic echo canceller are integrated in order to achieve a robust full-duplex speech interface between user and mobile robot. In particular, this paper addresses the situation in which an echo signal and the userpsilas voice are propagated from the same direction toward a linear array of microphones. We propose a cascading scheme that uses an acoustic echo canceller followed by an adaptive beamformer. By using the system output to adapt echo canceller, and by controlling the adaptive mode of each adaptive filter in the beamformer and in the echo canceller, we can recover the original speech, which would otherwise have remained contaminated by both noise and echoes. In addition, a double-talk detector is proposed so that effective acoustic echo cancellation may be attained without generating error divergence when the userpsilas voice and the echo are present simultaneously. Representative experimental results with real data demonstrate the validity of the proposed scheme. Jounghoon Beh, Taekjin Lee, Sungjoo Ahn, Hanseok Ko |
IROS | 6 |
| 2008 | Topological Mappings of Video and Audio DataabstractWe review a new form of self-organizing map which is based on a nonlinear projection of latent points into data space, identical to that performed in the Generative Topographic Mapping (GTM).(1) But whereas the GTM is an extension of a mixture of experts, this model is an extension of a product of experts.(2) We show visualisation and clustering results on a data set composed of video data of lips uttering 5 Korean vowels. Finally we note that we may dispense with the probabilistic underpinnings of the product of experts and derive the same algorithm as a minimisation of mean squared error between the prototypes and the data. This leads us to suggest a new algorithm which incorporates local and global information in the clustering. Both ot the new algorithms achieve better results than the standard Self-Organizing Map. Colin Fyfe, Wesam Barbakh, Wei Chuang Ooi, Hanseok Ko |
Int. J. Neural Syst. | 4 |
| 2008 | Gradient-based local affine invariant feature extraction for mobile robot localization in indoor environments
Jihyo Lee, Hanseok Ko |
Pattern Recognit. Lett. | 2 |
| 2008 | Real-Time Continuous Phoneme Recognition System Using Class-Dependent Tied-Mixture HMM With HBT Structure for Speech-Driven Lip-SyncabstractThis work describes a real-time lip-sync method using which an avatar's lip shape is synchronized with the corresponding speech signal. Phoneme recognition is generally regarded as an important task in the operation of a real-time lip-sync system. In this work, the use of the Head-Body-Tail (HBT) model is proposed for the purpose of more efficiently recognizing phonemes which are variously uttered due to co-articulation effects. The HBT model effectively deals with the transition parts of context-dependent models for small-sized vocabulary tasks. These models provide better recognition performance than general context-dependent or context-independent models for the task of digit or vowel recognition. Moreover, each phoneme is categorized into one among four classes and the class-dependent codebook is generated to further improve the performance. Additionally, for the clear representation of the context dependency information in the transient parts, some Gaussians are excluded from class-dependent codebook. The proposed method leads to a lip-sync system that performs at a level that is similar to previous designs based on HBT and continuous hidden Markov models (CHMMs). However, our method reduces the number of model parameters by one-third and enables real-time operation. Hanseok Ko |
IEEE Trans. Multim. | 2 |
| 2007 | Combination of self-organization map and kernel mutual subspace method for video surveillanceabstractThis paper addresses the video surveillance issue of automatically identifying moving vehicles and people from continuous observation of image sequences. With a single far-field surveillance camera, moving objects are first segmented by simple background subtraction. To reduce the redundancy and select the representative prototypes from input video streams, the Self-organizing Feature Map (SOM) is applied for both training and testing sequences. The recognition scheme is designed based on the recently proposed Kernel Mutual Subspace (KMS) model. As an alternative to some probability-based models, KMS does not make assumptions about the data sampling processing and offers an efficient and robust classifier. Experiments demonstrated a highly accurate recognition result, showing the model’s applicability in real-world surveillance system. Junbum Park, Hanseok Ko |
AVSS | 3 |
| 2007 | Visualising and Clustering Video Data
Colin Fyfe, Wei Chuang Ooi, Hanseok Ko |
IDEAL | 3 |
| 2007 | Enabling directional human-robot speech interface via adaptive beamforming and spatial noise reductionabstractThis paper introduces a home robot application of multi-channel based spatial noise reduction for creating human-robot speech interfaces. A microphone array is employed first to create a speech-only directional conduit, which is realized through adaptive beamforming. Through the directional conduit, the intended speech signal from the desired direction is processed for detection and recognition, while unintended speech-like-sources or undesirable noise from other angles is suppressed. If speech signal is absent among the incoming signals through the conduit, further attenuation of undesirable signals is achieved by using a spatial noise reduction filter. Experimental validation of the technique was conducted using a computer simulation and also an online Samsung AnyBot test. Although the environments exhibited highly non-stationary noise, the method achieved an average speech recognition rate of 87.4% in the case of the computer simulation and 81.6% for the online Samsung AnyBot test. From the cases tested so far, the proposed implementation seems to be effective for practical robot applications in highly non-stationary noise environment. Jounghoon Beh, Taekjin Lee, Sungjoo Ahn, David K. Han, Hanseok Ko |
IROS | 6 |
| 2006 | A new state-dependent phonetic tied-mixture model with head-body-tail structured HMM for real-time continuous phoneme recognition system
Hanseok Ko |
INTERSPEECH | 2 |
| 2006 | Svm-based Phoneme Classification and Lip Shape Refinement in Real-time Lip-synch SystemabstractIn this paper, we present a real time lip-synch system that activates 2-D avatar's lip motion in synch with incoming speech utterance. To achieve the real time operation of the system, the processing time was minimized by "merge and split" procedures resulting in coarse-to-fine phoneme classification. At each stage of phoneme classification, the support vector machine (SVM) method was applied to reduce the computational load while maintaining the desired accuracy. The coarse-to-fine phoneme classification, is accomplished via two_stages of feature extraction: in the first stage, each speech frame is acoustically analyzed for three classes of lip opening using Mel Frequency Cepstral Coefficients (MFCC) as a feature; in the second stage, each frame is further refined for detailed lip shape using formant information. The method was implemented in 2-D lip animation and it was demonstrated that the system was effective in accomplishing real-time lip-synch. This approach was tested on a PC using the Microsoft Visual Studio with an Intel Pentium IV 1.4 Giga Hz CPU and 384 MB RAM. It was observed that the methods of phoneme merging and SVM achieved about twice the speed in recognition than the method employing the Hidden Markov Model (HMM). A typical latency time per a single frame observed using the proposed method was in the order of 18.22 milliseconds while an HMM method under identical conditions resulted about 30.67 milliseconds. Hanseok Ko, David K. Han |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2006 | Prediction Based Occluded Multitarget Tracking Using Spatio-temporal AttentionabstractThis paper proposes the prediction based occluded multitarget tracking method using spatio-temporal attention mechanism. To cope with occlusion between targets, the proposed method provides an efficient method for more complex analysis by combining object association with partial probability model in spatially attentive window and occlusion activity detection in predicted temporal location. While multiple objects are moving or occluding between them in areas of visual field, a simultaneous tracking of multiple objects tends to fail. This is due to the fact that incompletely estimated feature vectors such as location, color, velocity, and acceleration of a target can provide only ambiguous and missing information. Thus, the spatially and temporally considered mechanism is proposed to track each target before, during, and after occlusion. Robustness of the proposed method is demonstrated with representative simulations. Heungkyu Lee, June Kim, Hanseok Ko |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2006 | Competing models-based text-prompted speaker independent verification algorithm
Heungkyu Lee, Hanseok Ko |
Speech Commun. | 2 |
| 2006 | Achieving a reliable compact acoustic model for embedded speech recognition system with high confusion frequency model handling
Hanseok Ko |
Speech Commun. | 2 |
| 2005 | Speaker Adaptive Confidence Scoring Using Bayesian CombiningabstractBayesian combining of confidence measures is proposed for speech recognition. Bayesian combining is achieved by the estimation of joint pdf of confidence feature vector in correct and incorrect hypothesis classes. If the joint pdf in the two classes are correctly estimated, this method guarantees an optimal combining in the minimum Bayes risk sense. Investigating the distribution of confidence features, we found out that the pdf are well estimated by the Gaussian mixture model with full covariance matrix in combining small number of features. In addition, the adaptation of a confidence score by adapting the joint pdf is presented. The proposed methods reduced the classification error rate by 17% from the conventional single feature based confidence scoring method in an isolated word out-of-vocabulary rejection test. Hanseok Ko |
ICASSP (1) | 2 |
| 2005 | Environment-independent mask estimation for missing-feature reconstructionabstractIn this paper, we propose an effective mask-estimation method for missing-feature reconstruction in order to achieve robust speech recognition in unknown noise environments. In previous work, it was found that training a model for mask estimation on speech corrupted by white noise did not provide environment-independent recognition accuracy. In this paper we describe a training method based on bands of colored noise that is more effective in reflecting spectral variations across neighboring frames and subbands. We also achieved further improvement in recognition accuracy by reconsidering frames that appeared to be unvoiced in the initial pitch analysis. Performance is evaluated using the Aurora 2.0 database in the presence of various types of noise maskers. Experimental results indicate that the proposed methods are effective in estimating masks for missing-feature reconstruction while remaining more independent of the noise conditions. 1. Wooil Kim, Richard M. Stern, Hanseok Ko |
INTERSPEECH | 3 |
| 2005 | Effective acoustic model clustering via decision-tree with supervised learning
Hanseok Ko |
Speech Commun. | 2 |
| 2005 | Bayesian fusion of confidence measures for speech recognitionabstractThe application of Bayesian fusion of confidence measures to speech recognition is proposed. Feature level, decision level, and hybrid fusion are considered under the Bayesian framework. The use of speaker-adapted feature-level Bayesian fusion reduced the error rate by 19.4% as compared to the conventional single feature-based confidence scoring in an isolated word out-of-vocabulary rejection test. The decision-level Bayesian fusion also showed better performance than the majority rule. Finally, hybrid Bayesian fusion, which can combine both confidence measure features and local decisions, achieved the best performance. Hanseok Ko |
IEEE Signal Process. Lett. | 2 |
| 2004 | PCMM-based feature compensation schemes using model interpolation and mixture sharingabstractIn this paper, we propose an effective feature compensation scheme based on the speech model in order to achieve robust speech recognition. The proposed feature compensation method is based on parallel combined mixture model (PCMM). The previous PCMM works require a highly sophisticated procedure for estimation of the combined mixture model in order to reflect the time-varying noisy conditions at every utterance. The proposed schemes can cope with the time-varying background noise by employing the interpolation method of the multiple mixture models. We apply the 'data-driven' method to PCMM for more reliable model combination and introduce a frame-synched version for estimation of environments a posteriori. In order to reduce the computational complexity due to multiple models, we propose a technique for mixture sharing. The statistically similar Gaussian components are selected and the smoothed versions are generated for sharing. The performance was examined over Aurora 2.0 and speech corpus recorded while car-driving. The experimental results indicate that the proposed schemes are effective in realizing robust speech recognition and reducing the computational complexities under both simulated environments and real-life conditions. Wooil Kim, Ohil Kwon, Hanseok Ko |
ICASSP (1) | 3 |
| 2004 | Face detection using support vector domain description in color imagesabstractWe present a system for face detection in color images using the support vector domain description (SVDD). Conventional face detection algorithms require a training procedure using both face and non-face images. In the SVDD, however, we employ only face images for training. We can detect faces in color images from the radius and center pairs of SVDD. We also use entropic threshold for extracting the facial feature and sliding window for improved performance while saving processing time. Experimental results indicate the effectiveness and efficiency of the proposed algorithm compared to the conventional PCA (principal component analysis) based methods. Jin Seo, Hanseok Ko |
ICASSP (5) | 2 |
| 2004 | Multi-eigenspace normalization for robust speech recognition in noisy environmentsabstractIn this paper, we propose an effective feature normalization scheme based on eigenspace normalization, for achieving robust speech recognition. In general, Mean and Variance Normalization (MVN) is implemented in cepstral domain. However, another MVN approach using eigenspace was recently introduced, in that the eigenspace normalization procedure performs normalization in a single eigenspace. This procedure consists of linear PCA matrix feature transformation followed by mean and variance normalization of the transformed cepstral feature. In the proposed scheme, we apply independent and unique eigenspaces to cepstra, delta and delta-delta cepstra respectively. We also normalize training data in eigenspace. In addition, a feature space rotation procedure is introduced to reduce the mismatch of training and test data distribution in noisy condition. As a result, we obtained a substantial improvement over the basic eigenspace normalization. Hanseok Ko |
INTERSPEECH | 2 |
| 2004 | Compact acoustic model for embedded implementationabstractAn acoustic model for an embedded speech recognition system must exhibit two desirable features; ability to minimize performance degradation in recognition while solving the memory problem under limited system resources. To cope with the challenges, we introduce the state-clustered tied-mixture (SCTM) HMM as an acoustic model optimization. The proposed SCTM modeling shows a significant improvement in recognition performance as well as a solution to sparse training data problem. Moreover, the state weight quantizing method achieves a drastic reduction in model size. In this paper, we describe the acoustic model optimization procedure for embedded speech recognition system and corresponding performance evaluation results. Hanseok Ko |
INTERSPEECH | 2 |
| 2004 | A New Feature Normalization Scheme Based on Eigenspace for Noisy Speech Recognition
Hanseok Ko |
SPIRE | 2 |
| 2003 | A novel spectral subtraction scheme for robust speech recognition: spectral subtraction using spectral harmonics of speechabstractThis paper addresses a novel noise-compensation scheme to solve the mismatch problem between training condition and testing condition for an automatic speech recognition (ASR) system, specifically in in-car environments. The conventional spectral subtraction schemes rely on the signal to noise ratio (SNR) such that attenuation is imposed on that part of the spectrum that appears to have low SNR, and accentuation is made on that part of high SNR. However, these schemes are based on the postulation that the power spectrum of noise is in general at the lower level in magnitude than that of speech. Therefore, while such postulation is adequate for high SNR environment, it is grossly inadequate for low SNR scenarios such as the in-car environment. This paper proposes an efficient spectral subtraction scheme focused to specifically low SNR noisy environments by distinguishing the speech-dominant segment from the noise-dominant segment in the speech spectrum. Representative experiments confirm the superior performance of the proposed method over conventional methods. The experiments are conducted using car noise-corrupted utterances of the Aurora2 corpus. Jounghoon Beh, Hanseok Ko |
ICASSP (1) | 2 |
| 2003 | A novel spectral subtraction scheme for robust speech recognition: spectral subtraction using spectral harmonics of speechabstractThis paper addresses a novel noise-compensation scheme to solve the mismatch problem between training condition and testing condition for the automatic speech recognition (ASR) system, specifically in the car environments. The conventional spectral subtraction schemes rely on the signal to noise ratio (SNR) such that attenuation is imposed on that part of the spectrum that appears to have low SNR, and accentuation is made on that part of high SNR. However, since these schemes are based on the postulation that the power spectrum of noise is in general at the lower level in magnitude than that of speech. Therefore, while such postulation is adequate for high SNR environment, it is grossly inadequate for low SNR scenarios such as a car environment. This paper proposes an efficient spectral subtraction scheme focused to specifically low SNR noisy environments by distinguishing the speech-dominant segment from the nose-dominant segment in speech spectrum. Representative experiments confirm the superior performance of the proposed method over conventional methods. The experiments are conducted using car noise-corrupted utterances of Aurora2 corpus. Jounghoon Beh, Hanseok Ko |
ICME | 2 |
| 2003 | Feature compensation scheme based on parallel combined mixture model
Wooil Kim, Sungjoo Ahn, Hanseok Ko |
INTERSPEECH | 3 |
| 2003 | Utterance verification under distributed detection and fusion framework
Hanseok Ko |
INTERSPEECH | 2 |
| 2002 | Improved acoustic modeling based on selective data-driven PMCabstractThis paper proposes an effective method to remedy the acoustic modeling problem inherent in the usual log-normal PMC intended for achieving robust speech recognition. In particular, the Gaussian kernels under the prescribed log-normal PMC cannot sufficiently express the corrupted speech distributions. The proposed scheme corrects this deficiency by judicially selecting the “fairly” corrupted component and by re-estimating it as a mixture of two distributions using data-driven PMC. As a result, some components become merged while equal number of components split. The determination for splitting or merging is achieved by means of measuring the similarity of corrupted speech model to those of clean model and noise model. The experimental results indicate that the suggested algorithm is effective in representing the corrupted speech distributions and attains consistent improvement over various SNR and noise cases. Wooil Kim, Hanseok Ko |
ICASSP | 2 |
| 2002 | Multiple vehicle tracking based on regional estimation in nighttime CCD imagesabstractIn this paper, we develop an image based tracking algorithm of multiple vehicles focused to effective detection and segmentation of moving objects for tracking under poor environmental conditions. In particular, we propose a novel image-tracking algorithm aimed at being robust to occlusion, false alarms, missed detection, and partial or multiple detection of target objects, adverse conditions considered as important issues in Intelligent Transportation System (ITS). Upon applying the Retinex algorithm as preprocessing to reduce the illumination effects at nighttime images, we apply a two-step tracking procedure, performing regional search and track. A regional estimation is first achieved based on a gating using probability data association, to initiate the search. We then invoke an object-oriented multiprocessing for multiple vehicle tracking under poor conditions. Representative experimental results show that the proposed method is effective in CCD images. Ilkwang Lee, Hanseok Ko, David K. Han |
ICASSP | 2 |
| 2002 | Achieving Real-Time Lip Synch via SVM-Based Phoneme Classification and Lip Shape RefinementabstractIn this paper, we develop a real time lip-synch system that activates a 2D avatar's lip motion in synch with incoming speech utterance. To realize "real time" operation of the system, we contain the processing time by invoking a merge and split procedure performing coarse-to-fine phoneme classification. At each stage of phoneme classification, we apply a support vector machine (SVM) to constrain the computational load while attaining desirable accuracy. Coarse-to-fine phoneme classification is accomplished via 2 stages of feature extraction, where each speech frame is acoustically analyzed first for 3 classes of lip opening using MFCC as the feature and then a further refined classification for detailed lip shape using formant information. We implemented the system with 2D lip animation that shows the effectiveness of the proposed 2-stage procedure accomplishing the real-time lip-synch task. Yongsung Kang, Hanseok Ko |
ICMI | 3 |
| 2002 | On effective speaker verification based on subword model
Sungjoo Ahn, Sunmee Kang, Hanseok Ko |
INTERSPEECH | 3 |
| 2002 | Construction of decision tree from data driven clustering
Hanseok Ko |
INTERSPEECH | 2 |
| 2001 | Tracking of mobile phone using IMM in CDMA environmentabstractThis paper proposes an effective method to localize mobile phones in a CDMA environment. This is to remedy the performance limitation inherent in the traditional localization algorithms, which make use of the present information only. If a Kalman filter is used that includes the previous information of location of a mobile unit, then the location error can be significantly reduced. Since it is difficult to represent the actual movement of a user by only one motion model, better error performance will be shown if an interacting multiple model (IMM) that uses several Kalman filter models is applied in place of just one Kalman filter. Performance analysis of the location error between the Kalman filter and IMM implementations confirm our postulation that the IMM significantly reduces location error. Jihyo Lee, Hanseok Ko |
ICASSP | 2 |
| 2001 | Model based stress decision methodabstractThis paper proposes an effective decision method focused to evaluate the “stress position”. Conventional methods usually extract the acoustic parameters and compare them to reference in absolute scale, adversely producing unstable results as testing condition changes. To cope with the environmental dependency, the proposed method is designed to be model-based that determines the stressed interval by making relative comparison over candidates. The stressed/unstressed models are then induced from normal phone models by adaptive training. The experimental results indicate that the proposed method is promising and that it is useful for automatic detection of stress positions. The results also show that generating the stressed/unstressed model by adaptive training is effective. Wooil Kim, Sungjoo Ahn, Hanseok Ko |
INTERSPEECH | 4 |
| 2000 | Effective speaker adaptations for speaker verificationabstractThis paper concerns effective speaker adaptation methods to solve the over-training problem in speaker verification, which frequently occurs when modeling a speaker with sparse training data. While various speaker adaptations have already been applied to speech recognition, these methods have not yet been formally considered in speaker verification. This paper proposes speaker adaptation methods using a combination of maximum a posteriori (MAP) and maximum likelihood linear regression (MLLR) adaptations, which are successfully used in speech recognition, and applies to speaker verification. Our aim is to remedy the small training data problem by investigating effective speaker adaptations for speaker modeling. Experimental results show that the speaker verification system using a weighted MAP and MLLR adaptation outperforms that of the conventional speaker models without adaptation by a factor of up to 5 times. From these results, we show that the speaker adaptation method achieves significantly better performance even when only small training data is available for speaker verification. Sungjoo Ahn, Sunmee Kang, Hanseok Ko |
ICASSP | 3 |
| 2000 | An effective acoustic modeling of names based on model inductionabstractIn a speech recognition based automatic directory assistance service, name modeling is an important issue that directly affects the overall system performance. In this paper, we propose an effective name modeling method considering the similarity property of names. In particular, we use explicit models to capture the common surnames (as in Korean names) while phone models are used to capture the less common first names. The proposed algorithm includes the surname model induction as a remedy to the insufficient training data problem caused by the model number increase. To efficiently induce the surname model, a model selection method based on the Bayesian information criterion (BIC) is introduced. Our experiment shows that the proposed name modeling method is effective and that the model induction method using BIC produces compact but accurate models. Sunmee Kang, Hanseok Ko |
ICASSP | 3 |
| 2000 | Dynamical behavior of autoassociative memory performing novelty filtering for signal enhancementabstractThis paper concerns the dynamical behavior, in probabilistic sense, of a simple perceptron network with sigmoidal output units performing autoassociation for novelty filtering. Networks of retinotopic topology having a one-to-one correspondence between input and output units can be readily trained using the delta learning rule, to perform autoassociative mappings. A novelty filter is obtained by subtracting the network output from the input vector. Then the presentation of a "familiar" pattern tends to evoke a null response; but any anomalous component is enhanced. Such a behavior exhibits a promising feature for enhancement of weak signals in additive noise. As an analysis of the novelty filtering, this paper shows that the probability density function of the weight converges to Gaussian when the input time series is statistically characterized by nonsymmetrical probability density functions. After output units are locally linearized, the recursive relation for updating the weight of the neural network is converted into a first-order random differential equation. Based on this equation it is shown that the probability density function of the weight satisfies the Fokker-Planck equation. By solving the Fokker-Planck equation, it is found that the weight is Gaussian distributed with time dependent mean and variance. Hanseok Ko, Garry M. Jacyna |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 1994 | Signal detectability enhancement with auto-associative backpropagation networks
Hanseok Ko, Robert H. Baran |
Neurocomputing | 1 |