Hoirin Kim

dblp:69/815 · also Hoi Rin Kim · DBLP profile ↗
← Back
49ranked-venue papers
0as first author
15since 2021 · last 2025
0000-0002-8787-6982ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 43 · 15 since 2021Artificial intelligence and machine learning · 31 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
abstract
This paper examines how linguistic similarity affects cross-lingual phonetic representation in speech processing for low-resource languages, emphasizing effective source language selection. Previous cross-lingual research has used various source languages to enhance performance for the target low-resource language without thorough consideration of selection. Our study stands out by providing an in-depth analysis of language selection, supported by a practical approach to assess phonetic proximity among multiple language families. We investigate how within-family similarity impacts performance in multilingual training, which aids in understanding language dynamics. We also evaluate the effect of using phonologically similar languages, regardless of family. For the phoneme recognition task, utilizing phonologically similar languages consistently achieves a relative improvement of 55.6% over monolingual training, even surpassing the performance of a large-scale self-supervised learning model. Multilingual training within the same language family demonstrates that higher phonological similarity enhances performance, while lower similarity results in degraded performance compared to monolingual training.
Minu Kim 0001, Kangwook Jang, Hoirin Kim
ICASSP3
2025 ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
Minu Kim 0001, Kangwook Jang, Hoirin Kim
INTERSPEECH3
2025 HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
Hyebin Ahn, Kangwook Jang, Hoirin Kim
INTERSPEECH3
2024 STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
abstract
Albeit great performance of Transformer-based speech self-supervised learning (SSL) models, their large parameter size and computational cost make them unfavorable to utilize. In this study, we propose to compress the speech SSL models by distilling speech temporal relation (STaR). Unlike previous works that directly match the representation for each speech frame, STaR distillation transfers temporal relation between speech frames, which is more suitable for lightweight student with limited capacity. We explore three STaR distillation objectives and select the best combination as the final STaR loss. Our model distilled from HuBERT Base achieves an overall score of 79.8 on SUPERB benchmark, the best performance among models with up to 27 million parameters. We show that our method is applicable across different speech SSL models and maintains robust performance with further reduced parameters.
Kangwook Jang, Sungnyun Kim, Hoirin Kim
ICASSP3
2024 One-class learning with adaptive centroid shift for audio deepfake detection
abstract
As speech synthesis systems continue to make remarkable advances in recent years, the importance of robust deepfake detection systems that perform well in unseen systems has grown.In this paper, we propose a novel adaptive centroid shift (ACS) method that updates the centroid representation by continually shifting as the weighted average of bonafide representations.Our approach uses only bonafide samples to define their centroid, which can yield a specialized centroid for one-class learning.Integrating our ACS with one-class learning gathers bonafide representations into a single cluster, forming wellseparated embeddings robust to unseen spoofing attacks.Our proposed method achieves an equal error rate (EER) of 2.19% on the ASVspoof 2021 deepfake dataset, outperforming all existing systems.Furthermore, the t-SNE visualization illustrates that our method effectively maps the bonafide embeddings into a single cluster and successfully disentangles the bonafide and spoof classes.
Hyun Myung Kim, Kangwook Jang, Hoirin Kim
INTERSPEECH3
2024 Utilizing Adaptive Global Response Normalization and Cluster-Based Pseudo Labels for Zero-Shot Voice Conversion
Ji Sub Um, Hoirin Kim
INTERSPEECH2
2024 Learning Video Temporal Dynamics With Cross-Modal Attention For Robust Audio-Visual Speech Recognition
abstract
Audio-visual speech recognition (AVSR) aims to transcribe human speech using both audio and video modalities. In practical environments with noise-corrupted audio, the role of video information becomes crucial. However, prior works have primarily focused on enhancing audio features in AVSR, overlooking the importance of video features. In this study, we strengthen the video features by learning three temporal dynamics in video data: context order, playback direction, and the speed of video frames. Cross-modal attention modules are introduced to enrich video features with audio information so that speech variability can be taken into account when training on the video temporal dynamics. Based on our approach, we achieve the state-of-the-art performance on the LRS2 and LRS3 AVSR benchmarks for the noise-dominant settings. Our approach excels in scenarios especially for babble and speech noise, indicating the ability to distinguish the speech signal that should be recognized from lip movements in the video modality. We support the validity of our methodology by offering the ablation experiments for the temporal dynamics losses and the cross-modal attention architecture design.
Sungnyun Kim, Kangwook Jang, Sangmin Bae, Hoirin Kim, Se-Young Yun
SLT4
2023 Recycle-and-Distill: Universal Compression Strategy for Transformer-based Speech SSL Models with Attention Map Reusing and Masking Distillation
abstract
Transformer-based speech self-supervised learning (SSL) models, such as HuBERT, show surprising performance in various speech processing tasks.However, huge number of parameters in speech SSL models necessitate the compression to a more compact model for wider usage in academia or small companies.In this study, we suggest to reuse attention maps across the Transformer layers, so as to remove key and query parameters while retaining the number of layers.Furthermore, we propose a novel masking distillation strategy to improve the student model's speech representation quality.We extend the distillation loss to utilize both masked and unmasked speech frames to fully leverage the teacher model's high-quality representation.Our universal compression strategy yields the student model that achieves phoneme error rate (PER) of 7.72% and word error rate (WER) of 9.96% on the SUPERB benchmark.
Kangwook Jang, Sungnyun Kim, Se-Young Yun, Hoirin Kim
INTERSPEECH4
2023 AdaMS: Deep Metric Learning with Adaptive Margin and Adaptive Scale for Acoustic Word Discrimination
Myunghun Jung, Hoirin Kim
INTERSPEECH2
2022 Anti-Spoofing Using Transfer Learning with Variational Information Bottleneck
abstract
Recent advances in sophisticated synthetic speech generated from text-to-speech (TTS) or voice conversion (VC) systems cause threats to the existing automatic speaker verification (ASV) systems. Since such synthetic speech is generated from diverse algorithms, generalization ability with using limited training data is indispensable for a robust anti-spoofing system. In this work, we propose a transfer learning scheme based on the wav2vec 2.0 pretrained model with variational information bottleneck (VIB) for speech anti-spoofing task. Evaluation on the ASVspoof 2019 logical access (LA) database shows that our method improves the performance of distinguishing unseen spoofed and genuine speech, outperforming current state-of-the-art anti-spoofing systems. Furthermore, we show that the proposed system improves performance in low-resource and cross-dataset settings of anti-spoofing task significantly, demonstrating that our system is also robust in terms of data size and data distribution.
Youngsik Eom, Yeonghyeon Lee, Ji Sub Um, Hoirin Kim
INTERSPEECH4
2022 Asymmetric Proxy Loss for Multi-View Acoustic Word Embeddings
abstract
Acoustic word embeddings (AWEs) are discriminative representations of speech segments, and learned embedding space reflects the phonetic similarity between words.With multi-view learning, where text labels are considered as supplementary input, AWEs are jointly trained with acoustically grounded word embeddings (AGWEs).In this paper, we expand the multiview approach into a proxy-based framework for deep metric learning by equating AGWEs with proxies.A simple modification in computing the similarity matrix allows the general pair weighting to formulate the data-to-proxy relationship.Under the systematized framework, we propose an asymmetric-proxy loss that combines different parts of loss functions asymmetrically while keeping their merits.It follows the assumptions that the optimal function for anchor-positive pairs may differ from one for anchor-negative pairs, and a proxy may have a different impact when it substitutes for different positions in the triplet.We present comparative experiments with various proxy-based losses including our asymmetric-proxy loss, and evaluate AWEs and AGWEs for word discrimination tasks on WSJ corpus.The results demonstrate the effectiveness of the proposed method.
Myunghun Jung, Hoirin Kim
INTERSPEECH2
2022 FitHuBERT: Going Thinner and Deeper for Knowledge Distillation of Speech Self-Supervised Models
Yeonghyeon Lee, Kangwook Jang, Jahyun Goo, Youngmoon Jung, Hoirin Kim
INTERSPEECH5
2022 ACNN-VC: Utilizing Adaptive Convolution Neural Network for One-Shot Voice Conversion
Ji Sub Um, Yeunju Choi, Hoirin Kim
INTERSPEECH3
2021 Neural MOS Prediction for Synthesized Speech Using Multi-Task Learning with Spoofing Detection and Spoofing Type Classification
abstract
Several studies have proposed deep-learning-based models to predict the mean opinion score (MOS) of synthesized speech, showing the possibility of replacing human raters. However, inter- and intra-rater variability in MOSs makes it hard to en-sure the high performance of the models. In this paper, we propose a multi-task learning (MTL) method to improve the performance of a MOS prediction model using the following two auxiliary tasks: spoofing detection (SD) and spoofing type classification (STC). Besides, we use the focal loss to maximize the synergy between SD and STC for MOS pre-diction. Experiments using the MOS evaluation results of the Voice Conversion Challenge 2018 show that proposed MTL with two auxiliary tasks improves MOS prediction. Our proposed model achieves up to 11.6% relative improvement in performance over the baseline model.
Yeunju Choi, Youngmoon Jung, Hoirin Kim
SLT3
2021 Supervised Attention for Speaker Recognition
abstract
The recently proposed self-attentive pooling (SAP) has shown good performance in several speaker recognition systems. In SAP systems, the context vector is trained end-to-end together with the feature extractor, where the role of context vector is to select the most discriminative frames for speaker recognition. However, the SAP underperforms compared to the temporal average pooling (TAP) baseline in some settings, which implies that the attention is not learnt effectively in end-to-end training. To tackle this problem, we introduce strategies for training the attention mechanism in a supervised manner, which learns the context vector using classified samples. With our proposed methods, context vector can be boosted to select the most informative frames. We show that our method outperforms existing methods in various experimental settings including short utterance speaker recognition, and achieves competitive performance over the existing baselines on the VoxCeleb datasets.
Seong Min Kye, Joon Son Chung, Hoirin Kim
SLT3
2020 Deep MOS Predictor for Synthetic Speech Using Cluster-Based Modeling
abstract
While deep learning has made impressive progress in speech synthesis and voice conversion, the assessment of the synthesized speech is still carried out by human participants. Several recent papers have proposed deep-learning-based assessment models and shown the potential to automate the speech quality assessment. To improve the previously proposed assessment model, MOSNet, we propose three models using cluster-based modeling methods: using a global quality token (GQT) layer, using an Encoding Layer, and using both of them. We perform experiments using the evaluation results of the Voice Conversion Challenge 2018 to predict the mean opinion score of synthesized speech and similarity score between synthesized speech and reference speech. The results show that the GQT layer helps to predict human assessment better by automatically learning the useful quality tokens for the task and that the Encoding Layer helps to utilize frame-level scores more precisely.
Yeunju Choi, Youngmoon Jung, Hoirin Kim
INTERSPEECH3
2020 Multi-Task Network for Noise-Robust Keyword Spotting and Speaker Verification Using CTC-Based Soft VAD and Global Query Attention
abstract
Keyword spotting (KWS) and speaker verification (SV) have been studied independently although it is known that acoustic and speaker domains are complementary. In this paper, we propose a multi-task network that performs KWS and SV simultaneously to fully utilize the interrelated domain information. The multi-task network tightly combines sub-networks aiming at performance improvement in challenging conditions such as noisy environments, open-vocabulary KWS, and short-duration SV, by introducing novel techniques of connectionist temporal classification (CTC)-based soft voice activity detection (VAD) and global query attention. Frame-level acoustic and speaker information is integrated with phonetically originated weights so that forms a word-level global representation. Then it is used for the aggregation of feature vectors to generate discriminative embeddings. Our proposed approach shows 4.06% and 26.71% relative improvements in equal error rate (EER) compared to the baselines for both tasks. We also present a visualization example and results of ablation experiments.
Myunghun Jung, Youngmoon Jung, Jahyun Goo, Hoirin Kim
INTERSPEECH4
2020 Improving Multi-Scale Aggregation Using Feature Pyramid Module for Robust Speaker Verification of Variable-Duration Utterances
abstract
Currently, the most widely used approach for speaker verification is the deep speaker embedding learning. In this approach, we obtain a speaker embedding vector by pooling single-scale features that are extracted from the last layer of a speaker feature extractor. Multi-scale aggregation (MSA), which utilizes multi-scale features from different layers of the feature extractor, has recently been introduced and shows superior performance for variable-duration utterances. To increase the robustness dealing with utterances of arbitrary duration, this paper improves the MSA by using a feature pyramid module. The module enhances speaker-discriminative information of features from multiple layers via a top-down pathway and lateral connections. We extract speaker embeddings using the enhanced features that contain rich speaker information with different time scales. Experiments on the VoxCeleb dataset show that the proposed module improves previous MSA methods with a smaller number of parameters. It also achieves better performance than state-of-the-art approaches for both short and long utterances.
Youngmoon Jung, Seong Min Kye, Yeunju Choi, Myunghun Jung, Hoirin Kim
INTERSPEECH5
2020 Meta-Learning for Short Utterance Speaker Recognition with Imbalance Length Pairs
abstract
In practical settings, a speaker recognition system needs to identify a speaker given a short utterance, while the enrollment utterance may be relatively long.However, existing speaker recognition models perform poorly with such short utterances.To solve this problem, we introduce a meta-learning framework for imbalance length pairs.Specifically, we use a Prototypical Networks and train it with a support set of long utterances and a query set of short utterances of varying lengths.Further, since optimizing only for the classes in the given episode may be insufficient for learning discriminative embeddings for unseen classes, we additionally enforce the model to classify both the support and the query set against the entire set of classes in the training set.By combining these two learning schemes, our model outperforms existing state-of-the-art speaker verification models learned with a standard supervised learning framework on short utterance (1-2 seconds) on the VoxCeleb datasets.We also validate our proposed model for unseen speaker identification, on which it also achieves significant performance gains over the existing approaches.The codes are available at https://github.com/seongmin-kye/meta-SR.
Seong Min Kye, Youngmoon Jung, Haebeom Lee, Sung Ju Hwang, Hoirin Kim
INTERSPEECH5
2020 Dual Attention in Time and Frequency Domain for Voice Activity Detection
abstract
Voice activity detection (VAD) is a challenging task in low signal-to-noise ratio (SNR) environment, especially in non-stationary noise. To deal with this issue, we propose a novel attention module that can be integrated in Long Short-Term Memory (LSTM). Our proposed attention module refines each LSTM layer's hidden states so as to make it possible to adaptively focus on both time and frequency domain. Experiments are conducted on various noisy conditions using Aurora 4 database. Our proposed method obtains the 95.58 % area under the ROC curve (AUC), achieving 22.05 % relative improvement compared to baseline, with only 2.44 % increase in the number of parameters. Besides, we utilize focal loss for alleviating the performance degradation caused by imbalance between speech and non-speech sections in training sets. The results show that the focal loss can improve the performance in various imbalance situations compared to the cross entropy loss, a commonly used loss function in VAD.
Joohyung Lee 0001, Youngmoon Jung, Hoirin Kim
INTERSPEECH3
2020 Interlayer Selective Attention Network for Robust Personalized Wake-Up Word Detection
abstract
Previous research methods on wake-up word detection (WWD) have been proposed with focus on finding a decent word representation that can well express the characteristics of a word. However, there are various obstacles such as noise and reverberation which make it difficult in real-world environments where WWD works. To tackle this, we propose a novel architecture called interlayer selective attention network (ISAN) which generates more robust word representation by introducing the concept of selective attention. Experiments in real-world scenarios demonstrated that the proposed ISAN outperformed several baseline methods as well as other attention methods. In addition, the effectiveness of ISAN was analyzed with visualizations.
Hyungjun Lim, Younggwan Kim, Jahyun Goo, Hoirin Kim
IEEE Signal Process. Lett.4
2020 Cross-Informed Domain Adversarial Training for Noise-Robust Wake-Up Word Detection
abstract
A proper representation that can well express the characteristics of a word plays an important role in wake-up word detection (WWD). However, it may be easily corrupted due to various types of environmental noise occurred in the place where WWD typically works, causing unreliable performance. To deal with this practical issue, we propose a novel strategy called cross-informed domain adversarial training (CiDAT) for noise-robust WWD. In the method, additional paths were introduced to conventional domain adversarial training (DAT) to encourage its ability to generate domain-invariant representation. Experiments on the Aurora4 corpus verified that CiDAT significantly outperformed the baselines as well as conventional DAT.
Hyungjun Lim, Younggwan Kim, Hoirin Kim
IEEE Signal Process. Lett.3
2019 Self-Adaptive Soft Voice Activity Detection Using Deep Neural Networks for Robust Speaker Verification
abstract
Voice activity detection (VAD), which classifies frames as speech or non-speech, is an important module in many speech applications including speaker verification. In this paper, we propose a novel method, called self-adaptive soft VAD, to incorporate a deep neural network (DNN)-based VAD into a deep speaker embedding system. The proposed method is a combination of the following two approaches. The first approach is soft VAD, which performs a soft selection of frame-level features extracted from a speaker feature extractor. The frame-level features are weighted by their corresponding speech posteriors estimated from the DNN-based VAD, and then aggregated to generate a speaker embedding. The second approach is self-adaptive VAD, which fine-tunes the pre-trained VAD on the speaker verification data to reduce the domain mismatch. Here, we introduce two unsupervised domain adaptation (DA) schemes, namely speech posterior-based DA (SP-DA) and joint learning-based DA (JL-DA). Experiments on a Korean speech database demonstrate that the verification performance is improved significantly in real-world environments by using self-adaptive soft VAD.
Youngmoon Jung, Yeunju Choi, Hoirin Kim
ASRU3
2019 Additional Shared Decoder on Siamese Multi-View Encoders for Learning Acoustic Word Embeddings
abstract
Acoustic word embeddings - fixed-dimensional vector representations of arbitrary-length words - have attracted increasing interest in query-by-example spoken term detection. Recently, on the fact that the orthography of text labels partly reflects the phonetic similarity between the words' pronunciation, a multi-view approach has been introduced that jointly learns acoustic and text embeddings. It showed that it is possible to learn discriminative embeddings by designing the objective which takes text labels as well as word segments. In this paper, we propose a network architecture that expands the multi-view approach by combining the Siamese multiview encoders with a shared decoder network to maximize the effect of the relationship between acoustic and text embeddings in embedding space. Discriminatively trained with multi-view triplet loss and decoding loss, our proposed approach achieves better performance on acoustic word discrimination task with the WSJ dataset, resulting in 11.1% relative improvement in average precision. We also present experimental results on cross-view word discrimination and word level speech recognition tasks.
Myunghun Jung, Hyungjun Lim, Jahyun Goo, Youngmoon Jung, Hoirin Kim
ASRU5
2019 Spatial Pyramid Encoding with Convex Length Normalization for Text-Independent Speaker Verification
abstract
In this paper, we propose a new pooling method called spatial pyramid encoding (SPE) to generate speaker embeddings for text-independent speaker verification. We first partition the output feature maps from a deep residual network (ResNet) into increasingly fine sub-regions and extract speaker embeddings from each sub-region through a learnable dictionary encoding layer. These embeddings are concatenated to obtain the final speaker representation. The SPE layer not only generates a fixed-dimensional speaker embedding for a variable-length speech segment, but also aggregates the information of feature distribution from multi-level temporal bins. Furthermore, we apply deep length normalization by augmenting the loss function with ring loss. By applying ring loss, the network gradually learns to normalize the speaker embeddings using model weights themselves while preserving convexity, leading to more robust speaker embeddings. Experiments on the VoxCeleb1 dataset show that the proposed system using the SPE layer and ring loss-based deep length normalization outperforms both i-vector and d-vector baselines.
Youngmoon Jung, Younggwan Kim, Hyungjun Lim, Yeunju Choi, Hoirin Kim
INTERSPEECH5
2018 Joint Learning Using Denoising Variational Autoencoders for Voice Activity Detection
Youngmoon Jung, Younggwan Kim, Yeunju Choi, Hoirin Kim
INTERSPEECH4
2018 Learning Self-Informed Feature Contribution for Deep Learning-Based Acoustic Modeling
abstract
In this paper, we introduce a new feature engineering approach for deep learning-based acoustic modeling, which utilizes input feature contributions. For this purpose, we propose an auxiliary deep neural network (DNN) called a feature contribution network (FCN) whose output layer is composed of sigmoid-based contribution gates. In our framework, the FCN tries to learn element-level discriminative contributions of input features and an acoustic model network (AMN) is trained by gated features generated by element-wise multiplication between contribution gate outputs and input features. In addition, we also propose a regularization method for the FCN, which helps the FCN to activate the minimum number of the gates. The proposed methods were evaluated on the TED-LIUM release 1 corpus. We applied the proposed methods to DNN- and long short-term memory-based AMNs. Experimental results results showed that AMNs with the FCNs consistently improved recognition performance compared with AMN-only frameworks.
Younggwan Kim, Myung Jong Kim, Jahyun Goo, Hoirin Kim
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Deep Least Squares Regression for Speaker Adaptation
Younggwan Kim, Hyungjun Lim, Jahyun Goo, Hoirin Kim
INTERSPEECH4
2016 Cross-acoustic transfer learning for sound event classification
abstract
A well-trained acoustic model that effectively captures the characteristics of sound events is a critical factor to develop more reliable system for sound event classification. Deep neural network (DNN) which has an ability to extract discriminative representation of features can be a good candidate for acoustic model of sound events. Compared to other data such as speech or image, the amount of sound database is often insufficient for learning the DNN properly, resulting in overfitting problems. In this paper, we propose a cross-acoustic transfer learning framework that can effectively train the DNN even with insufficient sound data by employing rich speech data. Three datasets are used to evaluate our proposed method; one sound dataset is from Real World Computing Partnership (RWCP) DB and two speech datasets are from Resource Management (RM) and Wall Street Journal (WSJ) DBs. A series of experimental results verify that cross-acoustic transfer learning performs significantly better than the baseline DNN which was trained only from sound data, achieving 26.24% relative classification error rate (CER) improvement over the DNN baseline system.
Hyungjun Lim, Myung Jong Kim, Hoirin Kim
ICASSP3
2016 Speaker Normalization Through Feature Shifting of Linearly Transformed i-Vector
Jahyun Goo, Younggwan Kim, Hyungjun Lim, Hoirin Kim
INTERSPEECH4
2016 Dysarthric Speech Recognition Using Kullback-Leibler Divergence-Based Hidden Markov Model
Myung Jong Kim, Jun Wang 0037, Hoirin Kim
INTERSPEECH3
2015 Speech emotion classification using tree-structured sparse logistic regression
Myung Jong Kim, Joohong Yoo, Younggwan Kim, Hoirin Kim
INTERSPEECH4
2015 Robust sound event classification using LBP-HOG based bag-of-audio-words feature representation
Hyungjun Lim, Myung Jong Kim, Hoirin Kim
INTERSPEECH3
2015 Probabilistic Class Histogram Equalization Based on Posterior Mean Estimation for Robust Speech Recognition
abstract
In this letter, we propose a new probabilistic class histogram equalization technique for noise robust speech recognition. To cope with the sparse data problem which is common in the case of short test data, the proposed histogram equalization technique employs the posterior mean estimator, a kind of the Bayesian estimator, for test CDF. Experiments on the Aurora-4 framework showed that the proposed method produces performance improvement over the conventional maximum likelihood estimation-based approach.
Youngjoo Suh, Hoirin Kim
IEEE Signal Process. Lett.2
2015 Automatic Intelligibility Assessment of Dysarthric Speech Using Phonologically-Structured Sparse Linear Model
abstract
This paper presents a new method for automatically assessing the speech intelligibility of patients with dysarthria, which is a motor speech disorder impeding the physical production of speech. The proposed method consists of two main steps: feature representation and prediction. In the feature representation step, the speech utterance is converted into a phone sequence using an automatic speech recognition technique and is then aligned with a canonical phone sequence from a pronunciation dictionary using a weighted finite state transducer to capture the pronunciation mappings such as match, substitution, and deletion. The histograms of the pronunciation mappings on a pre-defined word set are used for features. Next, in the prediction step, a structured sparse linear model incorporated with phonological knowledge that simultaneously addresses phonologically structured sparse feature selection and intelligibility prediction is proposed. Evaluation of the proposed method on a database of 109 speakers consisting of 94 dysarthric and 15 control speakers yielded a root mean square error of 8.14 compared to subjectively rated scores in the range of 0 to 100. This is a promising performance in which the system can be successfully applied to help speech therapists in diagnosing the degree of speech disorder.
Myung Jong Kim, Younggwan Kim, Hoirin Kim
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Constrained MLE-based speaker adaptation with L1 regularization
abstract
Maximum a posterior (MAP) adaptation is one of the popular and powerful methods for obtaining a speaker-specific acoustic model. Basically, MAP adaptation needs a data storage for speaker adaptive (SA) model as much as speaker independent (SI) model needs. Modern speech recognition systems have a huge number of parameters and deal with millions of users. To reduce the data storage for SA models, in this paper, we propose a constrained maximum likelihood estimation-based speaker adaptation with L1 regularization. By the proposed method, we can more efficiently perform the model adjustments for SA models without almost any loss of phone recognition performance than the conventional sparse MAP adaptation method.
Younggwan Kim, Hoirin Kim
ICASSP2
2013 Dysarthric speech recognition using dysarthria-severity-dependent and speaker-adaptive models
Myung Jong Kim, Joohong Yoo, Hoirin Kim
INTERSPEECH3
2012 Automatic Assessment of Dysarthric Speech Intelligibility Based on Selected Phonetic Quality Features
Myung Jong Kim, Hoirin Kim
ICCHP (2)2
2012 Combination of Multiple Speech Dimensions for Automatic Assessment of Dysarthric Speech Intelligibility
Myung Jong Kim, Hoirin Kim
INTERSPEECH2
2012 Multiple Acoustic Model-Based Discriminative Likelihood Ratio Weighting for Voice Activity Detection
abstract
In this letter, we propose a novel statistical voice activity detection (VAD) technique. The proposed technique employs probabilistically derived multiple acoustic models to effectively optimize the weights on frequency domain likelihood ratios with the discriminative training approach for more accurate voice activity detection. Experiments performed on various AURORA noisy environments showed that the proposed approach produces meaningful performance improvements over the single acoustic model-based conventional approaches.
Youngjoo Suh, Hoirin Kim
IEEE Signal Process. Lett.2
2012 Audio-Based Objectionable Content Detection Using Discriminative Transforms of Time-Frequency Dynamics
abstract
In this paper, the problem of detecting objectionable sounds, such as sexual screaming or moaning, to classify and block objectionable multimedia content is addressed. Objectionable sounds show distinctive characteristics, such as large temporal variations and fast spectral transitions, which are different from general audio signals, such as speech and music. To represent these characteristics, segment-based two-dimensional Mel-frequency cepstral coefficients and histograms of gradient directions are used as a feature set to characterize the time-frequency dynamics within a long-range segment of the target signal. After extracting the features, they are transformed to features with lower dimensions while preserving discriminative information using linear discriminant analysis based on a combination of global and local Fisher criteria. A Gaussian mixture model is adopted to statistically represent objectionable and non-objectionable sounds, and test sounds are classified by using a likelihood ratio test. Evaluation of the proposed feature extraction method on a database of several hundred objectionable and non-objectionable sound clips yielded precision/recall breakeven point of 91.25%, which is a promising performance which shows that the system can be applied to help an image-based approach to block such multimedia content.
Myung Jong Kim, Hoirin Kim
IEEE Trans. Multim.2
2010 Automatic detection of malicious sound using segmental two-dimensional mel-frequency cepstral coefficients and histograms of oriented gradients
abstract
This paper addresses the problem of recognizing malicious sounds, such as sexual scream or moan, to detect and block the objectionable multimedia contents. The malicious sounds show the distinct characteristics that have large temporal variations and fast spectral transitions. Therefore, extracting appropriate features to properly represent these characteristics is important in achieving a better performance. In this paper, we employ segment-based two-dimensional Mel-frequency cepstral coefficients and histograms of gradient directions as a feature set to characterize both the temporal variations and spectral transitions within a long-range segment of the target signal. Gaussian mixture model (GMM) is adopted to statistically represent the malicious and non-malicious sounds, and the test sounds are classified by a maximum a posterior probability (MAP) method. Evaluation of the proposed feature extraction method on a database of several hundred malicious and non-malicious sound clips yielded precision of 91.31% and recall of 94.27%. This result suggests that this approach could be used as an alternative to the image-based methods.
Myung Jong Kim, Younggwan Kim, JaeDeok Lim, Hoirin Kim
ACM Multimedia4
2010 Robust speaker recognition based on filtering in autocorrelation domain and sub-band feature recombination
Sungtak Kim, Miyoung Ji, Hoirin Kim
Pattern Recognit. Lett.3
2009 The effectiveness of histogram equalization on environmental model adaptation
abstract
In this paper, we introduce a new histogram equalization-based environmental model adaptation method for robust speech recognition in noise environments. The proposed method adapts initially-trained acoustic mean models of a speech recognizer into the environmentally matched models. The covariance models are adapted by using utterance-level local covariance matrices. We performed a series of experiments based on the Aurora2 framework to examine the effectiveness of the proposed environmental model adaptation technique. In both clean and multi-condition trainings, the proposed approach achieved substantial performance improvements over the baseline speech recognizers.
Youngjoo Suh, Hoirin Kim
ICASSP2
2009 Environmental Model Adaptation Based on Histogram Equalization
abstract
In this letter, a new environmental model adaptation method is proposed for robust speech recognition under noisy environments. The proposed method adapts initial acoustic models of a speech recognizer into environmentally matched models by utilizing the histogram equalization technique. Experiments performed on the Aurora noisy environment showed that the proposed technique provides substantial improvement over the baseline speech recognizer trained on the clean speech data.
Youngjoo Suh, Hoirin Kim
IEEE Signal Process. Lett.2
2007 Reliable Speaker Identification Using Multiple Microphones in Ubiquitous Robot Companion Environment
abstract
This paper presents a text-independent speaker identification system using multiple microphones on the robot, which is intended for use in human-robot interaction. For the purpose of the best possible classification rate in speaker identification, the individual identification results obtained from multiple microphones on the robot are combined by various combination schemes. The performance improvement has been achieved. Our ultimate goal is to enhance human-robot interaction by improving the recognition performance of speaker identification with multiple microphones on the robot side in adverse distant-talking environments. Various combination schemes obtained high classification accuracy in the ubiquitous robot companion (URC) environment, where the robot is connected to a server through extremely high broadband penetration rate. In conclusion, our speaker identification system can provide human-robot interaction with a reliable basic interface with high classification accuracy.
Mikyong Ji, Sungtak Kim, Hoirin Kim, Young-Jo Cho
RO-MAN3
2007 Probabilistic Class Histogram Equalization for Robust Speech Recognition
abstract
In this letter, a probabilistic class histogram equalization method is proposed to compensate for an acoustic mismatch in noise robust speech recognition. The proposed method aims not only to compensate for the acoustic mismatch between training and test environments but also to reduce the limitations of the conventional histogram equalization. It utilizes multiple class-specific reference and test cumulative distribution functions, classifies noisy test features into their corresponding classes by means of soft classification with a Gaussian mixture model, and equalizes the features by using their corresponding class-specific distributions. Experiments on the Aurora 2 task confirm the superiority of the proposed approach in acoustic feature compensation.
Youngjoo Suh, Mikyong Ji, Hoirin Kim
IEEE Signal Process. Lett.3
2006 A Music Summarization Scheme using Tempo Tracking and Two Stage Clustering
abstract
In this paper, we present effective methods for music summarization which automatically extract a representative portion of the music by signal processing technology. Our proposed method uses 2-dimensional similarity matrix, tempo tracking, and clustering techniques to extract several segments which have different moods or dissimilar semantic structure in the music. The segments extracted are combined to generate a complete music summary. The three main techniques used in this paper are well-known and widely used for extracting music summary. However, we use them in a different way, and experiments show the proposed method captures the main theme of the music more effectively than conventional methods. The experimental results also show that one of the proposed methods could be used for real-time application since the processing time in generating music summary is much faster than other methods
Sungtak Kim, Suk-Bong Kwon, Hoirin Kim
MMSP4
2006 Intelligent broadcasting system and services for personalized semantic contents consumption
Sung Ho Jin, Tae Meon Bae, Yong Man Ro, Hoirin Kim, Munchurl Kim
Expert Syst. Appl.4