EDBT 2026 Demo / reviewers in the wild / expert
Nam Soo Kim
dblp:22/792
· DBLP profile ↗
158ranked-venue papers
21as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 130 · 15 first-author · 27 since 2021Artificial intelligence and machine learning · 80 · 7 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 19.43-bit Effective Resolution and 3.9 kSPS SAR-Based Integrator-Residue Extended Counting First-Order ΔΣ ADC
Sanghwa Han, Hyunjoong Lee, Nam Soo Kim, Jaehoon Jun |
ISCAS | 3 |
| 2025 | Evidential-TTS: High Fidelity Zero-Shot Text-to-Speech Using Evidential Deep LearningabstractWe propose Evidential-TTS, a novel zero-shot text-to-speech (TTS) system based on evidential deep learning (EDL). The model includes a length regulator to ensure precise alignment between phonemes and acoustic tokens. This module allows the evidential token generator to convert the aligned phoneme sequence into acoustic tokens using iterative parallel decoding (IPD). However, IPD often suffers from overconfidence when using categorical probabilities as confidence scores. To address this, we introduce model uncertainty into the sampling process, quantified through EDL optimization. This uncertainty provides a more reliable sampling path for high-quality speech generation. Experimental results show that Evidential-TTS outperforms existing models in terms of speech naturalness and intelligibility. An ablation study further demonstrates the importance of uncertainty estimation in guiding the sampling trajectory of IPD. Myeonghun Jeong, Nam Soo Kim |
ICASSP | 4 |
| 2025 | FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep LearningabstractRecently, fake audio detection has gained significant attention, as advancements in speech synthesis and voice conversion have increased the vulnerability of automatic speaker verification (ASV) systems to spoofing attacks. A key challenge in this task is generalizing models to detect unseen, out-of-distribution (OOD) attacks. Although existing approaches have shown promising results, they inherently suffer from overconfidence issues due to the usage of softmax for classification, which can produce unreliable predictions when encountering unpredictable spoofing attempts. To deal with this limitation, we propose a novel framework called fake audio detection with evidential learning (FADEL). By modeling class probabilities with a Dirichlet distribution, FADEL incorporates model uncertainty into its predictions, thereby leading to more robust performance in OOD scenarios. Experimental results on the ASVspoof2019 Logical Access (LA) and ASVspoof2021 LA datasets indicate that the proposed method significantly improves the performance of baseline models. Furthermore, we demonstrate the validity of uncertainty estimation by analyzing a strong correlation between average uncertainty and equal error rate (EER) across different spoofing algorithms. Ju Yeon Kang, Jiwon Yoon 0002, Min Hyun Han, Nam Soo Kim |
ICASSP | 5 |
| 2025 | SNR-Aligned Consistent Diffusion for Adaptive Speech Enhancement
Yonghyeon Jun, Beomjun Woo, Myeonghun Jeong, Nam Soo Kim |
INTERSPEECH | 4 |
| 2025 | SegINR: Segment-Wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-SpeechabstractWe present SegINR, a novel approach to neural Text-to-Speech (TTS) that eliminates the need for either an auxiliary duration predictor or autoregressive (AR) sequence modeling for alignment. SegINR simplifies the TTS process by directly converting text sequences into frame-level features. Encoded text embeddings are transformed into segments of frame-level features with length regulation using a conditional implicit neural representation (INR). This method, termed Segment-wise INR (SegINR), captures temporal dynamics within each segment while autonomously defining segment boundaries, resulting in lower computational costs. Integrated into a two-stage TTS framework, SegINR is employed for semantic token prediction. Experiments in zero-shot adaptive TTS scenarios show that SegINR outperforms conventional methods in speech quality with computational efficiency. Myeonghun Jeong, Joun Yeop Lee, Nam Soo Kim |
IEEE Signal Process. Lett. | 4 |
| 2025 | Sampling-Based Pruned Knowledge Distillation for Training Lightweight RNN-TabstractWe present a novel training method for small-scale RNN-T models, widely used in real-world speech recognition applications. Despite efforts to scale down models for edge devices, the demand for even smaller and more compact speech recognition models persists to accommodate a broader range of devices. In this letter, we propose Sampling-based Pruned Knowledge Distillation (SP-KD) for training lightweight RNN-T models. In contrast to the conventional knowledge distillation techniques, the proposed method enables student models to distill knowledge from the distribution of teacher models, which is estimated by considering not only the best paths but also less likely paths. Additionally, we leverage pruning the output lattice of RNN-T to comprehensively transfer knowledge from teacher models to student models. Experimental results demonstrate that our proposed method outperforms the baseline in training tiny RNN-T models. Dongjune Lee, Ju Yeon Kang, Myeonghun Jeong, Nam Soo Kim |
IEEE Signal Process. Lett. | 5 |
| 2025 | Towards Maximum Likelihood Training for Transducer-Based Streaming Speech RecognitionabstractTransducer neural networks have emerged as the mainstream approach for streaming automatic speech recognition (ASR), offering state-of-the-art performance in balancing accuracy and latency. In the conventional framework, streaming transducer models are trained to maximize the likelihood function based on non-streaming recursion rules. However, this approach leads to a mismatch between training and inference, resulting in the issue of deformed likelihood and consequently suboptimal ASR accuracy. We introduce a mathematical quantification of the gap between the actual likelihood and the deformed likelihood, namely forward variable causal compensation (FoCC). We also present its estimator, FoCCE, as a solution to estimate the exact likelihood. Through experiments on the LibriSpeech dataset, we show that FoCCE training improves the accuracy of the streaming transducers. Hyeon Seung Lee, Jiwon Yoon 0002, Nam Soo Kim |
IEEE Signal Process. Lett. | 4 |
| 2024 | MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
Myeonghun Jeong, Hyeon Seung Lee, Byoung Jin Choi, Nam Soo Kim |
INTERSPEECH | 6 |
| 2024 | High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
Joun Yeop Lee, Myeonghun Jeong, Ji-Hyun Lee, Hoonyoung Cho, Nam Soo Kim |
INTERSPEECH | 6 |
| 2024 | HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
Jiwon Yoon 0002, Beom Jun Woo, Nam Soo Kim |
INTERSPEECH | 3 |
| 2024 | Variable-Length Speaker Conditioning in Flow-Based Text-to-SpeechabstractIn this letter, we propose a novel speaker conditioning technique that leverages a variable-length reference embedding sequence for flow-based text-to-speech (TTS) architecture in the context of zero-shot multi-speaker text-to-speech (ZSM-TTS). Unlike conventional ZSM-TTS methods, which usually rely on a single fixed-dimensional vector to represent the entire reference speech, our approach aims to extract variable-length embedding sequence for a more flexible and efficient conditioning. We enhance the current affine coupling function in flow-based TTS architecture by introducing an attentive speaker conditioning. This allows a local variation of the speaker conditioning. Our experiments demonstrate the effectiveness of the proposed method, highlighting improvements in terms of speaker similarity, speech naturalness, and speech intelligibility compared to the baseline methods. Byoung Jin Choi, Myeonghun Jeong, Nam Soo Kim |
IEEE Signal Process. Lett. | 4 |
| 2024 | Efficient Parallel Audio Generation Using Group Masked Language ModelingabstractWe present a fast and high-quality codec language model for parallel audio generation. While SoundStorm, a state-of-the-art parallel audio generation model, accelerates inference speed compared to autoregressive models, it still suffers from slow inference due to iterative sampling. To resolve this problem, we propose Group-Masked Language Modeling (G-MLM) and Group Iterative Parallel Decoding (G-IPD) for efficient parallel audio generation. Both the training and sampling schemes enable the model to synthesize high-quality audio with a small number of iterations by effectively modeling the group-wise conditional dependencies. In addition, our model employs a cross-attention-based architecture to capture the speaker style of the prompt voice and improves computational efficiency. Experimental results demonstrate that our proposed model outperforms the baselines in prompt-based audio generation. Myeonghun Jeong, Joun Yeop Lee, Nam Soo Kim |
IEEE Signal Process. Lett. | 4 |
| 2024 | Transfer Learning for Low-Resource, Multi-Lingual, and Zero-Shot Multi-Speaker Text-to-SpeechabstractThough neural text-to-speech (TTS) models show remarkable performance, they still require a large amount ofpaired dataset, which is expensive to collect. The heavy demand for collecting paired datasets makes the TTS models support only a small number of speakers and languages. To address this problem, we introduce a transfer learning framework for multi-lingual, zero-shot multi-speaker, and low-resource TTS. Firstly, we pretrain our model in an unsupervised manner with a multi-lingual multi-speaker speech-only dataset by leveraging the self-supervised speech representations as intermediate linguistic representations. Given this pretrained linguistic information, we then apply a supervised learning technique to the TTS model with a small amount of paired dataset. The pretrained linguistic representations extracted from the large-scale speech-only dataset facilitate phoneme-to-linguistic feature matching, which provides good guidance for supervised learning with a limited amount of labeled data. We evaluate the performance of our proposed model in low-resource, multi-lingual, and zero-shot multi-speaker TTS tasks. The experimental results demonstrate that our proposed method outperforms the baseline in terms of naturalness, intelligibility, and speaker similarity. Myeonghun Jeong, Byoung Jin Choi, Jaesam Yoon, Won Jang, Nam Soo Kim |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Transduce and Speak: Neural Transducer for Text-To-Speech with Semantic Token PredictionabstractWe introduce a text-to-speech (TTS) framework based on a neural transducer. We use discretized semantic tokens acquired from wav2vec 2.0 embeddings, which makes it easy to adopt a neural transducer for the TTS framework enjoying its monotonic alignment constraints. The proposed model first generates aligned semantic tokens using the neural transducer, then synthesizes a speech sample from the semantic tokens using a non-autoregressive (NAR) speech generator. This decoupled framework alleviates the training complexity of TTS and allows each stage to focus on 1) linguistic and alignment modeling and 2) fine-grained acoustic modeling, respectively. Experimental results on the zero-shot adaptive TTS show that the proposed model exceeds the baselines in speech quality and speaker similarity via objective and subjective measures. We also investigate the inference speed and prosody controllability of our proposed model, showing the potential of the neural transducer for TTS frameworks. Myeonghun Jeong, Byoung Jin Choi, Dongjune Lee, Nam Soo Kim |
ASRU | 5 |
| 2023 | Improving Learning Objectives for Speaker Verification from the Perspective of Score ComparisonabstractDeep speaker embedding systems are usually trained with classification-based or end-to-end learning objectives. Popular end-to-end approaches utilize deep metric learning, which can be viewed as a few-shot classification objective. In this paper, we investigate the limit of conventional learning objectives in speaker verification and propose a new learning objective designed from the perspective of similarity scores. The proposed method trains a network by score comparison unbound from the classification, which is more suitable for verification tasks. Experiments conducted with popular speaker embedding networks demonstrate the improvements on the VoxCeleb dataset using the proposed loss. Min Hyun Han, Sung Hwan Mun, Myeonghun Jeong, Sunghwan Ahn, Nam Soo Kim |
ICASSP | 6 |
| 2023 | Multi-Resolution Sequence Aggregation and Model-Agnostic Framework for Time-Series ForecastingabstractIn time-series forecasting, signals such as traffic volume collected in the real world are noisy and irregularly sampled due to sensor malfunctions, so it is difficult to make accurate prediction. To resolve such difficulty, downsampling can be used to reduce noise and allow capturing slow trend of signals. In addition, upsampling can fill the missing data of irregularly sampled signals to catch fine details. Although extracting multi-resolution temporal features such as down or upsampling can improve prediction accuracy, the existing time-series forecasting approaches have used the original and/or downsampled signals only, so they cannot detect fine details of upsampled one. Moreover, these methods merge multi-resolution inputs without carefully concern to chronological order of time-series, which is very important in the time-series. To overcome this challenge, we propose a framework that can fully utilize multi-resolution time-series signals in up, original, and downscale, and sequentially aggregate them, named multi-resolution sequence aggregation and model-agnostic (MAMA) framework. Note that i) MAMA aggregates the multi-resolution signals without breaking its sequential characteristics, whose effectiveness was verified by the experiment results, and ii) it can adopt any existing forecasting algorithms. From experiments with the real-world datasets, it was observed that the prediction accuracy of the well-known forecasting models (i.e., LSTNet, TCN, and Informer) were improved by 11.5% on average when the proposed architecture is used. In ablation study, we showed that a performance improvement of 1.5% was achieved with the help of sequential aggregation module. Juhyun Lyu, Jinseok Yang, Woohyung Lim, Wonbin Ahn, Dongwan Kang, Nam Soo Kim |
ICASSP | 8 |
| 2023 | EM-Network: Oracle Guided Self-distillation for Sequence LearningabstractWe introduce EM-Network, a novel self-distillation approach that effectively leverages target information for supervised sequence-to-sequence (seq2seq) learning. In contrast to conventional methods, it is trained with oracle guidance, which is derived from the target sequence. Since the oracle guidance compactly represents the target-side context that can assist the sequence model in solving the task, the EM-Network achieves a better prediction compared to using only the source input. To allow the sequence model to inherit the promising capability of the EM-Network, we propose a new self-distillation strategy, where the original sequence model can benefit from the knowledge of the EM-Network in a one-stage manner. We conduct comprehensive experiments on two types of seq2seq models: connectionist temporal classification (CTC) for speech recognition and attention-based encoder-decoder (AED) for machine translation. Experimental results demonstrate that the EM-Network significantly advances the current state-of-the-art approaches, improving over the best prior work on speech recognition and establishing state-of-the-art performance on WMT’14 and IWSLT’14. Jiwon Yoon 0002, Sunghwan Ahn, Hyeon Seung Lee, Seok Min Kim, Nam Soo Kim |
ICML | 6 |
| 2023 | Towards Single Integrated Spoofing-aware Speaker Verification Embeddings
Sung Hwan Mun, Hye-Jin Shim, Hemlata Tak, Xin Wang 0037, Xuechen Liu 0001, Md. Sahidullah, Myeonghun Jeong, Min Hyun Han, Massimiliano Todisco, Kong-Aik Lee, Junichi Yamagishi, Nicholas W. D. Evans, Tomi Kinnunen, Nam Soo Kim, Jee-Weon Jung |
INTERSPEECH | 14 |
| 2023 | MCR-Data2vec 2.0: Improving Self-supervised Speech Pre-training via Model-level Consistency Regularization
Jiwon Yoon 0002, Seok Min Kim, Nam Soo Kim |
INTERSPEECH | 3 |
| 2023 | Text Implicates Prosodic Ambiguity: A Corpus for Intention Identification of the Korean Spoken LanguageabstractPhonetic features are indispensable in understanding the spoken language. Especially in Korean, which is wh-in-situ and head-final, the addressee of spoken language sometimes finds it hard to discern the speaker’s original intention if not provided with the sentence prosody. However, acoustic information may not be guaranteed for all spoken language processing, due to the difficulty of managing and computing speech data. This article suggests a corpus that aims to distinguish utterances with ambiguous intention from clear-cut ones, utilizing the prosodic ambiguity of the text input. In detail, the resulting classification system decides whether the given text input is one of fragment, statement, question, command, rhetorical question/command, or indecisive, taking into account the intonation-dependency of the text. Based on an intuitive understanding of the Korean language engaged in the data annotation, we construct a corpus with seven intention categories, train classification systems, and validate the utility of our dataset with quantitative and qualitative analyses. Won-Ik Cho, Nam Soo Kim |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2023 | Oracle Teacher: Leveraging Target Information for Better Knowledge Distillation of CTC ModelsabstractKnowledge distillation (KD), best known as an effective method for model compression, aims at transferring the knowledge of a bigger network (teacher) to a much smaller network (student). Conventional KD methods usually employ the teacher model trained in a supervised manner, where output labels are treated only as targets. Extending this supervised scheme further, we introduce a new type of teacher model for connectionist temporal classification (CTC)-based sequence models, namely Oracle Teacher, that leverages both the source inputs and the output labels as the teacher model's input. Since the Oracle Teacher learns a more accurate CTC alignment by referring to the target information, it can provide the student with more optimal guidance. One potential risk for the proposed approach is a trivial solution that the model's output directly copies the target input. Based on a many-to-one mapping property of the CTC algorithm, we present a training strategy that can effectively prevent the trivial solution and thus enables utilizing both source and target inputs for model training. Extensive experiments are conducted on two sequence learning tasks: speech recognition and scene text recognition. From the experimental results, we empirically show that the proposed model improves the students across these tasks while achieving a considerable speed-up in the teacher model's training time. Jiwon Yoon 0002, Hyung Yong Kim, Hyeon Seung Lee, Sunghwan Ahn, Nam Soo Kim |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | "Feels Like I've Known You Forever": Empathy and Self-Awareness in Human Open-Domain Dialogs
Yoon Kyung Lee, Won-Ik Cho, Seoyeon Bae, Hyunwoo Choi, Jisang Park 0003, Nam Soo Kim, Sowon Hahn |
CogSci | 6 |
| 2022 | Transfer Learning Framework for Low-Resource Text-to-Speech using a Large-Scale Unlabeled Speech CorpusabstractTraining a text-to-speech (TTS) model requires a large scale text labeled speech corpus, which is troublesome to collect. In this paper, we propose a transfer learning framework for TTS that utilizes a large amount of unlabeled speech dataset for pre-training. By leveraging wav2vec2.0 representation, unlabeled speech can highly improve performance, especially in the lack of labeled speech. We also extend the proposed method to zero-shot multi-speaker TTS (ZS-TTS). The experimental results verify the effectiveness of the proposed method in terms of naturalness, intelligibility, and speaker generalization. We highlight that the single speaker TTS model fine-tuned on the only 10 minutes of labeled dataset outperforms the other baselines, and the ZS-TTS model fine-tuned on the only 30 minutes of single speaker dataset can generate the voice of the arbitrary speaker, by pre-training on unlabeled multi-speaker speech corpus. Myeonghun Jeong, Byoung Jin Choi, Sunghwan Ahn, Joun Yeop Lee, Nam Soo Kim |
INTERSPEECH | 6 |
| 2022 | StyleKQC: A Style-Variant Paraphrase Corpus for Korean Questions and CommandsabstractParaphrasing is often performed with less concern for controlled style conversion. Especially for questions and commands, style-variant paraphrasing can be crucial in tone and manner, which also matters with industrial applications such as dialog systems. In this paper, we attack this issue with a corpus construction scheme that simultaneously considers the core content and style of directives, namely intent and formality, for the Korean language. Utilizing manually generated natural language queries on six daily topics, we expand the corpus to formal and informal sentences by human rewriting and transferring. We verify the validity and industrial applicability of our approach by checking the adequate classification and inference performance that fit with conventional fine-tuning approaches, at the same time proposing a supervised formality transfer task. Won-Ik Cho, Sangwhan Moon, Jong In Kim, Seok Min Kim, Nam Soo Kim |
LREC | 5 |
| 2022 | OpenKorPOS: Democratizing Korean Tokenization with Voting-Based Open Corpus AnnotationabstractKorean is a language with complex morphology that uses spaces at larger-than-word boundaries, unlike other East-Asian languages. While morpheme-based text generation can provide significant semantic advantages compared to commonly used character-level approaches, Korean morphological analyzers only provide a sequence of morpheme-level tokens, losing information in the tokenization process. Two crucial issues are the loss of spacing information and subcharacter level morpheme normalization, both of which make the tokenization result challenging to reconstruct the original input string, deterring the application to generative tasks. As this problem originates from the conventional scheme used when creating a POS tagging corpus, we propose an improvement to the existing scheme, which makes it friendlier to generative tasks. On top of that, we suggest a fully-automatic annotation of a corpus by leveraging public analyzers. We vote the surface and POS from the outcome and fill the sequence with the selected morphemes, yielding tokenization with a decent quality that incorporates space information. Our scheme is verified via an evaluation done on an external corpus, and subsequently, it is adapted to Korean Wikipedia to construct an open, permissive resource. We compare morphological analyzer performance trained on our corpus with existing methods, then perform an extrinsic evaluation on a downstream task. Sangwhan Moon, Won-Ik Cho, Hye Joo Han, Naoaki Okazaki, Nam Soo Kim |
LREC | 5 |
| 2022 | Fully Unsupervised Training of Few-Shot Keyword SpottingabstractFor training a few-shot keyword spotting (FS-KWS) model, a large labeled dataset containing massive target keywords has known to be essential to generalize to arbitrary target keywords with only a few enrollment samples. To alleviate the expensive data collection with labeling, in this paper, we propose a novel FS-KWS system trained only on synthetic data. The proposed system is based on metric learning enabling target keywords to be detected using distance metrics. Exploiting the speech synthesis model that generates speech with pseudo phonemes instead of texts, we easily obtain a large collection of multi-view samples with the same semantics. These samples are sufficient for training, considering metric learning does not intrinsically necessitate labeled data. All of the components in our framework do not require any supervision, making our method unsupervised. Experimental results on real datasets show our proposed method is competitive even without any labeled and real datasets. Dongjune Lee, Sung Hwan Mun, Min Hyun Han, Nam Soo Kim |
SLT | 5 |
| 2022 | Frequency and Multi-Scale Selective Kernel Attention for Speaker VerificationabstractThe majority of recent state-of-the-art speaker verification architectures adopt multi-scale processing and frequency-channel attention mechanisms. Convolutional layers of these models typically have a fixed kernel size, e.g., 3 or 5. In this study, we further contribute to this line of research utilising a selective kernel attention (SKA) mechanism. The SKA mechanism allows each convolutional layer to adaptively select the kernel size in a data-driven fashion. It is based on an attention mechanism which exploits both frequency and channel domain. We first apply existing SKA module to our baseline. Then we propose two SKA variants where the first variant is applied in front of the ECAPA-TDNN model and the other is combined with the Res2net backbone block. Through extensive experiments, we demonstrate that our two proposed SKA variants consistently improves the performance and are complementary when tested on three different evaluation protocols. Sung Hwan Mun, Jee-Weon Jung, Min Hyun Han, Nam Soo Kim |
SLT | 4 |
| 2022 | Inter-KD: Intermediate Knowledge Distillation for CTC-Based Automatic Speech RecognitionabstractRecently, the advance in deep learning has brought a considerable improvement in the end-to-end speech recognition field, simplifying the traditional pipeline while producing promising results. Among the end-to-end models, the connectionist temporal classification (CTC)-based model has attracted research interest due to its non-autoregressive nature. However, such CTC models require a heavy computational cost to achieve outstanding performance. To mitigate the computational burden, we propose a simple yet effective knowledge distillation (KD) for the CTC framework, namely Inter-KD, that additionally transfers the teacher's knowledge to the intermediate CTC layers of the student network. From the experimental results on the LibriSpeech, we verify that the Inter-KD shows better achievements compared to the conventional KD methods. Without using any language model (LM) and data augmentation, Inter-KD improves the word error rate (WER) performance from 8.85% to 6.30% on the test-clean. Jiwon Yoon 0002, Beom Jun Woo, Sunghwan Ahn, Hyeon Seung Lee, Nam Soo Kim |
SLT | 5 |
| 2022 | A Controllable Multi-Lingual Multi-Speaker Multi-Style Text-to-Speech Synthesis With Multivariate Information MinimizationabstractIn this letter, we propose a multivariate information minimization method that disentangles three or more latent representations. We show that control factors can be disentangled by minimizing interactive dependency, which can be expressed as a sum of mutual information upper bound terms. Since the upper bound estimate converges from the early training stage, there is little performance degradation due to auxiliary loss. The proposed technique is applied to train a text-to-speech synthesizer with multi-lingual, multi-speaker, and multi-style corpora. Subjective listening tests validate that the proposed method can improve the synthesizer in terms of quality as well as controllability. Sung Jun Cheon, Byoung Jin Choi, Hyeon Seung Lee, Nam Soo Kim |
IEEE Signal Process. Lett. | 5 |
| 2022 | SNAC: Speaker-Normalized Affine Coupling Layer in Flow-Based Architecture for Zero-Shot Multi-Speaker Text-to-SpeechabstractZero-shot multi-speaker text-to-speech (ZSM-TTS) models aim to generate a speech sample with the voice characteristic of an unseen speaker. The main challenge of ZSM-TTS is to increase the overall speaker similarity for unseen speakers. One of the most successful speaker conditioning methods for flow-based multi-speaker text-to-speech (TTS) models is to utilize the functions which predict the scale and bias parameters of the affine coupling layers according to the given speaker embedding vector. In this letter, we improve on the previous speaker conditioning method by introducing a speaker-normalized affine coupling (SNAC) layer which allows for unseen speaker speech synthesis in a zero-shot manner leveraging a normalization-based conditioning technique. The newly designed coupling layer explicitly normalizes the input by the parameters predicted from a speaker embedding vector while training, enabling an inverse process of denormalizing for a new speaker embedding at inference. The proposed conditioning scheme yields the state-of-the-art performance in terms of the speech quality and speaker similarity in a ZSM-TTS setting. Byoung Jin Choi, Myeonghun Jeong, Joun Yeop Lee, Nam Soo Kim |
IEEE Signal Process. Lett. | 4 |
| 2022 | Neurally Optimized Decoder for Low Bitrate Speech CodecabstractRecently, a conventional neural decoder for speech codec has shown promising performance. However, it typically requires some prior knowledge of decoding such as bit allocation or dequantization scheme, which is not a universal solution for many different kinds of speech codecs. In order to address this limitation, we propose a neurally optimized decoder based on a generative model which can directly reconstruct the speech from the bitstream without a prior knowledge. The proposed decoder mainly consists of two components: 1) a dequantization model to group and dequantize related bits from the bitstream and 2) a generative model to restore the speech conditioned on the output of the dequantization model. Through experiments with mixed excitation linear prediction (MELP), Advanced multi-band excitation (AMBE), and SPEEX at around 2.4 kb/s, it is showed that the proposed model showed better performance in most of the objective and subjective evaluation compared to the conventional speech codecs. Hyung Yong Kim, Jiwon Yoon 0002, Won-Ik Cho, Nam Soo Kim |
IEEE Signal Process. Lett. | 4 |
| 2021 | Giving Space to Your Message: Assistive Word Segmentation for the Electronic Typing of Digital MinoritiesabstractFor readability and disambiguation of the written text, appropriate word segmentation is recommended for documentation, and it also holds for the digitized texts. If the language is agglutinative while far from scriptio continua, for instance in the Korean language, the problem becomes more significant. However, some device users these days find it challenging to communicate via key stroking, not only for handicap but also for being unskilled. In this study, we propose a real-time assistive technology that utilizes an automatic word segmentation, designed for digital minorities who are not familiar with electronic typing. We propose a data-driven system trained upon a spoken Korean language corpus with various non-canonical expressions and dialects, guaranteeing the comprehension of contextual information. Through quantitative and qualitative comparison with other text processing toolkits, we show the reliability of the proposed system and its fit with colloquial and non-normalized texts, which fulfills the aim of supportive technology. Won-Ik Cho, Sung Jun Cheon, Woo Hyun Kang, Ji Won Kim, Nam Soo Kim |
Conference on Designing Interactive Systems | 5 |
| 2021 | kosp2e: Korean Speech to English Translation CorpusabstractMost speech-to-text (S2T) translation studies use English speech as a source, which makes it difficult for non-English speakers to take advantage of the S2T technologies. For some languages, this problem was tackled through corpus construction, but the farther linguistically from English or the more under-resourced, this deficiency and underrepresentedness becomes more significant. In this paper, we introduce kosp2e (read as `kospi'), a corpus that allows Korean speech to be translated into English text in an end-to-end manner. We adopt open license speech recognition corpus, translation corpus, and spoken language corpora to make our dataset freely available to the public, and check the performance through the pipeline and training-based approaches. Using pipeline and various end-to-end schemes, we obtain the highest BLEU of 21.3 and 18.0 for each based on the English hypothesis, validating the feasibility of our data. We plan to supplement annotations for other target languages through community contributions in the future. Won-Ik Cho, Seok Min Kim, Hyunchang Cho, Nam Soo Kim |
Interspeech | 4 |
| 2021 | Diff-TTS: A Denoising Diffusion Model for Text-to-SpeechabstractAlthough neural text-to-speech (TTS) models have attracted a lot of attention and succeeded in generating human-like speech, there is still room for improvements to its naturalness and architectural efficiency.In this work, we propose a novel nonautoregressive TTS model, namely Diff-TTS, which achieves highly natural and efficient speech synthesis.Given the text, Diff-TTS exploits a denoising diffusion framework to transform the noise signal into a mel-spectrogram via diffusion time steps.In order to learn the mel-spectrogram distribution conditioned on the text, we present a likelihood-based optimization method for TTS.Furthermore, to boost up the inference speed, we leverage the accelerated sampling method that allows Diff-TTS to generate raw waveforms much faster without significantly degrading perceptual quality.Through experiments, we verified that Diff-TTS generates 28 times faster than the real-time with a single NVIDIA 2080Ti GPU. Myeonghun Jeong, Hyeongju Kim, Sung Jun Cheon, Byoung Jin Choi, Nam Soo Kim |
Interspeech | 5 |
| 2021 | Team02 Text-Independent Speaker Verification System for SdSV Challenge 2021abstractIn this paper, we provide description of our submitted systems to the Short Duration Speaker Verification (SdSV) Challenge 2021 Task 2. The challenge provides a difficult set of cross-language text-independent speaker verification trials. Our submissions employ ResNet-based embedding networks which are trained using various strategies exploiting both in-domain and out-of-domain datasets. The results show that using the recently proposed joint factor embedding (JFE) scheme can enhance the performance by disentangling the language-dependent information from the speaker embedding. However, upon analyzing the speaker embeddings, it was found that there exists a clear discrepancy between the in-domain and out-of-domain datasets. Therefore, among our submitted systems, the best performance was achieved by pre-training the embedding system using out-of-domain dataset and fine-tuning it with only the in-domain data, which resulted in a MinDCF of 0.142716 on the SdSV2021 evaluation set. Woo Hyun Kang, Nam Soo Kim |
Interspeech | 2 |
| 2021 | Expressive Text-to-Speech Using Style TagabstractAs recent text-to-speech (TTS) systems have been rapidly improved in speech quality and generation speed, many researchers now focus on a more challenging issue: expressive TTS. To control speaking styles, existing expressive TTS models use categorical style index or reference speech as style input. In this work, we propose StyleTagging-TTS (ST-TTS), a novel expressive TTS model that utilizes a style tag written in natural language. Using a style-tagged TTS dataset and a pre-trained language model, we modeled the relationship between linguistic embedding and speaking style domain, which enables our model to work even with style tags unseen during training. As style tag is written in natural language, it can control speaking style in a more intuitive, interpretable, and scalable way compared with style index or reference speech. In addition, in terms of model architecture, we propose an efficient non-autoregressive (NAR) TTS architecture with single-stage training. The experimental result shows that ST-TTS outperforms the existing expressive TTS model, Tacotron2-GST in speech quality and expressiveness. Sung Jun Cheon, Byoung Jin Choi, Jong Jin Kim, Nam Soo Kim |
Interspeech | 5 |
| 2021 | Gated Recurrent Context: Softmax-Free Attention for Online Encoder-Decoder Speech RecognitionabstractRecently, attention-based encoder-decoder (AED) models have shown state-of-the-art performance in automatic speech recognition (ASR). As the original AED models with global attentions are not capable of online inference, various online attention schemes have been developed to reduce ASR latency for better user experience. However, a common limitation of the conventional softmax-based online attention approaches is that they introduce an additional hyperparameter related to the length of the attention window, requiring multiple trials of model training for tuning the hyperparameter. In order to deal with this problem, we propose a novel softmax-free attention method and its modified formulation for online attention, which does not need any additional hyperparameter at the training phase. Through a number of ASR experiments, we demonstrate the tradeoff between the latency and performance of the proposed online attention technique can be controlled by merely adjusting a threshold at the test phase. Furthermore, the proposed methods showed competitive performance to the conventional global and online attentions in terms of word-error-rates (WERs). Hyeon Seung Lee, Woo Hyun Kang, Sung Jun Cheon, Hyeongju Kim, Nam Soo Kim |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | TutorNet: Towards Flexible Knowledge Distillation for End-to-End Speech RecognitionabstractIn recent years, there has been a great deal of research in developing end-to-end speech recognition models, which enable simplifying the traditional pipeline and achieving promising results. Despite their remarkable performance improvements, end-to-end models typically require expensive computational cost to show successful performance. To reduce this computational burden, knowledge distillation (KD), which is a popular model compression method, has been used to transfer knowledge from a deep and complex model (teacher) to a shallower and simpler model (student). Previous KD approaches have commonly designed the architecture of the student by reducing the width per layer or the number of layers of the teacher. This structural reduction scheme might limit the flexibility of model selection since the student model structure should be similar to that of the given teacher. To cope with this limitation, we propose a KD method for end-to-end speech recognition, namely TutorNet, that applies KD techniques across different types of neural networks at the hidden representation-level as well as the output-level. For concrete realizations, we firstly apply representation-level knowledge distillation (RKD) during the initialization step, and then apply the softmax-level knowledge distillation (SKD) combined with the original task learning. When the student is trained with RKD, we make use of frame weighting that points out the frames to which the teacher pays more attention. Through a number of experiments, it is verified that TutorNet not only distills the knowledge between networks with different topologies but also significantly contributes to improving the performance of the distilled student. Jiwon Yoon 0002, Hyeon Seung Lee, Hyung Yong Kim, Won-Ik Cho, Nam Soo Kim |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Text Matters but Speech Influences: A Computational Analysis of Syntactic Ambiguity Resolution
Jeonghwa Cho, Woo Hyun Kang, Nam Soo Kim |
CogSci | 3 |
| 2020 | Adaptive Knowledge Distillation Based on EntropyabstractKnowledge distillation (KD) approach is widely used in the deep learning field mainly for model size reduction. KD utilizes soft labels of teacher model, which contain the dark- knowledge that one-hot ground-truth does not have. This knowledge can improve the performance of already saturated student model. In case of multiple-teacher models, generally, the same weighted average (interpolated training) of multiple-teacher's labels is applied to KD training. However, if the knowledge characteristics among teachers are somewhat different, the interpolated training can be at risk of crushing each knowledge characteristics and can also raise noise component. In this paper, we propose an entropy based KD training, which utilizes the teacher model labels with lower entropy at a larger rate among the various teacher models. The proposed method shows a better performance than the conventional KD training scheme in automatic speech recognition. Kisoo Kwon, Hwidong Na, Hoshik Lee, Nam Soo Kim |
ICASSP | 4 |
| 2020 | Robust Front-End for Multi-Channel ASR using Flow-Based Density Estimation
Hyeongju Kim, Hyeon Seung Lee, Woo Hyun Kang, Hyung Yong Kim, Nam Soo Kim |
IJCAI | 5 |
| 2020 | Speech to Text Adaptation: Towards an Efficient Cross-Modal DistillationabstractSpeech is one of the most effective means of communication and is full of information that helps the transmission of utterer's thoughts. However, mainly due to the cumbersome processing of acoustic features, phoneme or word posterior probability has frequently been discarded in understanding the natural language. Thus, some recent spoken language understanding (SLU) modules have utilized end-to-end structures that preserve the uncertainty information. This further reduces the propagation of speech recognition error and guarantees computational efficiency. We claim that in this process, the speech comprehension can benefit from the inference of massive pre-trained language models (LMs). We transfer the knowledge from a concrete Transformer-based text LM to an SLU module which can face a data shortage, based on recent cross-modal distillation methodologies. We demonstrate the validity of our proposal upon the performance on Fluent Speech Command, an English SLU benchmark. Thereby, we experimentally verify our hypothesis that the knowledge could be shared from the top layer of the LM to a fully speech-based module, in which the abstracted speech is expected to meet the semantic representation. Won-Ik Cho, Donghyun Kwak, Jiwon Yoon 0002, Nam Soo Kim |
INTERSPEECH | 4 |
| 2020 | Reformer-TTS: Neural Speech Synthesis with Reformer Network
Hyeong Rae Ihm, Joun Yeop Lee, Byoung Jin Choi, Sung Jun Cheon, Nam Soo Kim |
INTERSPEECH | 5 |
| 2020 | Robust Text-Dependent Speaker Verification via Character-Level Information Preservation for the SdSV Challenge 2020abstractThis paper describes our submission to Task 1 of the Short-duration Speaker Verification (SdSV) challenge 2020. Task 1 is a text-dependent speaker verification task, where both the speaker and phrase are required to be verified. The submitted systems were composed of TDNN-based and ResNet-based front-end architectures, in which the frame-level features were aggregated with various pooling methods (e.g., statistical, self-attentive, ghostVLAD pooling). Although the conventional pooling methods provide embeddings with a sufficient amount of speaker-dependent information, our experiments show that these embeddings often lack phrase-dependent information. To mitigate this problem, we propose a new pooling and score compensation methods that leverage a CTC-based automatic speech recognition (ASR) model for taking the lexical content into account. Both methods showed improvement over the conventional techniques, and the best performance was achieved by fusing all the experimented systems, which showed 0.0785% MinDCF and 2.23% EER on the challenge's evaluation subset. Sung Hwan Mun, Woo Hyun Kang, Min Hyun Han, Nam Soo Kim |
INTERSPEECH | 4 |
| 2020 | Discourse Component to Sentence (DC2S): An Efficient Human-Aided Construction of Paraphrase and Sentence Similarity DatasetabstractAssessing the similarity of sentences and detecting paraphrases is an essential task both in theory and practice, but achieving a reliable dataset requires high resource. In this paper, we propose a discourse component-based paraphrase generation for the directive utterances, which is efficient in terms of human-aided construction and content preservation. All discourse components are expressed in natural language phrases, and the phrases are created considering both speech act and topic so that the controlled construction of the sentence similarity dataset is available. Here, we investigate the validity of our scheme using the Korean language, a language with diverse paraphrasing due to frequent subject drop and scramblings. With 1,000 intent argument phrases and thus generated 10,000 utterances, we make up a sentence similarity dataset of practically sufficient size. It contains five sentence pair types, including paraphrase, and displays a total volume of about 550K. To emphasize the utility of the scheme and dataset, we measure the similarity matching performance via conventional natural language inference models, also suggesting the multi-lingual extensibility. Won-Ik Cho, Jong In Kim, Young Ki Moon, Nam Soo Kim |
LREC | 4 |
| 2020 | SoftFlow: Probabilistic Framework for Normalizing Flow on ManifoldsabstractFlow-based generative models are composed of invertible transformations between two random variables of the same dimension. Therefore, flow-based models cannot be adequately trained if the dimension of the data distribution does not match that of the underlying target distribution. In this paper, we propose SoftFlow, a probabilistic framework for training normalizing flows on manifolds. To sidestep the dimension mismatch problem, SoftFlow estimates a conditional distribution of the perturbed input data instead of learning the data distribution directly. We experimentally show that SoftFlow can capture the innate structure of the manifold data and generate high-quality samples unlike the conventional flow-based models. Furthermore, we apply the proposed framework to 3D point clouds to alleviate the difficulty of forming thin structures for flow-based models. The proposed model for 3D point clouds, namely SoftPointFlow, can estimate the distribution of various shapes more accurately and achieves state-of-the-art performance in point cloud generation. Hyeongju Kim, Hyeon Seung Lee, Woo Hyun Kang, Joun Yeop Lee, Nam Soo Kim |
NeurIPS | 5 |
| 2020 | Pay Attention to Categories: Syntax-Based Sentence Modeling with Metadata Projection Matrix
Won-Ik Cho, Nam Soo Kim |
PACLIC | 2 |
| 2020 | Memory Attention: Robust Alignment Using Gating Mechanism for End-to-End Speech SynthesisabstractRecent end-to-end (e2e) speech synthesis systems usually employ attention techniques to align an input text sequence against a mel-spectrogram sequence. Attention-based e2e approach has shown state-of-the-art performance in speech synthesis. However, generating stable and robust attention alignment to avoid some serious failures such as repeating, missing, and mumbling phones is still an ongoing challenge. In order to mitigate these alignment failures, we propose a novel attention method called memory attention for e2e speech synthesis, which is inspired by the gating mechanism of the long-short term memory (LSTM). Leveraging the sequence modeling power of the gating techniques, memory attention can produce a stable alignment by controlling the amount of content-based and location-based information. For performance evaluation, we compared our proposed memory attention algorithm with various conventional attention techniques in single speaker and emotional speech synthesis scenarios. From the experimental results, we conclude that memory attention can robustly generate various stylish speech. Joun Yeop Lee, Sung Jun Cheon, Byoung Jin Choi, Nam Soo Kim |
IEEE Signal Process. Lett. | 4 |
| 2019 | End-to-End Multi-Channel Speech Enhancement Using Inter-Channel Time-Restricted Attention on Raw Waveform
Hyeon Seung Lee, Hyung Yong Kim, Woo Hyun Kang, Jeunghun Kim, Nam Soo Kim |
INTERSPEECH | 5 |
| 2018 | Acoustic Modeling Using Adversarially Trained Variational Recurrent Neural Network for Speech SynthesisabstractIn this paper, we propose a variational recurrent neural network (VRNN) based method for modeling and generating speech parameter sequences. In recent years. the performance of speech synthesis systems has been improved over conventional techniques thanks to deep learning-based acoustic models. Among the popular deep learning techniques, recurrent neural networks (RNNs) has been successful in modeling time-dependent sequential data efficiently. However, due to the deterministic nature of RNNs prediction, such models do not reflect the full complexity of highly structured data, like natural speech. In this regard, we propose adversarially trained variational recurrent neural network (AdVRNN) which use VRNN to better represent the variability of natural speech for acoustic modeling in speech synthesis. Also, we apply adversarial learning scheme in training AdVRNN to overcome oversmoothing problem. We conducted comparative experiments for the proposed VRNN with the conventional gated recurrent unit which is one of RNNs, for speech synthesis system. It is shown that the proposed AdVRNN based method performed better than the conventional GRU technique. Joun Yeop Lee, Sung Jun Cheon, Byoung Jin Choi, Nam Soo Kim, Eunwoo Song |
INTERSPEECH | 4 |
| 2017 | Integrated DNN-based model adaptation technique for noise-robust speech recognitionabstractSince the introduction of deep neural network (DNN)-based acoustic model, robust automatic speech recognition using DNN are being in research. Especially in model adaptation, the techniques utilizing auxiliary context features is known to be a promising technique. Recently, we proposed a technique which is called two-stage noise-aware training (TSNAT). The key idea of TS-NAT is to let the DNN clarify the relationship among noise estimate, noisy features and phonetic target through clean feature representation. However, although TS-NAT enhances the robustness of the DNN, we cannot be certain whether TS-NAT describes the clean feature representation sufficiently. In this paper, we extend TS-NAT using true noise feature and various DNN training techniques. It has been shown that the proposed technique outperforms the conventional DNN-based techniques on Aurora5-task and mismatched noise conditions. Kang Hyun Lee, Woo Hyun Kang, Tae Gyoon Kang, Nam Soo Kim |
ICASSP | 4 |
| 2017 | Audio Classification Using Class-Specific Learned Descriptors
Sukanya Sonowal, Tushar Sandhan, In Kyu Choi, Nam Soo Kim |
INTERSPEECH | 4 |
| 2017 | Robust Time-Delay Estimation for Acoustic Indoor Localization in Reverberant EnvironmentsabstractIn this letter, we propose a robust approach to time-delay estimation for acoustic indoor localization. Particularly, we focus on the acoustic indoor localization that works in real environments, which have been seldom addressed in previous studies. In the actual environments, it is difficult to estimate the correct time delays for localization due to multipath signals caused by acoustic reverberation. The proposed algorithm minimizes the effect of multipaths and determines the target position based on a novel reliability measure. Experiments conducted in both actual room and simulated environments showed good performance of the proposed technique. Jae Choi, Jeunghun Kim, Nam Soo Kim |
IEEE Signal Process. Lett. | 3 |
| 2016 | NMF-based source separation utilizing prior knowledge on encoding vectorabstractNon-negative matrix factorization (NMF) is an unsupervised technique to represents a nonnegative data matrix with a product of nonnegative basis and encoding matrices. The encoding matrix for the training phase contains information on the pattern of how each basis vector is utilized. The histogram for each row of this matrix corresponding to a specific basis turned out to be sparse, while the level of sparsity varied significantly in each basis. In this paper, the distribution of each component of an encoding vector is modeled as an independent exponential or gamma distribution, and a new objective function with the log-likelihood of the current encoding vector is proposed. Experimental results on audio source separation demonstrate that the utilization of the prior knowledge on the encoding matrix based on sparse statistical models can enhance the source separation performance. Kisoo Kwon, Jong Won Shin, Nam Soo Kim |
ICASSP | 3 |
| 2016 | Two-stage noise aware training using asymmetric deep denoising autoencoderabstractEver since the deep neural network (DNN)-based acoustic model appeared, the recognition performance of automatic speech recognition has been greatly improved. Due to this achievement, various researches on DNN-based technique for noise robustness are also in progress. Among these approaches, the noise-aware training (NAT) technique which aims to improve the inherent robustness of DNN using noise estimates has shown remarkable performance. However, despite the great performance, we cannot be certain whether NAT is an optimal method for sufficiently utilizing the inherent robustness of DNN. In this paper, we propose a novel technique which helps the DNN to address the complex connection between the input and target vectors of NAT smoothly. The proposed method outperformed the conventional NAT in Aurora-5 task. Kang Hyun Lee, Shin Jae Kang, Woo Hyun Kang, Nam Soo Kim |
ICASSP | 4 |
| 2016 | DNN-Based Feature Enhancement Using Joint Training Framework for Robust Multichannel Speech Recognition
Kang Hyun Lee, Tae Gyoon Kang, Woo Hyun Kang, Nam Soo Kim |
INTERSPEECH | 4 |
| 2015 | Reverberation-robust acoustic indoor localization
Jae Choi, Jeunghun Kim, Shin Jae Kang, Nam Soo Kim |
INTERSPEECH | 4 |
| 2015 | Speaker adaptation using relevance vector regression for HMM-based expressive TTS
Doo Hwa Hong, Joun Yeop Lee, Se Young Jang, Nam Soo Kim |
INTERSPEECH | 4 |
| 2015 | Discriminative nonnegative matrix factorization using cross-reconstruction error for source separation
Kisoo Kwon, Jong Won Shin, Hyung Yong Kim, Nam Soo Kim |
INTERSPEECH | 4 |
| 2015 | DNN-based residual echo suppression
Chul Min Lee, Jong Won Shin, Nam Soo Kim |
INTERSPEECH | 3 |
| 2015 | NMF-based Target Source Separation Using Deep Neural NetworkabstractNon-negative matrix factorization (NMF) is one of the most well-known techniques that are applied to separate a desired source from mixture data. In the NMF framework, a collection of data is factorized into a basis matrix and an encoding matrix. The basis matrix for mixture data is usually constructed by augmenting the basis matrices for independent sources. However, target source separation with the concatenated basis matrix turns out to be problematic if there exists some overlap between the subspaces that the bases for the individual sources span. In this letter, we propose a novel approach to improve encoding vector estimation for target signal extraction. Estimating encoding vectors from the mixture data is viewed as a regression problem and a deep neural network (DNN) is used to learn the mapping between the mixture data and the corresponding encoding vectors. To demonstrate the performance of the proposed algorithm, experiments were conducted in the speech enhancement task. The experimental results show that the proposed algorithm outperforms the conventional encoding vector estimation scheme. Tae Gyoon Kang, Kisoo Kwon, Jong Won Shin, Nam Soo Kim |
IEEE Signal Process. Lett. | 4 |
| 2015 | NMF-Based Speech Enhancement Using Bases UpdateabstractThis letter presents a speech enhancement technique combining statistical models and non-negative matrix factorization (NMF) with on-line update of speech and noise bases. The statistical model-based enhancement methods have been known to be less effective to non-stationary noises while the template-based enhancement techniques can deal with them quite well. However, the template-based enhancement techniques usually rely on a priori information. To overcome the shortcomings of both approaches, we propose a novel speech enhancement method that combines the statistical model-based enhancement scheme with the NMF-based gain function. For a better performance in time-varying noise environments, both the speech and noise bases of NMF are adapted simultaneously with the help of the estimated speech presence probability. Experimental results showed that the proposed method outperformed not only the statistical model-based and NMF approaches, but also their combination in various noise environments. Kisoo Kwon, Jong Won Shin, Nam Soo Kim |
IEEE Signal Process. Lett. | 3 |
| 2014 | Parametric multichannel noise reduction algorithm utilizing temporal correlations in reverberant environmentabstractIn this paper, we propose a parametric multichannel noise reduction algorithm utilizing temporal correlations in a noisy and reverberant environment. Under the reverberant condition, the received acoustic signal becomes highly correlated in the time domain and it makes successful noise reduction quite difficult. The proposed parametric noise reduction method takes account of interdependencies between components observed from different frames. Extended speech and noise power spectral density (PSD) matrices are estimated containing additional temporal information, and the parametric multichannel noise reduction filter based on these PSD matrices is applied to the input microphone array signal. According to the experimental results, the proposed algorithm has been found to show better performances compared with the conventional multiplicative filtering technique which considers the current input signals only. Yu Gwang Jin, Jong Won Shin, Chul Min Lee, Soo Hyun Bae, Nam Soo Kim |
ICASSP | 5 |
| 2014 | Reverberation and noise robust feature enhancement using multiple inputsabstractWe propose a novel approach to feature enhancement in multi-channel scenario. Our approach is based on the interacting multiple model (IMM), which was originally developed in single-channel scenario. We extend the single-channel IMM algorithm such that it can handle the multichannel inputs under the Bayesian framework. The multichannel IMM algorithm is capable of tracking time-varying room impulse responses and background noises by updating the relevant parameters in an on-line manner. In various environmental conditions, the performance gain of the proposed method has been confirmed. Shin Jae Kang, Tae Gyoon Kang, Kang Hyun Lee, Kiho Cho, Nam Soo Kim |
ICASSP | 5 |
| 2014 | Speech enhancement combining statistical models and NMF with update of speech and noise basesabstractSpeech enhancement based on statistical models has shown good performance, but the performance degrades when environment noise is highly non-stationary due to the stationary assumption. On the contrary, the template-based enhancement methods are more robust to non-stationary noise, but these are heavily dependent on a priori information present in training data. In order to get over both of the shortcomings, we propose a novel speech enhancement method which combines the statistical model-based enhancement scheme with the template-based enhancement. To reduce a dependency on a priori information, the speech and noise bases are updated simultaneously using the estimated speech presence probability, which is obtained from statistical model-based enhancement. Experimental results showed that the proposed method outperformed not only the statistical model-based and non-negative matrix factorization (NMF) approaches, but also their combination implemented with existing bases update rule in various kinds of noise. Kisoo Kwon, Jong Won Shin, Sukanya Sonowal, In Kyu Choi, Nam Soo Kim |
ICASSP | 5 |
| 2014 | Crossband filtering for stereophonic acoustic echo suppressionabstractIn this paper, we propose a novel stereophonic acoustic echo suppression (SAES) technique based on crossband filtering in the short-time Fourier transform (STFT) domain. The proposed algorithm considers spectral correlations among components in adjacent frequency bins, and estimates the extended power spectral density (PSD) matrices and cross PSD vectors from the signal statistics for more precise echo estimation. In the STFT domain, the echo spectra are estimated by performing the technique without any distinguishable double-talk detector. According to the experimental results, the proposed algorithm has been found to show better performances compared with the conventional SAES method. Chul Min Lee, Jong Won Shin, Yu Gwang Jin, Jeoung Hun Kim, Nam Soo Kim |
ICASSP | 5 |
| 2014 | NMF-based speech enhancement incorporating deep neural network
Tae Gyoon Kang, Kisoo Kwon, Jong Won Shin, Nam Soo Kim |
INTERSPEECH | 4 |
| 2014 | A data-driven approach to speech enhancement using Gaussian process
Sukanya Sonowal, Kisoo Kwon, Nam Soo Kim, Jong Won Shin |
INTERSPEECH | 3 |
| 2014 | Formant enhancement based speech watermarking for tampering detectionabstractUnauthorized tampering in speech signals has brought serious problems when verifying the originality and integrity of speech signals. Digital watermarking can effectively check if the original signals have been tampered by embedding digital data into them. This paper proposes a tampering detection scheme for speech signals based on formant enhancement-based watermarking. Watermarks are embedded as slight enhancement of formant by symmetrically controlling a pair of linear spectral frequencies (LSFs) of corresponding formant. We evaluated the proposed scheme with objective evaluations concerning three criteria that are required for tampering detection scheme: (i) inaudibility to human auditory system, (ii) robustness against meaningful processing, and (iii) fragility against tampering. The evaluation results showed that the proposed scheme could provide satisfactory performance in all the criteria and had the ability to detect tampering in speech signals. Index Terms: tampering detection, speech watermarking, formant enhancement, inaudibility, robustness, fragility Shengbei Wang, Masashi Unoki, Nam Soo Kim |
INTERSPEECH | 3 |
| 2014 | Spectro-Temporal Filtering for Multichannel Speech Enhancement in Short-Time Fourier Transform DomainabstractIn this letter, we propose a spectro-temporal filtering algorithm for multichannel speech enhancement in the short-time Fourier transform (STFT) domain. Compared with the traditional multiplicative filtering technique, the proposed method takes account of interdependencies between components in adjacent frames and frequency bins. For spectro-temporal filtering, speech and noise power spectral density (PSD) matrices are estimated based on an extended formulation utilizing temporal and spectral correlations, and the parametric noise reduction filter based on these PSD matrices is applied to the input microphone array signal. Moreover, multichannel speech presence probabilities are also estimated within a unified framework. A number of experimental results show that the proposed spectro-temporal filtering method improves the performance of multichannel speech enhancement. Yu Gwang Jin, Jong Won Shin, Nam Soo Kim |
IEEE Signal Process. Lett. | 3 |
| 2014 | Stereophonic Acoustic Echo Suppression Incorporating Spectro-Temporal CorrelationsabstractIn this letter, we propose an enhanced stereophonic acoustic echo suppression (SAES) algorithm incorporating spectral and temporal correlations in the short-time Fourier transform (STFT) domain. Unlike traditional stereophonic acoustic echo cancellation, SAES estimates the echo spectra in the STFT domain and uses a Wiener filter to suppress echo without performing any explicit double-talk detection. The proposed approach takes account of interdependencies among components in adjacent time frames and frequency bins, which enables more accurate estimation of the echo signals. Experimental results show that the proposed method yields improved performance compared to that of conventional SAES. Chul Min Lee, Jong Won Shin, Nam Soo Kim |
IEEE Signal Process. Lett. | 3 |
| 2013 | Factored maximum likelihood kernelized regression for HMM-based singing voice synthesis
June Sig Sung, Doo Hwa Hong, Hyun Woo Koo, Nam Soo Kim |
INTERSPEECH | 4 |
| 2013 | Reverberation and Noise Robust Feature Compensation Based on IMMabstractIn this paper, we propose a novel feature compensation approach based on the interacting multiple model (IMM) algorithm specially designed for joint processing of background noise and acoustic reverberation. Our approach to cope with the time-varying environmental parameters is to establish a switching linear dynamic model for the additive and convolutive distortions, such as the background noise and acoustic reverberation, in the log-spectral domain. We construct multiple state space models with the speech corruption process in which the log spectra of clean speech and log frequency response of acoustic reverberation are jointly handled as the state of our interest. The proposed approach shows significant improvements in the Aurora-5 automatic speech recognition (ASR) task which was developed to investigate the influence on the performance of ASR for a hands-free speech input in noisy room environments. Chang Woo Han, Shin Jae Kang, Nam Soo Kim |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Artificial stereo data generation for speech feature mappingabstractFeature mapping technique is widely used to eliminate the mismatch between the training and test conditions of speech recognition. In the feature mapping, a target (mismatched) feature vector sequence is mapped closer to the corresponding reference (matched) feature vector stream. The training of the mapping system is usually carried out based on a set of stereo data which consists of simultaneous recordings obtained in both the reference and target conditions. In this paper, we propose a novel approach to blind parameter estimation which does not require the reference feature vectors. The proposed approach is motivated by the hidden Markov model (HMM)-based speech synthesis algorithm. Chang Woo Han, Tae Gyoon Kang, Shin Jae Kang, June Sig Sung, Nam Soo Kim |
ICASSP | 5 |
| 2012 | Factored MLLR Adaptation Algorithm for HMM-based Expressive TTS
June Sig Sung, Doo Hwa Hong, Hyun Woo Koo, Nam Soo Kim |
INTERSPEECH | 4 |
| 2012 | Speech Feature Mapping Based on Switching Linear Dynamic SystemabstractSignals originated from the same speech source usually appear differently depending on a variety of acoustic effects such as the background noises, linear or nonlinear distortions incurred by the recording devices or reverberations. These acoustical effects result in mismatches between the trained speech recognition models and the input speech. One of the well-known approaches to reduce this mismatch is to map the distorted speech feature to its clean counterpart. The mapping function is usually trained based on a set of stereo data which consists of the simultaneous recordings obtained in both the reference and target conditions. In this paper, we propose the switching linear dynamic system (SLDS) as a useful model for speech feature sequence mapping. In contrast to the conventional vector-to-vector mapping algorithms, SLDS can describe sequence-to-sequence mapping in a systematic way. The proposed approach is applied to robust speech recognition in various environmental conditions and shows a dramatic improvement in recognition performance. Nam Soo Kim, Tae Gyoon Kang, Shin Jae Kang, Chang Woo Han, Doo Hwa Hong |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Switching linear dynamic transducer for stereo data based speech feature mappingabstractThe performance of a speech recognition system may be degraded even without any background noise because of the linear or non-linear distortions incurred by recording devices or reverberations. One of the well-known approaches to reduce this channel distortion is feature mapping which maps the distorted speech feature to its clean counterpart. The feature mapping rule is usually trained based on a set of stereo data which consists of the simultaneous recordings obtained in both the reference and target conditions. In this paper, we propose a novel approach to speech feature sequence mapping based on the switching linear dynamic transducer (SLDT). The proposed algorithm enables us a sequence-to-sequence mapping in a systematic way, instead of the traditional vector-to-vector mapping. The proposed approach is applied to compensate channel distortion in speech recognition and shows improvement in recognition performance. Chang Woo Han, Tae Gyoon Kang, Doo Hwa Hong, Nam Soo Kim, Kiwan Eom |
ICASSP | 4 |
| 2011 | A data-driven residual gain approach for two-stage speech enhancementabstractIn this paper, we propose a novel speech enhancement algorithm based on data-driven residual gain estimation. The system consists of two stages. A noisy input signal is processed at the first stage by a conventional speech enhancement module from which both the enhanced signal and several signal to-noise ratio (SNR)-related parameters are obtained. At the second stage, the residual gain, which is estimated by a data driven method, is applied to the enhanced signal to further adjust it. According to the experimental results, the proposed algorithm has been found to show better performances compared with the conventional speech enhancement technique based on soft decision as well as the data-driven approach using the SNR grid look-up table. Yu Gwang Jin, Chul Min Lee, Kiho Cho, Nam Soo Kim |
ICASSP | 4 |
| 2011 | Decision Tree-Based Clustering with Outlier Detection for HMM-Based Speech Synthesis
Kyung Hwan Oh, June Sig Sung, Doo Hwa Hong, Nam Soo Kim |
INTERSPEECH | 4 |
| 2011 | Factored MLLR Adaptation for Singing Voice GenerationabstractIn our previous study, we proposed factored MLLR (FMLLR) where each MLLR parameter is defined as a function of a control vector. We presented a method to train the FMLLR parameters based on a general framework of the expectationmaximization (EM) algorithm. In this paper, we extend the FMLLR structure from diagonal to unrestricted full matrix with a sophisticated algorithm for the training of relevant parameters. In the experiments on artificial generation of singing voice, we evaluate the performance of the FMLLR technique with two matrix structures and also compare with other approaches to parameter adaptation in HMM-based speech synthesis. Index Terms: Parameter adaptation, MLLR, HRHSMM, factored MLLR June Sig Sung, Doo Hwa Hong, Shin Jae Kang, Nam Soo Kim |
INTERSPEECH | 4 |
| 2011 | Factored MLLR AdaptationabstractOne of the most popular approaches to parameter adaptation in hidden Markov model (HMM) based systems is the maximum likelihood linear regression (MLLR) technique. In this letter, we extend MLLR to factored MLLR (FMLLR) in which the MLLR parameters depend on a continuous-valued control vector. Since it is practically impossible to estimate the MLLR parameters for each control vector separately, we propose a compact parametric form of the MLLR parameters. In the proposed approach, each MLLR parameter is represented as an inner product between a regression vector and transformed control vector. We present an algorithm to train the FMLLR parameters based on a general framework of the expectation-maximization (EM) algorithm. The proposed approach is applied to adapt the HMM parameters obtained from a database of reading-style speech to singing-style voices while treating the pitches and durations extracted from the musical notes as the control vectors. This enables to efficiently construct a singing voice synthesizer with only a small amount of singing data. Nam Soo Kim, June Sig Sung, Doo Hwa Hong |
IEEE Signal Process. Lett. | 1 |
| 2010 | Multichannel noise reduction using low order RTF estimate
Subhojit Chakladar, Nam Soo Kim, Yu Gwang Jin, Tae Gyoon Kang |
INTERSPEECH | 2 |
| 2010 | Phone mismatch penalty matrices for two-stage keyword spotting via multi-pass phone recognizerabstractIn this paper, we propose a novel approach to estimate three types of phone mismatch penalty matrices for two-state keyword spotting. When the output of a phone recognizer is given, text matching with the phone sequences provided by the specified keyword using the proposed phone mismatch penalty matrices is carried out to detect a specific keyword. The penalty matrices which is estimated from the training data through deliberate error generation are accounting for substitution, insertion and deletion errors. In comparative experiments on a Korean continuous speech recognition task, the proposed approach has shown a significant improvement. Chang Woo Han, Shin Jae Kang, Chul Min Lee, Nam Soo Kim |
INTERSPEECH | 4 |
| 2010 | Excitation modeling based on waveform interpolation for HMM-based speech synthesisabstractIt is generally known that a well-designed excitation produces high quality signals in hidden Markov model (HMM)-based speech synthesis systems. This paper proposes a novel tech- niques for generating excitation based on the waveform inter- polation (WI). For modeling WI parameters, we implemented statistical method like principal component analysis (PCA). The parameters of the proposed excitation modeling techniques can be easily combined with the conventional speech synthesis sys- tem under the HMM framework. From a number of experi- ments, the proposed method has been found to generate more naturally sounding speech. Index Terms: HMM-based speech synthesis, Waveform Inter- polation, Principal Component Analysis In this paper, we propose a novel approach to excitation modeling under the waveform interpolation (WI) framework. For parameterizing the excitation generation model, a charac- teristic waveform (CW) is extracted from each frame of LP residual signals. To derive a compact representation of each CW, we apply principal component analysis (PCA) to a collec- tion of the extracted CW's. Once PCA is done, each CW can be compactly approximated as a linear combination of a few PCA basis vectors. The statistical distribution of the linear com- bination coefficients and their dynamics can be efficiently de- scribed by means of HMM's for which the relevant parameters are estimated by following the conventional HMM training pro- cedure. Given a sentence we want to synthesize, the sequence of CW's can be generated from the trained HMM's according to the maximum likelihood (ML) criterion. The WI algorithm enables a smooth transition between adjacent CW's resulting in a more natural excitation signal. The major advantages of the proposed technique are twofold. First, instead of using a fixed set of waveforms such as the impulse train and the ran- dom noise, the proposed method finds CWs which represents the excitation waveforms from the various kinds of modeling in frequency domain. Second, the WI approach lets the excita- tion signal evolve smoothly, which may reduce the audible arti- facts of the synthesized speech. From a number of experiments on speech synthesis, it has been demonstrated that the propose technique enhances the quality of the synthesized speech. June Sig Sung, Doo Hwa Hong, Kyung Hwan Oh, Nam Soo Kim |
INTERSPEECH | 4 |
| 2010 | Voice activity detection based on statistical models and machine learning approaches
Jong Won Shin, Joon-Hyuk Chang, Nam Soo Kim |
Comput. Speech Lang. | 3 |
| 2010 | Robust Data Hiding for MCLT Based Acoustic Data TransmissionabstractAcoustic data transmission enables a short-range wireless communication between the loudspeaker and microphone. A transmitter embeds a data stream into a base audio signal such as music or commercial advertisement and broadcasts the data through the air by playing back the data-embedded sound using a loudspeaker. A receiver picks up the sound signal using microphone and extracts the hidden message. In our previous work, we proposed an acoustic data transmission system which takes advantage of the phase modification of the modulated complex lapped transform (MCLT) coefficients. In this letter, we propose several techniques to realize a more robust communication system in real noisy environment. Kiho Cho, Hwan Sik Yun, Nam Soo Kim |
IEEE Signal Process. Lett. | 3 |
| 2010 | Frequency-Domain Double-Talk Detection Based on the Gaussian Mixture ModelabstractIn this letter, we propose a novel frequency-domain approach to double-talk detection (DTD) based on the Gaussian mixture model (GMM). In contrast to a previous approach based on a simple and heuristic decision rule utilizing time-domain cross-correlations, GMM is applied to a set of feature vectors extracted from the frequency-domain cross-correlation coefficients. Performance of the proposed approach is evaluated through objective tests under various environments, and better results are obtained as compared to the time-domain method. Joon-Hyuk Chang, Nam Soo Kim, Yongserk Kim |
IEEE Signal Process. Lett. | 3 |
| 2010 | Acoustic Data Transmission Based on Modulated Complex Lapped TransformabstractAcoustic data transmission is a technique to embed the data in a sound wave imperceptibly and to detect it at the receiver. This letter proposes a novel acoustic data transmission system designed based on the modulated complex lapped transform (MCLT). In the proposed system, data is embedded in an audio file by modifying the phases of the original MCLT coefficients. The data can be transmitted by playing the embedded audio and extracting it from the received audio. By embedding the data in the MCLT domain, the perceived quality of the resulting audio could be kept almost similar as the original audio. The system can transmit data at several hundreds of bits per second (bps), which is sufficient to deliver some useful short messages. Hwan Sik Yun, Kiho Cho, Nam Soo Kim |
IEEE Signal Process. Lett. | 3 |
| 2009 | DCT based multiple hashing technique for robust audio fingerprintingabstractAudio fingerprinting techniques should successfully perform content-based audio identification even when the audio files are slightly or seriously distorted. In this paper, we present a novel audio fingerprinting technique based on combining fingerprint matching results for multiple hash tables in order to improve the robustness of hashing. Multiple hash tables are built based on the discrete cosine transform (DCT) which is applied to the time sequence of energies in each sub-band. Experimental results show that the recognition errors are significantly reduced compared with Philips Robust Hash (PRH) under various distortions. Kiho Cho, Hwan Sik Yun, Jong Won Shin, Nam Soo Kim |
ICASSP | 5 |
| 2009 | Speech reinforcement based on partial masking effectabstractPerceived quality of the speech signal deteriorates significantly in the presence of ambient noise. In this paper, based on the analysis that the partial masking effect is a main source of the quality degradation when interfering signals are present, we propose a novel approach to enhance the perceived quality of speech signal when the ambient noise cannot be directly controlled by reinforcing it so that it can be heard more clearly. To find a suitable reinforcement rule, the loudness perception model proposed by Moore et al. [1] is adopted with the consideration on the prevention of the hearing damage. Experimental results show that the perceived quality and intelligibility can be enhanced under various noise environments. Jong Won Shin, Yu Gwang Jin, Seung Seop Park, Nam Soo Kim |
ICASSP | 4 |
| 2009 | Global Soft Decision Employing Support Vector Machine For Speech EnhancementabstractIn this letter, we propose a novel speech enhancement technique based on global soft decision incorporating a support vector machine (SVM). Global soft decision in the proposed approach is performed employing the probabilistic outputs of the SVM rather than the conventional Bayes' rule. Actually, global speech absence probability (GSAP) is determined by the sigmoid function based on key parameters estimated by the model-trust minimization algorithm of the SVM output. Improved results are obtained in terms of speech quality measures for various types of noise and at different signal-to-noise ratio (SNR) levels when the proposed SVM is adopted in the global soft decision for speech enhancement. Joon-Hyuk Chang, Q-Haing Jo, Dong Kook Kim, Nam Soo Kim |
IEEE Signal Process. Lett. | 4 |
| 2009 | Audio Fingerprinting Based on Multiple Hashing in DCT DomainabstractAudio fingerprinting techniques aim at successfully performing content-based audio identification even when the audio signals are slightly or seriously distorted. In this letter, we propose a novel audio fingerprinting technique based on multiple hashing. In order to improve the robustness of hashing, multiple hash strings are generated through the discrete cosine transform (DCT) which is applied to the temporal energy sequence in each subband. Experimental results show that the proposed algorithm outperforms the Philips Robust Hash (PRH) algorithm under various distortions. Hwan Sik Yun, Nam Soo Kim |
IEEE Signal Process. Lett. | 3 |
| 2008 | Cepstral domain feature compensation based on diagonal approximationabstractIn this paper, we propose a novel approach to feature compensation performed in the cepstral domain. We apply the linear approximation method in the cepstral domain to simplify the relationship among clean speech, noise and noisy speech. Conventional log-spectral domain feature compensation methods usually assume that each log-spectral coefficient is independent, which is far from real observations. Processing in the cepstral domain has the advantage that the spectral correlation among different frequencies are taken into consideration. By using the diagonal covariance approximation, we can easily modify the conventional log-spectral domain feature compensation technique to fit to the cepstral domain. The proposed approach shows significant improvements in the AURORA2 speech recognition task. Woohyung Lim, Chang Woo Han, Jong Won Shin, Nam Soo Kim |
ICASSP | 4 |
| 2008 | Decision tree based frame mode selection for AMR-WB+
Jong Kyu Kim, Seung Seop Park, Chang Woo Han, Nam Soo Kim |
INTERSPEECH | 4 |
| 2008 | Voice Activity Detection Based on Conditional MAP CriterionabstractIn this letter, we propose a novel approach to voice activity detection (VAD) based on the modified maximum a posteriori (MAP) criterion conditioned on the voice activity decision made in the previous frame. To exploit the inter-frame correlation of voice activity, the probability of the voice presence conditioned on both the observed spectrum and the voice activity decision in the previous frame is employed instead of the conventional strategy that depends only on the current observation. The proposed conditional MAP criterion incorporating temporal correlations leads to two separate thresholds for the likelihood ratio test (LRT) depending on the previous VAD result. Experimental results show that the VAD based on the proposed conditional MAP criterion outperforms the VAD based on the conventional MAP criterion under various noise environments. Jong Won Shin, Hyuk Jin Kwon, Suk Ho Jin, Nam Soo Kim |
IEEE Signal Process. Lett. | 4 |
| 2008 | Analysis and Improvement of Speech/Music Classification for 3GPP2 SMV Based on GMMabstractIn this letter, a novel approach is proposed to improve the performance of speech/music classification for the selectable mode vocoder (SMV) of 3GPP2 using the Gaussian mixture model (GMM). An in-depth analysis of the features and classification method adopted in the conventional SMV is performed. Feature vectors applied to the GMM are then selected from the relevant parameters of the SMV for efficient speech/music classification. The performance of the proposed algorithm is evaluated under various conditions and yields better results compared with the conventional scheme implemented in the SMV. Ji-Hyun Song, Kye-Hwan Lee, Joon-Hyuk Chang, Jong Kyu Kim, Nam Soo Kim |
IEEE Signal Process. Lett. | 5 |
| 2007 | Feature Compensation using More Accurate Statistics of Modeling ErrorabstractIn this paper, we propose a novel approach to feature compensation for robust speech recognition in noisy environments. We analyze the statistics of the modeling error in the log mel magnitude spectrum domain, and model it as a Gaussian distribution. The mean and variance of the distribution are Gaussian functions of the SNR, which enables us to use the SNR dependency of the modeling error efficiently. The proposed feature compensation approach, which is based on the interacting multiple model (IMM) technique, incorporates the statistics of the modeling error and shows significant improvement in the AURORA2 speech recognition task. Woohyung Lim, Jong Kyu Kim, Nam Soo Kim |
ICASSP (4) | 3 |
| 2007 | A statistical model based post-filtering algorithm for residual echo suppression
Seung Yeol Lee, Jong Won Shin, Hwan Sik Yun, Nam Soo Kim |
INTERSPEECH | 4 |
| 2007 | A multiple-model based framework for automatic speech segmentation
Seung Seop Park, Jong Won Shin, Jong Kyu Kim, Nam Soo Kim |
INTERSPEECH | 4 |
| 2007 | Speech reinforcement based on partial specific loudness
Jong Won Shin, Woohyung Lim, June Sig Sung, Nam Soo Kim |
INTERSPEECH | 4 |
| 2007 | Multiple statistical models for soft decision in noisy speech enhancement
Joon-Hyuk Chang, Saeed Gazor, Nam Soo Kim, Sanjit K. Mitra |
Pattern Recognit. | 3 |
| 2007 | Voice activity detection based on a family of parametric distributions
Jong Won Shin, Joon-Hyuk Chang, Nam Soo Kim |
Pattern Recognit. Lett. | 3 |
| 2007 | A Statistical Model-Based Residual Echo SuppressionabstractIn this letter, we propose a novel residual echo suppression (RES) algorithm based on a statistical model constructed in the acoustic echo cancellation framework. In the proposed approach, all the possible near-end and far-end signal conditions are classified into four distinct hypotheses, and the power spectral density estimation is carried out according to the result of hypothesis testing. The distribution of each signal component is characterized by a parametric model, and the conventional likelihood ratio test is performed to make an optimal decision. The experimental results show that the proposed algorithm yields improved performance compared to that of the previous RES technique. Seung Yeol Lee, Nam Soo Kim |
IEEE Signal Process. Lett. | 2 |
| 2007 | Feature Compensation Incorporating Modeling Error StatisticsabstractIn this letter, we propose a novel approach to feature compensation for robust speech recognition in noisy environments. We analyze the error distribution of speech corruption model in the log spectral domain and represent the statistics as functions with respect to the signal-to-noise ratio. The proposed algorithm incorporates modeling error statistics into the interacting multiple model technique and shows a performance improvement over the AURORA2 speech recognition task. Woohyung Lim, Nam Soo Kim |
IEEE Signal Process. Lett. | 2 |
| 2007 | Perceptual Reinforcement of Speech Signal Based on Partial Specific LoudnessabstractIn the presence of background noise, the perceptual loudness of speech signal significantly decreases, resulting in the deterioration of intelligibility and clarity. In this letter, we propose a novel approach to enhance the perceived quality of the speech signal when the additive noise cannot be directly controlled. Instead of controlling the background noise, we propose to reinforce the speech signal so that it can be heard more clearly in noisy environments. To find a suitable reinforcement rule, the loudness perception model proposed by Moore et al. is adopted. Experimental results show that the loudness of the reinforced signal can be maintained at the level almost the same as that of the original noise-free speech, and the proposed algorithm can enhance the perceived speech quality under various noise environments. Jong Won Shin, Nam Soo Kim |
IEEE Signal Process. Lett. | 2 |
| 2007 | On Using Multiple Models for Automatic Speech SegmentationabstractIn this paper, we propose a novel approach to automatic speech segmentation for unit-selection based text-to-speech systems. Instead of using a single automatic segmentation machine (ASM), we make use of multiple independent ASMs to produce a final boundary time-mark. Specifically, given multiple boundary time-marks provided by separate ASMs, we first compensate for the potential ASM-specific context-dependent systematic error (or a bias) of each time-mark and then compute the weighted sum of the bias-removed time-marks, yielding the final time-mark. The bias and weight parameters required for the proposed method are obtained beforehand for each phonetic context (e.g., /p/-/a/) through a training procedure where manual segmentations are utilized as the references. For the training procedure, we first define a cost function in order to quantify the discrepancy between the automatic and manual segmentations (or the error) and then minimize the sum of costs with respect to bias and weight parameters. In case a squared error is used for the cost, the bias parameters are easily obtained by averaging the errors of each phonetic context and then, with the bias parameters fixed, the weight parameters are simultaneously optimized through a gradient projection method which is adopted to overcome a set of constraints imposed on the weight parameter space. A decision tree which clusters all the phonetic contexts is utilized to deal with the unseen phonetic contexts. Our experimental results indicate that the proposed method improves the percentage of boundaries that deviate less than 20 ms with respect to the reference boundary from 95.06% with a HMM-based procedure and 96.85% with a previous multiple-model based procedure to 97.07%. Seung Seop Park, Nam Soo Kim |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Signal modification incorporating perceptual weighting filterabstractAbstract In this paper, an improved preprocessor for low-bit-rate speech coding employing the perceptual weightingfilter is proposed. Speech modification in the proposedapproach is performed according to a criterion whichmakes a compromise between the modification and per-ceptual weighted quantization errors. For this, the per-ceptual weighting filter is expressed in terms of a trans-form domain matrix. The proposed approach is effec-tive in enhancing the speech signal at coder-decoder(CODEC) output through a number of listening tests. 1. Introduction Ingeneral,theperformanceofalow-bit-ratespeechcoderdegrades seriously under the presence of various inter-fering signals such as background noise, acoustic echo,music sounds or interfering speaker’s speech. This phe-nomenon is mainly due to the deviation from the as-sumed speech production model which is used in thecodebook training since a number of codebooks used inthe coder are trained based on a large amount of speechdata and the ranges for parameter search are specified tofit the pure speech signals. One of the successful appli-cations of the unwanted distortion reduction technique tolow-bit-rate coding is the speech enhancement technique[2, 3, 8, 1]. Even though aforementioned enhancementtechniques have been found effective in the presence of astationary background noise, they are not capable of han-dling such interfering signals as the acoustic echoes, mu-sicsoundsorco-talkers’speech. Thisismainlyduetothefactthattheconventionalapproachesadopttheopenloopanalysis which can not take advantage of speech codercharacteristics. An alternative method is the generalized Joon-Hyuk Chang, Woohyung Lim, Nam Soo Kim |
INTERSPEECH | 3 |
| 2006 | Clean speech feature estimation based on soft spectral masking
Woohyung Lim, Nam Soo Kim |
INTERSPEECH | 3 |
| 2006 | Automatic speech segmentation with multiple statistical models
Seung Seop Park, Jong Won Shin, Nam Soo Kim |
INTERSPEECH | 3 |
| 2006 | Speech enhancement based on residual noise shaping
Jong Won Shin, Seung Yeol Lee, Hwan Sik Yun, Nam Soo Kim |
INTERSPEECH | 4 |
| 2006 | Automatic Speech Segmentation Based on Boundary-Type Candidate SelectionabstractIn this letter, we propose a new approach to improve the performance of automatic speech segmentation techniques for concatenative text-to-speech synthesis. Instead of using a single automatic segmentation machine (ASM), we make use of multiple ASMs to draw the final boundary time marks. Given multiple ASMs, the best time mark is chosen among the results provided by the multiple separate ASMs depending on the contextual condition. The experimental results show that our approach dramatically improves the segmentation accuracy Seung Seop Park, Nam Soo Kim |
IEEE Signal Process. Lett. | 2 |
| 2006 | Signal modification for ADPCM based on analysis-by-synthesis frameworkabstractIn this letter, we propose a novel approach to improve the performance of the adaptive differential pulse code modulation (ADPCM) codec by modifying the input signal under the analysis-by-synthesis framework. Modification of the input signal is performed such that the ADPCM codec causes less quantization error. When applied to the ITU-T G.726 ADPCM coder, the proposed algorithm improves the output signal-to-noise ratio up to 2.39 dB. Jong Won Shin, Nam Soo Kim |
IEEE Signal Process. Lett. | 2 |
| 2006 | A new structural approach in system identification with generalized analysis-by-synthesis for robust speech codingabstractIn this paper, we apply a new structural approach to generalized analysis-by-synthesis (GAbS) for system identification as a preprocessor of a low-bit-rate speech coder. In our approach, the coder-decoder (CODEC) system is separately estimated and then applied to modify the current input signal. This is different from that originally proposed where the CODEC system is sequentially estimated and then applied to the next input signal. The proposed estimation scheme is compared to the conventional method in terms of the signal modification approach under the various noise data and in several SNR conditions, and shows better performance. Joon-Hyuk Chang, Nam Soo Kim |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Voice Activity Detection based on Generalized Gamma DistributionabstractWe propose a voice activity detection (VAD) algorithm based on the generalized gamma distribution (G/spl Gamma/D). The distributions of noise spectra and noisy speech spectra, including speech-inactive intervals, are modeled by a set of G/spl Gamma/Ds and applied to the likelihood ratio test (LRT) for VAD. The parameters of G/spl Gamma/D are estimated through an on-line maximum likelihood (ML) estimation procedure where the global speech absence probability (GSAP) is incorporated under a forgetting scheme. Experimental results show that the proposed VAD algorithm, based on G/spl Gamma/D, outperformed the algorithms based on other statistical models. Jong Won Shin, Joon-Hyuk Chang, Hwan Sik Yun, Nam Soo Kim |
ICASSP (1) | 4 |
| 2005 | A new structural preprocessor for low-bit rate speech codingabstractIn this paper, we apply a new structural approach to generalized analysis-by-synthesis (GAbS) for system identification as a preprocessor of a low-bit-rate speech coder. In our approach, the coder-decoder (CODEC) system is separately estimated and then applied to modify the current input signal. This is different from that originally proposed where the CODEC system is sequentially estimated and then applied to the next input signal. The proposed estimation scheme is compared to the conventional method in terms of the signal modification approach under the various noise data and in several SNR conditions, and shows better performance. Joon-Hyuk Chang, Jong Won Shin, Seung Yeol Lee, Nam Soo Kim |
INTERSPEECH | 4 |
| 2005 | Feature compensation based on switching linear dynamic model and soft decisionabstractIn this paper, we present a new approach to feature compensation for robust speech recognition in noisy environments. We employ the switching linear dynamic model (SLDM) as a parametric model for the clean speech distribution, which enables us to utilize temporal correlations in speech signals. Both the background noise and clean speech components are simultaneously estimated by means of the interacting multiple model (IMM) algorithm. Moreover, we combine the SLDM algorithm with the spectral subtraction (SS) approach based on a soft decision. Performance of the presented compensation technique is evaluated through the experiments on AURORA 2 database. Woohyung Lim, Bong Kyoung Kim, Nam Soo Kim |
INTERSPEECH | 3 |
| 2005 | Pitch estimation of speech signal based on adaptive lattice notch filter
Joon-Hyuk Chang, Nam Soo Kim, Sanjit K. Mitra |
Signal Process. | 2 |
| 2005 | A new double-talk detector using echo path estimation
Hae Kyung Jung, Nam Soo Kim, Taejeong Kim |
Speech Commun. | 2 |
| 2005 | Image probability distribution based on generalized gamma functionabstractIn this letter, we propose results of distribution tests that indicate that for many natural images, the statistics of the discrete cosine transform (DCT) coefficients are best approximated by a generalized gamma function (G/spl Gamma/F), which includes the conventional Gaussian, Laplacian, and gamma probability density functions. The major parameter of the G/spl Gamma/F is estimated according to the maximum likelihood (ML) principle. Experimental results on a number of /spl chi//sup 2/ tests indicate that the G/spl Gamma/F can be used effectively for modeling the DCT coefficients compared to the conventional Laplacian and generalized Gaussian function (GGF). Joon-Hyuk Chang, Jong Won Shin, Nam Soo Kim, Sanjit K. Mitra |
IEEE Signal Process. Lett. | 3 |
| 2005 | Feature compensation based on switching linear dynamic modelabstractIn this letter, we propose a novel approach to feature compensation for robust speech recognition in noisy environments. We employ the switching linear dynamic model (SLDM) as a parametric model for the clean speech distribution, which enables us to exploit temporal correlations inherent in speech signals. Both the background noise and clean speech components are simultaneously estimated by means of the interacting multiple model (IMM) algorithm. Nam Soo Kim, Woohyung Lim, Richard M. Stern |
IEEE Signal Process. Lett. | 1 |
| 2005 | An approach to robust unsupervised speaker adaptationabstractIn this letter, we propose an approach to robust unsupervised speaker adaptation. Usually, recognition errors made on the adaptation utterances mislead parameter estimation when a speaker adaptation algorithm is operated in an unsupervised mode. In order to alleviate this problem, we first adapt a Gaussian mixture model (GMM) and then transform the hidden Markov model (HMM) parameters according to the information extracted from GMM adaptation. Nam Soo Kim, Dong Jin Seo, Woohyung Lim |
IEEE Signal Process. Lett. | 1 |
| 2005 | Statistical modeling of speech signals based on generalized gamma distributionabstractIn this letter, we propose a new statistical model, two-sided generalized gamma distribution (G/spl Gamma/D) for an efficient parametric characterization of speech spectra. G/spl Gamma/D forms a generalized class of parametric distributions, including the Gaussian, Laplacian, and Gamma probability density functions (pdfs) as special cases. We also propose a computationally inexpensive online maximum likelihood (ML) parameter estimation algorithm for G/spl Gamma/D. Likelihoods, coefficients of variation (CVs), and Kolmogorov-Smirnov (KS) tests show that G/spl Gamma/D can model the distribution of the real speech signal more accurately than the conventional Gaussian, Laplacian, Gamma, or generalized Gaussian distribution (GGD). Jong Won Shin, Joon-Hyuk Chang, Nam Soo Kim |
IEEE Signal Process. Lett. | 3 |
| 2005 | Rapid online adaptation based on transformation space model evolutionabstractThis paper presents a new approach to online linear regression adaptation of continuous density hidden Markov models based on transformation space model (TSM) evolution. The TSM which characterizes the a priori knowledge of the training speakers associated with maximum likelihood linear regression matrix parameters is effectively described in terms of the latent variable models such as the factor analysis or probabilistic principal component analysis. The TSM provides various sources of information such as the correlation information, the prior distribution, and the prior knowledge of the regression parameters that are very useful for rapid adaptation. The quasi-Bayes estimation algorithm is formulated to incrementally update the hyperparameters of the TSM and regression matrices simultaneously. The proposed TSM evolution is a general framework with batch TSM adaptation as a special case. Experiments on supervised speaker adaptation demonstrate that the proposed approach is more effective compared with the conventional quasi-Bayes linear regression technique when a small amount of adaptation data is available. Dong Kook Kim, Nam Soo Kim |
IEEE Trans. Speech Audio Process. | 2 |
| 2004 | Inner product based-multiband vector quantization for wideband speech coding at 16 kbps
Seung Yeol Lee, Nam Soo Kim, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2004 | Speech probability distribution based on generalized gama distribution
Jong Won Shin, Joon-Hyuk Chang, Nam Soo Kim |
INTERSPEECH | 3 |
| 2004 | Maximum a posteriori adaptation of HMM parameters based on speaker space projection
Dong Kook Kim, Nam Soo Kim |
Speech Commun. | 2 |
| 2004 | Rapid online adaptation using speaker space model evolution
Dong Kook Kim, Nam Soo Kim |
Speech Commun. | 2 |
| 2004 | Feature compensation based on soft decisionabstractIn this letter, we propose a novel approach to feature compensation for robust speech recognition in noisy environments. Our approach combines the interacting multiple model (IMM) and spectral subtraction (SS) techniques based on a soft decision for speech presence. The proposed approach shows 13.56% of average relative improvement compared to the IMM algorithm in the speech recognition experiments performed on the AURORA2 database when clean condition training is applied. Nam Soo Kim |
IEEE Signal Process. Lett. | 1 |
| 2004 | Discriminative training for concatenative speech synthesisabstractIn this letter, we propose an approach to train the cost functions used for unit selection in concatenative speech synthesis. We first view the unit selection as a classification problem, and we apply the discriminative training technique, which is found to be an efficient way to perform parameter estimation in speech recognition. Instead of defining an objective function that accounts for the subjective speech quality, we take the classification error as the objective function to be optimized. The classification error is approximated by a smooth function, and the relevant parameters are updated by means of the gradient descent technique. Nam Soo Kim, Seung Seop Park |
IEEE Signal Process. Lett. | 1 |
| 2004 | Signal modification for robust speech codingabstractUsually, the performance of a low-bit-rate speech coder degrades seriously in the presence of various interfering signals such as the background noise, acoustic echo, co-talkers' speech and other unwanted signals. This comes from the mismatch between the input signal and the assumed speech production model on which the design of the given speech coder is based. In this paper, we present an approach to modify the input signal such that it can be coded more effectively within the generalized analysis-by-synthesis framework. Signal modification in the presented approach is performed according to a criterion which makes a compromise between the modification and coder quantization errors. The coder-decoder (CODEC) characteristic is described in terms of a transfer matrix, and an on-line method using the recursive least square (RLS) technique is proposed to estimate it. Since each part of the speech signal is differently affected by the modification, we also devise an adaptive method based on the signal-to-quantization noise ratio (SQNR). In contrast to the conventional modification techniques, our approach can be implemented as a simple front-end for any analysis-by-synthesis type coders. Nam Soo Kim, Joon-Hyuk Chang |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | Online adaptation using speatransformation space model evolutionabstractThis paper presents a new approach to online speaker adaptation based on transformation space model evolution. This approach extends the previous idea of speaker space model evolution by applying the a priori knowledge of training speakers to the speaker-dependent maximum likelihood linear regression (MLLR) matrix parameters. A quasi-Bayes (QB) estimation algorithm is devised to incrementally update the hyperparameters of the transformation space model and the regression matrices simultaneously. Experiments on supervised speaker adaptation demonstrate that the proposed approach is more effective compared with the conventional quasi-Bayes linear regression (QBLR) technique when a small amount of adaptation data is available. Dong Kook Kim, Woohyung Lim, Nam Soo Kim |
ICASSP (1) | 4 |
| 2003 | Likelihood ratio test with complex laplacian model for voice activity detectionabstractThis paper proposes a voice activity detector (VAD) based on the complex Laplacian model. With the use of a goodness-of-fit (GOF) test, it is discovered that the Laplacian model is more suitable to describe noisy speech distribution than the conventional Gaussian model. The likelihood ratio (LR) based on the Laplacian model is computed and then applied to the VAD operation. According to the experimental results, we can find that the Laplacian statistical model is more suitable for the VAD algorithm compared to the Gaussian model. Joon-Hyuk Chang, Jong Won Shin, Nam Soo Kim |
INTERSPEECH | 3 |
| 2003 | Feature compensation technique for robust speech recognition in noisy environments
Woohyung Lim, Nam Soo Kim |
INTERSPEECH | 4 |
| 2003 | Discriminative weight training for unit-selection based speech synthesis
Seung Seop Park, Chong Kyu Kim, Nam Soo Kim |
INTERSPEECH | 3 |
| 2002 | A new double-talk detector using echo path estimationabstractThis paper presents a new double-talk detector (DTD) based on echo path estimation. The proposed algorithm consists of two stages. In the £rst stage, single-talk periods are distinguished from the double-talks or echo path changes based on the energy level of the echo path estimate. An accurate distinction between the double-talk and echo path change is made in the second stage based on the gradient of the energy level in the estimated echo path. By experiments, it is found that the proposed approach is effective in detecting double-talk periods. Moreover, the required decision delay is shorter than that of the conventional methods. Hae Kyung Jung, Nam Soo Kim, Taejeong Kim |
ICASSP | 2 |
| 2002 | Generalized analysis-by-synthesis based on system identificationabstractIn this paper, we propose an approach to modify the input signal such that it can be coded more effectively within the generalized analysis-by-synthesis framework. Since most of the low-bit-rate speech coders are designed based on the human speech production mechanism, the perceived quality of the speech reconstructed in the decoder degrades seriously if the original input signal deviates from the pure speech. In order to alleviate this problem, we introduce a criterion which compromises the quantization error with the distortion incurred due to the modification. The coder-decoder (CODEC) characteristic is described in terms of a transfer matrix, and it is estimated according to the least squares criterion. In contrast to the conventional modification techniques, our approach can be implemented as a simple front-end for any analysis-by-synthesis type coders. The proposed approach is found effective in reducing audible distortions through a number of listening tests. Nam Soo Kim, Joon-Hyuk Chang |
ICASSP | 1 |
| 2002 | Markov models based on speaker space model evolution
Dong Kook Kim, Nam Soo Kim |
INTERSPEECH | 2 |
| 2002 | Feature domain compensation of nonstationary noise for robust speech recognition
Nam Soo Kim |
Speech Commun. | 1 |
| 2002 | A preprocessor for low-bit-rate speech codingabstractIn this letter, we propose a preprocessor that modifies the input signal such that it can be coded more effectively in a low-bit-rate speech coder. Since most of the low-bit-rate speech coders are designed based on the human speech production mechanism, the perceived quality of the speech reconstructed in the decoder degrades seriously if the original input signal deviates from the pure speech. In order to alleviate this problem, we introduce a criterion that compromises the quantization error with the distortion incurred due to the modification. The coder-decoder characteristic is described in terms of a transfer matrix, and it is estimated according to the least squares criterion. The proposed approach is found to be effective in reducing audible distortions through a number of listening tests. Nam Soo Kim, Joon-Hyuk Chang |
IEEE Signal Process. Lett. | 1 |
| 2001 | EMAP-based speaker adaptation with robust correlation estimationabstractWe propose a method to enhance the performance of the extended maximum a posteriori (EMAP) estimation using the probabilistic principal component analysis (PPCA). PPCA is used to robustly estimate the correlation matrix among separate hidden Markov model (HMM) parameters. The correlation matrix is then applied to the EMAP scheme for speaker adaptation. PPCA is efficient to compute, and shows better performance compared to the method previously used for EMAP. Through various experiments on continuous digit recognition, it is shown that the EMAP approach based on the PPCA gives enhanced performance especially for a small amount of adaptation data. Eugene Jon, Dong Kook Kim, Nam Soo Kim |
ICASSP | 3 |
| 2001 | Robust correlation estimation for EMAP-based speaker adaptationabstractIn this letter, we propose a method to enhance the performance of the extended maximum a posteriori (EMAP) estimation using the probabilistic principal component analysis (PPCA). PPCA is used to robustly estimate the correlation matrix among separate hidden Markov model (HMM) parameters. The correlation matrix is then applied to the EMAP scheme for speaker adaptation. PPCA is efficient to compute and shows better performance compared to the method previously used for EMAP. Through various experiments on continuous digit recognition, it is shown that the EMAP approach based on the PPCA gives enhanced performance, especially for a small amount of adaptation data. Eugene Jon, Dong Kook Kim, Nam Soo Kim |
IEEE Signal Process. Lett. | 3 |
| 2001 | Rapid speaker adaptation using probabilistic principal component analysisabstractIn this letter, we propose a rapid speaker adaptation technique based on the probabilistic principal component analysis (PPCA). The PPCA is employed to obtain the canonical speaker models that provide the a priori knowledge of the training speakers. The proposed approach is conveniently incorporated into the Bayesian adaptation framework, where the parameters are adapted to the new speaker's speech according to the maximum a posteriori (MAP) criterion. Through a number of continuous digit recognition experiments, we can find the effectiveness of the PPCA-based approach compared to the other adaptation approaches with a small amount of adaptation data. Dong Kook Kim, Nam Soo Kim |
IEEE Signal Process. Lett. | 2 |
| 2000 | Speech enhancement: new approaches to soft decision
Joon-Hyuk Chang, Nam Soo Kim |
INTERSPEECH | 2 |
| 2000 | Bayesian speaker adaptation based on probabilistic principal component analysis
Dong Kook Kim, Nam Soo Kim |
INTERSPEECH | 2 |
| 2000 | Spectral enhancement based on global soft decisionabstractIn this letter, we propose a novel speech enhancement technique based on global soft decision. The proposed approach provides a unified framework for such procedures as speech absence probability (SAP) computation, spectral gain modification, and noise spectrum estimation using the same statistical model assumption. Performances of the proposed enhancement algorithm are evaluated by subjective tests under various environments and show better results compared with the IS-127 standard enhancement method. Nam Soo Kim, Joon-Hyuk Chang |
IEEE Signal Process. Lett. | 1 |
| 2000 | Filtering on hidden Markov modelsabstractIn this letter, we propose a novel approach to adapt the hidden Markov model (HMM) parameters when the original feature vector sequences are transformed by a causal finite impulse response (FIR) filter. Our approach enables us to be free from the requirement of retraining the whole recognition parameters when the feature vectors are changed and makes it sufficient to adapt the parameters to the given FIR filter coefficients. Performance of the proposed technique is evaluated and compared to that of the retrained HMM parameters based on the continuous digit recognition experiments. Nam Soo Kim, Dong Kook Kim |
IEEE Signal Process. Lett. | 1 |
| 1999 | Time-varying noise compensation using multiple Kalman filtersabstractThe environmental conditions in which a speech recognition system should be operating are usually nonstationary. We present an approach to compensate for the effects of time-varying noise using a bank of Kalman filters. The presented method is based on the interacting multiple model (IMM) technique well-known in the area of multiple target tracking. Moreover, we propose a way to get fixed-interval smoothed estimates for the environmental parameters. The performances of the proposed approaches are evaluated in the continuous digit recognition experiments where not only the slowly evolving noise but also the rapidly varying noise sources are added to simulate the noisy environments. Nam Soo Kim |
ICASSP | 1 |
| 1999 | A statistical model-based voice activity detectionabstractIn this letter, we develop a robust voice activity detector (VAD) for the application to variable-rate speech coding. The developed VAD employs the decision-directed parameter estimation method for the likelihood ratio test. In addition, we propose an effective hang-over scheme which considers the previous observations by a first-order Markov process modeling of speech occurrences. According to our simulation results, the proposed VAD shows significantly better performances than the G.729B VAD in low signal-to-noise ratio (SNR) and vehicular noise environments. Jongseo Sohn, Nam Soo Kim, Wonyong Sung |
IEEE Signal Process. Lett. | 2 |
| 1998 | Speech recognition in noisy environments using first-order vector Taylor series
Do Yeong Kim, Chong Kwan Un, Nam Soo Kim |
Speech Commun. | 3 |
| 1998 | Statistical linear approximation for environment compensationabstractThe statistical linear approximation (SLA) method is proposed as a novel way to approximate a nonlinear function by a linearized model. In the proposed method, an optimization criterion for approximation is defined in terms of statistical expectation. The SLA is applied to environment compensation where the speech contamination rule appears as a highly nonlinear function of the relevant variables. Nam Soo Kim |
IEEE Signal Process. Lett. | 1 |
| 1998 | Nonstationary environment compensation based on sequential estimationabstractSequential approaches are proposed to compensate for the effects of the nonstationary environment for robust speech recognition. Unlike the batch approaches, the proposed methods derive a different parameter estimate for each time using the sequential expectation maximization (EM) algorithm. Moreover, we also propose the forward-backward estimation scheme as an improvement of the sequential parameter estimation. Nam Soo Kim |
IEEE Signal Process. Lett. | 1 |
| 1998 | IMM-based estimation for slowly evolving environmentsabstractWe propose a new approach to environmental parameter estimation for robust speech recognition in adverse conditions. The proposed method is based on the interacting multiple model (IMM) technique widely used in the area of multiple target tracking. Through a number of continuous digit recognition experiments, we can find the effectiveness of the IMM-based approach in slowly evolving environment conditions. Nam Soo Kim |
IEEE Signal Process. Lett. | 1 |
| 1998 | Deleted strategy for MMI-based HMM trainingabstractWe apply the maximum mutual information (MMI) criterion to discriminative training of hidden Markov model (HMM) parameters. In contrast to the conventional MMI training approach, we adopt the cross-validatory strategy with which the parameters are estimated on a part and assessed on the other parts of the training data. For this purpose, we propose the deleted MMI training method, which performs cross-validatory parameter updating while maintaining the converging behavior of the conventional MMI-based algorithm. The proposed method is compared to the conventional MMI approach in classification of artificial data and in speaker-independent continuous speech recognition, and shows better performance. Nam Soo Kim, Chong Kwan Un |
IEEE Trans. Speech Audio Process. | 1 |
| 1997 | Model-based approach for robust speech recognition in noisy environements with multiple noise sourcesabstractIn this paper, we consider the hidden Markov model(HMM) parameter compensation in noisy environments with multiple noise sources based on the vector Taylor series(VTS) approach. General formulations for multiple environmental variables are derived and systematic expectation-maximization(EM) solutions are presented in maximum likelihood(ML) sense. It is assumed that each noise source is independent and having Gaussian distribution. To evaluate proposed method, we conduct speaker independent isolated word recognition experiments in various noisy environments. Experimental results show that proposed algorithm ahieves significant improvement. Especially, the proposed method is consistently more effective than the parallel model combination(PMC) based on log-normal approximation. Do Yeong Kim, Nam Soo Kim, Chong Kwan Un |
EUROSPEECH | 2 |
| 1997 | Frame-correlated hidden Markov model based on extended logarithmic poolabstractWe present a novel method to incorporate temporal correlations into a speech recognition system based on conventional hidden Markov models (HMMs). The temporal correlations are considered to be useful for recognition because of the fact that the speech features of the present frame are highly informative about the feature characteristics of neighboring frames. In this paper, by treating these correlations in the form of conditional probability distributions (PDs), we propose a new technique for incorporating frame correlations. With the proposed method called the extended logarithmic pool (ELP), we approximate a joint conditional PD by separate conditional PDs associated with respective conditions. We provide a constrained optimization algorithm with which we can find the optimal value for the pooling weights. For practical purposes, we also suggest methods to get robust PD estimates for characterizing frame correlation. In addition, to improve model discriminability, a technique to combine two kinds of PDs through the exponents is introduced. The results in the experiments of speaker-independent continuous speech recognition with the proposed approaches show error reduction up to 20.5% as compared to that with the conventional bigram-constrained (BC) HMM method. Nam Soo Kim, Chong Kwan Un |
IEEE Trans. Speech Audio Process. | 1 |
| 1997 | Statistically reliable deleted interpolationabstractThe output probability distributions (PDs) in each state of a discrete HMM suffer from sparseness, causing inaccurate modeling of probabilistic characteristics of speech features within the state. A desirable solution to the problem arising from insufficient training data is to interpolate a maximum likelihood (ML) estimate of a PD with some other estimates that are, to some extent, able to strengthen the robustness of the PD with respect to unseen data. We propose a statistically reliable deleted interpolation (DI) approach. The DI is an efficient technique for interpolating several probability distribution (PD) estimates, and usually different interpolating weights are used for each predetermined range of PD counts. Our approach attempts to piecewise linearly approximate the interpolating weight curve based on some reasoning concerned with statistical reliability of sample-based estimates. Nam Soo Kim, Chong Kwan Un |
IEEE Trans. Speech Audio Process. | 1 |
| 1995 | On estimating robust probability distribution in HMM-based speech recognitionabstractWe present various methods for estimating a robust output probability distribution (PD) in speech recognition based on the discrete hidden Markov model (HMM). In speech recognition, we encounter the problem of an insufficient amount of training data, which may cause inaccurate modeling of the HMM parameters, especially the output PD's. In this paper, to enhance the robustness of the output PD's with respect to unseen data, we study two approaches: smoothing and tying of the PD's. We introduce a new algorithm to smooth a PD, where a smoothing matrix is estimated by following the strategy of cross-validation as used in deleted interpolation. As for tying, we derive a number of state classes based on a clustering tree which achieves a good compromise between robustness and detail of the tied PD's in specifying speech feature characteristics. In addition to providing an efficient method for constructing the clustering tree, we suggest a measure that accounts for the variation of estimated PD's under various situations. The performances of the proposed methods are evaluated by speaker-independent isolated word recognition experiments and are shown to be better in recognition accuracy than that of the PD's based on the maximum likelihood criterion.> Nam Soo Kim, Chong Kwan Un |
IEEE Trans. Speech Audio Process. | 1 |
| 1990 | Generalized training of hidden Markov model parameters for speech recognition
Nam Soo Kim, Chong Kwan Un |
ICSLP | 1 |