Hong Kook Kim

dblp:94/2710 · DBLP profile ↗
← Back
64ranked-venue papers
19as first author
9since 2021 · last 2026
0000-0002-0105-6693ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 48 · 12 first-author · 3 since 2021Artificial intelligence and machine learning · 30 · 10 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
9 papers
Audio and music processing · 98% Geometric modeling and processing · 2%
Artificial intelligence
2 papers
Speech recognition and synthesis · 100%
Human-computer interaction and pervasive computing
1 paper
Haptics and multimodal interaction · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
speech enhancement
0.632021
TAU-Net: Temporal Activation U-Net Shared With Nonnegative Matrix Factorization for Speech Enhancement in Unseen Noise Environments · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Cepstrum-Domain Model Combination Based on Decomposition of Speech and Noise Using MMSE-LSA for ASR in Noisy Environments · IEEE Trans. Speech Audio Process. 2009
Cepstrum-domain acoustic feature compensation based on decomposition of speech and noise for ASR in noisy environments · IEEE Trans. Speech Audio Process. 2003
Audio and music processing › speech enhancement
deep neural network-based speech enhancement
0.512021
TAU-Net: Temporal Activation U-Net Shared With Nonnegative Matrix Factorization for Speech Enhancement in Unseen Noise Environments · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Audio and music processing
speech recognition
0.252009
Cepstrum-Domain Model Combination Based on Decomposition of Speech and Noise Using MMSE-LSA for ASR in Noisy Environments · IEEE Trans. Speech Audio Process. 2009
Cepstrum-domain acoustic feature compensation based on decomposition of speech and noise for ASR in noisy environments · IEEE Trans. Speech Audio Process. 2003
A bitstream-based front-end for wireless speech recognition on IS-136 communications system · IEEE Trans. Speech Audio Process. 2001
Natural language and speech › Speech recognition and synthesis › microphone array processing
direction-of-arrival estimation
0.212014
Direction-of-arrival based SNR estimation for dual-microphone speech enhancement · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Natural language and speech › Speech recognition and synthesis
signal-to-noise ratio estimation
0.212014
Direction-of-arrival based SNR estimation for dual-microphone speech enhancement · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Natural language and speech › Speech recognition and synthesis
speech enhancement
0.212014
Direction-of-arrival based SNR estimation for dual-microphone speech enhancement · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Audio and music processing › speech recognition
robust speech recognition
0.122009
Cepstrum-Domain Model Combination Based on Decomposition of Speech and Noise Using MMSE-LSA for ASR in Noisy Environments · IEEE Trans. Speech Audio Process. 2009
Cepstrum-domain acoustic feature compensation based on decomposition of speech and noise for ASR in noisy environments · IEEE Trans. Speech Audio Process. 2003
Haptics and multimodal interaction
multimodal interaction
0.112010
Immersive modeling system (IMMS) for personal electronic products using a multi-modal interface · Comput. Aided Des. 2010
Audio and music processing
speech coding
0.132003
Improving the transcoding capability of speech coders · IEEE Trans. Multim. 2003
On approximating line spectral frequencies to LPC cepstral coefficients · IEEE Trans. Speech Audio Process. 2000
Interlacing properties of line spectrum pair frequencies · IEEE Trans. Speech Audio Process. 1999
Audio and music processing
linear prediction
0.021999
Use of spectral autocorrelation in spectral envelope linear prediction for speech recognition · IEEE Trans. Speech Audio Process. 1999
Interlacing properties of line spectrum pair frequencies · IEEE Trans. Speech Audio Process. 1999
Audio and music processing
speech analysis
0.021999
Use of spectral autocorrelation in spectral envelope linear prediction for speech recognition · IEEE Trans. Speech Audio Process. 1999
Interlacing properties of line spectrum pair frequencies · IEEE Trans. Speech Audio Process. 1999
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
robust speech recognition
0.012002
Performance improvement of a bitstream-based front-end for wireless speech recognition in adverse environments · IEEE Trans. Speech Audio Process. 2002
Geometric modeling and processing › computer-aided design
product modeling
0.012010
Immersive modeling system (IMMS) for personal electronic products using a multi-modal interface · Comput. Aided Des. 2010
Audio and music processing › audio feature extraction
cepstral features
0.012000
On approximating line spectral frequencies to LPC cepstral coefficients · IEEE Trans. Speech Audio Process. 2000
Audio and music processing › speech coding
linear predictive coding
0.012000
On approximating line spectral frequencies to LPC cepstral coefficients · IEEE Trans. Speech Audio Process. 2000
Audio and music processing
audio feature extraction
0.011999
Use of spectral autocorrelation in spectral envelope linear prediction for speech recognition · IEEE Trans. Speech Audio Process. 1999
Audio and music processing › speech analysis
formant tracking
0.011999
Interlacing properties of line spectrum pair frequencies · IEEE Trans. Speech Audio Process. 1999
Audio and music processing › speech coding
line-spectral frequencies
0.011999
Interlacing properties of line spectrum pair frequencies · IEEE Trans. Speech Audio Process. 1999

Methods — techniques the papers use, named apart from their topics

u-net · 0.5online dictionary learning · 0.5nonnegative matrix factorization · 0.5multimodal interface · 0.2wiener filter · 0.2log-likelihood ratio test · 0.2decision-directed approach · 0.2beamforming · 0.2MMSE-LSA · 0.1hidden markov model · 0.1parallel model combination · 0.1speech enhancement · 0.1codebook gain re-estimation · 0.1perceptual subjective quality measure · 0.0mel-filterbank cepstrum coefficients · 0.0bitstream mapping · 0.0
YearPublicationVenuePosition
2026 PS-TTS: Phonetic Synchronization in Text-To-Speech for Achieving Natural Automated Dubbing
Changi Hong, Yoonah Song, Hwayoung Park, Chaewoon Bang, Dayeon Ku, Do Hyun Lee, Hong Kook Kim
ICPR (9)7
2026 Revisiting data imbalance in token-based self-supervised learning
Daeyoung Han, Hyung Rok Jung, Tianhong Li, Dina Katabi, Jeany Son, Hong Kook Kim, Moongu Jeon
Neurocomputing6
2024 Optimization for Low-Resource Speaker Adaptation in End-to-End Text-to-Speech
abstract
Fine tuning of an end-to-end text-to-speech (TTS) model is one of the most common methods for optimizing the performance of target speakers. However, the fine-tuning process is time consuming and necessitates the storage of large number of model parameters. Therefore, to reduce the storage of a large number of parameters, the optimization of only partial components of the entire TTS model needs to be considered. This paper proposes an optimization method for low-resource speaker adaptation on a TTS model, i.e., the Variational Inference with Adversarial Learning for End-to-End Text-to-Speech (VITS) model. Particu-larly, the encoder modules of the VITS model, flow network, speaker embedding layer, and projection layer of the text encoder are optimized. The performances of fine-tuned models with varying number of optimized parameters are compared based on speaker embedding cosine similarity (SECS), word error rate (WER), and deep noise suppression mean opinion score (DNSMOS). The results show that the proposed optimization method exhibits reasonable quality compared to the fully fine-tuned model with the tuning of only 9% of the model parameters with SECS of 0.630, WER of 2.4%, and DNSMOS of 3.44
Changi Hong, Jung Hyuk Lee, Moongu Jeon, Hong Kook Kim
CCNC4
2023 Non-Parallel Voice Conversion Using Cycle-Consistent Adversarial Networks with Self-Supervised Representations
abstract
Numerous voice conversion techniques using non-parallel data have been presented. Among these, there are many algorithms related to style transfer. This is because the voice conversion problem can be determined as a style transfer problem, where the linguistic and speaker information can be regarded as domains and styles, respectively. Here, the group of CycleGAN-VC series has considerable achievement, and thus we examine the feasibility of CycleGAN-VC for self-supervised representations. In other words, we incorporate analysis features extracted from wav2vec into the CycleGAN-VC model. Objective experiments showed that the quality of the converted speech is comparable to that of the original speech, and the source speech was successfully transformed into the voice of the target speech while preserving the linguistic information.
Chanjun Chun, Young Han Lee, Geon Woo Lee, Moongu Jeon, Hong Kook Kim
CCNC5
2023 Adversarial Continual Learning to Transfer Self-Supervised Speech Representations for Voice Pathology Detection
abstract
In recent years, voice pathology detection (VPD) has received considerable attention because of the increasing risk of voice problems. Several methods, such as support vector machine and convolutional neural network-based models, achieve good VPD performance. To further improve the performance, we use a self-supervised pretrained model as feature representation instead of explicit speech features. When the pretrained model is fine-tuned for VPD, an overfitting problem occurs due to a domain shift from conversation speech to the VPD task. To mitigate this problem, we propose an adversarial task adaptive pretraining (A-TAPT) approach by incorporating adversarial regularization during the continual learning process. Experiments on VPD using the Saarbrucken Voice Database show that the proposed A-TAPT improves the unweighted average recall (UAR) by an absolute increase of 12.36% and 15.38% compared with SVM and ResNet50, respectively. It is also shown that the proposed A-TAPT achieves a UAR that is 2.77% higher than that of conventional TAPT learning.
Dongkeon Park, Yechan Yu, Dina Katabi, Hong Kook Kim
IEEE Signal Process. Lett.4
2022 Sound Event Detection Using Attention and Aggregation-Based Feature Pyramid Network
abstract
This paper proposes a sound event detection (SED) model using an EfficientNet-B2 and an attention and aggregation-based feature pyramid network (A2-FPN). In particular, the EfficientNet-B2 is first obtained from the pretrained model on the basis of the pretraining, sampling, labeling, and aggregation (PSLA) framework. Then, the A2-FPN module is applied to the outputs of the layers of the EfficientNet-B2 to deal with the different time and frequency resolutions from acoustic features. The aggregated feature map from the A2-FPN module is used as input features to two bidirectional gated recurrent unit layers. Specifically, the proposed A2-FPN-based SED model is trained by the mean-teacher approach to utilize weakly labeled and unlabeled data. Finally, the proposed A2-FPN-based SED model is applied to the detection and classification of acoustic scenes and events (DCASE) 2021 Challenge Task 4. Consequently, it is shown that the polyphonic sound event detection score (PSDS) 1 and 2 of the proposed A2-FPN-based SED model are the higher of 0.03 and 0.172, respectively, than those of the DCASE 2021 Challenge Task 4 baseline.
Ji Won Kim, Geon Woo Lee, Hong Kook Kim, Nam Kyun Kim
APCC3
2022 Auxiliary Loss of Transformer with Residual Connection for End-to-End Speaker Diarization
abstract
End-to-end neural diarization (EEND) with self-attention directly predicts speaker labels from inputs and enables the handling of overlapped speech. Although the EEND outperforms clustering-based speaker diarization (SD), it cannot be further improved by simply increasing the number of encoder blocks because the last encoder block is dominantly supervised compared with lower blocks. This paper proposes a new residual auxiliary EEND (RX-EEND) learning architecture for transformers to enforce the lower encoder blocks to learn more accurately. The auxiliary loss is applied to the output of each encoder block, including the last encoder block. The effect of auxiliary loss on the learning of the encoder blocks can be further increased by adding a residual connection between the encoder blocks of the EEND. Performance evaluation and ablation study reveal that the auxiliary loss in the proposed RX-EEND provides relative reductions in the diarization error rate (DER) by 50.3% and 21.0% on the simulated and CALLHOME (CH) datasets, respectively, compared with self-attentive EEND (SA-EEND). Furthermore, the residual connection used in RX-EEND further relatively reduces the DER by 8.1% for CH dataset.
Yechan Yu, Dongkeon Park, Hong Kook Kim
ICASSP3
2022 DenseBert4Ret: Deep bi-modal for image retrieval
Zafran Khan, Bushra Latif, Joonmo Kim, Hong Kook Kim, Moongu Jeon
Inf. Sci.4
2021 TAU-Net: Temporal Activation U-Net Shared With Nonnegative Matrix Factorization for Speech Enhancement in Unseen Noise Environments
abstract
In this paper, a novel speech enhancement method based on a hybrid machine-learning architecture consisting of U-Net and nonnegative matrix factorization (NMF) is proposed. The proposed method attempts to take advantage of both the accurate separation for known noise environments by U-Net and the adaptation to unseen noises by an NMF with an online dictionary learning technique. To merge the two different architectures, a modified U-Net with a temporal activation layer (TAU-Net) is jointly optimized with NMF models that represent universal speech and noise. The proposed method first estimates the temporal activations from the encoder of the proposed TAU-Net. Then, an NMF with online dictionary learning adjusts the initially given temporal activations to suppress their cross-activations due to unseen noises that are unknown in the training phase of TAU-Net. Finally, clean speech is obtained by adjusting temporal activations to the TAU-Net decoder. The effectiveness of the proposed TAU-Net-based speech enhancement method is evaluated in various unseen noise environments. Consequently, the proposed method achieves a substantial improvement with average signal-to-distortion ratios of 2.32 dB and 5.68 dB, which are higher than those of the baseline methods such asspeech enhancement generative adversarial network (SEGAN) and U-Net, respectively.
Kwang Myung Jeon, Geon Woo Lee, Nam Kyun Kim, Hong Kook Kim
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Two-Stage Polyphonic Sound Event Detection Based on Faster R-CNN-LSTM with Multi-Token Connectionist Temporal Classification
In Young Park, Hong Kook Kim
INTERSPEECH2
2019 Directional Audio Rendering Using a Neural Network Based Personalized HRTF
Geon Woo Lee, Jung Hyuk Lee, Seong Ju Kim, Hong Kook Kim
INTERSPEECH4
2017 Low-Frequency Ultrasonic Communication for Speech Broadcasting in Public Transportation
Kwang Myung Jeon, Nam Kyun Kim, Chan Woong Kwak, Jung Min Moon, Hong Kook Kim
INTERSPEECH5
2016 Local Sparsity Based Online Dictionary Learning for Environment-Adaptive Speech Enhancement with Nonnegative Matrix Factorization
Kwang Myung Jeon, Hong Kook Kim
INTERSPEECH2
2014 Single-channel speech enhancement based on non-negative matrix factorization and online noise adaptation
Kwang Myung Jeon, Chanjun Chun, Woo Kyeong Seong, Hong Kook Kim, Myung Kyu Choi
INTERSPEECH4
2014 Direction-of-arrival based SNR estimation for dual-microphone speech enhancement
abstract
In this paper, we propose a method for estimating target speech by exploring the spatial cues in adverse noise environments. This method is able to reliably estimate the signal-to-noise ratio (SNR) using the phase difference obtained from dual-microphone signals. To this end, spatial cues such as the phase difference are used to estimate the target-to-non-target directional signal ratio (TNR). Based on the estimated TNR, a direction-of-arrival (DOA)-based SNR is then estimated by using a statistical model-based log-likelihood ratio test for the target speech activity decision followed by a decision-directed approach. The estimate is then incorporated into a Wiener filter in order to obtain a spectral-gain attenuator. The perceptual evaluation of speech quality shows that the performance of a dual-microphone speech enhancement system employing the proposed estimation method outperforms single- and dual-microphone speech enhancement systems that use conventional methods such as Wiener filtering, beamforming, or phase-error-based filtering under noise conditions whose SNR ranges from 0 to 20 dB.
Seon Man Kim, Hong Kook Kim
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Target-to-non-target directional ratio estimation based on dual-microphone phase differences for target-directional speech enhancement
Seon Man Kim, Hong Kook Kim
INTERSPEECH2
2012 Adaptation mode control with residual noise estimation for beamformer-based multi-channel speech enhancement
abstract
In this paper, we propose a new adaptation mode controller (AMC) for a generalized sidelobe canceller (GSC) having prior knowledge of the direction-of-arrival (DOA) of a desired speech source. In order to optimize the adaptation mode of a GSC, the residual noise remaining in the GSC output must be employed for adapting the AMC. The residual noise in the GSC output is estimated by using a short-time Fourier transform (STFT)-based Wiener filter, where a priori signal-to-noise ratio (SNR) and a posteriori target-to-non-target-directional signal ratio (TNR) are estimated based on a decision-directed approach and a DOA-based approach, respectively. The estimated residual noise is finally incorporated as a control parameter into the adaptive filters in the AMC. The performance of the proposed AMC is evaluated by measuring the perceptual evaluation of speech quality (PESQ) scores and cepstral distortion in car noise environments with SNRs from 0 to 20 dB. Experimental results show that the proposed AMC performs better than the conventional AMCs.
Seon Man Kim, Hong Kook Kim, Sung Joo Lee
ICASSP2
2012 Dysarthric Speech Recognition Error Correction Using Weighted Finite State Transducers Based on Context-Dependent Pronunciation Variation
Woo Kyeong Seong, Ji Hun Park, Hong Kook Kim
ICCHP (2)3
2011 Hybrid probabilistic adaptation mode controller for generalized sidelobe canceller-based target-directional speech enhancement
abstract
In this paper, we propose a new adaptation mode controller (AMC) for a generalized sidelobe canceller (GSC) having prior knowledge of the direction-of-arrival (DOA) of a desired signal source. To this end, the DOA-based a posteriori target-to-non-target-directional signal ratio (TNR) is first estimated from spatial cues such as phase differences among multi-microphone signals. Next, the estimated TNR is utilized to estimate the target-directional-signal absence probability (TSAP) and presence probability (TSPP), which include global and local terms. The probabilities are then applied to control parameters of adaptive filters via AMC. The performance evaluation of target-directional speech by the proposed approach is carried out by the perceptual evaluation of speech quality (PESQ) scores, noise reduction (NR), and average signal-to-noise ratio (SNR) under car noise conditions with various SNRs from −5 to 20 dB. It is shown from the experiments that the proposed approach provides better results than the conventional AMC.
Seon Man Kim, Hong Kook Kim
ICASSP2
2010 On the use of feature-space MLLR adaptation for non-native speech recognition
abstract
In this paper, we address issues associated with a feature-space maximum likelihood linear regression (fMLLR) adaptation method applied to non-native speech recognition. In particular, fMLLR smoothing is proposed here to compensate for mismatches between adaptation and test data, caused by the various disfluencies of non-native speakers. The proposed fMLLR smoothing is performed with a Viterbi decoding procedure and implemented at two levels: a Gaussian mixture probability density function (mpdf) level and an observation probability density function (opdf) level. The mpdf-level smoothing is performed by comparing the pdf of each Gaussian mixture component of an original speech feature vector with that transformed by the fMLLR. On the other hand, the opdf-level smoothing compares the Gaussian mixture probabilities between the original and its fMLLR transformed feature vectors. It is shown from non-native automatic speech recognition experiments on a Korean-spoken English continuous speech corpus that an ASR system employing the proposed mpdf-level and opdf-level fMLLR smoothing methods can relatively reduce the average word error rate by 30.65% and 29.82%, respectively, when compared to a traditional fMLLR adaptation method.
Yoo Rhee Oh, Hong Kook Kim
ICASSP2
2010 SNR-based mask compensation for computational auditory scene analysis applied to speech recognition in a car environment
Ji Hun Park, Seon Man Kim, Jae Sam Yoon, Hong Kook Kim, Sung Joo Lee
INTERSPEECH4
2010 Immersive modeling system (IMMS) for personal electronic products using a multi-modal interface
Yong-Gu Lee, Hyungjun Park, Woontack Woo, Jeha Ryu, Hong Kook Kim, Sung Wook Baik, Kwang Hee Ko, Han Kyun Choi, Sun-Uk Hwang, Duck Bong Kim, Hyun Soo Kim, Kwan H. Lee
Comput. Aided Des.5
2010 Despeckling of medical ultrasound images using Daubechies complex wavelet transform
Ashish Khare, Manish Khare, Yongyeon Jeong, Hong Kook Kim, Moongu Jeon
Signal Process.4
2010 Entropy coding of compressed feature parameters for distributed speech recognition
Young Han Lee, Hong Kook Kim
Speech Commun.2
2009 MLLR/MAP adaptation using pronunciation variation for non-native speech recognition
abstract
In this paper, we propose an acoustic model adaptation method based on a maximum likelihood linear regression (MLLR) and a maximum a posteriori (MAP) adaptation using pronunciation variations for non-native speech recognition. To this end, we first obtain pronunciation variations using an indirect data-driven approach. Next, we generate two sets of regression classes: one composed of regression classes for all pronunciations and the other of classes for pronunciation variations. The former are referred to as overall regression classes and the latter as pronunciation variation regression classes. Next, we sequentially apply the two adaptations to non-native speech using the overall regression classes, while the acoustic models associated with the pronunciation variations are adapted using the pronunciation variation regression classes. In the final step, both sets of adapted acoustic models are merged. Thus, the resultant acoustic models can cover the characteristics of non-native speakers as well as the pronunciation variations of non-native speech. It is shown from non-native automatic speech recognition experiments for Korean spoken English continuous speech that an ASR system employing the proposed adaptation method can relatively reduce the average word error rate by 9.43% when compared to a traditional MLLR/MAP adaptation method.
Yoo Rhee Oh, Hong Kook Kim
ASRU2
2009 Class-dependent and differential Huffman coding of compressed feature parameters for distributed speech recognition
abstract
In this paper, we propose an entropy coding method for compressing quantized mel-frequency cepstral coefficients (MFCCs) used for distributed speech recognition (DSR). In the European Telecommunication Standards Institute (ETSI) extended DSR standard, MFCCs are compressed with additional parameters such as pitch and voicing class. The entropy of compressed MFCCs in each analysis frame varies according to the voicing class of the frame, thereby enabling the design of different Huffman trees for MFCCs according to voicing class, referred to here as class-dependent Huffman coding. In addition to the voicing class, the correlation in subvector-wise is utilized for Huffman coding, which is called subvector-wise Huffman coding. It is also explored that differential Huffman coding can further enhance a coding gain against class-dependent Huffman coding and subvector-wise Huffman coding. Based on the benefits above, hybrid types of Huffman coding by combining class-dependent and subvector-wise with differential Huffman coding are compared in this paper. Subsequent experiments show that the average bitrate of subvector-wise differential Huffman coding is measured at 33.93 bits/frame, whereas that of a traditional Huffman coding which does not consider voicing class and encodes with a single Huffman coding tree for all the subvectors is at 42.22 bits/frame.
Young Han Lee, Deok Su Kim, Hong Kook Kim
ICASSP3
2009 Acoustic model combination to compensate for residual noise in multi-channel source separation
abstract
In this paper, we propose an acoustic model combination technique for reducing a mismatch in a multi-channel noisy environment. To this end, we first apply a mask-based multi-channel source separation method, typically computational auditory scene analysis (CASA), to separate the speech source from noise. However, a certain degree of noise remains in the separated speech source, especially under low signal-to-noise ratio (SNR) conditions since the estimated mask is not ideal. Thus, the performance of automatic speech recognition (ASR) is limited. To improve ASR performance, the remaining noise can be further compensated in the acoustic model domain under a framework of parallel model combination. In particular, a noise model for PMC is estimated from the noise remained after application of the mask-based source separation, and SNR for PMC is also estimated based on the average of relative magnitude of mask along the utterance. It is shown from the experiments that the proposed acoustic model combination method relatively reduces the word error rate by 52.14% compared to mask-based source separation alone.
Jae Sam Yoon, Ji Hun Park, Hong Kook Kim
ICASSP3
2009 A media-specific FEC based on huffman coding for distributed speech recognition
Young Han Lee, Hong Kook Kim
INTERSPEECH2
2009 Cepstrum-Domain Model Combination Based on Decomposition of Speech and Noise Using MMSE-LSA for ASR in Noisy Environments
abstract
This paper presents an efficient method for combining models of speech and noise for robust speech recognition applications in noisy environments. This method decomposes the cepstrum domain representation of noise-corrupted speech into clean speech cepstrum and background noise cepstrum components using a minimum mean squared error-log spectral amplitude (MMSE-LSA) criterion. Speech recognition is then performed on noisy cepstrum domain observations using a model that is formed by parallel combination of cepstrum domain clean speech distributions and background noise distributions estimated using this MMSE-LSA based noise decomposition. This method is far more efficient than other parallel model combination (PMC) procedures because model combination is performed directly in the cepstrum domain rather than in the linear spectral domain. Whereas background noise model estimation is addressed as a separate issue in existing PMC procedures, this method explicitly incorporates a mechanism to continually update background noise models and signal-to-noise ratio (SNR) estimates over time. The performance of the proposed cepstrum-domain model combination method is compared with a well known implementation of PMC which uses a log-normal approximation when combining speech and background noise model means and variances on a connected digit string recognition task which is subjected to mismatched channel and environment conditions. As a result, it is shown that the proposed model combination technique gives a word error rate that is comparable to PMC when background noise information and SNR are known prior to estimation. The paper will also present the results of experiments where a combination of cepstrum-domain feature compensation and model combination are applied to this task.
Hong Kook Kim, Richard C. Rose
IEEE Trans. Speech Audio Process.1
2008 Acoustic and pronunciation model adaptation for context-independent and context-dependent pronunciation variability of non-native speech
abstract
In this paper, we propose an acoustic and pronunciation model adaptation method for context-independent (CI) and context-dependent (CD) pronunciation variability to improve the performance of a non-native automatic speech recognition (ASR) system. The proposed adaptation method is performed in three steps. First, we perform phone recognition to obtain an n-best list of phoneme sequences and derive pronunciation variant rules by using a decision tree. Second, the pronunciation variant rules are decomposed into CI and CD pronunciation variation on the basis of context dependency. That is, some pronunciation variant rules that are dedicated to the specific phoneme sequences is classified into CI pronunciation variation, but others are classified into CD one. It is assumed here that CI and CD pronunciation variabilities are invoked by a different pronunciation space from the mother tongue of a non-native speaker and the coarticulation effects in a context, respectively. Third, the acoustic model adaptation is performed in a state-tying step for the CI pronunciation variability from an indirect data-driven method. In addition, the pronunciation model adaptation is completed by constructing a multiple pronunciation dictionary using the CD pronunciation variability. It is shown from the continuous Korean-English ASR experiments that the proposed method can reduce the average word error rate (WER) by 16.02% when compared with the baseline ASR system that is trained by native speech. Moreover, an ASR system using the proposed method provides average WER reductions of 8.95% and 3.67% when compared to the only acoustic model adaptation and the only pronunciation model adaptation, respectively.
Yoo Rhee Oh, Hong Kook Kim
ICASSP3
2008 Mask estimation incorporating time-frequency trajectories for a CASA-based ASR front-end
Ji Hun Park, Jae Sam Yoon, Hong Kook Kim
INTERSPEECH3
2008 Gammatone-domain model combination for consonant recognition in noisy environments
Jae Sam Yoon, Ji Hun Park, Hong Kook Kim
INTERSPEECH3
2008 Cepstral domain interpretations of line spectral frequencies
Hong Kook Kim, Seung Ho Choi
Signal Process.1
2007 Non-native pronunciation variation modeling using an indirect data driven method
abstract
In this paper, we propose a pronunciation variation modeling method for improving the performance of a non-native automatic speech recognition (ASR) system that does not degrade the performance of a native ASR system. The proposed method is based on an indirect data-driven approach, where pronunciation variability is investigated from the training speech data, and variant rules are subsequently derived and applied to compensate for variability in the ASR pronunciation dictionary. To this end, native utterances are first recognized by using a phoneme recognizer, and then the variant phoneme patterns of native speech are obtained by aligning the recognized and reference phonetic sequences. The reference sequences are transcribed by using each of canonical, knowledge-based, and hand-labeled methods. Similar to non-native speech, the variant phoneme patterns of non-native speech can also be obtained by recognizing non-native utterances and comparing the recognized phoneme sequences and reference phonetic transcriptions. Finally, variant rules are derived from native and non-native variant phoneme patterns using decision trees and applied to the adaptation of a dictionary for non-native and native ASR systems. In this paper, Korean spoken by Chinese native speakers is considered as the non-native speech. It is shown from non-native ASR experiments that an ASR system using the dictionary constructed by the proposed pronunciation variation modeling method can relatively reduce the average word error rate (WER) by 18.5% when compared to the baseline ASR system using a canonical transcribed dictionary. In addition, the WER of a native ASR system using the proposed dictionary is also relatively reduced by 1.1%, as compared to the baseline native ASR system with a canonical constructed dictionary.
Yoo Rhee Oh, Hong Kook Kim
ASRU3
2007 Acoustic model adaptation based on pronunciation variability analysis for non-native speech recognition
Yoo Rhee Oh, Jae Sam Yoon, Hong Kook Kim
Speech Commun.3
2006 A Highly Adaptive Acoustic Echo Cancellation Solution for VoIP Conferencing Systems
abstract
Teleconferencing solutions such as VoIP conferenc- ing systems employ acoustic echo cancellers to reduce echoes that are resulted from the coupling between loudspeaker and microphone. While echo cancellation in itself is a very delicate and complex task, requiring an exact modeling of the echo generating environment and a great amount of fine tuning in real-time, the task becomes even more complex in the case of a multi-party conferencing scenario. We propose a highly adaptive and efficient acoustic echo cancellation solution for multi-party VoIP conferencing systems. The proposed solution is based on a novel VoIP conferencing framework for high quality echo cancellation in a multi-party VoIP conferencing environment. All the building blocks of the proposed solution are completely specified and elaborated using a model VoIP conference call scenario. The proposed acoustic echo cancellation solution is simulated, tested, and verified using MATLAB for its high efficiency and quick adaptability. Results show that our proposed solution offers a unique conferencing experience for multi-party VoIP conferencing systems by delivering extremely good voice quality. I. INTRODUCTION In recent years the need for echo cancellation has appeared in a new area, namely hands-free telephony for which the demand is expected to increase enormously in the years to come. In hands-free telephony, there is an additional delay for the signals to travel via the telephone network. This can be quite long, depending on the telephone setup and connection, and the echoes may be noticeable. Therefore, in some way the echoes from the loudspeaker output that are picked up by the microphone must be removed from the microphone signal. A. Echo Cancellation in VoIP Echo cancellation is critical in achieving high quality voice transmissions over packet networks, which typically face transmission delays in excess of 30 to 40 ms. Such a long delay makes echo readily apparent to listeners, and must be eliminated for a viable telephony service. In this paper, we propose a highly adaptive and efficient acoustic echo cancellation solution for multi-party VoIP con- ferencing systems. The proposed solution is based on a novel VoIP conferencing framework for high quality echo cancel- lation in a multi-party VoIP conferencing environment. The proposed acoustic echo cancellation solution is implemented and its performance is evaluated using extensive simulations for its high efficiency and quick adaptability.
Umar Iqbal Choudhry, Jongwon Kim 0001, Hong Kook Kim
AICCSA3
2006 Acoustic Model Adaptation Based on Pronunciation Variability Analysis for Non-Native Speech Recognition
abstract
In this paper, we investigate the pronunciation variability between native and non-native speakers and propose an acoustic model adaptation method based on the variability analysis in order to improve the performance of a non-native speech recognition system. The proposed acoustic model adaptation is performed in two steps. First, we construct baseline acoustic models from native speech, and perform phone recognition by using the baseline acoustic models to identify most informative variant phonetic units from native to non-native. Next, the acoustic model corresponding to each informative variant phonetic unit is adapted so that the state tying of the acoustic model for non-native speech reflects such a phonetic variability. For further improvement, the traditional acoustic model adaptation such as MLLR or MAP could be applied on the system that is adapted with the proposed method. In this work, we select English as a target language and non-native speakers are all Korean. It is shown from the continuous Korean-English speech recognition experiments that the proposed method can achieve the average word error rate reduction by 12.75% when compared with the speech recognition system with the baseline acoustic models trained by native speech. Moreover, the reduction of 57.12% in the average word error rate is obtained by applying MLLR or MAP adaptation to the adapted acoustic models by the proposed method.
Yoo Rhee Oh, Jae Sam Yoon, Hong Kook Kim
ICASSP (1)3
2005 Error Prediction in Spoken Dialog: From Signal-to-Noise Ratio to Semantic Confidence Scores
abstract
Spoken dialog systems aim to interpret the meanings of users' utterances and respond to them accordingly. The users' utterances are first recognized by an automatic speech recognizer (ASR) and the intents of the users are extracted by the spoken language understanding (SLU) unit. Both ASR and SLU are noisy and in general their noise statistics are not correlated. Our goal is to exploit the signal-to-noise information and ASR lattice-based and semantic confidence scores for SLU error prediction and prevention of these by rejecting erroneous utterances, or asking confirmation questions. In our experiments, we have shown up to 80% relative decrease in the error rate of the accepted utterances collected using the AT&T How May I Help You/spl trade/ spoken dialog system used for customer care.
Dilek Hakkani-Tür, Gökhan Tür, Giuseppe Riccardi, Hong Kook Kim
ICASSP (1)4
2005 A MFCC-based CELP speech coder for server-based speech recognition in network environments
Gil Ho Lee, Jae Sam Yoon, Hong Kook Kim
INTERSPEECH3
2004 Why speech recognizers make errors ? a robustness view
abstract
The performance of large vocabulary speech recognizers often varies depending on the input speech and the quality of the trained models. The particular attributes that cause recognition errors are a research area that has not been well studied. This paper addresses this issue from a robustness perspective using a large amount of field data collected from natural language dialog services. In particular, we present a method for tracking time-varying or nonstationary extraneous events, such as music, background noise, etc. We show that this measure is a better predictor of recognition errors than a standard measure of stationary signal-to-noise ratio (SNR). Combining the two measures provides a data selection algorithm for detecting problematic speech.
Hong Kook Kim, Mazin G. Rahim
INTERSPEECH1
2004 Robust speech recognition in client-server scenarios
abstract
This paper addresses issues that are specific to the implementation of automatic speech recognition (ASR) applications and services in client-server scenarios. It is assumed in all of these scenarios that functionality in a human-machine dialog system is distributed between mobile client devices and network based multi-user media and application servers. It is argued that, while there has already been a great deal of research addressing issues relating to the communications channels associated with these scenarios, there are many additional problems that have received relatively little attention. These include issues of how environmental and speaker robustness algorithms are implemented in mobile domains and how multiple ASR channels can be implemented more efficiently in multi-user deployments. Preliminary results are summarized showing the effect of user specific unsupervised adaptation and normalization algorithms on ASR performance in mobile domains. Results are also presented demonstrating the efficiencies that are obtainable from using intelligent algorithms for assigning ASR decoders to computation servers in multi-user deployments.
Richard C. Rose, Hong Kook Kim
INTERSPEECH2
2003 Cepstrum-domain acoustic feature compensation based on decomposition of speech and noise for ASR in noisy environments
abstract
This paper presents a set of acoustic feature pre-processing techniques that are applied to improving automatic speech recognition (ASR) performance on noisy speech recognition tasks. The principal contribution of this paper is an approach for cepstrum-domain feature compensation in ASR which is motivated by techniques for decomposing speech and noise that were originally developed for noisy speech enhancement. This approach is applied in combination with other feature compensation algorithms to compensating ASR features obtained from a mel-filterbank cepstrum coefficient front-end. Performance comparisons are made with respect to the application of the minimum mean squared error log spectral amplitude (MMSE-LSA) estimator based speech enhancement algorithm prior to feature analysis. An experimental study is presented where the feature compensation approaches described in the paper are found to greatly reduce ASR word error rate compared to uncompensated features under environmental and channel mismatched conditions.
Hong Kook Kim, Richard C. Rose
IEEE Trans. Speech Audio Process.1
2003 Improving the transcoding capability of speech coders
abstract
With the trend of merging various communication networks, a need arises to provide transcoding between different speech coding formats. Presently this means a cross tandem between the two coders in each case. This results in both quality loss and extra delay. A possible alternative is using a bitstream mapping approach that directly converts parameter values. For several standard coders having a similar coding structure, it should be possible to generate comparable or better quality without adding much delay or complexity. This paper proposes a bitstream mapping method between ITU-T Recommendation G.729 and TIA IS-641. Informal listening tests and the perceptual subjective quality measure (PSQM) scores show that the proposed method has better quality than the cross tandem method, while it has at least 5 ms less delay and six times less computation.
Hong-Goo Kang, Hong Kook Kim, Richard V. Cox
IEEE Trans. Multim.2
2002 A phase generation method for speech reconstruction from spectral envelope and pitch intervals
abstract
In this paper, we propose a new speech reconstruction method from spectral envelope and pitch intervals, which is applicable to the network side of a distributed speech recognition system as a play-back function. The spectral envelope of speech is represented as a set of mel-frequency cepstral coefficients that is a well-known recognition parameter. First, a sinusoidal synthesis with a zero-phase model is used to obtain a pitch-based waveform. To enhance the naturalness of the speech we replace the zero phase information with pre-stored linear and random codebooks. The ultimate phase information is determined depending on the energy ratio between linear and random components. Unlike the classic low bit-rate speech coding, however, the energy ratio is estimated in the decoding stage from a time-frequency filter applied to the pitch-based synthesized signal. Thus, the phase information is not a feature parameter from the encoder side. The proposed phase generation method uses the knowledge that pitch variation is a main cause of the mixed characteristics in speech signals. An informal listening test verifies that the quality of the proposed method is much better than that of the synthetic quality.
Hong-Goo Kang, Hong Kook Kim
ICASSP2
2002 Cepstrum-domain model combination based on decomposition of speech and noise for noisy speech recognition
abstract
We propose a cepstrum-domain model combination method for automatic speech recognition in noisy environments. The distinguishing aspect of the method is that noise-corrupted speech is decomposed into clean speech and noise components directly in the cepstrum domain without having to transform to the linear spectrum domain as is necessary for many existing model combination approaches. This is accomplished by exploiting the properties of the minimum mean squared error-log spectral amplitude (MMSE-LSA) based speech enhancement algorithm. As a result, a clean speech hidden Markov model (HMM) is easily compensated for a noise-corrupted domain by adding the means and covariance matrices of the clean speech HMM and those of an estimated noise model. The complexity of the proposed model combination procedure is significantly reduced with respect to conventional parallel model combination. The procedure was applied to a noisy connected digit recognition task. A 40% reduction in word error rate was achieved when it was combined with acoustic feature compensation techniques under mismatched environmental and channel conditions.
Hong Kook Kim, Richard C. Rose
ICASSP1
2002 Algorithms for distributed speech recognition in a noisy automobile environment
Hong Kook Kim, Richard C. Rose
INTERSPEECH1
2002 An adaptive short-term postfilter based on pseudo-cepstral representation of line spectral frequencies
Hong Kook Kim, Hong-Goo Kang
Speech Commun.1
2002 Performance improvement of a bitstream-based front-end for wireless speech recognition in adverse environments
abstract
We propose a feature enhancement algorithm for wireless speech recognition in adverse acoustic environments. A speech recognition system is realized at the network side of a wireless communications system and feature parameters are extracted directly from the bitstream of the speech coder employed in the system, where the feature parameters are composed of spectral envelope information and coder-specific information. The coder-specific information is apt to be affected by environmental noise because the speech coder fails to generate high quality speech in noisy environments. We first found that enhancing noisy speech prior to speech coding improves the recognizer's performance. However, our aim was to develop a robust front-end operating at the network side of a wireless communications system without regard to whether speech enhancement was applied at the sender side. We investigated the effect of a speech enhancement algorithm on the bitstream-based feature parameters. Consequently, a feature enhancement algorithm is proposed which incorporates feature parameters obtained from the decoded speech and a noise suppressed version of the decoded speech. The coder-specific information can also be improved by re-estimating the codebook gains and residual energy from the enhanced residual signal. HMM-based connected digit recognition experiments show that the proposed feature enhancement algorithm significantly improves recognition performance at low signal-to-noise ratio (SNR) without causing poorer performance at high SNR. From large vocabulary speech recognition experiments with far-field microphone speech signals recorded in an office environment, we show that the feature enhancement algorithm greatly improves word recognition accuracy.
Hong Kook Kim, Richard V. Cox, Richard C. Rose
IEEE Trans. Speech Audio Process.1
2001 Feature enhancement for a bitstream-based front-end in wireless speech recognition
abstract
We propose a feature enhancement algorithm for wireless speech recognition in adverse acoustic environments. A speech recognition system is realized at the receiver side of a wireless communications system and feature parameters are extracted directly from the bitstream of the speech coder employed in the system. The feature parameters are composed of spectral envelope and coder-specific information. The proposed feature enhancement algorithm incorporates feature parameters obtained from the decoded speech and an enhanced version into the bitstream-based feature parameters. Moreover, the coder-specific parameters are improved by reestimating the codebook gains and residual energy from the enhanced residual signal. HMM-based connected digit recognition experiments show that the proposed feature enhancement algorithm significantly improves recognition accuracy at low SNR without causing poorer performance at high SNR.
Hong Kook Kim, Richard V. Cox
ICASSP1
2001 Acoustic feature compensation based on decomposition of speech and noise for ASR in noisy environments
abstract
This paper presents a set of acoustic feature pre–processing techniques that are applied to improving automatic speech recognition (ASR) performance on the Aurora 2 noisy speech recognition task. The principal contribution of this paper is an approach for cepstrum domain feature compensation in ASR which is motivated by techniques for decomposing speech and noise that were originally developed for noisy speech enhancement. This approach is applied in combination with other feature compensation algorithms to compensating ASR features obtained from a mel–filterbank cepstrum coefficient (MFCC) front–end. Performance comparisons are made with respect to the application of the minimum mean squared error log spectral amplitude estimator (MMSE–LSA) based speech enhancement algorithm prior to feature analysis. An experimental study is presented where the feature compensation approaches described in the paper are found to reduce ASR word error rate by as much as 31% relative to uncompensated features under simulated environmental and channel mismatched conditions.
Hong Kook Kim, Richard C. Rose, Hong-Goo Kang
INTERSPEECH1
2001 Robust speech recognition techniques applied to a speech in noise task
abstract
This paper describes the design and evaluation of an automatic speech recognition (ASR) system on the Naval Research Laboratory Speech In Noise (SPINE) speech corpus. This corpus represents a task which involves human-human interaction on a constrained problem solving scenario under six di erent simulated noisy environments. Acoustic and language modeling were performed using a small dataset taken entirely from a subset of the acoustic environments. Speech recognition was performed on continuous conversations by detecting speech utterances, performing acoustic feature analysis and normalization, and adapting HMMmodels in multiple passes over each conversation-side. The ASR word accuracy (WAC) ranged from 77 percent in an o ce environment to 61 percent in conditions that include signi cant levels of background speech and noise.
Richard C. Rose, Hong Kook Kim, Donald Hindle
INTERSPEECH2
2001 A new distortion measure for spectral quantization based on the LSF intermodel interlacing property
Mi Suk Lee, Hong Kook Kim, Hwang Soo Lee
Speech Commun.2
2001 A frame erasure concealment algorithm based on gain parameter re-estimation for CELP coders
abstract
In this paper, we propose a frame erasure concealment algorithm based on reestimating gain parameters for a code-excited linear prediction (CELP) coder. When a frame is detected as being erased, the coding parameters, especially the adaptive codebook gain and fixed codebook gain, of the erased and subsequent frames, are reestimated by a gain-matching procedure. By doing this, we can reduce the abrupt change caused in the decoded excitation signal by a simple scaling down procedure. We have applied this technique to the IS-63-1 speech coder and found that the proposed algorithm improves the speech quality under various channel conditions compared with the conventional extrapolation-based concealment algorithm.
Hong Kook Kim, Hong-Goo Kang
IEEE Signal Process. Lett.1
2001 A bitstream-based front-end for wireless speech recognition on IS-136 communications system
abstract
We propose a feature extraction method for a speech recognizer that operates in digital communication networks. The feature parameters are basically extracted by converting the quantized spectral information of a speech coder into a cepstrum. We also include the voiced/unvoiced information obtained from the bitstream of the speech coder in the recognition feature set. We performed speaker-independent connected digit HMM recognition experiments under clean, background noise, and channel impairment conditions. From these results, we found that the speech recognition system employing the proposed bitstream-based front-end gives superior word and string accuracies over a recognizer constructed from decoded speech signals. Its performance is comparable to that of a wireline recognition system that uses the cepstrum as a feature set. Next, we extended the evaluation of the proposed bitstream-based front-end to large vocabulary speech recognition with a name database. The recognition results proved that the proposed bitstream-based front-end also gives a comparable performance to the conventional wireline front-end.
Hong Kook Kim, Richard V. Cox
IEEE Trans. Speech Audio Process.1
2000 Bitstream-based feature extraction for wireless speech recognition
abstract
In this paper, we propose a feature extraction method for a speech recognizer that operates in digital communication networks. The feature parameters are basically extracted by converting the quantized spectral information of a speech coder into a cepstrum. We also combine the voiced/unvoiced information obtained from the bitstream of the speech coder into the recognition feature set. From speaker-independent connected digit HMM recognition, we find that the speech recognition system employing the proposed bitstream-based front-end gives superior word and string accuracies over a recognizer constructed from decoded speech signals. Its performance is comparable to that of the wireline recognition system that uses only the cepstrum as a feature set.
Hong Kook Kim, Richard V. Cox
ICASSP1
2000 Speech recognition using quantized LSP parameters and their transformations in digital communication
Seung Ho Choi, Hong Kook Kim, Hwang Soo Lee
Speech Commun.2
2000 On approximating line spectral frequencies to LPC cepstral coefficients
abstract
We propose an approximation of the line spectral frequencies (LSFs) to the LPC cepstral coefficients (LPCCs). A direct relationship between LPCC and LSF is derived, and new parameters, named pseudo-cepstral coefficients, are obtained by reducing some factors in the relationship. We also address the statistical properties of the pseudo-cepstral coefficients and show their useful application to speech recognition.
Hong Kook Kim, Seung Ho Choi, Hwang Soo Lee
IEEE Trans. Speech Audio Process.1
1999 LSP weighting functions based on spectral sensitivity and mel-frequency warping for speech recognition in digital communication
abstract
In digital communication networks, a speech recognition system extracts feature parameters after reconstructing speech signals. In this paper, we consider a useful approach of incorporating speech coding parameters into a speech recognizer. Most speech coders employ line spectrum pairs (LSPs) to represent spectral parameters. We introduce weighted distance measures to improve the recognition performance of an LSP-based speech recognizer. Experiments on speaker-independent connected-digit recognition showed that weighted distance measures provide better recognition accuracy than unweighted distance measures do. Compared with a conventional method employing mel-frequency cepstral coefficients, the proposed method achieved higher performance in terms of a recognition accuracy.
Seung Ho Choi, Hong Kook Kim, Hwang Soo Lee
ICASSP2
1999 A 4 kbps adaptive fixed code-excited linear prediction speech coder
abstract
We propose an adaptive fixed code-excited linear prediction (AF-CELP) speech coder operating at 4 kbps. By exploiting the fact that a fixed codebook contribution to the speech signal is also periodic as the corresponding adaptive codebook contribution, the adaptive fixed codebook model efficiently represents excitation signals. In order to overcome the quality degradation caused by the coarse quantization of excitation, a paired pulse algebraic codebook structure is also applied to the excitation model. Additionally, a pitch prefiltering, a noise spreading, and a harmonic enhancement technique are adopted in the decoding process. The spectrogram reading and informal listening tests proved that the AF-CELP reproduces high quality speech.
Hong Kook Kim, Mi Suk Lee, Hwang Soo Lee
ICASSP1
1999 Interlacing properties of line spectrum pair frequencies
abstract
An interlacing property of the line spectrum pair frequency (LSF) is proved on the basis of the logarithmic spectral difference function defined by the autoregressive models of successive orders. The property that the LSFs of an order are interlaced with those of lower order, provides a tight bound on the formant frequency region.
Hong Kook Kim, Hwang Soo Lee
IEEE Trans. Speech Audio Process.1
1999 Use of spectral autocorrelation in spectral envelope linear prediction for speech recognition
abstract
This paper proposes a linear predictive (LP) analysis method where sample autocorrelations are estimated from the spectral envelope of a speech signal on the basis of the spectral autocorrelation. The spectral autocorrelation is defined as discrete quantities of speech spectrum with spectral resolution identical to the discrete Fourier transform (DFT) used to obtain the speech spectrum. From analytical and empirical derivation of its properties, we can estimate the fundamental frequency and the maximally correlated frequency for voiced and unvoiced speech, respectively, and then obtain the spectral envelope by sampling at a rate of the estimated frequency. A frequency normalization can be applied to the estimated spectral envelope because the number of samples of the spectral envelope usually differs from frame to frame. The spectral envelope is warped into the mel-frequency scale and the inverse DFT is applied to extract the estimate of sample autocorrelations. From the result of LP analysis on the sample autocorrelations, we finally obtain the spectral envelope cepstral coefficients (SECC). Hidden Markov model (HMM) recognition experiments show that SECC significantly improves the performance of a recognizer at low signal-to-noise ratios (SNRs) over several other representations.
Hong Kook Kim, Hwang Soo Lee
IEEE Trans. Speech Audio Process.1
1998 Adaptive encoding of fixed codebook in CELP coders
abstract
We propose an adaptive encoding method of fixed codebook in CELP coders and implement an adaptive fixed code excited linear prediction (AF-CELP) speech coder. The AF-CELP exploits the fact that the fixed codebook contribution to the speech signal is also periodic as the adaptive codebook (or pitch filter) contribution. By modeling the fixed codebook with the pitch lag and the gain from the adaptive codebook, the AF-CELP can be implemented at low bit rates as well as low complexity. Listening tests show that a 6.4 kbit/s AF-CELP has a comparable quality to the 8 kbit/s CS-ACELP.
Hong Kook Kim
ICASSP1
1997 A 4 kbit/s renewal code excited linear prediction speech coder
abstract
This paper proposes a new 4 kbit/s speech coder based on CELP structure with 45 ms total codec delay. The coder is mainly featured by the renewal codebook of the excitation signal and the linked split-vector quantizer of line spectrum pair parameters which enable the coder to get high quality speech at low bit rate. In addition, techniques of formant enhancement in the spectral envelop and harmonic recovery in the transient region are also introduced to reduce buzzy and hoarse sounds, respectively. From the intensive listening test with intermediated response system (IRS) speech, we obtained a comparable subjective quality to 32 kbit/s ADPCM (ITU Recommendation G.726) under a nominal speech input level of -26 dB overload.
Hong Kook Kim, Yong Duk Cho, Sang Ryong Kim
ICASSP1
1997 Joint estimation of pitch, band magnitudes, and v\UV decisions for MBE vocoder
Yong Duk Cho, Hong Kook Kim, Sang Ryong Kim
EUROSPEECH2