Hong-Goo Kang

dblp:04/6605 · DBLP profile ↗
← Back
123ranked-venue papers
4as first author
35since 2021 · last 2026
0000-0002-6554-0783ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 113 · 4 first-author · 34 since 2021Artificial intelligence and machine learning · 60 · 1 first-author · 22 since 2021
YearPublicationVenuePosition
2026 Content-Aware Style Augmentation for Zero-Shot Voice Conversion With Short Target Speech
abstract
In this letter, we propose a neural zero-shot voice conversion (ZS-VC) system that simultaneously achieves high speaker similarity and speech intelligibility by incorporating a content-aware style generation module. Although recent neural ZS-VC systems have shown strong performance in either speaker similarity or speech intelligibility, attaining high performance in both remains challenging, especially when only a short target speech sample is available. We attribute this limitation to the insufficient content problem—where the linguistic content of the target speech fails to fully cover that of the source speech. To address this issue, we introduce a method that augments the target speaker's style features for underrepresented content using self-supervised feature generation. Experimental results demonstrate that the proposed system, when integrated with the feature matching-based approach kNN-VC, outperforms existing methods in both key metrics. Demo samples are available at https://hyeonjincha.github.io/.
Hyeonjin Cha, Seyun Um, Miseul Kim, Seungshin Lee, Hong-Goo Kang
IEEE Signal Process. Lett.6
2026 GoP-Based Quality Enhancement on Video Compression
abstract
With recent increases in the demand for high-resolution video content, it has become increasingly challenging to transmit video data within the constraints of limited bandwidth. Due to the time-consuming nature of developing and disseminating new standard codecs, a large body of research has addressed improving low-quality videos through post-processing techniques. Previous studies have primarily concentrated on enhancing the quality of compressed video by addressing the temporal consistency of adjacent frames over short durations. However, these approaches often overlook specific characteristics of the video coding framework, such as notable variations in codec artifact patterns occurring at the Group of Pictures (GoP) level, which can result in considerable viewer discomfort. In this paper, we propose GoP-based Quality Enhancement (GQE), which aims to improve the quality of compressed videos by addressing issues at the GoP level. First, we present a GoP Guided Feature Propagation (GGFP) module, which addresses the root cause of the GoP level issue by propagating features from the I-frame of a different GoP to the frames currently undergoing enhancement. Then, we introduce a Temporal Aggregation (TA) module to efficiently and effectively aggregate features from the I-frame and the current frame. We extensively evaluate our model using diverse test sequences across a range of codecs, including HEVC, VP9, and AV1. Our approach not only achieves a significant reduction in the pattern shifts of GoP-level artifacts, but also demonstrates a substantial improvement in overall video quality.
Chajin Shin, Hong-Goo Kang, Sangyoun Lee
IEEE Trans. Image Process.3
2025 LAMA-UT: Language Agnostic Multilingual ASR Through Orthography Unification and Language-Specific Transliteration
abstract
Building a universal multilingual automatic speech recognition (ASR) model that performs equitably across languages has long been a challenge due to its inherent difficulties. To address this task we introduce a Language-Agnostic Multilingual ASR pipeline through orthography Unification and language-specific Transliteration (LAMA-UT). LAMA-UT operates without any language-specific modules while matching the performance of state-of-the-art models trained on a minimal amount of data. Our pipeline consists of two key steps. First, we utilize a universal transcription generator to unify orthographic features into Romanized form and capture common phonetic characteristics across diverse languages. Second, we utilize a universal converter to transform these universal transcriptions into language-specific ones. In experiments, we demonstrate the effectiveness of our proposed method leveraging universal transcriptions for massively multilingual ASR. Our pipeline achieves a relative error reduction rate of 45% when compared to Whisper and performs comparably to MMS, despite being trained on only 0.1% of Whisper's training data. Furthermore, our pipeline does not rely on any language-specific modules. However, it performs on par with zero-shot ASR approaches which utilize additional language-specific lexicons and language models. We expect this framework to serve as a cornerstone for flexible multilingual ASR systems that are generalizable even to unseen languages.
Woo-Jin Chung, Hong-Goo Kang
AAAI3
2025 StableQuant: Layer Adaptive Post-Training Quantization for Speech Foundation Models
abstract
In this paper, we propose StableQuant, a novel adaptive post-training quantization (PTQ) algorithm for widely used speech foundation models (SFMs). While PTQ has been successfully employed for compressing large language models (LLMs) due to its ability to bypass additional fine-tuning, directly applying these techniques to SFMs may not yield optimal results, as SFMs utilize distinct network architecture for feature extraction. StableQuant demonstrates optimal quantization performance regardless of the network architecture type, as it adaptively determines the quantization range for each layer by analyzing both the scale distributions and overall performance. We evaluate our algorithm on two SFMs, HuBERT and wav2vec2.0, for an automatic speech recognition (ASR) task, and achieve superior performance compared to traditional PTQ methods. StableQuant successfully reduces the sizes of SFM models to a quarter and doubles the inference speed while limiting the word error rate (WER) performance drop to less than 0.3% with 8-bit quantization.
Yeona Hong, Hyewon Han, Woo-Jin Chung, Hong-Goo Kang
ICASSP4
2025 Neural Spectral Band Generation for Audio Coding
Woongjib Choi, Byeong Hyeon Kim, Hyungseob Lim, Inseon Jang, Hong-Goo Kang
INTERSPEECH5
2025 Quadruple Path Modeling with Latent Feature Transfer for Permutation-free Continuous Speech Separation
Hyewon Han, Jonguk Yoo, Chang Woo Han, Jeongook Song, Hoonyoung Cho, Hong-Goo Kang
INTERSPEECH9
2025 Towards an Ultra-Low-Delay Neural Audio Coding with Computational Efficiency
Byeong Hyeon Kim, Hyungseob Lim, Inseon Jang, Hong-Goo Kang
INTERSPEECH4
2025 SpeechMLC: Speech Multi-label Classification
Miseul Kim, Seyun Um, Hyeonjin Cha, Hong-Goo Kang
INTERSPEECH4
2024 On Fine-Tuning Pre-Trained Speech Models With EMA-Target Self-Supervised Loss
abstract
Representation models pre-trained on self-supervised objectives are often fine-tuned for solving downstream tasks. However, fine-tuning can degrade the general knowledge that was originally built up by the pre-training, which could help prevent the model from overfitting given sparse fine-tuning data or bridge gaps between different domains. We hypothesize that preserving this general knowledge in pre-trained models is crucial for improving performance on downstream tasks. Based on this idea, we propose a novel method for fine-tuning self-supervised speech models that utilizes a self-supervised loss over the course of fine-tuning. Then, an Exponential Moving Average (EMA) technique is applied to smoothly transition the domain of the model from the generalized to the task-oriented one. We perform various downstream tasks using the proposed method, finding that our method improves performance on most of the tasks. Results show that our method induces the generalization ability of the model to be retained without overshadowing the downstream task performance.
Hejung Yang, Hong-Goo Kang
ICASSP2
2024 Speaker-Independent Acoustic-to-Articulatory Inversion through Multi-Channel Attention Discriminator
abstract
We present a novel speaker-independent acoustic-toarticulatory inversion (AAI) model, overcoming the limitations observed in conventional AAI models that rely on acoustic features derived from restricted datasets.To address these challenges, we leverage representations from a pre-trained selfsupervised learning (SSL) model to more effectively estimate the global, local, and kinematic pattern information in Electromagnetic Articulography (EMA) signals during the AAI process.We train our model using an adversarial approach and introduce an attention-based Multi-duration phoneme discriminator (MDPD) designed to fully capture the intricate relationship among multi-channel articulatory signals.Our method achieves a Pearson correlation coefficient of 0.847, marking state-of-theart performance in speaker-independent AAI models.The implementation details and code can be found online 1 .
Woo-Jin Chung, Hong-Goo Kang
INTERSPEECH2
2024 Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation
Miseul Kim, Soo-Whan Chung, Youna Ji, Hong-Goo Kang, Min-Seok Choi
INTERSPEECH4
2024 Enhanced Deep Speech Separation in Clustered Ad Hoc Distributed Microphone Environments
abstract
Ad-hoc distributed microphone environments, where microphone locations and numbers are unpredictable, present a challenge to traditional deep learning models, which typically require fixed architectures. To tailor deep learning models to accommodate arbitrary array configurations, the Transform-Average-Concatenate (TAC) layer was previously introduced. In this work, we integrate TAC layers with dual-path transformers for speech separation from two simultaneous talkers in realistic settings. However, the distributed nature makes it hard to fuse information across microphones efficiently. Therefore, we explore the efficacy of blindly clustering microphones around sources of interest prior to enhancement. Experimental results show that this deep cluster-informed approach significantly improves the system's capacity to cope with the inherent variability observed in ad-hoc distributed microphone environments.
Stijn Kindt, Nilesh Madhu, Hong-Goo Kang
INTERSPEECH4
2024 PARAN: Variational Autoencoder-based End-to-End Articulation-to-Speech System for Speech Intelligibility
Seyun Um, Hong-Goo Kang
INTERSPEECH3
2024 UNIQUE : Unsupervised Network for Integrated Speech Quality Evaluation
Juhwan Yoon, WooSeok Ko, Seyun Um, Sungwoong Hwang, Soojoong Hwang, Hong-Goo Kang
INTERSPEECH7
2024 Disentangled Representations in Local-Global Contexts for Arabic Dialect Identification
abstract
In this article, we propose a locally and globally informed disentanglement network for Arabic dialect identification (ADI). Our proposed disentanglement network aims to detach all irrelevant information (e.g., speaker, gender and channel) from the source utterance and extract only dialect-related representations fitted for the ADI problem. The proposed network consists of local convolutional backbone modules to learn low-resolution feature maps and self-attention-based bottleneck transformers to efficiently aggregate the local information to represent the global context as the learned dialect embeddings. We propose a novel supervised clustering loss to minimize intra-class variations and maximize inter-class variations in a latent space. Our model achieves state-of-the-art results in qualitative and quantitative evaluations by outperforming other competitive solutions on ADI-17 datasets. Specifically, we have shown that local-global awareness from our proposed network boosts feature representation and enhances identification performance.
Zainab Alhakeem, Se-In Jang, Hong-Goo Kang
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Progressive Multi-Stage Neural Audio Codec with Psychoacoustic Loss and Discriminator
abstract
In this paper, we improve the efficiency of the progressive multi-stage neural audio codec (PR-Codec) by utilizing perceptually motivated training criteria. Although our baseline PR-Codec successfully reconstructs full-band signals by progressively decoding the pre-defined subband signals, transparent quality can only be guaranteed in high bit-rates. To reduce bit-rates while maintaining perceptually transparent quality, we adopt a psychoacoustic model (PAM)-based loss and propose a perceptual weighting discriminator (PWD), which enables us to synthesize and discriminate audio signals in the perceptually motivated domain. We also introduce a scalar quantization with an entropy model to further enhance the quantization efficiency. Our experimental results show that our proposed model significantly improves perceptual reconstruction quality at the expense of the waveform disparity in the time-domain, compared to our previous model.
Byeong Hyeon Kim, Hyungseob Lim, Inseon Jang, Hong-Goo Kang
ICASSP5
2023 Style Modeling for Multi-Speaker Articulation-to-Speech
abstract
In this paper, we propose a neural articulation-to-speech (ATS) framework that synthesizes high-quality speech from articulatory signal in a multi-speaker situation. Most conventional ATS approaches only focus on modeling contextual information of speech from a single speaker’s articulatory features. To explicitly represent each speaker’s speaking style as well as the contextual information, our proposed model estimates style embeddings, guided from the essential speech style attributes such as pitch and energy. We adopt convolutional layers and transformer-based attention layers for our model to fully utilize both local and global information of articulatory signals, measured by electromagnetic articulogra-phy (EMA). Our model significantly improves the quality of synthesized speech compared to the baseline in terms of objective and subjective measurements in the Haskins dataset.
Miseul Kim, Zhenyu Piao, Hong-Goo Kang
ICASSP4
2023 End-to-End Neural Audio Coding in the MDCT Domain
abstract
Modern deep neural network (DNN)-based audio coding approaches utilize complicated non-linear functions (e.g., convolutional neural networks and non-linear activations), which leads to high complexity and memory usage. However, their decoded audio quality is still not much higher than that of signal processing-based legacy codecs. In this paper, we propose an effective frequency-domain neural audio coding paradigm that adopts the modified discrete cosine transform (MDCT) for analysis and synthesis and DNNs for the quantization of variables. It includes an efficient method to encode MDCT bins as well as a mechanism to adapt the quantization level of each bin. Our neural audio codec is trained in an end-to-end manner with the help of psychoacoustics-based perceptual loss, removing the burden of module-by-module fine-tuning. Experimental results show that our proposed model’s performance is comparable with the MP3 codec at around 64 and 48 kbps bit-rates for mono signals.
Hyungseob Lim, Byeong Hyeon Kim, Inseon Jang, Hong-Goo Kang
ICASSP5
2023 HappyQuokka System for ICASSP 2023 Auditory EEG Challenge
abstract
This report describes our submission to Task 2 of the Auditory EEG Decoding Challenge at ICASSP 2023 Signal Processing Grand Challenge (SPGC). Task 2 is a regression problem that focuses on reconstructing a speech envelope from an EEG signal. For the task, we propose a pre-layer normalized feedforward transformer (FFT) architecture. For within-subjects generation, we additionally utilize an auxiliary global conditioner which provides our model with additional information about seen individuals. Experimental results show that our proposed method outperforms the VLAAI baseline and all other submitted systems. Notably, it demonstrates significant improvements on the within-subjects task, likely thanks to our use of the auxiliary global conditioner. In terms of evaluation metrics set by the challenge, we obtain Pearson correlation values of 0.1895 ± 0.0869 for the within-subjects generation test and 0.0976 ± 0.0444 for the heldout-subjects test. We release the training code for our model online.1
Zhenyu Piao, Miseul Kim, Hyungchan Yoon, Hong-Goo Kang
ICASSP4
2023 MF-PAM: Accurate Pitch Estimation through Periodicity Analysis and Multi-level Feature Fusion
Woo-Jin Chung, Soo-Whan Chung, Hong-Goo Kang
INTERSPEECH4
2023 HD-DEMUCS: General Speech Restoration with Heterogeneous Decoders
Soo-Whan Chung, Hyewon Han, Youna Ji, Hong-Goo Kang
INTERSPEECH5
2023 Contrastive Learning based Deep Latent Masking for Music Source Separation
Hong-Goo Kang
INTERSPEECH2
2023 Feature Normalization for Fine-tuning Self-Supervised Models in Speech Enhancement
Hejung Yang, Hong-Goo Kang
INTERSPEECH2
2023 Pruning Self-Attention for Zero-Shot Multi-Speaker Text-to-Speech
abstract
For personalized speech generation, a neural text-to-speech (TTS) model must be successfully implemented with limited data from a target speaker. To this end, the baseline TTS model needs to be amply generalized to out-of-domain data (i.e., target speaker's speech). However, approaches to address this out-of-domain generalization problem in TTS have yet to be thoroughly studied. In this work, we propose an effective pruning method for a transformer known as sparse attention, to improve the TTS model's generalization abilities. In particular, we prune off redundant connections from self-attention layers whose attention weights are below the threshold. To flexibly determine the pruning strength for searching optimal degree of generalization, we also propose a new differentiable pruning method that allows the model to automatically learn the thresholds. Evaluations on zero-shot multi-speaker TTS verify the effectiveness of our method in terms of voice quality and speaker similarity.
Hyungchan Yoon, Eunwoo Song, Hyun-Wook Yoon, Hong-Goo Kang
INTERSPEECH5
2023 Adversarial Learning of Intermediate Acoustic Feature for End-to-End Lightweight Text-to-Speech
abstract
To simplify the generation process, several text-to-speech (TTS) systems implicitly learn intermediate latent representations instead of relying on predefined features (e.g., mel-spectrogram).However, their generation quality is unsatisfactory as these representations lack speech variances.In this paper, we improve TTS performance by adding prosody embeddings to the latent representations.During training, we extract reference prosody embeddings from mel-spectrograms, and during inference, we estimate these embeddings from text using generative adversarial networks (GANs).Using GANs, we reliably estimate the prosody embeddings in a fast way, which have complex distributions due to the dynamic nature of speech.We also show that the prosody embeddings work as efficient features for learning a robust alignment between text and acoustic features.Our proposed model surpasses several publicly available models with less parameters and computational complexity in comparative experiments.
Hyungchan Yoon, Seyun Um, Hong-Goo Kang
INTERSPEECH4
2023 SC-CNN: Effective Speaker Conditioning Method for Zero-Shot Multi-Speaker Text-to-Speech Systems
abstract
This letter proposes an effective speaker-conditioning method that is applicable to zero-shot multi-speaker text-to-speech (ZSM-TTS) systems. Based on the inductive bias in the speech generation task, in which local context information in text/phoneme sequences heavily affect the speaker characteristics of the output speech, we propose a Speaker-Conditional Convolutional Neural Network (SC-CNN) for the ZSM-TTS task. SC-CNN first predicts convolutional kernels from each learned speaker embedding, then applies 1-D convolutions to phoneme sequences with the predicted kernels. It utilizes the aforementioned inductive bias and effectively models the characteristic of speech by providing the speaker-specific local context in phonetic domain. We also build both FastSpeech2 and VITS-based ZSM-TTS systems to verify its superiority over conventional speaker conditioning methods. The results confirm that the models with SC-CNN outperform the recent ZSM-TTS models in terms of both subjective and objective measurements.
Hyungchan Yoon, Seyun Um, Hyun-Wook Yoon, Hong-Goo Kang
IEEE Signal Process. Lett.5
2022 Phase Continuity: Learning Derivatives of Phase Spectrum for Speech Enhancement
abstract
Modern neural speech enhancement models usually include various forms of phase information in their training loss terms, either explicitly or implicitly. However, these loss terms are typically designed to reduce the distortion of phase spectrum values at specific frequencies, which ensures they do not significantly affect the quality of the enhanced speech. In this paper, we propose an effective phase reconstruction strategy for neural speech enhancement that can operate in noisy environments. Specifically, we introduce a phase continuity loss that considers relative phase variations across the time and frequency axes. By including this phase continuity loss in a state-of-the-art neural speech enhancement system trained with reconstruction loss and a number of magnitude spectral losses, we show that our proposed method further improves the quality of enhanced speech signals over the baseline, especially when training is done jointly with a magnitude spectrum loss.
Hyewon Han, Hyeon-Kyeong Shin, Soo-Whan Chung, Hong-Goo Kang
ICASSP5
2022 Progressive Multi-Stage Neural Audio Coding with Guided References
abstract
In this paper, we propose an effective multi-stage neural audio coding algorithm that encodes full-band audio signals (up to 20 kHz) using an end-to-end training criterion. By predefining several dyadic subband signals as training targets, we progressively encode input audio signals in each stage such that deeper stages of the network encode the residual error terms from the previous encoding stage. Our proposed audio codec successfully decodes full-band audio signals by using an effective multi-stage vector quantization scheme to represent key encoding features extracted in the latent space. Subjective listening tests show that the decoded outputs of the proposed audio codec achieve almost transparent quality at an average bitrate of 132 kbps.
Chanwoo Lee, Hyungseob Lim, Inseon Jang, Hong-Goo Kang
ICASSP5
2022 Adversarial Audio Synthesis Using a Harmonic-Percussive Discriminator
abstract
In this paper, we propose a discriminator design scheme for generative adversarial network-based audio signal generation. Unlike conventional discriminators that take an entire signal as input, our discriminator separates the audio signal into harmonic and percussive components and analyzes each component independently. The rationale behind this idea is that conventional discriminators cannot reliably capture subtle distortions in audio signals, which have complicated time-frequency characteristics. By considering the time-frequency resolution of audio signals, our proposed method encourages the generator to better reconstruct harmonic and percussive features, both of which are critical for the quality of the generated signals. Listening tests show that our framework significantly enhances the stability of pitches and generates clearer piano samples compared to a baseline.
Hyungseob Lim, Chanwoo Lee, Inseon Jang, Hong-Goo Kang
ICASSP5
2022 Light-Weight Speaker Verification with Global Context Information
Miseul Kim, Zhenyu Piao, Seyun Um, Ran Lee, Jaemin Joh, Seungshin Lee, Hong-Goo Kang
INTERSPEECH7
2022 FluentTTS: Text-dependent Fine-grained Style Control for Multi-style TTS
Seyun Um, Hyungchan Yoon, Hong-Goo Kang
INTERSPEECH4
2022 Learning Audio-Text Agreement for Open-vocabulary Keyword Spotting
abstract
In this paper, we propose a novel end-to-end user-defined keyword spotting method that utilizes linguistically corresponding patterns between speech and text sequences. Unlike previous approaches requiring speech keyword enrollment, our method compares input queries with an enrolled text keyword sequence. To place the audio and text representations within a common latent space, we adopt an attention-based cross-modal matching approach that is trained in an end-to-end manner with monotonic matching loss and keyword classification loss. We also utilize a de-noising loss for the acoustic embedding network to improve robustness in noisy environments. Additionally, we introduce the LibriPhrase dataset, a new short-phrase dataset based on LibriSpeech for efficiently training keyword spotting models. Our proposed method achieves competitive results on various evaluation sets compared to other single-modal and cross-modal baselines.
Hyeon-Kyeong Shin, Hyewon Han, Soo-Whan Chung, Hong-Goo Kang
INTERSPEECH5
2022 Two-Stage Refinement of Magnitude and Complex Spectra for Real-Time Speech Enhancement
abstract
In this letter, we propose a two-stage network for performing speech enhancement that predicts magnitude spectra in the first stage and complex spectra in the second stage. To maximize the model's performance at each stage, we propose two convolutional modules: magnitude spectral masking (MSM) and complex spectra refinement (CSR). Each module is designed to take into account the specific characteristics of the signal type it handles. The MSM estimates multiplicative masks to remove noise in the magnitude component of the convolutional features, and the CSR refines the complex component of the convolutional features using additive features. By using these modules, our proposed two-stage enhancement model shows higher performance than previously proposed state-of-the-art algorithms. In addition, the number of parameters of our model is only 2.63 million, and it can operate in real time thanks to its causal characteristics and low computational complexity.
Hong-Goo Kang
IEEE Signal Process. Lett.2
2021 Looking Into Your Speech: Learning Cross-Modal Affinity for Audio-Visual Speech Separation
abstract
In this paper, we address the problem of separating individual speech signals from videos using audio-visual neural processing. Most conventional approaches utilize frame-wise matching criteria to extract shared information between co-occurring audio and video. Thus, their performance heavily depends on the accuracy of audio-visual synchronization and the effectiveness of their representations. To overcome the frame discontinuity problem between two modalities due to transmission delay mismatch or jitter, we propose a cross-modal affinity network (CaffNet) that learns global correspondence as well as locally-varying affinities between audio and visual streams. Given that the global term provides stability over a temporal sequence at the utterance-level, this resolves the label permutation problem characterized by inconsistent assignments. By extending the proposed cross-modal affinity on the complex network, we further improve the separation performance in the complex spectral domain. Experimental results verify that the proposed methods outperform conventional ones on various datasets, demonstrating their advantages in real-world scenarios.
Jiyoung Lee 0005, Soo-Whan Chung, Sunok Kim, Hong-Goo Kang, Kwanghoon Sohn
CVPR4
2021 LiteTTS: A Lightweight Mel-Spectrogram-Free Text-to-Wave Synthesizer Based on Generative Adversarial Networks
Huu-Kim Nguyen, Kihyuk Jeong, Seyun Um, Min-Jae Hwang, Eunwoo Song, Hong-Goo Kang
Interspeech6
2020 Improving LPCNET-Based Text-to-Speech with Linear Prediction-Structured Mixture Density Network
abstract
In this paper, we propose an improved LPCNet vocoder using a linear prediction (LP)-structured mixture density network (MDN). The recently proposed LPCNet vocoder has successfully achieved high-quality and lightweight speech synthesis systems by combining a vocal tract LP filter with a WaveRNN-based vocal source (i.e., excitation) generator. However, the quality of synthesized speech is often unstable because the vocal source component is insufficiently represented by the μ-law quantization method, and the model is trained without considering the entire speech production mechanism. To address this problem, we first introduce LP-MDN, which enables the autoregressive neural vocoder to structurally represent the interactions between the vocal tract and vocal source components. Then, we propose to incorporate the LP-MDN to the LPCNet vocoder by replacing the conventional discretized output with continuous density distribution. The experimental results verify that the proposed system provides high quality synthetic speech by achieving a mean opinion score of 4.41 within a text-to-speech framework.
Min-Jae Hwang, Eunwoo Song, Ryuichi Yamamoto, Frank K. Soong, Hong-Goo Kang
ICASSP5
2020 Emotional Speech Synthesis with Rich and Granularized Control
abstract
This paper proposes an effective emotion control method for an end-to-end text-to-speech (TTS) system. To flexibly control the distinct characteristic of a target emotion category, it is essential to determine embedding vectors representing the TTS input. We introduce an inter-to-intra emotional distance ratio algorithm to the embedding vectors that can minimize the distance to the target emotion category while maximizing its distance to the other emotion categories. To further enhance the expressiveness of a target speech, we also introduce an effective interpolation technique that enables the intensity of a target emotion to be gradually changed to that of neutral speech. Subjective evaluation results in terms of emotional expressiveness and controllability show the superiority of the proposed algorithm to the conventional methods.
Seyun Um, Sangshin Oh, Kyungguen Byun, Inseon Jang, Chunghyun Ahn, Hong-Goo Kang
ICASSP6
2020 FaceFilter: Audio-Visual Speech Separation Using Still Images
abstract
The objective of this paper is to separate a target speaker's speech from a mixture of two speakers using a deep audio-visual speech separation network. Unlike previous works that used lip movement on video clips or pre-enrolled speaker information as an auxiliary conditional feature, we use a single face image of the target speaker. In this task, the conditional feature is obtained from facial appearance in cross-modal biometric task, where audio and visual identity representations are shared in latent space. Learnt identities from facial images enforce the network to isolate matched speakers and extract the voices from mixed speech. It solves the permutation problem caused by swapped channel outputs, frequently occurred in speech separation tasks. The proposed method is far more practical than video-based speech separation since user profile images are readily available on many platforms. Also, unlike speaker-aware separation methods, it is applicable on separation with unseen speakers who have never been enrolled before. We show strong qualitative and quantitative results on challenging real-world examples.
Soo-Whan Chung, Soyeon Choe, Joon Son Chung, Hong-Goo Kang
INTERSPEECH4
2020 Seeing Voices and Hearing Voices: Learning Discriminative Embeddings Using Cross-Modal Self-Supervision
abstract
The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effective representations can be learnt from natural cross-modal synchrony. We build on earlier work to train embeddings that are more discriminative for uni-modal downstream tasks. To this end, we propose a novel training strategy that not only optimises metrics across modalities, but also enforces intra-class feature separation within each of the modalities. The effectiveness of the method is demonstrated on two downstream tasks: lip reading using the features trained on audio-visual synchronisation, and speaker recognition using the features trained for cross-modal biometric matching. The proposed method outperforms state-of-the-art self-supervised baselines by a signficant margin.
Soo-Whan Chung, Hong-Goo Kang, Joon Son Chung
INTERSPEECH2
2020 MIRNet: Learning Multiple Identities Representations in Overlapped Speech
abstract
Many approaches can derive information about a single speaker's identity from the speech by learning to recognize consistent characteristics of acoustic parameters.However, it is challenging to determine identity information when there are multiple concurrent speakers in a given signal.In this paper, we propose a novel deep speaker representation strategy that can reliably extract multiple speaker identities from an overlapped speech.We design a network that can extract a highlevel embedding that contains information about each speaker's identity from a given mixture.Unlike conventional approaches that need reference acoustic features for training, our proposed algorithm only requires the speaker identity labels of the overlapped speech segments.We demonstrate the effectiveness and usefulness of our algorithm in a speaker verification task and a speech separation system conditioned on the target speaker embeddings obtained through the proposed method.
Hyewon Han, Soo-Whan Chung, Hong-Goo Kang
INTERSPEECH3
2020 A Cross-Channel Attention-Based Wave-U-Net for Multi-Channel Speech Enhancement
Minh Tri Ho, Bong-Ki Lee, Dong Hoon Yi, Hong-Goo Kang
INTERSPEECH5
2020 Intra-Class Variation Reduction of Speaker Representation in Disentanglement Framework
abstract
In this paper, we propose an effective training strategy to extract robust speaker representations from a speech signal.One of the key challenges in speaker recognition tasks is to learn latent representations or embeddings containing solely speaker characteristic information in order to be robust in terms of intraspeaker variations.By modifying the network architecture to generate both speaker-related and speaker-unrelated representations, we exploit a learning criterion which minimizes the mutual information between these disentangled embeddings.We also introduce an identity change loss criterion which utilizes a reconstruction error to different utterances spoken by the same speaker.Since the proposed criteria reduce the variation of speaker characteristics caused by changes in background environment or spoken content, the resulting embeddings of each speaker become more consistent.The effectiveness of the proposed method is demonstrated through two tasks; disentanglement performance, and improvement of speaker recognition accuracy compared to the baseline model on a benchmark dataset, VoxCeleb1.Ablation studies also show the impact of each criterion on overall performance.
Yoohwan Kwon, Soo-Whan Chung, Hong-Goo Kang
INTERSPEECH3
2020 Speaker-Adaptive Neural Vocoders for Parametric Speech Synthesis Systems
abstract
This paper proposes speaker-adaptive neural vocoders for parametric text-to-speech (TTS) systems. Recently proposed WaveNet-based neural vocoding systems successfully generate a time sequence of speech signal with an autoregressive framework. However, it remains a challenge to synthesize high-quality speech when the amount of a target speaker's training data is insufficient. To generate more natural speech signals with the constraint of limited training data, we propose a speaker adaptation task with an effective variation of neural vocoding models. In the proposed method, a speaker-independent training method is applied to capture universal attributes embedded in multiple speakers, and the trained model is then optimized to represent the specific characteristics of the target speaker. Experimental results verify that the proposed TTS systems with speaker-adaptive neural vocoders outperform those with traditional source-filter model-based vocoders and those with WaveNet vocoders, trained either speaker-dependently or speaker-independently. In particular, our TTS system achieves 3.80 and 3.77 MOS for the Korean male and Korean female speakers, respectively, even though we use only ten minutes' speech corpus for training the model.
Eunwoo Song, Jin-Seob Kim, Kyungguen Byun, Hong-Goo Kang
MMSP4
2019 Perfect Match: Improved Cross-modal Embeddings for Audio-visual Synchronisation
abstract
This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronisation. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant audio segment given a short video clip. The method builds on the recent advances in learning representations from cross-modal self-supervision. The main contributions of this paper are as follows: (1) we propose a new learning strategy where the embeddings are learnt via a multi-way matching problem, as opposed to a binary classification (matching or non-matching) problem as proposed by recent papers; (2) we demonstrate that performance of this method far exceeds the existing baselines on the synchronisation task; (3) we use the learnt embeddings for visual speech recognition in self-supervision, and show that the performance matches the representations learnt end-to-end in a fully-supervised manner.
Soo-Whan Chung, Joon Son Chung, Hong-Goo Kang
ICASSP3
2019 Gradient-based Active Learning Query Strategy for End-to-end Speech Recognition
abstract
In this paper, we propose an effective active learning query strategy for an automatic speech recognition system with the aim of reducing the training cost. Generally, training a deep neural network with supervised learning requires a massive amount of labeled data to obtain excellent performance. However, labeling data is tedious and costly manual work. Active learning can solve this problem by choosing and only annotating informative instances, which presents better results even with less transcribed data. In this approach it is vitally important to accurately select informative samples. Based on the preliminary experiment results that true gradient length has the best performance in terms of measuring sample informativeness in ideal conditions, we propose utilizing both uncertainty and the expected gradient length criterion to approximate the true gradient length using a neural network. The experiment results show that our proposed method is superior to the conventional individual criterion when applied to a phoneme-based speech recognition system, and it has both a faster convergence speed and the greatest loss reduction in both clean and noisy conditions.
Soo-Whan Chung, Hong-Goo Kang
ICASSP3
2019 Parameter Enhancement for MELP Speech Codec in Noisy Communication Environment
abstract
In this paper, we propose a deep learning (DL)-based parameter enhancement method for a mixed excitation linear prediction (MELP) speech codec in noisy communication environment.Unlike conventional speech enhancement modules that are designed to obtain clean speech signal by removing noise components before speech codec processing, the proposed method directly enhances codec parameters on either the encoder or decoder side.As the proposed method has been implemented by a small network without any additional processes required in conventional enhancement systems, e.g., time-frequency (T-F) analysis/synthesis modules, its computational complexity is very low.By enhancing the noise-corrupted codec parameters with the proposed DL framework, we achieved an enhancement system that is much simpler and faster than conventional T-F mask-based speech enhancement methods, while the quality of its performance remains similar.
Min-Jae Hwang, Hong-Goo Kang
INTERSPEECH2
2019 An Effective Style Token Weight Control Technique for End-to-End Emotional Speech Synthesis
abstract
In this letter, we propose a high-quality emotional speech synthesis system, using emotional vector space, i.e., the weighted sum of global style tokens (GSTs). Our previous research verified the feasibility of GST-based emotional speech synthesis in an end-to-end text-to-speech synthesis framework. However, selecting appropriate reference audio (RA) signals to extract emotion embedding vectors to the specific types of target emotions remains problematic. To ameliorate the selection problem, we propose an effective way of generating emotion embedding vectors by utilizing the trained GSTs. By assuming that the trained GSTs represent an emotional vector space, we first investigate the distribution of all the training samples depending on the type of each emotion. We then regard the centroid of the distribution as an emotion-specific weighting value, which effectively controls the expressiveness of synthesized speech, even without using the RA for guidance, as it did before. Finally, we confirm that the proposed controlled weight-based method is superior to the conventional emotion label-based methods in terms of perceptual quality and emotion classification accuracy.
Ohsung Kwon, Inseon Jang, Chunghyun Ahn, Hong-Goo Kang
IEEE Signal Process. Lett.4
2019 A Joint Learning Algorithm for Complex-Valued T-F Masks in Deep Learning-Based Single-Channel Speech Enhancement Systems
abstract
This paper presents a joint learning algorithm for complex-valued time-frequency (T-F) masks in single-channel speech enhancement systems. Most speech enhancement algorithms operating in a single-channel microphone environment aim to enhance the magnitude component in a T-F domain, while the input noisy phase component is used directly without any processing. Consequently, the mismatch between the processed magnitude and the unprocessed phase degrades the sound quality. To address this issue, a learning method of targeting a T-F mask that is defined in a complex domain has recently been proposed. However, due to a wide dynamic range and an irregular spectrogram pattern of the complex-valued T-F mask, the learning process is difficult even with a large-scale deep learning network. Moreover, the learning process targeting the T-F mask itself does not directly minimize the distortion in spectra or time domains. In order to address these concerns, we focus on three issues: 1) an effective estimation of complex numbers with a wide dynamic range; 2) a learning method that is directly related to speech enhancement performance; and 3) a way to resolve the mismatch between the estimated magnitude and phase spectra. In this study, we propose objective functions that can solve each of these issues and train the network by minimizing them with a joint learning framework. The evaluation results demonstrate that the proposed learning algorithm achieves significant performance improvement in various objective measures and subjective preference listening test.
Jinkyu Lee 0002, Hong-Goo Kang
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Modeling-By-Generation-Structured Noise Compensation Algorithm for Glottal Vocoding Speech Synthesis System
abstract
This paper proposes a novel noise compensation algorithm for a glottal excitation model in a deep learning (DL)-based speech synthesis system. To generate high-quality speech synthesis outputs, the balance between harmonic and noise components of the glottal excitation signal should be well-represented by the DL network. However, it is hard to accurately model the noise component because the DL training process inevitably results in statistically smoothed outputs; thus, it is essential to introduce an additional noise compensation process. We propose a modeling-by-generation structure-based noise compensation method that the missing noise component in the generated glottal signal is directly extracted and parameterized during the entire training process. By modeling the noise component using the additional DL network, the proposed system successfully compensates the missing noise component. Objective and subjective test results confirm that the synthesized speech with the proposed noise compensation method is superior to that with conventional methods.
Min-Jae Hwang, Eunwoo Song, Kyungguen Byun, Hong-Goo Kang
ICASSP4
2018 Dnn-Based Wireless Positioning in an Outdoor Environment
abstract
In this paper, we propose a deep learning based algorithm to estimate the position of an user by utilizing reference signal received power (RSRP) and the location of base stations. To obtain reliable results in a real communication environment, parameters were measured using commercially available base stations and mobile phones within a LTE network. Since the structure of the measured data changes in accordance with the number of connected base stations, it is necessary to work on data uniformity processing before running the deep learning network. Therefore, we extract only the case in which three base stations are connected, using it as a feature of deep learning network. The experimental results reveal that the performance of the proposed algorithm is much better than that of the conventional fingerprint method. The average distance error is reduced from 71.04 meters for the fingerprint-based method to 43.51 meters for the proposed deep learning-based method.
Chahyeon Eom, Youngsu Kwak, Hong-Goo Kang, Chungyong Lee
ICASSP4
2018 A Unified Framework for the Generation of Glottal Signals in Deep Learning-based Parametric Speech Synthesis Systems
Min-Jae Hwang, Eunwoo Song, Jin-Seob Kim, Hong-Goo Kang
INTERSPEECH4
2018 AVSU: Workshop on Audio-Visual Scene Understanding for Immersive Multimedia
abstract
This workshop aims to provide a forum to exchange ideas in scene understanding techniques researched in audio and visual communities, and to ultimately unlock the creative potential of joint audio-visual signal processing to deliver a step change in various multimedia applications. Papers and talks presented in this workshop will contribute to the emerging technology for audio and visual information that can improve traditional approaches for multimedia content production and reproduction. The goals of this workshop are to (1) present and discuss the latest trends in audio and computer vision fields for the common research goals, (2) understand state-of-the-art techniques and bottlenecks in the other's discipline for the common topics, (3) investigate research opportunities of joint audio-visual scene understandings in multimedia content production. This workshop will be a good opportunity to bring together leading experts in audio processing and computer vision, and will bridge the gap between two research fields in multimedia content production and reproduction.
Adrian Hilton 0001, Hong-Goo Kang, Hansung Kim 0001, Kwanghoon Sohn
ACM Multimedia2
2018 Phase-Sensitive Joint Learning Algorithms for Deep Learning-Based Speech Enhancement
abstract
This letter presents a phase-sensitive joint learning algorithm for single-channel speech enhancement. Although a deep learning framework that estimates the time-frequency (T-F) domain ideal ratio masks demonstrates a strong performance, it is limited in the sense that the enhancement process is performed only in the magnitude domain, while the phase spectra remain unchanged. Thus, recent studies have been conducted to involve phase spectra in speech enhancement systems. A phase-sensitive mask (PSM) is a T-F mask that implicitly represents phase-related information. However, since the PSM has an unbounded value, the networks are trained to target its truncated values rather than directly estimating it. To effectively train the PSM, we first approximate it to have a bounded dynamic range under the assumption that speech and noise are uncorrelated. We then propose a joint learning algorithm that trains the approximated value through its parameterized variables in order to minimize the inevitable error caused by the truncation process. Specifically, we design a network that explicitly targets three parameterized variables: 1) speech magnitude spectra; 2) noise magnitude spectra; and 3) phase difference of clean to noisy spectra. To further improve the performance, we also investigate how the dynamic range of magnitude spectra controlled by a warping function affects the final performance in joint learning algorithms. Finally, we examined how the proposed additional constraint that preserves the sum of the estimated speech and noise power spectra affects the overall system performance. The experimental results show that the proposed learning algorithm outperforms the conventional learning algorithm with the truncated phase-sensitive approximation.
Jinkyu Lee 0002, Jan Skoglund, Turaj Zakizadeh Shabestary, Hong-Goo Kang
IEEE Signal Process. Lett.4
2018 SVD-Based Adaptive QIM Watermarking on Stereo Audio Signals
abstract
This paper proposes a blind digital audio water- marking algorithm that utilizes the quantization index modulation (QIM) and the singular value decomposition (SVD) of stereo audio signals. Conventional SVD-based blind audio watermarking algorithms lack physical interpretation since the matrix construction method for the input matrix for SVD is heuristically defined. However, in the proposed approach, because the SVD is directly applied to the stereo input signals, the resulting decomposed elements convey a conceptually meaningful inter- pretation of the original audio signal. As the proposed approach effectively utilizes the ratio of singular values, the embedded watermark is highly imperceptible and robust against volumetric scaling attacks; most QIM-based watermarking schemes are weak to these types of attacks. Experimental results under well-known practical attacks, such as compressions, resampling, and various types of signal processing, confirm that the proposed algorithm performs well compared to conventional audio watermarking algorithms.
Min-Jae Hwang, JeeSok Lee, MiSuk Lee, Hong-Goo Kang
IEEE Trans. Multim.4
2017 Perceptual quality and modeling accuracy of excitation parameters in DLSTM-based speech synthesis systems
abstract
This paper investigates how the perceptual quality of the synthesized speech is affected by reconstruction errors in excitation signals generated by a deep learning-based statistical model. In this framework, the excitation signal obtained by an LPC inverse filter is first decomposed into harmonic and noise components using an improved time-frequency trajectory excitation (ITFTE) scheme, then they are trained and generated by a deep long short-term memory (DLSTM)-based speech synthesis system. By controlling the parametric dimension of the ITFTE vocoder, we analyze the impact of the harmonic and noise components to the perceptual quality of the synthesized speech. Both objective and subjective experimental results confirm that the maximum perceptually allowable spectral distortion for the harmonic spectrum of the generated excitation is ~0.08 dB. On the other hand, the absolute spectral distortion in the noise components is meaningless, and only the spectral envelope is relevant to the perceptual quality.
Eunwoo Song, Frank K. Soong, Hong-Goo Kang
ASRU3
2017 Effective Spectral and Excitation Modeling Techniques for LSTM-RNN-Based Speech Synthesis Systems
abstract
In this paper, we report research results on modeling the parameters of an improved time-frequency trajectory excitation (ITFTE) and spectral envelopes of an LPC vocoder with a long short-term memory (LSTM)-based recurrent neural network (RNN) for high-quality text-to-speech (TTS) systems. The ITFTE vocoder has been shown to significantly improve the perceptual quality of statistical parameter-based TTS systems in our prior works. However, a simple feed-forward deep neural network (DNN) with a finite window length is inadequate to capture the time evolution of the ITFTE parameters. We propose to use the LSTM to exploit the time-varying nature of both trajectories of the excitation and filter parameters, where the LSTM is implemented to use the linguistic text input and to predict both ITFTE and LPC parameters holistically. In the case of LPC parameters, we further enhance the generated spectrum by applying LP bandwidth expansion and line spectral frequency-sharpening filters. These filters are not only beneficial for reducing unstable synthesis filter conditions but also advantageous toward minimizing the muffling problem in the generated spectrum. Experimental results have shown that the proposed LSTM-RNN system with the ITFTE vocoder significantly outperforms both similarly configured band aperiodicity-based systems and our best prior DNN-trainecounterpart, both objectively and subjectively.
Eunwoo Song, Frank K. Soong, Hong-Goo Kang
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Improved Time-Frequency Trajectory Excitation Vocoder for DNN-Based Speech Synthesis
Eunwoo Song, Frank K. Soong, Hong-Goo Kang
INTERSPEECH3
2015 Coherent channel based subband multichannel dereverberation
abstract
This paper presents a multichannel dereverberation algorithm that only uses coherent acoustic channels. In the framework of multi-input/output inverse theorem (MINT), the equalization performance varies depending on the length of the input acoustic channels. However, only the portion of observed channel that resemble the true acoustic channel contributes to performance enhancement when measurement error is accounted. Hence, the proposed algorithm derives the frequency dependent viable channel length (VCL) from the coherence analysis of Monte Carlo observations of a single acoustic channel. The VCL of the room impulse response (RIR) is determined by the portion where the stochastic characteristic of multiple observations is highly coherent. Experiments are conducted to compare the equalization performance of the subband MINT algorithm depending on the length of the input RIR. The equalization performance using frequency dependent VCL is as good as the one obtained using the maximum length of the measured channel, while its complexity is significantly reduced.
JeeSok Lee, Sejin Oh, Hong-Goo Kang
ICASSP3
2015 Improved time-frequency trajectory excitation modeling for a statistical parametric speech synthesis system
abstract
This paper proposes an improved time-frequency trajectory excitation (TFTE) modeling method for a statistical parametric speech synthesis system. The proposed approach overcomes the dimensional variation problem of the training process caused by the inherent nature of the pitch-dependent analysis paradigm. By reducing the redundancies of the parameters using predicted average block coefficients (PABC), the proposed algorithm efficiently models excitation, even if its dimension is varied. Objective and subjective test results verify that the proposed algorithm provides not only robustness to the training process but also naturalness to the synthesized speech.
Eunwoo Song, Young-Sun Joo, Hong-Goo Kang
ICASSP3
2015 Systematic integration of acoustic echo canceller and noise reduction modules for voice communication systems
Hyeonjoo Kang, JeeSok Lee, Soonho Baek, Hong-Goo Kang
INTERSPEECH4
2015 Deep neural network-based statistical parametric speech synthesis system using improved time-frequency trajectory excitation model
Eunwoo Song, Hong-Goo Kang
INTERSPEECH2
2015 A Priori SNR Estimation Using Air- and Bone-Conduction Microphones
abstract
This paper proposes an a priori signal-to-noise ratio (SNR) estimator using an air-conduction (AC) and a bone-conduction (BC) microphone. Among various ways of combining AC and BC microphones for speech enhancement, it is shown that the total enhancement performance can be maximized if the BC microphone is utilized for estimating the power spectral density (PSD) of the desired speech signal. Considering the fact that a small deviation in the speech PSD estimation process brings severe spectral distortion, this paper focuses on controlling weighting factors while estimating the a priori SNR with the decision-directed approach framework. The time–frequency varying weighting factor that is determined by taking a minimum mean square error criterion improves the capability of eliminating residual noise and minimizing speech distortion. Since the weighting factors are also adjusted by measuring the usefulness of the AC and BC microphones, the proposed approach is suitable for tracking the parameter even if the characteristic of environment changes rapidly. The simulation results confirm the superiority of the proposed algorithm to conventional algorithms in high noise environments.
Ho Seon Shin, Tim Fingscheidt, Hong-Goo Kang
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Mean normalization of power function based cepstral coefficients for robust speech recognition in noisy environment
abstract
This paper presents the effect of mean normalization to various types of cepstral coefficients for robust speech recognition in noisy environments. Although the cepstral mean normalization (CMN) technique was originally designed to compensate channel distortion, it has also been proved that the CMN also improves recognition accuracy in additive noisy environment. However, no one has yet considered the interaction of CMN with spectral mapping functions required for extracting cepstral features. This paper investigates the impact of CMN to the speech recognition system depending on the types of spectral mapping function by mathematically analyzing the amount of spectral distortion between clean and noisy conditions. The analytic result is also confirmed by comparing the type of recognition error patterns in automatic speech recognition experiment with Aurora 2 database. Experimental results show that the performance improvement by adopting CMN becomes significant if the logarithmic function is replaced with the appropriate setting of fractional power mapping function. Especially, the deletion errors are dramatically reduced.
Soonho Baek, Hong-Goo Kang
ICASSP2
2014 Detecting pathological speech using contour modeling of harmonic-to-noise ratio
abstract
This paper proposes a new feature extraction method for automatically detecting pathological voice in a normal conversation scenario. Unlike conventional approaches that utilize the static harmonic-to-noise ratio (HNR) characteristics of sustained vowel, the proposed method considers the dynamic movements of articulatory organs depending on the types of phonations. Assuming those movements reflect the health status of subjects, the proposed method utilizes the characteristics of HNR contour within a single sentence-level speech signal. Experimental results show that the proposed method reduces the classification error rate by 35.2 % (relative) compared to the conventional method.
Jung-Won Lee, Samuel Kim, Hong-Goo Kang
ICASSP3
2014 Factored adaptation of speaker and environment using orthogonal subspace transforms
abstract
This paper presents a subspace-based acoustic factorization framework to transform-based adaptation in speech recognition. In the proposed method, adaptation transforms are projected onto factor-dependent low-rank subspaces in a way that decouples the combined extrinsic factors affecting the speech signals. Usually, mismatch between the observed speech and the acoustic models is caused by multiple acoustic factors simultaneously, such as the speaker and environment. Data-driven adaptation methods, such as constrained MLLR, compensate for all sources of mismatch jointly. In many scenarios, however, it is highly desirable to separate the sources of mismatch in order to adapt to speaker and environment variability independently. This adds flexibility to the model adaptation framework. For example, a speaker transform obtained in one environment can be reused when the same speaker is in different environments. Or, an environment transform obtained during training, independently of speaker identities, can be applied to a speaker in deployment. One way to achieve this factorization is to construct each set of transforms such that they are orthogonal to each other, so that any change in one acoustic factor keeps other factors intact. The proposed subspace approach provides a straightforward factor analysis framework while allows us to explicitly formulate the independence among the estimated factor transforms. A series of experiments performed on the Aurora 4 corpus validates our approach.
Hyunson Seo, Hong-Goo Kang, Michael L. Seltzer
ICASSP2
2014 A maximum a Posterior-based reconstruction approach to speech bandwidth expansion in noise
abstract
We propose a novel bandwidth expansion algorithm for extending narrowband speech signal to wideband by exploiting segment examples pre-stored in a speaker independent database. Both narrowband and wideband representation of speech signals are pre-stored in the corpus and they are dynamically chopped into variable length segments. Narrowband segments are used dynamically to explain a given narrowband input sentence while the wideband expanded version of the input sentence is constructed correspondingly. The matching process in the narrowband favors a longer segment patch by the chosen Maximum A Posterior (MAP) criterion. As a result, the multiple choices in matching process are significantly reduced with the MAP criterion in decoding. The approach is further generalized to deal with noise corrupted narrowband input signals and the well-known Vector Taylor Series (VTS) noise adaptation algorithm is incorporated into the matching and bandwidth expansion process. A series of experiments is performed to validate the approach on both clean and noise corrupted narrowband speech where both car noise and babble noise corrupted samples are tested.
Hyunson Seo, Hong-Goo Kang, Frank K. Soong
ICASSP2
2014 An Efficient Multichannel Linear Prediction-Based Blind Equalization Algorithm in Near Common Zeros Condition
abstract
This letter proposes an efficient multichannel acoustic channel equalization method under insufficient channel diversity conditions. To overcome an ill-posed problem caused by near common zeros (NCZs) conditions between different channels, a regularization method that restricts the filter norm has been investigated. However, direct application of this method to the linear-predictive multi-input equalization (LIME) method is not effective. To address this situation, this letter puts forth a novel method to disregard the erroneous term of the LIME solution matrix and to increase forced channel diversity (FCD). The accuracy of the proposed equalization filter is compared to that of the conventional regularization method. Experimental results confirm that the NCZs problem can be solved by adopting the proposed methods.
Jae-Mo Yang, Hong-Goo Kang
IEEE Signal Process. Lett.2
2014 Online Speech Dereverberation Algorithm Based on Adaptive Multichannel Linear Prediction
abstract
This paper proposes a real-time acoustic channel equalization method that uses an adaptive multichannel linear prediction technique. In general, multichannel equalization algorithms can eliminate reverberation if they meet the following specific conditions including: the co-primeness between channels and sufficient filter length. It also requires the characteristic of correct channel information, however, it is difficult to estimate accurate acoustic channels in a practical system. The proposed method utilizes a theoretically perfect channel equalization algorithm and considers problems that may arise in the actual system. Linear-predictive multi-input equalization (LIME) is also an appropriate attempt at blind dereverberation by assuring the theoretical basis. However, a huge computational cost is incurred by calculating the large dimensions of a covariance matrix and its inversion. The proposed equalizer is developed as a multichannel linear prediction (MLP) oriented structure with a new formula that is optimized to time-varying acoustical room environments. Moreover, experimental results show that the proposed method works well even if the channel characteristics of each microphone are similar. The results of experiments using various room impulse response (RIR) models, including both the synthesized and real room environments, show that the proposed method is superior to conventional methods.
Jae-Mo Yang, Hong-Goo Kang
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Vector Taylor series based HMM adaptation for generalized cepstrum in noisy environment
abstract
This paper proposes a novel HMM adaptation algorithm for robust automatic speech recognition (ASR) system in noisy environments. The HMM adaptation using vector Taylor series (VTS) significantly improves the ASR performance in noisy environments. Recently, the power normalized cepstral coefficient (PNCC) that replaces a logarithmic mapping function with a power mapping function has been proposed and it is proved that the replacement of the mapping function is robust to additive noise. In this paper, we extend the VTS based approach to the cepstral coefficients obtained by using a power mapping function instead of a logarithmic mapping function. Experimental results indicate that HMM adaptation in the cepstrum obtained by using a power mapping function improves the ASR performance comparing the VTS based conventional approach for mel-frequency cepstral coefficients (MFCCs).
Soonho Baek, Hong-Goo Kang
ASRU2
2013 Enhancement of spectral clarity for HMM-based text-to-speech systems
abstract
This paper proposes a method to enhance the spectral clarity of hidden Markov model (HMM)-based text-to-speech (TTS) systems. A simple way of enhancing spectral clarity is increasing the order of spectral parameters in the speech analysis/synthesis stage, but the method has an inherent statistical modeling problem. The proposed algorithm takes a low-to-high-order spectral parameter mapping approach that adopts low-order parameters for HMM training but does high-order parameters for the actual synthesis step. Various ways of mapping criterion to find appropriate high-order parameters are investigated to further enhance the quality of synthesized speech. Performance evaluation results verify the superiority of the proposed method compared to the conventional one.
Young-Sun Joo, Chi-Sang Jung, Hong-Goo Kang
ICASSP3
2013 A source-filter based adaptive harmonic model and its application to speech prosody modification
JeeSok Lee, Frank K. Soong, Hong-Goo Kang
INTERSPEECH3
2012 Waveform Interpolation-Based Speech Analysis/Synthesis for HMM-Based TTS Systems
abstract
This letter proposes an HMM-based Text-to-Speech (TTS) system using waveform interpolation (WI)-based speech analysis and synthesis. The synthesized speech quality of the proposed system is significantly improved due to adopting an enhanced excitation modeling technique. The decomposition of characteristic waveform (CW) into slowly evolving waveform (SEW) and rapidly evolving waveform (REW) is efficient not only for excitation modeling but also for training process of HMMs. Objective and subjective test results verify the superiority of the proposed approach to conventional ones.
Chi-Sang Jung, Young-Sun Joo, Hong-Goo Kang
IEEE Signal Process. Lett.3
2011 Enhanced long-term predictor for Unified Speech and Audio Coding
abstract
Unified Speech and Audio Coding (USAC) is an emerging MPEG audio standard striving for efficiently representing both speech and music signals even in very low bitrate ranges. The reference codec takes an approach of unifying two state-of-the-art speech and audio coding structures in a single platform. This paper proposes an enhanced long term predictor (eLTP) that effectively utilizes periodic redundancies of inter- and intra- time frames. Experimental results with various types of input signals confirm the superiority of the proposed algorithm compared to the reference codec.
Jeongook Song, Hyen-O Oh, Hong-Goo Kang
ICASSP3
2011 Classification of Fricatives Using Feature Extrapolation of Acoustic-Phonetic Features in Telephone Speech
Jung-Won Lee, Jeung-Yoon Choi, Hong-Goo Kang
INTERSPEECH3
2011 Estimating redundancy information of selected features in multi-dimensional pattern classification
Chi-Sang Jung, Hyunson Seo, Hong-Goo Kang
Pattern Recognit. Lett.3
2011 A Two-Channel Noise Estimator for Speech Enhancement in a Highly Nonstationary Environment
abstract
This paper proposes a two-channel noise estimator for speech enhancement in a highly nonstationary environment. The proposed noise estimator utilizes a spatial filter which has a capability of extracting noise information even in a speech presence region. We exploit a first-order recursion method with time-frequency varying smoothing coefficients to accurately estimate a noise power spectral density (PSD) in both slowly and rapidly varying regions. The smoothing coefficients are determined by measuring the nonstationarity factor of noise, e.g., degree of noise variation. The nonstationarity factor is derived through a statistical assumption of stationary background noise, which does not need any assumption on the type of nonstationary noise. Since the proposed method efficiently estimates the noise PSD both in stationary and nonstationary regions, the enhanced speech obtained by applying the proposed algorithm to the two-channel enhancement system shows superior performance to conventional approaches in various noise environments.
Min-Seok Choi, Hong-Goo Kang
IEEE ACM Trans. Audio Speech Lang. Process.2
2011 Robust Session Variability Compensation for SVM Speaker Verification
abstract
This paper presents an enhanced nuisance attribute projection (NAP) method to improve the performance of speaker verification systems in mismatched train and test conditions. Unlike the conventional NAP training method that does not take any scheme to discriminate the source of nuisance, the proposed method quantitatively estimates the source of nuisance based on the statistics of given background speakers' eigenvalues. The estimated values are used for defining a discriminative weight for each of background speakers and selectively including the statistics of between-class scatter or of within-class scatter from them. Through the scheme, we intend to design a more robust projection matrix which involves less speaker-dependent or speaker-intrinsic variability while including more latent nuisance factors beyond the common within-class scatter of backgrounds. Experimental results on the recent NIST SRE evaluations demonstrate that the proposed algorithms produce consistent improvement over the previous NAP approaches.
Hyunson Seo, Chi-Sang Jung, Hong-Goo Kang
IEEE Trans. Speech Audio Process.3
2011 An Interactive 3-D Audio System With Loudspeakers
abstract
Traditional 3-D audio systems using two loudspeakers often have a limited sweet spot and may suffer from poor performance in reverberant environments. This paper presents a novel binaural 3-D audio system that actively combines head tracking and room modeling into 3-D audio synthesis. The user's head position and orientation are first tracked by a webcam-based 3-D head tracker. The system then improves its robustness to head movement and strong early reflections by incorporating the tracking information and an explicit room model into the binaural synthesis and crosstalk cancellation process. Sensitivity analysis on the room model shows that the method is reasonably robust to modeling errors. Subjective listening tests confirm that the proposed 3-D audio system significantly improves the users' perception and ability for localization.
Myung-Suk Song, Cha Zhang, Dinei A. F. Florêncio, Hong-Goo Kang
IEEE Trans. Multim.4
2010 Binaural loudness based speech reinforcement with a closed-form solution
abstract
This paper addresses a perceptual signal processing to far-end speech signal in communication systems under near-end environmental noise conditions. Based on the binaural perceptual loudness model, the proposed speech reinforcement system achieves better speech quality and clearness. To effectively reflect the noise influence to both ears, the proposed method utilizes a noise level difference between open and receiver side ear. Its computational complexity is also reduced by deriving an approximated closed-form solution while computing frequency-dependent gain factors. Test results confirm that the proposed system significantly enhances the clearness of target speech while maintaining speech quality compared to conventional monaural-based one.
Ho Seon Shin, Min-Seok Choi, Taesu Kim, Hong-Goo Kang
ICASSP4
2010 Personal 3D audio system with loudspeakers
abstract
Traditional 3D audio systems often have a limited sweet spot for the user to perceive 3D effects successfully. In this paper, we present a personal 3D audio system with loudspeakers that has unlimited sweet spots. The idea is to have a camera track the user's head movement, and recompute the crosstalk canceller filters accordingly. As far as the authors are aware of, our system is the first non-intrusive 3D audio system that adapts to both the head position and orientation with six degrees of freedom. The effectiveness of the proposed system is demonstrated with subjective listening tests comparing our system against traditional non-adaptive systems.
Myung-Suk Song, Cha Zhang, Dinei A. F. Florêncio, Hong-Goo Kang
ICME4
2010 A variable frame length and rate algorithm based on the spectral kurtosis measure for speaker verification
Chi-Sang Jung, Kyu Jeong Han, Hyunson Seo, Shri Narayanan, Hong-Goo Kang
INTERSPEECH5
2010 Enhancing loudspeaker-based 3D audio with room modeling
abstract
For many years, spatial (3D) sound using headphones has been widely used in a number of applications. A rich spatial sensation is obtained by using head related transfer functions (HRTF) and playing the appropriate sound through headphones. In theory, loudspeaker audio systems would be capable of rendering 3D sound fields almost as rich as headphones, as long as the room impulse responses (RIRs) between the loudspeakers and the ears are known. In practice, however, obtaining these RIRs is hard, and the performance of loudspeaker based systems is far from perfect. New hope has been recently raised by a system that tracks the user's head position and orientation, and incorporates them into the RIRs estimates in real time. That system made two simplifying assumptions: it used generic HRTFs, and it ignored room reverberation. In this paper we tackle the second problem: we incorporate a room reverberation estimate into the RIRs. Note that this is a nontrivial task: RIRs vary significantly with the listener's positions, and even if one could measure them at a few points, they are notoriously hard to interpolate. Instead, we take an indirect approach: we model the room, and from that model we obtain an estimate of the main reflections. Position and characteristics of walls do not vary with the users' movement, yet they allow to quickly compute an estimate of the RIR for each new user position. Of course the key question is whether the estimates are good enough. We show an improvement in localization perception of up to 32% (i.e., reducing average error from 23.5° to 15.9°).
Myung-Suk Song, Cha Zhang, Dinei A. F. Florêncio, Hong-Goo Kang
MMSP4
2010 Selecting Feature Frames for Automatic Speaker Recognition Using Mutual Information
abstract
In this paper, an information theoretic approach to selecting feature frames for speaker recognition systems is proposed. A conventional approach in which the frame shift is fixed to around half of the frame length may not be the best choice, because the characteristics of the speech signal may rapidly change, especially at phonetic boundaries. Experimental results show that the recognition accuracy increases if the frame interval is directly controlled using phonetic information. By applying these results to the well-known fact that the recognition accuracy is directly correlated with the amount of mutual information, this paper suggests a novel feature frame selection method for speaker recognition. Specifically, feature frames are chosen to have minimum-redundancy within selected feature frames, but maximum-relevancy to speaker models. It is verified by experiments that the proposed method produces consistent improvement, especially in a speaker verification system. It is also robust against variations in acoustic environment.
Chi-Sang Jung, Hong-Goo Kang
IEEE Trans. Speech Audio Process.3
2009 Normalized minimum-redundancy and maximum-relevancy based feature selection for speaker verification systems
abstract
In this paper, an information theoretical approach to select features for speaker recognition systems is proposed. Conventional approaches having a fixed interval of analysis frames are not appropriate to represent dynamically varying characteristics of speech signals. To maximize the speaker-related information varied by the characteristics of speech signals, we propose an information theory based feature selection method where features are selected to have minimum-redundancy with in selected features but maximum-relevancy to training speaker models. Experimental results verify that the proposed method reduces the error rates of speaker verification systems by 27.37 % in NIST 2002 database.
Chi-Sang Jung, Hong-Goo Kang
ICASSP3
2009 On the Study of Noise Allocation for Speech Signal in Low Bit-Rate Audio Coding
abstract
This letter proposes a new masking threshold adjustment method to improve the quality for the speech signals in low bit-rate audio coding. The Enhanced aacPlus (EAAC) audio codec increases the masking threshold of all frequency bands to be suitable for the given encoding rate by considering equal loudness noises only, which is a representative way for implementing the adjustment technique. The proposed method, however, dynamically adjusts the masking threshold of each frequency band based on the energy ratio of each band to the average band energy. More quantization noises are added to formant regions that have relatively large energy ratio values, but less distortion is allowed in spectral valley regions, which eventually helps to enhance perceptual quality for speech signals. The proposed idea reflects the spectral weighting criterion in searching optimal excitation codebooks used in many speech coding algorithms. Simulation results confirm that the proposed method implemented on the EAAC coder improves quality for the speech input signals at the same bit-rate while keeping equivalent quality for music contents.
Chang-Heon Lee, Hyen-O Oh, Hong-Goo Kang
IEEE Signal Process. Lett.3
2008 Designing a unified speech/audio codec by adopting a single channel harmonic source separation module
abstract
This paper propose a unified speech/audio codec by adopting a single channel harmonic separation module as a pre-processor. A modulation frequency analysis method is used for harmonic separation, and the separated components are first encoded by an appropriate codec, e.g. speech codec. The error between input and the encoded signal is recorded by another codec. Though any type of codec can be used for the purpose, we adopt two state-of-the-art international standards (AMR-WB and HE-AAC) to provide an interoperability option. The amount of allocated bits to each stage is controlled by a power ratio of separated harmonic componets to input signal. Subjective listening tests verify the consistency of the proposed method in speech, music and mixed signal inputs.
Sang-Wook Shin, Chang-Heon Lee, Hyen-O Oh, Hong-Goo Kang
ICASSP4
2008 Speech Bandwidth Extension Using Temporal Envelope Modeling
abstract
Speech bandwidth extension (SBE) assumes that high-frequency components of a speech signal, e.g., the frequency band of 47 kHz, can be estimated by parameters extracted from the narrowband signal (04 kHz). Therefore, it is very important to understand the characteristics of the highband signal as well as perceptual cues to represent the highband signal. This letter proposes a new SBE algorithm using a temporal envelope model. The temporal envelope model considers band-limited temporal envelopes as the perceptual cue of the 47 kHz band signal while it deemphasizes the importance of rapidly varying components. To implement the SBE with no additional bits, the proposed method adopts a Gaussian mixture model (GMM) to estimate the temporal envelope of the highband signal from that of the narrowband one. Simulation results confirm that the proposed SBE algorithm shows better perceptual quality than a conventional source-filter model-based approach.
Kyung-Tae Kim, Min-Ki Lee, Hong-Goo Kang
IEEE Signal Process. Lett.3
2007 A Soft-Decision Adaptation Mode Controller for an Efficient Frequency-Domain Generalized Sidelobe Canceller
abstract
In this paper, we propose a new soft-decision adaptation mode controller (SD-AMC) for frequency domain generalized sidelobe canceller (GSC) as a speech enhancement system. Contrarily to conventional systems that update filter coefficients in a hard-decision manner using voice activity detection (VAD), the proposed method flexibly controls the step-sizes of adaptive filters depending on the probability of speech presence in each frequency bin. Therefore, it further improves the system performance for various environments without much consideration on noise type and signal to noise ratio (SNR) of input signal. It also improves the robustness of GSC system by avoiding the miss-classification error by the hard-decision logic. Experimental results with speech recognition systems verify that the SD-AMC shows higher performance than ideally designed hard-decision approaches.
Min-Seok Choi, Chang-Hyun Baik, Young-Cheol Park, Hong-Goo Kang
ICASSP (4)4
2007 Speech quality estimation using packet loss effects in CELP-type speech coders
Min-Ki Lee, Kyung-Tae Kim, Hong-Goo Kang, Dae Hee Youn
INTERSPEECH3
2007 Applying a Speaker-Dependent Speech Compression Technique to Concatenative TTS Synthesizers
abstract
This paper proposes a new speaker-dependent coding algorithm to efficiently compress a large speech database for corpus-based concatenative text-to-speech (TTS) engines while maintaining high fidelity. To achieve a high compression ratio and meet the fundamental requirements of concatenative TTS synthesizers, such as partial segment decoding and random access capability, we adopt a nonpredictive analysis-by-synthesis scheme for speaker-dependent parameter estimation and quantization. The spectral coefficients are quantized by using a memoryless split vector quantization (VQ) approach that does not use frame correlation. Considering that excitation signals of a specific speaker show low intra-variation especially in the voiced regions, the conventional adaptive codebook for pitch prediction is replaced by a speaker-dependent pitch-pulse codebook trained by a corpus of single-speaker speech signals. To further improve the coding efficiency, the proposed coder flexibly combines nonpredictive and predictive type method considering the structure of the TTS system. By applying the proposed algorithm to a Korean TTS system, we could obtain comparable quality to the G.729 speech coder and satisfy all the requirements that TTS system needs. The results are verified by both objective and subjective quality measurements. In addition, the decoding complexity of the proposed coder is around 55% lower than that of G.729 annex A
Chang-Heon Lee, Sung-Kyo Jung, Hong-Goo Kang
IEEE Trans. Speech Audio Process.3
2006 On the Use of Voting Methods for Speaker Identification Based on Various Resolution Filterbanks
abstract
This paper proposes a novel speaker identification system based on score fusion of various resolution filterbanks. The proposed system uses multiple features which are extracted from filterbanks having various spectral resolutions. Each speaker model is constructed by independent feature set, but the system makes final decision by combining the outcome of each model. We introduce several well-known voting methods for decision. Simulation results using TIMIT database show that the proposed score fusion method significantly improves speaker identification performance compared to single model one. Especially, 59.28% of relative improvement is achieved by using a product rule.
Bong-Jin Lee, Sung-Wan Yoon, Hong-Goo Kang, Dae Hee Youn
ICASSP (1)3
2006 An efficient segment-based speech compression technique for hand-held TTS systems
Chang-Heon Lee, Sung-Kyo Jung, Thomas Eriksson, Won-Suk Jun, Hong-Goo Kang
INTERSPEECH5
2006 Performance analysis of various single channel speech enhancement algorithms for automatic speech recognition
Myung-Suk Song, Chang-Heon Lee, Hong-Goo Kang
INTERSPEECH3
2005 An improved estimation of a priori speech absence probability for speech enhancement : in perspective of speech perception
abstract
The purpose of this paper is to improve the perceptual quality of a single channel speech enhancement algorithm using MMSE LSA estimator. The proposed algorithm uses a nonlinear decision rule and an adaptive recursive averaging factor for tracking a priori speech absence probability (SAP) fast. We also introduce one-third of approximated critical bandwidth to efficiently smooth the a priori SAP and final gain term, which successfully eliminates the musical noise without much distortion of signal. The performance of the proposed algorithm is evaluated by performing subjective AB listening tests and measuring spectral distance. Simulation results verify the effectiveness of the proposed algorithm compared to conventional algorithms.
Min-Seok Choi, Hong-Goo Kang
ICASSP (1)2
2005 A noise-robust pitch synchronous feature extraction algorithm for speaker recognition systems
abstract
A noise-robust pitch synchronous feature extraction algorithm for speaker recognition systems is proposed in this paper. Since the pitch synchronous algorithms utilize pitch information, which is meaningful only for periodic segments, we propose a new scheme to deal with non-periodic ones such as unvoiced and noise-corrupted. The experimental results show that the proposed algorithm outperforms the conventional algorithm us- ing fixed length of analysis window in actual identification tasks even in low SNR noisy environments.
Samuel Kim, Sung-Wan Yoon, Thomas Eriksson, Hong-Goo Kang, Dae Hee Youn
INTERSPEECH4
2005 An information-theoretic perspective on feature selection in speaker recognition
abstract
This letter studies feature selection in speaker recognition from an information-theoretic view. We closely tie the performance, in terms of the expected classification error probability, to the mutual information between speaker identity and features. Information theory can then help us to make qualitative statements about feature selection and performance. We study various common features used for speaker recognition, such as mel-warped cepstrum coefficients and various parameterizations of linear prediction coefficients. The theory and experiments give valuable insights in feature selection and performance of speaker-recognition applications.
Thomas Eriksson, Samuel Kim, Hong-Goo Kang, Chungyong Lee
IEEE Signal Process. Lett.3
2005 A fast adaptive-codebook search algorithm for G.723.1 speech coder
abstract
This letter presents a new fast search algorithm for the multitap adaptive codebook used in the G.723.1 standard speech coder. In contrast with the standard method that a closed-loop pitch lag and gains for a fifth-order pitch predictor are searched simultaneously, the proposed algorithm adopts a sequential and restricted approach to determine the parameters. In other words, the proposed scheme first determines a couple of pitch lag candidates using a first-order pitch predictor and then computes the pitch gains of the fifth-order predictor within a restricted search area. Experimental results confirm that the proposed algorithm reduces the total complexity by 30.69% in the encoding process and provides speech quality equivalent to the standard method.
Sung-Kyo Jung, Kyung-Tae Kim, Young-Cheol Park, Hong-Goo Kang
IEEE Signal Process. Lett.4
2004 Improvement issues on transcoding algorithms: for the flexible usage to the various pairs of speech codec
abstract
The paper describes important issues on transcoding between different speech codecs by considering the paradigms of source and target coders. Conventional transcoding algorithms on LSP, pitch and adaptive/fixed codebook conversion are refined with regard to the structure of the coders. In addition, a new perceptual weighting filter, that plays a role in the post-filter and perceptual weighting filter together, is proposed to improve the performance further. The performance of the proposed algorithms is verified in a step-by-step manner with examples of transcoding between AMR, G.723.1 and G.729. By applying the proposed algorithm to the transcoders, the complexity is reduced by about 20-76.88% and quality is also improved compared to conventional approaches.
Jin-Kyu Choi, Chang-Heon Lee, Hong-Goo Kang, Young-Cheol Park, Dae Hee Youn
ICASSP (1)3
2004 A bit-rate/bandwidth scalable speech coder based on ITU-T G.723.1 standard
abstract
The paper presents a new scalable coder based on the ITU-T G.723.1 standard which is one of the most famous speech coders for VoIP applications. In order to support both bit-rate scalability and bandwidth scalability, the proposed coder adopts a split-band approach, where the input signal, sampled at 16 kHz, is decomposed into two equal frequency bands. The lower-band speech is coded with a standard coder, such as the G.723.1 standard. In addition, the low-band enhancement layer for lower-band speech improves the perceptual quality of decoded speech by employing additional coding units based on a cascaded codebook approach. The higher-band signal is encoded using an MDCT-based transform coding scheme. The proposed coder, at a bit-rate of 19.4 kbit/s, provides speech quality comparable to the ITU-T 24 kbit/s G.722.1 coder, while it also has interoperability with G.723.1.
Sung-Kyo Jung, Kyung-Tae Kim, Hong-Goo Kang
ICASSP (1)3
2004 A pitch synchronous feature extraction method for speaker recognition
abstract
The paper presents a novel feature extraction method to improve the performance of speaker identification systems. The proposed feature has the form of a typical conventional feature, Mel frequency cepstral coefficients (MFCC), but a flexible segmentation to reduce spectral mismatch between training and testing processes. Specifically, the length and shift size of the analysis frame are determined by a pitch synchronous method, pitch synchronous MFCC (PSMFCC). To verify the performance of the new feature, we measure the cepstral distortion between training and testing and also perform closed set speaker identification tests. With text-independent and text-dependent experiments, the proposed algorithm provides 44.3% and 26.7% relative improvement, respectively.
Samuel Kim, Thomas Eriksson, Hong-Goo Kang, Dae Hee Youn
ICASSP (1)3
2004 Theory for speaker recognition over IP
Thomas Eriksson, Samuel Kim, Hong-Goo Kang, Chungyong Lee
INTERSPEECH3
2004 Performance analysis of transcoding algorithms in packet-loss environments
Sung-Kyo Jung, Hong-Goo Kang, Dae Hee Youn, Chang-Heon Lee
INTERSPEECH2
2004 On the time variability of vocal tract for speaker recognition
Samuel Kim, Thomas Eriksson, Hong-Goo Kang
INTERSPEECH3
2004 Temporal normalization techniques for transform-type speech coding and application to split-band wideband coders
abstract
In this paper we present an efficient coding method for the upper band(4-7kHz) of wideband(0.5-7kHz) speech coding based on a band-split approach. Due to the impulselike characteristics in upper band signal, it is very difficult to efficiently quantize the signal at low bit-rate when we use transform coding techniques. We propose two temporal normalization techniques, direct temporal energy normalization and frequency domain linear prediction, to reduce the extremely noticeable artifacts. Simulation results show that the proposed algorithm successfully encodes the upper band signal, and the new split-band type wideband coder adopting the proposed technology provides better quality than 56 kbit/s ITU-T G.722 at the bitrate of 20 kbit/s.
Kyung-Tae Kim, Sung-Kyo Jung, MiSuk Lee, Hong-Goo Kang, Dae Hee Youn
INTERSPEECH4
2004 An efficient transcoding algorithm for G.723.1 and G.729A speech coders: interoperability between mobile and IP network
Sung-Wan Yoon, Hong-Goo Kang, Young-Cheol Park, Dae Hee Youn
Speech Commun.2
2003 A cascaded algebraic codebook structure to improve the performance of speech coder
abstract
This paper presents a cascade structure of an algebraic codebook to improve the performance of low bit-rate speech coder. A codeword of an algebraic codebook consists of a set of pulse amplitudes and positions. In general, the amplitude of each pulse is constrained to be either +1 or -1 due to the limitations of bit-rate and complexity. Thus, the performance of the codebook is varied depending on the characteristic of input target vectors. In this paper, we extend the algebraic codebook structure to two stages in order to provide flexible pulse combinations. While all pulses, M, are simultaneously selected in a classical one-stage algebraic codebook, the cascade structure searches the pulses with a two step procedure, i.e., L pulses at the first stage and (M-L) pulses at the second stage. Experiments confirm that our algorithm provides higher quality than the conventional scheme when the total number of pulses is same. In case of assigning 24 pulses per 8-ms subframe, a segmental SNR between target and synthesized signal increases 1.04 dB. In addition, at the same environment, the complexity of fixed codebook search is reduced by about 32%.
Sung-Kyo Jung, Kyoung-Tae Kim, Hong-Goo Kang, Dae Hee Youn
ICASSP (2)3
2003 A packet loss concealment algorithm based on time-scale modification for CELP-type speech coders
abstract
We propose a packet loss concealment algorithm for a code-excited linear prediction (CELP) speech coder. We perform a time-scale modification (TSM) using a waveform similarity overlap-add (WSOLA) technique to reconstruct the excitation signal of the lost or dropped frames. In addition, when a lost frame is classified as a voiced, an adaptive codebook gain and a fixed codebook gain are estimated by a modified gain parameter re-estimation (GRE) technique. By applying these techniques, we can reduce quality degradation of the decoded speech and error propagation effect through the adaptive codebook memory. We apply the proposed scheme to the ITU-T G.729 standard speech coder to evaluate the performance of the proposed method. The perceptual evaluation of speech quality (PESQ) and AB preference tests under various packet loss conditions verify that the proposed algorithm is superior to the concealment algorithm embedded in the G.729.
Moon-Keun Lee, Sung-Kyo Jung, Hong-Goo Kang, Young-Cheol Park, Dae Hee Youn
ICASSP (1)3
2003 Transcoding algorithm for g.723.1 and AMR speech coders: for interoperability between voIP and mobile networks
Sung-Wan Yoon, Jin-Kyu Choi, Hong-Goo Kang, Dae Hee Youn
INTERSPEECH3
2003 Improving the transcoding capability of speech coders
abstract
With the trend of merging various communication networks, a need arises to provide transcoding between different speech coding formats. Presently this means a cross tandem between the two coders in each case. This results in both quality loss and extra delay. A possible alternative is using a bitstream mapping approach that directly converts parameter values. For several standard coders having a similar coding structure, it should be possible to generate comparable or better quality without adding much delay or complexity. This paper proposes a bitstream mapping method between ITU-T Recommendation G.729 and TIA IS-641. Informal listening tests and the perceptual subjective quality measure (PSQM) scores show that the proposed method has better quality than the cross tandem method, while it has at least 5 ms less delay and six times less computation.
Hong-Goo Kang, Hong Kook Kim, Richard V. Cox
IEEE Trans. Multim.1
2002 A phase generation method for speech reconstruction from spectral envelope and pitch intervals
abstract
In this paper, we propose a new speech reconstruction method from spectral envelope and pitch intervals, which is applicable to the network side of a distributed speech recognition system as a play-back function. The spectral envelope of speech is represented as a set of mel-frequency cepstral coefficients that is a well-known recognition parameter. First, a sinusoidal synthesis with a zero-phase model is used to obtain a pitch-based waveform. To enhance the naturalness of the speech we replace the zero phase information with pre-stored linear and random codebooks. The ultimate phase information is determined depending on the energy ratio between linear and random components. Unlike the classic low bit-rate speech coding, however, the energy ratio is estimated in the decoding stage from a time-frequency filter applied to the pitch-based synthesized signal. Thus, the phase information is not a feature parameter from the encoder side. The proposed phase generation method uses the knowledge that pitch variation is a main cause of the mixed characteristics in speech signals. An informal listening test verifies that the quality of the proposed method is much better than that of the synthetic quality.
Hong-Goo Kang, Hong Kook Kim
ICASSP1
2002 An adaptive short-term postfilter based on pseudo-cepstral representation of line spectral frequencies
Hong Kook Kim, Hong-Goo Kang
Speech Commun.2
2001 A candidate for the ITU-T 4 kbit/s speech coding standard
abstract
This paper presents the 4 kbit/s speech coding candidate submitted by AT&T, Conexant, Deutsche Telekom, France Telecom, Matsushita, and NTT for the ITU-T 4 kbit/s selection phase. The algorithm was developed jointly based on the qualification version of Conexant. This paper focuses on the development carried out during the collaboration in order to narrow the gap to the requirements in an attempt to provide toll quality at 4 kbit/s. This objective is currently being verified in independent subjective tests coordinated by ITU-T and carried out in multiple languages. Subjective tests carried out during the development indicate that the collaboration work has been successful in improving the quality, and that meeting a majority of the requirements in the extensive selection phase test is a realistic goal.
Jes Thyssen, Adil Benyassine, Eyal Shlomot, Carlo Murgia, Huan-yu Su, Kazunori Mano, Yusuke Hiwasaki, Hiroyuki Ehara, Kazutoshi Yasunaga, Claude Lamblin, Balázs Kövesi, Joachim Stegmann, Hong-Goo Kang
ICASSP14
2001 Acoustic feature compensation based on decomposition of speech and noise for ASR in noisy environments
abstract
This paper presents a set of acoustic feature pre–processing techniques that are applied to improving automatic speech recognition (ASR) performance on the Aurora 2 noisy speech recognition task. The principal contribution of this paper is an approach for cepstrum domain feature compensation in ASR which is motivated by techniques for decomposing speech and noise that were originally developed for noisy speech enhancement. This approach is applied in combination with other feature compensation algorithms to compensating ASR features obtained from a mel–filterbank cepstrum coefficient (MFCC) front–end. Performance comparisons are made with respect to the application of the minimum mean squared error log spectral amplitude estimator (MMSE–LSA) based speech enhancement algorithm prior to feature analysis. An experimental study is presented where the feature compensation approaches described in the paper are found to reduce ASR word error rate by as much as 31% relative to uncompensated features under simulated environmental and channel mismatched conditions.
Hong Kook Kim, Richard C. Rose, Hong-Goo Kang
INTERSPEECH3
2001 A frame erasure concealment algorithm based on gain parameter re-estimation for CELP coders
abstract
In this paper, we propose a frame erasure concealment algorithm based on reestimating gain parameters for a code-excited linear prediction (CELP) coder. When a frame is detected as being erased, the coding parameters, especially the adaptive codebook gain and fixed codebook gain, of the erased and subsequent frames, are reestimated by a gain-matching procedure. By doing this, we can reduce the abrupt change caused in the decoded excitation signal by a simple scaling down procedure. We have applied this technique to the IS-63-1 speech coder and found that the proposed algorithm improves the speech quality under various channel conditions compared with the conventional extrapolation-based concealment algorithm.
Hong Kook Kim, Hong-Goo Kang
IEEE Signal Process. Lett.2
2000 Low-rate quantization of spectrum parameters
abstract
In this paper, we generalize the standard blockwise linear predictive (LP) coding by introducing low-pass filtering and downsampling of the LPC vectors at the encoder side, accompanied by interpolation at the decoder. Several concepts in LP coding, such as overlapping frames, interpolation and long analysis frames, can be described in the proposed framework. We also note that the proposed methods are in agreement with previous work on spectral dynamics, which indicates that spectral dynamics is (at least) as important as spectral distortion for speech quality. We have applied the proposed method to low-rate quantization of spectrum parameters. Objective and subjective performance of the new approach is compared to a standard blockwise LP coding, at various rates. We show that both the subjective and objective quality of the new method compare favorably with standard methods. Even at as low rate as 500 bit/s, the proposed method has a good quality.
Thomas Eriksson, Hong-Goo Kang, Per Hedelin
ICASSP2
1999 Pitch quantization in low bit-rate speech coding
abstract
This paper describes a new pitch quantization method for low bit-rate speech coding systems. The logarithm of the pitch period is quantized in a combination of two uniform quantizers, one working directly on logarithmic pitch values and the other working on the difference between current and previous logarithmic pitch. The best of the two output values is transmitted to the receiver. This scheme can exploit: both redundancy in the signal and properties of the ear to achieve an efficient quantization. Listening tests show that the proposed scheme allows the pitch parameter to be quantized using 4 bits, with no degradation in audible quality.
Thomas Eriksson, Hong-Goo Kang
ICASSP2
1999 Phase adjustment in waveform interpolation
abstract
This paper describes a method of improving the quality of the waveform interpolation (WI) speech coder by adjustment of the phase information. In WI, a slowly-evolving waveform (SEW) and a rapidly-evolving waveform (REW) represent the periodic and the non-periodic part of the signal. The phase of the synthesized signal is determined by the SEW and REW, and thus the correct quantization of these parameters is important to producing natural speech quality. A method is described, whereby the phase of the synthesized signal is adjusted by modifying the quantized REW spectrum as a function of the fundamental frequency. This essentially attempts to correct the discrepancies in phase that arise due to variation in pitch and also accounts for the difference in noise sensitivity between female and male speech. The overall effect would be the same if multiple codebooks (depending on pitch) were used to code the REW spectrum. Experimental results confirm that the new method results in significantly improved performance.
Hong-Goo Kang, Dipanjan Sen
ICASSP1
1999 Low delay analysis/synthesis schemes for joint speech enhancement and low bit rate speech coding
abstract
A corpus of spontaneous route descriptions was collected from 8 speakers (5 males and 3 females). The corpus was labelled according to the ToBI standard and a discourse analysis was completed. Four discourse tour acts were identified and these were found to occur mainly in non-embedded linear sequences. In general, the intonation of each route description was characterised by a single intonational phrase containing many intermediate phrases. There was a tendency for boundaries between tour acts and intermediate phrases to coincide, but there is usually more than one intermediate phrase to one tour act.
Rainer Martin 0001, Hong-Goo Kang, Richard V. Cox
EUROSPEECH2
1998 Quantization of the spectral envelope for sinusoidal coders
abstract
In an effort to efficiently code the spectral envelope of speech signals for wideband speech coding based on sinusoidal models, a robust computation of the discrete cepstrum coefficients and their quantization is investigated. A parameterization of the spectral envelope has been proposed which is based on discrete cepstral coefficients using regularization techniques. This paper presents an efficient quantization scheme for these coefficients in order to use them in applications like speech coding. We present results which show a 35% reduction in the bit rate when compared to simple scalar quantization. To verify the efficiency of the proposed quantization schemes, informal listening tests were performed in the context of a sinusoidal coder.
Thomas Eriksson, Hong-Goo Kang, Yannis Stylianou
ICASSP2
1997 A 3 channel digital CVSD bit-rate conversion system using a general purpose DSP
Yong-Soo Choi, Hong-Goo Kang, Sung-Youn Kim, Young-Cheol Park, Dae Hee Youn
EUROSPEECH2
1997 Improved regular pulse VSELP coding of speech at low bit-rates
Yong-Soo Choi, Hong-Goo Kang, Jae-Ha Yoo, Dae Hee Youn
EUROSPEECH2
1996 A fast VSELP speech coder based on mutually orthonormal regular pulse vectors
abstract
A new vector sum excited linear prediction (VSELP) speech coder employing mutually orthonormal regular pulse vectors is presented. Since the algorithm uses an efficient vector-sum codebook consisting of mutually orthonormal regular pulse basis vectors, computational load for codebook search can be significantly reduced, while the reconstructed speech quality is equivalent to that of the conventional VSELP. The method, referred to as regular pulse VSELP (RP-VSELP), employs the Gram-Schmidt (GS) procedure to orthonormalize the basis vectors designed in a form of regular pulses. To enhance the SNR performance, the basis vectors are optimized using an iterative closed-loop training process.
Yong-Soo Choi, Hong-Goo Kang, Dae Hee Youn
ICASSP2
1995 A low bit-rate speech coder using the perceptual properties of the human ear
Hong-Goo Kang, Jeong Tae Seo, Il-Whan Cha, Dae Hee Youn
EUROSPEECH1