Biing-Hwang Juang

dblp:99/2494 · also Biing-Hwang Fred Juang, Fred Juang · DBLP profile ↗
← Back
197ranked-venue papers
20as first author
7since 2021 · last 2023
0000-0002-5773-5679ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 149 · 11 first-author · 5 since 2021Artificial intelligence and machine learning · 73 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorComputer networks · 5Systems, architecture and hardware · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Theory of computation · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2023 Towards Zero-Shot Multilingual Transfer for Code-Switched Responses
abstract
Ting-Wei Wu, Changsheng Zhao, Ernie Chang, Yangyang Shi, Pierce Chuang, Vikas Chandra, Biing Juang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ting-Wei Wu, Changsheng Zhao 0002, Ernie Chang, Yangyang Shi, Pierce Chuang, Vikas Chandra, Biing-Hwang Juang
ACL (1)7
2023 Choice Fusion As Knowledge For Zero-Shot Dialogue State Tracking
abstract
With the demanding need for deploying dialogue systems in new domains with less cost, zero-shot dialogue state tracking (DST), which tracks user’s requirements in task-oriented dialogues without training on desired domains, draws attention increasingly. Although prior works have leveraged question-answering (QA) data to reduce the need for in-domain training in DST, they fail to explicitly model knowledge transfer and fusion for tracking dialogue states. To address this issue, we propose CoFunDST, which is trained on domain-agnostic QA datasets and directly uses candidate choices of slot-values as knowledge for zero-shot dialogue-state generation, based on a T5 pre-trained language model. Specifically, CoFunDST selects highly-relevant choices to the reference context and fuses them to initialize the decoder to constrain the model outputs. Our experimental results show that our proposed model achieves outperformed joint goal accuracy compared to existing zero-shot DST approaches in most domains on the MultiWOZ 2.1. Extensive analyses demonstrate the effectiveness of our proposed approach for improving zero-shot DST learning from QA.
Ruolin Su, Jingfeng Yang 0001, Ting-Wei Wu, Biing-Hwang Juang
ICASSP4
2022 Knowledge Augmented Bert Mutual Network in Multi-Turn Spoken Dialogues
abstract
Modern spoken language understanding (SLU) systems rely on sophisticated semantic notions revealed in single utterances to detect intents and slots. However, they lack the capability of modeling multi-turn dynamics within a dialogue particularly in long-term slot contexts. Without external knowledge, depending on limited linguistic legitimacy within a word sequence may overlook deep semantic information across dialogue turns. In this paper, we propose to equip a BERT-based joint model with a knowledge attention module to mutually leverage dialogue contexts between two SLU tasks. A gating mechanism is further utilized to filter out irrelevant knowledge triples and to circumvent distracting comprehension. Experimental results in two complicated multi-turn dialogue datasets have demonstrate by mutually modeling two SLU tasks with filtered knowledge and dialogue contexts, our approach has considerable improvements compared with several competitive baselines.
Ting-Wei Wu, Biing-Hwang Juang
ICASSP2
2022 Induce Spoken Dialog Intents via Deep Unsupervised Context Contrastive Clustering
Ting-Wei Wu, Biing-Hwang Juang
INTERSPEECH2
2021 A Label-Aware BERT Attention Network for Zero-Shot Multi-Intent Detection in Spoken Language Understanding
abstract
With the early success of query-answer assistants such as Alexa and Siri, research attempts to expand system capabilities of handling service automation are now abundant.However, preliminary systems have quickly found the inadequacy in relying on simple classification techniques to effectively accomplish the automation task.The main challenge is that the dialogue often involves complexity in user's intents (or purposes) which are multiproned, subject to spontaneous change, and difficult to track.Furthermore, public datasets have not considered these complications and the general semantic annotations are lacking which may result in zero-shot problem.Motivated by the above, we propose a Label-Aware BERT Attention Network (LABAN) for zeroshot multi-intent detection.We first encode input utterances with BERT and construct a label embedded space by considering embedded semantics in intent labels.An input utterance is then classified based on its projection weights on each intent embedding in this embedded space.We show that it successfully extends to few/zero-shot setting where part of intent labels are unseen in training data, by also taking account of semantics in these unseen intent labels.Experimental results show that our approach is capable of detecting many unseen intent labels correctly.It also achieves the state-of-the-art performance on five multiintent datasets in normal cases.
Ting-Wei Wu, Ruolin Su, Biing-Hwang Juang
EMNLP (1)3
2021 Act-Aware Slot-Value Predicting in Multi-Domain Dialogue State Tracking
abstract
As an essential component in task-oriented dialogue systems, dialogue state tracking (DST) aims to track human-machine interactions and generate state representations for managing the dialogue. Representations of dialogue states are dependent on the domain ontology and the user's goals. In several task-oriented dialogues with a limited scope of objectives, dialogue states can be represented as a set of slot-value pairs. As the capabilities of dialogue systems expand to support increasing naturalness in communication, incorporating dialogue act processing into dialogue model design becomes essential. The lack of such consideration limits the scalability of dialogue state tracking models for dialogues having specific objectives and ontology. To address this issue, we formulate and incorporate dialogue acts, and leverage recent advances in machine reading comprehension to predict both categorical and non-categorical types of slots for multi-domain dialogue state tracking. Experimental results show that our models can improve the overall accuracy of dialogue state tracking on the MultiWOZ 2.1 dataset, and demonstrate that incorporating dialogue acts can guide dialogue state design for future task-oriented dialogue systems.
Ruolin Su, Ting-Wei Wu, Biing-Hwang Juang
Interspeech3
2021 A Context-Aware Hierarchical BERT Fusion Network for Multi-Turn Dialog Act Detection
abstract
The success of interactive dialog systems is usually associated with the quality of the spoken language understanding (SLU) task, which mainly identifies the corresponding dialog acts and slot values in each turn.By treating utterances in isolation, most SLU systems often overlook the semantic context in which a dialog act is expected.The act dependency between turns is nontrivial and yet critical to the identification of the correct semantic representations.Previous works with limited context awareness have exposed the inadequacy of dealing with complexity in multiproned user intents, which are subject to spontaneous change during turn transitions.In this work, we propose to enhance SLU in multi-turn dialogs, employing a context-aware hierarchical BERT fusion Network (CaBERT-SLU) to not only discern context information within a dialog but also jointly identify multiple dialog acts and slots in each utterance.Experimental results show that our approach reaches new state-of-the-art (SOTA) performances in two complicated multi-turn dialogue datasets with considerable improvements compared with previous methods, which only consider single utterances for multiple intents and slot filling.
Ting-Wei Wu, Ruolin Su, Biing-Hwang Juang
Interspeech3
2020 Resource Allocation based on Graph Neural Networks in Vehicular Communications
abstract
In this article, we investigate spectrum allocation in vehicle-to-everything (V2X) network. We first express the V2X network into a graph, where each vehicle-to-vehicle (V2V) link is a node in the graph. We apply a graph neural network (GNN) to learn the low-dimensional feature of each node based on the graph information. According to the learned feature, multi-agent reinforcement learning (RL) is used to make spectrum allocation. Deep Q-network is utilized to learn to optimize the sum capacity of the V2X network. Simulation results show that the proposed allocation scheme can achieve near-optimal performance.
Ziyan He, Liang Wang 0014, Hao Ye 0004, Geoffrey Ye Li, Biing-Hwang Juang
GLOBECOM5
2020 Deep Learning based Semantic Communications: An Initial Investigation
abstract
Recently, deep learned enabled end-to-end (E2E) communication systems have been developed to merge all physical layer blocks in the traditional communication systems, which makes joint transceiver optimization possible. Powered by deep learning, natural language processing (NLP) has achieved great success in analyzing and understanding large amounts of language texts. Inspired by research results in both areas, we aim to provide a new view on communication systems from the semantic level. Particularly, we propose a deep learning based semantic communication system, named DeepSC, for text transmission. Based on the Transformer, the DeepSC aims at maximizing the system capacity and minimizing the semantic errors by recovering the meaning of sentences, rather than bit- or symbol-errors in traditional communications. Compared with the traditional communication system without considering semantic information exchange, the proposed DeepSC is more robust to channel variation and can achieve better performance, especially in the low signal-to-noise ratio (SNR) regime, as demonstrated by the extensive simulation results.
Huiqiang Xie, Zhijin Qin, Geoffrey Ye Li, Biing-Hwang Juang
GLOBECOM4
2020 Deep Over-the-Air Computation
abstract
As an efficient data fusion method, over-the-air computation integrates computation and communication by exploiting the superposition property of multiple access channels. In this paper, a framework on deep learning enabled over-the-air computation is proposed, where both the pre-processing and post-processing functions are represented by deep neural networks (DNNs). In this way, the over-the-air computation can approximate any function via learning through the data. The deep over-the-air framework is useful to a variety of machine learning applications on the Internet-of-Things (IoT). The experiments on distribution regression and anomaly detection have shown the effectiveness of the proposed method.
Hao Ye 0004, Geoffrey Ye Li, Biing-Hwang Juang
GLOBECOM3
2020 Bilinear Convolutional Auto-encoder based Pilot-free End-to-end Communication Systems
abstract
Recently, deep learning based end-to-end communication systems have been developed, where both the transmitter and the receiver are represented as deep neural networks (DNN) and an end-to-end loss is optimized directly. In this paper, we address the effects of the more general wireless channels to the end-to-end framework. We formulate this problem as training a deep auto-encoder system with an adversarial convolutional layer and propose a training procedure with mini-batches of input samples and channels. Instead of using pilots to explicitly estimate the unknown channel, the auto-encoder learns to address the channel effects without any pilot information. In particular, the receiver contains two modules, designed for channel information extraction and data recovery, respectively. The features obtained from the channel information extraction module are combined with received signals by a bilinear production and then processed by the data recovery module to reconstruct the original input data. The experimental results show a performance improvement compared with the traditional methods in commonly seen wireless channels, including frequency-selective channels and multi-input multi-output (MIMO) channels.
Hao Ye 0004, Geoffrey Ye Li, Biing-Hwang Juang
ICC3
2020 Deep Learning-Based End-to-End Wireless Communication Systems With Conditional GANs as Unknown Channels
abstract
In this article, we develop an end-to-end wireless communication system using deep neural networks (DNNs), where DNNs are employed to perform several key functions, including encoding, decoding, modulation, and demodulation. However, an accurate estimation of instantaneous channel transfer function, i.e., channel state information (CSI), is needed in order for the transmitter DNN to learn to optimize the receiver gain in decoding. This is very much a challenge since CSI varies with time and location in wireless communications and is hard to obtain when designing transceivers. We propose to use a conditional generative adversarial net (GAN) to represent channel effects and to bridge the transmitter DNN and the receiver DNN so that the gradient of the transmitter DNN can be back-propagated from the receiver DNN. In particular, a conditional GAN is employed to model the channel effects in a data-driven way, where the received signal corresponding to the pilot symbols is added as a part of the conditioning information of the GAN. To address the curse of dimensionality when the transmit symbol sequence is long, convolutional layers are utilized. From the simulation results, the proposed method is effective on additive white Gaussian noise (AWGN) channels, Rayleigh fading channels, and frequency-selective channels, which opens a new door for building data-driven DNNs for end-to-end communication systems.
Hao Ye 0004, Le Liang, Geoffrey Ye Li, Biing-Hwang Juang
IEEE Trans. Wirel. Commun.4
2018 Speaker-Invariant Training Via Adversarial Learning
abstract
We propose a novel adversarial multi-task learning scheme, aiming at actively curtailing the inter-talker feature variability while maximizing its senone discriminability so as to enhance the performance of a deep neural network (DNN) based ASR system. We call the scheme speaker-invariant training (SIT). In SIT, a DNN acoustic model and a speaker classifier network are jointly optimized to minimize the senone (tied triphone state) classification loss, and simultaneously mini-maximize the speaker classification loss. A speaker-invariant and senone-discriminative deep feature is learned through this adversarial multi-task learning. With SIT, a canonical DNN acoustic model with significantly reduced variance in its output probabilities is learned with no explicit speaker-independent (SI) transformations or speaker-specific representations used in training or testing. Evaluated on the CHiME-3 dataset, the SIT achieves 4.99% relative word error rate (WER) improvement over the conventional SI acoustic model. With additional unsupervised speaker adaptation, the speaker-adapted (SA) SIT model achieves 4.86% relative WER gain over the SA SI acoustic model.
Zhong Meng, Jinyu Li 0001, Zhuo Chen 0006, Yang Zhao 0002, Vadim Mazalov, Yifan Gong 0001, Biing-Hwang Juang
ICASSP7
2018 Adversarial Teacher-Student Learning for Unsupervised Domain Adaptation
abstract
The teacher-student (T/S) learning has been shown effective in unsupervised domain adaptation [1]. It is a form of transfer learning, not in terms of the transfer of recognition decisions, but the knowledge of posteriori probabilities in the source domain as evaluated by the teacher model. It learns to handle the speaker and environment variability inherent in and restricted to the speech signal in the target domain without proactively addressing the robustness to other likely conditions. Performance degradation may thus ensue. In this work, we advance T/S learning by proposing adversarial T/S learning to explicitly achieve condition-robust unsupervised domain adaptation. In this method, a student acoustic model and a condition classifier are jointly optimized to minimize the Kullback-Leibler divergence between the output distributions of the teacher and student models, and simultaneously, to min-maximize the condition classification loss. A condition-invariant deep feature is learned in the adapted student model through this procedure. We further propose multi-factorial adversarial T/S learning which suppresses condition variabilities caused by multiple factors simultaneously. Evaluated with the noisy CHiME-3 test set, the proposed methods achieve relative word error rate improvements of 44.60% and 5.38%, respectively, over a clean source model and a strong T/S learning baseline model.
Zhong Meng, Jinyu Li 0001, Yifan Gong 0001, Biing-Hwang Juang
ICASSP4
2018 Cycle-Consistent Speech Enhancement
abstract
Feature mapping using deep neural networks is an effective approach for single-channel speech enhancement. Noisy features are transformed to the enhanced ones through a mapping network and the mean square errors between the enhanced and clean features are minimized. In this paper, we propose a cycle-consistent speech enhancement (CSE) in which an additional inverse mapping network is introduced to reconstruct the noisy features from the enhanced ones. A cycle-consistent constraint is enforced to minimize the reconstruction loss. Similarly, a backward cycle of mappings is performed in the opposite direction with the same networks and losses. With cycle-consistency, the speech structure is well preserved in the enhanced features while noise is effectively reduced such that the feature-mapping network generalizes better to unseen data. In cases where only unparalleled noisy and clean data is available for training, two discriminator networks are used to distinguish the enhanced and noised features from the clean and noisy ones. The discrimination losses are jointly optimized with reconstruction losses through adversarial multi-task learning. Evaluated on the CHiME-3 dataset, the proposed CSE achieves 19.60% and 6.69% relative word error rate improvements respectively when using or without using parallel clean and noisy speech data.
Zhong Meng, Jinyu Li 0001, Yifan Gong 0001, Biing-Hwang Juang
INTERSPEECH4
2018 Adversarial Feature-Mapping for Speech Enhancement
abstract
Feature-mapping with deep neural networks is commonly used for single-channel speech enhancement, in which a feature-mapping network directly transforms the noisy features to the corresponding enhanced ones and is trained to minimize the mean square errors between the enhanced and clean features. In this paper, we propose an adversarial feature-mapping (AFM) method for speech enhancement which advances the feature-mapping approach with adversarial learning. An additional discriminator network is introduced to distinguish the enhanced features from the real clean ones. The two networks are jointly optimized to minimize the feature-mapping loss and simultaneously mini-maximize the discrimination loss. The distribution of the enhanced features is further pushed towards that of the clean features through this adversarial multi-task training. To achieve better performance on ASR task, senone-aware (SA) AFM is further proposed in which an acoustic model network is jointly trained with the feature-mapping and discriminator networks to optimize the senone classification loss in addition to the AFM losses. Evaluated on the CHiME-3 dataset, the proposed AFM achieves 16.95% and 5.27% relative word error rate (WER) improvements over the real noisy data and the feature-mapping baseline respectively and the SA-AFM achieves 9.85% relative WER improvement over the multi-conditional acoustic model.
Zhong Meng, Jinyu Li 0001, Yifan Gong 0001, Biing-Hwang Juang
INTERSPEECH4
2017 Minimum Semantic Error Cost Training of Deep Long Short-Term Memory Networks for Topic Spotting on Conversational Speech
abstract
The topic spotting performance on spontaneous conversational speech can be significantly improved by operating a support vector machine with a latent semantic rational kernel (LSRK) on the decoded word lattices (i.e., weighted finite-state transducers) of the speech [1]. In this work, we propose the minimum semantic error cost (MSEC) training of a deep bidirectional long short-term memory (BLSTM)-hidden Markov model acoustic model for generating lattices that are semantically accurate and are better suited for topic spotting with LSRK. With the MSEC training, the expected semantic error cost of all possible word sequences on the lattices is minimized given the reference. The word-word semantic error cost is first computed from either the latent semantic analysis or distributed vector-space word representations learned from the recurrent neural networks and is then accumulated to form the expected semantic error cost of the hypothesized word sequences. The proposed method achieves 3.5%-4.5% absolute topic classification accuracy improvement over the baseline BLSTM trained with cross-entropy on Switchboard-1 Release 2 dataset.
Zhong Meng, Biing-Hwang Juang
INTERSPEECH2
2017 Non-Uniform MCE Training of Deep Long Short-Term Memory Recurrent Neural Networks for Keyword Spotting
abstract
It has been shown in [1, 2] that improved performance can be achieved by formulating the keyword spotting as a non-uniform error automatic speech recognition problem. In this work, we discriminatively train a deep bidirectional long short-term memory (BLSTM) – hidden Markov model (HMM) based acoustic model with non-uniform boosted minimum classification error (BMCE) criterion which imposes more significant error cost on the keywords than those on the non-keywords. By introducing the BLSTM, the context information in both the past and the future are stored and updated to predict the desired output and the long-term dependencies within the speech signal are well captured. With non-uniform BMCE objective, the BLSTM is trained so that the recognition errors related to the keywords are remarkably reduced. The BLSTM is optimized using backpropagation through time and stochastic gradient descent. The keyword spotting system is implemented within weighted finite state transducer framework. The proposed method achieves 5.49% and 7.37% absolute figure-of-merit improvements respectively over the BLSTM and the feedforward deep neural network baseline systems trained with cross-entropy criterion for the keyword spotting task on Switchboard-1 Release 2 dataset.
Zhong Meng, Biing-Hwang Juang
INTERSPEECH2
2017 A comparative study of noise estimation algorithms for nonlinear compensation in robust speech recognition
Yong Zhao 0008, Biing-Hwang Juang
Speech Commun.2
2016 Non-Uniform Boosted MCE Training of Deep Neural Networks for Keyword Spotting
abstract
Keyword spotting can be formulated as a non-uniform error automatic speech recognition (ASR) problem. It has been demonstrated [1] that this new formulation with the nonuniform MCE training technique can lead to improved system performance in keyword spotting applications. In this paper, we demonstrate that deep neural networks (DNNs) can be successfully trained on the non-uniform minimum classification error (MCE) criterion which weighs the errors on keywords much more significantly than those on non-keywords in an ASR task. The integration with a DNN-HMM system enables modeling of multi-frame distributions, which conventional systems find difficult to accomplish. To further improve the performance, more confusable data is generated by boosting the likelihood of the sentences that have more errors. The keyword spotting system is implemented within a weighted finite state transducer (WFST) framework and the DNN is optimized using standard backpropagation and stochastic gradient decent. We evaluate the performance of the proposed framework on a large vocabulary spontaneous conversational telephone speech dataset (Switchboard-1 Release 2). The proposed approach achieves an absolute figure of merit improvement of 3.65% over the baseline system.
Zhong Meng, Biing-Hwang Juang
INTERSPEECH2
2016 Statistical Modeling of Speaker's Voice with Temporal Co-Location for Active Voice Authentication
abstract
Active voice authentication (AVA) is a new mode of talker authentication, in which the authentication is performed continuously on very short segments of the voice signal, which may have instantaneously undergone change of talker. AVA is necessary in providing real-time monitoring of a device authorized for a particular user. The authentication test thus cannot rely on a long history of the voice data nor any past decisions. Most conventional voice authentication techniques that operate on the assumption that the entire test utterance is from only one talker with a claimed identity (including i-vector) fail to meet this stringent requirement. This paper presents a different signal modeling technique, within a conditional vector-quantization framework and with matching short-time statistics that take into account the co-located speech codes to meet the new challenge. As one variation, the temporally co-located VQ (TC-VQ) associates each codeword with a set of Gaussian mixture models to account for the co-located distributions and a temporally colocated hidden Markov model (TC-HMM) is built upon the TCVQ. The proposed technique achieves an window-based equal error rate in the range of 3-5% and a relative gain of 4-25% over a baseline system using traditional HMMs on the AVA database.
Zhong Meng, Biing-Hwang Juang
INTERSPEECH2
2016 Air-Writing Recognition - Part I: Modeling and Recognition of Characters, Words, and Connecting Motions
abstract
Air-writing refers to writing of linguistic characters or words in a free space by hand or finger movements. Air-writing differs from conventional handwriting; the latter contains the pen-up-pen-down motion, while the former lacks such a delimited sequence of writing events. We address air-writing recognition problems in a pair of companion papers. In Part I, recognition of characters or words is accomplished based on six-degree-of-freedom hand motion data. We address air-writing on two levels: motion characters and motion words. Isolated air-writing characters can be recognized similar to motion gestures although with increased sophistication and variability. For motion word recognition in which letters are connected and superimposed in the same virtual box in space, we build statistical models for words by concatenating clustered ligature models and individual letter models. A hidden Markov model is used for air-writing modeling and recognition. We show that motion data along dimensions beyond a 2-D trajectory can be beneficially discriminative for air-writing recognition. We investigate the relative effectiveness of various feature dimensions of optical and inertial tracking signals and report the attainable recognition performance correspondingly. The proposed system achieves a word error rate of 0.8% for word-based recognition and 1.9% for letter-based recognition. We also subjectively and objectively evaluate the effectiveness of air-writing and compare it with text input using a virtual keyboard. The words per minute of air-writing and virtual keyboard are 5.43 and 8.42, respectively.
Mingyu Chen 0003, Ghassan Al-Regib, Biing-Hwang Juang
IEEE Trans. Hum. Mach. Syst.3
2016 Air-Writing Recognition - Part II: Detection and Recognition of Writing Activity in Continuous Stream of Motion Data
abstract
Air-writing refers to writing of characters or words in the free space by hand or finger movements. We address air-writing recognition problems in two companion papers. Part 2 addresses detecting and recognizing air-writing activities that are embedded in a continuous motion trajectory without delimitation. Detection of intended writing activities among superfluous finger movements unrelated to letters or words presents a challenge that needs to be treated separately from the traditional problem of pattern recognition. We first present a dataset that contains a mixture of writing and nonwriting finger motions in each recording. The LEAP from Leap Motion is used for marker-free and glove-free finger tracking. We propose a window-based approach that automatically detects and extracts the air-writing event in a continuous stream of motion data, containing stray finger movements unrelated to writing. Consecutive writing events are converted into a writing segment. The recognition performance is further evaluated based on the detected writing segment. Our main contribution is to build an air-writing system encompassing both detection and recognition stages and to give insights into how the detected writing segments affect the recognition result. With leave-one-out cross validation, the proposed system achieves an overall segment error rate of 1.15% for word-based recognition and 9.84% for letter-based recognition.
Mingyu Chen 0003, Ghassan Al-Regib, Biing-Hwang Juang
IEEE Trans. Hum. Mach. Syst.3
2015 Discriminative Training Using Non-Uniform Criteria for Keyword Spotting on Spontaneous Speech
abstract
In this work, we formulate the problem of keyword spotting as a non-uniform error automatic speech recognition (ASR) problem and propose a model training methodology based on the non-uniform minimum classification error (MCE) approach. The main idea is to adapt the fundamental MCE criteria to reflect the cost-sensitive notion in that errors on keywords are much more significant than errors on non-keywords in an automatic speech recognition task. The notion of cost sensitivity leads to emphasis of keyword models in parameter optimization. Then we present a system which takes advantage of the weighted finite-state transducer (WFST) framework to efficiently implement the non-uniform MCE. To enhance the approach of non-uniform error cost minimization for keyword spotting, we further formulate a technique called ”adaptive boosted non-uniform MCE” which incorporates the idea of boosting. We validate the proposed framework on two challenging large-scale spontaneous conversational telephone speech (CTS) datasets in two different languages (English and Mandarin). Experimental results show our framework can achieve consistent and significant spotting performance gains over both the maximum likelihood estimation (MLE) baseline and conventional discriminatively-trained systems with uniform error cost.
Chao Weng, Biing-Hwang Juang
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Recurrent deep neural networks for robust speech recognition
abstract
In this work, we propose recurrent deep neural networks (DNNs) for robust automatic speech recognition (ASR). Full recurrent connections are added to certain hidden layer of a conventional feedforward DNN and allow the model to capture the temporal dependency in deep representations. A new backpropagation through time (BPTT) algorithm is introduced to make the minibatch stochastic gradient descent (SGD) on the proposed recurrent DNNs more efficient and effective. We evaluate the proposed recurrent DNN architecture under the hybrid setup on both the 2ndCHiME challenge (track 2) and Aurora-4 tasks. Experimental results on the CHiME challenge data show that the proposed system can obtain consistent 7% relative WER improvements over the DNN systems, achieving state-of-the-art performance without front-end preprocessing, speaker adaptive training or multiple decoding passes. For the experiments on Aurora-4, the proposed system achieves 4% relative WER improvement over a strong DNN baseline system.
Chao Weng, Dong Yu 0001, Shinji Watanabe 0001, Biing-Hwang Juang
ICASSP4
2014 Latent semantic rational kernels for topic spotting on conversational speech
abstract
In this work, we propose latent semantic rational kernels (LSRK) for topic spotting on conversational speech. Rather than mapping the input weighted finite-state transducers (WFSTs) onto a high dimensional n-gram feature space as in n-gram rational kernels, the proposed LSRK maps the WFSTs onto a latent semantic space. With the proposed LSRK, all available external knowledge and techniques can be flexibly integrated into a unified WFST based framework to boost the topic spotting performance. We present how to generalize the LSRK using tf-idf weighting, latent semantic analysis, WordNet and probabilistic topic models. To validate the proposed LSRK framework, we conduct the topic spotting experiments on two datasets, Switchboard and AT&T HMIHY0300 initial collection. The experimental results show that with the proposed LSRK we can achieve significant and consistent topic spotting performance gains over the n-gram rational kernels.
Chao Weng, David L. Thomson, Patrick Haffner, Biing-Hwang Juang
IEEE ACM Trans. Audio Speech Lang. Process.4
2013 Modeling heterogeneous data sources for speech recognition using synchronous hidden Markov models
abstract
In this paper, we propose a novel acoustic modeling framework, synchronous HMM, which takes full advantage of the capacity of the heterogeneous data sources and achieves an optimal balance between modeling accuracy and robustness. The synchronous HMM introduces an additional layer of substates between the HMM states and the Gaussian component variables. The substates have the capability to register long-span non-phonetic attributes, which are integrally called speech scenes in this study. The hierarchical modeling scheme allows an accurate description of probability distribution of speech units in different speech scenes. To address the data sparsity problem, a decision-based clustering algorithm is presented to determine the set of speech scenes and to tie the substate parameters. Moreover, we propose the multiplex Viterbi algorithm to efficiently decode the synchronous HMMs within a search space of the same size as for the standard HMMs. The experiments on the Aurora 2 task show that the synchronous HMMs produce a significant improvement in recognition performance over the HMM baseline at the expense of a moderate increase in the memory requirement and computational complexity.
Yong Zhao 0008, Biing-Hwang Juang
ICASSP2
2013 Perceptually motivated temporal modeling of footsteps in a cross-environmental detection task
abstract
Real world sounds are ubiquitous and form an important part of the edifice of our cognitive abilities. Their perception combines signatures from spectral and temporal domains, among others, yet traditionally their analysis is focused on the frame based spectral properties. We consider the problem of sound analysis from perceptual perspective and investigate the temporal properties of a “footsteps” sound, which is a particularly challenging from the time-frequency analysis viewpoint. We identify the irregular repetition of the self similarity and the sense of duration as significant to its perceptual quality and extract features using the Teager-Kaiser energy operator. We build an acoustic event detection system for “footsteps” which shows promising results for detection in cross-environmental conditions when compared with conventional approach.
M. Umair Bin Altaf, Taras Butko, Biing-Hwang Juang
ICASSP3
2013 Adaptive boosted non-uniform mce for keyword spotting on spontaneous speech
abstract
In this work, we present a complete framework of discriminative training using non-uniform criteria for keyword spotting, adaptive boosted non-uniform minimum classification error (MCE) for keyword spotting on spontaneous speech. To further boost the spotting performance and tackle the potential issue of over-training in the non-uniform MCE proposed in our prior work, we make two improvements to the fundamental MCE optimization procedure. Furthermore, motivated by AdaBoost, we introduce an adaptive scheme to embed error cost functions together with model combinations during the decoding stage. The proposed framework is comprehensively validated on two challenging large-scale spontaneous conversational telephone speech (CTS) tasks in different languages (English and Mandarin) and the experimental results show it can achieve significant and consistent figure of merit (FOM) gains over both ML and discriminatively trained systems.
Chao Weng, Biing-Hwang Juang
ICASSP2
2013 Latent semantic rational kernels for topic spotting on spontaneous conversational speech
abstract
In this work, we propose latent semantic rational kernels (LSRK) for topic spotting on spontaneous conversational speech. Rather than mapping the input weighted finite-state transducers (WFSTs) onto a high dimensional n-gram feature space as in n-gram rational kernels, the proposed LSRK maps the WFSTs onto a latent semantic space. Moreover, with the LSRK framework, all available external knowledge can be flexibly incorporated to boost the topic spotting performance. The experiments we conducted on a spontaneous conversational task, Switchboard, show that our method can achieve significant performance gain over the baselines from 27.33% to 57.56% accuracy and almost double the classification accuracy over the n-gram rational kernels in all cases.
Chao Weng, Biing-Hwang Juang
ICASSP2
2013 On the performance of the robust acoustic echo cancellation system with decorrelation by sub-band resampling
abstract
This paper examines the effect of inter-channel decorrelation by sub-band resampling (SBR) on the performance of the robust acoustic echo cancellation (AEC) system based on the residual echo enhancement technique. Due to the flexibility of SBR, the decorrelation performance as measured by the coherence can be matched with other conventional decorrelation procedures. Given the same degree of decorrelation, we have shown previously that SBR achieves superior audio quality compared to other procedures. We show in this paper that SBR also provides higher stereophonic AEC performance in a very noisy condition, where the performance is evaluated by decomposing the true echo return loss enhancement and the misalignment per sub-band to better demonstrate the superiority of our decorrelation procedure over other methods.
Jason Wung, Ted S. Wada, Biing-Hwang Juang
ICASSP3
2013 Person identification using biometric markers from footsteps sound
M. Umair Bin Altaf, Taras Butko, Biing-Hwang Juang
INTERSPEECH3
2013 Feature Processing and Modeling for 6D Motion Gesture Recognition
abstract
A 6D motion gesture is represented by a 3D spatial trajectory and augmented by another three dimensions of orientation. Using different tracking technologies, the motion can be tracked explicitly with the position and orientation or implicitly with the acceleration and angular speed. In this work, we address the problem of motion gesture recognition for command-and-control applications. Our main contribution is to investigate the relative effectiveness of various feature dimensions for motion gesture recognition in both user-dependent and user-independent cases. We introduce a statistical feature-based classifier as the baseline and propose an HMM-based recognizer, which offers more flexibility in feature selection and achieves better performance in recognition accuracy than the baseline system. Our motion gesture database which contains both explicit and implicit motion information allows us to compare the recognition performance of different tracking signals on a common ground. This study also gives an insight into the attainable recognition rate with different tracking devices, which is valuable for the system designer to choose the proper tracking technology.
Mingyu Chen 0003, Ghassan Al-Regib, Biing-Hwang Juang
IEEE Trans. Multim.3
2012 6D motion gesture recognition using spatio-temporal features
abstract
Depending on the tracking technology in use, a 6D motion gesture can be tracked and represented explicitly by the position and orientation or implicitly by the acceleration and angular speed. In this work, we first present the reasoning for the definition and recognition of motion gestures. Five basic feature vectors are then derived from the 6D motion data. Our main contribution is to investigate the relative effectiveness of various feature dimensions for motion gesture recognition in both user dependent and user independent cases. We also propose a feature normalization procedure and prove its effectiveness in achieving “scale” invariance especially in the user independent case. Our study gives an insight into the attainable recognition rate with different tracking devices.
Mingyu Chen 0003, Ghassan Al-Regib, Biing-Hwang Juang
ICASSP3
2012 A comparative study of discriminative training using non-uniform criteria for cross-layer acoustic modeling
abstract
This work focuses on a comparative study of discriminative training using non-uniform criteria for cross-layer acoustic modeling. Two kinds of discriminative training (DT) frameworks, minimum classification error like (MCE-like) and minimum phone error like (MPE-like) DT frameworks, are augmented to allow the error cost embedding at the phoneme (model) level respectively. To facilitate this comparative study, we implement both augmented DT frameworks under the same umbrella, using the error cost derived from the same cross-layer confusion matrix. Experiments on a large vocabulary task WSJ0 demonstrated the effectiveness of both DT frameworks with the formulated non-uniform error cost embedded. Several preliminary investigations on the effect of the dynamic range of error cost are also presented.
Chao Weng, Biing-Hwang Juang
ICASSP2
2012 Inter-channel decorrelation by sub-band resampling in frequency domain
abstract
This paper presents a novel decorrelation procedure by frequency-domain resampling in sub-bands. The new procedure expands on the idea of resampling in the frequency domain that efficiently and effectively alleviates the non-uniqueness problem for a multi-channel acoustic echo cancellation system while introducing minimal distortion to the signal. We show in theory and verify experimentally that the amount of decorrelation in each sub-band, measured in terms of the coherence, can be controlled arbitrarily by varying the resampling ratio per frequency bin. For perceptual evaluation, we adjust the sub-band resampling ratios to match the coherence given by other decorrelation procedures. The speech quality (PESQ) score from the proposed decorrelation procedure remains high at around 4.5, which is about the highest possible PESQ score after signal modification.
Jason Wung, Ted S. Wada, Biing-Hwang Juang
ICASSP3
2012 Stranded Gaussian mixture hidden Markov models for robust speech recognition
abstract
Gaussian mixture (GMM)-HMMs, though being the predominant modeling technique for speech recognition, are often criticized as being inaccurate to model heterogeneous data sources. In this work, we propose the stranded Gaussian mixture (SGMM)-HMM, an extension of the GMM-HMM, to explicitly model the dependence among the mixture components, i.e., each mixture component is assumed to depend on the previous mixture component in addition to the state that generates it. In the evaluation over the Aurora 2 database, the proposed 20-mixture SGMM system obtains WER of 8.07%, 10% relative improvement over the baseline GMM system. The experiments demonstrate the discriminating power that would be possessed by the mixture weights in their advanced form.
Yong Zhao 0008, Biing-Hwang Juang
ICASSP2
2012 A general discriminative training algorithm for speech recognition using weighted finite-state transducers
abstract
In this paper, we present a general algorithmic framework based on WFSTs for implementing a variety of discriminative training methods, such as MMI, MCE, and MPE/MWE. In contrast to the ordinary word lattices, the transducer-based lattices are more amenable to representing and manipulating the underlying hypothesis space and have a finer granularity at the HMM-state level. The transducers are processed into a two-layer hierarchy: at a high level, it is analogous to a word lattice, and each word transition embodies an HMM-state subgraph for that word at a lower level. This hierarchy combined with the appropriate customization of the transducers leads to a flexible implementation for all of the training criteria being discussed. The effectiveness of the framework is verified on two speech recognition tasks: Resource Management, and AT&T SCANMail, an internal voicemail-to-text task.
Yong Zhao 0008, Andrej Ljolje, Diamantino Caseiro, Biing-Hwang Juang
ICASSP4
2012 Discriminative Training Using Non-uniform Criteria for Keyword Spotting on Spontaneous Speech
Chao Weng, Biing-Hwang Juang, Daniel Povey
INTERSPEECH2
2012 6DMG: a new 6D motion gesture database
abstract
Motion-based control is gaining popularity, and motion gestures form a complementary modality in human-computer interactions. To achieve more robust user-independent motion gesture recognition in a manner analogous to automatic speech recognition, we need a deeper understanding of the motions in gesture, which arouses the need for a 6D motion gesture database. In this work, we present a database that contains comprehensive motion data, including the position, orientation, acceleration, and angular speed, for a set of common motion gestures performed by different users. We hope this motion gesture database can be a useful platform for researchers and developers to build their recognition algorithms as well as a common test bench for performance comparisons.
Mingyu Chen 0003, Ghassan Al-Regib, Biing-Hwang Juang
MMSys3
2012 Index-based incremental language model for scalable directory assistance
Antonio Moreno-Daniel, Jay G. Wilpon, Biing-Hwang Juang
Speech Commun.3
2012 Automatic Speech Recognition Based on Non-Uniform Error Criteria
abstract
The Bayes decision theory is the foundation of the classical statistical pattern recognition approach, with the expected error as the performance objective. For most pattern recognition problems, the “error” is conventionally assumed to be binary, i.e., 0 or 1, equivalent to error counting, independent of the specifics of the error made by the system. The term “error rate” is thus long considered the prevalent system performance measure. This performance measure, nonetheless, may not be satisfactory in many practical applications. In automatic speech recognition, for example, it is well known that some errors are more detrimental (e.g., more likely to lead to misunderstanding of the spoken sentence) than others. In this paper, we propose an extended framework for the speech recognition problem with non-uniform classification/recognition error cost which can be controlled by the system designer. In particular, we address the issue of system model optimization when the cost of a recognition error is class dependent. We formulate the problem in the framework of the minimum classification error (MCE) method, after appropriate generalization to integrate the class-dependent error cost into one consistent objective function for optimization. We present a variety of training scenarios for automatic speech recognition under this extended framework. Experimental results for continuous speech recognition are provided to demonstrate the effectiveness of the new approach.
Qiang Fu 0008, Yong Zhao 0008, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.3
2012 Enhancement of Residual Echo for Robust Acoustic Echo Cancellation
abstract
This paper examines the technique of using a noise-suppressing nonlinearity in the adaptive filter error feedback-loop of an acoustic echo canceler (AEC) based on the least mean square (LMS) algorithm when there is an interference at the near end. The source of distortion may be linear, such as local speech or background noise, or nonlinear due to speech coding used in the telecommunication networks. Detailed derivation of the error recovery nonlinearity (ERN), which “enhances” the filter estimation error prior to the adaptation in order to assist the linear adaptation process, will be provided. Connections to other existing AEC and signal enhancement techniques will be revealed. In particular, the error enhancement technique is well-founded in the information-theoretic sense and has strong ties to independent component analysis (ICA), which is the basis for blind source separation (BSS) that permits unsupervised adaptation in the presence of multiple interfering signals. The single-channel AEC problem can be viewed as a special case of semi-blind source separation (SBSS) where one of the source signals is partially known, i.e., the far-end microphone signal that generates the near-end acoustic echo. The system approach to robust AEC will be motivated, where a proper integration of the LMS algorithm with the ERN into the AEC “system” allows for continuous and stable adaptation even during double talk without precise estimation of the signal statistics. The error enhancement paradigm encompasses many traditional signal enhancement techniques and opens up an entirely new avenue for solving the AEC problem in a real-world setting.
Ted S. Wada, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.2
2012 Nonlinear Compensation Using the Gauss-Newton Method for Noise-Robust Speech Recognition
abstract
In this paper, we present the Gauss-Newton method as a unified approach to estimating noise parameters of the prevalent nonlinear compensation models, such as vector Taylor series (VTS), data-driven parallel model combination (DPMC), and unscented transform (UT), for noise-robust speech recognition. While iterative estimation of noise means in a generalized EM framework has been widely known, we demonstrate that such approaches are variants of the Gauss-Newton method. Furthermore, we propose a novel noise variance estimation algorithm that is consistent with the Gauss-Newton principle. The formulation of the Gauss-Newton method reduces the noise estimation problem to determining the Jacobians of the corrupted speech parameters. For sampling-based compensations, we present two methods, sample Jacobian average (SJA) and cross-covariance (XCOV), to evaluate these Jacobians. The proposed noise estimation algorithm is evaluated for various compensation models on two tasks. The first is to fit a Gaussian mixture model (GMM) model to artificially corrupted samples, and the second is to perform speech recognition on the Aurora 2 database. The significant performance improvements confirm the efficacy of the Gauss-Newton method to estimating the noise parameters of the nonlinear compensation models.
Yong Zhao 0008, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.2
2011 Audio signal classification with temporal envelopes
abstract
The conventional approach to audio processing, based on the short-time power spectrum model, is not adequate when it comes to general audio signals. We propose an approach, justified by studies from psycho-acoustics and neuroimaging, which uses the magnitude and frequency envelope of the audio signal in the from of AM-FM modulations to build an ARMA model which is then fed to a GMM to classify into various audio classes. We show that it makes explicit certain aspects of the signal which are overlooked when processing is limited to the spectral domain.
M. Umair Bin Altaf, Biing-Hwang Juang
ICASSP2
2011 Trajectory triangulation: 3D motion reconstruction with ℓ1 optimization
abstract
In this paper, we first explain the formulation of the trajectory triangulation: 3D reconstruction of a moving point from a series of 2D projections. The system has to be overconstrained to be solved by least squares techniques. We take advantage of the sparseness of real-world motions in the transformed domain, and borrow the concept of compressive sampling to reformulate the problem with ℓ1optimization so that it is possible to reconstruct the trajectory even in an underconstrained system. Thus, fewer measurements are needed to reconstruct a 3D trajectory of even larger bandwidth coverage. We conduct experiments on both synthetic and real-world motion data to verify our proposed method, and compare the reconstruction results based on ℓ1and ℓ2optimization.
Mingyu Chen 0003, Ghassan Al-Regib, Biing-Hwang Juang
ICASSP3
2011 Variability regularization in large-margin classification
abstract
This paper introduces a novel regularization strategy to address the generalization issues for large-margin classifiers from the Empirical Risk Minimization (ERM) perspective. First, the ERM principle is argued to be more flexible than the Structural Risk Minimization (SRM) principle by reviewing the difference between the two strategies as the fundamental principles for large-margin classifier design. Second, after studying the large-margin classifier design based on the SRM principle, a realization of the ERM principle is proposed in the form of a bias-variance criterion instead of the conventional expected error criterion. The bias-variance criterion is shown to have the regularization capability needed by a large-margin classifier designed according to the ERM principle. Finally, a mathematical programming procedure is used to efficiently achieve the best regularization policy. The new regularization strategy based on the ERM principle is evaluated on a set of machine learning experiments. Experimental results clearly demonstrate the strength of the proposed regularization strategy to achieve the minimum error rate performance measure.
Dwi Sianto Mansjur, Ted S. Wada, Biing-Hwang Juang
ICASSP3
2011 Discriminative Training for direct minimization of deletion, insertion and substitution errors
abstract
In this paper, we follow the minimum error principle for acoustic modeling and formulate error objectives in insertion, deletion, and substitution separately for minimization during training. This new training paradigm generalized from the MVE criterion can explain the direct relationship between recognition errors and detection errors by re-interpreting deletion, insertion, and substitution errors as miss, false alarm, and miss/false-alarm errors happening together. Under the MVE criterion, by applying two mis-verification measures for miss and false alarm errors selectively along with the types of recognition error definition, we developed three individual objective training criteria, minimum deletion error (MDE), minimum insertion error (MIE), and minimum substitution error (MSE), of which each objective function can directly minimize each of the three types of the recognition errors. In the TIMIT phone recognition task, the experimental results confirm that each objective criterion of MDE, MIE, and MSE results in primarily minimizing its target error type, respectively. Furthermore, a simple combination of the individual objective criteria outperforms the conventional string-based MCE in the overall recognition error rate.
Sunghwan Shin, Ho-Young Jung, Biing-Hwang Juang
ICASSP3
2011 Recent development of discriminative training using non-uniform criteria for cross-level acoustic modeling
abstract
In this paper, we extend our previous study on discriminative training using non-uniform criteria for speech recognition. The work will put emphasis on how the acoustic modeling interacts with the risk at a higher level, which is more relevant to the most used evaluation measures, e.g., word error rate (WER). To be specific, the non-uniform error cost is first derived at the word level to minimize the risk w.r.t. WER and then computed on the word lattice using the forward-backward algorithm. With the statistics obtained from the forward-backward algorithm, the competing hypotheses for each label word are searched by performing dynamic programming between the label word sequence and the word lattice at the phone level. In order to alleviate the level inconsistency between the acoustic model (phone level) and the evaluation measure (word level), the derived error cost is embedded into the overall objective function in a cross-level fashion. Experiments on a large vocabulary task WSJO demonstrate the effectiveness of the overall approach, which show it outperforms two prevalent discriminative training methods and achieves about 13% relative improvement over the baseline system.
Chao Weng, Biing-Hwang Juang
ICASSP2
2011 A system approach to residual echo suppression in robust hands-free teleconferencing
abstract
This paper presents a system approach to the residual echo suppression (RES) problem in a noisy acoustic environment. We propose a method that takes advantage of our existing robust acoustic echo cancellation system in order to obtain a residual echo estimate that closely resembles the true, noise-free residual echo. To achieve improved RES during strong near-end interference (e.g., double talk), a psychoacoustic postfilter is also used. The simulation results show that our RES based on the system approach outperforms a conventional estimation method. Comparing the postfiltered output to the unprocessed one indicates that our proposed RES approach can raise the PESQ score by more than half a point.
Jason Wung, Ted S. Wada, Biing-Hwang Juang, Bowon Lee, Ton Kalker, Ronald W. Schafer
ICASSP3
2011 Non-linear noise compensation for robust speech recognition using Gauss-Newton method
abstract
In this paper, we present the Gauss-Newton method as a unified approach to optimizing non-linear noise compensation models, such as vector Taylor series (VTS), data-driven parallel model combination (DPMC), and unscented transform (UT). We demonstrate that the commonly used approaches that iteratively approximate the noise parameters in an EM framework are variants of the Gauss-Newton method. Through the formulation of the Gauss-Newton method for estimating noise means and variances, the noise estimation problems are reduced to determining the Jacobians of the noisy speech distributions. For the sampling-based compensations, we present two methods, sample Jacobian average (SJA) and cross-covariance (XCOV), to evaluate the Jacobians. Experiments on the Aurora 2 database verify the efficacy of the Gauss-Newton method to these noise compensation models.
Yong Zhao 0008, Biing-Hwang Juang
ICASSP2
2011 Large Margin - Minimum Classification Error Using Sum of Shifted Sigmoids as the Loss Function
Madhavi Vedula Ratnagiri, Biing-Hwang Juang, Lawrence R. Rabiner
INTERSPEECH2
2011 Individual Error Minimization Learning Framework and its Applications to Speech Recognition and Utterance Verification
Sunghwan Shin, Ho-Young Jung, Biing-Hwang Juang
INTERSPEECH3
2011 Model Adaptation for Automatic Speech Recognition Based on Multiple Time Scale Evolution
abstract
The change in speech characteristics is originated from various factors, at various (temporal) rates in a real world conversation. These temporal changes have their own dynamics and therefore, we propose to extend the single (time-) incremental adaptations to a multiscale adaptation, which has the potential of greatly increasing the model’s robustness as it will include adaptation mechanism to approximate the nature of the characteristic change. The formulation of the incremental adaptation assumes a time evolution system of the model, where the posterior distributions, used in the decision process, are successively updated based on a macroscopic time scale in accordance with the Kalman filter theory. In this paper, we extend the original incremental adaptation scheme, based on a single time scale, to multiple time scales, and apply the method to the adaptation of both the acoustic model and the language model. We further investigate methods to integrate the multi-scale adaptation scheme to realize the robust speech recognition performance. Large vocabulary continuous speech recognition experiments for English and Japanese lectures revealed the importance of modeling multiscale properties in speech recognition. Index Terms: speech recognition, incremental adaptation, multiscale, time evolution system
Shinji Watanabe 0001, Atsushi Nakamura, Biing-Hwang Juang
INTERSPEECH3
2011 Batch-Online Semi-Blind Source Separation Applied to Multi-Channel Acoustic Echo Cancellation
abstract
Semi-blind source separation (SBSS) is a special case of the well-known blind source separation (BSS) when some partial knowledge of the source signals is available to the system. In particular, a batch adaptation in the frequency domain based on independent component analysis (ICA) can be effectively used to jointly perform source separation and multichannel acoustic echo cancellation (MCAEC) through SBSS without double-talk detection. Many issues related to the implementation of an SBSS system are discussed in this paper. After a deep analysis of the structure of the SBSS adaptation, we propose a constrained batch-online implementation that stabilizes the convergence behavior even in the worst case scenario of a single far-end talker along with the non-uniqueness condition on the far-end mixing system. Specifically, a matrix constraint is proposed to reduce the effect of the non-uniqueness problem caused by highly correlated far-end reference signals during MCAEC. Experimental results show that high echo cancellation can be achieved just as the misalignment remains relatively low without any preprocessing procedure to decorrelate the far-end signals even for the single far-end talker case.
Francesco Nesta, Ted S. Wada, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.3
2010 Feature extraction by incremental parsing for music indexing
abstract
In this paper, we employ a linguistic-processing approach to the content-based retrieval of music information. Central to the approach is the use of a lossy version of the Lempel-Ziv incremental parsing (LZIP) algorithm, which constructs a dictionary by incrementally parsing music feature vectors. LZIP is adopted as a source characterization technique owing to it's universal-coding nature, and asymptotic convergence to the entropy of the source. The dictionary is composed of variable-length parsed representations, which are used to construct a highly sparse co-occurrence matrix, which counts the occurrence of the parsed representations in each music. As a feature analysis framework, Latent Semantic Analysis (LSA) is then applied to the co-occurrence matrix to generate a lower-dimensional approximation that exposes the most salient features of the represented audio documents. The aforementioned approach, in addition to adopting reduced sampling rates and quantized feature vectors, yields a system with reduced requirements in terms of processing and storage, and increases the tolerance to noisy queries. We demonstrate the performance of the system in the music genre classification problem, and analyze its robustness to perturbed queries. Moreover, we demonstrate that using the incremental parsing algorithm in forming the audio dictionary has superior retrieval performance compared to techniques yielding a dictionary with fixed-length entries such as vector quantization.
Nawaf I. Almoosa, Soo Hyun Bae, Biing-Hwang Juang
ICASSP3
2010 Multiple acoustic source localization based on multiple hypotheses testing using particle approach
abstract
Localization of multiple acoustic sources in a non-ideal environment has a number of difficulties, among which are accurate acoustic feature estimation for multiple sources and association uncertainty between measurements and their corresponding sources. This paper focuses more on the latter and proposes an algorithm based on a multiple-hypothesis framework for both a measurement model and a measurement association model to localize multiple sources. A conditional data likelihood model based on a measurement hypothesis is proposed and implemented using particles. Simulation results demonstrate that the proposed algorithm is capable of localizing the positions of multiple sources with a small number of microphones without any prior knowledge when the amount of reverberation is moderate.
Yeongseon Lee, Ted S. Wada, Biing-Hwang Juang
ICASSP3
2010 Discriminative linear-transform based adaptation using minimum verification error
abstract
This paper presents an investigation of the minimum verification error linear regression (MVELR) method for discriminative linear-transform based adaptation. The MVE criterion is employed to estimate a set of discriminative linear transformations which achieve the smallest empirical average loss with the given adaptation data. The MVELR directly minimizes the total detection errors, some of which are results of characteristic mismatch in the given adaptation data. In this study, segment-based phonetic detectors reflecting an important processing layer in speech event detection are initially trained via the conventional maximum likelihood (ML) method and then refined via the general MVE method using the original training data. Then, the initial MVE-trained detectors are adapted by two kinds of adaption techniques, MLLR and MVELR, respectively, with the given adaptation data for comparison. The experiments are performed on a supervised adaptation scenario and the effectiveness of the adapted detectors is evaluated based on the total detection error. Experimental results confirm the proposed MVELR method considerably reduces the total error rate over all categories of the detectors compared to the MLLR.
Sunghwan Shin, Ho-Young Jung, Biing-Hwang Juang
ICASSP4
2010 On noise estimation for robust speech recognition using vector Taylor series
abstract
In this paper, we propose a novel noise variance estimation method using the fixed point method for the VTS-based robust speech recognition. Noise parameters are re-estimated over a given utterance using an EM algorithm. The derivative of the auxiliary function with respect to the noise variance is resolved, and the fixed point algorithm estimates the noise variance by recursively approximating the root of the resulting derivative. The method leads to a re-estimation formula with a flavor like the standard ML variance estimation, and the iteration procedure is step-size free. We also investigate improving the noise estimation for efficient VTS adaptation. Several fast noise estimation methods are examined including estimation from non-speech areas and incremental adaptation. In the evaluation over Aurora 2 database, the proposed noise variance estimation method obtains a significant improvement in recognition accuracy over the method using sample variance. Further experiments show that the VTS ML estimation over non-speech areas is an effective fast adaptation method. The final refined approach achieves 8.75% WER, 13% relative improvement over the conventional VTS adaptation.
Yong Zhao 0008, Biing-Hwang Juang
ICASSP2
2010 Multi-Class Classification Using a New Sigmoid Loss Function for Minimum Classification Error (MCE)
abstract
A new loss function has been introduced for Minimum Classification Error, that approaches optimal Bayes’ risk and also gives an improvement in performance over standard MCE systems when evaluated on the Aurora connected digits database.
Madhavi Vedula Ratnagiri, Lawrence R. Rabiner, Biing-Hwang Juang
ICMLA3
2010 Feature Transformation and Model Design Using Minimum Classification Error
abstract
A Minimum Classification Error (MCE) based recognition system that also estimates a global feature transformation matrix has been implemented. Unlike earlier studies, we make the explicit assumption that the covariance matrix of the Gaussian mixtures is diagonal when estimating the transformation matrix. This is necessary for mathematical consistency between the model and the transformation matrix estimates. Experimental results show a reduction of up to 50% in the word error rate as compared to Maximum Likelihood estimation.
Madhavi Vedula Ratnagiri, Lawrence R. Rabiner, Biing-Hwang Juang
ICMLA3
2010 A comparative study of noise estimation algorithms for VTS-based robust speech recognition
abstract
We conduct a comparative study to investigate two noise es-timation approaches for robust speech recognition using vec-tor Taylor series (VTS) developed in the past few years. The first approach, iterative root finding (IRF), directly differenti-ates the EM auxiliary function and approximates the root of the derivative function through recursive refinements. The second approach, twofold expectation maximization (TEM), estimates noise distributions by regarding them as hidden variables in a modified EM fashion. Mathematical derivations reveal the sub-stantial connection between the two approaches. Two experi-ments are performed in evaluating the performance and conver-gence rate of the algorithms. The first is to fit a GMM model to artificially corrupted samples that are generated through Monte Carlo simulation. The second is to perform speech recognition on the Aurora 2 database. Index Terms: Robust speech recognition, vector Taylor series, noise estimation
Yong Zhao 0008, Biing-Hwang Juang
INTERSPEECH2
2010 Introduction to the Special Issue on Processing Reverberant Speech: Methodologies and Applications
abstract
The 17 papers in this special issue focus on the methodologies and applications of processing reverberant speech. The issue highlights some major aspects of the recent progress in the field.
Tomohiro Nakatani, Walter Kellermann, Patrick A. Naylor, Masato Miyoshi, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.5
2010 Speech Dereverberation Based on Variance-Normalized Delayed Linear Prediction
abstract
This paper proposes a statistical model-based speech dereverberation approach that can cancel the late reverberation of a reverberant speech signal captured by distant microphones without prior knowledge of the room impulse responses. With this approach, the generative model of the captured signal is composed of a source process, which is assumed to be a Gaussian process with a time-varying variance, and an observation process modeled by a delayed linear prediction (DLP). The optimization objective for the dereverberation problem is derived to be the sum of the squared prediction errors normalized by the source variances; hence, this approach is referred to as variance-normalized delayed linear prediction (NDLP). Inheriting the characteristic of DLP, NDLP can robustly estimate an inverse system for late reverberation in the presence of noise without greatly distorting a direct speech signal. In addition, owing to the use of variance normalization, NDLP allows us to improve the dereverberation result especially with relatively short (of the order of a few seconds) observations. Furthermore, NDLP can be implemented in a computationally efficient manner in the time-frequency domain. Experimental results demonstrate the effectiveness and efficiency of the proposed approach in comparison with two existing approaches.
Tomohiro Nakatani, Takuya Yoshioka, Keisuke Kinoshita, Masato Miyoshi, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.5
2010 IPSILON: Incremental Parsing for Semantic Indexing of Latent Concepts
abstract
A new framework for content-based image retrieval, which takes advantage of the source characterization property of a universal source coding scheme, is investigated. Based upon a new class of multidimensional incremental parsing algorithm, extended from the Lempel-Ziv incremental parsing code, the proposed method captures the occurrence pattern of visual elements from a given image. A linguistic processing technique, namely the latent semantic analysis (LSA) method, is then employed to identify associative ensembles of visual elements, which lay the foundation for intelligent visual information analysis. In 2-D applications, incremental parsing decomposes an image into elementary patches that are different from the conventional fixed square-block type patches. When used in compressive representations, it is amenable in schemes that do not rely on average distortion criteria, a methodology that is a departure from the conventional vector quantization. We call this methodology a parsed representation. In this article, we present our implementations of an image retrieval system, called IPSILON, with parsed representations induced by different perceptual distortion thresholds. We evaluate the effectiveness of the use of the parsed representations by comparing their performance with that of four image retrieval systems, one using the conventional vector quantization for visual information analysis under the same LSA paradigm, another using a method called SIMPLIcity which is based upon an image segmentation and integrated region matching, and the other two based upon query-by-semantic-example and query-by-visual-example. The first two of them were tested with 20,000 images of natural scenes, and the others were tested with a portion of the images. The experimental results show that the proposed parsed representation efficiently captures the salient features in visual images and the IPSILON systems outperform other systems in terms of retrieval precision and distortion robustness.
Soo Hyun Bae, Biing-Hwang Juang
IEEE Trans. Image Process.2
2009 Aspect modeling of parsed representation for image retrieval
abstract
A probabilistic framework based on a universal source coding for content-based image retrieval is proposed. By a multidimensional incremental parsing technique, which is an extension of the Lempel-Ziv incremental parsing algorithm, a given image is parsed into a number of variable-size rectangular blocks, called parsed representations. To achieve a semantically relevant pattern matching, we introduce a new similarity measure from the first- and second-order statistics of given image patches. Once the occurrence patterns of images in the corpus are analyzed, the term-document joint distribution is estimated by an aspect modeling technique under the assumption of latent aspects. To compare the performance of the proposed image retrieval framework based on the parsed representations, we implement a benchmark system based on the fixed-shape block representations trained by vector quantization. In addition to these two systems, we bring two content-based image retrieval systems into the performance evaluation. The experimental results on a database of 20,000 natural scene images demonstrate that the proposed image retrieval system significantly outperforms other existing and the benchmark systems.
Soo Hyun Bae, Biing-Hwang Juang
ICASSP2
2009 Kernel-based nonlinear independent component analysis for underdetermined blind source separation
abstract
In this paper we propose a new unsupervised training method for nonlinear spatial filter using a new independent component analysis based on kernel infomax. The nonlinearity of the spatial filter used in this paper is equivalent to the integration of beamforming and spectral subtraction, and the whole structure is optimized by independent component analysis in the reproducing kernel Hilbert space. The optimized filter is shown to be capable of achieving better quality output than the conventional method based on time-frequency binary masking.
Shigeki Miyabe, Biing-Hwang Juang, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP2
2009 A scalable method for voice search to nationwide business listings
abstract
Voice search or 411-service is the task that finds a ranked set of directory listings that match a spoken query, where the target entries in the listing database and the spoken query may differ moderately in their syntactic form. While the conventional paradigm uses a two-box input (location + name), a single-box paradigm to voice search can allow users to provide all the information in a single utterance, thereby increasing query efficiency. Furthermore, the scalability of traditional methods used in the two-box paradigm is infeasible, and alternative strategies that sacrifice accuracy are normally adopted. This work presents a scalable algorithm for directory search over a nationwide database of listings (millions of entries) without compromising recognition accuracy.
Antonio Moreno-Daniel, Biing-Hwang Juang, Jay G. Wilpon
ICASSP2
2009 Real-time speech enhancement in noisy reverberant multi-talker environments based on a location-independent room acoustics model
abstract
This paper describes a new real-time speech enhancement method that reduces signal distortion caused by stationary noise and late reflections of reverberation in speech signals captured by a single distant microphone under multi-talker conditions. A major problem here is how to estimate the energy of the late reflections in real time when the room impulse responses from individual talkers to the microphone are not given or fixed in advance. To solve this problem, we introduce a probabilistic room acoustics model, and provide a method for estimating the energy of late reflections based on this model. In this method, parameters of the model for a room can be fixed in advance only from a few seconds of observation. By incorporating the proposed approach into a conventional frequency domain noise reduction scheme, we realize an integrated real-time speech enhancement framework. The effectiveness of the proposed method is confirmed experimentally for a case where there are two talkers in a room.
Tomohiro Nakatani, Takuya Yoshioka, Keisuke Kinoshita, Masato Miyoshi, Biing-Hwang Juang
ICASSP5
2009 Speech enhancement using minimum mean-square error estimation and a post-filter derived from vector quantization of clean speech
abstract
In this paper, a novel post-filtering method applied after the logSTSA filter is proposed. Since the post-filter is derived from vector quantization of clean speech database, it has an equivalent effect of imposing clean source spectral constraints on the enhanced speech. When combined with the logSTSA filter, the additional filter can noticeably suppress residual artifacts by effectively lowering the residual white noise of decision-directed estimation as well as reducing the musical noise of maximum likelihood estimation. Compared to the logSTSA enhanced speech, the overall enhanced speech is able to raise the PESQ score by nearly half a point.
Jason Wung, Shigeki Miyabe, Biing-Hwang Juang
ICASSP3
2009 A study on recognizing distorted speech over local distributed transducer networks
abstract
In a collaborative scenario, a multiplicity of portable devices may constitute a network of distributed microphones, without a clearly defined geometric configuration or synchronization that can be taken advantage of for traditional microphone array processing to enhance the acquired signal. This application scenario represents a severe, but interesting challenge for automatic speech recognition systems. In this paper, we investigate a variety of robust speech recognition techniques with a focus on the distributed transducer scenario. We also report some important study results that lead to new thinking in the design of robust speech recognition for broadened applications. Two issues that are inherent to distributed transducer networks are specially investigated. First, we study the effect of the sampling rate skew of microphones to the system performance; second, we explore the possibility of combining recognition hypotheses from multiple transducer channels for improved recognition accuracy.
Yong Zhao 0008, Sunghwan Shin, Enrique Robledo-Arnuncio, Biing-Hwang Juang
ICASSP4
2009 Using Kernel Density Classifier with Topic Model and Cost Sensitive Learning for Automatic Text Categorization
abstract
This paper proposes a novel framework for automatic text categorization problem based on the kernel density classifier. The overall goal is to tackle two main issues in automatic text categorization problems: the interpretability and the performance. Specifically, to solve the interpretability issue, the latent semantic analysis technique is used to construct a topic space, in which each dimension represents a single topic. The text features are extracted directly from this topic space. To solve the performance issue, classifierspsila parameters are optimized for either cost-sensitive or non-cost-sensitive categorization. We have experimentally evaluated the proposed framework by using a corpus of twenty newsgroups. The experimental results confirm the effectiveness of the framework to utilize the features from the topic model for cost-sensitive categorization.
Dwi Sianto Mansjur, Ted S. Wada, Biing-Hwang Juang
ICDAR3
2009 Signal Processing in Cognitive Radio
abstract
Cognitive radio allows for usage of licensed frequency bands by unlicensed users. However, these unlicensed (cognitive) users need to monitor the spectrum continuously to avoid possible interference with the licensed (primary) users. Apart from this, cognitive radio is expected to learn from its surroundings and perform functions that best serve its users. Such an adaptive technology naturally presents unique signal-processing challenges. In this paper, we describe the fundamental signal-processing aspects involved in developing a fully functional cognitive radio network, including spectrum sensing and spectrum sculpting.
Jun Ma 0007, Geoffrey Ye Li, Biing-Hwang Juang
Proc. IEEE3
2009 Subjective Evaluation of Spatial Resolution and Quantization Noise Tradeoffs
abstract
Most full-reference fidelity/quality metrics compare the original image to a distorted image at the same resolution assuming a fixed viewing condition. However, in many applications, such as video streaming, due to the diversity of channel capacities and display devices, the viewing distance and the spatiotemporal resolution of the displayed signal may be adapted in order to optimize the perceived signal quality. For example, at low bitrate coding applications an observer may prefer to reduce the resolution or increase the viewing distance to reduce the visibility of the compression artifacts. The tradeoff between resolution/viewing conditions and visibility of compression artifacts requires new approaches for the evaluation of image quality that account for both image distortions and image size. In order to better understand such tradeoffs, we conducted subjective tests using two representative still image coders, JPEG and JPEG 2000. Our results indicate that an observer would indeed prefer a lower spatial resolution (at a fixed viewing distance) in order to reduce the visibility of the compression artifacts, but not all the way to the point where the artifacts are completely invisible. Moreover, the observer is willing to accept more artifacts as the image size decreases. The subjective test results we report can be used to select viewing conditions for coding applications. They also set the stage for the development of novel fidelity metrics. The focus of this paper is on still images, but it is expected that similar tradeoffs apply to video.
Soo Hyun Bae, Thrasyvoulos N. Pappas, Biing-Hwang Juang
IEEE Trans. Image Process.3
2008 Toward robust moment invariants for image registration
abstract
We apply pattern recognition techniques to enhance the robustness of moment-invariants-based image classifiers. Moment invariants exhibit variations under transformations that do not preserve the original image function, such as geometrical transformations involving interpolation. Such variations degrade the performance of classifiers due to the errors in the nearest neighbor search stage. We propose the use of linear discriminant analysis (LDA) and principal component analysis (PCA) to alleviate the variations and enhance the robustness of classification. We demonstrate the improved performance in image registration applications under spatial scaling and rotation transformations.
Nawaf I. Almoosa, Soo Hyun Bae, Biing-Hwang Juang
ICASSP3
2008 Non-Uniform error criteria for automatic pattern and speech recognition
abstract
The classical Bayes decision theory [1] is the foundation of statistical pattern recognition. Conventional applications of the Bayes decision theory result in ubiquitous use of the maximum a posteriori probability (MAP) decision policy and the paradigm of distribution estimation as practice in the design of a statistical pattern recognition system. In this paper, we address the issue of non-uniform error criteria in statistical pattern recognition, and generalize the Bayes decision theory for pattern recognition tasks where errors over different classes have different degrees of significance. We further propose extensions of the method of minimum classification error (MCE) [2] for a practical design of a statistical pattern recognition system to achieve empirical optimality when non-uniform error criteria are prescribed. In addition, we apply our method upon speech recognition tasks. In the context of automatic speech recognition (ASR), we present a variety of training scenarios and weighting strategies under our framework. The experimental demonstrations for both general pattern recognition and continuous speech recognition are provided to support the effectiveness of our new approach.
Qiang Fu 0008, Dwi Sianto Mansjur, Biing-Hwang Juang
ICASSP3
2008 Blind speech dereverberation with multi-channel linear prediction based on short time fourier transform representation
abstract
It has recently been shown that the use of the time-varying nature of speech signals allows us to achieve high quality speech dereverberation based on multi-channel linear prediction (MCLP). However, this approach requires a huge computing cost for calculating large covariance matrices in the time domain. In addition, we face the important problem of how to combine the speech dereverberation efficiently with many other useful speech enhancement techniques in the short time Fourier transform (STFT) domain. As the first step to overcoming these problems, this paper presents methods for implementing MCLP based speech dereverberation that allow it to work in the STFT domain with much less computing cost. The effectiveness of the present methods is confirmed by experiments in terms of the recovered signal quality and the computing time.
Tomohiro Nakatani, Takuya Yoshioka, Keisuke Kinoshita, Masato Miyoshi, Biing-Hwang Juang
ICASSP5
2008 Towards robust acoustic echo cancellation during double-talk and near-end background noise via enhancement of residual echo
abstract
This paper examines the technique of using a noise suppressing nonlinearity in the adaptive filter error feedback loop of the acoustic echo canceler (AEC) based on the least mean square (LMS) algorithm when there are both double-talk and white background noise at the near-end. By combining the previously introduced noise suppressing technique with a compressive nonlinearity derived from the theory of robust statistics, consistently better results are obtained during double-talk as well as during single-talk when compared to the traditional approach of using only the compressive nonlinearity. It is shown that a compressive form of noise reducing nonlinearity can be derived also from the signal enhancement point of view when the noise probability density (pdf) is tailed more heavily and has a higher kurtosis than the Gaussian pdf. A combination of such a noise compressing nonlinearity and a noise suppressing nonlinearity is capable of producing results that are similar to that of the robust statistics approach during double-talk along with an added benefit of increased robustness during single-talk when there is only the background noise.
Ted S. Wada, Biing-Hwang Juang
ICASSP2
2008 Incremental parsing for latent semantic indexing of images
abstract
A generalized latent semantic analysis framework using a universal source coding algorithm for content-based image retrieval is proposed. By the multidimensional incremental parsing algorithm which is considered as a multidimensional extension of the Lempel-Ziv data compression method, a given image is compressed at a moderate bitrate while constructing the dictionary which implicitly embeds source statistics. Instead of concatenating all the corresponding dictionaries of an image corpus, we sequentially compress images using a previously constructed dictionary and end up with a visual lexicon which contains the least number of visual words covering all the images in the corpus. From the latent semantic analysis of the co-occurrence pattern of visual words over the images, a similarity between a given query and an image from the corpus is measured. An application of the proposed technique on a database of 20,000 natural scene images has demonstrated that the performance of the proposed system is favorable to that of existing approaches.
Soo Hyun Bae, Biing-Hwang Juang
ICIP2
2008 An Investigation of Non-Uniform Error Cost Function Design in Automatic Speech Recognition
abstract
The classical Bayes decision theory [3] is the foundation of statistical pattern recognition. In [4], we have addressed the issue of non-uniform error criteria in statistical pattern recognition, and generalized the Bayes decision theory for pattern recognition tasks where errors over different classes have varying degrees of significance. We further introduced the weighted minimum classification error (MCE) method for a practical design of a statistical pattern recognition system to achieve empirical optimality when non-uniform error criteria are prescribed. However, one key issue in the weighted MCE method, the methodology of building a suitable non-uniform error cost function given the userpsilas requirements, has not been addressed yet. In this paper, we propose some viable techniques for the design of the non-uniform error cost function in the context of automatic speech recognition (ASR) according to different training scenarios. The experimental results on the TIDIGITS database [8] are presented to demonstrate the effectiveness of our methodologies.
Qiang Fu 0008, Biing-Hwang Juang
ICMLA2
2008 Improving Kernel Density Classifier Using Corrective Bandwidth Learning with Smooth Error Loss Function
abstract
In this paper, we propose a corrective bandwidth learning algorithm for Kernel Density Estimation (KDE)-based classifiers. The objective of the corrective bandwidth learning algorithm is to minimize the expected error-rate. It utilizes a gradient descent technique to obtain the appropriate bandwidths. The proposed classifier is called the "Empirical Mixture Model" (EMM) classifier. Experiments were conducted on a set of multivariate multi-class classification problems with various data sizes. The proposed classifier has an error-rate closer to the true model compared to conventional KDE-based classifiers for both small and large data sizes. Additional experiments on standard machine learning datasets showed that the proposed bandwidth learning algorithm performed very well in gen-eral.
Dwi Sianto Mansjur, Biing-Hwang Juang
ICMLA2
2008 Utilizing non-uniform cost learning for active control of inter-class confusion
abstract
In this paper, we demonstrate the use of learning with non-uniform error-cost as a novel technique to design a multiclass cost-sensitive classifier. We investigate two important aspects of the design. First, we show that the learning is effective enough for active control of the multiclass confusion matrix using the cost-matrix. Second, we study the cases when the classifiers have mild model mismatch problems, and conclude that our design still have better performance compared to the conventional cost-sensitive classifier design.
Dwi Sianto Mansjur, Qiang Fu 0008, Biing-Hwang Juang
ICPR3
2008 Incremental learning of mixture models for simultaneous estimation of class distribution and inter-class decision boundaries
abstract
In this paper, we propose a novel design of high performance Bayes classifier from a small number of observations. The two main challenges to obtain the classifier are the lack of the true functional form of the class-conditional density and the lack of enough data to estimate the parameters of the classifiers. Incremental learning of Gaussian mixture model (GMM) is used to mitigate the lack of the true functional form. Moreover, the classifier uses the training samples from all classes to evaluate the goodness of a particular mixture to be used as the classifier for a specific class. This selection process eases the difficulty of the accurate parameter estimation. Thus, the important trait of the proposed classifier is being able to estimate simultaneously class-conditional density and inter-class boundaries to arbitrary precision. Our experimental results show that the proposed classifier not only has better performance than the conventional classifiers but also requires fewer parameters.
Dwi Sianto Mansjur, Biing-Hwang Juang
ICPR2
2008 Towards the integration of automatic speech recognition and information retrieval for spoken query processing
Antonio Moreno-Daniel, Jay G. Wilpon, Biing-Hwang Juang, Sarangarajan Parthasarathy
INTERSPEECH3
2008 Speech Dereverberation Based on Maximum-Likelihood Estimation With Time-Varying Gaussian Source Model
abstract
Distant acquisition of acoustic signals in an enclosed space often produces reverberant components due to acoustic reflections in the room. Speech dereverberation is in general desirable when the signal is acquired through distant microphones in such applications as hands-free speech recognition, teleconferencing, and meeting recording. This paper proposes a new speech dereverberation approach based on a statistical speech model. A time-varying Gaussian source model (TVGSM) is introduced as a model that represents the dynamic short time characteristics of nonreverberant speech segments, including the time and frequency structures of the speech spectrum. With this model, dereverberation of the speech signal is formulated as a maximum-likelihood (ML) problem based on multichannel linear prediction, in which the speech signal is recovered by transforming the observed signal into one that is probabilistically more like nonreverberant speech. We first present a general ML solution based on TVGSM, and derive several dereverberation algorithms based on various source models. Specifically, we present a source model consisting of a finite number of states, each of which is manifested by a short time speech spectrum, defined by a corresponding autocorrelation (AC) vector. The dereverberation algorithm based on this model involves a finite collection of spectral patterns that form a codebook. We confirm experimentally that both the time and frequency characteristics represented in the source models are very important for speech dereverberation, and that the prior knowledge represented by the codebook allows us to further improve the dereverberated speech quality. We also confirm that the quality of reverberant speech signals can be greatly improved in terms of the spectral shape and energy time-pattern distortions from simply a short speech signal using a speaker-independent codebook.
Tomohiro Nakatani, Biing-Hwang Juang, Takuya Yoshioka, Keisuke Kinoshita, Marc Delcroix, Masato Miyoshi
IEEE Trans. Speech Audio Process.2
2008 Multidimensional Incremental Parsing for Universal Source Coding
abstract
A multidimensional incremental parsing algorithm (MDIP) for multidimensional discrete sources, as a generalization of the Lempel-Ziv coding algorithm, is investigated. It consists of three essential component schemes, maximum decimation matching, hierarchical structure of multidimensional source coding, and dictionary augmentation. As a counterpart of the longest match search in the Lempel-Ziv algorithm, two classes of maximum decimation matching are studied. Also, an underlying behavior of the dictionary augmentation scheme for estimating the source statistics is examined. For an m-dimensional source, m augmentative patches are appended into the dictionary at each coding epoch, thus requiring the transmission of a substantial amount of information to the decoder. The property of the hierarchical structure of the source coding algorithm resolves this issue by successively incorporating lower dimensional coding procedures in the scheme. In regard to universal lossy source coders, we propose two distortion functions, the local average distortion and the local minimax distortion with a set of threshold levels for each source symbol. For performance evaluation, we implemented three image compression algorithms based upon the MDIP; one is lossless and the others are lossy. The lossless image compression algorithm does not perform better than the Lempel-Ziv-Welch coding, but experimentally shows efficiency in capturing the source structure. The two lossy image compression algorithms are implemented using the two distortion functions, respectively. The algorithm based on the local average distortion is efficient at minimizing the signal distortion, but the images by the one with the local minimax distortion have a good perceptual fidelity among other compression algorithms. Our insights inspire future research on feature extraction of multidimensional discrete sources.
Soo Hyun Bae, Biing-Hwang Juang
IEEE Trans. Image Process.2
2007 Automatic speech recognition based on weighted minimum classification error (W-MCE) training method
abstract
The Bayes decision theory is the foundation of the classical statistical pattern recognition approach. For most of pattern recognition problems, the Bayes decision theory is employed assuming that the system performance metric is defined as the simple error counting, which assigns identical cost to each recognition error. However, this prevalent performance metric is not desirable in many practical applications. For example, the cost of "recognition" error is required to be differentiated in keyword spotting systems. In this paper, we propose an extended framework for the speech recognition problem with non-uniform classification/recognition error cost. As the system performance metric, the recognition error is weighted based on the task objective. The Bayes decision theory is employed according to this performance metric and the decision rule with a non-uniform error cost function is derived. We argue that the minimum classification error (MCE) method, after appropriate generalization, is the most suitable training algorithm for the "optimal" classifier design to minimize the weighted error rate. We formulate the weighted MCE (W-MCE) algorithm based on the conventional MCE infrastructure by integrating the error cost and the recognition error count into one objective function. In the context of automatic speech recognition (ASR), we present a variety of training scenarios and weighting strategies under this extended framework. The experimental demonstration for large vocabulary continuous speech recognition is provided to support the effectiveness of our approach.
Qiang Fu 0008, Biing-Hwang Juang
ASRU2
2007 A study on rescoring using HMM-based detectors for continuous speech recognition
abstract
This paper presents an investigation of the rescoring performance using hidden Markov model (HMM) based attribute detectors. The minimum verification error (MVE) criterion is employed to enhance the reliability of the detectors in continuous speech recognition. The HMM-based detectors are applied on the possible recognition candidates, which are generated from the conventional decoder and organized in phone/word graphs. We focus on the study of rescoring performance with the detectors trained on the tokens produced by the decoder but labeled in broad phonetic categories rather than the phonetic identities. Various training criteria and knowledge fusion methods are investigated under various semantic level rescoring scenarios. This research demonstrates various possibilities of embedding auxiliary information into the current automatic speech recognition (ASR) framework for improved results. It also represents an intermediate step towards the construction of a true detection-based ASR paradigm.
Qiang Fu 0008, Biing-Hwang Juang
ASRU2
2007 Acoustic Model Enhancement: An Adaptation Technique for Speaker Verification Under Noisy Environments
abstract
This work presents an acoustic model adaptation method for speaker verification (SV) in environments with additive noise. In contrast to traditional acoustic model adaptation techniques that adapt the models parameters based on a model of the noise, acoustic model enhancement (AME) belongs to a new scheme in which the models are adapted to the speech enhancement strategy. The theoretical framework is presented for spectral subtraction (SS) as the enhancement technique and GMM as the acoustic models. In order to study the effect of additive noise only, a modified TIMIT dataset was used. The experimental setup uses two types of noise: one with fixed spectrum that helps as a proof of concept, and another with time-varying spectrum as a more realistic performance reference for AME. The results for this latter type show that at 20 dB SNR, the equal error rate (EER) dropped from 17% to around 8.9% when the noisy speech was enhanced with SS, whereas it further dropped to 8.1% with AME.
Antonio Moreno-Daniel, Juan A. Nolazco-Flores, Ted S. Wada, Biing-Hwang Juang
ICASSP (4)4
2007 Spoken Query Processing for Information Retrieval
abstract
This work proposes a way to integrate an information retrieval (IR) system with an automatic speech recognition (ASR) engine to support natural spoken queries. A broader interaction between the two modules is achieved by transmitting a lattice of terms to the IR system. This is in contrast with conventional systems where only the best-path recognition output is transmitted. Acoustic scores associated with the term-lattice are used to weigh the terms. A latent semantic indexing (LSI) scheme in which documents and terms are mapped to a single reduced feature-space with 400 semantic components is used. The conventional LSI method is nevertheless modified to allow the aforementioned broader interaction between acoustic hypothesis and semantic determination. The results show that the proposed method moderately outperforms the traditional approach for spoken queries formulated as casual phrases.
Antonio Moreno-Daniel, Sarangarajan Parthasarathy, Biing-Hwang Juang, Jay G. Wilpon
ICASSP (4)3
2007 Study on Speech Dereverberation with Autocorrelation Codebook
abstract
This paper proposes a new speech dereverberation approach based on a statistical speech model. An autocorrelation codebook is introduced as a model that can represent time-varying short-time speech characteristics corresponding to the cepstrum and harmonics. The speech dereverberation is formulated as a likelihood maximization problem, in which the quality of a speech signal is recovered by turning the signal into one that is probabilistically more like a clean speech. Two dereverberation algorithms are derived based on different scenarios, regularized inversion and inverse filter estimation. Experimental results show that the proposed approach allows us to reduce both reverberation and noise with the regularized inversion, and to estimate inverse filters that can dereverberate signals effectively from just a small number of observed signals.
Tomohiro Nakatani, Biing-Hwang Juang, Takafumi Hikichi, Takuya Yoshioka, Keisuke Kinoshita, Marc Delcroix, Masato Miyoshi
ICASSP (1)2
2007 Blind Source Separation of Acoustic Mixtures with Distributed Microphones
abstract
The problem of blind source separation of acoustic mixtures is often addressed using independent component analysis in the frequency domain. One problem with this approach is the inconsistency across frequency in the permutation of the source estimates. Solutions to this problem have been proposed that exploit known properties of both the source signals and the mixing system, but require the microphones to be in a constrained geometry. In this paper a solution is presented that avoids this constraint by extracting information from the magnitude of the mixing system instead of its phase. The new method combines that information with information from the source estimates to provide a reliable permutation alignment.
Enrique Robledo-Arnuncio, Biing-Hwang Juang
ICASSP (3)2
2007 An overview on automatic speech attribute transcription (ASAT)
abstract
Automatic Speech Attribute Transcription (ASAT), an ITR project sponsored under the NSF grant (IIS-04-27113), is a cross-institute effort involving Georgia Institute of Technology, The Ohio State University, University of California at Berkeley, and Rutgers University. This project approaches speech recognition from a more linguistic perspective: unlike traditional ASR systems, humans detect acoustic and auditory cues, weigh and combine them to form theories, and then process these cognitive hypotheses until linguistically and pragmatically consistent speech understanding is achieved. A major goal of the ASAT paradigm is to develop a detection-based approach to automatic speech recognition (ASR) based on attribute detection and knowledge integration. We report on progress of the ASAT project, present a sharable platform for community collaboration, and highlight areas of potential interdisciplinary ASR research. Index Terms: attributes, events, features, detection, speech recognition, speech attribute transcription, utterance verification
Mark A. Clements, Sorin Dusan, Eric Fosler-Lussier, Keith Johnson, Biing-Hwang Juang, Lawrence R. Rabiner
INTERSPEECH6
2007 Joint Source-Channel Modeling and Estimation for Speech Dereverberation
abstract
Speech dereverberation is an important challenge in acoustic signal processing because of the detrimental effects upon the signal quality brought by the reverberant components due to acoustic reflections from the walls of the room enclosure. Most of the methods proposed so far such as microphone array beamforming and room impulse response inversion for deconvolution make very little use of the knowledge about the characteristics of the source signal. This paper presents a formulation of speech dereverberation as a probabilistic modeling problem in which joint estimation of the source and the channel (or its inverse) can be accomplished. We discuss reasonable representations of the source prior for inclusion in such a formulation and propose the corresponding solutions.
Biing-Hwang Juang, Tomohiro Nakatani
ISCAS1
2007 Robust blind dereverberation of speech signals based on characteristics of short-time speech segments
abstract
This paper addresses blind dereverberation techniques based on the inherent characteristics of speech signals. Two challenging issues for speech dereverberation involve decomposing reverberant observed signals into colored sources and room transfer functions (RTFs), and making the inverse filtering robust as regards acoustic and system noise. We show that short-time speech characteristics are very important for this task, and that multi-channel linear prediction (MCLP) is a useful tool for achieving robust inverse filtering. As examples, we detail three recently proposed robust dereverberation methods. By assuming the source to be a small order autoregressive process, we can present an efficient source estimation method that reduces late reverberation reflections of the reverberation using multi-step linear prediction. By exploiting the time-varying nature of the speech signals, we can also develop a method that can estimate both the source and the inverse filters of the RTFs. Furthermore, we can achieve high quality speech dereverberation by formulating the problem as a likelihood maximization problem using a statistical speech model that represents the spectral characteristics of short-time speech segments including harmonicity.
Tomohiro Nakatani, Takafumi Hikichi, Keisuke Kinoshita, Takuya Yoshioka, Marc Delcroix, Masato Miyoshi, Biing-Hwang Juang
ISCAS7
2007 Speech Analysis in a Model of the Central Auditory System
abstract
Recently, there is a significant increase in research interest in the area of biologically inspired systems, which, in the context of speech communications, attempt to learn from human's auditory perception and cognition capabilities so as to derive the knowledge and benefits currently unavailable in practice. One particular pursuit is to understand why the human auditory system generally performs with much more robustness than an engineering system, say a state-of-the-art automatic speech recognizer. In this study, we adopt a computational model of the mammalian central auditory system and develop a methodology to analyze and interpret its behavior for an enhanced understanding of its end product, which is a data-redundant, dimension-expanded representation of neural firing rates in the primary auditory cortex (A1). Our first approach is to reinterpret the well-known Mel-frequency cepstral coefficients (MFCCs) in the context of the auditory model. We then present a framework for interpreting the cortical response as a place-coding of speech information, and identify some key advantages of the model's dimension expansion. The framework consists of a model of ldquosourcerdquo-invariance that predicts how speech information is encoded in a class-dependent manner, and a model of ldquoenvironmentrdquo-invariance that predicts the noise-robustness of class-dependent signal-respondent neurons. The validity of these ideas are experimentally assessed under existing recognition framework by selecting features that demonstrate their effects and applying them in a conventional phoneme classification task. The results are quantitatively and qualitatively discussed, and our insights inspire future research on category-dependent features and speech classification using the auditory model.
Woojay Jeon, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.2
2006 Spatial Resolution and Quantization Noise Tradeoffs for Scalable Image Compression
abstract
Most full-reference quality metrics compare the original image to a distorted image at the same level of resolution assuming a fixed viewing distance. In video streaming applications, however, the transmitted or received signal may differ from the original in compression as well as spatiotemporal resolution. For example, at low bitrate coding applications the compressed image may be too distorted, and hence the observer may prefer to reduce the resolution or increase the viewing distance in order to reduce the visibility of the compression artifacts. The selection of the best tradeoff between resolution/viewing distance and visibility of compression artifacts requires a quality metric that accounts for both image distortions and image size. Such tradeoffs are not reflected in existing quality metrics, which ignore the signal visibility and only measure the visibility of compression distortions, which decrease with image size. In order to better understand such tradeoffs, with the goal of developing better quality metrics, we conducted subjective tests using a number of existing still image coders (JPEG2000 SPHIT, and JPEG). Our results indicate that the objective quality (perceptually weighted PSNR) of the images that the viewers select decreases with resolution, that is, the viewers are willing to accept more artifacts as image size decreases
Soo Hyun Bae, Thrasyvoulos N. Pappas, Biing-Hwang Juang
ICASSP (2)3
2006 Separation of Snr Via Dimension Expansion in a Model of the Central Auditory System
abstract
In this study, we provide a theoretical approach for analyzing signal and noise separation and the noise-robustness of class-dependent activation areas in a model of the primary auditory cortex in the central auditory system. Specifically, we interpret the auditory model as a system of localized matched filters that act as a place-coding mechanism for mapping signal and noise spectra into separate regions in the three-dimensional cortical space. This framework allows us to analyse the noise robustness of signal-respondent neurons by computing their signal-to-noise ratio (SNR)'s without having to explicitly consider the complex mathematical expressions for the auditory model. The framework is also fundamentally consistent with the notion of category-dependence proposed in our previous work. Our theoretical developments of the place-coding effect and the separation of SNR is also demonstrated experimentally
Woojay Jeon, Biing-Hwang Juang
ICASSP (1)2
2006 Speech Dereverberation Based on Probabilistic Models of Source and Room Acoustics
abstract
This paper proposes a new single channel speech dereverberation method, in which the features of source signals and room acoustics are represented by probabilistic density functions (pdf) and the source signals are estimated by maximizing a likelihood function defined based on the pdfs. Two types of pdfs are introduced for the source signals, based on two essential speech signal features, harmonicity and sparseness, while the pdf for the room acoustics is defined based on an inverse filtering operation. The EM algorithm is used to solve this maximum likelihood problem efficiently. The resultant algorithm elaborates the initial source signal estimate given solely based on its source signal features by integrating them with the room acoustics feature through the EM iteration. The effectiveness of the present method is shown in terms of the energy decay curves of the dereverberated impulse responses
Tomohiro Nakatani, Biing-Hwang Juang, Keisuke Kinoshita, Masato Miyoshi
ICASSP (1)2
2006 Subjective Image Quality Tradeoffs Between Spatial Resolution and Quantization Noise
abstract
The importance of tradeoffs between spatial resolution and quantization noise has been examined in our previous work. Subjective experiments indicate that as the bitrate decreases, human observers generally prefer to reduce image resolution in order to maintain image quality, but the amount of distortion they are willing to accept increases with decreasing resolution. In this paper, we conducted further experiments with several images, different encoders, and a finer set of bitrates to determine the preferred resolution at each bit rate, and also the resolution at which there are no visible coding artifacts. Analysis of the subjective results using a wavelet-based perceptual quality metric verifies our earlier conclusion that human observers tend to reduce resolution in order to maintain image quality, but are willing to accept more artifacts as image size decreases.
Soo Hyun Bae, Thrasyvoulos N. Pappas, Biing-Hwang Juang
ICIP3
2006 Design of a Transmission Protocol for a CVE
abstract
Virtual reality is a useful tool to train individuals for various different situations or scenarios. A single person utilizing a virtual environment can greatly enhance their skill set, but the greatest gain in skill is obtained when multiple users realistically interact in the virtual environment. To accomplish this, a network protocol must be designed to efficiently communicate vital information to participants in the virtual environment. The purpose of this paper is to document the design of a transmission protocol for a collaborative virtual environment (CVE). In addition, the network protocol was tested with a network emulator to understand how well the protocol functioned under various network conditions. When the jitter level was below 50 ms the transmission protocol was able to reproduce the smooth avatar movement even at transmission intervals as high as 100 ms. However as the jitter increased the transmission interval had to be kept as low as 10 ms to reproduce realistic human movement.
Fred Stakem, Ghassan Al-Regib, Biing-Hwang Juang, Mourad Bouzit
ICIP3
2006 Investigation on rescoring using minimum verification error (MVE) detectors
Qiang Fu 0008, Biing-Hwang Juang
INTERSPEECH2
2006 Generalization of the minimum classification error (MCE) training based on maximizing generalized posterior probability (GPP)
Qiang Fu 0008, Antonio Moreno-Daniel, Biing-Hwang Juang, Jian-Lai Zhou, Frank K. Soong
INTERSPEECH3
2006 Study of a Fast Discriminative Training Algorithm for Pattern Recognition
abstract
Discriminative training refers to an approach to pattern recognition based on direct minimization of a cost function commensurate with the performance of the recognition system. This is in contrast to the procedure of probability distribution estimation as conventionally required in Bayes' formulation of the statistical pattern recognition problem. Currently, most discriminative training algorithms for nonlinear classifier designs are based on gradient-descent (GD) methods for cost minimization. These algorithms are easy to derive and effective in practice, but are slow in training speed and have difficulty selecting the learning rates. To address the problem, we present our study on a fast discriminative training algorithm. The algorithm initializes the parameters by the expectation-maximization (EM) algorithm, and then uses a set of closed-form formulas derived in this paper to further optimize a proposed objective of minimizing error rate. Experiments in speech applications show that the algorithm provides better recognition accuracy in a fewer iterations than the EM algorithm and a neural network trained by hundreds of GD iterations. Although some convergent properties need further research, the proposed objective and derived formulas can benefit further study of the problem.
Biing-Hwang Juang
IEEE Trans. Neural Networks2
2005 3CCD interpolation using selective projection
abstract
The emergence of HDTV accelerates the evolution of high-resolution imaging systems. A 3CCD digital camera system has been developed for higher resolution than one CCD imaging has. From the pixel correlation caused by a half-pixel shift of the green channel, we can interpolate pixels and get four times higher resolution of the color image. The proposed method involves three projection operators. The first is to reduce aliasing of image regions by selective projection in subband channels. The second projection makes an inverse of the MTF which generates blurring over the entire image. The last operator works for fast convergence. From experimental results, the proposed algorithm shows suppression of jagging effects and restoration of aliased image regions. It is experimentally shown that the projection process converges and is almost finished at the first iteration.
Soo Hyun Bae, Moon-Cheol Kim, Biing-Hwang Juang
ICASSP (2)3
2005 A Study of Auditory Modeling and Processing for Speech Signals
abstract
We study a modified version of a computational model of the human peripheral and central auditory system (Wang, K. and Shamma, S.A., 1995; Yang, X. et al., 1992), and examine the validity of its output from two practical perspectives. One considers the well-known Mel-frequency cepstral coefficients (MFCC) as an approximate representation of the physiology-based early auditory processing result. The other allows the derivation of feature vectors from the dimension expanded cortical response of the central auditory system for use in a conventional phoneme recognition task. In addition to confirming the relevancy of the model under an existing statistical speech recognition framework, we conduct a preliminary study of the cortical response in connection with known physiological studies, to find new possibilities in using the auditory model to perform cognitive functions based on a better understanding of the human auditory system. In particular, the cortical response may be a place-coded data set where sounds are categorized according to the regions containing their most distinguishing features. The results of this study encourage us to develop hierarchical, detection-based methods in which this mechanism may be utilized to simulate a variety of human perceptual and cognitive functions.
Woojay Jeon, Biing-Hwang Juang
ICASSP (1)2
2005 Robustness of Bit-stream Based Features for Speaker Verification
abstract
The paper presents a speaker verification system that uses the YOHO database which has been coded to the ITU-T G.729 standard. A set of bitstream based features, consisting of 16 LPC cepstral coefficients and MFCC derived from the quantized line spectral pairs as well as residual information in the form of pitch, was utilized to construct the speakers' models, and their robustness was studied under white noise conditions. Results suggest that, using a cohort model, MFCC are more robust under noise conditions than LPC cepstral coefficients; the addition of pitch to the feature vector contributes from a 16% to a 29% of improvement in verification performance under different noise conditions.
Antonio Moreno-Daniel, Biing-Hwang Juang, Juan A. Nolazco-Flores
ICASSP (1)2
2005 Issues in frequency domain blind source separation - a critical revisit
abstract
One of the most important problems in frequency domain blind source separation (FDBSS) is the inconsistency across frequency in the permutation of the source estimates. According to previous studies, this problem can be reduced significantly by constraining the length of the unmixing filters. This improvement has been attributed to the smoothing of the unmixing frequency response. In this paper we study the effect of modifying these length constraints taking into account the circularity of the IDFT, and we show that the smoothing of the unmixing frequency response alone can not account for the improvements in performance.
Enrique Robledo-Arnuncio, Biing-Hwang Juang
ICASSP (5)2
2005 Segment-based phonetic class detection using minimum verification error (MVE) training
Qiang Fu 0008, Biing-Hwang Juang
INTERSPEECH2
2005 A category-dependent feature selection method for speech signals
Woojay Jeon, Biing-Hwang Juang
INTERSPEECH2
2005 Using inter-frequency decorrelation to reduce the permutation inconsistency problem in blind source separation
Enrique Robledo-Arnuncio, Biing-Hwang Juang
INTERSPEECH2
2004 Speaker Verification Using Coded Speech
Antonio Moreno-Daniel, Biing-Hwang Juang, Juan A. Nolazco-Flores
CIARP2
2003 Fast discriminative training for sequential observations with application to speaker identification
abstract
This paper presents a fast discriminative training algorithm for sequences of observations. It considers a sequence of feature vectors as one single composite token in training or testing. In contrast to the traditional EM algorithm, this algorithm is derived from a discriminative objective, aiming at directly minimizing the recognition error. Compared to the gradient-descent algorithms for discriminative training, this algorithm invokes a mild assumption which leads to closed-form formulas for re-estimation, rather than relying on gradient search, without sacrificing the algorithmic rigor. As such, it is in general much faster than a descent based algorithm and does not need to determine the learning rate or step size. Our experiment shows that the proposed algorithm reduces error rate by 14.65, 66.46, and 100.00% for 1, 5, and 10 seconds of testing data respectively, in a speaker identification application.
Biing-Hwang Juang
ICASSP (2)2
2002 A new algorithm for fast discriminative training
abstract
Currently, almost all discriminative training algorithms for nonlinear classifier design are based on gradient-descent methods, such as the backpropagation and the generalized probabilistic descent algorithm. These algorithms are easy to derive and effective in applications. However, a drawback for the gradient-descent approaches is the slow training speed, which limits their applications in large training problems, such as large vocabulary speech recognition and other applications. For hidden Markov models, some training algorithms, such as the reestimation (or expectation-maximization) algorithm for maximum likelihood estimation (MLE), are fast, but they are not readily extendible to discriminative training for recognition performance improvements. To address the problem, we proposed a fast discriminative training algorithm in this paper. It is a batch-mode algorithm derived for the objective function of minimal error rate. The significant advantage is its closed-form solution for parameter estimation during iterations, instead of incremental search in the direction of gradient, as conventionally done. We experimentally show that the algorithm requires only a few iterations to achieve the optimization objective and that the estimated results lead to better recognition performance than a traditional MLE.
Biing-Hwang Juang
ICASSP2
2002 Classifier design for verification of multi-class recognition decision
abstract
This paper investigates a 2-class classifier approach with the aim of improving the word verification performance. The classifier operates on a discriminant function which is a linear combination of the smoothed likelihood ratios for the N-best candidates and the background (BG) and out-of-vocabulary (OOV) filler models, and is optimized using discriminative training to minimize the classification error. This paper discusses several strategies involving the likelihood ratio based formulation and the use of N-best candidates and the BG and OOV models in the classifier. In word verification experiments using a connected-digit database containing utterances recorded in a moving car with a hands-free microphone, the likelihood ratio based formulation achieved a relative error reduction of 35% in comparison with a likelihood based formulation. In addition, we observed that the use of N-best candidates and the BG and OOV models improved the performance with a relative error reduction of roughly 10%.
Tomoko Matsui, Frank K. Soong, Biing-Hwang Juang
ICASSP3
2002 Multiple description speech coding with diversities
abstract
Multiple description coding (MDC) allows management of tradeoff in achievable rate-distortion performance among various components in a processing (coding-decoding) system or channels in a transmission scheme. Current MDC methods first fix performance of the central decoder at a conventional single description (SD) level. and then inject inter-channel redundancy to side coders to provide the needed robustness in case of information loss during transmission. In this work. we relegate the SD structure and quality to side decoders, and concentrate on improving the central decoder beyond the known SD quality. We provide justification for this alternative, state its inherent nature and challenge, and propose an efficient method to design MD coder under such criterion.
Biing-Hwang Juang
ICASSP2
2002 Technical advances in digital audio radio broadcasting
abstract
The move to digital is a natural progression taking place in all aspects of broadcast media applications from document processing in newspapers to video processing in television distribution. This is no less true for audio broadcasting which has taken a unique development path in the United States. This path has been heavily influenced by a combination of regulatory and migratory requirements specific to the U.S. market. In addition, competition between proposed terrestrial and satellite systems combined with increasing consumer expectations have set ambitious, and often changing, requirements for the systems. The result has been a unique set of evolving requirements on source coding, channel coding, and modulation technologies to make these systems a reality. This paper outlines the technical development of the terrestrial wireless and satellite audio broadcasting systems in the U.S., providing details on specific source and channel coding designs and adding perspective on why specific designs were selected in the final systems. These systems are also compared to other systems such as Eureka-147, DRM, and Worldspace, developed under different requirements.
Christoph Faller, Biing-Hwang Juang, Peter Kroon, Hui-Ling Lou, Sean A. Ramprashad, Carl-Erik W. Sundberg
Proc. IEEE2
2001 Introduction: A Simple Complex in Artificial Intelligence and Machine Learning
Biing-Hwang Juang
Int. J. Pattern Recognit. Artif. Intell.1
2001 An application of discriminative feature extraction to filter-bank-based speech recognition
abstract
A pattern recognizer is usually a modular system which consists of a feature extractor module and a classifier module. Traditionally, these two modules have been designed separately, which may not result in an optimal recognition accuracy. To alleviate this fundamental problem, the authors have developed a design method, named discriminative feature extraction (DFE), that enables one to design the overall recognizer, i.e., both the feature extractor and the classifier, in a manner consistent with the objective of minimizing recognition errors. This paper investigates the application of this method to designing a speech recognizer that consists of a filter-hank feature extractor and a multi-prototype distance classifier. Carefully investigated experiments demonstrate that DFE achieves the design of a better recognizer and provides an innovative recognition-oriented analysis of the filter-bank, as an alternative to conventional analysis based on psychoacoustic expertise or heuristics.
Alain Biem, Shigeru Katagiri, Erik McDermott, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.4
2001 Why speech synthesis? (in memory of Prof. Jonathan Allen, 1934-2000)
abstract
Normally, this transactions has a history of attracting world-class papers in speech processing, mostly in areas such as analysis, enhancement, coding, and recognition, and, to a lesser extent, synthesis. The few published papers that we believe are relevant to speech synthesis are mainly concerned with speech modifications and speech resynthesis from given speech parameter vectors. Linguistic and system issues, synthesizers, "letter-to-sound" (pronunciation) rules, generation (from text) of proper prosody (contours of fundamental frequency, phoneme durations, and amplitudes) to make the synthetic speech pragmatically natural, to name a few, were only sparsely addressed. The editors go about explaining what has changed so that they can justify a Special Issue on Speech Synthesis, and explain what "signal processing" has to do with it. In addition, after summarizing the specila issue contents, it is nooted that this special issue is dedicated to the memory of Prof. Jonathan Allen who died on April 24, 2000. Prof. Allen was Director of the Research Laboratory of Electronics (RLE), Massachusetts Institute of Technology (MIT), Cambridge, from 1981 until his death. Many knew him as a friend and a leader in speech research, and, most notably and specifically, as a true pioneer in speech synthesis. A brief biography is given highlighting his professional achievements.
Biing-Hwang Juang
IEEE Trans. Speech Audio Process.1
2001 Speech recognition and utterance verification based on a generalized confidence score
abstract
In this paper, we introduce a generalized confidence score (GCS) function that enables a framework to integrate different confidence scores in speech recognition and utterance verification. A modified decoder based on the GCS is then proposed. The GCS is defined as a combination of various confidence scores obtained by exponential weighting from various confidence information sources, such as likelihood, likelihood ratio, duration, language model probabilities, etc. We also propose the use of a confidence preprocessor to transform raw scores into manageable terms for easy integration. We consider two kinds of hybrid decoders, an ordinary hybrid decoder and an extended hybrid decoder, as implementation examples based on the generalized confidence score. The ordinary hybrid decoder uses a frame-level likelihood ratio in addition to a frame-level likelihood, while a conventional decoder uses only the frame likelihood or likelihood ratio. The extended hybrid decoder uses not only the frame-level likelihood but also multilevel information such as frame-level, phone-level, and word-level confidence scores based on the likelihood ratios. Our experimental evaluation shows that the proposed hybrid decoders give better results than those obtained by the conventional decoders, especially in dealing with ill-formed utterances that contain out-of-vocabulary words and phrases.
Myoung-Wan Koo, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.3
2000 Special issue on spoken language processing
Biing-Hwang Juang, Sadaoki Furui
Proc. IEEE1
2000 Automatic recognition and understanding of spoken language - a first step toward natural human-machine communication
abstract
The promise of a powerful computing device to help people in productivity as well as in recreation can only be realized with proper human-machine communication. Automatic recognition and understanding of spoken language is the first step toward natural human-machine interaction. Research in this field has produced remarkable results, leading to many exciting expectations and new challenges. We summarize the development of the spoken language technology from both a vertical (chronology) and a horizontal (spectrum of technical approaches) perspective. We highlight the introduction of statistical methods in dealing with language-related problems, as this represents a paradigm shift in the research field of spoken language processing. Statistical methods are designed to allow the machine to learn structural regularities in the speech signal, directly from data, for the purpose of automatic speech recognition and understanding. Research results in spoken language processing have led to a number of successful applications, ranging from dictation software for personal computers and telephone-call processing systems for automatic call routing, to automatic sub-captioning for television broadcasts. We analyze the technical successes that support these applications. Along with an assessment of the state of the art in this broad technical field, we also discuss the limitations of the current technology, and point out the challenges that are ahead. This paper presents an accurate overview of spoken language technology as a basis to inspire future advances.
Biing-Hwang Juang, Sadaoki Furui
Proc. IEEE1
2000 Automatic verbal information verification for user authentication
abstract
Traditional speaker authentication focuses on speaker verification (SV) and speaker identification, which is accomplished by matching the speaker's voice with his or her registered speech patterns. In this paper, we propose a new technique, verbal information verification (VIV), in which spoken utterances of a claimed speaker are verified against the key (usually confidential) information in the speaker's registered profile automatically; to decide whether the claimed identity should be accepted or rejected. Using the proposed sequential procedure involving three question-response turns, we achieved an error-free result in a telephone speaker authentication experiment with 100 speakers. We further propose a speaker authentication system by combining VIV with SV. In the system, a user is verified by VIV in the first four to five accesses, usually from different acoustic environments. During these uses, one of the key questions pertains to a pass-phrase for SV. The VIV system collects and verifies the pass-phrase utterance for use as training data for speaker model construction. After a speaker-dependent model is constructed, the system then migrates to SV. This approach avoids the inconvenience of a formal enrollment procedure, ensures the quality of the training data for SV, and mitigates the mismatch caused by different acoustic environments between training and testing. Experiments showed that the proposed system improved the SV performance by over 40% in equal-error rate compared to a conventional SV system.
Biing-Hwang Juang
IEEE Trans. Speech Audio Process.2
1999 A block least squares approach to acoustic echo cancellation
abstract
We propose an efficient block least squares (BLS) algorithm for acoustic echo cancellation. The high computation and memory requirements associated with a long room echo make the simple, gradient-based LMS filter a more acceptable commercial solution than a full-fledged LS canceler. However, the LMS echo canceler has slower convergence and worse steady-state performance than its LS counterpart. In the proposed BLS approach, the autocorrelation and cross-correlation of the source and echo, required in solving the LS normal equations, are performed once per block using FFTs. With appropriate data windowing the autocorrelation matrix is constrained to be Toeplitz, allowing the corresponding normal equations to be solved efficiently. The positive definiteness of the autocorrelation function eliminates the stability problems of other fast LS algorithms. BLS can reduce the echo residual to the level of background noise, allowing a residual power based, statistical near-end speech detector to be devised. Performance in real environments under various settings of filter length, SNR, near-end speech presence, etc., is investigated.
Eric A. Woudenberg, Frank K. Soong, Biing-Hwang Juang
ICASSP3
1998 A new decoder based on a generalized confidence score
abstract
We propose a new decoder based on a generalized confidence score. The generalized confidence score is defined as a product of confidence scores obtained from confidence information sources such as likelihood, likelihood ratio, duration, duration ratio, language model probabilities, supra-segmental information, etc. All confidence information sources are converted into confidence scores by a confidence pre-processor. We show an extended hybrid decoder as an example of the decoder based on the generalized confidence score. The extended hybrid decoder uses multi-level confidence scores such as frame-level, phone-level, and word-level likelihood ratios, while the conventional hybrid decoder uses the frame-level confidence score. The experimental result shows that the extended decoder gives a better result than the conventional hybrid decoder, particularly in dealing with out-of-vocabulary words or out-of-task sentences.
Myoung-Wan Koo, Biing-Hwang Juang
ICASSP3
1998 Speaker verification using verbal information verification for automatic enrolment
abstract
A conventional speaker verification (SV) system needs an enrolment session to collect the training data. Li et al. (1997) introduced a speaker authentication method called verbal information verification (VIV) which verifies a speaker by verbal contents instead of speech characteristics. Such a system does not need an enrolment session. In this paper, VIV is combined with SV. We propose a system which uses VIV to collect training data during the first few accesses automatically, which are often from different acoustic environments. Then, a speaker dependent model is trained and speaker authentication can be performed by SV. This approach not only avoid formal enrolment session which brings convenience to the user, but mitigates the mismatch problem causing by different acoustic environments between training and test sessions. Our experiments show that the proposed system improved the SV performance over 40% compared to the conventional SV system.
Biing-Hwang Juang
ICASSP2
1998 Statistical modeling of pronunciation and production variations for speech recognition
abstract
In this paper, we propose a procedure for training a pronunciation network with criteria consistent with the optimality objectives for speech recognition systems. In particular, we describe a framework for using maximum likelihood(ML) and minimum classi cation error(MCE) criteria for pronunciation network optimization. The ML criterion is used to obtain an optimal structure for the pronunciation network based on statistically-derived phonological rules. Discrimination among di erent pronunciation networks is achieved by weighting of the pronunciation networks, optimized by applying the MCE criterion. Experinent results demonstrate improvements in speech recognition accuracy after applying statistically derived phonological rules. It is shown that the impact of the pronunciation network weighting on the recognition performance is determined by the size of the recognition vocabulary. 1.
Filipp Korkmazskiy, Biing-Hwang Juang
ICSLP2
1998 Context dependent anti subword modeling for utterance verification
Padma Ramesh, Biing-Hwang Juang
ICSLP3
1998 Pattern recognition using a family of design algorithms based upon the generalized probabilistic descent method
abstract
This paper provides a comprehensive introduction to a novel approach to pattern recognition which is based on the generalized probabilistic descent method (GPD) and its related design algorithms. The paper contains a survey of recent recognizer design techniques, the formulation of GPD, the concept of minimum classification error learning that is closely related to the GPD formalization, a relational analysis between GPD and other important design methods, and various embodiments of GPD-based design, including segmental-GPD, minimum spotting error training, discriminative utterance verification, and discriminative feature extraction. GPD development has its origins in basic pattern recognition and Bayes decision theory. It represents a simple but careful re-investigation of the classical theory and successfully leads to an innovative framework. For clarity of presentation, detailed discussions about its embodiments are provided for examples of speech pattern recognition tasks that use a distance-based classifier. Experimental results in speech pattern recognition tasks clearly demonstrate the remarkable utility of the family of GPD-based design algorithms.
Shigeru Katagiri, Biing-Hwang Juang
Proc. IEEE2
1998 Flexible speech understanding based on combined key-phrase detection and verification
abstract
We propose a novel speech understanding strategy based on combined detection and verification of semantically tagged key-phrases in spontaneous spoken utterances. Key-phrases are defined in a top-down manner so as to constitute semantic slots. Their detection directly leads to robust understanding. A phrase network realizes both a wide coverage and a reasonable constraint for detection. A subword-based verifier is then incorporated to reduce false alarms in detection and attach confidence measures of the detected phrases. This set of phrase confidence measures, when incorporated in a spoken dialogue system, forms a basis for designing intelligent speech interfaces that accept only verified key-phrases and reprompt users to clarify unspecified or unrecognized portions. Several forms of confidence measures based on subword-level tests are investigated. The proposed approach was tested on field data collected from real-world trial applications. The combined detection and verification strategy drastically improves the accuracy in handling out-of-grammar utterances over the conventional decoding approaches while maintaining the performance for in-grammar utterances.
Tatsuya Kawahara, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.3
1997 Combining key-phrase detection and subword-based verification for flexible speech understanding
abstract
A flexible speech understanding framework combining key-phrase detection and verification is presented. Detection of semantically-tagged key-phrases directly leads to robust understanding. In order to select reliable detection and eliminate false alarms, utterance verification technique is incorporated. A phrase verifier combines subword-based likelihood ratios of correct models and anti-subword alternate models. A confidence measure that focuses on mis-matched subwords is proposed and demonstrated as the most effective. The combined strategy drastically improves the semantic accuracy for out-of-grammar utterances, while maintaining the performance for in-grammar samples. We also found that utterance verification applied after grammar-based decoding is not so effective as the proposed detection and verification strategy.
Tatsuya Kawahara, Biing-Hwang Juang
ICASSP3
1997 Generalized mixture of HMMs for continuous speech recognition
abstract
This paper presents a new technique for modeling heterogeneous data sources such as speech signals received via distinctly different channels. Such a scenario arises when an automatic speech recognition system is deployed in wireless telephony in which highly heterogeneous channels coexist and interoperate. The problem is that a simple model may become inadequate to describe accurately the diversity of the signal, resulting in an unsatisfactory recognition performance. To deal with such a problem, we propose a generalized mixture model (GMM) approach. For speech signals, in particular, we use mixtures of hidden Markov models (i.e., GMHMM, generalized mixture of HMMs). By applying discriminative training for GMHMM we obtained 1.0% word error rate for the recognition of the digits strings from the wireless database, comparing to 1.4% word error rate for the conventional HMM based discriminative technique.
Filipp Korkmazskiy, Biing-Hwang Juang, Frank K. Soong
ICASSP2
1997 Verbal information verification
Biing-Hwang Juang, Qiru Zhou
EUROSPEECH2
1997 Tribute to James L. Flanagan
Biing-Hwang Juang, Christel Sorin, Sadaoki Furui, Louis C. W. Pols
Speech Commun.1
1997 Filtering the time sequences of spectral parameters for speech recognition,
Climent Nadeu, Pau Pachès-Leal, Biing-Hwang Juang
Speech Commun.3
1997 Selective feature extraction via signal decomposition
abstract
In this article, a mathematical framework that jointly optimizes the parameters of classifier and feature extractor is presented. In this approach, feature extraction is formulated as a process of projecting the signals onto a smaller subspace in which the statistical properties of the signal can be efficiently modeled. An algorithm, called statistical matching pursuit (SMP), is proposed to learn from the training data the optimal projection dimensions and the extent of signal reduction. The algorithm is designed to achieve unconditional convergence and can be seamlessly incorporated into the expectation-maximization (EM) algorithm employed to train the classifier. Finally, we report some experimental results on speech recognition and elaborate the potential of the proposed method.
Kuansan Wang, Biing-Hwang Juang
IEEE Signal Process. Lett.3
1997 Minimum classification error rate methods for speech recognition
abstract
A critical component in the pattern matching approach to speech recognition is the training algorithm, which aims at producing typical (reference) patterns or models for accurate pattern comparison. In this paper, we discuss the issue of speech recognizer training from a broad perspective with root in the classical Bayes decision theory. We differentiate the method of classifier design by way of distribution estimation and the discriminative method of minimizing classification error rate based on the fact that in many realistic applications, such as speech recognition, the real signal distribution form is rarely known precisely. We argue that traditional methods relying on distribution estimation are suboptimal when the assumed distribution form is not the true one, and that "optimality" in distribution estimation does not automatically translate into "optimality" in classifier design. We compare the two different methods in the context of hidden Markov modeling for speech recognition. We show the superiority of the minimum classification error (MCE) method over the distribution estimation method by providing the results of several key speech recognition experiments. In general, the MCE method provides a significant reduction of recognition error rate.
Biing-Hwang Juang, Wu Hou
IEEE Trans. Speech Audio Process.1
1997 Discriminative utterance verification for connected digits recognition
abstract
Utterance verification represents an important technology in the design of user-friendly speech recognition systems. It involves the recognition of keyword strings and the rejection of nonkeyword strings. This paper describes a hidden Markov model-based (HMM-based) utterance verification system using the framework of statistical hypothesis testing. The two major issues on how to design keyword and string scoring criteria are addressed. For keyword verification, different alternative hypotheses are proposed based on the scores of antikeyword models and a general acoustic filler model. For string verification, different measures are proposed with the objective of detecting nonvocabulary word strings and possibly erroneous strings (so-called putative errors). This paper also motivates the need for discriminative hypothesis testing in verification. One such approach based on minimum classification error training is investigated in detail. When the proposed verification technique was integrated into a state-of-the-art connected digit recognition system, the string error rate for valid digit strings was found to decrease by 57% when setting the rejection rate to 5%. Furthermore, the system was able to correctly reject over 99.9% of nonvocabulary word strings.
Mazin G. Rahim, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.3
1996 Discriminative utterance verification using minimum string verification error (MSVE) training
abstract
This paper focuses on one aspect of achieving flexible speech recognition, namely, improving the ability to cope with naturally spoken utterances through discriminative utterance verification. We propose an algorithm for training utterance verification systems based on the minimum verification error (MVE) training framework. Experimental results on speaker-independent connected digits, show a significant improvement in verification accuracy when the discriminant function used in MVE training is made consistent with the confidence measure used in utterance verification. At a 10% rejection rate, for example, the new proposed method reduces the string error rate by a further 22.7% over our previously reported results in which the MVE based discriminative training was not incorporated.
Mazin G. Rahim, Biing-Hwang Juang, Wu Chou
ICASSP3
1996 Key-phrase detection and verification for flexible speech understanding
Tatsuya Kawahara, Biing-Hwang Juang
ICSLP3
1996 Discriminative adaptation for speaker verification
Filipp Korkmazskiy, Biing-Hwang Juang
ICSLP2
1996 A study on task-independent subword selection and modeling for speech recognition
Biing-Hwang Juang, Wu Chou, J. J. Molina-Perez
ICSLP2
1996 Maximum likelihood learning of auditory feature maps for stationary vowels
Kuansan Wang, Biing-Hwang Juang
ICSLP3
1996 Time-frequency analysis and auditory modeling for automatic recognition of speech
abstract
Modern speech processing research may be categorized into three broad areas: statistical, physiological, and perceptual. Statistical research investigates the nature of the variability of the speech waveform from a signal processing viewpoint. This approach relates to the processing of speech in order to obtain measurements of speech characteristics which demonstrate manageable variabilities across a wide range of the talker population, in the presence of noise or competing speakers as well as the interaction of speech with the channel through which it is transmitted, and under the inherent interaction of the information content of speech itself (i.e., the contextual factor). Physiological research aims at constructing accurate models of the articulatory and auditory process, helping to limit the signal space for speech processing. In the perceptual realm, work focuses on understanding the psychoacoustic and possibly the psycholinguistic aspects of the speech communication process that the human so conveniently conducts. By studying this working analysis/recognition system, insights may be garnered that will lead to improved methods of speech processing. Conversely by studying the limitations of this system, particularly how it reduces the information rate of the received signal through, for example, masking and adaptation improvements may be made in the efficiency of speech coding schemes without impacting the quality of the reconstructed speech. Thus comprehension of speech production and perception impacts methods of speech processing, and vice-versa. This paper enunciates such a position, focusing on how modern time-frequency signal analysis methods could help expedite needed advances in these areas.
James W. Pitton, Kuansan Wang, Biing-Hwang Juang
Proc. IEEE3
1996 Signal conditioning techniques for robust speech recognition
abstract
Acoustic mismatch encountered in various training and testing conditions of hidden Markov model (HMM) based systems often causes severe degradation in speech recognition performance. For telephone based speech recognition tasks, acoustic mismatch can arise from various sources, such as variations in telephone handsets, ambient noise, and channel distortions. This paper presents three techniques for blind channel equalization, namely, cepstral mean subtraction (CMS), signal bias removal (SBR) and hierarchical signal bias removal (HSBR). Experimental results on various connected digits databases show a reduction in the digit error rate by 16%, 21%, and 28% when employing CMS, SBR, and HSBR, respectively. Our results also demonstrate that the HSBR technique outperforms SBR and CMS on every sub-data collection and exhibits consistent improvements even for short utterances.
Mazin G. Rahim, Biing-Hwang Juang, Wu Chou, Eric R. Buhrke
IEEE Signal Process. Lett.2
1996 Signal bias removal by maximum likelihood estimation for robust telephone speech recognition
abstract
An acoustical mismatch between the training and the testing conditions of hidden Markov model (HMM)-based speech recognition systems often causes a severte degradation in the recognition performance. In telephone speech recognition, for example, undesirable signal components due to ambient noise and channel distortion, as well as due to different variations of telephone handsets render the recognizer unusable for real- world applications. This paper presents a signal bias removal (SBR) method based on maximum likelihood1 estimation for the minimization of these undesirable effects. The proposed method is readily applicable in various architectures, i.e., dis- crete (vector-quantization based), semicontinuous and continuous density HMM. In this paper, the SBR method, integrated into a discrete density HMM, is applied to telephone speech recognition where the contamination due to extraneous signal components is assumed to be unknown. To enable real-time implementation, a sequential method for the estimation of the bias is presented. Experimental results for speaker-independent connected digit recognition show a reduction in the per digit error rate by up to 41% and 14% during mismatched and matclhed training and testing conditions, respectively.
Biing-Hwang Juang, Mazin G. Rahim
IEEE Trans. Speech Audio Process.1
1995 Robust utterance verification for connected digits recognition
abstract
Utterance verification represents an important technology in the design of user-friendly speech recognition systems. This paper addresses the issue of robustness in utterance verification. Four different approaches to robustness have been investigated: a string based likelihood measure for the detection of non-vocabulary words and "putative" errors, a signal bias removal method for channel normalization, on-line adaptation technique for achieving desirable trade-off between false rejection and false alarms, and a discriminative training method for the minimization of the expected string error rate. When these techniques were all integrated into a state-of-the-art connected digit recognition system, the string error rate was found to decrease by up to 57% at a rejection rate of 5%. For non-vocabulary word strings, the proposed utterance verification system rejected over 99.9% of extraneous speech.
Mazin G. Rahim, Biing-Hwang Juang
ICASSP3
1995 A training procedure for verifying string hypotheses in continuous speech recognition
abstract
A procedure is proposed for verifying the occurrence of string hypotheses produced by a hidden Markov model (HMM) based continuous speech recognizer. Most existing procedures verify word hypotheses through likelihood ratio scoring procedures computed using ad hoc approximations for the density of the alternative hypothesis in the denominator of the likelihood ratio statistic. The discriminative training procedure described in this paper attempts to adjust the parameters of the null hypothesis and the alternate hypothesis models to increase the power of a hypothesis test for utterance verification. The training procedure was evaluated for its ability to detect a twenty word vocabulary in a subset of the Switchboard conversational speech corpus. Experimental results show that the use of this procedure results in significant improvement in the word verification operating characteristic, as well as an improvement in the overall system performance.
Richard C. Rose, Biing-Hwang Juang
ICASSP2
1995 Filtering the time sequence of spectral parameters for speaker-independent CDHMM word recognition
abstract
In this work, we show how speaker-independent CDHMM word recognition performance can be significantly improved for clean speech by filtering the time sequence of spectral parameters to enhance its time dynamics. Experimental results with the standard TI connected digits database show the filter can achieve more than 30% reduction of string recognition error. As shown in this paper, that improvement is partially due to the speaker variability reduction obtained by attenuating the very low modulation frequencies. The widely used cepstral mean subtraction technique also improves the recognition rate, but it can not achieve such a noticeable improvement as the parameter filter. In fact, the best results are obtained when the peak of the long-term spectrum of the filter output is at around 3 Hz, a frequency which corresponds to the average syllable rate of the employed database.
Climent Nadeu, Pau Pachès-Leal, Biing-Hwang Juang
EUROSPEECH3
1995 Discriminative utterance verification for connected digits recognition
Mazin G. Rahim, Biing-Hwang Juang
EUROSPEECH3
1995 A vocabulary independent discriminatively trained method for rejection of non-keywords in sub word based speech recognition
Rafid A. Sukkar, Biing-Hwang Juang
EUROSPEECH3
1994 An algorithm of high resolution and efficient multiple string hypothesization for continuous speech recognition using inter-word models
abstract
We propose a new accurate string hypothesization algorithm to find the N-best multiple string hypotheses in continuous speech recognition. The algorithm differs from the conventional N-best search algorithms in that it allows the use of the same set of long term language model scores and the detailed context-dependent subword models such as inter-word context dependent triphone models in both forward and backward search for high performance speech recognition. It is an extension of the tree-trellis N-best search algorithm[1]. The inter-word context dependency is exactly preserved in both forward partial path map preparation and the proposed backward N-best multiple string hypothesis tree search. The search efficiency is maximized by applying the same high resolution acoustic and language models in both search directions. When search heuristics are used, the proposed approach provides a more accurate string model matching than that of the conventional frame-synchronous Viterbi beam search decoder.>
Wu Chou, Tatsuo Matsuoka, Biing-Hwang Juang
ICASSP (2)3
1994 Speaker recognition based on minimum error discriminative training
abstract
We study the use of discriminative training to construct speaker models for speaker verification and speaker identification. As opposed to conventional training which estimates a speaker's model based only on the training utterances from the same speaker, we use a discriminative training approach which takes into account the models of other competing speakers and formulates the optimization criterion such that speaker recognition error rate on the training data is directly minimized. We also propose a normalized score function which makes the verification formulation consistent with the minimum error training objective. We show that the speaker recognition performance is significantly improved when discriminative training is incorporated.>
Chi-Shi Liu, Biing-Hwang Juang, Aaron E. Rosenberg
ICASSP (1)3
1994 Signal bias removal for robust telephone based speech recognition in adverse environments
abstract
A speech signal transmitted through a telephone channel often encounters variable conditions which significantly deteriorate the performance of state-of-the-art HMM-based speech recognition systems. Undesirable components due to ambient noise and channel interference, as well as different sound pick-up equipments, render the recognizer unsuitable for real-world applications. This paper presents a signal bias removal method based on the maximum likelihood estimation for the minimization of these undesirable effects. The proposed method, integrated into a discrete density HMM, is applied to telephone speech recognition where the contamination due to ambient noise and channel distortion are assumed to be unknown. Experimental results are presented for speaker-independent connected digit recognition which demonstrate a reduction in the word error rate by up to 40%.>
Mazin G. Rahim, Biing-Hwang Juang
ICASSP (1)2
1994 Minimum error rate training of inter-word context dependent acoustic model units in speech recognition
Wu Chou, C.-E. Lee, Biing-Hwang Juang
ICSLP3
1994 Recent technology developments in connected digit speech recognition
Biing-Hwang Juang, Jay G. Wilpon
ICSLP1
1994 Filtering of spectral parameters for speech recognition
abstract
The time sequences of speech parameters resulting from current short-time spectral estimators show a tradeoff between estimation error variance and time and frequency resolution. In this paper, we apply frequency analysis and linear filtering to these sequences to gain insights into their limitations and to provide an interpretation framework for several parameter processing techniques proposed in the past. Particularly, the observation of their long-term spectrum reveals the importance of band equalization for improving discrimination in speech recognition. Based on that, we propose a method of filtering the sequences that includes an explicit equalization and incorporates a bandwidth parameter. By using Slepian sequences in the design of the filters, good results were obtained in our preliminary word recognition tests.
Climent Nadeu, Biing-Hwang Juang
ICSLP2
1994 A Minimum Error Rate Pattern Recognition Approach to Speech Recognition
abstract
In this paper, a minimum error rate pattern recognition approach to speech recognition is studied with particular emphasis on the speech recognizer designs based on hidden Markov models (HMMs) and Viterbi decoding. This approach differs from the traditional maximum likelihood based approach in that the objective of the recognition error rate minimization is established through a specially designed loss function, and is not based on the assumptions made about the speech generation process. Various theoretical and practical issues concerning this minimum error rate pattern recognition approach in speech recognition are investigated. The formulation and the algorithmic structures of several minimum error rate training algorithms for an HMM-based speech recognizer are discussed. The tree-trellis based N-best decoding method and a robust speech recognition scheme based on the combined string models are described. This approach can be applied to large vocabulary, continuous speech recognition tasks and to speech recognizers using word or subword based speech recognition units. Various experimental results have shown that significant error rate reduction can be achieved through the proposed approach.
Wu Chou, Biing-Hwang Juang, Frank K. Soong
Int. J. Pattern Recognit. Artif. Intell.3
1993 Minimum error rate training based on N-best string models
Wu Chou, Biing-Hwang Juang
ICASSP (2)3
1993 Discriminative analysis of distortion sequences in speech recognition
abstract
In a traditional speech recognition system, the distance score between a test token and a reference pattern is obtained by simply averaging the distortion sequence resulted from the matching of the two patterns through a dynamic programming procedure. The final decision is made by choosing the one with the minimal average distance score. If one views the distortion sequence as a form of observed features, a decision rule based on a specific discriminant function designed for the distortion sequence obviously will perform better than that based on the simple average distortion. The authors therefore, suggest a linear discriminant function of the form triangle = Sigma /sub (i1)/T w(i)* d(i) to compute the distance score triangle instead of a direct average triangle =1/T Sigma /sub (i1)/T d(i). Several adaptive algorithms are proposed to learn the discriminant weighting function. These include one heuristic method, two methods based on the error propagation algorithm, and one method based on the generalized probabilistic descent algorithm (GPD). They study these methods in a speaker-independent speech recognition task involving utterances of the highly confusible English E-set (b,c,d,e,g,p,t,v,z). The results show that the best performance is obtained by using the GPD-method which achieved a 78.1% accuracy, compared to 67.6% with the traditional unweighted average method. Besides the experimental comparisons, an analytical discussion of various training algorithms is also provided.>
Pao-Chung Chang, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.3
1993 Discriminative training of dynamic programming based speech recognizers
abstract
A new minimum recognition error formulation and a generalized probabilistic descent (GPD) algorithm are analyzed and used to accomplish discriminative training of a conventional dynamic-programming-based speech recognizer. The objective of discriminative training here is to directly minimize the recognition error rate. To achieve this, a formulation that allows controlled approximation of the exact error rate and renders optimization possible is used. The GPD method is implemented in a dynamic-time-warping (DTW)-based system. A linear discriminant function on the DTW distortion sequence is used to replace the conventional average DTW path distance. A series of speaker-independent recognition experiments using the highly confusible English E-set as the vocabulary showed a recognition rate of 84.4% compared to approximately 60% for traditional template training via clustering. The experimental results verified that the algorithm converges to a solution that achieves minimum error rate.>
Pao-Chung Chang, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.2
1993 Optimal quantization of LSP parameters
abstract
Two nonuniform aspects of the line spectrum pair (LSP) linear predictive coding (LPC) parameters are investigated, including nonuniform statistical distributions and spectral sensitivities of adjacent LSP frequency differences. Based upon these two nonuniform properties, a globally optimal scalar quantizer is designed for each differential LSP frequency. The design algorithm is dynamic programming based and minimization of a nontrivial data dependent spectral distortion is adopted as the optimality criterion. At 32 bits/frame, the new LSP quantizer achieves a 1-dB average log spectral distortion, a commonly accepted level for reproducing perceptually transparent spectral information. The quantization performance has also been shown to be robust across different speakers and databases.>
Frank K. Soong, Biing-Hwang Juang
IEEE Trans. Speech Audio Process.2
1992 Discriminative template training for dynamic programming speech recognition
abstract
A newly proposed minimum recognition error formulation and a generalized probabilistic descent (GPD) algorithm are analyzed and used to accomplish discriminative training of a conventional dynamic programming based speech recognizer. Unlike many other approaches, the objective of discriminative training the new framework is to directly minimize the recognition error rate. A series of speaker independent recognition experiments using the highly confusing English E-set as the vocabulary was conducted to examine the characteristics of the GPD method for discriminative training. Without ad hoc supplementary schemes, the method achieved a recognition rate of 83.7%, a remarkable performance improvement compared to 63.8% with the traditional template training via clustering. The experimental results verify that the GPD algorithm with the new minimum recognition error formulation indeed converges to a solution that accomplishes the objective of minimum error rate.>
Pao-Chung Chang, Biing-Hwang Juang
ICASSP2
1992 Segmental GPD training of HMM based speech recognizer
abstract
A novel training algorithm, segmental GPD (generalized probabilistic descent) training, for a hidden Markov model (HMM)-based speech recognizer using Viterbi decoding is proposed. This algorithm is based on the principle of minimum recognition error rate in which segmentation and discriminative training are jointly optimized. Various issues related to the special structure of HMM in segmental GPD training are studied. The authors tested this algorithm on two speaker-independent recognition tasks. The first experiment involves English E-set. Segmental GPD training was directly applied to HMM generated from nonoptimal uniform segmentation. A recognition rate of 88.7% was achieved on English E-set with whole word HMM. The second experiment involves the connected digits TI-database. Segmental GPD training was applied to HMM which were already trained using conventional training methods. A string recognition rate of 98.8% was achieved on 10-state word based HMM through segmental GPD training.>
Wu Chou, Biing-Hwang Juang
ICASSP2
1992 Vector equalization in hidden Markov models for noisy speech recognition
abstract
Speech recognizers often experience serious performance degradation when deployed in an unknown acoustic (particularly, noise contaminated) environment. To combat this problem, the authors proposed in a previous study a distortion measure that takes into account the norm shrinkage bias in the noisy cepstrum. The authors incorporate a first-order equalization mechanism, specifically aimed at avoiding the norm shrinkage problem, in a hidden Markov model (HMM) framework to model the speech cepstral sequence. Such a modeling technique requires special care as the formulation inevitably involves parameter estimation from a set of data with singular dispersion. The authors provide solutions to this HMM stochastic modeling problem and give algorithms for estimating the necessary model parameters. They experimentally show that incorporation of the first-order normal equalization model makes the HMM-based speech recognizer robust to noise. With respect to a conventional HMM recognizer, this leads to an improvement in recognition performance which is equivalent to about 15-20-dB gain in signal-to-noise ratio.>
Biing-Hwang Juang, Kuldip K. Paliwal
ICASSP1
1992 The use of cohort normalized scores for speaker verification
Aaron E. Rosenberg, Joel DeLong, Biing-Hwang Juang, Frank K. Soong
ICSLP4
1991 Discriminative analysis of distortion sequences in speech recognition
abstract
The authors suggest a linear discriminant function to complete the distance score instead of a conventional average distance. Several discriminative algorithms are proposed to learn the discriminant function. These include one heuristic method, two methods based on the error propagation algorithm, and one method based on the generalized probabilistic descent (GPD) algorithm. The authors study these methods in a speaker-independent speech recognition task involving utterances of the highly confusable English E-set. The results show that the best performance is obtained by using the GPD method, which achieved a 78.1% accuracy, compared to 67.6% with the traditional average method.>
Pao-Chung Chang, Sin-Horng Chen, Biing-Hwang Juang
ICASSP3
1990 Statistical segmentation and word modeling techniques in isolated word recognition
abstract
A speech recognition system is described using a combination of statistical segment and word modeling. Segment models are constructed by first segmenting training data automatically and then grouping the resultant segments into clusters. Mixtures of Gaussian densities are used to model each segment cluster. In order to integrate the segment models into word models, a generalization of the hidden Markov model approach is proposed. Experimental results on a multispeaker recognition system for alpha-digits demonstrate that the new approach improved the performance of conventional whole-word-based models. In particular, the word models show good discrimination abilities for differentiating phonetically similar words such as the E-set alphabet.>
S. A. Euler, Biing-Hwang Juang, Frank K. Soong
ICASSP2
1990 Speaker recognition based on source coding approaches
abstract
The use of nonmemoryless source coders in speaker recognition problems is studied, and the effects of source variations, including speaking inconsistency and channel mismatch, in source coder designs for the intended application are discussed. It is found that incorporation of memory in source coders in general enhances the speaker recognition accuracy but that more remarkable improvements can be accomplished by properly including potential source variations in the coder design/training. An experiment with a 100-speaker database shows a 99.5% recognition accuracy.>
Biing-Hwang Juang, Frank K. Soong
ICASSP1
1990 A study on speaker adaptation of continuous density HMM parameters
abstract
For a speech recognition system based on a continuous density hidden Markov model (CDHMM), it is shown that speaker adaptation of the parameters of the CDHMM can be formulated as a Bayesian learning procedure and it can be integrated into the segmental k-means training algorithm. Some results are reported for adapting both the mean and the diagonal covariance matrix of the Gaussian state observation densities of a CDHMM. When the speaker adaptation procedure is tested on a 39-word English alpha-digit vocabulary in isolated word mode, the results indicate that the procedure achieves better performance than a speaker-independent system, when only one training token from each word is used to perform speaker adaptation. It is also shown that much better performance can be achieved when two or more training tokens are used for speaker adaptation.>
Chih-Heng Lin, Biing-Hwang Juang
ICASSP3
1990 Optimal quantization of LSP parameters using delayed decisions
abstract
A previously published study by the authors (Proc. ICASSP, p.394-7, 1988) of optimal quantization of line spectral pair (LSP) parameters is extended by incorporating delayed decisions coding (in frequency). The A* algorithm is proposed for finding the best quantization bit pattern of LSP frequency differences. The best coding pattern is obtained efficiently without an exhaustive, hence prohibitive, search. The proposed search achieves a better rate-distortion performance than the best results obtained in the previous study. At 30 bits/frame, a net gain of 2 bits/frame over the previous results, the novel method achieves 1-dB average spectral distortion. Most importantly, the number of frames with large spectral distortions (>2 dB), which can be subjectively disturbing and degrade the perceived quality of a speech coder, is significantly reduced. The search complexity of the A* algorithm is moderate. While the peak load is comparable to a nonoptimal M-algorithm, the average load is about an order of magnitude lower.>
Frank K. Soong, Biing-Hwang Juang
ICASSP2
1990 Recent developments in speech recognition under adverse conditions
Biing-Hwang Juang
ICSLP1
1989 Speech enhancement based upon hidden Markov modeling
abstract
A maximum a posteriori approach for enhancing speech signals which have been degraded by statistically independent additive noise is proposed. The approach is based upon statistical modeling of the clean speech signal and the noise process using long training sequences from the two processes. Hidden Markov models (HMMs) with mixtures of Gaussian autoregressive (AR) output probability distributions are used to model the clean speech signal. A low-order Gaussian AR model is used for the wideband Gaussian noise considered here. The parameter set of the HMM is estimated using the Baum or the EM (estimation-maximization) algorithm. The enhancement of the noisy speech is done by means of reestimation of the clean speech waveform using the EM algorithm. An approximate improvement of 4.0-6.0 dB in signal-to-noise ratio (SNR) is achieved at 10 dB input SNR.>
Yariv Ephraim, David Malah, Biing-Hwang Juang
ICASSP3
1989 Word recognition using whole word and subword models
abstract
The problem of how to select and construct a set of fundamental unit statistical models suitable for speech recognition is addressed. A unified framework is discussed which can be used to accomplish the goal of creating effective basic models of speech. The performances of three types of fundamental units, namely whole word, phoneme-like, and acoustic segment units, in a 1109-word vocabulary speech recognition task are compared. The authors point out the relative advantages of each type of speech unit based on the results of a series of recognition experiments.>
Biing-Hwang Juang, Frank K. Soong, Lawrence R. Rabiner
ICASSP2
1989 HMM clustering for connected word recognition
abstract
The authors describe an HMM (hidden Markov model) clustering procedure and discuss its application to connected-word systems and to large-vocabulary recognition based on phonelike units. It is shown that the conventional approach of maximizing likelihood is easily implemented but does not work well in practice, as it tends to give improved models of tokens for which the initial model was generally quite good, but does not improve tokens which are poorly represented by the initial model. The authors have developed a splitting procedure which initializes each new cluster (statistical model) by splitting off all tokens in the training set which were poorly represented by the current set of models. This procedure is highly efficient and gives excellent recognition performance in connected-word tasks. In particular, for speaker-independent connected-digit recognition, using two HMM-clustered models, the recognition performance is as good as or better than previous results using 4-6 models/digit obtained from template-based clustering.>
Lawrence R. Rabiner, Biing-Hwang Juang, Jay G. Wilpon
ICASSP3
1988 On the application of hidden Markov models for enhancing noisy speech
abstract
An algorithm is proposed for enhancing noisy speech which has been degraded by statistically independent additive noise. The algorithm is based on modeling the clean speech as a hidden Markov process with mixtures of Gaussian autoregressive (AR) output processes and modeling the noise as a sequence of stationary, statistically independent, Gaussian AR vectors. The parameter sets of the models are estimated using training sequences from the clean speech and the noise process. The parameter set of the hidden Markov model is estimated by the segmental k-means algorithm. Given the estimated models, the enhancement of the noisy speech is done by alternate maximization of the likelihood function of the noisy speech, one over all sequences of states and mixture components assuming that the clean speech signal is given, and then over all vectors of the original speech using the resulting most probable sequence of states and mixture components. This alternating maximization is equivalent to first estimating the most probable sequence of AR models for the speech signal using the Viterbi algorithm, and then applying these AR models for constructing a sequence of Wiener filters which are used to enhance the noisy speech.>
Yariv Ephraim, David Malah, Biing-Hwang Juang
ICASSP3
1988 A segment model based approach to speech recognition
abstract
Proposes a global acoustic segment model for characterizing fundamental speech sound units and their interactions based upon a general framework of hidden Markov models (HMM). Each segment model represents a class of acoustically similar sounds. The intra-segment variability of each sound class is modeled by an HMM, and the sound-to-sound transition rules are characterized by a probabilistic intersegment transition matrix. An acoustically-derived lexicon is used to construct word models based upon subword segment models. The proposed segment model was tested on a speaker-trained, isolated word, speech recognition task with a vocabulary of 1109 basic English words. In the current study, only 128 segment models were used, and recognition was performed by optimally aligning the test utterance with all acoustic lexicon entries using a maximum likelihood Viterbi decoding algorithm. Based upon a database of three male speakers, the average word recognition accuracy for the top candidate was 85% and increased to 96% and 98% for the top 3 and top 5 candidates, respectively.>
Frank K. Soong, Biing-Hwang Juang
ICASSP3
1988 A family of distortion measures based upon projection operation for robust speech recognition
abstract
The authors aim at the formulation of similarity measures for robust speech recognition. Their consideration focuses on the speech cepstrum derived from linear prediction coefficients (the LPC cepstrum). By using common models for noisy speech, they analytically and empirically show how the ambient noise can affect some important attributes of the LPC cepstrum such as the vector norm, coefficient order, and the direction perturbation. The new findings led them to propose a family of distortion measures based on the projection between two cepstral vectors. Performance evaluation of these measures has been conducted in both speaker-dependent and speaker-independent isolated word recognition tasks. Experimental results show that the new measures cause no degradation in recognition accuracy at high SNR, but perform significantly better when tested under noisy conditions using only clean reference templates. At an SNR of 5 dB, the new measures are shown to be able to achieve a recognition rate equivalent to that obtained by the filtered cepstral measure at 20 dB SNR, demonstrating a gain of 15 dB.>
David Mansour, Biing-Hwang Juang
ICASSP2
1988 The short-time modified coherence representation and its application for noisy speech recognition
abstract
A technique for robust spectral representation of all-pole sequences is proposed. It is shown that the autocorrelation of an all-pole sequence, obtained by passing white noise through an all-pole filter 1/A(z), is an all-pole sequence of the form 1/A/sup 2/(z). The short-time modified coherence (SCM) representation, proposed here, is an all-pole modeling of the autocorrelation sequence followed by a spectral shaper. The spectral shaper, essentially a square root operator in the frequency domain, compensates for the inherent spectral distortion introduced by the autocorrelation operation on the autocorrelation sequence of the signal. The properties of the SMC representation, especially its robustness to additive white noise, are analyzed. Initial implementation of the SMC in a speaker-dependent isolated-word recognizer shows a considerable improvement over the standard linear predictive coding (LPC) representation. The SMC recognizer achieved an improvement in recognition accuracy equivalent to an increase in input SNR of approximately 13 dB, as compared to the LPC recognizer.>
David Mansour, Biing-Hwang Juang
ICASSP2
1987 Signal restoration by spectral mapping
abstract
Traditional approaches to the problem of noise suppression or signal restoration have been almost entirely based upon the methodology and theory of signal estimation. In this paper, we treat signal restoration as a problem in signal detection. Instead of estimating the characteristics of the signal and/or the noise, we establish a correspondence between the clean and the noisy signal through spectral mapping. In the procedure, we collect separate samples of both the clean signal and the noise. When the noise is additive, the (simulated) noisy signal is obtained by adding the noise to the clean signal. The sequence of short time spectra of the clean signal and that of the noisy signal form a one-to-one correspondence. The noisy spectral sequence is then used as a detection reference, to which the short time spectrum of an unknown noisy observation is compared, resulting in a detected occurrence of a particular group of spectra in the noisy sequence. Through the (inverse) mapping, the clean spectra that correspond to the detected noisy spectra are selected and processed to produce the restored spectrum. One important notion of the approach is that it is not limited to the usual least squares or minimum mean square framework. Our preliminary results show that when the mapping (detection) is based upon the likelihood ratio distortion measure, an SNR improvement of approximately 10 dB is obtainable for a 14 dB SNR noisy signal. Under the same condition, an improvement of approximately 8.5 dB can be obtained using a truncated cepstral distance measure.
Biing-Hwang Juang, Lawrence R. Rabiner
ICASSP1
1987 A performance evaluation of a connected digit recognizer
abstract
In this paper we discuss a system for automatically recognizing fluently spoken digit strings based on whole word reference units. The system that we will describe can use either hidden Markov model (HMM) technology or template-based technology. The training procedure derives the digit reference patterns (either templates or statistical models) from connected digit strings. To evaluate the performance of the overall connected digit recognizer, a set of 50 people (25 men, 25 women), from the non-technical local population, was each asked to record 1200 random connected digit strings over local dialed-up telephone lines. Both a speaker trained and a multispeaker training set was created, and a full performance evaluation was made. Results show that the average string accuracy for unknown and known length strings, in the speaker trained mode, was 98% and 99% respectively; in the multi-speaker mode the average string accuracies were 94% and 96.6% respectively.
Lawrence R. Rabiner, Jay G. Wilpon, Biing-Hwang Juang
ICASSP3
1987 An investigation on the use of acoustic sub-word units for automatic speech recognition
abstract
An approach to automatic speech recognition is described which attempts to link together ideas from pattern recognition such as dynamic time warping and hidden Markov modeling, with ideas from linguistically motivated approaches. In this approach, the basic sub-word units are defined acoustically, but not necessarily phonetically. An algorithm was developed which automatically decomposed speech into multiple sub-word segments, based solely upon strict acoustic criteria, without any reference to linguistic content. By repeating this procedure on a large corpus of speech data we obtained an extensive pool of unlabeled sub-word speech segments. Then using well defined clustering techniques, a small set of representative acoustic sub-word units (e.g. an inventory of units) was created. This process is fast, easy to use, and required no human intervention. The interpretation of these sub-word units, in a linguistic sense, in the context of word decoding is an important issue which must be addressed for them to be useful in a large vocabulary system. We have not yet addressed this issue; instead a couple of simple experiments were performed to determine if these acoustic sub-word units had any potential value for speech recognition. For these experiments we used a connected digits database from a single female talker. A 25 sub-word unit codebook of acoustic segments was created from about 1600 segments drawn from 100 connected digit strings. A simple isolated digit recognition system, designed using the statistics of the codewords in the acoustic sub-word unit codebook had a recognition accuracy of 100%. In another experiment a connected digit recognition system was created with representative digit templates created by concatenating the sub-word units in an appropriate manner. The system had a string recognition accuracy of 96%.
Jay G. Wilpon, Biing-Hwang Juang, Lawrence R. Rabiner
ICASSP2
1986 Design and performance of trellis vector quantizers for speech signals
abstract
In this paper, we present design algorithms for linear predictive vector quantizers with trellis structures that when combined with appropriate search procedures can achieve lower distortions than conventional memoryless vector quantizers. We consider two types of trellis structures: shift register and minimum degradation network. In the shift register case (SR-TVQ), the state of the trellis encoder corresponds to the contents of the shift register; it therefore does not require an explicit state transition matrix. The minimum degradation network is a state transition matrix obtained through a pruning procedure (P-TVQ) to give minimum distortion degradation from an omni search (full rate) vector quantizer. The role of search procedure, the delay requirements of each type of the encoders, the performance of the designs, and a comparison of the two methods are discussed.
Biing-Hwang Juang
ICASSP1
1986 Mixture autoregressive hidden Markov models for speaker independent isolated word recognition
abstract
In this paper a signal modeling technique based upon finite mixture autoregressive probabilistic functions of Markov chains is developed and applied to the problem of speech recognition, particularly speaker-independent recognition of isolated digits. Two types of mixture probability densities are investigated: finite mixtures of Gaussian autoregressive densities (GAM) and nearest-neighbor partitioned finite mixtures of Gaussian autoregressive densities (PGAM). In the former (GAM), the observation density in each Markov state is simply a (stochastically constrained) weighted sum of Gaussian autoregressive densities, while in the latter (PGAM) it involves nearest-neighbor decoding which, in effect, defines a set of partitions on the observation space. In this paper we discuss the signal modeling methodology and give experimental results on speaker independent recognition of isolated digits.
Biing-Hwang Juang, Lawrence R. Rabiner
ICASSP1
1986 On the use of bandpass liftering in speech recognition
abstract
In this paper, we extend the interpretation of distortion measures, based upon the observation that measurements of speech spectral envelopes (as normally obtained from analysis procedures) are prone to statistical variations due to window position fluctuations, excitation interference, measurement noise, etc. and may possess spurious characteristics because of analysis model constraints. We have found that these undesirable spectral measurement variations can be controlled (i.e. reduced in the level of variation) through proper cepstral processing and that a statistical model can be established to predict the variances of the cepstral coefficient measurements. The findings lead to the use of a bandpass "liftering" process aimed at reducing the variability of the statistical components of spectral measurements. We have applied this liftering process to various speech recognition problems; in particular, vowel recognition and isolated word recognition. With the liftering process, we have been able to achieve an average digit error rate of 1%, which is about half of the previously reported best results, with dynamic time warping in a speaker-independent isolated digit test.
Biing-Hwang Juang, Lawrence R. Rabiner, Jay G. Wilpon
ICASSP1
1986 A continuous training procedure for connected digit recognition
abstract
Algorithms for recognizing strings of connected words from whole word patterns (either templates or statistical models) have advanced to the point of high efficiency and accuracy. Although the computation rate of these connected word recognition algorithms remains high, advances in VLSI hardware make even the most ambitious connected word recognition tasks practical with todays technology. The greatest impediment to the successful utilization of connected word recognizers is the difficulty in extracting reliable, robust whole word reference patterns. In the past, connected word recognizers have relied on either isolated word reference patterns (which are trivially obtained), or reference patterns derived from limited context strings of words (e.g. the middle digit from strings of 3 digits). The resulting whole word reference patterns were adequate for slow rates of speech articulation, but proved inadequate when users spoke strings of words at high rates (e.g. on the order of 200-300 words per minute). To alleviate this difficulty, a training procedure for extracting whole word patterns from naturally spoken word strings has been implemented and is described here. The training procedure is essentially a k-means loop in which a set of known word strings is segmented into individual words based on matching an initial set of word reference patterns (typically a speaker independent set of isolated word reference patterns is used). The segmented words are then used to create an updated set of word reference patterns (either via clustering methods, for templates or via statistical techniques, for word models), which are then used in the segmental loop to give an updated set of word tokens from the labelled training set. This procedure is iterated until a stable set of whole word reference patterns is obtained. The training procedure was implemented and tested in a connected digits recognition task. For this task, string accuracies (on variable length strings with from 1-7 digits) on the order of 98-99% were obtained.
Lawrence R. Rabiner, Jay G. Wilpon, Biing-Hwang Juang
ICASSP3
1986 Maximum likelihood estimation for multivariate mixture observations of markov chains
abstract
To use probabilistic functions of a Markov chain to model certain parameterizations of the speech signal, we extend an estimation technique of Liporace to the eases of multivariate mixtures, such as Gaussian sums, and products of mixtures. We also show how these problems relate to Liporace's original framework.
Biing-Hwang Juang, Stephen E. Levinson, Man Mohan Sondhi
IEEE Trans. Inf. Theory1
1985 Recent developments in the application of hidden Markov models to speaker-independent isolated word recognition
abstract
In this paper we extend previous work on isolated word recognition based on hidden Markov models by replacing the discrete symbol representation of the speech signal by a continuous Gaussian mixture density. In this manner the inherent quantization error introduced by the discrete representation is essentially eliminated. The resulting recognizer was tested on a vocabulary of the 10 digits across a wide range of talkers and test conditions, and shown to have an error rate at least comparable to that of the best template recognizers and significantly lower than that of the discrete symbol hidden Markov model system. Several issues involved in the training of the continuous density models and in the implementation of the recognizer are discussed.
Biing-Hwang Juang, Lawrence R. Rabiner, Stephen E. Levinson, Man Mohan Sondhi
ICASSP1
1985 A vector quantization approach to speaker recognition
abstract
In this study a vector quantization (VQ) codebook was used as an efficient means of characterizing the short-time spectral features of a speaker. A set of such codebooks were then used to recognize the identity of an unknown speaker from his/her unlabelled spoken utterances based on a minimum distance (distortion) classification rule. A series of speaker recognition experiments was performed using a 100-talker (50 male and 50 female) telephone recording database consisting of isolated digit utterances. For ten random but different isolated digits, over 98% speaker identification accuracy was achieved. The effects, on performance, of different system parameters such as codebook sizes, the number of test digits, phonetic richness of the text, and difference in recording sessions were also studied in detail.
Frank K. Soong, Aaron E. Rosenberg, Lawrence R. Rabiner, Biing-Hwang Juang
ICASSP4
1984 Line spectrum pair (LSP) and speech data compression
abstract
Line Spectrum Pair (LSP) was first introduced by Itakura [1,2] as an alternative LPC spectral representations. It was found that this new representation has such interesting properties as (1) all zeros of LSP polynomials are on the unit circle, (2) the corresponding zeros of the symmetric and anti-symmetric LSP polynomials are interlaced, and (3) the reconstructed LPC all-pole filter preserves its minimum phase property if (1) and (2) are kept intact through a quantization procedure. In this paper we prove all these properties via a "phase function." The statistical characteristics of LSP frequencies are investigated by analyzing a speech data base. In addition, we derive an expression for spectral sensitivity with respect to single LSP frequency deviation such that some insight on their quantization effects can be obtained. Results on multi-pulse LPC using LSP for spectral information compression are finally presented.
Frank K. Soong, Biing-Hwang Juang
ICASSP2
1983 Speech enhancement with harmonic synthesis
abstract
Development and tests on an algorithm to enhance the intelligibility of speech degraded by an interfering talker is reported. This paper discusses the formulation of the problem, the techniques developed, and the results of a limited-scale intelligibility test. While the test results indicate that no intelligibility improvement is obtained from the processing, several promising new directions for this problem have been identified.
Brian A. Hanson, David Y. Wong, Biing-Hwang Juang
ICASSP3
1983 Very low data rate speech compression with LPC vector and matrix quantization
abstract
Frame predictive vector quantization is developed to compress the bit rate for coding the LPC filter coefficients to under 250 bits/sec. An innovative LPC compression technique, matrix quantization, is also developed to compress the LPC filter coefficients to a rate under 150 bits/sec. Subjective evaluation with the diagnostic rhyme test (DRT) finds the proposed techniques to be feasible for intelligible speech transmission at bit rates between 400 bits/sec and 200 bits/sec.
David Y. Wong, Biing-Hwang Juang, D. Y. Cheng
ICASSP2
1982 Multiple stage vector quantization for speech coding
abstract
In this paper, we present a multiple stage vector quantization technique which allows easy expansion of the original vector quantizer design to operate at higher bit rates for lower distortion. The computation and storage reduction is achieved by the fact that the overall requirements are the sum of the requirements of each stage instead of an exponentially increasing function of the bit rate as in the original one stage design. In the case of Euclidean distance measures such as the log area ratio measure, experimental results show that the quantizer performance is very close to a theoretically predicted asymptotically optimal rate distortion relationship.
Biing-Hwang Juang, Augustine H. Gray Jr.
ICASSP1
1982 Voice coding at 800 BPS and lower data rates with LPC vector quantization
abstract
Design of an 800 bps LPC vocoder based on vector quantization is presented. Subjective evaluation under different channel error and acoustic-ambient noise conditions are discussed. The results indicate that it preserves much of the intelligibility as well as robustness of LPC. Further reduction in bit rate is achieved by eliminating frame to frame redundancy in the vocoder parameters. Techniques include frame repeat coding and the newly developed matrix coding technique.
David Y. Wong, Biing-Hwang Juang
ICASSP2
1982 Vector quantization for linear prediction voice coding (Ph.D. Thesis abstr.)
Biing-Hwang Juang
IEEE Trans. Inf. Theory1
1981 Recent developments in vector quantization for speech processing
abstract
Vector Quantization is applied to modify a 2400 bps LPC vocoder to operate at 800 bps, while retaining acceptable intelligibility and naturalness of quality. The design of this speech compression system is discussed and compared to other very low bit rate vocoders. Advantages of vector quantization over a scalar technique are examined in detail, and several new properties are presented.
David Y. Wong, Biing-Hwang Juang, Augustine H. Gray Jr.
ICASSP2