Ramesh A. Gopinath

dblp:11/6392 · also Ramesh Gopinath · DBLP profile ↗
← Back
62ranked-venue papers
15as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 54 · 13 first-authorArtificial intelligence and machine learning · 24 · 1 first-authorSystems, architecture and hardware · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Speech recognition and synthesis · 70% Probabilistic and Bayesian machine learning · 19% Representation and self-supervised learning · 7%
Computer graphics and multimedia
1 paper
Image and video coding · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
acoustic modeling
0.242007
Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2007
Subspace constrained Gaussian mixture models for speech recognition · IEEE Trans. Speech Audio Process. 2005
Modeling inverse covariance matrices by basis expansion · IEEE Trans. Speech Audio Process. 2004
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.232007
Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2007
Subspace constrained Gaussian mixture models for speech recognition · IEEE Trans. Speech Audio Process. 2005
A robust high accuracy speech recognition system for mobile applications · IEEE Trans. Speech Audio Process. 2002
Natural language and speech › Speech recognition and synthesis › acoustic modeling
subspace constrained gaussian mixture models
0.122007
Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2007
Subspace constrained Gaussian mixture models for speech recognition · IEEE Trans. Speech Audio Process. 2005
Natural language and speech › Speech recognition and synthesis › acoustic modeling
discriminative acoustic model training
0.112007
Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2007
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture model
0.122004
Modeling inverse covariance matrices by basis expansion · IEEE Trans. Speech Audio Process. 2004
Gaussianization · NIPS 2000
Machine learning › Probabilistic and Bayesian machine learning
covariance modeling
0.012004
Modeling inverse covariance matrices by basis expansion · IEEE Trans. Speech Audio Process. 2004
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation
0.012000
Gaussianization · NIPS 2000
Machine learning › Generative modeling › normalizing flow
gaussianization
0.012000
Gaussianization · NIPS 2000
Machine learning › Representation and self-supervised learning › blind source separation
independent component analysis
0.012000
Gaussianization · NIPS 2000
Machine learning › Representation and self-supervised learning › blind source separation › independent component analysis
nonlinear ICA
0.012000
Gaussianization · NIPS 2000
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
hidden markov model
0.012005
Subspace constrained Gaussian mixture models for speech recognition · IEEE Trans. Speech Audio Process. 2005
Image and video coding › image compression
wavelet-based image coding
0.011995
On cosine-modulated wavelet orthonormal bases · IEEE Trans. Image Process. 1995
Information theory › signal processing › filtering
filter banks
0.011995
On cosine-modulated wavelet orthonormal bases · IEEE Trans. Image Process. 1995

Methods — techniques the papers use, named apart from their topics

maximum mutual information · 0.1error-weighted training · 0.1maximum likelihood estimation · 0.1expectation-maximization · 0.1generalized EM algorithm · 0.0basis expansion · 0.0multichannel CDCN · 0.0hidden markov model · 0.0discriminative training · 0.0bayesian information criterion · 0.0total variation smoothness · 0.0k-regularity design · 0.0cosine modulation · 0.0
YearPublicationVenuePosition
2007 Efficient, Low Latency Adaptation for Speech Recognition
abstract
Constrained or feature space maximum likelihood linear regression (FMLLR) is known to be an effective algorithm for adaptation to a new speaker or environment. It employs a single transformation matrix and bias vector to linearly transform the test speaker's features. FMLLR makes no assumption on the underlying noise, environment or speaker and estimates parameters to maximize likelihood of the test data. The standard implementation needs considerable computational power, requires significant amounts of storage, and requires a first pass decoding before adaptation can begin. In this paper, we propose a simplified implementation of FMLLR for embedded applications to address these problems. Here, we employ a simple speech/silence segmentation to estimate parameters. We operate in the 13 dimensional cepstral space, hence resource requirements are low. The algorithm does not require a first pass decoding (parameter estimation is accomplished entirely in the front end) and can be applied with low latency as compared to FMLLR. The algorithms we describe here provide an attractive tradeoff between the power of FMLLR and the computational simplicity of Cepstral Mean Subtraction. With minimal cost, we achieve nearly 15% relative gains on an embedded speech recognition task.
Suleyman Serdar Kozat, Karthik Visweswariah, Ramesh A. Gopinath
ICASSP (4)3
2007 Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech Recognition
abstract
In this paper, we study discriminative training of acoustic models for speech recognition under two criteria: maximum mutual information (MMI) and a novel "error-weighted" training technique. We present a proof that the standard MMI training technique is valid for a very general class of acoustic models with any kind of parameter tying. We report experimental results for subspace constrained Gaussian mixture models (SCGMMs), where the exponential model weights of all Gaussians are required to belong to a common "tied" subspace, as well as for subspace precision and mean (SPAM) models which impose separate subspace constraints on the precision matrices (i.e., inverse covariance matrices) and means. It has been shown previously that SCGMMs and SPAM models generalize and yield significant error rate improvements over previously considered model classes such as diagonal models, models with semitied covariances, and extended maximum likelihood linear transformation (EMLLT) models. We show here that MMI and error-weighted training each individually result in over 20% relative reduction in word error rate on a digit task over maximum-likelihood (ML) training. We also show that a gain of as much as 28% relative can be achieved by combining these two discriminative estimation techniques
Scott Axelrod, Vaibhava Goel, Ramesh A. Gopinath, Peder A. Olsen, Karthik Visweswariah
IEEE Trans. Speech Audio Process.3
2006 Feature Adaptation Based on Gaussian Posteriors
abstract
In this paper we consider the use of non-linear methods for feature adaptation to reduce the mismatch between test and training conditions. The non-linearity is introduced by using the posteriors of a set of Gaussians to (softly) partition the observation space for feature adaptation. The modeling framework used is based on the fMPE models (D. Povey et al., 2005) applied to FMLLR matrices directly. However, the parameters are estimated to maximize the likelihood of the test data. We observe a relative gain of 14% on top of FMLLR, which was a 42% relative gain over the baseline
Suleyman Serdar Kozat, Karthik Visweswariah, Ramesh A. Gopinath
ICASSP (1)3
2006 Dynamic Noise Adaptation
abstract
We consider the problem of robust speech recognition in the car environment. We present a new dynamic noise adaptation algorithm, called DNA, for the robust front-end compensation of evolving semi-stationary noise as typically encountered in the car setting. A large dataset of in-car noise was collected for the evaluation of the new algorithm. This dataset was combined with the Aurora II framework to produce a new, publicly available framework, called DNA + AURORA II, for the evaluation of adaptive noise compensation algorithms. We show that DNA consistently outperforms several existing, related state-of-the-art front-end denoising techniques
Steven J. Rennie, Trausti T. Kristjansson, Peder A. Olsen, Ramesh A. Gopinath
ICASSP (1)4
2006 On designing context sensitive language models for spoken dialog systems
Vaibhava Goel, Ramesh A. Gopinath
INTERSPEECH2
2006 Super-human multi-talker speech recognition: the IBM 2006 speech separation challenge system
abstract
We describe a system for model based speech separation which achieves super-human recognition performance when two talkers speak at similar levels. The system can separate the speech of two speakers from a single channel recording with remarkable results. It incorporates a novel method for performing two-talker speaker identification and gain estimation. We extend the method of model based high resolution signal reconstruction to incorporate temporal dynamics. We report on two methods for introducing dynamics; the first uses dynamics in the acoustic model space, the second incorporates dynamics based on sentence grammar. The addition of temporal constraints leads to dramatic improvements in the separation performance. Once the signals have been separated they are then recognized using speaker dependent labeling. 1.
Trausti T. Kristjansson, John R. Hershey, Peder A. Olsen, Steven J. Rennie, Ramesh A. Gopinath
INTERSPEECH5
2005 Initializing Subspace Constrained Gaussian Mixture Models
abstract
A recent series of papers [1, 2, 3, 4] introduced subspace constrained Gaussian mixture models (SCGMM) and showed that SCGMM can very efficiently approximate full covariance Gaussian mixture models (FCGMM); a significant reduction in the number of parameters is achieved with little loss in the accuracy of the model. SCGMM were arrived at as a sequence of generalizations of diagonal covariance GMM. As an artifact of this process the initialization of SCGMM parameters in that work is complex, i.e., relies on best parameter settings of less general models. This paper overcomes this problem by showing how an FCGMM can be used to give a simple and direct initialization of an SCGMM. The initialization scheme is powerful enough that as the number of parameters in an SCGMM approaches that of an FCGMM (i.e., large SCGMM) further training of the SCGMM is unnecessary.
Peder A. Olsen, Karthik Visweswariah, Ramesh A. Gopinath
ICASSP (1)3
2005 Subspace constrained Gaussian mixture models for speech recognition
abstract
A standard approach to automatic speech recognition uses hidden Markov models whose state dependent distributions are Gaussian mixture models. Each Gaussian can be viewed as an exponential model whose features are linear and quadratic monomials in the acoustic vector. We consider here models in which the weight vectors of these exponential models are constrained to lie in an affine subspace shared by all the Gaussians. This class of models includes Gaussian models with linear constraints placed on the precision (inverse covariance) matrices (such as diagonal covariance, maximum likelihood linear transformation, or extended maximum likelihood linear transformation), as well as the LDA/HLDA models used for feature selection which tie the part of the Gaussians in the directions not used for discrimination. In this paper, we present algorithms for training these models using a maximum likelihood criterion. We present experiments on both small vocabulary, resource constrained, grammar-based tasks, as well as large vocabulary, unconstrained resource tasks to explore the rather large parameter space of models that fit within our framework. In particular, we demonstrate significant improvements can be obtained in both word error rate and computational complexity.
Scott Axelrod, Vaibhava Goel, Ramesh A. Gopinath, Peder A. Olsen, Karthik Visweswariah
IEEE Trans. Speech Audio Process.3
2004 Adaptation of front end parameters in a speech recognizer
abstract
In this paper we consider the problem of adapting parameters of the algorithm used for extraction of features. Typical speech recognition systems use a sequence of modules to extract features which are then used for recognition. We present a method to adapt the parameters in these modules under a variety of criteria, e.g maximum likelihood, maximum mutual information. This method works under the assumption that the functions that the modules implement are differentiable with respect to their inputs and parameters. We use this framework to optimize a linear transform preceding the linear discriminant analysis (LDA) matrix and show that it gives significantly better performance than a linear transform after the LDA matrix with small amounts of data. We show that linear transforms can be estimated by directly optimizing likelihood or the MMI objective without using auxiliary functions. We also apply the method to optimize the Mel bins, and the compression power in a system that uses power law compression. 1.
Karthik Visweswariah, Ramesh A. Gopinath
INTERSPEECH2
2004 Task adaptation of acoustic and language models based on large quantities of data
abstract
We investigate use of large amounts, over 1500 hours, of untranscribed data recorded from a deployed conversational system to improve the acoustic and language models. The system that we considered allows users to perform transactions on their retirement accounts. Using all the untranscribed data we get over 19 % relative improvement in word error rate over a baseline system. In contrast, a system built using 70 hours of transcribed data results in over 31 % relative improvement. 1.
Karthik Visweswariah, Ramesh A. Gopinath, Vaibhava Goel
INTERSPEECH2
2004 Modeling inverse covariance matrices by basis expansion
abstract
This paper proposes a new covariance modeling technique for Gaussian mixture models. Specifically the inverse covariance (precision) matrix of each Gaussian is expanded in a rank-1 basis i.e., /spl Sigma//sub j//sup -1/=P/sub j/=/spl Sigma//sub k=1//sup D//spl lambda//sub k//sup j/a/sub k/a/sub k//sup T/, /spl lambda//sub k//sup j//spl isin//spl Ropf/,a/sub k//spl isin//spl Ropf//sup d/. A generalized EM algorithm is proposed to obtain maximum likelihood parameter estimates for the basis set {a/sub k/a/sub k//sup T/}/sub k=1//sup D/ and the expansion coefficients {/spl lambda//sub k//sup j/}. This model, called the extended maximum likelihood linear transform (EMLLT) model, is extremely flexible: by varying the number of basis elements from D=d to D=d(d+1)/2 one gradually moves from a maximum likelihood linear transform (MLLT) model to a full-covariance model. Experimental results on two speech recognition tasks show that the EMLLT model can give relative gains of up to 35% in the word error rate over a standard diagonal covariance model, 30% over a standard MLLT model.
Peder A. Olsen, Ramesh A. Gopinath
IEEE Trans. Speech Audio Process.2
2003 Dimensional reduction, covariance modeling, and computational complexity in ASR systems
abstract
We study acoustic modeling for speech recognition using mixtures of exponential models with linear and quadratic features tied across all context dependent states. These models are one version of the SPAM models introduced by Axelrod, Gopinath and Olsen (see Proc. ICSLP, 2002). They generalize diagonal covariance, MLLT, EMLLT, and full covariance models. Reduction of the dimension of the acoustic vectors using LDA/HDA projections corresponds to a special case of reducing the exponential model feature space. We see, in one speech recognition task, that SPAM models on an LDA projected space of varying dimensions achieve a significant fraction of the WER improvement in going from MLLT to full covariance modeling, while maintaining the low computational cost of the MLLT models. Further, the feature precomputation cost can be minimized using the hybrid feature technique of Visweswariah, Olsen, Gopinath and Axelrod (see ICASSP 2003); and the number of Gaussians one needs to compute can be greatly reducing using hierarchical clustering of the Gaussians (with fixed feature space). Finally, we show that reducing the quadratic and linear feature spaces separately produces models with better accuracy, but comparable computational complexity, to LDA/HDA based models.
Scott Axelrod, Ramesh A. Gopinath, Peder A. Olsen, Karthik Visweswariah
ICASSP (1)2
2003 Maximum likelihood training of subspaces for inverse covariance modeling
abstract
Speech recognition systems typically use mixtures of diagonal Gaussians to model the acoustics. Using Gaussians with a more general covariance structure can give improved performance; EM-LLT and SPAM models give improvements by restricting the inverse covariance to a linear/affine subspace spanned by rank one and full rank matrices respectively. We consider training these subspaces to maximize likelihood. For EMLLT ML training the subspace results in significant gains over the scheme proposed by Olsen and Gopinath (see Proceedings of ICASSP, 2002). For SPAM ML training of the subspace slightly improves performance over the method reported by Axelrod, Gopinath and Olsen (see Proceedings of ICSLP, 2002). For the same subspace size an EMLLT model is more efficient computationally than a SPAM model, while the SPAM model is more accurate. This paper proposes a hybrid method of structuring the inverse covariances that both has good accuracy and is computationally efficient.
Karthik Visweswariah, Peder A. Olsen, Ramesh A. Gopinath, Scott Axelrod
ICASSP (1)3
2003 Large vocabulary conversational speech recognition with a subspace constraint on inverse covariance matrices
abstract
This paper applies the recently proposed SPAM models for acoustic modeling in a Speaker Adaptive Training (SAT) context on large vocabulary conversational speech databases, including the Switchboard database. SPAM models are Gaussian mixture models in which a subspace constraint is placed on the precision and mean matrices (although this paper focuses on the case of unconstrained means). They include diagonal covariance, full covariance, MLLT, and EMLLT models as special cases. Adaptation is carried out with maximum likelihood estimation of the means and feature-space under the SPAM model. This paper shows the first experimental evidence that the SPAM models can achieve significant word-error-rate improvements over state-of-the-art diagonal covariance models, even when those diagonal models are given the benefit of choosing the optimal number of Gaussians (according to the Bayesian Information Criterion). This paper also is the first to apply SPAM models in a SAT context. All experiments are performed on the IBM "Superhuman" speech corpus, which is a challenging and diverse conversational speech test set that includes the Switchboard portion of the 1998 Hub5 evaluation data set.
Scott Axelrod, Vaibhava Goel, Brian Kingsbury, Karthik Visweswariah, Ramesh A. Gopinath
INTERSPEECH5
2003 Discriminative estimation of subspace precision and mean (SPAM) models
abstract
The SPAM model was recently proposed as a very general method for modeling Gaussians with constrained means and covariances. It has been shown to yield significant error rate improvements over other methods of constraining covariances such as diagonal, semi-tied covariances, and extended maximum likelihood linear transformations. In this paper we address the problem of discriminative estimation of SPAM model parameters, in an attempt to further improve its performance. We present discriminative estimation under two criteria: maximum mutual information (MMI) and an "error-weighted" training. We show that both these methods individually result in over 20% relative reduction in word error rate on a digit task over maximum likelihood (ML) estimated SPAM model parameters. We also show that a gain of as much as 28% relative can be achieved by combining these two discriminative estimation techniques. The techniques developed in this paper also apply directly to an extension of SPAM called subspace constrained exponential models.
Vaibhava Goel, Scott Axelrod, Ramesh A. Gopinath, Peder A. Olsen, Karthik Visweswariah
INTERSPEECH3
2003 Acoustic modeling with mixtures of subspace constrained exponential models
abstract
Gaussian distributions are usually parameterized with their natural parameters: the mean and the covariance #. They can also be re-parameterized as exponential models with canonical parameters P = # -1 and # = P. In this paper we consider modeling acoustics with mixtures of Gaussians parameterized with canonical parameters where the parameters are constrained to lie in a shared affine subspace. This class of models includes Gaussian models with various constraints on its parameters: diagonal covariances, MLLT models, and the recently proposed EMLLT and SPAM models. We describe how to perform maximum likelihood estimation of the subspace and parameters within a fixed subspace. In speech recognition experiments, we show that this model improves upon all of the above classes of models with roughly the same number of parameters and with little computational overhead. In particular we get 30-40% relative improvement over LDA+MLLT models when using roughly the same number of parameters.
Karthik Visweswariah, Scott Axelrod, Ramesh A. Gopinath
INTERSPEECH3
2003 Least squared error FIR filters with flat amplitude or group delay constraints
abstract
Finite impulse response lowpass filters with prescribed number of zeros at /spl omega/=/spl pi/ and either prescribed group delay flatness or prescribed amplitude flatness and linear phase are designed by minimizing a squared error. A simple orthogonal projection of least squared error filters (that are often available in closed form) onto a linear subspace determined via Thiran filters (for group delay flatness) and via Herrmann filters (for amplitude flatness) gives the desired filters.
Ramesh A. Gopinath
IEEE Signal Process. Lett.1
2002 Digit recognition in noisy environments via a sequential GMM/SVM system
abstract
This paper exploits the fact that when GMM and SVM classifiers with roughly the same level of performance exhibit uncorrelated errors they can be combined to produce a better classifier. The gain accrues from combining the descriptive strength of GMM models with the discriminative power of SVM classifiers. This idea, first exploited in the context of speaker recognition [1, 2], is applied to speech recognition - specifically to a digit recognition task in a noisy environment - with significant gains in performance.
Shai Fine, George Saon, Ramesh A. Gopinath
ICASSP3
2002 Rapid adaptation with linear combinations of rank-one matrices
abstract
Linear transforms are often used to adapt the acoustic models in speech recognition systems. When there is very little (5–10 sees.) acoustic data adaptation suffers from unreliable parameter estimation. Typically this problem is handled by imposing a diagonal or block diagonal structure on the transform. This paper proposes using transforms that are linear combinations of rank-one matrices. This approach is applied to the adaptation of the Gaussian means, Gaussian covariances and the acoustic features. Experimental results with varying amounts of adaptation data indicate that for the same number of parameters, our new parameterization performs significantly better than simpler transform parameterizations (diagonal and/or block-diagonal).
Vaibhava Goel, Karthik Visweswariah, Ramesh A. Gopinath
ICASSP3
2002 Adaptation experiments on the SPINE database with the Extended Maximum Likelihood Linear Transformation (EMLLT) model
abstract
This paper applies the recently proposed Extended Maximum Likelihood Linear Transformation (EMLLT) model for inverse covariances in a Speaker Adaptive Training (SAT) context. The paper adapts standard algorithms for maximum likelihood estimation of linear transforms for mean, variance and feature space adaptation respectively, to the EMLLT model. Experimental results showing word-error-rate improvements are reported on the SPINE2 database. The system described here is the best-performing system submitted by IBM in the SPINE2 evaluation conducted by NIST in October 2001.
Ramesh A. Gopinath, Vaibhava Goel, Karthik Visweswariah, Peder A. Olsen
ICASSP1
2002 Modeling inverse covariance matrices by basis expansion
abstract
This paper proposes a new covariance modeling technique for Gaussian Mixture Models. Specifically the inverse covariance (precision) matrix of each Gaussian is expanded in a rank-1 basis i.e., Σj−1= Pj= Σk = 1DλkjakakT, λkj∈ ℝd. A generalized EM algorithm is proposed to obtain maximum likelihood parameter estimates for the basis set {akakT} and the expansion coefficients {λkj}. This model, called the Extended Maximum Likelihood Linear Transform (EMLLT) model, is extremely flexible: by varying the number of basis elements from d to d(d + 1)/2 one gradually moves from a Maximum Likelihood Linear Transform (MLLT) model to a full-covariance model. Experimental results on two speech recognition tasks show that the EMLLT model can give relative gains of up to 35% in the word error rate over a standard diagonal covariance model.
Peder A. Olsen, Ramesh A. Gopinath
ICASSP2
2002 Structuring linear transforms for adaptation using training time information
abstract
Linear transforms are often used for adaptation to test data in speech recognition systems. However, when used with small amounts of test data, these techniques provide limited improvements if any. This paper proposes a two-step Bayesian approach where a) the transforms lie in a subspace obtained at training time and b) the expansion coefficients of the transform are obtained using MAP. Estimation algorithms are given for adaptation transforms for means, covariances, and feature spaces. Experimental results indicate that our method gives a significant improvement in performance over other methods.
Karthik Visweswariah, Vaibhava Goel, Ramesh A. Gopinath
ICASSP3
2002 Short-time Gaussianization for robust speaker verification
abstract
In this paper, a novel approach for robust speaker verification, namely short-time Gaussianization, is proposed. Short-time Gaussianization is initiated by a global linear transformation of the features, followed by a short-time windowed cumulative distribution function (CDF) matching. First, the linear transformation in the feature space leads to local independence or decorrelation. Then the CDF matching is applied to segments of speech localized in time and tries to warp a given feature so that its CDF matches normal distribution. It is shown that one of the recent techniques used for speaker recognition, feature warping [l] can be formulated within the framework of Gaussianization. Compared to the baseline system with cepstral mean subtraction (CMS), around 20% relative improvement in both equal error rate(EER) and minimum detection cost function (DCF) is obtained on NIST 2001 cellular phone data evaluation.
Bing Xiang, Upendra V. Chaudhari, Jirí Navrátil 0001, Ganesh N. Ramaswamy, Ramesh A. Gopinath
ICASSP5
2002 Modeling with a subspace constraint on inverse covariance matrices
abstract
We consider a family of Gaussian mixture models for use in HMM based speech recognition system. These "SPAM" models have state independent choices of subspaces to which the precision (inverse covariance) matrices and means are restricted to belong. They provide a flexible tool for robust, compact, and fast acoustic modeling. The focus of this paper is on the case where the means are unconstrained. The models in the case already generalize the recently introduced EMLLT models, which themselves interpolate between MLLT and full covariance models. We describe an algorithm to train both the state-dependent and state-independent parameters. Results are reported on one speech recognition task. The SPAM models are seen to yield significant improvements in accuracy over EMLLT models with comparable model size and runtime speed. We find a 10% relative reduction in error rate over an MLLT model can be obtained while decreasing the acoustic modeling time by 20%.
Scott Axelrod, Ramesh A. Gopinath, Peder A. Olsen
INTERSPEECH2
2002 Large vocabulary conversational speech recognition with the extended maximum likelihood linear transformation (EMLLT) model
abstract
This paper applies the recently proposed Extended Maximum Likelihood Linear Transformation (EMLLT) model in a Speaker Adaptive Training (SAT) context on the Switchboard database. Adaptation is carried out with maximum likelihood estimation of linear transforms for the means, precisions (inverse covariances) and the feature-space under the EMLLT model. This paper shows the first experimental evidence that significant word-error-rate improvements can be achieved with the EMLLT model (in both VTL and VTL+SAT training contexts) over a state-of-the-art diagonal covariance model in a difficult large-vocabulary conversational speech recognition task. The improvements were of the order of 1% absolute in multiple scenarios.
Jing Huang 0019, Vaibhava Goel, Ramesh A. Gopinath, Brian Kingsbury, Peder A. Olsen, Karthik Visweswariah
INTERSPEECH3
2002 An EM algorithm for convolutive independent component analysis
Sabine Deligne, Ramesh A. Gopinath
Neurocomputing2
2002 Automatic transcription of Broadcast News
Scott Saobing Chen, Ellen Eide, Mark J. F. Gales, Ramesh A. Gopinath, D. Kanvesky, Peder A. Olsen
Speech Commun.4
2002 A robust high accuracy speech recognition system for mobile applications
abstract
This paper describes a robust, accurate, efficient, low-resource, medium-vocabulary, grammar-based speech recognition system using hidden Markov models for mobile applications. Among the issues and techniques we explore are improving robustness and efficiency of the front-end, using multiple microphones for removing extraneous signals from speech via a new multichannel CDCN technique, reducing computation via silence detection, applying the Bayesian information criterion (BIC) to build smaller and better acoustic models, minimizing finite state grammars, using hybrid maximum likelihood and discriminative models, and automatically generating baseforms from single new-word utterances.
Sabine Deligne, Satya Dharanipragada, Ramesh A. Gopinath, Benoît Maison, Peder A. Olsen, Harry Printz
IEEE Trans. Speech Audio Process.3
2001 The IBM Personal Speech Assistant
abstract
We describe technology and experience with an experimental personal information manager, which interacts with the user primarily but not exclusively through speech recognition and synthesis. This device, which controls a client PDA, is known as the personal speech assistant (PSA). The PSA contains complete speech recognition, speech synthesis and dialog management systems. Packaged in a hand-sized enclosure, of size and physical design to mate with the popular Palm III personal digital assistant, the PSA includes its own battery, microphone, speaker, audio input and output amplifiers, processor and memory. The PSA supports speaker-independent English speech recognition using a 500-word vocabulary, and English speech synthesis on an arbitrary vocabulary. We survey the technical issues we encountered in building the hardware and software for this device, and the solutions we implemented, including audio system design, power and space budget, speech recognition in adverse acoustic environments with constrained processing resources, dialog management, appealing applications, and overall system architecture.
Liam Comerford, David Frank, Ponani S. Gopalakrishnan, Ramesh A. Gopinath, Jan Sedivý
ICASSP4
2001 Automatic generation and selection of multiple pronunciations for dynamic vocabularies
abstract
We present a scheme for the acoustic modeling of speech recognition applications requiring dynamic vocabularies. It applies especially to the acoustic modeling of out-of-vocabulary words which need to be added to a recognition lexicon based on the observation of a few (say one or two) speech utterances of these words. Standard approaches to this problem derive a single pronunciation from each speech utterance by combining acoustic and phone transition scores. In our scheme, multiple pronunciations are generated from each speech utterance of a word to enroll by varying the relative weights assigned to the acoustic and phone transition models. In our experiments, the use of these multiple baseforms dramatically outperforms the standard approach with a relative decrease of the word error rate ranging from 20% to 40% on all our test sets.
Sabine Deligne, Benoît Maison, Ramesh A. Gopinath
ICASSP3
2001 A hybrid GMM/SVM approach to speaker identification
abstract
Proposes a classification scheme that incorporates statistical models and support vector machines. A hybrid system which appropriately combines the advantages of both the generative and discriminant model paradigms is described and experimentally evaluated on a text-independent speaker recognition task in matched and mismatched training and test conditions. Our results prove that the combination is beneficial in terms of performance and practical in terms of computation. We report relative improvements of up to 25% reduction in identification error rate compared to the baseline statistical model.
Shai Fine, Jirí Navrátil 0001, Ramesh A. Gopinath
ICASSP3
2001 Multiple linear transforms
abstract
Heteroscedastic discriminant analysis (HDA) has been proposed as a replacement for linear discriminant analysis (LDA) in speech recognition systems that use mixtures of diagonal covariance Gaussians to model the data. Typically HDA and LDA involve a dimension reduction of the feature space. A specific version HDA that involves no dimension reduction; and is popularly known as maximum likelihood linear transform (MLLT) is often used on the feature space to give significant improvements in performance. MLLT approximately diagonalizes the class covariances, and in effect, tries to approximate the performance of a full-covariance-system. However, the performance of a full-covariance system could in some cases be much better than using MLLT-based diagonal covariance system. We propose the method of multiple linear transforms, that bridges this gap in performance, while maintaining the speed efficiency of a diagonal covariance system. This technique improves the performance of a diagonal covariance system, over what could be obtained from HDA or MLLT.
Nagendra K. Goel, Ramesh A. Gopinath
ICASSP2
2001 Robust confidence annotation and rejection for continuous speech recognition
abstract
We are looking for confidence scoring techniques that perform well on a broad variety of tasks. Our main focus is on word-level error rejection, but most results apply to other scenarios as well. A variation of the normalized cross entropy that is adapted to that purpose is introduced. It is successfully used to automatically select features and optimize the word-level confidence measure on several test sets. Sentence-level confidence geared toward the rejection of out-of-grammar utterances is also investigated. The combination of a word graph based technique and the acoustic score shows excellent performance across all the tasks we considered.
Benoît Maison, Ramesh A. Gopinath
ICASSP2
2001 Low-resource hidden Markov model speech recognition
Sabine Deligne, Ellen Eide, Ramesh A. Gopinath, Dimitri Kanevsky, Benoît Maison, Peder A. Olsen, Harry Printz, Jan Sedivý
INTERSPEECH3
2001 Enhancing GMM scores using SVM "hints"
abstract
This paper proposes a classification scheme that combines statistical models and support vector machines. It exploits the fact (observed in [1]) that GMM and SVM classifiers with roughly the same level of performance produce uncorrelated errors. We describe a novel scheme which employs an SVM classifier as an “advisor ” to the GMM classifier in uncertain cases. The utility of the combined generative/discriminative approach is demonstrated on standard text-independent speaker verification and speaker identification tasks in matched and mismatched training and test conditions. Results indicate significant improvements in performance without much computational overhead. 1.
Shai Fine, Jirí Navrátil 0001, Ramesh A. Gopinath
INTERSPEECH3
2000 Maximum likelihood discriminant feature spaces
abstract
Linear discriminant analysis (LDA) is known to be inappropriate for the case of classes with unequal sample covariances. There has been an interest in generalizing LDA to heteroscedastic discriminant analysis (HDA) by removing the equal within-class covariance constraint. This paper presents a new approach to HDA by defining an objective function which maximizes the class discrimination in the projected subspace while ignoring the rejected dimensions. Moreover, we investigate the link between discrimination and the likelihood of the projected samples and show that HDA can be viewed as a constrained ML projection for a full covariance Gaussian model, the constraint being given by the maximization of the projected between-class scatter volume. It is shown that, under diagonal covariance Gaussian modeling constraints, applying a diagonalizing linear transformation (MLLT) to the HDA space results in increased classification accuracy even though HDA alone actually degrades the recognition performance. Experiments performed on the Switchboard and Voicemail databases show a 10%-13% relative improvement in the word error rate over standard cepstral processing.
George Saon, Mukund Padmanabhan, Ramesh A. Gopinath, Scott Saobing Chen
ICASSP3
2000 Transformation enhanced multi-grained modeling for text-independent speaker recognition
abstract
We describe our formulation of transformation enhanced data modeling used to develop a multi-grained data analysis approach to text independent speaker recognition. The broad goal is to address difficulties caused by sparse training and test data. First, our development of maximum likelihood transformation based recognition with diagonally constrained Gaussian mixture models is detailed. We give results to show its robustness to decreasing training data. Then using the these models as building blocks, a multigrained model structure is developed. For this, the training data must be labeled, e.g. with an HMM based phone labeler. A graduated phone class structure is then used to train the speaker model at various levels of detail. This structure is a tree with the root node containing all the phones. Subsequent levels partition the phones into increasingly finer grained linguistic classes. We demonstrate the effectiveness of the modeling with identification and verification experiments. 1.
Upendra V. Chaudhari, Jirí Navrátil 0001, Stéphane H. Maes, Ramesh A. Gopinath
INTERSPEECH4
2000 Transcription of broadcast news with a time constraint: IBM's 10xRT HUB4 system
abstract
We describe a system which automatically transcribes broadcast news in less than 10 times real-time. We detail the system architecture of this system, which was used by IBM in the 1999 HUB4 10xRT evaluation, and show that the performance of this system is over 20 percent more accurate at the same speed than the system we used in the 1998 evaluation. Furthermore, we have closed the gap in word recognition accuracy between an unlimited resource system and this which runs in under 10 times real time from 45 percent to 14 percent.
Ellen Eide, Benoît Maison, Dimitri Kanevsky, Peder A. Olsen, Scott Saobing Chen, Lidia Mangu, Mark J. F. Gales, Miroslav Novak, Ramesh A. Gopinath
INTERSPEECH9
2000 Gaussianization
abstract
High dimensional data modeling is difficult mainly because the so-called "curse of dimensionality". We propose a technique called "Gaussianiza(cid:173) tion" for high dimensional density estimation, which alleviates the curse of dimensionality by exploiting the independence structures in the data. Gaussianization is motivated from recent developments in the statistics literature: projection pursuit, independent component analysis and Gaus(cid:173) sian mixture models with semi-tied covariances. We propose an iter(cid:173) ative Gaussianization procedure which converges weakly: at each it(cid:173) eration, the data is first transformed to the least dependent coordinates and then each coordinate is marginally Gaussianized by univariate tech(cid:173) niques. Gaussianization offers density estimation sharper than traditional kernel methods and radial basis function methods. Gaussianization can be viewed as efficient solution of nonlinear independent component anal(cid:173) ysis and high dimensional projection pursuit.
Scott Saobing Chen, Ramesh A. Gopinath
NIPS2
1999 Recent improvements to IBM's speech recognition system for automatic transcription of broadcast news
abstract
We describe extensions and improvements to IBM's system for automatic transcription of broadcast news. The speech recognizer uses a total of 160 hours of acoustic training data, 80 hours more than for the system described in Chen et al. (1998). In addition to improvements obtained in 1997 we made a number of changes and algorithmic enhancements. Among these were changing the acoustic vocabulary, reducing the number of phonemes, insertion of short pauses, mixture models consisting of non-Gaussian components, pronunciation networks, factor analysis (FACILT) and Bayesian information criteria (BIC) applied to choosing the number of components in a Gaussian mixture model. The models were combined in a single system using NIST's script voting machine known as rover (Fiscus 1997).
Scott Saobing Chen, Ellen Eide, Mark J. F. Gales, Ramesh A. Gopinath, Dimitri Kanevsky, Peder A. Olsen
ICASSP4
1999 Model selection in acoustic modeling
abstract
Recently several classes of models have been suggested for use in continuous density HMMs for speech recognition. This paper proposes to choose both the model type and model size (number of parameters) by optimizing the Bayesian information criterion. Specically we apply this to Gaussian mixture density estimation to determine both the number of Gaussians and the covariance structure of each Gaussian, and decision tree clustering of HMM states. A numerical algorithm similar to the EM algorithm for mixture density estimation is proposed for optimizing BIC. 1
Scott Saobing Chen, Ramesh A. Gopinath
EUROSPEECH2
1999 Improved speaker segmentation and segments clustering using the bayesian information criterion
Alain Tritschler, Ramesh A. Gopinath
EUROSPEECH2
1998 Maximum likelihood modeling with Gaussian distributions for classification
abstract
Maximum likelihood (ML) modeling of multiclass data for classification often suffers from the following problems: (a) data insufficiency implying overtrained or unreliable models, (b) large storage requirement, (c) large computational requirement and/or (d) the ML is not discriminating between classes. Sharing parameters across classes (or constraining the parameters) clearly tends to alleviate the first three problems. We show that in some cases it can also lead to better discrimination (as evidenced by reduced misclassification error). The parameters considered are the means and variances of the Gaussians and linear transformations of the feature space (or equivalently the Gaussian means). Some constraints on the parameters are shown to lead to linear discrimination analysis (a well-known result) while others are shown to lead to optimal feature spaces (a relatively new result). Applications of some of these ideas to the speech recognition problem are also given.
Ramesh A. Gopinath
ICASSP1
1998 Transcription of broadcast news-some recent improvements to IBM's LVCSR system
abstract
This paper describes extensions and improvements to IBM's large vocabulary continuous speech recognition (LVCSR) system for transcription of broadcast news. The recognizer uses an additional 35 hours of training data over the one used in the 1996 Hub4 evaluation. It includes a number of new features: optimal feature space for acoustic modeling (in training and/or testing), filler-word modeling, Bayesian information criterion (BIC) based segment clustering, an improved implementation of iterative MLLR and 4-gram language models. Results using the 1996 DARPA Hub4 evaluation data set are presented.
Lazaros Polymenakos, Peder A. Olsen, D. Kanvesky, Ramesh A. Gopinath, Ponani S. Gopalakrishnan, Scott Saobing Chen
ICASSP4
1998 Techniques for capturing temporal variations in speech signals with fixed-rate processing
Satya Dharanipragada, Ramesh A. Gopinath, Bhaskar D. Rao
ICSLP2
1998 Factor analysis invariant to linear transformations of data
abstract
Modeling data with Gaussian distributions is an important statistical problem. To obtain robust models one imposes constraints the means and covariances of these distributions [6, 4, 10, 8]. Constrained ML modeling implies the existence of optimal feature spaces where the constraints are more valid [2, 3]. This paper introduces one such constrained ML modeling technique called factor analysis invariant to linear transformations (FACILT) which is essentially factor analysis in optimal feature spaces. FACILT is a generalization of several existing methods for modeling covariances. This paper presents an EM algorithm for FACILT modeling. 1.
Ramesh A. Gopinath, Bhuvana Ramabhadran, Satya Dharanipragada
ICSLP1
1997 Transcription of broadcast news-system robustness issues and adaptation techniques
abstract
This paper describes some of the main problems and issues specific to the transcription of broadcast news and describes some of the methods for solving them that have been incorporated into the IBM Large Vocabulary Continuous Speech Recognition System.
Raimo Bakis, Scott Saobing Chen, Ponani S. Gopalakrishnan, Ramesh A. Gopinath, Stéphane H. Maes, Lazaros Polymenakos
ICASSP4
1997 New methods in continuous Mandarin speech recognition
C. Julian Chen, Ramesh A. Gopinath, Michael D. Monkowski, Michael Picheny, Katherine Shen
EUROSPEECH2
1996 Modulated filter banks and wavelets-a general unified theory
abstract
This paper generalizes and unifies well-known results on modulated filter banks (MFBs) and modulated wavelet tight frames (MWTFs). It classifies MFBs based on the discrete cosine or sine transforms that they are associated with. By proper choice of the form of modulation the perfect reconstruction (PR) conditions are seen to be (surprisingly) identical for all classes of MFBs. This has the interesting consequence that optimal MFB prototype designs can be shared across MFB classes. For some classes of MFBs associated MWTFs do not exist, while for others they do. The results cover both orthogonal and biorthogonal MFBs; and the filters could be arbitrary sequences in l/sup 2/(Z).
Ramesh A. Gopinath
ICASSP1
1995 Nonlinear recovery of sparse signals from narrowband data
abstract
This paper describes the connection between a certain signal recovery problem and the decoding of Reed-Solomon codes. It is shown that any algorithm for decoding Reed-Solomon codes (over finite fields) can be used to recover wide-band signals (over the real/complex field) from narrow-band information. It also shows that a signal with at most N/sub t/ frequency samples can be recovered from any contiguous band of 2N/sub t/ frequency samples.
Ramesh A. Gopinath
ICASSP1
1995 On cosine-modulated wavelet orthonormal bases
abstract
Multiplicity M, K-regular, orthonormal wavelet bases (that have implications in transform coding applications) have previously been constructed by several authors. The paper describes and parameterizes the cosine-modulated class of multiplicity M wavelet tight frames (WTFs). In these WTFs, the scaling function uniquely determines the wavelets. This is in contrast to the general multiplicity M case, where one has to, for any given application, design the scaling function and the wavelets. Several design techniques for the design of K regular cosine-modulated WTFs are described and their relative merits discussed. Wavelets in K-regular WTFs may or may not be smooth, Since coding applications use WTFs with short length scaling and wavelet vectors (since long filters produce ringing artifacts, which is undesirable in, say, image coding), many smooth designs of K regular WTFs of short lengths are presented. In some cases, analytical formulas for the scaling and wavelet vectors are also given. In many applications, smoothness of the wavelets is more important than K regularity. The authors define smoothness of filter banks and WTFs using the concept of total variation and give several useful designs based on this smoothness criterion. Optimal design of cosine-modulated WTFs for signal representation is also described. All WTFs constructed in the paper are orthonormal bases.
Ramesh A. Gopinath, C. Sidney Burrus
IEEE Trans. Image Process.1
1994 Factorization approach to time-varying filter banks and wavelets
abstract
A complete factorization of all optimal (in terms of quick transition) time-varying FIR unitary filter bank tree topologies is obtained. This has applications in adaptive subband coding, tiling of the time-frequency plane and the construction of orthonormal wavelet and wavelet packet bases for the half-line and interval [8, ?, 11, 2]. For an M-channel filter bank the factorization allows one to construct entry/exit filters that allow the filter bank to be used on finite signals without distortion at the boundaries. One of the advantages of our approach is that an efficient implementation algorithm comes with the factorization. The factorization can be used to generate filter bank tree-structures where the tree topology changes over time. Explicit formulas for the transition filters are obtained for arbitrary tree transitions. The results hold for tree structures where filter banks with any number of channels or filters of any length are used. Time-varying wavelet and wavelet packet bases are also constructed using these filter bank structures.>
Ramesh A. Gopinath
ICASSP (3)1
1994 Wavelet-Based Post-Processing of Low Bit Rate Transform Coded Images
abstract
We propose a novel method based on wavelet thresholding for enhancement of decompressed transform coded images. Transform coding at low bit rates typically introduces artifacts associated with the basis functions of the transform. In particular, the method works remarkably well in "deblocking" of DCT compressed images. The method is nonlinear, computationally efficient, and spatially adaptive and has the distinct feature that it removes artifacts yet retain sharp features in the images. An important implication of this result is that images coded using the JPEG standard can efficiently be postprocessed to give significantly improved visual quality in the images. The algorithm can use a conventional JPEG encoder and decoder for which VLSI chips are available.>
Ramesh A. Gopinath, Markus Lang, Haitao Guo, Jan E. Odegard
ICIP (2)1
1994 Wavelet Based Speckle Reduction with Application to SAR Based ATD/R
abstract
The paper introduces a novel speckle reduction method based on thresholding the wavelet coefficients of the logarithmically transformed image. The method is computational efficient and can significantly reduce the speckle while preserving the resolution of the original image. Both soft and hard thresholding schemes are studied and the results are compared. When fully polarimetric SAR images are available, the authors propose several approaches to combine the data from different polarizations to achieve even better performance. Wavelet processed imagery is shown to provide better detection performance for the synthetic-aperture radar (SAR) based automatic target detection/recognition (ATD/R) problem.>
Haitao Guo, Jan E. Odegard, Markus Lang, Ramesh A. Gopinath, Ivan W. Selesnick, C. Sidney Burrus
ICIP (1)4
1994 Some Thoughts on Least Squared Error Optimal Windows
abstract
Windowing methods give simple designs of FIR filters. Window designs are generally not considered optimal in any meaningful sense. A typical desired FIR lowpass filter is specified by giving the passband, stopband and transitionband responses of the filter. A particular class of transition responses viz., spline transition functions, give rise to windows that are optimal in a least squared sense. Do all transition functions have associated windows that are least squared optimal? More importantly, given a window does there exist a (meaningful) transition function with respect to which the window is least squared optimal? In trying to answer the second question this note also characterizes all possible lowpass extensions of a given sequence and exhibits the unique minimum norm extension. This is used to show that a given finitely supported window w/sub b/(n), and a desired response with prescribed transition band edges (but no transition function) there exist infinitely many transition functions with respect to which the windowed FIR filter (obtained by windowing an ideal frequency response) is least squared optimal. The unique minimum norm lowpass extension of w/sub b/(n) is used to exhibit a particular transition function with an extremal property. Since it is desirable to have a monotone transition function we ask the question: can a given window be least squared optimal with respect to a monotone desired transition response? This question leads to a few open problems on non-negative definite lowpass extensions of sequences.>
Ramesh A. Gopinath
ISCAS1
1993 Theory of modulated filter banks and modulated wavelet tight frames
Ramesh A. Gopinath, C. Sidney Burrus
ICASSP (3)1
1993 A tutorial overview of filter banks, wavelets and interrelations
Ramesh A. Gopinath, C. Sidney Burrus
ISCAS1
1992 Wavelet-based lowpass/bandpass interpolation
abstract
Wavelet-based lowpass and bandpass interpolation schemes that are exact for certain classes of signals including polynomials of arbitrarily large degree are discussed. The interpolation technique is studied in the context of wavelet-Galerkin approximation of the shift operator. A recursive dyadic interpolation algorithm makes it an attractive alternative to other schemes. It turns out that the Fourier transform of the lowpass interpolatory function is also (a positive) interpolatory function. The nature of the corresponding interpolating class is not well understood. Extension to the case of multiplicity M orthonormal wavelet bases, where there is an efficient M-adic interpolation scheme, is also given.>
Ramesh A. Gopinath, C. Sidney Burrus
ICASSP1
1992 Optimal wavelets for signal decomposition and the existence of scale-limited signals
abstract
Wavelet methods give a flexible alternative to Fourier methods in nonstationary signal analysis. The concept of band-limitedness plays a fundamental role in Fourier analysis. Since wavelet theory replaces frequency with scale, a natural question is whether there exists a useful concept of scale-limitedness. Obvious definitions of scale-limitedness are too restrictive, in that there would be few or no useful scale-limited signals. The authors introduce a viable definition for scale-limited signals, and show that the class is rich enough to include bandlimited signals, and impulse trains, among others. Moreover, for a wide choice of criteria, it is shown how to design the optimal wavelet for representing a given signal, and how to design robust wavelets that optimally represent certain classes of signals.>
Jan E. Odegard, Ramesh A. Gopinath, C. Sidney Burrus
ICASSP2
1991 Wavelet-Galerkin approximation of linear translation invariant operators
abstract
It is shown that the wavelet-Galerkin discretization of linear translation invariant (LTI) operators has good numerical properties, arising from the vanishing moments property of wavelets. If a wavelet has M vanishing moments, then it can have at most M-1 continuous derivatives, and hence operators of the form d/sup p//dx/sup p/, where p>M, have to be considered as generalized derivatives. Even in this case the approximation results derived hold. Also, the virtual expansion theorem is useful in the sense that there is no need to compute the expansion coefficients of the function at some level V/sub triangle x/.>
Ramesh A. Gopinath, Wayne Lawton, C. Sidney Burrus
ICASSP1
1990 Least squared error FIR filter design with spline transition functions
abstract
The authors propose the use of transition bands and transition functions in the ideal amplitude frequency response of a linear finite impulse response phase (FIR) digital filter to allow the analytical design of optimal least integral squared error filters with an explicit control of the transition band edges. Design formulas for approximations to ideal frequency responses which use a pth-order spline transition function are derived and analyzed. A variable-order spline is used with a method for optimally choosing the order. The results of this method are compared with those of other optimal methods.>
C. Sidney Burrus, Admadji W. Soewito, Ramesh A. Gopinath
ICASSP3
1990 Efficient computation of the wavelet transforms
abstract
An efficient algorithm is presented for the computation of the continuous wavelet transform (CWT). CWTs arise in the analysis of bandpass signals that are scale-time perturbed. To detect scale-time perturbed signals, the CWT of the received signal has to be computed with respect to the signals being detected. Efficient algorithms exist for the computation of the discrete wavelet transform, because the DWT wavelets give rise to quadrature mirror filters which can be implemented using lattice structures. Computation of the CWT with respect to an arbitrary wavelet (not necessarily a DWT wavelet), is time consuming. The DWT wavelets are used to compute the CWT of a signal with respect to an arbitrary wavelet. This is accomplished by using the computational properties of compactly supported DWT wavelet bases.>
Ramesh A. Gopinath, C. Sidney Burrus
ICASSP1