VLDB 2026 Research / reviewers in the wild / expert
Peder A. Olsen
dblp:13/5978
· DBLP profile ↗
76ranked-venue papers
14as first author
4since 2021 · last 2025
0000-0002-3836-8017ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 9 first-authorArtificial intelligence and machine learning · 44 · 8 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorTheory of computation · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Trustworthy machine learning · 36% Speech recognition and synthesis · 24% Probabilistic and Bayesian machine learning · 13% | |
| Theoretical computer science
4 papers |
Mathematical optimization · 100% | |
| Computer graphics and multimedia
3 papers |
Image and video coding · 50% Image and video processing · 43% Audio and music processing · 7% | |
| Computer networks
1 paper |
Vehicular, aerial and satellite networks · 50% Cellular and mobile networks · 50% |
Topics — the 30 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video coding
image compression |
0.9 | 1 | 2025 | Earth+: On-Board Satellite Imagery Compression Leveraging Historical Earth Observations · ASPLOS (1) 2025 |
Computer vision › Image recognition and object detection › object counting
crowd counting |
0.4 | 1 | 2020 | Crowd Counting with Decomposed Uncertainty · AAAI 2020 |
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty decomposition |
0.4 | 1 | 2020 | Crowd Counting with Decomposed Uncertainty · AAAI 2020 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.4 | 1 | 2020 | Crowd Counting with Decomposed Uncertainty · AAAI 2020 |
Mathematical optimization › continuous optimization
convex optimization |
0.4 | 2 | 2014 | QUIC & DIRTY: A Quadratic Approximation Approach for Dirty Statistical Models · NIPS 2014 Nuclear Norm Minimization via Active Subspace Selection · ICML 2014 |
Machine learning › Trustworthy machine learning › interpretability
explainable AI |
0.3 | 1 | 2018 | Improving Simple Models with Confidence Profiles · NeurIPS 2018 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2018 | Improving Simple Models with Confidence Profiles · NeurIPS 2018 |
Machine learning › Transfer learning and domain adaptation
knowledge transfer |
0.3 | 1 | 2018 | Improving Simple Models with Confidence Profiles · NeurIPS 2018 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2018 | Improving Simple Models with Confidence Profiles · NeurIPS 2018 |
Natural language and speech › Speech recognition and synthesis
acoustic modeling |
0.3 | 5 | 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced Languages · IEEE Trans. Speech Audio Process. 2012 Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2007 Subspace constrained Gaussian mixture models for speech recognition · IEEE Trans. Speech Audio Process. 2005 |
Mathematical optimization › numerical computation › numerical optimization
second-order methods |
0.3 | 2 | 2013 | Second Order Methods for Optimizing Convex Matrix Functions and Sparse Covariance Clustering · IEEE Trans. Speech Audio Process. 2013 Newton-Like Methods for Sparse Inverse Covariance Estimation · NIPS 2012 |
Cellular and mobile networks › downlink transmission
downlink capacity |
0.3 | 1 | 2025 | Earth+: On-Board Satellite Imagery Compression Leveraging Historical Earth Observations · ASPLOS (1) 2025 |
Vehicular, aerial and satellite networks
satellite constellation |
0.3 | 1 | 2025 | Earth+: On-Board Satellite Imagery Compression Leveraging Historical Earth Observations · ASPLOS (1) 2025 |
Image and video processing › image restoration › adverse weather image restoration
cloud removal |
0.2 | 1 | 2016 | Removing Clouds and Recovering Ground Observations in Satellite Image Sequences via Temporally Contiguous Robust Matrix Completion · CVPR 2016 |
Image and video processing
image restoration |
0.2 | 1 | 2016 | Removing Clouds and Recovering Ground Observations in Satellite Image Sequences via Temporally Contiguous Robust Matrix Completion · CVPR 2016 |
Image and video processing › image restoration
matrix completion |
0.2 | 1 | 2016 | Removing Clouds and Recovering Ground Observations in Satellite Image Sequences via Temporally Contiguous Robust Matrix Completion · CVPR 2016 |
Mathematical optimization › continuous optimization › matrix optimization › matrix recovery
matrix completion |
0.2 | 1 | 2014 | Nuclear Norm Minimization via Active Subspace Selection · ICML 2014 |
Mathematical optimization › continuous optimization
nonsmooth optimization |
0.2 | 1 | 2014 | Nuclear Norm Minimization via Active Subspace Selection · ICML 2014 |
Mathematical optimization › continuous optimization › convex optimization › norm optimization
nuclear norm minimization |
0.2 | 1 | 2014 | Nuclear Norm Minimization via Active Subspace Selection · ICML 2014 |
Mathematical optimization › continuous optimization › convex optimization › proximal methods
proximal gradient method |
0.2 | 1 | 2014 | Nuclear Norm Minimization via Active Subspace Selection · ICML 2014 |
Mathematical optimization
quadratic approximation |
0.2 | 1 | 2014 | QUIC & DIRTY: A Quadratic Approximation Approach for Dirty Statistical Models · NIPS 2014 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.2 | 3 | 2007 | Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2007 Subspace constrained Gaussian mixture models for speech recognition · IEEE Trans. Speech Audio Process. 2005 A robust high accuracy speech recognition system for mobile applications · IEEE Trans. Speech Audio Process. 2002 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
covariance estimation |
0.1 | 1 | 2012 | Newton-Like Methods for Sparse Inverse Covariance Estimation · NIPS 2012 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.1 | 1 | 2012 | Newton-Like Methods for Sparse Inverse Covariance Estimation · NIPS 2012 |
Natural language and speech › Speech recognition and synthesis › acoustic modeling
hidden markov model acoustic modeling |
0.1 | 1 | 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced Languages · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
low-resource speech recognition |
0.1 | 1 | 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced Languages · IEEE Trans. Speech Audio Process. 2012 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation › covariance estimation
sparse inverse covariance estimation |
0.1 | 1 | 2012 | Newton-Like Methods for Sparse Inverse Covariance Estimation · NIPS 2012 |
Mathematical optimization
continuous optimization |
0.1 | 1 | 2012 | Newton-Like Methods for Sparse Inverse Covariance Estimation · NIPS 2012 |
Natural language and speech › Speech recognition and synthesis › acoustic modeling
subspace constrained gaussian mixture models |
0.1 | 2 | 2007 | Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2007 Subspace constrained Gaussian mixture models for speech recognition · IEEE Trans. Speech Audio Process. 2005 |
Machine learning › Efficient and distributed learning
model deployment |
0.1 | 1 | 2018 | Improving Simple Models with Confidence Profiles · NeurIPS 2018 |
Methods — techniques the papers use, named apart from their topics
reference-based compression · 1.7change detection · 1.7temporal smoothness · 0.5robust matrix completion · 0.5low-rank optimization · 0.5neural network · 0.4bootstrap ensemble · 0.4quadratic approximation · 0.4sample weighting · 0.3linear probe · 0.3confidence profiling · 0.3conjugate gradient · 0.3FISTA · 0.3second-order optimization · 0.2proximal gradient · 0.2alternating least squares · 0.2active subspace selection · 0.2l1 regularization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Earth+: On-Board Satellite Imagery Compression Leveraging Historical Earth ObservationsabstractDue to limited downlink (satellite-to-ground) capacity, over 90% of the images captured by the earth-observation satellites are not downloaded to the ground. To overcome the downlink limitation, we present Earth+, a new on-board satellite imagery compression system that identifies and downloads only changed areas in each image compared to latest on-board reference images of the same location. The key of Earth+ is that it obtains latest on-board reference images by letting the ground stations upload images recently captured by all satellites in the constellation. To our best knowledge, Earth+ is the first system that leverages images across an entire satellite constellation to enable more images to be downloaded to the ground (by better satellite imagery compression). Our evaluation shows that to download images of the same area, Earth+ can reduce the downlink usage by 3.3× compared to state-of-the-art on-board image compression techniques without sacrificing imagery quality or using more resources (downlink, computation or storage). Kuntai Du, Yihua Cheng, Peder A. Olsen, Shadi A. Noghabi, Junchen Jiang |
ASPLOS (1) | 3 |
| 2023 | An Iterative Method for Hyperspectral Pixel Unmixing Leveraging Latent Dirichlet Variational AutoencoderabstractWe develop a hyperspectral pixel unmixing method that uses a Latent Variational Autoencoder within an analysis-synthesis loop to (1) construct pure spectra of the materials present in an image and (2) infer the mixing ratios of these materials in hyperspectral pixels without the need of labelled data. On OnTech-HIS-Syn-6em synthetic dataset that contains pixel unmixing groundtruth, the proposed method achieves acc = 100%, SAD = 0.0582 and RMSE = 0.0695 for segmentation, endmember extraction and abundance estimation, respectively. On HYDICE Urban benchmark, the proposed method achieves acc = 72.4%, SAD = 0.1669 and RMSE = 0.1984 for segmentation, endmember extraction and abundance estimation, respectively. Additionally, we applied this technique for crop analysis on hyperspectral data collected by the United States Department of Agriculture and achieved a coefficient of determination R2= 0.7 with respect to the ground truth. These results confirm that the proposed method is able to perform pixel unmixing without using labelled data. Kiran Mantripragada, Paul R. Adler, Peder A. Olsen, Faisal Z. Qureshi |
IGARSS | 3 |
| 2023 | Seeing Through Clouds in Satellite ImagesabstractThis article presents a neural network-based solution to recover pixels occluded by clouds in satellite images. We leverage radio frequency (RF) signals in the ultrahigh-/superhigh-frequency band that penetrates clouds to help reconstruct the occluded regions in multispectral images. We introduce the first multimodal multitemporal method for cloud removal. Our model uses publicly available satellite observations and produces daily cloud-free images. Experimental results show that our system outperforms several baselines on multiple metrics. We also demonstrate use cases of our system in digital agriculture, flood monitoring, and wildfire detection. Mingmin Zhao, Peder A. Olsen, Ranveer Chandra |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Location Aware Super-Resolution for Satellite Data FusionabstractSatellite data fusion involves images with different spatial, temporal, and spectral resolution. These images are taken under different illumination conditions, with different sensors and atmospheric noise. We use classic super-resolution algorithms to synthesize commercial satellite images (Pléiades) from a public satellite source (Sentinel-2). Each super-resolution method is then further improved by adaptive sharpening to the location by use of matrix completion (regression with missing pixels). Finally, we consider ensemble systems and a residual channel attention dual network with stochastic dropout. The resulting systems are visibly less blurry with higher fidelity and yield improved performance. Olaoluwa Adigun, Peder A. Olsen, Ranveer Chandra |
IGARSS | 2 |
| 2020 | Crowd Counting with Decomposed UncertaintyabstractResearch in neural networks in the field of computer vision has achieved remarkable accuracy for point estimation. However, the uncertainty in the estimation is rarely addressed. Uncertainty quantification accompanied by point estimation can lead to a more informed decision, and even improve the prediction quality. In this work, we focus on uncertainty estimation in the domain of crowd counting. With increasing occurrences of heavily crowded events such as political rallies, protests, concerts, etc., automated crowd analysis is becoming an increasingly crucial task. The stakes can be very high in many of these real-world applications. We propose a scalable neural network framework with quantification of decomposed uncertainty using a bootstrap ensemble. We demonstrate that the proposed uncertainty quantification method provides additional insight to the crowd counting problem and is simple to implement. We also show that our proposed method exhibits state-of-the-art performances in many benchmark crowd counting datasets. Min-hwan Oh, Peder A. Olsen, Karthikeyan Natesan Ramamurthy |
AAAI | 2 |
| 2018 | Detecting and Counting Panicles in Sorghum ImagesabstractPhenotyping, the process of measuring plant traits, plays a central role in plant breeding. However, traditional approaches are labor-intensive, time-consuming, costly, and error prone. Accurate, automated, high-throughput phenotyping can relieve a huge burden in the breeding pipeline. In this paper, we propose computer vision systems and approaches to annotate, detect, and count panicles (heads), a key phenotype, from aerial images of Sorghum crops. The annotation system allows the users to label panicles in Sorghum aerial images. This annotated data is used for learning by the panicle detection and counting algorithms. The proposed approaches were used with aerial imagery of 18 varieties of Sorghum crop collected at 6 different dates in the Midwestern United States. The detector has an AUC of over 0.98 and the counter has a mean absolute error of 2.66 without adapting to variety and 1.88 when using variety specific information. Our approaches are being adopted into a high-throughput phenotyping pipeline for accelerating Sorghum breeding. Peder A. Olsen, Karthikeyan Natesan Ramamurthy, Javier Ribera, Yuhao Chen 0001, Addie M. Thompson, Ronny Luss, Mitchell R. Tuinstra, Naoki Abe |
DSAA | 1 |
| 2018 | Improving Simple Models with Confidence ProfilesabstractIn this paper, we propose a new method called ProfWeight for transferring information from a pre-trained deep neural network that has a high test accuracy to a simpler interpretable model or a very shallow network of low complexity and a priori low test accuracy. We are motivated by applications in interpretability and model deployment in severely memory constrained environments (like sensors). Our method uses linear probes to generate confidence scores through flattened intermediate representations. Our transfer method involves a theoretically justified weighting of samples during the training of the simple model using confidence scores of these intermediate layers. The value of our method is first demonstrated on CIFAR-10, where our weighting method significantly improves (3-4\%) networks with only a fraction of the number of Resnet blocks of a complex Resnet model. We further demonstrate operationally significant results on a real manufacturing problem, where we dramatically increase the test accuracy of a CART model (the domain standard) by roughly $13\%$. Amit Dhurandhar, Karthikeyan Shanmugam 0001, Ronny Luss, Peder A. Olsen |
NeurIPS | 4 |
| 2016 | Removing Clouds and Recovering Ground Observations in Satellite Image Sequences via Temporally Contiguous Robust Matrix CompletionabstractWe consider the problem of removing and replacing clouds in satellite image sequences, which has a wide range of applications in remote sensing. Our approach first detects and removes the cloud-contaminated part of the image sequences. It then recovers the missing scenes from the clean parts using the proposed "TECROMAC" (TEmporally Contiguous RObust MAtrix Completion) objective. The objective function balances temporal smoothness with a low rank solution while staying close to the original observations. The matrix whose the rows are pixels and columns are days corresponding to the image, has low-rank because the pixels reflect land-types such as vegetation, roads and lakes and there are relatively few variations as a result. We provide efficient optimization algorithms for TECROMAC, so we can exploit images containing millions of pixels. Empirical results on real satellite image sequences, as well as simulated data, demonstrate that our approach is able to recover underlying images from heavily cloud-contaminated observations. Peder A. Olsen, Andrew Conn 0001, Aurélie C. Lozano |
CVPR | 2 |
| 2015 | Managing healthcare costs by peer-group modeling
Sholom M. Weiss, Casimir A. Kulikowski, Robert S. Galen, Peder A. Olsen, Ramesh Natarajan |
Appl. Intell. | 4 |
| 2014 | Nuclear Norm Minimization via Active Subspace SelectionabstractWe describe a novel approach to optimizing matrix problems involving nuclear norm regularization and apply it to the matrix completion problem. We combine methods from non-smooth and smooth optimization. At each step we use the proximal gradient to select an active subspace. We then find a smooth, convex relaxation of the smaller subspace problems and solve these using second order methods. We apply our methods to matrix completion problems including Netflix dataset, and show that they are more than 6 times faster than state-of-the-art nuclear norm solvers. Also, this is the first paper to scale nuclear norm solvers to the Yahoo-Music dataset, and the first time in the literature that the efficiency of nuclear norm solvers can be compared and even compete with non-convex solvers like Alternating Least Squares (ALS). Cho-Jui Hsieh, Peder A. Olsen |
ICML | 2 |
| 2014 | QUIC & DIRTY: A Quadratic Approximation Approach for Dirty Statistical Models
Cho-Jui Hsieh, Inderjit S. Dhillon, Pradeep Ravikumar, Stephen Becker, Peder A. Olsen |
NIPS | 5 |
| 2014 | Graphical Models for Identifying Fraud and Waste in Healthcare ClaimsabstractWe describe graphical model based methods for analyzing prescription and medical claims data in order to identify fraud and waste. Our approach draws on ideas from speech recognition and language modeling to identify patients, doctors and pharmacies whose prescription encounters show significant departure from normative behavior. We have analyzed claims data from a large healthcare provider, consisting of over 53 million individual prescription claims in the calendar year 2011. Peder A. Olsen, Ramesh Natarajan, Sholom M. Weiss |
SDM | 1 |
| 2014 | A variational approach to stable principal component pursuit
Aleksandr Y. Aravkin, Stephen Becker, Volkan Cevher, Peder A. Olsen |
UAI | 4 |
| 2013 | State of the art discriminative training of subspace constrained Gaussian mixture models in big training corporaabstractDiscriminatively trained full-covariance Gaussian mixture models have been shown to outperform its corresponding diagonal-covariance models on large vocabulary speech recognition tasks. However, the size of full-covariance model is much larger than that of diagonal-covariance model and is therefore not practical for use in a real system. In this paper, we present a method to build a large discriminatively trained full-covariance model with large (over 9000 hours) training corpora and still improve performance over the diagonal-covariance model. We then reduce the size of the full-covariance model to the size of its baseline diagonal-covariance model by using subspace constrained Gaussian mixture model (SCGMM). The resulting discriminatively trained SCGMM still retains the performance of its corresponding full-covariance model, and improves 5% relative over the same size diagonal-covariance model on a large vocabulary speech recognition task. Jing Huang 0019, Peder A. Olsen, Vaibhava Goel |
ICASSP | 2 |
| 2013 | Second Order Methods for Optimizing Convex Matrix Functions and Sparse Covariance ClusteringabstractA variety of first-order methods have recently been proposed for solving matrix optimization problems arising in machine learning. The premise for utilizing such algorithms is that second order information is too expensive to employ, and so simple first-order iterations are likely to be optimal. In this paper, we argue that second-order information is in fact efficiently accessible in many matrix optimization problems, and can be effectively incorporated into optimization algorithms. We begin by reviewing how certain Hessian operations can be conveniently represented in a wide class of matrix optimization problems, and provide the first proofs for these results. Next we consider a concrete problem, namely the minimization of the ℓ1regularized Jeffreys divergence, and derive formulae for computing Hessians and Hessian vector products. This allows us to propose various second order methods for solving the Jeffreys divergence problem. We present extensive numerical results illustrating the behavior of the algorithms and apply the methods to a speech recognition problem. We compress full covariance Gaussian mixture models utilized for acoustic models in automatic speech recognition. By discovering clusters of (sparse inverse) covariance matrices, we can compress the number of covariance parameters by a factor exceeding 200, while still outperforming the word error rate (WER) performance of a diagonal covariance model that has 20 times less covariance parameters than the original acoustic model. Gillian M. Chin, Jorge Nocedal, Peder A. Olsen, Steven J. Rennie |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Affine invariant sparse maximum a posteriori adaptationabstractModern speech applications utilize acoustic models with billions of parameters, and serve millions of users. Storing an acoustic model for each user is costly. We show through the use of sparse regularization, that it is possible to obtain competitive adaptation performance by changing only a small fraction of the parameters of an acoustic model. This allows for the compression of speaker-dependent models: a capability that has important implications for systems with millions of users. We achieve a performance comparable to the best Maximum A Posteriori (MAP) adaptation models while only adapting 5% of the acoustic model parameters. Thus it is possible to compress the speaker dependent acoustic models by close to a factor of 20. The proposed sparse adaptation criterion improves three aspects of previous work: It combines ℓ0and ℓ1penalties, have different adaptation rates for mean and variance parameters and is invariant to affine transformations. Peder A. Olsen, Jing Huang 0019, Steven J. Rennie, Vaibhava Goel |
ICASSP | 1 |
| 2012 | Newton-Like Methods for Sparse Inverse Covariance EstimationabstractWe propose two classes of second-order optimization methods for solving the sparse inverse covariance estimation problem. The first approach, which we call the Newton-LASSO method, minimizes a piecewise quadratic model of the objective function at every iteration to generate a step. We employ the fast iterative shrinkage thresholding method (FISTA) to solve this subproblem. The second approach, which we call the Orthant-Based Newton method, is a two-phase algorithm that first identifies an orthant face and then minimizes a smooth quadratic approximation of the objective function using the conjugate gradient method. These methods exploit the structure of the Hessian to efficiently compute the search direction and to avoid explicitly storing the Hessian. We show that quasi-Newton methods are also effective in this context, and describe a limited memory BFGS variant of the orthant-based Newton method. We present numerical results that suggest that all the techniques described in this paper have attractive properties and constitute useful tools for solving the sparse inverse covariance estimation problem. Comparisons with the method implemented in the QUIC software package are presented. Peder A. Olsen, Figen Öztoprak, Jorge Nocedal, Steven J. Rennie |
NIPS | 1 |
| 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced LanguagesabstractThis paper proposes an acoustic modeling approach based on bootstrap and restructuring to dealing with data sparsity for low-resourced languages. The goal of the approach is to improve the statistical reliability of acoustic modeling for automatic speech recognition (ASR) in the context of speed, memory and response latency requirements for real-world applications. In this approach, randomized hidden Markov models (HMMs) estimated from the bootstrapped training data are aggregated for reliable sequence prediction. The aggregation leads to an HMM with superior prediction capability at cost of a substantially larger size. For practical usage the aggregated HMM is restructured by Gaussian clustering followed by model refinement. The restructuring aims at reducing the aggregated HMM to a desirable model size while maintaining its performance close to the original aggregated HMM. To that end, various Gaussian clustering criteria and model refinement algorithms have been investigated in the full covariance model space before the conversion to the diagonal covariance model space in the last stage of the restructuring. Large vocabulary continuous speech recognition (LVCSR) experiments on Pashto and Dari have shown that acoustic models obtained by the proposed approach can yield superior performance over the conventional training procedure with almost the same run-time memory consumption and decoding speed. Peder A. Olsen, Pierre L. Dognin, Upendra V. Chaudhari, John R. Hershey, Bowen Zhou 0006 |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | Sparse Maximum A Posteriori adaptationabstractMaximum A Posteriori (MAP) adaptation is a powerful tool for building speaker specific acoustic models. Modern speech applications utilize acoustic models with millions of parameters, and serve millions of users. Storing an acoustic model for each user in such settings is costly. However, speaker specific acoustic models are generally similar to the acoustic model being adapted. By imposing sparseness constraints, we can save significantly on storage, and even improve the quality of the resulting speaker-dependent model. In this paper we utilize the ℓ1or ℓ0norm as a regularizer to induce sparsity. We show that we can obtain up to 95% sparsity with negligible loss in recognition accuracy, with both penalties. By removing small differences, which constitute “adaptation noise”, sparse MAP is actually able to improve upon MAP adaptation. Sparse MAP reduces the MAP word error rate by 2% relative at 89% sparsity. Peder A. Olsen, Jing Huang 0019, Vaibhava Goel, Steven J. Rennie |
ASRU | 1 |
| 2011 | Clustering of bootstrapped acoustic model with full covarianceabstractHMM-based acoustic models built from bootstrap are generally very large, especially when full covariance matrices are used for Gaussians. Therefore, clustering is needed to compact the acoustic model to a reasonable size for practical applications. This paper discusses and investigates multiple distance measurements and algorithms for the clustering. The distance measurements include Entropy, KL, Bhattacharyya, Chernoff and their weighted versions. For clustering algorithms, besides conventional greedy bottom-up, algorithms such as N-Best distance Refinement (NBR), K-step Look-Ahead (KLA), Breadth-First Searched (BFS) best path are proposed. A two-pass optimization approach is also proposed to improve the model structure. Experiments in the Bootstrap and Restructuring (B SRS) frame work on Pashto show that the discussed clustering approach can lead to better quality of the restructured model. It also shows that final acoustic model that is diagonalized from the full covariance yields good improvement over BSRS model directly with diagonal model and yields significant improvement over the conventional diagonal model. Peder A. Olsen, John R. Hershey, Bowen Zhou 0006, Yunxin Zhao |
ICASSP | 4 |
| 2011 | Front-end feature transforms with context filtering for speaker adaptationabstractFeature-space transforms such as feature-space maximum likelihood linear regression (FMLLR) are very effective speaker adaptation technique, especially on mismatched test data. In this study, we extend the full-rank square matrix of FMLLR to a non-square matrix that uses neighboring feature vectors in estimating the adapted central feature vector. Through optimizing an appropriate objective function we aim to filter out and transform features through the correlation of the feature context. We compare to FMLLR that just con sider the current feature vector only. Our experiments are conducted on the automobile data with different speed conditions. Results show that context filtering improves 23% on word error rate over conventional FMLLR on noisy 60mph data with adapted ML model, and 7%/9% improvement over the discriminatively trained FMMI/BMMI models. Jing Huang 0019, Karthik Visweswariah, Peder A. Olsen, Vaibhava Goel |
ICASSP | 3 |
| 2011 | A-Functions: A generalization of Extended Baum-Welch transformations to convex optimizationabstractWe introduce the Line Search A-Function (LSAF) technique that generalizes the Extended-Baum Welch technique in order to provide an effective optimization technique for a broader set of functions. We show how LSAF can be applied to functions of various probability density and distribution functions by demonstrating that these probability functions have an A-function. We also show that sparse representation problems (SR) that use 11 or combination of 11/12 regularization norms can also be efficiently optimized through an A-function derived for their objective functions. We will demonstrate the efficiency of LSAF for SR problems through simulations by comparing it with Approximate Bayesian Compressive Sensing method that we recently applied to speech recognition. Dimitri Kanevsky, David Nahamoo, Tara N. Sainath, Bhuvana Ramabhadran, Peder A. Olsen |
ICASSP | 5 |
| 2011 | Discriminative training for full covariance modelsabstractIn this paper we revisit discriminative training of full covariance acoustic models for automatic speech recognition. One of the difficult aspects of discriminative training is how to set the constant D that appears in the parameter updates. For diagonal covariance models, this constant D is set based on knowing the smallest value of D, D*, for which the resulting covariances remain positive definite. In this paper we show how to compute D* analytically, and show empirically that knowing this smallest value is important. Our baseline speech recognition models are state of the art broadcast news systems, built using the boosted Maximum Mutual Information criterion and feature space Maximum Mutual Information for feature selection. We show that discriminatively built full covariance models outperform our best diagonal covariance models. Moreover, full covariance models at optimal performance can be obtained by only a few discriminative iterations starting with a diagonal covariance model. The experiments also show that systems utilizing full covariance models are less sensitive to the choice of the number of gaussians. Peder A. Olsen, Vaibhava Goel, Steven J. Rennie |
ICASSP | 1 |
| 2011 | Rapid feature space MLLR speaker adaptation with bilinear modelsabstractIn this paper, we propose a novel method for rapid feature space Maximum Likelihood Linear Regression (FMLLR) speaker adaptation based on bilinear models. When the amount of adaptation data is limited, the conventional FMLLR transforms can be easily over-trained and can even degrade the performance. In such cases, usually by introducing structural constraints on the FMLLR transformation, the original FMLLR adaptation method can be modified for rapid adaptation. The objective of our bilinear model is to introduce a prior knowledge analysis on the training speakers based on Singular Vector Decomposition (SVD), and to incorporate it in the decoding process. This can effectively reduce the number of free parameters of FMLLR transformation and achieve performance improvements even with limited adaptation data. The efficiency of the proposed algorithm is demonstrated with experiments on the Mandarin digital dataset and the Mandarin voice search dataset respectively. Shilei Zhang, Peder A. Olsen, Yong Qin 0001 |
ICASSP | 2 |
| 2011 | Acoustic Modeling with Bootstrap and Restructuring Based on Full Covariance
Peder A. Olsen, John R. Hershey, Bowen Zhou 0006 |
INTERSPEECH | 4 |
| 2010 | Restructuring exponential family mixture modelsabstractVariational KL (varKL) divergence minimization was previously applied to restructuring acoustic models (AMs) using Gaussian mixture models by reducing their size while preserving their accuracy. In this paper, we derive a related varKL for exponential family mixture models (EMMs) and test its accuracy using the weighted local maximum likelihood agglomerative clustering technique. Minimizing varKL between a reference and a restructured AM led previously to the variational expectation maximization (varEM) algorithm; which we extend to EMMs. We present results on a clustering task using AMs trained on 50 hrs of Broadcast News (BN). EMMs are trained on fMMI-PLP features combined with frame level phone posterior probabilities given by the recently introduced sparse representation phone identification process. As we reduce model size, we test the word error rate using the standard BN test set and compare with baseline models of the same size, trained directly from data. Index Terms: KL divergence, variational approximation, variational expectation-maximization, exponential family distributions, acoustic model clustering. Pierre L. Dognin, John R. Hershey, Vaibhava Goel, Peder A. Olsen |
INTERSPEECH | 4 |
| 2010 | Incorporating sparse representation phone identification features in automatic speech recognition using exponential familiesabstractSparse representation phone identification features (SPIF) is a recently developed technique to obtain an estimate of phone posterior probabilities conditioned on an acoustic feature vector. In this paper, we explore incorporating SPIF phone posterior probability estimates in large vocabulary continuous speech recognition (LVCSR) task by including them as additional features of exponential densities that model the HMM state emission likelihoods. We compare our proposed approach to a number of other well known methods of combining feature streams or multiple LVCSR systems. Our experiments show that using exponential models to combine features results in a word error rate reduction of 0.5% absolute (18.7% down to 18.2%); this is comparable to best error rate reduction obtained from system combination methods, but without having to build multiple systems or tune the system combination weights. Vaibhava Goel, Tara N. Sainath, Bhuvana Ramabhadran, Peder A. Olsen, David Nahamoo, Dimitri Kanevsky |
INTERSPEECH | 4 |
| 2010 | Signal interaction and the devil functionabstractAbstract It is common in signal processing to model signals in the logpower spectrum domain. In this domain, when multiple signalsare present, they combine in a nonlinear way. If the phases ofthe signals are independent, then we can analyze the interactionin terms of a probability density we call the “devil function,”after its treacherous form. This paper derives an analytical ex-pression for the devil function, and discusses its properties withrespect to model-based signal enhancement. Exact inference inthis problem requires integrals involving the devil function thatare intractable. Previous methods have used approximations toderive closed-form solutions. However it is unknown how theseapproximations differ from the true interaction function in termsof performance. We propose Monte-Carlo methods for approx-imating the required integrals. Tests are conducted on a speechseparation and recognition problem to compare these methodswith past approximations. 1. Introduction Signals are often analyzed and classified using models of theirlog power spectra. In the context of noise, or any other inter-fering signal, such model-based classifiers must compensate forthe noise in some way. Model-based noise compensation re-quires a signal interaction model, which describes the effect ofadding two signals on the resulting acoustic features. Tradition-ally the influence of phase has either been ignored through theuse of approximate interaction models, or has been diminishedby averaging, especially when working in the log spectrum do-main.We describe and illustrate the signal interaction model, whichwe call the “devil function.” Exact inference using this func-tion is difficult because the required integrals are intractable.Even efficient approximate inference can be elusive. Just howimportant it is to accurately model this signal interaction is anempirical question. To address this question we perform exper-iments using accelerated Monte Carlo simulations that can ap-proximate inference arbitrarily well if enough samples are used,and compare the results to simpler approximations on a speechseparation and recognition task. John R. Hershey, Peder A. Olsen, Steven J. Rennie |
INTERSPEECH | 2 |
| 2010 | Modeling posterior probabilities using the linear exponential familyabstractA commonly used distribution on the probability simplex is the Dirichlet distribution. In this paper we present the linear exponential family as an alternative. The distribution is known in the statistics community, but we present in this paper a numerically stable method to compute its parameters. Although the Dirichlet distribution is known to be a good Bayesian prior for probabilities we believe this paper shows that the linear exponential model offers a good alternative in other contexts, such as when we want to use posterior probabilities as features for automatic speech recognition. We show how to incorporate posterior probabilities as additional features to an existing GMM, and show that the resulting model gives a 0.6% relative gain on a broadcast news speech recognition system. Index Terms: Linear Exponential Family, Simplex, Divided Difference, Partition Function Peder A. Olsen, Vaibhava Goel, Charles A. Micchelli, John R. Hershey |
INTERSPEECH | 1 |
| 2010 | Super-human multi-talker speech recognition: A graphical modeling approach
John R. Hershey, Steven J. Rennie, Peder A. Olsen, Trausti T. Kristjansson |
Comput. Speech Lang. | 3 |
| 2009 | Optimal quantization and bit allocation for compressing large discriminative feature space transformsabstractDiscriminative training of the feature space using the minimum phone error (MPE) objective function has been shown to yield remarkable accuracy improvements. These gains, however, come at a high cost of memory required to store the transform. In a previous paper we reduced this memory requirement by 94% by quantizing the transform parameters. We used dimension dependent quantization tables and learned the quantization values with a fixed assignment of transform parameters to quantization values. In this paper we refine and extend the techniques to attain a further 35% reduction in memory with no degradation in sentence error rate. We discuss a principled method to assign the transform parameters to quantization values. We also show how the memory can be gradually reduced using a Viterbi algorithm to optimally assign variable number of bits to dimension dependent quantization tables. The techniques described could also be applied to the quantization of general linear transforms - a problem that should be of wider interest. Etienne Marcheret, Vaibhava Goel, Peder A. Olsen |
ASRU | 3 |
| 2009 | Hierarchical variational loopy belief propagation for multi-talker speech recognitionabstractWe present a new method for multi-talker speech recognition using a single-channel that combines loopy belief propagation and variational inference methods to control the complexity of inference. The method models each source using an HMM with a hierarchical set of acoustic states, and uses the max model to approximate how the sources interact to generate mixed data. Inference involves inferring a set of probabilistic time-frequency masks to separate the speakers. By conditioning these masks on the hierarchical acoustic states of the speakers, the fidelity and complexity of acoustic inference can be precisely controlled. Acoustic inference using the algorithm scales linearly with the number of probabilistic time-frequency masks, and temporal inference scales linearly with LM size. Results on the monaural speech separation task (SSC) data demonstrate that the presented hierarchical variational max-sum product algorithm (HVMSP) outperforms VMSP by over 2% absolute using 4 times fewer probablistic masks. HVMSP furthermore performs on-par with the MSP algorithm, which utilizes exact conditional marginal likelihoods, using 256 times less time-frequency masks. Steven J. Rennie, John R. Hershey, Peder A. Olsen |
ASRU | 3 |
| 2009 | A fast, accurate approximation to log likelihood of Gaussian mixture modelsabstractIt has been a common practice in speech recognition and elsewhere to approximate the log likelihood of a Gaussian mixture model (GMM) with the maximum component log likelihood. While often a computational necessity, the max approximation comes at a price of inferior modeling when the Gaussian components significantly overlap. This paper shows how the approximation error can be reduced by changing component priors. In our experiments the loss in word error rate due to max approximation, albeit small, is reduced by 50-100% at no cost in computational efficiency. Furthermore, we expect acoustic models will become larger with time and increase component overlap and word error rate loss. This makes reducing the approximation error more relevant. The techniques considered do not use the original data and can easily be applied as a post-processing step to any GMM. Pierre L. Dognin, Vaibhava Goel, John R. Hershey, Peder A. Olsen |
ICASSP | 4 |
| 2009 | Refactoring acoustic models using variational density approximationabstractIn model-based pattern recognition it is often useful to change the structure, or refactor, a model. For example, we may wish to find a Gaussian mixture model (GMM) with fewer components that best approximates a reference model. One application for this arises in speech recognition, where a variety of model size requirements exists for different platforms. Since the target size may not be known a priori, one strategy is to train a complex model and subsequently derive models of lower complexity. We present methods for reducing model size without training data, following two strategies: GMM-approximation and Gaussian clustering based on divergences. A variational expectation-maximization algorithm is derived that unifies these two approaches. The resulting algorithms reduce the model size by 50% with less than 4% increase in error rate relative to the same-sized model trained on data. In fact, for up to 35% reduction in size, the algorithms can improve accuracy relative to baseline. Pierre L. Dognin, John R. Hershey, Vaibhava Goel, Peder A. Olsen |
ICASSP | 4 |
| 2009 | Single-channel speech separation and recognition using loopy belief propagationabstractWe address the problem of single-channel speech separation and recognition using loopy belief propagation in a way that enables efficient inference for an arbitrary number of speech sources. The graphical model consists of a set of N Markov chains, each of which represents a language model or grammar for a given speaker. A Gaussian mixture model with shared states is used to model the hidden acoustic signal for each grammar state of each source. The combination of sources is modeled in the log spectrum domain using non-linear interaction functions. Previously, temporal inference in such a model has been performed using an N-dimensional Viterbi algorithm that scales exponentially with the number of sources. In this paper, we describe a loopy message passing algorithm that scales linearly with language model size. The algorithm achieves human levels of performance, and is an order of magnitude faster than competitive systems for two speakers. Steven J. Rennie, John R. Hershey, Peder A. Olsen |
ICASSP | 3 |
| 2009 | Refactoring acoustic models using variational expectation-maximization
Pierre L. Dognin, John R. Hershey, Vaibhava Goel, Peder A. Olsen |
INTERSPEECH | 4 |
| 2009 | Acoustic modeling using exponential families
Vaibhava Goel, Peder A. Olsen |
INTERSPEECH | 2 |
| 2009 | Compacting discriminative feature space transforms for embedded devices
Etienne Marcheret, Petr Fousek, Peder A. Olsen, Vaibhava Goel |
INTERSPEECH | 4 |
| 2009 | Variational loopy belief propagation for multi-talker speech recognitionabstractWe address single-channel speech separation and recognition by combining loopy belief propagation and variational inference methods. Inference is done in a graphical model consisting of an HMM for each speaker combined with the max interaction model of source combination. We present a new variational inference algorithm that exploits the structure of the max model to compute an arbitrarily tight bound on the probability of the mixed data. The variational parameters are chosen so that the algorithm scales linearly in the size of the language and acoustic models, and quadratically in the number of sources. The algorithm scores 30.7% on the SSC task [1], which is the best published result by a method that scales linearly with speaker model complexity to date. The algorithm achieves average recognition error rates of 27%, 35%, and 51% on small datasets of SSC-derived speech mixtures containing two, three, and four sources, respectively, using a single audio channel. Index Terms: Speech separation, variational inference, loopy belief propagation, factorial hidden Markov models, ASR, Iroquois, Max model. Steven J. Rennie, John R. Hershey, Peder A. Olsen |
INTERSPEECH | 3 |
| 2008 | Accelerated Monte Carlo for Kullback-Leibler divergence between Gaussian mixture modelsabstractKullback Leibler (KL) divergence is widely used as a measure of dissimilarity between two probability distributions; however, the required integral is not tractable for gaussian mixture models (GMMs), and naive Monte-Carlo sampling methods can be expensive. Our work aims to improve the estimation of KL divergence for GMMs by sampling methods. We show how to accelerate Monte-Carlo sampling using variational approximations of the KL divergence. To this end we employ two different methodologies, control variates, and importance sampling. With control variates we use sampling to estimate the difference between the variational approximation and the the unknown KL divergence. With importance sampling, we estimate the KL divergence directly, using a sampling distribution derived from the variational approximation. We show that with these techniques we can achieve improvements in accuracy equivalent to using a factor of 30 times more samples. John R. Hershey, Peder A. Olsen, Emmanuel Yashchin |
ICASSP | 3 |
| 2008 | Variational Bhattacharyya divergence for hidden Markov modelsabstractMany applications require the use of divergence measures between probability distributions. Several of these, such as the Kullback-Leibler (KL) divergence and the Bhattacharyya divergence, are tractable for simple distributions such as Gaussians, but are intractable for more complex distributions such as hidden Markov models (HMMs) used in speech recognizers. For tasks related to classification error, the Bhattacharyya divergence is of special importance, due to its relationship with the Bayes error. Here we derive novel variational approximations to the Bhattacharyya divergence for HMMs. Remarkably the variational Bhattacharyya divergence can be computed in a simple closed-form expression for a given sequence length. One of the approximations can even be integrated over all possible sequence lengths in a closed-form expression. We apply the variational Bhattacharyya divergence for HMMs to word confusability, the problem of estimating the probability of mistaking one spoken word for another. John R. Hershey, Peder A. Olsen |
ICASSP | 2 |
| 2008 | Optimizing speech recognition grammars using a measure of similarity between hidden Markov modelsabstractIn this paper we discuss a method of optimizing weights in a stochastic finite state grammar using a measure of similarity between hidden Markov models. We compute the similarity using an edit distance and weights that are derived from the Bhattacharyya error between pairs of Gaussian mixture models. Forward-backward procedures are used to carry out the similarity computation, and to obtain the derivatives needed in gradient descent based optimization. We apply this procedure to the problem of estimating parameters of garbage models that are often included in SRGS grammars. Experimental results indicate that the method improves the garbage models and naturally results in models that are a function of their context in the grammar. Binit Mohanty, John R. Hershey, Peder A. Olsen, Suleyman Serdar Kozat, Vaibhava Goel |
ICASSP | 3 |
| 2008 | Efficient model-based speech separation and denoising using non-negative subspace analysisabstractWe present a new probabilistic architecture for analyzing composite non-negative data, called Non-negative Subspace Analysis (NSA). The NSA model provides a framework for understanding the relationships between sparse subspace and mixture model based approaches, and encompasses a range of models, including Sparse Non-negative Matrix Factorization (SNMF) [1] and mixture-model based analysis as special cases. We present a convenient instantiation of the NSA model, and an efficient variational approximate learning and inference algorithm that combines the advantages of SNMF and mixture model-based approaches. Preliminary recognition results on the Pascal Speech Separation Challenge 2006 test set [2], based on NSA separation results, are presented. The results fall short of those achieved by Algonquin [3], a state-of-the-art mixture-model based method, but considering that NSA runs an order of magnitude faster, the results are impressive. NSA outperforms SNMF in terms of word error rate (WER) on the task by a significant margin of over 9% absolute. Steven J. Rennie, John R. Hershey, Peder A. Olsen |
ICASSP | 3 |
| 2007 | Variational Kullback-Leibler divergence for Hidden Markov modelsabstractDivergence measures are widely used tools in statistics and pattern recognition. The Kullback-Leibler (KL) divergence between two hidden Markov models (HMMs) would be particularly useful in the fields of speech and image recognition. Whereas the KL divergence is tractable for many distributions, including Gaussians, it is not in general tractable for mixture models or HMMs. Recently, variational approximations have been introduced to efficiently compute the KL divergence and Bhattacharyya divergence between two mixture models, by reducing them to the divergences between the mixture components. Here we generalize these techniques to approach the divergence between HMMs using a recursive backward algorithm. Two such methods are introduced, one of which yields an upper bound on the KL divergence, the other of which yields a recursive closed-form solution. The KL and Bhattacharyya divergences, as well as a weighted edit-distance technique, are evaluated for the task of predicting the confusability of pairs of words. John R. Hershey, Peder A. Olsen, Steven J. Rennie |
ASRU | 2 |
| 2007 | Approximating the Kullback Leibler Divergence Between Gaussian Mixture ModelsabstractThe Kullback Leibler (KL) divergence is a widely used tool in statistics and pattern recognition. The KL divergence between two Gaussian mixture models (GMMs) is frequently needed in the fields of speech and image recognition. Unfortunately the KL divergence between two GMMs is not analytically tractable, nor does any efficient computational algorithm exist. Some techniques cope with this problem by replacing the KL divergence with other functions that can be computed efficiently. We introduce two new methods, the variational approximation and the variational upper bound, and compare them to existing methods. We discuss seven different techniques in total and weigh the benefits of each one against the others. To conclude we evaluate the performance of each one through numerical experiments. John R. Hershey, Peder A. Olsen |
ICASSP (4) | 2 |
| 2007 | Word confusability - measuring hidden Markov model similarityabstractWe address the problem of word confusability in speech recognition by measuring the similarity between Hidden Markov Models (HMMs) using a number of recently developed techniques. The focus is on defining a word confusability that is accurate, in the sense of predicting artificial speech recognition errors, and computationally efficient when applied to speech recognition applications. It is shown by using the edit distance framework for HMMs that we can use statistical information measures of distances between probability distribution functions to define similarity or distance measures between HMMs. We use correlation between errors in a real speech recognizer and the HMM similarities to measure how well each technique works. We demonstrate significant improvements relative to traditional phone confusion weighted edit distance measures by use of a Bhattacharyya divergence-based edit distance. Index Terms: Bayes Error, Bhattacharyya divergence, variational methods, gaussian mixture models, unscented transformation, Kullback‐Leibler distance rate. Peder A. Olsen, John R. Hershey |
INTERSPEECH | 2 |
| 2007 | Bhattacharyya error and divergence using variational importance samplingabstractMany applications require the use of divergence measures between probability distributions. Several of these, such as the Kullback Leibler (KL) divergence and the Bhattacharyya divergence,aretractableforsingleGaussians,butintractableforcomplex distributions such as Gaussian mixture models (GMMs) used in speech recognizers. For tasks related to classification error, theBhattacharyyadivergenceisofspecialimportance. Here we derive efficient approximations to the Bhattacharyya divergence for GMMs, using novel variational methods and importance sampling. We introduce a combination of the two, variational importance sampling (VISa), which performs importance sampling using a proposal distribution derived from the variational approximation. VISa achieves the same accuracy as naive importance sampling at a fraction of the computation. Finally we apply the Bhattacharyya divergence to compute word confusability and compare the corresponding estimates using the KL divergence. Index Terms: Variational importance sampling, Bhattacharyya divergence, variational methods, Gaussian mixture models. Peder A. Olsen, John R. Hershey |
INTERSPEECH | 1 |
| 2007 | Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech RecognitionabstractIn this paper, we study discriminative training of acoustic models for speech recognition under two criteria: maximum mutual information (MMI) and a novel "error-weighted" training technique. We present a proof that the standard MMI training technique is valid for a very general class of acoustic models with any kind of parameter tying. We report experimental results for subspace constrained Gaussian mixture models (SCGMMs), where the exponential model weights of all Gaussians are required to belong to a common "tied" subspace, as well as for subspace precision and mean (SPAM) models which impose separate subspace constraints on the precision matrices (i.e., inverse covariance matrices) and means. It has been shown previously that SCGMMs and SPAM models generalize and yield significant error rate improvements over previously considered model classes such as diagonal models, models with semitied covariances, and extended maximum likelihood linear transformation (EMLLT) models. We show here that MMI and error-weighted training each individually result in over 20% relative reduction in word error rate on a digit task over maximum-likelihood (ML) training. We also show that a gain of as much as 28% relative can be achieved by combining these two discriminative estimation techniques Scott Axelrod, Vaibhava Goel, Ramesh A. Gopinath, Peder A. Olsen, Karthik Visweswariah |
IEEE Trans. Speech Audio Process. | 4 |
| 2006 | Dynamic Noise AdaptationabstractWe consider the problem of robust speech recognition in the car environment. We present a new dynamic noise adaptation algorithm, called DNA, for the robust front-end compensation of evolving semi-stationary noise as typically encountered in the car setting. A large dataset of in-car noise was collected for the evaluation of the new algorithm. This dataset was combined with the Aurora II framework to produce a new, publicly available framework, called DNA + AURORA II, for the evaluation of adaptive noise compensation algorithms. We show that DNA consistently outperforms several existing, related state-of-the-art front-end denoising techniques Steven J. Rennie, Trausti T. Kristjansson, Peder A. Olsen, Ramesh A. Gopinath |
ICASSP (1) | 3 |
| 2006 | Super-human multi-talker speech recognition: the IBM 2006 speech separation challenge systemabstractWe describe a system for model based speech separation which achieves super-human recognition performance when two talkers speak at similar levels. The system can separate the speech of two speakers from a single channel recording with remarkable results. It incorporates a novel method for performing two-talker speaker identification and gain estimation. We extend the method of model based high resolution signal reconstruction to incorporate temporal dynamics. We report on two methods for introducing dynamics; the first uses dynamics in the acoustic model space, the second incorporates dynamics based on sentence grammar. The addition of temporal constraints leads to dramatic improvements in the separation performance. Once the signals have been separated they are then recognized using speaker dependent labeling. 1. Trausti T. Kristjansson, John R. Hershey, Peder A. Olsen, Steven J. Rennie, Ramesh A. Gopinath |
INTERSPEECH | 3 |
| 2006 | Single Channel Speech Separation Using Factorial DynamicsabstractHuman listeners have the extraordinary ability to hear and recognize speech even when more than one person is talking. Their machine counterparts have historically been unable to compete with this ability, until now. We present a modelbased system that performs on par with humans in the task of separating speech of two talkers from a single-channel recording. Remarkably, the system surpasses human recognition performance in many conditions. The models of speech use temporal dynamics to help infer the source speech signals, given mixed speech signals. The estimated source signals are then recognized using a conventional speech recognition system. We demonstrate that the system achieves its best performance when the model of temporal dynamics closely captures the grammatical constraints of the task. One of the hallmarks of human perception is our ability to solve the auditory cocktail party problem: we can direct our attention to a given speaker in the presence of interfering speech, and understand what was said remarkably well. Until now the same could not be said for automatic speech recognition systems. However, we have recently introduced a system which in many conditions performs this task better than humans [1][2]. The model addresses the Pascal Speech Separation Challenge task [3], and outperforms all other published results by more than 10% word error rate (WER). In this model, dynamics are modeled using a layered combination of one or two Markov chains: one for long-term dependencies and another for short-term dependencies. The combination of the two speakers was handled via an iterative Laplace approximation method known as Algonquin [4]. Here we describe experiments that show better performance on the same task with a simpler version of the model. The task we address is provided by the PASCAL Speech Separation Challenge [3], which provides standard training, development, and test data sets of single-channel speech mixtures following an arbitrary but simple grammar. In addition, the challenge organizers have conducted human-listening experiments to provide an interesting baseline for comparison of computational techniques. The overall system we developed is composed of the three components: a speaker identification and gain estimation component, a signal separation component, and a speech recognition system. In this paper we focus on the signal separation component, which is composed of the acoustic and grammatical models. The details of the other components are discussed in [2]. Single-channel speech separation has previously been attempted using Gaussian mixture models (GMMs) on individual frames of acoustic features. However such models tend to perform well only when speakers are of different gender or have rather different voices [4]. When speakers have similar voices, speaker-dependent mixture models cannot unambiguously identify the component speakers. In such cases it is helpful to model the temporal dynamics of the speech. Several models in the literature have attempted to do so either for recognition [5, 6] or enhancement [7, 8] of speech. Such models have typically been based on a discrete-state hidden Markov model (HMM) operating on a frame-based acoustic feature vector. Modeling the dynamics of the log spectrum of speech is challenging in that different speech components evolve at different time-scales. For example the excitation, which carries mainly pitch, versus the filter, which consists of the formant structure, are somewhat independent of each other. The formant structure closely follows the sequences of phonemes in each word, which are pronounced at a rate of several per second. In non-tonal languages such as English, the pitch fluctuates with prosody over the course of a sentence, and is not directly coupled with the words being spoken. Nevertheless, it seems to be important in separating speech, because the pitch harmonics carry predictable structure that stands out against the background. We address the various dynamic components of speech by testing different levels of dynamic constraints in our models. We explore four different levels of dynamics: no dynamics, low-level acoustic dynamics, high-level grammar dynamics, and a layered combination, dual dynamics, of the acoustic and grammar dynamics. The grammar dynamics and dual dynamics models perform the best in our experiments. The acoustic models are combined to model mixtures of speech using two methods: a nonlinear model known as Algonquin, which models the combination of log-spectrum models as a sum in the power spectrum, and a simpler max model that combines two log spectra using the max function. It turns out that whereas Algonquin works well, our formulation of the max model does better overall. With the combination of the max model and grammar-level dynamics, the model produces remarkable results: it is often able to extract two utterances from a mixture even when they are from the same speaker 1 . Overall results are given in Table 1, which shows that our closest competitors are human listeners. Table 1: Overall word error rates across all conditions on the challenge task. Human: average human error rate, IBM : our best result, Next Best: the best of the eight other published results on this task, and Chance: the theoretical error rate for random guessing. System: Word Error Rate: Human 22.3% IBM 22.6% Next Best 34.2% Chance 93.0% John R. Hershey, Trausti T. Kristjansson, Steven J. Rennie, Peder A. Olsen |
NIPS | 4 |
| 2005 | Initializing Subspace Constrained Gaussian Mixture ModelsabstractA recent series of papers [1, 2, 3, 4] introduced subspace constrained Gaussian mixture models (SCGMM) and showed that SCGMM can very efficiently approximate full covariance Gaussian mixture models (FCGMM); a significant reduction in the number of parameters is achieved with little loss in the accuracy of the model. SCGMM were arrived at as a sequence of generalizations of diagonal covariance GMM. As an artifact of this process the initialization of SCGMM parameters in that work is complex, i.e., relies on best parameter settings of less general models. This paper overcomes this problem by showing how an FCGMM can be used to give a simple and direct initialization of an SCGMM. The initialization scheme is powerful enough that as the number of parameters in an SCGMM approaches that of an FCGMM (i.e., large SCGMM) further training of the SCGMM is unnecessary. Peder A. Olsen, Karthik Visweswariah, Ramesh A. Gopinath |
ICASSP (1) | 1 |
| 2005 | Voicing features for robust speech detection
Trausti T. Kristjansson, Sabine Deligne, Peder A. Olsen |
INTERSPEECH | 3 |
| 2005 | Feature adaptation using projection of Gaussian posteriors
Karthik Visweswariah, Peder A. Olsen |
INTERSPEECH | 2 |
| 2005 | Subspace constrained Gaussian mixture models for speech recognitionabstractA standard approach to automatic speech recognition uses hidden Markov models whose state dependent distributions are Gaussian mixture models. Each Gaussian can be viewed as an exponential model whose features are linear and quadratic monomials in the acoustic vector. We consider here models in which the weight vectors of these exponential models are constrained to lie in an affine subspace shared by all the Gaussians. This class of models includes Gaussian models with linear constraints placed on the precision (inverse covariance) matrices (such as diagonal covariance, maximum likelihood linear transformation, or extended maximum likelihood linear transformation), as well as the LDA/HLDA models used for feature selection which tie the part of the Gaussians in the directions not used for discrimination. In this paper, we present algorithms for training these models using a maximum likelihood criterion. We present experiments on both small vocabulary, resource constrained, grammar-based tasks, as well as large vocabulary, unconstrained resource tasks to explore the rather large parameter space of models that fit within our framework. In particular, we demonstrate significant improvements can be obtained in both word error rate and computational complexity. Scott Axelrod, Vaibhava Goel, Ramesh A. Gopinath, Peder A. Olsen, Karthik Visweswariah |
IEEE Trans. Speech Audio Process. | 4 |
| 2004 | Fast clustering of Gaussians and the virtue of representing Gaussians in exponential model format
Peder A. Olsen, Karthik Visweswariah |
INTERSPEECH | 1 |
| 2004 | Modeling inverse covariance matrices by basis expansionabstractThis paper proposes a new covariance modeling technique for Gaussian mixture models. Specifically the inverse covariance (precision) matrix of each Gaussian is expanded in a rank-1 basis i.e., /spl Sigma//sub j//sup -1/=P/sub j/=/spl Sigma//sub k=1//sup D//spl lambda//sub k//sup j/a/sub k/a/sub k//sup T/, /spl lambda//sub k//sup j//spl isin//spl Ropf/,a/sub k//spl isin//spl Ropf//sup d/. A generalized EM algorithm is proposed to obtain maximum likelihood parameter estimates for the basis set {a/sub k/a/sub k//sup T/}/sub k=1//sup D/ and the expansion coefficients {/spl lambda//sub k//sup j/}. This model, called the extended maximum likelihood linear transform (EMLLT) model, is extremely flexible: by varying the number of basis elements from D=d to D=d(d+1)/2 one gradually moves from a maximum likelihood linear transform (MLLT) model to a full-covariance model. Experimental results on two speech recognition tasks show that the EMLLT model can give relative gains of up to 35% in the word error rate over a standard diagonal covariance model, 30% over a standard MLLT model. Peder A. Olsen, Ramesh A. Gopinath |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | Dimensional reduction, covariance modeling, and computational complexity in ASR systemsabstractWe study acoustic modeling for speech recognition using mixtures of exponential models with linear and quadratic features tied across all context dependent states. These models are one version of the SPAM models introduced by Axelrod, Gopinath and Olsen (see Proc. ICSLP, 2002). They generalize diagonal covariance, MLLT, EMLLT, and full covariance models. Reduction of the dimension of the acoustic vectors using LDA/HDA projections corresponds to a special case of reducing the exponential model feature space. We see, in one speech recognition task, that SPAM models on an LDA projected space of varying dimensions achieve a significant fraction of the WER improvement in going from MLLT to full covariance modeling, while maintaining the low computational cost of the MLLT models. Further, the feature precomputation cost can be minimized using the hybrid feature technique of Visweswariah, Olsen, Gopinath and Axelrod (see ICASSP 2003); and the number of Gaussians one needs to compute can be greatly reducing using hierarchical clustering of the Gaussians (with fixed feature space). Finally, we show that reducing the quadratic and linear feature spaces separately produces models with better accuracy, but comparable computational complexity, to LDA/HDA based models. Scott Axelrod, Ramesh A. Gopinath, Peder A. Olsen, Karthik Visweswariah |
ICASSP (1) | 3 |
| 2003 | Maximum likelihood training of subspaces for inverse covariance modelingabstractSpeech recognition systems typically use mixtures of diagonal Gaussians to model the acoustics. Using Gaussians with a more general covariance structure can give improved performance; EM-LLT and SPAM models give improvements by restricting the inverse covariance to a linear/affine subspace spanned by rank one and full rank matrices respectively. We consider training these subspaces to maximize likelihood. For EMLLT ML training the subspace results in significant gains over the scheme proposed by Olsen and Gopinath (see Proceedings of ICASSP, 2002). For SPAM ML training of the subspace slightly improves performance over the method reported by Axelrod, Gopinath and Olsen (see Proceedings of ICSLP, 2002). For the same subspace size an EMLLT model is more efficient computationally than a SPAM model, while the SPAM model is more accurate. This paper proposes a hybrid method of structuring the inverse covariances that both has good accuracy and is computationally efficient. Karthik Visweswariah, Peder A. Olsen, Ramesh A. Gopinath, Scott Axelrod |
ICASSP (1) | 2 |
| 2003 | Discriminative estimation of subspace precision and mean (SPAM) modelsabstractThe SPAM model was recently proposed as a very general method for modeling Gaussians with constrained means and covariances. It has been shown to yield significant error rate improvements over other methods of constraining covariances such as diagonal, semi-tied covariances, and extended maximum likelihood linear transformations. In this paper we address the problem of discriminative estimation of SPAM model parameters, in an attempt to further improve its performance. We present discriminative estimation under two criteria: maximum mutual information (MMI) and an "error-weighted" training. We show that both these methods individually result in over 20% relative reduction in word error rate on a digit task over maximum likelihood (ML) estimated SPAM model parameters. We also show that a gain of as much as 28% relative can be achieved by combining these two discriminative estimation techniques. The techniques developed in this paper also apply directly to an extension of SPAM called subspace constrained exponential models. Vaibhava Goel, Scott Axelrod, Ramesh A. Gopinath, Peder A. Olsen, Karthik Visweswariah |
INTERSPEECH | 4 |
| 2003 | An efficient integrated gender detection scheme and time mediated averaging of gender dependent acoustic modelsabstractThis paper discusses building gender dependent gaussian mixture models (GMMs) and how to integrate these with an efficient gender detection scheme. Gender specific acoustic models of half the size of a corresponding gender independent acoustic model substantially outperform the larger gender independent acoustic models. With perfect gender detection, gender dependent modeling should therefore yield higher recognition accuracy without consuming more memory. Furthermore, as certain phonemes are inherently gender independent (e.g. silence) much of the male and female specific acoustic models can be shared. This paper proposes how to discover which phonemes are inherently similar for male and female speakers and how to efficiently share this information between gender dependent GMMs. A highly accurate gender detection scheme is suggested that takes advantage of computations inherently done in the speech recognizer to detect the gender at a computational cost that is negligible. By making the gender assignment probabilistic an increase in word error rate (WER) seen for erroneously gender labeled speakers is avoided. The method of gender detection and probabilistic use of gender is novel and should be of interest beyond mere gender detection. The only requirement for the method to work is that the training data be appropriately labeled. 1. Peder A. Olsen, Satya Dharanipragada |
INTERSPEECH | 1 |
| 2002 | Adaptation experiments on the SPINE database with the Extended Maximum Likelihood Linear Transformation (EMLLT) modelabstractThis paper applies the recently proposed Extended Maximum Likelihood Linear Transformation (EMLLT) model for inverse covariances in a Speaker Adaptive Training (SAT) context. The paper adapts standard algorithms for maximum likelihood estimation of linear transforms for mean, variance and feature space adaptation respectively, to the EMLLT model. Experimental results showing word-error-rate improvements are reported on the SPINE2 database. The system described here is the best-performing system submitted by IBM in the SPINE2 evaluation conducted by NIST in October 2001. Ramesh A. Gopinath, Vaibhava Goel, Karthik Visweswariah, Peder A. Olsen |
ICASSP | 4 |
| 2002 | Modeling inverse covariance matrices by basis expansionabstractThis paper proposes a new covariance modeling technique for Gaussian Mixture Models. Specifically the inverse covariance (precision) matrix of each Gaussian is expanded in a rank-1 basis i.e., Σj−1= Pj= Σk = 1DλkjakakT, λkj∈ ℝd. A generalized EM algorithm is proposed to obtain maximum likelihood parameter estimates for the basis set {akakT} and the expansion coefficients {λkj}. This model, called the Extended Maximum Likelihood Linear Transform (EMLLT) model, is extremely flexible: by varying the number of basis elements from d to d(d + 1)/2 one gradually moves from a Maximum Likelihood Linear Transform (MLLT) model to a full-covariance model. Experimental results on two speech recognition tasks show that the EMLLT model can give relative gains of up to 35% in the word error rate over a standard diagonal covariance model. Peder A. Olsen, Ramesh A. Gopinath |
ICASSP | 1 |
| 2002 | Modeling with a subspace constraint on inverse covariance matricesabstractWe consider a family of Gaussian mixture models for use in HMM based speech recognition system. These "SPAM" models have state independent choices of subspaces to which the precision (inverse covariance) matrices and means are restricted to belong. They provide a flexible tool for robust, compact, and fast acoustic modeling. The focus of this paper is on the case where the means are unconstrained. The models in the case already generalize the recently introduced EMLLT models, which themselves interpolate between MLLT and full covariance models. We describe an algorithm to train both the state-dependent and state-independent parameters. Results are reported on one speech recognition task. The SPAM models are seen to yield significant improvements in accuracy over EMLLT models with comparable model size and runtime speed. We find a 10% relative reduction in error rate over an MLLT model can be obtained while decreasing the acoustic modeling time by 20%. Scott Axelrod, Ramesh A. Gopinath, Peder A. Olsen |
INTERSPEECH | 3 |
| 2002 | Large vocabulary conversational speech recognition with the extended maximum likelihood linear transformation (EMLLT) modelabstractThis paper applies the recently proposed Extended Maximum Likelihood Linear Transformation (EMLLT) model in a Speaker Adaptive Training (SAT) context on the Switchboard database. Adaptation is carried out with maximum likelihood estimation of linear transforms for the means, precisions (inverse covariances) and the feature-space under the EMLLT model. This paper shows the first experimental evidence that significant word-error-rate improvements can be achieved with the EMLLT model (in both VTL and VTL+SAT training contexts) over a state-of-the-art diagonal covariance model in a difficult large-vocabulary conversational speech recognition task. The improvements were of the order of 1% absolute in multiple scenarios. Jing Huang 0019, Vaibhava Goel, Ramesh A. Gopinath, Brian Kingsbury, Peder A. Olsen, Karthik Visweswariah |
INTERSPEECH | 5 |
| 2002 | Theory and practice of acoustic confusability
Harry Printz, Peder A. Olsen |
Comput. Speech Lang. | 2 |
| 2002 | Automatic transcription of Broadcast News
Scott Saobing Chen, Ellen Eide, Mark J. F. Gales, Ramesh A. Gopinath, D. Kanvesky, Peder A. Olsen |
Speech Commun. | 6 |
| 2002 | A robust high accuracy speech recognition system for mobile applicationsabstractThis paper describes a robust, accurate, efficient, low-resource, medium-vocabulary, grammar-based speech recognition system using hidden Markov models for mobile applications. Among the issues and techniques we explore are improving robustness and efficiency of the front-end, using multiple microphones for removing extraneous signals from speech via a new multichannel CDCN technique, reducing computation via silence detection, applying the Bayesian information criterion (BIC) to build smaller and better acoustic models, minimizing finite state grammars, using hybrid maximum likelihood and discriminative models, and automatically generating baseforms from single new-word utterances. Sabine Deligne, Satya Dharanipragada, Ramesh A. Gopinath, Benoît Maison, Peder A. Olsen, Harry Printz |
IEEE Trans. Speech Audio Process. | 5 |
| 2001 | Speech recognition for DARPA CommunicatorabstractWe report the results of investigations in acoustic modeling, language modeling and decoding techniques, for the DARPA Communicator, a speaker-independent, telephone-based dialog system. By a combination of methods, including enlarging the acoustic model, augmenting the recognizer vocabulary, conditioning the language model upon the dialog state, and applying a post-processing decoding method, we lowered the overall word error rate from 21.9% to 15.0%, a gain of 6.9% absolute and 31.5% relative. Andrew Aaron, Scott Saobing Chen, Paul S. Cohen, Satya Dharanipragada, Ellen Eide, Martin Franz, Jean-Michel LeRoux, X. Luo, Benoît Maison, Lidia Mangu, T. Mathes, Miroslav Novak, Peder A. Olsen, Michael Picheny, Harry Printz, Bhuvana Ramabhadran, Andrej Sakrajda, George Saon, Borivoj Tydlitát, Karthik Visweswariah, D. Yuk |
ICASSP | 13 |
| 2001 | Low-resource hidden Markov model speech recognition
Sabine Deligne, Ellen Eide, Ramesh A. Gopinath, Dimitri Kanevsky, Benoît Maison, Peder A. Olsen, Harry Printz, Jan Sedivý |
INTERSPEECH | 6 |
| 2000 | Transcription of broadcast news with a time constraint: IBM's 10xRT HUB4 systemabstractWe describe a system which automatically transcribes broadcast news in less than 10 times real-time. We detail the system architecture of this system, which was used by IBM in the 1999 HUB4 10xRT evaluation, and show that the performance of this system is over 20 percent more accurate at the same speed than the system we used in the 1998 evaluation. Furthermore, we have closed the gap in word recognition accuracy between an unlimited resource system and this which runs in under 10 times real time from 45 percent to 14 percent. Ellen Eide, Benoît Maison, Dimitri Kanevsky, Peder A. Olsen, Scott Saobing Chen, Lidia Mangu, Mark J. F. Gales, Miroslav Novak, Ramesh A. Gopinath |
INTERSPEECH | 4 |
| 1999 | Maximum likelihood estimates for exponential type density familiesabstractWe consider a parametric family of density functions of the type exp(-|x|/sup /spl alpha//2/) for modeling acoustic feature vectors used in automatic recognition of speech. The parameter /spl alpha/ is a measure of the impulsiveness as well as the nongaussian nature of the data. While previous work has focused on estimating the mean and the variance of the data here we attempt to estimate the impulsiveness /spl alpha/ from the data on a maximum likelihood basis. We show that there is a balance between /spl alpha/ and the number of data points N that must be satisfied before maximum likelihood estimation is carried out. Numerical experiments are performed on multidimensional vectors obtained from speech data. Sankar Basu, Charles A. Micchelli, Peder A. Olsen |
ICASSP | 3 |
| 1999 | Recent improvements to IBM's speech recognition system for automatic transcription of broadcast newsabstractWe describe extensions and improvements to IBM's system for automatic transcription of broadcast news. The speech recognizer uses a total of 160 hours of acoustic training data, 80 hours more than for the system described in Chen et al. (1998). In addition to improvements obtained in 1997 we made a number of changes and algorithmic enhancements. Among these were changing the acoustic vocabulary, reducing the number of phonemes, insertion of short pauses, mixture models consisting of non-Gaussian components, pronunciation networks, factor analysis (FACILT) and Bayesian information criteria (BIC) applied to choosing the number of components in a Gaussian mixture model. The models were combined in a single system using NIST's script voting machine known as rover (Fiscus 1997). Scott Saobing Chen, Ellen Eide, Mark J. F. Gales, Ramesh A. Gopinath, Dimitri Kanevsky, Peder A. Olsen |
ICASSP | 6 |
| 1999 | Tail distribution modelling using the richter and power exponential distributions
Mark J. F. Gales, Peder A. Olsen |
EUROSPEECH | 2 |
| 1998 | Transcription of broadcast news-some recent improvements to IBM's LVCSR systemabstractThis paper describes extensions and improvements to IBM's large vocabulary continuous speech recognition (LVCSR) system for transcription of broadcast news. The recognizer uses an additional 35 hours of training data over the one used in the 1996 Hub4 evaluation. It includes a number of new features: optimal feature space for acoustic modeling (in training and/or testing), filler-word modeling, Bayesian information criterion (BIC) based segment clustering, an improved implementation of iterative MLLR and 4-gram language models. Results using the 1996 DARPA Hub4 evaluation data set are presented. Lazaros Polymenakos, Peder A. Olsen, D. Kanvesky, Ramesh A. Gopinath, Ponani S. Gopalakrishnan, Scott Saobing Chen |
ICASSP | 2 |
| 1992 | A note on irregular discrete wavelet transformsabstractEstimates of frame bounds for wavelet frames generated by discrete sets in phase space satisfying only a certain density condition are deduced. Numerical examples show that these estimates, which are the sharpest results of this kind known to the authors, are rather poor compared to Daubechies' estimates (see ibid., vol.36, no.5, p.961-1005, 1990) for the usual discrete wavelet transforms.> Peder A. Olsen, Kristian Seip |
IEEE Trans. Inf. Theory | 1 |