VLDB 2026 Research / reviewers in the wild / expert
Matt Shannon
dblp:38/9012
· DBLP profile ↗
15ranked-venue papers
7as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Speech recognition and synthesis · 62% Trustworthy machine learning · 16% Learning paradigms · 12% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.6 | 1 | 2022 | Global Normalization for Streaming Speech Recognition in a Modular Framework · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › dataset bias
label bias |
0.6 | 1 | 2022 | Global Normalization for Streaming Speech Recognition in a Modular Framework · NeurIPS 2022 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › real-time speech recognition
streaming speech recognition |
0.6 | 1 | 2022 | Global Normalization for Streaming Speech Recognition in a Modular Framework · NeurIPS 2022 |
Natural language and speech › Speech recognition and synthesis › speech synthesis
controllable speech synthesis |
0.4 | 1 | 2020 | Semi-Supervised Generative Modeling for Controllable Speech Synthesis · ICLR 2020 |
Machine learning › Learning paradigms › semi-supervised learning
generative semi-supervised learning |
0.4 | 1 | 2020 | Semi-Supervised Generative Modeling for Controllable Speech Synthesis · ICLR 2020 |
Natural language and speech › Speech recognition and synthesis
speech synthesis |
0.4 | 1 | 2020 | Semi-Supervised Generative Modeling for Controllable Speech Synthesis · ICLR 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › hidden markov model
autoregressive hidden markov model |
0.2 | 1 | 2013 | Autoregressive Models for Statistical Parametric Speech Synthesis · IEEE Trans. Speech Audio Process. 2013 |
Natural language and speech › Speech recognition and synthesis › speech synthesis
statistical parametric speech synthesis |
0.2 | 1 | 2013 | Autoregressive Models for Statistical Parametric Speech Synthesis · IEEE Trans. Speech Audio Process. 2013 |
Methods — techniques the papers use, named apart from their topics
globally normalized autoregressive transducer · 0.6semi-supervised learning · 0.4generative modeling · 0.4global variance parameter generation · 0.2expectation-maximization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-SpeechabstractEric Battenberg, RJ Skerry-Ryan, Daisy Stanton, Soroosh Mariooryad, Matt Shannon, Julian Salazar, David Teh-Hwa Kao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Eric Battenberg, R. J. Skerry-Ryan, Daisy Stanton, Soroosh Mariooryad, Matt Shannon, Julian Salazar, David Kao |
NAACL (Long Papers) | 5 |
| 2022 | Speaker GenerationabstractThis work explores the task of synthesizing speech in non-existent human-sounding voices. We call this task "speaker generation", and present TacoSpawn, a system that performs competitively at this task. TacoSpawn is a recurrent attention-based text-to-speech model that learns a distribution over a speaker embedding space, which enables sampling of novel and diverse speakers. Our method is easy to implement, and does not require transfer learning from speaker ID systems. We present objective and subjective metrics for evaluating performance on this task, and demonstrate that our proposed objective metrics correlate with human perception of speaker similarity. Audio samples are available on our demo page1. Daisy Stanton, Matt Shannon, Soroosh Mariooryad, R. J. Skerry-Ryan, Eric Battenberg, Tom Bagby, David Kao |
ICASSP | 2 |
| 2022 | Global Normalization for Streaming Speech Recognition in a Modular FrameworkabstractWe introduce the Globally Normalized Autoregressive Transducer (GNAT) for addressing the label bias problem in streaming speech recognition. Our solution admits a tractable exact computation of the denominator for the sequence-level normalization. Through theoretical and empirical results, we demonstrate that by switching to a globally normalized model, the word error rate gap between streaming and non-streaming speech-recognition models can be greatly reduced (by more than 50% on the Librispeech dataset). This model is developed in a modular framework which encompasses all the common neural speech recognition models. The modularity of this framework enables controlled comparison of modelling choices and creation of new models. A JAX implementation of our models has been open sourced. Ehsan Variani, Michael Riley 0001, David Rybach, Matt Shannon, Cyril Allauzen |
NeurIPS | 5 |
| 2020 | Location-Relative Attention Mechanisms for Robust Long-Form Speech SynthesisabstractDespite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in frequency for out-of-domain text. We show that these failures can be addressed using simple location-relative attention mechanisms that do away with content-based query/key comparisons. We compare two families of attention mechanisms: location-relative GMM-based mechanisms and additive energy-based mechanisms. We suggest simple modifications to GMM-based attention that allow it to align quickly and consistently during training, and introduce a new location-relative attention mechanism to the additive energy-based family, called Dynamic Convolution Attention (DCA). We compare the various mechanisms in terms of alignment speed and consistency during training, naturalness, and ability to generalize to long utterances, and conclude that GMM attention and DCA can generalize to very long utterances, while preserving naturalness for shorter, in-domain utterances. Eric Battenberg, R. J. Skerry-Ryan, Soroosh Mariooryad, Daisy Stanton, David Kao, Matt Shannon, Tom Bagby |
ICASSP | 6 |
| 2020 | Semi-Supervised Generative Modeling for Controllable Speech Synthesis
Raza Habib, Soroosh Mariooryad, Matt Shannon, Eric Battenberg, R. J. Skerry-Ryan, Daisy Stanton, David Kao, Tom Bagby |
ICLR | 3 |
| 2017 | Acoustic Modeling for Google Home
Bo Li 0028, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K. Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim 0001, Olivier Siohan, Mitch Weintraub, Erik McDermott, Richard Rose, Matt Shannon |
INTERSPEECH | 20 |
| 2017 | Recurrent Neural Aligner: An Encoder-Decoder Neural Network Model for Sequence to Sequence Mapping
Hasim Sak, Matt Shannon, Kanishka Rao, Françoise Beaufays |
INTERSPEECH | 2 |
| 2017 | Optimizing Expected Word Error Rate via Sampling for Speech RecognitionabstractState-level minimum Bayes risk (sMBR) training has become the de facto standard for sequence-level training of speech recognition acoustic models. It has an elegant formulation using the expectation semiring, and gives large improvements in word error rate (WER) over models trained solely using cross-entropy (CE) or connectionist temporal classification (CTC). sMBR training optimizes the expected number of frames at which the reference and hypothesized acoustic states differ. It may be preferable to optimize the expected WER, but WER does not interact well with the expectation semiring, and previous approaches based on computing expected WER exactly involve expanding the lattices used during training. In this paper we show how to perform optimization of the expected WER by sampling paths from the lattices used during conventional sMBR training. The gradient of the expected WER is itself an expectation, and so may be approximated using Monte Carlo sampling. We show experimentally that optimizing WER during acoustic model training gives 5% relative improvement in WER over a well-tuned sMBR baseline on a 2-channel query recognition task (Google Home). Matt Shannon |
INTERSPEECH | 1 |
| 2017 | Improved End-of-Query Detection for Streaming Speech Recognition
Matt Shannon, Gabor Simko, Shuo-Yiin Chang, Carolina Parada |
INTERSPEECH | 1 |
| 2014 | Measuring the perceptual effects of modelling assumptions in speech synthesis using stimuli constructed from repeated natural speechabstractAcoustic models used for statistical parametric speech synthe-sis typically incorporate many modelling assumptions. It is an open question to what extent these assumptions limit the natu-ralness of synthesised speech. To investigate this question, we recorded a speech corpus where each prompt was read aloud multiple times. By combining speech parameter trajectories ex-tracted from different repetitions, we were able to quantify the perceptual effects of certain commonly used modelling assump-tions. Subjective listening tests show that taking the source and filter parameters to be conditionally independent, or using di-agonal covariance matrices, significantly limits the naturalness that can be achieved. Our experimental results also demonstrate the shortcomings of mean-based parameter generation. Index terms: speech synthesis, acoustic modelling, stream in-dependence, diagonal covariance matrices, repeated speech 1. Gustav Eje Henter, Thomas Merritt, Matt Shannon, Catherine Mayo, Simon King 0001 |
INTERSPEECH | 3 |
| 2013 | Fast, low-artifact speech synthesis considering global varianceabstractSpeech parameter generation considering global variance (GV generation) is widely acknowledged to dramatically improve the quality of synthetic speech generated by HMM-based systems. However it is slower and has higher latency than the standard speech parameter generation algorithm. In addition it is known to produce artifacts, though existing approaches to prevent artifacts are effective. We present a simple new theoretical analysis of speech parameter generation considering global variance based on Lagrange multipliers. This analysis sheds light on one source of artifacts and suggests a way to reduce their occurrence. It also suggests an approximation to exact GV generation that allows fast, low latency synthesis. In a subjective evaluation our fast approximation shows no degradation in naturalness compared to conventional GV generation. Matt Shannon, William J. Byrne |
ICASSP | 1 |
| 2013 | Autoregressive Models for Statistical Parametric Speech SynthesisabstractWe propose using the autoregressive hidden Markov model (HMM) for speech synthesis. The autoregressive HMM uses the same model for parameter estimation and synthesis in a consistent way, in contrast to the standard approach to statistical parametric speech synthesis. It supports easy and efficient parameter estimation using expectation maximization, in contrast to the trajectory HMM. At the same time its similarities to the standard approach allow use of established high quality synthesis algorithms such as speech parameter generation considering global variance. The autoregressive HMM also supports a speech parameter generation algorithm not available for the standard approach or the trajectory HMM and which has particular advantages in the domain of real-time, low latency synthesis. We show how to do efficient parameter estimation and synthesis with the autoregressive HMM and look at some of the similarities and differences between the standard approach, the trajectory HMM and the autoregressive HMM. We compare the three approaches in subjective and objective evaluations. We also systematically investigate which choices of parameters such as autoregressive order and number of states are optimal for the autoregressive HMM. Matt Shannon, Heiga Zen, William J. Byrne |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | The Effect of Using Normalized Models in Statistical Speech SynthesisabstractThe standard approach to HMM-based speech synthesis is inconsistent in the enforcement of the deterministic constraints between static and dynamic features. The trajectory HMM and autoregressive HMM have been proposed as normalized models which rectify this inconsistency. This paper investigates the practical effects of using these normalized models, and examines the strengths and weaknesses of the different models as probabilistic models of speech. The most striking difference observed is that the standard approach greatly underestimates predictive variance. We argue that the normalized models have better predictive distributions than the standard approach, but that all the models we consider are still far from satisfactory probabilistic models of speech. We also present evidence that better intra-frame correlation modelling goes some way towards improving existing normalized models. Index terms: HMM-based speech synthesis, acoustic modelling, autoregressive HMM, trajectory HMM, normalization 1. Matt Shannon, Heiga Zen, William J. Byrne |
INTERSPEECH | 1 |
| 2010 | Autoregressive clustering for HMM speech synthesisabstractThe autoregressive HMM has been shown to provide efficient parameter estimation and high-quality synthesis, but in previous experiments decision trees derived from a non-autoregressive system were used. In this paper we investigate the use of autoregressive clustering for autoregressive HMM-based speech synthesis. We describe decision tree clustering for the autoregressive HMM and highlight differences to the standard clustering procedure. Subjective listening evaluation results suggest that autoregressive clustering improves the naturalness of the resulting speech. We find that the standard minimum description length (MDL) criterion for selecting model complexity is inappropriate for the autoregressive HMM. Investigating the effect of model complexity on naturalness, we find that a large degree of overfitting is tolerated without a substantial decrease in naturalness. Index terms: HMM-based speech synthesis, decision tree clustering, autoregressive HMM 1. Matt Shannon, William J. Byrne |
INTERSPEECH | 1 |
| 2009 | Autoregressive HMMs for speech synthesisabstractWe propose the autoregressive HMM for speech synthesis. We show that the autoregressive HMM supports efficient EM parameter estimation and that we can use established effective synthesis techniques such as synthesis considering global variance with minimal modification. The autoregressive HMM uses the same model for parameter estimation and synthesis in a consistent way, in contrast to the standard HMM synthesis framework, and supports easy and efficient parameter estimation, in contrast to the trajectory HMM. We find that the autoregressive HMM gives performance comparable to the standard HMM synthesis framework on a Blizzard Challenge-style naturalness evaluation. Matt Shannon, William J. Byrne |
INTERSPEECH | 1 |