Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Matt Shannon

dblp:38/9012 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Speech recognition and synthesis · 62% Trustworthy machine learning · 16% Learning paradigms · 12%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.612022
Global Normalization for Streaming Speech Recognition in a Modular Framework · NeurIPS 2022
Machine learning › Trustworthy machine learning › dataset bias
label bias
0.612022
Global Normalization for Streaming Speech Recognition in a Modular Framework · NeurIPS 2022
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › real-time speech recognition
streaming speech recognition
0.612022
Global Normalization for Streaming Speech Recognition in a Modular Framework · NeurIPS 2022
Natural language and speech › Speech recognition and synthesis › speech synthesis
controllable speech synthesis
0.412020
Semi-Supervised Generative Modeling for Controllable Speech Synthesis · ICLR 2020
Machine learning › Learning paradigms › semi-supervised learning
generative semi-supervised learning
0.412020
Semi-Supervised Generative Modeling for Controllable Speech Synthesis · ICLR 2020
Natural language and speech › Speech recognition and synthesis
speech synthesis
0.412020
Semi-Supervised Generative Modeling for Controllable Speech Synthesis · ICLR 2020
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › hidden markov model
autoregressive hidden markov model
0.212013
Autoregressive Models for Statistical Parametric Speech Synthesis · IEEE Trans. Speech Audio Process. 2013
Natural language and speech › Speech recognition and synthesis › speech synthesis
statistical parametric speech synthesis
0.212013
Autoregressive Models for Statistical Parametric Speech Synthesis · IEEE Trans. Speech Audio Process. 2013

Methods — techniques the papers use, named apart from their topics

globally normalized autoregressive transducer · 0.6semi-supervised learning · 0.4generative modeling · 0.4global variance parameter generation · 0.2expectation-maximization · 0.2
YearPublicationVenuePosition
2025 Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
abstract
Eric Battenberg, RJ Skerry-Ryan, Daisy Stanton, Soroosh Mariooryad, Matt Shannon, Julian Salazar, David Teh-Hwa Kao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Eric Battenberg, R. J. Skerry-Ryan, Daisy Stanton, Soroosh Mariooryad, Matt Shannon, Julian Salazar, David Kao
NAACL (Long Papers)5
2022 Speaker Generation
abstract
This work explores the task of synthesizing speech in non-existent human-sounding voices. We call this task "speaker generation", and present TacoSpawn, a system that performs competitively at this task. TacoSpawn is a recurrent attention-based text-to-speech model that learns a distribution over a speaker embedding space, which enables sampling of novel and diverse speakers. Our method is easy to implement, and does not require transfer learning from speaker ID systems. We present objective and subjective metrics for evaluating performance on this task, and demonstrate that our proposed objective metrics correlate with human perception of speaker similarity. Audio samples are available on our demo page1.
Daisy Stanton, Matt Shannon, Soroosh Mariooryad, R. J. Skerry-Ryan, Eric Battenberg, Tom Bagby, David Kao
ICASSP2
2022 Global Normalization for Streaming Speech Recognition in a Modular Framework
abstract
We introduce the Globally Normalized Autoregressive Transducer (GNAT) for addressing the label bias problem in streaming speech recognition. Our solution admits a tractable exact computation of the denominator for the sequence-level normalization. Through theoretical and empirical results, we demonstrate that by switching to a globally normalized model, the word error rate gap between streaming and non-streaming speech-recognition models can be greatly reduced (by more than 50% on the Librispeech dataset). This model is developed in a modular framework which encompasses all the common neural speech recognition models. The modularity of this framework enables controlled comparison of modelling choices and creation of new models. A JAX implementation of our models has been open sourced.
Ehsan Variani, Michael Riley 0001, David Rybach, Matt Shannon, Cyril Allauzen
NeurIPS5
2020 Location-Relative Attention Mechanisms for Robust Long-Form Speech Synthesis
abstract
Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in frequency for out-of-domain text. We show that these failures can be addressed using simple location-relative attention mechanisms that do away with content-based query/key comparisons. We compare two families of attention mechanisms: location-relative GMM-based mechanisms and additive energy-based mechanisms. We suggest simple modifications to GMM-based attention that allow it to align quickly and consistently during training, and introduce a new location-relative attention mechanism to the additive energy-based family, called Dynamic Convolution Attention (DCA). We compare the various mechanisms in terms of alignment speed and consistency during training, naturalness, and ability to generalize to long utterances, and conclude that GMM attention and DCA can generalize to very long utterances, while preserving naturalness for shorter, in-domain utterances.
Eric Battenberg, R. J. Skerry-Ryan, Soroosh Mariooryad, Daisy Stanton, David Kao, Matt Shannon, Tom Bagby
ICASSP6
2020 Semi-Supervised Generative Modeling for Controllable Speech Synthesis
Raza Habib, Soroosh Mariooryad, Matt Shannon, Eric Battenberg, R. J. Skerry-Ryan, Daisy Stanton, David Kao, Tom Bagby
ICLR3
2017 Acoustic Modeling for Google Home
Bo Li 0028, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K. Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim 0001, Olivier Siohan, Mitch Weintraub, Erik McDermott, Richard Rose, Matt Shannon
INTERSPEECH20
2017 Recurrent Neural Aligner: An Encoder-Decoder Neural Network Model for Sequence to Sequence Mapping
Hasim Sak, Matt Shannon, Kanishka Rao, Françoise Beaufays
INTERSPEECH2
2017 Optimizing Expected Word Error Rate via Sampling for Speech Recognition
abstract
State-level minimum Bayes risk (sMBR) training has become the de facto standard for sequence-level training of speech recognition acoustic models. It has an elegant formulation using the expectation semiring, and gives large improvements in word error rate (WER) over models trained solely using cross-entropy (CE) or connectionist temporal classification (CTC). sMBR training optimizes the expected number of frames at which the reference and hypothesized acoustic states differ. It may be preferable to optimize the expected WER, but WER does not interact well with the expectation semiring, and previous approaches based on computing expected WER exactly involve expanding the lattices used during training. In this paper we show how to perform optimization of the expected WER by sampling paths from the lattices used during conventional sMBR training. The gradient of the expected WER is itself an expectation, and so may be approximated using Monte Carlo sampling. We show experimentally that optimizing WER during acoustic model training gives 5% relative improvement in WER over a well-tuned sMBR baseline on a 2-channel query recognition task (Google Home).
Matt Shannon
INTERSPEECH1
2017 Improved End-of-Query Detection for Streaming Speech Recognition
Matt Shannon, Gabor Simko, Shuo-Yiin Chang, Carolina Parada
INTERSPEECH1
2014 Measuring the perceptual effects of modelling assumptions in speech synthesis using stimuli constructed from repeated natural speech
abstract
Acoustic models used for statistical parametric speech synthe-sis typically incorporate many modelling assumptions. It is an open question to what extent these assumptions limit the natu-ralness of synthesised speech. To investigate this question, we recorded a speech corpus where each prompt was read aloud multiple times. By combining speech parameter trajectories ex-tracted from different repetitions, we were able to quantify the perceptual effects of certain commonly used modelling assump-tions. Subjective listening tests show that taking the source and filter parameters to be conditionally independent, or using di-agonal covariance matrices, significantly limits the naturalness that can be achieved. Our experimental results also demonstrate the shortcomings of mean-based parameter generation. Index terms: speech synthesis, acoustic modelling, stream in-dependence, diagonal covariance matrices, repeated speech 1.
Gustav Eje Henter, Thomas Merritt, Matt Shannon, Catherine Mayo, Simon King 0001
INTERSPEECH3
2013 Fast, low-artifact speech synthesis considering global variance
abstract
Speech parameter generation considering global variance (GV generation) is widely acknowledged to dramatically improve the quality of synthetic speech generated by HMM-based systems. However it is slower and has higher latency than the standard speech parameter generation algorithm. In addition it is known to produce artifacts, though existing approaches to prevent artifacts are effective. We present a simple new theoretical analysis of speech parameter generation considering global variance based on Lagrange multipliers. This analysis sheds light on one source of artifacts and suggests a way to reduce their occurrence. It also suggests an approximation to exact GV generation that allows fast, low latency synthesis. In a subjective evaluation our fast approximation shows no degradation in naturalness compared to conventional GV generation.
Matt Shannon, William J. Byrne
ICASSP1
2013 Autoregressive Models for Statistical Parametric Speech Synthesis
abstract
We propose using the autoregressive hidden Markov model (HMM) for speech synthesis. The autoregressive HMM uses the same model for parameter estimation and synthesis in a consistent way, in contrast to the standard approach to statistical parametric speech synthesis. It supports easy and efficient parameter estimation using expectation maximization, in contrast to the trajectory HMM. At the same time its similarities to the standard approach allow use of established high quality synthesis algorithms such as speech parameter generation considering global variance. The autoregressive HMM also supports a speech parameter generation algorithm not available for the standard approach or the trajectory HMM and which has particular advantages in the domain of real-time, low latency synthesis. We show how to do efficient parameter estimation and synthesis with the autoregressive HMM and look at some of the similarities and differences between the standard approach, the trajectory HMM and the autoregressive HMM. We compare the three approaches in subjective and objective evaluations. We also systematically investigate which choices of parameters such as autoregressive order and number of states are optimal for the autoregressive HMM.
Matt Shannon, Heiga Zen, William J. Byrne
IEEE Trans. Speech Audio Process.1
2011 The Effect of Using Normalized Models in Statistical Speech Synthesis
abstract
The standard approach to HMM-based speech synthesis is inconsistent in the enforcement of the deterministic constraints between static and dynamic features. The trajectory HMM and autoregressive HMM have been proposed as normalized models which rectify this inconsistency. This paper investigates the practical effects of using these normalized models, and examines the strengths and weaknesses of the different models as probabilistic models of speech. The most striking difference observed is that the standard approach greatly underestimates predictive variance. We argue that the normalized models have better predictive distributions than the standard approach, but that all the models we consider are still far from satisfactory probabilistic models of speech. We also present evidence that better intra-frame correlation modelling goes some way towards improving existing normalized models. Index terms: HMM-based speech synthesis, acoustic modelling, autoregressive HMM, trajectory HMM, normalization 1.
Matt Shannon, Heiga Zen, William J. Byrne
INTERSPEECH1
2010 Autoregressive clustering for HMM speech synthesis
abstract
The autoregressive HMM has been shown to provide efficient parameter estimation and high-quality synthesis, but in previous experiments decision trees derived from a non-autoregressive system were used. In this paper we investigate the use of autoregressive clustering for autoregressive HMM-based speech synthesis. We describe decision tree clustering for the autoregressive HMM and highlight differences to the standard clustering procedure. Subjective listening evaluation results suggest that autoregressive clustering improves the naturalness of the resulting speech. We find that the standard minimum description length (MDL) criterion for selecting model complexity is inappropriate for the autoregressive HMM. Investigating the effect of model complexity on naturalness, we find that a large degree of overfitting is tolerated without a substantial decrease in naturalness. Index terms: HMM-based speech synthesis, decision tree clustering, autoregressive HMM 1.
Matt Shannon, William J. Byrne
INTERSPEECH1
2009 Autoregressive HMMs for speech synthesis
abstract
We propose the autoregressive HMM for speech synthesis. We show that the autoregressive HMM supports efficient EM parameter estimation and that we can use established effective synthesis techniques such as synthesis considering global variance with minimal modification. The autoregressive HMM uses the same model for parameter estimation and synthesis in a consistent way, in contrast to the standard HMM synthesis framework, and supports easy and efficient parameter estimation, in contrast to the trajectory HMM. We find that the autoregressive HMM gives performance comparable to the standard HMM synthesis framework on a Blizzard Challenge-style naturalness evaluation.
Matt Shannon, William J. Byrne
INTERSPEECH1