Matthew P. Aylett

dblp:28/4033 · also Matthew Peter Aylett · DBLP profile ↗
← Back
42ranked-venue papers
17as first author
10since 2021 · last 2026
0000-0001-7057-0525ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 14 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 11 first-authorHuman-computer interaction and ubiquitous computing · 17 · 5 first-author · 8 since 2021
YearPublicationVenuePosition
2026 ARTICULATE: Science in your Own Language
abstract
The ARTICULATE project is an ambitious and interdisciplinary initiative funded by the CHIST-ERA call 2025. Its vision is to revolutionize science education and democratize scientific knowledge beyond academia and English-speaking audiences through the integration of AI with self-regulated learning. The aim is to translate science not just across language but across language style, to create engaging spoken digital experiences. We present an introduction to this project, an overview of the consortium and research approach, and a number of expected impacts.
Yolanda Vazquez-Alvarez, Matthew P. Aylett, Benjamin R. Cowan, Justin Edwards, Sanna Järvelä, Ioannis Konstas, Madeleine Steeds
EAMT (2)2
2026 "What do you expect? You're part of the internet": Analyzing Western Celebrities' Experiences as Usees of Deepfake Technology
abstract
Deepfake technology is often used to create non-consensual synthetic intimate imagery (NSII), mainly of celebrity women. Through Critical Discursive Psychological analysis, we ask; i) how celebrities construct being targeted by deepfakes and ii) how they navigate infrastructural and social obstacles when seeking recourse. In this paper, we adopt Baumer’s concept of “Usees” (stakeholders who are non-consenting, unaware and directly targeted by technology), to understand public statements made by eight celebrity women and one non-binary individual targeted with NSII. Celebrities describe harms of being non-consensually targeted by deepfakes and they describe various infrastructural/social factors) which hinder activism and recourse. This work has implications in recognizing the roles of infrastructures underlying deepfake abuse and the potential of human-computer interaction to improve existing recourses for NSII. We also contribute to understanding false victim-blaming rhetoric about deepfake abuse. Future work should involve interventions which challenge values and false beliefs which motivate NSII creation/dissemination.
John Twomey, Sarah Foley, Sarah Robinson, Michael Quayle, Matthew P. Aylett, Conor Linehan, Gillian Murphy
Int. J. Hum. Comput. Stud.5
2025 Lost in the Story: The Impact of Narrative with a Direction-Giving Robot
abstract
Sharing a story alongside an expository response is inherently human, often enhancing communication by adding personal details based on our unique experiences to what we say. When used in a task environment, narratives may be used to exploit measurable effects, such as on memory recall or interaction engagement. With the increasing presence of social robots in everyday environments, it remains unclear whether narrative communication from robots (e.g. “This picture shows a family who recently...”) instead of a factual description yields similar benefits to those observed in human-human interactions. In this paper, we develop and study a direction-giving robot, comparing three styles of navigation instruction: narrative with landmarks, landmarks only, and baseline without landmarks. We evaluate the effects of these conditions on recall, task success, and social acceptability factors (N=38) using a Furhat robot receptionist in a lab environment.
Bruce W. Wilson, Mei Yii Lim, Helen Hastie, Matthew P. Aylett
HAI4
2024 Developing Fictive Dialogs for a Classroom Language Learning Conversational Interface
abstract
The use of Educational Conversational Agents (ECA) offers benefits for language learning but needs to overcome challenges such as student distraction and lack of engagement. Our aim is to enhance learning outcomes by designing an ECA to serve as both teacher and peer. Integrating RASA, a conversational agent from Rapport, and Automatic Speech Recognition (ASR) from Deepgram, the ECA engages in interactive sessions with a human confederate and student participant. During teaching, speech synthesis is used to read out an educational text, while interruptions from the confederate trigger semi-scripted dialogues, fostering students’ engagement and understanding. Evaluation includes conditions with and without dialogue elements, assessed through test questionnaires and Godspeed and NASA-TLX surveys to explore user satisfaction. Results show that more dialogues between the agent and the human confederate lead to better learning outcomes for students, but direct conversations between the agent and the student may lower the performance.
Matthew P. Aylett, Xuanchen Li, Chenxi Meng, Zewen Qu, Sirui Wang 0006, Zechen Yang
HAI1
2024 Follow the Yellow or Red Brick Road? Investigating the Impact of Narratives in a Guided Navigation Task
abstract
Sharing casual stories in conversations is a natural human behaviour that adds personal and unique details, while also offering measurable benefits such as improved memory recall and more positive social human-human interactions. However, as social robots become more common in everyday environments, it is unclear whether these stories, or narratives, have similar effects on human-robot interactions with embodied agents.
Bruce W. Wilson, Shivaanee Eswaran, Omar Riyaz, Francisco Javier Chiyah Garcia, Matthew P. Aylett
HAI5
2024 Case study in choosing a graphical character to support reminiscence therapy for those living with dementia
abstract
The present study describes the process employed to identify a suitable intelligent virtual agent (IVA) for the AMPER App supporting reminiscence therapy for those living with dementia through the use of a facilitating agent. This included three distinct phases: 1) co-creation with project stakeholders and identification of a set of desirable IVA traits; 2) a blind internal team rating process to select a subset from available IVAs; 3) a character survey with healthy older adults to select a final 4 IVAs (2 male, 2 female). We analyse the results, assess inter-subject agreement, and suggest guidelines for IVA selection.
Matthew P. Aylett, Katerina Pappa, Mei Yii Lim, Ruth Aylett, Bruce W. Wilson, Mario A. Parra
IVA1
2024 Demonstration of the AMPER System for Individuals with Alzheimer's Disease
abstract
We present a demonstration of the AMPER system: an Android application that guides individuals with Alzheimer’s disease and their carers through reminiscence therapy with the use of a virtual agent. The application supports a novel method of reminiscence story selection, using material metadata, user data, and a spreading activation algorithm. This aims to present relevant material to the user, in an order that attempts to mimic a human autobiographical memory.
Bruce W. Wilson, Mei Yii Lim, Katerina Pappa, Matthew P. Aylett, Mario A. Parra, Ruth Aylett
IVA4
2023 Creating Inclusive Voices for the 21st Century: A Non-Binary Text-to-Speech for Conversational Assistants
abstract
As voice assistant usage continues to grow, their homogeneity becomes even more problematic with the UNESCO report, “I’d Blush if I could” showing that designing only feminine voice assistants encourages negative behavior, both with virtual assistants and with real people [3]. While masculine text-to-speech (TTS) voices exist, ones that cover the full range of gender presentations, such as non-binary or gender-ambiguous voices are largely missing. In this paper, we present a method of creating a non-binary TTS voice and an example voice, Sam, created with input from the non-binary and transgender communities. We have open-sourced the resulting voice, along with the process and data used to create it. Finally, we present results from a large-scale survey showing that non-binary individuals are more likely to prefer a non-binary voice assistant compared to cisgendered individuals and discuss differences across age and gender.
Andreea Danielescu 0001, Sharone Horowit-Hendler, Alexandria Pabst, Kenneth Michael Stewart, Eric M. Gallo, Matthew P. Aylett
CHI6
2023 Why is my Agent so Slow? Deploying Human-Like Conversational Turn-Taking
abstract
The emphasis on one-to-one speak/wait spoken conversational interaction with intelligent agents leads to long pauses between conversational turns, undermines the flow and naturalness of the interaction, and undermines the user experience. Despite ground breaking advances in the area of generating and understanding natural language with techniques such as LLMs, conversational interaction has remained relatively overlooked. In this workshop we will discuss and review the challenges, recent work and potential impact of improving conversational interaction with artificial systems. We hope to share experiences of poor human/system interaction, best practices with third party tools, and generate design guidance for the community.
Matthew P. Aylett, Éva Székely, Donald McMillan, Gabriel Skantze, Marta Romeo, Joel E. Fischer, Gisela Reyes-Cruz
HAI1
2023 Synthesising Personality with Neural Speech Synthesis
abstract
Matching the personality of conversational agents to the personality of the user can significantly improve the user experience, with many successful examples in text-based chatbots.It is also important for a voice-based system to be able to alter the personality of the speech as perceived by the users.In this pilot study, fifteen voices were rated using Big Five personality traits.Five content-neutral sentences were chosen for the listening tests.The audio data, together with two rated traits (Extroversion and Agreeableness), were used to train a neural speech synthesiser based on one male and one female voices.The effect of altering the personality trait features was evaluated by a second listening test.Both perceived extroversion and agreeableness in the synthetic voices were affected significantly.The controllable range was limited due to a lack of variance in the source audio data.The perceived personality traits correlated with each other and with the naturalness of the speech.
Shilin Gao, Matthew P. Aylett, David A. Braude, Catherine Lai
SIGDIAL2
2020 Speech Synthesis for the Generation of Artificial Personality
abstract
A synthetic voice personifies the system using it. In this work we examine the impact text content, voice quality and synthesis system have on the perceived personality of two synthetic voices. Subjects rated synthetic utterances based on the Big-Five personality traits and naturalness. The naturalness rating of synthesis output did not correlate significantly with any Big-Five characteristic except for a marginal correlation with openness. Although text content is dominant in personality judgments, results showed that voice quality change implemented using a unit selection synthesis system significantly affected the perception of the Big-Five, for example tense voice being associated with being disagreeable and lax voice with lower conscientiousness. In addition a comparison between a parametric implementation and unit selection implementation of the same voices showed that parametric voices were rated as significantly less neurotic than both the text alone and the unit selection system, while the unit selection was rated as more open than both the text alone and the parametric system. The results have implications for synthesis voice and system type selection for applications such as personal assistants and embodied conversational agents where developing an emotional relationship with the user, or developing a branding experience is important.
Matthew P. Aylett, Alessandro Vinciarelli, Mirjam Wester
IEEE Trans. Affect. Comput.1
2019 All Together Now: The Living Audio Dataset
David A. Braude, Matthew P. Aylett, Caoimhín Laoide-Kemp, Simone Ashby, Kristen M. Scott, Brian Ó Raghallaigh, Anna Braudo, Alex Brouwer, Adriana Cornelia Stan
INTERSPEECH2
2019 The State of Speech in HCI: Trends, Themes and Challenges
abstract
Abstract Speech interfaces are growing in popularity. Through a review of 99 research papers this work maps the trends, themes, findings and methods of empirical research on speech interfaces in the field of human–computer interaction (HCI). We find that studies are usability/theory-focused or explore wider system experiences, evaluating Wizard of Oz, prototypes or developed systems. Measuring task and interaction was common, as was using self-report questionnaires to measure concepts like usability and user attitudes. A thematic analysis of the research found that speech HCI work focuses on nine key topics: system speech production, design insight, modality comparison, experiences with interactive voice response systems, assistive technology and accessibility, user speech production, using speech technology for development, peoples’ experiences with intelligent personal assistants and how user memory affects speech interface interaction. From these insights we identify gaps and challenges in speech research, notably taking into account technological advancements, the need to develop theories of speech interface interaction, grow critical mass in this domain, increase design work and expand research from single to multiple user interaction contexts so as to reflect current use contexts. We also highlight the need to improve measure reliability, validity and consistency, in the wild deployment and reduce barriers to building fully functional speech interfaces for research. RESEARCH HIGHLIGHTS Most papers focused on usability/theory-based or wider system experience research with a focus on Wizard of Oz and developed systems Questionnaires on usability and user attitudes often used but few were reliable or validated Thematic analysis showed nine primary research topics Challenges identified in theoretical approaches and design guidelines, engaging with technological advances, multiple user and in the wild contexts, critical research mass and barriers to building speech interfaces
Leigh Clark, Philip R. Doyle, Diego Garaialde, Emer Gilmartin, Stephan Schlögl, Jens Edlund, Matthew P. Aylett, João P. Cabral, Cosmin Munteanu, Justin Edwards, Benjamin R. Cowan
Interact. Comput.7
2018 A life story in three parts: the use of triptychs to make sense of personal digital data
abstract
Many social media platforms support the curation of personal digital data, and, more recently, the use of that data for review and reflection. We explored the process of reflection by asking users to create a meaningful ‘triptych’ of photographs drawn from their Facebook accounts. In a first study, we asked participants to manually trawl their own accounts and select three relevant images, which we then framed and used as an interview probe. In a second study, we designed an automated triptych generation system and assessed participants’ experiences of using this system. We conducted qualitative analyses of participant interviews from both studies. Consistent with other ‘slow technology’ work, we found the act of creating a physical artefact from social media data gave that data new meaning, albeit with notable differences between manual versus automatically generated triptychs. We conclude by discussing possible improvements to the design of the automated triptych system.
Lisa Thomas 0001, Elaine Farrow, Matthew P. Aylett, Pamela Briggs
Pers. Ubiquitous Comput.3
2017 Bot or not: exploring the fine line between cyber and human identity
abstract
Speech technology is rapidly entering the everyday through the large scale commercial impact of systems such as Apple Siri and Amazon Echo. Meanwhile technology that allows voice cloning, voice modification, speech recognition, speech analytics and expressive speech synthesis has changed dramatically over recent years. The demonstration, described in this paper, is an educational tool in the form of an online quiz called `Bot or Not'. Using the quiz we have gathered impressions of what people realise is possible with current speech synthesis technology. The opinions of various groups regarding the synthesis of famous voices, sounding like a robot, and the difference between synthesis and voice modification were collected.
Mirjam Wester, Matthew P. Aylett, David A. Braude
ICMI2
2017 Beyond the Listening Test: An Interactive Approach to TTS Evaluation
abstract
Traditionally, subjective text-To-speech (TTS) evaluation is performed through audio-only listening tests, where participants evaluate unrelated, context-free utterances. The ecological validity of ...
Joseph Mendelson, Matthew P. Aylett
INTERSPEECH2
2017 Real-Time Reactive Speech Synthesis: Incorporating Interruptions
Mirjam Wester, David A. Braude, Blaise Potard, Matthew P. Aylett, Francesca Shaw
INTERSPEECH4
2016 Ask Alice: an artificial retrieval of information agent
abstract
We present a demonstration of the ARIA framework, a modular approach for rapid development of virtual humans for information retrieval that have linguistic, emotional, and social skills and a strong personality. We demonstrate the framework's capabilities in a scenario where `Alice in Wonderland', a popular English literature book, is embodied by a virtual human representing Alice. The user can engage in an information exchange dialogue, where Alice acts as the expert on the book, and the user as an interested novice. Besides speech recognition, sophisticated audio-visual behaviour analysis is used to inform the core agent dialogue module about the user's state and intentions, so that it can go beyond simple chat-bot dialogue. The behaviour generation module features a unique new capability of being able to deal gracefully with interruptions of the agent.
Michel F. Valstar, Tobias Baur 0001, Angelo Cafaro, Alexandru Ghitulescu, Blaise Potard, Johannes Wagner 0001, Elisabeth André, Laurent Durieu, Matthew P. Aylett, Soumia Dermouche, Catherine Pelachaud, Eduardo Coutinho, Björn W. Schuller, Yue Zhang 0014, Dirk Heylen, Mariët Theune, Jelte van Waterschoot
ICMI9
2016 Idlak Tangle: An Open Source Kaldi Based Parametric Speech Synthesiser Based on DNN
Blaise Potard, Matthew P. Aylett, David A. Baude, Petr Motlícek
INTERSPEECH2
2016 Cross Modal Evaluation of High Quality Emotional Speech Synthesis with the Virtual Human Toolkit
Blaise Potard, Matthew P. Aylett, David A. Baude
IVA2
2016 Designing Interactions with Multilevel Auditory Displays in Mobile Audio-Augmented Reality
abstract
Auditory interfaces offer a solution to the problem of effective eyes-free mobile interactions. In this article, we investigate the use of multilevel auditory displays to enable eyes-free mobile interaction with indoor location-based information in non-guided audio-augmented environments. A top-level exocentric sonification layer advertises information in a gallery-like space. A secondary interactive layer is used to evaluate three different conditions that varied in the presentation (sequentialversussimultaneous) and spatialisation (non-spatialisedversusegocentric/exocentric spatialisation) of multiple auditory sources. Our findings show that (1) participants spent significantly more time interacting with spatialised displays; (2) using the same design for primary and interactive secondary display (simultaneous exocentric) showed a negative impact on the user experience, an increase in workload and substantially increased participant movement; and (3) the other spatial interactive secondary display designs (simultaneous egocentric, sequential egocentric, and sequential exocentric) showed an increase in time spent stationary but no negative impact on the user experience, suggesting a more exploratory experience. A follow-up qualitative and quantitative analysis of user behaviour support these conclusions. These results provide practical guidelines for designing effective eyes-free interactions for far richer auditory soundscapes.
Yolanda Vazquez-Alvarez, Matthew P. Aylett, Stephen A. Brewster, Rocio von Jungenfeld, Antti Virolainen
ACM Trans. Comput. Hum. Interact.2
2015 Generating Narratives from Personal Digital Data: Using Sentiment, Themes, and Named Entities to Construct Stories
Elaine Farrow, Thomas Dickinson, Matthew P. Aylett
INTERACT (4)3
2015 Artificial personality and disfluency
abstract
The focus of this paper is artificial voices with different person-alities. Previous studies have shown links between an individ-ual’s use of disfluencies in their speech and their perceived per-sonality. Here, filled pauses (uh and um) and discourse markers (like, you know, I mean) have been included in synthetic speech as a way of creating an artificial voice with different personali-ties. We discuss the automatic insertion of filled pauses and dis-course markers (i.e., fillers) into otherwise fluent texts. The au-tomatic system is compared to a ground truth of human “acted” filler insertion. Perceived personality (as defined by the big five personality dimensions) of the synthetic speech is assessed by means of a standardised questionnaire. Synthesis without fillers is compared to synthesis with either spontaneous or synthetic fillers. Our findings explore how the inclusion of disfluencies influences the way in which subjects rate the perceived person-ality of an artificial voice. Index Terms: artificial personality, TTS, disfluency 1.
Mirjam Wester, Matthew P. Aylett, Marcus Tomalin, Rasmus Dall
INTERSPEECH2
2014 A flexible front-end for HTS
abstract
Parametric speech synthesis techniques depend on full context acoustic models generated by language front-ends, which anal-yse linguistic and phonetic structure. HTS, the leading paramet-ric synthesis system, can use a number of different front-ends to generate full context models for synthesis and training. In this paper we explore the use of a new text processing front-end that has been added to the speech recognition toolkit Kaldi as part of an ongoing project to produce a new parametric speech synthesis system, Idlak. The use of XML specification files, a modular design, and modern coding and testing approaches, make the Idlak front-end ideal for adding, altering and experi-menting with the contexts used in full context acoustic models. The Idlak front-end was evaluated against the standard Festival front-end in the HTS system. Results from the Idlak front-end compare well with the more mature Festival front-end (Idlak-2.83 MOS vs Festival- 2.85 MOS), although a slight reduction in naturalness perceived by non-native English speakers can be attributed to Festival’s insertion of non-punctuated pauses. Index Terms: speech synthesis, text processing, parametric synthesis, Kaldi, Idlak
Matthew P. Aylett, Rasmus Dall, Arnab Ghoshal, Gustav Eje Henter, Thomas Merritt
INTERSPEECH1
2014 Phonetic feature extraction for context-sensitive glottal source processing
John Kane 0002, Matthew P. Aylett, Irena Yanushevskaya, Christer Gobl
Speech Commun.2
2013 Speaker and language independent voice quality classification applied to unlabelled corpora of expressive speech
abstract
Voice quality plays a pivotal role in speech style variation. Therefore, control and analysis of voice quality is critical for many areas of speech technology. Until now, most work has focused on small purpose built corpora. In this paper we apply state-of-the-art voice quality analysis to large speech corpora built for expressive speech synthesis. A fuzzy-input fuzzy-output support vector machine classifier is trained and validated using features extracted from these corpora. We then apply this classifier to freely available audiobook data and demonstrate a clustering of the voice qualities that approximates the performance of human perceptual ratings. The ability to detect voice quality variation in these widely available unlabelled audiobook corpora means that the proposed method may be used as a valuable resource in expressive speech synthesis.
John Kane 0002, Stefan Scherer, Matthew P. Aylett, Louis-Philippe Morency, Christer Gobl
ICASSP3
2012 Proper Name Splicing in Computer Games with TTS
abstract
Building high quality synthesis systems with open domain vocabulary and a small audio database is a challenging problem, even when the targeted application is well constrained. Monophone unit concatenation (as opposed to diphone) is an approach that can compensate for the poor unit coverage that a small database implies. However, joining at phone boundaries is a delicate task that requires accurate targeting. In this paper, we present an automatically trained targeting system based on the parametric synthesiser HTS, and compare it to a concatenative monophone system and a baseline concatenative diphone system. We apply a novel evaluation methodology which includes a qualitative component, and allows for fast incremental development of synthesis systems. Preliminary results show that although the hybrid system performed significantly more poorly on out of database items, it is less affected by segmentation errors than the monophone system. Index Terms: hybrid speech synthesis, unit selection, evaluation of TTS systems
Blaise Potard, Matthew P. Aylett, Christopher J. Pidcock
INTERSPEECH2
2012 Synthesising and Evaluating Cross-Modal Emotional Ambiguity in Virtual Agents
Matthew P. Aylett, Blaise Potard
IVA1
2011 The Romanian speech synthesis (RSS) corpus: Building a high quality HMM-based speech synthesis system using a high sampling rate
Adriana Cornelia Stan, Junichi Yamagishi, Simon King 0001, Matthew P. Aylett
Speech Commun.4
2009 Speech synthesis without a phone inventory
abstract
In speech synthesis the unit inventory is decided using phonological and phonetic expertise. This process is resource intensive and potentially sub-optimal. In this paper we investigate how acoustic clustering, together with lexicon constraints, can be used to build a self-organised inventory. Six English speech synthesis systems were built using two frameworks, unit selection and parametric HTS for three inventory conditions: 1) a traditional phone set, 2) a system using orthographic units, and 3) a self-organised inventory. A listening test showed a strong preference for the classic system, and for the orthographic system over the self-organised system. Results also varied by letter to sound complexity and database coverage. This suggests the self-organised approach failed to generalise pronunciation as well as introducing noise above and beyond that caused by orthographic sound mismatch. Index Terms: speech synthesis, unit selection, parametric synthesis, phone inventory, orthographic synthesis
Matthew P. Aylett, Simon King 0001, Junichi Yamagishi
INTERSPEECH1
2007 The CereVoice Characterful Speech Synthesiser SDK
Matthew P. Aylett, Christopher J. Pidcock
IVA1
2006 Detecting High Level Dialog Structure Without Lexical Information
abstract
The potentially enormous audio resources now available to both organizations, and on the Internet, present a serious challenge to audio browsing technology. In this paper we outline a set of techniques that can be used to determine high level dialog structure without the requirement of resource intensive automatic speech recognition (ASR). Using syllable finding algorithms based on band pass energy together with prosodic feature extraction, we show that a sub-lexical approach to prosodic analysis can outperform results based on ASR and even those based on a word alignment which requires a complete transcription. We consider how these techniques could be integrated into ASR technology and suggest a framework for extending this type of sub-lexical prosodic analysis
Matthew P. Aylett
ICASSP (1)1
2005 Synthesising hyperarticulation in unit selection TTS
abstract
Within speech synthesis we often wish to give extra focus to words which carry important information, such as names, dates and amounts. In this paper we look carefully at cost functions that can be used to bias unit selection in favour of hyper-articulated speech in order to give this impression of focus. Hyper-articulated speech tends to be accented, emphatic and requires more articulatory effort. We apply two cost functions to try to force the selection of hyper-articulated speech. The first operates on the duration of units in the unit selection database, the second on the language redundancy (word trigram predictability) of the word containing the unit. We estimate their relative importance in selecting hyperarticulated speech in unit selection speech synthesis. A listening test was carried out where these cost functions were applied to one random content word in a haskins anomalous sentence. Listeners were asked to select the two clearest and most focused words from the sentence. The duration increasing cost function was significantly related to an increase in perceived prominence whereas low redundancy, and a combination of both approaches did not produce significant results. Thus, although a significant correlation exists between the average duration and redundancy of diphones and perceived prominence, such a correlation was not smoothly translated into error free method for altering such perceived prominence.
Matthew P. Aylett
INTERSPEECH1
2003 My voice, your prosody: sharing a speaker specific prosody model across speakers in unit selection TTS
abstract
Data sparsity is a major problem for data driven prosodic models. Being able to share prosodic data across speakers is a potential solution to this problem. This paper explores this potential solution by addressing two questions: 1) Does a larger less sparse model from a different speaker produce more natural speech than a small sparse model built from the original speaker? 2)Does a different speaker's larger model generate more unit selection errors than a small sparse model built from the original speaker? A unit selection approach is used to produce a lazy learning model of three English RP speaker's f0 and durational parameters. Speaker 1 (the target speaker) had a much smaller database (approximately one quarter to one fifth the size) of the other two. Speaker 2 was a female speaker with frequent mid phrase rises. Speaker 3 was a male speaker with a similar f0 range to speaker 1 and with a measured prosodic style suitable for news and financial text. We apply the models created for speaker 2 (an inappropriate model) and speaker 3 (an appropriate model) to speaker 1 and compare the results. Three passages (of three to four sentences in length) from challenging prosodic genres (news report, poetry and personal email) were synthesised using the target speaker and each of the three models. The synthesised utterances were played to 15 native english subjects and rated using a 5 point MOS scale. In addition, 7 experienced speech engineers rated each word for errors on a three point scale: 1. Acceptable, 2. Poor, 3. Unacceptable. The results suggest that a large model from an appropriate speaker does not sound more natural or produce fewer errors than a smaller model generated from the individual speaker's own data. In addition it shows that an inappropriate model does produce both less natural and more errors in the speech. High variance in both subject and materials analysis suggest both tests are far from ideal and that evaluation techniques for both error rate and naturalness need to improve.
Matthew P. Aylett, Justin Fackrell, Peter Rutten
INTERSPEECH1
2002 Stochastic suprasegmentals: relationship between the spectral characteristics of vowels, redundancy and prosodic structure
abstract
Previous work has shown a relationship between syllabic duration, redundancy within speech, and prosodic structure [1].In addition, a spectral care of articulation measure of vowels in spontaneous speech has supported these duration results and suggest that care of articulation varies inversely with redundancy and conversely with prosodic prominence.However, these spectral measures remain inconclusive due to measurement difficulties.In this paper a simpler spectral measurement is presented as a metric of care of articulation and applied to three vowels from each corner of the vowel triangle.Prosodic and redundancy factors can predict up to 5% of the variance of this new measurement, supporting the inverse redundancy result.However, in contrast to syllabic duration, whether prosodic boundaries are controlled for or not, the predictive power of prosodic factors and redundancy factors remain relatively independent.This suggests that 1. phrase final syllables, despite lengthening, do not show increased care of articulation if measured spectrally, and 2. unlike duration, redundancy factors affect the spectral characteristics of vowels independently of prosodic structure.
Matthew P. Aylett
INTERSPEECH1
2002 A statistically motivated database pruning technique for unit selection synthesis
Peter Rutten, Matthew P. Aylett, Justin Fackrell
INTERSPEECH2
2001 Modelling care of articulation with HMMs is dangerous
abstract
Changes in care of articulation (COA) affect both the spectral and durational characteristics of speech. This can have severe repercussions on both the success of speech recognition, and the quality of speech synthesis. Although auto-segmentation has proven useful for measuring the durational effects of COA, an automatic spectral measurement has proven more problematic. In this paper, we will explore the use of the acoustic log likelihoods generated by HMM auto-segmentation as a measure of these changes in comparison with two phonetically motivated modeling systems based on vocalic F1/F2 values. When duration variation is controlled, the HMM output does not correlate with the human perception of vowel goodness, whereas, the phonetically motivated models do.
Matthew P. Aylett
INTERSPEECH1
2000 Stochastic suprasegmentals: relationships between redundancy, prosodic structure and care of articulation in spontaneous speech
abstract
Within spontaneous speech there are wide variations in the articulation of the same word by the same speaker. This paper explores two related factors which influence variation in articulation, prosodic structure and redundancy. We argue that the constraint of producing robust communication while efficiently expending articulatory effort leads to an inverse relationship between language redundancy and care of articulation. The inverse relationship improves robustness by spreading the information more evenly across the speech signal leading to a smoother signal redundancy profile. We argue that prosodic prominence is a linguistic means of achieving smooth signal redundancy. Prosodic prominence increases care of articulation and coincides with unpredictable sections of speech. By doing so, prosodic prominence leads to a smoother signal redundancy. Results confirm the strong relationship between prosodic prominence and care of articulation as well as an inverse relationship between language redundancy and care of articulation. In addition, when variation in prosodic boundaries is controlled for, language redundancy can predict up to 65% of the variance in raw syllabic duration. This is comparable with 64% predicted by prosodic prominence (accent, lexical stress and vowel type). Moreover most (62%) of this predictive power is shared. This suggests that, in English, prosodic structure is the means with which constraints caused by a robust signal requirement are expressed in spontaneous speech.
Matthew P. Aylett
INTERSPEECH1
1998 Building a statistical model of the vowel space for phoneticians
abstract
Vowel space data (A two dimensional F1/F2 plot) is of interest to phoneticians for the purpose of comparing different accents, languages, speaker styles and individual speakers. Current automatic methods used by speech technologists do not generally produce traditional vowel space models (See [6] for an overview); instead they tend to produce hyper dimensional code books covering the entire speakers speech stream. This makes it difficult to relate results generated by these methods to observations in laboratory phonetics. In order to address these problems a model was developed based on a mixture Gaussian density function fitted using expectation maximisation on F1/F2 data producing a probability distribution in F1/F2 space. Speech was pre-processed using voicing to automatically excerpt vowel data without any need for segmentation and a parametric fit algorithm [7] was applied to calculate likely vowel targets. The result was a clear visualisation of a speaker's vowel space requiring ...
Matthew P. Aylett
ICSLP1
1998 The automatic marking of prominence in spontaneous speech using duration and part of speech information
abstract
The work reported in this paper was the result of the need to label a large corpus of spontaneous, task-oriented dialogue with prosodic prominences. A computational model using only word duration, part of speech and a dictionary lookup of each word's canonical phonemic contents was trained against the results of a human coder marking prominence. Because word durations were normalised, it was possible to set a common threshold for all members of a form class above which the lexically stressed syllables were classed as prominent. The method used is presented and the relative importance of duration information, phonemic contents, syllabic context and part of speech information is explored. The automatic coder was validated against unseen material and achieved a 58% agreement with a human coder. Further investigation showed that three humans coders agreed no better with each other than each agreed with the computational model. Thus, although the automatic system did not conform very well t...
Matthew P. Aylett, Matthew Bull
ICSLP1
1998 Vowel quality in spontaneous speech: what makes a good vowel?
abstract
Clear speech is characterised by longer segmental durations and less target undershoot [9] which results in more extreme spectral features.This paper deals with the clarity of vowels produced in spontaneous speech in a large corpus of task-oriented dialogues.We present an automatic technique for measuring vowel clarity on the basis of a vowel's spectral characteristics.This technique was evaluated using a perceptual test.Subjects rated the 'goodness' of vowels with different spectral characteristics with controlled duration and amplitude and these results were compared with an automatic rating.Results indicated that although agreement between subjects and the automatic measurement was poor it was as poor as the agreement between subjects.On the basis of these results we address the following questions:1. Can subjects reliably judge the clarity of vowels excerpted from spontaneous speech without duration cues? 2. Can a statistical model [3] reliably predict the subjects' response to such vowels?
Matthew P. Aylett, Alice Turk
ICSLP1
1998 An analysis of the timing of turn-taking in a corpus of goal-oriented dialogue
abstract
This paper presents a context-based analysis of the intervals between different speakers ’ utterances in a corpus of taskoriented dialogue (the Human Communication Research Centre’s Map Task Corpus. See Anderson et al. 1991). In the analysis, we assessed the relationship between inter-speaker intervals and various contextual factors, such as the effects of eye contact, the presence of conversational game boundaries, the category of move in an utterance, and the degree of experience with the task in hand. The results of the analysis indicated that the main factors which gave rise to significant differences in inter-speaker intervals were those which related to decision-making and planning- the greater the amount of planning, the greater the inter-speaker interval. Differences between speakers were also found to be significant, although this effect did not necessarily interact with all other effects. These results provide unique and useful data for the improved effectiveness of dialogue systems. 1.
Matthew Bull, Matthew P. Aylett
ICSLP2