VLDB 2026 Research / reviewers in the wild / expert
Matthew P. Aylett
dblp:28/4033 · also Matthew Peter Aylett
· DBLP profile ↗
42ranked-venue papers
17as first author
10since 2021 · last 2026
0000-0001-7057-0525ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 14 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 11 first-authorHuman-computer interaction and ubiquitous computing · 17 · 5 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARTICULATE: Science in your Own LanguageabstractThe ARTICULATE project is an ambitious and interdisciplinary initiative funded by the CHIST-ERA call 2025. Its vision is to revolutionize science education and democratize scientific knowledge beyond academia and English-speaking audiences through the integration of AI with self-regulated learning. The aim is to translate science not just across language but across language style, to create engaging spoken digital experiences. We present an introduction to this project, an overview of the consortium and research approach, and a number of expected impacts. Yolanda Vazquez-Alvarez, Matthew P. Aylett, Benjamin R. Cowan, Justin Edwards, Sanna Järvelä, Ioannis Konstas, Madeleine Steeds |
EAMT (2) | 2 |
| 2026 | "What do you expect? You're part of the internet": Analyzing Western Celebrities' Experiences as Usees of Deepfake TechnologyabstractDeepfake technology is often used to create non-consensual synthetic intimate imagery (NSII), mainly of celebrity women. Through Critical Discursive Psychological analysis, we ask; i) how celebrities construct being targeted by deepfakes and ii) how they navigate infrastructural and social obstacles when seeking recourse. In this paper, we adopt Baumer’s concept of “Usees” (stakeholders who are non-consenting, unaware and directly targeted by technology), to understand public statements made by eight celebrity women and one non-binary individual targeted with NSII. Celebrities describe harms of being non-consensually targeted by deepfakes and they describe various infrastructural/social factors) which hinder activism and recourse. This work has implications in recognizing the roles of infrastructures underlying deepfake abuse and the potential of human-computer interaction to improve existing recourses for NSII. We also contribute to understanding false victim-blaming rhetoric about deepfake abuse. Future work should involve interventions which challenge values and false beliefs which motivate NSII creation/dissemination. John Twomey, Sarah Foley, Sarah Robinson, Michael Quayle, Matthew P. Aylett, Conor Linehan, Gillian Murphy |
Int. J. Hum. Comput. Stud. | 5 |
| 2025 | Lost in the Story: The Impact of Narrative with a Direction-Giving RobotabstractSharing a story alongside an expository response is inherently human, often enhancing communication by adding personal details based on our unique experiences to what we say. When used in a task environment, narratives may be used to exploit measurable effects, such as on memory recall or interaction engagement. With the increasing presence of social robots in everyday environments, it remains unclear whether narrative communication from robots (e.g. “This picture shows a family who recently...”) instead of a factual description yields similar benefits to those observed in human-human interactions. In this paper, we develop and study a direction-giving robot, comparing three styles of navigation instruction: narrative with landmarks, landmarks only, and baseline without landmarks. We evaluate the effects of these conditions on recall, task success, and social acceptability factors (N=38) using a Furhat robot receptionist in a lab environment. Bruce W. Wilson, Mei Yii Lim, Helen Hastie, Matthew P. Aylett |
HAI | 4 |
| 2024 | Developing Fictive Dialogs for a Classroom Language Learning Conversational InterfaceabstractThe use of Educational Conversational Agents (ECA) offers benefits for language learning but needs to overcome challenges such as student distraction and lack of engagement. Our aim is to enhance learning outcomes by designing an ECA to serve as both teacher and peer. Integrating RASA, a conversational agent from Rapport, and Automatic Speech Recognition (ASR) from Deepgram, the ECA engages in interactive sessions with a human confederate and student participant. During teaching, speech synthesis is used to read out an educational text, while interruptions from the confederate trigger semi-scripted dialogues, fostering students’ engagement and understanding. Evaluation includes conditions with and without dialogue elements, assessed through test questionnaires and Godspeed and NASA-TLX surveys to explore user satisfaction. Results show that more dialogues between the agent and the human confederate lead to better learning outcomes for students, but direct conversations between the agent and the student may lower the performance. Matthew P. Aylett, Xuanchen Li, Chenxi Meng, Zewen Qu, Sirui Wang 0006, Zechen Yang |
HAI | 1 |
| 2024 | Follow the Yellow or Red Brick Road? Investigating the Impact of Narratives in a Guided Navigation TaskabstractSharing casual stories in conversations is a natural human behaviour that adds personal and unique details, while also offering measurable benefits such as improved memory recall and more positive social human-human interactions. However, as social robots become more common in everyday environments, it is unclear whether these stories, or narratives, have similar effects on human-robot interactions with embodied agents. Bruce W. Wilson, Shivaanee Eswaran, Omar Riyaz, Francisco Javier Chiyah Garcia, Matthew P. Aylett |
HAI | 5 |
| 2024 | Case study in choosing a graphical character to support reminiscence therapy for those living with dementiaabstractThe present study describes the process employed to identify a suitable intelligent virtual agent (IVA) for the AMPER App supporting reminiscence therapy for those living with dementia through the use of a facilitating agent. This included three distinct phases: 1) co-creation with project stakeholders and identification of a set of desirable IVA traits; 2) a blind internal team rating process to select a subset from available IVAs; 3) a character survey with healthy older adults to select a final 4 IVAs (2 male, 2 female). We analyse the results, assess inter-subject agreement, and suggest guidelines for IVA selection. Matthew P. Aylett, Katerina Pappa, Mei Yii Lim, Ruth Aylett, Bruce W. Wilson, Mario A. Parra |
IVA | 1 |
| 2024 | Demonstration of the AMPER System for Individuals with Alzheimer's DiseaseabstractWe present a demonstration of the AMPER system: an Android application that guides individuals with Alzheimer’s disease and their carers through reminiscence therapy with the use of a virtual agent. The application supports a novel method of reminiscence story selection, using material metadata, user data, and a spreading activation algorithm. This aims to present relevant material to the user, in an order that attempts to mimic a human autobiographical memory. Bruce W. Wilson, Mei Yii Lim, Katerina Pappa, Matthew P. Aylett, Mario A. Parra, Ruth Aylett |
IVA | 4 |
| 2023 | Creating Inclusive Voices for the 21st Century: A Non-Binary Text-to-Speech for Conversational AssistantsabstractAs voice assistant usage continues to grow, their homogeneity becomes even more problematic with the UNESCO report, “I’d Blush if I could” showing that designing only feminine voice assistants encourages negative behavior, both with virtual assistants and with real people [3]. While masculine text-to-speech (TTS) voices exist, ones that cover the full range of gender presentations, such as non-binary or gender-ambiguous voices are largely missing. In this paper, we present a method of creating a non-binary TTS voice and an example voice, Sam, created with input from the non-binary and transgender communities. We have open-sourced the resulting voice, along with the process and data used to create it. Finally, we present results from a large-scale survey showing that non-binary individuals are more likely to prefer a non-binary voice assistant compared to cisgendered individuals and discuss differences across age and gender. Andreea Danielescu 0001, Sharone Horowit-Hendler, Alexandria Pabst, Kenneth Michael Stewart, Eric M. Gallo, Matthew P. Aylett |
CHI | 6 |
| 2023 | Why is my Agent so Slow? Deploying Human-Like Conversational Turn-TakingabstractThe emphasis on one-to-one speak/wait spoken conversational interaction with intelligent agents leads to long pauses between conversational turns, undermines the flow and naturalness of the interaction, and undermines the user experience. Despite ground breaking advances in the area of generating and understanding natural language with techniques such as LLMs, conversational interaction has remained relatively overlooked. In this workshop we will discuss and review the challenges, recent work and potential impact of improving conversational interaction with artificial systems. We hope to share experiences of poor human/system interaction, best practices with third party tools, and generate design guidance for the community. Matthew P. Aylett, Éva Székely, Donald McMillan, Gabriel Skantze, Marta Romeo, Joel E. Fischer, Gisela Reyes-Cruz |
HAI | 1 |
| 2023 | Synthesising Personality with Neural Speech SynthesisabstractMatching the personality of conversational agents to the personality of the user can significantly improve the user experience, with many successful examples in text-based chatbots.It is also important for a voice-based system to be able to alter the personality of the speech as perceived by the users.In this pilot study, fifteen voices were rated using Big Five personality traits.Five content-neutral sentences were chosen for the listening tests.The audio data, together with two rated traits (Extroversion and Agreeableness), were used to train a neural speech synthesiser based on one male and one female voices.The effect of altering the personality trait features was evaluated by a second listening test.Both perceived extroversion and agreeableness in the synthetic voices were affected significantly.The controllable range was limited due to a lack of variance in the source audio data.The perceived personality traits correlated with each other and with the naturalness of the speech. Shilin Gao, Matthew P. Aylett, David A. Braude, Catherine Lai |
SIGDIAL | 2 |
| 2020 | Speech Synthesis for the Generation of Artificial PersonalityabstractA synthetic voice personifies the system using it. In this work we examine the impact text content, voice quality and synthesis system have on the perceived personality of two synthetic voices. Subjects rated synthetic utterances based on the Big-Five personality traits and naturalness. The naturalness rating of synthesis output did not correlate significantly with any Big-Five characteristic except for a marginal correlation with openness. Although text content is dominant in personality judgments, results showed that voice quality change implemented using a unit selection synthesis system significantly affected the perception of the Big-Five, for example tense voice being associated with being disagreeable and lax voice with lower conscientiousness. In addition a comparison between a parametric implementation and unit selection implementation of the same voices showed that parametric voices were rated as significantly less neurotic than both the text alone and the unit selection system, while the unit selection was rated as more open than both the text alone and the parametric system. The results have implications for synthesis voice and system type selection for applications such as personal assistants and embodied conversational agents where developing an emotional relationship with the user, or developing a branding experience is important. Matthew P. Aylett, Alessandro Vinciarelli, Mirjam Wester |
IEEE Trans. Affect. Comput. | 1 |
| 2019 | All Together Now: The Living Audio Dataset
David A. Braude, Matthew P. Aylett, Caoimhín Laoide-Kemp, Simone Ashby, Kristen M. Scott, Brian Ó Raghallaigh, Anna Braudo, Alex Brouwer, Adriana Cornelia Stan |
INTERSPEECH | 2 |
| 2019 | The State of Speech in HCI: Trends, Themes and ChallengesabstractAbstract Speech interfaces are growing in popularity. Through a review of 99 research papers this work maps the trends, themes, findings and methods of empirical research on speech interfaces in the field of human–computer interaction (HCI). We find that studies are usability/theory-focused or explore wider system experiences, evaluating Wizard of Oz, prototypes or developed systems. Measuring task and interaction was common, as was using self-report questionnaires to measure concepts like usability and user attitudes. A thematic analysis of the research found that speech HCI work focuses on nine key topics: system speech production, design insight, modality comparison, experiences with interactive voice response systems, assistive technology and accessibility, user speech production, using speech technology for development, peoples’ experiences with intelligent personal assistants and how user memory affects speech interface interaction. From these insights we identify gaps and challenges in speech research, notably taking into account technological advancements, the need to develop theories of speech interface interaction, grow critical mass in this domain, increase design work and expand research from single to multiple user interaction contexts so as to reflect current use contexts. We also highlight the need to improve measure reliability, validity and consistency, in the wild deployment and reduce barriers to building fully functional speech interfaces for research. RESEARCH HIGHLIGHTS Most papers focused on usability/theory-based or wider system experience research with a focus on Wizard of Oz and developed systems Questionnaires on usability and user attitudes often used but few were reliable or validated Thematic analysis showed nine primary research topics Challenges identified in theoretical approaches and design guidelines, engaging with technological advances, multiple user and in the wild contexts, critical research mass and barriers to building speech interfaces Leigh Clark, Philip R. Doyle, Diego Garaialde, Emer Gilmartin, Stephan Schlögl, Jens Edlund, Matthew P. Aylett, João P. Cabral, Cosmin Munteanu, Justin Edwards, Benjamin R. Cowan |
Interact. Comput. | 7 |
| 2018 | A life story in three parts: the use of triptychs to make sense of personal digital dataabstractMany social media platforms support the curation of personal digital data, and, more recently, the use of that data for review and reflection. We explored the process of reflection by asking users to create a meaningful ‘triptych’ of photographs drawn from their Facebook accounts. In a first study, we asked participants to manually trawl their own accounts and select three relevant images, which we then framed and used as an interview probe. In a second study, we designed an automated triptych generation system and assessed participants’ experiences of using this system. We conducted qualitative analyses of participant interviews from both studies. Consistent with other ‘slow technology’ work, we found the act of creating a physical artefact from social media data gave that data new meaning, albeit with notable differences between manual versus automatically generated triptychs. We conclude by discussing possible improvements to the design of the automated triptych system. Lisa Thomas 0001, Elaine Farrow, Matthew P. Aylett, Pamela Briggs |
Pers. Ubiquitous Comput. | 3 |
| 2017 | Bot or not: exploring the fine line between cyber and human identityabstractSpeech technology is rapidly entering the everyday through the large scale commercial impact of systems such as Apple Siri and Amazon Echo. Meanwhile technology that allows voice cloning, voice modification, speech recognition, speech analytics and expressive speech synthesis has changed dramatically over recent years. The demonstration, described in this paper, is an educational tool in the form of an online quiz called `Bot or Not'. Using the quiz we have gathered impressions of what people realise is possible with current speech synthesis technology. The opinions of various groups regarding the synthesis of famous voices, sounding like a robot, and the difference between synthesis and voice modification were collected. Mirjam Wester, Matthew P. Aylett, David A. Braude |
ICMI | 2 |
| 2017 | Beyond the Listening Test: An Interactive Approach to TTS EvaluationabstractTraditionally, subjective text-To-speech (TTS) evaluation is performed through audio-only listening tests, where participants evaluate unrelated, context-free utterances. The ecological validity of ... Joseph Mendelson, Matthew P. Aylett |
INTERSPEECH | 2 |
| 2017 | Real-Time Reactive Speech Synthesis: Incorporating Interruptions
Mirjam Wester, David A. Braude, Blaise Potard, Matthew P. Aylett, Francesca Shaw |
INTERSPEECH | 4 |
| 2016 | Ask Alice: an artificial retrieval of information agentabstractWe present a demonstration of the ARIA framework, a modular approach for rapid development of virtual humans for information retrieval that have linguistic, emotional, and social skills and a strong personality. We demonstrate the framework's capabilities in a scenario where `Alice in Wonderland', a popular English literature book, is embodied by a virtual human representing Alice. The user can engage in an information exchange dialogue, where Alice acts as the expert on the book, and the user as an interested novice. Besides speech recognition, sophisticated audio-visual behaviour analysis is used to inform the core agent dialogue module about the user's state and intentions, so that it can go beyond simple chat-bot dialogue. The behaviour generation module features a unique new capability of being able to deal gracefully with interruptions of the agent. Michel F. Valstar, Tobias Baur 0001, Angelo Cafaro, Alexandru Ghitulescu, Blaise Potard, Johannes Wagner 0001, Elisabeth André, Laurent Durieu, Matthew P. Aylett, Soumia Dermouche, Catherine Pelachaud, Eduardo Coutinho, Björn W. Schuller, Yue Zhang 0014, Dirk Heylen, Mariët Theune, Jelte van Waterschoot |
ICMI | 9 |
| 2016 | Idlak Tangle: An Open Source Kaldi Based Parametric Speech Synthesiser Based on DNN
Blaise Potard, Matthew P. Aylett, David A. Baude, Petr Motlícek |
INTERSPEECH | 2 |
| 2016 | Cross Modal Evaluation of High Quality Emotional Speech Synthesis with the Virtual Human Toolkit
Blaise Potard, Matthew P. Aylett, David A. Baude |
IVA | 2 |
| 2016 | Designing Interactions with Multilevel Auditory Displays in Mobile Audio-Augmented RealityabstractAuditory interfaces offer a solution to the problem of effective eyes-free mobile interactions. In this article, we investigate the use of multilevel auditory displays to enable eyes-free mobile interaction with indoor location-based information in non-guided audio-augmented environments. A top-level exocentric sonification layer advertises information in a gallery-like space. A secondary interactive layer is used to evaluate three different conditions that varied in the presentation (sequentialversussimultaneous) and spatialisation (non-spatialisedversusegocentric/exocentric spatialisation) of multiple auditory sources. Our findings show that (1) participants spent significantly more time interacting with spatialised displays; (2) using the same design for primary and interactive secondary display (simultaneous exocentric) showed a negative impact on the user experience, an increase in workload and substantially increased participant movement; and (3) the other spatial interactive secondary display designs (simultaneous egocentric, sequential egocentric, and sequential exocentric) showed an increase in time spent stationary but no negative impact on the user experience, suggesting a more exploratory experience. A follow-up qualitative and quantitative analysis of user behaviour support these conclusions. These results provide practical guidelines for designing effective eyes-free interactions for far richer auditory soundscapes. Yolanda Vazquez-Alvarez, Matthew P. Aylett, Stephen A. Brewster, Rocio von Jungenfeld, Antti Virolainen |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2015 | Generating Narratives from Personal Digital Data: Using Sentiment, Themes, and Named Entities to Construct Stories
Elaine Farrow, Thomas Dickinson, Matthew P. Aylett |
INTERACT (4) | 3 |
| 2015 | Artificial personality and disfluencyabstractThe focus of this paper is artificial voices with different person-alities. Previous studies have shown links between an individ-ual’s use of disfluencies in their speech and their perceived per-sonality. Here, filled pauses (uh and um) and discourse markers (like, you know, I mean) have been included in synthetic speech as a way of creating an artificial voice with different personali-ties. We discuss the automatic insertion of filled pauses and dis-course markers (i.e., fillers) into otherwise fluent texts. The au-tomatic system is compared to a ground truth of human “acted” filler insertion. Perceived personality (as defined by the big five personality dimensions) of the synthetic speech is assessed by means of a standardised questionnaire. Synthesis without fillers is compared to synthesis with either spontaneous or synthetic fillers. Our findings explore how the inclusion of disfluencies influences the way in which subjects rate the perceived person-ality of an artificial voice. Index Terms: artificial personality, TTS, disfluency 1. Mirjam Wester, Matthew P. Aylett, Marcus Tomalin, Rasmus Dall |
INTERSPEECH | 2 |
| 2014 | A flexible front-end for HTSabstractParametric speech synthesis techniques depend on full context acoustic models generated by language front-ends, which anal-yse linguistic and phonetic structure. HTS, the leading paramet-ric synthesis system, can use a number of different front-ends to generate full context models for synthesis and training. In this paper we explore the use of a new text processing front-end that has been added to the speech recognition toolkit Kaldi as part of an ongoing project to produce a new parametric speech synthesis system, Idlak. The use of XML specification files, a modular design, and modern coding and testing approaches, make the Idlak front-end ideal for adding, altering and experi-menting with the contexts used in full context acoustic models. The Idlak front-end was evaluated against the standard Festival front-end in the HTS system. Results from the Idlak front-end compare well with the more mature Festival front-end (Idlak-2.83 MOS vs Festival- 2.85 MOS), although a slight reduction in naturalness perceived by non-native English speakers can be attributed to Festival’s insertion of non-punctuated pauses. Index Terms: speech synthesis, text processing, parametric synthesis, Kaldi, Idlak Matthew P. Aylett, Rasmus Dall, Arnab Ghoshal, Gustav Eje Henter, Thomas Merritt |
INTERSPEECH | 1 |
| 2014 | Phonetic feature extraction for context-sensitive glottal source processing
John Kane 0002, Matthew P. Aylett, Irena Yanushevskaya, Christer Gobl |
Speech Commun. | 2 |
| 2013 | Speaker and language independent voice quality classification applied to unlabelled corpora of expressive speechabstractVoice quality plays a pivotal role in speech style variation. Therefore, control and analysis of voice quality is critical for many areas of speech technology. Until now, most work has focused on small purpose built corpora. In this paper we apply state-of-the-art voice quality analysis to large speech corpora built for expressive speech synthesis. A fuzzy-input fuzzy-output support vector machine classifier is trained and validated using features extracted from these corpora. We then apply this classifier to freely available audiobook data and demonstrate a clustering of the voice qualities that approximates the performance of human perceptual ratings. The ability to detect voice quality variation in these widely available unlabelled audiobook corpora means that the proposed method may be used as a valuable resource in expressive speech synthesis. John Kane 0002, Stefan Scherer, Matthew P. Aylett, Louis-Philippe Morency, Christer Gobl |
ICASSP | 3 |
| 2012 | Proper Name Splicing in Computer Games with TTSabstractBuilding high quality synthesis systems with open domain vocabulary and a small audio database is a challenging problem, even when the targeted application is well constrained. Monophone unit concatenation (as opposed to diphone) is an approach that can compensate for the poor unit coverage that a small database implies. However, joining at phone boundaries is a delicate task that requires accurate targeting. In this paper, we present an automatically trained targeting system based on the parametric synthesiser HTS, and compare it to a concatenative monophone system and a baseline concatenative diphone system. We apply a novel evaluation methodology which includes a qualitative component, and allows for fast incremental development of synthesis systems. Preliminary results show that although the hybrid system performed significantly more poorly on out of database items, it is less affected by segmentation errors than the monophone system. Index Terms: hybrid speech synthesis, unit selection, evaluation of TTS systems Blaise Potard, Matthew P. Aylett, Christopher J. Pidcock |
INTERSPEECH | 2 |
| 2012 | Synthesising and Evaluating Cross-Modal Emotional Ambiguity in Virtual Agents
Matthew P. Aylett, Blaise Potard |
IVA | 1 |
| 2011 | The Romanian speech synthesis (RSS) corpus: Building a high quality HMM-based speech synthesis system using a high sampling rate
Adriana Cornelia Stan, Junichi Yamagishi, Simon King 0001, Matthew P. Aylett |
Speech Commun. | 4 |
| 2009 | Speech synthesis without a phone inventoryabstractIn speech synthesis the unit inventory is decided using phonological and phonetic expertise. This process is resource intensive and potentially sub-optimal. In this paper we investigate how acoustic clustering, together with lexicon constraints, can be used to build a self-organised inventory. Six English speech synthesis systems were built using two frameworks, unit selection and parametric HTS for three inventory conditions: 1) a traditional phone set, 2) a system using orthographic units, and 3) a self-organised inventory. A listening test showed a strong preference for the classic system, and for the orthographic system over the self-organised system. Results also varied by letter to sound complexity and database coverage. This suggests the self-organised approach failed to generalise pronunciation as well as introducing noise above and beyond that caused by orthographic sound mismatch. Index Terms: speech synthesis, unit selection, parametric synthesis, phone inventory, orthographic synthesis Matthew P. Aylett, Simon King 0001, Junichi Yamagishi |
INTERSPEECH | 1 |
| 2007 | The CereVoice Characterful Speech Synthesiser SDK
Matthew P. Aylett, Christopher J. Pidcock |
IVA | 1 |
| 2006 | Detecting High Level Dialog Structure Without Lexical InformationabstractThe potentially enormous audio resources now available to both organizations, and on the Internet, present a serious challenge to audio browsing technology. In this paper we outline a set of techniques that can be used to determine high level dialog structure without the requirement of resource intensive automatic speech recognition (ASR). Using syllable finding algorithms based on band pass energy together with prosodic feature extraction, we show that a sub-lexical approach to prosodic analysis can outperform results based on ASR and even those based on a word alignment which requires a complete transcription. We consider how these techniques could be integrated into ASR technology and suggest a framework for extending this type of sub-lexical prosodic analysis Matthew P. Aylett |
ICASSP (1) | 1 |
| 2005 | Synthesising hyperarticulation in unit selection TTSabstractWithin speech synthesis we often wish to give extra focus to words which carry important information, such as names, dates and amounts. In this paper we look carefully at cost functions that can be used to bias unit selection in favour of hyper-articulated speech in order to give this impression of focus. Hyper-articulated speech tends to be accented, emphatic and requires more articulatory effort. We apply two cost functions to try to force the selection of hyper-articulated speech. The first operates on the duration of units in the unit selection database, the second on the language redundancy (word trigram predictability) of the word containing the unit. We estimate their relative importance in selecting hyperarticulated speech in unit selection speech synthesis. A listening test was carried out where these cost functions were applied to one random content word in a haskins anomalous sentence. Listeners were asked to select the two clearest and most focused words from the sentence. The duration increasing cost function was significantly related to an increase in perceived prominence whereas low redundancy, and a combination of both approaches did not produce significant results. Thus, although a significant correlation exists between the average duration and redundancy of diphones and perceived prominence, such a correlation was not smoothly translated into error free method for altering such perceived prominence. Matthew P. Aylett |
INTERSPEECH | 1 |
| 2003 | My voice, your prosody: sharing a speaker specific prosody model across speakers in unit selection TTSabstractData sparsity is a major problem for data driven prosodic models. Being able to share prosodic data across speakers is a potential solution to this problem. This paper explores this potential solution by addressing two questions: 1) Does a larger less sparse model from a different speaker produce more natural speech than a small sparse model built from the original speaker? 2)Does a different speaker's larger model generate more unit selection errors than a small sparse model built from the original speaker? A unit selection approach is used to produce a lazy learning model of three English RP speaker's f0 and durational parameters. Speaker 1 (the target speaker) had a much smaller database (approximately one quarter to one fifth the size) of the other two. Speaker 2 was a female speaker with frequent mid phrase rises. Speaker 3 was a male speaker with a similar f0 range to speaker 1 and with a measured prosodic style suitable for news and financial text. We apply the models created for speaker 2 (an inappropriate model) and speaker 3 (an appropriate model) to speaker 1 and compare the results. Three passages (of three to four sentences in length) from challenging prosodic genres (news report, poetry and personal email) were synthesised using the target speaker and each of the three models. The synthesised utterances were played to 15 native english subjects and rated using a 5 point MOS scale. In addition, 7 experienced speech engineers rated each word for errors on a three point scale: 1. Acceptable, 2. Poor, 3. Unacceptable. The results suggest that a large model from an appropriate speaker does not sound more natural or produce fewer errors than a smaller model generated from the individual speaker's own data. In addition it shows that an inappropriate model does produce both less natural and more errors in the speech. High variance in both subject and materials analysis suggest both tests are far from ideal and that evaluation techniques for both error rate and naturalness need to improve. Matthew P. Aylett, Justin Fackrell, Peter Rutten |
INTERSPEECH | 1 |
| 2002 | Stochastic suprasegmentals: relationship between the spectral characteristics of vowels, redundancy and prosodic structureabstractPrevious work has shown a relationship between syllabic duration, redundancy within speech, and prosodic structure [1].In addition, a spectral care of articulation measure of vowels in spontaneous speech has supported these duration results and suggest that care of articulation varies inversely with redundancy and conversely with prosodic prominence.However, these spectral measures remain inconclusive due to measurement difficulties.In this paper a simpler spectral measurement is presented as a metric of care of articulation and applied to three vowels from each corner of the vowel triangle.Prosodic and redundancy factors can predict up to 5% of the variance of this new measurement, supporting the inverse redundancy result.However, in contrast to syllabic duration, whether prosodic boundaries are controlled for or not, the predictive power of prosodic factors and redundancy factors remain relatively independent.This suggests that 1. phrase final syllables, despite lengthening, do not show increased care of articulation if measured spectrally, and 2. unlike duration, redundancy factors affect the spectral characteristics of vowels independently of prosodic structure. Matthew P. Aylett |
INTERSPEECH | 1 |
| 2002 | A statistically motivated database pruning technique for unit selection synthesis
Peter Rutten, Matthew P. Aylett, Justin Fackrell |
INTERSPEECH | 2 |
| 2001 | Modelling care of articulation with HMMs is dangerousabstractChanges in care of articulation (COA) affect both the spectral and durational characteristics of speech. This can have severe repercussions on both the success of speech recognition, and the quality of speech synthesis. Although auto-segmentation has proven useful for measuring the durational effects of COA, an automatic spectral measurement has proven more problematic. In this paper, we will explore the use of the acoustic log likelihoods generated by HMM auto-segmentation as a measure of these changes in comparison with two phonetically motivated modeling systems based on vocalic F1/F2 values. When duration variation is controlled, the HMM output does not correlate with the human perception of vowel goodness, whereas, the phonetically motivated models do. Matthew P. Aylett |
INTERSPEECH | 1 |
| 2000 | Stochastic suprasegmentals: relationships between redundancy, prosodic structure and care of articulation in spontaneous speechabstractWithin spontaneous speech there are wide variations in the articulation of the same word by the same speaker. This paper explores two related factors which influence variation in articulation, prosodic structure and redundancy. We argue that the constraint of producing robust communication while efficiently expending articulatory effort leads to an inverse relationship between language redundancy and care of articulation. The inverse relationship improves robustness by spreading the information more evenly across the speech signal leading to a smoother signal redundancy profile. We argue that prosodic prominence is a linguistic means of achieving smooth signal redundancy. Prosodic prominence increases care of articulation and coincides with unpredictable sections of speech. By doing so, prosodic prominence leads to a smoother signal redundancy. Results confirm the strong relationship between prosodic prominence and care of articulation as well as an inverse relationship between language redundancy and care of articulation. In addition, when variation in prosodic boundaries is controlled for, language redundancy can predict up to 65% of the variance in raw syllabic duration. This is comparable with 64% predicted by prosodic prominence (accent, lexical stress and vowel type). Moreover most (62%) of this predictive power is shared. This suggests that, in English, prosodic structure is the means with which constraints caused by a robust signal requirement are expressed in spontaneous speech. Matthew P. Aylett |
INTERSPEECH | 1 |
| 1998 | Building a statistical model of the vowel space for phoneticiansabstractVowel space data (A two dimensional F1/F2 plot) is of interest to phoneticians for the purpose of comparing different accents, languages, speaker styles and individual speakers. Current automatic methods used by speech technologists do not generally produce traditional vowel space models (See [6] for an overview); instead they tend to produce hyper dimensional code books covering the entire speakers speech stream. This makes it difficult to relate results generated by these methods to observations in laboratory phonetics. In order to address these problems a model was developed based on a mixture Gaussian density function fitted using expectation maximisation on F1/F2 data producing a probability distribution in F1/F2 space. Speech was pre-processed using voicing to automatically excerpt vowel data without any need for segmentation and a parametric fit algorithm [7] was applied to calculate likely vowel targets. The result was a clear visualisation of a speaker's vowel space requiring ... Matthew P. Aylett |
ICSLP | 1 |
| 1998 | The automatic marking of prominence in spontaneous speech using duration and part of speech informationabstractThe work reported in this paper was the result of the need to label a large corpus of spontaneous, task-oriented dialogue with prosodic prominences. A computational model using only word duration, part of speech and a dictionary lookup of each word's canonical phonemic contents was trained against the results of a human coder marking prominence. Because word durations were normalised, it was possible to set a common threshold for all members of a form class above which the lexically stressed syllables were classed as prominent. The method used is presented and the relative importance of duration information, phonemic contents, syllabic context and part of speech information is explored. The automatic coder was validated against unseen material and achieved a 58% agreement with a human coder. Further investigation showed that three humans coders agreed no better with each other than each agreed with the computational model. Thus, although the automatic system did not conform very well t... Matthew P. Aylett, Matthew Bull |
ICSLP | 1 |
| 1998 | Vowel quality in spontaneous speech: what makes a good vowel?abstractClear speech is characterised by longer segmental durations and less target undershoot [9] which results in more extreme spectral features.This paper deals with the clarity of vowels produced in spontaneous speech in a large corpus of task-oriented dialogues.We present an automatic technique for measuring vowel clarity on the basis of a vowel's spectral characteristics.This technique was evaluated using a perceptual test.Subjects rated the 'goodness' of vowels with different spectral characteristics with controlled duration and amplitude and these results were compared with an automatic rating.Results indicated that although agreement between subjects and the automatic measurement was poor it was as poor as the agreement between subjects.On the basis of these results we address the following questions:1. Can subjects reliably judge the clarity of vowels excerpted from spontaneous speech without duration cues? 2. Can a statistical model [3] reliably predict the subjects' response to such vowels? Matthew P. Aylett, Alice Turk |
ICSLP | 1 |
| 1998 | An analysis of the timing of turn-taking in a corpus of goal-oriented dialogueabstractThis paper presents a context-based analysis of the intervals between different speakers ’ utterances in a corpus of taskoriented dialogue (the Human Communication Research Centre’s Map Task Corpus. See Anderson et al. 1991). In the analysis, we assessed the relationship between inter-speaker intervals and various contextual factors, such as the effects of eye contact, the presence of conversational game boundaries, the category of move in an utterance, and the degree of experience with the task in hand. The results of the analysis indicated that the main factors which gave rise to significant differences in inter-speaker intervals were those which related to decision-making and planning- the greater the amount of planning, the greater the inter-speaker interval. Differences between speakers were also found to be significant, although this effect did not necessarily interact with all other effects. These results provide unique and useful data for the improved effectiveness of dialogue systems. 1. Matthew Bull, Matthew P. Aylett |
ICSLP | 2 |