EDBT 2026 Demo / reviewers in the wild / expert
Catherine I. Watson
dblp:226/1933 · also Catherine Inez Watson
· DBLP profile ↗
29ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0001-9010-5188ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 4 first-author · 8 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TTSVowelViz: A Tool for Visualising Text-to-Speech Model Training via Vowel SpacesabstractIn text-to-speech (TTS) model training, the saturation of the loss curve indicates how well a model learns the characteristics of the training dataset. But it does not reveal the linguistic properties learned by the model. Existing TTS approaches miss the potential to incorporate linguistic insights into model training. We introduce TTSVowelViz, a novel tool that visualises static and dynamic vowel spaces during model training, bridging linguistic knowledge and TTS model development. It helps identify which vowel sounds are accurately learned and how the vowel spaces are evolved during training. To assess TTSVowelViz, we fine-tuned a TTS model from General American English to New Zealand English and conducted a perception test. Our results show that the formants of specific vowels in the vowel spaces generated by TTSVowelViz align with human perception, effectively visualising the perceived accent shift. This work highlights vowel space visualisation as a valuable interpretability tool for TTS training. Pasindu Udawatta, Jesin James, Balamurali B. T, Catherine I. Watson, Ake Nicholas, Binu Abeysinghe |
LREC | 4 |
| 2026 | A review on speech emotion recognition for low-resource and Indigenous languagesabstractSpeech emotion recognition (SER) is an emerging field in human–computer interaction. Although numerous studies have focused on SER for well-resourced languages, the literature reveals a significant gap in research on low-resource and Indigenous (LRI) languages. This paper presents a comprehensive review of the existing literature on SER in the context of LRI languages, analysing critical factors to consider at each stage of designing an SER system. The review indicates that most studies on SER for LRI languages adopt emotion categories established for well-resourced languages, often assuming the universality of emotions. However, the literature suggests that this approach may be limited due to emotional disparities influenced by cultural variations. Additionally, the review underscores that current SER systems typically lack community-oriented methodologies in the development of technology for LRI languages. The importance of feature selection is highlighted, with evidence suggesting that a combination of traditional machine learning methods and carefully selected acoustic features may offer viable options for SER in these languages. Furthermore, the review identifies a need for further exploration of semi-supervised and unsupervised approaches to enhance SER capabilities in LRI contexts. Overall, current SER systems for LRI languages lag behind state-of-the-art standards due to the lack of resources, indicating that there is still much work to be done in this area. Himashi Rathnayake, Jesin James, Gianna Leoni, Ake Nicholas, Catherine I. Watson, Peter Keegan |
Speech Commun. | 5 |
| 2025 | Introducing EMOPARKNZ: the Emotional Speech Database from New Zealand English Speakers with Parkinson's Disease
Itay Ben-Dom, Catherine I. Watson, Clare M. McCann |
INTERSPEECH | 2 |
| 2025 | Perception of Long and Short Vowel Contrast in Te Reo Māori in Clean and Everyday Listening EnvironmentsabstractTe reo Māori (the Māori language) is the language of the indigenous people in New Zealand and has a long-short vowel contrast. This study first investigates the cues used to perceive vowel length for Māori fluent users and learners. Secondly, it explores the effect of an everyday listening environment in the form of a traditional meeting house (wharenui) on cue-weighting between duration and stress for identifying long and short vowels when the stimuli are rendered with the wharenui room acoustics. An identification test was carried out with three pairs of words that differed in stress location and were manipulated in vowel duration. We found both groups to mainly use duration as a cue, even though the advanced listeners commented that they were listening for 'intonation'. For the wharenui acoustics, there were no differences observed between the groups and only a small categorical shift for the /a:/ vowel for the learners. C. T. Justine Hui, Jenice Kuzhikombil, Isabella Shields, Hiraia Haami-Wells, Catherine I. Watson, Peter Keegan |
INTERSPEECH | 5 |
| 2025 | Role of language familiarity in understanding speech in noise under various acoustic environmentsabstractWe communicate in complex acoustic environments in everyday life but our familiarity with the language can affect how well we can understand speech in these environments. The current study examines the role of language familiarity in understanding speech in varying acoustic environments via a speech intelligibility test conducted under anechoic and reverberant conditions with various speech-noise separation angles. Four groups were recruited with differing level of language familiarity: first language (L1) New Zealand English (NZE) listeners, second language (L2) Japanese native listeners with exposure to NZE, L2 Japanese native listeners with overseas English experiences without exposure to NZE, and Japanese native listeners who have learnt English as a foreign language (FL) without overseas English experiences. The L1 group performed better in overall speech intelligibility performance compared to the 3 Japanese native groups. Contrary to previous literature where non-native listeners were found to have a similar benefit from spatial separation to native listeners, this was not the case for the FL group, suggesting that this benefit is only available for listeners with a certain level of language familiarity. While there were differences between L2 and FL groups in the anechoic condition, these differences become marginal in the reverberant conditions for the two groups with little exposure to NZE. This suggests that familiarity to the specific language variety has an advantage in acoustically adverse environments. C. T. Justine Hui, Hinako Masuda, Eri Osawa, Takayuki Arai, Catherine I. Watson, Yusuke Hioka |
Speech Commun. | 5 |
| 2024 | Performance of single-channel speech enhancement algorithms on Mandarin listeners with different immersion conditions in New Zealand EnglishabstractSpeech enhancement (SE) is a widely used technology to improve the quality and intelligibility of noisy speech. So far, SE algorithms were designed and evaluated on native listeners only, but not on non-native listeners who are known to be more disadvantaged when listening in noisy environments. This paper investigates the performance of five widely used single-channel SE algorithms on early-immersed New Zealand English (NZE) listeners and native Mandarin listeners with different immersion conditions in NZE under negative input signal-to-noise ratio (SNR) by conducting a subjective listening test in NZE sentences. The performance of the SE algorithms in terms of speech intelligibility in the three participant groups was investigated. The result showed that the early-immersed group always achieved the highest intelligibility. The late-immersed group outperformed the non-immersed group for higher input SNR conditions, possibly due to the increasing familiarity with the NZE accent, whereas this advantage disappeared at the lowest tested input SNR conditions. The SE algorithms tested in this study failed to improve and rather degraded the speech intelligibility, indicating that these SE algorithms may not be able to reduce the perception gap between early-, late- and non-immersed listeners, nor able to improve the speech intelligibility under negative input SNR in general. These findings have implications for the future development of SE algorithms tailored to Mandarin listeners, and for understanding the impact of language immersion on speech perception in noise. Yunqi C. Zhang, Yusuke Hioka, C. T. Justine Hui, Catherine I. Watson |
Speech Commun. | 4 |
| 2022 | Visualising Model Training via Vowel Space for Text-To-Speech SystemsabstractWith the recent developments in speech synthesis via machine learning, this study explores incorporating linguistics knowledge to visualise and evaluate synthetic speech model training.If changes to the first and second formant (in turn, the vowel space) can be seen and heard in synthetic speech, this knowledge can inform speech synthesis technology developers.A speech synthesis model trained on a large General American English database was fine-tuned into a New Zealand English voice to identify if the changes in the vowel space of synthetic speech could be seen and heard.The vowel spaces at different intervals during the fine-tuning were analysed to determine if the model learned the New Zealand English vowel space.Our findings based on vowel space analysis show that we can visualise how a speech synthesis model learns the vowel space of the database it is trained on.Perception tests confirmed that humans could perceive when a speech synthesis model has learned characteristics of the speech database it is training on.Using the vowel space as an intermediary evaluation helps understand what sounds are to be added to the training database and build speech synthesis models based on linguistics knowledge. Binu Abeysinghe, Jesin James, Catherine I. Watson, Felix Marattukalam |
INTERSPEECH | 3 |
| 2022 | Differences between listeners with early and late immersion age in spatial release from masking in various acoustic environments
C. T. Justine Hui, Yusuke Hioka, Hinako Masuda, Catherine I. Watson |
Speech Commun. | 4 |
| 2021 | Comparing Speech Enhancement Techniques for Voice Adaptation-Based Speech SynthesisabstractThis study investigates the use of speech enhancement techniques in creating text-to-speech voices with degraded or noisy speech. A number of synthetic voices were created using speech that was first degraded by different noise types at various signal-to-noise ratios (SNRs), then enhanced through four speech enhancement algorithms: Subspace, Wiener filter, SEGAN and a DNN-based method. Subjective listening tests show that the quality of the synthetic voices produced by subspace and the DNN-based method enhanced speech outperforms the quality of the voices created using Wiener filter or SEGAN enhanced speech at low SNRs, and speech enhanced by the subspace method results in higher quality synthetic speech at higher SNRs. Nicholas Eng, C. T. Justine Hui, Yusuke Hioka, Catherine I. Watson |
Interspeech | 4 |
| 2019 | Effects of sentence structure and word complexity on intelligibility in machine-to-human communications
C. T. Justine Hui, Sahil Jain, Catherine I. Watson |
Comput. Speech Lang. | 3 |
| 2018 | An Open Source Emotional Speech Corpus for Human Robot Interaction ApplicationsabstractFor further understanding the wide array of emotions embedded in human speech, we are introducing a strictly-guided simulated emotional speech corpus. In contrast to existing speech corpora, this was constructed by maintaining an equal distribution of 4 long vowels in New Zealand English. This balance is to facilitate emotion related formant and glottal source feature comparison studies. Also, the corpus has 5 secondary emotions and 5 primary emotions. Secondary emotions are important in Human-Robot Interaction (HRI) to model natural conversations among humans and robots. But there are few existing speech resources to study these emotions, which has motivated the creation of this corpus. A large scale perception test with 120 participants showed that the corpus has approximately 70% and 40% accuracy in the correct classification of primary and secondary emotions respectively. The reasons behind the differences in perception accuracies of the two emotion types is further investigated. A preliminary prosodic analysis of corpus shows significant differences among the emotions. The corpus is made public at: github.com/tli725/JL-Corpus. Jesin James, Catherine I. Watson |
INTERSPEECH | 3 |
| 2018 | Artificial Empathy in Social Robots: An analysis of Emotions in SpeechabstractArtificial speech developed using speech synthesizers has been used as the voice for robots in Human Robot Interaction (HRI). As humans anthropomorphize robots, an empathetically interacting robot is expected to increase the level of acceptance of social robots. Here, a human perception experiment evaluates whether human subjects perceive empathy in robot speech. For this experiment, empathy is expressed only by adding appropriate emotions to the words in speech. Also, humans' preferences for a robot interacting with empathetic speech versus a standard robotic voice are also assessed. The results show that humans are able to perceive empathy and emotions in robot speech, and prefer it over the standard robotic voice. It is important for the emotions in empathetic speech to be consistent with the language content of what is being said, and with the human users' emotional state. Analyzing emotions in empathetic speech using valence-arousal model has revealed the importance of secondary emotions in developing empathetically speaking social robots. Jesin James, Catherine I. Watson, Bruce A. MacDonald |
RO-MAN | 2 |
| 2017 | The Motivation and Development of MPAi, a Māori Pronunciation AidabstractThis paper outlines the motivation and development of a pronunciation aid (MPAi) for the Māori language, the language of the indigenous people of New Zealand. Māori is threatened and after a break in transmission the language is currently undergoing revitalization. The data for the aid has come from a corpus of 60 speakers (men and women). The language aid allows users to model their speech against exemplars from young speakers or older speakers of Māori. This is important, because of the status of the elders in the Māori speaking community, but it also recognizes that Māori is undergoing substantial vowel change. The pronunciation aid gives feedback on vowel production via formant analysis, and selected words via speech recognition. The evaluation of the aid by 22 language teachers is presented and the resulting changes are discussed. Catherine I. Watson, Peter Keegan, Margaret Maclagan, Ray Harlow, J. King |
INTERSPEECH | 1 |
| 2014 | Mappings between vocal tract area functions, vocal tract resonances and speech formants for multiple speakers
Catherine I. Watson |
INTERSPEECH | 1 |
| 2011 | Phrases, Pitch and Perceived Prominence in Maori
Catherine I. Watson, Ray Harlow, Jeanette King, Margaret Maclagan, Helen Charters, Peter Keegan |
INTERSPEECH | 2 |
| 2010 | The effect of audience familiarity on the perception of modified accent
Jonathan Teutenberg, Catherine I. Watson |
INTERSPEECH | 2 |
| 2010 | Deployment of a service robot to help older peopleabstractThis paper presents the first version of a mobile service robot designed for older people. Six service application modules were developed with the key objective being successful interaction between the robot and the older people. A series of trials were conducted in an independent living facility at a retirement village, with the participation of 32 residents and 21 staff. In this paper, challenges of deploying the robot and lessons learned are discussed. Results show that the robot could successfully interact with people and gain their acceptance. Chandimal Jayawardena, I-Han Kuo, Ulrike Unger, Aleksandar Igic, Richie Wong, Catherine I. Watson, Rebecca Q. Stafford, Elizabeth Broadbent, Priyesh Tiwari, Joochan Sohn, Bruce A. MacDonald |
IROS | 6 |
| 2010 | Improved robot attitudes and emotions at a retirement home after meeting a robotabstractThis study investigated whether attitudes and emotions towards robots predicted acceptance of a healthcare robot in a retirement village population. Residents (n = 32) and staff (n = 21) at a retirement village interacted with a robot for approximately 30 minutes. Prior to meeting the robot, participants had their heart rate and blood pressure measured. The robot greeted the participants, assisted them in taking their vital signs, performed a hydration reminder, told a joke, played a music video, and asked some questions about falls and medication management. Participants were given two questionnaires; one before and one after interacting with the robot. Measures included in both questionnaires were the Robot Attitude Scale (RAS) and the Positive and Negative Affect Schedule (PANAS). After using the robot, participants rated the overall quality of the robot interaction. Both residents and staff reported more favourable attitudes (p <; .05) and decreases in negative affect (p <; .05) towards the robot after meeting it, compared with before meeting it. Pre-interaction emotions and robot attitudes, combined with post-interaction changes in emotions and robot attitudes, were highly predictive of participants' robot evaluations (R = .88, p <; .05). The results suggest both pre-interaction emotions and attitudes towards robots, as well as experience with the robot, are important areas to monitor and address in influencing acceptance of healthcare robots in retirement village residents and staff. The results support an active cognition model that incorporates a feedback loop based on re-evaluation after experience. Rebecca Q. Stafford, Elizabeth Broadbent, Chandimal Jayawardena, Ulrike Unger, I-Han Kuo, Aleksandar Igic, Richie Wong, Ngaire Kerse, Catherine I. Watson, Bruce A. MacDonald |
RO-MAN | 9 |
| 2009 | Investigating changes in the rhythm of maori over time
Margaret Maclagan, Catherine I. Watson, Jeanette King, Ray Harlow, Peter Keegan |
INTERSPEECH | 2 |
| 2009 | Expressive facial speech synthesis on a robotic platformabstractThis paper presents our expressive facial speech synthesis system Eface, for a social or service robot. Eface aims at enabling a robot to deliver information clearly with empathetic speech and an expressive virtual face. The empathetic speech is built on the Festival speech synthesis system and provides robots the capability to speak with different voices and emotions. Two versions of a virtual face have been implemented to display the robot's expressions. One with just over 100 polygons has a lower hardware requirement but looks less natural. The other has over 1000 polygons; it looks realistic, but costs more CPU resource and requires better video hardware. The whole system is incorporated into the popular open source robot interface Player, which makes client programs easy to write and debug. Also, it is convenient to use the same system with different robot platforms. We have implemented this system on a physical robot and tested it with a robotic nurse assistant scenario. Xingyan Li, Bruce A. MacDonald, Catherine I. Watson |
IROS | 3 |
| 2008 | Modelling and synthesising F0 contours with the discrete cosine transformabstractThe discrete cosine transform is proposed as a basis for representing fundamental frequency (F0) contours of speech. The advantages over existing representations include deterministic algorithms for both analysis and synthesis and a simple distance measure in the parameter space. A two-tier model using the DCT is shown to be able to model F0 contours to around 10 Hz RMS error. A proof-of-concept system for synthesising DCT parameters is evaluated, showing that the benefits do not come at the expense of speech synthesis applications. Jonathan Teutenberg, Catherine I. Watson, Patricia J. Riddle |
ICASSP | 2 |
| 2008 | A Niuean variant of New Zealand English?
Donna Starks, Catherine I. Watson |
INTERSPEECH | 3 |
| 2008 | The English pronunciation of successive groups of Maori speakers
Catherine I. Watson, Margaret Maclagan, Jeanette King, Ray Harlow |
INTERSPEECH | 1 |
| 2007 | Age-related changes in fundamental frequency and formants: a longitudinal study of four speakersabstractThe study is concerned with a longitudinal acoustic analysis of two sets of recordings from the same four speakers over an interval of between 29 and 50 years. The aim was to determine whether there is any evidence for age-related acoustic changes. Our analysis showed that the same speakers have lower f0, a lower F1, a marginally lower F2, and an unchanging or sometimes higher F3 in their later recordings. There is some suggestion from these data that the change in F1-f0 in Bark from earlier to late recordings is proportional to the change in F3-F2 in Bark. This suggests that there is shift in the speaker space roughly along a diagonal in the phonetic height x backness plane with increasing age. 1. Jonathan Harrington, Sallyanne Palethorpe, Catherine I. Watson |
INTERSPEECH | 3 |
| 2000 | Matching a tone-based and tune-based approach to English intonation for concept-to-speech generation
Elke Teich, Catherine I. Watson, Cécile Pereira |
COLING | 2 |
| 1999 | A profile of the discourse and intonational structures of route descriptionsabstractThe problem of how to estimate variance parameters in client models from scarce data is addressed in the context of text-dependent, HMM-based, automatic speaker verification. Variance flooring and variance scaling are investigated as two alternative estimation techniques and are used with or without variance tying on the state level to reduce the number of parameters to estimate. The best results are achieved with no tying and a variance flooring method where the floor to a variance vector in a client model is proportional to the corresponding variance vector in a gender-dependent, multispeaker, non-client model. Further, variance tying reduces storage requirements considerably without much loss in recognition accuracy. It is also confirmed from a previous study that re-using non-client variances has comparable performance to variance flooring and is much simpler. Comparisons are made on three large telephone quality speech corpora. Sandra Williams, Catherine I. Watson |
EUROSPEECH | 2 |
| 1998 | Dynamic features in children's vowels
Steve Cassidy, Catherine I. Watson |
ICSLP | 2 |
| 1998 | Some acoustic characteristics of emotion
Cécile Pereira, Catherine I. Watson |
ICSLP | 2 |
| 1998 | A kinematic analysis of new zealand and australian English vowel spaces
Catherine I. Watson, Jonathan Harrington, Sallyanne Palethorpe |
ICSLP | 1 |