Martin Cooke

dblp:33/5151 · also Martin P. Cooke · DBLP profile ↗
← Back
85ranked-venue papers
22as first author
8since 2021 · last 2025
0000-0002-9514-3332ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 75 · 17 first-author · 8 since 2021Artificial intelligence and machine learning · 63 · 14 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Prediction of listening effort ratings for habitual and clear-Lombard speech presented in noise
abstract
Contains fulltext : 322635.pdf (Publisher’s version ) (Open Access)
Esther Janse, Martin Cooke
INTERSPEECH3
2024 Listeners' F0 preferences in quiet and stationary noise
Olympia Simantiraki, Martin Cooke
INTERSPEECH2
2023 Exploring the mutual intelligibility breakdown caused by sculpting speech from a competing speech signal
Martin Cooke, María Luisa García Lecumberri
INTERSPEECH1
2023 The effect of masking noise on listeners' spectral tilt preferences
Olympia Simantiraki, Yannis Pantazis, Martin Cooke
INTERSPEECH3
2022 Generating iso-accented stimuli for second language research: methodology and a dataset for Spanish-accented English
Rubén Pérez Ramón, Martin Cooke, María Luisa García Lecumberri
INTERSPEECH2
2022 The Lombard intelligibility benefit of native and non-native speech for native and non-native listeners
abstract
Speech produced in noise (Lombard speech) is more intelligible than speech produced in quiet (plain speech). Previous research on the Lombard intelligibility benefit focused almost entirely on how native speakers produce and perceive Lombard speech. In this study, we investigate the size of the Lombard intelligibility benefit of both native (American-English) and non-native (native Dutch) English for native and non-native listeners (Dutch and Spanish). We used a glimpsing metric to measure the energetic masking potential of speech, which predicted that both native and non-native Lombard speech could withstand greater amounts of masking to a similar extent, compared to plain speech. In an intelligibility experiment, native English, Spanish, and Dutch listeners listened to the same words, mixed with noise. While the non-native listeners appeared to benefit more from Lombard speech than the native listeners did, each listener group experienced a similar benefit for native and non-native Lombard speech. Energetic masking, as captured by the glimpsing metric, only accounted for part of the Lombard benefit, indicating that the Lombard intelligibility benefit does not only result from a shift in spectral distribution. Despite subtle native language influences on non-native Lombard speech, both native and non-native speech provides a Lombard benefit.
Katherine Marcoux, Martin Cooke, Benjamin V. Tucker, Mirjam Ernestus
Speech Commun.2
2022 Foreign accent strength and intelligibility at the segmental level
abstract
The relationship between strength of foreign accent and intelligibility is not straightforward. This relationship resists a simple characterisation due in part to the multiplicity of cues that carry accent in the word- and sentence-level materials typically used in the study of accent. One of the principal conveyors of accent is the phonetic segment. The current study attempts to isolate this segmental contribution to foreign accent and consequently measure the relationship between segmental accent and intelligibility for listeners with differing linguistic correspondence to the target and accented language. English, Spanish and Czech listeners identified English words in which the initial consonant was either intact, or had been replaced by a Spanish-accented counterpart; in a second task, they rated the accent strength of the same tokens. All speech material was produced by an English-Spanish bilingual talker. Overall, Spanish listeners displayed a smaller loss of intelligibility due to the accented segment than native English listeners, while the Czech cohort experienced the largest intelligibility loss. However, the relationship between accent strength and intelligibility loss was not linear, varying with phoneme identity and its role in a listener's first language. These findings suggest that how accented and intelligible a sound is depends strongly on the interactions between the phonological systems of speakers and listeners.
Rubén Pérez Ramón, María Luisa García Lecumberri, Martin Cooke
Speech Commun.3
2021 SpeechAdjuster: A Tool for Investigating Listener Preferences and Speech Intelligibility
Olympia Simantiraki, Martin Cooke
Interspeech2
2020 Effects of Spectral Tilt on Listeners' Preferences And Intelligibility
abstract
High intelligibility can be achieved when listening to synthetic or artificially-produced speech under adverse conditions. But can listener preferences reveal any extra information when intelligibility is at ceiling? This paper describes a real-time speech modification technique which allows evaluation of the impact of individual speech properties on listeners' preferences. The current study investigates spectral tilt, a feature which also changes naturally in different speaking styles such as Lombard speech. In the listening experiment, participants were asked to adjust the spectral tilt in masked conditions; subsequently, intelligibility was assessed. Listeners preferred flatter spectral tilts as SNR decreased. Additionally, tilt preferences were evident even while intelligibility was at or close to ceiling levels, suggesting that preferences provide information over and above that measured in traditional intelligibility-based studies. On the basis of these findings a method for probabilistic modelling of listener preferences, useful in the development of speech enrichment algorithms, is proposed.
Olympia Simantiraki, Martin Cooke, Yannis Pantazis
ICASSP2
2020 The Effect of Language Proficiency on the Perception of Segmental Foreign Accent
Rubén Pérez Ramón, María Luisa García Lecumberri, Martin Cooke
INTERSPEECH3
2020 Intelligibility-Enhancing Speech Modifications - The Hurricane Challenge 2.0
abstract
S.1341-1345
Jan Rennies, Henning F. Schepker, Cassia Valentini-Botinhao, Martin Cooke
INTERSPEECH4
2020 Exploring Listeners' Speech Rate Preferences
abstract
Fast speech may reduce intelligibility, but there is little agree-ment as to whether listeners benefit from slower speech in noisyconditions. The current study explored the relationship betweenspeech rate and masker properties using a listening preferencetechnique in which participants were able to control speech ratein real time. Spanish listeners adjusted speech rate while lis-tening to word sequences in quiet, in stationary noise at signal-to-noise ratios of 0, +6 and +12dB, and in modulated noisefor 5 envelope modulation rates. Following selection of a pre-ferred rate, participants went on to identify words presented atthat rate. Listeners favoured faster speech in quiet, chose in-creasingly slower rates in increasing levels of stationary noise,and showed a preference for speech rates that led to a contrastwith masker envelope modulation rates. Participants showeddistinct preferences even when intelligibility was near ceilinglevels. These outcomes suggest that individuals attempt to com-pensate for the decrement in cognitive resources availability inmore adverse conditions by reducing speech rate and are able toexploit differences in modulation properties of the target speechand masker. The listening preference approach provides in-sights into factors such as listening effort that are not measuredin intelligibility-based metrics.
Olympia Simantiraki, Martin Cooke
INTERSPEECH2
2020 Is segmental foreign accent perceived categorically?
Rubén Pérez Ramón, Martin Cooke, María Luisa García Lecumberri
Speech Commun.2
2019 Combining spectral and temporal modification techniques for speech intelligibility enhancement
Martin Cooke, Vincent Aubanel, María Luisa García Lecumberri
Comput. Speech Lang.1
2018 Impact of Different Speech Types on Listening Effort
abstract
Listeners are exposed to different types of speech in everyday life, from natural speech to speech that has undergone modifications or has been generated synthetically. While many studies have focused on measuring the intelligibility of these distinct speech types, their impact on listening effort is not known. The current study combined an objective measure of intelligibility, a physiological measure of listening effort (pupil size) and listeners’ subjective judgements, to examine the impact of four speech types: plain (natural) speech, speech produced in noise (Lombard speech), speech enhanced to promote intelligibility, and synthetic speech. For each speech type, listeners responded to sentences presented in one of three levels of speech-shaped noise. Subjective effort ratings and intelligibility scores showed an inverse ranking across speech types, with synthetic speech being the most demanding and enhanced speech the least. Pupil size measures indicated an increase in listening effort with decreasing signal-to-noise ratio for all speech types apart from\nsynthetic speech, which required significantly more effort at the most favourable noise level. Naturally and artificially modified speech were less effortful than plain speech at the more adverse noise levels. These outcomes indicate a clear impact of speech type on the cognitive demands required for comprehension.
Olympia Simantiraki, Martin Cooke, Simon King 0001
INTERSPEECH2
2018 Learning static spectral weightings for speech intelligibility enhancement in noise
Martin Cooke
Comput. Speech Lang.2
2017 Glottal Source Features for Automatic Speech-Based Depression Assessment
abstract
Depression is one of the most prominent mental disorders, with an increasing rate that makes it the fourth cause of disability worldwide. The field of automated depression assessment has emerged to aid clinicians in the form of a decision support system. Such a system could assist as a pre-screening tool, or even for monitoring high risk populations. Related work most commonly involves multimodal approaches, typically combining audio and visual signals to identify depression presence and/or severity. The current study explores categorical assessment of depression using audio features alone. Specifically, since depression-related vocal characteristics impact the glottal source signal, we examine Phase Distortion Deviation which has previously been applied to the recognition of voice qualities such as hoarseness, breathiness and creakiness, some of which are thought to be features of depressed speech. The proposed method uses as features DCT-coefficients of the Phase Distortion Deviation for each frequency band. An automated machine learning tool, Just Add Data, is used to classify speech samples. The method is evaluated on a benchmark dataset (AVEC2014), in two conditions: read-speech and spontaneous-speech. Our findings indicate that Phase Distortion Deviation is a promising audio-only feature for automated detection and assessment of depressed speech.
Olympia Simantiraki, Paulos Charonyktakis, Anastasia Pampouchidou, Manolis Tsiknakis, Martin Cooke
INTERSPEECH5
2016 The Effects of Modified Speech Styles on Intelligibility for Non-Native Listeners
Martin Cooke, María Luisa García Lecumberri
INTERSPEECH1
2016 Can Intensive Exposure to Foreign Language Sounds Affect the Perception of Native Sounds?
María Luisa García Lecumberri, Martin Cooke
INTERSPEECH3
2016 Language Effects in Noise-Induced Word Misperceptions
abstract
International audience
María Luisa García Lecumberri, Jon Barker, Ricard Marxer, Martin Cooke
INTERSPEECH4
2016 Glimpse-Based Metrics for Predicting Speech Intelligibility in Additive Noise Conditions
abstract
The glimpsing model of speech perception in noise operates by recognising those speech-dominant spectro-temporal regions, or glimpses, that survive energetic masking; hence, a speech recognition component is an integral part of the model. The current study evaluates whether a simpler family of metrics based solely on quantifying the amount of supra-threshold target speech available after energetic masking can account for subjective intelligibility. The predictive power of glimpse-based metrics is compared for natural, processed and synthetic speech in the presence of stationary and fluctuating maskers. These metrics are raw glimpse proportion, extended glimpse proportion, and two further refinements: one, FMGP, incorporates a component simulating the effect of forward masking; the other, HEGP, selects speech-dominant spectro-temporal regions with above-average energy on the noisy speech. The metrics are compared alongside a state-of-the-art non-glimpsing metric, using three large datasets of listener scores. Both FMGP and HEGP equal or improve upon the predictive power of the raw and extended metrics, with across-masker correlations ranging from 0.81--0.92; both metrics equal or exceed the state-of-the-art metric in all conditions. These outcomes suggests that easily-computed measures of unmasked, supra-threshold speech can serve as robust proxies for intelligibility across a range of speech styles and additive masking conditions.
Martin Cooke
INTERSPEECH2
2016 Undoing Misperceptions: A Microscopic Analysis of Consistent Confusions Through Signal Modifications
Máté Attila Tóth, Martin Cooke
INTERSPEECH2
2016 Misperceptions Arising from Speech-in-Babble Interactions
Máté Attila Tóth, Martin Cooke, Jon Barker
INTERSPEECH2
2016 Evaluating the predictions of objective intelligibility metrics for modified and synthetic speech
Martin Cooke, Cassia Valentini-Botinhao
Comput. Speech Lang.2
2015 A framework for the evaluation of microscopic intelligibility models
abstract
International audience
Ricard Marxer, Martin Cooke, Jon Barker
INTERSPEECH2
2015 A glimpse-based approach for predicting binaural intelligibility with single and multiple maskers in anechoic conditions
abstract
A distortion-weighted glimpsing metric developed for estimating monaural speech intelligibility is extended to predict binaural speech intelligibility in noise. Two aspects of binaural listen- ing, the better ear effect and the binaural advantage, are taken into account in the new metric, which predicts intelligibility using monaural target and masker signals and their location, and is therefore able to provide intelligibility estimates in situations where binaural signals are not readily available. Perceptual listening experiments were conducted to evaluate the predictive power of the proposed metric for speech in the presence of single and multiple maskers in anechoic conditions, for a range of source/masker azimuth combinations. The binaural metric is highly correlated (ρ > 0.9) with listeners’ performance in all conditions tested, but overestimates intelligibility somewhat in conditions where multiple maskers are present and the target speech source location is unknown.
Martin Cooke, Bruno Fazenda, Trevor J. Cox
INTERSPEECH2
2015 A quantitative model of first language influence in second language consonant learning
Martin Cooke, María Luisa García Lecumberri
Speech Commun.2
2014 Generating segmental foreign accent
abstract
For most of us, speaking in a non-native language involves deviating to some extent from native pronunciation norms. However, the detailed basis for foreign accent (FA) remains elusive, in part due to methodological challenges in isolating segmental from suprasegmental factors. The current study examines the role of segmental features in conveying FA through the use of a generative approach in which accent is localised to single consonantal segments. Three techniques are evaluated: the first requires a highly-proficiency bilingual to produce words with isolated accented segments; the second uses cross-splicing of context-dependent consonants from the non-native language into native words; the third employs hidden Markov model synthesis to blend voice models for both languages. Using English and Spanish as the native/non-native languages respectively, listener cohorts from both languages identified words and rated their degree of FA. All techniques were capable of generating accented words, but to differing degrees. Naturally-produced speech led to the strongest FA ratings and synthetic speech the weakest, which we interpret as the outcome of over-smoothing. Nevertheless, the flexibility offered by synthesising localised accent encourages further development of the method.
María Luisa García Lecumberri, Roberto Barra-Chicote, Rubén Pérez Ramón, Junichi Yamagishi, Martin Cooke
INTERSPEECH5
2014 DIAPIX-FL: a symmetric corpus of problem-solving dialogues in first and second languages
abstract
This paper describes a corpus of conversations recorded using an extension of the DiapixUK task: the Diapix Foreign Language corpus (DIAPIX-FL) . English and Spanish native talkers were recorded speaking both English and Spanish. The bidirectionality of the corpus makes it possible to separate language (English or Spanish) from speaking in a first language (L1) or second language (L2). An acoustic analysis was carried out to analyse changes in F0, voicing, intensity, spectral tilt and formants that might result from speaking in an L2. The effect of L1 and nativeness on turn types was also studied. Factors that were investigated were pausing, elongations, and incomplete words. Speakers displayed certain patterns that suggest an on-going process of L2 phonological acquisition, such as the overall percentage of voicing in their speech. Results also show an increase in hesitation phenomena (pauses, elongations, incomplete turns), a decrease in produced speech and speech rate, a reduction of F0 range, raising of minimum F0 when speaking in the non-native language which are consistent with more tentative speech and may be used as indicators of non-nativeness.
Mirjam Wester, María Luisa García Lecumberri, Martin Cooke
INTERSPEECH3
2014 The listening talker: A review of human and algorithmic context-induced modifications of speech
Martin Cooke, Simon King 0001, Maëva Garnier, Vincent Aubanel
Comput. Speech Lang.1
2014 Introduction to the Special Issue on The listening talker: context-dependent speech production and perception
Martin Cooke, Simon King 0001, W. Bastiaan Kleijn, Yannis Stylianou
Comput. Speech Lang.1
2013 Information-preserving temporal reallocation of speech in the presence of fluctuating maskers
abstract
How can speech be retimed so as to maximise its intelligibility in the face of competing speech? We present a general strategy which modifies local speech rate to minimise overlap with a known fluctuating masker. Continuous time-scale factors are derived in an optimisation procedure which seeks to minimise overall energetic masking of the speech by the masker while additionally unmasking those speech regions potentially most important for speech recognition. Intelligibility increases are evaluated with both objective and subjective measures and show significant gains over an unmodified baseline, with larger benefits at lower signal-to-noise ratios. The retiming approach does not lead to benefits for speech mixed with stationary maskers, suggesting that the gains observed for the fluctuating masker are not simply due to durational expansion. Index Terms: speech intelligibility, temporal modification, energetic and informal masking
Vincent Aubanel, Martin Cooke
INTERSPEECH2
2013 Intelligibility-enhancing speech modifications: the hurricane challenge
abstract
Speech output is used extensively, including in situations where correct message reception is threatened by adverse listening conditions. Recently, there has been a growing interest in algorithmic modifications that aim to increase the intelligibility of both natural and synthetic speech when presented in noise. The Hurricane Challenge is the first large-scale open evaluation of algorithms designed to enhance speech intelligibility. Eighteen systems operating on a common data set were subjected to extensive listening tests and compared to unmodified natural and text-to-speech (TTS) baselines. The best-performing systems achieved gains over unmodified natural speech of 4.4 and 5.1 dB in competing speaker and stationary noise respectively, while TTS systems made gains of 5.6 and 5.1 dB over their baseline. Surprisingly, for most conditions the largest gains were observed for noise-independent algorithms, suggesting that performance in this task can be further improved by exploiting information in the masking signal. Index Terms: intelligibility, speech modification, TTS 1.
Martin Cooke, Catherine Mayo, Cassia Valentini-Botinhao
INTERSPEECH1
2013 Elicitation and analysis of a corpus of robust noise-induced word misperceptions in Spanish
abstract
Slips of the ear are of great relevance in the study of how listeners process speech. Our interest in speech misperceptions comes from their value as diagnostic stimuli in evaluating computational models of speech perception in noise. Previous corpora of misperceptions have largely been recorded based on reports of isolated occurrences ‘in the wild’, and consequently are not available for further analysis or replication. The current study involves the elicitation in the laboratory of a corpus of over one thousand robust misperceptions of Spanish words induced by stationary and non-stationary maskers. Misperceptions are analysed into single-phoneme substitutions, insertions and deletions, dual vowel/consonant changes, syllable insertions/deletions, compound reformulations and eccentric cases which defy simple explanation. A novel categorisation scheme based on the interaction between the background and foreground is introduced. The new corpus will permit the evaluation of speech perception models that make detailed predictions of listeners’ responses to specific speech-in-noise tokens.
María Luisa García Lecumberri, Máté Attila Tóth, Martin Cooke
INTERSPEECH4
2013 Evaluating the intelligibility benefit of speech modifications in known noise conditions
Martin Cooke, Catherine Mayo, Cassia Valentini-Botinhao, Yannis Stylianou, Bastian Sauert
Speech Commun.1
2012 Speech Communication in the Wild
Martin Cooke
EACL1
2012 Effects of the availability of visual information and presence of competing conversations on speech production
abstract
International audience
Vincent Aubanel, Martin Cooke, Emma Foster, María Luisa García Lecumberri, Cassie Mayo
INTERSPEECH2
2012 Effect of prosodic changes on speech intelligibility
Catherine Mayo, Vincent Aubanel, Martin Cooke
INTERSPEECH3
2012 Optimised spectral weightings for noise-dependent speech intelligibility enhancement
abstract
Natural or synthetic speech is increasingly used in less-thanideal listening conditions. Maximising the likelihood of correct message reception in such situations often leads to a strategy of loud and repetitive renditions of output speech. An alternative approach is to modify the speech signal in ways which increase intelligibility in noise without increasing signal level or duration. The current study focused on the design of stationary spectral modifications whose effect is to reallocate speech energy across frequency bands. Frequency band weights were selected using a genetic algorithm-based optimisation procedure, with glimpse proportion as the objective intelligibility metric, for a range of noise types and levels. As expected, a clear dependence of noise type and global signal-to-noise ratio on energy reallocation was found. One unanticipated outcome was the consistent discovery of sparse, highly-selective spectral energy weightings, particularly in high noise conditions. In a subjective test using stationary noise and competing speech maskers, listeners were able to identify significantly more words in sentences as a result of spectral weighting, with increases of up to 15 percentage points. These findings suggest that contextdependent speech output can be used to maintain intelligibility at lower sound output levels.
Martin Cooke
INTERSPEECH2
2012 Maximising objective speech intelligibility by local f0 modulation
abstract
We investigated the effect on objective speech intelligibility of scaling the fundamental frequency (f0) of voiced regions in a set of utterances. The frequency scaling was driven by maximising the glimpse proportion in voiced epochs, inspired by musical consonance maximisation techniques. Results show that depending on the energetic masker and the signal to noise ratio, f0 modifications increased the mean glimpse proportion by up to 15 %. On average, lower mean f0 changes resulted in greater glimpse proportions. It was also found that the glimpse proportion could be a good predictor of music consonance. Index Terms: roughness, glimpse proportion, objective speech intelligibility, musical consonance, fundamental frequency.
Julián Villegas, Martin Cooke
INTERSPEECH2
2011 Conversing in the Presence of a Competing Conversation: Effects on Speech Production
abstract
International audience
Vincent Aubanel, Martin Cooke, Julián Villegas, María Luisa García Lecumberri
INTERSPEECH2
2011 The Role of Word-Initial Glottal Stops in Recognizing English Words
abstract
English word-initial vowels in natural continuous speech are optionally preceded by glottal stops or functionally equivalent glottalizations. It may be claimed that these glottal elements disturb the smooth flow of speech. However, they clearly mark word boundaries, which may potentially facilitate speech processing in the brain of the listener. The present study utilizes the word-monitoring paradigm to determine whether listeners react faster to words with or without glottalizations. Three groups of subjects were compared: Czech and Spanish learners of English and native English speakers. The results indicate that perceptual use of glottalization for word segmentation is not entirely governed by universal rules and reflects the mother tongue of the listener as well as the status (L1/L2) of the target language. Index Terms: foreign accent, second language, glottalization, reaction times, word-initial vowels
Maria Paola Bissiri, María Luisa García Lecumberri, Martin Cooke, Jan Volín
INTERSPEECH3
2011 Crowdsourcing for Word Recognition in Noise
abstract
Access to large samples of listeners is an appealing prospect for speech perception researchers, but lack of control over key factors such as listeners ’ linguistic backgrounds and quality of stimulus delivery is a formidable barrier to the application of crowdsourcing. We describe the outcome of a web-based listening experiment designed to discover consistent confusions amongst words presented in noise, alongside an identical task carried out using traditional laboratory methods. Web listeners were graded according based on information they provided as well as via their responses to tokens recognised robustly by a majority of participants. While overall word identification scores even for the best-performing web subset were well below those obtained in the laboratory, word confusions with high levels of cross-listener agreement were obtained nevertheless, suggesting that focused application of crowdsourcing in speech perception can provide useful data for scientific analysis. Index Terms: speech perception, noise, web experiment 1.
Martin Cooke, Jon Barker, María Luisa García Lecumberri, Krzysztof Wasilewski
INTERSPEECH1
2011 Subjective and Objective Evaluation of Speech Intelligibility Enhancement Under Constant Energy and Duration Constraints
abstract
Speakers appear to adopt strategies to improve speech intelligibility for interlocutors in adverse acoustic conditions. Generated speech, whether synthetic, recorded or live, may also benefit from context-sensitive modifications in challenging situations. The current study measured the effect on intelligibility of six spectral and temporal modifications operating under global constraints of constant input-output energy and duration. Reallocation of energy from mid-frequency regions with high local SNR produced the largest intelligibility benefits, while other approaches such as pause insertion or maintenance of a constant segmental SNR actually led to a deterioration in intelligibility. Listener scores correlated only moderately well with recent objective intelligibility estimators, suggesting that further development of intelligibility models is required to improve predictions for modified speech. Index Terms: speech intelligibility, objective measures, energy reallocation
Martin Cooke
INTERSPEECH2
2011 Mtrans: A Multi-Channel, Multi-Tier Speech Annotation Tool
abstract
International audience
Julián Villegas, Martin Cooke, Vincent Aubanel, Marco Aldo Piccolino Boniforti
INTERSPEECH2
2011 Motion strategies for binaural localisation of speech sources in azimuth and distance by artificial listeners
Yan-Chen Lu, Martin Cooke
Speech Commun.2
2010 Energy reallocation strategies for speech enhancement in known noise conditions
abstract
Speech output, whether live, recorded or synthetic, is often employed in difficult listening conditions. Context-sensitive speech modifications aim to promote intelligibility while maintaining quality and listener comfort. The current study used objective measures of intelligibility and quality to compare five energy reallocation strategies operating under equal energy and preserved duration constraints. Results in both stationary and highly-nonstationary backgrounds suggest that time-varying modifications lead to large increases in objective intelligibility, but that speech quality is best preserved by time-invariant modifications. Selective amplification of time-frequency regions with low a priori SNR produced the highest objective intelligibility without severe disruption to quality.
Martin Cooke
INTERSPEECH2
2010 Speech fragment decoding techniques for simultaneous speaker identification and speech recognition
Jon Barker, Ning Ma 0002, André Coy, Martin Cooke
Comput. Speech Lang.4
2010 Monaural speech separation and recognition challenge
Martin Cooke, John R. Hershey, Steven J. Rennie
Comput. Speech Lang.1
2010 Language-independent processing in speech perception: Identification of English intervocalic consonants by speakers of eight European languages
Martin Cooke, María Luisa García Lecumberri, Odette Scharenborg, Wim A. van Dommelen
Speech Commun.1
2010 Preface
Anne Cutler, Martin Cooke, María Luisa García Lecumberri
Speech Commun.2
2010 Non-native speech perception in adverse conditions: A review
María Luisa García Lecumberri, Martin Cooke, Anne Cutler
Speech Commun.2
2010 Binaural Estimation of Sound Source Distance via the Direct-to-Reverberant Energy Ratio for Static and Moving Sources
abstract
One of the principal cues believed to be used by listeners to estimate the distance to a sound source is the ratio of energies along the direct and indirect paths to the receiver. In essence, this “direct-to-reverberant” energy ratio reveals the absolute distance component of the direct energy by normalizing by what is assumed to be distance-independent reverberant energy. Earlier approaches to direct-to-reverberant energy ratio calculation made use of the estimated room impulse response, but these techniques are computationally expensive and inaccurate in practice. This paper proposes and evaluates an alternative approach which uses binaural signals to segregate energy arriving from the estimated direction of the direct source from that arriving from other directions, employing a novel binaural equalization-cancellation technique. The system is integrated with a probabilistic inference framework, particle filtering, to handle the nonstationarity of energy-based measurements. The algorithm is capable of using reverberation to estimate source distance in large rooms with errors of less than 1 m for static sources and 1.5-3.5 m for sources with varying degrees of motion complexity. Model performance can be accounted for largely in terms of a competition between auditory horizon and source energy fluctuation effects.
Yan-Chen Lu, Martin Cooke
IEEE Trans. Speech Audio Process.2
2009 Discovering consistent word confusions in noise
abstract
Listeners make mistakes when communicating under adverse conditions, with overall error rates reasonably well-predicted by existing speech intelligibility metrics. However, a detailed examination of confusions made by a majority of listeners is more likely to provide insights into processes of normal word recognition. The current study measured the rate at which robust misperceptions occurred for highly-confusable words embedded in noise. In a second experiment, confusions discovered in the first listening test were subjected to a range of manipulations designed to help identify their cause. These experiments reveal that while majority confusions are quite rare, they occur sufficiently often to make large-scale discovery worthwhile. Surprisingly few misperceptions were due solely to energetic masking by the noise, suggesting that speech and noise “react” in complex ways which are not well-described by traditional masking concepts. Index Terms: speech perception, word confusions, noise 1.
Martin Cooke
INTERSPEECH1
2009 Speaking in the presence of a competing talker
abstract
How do speakers cope with a competing talker? This study investigated the possibility that speakers are able to retime their contributions to take advantages of temporal fluctuations in the background, reducing any adverse effects for an interlocutor. Speech was produced in quiet, competing talker, modulated noise and stationary backgrounds, with and without a communicative task. An analysis of the timing of contributions relative to the background indicated a significantly reduced chance of overlapping for the modulated noise backgrounds relative to quiet, with competing speech resulting in the least overlap. Strong evidence for an active overlap avoidance strategy is presented. Index Terms: competing talker, speech production, Lombard effect, glimpsing, temporal overlap
Youyi Lu, Martin Cooke
INTERSPEECH2
2009 The contribution of changes in F0 and spectral tilt to increased intelligibility of speech produced in noise
Youyi Lu, Martin Cooke
Speech Commun.2
2008 The interspeech 2008 consonant challenge
abstract
Listeners outperform automatic speech recognition systems at every level, including the very basic level of consonant identification. What is not clear is where the human advantage originates. Does the fault lie in the acoustic representations of speech or in the recognizer architecture, or in a lack of compatibility between the two? Many insights can be gained by carrying out a detailed human-machine comparison. The purpose of the Interspeech 2008 Consonant Challenge is to promote focused comparisons on a task involving intervocalic consonant identification in noise, with all participants using the same training and test data. This paper describes the Challenge, listener results and baseline ASR performance. Index Terms: consonant perception, VCV, humanmachine performance comparisons
Martin Cooke, Odette Scharenborg
INTERSPEECH1
2008 The non-native consonant challenge for european languages
abstract
This paper reports on a multilingual investigation into the effects of different masker types on native and non-native perception in a VCV consonant recognition task. Native listeners outperformed 7 other language groups, but all groups showed a similar ranking of maskers. Strong first language (L1) interference was observed, both from the sound system and from the L1 orthography. Universal acoustic-perceptual tendencies are also at work in both native and non-native sound identifications in noise. The effect of linguistic distance, however, was less clear: in large multilingual studies, listener variables may overpower other factors.
María Luisa García Lecumberri, Martin Cooke, Francesco Cutugno, Mircea Giurgiu, Bernd T. Meyer, Odette Scharenborg, Wim A. van Dommelen, Jan Volín
INTERSPEECH2
2007 L2 consonant identification in noise: cross-language comparisons
abstract
The difficulty of listening to speech in noise is exacerbated\nwhen the speech is in the listener’s L2 rather than L1. In this\nstudy, Spanish and Dutch users of English as an L2 identified\nAmerican English consonants in a constant intervocalic\ncontext. Their performance was compared with that of L1\n(British English) listeners, under quiet conditions and when\nthe speech was masked by speech from another talker or by\nnoise. Masking affected performance more for the Spanish\nlisteners than for the L1 listeners, but not for the Dutch\nlisteners, whose performance was worse than the L1 case to\nabout the same degree in all conditions. There were, however,large differences in the pattern of results across individual consonants, which were consistent with differences in how consonants are identified in the respective L1s.
Anne Cutler, Martin Cooke, María Luisa García Lecumberri, Dennis Pasveer
INTERSPEECH2
2007 Model-driven detection of clean speech patches in noise
abstract
Listeners may be able to recognise speech in adverse conditions by “glimpsing” time-frequency regions where the target speech is dominant. Previous computational attempts to identify such regions have been source-driven, using primitive cues. This paper describes a model-driven approach in which the likelihood of spectro-temporal patches of a noisy mixture representing speech is given by a generative model. The focus is on patch size and patch modelling. Small patches lead to a lack of discrimination, while large patches are more likely to contain contributions from other sources. A “cleanness” measure reveals that a good patch size is one which extends over a quarter of the speech frequency range and lasts for 40 ms. Gaussian mixture models are used to represent patches. A compact representation based on a 2D discrete cosine transform leads to reasonable speech/background discrimination.
Jonathan Laidler, Martin Cooke, Neil D. Lawrence
INTERSPEECH2
2007 Active binaural distance estimation for dynamic sources
abstract
A method for estimating sound source distance in dynamic auditory „scenes‟ using binaural data is presented. The technique requires little prior knowledge of the acoustic environment. It consists of feature extraction for two dynamic distance cues, motion parallax and acoustic τ, coupled with an inference framework for distance estimation. Sequential and nonsequential models are evaluated using simulated anechoic and reverberant spaces. Sequential approaches based on particle filtering more than half the distance estimation error in all conditions relative to the non-sequential models. These results confirm the value of active behaviour and probabilistic reasoning in auditorily-inspired models of distance perception.
Yan-Chen Lu, Martin Cooke, Heidi Christensen
INTERSPEECH2
2007 Modelling speaker intelligibility in noise
Jon Barker, Martin Cooke
Speech Commun.2
2006 Recent advances in speech fragment decoding techniques
abstract
This paper addresses the problem of recognising speech in the presence of a competing speaker. We employ a speech fragment decoding technique that treats segregation and recognition as coupled problems. Data-driven techniques are used to segment a spectro-temporal representation into a set of spectro-temporal fragments, such that each fragment is dominated by one or other of the speech sources. A speech fragment decoder is used which employs missing data techniques and clean speech models to simultaneously search for the set of fragments and the word sequence that best matches the target speaker model. The paper reports recent advances in this technique, and presents an evaluation based on artificially mixed speech utterances. The fragment decoder produces significantly lower error rates than a conventional recogniser, and mimics the pattern of human performance whereby performance increases as the target-masker ratio is reduced below -3 dB. Index Terms: speech recognition, speech separation, simultaneous speech, auditory scene analysis, noise robustness.
Jon Barker, André Coy, Ning Ma 0002, Martin Cooke
INTERSPEECH4
2005 Decoding speech in the presence of other sources
Jon Barker, Martin Cooke, Daniel P. W. Ellis
Speech Commun.2
2004 Introduction to the special issue on the recognition and organization of real-world sound
Martin Cooke, Daniel P. W. Ellis
Speech Commun.1
2001 Robust ASR based on clean speech models: an evaluation of missing data techniques for connected digit recognition in noise
abstract
In this study, techniques for classification with missing or unreliable data are applied to the problem of noise-robustness in Automatic Speech Recognition (ASR). The techniques described make minimal assumptions about any noise background and rely instead on what is known about clean speech. A system is evaluated using the Aurora 2 connected digit recognition task. Using models trained on clean speech we obtain a 65% relative improvement over the Aurora clean training baseline system, a performance comparable with the Aurora baseline for multicondition training.
Jon Barker, Martin Cooke, Phil D. Green
INTERSPEECH2
2001 A tool for automatic feedback on phonemic transcription
Martin Cooke, María Luisa García Lecumberri, John A. Maidment
INTERSPEECH1
2001 The auditory organization of speech and other sources in listeners and computational models
Martin Cooke, Daniel P. W. Ellis
Speech Commun.1
2001 Robust automatic speech recognition with missing and unreliable acoustic data
Martin Cooke, Phil D. Green, Ljubomir Josifovski, Ascension Vizinho
Speech Commun.1
2000 Decoding speech in the presence of other sound sources
abstract
Conventional speech recognition is notoriously vulnerable to additive noise, and even the best compensation methods are defeated if the noise is nonstationary.To address this problem, we propose a new integration of bottom-up techniques to identify 'coherent fragments' of spectro-temporal energy (based on local features), with the top-down hypothesis search of conventional speech recognition, extended to search also across possible assignments of each fragment as speech or interference.Initial tests demonstrate the feasibility of this approach, and achieve a reduction in word error rate of more than 25% relative at 5 dB SNR over stationary noise missing data recognition.
Jon Barker, Martin Cooke, Daniel P. W. Ellis
INTERSPEECH2
2000 Soft decisions in missing data techniques for robust automatic speech recognition
abstract
In previous work we have developed the theory and demonstrated the promise of the Missing Data approach to robust Automatic Speech Recognition. This technique is based on hard decisions as to whether each time-frequency "pixel" is either reliable or unreliable. In this paper we replace these discrete decisions with soft estimates of the probability that each "pixel" is reliable. We adapt the probability calculation to use these estimates as weighting factors for the complementary reliable/unreliable interpretations for each feature vector component. Experiments using the TIDigits connected digit recognition task demonstrate that this technique affords significant performance improvements at low SNRs. 1. INTRODUCTION In previous work [2, 5, 6] we have developed the theory and demonstrated the promise of the Missing Data approach to robust Automatic Speech Recognition. In this technique, spectral-temporal regions uncontaminated by noise are identified and CDHMM recognition methods are ...
Jon Barker, Ljubomir Josifovski, Martin Cooke, Phil D. Green
INTERSPEECH3
2000 A neural network for classification with incomplete data: application to robust ASR
abstract
If the data vector for input to an automatic classifier is incomplete, the optimal estimate for each class probability must be calculated as the expected value of the classifier output. We identify a form of RBF classifier whose expected outputs can easily be evaluated in terms of the original function parameters. We then describe two ways in which this classifier can be applied to robust automatic speech recognition, depending on whether or not the position of missing data is known
Andrew C. Morris, Ljubomir Josifovski, Hervé Bourlard, Martin Cooke, Phil D. Green
INTERSPEECH4
1999 The interactive auditory demonstrations project
abstract
Topics in speech and hearing are well-suited to demonstrations using media other than the printed word. Currently, educators rely largely on passive formats such as the CD collections for general auditory psychophysics [6], auditory scene analysis [2] and cochlear damage [10]. Progress in programming tools and cheap, multimedia hardware now presents the potential to go much further. The interactive auditory demonstrations project aims to provide the user with an environment in which to explore the many phenomena and processes associated with speech and hearing. This promotes a much richer space of parameter manipulation than is possible via passive media. Further, the ability to initiate actions, repeat procedures and benefit from practically any kind of multimodal feedback enables a much wider range of learning possibilities. This paper focusses on the interface issues which are revealed by interactive exploration of the domain. A demonstration of linear prediction is presented to illustrate these issues.
Martin Cooke, Helen Parker, Guy J. Brown, Stuart N. Wrigley
EUROSPEECH1
1999 State based imputation of missing data for robust speech recognition and speech enhancement
Ljubomir Josifovski, Martin Cooke, Phil D. Green, Ascension Vizinho
EUROSPEECH2
1999 Missing data theory, spectral subtraction and signal-to-noise estimation for robust ASR: an integrated study
abstract
In this paper we are presenting a new algorithm for determining the fundamental frequency.The evaluation of pitch is a very difficult problem mainly because of the great variability and irregularity of the speech signals.The algorithm we are presenting is original so far as it relies on the implicit calculation of the autocorrelation of the temporal excitation signal.We have tested our algorithm on the Bagshaw database, created at the Center for Speech Technology Research at Edinburgh, which is primarily dedicated to the evaluation of algorithms estimating the fundamental frequency of speech.The results of our experiments show that our approach is very reliable.
Ascension Vizinho, Phil D. Green, Martin Cooke, Ljubomir Josifovski
EUROSPEECH3
1999 Is the sine-wave speech cocktail party worth attending?
Jon Barker, Martin Cooke
Speech Commun.2
1998 Some solution to the missing feature problem in data classification, with application to noise robust ASR
abstract
We address the theoretical and practical issues involved in automatic speech recognition (ASR) when some of the observation data for the target signal is masked by other signals. Techniques discussed range from simple missing data imputation to Bayesian optimal classification. We have developed the Bayesian approach because this allows prior knowledge to be incorporated naturally into the recognition process, thereby permitting us to go beyond the simple "integrate over missing data" or "marginals" approach reported elsewhere, which we show to be inadequate for dealing with realistic patterns of missing data. After deriving general techniques for recognition with missing data, these techniques are formulated in the context of an HMM based CSR system. This scheme is evaluated under both random and more realistic patterns of missing data, with speech from the DARPA RM corpus and noise from NOISEX. We find that a key problem in real world recognition with missing data is that efficient ASR requires data vector components to be independent, and incomplete data cannot be orthogonalised in the usual way by projection. We show that use of spectral peaks only can provide an effective solution to this problem.
Andrew C. Morris, Martin Cooke, Phil D. Green
ICASSP2
1997 Missing data techniques for robust speech recognition
abstract
In noisy listening conditions, the information available on which to base speech recognition decisions is necessarily incomplete: some spectro-temporal regions are dominated by other sources. We report on the application of a variety of techniques for missing data in speech recognition. These techniques may be based on marginal distributions or on reconstruction of missing parts of the spectrum. Application of these ideas in the resource management task shows a performance which is robust to random removal of up to 80% of the frequency channels, but falls off rapidly with deletions which more realistically simulate masked speech. We report on a vowel classification experiment designed to isolate some of the RM problems for more detailed exploration. The results of this experiment confirm the general superiority of marginals-based schemes, demonstrate the viability of shared covariance statistics, and suggest several ways in which performance improvements on the larger task may be obtained.
Martin Cooke, Andrew C. Morris, Phil D. Green
ICASSP1
1997 Modelling the recognition of spectrally reduced speech
abstract
Progress in robust automatic speech recognition may benefit from a fuller account of the mechanisms and representations used by listeners in processing distorted speech. This paper reports on a number of studies which consider how recognisers trained on clean speech can be adapted to cope with a particular form of spectral distortion, namely reduction of clean speech to sine-wave replicas. Using the Resource Management corpus, the first set of recognition experiments confirm the high information content of sine-wave replicas by demonstrating that such tokens can be recognised at levels approaching those for natural speech if matched conditions apply during training. Further recognition tests show that sine-wave speech can be recognised using natural speech models if a spectral peak representation is employed in concert with occluded speech recognition techniques. 1. INTRODUCTION Clean speech and speech with additive noise have been the primary conditions employed in most ASR studies....
Jon Barker, Martin Cooke
EUROSPEECH2
1995 Auditory scene analysis and hidden Markov model recognition of speech in noise
abstract
We describe a novel paradigm for automatic speech recognition in noisy environments in which an initial stage of auditory scene analysis separates out the evidence for the speech to be recognised from the evidence for other sounds. In general, this evidence will be incomplete, since intruding sound sources will dominate some spectro-temporal regions. We generalise continuous-density hidden Markov model recognition to this 'occluded speech' case. The technique is based on estimating the probability that a Gaussian mixture density distribution for an auditory firing rate map will generate an observation such that the separated components are at their observed values and the remaining components are not greater than their values in the acoustic mixture. Experiments on isolated digit recognition in noise demonstrate the potential of the new approach to yield performance comparable to that of listeners.
Phil D. Green, Martin Cooke, M. D. Crawford
ICASSP2
1994 Handling missing data in speech recognition
abstract
In this paper, we propose a new paradigm for robust ASR based on auditory scene analysis. In previous work, we have shown how models of auditory processing and grouping principles can be used to separate the evidence for a speech signal from arbitrary intrusions. However, this evidence will generally be incomplete since some spectrotemporal regions will be dominated by the other sources. Here, we address the problem of recognising such `occluded' speech. Two investigations are reported: the first applies unsupervised learning and subsequent recognition to spectral vectors with missing components. The second adapts the Viterbi algorithm for HMM-based ASR to the occluded speech case. Both techniques are encouragingly robust: for instance, more than half of the observation vector can be obscured without appreciable deterioration in recognition performance. Additionally, our demonstration that it is possible to learn to recognise speech from partial information suggests a model for the for...
Martin Cooke, Phil D. Green, Malcolm Crawford
ICSLP1
1994 Computational auditory scene analysis
Guy J. Brown, Martin Cooke
Comput. Speech Lang.2
1993 Computational auditory scene analysis: Exploiting principles of perceived continuity
Martin Cooke, Guy J. Brown
Speech Commun.1
1992 A computational model of auditory scene analysis
Guy J. Brown, Martin Cooke
ICSLP2
1986 A computer model of peripheral auditory processing incorporating phase-locking, suppression and adaptation effects
Martin Cooke
Speech Commun.1