EDBT 2026 Demo / reviewers in the wild / expert
Julie Carson-Berndsen
dblp:81/2648
· DBLP profile ↗
52ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-1851-3643ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 2 first-author · 9 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Under the hood: Phonemic Restoration in transformer-based automatic speech recognitionabstractThis study investigates how the automatic speech recognition (ASR) models wav2vec 2.0 large-960h-lv60-self and Whisper large-v3 perform when segment-level signal perturbations (added noise, noisy gaps, and two types of silent gaps) are introduced in English words and pseudowords. We probed the speech embeddings throughout their encoder transformer layers to examine how they encode articulatory features (place and manner of articulation, and voicing). We found that wav2vec 2.0 was more successful than Whisper at restoring perturbed segments across conditions. For wav2vec 2.0 embeddings, classification accuracy was higher in words than in pseudowords. The articulatory features encoding of both ASR models was least disturbed by added noise, and most disturbed by noisy gaps, with silent gaps falling in between. Coarticulatory cues improved classification of articulatory features and classification accuracy increased from early to late layers for both models. Among the examined target sounds, [n] stood out from [m], , and [l], as it was classified particularly well under all conditions. We compare ASR performance to the Phonemic Restoration Effect in human speech perception and discuss potential reasons for the performance differences between the two ASR models. This approach aims to foster a better understanding of otherwise opaque systems. Iona Gessinger, Erfan A. Shams, Julie Carson-Berndsen |
Comput. Speech Lang. | 3 |
| 2025 | Improving Linguistic Diversity of Large Language Models with Possibility Exploration Fine-Tuning
Long Mai, Julie Carson-Berndsen |
INTERSPEECH | 2 |
| 2024 | Following the Embedding: Identifying Transition Phenomena in Wav2vec 2.0 Representations of Speech AudioabstractAlthough transformer-based models have improved the state-of-the-art in speech recognition, it is still not well understood what information from the speech signal these models encode in their latent representations. This study investigates the potential of using labelled data (TIMIT) to probe wav2vec 2.0 embeddings for insights into the encoding and visualisation of speech signal information at phone boundaries. Our experiment involves training probing models to detect phone-specific articulatory features in the hidden layers based on IPA classifications. Furthermore, we propose an analysis framework for visualising the probabilities of the detected articulatory features in every layer and frame vector. Our primary focus is to probe and better understand the structure of speech signal information in the embeddings learned by unsupervised transformers, with a view to contributing to more explainable speech processing systems. Patrick Cormac English, Erfan A. Shams, John D. Kelleher, Julie Carson-Berndsen |
ICASSP | 4 |
| 2024 | Enhancing Conversation Smoothness in Language Learning Chatbots: An Evaluation of GPT4 for ASR Error CorrectionabstractThe integration of natural language processing (NLP) technologies into educational applications has shown promising results, particularly in the language learning domain. Many spoken open-domain chatbots have been used as speaking partners, helping language learners improve their language skills. However, one of the significant challenges is the high word-error-rate (WER) when recognising non-native/non-fluent speech, which interrupts conversation flow and leads to disappointment for learners. This paper explores the use of GPT4 for ASR error correction in conversational settings. In addition to WER, we propose to use semantic textual similarity (STS) and next response sensibility (NRS) metrics to evaluate the impact of correction models on conversation smoothness. We find that transcriptions corrected by GPT4 lead to higher conversation smoothness, despite an increase in WER. GPT4 also outperforms standard error correction methods without the need for in-domain training data. Long Mai, Julie Carson-Berndsen |
ICASSP | 2 |
| 2024 | Searching for Structure: Appraising the Organisation of Speech Features in wav2vec 2.0 Embeddings
Patrick Cormac English, John D. Kelleher, Julie Carson-Berndsen |
INTERSPEECH | 3 |
| 2024 | PhoneViz: exploring alignments at a glance
Margot Masson, Erfan A. Shams, Iona Gessinger, Julie Carson-Berndsen |
INTERSPEECH | 4 |
| 2024 | Are Articulatory Feature Overlaps Shrouded in Speech Embeddings?
Erfan A. Shams, Iona Gessinger, Patrick Cormac English, Julie Carson-Berndsen |
INTERSPEECH | 4 |
| 2023 | Discovering Phonetic Feature Event Patterns in Transformer Embeddings
Patrick Cormac English, John D. Kelleher, Julie Carson-Berndsen |
INTERSPEECH | 3 |
| 2022 | Unsupervised domain adaptation for speech recognition with unsupervised error correctionabstractThe transcription quality of automatic speech recognition (ASR) systems degrades significantly when transcribing audios coming from unseen domains.We propose an unsupervised error correction method for unsupervised ASR domain adaption, aiming to recover transcription errors caused by domain mismatch.Unlike existing correction methods that rely on transcribed audios for training, our approach requires only unlabeled data of the target domains in which a pseudo-labeling technique is applied to generate correction training samples.To reduce over-fitting to the pseudo data, we also propose an encoder-decoder correction model that can take into account additional information such as dialogue context and acoustic features.Experiment results show that our method obtains a significant word error rate (WER) reduction over non-adapted ASR systems.The correction model can also be applied on top of other adaptation approaches to bring an additional improvement of 10% relatively. Long Mai, Julie Carson-Berndsen |
INTERSPEECH | 2 |
| 2022 | Production characteristics of obstruents in WaveNet and older TTS systems
Ayushi Pandey, Sébastien Le Maguer, Julie Carson-Berndsen, Naomi Harte |
INTERSPEECH | 3 |
| 2021 | The Influence of Regional Pronunciation Variation on Children's Spelling and the Potential Benefits of Accent Adapted SpellcheckersabstractA child who is unfamiliar with the correct spelling of a word often employs a "sound it out" approach: breaking the word down into its constituent sounds and then choosing letters to represent the identified sounds.This often results in a misspelling that is orthographically very different to the intended target.Recently, efforts have been made to develop phonetic based spellcheckers to tackle the more deviant nature of children's misspellings.However, little work has been done to investigate the potential of spelling correction tools that incorporate regional pronunciation variation.If a child must first identify the sounds that make up a word, it stands to reason their pronunciation would influence this process.We investigate this hypothesis along with the feasibility and potential benefits of adapting spelling correction tools to more specific language variantsparticularly Irish Accented English.We use misspelling data from schoolchildren across Ireland to adapt an existing English phoneticbased spellchecker and demonstrate improvements in performance.These results not only prompt consideration of language varieties in the development of spellcheckers but also contribute to existing literature on the role of regional accent in the acquisition of writing proficiency. Emma O'Neill, Joe Kenny, Anthony Ventresque, Julie Carson-Berndsen |
CoNLL | 4 |
| 2019 | The Effect of Phoneme Distribution on Perceptual Similarity in EnglishabstractThe 20th Annual Conference of the International Speech Communication Association (Interspeech 2019), Graz, Austria, 15-19 September 2019 Emma O'Neill, Julie Carson-Berndsen |
INTERSPEECH | 2 |
| 2016 | Enhancing Data-Driven Phone Confusions Using Restricted Recognition
Mark Kane, Julie Carson-Berndsen |
INTERSPEECH | 2 |
| 2015 | The effect of soft, modal and loud voice levels on entrainment in noisy conditionsabstractConversation partners have a tendency to adapt their vocal in- tensity to each other and to other social and environmental fac- tors. A socially adequate vocal intensity level by a speech syn- thesiser that goes beyond mere volume adjustment is highly de- sirable for a rewarding and successful human-machine or ma- chine mediated human-human interaction. This paper examines the interaction of the Lombard effect and speaker entrainment in a controlled experiment conducted with a confederate inter- locutor. The interlocutor was asked to maintain either a soft, a modal or a loud voice level during the dialogues. Through half of the trials, subjects were exposed to a cocktail party noise through headphones. The analytical results suggest that both the background noise and the interlocutor’s voice level affect the dynamics of speaker entrainment. Speakers appear to still en- train to the voice level of their interlocutor in noisy conditions, though to a lesser extent, as strategies of ensuring intelligibility affect voice levels as well. These findings could be leveraged in spoken dialogue systems and speech generating devices to help choose a vocal effort level for the synthetic voice that is both intelligible and socially suited to a specific interaction. Éva Székely, Mark T. Keane, Julie Carson-Berndsen |
INTERSPEECH | 3 |
| 2014 | Predicting synthetic voice style from facial expressions. An application for augmented conversations
Éva Székely, Shannon Hennig, João P. Cabral, Julie Carson-Berndsen |
Speech Commun. | 5 |
| 2012 | Exemplar-based pitch contour generation using DOP for syntactic tree decompositionabstractThe generation of a pitch contour from linguistic information has long been recognised as a requirement for natural sounding speech synthesis. This paper investigates the use of an exemplar-based model for pitch contour generation. The main drawbacks of previous unit selection-based approaches for pitch contour generation is determining the size of the unit, and to guarantee that only prosodic and linguistically related units will be selected. The work presented in this paper overcomes these drawbacks by using only prosodic-syntactic correlated data, and a dynamic unit size model using data-oriented parsing. An AB comparison perceptual test showed 58% preference for the exemplar-based model, 25% for a HTS model, and 17% find both the same in terms of naturalness and pitch. In a MOS test, exemplar-based model achieved higher scores than that the HTS model achieved. Mohamed Abou-Zleikha, Peter Cahill, Julie Carson-Berndsen |
ICASSP | 3 |
| 2012 | Detecting a targeted voice style in an audiobook using voice quality featuresabstractAudiobooks are known to contain a variety of expressive speaking styles that occur as a result of the narrator mimicking a character in a story, or expressing affect. An accurate modeling of this variety is essential for the purposes of speech synthesis from an audiobook. Voice quality differences are important features characterizing these different speaking styles, which are realized on a gradient and are often difficult to predict from the text. The present study uses a parameter characterizing breathy to tense voice qualities using features of the wavelet transform, and a measure for identifying creaky segments in an utterance. Based on these features, a combination of supervised and unsupervised classification is used to detect the regions in an audiobook, where the speaker changes his regular voice quality to a particular voice style. The target voice style candidates are selected based on the agreement of the supervised classifier ensemble output, and evaluated in a listening test. Éva Székely, John Kane 0002, Stefan Scherer, Christer Gobl, Julie Carson-Berndsen |
ICASSP | 5 |
| 2012 | Rapidly Testing the Interaction Model of a Pronunciation Training System via Wizard-of-Oz
João P. Cabral, Mark Kane, Mohamed Abou-Zleikha, Éva Székely, Amalia Zahra, Kalu U. Ogbureke, Peter Cahill, Julie Carson-Berndsen, Stephan Schlögl |
LREC | 9 |
| 2012 | Evaluating expressive speech synthesis from audiobook corpora for conversational phrases
Éva Székely, João P. Cabral, Mohamed Abou-Zleikha, Peter Cahill, Julie Carson-Berndsen |
LREC | 5 |
| 2012 | English to Indonesian Transliteration to Support English Pronunciation Practice
Amalia Zahra, Julie Carson-Berndsen |
LREC | 2 |
| 2012 | Synthesizing expressive speech from amateur audiobook recordingsabstractFreely available audiobooks are a rich resource of expressive speech recordings that can be used for the purposes of speech synthesis. Natural sounding, expressive synthetic voices have previously been built from audiobooks that contained large amounts of highly expressive speech recorded from a professionally trained speaker. The majority of freely available audiobooks, however, are read by amateur speakers, are shorter and contain less expressive (less emphatic, less emotional, etc.) speech both in terms of quality and quantity. Synthesizing expressive speech from a typical online audiobook therefore poses many challenges. In this work we address these challenges by applying a method consisting of minimally supervised techniques to align the text with the recorded speech, select groups of expressive speech segments and build expressive voices for hidden Markov-model based synthesis using speaker adaptation. Subjective listening tests have shown that the expressive synthetic speech generated with this method is often able to produce utterances suited to an emotional message. We used a restricted amount of speech data in our experiment, in order to show that the method is generally applicable to most typical audiobooks widely available online. Éva Székely, Tamás Gábor Csapó, Bálint Tóth, Péter Mihajlik, Julie Carson-Berndsen |
SLT | 5 |
| 2011 | Automatic Rule Extraction for Modeling Pronunciation Variation
Julie Carson-Berndsen |
CICLing (2) | 2 |
| 2011 | Multiple Source Phoneme Recognition Aided by Articulatory Features
Mark Kane, Julie Carson-Berndsen |
IEA/AIE (2) | 2 |
| 2011 | Correlating Text with ProsodyabstractThe prediction of prosody from text information has long been recognised as a requirement for natural sounding speech synthesis. While an examination of the relationship between text information and prosody typically focuses on the role of accent, duration and phrasing both from a statistical and rulebased perspective, this paper investigates the correlation between the similarities calculated with respect to text information and those calculated with respect to prosody from an exemplarbased perspective. Two text features are examined, the syntactic tree and the dependency tree, along with two prosody features, pitch and intensity. The work in this paper investigates 1) the correlation between text information and prosody information 2) the conditional membership probability between text information and prosodic information, and 3) the effect of the number of exemplars on the conditional membership probability. Index Terms: prosody prediction, syntactic tree, dependency tree, prosody text correlation, prosody text similarity Mohamed Abou-Zleikha, Julie Carson-Berndsen |
INTERSPEECH | 2 |
| 2011 | Evaluation of Glottal Epoch Detection Algorithms on Different Voice TypesabstractAccording to the source-filter model of speech production, speech can be represented by passing the excitation signal through the vocal tract filter. The epoch or instant of maximum excitation corresponds to the glottal closure instant. Several speech processing applications require robust epoch detection but this can be a difficult task. Although state-of-the-art epoch estimation methods can produce reliable results, they are generally evaluated using speech recorded with a neutral voice quality (modal voice). This paper reviews and evaluates six popular algorithms for the calculation of glottal closure instants on speech spoken with modal voice and seven additional voice qualities. Results show that the performance of each method is affected by the voice type and that some methods perform better than others for each voice quality. Index Terms: GCI, epoch detection, glottal source João P. Cabral, John Kane 0002, Christer Gobl, Julie Carson-Berndsen |
INTERSPEECH | 4 |
| 2011 | Clustering Expressive Speech Styles in Audiobooks Using Glottal Source ParametersabstractA great challenge for text-to-speech synthesis is to produce expressive speech. The main problem is that it is difficult to synthesise high-quality speech using expressive corpora. With the increasing interest in audiobook corpora for speech synthesis, there is a demand to synthesise speech which is rich in prosody, emotions and voice styles. In this work, Self-Organising Feature Maps (SOFM) are used for clustering the speech data using voice quality parameters of the glottal source, in order to map out the variety of voice styles in the corpus. Subjective evaluation showed that this clustering method successfully separated the speech data into groups of utterances associated with different voice characteristics. This work can be applied in unitselection synthesis by selecting appropriate data sets to synthesise utterances with specific voice styles. It can also be used in parametric speech synthesis to model different voice styles separately. Index Terms: expressive speech, voice quality, audiobook, speech synthesis Éva Székely, João P. Cabral, Peter Cahill, Julie Carson-Berndsen |
INTERSPEECH | 4 |
| 2011 | Phonetic Representation-Based Speech Translation
Jie Jiang 0002, Julie Carson-Berndsen, Peter Cahill, Andy Way |
MTSummit | 3 |
| 2010 | Lattice Score Based Data Cleaning for Phrase-Based Statistical Machine Translation
Jie Jiang 0002, Julie Carson-Berndsen, Andy Way |
EAMT | 2 |
| 2010 | Framework for cross-language automatic phonetic segmentationabstractAnnotation of large multilingual corpora remains a challenge to the data-driven approach to speech research, especially for under-resourced languages. This paper presents cross-language automatic phonetic segmentation using Hidden Markov Models (HMMs). The underlying notion is segmentation based on articulation (manner and place) so as to provide extensive models that will be applicable across languages. A test on the Appen Spanish speech corpus gives phone recognition accuracy of 61.15% when bootstrapped with acoustic models trained on the TIMIT as compared with a baseline result of 54.63% for flat start initialization of the monophone models. Kalu U. Ogbureke, Julie Carson-Berndsen |
ICASSP | 2 |
| 2010 | Hidden Markov models with context-sensitive observations for grapheme-to-phoneme conversionabstractHidden Markov models (HMMs) have proven useful in various aspects of speech technology from automatic speech recognition through speech synthesis, speech segmentation and grapheme-to-phoneme conversion to part-of-speech tagging. Traditionally, context is modelled at the hidden states in the form of context-dependent models. This paper constitutes an extension to this approach; the underlying concept is to model context at the observations for HMMs with discrete observations and discrete probability distributions. The HMMs emit context-sensitive discrete observations and are evaluated with a grapheme-to-phoneme conversion system. Index Terms: hidden Markov models, grapheme-to-phoneme conversion, machine learning Kalu U. Ogbureke, Peter Cahill, Julie Carson-Berndsen |
INTERSPEECH | 3 |
| 2010 | Muse: An open source speech technology research platformabstractThis paper introduces the open source muster speech engine (Muse) for speech technology research. The Muse platform abstracts common data types and software as used by speech technology researchers. It is designed to assist researchers in making repeatable experiments that are not hard coded to a specific platform, language, algorithm, or corpus. It contains a script language and a shell where users can interact with various components. The presentation of this paper will be accompanied by a demo at the SLT workshop. Peter Cahill, Julie Carson-Berndsen |
SLT | 2 |
| 2009 | Using same-language machine translation to create alternative target sequences for text-to-speech synthesisabstractModern speech synthesis systems attempt to produce\nspeech utterances from an open domain of words. In some situations, the synthesiser will not have the appropriate units to pronounce some words or phrases accurately but it still must attempt to pronounce them. This paper presents a hybrid machine translation and unit selection speech synthesis system. The machine translation system was trained with English as the source and target language. Rather than the synthesiser only saying the input text as would happen in conventional synthesis systems, the synthesiser may say an alternative utterance with the same\nmeaning. This method allows the synthesiser to overcome the\nproblem of insufficient units in runtime. Peter Cahill, Jinhua Du, Andy Way, Julie Carson-Berndsen |
INTERSPEECH | 4 |
| 2009 | Improving initial boundary estimation for HMM-based automatic phonetic segmentationabstractThis paper presents an approach to boundary estimation for automatic segmentation of speech given a phone (sound) sequence. The technique presented represents an extension to existing approaches to Hidden Markov Model based automatic segmentation which modifies the topology of the model to control for duration. An HMM system trained with this modified topology places 77.10%, 86.72 % and 91.15 % of the boundaries, on the TIMIT speech test corpus annotations, within 10, 15 and 20 ms respectively as compared with manual annotations. This represents an improvement over the baseline result of 70.99%, 83.50 % and 89.18 % for initial boundary estimation. Index Terms: automatic phonetic segmentation, hidden markov models, gaussian mixture models Kalu U. Ogbureke, Julie Carson-Berndsen |
INTERSPEECH | 2 |
| 2007 | Articulatory acoustic feature applications in speech synthesisabstractThe quality of unit selection speech synthesisers depends significantly on the content of the speech database being used. In this paper a technique is introduced that can highlight mispronunciations and abnormal units in the speech synthesis voice database through the use of articulatory acoustic feature extraction to obtain an additional layer of annotation. A set of articulatory acoustic feature classifiers help minimise the selection of inappropriate units in the speech database and are shown to significantly improve the word error rate of a diphone synthesiser. Index Terms: speech synthesis, unit selection, articulatory acoustic feature extraction Peter Cahill, Daniel Aioanei, Julie Carson-Berndsen |
INTERSPEECH | 3 |
| 2006 | Diagnostic Evaluation of Phonetic Feature Extraction Engines: A Case Study with the Time Map Model
Daniel Aioanei, Julie Carson-Berndsen, Supphanat Kanokphara |
IEA/AIE | 2 |
| 2006 | Comparative Study: HMM and SVM for Automatic Articulatory Feature Extraction
Supphanat Kanokphara, Jan Macek, Julie Carson-Berndsen |
IEA/AIE | 3 |
| 2005 | Feature-Table-Based Automatic Question Generation for Tree-Based State Tying: A Practical Implementation
Supphanat Kanokphara, Julie Carson-Berndsen |
IEA/AIE | 2 |
| 2005 | An agent-based framework for speech investigationabstractThe 9th International Conference on Speech Communication and Technology (Interspeech 2005), Lisbon, Portugal, September, 2005 Michael Walsh 0001, Gregory M. P. O'Hare, Julie Carson-Berndsen |
INTERSPEECH | 3 |
| 2004 | A Multilingual Phonological Resource Toolkit for Ubiquitous Speech Technology
Daniel Aioanei, Julie Carson-Berndsen, Anja Geumann, Robert Kelly, Moritz Neugebauer |
LREC | 2 |
| 2004 | Acquiring Reusable Multilingual Phonotactic Resources
Julie Carson-Berndsen, Robert Kelly |
LREC | 1 |
| 2003 | Multi-linear HMM based system for articulatory feature extractionabstractA novel system for automatic articulatory feature extraction has been developed. The system defines an autosegmental multi-linear representation of features and uses multiple hidden Markov model based recognisers to extract these feature classes. Overlap and precedence relations among features on different tiers can be extracted and then presented to a phonological parser for further recognition. The system thus accounts for coarticulation phenomena. The system was implemented using a novel modification of the HTK (HMM toolkit) which allows it to perform multi-thread multi-feature recognition. The system performance is extremely promising. Among the highest accuracies achieved are 98% for vowels and 93% for rhotic sounds. Current work investigates interdependencies of extracting different feature types. Tarek Abu-Amer, Julie Carson-Berndsen |
ICASSP (2) | 2 |
| 2003 | A Multi-Agent Computational Linguistic Approach to Speech Recognition
Michael Walsh 0001, Robert Kelly, Gregory M. P. O'Hare, Julie Carson-Berndsen, Tarek Abu-Amer |
IJCAI | 4 |
| 2003 | HARTFEX: a multi-dimensional system of HMM based recognisers for articulatory features extractionabstractHARTFEX is a novel system that employs several tiers of HMMs recognisers that work in parallel to extract multidimentions of articulatory features. The features segments on the different tiers overlap to account for the coarticulation phenomena. The overlap and precedence relation among features are applied to a phonological parser for further processing. HARTFEX system is built on a modified version of HTK toolkit that allows it to perform multi-thread multi-feature recognition. The system testing results are highly promising. The recognition accuracy for vowel is 98\% and for rhotic is 93%. Current work investigates inherited interdependencies of extracting different feature sets. Tarek Abu-Amer, Julie Carson-Berndsen |
INTERSPEECH | 2 |
| 2003 | Computational Linguistic Motivations for a Finite-State Machine Hierarchy
Robert Kelly, Julie Carson-Berndsen |
CIAA | 2 |
| 2001 | A testbed for developing multilingual phonotactic descriptionsabstractThis paper presents a testbed for developing multilingual phonotactic descriptions that employs finite state methods to represent the phonotactics of one or more languages. The motivation for this work is to make an extensive range of phonotactic descriptions of varying granularity available for speech technology applications. We discuss the design of the phonotactic testbed and how various modules may be used to generate finite state phonotactic descriptions. We provide an example multilingual application drawn from a partial sample of onset clusters spanning four language families, demonstrating how the commonalities of a broad spectrum of languages can be expressed using individual and generic phonotactic automata. We then discuss how these representations are extended via a three-tiered model to provide the basis for the feature- and event-based phonotactic automata. Simone Ashby, Julie Carson-Berndsen, Gina Joue |
INTERSPEECH | 2 |
| 2001 | Defining constraints for multilinear speech processingabstractThis paper presents a constraint model for the interpretation of multilinear representations of speech utterances which can provide important fine-grained information for speech recognition applications. The model uses explicit structural constraints specifying overlap and precedence relations between features in both the phonological and the phonetic domains in order to recognise well-formed syllable structures. In the phonological domain, these constraints together form a complete phonotactic description of the language, while in the phonetic domain, the constraints define the internal structure of phonologial features based on phonetic realisations. The constraints are enhanced by a constraint relaxation procedure to cater for underspecified input and allows output representations to be extrapolated based on the phonetic and phonological information contained in the constraints and the rankings which have been assigned to them. This approach thus addresses issues of robustness in sp eech recognition. Julie Carson-Berndsen, Michael Walsh 0001 |
INTERSPEECH | 1 |
| 2001 | An embodiment paradigm for speech recognition systemsabstractThe problems of conventional speech recognition approaches include incomplete linguistic knowledge and inability to deal with underspecification. These issues can be addressed by understanding the constraints of speech to predict speech tendencies. We believe that understanding what constraints exist requires an embodied view of speech and that the traditional disembodied view of speech is the fundamental limitation on the robustness of many speech systems. We argue that viewing speech as a form of embodied cognition, or within context of its production and use, provides important insights in speech structure and speech recognition. In making this claim, this paper briefly outlines a strongly embodied account of cognition and develops from that an embodiment paradigm for speech recognition. The embodiment paradigm proposed leads to both an explanatory and descriptive account of linguistic structure. It simplifies the view of speech structure for automatic speech recognisers, by considering only the most directly relevant motivations or constraints influencing communication and thus speech. Gina Joue, Julie Carson-Berndsen |
INTERSPEECH | 2 |
| 1999 | A generic lexicon tool for word model definition in multimodal applicationsabstractThis paper describes a generic lexicon tool which uses lexical representations and finite state transducers enhanced by arithmetic operations in DATR to generate individual output formats from a general phonological feature based representation.The tool was developed in connection with the lexicon component of a diagnostic evaluation toolkit, BEETLE, for a linguistic word recognition system.This lexicon is used online by the system to distinguish between actual and potential syllables and used offline for evaluation purposes with respect to a particular corpus.Rather than design a syllable lexicon which can only be used for these two tasks, it was decided to develop a more generic lexicon from which specific lexica can be generated on the fly; BEETLE is, therefore, only one application of the generic lexicon tool which can be used to generate output formats for other speech technology and multimodal applications. Julie Carson-Berndsen |
EUROSPEECH | 1 |
| 1994 | Phoneme recognition using acoustic eventsabstractThis paper presents a new approach to phoneme recognition using nonsequential sub{phoneme units.These units are called acoustic events and are phonologically meaningful as well as recognizable from speech signals.Acoustic events form a phonologically incomplete representation as compared to distinctive features.This problem may partly be overcome by incorporating phonological constraints.Currently, 24 binary events describing manner and place of articulation, vowel quality and voicing are used to recognize all German phonemes.Phoneme recognition in this paradigm consists of two steps: After the acoustic events have been determined from the speech signal, a phonological parser is used to generate syllable and phoneme hypotheses from the event lattice.Results obtained on a speaker{dependent corpus are presented. Kai Hübener, Julie Carson-Berndsen |
ICSLP | 2 |
| 1992 | Event Relations At The Phonetics/Phonology Interface
Julie Carson-Berndsen, Dafydd Gibbon |
COLING | 1 |
| 1990 | Phonological Processing of Speech Variants
Julie Carson-Berndsen |
COLING | 1 |
| 1988 | Unification and transduction in computational phonology
Julie Carson-Berndsen |
COLING | 1 |