EDBT 2026 Demo / reviewers in the wild / expert
Sharon Goldwater
dblp:75/5799
· DBLP profile ↗
74ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0002-7298-0947ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 7 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Measuring the Effects of Visual Salience in Human and AI Descriptions with Image EditingabstractHow does our perception of the world influence the way we talk about it? Psycholinguistic studies have investigated whether visual salience correlates with entity mention and ordering, but often disregarded its effect on grammar or relied on simplistic images or artificial cues. In this study, we explore the use of generative AI to better control for salience in visual stimuli while keeping them realistic, and to serve as a proxy for human participants in studying how different types of salience impact image descriptions.We consider three salience types: perceptual (e.g. relative size in the image), inherent (e.g. animacy), and relational (e.g. human–object interaction). We first analyze human- and AI-generated captions for natural images to examine how salience correlates with how early, and in what grammatical role, an entity is mentioned. We find strong correlations between models and humans in this observational study, justifying the use of AI models alone in a further causal study. For this second study, we created datasets composed of pairs of images, where we used an image-editing model to intervene on the salience of a target entity. We show that relational and perceptual salience lead to the entity being mentioned earlier in captions and being mapped to more prominent grammatical roles. The magnitude of this effect varies across entity types, with animate entities (high inherent salience) showing a particularly distinct pattern. Nina Gregorio, Edoardo Maria Ponti, Sharon Goldwater |
CoNLL | 3 |
| 2026 | When transformers learn "impossible" languages, what do they learn?abstractRecent work suggests that transformer language models show a bias towards human languages over unnatural ("impossible") languages argued to be unacquirable by humans.However, this literature has largely based these claims on differences in sample efficiency and test-set perplexity, rather than on direct evaluations of the linguistic capacities that could plausibly explain non-attestation in human languages.We evaluate two theoretically motivated linking hypotheses: impossibility arising from deficiencies in grammatical sensitivity or generative production.Using GPT-2 style models trained on perturbed "impossible" variants of English, we measure sensitivity to grammaticality using BLiMP minimal pairs, finding that model performance exhibits only gradual degredation, mediated by the language's information locality.In contrast, these models exhibited pronounced failures in generation, producing substantially fewer high-quality sentences at longer lengths.Together, these results suggest generative deficiency and transmission failures as a plausible linking hypothesis between language model behaviour and non-attestation of impossible languages. Ram Janarthan, Coleman Haley, Sharon Goldwater |
CoNLL | 3 |
| 2026 | A framework for analyzing concept representations in neural modelsabstractUnderstanding how neural models represent human-interpretable concepts is challenging.Prior work has explored linear concept subspaces from diverse perspectives, such as probing and concept erasure.We introduce a unified framework to study these subspaces along two axes: containment, which tests if a concept is fully represented in a subspace but not outside it, and disentanglement, which tests for isolation from other concepts.In experiments on both text and speech models, we first highlight that concept subspaces may not be uniquely determined, and discuss the implications for concept subspace analysis.Then, we compare properties of concept subspaces estimated using five estimators, proposed in different communities.We find that (1) the choice of estimator impacts the containment and disentanglement properties; (2) the state-of-theart concept erasure method, LEACE, performs well on both testing axes, but still struggles to generalize to unseen data; and (3) in HuBERT speech representations, phone information is both contained and disentangled from speaker information, while speaker information is hard to contain in a compact subspace, despite being disentangled from phones. 1 1 We release the source code at https://github.com/ burin-n/concept_space. Burin Naowarat, Hao Tang 0002, Sharon Goldwater |
CoNLL | 3 |
| 2025 | The Cross-linguistic Role of Animacy in Grammar StructuresabstractAnimacy is a semantic feature of nominals and follows a hierarchy: personal pronouns > human > animate > inanimate.In several languages, animacy imposes hard constraints on grammar.While it has been argued that these constraints may emerge from universal soft tendencies, it has been difficult to provide empirical evidence for this conjecture due to the lack of data annotated with animacy classes.In this work, we first propose a method to reliably classify animacy classes of nominals in 11 languages from 5 families, leveraging multilingual large language models (LLMs) and word sense disambiguation datasets.Then, through this newly acquired data, we verify that animacy displays consistent cross-linguistic tendencies in terms of preferred morphosyntactic constructions, although not always in line with received wisdom: animacy in nouns correlates with the alignment role of agent, early positions in a clause, and syntactic pivot (e.g., for relativisation), but not necessarily with grammatical subjecthood.Furthermore, the behaviour of personal pronouns in the hierarchy is idiosyncratic as they are rarely plural and relativised, contrary to high-animacy nouns. Nina Gregorio, Matteo Gay, Sharon Goldwater, Edoardo Maria Ponti |
ACL (1) | 3 |
| 2025 | Revisiting Common Assumptions about Arabic Dialects in NLPabstractArabic has diverse dialects, where one dialect can be substantially different from the others.In the NLP literature, some assumptions about these dialects are widely adopted (e.g., "Arabic dialects can be grouped into distinguishable regional dialects") and are manifested in different computational tasks such as Arabic Dialect Identification (ADI).However, these assumptions are not quantitatively verified.We identify four of these assumptions and examine them by extending and analyzing a multi-label dataset, where the validity of each sentence in 11 different country-level dialects is manually assessed by speakers of these dialects.Our analysis indicates that the four assumptions oversimplify reality, and some of them are not always accurate.This in turn might be hindering further progress in different Arabic NLP tasks. Amr Keleg, Sharon Goldwater, Walid Magdy |
ACL (1) | 2 |
| 2025 | Visual groundedness as an organizing principle for word class: Evidence from Japanese
Coleman Haley, Sharon Goldwater |
CogSci | 2 |
| 2025 | Effective Context in Neural Speech ModelsabstractModern neural speech models benefit from having longer context, and many approaches have been proposed to increase the maximum context a model can use. However, few have attempted to measure how much context these models actually use, i.e., the effective context. Here, we propose two approaches to measuring the effective context, and use them to analyze different speech Transformers. For supervised models, we find that the effective context correlates well with the nature of the task, with fundamental frequency tracking, phone classification, and word classification requiring increasing amounts of effective context. For self-supervised models, we find that effective context increases mainly in the early layers, and remains relatively short---similar to the supervised phone model. Given that these models do not use a long context during prediction, we show that HuBERT can be run in streaming mode without modification to the architecture and without further fine-tuning. Yen Meng, Sharon Goldwater, Hao Tang 0002 |
INTERSPEECH | 2 |
| 2025 | A Grounded Typology of Word ClassesabstractColeman Haley, Sharon Goldwater, Edoardo Ponti. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Coleman Haley, Sharon Goldwater, Edoardo Maria Ponti |
NAACL (Long Papers) | 2 |
| 2024 | A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
Oli Danyi Liu, Hao Tang 0002, Naomi Feldman, Sharon Goldwater |
CogSci | 4 |
| 2024 | Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representationsabstractSelf-supervised speech representations can hugely benefit downstream speech technologies, yet the properties that make the museful are still poorly understood. Two candidate properties related to the geometry of the representation space have been hypothesized to correlate well with downstream tasks: (1) the degree of orthogonality between the subspaces spanned by the speaker centroids and phone centroids, and (2) the isotropy of the space, i.e., the degree to which all dimensions are effectively utilized. To study them, we introduce a new measure, Cumulative Residual Variance (CRV), which can be used to assess both properties. Using linear classifiers for speaker and phone ID to probe the representations of six different self-supervised models and two untrained baselines, we ask whether either orthogonality or isotropy correlate with linear probing accuracy. We find that both measures correlate with phonetic probing accuracy, though our results on isotropy are more nuanced. Mukhtar Mohamed, Oli Danyi Liu, Hao Tang 0002, Sharon Goldwater |
INTERSPEECH | 4 |
| 2023 | ALDi: Quantifying the Arabic Level of Dialectness of TextabstractTranscribed speech and user-generated text in Arabic typically contain a mixture of Modern Standard Arabic (MSA), the standardized language taught in schools, and Dialectal Arabic (DA), used in daily communications.To handle this variation, previous work in Arabic NLP has focused on Dialect Identification (DI) on the sentence or the token level.However, DI treats the task as binary, whereas we argue that Arabic speakers perceive a spectrum of dialectness, which we operationalize at the sentence level as the Arabic Level of Dialectness (ALDi), a continuous linguistic variable.We introduce the AOC-ALDi dataset (derived from the AOC dataset), containing 127,835 sentences (17% from news articles and 83% from user comments on those articles) which are manually labeled with their level of dialectness.We provide a detailed analysis of AOC-ALDi and show that a model trained on it can effectively identify levels of dialectness on a range of other corpora (including dialects and genres not included in AOC-ALDi), providing a more nuanced picture than traditional DI systems.Through case studies, we illustrate how ALDi can reveal Arabic speakers' stylistic choices in different situations, a useful property for sociolinguistic analyses. Amr Keleg, Sharon Goldwater, Walid Magdy |
EMNLP | 2 |
| 2023 | Analyzing Acoustic Word Embeddings from Pre-Trained Self-Supervised Speech ModelsabstractGiven the strong results of self-supervised models on various tasks, there have been surprisingly few studies exploring self-supervised representations for acoustic word embeddings (AWE), fixed-dimensional vectors representing variable-length spoken word segments. In this work, we study several pre-trained models and pooling methods for constructing AWEs with self-supervised representations. Owing to the contextualized nature of self-supervised representations, we hy-pothesize that simple pooling methods, such as averaging, might already be useful for constructing AWEs. When evaluating on a standard word discrimination task, we find that HuBERT representations with mean-pooling rival the state of the art on English AWEs. More surprisingly, despite being trained only on English, HuBERT representations evaluated on Xitsonga, Mandarin, and French consistently outperform the multilingual model XLSR-53 (as well as Wav2Vec 2.0 trained on English). Ramon Sanabria, Hao Tang 0002, Sharon Goldwater |
ICASSP | 3 |
| 2023 | Self-supervised Predictive Coding Models Encode Speaker and Phonetic Information in Orthogonal SubspacesabstractSelf-supervised speech representations are known to encode both speaker and phonetic information, but how they are distributed in the high-dimensional space remains largely unexplored. We hypothesize that they are encoded in orthogonal subspaces, a property that lends itself to simple disentanglement. Applying principal component analysis to representations of two predictive coding models, we identify two subspaces that capture speaker and phonetic variances, and confirm that they are nearly orthogonal. Based on this property, we propose a new speaker normalization method which collapses the subspace that encodes speaker information, without requiring transcriptions. Probing experiments show that our method effectively eliminates speaker information and outperforms a previous baseline in phone discrimination tasks. Moreover, the approach generalizes and can be used to remove information of unseen speakers. Oli Danyi Liu, Hao Tang 0002, Sharon Goldwater |
INTERSPEECH | 3 |
| 2023 | Parsing dialog turns with prosodic features in EnglishabstractParsing spoken dialogue presents challenges that parsing text does not, including a lack of clear sentence boundaries.We know from previous work that prosody helps in parsing single sentences [1], but we want to show the effect of prosody on parsing speech that isn't segmented into sentences.In experiments on the English Switchboard corpus, we find prosody helps our model both with parsing and with accurately identifying sentence boundaries.However, we find that the bestperforming parser is not necessarily the parser that produces the best sentence segmentation performance.We suggest that the best parses instead come from modelling sentence boundaries jointly with other syntactic boundaries. Elizabeth Nielsen, Mark Steedman, Sharon Goldwater |
INTERSPEECH | 3 |
| 2023 | Acoustic Word Embeddings for Untranscribed Target Languages with Continued Pretraining and Learned PoolingabstractAcoustic word embeddings are typically created by training a pooling function using pairs of word-like units. For unsupervised systems, these are mined using k-nearest neighbor (KNN) search, which is slow. Recently, mean-pooled representations from a pre-trained self-supervised English model were suggested as a promising alternative, but their performance on target languages was not fully competitive. Here, we explore improvements to both approaches: we use continued pre-training to adapt the self-supervised model to the target language, and we use a multilingual phone recognizer (MPR) to mine phone n-gram pairs for training the pooling function. Evaluating on four languages, we show that both methods outperform a recent approach on word discrimination. Moreover, the MPR method is orders of magnitude faster than KNN, and is highly data efficient. We also show a small improvement from performing learned pooling on top of the continued pre-trained representations. Ramon Sanabria, Ondrej Klejch, Hao Tang 0002, Sharon Goldwater |
INTERSPEECH | 4 |
| 2022 | Regularization or lexical probability-matching? How German speakers generalize plural morphology
Kate McCurdy, Sharon Goldwater, Adam Lopez |
CogSci | 2 |
| 2021 | Prosodic segmentation for parsing spoken dialogueabstractElizabeth Nielsen, Mark Steedman, Sharon Goldwater. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Elizabeth Nielsen, Mark Steedman, Sharon Goldwater |
ACL/IJCNLP (1) | 3 |
| 2021 | A phonetic model of non-native spoken word processingabstractYevgen Matusevych, Herman Kamper, Thomas Schatz, Naomi Feldman, Sharon Goldwater. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Yevgen Matusevych, Herman Kamper, Thomas Schatz, Naomi Feldman, Sharon Goldwater |
EACL | 5 |
| 2021 | Multilingual and unsupervised subword modeling for zero-resource languages
Enno Hermann, Herman Kamper, Sharon Goldwater |
Comput. Speech Lang. | 3 |
| 2021 | Black or White but Never Neutral: How Readers Perceive Identity from Yellow or Skin-toned EmojiabstractResearch in sociology and linguistics shows that people use language not only to express their own identity but to understand the identity of others. Recent work established a connection between expression of identity and emoji usage on social media, through use of emoji skin tone modifiers. Motivated by that finding, this work asks if, as with language, readers are sensitive to such acts of self-expression and use them to understand the identity of authors. In behavioral experiments (n=488), where text and emoji content of social media posts were carefully controlled before being presented to participants, we find in the affirmative - emoji are a salient signal of author identity. That signal is distinct from, and complementary to, the one encoded in language. Participant groups (based on self-identified ethnicity) showed no differences in how they perceive this signal, except in the case of the default yellow emoji. While both groups associate this with a White identity, the effect was stronger in White participants. Our finding that emoji can index social variables will have experimental applications for researchers but also implications for designers: supposedly neutral defaults may be more representative of some users than others. Alexander Robertson, Walid Magdy, Sharon Goldwater |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2021 | Improved Acoustic Word Embeddings for Zero-Resource Languages Using Multilingual TransferabstractAcoustic word embeddings are fixed-dimensional representations of variable-length speech segments. Such embeddings can form the basis for speech search, indexing and discovery systems when conventional speech recognition is not possible. In zero-resource settings where unlabelled speech is the only available resource, we need a method that gives robust embeddings on an arbitrary language. Here we explore multilingual transfer: we train a single supervised embedding model on labelled data from multiple well-resourced languages and then apply it to unseen zero-resource languages. We consider three multilingual recurrent neural network (RNN) models: a classifier trained on the joint vocabularies of all training languages; a Siamese RNN trained to discriminate between same and different words from multiple languages; and a correspondence autoencoder (CAE) RNN trained to reconstruct word pairs. In a word discrimination task on six target languages, all of these models outperform state-of-the-art unsupervised models trained on the zero-resource languages themselves, giving relative improvements of more than 30% in average precision. When using only a few training languages, the multilingual CAE-RNN performs better, but with more training languages the other multilingual models perform similarly. Using more training languages is generally beneficial, but improvements are marginal on some languages. We present probing experiments which show that the CAE-RNN encodes more phonetic, word duration, language identity and speaker information than the other multilingual models. Herman Kamper, Yevgen Matusevych, Sharon Goldwater |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Inflecting When There's No Majority: Limitations of Encoder-Decoder Neural Networks as Cognitive Models for German PluralsabstractCan artificial neural networks learn to represent inflectional morphology and generalize to new words as human speakers do?Kirov and Cotterell (2018) argue that the answer is yes: modern Encoder-Decoder (ED) architectures learn human-like behavior when inflecting English verbs, such as extending the regular past tense form /-(e)d/ to novel words.However, their work does not address the criticism raised by Marcus et al. (1995): that neural models may learn to extend not the regular, but the most frequent class -and thus fail on tasks like German number inflection, where infrequent suffixes like /-s/ can still be productively generalized.To investigate this question, we first collect a new dataset from German speakers (production and ratings of plural forms for novel nouns) that is designed to avoid sources of information unavailable to the ED model.The speaker data show high variability, and two suffixes evince 'regular' behavior, appearing more often with phonologically atypical inputs.Encoder-decoder models do generalize the most frequently produced plural class, but do not show human-like variability or 'regular' extension of these other plural markers.We conclude that modern neural models may still struggle with minority-class generalization. Kate McCurdy, Sharon Goldwater, Adam Lopez |
ACL | 2 |
| 2020 | Input matters in the modeling of early phonetic learning
Ruolan Li, Thomas Schatz, Yevgen Matusevych, Sharon Goldwater, Naomi Feldman |
CogSci | 4 |
| 2020 | Evaluating computational models of infant phonetic learning across languages
Yevgen Matusevych, Thomas Schatz, Herman Kamper, Naomi Feldman, Sharon Goldwater |
CogSci | 5 |
| 2020 | The role of context in neural pitch accent detection in EnglishabstractProsody is a rich information source in natural language, serving as a marker for phenomena such as contrast.In order to make this information available to downstream tasks, we need a way to detect prosodic events in speech.We propose a new model for pitch accent detection, inspired by the work of Stehwien et al. (2018), who presented a CNN-based model for this task.Our model makes greater use of context by using full utterances as input and adding an LSTM layer.We find that these innovations lead to an improvement from 87.5 percent to 88.7 percent accuracy on pitch accent detection on American English speech in the Boston University Radio News Corpus, a state-of-the-art result.We also find that a simple baseline that just predicts a pitch accent on every content word yields 82.2 percent accuracy, and we suggest that this is the appropriate baseline for this task.Finally, we conduct ablation tests that show pitch is the most important acoustic feature for this task and this corpus. Elizabeth Nielsen, Mark Steedman, Sharon Goldwater |
EMNLP (1) | 3 |
| 2020 | Cross-Lingual Topic Prediction For Speech Using TranslationsabstractGiven a large amount of unannotated speech in a low-resource language, can we classify the speech utterances by topicƒ We consider this question in the setting where a small amount of speech in the low-resource language is paired with text translations in a high-resource language. We develop an effective cross-lingual topic classifier by training on just 20 hours of translated speech, using a recent model for direct speech-to-text translation. While the translations are poor, they are still good enough to correctly classify the topic of 1-minute speech segments over 70% of the time—a 20% improvement over a majority-class baseline. Such a system could be useful for humanitarian applications like crisis response, where incoming speech in a foreign low-resource language must be quickly assessed for further action. Sameer Bansal, Herman Kamper, Adam Lopez, Sharon Goldwater |
ICASSP | 4 |
| 2020 | Multilingual Acoustic Word Embedding Models for Processing Zero-resource LanguagesabstractAcoustic word embeddings are fixed-dimensional representations of variable-length speech segments. In settings where unlabelled speech is the only available resource, such embeddings can be used in "zero-resource" speech search, indexing and discovery systems. Here we propose to train a single supervised embedding model on labelled data from multiple well-resourced languages and then apply it to unseen zero-resource languages. For this transfer learning approach, we consider two multilingual recurrent neural network models: a discriminative classifier trained on the joint vocabularies of all training languages, and a correspondence autoencoder trained to reconstruct word pairs. We test these using a word discrimination task on six target zero-resource languages. When trained on seven well-resourced languages, both models perform similarly and outperform unsupervised models trained on the zero-resource languages. With just a single training language, the second model works better, but performance depends more on the particular training-testing language pair. Herman Kamper, Yevgen Matusevych, Sharon Goldwater |
ICASSP | 3 |
| 2020 | Analyzing ASR Pretraining for Low-Resource Speech-to-Text TranslationabstractPrevious work has shown that for low-resource source languages, automatic speech-to-text translation (AST) can be improved by pre-training an end-to-end model on automatic speech recognition (ASR) data from a high-resource language. However, it is not clear what factors - e.g., language relatedness or size of the pretraining data - yield the biggest improvements, or whether pretraining can be effectively combined with other methods such as data augmentation. Here, we experiment with pretraining on datasets of varying sizes, including languages related and unrelated to the AST source language. We find that the best predictor of final AST performance is the word error rate of the pretrained ASR model, and that differences in ASR/AST performance correlate with how phonetic information is encoded in the later RNN layers of our model. We also show that pretraining and data augmentation yield complementary benefits for AST. Mihaela Catalina Stoian, Sameer Bansal, Sharon Goldwater |
ICASSP | 3 |
| 2019 | Are we there yet? Encoder-decoder neural networks as cognitive models of English past tense inflectionabstractThe cognitive mechanisms needed to account for the English past tense have long been a subject of debate in linguistics and cognitive science.Neural network models were proposed early on, but were shown to have clear flaws.Recently, however, Kirov and Cotterell (2018) showed that modern encoder-decoder (ED) models overcome many of these flaws.They also presented evidence that ED models demonstrate humanlike performance in a nonce-word task.Here, we look more closely at the behaviour of their model in this task.We find that (1) the model exhibits instability across multiple simulations in terms of its correlation with human data, and (2) even when results are aggregated across simulations (treating each simulation as an individual human participant), the fit to the human data is not strong-worse than an older rule-based model.These findings hold up through several alternative training regimes and evaluation measures.Although other neural architectures might do better, we conclude that there is still insufficient evidence to claim that neural nets are a good cognitive model for this task. Maria Corkery, Yevgen Matusevych, Sharon Goldwater |
ACL (1) | 3 |
| 2018 | Self-Representation on Twitter Using Emoji Skin Color Modifiers
Alexander Robertson, Walid Magdy, Sharon Goldwater |
ICWSM | 3 |
| 2018 | Low-Resource Speech-to-Text TranslationabstractSpeech-to-text translation has many potential applications for low-resource languages, but the typical approach of cascading speech recognition with machine translation is often impossible, since the transcripts needed to train a speech recognizer are usually not available for low-resource languages. Recent work has found that neural encoder-decoder models can learn to directly translate foreign speech in high-resource scenarios, without the need for intermediate transcription. We investigate whether this approach also works in settings where both data and computation are limited. To make the approach efficient, we make several architectural changes, including a change from character-level to word-level decoding. We find that this choice yields crucial speed improvements that allow us to train with fewer computational resources, yet still performs well on frequent words. We explore models trained on between 20 and 160 hours of data, and find that although models trained on less data have considerably lower BLEU scores, they can still predict words with relatively high precision and recall---around 50% for a model trained on 50 hours of data, versus around 60% for the full 160 hour model. Thus, they may still be useful for some low-resource scenarios. Sameer Bansal, Herman Kamper, Karen Livescu, Adam Lopez, Sharon Goldwater |
INTERSPEECH | 5 |
| 2018 | Multilingual Bottleneck Features for Subword Modeling in Zero-resource LanguagesabstractHow can we effectively develop speech technology for languages where no transcribed data is available? Many existing approaches use no annotated resources at all, yet it makes sense to leverage information from large annotated corpora in other languages, for example in the form of multilingual bottleneck features (BNFs) obtained from a supervised speech recognition system. In this work, we evaluate the benefits of BNFs for subword modeling (feature extraction) in six unseen languages on a word discrimination task. First we establish a strong unsupervised baseline by combining two existing methods: vocal tract length normalisation (VTLN) and the correspondence autoencoder (cAE). We then show that BNFs trained on a single language already beat this baseline; including up to 10 languages results in additional improvements which cannot be matched by just adding more data from a single language. Finally, we show that the cAE can improve further on the BNFs if high-quality same-word pairs are available. Enno Hermann, Sharon Goldwater |
INTERSPEECH | 2 |
| 2018 | Context Sensitive Neural Lemmatization with LematusabstractToms Bergmanis, Sharon Goldwater. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Toms Bergmanis, Sharon Goldwater |
NAACL-HLT | 2 |
| 2017 | An embedded segmental K-means model for unsupervised segmentation and clustering of speechabstractUnsupervised segmentation and clustering of unlabelled speech are core problems in zero-resource speech processing. Most approaches lie at methodological extremes: some use probabilistic Bayesian models with convergence guarantees, while others opt for more efficient heuristic techniques. Despite competitive performance in previous work, the full Bayesian approach is difficult to scale to large speech corpora. We introduce an approximation to a recent Bayesian model that still has a clear objective function but improves efficiency by using hard clustering and segmentation rather than full Bayesian inference. Like its Bayesian counterpart, this embedded segmental K-means model (ES-KMeans) represents arbitrary-length word segments as fixed-dimensional acoustic word embeddings. We first compare ES-KMeans to previous approaches on common English and Xitsonga data sets (5 and 2.5 hours of speech): ES-KMeans outperforms a leading heuristic method in word segmentation, giving similar scores to the Bayesian model while being 5 times faster with fewer hyperparameters. However, its clusters are less pure than those of the other models. We then show that ES-KMeans scales to larger corpora by applying it to the 5 languages of the Zero Resource Speech Challenge 2017 (up to 45 hours), where it performs competitively compared to the challenge baseline. Herman Kamper, Karen Livescu, Sharon Goldwater |
ASRU | 3 |
| 2017 | From Segmentation to Analyses: a Probabilistic Model for Unsupervised Morphology InductionabstractA major motivation for unsupervised morphological analysis is to reduce the sparse data problem in under-resourced languages.Most previous work focuses on segmenting surface forms into their constituent morphs (e.g., taking: tak +ing), but surface form segmentation does not solve the sparse data problem as the analyses of take and taking are not connected to each other.We extend the MorphoChains system (Narasimhan et al., 2015) to provide morphological analyses that can abstract over spelling differences in functionally similar morphs.These analyses are not required to use all the orthographic material of a word (stopping: stop +ing), nor are they limited to only that material (acidified: acid +ify +ed).On average across six typologically varied languages our system has a similar or better F-score on EMMA (a measure of underlying morpheme accuracy) than three strong baselines; moreover, the total number of distinct morphemes identified by our system is on average 12.8% lower than for Morfessor (Virpioja et al., 2013), a stateof-the-art surface segmentation system. Toms Bergmanis, Sharon Goldwater |
EACL (1) | 2 |
| 2017 | Aye or naw, whit dae ye hink? Scottish independence and linguistic identity on social mediaabstractPhilippa Shoemark, Debnil Sur, Luke Shrimpton, Iain Murray, Sharon Goldwater. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Philippa Shoemark, Debnil Sur, Luke Shrimpton, Iain Murray 0001, Sharon Goldwater |
EACL (1) | 5 |
| 2017 | Weakly supervised spoken term discovery using cross-lingual side informationabstractRecent work on unsupervised term discovery (UTD) aims to identify and cluster repeated word-like units from audio alone. These systems are promising for some very low-resource languages where transcribed audio is unavailable, or where no written form of the language exists. However, in some cases it may still be feasible (e.g., through crowdsourcing) to obtain (possibly noisy) text translations of the audio. If so, this information could be used as a source of side information to improve UTD. Here, we present a simple method for rescoring the output of a UTD system using text translations, and test it on a corpus of Spanish audio with English translations. We show that it greatly improves the average precision of the results over a wide range of system configurations and data preprocessing methods. Sameer Bansal, Herman Kamper, Sharon Goldwater, Adam Lopez |
ICASSP | 3 |
| 2017 | A segmental framework for fully-unsupervised large-vocabulary speech recognition
Herman Kamper, Aren Jansen, Sharon Goldwater |
Comput. Speech Lang. | 3 |
| 2016 | Unsupervised Word Segmentation and Lexicon Discovery Using Acoustic Word EmbeddingsabstractIn settings where only unlabeled speech data is available, speech technology needs to be developed without transcriptions, pronunciation dictionaries, or language modelling text. A similar problem is faced when modeling infant language acquisition. In these cases, categorical linguistic structure needs to be discovered directly from speech audio. We present a novel unsupervised Bayesian model that segments unlabeled speech and clusters the segments into hypothesized word groupings. The result is a complete unsupervised tokenization of the input speech in terms of discovered word types. In our approach, a potential word segment (of arbitrary length) is embedded in a fixed-dimensional acoustic vector space. The model, implemented as a Gibbs sampler, then builds a whole-word acoustic model in this space while jointly performing segmentation. We report word error rates in a small-vocabulary connected digit recognition task by mapping the unsupervised decoded output to ground truth transcriptions. The model achieves around 20% error rate, outperforming a previous HMM-based system by about 10% absolute. Moreover, in contrast to the baseline, our model does not require a pre-specified vocabulary size. Herman Kamper, Aren Jansen, Sharon Goldwater |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Unsupervised neural network based feature extraction using weak top-down constraintsabstractDeep neural networks (DNNs) have become a standard component in supervised ASR, used in both data-driven feature extraction and acoustic modelling. Supervision is typically obtained from a forced alignment that provides phone class targets, requiring transcriptions and pronunciations. We propose a novel unsupervised DNN-based feature extractor that can be trained without these resources in zero-resource settings. Using unsupervised term discovery, we find pairs of isolated word examples of the same unknown type; these provide weak top-down supervision. For each pair, dynamic programming is used to align the feature frames of the two words. Matching frames are presented as input-output pairs to a deep autoencoder (AE) neural network. Using this AE as feature extractor in a word discrimination task, we achieve 64% relative improvement over a previous state-of-the-art system, 57% improvement relative to a bottom-up trained deep AE, and come to within 23% of a supervised system. Herman Kamper, Micha Elsner, Aren Jansen, Sharon Goldwater |
ICASSP | 4 |
| 2015 | Fully unsupervised small-vocabulary speech recognition using a segmental Bayesian modelabstractCurrent supervised speech technology relies heavily on tran-scribed speech and pronunciation dictionaries. In settings where unlabelled speech data alone is available, unsupervised methods are required to discover categorical linguistic structure directly from the audio. We present a novel Bayesian model which seg-ments unlabelled input speech into word-like units, resulting in a complete unsupervised transcription of the speech in terms of discovered word types. In our approach, a potential word segment (of arbitrary length) is embedded in a fixed-dimensional space; the model (implemented as a Gibbs sampler) then builds a whole-word acoustic model in this space while jointly doing seg-mentation. We report word error rates in a connected digit recog-nition task by mapping the unsupervised output to ground truth transcriptions. Our model outperforms a previously developed HMM-based system, even when the model is not constrained to discover only the 11 word types present in the data. Index Terms: unsupervised speech processing, word discovery, speech segmentation, unsupervised learning, segmental models 1. Herman Kamper, Aren Jansen, Sharon Goldwater |
INTERSPEECH | 3 |
| 2015 | A comparison of neural network methods for unsupervised representation learning on the zero resource speech challengeabstractThe success of supervised deep neural networks (DNNs) in speech recognition cannot be transferred to zero-resource languages where the requisite transcriptions are unavailable. We investigate unsupervised neural network based methods for learning frame-level representations. Good frame representations eliminate differences in accent, gender, channel characteristics, and other factors to model subword units for within- and across-speaker phonetic discrimination. We enhance the correspondence autoencoder (cAE) and show that it can transform Mel Frequency Cepstral Coefficients (MFCCs) into more effective frame representations given a set of matched word pairs from an unsupervised term discovery (UTD) system. The cAE combines the feature extraction power of autoencoders with the weak supervision signal from UTD pairs to better approximate the extrinsic task’s objective during training. We use the Zero Resource Speech Challenge’s minimal triphone pair ABX discrimination task to evaluate our methods. Optimizing a cAE architecture on English and applying it to a zero-resource language, Xitsonga, we obtain a relative error rate reduction of 35% compared to the original MFCCs. We also show that Xitsonga frame representations extracted from the bottleneck layer of a supervised DNN trained on English can be further enhanced by the cAE, yielding a relative error rate reduction of 39%. Daniel Renshaw, Herman Kamper, Aren Jansen, Sharon Goldwater |
INTERSPEECH | 4 |
| 2014 | Weak semantic context helps phonetic learning in a model of infant language acquisitionabstractLearning phonetic categories is one of the first steps to learning a language, yet is hard to do using only distributional phonetic information.Semantics could potentially be useful, since words with different meanings have distinct phonetics, but it is unclear how many word meanings are known to infants learning phonetic categories.We show that attending to a weaker source of semantics, in the form of a distribution over topics in the current context, can lead to improvements in phonetic category learning.In our model, an extension of a previous model of joint word-form and phonetic category inference, the probability of word-forms is topic-dependent, enabling the model to find significantly better phonetic vowel categories and word-forms than a model with no semantic knowledge. Stella Frank, Naomi Feldman, Sharon Goldwater |
ACL (1) | 3 |
| 2014 | Unsupervised lexical clustering of speech segments using fixed-dimensional acoustic embeddingsabstractUnsupervised speech processing methods are essential for applications ranging from zero-resource speech technology to modelling child language acquisition. One challenging problem is discovering the word inventory of the language: the lexicon. Lexical clustering is the task of grouping unlabelled acoustic word tokens according to type. We propose a novel lexical clustering model: variable-length word segments are embedded in a fixed-dimensional acoustic space in which clustering is then performed. We evaluate several clustering algorithms and find that the best methods produce clusters with wide variation in sizes, as observed in natural language. The best probabilistic approach is an infinite Gaussian mixture model (IGMM), which automatically chooses the number of clusters. Performance is comparable to that of non-probabilistic Chinese Whispers and average-linkage hierarchical clustering. We conclude that IGMM clustering of fixed-dimensional embeddings holds promise as the lexical clustering component in unsupervised speech processing systems. Herman Kamper, Aren Jansen, Simon King 0001, Sharon Goldwater |
SLT | 4 |
| 2013 | A Joint Learning Model of Word Segmentation, Lexical Acquisition, and Phonetic VariabilityabstractWe present a cognitive model of early lexical acquisition which jointly performs word segmentation and learns an explicit model of phonetic variation.We define the model as a Bayesian noisy channel; we sample segmentations and word forms simultaneously from the posterior, using beam sampling to control the size of the search space.Compared to a pipelined approach in which segmentation is performed first, our model is qualitatively more similar to human learners.On data with variable pronunciations, the pipelined approach learns to treat syllables or morphemes as words.In contrast, our joint model, like infant learners, tends to learn multiword collocations.We also conduct analyses of the phonetic variations that the model learns to accept and its patterns of word recognition errors, and relate these to developmental evidence. Micha Elsner, Sharon Goldwater, Naomi Feldman, Frank D. Wood |
EMNLP | 2 |
| 2013 | Exploring the Utility of Joint Morphological and Syntactic Learning from Child-directed SpeechabstractChildren learn various levels of linguistic structure concurrently, yet most existing models of language acquisition deal with only a single level of structure, implicitly assuming a sequential learning process.Developing models that learn multiple levels simultaneously can provide important insights into how these levels might interact synergistically during learning.Here, we present a model that jointly induces syntactic categories and morphological segmentations by combining two well-known models for the individual tasks.We test on child-directed utterances in English and Spanish and compare to single-task baselines.In the morphologically poorer language (English), the model improves morphological segmentation, while in the morphologically richer language (Spanish), it leads to better syntactic categorization.These results provide further evidence that joint learning is useful, but also suggest that the benefits may be different for typologically different languages. Stella Frank, Frank Keller, Sharon Goldwater |
EMNLP | 3 |
| 2013 | A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisitionabstractWe summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding zero resource (unsupervised) speech technologies and related models of early language acquisition. Centered around the tasks of phonetic and lexical discovery, we consider unified evaluation metrics, present two new approaches for improving speaker independence in the absence of supervision, and evaluate the application of Bayesian word segmentation algorithms to automatic subword unit tokenizations. Finally, we present two strategies for integrating zero resource techniques into supervised settings, demonstrating the potential of unsupervised methods to improve mainstream technologies. Aren Jansen, Emmanuel Dupoux, Sharon Goldwater, Mark Johnson 0001, Sanjeev Khudanpur, Kenneth Church 0001, Naomi Feldman, Hynek Hermansky, Florian Metze, Richard C. Rose, Mike Seltzer, Pascal Clark, Ian McGraw, Balakrishnan Varadarajan, Erin D. Bennett, Benjamin Börschinger, Justin T. Chiu, Ewan Dunbar, Abdellah Fourtassi, David F. Harwath, Chia-ying Lee, Keith D. Levin, Atta Norouzian, Vijayaditya Peddinti, Rachael Richardson, Thomas Schatz, Samuel Thomas 0001 |
ICASSP | 3 |
| 2013 | Unsupervised Dependency Parsing with Acoustic CuesabstractUnsupervised parsing is a difficult task that infants readily perform. Progress has been made on this task using text-based models, but few computational approaches have considered how infants might benefit from acoustic cues. This paper explores the hypothesis that word duration can help with learning syntax. We describe how duration information can be incorporated into an unsupervised Bayesian dependency parser whose only other source of information is the words themselves (without punctuation or parts of speech). Our results, evaluated on both adult-directed and child-directed utterances, show that using word duration can improve parse quality relative to words-only baselines. These results support the idea that acoustic cues provide useful evidence about syntactic structure for language-learning infants, and motivate the use of word duration cues in NLP tasks with speech. John K. Pate, Sharon Goldwater |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | Minimally-Supervised Morphological Segmentation using Adaptor GrammarsabstractThis paper explores the use of Adaptor Grammars, a nonparametric Bayesian modelling framework, for minimally supervised morphological segmentation. We compare three training methods: unsupervised training, semi-supervised training, and a novel model selection method. In the model selection method, we train unsupervised Adaptor Grammars using an over-articulated metagrammar, then use a small labelled data set to select which potential morph boundaries identified by the metagrammar should be returned in the final output. We evaluate on five languages and show that semi-supervised training provides a boost over unsupervised training, while the model selection method yields the best average results over all languages and is competitive with state-of-the-art semi-supervised systems. Moreover, this method provides the potential to tune performance according to different evaluation metrics or downstream tasks. Kairit Sirts, Sharon Goldwater |
Trans. Assoc. Comput. Linguistics | 2 |
| 2012 | Bootstrapping a Unified Model of Lexical and Phonetic Acquisition
Micha Elsner, Sharon Goldwater, Jacob Eisenstein |
ACL (1) | 2 |
| 2012 | Semantic Parsing with Bayesian Tree Transducers
Bevan K. Jones, Mark Johnson 0001, Sharon Goldwater |
ACL (1) | 3 |
| 2012 | A Probabilistic Model of Syntactic and Semantic Acquisition from Child-Directed Utterances and their Meanings
Tom Kwiatkowski, Sharon Goldwater, Luke Zettlemoyer, Mark Steedman |
EACL | 2 |
| 2011 | Unsupervised Extraction of Recurring Words from Infant-Directed Speech
Fergus R. McInnes, Sharon Goldwater |
CogSci | 2 |
| 2011 | Predictability effects in adult-directed and infant-directed speech: Does the listener matter?
John K. Pate, Sharon Goldwater |
CogSci | 2 |
| 2011 | A Bayesian Mixture Model for PoS Induction Using Multiple Features
Christos Christodoulopoulos 0001, Sharon Goldwater, Mark Steedman |
EMNLP | 2 |
| 2011 | Lexical Generalization in CCG Grammar Induction for Semantic Parsing
Tom Kwiatkowski, Luke Zettlemoyer, Sharon Goldwater, Mark Steedman |
EMNLP | 3 |
| 2011 | Computational Modeling of Human Language Acquisition Afra Alishahi (University of the Saarland) Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 11), 2010, xiv+93 pp; paperbound, ISBN 978-1-60845-339-9, $40.00; ebook, ISBN 978-1-60845-340-5, $30.00 or by subscription
Sharon Goldwater |
Comput. Linguistics | 1 |
| 2011 | Producing Power-Law Distributions and Damping Word Frequencies with Two-Stage Language Models
Sharon Goldwater, Thomas L. Griffiths 0001, Mark Johnson 0001 |
J. Mach. Learn. Res. | 1 |
| 2010 | Two Decades of Unsupervised POS Induction: How Far Have We Come?
Christos Christodoulopoulos 0001, Sharon Goldwater, Mark Steedman |
EMNLP | 2 |
| 2010 | Inducing Probabilistic CCG Grammars from Logical Form with Higher-Order Unification
Tom Kwiatkowski, Luke Zettlemoyer, Sharon Goldwater, Mark Steedman |
EMNLP | 3 |
| 2010 | Inducing Tree-Substitution Grammars
Trevor Cohn, Phil Blunsom, Sharon Goldwater |
J. Mach. Learn. Res. | 3 |
| 2010 | Which words are hard to recognize? Prosodic, lexical, and disfluency factors that increase speech recognition error rates
Sharon Goldwater, Daniel Jurafsky, Christopher D. Manning |
Speech Commun. | 1 |
| 2009 | Improving Morphology Induction by Learning Spelling Rules
Jason Naradowsky, Sharon Goldwater |
IJCAI | 2 |
| 2009 | Inducing Compact but Accurate Tree-Substitution Grammars
Trevor Cohn, Sharon Goldwater, Phil Blunsom |
HLT-NAACL | 2 |
| 2009 | Improving nonparameteric Bayesian inference: experiments on unsupervised word segmentation with adaptor grammars
Mark Johnson 0001, Sharon Goldwater |
HLT-NAACL | 2 |
| 2008 | Which Words Are Hard to Recognize? Prosodic, Lexical, and Disfluency Factors that Increase ASR Error Rates
Sharon Goldwater, Daniel Jurafsky, Christopher D. Manning |
ACL | 1 |
| 2007 | A fully Bayesian approach to unsupervised part-of-speech tagging
Sharon Goldwater, Thomas L. Griffiths 0001 |
ACL | 1 |
| 2007 | Bayesian Inference for PCFGs via Markov Chain Monte Carlo
Mark Johnson 0001, Thomas L. Griffiths 0001, Sharon Goldwater |
HLT-NAACL | 3 |
| 2006 | Contextual Dependencies in Unsupervised Word SegmentationabstractDeveloping better methods for segmenting continuous text into words is important for improving the processing of Asian languages, and may shed light on how humans learn to segment speech. We propose two new Bayesian word segmentation methods that assume unigram and bigram models of word dependencies respectively. The bigram model greatly outperforms the unigram model (and previous probabilistic models), demonstrating the importance of such dependencies for word segmentation. We also show that previous probabilistic models rely crucially on sub-optimal search procedures. Sharon Goldwater, Thomas L. Griffiths 0001, Mark Johnson 0001 |
ACL | 1 |
| 2006 | Adaptor Grammars: A Framework for Specifying Compositional Nonparametric Bayesian ModelsabstractThis paper introduces adaptor grammars, a class of probabilistic models of lan- guage that generalize probabilistic context-free grammars (PCFGs). Adaptor grammars augment the probabilistic rules of PCFGs with “adaptors” that can in- duce dependencies among successive uses. With a particular choice of adaptor, based on the Pitman-Yor process, nonparametric Bayesian models of language using Dirichlet processes and hierarchical Dirichlet processes can be written as simple grammars. We present a general-purpose inference algorithm for adaptor grammars, making it easy to define and use such models, and illustrate how several existing nonparametric Bayesian models can be expressed within this framework. Mark Johnson 0001, Thomas L. Griffiths 0001, Sharon Goldwater |
NIPS | 3 |
| 2005 | Representational Bias in Unsupervised Learning of Syllable Structure
Sharon Goldwater, Mark Johnson 0001 |
CoNLL | 1 |
| 2005 | Interpolating between types and tokens by estimating power-law generatorsabstractStandard statistical models of language fail to capture one of the most striking properties of natural languages: the power-law distribution in the frequencies of word tokens. We present a framework for developing statistical models that generically produce power-laws, augmenting stan- dard generative models with an adaptor that produces the appropriate pattern of token frequencies. We show that taking a particular stochastic process – the Pitman-Yor process – as an adaptor justifies the appearance of type frequencies in formal analyses of natural language, and improves the performance of a model for unsupervised learning of morphology. Sharon Goldwater, Thomas L. Griffiths 0001, Mark Johnson 0001 |
NIPS | 1 |
| 2003 | A Type System for Statically Detecting Spreadsheet ErrorsabstractWe describe a methodology for detecting user errors in spreadsheets, using the notion of units as our basic elements of checking. We define the concept of a header and discuss two types of relationships between headers, namely is-a and has-a relationships. With these, we develop a set of rules to assign units to cells in the spreadsheet. We check for errors by ensuring that every cell has a well-formed unit. We describe an implementation of the system that allows the user to check Microsoft Excel spreadsheets. We have run our system on practical examples, and even found errors in published spreadsheets. Yanif Ahmad, Tudor Antoniu, Sharon Goldwater, Shriram Krishnamurthi |
ASE | 3 |
| 2000 | Compiling Language Models from a Linguistically Motivated Unification Grammar
Manny Rayner, Beth Ann Hockey, Frankie James, Elizabeth Owen Bratt, Sharon Goldwater, Jean Mark Gawron |
COLING | 5 |