EDBT 2026 Demo / reviewers in the wild / expert
Ben Bergen 0001
dblp:12/3783-1 · also Benjamin Bergen 0001, Benjamin K. Bergen 0001
· DBLP profile ↗
49ranked-venue papers
1as first author
31since 2021 · last 2026
0000-0002-9395-9151ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 1 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discovering Lexical Gaps Using Embeddings from Multilingual LLMsabstractLexical gaps are words that do not exist in certain languages.They pose challenges for building multilingual lexical resources, for machine translation, and for cross-lingual transfer.Existing lexical gap detection relies on human judgments or fixed conceptual taxonomies.We propose a data-driven framework for identifying cross-lingual lexical gaps.We extracted contextualized embeddings from Korean-English bilingual LLMs for Korean-to-English and English-to-Korean translation pairs.Combinations of LLMs, embedding types, dimensionality, and orthogonal transformations across 100 train-test splits yielded 4000 distinct embedding spaces in each source language.In each space, we computed the semantic similarity between each source word and its nearest neighbor in the target language, and compared their distribution for gap words versus non-gap words.In 94% (Korean-to-English) and 97% (English-to-Korean) of embedding spaces, gap words showed weaker cross-lingual semantic alignment than non-gap words.Logistic classifiers trained on unaligned embedding spaces can reliably separate gap words from non-gap words, achieving AUCs of 0.81 (Korean-to-English) and 0.76 (English-to-Korean) and retrieving 18/19 Korean and 26/27 English gap words.This approach provides a languageagnostic and taxonomy-free method for scalable lexical gap identification. Yoonwon Jung, Aaron S. Cohen, Ben Bergen 0001 |
CoNLL | 3 |
| 2026 | Goldfish: Monolingual Language Models for 350 LanguagesabstractFor many low-resource languages, the only available language models are large multilingual models trained on many languages simultaneously. Despite state-of-the-art performance on reasoning tasks, we find that these models still struggle with basic grammatical text generation in many languages. First, large multilingual models perform worse than bigrams for many languages (e.g. 24% of languages in XGLM 4.5B; 43% in BLOOM 7.1B) using FLORES perplexity as an evaluation metric. Second, when we train small monolingual models with only 125M parameters on 1GB or less data for 350 languages, these small models outperform large multilingual models both in perplexity and on a massively multilingual grammaticality benchmark. To facilitate future work on low-resource language modeling, we release Goldfish, a suite of over 1,000 small monolingual language models trained comparably for 350 languages. These models represent the first publicly-available monolingual language models for 215 of the languages included. Tyler A. Chang, Catherine Arnett, Zhuowen Tu, Ben Bergen 0001 |
LREC | 4 |
| 2025 | On the Acquisition of Shared Grammatical Representations in Bilingual Language ModelsabstractWhile crosslingual transfer is crucial to contemporary language models’ multilingual capabilities, how it occurs is not well understood. In this paper, we ask what happens to a monolingual language model when it begins to be trained on a second language. Specifically, we train small bilingual models for which we control the amount of data for each language and the order of language exposure. To find evidence of shared multilingual representations, we turn to structural priming, a method used to study grammatical representations in humans. We first replicate previous crosslingual structural priming results and find that after controlling for training data quantity and language exposure, there are asymmetrical effects across language pairs and directions. We argue that this asymmetry may shape hypotheses about human structural priming effects. We also find that structural priming effects are less robust for less similar language pairs, highlighting potential limitations of crosslingual transfer learning and shared representations for typologically diverse languages. Catherine Arnett, Tyler A. Chang, James A. Michaelov, Ben Bergen 0001 |
ACL (1) | 4 |
| 2025 | Dissecting the Ullman Variations with a SCALPEL: Why do LLMs fail at Trivial Alterations to the False Belief Task?
Zhiqiang Pi, Annapurna Vadaparty, Ben Bergen 0001, Cameron R. Jones |
CogSci | 3 |
| 2025 | Judging the Judges: Displacing and Inverting the Turing test to Investigate the Interrogator
Ishika Rathi, Ben Bergen 0001, Cameron R. Jones |
CogSci | 2 |
| 2025 | Why do language models perform worse for morphologically complex languages?abstractLanguage models perform differently across languages. It has been previously suggested that morphological typology may explain some of this variability (Cotterell et al., 2018). We replicate previous analyses and find additional new evidence for a performance gap between agglutinative and fusional languages, where fusional languages, such as English, tend to have better language modeling performance than morphologically more complex languages like Turkish. We then propose and test three possible causes for this performance gap: morphological alignment of tokenizers, tokenization quality, and disparities in dataset sizes and measurement. To test the morphological alignment hypothesis, we present MorphScore, a tokenizer evaluation metric, and supporting datasets for 22 languages. We find some evidence that tokenization quality explains the performance gap, but none for the role of morphological alignment. Instead we find that the performance gap is most reduced when training datasets are of equivalent size across language types, but only when scaled according to the so-called “byte-premium”—the different encoding efficiencies of different languages and orthographies. These results suggest that languages of particular morphological types are not intrinsically advantaged or disadvantaged in language modeling. Differences in performance can be attributed to disparities in dataset size. These findings bear on ongoing efforts to improve performance for low-performing and under-resourced languages. Catherine Arnett, Ben Bergen 0001 |
COLING | 2 |
| 2025 | Are explicit belief representations necessary? A comparison between Large Language Models and Bayesian probabilistic modelsabstractDingyi Pan, Ben Bergen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dingyi Pan, Ben Bergen 0001 |
NAACL (Long Papers) | 2 |
| 2025 | Explaining and Mitigating Crosslingual Tokenizer InequitiesabstractThe number of tokens it takes to encode parallel text in different languages is known to vary. These disparities are called *token premiums*. Having high token premiums leads to less throughput during training and increases costs at inference.
In this paper, we show that even after controlling for dataset size, vocabulary size, and data content, monolingual tokenizers exhibit a wide range of token premiums across languages. To understand the cross-linguistic differences that cause these token premiums,
we train a suite of approximately 7,000 comparable monolingual tokenizers for 97 languages, manipulating tokenization algorithm vocabulary size, and dataset size. We measure token premiums and test for a relationship between factors such as data similarity (between tokenizer training and evaluation), vocabulary size, and pre-tokenization. We also investigate the role of language-specific features such as writing system and word length. We find that similarity between training and test data does not impact token premiums, but vocabulary size and pre-tokenization do. While simply increasing vocabulary size does not lead to reduced token premium effects, we can determine an "optimal" vocabulary size for each language to achieve significantly reduced token premium effects. We also train superword tokenizers which allow merges over whitespaces, and we find that they both reduce token premium effects and improve compression overall. Thus, intervening on the vocabulary size or the pre-tokenizer significantly reduces crosslingual token premium effects. Catherine Arnett, Tyler A. Chang, Stella Biderman, Ben Bergen 0001 |
NeurIPS | 4 |
| 2025 | Bigram Subnetworks: Mapping to Next Tokens in Transformer Language ModelsabstractIn Transformer language models, activation vectors transform from current token embeddings to next token predictions as they pass through the model. To isolate a minimal form of this transformation, we identify language model subnetworks that make bigram predictions, naive next token predictions based only on the current token. We find that bigram subnetworks can be found in fully trained language models up to 1B parameters, and these subnetworks are critical for model performance even when they consist of less than 0.2% of model parameters. Bigram subnetworks are concentrated in the first Transformer MLP layer, and they overlap significantly with subnetworks trained to optimally prune a given model. Mechanistically, the bigram subnetworks often recreate a pattern from the full models where the first layer induces a sharp change that aligns activations with next token predictions rather than current token representations. Our results demonstrate that bigram subnetworks comprise a minimal subset of parameters that are both necessary and sufficient for basic next token predictions in language models, and they help drive the transformation from current to next token activations in the residual stream. These subnetworks can lay a foundation for studying more complex language model circuits by building up from a minimal circuit. Tyler A. Chang, Ben Bergen 0001 |
NeurIPS | 2 |
| 2025 | Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and ScaleabstractWe show that across architecture (Transformer vs. Mamba vs. RWKV), training dataset (OpenWebText vs. The Pile), and scale (14 million parameters to 12 billion parameters), autoregressive language models exhibit highly consistent patterns of change in their behavior over the course of pretraining. Based on our analysis of over 1,400 language model checkpoints on over 110,000 tokens of English, we find that up to 98% of the variance in language model behavior at the word level can be explained by three simple heuristics: the unigram probability (frequency) of a given word, the $n$-gram probability of the word, and the semantic similarity between the word and its context. Furthermore, we see consistent behavioral phases in all language models, with their predicted probabilities for words overfitting to those words' $n$-gram probabilities for increasing $n$ over the course of training. Taken together, these results suggest that learning in neural language models may follow a similar trajectory irrespective of model details. James A. Michaelov, Roger Levy, Ben Bergen 0001 |
NeurIPS | 3 |
| 2024 | Does reading words help you to read minds? A comparison of humans and LLMs at a recursive mindreading task
Cameron R. Jones, Sean Trott, Ben Bergen 0001 |
CogSci | 3 |
| 2024 | Correlations between Multilingual Language Model Geometry and Crosslingual Transfer PerformanceabstractA common approach to interpreting multilingual language models is to evaluate their internal representations. For example, studies have found that languages occupy distinct subspaces in the models’ representation spaces, and geometric distances between languages often reflect linguistic properties such as language families and typological features. In our work, we investigate whether geometric distances between language representations correlate with zero-shot crosslingual transfer performance for POS-tagging and NER in three multilingual language models. We consider four distance metrics, including new metrics that identify a basis for a multilingual representation space that sorts axes based on their language-separability. We find that each distance metric either only moderately correlates or does not correlate with crosslingual transfer performance, and metrics do not generalize well across models, layers, and tasks. Although pairwise language separability is a reasonable predictor of crosslingual transfer, representational geometry overall is an inconsistent predictor for the crosslingual performance of multilingual language models. Cheril Shah, Yashashree Chandak, Atharv Mahesh Mane, Ben Bergen 0001, Tyler A. Chang |
LREC/COLING | 4 |
| 2024 | When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource LanguagesabstractMultilingual language models are widely used to extend NLP systems to low-resource languages.However, concrete evidence for the effects of multilinguality on language modeling performance in individual languages remains scarce.Here, we pre-train over 10,000 monolingual and multilingual language models for over 250 languages, including multiple language families that are under-studied in NLP.We assess how language modeling performance in each language varies as a function of (1) monolingual dataset size, (2) added multilingual dataset size, (3) linguistic similarity of the added languages, and (4) model size (up to 45M parameters).We find that in moderation, adding multilingual data improves low-resource language modeling performance, similar to increasing low-resource dataset sizes by up to 33%.Improvements depend on the syntactic similarity of the added multilingual data, with marginal additional effects of vocabulary overlap.However, high-resource languages consistently perform worse in multilingual pre-training scenarios.As dataset sizes increase, adding multilingual data begins to hurt performance for both low-resource and highresource languages, likely due to limited model capacity (the "curse of multilinguality").These results suggest that massively multilingual pretraining may not be optimal for any languages involved, but that more targeted models can significantly improve performance. Tyler A. Chang, Catherine Arnett, Zhuowen Tu, Ben Bergen 0001 |
EMNLP | 4 |
| 2024 | Does GPT-4 pass the Turing test?abstractWe evaluated GPT-4 in a public online Turing test.The best-performing GPT-4 prompt passed in 49.7% of games, outperforming ELIZA (22%) and GPT-3.5 (20%), but falling short of the baseline set by human participants (66%).Participants' decisions were based mainly on linguistic style (35%) and socioemotional traits (27%), supporting the idea that intelligence, narrowly conceived, is not sufficient to pass the Turing test.Participant knowledge about LLMs and number of games played positively correlated with accuracy in detecting AI, suggesting learning and practice as possible strategies to mitigate deception.Despite known limitations as a test of intelligence, we argue that the Turing test continues to be relevant as an assessment of naturalistic communication and deception.AI models with the ability to masquerade as humans could have widespread societal consequences, and we analyse the effectiveness of different strategies and criteria for judging humanlikeness. Cameron R. Jones, Ben Bergen 0001 |
NAACL-HLT | 2 |
| 2024 | Language Model Behavior: A Comprehensive SurveyabstractAbstract Transformer language models have received widespread public attention, yet their generated text is often surprising even to NLP researchers. In this survey, we discuss over 250 recent studies of English language model behavior before task-specific fine-tuning. Language models possess basic capabilities in syntax, semantics, pragmatics, world knowledge, and reasoning, but these capabilities are sensitive to specific inputs and surface features. Despite dramatic increases in generated text quality as models scale to hundreds of billions of parameters, the models are still prone to unfactual responses, commonsense errors, memorized text, and social biases. Many of these weaknesses can be framed as over-generalizations or under-generalizations of learned patterns in text. We synthesize recent results to highlight what is currently known about large language model capabilities, thus providing a resource for applied work and for research in adjacent fields that use language models. Tyler A. Chang, Ben Bergen 0001 |
Comput. Linguistics | 2 |
| 2024 | Do Multimodal Large Language Models and Humans Ground Language Similarly?abstractAbstract Large Language Models (LLMs) have been criticized for failing to connect linguistic meaning to the world—for failing to solve the “symbol grounding problem.” Multimodal Large Language Models (MLLMs) offer a potential solution to this challenge by combining linguistic representations and processing with other modalities. However, much is still unknown about exactly how and to what degree MLLMs integrate their distinct modalities—and whether the way they do so mirrors the mechanisms believed to underpin grounding in humans. In humans, it has been hypothesized that linguistic meaning is grounded through “embodied simulation,” the activation of sensorimotor and affective representations reflecting described experiences. Across four pre-registered studies, we adapt experimental techniques originally developed to investigate embodied simulation in human comprehenders to ask whether MLLMs are sensitive to sensorimotor features that are implied but not explicit in descriptions of an event. In Experiment 1, we find sensitivity to some features (color and shape) but not others (size, orientation, and volume). In Experiment 2, we identify likely bottlenecks to explain an MLLM’s lack of sensitivity. In Experiment 3, we find that despite sensitivity to implicit sensorimotor features, MLLMs cannot fully account for human behavior on the same task. Finally, in Experiment 4, we compare the psychometric predictive power of different MLLM architectures and find that ViLT, a single-stream architecture, is more predictive of human responses to one sensorimotor feature (shape) than CLIP, a dual-encoder architecture—despite being trained on orders of magnitude less data. These results reveal strengths and limitations in the ability of current MLLMs to integrate language with other modalities, and also shed light on the likely mechanisms underlying human language comprehension. Cameron R. Jones, Ben Bergen 0001, Sean Trott |
Comput. Linguistics | 2 |
| 2024 | SAM-Net: Self-Attention based Feature Matching with Spatial Transformers and Knowledge DistillationabstractIn this research paper, we introduce a novel approach to enhance the performance of 2D feature matching and pose estimation through the integration of a hierarchical attention mechanism and knowledge distillation. Our proposed hierarchical attention mechanism operates at multiple scales, enabling both global context awareness and precise matching of 2D features, which is crucial for various computer vision tasks. To further improve our model’s performance, we incorporate insights from an existing model PixLoc (Sarlin et al., 2021) through knowledge distillation, effectively acquiring its behavior and capabilities by ignoring dynamic objects. SAM-Net outperforms state-of-the-art methods, validated on both indoor and outdoor public datasets. For the indoor dataset, our approach achieves remarkable AUC (5°/10°/20°) scores of 55.31/71.70/83.37. Similarly, for the outdoor dataset, we demonstrate outstanding AUC values of 26.01/46.44/63.61. Furthermore, SAM-Net achieves top ranking among published methods in two public visual localization benchmarks, highlighting the real benefits of the proposed method. The code and test suite can be accessed at link.1 Ben Bergen 0001, Victor Domsa, Levente Tamas |
Expert Syst. Appl. | 1 |
| 2024 | Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and StabilityabstractAbstract How do language models learn to make predictions during pre-training? To study this, we extract learning curves from five autoregressive English language model pre-training runs, for 1M unseen tokens in context. We observe that the language models generate short repetitive phrases before learning to generate longer and more coherent text. We also find that individual tokens often exhibit sudden increases or decreases in loss that are surprisingly consistent across pre-training runs. To better understand these fluctuations, we quantify the final surprisal, within-run variability, age of acquisition, forgettability, and cross-run variability of learning curves for individual tokens in context. More frequent tokens reach lower final surprisals, exhibit less variability within and across pre-training runs, are learned earlier, and are less likely to be “forgotten” during pre-training. Higher n-gram probabilities further accentuate these effects. Independent of the target token, shorter and more frequent contexts correlate with marginally more stable and quickly acquired predictions. Based on our results, we argue for the existence of sequential learning dependencies between different model capabilities, and we characterize language model learning as early n-gram learning before gradual refinement of tail n-gram predictions. Tyler A. Chang, Zhuowen Tu, Ben Bergen 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | Comparing Humans and Large Language Models on an Experimental Protocol Inventory for Theory of Mind Evaluation (EPITOME)abstractAbstract We address a growing debate about the extent to which large language models (LLMs) produce behavior consistent with Theory of Mind (ToM) in humans. We present EPITOME: a battery of six experiments that tap diverse ToM capacities, including belief attribution, emotional inference, and pragmatic reasoning. We elicit a performance baseline from human participants for each task. We use the dataset to ask whether distributional linguistic information learned by LLMs is sufficient to explain ToM in humans. We compare performance of five LLMs to a baseline of responses from human comprehenders. Results are mixed. LLMs display considerable sensitivity to mental states and match human performance in several tasks. Yet, they commit systematic errors in others, especially those requiring pragmatic reasoning on the basis of mental state information. Such uneven performance indicates that human-level ToM may require resources beyond distributional information. Cameron R. Jones, Sean Trott, Ben Bergen 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | Can Peanuts Fall in Love with Distributional Semantics?
James A. Michaelov, Seana Coulson, Ben Bergen 0001 |
CogSci | 3 |
| 2023 | Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language Modelsabstractgrammatical knowledge-of parts of speech and grammatical patterns-is key to the capacity for linguistic generalization in humans.But how abstract is grammatical knowledge in large language models?In the human literature, compelling evidence for grammatical abstraction comes from structural priming.A sentence that shares the same grammatical structure as a preceding sentence is processed and produced more readily.Because confounds exist when using stimuli in a single language, evidence of abstraction is even more compelling from crosslingual structural priming, where use of a syntactic structure in one language primes an analogous structure in another language.We measure crosslingual structural priming in large language models, comparing model behavior to human experimental results from eight crosslingual experiments covering six languages, and four monolingual structural priming experiments in three non-English languages.We find evidence for abstract monolingual and crosslingual grammatical representations in the models that function similarly to those found in humans.These results demonstrate that grammatical representations in multilingual language models are not only similar across languages, but they can causally influence text produced in different languages. James A. Michaelov, Catherine Arnett, Tyler A. Chang, Ben Bergen 0001 |
EMNLP | 4 |
| 2022 | Does Contextual Diversity Hinder Early Word Acquisition?
Tyler A. Chang, Ben Bergen 0001 |
CogSci | 2 |
| 2022 | Distrubutional Semantics Still Can't Account for Affordances
Cameron R. Jones, Tyler A. Chang, Seana Coulson, James A. Michaelov, Sean Trott, Ben Bergen 0001 |
CogSci | 6 |
| 2022 | Can a pressure against homophones explain phonological neighborhoods?
Sean Trott, Ben Bergen 0001 |
CogSci | 2 |
| 2022 | Do Language Models Make Human-like Predictions about the Coreferents of Italian Anaphoric Zero Pronouns?abstractSome languages allow arguments to be omitted in certain contexts. Yet human language comprehenders reliably infer the intended referents of these zero pronouns, in part because they construct expectations about which referents are more likely. We ask whether Neural Language Models also extract the same expectations. We test whether 12 contemporary language models display expectations that reflect human behavior when exposed to sentences with zero pronouns from five behavioral experiments conducted in Italian by Carminati (2005). We find that three models - XGLM 2.9B, 4.5B, and 7.5B - capture the human behavior from all the experiments, with others successfully modeling some of the results. This result suggests that human expectations about coreference can be derived from exposure to language, and also indicates features of language models that allow them to better reflect human behavior. James A. Michaelov, Ben Bergen 0001 |
COLING | 2 |
| 2022 | Collateral facilitation in humans and language modelsabstractAre the predictions of humans and language models affected by similar things?Research suggests that while comprehending language, humans make predictions about upcoming words, with more predictable words being processed more easily.However, evidence also shows that humans display a similar processing advantage for highly anomalous words when these words are semantically related to the preceding context or to the most probable continuation.Using stimuli from 3 psycholinguistic experiments, we find that this is also almost always also the case for 8 contemporary transformer language models (BERT, ALBERT, RoBERTa, XLM-R, GPT-2, GPT-Neo, GPT-J, and XGLM).We then discuss the implications of this phenomenon for our understanding of both human language comprehension and the predictions made by language models. James A. Michaelov, Ben Bergen 0001 |
CoNLL | 2 |
| 2022 | The Geometry of Multilingual Language Model RepresentationsabstractWe assess how multilingual language models maintain a shared multilingual representation space while still encoding language-sensitive information in each language.Using XLM-R as a case study, we show that languages occupy similar linear subspaces after mean-centering, evaluated based on causal effects on language modeling performance and direct comparisons between subspaces for 88 languages.The subspace means differ along language-sensitive axes that are relatively stable throughout middle layers, and these axes encode information such as token vocabularies.Shifting representations by language means is sufficient to induce token predictions in different languages.However, we also identify stable languageneutral axes that encode information such as token positions and part-of-speech.We visualize representations projected onto languagesensitive and language-neutral axes, identifying language family and part-of-speech clusters, along with spirals, toruses, and curves representing token position information.These results demonstrate that multilingual language models encode information along orthogonal language-sensitive and language-neutral axes, allowing the models to extract a variety of features for downstream tasks and cross-lingual transfer learning. Tyler A. Chang, Zhuowen Tu, Ben Bergen 0001 |
EMNLP | 3 |
| 2022 | Word Acquisition in Neural Language ModelsabstractAbstract We investigate how neural language models acquire individual words during training, extracting learning curves and ages of acquisition for over 600 words on the MacArthur-Bates Communicative Development Inventory (Fenson et al., 2007). Drawing on studies of word acquisition in children, we evaluate multiple predictors for words’ ages of acquisition in LSTMs, BERT, and GPT-2. We find that the effects of concreteness, word length, and lexical class are pointedly different in children and language models, reinforcing the importance of interaction and sensorimotor experience in child language acquisition. Language models rely far more on word frequency than children, but, like children, they exhibit slower learning of words in longer utterances. Interestingly, models follow consistent patterns during training for both unidirectional and bidirectional models, and for both LSTM and Transformer architectures. Models predict based on unigram token frequencies early in training, before transitioning loosely to bigram probabilities, eventually converging on more nuanced predictions. These results shed light on the role of distributional learning mechanisms in children, while also providing insights for more human-like language acquisition in language models. Tyler A. Chang, Ben Bergen 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2021 | RAW-C: Relatedness of Ambiguous Words in Context (A New Lexical Resource for English)abstractSean Trott, Benjamin Bergen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sean Trott, Ben Bergen 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | The Role of Physical Inference in Pronoun Resolution
Cameron R. Jones, Ben Bergen 0001 |
CogSci | 2 |
| 2021 | Different kinds of cognitive plausibility: why are transformers better than RNNs at predicting N400 amplitude?
James A. Michaelov, Megan D. Bardolph, Seana Coulson, Ben Bergen 0001 |
CogSci | 4 |
| 2020 | Effects of Battle and Journey Metaphors on Charitable Donations for Cancer Patients
Alex Liebscher, Sean Trott, Ben Bergen 0001 |
CogSci | 3 |
| 2020 | How well does surprisal explain N400 amplitude under different experimental conditions?abstractWe investigate the extent to which word surprisal can be used to predict a neural measure of human language processing difficulty - the N400. To do this, we use recurrent neural networks to calculate the surprisal of stimuli from previously published neurolinguistic studies of the N400. We find that surprisal can predict N400 amplitude in a wide range of cases, and the cases where it cannot do so provide valuable insight into the neurocognitive processes underlying the response. James A. Michaelov, Ben Bergen 0001 |
CoNLL | 2 |
| 2019 | Prosodic cues signal the intent of potential indirect requests
Sean Trott, Stefanie Reed, Victor Ferreira, Ben Bergen 0001 |
CogSci | 4 |
| 2019 | Sub-morphemic form-meaning systematicity: the impact of onset phones on word concreteness
Sean Trott, Arturs Semenuks, Ben Bergen 0001 |
CogSci | 3 |
| 2017 | When metaphors in the mind become metaphors in the mouth: Documenting the emergence of a new system of linguistic metaphors for time
Rose Hendricks, Tyler Marghetis, Ben Bergen 0001 |
CogSci | 3 |
| 2017 | Listeners integrate speech, gesture, and discourse structure to interpret the temporal structure of complex events
Andrea Nishimi, Esther Walker, Ben Bergen 0001, Tyler Marghetis |
CogSci | 3 |
| 2016 | Finding Non-Arbitrary Form-Meaning Systematicity Using String-Metric Learning for Kernel RegressionabstractArbitrariness of the sign-the notion that the forms of words are unrelated to their meanings-is an underlying assumption of many linguistic theories.Two lines of research have recently challenged this assumption, but they produce differing characterizations of non-arbitrariness in language.Behavioral and corpus studies have confirmed the validity of localized form-meaning patterns manifested in limited subsets of the lexicon.Meanwhile, global (lexicon-wide) statistical analyses instead find diffuse form-meaning systematicity across the lexicon as a whole.We bridge the gap with an approach that can detect both local and global formmeaning systematicity in language.In the kernel regression formulation we introduce, form-meaning relationships can be used to predict words' distributional semantic vectors from their forms.Furthermore, we introduce a novel metric learning algorithm that can learn weighted edit distances that minimize kernel regression error.Our results suggest that the English lexicon exhibits far more global form-meaning systematicity than previously discovered, and that much of this systematicity is focused in localized formmeaning patterns. E. Dario Gutiérrez, Roger Levy, Ben Bergen 0001 |
ACL (1) | 3 |
| 2016 | Literal and Metaphorical Senses in Compositional Distributional Semantic ModelsabstractMetaphorical expressions are pervasive in natural language and pose a substantial challenge for computational semantics.The inherent compositionality of metaphor makes it an important test case for compositional distributional semantic models (CDSMs).This paper is the first to investigate whether metaphorical composition warrants a distinct treatment in the CDSM framework.We propose a method to learn metaphors as linear transformations in a vector space and find that, across a variety of semantic domains, explicitly modeling metaphor improves the resulting semantic representations.We then use these representations in a metaphor identification task, achieving a high performance of 0.82 in terms of F-score. E. Dario Gutiérrez, Ekaterina Shutova, Tyler Marghetis, Ben Bergen 0001 |
ACL (1) | 4 |
| 2016 | Left-right mental timeline is robust to visuospatial and verbal interference
Rose Hendricks, Esther Walker, Ben Bergen 0001, Lera Boroditsky, Rafael E. Núñez |
CogSci | 3 |
| 2015 | The mental number-line spreads by gestural contagion
Tyler Marghetis, Luke Eberle, Ben Bergen 0001 |
CogSci | 3 |
| 2014 | Origins of time: New insights into the psychological foundations of time
Katharine Tillman, Esther Walker, Tyler Marghetis, Andrea Bender, Sieghard Beller, Mahesh Srinivasan, David Barner, Julio Santiago, Ben Bergen 0001, Rafael E. Núñez, Daniel Casasanto, Lera Boroditsky |
CogSci | 9 |
| 2014 | Does beat perception rely on the covert use of the motor system?
Esther Walker, Benjamin Stillerman, John Iversen, Aniruddh Patel, Ben Bergen 0001 |
CogSci | 5 |
| 2013 | When Tuesday comes before Threesday: Cross-linguistic differences in numerical transparency of time words predicts temporal reasoning strategy and performance
Ben Bergen 0001 |
CogSci | 2 |
| 2013 | Placing Numbers in Behavioral Space: Activity-Specific Interactions between Number and Space with a Single Response Button
Tyler Marghetis, Jasmeen Kanwal, Ben Bergen 0001 |
CogSci | 3 |
| 2013 | Later events lie behind her, but not behind you: Compatibility effects for temporal sequences along the sagittal axis depend on perspective
Esther Walker, Ben Bergen 0001, Rafael E. Núñez |
CogSci | 2 |
| 2012 | Towards a cognitive science of literary style: Perspective-taking in processing omniscient versus objective voice
Manami Sato, Hiromu Sakai, Jennifer Wu, Ben Bergen 0001 |
CogSci | 4 |
| 2011 | Making SNAP Judgments: Rethinking the Spatial Representation of Number
Tyler Marghetis, Esther Walker, Ben Bergen 0001, Rafael E. Núñez |
CogSci | 3 |
| 2011 | Grammatical aspect in language production: Using gesture to reveal event representations
Fey Parrill, Ben Bergen 0001, Patricia Lichtenstein |
CogSci | 2 |