EDBT 2026 Demo / reviewers in the wild / expert
Edward Gibson
dblp:93/5649
· DBLP profile ↗
40ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-5912-883XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 4 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Readers make targeted regressions to plausible errors in reanalysis of "noisy-channel garden-path" sentencesabstractA key question in psycholinguistics is how inferences about the meaning of linguistic input unfold incrementally a comprehender's mind.In this work, we study reading dynamics for "noisy-channel garden-path" sentences, which temporarily appear well-formed but feature late-appearing violations of expectation that can be resolved not by inferring an alternative syntactic structure, but by inferring the presence of an error.We find evidence for targeted regressions -eye movements towards regions that are promising loci of possible errors in light of later-arriving information, showing patterns consistent with the posterior inferences of a model of noisy-channel processing with reanalysis.We discuss the implications of these findings for theories of noisy-channel language comprehension and information-theoretic explanations of reading dynamics. Thomas Hikaru Clark, Roger Levy, Edward Gibson |
CoNLL | 3 |
| 2025 | A Model of Approximate and Incremental Noisy-Channel Language Processing
Thomas Hikaru Clark, Jacob Hoover Vigly, Edward Gibson, Roger Levy |
CogSci | 3 |
| 2025 | How grammatical gender supports efficient communication
Dorothée B. Hoppe, Edward Gibson, Jacolien van Rij, Petra Hendriks, Michael Ramscar |
CogSci | 2 |
| 2025 | Resource-Rational Noisy-Channel Language Processing: Testing the Effect of Algorithmic Constraints on InferencesabstractHuman language use is robust to errors: comprehenders can and do mentally correct utterances that are implausible or anomalous.How are humans able to solve these problems in real time, picking out alternatives from an unbounded space of options using limited cognitive resources?And can language models trained on next-word prediction for typical language be augmented to handle language anomalies in a human-like way?Using a language model as a prior and an error model to encode likelihoods, we use Sequential Monte Carlo with optional rejuvenation to perform incremental and approximate probabilistic inference over intended sentences and production errors.We demonstrate that the model captures previously established patterns in human sentence processing, and that a trade-off between human-like noisy-channel inferences and computational resources falls out of this model.From a psycholinguistic perspective, our results offer a candidate algorithmic model of rational inference in language processing.From an NLP perspective, our results showcase how to elicit human-like noisy-channel inference behavior from a relatively small LLM while controlling the amount of computation available during inference.Our model is implemented in the Gen.jl probabilistic programming language, and our code is available at https://github. com/thomashikaru/noisy_channel_model. Thomas Hikaru Clark, Jacob Hoover Vigly, Edward Gibson, Roger Levy |
EMNLP | 3 |
| 2024 | Availability, informatively and burstiness: Why average corpus measures are an inaccurate guide to surprisal in language
Edward Gibson, Michael Ramscar |
CogSci | 2 |
| 2024 | Inferring errors and intended meanings with a generative model of language production in aphasia
Thomas Hikaru Clark, Edward Gibson, Roger Levy |
CogSci | 2 |
| 2024 | Even Laypeople Use Legalese
Eric Martinez, Francis Mollica, Edward Gibson |
CogSci | 3 |
| 2024 | Mis-Heard Lyrics: an Ecologically-Valid Test of Noisy Channel Processing
Moshe Poliak, Hannah Kimura, Edward Gibson |
CogSci | 3 |
| 2023 | A fine-grained comparison of pragmatic language understanding in humans and language modelsabstractPragmatics and non-literal language understanding are essential to human communication, and present a long-standing challenge for artificial language models.We perform a finegrained comparison of language models and humans on seven pragmatic phenomena, using zero-shot prompting on an expert-curated set of English materials.We ask whether models (1) select pragmatic interpretations of speaker utterances, (2) make similar error patterns as humans, and (3) use similar linguistic cues as humans to solve the tasks.We find that the largest models achieve high accuracy and match human error patterns: within incorrect responses, models favor literal interpretations over heuristic-based distractors.We also find preliminary evidence that models and humans are sensitive to similar linguistic cues.Our results suggest that pragmatic behaviors can emerge in models without explicitly constructed representations of mental states.However, models tend to struggle with phenomena relying on social expectation violations. Jennifer Hu 0001, Sammy Floyd, Olessia Jouravlev, Evelina Fedorenko, Edward Gibson |
ACL (1) | 5 |
| 2023 | Transformer-Maze: An Accessible Incremental Processing Measurement Tool
Annika Heuser, Edward Gibson |
CogSci | 2 |
| 2023 | Can Language Models Be Tricked by Language Illusions? Easier with Syntax, Harder with SemanticsabstractLanguage models (LMs) have been argued to overlap substantially with human beings in grammaticality judgment tasks.But when humans systematically make errors in language processing, should we expect LMs to behave like cognitive models of language and mimic human behavior?We answer this question by investigating LMs' more subtle judgments associated with "language illusions" -sentences that are vague in meaning, implausible, or ungrammatical but receive unexpectedly high acceptability judgments by humans.We looked at three illusions: the comparative illusion (e.g."More people have been to Russia than I have"), the depth-charge illusion (e.g."No head injury is too trivial to be ignored"), and the negative polarity item (NPI) illusion (e.g."The hunter who no villager believed to be trustworthy will ever shoot a bear").We found that probabilities represented by LMs were more likely to align with human judgments of being "tricked" by the NPI illusion which examines a structural dependency, compared to the comparative and the depth-charge illusions which require sophisticated semantic understanding.No single LM or metric yielded results that are entirely consistent with human behavior.Ultimately, we show that LMs are limited both in their construal as cognitive models of human language processing and in their capacity to recognize nuanced but critical information in complicated language materials. Edward Gibson, Forrest Davis |
CoNLL | 2 |
| 2022 | Evidence for Availability Effects on Speaker Choice in the Russian Comparative Alternation
Thomas Hikaru Clark, Ethan Wilcox, Edward Gibson, Roger Levy |
CogSci | 3 |
| 2022 | Hering's opponent-colors theory fails a key test in a non-Western culture
Bevil R. Conway, Saima Malik-Moraleda, Edward Gibson |
CogSci | 3 |
| 2022 | So much for plain language: An analysis of the accessibility of United States federal laws (1951-2009)
Eric Martinez, Francis Mollica, Edward Gibson |
CogSci | 3 |
| 2022 | Dimensions of Diversity in Spatial Cognition: Culture, Context, Age, and Ability
Benjamin Pitt, Holly Huey, Matthew Jordan, Yuval Hart, Moira R. Dillon, Roberto Bottini, Alexandra Carstensen, Isabelle Boni, Steve Piantadosi, Edward Gibson, Tyler Marghetis, Kevin J. Holmes, Maya Star-Lack, Sandra Chacon |
CogSci | 10 |
| 2022 | Rational Inference from Number Agreement Mismatch
Edward Gibson, Roger Levy |
CogSci | 2 |
| 2021 | What did I sign? A study of the impenetrability of legalese in contracts
Eric Martinez, Francis Mollica, Anita Podrug, Edward Gibson |
CogSci | 5 |
| 2021 | Variation in spatial concepts: Different frames of reference on different axes
Benjamin Pitt, Alexandra Carstensen, Edward Gibson, Steve Piantadosi |
CogSci | 3 |
| 2020 | Multi-directional mappings in the minds of the Tsimane': Size, time, and number on three spatial axes
Benjamin Pitt, Daniel Casasanto, Stephen Ferrigno, Edward Gibson, Steve Piantadosi |
CogSci | 4 |
| 2020 | Exact number concepts depend on language
Benjamin Pitt, Edward Gibson, Steve Piantadosi |
CogSci | 2 |
| 2020 | Procedures and principles of number: Evidence from the Tsimane'
Benjamin Pitt, Rose M. Schneider, Stephen Ferrigno, Edward Gibson, David Barner, Steve Piantadosi |
CogSci | 4 |
| 2019 | Math ability varies independently of number estimation in the Tsimané
Samuel J. Cheyette, Benjamin Pitt, Steve Piantadosi, Edward Gibson |
CogSci | 4 |
| 2019 | Verb Frequency Explains the Unacceptability of Factive and Manner-of-speaking Islands in English
Yingtong Liu, Rachel Ryskin, Richard Futrell, Edward Gibson |
CogSci | 4 |
| 2018 | Percepts and Concepts Across Cultures
Asifa Majid, Edward Gibson, Tanya M. Luhrmann, Josh H. McDermott, Artin Arshamian |
CogSci | 2 |
| 2018 | The Natural Stories Corpus
Richard Futrell, Edward Gibson, Hal Tily, Idan A. Blank, Anastasia Vishnevetsky, Steve Piantadosi, Evelina Fedorenko |
LREC | 2 |
| 2017 | Comprehenders Model the Nature of Noise in the Environment
Rachel Ryskin, Richard Futrell, Edward Gibson |
CogSci | 3 |
| 2015 | Experiments with Generative Models for Dependency Tree LinearizationabstractWe present experiments with generative models for linearization of unordered labeled syntactic dependency trees (Belz et al., 2011;Rajkumar and White, 2014).Our linearization models are derived from generative models for dependency structure (Eisner, 1996).We present a series of generative dependency models designed to capture successively more information about ordering constraints among sister dependents.We give a dynamic programming algorithm for computing the conditional probability of word orders given tree structures under these models.The models are tested on corpora of 11 languages using test-set likelihood, and human ratings for generated forms are collected for English.Our models benefit from representing local order constraints among sisters and from backing off to less sparse distributions, including distributions not conditioned on the head. Richard Futrell, Edward Gibson |
EMNLP | 2 |
| 2014 | Language for Communication: Language as Rational Inference
Edward Gibson |
COLING | 1 |
| 2012 | Verb omission errors: Evidence of rational processing of noisy language inputs
Leon Bergen, Roger Levy, Edward Gibson |
CogSci | 3 |
| 2011 | Storage and computation in syntax: Evidence from relative clause priming
Melissa Troyer, Timothy J. O'Donnell, Evelina Fedorenko, Edward Gibson |
CogSci | 4 |
| 2011 | Thinking for Seeing: Enculturation of Visual-Referential Expertise as Demonstrated by Photo-Triggered Perceptual Reorganization of Two-Tone "Mooney" Images
Jennifer M. D. Yoon, Nathan Witthoft, Jonathan Winawer, Michael C. Frank, Edward Gibson, Ellen M. Markman |
CogSci | 5 |
| 2008 | Sentence and Text Comprehension: Evidence from Human Language Processing
Edward Gibson |
NLDB | 1 |
| 2006 | A comparison of inter-transcriber reliability for two systems of prosodic annotation: rap (rhythm and pitch) and toBI (tones and break indices)abstractAgreement was investigated among five labelers for the use of two prosodic annotation systems: the ToBI (Tones and Break Indices) system [1,2] and the RaP (Rhythm and Pitch) system [3]. Each system permits the labeling of pitch accents and two levels of phrasal boundaries; RaP also permits labeling of speech rhythm and distinguishes multiple levels of prominence on syllables. After training with computerized materials and getting expert feedback, coders applied each system to a corpus of read and spontaneous speech (36 minutes for ToBI and 19 for RaP). Inter-coder reliability was computed according to two metrics: transcriber-syllable-pairs and the kappa statistic. High agreement was obtained for both systems for pitch accent presence, pitch accent type, boundary presence, boundary type, and, for RaP, presence and strength of metrical prominences. Agreement levels for ToBI were similar to those of previous studies [4,5], indicating that participants were proficient coders. Moreover, the high level of agreement demonstrated for the RaP system indicates that RaP is a viable alternative to ToBI for prosodic labeling of large speech corpora. Laura Dilley, Mara Breen, Marti Bolivar, John Kraemer, Edward Gibson |
INTERSPEECH | 5 |
| 2005 | Measuring human readability of machine generated text: three case studies in speech recognition and machine translationabstractWe present highlights from three experiments that test the readability of current state-of-the art system output from: (1) an automated English speech-to-text (SST) system; (2) a text-based Arabic-to-English machine translation (MT) system; and (3) an audio-based Arabic-to-English MT process. We measure readability in terms of reaction time and passage comprehension in each case, applying standard psycholinguistic testing procedures and a modified version of the standard defense language proficiency test for Arabic called the DLPT*. We learned that: (1) subjects are slowed down by about 25% when reading system STT output; (2) text-based MT systems enable an English speaker to pass Arabic Level 2 on the DLPT*; and (3) audio-based MT systems do not enable English speakers to pass Arabic Level 2. We intend for these generic measures of readability to predict performance of more application-specific tasks. Douglas A. Jones, Edward Gibson, Wade Shen, Neil Granoien, Martha Herzog, Douglas A. Reynolds, Clifford J. Weinstein |
ICASSP (5) | 2 |
| 2005 | Representing Discourse Coherence: A Corpus-Based StudyabstractThis article aims to present a set of discourse structure relations that are easy to code and to develop criteria for an appropriate data structure for representing these relations. Discourse structure here refers to informational relations that hold between sentences in a discourse. The set of discourse relations introduced here is based on Hobbs (1985). We present a method for annotating discourse coherence structures that we used to manually annotate a database of 135 texts from the Wall Street Journal and the AP Newswire. Alltexts were independently annotated by two annotators. Kappa values of greater than 0.8 indicated good interannotator agreement. We furthermore present evidence that trees are not a descriptively adequate data structure for representing discourse structure: In coherence structures of naturally occurring texts, we found many different kinds of crossed dependencies, as well as many nodes with multiple parents. The claims are supported by statistical results from our hand-annotated database of 135 texts. Florian Wolf 0003, Edward Gibson |
Comput. Linguistics | 2 |
| 2004 | Paragraph-, Word-, and Coherence-based Approaches to Sentence Ranking: A Comparison of Algorithm and Human PerformanceabstractSentence ranking is a crucial part of generating text summaries. We compared human sentence rankings obtained in a psycholinguistic experiment to three different approaches to sentence ranking: A simple paragraph-based approach intended as a baseline, two word-based approaches, and two coherence-based approaches. In the paragraph-based approach, sentences in the beginning of paragraphs received higher importance ratings than other sentences. The word-based approaches determined sentence rankings based on relative word frequencies (Luhn (1958); Salton & Buckley (1988)). Coherence-based approaches determined sentence rankings based on some property of the coherence structure of a text (Marcu (2000); Page et al. (1998)). Our results suggest poor performance for the simple paragraph-based approach, whereas word-based approaches perform remarkably well. The best performance was achieved by a coherence-based approach where coherence structures are represented in a non-tree structure. Most approaches also outperformed the commercially available MSWord summarizer. Florian Wolf 0003, Edward Gibson |
ACL | 2 |
| 2004 | Representing discourse coherence: A corpus-based analysis
Florian Wolf 0003, Edward Gibson |
COLING | 2 |
| 2003 | Measuring the readability of automatic speech-to-text transcriptsabstractAbstract • This paper reports initial results from a novel psycholinguistic study that measures the readability of several types of speech transcripts. We define a four-part figure of merit to measure readability: accuracy of answers to comprehension questions, reaction-time for passage reading, reaction-time for question answering and a subjective rating of passage difficulty. We present results from an experiment with 28 test subjects reading transcripts in four experimental conditions. 1. Douglas A. Jones, Florian Wolf 0003, Edward Gibson, Elliott Williams, Evelina Fedorenko, Douglas A. Reynolds, Marc A. Zissman |
INTERSPEECH | 3 |
| 1990 | Memory Capacity and Sentence ProcessingabstractThe limited capacity of working memory is intrinsic to human sentence processing, and therefore must be addressed by any theory of human sentence processing. This paper gives a theory of garden-path effects and processing overload that is based on simple assumptions about human short term memory capacity. Edward Gibson |
ACL | 1 |
| 1990 | A Computational Theory of Processing Overload and Garden-Path Effects
Edward Gibson |
COLING | 1 |