EDBT 2026 Demo / reviewers in the wild / expert
Jennifer Hu 0001
dblp:217/1862
· DBLP profile ↗
21ranked-venue papers
10as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 10 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On Emergent Social World Models - Evidence for Functional Integration of Theory of Mind and Pragmatic Reasoning in Language ModelsabstractThis paper investigates whether LMs recruit shared computational mechanisms for general Theory of Mind (ToM) and language-specific pragmatic reasoning in order to contribute to the general question of whether LMs may be said to have emergent "social world models," i.e., representations of mental states that are repurposed across tasks (the functional integration hypothesis).Using behavioral evaluations and causal-mechanistic experiments via functional localization methods inspired by cognitive neuroscience, we analyze LMs' performance across seven subcategories of ToM abilities (Beaudoin et al., 2020) on a substantially larger localizer dataset than used in prior like-minded work.Results from stringent hypothesis-driven statistical testing offer suggestive evidence for the functional integration hypothesis, indicating that LMs may develop interconnected "social world models" rather than isolated competencies.This work contributes novel ToM localizer data, methodological refinements to functional localization techniques, and empirical insights into the emergence of social cognition in artificial systems. Polina Tsvilodub, Jan-Felix Klumpp, Amir Pour, Jennifer Hu 0001, Michael Franke |
ACL (1) | 4 |
| 2026 | What Can String Probability Tell Us About Grammaticality?abstractAbstract What have language models (LMs) learned about grammar? This question remains hotly debated, with major ramifications for linguistic theory. However, since probability and grammaticality are distinct notions in linguistics, it is not obvious what string probabilities can reveal about an LM’s underlying grammatical knowledge. We present a theoretical analysis of the relationship between grammar, meaning, and string probability, based on simple assumptions about the generative process of corpus data. Our framework makes three predictions, which we validate empirically using 280K sentence pairs in English and Chinese: (1) correlation between the probability of strings within minimal pairs, i.e., string pairs with minimal semantic differences; (2) correlation between models’ and humans’ deltas within minimal pairs; and (3) poor separation in probability space between unpaired grammatical and ungrammatical strings. Our analyses give theoretical grounding for using probability to learn about LMs’ structural knowledge, and suggest directions for future work in LM grammatical evaluation. Jennifer Hu 0001, Ethan Wilcox, Siyuan Song, Kyle Mahowald, Roger Levy |
Trans. Assoc. Comput. Linguistics | 1 |
| 2025 | Making Sense of Nonsense
Jennifer Hu 0001, Felix Sosa, Tomer D. Ullman |
CogSci | 1 |
| 2025 | Language production is harder than comprehension for children and language models
Jennifer Hu 0001, Alvin Wei Ming Tan, Steven Y. Feng, Michael C. Frank |
CogSci | 1 |
| 2025 | The Uncanny Valley meets the Humorous Hill: Things are funny when they match a pattern but fall short on quality
Antara Raaghavi Bhattacharya, Jennifer Hu 0001, Tomer D. Ullman |
CogSci | 2 |
| 2025 | One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversityabstractSonia Krishna Murthy, Tomer Ullman, Jennifer Hu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sonia K. Murthy, Tomer D. Ullman, Jennifer Hu 0001 |
NAACL (Long Papers) | 3 |
| 2025 | Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models
Anna A. Ivanova, Aalok Sathe, Benjamin Lipkin, Unnathi U. Kumar, Setayesh Radkani, Thomas Hikaru Clark, Carina Kauf, Jennifer Hu 0001, R. T. Pramod, Gabriel Grand, Vivian C. Paulun, Maria Ryskina, Ekin Akyürek, Ethan Wilcox, Nafisa Rashid, Leshem Choshen, Roger Levy, Evelina Fedorenko, Josh Tenenbaum, Jacob Andreas |
Trans. Assoc. Comput. Linguistics | 8 |
| 2024 | Shades of Zero: Distinguishing impossibility from inconceivability
Jennifer Hu 0001, Felix Sosa, Tomer D. Ullman |
CogSci | 1 |
| 2024 | The Task Task: Creative problem generation in humans and language models
Junyi Chu, Jennifer Hu 0001, Tomer D. Ullman |
CogSci | 2 |
| 2023 | A fine-grained comparison of pragmatic language understanding in humans and language modelsabstractPragmatics and non-literal language understanding are essential to human communication, and present a long-standing challenge for artificial language models.We perform a finegrained comparison of language models and humans on seven pragmatic phenomena, using zero-shot prompting on an expert-curated set of English materials.We ask whether models (1) select pragmatic interpretations of speaker utterances, (2) make similar error patterns as humans, and (3) use similar linguistic cues as humans to solve the tasks.We find that the largest models achieve high accuracy and match human error patterns: within incorrect responses, models favor literal interpretations over heuristic-based distractors.We also find preliminary evidence that models and humans are sensitive to similar linguistic cues.Our results suggest that pragmatic behaviors can emerge in models without explicitly constructed representations of mental states.However, models tend to struggle with phenomena relying on social expectation violations. Jennifer Hu 0001, Sammy Floyd, Olessia Jouravlev, Evelina Fedorenko, Edward Gibson |
ACL (1) | 1 |
| 2023 | I Cast Detect Thoughts: Learning to Converse and Guide with Intents and Theory-of-Mind in Dungeons and DragonsabstractPei Zhou, Andrew Zhu, Jennifer Hu, Jay Pujara, Xiang Ren, Chris Callison-Burch, Yejin Choi, Prithviraj Ammanabrolu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Andrew Zhu, Jennifer Hu 0001, Jay Pujara, Xiang Ren 0001, Chris Callison-Burch, Yejin Choi 0001, Prithviraj Ammanabrolu |
ACL (1) | 3 |
| 2023 | Prompting is not a substitute for probability measurements in large language modelsabstractPrompting is now a dominant method for evaluating the linguistic knowledge of large language models (LLMs).While other methods directly read out models' probability distributions over strings, prompting requires models to access this internal information by processing linguistic input, thereby implicitly testing a new type of emergent ability: metalinguistic judgment.In this study, we compare metalinguistic prompting and direct probability measurements as ways of measuring models' linguistic knowledge.Broadly, we find that LLMs' metalinguistic judgments are inferior to quantities directly derived from representations.Furthermore, consistency gets worse as the prompt query diverges from direct measurements of next-word probabilities.Our findings suggest that negative results relying on metalinguistic prompts cannot be taken as conclusive evidence that an LLM lacks a particular linguistic generalization.Our results also highlight the value that is lost with the move to closed APIs where access to probability distributions is limited. Jennifer Hu 0001, Roger Levy |
EMNLP | 1 |
| 2023 | Expectations over Unspoken Alternatives Predict Pragmatic InferencesabstractAbstract Scalar inferences (SI) are a signature example of how humans interpret language based on unspoken alternatives. While empirical studies have demonstrated that human SI rates are highly variable—both within instances of a single scale, and across different scales—there have been few proposals that quantitatively explain both cross- and within-scale variation. Furthermore, while it is generally assumed that SIs arise through reasoning about unspoken alternatives, it remains debated whether humans reason about alternatives as linguistic forms, or at the level of concepts. Here, we test a shared mechanism explaining SI rates within and across scales: context-driven expectations about the unspoken alternatives. Using neural language models to approximate human predictive distributions, we find that SI rates are captured by the expectedness of the strong scalemate as an alternative. Crucially, however, expectedness robustly predicts cross-scale variation only under a meaning-based view of alternatives. Our results suggest that pragmatic inferences arise from context-driven expectations over alternatives, and these expectations operate at the level of concepts.1 Jennifer Hu 0001, Roger Levy, Judith Degen, Sebastian Schuster 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2022 | Teasing apart models of pragmatics using optimal reference game design
Irene Zhou, Jennifer Hu 0001, Roger Levy, Noga Zaslavsky |
CogSci | 2 |
| 2021 | Competition from novel features drives scalar inferences in reference games
Jennifer Hu 0001, Noga Zaslavsky, Roger Levy |
CogSci | 1 |
| 2021 | Empirical Support for a Rate-Distortion Account of Pragmatic Reasoning
Irene Zhou, Jennifer Hu 0001, Roger Levy, Noga Zaslavsky |
CogSci | 2 |
| 2021 | Controlled Evaluation of Grammatical Knowledge in Mandarin Chinese Language ModelsabstractPrior work has shown that structural supervision helps English language models learn generalizations about syntactic phenomena such as subject-verb agreement.However, it remains unclear if such an inductive bias would also improve language models' ability to learn grammatical dependencies in typologically different languages.Here we investigate this question in Mandarin Chinese, which has a logographic, largely syllable-based writing system; different word order; and sparser morphology than English.We train LSTMs, Recurrent Neural Network Grammars, Transformer language models, and Transformerparameterized generative parsing models on two Mandarin Chinese datasets of different sizes.We evaluate the models' ability to learn different aspects of Mandarin grammar that assess syntactic and semantic relationships.We find suggestive evidence that structural supervision helps with representing syntactic state across intervening content and improves performance in low-data settings, suggesting that the benefits of hierarchical inductive biases in acquiring dependency relationships may extend beyond English. Jennifer Hu 0001, Roger Levy |
EMNLP (1) | 2 |
| 2020 | A Systematic Assessment of Syntactic Generalization in Neural Language ModelsabstractWhile state-of-the-art neural network models continue to achieve lower perplexity scores on language modeling benchmarks, it remains unknown whether optimizing for broad-coverage predictive performance leads to human-like syntactic knowledge.Furthermore, existing work has not provided a clear picture about the model properties required to produce proper syntactic generalizations.We present a systematic evaluation of the syntactic knowledge of neural language models, testing 20 combinations of model types and data sizes on a set of 34 English-language syntactic test suites.We find substantial differences in syntactic generalization performance by model architecture, with sequential models underperforming other architectures.Factorially manipulating model architecture and training dataset size (1M-40M words), we find that variability in syntactic generalization performance is substantially greater by architecture than by dataset size for the corpora tested in our experiments.Our results also reveal a dissociation between perplexity and syntactic generalization performance. Jennifer Hu 0001, Jon Gauthier, Ethan Wilcox, Roger Levy |
ACL | 1 |
| 2020 | On the Predictive Power of Neural Language Models for Human Real-Time Comprehension Behavior
Ethan Wilcox, Jon Gauthier, Jennifer Hu 0001, Roger Levy |
CogSci | 3 |
| 2019 | Separating object resonance and room reverberation in impact sounds
Jennifer Hu 0001, James Traer, Josh H. McDermott |
CogSci | 1 |
| 2018 | Generating Bilingual Pragmatic Color ReferencesabstractWill Monroe, Jennifer Hu, Andrew Jong, Christopher Potts. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Will Monroe, Jennifer Hu 0001, Andrew Jong, Christopher Potts |
NAACL-HLT | 2 |