Janet B. Pierrehumbert

dblp:60/5814 · DBLP profile ↗
← Back
32ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0002-5989-3574ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks
abstract
Fangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann, Adrian de Wynter, Xun Wang, Si-Qing Chen, Michael J. Wooldridge, Janet B. Pierrehumbert, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Fangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann, Adrian de Wynter, Xun Wang 0012, Michael J. Wooldridge, Janet B. Pierrehumbert, Furu Wei
ACL (1)9
2025 ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts
abstract
Scientific fact-checking has largely focused on textual and tabular sources, neglecting scientific charts-a primary medium for conveying quantitative evidence and supporting statistical reasoning in research communication.We introduce CLIMATEVIZ, the first largescale benchmark for scientific fact-checking grounded in real-world, expert-curated scientific charts.CLIMATEVIZ comprises 49,862 claims paired with 2,896 visualizations, each labeled as support, refute, or not enough information.To enable interpretable verification, each instance includes structured knowledge graph explanations that capture statistical patterns, temporal trends, spatial comparisons, and causal relations.We conduct a comprehensive evaluation of state-of-the-art multimodal large language models, including proprietary and open-source systems, under zero-shot and few-shot settings.Our results show that current models struggle to perform fact-checking when statistical reasoning over charts is required: even the best-performing systems, such as Gemini 2.5 and InternVL 2.5, achieve only 76.2-77.8%accuracy in label-only output settings, which is far below human performance (89.3% and 92.7%).While few-shot prompting yields limited improvements, explanationaugmented outputs significantly enhance performance in some closed-source models, notably o3 and Gemini 2.5.We released our dataset and code alongside the paper. 1(c) Subgraph of Relevant Facts Caption: Cumulative mass loss of the Greenland Ice Sheet from 1972 to 2022, showing accelerating ice loss and corresponding sea level rise.
Ruiran Su, Jiasheng Si, Zhijiang Guo, Janet B. Pierrehumbert
EMNLP4
2024 Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics
abstract
Scalar adjectives pertain to various domain scales and vary in intensity within each scale (e.g. certain is more intense than likely on the likelihood scale). Scalar implicatures arise from the consideration of alternative statements which could have been made. They can be triggered by scalar adjectives and require listeners to reason pragmatically about them. Some scalar adjectives are more likely to trigger scalar implicatures than others. This phenomenon is referred to as scalar diversity. In this study, we probe different families of Large Language Models such as GPT-4 for their knowledge of the lexical semantics of scalar adjectives and one specific aspect of their pragmatics, namely scalar diversity. We find that they encode rich lexical-semantic information about scalar adjectives. However, the rich lexical-semantic knowledge does not entail a good understanding of scalar diversity. We also compare current models of different sizes and complexities and find that larger models are not always better. Finally, we explain our probing results by leveraging linguistic intuitions and model training objectives.
Fangru Lin, Daniel Altshuler, Janet B. Pierrehumbert
LREC/COLING3
2024 STEntConv: Predicting Disagreement between Reddit Users with Stance Detection and a Signed Graph Convolutional Network
abstract
The rise of social media platforms has led to an increase in polarised online discussions, especially on political and socio-cultural topics such as elections and climate change. We propose a simple and entirely novel unsupervised method to better predict whether the authors of two posts agree or disagree, leveraging user stances about named entities obtained from their posts. We present STEntConv, a model which builds a graph of users and named entities weighted by stance and trains a Signed Graph Convolutional Network (SGCN) to detect disagreement between comment and reply posts. We run experiments and ablation studies and show that including this information improves disagreement detection performance on a dataset of Reddit posts for a range of controversial subreddit topics, without the need for platform-specific features or user history
Isabelle Lorge, Li Zhang 0131, Xiaowen Dong 0001, Janet B. Pierrehumbert
LREC/COLING4
2024 Graph-enhanced Large Language Models in Asynchronous Plan Reasoning
abstract
Planning is a fundamental property of human intelligence. Reasoning about asynchronous plans is challenging since it requires sequential and parallel planning to optimize time costs. Can large language models (LLMs) succeed at this task? Here, we present the first large-scale study investigating this question. We find that a representative set of closed and open-source LLMs, including GPT-4 and LLaMA-2, behave poorly when not supplied with illustrations about the task-solving process in our benchmark AsyncHow. We propose a novel technique called *Plan Like a Graph* (PLaG) that combines graphs with natural language prompts and achieves state-of-the-art results. We show that although PLaG can boost model performance, LLMs still suffer from drastic degradation when task complexity increases, highlighting the limits of utilizing LLMs for simulating digital devices. We see our study as an exciting step towards using LLMs as efficient autonomous agents. Our code and data are available at https://github.com/fangru-lin/graph-llm-asynchow-plan.
Fangru Lin, Emanuele La Malfa, Valentin Hofmann, Elle Michelle Yang, Anthony G. Cohn 0001, Janet B. Pierrehumbert
ICML6
2024 Geographic Adaptation of Pretrained Language Models
abstract
Abstract While pretrained language models (PLMs) have been shown to possess a plethora of linguistic knowledge, the existing body of research has largely neglected extralinguistic knowledge, which is generally difficult to obtain by pretraining on text alone. Here, we contribute to closing this gap by examining geolinguistic knowledge, i.e., knowledge about geographic variation in language. We introduce geoadaptation, an intermediate training step that couples language modeling with geolocation prediction in a multi-task learning setup. We geoadapt four PLMs, covering language groups from three geographic areas, and evaluate them on five different tasks: fine-tuned (i.e., supervised) geolocation prediction, zero-shot (i.e., unsupervised) geolocation prediction, fine-tuned language identification, zero-shot language identification, and zero-shot prediction of dialect features. Geoadaptation is very successful at injecting geolinguistic knowledge into the PLMs: The geoadapted PLMs consistently outperform PLMs adapted using only language modeling (by especially wide margins on zero-shot prediction tasks), and we obtain new state-of-the-art results on two benchmarks for geolocation prediction and language identification. Furthermore, we show that the effectiveness of geoadaptation stems from its ability to geographically retrofit the representation space of the PLMs.
Valentin Hofmann, Goran Glavas, Nikola Ljubesic, Janet B. Pierrehumbert, Hinrich Schütze
Trans. Assoc. Comput. Linguistics4
2022 Unsupervised Detection of Contextualized Embedding Bias with Application to Ideology
abstract
We propose a fully unsupervised method to detect bias in contextualized embeddings. The method leverages the assortative information latently encoded by social networks and combines orthogonality regularization, structured sparsity learning, and graph neural networks to find the embedding subspace capturing this information. As a concrete example, we focus on the phenomenon of ideological bias: we introduce the concept of an ideological subspace, show how it can be found by applying our method to online discussion forums, and present techniques to probe it. Our experiments suggest that the ideological subspace encodes abstract evaluative semantics and reflects changes in the political left-right spectrum during the presidency of Donald Trump.
Valentin Hofmann, Janet B. Pierrehumbert, Hinrich Schütze
ICML2
2022 The Reddit Politosphere: A Large-Scale Text and Network Resource of Online Political Discourse
Valentin Hofmann, Hinrich Schütze, Janet B. Pierrehumbert
ICWSM3
2022 Forecasting COVID-19 Caseloads Using Unsupervised Embedding Clusters of Social Media Posts
abstract
Felix Drinkall, Stefan Zohren, Janet Pierrehumbert. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Felix Drinkall, Stefan Zohren, Janet B. Pierrehumbert
NAACL-HLT3
2022 Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks
abstract
Paul Rottger, Bertie Vidgen, Dirk Hovy, Janet Pierrehumbert. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Paul Röttger, Bertie Vidgen, Dirk Hovy, Janet B. Pierrehumbert
NAACL-HLT4
2021 Superbizarre Is Not Superb: Derivational Morphology Improves BERT's Interpretation of Complex Words
abstract
Valentin Hofmann, Janet Pierrehumbert, Hinrich Schütze. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Valentin Hofmann, Janet B. Pierrehumbert, Hinrich Schütze
ACL/IJCNLP (1)2
2021 Dynamic Contextualized Word Embeddings
abstract
Valentin Hofmann, Janet Pierrehumbert, Hinrich Schütze. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Valentin Hofmann, Janet B. Pierrehumbert, Hinrich Schütze
ACL/IJCNLP (1)2
2021 HateCheck: Functional Tests for Hate Speech Detection Models
abstract
Detecting online hate is a difficult task that even state-of-the-art models struggle with. Typically, hate speech detection models are evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1 score. However, this approach makes it difficult to identify specific model weak points. It also risks overestimating generalisable model performance due to increasingly well-evidenced systematic gaps and biases in hate speech datasets. To enable more targeted diagnostic insights, we introduce HateCheck, a suite of functional tests for hate speech detection models. We specify 29 model functionalities motivated by a review of previous research and a series of interviews with civil society stakeholders. We craft test cases for each functionality and validate their quality through a structured annotation process. To illustrate HateCheck's utility, we test near-state-of-the-art transformer models as well as two popular commercial models, revealing critical model weaknesses.
Paul Röttger, Bertie Vidgen, Dong Nguyen 0002, Zeerak Talat, Helen Z. Margetts, Janet B. Pierrehumbert
ACL/IJCNLP (1)6
2020 Predicting the Growth of Morphological Families from Social and Linguistic Factors
abstract
We present the first study that examines the evolution of morphological families, i.e., sets of morphologically related words such as "trump", "antitrumpism", and "detrumpify", in social media.We introduce the novel task of Morphological Family Expansion Prediction (MFEP) as predicting the increase in the size of a morphological family.We create a ten-year Reddit corpus as a benchmark for MFEP and evaluate a number of baselines on this benchmark.Our experiments demonstrate very good performance on MFEP.
Valentin Hofmann, Janet B. Pierrehumbert, Hinrich Schütze
ACL2
2020 A Graph Auto-encoder Model of Derivational Morphology
abstract
There has been little work on modeling the morphological well-formedness (MWF) of derivatives, a problem judged to be complex and difficult in linguistics (Bauer, 2019).We present a graph auto-encoder that learns embeddings capturing information about the compatibility of affixes and stems in derivation.The auto-encoder models MWF in English surprisingly well by combining syntactic and semantic information with associative information from the mental lexicon.
Valentin Hofmann, Hinrich Schütze, Janet B. Pierrehumbert
ACL3
2020 DagoBERT: Generating Derivational Morphology with a Pretrained Language Model
abstract
Can pretrained language models (PLMs) generate derivationally complex words?We present the first study investigating this question, taking BERT as the example PLM.We examine BERT's derivational capabilities in different settings, ranging from using the unmodified pretrained model to full finetuning.Our best model, DagoBERT (Derivationally and generatively optimized BERT), clearly outperforms the previous state of the art in derivation generation (DG).Furthermore, our experiments show that the input segmentation crucially impacts BERT's derivational knowledge, suggesting that the performance of PLMs could be further improved if a morphologically informed vocabulary of units were used.
Valentin Hofmann, Janet B. Pierrehumbert, Hinrich Schütze
EMNLP (1)2
2020 The cognitive status of simple and complex models
Janet B. Pierrehumbert
INTERSPEECH1
2017 Prior Expectations in Linguistic Learning: A Stochastic Model of Individual Differences
R. Alexander Schumacher, Janet B. Pierrehumbert
CogSci2
2016 Using Pronunciation-Based Morphological Subword Units to Improve OOV Handling in Keyword Search
abstract
Out-of-vocabulary (OOV) keywords present a challenge for keyword search (KWS) systems especially in the low-resource setting. Previous research has centered around approaches that use a variety of subword units to recover OOV words. This paper systematically investigates morphology-based subword modeling approaches on seven low-resource languages. We show that using morphological subword units (morphs) in speech recognition decoding is substantially better than expanding word-decoded lattices into subword units including phones, syllables and morphs. As alternatives to grapheme-based morphs, we apply unsupervised morphology learning to sequences of phonemes, graphones, and syllables. Using one of these phone-based morphs is almost always better than using the grapheme-based morphs, but the particular choice varies with the language. By combining the different methods, a substantial gain is obtained over the best single case for all languages, especially for OOV performance.
Yanzhang He, Peter Baumann 0003, Hao Fang 0002, Brian Hutchinson, Aaron Jaech, Mari Ostendorf, Eric Fosler-Lussier, Janet B. Pierrehumbert
IEEE ACM Trans. Audio Speech Lang. Process.8
2015 Exponential Language Modeling Using Morphological Features and Multi-Task Learning
abstract
For languages with fast vocabulary growth and limited resources, data sparsity leads to challenges in training a language model. One strategy for addressing this problem is to leverage morphological structure as features in the model. This paper explores different uses of unsupervised morphological features in both the history and prediction space for three word-based exponential models (maximum entropy, logbilinear, and recurrent neural net (RNN)). Multi-task training is introduced as a regularizing mechanism to improve performance in the continuous-space approaches. The models are compared to non-parametric baselines. From using the RNN with morphological features and multi-task learning, experiments with conversational speech from four languages show we can obtain consistent gains of 7-11% in perplexity reduction in a limited-resource scenario (10 hrs speech), and 12-18% when the training size is increased ( 80 hrs ). Results are mixed for all other approaches, compared to a modified Kneser-Ney baseline, but morphology is useful in continuous-space models compared to their word-only baseline. Multi-task learning improves both continuous-space models.
Hao Fang 0002, Mari Ostendorf, Peter Baumann 0003, Janet B. Pierrehumbert
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Real Words, Possible Words, and New Words
Janet B. Pierrehumbert
CogSci1
2014 Reconciling Inconsistency in Encoded Morphological Distinctions in an Artificial Language
R. Alexander Schumacher, Janet B. Pierrehumbert, Patrick Lashell
CogSci2
2014 Subword-based modeling for handling OOV words inkeyword spotting
abstract
This work compares ASR decoding at different subword levels crossed with alternative keyword search strategies to handle the OOV issue for keyword spotting in the low-resource setting. We show that a morpheme-based subword modeling approach is effective in recovering OOV keywords within a Turkish low-resource keyword spotting task, where mixed word and morpheme decoding approach outperforms the traditional subword-based search from word-decoded lattices that are broken down to subword lattices. Furthermore, unsupervised learning of morphology works almost as well as a rule-based system designed for the language despite the low-resource condition. A staged keyword search strategy benefits from both methods of morphological analysis.
Yanzhang He, Brian Hutchinson, Peter Baumann 0003, Mari Ostendorf, Eric Fosler-Lussier, Janet B. Pierrehumbert
ICASSP6
2014 Using Resource-Rich Languages to Improve Morphological Analysis of Under-Resourced Languages
Peter Baumann 0003, Janet B. Pierrehumbert
LREC2
2010 Audio-visual anticipatory coarticulation modeling by human and machine
abstract
The phenomenon of anticipatory coarticulation provides a ba-sis for the observed asynchrony between the acoustic and vi-sual onsets of phones in certain linguistic contexts. This type of asynchrony is typically not explicitly modeled in audio-visual speech models. In this work, we study within-word audio-visual asynchrony using manual labels of words in which theory suggests that audio-visual asynchrony should occur, and show that these hand labels confirm the theory. We then introduce a new statistical model of audio-visual speech, the asynchrony-dependent transition (ADT) model. This model allows asyn-chrony between audio and video states within word boundaries, where the audio and video state transitions depend not only on the state of that modality, but also on the instantaneous asyn-chrony. The ADT model outperforms a baseline synchronous model in mimicking the hand labels in a forced alignment task, and its behavior as parameters are changed conforms to our ex-pectations about anticipatory coarticulation. The same model could be used for speech recognition, although here we consider it only for the task of forced alignment for linguistic analysis. Index Terms: audio-visual speech recognition, audio-visual asynchrony, anticipatory coarticulation, dynamic Bayesian net-works 1.
Louis H. Terry, Karen Livescu, Janet B. Pierrehumbert, Aggelos K. Katsaggelos
INTERSPEECH3
2007 Much ado about nothing: A social network model of Russian paradigmatic gaps
Robert Daland, Andrea D. Sims, Janet B. Pierrehumbert
ACL3
1992 TOBI: a standard for labeling English prosody
Kim E. A. Silverman, Mary E. Beckman, John F. Pitrelli, Mari Ostendorf, Colin W. Wightman, Patti Price, Janet B. Pierrehumbert, Julia Hirschberg
ICSLP7
1987 Intonation and the Intentional Structure of Discourse
Julia Hirschberg, Diane J. Litman, Janet B. Pierrehumbert, G. Ward
IJCAI3
1986 Japanese prosodic phrasing and intonation Synthesis
abstract
A computer program for synthesizing Japanese fundamental frequency contours implements our theory of Japanese intonation. This theory provides a complete qualitative description of the known characteristics of Japanese intonation, as well as a quantitative model of tone-scaling and timing precise enough to translate straightforwardly into a computational algorithm. An important aspect of the description is that various features of the intonation pattern are designated to be phonological properties of different types of phrasal units in a hierarchical organization. This phrasal organization is known to play an important role in parsing speech. Our research shows it also to be one reflex of intonational prominence, and hence of focus and other discourse structures. The qualitative features of each phrasal level and their implementation in the synthesis program are described.
Mary E. Beckman, Janet B. Pierrehumbert
ACL2
1986 The intonational Structuring of Discourse
abstract
We propose a mapping between prosodic phenomena and semantico-pragmatic effects based upon the hypothesis that intonation conveys information about the intentional as well as the attentional structure of discourse. In particular, we discuss how variations in pitch range and choice of accent and tune can help to convey such information as: discourse segmentation and topic structure, appropriate choice of referent, the distinction between 'given' and 'new' information, conceptual contrast or parallelism between mentioned items, and subordination relationships between propositions salient in the discourse. Our goals for this research are practical as well as theoretical. In particular, we are investigating the problem of intonational assignment in synthetic speech.
Julia Hirschberg, Janet B. Pierrehumbert
ACL2
1984 Synthesis by rule of english intonation patterns
abstract
This papet reports work on synthesizing English F0 contours. One motivation for this work is to improve the naturalness and liveliness of the prosody in speech synthesis systems. However, our main goal is to develop a theory of the dimensions of variation controlling intonation, and of their interaction.
Janet B. Pierrehumbert, Mark Y. Liberman
ICASSP2
1983 Automatic Recognition of Intonation Patterns
abstract
Americanae nace como un proyecto conjunto que surge dentro de la Red Europea de Información y Documentación sobre América Latina (REDIAL), y que ha afrontado la Biblioteca de la Agencia Española de Cooperación Internacional para el Desarrollo (AECID). Esta nueva biblioteca virtual hace más accesibles los libros digitales de tema americanista a los investigadores y usuarios interesados de cualquier parte del mundo.
Janet B. Pierrehumbert
ACL1