David Bamman

dblp:39/5799 · DBLP profile ↗
← Back
24ranked-venue papers
7as first author
9since 2021 · last 2025
0009-0003-1171-9408ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 6 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Culture is Not Trivia: Sociocultural Theory for Cultural NLP
abstract
The field of cultural NLP has recently experienced rapid growth, driven by a pressing need to ensure that language technologies are effective and safe across a pluralistic user base.This work has largely progressed without a shared conception of culture, instead choosing to rely on a wide array of cultural proxies.However, this leads to a number of recurring limitations: coarse national boundaries fail to capture nuanced differences that lay within them, limited coverage restricts datasets to only a subset of usually highly-represented cultures, and a lack of dynamicity results in static cultural benchmarks that do not change as culture evolves.In this position paper, we argue that these methodological limitations are symptomatic of a theoretical gap.We draw on a well-developed theory of culture from sociocultural linguistics to fill this gap by 1) demonstrating in a case study how it can clarify methodological constraints and affordances, 2) offering theoretically-motivated paths forward to achieving cultural competence, and 3) arguing that localization is a more useful framing for the goals of much current work in cultural NLP.
Naitian Zhou, David Bamman, Isaac L. Bleaman
ACL (1)2
2024 AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters
abstract
Li Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren Klein, Jesse Dodge. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Li Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren F. Klein, Jesse Dodge
ACL (1)5
2024 Social Meme-ing: Measuring Linguistic Variation in Memes
abstract
Naitian Zhou, David Jurgens, David Bamman. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Naitian Zhou, David Jurgens, David Bamman
NAACL-HLT3
2023 Grounding Characters and Places in Narrative Text
abstract
Tracking characters and locations throughout a story can help improve the understanding of its plot structure.Prior research has analyzed characters and locations from text independently without grounding characters to their locations in narrative time.Here, we address this gap by proposing a new spatial relationship categorization task.The objective of the task is to assign a spatial relationship category for every character and location co-mention within a window of text, taking into consideration linguistic context, narrative tense, and temporal scope.To this end, we annotate spatial relationships in approximately 2500 book excerpts and train a model using contextual embeddings as features to predict these relationships.When applied to a set of books, this model allows us to test several hypotheses on mobility and domestic space, revealing that protagonists are more mobile than non-central characters and that women as characters tend to occupy more interior space than men.Overall, our work is the first step towards joint modeling and analysis of characters and places in narrative text.
Sandeep Soni, Amanpreet Sihra, Elizabeth F. Evans, Matthew Wilkens, David Bamman
ACL (1)5
2023 Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4
abstract
In this work, we carry out a data archaeology to infer books that are known to ChatGPT and GPT-4 using a name cloze membership inference query.We find that OpenAI models have memorized a wide collection of copyrighted materials, and that the degree of memorization is tied to the frequency with which passages of those books appear on the web.The ability of these models to memorize an unknown set of books complicates assessments of measurement validity for cultural analytics by contaminating test data; we show that models perform much better on memorized books than on nonmemorized books for downstream tasks.We argue that this supports a case for open models whose training data is known.
Kent K. Chang, Mackenzie Cramer, Sandeep Soni, David Bamman
EMNLP4
2022 Discovering Differences in the Representation of People using Contextualized Semantic Axes
abstract
A common paradigm for identifying semantic differences across social and temporal contexts is the use of static word embeddings and their distances.In particular, past work has compared embeddings against "semantic axes" that represent two opposing concepts.We extend this paradigm to BERT embeddings, and construct contextualized axes that mitigate the pitfall where antonyms have neighboring representations.We validate and demonstrate these axes on two people-centric datasets: occupations from Wikipedia, and multi-platform discussions in extremist, men's communities over fourteen years.In both studies, contextualized semantic axes can characterize differences among instances of the same word type.In the latter study, we show that references to women and the contexts around them have become more detestable over time.
Li Lucy, Divya Tadimeti, David Bamman
EMNLP3
2021 Narrative Theory for Computational Narrative Understanding
abstract
Over the past decade, the field of natural language processing has developed a wide array of computational methods for reasoning about narrative, including summarization, commonsense inference, and event detection.While this work has brought an important empirical lens for examining narrative, it is by and large divorced from the large body of theoretical work on narrative within the humanities, social and cognitive sciences.In this position paper, we introduce the dominant theoretical frameworks to the NLP community, situate current research in NLP within distinct narratological traditions, and argue that linking computational work in NLP to theory opens up a range of new empirical questions that would both help advance our understanding of narrative and open up new practical applications.
Andrew Piper, Richard Jean So, David Bamman
EMNLP (1)3
2021 Robust Laughter Detection in Noisy Environments
Jon Gillick, Wesley Deng, Kimiko Ryokai, David Bamman
Interspeech4
2021 Characterizing English Variation across Social Media Communities with BERT
abstract
Abstract Much previous work characterizing language variation across Internet social groups has focused on the types of words used by these groups. We extend this type of study by employing BERT to characterize variation in the senses of words as well, analyzing two months of English comments in 474 Reddit communities. The specificity of different sense clusters to a community, combined with the specificity of a community’s unique word types, is used to identify cases where a social group’s language deviates from the norm. We validate our metrics using user-created glossaries and draw on sociolinguistic theories to connect language variation with trends in community behavior. We find that communities with highly distinctive language are medium-sized, and their loyal and highly engaged users interact in dense networks.
Li Lucy, David Bamman
Trans. Assoc. Comput. Linguistics2
2020 Measuring Information Propagation in Literary Social Networks
abstract
We present the task of modeling information propagation in literature, in which we seek to identify pieces of information passing from character A to character B to character C, only given a description of their activity in text.We describe a new pipeline for measuring information propagation in this domain and publish a new dataset for speaker attribution, enabling the evaluation of an important component of this pipeline on a wider range of literary texts than previously studied.Using this pipeline, we analyze the dynamics of information propagation in over 5,000 works of English fiction, finding that information flows through characters that fill structural holes connecting different communities, and that characters who are women are depicted as filling this role much more frequently than characters who are men.
Matthew Sims, David Bamman
EMNLP (1)2
2020 An Annotated Dataset of Coreference in English Literature
abstract
We present in this work a new dataset of coreference annotations for works of literature in English, covering 29,103 mentions in 210,532 tokens from 100 works of fiction published between 1719 and 1922. This dataset differs from previous coreference corpora in containing documents whose average length (2,105.3 words) is four times longer than other benchmark datasets (463.7 for OntoNotes), and contains examples of difficult coreference problems common in literature. This dataset allows for an evaluation of cross-domain performance for the task of coreference resolution, and analysis into the characteristics of long-distance within-document coreference.
David Bamman, Olivia Lewke, Anya Mansoor
LREC1
2019 Literary Event Detection
abstract
In this work we present a new dataset of literary events—events that are depicted as taking place within the imagined space of a novel. While previous work has focused on event detection in the domain of contemporary news, literature poses a number of complications for existing systems, including complex narration, the depiction of a broad array of mental states, and a strong emphasis on figurative language. We outline the annotation decisions of this new dataset and compare several models for predicting events; the best performing model, a bidirectional LSTM with BERT token representations, achieves an F1 score of 73.9. We then apply this model to a corpus of novels split across two dimensions—prestige and popularity—and demonstrate that there are statistically significant differences in the distribution of events for prestige.
Matthew Sims, Jong Ho Park, David Bamman
ACL (1)3
2019 Learning to Groove with Inverse Sequence Transformations
abstract
We explore models for translating abstract musical ideas (scores, rhythms) into expressive performances using seq2seq and recurrent variational information bottleneck (VIB) models. Though seq2seq models usually require painstakingly aligned corpora, we show that it is possible to adapt an approach from the Generative Adversarial Network (GAN) literature (e.g. Pix2Pix, Vid2Vid) to sequences, creating large volumes of paired data by performing simple transformations and training generative models to plausibly invert these transformations. Music, and drumming in particular, provides a strong test case for this approach because many common transformations (quantization, removing voices) have clear semantics, and learning to invert them has real-world applications. Focusing on the case of drum set players, we create and release a new dataset for this purpose, containing over 13 hours of recordings by professional drummers aligned with fine-grained timing and dynamics information. We also explore some of the creative potential of these models, demonstrating improvements on state-of-the-art methods for Humanization (instantiating a performance from a musical score).
Jon Gillick, Adam Roberts, Jesse H. Engel, Douglas Eck, David Bamman
ICML5
2018 Capturing, Representing, and Interacting with Laughter
abstract
We investigate a speculative future in which we celebrate happiness by capturing laughter and representing it in tangible forms. We explored technologies for capturing naturally occurring laughter as well as various physical representations of it. For several weeks, our participants collected audio samples of everyday conversations with their loved ones. We processed those samples through a machine learning algorithm and shared the resulting tangible representations (e.g., physical containers and edible displays) with our participants. In collecting, listening to, interacting with, and sharing their laughter with loved ones, participants described both joy in preserving and interacting with laughter and tension in collecting it. This study revealed that the tangibility of laughter representations matters, especially its symbolism and material quality. We discuss design implications of giving permanent forms to laughter and consider the sound of laughter as a part of our personal past that we might seek to preserve and reflect upon.
Kimiko Ryokai, Elena Durán, Noura Howell, Jon Gillick, David Bamman
CHI5
2018 Please Clap: Modeling Applause in Campaign Speeches
abstract
This work examines the rhetorical techniques that speakers employ during political campaigns.We introduce a new corpus of speeches from campaign events in the months leading up to the 2016 U.S. presidential election and develop new models for predicting moments of audience applause.In contrast to existing datasets, we tackle the challenge of working with transcripts that derive from uncorrected closed captioning, using associated audio recordings to automatically extract and align labels for instances of audience applause.In prediction experiments, we find that lexical features carry the most information, but that a variety of features are predictive, including prosody, long-term contextual dependencies, and theoretically motivated features designed to capture rhetorical techniques.
Jon Gillick, David Bamman
NAACL-HLT2
2017 The Labeled Segmentation of Printed Books
abstract
We introduce the task of book structure labeling: segmenting and assigning a fixed category (such as TABLE OF CONTENTS, PREFACE, INDEX) to the document structure of printed books.We manually annotate the page-level structural categories for a large dataset totaling 294,816 pages in 1,055 books evenly sampled from 1750-1922, and present empirical results comparing the performance of several classes of models.The best-performing model, a bidirectional LSTM with rich features, achieves an overall accuracy of 95.8 and a class-balanced macro F-score of 71.4.
Lara McConnaughey, Jennifer Dai, David Bamman
EMNLP3
2017 Adversarial Training for Relation Extraction
abstract
Adversarial training is a mean of regularizing classification algorithms by generating adversarial noise to the training data.We apply adversarial training in relation extraction within the multi-instance multi-label learning framework.We evaluate various neural network architectures on two different datasets.Experimental results demonstrate that adversarial training is generally effective for both CNN and RNN models and significantly improves the precision of predicted relations.
Yi Wu 0013, David Bamman, Stuart Russell 0001
EMNLP2
2016 Beyond Canonical Texts: A Computational Analysis of Fanfiction
abstract
While much computational work on fiction has focused on works in the literary canon, user-created fanfiction presents a unique opportunity to study an ecosystem of literary production and consumption, embodying qualities both of large-scale literary data (55 billion tokens) and also a social network (with over 2 million users).We present several empirical analyses of this data in order to illustrate the range of affordances it presents to research in NLP, computational social science and the digital humanities.We find that fanfiction deprioritizes main protagonists in comparison to canonical texts, has a statistically significant difference in attention allocated to female characters, and offers a framework for developing models of reader reactions to stories.
Smitha Milli, David Bamman
EMNLP2
2015 Open Extraction of Fine-Grained Political Statements
abstract
Text data has recently been used as evidence in estimating the political ideologies of individuals, including political elites and social media users.While inferences about people are often the intrinsic quantity of interest, we draw inspiration from open information extraction to identify a new task: inferring the political import of propositions like OBAMA IS A SOCIAL-IST.We present several models that exploit the structure that exists between people and the assertions they make to learn latent positions of people and propositions at the same time, and we evaluate them on a novel dataset of propositions judged on a political spectrum.
David Bamman, Noah A. Smith
EMNLP1
2015 Contextualized Sarcasm Detection on Twitter
David Bamman, Noah A. Smith
ICWSM1
2014 A Bayesian Mixed Effects Model of Literary Character
abstract
We consider the problem of automatically inferring latent character types in a collection of 15,099 English novels published between 1700 and 1899.Unlike prior work in which character types are assumed responsible for probabilistically generating all text associated with a character, we introduce a model that employs multiple effects to account for the influence of extra-linguistic information (such as author).In an empirical evaluation, we find that this method leads to improved agreement with the preregistered judgments of a literary scholar, complementing the results of alternative models.
David Bamman, Ted Underwood, Noah A. Smith
ACL (1)1
2014 Unsupervised Discovery of Biographical Structure from Text
abstract
We present a method for discovering abstract event classes in biographies, based on a probabilistic latent-variable model. Taking as input timestamped text, we exploit latent correlations among events to learn a set of event classes (such as Born, Graduates High School, and Becomes Citizen), along with the typical times in a person’s life when those events occur. In a quantitative evaluation at the task of predicting a person’s age for a given event, we find that our generative model outperforms a strong linear regression baseline, along with simpler variants of the model that ablate some features. The abstract event classes that we learn allow us to perform a large-scale analysis of 242,970 Wikipedia biographies. Though it is known that women are greatly underrepresented on Wikipedia—not only as editors (Wikipedia, 2011) but also as subjects of articles (Reagle and Rhue, 2011)—we find that there is a bias in their characterization as well, with biographies of women containing significantly more emphasis on events of marriage and divorce than biographies of men.
David Bamman, Noah A. Smith
Trans. Assoc. Comput. Linguistics1
2013 Learning Latent Personas of Film Characters
David Bamman, Brendan T. O'Connor 0001, Noah A. Smith
ACL (1)1
2008 The Annotation Guidelines of the Latin Dependency Treebank and Index Thomisticus Treebank: the Treatment of some specific Syntactic Constructions in Latin
David Bamman, Marco Passarotti, Roberto Busa, Gregory R. Crane
LREC1