Justin Grimmer

dblp:165/0782 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2024
0000-0001-6642-9799ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 73% Information extraction and text analysis · 27%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 4 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
argument mining
0.812024
AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments · EMNLP 2024
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.712023
Naive regression requires weaker assumptions than factor models to adjust for multiple cause confounding · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › causal inference
deconfounding
0.712023
Naive regression requires weaker assumptions than factor models to adjust for multiple cause confounding · J. Mach. Learn. Res. 2023
Computational social science and digital humanities › causal inference
causal inference from text
0.212016
Discovery of Treatments from Text Corpora · ACL (1) 2016

Methods — techniques the papers use, named apart from their topics

topic model · 1.5human evaluation · 1.5causal inference · 1.5semiparametric regression · 1.3factor analysis · 1.3supervised indian buffet process · 0.2statistical model · 0.2experimental design · 0.2
YearPublicationVenuePosition
2024 AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments
abstract
We introduce AutoPersuade, a three-part framework for constructing persuasive messages.First, we curate a large dataset of arguments with human evaluations.Next, we develop a novel topic model to identify argument features that influence persuasiveness.Finally, we use this model to predict the effectiveness of new arguments and assess the causal impact of different components to provide explanations.We validate AutoPersuade through an experimental study on arguments for veganism, demonstrating its effectiveness with human studies and out-of-sample predictions.
Till Saenger, Musashi Hinck, Justin Grimmer, Brandon M. Stewart
EMNLP3
2023 Naive regression requires weaker assumptions than factor models to adjust for multiple cause confounding
abstract
The empirical practice of using factor models to adjust for shared, unobserved confounders, $\boldsymbol{Z}$, in observational settings with multiple treatments, $\boldsymbol{A}$, is widespread in fields including genetics, networks, medicine, and politics. Wang and Blei (2019, WB) generalize these procedures to develop the “deconfounder,” a causal inference method using factor models of $\boldsymbol{A}$ to estimate “substitute confounders,” $\widehat{\boldsymbol{Z}}$, then estimating treatment effects---regressing the outcome, $\boldsymbol{Y}$, on part of $\boldsymbol{A}$ while adjusting for $\widehat{\boldsymbol{Z}}$. WB claim the deconfounder is unbiased when (among other assumptions) there are no single-cause confounders and $\widehat{\boldsymbol{Z}}$ is “pinpointed.” We clarify pinpointing requires each confounder to affect infinitely many treatments. We prove that when the conditions hold for the deconfounder to be asymptotically unbiased, a naive semiparametric regression of $\boldsymbol{Y}$ on $\boldsymbol{A}$ which ignores confounding is also asymptotically unbiased. We provide bias formulas for finite numbers of treatments and show that different deconfounders exhibit different kinds of bias. We replicate every deconfounder analysis with available data and find that neither the naive regression nor the deconfounder consistently outperform the other. In practice, the deconfounder produces implausible estimates in WB's case study of movie earnings: estimates suggest comic author Stan Lee's cameo appearances causally contributed $15.5 billion, most of Marvel movie revenue. We conclude neither approach is a viable substitute for careful research design in real-world applications.
Justin Grimmer, Dean Knox, Brandon M. Stewart
J. Mach. Learn. Res.1
2022 Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and Beyond
abstract
Abstract A fundamental goal of scientific research is to learn about causal relationships. However, despite its critical role in the life and social sciences, causality has not had the same importance in Natural Language Processing (NLP), which has traditionally placed more emphasis on predictive tasks. This distinction is beginning to fade, with an emerging area of interdisciplinary research at the convergence of causal inference and language processing. Still, research on causality in NLP remains scattered across domains without unified definitions, benchmark datasets and clear articulations of the challenges and opportunities in the application of causal inference to the textual domain, with its unique properties. In this survey, we consolidate research across academic areas and situate it in the broader NLP landscape. We introduce the statistical challenge of estimating causal effects with text, encompassing settings where text is used as an outcome, treatment, or to address confounding. In addition, we explore potential uses of causal inference to improve the robustness, fairness, and interpretability of NLP models. We thus provide a unified overview of causal inference for the NLP community.1
Amir Feder, Katherine A. Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E. Roberts, Brandon M. Stewart, Victor Veitch, Diyi Yang
Trans. Assoc. Comput. Linguistics8
2016 Discovery of Treatments from Text Corpora
abstract
An extensive literature in computational social science examines how features of messages, advertisements, and other corpora affect individuals' decisions, but these analyses must specify the relevant features of the text before the experiment.Automated text analysis methods are able to discover features of text, but these methods cannot be used to obtain the estimates of causal effects-the quantity of interest for applied researchers.We introduce a new experimental design and statistical model to simultaneously discover treatments in a corpora and estimate causal effects for these discovered treatments.We prove the conditions to identify the treatment effects of texts and introduce the supervised Indian Buffet process to discover those treatments.Our method enables us to discover treatments in a training set using a collection of texts and individuals' responses to those texts, and then estimate the effects of these interventions in a test set of new texts and survey respondents.We apply the model to an experiment about candidate biographies, recovering intuitive features of voters' decisions and revealing a penalty for lawyers and a bonus for military service.
Christian J. Fong, Justin Grimmer
ACL (1)2
2015 TopicCheck: Interactive Alignment for Assessing Topic Model Stability
abstract
Jason Chuang, Margaret E. Roberts, Brandon M. Stewart, Rebecca Weiss, Dustin Tingley, Justin Grimmer, Jeffrey Heer. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Jason Chuang, Margaret E. Roberts, Brandon M. Stewart, Rebecca Weiss, Dustin Tingley, Justin Grimmer, Jeffrey Heer
HLT-NAACL6