VLDB 2026 Research / reviewers in the wild / expert
Katherine A. Keith
dblp:203/8348
· DBLP profile ↗
9ranked-venue papers
5as first author
4since 2021 · last 2026
0000-0002-8101-4572ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Probabilistic and Bayesian machine learning · 68% Information extraction and text analysis · 23% Language models and text generation · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Computational social science and digital humanities · 56% Computational finance and economics · 28% Bioinformatics and computational biology · 16% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.8 | 1 | 2024 | Proximal Causal Inference With Text Data · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
deconfounding |
0.8 | 1 | 2024 | Proximal Causal Inference With Text Data · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
proximal causal inference |
0.8 | 1 | 2024 | Proximal Causal Inference With Text Data · NeurIPS 2024 |
Computational social science and digital humanities
causal inference |
0.4 | 1 | 2020 | Text and Causal Inference: A Review of Using Text to Remove Confounding from Causal Estimates · ACL 2020 |
Bioinformatics and computational biology
confounder adjustment |
0.4 | 1 | 2020 | Text and Causal Inference: A Review of Using Text to Remove Confounding from Causal Estimates · ACL 2020 |
Computational finance and economics
financial decision-making |
0.4 | 1 | 2019 | Modeling Financial Analysts' Decision Making via the Pragmatics and Semantics of Earnings Calls · ACL (1) 2019 |
Computational finance and economics › financial data analysis
financial text analysis |
0.4 | 1 | 2019 | Modeling Financial Analysts' Decision Making via the Pragmatics and Semantics of Earnings Calls · ACL (1) 2019 |
Data mining › statistical analysis › statistical estimation
quantification |
0.3 | 1 | 2018 | Uncertainty-aware generative models for inferring document class prevalence · EMNLP 2018 |
Natural language and speech › Language models and text generation
prompting |
0.3 | 1 | 2026 | What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis
text classification |
0.3 | 1 | 2026 | What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis › event extraction
joint entity and event extraction |
0.3 | 1 | 2017 | Identifying civilians killed by police with distantly supervised entity-event extraction · EMNLP 2017 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.1 | 1 | 2018 | Uncertainty-aware generative models for inferring document class prevalence · EMNLP 2018 |
Methods — techniques the papers use, named apart from their topics
statistical inference · 2.0simulation · 2.0zero-shot models · 0.8proximal g-formula · 0.8odds ratio falsification · 0.8generative probabilistic modeling · 0.7discriminative classifier re-interpretation · 0.7bayesian inference · 0.7convolutional neural network · 0.6EM-based distant supervision · 0.6text analysis · 0.4semantic features · 0.4pragmatic features · 0.4logistic regression · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classificationabstractGenerative large language models (LLMs) are now used extensively for text classification in computational social science (CSS).In this work, we focus on the steps before and after LLM prompting: conceptualization of the categories to classify and using LLM predictions in downstream statistical inference.We argue these steps have been overlooked in much of LLM-era CSS and LLMs can tempt analysts to skip conceptualization altogether.For example, a political scientist classifying "protest" with LLMs may never be forced to craft a definition: unlike human annotators who would ask clarifying questions, an LLM can silently accept an underspecified concept to classify and return plausible-looking labels.Using simulations, we show that conceptualization failures induce downstream inferential bias that cannot be corrected solely by a more accurate LLM or posthoc bias correction methods.We conclude by reminding CSS analysts that conceptualization is still a first-order concern in the LLM-era and provide concrete advice for pursuing low-cost, unbiased, low-variance downstream estimates.4 §A provides greater details on this terminology.2044 1.Background Concept 2. Systematized Concept 3. Andrew Halterman, Katherine A. Keith |
ACL (1) | 2 |
| 2024 | Proximal Causal Inference With Text DataabstractRecent text-based causal methods attempt to mitigate confounding bias by estimating proxies of confounding variables that are partially or imperfectly measured from unstructured text data. These approaches, however, assume analysts have supervised labels of the confounders given text for a subset of instances, a constraint that is sometimes infeasible due to data privacy or annotation costs. In this work, we address settings in which an important confounding variable is completely unobserved. We propose a new causal inference method that uses two instances of pre-treatment text data, infers two proxies using two zero-shot models on the separate instances, and applies these proxies in the proximal g-formula. We prove, under certain assumptions about the instances of text and accuracy of the zero-shot predictions, that our method of inferring text-based proxies satisfies identification conditions of the proximal g-formula while other seemingly reasonable proposals do not. To address untestable assumptions associated with our method and the proximal g-formula, we further propose an odds ratio falsification heuristic that flags when to proceed with downstream effect estimation using the inferred proxies. We evaluate our method in synthetic and semi-synthetic settings---the latter with real-world clinical notes from MIMIC-III and open large language models for zero-shot prediction---and find that our method produces estimates with low bias. We believe that this text-based design of proxies allows for the use of proximal causal inference in a wider range of scenarios, particularly those for which obtaining suitable proxies from structured data is difficult. Jacob M. Chen, Rohit Bhattacharya, Katherine A. Keith |
NeurIPS | 3 |
| 2022 | Paying Attention to the Algorithm Behind the Curtain: Bringing Transparency to YouTube's Demonetization AlgorithmsabstractYouTube has long been a top-choice destination for independent video content creators to share their work. A large part of YouTube's appeal is owed to its practice of sharing advertising revenue with qualifying content creators through the YouTube Partner Program (YPP). In recent years, changes to the monetization policies and the introduction of algorithmic systems for making monetization decisions have been a source of controversy and tension between content creators and the platform. There have been numerous accusations suggesting that the underlying monetization algorithms engage in preferential treatment of larger channels and effectively censor minority voices by demonetizing their content. In this paper, we conduct a measurement of the YouTube monetization algorithms. We begin by measuring the incidence rates of different monetization decisions and the time taken to reach them. Next, we analyze the relationships between video content, channel popularity and these decisions. Finally, we explore the relationship between demonetization and a channel's view growth rate. Taken all together, our work suggests that demonetization after a video is publicly listed is not a common occurrence, the characteristics of the process are associated with channel size and (in unexplainable ways) video topic, and demonetization appears to have a harsh influence on the growth rate of smaller channels. We also highlight the challenges associated with conducting large-scale algorithm audits such as ours and make an argument for more transparency in algorithmic decision-making. Arun Dunna, Katherine A. Keith, Ethan Zuckerman, Narseo Vallina-Rodriguez, Brendan T. O'Connor 0001, Rishab Nithyanand |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2022 | Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and BeyondabstractAbstract A fundamental goal of scientific research is to learn about causal relationships. However, despite its critical role in the life and social sciences, causality has not had the same importance in Natural Language Processing (NLP), which has traditionally placed more emphasis on predictive tasks. This distinction is beginning to fade, with an emerging area of interdisciplinary research at the convergence of causal inference and language processing. Still, research on causality in NLP remains scattered across domains without unified definitions, benchmark datasets and clear articulations of the challenges and opportunities in the application of causal inference to the textual domain, with its unique properties. In this survey, we consolidate research across academic areas and situate it in the broader NLP landscape. We introduce the statistical challenge of estimating causal effects with text, encompassing settings where text is used as an outcome, treatment, or to address confounding. In addition, we explore potential uses of causal inference to improve the robustness, fairness, and interpretability of NLP models. We thus provide a unified overview of causal inference for the NLP community.1 Amir Feder, Katherine A. Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E. Roberts, Brandon M. Stewart, Victor Veitch, Diyi Yang |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | Text and Causal Inference: A Review of Using Text to Remove Confounding from Causal EstimatesabstractMany applications of computational social science aim to infer causal conclusions from nonexperimental data.Such observational data often contains confounders, variables that influence both potential causes and potential effects.Unmeasured or latent confounders can bias causal estimates, and this has motivated interest in measuring potential confounders from observed text.For example, an individual's entire history of social media posts or the content of a news article could provide a rich measurement of multiple confounders.Yet, methods and applications for this problem are scattered across different communities and evaluation practices are inconsistent.This review is the first to gather and categorize these examples and provide a guide to dataprocessing and evaluation decisions.Despite increased attention on adjusting for confounding using text, there are still many open problems, which we highlight in this paper. Katherine A. Keith, David D. Jensen, Brendan T. O'Connor 0001 |
ACL | 1 |
| 2019 | Modeling Financial Analysts' Decision Making via the Pragmatics and Semantics of Earnings CallsabstractEvery fiscal quarter, companies hold earnings calls in which company executives respond to questions from analysts.After these calls, analysts often change their price target recommendations, which are used in equity research reports to help investors make decisions.In this paper, we examine analysts' decision making behavior as it pertains to the language content of earnings calls.We identify a set of 20 pragmatic features of analysts' questions which we correlate with analysts' pre-call investor recommendations.We also analyze the degree to which semantic and pragmatic features from an earnings call complement market data in predicting analysts' post-call changes in price targets.Our results show that earnings calls are moderately predictive of analysts' decisions even though these decisions are influenced by a number of other factors including private communication with company executives and market conditions.A breakdown of model errors indicates disparate performance on calls from different market sectors. Katherine A. Keith, Amanda Stent |
ACL (1) | 1 |
| 2018 | Uncertainty-aware generative models for inferring document class prevalenceabstractPrevalence estimation is the task of inferring the relative frequency of classes of unlabeled examples in a group-for example, the proportion of a document collection with positive sentiment.Previous work has focused on aggregating and adjusting discriminative individual classifiers to obtain prevalence point estimates.But imperfect classifier accuracy ought to be reflected in uncertainty over the predicted prevalence for scientifically valid inference.In this work, we present (1) a generative probabilistic modeling approach to prevalence estimation, and (2) the construction and evaluation of prevalence confidence intervals; in particular, we demonstrate that an off-theshelf discriminative classifier can be given a generative re-interpretation, by backing out an implicit individual-level likelihood function, which can be used to conduct fast and simple group-level Bayesian inference.Empirically, we demonstrate our approach provides better confidence interval coverage than an alternative, and is dramatically more robust to shifts in the class prior between training and testing. 1 Katherine A. Keith, Brendan T. O'Connor 0001 |
EMNLP | 1 |
| 2018 | Monte Carlo Syntax Marginals for Exploring and Using Dependency ParsesabstractKatherine Keith, Su Lin Blodgett, Brendan O’Connor. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Katherine A. Keith, Su Lin Blodgett, Brendan T. O'Connor 0001 |
NAACL-HLT | 1 |
| 2017 | Identifying civilians killed by police with distantly supervised entity-event extractionabstractWe propose a new, socially-impactful task for natural language processing: from a news corpus, extract names of persons who have been killed by police.We present a newly collected police fatality corpus, which we release publicly, and present a model to solve this problem that uses EM-based distant supervision with logistic regression and convolutional neural network classifiers.Our model outperforms two off-the-shelf event extractor systems, and it can suggest candidate victim names in some cases faster than one of the major manually-collected police fatality databases. Katherine A. Keith, Abram Handler, Michael Pinkham, Cara Magliozzi, Joshua McDuffie, Brendan T. O'Connor 0001 |
EMNLP | 1 |