VLDB 2026 Research / reviewers in the wild / expert
Kawin Ethayarajh
dblp:198/6540
· DBLP profile ↗
12ranked-venue papers
9as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 9 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Language models and text generation · 38% Representation and self-supervised learning · 29% Trustworthy machine learning · 21% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › word representation
word embedding |
1.1 | 3 | 2019 | Rotate King to get Queen: Word Relationships as Orthogonal Transformations in Embedding Space · EMNLP/IJCNLP (1) 2019 Towards Understanding Linear Word Analogies · ACL (1) 2019 Understanding Undesirable Word Embedding Associations · ACL (1) 2019 |
Natural language and speech › Language models and text generation
text generation evaluation |
1.0 | 2 | 2022 | The Authenticity Gap in Human Evaluation · EMNLP 2022 Utility is in the Eye of the User: A Critique of NLP Leaderboards · EMNLP (1) 2020 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 2 | 2020 | Is Your Classifier Actually Biased? Measuring Fairness under Uncertainty with Bernstein Bounds · ACL 2020 Understanding Undesirable Word Embedding Associations · ACL (1) 2019 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Model Alignment as Prospect Theoretic Optimization · ICML 2024 |
Natural language and speech › Language models and text generation
preference optimization |
0.8 | 1 | 2024 | Model Alignment as Prospect Theoretic Optimization · ICML 2024 |
Machine learning › Representation and self-supervised learning
probing |
0.5 | 1 | 2021 | Conditional probing: measuring usable information beyond a baseline · EMNLP (1) 2021 |
Performance modeling and evaluation
benchmarking |
0.5 | 1 | 2021 | Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking · NeurIPS 2021 |
Natural language and speech › Language models and text generation › large language model evaluation › large language model benchmarking
leaderboard evaluation |
0.4 | 1 | 2020 | Utility is in the Eye of the User: A Critique of NLP Leaderboards · EMNLP (1) 2020 |
Machine learning › Trustworthy machine learning › fairness › fairness in NLP
bias in word embeddings |
0.4 | 1 | 2019 | Understanding Undesirable Word Embedding Associations · ACL (1) 2019 |
Machine learning › Representation and self-supervised learning › word representation
contextualized word representation |
0.4 | 1 | 2019 | How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings · EMNLP/IJCNLP (1) 2019 |
Machine learning › Trustworthy machine learning
debiasing |
0.4 | 1 | 2019 | Understanding Undesirable Word Embedding Associations · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis › lexical semantics
word analogy |
0.4 | 1 | 2019 | Rotate King to get Queen: Word Relationships as Orthogonal Transformations in Embedding Space · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation › text generation
story generation |
0.2 | 1 | 2022 | The Authenticity Gap in Human Evaluation · EMNLP 2022 |
Machine learning › Representation and self-supervised learning › representation analysis
neural representation analysis |
0.1 | 1 | 2021 | Conditional probing: measuring usable information beyond a baseline · EMNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
prospect theory · 0.8utility theory · 0.6v-information · 0.5utility-based aggregation · 0.5probing · 0.5microeconomic utility theory · 0.4confidence intervals · 0.4bernstein bounds · 0.4subspace projection · 0.4matrix factorization · 0.4WEAT · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Anchor Points: Benchmarking Models with Much Fewer ExamplesabstractModern language models often exhibit powerful but brittle behavior, leading to the development of larger and more diverse benchmarks to reliably assess their behavior.Here, we suggest that model performance can be benchmarked and elucidated with much smaller evaluation sets.We first show that in six popular language classification benchmarks, model confidence in the correct class on many pairs of points is strongly correlated across models.We build upon this phenomenon to propose Anchor Point Selection, a technique to select small subsets of datasets that capture model behavior across the entire dataset.Anchor points reliably rank models: across 87 diverse language model-prompt pairs, evaluating models using 1-30 anchor points outperforms uniform sampling and other baselines at accurately ranking models.Moreover, just a dozen anchor points can be used to estimate model per-class predictions on all other points in a dataset with low error, sufficient for gauging where the model is likely to fail.Lastly, we present Anchor Point Maps for visualizing these insights and facilitating comparisons of the performance of different models on various regions within the dataset distribution. Rajan Vivek, Kawin Ethayarajh, Diyi Yang, Douwe Kiela |
EACL (1) | 2 |
| 2024 | Model Alignment as Prospect Theoretic OptimizationabstractKahneman & Tversky’s $\textit{prospect theory}$ tells us that humans perceive random variables in a biased but well-defined manner (1992); for example, humans are famously loss-averse. We show that objectives for aligning LLMs with human feedback implicitly incorporate many of these biases—the success of these objectives (e.g., DPO) over cross-entropy minimization can partly be ascribed to them belonging to a family of loss functions that we call $\textit{human-aware losses}$ (HALOs). However, the utility functions these methods attribute to humans still differ from those in the prospect theory literature. Using a Kahneman-Tversky model of human utility, we propose a HALO that directly maximizes the utility of generations instead of maximizing the log-likelihood of preferences, as current methods do. We call this approach KTO, and it matches or exceeds the performance of preference-based methods at scales from 1B to 30B, despite only learning from a binary signal of whether an output is desirable. More broadly, our work suggests that there is no one HALO that is universally superior; the best loss depends on the inductive biases most appropriate for a given setting, an oft-overlooked consideration. Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Daniel Jurafsky, Douwe Kiela |
ICML | 1 |
| 2022 | The Authenticity Gap in Human EvaluationabstractHuman ratings are the gold standard in NLG evaluation.The standard protocol is to collect ratings of generated text, average across annotators, and rank NLG systems by their average scores.However, little consideration has been given as to whether this approach faithfully captures human preferences.Analyzing this standard protocol through the lens of utility theory in economics, we identify the implicit assumptions it makes about annotators.These assumptions are often violated in practice, in which case annotator ratings cease to reflect their preferences.The most egregious violations come from using Likert scales, which provably reverse the direction of the true preference in certain cases.We suggest improvements to the standard protocol to make it more theoretically sound, but even in its improved form, it cannot be used to evaluate open-ended tasks like story generation.For the latter, we propose a new human evaluation protocol called systemlevel probabilistic assessment (SPA).When human evaluation of stories is done with SPA, we can recover the ordering of GPT-3 models by size, with statistically significant results.However, when human evaluation is done with the standard protocol, less than half of the expected preferences can be recovered (e.g., there is no significant difference between curie and davinci, despite using a highly powered test). Kawin Ethayarajh, Daniel Jurafsky |
EMNLP | 1 |
| 2022 | Understanding Dataset Difficulty with V-Usable Information
Kawin Ethayarajh, Yejin Choi 0001, Swabha Swayamdipta |
ICML | 1 |
| 2021 | Conditional probing: measuring usable information beyond a baselineabstractProbing experiments investigate the extent to which neural representations make properties-like part-of-speech-predictable.One suggests that a representation encodes a property if probing that representation produces higher accuracy than probing a baseline representation like non-contextual word embeddings.Instead of using baselines as a point of comparison, we're interested in measuring information that is contained in the representation but not in the baseline.For example, current methods can detect when a representation is more useful than the word identity (a baseline) for predicting part-ofspeech; however, they cannot detect when the representation is predictive of just the aspects of part-of-speech not explainable by the word identity.In this work, we extend a theory of usable information called V-information and propose conditional probing, which explicitly conditions on the information in the baseline.In a case study, we find that after conditioning on non-contextual word embeddings, properties like part-of-speech are accessible at deeper layers of a network than previously thought. John Hewitt, Kawin Ethayarajh, Percy Liang, Christopher D. Manning |
EMNLP (1) | 2 |
| 2021 | Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation BenchmarkingabstractWe introduce Dynaboard, an evaluation-as-a-service framework for hosting benchmarks and conducting holistic model comparison, integrated with the Dynabench platform. Our platform evaluates NLP models directly instead of relying on self-reported metrics or predictions on a single dataset. Under this paradigm, models are submitted to be evaluated in the cloud, circumventing the issues of reproducibility, accessibility, and backwards compatibility that often hinder benchmarking in NLP. This allows users to interact with uploaded models in real time to assess their quality, and permits the collection of additional metrics such as memory use, throughput, and robustness, which -- despite their importance to practitioners -- have traditionally been absent from leaderboards. On each task, models are ranked according to the Dynascore, a novel utility-based aggregation of these statistics, which users can customize to better reflect their preferences, placing more/less weight on a particular axis of evaluation or dataset. As state-of-the-art NLP models push the limits of traditional benchmarks, Dynaboard offers a standardized solution for a more diverse and comprehensive evaluation of model quality. Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, Douwe Kiela |
NeurIPS | 2 |
| 2020 | Is Your Classifier Actually Biased? Measuring Fairness under Uncertainty with Bernstein BoundsabstractMost NLP datasets are not annotated with protected attributes such as gender, making it difficult to measure classification bias using standard measures of fairness (e.g., equal opportunity).However, manually annotating a large dataset with a protected attribute is slow and expensive.Instead of annotating all the examples, can we annotate a subset of them and use that sample to estimate the bias?While it is possible to do so, the smaller this annotated sample is, the less certain we are that the estimate is close to the true bias.In this work, we propose using Bernstein bounds to represent this uncertainty about the bias estimate as a confidence interval.We provide empirical evidence that a 95% confidence interval derived this way consistently bounds the true bias.In quantifying this uncertainty, our method, which we call Bernstein-bounded unfairness, helps prevent classifiers from being deemed biased or unbiased when there is insufficient evidence to make either claim.Our findings suggest that the datasets currently used to measure specific biases are too small to conclusively identify bias except in the most egregious cases.For example, consider a coreference resolution system that is 5% more accurate on gender-stereotypical sentences -to claim it is biased with 95% confidence, we need a bias-specific dataset that is 3.8 times larger than WinoBias, the largest available. Kawin Ethayarajh |
ACL | 1 |
| 2020 | Utility is in the Eye of the User: A Critique of NLP LeaderboardsabstractBenchmarks such as GLUE have helped drive advances in NLP by incentivizing the creation of more accurate models.While this leaderboard paradigm has been remarkably successful, a historical focus on performance-based evaluation has been at the expense of other qualities that the NLP community values in models, such as compactness, fairness, and energy efficiency.In this opinion paper, we study the divergence between what is incentivized by leaderboards and what is useful in practice through the lens of microeconomic theory.We frame both the leaderboard and NLP practitioners as consumers and the benefit they get from a model as its utility to them.With this framing, we formalize how leaderboards -in their current form -can be poor proxies for the NLP community at large.For example, a highly inefficient model would provide less utility to practitioners but not to a leaderboard, since it is a cost that only the former must bear.To allow practitioners to better estimate a model's utility to them, we advocate for more transparency on leaderboards, such as the reporting of statistics that are of practical concern (e.g., model size, energy efficiency, and inference latency). Kawin Ethayarajh, Daniel Jurafsky |
EMNLP (1) | 1 |
| 2019 | Understanding Undesirable Word Embedding AssociationsabstractWord embeddings are often criticized for capturing undesirable word associations such as gender stereotypes.However, methods for measuring and removing such biases remain poorly understood.We show that for any embedding model that implicitly does matrix factorization, debiasing vectors post hoc using subspace projection (Bolukbasi et al., 2016) is, under certain conditions, equivalent to training on an unbiased corpus.We also prove that WEAT, the most common association test for word embeddings, systematically overestimates bias.Given that the subspace projection method is provably effective, we use it to derive a new measure of association called the relational inner product association (RIPA).Experiments with RIPA reveal that, on average, skipgram with negative sampling (SGNS) does not make most words any more gendered than they are in the training corpus.However, for gender-stereotyped words, SGNS actually amplifies the gender association in the corpus. Kawin Ethayarajh, David Duvenaud, Graeme Hirst |
ACL (1) | 1 |
| 2019 | Towards Understanding Linear Word AnalogiesabstractA surprising property of word vectors is that word analogies can often be solved with vector arithmetic.However, it is unclear why arithmetic operators correspond to non-linear embedding models such as skip-gram with negative sampling (SGNS).We provide a formal explanation of this phenomenon without making the strong assumptions that past theories have made about the vector space and word distribution.Our theory has several implications.Past work has conjectured that linear substructures exist in vector spaces because relations can be represented as ratios; we prove that this holds for SGNS.We provide novel justification for the addition of SGNS word vectors by showing that it automatically downweights the more frequent word, as weighting schemes do ad hoc.Lastly, we offer an information theoretic interpretation of Euclidean distance in vector spaces, justifying its use in capturing word dissimilarity. Kawin Ethayarajh, David Duvenaud, Graeme Hirst |
ACL (1) | 1 |
| 2019 | How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 EmbeddingsabstractReplacing static word embeddings with contextualized word representations has yielded significant improvements on many NLP tasks.However, just how contextual are the contextualized representations produced by models such as ELMo and BERT?Are there infinitely many context-specific representations for each word, or are words essentially assigned one of a finite number of word-sense representations?For one, we find that the contextualized representations of all words are not isotropic in any layer of the contextualizing model.While representations of the same word in different contexts still have a greater cosine similarity than those of two different words, this self-similarity is much lower in upper layers.This suggests that upper layers of contextualizing models produce more context-specific representations, much like how upper layers of LSTMs produce more task-specific representations.In all layers of ELMo, BERT, and GPT-2, on average, less than 5% of the variance in a word's contextualized representations can be explained by a static embedding for that word, providing some justification for the success of contextualized representations. Kawin Ethayarajh |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Rotate King to get Queen: Word Relationships as Orthogonal Transformations in Embedding SpaceabstractA notable property of word embeddings is that word relationships can exist as linear substructures in the embedding space.For example, gender corresponds to womanman and queenking.This, in turn, allows word analogies to be solved arithmetically: kingman + woman ≈ queen.This property is notable because it suggests that models trained on word embeddings can easily learn such relationships as geometric translations.However, there is no evidence that models exclusively represent relationships in this manner.We document an alternative way in which downstream models might learn these relationships: orthogonal and linear transformations.For example, given a translation vector for gender, we can find an orthogonal matrix R, representing a rotation and reflection, such that R( king) ≈ queen and R( man) ≈ woman.Analogical reasoning using orthogonal transformations is almost as accurate as using vector arithmetic; using linear transformations is more accurate than both.Our findings suggest that these transformations can be as good a representation of word relationships as translation vectors. Kawin Ethayarajh |
EMNLP/IJCNLP (1) | 1 |