Kawin Ethayarajh

dblp:198/6540 · DBLP profile ↗
← Back
12ranked-venue papers
9as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 9 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Language models and text generation · 38% Representation and self-supervised learning · 29% Trustworthy machine learning · 21%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › word representation
word embedding
1.132019
Rotate King to get Queen: Word Relationships as Orthogonal Transformations in Embedding Space · EMNLP/IJCNLP (1) 2019
Towards Understanding Linear Word Analogies · ACL (1) 2019
Understanding Undesirable Word Embedding Associations · ACL (1) 2019
Natural language and speech › Language models and text generation
text generation evaluation
1.022022
The Authenticity Gap in Human Evaluation · EMNLP 2022
Utility is in the Eye of the User: A Critique of NLP Leaderboards · EMNLP (1) 2020
Machine learning › Trustworthy machine learning
fairness
0.822020
Is Your Classifier Actually Biased? Measuring Fairness under Uncertainty with Bernstein Bounds · ACL 2020
Understanding Undesirable Word Embedding Associations · ACL (1) 2019
Natural language and speech › Language models and text generation
alignment
0.812024
Model Alignment as Prospect Theoretic Optimization · ICML 2024
Natural language and speech › Language models and text generation
preference optimization
0.812024
Model Alignment as Prospect Theoretic Optimization · ICML 2024
Machine learning › Representation and self-supervised learning
probing
0.512021
Conditional probing: measuring usable information beyond a baseline · EMNLP (1) 2021
Performance modeling and evaluation
benchmarking
0.512021
Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking · NeurIPS 2021
Natural language and speech › Language models and text generation › large language model evaluation › large language model benchmarking
leaderboard evaluation
0.412020
Utility is in the Eye of the User: A Critique of NLP Leaderboards · EMNLP (1) 2020
Machine learning › Trustworthy machine learning › fairness › fairness in NLP
bias in word embeddings
0.412019
Understanding Undesirable Word Embedding Associations · ACL (1) 2019
Machine learning › Representation and self-supervised learning › word representation
contextualized word representation
0.412019
How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings · EMNLP/IJCNLP (1) 2019
Machine learning › Trustworthy machine learning
debiasing
0.412019
Understanding Undesirable Word Embedding Associations · ACL (1) 2019
Natural language and speech › Information extraction and text analysis › lexical semantics
word analogy
0.412019
Rotate King to get Queen: Word Relationships as Orthogonal Transformations in Embedding Space · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation › text generation
story generation
0.212022
The Authenticity Gap in Human Evaluation · EMNLP 2022
Machine learning › Representation and self-supervised learning › representation analysis
neural representation analysis
0.112021
Conditional probing: measuring usable information beyond a baseline · EMNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

prospect theory · 0.8utility theory · 0.6v-information · 0.5utility-based aggregation · 0.5probing · 0.5microeconomic utility theory · 0.4confidence intervals · 0.4bernstein bounds · 0.4subspace projection · 0.4matrix factorization · 0.4WEAT · 0.4
YearPublicationVenuePosition
2024 Anchor Points: Benchmarking Models with Much Fewer Examples
abstract
Modern language models often exhibit powerful but brittle behavior, leading to the development of larger and more diverse benchmarks to reliably assess their behavior.Here, we suggest that model performance can be benchmarked and elucidated with much smaller evaluation sets.We first show that in six popular language classification benchmarks, model confidence in the correct class on many pairs of points is strongly correlated across models.We build upon this phenomenon to propose Anchor Point Selection, a technique to select small subsets of datasets that capture model behavior across the entire dataset.Anchor points reliably rank models: across 87 diverse language model-prompt pairs, evaluating models using 1-30 anchor points outperforms uniform sampling and other baselines at accurately ranking models.Moreover, just a dozen anchor points can be used to estimate model per-class predictions on all other points in a dataset with low error, sufficient for gauging where the model is likely to fail.Lastly, we present Anchor Point Maps for visualizing these insights and facilitating comparisons of the performance of different models on various regions within the dataset distribution.
Rajan Vivek, Kawin Ethayarajh, Diyi Yang, Douwe Kiela
EACL (1)2
2024 Model Alignment as Prospect Theoretic Optimization
abstract
Kahneman & Tversky’s $\textit{prospect theory}$ tells us that humans perceive random variables in a biased but well-defined manner (1992); for example, humans are famously loss-averse. We show that objectives for aligning LLMs with human feedback implicitly incorporate many of these biases—the success of these objectives (e.g., DPO) over cross-entropy minimization can partly be ascribed to them belonging to a family of loss functions that we call $\textit{human-aware losses}$ (HALOs). However, the utility functions these methods attribute to humans still differ from those in the prospect theory literature. Using a Kahneman-Tversky model of human utility, we propose a HALO that directly maximizes the utility of generations instead of maximizing the log-likelihood of preferences, as current methods do. We call this approach KTO, and it matches or exceeds the performance of preference-based methods at scales from 1B to 30B, despite only learning from a binary signal of whether an output is desirable. More broadly, our work suggests that there is no one HALO that is universally superior; the best loss depends on the inductive biases most appropriate for a given setting, an oft-overlooked consideration.
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Daniel Jurafsky, Douwe Kiela
ICML1
2022 The Authenticity Gap in Human Evaluation
abstract
Human ratings are the gold standard in NLG evaluation.The standard protocol is to collect ratings of generated text, average across annotators, and rank NLG systems by their average scores.However, little consideration has been given as to whether this approach faithfully captures human preferences.Analyzing this standard protocol through the lens of utility theory in economics, we identify the implicit assumptions it makes about annotators.These assumptions are often violated in practice, in which case annotator ratings cease to reflect their preferences.The most egregious violations come from using Likert scales, which provably reverse the direction of the true preference in certain cases.We suggest improvements to the standard protocol to make it more theoretically sound, but even in its improved form, it cannot be used to evaluate open-ended tasks like story generation.For the latter, we propose a new human evaluation protocol called systemlevel probabilistic assessment (SPA).When human evaluation of stories is done with SPA, we can recover the ordering of GPT-3 models by size, with statistically significant results.However, when human evaluation is done with the standard protocol, less than half of the expected preferences can be recovered (e.g., there is no significant difference between curie and davinci, despite using a highly powered test).
Kawin Ethayarajh, Daniel Jurafsky
EMNLP1
2022 Understanding Dataset Difficulty with V-Usable Information
Kawin Ethayarajh, Yejin Choi 0001, Swabha Swayamdipta
ICML1
2021 Conditional probing: measuring usable information beyond a baseline
abstract
Probing experiments investigate the extent to which neural representations make properties-like part-of-speech-predictable.One suggests that a representation encodes a property if probing that representation produces higher accuracy than probing a baseline representation like non-contextual word embeddings.Instead of using baselines as a point of comparison, we're interested in measuring information that is contained in the representation but not in the baseline.For example, current methods can detect when a representation is more useful than the word identity (a baseline) for predicting part-ofspeech; however, they cannot detect when the representation is predictive of just the aspects of part-of-speech not explainable by the word identity.In this work, we extend a theory of usable information called V-information and propose conditional probing, which explicitly conditions on the information in the baseline.In a case study, we find that after conditioning on non-contextual word embeddings, properties like part-of-speech are accessible at deeper layers of a network than previously thought.
John Hewitt, Kawin Ethayarajh, Percy Liang, Christopher D. Manning
EMNLP (1)2
2021 Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking
abstract
We introduce Dynaboard, an evaluation-as-a-service framework for hosting benchmarks and conducting holistic model comparison, integrated with the Dynabench platform. Our platform evaluates NLP models directly instead of relying on self-reported metrics or predictions on a single dataset. Under this paradigm, models are submitted to be evaluated in the cloud, circumventing the issues of reproducibility, accessibility, and backwards compatibility that often hinder benchmarking in NLP. This allows users to interact with uploaded models in real time to assess their quality, and permits the collection of additional metrics such as memory use, throughput, and robustness, which -- despite their importance to practitioners -- have traditionally been absent from leaderboards. On each task, models are ranked according to the Dynascore, a novel utility-based aggregation of these statistics, which users can customize to better reflect their preferences, placing more/less weight on a particular axis of evaluation or dataset. As state-of-the-art NLP models push the limits of traditional benchmarks, Dynaboard offers a standardized solution for a more diverse and comprehensive evaluation of model quality.
Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, Douwe Kiela
NeurIPS2
2020 Is Your Classifier Actually Biased? Measuring Fairness under Uncertainty with Bernstein Bounds
abstract
Most NLP datasets are not annotated with protected attributes such as gender, making it difficult to measure classification bias using standard measures of fairness (e.g., equal opportunity).However, manually annotating a large dataset with a protected attribute is slow and expensive.Instead of annotating all the examples, can we annotate a subset of them and use that sample to estimate the bias?While it is possible to do so, the smaller this annotated sample is, the less certain we are that the estimate is close to the true bias.In this work, we propose using Bernstein bounds to represent this uncertainty about the bias estimate as a confidence interval.We provide empirical evidence that a 95% confidence interval derived this way consistently bounds the true bias.In quantifying this uncertainty, our method, which we call Bernstein-bounded unfairness, helps prevent classifiers from being deemed biased or unbiased when there is insufficient evidence to make either claim.Our findings suggest that the datasets currently used to measure specific biases are too small to conclusively identify bias except in the most egregious cases.For example, consider a coreference resolution system that is 5% more accurate on gender-stereotypical sentences -to claim it is biased with 95% confidence, we need a bias-specific dataset that is 3.8 times larger than WinoBias, the largest available.
Kawin Ethayarajh
ACL1
2020 Utility is in the Eye of the User: A Critique of NLP Leaderboards
abstract
Benchmarks such as GLUE have helped drive advances in NLP by incentivizing the creation of more accurate models.While this leaderboard paradigm has been remarkably successful, a historical focus on performance-based evaluation has been at the expense of other qualities that the NLP community values in models, such as compactness, fairness, and energy efficiency.In this opinion paper, we study the divergence between what is incentivized by leaderboards and what is useful in practice through the lens of microeconomic theory.We frame both the leaderboard and NLP practitioners as consumers and the benefit they get from a model as its utility to them.With this framing, we formalize how leaderboards -in their current form -can be poor proxies for the NLP community at large.For example, a highly inefficient model would provide less utility to practitioners but not to a leaderboard, since it is a cost that only the former must bear.To allow practitioners to better estimate a model's utility to them, we advocate for more transparency on leaderboards, such as the reporting of statistics that are of practical concern (e.g., model size, energy efficiency, and inference latency).
Kawin Ethayarajh, Daniel Jurafsky
EMNLP (1)1
2019 Understanding Undesirable Word Embedding Associations
abstract
Word embeddings are often criticized for capturing undesirable word associations such as gender stereotypes.However, methods for measuring and removing such biases remain poorly understood.We show that for any embedding model that implicitly does matrix factorization, debiasing vectors post hoc using subspace projection (Bolukbasi et al., 2016) is, under certain conditions, equivalent to training on an unbiased corpus.We also prove that WEAT, the most common association test for word embeddings, systematically overestimates bias.Given that the subspace projection method is provably effective, we use it to derive a new measure of association called the relational inner product association (RIPA).Experiments with RIPA reveal that, on average, skipgram with negative sampling (SGNS) does not make most words any more gendered than they are in the training corpus.However, for gender-stereotyped words, SGNS actually amplifies the gender association in the corpus.
Kawin Ethayarajh, David Duvenaud, Graeme Hirst
ACL (1)1
2019 Towards Understanding Linear Word Analogies
abstract
A surprising property of word vectors is that word analogies can often be solved with vector arithmetic.However, it is unclear why arithmetic operators correspond to non-linear embedding models such as skip-gram with negative sampling (SGNS).We provide a formal explanation of this phenomenon without making the strong assumptions that past theories have made about the vector space and word distribution.Our theory has several implications.Past work has conjectured that linear substructures exist in vector spaces because relations can be represented as ratios; we prove that this holds for SGNS.We provide novel justification for the addition of SGNS word vectors by showing that it automatically downweights the more frequent word, as weighting schemes do ad hoc.Lastly, we offer an information theoretic interpretation of Euclidean distance in vector spaces, justifying its use in capturing word dissimilarity.
Kawin Ethayarajh, David Duvenaud, Graeme Hirst
ACL (1)1
2019 How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings
abstract
Replacing static word embeddings with contextualized word representations has yielded significant improvements on many NLP tasks.However, just how contextual are the contextualized representations produced by models such as ELMo and BERT?Are there infinitely many context-specific representations for each word, or are words essentially assigned one of a finite number of word-sense representations?For one, we find that the contextualized representations of all words are not isotropic in any layer of the contextualizing model.While representations of the same word in different contexts still have a greater cosine similarity than those of two different words, this self-similarity is much lower in upper layers.This suggests that upper layers of contextualizing models produce more context-specific representations, much like how upper layers of LSTMs produce more task-specific representations.In all layers of ELMo, BERT, and GPT-2, on average, less than 5% of the variance in a word's contextualized representations can be explained by a static embedding for that word, providing some justification for the success of contextualized representations.
Kawin Ethayarajh
EMNLP/IJCNLP (1)1
2019 Rotate King to get Queen: Word Relationships as Orthogonal Transformations in Embedding Space
abstract
A notable property of word embeddings is that word relationships can exist as linear substructures in the embedding space.For example, gender corresponds to womanman and queenking.This, in turn, allows word analogies to be solved arithmetically: kingman + woman ≈ queen.This property is notable because it suggests that models trained on word embeddings can easily learn such relationships as geometric translations.However, there is no evidence that models exclusively represent relationships in this manner.We document an alternative way in which downstream models might learn these relationships: orthogonal and linear transformations.For example, given a translation vector for gender, we can find an orthogonal matrix R, representing a rotation and reflection, such that R( king) ≈ queen and R( man) ≈ woman.Analogical reasoning using orthogonal transformations is almost as accurate as using vector arithmetic; using linear transformations is more accurate than both.Our findings suggest that these transformations can be as good a representation of word relationships as translation vectors.
Kawin Ethayarajh
EMNLP/IJCNLP (1)1