Esin Durmus

dblp:219/6227 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
13since 2021 · last 2025
0009-0009-7331-8160ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 6 first-author · 13 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 SafeArena: Evaluating the Safety of Autonomous Web Agents
abstract
LLM-based agents are becoming increasingly proficient at solving web-based tasks. With this capability comes a greater risk of misuse for malicious purposes, such as posting misinformation in an online forum or selling illicit substances on a website. To evaluate these risks, we propose SafeArena, a benchmark focused on the deliberate misuse of web agents. SafeArena comprises 250 safe and 250 harmful tasks across four websites. We classify the harmful tasks into five harm categories—misinformation, illegal activity, harassment, cybercrime, and social bias, designed to assess realistic misuses of web agents. We evaluate leading LLM-based web agents, including GPT-4o, Claude-3.5 Sonnet, Qwen-2-VL 72B, and Llama-3.2 90B, on our benchmark. To systematically assess their susceptibility to harmful tasks, we introduce the Agent Risk Assessment framework that categorizes agent behavior across four risk levels. We find agents are surprisingly compliant with malicious requests, with GPT-4o and Qwen-2 completing 34.7% and 27.3% of harmful requests, respectively. Our findings highlight the urgent need for safety alignment procedures for web agents.
Ada Defne Tur, Nicholas Meade, Xing Han Lù, Alejandra Zambrano, Arkil Patel, Esin Durmus, Spandana Gella, Karolina Stanczak, Siva Reddy
ICML6
2024 Towards Understanding Sycophancy in Language Models
abstract
Reinforcement learning from human feedback (RLHF) is a popular technique for training high-quality AI assistants. However, RLHF may also encourage model responses that match user beliefs over truthful responses, a behavior known as sycophancy. We investigate the prevalence of sycophancy in RLHF-trained models and whether human preference judgments are responsible. We first demonstrate that five state-of-the-art AI assistants consistently exhibit sycophancy behavior across four varied free-form text-generation tasks. To understand if human preferences drive this broadly observed behavior of RLHF models, we analyze existing human preference data. We find that when a response matches a user's views, it is more likely to be preferred. Moreover, both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time. Optimizing model outputs against PMs also sometimes sacrifices truthfulness in favor of sycophancy. Overall, our results indicate that sycophancy is a general behavior of RLHF models, likely driven in part by human preference judgments favoring sycophantic responses.
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Miranda Zhang, Ethan Perez
ICLR7
2024 NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps
abstract
Kristina Gligoric, Myra Cheng, Lucia Zheng, Esin Durmus, Dan Jurafsky. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Kristina Gligoric, Myra Cheng, Lucia Zheng, Esin Durmus, Daniel Jurafsky
NAACL-HLT4
2024 Many-shot Jailbreaking
abstract
We investigate a family of simple long-context attacks on large language models: prompting with hundreds of demonstrations of undesirable behavior. This attack is newly feasible with the larger context windows recently deployed by language model providers like Google DeepMind, OpenAI and Anthropic. We find that in diverse, realistic circumstances, the effectiveness of this attack follows a power law, up to hundreds of shots. We demonstrate the success of this attack on the most widely used state-of-the-art closed-weight models, and across various tasks. Our results suggest very long contexts present a rich new attack surface for LLMs.
Cem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Meg Tong, Jesse Mu, Daniel Ford, Francesco Mosconi, Rajashree Agrawal, Rylan Schaeffer, Naomi Bashkansky, Samuel Svenningsen, Mike Lambert, Ansh Radhakrishnan, Carson Denison, Evan Hubinger, Yuntao Bai, Trenton Bricken, Timothy Maxwell, Nicholas Schiefer, James Sully, Alex Tamkin, Tamera Lanham, Karina Nguyen, Tomek Korbak, Jared Kaplan, Deep Ganguli, Samuel R. Bowman, Ethan Perez, Roger B. Grosse, David Duvenaud
NeurIPS2
2024 Benchmarking Large Language Models for News Summarization
abstract
Abstract Large language models (LLMs) have shown promise for automatic summarization but the reasons behind their successes are poorly understood. By conducting a human evaluation on ten LLMs across different pretraining methods, prompts, and model scales, we make two important observations. First, we find instruction tuning, not model size, is the key to the LLM’s zero-shot summarization capability. Second, existing studies have been limited by low-quality references, leading to underestimates of human performance and lower few-shot and finetuning performance. To better evaluate LLMs, we perform human evaluation over high-quality summaries we collect from freelance writers. Despite major stylistic differences such as the amount of paraphrasing, we find that LLM summaries are judged to be on par with human written summaries.
Faisal Ladhak, Esin Durmus, Percy Liang, Kathy McKeown, Tatsunori B. Hashimoto
Trans. Assoc. Comput. Linguistics3
2023 Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models
abstract
To recognize and mitigate harms from large language models (LLMs), we need to understand the prevalence and nuances of stereotypes in LLM outputs.Toward this end, we present Marked Personas, a prompt-based method to measure stereotypes in LLMs for intersectional demographic groups without any lexicon or data labeling.Grounded in the sociolinguistic concept of markedness (which characterizes explicitly linguistically marked categories versus unmarked defaults), our proposed method is twofold: 1) prompting an LLM to generate personas, i.e., natural language descriptions, of the target demographic group alongside personas of unmarked, default groups; 2) identifying the words that significantly distinguish personas of the target group from corresponding unmarked ones.We find that the portrayals generated by GPT-3.5 and GPT-4 contain higher rates of racial stereotypes than human-written portrayals using the same prompts.The words distinguishing personas of marked (non-white, non-male) groups reflect patterns of othering and exoticizing these demographics.An intersectional lens further reveals tropes that dominate portrayals of marginalized groups, such as tropicalism and the hypersexualization of minoritized women.These representational harms have concerning implications for downstream applications like story generation.
Myra Cheng, Esin Durmus, Daniel Jurafsky
ACL (1)2
2023 Contrastive Error Attribution for Finetuned Language Models
abstract
Recent work has identified noisy and misannotated data as a core cause of hallucinations and unfaithful outputs in Natural Language Generation (NLG) tasks.Consequently, identifying and removing these examples is a key open challenge in creating reliable NLG systems.In this work, we introduce a framework to identify and remove low-quality training instances that lead to undesirable outputs, such as faithfulness errors in text summarization.We show that existing approaches for error tracing, such as gradient-based influence measures, do not perform reliably for detecting faithfulness errors in NLG datasets.We overcome the drawbacks of existing error tracing methods through a new, contrast-based estimate that compares undesired generations to human-corrected outputs.Our proposed method can achieve a mean average precision of 0.93 at detecting known data errors across synthetic tasks with known ground truth, substantially outperforming existing approaches.Using this approach and re-training models on cleaned data leads to a 70% reduction in entity hallucinations on the NYT dataset and a 55% reduction in semantic errors on the E2E dataset.
Faisal Ladhak, Esin Durmus, Tatsunori B. Hashimoto
ACL (1)2
2023 When Do Pre-Training Biases Propagate to Downstream Tasks? A Case Study in Text Summarization
abstract
Faisal Ladhak, Esin Durmus, Mirac Suzgun, Tianyi Zhang, Dan Jurafsky, Kathleen McKeown, Tatsunori Hashimoto. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Faisal Ladhak, Esin Durmus, Mirac Suzgun, Daniel Jurafsky, Kathy McKeown, Tatsunori B. Hashimoto
EACL2
2023 Whose Opinions Do Language Models Reflect?
abstract
Language models (LMs) are increasingly being used in open-ended contexts, where the opinions they reflect in response to subjective queries can have a profound impact, both on user satisfaction, and shaping the views of society at large. We put forth a quantitative framework to investigate the opinions reflected by LMs – by leveraging high-quality public opinion polls. Using this framework, we create OpinionQA, a dataset for evaluating the alignment of LM opinions with those of 60 US demographic groups over topics ranging from abortion to automation. Across topics, we find substantial misalignment between the views reflected by current LMs and those of US demographic groups: on par with the Democrat-Republican divide on climate change. Notably, this misalignment persists even after explicitly steering the LMs towards particular groups. Our analysis not only confirms prior observations about the left-leaning tendencies of some human feedback-tuned LMs, but also surfaces groups whose opinions are poorly reflected by current LMs (e.g., 65+ and widowed individuals).
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, Tatsunori B. Hashimoto
ICML2
2022 Spurious Correlations in Reference-Free Evaluation of Text Generation
abstract
Model-based, reference-free evaluation metrics have been proposed as a fast and cost-effective approach to evaluate Natural Language Generation (NLG) systems.Despite promising recent results, we find evidence that reference-free evaluation metrics of summarization and dialog generation may be relying on spurious correlations with measures such as word overlap, perplexity, and length.We further observe that for text summarization, these metrics have high error rates when ranking current state-ofthe-art abstractive summarization systems.We demonstrate that these errors can be mitigated by explicitly designing evaluation metrics to avoid spurious features in reference-free evaluation.
Esin Durmus, Faisal Ladhak, Tatsunori B. Hashimoto
ACL (1)1
2022 Faithful or Extractive? On Mitigating the Faithfulness-Abstractiveness Trade-off in Abstractive Summarization
abstract
Despite recent progress in abstractive summarization, systems still suffer from faithfulness errors.While prior work has proposed models that improve faithfulness, it is unclear whether the improvement comes from an increased level of extractiveness of the model outputs as one naive way to improve faithfulness is to make summarization models more extractive.In this work, we present a framework for evaluating the effective faithfulness of summarization systems, by generating a faithfulnessabstractiveness trade-off curve that serves as a control at different operating points on the abstractiveness spectrum.We then show that the baseline system as well as recently proposed methods for improving faithfulness, fail to consistently improve over the control at the same level of abstractiveness.Finally, we learn a selector to identify the most faithful and abstractive summary for a given document, and show that this system can attain higher faithfulness scores in human evaluations while being more abstractive than the baseline system on two datasets.Moreover, we show that our system is able to achieve a better faithfulnessabstractiveness trade-off than the control at the same level of abstractiveness.
Faisal Ladhak, Esin Durmus, He He 0001, Claire Cardie, Kathy McKeown
ACL (1)2
2022 Improving Faithfulness by Augmenting Negative Summaries from Fake Documents
abstract
Current abstractive summarization systems tend to hallucinate content that is unfaithful to the source document, posing a risk of misinformation.To mitigate hallucination, we must teach the model to distinguish hallucinated summaries from faithful ones.However, the commonly used maximum likelihood training does not disentangle factual errors from other model errors.To address this issue, we propose a back-translation-style approach to augment negative samples that mimic factual errors made by the model.Specifically, we train an elaboration model that generates hallucinated documents given the reference summaries, and then generates negative summaries from the fake documents.We incorporate the negative samples into training through a controlled generator, which produces faithful/unfaithful summaries conditioned on the control codes.Additionally, we find that adding textual entailment data through multitasking further boosts the performance.Experiments on three datasets (XSum, GigaWord, and WikiHow) show that our method consistently improves faithfulness without sacrificing informativeness according to both human and automatic evaluation. 1
Faisal Ladhak, Esin Durmus, He He 0001
EMNLP3
2022 Language modeling via stochastic processes
Rose E. Wang, Esin Durmus, Noah D. Goodman, Tatsunori B. Hashimoto
ICLR2
2020 FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive Summarization
abstract
Neural abstractive summarization models are prone to generate content inconsistent with the source document, i.e. unfaithful.Existing automatic metrics do not capture such mistakes effectively.We tackle the problem of evaluating faithfulness of a generated summary given its source document.We first collected human annotations of faithfulness for outputs from numerous models on two datasets.We find that current models exhibit a trade-off between abstractiveness and faithfulness: outputs with less word overlap with the source document are more likely to be unfaithful.Next, we propose an automatic question answering (QA) based metric for faithfulness, FEQA, 1 which leverages recent advances in reading comprehension.Given questionanswer pairs generated from the summary, a QA model extracts answers from the document; non-matched answers indicate unfaithful information in the summary.Among metrics based on word overlap, embedding similarity, and learned language understanding models, our QA-based metric has significantly higher correlation with human faithfulness scores, especially on highly abstractive summaries.* Most of the work is done while the authors were at Amazon Web Services AI.1 Faithfulness Evaluation with Question Answering.
Esin Durmus, He He 0001, Mona T. Diab
ACL1
2020 Exploring the Role of Argument Structure in Online Debate Persuasion
abstract
Online debate forums provide users a platform to express their opinions on controversial topics while being exposed to opinions from diverse set of viewpoints.Existing work in Natural Language Processing (NLP) has shown that linguistic features extracted from the debate text and features encoding the characteristics of the audience are both critical in persuasion studies.In this paper, we aim to further investigate the role of discourse structure of the arguments from online debates in their persuasiveness.In particular, we use the factor graph model to obtain features for the argument structure of debates from an online debating platform and incorporate these features to an LSTM-based model to predict the debater that makes the most convincing arguments.We find that incorporating argument structure features play an essential role in achieving the better predictive performance in assessing the persuasiveness of the arguments in online debates.
Jialu Li 0001, Esin Durmus, Claire Cardie
EMNLP (1)2
2019 A Corpus for Modeling User and Language Effects in Argumentation on Online Debating
abstract
Existing argumentation datasets have succeeded in allowing researchers to develop computational methods for analyzing the content, structure and linguistic features of argumentative text.They have been much less successful in fostering studies of the effect of "user" traits -characteristics and beliefs of the participants -on the debate/argument outcome as this type of user information is generally not available.This paper presents a dataset of 78, 376 debates generated over a 10-year period along with surprisingly comprehensive participant profiles.We also complete an example study using the dataset to analyze the effect of selected user traits on the debate outcome in comparison to the linguistic features typically employed in studies of this kind.
Esin Durmus, Claire Cardie
ACL (1)1
2019 Determining Relative Argument Specificity and Stance for Complex Argumentative Structures
abstract
Systems for automatic argument generation and debate require the ability to (1) determine the stance of any claims employed in the argument and (2) assess the specificity of each claim relative to the argument context.Existing work on understanding claim specificity and stance, however, has been limited to the study of argumentative structures that are relatively shallow, most often consisting of a single claim that directly supports or opposes the argument thesis.In this paper, we tackle these tasks in the context of complex arguments on a diverse set of topics.In particular, our dataset consists of manually curated argument trees for 741 controversial topics covering 95,312 unique claims; lines of argument are generally of depth 2 to 6.We find that as the distance between a pair of claims increases along the argument path, determining the relative specificity of a pair of claims becomes easier and determining their relative stance becomes harder.
Esin Durmus, Faisal Ladhak, Claire Cardie
ACL (1)1
2019 The Role of Pragmatic and Discourse Context in Determining Argument Impact
abstract
Esin Durmus, Faisal Ladhak, Claire Cardie. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Esin Durmus, Faisal Ladhak, Claire Cardie
EMNLP/IJCNLP (1)1
2019 Modeling the Factors of User Success in Online Debate
abstract
Debate is a process that gives individuals the opportunity to express, and to be exposed to, diverging viewpoints on controversial issues; and the existence of online debating platforms makes it easier for individuals to participate in debates and obtain feedback on their debating skills. But understanding the factors that contribute to a user's success in debate is complicated: while success depends, in part, on the characteristics of the language they employ, it is also important to account for the degree to which their beliefs and personal traits are compatible with that of the audience. Friendships and previous interactions among users on the platform may further influence success.
Esin Durmus, Claire Cardie
WWW1
2018 Exploring the Role of Prior Beliefs for Argument Persuasion
abstract
Esin Durmus, Claire Cardie. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Esin Durmus, Claire Cardie
NAACL-HLT1