EDBT 2026 Demo / reviewers in the wild / expert
Eunsol Choi
dblp:116/2765
· DBLP profile ↗
54ranked-venue papers
7as first author
40since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 7 first-author · 35 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real UsersabstractNishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan Lee Boyd-Graber, Aakanksha Naik. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan L. Boyd-Graber, Aakanksha Naik |
ACL (1) | 8 |
| 2026 | BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice BenchmarksabstractNishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie, Atrey Desai, Vipul Gupta, Steven James Moore, Eunsol Choi, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Nishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie, Atrey Desai, Steven James Moore, Eunsol Choi, Rachel Rudinger, Jordan L. Boyd-Graber |
ACL (1) | 8 |
| 2026 | Envisioning Future Food, Technology, and Health: Teens' Perspectives Through Design FictionabstractTeenage years are a critical period for shaping food practices and health behaviors, yet teens remain underrepresented in research on future food and health technologies. This paper reports on design fiction workshops where 20 teens speculated on artifacts such as health-tracking mirrors and food-making machines while reflecting on their eating habits and values. Our analysis shows how teens negotiate their human stances in a techno-centric world, balancing excitement about innovation with concerns over health, control, and responsibility. We contribute to HCI by: (1) identifying teen-specific design considerations for reflective healthy-eating technologies, highlighting “absence” as a protective affordance; (2) reframing food technologies for teens as sites of family care and playful identity work, opening a design space at the intersection of reflection and play; and (3) advancing methodological understanding of youth-centered design fiction by showing how dual-mode speculative workshops can move teen participants beyond solutionist or anti-solutionist positions toward grounded negotiation with sociotechnical systems. Chun-Han Ariel Wang, Eunsol Choi, Norman Makoto Su, Chia-Fang Chung |
CHI | 2 |
| 2025 | CaLMQA: Exploring culturally specific long-form question answering across 23 languagesabstractShane Arora, Marzena Karpinska, Hung-Ting Chen, Ipsita Bhattacharjee, Mohit Iyyer, Eunsol Choi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shane Arora, Marzena Karpinska, Hung-Ting Chen, Ipsita Bhattacharjee, Mohit Iyyer, Eunsol Choi |
ACL (1) | 6 |
| 2025 | RefreshKV: Updating Small KV Cache During Long-form GenerationabstractGenerating long sequences of tokens given a long-context input is a very compute-intensive inference scenario for large language models (LLMs).One prominent inference speed-up approach is to construct a smaller key-value (KV) cache, relieving LLMs from computing attention over a long sequence of tokens.While such methods work well to generate short sequences, their performance degrades rapidly for longform generation.Most KV compression happens once, prematurely removing tokens that can be useful later in the generation.We propose a new inference method, RefreshKV, that flexibly alternates between full context attention and attention over a subset of input tokens during generation.After each full attention step, we update the smaller KV cache based on the attention pattern over the entire input.Applying our method to off-the-shelf LLMs achieves comparable speedup to eviction-based methods while improving performance for various long-form generation tasks.Lastly, we show that continued pretraining with our inference setting brings further gains in performance. Tanya Goyal, Eunsol Choi |
ACL (1) | 3 |
| 2025 | Scaling Rich Style-Prompted Text-to-Speech DatasetsabstractWe introduce Paralinguistic Speech Captions (ParaSpeechCaps), a large-scale dataset that annotates speech utterances with rich style captions.While rich abstract tags (e.g.guttural, nasal, pained) have been explored in small-scale human-annotated datasets, existing large-scale datasets only cover basic tags (e.g.low-pitched, slow, loud).We combine off-the-shelf text and speech embedders, classifiers and an audio language model to automatically scale rich tag annotations for the first time.ParaSpeechCaps covers a total of 59 style tags, including both speaker-level intrinsic tags and utterance-level situational tags.It consists of 282 hours of human-labelled data (PSC-Base) and 2427 hours of automatically annotated data (PSC-Scaled).We finetune Parler-TTS, an open-source style-prompted TTS model, on ParaSpeechCaps, and achieve improved style consistency (+7.9% Consistency MOS) and speech quality (+15.5% Naturalness MOS) over the best performing baseline that combines existing rich style tag datasets.We ablate several of our dataset design choices to lay the foundation for future work in this space.Our dataset, models and code are released at https://github. com/ajd12342/paraspeechcaps. Anuj Diwan, Zhisheng Zheng, David F. Harwath, Eunsol Choi |
EMNLP | 4 |
| 2025 | User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning SignalabstractOnce language models (LMs) are deployed, they can interact with users long-term, ideally evolving based on their feedback.Asking for direct user feedback can be disruptive; thus, we study harvesting implicit user feedback from user-LM interaction logs.We study two user-LM interaction datasets (WildChat and LMSYS).First, we analyze user feedback in the user-LLM conversation logs, providing insights into when and why such feedback occurs.Second, we study harvesting learning signals from such implicit user feedback.Specifically, we study whether incorporating the contents of user feedback (e.g., user wanted clarification), in addition to the polarity of the feedback, can improve the model performance.We observe mixed results, showing this helps in short human-designed questions (MTBench) but not on longer and more complex questions (Wild-Bench).Together, we provide an in-depth study of implicit user feedback, showing its potential and limitations. Michael JQ Zhang, Eunsol Choi |
EMNLP | 3 |
| 2025 | Modeling Future Conversation Turns to Teach LLMs to Ask Clarifying QuestionsabstractLarge language models (LLMs) must often respond to highly ambiguous user requests. In such cases, the LLM's best response may be to ask a clarifying question to elicit more information. Existing LLMs often respond by presupposing a single interpretation of such ambiguous requests, frustrating users who intended a different interpretation. We speculate this is caused by current preference data labeling practice, where LLM responses are evaluated only on their prior contexts. To address this, we assign preference labels by simulating their expected outcomes in future turns. This allows LLMs to learn to ask clarifying questions when it can generate responses that are tailored to each user interpretation in future turns. On open-domain QA datasets with multiple annotations, we evaluate systems based on their ability to ask clarifying questions to recover each user's interpretation and expected answer. We compare systems trained using our proposed preference labeling methods against standard methods, which assign preferences based on only prior context. Our method achieves a 5% improvement in F1 measured against the answer set from different interpretations of each query, showing the value of modeling future conversation turns. We further demonstrate that our method can be used to train models to judiciously determine when to ask clarifying questions, directly answering the question when clarification is unnecessary. In our experiments, we find that our method achives a 3% improvement in accuracy of such judgments over existing methods. Michael J. Q. Zhang, W. Bradley Knox, Eunsol Choi |
ICLR | 3 |
| 2025 | Diverging Preferences: When do Annotators Disagree and do Models Know?abstractWe examine diverging preferences in human-labeled preference datasets. We develop a taxonomy of disagreement sources spanning ten categories across four high-level classes and find that the majority of disagreements are due to factors such as task underspecification or response style. Our findings challenge a standard assumption in reward modeling methods that annotator disagreements can be attributed to simple noise. We then explore how these findings impact two areas of LLM development: reward modeling training and evaluation. In our experiments, we demonstrate how standard reward modeling (e.g., Bradley-Terry) and LLM-as-Judge evaluation methods fail to account for divergence between annotators. These findings highlight challenges in LLM evaluations, which are greatly influenced by divisive features like response style, and in developing pluralistically aligned LLMs. To address these issues, we develop methods for identifying diverging preferences to mitigate their influence in evaluations and during LLM training. Michael J. Q. Zhang, Zhilin Wang, Jena D. Hwang, Yi Dong 0003, Olivier Delalleau, Yejin Choi 0001, Eunsol Choi, Xiang Ren 0001, Valentina Pyatkin |
ICML | 7 |
| 2025 | Open-World Evaluation for Retrieving Diverse PerspectivesabstractHung-Ting Chen, Eunsol Choi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hung-Ting Chen, Eunsol Choi |
NAACL (Long Papers) | 2 |
| 2025 | From Distributional to Overton Pluralism: Investigating Large Language Model AlignmentabstractThom Lake, Eunsol Choi, Greg Durrett. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Thom Lake, Eunsol Choi, Greg Durrett |
NAACL (Long Papers) | 2 |
| 2024 | Aligning Data with the Goals of an Organization and Its Workers: Designing Data Labeling for Social Service Case NotesabstractThe challenges of data collection in nonprofits for performance and funding reports are well-established in HCI research. Few studies, however, delve into improving the data collection process. Our study proposes ideas to improve data collection by exploring challenges that social workers experience when labeling their case notes. Through collaboration with an organization that provides intensive case management to those experiencing homelessness in the U.S., we conducted interviews with caseworkers and held design sessions where caseworkers, managers, and program analysts examined storyboarded ideas to improve data labeling. Our findings suggest several design ideas on how data labeling practices can be improved: Aligning labeling with caseworker goals, enabling shared control on data label design for a comprehensive portrayal of caseworker contributions, improving the synthesis of qualitative and quantitative data, and making labeling user-friendly. We contribute design implications for data labeling to better support multiple stakeholder goals in social service contexts. Apoorva Gondimalla, Varshinee Sreekanth, Govind Joshi, Whitney Nelson, Eunsol Choi, Stephen C. Slota, Sherri R. Greenberg, Kenneth R. Fleischmann, Min Kyung Lee |
CHI | 5 |
| 2024 | Learning to Reject with a Fixed Predictor: Application to DecontextualizationabstractWe study the problem of classification with a reject option for a fixed predictor, crucial to natural language processing. We introduce a new problem formulation for this scenario, and an algorithm minimizing a new surrogate loss function. We provide a complete theoretical analysis of the surrogate loss function with a strong $H$-consistency guarantee. For evaluation, we choose the \textit{decontextualization} task, and provide a manually-labelled dataset of $2\mathord,000$ examples. Our algorithm significantly outperforms the baselines considered, with a $\sim 25$% improvement in coverage when halving the error rate, which is only $\sim 3$% away from the theoretical limit. Christopher Mohri, Daniel Andor, Eunsol Choi, Michael Collins 0001, Anqi Mao, Yutao Zhong 0002 |
ICLR | 3 |
| 2024 | RECOMP: Improving Retrieval-Augmented LMs with Context Compression and Selective AugmentationabstractRetrieval-augmented language models improve language models (LMs) by retrieving documents and prepending them in-context.
However, these documents, often spanning hundreds of words, make inference substantially less efficient. We propose compressing the retrieved documents into textual summaries prior to in-context integration. This not only reduces the computational costs but also relieve the burden of LMs to identify relevant information in long retrieved documents. We present two compressors -- an extractive compressor which selects useful sentences from retrieved documents and an abstractive compressor which generates summary by synthesizing information from multiple documents. Both are trained to achieve performance gain in LMs when we prepend the generated summary from the compressor to LMs' input, while minimizing the summary length. When retrieved documents are irrelevant to the input or offer no additional information to LM, our compressors output an empty string, enabling selective augmentation. We evaluate our approach on the language modeling task and open domain question answering task. We achieve a compression rate of as low as 6% with minimal loss in performance for both tasks, significantly outperforming the off-the-shelf summarization models. We show that our compressors trained for one LM can transfer to other LMs on the language modeling task and provide a summary largely faithful to the retrieved documents. Eunsol Choi |
ICLR | 3 |
| 2024 | BAT: Learning to Reason about Spatial Sounds with Large Language ModelsabstractSpatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed SpatialSoundQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or Spatial-AST, which by itself achieves strong performance across sound event detection, spatial localization, and distance estimation. By integrating Spatial-AST with LLaMA-2 7B model, BAT transcends standard Sound Event Localization and Detection (SELD) tasks, enabling the model to reason about the relationships between the sounds in its environment. Our experiments demonstrate BAT’s superior performance on both spatial sound perception and reasoning, showcasing the immense potential of LLMs in navigating and interpreting complex spatial audio environments. Zhisheng Zheng, Puyuan Peng, Ziyang Ma 0001, Xie Chen 0001, Eunsol Choi, David F. Harwath |
ICML | 5 |
| 2024 | Complex Claim Verification with Evidence Retrieved in the WildabstractJifan Chen, Grace Kim, Aniruddh Sriram, Greg Durrett, Eunsol Choi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jifan Chen, Grace Kim, Aniruddh Sriram, Greg Durrett, Eunsol Choi |
NAACL-HLT | 5 |
| 2024 | SVFT: Parameter-Efficient Fine-Tuning with Singular VectorsabstractPopular parameter-efficient fine-tuning (PEFT) methods, such as LoRA and its variants, freeze pre-trained model weights $\(\mathbf{W}\)$ and inject learnable matrices $\(\mathbf{\Delta W}\)$. These $\(\mathbf{\Delta W}\)$ matrices are structured for efficient parameterization, often using techniques like low-rank approximations or scaling vectors. However, these methods typically exhibit a performance gap compared to full fine-tuning. While recent PEFT methods have narrowed this gap, they do so at the expense of additional learnable parameters. We propose SVFT, a *simple* approach that structures $\(\mathbf{\Delta W}\)$ based on the specific weight matrix $\(\mathbf{W}\)$. SVFT updates $\(\mathbf{W}\)$ as a sparse combination $\(M\)$ of outer products of its singular vectors, training only the coefficients of these combinations. Crucially, we make additional off-diagonal elements in $M$ learnable, enabling a smooth trade-off between trainable parameters and expressivity—an aspect that distinctly sets our approach apart from previous works leveraging singular values. Extensive experiments on language and vision benchmarks show that SVFT recovers up to **96%** of full fine-tuning performance while training only **0.006 to 0.25%** of parameters, outperforming existing methods that achieve only up to **{85\%}** performance with **0.03 to 0.8%** of the trainable parameter budget. Vijay Lingam, Atula Neerkaje, Aditya Vavre, Aneesh Shetty, Gautham Krishna Gudur, Joydeep Ghosh, Eunsol Choi, Alexandros G. Dimakis, Aleksandar Bojchevski, Sujay Sanghavi |
NeurIPS | 7 |
| 2023 | Can LMs Learn New Entities from Descriptions? Challenges in Propagating Injected KnowledgeabstractPre-trained language models (LMs) are used for knowledge intensive tasks like question answering, but their knowledge gets continuously outdated as the world changes.Prior work has studied targeted updates to LMs, injecting individual facts and evaluating whether the model learns these facts while not changing predictions on other contexts.We take a step forward and study LMs' abilities to make inferences based on injected facts (or propagate those facts): for example, after learning that something is a TV show, does an LM predict that you can watch it?We study this with two clozestyle tasks: an existing dataset of real-world sentences about novel entities (ECBD) as well as a new controlled benchmark with manually designed templates requiring varying levels of inference about injected knowledge.Surprisingly, we find that existing methods for updating knowledge (gradient-based fine-tuning and modifications of this approach) show little propagation of injected knowledge.These methods improve performance on cloze instances only when there is lexical overlap between injected facts and target inferences.Yet, prepending entity definitions in an LM's context improves performance across all settings, suggesting that there is substantial headroom for parameterupdating approaches for knowledge injection. Yasumasa Onoe, Michael J. Q. Zhang, Shankar Padmanabhan, Greg Durrett, Eunsol Choi |
ACL (1) | 5 |
| 2023 | Concise Answers to Complex Questions: Summarization of Long-form AnswersabstractLong-form question answering systems provide rich information by presenting paragraph-level answers, often containing optional background or auxiliary information.While such comprehensive answers are helpful, not all information is required to answer the question (e.g.users with domain knowledge do not need an explanation of background).Can we provide a concise version of the answer by summarizing it, while still addressing the question?We conduct a user study on summarized answers generated from state-of-the-art models and our newly proposed extract-and-decontextualize approach.We find a large proportion of long-form answers (over 90%) in the ELI5 domain can be adequately summarized by at least one system, while complex and implicit answers are challenging to compress.We observe that decontextualization improves the quality of the extractive summary, exemplifying its potential in the summarization task.To promote future work, we provide an extractive summarization dataset covering 1K long-form answers and our user study annotations.Together, we present the first study on summarizing long-form answers, taking a step forward for QA agents that can provide answers at multiple granularities. Abhilash Potluri, Eunsol Choi |
ACL (1) | 3 |
| 2023 | A Critical Evaluation of Evaluations for Long-form Question AnsweringabstractLong-form question answering (LFQA) enables answering a wide range of questions, but its flexibility poses enormous challenges for evaluation.We perform the first targeted study of the evaluation of long-form answers, covering both human and automatic evaluation practices.We hire domain experts in seven areas to provide preference judgments over pairs of answers, along with free-form justifications for their choices.We present a careful analysis of experts' evaluation, which focuses on new aspects such as the comprehensiveness of the answer.Next, we examine automatic text generation metrics, finding that no existing metrics are predictive of human preference judgments.However, some metrics correlate with fine-grained aspects of answers (e.g., coherence).We encourage future work to move away from a single "overall score" of the answer and adopt a multi-faceted evaluation, targeting aspects such as factuality and completeness.We publicly release all of our annotations and code to spur future work into LFQA evaluation. 1 Yixiao Song, Mohit Iyyer, Eunsol Choi |
ACL (1) | 4 |
| 2023 | DiffQG: Generating Questions to Summarize Factual ChangesabstractJeremy R. Cole, Palak Jain, Julian Martin Eisenschlos, Michael J.Q. Zhang, Eunsol Choi, Bhuwan Dhingra. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Jeremy R. Cole, Palak Jain 0006, Julian Martin Eisenschlos, Michael J. Q. Zhang, Eunsol Choi, Bhuwan Dhingra |
EACL | 5 |
| 2023 | Continually Improving Extractive QA via Human FeedbackabstractWe study continually improving an extractive question answering (QA) system via human user feedback.We design and deploy an iterative approach, where information-seeking users ask questions, receive model-predicted answers, and provide feedback.We conduct experiments involving thousands of user interactions under diverse setups to broaden the understanding of learning from feedback over time.Our experiments show effective improvement from user feedback of extractive QA models over time across different data regimes, including significant potential for domain adaptation.* Equal contribution. 1 The term continual learning is at times used to refer to a scenario where models adapt to new tasks over time.We study improving the model continually on its original task. Hung-Ting Chen, Yoav Artzi, Eunsol Choi |
EMNLP | 4 |
| 2023 | Mitigating Temporal Misalignment by Discarding Outdated FactsabstractWhile large language models are able to retain vast amounts of world knowledge seen during pretraining, such knowledge is prone to going out of date and is nontrivial to update.Furthermore, these models are often used under temporal misalignment, tasked with answering questions about the present, despite having only been trained on data collected in the past.To mitigate the effects of temporal misalignment, we propose fact duration prediction: the task of predicting how long a given fact will remain true.In our experiments, we demonstrate that identifying which facts are prone to rapid change can help models avoid reciting outdated information and determine which predictions require seeking out up-to-date knowledge sources.We also show how modeling fact duration improves calibration for knowledgeintensive tasks, such as open-retrieval question answering, under temporal misalignment, by discarding volatile facts. Michael J. Q. Zhang, Eunsol Choi |
EMNLP | 2 |
| 2023 | Continual Learning for On-Device Speech Recognition Using Disentangled ConformersabstractAutomatic speech recognition research focuses on training and evaluating on static datasets. Yet, as speech models are increasingly deployed on personal devices, such models encounter user-specific distributional shifts. To simulate this real-world scenario, we introduce LibriContinual, a continual learning benchmark for speaker-specific domain adaptation derived from LibriVox audiobooks, with data corresponding to 118 individual speakers and 6 train splits per speaker of different sizes. Additionally, current speech recognition models and continual learning algorithms are not optimized to be compute-efficient. We adapt a general-purpose training algorithm NetAug for ASR and create a novel Conformer variant called the DisConformer (Disentangled Conformer). This algorithm produces ASR models consisting of a frozen ‘core’ network for general-purpose use and several tunable ‘augment’ networks for speaker-specific tuning. Using such models, we propose a novel compute-efficient continual learning algorithm called DisentangledCL. Our experiments show that the DisConformer models significantly outperform base-lines on general ASR i.e. LibriSpeech (15.58% rel. WER on test-other). On speaker-specific LibriContinual they significantly outper-form trainable-parameter-matched baselines (by 20.65% rel. WER on test) and even match fully finetuned baselines in some settings. Anuj Diwan, Ching-Feng Yeh, Wei-Ning Hsu, Paden Tomasello, Eunsol Choi, David F. Harwath, Abdel-rahman Mohamed |
ICASSP | 5 |
| 2023 | Propagating Knowledge Updates to LMs Through DistillationabstractModern language models have the capacity to store and use immense amounts of knowledge about real-world entities, but it remains unclear how to update such knowledge stored in model parameters. While prior methods for updating knowledge in LMs successfully inject atomic facts, updated LMs fail to make inferences based on injected facts. In this work, we demonstrate that a context distillation-based approach can both impart knowledge about entities \emph{and} propagate that knowledge to enable broader inferences. Our approach consists of two stages: transfer set generation and distillation on the transfer set. We first generate a transfer set by prompting a language model to generate continuations from the entity definition. Then, we update the model parameters so that the distribution of the LM (the 'student') matches the distribution of the LM conditioned on the definition (the 'teacher') on the transfer set. Our experiments demonstrate that this approach is more effective at propagating knowledge updates than fine-tuning and other gradient-based knowledge-editing methods. Moreover, it does not compromise performance in other contexts, even when injecting the definitions of up to 150 entities at once. Shankar Padmanabhan, Yasumasa Onoe, Michael J. Q. Zhang, Greg Durrett, Eunsol Choi |
NeurIPS | 5 |
| 2023 | Understanding Postpartum Parents' Experiences via Two Digital PlatformsabstractDigital platforms, including online forums and helplines, have emerged as avenues of support for caregivers suffering from postpartum mental health distress. Understanding support seekers' experiences as shared on these platforms could provide crucial insight into caregivers' needs during this vulnerable time. In the current work, we provide a descriptive analysis of the concerns, psychological states, and motivations shared by healthy and distressed postpartum support seekers on two digital platforms, a one-on-one digital helpline and a publicly available online forum. Using a combination of human annotations, dictionary models and unsupervised techniques, we find stark differences between the experiences of distressed and healthy mothers. Distressed mothers described interpersonal problems and a lack of support, with 8.60% - 14.56% reporting severe symptoms including suicidal ideation. In contrast, the majority of healthy mothers described childcare issues, such as questions about breastfeeding or sleeping, and reported no severe mental health concerns. Across the two digital platforms, we found that distressed mothers shared similar content. However, the patterns of speech and affect shared by distressed mothers differed between the helpline vs. the online forum, suggesting the design of these platforms may shape meaningful measures of their support-seeking experiences. Our results provide new insight into the experiences of caregivers suffering from postpartum mental health distress. We conclude by discussing methodological considerations for understanding content shared by support seekers and design considerations for the next generation of support tools for postpartum parents. Xuewen Yao, Miriam Mikhelson, Megan Micheletti, Eunsol Choi, S. Craig Watkins, Edison Thomaz, Kaya de Barbaro |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2022 | Misinfo Reaction Frames: Reasoning about Readers' Reactions to News HeadlinesabstractSaadia Gabriel, Skyler Hallinan, Maarten Sap, Pemi Nguyen, Franziska Roesner, Eunsol Choi, Yejin Choi. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Saadia Gabriel, Skyler Hallinan, Maarten Sap, Pemi Nguyen, Franziska Roesner, Eunsol Choi, Yejin Choi 0001 |
ACL (1) | 6 |
| 2022 | Simulating Bandit Learning from User Feedback for Extractive Question AnsweringabstractWe study learning from user feedback for extractive question answering by simulating feedback using supervised data.We cast the problem as contextual bandit learning, and analyze the characteristics of several learning scenarios with focus on reducing data annotation.We show that systems initially trained on a small number of examples can dramatically improve given feedback from users on modelpredicted answers, and that one can use existing datasets to deploy systems in new domains without any annotation, but instead improving the system on-the-fly via user feedback. Eunsol Choi, Yoav Artzi |
ACL (1) | 2 |
| 2022 | How Do We Answer Complex Questions: Discourse Structure of Long-form AnswersabstractLong-form answers, consisting of multiple sentences, can provide nuanced and comprehensive answers to a broader set of questions.To better understand this complex and understudied task, we study the functional structure of long-form answers collected from three datasets, ELI5 (Fan et al., 2019), We-bGPT (Nakano et al., 2021) and Natural Questions (Kwiatkowski et al., 2019).Our main goal is to understand how humans organize information to craft complex answers.We develop an ontology of six sentence-level functional roles for long-form answers, and annotate 3.9k sentences in 640 answer paragraphs.Different answer collection methods manifest in different discourse structures.We further analyze model-generated answers -finding that annotators agree less with each other when annotating model-generated answers compared to annotating human-written answers.Our annotated data enables training a strong classifier that can be used for automatic analysis.We hope our work can inspire future research on discourselevel modeling and evaluation of long-form QA systems. 1 Junyi Jessy Li, Eunsol Choi |
ACL (1) | 3 |
| 2022 | Generating Literal and Implied Subquestions to Fact-check Complex ClaimsabstractVerifying political claims is a challenging task, as politicians can use various tactics to subtly misrepresent the facts for their agenda.Existing automatic fact-checking systems fall short here, and their predictions like "half-true" are not very useful in isolation, since it is unclear which parts of a claim are true or false.In this work, we focus on decomposing a complex claim into a comprehensive set of yes-no subquestions whose answers influence the veracity of the claim.We present CLAIMDECOMP, a dataset of decompositions for over 1000 claims.Given a claim and its verification paragraph written by fact-checkers, our trained annotators write subquestions covering both explicit propositions of the original claim and its implicit facets, such as additional political context that changes our view of the claim's veracity.We study whether state-of-the-art pre-trained models can learn to generate such subquestions.Our experiments show that these models generate reasonable questions, but predicting implied subquestions based only on the claim (without consulting other evidence) remains challenging.Nevertheless, we show that predicted subquestions can help identify relevant evidence to fact-check the full claim and derive the veracity through their answers, suggesting that claim decomposition can be a useful piece of a fact-checking pipeline.1 1 We release our code and dataset: https://jifan-chen. github.io/ClaimDecompJoe Biden stated on August 31, 2020 in a speech: "When I was vice president, violent crime fell 15% in this country.... The murder rate now is up 26% across the nation this year under Donald Trump."Claim Decomposi-on: focus of this work Claim Q1: Did the crime rate fall by 15% during Joe Biden's presidency?Q2: Did the murder rate in 2020 increase by 26% from 2019?Q3: Is Biden comparing crime rates from the same time interval in his statement?Q4: Is violent crime rate and murder rate directly comparable?Literal Implied Jifan Chen, Aniruddh Sriram, Eunsol Choi, Greg Durrett |
EMNLP | 3 |
| 2022 | Rich Knowledge Sources Bring Complex Knowledge Conflicts: Recalibrating Models to Reflect Conflicting EvidenceabstractQuestion answering models can use rich knowledge sources -up to one hundred retrieved passages and parametric knowledge in the large-scale language model (LM).Prior work assumes information in such knowledge sources is consistent with each other, paying little attention to how models blend information stored in their LM parameters with that from retrieved evidence documents.In this paper, we simulate knowledge conflicts (i.e., where parametric knowledge suggests one answer and different passages suggest different answers) and examine model behaviors.We find retrieval performance heavily impacts which sources models rely on, and current models mostly rely on non-parametric knowledge in their best-performing settings.We discover a troubling trend that contradictions among knowledge sources affect model confidence only marginally.To address this issue, we present a new calibration study, where models are discouraged from presenting any single answer when presented with multiple conflicting answer candidates in retrieved evidences. Hung-Ting Chen, Michael J. Q. Zhang, Eunsol Choi |
EMNLP | 3 |
| 2022 | Why is Winoground Hard? Investigating Failures in Visuolinguistic CompositionalityabstractRecent visuolinguistic pre-trained models show promising progress on various end tasks such as image retrieval and video captioning.Yet, they fail miserably on the recently proposed Winoground dataset (Thrush et al., 2022), which challenges models to match paired images and English captions, with items constructed to overlap lexically but differ in meaning (e.g., "there is a mug in some grass" vs. "there is some grass in a mug").By annotating the dataset using new fine-grained tags, we show that solving the Winoground task requires not just compositional language understanding, but a host of other abilities like commonsense reasoning or locating small, out-of-focus objects in low-resolution images.In this paper, we identify the dataset's main challenges through a suite of experiments on related tasks (probing task, image retrieval task), data augmentation, and manual inspection of the dataset.Our analysis suggests that a main challenge in visuolinguistic models may lie in fusing visual and textual representations, rather than in compositional language understanding.We release our annotation and code at https://github. com/ajd12342/why-winoground-hard. Anuj Diwan, Layne Berry, Eunsol Choi, David F. Harwath, Kyle Mahowald |
EMNLP | 3 |
| 2022 | Modeling Exemplification in Long-form Question Answering via RetrievalabstractShufan Wang, Fangyuan Xu, Laure Thompson, Eunsol Choi, Mohit Iyyer. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Shufan Wang, Laure Thompson, Eunsol Choi, Mohit Iyyer |
NAACL-HLT | 4 |
| 2021 | Challenges in Information-Seeking QA: Unanswerable Questions and Paragraph RetrievalabstractAkari Asai, Eunsol Choi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Akari Asai, Eunsol Choi |
ACL/IJCNLP (1) | 2 |
| 2021 | SituatedQA: Incorporating Extra-Linguistic Contexts into QAabstractAnswers to the same question may change depending on the extra-linguistic contexts (when and where the question was asked).To study this challenge, we introduce SITUATEDQA, an open-retrieval QA dataset where systems must produce the correct answer to a question given the temporal or geographical context.To construct SITUATEDQA, we first identify such questions in existing QA datasets.We find that a significant proportion of information seeking questions have context-dependent answers (e.g.roughly 16.5% of NQ-Open).For such context-dependent questions, we then crowdsource alternative contexts and their corresponding answers.Our study shows that existing models struggle with producing answers that are frequently updated or from uncommon locations.We further quantify how existing models, which are trained on data collected in the past, fail to generalize to answering questions asked in the present, even when provided with an updated evidence corpus (a roughly 15 point drop in accuracy).Our analysis suggests that open-retrieval QA benchmarks should incorporate extra-linguistic context to stay relevant globally and in the future.Our data, code, and datasheet are available at https: //situatedqa.github.io/. Michael J. Q. Zhang, Eunsol Choi |
EMNLP (1) | 2 |
| 2021 | Learning with Different Amounts of Annotation: From Zero to Many LabelsabstractTraining NLP systems typically assumes access to annotated data that has a single human label per example.Given imperfect labeling from annotators and inherent ambiguity of language, we hypothesize that single label is not sufficient to learn the spectrum of language interpretation.We explore new annotation distribution schemes, assigning multiple labels per example for a small subset of training examples.Introducing such multi label examples at the cost of annotating fewer examples brings clear gains on natural language inference task and entity typing task, even when we simply first train with a single label data and then fine tune with multi label examples.Extending a MixUp data augmentation framework, we propose a learning algorithm that can learn from training examples with different amount of annotation (with zero, one, or multiple labels).This algorithm efficiently combines signals from uneven training data and brings additional gains in low annotation budget and cross domain settings.Together, our method achieves consistent gains in two tasks, suggesting distributing labels unevenly among training examples can be beneficial for many NLP tasks. 1 Shujian Zhang, Chengyue Gong, Eunsol Choi |
EMNLP (1) | 3 |
| 2021 | What Happens to My Instagram Account After I Die? Re-imagining Social Media as a Commemorative Space for Remembrance and Recovery
Soonho Kwon, Eunsol Choi, Sunah Hwang, Younah Kang |
INTERACT (2) | 2 |
| 2021 | XOR QA: Cross-lingual Open-Retrieval Question AnsweringabstractAkari Asai, Jungo Kasai, Jonathan Clark, Kenton Lee, Eunsol Choi, Hannaneh Hajishirzi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Akari Asai, Jungo Kasai, Jonathan H. Clark, Kenton Lee, Eunsol Choi, Hannaneh Hajishirzi |
NAACL-HLT | 5 |
| 2021 | Decontextualization: Making Sentences Stand-AloneabstractAbstract Models for question answering, dialogue agents, and summarization often interpret the meaning of a sentence in a rich context and use that meaning in a new context. Taking excerpts of text can be problematic, as key pieces may not be explicit in a local window. We isolate and define the problem of sentence decontextualization: taking a sentence together with its context and rewriting it to be interpretable out of context, while preserving its meaning. We describe an annotation procedure, collect data on the Wikipedia corpus, and use the data to train models to automatically decontextualize sentences. We present preliminary studies that show the value of sentence decontextualization in a user-facing task, and as preprocessing for systems that perform document understanding. We argue that decontextualization is an important subtask in many downstream applications, and that the definitions and resources provided can benefit tasks that operate on sentences that occur in a richer context. Eunsol Choi, Jennimaria Palomaki, Matthew Lamm, Tom Kwiatkowski, Dipanjan Das 0001, Michael Collins 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | QED: A Framework and Dataset for Explanations in Question AnsweringabstractA question answering system that in addition to providing an answer provides an explanation of the reasoning that leads to that answer has potential advantages in terms of debuggability, extensibility, and trust. To this end, we propose QED, a linguistically informed, extensible framework for explanations in question answering. A QED explanation specifies the relationship between a question and answer according to formal semantic notions such as referential equality, sentencehood, and entailment. We describe and publicly release an expert-annotated dataset of QED explanations built upon a subset of the Google Natural Questions dataset, and report baseline models on two tasks—post- hoc explanation generation given an answer, and joint question answering and explanation generation. In the joint setting, a promising result suggests that training on a relatively small amount of QED data can improve question answering. In addition to describing the formal, language-theoretic motivations for the QED approach, we describe a large user study showing that the presence of QED explanations significantly improves the ability of untrained raters to spot errors made by a strong neural QA baseline. Matthew Lamm, Jennimaria Palomaki, Christopher Alberti, Daniel Andor, Eunsol Choi, Livio B. Soares, Michael Collins 0001 |
Trans. Assoc. Comput. Linguistics | 5 |
| 2020 | Entities as Experts: Sparse Memory Access with Entity SupervisionabstractWe focus on the problem of capturing declarative knowledge about entities in the learned parameters of a language model.We introduce a new model-Entities as Experts (EAE)that can access distinct memories of the entities mentioned in a piece of text.Unlike previous efforts to integrate entity knowledge into sequence models, EAE's entity representations are learned directly from text.We show that EAE's learned representations capture sufficient knowledge to answer TriviaQA questions such as "Which Dr. Who villain has been played by Roger Delgado, Anthony Ainley, Eric Roberts?", outperforming an encodergenerator Transformer model with 10× the parameters.According to the LAMA knowledge probes, EAE contains more factual knowledge than a similarly sized BERT, as well as previous approaches that integrate external sources of entity knowledge.Because EAE associates parameters with specific entities, it only needs to access a fraction of its parameters at inference time, and we show that the correct identification and representation of entities is essential to EAE's performance. * Work done during Google AI residency. Thibault Févry, Livio B. Soares, Nicholas FitzGerald, Eunsol Choi, Tom Kwiatkowski |
EMNLP (1) | 4 |
| 2020 | TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse LanguagesabstractConfidently making progress on multilingual modeling requires challenging, trustworthy evaluations. We present TyDi QA—a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs. The languages of TyDi QA are diverse with regard to their typology—the set of linguistic features each language expresses—such that we expect models performing well on this set to generalize across a large number of the world’s languages. We present a quantitative analysis of the data quality and example-level qualitative linguistic analyses of observed language phenomena that would not be found in English-only corpora. To provide a realistic information-seeking task and avoid priming effects, questions are written by people who want to know the answer, but don’t know the answer yet, and the data is collected directly in each language without the use of translation. Jonathan H. Clark, Jennimaria Palomaki, Vitaly Nikolaev, Eunsol Choi, Dan Garrette, Michael Collins 0001, Tom Kwiatkowski |
Trans. Assoc. Comput. Linguistics | 4 |
| 2019 | FlowQA: Grasping Flow in History for Conversational Machine Comprehension
Hsin-Yuan Huang, Eunsol Choi, Scott Yih |
ICLR (Poster) | 2 |
| 2018 | Ultra-Fine Entity TypingabstractWe introduce a new entity typing task: given a sentence with an entity mention, the goal is to predict a set of free-form phrases (e.g.skyscraper, songwriter, or criminal) that describe appropriate types for the target entity.This formulation allows us to use a new type of distant supervision at large scale: head words, which indicate the type of the noun phrases they appear in.We show that these ultra-fine types can be crowd-sourced, and introduce new evaluation sets that are much more diverse and fine-grained than existing benchmarks.We present a model that can predict open types, and is trained using a multitask objective that pools our new head-word supervision with prior supervision from entity linking.Experimental results demonstrate that our model is effective in predicting entity types at varying granularity; it achieves state of the art performance on an existing fine-grained entity typing benchmark, and sets baselines for our newly-introduced datasets.1 Eunsol Choi, Omer Levy, Yejin Choi 0001, Luke Zettlemoyer |
ACL (1) | 1 |
| 2018 | QuAC: Question Answering in ContextabstractWe present QuAC, a dataset for Question Answering in Context that contains 14K information-seeking QA dialogs (100K questions in total).The dialogs involve two crowd workers: (1) a student who poses a sequence of freeform questions to learn as much as possible about a hidden Wikipedia text, and (2) a teacher who answers the questions by providing short excerpts from the text.QuAC introduces challenges not found in existing machine comprehension datasets: its questions are often more open-ended, unanswerable, or only meaningful within the dialog context, as we show in a detailed qualitative evaluation.We also report results for a number of reference models, including a recently state-ofthe-art reading comprehension architecture extended to model dialog context.Our best model underperforms humans by 20 F1, suggesting that there is significant room for future work on this data.Dataset, baseline, and leaderboard available at http://quac.ai.How was perversion handled?How long was he there?How popular did she become?How did Mark Felt contact Woodword?How did the meeting go?How did it do on the charts?When was she born?When was it founded?When was the breakup? Eunsol Choi, He He 0001, Mohit Iyyer, Mark Yatskar, Scott Yih, Yejin Choi 0001, Percy Liang, Luke Zettlemoyer |
EMNLP | 1 |
| 2018 | Neural Metaphor Detection in ContextabstractWe present end-to-end neural models for detecting metaphorical word use in context.We show that relatively standard BiLSTM models which operate on complete sentences work well in this setting, in comparison to previous work that used more restricted forms of linguistic context.These models establish a new state-of-the-art on existing verb metaphor detection benchmarks, and show strong performance on jointly predicting the metaphoricity of all words in a running text. Eunsol Choi, Yejin Choi 0001, Luke Zettlemoyer |
EMNLP | 2 |
| 2017 | Coarse-to-Fine Question Answering for Long DocumentsabstractEunsol Choi, Daniel Hewlett, Jakob Uszkoreit, Illia Polosukhin, Alexandre Lacoste, Jonathan Berant. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Eunsol Choi, Daniel Hewlett, Jakob Uszkoreit, Illia Polosukhin, Alexandre Lacoste, Jonathan Berant |
ACL (1) | 1 |
| 2017 | TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading ComprehensionabstractWe present TriviaQA, a challenging reading comprehension dataset containing over 650K question-answer-evidence triples.TriviaQA includes 95K questionanswer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision for answering the questions.We show that, in comparison to other recently introduced large-scale datasets, TriviaQA (1) has relatively complex, compositional questions, (2) has considerable syntactic and lexical variability between questions and corresponding answer-evidence sentences, and (3) requires more cross sentence reasoning to find answers.We also present two baseline algorithms: a featurebased classifier and a state-of-the-art neural network, that performs well on SQuAD reading comprehension.Neither approach comes close to human performance (23% and 40% vs. 80%), suggesting that Trivi-aQA is a challenging testbed that is worth significant future study. 1 Mandar Joshi, Eunsol Choi, Daniel S. Weld, Luke Zettlemoyer |
ACL (1) | 2 |
| 2017 | Zero-Shot Relation Extraction via Reading ComprehensionabstractWe show that relation extraction can be reduced to answering simple reading comprehension questions, by associating one or more natural-language questions with each relation slot.This reduction has several advantages: we can (1) learn relationextraction models by extending recent neural reading-comprehension techniques, (2) build very large training sets for those models by combining relation-specific crowd-sourced questions with distant supervision, and even (3) do zero-shot learning by extracting new relation types that are only specified at test-time, for which we have no labeled training examples.Experiments on a Wikipedia slot-filling task demonstrate that the approach can generalize to new questions for known relation types with high accuracy, and that zero-shot generalization to unseen relation types is possible, at lower accuracy levels, setting the bar for future work on this task. Omer Levy, Minjoon Seo, Eunsol Choi, Luke Zettlemoyer |
CoNLL | 3 |
| 2017 | Truth of Varying Shades: Analyzing Language in Fake News and Political Fact-CheckingabstractWe present an analytic study on the language of news media in the context of political fact-checking and fake news detection.We compare the language of real news with that of satire, hoaxes, and propaganda to find linguistic characteristics of untrustworthy text.To probe the feasibility of automatic political fact-checking, we also present a case study based on PolitiFact.com using their factuality judgments on a 6-point scale.Experiments show that while media fact-checking remains to be an open research question, stylistic cues can help determine the truthfulness of text. Hannah Rashkin, Eunsol Choi, Jin Yea Jang, Svitlana Volkova, Yejin Choi 0001 |
EMNLP | 2 |
| 2016 | Document-level Sentiment Inference with Social, Faction, and Discourse ContextabstractWe present a new approach for documentlevel sentiment inference, where the goal is to predict directed opinions (who feels positively or negatively towards whom) for all entities mentioned in a text.To encourage more complete and consistent predictions, we introduce an ILP that jointly models (1) sentence-and discourse-level sentiment cues, (2) factual evidence about entity factions, and (3) global constraints based on social science theories such as homophily, social balance, and reciprocity.Together, these cues allow for rich inference across groups of entities, including for example that CEOs and the companies they lead are likely to have similar sentiment towards others.We evaluate performance on new, densely labeled data that provides supervision for all pairs, complementing previous work that only labeled pairs mentioned in the same sentence.Experiments demonstrate that the global model outperforms sentence-level baselines, by providing more coherent predictions across sets of related entities. Eunsol Choi, Hannah Rashkin, Luke Zettlemoyer, Yejin Choi 0001 |
ACL (1) | 1 |
| 2016 | Extracting Structured Scholarly Information from the Machine Translation Literature
Eunsol Choi, Matic Horvat, Jonathan May, Kevin Knight, Daniel Marcu |
LREC | 1 |
| 2015 | Scalable Semantic Parsing with Partial OntologiesabstractEunsol Choi, Tom Kwiatkowski, Luke Zettlemoyer. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Eunsol Choi, Tom Kwiatkowski, Luke Zettlemoyer |
ACL (1) | 1 |
| 2013 | Scaling Semantic Parsers with On-the-Fly Ontology MatchingabstractWe consider the challenge of learning semantic parsers that scale to large, open-domain problems, such as question answering with Freebase.In such settings, the sentences cover a wide variety of topics and include many phrases whose meaning is difficult to represent in a fixed target ontology.For example, even simple phrases such as 'daughter' and 'number of people living in' cannot be directly represented in Freebase, whose ontology instead encodes facts about gender, parenthood, and population.In this paper, we introduce a new semantic parsing approach that learns to resolve such ontological mismatches.The parser is learned from question-answer pairs, uses a probabilistic CCG to build linguistically motivated logicalform meaning representations, and includes an ontology matching model that adapts the output logical forms for each target ontology.Experiments demonstrate state-of-the-art performance on two benchmark semantic parsing datasets, including a nine point accuracy improvement on a recent Freebase QA corpus. Tom Kwiatkowski, Eunsol Choi, Yoav Artzi, Luke Zettlemoyer |
EMNLP | 2 |