EDBT 2026 Demo / reviewers in the wild / expert
Yulia Tsvetkov
dblp:75/8157
· DBLP profile ↗
104ranked-venue papers
13as first author
68since 2021 · last 2026
0000-0002-4634-7128ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 98 · 12 first-author · 64 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When One LLM Drools, Multi-LLM Collaboration RulesabstractShangbin Feng, Wenxuan Ding, Alisa Liu, Zifeng Wang, Weijia Shi, Yike Wang, Shannon Zejiang Shen, Xiaochuang Han, Hunter Lang, Chen-Yu Lee, Tomas Pfister, Yejin Choi, Yulia Tsvetkov. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shangbin Feng, Wenxuan Ding 0001, Alisa Liu, Zifeng Wang 0002, Yike Wang 0002, Shannon Shen 0001, Xiaochuang Han, Hunter Lang, Chen-Yu Lee, Tomas Pfister, Yejin Choi 0001, Yulia Tsvetkov |
ACL (1) | 13 |
| 2026 | Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration SystemsabstractLanguage models (LMs) are increasingly used in collaboration: multiple LMs trained by different parties collaborate through routing systems, multi-agent debate, model merging, and more.Critical safety risks remain in this decentralized paradigm: what if some of the models in multi-LLM systems are compromised or malicious?We first quantify the impact of malicious models by engineering four categories of malicious LMs, plug them into four types of popular model collaboration systems, and evaluate the compromised system across 10 datasets.We find that malicious models have a severe impact on the multi-LLM systems, especially for reasoning and safety domains where performance is lowered by 7.12% and 7.94% on average.We then propose mitigation strategies to alleviate the impact of malicious components, by employing external supervisors that oversee model collaboration to disable/mask them out to reduce their influence.On average, these strategies recover 95.31% of the initial performance, while making model collaboration systems fully resistant to malicious models remains an open research question. Wenxuan Ding 0001, Shangbin Feng, Yulia Tsvetkov |
ACL (1) | 4 |
| 2026 | Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language ModelsabstractThe output quality of large language models (LLMs) can be improved via “reasoning”: generating segments of chain-of-thought (CoT) content to further condition the model prior to producing user-facing output. While these chains contain valuable information, they are verbose and lack explicit organization, making them tedious to review. Moreover, they lack opportunities for user feedback, such as removing unwanted considerations, adding desired ones, or clarifying unclear assumptions. We introduce Interactive Reasoning, an interaction design that visualizes chain-of-thought outputs as a hierarchy of topics and enables user review and modification. We implement interactive reasoning in Hippo, a prototype for AI-assisted decision making in the face of uncertain trade-offs. In a user study with 16 participants, we find that interactive reasoning in Hippo allows users to quickly identify and interrupt erroneous generations, efficiently steer the model towards customized responses, and better understand both model reasoning and model outputs. Our work contributes to a new paradigm that incorporates user oversight into LLM reasoning processes. Rock Yuren Pang, K. J. Kevin Feng, Shangbin Feng, Chu Li 0001, Yulia Tsvetkov, Jeffrey Heer, Katharina Reinecke |
IUI | 6 |
| 2025 | CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-TeamingabstractYu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yu Ying Chiu, Bill Y. Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi 0001 |
ACL (1) | 9 |
| 2025 | Biased LLMs can Influence Political Decision-MakingabstractJillian Fisher, Shangbin Feng, Robert Aron, Thomas Richardson, Yejin Choi, Daniel W Fisher, Jennifer Pan, Yulia Tsvetkov, Katharina Reinecke. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jillian Fisher, Shangbin Feng, Robert Aron, Yejin Choi 0001, Daniel W. Fisher, Jennifer Pan, Yulia Tsvetkov, Katharina Reinecke |
ACL (1) | 8 |
| 2025 | Explore Theory of Mind: program-guided adversarial data generation for theory of mind reasoningabstractDo large language models (LLMs) have theory of mind? A plethora of papers and benchmarks have been introduced to evaluate if current models have been able to develop this key ability of social intelligence. However, all rely on limited datasets with simple patterns that can potentially lead to problematic blind spots in evaluation and an overestimation of model capabilities. We introduce ExploreToM, the first framework to allow large-scale generation of diverse and challenging theory of mind data for robust training and evaluation. Our approach leverages an A* search over a custom domain-specific language to produce complex story structures and novel, diverse, yet plausible scenarios to stress test the limits of LLMs. Our evaluation reveals that state-of-the-art LLMs, such as Llama-3.1-70B and GPT-4o, show accuracies as low as 0% and 9% on ExploreToM-generated data, highlighting the need for more robust theory of mind evaluation. As our generations are a conceptual superset of prior work, fine-tuning on our data yields a 27-point accuracy improvement on the classic ToMi benchmark (Le et al., 2019). ExploreToM also enables uncovering underlying skills and factors missing for models to show theory of mind, such as unreliable state tracking or data imbalances, which may contribute to models' poor performance on benchmarks. Melanie Sclar, Jane Dwivedi-Yu, Maryam Fazel-Zarandi, Yulia Tsvetkov, Yonatan Bisk, Yejin Choi 0001, Asli Celikyilmaz |
ICLR | 4 |
| 2025 | Varying Shades of Wrong: Aligning LLMs with Wrong Answers OnlyabstractIn the absence of abundant reliable annotations for challenging tasks and contexts, how can we expand the frontier of LLM capabilities with potentially wrong answers? We focus on two research questions: (1) Can LLMs generate reliable preferences among wrong options? And if so, (2) Would alignment with such wrong-over-wrong preferences be helpful? We employ methods based on self-consistency, token probabilities, and LLM-as-a-judge to elicit wrong-over-wrong preferences, and fine-tune language models with preference optimization approaches using these synthesized preferences. Extensive experiments with seven LLMs and eight datasets demonstrate that (1) LLMs do have preliminary capability in distinguishing various shades of wrong, achieving up to 20.9% higher performance than random guess; (2) Alignment with wrong-over-wrong preferences helps LLMs to produce less wrong and sometimes even outright correct answers, while improving overall model calibration. Code and data are publicly available at https://github.com/yaojh18/Varying-Shades-of-Wrong. Jihan Yao, Wenxuan Ding 0001, Shangbin Feng, Lucy Lu Wang, Yulia Tsvetkov |
ICLR | 5 |
| 2025 | Model Swarms: Collaborative Search to Adapt LLM Experts via Swarm IntelligenceabstractWe propose Model Swarms, a collaborative search algorithm to adapt LLMs via swarm intelligence, the collective behavior guiding individual systems. Specifically, Model Swarms starts with a pool of LLM experts and a utility function. Guided by the best-found checkpoints across models, diverse LLM experts collaboratively move in the weight space and optimize a utility function representing model adaptation objectives. Compared to existing model composition approaches, Model Swarms offers tuning-free model adaptation, works in low-data regimes with as few as 200 examples, and does not require assumptions about specific experts in the swarm or how they should be composed. Extensive experiments demonstrate that Model Swarms could flexibly adapt LLM experts to a single task, multi-task domains, reward models, as well as diverse human interests, improving over 12 model composition baselines by up to 21.0% across tasks and contexts. Further analysis reveals that LLM experts discover previously unseen capabilities in initial checkpoints and that Model Swarms enable the weak-to-strong transition of experts through the collaborative search process. Shangbin Feng, Zifeng Wang 0002, Yike Wang 0002, Sayna Ebrahimi, Hamid Palangi, Lesly Miculicich, Achin Kulshrestha, Nathalie Rauschmayr, Yejin Choi 0001, Yulia Tsvetkov, Chen-Yu Lee, Tomas Pfister |
ICML | 10 |
| 2025 | ALPACA AGAINST VICUNA: Using LLMs to Uncover Memorization of LLMsabstractAly M. Kassem, Omar Mahmoud, Niloofar Mireshghallah, Hyunwoo Kim, Yulia Tsvetkov, Yejin Choi, Sherif Saad, Santu Rana. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Aly M. Kassem, Omar Mahmoud 0001, Niloofar Mireshghallah, Hyunwoo Kim 0002, Yulia Tsvetkov, Yejin Choi 0001, Sherif Saad, Santu Rana |
NAACL (Long Papers) | 5 |
| 2025 | ComPO: Community Preferences for Language Model PersonalizationabstractSachin Kumar, Chan Young Park, Yulia Tsvetkov, Noah A. Smith, Hannaneh Hajishirzi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sachin Kumar 0009, Chan Young Park, Yulia Tsvetkov, Noah A. Smith, Hannaneh Hajishirzi |
NAACL (Long Papers) | 3 |
| 2025 | Heterogeneous Swarms: Jointly Optimizing Model Roles and Weights for Multi-LLM SystemsabstractWe propose Heterogeneous Swarms, an algorithm to design multi-LLM systems by jointly optimizing model roles and weights. We represent multi-LLM systems as directed acyclic graphs (DAGs) of LLMs with topological message passing for collaborative generation. Given a pool of LLM experts and a utility function, Heterogeneous Swarms employs two iterative steps: role-step and weight-step. For role-step, we interpret model roles as learning a DAG that specifies the flow of inputs and outputs between LLMs. Starting from a swarm of random continuous adjacency matrices, we decode them into discrete DAGs, call the LLMs in topological order, evaluate on the utility function (e.g. accuracy on a task), and optimize the adjacency matrices with particle swarm optimization based on the utility score. For weight-step, we assess the contribution of individual LLMs in the multi-LLM systems and optimize model weights with swarm intelligence. We propose JFK-score to quantify the individual contribution of each LLM in the best-found DAG of the role-step, then optimize model weights with particle swarm optimization based on the JFK-score. Experiments demonstrate that Heterogeneous Swarms outperforms 17 role- and/or weight-based baselines by 18.5% on average across 12 tasks. Further analysis reveals that Heterogeneous Swarms discovers multi-LLM systems with heterogeneous model roles and substantial collaborative gains, and benefits from the diversity of language models. Shangbin Feng, Zifeng Wang 0002, Palash Goyal, Yike Wang 0002, Huang Xia, Hamid Palangi, Luke Zettlemoyer, Yulia Tsvetkov, Chen-Yu Lee, Tomas Pfister |
NeurIPS | 9 |
| 2025 | Precise Information Control in Long-Form Text GenerationabstractA central challenge in language models (LMs) is faithfulness hallucination: the generation of information unsubstantiated by input context. To study this problem, we propose Precise Information Control (PIC), a new task formulation that requires models to generate long-form outputs grounded in a provided set of short self-contained statements, without adding any unsupported ones. PIC includes a full setting that tests a model’s ability to include exactly all input claims, and a partial setting that requires the model to selectively incorporate only relevant claims. We present PIC-Bench, a benchmark of eight long-form generation tasks (e.g., summarization, biography generation) adapted to the PIC setting, where LMs are supplied with well-formed, verifiable input claims. Our evaluation of a range of open and proprietary LMs on PIC-Bench reveals that, surprisingly, state-of-the-art LMs still hallucinate against user-provided input in over 70% of generations. To alleviate this lack of faithfulness, we introduce a post-training framework that uses a weakly supervised preference data construction method to train an 8B PIC-LM with stronger PIC ability—improving from 69.1% to 91.0% F1 in the full PIC setting. When integrated into end-to-end factual generation pipelines, PIC-LM improves exact match recall by 17.1% on ambiguous QA with retrieval, and factual precision by 30.5% on a birthplace fact-checking task, underscoring the potential of precisely grounded generation. Jacqueline He, Howard Yen, Margaret Li, Shuyue Stella Li, Yulia Tsvetkov, Danqi Chen 0001, Pang Wei W. Koh, Luke Zettlemoyer |
NeurIPS | 7 |
| 2025 | Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)abstractLarge language models (LMs) often struggle to generate diverse, human-like creative content, raising concerns about the long-term homogenization of human thought through repeated exposure to similar outputs. Yet scalable methods for evaluating LM output diversity remain limited, especially beyond narrow tasks such as random number or name generation, or beyond repeated sampling from a single model. To address this gap, we introduce Infinity-Chat, a large-scale dataset of 26K diverse, real-world, open-ended user queries that admit a wide range of plausible answers with no single ground truth. We introduce the first comprehensive taxonomy for characterizing the full spectrum of open-ended prompts posed to LMs, comprising 6 top-level categories (e.g., creative content generation, brainstorm & ideation) that further breaks down to 17 subcategories. Using Infinity-Chat, we present a large-scale study of mode collapse in LMs, revealing a pronounced Artificial Hivemind effect in open-ended generation of LMs, characterized by (1) intra-model repetition, where a single model consistently generates similar responses, and more so (2) inter-model homogeneity, where different models produce strikingly similar outputs. Infinity-Chat also includes 31,250 human annotations, across absolute ratings and pairwise preferences, with 25 independent human annotations per example. This enables studying collective and individual-specific human preferences in response to open-ended queries. Our findings show that state-of-the-art LMs, reward models, and LM judges are less well calibrated to human ratings on model generations that elicit differing idiosyncratic annotator preferences, despite maintaining comparable overall quality. Overall, INFINITY-CHAT presents the first large-scale resource for systematically studying real-world open-ended queries to LMs, revealing critical insights to guide future research for mitigating long-term AI safety risks posed by the Artificial Hivemind. Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yulia Tsvetkov, Maarten Sap, Yejin Choi 0001 |
NeurIPS | 7 |
| 2025 | Sparta Alignment: Collectively Aligning Multiple Language Models through CombatabstractWe propose Sparta Alignment, an algorithm to collectively align multiple LLMs through competition and combat. To complement a single model's lack of diversity in generation and biases in evaluation, multiple LLMs form a 'sparta tribe' to compete against each other in fulfilling instructions while serving as judges for the competition of others. For each iteration, one instruction and two models are selected for a duel, the other models evaluate the two responses, and their evaluation scores are aggregated through a adapted elo-ranking based reputation system, where winners/losers of combat gain/lose weight in evaluating others. The peer-evaluated combat results then become preference pairs where the winning response is preferred over the losing one, and all models learn from these preferences at the end of each iteration. Sparta Alignment enables the self-evolution of multiple LLMs in an iterative and collective competition process. Extensive experiments demonstrate that Sparta Alignment outperforms initial models and 4 self-alignment baselines across 10 out of 12 tasks and datasets with 7.0\% average improvement. Further analysis reveals that Sparta Alignment generalizes more effectively to unseen tasks and leverages the expertise diversity of participating models to produce more logical, direct and informative outputs. Yuru Jiang, Wenxuan Ding 0001, Shangbin Feng, Greg Durrett, Yulia Tsvetkov |
NeurIPS | 5 |
| 2025 | Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?abstractSpurious correlations occur when models rely on non-essential features that coincidentally co-vary with target labels, leading to incorrect reasoning under distribution shift. We consider spurious correlations in multi-modal Large Vision Language Models (LVLMs) pretrained on extensive and diverse datasets without explicit task supervision. We develop a benchmark by sourcing GPT-4o errors on real-world visual-question-answering (VQA) benchmarks, then curating a subset through LVLM-human annotation and synthetic counterfactual evaluation to identify errors caused by spurious correlations. This process yields SpuriVerse, a novel benchmark comprised of 124 distinct types of spurious correlations extracted from real-world datasets, each containing 1 realistic and 10 synthetic VQA samples for a total of 1364 multiple choice questions. We evaluate 15 open and closed-source LVLMs on SpuriVerse, finding that even state-of-the-art closed-source models struggle significantly, achieving at best only 35.0\% accuracy. Fine-tuning on synthetic examples that emphasize the spurious correlation improves performance to 78.4\%, suggesting that training on diverse spurious patterns generalizes to unseen situations: models appear to learn to avoid "shortcuts" and attend to the overall image context. Yiwei Yang 0009, Chung Peng Lee, Shangbin Feng, Dora Zhao, Bingbing Wen, Anthony Z. Liu, Yulia Tsvetkov, Bill Howe |
NeurIPS | 7 |
| 2025 | Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in TransformersabstractAbstract Transformers trained on natural language data have been shown to exhibit hierarchical generalization without explicitly encoding any structural bias. In this work, we investigate sources of inductive bias in transformer models and their training that could cause such preference for hierarchical generalization. We extensively experiment with transformers trained on five synthetic, controlled datasets using several training objectives and show that, while objectives such as sequence-to-sequence modeling, classification, etc., often fail to lead to hierarchical generalization, the language modeling objective consistently leads to transformers generalizing hierarchically. We then study how different generalization behaviors emerge during the training by conducting pruning experiments that reveal the joint existence of subnetworks within the model implementing different generalizations. Finally, we take a Bayesian perspective to understand transformers’ preference for hierarchical generalization: We establish a correlation between whether transformers generalize hierarchically on a dataset and if the simplest explanation of that dataset is provided by a hierarchical grammar compared to regular grammars exhibiting linear generalization. Overall, our work presents new insights on the origins of hierarchical generalization in transformers and provides a theoretical framework for studying generalization in language models. Kabir Ahuja, Vidhisha Balachandran, Madhur Panwar, Tianxing He, Noah A. Smith, Navin Goyal, Yulia Tsvetkov |
Trans. Assoc. Comput. Linguistics | 7 |
| 2025 | Know Your Limits: A Survey of Abstention in Large Language ModelsabstractAbstract Abstention, the refusal of large language models (LLMs) to provide an answer, is increasingly recognized for its potential to mitigate hallucinations and enhance safety in LLM systems. In this survey, we introduce a framework to examine abstention from three perspectives: the query, the model, and human values. We organize the literature on abstention methods, benchmarks, and evaluation metrics using this framework, and discuss merits and limitations of prior work. We further identify and motivate areas for future research, such as whether abstention can be achieved as a meta-capability that transcends specific tasks or domains, and opportunities to optimize abstention abilities in specific contexts. In doing so, we aim to broaden the scope and impact of abstention methodologies in AI systems.1 Bingbing Wen, Jihan Yao, Shangbin Feng, Chenjun Xu, Yulia Tsvetkov, Bill Howe, Lucy Lu Wang |
Trans. Assoc. Comput. Linguistics | 5 |
| 2024 | DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related LanguagesabstractFahim Faisal, Orevaoghene Ahia, Aarohi Srivastava, Kabir Ahuja, David Chiang, Yulia Tsvetkov, Antonios Anastasopoulos. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Fahim Faisal, Orevaoghene Ahia, Aarohi Srivastava, Kabir Ahuja, David Chiang 0001, Yulia Tsvetkov, Antonios Anastasopoulos |
ACL (1) | 6 |
| 2024 | Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM CollaborationabstractShangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding, Vidhisha Balachandran, Yulia Tsvetkov. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shangbin Feng, Yike Wang 0002, Wenxuan Ding 0001, Vidhisha Balachandran, Yulia Tsvetkov |
ACL (1) | 6 |
| 2024 | What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot DetectionabstractSocial media bot detection has always been an arms race between advancements in machine learning bot detectors and adversarial bot strategies to evade detection.In this work, we bring the arms race to the next level by investigating the opportunities and risks of state-of-the-art large language models (LLMs) in social bot detection.To investigate the opportunities, we design novel LLM-based bot detectors by proposing a mixture-of-heterogeneous-experts framework to divide and conquer diverse user information modalities.To illuminate the risks, we explore the possibility of LLM-guided manipulation of user textual and structured information to evade detection.Extensive experiments with three LLMs on two datasets demonstrate that instruction tuning on merely 1,000 annotated examples produces specialized LLMs that outperform state-of-the-art bot detection baselines by up to 9.1% on both datasets.On the other hand, LLM-guided manipulation strategies could significantly bring down the performance of existing bot detectors by up to 29.6% and harm the calibration and reliability of bot detection systems.Ultimately, this works identifies LLMs as the new frontier of social bot detection research.1 Shangbin Feng, Herun Wan, Ningnan Wang, Zhaoxuan Tan, Minnan Luo, Yulia Tsvetkov |
ACL (1) | 6 |
| 2024 | Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under AttacksabstractYichen Wang, Shangbin Feng, Abe Hou, Xiao Pu, Chao Shen, Xiaoming Liu, Yulia Tsvetkov, Tianxing He. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yichen Wang 0002, Shangbin Feng, Abe Bohan Hou, Xiao Pu 0003, Chao Shen 0001, Xiaoming Liu 0001, Yulia Tsvetkov, Tianxing He |
ACL (1) | 7 |
| 2024 | Voices Unheard: NLP Resources and Models for Yorùbá Regional DialectsabstractOrevaoghene Ahia, Anuoluwapo Aremu, Diana Abagyan, Hila Gonen, David Ifeoluwa Adelani, Daud Abolade, Noah A. Smith, Yulia Tsvetkov. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Orevaoghene Ahia, Aremu Anuoluwapo, Diana Abagyan, Hila Gonen, David Ifeoluwa Adelani, Daud Abolade, Noah A. Smith, Yulia Tsvetkov |
EMNLP | 8 |
| 2024 | Teaching LLMs to Abstain across Languages via Multilingual FeedbackabstractShangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding, Orevaoghene Ahia, Shuyue Stella Li, Vidhisha Balachandran, Sunayana Sitaram, Yulia Tsvetkov. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Shangbin Feng, Yike Wang 0002, Wenxuan Ding 0001, Orevaoghene Ahia, Shuyue Stella Li, Vidhisha Balachandran, Sunayana Sitaram, Yulia Tsvetkov |
EMNLP | 9 |
| 2024 | Modular Pluralism: Pluralistic Alignment via Multi-LLM CollaborationabstractWhile existing alignment paradigms have been integral in developing large language models (LLMs), LLMs often learn an averaged human preference and struggle to model diverse preferences across cultures, demographics, and communities.We propose MODULAR PLU-RALISM, a modular framework based on multi-LLM collaboration for pluralistic alignment: it "plugs into" a base LLM a pool of smaller but specialized community LMs, where models collaborate in distinct modes to flexibility support three modes of pluralism: Overton, steerable, and distributional (Sorensen et al., 2024b).MODULAR PLURALISM is uniquely compatible with black-box LLMs and offers the modular control of adding new community LMs for previously underrepresented communities.We evaluate MODULAR PLURAL-ISM with six tasks and four datasets featuring questions/instructions with value-laden and perspective-informed responses.Extensive experiments demonstrate that MODULAR PLU-RALISM advances the three pluralism objectives across six black-box and open-source LLMs.Further analysis reveals that LLMs are generally faithful to the inputs from smaller community LLMs, allowing seamless patching by adding a new community LM to better cover previously underrepresented communities.1 Shangbin Feng, Taylor Sorensen, Jillian Fisher, Chan Young Park, Yejin Choi 0001, Yulia Tsvetkov |
EMNLP | 7 |
| 2024 | Locating Information Gaps and Narrative Inconsistencies Across Languages: A Case Study of LGBT People Portrayals on WikipediaabstractTo explain social phenomena and identify systematic biases, much research in computational social science focuses on comparative text analyses.These studies often rely on coarse corpuslevel statistics or local word-level analyses, mainly in English.We introduce the INFOGAP method-an efficient and reliable approach to locating information gaps and inconsistencies in articles at the fact level, across languages.We evaluate INFOGAP by analyzing LGBT people's portrayals, across 2.7K biography pages on English, Russian, and French Wikipedias.We find large discrepancies in factual coverage across the languages.Moreover, our analysis reveals that biographical facts carrying negative connotations are more likely to be highlighted in Russian Wikipedia.Crucially, INFOGAP both facilitates large scale analyses, and pinpoints local document-and fact-level information gaps, laying a new foundation for targeted and nuanced comparative language analysis at scale. 1 Farhan Samir, Chan Young Park, Anjalie Field, Vered Shwartz, Yulia Tsvetkov |
EMNLP | 5 |
| 2024 | Gen-Z: Generative Zero-Shot Text Classification with Contextualized Label DescriptionsabstractLanguage model (LM) prompting—a popular paradigm for solving NLP tasks—has been shown to be susceptible to miscalibration and brittleness to slight prompt variations, caused by its discriminative prompting approach, i.e., predicting the label given the input. To address these issues, we propose Gen-Z—a generative prompting framework for zero-shot text classification. GEN-Z is generative, as it measures the LM likelihood of input text, conditioned on natural language descriptions of labels. The framework is multivariate, as label descriptions allow us to seamlessly integrate additional contextual information about the labels to improve task performance. On various standard classification benchmarks, with six open-source LM families, we show that zero-shot classification with simple contextualization of the data source of the evaluation set consistently outperforms both zero-shot and few-shot baselines while improving robustness to prompt variations. Further, our approach enables personalizing classification in a zero-shot manner by incorporating author, subject, or reader information in the label descriptions. Sachin Kumar 0009, Chan Young Park, Yulia Tsvetkov |
ICLR | 3 |
| 2024 | Knowledge Card: Filling LLMs' Knowledge Gaps with Plug-in Specialized Language ModelsabstractBy design, large language models (LLMs) are static general-purpose models, expensive to retrain or update frequently. As they are increasingly adopted for knowledge-intensive tasks, it becomes evident that these design choices lead to failures to generate factual, relevant, and up-to-date knowledge. To this end, we propose Knowledge Card, a modular framework to plug in new factual and relevant knowledge into general-purpose LLMs. We first introduce knowledge cards---specialized language models trained on corpora from specific domains and sources. Knowledge cards serve as parametric repositories that are selected at inference time to generate background knowledge for the base LLM. We then propose three content selectors to dynamically select and retain information in documents generated by knowledge cards, specifically controlling for relevance, brevity, and factuality of outputs. Finally, we propose two complementary integration approaches to augment the base LLM with the (relevant, factual) knowledge curated from the specialized LMs. Through extensive experiments, we demonstrate that Knowledge Card achieves state-of-the-art performance on six benchmark datasets. Ultimately, Knowledge Card framework enables dynamic synthesis and updates of knowledge from diverse domains. Its modularity will ensure that relevant knowledge can be continuously updated through the collective efforts of the research community. Shangbin Feng, Yuyang Bai, Vidhisha Balachandran, Tianxing He, Yulia Tsvetkov |
ICLR | 6 |
| 2024 | Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryabstractExisting efforts on quantifying privacy implications for large language models (LLMs) solely focus on measuring leakage of training data. In this work, we shed light on the often-overlooked interactive settings where an LLM receives information from multiple sources and generates an output to be shared with other entities, creating the potential of exposing sensitive input data in inappropriate contexts. In these scenarios, humans nat- urally uphold privacy by choosing whether or not to disclose information depending on the context. We ask the question “Can LLMs demonstrate an equivalent discernment and reasoning capability when considering privacy in context?” We propose CONFAIDE, a benchmark grounded in the theory of contextual integrity and designed to identify critical weaknesses in the privacy reasoning capabilities of instruction-tuned LLMs. CONFAIDE consists of four tiers, gradually increasing in complexity, with the final tier evaluating contextual privacy reasoning and theory of mind capabilities. Our experiments show that even commercial models such as GPT-4 and ChatGPT reveal private information in contexts that humans would not, 39% and 57% of the time, respectively, highlighting the urgent need for a new direction of privacy-preserving approaches as we demonstrate a larger underlying problem stemmed in the models’ lack of reasoning capabilities. Niloofar Mireshghallah, Hyunwoo Kim 0002, Yulia Tsvetkov, Maarten Sap, Reza Shokri, Yejin Choi 0001 |
ICLR | 4 |
| 2024 | Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingabstractAs large language models (LLMs) are adopted as a fundamental component of language technologies, it is crucial to accurately characterize their performance. Because choices in prompt design can strongly influence model behavior, this design process is critical in effectively using any modern pre-trained generative language model. In this work, we focus on LLM sensitivity to a quintessential class of meaning-preserving design choices: prompt formatting. We find that several widely used open-source LLMs are extremely sensitive to subtle changes in prompt formatting in few-shot settings, with performance differences of up to 76 accuracy points when evaluated using LLaMA-2-13B. Sensitivity remains even when increasing model size, the number of few-shot examples, or performing instruction tuning. Our analysis suggests that work evaluating LLMs with prompting-based methods would benefit from reporting a range of performance across plausible prompt formats, instead of the currently-standard practice of reporting performance on a single format. We also show that format performance only weakly correlates between models, which puts into question the methodological validity of comparing models with an arbitrarily chosen, fixed prompt format. To facilitate systematic analysis we propose FormatSpread, an algorithm that rapidly evaluates a sampled set of plausible prompt formats for a given task, and reports the interval of expected performance without accessing model weights. Furthermore, we present a suite of analyses that characterize the nature of this sensitivity, including exploring the influence of particular atomic perturbations and the internal representation of particular formats. Melanie Sclar, Yejin Choi 0001, Yulia Tsvetkov, Alane Suhr |
ICLR | 3 |
| 2024 | BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual TransferabstractAkari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Akari Asai, Sneha Reddy Kudugunta, Xinyan Yu 0001, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi |
NAACL-HLT | 7 |
| 2024 | David helps Goliath: Inference-Time Collaboration Between Small Specialized and Large General Diffusion LMsabstractXiaochuang Han, Sachin Kumar, Yulia Tsvetkov, Marjan Ghazvininejad. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xiaochuang Han, Sachin Kumar 0009, Yulia Tsvetkov, Marjan Ghazvininejad |
NAACL-HLT | 3 |
| 2024 | SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text GenerationabstractAbe Hou, Jingyu Zhang, Tianxing He, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, Yulia Tsvetkov. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Abe Bohan Hou, Tianxing He, Yichen Wang 0002, Yung-Sung Chuang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, Yulia Tsvetkov |
NAACL-HLT | 10 |
| 2024 | P³Sum: Preserving Author's Perspective in News Summarization with Diffusion Language ModelsabstractYuhan Liu, Shangbin Feng, Xiaochuang Han, Vidhisha Balachandran, Chan Young Park, Sachin Kumar, Yulia Tsvetkov. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Shangbin Feng, Xiaochuang Han, Vidhisha Balachandran, Chan Young Park, Sachin Kumar 0009, Yulia Tsvetkov |
NAACL-HLT | 7 |
| 2024 | MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based TokenizationabstractIn multilingual settings, non-Latin scripts and low-resource languages are usually disadvantaged in terms of language models’ utility, efficiency, and cost. Specifically, previous studies have reported multiple modeling biases that the current tokenization algorithms introduce to non-Latin script languages, the main one being over-segmentation. In this work, we propose MAGNET— multilingual adaptive gradient-based tokenization—to reduce over-segmentation via adaptive gradient-based subword tokenization. MAGNET learns to predict segment boundaries between byte tokens in a sequence via sub-modules within the model, which act as internal boundary predictors (tokenizers). Previous gradient-based tokenization methods aimed for uniform compression across sequences by integrating a single boundary predictor during training and optimizing it end-to-end through stochastic reparameterization alongside the next token prediction objective. However, this approach still results in over-segmentation for non-Latin script languages in multilingual settings. In contrast, MAGNET offers a customizable architecture where byte-level sequences are routed through language-script-specific predictors, each optimized for its respective language script. This modularity enforces equitable segmentation granularity across different language scripts compared to previous methods. Through extensive experiments, we demonstrate that in addition to reducing segmentation disparities, MAGNET also enables faster language modeling and improves downstream utility. Orevaoghene Ahia, Sachin Kumar 0009, Hila Gonen, Valentin Hofmann, Tomasz Limisiewicz, Yulia Tsvetkov, Noah A. Smith |
NeurIPS | 6 |
| 2024 | The Art of Saying No: Contextual Noncompliance in Language ModelsabstractChat-based language models are designed to be helpful, yet they should not comply with every user request. While most existing work primarily focuses on refusal of ``unsafe'' queries, we posit that the scope of noncompliance should be broadened. We introduce a comprehensive taxonomy of contextual noncompliance describing when and how models should not comply with user requests. Our taxonomy spans a wide range of categories including incomplete, unsupported, indeterminate, and humanizing requests (in addition to unsafe requests). To test noncompliance capabilities of language models, we use this taxonomy to develop a new evaluation suite of 1000 noncompliance prompts. We find that most existing models show significantly high compliance rates in certain previously understudied categories with models like GPT-4 incorrectly complying with as many as 30\% of requests.To address these gaps, we explore different training strategies using a synthetically-generated training set of requests and expected noncompliant responses. Our experiments demonstrate that while direct finetuning of instruction-tuned models can lead to both over-refusal and a decline in general capabilities, using parameter efficient methods like low rank adapters helps to strike a good balance between appropriate noncompliance and other capabilities. Faeze Brahman, Sachin Kumar 0009, Vidhisha Balachandran, Pradeep Dasigi, Valentina Pyatkin, Abhilasha Ravichander, Sarah Wiegreffe, Nouha Dziri, Khyathi Raghavi Chandu, Jack Hessel, Yulia Tsvetkov, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi |
NeurIPS | 11 |
| 2024 | MatFormer: Nested Transformer for Elastic InferenceabstractFoundation models are applied in a broad spectrum of settings with different inference constraints, from massive multi-accelerator clusters to resource-constrained standalone mobile devices. However, the substantial costs associated with training these models often limit the number of unique model sizes that can be offered. Consequently, practitioners are compelled to select a model that may not be optimally aligned with their specific latency and cost requirements. We present MatFormer, a novel Transformer architecture designed to provide elastic inference across diverse deployment constraints. MatFormer achieves this by incorporating a nested Feed Forward Network (FFN) block structure within a standard Transformer model. During training, we optimize the parameters of multiple nested FFN blocks with varying sizes, enabling the extraction of hundreds of accurate smaller models without incurring additional computational costs. We empirically validate the efficacy of MatFormer across different model classes (decoders and encoders) and modalities (language and vision), demonstrating its potential for real-world deployment. We show that a 850M decoder-only MatFormer language model (MatLM) allows us to extract multiple smaller models spanning from 582M to 850M parameters, each exhibiting better validation loss and one-shot downstream evaluations than independently trained counterparts. Furthermore, we observe that smaller encoders extracted from a universal MatFormer-based ViT (MatViT) encoder preserve the metric-space structure for adaptive large-scale retrieval. Finally, we showcase that speculative decoding with the accurate and consistent submodels extracted from MatFormer can lead to significant reduction in inference latency. Devvrit, Sneha Reddy Kudugunta, Aditya Kusupati, Tim Dettmers, Kaifeng Chen, Inderjit S. Dhillon, Yulia Tsvetkov, Hannaneh Hajishirzi, Sham M. Kakade, Ali Farhadi, Prateek Jain 0002 |
NeurIPS | 7 |
| 2024 | MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningabstractUsers typically engage with LLMs interactively, yet most existing benchmarks evaluate them in a static, single-turn format, posing reliability concerns in interactive scenarios. We identify a key obstacle towards reliability: LLMs are trained to answer any question, even with incomplete context or insufficient knowledge. In this paper, we propose to change the static paradigm to an interactive one, develop systems that proactively ask questions to gather more information and respond reliably, and introduce an benchmark—MEDIQ—to evaluate question-asking ability in LLMs. MEDIQ simulates clinical interactions consisting of a Patient System and an adaptive Expert System; with potentially incomplete initial information, the Expert refrains from making diagnostic decisions when unconfident, and instead elicits missing details via follow-up questions. We provide a pipeline to convert single-turn medical benchmarks into an interactive format. Our results show that directly prompting state-of-the-art LLMs to ask questions degrades performance, indicating that adapting LLMs to proactive information-seeking settings is nontrivial. We experiment with abstention strategies to better estimate model confidence and decide when to ask questions, improving diagnostic accuracy by 22.3%; however, performance still lags compared to an (unrealistic in practice) upper bound with complete information upfront. Further analyses show improved interactive performance with filtering irrelevant contexts and reformatting conversations. Overall, we introduce a novel problem towards LLM reliability, an interactive MEDIQ benchmark and a novel question-asking system, and highlight directions to extend LLMs’ information-seeking abilities in critical domains. Shuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen, Emma Pierson, Pang Wei W. Koh, Yulia Tsvetkov |
NeurIPS | 7 |
| 2024 | KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language ModelsabstractLarge language models (LLMs) demonstrate remarkable performance on knowledge-intensive tasks, suggesting that real-world knowledge is encoded in their model parameters. However, besides explorations on a few probing tasks in limited knowledge domains, it is not well understood how to evaluate LLMs' knowledge systematically and how well their knowledge abilities generalize, across a spectrum of knowledge domains and progressively complex task formats. To this end, we propose KGQuiz, a knowledge-intensive benchmark to comprehensively investigate the knowledge generalization abilities of LLMs. KGQuiz is a scalable framework constructed from triplet-based knowledge, which covers three knowledge domains and consists of five tasks with increasing complexity: true-or-false, multiple-choice QA, blank filling, factual editing, and open-ended knowledge generation. To gain a better understanding of LLMs' knowledge abilities and their generalization, we evaluate 10 open-source and black-box LLMs on the KGQuiz benchmark across the five knowledge-intensive tasks and knowledge domains. Extensive experiments demonstrate that LLMs achieve impressive performance in straightforward knowledge QA tasks, while settings and contexts requiring more complex reasoning or employing domain-specific facts still present significant challenges. We envision KGQuiz as a testbed to analyze such nuanced variations in performance across domains and task formats, and ultimately to understand, evaluate, and improve LLMs' knowledge abilities across a wide spectrum of knowledge domains and tasks. Yuyang Bai, Shangbin Feng, Vidhisha Balachandran, Zhaoxuan Tan, Shiqi Lou, Tianxing He, Yulia Tsvetkov |
WWW | 7 |
| 2023 | From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsabstractLanguage models (LMs) are pretrained on diverse data sources, including news, discussion forums, books, and online encyclopedias.A significant portion of this data includes opinions and perspectives which, on one hand, celebrate democracy and diversity of ideas, and on the other hand are inherently socially biased.Our work develops new methods to (1) measure political biases in LMs trained on such corpora, along social and economic axes, and (2) measure the fairness of downstream NLP models trained on top of politically biased LMs.We focus on hate speech and misinformation detection, aiming to empirically quantify the effects of political (social, economic) biases in pretraining data on the fairness of high-stakes social-oriented tasks.Our findings reveal that pretrained LMs do have political leanings that reinforce the polarization present in pretraining corpora, propagating social biases into hate speech predictions and misinformation detectors.We discuss the implications of our findings for NLP research and propose future directions to mitigate unfairness. 1 Warning: This paper contains examples of hate speech. Shangbin Feng, Chan Young Park, Yulia Tsvetkov |
ACL (1) | 4 |
| 2023 | KALM: Knowledge-Aware Integration of Local, Document, and Global Contexts for Long Document UnderstandingabstractWith the advent of pretrained language models (LMs), increasing research efforts have been focusing on infusing commonsense and domain-specific knowledge to prepare LMs for downstream tasks.These works attempt to leverage knowledge graphs, the de facto standard of symbolic knowledge representation, along with pretrained LMs.While existing approaches have leveraged external knowledge, it remains an open question how to jointly incorporate knowledge graphs representing varying contexts-from local (e.g., sentence), to document-level, to global knowledge-to enable knowledge-rich exchange across these contexts.Such rich contextualization can be especially beneficial for long document understanding tasks since standard pretrained LMs are typically bounded by the input sequence length.In light of these challenges, we propose KALM, a Knowledge-Aware Language Model that jointly leverages knowledge in local, document-level, and global contexts for long document understanding.KALM first encodes long documents and knowledge graphs into the three knowledge-aware context representations.It then processes each context with context-specific layers, followed by a "context fusion" layer that facilitates knowledge exchange to derive an overarching document representation.Extensive experiments demonstrate that KALM achieves state-of-the-art performance on six long document understanding tasks and datasets.Further analyses reveal that the three knowledge-aware contexts are complementary and they all contribute to model performance, while the importance and information exchange patterns of different contexts vary with respect to different tasks and datasets. Shangbin Feng, Zhaoxuan Tan, Zhenyu Lei 0004, Yulia Tsvetkov |
ACL (1) | 5 |
| 2023 | SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular ControlabstractDespite the growing success of diffusion models in continuous-valued domains (e.g., images), similar efforts for discrete domains such as text have yet to match the performance of autoregressive language models.In this work, we present SSD-LM-a diffusion-based language model with two key design choices.First, SSD-LM is semi-autoregressive, iteratively generating blocks of text, allowing for flexible output length at decoding time while enabling local bidirectional context updates.Second, it is simplex-based, performing diffusion on the natural vocabulary space rather than a learned latent space, allowing us to incorporate classifier guidance and modular control using offthe-shelf classifiers without any adaptation.We evaluate SSD-LM on unconstrained text generation benchmarks, and show that it matches or outperforms strong autoregressive GPT-2 models across standard quality and diversity metrics, while vastly outperforming diffusionbased baselines.On controlled text generation, SSD-LM also outperforms competitive baselines, with an extra advantage in modularity. 1 Xiaochuang Han, Sachin Kumar 0009, Yulia Tsvetkov |
ACL (1) | 3 |
| 2023 | Understanding In-Context Learning via Supportive Pretraining DataabstractXiaochuang Han, Daniel Simig, Todor Mihaylov, Yulia Tsvetkov, Asli Celikyilmaz, Tianlu Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Xiaochuang Han, Daniel Simig, Todor Mihaylov, Yulia Tsvetkov, Asli Celikyilmaz |
ACL (1) | 4 |
| 2023 | On the Blind Spots of Model-Based Evaluation Metrics for Text GenerationabstractTianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James Glass, Yulia Tsvetkov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Tianxing He, Tianle Wang 0003, Sachin Kumar 0009, Kyunghyun Cho, James R. Glass, Yulia Tsvetkov |
ACL (1) | 7 |
| 2023 | Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief TrackerabstractTheory of Mind (ToM)-the ability to reason about the mental states of other people-is a key element of our social intelligence.Yet, despite their ever more impressive performance, large-scale neural language models still lack basic theory of mind capabilities out-of-the-box.We posit that simply scaling up models will not imbue them with theory of mind due to the inherently symbolic and implicit nature of the phenomenon, and instead investigate an alternative: can we design a decoding-time algorithm that enhances theory of mind of off-the-shelf neural language models without explicit supervision?We present SYMBOLICTOM, a plug-andplay approach to reason about the belief states of multiple characters in reading comprehension tasks via explicit symbolic representation.More concretely, our approach tracks each entity's beliefs, their estimation of other entities' beliefs, and higher-order levels of reasoning, all through graphical representations, allowing for more precise and interpretable reasoning than previous approaches.Empirical results on the well-known ToMi benchmark (Le et al., 2019) demonstrate that SYMBOLICTOM dramatically enhances off-the-shelf neural networks' theory of mind in a zero-shot setting while showing robust out-of-distribution performance compared to supervised baselines.Our work also reveals spurious patterns in existing theory of mind benchmarks, emphasizing the importance of out-of-distribution evaluation and methods that do not overfit a particular dataset. Melanie Sclar, Sachin Kumar 0009, Peter West, Alane Suhr, Yejin Choi 0001, Yulia Tsvetkov |
ACL (1) | 6 |
| 2023 | Language Generation Models Can Cause Harm: So What Can We Do About It? An Actionable SurveyabstractSachin Kumar, Vidhisha Balachandran, Lucille Njoo, Antonios Anastasopoulos, Yulia Tsvetkov. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Sachin Kumar 0009, Vidhisha Balachandran, Lucille Njoo, Antonios Anastasopoulos, Yulia Tsvetkov |
EACL | 5 |
| 2023 | Do All Languages Cost the Same? Tokenization in the Era of Commercial Language ModelsabstractLanguage models have graduated from being research prototypes to commercialized products offered as web APIs, and recent works have highlighted the multilingual capabilities of these products.The API vendors charge their users based on usage, more specifically on the number of "tokens" processed or generated by the underlying language models.What constitutes a token, however, is training data and model dependent with a large variance in the number of tokens required to convey the same information in different languages.In this work, we analyze the effect of this nonuniformity on the fairness of an API's pricing policy across languages.We conduct a systematic analysis of the cost and utility of OpenAI's language model API on multilingual benchmarks in 22 typologically diverse languages.We show evidence that speakers of a large number of the supported languages are overcharged while obtaining poorer results.These speakers tend to also come from regions where the APIs are less affordable to begin with.Through these analyses, we aim to increase transparency around language model APIs' pricing policies and encourage the vendors to make them more equitable. Orevaoghene Ahia, Sachin Kumar 0009, Hila Gonen, Jungo Kasai, David R. Mortensen, Noah A. Smith, Yulia Tsvetkov |
EMNLP | 7 |
| 2023 | FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual KnowledgeabstractEvaluating the factual consistency of automatically generated summaries is essential for the progress and adoption of reliable summarization systems.Despite recent advances, existing factuality evaluation models are not robust, being especially prone to entity and relation errors in new domains.We propose FAC-TKB-a simple new approach to factuality evaluation that is generalizable across domains, in particular with respect to entities and relations.FACTKB is based on language models pretrained using facts extracted from external knowledge bases.We introduce three types of complementary factuality pretraining objectives based on entity-specific facts, facts extracted from auxiliary knowledge about entities, and facts constructed compositionally through knowledge base walks.The resulting factuality evaluation model achieves state-of-the-art performance on two in-domain news summarization benchmarks as well as on three outof-domain scientific literature datasets.Further analysis of FACTKB shows improved ability to detect erroneous entities and relations in summaries and is robust and easily generalizable across domains.Code and data are available at https://github.com/BunsenFeng/FactKB. Shangbin Feng, Vidhisha Balachandran, Yuyang Bai, Yulia Tsvetkov |
EMNLP | 4 |
| 2023 | GlobalBench: A Benchmark for Global Progress in Natural Language ProcessingabstractYueqi Song, Simran Khanuja, Pengfei Liu, Fahim Faisal, Alissa Ostapenko, Genta Winata, Alham Aji, Samuel Cahyawijaya, Yulia Tsvetkov, Antonios Anastasopoulos, Graham Neubig. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yueqi Song, Simran Khanuja, Pengfei Liu 0003, Fahim Faisal, Alissa Ostapenko, Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Yulia Tsvetkov, Antonios Anastasopoulos, Graham Neubig |
EMNLP | 9 |
| 2023 | Can Language Models Solve Graph Problems in Natural Language?abstractLarge language models (LLMs) are increasingly adopted for a variety of tasks with implicit graphical structures, such as planning in robotics, multi-hop question answering or knowledge probing, structured commonsense reasoning, and more. While LLMs have advanced the state-of-the-art on these tasks with structure implications, whether LLMs could explicitly process textual descriptions of graphs and structures, map them to grounded conceptual spaces, and perform structured operations remains underexplored. To this end, we propose NLGraph (Natural Language Graph), a comprehensive benchmark of graph-based problem solving designed in natural language. NLGraph contains 29,370 problems, covering eight graph reasoning tasks with varying complexity from simple tasks such as connectivity and shortest path up to complex problems such as maximum flow and simulating graph neural networks. We evaluate LLMs (GPT-3/4) with various prompting approaches on the NLGraph benchmark and find that 1) language models do demonstrate preliminary graph reasoning abilities, 2) the benefit of advanced prompting and in-context learning diminishes on more complex graph problems, while 3) LLMs are also (un)surprisingly brittle in the face of spurious correlations in graph and problem settings. We then propose Build-a-Graph Prompting and Algorithmic Prompting, two instruction-based approaches to enhance LLMs in solving natural language graph problems. Build-a-Graph and Algorithmic prompting improve the performance of LLMs on NLGraph by 3.07% to 16.85% across multiple tasks and settings, while how to solve the most complicated graph reasoning tasks in our setup with language models remains an open research question. Heng Wang 0008, Shangbin Feng, Tianxing He, Zhaoxuan Tan, Xiaochuang Han, Yulia Tsvetkov |
NeurIPS | 6 |
| 2022 | Speaker Information Can Guide Models to Better Inductive Biases: A Case Study On Predicting Code-SwitchingabstractNatural language processing (NLP) models trained on people-generated data can be unreliable because, without any constraints, they can learn from spurious correlations that are not relevant to the task.We hypothesize that enriching models with speaker information in a controlled, educated way can guide them to pick up on relevant inductive biases.For the speakerdriven task of predicting code-switching points in English-Spanish bilingual dialogues, we show that adding sociolinguistically-grounded speaker features as prepended prompts significantly improves accuracy.We find that by adding influential phrases to the input, speakerinformed models learn useful and explainable linguistic information.To our knowledge, we are the first to incorporate speaker characteristics in a neural model for code-switching, and more generally, take a step towards developing transparent, personalized models that use speaker information in a controlled way. Prompt Speaker Description Example ListASH is first speaker, older, female, from Spanish speaking country, between English and Spanish prefers both, rarely switches languages.JAC is second speaker, older, male, from Spanish speaking country, between English and Spanish prefers both, never switches languages.Sentence ASH is a middle-aged woman from a Spanish speaking country.Between English and Spanish she prefers both, and she rarely switches languages.ASH speaks first.JAC is a middle-aged man from a Spanish speaking country.Between English and Spanish he prefers both, and he never switches languages.JAC speaks second.Partner ASH, JAC are all middle-aged from a Spanish speaking country.Between English and Spanish they prefer both.ASH is a woman and rarely switches languages.JAC is a man and never switches languages.ASH speaks first.* About what happened in reality with, this guy, uh, who foresees the future. *… these types of movies confuse me. But …* is a … it was a documentary.Eh eh esa clase de películas me confunden.Pero* like I watch them with my girlfriend and she explains. Alissa Ostapenko, Shuly Wintner, Melinda Fricke, Yulia Tsvetkov |
ACL (1) | 4 |
| 2022 | Threat Scenarios and Best Practices to Detect Neural Fake NewsabstractIn this work, we discuss different threat scenarios from neural fake news generated by state-of-the-art language models. Through our experiments, we assess the performance of generated text detection systems under these threat scenarios. For each scenario, we also identify the minimax strategy for the detector that minimizes its worst-case performance. This constitutes a set of best practices that practitioners can rely on. In our analysis, we find that detectors are prone to shortcut learning (lack of out-of-distribution generalization) and discuss approaches to mitigate this problem and improve detectors more broadly. Finally, we argue that strong detectors should be released along with new generators. Artidoro Pagnoni, Martin Graciarena, Yulia Tsvetkov |
COLING | 3 |
| 2022 | Correcting Diverse Factual Errors in Abstractive Summarization via Post-Editing and Language Model InfillingabstractAbstractive summarization models often generate inconsistent summaries containing factual errors or hallucinated content.Recent works focus on correcting factual errors in generated summaries via post-editing.Such correction models are trained using adversarial nonfactual summaries constructed using heuristic rules for injecting errors.However, generating non-factual summaries using heuristics often does not generalize well to actual model errors.In this work, we propose to generate hard, representative synthetic examples of nonfactual summaries through infilling language models.With this data, we train a more robust fact-correction model to post-edit the summaries to improve factual consistency.Through quantitative and qualitative experiments on two popular summarization datasets-CNN/DM and XSum-we show that our approach vastly outperforms prior methods in correcting erroneous summaries.Our model-FACTEDITimproves factuality scores by over ∼11 points on CNN/DM and over ∼31 points on XSum on average across multiple summarization models, producing more factual summaries while maintaining competitive summarization quality. 1 The first vaccine for Ebola was approved by the FDA in 2019 in the US, five years after the initial outbreak in 2014.To produce the vaccine, scientists had to sequence the DNA of Ebola, then identify possible vaccines, and finally show successful clinical trials.Scientists say a vaccine for COVID-19 is unlikely to be ready this year, although clinical trials have already started.Scientists believe a vaccine for Covid-19 might not be ready this year.The first vaccine for Ebola took 5 years to be approved by the FDA.Scientists believe a vaccine for Ebola might not be ready this year.The first vaccine for Ebola took 5 years to be produced by the CBP. Vidhisha Balachandran, Hannaneh Hajishirzi, William W. Cohen, Yulia Tsvetkov |
EMNLP | 4 |
| 2022 | Gradient-based Constrained Sampling from Language ModelsabstractLarge pretrained language models generate fluent text but are notoriously hard to controllably sample from.In this work, we study constrained sampling from such language models: generating text that satisfies user-defined constraints, while maintaining fluency and model's performance in a downstream task.We propose MUCOLA-a sampling procedure that combines the log-likelihood of the language model with arbitrary (differentiable) constraints in a single energy function, and then generates samples in a non-autoregressive manner.Specifically, it initializes the entire output sequence with noise and follows a Markov chain defined by Langevin Dynamics using the gradients of the energy function.We evaluate MUCOLA on text generation with soft and hard constraints as well as their combinations obtaining significant improvements over competitive baselines for toxicity avoidance, sentiment control, and keyword-guided generation. 1 Sachin Kumar 0009, Biswajit Paria, Yulia Tsvetkov |
EMNLP | 3 |
| 2022 | Gendered Mental Health Stigma in Masked Language ModelsabstractMental health stigma prevents many individuals from receiving the appropriate care, and social psychology studies have shown that mental health tends to be overlooked in men.In this work, we investigate gendered mental health stigma in masked language models.In doing so, we operationalize mental health stigma by developing a framework grounded in psychology research: we use clinical psychology literature to curate prompts, then evaluate the models' propensity to generate gendered words.We find that masked language models capture societal stigma about gender in mental health: models are consistently more likely to predict female subjects than male in sentences about having a mental health condition (32% vs. 19%), and this disparity is exacerbated for sentences that indicate treatment-seeking behavior.Furthermore, we find that different models capture dimensions of stigma differently for men and women, associating stereotypes like anger, blame, and pity more with women with mental health conditions than with men.In showing the complex nuances of models' gendered mental health stigma, we demonstrate that context and overlapping dimensions of identity are important considerations when assessing computational models' social biases. Inna Wanyin Lin, Lucille Njoo, Anjalie Field, Ashish Sharma 0004, Katharina Reinecke, Tim Althoff, Yulia Tsvetkov |
EMNLP | 7 |
| 2022 | Referee: Reference-Free Sentence Summarization with Sharper Controllability through Symbolic Knowledge DistillationabstractWe present REFEREE, a novel framework for sentence summarization that can be trained reference-free (i.e., requiring no gold summaries for supervision), while allowing direct control for compression ratio.Our work is the first to demonstrate that reference-free, controlled sentence summarization is feasible via the conceptual framework of Symbolic Knowledge Distillation (West et al., 2022), where latent knowledge in pre-trained language models is distilled via explicit examples sampled from the teacher models, further purified with three types of filters: length, fidelity, and Information Bottleneck.Moreover, we uniquely propose iterative distillation of knowledge, where student models from the previous iteration of distillation serve as teacher models in the next iteration.Starting off from a relatively modest set of GPT3-generated summaries, we demonstrate how iterative knowledge distillation can lead to considerably smaller, but better summarizers with sharper controllability.A useful by-product of this iterative distillation process is a high-quality dataset of sentence-summary pairs with varying degrees of compression ratios.Empirical results demonstrate that the final student models vastly outperform the much larger GPT3-Instruct model in terms of the controllability of compression ratios, without compromising the quality of resulting summarization. 1 Melanie Sclar, Peter West, Sachin Kumar 0009, Yulia Tsvetkov, Yejin Choi 0001 |
EMNLP | 4 |
| 2022 | SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, Yuan Cao 0007 |
ICLR | 5 |
| 2022 | Controlled Analyses of Social Biases in Wikipedia BiosabstractSocial biases on Wikipedia, a widely-read global platform, could greatly influence public opinion. While prior research has examined man/woman gender bias in biography articles, possible influences of other demographic attributes limit conclusions. In this work, we present a methodology for analyzing Wikipedia pages about people that isolates dimensions of interest (e.g., gender), from other attributes (e.g., occupation). Given a target corpus for analysis (e.g. biographies about women), we present a method for constructing a comparison corpus that matches the target corpus in as many attributes as possible, except the target one. We develop evaluation metrics to measure how well the comparison corpus aligns with the target corpus and then examine how articles about gender and racial minorities (cis. women, non-binary people, transgender women, and transgender men; African American, Asian American, and Hispanic/Latinx American people) differ from other articles. In addition to identifying suspect social biases, our results show that failing to control for covariates can result in different conclusions and veil biases. Our contributions include methodology that facilitates further analyses of bias in Wikipedia articles, findings that can aid Wikipedia editors in reducing biases, and a framework and evaluation metrics to guide future work in this area. Anjalie Field, Chan Young Park, Kevin Z. Lin, Yulia Tsvetkov |
WWW | 4 |
| 2021 | A Survey of Race, Racism, and Anti-Racism in NLPabstractAnjalie Field, Su Lin Blodgett, Zeerak Waseem, Yulia Tsvetkov. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Anjalie Field, Su Lin Blodgett, Zeerak Talat, Yulia Tsvetkov |
ACL/IJCNLP (1) | 4 |
| 2021 | StructSum: Summarization via Structured RepresentationsabstractVidhisha Balachandran, Artidoro Pagnoni, Jay Yoon Lee, Dheeraj Rajagopal, Jaime Carbonell, Yulia Tsvetkov. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Vidhisha Balachandran, Artidoro Pagnoni, Jay-Yoon Lee, Dheeraj Rajagopal, Jaime G. Carbonell, Yulia Tsvetkov |
EACL | 6 |
| 2021 | Cross-Cultural Similarity Features for Cross-Lingual Transfer Learning of Pragmatically Motivated TasksabstractJimin Sun, Hwijeen Ahn, Chan Young Park, Yulia Tsvetkov, David R. Mortensen. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Jimin Sun, Hwijeen Ahn, Chan Young Park, Yulia Tsvetkov, David R. Mortensen |
EACL | 4 |
| 2021 | Evaluating the Morphosyntactic Well-formedness of Generated TextsabstractAdithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary, David R. Mortensen, Graham Neubig, Yulia Tsvetkov. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Adithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary, David R. Mortensen, Graham Neubig, Yulia Tsvetkov |
EMNLP (1) | 7 |
| 2021 | SELFEXPLAIN: A Self-Explaining Architecture for Neural Text ClassifiersabstractWe introduce SELFEXPLAIN, a novel selfexplaining model that explains a text classifier's predictions using phrase-based concepts.SELFEXPLAIN augments existing neural classifiers by adding (1) a globally interpretable layer that identifies the most influential concepts in the training set for a given sample and (2) a locally interpretable layer that quantifies the contribution of each local input concept by computing a relevance score relative to the predicted label.Experiments across five text-classification datasets show that SELFEX-PLAIN facilitates interpretability without sacrificing performance.Most importantly, explanations from SELFEXPLAIN show sufficiency for model predictions and are perceived as adequate, trustworthy and understandable by human judges compared to existing widely-used baselines.1 Dheeraj Rajagopal, Vidhisha Balachandran, Eduard H. Hovy, Yulia Tsvetkov |
EMNLP (1) | 4 |
| 2021 | DialoGraph: Incorporating Interpretable Strategy-Graph Networks into Negotiation Dialogues
Rishabh Joshi, Vidhisha Balachandran, Shikhar Vashishth, Alan W. Black, Yulia Tsvetkov |
ICLR | 5 |
| 2021 | Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual Models
Yulia Tsvetkov, Orhan Firat, Yuan Cao 0007 |
ICLR | 2 |
| 2021 | Multilingual Contextual Affective Analysis of LGBT People Portrayals in Wikipedia
Chan Young Park, Xinru Yan, Anjalie Field, Yulia Tsvetkov |
ICWSM | 4 |
| 2021 | Controlling Dialogue Generation with Semantic ExemplarsabstractPrakhar Gupta, Jeffrey Bigham, Yulia Tsvetkov, Amy Pavel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Prakhar Gupta, Jeffrey P. Bigham, Yulia Tsvetkov, Amy Pavel |
NAACL-HLT | 3 |
| 2021 | Understanding Factuality in Abstractive Summarization with FRANK: A Benchmark for Factuality MetricsabstractArtidoro Pagnoni, Vidhisha Balachandran, Yulia Tsvetkov. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Artidoro Pagnoni, Vidhisha Balachandran, Yulia Tsvetkov |
NAACL-HLT | 3 |
| 2021 | Controlled Text Generation as Continuous Optimization with Multiple ConstraintsabstractAs large-scale language model pretraining pushes the state-of-the-art in text generation, recent work has turned to controlling attributes of the text such models generate. While modifying the pretrained models via fine-tuning remains the popular approach, it incurs a significant computational cost and can be infeasible due to a lack of appropriate data. As an alternative, we propose \textsc{MuCoCO}---a flexible and modular algorithm for controllable inference from pretrained models. We formulate the decoding process as an optimization problem that allows for multiple attributes we aim to control to be easily incorporated as differentiable constraints. By relaxing this discrete optimization to a continuous one, we make use of Lagrangian multipliers and gradient-descent-based techniques to generate the desired text. We evaluate our approach on controllable machine translation and style transfer with multiple sentence-level attributes and observe significant improvements over baselines. Sachin Kumar 0009, Eric Malmi, Aliaksei Severyn, Yulia Tsvetkov |
NeurIPS | 4 |
| 2020 | Explaining Black Box Predictions and Unveiling Data Artifacts through Influence FunctionsabstractModern deep learning models for NLP are notoriously opaque.This has motivated the development of methods for interpreting such models, e.g., via gradient-based saliency maps or the visualization of attention weights.Such approaches aim to provide explanations for a particular model prediction by highlighting important words in the corresponding input text.While this might be useful for tasks where decisions are explicitly influenced by individual tokens in the input, we suspect that such highlighting is not always suitable for tasks where model decisions should be driven by more complex reasoning.In this work, we investigate the use of influence functions for NLP, providing an alternative approach to interpreting neural text classifiers.Influence functions explain the decisions of a model by identifying influential training examples.Despite the promise of this approach, influence functions have not yet been extensively evaluated in the context of NLP, a gap addressed by this work.We conduct a comparison between influence functions and common word-saliency methods on representative tasks.As suspected, we find that influence functions are particularly useful for natural language inference, a task in which 'saliency maps' may not provide clear interpretation.Furthermore, we develop a new quantitative measure based on influence functions that can reveal artifacts in training data. Xiaochuang Han, Byron C. Wallace, Yulia Tsvetkov |
ACL | 3 |
| 2020 | Balancing Training for Multilingual Neural Machine TranslationabstractWhen training multilingual machine translation (MT) models that can translate to/from multiple languages, we are faced with imbalanced training sets: some languages have much more training data than others.Standard practice is to up-sample less resourced languages to increase representation, and the degree of up-sampling has a large effect on the overall performance.In this paper, we propose a method that instead automatically learns how to weight training data through a data scorer that is optimized to maximize performance on all test languages.Experiments on two sets of languages under both one-to-many and manyto-one MT settings show our method not only consistently outperforms heuristic baselines in terms of average performance, but also offers flexible control over the performance of which languages are optimized.1 Xinyi Wang 0001, Yulia Tsvetkov, Graham Neubig |
ACL | 2 |
| 2020 | Understanding Linguistic Accommodation in Code-Switched Human-Machine DialoguesabstractCode-switching is a ubiquitous phenomenon in multilingual communities. Natural language technologies that wish to communicate like humans must therefore adaptively incorporate code-switching techniques when they are deployed in multilingual settings. To this end, we propose a Hindi-English human-machine dialogue system that elicits code-switching conversations in a controlled setting. It uses different code-switching agent strategies to understand how users respond and accommodate to the agent's language choice. Through this system, we collect and release a new dataset CommonDost, comprising of 439 human-machine multilingual conversations. We adapt pre-defined metrics to discover linguistic accommodation from users to agents. Finally, we compare these dialogues with Spanish-English dialogues collected in a similar setting, and analyze the impact of linguistic and socio-cultural factors on code-switching patterns across the two language pairs. Tanmay Parekh, Emily P. Ahn, Yulia Tsvetkov, Alan W. Black |
CoNLL | 3 |
| 2020 | Automatic Extraction of Rules Governing Morphological AgreementabstractAditi Chaudhary, Antonios Anastasopoulos, Adithya Pratapa, David R. Mortensen, Zaid Sheikh, Yulia Tsvetkov, Graham Neubig. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Aditi Chaudhary, Antonios Anastasopoulos, Adithya Pratapa, David R. Mortensen, Zaid Sheikh, Yulia Tsvetkov, Graham Neubig |
EMNLP (1) | 6 |
| 2020 | Unsupervised Discovery of Implicit Gender BiasabstractDespite their prevalence in society, social biases are difficult to identify, primarily because human judgements in this domain can be unreliable.We take an unsupervised approach to identifying gender bias against women at a comment level and present a model that can surface text likely to contain bias.Our main challenge is forcing the model to focus on signs of implicit bias, rather than other artifacts in the data.Thus, our methodology involves reducing the influence of confounds through propensity matching and adversarial learning.Our analysis shows how biased comments directed towards female politicians contain mixed criticisms, while comments directed towards other female public figures focus on appearance and sexualization.Ultimately, our work offers a way to capture subtle biases in various domains without relying on subjective human judgements. Anjalie Field, Yulia Tsvetkov |
EMNLP (1) | 2 |
| 2020 | Fortifying Toxic Speech Detectors Against Veiled ToxicityabstractModern toxic speech detectors are incompetent in recognizing disguised offensive language, such as adversarial attacks that deliberately avoid known toxic lexicons, or manifestations of implicit bias.Building a large annotated dataset for such veiled toxicity can be very expensive.In this work, we propose a framework aimed at fortifying existing toxic speech detectors without a large labeled corpus of veiled toxicity.Just a handful of probing examples are used to surface orders of magnitude more disguised offenses.We augment the toxic speech detector's training data with these discovered offensive examples, thereby making it more robust to veiled toxicity while preserving its utility in detecting overt toxicity.1 Warning: this paper contains examples that may be offensive or upsetting. Xiaochuang Han, Yulia Tsvetkov |
EMNLP (1) | 2 |
| 2020 | On Negative Interference in Multilingual Models: Findings and A Meta-Learning TreatmentabstractModern multilingual models are trained on concatenated text from multiple languages in hopes of conferring benefits to each (positive transfer), with the most pronounced benefits accruing to low-resource languages.However, recent work has shown that this approach can degrade performance on high-resource languages, a phenomenon known as negative interference.In this paper, we present the first systematic study of negative interference.We show that, contrary to previous belief, negative interference also impacts low-resource languages.While parameters are maximally shared to learn language-universal structures, we demonstrate that language-specific parameters do exist in multilingual models and they are a potential cause of negative interference.Motivated by these observations, we also present a meta-learning algorithm that obtains better cross-lingual transferability and alleviates negative interference, by adding languagespecific layers as meta-parameters and training them in a manner that explicitly improves shared layers' generalization on all languages.Overall, our results show that negative interference is more common than previously known, suggesting new directions for improving multilingual representations. 1 Model NER (F1) POS (F1) ar fr ru hi sw te avg Zachary C. Lipton, Yulia Tsvetkov |
EMNLP (1) | 3 |
| 2020 | Augmenting Non-Collaborative Dialog Systems with Explicit Semantic and Strategic Dialog History
Yiheng Zhou, Yulia Tsvetkov, Alan W. Black, Zhou Yu 0005 |
ICLR | 2 |
| 2019 | Entity-Centric Contextual Affective AnalysisabstractWhile contextualized word representations have improved state-of-the-art benchmarks in many NLP tasks, their potential usefulness for social-oriented tasks remains largely unexplored.We show how contextualized word embeddings can be used to capture affect dimensions in portrayals of people.We evaluate our methodology quantitatively, on held-out affect lexicons, and qualitatively, through case examples.We find that contextualized word representations do encode meaningful affect information, but they are heavily biased towards their training data, which limits their usefulness to in-domain analyses.We ultimately use our method to examine differences in portrayals of men and women. Anjalie Field, Yulia Tsvetkov |
ACL (1) | 2 |
| 2019 | Finding Microaggressions in the Wild: A Case for Locating Elusive Phenomena in Social Media PostsabstractLuke Breitfeller, Emily Ahn, David Jurgens, Yulia Tsvetkov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Luke Breitfeller, Emily P. Ahn, David Jurgens, Yulia Tsvetkov |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Topics to Avoid: Demoting Latent Confounds in Text ClassificationabstractSachin Kumar, Shuly Wintner, Noah A. Smith, Yulia Tsvetkov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sachin Kumar 0009, Shuly Wintner, Noah A. Smith, Yulia Tsvetkov |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Von Mises-Fisher Loss for Training Sequence to Sequence Models with Continuous Outputs
Sachin Kumar 0009, Yulia Tsvetkov |
ICLR (Poster) | 2 |
| 2019 | Contextual Affective Analysis: A Case Study of People Portrayals in Online #MeToo Stories
Anjalie Field, Gayatri Bhat, Yulia Tsvetkov |
ICWSM | 3 |
| 2019 | A Dynamic Strategy Coach for Effective NegotiationabstractNegotiation is a complex activity involving strategic reasoning, persuasion, and psychology.An average person is often far from an expert in negotiation.Our goal is to assist humans to become better negotiators through a machine-in-the-loop approach that combines machine's advantage at data-driven decisionmaking and human's language generation ability.We consider a bargaining scenario where a seller and a buyer negotiate the price of an item for sale through a text-based dialog.Our negotiation coach monitors messages between them and recommends tactics in real time to the seller to get a better deal (e.g., "reject the proposal and propose a price", "talk about your personal experience with the product").The best strategy and tactics largely depend on the context (e.g., the current price, the buyer's attitude).Therefore, we first identify a set of negotiation tactics, then learn to predict the best strategy and tactics in a given dialog context from a set of human-human bargaining dialogs.Evaluation on human-human dialogs shows that our coach increases the profits of the seller by almost 60%. 1 Yiheng Zhou, He He 0001, Alan W. Black, Yulia Tsvetkov |
SIGdial | 4 |
| 2018 | Style Transfer Through Back-TranslationabstractStyle transfer is the task of rephrasing the text to contain specific stylistic properties without changing the intent or affect within the context.This paper introduces a new method for automatic style transfer.We first learn a latent representation of the input sentence which is grounded in a language translation model in order to better preserve the meaning of the sentence while reducing stylistic properties.Then adversarial generation techniques are used to make the output match the desired style.We evaluate this technique on three different style transformations: sentiment, gender and political slant.Compared to two state-of-the-art style transfer modeling techniques we show improvements both in automatic evaluation of style transfer and in manual evaluation of meaning preservation and fluency. Shrimai Prabhumoye, Yulia Tsvetkov, Ruslan Salakhutdinov, Alan W. Black |
ACL (1) | 2 |
| 2018 | Framing and Agenda-Setting in Russian News: a Computational Analysis of Intricate Political StrategiesabstractAmidst growing concern over media manipulation, NLP attention has focused on overt strategies like censorship and "fake news".Here, we draw on two concepts from the political science literature to explore subtler strategies for government media manipulation: agenda-setting (selecting what topics to cover) and framing (deciding how topics are covered).We analyze 13 years (100K articles) of the Russian newspaper Izvestia and identify a strategy of distraction: articles mention the U.S. more frequently in the month directly following an economic downturn in Russia.We introduce embedding-based methods for cross-lingually projecting English frames to Russian, and discover that these articles emphasize U.S. moral failings and threats to the U.S. Our work offers new ways to identify subtle media manipulation strategies at the intersection of agenda-setting and framing. Anjalie Field, Doron Kliger, Shuly Wintner, Jennifer Pan, Daniel Jurafsky, Yulia Tsvetkov |
EMNLP | 6 |
| 2018 | RtGender: A Corpus for Studying Differential Responses to Gender
Rob Voigt, David Jurgens, Vinodkumar Prabhakaran, Daniel Jurafsky, Yulia Tsvetkov |
LREC | 5 |
| 2018 | Native Language Cognate Effects on Second Language Lexical ChoiceabstractWe present a computational analysis of cognate effects on the spontaneous linguistic productions of advanced non-native speakers. Introducing a large corpus of highly competent non-native English speakers, and using a set of carefully selected lexical items, we show that the lexical choices of non-natives are affected by cognates in their native language. This effect is so powerful that we are able to reconstruct the phylogenetic language tree of the Indo-European language family solely from the frequencies of specific lexical items in the English of authors with various native languages. We quantitatively analyze non-native lexical choice, highlighting cognate facilitation as one of the important phenomena shaping the language of non-native speakers. Ella Rabinovich, Yulia Tsvetkov, Shuly Wintner |
Trans. Assoc. Comput. Linguistics | 2 |
| 2016 | Learning the Curriculum with Bayesian Optimization for Task-Specific Word Representation LearningabstractWe use Bayesian optimization to learn curricula for word representation learning, optimizing performance on downstream tasks that depend on the learned representations as features.The curricula are modeled by a linear ranking function which is the scalar product of a learned weight vector and an engineered feature vector that characterizes the different aspects of the complexity of each instance in the training corpus.We show that learning the curriculum improves performance on a variety of downstream tasks over random orders and in comparison to the natural corpus order. Yulia Tsvetkov, Manaal Faruqui, Wang Ling, Brian MacWhinney, Chris Dyer |
ACL (1) | 1 |
| 2016 | Morphological Inflection Generation Using Character Sequence to Sequence LearningabstractManaal Faruqui, Yulia Tsvetkov, Graham Neubig, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Manaal Faruqui, Yulia Tsvetkov, Graham Neubig, Chris Dyer |
HLT-NAACL | 2 |
| 2016 | Polyglot Neural Language Models: A Case Study in Cross-Lingual Phonetic Representation LearningabstractYulia Tsvetkov, Sunayana Sitaram, Manaal Faruqui, Guillaume Lample, Patrick Littell, David Mortensen, Alan W Black, Lori Levin, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Yulia Tsvetkov, Sunayana Sitaram, Manaal Faruqui, Guillaume Lample, Patrick Littell, David R. Mortensen, Alan W. Black, Lori S. Levin, Chris Dyer |
HLT-NAACL | 1 |
| 2016 | Cross-Lingual Bridges with Models of Lexical BorrowingabstractLinguistic borrowing is the phenomenon of transferring linguistic constructions (lexical, phonological, morphological, and syntactic) from a donor language to a recipient language as a result of contacts between communities speaking different languages. Borrowed words are found in all languages, andin contrast to cognate relationshipsborrowing relationships may exist across unrelated languages (for example, about 40% of Swahilis vocabulary is borrowed from the unrelated language Arabic). In this work, we develop a model of morpho-phonological transformations across languages. Its features are based on universal constraints from Optimality Theory (OT), and we show that compared to several standardbut linguistically more naïvebaselines, our OT-inspired model obtains good performance at predicting donor forms from borrowed forms with only a few dozen training examples, making this a cost-effective strategy for sharing lexical information across languages. We demonstrate applications of the lexical borrowing model in machine translation, using resource-rich donor language to obtain translations of out-of-vocabulary loanwords in a lower resource language. Our framework obtains substantial improvements (up to 1.6 BLEU) over standard baselines. Yulia Tsvetkov, Chris Dyer |
J. Artif. Intell. Res. | 1 |
| 2015 | Sparse Overcomplete Word Vector RepresentationsabstractManaal Faruqui, Yulia Tsvetkov, Dani Yogatama, Chris Dyer, Noah A. Smith. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Manaal Faruqui, Yulia Tsvetkov, Dani Yogatama, Chris Dyer, Noah A. Smith |
ACL (1) | 2 |
| 2015 | Not All Contexts Are Created Equal: Better Word Representations with Variable AttentionabstractWang Ling, Yulia Tsvetkov, Silvio Amir, Ramón Fermandez, Chris Dyer, Alan W Black, Isabel Trancoso, Chu-Cheng Lin. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Wang Ling, Yulia Tsvetkov, Silvio Amir, Ramon Fermandez, Chris Dyer, Alan W. Black, Isabel Trancoso, Chu-Cheng Lin |
EMNLP | 2 |
| 2015 | Evaluation of Word Vector Representations by Subspace AlignmentabstractUnsupervisedly learned word vectors have proven to provide exceptionally effective features in many NLP tasks.Most common intrinsic evaluations of vector quality measure correlation with similarity judgments.However, these often correlate poorly with how well the learned representations perform as features in downstream evaluation tasks.We present QVEC-a computationally inexpensive intrinsic evaluation measure of the quality of word embeddings based on alignment to a matrix of features extracted from manually crafted lexical resources-that obtains strong correlation with performance of the vectors in a battery of downstream semantic evaluation tasks. 1 Yulia Tsvetkov, Manaal Faruqui, Wang Ling, Guillaume Lample, Chris Dyer |
EMNLP | 1 |
| 2015 | Constraint-Based Models of Lexical BorrowingabstractLinguistic borrowing is the phenomenon of transferring linguistic constructions (lexical, phonological, morphological, and syntactic) from a "donor" language to a "recipient" language as a result of contacts between communities speaking different languages.Borrowed words are found in all languages, and-in contrast to cognate relationships-borrowing relationships may exist across unrelated languages (for example, about 40% of Swahili's vocabulary is borrowed from Arabic).In this paper, we develop a model of morpho-phonological transformations across languages with features based on universal constraints from Optimality Theory (OT).Compared to several standardbut linguistically naïve-baselines, our OTinspired model obtains good performance with only a few dozen training examples, making this a cost-effective strategy for sharing lexical information across languages. Yulia Tsvetkov, Waleed Ammar, Chris Dyer |
HLT-NAACL | 1 |
| 2014 | Metaphor Detection with Cross-Lingual Model TransferabstractWe show that it is possible to reliably dis-criminate whether a syntactic construction is meant literally or metaphorically using lexical semantic features of the words that participate in the construction. Our model is constructed using English resources, and we obtain state-of-the-art performance relative to previous work in this language. Using a model transfer approach by piv-oting through a bilingual dictionary, we show our model can identify metaphoric expressions in other languages. We pro-vide results on three new test sets in Span-ish, Farsi, and Russian. The results sup-port the hypothesis that metaphors are conceptual, rather than lexical, in nature. 1 Yulia Tsvetkov, Leonid Boytsov, Anatole Gershman, Eric Nyberg, Chris Dyer |
ACL (1) | 1 |
| 2014 | Automatic Classification of Communicative Functions of Definiteness
Archna Bhatia, Chu-Cheng Lin, Nathan Schneider 0001, Yulia Tsvetkov, Fatima Talib Al-Raisi, Laleh Roostapour, Jordan Bender, Abhimanu Kumar, Lori S. Levin, Mandy Simons, Chris Dyer |
COLING | 4 |
| 2014 | Augmenting Translation Models with Simulated Acoustic Confusions for Improved Spoken Language TranslationabstractWe propose a novel technique for adapting text-based statistical machine translation to deal with input from automatic speech recognition in spoken language translation tasks.We simulate likely misrecognition errors using only a source language pronunciation dictionary and language model (i.e., without an acoustic model), and use these to augment the phrase table of a standard MT system.The augmented system can thus recover from recognition errors during decoding using synthesized phrases.Using the outputs of five different English ASR systems as input, we find consistent and significant improvements in translation quality.Our proposed technique can also be used in conjunction with lattices as ASR output, leading to further improvements. Yulia Tsvetkov, Florian Metze, Chris Dyer |
EACL | 1 |
| 2014 | A Unified Annotation Scheme for the Semantic/Pragmatic Components of Definiteness
Archna Bhatia, Mandy Simons, Lori S. Levin, Yulia Tsvetkov, Chris Dyer, Jordan Bender |
LREC | 4 |
| 2014 | Augmenting English Adjective Senses with Supersenses
Yulia Tsvetkov, Nathan Schneider 0001, Dirk Hovy, Archna Bhatia, Manaal Faruqui, Chris Dyer |
LREC | 1 |
| 2014 | Identification of Multiword Expressions by Combining Multiple Linguistic Information SourcesabstractWe propose a framework for using multiple sources of linguistic information in the task of identifying multiword expressions in natural language texts. We define various linguistically motivated classification features and introduce novel ways for computing them. We then manually define interrelationships among the features, and express them in a Bayesian network. The result is a powerful classifier that can identify multiword expressions of various types and multiple syntactic constructions in text corpora. Our methodology is unsupervised and language-independent; it requires relatively few language resources and is thus suitable for a large number of languages. We report results on English, French, and Hebrew, and demonstrate a significant improvement in identification accuracy, compared with less sophisticated baselines. Yulia Tsvetkov, Shuly Wintner |
Comput. Linguistics | 1 |
| 2013 | Identification and modeling of word fragments in spontaneous speechabstractThis paper presents a novel approach to handling disfluencies, word fragments and self-interruption points in Cantonese conversational speech. We train a classifier that exploits lexical and acoustic information to automatically identify disfluencies during training of a speech recognition system on conversational speech, and then use this classifier to augment reference annotations used for acoustic model training. We experiment with approaches to modeling disfluencies in the pronunciation dictionary, and their effect on the polyphonic decision tree clustering. We achieve automatic detection of disfluencies with 88% accuracy, which leads to a reduction in character error rate of 1.9% absolute. While the high baseline error rates are due to the task we are currently working on, we demonstrate that this approach works well on the Switchboard corpus, for which the conversational nature of speech is also a major problem. Yulia Tsvetkov, Zaid Sheikh, Florian Metze |
ICASSP | 1 |
| 2012 | Extraction of multi-word expressions from small parallel corporaabstractAbstract We present a general, novel methodology for extracting multi-word expressions (MWEs) of various types, along with their translations, from small, word-aligned parallel corpora. Unlike existing approaches, we focus on misalignments; these typically indicate expressions in the source language that are translated to the target in a non-compositional way. We introduce a simple algorithm that proposes MWE candidates based on such misalignments, relying on 1:1 alignments as anchors that delimit the search space. We use a large monolingual corpus to rank and filter these candidates. Evaluation of the quality of the extraction algorithm reveals significant improvements over naïve alignment-based methods. The extracted MWEs, with their translations, are used in the training of a statistical machine translation system, showing a small but significant improvement in its performance. Yulia Tsvetkov, Shuly Wintner |
Nat. Lang. Eng. | 1 |
| 2011 | Identification of Multi-word Expressions by Combining Multiple Linguistic Information Sources
Yulia Tsvetkov, Shuly Wintner |
EMNLP | 1 |
| 2010 | Automatic Acquisition of Parallel Corpora from Websites with Dynamic Content
Yulia Tsvetkov, Shuly Wintner |
LREC | 1 |