Alan Ritter

dblp:47/3133 · DBLP profile ↗
← Back
61ranked-venue papers
9as first author
33since 2021 · last 2026
0009-0001-4602-138XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 53 · 7 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 GeoRC: A Benchmark for Geolocation Reasoning Chains
abstract
Mohit Talreja, Joshua Diao, Jim James, Radu Casapu, Tejas Santanam, Ethan Mendes, Alan Ritter, Wei Xu, James Hays. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mohit Talreja, Joshua Diao, Jim James, Radu Casapu, Tejas Santanam, Ethan Mendes, Alan Ritter, Wei Xu 0004, James Hays
ACL (1)7
2026 Supporting Informed Self-Disclosure: Design Recommendations for Presenting AI-Estimates of Privacy Risks to Users
abstract
People candidly discuss sensitive topics online under the perceived safety of anonymity; yet, for many, this perceived safety is tenuous, as miscalibrated risk perceptions can lead to over-disclosure. Recent advances in Natural Language Processing (NLP) afford an unprecedented opportunity to present users with quantified disclosure-based re-identification risk — i.e., “population risk estimates” (PREs). How can PREs be presented to users in a way that promotes informed decision-making, mitigating risk without encouraging unnecessary self-censorship? Using design fictions and comic-boarding, we story-boarded five design concepts for presenting PREs to users and evaluated them through an online survey with N = 44 Reddit users. We found participants had detailed conceptions of how PREs may impact risk awareness and motivation, but envisioned needing additional context and support to effectively interpret and act on risks. We distill our findings into four key design recommendations for how best to present users with quantified privacy risks to support informed disclosure decision-making.
Isadora Krsek, Meryl Ye, Wei Xu 0004, Alan Ritter, Laura A. Dabbish, Sauvik Das
CHI4
2025 CROSSNEWS: A Cross-Genre Authorship Verification and Attribution Benchmark
abstract
Authorship models have historically generalized poorly to new domains because of the wide distribution of author-identifying signals across domains. In particular, the effects of topic and genre are highly domain-dependent and impact authorship analysis performance greatly. This paper addresses the existing data gap in authorship for these resources by introducing CROSSNEWS, a novel cross-genre dataset that connects formal journalistic articles and casual social media posts. CROSSNEWS is the largest authorship dataset of its kind for supporting both verification and attribution tasks, with comprehensive topic and genre annotations. We use CROSSNEWS to demonstrate that current models exhibit poor performance in genre transfer scenarios, underscoring the need for authorship models robust to genre-specific effects. We also explore SELMA, a new LLM embedding approach for large-scale authorship setups that outperforms existing models in both same-genre and cross-genre settings.
Marcus Ma, Duong Minh Le, Junmo Kang, Yao Dou, John Cadigan, Dayne Freitag, Alan Ritter, Wei Xu 0004
AAAI7
2025 Translation and Fusion Improves Cross-lingual Information Extraction
abstract
Large language models (LLMs) combined with instruction tuning have shown significant progress in information extraction (IE) tasks, exhibiting strong generalization capabilities to unseen datasets by following annotation guidelines.However, their applicability to lowresource languages remains limited due to lack of both labeled data for fine-tuning, and unlabeled text for pre-training.In this paper, we propose TransFusion, a framework in which models are fine-tuned to use English translations of low-resource language data, enabling more precise predictions through annotation fusion.Based on TransFusion, we introduce GoLLIE-TF, a cross-lingual instruction-tuned LLM for IE tasks, designed to close the performance gap between high and low-resource languages.Our experiments across twelve multilingual IE datasets spanning 50 languages demonstrate that GoLLIE-TF achieves better cross-lingual transfer over the base model.In addition, we show that TransFusion significantly improves low-resource language named entity recognition when applied to proprietary models such as GPT-4 (+5 F1) with a prompting approach, or fine-tuning different language models including decoder-only (+14 F1) and encoder-only (+13 F1) architectures.
Yang Chen 0065, Vedaant Shah, Alan Ritter
ACL (1)3
2025 Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
abstract
The surge of LLM studies makes synthesizing their findings challenging. Analysis of experimental results from literature can uncover important trends across studies, but the time-consuming nature of manual data extraction limits its use.Our study presents a semi-automated approach for literature analysis that accelerates data extraction using LLMs.It automatically identifies relevant arXiv papers, extracts experimental results and related attributes, and organizes them into a structured dataset, LLMEvalDB.We then conduct an automated literature analysis of frontier LLMs, reducing the effort of paper surveying and data extraction by more than 93% compared to manual approaches.We validate LLMEvalDB by showing that it reproduces key findings from a recent manual analysis of Chain-of-Thought (CoT) reasoning and also uncovers new insights that go beyond it, showing, for example, that in-context examples benefit coding & multimodal tasks but offer limited gains in math reasoning tasks compared to zero-shot CoT.Our automatically updatable dataset enables continuous tracking of target models by extracting evaluation studies as new data becomes available. Through LLMEvalDB and empirical analysis, we provide insights into LLMs while facilitating ongoing literature analyses of their behavior.
Jungsoo Park, Junmo Kang, Gabriel Stanovsky, Alan Ritter
ACL (1)4
2025 Balancing the Budget: Understanding Trade-offs Between Supervised and Preference-Based Finetuning
abstract
Post-training of Large Language Models often involves a pipeline of Supervised Finetuning (SFT) followed by Preference Finetuning (PFT) using methods like Direct Preference Optimization. Both stages require annotated data that are very different in structure and costs. We study how to optimally allocate a fixed training data budget between the two stages, through extensive experiments spanning four diverse tasks, multiple model sizes and various data annotation costs. Our findings reveal that just SFT on the base model dominates performance in low-data regimes (<1,000 annotated examples). With larger data-budgets, we observe that a combination of SFT and PFT, often with increasing portions allocated towards preference data yields optimal performance. However, completely eliminating SFT and running PFT directly on the base model yields suboptimal performance, described as the cold start problem on tasks like mathematics. We observe that this is due to the distribution shift arising from using DPO directly on the base model to elicit step-by-step reasoning. This limitation can be effectively addressed by allocating even a small portion (<10%) of the budget to SFT first, resulting in performance improvements of 15-20% on analytical benchmarks like GSM8k. These results provide actionable insights for researchers and practitioners optimizing model development under budget constraints, where high-quality data curation often represents a significant portion of the total costs of model development.
Mohit Raghavendra, Junmo Kang, Alan Ritter
ACL (1)3
2025 SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?
abstract
Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu, Jianfeng Gao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu 0004, Jianfeng Gao 0001
EMNLP6
2025 CARE: Multilingual Human Preference Learning for Cultural Awareness
abstract
Language Models (LMs) are typically tuned with human preferences to produce helpful responses, but the impact of preference tuning on the ability to handle culturally diverse queries remains understudied.In this paper, we systematically analyze how native human cultural preferences can be incorporated into the preference learning process to train more culturally aware LMs.We introduce CARE, a multilingual resource containing 3,490 culturally specific questions and 31.7kresponses with human judgments.We demonstrate how a modest amount of high-quality native preferences improves cultural awareness across various LMs, outperforming larger generic preference data.Our analyses reveal that models with stronger initial cultural performance benefit more from alignment, leading to gaps among models developed in different regions with varying access to culturally relevant data.CARE is publicly available at https://github.com/Guochry/ CARE.
Geyang Guo, Tarek Naous, Hiromi Wakaki, Yukiko Nishimura, Yuki Mitsufuji, Alan Ritter, Wei Xu 0004
EMNLP6
2025 How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation
abstract
As Large Language Models (LLMs) are widely deployed in diverse scenarios, the extent to which they could tacitly spread misinformation emerges as a critical safety concern.Current research primarily evaluates LLMs on explicit false statements, overlooking how misinformation often manifests subtly as unchallenged premises in real-world interactions.We curated ECHOMIST, the first comprehensive benchmark for implicit misinformation, where false assumptions are embedded in the query to LLMs.ECHOMIST targets circulated, harmful, and ever-evolving implicit misinformation from diverse sources, including realistic human-AI conversations and social media interactions.Through extensive empirical studies on 15 stateof-the-art LLMs, we find that current models perform alarmingly poorly on this task, often failing to detect false premises and generating counterfactual explanations.We also investigate two mitigation methods, i.e., Self-Alert and RAG, to enhance LLMs' capability to counter implicit misinformation.Our findings indicate that ECHOMIST remains a persistent challenge and underscore the critical need to safeguard against the risk of implicit misinformation. 1
Ruohao Guo, Wei Xu 0004, Alan Ritter
EMNLP3
2025 What are Foundation Models Cooking in the Post-Soviet World?
abstract
The culture of the Post-Soviet states is complex, shaped by a turbulent history that continues to influence current events.In this study, we investigate the Post-Soviet cultural food knowledge of foundation models by constructing BORSCH, a multimodal dataset encompassing 1147 and 823 dishes in the Russian and Ukrainian languages, centered around the Post-Soviet region.We demonstrate that leading models struggle to correctly identify the origins of dishes from Post-Soviet nations in both text-only and multimodal Question Answering (QA), instead over-predicting countries linked to the language the question is asked in.Through analysis of pretraining data, we show that these results can be explained by misleading dish-origin co-occurrences, along with linguistic phenomena such as Russian-Ukrainian code mixing.Finally, to move beyond QA-based assessments, we test models' abilities to produce accurate visual descriptions of dishes.The weak correlation between this task and QA suggests that QA alone may be insufficient as an evaluation of cultural understanding.To foster further research, we will make BORSCH publicly available at github.com/alavrouk/BORSch.
Anton Lavrouk, Tarek Naous, Alan Ritter, Wei Xu 0004
EMNLP3
2025 Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
abstract
We present Self-MoE, an approach that transforms a monolithic LLM into a compositional, modular system of self-specialized experts, named MiXSE (MiXture of Self-specialized Experts). Our approach leverages self-specialization, which constructs expert modules using self-generated synthetic data, each equipping a shared base LLM with distinct domain-specific capabilities, activated via self-optimized routing. This allows for dynamic and capability-specific handling of various target tasks, enhancing overall capabilities, without extensive human-labeled data and added parameters. Our empirical results reveal that specializing LLMs may exhibit potential trade-offs in performances on non-specialized tasks. On the other hand, our Self-MoE demonstrates substantial improvements (6.5%p on average) over the base LLM across diverse benchmarks such as knowledge, reasoning, math, and coding. It also consistently outperforms other methods, including instance merging and weight merging, while offering better flexibility and interpretability by design with semantic experts and routing. Our findings highlight the critical role of modularity, the applicability of Self-MoE to multiple base LLMs, and the potential of self-improvement in achieving efficient, scalable, and adaptable systems.
Junmo Kang, Leonid Karlinsky, Hongyin Luo, Zhen Wang 0041, Jacob A. Hansen, James R. Glass, David D. Cox, Rameswar Panda, Rogério Feris, Alan Ritter
ICLR10
2025 Language Models can Self-Improve at State-Value Estimation for Better Search
abstract
Collecting ground-truth rewards or human demonstrations for multi-step reasoning tasks is often prohibitively expensive, especially in interactive domains such as web tasks. We introduce Self-Taught Lookahead (STL), a reward-free framework that improves language model–based value functions by reasoning explicitly about state transitions. STL can be viewed as a chain-of-thought analogue of the value iteration algorithm: instead of regressing directly on numeric values, a value LLM is trained to simulate a step of lookahead in natural language—predicting the next action, resulting state, and rationale for its value. This process refines value estimates without any labeled data. The self-supervised procedure yields more accurate state-value predictions, which in turn enable lightweight search algorithms to expand fewer states while maintaining strong performance. Empirically, STL-trained value models built on moderately sized (8B-parameter) open-weight LLMs boost web agent success rates by over 39%, achieving performance comparable to proprietary models. STL also generalizes to multi-hop question answering and math puzzles. Overall, STL enables small open-source models to guide efficient search, reducing inference costs by integrating explicit reasoning with value learning.
Ethan Mendes, Alan Ritter
NeurIPS2
2025 Probabilistic Reasoning with LLMs for Privacy Risk Estimation
abstract
Probabilistic reasoning is a key aspect of both human and artificial intelligence that allows for handling uncertainty and ambiguity in decision-making. In this paper, we introduce a new numerical reasoning task under uncertainty for large language models, focusing on estimating the privacy risk of user-generated documents containing privacy-sensitive information. We propose BRANCH, a new LLM methodology that estimates the $k$-privacy value of a text—the size of the population matching the given information. BRANCH factorizes a joint probability distribution of personal information as random variables. The probability of each factor in a population is estimated separately using a Bayesian network and combined to compute the final $k$-value. Our experiments show that this method successfully estimates the $k$-value 73% of the time, a 13% increase compared to o3-mini with chain-of-thought reasoning. We also find that LLM uncertainty is a good indicator for accuracy, as high variance predictions are 37.47% less accurate on average.
Jonathan Zheng, Alan Ritter, Sauvik Das, Wei (Coco) Xu
NeurIPS2
2025 Measuring, Modeling, and Helping People Account for Privacy Risks in Online Self-Disclosures with AI
abstract
In pseudonymous online fora like Reddit, the benefits of self-disclosure are often apparent to users (e.g., I can vent about my in-laws to understanding strangers), but the privacy risks are more abstract (e.g., will my partner be able to tell that this is me?). Prior work has sought to develop natural language processing (NLP) tools that help users identify potentially risky self-disclosures in their text, but none have been designed for or evaluated with the users they hope to protect. Absent this assessment, these tools will be limited by the social-technical gap: users need assistive tools that help them make informed decisions, not paternalistic tools that tell them to avoid self-disclosure altogether. To bridge this gap, we conducted a study with N =21 Reddit users; we had them use a state-of-the-art NLP disclosure detection model on two of their authored posts and asked them questions to understand if and how the model helped, where it fell short, and how it could be improved to help them make more informed decisions. Despite its imperfections, users responded positively to the model and highlighted its use as a tool that can help them catch mistakes, inform them of risks they were unaware of, and encourage self-reflection. However, our work also shows how, to be useful and usable, AI for supporting privacy decision-making must account for posting context, disclosure norms, and users' lived threat models, and provide explanations that help contextualize detected risks.
Isadora Krsek, Anubha Kabra, Yao Dou, Tarek Naous, Laura A. Dabbish, Alan Ritter, Wei Xu 0004, Sauvik Das
Proc. ACM Hum. Comput. Interact.6
2024 Reducing Privacy Risks in Online Self-Disclosures with Language Models
abstract
Yao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra, Sauvik Das, Alan Ritter, Wei Xu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra, Sauvik Das, Alan Ritter, Wei Xu 0004
ACL (1)6
2024 Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
abstract
Language style is often used by writers to convey their intentions, identities, and mastery of language.In this paper, we show that current large language models struggle to capture some language styles without fine-tuning.To address this challenge, we investigate whether LLMs can be meta-trained based on representative lexicons to recognize new styles they have not been fine-tuned on.Experiments on 13 established style classification tasks, as well as 63 novel tasks generated using LLMs, demonstrate that meta-training with style lexicons consistently improves zero-shot transfer across styles.We release the code and data at https: //github.com/octaviaguo/Style-LLM.Instruction: Classify a sentence as "formal" if its style is similar to the words "albeit, lest, herein, insofar" or as "informal" if its style is similar to the words "imo, kinda, argh, omg".Here is the sentence: "I think she is unvirtuous."
Ruohao Guo, Wei Xu 0004, Alan Ritter
ACL (1)3
2024 Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
abstract
As the reach of large language models (LMs) expands globally, their ability to cater to diverse cultural contexts becomes crucial.Despite advancements in multilingual capabilities, models are not designed with appropriate cultural nuances.In this paper, we show that multilingual and Arabic monolingual LMs exhibit bias towards entities associated with Western culture.We introduce CAMeL, a novel resource of 628 naturally-occurring prompts and 20,368 entities spanning eight types that contrast Arab and Western cultures.CAMeL provides a foundation for measuring cultural biases in LMs through both extrinsic and intrinsic evaluations.Using CAMeL, we examine the cross-cultural performance in Arabic of 16 different LMs on tasks such as story generation, NER, and sentiment analysis, where we find concerning cases of stereotyping and cultural unfairness.We further test their text-infilling performance, revealing the incapability of appropriate adaptation to Arab cultural contexts.Finally, we analyze 6 Arabic pre-training corpora and find that commonly used sources such as Wikipedia may not be best suited to build culturally aware LMs, if used as they are without adjustment.We will make CAMeL publicly available at: https://github.com/tareknaous/camel
Tarek Naous, Michael J. Ryan, Alan Ritter, Wei Xu 0004
ACL (1)3
2024 NEO-BENCH: Evaluating Robustness of Large Language Models with Neologisms
abstract
The performance of Large Language Models (LLMs) degrades from the temporal drift between data used for model training and newer text seen during inference.One understudied avenue of language change causing data drift is the emergence of neologisms -new word forms -over time.We create a diverse resource of recent English neologisms by using several popular collection methods.We analyze temporal drift using neologisms by comparing sentences containing new words with near-identical sentences that replace neologisms with existing substitute words.Model performance is nearly halved in machine translation when a single neologism is introduced in a sentence.Motivated by these results, we construct a benchmark to evaluate LLMs' ability to generalize to neologisms with various natural language understanding tasks and model perplexity.Models with later knowledge cutoff dates yield lower perplexities and perform better in downstream tasks.LLMs are also affected differently based on the linguistic origins of words, indicating that neologisms are complex for static LLMs to address.We will release our benchmark and code for reproducing our experiments. Oracle Ensemble Microsoft Bing
Jonathan Zheng, Alan Ritter, Wei Xu 0004
ACL (1)2
2024 UniIR: Training and Benchmarking Universal Multimodal Information Retrievers
Cong Wei 0001, Yang Chen 0065, Hexiang Hu, Ge Zhang 0009, Jie Fu 0001, Alan Ritter, Wenhu Chen
ECCV (87)7
2024 Granular Privacy Control for Geolocation with Vision Language Models
abstract
Vision Language Models (VLMs) are rapidly advancing in their capability to answer information-seeking questions. As these models are widely deployed in consumer applications, they could lead to new privacy risks due to emergent abilities to identify people in photos, geolocate images, etc. As we demonstrate, somewhat surprisingly, current open-source and proprietary VLMs are very capable image geolocators, making widespread geolocation with VLMs an immediate privacy risk, rather than merely a theoretical future concern. As a first step to address this challenge, we develop a new benchmark, GPTGeoChat, to test the capability of VLMs to moderate geolocation dialogues with users. We collect a set of 1,000 image geolocation conversations between in-house annotators and GPT-4v, which are annotated with the granularity of location information revealed at each turn. Using this new dataset we evaluate the ability of various VLMs to moderate GPT-4v geolocation conversations by determining when too much location information has been revealed. We find that custom fine-tuned models perform on par with prompted API-based models when identifying leaked location information at the country or city level, however fine-tuning on supervised data appears to be needed to accurately moderate finer granularities, such as the name of a restaurant or building.
Ethan Mendes, Yang Chen 0065, James Hays, Sauvik Das, Wei Xu 0004, Alan Ritter
EMNLP6
2024 Constrained Decoding for Cross-lingual Label Projection
abstract
Zero-shot cross-lingual transfer utilizing multilingual LLMs has become a popular learning paradigm for low-resource languages with no labeled training data. However, for NLP tasks that involve fine-grained predictions on words and phrases, the performance of zero-shot cross-lingual transfer learning lags far behind supervised fine-tuning methods. Therefore, it is common to exploit translation and label projection to further improve the performance by (1) translating training data that is available in a high-resource language (e.g., English) together with the gold labels into low-resource languages, and/or (2) translating test data in low-resource languages to a high-source language to run inference on, then projecting the predicted span-level labels back onto the original test data. However, state-of-the-art marker-based label projection methods suffer from translation quality degradation due to the extra label markers injected in the input to the translation model. In this work, we explore a new direction that leverages constrained decoding for label projection to overcome the aforementioned issues. Our new method not only can preserve the quality of translated texts but also has the versatility of being applicable to both translating training and translating test data strategies. This versatility is crucial as our experiments reveal that translating test data can lead to a considerable boost in performance compared to translating only training data. We evaluate on two cross-lingual transfer tasks, namely Named Entity Recognition and Event Argument Extraction, spanning 20 languages. The results demonstrate that our approach outperforms the state-of-the-art marker-based method by a large margin and also shows better performance than other label projection methods that rely on external word alignment.
Duong Minh Le, Yang Chen 0065, Alan Ritter, Wei Xu 0004
ICLR3
2024 Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
abstract
While Large Language Models (LLMs) are increasingly being used in real-world applications, they remain vulnerable to *prompt injection attacks*: malicious third party prompts that subvert the intent of the system designer. To help researchers study this problem, we present a dataset of over 563,000 prompt injection attacks and 118,000 prompt-based "defenses" against prompt injection, all created by players of an online game called Tensor Trust. To the best of our knowledge, this is the first dataset that includes both human-generated attacks and defenses for instruction-following LLMs. The attacks in our dataset have easily interpretable structure, and shed light on the weaknesses of LLMs. We also use the dataset to create a benchmark for resistance to two types of prompt injection, which we refer to as *prompt extraction* and *prompt hijacking*. Our benchmark results show that many models are vulnerable to the attack strategies in the Tensor Trust dataset. Furthermore, we show that some attack strategies from the dataset generalize to deployed LLM-based applications, even though they have a very different set of constraints to the game. We release data and code at [tensortrust.ai/paper](https://tensortrust.ai/paper)
Sam Toyer, Olivia Watkins, Ethan Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, Stuart Russell 0001
ICLR11
2023 Distill or Annotate? Cost-Efficient Fine-Tuning of Compact Models
abstract
Fine-tuning large models is highly effective, however, inference can be expensive and produces carbon emissions.Knowledge distillation has been shown to be a practical solution to reduce inference costs, but the distillation process itself requires significant computational resources.Rather than buying or renting GPUs to fine-tune, then distill a large model, an NLP practitioner might instead choose to allocate the available budget to hire annotators and manually label additional fine-tuning data.In this paper, we investigate how to most efficiently use a fixed budget to build a compact model.Through extensive experiments on six diverse tasks, we show that distilling from T5-XXL (11B) to T5-Small (60M) is almost always a cost-efficient strategy compared to annotating more data to directly train a compact model (T5-Small).We further investigate how the optimal budget allocated towards computation varies across scenarios.We will make our code, datasets, annotation cost estimates, and baseline models available as a benchmark to support further work on cost-efficient training of compact models.
Junmo Kang, Wei Xu 0004, Alan Ritter
ACL (1)3
2023 Improved Instruction Ordering in Recipe-Grounded Conversation
abstract
In this paper, we study the task of instructional dialogue and focus on the cooking domain.Analyzing the generated output of the GPT-J model, we reveal that the primary challenge for a recipe-grounded dialog system is how to provide the instructions in the correct order.We hypothesize that this is due to the model's lack of understanding of user intent and inability to track the instruction state (i.e., which step was last instructed).Therefore, we propose to explore two auxiliary subtasks, namely User Intent Detection and Instruction State Tracking, to support Response Generation with improved instruction grounding.Experimenting with our newly collected dataset, ChattyChef, shows that incorporating user intent and instruction state information helps the response generation model mitigate the incorrect order issue.Furthermore, to investigate whether Chat-GPT has completely solved this task, we analyze its outputs and find that it also makes mistakes (10.7% of the responses), about half of which are out-of-order instructions.We will release ChattyChef to facilitate further research in this area at: https://github. com/octaviaguo/ChattyChef.
Duong Minh Le, Ruohao Guo, Wei Xu 0004, Alan Ritter
ACL (1)4
2023 Do CoNLL-2003 Named Entity Taggers Still Work Well in 2023?
abstract
The CoNLL-2003 English named entity recognition (NER) dataset has been widely used to train and evaluate NER models for almost 20 years.However, it is unclear how well models that are trained on this 20-year-old data and developed over a period of decades using the same test set will perform when applied on modern data.In this paper, we evaluate the generalization of over 20 different models trained on CoNLL-2003, and show that NER models have very different generalization.Surprisingly, we find no evidence of performance degradation in pre-trained Transformers, such as RoBERTa and T5, even when fine-tuned using decades-old data.We investigate why some models generalize well to new data while others do not, and attempt to disentangle the effects of temporal drift and overfitting due to test reuse.Our analysis suggests that most deterioration is due to temporal mismatch between the pre-training corpora and the downstream test sets.We found that four factors are important for good generalization: model architecture, number of parameters, time period of the pre-training corpus, in addition to the amount of fine-tuning data.We suggest current evaluation methods have, in some sense, underestimated progress on NER over the past 20 years, as NER models have not only improved on the original CoNLL-2003 test set, but improved even more on modern data.Our datasets can be found at https:// github.com/ShuhengL/acl2023_conllpp.
Shuheng Liu 0002, Alan Ritter
ACL (1)2
2023 Human-in-the-loop Evaluation for Early Misinformation Detection: A Case Study of COVID-19 Treatments
abstract
We present a human-in-the-loop evaluation framework for fact-checking novel misinformation claims and identifying social media messages that support them.Our approach extracts check-worthy claims, which are aggregated and ranked for review.Stance classifiers are then used to identify tweets supporting novel misinformation claims, which are further reviewed to determine whether they violate relevant policies.To demonstrate the feasibility of our approach, we develop a baseline system based on modern NLP methods for human-in-the-loop fact-checking in the domain of COVID-19 treatments.We make our data 1 and detailed annotation guidelines available to support the evaluation of human-in-the-loop systems that identify novel misinformation directly from raw usergenerated content.
Ethan Mendes, Yang Chen 0065, Wei Xu 0004, Alan Ritter
ACL (1)4
2023 Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
abstract
Pre-trained vision and language models (Chen et al., 2023b,a;Dai et al., 2023; Li et al., 2023b) have demonstrated state-of-the-art capabilities over existing tasks involving images and texts, including visual question answering.However, it remains unclear whether these models possess the capability to answer questions that are not only querying visual content but knowledge-intensive and informationseeking.In this study, we introduce INFOS-EEK 1 , a visual question answering dataset tailored for information-seeking questions that cannot be answered with only common sense knowledge.Using INFOSEEK, we analyze various pre-trained visual question answering models and gain insights into their characteristics.Our findings reveal that state-of-the-art pre-trained multi-modal models (e.g., PaLI-X, BLIP2, etc.) face challenges in answering visual information-seeking questions, but finetuning on the INFOSEEK dataset elicits models to use fine-grained knowledge that was learned during their pre-training.Furthermore, we show that accurate visual entity recognition can be used to improve performance on INFOSEEK by retrieving relevant documents, showing a significant space for improvement.* Work done when interned at Google 1 Our dataset is available at https:// open-vision-language.github.io/infoseek/.Dataset OK-VQA ViQuAE INFOSEEK PaLM (Q-only) 23.8 31.5 5.6 Current SotA 66.1 22.1 18.2 Require Knowledge † 29.2% 95.2% 95.6% † :% of questions that require knowledge to answer.PaLM (Q-only): a question-only baseline using PaLM.
Yang Chen 0065, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, Ming-Wei Chang
EMNLP6
2022 Extracting a Knowledge Base of COVID-19 Events from Social Media
abstract
We present a manually annotated corpus of 10,000 tweets containing public reports of five COVID-19 events, including positive and negative tests, deaths, denied access to testing, claimed cures and preventions. We designed slot-filling questions for each event type and annotated a total of 28 fine-grained slots, such as the location of events, recent travel, and close contacts. We show that our corpus can support fine-tuning BERT-based classifiers to automatically extract publicly reported events, which can be further collected for building a knowledge base. Our knowledge base is constructed over Twitter data covering two years and currently covers over 4.2M events. It can answer complex queries with high precision, such as “Which organizations have employees that tested positive in Philadelphia?” We believe our proposed methodology could be quickly applied to develop knowledge bases for new domains in response to an emerging crisis, including natural disasters or future disease outbreaks.
Shi Zong, Ashutosh Baheti, Wei Xu 0004, Alan Ritter
COLING4
2022 Stanceosaurus: Classifying Stance Towards Multicultural Misinformation
abstract
We present Stanceosaurus, a new corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims.As far as we are aware, it is the largest corpus annotated with stance towards misinformation claims.The claims in Stanceosaurus originate from 15 fact-checking sources that cover diverse geographical regions and cultures.Unlike existing stance datasets, we introduce a more fine-grained 5class labeling strategy with additional subcategories to distinguish implicit stance.Pretrained transformer-based stance classifiers that are fine-tuned on our corpus show good generalization on unseen claims and regional claims from countries outside the training data.Cross-lingual experiments demonstrate Stanceosaurus' capability of training multilingual models, achieving 53.1 F1 on Hindi and 50.4 F1 on Arabic without any targetlanguage fine-tuning.Finally, we show how a domain adaptation method can be used to improve performance on Stanceosaurus using additional RumourEval-2019 data.We make Stanceosaurus publicly available to the research community and hope it will encourage further work on misinformation identification across languages and cultures.1 Source Country & Regions Lang #Claims #Tweets Irr.Sup.Ref. Dis.Que.Snopes USA (80%), INT'L (16.7%),Other (3.3%) en 30 3197 1051 428 229 1447 42 Poynter Europe (5%), INT'L (90%), Other (5%) en 20 2197 949 274 97 844 33 FullFact UK (30%), INT'L (55%), Other (15%) en 20 2379 806 300 179 1057 37 AFP Fact Check CAN Canada (55%), INT'L (30%), Other (15%) en 20 2078 746 252 130 910 40 AAP Fact Check Australia (10%), INT'L (65%), Other (25%) en 20 2302 739 374 136 1019 34 AFP Fact Check NZ New Zealand (15%), INT'L (75%), Other (10%) en 20 2227 879 194 81 1044 29 Blackdotresearch Singapore (30%), INT'L (55%), Other (15%) en 20 2307 842 248 113 1076 28 Factly India (45%), INT'L (55%) en 20 1979 889 190 117 734 49 Politifact USA (20%), INT'L (35%), Other (45%) en 20 2041 984 289 8 753 7 Alt News India (90.4%),INT'L (4.8%),Other (4.8%) hi 21 1730 550 489 172 500 19 Aajtak India (67%), Other (33%) hi 9 806 456 110 40 193 7 Hindi Newschecker India (56%), Other (44%) hi 9 781 195 313 46 219 8 MISBAR Arab World (58.3%),INT'L (8.3%),Other (33.4%) ar 12 2283 454 514 203 1031 81 Fatabyyano Arab World (28.5%),INT'L (57.1%),Other (14.4%) ar 7 986 234 163 49 522 18 Maharat Fact-o-meter INT'L (100%) ar 3 740
Jonathan Zheng, Ashutosh Baheti, Tarek Naous, Wei Xu 0004, Alan Ritter
EMNLP5
2021 Process-Level Representation of Scientific Protocols with Interactive Annotation
abstract
We develop Process Execution Graphs (PEG), a document-level representation of real-world wet lab biochemistry protocols, addressing challenges such as cross-sentence relations, long-range coreference, grounding, and implicit arguments.We manually annotate PEGs in a corpus of complex lab protocols with a novel interactive textual simulator that keeps track of entity traits and semantic constraints during annotation.We use this data to develop graph-prediction models, finding them to be good at entity identification and local relation extraction, while our corpus facilitates further exploration of challenging long-range relations.1
Ronen Tamari, Fan Bai 0006, Alan Ritter, Gabriel Stanovsky
EACL3
2021 Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive Contexts
abstract
Dialogue models trained on human conversations inadvertently learn to generate toxic responses.In addition to producing explicitly offensive utterances, these models can also implicitly insult a group or individual by aligning themselves with an offensive statement.To better understand the dynamics of contextually offensive language, we investigate the stance of dialogue model responses in offensive Reddit conversations.Specifically, we create TOXICHAT, a crowd-annotated dataset of 2,000 Reddit threads and model responses labeled with offensive language and stance.Our analysis reveals that 42% of human responses agree with toxic comments, whereas only 13% agree with safe comments.This undesirable behavior is learned by neural dialogue models, such as DialoGPT, which we show are two times more likely to agree with offensive comments.To enable automatic detection of offensive language, we fine-tuned transformerbased classifiers on TOXICHAT that achieve 0.71 F 1 for offensive labels and 0.53 Macro-F 1 for stance labels.Finally, we quantify the effectiveness of controllable text generation (CTG) methods to mitigate the tendency of neural dialogue models to agree with offensive comments.Compared to the baseline, our best CTG model achieves a 19% reduction in agreement with offensive comments and produces 29% fewer offensive replies.Our work highlights the need for further efforts to characterize and analyze inappropriate behavior in dialogue models, in order to help make them safer. 1
Ashutosh Baheti, Maarten Sap, Alan Ritter, Mark O. Riedl
EMNLP (1)3
2021 Pre-train or Annotate? Domain Adaptation with a Constrained Budget
abstract
Recent work has demonstrated that pretraining in-domain language models can boost performance when adapting to a new domain.However, the costs associated with pretraining raise an important question: given a fixed budget, what steps should an NLP practitioner take to maximize performance?In this paper, we view domain adaptation with a constrained budget as a consumer choice problem, where the goal is to select an optimal combination of data annotation and pre-training.We measure annotation costs of three procedural text datasets, along with the pre-training costs of several in-domain language models.The utility of different combinations of pretraining and data annotation are evaluated under varying budget constraints to assess which combination strategy works best.We find that for small budgets, spending all funds on annotation leads to the best performance; once the budget becomes large enough, however, a combination of data annotation and in-domain pre-training yields better performance.Our experiments suggest task-specific data annotation should be part of an economical strategy when adapting an NLP model to a new domain.1
Fan Bai 0006, Alan Ritter, Wei Xu 0004
EMNLP (1)2
2021 Model Selection for Cross-lingual Transfer
abstract
Transformers that are pre-trained on multilingual corpora, such as, mBERT and XLM-RoBERTa, have achieved impressive crosslingual transfer capabilities.In the zero-shot transfer setting, only English training data is used, and the fine-tuned model is evaluated on another target language.While this works surprisingly well, substantial variance has been observed in target language performance between different fine-tuning runs, and in the zero-shot setup, no target-language development data is available to select among multiple fine-tuned models.Prior work has relied on English dev data to select among models that are fine-tuned with different learning rates, number of steps and other hyperparameters, often resulting in suboptimal choices.In this paper, we show that it is possible to select consistently better models when small amounts of annotated data are available in auxiliary pivot languages.We propose a machine learning approach to model selection that uses the finetuned model's own internal representations to predict its cross-lingual capabilities.In extensive experiments we find that this method consistently selects better models than English validation data across twenty five languages (including eight low-resource languages), and often achieves results that are comparable to model selection using target language development data. 1
Yang Chen 0065, Alan Ritter
EMNLP (1)2
2020 Fluent Response Generation for Conversational Question Answering
abstract
Question answering (QA) is an important aspect of open-domain conversational agents, garnering specific research focus in the conversational QA (ConvQA) subtask.One notable limitation of recent ConvQA efforts is the response being answer span extraction from the target corpus, thus ignoring the natural language generation (NLG) aspect of high-quality conversational agents.In this work, we propose a method for situating QA responses within a SEQ2SEQ NLG approach to generate fluent grammatical answer responses while maintaining correctness.From a technical perspective, we use data augmentation to generate training data for an end-to-end system.Specifically, we develop Syntactic Transformations (STs) to produce question-specific candidate answer responses and rank them using a BERT-based classifier (Devlin et al., 2019).Human evaluation on SQuAD 2.0 data (Rajpurkar et al., 2018) demonstrate that the proposed model outperforms baseline CoQA and QuAC models in generating conversational responses.We further show our model's scalability by conducting tests on the CoQA dataset. 1
Ashutosh Baheti, Alan Ritter, Kevin Small
ACL2
2020 Code and Named Entity Recognition in StackOverflow
abstract
There is an increasing interest in studying natural language and computer code together, as large corpora of programming texts become readily available on the Internet.For example, StackOverflow currently has over 15 million programming related questions written by 8.5 million users.Meanwhile, there is still a lack of fundamental NLP techniques for identifying code tokens or software-related named entities that appear within natural language sentences.In this paper, we introduce a new named entity recognition (NER) corpus for the computer programming domain, consisting of 15,372 sentences annotated with 20 fine-grained entity types.We trained indomain BERT representations (BERTOverflow) on 152 million sentences from Stack-Overflow, which lead to an absolute increase of +10 F 1 score over off-the-shelf BERT.We also present the SoftNER model which achieves an overall 79.10 F 1 score for code and named entity recognition on StackOverflow data.Our SoftNER model incorporates a context-independent code token classifier with corpus-level features to improve the BERTbased tagging model. 1
Jeniya Tabassum, Mounica Maddela, Wei Xu 0004, Alan Ritter
ACL4
2020 Measuring Forecasting Skill from Text
abstract
People vary in their ability to make accurate predictions about the future.Prior studies have shown that some individuals can predict the outcome of future events with consistently better accuracy.This leads to a natural question: what makes some forecasters better than others?In this paper we explore connections between the language people use to describe their predictions and their forecasting skill.Datasets from two different forecasting domains are explored: (1) geopolitical forecasts from Good Judgment Open, an online prediction forum and (2) a corpus of company earnings forecasts made by financial analysts.We present a number of linguistic metrics which are computed over text associated with people's predictions about the future including: uncertainty, readability, and emotion.By studying linguistic factors associated with predictions, we are able to shed some light on the approach taken by skilled forecasters.Furthermore, we demonstrate that it is possible to accurately predict forecasting skill using a model that is based solely on language.This could potentially be useful for identifying accurate predictions or potentially skilled forecasters earlier. 1
Shi Zong, Alan Ritter, Eduard H. Hovy
ACL2
2020 An Empirical Study of Pre-trained Transformers for Arabic Information Extraction
abstract
Multilingual pre-trained Transformers, such as mBERT (Devlin et al., 2019) and XLM-RoBERTa (Conneau et al., 2020a), have been shown to enable the effective cross-lingual zero-shot transfer.However, their performance on Arabic information extraction (IE) tasks is not very well studied.In this paper, we pre-train a customized bilingual BERT, dubbed GigaBERT, that is designed specifically for Arabic NLP and English-to-Arabic zero-shot transfer learning.We study Giga-BERT's effectiveness on zero-short transfer across four IE tasks: named entity recognition, part-of-speech tagging, argument role labeling, and relation extraction.Our best model significantly outperforms mBERT, XLM-RoBERTa, and AraBERT (Antoun et al., 2020) in both the supervised and zero-shot transfer settings.We have made our pre-trained models publicly available at https://github.com
Wuwei Lan, Yang Chen 0065, Wei Xu 0004, Alan Ritter
EMNLP (1)4
2019 Visual Exploration of Neural Document Embedding in Information Retrieval: Semantics and Feature Selection
abstract
Neural embeddings are widely used in language modeling and feature generation with superior computational power. Particularly, neural document embedding - converting texts of variable-length to semantic vector representations - has shown to benefit widespread downstream applications, e.g., information retrieval (IR). However, the black-box nature makes it difficult to understand how the semantics are encoded and employed. We propose visual exploration of neural document embedding to gain insights into the underlying embedding space, and promote the utilization in prevalent IR applications. In this study, we take an IR application-driven view, which is further motivated by biomedical IR in healthcare decision-making, and collaborate with domain experts to design and develop a visual analytics system. This system visualizes neural document embeddings as a configurable document map and enables guidance and reasoning; facilitates to explore the neural embedding space and identify salient neural dimensions (semantic features) per task and domain interest; and supports advisable feature selection (semantic analysis) along with instant visual feedback to promote IR performance. We demonstrate the usefulness and effectiveness of this system and present inspiring findings in use cases. This work will help designers/developers of downstream applications gain insights and confidence in neural document embedding, and exploit that to achieve more favorable performance in application domains.
Xiaonan Ji, Han-Wei Shen, Alan Ritter, Raghu Machiraju, Po-Yin Yen
IEEE Trans. Vis. Comput. Graph.3
2018 Generating More Interesting Responses in Neural Conversation Models with Distributional Constraints
abstract
Neural conversation models tend to generate safe, generic responses for most inputs.This is due to the limitations of likelihoodbased decoding objectives in generation tasks with diverse outputs, such as conversation.To address this challenge, we propose a simple yet effective approach for incorporating side information in the form of distributional constraints over the generated responses.We propose two constraints that help generate more content rich responses that are based on a model of syntax and topics (Griffiths et al., 2005) and semantic similarity (Arora et al., 2016).We evaluate our approach against a variety of competitive baselines, using both automatic metrics and human judgments, showing that our proposed approach generates responses that are much less generic without sacrificing plausibility.A working demo of our code can be found at https://github.com/abaheti95/ DC-NeuralConversation.
Ashutosh Baheti, Alan Ritter, Jiwei Li 0001, William B. Dolan
EMNLP2
2017 Adversarial Learning for Neural Dialogue Generation
abstract
In this paper, drawing intuition from the Turing test, we propose using adversarial training for open-domain dialogue generation: the system is trained to produce sequences that are indistinguishable from human-generated dialogue utterances.We cast the task as a reinforcement learning (RL) problem where we jointly train two systems, a generative model to produce response sequences, and a discriminator-analagous to the human evaluator in the Turing test-to distinguish between the human-generated dialogues and the machine-generated ones.The outputs from the discriminator are then used as rewards for the generative model, pushing the system to generate dialogues that mostly resemble human dialogues.In addition to adversarial training we describe a model for adversarial evaluation that uses success in fooling an adversary as a dialogue evaluation metric, while avoiding a number of potential pitfalls.Experimental results on several metrics, including adversarial evaluation, demonstrate that the adversarially-trained system generates higher-quality responses than previous baselines.
Jiwei Li 0001, Will Monroe, Tianlin Shi, Sébastien Jean, Alan Ritter, Daniel Jurafsky
EMNLP5
2017 "i have a feeling trump will win..................": Forecasting Winners and Losers from User Predictions on Twitter
abstract
Social media users often make explicit predictions about upcoming events.Such statements vary in the degree of certainty the author expresses toward the outcome: "Leonardo DiCaprio will win Best Actor" vs. "Leonardo DiCaprio may win" or "No way Leonardo wins!".Can popular beliefs on social media predict who will win?To answer this question, we build a corpus of tweets annotated for veridicality on which we train a log-linear classifier that detects positive veridicality with high precision.1 We then forecast uncertain outcomes using the wisdom of crowds, by aggregating users' explicit predictions.Our method for forecasting winners is fully automated, relying only on a set of contenders as input.It requires no training data of past outcomes and outperforms sentiment and tweet volume baselines on a broad range of contest prediction tasks.We further demonstrate how our approach can be used to measure the reliability of individual accounts' predictions and retrospectively identify surprise outcomes.
Sandesh Swamy, Alan Ritter, Marie-Catherine de Marneffe
EMNLP2
2017 Learning to Extract Events from Knowledge Base Revisions
abstract
Broad-coverage knowledge bases (KBs) such as Wikipedia, Freebase, Microsoft's Satori and Google's Knowledge Graph contain structured data describing real-world entities. These data sources have become increasingly important for a wide range of intelligent systems: from information retrieval and question answering, to Facebook's Graph Search, IBM's Watson, and more. Previous work on learning to populate knowledge bases from text has, for the most part, made the simplifying assumption that facts remain constant over time. But this is inaccurate -- we live in a rapidly changing world. Knowledge should not be viewed as a static snapshot, but instead a rapidly evolving set of facts that must change as the world changes.
Alexander Konovalov 0003, Benjamin Strauss, Alan Ritter, Brendan T. O'Connor 0001
WWW3
2017 Using ontology-based semantic similarity to facilitate the article screening process for systematic reviews
Xiaonan Ji, Alan Ritter, Po-Yin Yen
J. Biomed. Informatics2
2016 Deep Reinforcement Learning for Dialogue Generation
abstract
Recent neural models of dialogue generation offer great promise for generating responses for conversational agents, but tend to be shortsighted, predicting utterances one at a time while ignoring their influence on future outcomes.Modeling the future direction of a dialogue is crucial to generating coherent, interesting dialogues, a need which led traditional NLP models of dialogue to draw on reinforcement learning.In this paper, we show how to integrate these goals, applying deep reinforcement learning to model future reward in chatbot dialogue.The model simulates dialogues between two virtual agents, using policy gradient methods to reward sequences that display three useful conversational properties: informativity, coherence, and ease of answering (related to forward-looking function).We evaluate our model on diversity, length as well as with human judges, showing that the proposed algorithm generates more interactive responses and manages to foster a more sustained conversation in dialogue simulation.This work marks a first step towards learning a neural conversational model based on the long-term success of dialogues.
Jiwei Li 0001, Will Monroe, Alan Ritter, Daniel Jurafsky, Michel Galley, Jianfeng Gao 0001
EMNLP3
2016 TweeTime : A Minimally Supervised Method for Recognizing and Normalizing Time Expressions in Twitter
abstract
We describe TweeTIME, a temporal tagger for recognizing and normalizing time expressions in Twitter.Most previous work in social media analysis has to rely on temporal resolvers that are designed for well-edited text, and therefore suffer from reduced performance due to domain mismatch.We present a minimally supervised method that learns from large quantities of unlabeled data and requires no hand-engineered rules or hand-annotated training corpora.TweeTIME achieves 0.68 F1 score on the end-to-end task of resolving date expressions, outperforming a broad range of state-of-the-art systems. 1
Jeniya Tabassum, Alan Ritter, Wei Xu 0004
EMNLP2
2015 Never-Ending Learning
abstract
Whereas people learn many different types of knowledge from diverse experiences over many years, most current machine learning systems acquire just a single function or data model from just a single data set. We propose a never-ending learning paradigm for machine learning, to better reflect the more ambitious and encompassing type of learning performed by humans. As a case study, we describe the Never-Ending Language Learner (NELL), which achieves some of the desired properties of a never-ending learner, and we discuss lessons learned. NELL has been learning to read the web 24 hours/day since January 2010, and so far has acquired a knowledge base with over 80 million confidence-weighted beliefs (e.g., servedWith(tea, biscuits)). NELL has also learned millions of features and parameters that enable it to read these beliefs from the web. Additionally, it has learned to reason over these beliefs to infer new beliefs, and is able to extend its ontology by synthesizing new relational predicates. NELL can be tracked online at http://rtw.ml.cmu.edu, and followed on Twitter at @CMUNELL.
Tom M. Mitchell, William W. Cohen, Estevam Hruschka, Partha P. Talukdar, Justin Betteridge, Andrew Carlson, Bhavana Dalvi, Matt Gardner 0001, Bryan Kisiel, Jayant Krishnamurthy, Ni Lao, Kathryn Mazaitis, Thahir Mohamed, Ndapandula Nakashole, Emmanouil A. Platanios, Alan Ritter, Mehdi Samadi, Burr Settles, Richard C. Wang, Derry Wijaya, Abhinav Gupta 0001, Xinlei Chen, Abulhair Saparov, Malcolm Greaves, Joel Welling
AAAI16
2015 Examining the Distribution, Modularity, and Community Structure in Article Networks for Systematic Reviews
Xiaonan Ji, Raghu Machiraju, Alan Ritter, Po-Yin Yen
AMIA3
2015 Sense discovery via co-clustering on images and text
abstract
We present a co-clustering framework that can be used to discover multiple semantic and visual senses of a given Noun Phrase (NP). Unlike traditional clustering approaches which assume a one-to-one mapping between the clusters in the text-based feature space and the visual space, we adopt a one-to-many mapping between the two spaces. This is primarily because each semantic sense (concept) can correspond to different visual senses due to viewpoint and appearance variations. Our structure-EM style optimization not only extracts the multiple senses in both semantic and visual feature space, but also discovers the mapping between the senses. We introduce a challenging dataset (CMU Polysemy-30) for this problem consisting of 30 NPs (∼5600 labeled instances out of ∼22K total instances). We have also conducted a large-scale experiment that performs sense disambiguation for ∼2000 NPs.
Xinlei Chen, Alan Ritter, Abhinav Gupta 0001, Tom M. Mitchell
CVPR2
2015 Weakly Supervised Extraction of Computer Security Events from Twitter
abstract
Twitter contains a wealth of timely information, however staying on top of breaking events requires that an information analyst constantly scan many sources, leading to information overload. For example, a user might wish to be made aware whenever an infectious disease outbreak takes place, when a new smartphone is announced or when a distributed Denial of Service (DoS) attack might affect an organization's network connectivity. There are many possible event categories an analyst may wish to track, making it impossible to anticipate all those of interest in advance. We therefore propose a weakly supervised approach, in which extractors for new categories of events are easy to define and train, by specifying a small number of seed examples. We cast seed-based event extraction as a learning problem where only positive and unlabeled data is available. Rather than assuming unlabeled instances are negative, as is common in previous work, we propose a learning objective which regularizes the label distribution towards a user-provided expectation. Our approach greatly outperforms heuristic negatives, used in most previous work, in experiments on real-world data. Significant performance gains are also demonstrated over two novel and competitive baselines: semi-supervised EM and one-class support-vector machines. We investigate three security-related events breaking on Twitter: DoS attacks, data breaches and account hijacking. A demonstration of security events extracted by our system is available at: http://kb1.cse.ohio-state.edu:8123/events/hacked
Alan Ritter, Evan Wright, William Casey, Tom M. Mitchell
WWW1
2014 Weakly Supervised User Profile Extraction from Twitter
abstract
While user attribute extraction on social media has received considerable attention, existing approaches, mostly supervised, encounter great difficulty in obtaining gold standard data and are therefore limited to predicting unary predicates (e.g., gender).In this paper, we present a weaklysupervised approach to user profile extraction from Twitter.Users' profiles from social media websites such as Facebook or Google Plus are used as a distant source of supervision for extraction of their attributes from user-generated text.In addition to traditional linguistic features used in distant supervision for information extraction, our approach also takes into account network information, a unique opportunity offered by social media.We test our algorithm on three attribute domains: spouse, education and job; experimental results demonstrate our approach is able to make accurate predictions for users' attributes based on their tweets.1• We experimentally demonstrate the effectiveness of our approach on 3 relations: SPOUSE, JOB and EDUCATION.The remainder of this paper is organized as follows: We summarize related work in Section 2. The creation of our dataset is described in Section 3. The details of our model are presented in Section 4. We present experimental results in Section 5 and conclude in Section 6.
Jiwei Li 0001, Alan Ritter, Eduard H. Hovy
ACL (1)2
2014 Major Life Event Extraction from Twitter based on Congratulations/Condolences Speech Acts
abstract
Social media websites provide a platform for anyone to describe significant events taking place in their lives in realtime.Currently, the majority of personal news and life events are published in a textual format, motivating information extraction systems that can provide a structured representations of major life events (weddings, graduation, etc. . .).This paper demonstrates the feasibility of accurately extracting major life events.Our system extracts a fine-grained description of users' life events based on their published tweets.We are optimistic that our system can help Twitter users more easily grasp information from users they take interest in following and also facilitate many downstream applications, for example realtime friend recommendation.
Jiwei Li 0001, Alan Ritter, Claire Cardie, Eduard H. Hovy
EMNLP2
2014 Extracting Lexically Divergent Paraphrases from Twitter
abstract
We present MultiP (Multi-instance Learning Paraphrase Model), a new model suited to identify paraphrases within the short messages on Twitter. We jointly model paraphrase relations between word and sentence pairs and assume only sentence-level annotations during learning. Using this principled latent variable model alone, we achieve the performance competitive with a state-of-the-art method which combines a latent space model with a feature-based supervised classifier. Our model also captures lexically divergent paraphrases that differ from yet complement previous methods; combining our model with previous work significantly outperforms the state-of-the-art. In addition, we present a novel annotation methodology that has allowed us to crowdsource a paraphrase corpus from Twitter. We make this new dataset available to the research community.
Wei Xu 0004, Alan Ritter, Chris Callison-Burch, William B. Dolan, Yangfeng Ji
Trans. Assoc. Comput. Linguistics2
2013 Modeling Missing Data in Distant Supervision for Information Extraction
abstract
Distant supervision algorithms learn information extraction models given only large readily available databases and text collections. Most previous work has used heuristics for generating labeled data, for example assuming that facts not contained in the database are not mentioned in the text, and facts in the database must be mentioned at least once. In this paper, we propose a new latent-variable approach that models missing data. This provides a natural way to incorporate side information, for instance modeling the intuition that text will often mention rare entities which are likely to be missing in the database. Despite the added complexity introduced by reasoning about missing data, we demonstrate that a carefully designed local search approach to inference is very accurate and scales to large datasets. Experiments demonstrate improved performance for binary and unary relation extraction when compared to learning with heuristic labels, including on average a 27% increase in area under the precision recall curve in the binary case.
Alan Ritter, Luke Zettlemoyer, Mausam, Oren Etzioni
Trans. Assoc. Comput. Linguistics1
2012 Paraphrasing for Style
Wei Xu 0004, Alan Ritter, William B. Dolan, Ralph Grishman, Colin Cherry
COLING2
2012 Open domain event extraction from twitter
abstract
Tweets are the most up-to-date and inclusive stream of in- formation and commentary on current events, but they are also fragmented and noisy, motivating the need for systems that can extract, aggregate and categorize important events. Previous work on extracting structured representations of events has focused largely on newswire text; Twitter's unique characteristics present new challenges and opportunities for open-domain event extraction. This paper describes TwiCal-- the first open-domain event-extraction and categorization system for Twitter. We demonstrate that accurately extracting an open-domain calendar of significant events from Twitter is indeed feasible. In addition, we present a novel approach for discovering important event categories and classifying extracted events based on latent variable models. By leveraging large volumes of unlabeled data, our approach achieves a 14% increase in maximum F1 over a supervised baseline. A continuously updating demonstration of our system can be viewed at http://statuscalendar.com; Our NLP tools are available at http://github.com/aritter/ twitter_nlp.
Alan Ritter, Mausam, Oren Etzioni, Sam Clark
KDD1
2011 Data-Driven Response Generation in Social Media
Alan Ritter, Colin Cherry, William B. Dolan
EMNLP1
2011 Named Entity Recognition in Tweets: An Experimental Study
Alan Ritter, Sam Clark, Mausam, Oren Etzioni
EMNLP1
2010 A Latent Dirichlet Allocation Method for Selectional Preferences
Alan Ritter, Mausam, Oren Etzioni
ACL1
2010 Unsupervised Modeling of Twitter Conversations
Alan Ritter, Colin Cherry, William B. Dolan
HLT-NAACL1
2009 Learning to generalize for complex selection tasks
abstract
Selection tasks are common in modern computer interfaces: we are often required to select a set of files, emails, data entries, and the like. File and data browsers have sorting and block selection facilities to make these tasks easier, but for complex selections there is little to aid the user without writing complex search queries. We propose an interactive machine learning solution to this problem called "smart selection," in which the user selects and deselects items as inputs to a selection classifier which attempts at each step to correctly generalize to the user's target state. Furthermore, we take advantage of our data on how users perform selection tasks over many sessions, and use it to train a label regressor that models their generalization behavior: we call this process learning to generalize. We then combine the user's explicit labels as well the label regressor outputs in the selection classifier to predict the user's desired selections. We show that the selection classifier alone takes dramatically fewer mouse clicks than the standard file browser, and when used in conjunction with the label regressor, the predictions of the classifier are significantly more accurate with respect to the target selection state.
Alan Ritter, Sumit Basu
IUI1
2008 It's a Contradiction - no, it's not: A Case Study using Functional Relations
Alan Ritter, Stephen Soderland, Doug Downey, Oren Etzioni
EMNLP1