EDBT 2026 Demo / reviewers in the wild / expert
Wei Xu 0004
dblp:32/1213-4
· DBLP profile ↗
62ranked-venue papers
7as first author
39since 2021 · last 2026
0000-0002-7044-3232ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 58 · 7 first-author · 36 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GeoRC: A Benchmark for Geolocation Reasoning ChainsabstractMohit Talreja, Joshua Diao, Jim James, Radu Casapu, Tejas Santanam, Ethan Mendes, Alan Ritter, Wei Xu, James Hays. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mohit Talreja, Joshua Diao, Jim James, Radu Casapu, Tejas Santanam, Ethan Mendes, Alan Ritter, Wei Xu 0004, James Hays |
ACL (1) | 8 |
| 2026 | Supporting Informed Self-Disclosure: Design Recommendations for Presenting AI-Estimates of Privacy Risks to UsersabstractPeople candidly discuss sensitive topics online under the perceived safety of anonymity; yet, for many, this perceived safety is tenuous, as miscalibrated risk perceptions can lead to over-disclosure. Recent advances in Natural Language Processing (NLP) afford an unprecedented opportunity to present users with quantified disclosure-based re-identification risk — i.e., “population risk estimates” (PREs). How can PREs be presented to users in a way that promotes informed decision-making, mitigating risk without encouraging unnecessary self-censorship? Using design fictions and comic-boarding, we story-boarded five design concepts for presenting PREs to users and evaluated them through an online survey with N = 44 Reddit users. We found participants had detailed conceptions of how PREs may impact risk awareness and motivation, but envisioned needing additional context and support to effectively interpret and act on risks. We distill our findings into four key design recommendations for how best to present users with quantified privacy risks to support informed disclosure decision-making. Isadora Krsek, Meryl Ye, Wei Xu 0004, Alan Ritter, Laura A. Dabbish, Sauvik Das |
CHI | 3 |
| 2025 | CROSSNEWS: A Cross-Genre Authorship Verification and Attribution BenchmarkabstractAuthorship models have historically generalized poorly to new domains because of the wide distribution of author-identifying signals across domains. In particular, the effects of topic and genre are highly domain-dependent and impact authorship analysis performance greatly. This paper addresses the existing data gap in authorship for these resources by introducing CROSSNEWS, a novel cross-genre dataset that connects formal journalistic articles and casual social media posts. CROSSNEWS is the largest authorship dataset of its kind for supporting both verification and attribution tasks, with comprehensive topic and genre annotations. We use CROSSNEWS to demonstrate that current models exhibit poor performance in genre transfer scenarios, underscoring the need for authorship models robust to genre-specific effects. We also explore SELMA, a new LLM embedding approach for large-scale authorship setups that outperforms existing models in both same-genre and cross-genre settings. Marcus Ma, Duong Minh Le, Junmo Kang, Yao Dou, John Cadigan, Dayne Freitag, Alan Ritter, Wei Xu 0004 |
AAAI | 8 |
| 2025 | SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?abstractYao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu, Jianfeng Gao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu 0004, Jianfeng Gao 0001 |
EMNLP | 8 |
| 2025 | CARE: Multilingual Human Preference Learning for Cultural AwarenessabstractLanguage Models (LMs) are typically tuned with human preferences to produce helpful responses, but the impact of preference tuning on the ability to handle culturally diverse queries remains understudied.In this paper, we systematically analyze how native human cultural preferences can be incorporated into the preference learning process to train more culturally aware LMs.We introduce CARE, a multilingual resource containing 3,490 culturally specific questions and 31.7kresponses with human judgments.We demonstrate how a modest amount of high-quality native preferences improves cultural awareness across various LMs, outperforming larger generic preference data.Our analyses reveal that models with stronger initial cultural performance benefit more from alignment, leading to gaps among models developed in different regions with varying access to culturally relevant data.CARE is publicly available at https://github.com/Guochry/ CARE. Geyang Guo, Tarek Naous, Hiromi Wakaki, Yukiko Nishimura, Yuki Mitsufuji, Alan Ritter, Wei Xu 0004 |
EMNLP | 7 |
| 2025 | How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit MisinformationabstractAs Large Language Models (LLMs) are widely deployed in diverse scenarios, the extent to which they could tacitly spread misinformation emerges as a critical safety concern.Current research primarily evaluates LLMs on explicit false statements, overlooking how misinformation often manifests subtly as unchallenged premises in real-world interactions.We curated ECHOMIST, the first comprehensive benchmark for implicit misinformation, where false assumptions are embedded in the query to LLMs.ECHOMIST targets circulated, harmful, and ever-evolving implicit misinformation from diverse sources, including realistic human-AI conversations and social media interactions.Through extensive empirical studies on 15 stateof-the-art LLMs, we find that current models perform alarmingly poorly on this task, often failing to detect false premises and generating counterfactual explanations.We also investigate two mitigation methods, i.e., Self-Alert and RAG, to enhance LLMs' capability to counter implicit misinformation.Our findings indicate that ECHOMIST remains a persistent challenge and underscore the critical need to safeguard against the risk of implicit misinformation. 1 Ruohao Guo, Wei Xu 0004, Alan Ritter |
EMNLP | 2 |
| 2025 | What are Foundation Models Cooking in the Post-Soviet World?abstractThe culture of the Post-Soviet states is complex, shaped by a turbulent history that continues to influence current events.In this study, we investigate the Post-Soviet cultural food knowledge of foundation models by constructing BORSCH, a multimodal dataset encompassing 1147 and 823 dishes in the Russian and Ukrainian languages, centered around the Post-Soviet region.We demonstrate that leading models struggle to correctly identify the origins of dishes from Post-Soviet nations in both text-only and multimodal Question Answering (QA), instead over-predicting countries linked to the language the question is asked in.Through analysis of pretraining data, we show that these results can be explained by misleading dish-origin co-occurrences, along with linguistic phenomena such as Russian-Ukrainian code mixing.Finally, to move beyond QA-based assessments, we test models' abilities to produce accurate visual descriptions of dishes.The weak correlation between this task and QA suggests that QA alone may be insufficient as an evaluation of cultural understanding.To foster further research, we will make BORSCH publicly available at github.com/alavrouk/BORSch. Anton Lavrouk, Tarek Naous, Alan Ritter, Wei Xu 0004 |
EMNLP | 4 |
| 2025 | Generating CAD Code with Vision-Language Models for 3D DesignsabstractGenerative AI has transformed the fields of Design and Manufacturing by providing
efficient and automated methods for generating and modifying 3D objects. One
approach involves using Large Language Models (LLMs) to generate Computer-
Aided Design (CAD) scripting code, which can then be executed to render a 3D
object; however, the resulting 3D object may not meet the specified requirements.
Testing the correctness of CAD generated code is challenging due to the complexity
and structure of 3D objects (e.g., shapes, surfaces, and dimensions) that are not
feasible in code. In this paper, we introduce CADCodeVerify, a novel approach to
iteratively verify and improve 3D objects generated from CAD code. Our approach
works by producing ameliorative feedback by prompting a Vision-Language Model
(VLM) to generate and answer a set of validation questions to verify the generated
object and prompt the VLM to correct deviations. To evaluate CADCodeVerify, we
introduce, CADPrompt, the first benchmark for CAD code generation, consisting of
200 natural language prompts paired with expert-annotated scripting code for 3D
objects to benchmark progress. Our findings show that CADCodeVerify improves
VLM performance by providing visual feedback, enhancing the structure of the 3D
objects, and increasing the success rate of the compiled program. When applied to
GPT-4, CADCodeVerify achieved a 7.30% reduction in Point Cloud distance and a
5.0% improvement in success rate compared to prior work. Kamel Alrashedy, Pradyumna Tambwekar, Zulfiqar Zaidi, Megan Langwasser, Wei Xu 0004, Matthew C. Gombolay |
ICLR | 5 |
| 2025 | On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic PhenomenaabstractTarek Naous, Wei Xu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tarek Naous, Wei Xu 0004 |
NAACL (Long Papers) | 2 |
| 2025 | The Impact of Visual Information in Chinese Characters: Evaluating Large Models' Ability to Recognize and Utilize RadicalsabstractXiaofeng Wu, Karl Stratos, Wei Xu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Karl Stratos, Wei Xu 0004 |
NAACL (Long Papers) | 3 |
| 2025 | Measuring, Modeling, and Helping People Account for Privacy Risks in Online Self-Disclosures with AIabstractIn pseudonymous online fora like Reddit, the benefits of self-disclosure are often apparent to users (e.g., I can vent about my in-laws to understanding strangers), but the privacy risks are more abstract (e.g., will my partner be able to tell that this is me?). Prior work has sought to develop natural language processing (NLP) tools that help users identify potentially risky self-disclosures in their text, but none have been designed for or evaluated with the users they hope to protect. Absent this assessment, these tools will be limited by the social-technical gap: users need assistive tools that help them make informed decisions, not paternalistic tools that tell them to avoid self-disclosure altogether. To bridge this gap, we conducted a study with N =21 Reddit users; we had them use a state-of-the-art NLP disclosure detection model on two of their authored posts and asked them questions to understand if and how the model helped, where it fell short, and how it could be improved to help them make more informed decisions. Despite its imperfections, users responded positively to the model and highlighted its use as a tool that can help them catch mistakes, inform them of risks they were unaware of, and encourage self-reflection. However, our work also shows how, to be useful and usable, AI for supporting privacy decision-making must account for posting context, disclosure norms, and users' lived threat models, and provide explanations that help contextualize detected risks. Isadora Krsek, Anubha Kabra, Yao Dou, Tarek Naous, Laura A. Dabbish, Alan Ritter, Wei Xu 0004, Sauvik Das |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2024 | Reducing Privacy Risks in Online Self-Disclosures with Language ModelsabstractYao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra, Sauvik Das, Alan Ritter, Wei Xu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra, Sauvik Das, Alan Ritter, Wei Xu 0004 |
ACL (1) | 7 |
| 2024 | Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style UnderstandingabstractLanguage style is often used by writers to convey their intentions, identities, and mastery of language.In this paper, we show that current large language models struggle to capture some language styles without fine-tuning.To address this challenge, we investigate whether LLMs can be meta-trained based on representative lexicons to recognize new styles they have not been fine-tuned on.Experiments on 13 established style classification tasks, as well as 63 novel tasks generated using LLMs, demonstrate that meta-training with style lexicons consistently improves zero-shot transfer across styles.We release the code and data at https: //github.com/octaviaguo/Style-LLM.Instruction: Classify a sentence as "formal" if its style is similar to the words "albeit, lest, herein, insofar" or as "informal" if its style is similar to the words "imo, kinda, argh, omg".Here is the sentence: "I think she is unvirtuous." Ruohao Guo, Wei Xu 0004, Alan Ritter |
ACL (1) | 2 |
| 2024 | FactPICO: Factuality Evaluation for Plain Language Summarization of Medical EvidenceabstractSebastian Joseph, Lily Chen, Jan Trienes, Hannah Göke, Monika Coers, Wei Xu, Byron Wallace, Junyi Jessy Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Sebastian Joseph, Lily Chen, Jan Trienes, Hannah Louisa Göke, Monika Coers, Wei Xu 0004, Byron C. Wallace, Junyi Jessy Li |
ACL (1) | 6 |
| 2024 | Having Beer after Prayer? Measuring Cultural Bias in Large Language ModelsabstractAs the reach of large language models (LMs) expands globally, their ability to cater to diverse cultural contexts becomes crucial.Despite advancements in multilingual capabilities, models are not designed with appropriate cultural nuances.In this paper, we show that multilingual and Arabic monolingual LMs exhibit bias towards entities associated with Western culture.We introduce CAMeL, a novel resource of 628 naturally-occurring prompts and 20,368 entities spanning eight types that contrast Arab and Western cultures.CAMeL provides a foundation for measuring cultural biases in LMs through both extrinsic and intrinsic evaluations.Using CAMeL, we examine the cross-cultural performance in Arabic of 16 different LMs on tasks such as story generation, NER, and sentiment analysis, where we find concerning cases of stereotyping and cultural unfairness.We further test their text-infilling performance, revealing the incapability of appropriate adaptation to Arab cultural contexts.Finally, we analyze 6 Arabic pre-training corpora and find that commonly used sources such as Wikipedia may not be best suited to build culturally aware LMs, if used as they are without adjustment.We will make CAMeL publicly available at: https://github.com/tareknaous/camel Tarek Naous, Michael J. Ryan, Alan Ritter, Wei Xu 0004 |
ACL (1) | 4 |
| 2024 | InfoLossQA: Characterizing and Recovering Information Loss in Text SimplificationabstractJan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert, Kyle Lo, Wei Xu, Byron Wallace, Junyi Jessy Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert, Kyle Lo, Wei Xu 0004, Byron C. Wallace, Junyi Jessy Li |
ACL (1) | 6 |
| 2024 | NEO-BENCH: Evaluating Robustness of Large Language Models with NeologismsabstractThe performance of Large Language Models (LLMs) degrades from the temporal drift between data used for model training and newer text seen during inference.One understudied avenue of language change causing data drift is the emergence of neologisms -new word forms -over time.We create a diverse resource of recent English neologisms by using several popular collection methods.We analyze temporal drift using neologisms by comparing sentences containing new words with near-identical sentences that replace neologisms with existing substitute words.Model performance is nearly halved in machine translation when a single neologism is introduced in a sentence.Motivated by these results, we construct a benchmark to evaluate LLMs' ability to generalize to neologisms with various natural language understanding tasks and model perplexity.Models with later knowledge cutoff dates yield lower perplexities and perform better in downstream tasks.LLMs are also affected differently based on the linguistic origins of words, indicating that neologisms are complex for static LLMs to address.We will release our benchmark and code for reproducing our experiments. Oracle Ensemble Microsoft Bing Jonathan Zheng, Alan Ritter, Wei Xu 0004 |
ACL (1) | 3 |
| 2024 | Design and Evaluation of an Automatic Text Simplification Prototype with Deaf and Hard-of-hearing ReadersabstractResearch has observed benefits from providing lexical and syntactic approaches to Automatic Text Simplification (ATS) to Deaf and Hard-of-hearing (DHH) readers. However, little research has explored DHH readers’ design preferences and interactions with these approaches. This work first explores the design space of ATS systems with DHH readers, identifying potential design configurations for evaluation. Open-ended discussion of participants’ design preferences reveal values informing those preferences, including maintaining reading fluency and efficiency, and control over the tool. Using popular design choices from our formative study, we evaluated a prototype that provides various simplification types to explore DHH readers’ interactions with the system. We observed potential conflicts between participants’ values and design preferences, such as the prototype’s impact on participants’ reading speed and participants’ perceived need to reread simplifications suggested by the tool. However, participants found the tool useful, showing a nuanced preference towards world-level lexical simplifications using pop-ups. Our findings highlight the importance of the tool’s design on users’ reading experiences, and provide implications for the design and evaluation of ATS prototypes with target readers. Oliver Alonzo, Sooyeon Lee, Akhter Al Amin, Mounica Maddela, Wei Xu 0004, Matt Huenerfauth |
ASSETS | 5 |
| 2024 | Improving Minimum Bayes Risk Decoding with Multi-PromptabstractWhile instruction fine-tuned LLMs are effective text generators, sensitivity to prompt construction makes performance unstable and suboptimal in practice.Relying on a single 'best' prompt cannot capture all differing approaches to a generation problem.Using this observation, we propose multi-prompt decoding, where many candidate generations are decoded from a prompt bank at inference-time.To ensemble candidates, we use Minimum Bayes Risk (MBR) decoding, which selects a final output using a trained value metric.We show multiprompt improves MBR across a comprehensive set of conditional generation tasks (Figure 1), and show this is a result of estimating a more diverse and higher quality candidate space than that of a single prompt.Further experiments confirm multi-prompt improves generation across tasks, models and metrics. 1 David Heineman, Yao Dou, Wei Xu 0004 |
EMNLP | 3 |
| 2024 | MedReadMe: A Systematic Study for Fine-grained Sentence Readability in Medical DomainabstractMedical texts are notoriously challenging to read. Properly measuring their readability is the first step towards making them more accessible. In this paper, we present a systematic study on fine-grained readability measurements in the medical domain at both sentence-level and span-level. We introduce a new dataset MedReadMe, which consists of manually annotated readability ratings and fine-grained complex span annotation for 4,520 sentences, featuring two novel "Google-Easy" and "Google-Hard" categories. It supports our quantitative analysis, which covers 650 linguistic features and automatic complex word and jargon identification. Enabled by our high-quality annotation, we benchmark and improve several state-of-the-art sentence-level readability metrics for the medical domain specifically, which include unsupervised, supervised, and prompting-based methods using recently developed large language models (LLMs). Informed by our fine-grained complex span annotation, we find that adding a single feature, capturing the number of jargon spans, into existing readability formulas can significantly improve their correlation with human judgments. We will publicly release the dataset and code. Wei Xu 0004 |
EMNLP | 2 |
| 2024 | Granular Privacy Control for Geolocation with Vision Language ModelsabstractVision Language Models (VLMs) are rapidly advancing in their capability to answer information-seeking questions. As these models are widely deployed in consumer applications, they could lead to new privacy risks due to emergent abilities to identify people in photos, geolocate images, etc. As we demonstrate, somewhat surprisingly, current open-source and proprietary VLMs are very capable image geolocators, making widespread geolocation with VLMs an immediate privacy risk, rather than merely a theoretical future concern. As a first step to address this challenge, we develop a new benchmark, GPTGeoChat, to test the capability of VLMs to moderate geolocation dialogues with users. We collect a set of 1,000 image geolocation conversations between in-house annotators and GPT-4v, which are annotated with the granularity of location information revealed at each turn. Using this new dataset we evaluate the ability of various VLMs to moderate GPT-4v geolocation conversations by determining when too much location information has been revealed. We find that custom fine-tuned models perform on par with prompted API-based models when identifying leaked location information at the country or city level, however fine-tuning on supervised data appears to be needed to accurately moderate finer granularities, such as the name of a restaurant or building. Ethan Mendes, Yang Chen 0065, James Hays, Sauvik Das, Wei Xu 0004, Alan Ritter |
EMNLP | 5 |
| 2024 | ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability AssessmentabstractWe present a comprehensive evaluation of large language models for multilingual readability assessment. Existing evaluation resources lack domain and language diversity, limiting the ability for cross-domain and cross-lingual analyses. This paper introduces ReadMe++, a multilingual multi-domain dataset with human annotations of 9757 sentences in Arabic, English, French, Hindi, and Russian, collected from 112 different data sources. This benchmark will encourage research on developing robust multilingual readability assessment methods. Using ReadMe++, we benchmark multilingual and monolingual language models in the supervised, unsupervised, and few-shot prompting settings. The domain and language diversity in ReadMe++ enable us to test more effective few-shot prompting, and identify shortcomings in state-of-the-art unsupervised methods. Our experiments also reveal exciting results of superior domain generalization and enhanced cross-lingual transfer capabilities by models trained on ReadMe++. We will make our data publicly available and release a python package tool for multilingual sentence readability prediction using our trained models at: https://github.com/tareknaous/readme. Tarek Naous, Michael J. Ryan, Anton Lavrouk, Mohit Chandra, Wei Xu 0004 |
EMNLP | 5 |
| 2024 | GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-ExplanationabstractResearch on jailbreaking has been valuable for testing and understanding the safety and security issues of large language models (LLMs).In this paper, we introduce Iterative Refinement Induced Self-Jailbreak (IRIS), a novel approach that leverages the reflective capabilities of LLMs for jailbreaking with only blackbox access.Unlike previous methods, IRIS simplifies the jailbreaking process by using a single model as both the attacker and target.This method first iteratively refines adversarial prompts through self-explanation, which is crucial for ensuring that even well-aligned LLMs obey adversarial instructions.IRIS then rates and enhances the output given the refined prompt to increase its harmfulness.We find that IRIS achieves jailbreak success rates of 98% on GPT-4, 92% on GPT-4 Turbo, and 94% on Llama-3.1-70B in under 7 queries.It significantly outperforms prior approaches in automatic, black-box, and interpretable jailbreaking, while requiring substantially fewer queries, thereby establishing a new standard for interpretable jailbreaking methods. Govind Ramesh, Yao Dou, Wei Xu 0004 |
EMNLP | 3 |
| 2024 | Constrained Decoding for Cross-lingual Label ProjectionabstractZero-shot cross-lingual transfer utilizing multilingual LLMs has become a popular learning paradigm for low-resource languages with no labeled training data. However, for NLP tasks that involve fine-grained predictions on words and phrases, the performance of zero-shot cross-lingual transfer learning lags far behind supervised fine-tuning methods. Therefore, it is common to exploit translation and label projection to further improve the performance by (1) translating training data that is available in a high-resource language (e.g., English) together with the gold labels into low-resource languages, and/or (2) translating test data in low-resource languages to a high-source language to run inference on, then projecting the predicted span-level labels back onto the original test data. However, state-of-the-art marker-based label projection methods suffer from translation quality degradation due to the extra label markers injected in the input to the translation model. In this work, we explore a new direction that leverages constrained decoding for label projection to overcome the aforementioned issues. Our new method not only can preserve the quality of translated texts but also has the versatility of being applicable to both translating training and translating test data strategies. This versatility is crucial as our experiments reveal that translating test data can lead to a considerable boost in performance compared to translating only training data. We evaluate on two cross-lingual transfer tasks, namely Named Entity Recognition and Event Argument Extraction, spanning 20 languages. The results demonstrate that our approach outperforms the state-of-the-art marker-based method by a large margin and also shows better performance than other label projection methods that rely on external word alignment. Duong Minh Le, Yang Chen 0065, Alan Ritter, Wei Xu 0004 |
ICLR | 4 |
| 2023 | Distill or Annotate? Cost-Efficient Fine-Tuning of Compact ModelsabstractFine-tuning large models is highly effective, however, inference can be expensive and produces carbon emissions.Knowledge distillation has been shown to be a practical solution to reduce inference costs, but the distillation process itself requires significant computational resources.Rather than buying or renting GPUs to fine-tune, then distill a large model, an NLP practitioner might instead choose to allocate the available budget to hire annotators and manually label additional fine-tuning data.In this paper, we investigate how to most efficiently use a fixed budget to build a compact model.Through extensive experiments on six diverse tasks, we show that distilling from T5-XXL (11B) to T5-Small (60M) is almost always a cost-efficient strategy compared to annotating more data to directly train a compact model (T5-Small).We further investigate how the optimal budget allocated towards computation varies across scenarios.We will make our code, datasets, annotation cost estimates, and baseline models available as a benchmark to support further work on cost-efficient training of compact models. Junmo Kang, Wei Xu 0004, Alan Ritter |
ACL (1) | 2 |
| 2023 | Improved Instruction Ordering in Recipe-Grounded ConversationabstractIn this paper, we study the task of instructional dialogue and focus on the cooking domain.Analyzing the generated output of the GPT-J model, we reveal that the primary challenge for a recipe-grounded dialog system is how to provide the instructions in the correct order.We hypothesize that this is due to the model's lack of understanding of user intent and inability to track the instruction state (i.e., which step was last instructed).Therefore, we propose to explore two auxiliary subtasks, namely User Intent Detection and Instruction State Tracking, to support Response Generation with improved instruction grounding.Experimenting with our newly collected dataset, ChattyChef, shows that incorporating user intent and instruction state information helps the response generation model mitigate the incorrect order issue.Furthermore, to investigate whether Chat-GPT has completely solved this task, we analyze its outputs and find that it also makes mistakes (10.7% of the responses), about half of which are out-of-order instructions.We will release ChattyChef to facilitate further research in this area at: https://github. com/octaviaguo/ChattyChef. Duong Minh Le, Ruohao Guo, Wei Xu 0004, Alan Ritter |
ACL (1) | 3 |
| 2023 | LENS: A Learnable Evaluation Metric for Text SimplificationabstractTraining learnable metrics using modern language models has recently emerged as a promising method for the automatic evaluation of machine translation.However, existing human evaluation datasets for text simplification have limited annotations that are based on unitary or outdated models, making them unsuitable for this approach.To address these issues, we introduce the SIMPEVAL corpus that contains: SIMPEVAL PAST , comprising 12K human ratings on 2.4K simplifications of 24 past systems, and SIMPEVAL 2022 , a challenging simplification benchmark consisting of over 1K human ratings of 360 simplifications including GPT-3.5 generated text.Training on SIMPEVAL, we present LENS, a Learnable Evaluation Metric for Text Simplification.Extensive empirical results show that LENS correlates much better with human judgment than existing metrics, paving the way for future progress in the evaluation of text simplification.We also introduce RANK & RATE, a human evaluation framework that rates simplifications from several models in a list-wise manner using an interactive interface, which ensures both consistency and accuracy in the evaluation process and is used to create the SIMPEVAL datasets. Mounica Maddela, Yao Dou, David Heineman, Wei Xu 0004 |
ACL (1) | 4 |
| 2023 | Human-in-the-loop Evaluation for Early Misinformation Detection: A Case Study of COVID-19 TreatmentsabstractWe present a human-in-the-loop evaluation framework for fact-checking novel misinformation claims and identifying social media messages that support them.Our approach extracts check-worthy claims, which are aggregated and ranked for review.Stance classifiers are then used to identify tweets supporting novel misinformation claims, which are further reviewed to determine whether they violate relevant policies.To demonstrate the feasibility of our approach, we develop a baseline system based on modern NLP methods for human-in-the-loop fact-checking in the domain of COVID-19 treatments.We make our data 1 and detailed annotation guidelines available to support the evaluation of human-in-the-loop systems that identify novel misinformation directly from raw usergenerated content. Ethan Mendes, Yang Chen 0065, Wei Xu 0004, Alan Ritter |
ACL (1) | 3 |
| 2023 | Revisiting non-English Text Simplification: A Unified Multilingual BenchmarkabstractRecent advancements in high-quality, largescale English resources have pushed the frontier of English Automatic Text Simplification (ATS) research.However, less work has been done on multilingual text simplification due to the lack of a diverse evaluation benchmark that covers complex-simple sentence pairs in many languages.This paper introduces the MULTI-SIM benchmark, a collection of 27 resources in 12 distinct languages containing over 1.7 million complex-simple sentence pairs.This benchmark will encourage research in developing more effective multilingual text simplification models and evaluation metrics.Our experiments using MULTISIM with pre-trained multilingual language models reveal exciting performance improvements from multilingual training in non-English settings.We observe strong performance from Russian in zero-shot crosslingual transfer to low-resource languages.We further show that few-shot prompting with BLOOM-176b achieves comparable quality to reference simplifications outperforming finetuned models in most languages.We validate these findings through human evaluation. Michael J. Ryan, Tarek Naous, Wei Xu 0004 |
ACL (1) | 3 |
| 2023 | Dancing Between Success and Failure: Edit-level Simplification Evaluation using SALSAabstractLarge language models (e.g., GPT-4) are uniquely capable of producing highly rated text simplification, yet current human evaluation methods fail to provide a clear understanding of systems' specific strengths and weaknesses.To address this limitation, we introduce SALSA, an edit-based human annotation framework that enables holistic and fine-grained text simplification evaluation.We develop twenty one linguistically grounded edit types, covering the full spectrum of success and failure across dimensions of conceptual, syntactic and lexical simplicity.Using SALSA, we collect 19K edit annotations on 840 simplifications, revealing discrepancies in the distribution of simplification strategies performed by fine-tuned models, prompted LLMs and humans, and find GPT-3.5 performs more quality edits than humans, but still exhibits frequent errors.Using our finegrained annotations, we develop LENS-SALSA, a reference-free automatic simplification metric, trained to predict sentence-and word-level quality simultaneously.Additionally, we introduce word-level quality estimation for simplification and report promising baseline results.Our data, new metric, and annotation toolkit are available at https://salsa-eval.com. EXAMPLEZero-shot GPT-3. David Heineman, Yao Dou, Mounica Maddela, Wei Xu 0004 |
EMNLP | 4 |
| 2023 | Multilingual Simplification of Medical TextsabstractAutomated text simplification aims to produce simple versions of complex texts.This task is especially useful in the medical domain, where the latest medical findings are typically communicated via complex, technical articles.This creates barriers for laypeople seeking access to up-to-date medical findings, consequently impeding progress on health literacy.Most existing work on medical text simplification has focused on monolingual settings, with the result that such evidence would be available only in just one language (most often, English).This work addresses this limitation via multilingual simplification, i.e., directly simplifying complex texts into simplified texts in multiple languages.We introduce MULTICOCHRANE, the first sentence-aligned multilingual text simplification dataset for the medical domain in four languages: English, Spanish, French, and Farsi.We evaluate fine-tuned and zero-shot models across these languages with extensive human assessments and analyses.Although models can generate viable simplified texts, we identify several outstanding challenges that this dataset might be used to address. MultiCochrane Complex SimplePreclinical studies have suggested that RIC may have beneficial effects in ischaemic stroke patients and those at risk of ischaemic stroke.English: Studies have suggested that RIC may have beneficial effects for preventing and treating ischaemic stroke.Spanish: Los estudios han indicado que el CIR puede tener efectos beneficiosos en la prevención y el tratamiento del accidente cerebrovascular isquémico.French: Des études ont suggéré que le CID pourrait avoir des effets bénéfiques sur la prévention et le traitement de l'AVC ischémique.Farsi: ﮐﮫ اﻧد ﮐرده ﭘﯾﺷﻧﮭﺎد طﺎﻟﻌﺎت RIC درﻣﺎن و ﭘﯾﺷﮕﯾری ﺑرای ﻣﻔﯾدی اﺛرات اﺳت ﻣﻣﮑن ﺑﺎﺷد داﺷﺗﮫ اﯾﺳﮑﻣﯾﮏ ﻣﻐزی .ﺳﮑﺗﮫ Human evaluation of system outputsEnglish: Interventions have suggested that RAP may have beneficial effects in ischaemic stroke patients, those at risk of stroke.Spanish: Las intervenciones pueden ser efectivas para los pacientes que se accidente cerebrovascular isquémico y los que se encuentran en riesgo del accidente cerebrovascular isquémico.Gloss: The interventions can be effective for the patients that accident themselves ischemic stroke and those that find themselves at risk of the ischemic stroke. Sebastian Joseph, Kathryn Kazanas, Keziah Reina, Vishnesh J. Ramanathan, Wei Xu 0004, Byron C. Wallace, Junyi Jessy Li |
EMNLP | 5 |
| 2022 | Extracting a Knowledge Base of COVID-19 Events from Social MediaabstractWe present a manually annotated corpus of 10,000 tweets containing public reports of five COVID-19 events, including positive and negative tests, deaths, denied access to testing, claimed cures and preventions. We designed slot-filling questions for each event type and annotated a total of 28 fine-grained slots, such as the location of events, recent travel, and close contacts. We show that our corpus can support fine-tuning BERT-based classifiers to automatically extract publicly reported events, which can be further collected for building a knowledge base. Our knowledge base is constructed over Twitter data covering two years and currently covers over 4.2M events. It can answer complex queries with high precision, such as “Which organizations have employees that tested positive in Philadelphia?” We believe our proposed methodology could be quickly applied to develop knowledge bases for new domains in response to an emerging crisis, including natural disasters or future disease outbreaks. Shi Zong, Ashutosh Baheti, Wei Xu 0004, Alan Ritter |
COLING | 3 |
| 2022 | Improving Large-scale Paraphrase Acquisition and GenerationabstractThis paper addresses the quality issues in existing Twitter-based paraphrase datasets, and discusses the necessity of using two separate definitions of paraphrase for identification and generation tasks.We present a new Multi-Topic Paraphrase in Twitter (MULTIPIT) corpus that consists of a total of 130k sentence pairs with crowdsoursing (MULTIPIT CROWD ) and expert (MULTIPIT EXPERT ) annotations using two different paraphrase definitions for paraphrase identification, in addition to a multi-reference test set (MULTIPIT NMR ) and a large automatically constructed training set (MULTIPIT AUTO ) for paraphrase generation.With improved data annotation quality and task-specific paraphrase definition, the best pre-trained language model fine-tuned on our dataset achieves the stateof-the-art performance of 84.2 F 1 for automatic paraphrase identification.Furthermore, our empirical results also demonstrate that the paraphrase generation models trained on MUL-TIPIT AUTO generate more diverse and highquality paraphrases compared to their counterparts fine-tuned on other corpora such as Quora, MSCOCO, and ParaNMT.Topic Domains #Train #Dev #Test Sent/Tweet Len %Paraphrase #Trends/URLs #Uniq Sent %Multi-Ref Our Multi-Topic Paraphrase in Twitter (MULTIPIT CROWD ) Dataset Trends Sports 25,255 3,157 3,157 10.24 / 13.79 40.52% 1,201 34,786 17.89% Entertainment 11,547 1,443 1,444 10.44 / 13.80 62.33% 610 15,784 18.11% Event 8,624 1,078 1,079 10.86 / 15.32 82.83% 359 11,746 17.75% Others 17,751 2,219 2,219 10.41 / 14.56 67. Yao Dou, Wei Xu 0004 |
EMNLP | 3 |
| 2022 | arXivEdits: Understanding the Human Revision Process in Scientific WritingabstractScientific publications are the primary means to communicate research discoveries, where the writing quality is of crucial importance.However, prior work studying the human editing process in this domain mainly focused on the abstract or introduction sections, resulting in an incomplete picture.In this work, we provide a complete computational framework for studying text revision in scientific writing.We first introduce ARXIVEDITS, a new annotated corpus of 751 full papers from arXiv with gold sentence alignment across their multiple versions of revision, as well as fine-grained span-level edits and their underlying intentions for 1,000 sentence pairs.It supports our datadriven analysis to unveil the common strategies practiced by researchers for revising their papers.To scale up the analysis, we also develop automatic methods to extract revision at document-, sentence-, and word-levels.A neural CRF sentence alignment model trained on our corpus achieves 93.8 F1, enabling the reliable matching of sentences between different versions.We formulate the edit extraction task as a span alignment problem, and our proposed method extracts more fine-grained and explainable edits, compared to the commonly used diff algorithm.An intention classifier trained on our dataset achieves 78.9 F1 on the finegrained intent classification task.Our data and system are released at tiny.one/arxivedits. Wei Xu 0004, Samuel Stevens 0001 |
EMNLP | 2 |
| 2022 | Stanceosaurus: Classifying Stance Towards Multicultural MisinformationabstractWe present Stanceosaurus, a new corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims.As far as we are aware, it is the largest corpus annotated with stance towards misinformation claims.The claims in Stanceosaurus originate from 15 fact-checking sources that cover diverse geographical regions and cultures.Unlike existing stance datasets, we introduce a more fine-grained 5class labeling strategy with additional subcategories to distinguish implicit stance.Pretrained transformer-based stance classifiers that are fine-tuned on our corpus show good generalization on unseen claims and regional claims from countries outside the training data.Cross-lingual experiments demonstrate Stanceosaurus' capability of training multilingual models, achieving 53.1 F1 on Hindi and 50.4 F1 on Arabic without any targetlanguage fine-tuning.Finally, we show how a domain adaptation method can be used to improve performance on Stanceosaurus using additional RumourEval-2019 data.We make Stanceosaurus publicly available to the research community and hope it will encourage further work on misinformation identification across languages and cultures.1 Source Country & Regions Lang #Claims #Tweets Irr.Sup.Ref. Dis.Que.Snopes USA (80%), INT'L (16.7%),Other (3.3%) en 30 3197 1051 428 229 1447 42 Poynter Europe (5%), INT'L (90%), Other (5%) en 20 2197 949 274 97 844 33 FullFact UK (30%), INT'L (55%), Other (15%) en 20 2379 806 300 179 1057 37 AFP Fact Check CAN Canada (55%), INT'L (30%), Other (15%) en 20 2078 746 252 130 910 40 AAP Fact Check Australia (10%), INT'L (65%), Other (25%) en 20 2302 739 374 136 1019 34 AFP Fact Check NZ New Zealand (15%), INT'L (75%), Other (10%) en 20 2227 879 194 81 1044 29 Blackdotresearch Singapore (30%), INT'L (55%), Other (15%) en 20 2307 842 248 113 1076 28 Factly India (45%), INT'L (55%) en 20 1979 889 190 117 734 49 Politifact USA (20%), INT'L (35%), Other (45%) en 20 2041 984 289 8 753 7 Alt News India (90.4%),INT'L (4.8%),Other (4.8%) hi 21 1730 550 489 172 500 19 Aajtak India (67%), Other (33%) hi 9 806 456 110 40 193 7 Hindi Newschecker India (56%), Other (44%) hi 9 781 195 313 46 219 8 MISBAR Arab World (58.3%),INT'L (8.3%),Other (33.4%) ar 12 2283 454 514 203 1031 81 Fatabyyano Arab World (28.5%),INT'L (57.1%),Other (14.4%) ar 7 986 234 163 49 522 18 Maharat Fact-o-meter INT'L (100%) ar 3 740 Jonathan Zheng, Ashutosh Baheti, Tarek Naous, Wei Xu 0004, Alan Ritter |
EMNLP | 4 |
| 2021 | Neural semi-Markov CRF for Monolingual Word AlignmentabstractWuwei Lan, Chao Jiang, Wei Xu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wuwei Lan, Wei Xu 0004 |
ACL/IJCNLP (1) | 3 |
| 2021 | Pre-train or Annotate? Domain Adaptation with a Constrained BudgetabstractRecent work has demonstrated that pretraining in-domain language models can boost performance when adapting to a new domain.However, the costs associated with pretraining raise an important question: given a fixed budget, what steps should an NLP practitioner take to maximize performance?In this paper, we view domain adaptation with a constrained budget as a consumer choice problem, where the goal is to select an optimal combination of data annotation and pre-training.We measure annotation costs of three procedural text datasets, along with the pre-training costs of several in-domain language models.The utility of different combinations of pretraining and data annotation are evaluated under varying budget constraints to assess which combination strategy works best.We find that for small budgets, spending all funds on annotation leads to the best performance; once the budget becomes large enough, however, a combination of data annotation and in-domain pre-training yields better performance.Our experiments suggest task-specific data annotation should be part of an economical strategy when adapting an NLP model to a new domain.1 Fan Bai 0006, Alan Ritter, Wei Xu 0004 |
EMNLP (1) | 3 |
| 2021 | BiSECT: Learning to Split and Rephrase Sentences with BitextsabstractAn important task in NLP applications such as sentence simplification is the ability to take a long, complex sentence and split it into shorter sentences, rephrasing as necessary.We introduce a novel dataset and a new model for this 'split and rephrase' task.Our BISECT training data consists of 1 million long English sentences paired with shorter, meaning-equivalent English sentences.We obtain these by extracting 1-2 sentence alignments in bilingual parallel corpora and then using machine translation to convert both sides of the corpus into the same language.BISECT contains higher quality training examples than previous Split and Rephrase corpora, with sentence splits that require more significant modifications.We categorize examples in our corpus, and use these categories in a novel model that allows us to target specific regions of the input sentence to be split and edited.Moreover, we show that models trained on BISECT can perform a wider variety of split operations and improve upon previous state-of-the-art approaches in automatic and human evaluations.1 Joongwon Kim, Mounica Maddela, Reno Kriz, Wei Xu 0004, Chris Callison-Burch |
EMNLP (1) | 4 |
| 2021 | Controllable Text Simplification with Explicit ParaphrasingabstractText Simplification improves the readability of sentences through several rewriting transformations, such as lexical paraphrasing, deletion, and splitting.Current simplification systems are predominantly sequence-to-sequence models that are trained end-to-end to perform all these operations simultaneously.However, such systems limit themselves to mostly deleting words and cannot easily adapt to the requirements of different target audiences.In this paper, we propose a novel hybrid approach that leverages linguistically-motivated rules for splitting and deletion, and couples them with a neural paraphrasing model to produce varied rewriting styles.We introduce a new data augmentation method to improve the paraphrasing capability of our model.Through automatic and manual evaluations, we show that our proposed model establishes a new state-ofthe art for the task, paraphrasing more often than the existing systems, and can control the degree of each simplification operation applied to the input texts. 1 Mounica Maddela, Fernando Alva-Manchego, Wei Xu 0004 |
NAACL-HLT | 3 |
| 2020 | Discourse Level Factors for Sentence Deletion in Text SimplificationabstractThis paper presents a data-driven study focusing on analyzing and predicting sentence deletion — a prevalent but understudied phenomenon in document simplification — on a large English text simplification corpus. We inspect various document and discourse factors associated with sentence deletion, using a new manually annotated sentence alignment corpus we collected. We reveal that professional editors utilize different strategies to meet readability standards of elementary and middle schools. To predict whether a sentence will be deleted during simplification to a certain level, we harness automatically aligned data to train a classification model. Evaluated on our manually annotated data, our best models reached F1 scores of 65.2 and 59.7 for this task at the levels of elementary and middle school, respectively. We find that discourse level factors contribute to the challenging task of predicting sentence deletion for simplification. Wei Xu 0004, Junyi Jessy Li |
AAAI | 3 |
| 2020 | Neural CRF Model for Sentence Alignment in Text SimplificationabstractThe success of a text simplification system heavily depends on the quality and quantity of complex-simple sentence pairs in the training corpus, which are extracted by aligning sentences between parallel articles.To evaluate and improve sentence alignment quality, we create two manually annotated sentence-aligned datasets from two commonly used text simplification corpora, Newsela and Wikipedia.We propose a novel neural CRF alignment model which not only leverages the sequential nature of sentences in parallel documents but also utilizes a neural sentence pair model to capture semantic similarity.Experiments demonstrate that our proposed approach outperforms all the previous work on monolingual sentence alignment task by more than 5 points in F1.We apply our CRF aligner to construct two new text simplification datasets, NEWSELA-AUTO and WIKI-AUTO, which are much larger and of better quality compared to the existing datasets.A Transformer-based seq2seq model trained on our datasets establishes a new state-of-the-art for text simplification in both automatic and human evaluation. 1 Mounica Maddela, Wuwei Lan, Wei Xu 0004 |
ACL | 5 |
| 2020 | Generalizing Natural Language Analysis through Span-relation RepresentationsabstractNatural language processing covers a wide variety of tasks predicting syntax, semantics, and information content, and usually each type of output is generated with specially designed architectures.In this paper, we provide the simple insight that a great variety of tasks can be represented in a single unified format consisting of labeling spans and relations between spans, thus a single task-independent model can be used across different tasks.We perform extensive experiments to test this insight on 10 disparate tasks spanning dependency parsing (syntax), semantic role labeling (semantics), relation extraction (information content), aspect based sentiment analysis (sentiment), and many others, achieving performance comparable to state-of-the-art specialized models.We further demonstrate benefits of multi-task learning, and also show that the proposed method makes it easy to analyze differences and similarities in how the model handles different tasks.Finally, we convert these datasets into a unified format to build a benchmark, which provides a holistic testbed for evaluating future models for generalized natural language analysis. Zhengbao Jiang, Wei Xu 0004, Jun Araki, Graham Neubig |
ACL | 2 |
| 2020 | Code and Named Entity Recognition in StackOverflowabstractThere is an increasing interest in studying natural language and computer code together, as large corpora of programming texts become readily available on the Internet.For example, StackOverflow currently has over 15 million programming related questions written by 8.5 million users.Meanwhile, there is still a lack of fundamental NLP techniques for identifying code tokens or software-related named entities that appear within natural language sentences.In this paper, we introduce a new named entity recognition (NER) corpus for the computer programming domain, consisting of 15,372 sentences annotated with 20 fine-grained entity types.We trained indomain BERT representations (BERTOverflow) on 152 million sentences from Stack-Overflow, which lead to an absolute increase of +10 F 1 score over off-the-shelf BERT.We also present the SoftNER model which achieves an overall 79.10 F 1 score for code and named entity recognition on StackOverflow data.Our SoftNER model incorporates a context-independent code token classifier with corpus-level features to improve the BERTbased tagging model. 1 Jeniya Tabassum, Mounica Maddela, Wei Xu 0004, Alan Ritter |
ACL | 3 |
| 2020 | An Empirical Study of Pre-trained Transformers for Arabic Information ExtractionabstractMultilingual pre-trained Transformers, such as mBERT (Devlin et al., 2019) and XLM-RoBERTa (Conneau et al., 2020a), have been shown to enable the effective cross-lingual zero-shot transfer.However, their performance on Arabic information extraction (IE) tasks is not very well studied.In this paper, we pre-train a customized bilingual BERT, dubbed GigaBERT, that is designed specifically for Arabic NLP and English-to-Arabic zero-shot transfer learning.We study Giga-BERT's effectiveness on zero-short transfer across four IE tasks: named entity recognition, part-of-speech tagging, argument role labeling, and relation extraction.Our best model significantly outperforms mBERT, XLM-RoBERTa, and AraBERT (Antoun et al., 2020) in both the supervised and zero-shot transfer settings.We have made our pre-trained models publicly available at https://github.com Wuwei Lan, Yang Chen 0065, Wei Xu 0004, Alan Ritter |
EMNLP (1) | 3 |
| 2019 | Multi-task Pairwise Neural Ranking for Hashtag SegmentationabstractHashtags are often employed on social media and beyond to add metadata to a textual utterance with the goal of increasing discoverability, aiding search, or providing additional semantics.However, the semantic content of hashtags is not straightforward to infer as these represent ad-hoc conventions which frequently include multiple words joined together and can include abbreviations and unorthodox spellings.We build a dataset of 12,594 hashtags split into individual segments and propose a set of approaches for hashtag segmentation by framing it as a pairwise ranking problem between candidate segmentations. 1 Our novel neural approaches demonstrate 24.6% error reduction in hashtag segmentation accuracy compared to the current state-of-the-art method.Finally, we demonstrate that a deeper understanding of hashtag semantics obtained through segmentation is useful for downstream applications such as sentiment analysis, for which we achieved a 2.6% increase in average recall on the Se-mEval 2017 sentiment analysis dataset. Mounica Maddela, Wei Xu 0004, Daniel Preotiuc-Pietro |
ACL (1) | 2 |
| 2018 | Neural Network Models for Paraphrase Identification, Semantic Textual Similarity, Natural Language Inference, and Question AnsweringabstractIn this paper, we analyze several neural network designs (and their variations) for sentence pair modeling and compare their performance extensively across eight datasets, including paraphrase identification, semantic textual similarity, natural language inference, and question answering tasks. Although most of these models have claimed state-of-the-art performance, the original papers often reported on only one or two selected datasets. We provide a systematic study and show that (i) encoding contextual information by LSTM and inter-sentence interactions are critical, (ii) Tree-LSTM does not help as much as previously claimed but surprisingly improves performance on Twitter datasets, (iii) the Enhanced Sequential Inference Model is the best so far for larger datasets, while the Pairwise Word Interaction Model achieves the best performance when less data is available. We release our implementations as an open-source toolkit. Wuwei Lan, Wei Xu 0004 |
COLING | 2 |
| 2018 | A Word-Complexity Lexicon and A Neural Readability Ranking Model for Lexical SimplificationabstractCurrent lexical simplification approaches rely heavily on heuristics and corpus level features that do not always align with human judgment.We create a human-rated wordcomplexity lexicon of 15,000 English words and propose a novel neural readability ranking model with a Gaussian-based feature vectorization layer that utilizes these human ratings to measure the complexity of any given word or phrase.Our model performs better than the state-of-the-art systems for different lexical simplification tasks and evaluation datasets.Additionally, we also produce SimplePPDB++, a lexical resource of over 10 million simplifying paraphrase rules, by applying our model to the Paraphrase Database (PPDB). 1 Mounica Maddela, Wei Xu 0004 |
EMNLP | 2 |
| 2017 | A Continuously Growing Dataset of Sentential ParaphrasesabstractA major challenge in paraphrase research is the lack of parallel corpora.In this paper, we present a new method to collect large-scale sentential paraphrases from Twitter by linking tweets through shared URLs.The main advantage of our method is its simplicity, as it gets rid of the classifier or human in the loop needed to select data before annotation and subsequent application of paraphrase identification algorithms in the previous work.We present the largest human-labeled paraphrase corpus to date of 51,524 sentence pairs and the first cross-domain benchmarking for automatic paraphrase identification.In addition, we show that more than 30,000 new sentential paraphrases can be easily and continuously captured every month at ∼70% precision, and demonstrate their utility for downstream NLP tasks through phrasal paraphrase extraction.We make our code and data freely available.1 Wuwei Lan, Siyu Qiu, Wei Xu 0004 |
EMNLP | 4 |
| 2016 | Discovering User Attribute Stylistic Differences via ParaphrasingabstractUser attribute prediction from social media text has proven successful and useful for downstream tasks. In previous studies, differences in user trait language use have been limited primarily to the presence or absence of words that indicate topical preferences. In this study, we aim to find linguistic style distinctions across three different user attributes: gender, age and occupational class. By combining paraphrases with a simple yet effective method, we capture a wide set of stylistic differences that are exempt from topic bias. We show their predictive power in user profiling, conformity with human perception and psycholinguistic hypotheses, and potential use in generating natural language tailored to specific user traits. Daniel Preotiuc-Pietro, Wei Xu 0004, Lyle H. Ungar |
AAAI | 2 |
| 2016 | TweeTime : A Minimally Supervised Method for Recognizing and Normalizing Time Expressions in TwitterabstractWe describe TweeTIME, a temporal tagger for recognizing and normalizing time expressions in Twitter.Most previous work in social media analysis has to rely on temporal resolvers that are designed for well-edited text, and therefore suffer from reduced performance due to domain mismatch.We present a minimally supervised method that learns from large quantities of unlabeled data and requires no hand-engineered rules or hand-annotated training corpora.TweeTIME achieves 0.68 F1 score on the end-to-end task of resolving date expressions, outperforming a broad range of state-of-the-art systems. 1 Jeniya Tabassum, Alan Ritter, Wei Xu 0004 |
EMNLP | 3 |
| 2016 | Optimizing Statistical Machine Translation for Text SimplificationabstractMost recent sentence simplification systems use basic machine translation models to learn lexical and syntactic paraphrases from a manually simplified parallel corpus. These methods are limited by the quality and quantity of manually simplified corpora, which are expensive to build. In this paper, we conduct an in-depth adaptation of statistical machine translation to perform text simplification, taking advantage of large-scale paraphrases learned from bilingual texts and a small amount of manual simplifications with multiple references. Our work is the first to design automatic metrics that are effective for tuning and evaluating simplification systems, which will facilitate iterative development for this task. Wei Xu 0004, Courtney Napoles, Ellie Pavlick, Quanze Chen, Chris Callison-Burch |
Trans. Assoc. Comput. Linguistics | 1 |
| 2015 | Cost Optimization in Crowdsourcing Translation: Low cost translations made even cheaper
Mingkun Gao, Wei Xu 0004, Chris Callison-Burch |
HLT-NAACL | 2 |
| 2015 | Problems in Current Text Simplification Research: New Data Can HelpabstractSimple Wikipedia has dominated simplification research in the past 5 years. In this opinion paper, we argue that focusing on Wikipedia limits simplification research. We back up our arguments with corpus analysis and by highlighting statements that other researchers have made in the simplification literature. We introduce a new simplification dataset that is a significant improvement over Simple Wikipedia, and present a novel quantitative-comparative approach to study the quality of simplification data resources. Wei Xu 0004, Chris Callison-Burch, Courtney Napoles |
Trans. Assoc. Comput. Linguistics | 1 |
| 2014 | Poetry of the Crowd: A Human Computation Algorithm to Convert Prose into Rhyming VerseabstractPoetry composition is a very complex task that requires a poet to satisfy multiple constraints concurrently. We believe that the task can be augmented by combining the creative abilities of humans with computational algorithms that efficiently constrain and permute available choices. We present a hybrid method for generating poetry from prose that combines crowdsourcing with natural language processing (NLP) machinery. We test the ability of crowd workers to accomplish the technically challenging and creative task of composing poems. Quanze Chen, Chenyang Lei, Wei Xu 0004, Ellie Pavlick, Chris Callison-Burch |
HCOMP | 3 |
| 2014 | Extracting Lexically Divergent Paraphrases from TwitterabstractWe present MultiP (Multi-instance Learning Paraphrase Model), a new model suited to identify paraphrases within the short messages on Twitter. We jointly model paraphrase relations between word and sentence pairs and assume only sentence-level annotations during learning. Using this principled latent variable model alone, we achieve the performance competitive with a state-of-the-art method which combines a latent space model with a feature-based supervised classifier. Our model also captures lexically divergent paraphrases that differ from yet complement previous methods; combining our model with previous work significantly outperforms the state-of-the-art. In addition, we present a novel annotation methodology that has allowed us to crowdsource a paraphrase corpus from Twitter. We make this new dataset available to the research community. Wei Xu 0004, Alan Ritter, Chris Callison-Burch, William B. Dolan, Yangfeng Ji |
Trans. Assoc. Comput. Linguistics | 1 |
| 2012 | Paraphrasing for Style
Wei Xu 0004, Alan Ritter, William B. Dolan, Ralph Grishman, Colin Cherry |
COLING | 1 |
| 2011 | Exploiting Syntactic and Distributional Information for Spelling Correction with Web-Scale N-gram Models
Wei Xu 0004, Joel R. Tetreault, Martin Chodorow, Ralph Grishman |
EMNLP | 1 |
| 2011 | Passage Retrieval for Information Extraction using Distant Supervision
Wei Xu 0004, Ralph Grishman |
IJCNLP | 1 |
| 2009 | Who, What, When, Where, Why? Comparing Multiple Approaches to the Cross-Lingual 5W Task
Kristen Parton, Kathy McKeown, Bob Coyne, Mona T. Diab, Ralph Grishman, Dilek Hakkani-Tür, Mary P. Harper, Heng Ji 0001, Wei-Yun Ma, Adam Meyers 0001, Sara Stolbach, Ang Sun, Gökhan Tür, Wei Xu 0004, Sibel Yaman |
ACL/IJCNLP | 14 |
| 2007 | Using Non-Local Features to Improve Named Entity Recognition Recall
Xinnian Mao, Wei Xu 0004, Saike He, Haila Wang |
PACLIC | 2 |
| 2006 | Extractive Summarization using Inter- and Intra- Event RelevanceabstractEvent-based summarization attempts to select and organize the sentences in a summary with respect to the events or the sub-events that the sentences describe. Each event has its own internal structure, and meanwhile often relates to other events semantically, temporally, spatially, causally or conditionally. In this paper, we define an event as one or more event terms along with the named entities associated, and present a novel approach to derive intra- and inter- event relevance using the information of internal association, semantic relatedness, distributional similarity and named entity clustering. We then apply PageRank ranking algorithm to estimate the significance of an event for inclusion in a summary from the event relevance derived. Experiments on the DUC 2001 test data shows that the relevance of the named entities involved in events achieves better result when their relevance is derived from the event terms they associate. It also reveals that the topic-specific relevance from documents themselves outperforms the semantic relevance from a general purpose knowledge base like Word-Net. Wenjie Li 0002, Qin Lu 0001, Wei Xu 0004, Chunfa Yuan |
ACL | 4 |
| 2006 | Deriving Event Relevance from the Ontology Constructed with Formal Concept Analysis
Wei Xu 0004, Wenjie Li 0002, Wei Li 0051, Chunfa Yuan |
CICLing | 1 |