VLDB 2026 Research / reviewers in the wild / expert
Junyi Jessy Li
dblp:148/9553
· DBLP profile ↗
62ranked-venue papers
9as first author
34since 2021 · last 2026
0000-0002-2550-5262ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 9 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving the Distributional Alignment of LLMs using SupervisionabstractGauri Kambhatla, Sanjana Gautam, Angela Zhang, Alexander Liu, Ravi Srinivasan, Junyi Jessy Li, Matthew Lease. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Gauri Kambhatla, Sanjana Gautam, Angela Zhang, Ravi Srinivasan, Junyi Jessy Li, Matthew Lease |
ACL (1) | 6 |
| 2025 | SPRI: Aligning Large Language Models with Context-Situated PrinciplesabstractAligning Large Language Models to integrate and reflect human values, especially for tasks that demand intricate human oversight, is arduous since it is resource-intensive and time-consuming to depend on human expertise for context-specific guidance. Prior work has utilized predefined sets of rules or principles to steer the behavior of models (Bai et al., 2022; Sun et al., 2023). However, these principles tend to be generic, making it challenging to adapt them to each individual input query or context. In this work, we present Situated-PRInciples (SPRI), a framework requiring minimal or no human effort that is designed to automatically generate guiding principles in real-time for each input query and utilize them to align each response. We evaluate SPRI on three tasks, and show that 1) SPRI can derive principles in a complex domain-specific task that leads to on-par performance as expert-crafted ones; 2) SPRI-generated principles lead to instance-specific rubrics that outperform prior LLM-as-a-judge frameworks; 3) using SPRI to generate synthetic SFT data leads to substantial improvement on truthfulness. Hongli Zhan, Muneeza Azmat, Raya Horesh, Junyi Jessy Li, Mikhail Yurochkin |
ICML | 4 |
| 2025 | exLong: Generating Exceptional Behavior Tests with Large Language ModelsabstractMany popular programming languages, including C#, Java, and Python, support exceptions. Exceptions are thrown during program execution if an unwanted event happens, e.g., a method is invoked with an illegal argument value. Software developers write exceptional behavior tests (EBTs) to check that their code detects unwanted events and throws appropriate exceptions. Prior research studies have shown the importance of EBTs, but those studies also highlighted that developers put most of their efforts on “happy paths”, e.g., paths without unwanted events. To help developers fill the gap, we present the first framework, dubbed ExLoNG, that automatically generates EBTs. ExLONG is a large language model instruction fine-tuned from CodeLlama and embeds reasoning about traces that lead to throw statements, conditional expressions that guard throw statements, and non-exceptional behavior tests that execute similar traces. We compare ExLONG with the state-of-the-art models for test generation (CAT-LM) and one of the strongest foundation models (GPT-4o), as well as with analysis-based tools for test generation (Randoop and EvoSuite). Our results show that ExLONG outperforms existing models and tools. Furthermore, we contributed several pull requests to open-source projects and 23 EBTs generated by ExLONG were already accepted. Jiyang Zhang 0003, Yu Liu 0079, Pengyu Nie 0001, Junyi Jessy Li, Milos Gligoric 0001 |
ICSE | 4 |
| 2025 | AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in AstronomyabstractLarge Language Models (LLMs) are being explored for applications in scientific research, including their capabilities to synthesize literature, answer research questions, generate research ideas, and even conduct computational experiments.Ultimately, our goal is for these to help scientists derive novel scientific insights. In many areas of science, such insights often arise from processing and visualizing data to understand its patterns. However, evaluating whether an LLM-mediated scientific workflow produces outputs conveying the correct scientific insights is challenging to evaluate and has not been addressed in past work.We introduce AstroVisBench, the first benchmark for both scientific computing and visualization in the astronomy domain.AstroVisBench judges a language model’s ability to both (1) create astronomy-specific workflows to process and analyze data and (2) visualize the results of these workflows through complex plots.Our evaluation of visualizations uses a novel LLM-as-a-judge workflow, which is validated against annotation by five professional astronomers.Using AstroVisBench we present an evaluation of state-of-the-art language models, showing a significant gap in their ability to engage in astronomy research as useful assistants.This evaluation provides a strong end-to-end evaluation for AI scientists that offers a path forward for the development of visualization-based workflows, which are central to a broad range of domains from physics to biology. Sebastian Joseph, Syed Murtaza Husain, Stella S. R. Offner, Stéphanie Juneau, Paul Torrey, Adam S. Bolton, Juan P. Farias, Niall Gaffney, Greg Durrett, Junyi Jessy Li |
NeurIPS | 10 |
| 2024 | Large Language Models Produce Responses Perceived to be EmpathicabstractLarge Language Models (LLMs) have demonstrated surprising performance on many tasks, including writing supportive messages that display empathy. Here, we had these models generate empathic messages in response to posts describing common life experiences, such as workplace situations, parenting, relationships, and other anxiety- and anger-eliciting situations. Across two studies (N=192, 202), we showed human raters a variety of responses written by several models (GPT4 Turbo, Llama2, and Mistral), and had people rate these responses on how empathic they seemed to be. We found that LLM-generated responses were consistently rated as more empathic than human-written responses. Linguistic analyses also show that these models write in distinct, predictable “styles”, in terms of their use of punctuation, emojis, and certain words. These results highlight the potential of using LLMs to enhance human peer support in contexts where empathy is important. Yoon Kyung Lee, Jina Suh, Hongli Zhan, Junyi Jessy Li, Desmond C. Ong |
ACII | 4 |
| 2024 | FactPICO: Factuality Evaluation for Plain Language Summarization of Medical EvidenceabstractSebastian Joseph, Lily Chen, Jan Trienes, Hannah Göke, Monika Coers, Wei Xu, Byron Wallace, Junyi Jessy Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Sebastian Joseph, Lily Chen, Jan Trienes, Hannah Louisa Göke, Monika Coers, Wei Xu 0004, Byron C. Wallace, Junyi Jessy Li |
ACL (1) | 8 |
| 2024 | InfoLossQA: Characterizing and Recovering Information Loss in Text SimplificationabstractJan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert, Kyle Lo, Wei Xu, Byron Wallace, Junyi Jessy Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert, Kyle Lo, Wei Xu 0004, Byron C. Wallace, Junyi Jessy Li |
ACL (1) | 8 |
| 2024 | Sarcasm Detection in a Disaster ContextabstractDuring natural disasters, people often use social media platforms such as Twitter to ask for help, to provide information about the disaster situation, or to express contempt about the unfolding event or public policies and guidelines. This contempt is in some cases expressed as sarcasm or irony. Understanding this form of speech in a disaster-centric context is essential to improving natural language understanding of disaster-related tweets. In this paper, we introduce HurricaneSARC, a dataset of 15,000 tweets annotated for intended sarcasm, and provide a comprehensive investigation of sarcasm detection using pre-trained language models. Our best model is able to obtain as much as 0.70 F1 on our dataset. We also demonstrate that the performance on HurricaneSARC can be improved by leveraging intermediate task transfer learning Tiberiu Sosea, Junyi Jessy Li, Cornelia Caragea |
LREC/COLING | 2 |
| 2024 | Detection and Measurement of Syntactic Templates in Generated TextabstractThe diversity of text can be measured beyond word-level features, however existing diversity evaluation focuses primarily on word-level features.Here we propose a method for evaluating diversity over syntactic features to characterize general repetition in models, beyond frequent n-grams.Specifically, we define syntactic templates (e.g., strings comprising parts-of-speech) and show that models tend to produce templated text in downstream tasks at a higher rate than what is found in human-reference texts We find that most (76%) templates in modelgenerated text can be found in pre-training data (compared to only 35% of human-authored text), and are not overwritten during fine-tuning or alignment processes such as RLHF.The connection between templates in generated text and the pre-training data allows us to analyze syntactic templates in models where we do not have the pre-training data.We also find that templates as features are able to differentiate between models, tasks, and domains, and are useful for qualitatively evaluating common model constructions.Finally, we demonstrate the use of templates as a useful tool for analyzing style memorization of training data in LLMs 1 . Chantal Shaib, Yanai Elazar, Junyi Jessy Li, Byron C. Wallace |
EMNLP | 3 |
| 2024 | Which questions should I answer? Salience Prediction of Inquisitive QuestionsabstractInquisitive questions -open-ended, curiositydriven questions people ask as they read -are an integral part of discourse processing (Van Kuppevelt, 1995;Onea, 2016;Kehler and Rohde, 2017) and comprehension (Prince, 2004).Recent work in NLP has taken advantage of question generation capabilities of LLMs to enhance a wide range of applications.But the space of inquisitive questions is vast: many potential questions can be evoked from a given context.So which of those should be prioritized to find answers?Linguistic theories, unfortunately, have not yet provided an answer.This paper presents QSALIENCE, a salience predictor of inquisitive questions.QSALIENCE is instruction-tuned over our dataset of linguistannotated salience scores of 1,766 (context, question) pairs.A question scores high on salience if answering it would greatly enhance the understanding of the text (Van Rooy, 2003).We show that highly salient questions are empirically more likely to be answered in the same article, bridging potential questions (Onea, 2016) with Questions Under Discussion (Roberts, 2012).We further validate our findings by showing that answering salient questions is an indicator of summarization quality in news. Yating Wu 0002, Ritika Mangla, Alexandros G. Dimakis, Greg Durrett, Junyi Jessy Li |
EMNLP | 5 |
| 2023 | Unsupervised Extractive Summarization of Emotion TriggersabstractUnderstanding what leads to emotions during large-scale crises is important as it can provide groundings for expressed emotions and subsequently improve the understanding of ongoing disasters.Recent approaches (Zhan et al., 2022) trained supervised models to both detect emotions and explain emotion triggers (events and appraisals) via abstractive summarization.However, obtaining timely and qualitative abstractive summaries is expensive and extremely time-consuming, requiring highlytrained expert annotators.In time-sensitive, high-stake contexts, this can block necessary responses.We instead pursue unsupervised systems that extract triggers from text.First, we introduce COVIDET-EXT, augmenting (Zhan et al., 2022)'s abstractive dataset (in the context of the COVID-19 crisis) with extractive triggers.Second, we develop new unsupervised learning models that can jointly detect emotions and summarize their triggers.Our best approach, entitled Emotion-Aware Pagerank, incorporates emotion information from external sources combined with a language understanding module, and outperforms strong baselines.We release our data and code at Tiberiu Sosea, Hongli Zhan, Junyi Jessy Li, Cornelia Caragea |
ACL (1) | 3 |
| 2023 | How people talk about each other: Modeling Generalized Intergroup Bias and EmotionabstractVenkata Subrahmanyan Govindarajan, Katherine Atwell, Barea Sinno, Malihe Alikhani, David I. Beaver, Junyi Jessy Li. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Venkata Subrahmanyan Govindarajan, Katherine Atwell, Barea Sinno, Malihe Alikhani, David Beaver, Junyi Jessy Li |
EACL | 6 |
| 2023 | Multilingual Simplification of Medical TextsabstractAutomated text simplification aims to produce simple versions of complex texts.This task is especially useful in the medical domain, where the latest medical findings are typically communicated via complex, technical articles.This creates barriers for laypeople seeking access to up-to-date medical findings, consequently impeding progress on health literacy.Most existing work on medical text simplification has focused on monolingual settings, with the result that such evidence would be available only in just one language (most often, English).This work addresses this limitation via multilingual simplification, i.e., directly simplifying complex texts into simplified texts in multiple languages.We introduce MULTICOCHRANE, the first sentence-aligned multilingual text simplification dataset for the medical domain in four languages: English, Spanish, French, and Farsi.We evaluate fine-tuned and zero-shot models across these languages with extensive human assessments and analyses.Although models can generate viable simplified texts, we identify several outstanding challenges that this dataset might be used to address. MultiCochrane Complex SimplePreclinical studies have suggested that RIC may have beneficial effects in ischaemic stroke patients and those at risk of ischaemic stroke.English: Studies have suggested that RIC may have beneficial effects for preventing and treating ischaemic stroke.Spanish: Los estudios han indicado que el CIR puede tener efectos beneficiosos en la prevención y el tratamiento del accidente cerebrovascular isquémico.French: Des études ont suggéré que le CID pourrait avoir des effets bénéfiques sur la prévention et le traitement de l'AVC ischémique.Farsi: ﮐﮫ اﻧد ﮐرده ﭘﯾﺷﻧﮭﺎد طﺎﻟﻌﺎت RIC درﻣﺎن و ﭘﯾﺷﮕﯾری ﺑرای ﻣﻔﯾدی اﺛرات اﺳت ﻣﻣﮑن ﺑﺎﺷد داﺷﺗﮫ اﯾﺳﮑﻣﯾﮏ ﻣﻐزی .ﺳﮑﺗﮫ Human evaluation of system outputsEnglish: Interventions have suggested that RAP may have beneficial effects in ischaemic stroke patients, those at risk of stroke.Spanish: Las intervenciones pueden ser efectivas para los pacientes que se accidente cerebrovascular isquémico y los que se encuentran en riesgo del accidente cerebrovascular isquémico.Gloss: The interventions can be effective for the patients that accident themselves ischemic stroke and those that find themselves at risk of the ischemic stroke. Sebastian Joseph, Kathryn Kazanas, Keziah Reina, Vishnesh J. Ramanathan, Wei Xu 0004, Byron C. Wallace, Junyi Jessy Li |
EMNLP | 7 |
| 2023 | QUDeval: The Evaluation of Questions Under Discussion Discourse ParsingabstractQuestions Under Discussion (QUD) is a versatile linguistic framework in which discourse progresses as continuously asking questions and answering them.Automatic parsing of a discourse to produce a QUD structure thus entails a complex question generation task: given a document and an answer sentence, generate a question that satisfies linguistic constraints of QUD and can be grounded in an anchor sentence in prior context.These questions are known to be curiosity-driven and open-ended.This work introduces the first framework for the automatic evaluation of QUD parsing, instantiating the theoretical constraints of QUD in a concrete protocol.We present QUDE-VAL, a dataset of fine-grained evaluation of 2,190 QUD questions generated from both finetuned systems and LLMs.Using QUDEVAL, we show that satisfying all constraints of QUD is still challenging for modern LLMs, and that existing evaluation metrics poorly approximate parser quality.Encouragingly, human-authored QUDs are scored highly by our human evaluators, suggesting that there is headroom for further progress on language modeling to improve both QUD parsing and QUD evaluation. Yating Wu 0002, Ritika Mangla, Greg Durrett, Junyi Jessy Li |
EMNLP | 4 |
| 2023 | Elaborative Simplification as Implicit Questions Under DiscussionabstractAutomated text simplification, a technique useful for making text more accessible to people such as children and emergent bilinguals, is often thought of as a monolingual translation task from complex to simplified text.This view fails to account for elaborative simplification, where new information is added into the simplified text.This paper proposes to view elaborative simplification through the lens of the Question Under Discussion (QUD) framework, providing a robust way to investigate what writers elaborate upon, how they elaborate, and how elaborations fit into the discourse context by viewing elaborations as explicit answers to implicit questions.We introduce ELABQUD, consisting of 1.3K elaborations accompanied with implicit QUDs, to study these phenomena.We show that explicitly modeling QUD (via question generation) not only provides essential understanding of elaborative simplification and how the elaborations connect with the rest of the discourse, but also substantially improves the quality of elaboration generation. Yating Wu 0002, William Sheffield, Kyle Mahowald, Junyi Jessy Li |
EMNLP | 4 |
| 2023 | Learning Deep Semantics for Test CompletionabstractWriting tests is a time-consuming yet essential task during software development. We propose to leverage recent advances in deep learning for text and code generation to assist developers in writing tests. We formalize the novel task of test completion to automatically complete the next statement in a test method based on the context of prior statements and the code under test. We develop TECo-a deep learning model using code semantics for test completion. The key insight underlying TECO is that predicting the next statement in a test method requires reasoning about code execution, which is hard to do with only syntax-level data that existing code completion models use. Teco extracts and uses six kinds of code semantics data, including the execution result of prior statements and the execution context of the test method. To provide a testbed for this new task, as well as to evaluate TECO, we collect a corpus of 130,934 test methods from 1,270 open-source Java projects. Our results show that Teco achieves an exact-match accuracy of 18, which is 29% higher than the best baseline using syntax-level data only. When measuring functional correctness of generated next statement, Teco can generate runnable code in 29% of the cases compared to 18% obtained by the best baseline. Moreover, Teco is sianificantly better than prior work on test oracle generation. Pengyu Nie 0001, Rahul Banerjee, Junyi Jessy Li, Raymond J. Mooney, Milos Gligoric 0001 |
ICSE | 3 |
| 2023 | Multilingual Code Co-evolution using Large Language ModelsabstractMany software projects implement APIs and algorithms in multiple programming languages. Maintaining such projects is tiresome, as developers have to ensure that any change (e.g., a bug fix or a new feature) is being propagated, timely and without errors, to implementations in other programming languages. In the world of ever-changing software, using rule-based translation tools (i.e., transpilers) or machine learning models for translating code from one language to another provides limited value. Translating each time the entire codebase from one language to another is not the way developers work. In this paper, we target a novel task: translating code changes from one programming language to another using large language models (LLMs). We design and implement the first LLM, dubbed Codeditor, to tackle this task. Codeditor explicitly models code changes as edit sequences and learns to correlate changes across programming languages. To evaluate Codeditor, we collect a corpus of 6,613 aligned code changes from 8 pairs of open-source software projects implementing similar functionalities in two programming languages (Java and C#). Results show that Codeditor outperforms the state-of-the-art approaches by a large margin on all commonly used automatic metrics. Our work also reveals that Codeditor is complementary to the existing generation-based models, and their combination ensures even greater performance. Jiyang Zhang 0003, Pengyu Nie 0001, Junyi Jessy Li, Milos Gligoric 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2022 | ProtoTEx: Explaining Model Decisions with Prototype TensorsabstractWe present PROTOTEX, a novel white-box NLP classification architecture based on prototype networks (Li et al., 2018).PROTOTEX faithfully explains model decisions based on prototype tensors that encode latent clusters of training examples.At inference time, classification decisions are based on the distances between the input text and the prototype tensors, explained via the training examples most similar to the most influential prototypes.We also describe a novel interleaved training algorithm that effectively handles classes characterized by the absence of indicative features.On a propaganda detection task, PROTOTEX accuracy matches BART-large and exceeds BERTlarge with the added benefit of providing faithful explanations.A user study also shows that prototype-based explanations help non-experts to better recognize propaganda in online news. Anubrata Das 0001, Chitrank Gupta, Venelin Kovatchev, Matthew Lease, Junyi Jessy Li |
ACL (1) | 5 |
| 2022 | Evaluating Factuality in Text Simplificationabstractmodels aim to make input texts more readable. Such methods have the potential to make complex information accessible to a wider audience, e.g., providing access to recent medical literature which might otherwise be impenetrable for a lay reader. However, such models risk introducing errors into automatically simplified texts, for instance by inserting statements unsupported by the corresponding original text, or by omitting key information. Providing more readable but inaccurate versions of texts may in many cases be worse than providing no such access at all. The problem of factual accuracy (and the lack thereof) has received heightened attention in the context of summarization models, but the factuality of automatically simplified texts has not been investigated. We introduce a taxonomy of errors that we use to analyze both references drawn from standard simplification datasets and state-of-the-art model outputs. We find that errors often appear in both that are not captured by existing evaluation metrics, motivating a need for research into ensuring the factual accuracy of automated simplification models. Ashwin Devaraj, William Sheffield, Byron C. Wallace, Junyi Jessy Li |
ACL (1) | 4 |
| 2022 | Impact of Evaluation Methodologies on Code SummarizationabstractThere has been a growing interest in developing machine learning (ML) models for code summarization tasks, e.g., comment generation and method naming.Despite substantial increase in the effectiveness of ML models, the evaluation methodologies, i.e., the way people split datasets into training, validation, and test sets, were not well studied.Specifically, no prior work on code summarization considered the timestamps of code and comments during evaluation.This may lead to evaluations that are inconsistent with the intended use cases.In this paper, we introduce the time-segmented evaluation methodology, which is novel to the code summarization research community, and compare it with the mixed-project and cross-project methodologies that have been commonly used.Each methodology can be mapped to some use cases, and the time-segmented methodology should be adopted in the evaluation of ML models for code summarization.To assess the impact of methodologies, we collect a dataset of (code, comment) pairs with timestamps to train and evaluate several recent ML models for code summarization.Our experiments show that different methodologies lead to conflicting evaluation results.We invite the community to expand the set of methodologies used in evaluations. Pengyu Nie 0001, Jiyang Zhang 0003, Junyi Jessy Li, Raymond J. Mooney, Milos Gligoric 0001 |
ACL (1) | 3 |
| 2022 | How Do We Answer Complex Questions: Discourse Structure of Long-form AnswersabstractLong-form answers, consisting of multiple sentences, can provide nuanced and comprehensive answers to a broader set of questions.To better understand this complex and understudied task, we study the functional structure of long-form answers collected from three datasets, ELI5 (Fan et al., 2019), We-bGPT (Nakano et al., 2021) and Natural Questions (Kwiatkowski et al., 2019).Our main goal is to understand how humans organize information to craft complex answers.We develop an ontology of six sentence-level functional roles for long-form answers, and annotate 3.9k sentences in 640 answer paragraphs.Different answer collection methods manifest in different discourse structures.We further analyze model-generated answers -finding that annotators agree less with each other when annotating model-generated answers compared to annotating human-written answers.Our annotated data enables training a strong classifier that can be used for automatic analysis.We hope our work can inspire future research on discourselevel modeling and evaluation of long-form QA systems. 1 Junyi Jessy Li, Eunsol Choi |
ACL (1) | 2 |
| 2022 | The Role of Context and Uncertainty in Shallow Discourse ParsingabstractDiscourse parsing has proven to be useful for a number of NLP tasks that require complex reasoning. However, over a decade since the advent of the Penn Discourse Treebank, predicting implicit discourse relations in text remains challenging. There are several possible reasons for this, and we hypothesize that models should be exposed to more context as it plays an important role in accurate human annotation; meanwhile adding uncertainty measures can improve model accuracy and calibration. To thoroughly investigate this phenomenon, we perform a series of experiments to determine 1) the effects of context on human judgments, and 2) the effect of quantifying uncertainty with annotator confidence ratings on model accuracy and calibration (which we measure using the Brier score (Brier et al, 1950)). We find that including annotator accuracy and confidence improves model accuracy, and incorporating confidence in the model’s temperature function can lead to models with significantly better-calibrated confidence measures. We also find some insightful qualitative results regarding human and model behavior on these datasets. Katherine Atwell, Remi Choi, Junyi Jessy Li, Malihe Alikhani |
COLING | 3 |
| 2022 | Text Simplification of College Admissions Instructions: A Professionally Simplified and Verified CorpusabstractAccess to higher education is critical for minority populations and emergent bilingual students. However, the language used by higher education institutions to communicate with prospective students is often too complex; concretely, many institutions in the US publish admissions application instructions far above the average reading level of a typical high school graduate, often near the 13th or 14th grade level. This leads to an unnecessary barrier between students and access to higher education. This work aims to tackle this challenge via text simplification. We present PSAT (Professionally Simplified Admissions Texts), a dataset with 112 admissions instructions randomly selected from higher education institutions across the US. These texts are then professionally simplified, and verified and accepted by subject-matter experts who are full-time employees in admissions offices at various institutions. Additionally, PSAT comes with manual alignments of 1,883 original-simplified sentence pairs. The result is a first-of-its-kind corpus for the evaluation and fine-tuning of text simplification systems in a high-stakes genre distinct from existing simplification resources. Zachary W. Taylor, Maximus H. Chu, Junyi Jessy Li |
COLING | 3 |
| 2022 | SNaC: Coherence Error Detection for Narrative SummarizationabstractProgress in summarizing long texts is inhibited by the lack of appropriate evaluation frameworks.A long summary that appropriately covers the facets of that text must also present a coherent narrative, but current automatic and human evaluation methods fail to identify gaps in coherence.In this work, we introduce SNAC, a narrative coherence evaluation framework for fine-grained annotations of long summaries.We develop a taxonomy of coherence errors in generated narrative summaries and collect spanlevel annotations for 6.6k sentences across 150 book and movie summaries.Our work provides the first characterization of coherence errors generated by state-of-the-art summarization models and a protocol for eliciting coherence judgments from crowdworkers.Furthermore, we show that the collected annotations allow us to benchmark past work in coherence modeling and train a strong classifier for automatically localizing coherence errors in generated summaries.Finally, our SNAC framework can support future work in long document summarization and coherence evaluation, including improved summarization modeling and posthoc summary correction. Tanya Goyal, Junyi Jessy Li, Greg Durrett |
EMNLP | 2 |
| 2022 | Discourse Comprehension: A Question Answering Framework to Represent Sentence ConnectionsabstractWhile there has been substantial progress in text comprehension through simple factoid question answering, more holistic comprehension of a discourse still presents a major challenge (Dunietz et al., 2020).Someone critically reflecting on a text as they read it will pose curiosity-driven, often open-ended questions, which reflect deep understanding of the content and require complex reasoning to answer (Ko et al., 2020;Westera et al., 2020).A key challenge in building and evaluating models for this type of discourse comprehension is the lack of annotated data, especially since collecting answers to such questions requires high cognitive load for annotators.This paper presents a novel paradigm that enables scalable data collection targeting the comprehension of news documents, viewing these questions through the lens of discourse.The resulting corpus, DCQA (Discourse Comprehension by Question Answering), captures both discourse and semantic links between sentences in the form of free-form, open-ended questions.On an evaluation set that we annotated on questions from Ko et al. (2020), we show that DCQA provides valuable supervision for answering openended questions.We additionally design pretraining methods utilizing existing questionanswering resources, and use synthetic data to accommodate unanswerable questions. Wei-Jen Ko, Cutter Dalton, Mark Simmons, Eliza Fisher, Greg Durrett, Junyi Jessy Li |
EMNLP | 6 |
| 2022 | Why Do You Feel This Way? Summarizing Triggers of Emotions in Social Media PostsabstractCrises such as the COVID-19 pandemic continuously threaten our world and emotionally affect billions of people worldwide in distinct ways.Understanding the triggers leading to people's emotions is of crucial importance.Social media posts can be a good source of such analysis, yet these texts tend to be charged with multiple emotions, with triggers scattering across multiple sentences.This paper takes a novel angle, namely, emotion detection and trigger summarization, aiming to both detect perceived emotions in text, and summarize events and their appraisals that trigger each emotion.To support this goal, we introduce COVIDET (Emotions and their Triggers during Covid-19), a dataset of ~1, 900 English Reddit posts related to COVID-19, which contains manual annotations of perceived emotions and abstractive summaries of their triggers described in the post.We develop strong baselines to jointly detect emotions and summarize emotion triggers.Our analyses show that COVIDET presents new challenges in emotion-specific summarization, as well as multi-emotion detection in long social media posts.* Hongli Zhan and Tiberiu Sosea contributed equally.Reddit Post 1: My sibling is 19 and she constantly goes places with her friends and to there houses and its honestly stressing me out.2: Our grandfather lives with us and he has dementia along with other health issues and my mom has diabetes and heart problems and I have autoimmune diseases & chronic health issues.3: She also has asthma.4: Its stressing me out because despite this she seems to not care about how badly it would affect all of us if we were to get the virus.5: And sadly I feel like its not much I can do she literally doesn't respect my mom and though I'm older she doesn't respect me either.6: Its so frustrating. Emotions and Abstractive Summaries of TriggersEmotion: anger Abstractive Summary of Trigger: My sister having absolutely no regard for any of our family's health coupled with the fact that I can't do anything about it is so aggravating to me.Emotion: fear Abstractive Summary of Trigger: My sibling, who, in spite of our family's myriad of issues that all make us high-risk people, continuously goes out and about, which makes her likely to get infected.I am scared for all of us right now. Hongli Zhan, Tiberiu Sosea, Cornelia Caragea, Junyi Jessy Li |
EMNLP | 4 |
| 2022 | CoditT5: Pretraining for Source Code and Natural Language EditingabstractPretrained language models have been shown to be effective in many software-related generation tasks; however, they are not well-suited for editing tasks as they are not designed to reason about edits. To address this, we propose a novel pretraining objective which explicitly models edits and use it to build CoditT5, a large language model for software-related editing tasks that is pretrained on large amounts of source code and natural language comments. We fine-tune it on various downstream editing tasks, including comment updating, bug fixing, and automated code review. By outperforming standard generation-based models, we demonstrate the generalizability of our approach and its suitability for editing tasks. We also show how a standard generation model and our edit-based model can complement one another through simple reranking strategies, with which we achieve state-of-the-art performance for the three downstream editing tasks. Jiyang Zhang 0003, Sheena Panthaplackel, Pengyu Nie 0001, Junyi Jessy Li, Milos Gligoric 0001 |
ASE | 4 |
| 2022 | Emotion analysis and detection during COVID-19abstractUnderstanding emotions that people express during large-scale crises helps inform policy makers and first responders about the emotional states of the population as well as provide emotional support to those who need such support. We present CovidEmo, a dataset of ~3,000 English tweets labeled with emotions and temporally distributed across 18 months. Our analyses reveal the emotional toll caused by COVID-19, and changes of the social narrative and associated emotions over time. Motivated by the time-sensitive nature of crises and the cost of large-scale annotation efforts, we examine how well large pre-trained language models generalize across domains and timeline in the task of perceived emotion prediction in the context of COVID-19. Our analyses suggest that cross-domain information transfers occur, yet there are still significant gaps. We propose semi-supervised learning as a way to bridge this gap, obtaining significantly better performance using unlabeled data from the target domain. Tiberiu Sosea, Chau Pham 0003, Alexander Tekle, Cornelia Caragea, Junyi Jessy Li |
LREC | 5 |
| 2022 | Political Ideology and Polarization: A Multi-dimensional ApproachabstractBarea Sinno, Bernardo Oviedo, Katherine Atwell, Malihe Alikhani, Junyi Jessy Li. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Barea Sinno, Bernardo Oviedo, Katherine Atwell, Malihe Alikhani, Junyi Jessy Li |
NAACL-HLT | 5 |
| 2021 | Deep Just-In-Time Inconsistency Detection Between Comments and Source CodeabstractNatural language comments convey key aspects of source code such as implementation, usage, and pre- and post-conditions. Failure to update comments accordingly when the corresponding code is modified introduces inconsistencies, which is known to lead to confusion and software bugs. In this paper, we aim to detect whether a comment becomes inconsistent as a result of changes to the corresponding body of code, in order to catch potential inconsistencies just-in-time, i.e., before they are committed to a code base. To achieve this, we develop a deep-learning approach that learns to correlate a comment with code changes. By evaluating on a large corpus of comment/code pairs spanning various comment types, we show that our model outperforms multiple baselines by significant margins. For extrinsic evaluation, we show the usefulness of our approach by combining it with a comment update model to build a more comprehensive automatic comment maintenance system which can both detect and resolve inconsistent comments based on code changes. Sheena Panthaplackel, Junyi Jessy Li, Milos Gligoric 0001, Raymond J. Mooney |
AAAI | 2 |
| 2021 | Paragraph-level Simplification of Medical TextsabstractWe consider the problem of learning to simplify medical texts. This is important because most reliable, up-to-date information in biomedicine is dense with jargon and thus practically inaccessible to the lay audience. Furthermore, manual simplification does not scale to the rapidly growing body of biomedical literature, motivating the need for automated approaches. Unfortunately, there are no large-scale resources available for this task. In this work we introduce a new corpus of parallel texts in English comprising technical and lay summaries of all published evidence pertaining to different clinical topics. We then propose a new metric based on likelihood scores from a masked language model pretrained on scientific texts. We show that this automated measure better differentiates between technical and lay summaries than existing heuristics. We introduce and evaluate baseline encoder-decoder Transformer models for simplification and propose a novel augmentation to these in which we explicitly penalize the decoder for producing 'jargon' terms; we find that this yields improvements over baselines in terms of readability. Ashwin Devaraj, Iain James Marshall, Byron C. Wallace, Junyi Jessy Li |
NAACL-HLT | 4 |
| 2021 | Did they answer? Subjective acts and intents in conversational discourseabstractElisa Ferracane, Greg Durrett, Junyi Jessy Li, Katrin Erk. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Elisa Ferracane, Greg Durrett, Junyi Jessy Li, Katrin Erk |
NAACL-HLT | 3 |
| 2021 | Where Are We in Discourse Relation Recognition?abstractDiscourse parsers recognize the intentional and inferential relationships that organize extended texts.They have had a great influence on a variety of NLP tasks as well as theoretical studies in linguistics and cognitive science.However it is often difficult to achieve good results from current discourse models, largely due to the difficulty of the task, particularly recognizing implicit discourse relations.Recent developments in transformer-based models have shown great promise on these analyses, but challenges still remain.We present a position paper which provides a systematic analysis of the state of the art discourse parsers.We aim to examine the performance of current discourse parsing models via gradual domain shift: within the same corpus, on in-domain texts, and on out-of-domain texts, and discuss the differences between the transformer-based models and the previous models in predicting different types of implicit relations both interand intra-sentential.We conclude by describing several shortcomings of the existing models and a discussion of how future work should approach this problem. Katherine Atwell, Junyi Jessy Li, Malihe Alikhani |
SIGDIAL | 2 |
| 2021 | Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue
Haizhou Li 0001, Gina-Anne Levow, Chitralekha Gupta, Berrak Sisman, Siqi Cai 0002, David Vandyke, Nina Dethlefs, Yan Wu 0002, Junyi Jessy Li |
SIGDIAL | 10 |
| 2020 | Associating Natural Language Comment and Source Code EntitiesabstractComments are an integral part of software development; they are natural language descriptions associated with source code elements. Understanding explicit associations can be useful in improving code comprehensibility and maintaining the consistency between code and comments. As an initial step towards this larger goal, we address the task of associating entities in Javadoc comments with elements in Java source code. We propose an approach for automatically extracting supervised data using revision histories of open source projects and present a manually annotated evaluation dataset for this task. We develop a binary classifier and a sequence labeling model by crafting a rich feature set which encompasses various aspects of code, comments, and the relationships between them. Experiments show that our systems outperform several baselines learning from the proposed supervision. Sheena Panthaplackel, Milos Gligoric 0001, Raymond J. Mooney, Junyi Jessy Li |
AAAI | 4 |
| 2020 | Discourse Level Factors for Sentence Deletion in Text SimplificationabstractThis paper presents a data-driven study focusing on analyzing and predicting sentence deletion — a prevalent but understudied phenomenon in document simplification — on a large English text simplification corpus. We inspect various document and discourse factors associated with sentence deletion, using a new manually annotated sentence alignment corpus we collected. We reveal that professional editors utilize different strategies to meet readability standards of elementary and middle schools. To predict whether a sentence will be deleted during simplification to a certain level, we harness automatically aligned data to train a classification model. Evaluated on our manually annotated data, our best models reached F1 scores of 65.2 and 59.7 for this task at the levels of elementary and middle school, respectively. We find that discourse level factors contribute to the challenging task of predicting sentence deletion for simplification. Wei Xu 0004, Junyi Jessy Li |
AAAI | 4 |
| 2020 | Detecting Perceived Emotions in Hurricane DisastersabstractNatural disasters (e.g., hurricanes) affect millions of people each year, causing widespread destruction in their wake. People have recently taken to social media websites (e.g., Twitter) to share their sentiments and feelings with the larger community. Consequently, these platforms have become instrumental in understanding and perceiving emotions at scale. In this paper, we introduce HurricaneEmo, an emotion dataset of 15,000 English tweets spanning three hurricanes: Harvey, Irma, and Maria. We present a comprehensive study of fine-grained emotions and propose classification tasks to discriminate between coarse-grained emotion groups. Our best BERT model, even after task-guided pre-training which leverages unlabeled Twitter data, achieves only 68% accuracy (averaged across all groups). HurricaneEmo serves not only as a challenging benchmark for models but also as a valuable resource for analyzing emotions in disaster-centric domains. Shrey Desai, Cornelia Caragea, Junyi Jessy Li |
ACL | 3 |
| 2020 | Learning to Update Natural Language Comments Based on Code ChangesabstractWe formulate the novel task of automatically updating an existing natural language comment based on changes in the body of code it accompanies.We propose an approach that learns to correlate changes across two distinct language representations, to generate a sequence of edits that are applied to the existing comment to reflect the source code modifications.We train and evaluate our model using a dataset that we collected from commit histories of open-source software projects, with each example consisting of a concurrent update to a method and its corresponding comment.We compare our approach against multiple baselines using both automatic metrics and human evaluation.Results reflect the challenge of this task and that our model outperforms baselines with respect to making edits. Sheena Panthaplackel, Pengyu Nie 0001, Milos Gligoric 0001, Junyi Jessy Li, Raymond J. Mooney |
ACL | 4 |
| 2020 | Help! Need Advice on Identifying AdviceabstractHumans use language to accomplish a wide variety of tasks - asking for and giving advice being one of them. In online advice forums, advice is mixed in with non-advice, like emotional support, and is sometimes stated explicitly, sometimes implicitly. Understanding the language of advice would equip systems with a better grasp of language pragmatics; practically, the ability to identify advice would drastically increase the efficiency of advice-seeking online, as well as advice-giving in natural language generation systems. We present a dataset in English from two Reddit advice forums - r/AskParents and r/needadvice - annotated for whether sentences in posts contain advice or not. Our analysis reveals rich linguistic phenomena in advice discourse. We present preliminary models showing that while pre-trained language models are able to capture advice better than rule-based systems, advice identification is challenging, and we identify directions for future research. Comments: To be presented at EMNLP 2020. Venkata Subrahmanyan Govindarajan, Benjamin T. Chen, Rebecca Warholic, Katrin Erk, Junyi Jessy Li |
EMNLP (1) | 5 |
| 2020 | Inquisitive Question Generation for High Level Text ComprehensionabstractInquisitive probing questions come naturally to humans in a variety of settings, but is a challenging task for automatic systems.One natural type of question to ask tries to fill a gap in knowledge during text comprehension, like reading a news article: we might ask about background information, deeper reasons behind things occurring, or more.Despite recent progress with data-driven approaches, generating such questions is beyond the range of models trained on existing datasets.We introduce INQUISITIVE, a dataset of ∼19K questions that are elicited while a person is reading through a document.Compared to existing datasets, INQUISITIVE questions target more towards high-level (semantic and discourse) comprehension of text.We show that readers engage in a series of pragmatic strategies to seek information.Finally, we evaluate question generation models based on GPT-2 (Radford et al., 2019) and show that our model is able to generate reasonable questions although the task is challenging, and highlight the importance of context to generate INQUIS-ITIVE questions. Wei-Jen Ko, Te-Yuan Chen, Yiyan Huang, Greg Durrett, Junyi Jessy Li |
EMNLP (1) | 5 |
| 2020 | Assessing Discourse Relations in Language Generation from GPT-2abstractRecent advances in NLP have been attributed to the emergence of large-scale pre-trained language models.GPT-2 (Radford et al., 2019), in particular, is suited for generation tasks given its left-to-right language modeling objective, yet the linguistic quality of its generated text has largely remain unexplored.Our work takes a step in understanding GPT-2's outputs in terms of discourse coherence.We perform a comprehensive study on the validity of explicit discourse relations in GPT-2's outputs under both organic generation and fine-tuned scenarios.Results show GPT-2 does not always generate text containing valid discourse relations; nevertheless, its text is more aligned with human expectation in the fine-tuned scenario.We propose a decoupled strategy to mitigate these problems and highlight the importance of explicitly modeling discourse information. Wei-Jen Ko, Junyi Jessy Li |
INLG | 2 |
| 2020 | An Annotated Dataset of Discourse Modes in Hindi StoriesabstractIn this paper, we present a new corpus consisting of sentences from Hindi short stories annotated for five different discourse modes argumentative, narrative, descriptive, dialogic and informative. We present a detailed account of the entire data collection and annotation processes. The annotations have a very high inter-annotator agreement (0.87 k-alpha). We analyze the data in terms of label distributions, part of speech tags, and sentence lengths. We characterize the performance of various classification algorithms on this dataset and perform ablation studies to understand the nature of the linguistic models suitable for capturing the nuances of the embedded discourse structures in the presented corpus. Swapnil Dhanwal, Hritwik Dutta, Hitesh Nankani, Nilay Shrivastava, Yaman Singla, Junyi Jessy Li, Debanjan Mahata, Rakesh Gosangi, Haimin Zhang 0003, Rajiv Ratn Shah, Amanda Stent |
LREC | 6 |
| 2020 | On the naturalness of hardware descriptionsabstractMining software repositories (MSR) has been shown effective for extracting data used to improve various software engineering tasks, including code completion, code repair, code search, and code summarization. Despite a large body of work on MSR, researchers have focused almost exclusively on repositories that contain code written in imperative programming languages, such as Java and C/C++. Unlike prior work, in this paper, we focus on mining publicly available hardware descriptions (HDs) written in hardware description languages (HDLs), such as VHDL. HDLs have unique syntax and semantics compared to popular imperative languages, and learning-based tools available to hardware designers are well behind those used in other application domains. We assembled large HD corpora consisting of source code written in several HDLs and report on their characteristics. Our language model evaluation reveals that HDs possess a high level of naturalness similar to software written in imperative languages. Further, by utilizing our corpora, we built several deep learning models for automated code completion in VHDL; our models take into account unique characteristics of HDLs, including similarities of nearby concurrent signal assignment statements, in-built concurrency, and the frequently used signal types. These characteristics led to more effective neural models, achieving a BLEU score of 37.3, an 8-14-point improvement over rule-based and neural baselines. Jaeseong Lee 0003, Pengyu Nie 0001, Junyi Jessy Li, Milos Gligoric 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2019 | Predicting and Analyzing Language Specificity in Social Media PostsabstractIn computational linguistics, specificity quantifies how much detail is engaged in text. It is an important characteristic of speaker intention and language style, and is useful in NLP applications such as summarization and argumentation mining. Yet to date, expert-annotated data for sentence-level specificity are scarce and confined to the news genre. In addition, systems that predict sentence specificity are classifiers trained to produce binary labels (general or specific).We collect a dataset of over 7,000 tweets annotated with specificity on a fine-grained scale. Using this dataset, we train a supervised regression model that accurately estimates specificity in social media posts, reaching a mean absolute error of 0.3578 (for ratings on a scale of 1-5) and 0.73 Pearson correlation, significantly improving over baselines and previous sentence specificity prediction systems. We also present the first large-scale study revealing the social, temporal and mental health factors underlying language specificity on social media. Daniel Preotiuc-Pietro, Junyi Jessy Li |
AAAI | 4 |
| 2019 | Domain Agnostic Real-Valued Specificity PredictionabstractSentence specificity quantifies the level of detail in a sentence, characterizing the organization of information in discourse. While this information is useful for many downstream applications, specificity prediction systems predict very coarse labels (binary or ternary) and are trained on and tailored toward specific domains (e.g., news). The goal of this work is to generalize specificity prediction to domains where no labeled data is available and output more nuanced realvalued specificity ratings.We present an unsupervised domain adaptation system for sentence specificity prediction, specifically designed to output real-valued estimates from binary training labels. To calibrate the values of these predictions appropriately, we regularize the posterior distribution of the labels towards a reference distribution. We show that our framework generalizes well to three different domains with 50%-68% mean absolute error reduction than the current state-of-the-art system trained for news sentence specificity. We also demonstrate the potential of our work in improving the quality and informativeness of dialogue generation systems. Wei-Jen Ko, Greg Durrett, Junyi Jessy Li |
AAAI | 3 |
| 2019 | Evaluating Discourse in Structured Text RepresentationsabstractDiscourse structure is integral to understanding a text and is helpful in many NLP tasks.Learning latent representations of discourse is an attractive alternative to acquiring expensive labeled discourse data.Liu and Lapata (2018) propose a structured attention mechanism for text classification that derives a tree over a text, akin to an RST discourse tree.We examine this model in detail, and evaluate on additional discourse-relevant tasks and datasets, in order to assess whether the structured attention improves performance on the end task and whether it captures a text's discourse structure.We find the learned latent trees have little to no structure and instead focus on lexical cues; even after obtaining more structured trees with proposed model modifications, the trees are still far from capturing discourse structure when compared to discourse dependency trees from an existing discourse parser.Finally, ablation studies show the structured attention provides little benefit, sometimes even hurting performance.1 Elisa Ferracane, Greg Durrett, Junyi Jessy Li, Katrin Erk |
ACL (1) | 3 |
| 2019 | Unsupervised Adversarial Domain Adaptation for Implicit Discourse Relation ClassificationabstractImplicit discourse relations are not only more challenging to classify, but also to annotate, than their explicit counterparts.We tackle situations where training data for implicit relations are lacking, and exploit domain adaptation from explicit relations (Ji et al., 2015).We present an unsupervised adversarial domain adaptive network equipped with a reconstruction component.Our system outperforms prior works and other adversarial benchmarks for unsupervised domain adaptation.Additionally, we extend our system to take advantage of labeled data if some are available. Hsin-Ping Huang, Junyi Jessy Li |
CoNLL | 2 |
| 2019 | Adaptive Ensembling: Unsupervised Domain Adaptation for Political Document AnalysisabstractShrey Desai, Barea Sinno, Alex Rosenfeld, Junyi Jessy Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shrey Desai, Barea Sinno, Alex Rosenfeld, Junyi Jessy Li |
EMNLP/IJCNLP (1) | 4 |
| 2019 | A framework for writing trigger-action todo comments in executable formatabstractNatural language elements, e.g., todo comments, are frequently used to communicate among developers and to describe tasks that need to be performed (actions) when specific conditions hold on artifacts related to the code repository (triggers), e.g., from the Apache Struts project: “remove expectedJDK15 and if() after switching to Java 1.6”. As projects evolve, development processes change, and development teams reorganize, these comments, because of their informal nature, frequently become irrelevant or forgotten. We present the first framework, dubbed TrigIt, to specify trigger-action todo comments in executable format. Thus, actions are executed automatically when triggers evaluate to true. TrigIt specifications are written in the host language (e.g., Java) and are evaluated as part of the build process. The triggers are specified as query statements over abstract syntax trees, abstract representation of build configuration scripts, issue tracking systems, and system clock time. The actions are either notifications to developers or code transformation steps. We implemented TrigIt for the Java programming language and migrated 44 existing trigger-action comments from several popular open-source projects. Evaluation of TrigIt, via a user study, showed that users find TrigIt easy to learn and use. TrigIt has the potential to enforce more discipline in writing and maintaining comments in large code repositories. Pengyu Nie 0001, Rishabh Rai, Junyi Jessy Li, Sarfraz Khurshid, Raymond J. Mooney, Milos Gligoric 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2018 | A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical LiteratureabstractWe present a corpus of 5,000 richly annotated abstracts of medical articles describing clinical randomized controlled trials. Annotations include demarcations of text spans that describe the Patient population enrolled, the Interventions studied and to what they were Compared, and the Outcomes measured (the 'PICO' elements). These spans are further annotated at a more granular level, e.g., individual interventions within them are marked and mapped onto a structured medical vocabulary. We acquired annotations from a diverse set of workers with varying levels of expertise and cost. We describe our data collection process and the corpus itself in detail. We then outline a set of challenging NLP tasks that would aid searching of the medical literature and the practice of evidence-based medicine. Benjamin E. Nye, Junyi Jessy Li, Roma Patel, Yinfei Yang, Iain James Marshall, Ani Nenkova, Byron C. Wallace |
ACL (1) | 2 |
| 2018 | Expressively vulgar: The socio-dynamics of vulgarity and its effects on sentiment analysis in social mediaabstractVulgarity is a common linguistic expression and is used to perform several linguistic functions. Understanding their usage can aid both linguistic and psychological phenomena as well as benefit downstream natural language processing applications such as sentiment analysis. This study performs a large-scale, data-driven empirical analysis of vulgar words using social media data. We analyze the socio-cultural and pragmatic aspects of vulgarity using tweets from users with known demographics. Further, we collect sentiment ratings for vulgar tweets to study the relationship between the use of vulgar words and perceived sentiment and show that explicitly modeling vulgar words can boost sentiment analysis performance. Isabel Cachola, Eric Holgate, Daniel Preotiuc-Pietro, Junyi Jessy Li |
COLING | 4 |
| 2018 | Why Swear? Analyzing and Inferring the Intentions of Vulgar ExpressionsabstractVulgar words are employed in language use for several different functions, ranging from expressing aggression to signaling group identity or the informality of the communication.This versatility of usage of a restricted set of words is challenging for downstream applications and has yet to be studied quantitatively or using natural language processing techniques.We introduce a novel data set of 7,800 tweets from users with known demographic traits where all instances of vulgar words are annotated with one of the six categories of vulgar word use.Using this data set, we present the first analysis of the pragmatic aspects of vulgarity and how they relate to social factors.We build a model able to predict the category of a vulgar word based on the immediate context it appears in with 67.4 macro F1 across six classes.Finally, we demonstrate the utility of modeling the type of vulgar word use in context by using this information to achieve state-of-the-art performance in hate speech detection on a benchmark data set. Eric Holgate, Isabel Cachola, Daniel Preotiuc-Pietro, Junyi Jessy Li |
EMNLP | 4 |
| 2017 | Aggregating and Predicting Sequence Labels from Crowd AnnotationsabstractDespite sequences being core to NLP, scant work has considered how to handle noisy sequence labels from multiple annotators for the same text. Given such annotations, we consider two complementary tasks: (1) aggregating sequential crowd labels to infer a best single set of consensus annotations; and (2) using crowd annotations as training data for a model that can predict sequences in unannotated text. For aggregation, we propose a novel Hidden Markov Model variant. To predict sequences in unannotated text, we propose a neural approach using Long Short Term Memory. We evaluate a suite of methods across two different applications and text genres: Named-Entity Recognition in news articles and Information Extraction from biomedical abstracts. Results show improvement over strong baselines. Our source code and data are available online. An T. Nguyen 0001, Byron C. Wallace, Junyi Jessy Li, Ani Nenkova, Matthew Lease |
ACL (1) | 3 |
| 2016 | Estimating Text Intelligibility via Information Packaging AnalysisabstractEffective communication through language involves organizing the content a person or system wishes to convey into text that flows naturally. There are many ways to render the same information, but those appropriate for one group of audience may not be intelligible to another. The goal of this thesis to analyze and address factors that influence the intelligibility of text from two aspects of information packaging: discourse structure and text specificity. Effective communication through language involves organizing the content a person or system wishes to convey into text that flows naturally. There are many ways to render the same information, but those appropriate for one group of audience may not be intelligible to another. The goal of this thesis to analyze and address factors that influence the intelligibility of text from two aspects of information packaging: discourse structure and text specificity. Junyi Jessy Li |
AAAI | 1 |
| 2016 | Improving the Annotation of Sentence Specificity
Junyi Jessy Li, Bridget O'Daniel, Wenli Zhao, Ani Nenkova |
LREC | 1 |
| 2016 | The Instantiation Discourse Relation: A Corpus Analysis of Its Properties and Improved DetectionabstractINSTANTIATION is a fairly common discourse relation and past work has suggested that it plays special roles in local coherence, in sentiment expression and in content selection in summarization.In this paper we provide the first systematic corpus analysis of the relation and show that relation-specific features can improve considerably the detection of the relation.We show that sentences involved in INSTANTIATION are set apart from other sentences by the use of gradable (subjective) adjectives, the occurrence of rare words and by different patterns in part-of-speech usage.Words across arguments of INSTANTI-ATION are connected through hypernym and meronym relations significantly more often than in other sentences and that they stand out in context by being significantly less similar to each other than other adjacent sentence pairs.These factors provide substantial predictive power that improves the identification of implicit INSTANTIATION relation by more than 5% F-measure. Junyi Jessy Li, Ani Nenkova |
HLT-NAACL | 1 |
| 2016 | The Role of Discourse Units in Near-Extractive SummarizationabstractAlthough human-written summaries of documents tend to involve significant edits to the source text, most automated summarizers are extractive and select sentences verbatim.In this work we examine how elementary discourse units (EDUs) from Rhetorical Structure Theory can be used to extend extractive summarizers to produce a wider range of human-like summaries.Our analysis demonstrates that EDU segmentation is effective in preserving human-labeled summarization concepts within sentences and also aligns with near-extractive summaries constructed by news editors.Finally, we show that using EDUs as units of content selection instead of sentences leads to stronger summarization performance in near-extractive scenarios, especially under tight budgets. Junyi Jessy Li, Kapil Thadani, Amanda Stent |
SIGDIAL Conference | 1 |
| 2015 | Fast and Accurate Prediction of Sentence SpecificityabstractRecent studies have demonstrated that specificity is an important characterization of texts potentially beneficial for a range of applications such as multi-document news summarization and analysis of science journalism. The feasibility of automatically predicting sentence specificity from a rich set of features has also been confirmed in prior work. In this paper we present a practical system for predicting sentence specificity which exploits only features that require minimum processing and is trained in a semi-supervised manner. Our system outperforms the state-of-the-art method for predicting sentence specificity and does not require part of speech tagging or syntactic parsing as the prior methods did. With the tool that we developed --- Speciteller --- we study the role of specificity in sentence simplification. We show that specificity is a useful indicator for finding sentences that need to be simplified and a useful objective for simplification, descriptive of the differences between original and simplified sentences. Junyi Jessy Li, Ani Nenkova |
AAAI | 1 |
| 2015 | Detecting Content-Heavy Sentences: A Cross-Language Case StudyabstractThe information conveyed by some sentences would be more easily understood by a reader if it were expressed in multiple sentences.We call such sentences content heavy: these are possibly grammatical but difficult to comprehend, cumbersome sentences.In this paper we introduce the task of detecting content-heavy sentences in cross-lingual context.Specifically we develop methods to identify sentences in Chinese for which English speakers would prefer translations consisting of more than one sentence.We base our analysis and definitions on evidence from multiple human translations and reader preferences on flow and understandability.We show that machine translation quality when translating content heavy sentences is markedly worse than overall quality and that this type of sentence are fairly common in Chinese news.We demonstrate that sentence length and punctuation usage in Chinese are not sufficient clues for accurately detecting heavy sentences and present a richer classification model that accurately identifies these sentences. Junyi Jessy Li, Ani Nenkova |
EMNLP | 1 |
| 2014 | Cross-lingual Discourse Relation Analysis: A corpus study and a semi-supervised classification system
Junyi Jessy Li, Marine Carpuat, Ani Nenkova |
COLING | 1 |
| 2014 | Addressing Class Imbalance for Improved Recognition of Implicit Discourse RelationsabstractIn this paper we address the problem of skewed class distribution in implicit dis-course relation recognition. We examine the performance of classifiers for both bi-nary classification predicting if a particu-lar relation holds or not and for multi-class prediction. We review prior work to point out that the problem has been addressed differently for the binary and multi-class problems. We demonstrate that adopting a unified approach can significantly im-prove the performance of multi-class pre-diction. We also propose an approach that makes better use of the full annotations in the training set when downsampling is used. We report significant absolute im-provements in performance in multi-class prediction, as well as significant improve-ment of binary classifiers for detecting the presence of implicit Temporal, Compari-son and Contingency relations. 1 Junyi Jessy Li, Ani Nenkova |
SIGDIAL Conference | 1 |
| 2014 | Reducing Sparsity Improves the Recognition of Implicit Discourse RelationsabstractThe earliest work on automatic detec-tion of implicit discourse relations relied on lexical features. More recently, re-searchers have demonstrated that syntactic features are superior to lexical features for the task. In this paper we re-examine the two classes of state of the art representa-tions: syntactic production rules and word pair features. In particular, we focus on the need to reduce sparsity in instance repre-sentation, demonstrating that different rep-resentation choices even for the same class of features may exacerbate sparsity issues and reduce performance. We present re-sults that clearly reveal that lexicalization of the syntactic features is necessary for good performance. We introduce a novel, less sparse, syntactic representation which leads to improvement in discourse rela-tion recognition. Finally, we demonstrate that classifiers trained on different repre-sentations, especially lexical ones, behave rather differently and thus could likely be combined in future systems. 1 Junyi Jessy Li, Ani Nenkova |
SIGDIAL Conference | 1 |