EDBT 2026 Demo / reviewers in the wild / expert
Smaranda Muresan
dblp:44/70
· DBLP profile ↗
58ranked-venue papers
8as first author
31since 2021 · last 2025
0000-0003-4532-0182ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 6 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 4Databases, data management, data science and information retrieval · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and ReasoningabstractWe introduce Browsing Lost Unformed Recollections, a tip-of-the-tongue known-item search and reasoning benchmark for general AI assistants. BLUR introduces a set of 573 real-world validated questions that demand searching and reasoning across multimodal and multilingual inputs, as well as proficient tool use, in order to excel on. Humans easily ace these questions (scoring on average 98%), while the best-performing system scores around 56%. To facilitate progress toward addressing this challenging and aspirational use case for general AI assistants, we release 350 questions through a public leaderboard, retain the answers to 250 of them, and have the rest as a private test set. Sky CH-Wang, Darshan Deshpande, Smaranda Muresan, Anand Kannappan, Rebecca Qian |
ACL (1) | 3 |
| 2025 | Latent Space Interpretation for Stylistic Analysis and Explainable Authorship AttributionabstractRecent state-of-the-art authorship attribution methods learn authorship representations of text in a latent, uninterpretable space, which hinders their usability in real-world applications. We propose a novel approach for interpreting learned embeddings by identifying representative points in the latent space and leveraging large language models to generate informative natural language descriptions of the writing style associated with each point. We evaluate the alignment between our interpretable and latent spaces and demonstrate superior prediction agreement over baseline methods. Additionally, we conduct a human evaluation to assess the quality of these style descriptions and validate their utility in explaining the latent space. Finally, we show that human performance on the challenging authorship attribution task improves by +20% on average when aided with explanations from our method. Milad Alshomary, Narutatsu Ri, Marianna Apidianaki, Ajay Patel, Smaranda Muresan, Kathy McKeown |
COLING | 5 |
| 2025 | Layered Insights: Generalizable Analysis of Human Authorial Style by Leveraging All Transformer LayersabstractWe propose a new approach for the authorship attribution task that leverages the various linguistic representations learned at different layers of pre-trained transformer-based models.We evaluate our approach on two popular authorship attribution models and three evaluation datasets, in in-domain and out-of-domain scenarios.We found that utilizing various transformer layers improves the robustness of authorship attribution models when tested on outof-domain data, resulting in a much stronger performance.Our analysis gives further insights into how our model's different layers get specialized in representing certain linguistic aspects that we believe benefit the model when tested out of the domain. Milad Alshomary, Nikhil Reddy Varimalla, Vishal Anand 0002, Smaranda Muresan, Kathy McKeown |
EMNLP | 4 |
| 2025 | Exploring Chain-of-Thought Reasoning for Steerable Pluralistic AlignmentabstractLarge Language Models (LLMs) are typically trained to reflect a relatively uniform set of values, which limits their applicability to tasks that require understanding of nuanced human perspectives.Recent research has underscored the importance of enabling LLMs to support steerable pluralism -the capacity to adopt a specific perspective and align generated outputs with it.In this work, we investigate whether Chain-of-Thought (CoT) reasoning techniques can be applied to building steerable pluralistic models.We explore several methods, including CoT prompting, fine-tuning on humanauthored CoT, fine-tuning on synthetic explanations, and Reinforcement Learning with Verifiable Rewards (RLVR).We evaluate these approaches using the Value Kaleidoscope and OpinionQA datasets.Among the methods studied, RLVR consistently outperforms others and demonstrates strong training sample efficiency.We further analyze the generated CoT traces with respect to faithfulness and safety. Kathy McKeown, Smaranda Muresan |
EMNLP | 3 |
| 2025 | Forecasting Conversation Derailments Through GenerationabstractForecasting conversation derailment can be useful in real-world settings such as online content moderation, conflict resolution, and business negotiations. However, despite language models’ success at identifying offensive speech present in conversations, they struggle to forecast future conversation derailments. In contrast to prior work that predicts conversation outcomes solely based on the past conversation history, our approach samples multiple future conversation trajectories conditioned on existing conversation history using a fine-tuned LLM. It predicts the conversation outcome based on the consensus of these trajectories. We also experimented with leveraging socio-linguistic attributes, which reflect turn-level conversation dynamics, as guidance when generating future conversations. Our method of future conversation trajectories surpasses state-of-the-art results on English conversation derailment prediction benchmarks and demonstrates significant accuracy gains in ablation studies. Kathy McKeown, Smaranda Muresan |
INLG | 3 |
| 2025 | Understanding Figurative Meaning through Explainable Visual EntailmentabstractArkadiy Saakyan, Shreyas Kulkarni, Tuhin Chakrabarty, Smaranda Muresan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Arkadiy Saakyan, Tuhin Chakrabarty, Smaranda Muresan |
NAACL (Long Papers) | 4 |
| 2025 | Navigating the Landscape of Hint Generation Research: From the Past to the FutureabstractAbstract Digital education has gained popularity in the last decade, especially after the COVID-19 pandemic. With the improving capabilities of large language models to reason and communicate with users, envisioning intelligent tutoring systems that can facilitate self-learning is not very far-fetched. One integral component to fulfill this vision is the ability to give accurate and effective feedback via hints to scaffold the learning process. In this survey article, we present a comprehensive review of prior research on hint generation, aiming to bridge the gap between research in education and cognitive science, and research in AI and Natural Language Processing. Informed by our findings, we propose a formal definition of the hint generation task, and discuss the roadmap of building an effective hint generation system aligned with the formal definition, including open challenges, future directions and ethical considerations. Anubhav Jangra, Jamshid Mozafari, Adam Jatowt, Smaranda Muresan |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | ICLEF: In-Context Learning with Expert Feedback for Explainable Style TransferabstractWhile state-of-the-art large language models (LLMs) can excel at adapting text from one style to another, current work does not address the explainability of style transfer models.Recent work has explored generating textual explanations from larger teacher models and distilling them into smaller student models.One challenge with such approach is that LLM outputs may contain errors that require expertise to correct, but gathering and incorporating expert feedback is difficult due to cost and availability.To address this challenge, we propose ICLEF, a novel human-AI collaboration approach to model distillation that incorporates scarce expert human feedback by combining in-context learning and model self-critique.We show that our method leads to generation of high-quality synthetic explainable style transfer datasets for formality (E-GYAFC) and subjective bias (E-WNC).Via automatic and human evaluation, we show that specialized student models finetuned on our datasets outperform generalist teacher models on the explainable style transfer task in one-shot settings, and perform competitively compared to few-shot teacher models, highlighting the quality of the data and the role of expert feedback.In an extrinsic task of authorship attribution, we show that explanations generated by smaller models fine-tuned on E-GYAFC are more predictive of authorship than explanations generated by few-shot teacher models. Arkadiy Saakyan, Smaranda Muresan |
ACL (1) | 2 |
| 2024 | Creativity Support in the Age of Large Language Models: An Empirical Study Involving Professional WritersabstractThe development of large language models (LLMs) capable of following instructions and engaging in conversational interactions has led to increased interest in their use across various support tools. We investigate the effectiveness of contemporary LLMs in assisting professional writers via an empirical user study (n=30). The design of our collaborative writing interface is grounded in the cognitive process model of writing [17]. This allows writers to obtain model help in each of the three non-linear cognitive activities in the writing process: planning, translating and reviewing. Participants write short fiction/non-fiction with model help and are subsequently asked to submit a post-completion survey to provide qualitative feedback on the potential and pitfalls of LLMs as writing collaborators. Upon analyzing the writer-LLM interactions, we find that while seeking help across all three types of cognitive activities, writers find LLMs more helpful in translation and reviewing. Our findings from analyzing both the interactions and the survey responses highlight future research directions in creative writing assistance using LLMs. Tuhin Chakrabarty, Vishakh Padmakumar, Faeze Brahman, Smaranda Muresan |
Creativity & Cognition | 4 |
| 2024 | Art or Artifice? Large Language Models and the False Promise of CreativityabstractResearchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by the Torrance Test of Creative Thinking (TTCT) [64], which measures creativity as a process, we use the Consensual Assessment Technique [3] and propose Torrance Test of Creative Writing (TTCW) to evaluate creativity as product. TTCW consists of 14 binary tests organized into the original dimensions of Fluency, Flexibility, Originality, and Elaboration. We recruit 10 creative writers and implement a human assessment of 48 stories written either by professional authors or LLMs using TTCW. Our analysis shows that LLM-generated stories pass 3-10X less TTCW tests than stories written by professionals. In addition, we explore the use of LLMs as assessors to automate the TTCW evaluation, revealing that none of the LLMs positively correlate with the expert assessments. Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, Chien-Sheng Wu |
CHI | 4 |
| 2024 | A Weak Supervision Approach for Few-Shot Aspect Based Sentiment AnalysisabstractRobert Vacareanu, Siddharth Varia, Kishaloy Halder, Shuai Wang, Giovanni Paolini, Neha Anna John, Miguel Ballesteros, Smaranda Muresan. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Robert Vacareanu, Siddharth Varia, Kishaloy Halder, Giovanni Paolini, Neha Anna John, Miguel Ballesteros, Smaranda Muresan |
EACL (1) | 8 |
| 2024 | Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word GameabstractThe New York Times Connections game has emerged as a popular and challenging pursuit for word puzzle enthusiasts.We collect 438 Connections games to evaluate the performance of state-of-the-art large language models (LLMs) against expert and novice human players.Our results show that even the bestperforming LLM, Claude 3.5 Sonnet, which has otherwise shown impressive reasoning abilities on a wide variety of benchmarks, can only fully solve 18% of the games.Novice and expert players perform better than Claude 3.5 Sonnet, with expert human players significantly outperforming it.We create a taxonomy of the knowledge types required to successfully cluster and categorize words in the Connections game.We find that while LLMs perform relatively well on categorizing words based on semantic relations they struggle with other types of knowledge such as Encyclopedic Knowledge, Multiword Expressions or knowledge that combines both Word Form and Meaning.Our results establish the New York Times Connections game as a challenging benchmark for evaluating abstract reasoning capabilities in AI systems. Prisha Samadarshi, Mariam Mustafa, Anushka Kulkarni, Raven Rothkopf, Tuhin Chakrabarty, Smaranda Muresan |
EMNLP | 6 |
| 2023 | NORMSAGE: Multi-Lingual Multi-Cultural Norm Discovery from Conversations On-the-FlyabstractKnowledge of norms is needed to understand and reason about acceptable behavior in human communication and interactions across sociocultural scenarios.Most computational research on norms has focused on a single culture, and manually built datasets, from nonconversational settings.We address these limitations by proposing a new framework, NORMSAGE 1 , to automatically extract culturespecific norms from multi-lingual conversations.NORMSAGE uses GPT-3 prompting to 1) extract candidate norms directly from conversations and 2) provide explainable selfverification to ensure correctness and relevance.Comprehensive empirical results show the promise of our approach to extract highquality culture-aware norms from multi-lingual conversations (English and Chinese), across several quality metrics.Further, our relevance verification can be extended to assess the adherence and violation of any norm with respect to a conversation on-the-fly, along with textual explanation.NORMSAGE achieves an AUC of 94.6% in this grounding setup, with generated explanations matching human-written quality.𝐚) 𝐈𝐧𝐢𝐭𝐢𝐚𝐥 𝐃𝐢𝐬𝐜𝐨𝐯𝐞𝐫𝒚: 𝒅𝒗𝒓(⋅) irrelevant contradict entail Correctness Verdict ( ! 𝑪 𝒗 ): Correctness Explanation ( ! 𝑪 𝒆 ): Yes, honesty is the foundation of trust, and strong family relationships are built on trust. Yi R. Fung 0001, Tuhin Chakrabarty, Owen Rambow, Smaranda Muresan, Heng Ji 0001 |
EMNLP | 5 |
| 2023 | Sociocultural Norm Similarities and Differences via Situational Alignment and Explainable Textual EntailmentabstractDesigning systems that can reason across cultures requires that they are grounded in the norms of the contexts in which they operate.However, current research on developing computational models of social norms has primarily focused on American society.Here, we propose a novel approach to discover and compare descriptive social norms across Chinese and American cultures.We demonstrate our approach by leveraging discussions on a Chinese Q&A platform-知乎 (Zhihu)-and the existing SOCIALCHEMISTRY dataset as proxies for contrasting cultural axes, align social situations cross-culturally, and extract social norms from texts using in-context learning.Embedding Chain-of-Thought prompting in a human-AI collaborative framework, we build a high-quality dataset of 3,069 social norms aligned with social situations across Chinese and American cultures alongside corresponding free-text explanations.To test the ability of models to reason about social norms across cultures, we introduce the task of explainable social norm entailment, showing that existing models under 3B parameters have significant room for improvement in both automatic and human evaluation.Further analysis of crosscultural norm differences based on our dataset shows empirical alignment with the social orientations framework, revealing several situational and descriptive nuances in norms across these cultures. Sky CH-Wang, Arkadiy Saakyan, Oliver Li, Smaranda Muresan |
EMNLP | 5 |
| 2023 | NormDial: A Comparable Bilingual Synthetic Dialog Dataset for Modeling Social Norm Adherence and ViolationabstractSocial norms fundamentally shape interpersonal communication.We present NORMDIAL, a high-quality dyadic dialogue dataset with turn-by-turn annotations of social norm adherences and violations for Chinese and American cultures.Introducing the task of social norm observance detection, our dataset is synthetically generated in both Chinese and English using a human-in-the-loop pipeline by prompting large language models with a small collection of expert-annotated social norms.We show that our generated dialogues are of high quality through human evaluation and further evaluate the performance of existing large language models on this task.Our findings point towards new directions for understanding the nuances of social norms as they manifest in conversational contexts that span across languages and cultures. Oliver Li, Mallika Subramanian, Arkadiy Saakyan, Sky CH-Wang, Smaranda Muresan |
EMNLP | 5 |
| 2023 | FeelingBlue: A Corpus for Understanding the Emotional Connotation of Color in ContextabstractAbstract While the link between color and emotion has been widely studied, how context-based changes in color impact the intensity of perceived emotions is not well understood. In this work, we present a new multimodal dataset for exploring the emotional connotation of color as mediated by line, stroke, texture, shape, and language. Our dataset, FeelingBlue, is a collection of 19,788 4-tuples of abstract art ranked by annotators according to their evoked emotions and paired with rationales for those annotations. Using this corpus, we present a baseline for a new task: Justified Affect Transformation. Given an image I, the task is to 1) recolor I to enhance a specified emotion e and 2) provide a textual justification for the change in e. Our model is an ensemble of deep neural networks which takes I, generates an emotionally transformed color palette p conditioned on I, applies p to I, and then justifies the color transformation in text via a visual-linguistic model. Experimental results shed light on the emotional connotation of color in context, demonstrating both the promise of our approach on this challenging task and the considerable potential for future investigations enabled by our corpus.1 Amith Ananthram, Olivia Winn, Smaranda Muresan |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | Multitask Instruction-based Prompting for Fallacy RecognitionabstractFallacies are used as seemingly valid arguments to support a position and persuade the audience about its validity.Recognizing fallacies is an intrinsically difficult task both for humans and machines.Moreover, a big challenge for computational models lies in the fact that fallacies are formulated differently across the datasets with differences in the input format (e.g., question-answer pair, sentence with fallacy fragment), genre (e.g., social media, dialogue, news), as well as types and number of fallacies (from 5 to 18 types per dataset).To move towards solving the fallacy recognition task, we approach these differences across datasets as multiple tasks and show how instruction-based prompting in a multitask setup based on the T5 model improves the results against approaches built for a specific dataset such as T5, BERT or GPT-3.We show the ability of this multitask prompting approach to recognize 28 unique fallacies across domains and genres and study the effect of model size and prompt choice by analyzing the per-class (i.e., fallacy type) results.Finally, we analyze the effect of annotation quality on model performance, and the feasibility of complementing this approach with external knowledge. Tariq Alhindi, Tuhin Chakrabarty, Elena Musi, Smaranda Muresan |
EMNLP | 4 |
| 2022 | Affective Idiosyncratic Responses to MusicabstractAffective responses to music are highly personal.Despite consensus that idiosyncratic factors play a key role in regulating how listeners emotionally respond to music, precisely measuring the marginal effects of these variables has proved challenging.To address this gap, we develop computational methods to measure affective responses to music from over 403M listener comments on a Chinese social music platform.Building on studies from music psychology in systematic and quasi-causal analyses, we test for musical, lyrical, contextual, demographic, and mental health effects that drive listener affective responses.Finally, motivated by the social phenomenon known as 网抑 云 (wǎng-yì-yún), we identify influencing factors of platform user self-disclosures, the social support they receive, and notable differences in discloser user activity. Sky CH-Wang, Evan Li, Oliver Li, Smaranda Muresan |
EMNLP | 4 |
| 2022 | FLUTE: Figurative Language Understanding through Textual ExplanationsabstractFigurative language understanding has been recently framed as a recognizing textual entailment (RTE) task (a.k.a.natural language inference, or NLI).However, similar to classical RTE/NLI datasets, the current benchmarks suffer from spurious correlations and annotation artifacts.To tackle this problem, work on NLI has built explanation-based datasets such as e-SNLI, allowing us to probe whether language models are right for the right reasons.Yet no such data exists for figurative language, making it harder to assess genuine understanding of such expressions.To address this issue, we release FLUTE, a dataset of 9,000 figurative NLI instances with explanations, spanning four categories: Sarcasm, Simile, Metaphor, and Idioms.We collect the data through a model-in-the-loop framework based on GPT-3, crowd workers, and expert annotators.We show how utilizing GPT-3 in conjunction with human annotators (novices and experts) can aid in scaling up the creation of datasets even for such complex linguistic phenomena as figurative language.The baseline performance of the T5 model fine-tuned on FLUTE shows that our dataset can bring us a step closer to developing models that understand figurative language through textual explanations. Tuhin Chakrabarty, Arkadiy Saakyan, Debanjan Ghosh, Smaranda Muresan |
EMNLP | 4 |
| 2022 | Fine-tuned Language Models are Continual LearnersabstractRecent work on large language models relies on the intuition that most natural language processing tasks can be described via natural language instructions and that models trained on these instructions show strong zero-shot performance on several standard datasets.However, these models even though impressive still perform poorly on a wide range of tasks outside of their respective training and evaluation sets.To address this limitation, we argue that a model should be able to keep extending its knowledge and abilities, without forgetting previous skills.In spite of the limited success of Continual Learning we show that Fine-tuned Language Models can be continual learners.We empirically investigate the reason for this success and conclude that Continual Learning emerges from self-supervision pre-training.Our resulting model Continual-T0 (CT0) is able to learn 8 new diverse language generation tasks, while still maintaining good performance on previous tasks, spanning in total 70 datasets.Finally, we show that CT0 is able to combine instructions in ways it was never trained for, demonstrating some level of instruction compositionality. Thomas Scialom, Tuhin Chakrabarty, Smaranda Muresan |
EMNLP | 3 |
| 2022 | BeSt: The Belief and Sentiment CorpusabstractWe present the BeSt corpus, which records cognitive state: who believes what (i.e., factuality), and who has what sentiment towards what. This corpus is inspired by similar source-and-target corpora, specifically MPQA and FactBank. The corpus comprises two genres, newswire and discussion forums, in three languages, Chinese (Mandarin), English, and Spanish. The corpus is distributed through the LDC. Jennifer Tracey, Owen Rambow, Claire Cardie, Adam Dalton 0001, Hoa Trang Dang, Mona T. Diab, Bonnie J. Dorr, Louise Guthrie, Magdalena Markowska, Smaranda Muresan, Vinodkumar Prabhakaran, Samira Shaikh, Tomek Strzalkowski |
LREC | 10 |
| 2022 | Unsupervised Stem-based Cross-lingual Part-of-Speech Tagging for Morphologically Rich Low-Resource LanguagesabstractUnsupervised cross-lingual projection for part-of-speech (POS) tagging relies on the use of parallel data to project POS tags from a source language for which a POS tagger is available onto a target language across word-level alignments. The projected tags then form the basis for learning a POS model for the target language. However, languages with rich morphology often yield sparse word alignments because words corresponding to the same citation form do not align well. We hypothesize that for morphologically complex languages, it is more efficient to use the stem rather than the word as the core unit of abstraction. Our contributions are: 1) we propose an unsupervised stem-based cross-lingual approach for POS tagging for low-resource languages of rich morphology; 2) we further investigate morpheme-level alignment and projection; and 3) we examine whether the use of linguistic priors for morphological segmentation improves POS tagging. We conduct experiments using six source languages and eight morphologically complex target languages of diverse typologies. Our results show that the stem-based approach improves the POS models for all the target languages, with an average relative error reduction of 10.3% in accuracy per target language, and outperforms the word-based approach that operates on three-times more data for about two thirds of the language pairs we consider. Moreover, we show that morpheme-level alignment and projection and the use of linguistic priors for morphological segmentation further improve POS tagging. Ramy Eskander, Cass Lowry, Sujay Khandagale, Judith L. Klavans, Maria Polinsky, Smaranda Muresan |
NAACL-HLT | 6 |
| 2021 | COVID-Fact: Fact Extraction and Verification of Real-World Claims on COVID-19 PandemicabstractArkadiy Saakyan, Tuhin Chakrabarty, Smaranda Muresan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Arkadiy Saakyan, Tuhin Chakrabarty, Smaranda Muresan |
ACL/IJCNLP (1) | 3 |
| 2021 | Metaphor Generation with Conceptual MappingsabstractKevin Stowe, Tuhin Chakrabarty, Nanyun Peng, Smaranda Muresan, Iryna Gurevych. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kevin Stowe, Tuhin Chakrabarty, Nanyun Peng 0001, Smaranda Muresan, Iryna Gurevych |
ACL/IJCNLP (1) | 4 |
| 2021 | "Laughing at you or with you": The Role of Sarcasm in Shaping the Disagreement SpaceabstractDetecting arguments in online interactions is useful to understand how conflicts arise and get resolved.Users often use figurative language, such as sarcasm, either as persuasive devices or to attack the opponent by an ad hominem argument.To further our understanding of the role of sarcasm in shaping the disagreement space, we present a thorough experimental setup using a corpus annotated with both argumentative moves (agree/disagree) and sarcasm.We exploit joint modeling in terms of (a) applying discrete features that are useful in detecting sarcasm to the task of argumentative relation classification (agree/disagree/none), and (b) multitask learning for argumentative relation classification and sarcasm detection using deep learning architectures (e.g., dual Long Short-Term Memory (LSTM) with hierarchical attention and Transformer-based architectures).We demonstrate that modeling sarcasm improves the argumentative relation classification task (agree/disagree/none) in all setups.* Equal Contribution.Arg.Rel.Turn Pairs Prior Turn: Today, no informed creationist would deny natural selection.Agree Current Turn: Seeing how this was proposed over a century and a half ago by Darwin, what took the creationists so long to catch up?Prior Turn: Personally I wouldn't own a gun for self defense because I am just not that big of a sissy.Disagree Current Turn: Because taking responsibility for ones own safety is certainly a sissy thing to do? Prior Turn: I'm not surprised that no one on your side of the debate would correct you, but wolves and dogs are both members of the same species.The Canid species.Current Turn: Wow, you 're even wrong when you get away from your precious Bible and try to sound scientific.Prior Turn: The hand of God kept me from serious harm.Maybe He has a plan for me.N one Current Turn: You better hurry up .Are n't you like 113 years old. Debanjan Ghosh, Ritvik Shrivastava, Smaranda Muresan |
EACL | 3 |
| 2021 | Don't Go Far Off: An Empirical Study on Neural Poetry TranslationabstractDespite constant improvements in machine translation quality, automatic poetry translation remains a challenging problem due to the lack of open-sourced parallel poetic corpora, and to the intrinsic complexities involved in preserving the semantics, style and figurative nature of poetry.We present an empirical investigation for poetry translation along several dimensions: 1) size and style of training data (poetic vs. non-poetic), including a zeroshot setup; 2) bilingual vs. multilingual learning; and 3) language-family-specific models vs. mixed-language-family models.To accomplish this, we contribute a parallel dataset of poetry translations for several language pairs.Our results show that multilingual fine-tuning on poetic text significantly outperforms multilingual fine-tuning on non-poetic text that is 35X larger in size, both in terms of automatic metrics (BLEU, BERTScore, COMET) and human evaluation metrics such as faithfulness (meaning and poetic style).Moreover, multilingual fine-tuning on poetic data outperforms bilingual fine-tuning on poetic data. Tuhin Chakrabarty, Arkadiy Saakyan, Smaranda Muresan |
EMNLP (1) | 3 |
| 2021 | Implicit Premise Generation with Discourse-aware Commonsense Knowledge ModelsabstractEnthymemes are defined as arguments where a premise or conclusion is left implicit.We tackle the task of generating the implicit premise in an enthymeme, which requires not only an understanding of the stated conclusion and premise, but also additional inferences that could depend on commonsense knowledge.The largest available dataset for enthymemes (Habernal et al., 2018) consists of 1.7k samples, which is not large enough to train a neural text generation model.To address this issue, we take advantage of a similar task and dataset: Abductive reasoning in narrative text (Bhagavatula et al., 2020).However, we show that simply using a state-of-the-art seq2seq model fine-tuned on this data might not generate meaningful implicit premises associated with the given enthymemes.We demonstrate that encoding discourse-aware commonsense during fine-tuning improves the quality of the generated implicit premises and outperforms all other baselines both in automatic and human evaluations on three different datasets. Tuhin Chakrabarty, Aadit Trivedi, Smaranda Muresan |
EMNLP (1) | 3 |
| 2021 | ENTRUST: Argument Reframing with Language Models and EntailmentabstractTuhin Chakrabarty, Christopher Hidey, Smaranda Muresan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tuhin Chakrabarty, Christopher Hidey, Smaranda Muresan |
NAACL-HLT | 3 |
| 2021 | MERMAID: Metaphor Generation with Symbolism and Discriminative DecodingabstractTuhin Chakrabarty, Xurui Zhang, Smaranda Muresan, Nanyun Peng. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tuhin Chakrabarty, Xurui Zhang, Smaranda Muresan, Nanyun Peng 0001 |
NAACL-HLT | 3 |
| 2021 | Emotion-Infused Models for Explainable Psychological Stress DetectionabstractElsbeth Turcan, Smaranda Muresan, Kathleen McKeown. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Elsbeth Turcan, Smaranda Muresan, Kathy McKeown |
NAACL-HLT | 2 |
| 2021 | What to Fact-Check: Guiding Check-Worthy Information Detection in News Articles through Argumentative Discourse StructureabstractMost existing methods for automatic factchecking start with a precompiled list of claims to verify.We investigate the understudied problem of determining what statements in news articles are worthy to factcheck.We annotate the argument structure of 95 news articles in the climate change domain that are fact-checked by climate scientists at climatefeedback.org.We release the first multi-layer annotated corpus for both argumentative discourse structure (argument components and relations) and for factchecked statements in news articles.We discuss the connection between argument structure and check-worthy statements and develop several baseline models for detecting checkworthy statements in the climate change domain.Our preliminary results show that using information about argumentative discourse structure shows slight but statistically significant improvement over a baseline of local discourse structure. Tariq Alhindi, Brennan Xavier McManus, Smaranda Muresan |
SIGDIAL | 3 |
| 2020 | R^3: Reverse, Retrieve, and Rank for Sarcasm Generation with Commonsense KnowledgeabstractWe propose an unsupervised approach for sarcasm generation based on a non-sarcastic input sentence. Our method employs a retrieve-and-edit framework to instantiate two major characteristics of sarcasm: reversal of valence and semantic incongruity with the context, which could include shared commonsense or world knowledge between the speaker and the listener. While prior works on sarcasm generation predominantly focus on context incongruity, we show that combining valence reversal and semantic incongruity based on the commonsense knowledge generates sarcasm of higher quality. Human evaluation shows that our system generates sarcasm better than humans 34% of the time, and better than a reinforced hybrid baseline 90% of the time. Tuhin Chakrabarty, Debanjan Ghosh, Smaranda Muresan, Nanyun Peng 0001 |
ACL | 3 |
| 2020 | DeSePtion: Dual Sequence Prediction and Adversarial Examples for Improved Fact-CheckingabstractChristopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia, Kriste Krstovski, Mona Diab, Smaranda Muresan. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Christopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia, Kriste Krstovski, Mona T. Diab, Smaranda Muresan |
ACL | 7 |
| 2020 | Fact vs. Opinion: the Role of Argumentation Features in News ClassificationabstractA 2018 study led by the Media Insight Project showed that most journalists think that a clear marking of what is news reporting and what is commentary or opinion (e.g., editorial, op-ed) is essential for gaining public trust.We present an approach to classify news articles into news stories (i.e., reporting of factual information) and opinion pieces using models that aim to supplement the article content representation with argumentation features.Our hypothesis is that the nature of argumentative discourse is important in distinguishing between news stories and opinion articles.We show that argumentation features outperform linguistic features used previously and improve on fine-tuned transformer-based models when tested on data from publishers unseen in training. Tariq Alhindi, Smaranda Muresan, Daniel Preotiuc-Pietro |
COLING | 2 |
| 2020 | To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence TaggingabstractKasturi Bhattacharjee, Miguel Ballesteros, Rishita Anubhai, Smaranda Muresan, Jie Ma, Faisal Ladhak, Yaser Al-Onaizan. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Kasturi Bhattacharjee, Miguel Ballesteros, Rishita Anubhai, Smaranda Muresan, Jie Ma 0005, Faisal Ladhak, Yaser Al-Onaizan |
EMNLP (1) | 4 |
| 2020 | Generating similes effortlessly like a Pro: A Style Transfer Approach for Simile GenerationabstractLiterary tropes, from poetry to stories, are at the crux of human imagination and communication.Figurative language, such as a simile, goes beyond plain expressions to give readers new insights and inspirations.We tackle the problem of simile generation.Generating a simile requires proper understanding for effective mapping of properties between two concepts.To this end, we first propose a method to automatically construct a parallel corpus by transforming a large number of similes collected from Reddit to their literal counterpart using structured common sense knowledge.We then fine-tune a pretrained sequence to sequence model, BART (Lewis et al., 2019), on the literal-simile pairs to generate novel similes given a literal sentence.Experiments show that our approach generates 88% novel similes that do not share properties with the training data.Human evaluation on an independent set of literal statements shows that our model generates similes better than two literary experts 37% 1 of the times, and three baseline systems including a recent metaphor generation model 71% 2 of the times when compared pairwise.3 We also show how replacing literal sentences with similes from our best model in machine generated stories improves evocativeness and leads to better acceptance by human judges.* The research was conducted when the author was at USC/ISI.1 We average 32.6% and 41.3% for 2 humans. 2 We average 82% ,63% and 68% for three baselines.3 The simile in the title is generated by our best model.Input: Generating similes effortlessly, output: Generating similes like a Pro. Tuhin Chakrabarty, Smaranda Muresan, Nanyun Peng 0001 |
EMNLP (1) | 2 |
| 2020 | Unsupervised Cross-Lingual Part-of-Speech Tagging for Truly Low-Resource ScenariosabstractWe describe a fully unsupervised cross-lingual transfer approach for part-of-speech (POS) tagging under a truly low resource scenario.We assume access to parallel translations between the target language and one or more source languages for which POS taggers are available.We use the Bible as parallel data in our experiments: small size, out-of-domain and covering many diverse languages.Our approach innovates in three ways: 1) a robust approach of selecting training instances via cross-lingual annotation projection that exploits best practices of unsupervised type and token constraints, word-alignment confidence and density of projected POS, 2) a Bi-LSTM architecture that uses contextualized word embeddings, affix embeddings and hierarchical Brown clusters, and 3) an evaluation on 12 diverse languages in terms of language family and morphological typology.In spite of the use of limited and out-of-domain parallel data, our experiments demonstrate significant improvements in accuracy over previous work.In addition, we show that using multi-source information, either via projection or output combination, improves the performance for most target languages. Ramy Eskander, Smaranda Muresan |
EMNLP (1) | 2 |
| 2020 | MorphAGram, Evaluation and Framework for Unsupervised Morphological SegmentationabstractComputational morphological segmentation has been an active research topic for decades as it is beneficial for many natural language processing tasks. With the high cost of manually labeling data for morphology and the increasing interest in low-resource languages, unsupervised morphological segmentation has become essential for processing a typologically diverse set of languages, whether high-resource or low-resource. In this paper, we present and release MorphAGram, a publicly available framework for unsupervised morphological segmentation that uses Adaptor Grammars (AG) and is based on the work presented by Eskander et al. (2016). We conduct an extensive quantitative and qualitative evaluation of this framework on 12 languages and show that the framework achieves state-of-the-art results across languages of different typologies (from fusional to polysynthetic and from high-resource to low-resource). Ramy Eskander, Francesca Callejas, Elizabeth Nichols, Judith L. Klavans, Smaranda Muresan |
LREC | 5 |
| 2019 | AMPERSAND: Argument Mining for PERSuAsive oNline DiscussionsabstractTuhin Chakrabarty, Christopher Hidey, Smaranda Muresan, Kathy McKeown, Alyssa Hwang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tuhin Chakrabarty, Christopher Hidey, Smaranda Muresan, Kathy McKeown, Alyssa Hwang |
EMNLP/IJCNLP (1) | 3 |
| 2018 | "With 1 Follower I Must Be AWESOME : P." Exploring the Role of Irony Markers in Irony Recognition
Debanjan Ghosh, Smaranda Muresan |
ICWSM | 2 |
| 2018 | A Multi-layer Annotated Corpus of Argumentative Text: From Argument Schemes to Discourse Relations
Elena Musi, Manfred Stede, Leonard Kriese, Smaranda Muresan, Andrea Rocci |
LREC | 4 |
| 2018 | Sarcasm Analysis Using Conversation ContextabstractComputational models for sarcasm detection have often relied on the content of utterances in isolation. However, the speaker’s sarcastic intent is not always apparent without additional context. Focusing on social media discussions, we investigate three issues: (1) does modeling conversation context help in sarcasm detection? (2) can we identify what part of conversation context triggered the sarcastic reply? and (3) given a sarcastic post that contains multiple sentences, can we identify the specific sentence that is sarcastic? To address the first issue, we investigate several types of Long Short-Term Memory (LSTM) networks that can model both the conversation context and the current turn. We show that LSTM networks with sentence-level attention on context and current turn, as well as the conditional LSTM network, outperform the LSTM model that reads only the current turn. As conversation context, we consider the prior turn, the succeeding turn, or both. Our computational models are tested on two types of social media platforms: Twitter and discussion forums. We discuss several differences between these data sets, ranging from their size to the nature of the gold-label annotations. To address the latter two issues, we present a qualitative analysis of the attention weights produced by the LSTM models (with attention) and discuss the results compared with human performance on the two tasks. Debanjan Ghosh, Alexander R. Fabbri, Smaranda Muresan |
Comput. Linguistics | 3 |
| 2017 | The Role of Conversation Context for Sarcasm Detection in Online InteractionsabstractComputational models for sarcasm detection have often relied on the content of utterances in isolation.However, speaker's sarcastic intent is not always obvious without additional context.Focusing on social media discussions, we investigate two issues: (1) does modeling of conversation context help in sarcasm detection and (2) can we understand what part of conversation context triggered the sarcastic reply.To address the first issue, we investigate several types of Long Short-Term Memory (LSTM) networks that can model both the conversation context and the sarcastic response.1 We show that the conditional LSTM network (Rocktäschel et al., 2015) and LSTM networks with sentence level attention on context and response outperform the LSTM model that reads only the response.To address the second issue, we present a qualitative analysis of attention weights produced by the LSTM models with attention and discuss the results compared with human performance on the task. Debanjan Ghosh, Alexander R. Fabbri, Smaranda Muresan |
SIGDIAL Conference | 3 |
| 2016 | Identification of nonliteral language in social media: A case study on sarcasmabstractWith the rapid development of social media, spontaneously user‐generated content such as tweets and forum posts have become important materials for tracking people's opinions and sentiments online. A major hurdle for current state‐of‐the‐art automatic methods for sentiment analysis is the fact that human communication often involves the use of sarcasm or irony, where the author means the opposite of what she/he says. Sarcasm transforms the polarity of an apparently positive or negative utterance into its opposite. Lack of naturally occurring utterances labeled for sarcasm is one of the key problems for the development of machine‐learning methods for sarcasm detection. We report on a method for constructing a corpus of sarcastic Twitter messages in which determination of the sarcasm of each message has been made by its author. We use this reliable corpus to compare sarcastic utterances in Twitter to utterances that express positive or negative attitudes without sarcasm. We investigate the impact of lexical and pragmatic factors on machine‐learning effectiveness for identifying sarcastic utterances and we compare the performance of machine‐learning techniques and human judges on this task. Smaranda Muresan, Roberto I. González-Ibáñez, Debanjan Ghosh, Nina Wacholder |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2015 | Mining a Written Values Affirmation Intervention to Identify the Unique Linguistic Features of Stigmatized Groups
Travis Riddle, Sowmya Bhagavatula, Weiwei Guo, Smaranda Muresan, Geoff Cohen, Jonathan E. Cook 0002, Valerie Purdie-Vaughns |
EDM | 4 |
| 2015 | Sarcastic or Not: Word Embeddings to Predict the Literal or Sarcastic Meaning of WordsabstractSarcasm is generally characterized as a figure of speech that involves the substitution of a literal by a figurative meaning, which is usually the opposite of the original literal meaning.We re-frame the sarcasm detection task as a type of word sense disambiguation problem, where the sense of a word is either literal or sarcastic.We call this the Literal/Sarcastic Sense Disambiguation (LSSD) task.We address two issues: 1) how to collect a set of target words that can have either literal or sarcastic meanings depending on context; and 2) given an utterance and a target word, how to automatically detect whether the target word is used in the literal or the sarcastic sense.For the latter, we investigate several distributional semantics methods and show that a Support Vector Machines (SVM) classifier with a modified kernel using word embeddings achieves a 7-10% F1 improvement over a strong lexical baseline. Debanjan Ghosh, Weiwei Guo, Smaranda Muresan |
EMNLP | 3 |
| 2013 | Inducing terminologies from text: A case study for the consumer health domainabstractSpecialized medical ontologies and terminologies, such asSNOMED CTand theUnifiedMedicalLanguageSystem (UMLS), have been successfully leveraged in medical information systems to provide a standard web‐accessible medium for interoperability, access, and reuse. However, these clinically oriented terminologies and ontologies cannot provide sufficient support when integrated into consumer‐oriented applications, because these applications must “understand” both technical and lay vocabulary. The latter is not part of these specialized terminologies and ontologies. In this article, we propose a two‐step approach for building consumer health terminologies from text: 1)automatic extractionof definitions from consumer‐oriented articles and web documents, which reflects language in use, rather than relying solely on dictionaries, and 2)learningto map definitions expressed in natural language to terminological knowledge by inducing a syntactic‐semantic grammar rather than using hand‐written patterns or grammars. We present quantitative and qualitative evaluations of our two‐step approach, which show that our framework could be used to induce consumer health terminologies from text. Smaranda Muresan, Judith L. Klavans |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2012 | Combining Social Cognitive Theories with Linguistic Features for Multi-genre Sentiment Analysis
Hao Li 0031, Yu Chen 0022, Heng Ji 0001, Smaranda Muresan, Dequan Zheng |
PACLIC | 4 |
| 2011 | Learning for Deep Language UnderstandingabstractThe paper addresses the problem of learning to parse sentences to logical representations of their underlying meaning, by inducing a syntactic-semantic grammar. The approach uses a class of grammars which has been proven to be learnable from representative examples. In this paper, we introduce tractable learning algorithms for learning this class of grammars, comparing them in terms of a-priori knowledge needed by the learner, hypothesis space and algorithm complexity. We present experimental results on learning tense, aspect, modality and negation of verbal constructions. Smaranda Muresan |
IJCAI | 1 |
| 2011 | Network based models of cognitive and social dynamics of human languages
Animesh Mukherjee 0001, Monojit Choudhury, Samer Hassan 0002, Smaranda Muresan |
Comput. Speech Lang. | 4 |
| 2010 | Ontology-Based Semantic Interpretation as Grammar Rule Constraints
Smaranda Muresan |
CICLing | 1 |
| 2008 | Generalizing Word Lattice Translation
Chris Dyer, Smaranda Muresan, Philip Resnik |
ACL | 2 |
| 2007 | Grammar Approximation by Representative Sublanguage: A New Model for Language Learning
Smaranda Muresan, Owen Rambow |
ACL | 1 |
| 2004 | Inducing Constraint-Based Grammars using a Domain Ontology
Smaranda Muresan |
AAAI | 1 |
| 2002 | A Method for Automatically Building and Evaluating Dictionary Resources
Smaranda Muresan, Judith L. Klavans |
LREC | 1 |
| 2001 | Evaluation of the DEFINDER system for fully automatic glossary construction
Judith L. Klavans, Smaranda Muresan |
AMIA | 2 |
| 2001 | Data Flow Coherence Criteria in ILP ToolsabstractIn this paper we present a new method that uses data flow coherence criteria in definite logic program generation. We outline three main advantages of these criteria supported by our results: (i) drastically pruning the search space (around 90%), (ii) reducing the set of positive examples and reducing or even removing the need for the set of negative examples, and (iii) allowing the induction of predicates that are difficult or even impossible to generate by other methods. Besides these criteria, the approach takes into consideration the program termination condition for recursive predicates. The paper outlines some theoretical issues and implementation aspects of our system for automatic logic program induction. Smaranda Muresan, Tudor Muresan, Rodica Potolea |
ICTAI | 1 |
| 2000 | DEFINDER: Rule-based Methods for the Extraction of Medical Terminology and their Associated Definitions from On-line Text
Judith L. Klavans, Smaranda Muresan |
AMIA | 2 |