VLDB 2026 Research / reviewers in the wild / expert
Rada Mihalcea
dblp:m/RadaMihalcea · also Rada Flavia Mihalcea
· DBLP profile ↗
201ranked-venue papers
39as first author
62since 2021 · last 2026
0000-0002-0767-6703ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 177 · 37 first-author · 54 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 17 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Security and privacy · 2Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Culture Affordance Atlas: Reconciling Object Diversity Through Functional MappingabstractCulture shapes the objects people use and for what purposes, yet mainstream Vision-Language (VL) datasets frequently exhibit cultural biases, disproportionately favoring higher-income, Western contexts. This imbalance reduces model generalizability and perpetuates performance disparities, especially impacting lower-income and non-Western communities. To address these disparities, we propose a novel function-centric framework that categorizes objects by the functions they fulfill, across diverse cultural and economic contexts. We implement this framework by creating the Culture Affordance Atlas, a re-annotated and culturally grounded restructuring of the Dollar Street dataset spanning 46 functions and 288 objects. Through extensive empirical analyses using the CLIP model, we demonstrate that function-centric labels substantially reduce socioeconomic performance gaps between high and low-income groups by a median of 6 pp (statistically significant), improving model effectiveness for lower income contexts. Furthermore, our analyses reveals numerous culturally essential objects that are frequently overlooked in prominent VL datasets. Our contributions offer a scalable pathway toward building inclusive VL datasets and equitable AI systems. Joan Nwatu, Longju Bai, Oana Ignat, Rada Mihalcea |
AAAI | 4 |
| 2026 | Evaluating Language Models for Assessing Counselor ReflectionsabstractReflective listening is a fundamental communication skill in behavioral health counseling. It enables counselors to demonstrate an understanding of and empathy for clients’ experiences and concerns. Training to acquire and refine reflective listening skills is essential for counseling proficiency. Yet, it faces significant barriers, notably the need for specialized and timely feedback to improve counseling skills. In this work, we evaluate and compare several computational models, including transformer-based architectures, for their ability to assess the quality of counselors’ reflective listening skills. We explore a spectrum of neural-based models, ranging from compact, specialized RoBERTa models to advanced large-scale language models such as Flan, Mistral, and GPT-3.5, to score psychotherapy reflections. We introduce a psychotherapy dataset that encompasses three basic levels of reflective listening skills. Through comparative experiments, we show that a fine-tuned small RoBERTa model with a custom learning objective (Prompt-Aware margIn Ranking (PAIR)) effectively provides constructive feedback to counselors in training. This study also highlights the potential of machine learning in enhancing the training process for motivational interviewing (MI) by offering scalable and effective feedback alternatives for counseling training. Do June Min, Verónica Pérez-Rosas, Kenneth Resnicow, Rada Mihalcea |
ACM Trans. Comput. Heal. | 4 |
| 2025 | Why AI Is WEIRD and Shouldn't Be This Way: Towards AI for Everyone, with Everyone, by EveryoneabstractThis paper presents a vision for creating AI systems that are inclusive at every stage of development, from data collection to model design and evaluation. We address key limitations in the current AI pipeline and its WEIRD* representation, such as lack of data diversity, biases in model performance, and narrow evaluation metrics. We also focus on the need for diverse representation among the developers of these systems, as well as incentives that are not skewed toward certain groups. We highlight opportunities to develop AI systems that are for everyone (with diverse stakeholders in mind), with everyone (inclusive of diverse data and annotators), and by everyone (designed and developed by a globally diverse workforce). *WEIRD = an acronym coined by Joseph Henrich to highlight the coverage limitations of many psychological studies, referring to populations that are Western, Educated, Industrialized, Rich, and Democratic; while we do not fully adopt this term for AI, as its current scope does not perfectly align with the WEIRD dimensions, we believe that today's AI has a similarly "weird" coverage, particularly in terms of who is involved in its development and who benefits from it. Rada Mihalcea, Oana Ignat, Longju Bai, Angana Borah, Luis Chiruzzo, Zhijing Jin 0001, Claude Kwizera, Joan Nwatu, Soujanya Poria, Thamar Solorio |
AAAI | 1 |
| 2025 | Are Language Models Consequentialist or Deontological Moral Reasoners?abstractKeenan Samway, Max Kleiman-Weiner, David Guzman Piedrahita, Rada Mihalcea, Bernhard Schölkopf, Zhijing Jin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Keenan Samway, Max Kleiman-Weiner, David Guzman Piedrahita, Rada Mihalcea, Bernhard Schölkopf, Zhijing Jin 0001 |
EMNLP | 4 |
| 2025 | Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?abstractThe value orientation of Large Language Models (LLMs) has been extensively studied, as it can shape user experiences across demographic groups.However, two key challenges remain:(1) the lack of systematic comparison across value probing strategies, despite the Multiple Choice Question (MCQ) setting being vulnerable to perturbations, and (2) the uncertainty over whether probed values capture in-context information or predict models' real-world actions.In this paper, we systematically compare three widely used value probing methods: token likelihood, sequence perplexity, and text generation.Our results show that all three methods exhibit large variances under non-semantic perturbations in prompts and option formats, with sequence perplexity being the most robust overall.We further introduce two tasks to assess expressiveness: demographic prompting, testing whether probed values adapt to cultural context; and value-action agreement, testing the alignment of probed values with value-based actions.We find that demographic context has little effect on the text generation method, and probed values only weakly correlate with action preferences across all methods.Our work highlights the instability and the limited expressive power of current value probing methods, calling for more reliable LLM value representations. Mehar Singh, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Rada Mihalcea |
EMNLP | 6 |
| 2025 | Language Model Alignment in Multilingual Trolley ProblemsabstractWe evaluate the moral alignment of large language models (LLMs) with human preferences in multilingual trolley problems. Building on the Moral Machine experiment, which captures over 40 million human judgments across 200+ countries, we develop a cross-lingual corpus of moral dilemma vignettes in over 100 languages called MultiTP. This dataset enables the assessment of LLMs' decision-making processes in diverse linguistic contexts. Our analysis explores the alignment of 19 different LLMs with human judgments, capturing preferences across six moral dimensions: species, gender, fitness, status, age, and the number of lives involved. By correlating these preferences with the demographic distribution of language speakers and examining the consistency of LLM responses to various prompt paraphrasings, our findings provide insights into cross-lingual and ethical biases of LLMs and their intersection. We discover significant variance in alignment across languages, challenging the assumption of uniform moral reasoning in AI systems and highlighting the importance of incorporating diverse perspectives in AI ethics. The results underscore the need for further research on the integration of multilingual dimensions in responsible AI research to ensure fair and equitable AI interactions worldwide. Zhijing Jin 0001, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu 0004, Fernando Gonzalez Adauto, Francesco Ortu, András Strausz, Mrinmaya Sachan, Rada Mihalcea, Yejin Choi 0001, Bernhard Schölkopf |
ICLR | 10 |
| 2025 | The Agent Paradox: Can Multi-Agent Systems Replicate the Complexity of Human Cognition and Social Behavior?
Rada Mihalcea |
AAMAS | 1 |
| 2025 | The Power of Many: Multi-Agent Multimodal Models for Cultural Image CaptioningabstractLongju Bai, Angana Borah, Oana Ignat, Rada Mihalcea. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Longju Bai, Angana Borah, Oana Ignat, Rada Mihalcea |
NAACL (Long Papers) | 4 |
| 2025 | Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal ModelsabstractJoan Nwatu, Oana Ignat, Rada Mihalcea. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Joan Nwatu, Oana Ignat, Rada Mihalcea |
NAACL (Long Papers) | 3 |
| 2025 | Speech-Integrated Modeling for Behavioral Coding in CounselingabstractComputational models of psychotherapy often ignore vocal cues by relying solely on text. To address this, we propose MISQ, a framework that integrates speech features directly into language models using a speech encoder and lightweight adapter. MISQ improves behavioral analysis in counseling conversations, achieving ~5% relative gains over text-only or indirect speech methods—underscoring the value of vocal signals like tone and prosody. Do June Min, Verónica Pérez-Rosas, Kenneth Resnicow, Rada Mihalcea |
SIGDIAL | 4 |
| 2024 | EmoBench: Evaluating the Emotional Intelligence of Large Language ModelsabstractSahand Sabour, Siyang Liu, Zheyuan Zhang, June Liu, Jinfeng Zhou, Alvionna Sunaryo, Tatia Lee, Rada Mihalcea, Minlie Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Sahand Sabour, Siyang Liu 0003, Zheyuan Zhang 0002, June M. Liu, Jinfeng Zhou, Alvionna S. Sunaryo, Tatia M. C. Lee, Rada Mihalcea, Minlie Huang |
ACL (1) | 8 |
| 2024 | Towards Dog Bark Decoding: Leveraging Human Speech Processing for Automated Bark ClassificationabstractSimilar to humans, animals make extensive use of verbal and non-verbal forms of communication, including a large range of audio signals. In this paper, we address dog vocalizations and explore the use of self-supervised speech representation models pre-trained on human speech to address dog bark classification tasks that find parallels in human-centered tasks in speech recognition. We specifically address four tasks: dog recognition, breed identification, gender classification, and context grounding. We show that using speech embedding representations significantly improves over simpler classification baselines. Further, we also find that models pre-trained on large human speech acoustics can provide additional performance boosts on several tasks. Artem Abzaliev, Humberto Pérez Espinosa, Rada Mihalcea |
LREC/COLING | 3 |
| 2024 | Analyzing Occupational Distribution Representation in Japanese Language ModelsabstractRecent advances in large language models (LLMs) have enabled users to generate fluent and seemingly convincing text. However, these models have uneven performance in different languages, which is also associated with undesirable societal biases toward marginalized populations. Specifically, there is relatively little work on Japanese models, despite it being the thirteenth most widely spoken language. In this work, we first develop three Japanese language prompts to probe LLMs’ understanding of Japanese names and their association between gender and occupations. We then evaluate a variety of English, multilingual, and Japanese models, correlating the models’ outputs with occupation statistics from the Japanese Census Bureau from the last 100 years. Our findings indicate that models can associate Japanese names with the correct gendered occupations when using constrained decoding. However, with sampling or greedy decoding, Japanese language models have a preference for a small set of stereotypically gendered occupations, and multilingual models, though trained on Japanese, are not always able to understand Japanese prompts. Katsumi Ibaraki, Winston Wu, Lu Wang 0008, Rada Mihalcea |
LREC/COLING | 4 |
| 2024 | Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation CostabstractCurrent foundation models have shown impressive performance across various tasks. However, several studies have revealed that these models are not effective for everyone due to the imbalanced geographical and economic representation of the data used in the training process. Most of this data comes from Western countries, leading to poor results for underrepresented countries. To address this issue, more data needs to be collected from these countries, but the cost of annotation can be a significant bottleneck. In this paper, we propose methods to identify the data to be annotated to balance model performance and annotation costs. Our approach first involves finding the countries with images of topics (objects and actions) most visually distinct from those already in the training datasets used by current large vision-language foundation models. Next, we identify countries with higher visual similarity for these topics and show that using data from these countries to supplement the training data improves model performance and reduces annotation costs. The resulting lists of countries and corresponding topics are made available at https://github.com/MichiganNLP/visual_diversity_budget. Oana Ignat, Longju Bai, Joan Nwatu, Rada Mihalcea |
LREC/COLING | 4 |
| 2024 | Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language ModelsabstractRecent progress in large language models (LLMs) has enabled the deployment of many generative NLP applications. At the same time, it has also led to a misleading public discourse that “it’s all been solved.” Not surprisingly, this has, in turn, made many NLP researchers – especially those at the beginning of their careers – worry about what NLP research area they should focus on. Has it all been solved, or what remaining questions can we work on regardless of LLMs? To address this question, this paper compiles NLP research directions rich for exploration. We identify fourteen different research areas encompassing 45 research directions that require new research and are not directly solvable by LLMs. While we identify many research areas, many others exist; we do not cover areas currently addressed by LLMs, but where LLMs lag behind in performance or those focused on LLM development. We welcome suggestions for other research directions to include: https://bit.ly/nlp-era-llm. Oana Ignat, Zhijing Jin 0001, Artem Abzaliev, Laura Biester, Santiago Castro, Naihao Deng, Xinyi Gao 0004, Aylin Gunal, Jacky He, Ashkan Kazemi, Muhammad Khalifa, Namho Koh, Andrew Lee 0001, Siyang Liu 0003, Do June Min, Shinka Mori, Joan Nwatu, Verónica Pérez-Rosas, Zekun Wang 0002, Winston Wu, Rada Mihalcea |
LREC/COLING | 22 |
| 2024 | Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection GenerationabstractIn this paper, we study the problem of multi-reward reinforcement learning to jointly optimize for multiple text qualities for natural language generation. We focus on the task of counselor reflection generation, where we optimize the generators to simultaneously improve the fluency, coherence, and reflection quality of generated counselor responses. We introduce two novel bandit methods, DynaOpt and C-DynaOpt, which rely on the broad strategy of combining rewards into a single value and optimizing them simultaneously. Specifically, we employ non-contextual and contextual multi-arm bandits to dynamically adjust multiple reward weights during training. Through automatic and manual evaluations, we show that our proposed techniques, DynaOpt and C-DynaOpt, outperform existing naive and bandit baselines, showcasing their potential for enhancing language models. Do June Min, Verónica Pérez-Rosas, Kenneth Resnicow, Rada Mihalcea |
LREC/COLING | 4 |
| 2024 | Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated DataabstractSynthetic data generation has the potential to impact applications and domains with scarce data. However, before such data is used for sensitive tasks such as mental health, we need an understanding of how different demographics are represented in it. In our paper, we analyze the potential of producing synthetic data using GPT-3 by exploring the various stressors it attributes to different race and gender combinations, to provide insight for future researchers looking into using LLMs for data generation. Using GPT-3, we develop HeadRoom, a synthetic dataset of 3,120 posts about depression-triggering stressors, by controlling for race, gender, and time frame (before and after COVID-19). Using this dataset, we conduct semantic and lexical analyses to (1) identify the predominant stressors for each demographic group; and (2) compare our synthetic data to a human-generated dataset. We present the procedures to generate queries to develop depression data using GPT-3, and conduct analyzes to uncover the types of stressors it assigns to demographic groups, which could be used to test the limitations of LLMs for synthetic data generation for depression data. Our findings show that synthetic data mimics some of the human-generated data distribution for the predominant depression stressors across diverse demographics. Shinka Mori, Oana Ignat, Andrew Lee 0001, Rada Mihalcea |
LREC/COLING | 4 |
| 2024 | A Comparative Multidimensional Analysis of Empathetic SystemsabstractRecently, empathetic dialogue systems have received significant attention.While some researchers have noted limitations, e.g., that these systems tend to generate generic utterances, no study has systematically verified these issues.We survey 21 systems, asking what progress has been made on the task.We observe multiple limitations of current evaluation procedures.Most critically, studies tend to rely on a single non-reproducible empathy score, which inadequately reflects the multidimensional nature of empathy.To better understand the differences between systems, we comprehensively analyze each system with automated methods that are grounded in a variety of aspects of empathy.We find that recent systems lack three important aspects of empathy: specificity, reflection levels, and diversity.Based on our results, we discuss problematic behaviors that may have gone undetected in prior evaluations, and offer guidance for developing future systems. 1 Andrew Lee 0001, Jonathan K. Kummerfeld, Lawrence C. An, Rada Mihalcea |
EACL (1) | 4 |
| 2024 | The Generation Gap: Exploring Age Bias in the Value Systems of Large Language ModelsabstractWe explore the alignment of values in Large Language Models (LLMs) with specific age groups, leveraging data from the World Value Survey across thirteen categories.Through a diverse set of prompts tailored to ensure response robustness, we find a general inclination of LLM values towards younger demographics, especially when compared to the US population.Although a general inclination can be observed, we also found that this inclination toward younger groups can be different across different value categories.Additionally, we explore the impact of incorporating age identity information in prompts and observe challenges in mitigating value discrepancies with different age cohorts.Our findings highlight the age bias in LLMs and provide insights for future work.Materials for our analysis are available at https://github.com/Michiga nNLP/Age-Bias-In-LLMs Siyang Liu 0003, Trisha Maturi, Bowen Yi 0001, Rada Mihalcea |
EMNLP | 5 |
| 2024 | Unsupervised Discrete Representations of American Sign LanguageabstractMany modalities are naturally represented as continuous signals, making it difficult to use them with models that expect discrete units, such as LLMs.In this paper, we explore the use of audio compression techniques for the discrete representation of the gestures used in sign language.We train a tokenizer for American Sign Language (ASL) fingerspelling, which discretizes sequences of fingerspelling signs into tokens.We also propose a loss function to improve the interpretability of these tokens such that they preserve both the semantic and the visual information of the signal.We show that the proposed method improves the performance of the discretized sequence on downstream tasks. Artem Abzaliev, Rada Mihalcea |
EMNLP | 2 |
| 2024 | Can Large Language Models Infer Causation from Correlation?abstractCausal inference is one of the hallmarks of human intelligence. While the field of CausalNLP has attracted much interest in the recent years, existing causal inference datasets in NLP primarily rely on discovering causality from empirical knowledge (e.g., commonsense knowledge). In this work, we propose the first benchmark dataset to test the pure causal inference skills of large language models (LLMs). Specifically, we formulate a novel task Corr2Cause, which takes a set of correlational statements and determines the causal relationship between the variables. We curate a large-scale dataset of more than 200K samples, on which we evaluate seventeen existing LLMs. Through our experiments, we identify a key shortcoming of LLMs in terms of their causal inference skills, and show that these models achieve almost close to random performance on the task. This shortcoming is somewhat mitigated when we try to re-purpose LLMs for this skill via finetuning, but we find that these models still fail to generalize – they can only perform causal inference in in-distribution settings when variable names and textual expressions used in the queries are similar to those in the training set, but fail in out-of-distribution settings generated by perturbing these queries. Corr2Cause is a challenging task for LLMs, and can be helpful in guiding future research on improving LLMs’ pure reasoning skills and generalizability. Our data is at https://huggingface.co/datasets/causalnlp/corr2cause. Our code is at https://github.com/causalNLP/corr2cause. Zhijing Jin 0001, Jiarui Liu 0004, Zhiheng Lyu, Spencer Poff, Mrinmaya Sachan, Rada Mihalcea, Mona T. Diab, Bernhard Schölkopf |
ICLR | 6 |
| 2024 | A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityabstractWhile alignment algorithms are commonly used to tune pre-trained language models towards user preferences, we lack explanations for the underlying mechanisms in which models become ``aligned'', thus making it difficult to explain phenomena like jailbreaks. In this work we study a popular algorithm, direct preference optimization (DPO), and the mechanisms by which it reduces toxicity. Namely, we first study how toxicity is represented and elicited in pre-trained language models (GPT2-medium, Llama2-7b). We then apply DPO with a carefully crafted pairwise dataset to reduce toxicity. We examine how the resulting models avert toxic outputs, and find that capabilities learned from pre-training are not removed, but rather bypassed. We use this insight to demonstrate a simple method to un-align the models, reverting them back to their toxic behavior. Andrew Lee 0001, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, Rada Mihalcea |
ICML | 6 |
| 2024 | Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference OptimizationabstractPeer Reviewed Navonil Majumder, Chia-Yu Hung, Deepanway Ghosal, Wei-Ning Hsu, Rada Mihalcea, Soujanya Poria |
ACM Multimedia | 5 |
| 2024 | Understanding the Capabilities and Limitations of Large Language Models for Cultural CommonsenseabstractSiqi Shen, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Soujanya Poria, Rada Mihalcea. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Soujanya Poria, Rada Mihalcea |
NAACL-HLT | 6 |
| 2024 | Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM AgentsabstractAs AI systems pervade human life, ensuring that large language models (LLMs) make safe decisions remains a significant challenge. We introduce the Governance of the Commons Simulation (GovSim), a generative simulation platform designed to study strategic interactions and cooperative decision-making in LLMs. In GovSim, a society of AI agents must collectively balance exploiting a common resource with sustaining it for future use. This environment enables the study of how ethical considerations, strategic planning, and negotiation skills impact cooperative outcomes. We develop an LLM-based agent architecture and test it with the leading open and closed LLMs. We find that all but the most powerful LLM agents fail to achieve a sustainable equilibrium in GovSim, with the highest survival rate below 54%. Ablations reveal that successful multi-agent communication between agents is critical for achieving cooperation in these cases. Furthermore, our analyses show that the failure to achieve sustainable cooperation in most LLMs stems from their inability to formulate and analyze hypotheses about the long-term effects of their actions on the equilibrium of the group. Finally, we show that agents that leverage "Universalization"-based reasoning, a theory of moral thinking, are able to achieve significantly better sustainability. Taken together, GovSim enables us to study the mechanisms that underlie sustainable self-government with specificity and scale. We open source the full suite of our research results, including the simulation environment, agent prompts, and a comprehensive web interface. Giorgio Piatti, Zhijing Jin 0001, Max Kleiman-Weiner, Bernhard Schölkopf, Mrinmaya Sachan, Rada Mihalcea |
NeurIPS | 6 |
| 2024 | CVQA: Culturally-diverse Multilingual Visual Question Answering BenchmarkabstractVisual Question Answering~(VQA) is an important task in multimodal AI, which requires models to understand and reason on knowledge present in visual and textual data. However, most of the current VQA datasets and models are primarily focused on English and a few major world languages, with images that are Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, some datasets extend the text to other languages, either via translation or some other approaches, but usually keep the same images, resulting in narrow cultural representation. To address these limitations, we create CVQA, a new Culturally-diverse Multilingual Visual Question Answering benchmark dataset, designed to cover a rich set of languages and regions, where we engage native speakers and cultural experts in the data collection process. CVQA includes culturally-driven images and questions from across 28 countries in four continents, covering 26 languages with 11 scripts, providing a total of 9k questions. We benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and we show that the dataset is challenging for the current state-of-the-art models. This benchmark will serve as a probing evaluation suite for assessing the cultural bias of multimodal models and hopefully encourage more research efforts towards increasing cultural awareness and linguistic diversity in this field. Chenyang Lyu, Haryo Akbarianto Wibowo, Santiago Góngora, Aishik Mandal, Sukannya Purkayastha, Jesús-Germán Ortiz-Barajas, Emilio Villa-Cueva, Jinheon Baek, Soyeong Jeong, Injy Hamed, Zheng Wei Lim, Paula Mónica Silva, Jocelyn Dunstan, Mélanie Jouitteau, David Le Meur, Joan Nwatu, Ganzorig Batnasan, Munkh-Erdene Otgonbold, Munkhjargal Gochoo, Guido Ivetta, Luciana Benotti, Laura Alonso Alemany, Hernán Maina, Jiahui Geng, Tiago Timponi Torrent, Frederico Belcavello, Marcelo Viridiano, Jan Christian Blaise Cruz, Dan John Velasco, Oana Ignat, Zara Burzo, Chenxi Whitehouse, Artem Abzaliev, Teresa Clifford, Grainne Caulfield, Teresa Lynn, Christian Salamea Palacios, Vladimir Araujo, Yova Kementchedjhieva, Mihail Mihaylov, Israel Abebe Azime, Henok Biadglign Ademtew, Bontu Fufa Balcha, Naome A. Etori, David Ifeoluwa Adelani, Rada Mihalcea, Atnafu Lambebo Tonja, Maria Camila Buitrago Cabrera, Gisela Vallejo, Holy Lovenia, Ruochen Zhang 0001, Marcos Estecha-Garitagoitia, Mario Rodríguez-Cantelar, Toqeer Ehsan, Rendi Chevi, Muhammad Farid Adilazuarda, Ryandito Diandaru, Samuel Cahyawijaya, Fajri Koto, Tatsuki Kuribayashi, Haiyue Song, Aditya Khandavally, Thanmay Jayakumar, Raj Dabre, Mohamed Fazli Mohamed Imam, Kumaranage Ravindu Yasas Nagasinghe, Alina Dragonetti, Luis Fernando D'Haro, Olivier Niyomugisha, Jay Gala, Pranjal A. Chitale, Fauzan Farooqui, Thamar Solorio, Alham Fikri Aji |
NeurIPS | 48 |
| 2023 | Empathy Identification Systems are not Accurately Accounting for ContextabstractUnderstanding empathy in text dialogue data is a difficult, yet critical, skill for effective human-machine interaction.In this work, we ask whether systems are making meaningful progress on this challenge.We consider a simple model that checks if an input utterance is similar to a small set of empathetic examples.Crucially, the model does not look at what the utterance is a response to, i.e., the dialogue context.This model performs comparably to prior work on standard benchmarks and even outperforms state-of-the-art models for empathetic rationale extraction by 16.7 points on T-F1 and 4.3 on IOU-F1.This indicates that current systems rely on the surface form of the response, rather than whether it is suitable in context.To confirm this, we create examples with dialogue contexts that change the interpretation of the response and show that current systems continue to label utterances as empathetic.We discuss the implications of our findings, including improvements for empathetic benchmarks and how our model can be an informative baseline. Andrew Lee 0001, Jonathan K. Kummerfeld, Lawrence C. An, Rada Mihalcea |
EACL | 4 |
| 2023 | Task-Adaptive Tokenization: Enhancing Long-Form Text Generation Efficacy in Mental Health and BeyondabstractWe propose task-adaptive tokenization 1 as a way to adapt the generation pipeline to the specifics of a downstream task and enhance long-form generation in mental health.Inspired by insights from cognitive science, our task-adaptive tokenizer samples variable segmentations from multiple outcomes, with sampling probabilities optimized based on taskspecific data.We introduce a strategy for building a specialized vocabulary and introduce a vocabulary merging protocol that allows for the integration of task-specific tokens into the pre-trained model's tokenization step.Through extensive experiments on psychological question-answering tasks in both Chinese and English, we find that our task-adaptive tokenization approach brings a significant improvement in generation performance while using up to 60% fewer tokens.Preliminary experiments point to promising results when using our tokenization approach with very large language models. Siyang Liu 0003, Naihao Deng, Sahand Sabour, Yilin Jia, Minlie Huang, Rada Mihalcea |
EMNLP | 6 |
| 2023 | Bridging the Digital Divide: Performance Variation across Socio-Economic Factors in Vision-Language ModelsabstractDespite the impressive performance of current AI models reported across various tasks, performance reports often do not include evaluations of how these models perform on the specific groups that will be impacted by these technologies.Among the minority groups underrepresented in AI, data from low-income households are often overlooked in data collection and model evaluation.We evaluate the performance of a state-of-the-art vision-language model (CLIP) on a geo-diverse dataset containing household images associated with different income values (Dollar Street) and show that performance inequality exists among households of different income levels.Our results indicate that performance for the poorer groups is consistently lower than the wealthier groups across various topics and countries.We highlight insights that can help mitigate these issues and propose actionable steps for economic-level inclusive AI development.Code is available at Analysis for Bridging the Digital Divide. Joan Nwatu, Oana Ignat, Rada Mihalcea |
EMNLP | 3 |
| 2023 | Cross-Cultural Analysis of Human Values, Morals, and Biases in Folk TalesabstractFolk tales are strong cultural and social influences in children's lives, and they are known to teach morals and values.However, existing studies on folk tales are largely limited to European tales.In our study, we compile a large corpus of over 1,900 tales originating from 27 diverse cultures across six continents.Using a range of lexicons and correlation analyses, we examine how human values, morals, and gender biases are expressed in folk tales across cultures.We discover differences between cultures in prevalent values and morals, as well as cross-cultural trends in problematic gender biases.Furthermore, we find trends of reduced value expression when examining public-domain fiction stories, extrinsically validate our analyses against the multicultural Schwartz Survey of Cultural Values, and find traditional gender biases associated with values, morals, and agency.This largescale cross-cultural study of folk tales paves the way for future studies on how literature influences and reflects cultural norms. Winston Wu, Lu Wang 0008, Rada Mihalcea |
EMNLP | 3 |
| 2023 | Non-Contact Based Modeling of EnervationabstractSignificant research is currently carried out with a focus on autonomous vehicles; research is starting to focus on areas such as the modeling of occupant states and behavioral elements. This paper contributes to this line of research by developing a pipeline that extracts physiological signals from thermal imagery and modeling occupant enervation using a fully non-contact based approach. These signals are obtained via a multimodal dataset of 36 subjects across multiple channels, including the thermal and physiological modalities. Moreover, we provide a comparative analysis of non-contact and contact based channels to model the enervation state of individuals. Our analysis indicates that non-contact physiological signals extracted from thermal imagery can reach and exceed the performance of contact-based physiological signals. In addition, modeling of enervation is possible using said non-contact physiological signals and thermal features, with an accuracy of up to 70% in identifying energized and enervated occupant states. Our findings provide a novel approach for future research and opens the possibility for integration of unrestrictive sensors in future automobiles. Kais Riani, Salem Sharak, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea, John Elson, Clay Maranville, Kwaku O. Prakah-Asante, Waqas Manzoor |
FG | 5 |
| 2023 | Evaluating Parameter-Efficient Transfer Learning Approaches on SURE Benchmark for Speech UnderstandingabstractFine-tuning is widely used as the default algorithm for transfer learning from pre-trained models. Parameter inefficiency can however arise when, during transfer learning, all the parameters of a large pre-trained model need to be updated for individual downstream tasks. As the number of parameters grows, fine-tuning is prone to overfitting and catastrophic forgetting. In addition, full fine-tuning can become prohibitively expensive when the model is used for many tasks. To mitigate this issue, parameter-efficient transfer learning algorithms, such as adapters and prefix tuning, have been proposed as a way to introduce a few trainable parameters that can be plugged into large pre-trained language models such as BERT, HuBERT. In this paper, we introduce the Speech UndeRstanding Evaluation (SURE) benchmark for parameter-efficient learning for various speech processing tasks. Additionally, we introduce a new adapter, ConvAdapter, based on 1D convolution. We show that ConvAdapter outperforms the standard adapters while showing comparable performance against prefix tuning and Low-Rank Adaptation with only 0.94% of trainable parameters. Yingting Li, Ambuj Mehrish, Rishabh Bhardwaj, Navonil Majumder, Bo Cheng 0001, Shuai Zhao 0001, Amir Zadeh 0001, Rada Mihalcea, Soujanya Poria |
ICASSP | 8 |
| 2023 | We Are in This Together: Quantifying Community Subjective Wellbeing and ResilienceabstractThe COVID-19 pandemic disrupted everyone's life across the world. In this work, we characterize the subjective wellbeing patterns of 112 cities across the United States during the pandemic prior to vaccine availability, as exhibited in subreddits corresponding to the cities. We quantify subjective wellbeing using positive and negative affect. We then measure the pandemic's impact by comparing a community's observed wellbeing with its expected wellbeing, as forecasted by time series models derived from prior to the pandemic. We show that general community traits reflected in language can be predictive of community resilience. We predict how the pandemic would impact the wellbeing of each community based on linguistic and interaction features from normal times before the pandemic. We find that communities with interaction characteristics corresponding to more closely connected users and higher engagement were less likely to be significantly impacted. Notably, we find that communities that talked more about social ties normally experienced in-person, such as friends, family, and affiliations, were actually more likely to be impacted. Additionally, we use the same features to also predict how quickly each community would recover after the initial onset of the pandemic. We similarly find that communities that talked more about family, affiliations, and identifying as part of a group had a slower recovery. MeiXing Dong, Ruixuan Sun, Laura Biester, Rada Mihalcea |
ICWSM | 4 |
| 2023 | Improving Mental Health Classifier Generalization with Pre-diagnosis DataabstractRecent work has shown that classifiers for depression detection often fail to generalize to new datasets. Most NLP models for this task are built on datasets that use textual reports of a depression diagnosis (e.g., statements on social media) to identify diagnosed users; this approach allows for collection of large-scale datasets, but leads to poor generalization to out-of-domain data. Notably, models tend to capture features that typify direct discussion of mental health rather than more subtle indications of depression symptoms. In this paper, we explore the hypothesis that building classifiers using exclusively social media posts from before a user's diagnosis will lead to less reliance on shortcuts and better generalization. We test our classifiers on a dataset that is based on an external survey rather than textual self-reports, and find that using pre-diagnosis data for training yields improved performance with many types of classifiers. Yujian Liu, Laura Biester, Rada Mihalcea |
ICWSM | 3 |
| 2023 | Query Rewriting for Effective Misinformation DiscoveryabstractAshkan Kazemi, Artem Abzaliev, Naihao Deng, Rui Hou, Scott Hale, Veronica Perez-Rosas, Rada Mihalcea. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ashkan Kazemi, Artem Abzaliev, Naihao Deng, Rui Hou 0007, Scott A. Hale, Verónica Pérez-Rosas, Rada Mihalcea |
IJCNLP (1) | 7 |
| 2023 | Reflection of Demographic Background on Word UsageabstractAbstract The availability of personal writings in electronic format provides researchers in the fields of linguistics, psychology, and computational linguistics with an unprecedented chance to study, on a large scale, the relationship between language use and the demographic background of writers, allowing us to better understand people across different demographics. In this article, we analyze the relation between language and demographics by developing cross-demographic word models to identify words with usage bias, or words that are used in significantly different ways by speakers of different demographics. Focusing on three demographic categories, namely, location, gender, and industry, we identify words with significant usage differences in each category and investigate various approaches of encoding a word’s usage, allowing us to identify language aspects that contribute to the differences. Our word models using topic-based features achieve at least 20% improvement in accuracy over the baseline for all demographic categories, even for scenarios with classification into 15 categories, illustrating the usefulness of topic-based features in identifying word usage differences. Further, we note that for location and industry, topics extracted from immediate context are the best predictors of word usages, hinting at the importance of word meaning and its grammatical function for these demographics, while for gender, topics obtained from longer contexts are better predictors for word usage. Aparna Garimella, Carmen Banea, Rada Mihalcea |
Comput. Linguistics | 3 |
| 2023 | Beneath the Tip of the Iceberg: Current Challenges and New Directions in Sentiment Analysis ResearchabstractSentiment analysis as a field has come a long way since it was first introduced as a task nearly 20 years ago. It has widespread commercial applications in various domains like marketing, risk management, market research, and politics, to name a few. Given its saturation in specific subtasks — such as sentiment polarity classification — and datasets, there is an underlying perception that this field has reached its maturity. In this article, we discuss this perception by pointing out the shortcomings and under-explored, yet key aspects of this field necessary to attaintruesentiment understanding. We analyze the significant leaps responsible for its current relevance. Further, we attempt to chart a possible course for this field that covers many overlooked and unanswered questions. Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Rada Mihalcea |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | FIBER: Fill-in-the-Blanks as a Challenging Video Understanding Evaluation FrameworkabstractSantiago Castro, Ruoyao Wang, Pingxuan Huang, Ian Stewart, Oana Ignat, Nan Liu, Jonathan Stroud, Rada Mihalcea. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Santiago Castro, Ruoyao Wang, Pingxuan Huang, Oana Ignat, Nan Liu 0010, Jonathan C. Stroud, Rada Mihalcea |
ACL (1) | 8 |
| 2022 | CICERO: A Dataset for Contextualized Commonsense Inference in DialoguesabstractThis paper addresses the problem of dialogue reasoning with contextualized commonsense inference.We curate CICERO, a dataset of dyadic conversations with five types of utterance-level reasoning-based inferences: cause, subsequent event, prerequisite, motivation, and emotional reaction.The dataset contains 53,105 of such inferences from 5,672 dialogues.We use this dataset to solve relevant generative and discriminative tasks: generation of cause and subsequent event; generation of prerequisite, motivation, and listener's emotional reaction; and selection of plausible alternatives.Our results ascertain the value of such dialogue-centric commonsense knowledge datasets.It is our hope that CI-CERO will open new research avenues into commonsense-based dialogue reasoning. Deepanway Ghosal, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
ACL (1) | 4 |
| 2022 | Knowledge Enhanced Reflection Generation for Counseling DialoguesabstractIn this paper, we study the effect of commonsense and domain knowledge while generating responses in counseling conversations using retrieval and generative methods for knowledge integration.We propose a pipeline that collects domain knowledge through web mining, and show that retrieval from both domainspecific and commonsense knowledge bases improves the quality of generated responses.We also present a model that incorporates knowledge generated by COMET using soft positional encoding and masked self-attention.We show that both retrieved and COMETgenerated knowledge improve the system's performance as measured by automatic metrics and by human evaluation.Lastly, we present a comparative study on the types of knowledge encoded by our system, showing that causal and intentional relationships benefit the generation task more than other types of commonsense relations. Verónica Pérez-Rosas, Charles Welch, Soujanya Poria, Rada Mihalcea |
ACL (1) | 5 |
| 2022 | Leveraging Similar Users for Personalized Language Modeling with Limited DataabstractCharles Welch, Chenxi Gu, Jonathan Kummerfeld, Veronica Perez-Rosas, Rada Mihalcea. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Charles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, Rada Mihalcea |
ACL (1) | 5 |
| 2022 | Towards Understanding the Relation between Gestures and LanguageabstractIn this paper, we explore the relation between gestures and language. Using a multimodal dataset, consisting of Ted talks where the language is aligned with the gestures made by the speakers, we adapt a semi-supervised multimodal model to learn gesture embeddings. We show that gestures are predictive of the native language of the speaker, and that gesture embeddings further improve language prediction result. In addition, gesture embeddings might contain some linguistic information, as we show by probing embeddings for psycholinguistic categories. Finally, we analyze the words that lead to the most expressive gestures and find that function words drive the expressiveness of gestures. Artem Abzaliev, Andrew Owens, Rada Mihalcea |
COLING | 3 |
| 2022 | In-the-Wild Video Question AnsweringabstractExisting video understanding datasets mostly focus on human interactions, with little attention being paid to the “in the wild” settings, where the videos are recorded outdoors. We propose WILDQA, a video understanding dataset of videos recorded in outside settings. In addition to video question answering (Video QA), we also introduce the new task of identifying visual support for a given question and answer (Video Evidence Selection). Through evaluations using a wide range of baseline models, we show that WILDQA poses new challenges to the vision and language research communities. The dataset is available at https://lit.eecs.umich.edu/wildqa/. Santiago Castro, Naihao Deng, Pingxuan Huang, Mihai Burzo, Rada Mihalcea |
COLING | 5 |
| 2022 | Two is Better than Many? Binary Classification as an Effective Approach to Multi-Choice Question AnsweringabstractWe propose a simple refactoring of multichoice question answering (MCQA) tasks as a series of binary classifications.The MCQA task is generally performed by scoring each (question, answer) pair normalized over all the pairs, and then selecting the answer from the pair that yield the highest score.For n answer choices, this is equivalent to an n-class classification setup where only one class (true answer) is correct.We instead show that classifying (question, true answer) as positive instances and (question, false answer) as negative instances is significantly more effective across various models and datasets.We show the efficacy of our proposed approach in different tasks -abductive reasoning, commonsense question answering, science question answering, and sentence completion.Our DeBERTa binary classification model reaches the top or close to the top performance on public leaderboards for these tasks.The source code of the proposed approach is available at https Deepanway Ghosal, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
EMNLP | 3 |
| 2022 | PAIR: Prompt-Aware margIn Ranking for Counselor Reflection Scoring in Motivational InterviewingabstractReflections are a core verbal skill used by mental health counselors to express understanding and acknowledgement of the client's experience and concerns.In this paper, we propose a system for the automatic evaluation of counselor reflections.Specifically, our system takes as input one dialog turn containing a client prompt likely leading to a reflection and a counselor response to it, and outputs a numeric score indicating the quality of the reflection made by the counselor.We compile a dataset consisting of reflections portraying different levels of reflective listening skills, and propose Prompt-Aware margIn Ranking (PAIR), a novel framework for reflection scoring that contrasts positive and negative prompt and response pairs using adhoc multi-gap and prompt-aware margin ranking losses.Through empirical evaluations and deployment of our system in a real-life educational environment, we show that our scoring model outperforms several baselines on different metrics, and can be used to provide useful feedback to counseling trainees. Do June Min, Verónica Pérez-Rosas, Kenneth Resnicow, Rada Mihalcea |
EMNLP | 4 |
| 2022 | Contact Versus Noncontact Detection of Driver's DrowsinessabstractWith an estimated number of injuries in the millions, accidents caused due to drowsy driving remain a significant source of financial costs and loss of life. Accurate detection of driver’s drowsiness could provide a clear avenue towards eliminating a great majority of the associated accidents and losses. Existing research on the subject could be defined as either contact-based or noncontact-based alertness detection. This paper utilizes a novel multimodal driver’s alertness dataset consisting of 45 subjects via seven recorded channels, including four contact-based and three noncontact-based channels, to investigate the performance of said modalities in detecting driver’s drowsiness as well as provide a novel comparison between the results of multiple contact and noncontact methods. Our results highlight the viability of noncontact methods to detect driver’s drowsiness as an implementable technology in automobiles. Salem Sharak, Kapotaksha Das, Kais Riani, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea |
ICPR | 6 |
| 2022 | Using Paraphrases to Study Properties of Contextual EmbeddingsabstractWe use paraphrases as a unique source of data to analyze contextualized embeddings, with a particular focus on BERT.Because paraphrases naturally encode consistent word and phrase semantics, they provide a unique lens for investigating properties of embeddings.Using the Paraphrase Database's alignments, we study words within paraphrases as well as phrase representations.We find that contextual embeddings effectively handle polysemous words, but give synonyms surprisingly different representations in many cases.We confirm previous findings that BERT is sensitive to word order, but find slightly different patterns than prior work in terms of the level of contextualization across BERT's layers. Laura Burdick, Jonathan K. Kummerfeld, Rada Mihalcea |
NAACL-HLT | 3 |
| 2022 | When to Make Exceptions: Exploring Language Models as Accounts of Human Moral JudgmentabstractAI systems are becoming increasingly intertwined with human life. In order to effectively collaborate with humans and ensure safety, AI systems need to be able to understand, interpret and predict human moral judgments and decisions. Human moral judgments are often guided by rules, but not always. A central challenge for AI safety is capturing the flexibility of the human moral mind — the ability to determine when a rule should be broken, especially in novel or unusual situations. In this paper, we present a novel challenge set consisting of moral exception question answering (MoralExceptQA) of cases that involve potentially permissible moral exceptions – inspired by recent moral psychology studies. Using a state-of-the-art large language model (LLM) as a basis, we propose a novel moral chain of thought (MoralCoT) prompting strategy that combines the strengths of LLMs with theories of moral reasoning developed in cognitive science to predict human moral judgments. MoralCoT outperforms seven existing LLMs by 6.2% F1, suggesting that modeling human reasoning might be necessary to capture the flexibility of the human moral mind. We also conduct a detailed error analysis to suggest directions for future work to improve AI safety using MoralExceptQA. Our data is open-sourced at https://huggingface.co/datasets/feradauto/MoralExceptQA and code at https://github.com/feradauto/MoralCoT. Zhijing Jin 0001, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, Bernhard Schölkopf |
NeurIPS | 7 |
| 2022 | How Well Do You Know Your Audience? Toward Socially-aware Question GenerationabstractWhen writing, a person may need to anticipate questions from their audience, but different social groups may ask very different types of questions.If someone is writing about a problem they want to resolve, what kind of follow-up question will a domain expert ask, and could the writer better address the expert's information needs by rewriting their original post?In this paper, we explore the task of socially-aware question generation.We collect a data set of questions and posts from social media, including background information about the question-askers' social groups.We find that different social groups, such as experts and novices, consistently ask different types of questions.We train several text-generation models that incorporate social information, and we find that a discrete social-representation model outperforms the text-only model when different social groups ask highly different questions from one another.Our work provides a framework for developing text generation models that can help writers anticipate the information expectations of highly different social groups. Rada Mihalcea |
SIGDIAL | 2 |
| 2022 | Deep Learning for Text Style Transfer: A SurveyabstractAbstract Text style transfer is an important task in natural language generation, which aims to control certain attributes in the generated text, such as politeness, emotion, humor, and many others. It has a long history in the field of natural language processing, and recently has re-gained significant attention thanks to the promising performance brought by deep neural models. In this article, we present a systematic survey of the research on neural text style transfer, spanning over 100 representative articles since the first neural text style transfer work in 2017. We discuss the task formulation, existing datasets and subtasks, evaluation, as well as the rich methodologies in the presence of parallel and non-parallel data. We also provide discussions on a variety of important topics regarding the future development of this task.1 Di Jin 0005, Zhijing Jin 0001, Zhiting Hu, Olga Vechtomova, Rada Mihalcea |
Comput. Linguistics | 5 |
| 2022 | Multimodal Deception Detection Using Real-Life Trial DataabstractHearings of witnesses and defendants play a crucial role when reaching court trial decisions. Given the high-stakes nature of trial outcomes, developing computational models that assist the decision-making process is an important research venue. In this article, we address the identification of deception in real-life trial data. We use a dataset consisting of videos collected from public court trials. We explore the use of verbal and non-verbal modalities to build a multimodal deception detection system that aims to discriminate between truthful and deceptive statements provided by defendants and witnesses. In particular, three complementary modalities (visual, acoustic and linguistic) are evaluated for the classification of deception at the subject level. The final classifier is obtained by combining the three modalities via score-level classification, achieving 83.05 percent accuracy in subject-level deceit detection. To place our results in perspective, we present a human deception detection study where we evaluate the human capability of detecting deception using different modalities and compare the results to the developed system. The results show that our system outperforms the average non-expert human capability of identifying deceit. Mehmet Umut Sen, Verónica Pérez-Rosas, Berrin A. Yanikoglu, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea |
IEEE Trans. Affect. Comput. | 6 |
| 2022 | Detection and Recognition of Driver Distraction Using Multimodal SignalsabstractDistracted driving is a leading cause of accidents worldwide. The tasks of distraction detection and recognition have been traditionally addressed as computer vision problems. However, distracted behaviors are not always expressed in a visually observable way. In this work, we introduce a novel multimodal dataset of distracted driver behaviors, consisting of data collected using twelve information channels coming from visual, acoustic, near-infrared, thermal, physiological and linguistic modalities. The data were collected from 45 subjects while being exposed to four different distractions (three cognitive and one physical). For the purposes of this paper, we performed experiments with visual, physiological, and thermal information to explore potential of multimodal modeling for distraction recognition. In addition, we analyze the value of different modalities by identifying specific visual, physiological, and thermal groups of features that contribute the most to distraction characterization. Our results highlight the advantage of multimodal representations and reveal valuable insights for the role played by the three modalities on identifying different types of driving distractions. Kapotaksha Das, Michalis Papakostas, Kais Riani, Andrew Brian Gasiorowski, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea |
ACM Trans. Interact. Intell. Syst. | 7 |
| 2022 | When Did It Happen? Duration-informed Temporal Localization of Narrated Actions in VlogsabstractWe consider the task of temporal human action localization in lifestyle vlogs. We introduce a novel dataset consisting of manual annotations of temporal localization for 13,000 narrated actions in 1,200 video clips. We present an extensive analysis of this data, which allows us to better understand how the language and visual modalities interact throughout the videos. We propose a simple yet effective method to localize the narrated actions based on their expected duration. Through several experiments and analyses, we show that our method brings complementary information with respect to previous methods, and leads to improvements over previous work for the task of temporal action localization. Oana Ignat, Santiago Castro, Jiajun Bao, Dandan Shan, Rada Mihalcea |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2021 | Humor Knowledge Enriched Transformer for Understanding Multimodal HumorabstractRecognizing humor from a video utterance requires understanding the verbal and non-verbal components as well as incorporating the appropriate context and external knowledge. In this paper, we propose Humor Knowledge enriched Transformer (HKT) that can capture the gist of a multimodal humorous expression by integrating the preceding context and external knowledge. We incorporate humor centric external knowledge into the model by capturing the ambiguity and sentiment present in the language. We encode all the language, acoustic, vision, and humor centric features separately using Transformer based encoders, followed by a cross attention layer to exchange information among them. Our model achieves 77.36% and 79.41% accuracy in humorous punchline detection on UR-FUNNY and MUStaRD datasets -- achieving a new state-of-the-art on both datasets with the margin of 4.93% and 2.94% respectively. Furthermore, we demonstrate that our model can capture interpretable, humor-inducing patterns from all modalities. Md. Kamrul Hasan 0003, Sangwu Lee, Wasifur Rahman, Amir Zadeh 0001, Rada Mihalcea, Louis-Philippe Morency, Mohammed E. Hoque 0001 |
AAAI | 5 |
| 2021 | Towards Classifying Human Circadian Rhythm Using Multiple ModalitiesabstractAutonomous vehicles represent one of the most active technologies currently being developed, with research areas addressing, among others, the modeling of the states and behavioral elements of the occupants. This paper contributes to this line of research by studying the circadian rhythm of individuals using a novel multimodal dataset of 36 subjects consisting of five information channels. These channels include visual, thermal, physiological, linguistic, and background data. Moreover, we propose a framework to explore whether the circadian rhythm can be modeled without continuous monitoring and investigate the hypothesis that multimodal features have a greater propensity for improved performance using data points specific to certain times during the day. Our analysis shows that multimodal fusion can lead to an accuracy of up to 77% on identifying energized and enervated states of the participants. Our findings highlight the validity of our hypothesis and present a novel approach for future research. Kais Riani, Salem Sharak, Kapotaksha Das, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea, John Elson, Clay Maranville, Kwaku O. Prakah-Asante, Waqas Manzoor |
ACII | 6 |
| 2021 | Hitting your MARQ: Multimodal ARgument Quality Assessment in Long Debate VideoabstractThe combination of gestures, intonations, and textual content plays a key role in argument delivery.However, the current literature mostly considers textual content while assessing the quality of an argument, and is limited to datasets containing short sequences (18-48 words).In this paper, we study argument quality assessment in a multimodal context, and experiment on DBATES, a publicly available dataset of long debate videos.First, we propose a set of interpretable debate-centric features such as clarity, content variation, body movement cues, and pauses, inspired by theories of argumentation quality.Second, we design the Multimodal ARgument Quality assessor (MARQ) -a hierarchical neural network model that summarizes the multimodal signals on long sequences and enriches the multimodal embedding with debate-centric features.Our proposed MARQ model achieves an accuracy of 81.91% on the argument quality prediction task and outperforms established baseline models with an error rate reduction of 22.7%.Through ablation studies, we demonstrate the importance of multimodal cues in modeling argument quality. Md. Kamrul Hasan 0003, James Spann, Masum Hasan, Md. Saiful Islam 0013, Kurtis Haut, Rada Mihalcea, Mohammed E. Hoque 0001 |
EMNLP (1) | 6 |
| 2021 | Analyzing the Surprising Variability in Word Embedding Stability Across LanguagesabstractWord embeddings are powerful representations that form the foundation of many natural language processing architectures, both in English and in other languages.To gain further insight into word embeddings, we explore their stability (e.g., overlap between the nearest neighbors of a word in different embedding spaces) in diverse languages.We discuss linguistic properties that are related to stability, drawing out insights about correlations with affixing, language gender systems, and other features.This has implications for embedding use, particularly in research that uses them to study language trends. Laura Burdick, Jonathan K. Kummerfeld, Rada Mihalcea |
EMNLP (1) | 3 |
| 2021 | STaCK: Sentence Ordering with Temporal Commonsense KnowledgeabstractSentence order prediction is the task of finding the correct order of sentences in a randomly ordered document.Correctly ordering the sentences requires an understanding of coherence with respect to the chronological sequence of events described in the text.Documentlevel contextual understanding and commonsense knowledge centered around these events are often essential in uncovering this coherence and predicting the exact chronological order.In this paper, we introduce STaCK -a framework based on graph neural networks and temporal commonsense knowledge to model global information and predict the relative order of sentences.Our graph network accumulates temporal evidence using knowledge of 'past' and 'future' and formulates sentence ordering as a constrained edge classification problem.We report results on five different datasets, and empirically show that the proposed method is naturally suitable for order prediction, thus demonstrating the role of temporal commonsense knowledge.The implementation of this work is available at: https://github.com/declare-lab/ sentence-ordering.Jennifer has her final exam tomorrow.She got so stressed, she pulled an Deepanway Ghosal, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
EMNLP (1) | 3 |
| 2021 | WhyAct: Identifying Action Reasons in Lifestyle VlogsabstractWe aim to automatically identify human action reasons in online videos.We focus on the widespread genre of lifestyle vlogs, in which people perform actions while verbally describing them.We introduce and make publicly available the WHYACT dataset, consisting of 1,077 visual actions manually annotated with their reasons.We describe a multimodal model that leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video. Oana Ignat, Santiago Castro, Hanwen Miao, Weiji Li, Rada Mihalcea |
EMNLP (1) | 5 |
| 2021 | Understanding Driving Distractions: A Multimodal Analysis on Distraction CharacterizationabstractDistracted driving is a leading cause of accidents worldwide. The tasks of distraction detection and recognition have been traditionally addressed as computer vision problems. However, distracted behaviors are not always expressed in a visually observable way. In this work, we introduce a novel multimodal dataset of distracted driver behaviors, consisting of data collected using twelve information channels coming from visual, acoustic, near-infrared, thermal, physiological and linguistic modalities. The data were collected from 45 subjects while being exposed to four different distractions (three cognitive and one physical). For the purposes of this paper, we experiment with visual and physiological information and explore the potential of multimodal modeling for distraction recognition. In addition, we analyze the value of different modalities by identifying specific visual and physiological groups of features that contribute the most to distraction characterization. Our results highlight the advantage of multimodal representations and reveal valuable insights for the role played by the two modalities on identifying different types of driving distractions. Michalis Papakostas, Kais Riani, Andrew Brian Gasiorowski, Mohamed Abouelenien, Rada Mihalcea, Mihai Burzo |
IUI | 6 |
| 2021 | MUSER: MUltimodal Stress detection using Emotion Recognition as an Auxiliary TaskabstractYiqun Yao, Michalis Papakostas, Mihai Burzo, Mohamed Abouelenien, Rada Mihalcea. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yiqun Yao, Michalis Papakostas, Mihai Burzo, Mohamed Abouelenien, Rada Mihalcea |
NAACL-HLT | 5 |
| 2021 | CIDER: Commonsense Inference for Dialogue Explanation and ReasoningabstractWell, I missed several buses.How on earth can you miss several buses?I, ah ..., I have got late.But there's a bus every ten minutes, and you are over 1 hour late.Have you got it now? Deepanway Ghosal, Pengfei Hong, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
SIGDIAL | 5 |
| 2020 | KinGDOM: Knowledge-Guided DOMain Adaptation for Sentiment AnalysisabstractCross-domain sentiment analysis has received significant attention in recent years, prompted by the need to combat the domain gap between different applications that make use of sentiment analysis.In this paper, we take a novel perspective on this task by exploring the role of external commonsense knowledge.We introduce a new framework, KinGDOM, which utilizes the ConceptNet knowledge graph to enrich the semantics of a document by providing both domain-specific and domain-general background concepts.These concepts are learned by training a graph convolutional autoencoder that leverages inter-domain concepts in a domain-invariant manner.Conditioning a popular domain-adversarial baseline method with these learned concepts helps improve its performance over state-of-the-art approaches, demonstrating the efficacy of our proposed framework. Deepanway Ghosal, Devamanyu Hazarika, Abhinaba Roy, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
ACL | 5 |
| 2020 | "Judge me by my size (noun), do you?" YodaLib: A Demographic-Aware Humor Generation FrameworkabstractThe subjective nature of humor makes computerized humor generation a challenging task.We propose an automatic humor generation framework for filling the blanks in Mad Libs R stories, while accounting for the demographic backgrounds of the desired audience.We collect a dataset consisting of such stories, which are filled in and judged by carefully selected workers on Amazon Mechanical Turk.We build upon the BERT platform to predict location-biased word fillings in incomplete sentences, and we fine-tune BERT to classify location-specific humor in a sentence.We leverage these components to produce YODALIB, a fully-automated Mad Libs style humor generation framework, which selects and ranks appropriate candidate words and sentences in order to generate a coherent and funny story tailored to certain demographics.Our experimental results indicate that YODALIB outperforms a previous semi-automated approach proposed for this task, while also surpassing human annotators in both qualitative and quantitative analyses. Aparna Garimella, Carmen Banea, Nabil Hossain, Rada Mihalcea |
COLING | 4 |
| 2020 | Biased TextRank: Unsupervised Graph-Based Content ExtractionabstractWe introduce Biased TextRank, a graph-based content extraction method inspired by the popular TextRank algorithm that ranks text spans according to their importance for language processing tasks and according to their relevance to an input "focus."Biased TextRank enables focused content extraction for text by modifying the random restarts in the execution of TextRank.The random restart probabilities are assigned based on the relevance of the graph nodes to the focus of the task.We present two applications of Biased TextRank: focused summarization and explanation extraction, and show that our algorithm leads to improved performance on two different datasets by significant ROUGE-N score margins.Much like its predecessor, Biased TextRank is unsupervised, easy to implement and orders of magnitude faster and lighter than current state-ofthe-art Natural Language Processing methods for similar tasks. Ashkan Kazemi, Verónica Pérez-Rosas, Rada Mihalcea |
COLING | 3 |
| 2020 | Exploring the Value of Personalized Word EmbeddingsabstractIn this paper, we introduce personalized word embeddings, and examine their value for language modeling.We compare the performance of our proposed prediction model when using personalized versus generic word representations, and study how these representations can be leveraged for improved performance.We provide insight into what types of words can be more accurately predicted when building personalized models.Our results show that a subset of words belonging to specific psycholinguistic categories tend to vary more in their representations across users and that combining generic and personalized word embeddings yields the best performance, with a 4.7% relative reduction in perplexity.Additionally, we show that a language model using personalized word embeddings can be effectively used for authorship attribution. Charles Welch, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, Rada Mihalcea |
COLING | 4 |
| 2020 | MIME: MIMicking Emotions for Empathetic Response GenerationabstractNavonil Majumder, Pengfei Hong, Shanshan Peng, Jiankun Lu, Deepanway Ghosal, Alexander Gelbukh, Rada Mihalcea, Soujanya Poria. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Navonil Majumder, Pengfei Hong, Shanshan Peng, Jiankun Lu, Deepanway Ghosal, Alexander F. Gelbukh, Rada Mihalcea, Soujanya Poria |
EMNLP (1) | 7 |
| 2020 | Compositional Demographic Word EmbeddingsabstractWord embeddings are usually derived from corpora containing text from many individuals, thus leading to general purpose representations rather than individually personalized representations.While personalized embeddings can be useful to improve language model performance and other language processing tasks, they can only be computed for people with a large amount of longitudinal data, which is not the case for new users.We propose a new form of personalized word embeddings that use demographic-specific word representations derived compositionally from full or partial demographic information for a user (i.e., gender, age, location, religion).We show that the resulting demographic-aware word representations outperform generic word representations on two tasks for English: language modeling and word associations.We further explore the trade-off between the number of available attributes and their relative effectiveness and discuss the ethical implications of using them. Charles Welch, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, Rada Mihalcea |
EMNLP (1) | 4 |
| 2020 | Improving Low Compute Language Modeling with In-Domain Embedding InitialisationabstractMany NLP applications, such as biomedical data and technical support, have 10-100 million tokens of in-domain data and limited computational resources for learning from it.How should we train a language model in this scenario?Most language modeling research considers either a small dataset with a closed vocabulary (like the standard 1 million token Penn Treebank), or the whole web with bytepair encoding.We show that for our target setting in English, initialising and freezing input embeddings using in-domain data can improve language model performance by providing a useful representation of rare words, and this pattern holds across several different domains.In the process, we show that the standard convention of tying input and output embeddings does not improve perplexity when initializing with embeddings trained on in-domain data. Charles Welch, Rada Mihalcea, Jonathan K. Kummerfeld |
EMNLP (1) | 2 |
| 2020 | LifeQA: A Real-life Dataset for Video Question AnsweringabstractWe introduce LifeQA, a benchmark dataset for video question answering that focuses on day-to-day real-life situations. Current video question answering datasets consist of movies and TV shows. However, it is well-known that these visual domains are not representative of our day-to-day lives. Movies and TV shows, for example, benefit from professional camera movements, clean editing, crisp audio recordings, and scripted dialog between professional actors. While these domains provide a large amount of data for training models, their properties make them unsuitable for testing real-life question answering systems. Our dataset, by contrast, consists of video clips that represent only real-life scenarios. We collect 275 such video clips and over 2.3k multiple-choice questions. In this paper, we analyze the challenging but realistic aspects of LifeQA, and we apply several state-of-the-art video question answering models to provide benchmarks for future research. The full dataset is publicly available at https://lit.eecs.umich.edu/lifeqa/. Santiago Castro, Mahmoud Azab, Jonathan C. Stroud, Cristina Noujaim, Ruoyao Wang, Jia Deng 0001, Rada Mihalcea |
LREC | 7 |
| 2020 | MuSE: a Multimodal Dataset of Stressed EmotionabstractEndowing automated agents with the ability to provide support, entertainment and interaction with human beings requires sensing of the users’ affective state. These affective states are impacted by a combination of emotion inducers, current psychological state, and various conversational factors. Although emotion classification in both singular and dyadic settings is an established area, the effects of these additional factors on the production and perception of emotion is understudied. This paper presents a new dataset, Multimodal Stressed Emotion (MuSE), to study the multimodal interplay between the presence of stress and expressions of affect. We describe the data collection protocol, the possible areas of use, and the annotations for the emotional content of the recordings. The paper also presents several baselines to measure the performance of multimodal features for emotion and stress classification. Mimansa Jaiswal, Cristian-Paul Bara, Yuanhang Luo, Mihai Burzo, Rada Mihalcea, Emily Mower Provost |
LREC | 5 |
| 2020 | Small Town or Metropolis? Analyzing the Relationship between Population Size and LanguageabstractThe variance in language used by different cultures has been a topic of study for researchers in linguistics and psychology, but often times, language is compared across multiple countries in order to show a difference in culture. As a geographically large country that is diverse in population in terms of the background and experiences of its citizens, the U.S. also contains cultural differences within its own borders. Using a set of over 2 million posts from distinct Twitter users around the country dating back as far as 2014, we ask the following question: is there a difference in how Americans express themselves online depending on whether they reside in an urban or rural area? We categorize Twitter users as either urban or rural and identify ideas and language that are more commonly expressed in tweets written by one population over the other. We take this further by analyzing how the language from specific cities of the U.S. compares to the language of other cities and by training predictive models to predict whether a user is from an urban or rural area. We publicly release the tweet and user IDs that can be used to reconstruct the dataset for future studies in this direction. Amy Rechkemmer, Steven R. Wilson 0001, Rada Mihalcea |
LREC | 3 |
| 2020 | Inferring Social Media Users' Mental Health Status from Multimodal InformationabstractWorldwide, an increasing number of people are suffering from mental health disorders such as depression and anxiety. In the United States alone, one in every four adults suffers from a mental health condition, which makes mental health a pressing concern. In this paper, we explore the use of multimodal cues present in social media posts to predict users’ mental health status. Specifically, we focus on identifying social media activity that either indicates a mental health condition or its onset. We collect posts from Flickr and apply a multimodal approach that consists of jointly analyzing language, visual, and metadata cues and their relation to mental health. We conduct several classification experiments aiming to discriminate between (1) healthy users and users affected by a mental health illness; and (2) healthy users and users prone to mental illness. Our experimental results indicate that using multiple modalities can improve the performance of this classification task as compared to the use of one modality at a time, and can provide important cues into a user’s mental status. Zhentao Xu, Verónica Pérez-Rosas, Rada Mihalcea |
LREC | 3 |
| 2020 | Counseling-Style Reflection Generation Using Generative Pretrained Transformers with Augmented ContextabstractIn this paper, we introduce a counseling dialogue system that provides real-time assistance to counseling trainees.The system generates sample counselors' reflections -i.e., responses that reflect back on what the client has said given the dialogue history.We build our model upon the recent generative pretrained transformer architecture and leverage context augmentation techniques inspired by traditional strategies used during counselor training to further enhance its performance.We show that the system incorporating these strategies outperforms the baseline models on the reflection generation task on multiple metrics.To confirm our findings, we present a human evaluation study that shows that the output of the enhanced system obtains higher ratings and is on par with human responses in terms of stylistic and grammatical correctness, as well as context-awareness. Charles Welch, Rada Mihalcea, Verónica Pérez-Rosas |
SIGdial | 3 |
| 2020 | Extending sparse text with induced domain-specific lexicons and embeddings: A case study on predicting donations
MeiXing Dong, Rada Mihalcea, Dragomir R. Radev |
Comput. Speech Lang. | 2 |
| 2019 | DialogueRNN: An Attentive RNN for Emotion Detection in ConversationsabstractEmotion detection in conversations is a necessary step for a number of applications, including opinion mining over chat history, social media threads, debates, argumentation mining, understanding consumer feedback in live conversations, and so on. Currently systems do not treat the parties in the conversation individually by adapting to the speaker of each utterance. In this paper, we describe a new method based on recurrent neural networks that keeps track of the individual party states throughout the conversation and uses this information for emotion classification. Our model outperforms the state-of-the-art by a significant margin on two different datasets. Navonil Majumder, Soujanya Poria, Devamanyu Hazarika, Rada Mihalcea, Alexander F. Gelbukh, Erik Cambria |
AAAI | 4 |
| 2019 | Towards Multimodal Sarcasm Detection (An _Obviously_ Perfect Paper)abstractSantiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, Soujanya Poria. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, Soujanya Poria |
ACL (1) | 5 |
| 2019 | Women's Syntactic Resilience and Men's Grammatical Luck: Gender-Bias in Part-of-Speech Tagging and Dependency ParsingabstractSeveral linguistic studies have shown the prevalence of various lexical and grammatical patterns in texts authored by a person of a particular gender, but models for part-of-speech tagging and dependency parsing have still not adapted to account for these differences.To address this, we annotate the Wall Street Journal part of the Penn Treebank with the gender information of the articles' authors, and build taggers and parsers trained on this data that show performance differences in text written by men and women.Further analyses reveal numerous part-of-speech tags and syntactic relations whose prediction performances benefit from the prevalence of a specific gender in the training data.The results underscore the importance of accounting for gendered differences in syntactic tasks, and outline future venues for developing more accurate taggers and parsers.We release our data to the research community. Aparna Garimella, Carmen Banea, Eduard H. Hovy, Rada Mihalcea |
ACL (1) | 4 |
| 2019 | Identifying Visible Actions in Lifestyle VlogsabstractWe consider the task of identifying human actions visible in online videos. We focus on the widely spread genre of lifestyle vlogs, which consist of videos of people performing actions while verbally describing them. Our goal is to identify if actions mentioned in the speech description of a video are visually present. We construct a dataset with crowdsourced manual annotations of visible actions, and introduce a multimodal algorithm that leverages information derived from visual and linguistic clues to automatically infer which actions are visible in a video. We demonstrate that our multimodal algorithm outperforms algorithms based only on one modality at a time. Oana Ignat, Laura Burdick, Jia Deng 0001, Rada Mihalcea |
ACL (1) | 4 |
| 2019 | What Makes a Good Counselor? Learning to Distinguish between High-quality and Low-quality Counseling ConversationsabstractThe quality of a counseling intervention relies highly on the active collaboration between clients and counselors.In this paper, we explore several linguistic aspects of the collaboration process occurring during counseling conversations.Specifically, we address the differences between high-quality and low-quality counseling.Our approach examines participants' turn-by-turn interaction, their linguistic alignment, the sentiment expressed by speakers during the conversation, as well as the different topics being discussed.Our results suggest important language differences in lowand high-quality counseling, which we further use to derive linguistic features able to capture the differences between the two groups.These features are then used to build automatic classifiers that can predict counseling quality with accuracies of up to 88%. Verónica Pérez-Rosas, Kenneth Resnicow, Rada Mihalcea |
ACL (1) | 4 |
| 2019 | MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in ConversationsabstractEmotion recognition in conversations (ERC) is a challenging task that has recently gained popularity due to its potential applications.Until now, however, there has been no largescale multimodal multi-party emotional conversational database containing more than two speakers per dialogue.To address this gap, we propose the Multimodal EmotionLines Dataset (MELD), an extension and enhancement of EmotionLines.MELD contains about 13,000 utterances from 1,433 dialogues from the TV-series Friends.Each utterance is annotated with emotion and sentiment labels, and encompasses audio, visual, and textual modalities.We propose several strong multimodal baselines and show the importance of contextual and multimodal information for emotion recognition in conversations.The full dataset is available for use at http:// affective-meld.github.io. Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, Rada Mihalcea |
ACL (1) | 6 |
| 2019 | Predicting Human Activities from User-Generated ContentabstractThe activities we do are linked to our interests, personality, political preferences, and decisions we make about the future.In this paper, we explore the task of predicting human activities from user-generated content.We collect a dataset containing instances of social media users writing about a range of everyday activities.We then use a state-of-the-art sentence embedding framework tailored to recognize the semantics of human activities and perform an automatic clustering of these activities.We train a neural network model to make predictions about which clusters contain activities that were performed by a given user based on the text of their previous posts and selfdescription.Additionally, we explore the degree to which incorporating inferred user traits into our model helps with this prediction task. Steven R. Wilson 0001, Rada Mihalcea |
ACL (1) | 2 |
| 2019 | Look Who's Talking: Inferring Speaker Attributes from Personal Longitudinal Dialog
Charles Welch, Verónica Pérez-Rosas, Jonathan K. Kummerfeld, Rada Mihalcea |
CICLing (2) | 4 |
| 2019 | Representing Movie Characters in DialoguesabstractWe introduce a new embedding model to represent movie characters and their interactions in a dialogue by encoding in the same representation the language used by these characters as well as information about the other participants in the dialogue.We evaluate the performance of these new character embeddings on two tasks: (1) character relatedness, using a dataset we introduce consisting of a dense character interaction matrix for 4,761 unique character pairs over 22 hours of dialogue from eighteen movies; and (2) character relation classification, for fine-and coarse-grained relations, as well as sentiment relations.Our experiments show that our model significantly outperforms the traditional Word2Vec continuous bag-of-words and skip-gram models, demonstrating the effectiveness of the character embeddings we introduce.We further show how these embeddings can be used in conjunction with a visual question answering system to improve over previous results. Mahmoud Azab, Noriyuki Kojima, Jia Deng 0001, Rada Mihalcea |
CoNLL | 4 |
| 2019 | Towards Extracting Medical Family History from Natural Language Interactions: A New Dataset and BaselinesabstractMahmoud Azab, Stephane Dadian, Vivi Nastase, Larry An, Rada Mihalcea. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mahmoud Azab, Stephane Dadian, Vivi Nastase, Lawrence C. An, Rada Mihalcea |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Muse-ing on the Impact of Utterance Ordering on Crowdsourced Emotion AnnotationsabstractEmotion recognition algorithms rely on data annotated with high quality labels. However, emotion expression and perception are inherently subjective. There is generally not a single annotation that can be unambiguously declared "correct." As a result, annotations are colored by the manner in which they were collected. In this paper, we conduct crowdsourcing experiments to investigate this impact on both the annotations themselves and on the performance of these algorithms. We focus on one critical question: the effect of context. We present a new emotion dataset, Multimodal Stressed Emotion (MuSE), and annotate the dataset using two conditions: randomized, in which annotators are presented with clips in random order, and contextualized, in which annotators are presented with clips in order. We find that contextual labeling schemes result in annotations that are more similar to a speaker's own self-reported labels and that labels generated from randomized schemes are most easily predictable by automated systems. Mimansa Jaiswal, Zakaria Aldeneh, Cristian-Paul Bara, Yuanhang Luo, Mihai Burzo, Rada Mihalcea, Emily Mower Provost |
ICASSP | 6 |
| 2019 | Towards Automatic Detection of Misinformation in Online Medical Videos
Rui Hou 0007, Verónica Pérez-Rosas, Stacy L. Loeb, Rada Mihalcea |
ICMI | 4 |
| 2018 | CASCADE: Contextual Sarcasm Detection in Online Discussion ForumsabstractThe literature in automated sarcasm detection has mainly focused on lexical-, syntactic- and semantic-level analysis of text. However, a sarcastic sentence can be expressed with contextual presumptions, background and commonsense knowledge. In this paper, we propose a ContextuAl SarCasm DEtector (CASCADE), which adopts a hybrid approach of both content- and context-driven modeling for sarcasm detection in online social media discussions. For the latter, CASCADE aims at extracting contextual information from the discourse of a discussion thread. Also, since the sarcastic nature and form of expression can vary from person to person, CASCADE utilizes user embeddings that encode stylometric and personality features of users. When used along with content-based feature extractors such as convolutional neural networks, we see a significant boost in the classification performance on a large Reddit corpus. Devamanyu Hazarika, Soujanya Poria, Sruthi Gorantla, Erik Cambria, Roger Zimmermann, Rada Mihalcea |
COLING | 6 |
| 2018 | Automatic Detection of Fake NewsabstractThe proliferation of misleading information in everyday access media outlets such as social media feeds, news blogs, and online newspapers have made it challenging to identify trustworthy news sources, thus increasing the need for computational tools able to provide insights into the reliability of online content. In this paper, we focus on the automatic identification of fake content in online news. Our contribution is twofold. First, we introduce two novel datasets for the task of fake news detection, covering seven different news domains. We describe the collection, annotation, and validation process in detail and present several exploratory analyses on the identification of linguistic differences in fake and legitimate news content. Second, we conduct a set of learning experiments to build accurate fake news detectors, and show that we can achieve accuracies of up to 76%. In addition, we provide comparative analyses of the automatic and manual identification of fake news. Verónica Pérez-Rosas, Bennett Kleinberg, Alexandra Lefevre, Rada Mihalcea |
COLING | 4 |
| 2018 | ICON: Interactive Conversational Memory Network for Multimodal Emotion DetectionabstractEmotion recognition in conversations is crucial for building empathetic machines.Current work in this domain do not explicitly consider the inter-personal influences that thrive in the emotional dynamics of dialogues.To this end, we propose Interactive COnversational memory Network (ICON), a multimodal emotion detection framework that extracts multimodal features from conversational videos and hierarchically models the selfand interspeaker emotional influences into global memories.Such memories generate contextual summaries which aid in predicting the emotional orientation of utterance-videos.Our model outperforms state-of-the-art networks on multiple classification and regression tasks in two benchmark datasets. Devamanyu Hazarika, Soujanya Poria, Rada Mihalcea, Erik Cambria, Roger Zimmermann |
EMNLP | 3 |
| 2018 | Analyzing the Quality of Counseling Conversations: the Tell-Tale Signs of High-quality Counseling
Verónica Pérez-Rosas, Xuetong Sun, Christy Li, Kenneth Resnicow, Rada Mihalcea |
LREC | 6 |
| 2018 | World Knowledge for Abstract Meaning Representation Parsing
Charles Welch, Jonathan K. Kummerfeld, Song Feng 0002, Rada Mihalcea |
LREC | 4 |
| 2018 | Speaker Naming in MoviesabstractMahmoud Azab, Mingzhe Wang, Max Smith, Noriyuki Kojima, Jia Deng, Rada Mihalcea. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Mahmoud Azab, Max Olan Smith, Noriyuki Kojima, Jia Deng 0001, Rada Mihalcea |
NAACL-HLT | 6 |
| 2018 | Factors Influencing the Surprising Instability of Word EmbeddingsabstractLaura Wendlandt, Jonathan K. Kummerfeld, Rada Mihalcea. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Laura Burdick, Jonathan K. Kummerfeld, Rada Mihalcea |
NAACL-HLT | 3 |
| 2018 | Possession identification in textabstractAbstract Just as industrialization matured from mass production to customization and personalization, so has the Web migrated from generic content to public disclosures of one’s most intimately held thoughts, opinions, and beliefs. This relatively new type of data is able to represent finer and more narrowly defined demographic slices. If until now researchers have primarily focused on leveraging personalized content to identify latent information such as gender, nationality, location, or age, this article seeks to establish a structured way of extracting possessions, or items that people own or are entitled to, as a way to ultimately provide insights into people’s behaviors and characteristics. We introduce the new task of ‘possession identification in text’, and release a novel dataset where possessions are marked at different confidence levels. We present experiments and results obtained when seeking to automatically identify and extract possessions from the text. Carmen Banea, Rada Mihalcea |
Nat. Lang. Eng. | 2 |
| 2017 | Grounded emotionsabstractEmotions are grounded in contextual experience. While natural language processing tools typically look at textual content to find clues pertaining to an author's emotional state, factors occurring throughout the day, such as weather or news exposure, may prime one toward a particular emotional response. In this paper, we explore five types of external factors and through extensive analyses show their impact and correlation with a user's emotional state. Ultimately, we show that when combining all extrinsic features, we are able to predict with an accuracy of 67% the emotional state of a user. Vicki Liu, Carmen Banea, Rada Mihalcea |
ACII | 3 |
| 2017 | Understanding and Predicting Empathic Behavior in Counseling TherapyabstractVerónica Pérez-Rosas, Rada Mihalcea, Kenneth Resnicow, Satinder Singh, Lawrence An. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Verónica Pérez-Rosas, Rada Mihalcea, Kenneth Resnicow, Satinder Singh 0001, Lawrence C. An |
ACL (1) | 2 |
| 2017 | Deception Detection: When Computers Become Better than HumansabstractWhether we like it or not, deception happens every day and everywhere: thousands of trials taking place daily around the world; little white lies: "I'm busy that day!" even if your calendar is blank; news "with a twist" (a.k.a. fake news) meant to attract the readers attraction, and get some advertisement clicks on the side; portrayed identities, on dating sites and elsewhere. Can a computer automatically detect deception in written accounts or in video recordings? In this talk, I will describe our work in building linguistic and multimodal algorithms for deception detection, targeting deceptive statements, trial videos, fake news, identity deceptions, and also going after deception in multiple cultures. I will also show how these algorithms can provide insights into what makes a good lie - and thus teach us how we can spot a liar. As it turns out, computers can be trained to identify lies in many different contexts, and they can do it much better than humans do! Rada Mihalcea |
CIKM | 1 |
| 2017 | Predicting Counselor Behaviors in Motivational Interviewing EncountersabstractVerónica Pérez-Rosas, Rada Mihalcea, Kenneth Resnicow, Satinder Singh, Lawrence An, Kathy J. Goggin, Delwyn Catley. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Verónica Pérez-Rosas, Rada Mihalcea, Kenneth Resnicow, Satinder Singh 0001, Lawrence C. An, Kathy J. Goggin, Delwyn Catley |
EACL (1) | 2 |
| 2017 | Demographic-aware word associationsabstractVariations of word associations across different groups of people can provide insights into people's psychologies and their world views.To capture these variations, we introduce the task of demographicaware word associations.We build a new gold standard dataset consisting of word association responses for approximately 300 stimulus words, collected from more than 800 respondents of different gender (male/female) and from different locations (India/United States), and show that there are significant variations in the word associations made by these groups.We also introduce a new demographic-aware word association model based on a neural net skip-gram architecture, and show how computational methods for measuring word associations that specifically account for writer demographics can outperform generic methods that are agnostic to such information. Aparna Garimella, Carmen Banea, Rada Mihalcea |
EMNLP | 3 |
| 2017 | Multimodal gender detectionabstractAutomatic gender classification is receiving increasing attention in the computer interaction community as the need for personalized, reliable, and ethical systems arises. To date, most gender classification systems have been evaluated on textual and audiovisual sources. This work explores the possibility of enhancing such systems with physiological cues obtained from thermography and physiological sensor readings. Using a multimodal dataset consisting of audiovisual, thermal, and physiological recordings of males and females, we extract features from five different modalities, namely acoustic, linguistic, visual, thermal, and physiological. We then conduct a set of experiments where we explore the gender prediction task using single and combined modalities. Experimental results suggest that physiological and thermal information can be used to recognize gender at reasonable accuracy levels, which are comparable to the accuracy of current gender prediction systems. Furthermore, we show that the use of non-contact physiological measurements, such as thermography readings, can enhance current systems that are based on audio or visual input. This can be particularly useful for scenarios where non-contact approaches are preferred, i.e., when data is captured under noisy audiovisual conditions or when video or speech data are not available due to ethical considerations. Mohamed Abouelenien, Verónica Pérez-Rosas, Rada Mihalcea, Mihai Burzo |
ICMI | 3 |
| 2017 | Identifying Usage Expression Sentences in Consumer Product ReviewsabstractIn this paper we introduce the problem of identifying usage expression sentences in a consumer product review. We create a human-annotated gold standard dataset of 565 reviews spanning five distinct product categories. Our dataset consists of more than 3,000 annotated sentences. We further introduce a classification system to label sentences according to whether or not they describe some “usage”. The system combines lexical, syntactic, and semantic features in a product-agnostic fashion to yield good classification performance. We show the effectiveness of our approach using importance ranking of features, error analysis, and cross-product classification experiments. Shibamouli Lahiri, V. G. Vinod Vydiswaran, Rada Mihalcea |
IJCNLP(1) | 3 |
| 2017 | Identity Deception DetectionabstractThis paper addresses the task of detecting identity deception in language. Using a novel identity deception dataset, consisting of real and portrayed identities from 600 individuals, we show that we can build accurate identity detectors targeting both age and gender, with accuracies of up to 88. We also perform an analysis of the linguistic patterns used in identity deception, which lead to interesting insights into identity portrayers. Verónica Pérez-Rosas, Quincy Davenport, Anna Mengdan Dai, Mohamed Abouelenien, Rada Mihalcea |
IJCNLP(1) | 5 |
| 2017 | Measuring Semantic Relations between Human ActivitiesabstractThe things people do in their daily lives can provide valuable insights into their personality, values, and interests. Unstructured text data on social media platforms are rich in behavioral content, and automated systems can be deployed to learn about human activity on a broad scale if these systems are able to reason about the content of interest. In order to aid in the evaluation of such systems, we introduce a new phrase-level semantic textual similarity dataset comprised of human activity phrases, providing a testbed for automated systems that analyze relationships between phrasal descriptions of people’s actions. Our set of 1,000 pairs of activities is annotated by human judges across four relational dimensions including similarity, relatedness, motivational alignment, and perceived actor congruence. We evaluate a set of strong baselines for the task of generating scores that correlate highly with human ratings, and we introduce several new approaches to the phrase-level similarity task in the domain of human activities. Steven R. Wilson 0001, Rada Mihalcea |
IJCNLP(1) | 2 |
| 2017 | Keyword extraction from emailsabstractAbstract Emails constitute an important genre of online communication. Many of us are often faced with the daunting task of sifting through increasingly large amounts of emails on a daily basis. Keywords extracted from emails can help us combat such information overload by allowing a systematic exploration of the topics contained in emails. Existing literature on keyword extraction has not covered the email genre, and no human-annotated gold standard datasets are currently available. In this paper, we introduce a new dataset for keyword extraction from emails, and evaluate supervised and unsupervised methods for keyword extraction from emails. The results obtained with our supervised keyword extraction system (38.99% F-score) improve over the results obtained with the best performing systems participating in theSemEval2010 keyword extraction task. Shibamouli Lahiri, Rada Mihalcea, Po-Hsiang Lai |
Nat. Lang. Eng. | 2 |
| 2017 | Coarse-Grained +/-Effect Word Sense Disambiguation for Implicit Sentiment AnalysisabstractRecent work has addressed opinion inferences that arise when opinions are expressed toward +/-effect events, events that positively or negatively affect entities. Many words have mixtures of senses with different +/-effect labels, and therefore word sense disambiguation is needed to exploit +/-effect information for sentiment analysis. This paper presents a knowledge-based +/-effect coarse-grained sense disambiguation method based on selectional preferences modeled via topic models. The method achieves an overall accuracy of 0.83, which represents a significant improvement over three competitive baselines. Yoonjung Choi, Janyce Wiebe, Rada Mihalcea |
IEEE Trans. Affect. Comput. | 3 |
| 2017 | Detecting Deceptive Behavior via Integration of Discriminative Features From Multiple ModalitiesabstractDeception detection has received an increasing amount of attention in recent years, due to the significant growth of digital media, as well as increased ethical and security concerns. Earlier approaches to deception detection were mainly focused on law enforcement applications and relied on polygraph tests, which had proved to falsely accuse the innocent and free the guilty in multiple cases. In this paper, we explore a multimodal deception detection approach that relies on a novel data set of 149 multimodal recordings, and integrates multiple physiological, linguistic, and thermal features. We test the system on different domains, to measure its effectiveness and determine its limitations. We also perform feature analysis using a decision tree model, to gain insights into the features that are most effective in detecting deceit. Our experimental results indicate that our multimodal approach is a promising step toward creating a feasible, non-invasive, and fully automated deception detection system. Mohamed Abouelenien, Verónica Pérez-Rosas, Rada Mihalcea, Mihai Burzo |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2016 | What's Hot in Human Language Technology: Highlights from NAACL HLT 2015abstractThis paper shows a few examples to highlight the trends observed at the NAACL HLT 2015 conference. Joyce Y. Chai, Anoop Sarkar, Rada Mihalcea |
AAAI | 3 |
| 2016 | Identifying Cross-Cultural Differences in Word UsageabstractPersonal writings have inspired researchers in the fields of linguistics and psychology to study the relationship between language and culture to better understand the psychology of people across different cultures. In this paper, we explore this relation by developing cross-cultural word models to identify words with cultural bias – i.e., words that are used in significantly different ways by speakers from different cultures. Focusing specifically on two cultures: United States and Australia, we identify a set of words with significant usage differences, and further investigate these words through feature analysis and topic modeling, shedding light on the attributes of language that contribute to these differences. Aparna Garimella, Rada Mihalcea, James W. Pennebaker |
COLING | 2 |
| 2016 | Targeted Sentiment to Understand Student CommentsabstractWe address the task of targeted sentiment as a means of understanding the sentiment that students hold toward courses and instructors, as expressed by students in their comments. We introduce a new dataset consisting of student comments annotated for targeted sentiment and describe a system that can both identify the courses and instructors mentioned in student comments, as well as label the students’ sentiment toward those entities. Through several comparative evaluations, we show that our system outperforms previous work on a similar task. Charles Welch, Rada Mihalcea |
COLING | 2 |
| 2016 | Structured Matching for Phrase Localization
Mahmoud Azab, Noriyuki Kojima, Rada Mihalcea, Jia Deng 0001 |
ECCV (8) | 4 |
| 2016 | Building a Dataset for Possessions Identification in Text
Carmen Banea, Rada Mihalcea |
LREC | 3 |
| 2015 | Mining semantic affordances of visual object categoriesabstractAffordances are fundamental attributes of objects. Affordances reveal the functionalities of objects and the possible actions that can be performed on them. Understanding affordances is crucial for recognizing human activities in visual data and for robots to interact with the world. In this paper we introduce the new problem of mining the knowledge of semantic affordance: given an object, determining whether an action can be performed on it. This is equivalent to connecting verb nodes and noun nodes in WordNet, or filling an affordance matrix encoding the plausibility of each action-object pair. We introduce a new benchmark with crowdsourced ground truth affordances on 20 PASCAL VOC object classes and 957 action classes. We explore a number of approaches including text mining, visual mining, and collaborative filtering. Our analyses yield a number of significant insights that reveal the most effective ways of collecting knowledge of semantic affordances. Yu-Wei Chao, Rada Mihalcea, Jia Deng 0001 |
CVPR | 3 |
| 2015 | Co-Training for Topic Classification of Scholarly DataabstractWith the exponential growth of scholarly data during the past few years, effective methods for topic classification are greatly needed.Current approaches usually require large amounts of expensive labeled data in order to make accurate predictions.In this paper, we posit that, in addition to a research article's textual content, its citation network also contains valuable information.We describe a co-training approach that uses the text and citation information of a research article as two different views to predict the topic of an article.We show that this method improves significantly over the individual classifiers, while also bringing a substantial reduction in the amount of labeled data required for training accurate classifiers. Cornelia Caragea, Florin Adrian Bulgarov, Rada Mihalcea |
EMNLP | 3 |
| 2015 | Verbal and Nonverbal Clues for Real-life Deception DetectionabstractDeception detection has been receiving an increasing amount of attention from the computational linguistics, speech, and multimodal processing communities. One of the major challenges encountered in this task is the availability of data, and most of the research work to date has been conducted on acted or artificially collected data. The generated deception models are thus lacking real-world evidence. In this paper, we explore the use of multimodal real-life data for the task of deception detection. We develop a new deception dataset consisting of videos from reallife scenarios, and build deception tools relying on verbal and nonverbal features. We achieve classification accuracies in the range of 77-82% when using a model that extracts and fuses features from the linguistic and visual modalities. We show that these results outperform the human capability of identifying deceit. Verónica Pérez-Rosas, Mohamed Abouelenien, Rada Mihalcea, C. J. Linton, Mihai Burzo |
EMNLP | 3 |
| 2015 | Experiments in Open Domain Deception DetectionabstractThe widespread use of deception in online sources has motivated the need for methods to automatically profile and identify deceivers.This work explores deception, gender and age detection in short texts using a machine learning approach.First, we collect a new open domain deception dataset also containing demographic data such as gender and age.Second, we extract feature sets including n-grams, shallow and deep syntactic features, semantic features, and syntactic complexity and readability metrics.Third, we build classifiers that aim to predict deception, gender, and age.Our findings show that while deception detection can be performed in short texts even in the absence of a predetermined domain, gender and age prediction in deceptive texts is a challenging task.We further explore the linguistic differences in deceptive content that relate to deceivers gender and age and find evidence that both age and gender play an important role in people's word choices when fabricating lies. Verónica Pérez-Rosas, Rada Mihalcea |
EMNLP | 2 |
| 2015 | Deception Detection using Real-life Trial DataabstractHearings of witnesses and defendants play a crucial role when reaching court trial decisions. Given the high-stake nature of trial outcomes, implementing accurate and effective computational methods to evaluate the honesty of court testimonies can offer valuable support during the decision making process. In this paper, we address the identification of deception in real-life trial data. We introduce a novel dataset consisting of videos collected from public court trials. We explore the use of verbal and non-verbal modalities to build a multimodal deception detection system that aims to discriminate between truthful and deceptive statements provided by defendants and witnesses. We achieve classification accuracies in the range of 60-75% when using a model that extracts and fuses features from the linguistic and gesture modalities. In addition, we present a human deception detection study where we evaluate the human capability of detecting deception in trial hearings. The results show that our system outperforms the human capability of identifying deceit. Verónica Pérez-Rosas, Mohamed Abouelenien, Rada Mihalcea, Mihai Burzo |
ICMI | 3 |
| 2015 | Values in Words: Using Language to Evaluate and Understand Personal Values
Ryan L. Boyd, Steven R. Wilson 0001, James W. Pennebaker, Michal Kosinski, David Stillwell, Rada Mihalcea |
ICWSM | 6 |
| 2015 | Using Word Semantics To Assist English as a Second Language LearnersabstractWe introduce an interactive interface that aims to help English as a Second Language (ESL) students overcome language related hindrances while reading a text.The interface allows the user to find supplementary information on selected difficult words.The interface is empowered by our lexical substitution engine that provides context-based synonyms for difficult words.We also provide a practical solution for a real-world usage scenario.We demonstrate using the lexical substitution engine -as a browser extension that can annotate and disambiguate difficult words on any webpage. Mahmoud Azab, Chris Hokamp, Rada Mihalcea |
HLT-NAACL | 3 |
| 2015 | Word from the editorsabstractGraph structures naturally model connections. In natural language processing (NLP) connections are ubiquitous, on anything between small and web scale. We find them between words – as grammatical, collocation or semantic relations – contributing to the overall meaning, and maintaining the cohesive structure of the text and the discourse unity. We find them between concepts in ontologies or other knowledge repositories – since the early ages of artificial intelligence, associative or semantic networks have been proposed and used as knowledge stores, because they naturally capture the language units and relations between them, and allow for a variety of inference and reasoning processes, simulating some of the functionalities of the human mind. We find them between complete texts or web pages, and between entities in a social network, where they model relations at the web scale. Beyond the more often encountered ‘regular’ graphs, hypergraphs have also appeared in our field to model relations between more than two units. Zornitsa Kozareva, Vivi Nastase, Rada Mihalcea |
Nat. Lang. Eng. | 3 |
| 2015 | A survey of graphs in natural language processingabstractAbstract Graphs are a powerful representation formalism that can be applied to a variety of aspects related to language processing. We provide an overview of how Natural Language Processing problems have been projected into the graph framework, focusing in particular on graph construction – a crucial step in modeling the data to emphasize the phenomena targeted. Vivi Nastase, Rada Mihalcea, Dragomir R. Radev |
Nat. Lang. Eng. | 2 |
| 2014 | Iterative Constrained Clustering for Subjectivity Word Sense DisambiguationabstractSubjectivity word sense disambiguation (SWSD) is a supervised and application-specific word sense disambiguation task disambiguating between subjective and objective senses of a word. Not sur-prisingly, SWSD suffers from the knowl-edge acquisition bottleneck. In this work, we use a “cluster and label ” strategy to generate labeled data for SWSD semi-automatically. We define a new algo-rithm called Iterative Constrained Cluster-ing (ICC) to improve the clustering purity and, as a result, the quality of the gener-ated data. Our experiments show that the Cem Akkaya, Janyce Wiebe, Rada Mihalcea |
EACL | 3 |
| 2014 | Deception detection using a multimodal approachabstractIn this paper we address the automatic identification of deceit by using a multimodal approach. We collect deceptive and truthful responses using a multimodal setting where we acquire data using a microphone, a thermal camera, as well as physiological sensors. Among all available modalities, we focus on three modalities namely, language use, physiological response, and thermal sensing. To our knowledge, this is the first work to integrate these specific modalities to detect deceit. Several experiments are carried out in which we first select representative features for each modality, and then we analyze joint models that integrate several modalities. The experimental results show that the combination of features from different modalities significantly improves the detection of deceptive behaviors as compared to the use of one modality at a time. Moreover, the use of non-contact modalities proved to be comparable with and sometimes better than existing contact-based methods. The proposed method increases the efficiency of detecting deceit by avoiding human involvement in an attempt to move towards a completely automated non-invasive deception detection process. Mohamed Abouelenien, Verónica Pérez-Rosas, Rada Mihalcea, Mihai Burzo |
ICMI | 3 |
| 2014 | Modeling Language Proficiency Using Implicit Feedback
Chris Hokamp, Rada Mihalcea, Peter Schuelke |
LREC | 2 |
| 2014 | Building a Dataset for Summarization and Keyword Extraction from Emails
Vanessa Loza, Shibamouli Lahiri, Rada Mihalcea, Po-Hsiang Lai |
LREC | 3 |
| 2014 | A Multimodal Dataset for Deception Detection
Verónica Pérez-Rosas, Rada Mihalcea, Alexis Narvaez, Mihai Burzo |
LREC | 2 |
| 2014 | Computational approaches to subjectivity and sentiment analysis: Present and envisaged methods and applications
Alexandra Balahur, Rada Mihalcea, Andrés Montoyo |
Comput. Speech Lang. | 2 |
| 2014 | Sense-level subjectivity in a multilingual setting
Carmen Banea, Rada Mihalcea, Janyce Wiebe |
Comput. Speech Lang. | 2 |
| 2014 | Explorations in lexical sample and all-words lexical substitutionabstractIn this paper, we experiment with several techniques to solve the problem of lexical substitution, both in alexical sampleas well as anall-wordssetting, and compare the benefits of combining multiple lexical resources using both unsupervised and supervised approaches. Overall in the lexical sample setting, the results obtained through the combination of several resources exceed the current state-of-the-art when selecting the best substitute for a given target word, and place second when selecting the top ten substitutes, thus demonstrating the usefulness of the approach. Further, we put forth a novel exploration in all-words lexical substitution and set ground for further explorations of this more generalized setting. Ravi Som Sinha, Rada Mihalcea |
Nat. Lang. Eng. | 2 |
| 2013 | Utterance-Level Multimodal Sentiment Analysis
Verónica Pérez-Rosas, Rada Mihalcea, Louis-Philippe Morency |
ACL (1) | 2 |
| 2013 | Automatic detection of deceit in verbal communicationabstractThis paper presents experiments in building a classifier for the automatic detection of deceit. Using a dataset of deceptive videos, we run several comparative evaluations focusing on the verbal component of these videos, with the goal of understanding the difference in deceit detection when using manual versus automatic transcriptions, as well as the difference between spoken and written lies. We show that using only the linguistic component of the deceptive videos, we can detect deception with accuracies in the range of 52-73%. Rada Mihalcea, Verónica Pérez-Rosas, Mihai Burzo |
ICMI | 1 |
| 2013 | Multilingual Word Sense Disambiguation Using Wikipedia
Bharath Dandala, Rada Mihalcea, Razvan C. Bunescu |
IJCNLP | 2 |
| 2013 | Sentiment analysis of online spoken reviewsabstractThis paper describes several experiments in building a sentiment analysis classifier for spoken reviews. We specifically focus on the linguistic component of these reviews, with the goal of understanding the difference in sentiment classification performance when using manual versus automatic transcriptions, as well as the difference between spoken and written reviews. We introduce a novel dataset, consisting of video reviews for two different domains (cellular phones and fiction books), and we show that using only the linguistic component of these reviews we can obtain sentiment classifiers with accuracies in the range of 65-75%. Verónica Pérez-Rosas, Rada Mihalcea |
INTERSPEECH | 2 |
| 2013 | Porting Multilingual Subjectivity Resources across LanguagesabstractSubjectivity analysis focuses on the automatic extraction of private states in natural language. In this paper, we explore methods for generating subjectivity analysis resources in a new language by leveraging on the tools and resources available in English. Given a bridge between English and the selected target language (e.g., a bilingual dictionary or a parallel corpus), the methods can be used to rapidly create tools for subjectivity analysis in the new language. Carmen Banea, Rada Mihalcea, Janyce Wiebe |
IEEE Trans. Affect. Comput. | 2 |
| 2012 | Lyrics, Music, and Emotions
Rada Mihalcea, Carlo Strapparava |
EMNLP-CoNLL | 1 |
| 2012 | Towards sensing the influence of visual narratives on human affectabstractIn this paper, we explore a multimodal approach to sensing affective state during exposure to visual narratives. Using four different modalities, consisting of visual facial behaviors, thermal imaging, heart rate measurements, and verbal descriptions, we show that we can effectively predict changes in human affect. Our experiments show that these modalities complement each other, and illustrate the role played by each of the four modalities in detecting human affect. Mihai Burzo, Daniel McDuff, Rada Mihalcea, Louis-Philippe Morency, Alexis Narvaez, Verónica Pérez-Rosas |
ICMI | 3 |
| 2012 | Towards multimodal deception detection - step 1: building a collection of deceptive videosabstractIn this paper, we introduce a novel crowdsourced dataset of deceptive videos. We describe the collection process and the characteristics of the dataset, and we validate it through initial experiments in the recognition of deceptive language. The collection, consisting of 140 truthful and deceptive videos, will enable future experiments in multimodal deceptive detection. Rada Mihalcea, Mihai Burzo |
ICMI | 1 |
| 2012 | Unsupervised Word Sense Disambiguation with Multilingual Representations
Erwin Fernandez-Ordonez, Rada Mihalcea, Samer Hassan 0002 |
LREC | 2 |
| 2012 | Learning Sentiment Lexicons in Spanish
Verónica Pérez-Rosas, Carmen Banea, Rada Mihalcea |
LREC | 3 |
| 2012 | A Parallel Corpus of Music and Lyrics Annotated with Emotions
Carlo Strapparava, Rada Mihalcea, Alberto Battocchi |
LREC | 2 |
| 2011 | Semantic Relatedness Using Salient Semantic AnalysisabstractThis paper introduces a novel method for measuring semantic relatedness using semantic profiles constructed from salient encyclopedic features. The model is built on the notion that the meaning of a word can be characterized by the salient concepts found in its immediate context. In addition to being computationally efficient, the new model has superior performance and remarkable consistency when compared to both knowledge-based and corpus-based state-of-the-art semantic relatedness models. Samer Hassan 0002, Rada Mihalcea |
AAAI | 2 |
| 2011 | A Comparison of Unsupervised Methods to Associate Colors with Words
Gözde Özbal, Carlo Strapparava, Rada Mihalcea, Daniele Pighin |
ACII (2) | 3 |
| 2011 | Learning to Grade Short Answer Questions using Semantic Similarity Measures and Dependency Graph Alignments
Michael Mohler, Razvan C. Bunescu, Rada Mihalcea |
ACL | 3 |
| 2011 | Improving the Impact of Subjectivity Word Sense Disambiguation on Contextual Opinion Analysis
Cem Akkaya, Janyce Wiebe, Alexander Conrad, Rada Mihalcea |
CoNLL | 4 |
| 2011 | Towards multimodal sentiment analysis: harvesting opinions from the webabstractWith more than 10,000 new videos posted online every day on social websites such as YouTube and Facebook, the internet is becoming an almost infinite source of information. One crucial challenge for the coming decade is to be able to harvest relevant information from this constant flow of multimodal data. This paper addresses the task of multimodal sentiment analysis, and conducts proof-of-concept experiments that demonstrate that a joint model that integrates visual, audio, and textual features can be effectively used to identify sentiment in Web videos. This paper makes three important contributions. First, it addresses for the first time the task of tri-modal sentiment analysis, and shows that it is a feasible task that can benefit from the joint exploitation of visual, audio and textual modalities. Second, it identifies a subset of audio-visual features relevant to sentiment analysis and present guidelines on how to integrate these features. Finally, it introduces a new dataset consisting of real online data, which will be useful for future research in this area. Louis-Philippe Morency, Rada Mihalcea, Payal Doshi |
ICMI | 2 |
| 2011 | Going Beyond Text: A Hybrid Image-Text Approach for Measuring Word Relatedness
Chee Wee Leong, Rada Mihalcea |
IJCNLP | 2 |
| 2010 | Computational Models for Incongruity Detection in Humour
Rada Mihalcea, Carlo Strapparava, Stephen G. Pulman |
CICLing | 1 |
| 2010 | Multilingual Subjectivity: Are More Languages Better?
Carmen Banea, Rada Mihalcea, Janyce Wiebe |
COLING | 2 |
| 2010 | Cross Language Text Classification by Model Translation and Semi-Supervised Learning
Rada Mihalcea, Mingjun Tian |
EMNLP | 2 |
| 2010 | Quantifying the Limits and Success of Extractive Summarization Systems Across Domains
Hakan Ceylan, Rada Mihalcea, Umut Ozertem, Elena Lloret, Manuel Palomar |
HLT-NAACL | 2 |
| 2009 | The Decomposition of Human-Written Book Summaries
Hakan Ceylan, Rada Mihalcea |
CICLing | 2 |
| 2009 | Linguistic Ethnography: Identifying Dominant Word Classes in Text
Rada Mihalcea, Stephen G. Pulman |
CICLing | 1 |
| 2009 | Using Encyclopedic Knowledge for Automatic Topic Identification
Kino Coursey, Rada Mihalcea, William E. Moen |
CoNLL | 2 |
| 2009 | Text-to-Text Semantic Similarity for Automatic Short Answer Grading
Michael Mohler, Rada Mihalcea |
EACL | 2 |
| 2009 | Subjectivity Word Sense Disambiguation
Cem Akkaya, Janyce Wiebe, Rada Mihalcea |
EMNLP | 3 |
| 2009 | Cross-lingual Semantic Relatedness Using Encyclopedic Knowledge
Samer Hassan 0002, Rada Mihalcea |
EMNLP | 2 |
| 2009 | A natural language interface for crime-related spatial queriesabstractWeb-based mapping applications such as Google Maps or Virtual Earth have become increasingly popular. However, current map search is still keyword-based and supports a limited number of spatial predicates. In this paper, we build towards a natural language query interface to spatial databases to answer crime-related spatial queries. The system has two main advantages compared with interfaces such as Google Maps: (1) It allows query conditions to be expressed in natural language, and (2) It supports a larger number of spatial predicates, such as ldquowithin 3 milesrdquo and ldquoclose tordquo. The system is evaluated using a set of crime-related queries run against a dataset that contains many spatial layers in the Denton, Texas area. The results show that our approach significantly outperforms Google Maps when processing complicated spatial queries. Yan Huang 0002, Rada Mihalcea, Hector Cuellar |
ISI | 3 |
| 2009 | Integrating Knowledge for Subjectivity Sense Labeling
Yaw Gyamfi, Janyce Wiebe, Rada Mihalcea, Cem Akkaya |
HLT-NAACL | 3 |
| 2008 | Linguistically Motivated Features for Enhanced Back-of-the-Book Indexing
Andras Csomai, Rada Mihalcea |
ACL | 2 |
| 2008 | Multilingual Subjectivity Analysis Using Machine Translation
Carmen Banea, Rada Mihalcea, Janyce Wiebe, Samer Hassan 0002 |
EMNLP | 2 |
| 2008 | How to Add a New Language on the NLP Map: Building Resources and Tools for Languages with Scarce Resources
Rada Mihalcea, Vivi Nastase |
IJCNLP | 1 |
| 2008 | A Bootstrapping Method for Building Subjectivity Lexicons for Languages with Scarce Resources
Carmen Banea, Rada Mihalcea, Janyce Wiebe |
LREC | 2 |
| 2008 | Babylon Parallel Text Builder: Gathering Parallel Texts for Low-Density Languages
Michael Mohler, Rada Mihalcea |
LREC | 2 |
| 2008 | The Text Mining Handbook: Advanced Approaches to Analyzing Unstructured Data Ronen Feldman and James Sanger (Bar-Ilan University and ABS Ventures) Cambridge, England: Cambridge University Press, 2007, xii+410 pp; hardbound, ISBN 0-521-83657-3
Rada Mihalcea |
Comput. Linguistics | 1 |
| 2008 | Toward communicating simple sentences using pictorial representations
Rada Mihalcea, Chee Wee Leong |
Mach. Transl. | 1 |
| 2007 | Learning Multilingual Subjective Language via Cross-Lingual Projections
Rada Mihalcea, Carmen Banea, Janyce Wiebe |
ACL | 1 |
| 2007 | Linking Educational Materials to Encyclopedic Knowledge
Andras Csomai, Rada Mihalcea |
AIED | 2 |
| 2007 | Characterizing Humour: An Exploration of Features in Humorous Texts
Rada Mihalcea, Stephen G. Pulman |
CICLing | 1 |
| 2007 | Wikify!: linking documents to encyclopedic knowledgeabstractThis paper introduces the use of Wikipedia as a resource for automatic keyword extraction and word sense disambiguation, and shows how this online encyclopedia can be used to achieve state-of-the-art results on both these tasks. The paper also shows how the two methods can be combined into a system able to automatically enrich a text with links to encyclopedic knowledge. Given an input document, the system identifies the important concepts in the text and automatically links these concepts to the corresponding Wikipedia pages. Evaluations of the system show that the automatic annotations are reliable and hardly distinguishable from manual annotations. Rada Mihalcea, Andras Csomai |
CIKM | 1 |
| 2007 | Explorations in Automatic Book Summarization
Rada Mihalcea, Hakan Ceylan |
EMNLP-CoNLL | 1 |
| 2007 | Of Men, Women, and Computers: Data-Driven Gender Modeling for Improved User Interfaces
Hugo Liu, Rada Mihalcea |
ICWSM | 2 |
| 2007 | Using Wikipedia for Automatic Word Sense Disambiguation
Rada Mihalcea |
HLT-NAACL | 1 |
| 2006 | Corpus-based and Knowledge-based Measures of Text Semantic Similarity
Rada Mihalcea, Courtney D. Corley, Carlo Strapparava |
AAAI | 1 |
| 2006 | Word Sense and SubjectivityabstractSubjectivity and meaning are both important properties of language. This paper explores their interaction, and brings empirical evidence in support of the hypotheses that (1) subjectivity is a property that can be associated with word senses, and (2) word sense disambiguation can directly benefit from subjectivity annotations. Janyce Wiebe, Rada Mihalcea |
ACL | 2 |
| 2006 | Creating a Testbed for the Evaluation of Automatically Generated Back-of-the-Book Indexes
Andras Csomai, Rada Mihalcea |
CICLing | 2 |
| 2006 | Random Walks on Text Structures
Rada Mihalcea |
CICLing | 1 |
| 2006 | NLP (Natural Language Processing) for NLP (Natural Language Programming)
Rada Mihalcea, Hugo Liu, Henry Lieberman |
CICLing | 1 |
| 2006 | Graph-based Algorithms for Natural Language Processing and Information Retrieval
Rada Mihalcea, Dragomir R. Radev |
HLT-NAACL | 1 |
| 2006 | Learning to Laugh (automatically): Computational Models for Humor RecognitionabstractHumor is one of the most interesting and puzzling aspects of human behavior. Despite the attention it has received in fields such as philosophy, linguistics, and psychology, there have been only few attempts to create computational models for humor recognition or generation. In this article, we bring empirical evidence that computational approaches can be successfully applied to the task of humor recognition. Through experiments performed on very large data sets, we show that automatic classification techniques can be effectively used to distinguish between humorous and non‐humorous texts, with significant improvements observed over a priori known baselines. Rada Mihalcea, Carlo Strapparava |
Comput. Intell. | 1 |
| 2005 | Language Independent Extractive Summarization
Rada Mihalcea |
AAAI | 1 |
| 2005 | Language Independent Extractive Summarization
Rada Mihalcea |
ACL | 1 |
| 2005 | SenseLearner: Word Sense Disambiguation for All Words in Unrestricted Text
Rada Mihalcea, Andras Csomai |
ACL | 1 |
| 2005 | Putting Pieces Together: Combining FrameNet, VerbNet and WordNet for Robust Semantic Parsing
Rada Mihalcea |
CICLing | 2 |
| 2005 | Parallel textsabstractParallel texts have become a vital element for natural language processing. We present a panorama of current research activities related to parallel texts, and offer some thoughts about the future of this rich field of investigation. Rada Mihalcea, Michel Simard |
Nat. Lang. Eng. | 1 |
| 2004 | PageRank on Semantic Networks, with Application to Word Sense Disambiguation
Rada Mihalcea, Paul Tarau, Elizabeth Figa |
COLING | 1 |
| 2004 | Co-training and Self-training for Word Sense Disambiguation
Rada Mihalcea |
CoNLL | 1 |
| 2004 | Voted Co-Training for Bootstrapping Sense Classifiers
Rada Mihalcea |
ECAI | 1 |
| 2004 | TextRank: Bringing Order into Text
Rada Mihalcea, Paul Tarau |
EMNLP | 1 |
| 2004 | Finding Semantic Associations on Express Lane
Vivi Nastase, Rada Mihalcea |
LREC | 2 |
| 2003 | Performance Analysis of a Part of Speech Tagging Task
Rada Mihalcea |
CICLing | 1 |
| 2003 | Turning WordNet into an Information Retrieval Resource: Systematic Polysemy and Conversion to Hierarchical CodesabstractThis paper addresses the problem of transforming WordNet into a resource tailored to Information Retrieval (IR) applications. We address two of the major drawbacks pointed out in previous literature in relation to this semantic network. One is the fine granularity of senses defined in WordNet, which proves useless from an IR perspective. To solve this problem, we propose a set of methods that enable the automatic transformation of WordNet into a coarse grained dictionary. The other drawback is the encoding used in this resource, and the methods for accessing related words across the semantic net. Due to the high number of connections among concepts, the simple computation of a path in this net, or the generation of related concepts may become a computationally intensive process. This effect is highly undesirable in time sensitive applications such as IR applications. We propose a methodology for hierarchical encoding that enables increased efficiency in WordNet-based IR systems. Rada Mihalcea |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2002 | Diacritics Restoration: Learning from Letters versus Learning from Words
Rada Mihalcea |
CICLing | 1 |
| 2002 | Instance Based Learning with Automatic Feature Selection Applied to Word Sense Disambiguation
Rada Mihalcea |
COLING | 1 |
| 2002 | Letter Level Learning for Language Independent Diacritics Restoration
Rada Mihalcea, Vivi Nastase |
CoNLL | 1 |
| 2002 | Bootstrapping Large Sense Tagged Corpora
Rada Mihalcea |
LREC | 1 |
| 2002 | Word sense disambiguation with pattern learning and automatic feature selectionabstractThis paper presents a novel approach for word sense disambiguation. The underlying algorithm has two main components: (1) pattern learning from available sense-tagged corpora (SemCor), from dictionary definitions (WordNet) and from a generated corpus (GenCor); and (2) instance based learning with automatic feature selection, when training data is available for a particular word. The ideas described in this paper were implemented in a system that achieves excellent performance on the data provided during the SENSEVAL-2 evaluation exercise, for both English all words and English lexical sample tasks. Rada Mihalcea |
Nat. Lang. Eng. | 1 |
| 2001 | The Role of Lexico-Semantic Feedback in Open-Domain Textual Question-AnsweringabstractThis paper presents an open-domain textual Question-Answering system that uses several feedback loops to enhance its performance. These feedback loops combine in a new way statistical results with syntactic, semantic or pragmatic information derived from texts and lexical databases. The paper presents the contribution of each feedback loop to the overall performance of 76% human-assessed precise answers. Sanda M. Harabagiu, Dan I. Moldovan, Marius Pasca, Rada Mihalcea, Mihai Surdeanu, Razvan C. Bunescu, Roxana Girju, Vasile Rus, Paul Morarescu |
ACL | 4 |
| 2001 | Word Semantics for Information Retrieval: Moving One Step Closer to the Semantic WebabstractThe goal of the Semantic Web is to create a new form of Web content meaningful to computers. The Semantic Web aims to provide greater functionality, via intelligent tools such as information extractors, brokers, reasoning services or question answering systems. Semantics can be addressed at several levels. In this paper, we focus on the lowest level-word semantics on which other higher levels such as concept, paragraph, or document levels can be based upon. This model, which we call Word Semantics (WS), does not include the rich set of tags proposed by the XML/RDF standards. Nevertheless, this simpler WS format comes with a big advantage: it is possible with existing technologies and resources. Practically, this new model relies on understanding word meanings, identifying important named entities such as person, organization and others, and linking all this information via an external general purpose ontology, namely WordNet. With these features, we regard the WS model as a short but strong step toward the long term goal of a Semantic Web. Rada Mihalcea, Silvana I. Mihalcea |
ICTAI | 1 |
| 2000 | The Structure and Performance of an Open-Domain Question Answering SystemabstractThis paper presents the architecture, operation and results obtained with the LASSO Question Answering system developed in the Natural Language Processing Laboratory at SMU. To find answers, the system relies on a combination of syntactic and semantic techniques. The search for the answer is based on a novel form of indexing called paragraph indexing. A score of 55.5% for short answers and 64.5% for long answers was achieved at the TREC-8 competition. Dan I. Moldovan, Sanda M. Harabagiu, Marius Pasca, Rada Mihalcea, Roxana Girju, Richard Goodrum, Vasile Rus |
ACL | 4 |
| 2000 | AutoASC - A System for Automatic Acquisition of Sense Tagged CorporaabstractMany natural language processing tasks, such as word sense disambiguation, knowledge acquisition, information retrieval, use semantically tagged corpora. Till recently, these corpus-based systems relied on text manually annotated with semantic tags; but the massive human intervention in this process has become a serious impediment in building robust systems. In this paper, we present AutoASC, a system which automatically acquires sense tagged corpora. It is based on (1) the information provided in WordNet, particularly the word definitions found within the glosses and (2) the information gathered from Internet using existing search engines. The system was tested on a set of 46 concepts, for which 2071 example sentences have been acquired; for these, a precision of 87% was observed. Rada Mihalcea, Dan I. Moldovan |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1999 | A Method for Word Sense Disambiguation of Unrestricted TextabstractSelecting the most appropriate sense for an ambiguous word in a sentence is a central problem in Natural Language Processing. In this paper, we present a method that attempts to disambiguate all the nouns, verbs, adverbs and adjectives in a text, using the senses provided in WordNet. The senses are ranked using two sources of information: (1) the Internet for gathering statistics for word-word cooccurrences and (2) WordNet for measuring the semantic density for a pair of words. We report an average accuracy of 80% for the first ranked sense, and 91% for the first two ranked senses. Extensions of this method for larger windows of more than two words are considered. Rada Mihalcea, Dan I. Moldovan |
ACL | 1 |