Malihe Alikhani

dblp:163/2171 · DBLP profile ↗
← Back
45ranked-venue papers
4as first author
39since 2021 · last 2026
0000-0002-1315-2228ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 4 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
abstract
Ziyi Wang, Yuxuan Lu, Wenbo Li, Amirali Amini, Bo Sun, Yakov Bart, Weimin Lyu, Jiri Gesi, Tian Wang, Jing Huang, Yu Su, Upol Ehsan, Malihe Alikhani, Toby Jia-Jun Li, Lydia Chilton, Dakuo Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuxuan Lu 0003, Amirali Amini, Yakov Bart, Weimin Lyu, Jiri Gesi, Upol Ehsan, Malihe Alikhani, Toby Jia-Jun Li, Lydia B. Chilton, Dakuo Wang
ACL (1)13
2026 How Pragmatics Shape Articulation: A Computational Case Study in STEM ASL Discourse
Saki Imai, Lee Kezar, Laurel Aichler, Mert Inan, Erin Walker, Alicia Wooten, Lorna C. Quandt, Malihe Alikhani
LREC8
2025 Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others from Conversational Cues
abstract
Typically, when evaluating Theory of Mind, we consider the beliefs of others to be binary: held or not held. But what if someone is unsure about their own beliefs? How can we quantify this uncertainty? We propose a new suite of tasks, challenging language models (LMs) to model the uncertainty of participants in a dialogue. We design these tasks around conversation forecasting, where the goal is to predict the probability of an unobserved conversation outcome. Uniquely, we view conversation agents themselves as forecasters, asking an LM to predict the uncertainty of an individual from their language use. We experiment with scaling methods, bagging, and demographic context for this regression task, conducting experiments on three dialogue corpora (social, negotiation, task-oriented) with eight LMs. While LMs can explain up to 7% variance in the uncertainty of others, we highlight the difficulty of the tasks and room for future work, especially in tasks that require explicit shifts in perspective.
Anthony Sicilia, Malihe Alikhani
ACL (1)2
2025 An Active Learning Framework for Inclusive Generation by Large Language Models
abstract
Ensuring that Large Language Models (LLMs) generate text representative of diverse sub-populations is essential, particularly when key concepts related to under-represented groups are scarce in the training data. We address this challenge with a novel clustering-based active learning framework, enhanced with knowledge distillation. The proposed framework transforms the intermediate outputs of the learner model, enabling effective active learning for generative tasks for the first time. Integration of clustering and knowledge distillation yields more representative models without prior knowledge of underlying data distribution and overbearing human efforts. We validate our approach in practice through case studies in counter-narration and style transfer. We construct two new datasets in tandem with model training, showing a performance improvement of 2%–10% over baseline models. Our results also show more consistent performance across various data subgroups and increased lexical diversity, underscoring our model’s resilience to skewness in available data. Further, our results show that the data acquired via our approach improves the performance of secondary models not involved in the learning loop, showcasing practical utility of the framework.
Sabit Hassan, Anthony Sicilia, Malihe Alikhani
COLING3
2025 Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation
abstract
Establishing shared goals is a fundamental step in human-AI communication.However, ambiguities can lead to outputs that seem correct but fail to reflect the speaker's intent.In this paper, we explore this issue with a focus on the data visualization domain, where ambiguities in natural language impact the generation of code that visualizes data.The availability of multiple views on the contextual (e.g. the intended plot and the code rendering the plot) allows for a unique and comprehensive analysis of diverse ambiguity types.We develop a taxonomy of types of ambiguity that arise in this task and propose metrics to quantify them.Using Matplotlib problems from the DS-1000 dataset, we demonstrate that our ambiguity metrics better correlate with human annotations than uncertainty baselines.Our work also explores how multi-turn dialogue can reduce ambiguity, and therefore, improve code accuracy by better matching user goals.We evaluate three pragmatic models to inform our dialogue strategies: Gricean Cooperativity, Discourse Representation Theory, and Questions under Discussion.A simulated user study reveals how pragmatic dialogues reduce ambiguity and enhance code accuracy, highlighting the value of multi-turn exchanges in code generation.
Mert Inan, Anthony Sicilia, Alex Xie, Saujas Vaduguru, Daniel Fried, Malihe Alikhani
EMNLP6
2025 Coherence-Driven Multimodal Safety Dialogue with Active Learning for Embodied Agents
Sabit Hassan, Hye-Young Chung, Xiang Zhi Tan, Malihe Alikhani
AAMAS4
2025 How to Align Multiple Signed Language Corpora for Better Sign-to-Sign Translations?
abstract
Mert Inan, Yang Zhong, Vidya Ganesh, Malihe Alikhani. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Mert Inan, Vidya Ganesh, Malihe Alikhani
NAACL (Long Papers)4
2024 Combining Discourse Coherence with Large Language Models for More Inclusive, Equitable, and Robust Task-Oriented Dialogue
abstract
Large language models (LLMs) are capable of generating well-formed responses, but using LLMs to generate responses on the fly is not yet feasible for many task-oriented systems. Modular architectures are often still required for safety and privacy guarantees on the output. We hypothesize that an offline generation approach using discourse theories, formal grammar rules, and LLMs can allow us to generate human-like, coherent text in a more efficient, robust, and inclusive manner within a task-oriented setting. To this end, we present the first discourse-aware multimodal task-oriented dialogue system that combines discourse theories with offline LLM generation. We deploy our bot as an app to the general public and keep track of the user ratings for six months. Our user ratings show an improvement from 2.8 to 3.5 out of 5 with the introduction of discourse coherence theories. We also show that our model reduces misunderstandings in the dialect of African-American Vernacular English from 93% to 57%. While terms of use prevent us from releasing our entire codebase, we release our code in a format that can be integrated into most existing dialogue systems.
Katherine Atwell, Mert Inan, Anthony Sicilia, Malihe Alikhani
LREC/COLING4
2024 Seeing Eye-to-Eye: Cross-Modal Coherence Relations Inform Eye-gaze Patterns During Comprehension & Production
abstract
Context influences how we engage with multimodal documents. Describing and processing the content of images is highly correlated with the goals of the discourse. It is known that these underlying cognitive processes can be tapped into by looking at eye movements, but the connection between discourse goals and eye movements is a missing link. In this study, we carry out both augmented reality and webcam-based eye-tracking experiments during comprehension and production tasks. We build on computational frameworks of coherence in text and images that study causal, logical, elaborative, and temporal inferences to understand how eye gaze patterns and coherence relations influence each other. No state-of-the-art techniques exist to analyze eye movements in multimodal language settings. So, we introduce a new eye gaze pattern ranking algorithm and a semantic gaze visualization technique to study this phenomenon better. Our results demonstrate that eye gaze durations are person-dependent, and during comprehension and production, ranked gaze patterns are significantly different for different types of coherence relations. We also present a case study of how Multimodal Large Language Models represent this connection of eye gaze patterns and coherence relations. We make all of our code and novel analysis tools available through https://github.com/Merterm/eye-gaze-coherence.
Mert Inan, Malihe Alikhani
LREC/COLING2
2024 HumBEL: A Human-in-the-Loop Approach for Evaluating Demographic Factors of Language Models in Human-Machine Conversations
abstract
While demographic factors like age and gender change the way people talk, and in particular, the way people talk to machines, there is little investigation into how large pre-trained language models (LMs) can adapt to these changes.To remedy this gap, we consider how demographic factors in LM language skills can be measured to determine compatibility with a target demographic.We suggest clinical techniques from Speech Language Pathology, which has norms for acquisition of language skills in humans.We conduct evaluation with a domain expert (i.e., a clinically licensed speech language pathologist), and also propose automated techniques to complement clinical evaluation at scale.Empirically, we focus on age, finding LM capability varies widely depending on task: GPT-3.5 mimics the ability of humans ranging from age 6-15 at tasks requiring inference, and simultaneously, outperforms a typical 21 year old at memorization.GPT-3.5 also has trouble with social language use, exhibiting less than 50% of the tested pragmatic skills.Findings affirm the importance of considering demographic alignment and conversational goals when using LMs as public-facing tools.Code, data, and a package will be available.
Anthony Sicilia, Jennifer C. Gates, Malihe Alikhani
EACL (1)3
2024 Studying and Mitigating Biases in Sign Language Understanding Models
abstract
Ensuring that the benefits of sign language technologies are distributed equitably among all community members is crucial.Thus, it is important to address potential biases and inequities that may arise from the design or use of these resources.Crowd-sourced sign language datasets, such as the ASL Citizen dataset, are great resources for improving accessibility and preserving linguistic diversity, but they must be used thoughtfully to avoid reinforcing existing biases.In this work, we utilize the rich information about participant demographics and lexical features present in the ASL Citizen dataset to study and document the biases that may result from models trained on crowd-sourced sign datasets.Further, we apply several bias mitigation techniques during model training, and find that these techniques reduce performance disparities without decreasing accuracy.With the publication of this work, we release the demographic information about the participants in the ASL Citizen dataset to encourage future bias mitigation work in this space.
Katherine Atwell, Danielle Bragg, Malihe Alikhani
EMNLP3
2023 Learning to Generate Equitable Text in Dialogue from Biased Training Data
abstract
The ingrained principles of fairness in a dialogue system's decision-making process and generated responses are crucial for user engagement, satisfaction, and task achievement.Absence of equitable and inclusive principles can hinder the formation of common ground, which in turn negatively impacts the overall performance of the system.For example, misusing pronouns in a user interaction may cause ambiguity about the intended subject.Yet, there is no comprehensive study of equitable text generation in dialogue.Aptly, in this work, we use theories of computational learning to study this problem.We provide formal definitions of equity in text generation, and further, prove formal connections between learning humanlikeness and learning equity: algorithms for improving equity ultimately reduce to algorithms for improving human-likeness (on augmented data).With this insight, we also formulate reasonable conditions under which text generation algorithms can learn to generate equitable text without any modifications to the biased training data on which they learn.To exemplify our theory in practice, we look at a group of algorithms for the GuessWhat?! visual dialogue game and, using this example, test our theory empirically.Our theory accurately predicts relative-performance of multiple algorithms in generating equitable text as measured by both human and automated evaluation.
Anthony Sicilia, Malihe Alikhani
ACL (1)2
2023 How people talk about each other: Modeling Generalized Intergroup Bias and Emotion
abstract
Venkata Subrahmanyan Govindarajan, Katherine Atwell, Barea Sinno, Malihe Alikhani, David I. Beaver, Junyi Jessy Li. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Venkata Subrahmanyan Govindarajan, Katherine Atwell, Barea Sinno, Malihe Alikhani, David Beaver, Junyi Jessy Li
EACL4
2023 PANCETTA: Phoneme Aware Neural Completion to Elicit Tongue Twisters Automatically
abstract
Tongue twisters are meaningful sentences that are difficult to pronounce.The process of automatically generating tongue twisters is challenging since the generated utterance must satisfy two conditions at once: phonetic difficulty and semantic meaning.Furthermore, phonetic difficulty is itself hard to characterize and is expressed in tongue twisters through a heterogeneous mix of phenomena such as alliteration and homophony.In this paper, we propose PANCETTA: Phoneme Aware Neural Completion to Elicit Tongue Twisters Automatically.We leverage phoneme representations to capture the notion of phonetic difficulty, and we train language models to generate original tongue twisters on two proposed task settings.To do this, we curate a dataset called TT-Corp, consisting of existing English tongue twisters.Through automatic and human evaluation, as well as qualitative analysis, we show that PANCETTA generates novel, phonetically difficult, fluent, and semantically meaningful tongue twisters.
Sedrick Keh, Steven Y. Feng, Varun Gangal, Malihe Alikhani, Eduard H. Hovy
EACL4
2023 Multilingual Content Moderation: A Case Study on Reddit
abstract
Meng Ye, Karan Sikka, Katherine Atwell, Sabit Hassan, Ajay Divakaran, Malihe Alikhani. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Meng Ye 0002, Karan Sikka, Katherine Atwell, Sabit Hassan, Ajay Divakaran, Malihe Alikhani
EACL6
2023 SODA: Million-scale Dialogue Distillation with Social Commonsense Contextualization
abstract
Hyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, Yejin Choi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Hyunwoo Kim 0002, Jack Hessel, Peter West, Ximing Lu, Youngjae Yu, Ronan Le Bras 0001, Malihe Alikhani, Gunhee Kim, Maarten Sap, Yejin Choi 0001
EMNLP9
2023 DisCGen: A Framework for Discourse-Informed Counterspeech Generation
abstract
Sabit Hassan, Malihe Alikhani. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Sabit Hassan, Malihe Alikhani
IJCNLP (1)2
2023 Learning Multimodal Cues of Children's Uncertainty
abstract
Qi Cheng, Mert Inan, Rahma Mbarki, Grace Grmek, Theresa Choi, Yiming Sun, Kimele Persaud, Jenny Wang, Malihe Alikhani. Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2023.
Mert Inan, Rahma Mbarki, Grace Grmek, Theresa Choi, Yiming Sun 0004, Kimele Persaud, Jenny Wang, Malihe Alikhani
SIGDIAL9
2022 Cross-Modal Coherence for Text-to-Image Retrieval
abstract
Common image-text joint understanding techniques presume that images and the associated text can universally be characterized by a single implicit model. However, co-occurring images and text can be related in qualitatively different ways, and explicitly modeling it could improve the performance of current joint understanding models. In this paper, we train a Cross-Modal Coherence Model for text-to-image retrieval task. Our analysis shows that models trained with image–text coherence relations can retrieve images originally paired with target text more often than coherence-agnostic models. We also show via human evaluation that images retrieved by the proposed coherence-aware model are preferred over a coherence-agnostic baseline by a huge margin. Our findings provide insights into the ways that different modalities communicate and the role of coherence relations in capturing commonsense inferences in text and imagery.
Malihe Alikhani, Fangda Han, Hareesh Ravi, Mubbasir Kapadia, Vladimir Pavlovic 0001, Matthew Stone
AAAI1
2022 Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation Models
abstract
We investigate the use of multimodal information contained in images as an effective method for enhancing the commonsense of Transformer models for text generation. We perform experiments using BART and T5 on concept-to-text generation, specifically the task of generative commonsense reasoning, or CommonGen. We call our approach VisCTG: Visually Grounded Concept-to-Text Generation. VisCTG involves captioning images representing appropriate everyday scenarios, and using these captions to enrich and steer the generation process. Comprehensive evaluation and analysis demonstrate that VisCTG noticeably improves model performance while successfully addressing several issues of the baseline generations, including poor commonsense, fluency, and specificity.
Steven Y. Feng, Zhuofu Tao, Malihe Alikhani, Teruko Mitamura, Eduard H. Hovy, Varun Gangal
AAAI4
2022 NAREOR: The Narrative Reordering Problem
abstract
Many implicit inferences exist in text depending on how it is structured that can critically impact the text's interpretation and meaning. One such structural aspect present in text with chronology is the order of its presentation. For narratives or stories, this is known as the narrative order. Reordering a narrative can impact the temporal, causal, event-based, and other inferences readers draw from it, which in turn can have strong effects both on its interpretation and interestingness. In this paper, we propose and investigate the task of Narrative Reordering (NAREOR) which involves rewriting a given story in a different narrative order while preserving its plot. We present a dataset, NAREORC, with human rewritings of stories within ROCStories in non-linear orders, and conduct a detailed analysis of it. Further, we propose novel task-specific training methods with suitable evaluation metrics. We perform experiments on NAREORC using state-of-the-art models such as BART and T5 and conduct extensive automatic and human evaluations. We demonstrate that although our models can perform decently, NAREOR is a challenging task with potential for further exploration. We also investigate two applications of NAREOR: generation of more interesting variations of stories and serving as adversarial sets for temporal/event-related tasks, besides discussing other prospective ones, such as for pedagogical setups related to language skills like essay writing and applications to medicine involving clinical narratives.
Varun Gangal, Steven Y. Feng, Malihe Alikhani, Teruko Mitamura, Eduard H. Hovy
AAAI3
2022 Learning variations in children's multimodal cues of uncertainty during a math-related numerical discrimination task
Theresa Choi, Rahma Mbarki, Grace Grmek, Kimele Persaud, Jinjing (Jenny) Wang, Malihe Alikhani
CogSci6
2022 Studying the Effect of Moderator Biases on the Diversity of Online Discussions: A Computational Cross-linguistic Study
Sabit Hassan, Katherine Atwell, Malihe Alikhani
CogSci3
2022 Learning cognitive and linguistic prosodic categories for automatic cross-lingual sign language understanding
Mert Inan, Sabit Hassan, Lorna C. Quandt, Malihe Alikhani
CogSci5
2022 Look For Adjectives In the Face: How Facial Expressions Contribute To Meaning In Signed Languages
Carla Viegas, Lorna C. Quandt, Malihe Alikhani
CogSci3
2022 The Role of Context and Uncertainty in Shallow Discourse Parsing
abstract
Discourse parsing has proven to be useful for a number of NLP tasks that require complex reasoning. However, over a decade since the advent of the Penn Discourse Treebank, predicting implicit discourse relations in text remains challenging. There are several possible reasons for this, and we hypothesize that models should be exposed to more context as it plays an important role in accurate human annotation; meanwhile adding uncertainty measures can improve model accuracy and calibration. To thoroughly investigate this phenomenon, we perform a series of experiments to determine 1) the effects of context on human judgments, and 2) the effect of quantifying uncertainty with annotator confidence ratings on model accuracy and calibration (which we measure using the Brier score (Brier et al, 1950)). We find that including annotator accuracy and confidence improves model accuracy, and incorporating confidence in the model’s temperature function can lead to models with significantly better-calibrated confidence measures. We also find some insightful qualitative results regarding human and model behavior on these datasets.
Katherine Atwell, Remi Choi, Junyi Jessy Li, Malihe Alikhani
COLING4
2022 APPDIA: A Discourse-aware Transformer-based Style Transfer Model for Offensive Social Media Conversations
abstract
Using style-transfer models to reduce offensiveness of social media comments can help foster a more inclusive environment. However, there are no sizable datasets that contain offensive texts and their inoffensive counterparts, and fine-tuning pretrained models with limited labeled data can lead to the loss of original meaning in the style-transferred text. To address this issue, we provide two major contributions. First, we release the first publicly-available, parallel corpus of offensive Reddit comments and their style-transferred counterparts annotated by expert sociolinguists. Then, we introduce the first discourse-aware style-transfer models that can effectively reduce offensiveness in Reddit text while preserving the meaning of the original text. These models are the first to examine inferential links between the comment and the text it is replying to when transferring the style of offensive Reddit text. We propose two different methods of integrating discourse relations with pretrained transformer models and evaluate them on our dataset of offensive comments from Reddit and their inoffensive counterparts. Improvements over the baseline with respect to both automatic metrics and human evaluation indicate that our discourse-aware models are better at preserving meaning in style-transferred text when compared to the state-of-the-art discourse-agnostic models.
Katherine Atwell, Sabit Hassan, Malihe Alikhani
COLING3
2022 PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personification Data for Learning Enhanced Generation
abstract
A personification is a figure of speech that endows inanimate entities with properties and actions typically seen as requiring animacy. In this paper, we explore the task of personification generation. To this end, we propose PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personification data for Learning Enhanced generation. We curate a corpus of personifications called PersonifCorp, together with automatically generated de-personified literalizations of these personifications. We demonstrate the usefulness of this parallel corpus by training a seq2seq model to personify a given literal input. Both automatic and human evaluations show that fine-tuning with PersonifCorp leads to significant gains in personification-related qualities such as animacy and interestingness. A detailed qualitative analysis also highlights key strengths and imperfections of PINEAPPLE over baselines, demonstrating a strong ability to generate diverse and creative personifications that enhance the overall appeal of a sentence.
Sedrick Keh, Varun Gangal, Steven Y. Feng, Harsh Jhamtani, Malihe Alikhani, Eduard H. Hovy
COLING6
2022 Including Signed Languages in Natural Language Processing (Extended Abstract)
abstract
Signed languages are the primary means of communication for many deaf and hard of hearing individuals. Since signed languages exhibit all the fundamental linguistic properties of natural language, we believe that tools and theories of Natural Language Processing (NLP) are crucial towards its modeling. However, existing research in Sign Language Processing (SLP) seldom attempt to explore and leverage the linguistic organization of signed languages. This position paper calls on the NLP community to include signed languages as a research area with high social and scientific impact. We first discuss the linguistic properties of signed languages to consider during their modeling. Then, we review the limitations of current SLP models and identify the open challenges to extend NLP to signed languages. Finally, we urge (1) the adoption of an efficient tokenization method; (2) the development of linguistically-informed models; (3) the collection of real-world signed language data; (4) the inclusion of local signed language communities as an active and leading voice in research.
Kayo Yin, Malihe Alikhani
IJCAI2
2022 Political Ideology and Polarization: A Multi-dimensional Approach
abstract
Barea Sinno, Bernardo Oviedo, Katherine Atwell, Malihe Alikhani, Junyi Jessy Li. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Barea Sinno, Bernardo Oviedo, Katherine Atwell, Malihe Alikhani, Junyi Jessy Li
NAACL-HLT4
2022 PAC-Bayesian domain adaptation bounds for multiclass learners
abstract
Multiclass neural networks are a common tool in modern unsupervised domain adaptation, yet an appropriate theoretical description for their non-uniform sample complexity is lacking in the adaptation literature. To fill this gap, we propose the first PAC-Bayesian adaptation bounds for multiclass learners. We facilitate practical use of our bounds by also proposing the first approximation techniques for the multiclass distribution divergences we consider. For divergences dependent on a Gibbs predictor, we propose additional PAC-Bayesian adaptation bounds which remove the need for inefficient Monte-Carlo estimation. Empirically, we test the efficacy of our proposed approximation techniques as well as some novel design-concepts which we include in our bounds. Finally, we apply our bounds to analyze a common adaptation algorithm that uses neural networks.
Anthony Sicilia, Katherine Atwell, Malihe Alikhani, Seong Jae Hwang
UAI3
2022 How to Ask for Donations? Learning User-Specific Persuasive Dialogue Policies through Online Interactions
abstract
Persuasive conversations are more effective when they are custom-tailored for the intended audience. Current persuasive dialogue systems rely heavily on advice-giving or focus on different framing policies in a constrained and less dynamic/flexible manner. In this paper, we argue for a new approach, in which the system can identify optimal persuasive strategies in context and persuade users through online interactions. We study two main questions (1) can a reinforcement-learning-based dialogue framework learn to exercise user-specific communicative strategies for persuading users? (2) How can we leverage the crowd-sourcing platforms to collect data for training, and evaluating such frameworks for human-AI(/machine) conversations? We describe a prototype system that interacts with users with the goal of persuading them to donate to a charity and use experiments with crowd workers and analyses of our learned policies to document that our approach leads to learning context-sensitive persuasive strategies that focus on user’s reactions towards donation and contribute to increasing dialogue success.
Nhat Tran, Malihe Alikhani, Diane J. Litman
UMAP2
2022 Modeling Non-Cooperative Dialogue: Theoretical and Empirical Insights
abstract
Abstract Investigating cooperativity of interlocutors is central in studying pragmatics of dialogue. Models of conversation that only assume cooperative agents fail to explain the dynamics of strategic conversations. Thus, we investigate the ability of agents to identify non-cooperative interlocutors while completing a concurrent visual-dialogue task. Within this novel setting, we study the optimality of communication strategies for achieving this multi-task objective. We use the tools of learning theory to develop a theoretical model for identifying non-cooperative interlocutors and apply this theory to analyze different communication strategies. We also introduce a corpus of non-cooperative conversations about images in the GuessWhat?! dataset proposed by De Vries et al. (2017). We use reinforcement learning to implement multiple communication strategies in this context and find that empirical results validate our theory.
Anthony Sicilia, Tristan Maidment, Pat Healy, Malihe Alikhani
Trans. Assoc. Comput. Linguistics4
2021 Including Signed Languages in Natural Language Processing
abstract
Kayo Yin, Amit Moryossef, Julie Hochgesang, Yoav Goldberg, Malihe Alikhani. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Kayo Yin, Amit Moryossef, Julie Hochgesang, Yoav Goldberg, Malihe Alikhani
ACL/IJCNLP (1)5
2021 Signed Coreference Resolution
abstract
Coreference resolution is key to many natural language processing tasks and yet has been relatively unexplored in Sign Language Processing.In signed languages, space is primarily used to establish reference.Solving coreference resolution for signed languages would not only enable higher-level Sign Language Processing systems, but also enhance our understanding of language in different modalities and of situated references, which are key problems in studying grounded language.In this paper, we: (1) introduce Signed Coreference Resolution (SCR), a new challenge for coreference modeling and Sign Language Processing; (2) collect an annotated corpus of German Sign Language with gold labels for coreference together with an annotation software for the task; (3) explore features of hand gesture, iconicity, and spatial situated properties and move forward to propose a set of linguistically informed heuristics and unsupervised models for the task; (4) put forward several proposals about ways to address the complexities of this challenge effectively 1 .
Kayo Yin, Kenneth DeHaan, Malihe Alikhani
EMNLP (1)3
2021 Examining Covert Gender Bias: A Case Study in Turkish and English Machine Translation Models
abstract
As Machine Translation (MT) has become increasingly more powerful, accessible, and widespread, the potential for the perpetuation of bias has grown alongside its advances.While overt indicators of bias have been studied in machine translation, we argue that covert biases expose a problem that is further entrenched.Through the use of the genderneutral language Turkish and the gendered language English, we examine cases of both overt and covert gender bias in MT models.Specifically, we introduce a method to investigate asymmetrical gender markings.We also assess bias in the attribution of personhood and examine occupational and personality stereotypes through overt bias indicators in MT models.Our work explores a deeper layer of bias in MT models and demonstrates the continued need for language-specific, interdisciplinary methodology in MT model development.
Chloe Ciora, Nur Iren, Malihe Alikhani
INLG3
2021 Towards Designing Enthusiastic AI Agents
abstract
Immersive virtual worlds are increasingly being used for education, training, and entertainment, and virtual humans that can interact with human users in these worlds play many important roles. Understating the emotional constructs of the user and generating multimodal forms of communications that are aligned with the user's needs and input is key to designing AI agents. Most virtual agents and communicative systems lack the ability to understand enthusiasm or generate multimodal enthusiastic communicative presentations. In this work, we argue for the importance of including enthusiasm in the design of human-AI collaboration and communication and review the existing datasets and models that can be used to bridge the gap in this area.
Carla Viegas, Malihe Alikhani
IVA2
2021 Where Are We in Discourse Relation Recognition?
abstract
Discourse parsers recognize the intentional and inferential relationships that organize extended texts.They have had a great influence on a variety of NLP tasks as well as theoretical studies in linguistics and cognitive science.However it is often difficult to achieve good results from current discourse models, largely due to the difficulty of the task, particularly recognizing implicit discourse relations.Recent developments in transformer-based models have shown great promise on these analyses, but challenges still remain.We present a position paper which provides a systematic analysis of the state of the art discourse parsers.We aim to examine the performance of current discourse parsing models via gradual domain shift: within the same corpus, on in-domain texts, and on out-of-domain texts, and discuss the differences between the transformer-based models and the previous models in predicting different types of implicit relations both interand intra-sentential.We conclude by describing several shortcomings of the existing models and a discussion of how future work should approach this problem.
Katherine Atwell, Junyi Jessy Li, Malihe Alikhani
SIGDIAL3
2021 ParsiNLU: A Suite of Language Understanding Challenges for Persian
abstract
Abstract Despite the progress made in recent years in addressing natural language understanding (NLU) challenges, the majority of this progress remains to be concentrated on resource-rich languages like English. This work focuses on Persian language, one of the widely spoken languages in the world, and yet there are few NLU datasets available for this language. The availability of high-quality evaluation datasets is a necessity for reliable assessment of the progress on different NLU tasks and domains. We introduce ParsiNLU, the first benchmark in Persian language that includes a range of language understanding tasks—reading comprehension, textual entailment, and so on. These datasets are collected in a multitude of ways, often involving manual annotations by native speakers. This results in over 14.5k new instances across 6 distinct NLU tasks. Additionally, we present the first results on state-of-the-art monolingual and multilingual pre-trained language models on this benchmark and compare them with human performance, which provides valuable insights into our ability to tackle natural language understanding challenges in Persian. We hope ParsiNLU fosters further research and advances in Persian language understanding.1
Daniel Khashabi, Arman Cohan, Siamak Shakeri, Pedram Hosseini, Pouya Pezeshkpour, Malihe Alikhani, Moin Aminnaseri, Marzieh Bitaab, Faeze Brahman, Sarik Ghazarian, Mozhdeh Gheini, Arman Kabiri, Rabeeh Karimi Mahabadi, Omid Memarrast, Ahmadreza Mosallanezhad, Erfan Noury, Shahab Raji, Mohammad Sadegh Rasooli, Sepideh Sadeghi, Erfan Sadeqi Azer, Niloofar Safi Samghabadi, Mahsa Shafaei, Saber Sheybani, Ali Tazarv, Yadollah Yaghoobzadeh
Trans. Assoc. Comput. Linguistics6
2020 That and There: Judging the Intent of Pointing Actions with Robotic Arms
abstract
Collaborative robotics requires effective communication between a robot and a human partner. This work proposes a set of interpretive principles for how a robotic arm can use pointing actions to communicate task information to people by extending existing models from the related literature. These principles are evaluated through studies where English-speaking human subjects view animations of simulated robots instructing pick-and-place tasks. The evaluation distinguishes two classes of pointing actions that arise in pick-and-place tasks: referential pointing (identifying objects) and locating pointing (identifying locations). The study indicates that human subjects show greater flexibility in interpreting the intent of referential pointing compared to locating pointing, which needs to be more deliberate. The results also demonstrate the effects of variation in the environment and task context on the interpretation of pointing. Our corpus, experiments and design principles advance models of context, common sense reasoning and communication in embodied communication.
Malihe Alikhani, Baber Khalid, Rahul Shome, Chaitanya Mitash, Kostas E. Bekris, Matthew Stone
AAAI1
2020 Cross-modal Coherence Modeling for Caption Generation
abstract
We use coherence relations inspired by computational models of discourse to study the information needs and goals of image captioning.Using an annotation protocol specifically devised for capturing image-caption coherence relations, we annotate 10,000 instances from publicly-available image-caption pairs.We introduce a new task for learning inferences in imagery and text, coherence relation prediction, and show that these coherence annotations can be exploited to learn relation classifiers as an intermediary step, and also train coherence-aware, controllable image captioning models.The results show a dramatic improvement in the consistency and quality of the generated captions with respect to information needs specified via coherence relations.
Malihe Alikhani, Piyush Sharma, Shengjie Li 0002, Radu Soricut, Matthew Stone
ACL1
2020 Combining Cognitive Modeling and Reinforcement Learning for Clarification in Dialogue
abstract
In many domains, dialogue systems need to work collaboratively with users to successfully reconstruct the meaning the user had in mind.In this paper, we show how cognitive models of users' communicative strategies can be leveraged in a reinforcement learning approach to dialogue planning to enable interactive systems to give targeted, effective feedback about the system's understanding.We describe a prototype system that collaborates on reference tasks that distinguish arbitrarily varying color patches from similar distractors, and use experiments with crowd workers and analyses of our learned policies to document that our approach leads to context-sensitive clarification strategies that focus on key missing information, elicit correct answers that the system understands, and contribute to increasing dialogue success.
Baber Khalid, Malihe Alikhani, Matthew Stone
COLING2
2020 Aspectuality Across Genre: A Distributional Semantics Approach
abstract
The interpretation of the lexical aspect of verbs in English plays a crucial role for recognizing textual entailment and learning discourse-level inferences.We show that two elementary dimensions of aspectual class, states vs. events, and telic vs. atelic events, can be modelled effectively with distributional semantics.We find that a verb's local context is most indicative of its aspectual class, and demonstrate that closed class words tend to be stronger discriminating contexts than content words.Our approach outperforms previous work on three datasets.Lastly, we contribute a dataset of human-human conversations annotated with lexical aspect and present experiments that show the correlation of telicity with genre and discourse goals.
Thomas Kober 0001, Malihe Alikhani, Matthew Stone, Mark Steedman
COLING2
2018 Arrows are the Verbs of Diagrams
abstract
Arrows are a key ingredient of schematic pictorial communication. This paper investigates the interpretation of arrows through linguistic, crowdsourcing and machine-learning methodology. Our work establishes a novel analogy between arrows and verbs: we advocate representing arrows in terms of qualitatively different structural and semantic frames, and resolving frames to specific interpretations using shallow world knowledge.
Malihe Alikhani, Matthew Stone
COLING1
2017 When is Likely Unlikely: Investigating the Variability of Vagueness
Kimele Persaud, Brian McMahan, Malihe Alikhani, Kevin Pei, Pernille Hemmer, Matthew Stone
CogSci3