Mona T. Diab

dblp:15/4305 · also Mona Talat Diab · DBLP profile ↗
← Back
85ranked-venue papers
12as first author
20since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 85 · 12 first-author · 20 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse Autoencoders
abstract
Sparse Autoencoders (SAEs) are a prominent tool in mechanistic interpretability (MI) for decomposing neural network activations into interpretable features. However, the aspiration to identify a canonical set of features is challenged by the observed inconsistency of learned SAE features across different training runs, undermining reproducibility and complicating model comparison. We study run-to-run feature consistency in SAEs and argue that it should be reported as a standard evaluation axis alongside reconstruction and sparsity. We propose the Pairwise Dictionary Mean Correlation Coefficient (PW-MCC) as an assignment-based metric to quantify consistency and demonstrate that high levels are achievable (PW-MCC ≈ 0.80 for TopK SAEs on LLM activations) with appropriate architectural choices.Our contributions include: (i) theoretical grounding for strong consistency in the idealized setting of TopK SAEs; (ii) synthetic validation using a model organism, which verifies PW-MCC as a reliable proxy for ground-truth recovery; and (iii) empirical analysis on LLM activations, where PW-MCC correlates with the similarity of automatically generated natural-language feature explanations.
Xiangchen Song, Aashiq Muhamed, Yujia Zheng 0001, Zeyu Tang 0002, Mona T. Diab, Virginia Smith, Kun Zhang 0001
ACL (1)6
2025 BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data
abstract
In this work, we tackle the challenge of embedding realistic human personality traits into LLMs.Previous approaches have primarily focused on prompt-based methods that describe the behavior associated with the desired personality traits, suffering from realism and validity issues.To address these limitations, we introduce BIG5-CHAT, a large-scale dataset containing 100,000 dialogues designed to ground models in how humans express their personality in language.Leveraging this dataset, we explore Supervised Fine-Tuning and Direct Preference Optimization as training-based methods to align LLMs more naturally with human personality patterns.Our methods outperform prompting on personality assessments such as BFI and IPIP-NEO, with trait correlations more closely matching human data.Furthermore, our experiments reveal that models trained to exhibit higher conscientiousness, higher agreeableness, lower extraversion, and lower neuroticism display better performance on reasoning tasks, aligning with psychological findings on how these traits impact human cognitive performance.To our knowledge, this work is the first comprehensive study to demonstrate how training-based methods can shape LLM personalities through learning from real human behaviors. Whenever I lay on my bed I get so tired.
Jiarui Liu 0004, Andy Liu, Mona T. Diab, Maarten Sap
ACL (1)5
2025 Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics
abstract
Jiarui Liu, Yueqi Song, Yunze Xiao, Mingqian Zheng, Lindia Tjuatja, Jana Schaich Borg, Mona T. Diab, Maarten Sap. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jiarui Liu 0004, Yueqi Song, Yunze Xiao, Mingqian Zheng, Lindia Tjuatja, Jana Schaich Borg, Mona T. Diab, Maarten Sap
EMNLP7
2025 Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design
abstract
Large Language Models (LLMs) increasingly exhibit anthropomorphism characteristicshuman-like qualities portrayed across their outlook, language, behavior, and reasoning functions.Such characteristics enable more intuitive and engaging human-AI interactions.However, current research on anthropomorphism remains predominantly risk-focused, emphasizing over-trust and user deception while offering limited design guidance.We argue that anthropomorphism should instead be treated as a concept of design that can be intentionally tuned to support user goals.Drawing from multiple disciplines, we propose that the anthropomorphism of an LLM-based artifact should reflect the interaction between artifact designers and interpreters.This interaction is facilitated by cues embedded in the artifact by the designers and the (cognitive) responses of the interpreters to the cues.Cues are categorized into four dimensions: perceptive, linguistic, behavioral, and cognitive.By analyzing the manifestation and effectiveness of each cue, we provide a unified taxonomy with actionable levers for practitioners.Consequently, we advocate for function-oriented evaluations of anthropomorphic design.
Yunze Xiao, Lynnette Hui Xian Ng, Jiarui Liu 0004, Mona T. Diab
EMNLP4
2025 Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders
abstract
Kshitish Ghate, Isaac Slaughter, Kyra Wilson, Mona T. Diab, Aylin Caliskan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kshitish Ghate, Isaac Slaughter, Kyra Wilson, Mona T. Diab, Aylin Caliskan
NAACL (Long Papers)4
2024 Investigating Cultural Alignment of Large Language Models
abstract
The intricate relationship between language and culture has long been a subject of exploration within the realm of linguistic anthropology.Large Language Models (LLMs), promoted as repositories of collective human knowledge, raise a pivotal question: do these models genuinely encapsulate the diverse knowledge adopted by different cultures?Our study reveals that these models demonstrate greater cultural alignment along two dimensions-firstly, when prompted with the dominant language of a specific culture, and secondly, when pretrained with a refined mixture of languages employed by that culture.We quantify cultural alignment by simulating sociological surveys, comparing model responses to those of actual survey participants as references.Specifically, we replicate a survey conducted in various regions of Egypt and the United States through prompting LLMs with different pretraining data mixtures in both Arabic and English with the personas of the real respondents and the survey questions.Further analysis reveals that misalignment becomes more pronounced for underrepresented personas and for culturally sensitive topics, such as those probing social values.Finally, we introduce Anthropological Prompting, a novel method leveraging anthropological reasoning to enhance cultural alignment.Our study emphasizes the necessity for a more balanced multilingual pretraining dataset to better represent the diversity of human experience and the plurality of different cultures with many implications on the topic of cross-lingual transfer.1
Badr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. Diab
ACL (1)4
2024 Depth-Wise Attention (DWAtt): A Layer Fusion Method for Data-Efficient Classification
abstract
Language Models pretrained on large textual data have been shown to encode different types of knowledge simultaneously. Traditionally, only the features from the last layer are used when adapting to new tasks or data. We put forward that, when using or finetuning deep pretrained models, intermediate layer features that may be relevant to the downstream task are buried too deep to be used efficiently in terms of needed samples or steps. To test this, we propose a new layer fusion method: Depth-Wise Attention (DWAtt), to help re-surface signals from non-final layers. We compare DWAtt to a basic concatenation-based layer fusion method (Concat), and compare both to a deeper model baseline—all kept within a similar parameter budget. Our findings show that DWAtt and Concat are more step- and sample-efficient than the baseline, especially in the few-shot setting. DWAtt outperforms Concat on larger data sizes. On CoNLL-03 NER, layer fusion shows 3.68 − 9.73% F1 gain at different few-shot sizes. The layer fusion models presented significantly outperform the baseline in various training scenarios with different data sizes, architectures, and training constraints.
Muhammad N. ElNokrashy, Badr AlKhamissi, Mona T. Diab
LREC/COLING3
2024 Towards a Responsible Thinking in the New Era of Gen AI: Walking the Walk
Mona T. Diab
DATA1
2024 GRASS: Compute Efficient Low-Memory LLM Training with Structured Sparse Gradients
abstract
Large language model (LLM) training and finetuning are often bottlenecked by limited GPU memory.While existing projection-based optimization methods address this by projecting gradients into a lower-dimensional subspace to reduce optimizer state memory, they typically rely on dense projection matrices, which can introduce computational and memory overheads.In this work, we propose GRASS (GRAdient Stuctured Sparsification), a novel approach that leverages sparse projections to transform gradients into structured sparse updates.This design not only significantly reduces memory usage for optimizer states but also minimizes gradient memory footprint, computation, and communication costs, leading to substantial throughput improvements.Extensive experiments on pretraining and finetuning tasks demonstrate that GRASS achieves competitive performance to full-rank training and existing projection-based methods.Notably, GRASS enables half-precision pretraining of a 13B parameter LLaMA model on a single 40GB A100 GPU-a feat infeasible for previous methodsand yields up to a 2× throughput improvement on an 8-GPU system.Code is released here 1 .P ← compute P (∇L(W (t) )) ▷ P ∈ R m×r 8: // [Optional] Update optimizer state 9:
Aashiq Muhamed, Oscar Li, David P. Woodruff, Mona T. Diab, Virginia Smith
EMNLP4
2024 Can Large Language Models Infer Causation from Correlation?
abstract
Causal inference is one of the hallmarks of human intelligence. While the field of CausalNLP has attracted much interest in the recent years, existing causal inference datasets in NLP primarily rely on discovering causality from empirical knowledge (e.g., commonsense knowledge). In this work, we propose the first benchmark dataset to test the pure causal inference skills of large language models (LLMs). Specifically, we formulate a novel task Corr2Cause, which takes a set of correlational statements and determines the causal relationship between the variables. We curate a large-scale dataset of more than 200K samples, on which we evaluate seventeen existing LLMs. Through our experiments, we identify a key shortcoming of LLMs in terms of their causal inference skills, and show that these models achieve almost close to random performance on the task. This shortcoming is somewhat mitigated when we try to re-purpose LLMs for this skill via finetuning, but we find that these models still fail to generalize – they can only perform causal inference in in-distribution settings when variable names and textual expressions used in the queries are similar to those in the training set, but fail in out-of-distribution settings generated by perturbing these queries. Corr2Cause is a challenging task for LLMs, and can be helpful in guiding future research on improving LLMs’ pure reasoning skills and generalizability. Our data is at https://huggingface.co/datasets/causalnlp/corr2cause. Our code is at https://github.com/causalNLP/corr2cause.
Zhijing Jin 0001, Jiarui Liu 0004, Zhiheng Lyu, Spencer Poff, Mrinmaya Sachan, Rada Mihalcea, Mona T. Diab, Bernhard Schölkopf
ICLR7
2024 Automatic Generation of Model and Data Cards: A Step Towards Responsible AI
abstract
Jiarui Liu, Wenkai Li, Zhijing Jin, Mona Diab. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jiarui Liu 0004, Zhijing Jin 0001, Mona T. Diab
NAACL-HLT4
2024 Analyzing the Role of Semantic Representations in the Era of Large Language Models
abstract
Zhijing Jin, Yuen Chen, Fernando Gonzalez Adauto, Jiarui Liu, Jiayi Zhang, Julian Michael, Bernhard Schölkopf, Mona Diab. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhijing Jin 0001, Yuen Chen, Fernando Gonzalez Adauto, Jiarui Liu 0004, Julian Michael, Bernhard Schölkopf, Mona T. Diab
NAACL-HLT8
2023 ALERT: Adapt Language Models to Reasoning Tasks
abstract
Ping Yu, Tianlu Wang, Olga Golovneva, Badr AlKhamissi, Siddharth Verma, Zhijing Jin, Gargi Ghosh, Mona Diab, Asli Celikyilmaz. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Olga Golovneva, Badr AlKhamissi, Siddharth Verma, Zhijing Jin 0001, Gargi Ghosh, Mona T. Diab, Asli Celikyilmaz
ACL (1)8
2023 Methods for Measuring, Updating, and Visualizing Factual Beliefs in Language Models
abstract
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Peter Hase, Mona T. Diab, Asli Celikyilmaz, Xian Li 0003, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer 0001
EACL2
2022 ToKen: Task Decomposition and Knowledge Infusion for Few-Shot Hate Speech Detection
abstract
Badr AlKhamissi, Faisal Ladhak, Srinivasan Iyer, Veselin Stoyanov, Zornitsa Kozareva, Xian Li, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona Diab. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Badr AlKhamissi, Faisal Ladhak, Srinivasan Iyer 0001, Veselin Stoyanov, Zornitsa Kozareva, Xian Li 0003, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona T. Diab
EMNLP10
2022 Efficient Large Scale Language Modeling with Mixtures of Experts
abstract
Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, Giridharan Anantharaman, Xian Li, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Xing Zhou, Punit Singh Koura, Brian O’Horo, Jeffrey Wang, Luke Zettlemoyer, Mona Diab, Zornitsa Kozareva, Veselin Stoyanov. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Mikel Artetxe, Shruti Bhosale, Naman Goyal 0001, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer 0001, Ramakanth Pasunuru, Giri Anantharaman, Xian Li 0003, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Punit Singh Koura, Brian O'Horo, Jeffrey Wang, Luke Zettlemoyer, Mona T. Diab, Zornitsa Kozareva, Veselin Stoyanov
EMNLP22
2022 Few-shot Learning with Multilingual Generative Language Models
abstract
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, Xian Li. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal 0001, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona T. Diab, Veselin Stoyanov, Xian Li 0003
EMNLP19
2022 BeSt: The Belief and Sentiment Corpus
abstract
We present the BeSt corpus, which records cognitive state: who believes what (i.e., factuality), and who has what sentiment towards what. This corpus is inspired by similar source-and-target corpora, specifically MPQA and FactBank. The corpus comprises two genres, newswire and discussion forums, in three languages, Chinese (Mandarin), English, and Spanish. The corpus is distributed through the LDC.
Jennifer Tracey, Owen Rambow, Claire Cardie, Adam Dalton 0001, Hoa Trang Dang, Mona T. Diab, Bonnie J. Dorr, Louise Guthrie, Magdalena Markowska, Smaranda Muresan, Vinodkumar Prabhakaran, Samira Shaikh, Tomek Strzalkowski
LREC6
2022 AnswerSumm: A Manually-Curated Dataset and Pipeline for Answer Summarization
abstract
Alexander Fabbri, Xiaojian Wu, Srini Iyer, Haoran Li, Mona Diab. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Alexander R. Fabbri, Xiaojian Wu, Srinivasan Iyer 0001, Haoran Li 0007, Mona T. Diab
NAACL-HLT5
2021 Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel Data
abstract
Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona Diab. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal 0001, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona T. Diab
ACL/IJCNLP (1)9
2020 A Multitask Learning Approach for Diacritic Restoration
abstract
In many languages like Arabic, diacritics are used to specify pronunciations as well as meanings. Such diacritics are often omitted in written text, increasing the number of possible pronunciations and meanings for a word. This results in a more ambiguous text making computational processing on such text more difficult. Diacritic restoration is the task of restoring missing diacritics in the written text. Most state-of-the-art diacritic restoration models are built on character level information which helps generalize the model to unseen data, but presumably lose useful information at the word level. Thus, to compensate for this loss, we investigate the use of multi-task learning to jointly optimize diacritic restoration with related NLP problems namely word segmentation, part-of-speech tagging, and syntactic diacritization. We use Arabic as a case study since it has sufficient data resources for tasks that we consider in our joint modeling. Our joint models significantly outperform the baselines and are comparable to the state-of-the-art models that are more complex relying on morphological analyzers and/or a lot more data (e.g. dialectal data).
Sawsan Alqahtani, Ajay Mishra, Mona T. Diab
ACL3
2020 FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive Summarization
abstract
Neural abstractive summarization models are prone to generate content inconsistent with the source document, i.e. unfaithful.Existing automatic metrics do not capture such mistakes effectively.We tackle the problem of evaluating faithfulness of a generated summary given its source document.We first collected human annotations of faithfulness for outputs from numerous models on two datasets.We find that current models exhibit a trade-off between abstractiveness and faithfulness: outputs with less word overlap with the source document are more likely to be unfaithful.Next, we propose an automatic question answering (QA) based metric for faithfulness, FEQA, 1 which leverages recent advances in reading comprehension.Given questionanswer pairs generated from the summary, a QA model extracts answers from the document; non-matched answers indicate unfaithful information in the summary.Among metrics based on word overlap, embedding similarity, and learned language understanding models, our QA-based metric has significantly higher correlation with human faithfulness scores, especially on highly abstractive summaries.* Most of the work is done while the authors were at Amazon Web Services AI.1 Faithfulness Evaluation with Question Answering.
Esin Durmus, He He 0001, Mona T. Diab
ACL3
2020 DeSePtion: Dual Sequence Prediction and Adversarial Examples for Improved Fact-Checking
abstract
Christopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia, Kriste Krstovski, Mona Diab, Smaranda Muresan. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Christopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia, Kriste Krstovski, Mona T. Diab, Smaranda Muresan
ACL6
2020 Detecting Urgency Status of Crisis Tweets: A Transfer Learning Approach for Low Resource Languages
abstract
We release an urgency dataset that consists of English tweets relating to natural crises.The set is annotated along with annotations of their corresponding urgency status.Additionally, we release evaluation datasets for two low-resource languages, i.e.Sinhala and Odia, and demonstrate an effective zero-shot transfer from English to these two languages by training cross-lingual classifiers.We adopt cross-lingual embeddings constructed using different methods to extract features of the tweets, including a few state-of-the-art contextual embeddings such as BERT, RoBERTa and XLM-R.We train a variety of classifier architectures, supervised and semi supervised, on the extracted features.We also further experiment with ensembling the various classifiers.With very limited amounts of labeled data in English and zero data in the low resource languages, we show a successful framework of training monolingual and cross-lingual classifiers using deep learning methods which are known to be data hungry.Specifically, we show that the recent deep contextual embeddings are also helpful when dealing with very small-scale datasets.Classifiers that incorporate RoBERTa yield the best performance for the English urgency detection task, with 25% F1 score absolute improvement over the baselines.For the zero-shot transfer to low resource languages, classifiers that use LASER features perform the best for Sinhala transfer while XLM-R features benefit the Odia transfer the most.
Efsun Sarioglu Kayi, Linyong Nan, Bohan Qu, Mona T. Diab, Kathy McKeown
COLING4
2020 Multitask Learning for Cross-Lingual Transfer of Broad-coverage Semantic Dependencies
abstract
We describe a method for developing broad-coverage semantic dependency parsers for languages for which no semantically annotated resource is available. We leverage a multitask learning framework coupled with annotation projection. We use syntactic parsing as the auxiliary task in our multitask setup. Our annotation projection experiments from English to Czech show that our multitask setup yields 3.1% (4.2%) improvement in labeled F1-score on in-domain (out-of-domain) test set compared to a single-task baseline.
Maryam Aminian, Mohammad Sadegh Rasooli, Mona T. Diab
EMNLP (1)3
2020 Data Paucity and Low Resource Scenarios: Challenges and Opportunities
abstract
In an era of unstructured data abundance, you would think that we have solved our data requirements for building robust systems for language processing. However, this is not the case if we think on a global scale with over 7000 languages where only a handful have digital resources. Systems at scale with good performance typically require annotated resources that cover the genres and domain divides. Moreover, the existence of a handful of resources in some languages is a reflection of the digital disparity in various societies leading to inadvertent biases in systems. In this talk I will show some solutions for low resource scenarios, both cross domain and genres as well as cross lingually.
Mona T. Diab
KDD1
2020 Diversity, Density, and Homogeneity: Quantitative Characteristic Metrics for Text Collections
abstract
Summarizing data samples by quantitative measures has a long history, with descriptive statistics being a case in point. However, as natural language processing methods flourish, there are still insufficient characteristic metrics to describe a collection of texts in terms of the words, sentences, or paragraphs they comprise. In this work, we propose metrics of diversity, density, and homogeneity that quantitatively measure the dispersion, sparsity, and uniformity of a text collection. We conduct a series of simulations to verify that each metric holds desired properties and resonates with human intuitions. Experiments on real-world datasets demonstrate that the proposed characteristic metrics are highly correlated with text classification performance of a renowned model, BERT, which could inspire future applications.
Yi-An Lai, Xuan Zhu 0002, Mona T. Diab
LREC4
2019 Efficient Sentence Embedding using Discrete Cosine Transform
abstract
Nada Almarwani, Hanan Aldarmaki, Mona Diab. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Nada AlMarwani, Hanan Aldarmaki, Mona T. Diab
EMNLP/IJCNLP (1)3
2019 Efficient Convolutional Neural Networks for Diacritic Restoration
abstract
Sawsan Alqahtani, Ajay Mishra, Mona Diab. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Sawsan Alqahtani, Ajay Mishra, Mona T. Diab
EMNLP/IJCNLP (1)3
2019 CASA-NLU: Context-Aware Self-Attentive Natural Language Understanding for Task-Oriented Chatbots
abstract
Arshit Gupta, Peng Zhang, Garima Lalwani, Mona Diab. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Arshit Gupta, Garima Lalwani, Mona T. Diab
EMNLP/IJCNLP (1)4
2019 Multi-Domain Goal-Oriented Dialogues (MultiDoGO): Strategies toward Curating and Annotating Large Scale Dialogue Data
abstract
Denis Peskov, Nancy Clarke, Jason Krone, Brigi Fodor, Yi Zhang, Adel Youssef, Mona Diab. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Denis Peskov, Nancy E. Clarke, Jason Krone, Brigi Fodor, Adel Youssef, Mona T. Diab
EMNLP/IJCNLP (1)7
2019 Understanding Cohesion in Writings and Speech of Schizophrenia Patients
abstract
Schizophrenia is one of the mental disorders that impacts a person's thinking, speech, and actions. It can reduce a person's ability to process auditory information and make decisions. Analyzing this disorder correctly is important because it might help with different ways of reducing its negative effects on its patients. Linguists and psychiatrists have been investigating language impairments and speech disorder in people with schizophrenia disorder which can be challenging. In this study, we attempt to address this issue by analyzing linguistic features i.e. cohesion in the writings and speech scripts of schizophrenia patients. Our results show that using referential cohesion with text easability or situation model features provides the best performance for speech whereas for writing dataset, readability or a combination of situation model and readability yield the best performance.
Amal AlQahtani, Efsun Sarioglu Kayi, Mona T. Diab
ICMLA3
2019 Investigating Input and Output Units in Diacritic Restoration
abstract
Diacritic restoration is the task of assigning diacritics (accents) for each character in a given segment. The typical input levels that have been previously used in diacritic restoration models are word and/or character units. In this paper, we investigate the use of subwords as input units along with their diacritic patterns (combinations of adjacent diacritics) as output, as an alternative to word or character-based models. Our experiments show that characters provide the optimal level of information for sequence-based diacritic restoration models across different languages. We additionally improved our diacritic restoration model by maximizing over the output diacritic sequence using a Conditional Random Field (CRF). Adding a CRF layer improves the performance on observed and unobserved words substantially for Arabic and marginally for Yoruba.
Sawsan Alqahtani, Mona T. Diab
ICMLA2
2018 Evaluation of Unsupervised Compositional Representations
abstract
We evaluated various compositional models, from bag-of-words representations to compositional RNN-based models, on several extrinsic supervised and unsupervised evaluation benchmarks. Our results confirm that weighted vector averaging can outperform context-sensitive models in most benchmarks, but structural features encoded in RNN models can also be useful in certain classification tasks. We analyzed some of the evaluation datasets to identify the aspects of meaning they measure and the characteristics of the various models that explain their performance variance.
Hanan Aldarmaki, Mona T. Diab
COLING2
2018 Emotion Detection and Classification in a Multigenre Corpus with Joint Multi-Task Deep Learning
abstract
Detection and classification of emotion categories expressed by a sentence is a challenging task due to subjectivity of emotion. To date, most of the models are trained and evaluated on single genre and when used to predict emotion in different genre their performance drops by a large margin. To address the issue of robustness, we model the problem within a joint multi-task learning framework. We train this model with a multigenre emotion corpus to predict emotions across various genre. Each genre is represented as a separate task, we use soft parameter shared layers across the various tasks. our experimental results show that this model improves the results across the various genres, compared to a single genre training in the same neural net architecture.
Shabnam Tafreshi, Mona T. Diab
COLING2
2018 WASA: A Web Application for Sequence Annotation
Fahad Alghamdi 0003, Mona T. Diab
LREC2
2018 Sentence and Clause Level Emotion Annotation, Detection, and Classification in a Multi-Genre Corpus
Shabnam Tafreshi, Mona T. Diab
LREC2
2018 Unsupervised Word Mapping Using Structural Similarities in Monolingual Embeddings
abstract
Most existing methods for automatic bilingual dictionary induction rely on prior alignments between the source and target languages, such as parallel corpora or seed dictionaries. For many language pairs, such supervised alignments are not readily available. We propose an unsupervised approach for learning a bilingual dictionary for a pair of languages given their independently-learned monolingual word embeddings. The proposed method exploits local and global structures in monolingual vector spaces to align them such that similar words are mapped to each other. We show empirically that the performance of bilingual correspondents that are learned using our proposed unsupervised method is comparable to that of using supervised bilingual correspondents from a seed dictionary.
Hanan Aldarmaki, Mahesh Mohan, Mona T. Diab
Trans. Assoc. Comput. Linguistics3
2016 LILI: A Simple Language Independent Approach for Language Identification
abstract
We introduce a generic Language Independent Framework for Linguistic Code Switch Point Detection. The system uses characters level 5-grams and word level unigram language models to train a conditional random fields (CRF) model for classifying input words into various languages. We test our proposed framework and compare it to the state-of-the-art published systems on standard data sets from several language pairs: English-Spanish, Nepali-English, English-Hindi, Arabizi (Refers to Arabic written using the Latin/Roman script)-English, Arabic-Engari (Refers to English written using Arabic script), Modern Standard Arabic(MSA)-Egyptian, Levantine-MSA, Gulf-MSA, one more English-Spanish, and one more MSA-EGY. The overall weighted average F-score of each language pair are 96.4%, 97.3%, 98.0%, 97.0%, 98.9%, 86.3%, 88.2%, 90.6%, 95.2%, and 85.0% respectively. The results show that our approach despite its simplicity, either outperforms or performs at comparable levels to state-of-the-art published systems.
Mohamed Al-Badrashiny, Mona T. Diab
COLING2
2016 Computational Approaches to Linguistic Code Switching
Mona T. Diab, Pascale Fung, Julia Hirschberg, Thamar Solorio
INTERSPEECH1
2016 SPLIT: Smart Preprocessing (Quasi) Language Independent Tool
Mohamed Al-Badrashiny, Arfath Pasha, Mona T. Diab, Nizar Habash, Owen Rambow, Wael Salloum, Ramy Eskander
LREC3
2016 Creating a Large Multi-Layered Representational Repository of Linguistic Code Switched Arabic Data
Mona T. Diab, Mahmoud Ghoneim, Abdelati Hawwari, Fahad Alghamdi 0003, Nada AlMarwani, Mohamed Al-Badrashiny
LREC1
2016 Explicit Fine grained Syntactic and Semantic Annotation of the Idafa Construction in Arabic
Abdelati Hawwari, Mahmoud Ghoneim, Mona T. Diab
LREC4
2016 Guidelines and Framework for a Large Scale Arabic Diacritized Corpus
Wajdi Zaghouani, Houda Bouamor, Abdelati Hawwari, Mona T. Diab, Ossama Obeid, Mahmoud Ghoneim, Sawsan Alqahtani, Kemal Oflazer
LREC4
2015 Tharawat: A Vision for a Comprehensive Resource for Arabic Computational Processing
Mona T. Diab
CICLing (1)1
2015 AIDA2: A Hybrid Approach for Token and Sentence Level Dialect Identification in Arabic
abstract
In this paper, we present a hybrid approach for performing token and sentence levels Dialect Identification in Arabic.Specifically we try to identify whether each token in a given sentence belongs to Modern Standard Arabic (MSA), Egyptian Dialectal Arabic (EDA) or some other class and whether the whole sentence is mostly EDA or MSA.The token level component relies on a Conditional Random Field (CRF) classifier that uses decisions from several underlying components such as language models, a named entity recognizer and and a morphological analyzer to label each word in the sentence.The sentence level component uses a classifier ensemble system that relies on two independent underlying classifiers that model different aspects of the language.Using a featureselection heuristic, we select the best set of features for each of these two classifiers.We then train another classifier that uses the class labels and the confidence scores generated by each of the two underlying classifiers to decide upon the final class for each sentence.The token level component yields a new state of the art F-score of 90.6% (compared to previous state of the art of 86.8%) and the sentence level component yields an accuracy of 90.8% (compared to 86.6% obtained by the best state of the art system).
Mohamed Al-Badrashiny, Heba Elfardy, Mona T. Diab
CoNLL3
2014 Fast Tweet Retrieval with Compact Binary Codes
Weiwei Guo, Wei Liu 0005, Mona T. Diab
COLING3
2014 SANA: A Large Scale Multi-Genre, Multi-Dialect Lexicon for Arabic Subjectivity and Sentiment Analysis
Muhammad Abdul-Mageed, Mona T. Diab
LREC2
2014 Tharwa: A Large Scale Dialectal Arabic - Standard Arabic - English Lexicon
Mona T. Diab, Mohamed Al-Badrashiny, Maryam Aminian, Heba Elfardy, Nizar Habash, Abdelati Hawwari, Wael Salloum, Pradeep Dasigi, Ramy Eskander
LREC1
2014 MADAMIRA: A Fast, Comprehensive Tool for Morphological Analysis and Disambiguation of Arabic
Arfath Pasha, Mohamed Al-Badrashiny, Mona T. Diab, Ahmed El Kholy, Ramy Eskander, Nizar Habash, Manoj Pooleery, Owen Rambow, Ryan Roth
LREC3
2014 SAMAR: Subjectivity and sentiment analysis for Arabic social media
Muhammad Abdul-Mageed, Mona T. Diab, Sandra Kübler
Comput. Speech Lang.2
2013 Linking Tweets to News: A Framework to Enrich Short Text Data in Social Media
Weiwei Guo, Hao Li 0031, Heng Ji 0001, Mona T. Diab
ACL (1)4
2013 Multiword Expressions in the Context of Statistical Machine Translation
Mahmoud Ghoneim, Mona T. Diab
IJCNLP2
2013 DIRA: Dialectal Arabic Information Retrieval Assistant
Arfath Pasha, Mohamed Al-Badrashiny, Mohamed Altantawy, Nizar Habash, Manoj Pooleery, Owen Rambow, Ryan Roth, Mona T. Diab
IJCNLP8
2013 Improving Lexical Semantics for Sentential Semantics: Modeling Selectional Preference and Similar Words in a Latent Variable Model
Weiwei Guo, Mona T. Diab
HLT-NAACL2
2013 Code Switch Point Detection in Arabic
Heba Elfardy, Mohamed Al-Badrashiny, Mona T. Diab
NLDB3
2013 ANEAR: Automatic Named Entity Aliasing Resolution
Ayah Zirikly, Mona T. Diab
NLDB2
2012 Subgroup Detection in Ideological Discussions
Amjad Abu-Jbara, Pradeep Dasigi, Mona T. Diab, Dragomir R. Radev
ACL (1)3
2012 Modeling Sentences in the Latent Space
Weiwei Guo, Mona T. Diab
ACL (1)2
2012 Who's (Really) the Boss? Perception of Situational Power in Written Interactions
Vinodkumar Prabhakaran, Owen Rambow, Mona T. Diab
COLING3
2012 AWATIF: A Multi-Genre Corpus for Modern Standard Arabic Subjectivity and Sentiment Analysis
Muhammad Abdul-Mageed, Mona T. Diab
LREC2
2012 Simplified guidelines for the creation of Large Scale Dialectal Arabic Annotations
Heba Elfardy, Mona T. Diab
LREC2
2012 Conventional Orthography for Dialectal Arabic
Nizar Habash, Mona T. Diab, Owen Rambow
LREC2
2012 Annotations for Power Relations on Email Threads
Vinodkumar Prabhakaran, Huzaifa Neralwala, Owen Rambow, Mona T. Diab
LREC4
2012 Arabic Dialect Processing Tutorial
Mona T. Diab, Nizar Habash
HLT-NAACL1
2012 Predicting Overt Display of Power in Written Dialogs
Vinodkumar Prabhakaran, Owen Rambow, Mona T. Diab
HLT-NAACL3
2011 Semantic Topic Models: Combining Word Distributional Statistics and Dictionary Definitions
Weiwei Guo, Mona T. Diab
EMNLP2
2011 CODACT: Towards Identifying Orthographic Variants in Dialectal Arabic
Pradeep Dasigi, Mona T. Diab
IJCNLP2
2011 Introduction to the Special Issue on Arabic Computational Linguistics
abstract
introduction Share on Introduction to the Special Issue on Arabic Computational Linguistics Editors: Graham Katz Georgetown University Georgetown UniversityView Profile , Mona Diab Columbia University Columbia UniversityView Profile Authors Info & Claims ACM Transactions on Asian Language Information ProcessingVolume 10Issue 1Article No.: 1pp 1–4https://doi.org/10.1145/1929908.1929909Published:01 March 2011Publication History 0citation297DownloadsMetricsTotal Citations0Total Downloads297Last 12 Months8Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Graham Katz, Mona T. Diab
ACM Trans. Asian Lang. Inf. Process.2
2010 Combining Orthogonal Monolingual and Multilingual Sources of Evidence for All Words WSD
Weiwei Guo, Mona T. Diab
ACL2
2010 Task-based Evaluation of Multiword Expressions: a Pilot Study in Statistical Machine Translation
Marine Carpuat, Mona T. Diab
HLT-NAACL2
2009 Who, What, When, Where, Why? Comparing Multiple Approaches to the Cross-Lingual 5W Task
Kristen Parton, Kathy McKeown, Bob Coyne, Mona T. Diab, Ralph Grishman, Dilek Hakkani-Tür, Mary P. Harper, Heng Ji 0001, Wei-Yun Ma, Adam Meyers 0001, Sara Stolbach, Ang Sun, Gökhan Tür, Wei Xu 0004, Sibel Yaman
ACL/IJCNLP4
2009 Unsupervised Classification of Verb Noun Multi-Word Expression Tokens
Mona T. Diab, Madhav Krishna
CICLing1
2009 Arabic Named Entity Recognition: A Feature-Driven Study
abstract
The named entity recognition task aims at identifying and classifying named entities within an open-domain text. This task has been garnering significant attention recently as it has been shown to help improve the performance of many natural language processing applications. In this paper, we investigate the impact of using different sets of features in three discriminative machine learning frameworks, namely, support vector machines, maximum entropy and conditional random fields for the task of named entity recognition. Our language of interest is Arabic. We explore lexical, contextual and morphological features and nine data-sets of different genres and annotations. We measure the impact of the different features in isolation and incrementally combine them in order to evaluate the robustness to noise of each approach. We achieve the highest performance using a combination of 15 features in conditional random fields using broadcast news data (Fbeta=1=83.34).
Yassine Benajiba, Mona T. Diab, Paolo Rosso
IEEE Trans. Speech Audio Process.2
2008 Semantic Role Labeling Systems for Arabic using Kernel Methods
Mona T. Diab, Alessandro Moschitti, Daniele Pighin
ACL1
2008 Arabic Named Entity Recognition using Optimized Feature Sets
Yassine Benajiba, Mona T. Diab, Paolo Rosso
EMNLP2
2008 A Pilot Arabic Propbank
Martha Palmer, Olga Babko-Malaya, Ann Bies, Mona T. Diab, Mohamed Maamouri, Aous Mansouri, Wajdi Zaghouani
LREC4
2007 Arabic diacritization in the context of statistical machine translation
Mona T. Diab, Mahmoud Ghoneim, Nizar Habash
MTSummit1
2007 Semi-automatic error analysis for large-scale statistical machine translation
Katrin Kirchhoff, Owen Rambow, Nizar Habash, Mona T. Diab
MTSummit4
2006 Unsupervised Induction of Modern Standard Arabic Verb Classes Using Syntactic Frames and LSA
Neal Snider, Mona T. Diab
ACL2
2006 Parsing Arabic Dialects
David Chiang 0001, Mona T. Diab, Nizar Habash, Owen Rambow, Safiullah Shareef
EACL2
2006 Developing and Using a Pilot Dialectal Arabic Treebank
Mohamed Maamouri, Ann Bies, Tim Buckwalter, Mona T. Diab, Nizar Habash, Owen Rambow, Dalila Tabessi
LREC4
2006 Unsupervised Induction of Modern Standard Arabic Verb Classes
Neal Snider, Mona T. Diab
HLT-NAACL2
2004 Relieving the data Acquisition Bottleneck in Word Sense Disambiguation
abstract
Supervised learning methods for WSD yield better performance than unsupervised methods. Yet the availability of clean training data for the former is still a severe challenge. In this paper, we present an unsupervised bootstrapping approach for WSD which exploits huge amounts of automatically generated noisy data for training within a supervised learning framework. The method is evaluated using the 29 nouns in the English Lexical Sample task of SENSEVAL 2. Our algorithm does as well as supervised algorithms on 31% of this test set, which is an improvement of 11% (absolute) over state-of-the-art bootstrapping WSD algorithms. We identify seven different factors that impact the performance of our system.
Mona T. Diab
ACL1
2002 An Unsupervised Method for Word Sense Tagging using Parallel Corpora
abstract
We present an unsupervised method for word sense disambiguation that exploits translation correspondences in parallel corpora. The technique takes advantage of the fact that cross-language lexicalizations of the same concept tend to be consistent, preserving some core element of its semantics, and yet also variable, reflecting differing translator preferences and the influence of context. Working with parallel corpora introduces an extra complication for evaluation, since it is difficult to find a corpus that is both sense tagged and parallel with another language; therefore we use pseudo-translations, created by machine translation systems, in order to make possible the evaluation of the approach against a standard test set. The results demonstrate that word-level translation correspondences are a valuable source of information for sense disambiguation.
Mona T. Diab, Philip Resnik
ACL1