VLDB 2026 Research / reviewers in the wild / expert
Mona T. Diab
dblp:15/4305 · also Mona Talat Diab
· DBLP profile ↗
85ranked-venue papers
12as first author
20since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 85 · 12 first-author · 20 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse AutoencodersabstractSparse Autoencoders (SAEs) are a prominent tool in mechanistic interpretability (MI) for decomposing neural network activations into interpretable features. However, the aspiration to identify a canonical set of features is challenged by the observed inconsistency of learned SAE features across different training runs, undermining reproducibility and complicating model comparison. We study run-to-run feature consistency in SAEs and argue that it should be reported as a standard evaluation axis alongside reconstruction and sparsity. We propose the Pairwise Dictionary Mean Correlation Coefficient (PW-MCC) as an assignment-based metric to quantify consistency and demonstrate that high levels are achievable (PW-MCC ≈ 0.80 for TopK SAEs on LLM activations) with appropriate architectural choices.Our contributions include: (i) theoretical grounding for strong consistency in the idealized setting of TopK SAEs; (ii) synthetic validation using a model organism, which verifies PW-MCC as a reliable proxy for ground-truth recovery; and (iii) empirical analysis on LLM activations, where PW-MCC correlates with the similarity of automatically generated natural-language feature explanations. Xiangchen Song, Aashiq Muhamed, Yujia Zheng 0001, Zeyu Tang 0002, Mona T. Diab, Virginia Smith, Kun Zhang 0001 |
ACL (1) | 6 |
| 2025 | BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded DataabstractIn this work, we tackle the challenge of embedding realistic human personality traits into LLMs.Previous approaches have primarily focused on prompt-based methods that describe the behavior associated with the desired personality traits, suffering from realism and validity issues.To address these limitations, we introduce BIG5-CHAT, a large-scale dataset containing 100,000 dialogues designed to ground models in how humans express their personality in language.Leveraging this dataset, we explore Supervised Fine-Tuning and Direct Preference Optimization as training-based methods to align LLMs more naturally with human personality patterns.Our methods outperform prompting on personality assessments such as BFI and IPIP-NEO, with trait correlations more closely matching human data.Furthermore, our experiments reveal that models trained to exhibit higher conscientiousness, higher agreeableness, lower extraversion, and lower neuroticism display better performance on reasoning tasks, aligning with psychological findings on how these traits impact human cognitive performance.To our knowledge, this work is the first comprehensive study to demonstrate how training-based methods can shape LLM personalities through learning from real human behaviors. Whenever I lay on my bed I get so tired. Jiarui Liu 0004, Andy Liu, Mona T. Diab, Maarten Sap |
ACL (1) | 5 |
| 2025 | Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion DynamicsabstractJiarui Liu, Yueqi Song, Yunze Xiao, Mingqian Zheng, Lindia Tjuatja, Jana Schaich Borg, Mona T. Diab, Maarten Sap. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jiarui Liu 0004, Yueqi Song, Yunze Xiao, Mingqian Zheng, Lindia Tjuatja, Jana Schaich Borg, Mona T. Diab, Maarten Sap |
EMNLP | 7 |
| 2025 | Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of DesignabstractLarge Language Models (LLMs) increasingly exhibit anthropomorphism characteristicshuman-like qualities portrayed across their outlook, language, behavior, and reasoning functions.Such characteristics enable more intuitive and engaging human-AI interactions.However, current research on anthropomorphism remains predominantly risk-focused, emphasizing over-trust and user deception while offering limited design guidance.We argue that anthropomorphism should instead be treated as a concept of design that can be intentionally tuned to support user goals.Drawing from multiple disciplines, we propose that the anthropomorphism of an LLM-based artifact should reflect the interaction between artifact designers and interpreters.This interaction is facilitated by cues embedded in the artifact by the designers and the (cognitive) responses of the interpreters to the cues.Cues are categorized into four dimensions: perceptive, linguistic, behavioral, and cognitive.By analyzing the manifestation and effectiveness of each cue, we provide a unified taxonomy with actionable levers for practitioners.Consequently, we advocate for function-oriented evaluations of anthropomorphic design. Yunze Xiao, Lynnette Hui Xian Ng, Jiarui Liu 0004, Mona T. Diab |
EMNLP | 4 |
| 2025 | Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language EncodersabstractKshitish Ghate, Isaac Slaughter, Kyra Wilson, Mona T. Diab, Aylin Caliskan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kshitish Ghate, Isaac Slaughter, Kyra Wilson, Mona T. Diab, Aylin Caliskan |
NAACL (Long Papers) | 4 |
| 2024 | Investigating Cultural Alignment of Large Language ModelsabstractThe intricate relationship between language and culture has long been a subject of exploration within the realm of linguistic anthropology.Large Language Models (LLMs), promoted as repositories of collective human knowledge, raise a pivotal question: do these models genuinely encapsulate the diverse knowledge adopted by different cultures?Our study reveals that these models demonstrate greater cultural alignment along two dimensions-firstly, when prompted with the dominant language of a specific culture, and secondly, when pretrained with a refined mixture of languages employed by that culture.We quantify cultural alignment by simulating sociological surveys, comparing model responses to those of actual survey participants as references.Specifically, we replicate a survey conducted in various regions of Egypt and the United States through prompting LLMs with different pretraining data mixtures in both Arabic and English with the personas of the real respondents and the survey questions.Further analysis reveals that misalignment becomes more pronounced for underrepresented personas and for culturally sensitive topics, such as those probing social values.Finally, we introduce Anthropological Prompting, a novel method leveraging anthropological reasoning to enhance cultural alignment.Our study emphasizes the necessity for a more balanced multilingual pretraining dataset to better represent the diversity of human experience and the plurality of different cultures with many implications on the topic of cross-lingual transfer.1 Badr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. Diab |
ACL (1) | 4 |
| 2024 | Depth-Wise Attention (DWAtt): A Layer Fusion Method for Data-Efficient ClassificationabstractLanguage Models pretrained on large textual data have been shown to encode different types of knowledge simultaneously. Traditionally, only the features from the last layer are used when adapting to new tasks or data. We put forward that, when using or finetuning deep pretrained models, intermediate layer features that may be relevant to the downstream task are buried too deep to be used efficiently in terms of needed samples or steps. To test this, we propose a new layer fusion method: Depth-Wise Attention (DWAtt), to help re-surface signals from non-final layers. We compare DWAtt to a basic concatenation-based layer fusion method (Concat), and compare both to a deeper model baseline—all kept within a similar parameter budget. Our findings show that DWAtt and Concat are more step- and sample-efficient than the baseline, especially in the few-shot setting. DWAtt outperforms Concat on larger data sizes. On CoNLL-03 NER, layer fusion shows 3.68 − 9.73% F1 gain at different few-shot sizes. The layer fusion models presented significantly outperform the baseline in various training scenarios with different data sizes, architectures, and training constraints. Muhammad N. ElNokrashy, Badr AlKhamissi, Mona T. Diab |
LREC/COLING | 3 |
| 2024 | Towards a Responsible Thinking in the New Era of Gen AI: Walking the Walk
Mona T. Diab |
DATA | 1 |
| 2024 | GRASS: Compute Efficient Low-Memory LLM Training with Structured Sparse GradientsabstractLarge language model (LLM) training and finetuning are often bottlenecked by limited GPU memory.While existing projection-based optimization methods address this by projecting gradients into a lower-dimensional subspace to reduce optimizer state memory, they typically rely on dense projection matrices, which can introduce computational and memory overheads.In this work, we propose GRASS (GRAdient Stuctured Sparsification), a novel approach that leverages sparse projections to transform gradients into structured sparse updates.This design not only significantly reduces memory usage for optimizer states but also minimizes gradient memory footprint, computation, and communication costs, leading to substantial throughput improvements.Extensive experiments on pretraining and finetuning tasks demonstrate that GRASS achieves competitive performance to full-rank training and existing projection-based methods.Notably, GRASS enables half-precision pretraining of a 13B parameter LLaMA model on a single 40GB A100 GPU-a feat infeasible for previous methodsand yields up to a 2× throughput improvement on an 8-GPU system.Code is released here 1 .P ← compute P (∇L(W (t) )) ▷ P ∈ R m×r 8: // [Optional] Update optimizer state 9: Aashiq Muhamed, Oscar Li, David P. Woodruff, Mona T. Diab, Virginia Smith |
EMNLP | 4 |
| 2024 | Can Large Language Models Infer Causation from Correlation?abstractCausal inference is one of the hallmarks of human intelligence. While the field of CausalNLP has attracted much interest in the recent years, existing causal inference datasets in NLP primarily rely on discovering causality from empirical knowledge (e.g., commonsense knowledge). In this work, we propose the first benchmark dataset to test the pure causal inference skills of large language models (LLMs). Specifically, we formulate a novel task Corr2Cause, which takes a set of correlational statements and determines the causal relationship between the variables. We curate a large-scale dataset of more than 200K samples, on which we evaluate seventeen existing LLMs. Through our experiments, we identify a key shortcoming of LLMs in terms of their causal inference skills, and show that these models achieve almost close to random performance on the task. This shortcoming is somewhat mitigated when we try to re-purpose LLMs for this skill via finetuning, but we find that these models still fail to generalize – they can only perform causal inference in in-distribution settings when variable names and textual expressions used in the queries are similar to those in the training set, but fail in out-of-distribution settings generated by perturbing these queries. Corr2Cause is a challenging task for LLMs, and can be helpful in guiding future research on improving LLMs’ pure reasoning skills and generalizability. Our data is at https://huggingface.co/datasets/causalnlp/corr2cause. Our code is at https://github.com/causalNLP/corr2cause. Zhijing Jin 0001, Jiarui Liu 0004, Zhiheng Lyu, Spencer Poff, Mrinmaya Sachan, Rada Mihalcea, Mona T. Diab, Bernhard Schölkopf |
ICLR | 7 |
| 2024 | Automatic Generation of Model and Data Cards: A Step Towards Responsible AIabstractJiarui Liu, Wenkai Li, Zhijing Jin, Mona Diab. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jiarui Liu 0004, Zhijing Jin 0001, Mona T. Diab |
NAACL-HLT | 4 |
| 2024 | Analyzing the Role of Semantic Representations in the Era of Large Language ModelsabstractZhijing Jin, Yuen Chen, Fernando Gonzalez Adauto, Jiarui Liu, Jiayi Zhang, Julian Michael, Bernhard Schölkopf, Mona Diab. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhijing Jin 0001, Yuen Chen, Fernando Gonzalez Adauto, Jiarui Liu 0004, Julian Michael, Bernhard Schölkopf, Mona T. Diab |
NAACL-HLT | 8 |
| 2023 | ALERT: Adapt Language Models to Reasoning TasksabstractPing Yu, Tianlu Wang, Olga Golovneva, Badr AlKhamissi, Siddharth Verma, Zhijing Jin, Gargi Ghosh, Mona Diab, Asli Celikyilmaz. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Olga Golovneva, Badr AlKhamissi, Siddharth Verma, Zhijing Jin 0001, Gargi Ghosh, Mona T. Diab, Asli Celikyilmaz |
ACL (1) | 8 |
| 2023 | Methods for Measuring, Updating, and Visualizing Factual Beliefs in Language ModelsabstractPeter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Peter Hase, Mona T. Diab, Asli Celikyilmaz, Xian Li 0003, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer 0001 |
EACL | 2 |
| 2022 | ToKen: Task Decomposition and Knowledge Infusion for Few-Shot Hate Speech DetectionabstractBadr AlKhamissi, Faisal Ladhak, Srinivasan Iyer, Veselin Stoyanov, Zornitsa Kozareva, Xian Li, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona Diab. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Badr AlKhamissi, Faisal Ladhak, Srinivasan Iyer 0001, Veselin Stoyanov, Zornitsa Kozareva, Xian Li 0003, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona T. Diab |
EMNLP | 10 |
| 2022 | Efficient Large Scale Language Modeling with Mixtures of ExpertsabstractMikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, Giridharan Anantharaman, Xian Li, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Xing Zhou, Punit Singh Koura, Brian O’Horo, Jeffrey Wang, Luke Zettlemoyer, Mona Diab, Zornitsa Kozareva, Veselin Stoyanov. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Mikel Artetxe, Shruti Bhosale, Naman Goyal 0001, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer 0001, Ramakanth Pasunuru, Giri Anantharaman, Xian Li 0003, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Punit Singh Koura, Brian O'Horo, Jeffrey Wang, Luke Zettlemoyer, Mona T. Diab, Zornitsa Kozareva, Veselin Stoyanov |
EMNLP | 22 |
| 2022 | Few-shot Learning with Multilingual Generative Language ModelsabstractXi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, Xian Li. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal 0001, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona T. Diab, Veselin Stoyanov, Xian Li 0003 |
EMNLP | 19 |
| 2022 | BeSt: The Belief and Sentiment CorpusabstractWe present the BeSt corpus, which records cognitive state: who believes what (i.e., factuality), and who has what sentiment towards what. This corpus is inspired by similar source-and-target corpora, specifically MPQA and FactBank. The corpus comprises two genres, newswire and discussion forums, in three languages, Chinese (Mandarin), English, and Spanish. The corpus is distributed through the LDC. Jennifer Tracey, Owen Rambow, Claire Cardie, Adam Dalton 0001, Hoa Trang Dang, Mona T. Diab, Bonnie J. Dorr, Louise Guthrie, Magdalena Markowska, Smaranda Muresan, Vinodkumar Prabhakaran, Samira Shaikh, Tomek Strzalkowski |
LREC | 6 |
| 2022 | AnswerSumm: A Manually-Curated Dataset and Pipeline for Answer SummarizationabstractAlexander Fabbri, Xiaojian Wu, Srini Iyer, Haoran Li, Mona Diab. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Alexander R. Fabbri, Xiaojian Wu, Srinivasan Iyer 0001, Haoran Li 0007, Mona T. Diab |
NAACL-HLT | 5 |
| 2021 | Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel DataabstractWei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona Diab. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal 0001, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona T. Diab |
ACL/IJCNLP (1) | 9 |
| 2020 | A Multitask Learning Approach for Diacritic RestorationabstractIn many languages like Arabic, diacritics are used to specify pronunciations as well as meanings. Such diacritics are often omitted in written text, increasing the number of possible pronunciations and meanings for a word. This results in a more ambiguous text making computational processing on such text more difficult. Diacritic restoration is the task of restoring missing diacritics in the written text. Most state-of-the-art diacritic restoration models are built on character level information which helps generalize the model to unseen data, but presumably lose useful information at the word level. Thus, to compensate for this loss, we investigate the use of multi-task learning to jointly optimize diacritic restoration with related NLP problems namely word segmentation, part-of-speech tagging, and syntactic diacritization. We use Arabic as a case study since it has sufficient data resources for tasks that we consider in our joint modeling. Our joint models significantly outperform the baselines and are comparable to the state-of-the-art models that are more complex relying on morphological analyzers and/or a lot more data (e.g. dialectal data). Sawsan Alqahtani, Ajay Mishra, Mona T. Diab |
ACL | 3 |
| 2020 | FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationabstractNeural abstractive summarization models are prone to generate content inconsistent with the source document, i.e. unfaithful.Existing automatic metrics do not capture such mistakes effectively.We tackle the problem of evaluating faithfulness of a generated summary given its source document.We first collected human annotations of faithfulness for outputs from numerous models on two datasets.We find that current models exhibit a trade-off between abstractiveness and faithfulness: outputs with less word overlap with the source document are more likely to be unfaithful.Next, we propose an automatic question answering (QA) based metric for faithfulness, FEQA, 1 which leverages recent advances in reading comprehension.Given questionanswer pairs generated from the summary, a QA model extracts answers from the document; non-matched answers indicate unfaithful information in the summary.Among metrics based on word overlap, embedding similarity, and learned language understanding models, our QA-based metric has significantly higher correlation with human faithfulness scores, especially on highly abstractive summaries.* Most of the work is done while the authors were at Amazon Web Services AI.1 Faithfulness Evaluation with Question Answering. Esin Durmus, He He 0001, Mona T. Diab |
ACL | 3 |
| 2020 | DeSePtion: Dual Sequence Prediction and Adversarial Examples for Improved Fact-CheckingabstractChristopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia, Kriste Krstovski, Mona Diab, Smaranda Muresan. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Christopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia, Kriste Krstovski, Mona T. Diab, Smaranda Muresan |
ACL | 6 |
| 2020 | Detecting Urgency Status of Crisis Tweets: A Transfer Learning Approach for Low Resource LanguagesabstractWe release an urgency dataset that consists of English tweets relating to natural crises.The set is annotated along with annotations of their corresponding urgency status.Additionally, we release evaluation datasets for two low-resource languages, i.e.Sinhala and Odia, and demonstrate an effective zero-shot transfer from English to these two languages by training cross-lingual classifiers.We adopt cross-lingual embeddings constructed using different methods to extract features of the tweets, including a few state-of-the-art contextual embeddings such as BERT, RoBERTa and XLM-R.We train a variety of classifier architectures, supervised and semi supervised, on the extracted features.We also further experiment with ensembling the various classifiers.With very limited amounts of labeled data in English and zero data in the low resource languages, we show a successful framework of training monolingual and cross-lingual classifiers using deep learning methods which are known to be data hungry.Specifically, we show that the recent deep contextual embeddings are also helpful when dealing with very small-scale datasets.Classifiers that incorporate RoBERTa yield the best performance for the English urgency detection task, with 25% F1 score absolute improvement over the baselines.For the zero-shot transfer to low resource languages, classifiers that use LASER features perform the best for Sinhala transfer while XLM-R features benefit the Odia transfer the most. Efsun Sarioglu Kayi, Linyong Nan, Bohan Qu, Mona T. Diab, Kathy McKeown |
COLING | 4 |
| 2020 | Multitask Learning for Cross-Lingual Transfer of Broad-coverage Semantic DependenciesabstractWe describe a method for developing broad-coverage semantic dependency parsers for languages for which no semantically annotated resource is available. We leverage a multitask learning framework coupled with annotation projection. We use syntactic parsing as the auxiliary task in our multitask setup. Our annotation projection experiments from English to Czech show that our multitask setup yields 3.1% (4.2%) improvement in labeled F1-score on in-domain (out-of-domain) test set compared to a single-task baseline. Maryam Aminian, Mohammad Sadegh Rasooli, Mona T. Diab |
EMNLP (1) | 3 |
| 2020 | Data Paucity and Low Resource Scenarios: Challenges and OpportunitiesabstractIn an era of unstructured data abundance, you would think that we have solved our data requirements for building robust systems for language processing. However, this is not the case if we think on a global scale with over 7000 languages where only a handful have digital resources. Systems at scale with good performance typically require annotated resources that cover the genres and domain divides. Moreover, the existence of a handful of resources in some languages is a reflection of the digital disparity in various societies leading to inadvertent biases in systems. In this talk I will show some solutions for low resource scenarios, both cross domain and genres as well as cross lingually. Mona T. Diab |
KDD | 1 |
| 2020 | Diversity, Density, and Homogeneity: Quantitative Characteristic Metrics for Text CollectionsabstractSummarizing data samples by quantitative measures has a long history, with descriptive statistics being a case in point. However, as natural language processing methods flourish, there are still insufficient characteristic metrics to describe a collection of texts in terms of the words, sentences, or paragraphs they comprise. In this work, we propose metrics of diversity, density, and homogeneity that quantitatively measure the dispersion, sparsity, and uniformity of a text collection. We conduct a series of simulations to verify that each metric holds desired properties and resonates with human intuitions. Experiments on real-world datasets demonstrate that the proposed characteristic metrics are highly correlated with text classification performance of a renowned model, BERT, which could inspire future applications. Yi-An Lai, Xuan Zhu 0002, Mona T. Diab |
LREC | 4 |
| 2019 | Efficient Sentence Embedding using Discrete Cosine TransformabstractNada Almarwani, Hanan Aldarmaki, Mona Diab. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Nada AlMarwani, Hanan Aldarmaki, Mona T. Diab |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Efficient Convolutional Neural Networks for Diacritic RestorationabstractSawsan Alqahtani, Ajay Mishra, Mona Diab. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sawsan Alqahtani, Ajay Mishra, Mona T. Diab |
EMNLP/IJCNLP (1) | 3 |
| 2019 | CASA-NLU: Context-Aware Self-Attentive Natural Language Understanding for Task-Oriented ChatbotsabstractArshit Gupta, Peng Zhang, Garima Lalwani, Mona Diab. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Arshit Gupta, Garima Lalwani, Mona T. Diab |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Multi-Domain Goal-Oriented Dialogues (MultiDoGO): Strategies toward Curating and Annotating Large Scale Dialogue DataabstractDenis Peskov, Nancy Clarke, Jason Krone, Brigi Fodor, Yi Zhang, Adel Youssef, Mona Diab. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Denis Peskov, Nancy E. Clarke, Jason Krone, Brigi Fodor, Adel Youssef, Mona T. Diab |
EMNLP/IJCNLP (1) | 7 |
| 2019 | Understanding Cohesion in Writings and Speech of Schizophrenia PatientsabstractSchizophrenia is one of the mental disorders that impacts a person's thinking, speech, and actions. It can reduce a person's ability to process auditory information and make decisions. Analyzing this disorder correctly is important because it might help with different ways of reducing its negative effects on its patients. Linguists and psychiatrists have been investigating language impairments and speech disorder in people with schizophrenia disorder which can be challenging. In this study, we attempt to address this issue by analyzing linguistic features i.e. cohesion in the writings and speech scripts of schizophrenia patients. Our results show that using referential cohesion with text easability or situation model features provides the best performance for speech whereas for writing dataset, readability or a combination of situation model and readability yield the best performance. Amal AlQahtani, Efsun Sarioglu Kayi, Mona T. Diab |
ICMLA | 3 |
| 2019 | Investigating Input and Output Units in Diacritic RestorationabstractDiacritic restoration is the task of assigning diacritics (accents) for each character in a given segment. The typical input levels that have been previously used in diacritic restoration models are word and/or character units. In this paper, we investigate the use of subwords as input units along with their diacritic patterns (combinations of adjacent diacritics) as output, as an alternative to word or character-based models. Our experiments show that characters provide the optimal level of information for sequence-based diacritic restoration models across different languages. We additionally improved our diacritic restoration model by maximizing over the output diacritic sequence using a Conditional Random Field (CRF). Adding a CRF layer improves the performance on observed and unobserved words substantially for Arabic and marginally for Yoruba. Sawsan Alqahtani, Mona T. Diab |
ICMLA | 2 |
| 2018 | Evaluation of Unsupervised Compositional RepresentationsabstractWe evaluated various compositional models, from bag-of-words representations to compositional RNN-based models, on several extrinsic supervised and unsupervised evaluation benchmarks. Our results confirm that weighted vector averaging can outperform context-sensitive models in most benchmarks, but structural features encoded in RNN models can also be useful in certain classification tasks. We analyzed some of the evaluation datasets to identify the aspects of meaning they measure and the characteristics of the various models that explain their performance variance. Hanan Aldarmaki, Mona T. Diab |
COLING | 2 |
| 2018 | Emotion Detection and Classification in a Multigenre Corpus with Joint Multi-Task Deep LearningabstractDetection and classification of emotion categories expressed by a sentence is a challenging task due to subjectivity of emotion. To date, most of the models are trained and evaluated on single genre and when used to predict emotion in different genre their performance drops by a large margin. To address the issue of robustness, we model the problem within a joint multi-task learning framework. We train this model with a multigenre emotion corpus to predict emotions across various genre. Each genre is represented as a separate task, we use soft parameter shared layers across the various tasks. our experimental results show that this model improves the results across the various genres, compared to a single genre training in the same neural net architecture. Shabnam Tafreshi, Mona T. Diab |
COLING | 2 |
| 2018 | WASA: A Web Application for Sequence Annotation
Fahad Alghamdi 0003, Mona T. Diab |
LREC | 2 |
| 2018 | Sentence and Clause Level Emotion Annotation, Detection, and Classification in a Multi-Genre Corpus
Shabnam Tafreshi, Mona T. Diab |
LREC | 2 |
| 2018 | Unsupervised Word Mapping Using Structural Similarities in Monolingual EmbeddingsabstractMost existing methods for automatic bilingual dictionary induction rely on prior alignments between the source and target languages, such as parallel corpora or seed dictionaries. For many language pairs, such supervised alignments are not readily available. We propose an unsupervised approach for learning a bilingual dictionary for a pair of languages given their independently-learned monolingual word embeddings. The proposed method exploits local and global structures in monolingual vector spaces to align them such that similar words are mapped to each other. We show empirically that the performance of bilingual correspondents that are learned using our proposed unsupervised method is comparable to that of using supervised bilingual correspondents from a seed dictionary. Hanan Aldarmaki, Mahesh Mohan, Mona T. Diab |
Trans. Assoc. Comput. Linguistics | 3 |
| 2016 | LILI: A Simple Language Independent Approach for Language IdentificationabstractWe introduce a generic Language Independent Framework for Linguistic Code Switch Point Detection. The system uses characters level 5-grams and word level unigram language models to train a conditional random fields (CRF) model for classifying input words into various languages. We test our proposed framework and compare it to the state-of-the-art published systems on standard data sets from several language pairs: English-Spanish, Nepali-English, English-Hindi, Arabizi (Refers to Arabic written using the Latin/Roman script)-English, Arabic-Engari (Refers to English written using Arabic script), Modern Standard Arabic(MSA)-Egyptian, Levantine-MSA, Gulf-MSA, one more English-Spanish, and one more MSA-EGY. The overall weighted average F-score of each language pair are 96.4%, 97.3%, 98.0%, 97.0%, 98.9%, 86.3%, 88.2%, 90.6%, 95.2%, and 85.0% respectively. The results show that our approach despite its simplicity, either outperforms or performs at comparable levels to state-of-the-art published systems. Mohamed Al-Badrashiny, Mona T. Diab |
COLING | 2 |
| 2016 | Computational Approaches to Linguistic Code Switching
Mona T. Diab, Pascale Fung, Julia Hirschberg, Thamar Solorio |
INTERSPEECH | 1 |
| 2016 | SPLIT: Smart Preprocessing (Quasi) Language Independent Tool
Mohamed Al-Badrashiny, Arfath Pasha, Mona T. Diab, Nizar Habash, Owen Rambow, Wael Salloum, Ramy Eskander |
LREC | 3 |
| 2016 | Creating a Large Multi-Layered Representational Repository of Linguistic Code Switched Arabic Data
Mona T. Diab, Mahmoud Ghoneim, Abdelati Hawwari, Fahad Alghamdi 0003, Nada AlMarwani, Mohamed Al-Badrashiny |
LREC | 1 |
| 2016 | Explicit Fine grained Syntactic and Semantic Annotation of the Idafa Construction in Arabic
Abdelati Hawwari, Mahmoud Ghoneim, Mona T. Diab |
LREC | 4 |
| 2016 | Guidelines and Framework for a Large Scale Arabic Diacritized Corpus
Wajdi Zaghouani, Houda Bouamor, Abdelati Hawwari, Mona T. Diab, Ossama Obeid, Mahmoud Ghoneim, Sawsan Alqahtani, Kemal Oflazer |
LREC | 4 |
| 2015 | Tharawat: A Vision for a Comprehensive Resource for Arabic Computational Processing
Mona T. Diab |
CICLing (1) | 1 |
| 2015 | AIDA2: A Hybrid Approach for Token and Sentence Level Dialect Identification in ArabicabstractIn this paper, we present a hybrid approach for performing token and sentence levels Dialect Identification in Arabic.Specifically we try to identify whether each token in a given sentence belongs to Modern Standard Arabic (MSA), Egyptian Dialectal Arabic (EDA) or some other class and whether the whole sentence is mostly EDA or MSA.The token level component relies on a Conditional Random Field (CRF) classifier that uses decisions from several underlying components such as language models, a named entity recognizer and and a morphological analyzer to label each word in the sentence.The sentence level component uses a classifier ensemble system that relies on two independent underlying classifiers that model different aspects of the language.Using a featureselection heuristic, we select the best set of features for each of these two classifiers.We then train another classifier that uses the class labels and the confidence scores generated by each of the two underlying classifiers to decide upon the final class for each sentence.The token level component yields a new state of the art F-score of 90.6% (compared to previous state of the art of 86.8%) and the sentence level component yields an accuracy of 90.8% (compared to 86.6% obtained by the best state of the art system). Mohamed Al-Badrashiny, Heba Elfardy, Mona T. Diab |
CoNLL | 3 |
| 2014 | Fast Tweet Retrieval with Compact Binary Codes
Weiwei Guo, Wei Liu 0005, Mona T. Diab |
COLING | 3 |
| 2014 | SANA: A Large Scale Multi-Genre, Multi-Dialect Lexicon for Arabic Subjectivity and Sentiment Analysis
Muhammad Abdul-Mageed, Mona T. Diab |
LREC | 2 |
| 2014 | Tharwa: A Large Scale Dialectal Arabic - Standard Arabic - English Lexicon
Mona T. Diab, Mohamed Al-Badrashiny, Maryam Aminian, Heba Elfardy, Nizar Habash, Abdelati Hawwari, Wael Salloum, Pradeep Dasigi, Ramy Eskander |
LREC | 1 |
| 2014 | MADAMIRA: A Fast, Comprehensive Tool for Morphological Analysis and Disambiguation of Arabic
Arfath Pasha, Mohamed Al-Badrashiny, Mona T. Diab, Ahmed El Kholy, Ramy Eskander, Nizar Habash, Manoj Pooleery, Owen Rambow, Ryan Roth |
LREC | 3 |
| 2014 | SAMAR: Subjectivity and sentiment analysis for Arabic social media
Muhammad Abdul-Mageed, Mona T. Diab, Sandra Kübler |
Comput. Speech Lang. | 2 |
| 2013 | Linking Tweets to News: A Framework to Enrich Short Text Data in Social Media
Weiwei Guo, Hao Li 0031, Heng Ji 0001, Mona T. Diab |
ACL (1) | 4 |
| 2013 | Multiword Expressions in the Context of Statistical Machine Translation
Mahmoud Ghoneim, Mona T. Diab |
IJCNLP | 2 |
| 2013 | DIRA: Dialectal Arabic Information Retrieval Assistant
Arfath Pasha, Mohamed Al-Badrashiny, Mohamed Altantawy, Nizar Habash, Manoj Pooleery, Owen Rambow, Ryan Roth, Mona T. Diab |
IJCNLP | 8 |
| 2013 | Improving Lexical Semantics for Sentential Semantics: Modeling Selectional Preference and Similar Words in a Latent Variable Model
Weiwei Guo, Mona T. Diab |
HLT-NAACL | 2 |
| 2013 | Code Switch Point Detection in Arabic
Heba Elfardy, Mohamed Al-Badrashiny, Mona T. Diab |
NLDB | 3 |
| 2013 | ANEAR: Automatic Named Entity Aliasing Resolution
Ayah Zirikly, Mona T. Diab |
NLDB | 2 |
| 2012 | Subgroup Detection in Ideological Discussions
Amjad Abu-Jbara, Pradeep Dasigi, Mona T. Diab, Dragomir R. Radev |
ACL (1) | 3 |
| 2012 | Modeling Sentences in the Latent Space
Weiwei Guo, Mona T. Diab |
ACL (1) | 2 |
| 2012 | Who's (Really) the Boss? Perception of Situational Power in Written Interactions
Vinodkumar Prabhakaran, Owen Rambow, Mona T. Diab |
COLING | 3 |
| 2012 | AWATIF: A Multi-Genre Corpus for Modern Standard Arabic Subjectivity and Sentiment Analysis
Muhammad Abdul-Mageed, Mona T. Diab |
LREC | 2 |
| 2012 | Simplified guidelines for the creation of Large Scale Dialectal Arabic Annotations
Heba Elfardy, Mona T. Diab |
LREC | 2 |
| 2012 | Conventional Orthography for Dialectal Arabic
Nizar Habash, Mona T. Diab, Owen Rambow |
LREC | 2 |
| 2012 | Annotations for Power Relations on Email Threads
Vinodkumar Prabhakaran, Huzaifa Neralwala, Owen Rambow, Mona T. Diab |
LREC | 4 |
| 2012 | Arabic Dialect Processing Tutorial
Mona T. Diab, Nizar Habash |
HLT-NAACL | 1 |
| 2012 | Predicting Overt Display of Power in Written Dialogs
Vinodkumar Prabhakaran, Owen Rambow, Mona T. Diab |
HLT-NAACL | 3 |
| 2011 | Semantic Topic Models: Combining Word Distributional Statistics and Dictionary Definitions
Weiwei Guo, Mona T. Diab |
EMNLP | 2 |
| 2011 | CODACT: Towards Identifying Orthographic Variants in Dialectal Arabic
Pradeep Dasigi, Mona T. Diab |
IJCNLP | 2 |
| 2011 | Introduction to the Special Issue on Arabic Computational Linguisticsabstractintroduction Share on Introduction to the Special Issue on Arabic Computational Linguistics Editors: Graham Katz Georgetown University Georgetown UniversityView Profile , Mona Diab Columbia University Columbia UniversityView Profile Authors Info & Claims ACM Transactions on Asian Language Information ProcessingVolume 10Issue 1Article No.: 1pp 1–4https://doi.org/10.1145/1929908.1929909Published:01 March 2011Publication History 0citation297DownloadsMetricsTotal Citations0Total Downloads297Last 12 Months8Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Graham Katz, Mona T. Diab |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2010 | Combining Orthogonal Monolingual and Multilingual Sources of Evidence for All Words WSD
Weiwei Guo, Mona T. Diab |
ACL | 2 |
| 2010 | Task-based Evaluation of Multiword Expressions: a Pilot Study in Statistical Machine Translation
Marine Carpuat, Mona T. Diab |
HLT-NAACL | 2 |
| 2009 | Who, What, When, Where, Why? Comparing Multiple Approaches to the Cross-Lingual 5W Task
Kristen Parton, Kathy McKeown, Bob Coyne, Mona T. Diab, Ralph Grishman, Dilek Hakkani-Tür, Mary P. Harper, Heng Ji 0001, Wei-Yun Ma, Adam Meyers 0001, Sara Stolbach, Ang Sun, Gökhan Tür, Wei Xu 0004, Sibel Yaman |
ACL/IJCNLP | 4 |
| 2009 | Unsupervised Classification of Verb Noun Multi-Word Expression Tokens
Mona T. Diab, Madhav Krishna |
CICLing | 1 |
| 2009 | Arabic Named Entity Recognition: A Feature-Driven StudyabstractThe named entity recognition task aims at identifying and classifying named entities within an open-domain text. This task has been garnering significant attention recently as it has been shown to help improve the performance of many natural language processing applications. In this paper, we investigate the impact of using different sets of features in three discriminative machine learning frameworks, namely, support vector machines, maximum entropy and conditional random fields for the task of named entity recognition. Our language of interest is Arabic. We explore lexical, contextual and morphological features and nine data-sets of different genres and annotations. We measure the impact of the different features in isolation and incrementally combine them in order to evaluate the robustness to noise of each approach. We achieve the highest performance using a combination of 15 features in conditional random fields using broadcast news data (Fbeta=1=83.34). Yassine Benajiba, Mona T. Diab, Paolo Rosso |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Semantic Role Labeling Systems for Arabic using Kernel Methods
Mona T. Diab, Alessandro Moschitti, Daniele Pighin |
ACL | 1 |
| 2008 | Arabic Named Entity Recognition using Optimized Feature Sets
Yassine Benajiba, Mona T. Diab, Paolo Rosso |
EMNLP | 2 |
| 2008 | A Pilot Arabic Propbank
Martha Palmer, Olga Babko-Malaya, Ann Bies, Mona T. Diab, Mohamed Maamouri, Aous Mansouri, Wajdi Zaghouani |
LREC | 4 |
| 2007 | Arabic diacritization in the context of statistical machine translation
Mona T. Diab, Mahmoud Ghoneim, Nizar Habash |
MTSummit | 1 |
| 2007 | Semi-automatic error analysis for large-scale statistical machine translation
Katrin Kirchhoff, Owen Rambow, Nizar Habash, Mona T. Diab |
MTSummit | 4 |
| 2006 | Unsupervised Induction of Modern Standard Arabic Verb Classes Using Syntactic Frames and LSA
Neal Snider, Mona T. Diab |
ACL | 2 |
| 2006 | Parsing Arabic Dialects
David Chiang 0001, Mona T. Diab, Nizar Habash, Owen Rambow, Safiullah Shareef |
EACL | 2 |
| 2006 | Developing and Using a Pilot Dialectal Arabic Treebank
Mohamed Maamouri, Ann Bies, Tim Buckwalter, Mona T. Diab, Nizar Habash, Owen Rambow, Dalila Tabessi |
LREC | 4 |
| 2006 | Unsupervised Induction of Modern Standard Arabic Verb Classes
Neal Snider, Mona T. Diab |
HLT-NAACL | 2 |
| 2004 | Relieving the data Acquisition Bottleneck in Word Sense DisambiguationabstractSupervised learning methods for WSD yield better performance than unsupervised methods. Yet the availability of clean training data for the former is still a severe challenge. In this paper, we present an unsupervised bootstrapping approach for WSD which exploits huge amounts of automatically generated noisy data for training within a supervised learning framework. The method is evaluated using the 29 nouns in the English Lexical Sample task of SENSEVAL 2. Our algorithm does as well as supervised algorithms on 31% of this test set, which is an improvement of 11% (absolute) over state-of-the-art bootstrapping WSD algorithms. We identify seven different factors that impact the performance of our system. Mona T. Diab |
ACL | 1 |
| 2002 | An Unsupervised Method for Word Sense Tagging using Parallel CorporaabstractWe present an unsupervised method for word sense disambiguation that exploits translation correspondences in parallel corpora. The technique takes advantage of the fact that cross-language lexicalizations of the same concept tend to be consistent, preserving some core element of its semantics, and yet also variable, reflecting differing translator preferences and the influence of context. Working with parallel corpora introduces an extra complication for evaluation, since it is difficult to find a corpus that is both sense tagged and parallel with another language; therefore we use pseudo-translations, created by machine translation systems, in order to make possible the evaluation of the approach against a standard test set. The results demonstrate that word-level translation correspondences are a valuable source of information for sense disambiguation. Mona T. Diab, Philip Resnik |
ACL | 1 |