Muhammad Abdul-Mageed

dblp:49/9389 · DBLP profile ↗
← Back
45ranked-venue papers
11as first author
32since 2021 · last 2026
0000-0002-8590-2040ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 9 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs
abstract
Abdellah EL Mekki, Samar M. Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-hibiri, Razan Saadie, Hamzah A. Alsayadi, Nadia Ghezaiel Hammouda, Alshima Mohammed Alkhazimi, Aya Hamod, Al-Yas Yaqoob Al-Ghafri, Wesam El-Sayed, Asila Ismail al Sharji, Mohamad Ballout, Anas Belfathi, Karim Ghaddar, Serry Sibaee, Alaa Aoun, Aeej Mohammed Aseri, Lina Abureesh, Ahlam Bashiti, Majdal Yousef, Abdulaziz Hafiz, Yehdih Mohamed, Emira Hamedtou, Brakehe Emehah, Rahaf Alhamouri, Youssef Nafea, Aya El Aatar, Walid Al-Dhabyani, Emhemed S. Hamed, Sara Shatnawi, Fakhraddin Alwajih, Khalid Elkhidir, Ashwag Alasmari, Abdurrahman Gerrio, Omar Said Alshahri, AbdelRahim A. Elmadany, Ismail Berrada, Amir Azad Adli Al-kathiri, Fadi Zaraket, Mustafa Jarrar, Yahya Mohamed EL Hadj, Hassan Alhuzali, Muhammad Abdul-Mageed. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Abdellah El Mekki, Samar Mohamed Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-Hibiri, Razan Saadie, Hamzah A. Alsayadi, Nadia Ghezaiel Hammouda, Alshima Alkhazimi, Aya Hamod, Al-Yas Al-Ghafri, Wesam El-Sayed, Asila Al Sharji, Mohamad Ballout, Anas Belfathi, Karim Ghaddar, Serry Sibaee, Alaa Aoun, Aeej Mohammed Aseri, Lina Abureesh, Ahlam Bashiti, Majdal Yousef, Abdulaziz Hafiz, Yehdih Mohamed, Emira Hamedtou, Brakehe Emehah, Rahaf Alhamouri, Youssef Nafea, Aya El Aatar, Walid Al-Dhabyani, Emhemed Hamed, Sara Shatnawi, Fakhraddin Alwajih, Khalid Elkhidir, Ashwag Alasmari, Abdurrahman Gerrio, Omar Alshahri, AbdelRahim A. Elmadany, Ismail Berrada, Amir Azad Adli Alkathiri, Fadi A. Zaraket, Mustafa Jarrar, Yahya Mohamed El Hadj, Hassan Alhuzali, Muhammad Abdul-Mageed
ACL (1)47
2026 Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMs
abstract
Large language models (LLMs) are widely deployed for open-ended communication, yet most bias evaluations still rely on English, classification-style tasks. We introduce \corpusname, a new multilingual, debate-style benchmark designed to reveal how narrative bias appears in realistic generative settings. Our dataset includes 8{,}400 structured debate prompts spanning four sensitive domains -- Women's Rights, Backwardness, Terrorism, and Religion -- across seven languages ranging from high-resource (English, Chinese) to low-resource (Swahili, Nigerian Pidgin). Using four flagship models (GPT-4o, Claude~3.5~Haiku, DeepSeek-Chat, and LLaMA-3-70B), we generate over 100{,}000 debate responses and automatically classify which demographic groups are assigned stereotyped versus modern roles. Results show that all models reproduce entrenched stereotypes despite safety alignment: Arabs are overwhelmingly linked to Terrorism and Religion ($\geq$89\%), Africans to socioeconomic ``backwardness'' (up to 77\%), and Western groups are consistently framed as modern or progressive. Biases grow sharply in lower-resource languages, revealing that alignment trained primarily in English does not generalize globally. Our findings highlight a persistent divide in multilingual fairness: current alignment methods reduce explicit toxicity but fail to prevent biased outputs in open-ended contexts. We release our \corpusname benchmark and analysis framework to support the next generation of multilingual bias evaluation and safer, culturally inclusive model alignment.
Muhammed Saeed, Muhammad Abdul-Mageed, Shady Shehata
LREC2
2025 Where Are We? Evaluating LLM Performance on African Languages
abstract
Ife Adebara, Hawau Olamide Toyin, Nahom Tesfu Ghebremichael, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ife Adebara, Hawau Olamide Toyin, Nahom Tesfu Ghebremichael, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed
ACL (1)5
2025 Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs
abstract
Fakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany, Omer Nacar, El Moatez Billah Nagoudi, Reem Abdel-Salam, Hanin Atwany, Youssef Nafea, Abdulfattah Mohammed Yahya, Rahaf Alhamouri, Hamzah A. Alsayadi, Hiba Zayed, Sara Shatnawi, Serry Sibaee, Yasir Ech-chammakhy, Walid Al-Dhabyani, Marwa Mohamed Ali, Imen Jarraya, Ahmed Oumar El-Shangiti, Aisha Alraeesi, Mohammed Anwar AL-Ghrawi, Abdulrahman S. Al-Batati, Elgizouli Mohamed, Noha Taha Elgindi, Muhammed Saeed, Houdaifa Atou, Issam Ait Yahia, Abdelhak Bouayad, Mohammed Machrouh, Amal Makouar, Dania Alkawi, Mukhtar Mohamed, Safaa Taher Abdelfadil, Amine Ziad Ounnoughene, Anfel Rouabhia, Rwaa Assi, Ahmed Sorkatti, Mohamedou Cheikh Tourad, Anis Koubaa, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Fakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany, Omer Nacar, El Moatez Billah Nagoudi, Reem Abdel-Salam, Hanin Atwany, Youssef Nafea, Abdulfattah Mohammed Yahya, Rahaf Alhamouri, Hamzah A. Alsayadi, Hiba Zayed, Sara Shatnawi, Serry Sibaee, Yasir Ech-Chammakhy, Walid Al-Dhabyani, Marwa Mohamed Ali, Imen Jarraya, Ahmed Oumar El-Shangiti, Aisha Alraeesi, Mohammed Anwar Al-Ghrawi, Abdulrahman S. Al-Batati, Elgizouli Mohamed, Noha Taha Elgindi, Muhammed Saeed, Houdaifa Atou, Issam Ait Yahia, Abdelhak Bouayad, Mohammed Machrouh, Amal Makouar, Dania Alkawi, Mukhtar Mohamed, Safaa Taher Abdelfadil, Amine Ziad Ounnoughene, Rouabhia Anfel, Rwaa Assi, Ahmed Sorkatti, Mohamedou Cheikh Tourad, Anis Koubaa, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed
ACL (1)44
2025 Voice of a Continent: Mapping Africa's Speech Technology Frontier
abstract
AbdelRahim A. Elmadany, Sang Yun Kwon, Hawau Olamide Toyin, Alcides Alcoba Inciarte, Hanan Aldarmaki, Muhammad Abdul-Mageed. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
AbdelRahim A. Elmadany, Sang Yun Kwon, Hawau Olamide Toyin, Alcides Alcoba Inciarte, Hanan Aldarmaki, Muhammad Abdul-Mageed
EMNLP6
2025 NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities
abstract
Translate to low-resource language LLM‫ھﺗﺎﻛل؟‬ ‫ﻣش‬ ‫ﻣﺻطﻔﻰ،‬ ‫ﯾﺎ‬ ‫إزﯾك‬ ‫أﺣﻣد:‬ ‫ده،‬ ‫اﻟﺷﺎرع‬ ‫أﻛل‬ ‫ﻣن‬ ‫ﺑﻘﻠﻖ‬ ‫اﻟﻣدام.‬‫ﺑﻛﻠم‬ ‫ﻛﻧت‬ ‫ﻣﻌﻠش،‬ ‫ﻣﺻطﻔﻰ:
Abdellah El Mekki, Houdaifa Atou, Omer Nacar, Shady Shehata, Muhammad Abdul-Mageed
EMNLP5
2025 EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
abstract
Large language models (LLMs) are transforming education by answering questions, explaining complex concepts, and generating content across a wide range of subjects.Despite strong performance on academic benchmarks, they often fail to tailor responses to students grade levels.This is a critical need in K-12 education, where age-appropriate vocabulary and explanation are essential for effective learning.Existing models frequently produce outputs that are too advanced or vague for younger learners, and there are no standardized benchmarks to evaluate their ability to adjust across cognitive and developmental stages.To address this gap, we introduce EDUADAPT, a benchmark of nearly 48k grade-labeled QA pairs across nine science subjects, spanning Grades 1-12 and grouped into four grade levels.We evaluate a diverse set of open-source LLMs on EDU-ADAPT and find that while larger models generally perform better, they still struggle with generating suitable responses for early-grade students (Grades 1-5).Our work presents the first dataset and evaluation framework for assessing grade-level adaptability in LLMs, aiming to foster more developmentally aligned educational AI systems through better training and prompting strategies.EDUADAPT code and datasets are publicly available.
Numaan Naeem, Abdellah El Mekki, Muhammad Abdul-Mageed
EMNLP3
2025 JAWAHER: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking
abstract
Samar Mohamed Magdy, Sang Yun Kwon, Fakhraddin Alwajih, Safaa Taher Abdelfadil, Shady Shehata, Muhammad Abdul-Mageed. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Samar Mohamed Magdy, Sang Yun Kwon, Fakhraddin Alwajih, Safaa Taher Abdelfadil, Shady Shehata, Muhammad Abdul-Mageed
NAACL (Long Papers)6
2025 uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
abstract
Abdul Waheed, Karima Kadaoui, Bhiksha Raj, Muhammad Abdul-Mageed. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Karima Kadaoui, Bhiksha Raj, Muhammad Abdul-Mageed
NAACL (Long Papers)4
2024 Cheetah: Natural Language Generation for 517 African Languages
abstract
Low-resource African languages pose unique challenges for natural language processing (NLP) tasks, including natural language generation (NLG).In this paper, we develop Cheetah, a massively multilingual NLG language model for African languages.Cheetah supports 517 African languages and language varieties, allowing us to address the scarcity of NLG resources and provide a solution to foster linguistic diversity.We demonstrate the effectiveness of Cheetah through comprehensive evaluations across six generation downstream tasks.In five of the six tasks, Cheetah significantly outperforms other models, showcasing its remarkable performance for generating coherent and contextually appropriate text in a wide range of African languages.We additionally conduct a detailed human evaluation to delve deeper into the linguistic capabilities of Cheetah.The findings of this study contribute to advancing NLP research in low-resource settings, enabling greater accessibility and inclusion for African languages in a rapidly expanding digital landscape.The GitHub repository for the Cheetah project is available at https://github.com
Ife Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed
ACL (1)3
2024 Peacock: A Family of Arabic Multimodal Large Language Models and Benchmarks
abstract
Fakhraddin Alwajih, El Moatez Billah Nagoudi, Gagan Bhatia, Abdelrahman Mohamed, Muhammad Abdul-Mageed. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Fakhraddin Alwajih, El Moatez Billah Nagoudi, Gagan Bhatia, Abdel-rahman Mohamed, Muhammad Abdul-Mageed
ACL (1)5
2024 To Distill or Not to Distill? On the Robustness of Robust Knowledge Distillation
abstract
Arabic is known to present unique challenges for Automatic Speech Recognition (ASR).On one hand, its rich linguistic diversity and wide range of dialects complicate the development of robust, inclusive models.On the other, current multilingual ASR models are compute-intensive and lack proper comprehensive evaluations.In light of these challenges, we distill knowledge from large teacher models into smaller student variants that are more efficient.We also introduce a novel human-annotated dataset covering five under-represented Arabic dialects for evaluation.We further evaluate both our models and existing SoTA multilingual models on both standard available benchmarks and our new dialectal data.Our best-distilled model's overall performance (45.0%WER) surpasses that of a SoTA model twice its size (SeamlessM4T-large-v2, WER=47.0%) and its teacher model (Whisper-large-v2, WER=55.1%), and its average performance on our new dialectal data (56.9%WER) outperforms all other models.To gain more insight into the poor performance of these models on dialectal data, we conduct an error analysis and report the main types of errors the different models tend to make.The GitHub repository for the project is available at https: //github.com/UBC-NLP/distill-whisper-ar.
Karima Kadaoui, Muhammad Abdul-Mageed
ACL (1)3
2024 LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions
abstract
Minghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, Alham Fikri Aji. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Minghao Wu, Muhammad Abdul-Mageed, Alham Fikri Aji
EACL (1)4
2024 DetoxLLM: A Framework for Detoxification with Explanations
abstract
Prior works on detoxification are scattered in the sense that they do not cover all aspects of detoxification needed in a real-world scenario. Notably, prior works restrict the task of developing detoxification models to only a seen subset of platforms, leaving the question of how the models would perform on unseen platforms unexplored. Additionally, these works do not address non-detoxifiability, a phenomenon whereby the toxic text cannot be detoxified without altering the meaning. We propose DetoxLLM, the first comprehensive end-to-end detoxification framework, which attempts to alleviate the aforementioned limitations. We first introduce a cross-platform pseudo-parallel corpus applying multi-step data processing and generation strategies leveraging ChatGPT. We then train a suite of detoxification models with our cross-platform corpus. We show that our detoxification models outperform the SoTA model trained with human-annotated parallel corpus. We further introduce explanation to promote transparency and trustworthiness. DetoxLLM additionally offers a unique paraphrase detector especially dedicated for the detoxification task to tackle the non-detoxifiable cases. Through experimental analysis, we demonstrate the effectiveness of our cross-platform corpus and the robustness of DetoxLLM against adversarial toxicity.
Md. Tawkat Islam Khondaker, Muhammad Abdul-Mageed, Laks V. S. Lakshmanan
EMNLP2
2024 Casablanca: Data and Models for Multidialectal Arabic Speech Recognition
abstract
Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah, Chafei Mohamed Chafei, Ahmed Oumar El-Shangiti, Hiba Zayed, Mohamedou Cheikh Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdelrahman Mohamed, Abdellah El Mekki, El Moatez Billah Nagoudi, Benelhadj Djelloul Mama Saadia, Hamzah A. Alsayadi, Walid Al-Dhabyani, Sara Shatnawi, Yasir Ech-chammakhy, Amal Makouar, Yousra Berrachedi, Mustafa Jarrar, Shady Shehata, Ismail Berrada, Muhammad Abdul-Mageed. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah, Chafei Mohamed Chafei, Ahmed Oumar El-Shangiti, Hiba Zayed, Mohamedou Cheikh Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdel-rahman Mohamed, Abdellah El Mekki, El Moatez Billah Nagoudi, Benelhadj Saadia, Hamzah A. Alsayadi, Walid Al-Dhabyani, Sara Shatnawi, Yasir Ech-Chammakhy, Amal Makouar, Yousra Berrachedi, Mustafa Jarrar, Shady Shehata, Ismail Berrada, Muhammad Abdul-Mageed
EMNLP27
2024 What Does it Take to Generalize SER Model Across Datasets? A Comprehensive Benchmark
Adham Ibrahim, Shady Shehata, Ajinkya Kulkarni, Mukhtar Mohamed, Muhammad Abdul-Mageed
INTERSPEECH5
2024 Interplay of Machine Translation, Diacritics, and Diacritization
abstract
Wei-Rui Chen, Ife Adebara, Muhammad Abdul-Mageed. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Wei-Rui Chen, Ife Adebara, Muhammad Abdul-Mageed
NAACL-HLT3
2024 EmbSum: Leveraging the Summarization Capabilities of Large Language Models for Content-Based Recommendations
abstract
Content-based recommendation systems play a crucial role in delivering personalized content to users in the digital world. In this work, we introduce EmbSum, a novel framework that enables offline pre-computations of users and candidate items while capturing the interactions within the user engagement history. By utilizing the pretrained encoder-decoder model and poly-attention layers, EmbSum derives User Poly-Embedding (UPE) and Content Poly-Embedding (CPE) to calculate relevance scores between users and candidate items. EmbSum actively learns the long user engagement histories by generating user-interest summary with supervision from large language model (LLM). The effectiveness of EmbSum is validated on two datasets from different domains, surpassing state-of-the-art (SoTA) methods with higher accuracy and fewer parameters. Additionally, the model’s ability to generate summaries of user interests serves as a valuable by-product, enhancing its usefulness for personalized content recommendations.
Yifei Sun 0010, Minghao Wu, Jie Lei 0006, Muhammad Abdul-Mageed, Rong Jin 0001, Angli Liu, Sem Park, Bo Long
RecSys6
2023 GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP
abstract
ChatGPT's emergence heralds a transformative phase in NLP, particularly demonstrated through its excellent performance on many English benchmarks.However, the model's efficacy across diverse linguistic contexts remains largely uncharted territory.This work aims to bridge this knowledge gap, with a primary focus on assessing ChatGPT's capabilities on Arabic languages and dialectal varieties.Our comprehensive study conducts a largescale automated and human evaluation of ChatGPT, encompassing 44 distinct language understanding and generation tasks on over 60 different datasets.To our knowledge, this marks the first extensive performance analysis of ChatGPT's deployment in Arabic NLP.Our findings indicate that, despite its remarkable performance in English, ChatGPT is consistently surpassed by smaller models that have undergone finetuning on Arabic.We further undertake a meticulous comparison of ChatGPT and GPT-4's Modern Standard Arabic (MSA) and Dialectal Arabic (DA), unveiling the relative shortcomings of both models in handling Arabic dialects compared to MSA.Although we further explore and confirm the utility of employing GPT-4 as a potential alternative for human evaluation, our work adds to a growing body of research underscoring the limitations of ChatGPT.
Md. Tawkat Islam Khondaker, El Moatez Billah Nagoudi, Muhammad Abdul-Mageed
EMNLP4
2023 JASMINE: Arabic GPT Models for Few-Shot Learning
abstract
El Moatez Billah Nagoudi, Muhammad Abdul-Mageed, AbdelRahim Elmadany, Alcides Inciarte, Md Tawkat Islam Khondaker. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
El Moatez Billah Nagoudi, Muhammad Abdul-Mageed, AbdelRahim A. Elmadany, Alcides Alcoba Inciarte, Md. Tawkat Islam Khondaker
EMNLP2
2023 The Skipped Beat: A Study of Sociopragmatic Understanding in LLMs for 64 Languages
abstract
Instruction tuned large language models (LLMs), such as ChatGPT, demonstrate remarkable performance in a wide range of tasks.Despite numerous recent studies that examine the performance of instruction-tuned LLMs on various NLP benchmarks, there remains a lack of comprehensive investigation into their ability to understand cross-lingual sociopragmatic meaning (SM), i.e., meaning embedded within social and interactive contexts.This deficiency arises partly from SM not being adequately represented in any of the existing benchmarks.To address this gap, we present SPARROW, an extensive multilingual benchmark specifically designed for SM understanding.SPARROW comprises 169 datasets covering 13 task types across six primary categories (e.g., anti-social language detection, emotion recognition).SPAR-ROW datasets encompass 64 different languages originating from 12 language families representing 16 writing scripts.We evaluate the performance of various multilingual pretrained language models (e.g., mT5) and instruction-tuned LLMs (e.g., BLOOMZ, ChatGPT) on SPARROW through fine-tuning, zero-shot, and/or few-shot learning.Our comprehensive analysis reveals that existing opensource instruction tuned LLMs still struggle to understand SM across various languages, performing close to a random baseline in some cases.We also find that although Chat-GPT outperforms many LLMs, it still falls behind task-specific finetuned models with a gap of 12.19 SPARROW score.Our benchmark is available
Khai Duy Doan, Qisheng Liao, Muhammad Abdul-Mageed
EMNLP4
2023 PACT: Pretraining with Adversarial Contrastive Learning for Text Classification
abstract
Md Tawkat Islam Khondaker, Muhammad Abdul-Mageed, Laks Lakshmanan, V.S.. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Md. Tawkat Islam Khondaker, Muhammad Abdul-Mageed, Laks V. S. Lakshmanan
IJCNLP (1)2
2023 ProMap: Effective Bilingual Lexicon Induction via Language Model Prompting
abstract
Abdellah El Mekki, Muhammad Abdul-Mageed, ElMoatez Billah Nagoudi, Ismail Berrada, Ahmed Khoumsi. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Abdellah El Mekki, Muhammad Abdul-Mageed, El Moatez Billah Nagoudi, Ismail Berrada, Ahmed Khoumsi
IJCNLP (1)2
2023 On the Robustness of Arabic Speech Dialect Identification
Peter Sullivan, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed
INTERSPEECH3
2023 N-Shot Benchmarking of Whisper on Diverse Arabic Speech Recognition
Bashar Talafha, Muhammad Abdul-Mageed
INTERSPEECH3
2022 Towards Afrocentric NLP for African Languages: Where We Are and Where We Can Go
abstract
Aligning with ACL 2022 special Theme on "Language Diversity: from Low Resource to Endangered Languages", we discuss the major linguistic and sociopolitical challenges facing development of NLP technologies for African languages.Situating African languages in a typological framework, we discuss how the particulars of these languages can be harnessed.To facilitate future research, we also highlight current efforts, communities, venues, datasets, and tools.Our main objective is to motivate and advocate for an Afrocentric approach to technology development.With this in mind, we recommend what technologies to build and how to build, evaluate, and deploy them based on the needs of local African communities.
Ife Adebara, Muhammad Abdul-Mageed
ACL (1)2
2022 AraT5: Text-to-Text Transformers for Arabic Language Generation
abstract
Transfer learning with a unified Transformer framework (T5) that converts all language problems into a text-to-text format was recently proposed as a simple and effective transfer learning approach.Although a multilingual version of the T5 model (mT5) was also introduced, it is not clear how well it can fare on non-English tasks involving diverse data.To investigate this question, we apply mT5 on a language with a wide variety of dialects-Arabic.For evaluation, we introduce a novel benchmark for ARabic language GENeration (ARGEN), covering seven important tasks.For model comparison, we pre-train three powerful Arabic T5-style models and evaluate them on ARGEN.Although pre-trained with ∼ 49% less data, our new models perform significantly better than mT5 on all ARGEN tasks (in 52 out of 59 test sets) and set several new SOTAs.Our models also establish new SOTA on the recently-proposed, large Arabic language understanding evaluation benchmark ARLUE (Abdul-Mageed et al., 2021).Our models are publicly available.We also link to individual ARGEN datasets through our public repository. 1
El Moatez Billah Nagoudi, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed
ACL (1)3
2022 Linguistically-Motivated Yorùbá-English Machine Translation
abstract
Translating between languages where certain features are marked morphologically in one but absent or marked contextually in the other is an important test case for machine translation. When translating into English which marks (in)definiteness morphologically, from Yorùbá which uses bare nouns but marks these features contextually, ambiguities arise. In this work, we perform fine-grained analysis on how an SMT system compares with two NMT systems (BiLSTM and Transformer) when translating bare nouns in Yorùbá into English. We investigate how the systems what extent they identify BNs, correctly translate them, and compare with human translation patterns. We also analyze the type of errors each model makes and provide a linguistic description of these errors. We glean insights for evaluating model performance in low-resource settings. In translating bare nouns, our results show the transformer model outperforms the SMT and BiLSTM models for 4 categories, the BiLSTM outperforms the SMT model for 3 categories while the SMT outperforms the NMT models for 1 category.
Ife Adebara, Muhammad Abdul-Mageed, Miikka Silfverberg
COLING2
2022 AfroLID: A Neural Language Identification Tool for African Languages
abstract
Language identification (LID) is a crucial precursor for NLP, especially for mining web data.Problematically, most of the world's 7000+ languages today are not covered by LID technologies.We address this pressing issue for Africa by introducing AfroLID, a neural LID toolkit for 517 African languages and varieties.AfroLID exploits a multi-domain web dataset manually curated from across 14 language families utilizing five orthographic systems.When evaluated on our blind Test set, AfroLID achieves 95.89 F 1 -score.We also compare AfroLID to five existing LID tools that each cover a small number of African languages, finding it to outperform them on most languages.We further show the utility of AfroLID in the wild by testing it on the acutely under-served Twitter domain.Finally, we offer a number of controlled case studies and perform a linguistically-motivated error analysis that allow us to both showcase AfroLID's powerful capabilities and limitations.1
Ife Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed, Alcides Alcoba Inciarte
EMNLP3
2021 ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic
abstract
Muhammad Abdul-Mageed, AbdelRahim Elmadany, El Moatez Billah Nagoudi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Muhammad Abdul-Mageed, AbdelRahim A. Elmadany, El Moatez Billah Nagoudi
ACL/IJCNLP (1)1
2021 Mega-COV: A Billion-Scale Dataset of 100+ Languages for COVID-19
abstract
Muhammad Abdul-Mageed, AbdelRahim Elmadany, El Moatez Billah Nagoudi, Dinesh Pabbi, Kunal Verma, Rannie Lin. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Muhammad Abdul-Mageed, AbdelRahim A. Elmadany, El Moatez Billah Nagoudi, Dinesh Pabbi, Kunal Verma, Rannie Lin
EACL1
2021 Self-Training Pre-Trained Language Models for Zero- and Few-Shot Multi-Dialectal Arabic Sequence Labeling
abstract
A sufficient amount of annotated data is usually required to fine-tune pre-trained language models for downstream tasks.Unfortunately, attaining labeled data can be costly, especially for multiple language varieties and dialects.We propose to self-train pre-trained language models in zero-and few-shot scenarios to improve performance on data-scarce varieties using only resources from data-rich ones.We demonstrate the utility of our approach in the context of Arabic sequence labeling by using a language model fine-tuned on Modern Standard Arabic (MSA) only to predict named entities (NE) and part-of-speech (POS) tags on several dialectal Arabic (DA) varieties.We show that self-training is indeed powerful, improving zero-shot MSA-to-DA transfer by as large as 10% F 1 (NER) and 2% accuracy (POS tagging).We acquire even better performance in few-shot scenarios with limited amounts of labeled data.We conduct an ablation study and show that the performance boost observed directly results from training data augmentation possible with DA examples via self-training.This opens up opportunities for developing DA models exploiting only MSA resources.Our approach can also be extended to other languages and tasks. 1
Muhammad Khalifa, Muhammad Abdul-Mageed, Khaled Shaalan
EACL2
2020 Automatic Detection of Machine Generated Text: A Critical Survey
abstract
Text generative models (TGMs) excel in producing text that matches the style of human language reasonably well.Such TGMs can be misused by adversaries, e.g., by automatically generating fake news and fake product reviews that can look authentic and fool humans.Detectors that can distinguish text generated by TGM from human written text play a vital role in mitigating such misuse of TGMs.Recently, there has been a flurry of works from both natural language processing (NLP) and machine learning (ML) communities to build accurate detectors for English.Despite the importance of this problem, there is currently no work that surveys this fast-growing literature and introduces newcomers to important research challenges.In this work, we fill this void by providing a critical survey and review of this literature to facilitate a comprehensive understanding of this problem.We conduct an in-depth error analysis of the state-of-the-art detector and discuss research directions to guide future work in this exciting area.
Ganesh Jawahar, Muhammad Abdul-Mageed, Laks V. S. Lakshmanan
COLING2
2020 Toward Micro-Dialect Identification in Diaglossic and Code-Switched Environments
abstract
Although prediction of dialects is an important language processing task, with a wide range of applications, existing work is largely limited to coarse-grained varieties.Inspired by geolocation research, we propose the novel task of Micro-Dialect Identification (MDI) and introduce MARBERT, a new language model with striking abilities to predict a fine-grained variety (as small as that of a city) given a single, short message.For modeling, we offer a range of novel spatially and linguistically-motivated multi-task learning models.To showcase the utility of our models, we introduce a new, large-scale dataset of Arabic micro-varieties (low-resource) suited to our tasks.MARBERT predicts micro-dialects with 9.9% F 1 , ∼ 76× better than a majority class baseline.Our new language model also establishes new state-ofthe-art on several external tasks. 1
Muhammad Abdul-Mageed, AbdelRahim A. Elmadany, Lyle H. Ungar
EMNLP (1)1
2019 Deep Learning the EEG Manifold for Phonological Categorization from Active Thoughts
abstract
Speech-related Brain Computer Interfaces (BCI) aim primarily at finding an alternative vocal communication pathway for people with speaking disabilities. As a step towards full decoding of imagined speech from active thoughts, we present a BCI system for subject-independent classification of phonological categories exploiting a novel deep learning based hierarchical feature extraction scheme. To better capture the complex representation of high-dimensional electroencephalography (EEG) data, we compute the joint variability of EEG electrodes into a channel cross-covariance matrix. We then extract the spatio-temporal information encoded within the matrix using a mixed deep neural network strategy. Our model framework is composed of a convolutional neural network (CNN), a long-short term network (LSTM), and a deep autoencoder. We train the individual networks hierarchically, feeding their combined outputs in a final gradient boosting classification step. Our best models achieve an average accuracy of 77.9% across five different binary classification tasks, providing a significant 22.5% improvement over previous methods. As we also show visually, our work demonstrates that the speech imagery EEG possesses significant discriminative information about the intended articulatory movements responsible for natural speech synthesis.
Pramit Saha, Sidney S. Fels, Muhammad Abdul-Mageed
ICASSP3
2019 SPEAK YOUR MIND! Towards Imagined Speech Recognition with Hierarchical Deep Learning
abstract
Speech-related Brain Computer Interface (BCI) technologies provide effective vocal communication strategies for controlling devices through speech commands interpreted from brain signals. In order to infer imagined speech from active thoughts, we propose a novel hierarchical deep learning BCI system for subject-independent classification of 11 speech tokens including phonemes and words. Our novel approach exploits predicted articulatory information of six phonological categories (e.g., nasal, bilabial) as an intermediate step for classifying the phonemes and words, thereby finding discriminative signal responsible for natural speech synthesis. The proposed network is composed of hierarchical combination of spatial and temporal CNN cascaded with a deep autoencoder. Our best models on the KARA database achieve an average accuracy of 83.42% across the six different binary phonological classification tasks, and 53.36% for the individual token identification task, significantly outperforming our baselines. Ultimately, our work suggests the possible existence of a brain imagery footprint for the underlying articulatory movement related to different sounds that can be used to aid imagined speech decoding.
Pramit Saha, Muhammad Abdul-Mageed, Sidney S. Fels
INTERSPEECH2
2019 Modeling Arabic subjectivity and sentiment in lexical space
Muhammad Abdul-Mageed
Inf. Process. Manag.1
2018 You Tweet What You Speak: A City-Level Dataset of Arabic Dialects
Muhammad Abdul-Mageed, Hassan Alhuzali, Mohamed Elaraby
LREC1
2017 EmoNet: Fine-Grained Emotion Detection with Gated Recurrent Neural Networks
abstract
Accurate detection of emotion from natural language has applications ranging from building emotional chatbots to better understanding individuals and their lives.
Muhammad Abdul-Mageed, Lyle H. Ungar
ACL (1)1
2017 Recognizing Pathogenic Empathy in Social Media
Muhammad Abdul-Mageed, Anneke Buffone, Johannes C. Eichstaedt, Lyle H. Ungar
ICWSM1
2016 Does 'well-being' translate on Twitter?
abstract
Laura Smith, Salvatore Giorgi, Rishi Solanki, Johannes Eichstaedt, H. Andrew Schwartz, Muhammad Abdul-Mageed, Anneke Buffone, Lyle Ungar. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2016.
Salvatore Giorgi, Rishi Solanki, Johannes C. Eichstaedt, H. Andrew Schwartz, Muhammad Abdul-Mageed, Anneke Buffone, Lyle H. Ungar
EMNLP6
2014 SANA: A Large Scale Multi-Genre, Multi-Dialect Lexicon for Arabic Subjectivity and Sentiment Analysis
Muhammad Abdul-Mageed, Mona T. Diab
LREC1
2014 SAMAR: Subjectivity and sentiment analysis for Arabic social media
Muhammad Abdul-Mageed, Mona T. Diab, Sandra Kübler
Comput. Speech Lang.1
2012 AWATIF: A Multi-Genre Corpus for Modern Standard Arabic Subjectivity and Sentiment Analysis
Muhammad Abdul-Mageed, Mona T. Diab
LREC1
2011 Automatic Detection of Arabic Non-Anaphoric Pronouns for Improving Anaphora Resolution
abstract
Anaphora resolution is one of the most difficult tasks in NLP. The ability to identify non-referential pronouns before attempting an anaphora resolution task would be significant, since the system would not have to attempt resolving such pronouns and hence end up with fewer errors. In addition, the number of non-referential pronouns has been found to be non-trivial in many domains. The task of detecting non-referential pronouns could also be incorporated into a part-of-speech tagger or a parser, or treated as an initial step in semantic interpretation. In this article, I describe a machine learning method for identifying non-referential pronouns in an annotated subsegment of the Penn Arabic Treebank using three different feature settings. I achieve an accuracy of 97.22% with 52 different features extracted from a small window size of -5/+5 tokens surrounding each potentially non-referential pronoun.
Muhammad Abdul-Mageed
ACM Trans. Asian Lang. Inf. Process.1