VLDB 2026 Research / reviewers in the wild / expert
Muhammad Abdul-Mageed
dblp:49/9389
· DBLP profile ↗
45ranked-venue papers
11as first author
32since 2021 · last 2026
0000-0002-8590-2040ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 9 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMsabstractAbdellah EL Mekki, Samar M. Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-hibiri, Razan Saadie, Hamzah A. Alsayadi, Nadia Ghezaiel Hammouda, Alshima Mohammed Alkhazimi, Aya Hamod, Al-Yas Yaqoob Al-Ghafri, Wesam El-Sayed, Asila Ismail al Sharji, Mohamad Ballout, Anas Belfathi, Karim Ghaddar, Serry Sibaee, Alaa Aoun, Aeej Mohammed Aseri, Lina Abureesh, Ahlam Bashiti, Majdal Yousef, Abdulaziz Hafiz, Yehdih Mohamed, Emira Hamedtou, Brakehe Emehah, Rahaf Alhamouri, Youssef Nafea, Aya El Aatar, Walid Al-Dhabyani, Emhemed S. Hamed, Sara Shatnawi, Fakhraddin Alwajih, Khalid Elkhidir, Ashwag Alasmari, Abdurrahman Gerrio, Omar Said Alshahri, AbdelRahim A. Elmadany, Ismail Berrada, Amir Azad Adli Al-kathiri, Fadi Zaraket, Mustafa Jarrar, Yahya Mohamed EL Hadj, Hassan Alhuzali, Muhammad Abdul-Mageed. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Abdellah El Mekki, Samar Mohamed Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-Hibiri, Razan Saadie, Hamzah A. Alsayadi, Nadia Ghezaiel Hammouda, Alshima Alkhazimi, Aya Hamod, Al-Yas Al-Ghafri, Wesam El-Sayed, Asila Al Sharji, Mohamad Ballout, Anas Belfathi, Karim Ghaddar, Serry Sibaee, Alaa Aoun, Aeej Mohammed Aseri, Lina Abureesh, Ahlam Bashiti, Majdal Yousef, Abdulaziz Hafiz, Yehdih Mohamed, Emira Hamedtou, Brakehe Emehah, Rahaf Alhamouri, Youssef Nafea, Aya El Aatar, Walid Al-Dhabyani, Emhemed Hamed, Sara Shatnawi, Fakhraddin Alwajih, Khalid Elkhidir, Ashwag Alasmari, Abdurrahman Gerrio, Omar Alshahri, AbdelRahim A. Elmadany, Ismail Berrada, Amir Azad Adli Alkathiri, Fadi A. Zaraket, Mustafa Jarrar, Yahya Mohamed El Hadj, Hassan Alhuzali, Muhammad Abdul-Mageed |
ACL (1) | 47 |
| 2026 | Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMsabstractLarge language models (LLMs) are widely deployed for open-ended communication, yet most bias evaluations still rely on English, classification-style tasks. We introduce \corpusname, a new multilingual, debate-style benchmark designed to reveal how narrative bias appears in realistic generative settings. Our dataset includes 8{,}400 structured debate prompts spanning four sensitive domains -- Women's Rights, Backwardness, Terrorism, and Religion -- across seven languages ranging from high-resource (English, Chinese) to low-resource (Swahili, Nigerian Pidgin). Using four flagship models (GPT-4o, Claude~3.5~Haiku, DeepSeek-Chat, and LLaMA-3-70B), we generate over 100{,}000 debate responses and automatically classify which demographic groups are assigned stereotyped versus modern roles. Results show that all models reproduce entrenched stereotypes despite safety alignment: Arabs are overwhelmingly linked to Terrorism and Religion ($\geq$89\%), Africans to socioeconomic ``backwardness'' (up to 77\%), and Western groups are consistently framed as modern or progressive. Biases grow sharply in lower-resource languages, revealing that alignment trained primarily in English does not generalize globally. Our findings highlight a persistent divide in multilingual fairness: current alignment methods reduce explicit toxicity but fail to prevent biased outputs in open-ended contexts. We release our \corpusname benchmark and analysis framework to support the next generation of multilingual bias evaluation and safer, culturally inclusive model alignment. Muhammed Saeed, Muhammad Abdul-Mageed, Shady Shehata |
LREC | 2 |
| 2025 | Where Are We? Evaluating LLM Performance on African LanguagesabstractIfe Adebara, Hawau Olamide Toyin, Nahom Tesfu Ghebremichael, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ife Adebara, Hawau Olamide Toyin, Nahom Tesfu Ghebremichael, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed |
ACL (1) | 5 |
| 2025 | Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMsabstractFakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany, Omer Nacar, El Moatez Billah Nagoudi, Reem Abdel-Salam, Hanin Atwany, Youssef Nafea, Abdulfattah Mohammed Yahya, Rahaf Alhamouri, Hamzah A. Alsayadi, Hiba Zayed, Sara Shatnawi, Serry Sibaee, Yasir Ech-chammakhy, Walid Al-Dhabyani, Marwa Mohamed Ali, Imen Jarraya, Ahmed Oumar El-Shangiti, Aisha Alraeesi, Mohammed Anwar AL-Ghrawi, Abdulrahman S. Al-Batati, Elgizouli Mohamed, Noha Taha Elgindi, Muhammed Saeed, Houdaifa Atou, Issam Ait Yahia, Abdelhak Bouayad, Mohammed Machrouh, Amal Makouar, Dania Alkawi, Mukhtar Mohamed, Safaa Taher Abdelfadil, Amine Ziad Ounnoughene, Anfel Rouabhia, Rwaa Assi, Ahmed Sorkatti, Mohamedou Cheikh Tourad, Anis Koubaa, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Fakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany, Omer Nacar, El Moatez Billah Nagoudi, Reem Abdel-Salam, Hanin Atwany, Youssef Nafea, Abdulfattah Mohammed Yahya, Rahaf Alhamouri, Hamzah A. Alsayadi, Hiba Zayed, Sara Shatnawi, Serry Sibaee, Yasir Ech-Chammakhy, Walid Al-Dhabyani, Marwa Mohamed Ali, Imen Jarraya, Ahmed Oumar El-Shangiti, Aisha Alraeesi, Mohammed Anwar Al-Ghrawi, Abdulrahman S. Al-Batati, Elgizouli Mohamed, Noha Taha Elgindi, Muhammed Saeed, Houdaifa Atou, Issam Ait Yahia, Abdelhak Bouayad, Mohammed Machrouh, Amal Makouar, Dania Alkawi, Mukhtar Mohamed, Safaa Taher Abdelfadil, Amine Ziad Ounnoughene, Rouabhia Anfel, Rwaa Assi, Ahmed Sorkatti, Mohamedou Cheikh Tourad, Anis Koubaa, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed |
ACL (1) | 44 |
| 2025 | Voice of a Continent: Mapping Africa's Speech Technology FrontierabstractAbdelRahim A. Elmadany, Sang Yun Kwon, Hawau Olamide Toyin, Alcides Alcoba Inciarte, Hanan Aldarmaki, Muhammad Abdul-Mageed. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. AbdelRahim A. Elmadany, Sang Yun Kwon, Hawau Olamide Toyin, Alcides Alcoba Inciarte, Hanan Aldarmaki, Muhammad Abdul-Mageed |
EMNLP | 6 |
| 2025 | NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local CommunitiesabstractTranslate to low-resource language LLMھﺗﺎﻛل؟ ﻣش ﻣﺻطﻔﻰ، ﯾﺎ إزﯾك أﺣﻣد: ده، اﻟﺷﺎرع أﻛل ﻣن ﺑﻘﻠﻖ اﻟﻣدام.ﺑﻛﻠم ﻛﻧت ﻣﻌﻠش، ﻣﺻطﻔﻰ: Abdellah El Mekki, Houdaifa Atou, Omer Nacar, Shady Shehata, Muhammad Abdul-Mageed |
EMNLP | 5 |
| 2025 | EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMsabstractLarge language models (LLMs) are transforming education by answering questions, explaining complex concepts, and generating content across a wide range of subjects.Despite strong performance on academic benchmarks, they often fail to tailor responses to students grade levels.This is a critical need in K-12 education, where age-appropriate vocabulary and explanation are essential for effective learning.Existing models frequently produce outputs that are too advanced or vague for younger learners, and there are no standardized benchmarks to evaluate their ability to adjust across cognitive and developmental stages.To address this gap, we introduce EDUADAPT, a benchmark of nearly 48k grade-labeled QA pairs across nine science subjects, spanning Grades 1-12 and grouped into four grade levels.We evaluate a diverse set of open-source LLMs on EDU-ADAPT and find that while larger models generally perform better, they still struggle with generating suitable responses for early-grade students (Grades 1-5).Our work presents the first dataset and evaluation framework for assessing grade-level adaptability in LLMs, aiming to foster more developmentally aligned educational AI systems through better training and prompting strategies.EDUADAPT code and datasets are publicly available. Numaan Naeem, Abdellah El Mekki, Muhammad Abdul-Mageed |
EMNLP | 3 |
| 2025 | JAWAHER: A Multidialectal Dataset of Arabic Proverbs for LLM BenchmarkingabstractSamar Mohamed Magdy, Sang Yun Kwon, Fakhraddin Alwajih, Safaa Taher Abdelfadil, Shady Shehata, Muhammad Abdul-Mageed. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Samar Mohamed Magdy, Sang Yun Kwon, Fakhraddin Alwajih, Safaa Taher Abdelfadil, Shady Shehata, Muhammad Abdul-Mageed |
NAACL (Long Papers) | 6 |
| 2025 | uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data RegimesabstractAbdul Waheed, Karima Kadaoui, Bhiksha Raj, Muhammad Abdul-Mageed. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Karima Kadaoui, Bhiksha Raj, Muhammad Abdul-Mageed |
NAACL (Long Papers) | 4 |
| 2024 | Cheetah: Natural Language Generation for 517 African LanguagesabstractLow-resource African languages pose unique challenges for natural language processing (NLP) tasks, including natural language generation (NLG).In this paper, we develop Cheetah, a massively multilingual NLG language model for African languages.Cheetah supports 517 African languages and language varieties, allowing us to address the scarcity of NLG resources and provide a solution to foster linguistic diversity.We demonstrate the effectiveness of Cheetah through comprehensive evaluations across six generation downstream tasks.In five of the six tasks, Cheetah significantly outperforms other models, showcasing its remarkable performance for generating coherent and contextually appropriate text in a wide range of African languages.We additionally conduct a detailed human evaluation to delve deeper into the linguistic capabilities of Cheetah.The findings of this study contribute to advancing NLP research in low-resource settings, enabling greater accessibility and inclusion for African languages in a rapidly expanding digital landscape.The GitHub repository for the Cheetah project is available at https://github.com Ife Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed |
ACL (1) | 3 |
| 2024 | Peacock: A Family of Arabic Multimodal Large Language Models and BenchmarksabstractFakhraddin Alwajih, El Moatez Billah Nagoudi, Gagan Bhatia, Abdelrahman Mohamed, Muhammad Abdul-Mageed. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Fakhraddin Alwajih, El Moatez Billah Nagoudi, Gagan Bhatia, Abdel-rahman Mohamed, Muhammad Abdul-Mageed |
ACL (1) | 5 |
| 2024 | To Distill or Not to Distill? On the Robustness of Robust Knowledge DistillationabstractArabic is known to present unique challenges for Automatic Speech Recognition (ASR).On one hand, its rich linguistic diversity and wide range of dialects complicate the development of robust, inclusive models.On the other, current multilingual ASR models are compute-intensive and lack proper comprehensive evaluations.In light of these challenges, we distill knowledge from large teacher models into smaller student variants that are more efficient.We also introduce a novel human-annotated dataset covering five under-represented Arabic dialects for evaluation.We further evaluate both our models and existing SoTA multilingual models on both standard available benchmarks and our new dialectal data.Our best-distilled model's overall performance (45.0%WER) surpasses that of a SoTA model twice its size (SeamlessM4T-large-v2, WER=47.0%) and its teacher model (Whisper-large-v2, WER=55.1%), and its average performance on our new dialectal data (56.9%WER) outperforms all other models.To gain more insight into the poor performance of these models on dialectal data, we conduct an error analysis and report the main types of errors the different models tend to make.The GitHub repository for the project is available at https: //github.com/UBC-NLP/distill-whisper-ar. Karima Kadaoui, Muhammad Abdul-Mageed |
ACL (1) | 3 |
| 2024 | LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale InstructionsabstractMinghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, Alham Fikri Aji. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Minghao Wu, Muhammad Abdul-Mageed, Alham Fikri Aji |
EACL (1) | 4 |
| 2024 | DetoxLLM: A Framework for Detoxification with ExplanationsabstractPrior works on detoxification are scattered in the sense that they do not cover all aspects of detoxification needed in a real-world scenario. Notably, prior works restrict the task of developing detoxification models to only a seen subset of platforms, leaving the question of how the models would perform on unseen platforms unexplored. Additionally, these works do not address non-detoxifiability, a phenomenon whereby the toxic text cannot be detoxified without altering the meaning. We propose DetoxLLM, the first comprehensive end-to-end detoxification framework, which attempts to alleviate the aforementioned limitations. We first introduce a cross-platform pseudo-parallel corpus applying multi-step data processing and generation strategies leveraging ChatGPT. We then train a suite of detoxification models with our cross-platform corpus. We show that our detoxification models outperform the SoTA model trained with human-annotated parallel corpus. We further introduce explanation to promote transparency and trustworthiness. DetoxLLM additionally offers a unique paraphrase detector especially dedicated for the detoxification task to tackle the non-detoxifiable cases. Through experimental analysis, we demonstrate the effectiveness of our cross-platform corpus and the robustness of DetoxLLM against adversarial toxicity. Md. Tawkat Islam Khondaker, Muhammad Abdul-Mageed, Laks V. S. Lakshmanan |
EMNLP | 2 |
| 2024 | Casablanca: Data and Models for Multidialectal Arabic Speech RecognitionabstractBashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah, Chafei Mohamed Chafei, Ahmed Oumar El-Shangiti, Hiba Zayed, Mohamedou Cheikh Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdelrahman Mohamed, Abdellah El Mekki, El Moatez Billah Nagoudi, Benelhadj Djelloul Mama Saadia, Hamzah A. Alsayadi, Walid Al-Dhabyani, Sara Shatnawi, Yasir Ech-chammakhy, Amal Makouar, Yousra Berrachedi, Mustafa Jarrar, Shady Shehata, Ismail Berrada, Muhammad Abdul-Mageed. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah, Chafei Mohamed Chafei, Ahmed Oumar El-Shangiti, Hiba Zayed, Mohamedou Cheikh Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdel-rahman Mohamed, Abdellah El Mekki, El Moatez Billah Nagoudi, Benelhadj Saadia, Hamzah A. Alsayadi, Walid Al-Dhabyani, Sara Shatnawi, Yasir Ech-Chammakhy, Amal Makouar, Yousra Berrachedi, Mustafa Jarrar, Shady Shehata, Ismail Berrada, Muhammad Abdul-Mageed |
EMNLP | 27 |
| 2024 | What Does it Take to Generalize SER Model Across Datasets? A Comprehensive Benchmark
Adham Ibrahim, Shady Shehata, Ajinkya Kulkarni, Mukhtar Mohamed, Muhammad Abdul-Mageed |
INTERSPEECH | 5 |
| 2024 | Interplay of Machine Translation, Diacritics, and DiacritizationabstractWei-Rui Chen, Ife Adebara, Muhammad Abdul-Mageed. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Wei-Rui Chen, Ife Adebara, Muhammad Abdul-Mageed |
NAACL-HLT | 3 |
| 2024 | EmbSum: Leveraging the Summarization Capabilities of Large Language Models for Content-Based RecommendationsabstractContent-based recommendation systems play a crucial role in delivering personalized content to users in the digital world. In this work, we introduce EmbSum, a novel framework that enables offline pre-computations of users and candidate items while capturing the interactions within the user engagement history. By utilizing the pretrained encoder-decoder model and poly-attention layers, EmbSum derives User Poly-Embedding (UPE) and Content Poly-Embedding (CPE) to calculate relevance scores between users and candidate items. EmbSum actively learns the long user engagement histories by generating user-interest summary with supervision from large language model (LLM). The effectiveness of EmbSum is validated on two datasets from different domains, surpassing state-of-the-art (SoTA) methods with higher accuracy and fewer parameters. Additionally, the model’s ability to generate summaries of user interests serves as a valuable by-product, enhancing its usefulness for personalized content recommendations. Yifei Sun 0010, Minghao Wu, Jie Lei 0006, Muhammad Abdul-Mageed, Rong Jin 0001, Angli Liu, Sem Park, Bo Long |
RecSys | 6 |
| 2023 | GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLPabstractChatGPT's emergence heralds a transformative phase in NLP, particularly demonstrated through its excellent performance on many English benchmarks.However, the model's efficacy across diverse linguistic contexts remains largely uncharted territory.This work aims to bridge this knowledge gap, with a primary focus on assessing ChatGPT's capabilities on Arabic languages and dialectal varieties.Our comprehensive study conducts a largescale automated and human evaluation of ChatGPT, encompassing 44 distinct language understanding and generation tasks on over 60 different datasets.To our knowledge, this marks the first extensive performance analysis of ChatGPT's deployment in Arabic NLP.Our findings indicate that, despite its remarkable performance in English, ChatGPT is consistently surpassed by smaller models that have undergone finetuning on Arabic.We further undertake a meticulous comparison of ChatGPT and GPT-4's Modern Standard Arabic (MSA) and Dialectal Arabic (DA), unveiling the relative shortcomings of both models in handling Arabic dialects compared to MSA.Although we further explore and confirm the utility of employing GPT-4 as a potential alternative for human evaluation, our work adds to a growing body of research underscoring the limitations of ChatGPT. Md. Tawkat Islam Khondaker, El Moatez Billah Nagoudi, Muhammad Abdul-Mageed |
EMNLP | 4 |
| 2023 | JASMINE: Arabic GPT Models for Few-Shot LearningabstractEl Moatez Billah Nagoudi, Muhammad Abdul-Mageed, AbdelRahim Elmadany, Alcides Inciarte, Md Tawkat Islam Khondaker. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. El Moatez Billah Nagoudi, Muhammad Abdul-Mageed, AbdelRahim A. Elmadany, Alcides Alcoba Inciarte, Md. Tawkat Islam Khondaker |
EMNLP | 2 |
| 2023 | The Skipped Beat: A Study of Sociopragmatic Understanding in LLMs for 64 LanguagesabstractInstruction tuned large language models (LLMs), such as ChatGPT, demonstrate remarkable performance in a wide range of tasks.Despite numerous recent studies that examine the performance of instruction-tuned LLMs on various NLP benchmarks, there remains a lack of comprehensive investigation into their ability to understand cross-lingual sociopragmatic meaning (SM), i.e., meaning embedded within social and interactive contexts.This deficiency arises partly from SM not being adequately represented in any of the existing benchmarks.To address this gap, we present SPARROW, an extensive multilingual benchmark specifically designed for SM understanding.SPARROW comprises 169 datasets covering 13 task types across six primary categories (e.g., anti-social language detection, emotion recognition).SPAR-ROW datasets encompass 64 different languages originating from 12 language families representing 16 writing scripts.We evaluate the performance of various multilingual pretrained language models (e.g., mT5) and instruction-tuned LLMs (e.g., BLOOMZ, ChatGPT) on SPARROW through fine-tuning, zero-shot, and/or few-shot learning.Our comprehensive analysis reveals that existing opensource instruction tuned LLMs still struggle to understand SM across various languages, performing close to a random baseline in some cases.We also find that although Chat-GPT outperforms many LLMs, it still falls behind task-specific finetuned models with a gap of 12.19 SPARROW score.Our benchmark is available Khai Duy Doan, Qisheng Liao, Muhammad Abdul-Mageed |
EMNLP | 4 |
| 2023 | PACT: Pretraining with Adversarial Contrastive Learning for Text ClassificationabstractMd Tawkat Islam Khondaker, Muhammad Abdul-Mageed, Laks Lakshmanan, V.S.. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Md. Tawkat Islam Khondaker, Muhammad Abdul-Mageed, Laks V. S. Lakshmanan |
IJCNLP (1) | 2 |
| 2023 | ProMap: Effective Bilingual Lexicon Induction via Language Model PromptingabstractAbdellah El Mekki, Muhammad Abdul-Mageed, ElMoatez Billah Nagoudi, Ismail Berrada, Ahmed Khoumsi. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Abdellah El Mekki, Muhammad Abdul-Mageed, El Moatez Billah Nagoudi, Ismail Berrada, Ahmed Khoumsi |
IJCNLP (1) | 2 |
| 2023 | On the Robustness of Arabic Speech Dialect Identification
Peter Sullivan, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed |
INTERSPEECH | 3 |
| 2023 | N-Shot Benchmarking of Whisper on Diverse Arabic Speech Recognition
Bashar Talafha, Muhammad Abdul-Mageed |
INTERSPEECH | 3 |
| 2022 | Towards Afrocentric NLP for African Languages: Where We Are and Where We Can GoabstractAligning with ACL 2022 special Theme on "Language Diversity: from Low Resource to Endangered Languages", we discuss the major linguistic and sociopolitical challenges facing development of NLP technologies for African languages.Situating African languages in a typological framework, we discuss how the particulars of these languages can be harnessed.To facilitate future research, we also highlight current efforts, communities, venues, datasets, and tools.Our main objective is to motivate and advocate for an Afrocentric approach to technology development.With this in mind, we recommend what technologies to build and how to build, evaluate, and deploy them based on the needs of local African communities. Ife Adebara, Muhammad Abdul-Mageed |
ACL (1) | 2 |
| 2022 | AraT5: Text-to-Text Transformers for Arabic Language GenerationabstractTransfer learning with a unified Transformer framework (T5) that converts all language problems into a text-to-text format was recently proposed as a simple and effective transfer learning approach.Although a multilingual version of the T5 model (mT5) was also introduced, it is not clear how well it can fare on non-English tasks involving diverse data.To investigate this question, we apply mT5 on a language with a wide variety of dialects-Arabic.For evaluation, we introduce a novel benchmark for ARabic language GENeration (ARGEN), covering seven important tasks.For model comparison, we pre-train three powerful Arabic T5-style models and evaluate them on ARGEN.Although pre-trained with ∼ 49% less data, our new models perform significantly better than mT5 on all ARGEN tasks (in 52 out of 59 test sets) and set several new SOTAs.Our models also establish new SOTA on the recently-proposed, large Arabic language understanding evaluation benchmark ARLUE (Abdul-Mageed et al., 2021).Our models are publicly available.We also link to individual ARGEN datasets through our public repository. 1 El Moatez Billah Nagoudi, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed |
ACL (1) | 3 |
| 2022 | Linguistically-Motivated Yorùbá-English Machine TranslationabstractTranslating between languages where certain features are marked morphologically in one but absent or marked contextually in the other is an important test case for machine translation. When translating into English which marks (in)definiteness morphologically, from Yorùbá which uses bare nouns but marks these features contextually, ambiguities arise. In this work, we perform fine-grained analysis on how an SMT system compares with two NMT systems (BiLSTM and Transformer) when translating bare nouns in Yorùbá into English. We investigate how the systems what extent they identify BNs, correctly translate them, and compare with human translation patterns. We also analyze the type of errors each model makes and provide a linguistic description of these errors. We glean insights for evaluating model performance in low-resource settings. In translating bare nouns, our results show the transformer model outperforms the SMT and BiLSTM models for 4 categories, the BiLSTM outperforms the SMT model for 3 categories while the SMT outperforms the NMT models for 1 category. Ife Adebara, Muhammad Abdul-Mageed, Miikka Silfverberg |
COLING | 2 |
| 2022 | AfroLID: A Neural Language Identification Tool for African LanguagesabstractLanguage identification (LID) is a crucial precursor for NLP, especially for mining web data.Problematically, most of the world's 7000+ languages today are not covered by LID technologies.We address this pressing issue for Africa by introducing AfroLID, a neural LID toolkit for 517 African languages and varieties.AfroLID exploits a multi-domain web dataset manually curated from across 14 language families utilizing five orthographic systems.When evaluated on our blind Test set, AfroLID achieves 95.89 F 1 -score.We also compare AfroLID to five existing LID tools that each cover a small number of African languages, finding it to outperform them on most languages.We further show the utility of AfroLID in the wild by testing it on the acutely under-served Twitter domain.Finally, we offer a number of controlled case studies and perform a linguistically-motivated error analysis that allow us to both showcase AfroLID's powerful capabilities and limitations.1 Ife Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed, Alcides Alcoba Inciarte |
EMNLP | 3 |
| 2021 | ARBERT & MARBERT: Deep Bidirectional Transformers for ArabicabstractMuhammad Abdul-Mageed, AbdelRahim Elmadany, El Moatez Billah Nagoudi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Muhammad Abdul-Mageed, AbdelRahim A. Elmadany, El Moatez Billah Nagoudi |
ACL/IJCNLP (1) | 1 |
| 2021 | Mega-COV: A Billion-Scale Dataset of 100+ Languages for COVID-19abstractMuhammad Abdul-Mageed, AbdelRahim Elmadany, El Moatez Billah Nagoudi, Dinesh Pabbi, Kunal Verma, Rannie Lin. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Muhammad Abdul-Mageed, AbdelRahim A. Elmadany, El Moatez Billah Nagoudi, Dinesh Pabbi, Kunal Verma, Rannie Lin |
EACL | 1 |
| 2021 | Self-Training Pre-Trained Language Models for Zero- and Few-Shot Multi-Dialectal Arabic Sequence LabelingabstractA sufficient amount of annotated data is usually required to fine-tune pre-trained language models for downstream tasks.Unfortunately, attaining labeled data can be costly, especially for multiple language varieties and dialects.We propose to self-train pre-trained language models in zero-and few-shot scenarios to improve performance on data-scarce varieties using only resources from data-rich ones.We demonstrate the utility of our approach in the context of Arabic sequence labeling by using a language model fine-tuned on Modern Standard Arabic (MSA) only to predict named entities (NE) and part-of-speech (POS) tags on several dialectal Arabic (DA) varieties.We show that self-training is indeed powerful, improving zero-shot MSA-to-DA transfer by as large as 10% F 1 (NER) and 2% accuracy (POS tagging).We acquire even better performance in few-shot scenarios with limited amounts of labeled data.We conduct an ablation study and show that the performance boost observed directly results from training data augmentation possible with DA examples via self-training.This opens up opportunities for developing DA models exploiting only MSA resources.Our approach can also be extended to other languages and tasks. 1 Muhammad Khalifa, Muhammad Abdul-Mageed, Khaled Shaalan |
EACL | 2 |
| 2020 | Automatic Detection of Machine Generated Text: A Critical SurveyabstractText generative models (TGMs) excel in producing text that matches the style of human language reasonably well.Such TGMs can be misused by adversaries, e.g., by automatically generating fake news and fake product reviews that can look authentic and fool humans.Detectors that can distinguish text generated by TGM from human written text play a vital role in mitigating such misuse of TGMs.Recently, there has been a flurry of works from both natural language processing (NLP) and machine learning (ML) communities to build accurate detectors for English.Despite the importance of this problem, there is currently no work that surveys this fast-growing literature and introduces newcomers to important research challenges.In this work, we fill this void by providing a critical survey and review of this literature to facilitate a comprehensive understanding of this problem.We conduct an in-depth error analysis of the state-of-the-art detector and discuss research directions to guide future work in this exciting area. Ganesh Jawahar, Muhammad Abdul-Mageed, Laks V. S. Lakshmanan |
COLING | 2 |
| 2020 | Toward Micro-Dialect Identification in Diaglossic and Code-Switched EnvironmentsabstractAlthough prediction of dialects is an important language processing task, with a wide range of applications, existing work is largely limited to coarse-grained varieties.Inspired by geolocation research, we propose the novel task of Micro-Dialect Identification (MDI) and introduce MARBERT, a new language model with striking abilities to predict a fine-grained variety (as small as that of a city) given a single, short message.For modeling, we offer a range of novel spatially and linguistically-motivated multi-task learning models.To showcase the utility of our models, we introduce a new, large-scale dataset of Arabic micro-varieties (low-resource) suited to our tasks.MARBERT predicts micro-dialects with 9.9% F 1 , ∼ 76× better than a majority class baseline.Our new language model also establishes new state-ofthe-art on several external tasks. 1 Muhammad Abdul-Mageed, AbdelRahim A. Elmadany, Lyle H. Ungar |
EMNLP (1) | 1 |
| 2019 | Deep Learning the EEG Manifold for Phonological Categorization from Active ThoughtsabstractSpeech-related Brain Computer Interfaces (BCI) aim primarily at finding an alternative vocal communication pathway for people with speaking disabilities. As a step towards full decoding of imagined speech from active thoughts, we present a BCI system for subject-independent classification of phonological categories exploiting a novel deep learning based hierarchical feature extraction scheme. To better capture the complex representation of high-dimensional electroencephalography (EEG) data, we compute the joint variability of EEG electrodes into a channel cross-covariance matrix. We then extract the spatio-temporal information encoded within the matrix using a mixed deep neural network strategy. Our model framework is composed of a convolutional neural network (CNN), a long-short term network (LSTM), and a deep autoencoder. We train the individual networks hierarchically, feeding their combined outputs in a final gradient boosting classification step. Our best models achieve an average accuracy of 77.9% across five different binary classification tasks, providing a significant 22.5% improvement over previous methods. As we also show visually, our work demonstrates that the speech imagery EEG possesses significant discriminative information about the intended articulatory movements responsible for natural speech synthesis. Pramit Saha, Sidney S. Fels, Muhammad Abdul-Mageed |
ICASSP | 3 |
| 2019 | SPEAK YOUR MIND! Towards Imagined Speech Recognition with Hierarchical Deep LearningabstractSpeech-related Brain Computer Interface (BCI) technologies provide effective vocal communication strategies for controlling devices through speech commands interpreted from brain signals. In order to infer imagined speech from active thoughts, we propose a novel hierarchical deep learning BCI system for subject-independent classification of 11 speech tokens including phonemes and words. Our novel approach exploits predicted articulatory information of six phonological categories (e.g., nasal, bilabial) as an intermediate step for classifying the phonemes and words, thereby finding discriminative signal responsible for natural speech synthesis. The proposed network is composed of hierarchical combination of spatial and temporal CNN cascaded with a deep autoencoder. Our best models on the KARA database achieve an average accuracy of 83.42% across the six different binary phonological classification tasks, and 53.36% for the individual token identification task, significantly outperforming our baselines. Ultimately, our work suggests the possible existence of a brain imagery footprint for the underlying articulatory movement related to different sounds that can be used to aid imagined speech decoding. Pramit Saha, Muhammad Abdul-Mageed, Sidney S. Fels |
INTERSPEECH | 2 |
| 2019 | Modeling Arabic subjectivity and sentiment in lexical space
Muhammad Abdul-Mageed |
Inf. Process. Manag. | 1 |
| 2018 | You Tweet What You Speak: A City-Level Dataset of Arabic Dialects
Muhammad Abdul-Mageed, Hassan Alhuzali, Mohamed Elaraby |
LREC | 1 |
| 2017 | EmoNet: Fine-Grained Emotion Detection with Gated Recurrent Neural NetworksabstractAccurate detection of emotion from natural language has applications ranging from building emotional chatbots to better understanding individuals and their lives. Muhammad Abdul-Mageed, Lyle H. Ungar |
ACL (1) | 1 |
| 2017 | Recognizing Pathogenic Empathy in Social Media
Muhammad Abdul-Mageed, Anneke Buffone, Johannes C. Eichstaedt, Lyle H. Ungar |
ICWSM | 1 |
| 2016 | Does 'well-being' translate on Twitter?abstractLaura Smith, Salvatore Giorgi, Rishi Solanki, Johannes Eichstaedt, H. Andrew Schwartz, Muhammad Abdul-Mageed, Anneke Buffone, Lyle Ungar. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2016. Salvatore Giorgi, Rishi Solanki, Johannes C. Eichstaedt, H. Andrew Schwartz, Muhammad Abdul-Mageed, Anneke Buffone, Lyle H. Ungar |
EMNLP | 6 |
| 2014 | SANA: A Large Scale Multi-Genre, Multi-Dialect Lexicon for Arabic Subjectivity and Sentiment Analysis
Muhammad Abdul-Mageed, Mona T. Diab |
LREC | 1 |
| 2014 | SAMAR: Subjectivity and sentiment analysis for Arabic social media
Muhammad Abdul-Mageed, Mona T. Diab, Sandra Kübler |
Comput. Speech Lang. | 1 |
| 2012 | AWATIF: A Multi-Genre Corpus for Modern Standard Arabic Subjectivity and Sentiment Analysis
Muhammad Abdul-Mageed, Mona T. Diab |
LREC | 1 |
| 2011 | Automatic Detection of Arabic Non-Anaphoric Pronouns for Improving Anaphora ResolutionabstractAnaphora resolution is one of the most difficult tasks in NLP. The ability to identify non-referential pronouns before attempting an anaphora resolution task would be significant, since the system would not have to attempt resolving such pronouns and hence end up with fewer errors. In addition, the number of non-referential pronouns has been found to be non-trivial in many domains. The task of detecting non-referential pronouns could also be incorporated into a part-of-speech tagger or a parser, or treated as an initial step in semantic interpretation. In this article, I describe a machine learning method for identifying non-referential pronouns in an annotated subsegment of the Penn Arabic Treebank using three different feature settings. I achieve an accuracy of 97.22% with 52 different features extracted from a small window size of -5/+5 tokens surrounding each potentially non-referential pronoun. Muhammad Abdul-Mageed |
ACM Trans. Asian Lang. Inf. Process. | 1 |