Ahmed Abdelali

dblp:48/5636 · DBLP profile ↗
← Back
36ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0002-4160-8181ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Information retrieval · 79% Knowledge graphs · 20% Data mining · 0%
Artificial intelligence
4 papers
Language models and text generation · 60% Representation and self-supervised learning · 24% Machine translation · 9%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
AraEval: An Arabic Multi-Task Evaluation Suite for Large Language Models · EMNLP 2025
Information retrieval › search engines › semantic search
ontology-based retrieval
0.912025
OntologyRAG-Q: Resource Development and Benchmarking for Retrieval-Augmented Question Answering in Qur'anic Tafsir · EMNLP 2025
Knowledge graphs › ontology
ontology construction
0.912025
OntologyRAG-Q: Resource Development and Benchmarking for Retrieval-Augmented Question Answering in Qur'anic Tafsir · EMNLP 2025
Information retrieval
question answering and dialogue systems
0.912025
OntologyRAG-Q: Resource Development and Benchmarking for Retrieval-Augmented Question Answering in Qur'anic Tafsir · EMNLP 2025
Information retrieval
retrieval-augmented generation
0.912025
OntologyRAG-Q: Resource Development and Benchmarking for Retrieval-Augmented Question Answering in Qur'anic Tafsir · EMNLP 2025
Machine learning › Representation and self-supervised learning › representation matching › feature alignment
multilingual representation alignment
0.812024
Exploring Alignment in Shared Cross-lingual Spaces · ACL (1) 2024
Information retrieval
retrieval models
0.322025
OntologyRAG-Q: Resource Development and Benchmarking for Retrieval-Augmented Question Answering in Qur'anic Tafsir · EMNLP 2025
Cross-language information retrieval using PARAFAC2 · KDD 2007
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation
0.312025
AraEval: An Arabic Multi-Task Evaluation Suite for Large Language Models · EMNLP 2025
Information retrieval › retrieval models › neural retrieval
dense retrieval
0.312025
OntologyRAG-Q: Resource Development and Benchmarking for Retrieval-Augmented Question Answering in Qur'anic Tafsir · EMNLP 2025
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.212015
How to Avoid Unwanted Pregnancies: Domain Adaptation using Neural Network Models · EMNLP 2015
Natural language and speech › Machine translation
neural machine translation
0.212015
How to Avoid Unwanted Pregnancies: Domain Adaptation using Neural Network Models · EMNLP 2015
Information retrieval
cross-language information retrieval
0.122007
Cross-language information retrieval using PARAFAC2 · KDD 2007
Benefits of the 'Massively Parallel Rosetta Stone': Cross-Language Information Retrieval with over 30 Languages · ACL 2007
Natural language and speech › Machine translation
parallel corpora
0.112007
Benefits of the 'Massively Parallel Rosetta Stone': Cross-Language Information Retrieval with over 30 Languages · ACL 2007
Information retrieval › retrieval models › latent semantic models
latent semantic indexing
0.112007
Cross-language information retrieval using PARAFAC2 · KDD 2007
Natural language and speech › Information extraction and text analysis › multilingual NLP
multilingual language resources
0.012007
Benefits of the 'Massively Parallel Rosetta Stone': Cross-Language Information Retrieval with over 30 Languages · ACL 2007
Data mining
clustering
0.012007
Cross-language information retrieval using PARAFAC2 · KDD 2007
Information retrieval › document organization
multilingual document clustering
0.012007
Cross-language information retrieval using PARAFAC2 · KDD 2007

Methods — techniques the papers use, named apart from their topics

multi-task evaluation · 0.9large language model · 0.9embedding model · 0.9clustering · 0.8neural network joint model · 0.2cross-entropy regularization · 0.2parallel corpora · 0.1cross-language retrieval · 0.1latent semantic analysis · 0.1PARAFAC2 · 0.1
YearPublicationVenuePosition
2026 Arabic-based Agent for Authorship Style Transfer
abstract
Authorship style transfer focuses on modifying text to align with a specific author’s writing style while maintaining its original meaning. In this study, we investigate this task in the context of Arabic text. To the best of our knowledge, this is the first work addressing authorship style transfer in the Arabic domain. Our approach begins with the construction of a novel Arabic parallel dataset for this task. Subsequently, we adapt a set of language models to evaluate how well they can recognize authorship and rewrite input text in the stylistic voice of a target author without altering the underlying semantics. The evaluation includes three large language models (LLMs) alongside one non-LLM Arabic-specific model, allowing us to demonstrate that the presented framework remains effective across different model architectures. As an outcome of this research, we introduce an Arabic language agent capable of performing authorship style transfer with competitive results.
Shadi Abudalfa, Raed Mughaus, Ahmed Abdelali
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2025 OntologyRAG-Q: Resource Development and Benchmarking for Retrieval-Augmented Question Answering in Qur'anic Tafsir
abstract
This paper introduces essential resources for Qur’anic studies: an annotated Tafsir ontology, a dataset of approximately 4,200 question-answer pairs, and a collection of 15 structured Tafsir books available in two formats. We present a comprehensive framework for handling sensitive Qur’anic Tafsir data that spans the entire pipeline from dataset construction through evaluation and error analysis. Our work establishes new benchmarks for retrieval and question-answering tasks on Qur’anic content, comparing performance across state-of-the-art embedding models and large language models (LLMs).We introduce OntologyRAG-Q, a novel retrieval-augmented generation approach featuring our custom Ayat-Ontology chunking method that segments Tafsir content at the verse level using ontology-driven structure. Benchmarking reveals strong performance across various LLMs, with GPT-4 achieving the highest results, followed closely by ALLaM. Expert evaluations show our system achieves 69.52% accuracy and 74.36% correctness overall, though multi-hop and context-dependent questions remain challenging. Our analysis demonstrates that answer position within documents significantly impacts retrieval performance, and among the evaluation metrics tested, BERT-recall and BERT-F1 correlate most strongly with expert assessments. The resources developed in this study are publicly available at https://github.com/sazani/OntologyRAG-Q.git.
Sadam Al-Azani, Maad Alowaifeer, Alhanoof Alhunief, Ahmed Abdelali
EMNLP4
2025 AraEval: An Arabic Multi-Task Evaluation Suite for Large Language Models
abstract
Alhanoof Althnian, Norah A. Alzahrani, Shaykhah Z. Alsubaie, Eman Albilali, Ahmed Abdelali, Nouf M. Alotaibi, M Saiful Bari, Yazeed Alnumay, Abdulhamed Alothaimen, Maryam Saif, Shahad D. Alzaidi, Faisal Abdulrahman Mirza, Yousef Almushayqih, Mohammed Al Saleem, Ghadah Alabduljabbar, Abdulmohsen Al-Thubaity, Areeb Alowisheq, Nora Al-Twairesh. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Alhanoof Althnian, Norah A. Alzahrani, Shaykhah Alsubaie, Eman Albilali, Ahmed Abdelali, Nouf M. Alotaibi, Saiful Bari, Yazeed Alnumay, Abdulhamed Alothaimen, Maryam Saif, Shahad D. Alzaidi, Faisal Mirza, Yousef Almushayqih, Mohammed Al Saleem, Ghadah Alabduljabbar, AbdulMohsen Al-Thubaity, Areeb Alowisheq, Nora Al-Twairesh
EMNLP5
2024 Exploring Alignment in Shared Cross-lingual Spaces
abstract
Despite their remarkable ability to capture linguistic nuances across diverse languages, questions persist regarding the degree of alignment between languages in multilingual embeddings. Drawing inspiration from research on high-dimensional representations in neural language models, we employ clustering to uncover latent concepts within multilingual models. Our analysis focuses on quantifying the alignment and overlap of these concepts across various languages within the latent space. To this end, we introduce two metrics CALIGN and COLAP aimed at quantifying these aspects, enabling a deeper exploration of multilingual embeddings. Our study encompasses three multilingual models (mT5, mBERT, and XLM-R) and three downstream tasks (Machine Translation, Named Entity Recognition, and Sentiment Analysis). Key findings from our analysis include: i) deeper layers in the network demonstrate increased cross-lingual alignment due to the presence of language-agnostic concepts, ii) fine-tuning of the models enhances alignment within the latent space, and iii) such task-specific calibration helps in explaining the emergence of zero-shot capabilities in the models.
Basel Mousi, Nadir Durrani, Fahim Dalvi, Majd Hawasly, Ahmed Abdelali
ACL (1)5
2024 LAraBench: Benchmarking Arabic AI with Large Language Models
abstract
Ahmed Abdelali, Hamdy Mubarak, Shammur Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Youssef Elshahawy, Ahmed Ali, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Ahmed Abdelali, Hamdy Mubarak, Shammur Absar Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Yousseif Elshahawy, Ahmed Ali 0002, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam
EACL (1)1
2023 Towards Generalization of Machine Learning Models: A Case Study of Arabic Sentiment Analysis
abstract
The abundance of social media data in the Arab world, specifically on Twitter, enabled companies and entities to exploit such rich and beneficial data that could be mined and used to extract important information, including sentiments and opinions of people towards a topic or a merchandise. However, with this plenitude comes the issue of producing models that are able to deliver consistent outcomes when tested within various contexts. Although model generalization has been thoroughly investigated in many fields, it has not been heavily investigated in the Arabic context. To address this gap, we investigate the generalization of models and data in Arabic with application to sentiment analysis, by performing a battery of experiments and building different models that are tested on five independent test sets to understand their performance when presented with unseen data. In doing so, we detail different techniques that improve the generalization of machine learning models in Arabic sentiment analysis, and share a large versatile dataset consisting of approximately 1.64M Arabic tweets and their corresponding sentiment to be used for future research. Our experiments concluded that the most consistent model is trained using a dataset labelled by a cascaded approach of two models, one that labels neutral tweets and another that identifies positive/negative tweets based on the Arabic emoji lexicon after class balancing. Both the BERT and the SVM models trained using the refined data achieve an average F-1 score of 0.62 and 0.60, and standard deviation of 0.06 and 0.04 respectively, when evaluated on five diverse test sets, outperforming other models by at least 17% relative gain in F-1. Based on our experiments, we share recommendations to improve model generalization for classification tasks.
Samir Abdaljalil, Shaimaa Hassanein, Hamdy Mubarak, Ahmed Abdelali
ICWSM4
2022 Textual Data Augmentation for Arabic-English Code-Switching Speech Recognition
abstract
The pervasiveness of intra-utterance code-switching (CS) in spoken content requires that speech recognition (ASR) systems handle mixed language. Designing a CS-ASR system has many challenges, mainly due to data scarcity, grammatical structure complexity, and domain mismatch. The most common method for addressing CS is to train an ASR system with the available transcribed CS speech, along with monolingual data. In this work, we propose a zero-shot learning methodology for CS-ASR by augmenting the monolingual data with artificially generating CS text. We based our approach on random lexical replacements and Equivalence Constraint (EC) while exploiting aligned translation pairs to generate random and grammatically valid CS content. Our empirical results show a 65.5% relative reduction in language model perplexity, and 7.7% in ASR WER on two ecologically valid CS test sets. The human evaluation of the generated text using EC suggests that more than 80% is of adequate quality.
Amir Hussein, Shammur Absar Chowdhury, Ahmed Abdelali, Najim Dehak, Ahmed Ali 0002, Sanjeev Khudanpur
SLT3
2021 Fighting the COVID-19 Infodemic in Social Media: A Holistic Perspective and a Call to Arms
Firoj Alam, Fahim Dalvi, Shaden Shaar, Nadir Durrani, Hamdy Mubarak, Alex Nikolov, Giovanni Da San Martino, Ahmed Abdelali, Hassan Sajjad 0001, Kareem Darwish, Preslav Nakov
ICWSM8
2021 Towards One Model to Rule All: Multilingual Strategy for Dialectal Code-Switching Arabic ASR
abstract
With the advent of globalization, there is an increasing demand for multilingual automatic speech recognition (ASR), handling language and dialectal variation of spoken content. Recent studies show its efficacy over monolingual systems. In this study, we design a large multilingual end-to-end ASR using self-attention based conformer architecture. We trained the system using Arabic (Ar), English (En) and French (Fr) languages. We evaluate the system performance handling: (i) monolingual (Ar, En and Fr); (ii) multi-dialectal (Modern Standard Arabic, along with dialectal variation such as Egyptian and Moroccan); (iii) code-switching -- cross-lingual (Ar-En/Fr) and dialectal (MSA-Egyptian dialect) test cases, and compare with current state-of-the-art systems. Furthermore, we investigate the influence of different embedding/character representations including character vs word-piece; shared vs distinct input symbol per language. Our findings demonstrate the strength of such a model by outperforming state-of-the-art monolingual dialectal Arabic and code-switching Arabic ASR.
Shammur Absar Chowdhury, Amir Hussein, Ahmed Abdelali, Ahmed Ali 0002
Interspeech3
2021 Arabic Diacritic Recovery Using a Feature-rich biLSTM Model
abstract
Diacritics (short vowels) are typically omitted when writing Arabic text, and readers have to reintroduce them to correctly pronounce words. There are two types of Arabic diacritics: The first are core-word diacritics (CW), which specify the lexical selection, and the second are case endings (CE), which typically appear at the end of word stems and generally specify their syntactic roles. Recovering CEs is relatively harder than recovering core-word diacritics due to inter-word dependencies, which are often distant. In this article, we use feature-rich recurrent neural network model that use a variety of linguistic and surface-level features to recover both core word diacritics and case endings. Our model surpasses all previous state-of-the-art systems with a CW error rate (CWER) of 2.9% and a CE error rate (CEER) of 3.7% for Modern Standard Arabic (MSA) and CWER of 2.2% and CEER of 2.5% for Classical Arabic (CA). When combining diacritized word cores with case endings, the resultant word error rates are 6.0% and 4.3% for MSA and CA, respectively. This highlights the effectiveness of feature engineering for such deep neural models.
Kareem Darwish, Ahmed Abdelali, Hamdy Mubarak, Mohamed Eldesouki
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2020 AraBench: Benchmarking Dialectal Arabic-English Machine Translation
abstract
Low-resource machine translation suffers from the scarcity of training data and the unavailability of standard evaluation sets. While a number of research efforts target the former, the unavailability of evaluation benchmarks remain a major hindrance in tracking the progress in low-resource machine translation. In this paper, we introduce AraBench, an evaluation suite for dialectal Arabic to English machine translation. Compared to Modern Standard Arabic, Arabic dialects are challenging due to their spoken nature, non-standard orthography, and a large variation in dialectness. To this end, we pool together already available Dialectal Arabic-English resources and additionally build novel test sets. AraBench offers 4 coarse, 15 fine-grained and 25 city-level dialect categories, belonging to diverse genres, such as media, chat, religion and travel with varying level of dialectness. We report strong baselines using several training settings: fine-tuning, back-translation and data augmentation. The evaluation suite opens a wide range of research frontiers to push efforts in low-resource machine translation, particularly Arabic dialect translation. The evaluation suite and the dialectal system are publicly available for research purposes.
Hassan Sajjad 0001, Ahmed Abdelali, Nadir Durrani, Fahim Dalvi
COLING2
2020 A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection
abstract
Access to social media often enables users to engage in conversation with limited accountability. This allows a user to share their opinions and ideology, especially regarding public content, occasionally adopting offensive language. This may encourage hate crimes or cause mental harm to targeted individuals or groups. Hence, it is important to detect offensive comments in social media platforms. Typically, most studies focus on offensive commenting in one platform only, even though the problem of offensive language is observed across multiple platforms. Therefore, in this paper, we introduce and make publicly available a new dialectal Arabic news comment dataset, collected from multiple social media platforms, including Twitter, Facebook, and YouTube. We follow two-step crowd-annotator selection criteria for low-representative language annotation task in a crowdsourcing platform. Furthermore, we analyze the distinctive lexical content along with the use of emojis in offensive comments. We train and evaluate the classifiers using the annotated multi-platform dataset along with other publicly available data. Our results highlight the importance of multiple platform dataset for (a) cross-platform, (b) cross-domain, and (c) cross-dialect generalization of classifier performance.
Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Abdelali, Soon-Gyo Jung, Jim Jansen, Joni Salminen
LREC3
2020 Effective multi-dialectal arabic POS tagging
abstract
Abstract This work introduces robust multi-dialectal part of speech tagging trained on an annotated data set of Arabic tweets in four major dialect groups: Egyptian, Levantine, Gulf, and Maghrebi. We implement two different sequence tagging approaches. The first uses conditional random fields (CRFs), while the second combines word- and character-based representations in a deep neural network with stacked layers of convolutional and recurrent networks with a CRF output layer. We successfully exploit a variety of features that help generalize our models, such as Brown clusters and stem templates. Also, we develop robust joint models that tag multi-dialectal tweets and outperform uni-dialectal taggers. We achieve a combined accuracy of 92.4% across all dialects, with per dialect results ranging between 90.2% and 95.4%. We obtained the results using a train/dev/test split of 70/10/20 for a data set of 350 tweets per dialect.
Kareem Darwish, Hamdy Mubarak, Younes Samih, Ahmed Abdelali, Lluís Màrquez, Mohamed Eldesouki, Laura Kallmeyer
Nat. Lang. Eng.5
2019 The MGB-5 Challenge: Recognition and Dialect Identification of Dialectal Arabic Speech
abstract
This paper describes the fifth edition of the Multi-Genre Broadcast Challenge (MGB-5), an evaluation focused on Arabic speech recognition and dialect identification. MGB-5 extends the previous MGB-3 challenge in two ways: first it focuses on Moroccan Arabic speech recognition; second the granularity of the Arabic dialect identification task is increased from 5 dialect classes to 17, by collecting data from 17 Arabic speaking countries. Both tasks use YouTube recordings to provide a multi-genre multi-dialectal challenge in the wild. Moroccan speech transcription used about 13 hours of transcribed speech data, split across training, development, and test sets, covering 7-genres: comedy, cooking, family/kids, fashion, drama, sports, and science (TEDx). The fine-grained Arabic dialect identification data was collected from known YouTube channels from 17 Arabic countries. 3,000 hours of this data was released for training, and 57 hours for development and testing. The dialect identification data was divided into three sub-categories based on the segment duration: short (under 5 s), medium (5-20 s), and long (>20 s). Overall, 25 teams registered for the challenge, and 9 teams submitted systems for the two tasks. We outline the approaches adopted in each system and summarize the evaluation results.
Ahmed Ali 0002, Suwon Shon, Younes Samih, Hamdy Mubarak, Ahmed Abdelali, James R. Glass, Steve Renals, Khalid Choukri
ASRU5
2018 The WAW Corpus: The First Corpus of Interpreted Speeches and their Translations for English and Arabic
Ahmed Abdelali, Irina P. Temnikova, Samy Hedaya, Stephan Vogel
LREC1
2018 Part-of-Speech Tagging for Arabic Gulf Dialect Using Bi-LSTM
Randah Alharbi, Walid Magdy, Kareem Darwish, Ahmed Abdelali, Hamdy Mubarak
LREC4
2018 Multi-Dialect Arabic POS Tagging: A CRF Approach
Kareem Darwish, Hamdy Mubarak, Ahmed Abdelali, Mohamed Eldesouki, Younes Samih, Randah Alharbi, Walid Magdy, Laura Kallmeyer
LREC3
2017 Learning from Relatives: Unified Dialectal Arabic Segmentation
abstract
Younes Samih, Mohamed Eldesouki, Mohammed Attia, Kareem Darwish, Ahmed Abdelali, Hamdy Mubarak, Laura Kallmeyer. Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017). 2017.
Younes Samih, Mohamed Eldesouki, Kareem Darwish, Ahmed Abdelali, Hamdy Mubarak, Laura Kallmeyer
CoNLL5
2017 Domain adaptation using neural network joint model
Shafiq R. Joty, Nadir Durrani, Hassan Sajjad 0001, Ahmed Abdelali
Comput. Speech Lang.4
2016 A Deep Fusion Model for Domain Adaptation in Phrase-based MT
abstract
We present a novel fusion model for domain adaptation in Statistical Machine Translation. Our model is based on the joint source-target neural network Devlin et al., 2014, and is learned by fusing in- and out-domain models. The adaptation is performed by backpropagating errors from the output layer to the word embedding layer of each model, subsequently adjusting parameters of the composite model towards the in-domain data. On the standard tasks of translating English-to-German and Arabic-to-English TED talks, we observed average improvements of +0.9 and +0.7 BLEU points, respectively over a competition grade phrase-based system. We also demonstrate improvements over existing adaptation methods.
Nadir Durrani, Hassan Sajjad 0001, Shafiq R. Joty, Ahmed Abdelali
COLING4
2016 Arabic to English Person Name Transliteration using Twitter
Hamdy Mubarak, Ahmed Abdelali
LREC2
2016 Eyes Don't Lie: Predicting Machine Translation Quality Using Eye Movement
abstract
Hassan Sajjad, Francisco Guzmán, Nadir Durrani, Ahmed Abdelali, Houda Bouamor, Irina Temnikova, Stephan Vogel. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Hassan Sajjad 0001, Francisco Guzmán, Nadir Durrani, Ahmed Abdelali, Houda Bouamor, Irina P. Temnikova, Stephan Vogel
HLT-NAACL4
2015 How to Avoid Unwanted Pregnancies: Domain Adaptation using Neural Network Models
abstract
We present novel models for domain adaptation based on the neural network joint model (NNJM).Our models maximize the cross entropy by regularizing the loss function with respect to in-domain model.Domain adaptation is carried out by assigning higher weight to out-domain sequences that are similar to the in-domain data.In our alternative model we take a more restrictive approach by additionally penalizing sequences similar to the outdomain data.Our models achieve better perplexities than the baseline NNJM models and give improvements of up to 0.5 and 0.6 BLEU points in Arabic-to-English and English-to-German language pairs, on a standard task of translating TED talks.
Shafiq R. Joty, Hassan Sajjad 0001, Nadir Durrani, Kamla Al-Mannai, Ahmed Abdelali, Stephan Vogel
EMNLP5
2015 QAT2 - the QCRI advanced transcription and translation system
Ahmed Abdelali, Ahmed Ali 0002, Francisco Guzmán, Felix Stahlberg, Stephan Vogel
INTERSPEECH1
2015 Using joint models or domain adaptation in statistical machine translation
Nadir Durrani, Hassan Sajjad 0001, Shafiq R. Joty, Ahmed Abdelali, Stephan Vogel
MTSummit4
2014 Revisiting Arabic Part of Speech Tagsets
abstract
Assigning the appropriate grammatical category to a word given a context is very important step in major areas of natural language processing. A limited numbers of Part of Speech Taggers currently exist for Arabic. These taggers mainly adopt tagsets that were developed for languages such as English. In this paper we present an effort of proposing a revised categories for Arabic POS tags that would take into consideration the richness of the language as well as its compatibility for automatic processing.
Yahya O. Mohamed Elhadj, Ahmed Abdelali, Rachid Bouziane, Adel Ammar
AICCSA2
2014 The AMARA Corpus: Building Parallel Language Resources for the Educational Domain
Ahmed Abdelali, Francisco Guzmán, Hassan Sajjad 0001, Stephan Vogel
LREC1
2014 Using Stem-Templates to Improve Arabic POS and Gender/Number Tagging
Kareem Darwish, Ahmed Abdelali, Hamdy Mubarak
LREC2
2013 Toward an efficient Arabic Part of Speech Tagger
abstract
The task of tagging and allotting the correct Part of Speech (POS) to text given its context is not obvious and requires expertise and use of considerable resources. Automating such task and building tools that can carry such job is crucial and imperative to advance in major areas of natural language processing. A limited numbers of Part of Speech Taggers exist currently for Arabic and their availability is not trivial. In this paper we present an effort to design and build a POS tagger that would take into consideration the richness of the language as well as the efficiency in processing volumes of text. The Light Arabic Part of Speech Tagger (LAPOST) current output is very comparable to existing system but more effective from the processing perspective.
Ahmed Abdelali, Yahya O. Mohamed Elhadj, Rachid Bouziane
AICCSA1
2011 An information-theoretic, vector-space-model approach to cross-language information retrieval
abstract
Abstract In this article, we demonstrate several novel ways in which insights from information theory (IT) and computational linguistics (CL) can be woven into a vector-space-model (VSM) approach to information retrieval (IR). Our proposals focus, essentially, on three areas: pre-processing (morphological analysis), term weighting, and alternative geometrical models to the widely used term-by-document matrix. The latter include (1) PARAFAC2 decomposition of a term-by-document-by-language tensor, and (2) eigenvalue decomposition of a term-by-term matrix (inspired by Statistical Machine Translation). We evaluate all proposals, comparing them to a ‘standard’ approach based on Latent Semantic Analysis, on a multilingual document clustering task. The evidence suggests that proper consideration of IT within IR is indeed called for: in all cases, our best results are achieved using the information-theoretic variations upon the standard approach. Furthermore, we show that different information-theoretic options can be combined for still better results. A key function of language is to encode and convey information, and contributions of IT to the field of CL can be traced back a number of decades. We think that our proposals help bring IR and CL more into line with one another. In our conclusion, we suggest that the fact that our proposals yield empirical improvements is not coincidental given that they increase the theoretical transparency of VSM approaches to IR; on the contrary, they help shed light on why aspects of these approaches work as they do.
Peter A. Chew, Brett W. Bader, Stephen Helmreich, Ahmed Abdelali, Stephen J. Verzi
Nat. Lang. Eng.4
2008 Latent Morpho-Semantic Analysis: Multilingual Information Retrieval with Character N-Grams and Mutual Information
Peter A. Chew, Brett W. Bader, Ahmed Abdelali
COLING3
2008 The Effects of Language Relatedness on Multilingual Information Retrieval: A Case Study With Indo-European and Semitic Languages
Peter A. Chew, Ahmed Abdelali
IJCNLP2
2007 Benefits of the 'Massively Parallel Rosetta Stone': Cross-Language Information Retrieval with over 30 Languages
Peter A. Chew, Ahmed Abdelali
ACL2
2007 Cross-language information retrieval using PARAFAC2
abstract
A standard approach to cross-language information retrieval (CLIR) uses Latent Semantic Analysis (LSA) in conjunction with a multilingual parallel aligned corpus. This approach has been shown to be successful in identifying similar documents across languages - or more precisely, retrieving the most similar document in one language to a query in another language. However, the approach has severe drawbacks when applied to a related task, that of clustering documents "language-independently", so that documents about similar topics end up closest to one another in the semantic space regardless of their language. The problem is that documents are generally more similar to other documents in the same language than they are to documents in a different language, but on the same topic. As a result, when using multilingual LSA, documents will in practice cluster by language, not by topic.
Peter A. Chew, Brett W. Bader, Tamara G. Kolda, Ahmed Abdelali
KDD4
2007 Improving query precision using semantic expansion
Ahmed Abdelali, Jim Cowie, Hamdy S. Soliman
Inf. Process. Manag.1
2004 Localization in Modern Standard Arabic
abstract
Abstract Modern Standard Arabic (MSA) is the official language used in all Arabic countries. In this paper we describe an investigation of the uniformity of MSA across different countries. Many studies have been carried out locally or regionally on Arabic and its dialects. Here we look on a more global scale by studying language variations between countries. The source material used in this investigation was derived from national newspapers available on the Web, which provided samples of common media usage in each country. This corpus has been used to investigate the lexical characteristics of Modern Standard Arabic as found in 10 different Arabic speaking countries. We describe our collection methods, the types of lexical analysis performed, and the results of our investigations. With respect to newspaper articles, MSA seems to be very uniform across all the countries included in the study, but we have detected various types of differences, with implications for computational processing of MSA.
Ahmed Abdelali
J. Assoc. Inf. Sci. Technol.1