VLDB 2026 Research / reviewers in the wild / expert
Mustafa Jarrar
dblp:j/MustafaJarrar
· DBLP profile ↗
27ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0003-4351-4207ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 13 since 2021Databases, data management, data science and information retrieval · 11 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMsabstractAbdellah EL Mekki, Samar M. Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-hibiri, Razan Saadie, Hamzah A. Alsayadi, Nadia Ghezaiel Hammouda, Alshima Mohammed Alkhazimi, Aya Hamod, Al-Yas Yaqoob Al-Ghafri, Wesam El-Sayed, Asila Ismail al Sharji, Mohamad Ballout, Anas Belfathi, Karim Ghaddar, Serry Sibaee, Alaa Aoun, Aeej Mohammed Aseri, Lina Abureesh, Ahlam Bashiti, Majdal Yousef, Abdulaziz Hafiz, Yehdih Mohamed, Emira Hamedtou, Brakehe Emehah, Rahaf Alhamouri, Youssef Nafea, Aya El Aatar, Walid Al-Dhabyani, Emhemed S. Hamed, Sara Shatnawi, Fakhraddin Alwajih, Khalid Elkhidir, Ashwag Alasmari, Abdurrahman Gerrio, Omar Said Alshahri, AbdelRahim A. Elmadany, Ismail Berrada, Amir Azad Adli Al-kathiri, Fadi Zaraket, Mustafa Jarrar, Yahya Mohamed EL Hadj, Hassan Alhuzali, Muhammad Abdul-Mageed. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Abdellah El Mekki, Samar Mohamed Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-Hibiri, Razan Saadie, Hamzah A. Alsayadi, Nadia Ghezaiel Hammouda, Alshima Alkhazimi, Aya Hamod, Al-Yas Al-Ghafri, Wesam El-Sayed, Asila Al Sharji, Mohamad Ballout, Anas Belfathi, Karim Ghaddar, Serry Sibaee, Alaa Aoun, Aeej Mohammed Aseri, Lina Abureesh, Ahlam Bashiti, Majdal Yousef, Abdulaziz Hafiz, Yehdih Mohamed, Emira Hamedtou, Brakehe Emehah, Rahaf Alhamouri, Youssef Nafea, Aya El Aatar, Walid Al-Dhabyani, Emhemed Hamed, Sara Shatnawi, Fakhraddin Alwajih, Khalid Elkhidir, Ashwag Alasmari, Abdurrahman Gerrio, Omar Alshahri, AbdelRahim A. Elmadany, Ismail Berrada, Amir Azad Adli Alkathiri, Fadi A. Zaraket, Mustafa Jarrar, Yahya Mohamed El Hadj, Hassan Alhuzali, Muhammad Abdul-Mageed |
ACL (1) | 44 |
| 2026 | AdabNER: Arabic Digital Archive Books with Nested Entity RecognitionabstractMost studies on Arabic Named Entity Recognition (NER) have focused on news texts and social media posts, while the large and rich corpus of literary Arabic books has been underrepresented.We introduce AdabNer, the first large-scale nested NER dataset for Modern Standard Arabic (MSA) literary texts, comprising the first 6,000 words annotated from each of 138 books spanning ten literary genres, including history, biography, literary criticism, and travel literature, and covering works from the 1880s to the 2020s.The corpus comprises about 876K tokens, manually annotated using a nested 21 entity tag annotation scheme, yielding 78,530 entity mentions, 18.96% of which are nested.We fine-tuned five pre-trained Arabic BERT encoders in two settings: stratified and leavebook-out, achieving F 1 scores of 0.86 and 0.83 with AraBERTv2, respectively.We also evaluated five large language models through fewshot in-context learning, including open-source models and the closed-source Gemini 3 Pro, with Gemini 3 Pro achieving the highest LLM F 1 score of 0.59.Supervised results degraded under out-of-domain evaluation; however, joint multi-domain training reduced this gap to less than a 1% F 1 loss, demonstrating that domaindiverse training data is key to robust Arabic NER, though broader validation beyond the experiments reported is needed.AdabNer and its annotation guidelines are publicly available. Aya Mourad, Mustafa Jarrar |
ACL (1) | 2 |
| 2026 | AraREQ: A Dataset and End-to-End System for Conflict Detection and Resolution in Software Requirements
Tymaa Hammouda, Alaa Aljabari, Nagham Hamad, Mustafa Jarrar |
LREC | 4 |
| 2025 | Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMsabstractFakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany, Omer Nacar, El Moatez Billah Nagoudi, Reem Abdel-Salam, Hanin Atwany, Youssef Nafea, Abdulfattah Mohammed Yahya, Rahaf Alhamouri, Hamzah A. Alsayadi, Hiba Zayed, Sara Shatnawi, Serry Sibaee, Yasir Ech-chammakhy, Walid Al-Dhabyani, Marwa Mohamed Ali, Imen Jarraya, Ahmed Oumar El-Shangiti, Aisha Alraeesi, Mohammed Anwar AL-Ghrawi, Abdulrahman S. Al-Batati, Elgizouli Mohamed, Noha Taha Elgindi, Muhammed Saeed, Houdaifa Atou, Issam Ait Yahia, Abdelhak Bouayad, Mohammed Machrouh, Amal Makouar, Dania Alkawi, Mukhtar Mohamed, Safaa Taher Abdelfadil, Amine Ziad Ounnoughene, Anfel Rouabhia, Rwaa Assi, Ahmed Sorkatti, Mohamedou Cheikh Tourad, Anis Koubaa, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Fakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany, Omer Nacar, El Moatez Billah Nagoudi, Reem Abdel-Salam, Hanin Atwany, Youssef Nafea, Abdulfattah Mohammed Yahya, Rahaf Alhamouri, Hamzah A. Alsayadi, Hiba Zayed, Sara Shatnawi, Serry Sibaee, Yasir Ech-Chammakhy, Walid Al-Dhabyani, Marwa Mohamed Ali, Imen Jarraya, Ahmed Oumar El-Shangiti, Aisha Alraeesi, Mohammed Anwar Al-Ghrawi, Abdulrahman S. Al-Batati, Elgizouli Mohamed, Noha Taha Elgindi, Muhammed Saeed, Houdaifa Atou, Issam Ait Yahia, Abdelhak Bouayad, Mohammed Machrouh, Amal Makouar, Dania Alkawi, Mukhtar Mohamed, Safaa Taher Abdelfadil, Amine Ziad Ounnoughene, Rouabhia Anfel, Rwaa Assi, Ahmed Sorkatti, Mohamedou Cheikh Tourad, Anis Koubaa, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed |
ACL (1) | 42 |
| 2025 | WojoodRelations: Arabic Relation Extraction Corpus and ModelingabstractRelation extraction (RE) is a core task in natural language processing, crucial for semantic understanding, knowledge graph construction, and enhancing downstream applications.Existing work on Arabic RE remains limited due to the language's rich morphology and syntactic complexity, and the lack of large, highquality datasets.In this paper, we present Wojood Relations , the largest and most diverse Arabic RE corpus to date, containing over 33K sentences (∼ 550K tokens) annotated with ∼ 15K relation triples across 40 relation types.The corpus is built on top of Wojood NER dataset with manual relation annotations carried out by expert annotators, achieving a Cohen's κ of 0.92, indicating high reliability.In addition, we propose two methods: NLI-RE, which formulates RE as a binary natural language inference problem using relation-aware templates, and GPT-Joint, a few-shot LLM framework for joint entity and RE via relationaware retrieval.Finally, we benchmark the dataset using both supervised models and incontext learning with LLMs.Supervised models achieve 92.89% F1 for RE, while LLMs obtain 72.73% F1 for joint entity and RE.These results establish strong baselines, highlight key challenges, and provide a foundation for advancing Arabic RE research. Alaa Aljabari, Mohammad Khalilia, Mustafa Jarrar |
EMNLP | 3 |
| 2024 | Qabas: An Open-Source Arabic Lexicographic DatabaseabstractWe present Qabas, a novel open-source Arabic lexicon designed for NLP applications. The novelty of Qabas lies in its synthesis of 110 lexicons. Specifically, Qabas lexical entries (lemmas) are assembled by linking lemmas from 110 lexicons. Furthermore, Qabas lemmas are also linked to 12 morphologically annotated corpora (about 2M tokens), making it the first Arabic lexicon to be linked to lexicons and corpora. Qabas was developed semi-automatically, utilizing a mapping framework and a web-based tool. Compared with other lexicons, Qabas stands as the most extensive Arabic lexicon, encompassing about 58K lemmas (45K nominal lemmas, 12.5K verbal lemmas, and 473 functional-word lemmas). Qabas is open-source and accessible online at https://sina.birzeit.edu/qabas Mustafa Jarrar, Tymaa Hammouda |
LREC/COLING | 1 |
| 2024 | Casablanca: Data and Models for Multidialectal Arabic Speech RecognitionabstractBashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah, Chafei Mohamed Chafei, Ahmed Oumar El-Shangiti, Hiba Zayed, Mohamedou Cheikh Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdelrahman Mohamed, Abdellah El Mekki, El Moatez Billah Nagoudi, Benelhadj Djelloul Mama Saadia, Hamzah A. Alsayadi, Walid Al-Dhabyani, Sara Shatnawi, Yasir Ech-chammakhy, Amal Makouar, Yousra Berrachedi, Mustafa Jarrar, Shady Shehata, Ismail Berrada, Muhammad Abdul-Mageed. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah, Chafei Mohamed Chafei, Ahmed Oumar El-Shangiti, Hiba Zayed, Mohamedou Cheikh Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdel-rahman Mohamed, Abdellah El Mekki, El Moatez Billah Nagoudi, Benelhadj Saadia, Hamzah A. Alsayadi, Walid Al-Dhabyani, Sara Shatnawi, Yasir Ech-Chammakhy, Amal Makouar, Yousra Berrachedi, Mustafa Jarrar, Shady Shehata, Ismail Berrada, Muhammad Abdul-Mageed |
EMNLP | 24 |
| 2023 | Offensive Hebrew Corpus and Detection using BERTabstractOffensive language detection has been well studied in many languages, but it is lagging behind in low-resource languages, such as Hebrew. In this paper, we present a new offensive language corpus in Hebrew. A total of 15,881 tweets were retrieved from Twitter. Each was labeled with one or more of five classes (abusive, hate, violence, pornographic, or none offensive) by Arabic-Hebrew bilingual speakers. The annotation process was challenging as each annotator is expected to be familiar with the Israeli culture, politics, and practices to understand the context of each tweet. We fine-tuned two Hebrew BERT models, HeBERT and AlephBERT, using our proposed dataset and another published dataset. We observed that our data boosts HeBERT performance by 2% when combined with DOLaH. Fine-tuning AlephBERT on our data and testing on DOLaHyields 69% accuracy, while fine-tuning on DOLaHand testing on our data yields 57% accuracy, which may be an indication to the generalizability our data offers. Our dataset and fine-tuned models are available on GitHub and Huggingface. Nagham Hamad, Mustafa Jarrar, Mohammad Khalilia, Nadim Nashif |
AICCSA | 2 |
| 2023 | Lisan: Yemeni, Iraqi, Libyan, and Sudanese Arabic Dialect Corpora with Morphological AnnotationsabstractThis article presents morphologically-annotated Yemeni, Sudanese, Iraqi, and Libyan Arabic dialects (${\text{L}\hat{\text{i}}\text{sa}\bar{\text{n}}}$) corpora. ${\text{L}\hat{\text{i}}\text{sa}\bar{\text{n}}}$ features around 1.2 million tokens. We collected the content of the corpora from several social media platforms. The Yemeni corpus ($\tilde 1.05{\text{M}}$ tokens) was collected automatically from Twitter. The corpora of the other three dialects ($\tilde 50{\text{K}}$ tokens each) was manually collected from Facebook and YouTube posts and comments. Thirty-five (35) annotators who are native speakers of the target dialects carried out the annotations. The annotators segmented all words in the four corpora into prefixes, stems and suffixes and labeled each with different morphological features such as part of speech, lemma, and a gloss in English. We developed the Arabic Dialect Annotation Toolkit (ADAT) to assist the annotators and to ensure compatibility with SAMA and Curras tagsets. We trained annotators on a set of guidelines and on how to use ADAT. ADAT is open source, and the four corpora are available at https://sina.birzeit.edu/currasat. Mustafa Jarrar, Fadi A. Zaraket, Tymaa Hammouda, Daanish Masood Alavi, Martin Wählisch |
AICCSA | 1 |
| 2023 | Open-Source Thesaurus Development for Under-Resourced Languages: a Welsh Case Study
Nouran Khallaf, Elin Arfon, Mo El-Haj, Jonathan Morris, Dawn Knight, Paul Rayson, Tymaa Hammouda, Mustafa Jarrar |
LDK | 8 |
| 2023 | A Benchmark and Scoring Algorithm for Enriching Arabic SynonymsabstractThis paper addresses the task of extending a given synset with additional synonyms taking into account synonymy strength as a fuzzy value.Given a mono/multilingual synset and a threshold (a fuzzy value [0 -1]), our goal is to extract new synonyms above this threshold from existing lexicons.We present twofold contributions: an algorithm and a benchmark dataset.The dataset consists of 3K candidate synonyms for 500 synsets.Each candidate synonym is annotated with a fuzzy value by four linguists.The dataset is important for (i) understanding how much linguists (dis/)agree on synonymy, in addition to (ii) using the dataset as a baseline to evaluate our algorithm.Our proposed algorithm extracts synonyms from existing lexicons and computes a fuzzy value for each candidate.Our evaluations show that the algorithm behaves like a linguist and its fuzzy values are close to those proposed by linguists (using RMSE and MAE).The dataset and a demo page are publicly available at https: //portal.sina.birzeit.edu/synonyms. Sana Ghanem, Mustafa Jarrar, Radi Jarrar, Ibrahim Bounhas |
GWC | 2 |
| 2023 | Context-Gloss Augmentation for Improving Arabic Target Sense VerificationabstractArabic language lacks semantic datasets and sense inventories.The most common semantically-labeled dataset for Arabic is the ArabGlossBERT, a relatively small dataset that consists of 167K context-gloss pairs (about 60K positive and 107K negative pairs), collected from Arabic dictionaries.This paper presents an enrichment to the ArabGlossBERT dataset, by augmenting it using (Arabic-English-Arabic) machine back-translation.Augmentation increased the dataset size to 352K pairs (149K positive and 203K negative pairs).We measure the impact of augmentation using different data configurations to fine-tune BERT on target sense verification (TSV) task.Overall, the accuracy ranges between 78% to 84% for different data configurations.Although our approach performed at par with the baseline, we did observe some improvements for some POS tags in some experiments.Furthermore, our fine-tuned models are trained on a larger dataset covering larger vocabulary and contexts.We provide an in-depth analysis of the accuracy for each part-of-speech (POS). Sanad Malaysha, Mustafa Jarrar, Mohammad Khalilia |
GWC | 2 |
| 2022 | Curras + Baladi: Towards a Levantine CorpusabstractThis paper presents two-fold contributions: a full revision of the Palestinian morphologically annotated corpus (Curras), and a newly annotated Lebanese corpus (Baladi). Both corpora can be used as a more general Levantine corpus. Baladi consists of around 9.6K morphologically annotated tokens. Each token was manually annotated with several morphological features and using LDC’s SAMA lemmas and tags. The inter-annotator evaluation on most features illustrates 78.5% Kappa and 90.1% F1-Score. Curras was revised by refining all annotations for accuracy, normalization and unification of POS tags, and linking with SAMA lemmas. This revision was also important to ensure that both corpora are compatible and can help to bridge the nuanced linguistic gaps that exist between the two highly mutually intelligible dialects. Both corpora are publicly available through a web portal. Karim El Haff, Mustafa Jarrar, Tymaa Hammouda, Fadi A. Zaraket |
LREC | 2 |
| 2022 | Wojood: Nested Arabic Named Entity Corpus and Recognition using BERTabstractThis paper presents Wojood, a corpus for Arabic nested Named Entity Recognition (NER). Nested entities occur when one entity mention is embedded inside another entity mention. Wojood consists of about 550K Modern Standard Arabic (MSA) and dialect tokens that are manually annotated with 21 entity types including person, organization, location, event and date. More importantly, the corpus is annotated with nested entities instead of the more common flat annotations. The data contains about 75K entities and 22.5% of which are nested. The inter-annotator evaluation of the corpus demonstrated a strong agreement with Cohen’s Kappa of 0.979 and an F1-score of 0.976. To validate our data, we used the corpus to train a nested NER model based on multi-task learning using the pre-trained AraBERT (Arabic BERT). The model achieved an overall micro F1-score of 0.884. Our corpus, the annotation guidelines, the source code and the pre-trained model are publicly available. Mustafa Jarrar, Mohammad Khalilia, Sana Ghanem |
LREC | 1 |
| 2021 | Extracting Synonyms from Bilingual DictionariesabstractWe present our progress in developing a novel algorithm to extract synonyms from bilingual dictionaries.Identification and usage of synonyms play a significant role in improving the performance of information access applications.The idea is to construct a translation graph from translation pairs, then to extract and consolidate cyclic paths to form bilingual sets of synonyms.The initial evaluation of this algorithm illustrates promising results in extracting Arabic-English bilingual synonyms.In the evaluation, we first converted the synsets in the Arabic WordNet into translation pairs (i.e., losing word-sense memberships).Next, we applied our algorithm to rebuild these synsets.We compared the original and extracted synsets obtaining an F-Measure of 82.3% and 82.1% for Arabic and English synsets extraction, respectively. Mustafa Jarrar, Eman Karajah, Muhammad Khalifa, Khaled Shaalan |
GWC | 1 |
| 2019 | Usability Evaluation of Lexicographic e-ServicesabstractAlthough the field of usability evaluation is a well-established discipline, there are no studies on how the usability of lexicographic e-services can be evaluated. This includes, for examples efficiency, effectiveness and user satisfaction when looking up for synonyms, meanings, or translations using online lexicons. In this paper, we propose to combine two types of usability evaluations to assess the usability of such services: a subjective user-experience evaluation and a more objective controlled experiment - demonstrating how both methods complement each other. We applied our proposed approach to evaluate two important online lexicographic e-services: a lexicographic search engine developed at Birzeit University (https://ontology.birzeit.edu) as well as Google Translate. The user-experience evaluation was conducted through a survey that involved 622 users, and was designed to measure effectiveness, efficiency, satisfaction and learnability. The controlled experiment involved a set of defined tasks, which were carried out by four teams (12 people) in two laboratories, and their performance was monitored. The tasks were designed to measure effectiveness and efficiency. Diana Alhafi, Anton Deik, Elhadj Benkhelifa, Mustafa Jarrar |
AICCSA | 4 |
| 2019 | An Arabic-Multilingual Database with a Lexicographic Search Engine
Mustafa Jarrar, Hamzeh Amayreh |
NLDB | 1 |
| 2019 | Diacritic-Based Matching of Arabic WordsabstractWords in Arabic consist of letters and short vowel symbols called diacritics inscribed atop regular letters. Changing diacritics may change the syntax and semantics of a word; turning it into another. This results in difficulties when comparing words based solely on string matching. Typically, Arabic NLP applications resort to morphological analysis to battle ambiguity originating from this and other challenges. In this article, we introduce three alternative algorithms to compare two words with possibly different diacritics. We propose the Subsume knowledge-based algorithm, the Imply rule-based algorithm, and the Alike machine-learning-based algorithm. We evaluated the soundness, completeness, and accuracy of the algorithms against a large dataset of 86,886 word pairs. Our evaluation shows that the accuracy of Subsume (100%), Imply (99.32%), and Alike (99.53%). Although accurate, Subsume was able to judge only 75% of the data. Both Subsume and Imply are sound, while Alike is not. We demonstrate the utility of the algorithms using a real-life use case -- in lemma disambiguation and in linking hundreds of Arabic dictionaries. Mustafa Jarrar, Fadi A. Zaraket, Rami Asia, Hamzeh Amayreh |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2017 | Clustering Arabic Tweets for Sentiment AnalysisabstractThe focus of this study is to evaluate the impact of linguistic preprocessing and similarity functions for clustering Arabic Twitter tweets. The experiments apply an optimized version of the standard K-Means algorithm to assign tweets into positive and negative categories. The results show that root-based stemming has a significant advantage over light stemming in all settings. The Averaged Kullback-Leibler Divergence similarity function clearly outperforms the Cosine, Pearson Correlation, Jaccard Coefficient and Euclidean functions. The combination of the Averaged Kullback-Leibler Divergence and root-based stemming achieved the highest purity of 0.764 while the second-best purity was 0.719. These results are of importance as it is contrary to normal-sized documents where, in many information retrieval applications, light stemming performs better than root-based stemming and the Cosine function is commonly used. Diab Abuaiadah, Dileep Rajendran, Mustafa Jarrar |
AICCSA | 3 |
| 2016 | Effectiveness of Automatic Translations for Cross-Lingual Ontology MappingabstractAccessing or integrating data lexicalized in different languages is a challenge. Multilingual lexical resources play a fundamental role in reducing the language barriers to map concepts lexicalized in different languages. In this paper we present a large-scale study on the effectiveness of automatic translations to support two key cross-lingual ontology mapping tasks: the retrieval of candidate matches and the selection of the correct matches for inclusion in the final alignment. We conduct our experiments using four different large gold standards, each one consisting of a pair of mapped wordnets, to cover four different families of languages. We categorize concepts based on their lexicalization (type of words, synonym richness, position in a subconcept graph) and analyze their distributions in the gold standards. Leveraging this categorization, we measure several aspects of translation effectiveness, such as word-translation correctness, word sense coverage, synset and synonym coverage. Finally, we thoroughly discuss several findings of our study, which we believe are helpful for the design of more sophisticated cross-lingual mapping algorithms. Mamoun Abu Helou, Matteo Palmonari, Mustafa Jarrar |
J. Artif. Intell. Res. | 3 |
| 2015 | The Graph Signature: A Scalable Query Optimization Index for RDF Graph Databases Using Bisimulation and Trace Equivalence SummarizationabstractQuerying large data graphs has brought the attention of the research community. Many solutions were proposed, such as Oracle Semantic Technologies, Virtuoso, RDF3X, and C-Store, among others. Although such approaches have shown good performance in queries with medium complexity, they perform poorly when the complexity of the queries increases. In this paper, the authors propose the Graph Signature Index, a novel and scalable approach to index and query large data graphs. The idea is that they summarize a graph and instead of executing the query on the original graph, they execute it on the summaries. The authors' experiments with Yago (16M triples) have shown that e.g., a query with 4 levels costs 62 sec using Oracle but it only costs about 0.6 sec with their index. Their index can be implemented on top of any Graph database, but they chose to implement it as an extension to Oracle on top of the SEM_MATCH table function. The paper also introduces disk-based versions of the Trace Equivalence and Bisimilarity algorithms to summarize data graphs, and discusses their complexity and usability for RDF graphs. Mustafa Jarrar, Anton Deik |
Int. J. Semantic Web Inf. Syst. | 1 |
| 2014 | Towards Building Lexical Ontology via Cross-Language MatchingabstractIn this paper, we introduce a methodology for mapping linguistic ontologies lexicalized across different languages.We present a classification-based semantics for mappings of lexicalized concepts across different languages.We propose an experiment for validating the proposed cross-language mapping semantics, and discuss its role in creating a gold standard that can be used in assessing cross-language matching systems. Mamoun Abu Helou, Matteo Palmonari, Mustafa Jarrar, Christiane Fellbaum |
GWC | 3 |
| 2012 | A Query Formulation Language for the Data WebabstractWe present a query formulation language (called MashQL) in order to easily query and fuse structured data on the web. The main novelty of MashQL is that it allows people with limited IT skills to explore and query one (or multiple) data sources without prior knowledge about the schema, structure, vocabulary, or any technical details of these sources. More importantly, to be robust and cover most cases in practice, we do not assume that a data source should have - an offline or inline - schema. This poses several language-design and performance complexities that we fundamentally tackle. To illustrate the query formulation power of MashQL, and without loss of generality, we chose the Data web scenario. We also chose querying RDF, as it is the most primitive data model; hence, MashQL can be similarly used for querying relational databases and XML. We present two implementations of MashQL, an online mashup editor, and a Firefox add on. The former illustrates how MashQL can be used to query and mash up the Data web as simple as filtering and piping web feeds; and the Firefox add on illustrates using the browser as a web composer rather than only a navigator. To end, we evaluate MashQL on querying two data sets, DBLP and DBPedia, and show that our indexing techniques allow instant user interaction. Mustafa Jarrar, Marios D. Dikaiakos |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2011 | Special issue on querying the data web - Novel techniques for querying structured data on the web
Paolo Ceravolo, Chengfei Liu, Mustafa Jarrar, Kai-Uwe Sattler |
World Wide Web | 3 |
| 2007 | Towards Automated Reasoning on ORM Schemes
Mustafa Jarrar |
ER | 1 |
| 2006 | Position paper: towards the notion of gloss, and the adoption of linguistic resources in formal ontology engineeringabstractIn this paper, we first introduce the notion of gloss for ontology engineering purposes. We propose that each vocabulary in an ontology should have a gloss. A gloss basically is an informal description of the meaning of a vocabulary that is supposed to render factual and critical knowledge to understanding a concept, but that is unreasonable or very difficult to formalize and/or articulate formally. We present a set of guidelines on what should and should not be provided in a gloss. Second, we propose to incorporate linguistic resources in the ontology engineering process. We clarify the importance of using lexical resources as a "consensus reference" in ontology engineering, and so enabling the adoption of the glosses found in these resources. A linguistic resource (i.e. its list of terms and their definitions) shall be seen as a shared vocabulary space for ontologies. We present an ontology engineering software tool (called DogmaModeler), and illustrate its support of reusing of WordNet's terms and glosses in ontology modeling. Mustafa Jarrar |
WWW | 1 |
| 2004 | Using a Novel ORM-Based Ontology Modelling Method to Build an Experimental Innovation Router
Peter Spyns, Sven Van Acker, Marleen Wynants, Mustafa Jarrar, Andriy Lisovoy |
EKAW | 4 |