EDBT 2026 Demo / reviewers in the wild / expert
Nizar Habash
dblp:34/1998
· DBLP profile ↗
143ranked-venue papers
22as first author
37since 2021 · last 2026
0000-0002-1831-3457ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 141 · 22 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AIabstractYuxia Wang, Rui Xing, Jonibek Mansurov, Giovanni Puccetti, Zhuohan Xie, Minh Ngoc Ta, Jiahui Geng, Jinyan Su, Mervat Abassy, Saadeldine Eletter, Kareem Elozeiri, Nurkhan Laiyk, Maiya Goloburda, Tarek Mahmoud, Raj Vardhan Tomar, Alexander Aziz, Ryuto Koike, Masahiro Kaneko, Artem Shelmanov, Ekaterina Artemova, Vladislav Mikhailov, Akim Tsvigun, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, Preslav Nakov. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuxia Wang 0003, Rui Xing 0002, Jonibek Mansurov, Giovanni Puccetti 0002, Zhuohan Xie, Minh Ngoc Ta, Jiahui Geng, Jinyan Su, Mervat Abassy, Saadeldine Eletter, Kareem Ashraf Elozeiri, Nurkhan Laiyk, Maiya Goloburda, Tarek Mahmoud, Raj Vardhan Tomar, Alexander Aziz, Ryuto Koike, Masahiro Kaneko, Artem Shelmanov, Ekaterina Artemova, Vladislav Mikhailov, Akim Tsvigun, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, Preslav Nakov |
ACL (1) | 24 |
| 2026 | A Bilingual Bimodal Benchmark for Arabic-English NLP across Grammatical Correction, Essay Scoring, Morphological Tagging, and Speech Recognition
Bashar Alhafni, Injy Hamed, Fadhl Eryani, David Palfreyman, Nizar Habash |
LREC | 5 |
| 2026 | DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language ModelsabstractWe present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluation for Modern Standard Arabic (MSA), dialectal varieties remain underrepresented despite their prevalence in everyday communication. DialectalArabicMMLU extends the MMLU-Redux framework through manual translation and adaptation of 3K multiple-choice question-answer pairs into five major dialects (Syrian, Egyptian, Emirati, Saudi, and Moroccan), yielding a total of 15K QA pairs across 32 academic and professional domains (22K QA pairs when also including English and MSA). The benchmark enables systematic assessment of LLM reasoning and comprehension beyond MSA, supporting both task-based and linguistic analysis. We evaluate 19 open-weight Arabic and multilingual LLMs (1B-13B parameters) and report substantial performance variation across dialects, revealing persistent gaps in dialectal generalization. DialectalArabicMMLU provides the first unified, human-curated resource for measuring dialectal understanding in Arabic, thus promoting more inclusive evaluation and future model development. Malik H. Altakrori, Nizar Habash, Teresa Lynn, Younes Samih, Abed Alhakim Freihat, Kirill Chirkunov, Muhammed AbuOdeh, Radu Florian, Preslav Nakov, Alham Fikri Aji |
LREC | 2 |
| 2026 | A Large and Balanced Multi-Domain Arabic Corpus Annotated for Morphology, Syntax, and Readability
Khalid N. Elmadani, Adel Mahmoud Wizani, Hanada Taha-Thomure, Nizar Habash |
LREC | 4 |
| 2026 | Benchmarking Arabic Authorship Attribution and Style Transfer with Large Language Models
Injy Hamed, Bashar Alhafni, Nizar Habash, Thamar Solorio |
LREC | 3 |
| 2026 | Hazawi+: A Structured Corpus in Kuwaiti Arabic StoriesabstractThe availability of high-quality corpora is foundational for advancements in Natural Language Processing (NLP), enabling the training and rigorous evaluation of computational models. While rich textual resources exist for high-resource languages, a significant scarcity persists for many natural languages, particularly understudied Arabic dialects such as Kuwaiti Arabic (KA). This article introduces Hazawi+ , a multi-domain textual corpus comprising over 7 million tokens of KA dialectal stories and novels. Unlike social network texts, Hazawi+ is specifically designed to capture the rich linguistic features inherent in narrative texts, including morphological complexity, informal syntax, and pragmatic nuances, making it an invaluable resource for developing NLP models in low-resource settings. The entire corpus underwent automatic morphological annotation using CAMeL tools specialized for Gulf Arabic, with annotation quality subsequently validated through a rigorous manual review of a 105,770-token sample by two language experts. To demonstrate Hazawi+ ’s immediate usability, we present a complementary empirical study involving the programmatic generation of a synthetic dataset of KA stories, which is then utilized in a downstream task to train a classifier capable of distinguishing between human-written and bot-generated narratives. This experiment serves as a crucial proof-of-concept, underscoring Hazawi+ ’s potential to provide researchers with deep insights into dialectal linguistic patterns and to significantly enhance the precision of various language processing tasks for the Kuwaiti Arabic dialect. Fatemah Husain, Mohammad Alenezi, Nizar Habash |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2025 | Enhancing Text Editing for Grammatical Error Correction: Arabic as a Case StudyabstractText editing frames grammatical error correction (GEC) as a sequence tagging problem, where edit tags are assigned to input tokens, and applying these edits results in the corrected text.This approach has gained attention for its efficiency and interpretability.However, while extensively explored for English, text editing remains largely underexplored for morphologically rich languages like Arabic.In this paper, we introduce a text editing approach that derives edit tags directly from data, eliminating the need for language-specific edits.We demonstrate its effectiveness on Arabic, a diglossic and morphologically rich language, and investigate the impact of different edit representations on model performance.Our approach achieves SOTA results on two Arabic GEC benchmarks and performs on par with SOTA on two others.Additionally, our models are over six times faster than existing Arabic GEC systems, making our approach more practical for real-world applications.Finally, we explore ensemble models, demonstrating how combining different models leads to further performance improvements.We make our code, data, and pretrained models publicly available.1 < Bashar Alhafni, Nizar Habash |
ACL (1) | 2 |
| 2025 | A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future DirectionsabstractLanguage in the Arab world presents a complex diglossic and multilingual setting, involving the use of Modern Standard Arabic, various dialects and sub-dialects, as well as multiple European languages. This diverse linguistic landscape has given rise to code-switching, both within Arabic varieties and between Arabic and foreign languages. The widespread occurrence of code-switching across the region makes it vital to address these linguistic needs when developing language technologies. In this paper, we provide a review of the current literature in the field of code-switched Arabic NLP, offering a broad perspective on ongoing efforts, challenges, research gaps, and recommendations for future research directions. Injy Hamed, Caroline Sabty, Slim Abdennadher, Ngoc Thang Vu, Thamar Solorio, Nizar Habash |
COLING | 6 |
| 2025 | From Multiple-Choice to Extractive QA: A Case Study for English and ArabicabstractThe rapid evolution of Natural Language Processing (NLP) has favoured major languages such as English, leaving a significant gap for many others due to limited resources. This is especially evident in the context of data annotation, a task whose importance cannot be underestimated, but which is time-consuming and costly. Thus, any dataset for resource-poor languages is precious, in particular when it is task-specific. Here, we explore the feasibility of repurposing an existing multilingual dataset for a new NLP task: we repurpose a subset of the BELEBELE dataset (Bandarkar et al., 2023), which was designed for multiple-choice question answering (MCQA), to enable the more practical task of extractive QA (EQA) in the style of machine reading comprehension. We present annotation guidelines and a parallel EQA dataset for English and Modern Standard Arabic (MSA). We also present QA evaluation results for several monolingual and cross-lingual QA pairs including English, MSA, and five Arabic dialects. We aim to help others adapt our approach for the remaining 120 BELEBELE language variants, many of which are deemed under-resourced. We also provide a thorough analysis and share insights to deepen understanding of the challenges and opportunities in NLP task reformulation. Teresa Lynn, Malik H. Altakrori, Samar Mohamed Magdy, Rocktim Jyoti Das, Chenyang Lyu, Mohamed Nasr, Younes Samih, Kirill Chirkunov, Alham Fikri Aji, Preslav Nakov, Shantanu Godbole, Salim Roukos, Radu Florian, Nizar Habash |
COLING | 14 |
| 2025 | Lemmatization as a Classification Task: Results from Arabic across Multiple GenresabstractLemmatization is crucial for NLP tasks in morphologically rich languages with ambiguous orthography like Arabic, but existing tools face challenges due to inconsistent standards and limited genre coverage.This paper introduces two novel approaches that frame lemmatization as classification into a Lemma-POS-Gloss (LPG) tagset, leveraging machine translation and semantic clustering.We also present a new Arabic lemmatization test set covering diverse genres, standardized alongside existing datasets.We evaluate character-level sequenceto-sequence models, which perform competitively and offer complementary value, but are limited to lemma prediction (not LPG) and prone to hallucinating implausible forms.Our results show that classification and clustering yield more robust, interpretable outputs, setting new benchmarks for Arabic lemmatization.Data/tool Mostafa Saeed, Nizar Habash |
EMNLP | 2 |
| 2025 | The Arabic Generality Score: Another Dimension of Modeling Arabic DialectnessabstractArabic dialects form a diverse continuum, yet NLP models often treat them as discrete categories.Recent work addresses this issue by modeling dialectness as a continuous variable, notably through the Arabic Level of Dialectness (ALDi).However, ALDi reduces complex variation to a single dimension.We propose a complementary measure: the Arabic Generality Score (AGS), which quantifies how widely a word is used across dialects.We introduce a pipeline that combines word alignment, etymology-aware edit distance, and smoothing to annotate a parallel corpus with word-level AGS.A regression model is then trained to predict AGS in context.Our approach outperforms strong baselines, including state-of-the-art dialect ID systems, on a multi-dialect benchmark.AGS offers a scalable, linguistically grounded way to model lexical generality, enriching representations of Arabic dialectness. 1 Sanad Shaban, Nizar Habash |
EMNLP | 2 |
| 2024 | Arabic Diacritics in the Wild: Exploiting Opportunities for Improved DiacritizationabstractThe widespread absence of diacritical marks in Arabic text poses a significant challenge for Arabic natural language processing (NLP).This paper explores instances of naturally occurring diacritics, referred to as "diacritics in the wild," to unveil patterns and latent information across six diverse genres: news articles, novels, children's books, poetry, political documents, and ChatGPT outputs.We present a new annotated dataset that maps realworld partially diacritized words to their maximal full diacritization in context.Additionally, we propose extensions to the analyze-anddisambiguate approach in Arabic NLP to leverage these diacritics, resulting in notable improvements.Our contributions encompass a thorough analysis, valuable datasets, and an extended diacritization algorithm.We release our code and datasets as open source. Salman Elgamal, Ossama Obeid, Mhd Tameem Kabbani, Go Inoue, Nizar Habash |
ACL (1) | 5 |
| 2024 | M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text DetectionabstractYuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Osama Mohammed Afzal, Tarek Mahmoud, Giovanni Puccetti, Thomas Arnold, Alham Aji, Nizar Habash, Iryna Gurevych, Preslav Nakov. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yuxia Wang 0003, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Osama Mohammed Afzal, Tarek Mahmoud, Giovanni Puccetti 0002, Thomas Arnold 0002, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, Preslav Nakov |
ACL (1) | 12 |
| 2024 | Palmyra 3.0: A User-Friendly Cloud-Based Platform for Morphology and Dependency Syntax AnnotationabstractWe present Palmyra 3.0, a cloud-based, configurable, and user-friendly platform for morphology and syntax annotation through dependency-tree visualization. Palmyra 3.0 implements a robust system that stores data on the cloud. By default, Palmyra 3.0 comes with an Arabic dependency parser that generates highly accurate trees, but it is easily configurable to support dependency parsers in other languages. Palmyra 3.0 provides default configuration files for a number of predefined formalisms, such as UD and CATiB, and a number of user-friendly features to support annotators. Muhammed AbuOdeh, Long Phan, Ahmed Elshabrawy, Nizar Habash |
LREC/COLING | 4 |
| 2024 | The SAMER Arabic Text Simplification CorpusabstractWe present the SAMER Corpus, the first manually annotated Arabic parallel corpus for text simplification targeting school-aged learners. Our corpus comprises texts of 159K words selected from 15 publicly available Arabic fiction novels most of which were published between 1865 and 1955. Our corpus includes readability level annotations at both the document and word levels, as well as two simplified parallel versions for each text targeting learners at two different readability levels. We describe the corpus selection process, and outline the guidelines we followed to create the annotations and ensure their quality. Our corpus is publicly available to support and encourage research on Arabic text simplification, Arabic automatic readability assessment, and the development of Arabic pedagogical language technologies. Bashar Alhafni, Reem Hazim, Juan David Pineros Liberato, Muhamed Al-Khalil, Nizar Habash |
LREC/COLING | 5 |
| 2024 | ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech CorpusabstractWe present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm ideas for a certain topic and then discuss it with an Interlocutor. The meetings cover different topics and are divided into phases with different language setups. The corpus presents a challenging set for automatic speech recognition (ASR), including two languages (Arabic and English) with Arabic spoken in multiple variants (Modern Standard Arabic, Gulf Arabic, and Egyptian Arabic) and English used with various accents. Adding to the complexity of the corpus, there is also code-switching between these languages and dialects. As part of our work, we take inspiration from established sets of transcription guidelines to present a set of guidelines handling issues of conversational speech, code-switching and orthography of both languages. We further enrich the corpus with two layers of annotations; (1) dialectness level annotation for the portion of the corpus where mixing occurs between different variants of Arabic, and (2) automatic morphological annotations, including tokenization, lemmatization, and part-of-speech tagging. Injy Hamed, Fadhl Eryani, David Palfreyman, Nizar Habash |
LREC/COLING | 4 |
| 2024 | EMAD: A Bridge Tagset for Unifying Arabic POS AnnotationsabstractThere have been many attempts to model the morphological richness and complexity of Arabic, leading to numerous Part-of-Speech (POS) tagsets that differ in terms of (a) which morphological features they represent, (b) how they represent them, and (c) the degree of specification of said features. Tagset granularity plays an important role in determining how annotated data can be used and for what applications. Due to the diversity among existing tagsets, many annotated corpora for Arabic cannot be easily combined, which exacerbates the Arabic resource poverty situation. In this work, we propose an intermediate tagset designed to facilitate the conversion and unification of different tagsets used to annotate Arabic corpora. This new tagset acts as a bridge between different annotation schemes, simplifying the integration of annotated corpora and promoting collaboration across the projects using them. Omar Kallas, Go Inoue, Nizar Habash |
LREC/COLING | 3 |
| 2024 | Camel Morph MSA: A Large-Scale Open-Source Morphological Analyzer for Modern Standard ArabicabstractWe present Camel Morph MSA, the largest open-source Modern Standard Arabic morphological analyzer and generator. Camel Morph MSA has over 100K lemmas, and includes rarely modeled morphological features of Modern Standard Arabic with Classical Arabic origins. Camel Morph MSA can produce ∼1.45B analyses and ∼535M unique diacritizations, almost an order of magnitude larger than SAMA (Maamouri et al., 2010c), in addition to having ∼36% less OOV rate than SAMA on a 10B word corpus. Furthermore, Camel Morph MSA fills the gaps of many lemma paradigms by modeling linguistic phenomena consistently. Camel Morph MSA seamlessly integrates with the Camel Tools Python toolkit (Obeid et al., 2020), ensuring ease of use and accessibility. Christian Khairallah, Salam Khalifa, Reham Marzouk, Mayar Nassar, Nizar Habash |
LREC/COLING | 5 |
| 2024 | Cross-Lingual Transfer from Related Languages: Treating Low-Resource Maltese as Multilingual Code-SwitchingabstractKurt Micallef, Nizar Habash, Claudia Borg, Fadhl Eryani, Houda Bouamor. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Kurt Micallef, Nizar Habash, Claudia Borg, Fadhl Eryani, Houda Bouamor |
EACL (1) | 2 |
| 2024 | M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text DetectionabstractYuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Chenxi Whitehouse, Osama Mohammed Afzal, Tarek Mahmoud, Toru Sasaki, Thomas Arnold, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, Preslav Nakov. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yuxia Wang 0003, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Chenxi Whitehouse, Osama Mohammed Afzal, Tarek Mahmoud, Toru Sasaki, Thomas Arnold 0002, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, Preslav Nakov |
EACL (1) | 13 |
| 2024 | HelloThere: A Corpus of Annotated Dialogues and Knowledge Bases of Time-Offset AvatarsabstractA Time-Offset Interaction Application (TOIA) is a software system that allows people to engage in face-to-face dialogue with previously recorded videos of other people.There are two TOIA usage modes: (a) creation mode, where users pre-record video snippets of themselves representing their answers to possible questions someone may ask them, and (b) interaction mode, where other users of the system can choose to interact with created avatars.This paper presents the HelloThere corpus that has been collected from two user studies involving several people who recorded avatars and many more who engaged in dialogues with them.The interactions with avatars are annotated by people asking them questions through three modes (card selection, text search, and voice input) and rating the appropriateness of their answers on a 1 to 5 scale.The corpus, made available to the research community, comprises 26 avatars' knowledge bases and 317 dialogues between 64 interrogators and the avatars in text format. Alberto Chierici, Nizar Habash |
SIGDIAL | 2 |
| 2023 | Exploring Segmentation Approaches for Neural Machine Translation of Code-Switched Egyptian Arabic-English TextabstractData sparsity is one of the main challenges posed by code-switching (CS), which is further exacerbated in the case of morphologically rich languages.For the task of machine translation (MT), morphological segmentation has proven successful in alleviating data sparsity in monolingual contexts; however, it has not been investigated for CS settings.In this paper, we study the effectiveness of different segmentation approaches on MT performance, covering morphology-based and frequency-based segmentation techniques.We experiment on MT from code-switched Arabic-English to English.We provide detailed analysis, examining a variety of conditions, such as data size and sentences with different degrees of CS.Empirical results show that morphology-aware segmenters perform the best in segmentation tasks but under-perform in MT.Nevertheless, we find that the choice of the segmentation setup to use for MT is highly dependent on the data size.For extreme low-resource scenarios, a combination of frequency and morphology-based segmentations is shown to perform the best.For more resourced settings, such a combination does not bring significant improvements over the use of frequency-based segmentation. Marwa Gaser, Manuel Mager, Injy Hamed, Nizar Habash, Slim Abdennadher, Ngoc Thang Vu |
EACL | 4 |
| 2023 | Advancements in Arabic Grammatical Error Detection and Correction: An Empirical InvestigationabstractGrammatical error correction (GEC) is a wellexplored problem in English with many existing models and datasets.However, research on GEC in morphologically rich languages has been limited due to challenges such as data scarcity and language complexity.In this paper, we present the first results on Arabic GEC using two newly developed Transformer-based pretrained sequence-to-sequence models.We also define the task of multi-class Arabic grammatical error detection (GED) and present the first results on multi-class Arabic GED.We show that using GED information as an auxiliary input in GEC models improves GEC performance across three datasets spanning different genres.Moreover, we also investigate the use of contextual morphological preprocessing in aiding GEC systems.Our models achieve SOTA results on two Arabic GEC shared task datasets and establish a strong benchmark on a recently created dataset.We make our code, data, and pretrained models publicly available.1 Bashar Alhafni, Go Inoue, Christian Khairallah, Nizar Habash |
EMNLP | 4 |
| 2023 | Benchmarking Dialectal Arabic-Turkish Machine TranslationabstractDue to the significant influx of Syrian refugees in Turkey in recent years, the Syrian Arabic dialect has become increasingly prevalent in certain regions of Turkey. Developing a machine translation system between Turkish and Syrian Arabic would be crucial in facilitating communication between the Turkish and Syrian communities in these regions, which can have a positive impact on various domains such as politics, trade, and humanitarian aid. Such a system would also contribute positively to the growing Arab-focused tourism industry in Turkey. In this paper, we present the first research effort exploring translation between Syrian Arabic and Turkish. We use a set of 2,000 parallel sentences from the MADAR corpus containing 25 different city dialects from different cities across the Arab world, in addition to Modern Standard Arabic (MSA), English, and French. Additionally, we explore the translation performance into Turkish from other Arabic dialects and compare the results to the performance achieved when translating from Syrian Arabic. We build our MADAR-Turk data set by manually translating the set of 2,000 sentences from the Damascus dialect of Syria to Turkish with the help of two native Arabic speakers from Syria who are also highly fluent in Turkish. We evaluate the quality of the translations and report the results achieved. We make this first-of-a-kind data set publicly available to support research in machine translation between these important but less studied language pairs. Hasan Alkheder, Houda Bouamor, Nizar Habash, Ahmet Zengin |
MTSummit (1) | 3 |
| 2023 | Tell Me More, Tell Me More: AI-Generated Question Suggestions for the Creation of Interactive Video RecordingsabstractTime-Offset Interaction Applications (TOIAs) are narrative-sharing systems that use databases of previously recorded videos of real people to mimic conversations with them. These video databases comprise large (the larger, the better) collections of videos of answers paired with specific questions. This paper focuses on a solution to the challenge of creating such databases without exhausting their creators’ creativity, energy, and interest. We describe the design and development process of Question Suggester (QS)-an intelligent GPT-3-based service that generates suggested questions following up a conversation based on the history of recorded questions and answers. We conduct a user study to empirically evaluate the value of QS for reducing the effort to create a video database while creating an interaction that is enjoyable. The users’ average experience rating for QS is 4.6 compared to 4.0 when QS is not used (on a 1-5 scale, p-value<0.05). The experience with interactions so created is more enjoyable, too (3.7vs. 3.3, p-value<0.05). The usage metrics and qualitative feedback confirm that QS is essential for interactive video-recording systems and for increasing their adoption. Alberto Chierici, Nizar Habash |
RO-MAN | 2 |
| 2022 | The Bahrain Corpus: A Multi-genre Corpus of Bahraini ArabicabstractIn recent years, the focus on developing natural language processing (NLP) tools for Arabic has shifted from Modern Standard Arabic to various Arabic dialects. Various corpora of various sizes and representing different genres, have been created for a number of Arabic dialects. As far as Gulf Arabic is concerned, Gumar Corpus (Khalifa et al., 2016) is the largest corpus, to date, that includes data representing the dialectal Arabic of the six Gulf Cooperation Council countries (Bahrain, Kuwait, Saudi Arabia, Qatar, United Arab Emirates, and Oman), particularly in the genre of “online forum novels”. In this paper, we present the Bahrain Corpus. Our objective is to create a specialized corpus of the Bahraini Arabic dialect, which includes written texts as well as transcripts of audio files, belonging to a different genre (folktales, comedy shows, plays, cooking shows, etc.). The corpus comprises 620K words, carefully curated. We provide automatic morphological annotations of the full corpus using state-of-the-art morphosyntactic disambiguation for Gulf Arabic. We validate the quality of the annotations on a 7.6K word sample. We plan to make the annotated sample as well as the full corpus publicly available to support researchers interested in Arabic NLP. Dana Abdulrahim, Go Inoue, Latifa Shamsan, Salam Khalifa, Nizar Habash |
LREC | 5 |
| 2022 | The Arabic Parallel Gender Corpus 2.0: Extensions and AnalysesabstractGender bias in natural language processing (NLP) applications, particularly machine translation, has been receiving increasing attention. Much of the research on this issue has focused on mitigating gender bias in English NLP models and systems. Addressing the problem in poorly resourced, and/or morphologically rich languages has lagged behind, largely due to the lack of datasets and resources. In this paper, we introduce a new corpus for gender identification and rewriting in contexts involving one or two target users (I and/or You) – first and second grammatical persons with independent grammatical gender preferences. We focus on Arabic, a gender-marking morphologically rich language. The corpus has multiple parallel components: four combinations of 1st and 2nd person in feminine and masculine grammatical genders, as well as English, and English to Arabic machine translation output. This corpus expands on Habash et al. (2019)’s Arabic Parallel Gender Corpus (APGC v1.0) by adding second person targets as well as increasing the total number of sentences over 6.5 times, reaching over 590K words. Our new dataset will aid the research and development of gender identification, controlled text generation, and post-editing rewrite systems that could be used to personalize NLP applications and provide users with the correct outputs based on their grammatical gender preferences. We make the Arabic Parallel Gender Corpus (APGC v2.0) publicly available Bashar Alhafni, Nizar Habash, Houda Bouamor |
LREC | 2 |
| 2022 | Hierarchical Aggregation of Dialectal Data for Arabic Dialect IdentificationabstractArabic is a collection of dialectal variants that are historically related but significantly different. These differences can be seen across regions, countries, and even cities in the same countries. Previous work on Arabic Dialect identification has focused mainly on specific dialect levels (region, country, province, or city) using level-specific resources; and different efforts used different schemas and labels. In this paper, we present the first effort aiming at defining a standard unified three-level hierarchical schema (region-country-city) for dialectal Arabic classification. We map 29 different data sets to this unified schema, and use the common mapping to facilitate aggregating these data sets. We test the value of such aggregation by building language models and using them in dialect identification. We make our label mapping code and aggregated language models publicly available. Nurpeiis Baimukan, Houda Bouamor, Nizar Habash |
LREC | 3 |
| 2022 | UniMorph 4.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet. Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova |
LREC | 4 |
| 2022 | Camel Treebank: An Open Multi-genre Arabic Dependency TreebankabstractWe present the Camel Treebank (CAMELTB), a 188K word open-source dependency treebank of Modern Standard and Classical Arabic. CAMELTB 1.0 includes 13 sub-corpora comprising selections of texts from pre-Islamic poetry to social media online commentaries, and covering a range of genres from religious and philosophical texts to news, novels, and student essays. The texts are all publicly available (out of copyright, creative commons, or under open licenses). The texts were morphologically tokenized and syntactically parsed automatically, and then manually corrected by a team of trained annotators. The annotations follow the guidelines of the Columbia Arabic Treebank (CATiB) dependency representation. We discuss our annotation process and guideline extensions, and we present some initial observations on lexical and syntactic differences among the annotated sub-corpora. This corpus will be publicly available to support and encourage research on Arabic NLP in general and on new, previously unexplored genres that are of interest to a wider spectrum of researchers, from historical linguistics and digital humanities to computer-assisted language pedagogy. Nizar Habash, Muhammed AbuOdeh, Dima Taji, Reem Faraj, Jamila El Gizuli, Omar Kallas |
LREC | 1 |
| 2022 | ZAEBUC: An Annotated Arabic-English Bilingual Writer CorpusabstractWe present ZAEBUC, an annotated Arabic-English bilingual writer corpus comprising short essays by first-year university students at Zayed University in the United Arab Emirates. We describe and discuss the various guidelines and pipeline processes we followed to create the annotations and quality check them. The annotations include spelling and grammar correction, morphological tokenization, Part-of-Speech tagging, lemmatization, and Common European Framework of Reference (CEFR) ratings. All of the annotations are done on Arabic and English texts using consistent guidelines as much as possible, with tracked alignments among the different annotations, and to the original raw texts. For morphological tokenization, POS tagging, and lemmatization, we use existing automatic annotation tools followed by manual correction. We also present various measurements and correlations with preliminary insights drawn from the data and annotations. The publicly available ZAEBUC corpus and its annotations are intended to be the stepping stones for additional annotations. Nizar Habash, David Palfreyman |
LREC | 1 |
| 2022 | User-Centric Gender RewritingabstractIn this paper, we define the task of gender rewriting in contexts involving two users (I and/or You) -first and second grammatical persons with independent grammatical gender preferences.We focus on Arabic, a gendermarking morphologically rich language.We develop a multi-step system that combines the positive aspects of both rule-based and neural rewriting models.Our results successfully demonstrate the viability of this approach on a recently created corpus for Arabic gender rewriting, achieving 88.42 M 2 F 0.5 on a blind test set.Our proposed system improves over previous work on the first-person-only version of this task, by 3.05 absolute increase in M 2 F 0.5 .We demonstrate a use case of our gender rewriting system by using it to post-edit the output of a commercial MT system to provide personalized outputs based on the users' grammatical gender preferences.We make our code, data, and pretrained models publicly available.1 Bashar Alhafni, Nizar Habash, Houda Bouamor |
NAACL-HLT | 2 |
| 2022 | Benchmarking Evaluation Metrics for Code-Switching Automatic Speech RecognitionabstractCode-switching poses a number of challenges and opportunities for multilingual automatic speech recognition. In this paper, we focus on the question of robust and fair evaluation metrics. To that end, we develop a reference benchmark data set of code-switching speech recognition hypotheses with human judgments. We define clear guidelines for minimal editing of automatic hypotheses. We validate the guidelines using 4-way inter-annotator agreement. We evaluate a large number of metrics in terms of correlation with human judgments. The metrics we consider vary in terms of representation (orthographic, phonological, semantic), directness (intrinsic vs extrinsic), granularity (e.g. word, character), and similarity computation method. The highest correlation to human judgment is achieved using transliteration followed by text normalization. We release the first corpus for human acceptance of code-switching speech recognition results in dialectal Arabic/English conversation speech. Injy Hamed, Amir Hussein, Oumnia Chellah, Shammur Absar Chowdhury, Hamdy Mubarak, Sunayana Sitaram, Nizar Habash, Ahmed Ali 0002 |
SLT | 7 |
| 2022 | Unsupervised Arabic dialect segmentation for machine translationabstractAbstract Resource-limited and morphologically rich languages pose many challenges to natural language processing tasks. Their highly inflected surface forms inflate the vocabulary size and increase sparsity in an already scarce data situation. In this article, we present an unsupervised learning approach to vocabulary reduction through morphological segmentation. We demonstrate its value in the context of machine translation for dialectal Arabic (DA), the primarily spoken, orthographically unstandardized, morphologically rich and yet resource poor variants of Standard Arabic. Our approach exploits the existence of monolingual and parallel data. We show comparable performance to state-of-the-art supervised methods for DA segmentation. Wael Salloum, Nizar Habash |
Nat. Lang. Eng. | 2 |
| 2021 | Automatic Error Type Annotation for ArabicabstractWe present ARETA, an automatic error type annotation system for Modern Standard Arabic.We design ARETA to address Arabic's morphological richness and orthographic ambiguity.We base our error taxonomy on the Arabic Learner Corpus (ALC) Error Tagset with some modifications.ARETA achieves a performance of 85.8% (micro average F1 score) on a manually annotated blind test portion of ALC.We also demonstrate ARETA's usability by applying it to a number of submissions from the QALB 2014 shared task for Arabic grammatical error correction.The resulting analyses give helpful insights on the strengths and weaknesses of different submissions, which is more useful than the opaque M 2 scoring metrics used in the shared task.ARETA employs a large Arabic morphological analyzer, but is completely unsupervised otherwise.We make ARETA publicly available. Riadh Belkebir, Nizar Habash |
CoNLL | 2 |
| 2021 | Towards Automatic Narrative Coherence PredictionabstractResearch in Psychology has shown that stories people tell about themselves, and how they recall their experiences, reveal a lot about their individual characteristics and mental well-being. The Narrative Coherence Coding Scheme (NaCCS) is a set of guidelines established in psychology research for annotating the “coherence” of a narrative along three dimensions: context, chronology and theme. A significant correlation was found between a narrative’s coherence score and independently collected mental health markers of the narrator. Currently, all coherence annotations are done manually; a time consuming task which drains vital resources. In this paper, we propose an Artificial Intelligence based approach involving Natural Language Processing (NLP) to predict a narrative’s coherence score (4-class classification problem). We explore a number of techniques, ranging from traditional machine learning models such as Support Vector Machines (SVM) to pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers). BERT produced the best results for all dimensions in terms of accuracy: 53.7% (context), 71.8% (chronology), and 69.6% (theme). The location of information in the narratives (beginning, end, throughout) was helpful in improving predictions. Filip Bendevski, Jumana Ibrahim, Tina Krulec, Theodore Waters, Nizar Habash, Hanan Salam, Himadri Mukherjee, Christin Camia |
ICMI | 5 |
| 2021 | A Cloud-based User-Centered Time-Offset Interaction ApplicationabstractAlberto Chierici, Tyeece Kiana Fredorcia Hensley, Wahib Kamran, Kertu Koss, Armaan Agrawal, Erin Meekhof, Goffredo Puccetti, Nizar Habash. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021. Alberto Chierici, Tyeece Kiana Fredorcia Hensley, Wahib Kamran, Kertu Koss, Armaan Agrawal, Erin Meekhof, Goffredo Puccetti, Nizar Habash |
SIGDIAL | 8 |
| 2020 | The Paradigm Discovery ProblemabstractThis work treats the paradigm discovery problem (PDP)-the task of learning an inflectional morphological system from unannotated sentences.We formalize the PDP and develop evaluation metrics for judging systems.Using currently available resources, we construct datasets for the task.We also devise a heuristic benchmark for the PDP and report empirical results on five diverse languages.Our benchmark system first makes use of word embeddings and string similarity to cluster forms by cell and by paradigm.Then, we bootstrap a neural transducer on top of the clustered data to predict words to realize the empty paradigm slots.An error analysis of our system suggests clustering by cell across different inflection classes is the most pressing challenge for future work.Our code and data are available at https://github.com/ alexerdmann/ParadigmDiscovery. Alexander Erdmann, Micha Elsner, Ryan Cotterell, Nizar Habash |
ACL | 5 |
| 2020 | Joint Diacritization, Lemmatization, Normalization, and Fine-Grained Morphological TaggingabstractThe written forms of Semitic languages are both highly ambiguous and morphologically rich: a word can have multiple interpretations and is one of many inflected forms of the same concept or lemma.This is further exacerbated for dialectal content, which is more prone to noise and lacks a standard orthography.The morphological features can be lexicalized, like lemmas and diacritized forms, or non-lexicalized, like gender, number, and partof-speech tags, among others.Joint modeling of the lexicalized and non-lexicalized features can identify more intricate morphological patterns, which provide better context modeling, and further disambiguate ambiguous lexical choices.However, the different modeling granularity can make joint modeling more difficult.Our approach models the different features jointly, whether lexicalized (on the characterlevel), or non-lexicalized (on the word-level).We use Arabic as a test case, and achieve stateof-the-art results for Modern Standard Arabic with 20% relative error reduction, and Egyptian Arabic with 11% relative error reduction. Nasser Zalmout, Nizar Habash |
ACL | 2 |
| 2020 | Multitask Easy-First Dependency Parsing: Exploiting Complementarities of Different Dependency RepresentationsabstractWe present a parsing model for projective dependency trees which takes advantage of the existence of complementary dependency annotations for a language.This is the case for Arabic with the availability of CATiB and UD treebanks.Our system performs syntactic parsing according to both annotation types jointly as a sequence of arc-creating operations following the Easy-First approach, and partially created trees for one annotation type are also available to the other as features for the score function.This method gives error reduction of 9.9% on CATiB and 6.1% on UD compared to a single-task baseline, and ablation tests show that the main contribution of this reduction is given by sharing tree representation between tasks, and not simply sharing BiLSTM layers as is usually performed in NLP multitask systems. Yash Kankanampati, Joseph Le Roux, Nadi Tomeh, Dima Taji, Nizar Habash |
COLING | 5 |
| 2020 | Utilizing Subword Entities in Character-Level Sequence-to-Sequence Lemmatization ModelsabstractIn this paper we present a character-level sequence-to-sequence lemmatization model, utilizing several subword features in multiple configurations.In addition to generic n-gram embeddings (using FastText), we experiment with concatenative (stems) and templatic (roots and patterns) morphological subwords.We present several architectures that embed these features directly at the encoder side, or learn them jointly at the decoder side with a multitask learning architecture.The results indicate that using the generic n-gram embeddings (through FastText) outperform the other linguistically-driven subwords.We use Modern Standard Arabic and Egyptian Arabic as test cases, with up to 22% and 13% relative error reduction, respectively, from a strong baseline.An error analysis shows that our best system is even able to handle word/lemma pairs that are both unseen in the training data. Nasser Zalmout, Nizar Habash |
COLING | 2 |
| 2020 | A Large-Scale Leveled Readability Lexicon for Standard ArabicabstractWe present a large-scale 26,000-lemma leveled readability lexicon for Modern Standard Arabic. The lexicon was manually annotated in triplicate by language professionals from three regions in the Arab world. The annotations show a high degree of agreement; and major differences were limited to regional variations. Comparing lemma readability levels with their frequencies provided good insights in the benefits and pitfalls of frequency-based readability approaches. The lexicon will be publicly available. Muhamed Al-Khalil, Nizar Habash, Zhengyang Jiang |
LREC | 2 |
| 2020 | The Margarita Dialogue Corpus: A Data Set for Time-Offset Interactions and Unstructured Dialogue SystemsabstractTime-Offset Interaction Applications (TOIAs) are systems that simulate face-to-face conversations between humans and digital human avatars recorded in the past. Developing a well-functioning TOIA involves several research areas: artificial intelligence, human-computer interaction, natural language processing, question answering, and dialogue systems. The first challenges are to define a sensible methodology for data collection and to create useful data sets for training the system to retrieve the best answer to a user’s question. In this paper, we present three main contributions: a methodology for creating the knowledge base for a TOIA, a dialogue corpus, and baselines for single-turn answer retrieval. We develop the methodology using a two-step strategy. First, we let the avatar maker list pairs by intuition, guessing what possible questions a user may ask to the avatar. Second, we record actual dialogues between random individuals and the avatar-maker. We make the Margarita Dialogue Corpus available to the research community. This corpus comprises the knowledge base in text format, the video clips for each answer, and the annotated dialogues. Alberto Chierici, Nizar Habash, Margarita Bicec |
LREC | 2 |
| 2020 | A Spelling Correction Corpus for Multiple Arabic DialectsabstractArabic dialects are the non-standard varieties of Arabic commonly spoken – and increasingly written on social media – across the Arab world. Arabic dialects do not have standard orthographies, a challenge for natural language processing applications. In this paper, we present the MADAR CODA Corpus, a collection of 10,000 sentences from five Arabic city dialects (Beirut, Cairo, Doha, Rabat, and Tunis) represented in the Conventional Orthography for Dialectal Arabic (CODA) in parallel with their raw original form. The sentences come from the Multi-Arabic Dialect Applications and Resources (MADAR) Project and are in parallel across the cities (2,000 sentences from each city). This publicly available resource is intended to support research on spelling correction and text normalization for Arabic dialects. We present results on a bootstrapping technique we use to speed up the CODA annotation, as well as on the degree of similarity across the dialects before and after CODA annotation. Fadhl Eryani, Nizar Habash, Houda Bouamor, Salam Khalifa |
LREC | 2 |
| 2020 | Morphological Analysis and Disambiguation for Gulf Arabic: The Interplay between Resources and MethodsabstractIn this paper we present the first full morphological analysis and disambiguation system for Gulf Arabic. We use an existing state-of-the-art morphological disambiguation system to investigate the effects of different data sizes and different combinations of morphological analyzers for Modern Standard Arabic, Egyptian Arabic, and Gulf Arabic. We find that in very low settings, morphological analyzers help boost the performance of the full morphological disambiguation task. However, as the size of resources increase, the value of the morphological analyzers decreases. Salam Khalifa, Nasser Zalmout, Nizar Habash |
LREC | 3 |
| 2020 | CAMeL Tools: An Open Source Python Toolkit for Arabic Natural Language ProcessingabstractWe present CAMeL Tools, a collection of open-source tools for Arabic natural language processing in Python. CAMeL Tools currently provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and Sentiment Analysis. In this paper, we describe the design of CAMeL Tools and the functionalities it provides. Ossama Obeid, Nasser Zalmout, Salam Khalifa, Dima Taji, Mai Oudah, Bashar Alhafni, Go Inoue, Fadhl Eryani, Alexander Erdmann, Nizar Habash |
LREC | 10 |
| 2020 | A Link Prediction Approach for Accurately Mapping a Large-scale Arabic Lexical Resource to English WordNetabstractSuccess of Natural Language Processing (NLP) models, just like all advanced machine learning models, rely heavily on large -scale lexical resources. For English, English WordNet (EWN) is a leading example of a large-scale resource that has enabled advances in Natural Language Understanding (NLU) tasks such as word sense disambiguation, question answering, sentiment analysis, and emotion recognition. EWN includes sets of cognitive synonyms called synsets, which are interlinked by means of conceptual-semantic and lexical relations and where each synset expresses a distinct concept. However, other languages are still lagging behind in having large-scale and rich lexical resources similar to EWN. In this article, we focus on enabling the development of such resources for Arabic. While there have been efforts in developing an Arabic WordNet (AWN), the current version of AWN has its limitations in size and in lacking transliteration standards, which are important for compatibility with Arabic NLP tools. Previous efforts for extending AWN resulted in a lexicon, called ArSenL, that overcame the size and the transliteration standard limitation but was limited in accuracy due to the heuristic approach that only considered surface matching between the English definitions from the Standard Arabic Morphological Analyzer (SAMA) and EWN synset terms, and that resulted in inaccurate mapping of Arabic lemmas to EWN’s synsets. Furthermore, there has been limited exploration of other expansion methods due to expensive manual validation needed. To address these limitations of simultaneously having large-scale size with high accuracy and standard representations, the mapping problem is formulated as a link prediction problem between a large-scale Arabic lexicon and EWN, where a word in one lexicon is linked to a word in another lexicon if the two words are semantically related. We use a semi-supervised approach to create a training dataset by finding common terms in the large-scale Arabic resource and AWN. This set of data becomes implicitly linked to EWN and can be used for training and evaluating prediction models. We propose the use of a two-step Boosting method, where the first step aims at linking English translations of SAMA’s terms to EWN’s synsets. The second step uses surface similarity between SAMA’s glosses and EWN’s synsets. The method results in a new large-scale Arabic lexicon that we call ArSenL 2.0 as a sequel to the previously developed sentiment lexicon ArSenL. A comprehensive study covering both intrinsic and extrinsic evaluations shows the superiority of the method compared to several baseline and state-of-the-art link prediction methods. Compared to previously developed ArSenL, ArSenL 2.0 included a larger set of sentimentally charged adjectives and verbs. It also showed higher linking accuracy on the ground truth data compared to previous ArSenL. For extrinsic evaluation, ArSenL 2.0 was used for sentiment analysis and showed, here, too, higher accuracy compared to previous ArSenL. Gilbert Badaro, Hazem M. Hajj, Nizar Habash |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2019 | The Effectiveness of Simple Hybrid Systems for Hypernym DiscoveryabstractHypernymy modeling has largely been separated according to two paradigms, patternbased methods and distributional methods.However, recent works utilizing a mix of these strategies have yielded state-of-the-art results.This paper evaluates the contribution of both paradigms to hybrid success by evaluating the benefits of hybrid treatment of baseline models from each paradigm.Even with a simple methodology for each individual system, utilizing a hybrid approach establishes new stateof-the-art results on two domain-specific English hypernym discovery tasks and outperforms all non-hybrid approaches in a general English hypernym discovery task. William Held, Nizar Habash |
ACL (1) | 2 |
| 2019 | Adversarial Multitask Learning for Joint Multi-Feature and Multi-Dialect Morphological ModelingabstractMorphological tagging is challenging for morphologically rich languages due to the large target space and the need for more training data to minimize model sparsity.Dialectal variants of morphologically rich languages suffer more as they tend to be more noisy and have less resources.In this paper we explore the use of multitask learning and adversarial training to address morphological richness and dialectal variations in the context of full morphological tagging.We use multitask learning for joint morphological modeling for the features within two dialects, and as a knowledge-transfer scheme for crossdialectal modeling.We use adversarial training to learn dialect invariant features that can help the knowledge-transfer scheme from the high to low-resource variants.We work with two dialectal variants: Modern Standard Arabic (high-resource "dialect" 1 ) and Egyptian Arabic (low-resource dialect) as a case study.Our models achieve state-of-the-art results for both.Furthermore, adversarial training provides more significant improvement when using smaller training datasets in particular. Nasser Zalmout, Nizar Habash |
ACL (1) | 2 |
| 2019 | Towards Variability Resistant Dialectal Speech Evaluation
Ahmed Ali 0002, Salam Khalifa, Nizar Habash |
INTERSPEECH | 3 |
| 2019 | The Impact of Preprocessing on Arabic-English Statistical and Neural Machine Translation
Mai Oudah, Amjad Almahairi, Nizar Habash |
MTSummit (1) | 3 |
| 2019 | A Survey of Opinion Mining in Arabic: A Comprehensive System Perspective Covering Challenges and Advances in Tools, Resources, Models, Applications, and VisualizationsabstractOpinion-mining or sentiment analysis continues to gain interest in industry and academics. While there has been significant progress in developing models for sentiment analysis, the field remains an active area of research for many languages across the world, and in particular for the Arabic language, which is the fifth most-spoken language and has become the fourth most-used language on the Internet. With the flurry of research activity in Arabic opinion mining, several researchers have provided surveys to capture advances in the field. While these surveys capture a wealth of important progress in the field, the fast pace of advances in machine learning and natural language processing (NLP) necessitates a continuous need for a more up-to-date literature survey. The aim of this article is to provide a comprehensive literature survey for state-of-the-art advances in Arabic opinion mining. The survey goes beyond surveying previous works that were primarily focused on classification models. Instead, this article provides a comprehensive system perspective by covering advances in different aspects of an opinion-mining system, including advances in NLP software tools, lexical sentiment and corpora resources, classification models, and applications of opinion mining. It also presents future directions for opinion mining in Arabic. The survey also covers latest advances in the field, including deep learning advances in Arabic Opinion Mining. The article provides state-of-the-art information to help new or established researchers in the field as well as industry developers who aim to deploy an operational complete opinion-mining system. Key insights are captured at the end of each section for particular aspects of the opinion-mining system giving the reader a choice of focusing on particular aspects of interest. Gilbert Badaro, Ramy Baly, Hazem M. Hajj, Wassim El-Hajj, Khaled B. Shaban, Nizar Habash, Ahmad A. Al Sallab, Ali Hamdi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2018 | Utilizing Character and Word Embeddings for Text Normalization with Sequence-to-Sequence ModelsabstractText normalization is an important enabling technology for several NLP tasks.Recently, neural-network-based approaches have outperformed well-established models in this task.However, in languages other than English, there has been little exploration in this direction.Both the scarcity of annotated data and the complexity of the language increase the difficulty of the problem.To address these challenges, we use a sequence-to-sequence model with character-based attention, which in addition to its self-learned character embeddings, uses word embeddings pre-trained with an approach that also models subword information.This provides the neural model with access to more linguistic information especially suitable for text normalization, without large parallel corpora.We show that providing the model with word-level features bridges the gap for the neural network approach to achieve a state-of-the-art F 1 score on a standard Arabic language correction shared task dataset. Daniel Watson, Nasser Zalmout, Nizar Habash |
EMNLP | 3 |
| 2018 | A Leveled Reading Corpus of Modern Standard Arabic
Muhamed Al-Khalil, Hind Saddiki, Nizar Habash, Latifa Al-Sulaiti |
LREC | 3 |
| 2018 | The MADAR Arabic Dialect Corpus and Lexicon
Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, Kemal Oflazer |
LREC | 2 |
| 2018 | Unified Guidelines and Resources for Arabic Dialect Orthography
Nizar Habash, Fadhl Eryani, Salam Khalifa, Owen Rambow, Dana Abdulrahim, Alexander Erdmann, Reem Faraj, Wajdi Zaghouani, Houda Bouamor, Nasser Zalmout, Sara Hassan, Faisal Al-Shargi, Sakhar B. Alkhereyf, Basma Abdulkareem, Ramy Eskander, Mohammad Salameh, Hind Saddiki |
LREC | 1 |
| 2018 | A Parallel Corpus of Arabic-Japanese News Articles
Go Inoue, Nizar Habash, Yuji Matsumoto 0001, Hiroyuki Aoyama |
LREC | 2 |
| 2018 | Palmyra: A Platform Independent Dependency Annotation Tool for Morphologically Rich Languages
Talha Javed, Nizar Habash, Dima Taji |
LREC | 2 |
| 2018 | A Morphologically Annotated Corpus of Emirati Arabic
Salam Khalifa, Nizar Habash, Fadhl Eryani, Ossama Obeid, Dana Abdulrahim, Meera Al Kaabi |
LREC | 2 |
| 2018 | CoNLL-UL: Universal Morphological Lattices for Universal Dependency Parsing
Amir More, Özlem Çetinoglu, Çagri Çöltekin, Nizar Habash, Benoît Sagot, Djamé Seddah, Dima Taji, Reut Tsarfaty |
LREC | 4 |
| 2018 | MADARi: A Web Interface for Joint Arabic Morphological Annotation and Spelling Correction
Ossama Obeid, Salam Khalifa, Nizar Habash, Houda Bouamor, Wajdi Zaghouani, Kemal Oflazer |
LREC | 3 |
| 2018 | Noise-Robust Morphological Disambiguation for Dialectal ArabicabstractNasser Zalmout, Alexander Erdmann, Nizar Habash. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Nasser Zalmout, Alexander Erdmann, Nizar Habash |
NAACL-HLT | 3 |
| 2018 | A Bilingual Interactive Human Avatar Dialogue SystemabstractThis demonstration paper presents a bilingual (Arabic-English) interactive human avatar dialogue system.The system is named TOIA (time-offset interaction application), as it simulates face-to-face conversations between humans using digital human avatars recorded in the past.TOIA is a conversational agent, similar to a chat bot, except that it is based on an actual human being and can be used to preserve and tell stories.The system is designed to allow anybody, simply using a laptop, to create an avatar of themselves, thus facilitating cross-cultural and cross-generational sharing of narratives to wider audiences.The system currently supports monolingual and cross-lingual dialogues in Arabic and English, but can be extended to other languages. Dana Abu Ali, Muaz Ahmad, Hayat Al Hassan, Paula Dozsa, Jose Varias, Nizar Habash |
SIGDIAL Conference | 7 |
| 2017 | Don't Throw Those Morphological Analyzers Away Just Yet: Neural Morphological Disambiguation for ArabicabstractThis paper presents a model for Arabic morphological disambiguation based on Recurrent Neural Networks (RNN).We train Long Short-Term Memory (LSTM) cells in several configurations and embedding levels to model the various morphological features.Our experiments show that these models outperform state-of-theart systems without explicit use of feature engineering.However, adding learning features from a morphological analyzer to model the space of possible analyses provides additional improvement.We make use of the resulting morphological models for scoring and ranking the analyses of the morphological analyzer for morphological disambiguation.The results show significant gains in accuracy across several evaluation metrics.Our system results in 4.4% absolute increase over the state-of-the-art in full morphological analysis accuracy (30.6% relative error reduction), and 10.6% (31.5% relative error reduction) for out-of-vocabulary words. Nasser Zalmout, Nizar Habash |
EMNLP | 2 |
| 2017 | Low Resourced Machine Translation via Morpho-syntactic Modeling: The Case of Dialectal Arabic
Alexander Erdmann, Nizar Habash, Dima Taji, Houda Bouamor |
MTSummit (1) | 2 |
| 2017 | A Sentiment Treebank and Morphologically Enriched Recursive Deep Models for Effective Sentiment Analysis in ArabicabstractAccurate sentiment analysis models encode the sentiment of words and their combinations to predict the overall sentiment of a sentence. This task becomes challenging when applied to morphologically rich languages (MRL). In this article, we evaluate the use of deep learning advances, namely the Recursive Neural Tensor Networks (RNTN), for sentiment analysis in Arabic as a case study of MRLs. While Arabic may not be considered the only representative of all MRLs, the challenges faced and proposed solutions in Arabic are common to many other MRLs. We identify, illustrate, and address MRL-related challenges and show how RNTN is affected by the morphological richness and orthographic ambiguity of the Arabic language. To address the challenges with sentiment extraction from text in MRL, we propose to explore different orthographic features as well as different morphological features at multiple levels of abstraction ranging from raw words to roots. A key requirement for RNTN is the availability of a sentiment treebank; a collection of syntactic parse trees annotated for sentiment at all levels of constituency and that currently only exists in English. Therefore, our contribution also includes the creation of the first Arabic Sentiment Treebank (A r S en TB) that is morphologically and orthographically enriched. Experimental results show that, compared to the basic RNTN proposed for English, our solution achieves significant improvements up to 8% absolute at the phrase level and 10.8% absolute at the sentence level, measured by average F1 score. It also outperforms well-known classifiers including Support Vector Machines, Recursive Auto Encoders, and Long Short-Term Memory by 7.6%, 3.2%, and 1.6% absolute respectively, all models being trained with similar morphological considerations. Ramy Baly, Hazem M. Hajj, Nizar Habash, Khaled B. Shaban, Wassim El-Hajj |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2016 | Creating Resources for Dialectal Arabic from a Single Annotation: A Case Study on Egyptian and LevantineabstractArabic dialects present a special problem for natural language processing because there are few resources, they have no standard orthography, and have not been studied much. However, as more and more written dialectal Arabic is found in social media, NLP for Arabic dialects becomes an important goal. We present a methodology for creating a morphological analyzer and a morphological tagger for dialectal Arabic, and we illustrate it on Egyptian and Levantine Arabic. To our knowledge, these are the first analyzer and tagger for Levantine. Ramy Eskander, Nizar Habash, Owen Rambow, Arfath Pasha |
COLING | 2 |
| 2016 | Machine Translation Evaluation for Arabic using Morphologically-enriched EmbeddingsabstractEvaluation of machine translation (MT) into morphologically rich languages (MRL) has not been well studied despite posing many challenges. In this paper, we explore the use of embeddings obtained from different levels of lexical and morpho-syntactic linguistic analysis and show that they improve MT evaluation into an MRL. Specifically we report on Arabic, a language with complex and rich morphology. Our results show that using a neural-network model with different input representations produces results that clearly outperform the state-of-the-art for MT evaluation into Arabic, by almost over 75% increase in correlation with human judgments on pairwise MT evaluation quality task. More importantly, we demonstrate the usefulness of morpho-syntactic representations to model sentence similarity for MT evaluation and address complex linguistic phenomena of Arabic. Francisco Guzmán, Houda Bouamor, Ramy Baly, Nizar Habash |
COLING | 4 |
| 2016 | SPLIT: Smart Preprocessing (Quasi) Language Independent Tool
Mohamed Al-Badrashiny, Arfath Pasha, Mona T. Diab, Nizar Habash, Owen Rambow, Wael Salloum, Ramy Eskander |
LREC | 4 |
| 2016 | Morphologically Annotated Corpora and Morphological Analyzers for Moroccan and Sanaani Yemeni Arabic
Faisal Al-Shargi, Aidan Kaplan, Ramy Eskander, Nizar Habash, Owen Rambow |
LREC | 4 |
| 2016 | Exploiting Arabic Diacritization for High Quality Automatic Annotation
Nizar Habash, Anas Shahrour, Muhamed Al-Khalil |
LREC | 1 |
| 2016 | DALILA: The Dialectal Arabic Linguistic Learning Assistant
Salam Khalifa, Houda Bouamor, Nizar Habash |
LREC | 3 |
| 2016 | A Large Scale Corpus of Gulf Arabic
Salam Khalifa, Nizar Habash, Dana Abdulrahim, Sara Hassan |
LREC | 2 |
| 2016 | Applying the Cognitive Machine Translation Evaluation Approach to Arabic
Irina P. Temnikova, Wajdi Zaghouani, Stephan Vogel, Nizar Habash |
LREC | 4 |
| 2016 | Arabic Corpora for Credibility Analysis
Ayman Al Zaatari, Rim El Ballouli, Shady Elbassuoni, Wassim El-Hajj, Hazem M. Hajj, Khaled B. Shaban, Nizar Habash, Emad Yahya |
LREC | 7 |
| 2016 | Building an Arabic Machine Translation Post-Edited Corpus: Guidelines and Annotation
Wajdi Zaghouani, Nizar Habash, Ossama Obeid, Behrang Mohit, Houda Bouamor, Kemal Oflazer |
LREC | 2 |
| 2015 | Arabic Transliteration of Romanized Tunisian Dialect Text: A Preliminary Investigation
Abir Masmoudi 0001, Nizar Habash, Mariem Ellouze, Yannick Estève, Lamia Hadrich Belguith |
CICLing (1) | 2 |
| 2015 | Predicting the Structure of Cooking RecipesabstractCooking recipes exist in abundance; but due to their unstructured text format, they are hard to study quantitatively beyond treating them as simple bags of words. In this paper, we propose an ingredient-instruction dependency tree data structure to represent recipes. The proposed rep-resentation allows for more refined com-parison of recipes and recipe-parts, and is a step towards semantic representation of recipes. Furthermore, we build a parser that maps recipes into the proposed rep-resentation. The parser’s edge prediction accuracy of 93.5 % improves over a strong baseline of 85.7 % (54.5 % error reduction). 1 Jermsak Jermsurawong, Nizar Habash |
EMNLP | 2 |
| 2015 | Improving Arabic Diacritization through Syntactic AnalysisabstractWe present an approach to Arabic automatic diacritization that integrates syntactic analysis with morphological tagging through improving the prediction of case and state features.Our best system increases the accuracy of word diacritization by 2.5% absolute on all words, and 5.2% absolute on nominals over a state-of-theart baseline.Similar increases are shown on the full morphological analysis choice. Anas Shahrour, Salam Khalifa, Nizar Habash |
EMNLP | 3 |
| 2015 | Morphological constraints for phrase pivot statistical machine translation
Ahmed El Kholy, Nizar Habash |
MTSummit | 2 |
| 2014 | Unsupervised Morphology-Based Vocabulary ExpansionabstractWe present a novel way of generating unseen words, which is useful for certain applications such as automatic speech recognition or optical character recognition in low-resource languages. We test our vocabulary generator on seven low-resource languages by measuring the decrease in out-of-vocabulary word rate on a held-out test set. The languages we study have very different morphological properties; we show how our results differ depending on the morphological complexity of the language. In our best result (on Assamese), our approach can predict 29% of the token-based out-of-vocabulary with a small amount of unlabeled training data. Mohammad Sadegh Rasooli, Tom Lippincott, Nizar Habash, Owen Rambow |
ACL (1) | 3 |
| 2014 | Automatic Transliteration of Romanized Dialectal ArabicabstractIn this paper, we address the problem of converting Dialectal Arabic (DA) text that is written in the Latin script (called Arabizi) into Arabic script following the CODA convention for DA orthography. The presented system uses a finite state transducer trained at the character level to generate all possible transliterations for the input Arabizi words. We then filter the generated list using a DA morpholog-ical analyzer. After that we pick the best choice for each input word using a lan-guage model. We achieve an accuracy of 69.4 % on an unseen test set compared to 63.1 % using a system which represents a previously proposed approach. 1 Mohamed Al-Badrashiny, Ramy Eskander, Nizar Habash, Owen Rambow |
CoNLL | 3 |
| 2014 | Alignment symmetrisation optimization targeting phrase pivot statistical machine translation
Ahmed El Kholy, Nizar Habash |
EAMT | 2 |
| 2014 | Improving deep neural network acoustic modeling for audio corpus indexing under the IARPA babel program
Brian Kingsbury, Jia Cui, Bhuvana Ramabhadran, Andrew Rosenberg, Mohammad Sadegh Rasooli, Owen Rambow, Nizar Habash, Vaibhava Goel |
INTERSPEECH | 8 |
| 2014 | A Multidialectal Parallel Corpus of Arabic
Houda Bouamor, Nizar Habash, Kemal Oflazer |
LREC | 2 |
| 2014 | Tharwa: A Large Scale Dialectal Arabic - Standard Arabic - English Lexicon
Mona T. Diab, Mohamed Al-Badrashiny, Maryam Aminian, Heba Elfardy, Nizar Habash, Abdelati Hawwari, Wael Salloum, Pradeep Dasigi, Ramy Eskander |
LREC | 6 |
| 2014 | Developing an Egyptian Arabic Treebank: Impact of Dialectal Morphology on Annotation and Tool Development
Mohamed Maamouri, Ann Bies, Seth Kulick, Michael Ciul, Nizar Habash, Ramy Eskander |
LREC | 5 |
| 2014 | A Corpus and Phonetic Dictionary for Tunisian Arabic Speech Recognition
Abir Masmoudi 0001, Mariem Ellouze, Yannick Estève, Lamia Hadrich Belguith, Nizar Habash |
LREC | 5 |
| 2014 | MADAMIRA: A Fast, Comprehensive Tool for Morphological Analysis and Disambiguation of Arabic
Arfath Pasha, Mohamed Al-Badrashiny, Mona T. Diab, Ahmed El Kholy, Ramy Eskander, Nizar Habash, Manoj Pooleery, Owen Rambow, Ryan Roth |
LREC | 6 |
| 2014 | Large Scale Arabic Error Annotation: Guidelines and Framework
Wajdi Zaghouani, Behrang Mohit, Nizar Habash, Ossama Obeid, Nadi Tomeh, Alla Rozovskaya, Noura Farra, Sarah Alkuhlani, Kemal Oflazer |
LREC | 3 |
| 2014 | A Conventional Orthography for Tunisian Arabic
Inès Zribi, Rahma Boujelben, Abir Masmoudi 0001, Mariem Ellouze, Lamia Hadrich Belguith, Nizar Habash |
LREC | 6 |
| 2013 | Automatic Extraction of Morphological Lexicons from Morphologically Annotated CorporaabstractWe present a method for automatically learning inflectional classes and associated lemmas from morphologically annotated corpora.The method consists of a core languageindependent algorithm, which can be optimized for specific languages.The method is demonstrated on Egyptian Arabic and German, two morphologically rich languages.Our best method for Egyptian Arabic provides an error reduction of 55.6% over a simple baseline; our best method for German achieves a 66.7% error reduction. Ramy Eskander, Nizar Habash, Owen Rambow |
EMNLP | 2 |
| 2013 | Selective Combination of Pivot and Direct Statistical Machine Translation Models
Ahmed El Kholy, Nizar Habash, Gregor Leusch, Evgeny Matusov, Hassan Sawaf |
IJCNLP | 2 |
| 2013 | A Web-based Annotation Framework For Large-Scale Text Correction
Ossama Obeid, Wajdi Zaghouani, Behrang Mohit, Nizar Habash, Kemal Oflazer, Nadi Tomeh |
IJCNLP | 4 |
| 2013 | DIRA: Dialectal Arabic Information Retrieval Assistant
Arfath Pasha, Mohamed Al-Badrashiny, Mohamed Altantawy, Nizar Habash, Manoj Pooleery, Owen Rambow, Ryan Roth, Mona T. Diab |
IJCNLP | 4 |
| 2013 | Orthographic and Morphological Processing for Persian-to-English Statistical Machine Translation
Mohammad Sadegh Rasooli, Ahmed El Kholy, Nizar Habash |
IJCNLP | 3 |
| 2013 | The Effects of Factorizing Root and Pattern Mapping in Bidirectional Tunisian - Standard Arabic Machine Translation
Ahmed Hamdi, Rahma Boujelben, Nizar Habash, Alexis Nasr |
MTSummit | 3 |
| 2013 | Automatic Morphological Enrichment of a Morphologically Underspecified Treebank
Sarah Alkuhlani, Nizar Habash, Ryan Roth |
HLT-NAACL | 2 |
| 2013 | Processing Spontaneous Orthography
Ramy Eskander, Nizar Habash, Owen Rambow, Nadi Tomeh |
HLT-NAACL | 2 |
| 2013 | Morphological Analysis and Disambiguation for Dialectal Arabic
Nizar Habash, Ryan Roth, Owen Rambow, Ramy Eskander, Nadi Tomeh |
HLT-NAACL | 1 |
| 2013 | Dialectal Arabic to English Machine Translation: Pivoting through Modern Standard Arabic
Wael Salloum, Nizar Habash |
HLT-NAACL | 2 |
| 2013 | Dependency Parsing of Modern Standard Arabic with Lexical and Inflectional FeaturesabstractWe explore the contribution of lexical and inflectional morphology features to dependency parsing of Arabic, a morphologically rich language with complex agreement patterns. Using controlled experiments, we contrast the contribution of different part-of-speech (POS) tag sets and morphological features in two input conditions: machine-predicted condition (in which POS tags and morphological feature values are automatically assigned), and gold condition (in which their true values are known). We find that more informative (fine-grained) tag sets are useful in the gold condition, but may be detrimental in the predicted condition, where they are outperformed by simpler but more accurately predicted tag sets. We identify a set of features (definiteness, person, number, gender, and undiacritized lemma) that improve parsing quality in the predicted condition, whereas other features are more useful in gold. We are the first to show that functional features for gender and number (e.g., “broken plurals”), and optionally the related rationality (“humanness”) feature, are more helpful for parsing than form-based gender and number. We finally show that parsing quality in the predicted condition can dramatically improve by training in a combined gold+predicted condition. We experimented with two transition-based parsers, MaltParser and Easy-First Parser. Our findings are robust across parsers, models, and input conditions. This suggests that the contribution of the linguistic knowledge in the tag sets and features we identified goes beyond particular experimental settings, and may be informative for other parsers and morphologically rich languages. Yuval Marton, Nizar Habash, Owen Rambow |
Comput. Linguistics | 2 |
| 2012 | Identifying Broken Plurals, Irregular Gender, and Rationality in Arabic Text
Sarah Alkuhlani, Nizar Habash |
EACL | 2 |
| 2012 | Translate, Predict or Generate: Modeling Rich Morphology in Statistical Machine Translation
Ahmed El Kholy, Nizar Habash |
EAMT | 2 |
| 2012 | Can Automatic Post-Editing Make MT More Meaningful
Kristen Parton, Nizar Habash, Kathy McKeown, Gonzalo Iglesias, Adrià de Gispert |
EAMT | 2 |
| 2012 | Hebrew Morphological Preprocessing for Statistical Machine Translation
Nimesh Singh, Nizar Habash |
EAMT | 2 |
| 2012 | Rich Morphology Generation Using Statistical Machine Translation
Ahmed El Kholy, Nizar Habash |
INLG | 2 |
| 2012 | Conventional Orthography for Dialectal Arabic
Nizar Habash, Mona T. Diab, Owen Rambow |
LREC | 1 |
| 2012 | Arabic Dialect Processing Tutorial
Mona T. Diab, Nizar Habash |
HLT-NAACL | 2 |
| 2012 | Improved Arabic-to-English statistical machine translation by reordering post-verbal subjects for word alignment
Marine Carpuat, Yuval Marton, Nizar Habash |
Mach. Transl. | 3 |
| 2012 | Special issue on Machine Translation for Arabic: Preface
Nizar Habash, Hany Hassan |
Mach. Transl. | 1 |
| 2012 | Orthographic and morphological processing for English-Arabic statistical machine translation
Ahmed El Kholy, Nizar Habash |
Mach. Transl. | 2 |
| 2012 | Machine translation between Hebrew and Arabic
Reshef Shilon, Nizar Habash, Alon Lavie, Shuly Wintner |
Mach. Transl. | 2 |
| 2011 | Using Deep Morphology to Improve Automatic Error Detection in Arabic Handwriting Recognition
Nizar Habash, Ryan Roth |
ACL | 1 |
| 2011 | Improving Arabic Dependency Parsing with Form-based and Functional Morphological Features
Yuval Marton, Nizar Habash, Owen Rambow |
ACL | 2 |
| 2011 | Automatic Error Analysis for Morphologically Rich Languages
Ahmed El Kholy, Nizar Habash |
MTSummit | 2 |
| 2010 | Morphological Analysis and Generation of Arabic Nouns: A Morphemic Functional Approach
Mohamed Altantawy, Nizar Habash, Owen Rambow, Ibrahim Saleh |
LREC | 2 |
| 2010 | Morphological Annotation of Quranic Arabic
Kais Dukes, Nizar Habash |
LREC | 2 |
| 2010 | Interlingual annotation of parallel text corpora: a new framework for annotation and evaluationabstractAbstract This paper focuses on an important step in the creation of a system of meaning representation and the development of semantically annotated parallel corpora, for use in applications such as machine translation, question answering, text summarization, and information retrieval. The work described below constitutes the first effort of any kind to annotate multiple translations of foreign-language texts with interlingual content. Three levels of representation are introduced: deep syntactic dependencies (IL0), intermediate semantic representations (IL1), and a normalized representation that unifies conversives, nonliteral language, and paraphrase (IL2). The resulting annotated, multilingually induced, parallel corpora will be useful as an empirical basis for a wide range of research, including the development and evaluation of interlingual NLP systems and paraphrase-extraction systems as well as a host of other research and development efforts in theoretical and applied linguistics, foreign language pedagogy, translation studies, and other related disciplines. Bonnie J. Dorr, Rebecca J. Passonneau, David Farwell, Rebecca Green, Nizar Habash, Stephen Helmreich, Eduard H. Hovy, Lori S. Levin, Keith J. Miller, Teruko Mitamura, Owen Rambow, Advaith Siddharthan |
Nat. Lang. Eng. | 5 |
| 2009 | Improving the Arabic Pronunciation Dictionary for Phone and Word Recognition with Linguistically-Based Pronunciation Rules
Fadi Biadsy, Nizar Habash, Julia Hirschberg |
HLT-NAACL | 2 |
| 2009 | Symbolic-to-statistical hybridization: extending generation-heavy machine translationabstractThe last few years have witnessed an increasing interest in hybridizing surface-based statistical approaches and rule-based symbolic approaches to machine translation (MT). Much of that work is focused on extending statistical MT systems with symbolic knowledge and components. In the brand of hybridization discussed here, we go in the opposite direction: adding statistical bilingual components to a symbolic system. Our base system is Generation-heavy machine translation (GHMT), a primarily symbolic asymmetrical approach that addresses the issue of Interlingual MT resource poverty in source-poor/target-rich language pairs by exploiting symbolic and statistical target-language resources. GHMT’s statistical components are limited to target-language models, which arguably makes it a simple form of a hybrid system . We extend the hybrid nature of GHMT by adding statistical bilingual components. We also describe the details of retargeting it to Arabic–English MT. The morphological richness of Arabic brings several challenges to the hybridization task. We conduct an extensive evaluation of multiple system variants. Our evaluation shows that this new variant of GHMT—a primarily symbolic system extended with monolingual and bilingual statistical components—has a higher degree of grammaticality than a phrase-based statistical MT system, where grammaticality is measured in terms of correct verb-argument realization and long-distance dependency translation. Nizar Habash, Bonnie J. Dorr, Christof Monz |
Mach. Transl. | 1 |
| 2008 | Improving NER in Arabic Using a Morphological Tagger
Benjamin Farber, Dayne Freitag, Nizar Habash, Owen Rambow |
LREC | 3 |
| 2008 | Identification of Naturally Occurring Numerical Expressions in Arabic
Nizar Habash, Ryan Roth |
LREC | 1 |
| 2007 | Determining Case in Arabic: Learning Complex Linguistic Behavior Requires Complex Linguistic Features
Nizar Habash, Ryan Gabbard, Owen Rambow, Seth Kulick, Mitchell P. Marcus |
EMNLP-CoNLL | 1 |
| 2007 | Arabic diacritization in the context of statistical machine translation
Mona T. Diab, Mahmoud Ghoneim, Nizar Habash |
MTSummit | 3 |
| 2007 | Syntactic preprocessing for statistical machine translation
Nizar Habash |
MTSummit | 1 |
| 2007 | Semi-automatic error analysis for large-scale statistical machine translation
Katrin Kirchhoff, Owen Rambow, Nizar Habash, Mona T. Diab |
MTSummit | 3 |
| 2006 | MAGEAD: A Morphological Analyzer and Generator for the Arabic DialectsabstractWe present MAGEAD, a morphological analyzer and generator for the Arabic language family.Our work is novel in that it explicitly addresses the need for processing the morphology of the dialects.MAGEAD performs an on-line analysis to or generation from a root+pattern+features representation, it has separate phonological and orthographic representations, and it allows for combining morphemes from different dialects.We present a detailed evaluation of MAGEAD. Nizar Habash, Owen Rambow |
ACL | 1 |
| 2006 | Combination of Arabic Preprocessing Schemes for Statistical Machine TranslationabstractStatistical machine translation is quite robust when it comes to the choice of input representation. It only requires consistency between training and testing. As a result, there is a wide range of possible preprocessing choices for data used in statistical machine translation. This is even more so for morphologically rich languages such as Arabic. In this paper, we study the effect of different word-level preprocessing schemes for Arabic on the quality of phrase-based statistical machine translation. We also present and evaluate different methods for combining preprocessing schemes resulting in improved translation quality. Fatiha Sadat, Nizar Habash |
ACL | 2 |
| 2006 | Parsing Arabic Dialects
David Chiang 0001, Mona T. Diab, Nizar Habash, Owen Rambow, Safiullah Shareef |
EACL | 3 |
| 2006 | Design, Construction and Validation of an Arabic-English Conceptual Interlingua for Cross-lingual Information Retrieval
Nizar Habash, Clinton Mah, Sabiha Imran, Randall J. Calistri-Yeh, Paraic Sheridan |
LREC | 1 |
| 2006 | Developing and Using a Pilot Dialectal Arabic Treebank
Mohamed Maamouri, Ann Bies, Tim Buckwalter, Mona T. Diab, Nizar Habash, Owen Rambow, Dalila Tabessi |
LREC | 5 |
| 2006 | Inter-annotator Agreement on a Multilingual Semantic Annotation Task
Rebecca J. Passonneau, Nizar Habash, Owen Rambow |
LREC | 2 |
| 2006 | Parallel Syntactic Annotation of Multiple Languages
Owen Rambow, Bonnie J. Dorr, David Farwell, Rebecca Green, Nizar Habash, Stephen Helmreich, Eduard H. Hovy, Lori S. Levin, Keith J. Miller, Teruko Mitamura, Flo Reeder, Advaith Siddharthan |
LREC | 5 |
| 2006 | Arabic Preprocessing Schemes for Statistical Machine Translation
Nizar Habash, Fatiha Sadat |
HLT-NAACL | 1 |
| 2005 | Arabic Tokenization, Part-of-Speech Tagging and Morphological Disambiguation in One Fell SwoopabstractWe present an approach to using a morphological analyzer for tokenizing and morphologically tagging (including partof-speech tagging) Arabic words in one process.We learn classifiers for individual morphological features, as well as ways of using these classifiers to choose among entries from the output of the analyzer.We obtain accuracy rates on all tasks in the high nineties. Nizar Habash, Owen Rambow |
ACL | 1 |
| 2004 | The Use of a Structural N-gram Language Model in Generation-Heavy Hybrid Machine Translation
Nizar Habash |
INLG | 1 |
| 2003 | Matador: a large-scale Spanish-English GHMT systemabstractThis paper describes and evaluates Matador, an implemented large-scale Spanish-English MT system built in the Generation-Heavy Hybrid Machine Translation (GHMT) approach. An extensive evaluation shows that Matador has a higher degree of robustness and superior output quality, in terms of grammaticality and accuracy, when compared to a primarily statistical approach. Nizar Habash |
MTSummit | 1 |
| 2003 | A Categorial Variation Database for English
Nizar Habash, Bonnie J. Dorr |
HLT-NAACL | 1 |
| 2003 | Hybrid Natural Language Generation from Lexical Conceptual Structures
Nizar Habash, Bonnie J. Dorr, David R. Traum |
Mach. Transl. | 1 |
| 2003 | Rapid porting of DUSTer to HindiabstractThe frequent occurrence of divergences —structural differences between languages---presents a great challenge for statistical word-level alignment and machine translation. This paper describes the adaptation of DUSTer, a divergence unraveling package, to Hindi during the DARPA TIDES-2003 Surprise Language Exercise. We show that it is possible to port DUSTer to Hindi in under 3 days. Bonnie J. Dorr, Necip Fazil Ayan, Nizar Habash, Nitin Madnani, Rebecca Hwa |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2002 | Generation-Heavy Hybrid Machine Translation
Nizar Habash |
INLG | 1 |
| 2001 | Large scale language independent generation using thematic hierarchies
Nizar Habash, Bonnie J. Dorr |
MTSummit | 1 |