Lamia Hadrich Belguith

dblp:59/1682 · DBLP profile ↗
← Back
79ranked-venue papers
0as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 8 since 2021Databases, data management, data science and information retrieval · 11 · 2 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Explainable Federated Deep Reinforcement Learning for Distributed Trading Agents over 6G-Enabled Wireless Edge Networks
Achraf Wali, Isaac Lera, Lamia Hadrich Belguith
IWCMC3
2026 Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG
Jaafer Klila, Sondes Bannour Souihi, Rahma Boujelben, Nasredine Semmar, Lamia Hadrich Belguith
LREC5
2025 Automatic Code-switched Academic Tunisian Arabic Speech Recognition
abstract
Aiming to support researchers by answering their questions, this paper presents the first module of a virtual assistant for Tunisian Arabic (TA): Automatic Speech Recognition (ASR). Given the specialized lexicon used, which is characterized by a high rate of Code-Switching (CS), our primary objective is to create a new dataset, Tunacad, tailored to this context. Tunacad comprises $\mathbf{1 2. 5 6}$ hours of CS spontaneous speech in the academic domain. Additionally, we experiment with our proposed CS corpus on predefined ASR models, Whisper and Wav2vec2-XLR-S. To enhance ASR performance on our corpus, we apply several techniques, such as automatic orthographic correction, normalization, and prompt-based correction. Additionally, we incorporate human evaluation to assess transcription quality beyond surface-level accuracy. Through these experiments and CS errors analysis statistics, we highlight the persistent challenges in processing TA speech.
Fatma Zahra Besdouri, Inès Zribi, Lamia Hadrich Belguith
AICCSA3
2025 Tunisian Dialect Speech Corpus: Construction and Emotion Annotation
Latifa Iben Nasr, Abir Masmoudi 0001, Lamia Hadrich Belguith
ICAART (3)3
2025 Blockchain-Based Federated Learning for Enhanced Cyber-Threats Detection in Connected Vehicles
abstract
Over the past few years, there have been made significant strides in advancing the Internet of Vehicles (IoV), recognizing its strategic importance in Intelligent Transport Systems. The proliferation of connected and autonomous vehicles on the roads has propelled the IoV into the spotlight. However, addressing the specific demands of vehicular networks, such as low latency, high mobility, extensive connectivity of 5G/6G networks, and robust security, remains a substantial challenge. Therefore, there is a critical need for substantial progress in implementing a resilient Intrusion Detection System within the IoV ecosystem. This paper introduces VFed-IDS, a decentralized, secure, flexible, scalable, and robust Blockchain and Federated Learning-based intrusion detection system. VFed-IDS is designed to identify cyber threats in the IoV while preserving privacy in connected vehicles. The proposed architecture consists of three main layers: the central layer, the local layer, and the Blockchain layer. The central layer includes the SDN Controller, responsible for training and aggregating the global model. The local layer comprises vehicles training individual models based on their private local datasets. The Blockchain layer introduces the Smart Contract VFed-SC, which manages the list of authenticated and collaborating vehicles in the Federated Learning process. It also hashes trained local model updates before transmitting them as transactions between the central and local layers. Simulation results demonstrate that VFed-IDS achieves a high accuracy rate of 99%, effectively enhancing the autonomous behavior of connected vehicles against cyber threats.
Houda Amari, Zakaria Abou El Houda, Hajar Moudoud, Lyes Khoukhi, Lamia Hadrich Belguith
ICC5
2025 Enhancing Financial Forecasting with Large Language Models: Capabilities and Challenges
abstract
The growing integration of Artificial intelligence (AI) into financial markets offers unprecedented opportunities for advanced forecasting, automated decision-making, and strategic trading. However, despite notable progress, current systems often fail to balance predictive performance with critical requirements such as explainability, dynamic risk management, and regulatory alignment, especially in institutional contexts. This paper introduces the Explainable Deep Trading Approach (EDTA), a novel modular framework that combines Large Language Models (LLMs), multi-agent architectures, explainable AI (XAI) techniques, and adaptive risk management into a unified, transparent trading system. Unlike prior models focused solely on accuracy, EDTA embeds explainability and risk differentiation directly within its decision-making process, ensuring outputs that are both high-performing and regulatorily compliant under frameworks such as MiFID II. We conduct extensive empirical evaluations across diverse financial assets and temporal resolutions, demonstrating that EDTA achieves a Mean Directional Accuracy (MDA) of 72.4%, with an F1-score of 0.71, outperforming established benchmarks such as FinGPT, FinBERT, and DeepScalper. Furthermore, a domain expert panel rated EDTA’s generated explanations highly, with average scores of 4.6/5 for clarity, 4.4/5 for regulatory alignment, and 4.5/5 for practical relevance. These results highlight EDTA as a robust, next-generation framework that addresses both technological and regulatory demands, paving the way for more trustworthy, explainable, and effective AI-driven financial systems.
Achraf Wali, Isaac Lera, Lamia Hadrich Belguith, Antoni Jaume-i-Capó
KES3
2025 Semantic analysis based on ontology and deep learning for a chatbot to assist persons with personality disorders on Twitter
abstract
This paper presents a chatbot taking advantage of semantic analysis based on ontology and deep learning techniques for ensuring the monitoring of Twitter users with personality disorders during the period of COVID-19. The monitoring provided by our work consists of (i) removing inappropriate tweets from the newsfeed of the sick person according to their state, (ii) providing via a chatbot an answer to the sick person in the form of another tweet that can help him to overcome their concerns about a problem related to the epidemic. Our approach was started by detecting people having personality disorders on Twitter, followed by detecting their behaviour towards COVID-19 expressed in tweets posted in relation to this epidemic. After that, moving to perform the filtration and the recommendation tasks of tweets based on a semantic analysis. Our semantic analysis is achieved at first by querying an ontology based on a comparison taking into account concepts and behaviour expressed. Then, via a deep learning approach in order to resolve untreated cases by the ontology. For the evaluation part, we obtained an F-measure value equals to 72% for the task of filtering inappropriate tweets and 75% for the task of recommended tweet.
Mourad Ellouze, Lamia Hadrich Belguith
Behav. Inf. Technol.2
2025 Emotion Recognition from Spontaneous Tunisian Dialect Speech
abstract
Emotional expressions are a fundamental aspect of human communication, with speech being one of the most natural modes of interaction. Speech Emotion Recognition (SER) is a significant research topic in Natural Language Processing (NLP), aimed at identifying emotions such as satisfaction, frustration, and anger from speech audio using multiple classifiers. This article presents a method to emotion recognition from spontaneous Tunisian Dialect (TD) speech, marking the first work in the SER field to utilize spontaneous speech for emotion recognition in this dialect. The dataset was created from freely available YouTube videos across multiple domains and labeled with four perceived emotions: anger, satisfaction, frustration, and neutral. To address the data scarcity issue, we implemented data augmentation techniques, specifically Vocal Tract Length Perturbation (VTLP). The preprocessing of the speech signals involved cleaning the data from ambient and unwanted noises. We extracted and selected various spectral features, including Mel-Frequency Cepstral Coefficients (MFCCs) and Linear Prediction Cepstral Coefficients (LPCC). Subsequently, we applied several classification methods: Support Vector Machine (SVM), Bidirectional Long Short-Term Memory (BiLSTM), Long Short-Term Memory (LSTM), Convolutional Neural Network (CNN), and Random Forest. Our experiments demonstrated that the Random Forest classifier achieved the highest F-score of 58.75%. The results were thoroughly discussed, analyzed, and compared across the five models using different feature extractions. This study provides valuable insights and advancements in the SER field, particularly for the TD, for future research directions for improving emotion recognition systems.
Latifa Iben Nasr, Abir Masmoudi 0001, Lamia Hadrich Belguith
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2024 Dimensional Emotion Annotation for Spontaneous Tunisian Dialect Speech
abstract
Speech can convey a wide range of emotions, such as joy and anger, which can often coexist within the same expression. To accurately capture the dynamic and evolving nature of these emotions, enabling the identification of overlapping emotional nuances, our study employs a three-dimensional model—Valence, Arousal, and Dominance (VAD)—to annotate emotions. This approach provides a more nuanced understanding of emotional states within the realm of Speech Emotion Recognition (SER), specifically within the SERTUS (Speech Emotion Recognition in TUnisian Spontaneous) dataset. Notably, this research represents the first endeavor to apply this approach to annotate an Arabic language dataset in the field of SER. The paper evaluates interannotator agreement across the three dimensions, ensuring the robustness of the annotation method. This approach significantly enhances emotion annotation in Arabic SER, thereby advancing our understanding of emotional dynamics. This research is particularly aimed at SER researchers, offering a novel methodology for more precise emotion recognition.
Latifa Iben Nasr, Abir Masmoudi 0001, Lamia Hadrich Belguith
WETICE3
2024 TTK: A toolkit for Tunisian linguistic analysis
Asma Mekki, Inès Zribi, Mariem Ellouze, Lamia Hadrich Belguith
Comput. Speech Lang.4
2024 Arabic Automatic Speech Recognition: Challenges and Progress
Fatma Zahra Besdouri, Inès Zribi, Lamia Hadrich Belguith
Speech Commun.3
2024 Automatic Algerian Sarcasm Detection from Texts and Images
abstract
In recent years, the number of Algerian Internet users has significantly increased, providing a valuable opportunity for collecting and utilizing opinions and sentiments expressed online. They now post not just texts but also images. However, to benefit from this wealth of information, it is crucial to address the challenge of sarcasm detection, which poses a limitation in sentiment analysis. Sarcasm often involves the use of nonliteral and ambiguous language, making its detection complex. To enhance the quality and relevance of sentiment analysis, it is essential to develop effective methods for sarcasm detection. By overcoming this limitation, we can fully harness the expressed online opinions and benefit from their valuable insights for a better understanding of trends and sentiments among the Algerian public. In this work, our aim is to develop a comprehensive system that addresses sarcasm detection in Algerian dialect, encompassing both text and image analysis. We propose a hybrid approach that combines linguistic characteristics and machine learning techniques for text analysis. Additionally, for image analysis, we utilized the deep learning model VGG-19 for image classification, and employed the EasyOCR technique for Arabic text extraction. By integrating these approaches, we strive to create a robust system capable of detecting sarcasm in both textual and visual content in the Algerian dialect. Our system achieved an accuracy of 92.79% for the textual models and 89.28% for the visual model.
Kheira Zineb Bousmaha, Khaoula Hamadouche, Hadjer Djouabi, Lamia Hadrich Belguith
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2023 Natural Tunisian Speech Preprocessing for Features Extraction
abstract
In this paper, we describe the process of building a corpus for Tunisian Speech Emotion Recognition (SER). To the best of our knowledge, it is the first work in the SER field that uses spontaneous speech emotion in Tunisian dialect. SER represents an active research problem in the field of Natural Language processing (NLP). It aims to detect different emotions such as satisfaction, frustration and anger from audio speeches using various classifiers. Speech signal preprocessing is the first and the most important step in the SER process. Moreover, Pre-Processing of Speech is very crucial in the applications where silence or ambient noise is completely undesirable. Voice activity detection is a common procedure that plays a key role in preprocessing speech signals and noise cancellation. Pre-emphasis of speech helps the system be computationally more effective [1].This work proposed a preprocessing method to extract features from natural Tunisian speech. Speech preprocessing consists of cleaning the speech signal from ambient and unwanted noises, detecting speech activity and normalizing the length of the vocal tract.
Latifa Iben Nasr, Abir Masmoudi 0001, Lamia Hadrich Belguith
ICIS3
2023 A Medical Chatbot for Tunisian Dialect using a Rule-Based and Machine Learning Approach
abstract
People nowadays have hectic schedules, so they tend to neglect their health because traveling to and from a hospital takes a significant amount of time. Many people would rather buy medications from a pharmacy than see a doctor. As a result, in terms of time, cost, and convenience, interacting with chatbots to obtain useful medical information could be a viable solution for them to overcome the aforementioned issues. In this paper, we propose a method to build a medical chatbot for the Arabic language and more specifically the Tunisian Dialect (TD). For that we collected, from patients, a set of 356 pairs of question/responses in TD. A variety of heterogeneous methods developed from standard machine learning were trained and a combination between Machine Learning and Rule-based was proposed for the implementation of our chatbot. We experimented various Machine Learning models and managed to achieve good results, scoring an F1-score of 98.60% with the Random Forest algorithm.
Sonda Rekik, Maryam Elamine, Lamia Hadrich Belguith
AICCSA3
2023 Approach Based on Bayesian Network and Ontology for Identifying Factors Impacting the States of People with Psychological Problems from Data on Social Media
Mourad Ellouze, Lamia Hadrich Belguith
MEDI2
2023 AlgBERT: Automatic Construction of Annotated Corpus for Sentiment Analysis in Algerian Dialect
abstract
Nowadays, sentiment analysis is one of the most crucial research fields of Natural Language Processing (NLP), and it is widely applied in a variety of applications such as marketing and politics. However, the Arabic language still lacks sufficient language resources to enable the tasks of opinion and emotion analysis comparing to other language such as English. Additionally, manual annotation requires a lot of effort and time. In this article, we address this problem and propose a novel automated annotation platform for sentiment analysis called AlgBERT by providing annotated corpus and using deep learning technology that includes many automatic natural language processing algorithms, which is the basis for text classification and opinion analysis. We suggest using BERT model as a method; it is the abbreviation of Bidirectional Encoder Representations from Transformers, as it is one of the most effective technologies in terms of results in different world languages. We used around of 54K comments collected from social networking (Twitter, YouTube) written in Arabic and Algerian dialects. Our AlgBERT system obtained excellent results with an accuracy of 91.04%, and this is considered as one of the best results for opinion analysis in Algerian dialect.
Khaoula Hamadouche, Kheira Zineb Bousmaha, Mohamed Abdelwaret Bekkoucha, Lamia Hadrich Belguith
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2023 Hybrid Pipeline for Building Arabic Tunisian Dialect-standard Arabic Neural Machine Translation Model from Scratch
abstract
Deep Learning is one of the most promising technologies compared to other methods in the context of machine translation. It has been proven to achieve impressive results on large amounts of parallel data for well-endowed languages. Nevertheless, for low-resource languages such as the Arabic Dialects, Deep Learning models failed due to the lack of available parallel corpora. In this article, we present a method to create a parallel corpus to build an effective NMT model able to translate into MSA, Tunisian Dialect texts present in social networks. For this, we propose a set of data augmentation methods aiming to increase the size of the state-of-the-art parallel corpus. By evaluating the impact of this step, we noticed that it has effectively boosted both the size and the quality of the corpus. Then, using the resulted corpus, we compare the effectiveness of CNN, RNN and transformers models to translate Tunisian Dialect into MSA. Experiments show that a better translation is achieved by the transformer model with a BLEU score of 60 vs., respectively, 33.36 and 53.98 with RNN and CNN models.
Saméh Kchaou, Rahma Boujelben, Lamia Hadrich Belguith
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2023 Tokenization of Tunisian Arabic: A Comparison between Three Machine Learning Models
abstract
Tokenization represents the way of segmenting a piece of text into smaller units called tokens. Since Arabic is an agglutinating language by nature, this treatment becomes a crucial preprocessing step for many Natural Language Processing (NLP) applications such as morphological analysis, parsing, machine translation, information extraction, and so on. In this article, we investigate word tokenization task with a rewriting process to rewrite the orthography of the stem. For this task, we are using Tunisian Arabic (TA) text. To the best of the researchers’ knowledge, this is the first study that uses TA for word tokenization. Therefore, we start by collecting and preparing various TA corpora from different sources. Then, we present a comparison of three character-based tokenizers based on Conditional Random Fields (CRF), Support Vector Machines (SVM) and Deep Neural Networks (DNN). The best proposed model using CRF achieved an F-measure result of 88.9%.
Asma Mekki, Inès Zribi, Mariem Ellouze, Lamia Hadrich Belguith
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2022 Bottom-up approach to translate Tunisian dialect texts in Social Networks
abstract
Dialect translation is a motivating task for both industrial and academic fields. Indeed, migrating to a standard language facilitates communication between people throughout the world, and facilitates application of Natural Language Processing tasks such as automatic opinion analysis, semantic analysis, etc. We describe, in this work, an effort to build a Neural Machine Translation (NMT) model in order to translate the comments posted on social media intended for the Tunisian community. It is a question of dealing with the Tunisian dialect (TD) in Social Networks (SN). Due to the orthographic ambiguity presented by the TD, we experiment different configurations corpora and NMT models following a bottom-up approach. The best configuration resulted in building a translation model which achieved a BLEU score of 69.22% on a test corpus.
Saméh Kchaou, Rahma Boujelben, Lamia Hadrich Belguith
AICCSA3
2022 Deep Learning CNN-LSTM Approach for Identifying Twitter Users Suffering from Paranoid Personality Disorder
Mourad Ellouze, Seifeddine Mechti, Lamia Hadrich Belguith
ICSOFT3
2022 An Improved Hybrid Method for Sentiment Analysis
abstract
In recent years, the sentiment analysis task has received great attention from communities. To fulfill this task, we proposed in this paper a hybrid method for the French language that combines the strengths of the lexical-based method and those of deep learning-based and machine learning-based methods. This method consists principally of two steps. In the first step, we annotate the learning dataset using an enhanced lexical-based method which is mainly based on a set of new rules. These rules improve significantly the polarity detection by taking into account the contextual information and a set of specific words. Then, we classify the reviews using machine learning (Naïve Bayes, Support Vector Machine) and deep learning (Long Short-Term Memory) algorithms. To realize our method, we collected a dataset composed of 7000 reviews in French from the online purchase site Amazon.fr. We obtained an overall rate of F-measure equal to 96,14%.
Sarsabene Hammi, Souha Hammami, Lamia Hadrich Belguith
INISTA3
2022 Aspect Term Extraction Improvement Based on a Hybrid Method
Sarsabene Hammi, Souha Hammami, Lamia Hadrich Belguith
ISMIS3
2022 Sarcasm Detection in Tunisian Social Media Comments: Case of COVID-19
Asma Mekki, Inès Zribi, Mariem Ellouze, Lamia Hadrich Belguith
ISMIS4
2022 Prediction and detection model for hierarchical Software-Defined Vehicular Network
abstract
Vehicle Ad-hoc Network (VANET) is the main component of the intelligent transportation system. With the development of the next-generation intelligent vehicular networks, the latter aims to provide strategic and secure services and communications in roads and smart cities. Due to VANET’s unique characteristics, such as high mobility of its nodes, self-organization, distributed network, and frequently changing topology, security, data integrity, and users’ privacy information are major concerns. Also, attack prevention is still an open issue. Distributed Denial of Service (DDoS) is one of the most dangerous attacks in VANETs, which aims to flood the system’s bandwidth. In this article, we propose a hierarchical architecture for securing Software-Defined Vehicular Network (SDVN) and a security model for predicting and detecting DDoS attacks based on behavioral analysis of nodes achieved by a Markov stochastic process. Simulation results show that our model effectively mitigates DDoS attacks with a high-reliability rate.
Houda Amari, Lyes Khoukhi, Lamia Hadrich Belguith
LCN3
2022 Standardisation of Dialect Comments in Social Networks in View of Sentiment Analysis : Case of Tunisian Dialect
abstract
With the growing access to the internet, the spoken Arabic dialect language becomes informal languages written in social media. Most users post comments using their own dialect. This linguistic situation inhibits mutual understanding between internet users and makes difficult to use computational approaches since most Arabic resources are intended for the formal language: Modern Standard Arabic (MSA). In this paper, we present a pipeline to standardize the written texts in social networks by translating them to the standard language MSA. We fine-tun at first an identification bert-based model to select Tunisian Dialect (TD) from MSA and other dialects. Then, we learned transformer model to translate TD to MSA. The final system includes the translated TD text and the originally text written in MSA. Each of these steps was evaluated on the same test corpus. In order to test the effectiveness of the approach, we compared two opinion analysis models, the first intended for the Sentiment Analysis (SA) of dialect texts and the second for the MSA texts. We concluded that through standardization we obtain the best score.
Saméh Kchaou, Rahma Boujelben, Emna Fsih, Lamia Hadrich Belguith
LREC4
2022 Annotating Verbal Multiword Expressions in Arabic: Assessing the Validity of a Multilingual Annotation Procedure
abstract
This paper describes our efforts to extend the PARSEME framework to Modern Standard Arabic. Theapplicability of the PARSEME guidelines was tested by measuring the inter-annotator agreement in theearly annotation stage. A subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADTwas selected and annotated by two Arabic native speakers independently. Following their annotations, anew Arabic corpus with over 1,250 annotated VMWEs has been built. This corpus already exceeds thesmallest corpora of the PARSEME suite, and enables first observations. We discuss our annotation guide-line schema that shows full MWE annotation is realizable in Arabic where we get good inter-annotator agreement.
Najet Hadj Mohamed, Chérifa Ben Khelil, Agata Savary, Iskandar Keskes, Jean-Yves Antoine, Lamia Hadrich Belguith
LREC6
2022 Appraisal of Two Arabic Opinion Summarization Methods: Statistical Versus Machine Learning
abstract
Abstract In this paper, we propose to overcome the challenge of digesting opinions in a news article. Our objective is to provide a summary of opinions delivered by many sources about a main topic in an Arabic news article. In literature, several studies addressed issues related to opinion summarization. However, we noticed a lack of studies that address this problem in Arabic language. So, we have proposed two different methods: multi-criteria and machine learning-based methods. We proceed by comparing the results provided by the proposed methods for opinionated sentence extraction. The proposed methods were evaluated using two feature types: text-based features and opinion-specific features. Experimental results show the robustness of machine learning method to extract opinionated sentences with consideration of two sets of features.
Imen Touati, Mariem Ellouze, Marwa Graja, Lamia Hadrich Belguith
Comput. J.4
2022 A decision system for computational authors profiling: From machine learning to deep learning
abstract
Summary In this study, we tackle the problem of author profiling. The aim of the proposed approach is to determine the author's age and gender. Once the user connects to the company website, this company collects the available data about him (which is usually very limited). Then, the user receives a service recommendation according to his gender and age. Thus, a context‐specific decision‐making system based on these limited data is required to produce an efficient classification. Such a decision system allows companies to promote their marketing. To obtain the best categorization, machine learning (ML) and deep learning (DL) techniques have been applied in the literature. In this article, we apply both classical ML techniques and recently developed DL techniques. More precisely, we adopt the gated recurrent unit model. Our experiments show that our findings are positively comparable with the best state‐of‐the‐art methods.
Seifeddine Mechti, Moez Krichen, Dhouha Ben Noureddine, Lamia Hadrich Belguith
Concurr. Comput. Pract. Exp.4
2021 Towards a Historical Ontology for Arabic Language: Investigation and Future Directions
Rim Laatar, Ahlem Rhayem, Chafik Aloulou, Lamia Hadrich Belguith
ISDA4
2021 Approach Based on Ontology and Machine Learning for Identifying Causes Affecting Personality Disorder Disease on Twitter
Mourad Ellouze, Seifeddine Mechti, Lamia Hadrich Belguith
KSEM3
2021 Securing Software-Defined Vehicular Network Architecture against DDoS attack
abstract
In the recent decades, Intelligent Transport Systems (ITS) attracted researchers’ great attention. ITS plays a very important role in making citizens’ lives easier in term of mobility, safety, quality of life and security. Vehicular ad-hoc networks (VANETs) became an inseparable component of ITS. The current architecture has been facing many issues due to VANET’s characteristics such as high mobility of its nodes and it is still vulnerable to important security attacks which threatens its main security services such as availability, data integrity, authentication and privacy. We propose a new VANET architecture called FCSDVN-ML, in which we combine three emerging paradigms: Software-Defined Network (SDN), Fog Computing (FC) and Machine Learning (ML) to improve security in VANETs. In this paper, we described our architecture components and we discussed its potential performance against Distributed Denial of Service (DDoS) attack using the hierarchical firewalls.
Houda Amari, Wassef Louati, Lyes Khoukhi, Lamia Hadrich Belguith
LCN4
2020 Treebank Creation and Parser Generation for Tunisian Social Media Text
abstract
Tunisian Arabic (TA) is a morphologically and syntactically rich dialect, which presents an interesting challenge for Natural Language Processing (NLP) tasks such as part-of-speech tagging, parsing, semantic analysis, etc. It is classified as a low-resourced language. Tunisians use it in daily life communication, social media exchanges, etc. In this paper, we focus on Tunisian Arabic linguistic resources and tools creation. We present the creation and generation of Tunisian treebank and parser for social media texts. We use an existing state-of-the-art parser to build this treebank. Then, we investigate the effects of different data sizes and different combinations of Tunisian dialect forms in automatic parsing.
Asma Mekki, Inès Zribi, Mariem Ellouze, Lamia Hadrich Belguith
AICCSA4
2020 TCP Incast Solutions in Data Center Networks: Survey
Houda Amari, Wassef Louati, Lyes Khoukhi, Lamia Hadrich Belguith
HIS4
2020 Automatic profile recognition of authors on social media based on hybrid approach
abstract
In this paper, we propose a method to discover the profile of a user on social media by detecting their gender, age, and personality traits. The stated method takes advantage of NLP techniques, lexico-semantic resources and statistical measures. These techniques have allowed us to have a more general method which isn't limited to the words founded in the training corpus and preserve the principle of relativity by taking advantage of the fuzzy logic. The previously mentioned method was implemented and evaluated using the PAN 2015 corpus (Author profiling).We obtained as a result, an F-measure value which equals to 81%. The found results are indeed interesting when compared to other works.
Mourad Ellouze, Seifeddine Mechti, Lamia Hadrich Belguith
KES3
2020 Toward Qualitative Evaluation of Embeddings for Arabic Sentiment Analysis
abstract
In this paper, we propose several protocols to evaluate specific embeddings for Arabic sentiment analysis (SA) task. In fact, Arabic language is characterized by its agglutination and morphological richness contributing to great sparsity that could affect embedding quality. This work presents a study that compares embeddings based on words and lemmas in SA frame. We propose first to study the evolution of embedding models trained with different types of corpora (polar and non polar) and explore the variation between embeddings by observing the sentiment stability of neighbors in embedding spaces. Then, we evaluate embeddings with a neural architecture based on convolutional neural network (CNN). We make available our pre-trained embeddings to Arabic NLP research community with free to use. We provide also for free resources used to evaluate our embeddings. Experiments are done on the Large Arabic-Book Reviews (LABR) corpus in binary (positive/negative) classification frame. Our best result reaches 91.9%, that is higher than the best previous published one (91.5%).
Amira Barhoumi, Nathalie Camelin, Chafik Aloulou, Yannick Estève, Lamia Hadrich Belguith
LREC5
2020 Disambiguating Arabic Words According to Their Historical Appearance in the Document Based on Recurrent Neural Networks
abstract
How can we determine the semantic meaning of a word in relation to its context of appearance? We eventually have to grabble with this difficult question, as one of the paramount problems of Natural Language Processing (NLP). In other words, this issue is commonly defined as Word Sense Disambiguation (WSD). The latter is one of the crucial difficulties within the NLP field. In this respect, word vectors extracted from a neural network model have been successfully applied for resolving the WSD problem. Accordingly, this article presents an unprecedented method to disambiguate Arabic words according to both their contextual appearance in a source text and the era in which they emerged. In fact, in the few previous decades, many researchers have been grabbling with Arabic Word Sense Disambiguation. It should be noted that the Arabic language can be divided into three major historical periods: old Arabic, middle-age Arabic, and contemporary Arabic. Actually, contemporary Arabic has proved to be the greatest concern of many researchers. The main gist of our work is to disambiguate Arabic words according to the historical period in which they appeared. To perform such a task, we suggest a method that deploys contextualized word embeddings to better gather valid syntactic and semantic information of the same word by taking into account its contextual uses. The preponderant thing is to convert both the senses and the contextual uses of an ambiguous item to vectors, then determine which of the possible conceptual meanings of the target word is closer to the given context.
Rim Laatar, Chafik Aloulou, Lamia Hadrich Belguith
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2020 Transliteration of Arabizi into Arabic Script for Tunisian Dialect
abstract
The evolution of information and communication technology has markedly influenced communication between correspondents. This evolution has facilitated the transmission of information and has engendered new forms of written communication (email, chat, SMS, comments, etc.). Most of these messages and comments are written in Latin script, also called Arabizi . Moreover, the language used in social media and SMS messaging is characterized by the use of informal and non-standard vocabulary, such as repeated letters for emphasis, typos, non-standard abbreviations, and nonlinguistic content like emoticons. Since the Tunisian dialect suffers from the unavailability of basic tools and linguistic resources compared to Modern Standard Arabic, we resort to the use of these written sources as a starting point to build large corpora automatically. In the context of natural language processing and to benefit from these networks’ data, transliterating from Arabizi to Arabic script is a necessary step because most recently available tools for processing the Tunisian dialect expect Arabic script input. Indeed, the transliteration task can help construct and enrich parallel corpora and dictionaries for the Tunisian dialect and can be useful for developing various natural language processing applications such as sentiment analysis, opinion mining, topic detection, and machine translation. In this article, we focus on converting the Tunisian dialect text that is written in Latin script to Arabic script following the Conventional Orthography for Dialectal Arabic. Then, we propose two models to transliterate Arabizi into Arabic script for the Tunisian dialect, namely a rule-based model and a discriminative model as a sequence classification task based on conditional random fields). In the first model, we use a set of transliteration rules to convert the Tunisian dialect Arabizi texts to Arabic script. In the second model, transliteration is performed both at word and character levels. In the end, our models got a character error rate of 10.47%.
Abir Masmoudi 0001, Mariem Ellouze, Mourad Khrouf, Lamia Hadrich Belguith
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2019 An Unsupervised Method for Detecting Style Breaches in a Document
abstract
In this paper, we propose an unsupervised method for identifying style breaches in a given document, also known as intrinsic plagiarism detection. In fact, plagiarism is one of the major challenges in various domains. Supervised learning techniques fail to capture the stylistic changes in a text effectively and this is mainly due to the static segmentation of the text. For this reason, we present in this paper our proposed method for intrinsic plagiarism detection, we experimented with the unsupervised algorithm Kmeans and the similarity measure Cosine. Our results are comparable to best systems presented in PAN@CLEF competitive conference.
Maryam Elamine, Seifeddine Mechti, Lamia Hadrich Belguith
AICCSA3
2019 Extrinsic Plagiarism Detection for French Language with Word Embeddings
Maryam Elamine, Fethi Bougares, Seifeddine Mechti, Lamia Hadrich Belguith
ISDA4
2019 Automatic Diacritics Restoration for Tunisian Dialect
abstract
Modern Standard Arabic, as well as Arabic dialect languages, are usually written without diacritics. The absence of these marks constitute a real problem in the automatic processing of these data by NLP tools. Indeed, writing Arabic without diacritics introduces several types of ambiguity. First, a word without diacratics could have many possible meanings depending on their diacritization. Second, undiacritized surface forms of an Arabic word might have as many as 200 readings depending on the complexity of its morphology [12]. In fact, the agglutination property of Arabic might produce a problem that can only be resolved using diacritics. Third, without diacritics a word could have many possible parts of speech (POS) instead of one. This is the case with the words that have the same spelling and POS tag but a different lexical sense, or words that have the same spelling but different POS tags and lexical senses [8]. Finally, there is ambiguity at the grammatical level (syntactic ambiguity). In this article, we propose the first work that investigates the automatic diacritization of Tunisian Dialect texts. We first describe our annotation guidelines and procedure. Then, we propose two major models, namely a statistical machine translation (SMT) and a discriminative model as a sequence classification task based on Conditional Random Fields (CRF). In the second approach, we integrate POS features to influence the generation of diacritics. Diacritics restoration was performed at both the word and the character levels. The results showed high scores of automatic diacritization based on the CRF system (Word Error Rate (WER) 21.44% for CRF and WER 34.6% for SMT).
Abir Masmoudi 0001, Salima Mdhaffar, Rahma Sellami, Lamia Hadrich Belguith
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2018 Word Sense Disambiguation using Skip Gram Model to Create a Historical Dictionary for Arabic
abstract
The evolution of the Arabic language from antiquity to the present days has given birth to several linguistic registers ascribed to the great periods of the history of the Arabic language. They can be classified as: Old Arabic, Classical Arabic and Modern Standard Arabic. In this work, we propose a method that aims to disambiguate words in Modern Standard Arabic. This method consists of measuring the semantic relation between the context of use of the ambiguous word and its sense definitions. Within the context of creating a historical dictionary for Arabic, and to disambiguate a word, we need to take into consideration the historical period in which the word appeared. This method disambiguates Arabic words takes into account that a word may have an old meaning but appears in a modern document.
Rim Laatar, Chafik Aloulou, Lamia Hadrich Belguith
AICCSA3
2018 Improving Native Language Identification Model with Syntactic Features: Case of Arabic
Seifeddine Mechti, Nabil Khoufi, Lamia Hadrich Belguith
ISDA (2)3
2018 Intelligent Analysis in Question Answering System Based on an Arabic Temporal Resource
Mayssa Mtibaa, Zeineb Neji, Mariem Ellouze, Lamia Hadrich Belguith
ISDA (2)4
2018 Word2vec for Arabic Word Sense Disambiguation
Rim Laatar, Chafik Aloulou, Lamia Hadrich Belguith
NLDB3
2017 Hybrid Method for Multilingual Automatic Grouping of Writing Styles
abstract
In this paper, we tackle the task of automatic grouping of writing styles, also called author clustering. This task deals with identifying authorship links and single-authored groups of documents [3]. This task is very important for other domains like plagiarism detection and author diarization. In this paper, we present a hybrid method that combines lexical and syntactic features for multilingual automatic grouping of writing style. The proposed method has been evaluated on two publicly available corpora. The obtained results outperform the ones obtained by the best state-of-the-art methods.
Seifeddine Mechti, Maryam Elamine, Lamia Hadrich Belguith, Rim Faiz
AICCSA3
2017 The role of temporal inferences in understanding Arabic text
abstract
Inference approaches in Arabic question answering systems are in their first steps if we compare them with other languages. Evidently, any user is interested in obtaining a specific and precise answer to a specific question. Therefore, the challenge of developing a system capable of obtaining a relevant and concise answer is obviously of great benefit. This paper deals with answering questions about temporal information involving several forms of inference.
Hajer Omri, Zeineb Neji, Mariem Ellouze, Lamia Hadrich Belguith
KES4
2016 HAD, a platform to create a historical dictionary
abstract
This paper will help the linguists find an easy method to facilitate the creation of a standard Arabic historical dictionary in order to save time and to be up to date with the other languages. In this method, we propose a platform of Automatic Natural Language Processing (ANLP) tools which permits the automatic indexing and research from Arabic texts corpus. Some pretreatments are done before the indexation process: segmentation, normalization, and filtering, morphological analysis. The prototype that we've developed for the generation of standard Arabic historical dictionary permits to extract contexts from the entered corpus and to assign meaning from the user. The evaluation of our system shows that the results are reliable.
Faten Khalfallah, Chafik Aloulou, Lamia Hadrich Belguith
AICCSA3
2016 An empirical method using features combination for Arabic native language identification
abstract
In this paper, we focus on the detection of the Arabic learners' mother tongue. The proposed method is based on the automatic classification using some data statistically extracted from a source corpus. We present a hybrid method that combines surface analysis in texts with an automatic learning method. Unlike the few techniques found in the state of the art, the features selection phase allowed improving performances. Therefore, the obtained results outperformed those provided by the best methods used for Arabic native language detection.
Seifeddine Mechti, Ayoub Abbassi, Lamia Hadrich Belguith, Rim Faiz
AICCSA3
2016 Arabic Pronominal Anaphora Resolution Based on New Set of Features
Souha Hammami, Lamia Hadrich Belguith
CICLing (1)2
2016 A Framework for Language Resource Construction and Syntactic Analysis: Case of Arabic
Nabil Khoufi, Chafik Aloulou, Lamia Hadrich Belguith
CICLing (1)3
2016 Adaptation of a Term Extractor to Arabic Specialised Texts: First Experiments and Limits
Wafa Neifar, Thierry Hamon, Pierre Zweigenbaum, Mariem Ellouze, Lamia Hadrich Belguith
CICLing (1)5
2016 Pattern Based on Temporal Inference
Zeineb Neji, Mariem Ellouze, Lamia Hadrich Belguith
ICANN (1)3
2016 IQAS: Inference question answering system for handling temporal inference
abstract
One of the most crucial problems in any Natural Language Processing (NLP) task is the representation of time. This includes applications such as Information Retrieval techniques (IR), Information Extraction (IE) and Question/answering systems (QA). This paper deals with temporal information involving several forms of inference in Arabic language.
Zeineb Neji, Mariem Ellouze, Lamia Hadrich Belguith
INISTA3
2016 Conditional Random Fields for the Tunisian Dialect Grapheme-to-Phoneme Conversion
Abir Masmoudi 0001, Mariem Ellouze, Fethi Bougares, Yannick Estève, Lamia Hadrich Belguith
INTERSPEECH5
2016 A Platform for the Conceptualization of Arabic Texts Dedicated to the Design of the UML Class Diagram
Kheira Zineb Bousmaha, Mustapha Kamel Rahmouni, Belkacem Kouninef, Lamia Hadrich Belguith
NLDB4
2016 Automatic Evaluation of a Summary's Linguistic Quality
Samira Ellouze, Maher Jaoua, Lamia Hadrich Belguith
NLDB3
2016 A Platform Based ANLP Tools for the Construction of an Arabic Historical Dictionary
Faten Khalfallah, Handi Msadak, Chafik Aloulou, Lamia Hadrich Belguith
NLDB4
2016 PIRAT: A Personalized Information Retrieval System in Arabic Texts Based on a Hybrid Representation of a User Profile
Houssem Safi, Maher Jaoua, Lamia Hadrich Belguith
NLDB3
2016 Toward hybrid method for parsing Modern Standard Arabic
abstract
Parsing Arabic language is a difficult task given the specificities of the language and given the scarcity of linguistic resources. Linguistic resources such as grammars are very important to any natural language processing application. Unfortunately, the manual construction of these resources is laborious and time-consuming. The use of annotated corpora as a knowledge database might be a solution to a fast construction of a grammar for a given language. In this paper, we began by presenting an overview of our method to automatically induce a probabilistic context free grammar from an Arabic annotated corpus (The Penn Arabic TreeBank). Then we tested the obtained grammar in the parsing task and we expose the evaluation results. Finally we present our vision of a hybrid method for parsing Modern Standard Arabic (MSA) that we believe that it could enhance obtained results.
Nabil Khoufi, Chafik Aloulou, Lamia Hadrich Belguith
SNPD3
2015 Automatic Dialogue Act Annotation within Arabic Debates
Samira Ben Dbabis, Hatem Ghorbel, Lamia Hadrich Belguith, Mohamed Kallel
CICLing (1)3
2015 Arabic Transliteration of Romanized Tunisian Dialect Text: A Preliminary Investigation
Abir Masmoudi 0001, Nizar Habash, Mariem Ellouze, Yannick Estève, Lamia Hadrich Belguith
CICLing (1)5
2015 Sentiment Classification of Arabic Documents: Experiments with multi-type features and ensemble algorithms
Amine Bayoudhi, Hatem Ghorbel, Lamia Hadrich Belguith
PACLIC3
2015 Statistical Framework with Knowledge Base Integration for Robust Speech Understanding of the Tunisian Dialect
abstract
In this paper, we propose a hybrid method for the spoken Tunisian dialect understanding within a limited task. This method couples a discriminative statistical method with a domain ontology. The statistical method is based on conditional random field (CRF) models learned from a little size corpus to perform conceptual labeling task. These models are able to detect the semantic dependency between words. However, the domain ontology is used to add prior knowledge about the task. Our experiments are based on a real spoken Tunisian dialect corpus. The obtained results show that the proposed method is able to improve the performance of CRF models for speech understanding by the integration of the domain ontology. Our method can be exploited for under-resourced languages and Arabic dialects to overcome the lack of linguistic resources .
Marwa Graja, Maher Jaoua, Lamia Hadrich Belguith
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Question focus extraction and answer passage retrieval
abstract
Question Analysis is an important task in Question Answering Systems (QAS). It consists generally in identifying the semantic type of the question and extracting the main focus of the question. The goal is to better specify the required information by the question. In this context and as part of a framework aiming to implement an Arabic opinion QAS for political debates, this paper addresses the problem of defining the focus of opinion questions and proposes particularly an approach for extracting the focus of attitude questions. The proposed approach is based on semi-automatically constructed lexico-syntactic patterns. Furthermore, the paper presents an adapted Vector Space Model (VSM) based method to retrieve candidate answer passages from a transcribed TV political show. Several experiments were carried out and showed that the focus extraction approach has achieved over 72% as F1 score for holder and target extraction, and has improved the baseline passage retrieval task by over than 25%.
Amine Bayoudhi, Lamia Hadrich Belguith, Hatem Ghorbel
AICCSA2
2014 Chunking Arabic texts using Conditional Random Fields
abstract
Chunking or shallow syntactic parsing is proving to be a task of interest to many natural language processing applications. The problem gets worse for the Arabic language because of its specific features that make it quite different and even more ambiguous than other natural languages when processed. In this paper, we present a method for chunking Arabic texts based on supervised learning. We use the Conditional Random Fields algorithm and the Penn Arabic Treebank to train the model. For the experimentation, we use over than 10,100 sentences as training data and 2,524 sentences for the test. The evaluation of the method consists of the calculation of the generated model accuracy and the results are very encouraging.
Nabil Khoufi, Chafik Aloulou, Lamia Hadrich Belguith
AICCSA3
2014 A Corpus and Phonetic Dictionary for Tunisian Arabic Speech Recognition
Abir Masmoudi 0001, Mariem Ellouze, Yannick Estève, Lamia Hadrich Belguith, Nizar Habash
LREC4
2014 A Conventional Orthography for Tunisian Arabic
Inès Zribi, Rahma Boujelben, Abir Masmoudi 0001, Mariem Ellouze, Lamia Hadrich Belguith, Nizar Habash
LREC5
2014 Focus Definition and Extraction of Opinion Attitude Questions
Amine Bayoudhi, Hatem Ghorbel, Lamia Hadrich Belguith
NLDB3
2014 Fine-Grained POS Tagging of Spoken Tunisian Dialect Corpora
Rahma Boujelben, Mariem Mallek, Mariem Ellouze, Lamia Hadrich Belguith
NLDB4
2014 Splitting Arabic Texts into Elementary Discourse Units
abstract
In this article, we propose the first work that investigates the feasibility of Arabic discourse segmentation into elementary discourse units within the segmented discourse representation theory framework. We first describe our annotation scheme that defines a set of principles to guide the segmentation process. Two corpora have been annotated according to this scheme: elementary school textbooks and newspaper documents extracted from the syntactically annotated Arabic Treebank. Then, we propose a multiclass supervised learning approach that predicts nested units. Our approach uses a combination of punctuation, morphological, lexical, and shallow syntactic features. We investigate how each feature contributes to the learning process. We show that an extensive morphological analysis is crucial to achieve good results in both corpora. In addition, we show that adding chunks does not boost the performance of our system.
Iskandar Keskes, Farah Benamara, Lamia Hadrich Belguith
ACM Trans. Asian Lang. Inf. Process.3
2013 Orthographic Transcription for Spoken Tunisian Arabic
Inès Zribi, Marwa Graja, Mariem Ellouze, Maher Jaoua, Lamia Hadrich Belguith
CICLing (1)5
2013 Question Answering System for Dialogues: A New Taxonomy of Opinion Questions
Amine Bayoudhi, Hatem Ghorbel, Lamia Hadrich Belguith
FQAS3
2013 Mapping Rules for Building a Tunisian Dialect Lexicon and Generating Corpora
Rahma Boujelben, Mariem Ellouze, Lamia Hadrich Belguith
IJCNLP3
2013 Morphological Analysis of Tunisian Dialect
Inès Zribi, Mariem Ellouze, Lamia Hadrich Belguith
IJCNLP3
2012 Clause-based Discourse Segmentation of Arabic Texts
Iskandar Keskes, Farah Benamara, Lamia Hadrich Belguith
LREC3
2011 Towards Understanding Spoken Tunisian Dialect
Marwa Graja, Maher Jaoua, Lamia Hadrich Belguith
ICONIP (3)3
2010 An Automatic Definition Extraction in Arabic Language
Omar Trigui, Lamia Hadrich Belguith, Paolo Rosso
NLDB2
2002 MASPAR: A multi-agent system for parsing arabic
abstract
This paper deals with representation of syntactic information based on a formal work setting of the unification grammars. It shows the interest of using this grammar type (namely HPSG grammars) notably for Arabic parsing, and focuses on its role in order to obtain a robust parsing.
Chafik Aloulou, Lamia Hadrich Belguith, Abdelmajid Ben Hamadou
SMC (2)2
1998 Multilingual Robust Anaphora Resolution
Ruslan Mitkov, Lamia Hadrich Belguith, Malgorzata Stys
EMNLP2