Caroline Sabty

dblp:139/3278 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-3590-5737ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Prompting vs Ensemble Architectures for Arabic English Code-Switched Classification
Donia Ali, Salma Haytham, Sandra George, Caroline Sabty
DATA (1)4
2026 SwitchEmbed: Representation Learning for Arabic-English Code-Switched Text
Mariam Rizkallah, Amani Ghonim, Ahmed Sherif, Caroline Sabty
DATA (1)4
2026 Attention-Pruned SHAP: Accelerating SHAP-Based Explainability with Attention-Guided Feature Pruning
Hamza Shafik, Ahmed Sherif, Caroline Sabty
NLDB3
2026 Adapting GPT for Egyptian Arabic-English Code-Switched Sentiment Analysis Through Prompting, Retrieval, and Sentiment-Guided Fine-Tuning
Ahmed Sherif, Reem Hussein, Caroline Sabty
NLDB3
2025 A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions
abstract
Language in the Arab world presents a complex diglossic and multilingual setting, involving the use of Modern Standard Arabic, various dialects and sub-dialects, as well as multiple European languages. This diverse linguistic landscape has given rise to code-switching, both within Arabic varieties and between Arabic and foreign languages. The widespread occurrence of code-switching across the region makes it vital to address these linguistic needs when developing language technologies. In this paper, we provide a review of the current literature in the field of code-switched Arabic NLP, offering a broad perspective on ongoing efforts, challenges, research gaps, and recommendations for future research directions.
Injy Hamed, Caroline Sabty, Slim Abdennadher, Ngoc Thang Vu, Thamar Solorio, Nizar Habash
COLING2
2025 Explainable AI for NLP: Enhancing Transparency in Sentiment Analysis and Named Entity Recognition
Feras Elkharrat, Mohamed Ghoniem, Ahmed Sherif, Caroline Sabty
NLDB (2)4
2024 Bridging the Gap: Developing an Automatic Speech Recognition System for Egyptian Dialect Integration Into Chatbots
Mazen Nabil, Aya Abdalla, Nada Sharaf, Caroline Sabty
NLDB (2)4
2024 Enhanced Cognitive Distortions Detection and Classification Through Data Augmentation Techniques
Mohamad Rasmy, Caroline Sabty, Nourhan Sakr, Alia El Bolock
PRICAI (1)2
2024 Visualizing the Unseen: Arabic Image-to-Story Generation Using Deep Learning Techniques
Eman Saleh, Caroline Sabty
PRICAI (2)2
2023 Fact-in-a-Box: Hiding Educational Facts in Short Stories for Implicit Learning
Alia El Bolock, Caroline Sabty, Nour Eldin Awad, Slim Abdennadher
CSEDU (1)2
2023 IntrusionHunter: Detection of Cyber Threats in Big Data
Hashem Mohamed, Alia El Bolock, Caroline Sabty
DATA3
2022 Enhancing Deep Learning with Embedded Features for Arabic Named Entity Recognition
abstract
The introduction of word embedding models has remarkably changed many Natural Language Processing tasks. Word embeddings can automatically capture the semantics of words and other hidden features. Nonetheless, the Arabic language is highly complex, which results in the loss of important information. This paper uses Madamira, an external knowledge source, to generate additional word features. We evaluate the utility of adding these features to conventional word and character embeddings to perform the Named Entity Recognition (NER) task on Modern Standard Arabic (MSA). Our NER model is implemented using Bidirectional Long Short Term Memory and Conditional Random Fields (BiLSTM-CRF). We add morphological and syntactical features to different word embeddings to train the model. The added features improve the performance by different values depending on the used embedding model. The best performance is achieved by using Bert embeddings. Moreover, our best model outperforms the previous systems to the best of our knowledge.
Ali L. Hatab, Caroline Sabty, Slim Abdennadher
LREC2
2019 Automatic Infogram Generation for Online Journalism
abstract
Infographics is a tool for data visualization. It makes data easy to understand and interpret. An infographic is defined as a visual representation for data like a chart or a diagram. Infographics can help in many fields including education. This is due to the fact that information can be easily memorized and understood if was given in a visual form. Infographics are found everywhere and used in many different fields. Facebook timeline is considered as an infographic. On another hand, online journalism is increasingly gaining popularity. It is also considered as a source of big data that is rapidly expanding. Online newspapers and magazines provide a large population with daily important information. The aim of the work is to use data visualization, infographics and Natural Language Processing (NLP) techniques in online journalism. The aim of the work is to automatically visualize the information in an article in the form of infographics.
Farah Khouzam, Nada Sharaf, Madeleine Saad, Caroline Sabty, Slim Abdennadher
IV (1)4
2018 Arabic Named Entity Recognition Using Clustered Word Embedding
Caroline Sabty, Mohamed Elmahdy 0001, Slim Abdennadher
CICLing (1)1