Reem Suwaileh

dblp:178/5998 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0001-7341-1407ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 ThatiAR: Subjectivity Detection in Arabic News Sentences
abstract
In this study, we present the first large dataset, ThatiAR, for subjectivity detection in Arabic, consisting of ~3.6K manually annotated sentences, and GPT-4o based explanations. In addition, we include instructions (both in English and Arabic) to facilitate LLM based fine-tuning. We provide an in-depth analysis of the dataset, annotation process, and extensive benchmark results, including PLMs and LLMs. Our analysis of the annotation process highlights that annotators were strongly influenced by their political, cultural, and religious backgrounds, especially at the beginning of the annotation process. The experimental results suggest that LLMs with in-context learning provide better performance. We release the dataset and resources to the community.
Reem Suwaileh, Maram Hasanain, Fatema Hubail, Wajdi Zaghouani, Firoj Alam
ICWSM1
2024 The CLEF-2024 CheckThat! Lab: Check-Worthiness, Subjectivity, Persuasion, Roles, Authorities, and Adversarial Robustness
Alberto Barrón-Cedeño, Firoj Alam, Tanmoy Chakraborty 0002, Tamer Elsayed, Preslav Nakov, Piotr Przybyla, Julia Maria Struß, Fatima Haouari, Maram Hasanain, Federico Ruggeri, Xingyi Song, Reem Suwaileh
ECIR (5)12
2023 IDRISI-RA: The First Arabic Location Mention Recognition Dataset of Disaster Tweets
abstract
Extracting geolocation information from social media data enables effective disaster management, as it helps response authorities; for example, in locating incidents for planning rescue activities, and affected people for evacuation.Nevertheless, geolocation extraction is greatly understudied for the low resource languages such as Arabic.To fill this gap, we introduce IDRISI-RA, the first publicly-available Arabic Location Mention Recognition (LMR) dataset that provides human-and automatically-labeled versions in order of thousands and millions of tweets, respectively.It contains both location mentions and their types (e.g., district, city).Our extensive analysis shows the decent geographical, domain, location granularity, temporal, and dialectical coverage of IDRISI-RA.Furthermore, we establish baselines using the standard Arabic NER models and build two simple, yet effective, LMR models.Our rigorous experiments confirm the need for developing specific models for Arabic LMR in the disaster domain.Moreover, experiments show the promising domain and geographical generalizability of IDRISI-RA under zero-shot learning.
Reem Suwaileh, Muhammad Imran 0002, Tamer Elsayed
ACL (1)1
2023 IDRISI-RE: A generalizable dataset with benchmarks for location mention recognition on disaster tweets
abstract
While utilizing Twitter data for crisis management is of interest to different response authorities, a critical challenge that hinders the utilization of such data is the scarcity of automated tools that extract geolocation information. The limited focus on Location Mention Recognition (LMR) in tweets, specifically, is attributed to the lack of a standard dataset that enables research in LMR. To bridge this gap, we present IDRISI-RE, a large-scale human-labeled LMR dataset comprising around 20.5k tweets. The annotated location mentions within the tweets are also assigned location types (e.g., country, city, street, etc.). IDRISI-RE contains tweets from 19 disaster events of diverse types (e.g., flood and earthquake) covering a wide geographical area of 22 English-speaking countries. Additionally, IDRISI-RE contains about 56.6k automatically-labeled tweets that we offer as a silver dataset. To highlight the superiority of IDRISI-RE over past efforts, we present rigorous analyses on reliability, consistency, coverage, diversity, and generalizability. Furthermore, we benchmark IDRISI-RE using a representative set of LMR models to provide the community with baselines for future work. Our extensive empirical analysis shows the promising generalizability of IDRISI-RE compared to existing datasets. We show that models trained on IDRISI-RE better tackle domain shifts and are less susceptible to change in geographical areas.
Reem Suwaileh, Tamer Elsayed, Muhammad Imran 0002
Inf. Process. Manag.1
2020 Are We Ready for this Disaster? Towards Location Mention Recognition from Crisis Tweets
abstract
The widespread usage of Twitter during emergencies has provided a new opportunity and timely resource to crisis responders for various disaster management tasks.Geolocation information of pertinent tweets is crucial for gaining situational awareness and delivering aid.However, the majority of tweets do not come with geoinformation.In this work, we focus on the task of location mention recognition from crisis-related tweets.Specifically, we investigate the influence of different types of labeled training data on the performance of a BERT-based classification model.We explore several training settings such as combing in-and out-domain data from news articles and general-purpose and crisis-related tweets.Furthermore, we investigate the effect of geospatial proximity while training on near or far-away events from the target event.Using five different datasets, our extensive experiments provide answers to several critical research questions that are useful for the research community to foster research in this important direction.For example, results show that, for training a location mention recognition model, Twitter-based data is preferred over general-purpose data; and crisis-related data is preferred over generalpurpose Twitter data.Furthermore, training on data from geographically-nearby disaster events to the target event boosts the performance compared to training on distant events.
Reem Suwaileh, Muhammad Imran 0002, Tamer Elsayed, Hassan Sajjad 0001
COLING1
2020 CheckThat! at CLEF 2020: Enabling the Automatic Identification and Verification of Claims in Social Media
Alberto Barrón-Cedeño, Tamer Elsayed, Preslav Nakov, Giovanni Da San Martino, Maram Hasanain, Reem Suwaileh, Fatima Haouari
ECIR (2)6
2020 Time-Critical Geolocation for Social Good
Reem Suwaileh
ECIR (2)1
2020 ArTest: The First Test Collection for Arabic Web Search with Relevance Rationales
abstract
The scarcity of Arabic test collections has long hindered information retrieval (IR) research over the Arabic Web. In this work, we present ArTest, the first large-scale test collection designed for the evaluation of ad-hoc search over the Arabic Web. ArTest uses ArabicWeb16, a collection of around 150M Arabic Web pages as the document collection, and includes 50 topics, 10,529 relevance judgments, and (more importantly) a rationale behind each judgment. To our knowledge, this is also the first IR test collection that includes rationales of primary assessors (i.e., topic developers) for their relevance judgments, exhibiting a useful resource for understanding the relevance phenomena. Finally, ArTest is made publicly-available for the research community.
Maram Hasanain, Yassmine Barkallah, Reem Suwaileh, Mucahid Kutlu, Tamer Elsayed
SIGIR3
2019 CheckThat! at CLEF 2019: Automatic Identification and Verification of Claims
Tamer Elsayed, Preslav Nakov, Alberto Barrón-Cedeño, Maram Hasanain, Reem Suwaileh, Giovanni Da San Martino, Pepa Atanasova
ECIR (2)5
2018 DART: A Large Dataset of Dialectal Arabic Tweets
Israa Alsarsour, Esraa Mohamed, Reem Suwaileh, Tamer Elsayed
LREC3
2018 EveTAR: building a large-scale multi-task test collection over Arabic tweets
Maram Hasanain, Reem Suwaileh, Tamer Elsayed, Mucahid Kutlu, Hind A. Al-Merekhi
Inf. Retr. J.2
2016 ArabicWeb16: A New Crawl for Today's Arabic Web
abstract
Web crawls provide valuable snapshots of the Web which enable a wide variety of research, be it distributional analysis to characterize Web properties or use of language, content analysis in social science, or Information Retrieval (IR) research to develop and evaluate effective search algorithms. While many English-centric Web crawls exist, existing public Arabic Web crawls are quite limited, limiting research and development. To remedy this, we present ArabicWeb16, a new public Web crawl of roughly 150M Arabic Web pages with significant coverage of dialectal Arabic as well as Modern Standard Arabic. For IR researchers, we expect ArabicWeb16 to support various research areas: ad-hoc search, question answering, filtering, cross-dialect search, dialect detection, entity search, blog search, and spam detection. Combined use with a separate Arabic Twitter dataset we are also collecting may provide further value.
Reem Suwaileh, Mucahid Kutlu, Nihal Fathima, Tamer Elsayed, Matthew Lease
SIGIR1