VLDB 2026 Research / reviewers in the wild / expert
Sabit Hassan
dblp:234/1594
· DBLP profile ↗
12ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0001-7518-966XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Active Learning Framework for Inclusive Generation by Large Language ModelsabstractEnsuring that Large Language Models (LLMs) generate text representative of diverse sub-populations is essential, particularly when key concepts related to under-represented groups are scarce in the training data. We address this challenge with a novel clustering-based active learning framework, enhanced with knowledge distillation. The proposed framework transforms the intermediate outputs of the learner model, enabling effective active learning for generative tasks for the first time. Integration of clustering and knowledge distillation yields more representative models without prior knowledge of underlying data distribution and overbearing human efforts. We validate our approach in practice through case studies in counter-narration and style transfer. We construct two new datasets in tandem with model training, showing a performance improvement of 2%–10% over baseline models. Our results also show more consistent performance across various data subgroups and increased lexical diversity, underscoring our model’s resilience to skewness in available data. Further, our results show that the data acquired via our approach improves the performance of secondary models not involved in the learning loop, showcasing practical utility of the framework. Sabit Hassan, Anthony Sicilia, Malihe Alikhani |
COLING | 1 |
| 2025 | Coherence-Driven Multimodal Safety Dialogue with Active Learning for Embodied Agents
Sabit Hassan, Hye-Young Chung, Xiang Zhi Tan, Malihe Alikhani |
AAMAS | 1 |
| 2023 | Multilingual Content Moderation: A Case Study on RedditabstractMeng Ye, Karan Sikka, Katherine Atwell, Sabit Hassan, Ajay Divakaran, Malihe Alikhani. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Meng Ye 0002, Karan Sikka, Katherine Atwell, Sabit Hassan, Ajay Divakaran, Malihe Alikhani |
EACL | 4 |
| 2023 | DisCGen: A Framework for Discourse-Informed Counterspeech GenerationabstractSabit Hassan, Malihe Alikhani. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Sabit Hassan, Malihe Alikhani |
IJCNLP (1) | 1 |
| 2023 | Emojis as anchors to detect Arabic offensive language and hate speechabstractAbstract We introduce a generic, language-independent method to collect a large percentage of offensive and hate tweets regardless of their topics or genres. We harness the extralinguistic information embedded in the emojis to collect a large number of offensive tweets. We apply the proposed method on Arabic tweets and compare it with English tweets—analyzing key cultural differences. We observed a constant usage of these emojis to represent offensiveness throughout different timespans on Twitter. We manually annotate and publicly release the largest Arabic dataset for offensive, fine-grained hate speech, vulgar, and violence content. Furthermore, we benchmark the dataset for detecting offensiveness and hate speech using different transformer architectures and perform in-depth linguistic analysis. We evaluate our models on external datasets—a Twitter dataset collected using a completely different method, and a multi-platform dataset containing comments from Twitter, YouTube, and Facebook, for assessing generalization capability. Competitive results on these datasets suggest that the data collected using our method capture universal characteristics of offensive language. Our findings also highlight the common words used in offensive communications, common targets for hate speech, specific patterns in violence tweets, and pinpoint common classification errors that can be attributed to limitations of NLP models. We observe that even state-of-the-art transformer models may fail to take into account culture, background, and context or understand nuances present in real-world data such as sarcasm. Hamdy Mubarak, Sabit Hassan, Shammur Absar Chowdhury |
Nat. Lang. Eng. | 2 |
| 2022 | Studying the Effect of Moderator Biases on the Diversity of Online Discussions: A Computational Cross-linguistic Study
Sabit Hassan, Katherine Atwell, Malihe Alikhani |
CogSci | 1 |
| 2022 | Learning cognitive and linguistic prosodic categories for automatic cross-lingual sign language understanding
Mert Inan, Sabit Hassan, Lorna C. Quandt, Malihe Alikhani |
CogSci | 3 |
| 2022 | APPDIA: A Discourse-aware Transformer-based Style Transfer Model for Offensive Social Media ConversationsabstractUsing style-transfer models to reduce offensiveness of social media comments can help foster a more inclusive environment. However, there are no sizable datasets that contain offensive texts and their inoffensive counterparts, and fine-tuning pretrained models with limited labeled data can lead to the loss of original meaning in the style-transferred text. To address this issue, we provide two major contributions. First, we release the first publicly-available, parallel corpus of offensive Reddit comments and their style-transferred counterparts annotated by expert sociolinguists. Then, we introduce the first discourse-aware style-transfer models that can effectively reduce offensiveness in Reddit text while preserving the meaning of the original text. These models are the first to examine inferential links between the comment and the text it is replying to when transferring the style of offensive Reddit text. We propose two different methods of integrating discourse relations with pretrained transformer models and evaluate them on our dataset of offensive comments from Reddit and their inoffensive counterparts. Improvements over the baseline with respect to both automatic metrics and human evaluation indicate that our discourse-aware models are better at preserving meaning in style-transferred text when compared to the state-of-the-art discourse-agnostic models. Katherine Atwell, Sabit Hassan, Malihe Alikhani |
COLING | 2 |
| 2022 | Cross-lingual Emotion DetectionabstractEmotion detection can provide us with a window into understanding human behavior. Due to the complex dynamics of human emotions, however, constructing annotated datasets to train automated models can be expensive. Thus, we explore the efficacy of cross-lingual approaches that would use data from a source language to build models for emotion detection in a target language. We compare three approaches, namely: i) using inherently multilingual models; ii) translating training data into the target language; and iii) using an automatically tagged parallel corpus. In our study, we consider English as the source language with Arabic and Spanish as target languages. We study the effectiveness of different classification models such as BERT and SVMs trained with different features. Our BERT-based monolingual models that are trained on target language data surpass state-of-the-art (SOTA) by 4% and 5% absolute Jaccard score for Arabic and Spanish respectively. Next, we show that using cross-lingual approaches with English data alone, we can achieve more than 90% and 80% relative effectiveness of the Arabic and Spanish BERT models respectively. Lastly, we use LIME to analyze the challenges of training cross-lingual models for different language pairs. Sabit Hassan, Shaden Shaar, Kareem Darwish |
LREC | 1 |
| 2022 | ArCovidVac: Analyzing Arabic Tweets About COVID-19 VaccinationabstractThe emergence of the COVID-19 pandemic and the first global infodemic have changed our lives in many different ways. We relied on social media to get the latest information about COVID-19 pandemic and at the same time to disseminate information. The content in social media consisted not only health related advice, plans, and informative news from policymakers, but also contains conspiracies and rumors. It became important to identify such information as soon as they are posted to make an actionable decision (e.g., debunking rumors, or taking certain measures for traveling). To address this challenge, we develop and publicly release the first largest manually annotated Arabic tweet dataset, ArCovidVac, for COVID-19 vaccination campaign, covering many countries in the Arab region. The dataset is enriched with different layers of annotation, including, (i) Informativeness more vs. less importance of the tweets); (ii) fine-grained tweet content types (e.g., advice, rumors, restriction, authenticate news/information); and (iii) stance towards vaccination (pro-vaccination, neutral, anti-vaccination). Further, we performed in-depth analysis of the data, exploring the popularity of different vaccines, trending hashtags, topics, and presence of offensiveness in the tweets. We studied the data for individual types of tweets and temporal changes in stance towards vaccine. We benchmarked the ArCovidVac dataset using transformer architectures for informativeness, content types, and stance detection. Hamdy Mubarak, Sabit Hassan, Shammur Absar Chowdhury, Firoj Alam |
LREC | 2 |
| 2019 | An Oracle Hierarchy for Small One-Way Finite Automata
Malek Anabtawi, Sabit Hassan, Christos A. Kapoutsis, Mohammad Zakzok |
LATA | 2 |
| 2018 | Interactive Evaluation of Classifiers Under Limited ResourcesabstractIn this paper, we propose strategies to estimate the accuracy of classifiers on a dataset when resource limitations restrict the number of instances for which true labels can be obtained. Our target scenarios include situations where the classifier output labels, but no scores, e.g. when the "classifier" is not an automated classifier but an inexpert human labeller who only outputs labels. Our objective is to optimally select a subset of the data to obtain true labels for, such that they provide the best estimate of classifier accuracy. We use techniques based on stratified sampling to address this problem. However, stratified sampling poses two challenges: i) how best to stratify the data, and ii) how to allocate samples among the strata. We propose a method of stratifying data and then present two novel interactive algorithms to approximate optimal allocation of samples to the strata. Our proposed methods for stratification and allocation are seen to outperform other popular approaches to the problem. Sabit Hassan, Shaden Shaar, Bhiksha Raj, Saquib Razak |
ICMLA | 1 |