VLDB 2026 Research / reviewers in the wild / expert
Jamil Saquer
dblp:15/3503
· DBLP profile ↗
15ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0006-7948-8409ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-Tuned Transformers vs. Zero-Shot Large Language Models for Mental Health Detection on Social Media
Mahmoud Abusaqer, Jessica Nwobodo, Jamil Saquer |
COMPSAC | 3 |
| 2025 | A Comparative Analysis of Transformer and Traditional ML Models for Cyberbullying Detection on Twitter (now X)abstractCyberbullying on social media poses critical risks to mental health and public safety. This paper investigates advanced computational approaches for detecting and classifying cyberbullying in a large Twitter dataset of 47,692 tweets, labeled into six categories. We compare three transformer-based models (GPT-3.5, BERT, and RoBERTa) against three traditional machine learning algorithms (Naïve Bayes, SVM, and Random Forest), evaluating accuracy, precision, recall, F1-score, and computational efficiency. Our results indicate that RoBERTa achieves the highest overall performance (87–88% accuracy) but at a higher computational cost (13 hours on CPU), while Random Forest offers a strong balance between speed and performance (85.36% accuracy in 83 seconds). In contrast, the experiment using GPT-3.5 in a batched, zero-shot configuration achieved accuracy of 25.41%, an F1-score of 23.61%, and elapsed time of 5.14 hours, highlighting the challenges of applying generative models to cyberbullying detection without fine-tuning. These findings inform model selection for real-world deployment of cyberbullying detection systems, illuminating the trade-offs between transformer-based and traditional methods for automated social media monitoring. Muhammad Abusaqer, Jamil Saquer |
COMPSAC | 2 |
| 2025 | Advancing Mental Disorder Detection: A Comparative Evaluation of Transformer and LSTM Architectures on Social MediaabstractThe rising prevalence of mental health disorders necessitates the development of robust, automated tools for early detection and monitoring. Recent advances in Natural Language Processing (NLP), particularly transformer-based architectures, have demonstrated significant potential in text analysis. This study provides a comprehensive evaluation of state-of-the-art transformer models (BERT, RoBERTa, DistilBERT, ALBERT, and ELECTRA) against Long Short-Term Memory (LSTM) based approaches using different text embedding techniques for mental health disorder classification on Reddit. We construct a large annotated dataset, validating its reliability through statistical judgmental analysis and topic modeling. Experimental results demonstrate the superior performance of transformer models over traditional deep-learning approaches. RoBERTa achieved the highest classification performance, with a 99.54% F1 score on the hold-out test set and a 96.05% F1 score on the external test set. Notably, LSTM models augmented with BERT embeddings proved highly competitive, achieving F1 scores exceeding 94% on the external dataset while requiring significantly fewer computational resources. These findings highlight the effectiveness of transformer-based models for real-time, scalable mental health monitoring. We discuss the implications for clinical applications and digital mental health interventions, offering insights into the capabilities and limitations of state-of-the-art NLP methodologies in mental disorder detection. Khalid Hasan, Jamil Saquer, Mukulika Ghosh |
COMPSAC | 2 |
| 2025 | Multiclass Hate Speech Detection: Evaluating 303 Model Configurations Across Traditional Machine Learning, Deep Learning, and Transformer ApproachesabstractMulticlass hate speech detection across demographic categories remains challenging due to implicit targeting strategies and linguistic variability in social media content. This study provides an evaluation of 303 model configurations across three methodological paradigms: traditional machine learning (147 configurations), deep learning with pre-trained embeddings (144 configurations), and transformer models (12 implementations). Using 39,747 tweets spanning five demographic hate speech categories (age, ethnicity, gender, religion, other_hate), we conducted 5-fold cross-validation to establish performance benchmarks across all approaches. Results reveal a clear performance hierarchy with RoBERTa achieving 95.02% accuracy and 95.04% weighted F1-score, followed by Word2Vec-enhanced InceptionCNN at 94.31% accuracy and 94.35% F1-score. Traditional machine learning demonstrates exceptional efficiency, with SGDClassifier achieving 94.10% accuracy and 94.17% F1-score in only 11.5 seconds for a complete cross-validation cycle. Gender-based and other_hate speech categories prove most challenging across all methodologies, exhibiting distinct linguistic patterns that complicate automated detection. These findings provide empirical guidance for developing scalable hate speech detection systems capable of fine-grained demographic targeting identification. Mahmoud Abusaqer, Khalid Hasan, Jamil Saquer |
ICMLA | 3 |
| 2025 | Mental Multi-class Classification on Social Media: Benchmarking Transformer Architectures against LSTM ModelsabstractMillions of people openly share mental health struggles on social media, providing rich data for early detection of conditions such as depression, bipolar disorder, etc. However, most prior Natural Language Processing (NLP) research has focused on single-disorder identification, leaving a gap in understanding the efficacy of advanced NLP techniques for distinguishing among multiple mental health conditions. In this work, we present a large-scale comparative study of state-of-the-art transformer versus Long Short-Term Memory (LSTM)-based models to classify mental health posts into exclusive categories of mental health conditions. We first curate a large dataset of Reddit posts spanning six mental health conditions and a control group, using rigorous filtering and statistical exploratory analysis to ensure annotation quality. We then evaluate five transformer architectures (BERT, RoBERTa, DistilBERT, ALBERT, and ELECTRA) against several LSTM variants (with or without attention, using contextual or static embeddings) under identical conditions. Experimental results show that transformer models consistently outperform the alternatives, with RoBERTa achieving 91-99% F1-scores and accuracies across all classes. Notably, attention-augmented LSTMs with BERT embeddings approach transformer performance (up to 97% F1-score) while training 2-3.5 times faster, whereas LSTMs using static embeddings fail to learn useful signals. These findings represent the first comprehensive benchmark for multi-class mental health detection, offering practical guidance on model selection and highlighting an accuracy–efficiency trade-off for real-world deployment of mental health NLP systems. Khalid Hasan, Jamil Saquer |
ICMLA | 2 |
| 2025 | eRACANN: Modular Neural Agents for Interpretable Cloud Resource ForecastingabstractAs cloud computing becomes foundational to sectors like healthcare, finance, and artificial intelligence, accurate resource utilization forecasting has emerged as a critical challenge for ensuring efficiency, cost-effectiveness, and service reliability. This research introduces Runtime-Assembled Context-Specific Cooperative Artificial Neural Networks (RACANN), a novel modular neural framework designed to predict cloud resource utilization with improved interpretability and efficiency. Unlike monolithic deep learning models, RACANN dynamically assembles lightweight context-specific neural agents at runtime, enabling fine-grained temporal adaptability and significantly reducing computational overhead. We further propose eRACANN, an explainable extension that embeds a fuzzy logic-based linguistic layer, offering human-readable justifications for predictions. Experimental results across both structured and noisy real-world datasets demonstrate that while deep neural networks (DNNs) may achieve marginally lower error rates, eRACANN excels in modularity, interpretability, and contextual transparency which are critical properties for operational deployment in dynamic, mission-critical cloud environments. The RACANN framework offers a scalable path toward explainable, adaptive, and resource-efficient AI for cloud infrastructure management. Nathan Nelson, Shusmoy Chowdhury, Ajay K. Katangur, Siming Liu 0001, Jamil Saquer |
JCC | 5 |
| 2025 | Parameter-Efficient Hate Speech Detection: A Comprehensive Evaluation of LoRA-Adapted LLMs Across 18 ArchitecturesabstractThe proliferation of hate speech on social media platforms necessitates automated detection systems that balance accuracy with computational efficiency.This study evaluates 18 LoRA-adapted Large Language Models (LLMs) across three model families (Llama, Phi, Qwen) at scales ranging from 0.5B to 14B parameters for hate speech detection.Using a dataset of 48,049 samples, we investigate the relationship between model architecture, size, instruction tuning, and detection performance while focusing on parameter-efficient fine-tuning approaches.Our results demonstrate that Phi-4 achieves state-of-the-art performance of 91.84% accuracy and 92.79% F1-score despite using only 0.15% of its parameters for task-specific adaptation.Smaller models like Llama-3.2-1B show competitive performance (91.53% accuracy, 92.69% F1-score) with significantly faster training times.We find that architectural design often outweighs parameter count, with smaller models frequently outperforming larger counterparts within the same family.The inconsistent benefits of instruction tuning across architectures highlight the importance of task-specific adaptation for hate speech detection.These findings offer valuable insights for developing efficient and effective hate speech detection systems across diverse deployment scenarios, from high-performance servers to resource-constrained environments. Mahmoud Abusaqer, Jamil Saquer |
SEKE | 2 |
| 2025 | Beyond Architectures: Evaluating the Role of Contextual Embeddings in Detecting Bipolar Disorder on Social MediaabstractBipolar disorder is a chronic mental illness frequently underdiagnosed due to subtle early symptoms and social stigma.This paper explores the advanced natural language processing (NLP) models for recognizing signs of bipolar disorder based on user-generated social media text.We conduct a comprehensive evaluation of transformer-based models (BERT, RoBERTa, ALBERT, ELECTRA, DistilBERT) and Long Short Term Memory (LSTM) models based on contextualized (BERT) and static (GloVe, Word2Vec) word embeddings.Experiments were performed on a large, annotated dataset of Reddit posts after confirming their validity through sentiment variance and judgmental analysis.Our results demonstrate that RoBERTa achieves the highest performance among transformer models with an F1 score of ∼98% while LSTM models using BERT embeddings yield nearly identical results.In contrast, LSTMs trained on static embeddings fail to capture meaningful patterns, scoring near-zero F1.These findings underscore the critical role of contextual language modeling in detecting bipolar disorder.In addition, we report model training times and highlight that DistilBERT offers an optimal balance between efficiency and accuracy.In general, our study offers actionable insights for model selection in mental health NLP applications and validates the potential of contextualized language models to support early bipolar disorder screening. Khalid Hasan, Jamil Saquer |
SEKE | 2 |
| 2024 | A Comparative Analysis of Transformer and LSTM Models for Detecting Suicidal Ideation on RedditabstractSuicide is a critical global health problem involving more than 700,000 deaths yearly, particularly among young adults. Many people express their suicidal thoughts on social media platforms such as Reddit. This paper evaluates the effectiveness of the deep learning transformer-based models BERT, RoBERTa, DistilBERT, ALBERT, and ELECTRA and various Long Short-Term Memory (LSTM) based models in detecting suicidal ideation from user posts on Reddit. Toward this objective, we curated an extensive dataset from diverse subreddits and conducted linguistic, topic modeling, and statistical analyses to ensure data quality. Our results indicate that each model could reach high accuracy and F1 scores, but among them, RoBERTa emerged as the most effective model with an accuracy of 93.22% and F1 score of 93.14%. An LSTM model that uses attention and BERT embeddings performed as the second best, with an accuracy of 92.65 % and an F1 score of 92.69 %. Our findings show that transformer-based models have the potential to improve suicide ideation detection, thereby providing a path to develop robust mental health monitoring tools from social media. This research, therefore, underlines the undeniable prospect of advanced techniques in Natural Language Processing (NLP) while improving suicide prevention efforts. Khalid Hasan, Jamil Saquer |
ICMLA | 2 |
| 2024 | Neo4j's BFS and DFS Evaluation in GDS and APOC Libraries with SPL Feature Models (P)abstractThis paper evaluates the performance of Breadth-First Search (BFS) and Depth-First Search (DFS) algorithms in Neo4j using the Graph Data Science (GDS) and the Awesome Procedures on Cypher (APOC) libraries.We benchmark these algorithms on feature models of varying complexity to assess their efficiency and scalability.Our findings indicate that GDS significantly outperforms APOC in terms of traversal times, particularly for large and complex graphs.This study underscores the importance of algorithm optimization in graph databases and provides insights into practical applications and future directions for improving feature model management in software product lines. Hazim Shatnawi, Jamil Saquer |
SEKE | 2 |
| 2016 | A Virtual Xylophone for Music EducationabstractThis paper describes the implementation of a virtual xylophone. During a setup phase, the program registers a background depth image, generated by a Kinect sensor, and the user interacts with the program to identify the color of tone bars and to select a restricted track space for tracking mallet locations. During a play phase, the program tracks mallet heads by locating pixels that are in front of the pixels registered in the background image. The program can easily be modified to restrict the notes available to the player or to use pentatonic or other musical scales. Nikolas Burks, Lloyd Smith, Jamil Saquer |
ISM | 3 |
| 2004 | Monotone concepts for formal concept analysis
Jitender S. Deogun, Jamil Saquer |
Discret. Appl. Math. | 2 |
| 2003 | Approximating Monotone Concepts
Jamil Saquer, Jitender S. Deogun |
HIS | 1 |
| 2001 | Discovering Representative Episodal Association Rules from Event Sequences Using Frequent Closed Episode Sets and Event ConstraintsabstractDiscovering association rules from time-series data is an important data mining problem. The number of potential rules grows quickly as the number of items in the antecedent grows. It is therefore difficult for an expert to analyze the rules and identify the useful. An approach for generating representative association rules for transactions that uses only a subset of the set of frequent itemsets called frequent closed itemsets was presented by Saquer and Deogun (2000). We employ formal concept analysis to develop the notion of frequent closed episodes. The concept of representative association rules is formalized in the context of event sequences. Applying constraints to target highly, significant rules further reduces the number of rules. Our approach results in a significant reduction of the number of rules generated, while maintaining the minimum set of relevant association rules and retaining the ability to generate the entire set of association rules with respect to the given constraints. We show how our method can be used to discover associations in a drought risk management decision support system and use multiple climatology datasets related to automated weather stations. Sherri K. Harms, Jitender S. Deogun, Jamil Saquer, Tsegaye Tadesse |
ICDM | 3 |
| 2000 | Using Closed Itemsets for Discovering Representative Association Rules
Jamil Saquer, Jitender S. Deogun |
ISMIS | 1 |