VLDB 2026 Research / reviewers in the wild / expert
Abdelkader Baggag
dblp:12/5306
· DBLP profile ↗
9ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-8742-5519ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rapid disaster damage assessment using deep adversarial sliced Wasserstein domain adaptationabstractAbstract Disasters vary greatly in their nature, severity and underlying distribution. A streamlined model is crucial for assessing the impact of unexpected disasters. Previous methods rely on training with historical disaster data to evaluate the damage caused. However, in practice, this approach rarely achieves acceptable performance due to domain shift. Therefore, existing models need to be adapted to the emergent disaster quickly. A promising way to achieve this goal is through unsupervised domain adaptation (UDA). To this end, many advancements have been made toward better transferability of the model, including the Domain Adversarial Neural Network (DANN), which utilizes adversarial training to attain a domain-invariant feature extractor. The Adversarial sliced Wasserstein Domain Adaptation Network (AWDAN) further improves on DANN by using sliced Wasserstein distance as a measure between the features extracted from the source and target domains. Inspired by these advancements, we explore the utility of UDA approaches in the disaster response domain. Specifically, we perform extensive experiments on real-world images collected from Twitter during four major disasters. We train a total of 216 models to benchmark six methods across all possible source–target domain combinations. We improve the performance of the previous state-of-the-art DANN-based method for rapid damage assessment by enhancing it with a deeper backbone architecture to learn better feature representations. Furthermore, we adopt AWDAN to more effectively mitigate the distribution shift in data obtained from different disaster events. Experimental results demonstrate that the proposed approach achieves statistically significant performance gains, with up to 11.4% improvement in F1-score and 8.9% improvement in accuracy over the source-only model, consistently outperforming several state-of-the-art domain adaptation frameworks, including DANN, CORAL, MMD, and CDAN. Fatma AlNaimi, Abdulaziz Al-Homaid, Ferda Ofli, Abdelkader Baggag |
Neural Comput. Appl. | 4 |
| 2025 | ArnoldiGCL: Graph Contrastive Learning via Learnable Arnoldi-Based Guided Spectral Chebyshev Polynomial FiltersabstractGraph Contrastive Learning (GCL) emerged as a powerful paradigm in self-supervised graph representation learning. While earlier applications of GCL rely on homophily assumptions, spectral graph neural networks (GNNs) enhance the effectiveness of GCL on heterophilic graphs by incorporating both low-pass and high-pass filters. However, due to numerical considerations, existing approaches oversimplify low-pass and high-pass filters by modeling them as basic linear operations, failing to capture complex topological relationships. Mustafa Coskun, Abdelkader Baggag, Mehmet Koyutürk |
KDD (2) | 2 |
| 2025 | How Much Wearable Data is Enough for the Utility and Trust of Augmented Artificial Intelligence Systems? A Scenario-Based Interview with Medical ProfessionalsabstractThe paper explores the synergy between wearable data and augmented Artificial Intelligence (AI) through findings from two interconnected studies. The first study (study 1) focuses on medical professionals’ perceptions of wearable data and AI, and the second study (Study 2) extends it focuses on how differences in the level of granularity in the data presented affect the professionals’ understanding, interpretation, and trust in AI recommendations. This system allows medical professionals to view AI-generated recommendations for sleep and activity improvement and explanations of the underlying rationale. While each study has distinct research questions, Study 2 builds upon Study 1's foundation. Both studies employed scenario-based interviews. Thematic analysis of Study 1 identified trust as a crucial factor in the acceptance of wearable data and AI, influencing Study 2's exploration of factors affecting trust, such as explainability, data granularity, representativeness, and user interaction. Study 2 highlighted varying perspectives on information sufficiency and data sharing from the AI system linked to professionals’ roles and tasks. The work offers insights into data granularity’s impact on engagement with AI recommendations. Yasmin Abdelaal, Michaël Aupetit 0001, Abdelkader Baggag, Mohammed Bashir, Dena Al-Thani |
Int. J. Hum. Comput. Interact. | 3 |
| 2023 | OutSingle: a novel method of detecting and injecting outliers in RNA-Seq count data using the optimal hard threshold for singular valuesabstractMOTIVATION: Finding outliers in RNA-sequencing (RNA-Seq) gene expression (GE) can help in identifying genes that are aberrant and cause Mendelian disorders. Recently developed models for this task rely on modeling RNA-Seq GE data using the negative binomial distribution (NBD). However, some of those models either rely on procedures for inferring NBD's parameters in a nonbiased way that are computationally demanding and thus make confounder control challenging, while others rely on less computationally demanding but biased procedures and convoluted confounder control approaches that hinder interpretability. RESULTS: In this article, we present OutSingle (Outlier detection using Singular Value Decomposition), an almost instantaneous way of detecting outliers in RNA-Seq GE data. It uses a simple log-normal approach for count modeling. For confounder control, it uses the recently discovered optimal hard threshold (OHT) method for noise detection, which itself is based on singular value decomposition (SVD). Due to its SVD/OHT utilization, OutSingle's model is straightforward to understand and interpret. We then show that our novel method, when used on RNA-Seq GE data with real biological outliers masked by confounders, outcompetes the previous state-of-the-art model based on an ad hoc denoising autoencoder. Additionally, OutSingle can be used to inject artificial outliers masked by confounders, which is difficult to achieve with previous approaches. We describe a way of using OutSingle for outlier injection and proceed to show how OutSingle outperforms its competition on 16 out of 18 datasets that were generated from three real datasets using OutSingle's injection procedure with different outlier types and magnitudes. Our methods are applicable to other types of similar problems involving finding outliers in matrices under the presence of confounders. AVAILABILITY AND IMPLEMENTATION: The code for OutSingle is available at https://github.com/esalkovic/outsingle. Edin Salkovic, Mohammad Amin Sadeghi, Abdelkader Baggag, Ahmed Gamal Rashed Salem, Halima Bensmail |
Bioinform. | 3 |
| 2021 | Computational prediction and interpretation of both general and specific types of promoters in Escherichia coli by exploiting a stacked ensemble-learning frameworkabstractPromoters are short consensus sequences of DNA, which are responsible for transcription activation or the repression of all genes. There are many types of promoters in bacteria with important roles in initiating gene transcription. Therefore, solving promoter-identification problems has important implications for improving the understanding of their functions. To this end, computational methods targeting promoter classification have been established; however, their performance remains unsatisfactory. In this study, we present a novel stacked-ensemble approach (termed SELECTOR) for identifying both promoters and their respective classification. SELECTOR combined the composition of k-spaced nucleic acid pairs, parallel correlation pseudo-dinucleotide composition, position-specific trinucleotide propensity based on single-strand, and DNA strand features and using five popular tree-based ensemble learning algorithms to build a stacked model. Both 5-fold cross-validation tests using benchmark datasets and independent tests using the newly collected independent test dataset showed that SELECTOR outperformed state-of-the-art methods in both general and specific types of promoter prediction in Escherichia coli. Furthermore, this novel framework provides essential interpretations that aid understanding of model success by leveraging the powerful Shapley Additive exPlanation algorithm, thereby highlighting the most important features relevant for predicting both general and specific types of promoters and overcoming the limitations of existing 'Black-box' approaches that are unable to reveal causal relationships from large amounts of initially encoded features. Fuyi Li, ZongYuan Ge, Yanwei Yue, Morihiro Hayashida, Abdelkader Baggag, Halima Bensmail, Jiangning Song |
Briefings Bioinform. | 7 |
| 2021 | Fast computation of Katz index for efficient processing of link prediction queries
Mustafa Coskun, Abdelkader Baggag, Mehmet Koyutürk |
Data Min. Knowl. Discov. | 2 |
| 2021 | Learning Spatiotemporal Latent Factors of Traffic via Regularized Tensor Factorization: Imputing Missing Values and ForecastingabstractIntelligent transportation systems are a key component in smart cities, and the estimation and prediction of the spatiotemporal traffic state is critical to capture the dynamics of traffic congestion, i.e., its generation, propagation and mitigation, in order to increase operational efficiency and improve livability within smart cities. And while spatiotemporal data related to traffic is becoming common place due to the wide availability of cheap sensors and the rapid deployment of IoT platforms, the data still suffer some challenges related to sparsity, incompleteness, and noise which makes the traffic analytics difficult. In this article, we investigate the problem of missing data or noisy information in the context of real-time monitoring and forecasting of traffic congestion for road networks in a city. The road network is represented as a directed graph in which nodes are junctions (intersections) and edges are road segments. We assume that the city has deployed high-fidelity sensors for speed reading in a subset of edges; and the objective is to infer the speed readings for the remaining edges in the network; and to estimate the missing values in the segments for which sensors have stopped generating data due to technical problems (e.g., battery, network, etc.). We propose a tensor representation for the series of road network snapshots, and develop a regularized factorization method to estimate the missing values, while learning the latent factors of the network. The regularizer, which incorporates spatial properties of the road network, improves the quality of the results. The learned factors, with a graph-based temporal dependency, are then used in an autoregressive algorithm to predict the future state of the road network with a large horizon. Extensive numerical experiments with real traffic data from the cities of Doha (Qatar) and Aarhus (Denmark) demonstrate that the proposed approach is appropriate for imputing the missing data and predicting the traffic state. It is accurate and efficient and can easily be applied to other traffic datasets. Abdelkader Baggag, Sofiane Abbar, Ankit Sharma 0004, Tahar Zanouda, Abdulaziz Al-Homaid, Abhiraj Mohan, Jaideep Srivastava |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | Comprehensive review and assessment of computational methods for predicting RNA post-transcriptional modification sites from RNA sequencesabstractRNA post-transcriptional modifications play a crucial role in a myriad of biological processes and cellular functions. To date, more than 160 RNA modifications have been discovered; therefore, accurate identification of RNA-modification sites is fundamental for a better understanding of RNA-mediated biological functions and mechanisms. However, due to limitations in experimental methods, systematic identification of different types of RNA-modification sites remains a major challenge. Recently, more than 20 computational methods have been developed to identify RNA-modification sites in tandem with high-throughput experimental methods, with most of these capable of predicting only single types of RNA-modification sites. These methods show high diversity in their dataset size, data quality, core algorithms, features extracted and feature selection techniques and evaluation strategies. Therefore, there is an urgent need to revisit these methods and summarize their methodologies, in order to improve and further develop computational techniques to identify and characterize RNA-modification sites from the large amounts of sequence data. With this goal in mind, first, we provide a comprehensive survey on a large collection of 27 state-of-the-art approaches for predicting N1-methyladenosine and N6-methyladenosine sites. We cover a variety of important aspects that are crucial for the development of successful predictors, including the dataset quality, operating algorithms, sequence and genomic features, feature selection, model performance evaluation and software utility. In addition, we also provide our thoughts on potential strategies to improve the model performance. Second, we propose a computational approach called DeepPromise based on deep learning techniques for simultaneous prediction of N1-methyladenosine and N6-methyladenosine. To extract the sequence context surrounding the modification sites, three feature encodings, including enhanced nucleic acid composition, one-hot encoding, and RNA embedding, were used as the input to seven consecutive layers of convolutional neural networks (CNNs), respectively. Moreover, DeepPromise further combined the prediction score of the CNN-based models and achieved around 43% higher area under receiver-operating curve (AUROC) for m1A site prediction and 2-6% higher AUROC for m6A site prediction, respectively, when compared with several existing state-of-the-art approaches on the independent test. In-depth analyses of characteristic sequence motifs identified from the convolution-layer filters indicated that nucleotide presentation at proximal positions surrounding the modification sites contributed most to the classification, whereas those at distal positions also affected classification but to different extents. To maximize user convenience, a web server was developed as an implementation of DeepPromise and made publicly available at http://DeepPromise.erc.monash.edu/, with the server accepting both RNA sequences and genomic sequences to allow prediction of two types of putative RNA-modification sites. Zhen Chen 0009, Fuyi Li, Yanan Wang 0003, Alexander Ian Smith, Geoffrey I. Webb, Tatsuya Akutsu, Abdelkader Baggag, Halima Bensmail, Jiangning Song |
Briefings Bioinform. | 8 |
| 2017 | Advanced Computation of Sparse Precision Matrices for Big Data
Abdelkader Baggag, Halima Bensmail, Jaideep Srivastava |
PAKDD (2) | 1 |