VLDB 2026 Research / reviewers in the wild / expert
Mohamad Mahmoud Al Rahhal
dblp:210/0209 · also Mohamad Mahmoud Alrahhal
· DBLP profile ↗
14ranked-venue papers
3as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Classification of heart sound signals with Whisper modelabstractHeart sounds, or phonocardiograms (PCG), are important for diagnosing cardiovascular conditions, providing a non-invasive means to assess heart function through auscultation. Accurate classification of PCG signals can facilitate early detection of cardiac abnormalities, significantly improving patient outcomes. However, the complexity and variability of heart sound recordings present significant challenges for traditional classification methods, necessitating advanced approaches that can effectively handle the nuances of cardiac acoustics. This paper introduces a novel transfer learning approach that adapts OpenAI's Whisper model, originally designed for robust speech recognition, to the task of heart sound classification. In particular, we employ Whisper's encoder architecture to effectively capture acoustic features that generalize to cardiac auscultation, making it a promising candidate for PCG analysis. To tailor the model for this specialized task, we implement a modified encoder architecture optimized for heart sound characteristics. We process the input to the model using a Log-Mel spectrogram pipeline specifically designed to highlight the unique acoustic properties of PCG signals. Experimental results demonstrate that the adapted Whisper model achieves state-of-the-art performance, surpassing existing methods in both accuracy and robustness. Maryam Alotaibi, Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Nassim Ammour, Mansour Abdulaziz Al Zuair |
Connect. Sci. | 3 |
| 2025 | LoRA-CLIP: Efficient Low-Rank Adaptation of Large CLIP Foundation Model for Scene ClassificationabstractScene classification in optical remote sensing (RS) imagery has been extensively investigated using both learning-from-scratch approaches and fine-tuning of ImageNet pretrained models. Meanwhile, contrastive language-image pretraining (CLIP) has emerged as a powerful foundation model for vision-language tasks, demonstrating remarkable zero-shot capabilities across various domains. Its image encoder is a key component in many vision instruction-tuning models, enabling effective alignment of text and visual modalities for diverse tasks. However, its potential for supervised RS scene classification remains unexplored. This work investigates the efficient adaptation of large CLIP models (containing over 300 M parameters) through low-rank adaptation (LoRA), specifically targeting the attention layers. By applying LoRA to CLIP’s attention mechanisms, we can effectively adapt the vision model for specialized scene classification tasks with minimal computational overhead, requiring fewer training epochs than traditional fine-tuning methods. Our extensive experiments demonstrate the promising capabilities of LoRA-CLIP. By training only on a small set of additional parameters, LoRA-CLIP outperforms models pretrained on ImageNet, demonstrating the clear advantages of using image–text pretrained backbones for scene classification. Mohamad Mahmoud Al Rahhal, Yakoub Bazi, Mansour Abdulaziz Al Zuair |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | Enhancing Intrusion Detection in IoT Environments: An Advanced Ensemble Approach Using Kolmogorov-Arnold NetworksabstractIn recent years, the evolution of machine learning techniques has significantly impacted the field of intrusion de-tection, particularly within the context of the Internet of Things (IoT). As IoT networks expand, the need for robust security measures to counteract potential threats has become increasingly critical. This paper introduces a hybrid Intrusion Detection System (IDS) that synergistically combines Kolmogorov-Arnold Networks (KANs) with the XGBoost algorithm. Our proposed IDS leverages the unique capabilities of KANs, which utilize learnable activation functions to model complex relationships within data, alongside the powerful ensemble learning techniques of XGBoost, known for its high performance in classification tasks. This hybrid approach not only enhances the detection accuracy but also improves the interpretability of the model, making it suitable for dynamic and intricate IoT environments. Experimental evaluations demonstrate that our hybrid IDS achieves an impressive detection accuracy exceeding 99 % in dis-tinguishing between benign and malicious activities. Additionally, we were able to achieve F1-scores, precision, and recall that are exceeding 98%. Furthermore, we conduct a comparative analysis against traditional Multi-Layer Perceptron (MLP) networks, assessing performance metrics such as Precision, Recall, and F1-score. The results underscore the efficacy of integrating KANs with XGBoost, highlighting the potential of this innovative approach to significantly strengthen the security framework of IoT networks. Amar Amouri, Mohamad Mahmoud Al Rahhal, Yakoub Bazi, Ismail Butun, Imad Mahgoub |
ISNCC | 2 |
| 2024 | Text-to-Event Retrieval in Aerial VideosabstractRecent years have seen a rise in the use of video sensors in remote sensing (RS) as they represent a rich source for understanding earth dynamics and human activities. Accordingly, the volume of RS video data has rapidly increased. Accessing a video of interest from large repositories through text-to-video retrieval is preferable due to its flexibility and efficiency. In this letter, we propose a text-to-event retrieval model for aerial videos. The architecture of our model consists of two branches. The first is the video branch that extracts frame-level features from the video by using the vision transformer (ViT). Then, these features are concatenated into a unified representation and fed into a temporal attention module to incorporate the temporal aspects. The second branch is the text branch that extracts textual representations from the query by the bidirectional encoder representations from transformers (BERTs). The two branches are trained jointly on video and text pairs by minimizing a bidirectional contrastive loss. Experimental results on the CapERA dataset, which is an extension of the event recognition in aerial video (ERA) dataset, show the effectiveness of the proposed method. The dataset will be available athttps://www.github.com/yakoubbazi/CapEra. Laila Bashmal, Shima M. Al Mehmadi, Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Mansour Abdulaziz Al Zuair |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Open-Ended Visual Question Answering Model For Remote Sensing ImagesabstractIn this paper, we present an open-ended visual question answering (VQA) model for remote sensing images, where the answers can be given in the form of short sentences, unlike closed-ended VQA. This model uses a vision and natural language transformers for embedding the image and its related question. The feature representations obtained at the output are concatenated and fed to a light transformer decoder for generating the answer in an autoregressive way. The complete architecture is trained in an end-to-end manner via the backpropagation algorithm. In the experiments, we evaluate the model on a manually labeled open-ended VQA dataset termed TextRS composed of 6245 image-question pairs. Sara O. Alsaleh, Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Mansour Abdulaziz Al Zuair |
IGARSS | 3 |
| 2022 | Remote Sensing Image Retrieval Using Multilingual TextsabstractRecently, multimodal retrieval has attracted increasing attention in remote sensing community. in particular, text-image retrieval showed a promising research topic due to its ability to enable a flexible retrieval experience. To achieve this end, we propose a transformer-based multilingual text-image retrieval approach. Specifically, we employ the transformer encoder for both textual and visual modalities. At the text encoder, we jointly train four languages: English, Arabic, French, and Italian. We conduct experiments on fine-grained multimodal datasets named RSITMD. The experimental studies of the proposed method demonstrate its superior performance on both single and multiple languages modalities compared with the state-of-the-art methods. Norah A. Alsharif, Yakoub Bazi, Mohamad Mahmoud Al Rahhal |
IGARSS | 3 |
| 2022 | Convmixer with Selective Kernel Attention for Hyperspectral Image ClassificationabstractThis paper presents an efficient approach for hyperspectral image classification based on ConvMixer networks. To boost the capabilities of this network, we add an attention layer based on the idea of selective kernels (SK). This layer combines the information obtained by applying kernels of different sizes to the feature maps. The aim is to capture better the spatial and channel-wise relationships for an enhanced representation of the data. The experimental results obtained on two hyperspectral datasets: WHU-Hi-HanChuan and WHU-Hi-HongHu datasets, confirm the promising capabilities of the proposed method compared to the state-of-the-art. Bashair Alwadei, Mansour Abdulaziz Al Zuair, Mohamad Mahmoud Al Rahhal, Yakoub Bazi |
IGARSS | 3 |
| 2022 | Contrasting YOLOv5, Transformer, and EfficientDet Detectors for Crop Circle Detection in DesertabstractOngoing discoveries of water reserves have fostered an increasing adoption of crop circles in the desert in several countries. Automatically quantifying and surveying the layout of crop circles in remote areas can be of great use for stakeholders in managing the expansion of the farming land. This letter compares latest deep learning models for crop circle detection and counting, namely Detection Transformers, EfficientDet and YOLOv5 are evaluated. To this end, we build two datasets, via Google Earth Pro, corresponding to two large crop circle hot spots in Egypt and Saudi Arabia. The images were drawn at an altitude of 20 km above the targets. The models are assessed in within-domain and cross-domain scenarios, and yielded plausible detection potential and inference response. Mohamed Lamine Mekhalfi, Carlo Nicolò, Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Norah A. Alsharif, Eslam Al Maghayreh |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Bi-Modal Transformer-Based Approach for Visual Question Answering in Remote Sensing ImageryabstractRecently, vision-language models based on transformers are gaining popularity for joint modeling of visual and textual modalities. In particular, they show impressive results when transferred to several downstream tasks such as zero and few-shot classification. In this paper, we propose a visual question answering (VQA) approach for remote sensing images based on these models. The VQA task attempts to provide answers to image-related questions. While VQA has gained popularity in computer vision, in remote sensing it is not widespread. First, we use the contrastive language image pre-training (CLIP) network for embedding the image patches and question words into a sequence of visual and textual representations. Then, we learn attention mechanisms to capture the intra-and-inter dependencies within and between these representations. Afterward, we generate the final answer by averaging the predictions of two classifiers mounted on the top of the resulting contextual representations. In the experiments, we study the performance of the proposed approach on two datasets acquired with Sentinel-2 and aerial sensors. In particular, we demonstrate that our approach can achieve better results with reduced training size compared to the recent state-of-the-art. Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Mohamed Lamine Mekhalfi, Mansour Abdulaziz Al Zuair, Farid Melgani |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Deep Vision Transformers for Remote Sensing Scene ClassificationabstractIn this paper, we present a scene classification method based on vision transformers. These types of networks, which are now the standard models in natural language processing (NLP) do not rely on convolution block as in convolutional neural networks (CNNs). Alternatively, they are based on a mechanism known as multi-head self-attention (MSA), which captures the contextual relations between image pixels regardless of their spatial distance. At the first step, the images under analysis are split into patches, then converted to sequence by flattening and embedding. The embedding position is encoded and added to the sequence to preserve the order of the patches. Then, the resulting sequence is fed to several MSA layers for generating the final representation. To increase the classification performance, we employed several data augmentation strategies to expand the size and the diversity of the training data. Additionally, we show experimentally that we can compress the network by pruning half of its layers while keeping the competing performance. We further investigate the performance of the data-efficient image transformers (DeiT), a version of the model that is trained by knowledge distillation with less amount of data. Experimental results on two remote sensing datasets show that vision transformers can outperform state-of-the-art methods based on CNNs. Laila Bashmal, Yakoub Bazi, Mohamad Mahmoud Al Rahhal |
IGARSS | 3 |
| 2021 | Adversarial Learning for Knowledge Adaptation From Multiple Remote Sensing SourcesabstractIn this work, we introduce a neural architecture to unsupervised domain from multiple source domains. This architecture uses an EfficientNet as a feature extractor coupled with a set of Softmax classifiers equal to the number of source domains followed by an opportune fusion layer. To reduce the domain discrepancy between each source and target domain, we adopt a Minmax entropy approach that is based on the idea of optimizing in an adversarial manner the conditional entropy of the target samples with respect to each source classifier and minimizes it with respect to the feature extractor. As for the fusion module, we propose a weighted average fusion layer with learnable weights for aggregating the outputs of the different Softmax classifiers. Experiments on a multisource data set composed of images acquired by manned and unmanned aerial vehicles (MAVs/UAVs) over different locations are reported and discussed. Mohamad Mahmoud Al Rahhal, Yakoub Bazi, Huda Al-Hwiti, Haikel Salem Alhichri, Naif Alajlan |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Real-Time Mobile-Based Electrocardiogram System for Remote Monitoring of Patients with Cardiac ArrhythmiasabstractIn this study, we propose an electrocardiogram (ECG) system for the simultaneous and remote monitoring of multiple heart patients. It consists of three main components: patient, sever, and monitoring units. The patient unit uses a wearable miniature sensor that continuously measures ECG signals and sends them to a smart mobile phone via a Bluetooth connection. In the mobile device, the ECG signals can be stored, displayed on screen, and automatically transmitted to a distant server unit over the internet; the server stores ECG data from several patients. Health care stakeholders use a monitoring unit to retrieve the ECG signals of multiple patients at any time from the server for display and real-time automatic analysis. The analysis includes segmentation of the ECG signal into separate heartbeats followed by arrhythmia detection and classification. When compared to existing real-time ECG systems, where the detection of abnormalities is usually performed using simple rules, the proposed system implements a real-time classification module that is based on a support vector machine (SVM) classifier. Extensive experimental results on ECG data obtained from a TechPatientTMsimulator, a real person, and 20 records from the MIT arrhythmia database are reported and discussed. Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Haikel Salem Alhichri, Nassim Ammour, Naif Alajlan, Mansour Abdulaziz Al Zuair |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2018 | Asymmetric Adaptation of Deep Features for Cross-Domain Classification in Remote Sensing ImageryabstractIn this letter, we introduce an asymmetric adaptation neural network (AANN) method for cross-domain classification in remote sensing images. Before the adaptation process, we feed the features obtained from a pretrained convolutional neural network to a denoising autoencoder (DAE) to perform dimensionality reduction. Then the first hidden layer of AANN (placed on the top of DAE) maps the labeled source data to the target space, while the subsequent layers control the separation between the available land-cover classes. To learn its weights, the network minimizes an objective function composed of two losses related to the distance between the source and target data distributions and class separation. The results of experiments conducted on six scenarios built from three benchmark scene remote sensing data sets (i.e., Merced, KSA, and AID data sets) are reported and discussed. Nassim Ammour, Laila Bashmal, Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Mansour Abdulaziz Al Zuair |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2016 | Deep learning approach for active classification of electrocardiogram signals
Mohamad Mahmoud Al Rahhal, Yakoub Bazi, Haikel Salem Alhichri, Naif Alajlan, Farid Melgani, Ronald R. Yager |
Inf. Sci. | 1 |