EDBT 2026 Demo / reviewers in the wild / expert
Mansour Abdulaziz Al Zuair
dblp:162/2082 · also Mansour Zuair
· DBLP profile ↗
13ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0003-2490-5739ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Classification of heart sound signals with Whisper modelabstractHeart sounds, or phonocardiograms (PCG), are important for diagnosing cardiovascular conditions, providing a non-invasive means to assess heart function through auscultation. Accurate classification of PCG signals can facilitate early detection of cardiac abnormalities, significantly improving patient outcomes. However, the complexity and variability of heart sound recordings present significant challenges for traditional classification methods, necessitating advanced approaches that can effectively handle the nuances of cardiac acoustics. This paper introduces a novel transfer learning approach that adapts OpenAI's Whisper model, originally designed for robust speech recognition, to the task of heart sound classification. In particular, we employ Whisper's encoder architecture to effectively capture acoustic features that generalize to cardiac auscultation, making it a promising candidate for PCG analysis. To tailor the model for this specialized task, we implement a modified encoder architecture optimized for heart sound characteristics. We process the input to the model using a Log-Mel spectrogram pipeline specifically designed to highlight the unique acoustic properties of PCG signals. Experimental results demonstrate that the adapted Whisper model achieves state-of-the-art performance, surpassing existing methods in both accuracy and robustness. Maryam Alotaibi, Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Nassim Ammour, Mansour Abdulaziz Al Zuair |
Connect. Sci. | 5 |
| 2025 | LoRA-CLIP: Efficient Low-Rank Adaptation of Large CLIP Foundation Model for Scene ClassificationabstractScene classification in optical remote sensing (RS) imagery has been extensively investigated using both learning-from-scratch approaches and fine-tuning of ImageNet pretrained models. Meanwhile, contrastive language-image pretraining (CLIP) has emerged as a powerful foundation model for vision-language tasks, demonstrating remarkable zero-shot capabilities across various domains. Its image encoder is a key component in many vision instruction-tuning models, enabling effective alignment of text and visual modalities for diverse tasks. However, its potential for supervised RS scene classification remains unexplored. This work investigates the efficient adaptation of large CLIP models (containing over 300 M parameters) through low-rank adaptation (LoRA), specifically targeting the attention layers. By applying LoRA to CLIP’s attention mechanisms, we can effectively adapt the vision model for specialized scene classification tasks with minimal computational overhead, requiring fewer training epochs than traditional fine-tuning methods. Our extensive experiments demonstrate the promising capabilities of LoRA-CLIP. By training only on a small set of additional parameters, LoRA-CLIP outperforms models pretrained on ImageNet, demonstrating the clear advantages of using image–text pretrained backbones for scene classification. Mohamad Mahmoud Al Rahhal, Yakoub Bazi, Mansour Abdulaziz Al Zuair |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Text-to-Event Retrieval in Aerial VideosabstractRecent years have seen a rise in the use of video sensors in remote sensing (RS) as they represent a rich source for understanding earth dynamics and human activities. Accordingly, the volume of RS video data has rapidly increased. Accessing a video of interest from large repositories through text-to-video retrieval is preferable due to its flexibility and efficiency. In this letter, we propose a text-to-event retrieval model for aerial videos. The architecture of our model consists of two branches. The first is the video branch that extracts frame-level features from the video by using the vision transformer (ViT). Then, these features are concatenated into a unified representation and fed into a temporal attention module to incorporate the temporal aspects. The second branch is the text branch that extracts textual representations from the query by the bidirectional encoder representations from transformers (BERTs). The two branches are trained jointly on video and text pairs by minimizing a bidirectional contrastive loss. Experimental results on the CapERA dataset, which is an extension of the event recognition in aerial video (ERA) dataset, show the effectiveness of the proposed method. The dataset will be available athttps://www.github.com/yakoubbazi/CapEra. Laila Bashmal, Shima M. Al Mehmadi, Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Mansour Abdulaziz Al Zuair |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Open-Ended Visual Question Answering Model For Remote Sensing ImagesabstractIn this paper, we present an open-ended visual question answering (VQA) model for remote sensing images, where the answers can be given in the form of short sentences, unlike closed-ended VQA. This model uses a vision and natural language transformers for embedding the image and its related question. The feature representations obtained at the output are concatenated and fed to a light transformer decoder for generating the answer in an autoregressive way. The complete architecture is trained in an end-to-end manner via the backpropagation algorithm. In the experiments, we evaluate the model on a manually labeled open-ended VQA dataset termed TextRS composed of 6245 image-question pairs. Sara O. Alsaleh, Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Mansour Abdulaziz Al Zuair |
IGARSS | 4 |
| 2022 | Convmixer with Selective Kernel Attention for Hyperspectral Image ClassificationabstractThis paper presents an efficient approach for hyperspectral image classification based on ConvMixer networks. To boost the capabilities of this network, we add an attention layer based on the idea of selective kernels (SK). This layer combines the information obtained by applying kernels of different sizes to the feature maps. The aim is to capture better the spatial and channel-wise relationships for an enhanced representation of the data. The experimental results obtained on two hyperspectral datasets: WHU-Hi-HanChuan and WHU-Hi-HongHu datasets, confirm the promising capabilities of the proposed method compared to the state-of-the-art. Bashair Alwadei, Mansour Abdulaziz Al Zuair, Mohamad Mahmoud Al Rahhal, Yakoub Bazi |
IGARSS | 2 |
| 2022 | Blockchained service provisioning and malicious node detection via federated learning in scalable Internet of Sensor Things networks
Zain Abubaker, Nadeem Javaid, Ahmad S. Al-Mogren, Mariam Akbar, Mansour Abdulaziz Al Zuair, Jalel Ben-Othman |
Comput. Networks | 5 |
| 2022 | Bi-Modal Transformer-Based Approach for Visual Question Answering in Remote Sensing ImageryabstractRecently, vision-language models based on transformers are gaining popularity for joint modeling of visual and textual modalities. In particular, they show impressive results when transferred to several downstream tasks such as zero and few-shot classification. In this paper, we propose a visual question answering (VQA) approach for remote sensing images based on these models. The VQA task attempts to provide answers to image-related questions. While VQA has gained popularity in computer vision, in remote sensing it is not widespread. First, we use the contrastive language image pre-training (CLIP) network for embedding the image patches and question words into a sequence of visual and textual representations. Then, we learn attention mechanisms to capture the intra-and-inter dependencies within and between these representations. Afterward, we generate the final answer by averaging the predictions of two classifiers mounted on the top of the resulting contextual representations. In the experiments, we study the performance of the proposed approach on two datasets acquired with Sentinel-2 and aerial sensors. In particular, we demonstrate that our approach can achieve better results with reduced training size compared to the recent state-of-the-art. Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Mohamed Lamine Mekhalfi, Mansour Abdulaziz Al Zuair, Farid Melgani |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | A Lightweight Privacy-Aware IoT-Based Metering Scheme for Smart Industrial EcosystemsabstractThe smart grid emerges as a new era of the electronic power grid. It integrates advanced sensing technologies, communications, and controlling methods that tell how electricity travels from different generation points to consumers. In order to fulfill customers' satisfaction and two way communications, a huge number of smart meters are deployed in different countries for real-time consumption and presentation of the rigorous energy usage. The privacy of industrial ecosystems may require greater attention while considering minimum network load, lower computational resources, better energy efficiency, and accuracy of data. Different research works have been done to tackle customers' privacy, but at the cost of using more computational resources, communication overhead, and hiring of a trusting third party. In this article, we have proposed a symmetric encryption scheme for industrial ecosystems in the Internet of Thing (IoT) environment. The performance evaluation and security analysis demonstrate successful user privacy and integrity with lower computational resources and communication overhead. Ikram Ud Din, Ahmad S. Al-Mogren, Mohsen Guizani, Mansour Abdulaziz Al Zuair |
IEEE Trans. Ind. Informatics | 5 |
| 2020 | Real-Time Mobile-Based Electrocardiogram System for Remote Monitoring of Patients with Cardiac ArrhythmiasabstractIn this study, we propose an electrocardiogram (ECG) system for the simultaneous and remote monitoring of multiple heart patients. It consists of three main components: patient, sever, and monitoring units. The patient unit uses a wearable miniature sensor that continuously measures ECG signals and sends them to a smart mobile phone via a Bluetooth connection. In the mobile device, the ECG signals can be stored, displayed on screen, and automatically transmitted to a distant server unit over the internet; the server stores ECG data from several patients. Health care stakeholders use a monitoring unit to retrieve the ECG signals of multiple patients at any time from the server for display and real-time automatic analysis. The analysis includes segmentation of the ECG signal into separate heartbeats followed by arrhythmia detection and classification. When compared to existing real-time ECG systems, where the detection of abnormalities is usually performed using simple rules, the proposed system implements a real-time classification module that is based on a support vector machine (SVM) classifier. Extensive experimental results on ECG data obtained from a TechPatientTMsimulator, a real person, and 20 records from the MIT arrhythmia database are reported and discussed. Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Haikel Salem Alhichri, Nassim Ammour, Naif Alajlan, Mansour Abdulaziz Al Zuair |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2020 | PCAPooL: unsupervised feature learning for face recognition using PCA, LBP, and pyramid pooling
Amani A. Alahmadi, Muhammad Hussain 0001, Hatim A. Aboalsamh, Mansour Abdulaziz Al Zuair |
Pattern Anal. Appl. | 4 |
| 2018 | Asymmetric Adaptation of Deep Features for Cross-Domain Classification in Remote Sensing ImageryabstractIn this letter, we introduce an asymmetric adaptation neural network (AANN) method for cross-domain classification in remote sensing images. Before the adaptation process, we feed the features obtained from a pretrained convolutional neural network to a denoising autoencoder (DAE) to perform dimensionality reduction. Then the first hidden layer of AANN (placed on the top of DAE) maps the labeled source data to the target space, while the subsequent layers control the separation between the available land-cover classes. To learn its weights, the network minimizes an objective function composed of two losses related to the distance between the source and target data distributions and class separation. The results of experiments conducted on six scenarios built from three benchmark scene remote sensing data sets (i.e., Merced, KSA, and AID data sets) are reported and discussed. Nassim Ammour, Laila Bashmal, Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Mansour Abdulaziz Al Zuair |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2017 | Scalable regular pattern mining in evolving body sensor data
Syed Khairuzzaman Tanbeer, Mohammad Mehedi Hassan, Ahmad S. Al-Mogren, Mansour Abdulaziz Al Zuair, Byeong-Soo Jeong |
Future Gener. Comput. Syst. | 4 |
| 2017 | Domain Adaptation Network for Cross-Scene ClassificationabstractIn this paper, we present a domain adaptation network to deal with classification scenarios subjected to the data shift problem (i.e., labeled and unlabeled images acquired with different sensors and over completely different geographical areas). We rely on the power of pretrained convolutional neural networks (CNNs) to generate an initial feature representation of the labeled and unlabeled images under analysis, referred as source and target domains, respectively. Then we feed the resulting features to an extra network placed on the top of the pretrained CNN for further learning. During the fine-tuning phase, we learn the weights of this network by jointly minimizing three regularization terms, which are: 1) the cross-entropy error on the labeled source data; 2) the maximum mean discrepancy between the source and target data distributions; and 3) the geometrical structure of the target data. Furthermore, to obtain robust hidden representations we propose a mini-batch gradient-based optimization method with a dynamic sample size for the local alignment of the source and target distributions. To validate the method, in the experiments we use the University of California Merced data set and a new multisensor data set acquired over several regions of the Kingdom of Saudi Arabia. The experiments show that: 1) pretrained CNNs offer an interesting solution for image classification compared to state-of-the-art methods; 2) their performances can be degraded when dealing with data sets subjected to the data shift problem; and 3) how the proposed approach represents a promising solution for effectively handling this issue. Essam Othman, Yakoub Bazi, Farid Melgani, Haikel Salem Alhichri, Naif Alajlan, Mansour Abdulaziz Al Zuair |
IEEE Trans. Geosci. Remote. Sens. | 6 |