Ibrahim Almakky

dblp:220/1470 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0008-8802-7107ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MAFM3: Modular Adaptation of Foundation Models for Multi-Modal Medical AI
abstract
Foundational models are trained on extensive datasets to capture the general trends of a domain. However, in medical imaging, the scarcity of data makes pre-training for every domain, modality, or task challenging. Instead of building separate models, we propose MAFM3(Modular Adaptation of Foundation Models for Multi-Modal Medical AI), a framework that enables a single foundation model to expand into diverse domains, tasks, and modalities through lightweight modular components. These components serve as specialized skill sets that allow the system to flexibly activate the appropriate capability at the inference time, depending on the input type or clinical objective. Unlike conventional adaptation methods that treat each new task or modality in isolation, MAFM3provides a unified and expandable framework for efficient multitask and multimodality adaptation. Empirically, we validate our approach by adapting a chest CT foundation model initially trained for classification into prognosis and segmentation modules. Our results show improved performance on both tasks. Furthermore, by incorporating PET scans, MAFM3achieved an improvement in the Dice score 5% compared to the respective baselines. These findings establish that foundation models, when equipped with modular components, are not inherently constrained to their initial training scope but can evolve into multitask, multimodality systems for medical imaging. The code implementation of this work will be made available upon acceptance. The code implementation of this work can be found at Code
Qazi Mohammad Areeb, Munachiso S. Nwadike, Ibrahim Almakky, Mohammad Yaqub, Numan Saeed
WACV3
2025 MedNNS: Supernet-Based Medical Task-Adaptive Neural Network Search
Lotfi Abdelkrim Mecharbat, Ibrahim Almakky, Martin Takác 0001, Mohammad Yaqub
MICCAI (6)2
2024 FissionFusion: Fast Geometric Generation and Hierarchical Souping for Medical Image Analysis
Santosh Sanjeev, Nuren Zhaksylyk, Ibrahim Almakky, Anees Ur Rehman Hashmi, Qazi Mohammad Areeb, Mohammad Yaqub
MICCAI (12)3
2023 FedSIS: Federated Split Learning with Intermediate Representation Sampling for Privacy-preserving Generalized Face Presentation Attack Detection
abstract
Lack of generalization to unseen domains/attacks is the Achilles heel of most face presentation attack detection (FacePAD) algorithms. Existing attempts to enhance the generalizability of FacePAD solutions assume that data from multiple source domains are available with a single entity to enable centralized training. In practice, data from different source domains may be collected by diverse entities, who are often unable to share their data due to legal and privacy constraints. While collaborative learning paradigms such as federated learning (FL) can overcome this problem, standard FL methods are ill-suited for domain generalization because they struggle to surmount the twin challenges of handling non-iid client data distributions during training and generalizing to unseen domains during inference. In this work, a novel framework called Federated Split learning with Intermediate representation Sampling (FedSIS) is introduced for privacy-preserving domain generalization. In FedSIS, a hybrid Vision Transformer (ViT) architecture is learned using a combination of FL and split learning to achieve robustness against statistical heterogeneity in the client data distributions without any sharing of raw data (thereby preserving privacy). To further improve generalization to unseen domains, a novel feature augmentation strategy called intermediate representation sampling is employed, and discriminative information from intermediate blocks of a ViT is distilled using a shared adapter network. The FedSIS approach has been evaluated on two well-known benchmarks for cross-domain FacePAD to demonstrate that it is possible to achieve state-of-the-art generalization performance without data sharing. Code: https://github.com/Naiftt/FedSIS
Naif Alkhunaizi, Koushik Srivatsan, Faris Almalik, Ibrahim Almakky, Karthik Nandakumar
IJCB4
2023 Arabic Dysarthric Speech Recognition Using Adversarial and Signal-Based Augmentation
Massa Baali, Ibrahim Almakky, Shady Shehata, Fakhri Karray
INTERSPEECH2
2023 FeSViBS: Federated Split Learning of Vision Transformer with Block Sampling
Faris Almalik, Naif Alkhunaizi, Ibrahim Almakky, Karthik Nandakumar
MICCAI (2)3
2019 Detection of Diabetic Retinopathy and Maculopathy in Eye Fundus Images Using Deep Learning and Image Augmentation
Sarni Suhaila Rahim, Vasile Palade, Ibrahim Almakky, Andreas Holzinger
CD-MAKE3
2019 Deep Convolutional Neural Networks for Text Localisation in Figures From Biomedical Literature
abstract
Text contained within figures is an important source of information in biomedical literature. Despite this, end-to-end text extraction from biomedical figures remains a challenging task. This paper presents a novel approach to address the founding block of this task, text detection, not only from biomedical figures but also from images in general. Particularly, the paper proposes an approach that simplifies the text detection problem into a reconstruction problem using a deep convolutional neural network. Designed to overcome the specific challenges of text detection from biomedical figures, our proposed model reports promising results on the DETEXT dataset.
Ibrahim Almakky, Vasile Palade, Ariel Ruiz-Garcia
IJCNN1
2019 Deep Q-Learning for Illumination and Rotation Invariant Face Detection
abstract
The domain of automatic face detection is a challenging problem that has made great progress in the last two decades. Much of this progress is due to the continuous advancements in deep learning and computer vision. Nonetheless, contemporary state-of-the-art face detection models rely on exhaustive search or are unable to deal with changes in the data distribution. In this work, we propose a novel Deep Reinforcement Learning (DRL) approach for face detection on data with nonuniform conditions. More specifically, we address illumination and rotation invariance in face detection. Firstly, we train a Stacked Convolutional Autoencoder (SCAE) in a greedy layer-wise unsupervised fashion for illumination invariant feature extraction. We then train a deep Q-network on the illumination invariant features produced by the SCAE model, to learn an action-value policy that allows an agent to place a bounding box around a face. The proposed approach achieves state-of-the-art recognition rates on images with varying degrees of illumination and images that contain faces with some degree of rotation.
Ariel Ruiz-Garcia, Vasile Palade, Ibrahim Almakky, Mark Elshaw
IJCNN3
2018 Deep Learning for Illumination Invariant Facial Expression Recognition
abstract
In this work we propose a novel method to address illumination invariance for facial expression recognition. We propose a Deep Convolutional Network (CNN) pre-trained as a Deep Stacked Convolutional Autoencoder (SCAE) in a greedy layer-wise unsupervised fashion. The SCAE model learns to encode facial expression images and produce a feature vector with relatively similar illumination, regardless of the luminance level of the input image. Moreover, we propose fine-tuning the stacked shallow autoencoders after each one of these is trained greedily, rather than just at the end, and show that this approach significantly improves the set of illumination invariant features learnt by the SCAE. Finally, we propose the use of a variant rectifier linear unit transfer function that helps the SCAE model reduce or increase the illumination of images with high or low luminance, and show that the lower and upper bounds greatly influence classification performance. The method proposed provides an increase in classification accuracy of 4% on the KDEF dataset and 8% on the CK+ dataset.
Ariel Ruiz-Garcia, Vasile Palade, Mark Elshaw, Ibrahim Almakky
IJCNN4