VLDB 2026 Research / reviewers in the wild / expert
Ahmed Gomaa
dblp:10/1146
· DBLP profile ↗
10ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Residual Channel-attention (RCA) network for remote sensing image scene classificationabstractAbstract High-resolution remote sensing (HRRS) image scene classification has gained increasing importance in recent years, with convolutional neural networks (CNNs) showing particular promise due to their proficiency in extracting spatial features. However, traditional CNNs face significant limitations. Specifically, they struggle to capture complex semantic relationships between objects at varying scales, and they lack the ability to effectively capture long-distance dependencies between features. This limitation is especially problematic in HRRS images, where spatial relationships and semantic content are deeply intertwined. Additionally, traditional CNNs are limited in handling substantial intra-class variation and inter-class similarity, which are common in remote sensing images. To overcome these challenges, we introduce a novel Residual Channel-attention (RCA) network for scene classification. The RCA network introduces a lightweight residual structure to better capture multi-scale spatial features and incorporates a channel attention mechanism that selectively emphasizes relevant feature channels while suppressing irrelevant ones. To further refine the focus on critical image features, we integrate a squeeze-and-excitation (SE) mechanism as a self-attention component, which helps the network prioritize the most informative features and ignore background noise. We evaluated the RCA network on three public datasets: RSSCN7, PatternNet, and EuroSAT, achieving classification accuracies of 97%, 99%, and 96%, respectively. The results demonstrate that superior of the RCA network compared to state-of-the-art strategies in remote sensing image classification. Furthermore, visualization using the Grad-CAM++ algorithm highlights the effectiveness of our channel attention mechanism and underscores the RCA network’s robust feature representation capabilities. Ahmed Gomaa, Omar M. Saad |
Multim. Tools Appl. | 1 |
| 2025 | A novel intelligent healthcare system based on efficient deep learning models for brain stroke predictionabstractAbstract Smart hospitals should equip patients with wearable devices capable of monitoring and predicting brain strokes using deep learning (DL) techniques. The latest advancements in the accuracy of DL have the potential to make a substantial impact in addressing the issues associated with predicting brain strokes. However, further enhancements are required to achieve even higher levels of precision in DL methods. This research proposes an innovative intelligent healthcare system (IHS) that utilizes a real-time web application. The IHS is specifically engineered to monitor and forecast cerebral stroke occurrences in individuals. Also, this work presents three DL models: long short-term memory (LSTM), gated recurrent unit (GRU), and bidirectional long short-term memory (BiLSTM), which are used to predict brain stroke. The three models are trained using a dataset collected from 4,981 individuals. This dataset consists of two categories—normal and stroke, and it comprises eleven distinct characteristics: gender, age, heart disease (HD), hypertension, marital status (MS), type of residence (RT), average glucose level (AGL), type of work (TW), body mass index (BMI), smoking status (SS), and stroke. The DL models suggested in this study are constructed using the Keras library, employing a hyperparameter tuning technique to maximize accuracy. The effectiveness of the three models is evaluated using precision-recall (PR) curves and a normalized error matrix (NEM). The results indicate that the BiLSTM model outperforms both the GRU and LSTM models in terms of efficiency for predicting brain stroke. The BiLSTM achieves the most efficiency, with a testing accuracy (TA) of 100%, followed by the LSTM with a TA of 99.90%. However, the GRU exhibits the lowest TA of 99.80%. The testing loss (TL) rates for the LSTM, GRU, and BiLSTM models are 0.0075, 0.022, and 0.0001, respectively. Additionally, the BiLSTM model achieves sensitivity, accuracy, F1-score, and area under the PR curves of 100%. The proposed DL models with IHS can assist physicians in efficiently and precisely diagnosing persons with brain strokes, enabling them to make prompt and accurate decisions. Saeed Mohsen, Ahmed Gomaa |
Multim. Tools Appl. | 2 |
| 2024 | Deep Learning for Cancer Prognosis Prediction Using Portrait Photos by StyleGAN Embedding
Amr Hagag, Ahmed Gomaa, Dominik Kornek, Andreas K. Maier, Rainer Fietkau, Christoph Bert, Yixing Huang, Florian Putz |
MICCAI (5) | 2 |
| 2023 | Detection of Earthquake-Induced Building Damages Using Remote Sensing Data and Deep Learning: A Case Study of Mashiki Town, JapanabstractNatural disasters cause extensive economic losses every year. Rapid detection of earthquake-induced building damages is crucial for disaster response. Remote sensing (RS) has been widely used to assess the impacts of natural disasters i.e. earthquakes and its implications on building damages. Deep Learning (DL) techniques have become increasingly popular for detecting building damages from RS data and have achieved significant success in detecting disaster implications. This paper examines the ability of DL to detect building damages caused by Kumamoto earthquake in Mashiki town, Japan using RS data. The findings indicate that the newly trained model demonstrated effective performance in discriminating between different levels of building damages, including no damage, damage, and collapse.1 Muhammad Salem, Ahmed Gomaa, Naoki Tsurusaki |
IGARSS | 2 |
| 2022 | Supervised Contrastive Learning for Robust and Efficient Multi-modal Emotion and Sentiment AnalysisabstractExpression of human emotion and sentiment are often multi-modal consisting use of spoken speech, vision, and text. Combining multiple modalities allows learning-based models to benefit with the complementary information present across modalities to produce more accurate predictions. One of the bigger challenges in multi-modal affective computing is performance consistency in non-ideal scenarios. Most benchmarks fail to generalize in non-ideal scenarios where one of the modalities is missing or highly corrupted due to occlusion, sensor errors, or change of orientation. Consequently, various modality fusion approaches were proposed. However, most of these fusion approaches assume that each modality is equally useful. To address the challenge of performance consistency, in this work we propose to use supervised contrastive learning (SCL). We demonstrate through various experiments and comparison with state-of-the-art (SOTA) methods that the model robustness against corrupted and missing modalities improves when trained with SCL. Next, we use the Perceiver architecture [1] in order to efficiently combine the representations of different modalities. Its iterative attention mechanism allows to create a reduced latent representation in an efficient manner. We observe that it can accommodate a wide range of modality combinations, allowing for robust information fusion. Our approach allows reduction of model complexity and efficient fusion of different modalities, while maintaining the performance consistency and model robustness. We conduct ablation experiments to study the effect of each contribution in different scenarios, and we show that the proposed methods outperform the state-of-art, while simultaneously being robust to corrupted modalities. Our method also outperforms its counterparts and SOTA while using less numerical complexity (inference times and compute operations). Ahmed Gomaa, Andreas K. Maier, Ronak Kosti |
ICPR | 1 |
| 2022 | Faster CNN-based vehicle detection and counting strategy for fixed camera scenesabstractAbstract Automatic detection and counting of vehicles in a video is a challenging task and has become a key application area of traffic monitoring and management. In this paper, an efficient real-time approach for the detection and counting of moving vehicles is presented based on YOLOv2 and features point motion analysis. The work is based on synchronous vehicle features detection and tracking to achieve accurate counting results. The proposed strategy works in two phases; the first one is vehicle detection and the second is the counting of moving vehicles. Different convolutional neural networks including pixel by pixel classification networks and regression networks are investigated to improve the detection and counting decisions. For initial object detection, we have utilized state-of-the-art faster deep learning object detection algorithm YOLOv2 before refining them using K-means clustering and KLT tracker. Then an efficient approach is introduced using temporal information of the detection and tracking feature points between the framesets to assign each vehicle label with their corresponding trajectories and truly counted it. Experimental results on twelve challenging videos have shown that the proposed scheme generally outperforms state-of-the-art strategies. Moreover, the proposed approach using YOLOv2 increases the average time performance for the twelve tested sequences by 93.4% and 98.9% from 1.24 frames per second achieved using Faster Region-based Convolutional Neural Network (F R-CNN ) and 0.19 frames per second achieved using the background subtraction based CNN approach (BS-CNN ), respectively to 18.7 frames per second. Ahmed Gomaa, Tsubasa Minematsu, Moataz M. Abdelwahab, Mohammed Abo-Zahhad 0001, Rin-Ichiro Taniguchi |
Multim. Tools Appl. | 1 |
| 2022 | A wide axial-ratio beamwidth circularly-polarized oval patch antenna with sunlight-shaped slots for gnss and wimax applicationsabstractAbstract This paper proposes a quadruple band stacked oval patch antenna with sunlight-shaped slots supporting L1/L2/L5 GNSS bands and the 2.3 Ghz WiMAX band. The antenna produces right-hand circular polarization waves with wide axial-ratio beamwidth of 223/216 $$^{\circ }$$ ∘ and 231/203 $$^{\circ }$$ ∘ at two orthogonal cutplanes at L5 and L2 GNSS bands, respectively. Firstly, the resonant modes $$TM_{110}$$ T M 110 and $$TM_{210}$$ T M 210 are excited inside a single layer oval patch antenna, where resonance frequencies are calculated using Mathieu functions. Meanwhile, it is shown that another version of the mode $$TM_{110}$$ T M 110 with similar distribution but orthogonal direction is excitable inside the same oval patch. Then, a second stacked oval patch layer is added, which splits the resonance frequency of each of the modes $$TM_{110}$$ T M 110 and $$TM_{210}$$ T M 210 into two different values. Depending on the probe feed position and the separation between the two layers, the phase shifts between modes versions in the upper and the lower layers change. Thus, by fine-tuning the probe feed position and the separation between layers, spatially-orthogonal with quadrature-phase-shift versions of the mode $$TM_{110}$$ T M 110 are obtained, producing a circularly polarized waves at L2 and L5 bands. Furthermore, sunlight shaped slots are etched into the upper and lower layer patches to fine tune the phase shifts between different modes versions, which enhances the overall axial-ratio beamwidth. Despite the simplicity of the overall structure and the feeding mechanism utilized in the proposed design, wide axial-ratio beamwidths are obtained, as compared to previous works. The proposed antenna shows low reflection coefficient values at 1.14–1.29 GHz (L2/L5), 1.45–1.6 GHz (L1), and 2.26–2.4 GHz (WiMAX). The antenna gains are 5.9, 5.6, 6, and 6.5 dBi/dBic at L5, L2, L1, and WiMAX bands, respectively. The half-power beamwidths are 99/96 $$^{\circ }$$ ∘ , 102/96 $$^{\circ }$$ ∘ , 112/85 $$^{\circ }$$ ∘ , and 65/48 $$^{\circ }$$ ∘ at two orthogonal cutplanes at L5, L2, L1, and WiMAX bands, respectively. Ahmad Abdalrazik, Ahmed Gomaa, Ahmed A. Kishk |
Wirel. Networks | 2 |
| 2020 | Efficient vehicle detection and tracking strategy in aerial videos by employing morphological operations and feature points motion analysis
Ahmed Gomaa, Moataz M. Abdelwahab, Mohammed Abo-Zahhad 0001 |
Multim. Tools Appl. | 1 |
| 2005 | Adapting spatial constraints of composite multimedia objects to achieve universal accessabstractA composite multimedia object (cmo) is comprised of different media components such as text, video, audio and image, with a variety of constraints that must be adhered to. The constraints are 1) rendering constraints that comprise the temporal and spatial constraints between different components, and 2) behavioral constraints that include the security and fidelity constraints on each component. Different users have different 3Cs, which are: capabilities (e.g., monitor size), characteristics (e.g., age) and credentials (e.g., subscription to service). The focus of this paper is on addressing the problems of (1) specifying a consistent cmo that "automatically" adapts its spatial constraints to different user's devices. (2) Identifying the conflicts that might occur between the temporal and spatial constraints when having different monitor resolution that displays the cmo by means of reachability analysis of colored time Petri net (3) Resolving the identified conflicts automatically to render a cmo that is error-free when rendered at different user devices. Ahmed Gomaa, Nabil R. Adam, Vijayalakshmi Atluri |
IPCCC | 1 |
| 2005 | Color Time Petri Net for Interactive Adaptive Multimedia ObjectsabstractA composite multimedia object (cmo) is comprised of different media components such as text, video, audio and image, with a variety of constraints that must be adhered to. The constraints are 1) rendering relationships that comprise the temporal and spatial constraints between different components, 2) behavioral requirements that include the security and fidelity constraints on each component and, 3) user interactions on a set of related media components. Different users have different capabilities (e.g. age), characteristics (e.g. monitor size) and credentials (e.g. subscription to service). Our objective is to author an interactive adaptive cmo that renders itself correctly to different users. Therefore, it is important to guarantee the consistency of the cmo specifications in all possible scenarios. In this paper, we include the user interaction with temporal and spatio-temporal behavior in the specification of the adaptive cmo. We then check the consistency of user interaction specifications by transforming the specifications into a color time Petri net model. We perform a reachability analysis on the Petri net to identify inconsistencies. We then resolve the identified inconsistencies to have a consistent Petri net. A consistent Petri net presents an error-free interactive cmo that can adapt to different users, by guaranteeing that link user interactions are reachable for all eligible users. Ahmed Gomaa, Nabil R. Adam, Vijayalakshmi Atluri |
MMM | 1 |