EDBT 2026 Demo / reviewers in the wild / expert
Marwan Torki
dblp:98/8617
· DBLP profile ↗
53ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0002-6149-1718ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 20 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 7 since 2021Computer networks · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Differentially Private Datastore Generation for Retrieval-Augmented Inference
Abdelrahman Abouelenein, Marwan Torki |
ICPR (14) | 2 |
| 2026 | An Efficient Multi-Rater Setup Towards Personalized and Diversified Medical Image SegmentationabstractMulti-rater medical image segmentation addresses annotation ambiguities but typically requires costly multiple expert annotations per scan. We propose P-Diverse, a novel two-stage framework that minimizes the annotation needs while achieving state-of-the-art performance. Stage-I trains a modified nnU-Net with expert-specific embeddings throughout the network stages, generating personalized segmentations using as low as one annotation per scan. Stage-II freezes that network to synthesize the missing an-notations and trains a diversification model that captures multi-rater variability using the available and synthesized annotations. We evaluated on the public NPC dataset and QUBIQ2021 dataset (where the current SOTA method fails), P-Diverse establishes new SOTA performance using synthetic annotations on the diversification stage, significantly reducing clinical annotation burdens. Code: https://github.com/SajedHassan/P-Diverse. Sajed Almorsy, Ayman Khalafallah, Marwan Torki |
WACV | 3 |
| 2026 | Zero-LEAD: Source-Free Universal Domain Adaptation for Abdominal Multi-Organ SegmentationabstractCross-modality medical image segmentation is critical for diagnosis and treatment planning, yet domain shifts and source data restrictions pose significant challenges. This paper introduces Zero-LEAD, marking the first unified framework designed for source-free universal domain adaptation (SF-UniDA) in segmentation, addressing all four UDA scenarios, closed-set, partial-set, open-set, and universal-set, without access to source data. Zero-LEAD integrates (1) Label-Efficient Adaptive Decomposition (LEAD) to decompose features into source-known and source-unknown components, and (2) a zero-shot segmentation module leveraging anatomical priors and semantic attributes to segment novel target classes. Extensive experiments across four datasets, Synapse, CHAOS, BTCV, and FLARE22, demonstrate strong performance across all adaptation settings. Zero-LEAD achieves 0.9159 Dice in closed-set (Synapse→CHAOS), 0.8721 Dice in partial-set (BTCV→CHAOS), 0.7801 Dice in open-set (Synapse→BTCV), and 0.7716/0.6866 Dice in universal-set (BTCV↔FLARE22), significantly outperforming state-of-the-art baselines. Ablation studies confirm the complementary contributions of LEAD and zero-shot modules, and qualitative analysis highlights improved boundary precision and robustness under both domain and label shifts. Ahmed El-Sayed, Marwan Torki |
WACV | 2 |
| 2025 | Unlocking Additional Learning Capabilities of Whisper for Arabic Language Via Instruction Fine-tuningabstractEncoder-decoder models based on transformers showed remarkable results in different natural language processing (NLP) generative tasks, including automatic speech recognition (ASR). In this paper, we focus on one of the most popular encoder-decoder models, Whisper, to increase its capabilities in Arabic language using instruction fine-tuning. We propose an Arabic dataset featuring diverse tasks with different prompts directly applied to acoustic features. The proposed dataset is synthesized from popular Arabic speech datasets covering various tasks such as gender identification, audio environment identification, and audio-text alignment. Moreover, we apply instruction fine-tuning to Whisper on the synthesized dataset, and our results demonstrate that Whisper can acquire new capabilities with minimal or no significant increase in word error rate (WER) and character error rate (CER). Specifically, the large Whisper model achieves an F1-score of 0.98 for gender identification, 0.76 for audio environment identification, and an Intersection over Union (IoU) of 0.86 for audio-text alignment. These enhancements are achieved at the cost of a slight increase in WER by 0.0085 and in CER by 0.0038. Bassam Mattar, Marwan Torki, Ayman Khalafallah |
AICCSA | 2 |
| 2025 | HTEB: Hybrid Transformer Enhancement Block for Robust Image Flare SuppressionabstractImage flare artifacts, caused by internal light reflections and scattering in camera lenses, degrade image quality by introducing unwanted bright spots, ghosting, and haze. These artifacts distort scene details and negatively impact computer vision applications. To address this, we propose a deep learning model that effectively removes flares from single images while preserving critical visual content. Our method first estimates and refines a depth map of the input image using a Dense Vision Transformer (DPT). The refined depth map is combined with the original RGB image to form a 4 -channel input, which is processed by a U-shaped network. This network uses encoderdecoder blocks to progressively suppress flares, with a novel Hybrid Transformer Enhancement Block (HTEB) at its core to model global relationships and local details. Experiments on the Flare7K++ benchmark and real-world images demonstrate that our approach outperforms existing methods in flare removal accuracy and detail preservation, proving its robustness for practical use. This work underscores the potential of hybrid attention-frequency mechanisms in transformer-based restoration architectures. Mohamed Mostafa, Hicham G. Elmongui, Marwan Torki |
AICCSA | 3 |
| 2025 | Eswindnet: Image Demoiréing Using Multiscale Swin Transformer LayersabstractCapturing electronic screens with digital cameras introduces high-frequency artifacts, known as moiré patterns, degrading overall image quality and colors. This work proposes ESwinDNet, an image demoiréing model that combines an encoder-decoder architecture with multiscale Swin Transformer layers. These layers efficiently compute pixel-level attention, a crucial aspect for low-level vision tasks such as image demoiréing. The proposed ESwinDNet model achieves comparable results to the large variant of the baseline model ESDNet-L on the UHDM dataset, demonstrating its capabilities in the removal of moiré patterns in 4K images, with nearly half the number of parameters and floating point operations, yielding faster training and inference time. The code is available on Github. Karim Alaa, Marwan Torki |
ICIP | 2 |
| 2025 | Cross-Modality Abdominal Multi-Organ Segmentation via Source-Free Unsupervised Domain AdaptationabstractCross-modality organ segmentation is crucial in medical imaging, enabling the use of annotated data from one modality (e.g., CT) to segment structures in another (e.g., MRI). However, differences between modalities and limited source data often hinder standard supervised or unsupervised domain adaptation methods. A novel Source-Free Unsupervised Domain Adaptation (SF-UDA) framework is proposed for abdominal multi-organ segmentation that does not require access to the original source labeled data during adaptation. Our two-stage approach first generates pseudo-labels with a pre-trained source model and reduces domain shifts using a contrastive cycle-based style-transfer module, while simultaneously training multiple networks with segmentation and entropy minimization losses to enhance feature alignment. In the second stage, segmentation results are refined through cycle-based learning and mask refinement to ensure spatial consistency and improve boundary accuracy. Experiments show our framework surpasses state-of-the-art SFUDA methods on abdominal multi-organ datasets, achieving a mean Dice coefficient of 0.9182. Ahmed El-Sayed, Marwan Torki |
ICIP | 2 |
| 2025 | Enhancing 3D Semi-Supervised Teeth Segmentation with Voxel Entropy-Guided Self-LearningabstractTeeth segmentation is crucial for the analysis of maxillofacial diseases, with Cone Beam Computed Tomography (CBCT) scans offering superior 3D imaging compared to conventional Computed Tomography (CT). We propose a semi-supervised framework based on a multi-iterative self-learning approach, introducing a novel ranking technique for pseudo-label selection and a post-processing algorithm to eliminate false positives and ensure mask consistency. Pseudo-label ranking is guided by average voxel-level entropy, which quantifies uncertainty between foreground and background classes. High-confidence labels are then selected to refine training in subsequent iterations. Our post-processing algorithm enhances segmentation fidelity by identifying connected components, retaining the largest within a defined radius, and filtering out extraneous voxels. Additionally, expert-refined pseudo-labels are incorporated to assess their impact on performance. To further validate our approach, we established a baseline model and conducted a comprehensive evaluation on both the private test set of the STS-3D Challenge and our self-collected out-of-set data. The framework employs nnUNetV2 with the ResEnc M preset, optimizing feature extraction and dataset adaptation. Our approach achieves a Dice score of 0.8244 and an IoU of 0.851 on the private test set of the STS-3D Challenge, outperforming previous published methods on the same benchmark and demonstrating the effectiveness of our semi-supervised teeth segmentation strategy. Elbadry Elbadry, Mahmoud Gamal, Marwa Baraka, Marwan Torki |
ICMLA | 4 |
| 2025 | DiffuPT: Class Imbalance Mitigation for Glaucoma Detection via Diffusion Based Generation and Model PretrainingabstractGlaucoma is a progressive optic neuropathy character-ized by structural damage to the optic nerve head andfunctinoal changes in the visual field. Detecting glaucoma early is crucial to preventing loss of eyesight. However, med-ical datasets often suffer from class imbalances, making detection more difficult for deep-learning algorithms. We use a generative-based framework to enhance glaucoma di-agnosis, specifically addressing class imbalance through synthetic data generation. In addition, we collected the largest national dataset for glaucoma detection to support our study. The imbalance between normal and glaucoma-tous cases leads to performance degradation of classifier models. We created a more robust classifier training process by combining our proposed framework leveraging diffusion models with a pretraining approach. This training process results in a better-performing classifier. The proposed approach shows promising results in improving the harmonic mean “sensitivity and specificity” and AU C for the roc for the glaucoma classifier. We report an improvement in the harmonic mean metric from 89.09% to 92.59% on the test set of our Egyptian dataset. We examine our method against other methods to overcome imbalance through extensive experiments. We report similar improvements on the AIROGS dataset. This study highlights that diffusion-based generatino can be important in tackling class imbalances in medi-cal datasets to improve diagnostic performance. Youssof Nawar, Nouran Soliman, Moustafa Wassel, Mohamed ElHabebe, Noha Adly, Marwan Torki, Ahmed Elmassry, Islam Ahmed |
WACV | 6 |
| 2024 | Utilizing Multi-Step Loss for Single Image Reflection RemovalabstractImage reflection removal is crucial for restoring image quality. Distorted images can negatively impact tasks like object detection and image segmentation. In this paper, we present a novel approach for image reflection removal using a single image. Instead of focusing on model architecture, we introduce a new training technique that can be generalized to image-to-image problems, with input and output being similar in nature. This technique is embodied in our multi-step loss mechanism, which has proven effective in the reflection removal task. Additionally, we address the scarcity of reflection removal training data by synthesizing a high-quality, non-linear synthetic dataset called RefGAN using Pix2Pix GAN. This dataset significantly enhances the model's ability to learn better patterns for reflection removal. We also utilize a ranged depth map, extracted from the depth estimation of the ambient image, as an auxiliary feature, leveraging its property of lacking depth estimations for reflections. Our approach demonstrates superior performance on the SIR2benchmark and other real-world datasets, proving its effectiveness by out-performing other state-of-the-art models. Abdelrahman Elnenaey, Marwan Torki |
AICCSA | 2 |
| 2024 | Flare-Free Vision: Empowering Uformer with Depth InsightsabstractImage flare is a common problem that occurs when a camera lens is pointed at a strong light source. It can manifest as ghosting, blooming, or other artifacts that can degrade the image quality. We propose a novel deep learning approach for flare removal that uses a combination of depth estimation and image restoration. We use a Dense Vision Transformer to estimate the depth of the scene. This depth map is then concatenated to the input image, which is then fed into a Uformer, a general U-shaped transformer for image restoration. Our proposed method demonstrates state-of-the-art performance on the Flare7K++ test dataset, demonstrating its effectiveness in removing flare artifacts from images. Our approach also demonstrates robustness and generalization to real-world images with various types of flare. We believe that our work opens up new possibilities for using depth information for image restoration. The code is available on GitHub Yousef Kotp, Marwan Torki |
ICASSP | 2 |
| 2024 | Mask-Aware Transformer for Crowd Counting
Sarah Jad, Marwan Torki, Ayman Khalafallah |
ICPR (18) | 2 |
| 2024 | Temporal Divide-and-Conquer Anomaly Actions Localization in Semi-supervised Videos with Hierarchical Transformer
Nada Osman, Marwan Torki |
ICPR (15) | 2 |
| 2024 | C2F-CHART: A Curriculum Learning Approach to Chart Classification
Nour Shaheen, Tamer Elsharnouby, Marwan Torki |
ICPR (1) | 3 |
| 2024 | AraEGI: Arabic Gender Identification Using Transformers for Egyptian DialectabstractWith the rise of social media, accurately identifying genders in the text has become crucial for various domains. It is used in a wide range of academic/commercial applications such as personalization of user experiences, analyzing and interpreting text more accurately in Natural Language Understanding (NLU), and identifying and mitigating biases in datasets and models. The importance of this task is reflected in the extensive research and contribution to it, specifically in Arabic text. Existing work addresses this problem as a sub-task of the Author Profiling task, by detecting the author’s gender using multiple texts of the same author. Although current approaches achieve satisfactory results, they still fall behind in the detection of genders based on a single short sentence. In this paper, we provide a newly tailored dataset, AraEGI, to tackle Gender Identification, capable of reaching robust results on the identification of gender in tweets. Our dataset consists of 3 sets, with tweet-level labels showcasing the genders of both the speaker and the listener. We experimented on our dataset using the latest Arabic transformer models to provide a benchmark for future research. Ahmed Abdelmaguid, Marwan Torki, Ayman Khalafallah |
ISCC | 2 |
| 2024 | Automatic Mandibular Semantic Segmentation of Teeth Pulp Cavity and Root Canals, and Inferior Alveolar Nerve on Pulpy3D DatasetabstractAccurate segmentation of the pulp cavity, root canals, and inferior alveolar nerve (IAN) in dental imaging is essential for effective orthodontic interventions. Despite the availability of numerous Cone Beam Computed Tomography (CBCT) scans annotated for individual dental-anatomical structures, there is a lack of a comprehensive dataset covering all necessary parts. As a result, existing deep learning models have encountered challenges due to the scarcity of comprehensive datasets encompassing all relevant anatomical structures. We present our novel Pulpy3D dataset, specifically curated to address dental-anatomical structures’ segmentation and identification needs. Additionally, we noticed that many current deep learning methods in dental imaging prefer 2D segmentation, missing out on the benefits of 3D segmentation. Our study suggests a UNet-based approach capable of segmenting dental structures using 3D volume segmentation, providing a better understanding of spatial relationships and more precise dental anatomy representation. Pulpy3D contributed in creating the seeding model from 150 scans, which helped complete the remainder of the dataset. Other modifications in the architecture, such as using separate networks, one semantic network, and a multi-task network, were highlighted in the model description to show how versatile the Pulpy3D dataset is and how different models, architectures, and tasks can run on the dataset. Additionally, we stress the lack of attention to pulp segmentation tasks in existing studies, underlining the need for specialized methods in this area. The code and Pulpy3D links can be found at https://github.com/mahmoudgamal0/Pulpy3D . Mahmoud Gamal, Marwa Baraka, Marwan Torki |
MICCAI (8) | 3 |
| 2024 | DR10K: Transfer Learning Using Weak Labels for Grading Diabetic Retinopathy on DR10K DatasetabstractIn this paper, we contrast the usage of two deep-learning approaches for the automatic grading of diabetic retinopathy (DR) and diabetic macular edema (DME) in retinal fundus photographs using a relatively small novel dataset. We developed a telemedicine system to collect and humanly grade 11,109 diabetic patients. The certified graders annotated the level of DR as well as the existence of a referable DME in the macula-centered fundus images only. We use EfficientNet to build an AI-based model for both problems. To examine the transfer learning validity, the model was trained on an external dataset (EyePacs) and then finetuned on the egyptian data for the DR and DME grading problems. Firstly, we use the macula-centered images only in fine-tuning. Secondly, we use optic-disc-centered images in addition to macula-centered images. We obtained the labels for the optic-disc-centered images directly from the corresponding macula-centered labels as weak labels. Then, both types of images are used in fine-tuning. We found an increase in the DR performance using the second approach in both accuracy and quadratic weighted kappa(QWK). Notably, QWK increased from 90.23% to 91.3% using additional weakly labeled optic-disc-centered fundus images. Mohamed ElHabebe, Shereen Elkordi, Ahmed Gamal-Eldin, Noha Adly, Marwan Torki, Ahmed Elmasry, Islam SH Ahmed |
WACV | 5 |
| 2024 | Narrator identification by querying Sanad graph and utilizing the NarratorsKG on AR-Sanad 280K-v2 dataset
Somaia Mahmoud, Emad Nabil, Omar Saif, Marwan Torki |
Neural Comput. Appl. | 4 |
| 2023 | ConvMixer-UNet: A Lightweight Network for Breast Lesion Segmentation in Ultrasound ImagesabstractBreast cancer is one of the most common types of cancer among women. It occurs when abnormal cells in the breast grow and divide uncontrollably. Early diagnosis and treatment are crucial in preventing its spread to the rest of the body. In this paper, we propose a ConvMixer-UNet network for ultrasound image segmentation. The objective is to identify the lesion in the ultrasound image. We design our network that consists of convolutional layers at the early level and ConvMixer layers at the latent level. ConvMixer is an extremely simple and parameter-efficient module that incorporates depthwise and pointwise convolutional layers. This model was evaluated using a breast ultrasound dataset (BUSI); it achieved an improvement in the value of Intersection over Union (IoU). We achieved 68.17% IoU and 80.60% Dice score. These scores are obtained via careful tuning for the network hyperparameters. Quantitative and qualitative comparisons ensure the value of our proposed network. Moreover, ConvMixer-UNet is considered a lightweight network compared to the leading medical segmentation network UNet and its extensions. We show that our network provides a significant reduction in the number of parameters to only 1.77 M parameters, in contrast to UNet which has 31.1 M parameters. Sara AbdElhakem, Sohier Basiony, Marwan Torki |
AICCSA | 3 |
| 2023 | AraPunc: Arabic Punctuation Restoration Using TransformersabstractAdding punctuation to Arabic text enhances readability and clarity. This is very clear in many applications such as automatic speech recognition (ASR) and machine translation (MT) systems. In this paper, we introduce a new punctuation dataset. Our AraPunc dataset is based on the pre-processing of the Tashkeela "Arabic diacritization corpus". We keep six classes: space ‘0’, full-stop ‘.’, comma ‘,’, the colon’:’, semicolon ‘;’, and question mark ‘?’. We treat the punctuation restoration task as a token-wise classification problem that assigns a class (one of the six classes) to each word on the input sentence. We train different transformer-based language models on our new dataset. We found that XLM-RoBERTa outperforms other transformer-based models with a macro-average F1-score of 0.7851 on the AraPunc test set. We also allowed cross-finetuning between QCRI Aljazeera Speech Recognition (QASR) dataset and our novel AraPunc dataset. We managed to achieve a macro average F1-score of 0.7050 on the QASR test set, after training the model first using the AraPunc dataset. Our experiments revealed that AraPunc provides better representations which makes it more suitable to fine-tune models for punctuation restoration task. We release our dataset and code to facilitate future research on this topic1. Abdelrahman Sakr, Marwan Torki |
AICCSA | 2 |
| 2023 | Enhancing LiDAR Semantic Segmentation Using Model SoupsabstractSemantic segmentation of Light Detection and Ranging (LiDAR) is a task that requires high efficiency and accuracy. To our knowledge, this work is the first to apply model soups to the LiDAR semantic segmentation task, showcasing their potential impact on the domain. Our contributions in this work are twofold: First, we successfully extend the application of model soups to LiDAR semantic segmentation. Second, we introduce an optimized and efficient version of the existing greedy soup, further enhancing the overall performance of the approach. In our method, We augment the state-of-the-art open-source code for 2DPASS semantic segmentation with our technique, retaining the original model structure and ensuring no increase in prediction time. The efficiency of our approach is demonstrated using Mean Intersection Over Union (MIoU) as the primary evaluation metric. Our experiments on the SemanticKITTI and NuScenes datasets demonstrate significant improvements. Specifically, we achieve a higher MIoU without any increase in prediction time. Our results, through comprehensive experiments and rigorous evaluations, open up new possibilities for enhanced perception systems in autonomous driving and related fields, highlighting the significance of our iterative uniform greedy model soup in advancing LiDAR semantic segmentation. Omar Wasfy, Sohier Basiony, Marwan Torki |
AICCSA | 3 |
| 2023 | STACKMAPS: A Visualization Technique for Diabetic Retinopathy GradingabstractConvolution Neural Networks (CNN) excelled humans in many classification tasks, including medical imaging applications. However, model interpretation is still an active research area. In this paper, we address the model interpretation problem for the Diabetic Retinopathy grading task. We propose a novel visualization method called StackMaps. Our proposed method is class-agnostic, which fits the diabetic retinopathy grading problem better than other alternatives. Moreover, unlike previous visualization methods, StackMaps gets rid of the dependency on gradients. Our StackMaps technique gets the most significant feature maps at the last convolutional layer after one forward pass. We can also get the significant feature maps from lower layers using beam search guided by the feature maps obtained from the previous layer. Finally, we compute the final map as the sum of the significant feature maps obtained at each layer. We evaluate StackMaps against other state-of-the-art visualization methods qualitatively and quantitatively. We used the FGADR dataset to define our experimental setup. We show that StackMaps achieves better visual interpretation and lesion localization. Ismail El-Yamany, Abdelrahman Wael, Noha Adly, Marwan Torki |
ICASSP | 4 |
| 2023 | Lit the Darkness: Three-Stage Zero-Shot Learning for Low-Light Enhancement with Multi-Neighbor Enhancement FactorsabstractLow-light images represent an obstacle for computer vision tasks due to the lack of perceptual quality. Also, it is challenging to enhance images and adapt to different illumination conditions. To address this problem, we introduce a zero-shot learning approach. We use a 3-stage model trained without the need for paired or unpaired images to improve the lighting and texture of images. The first stage extracts an enhancement pixel-wise map using depth-wise separable convolution. It also tries to extract enhancement factors so that it can consider results e.g., neighbors from other steps in the following stage. The second stage is a recurrent network that enhances the image iteratively while keeping a small model size. The third stage represents an unsupervised network to preserve semantic information and benefit from it during training. We show extensive experiments on benchmark datasets to compare our model with previous state-of-the-art models quantitatively and qualitatively. Mariam Saeed, Marwan Torki |
ICASSP | 2 |
| 2023 | DAUT: Underwater Image Enhancement Using Depth Aware U-shape TransformerabstractImages captured underwater are subject to different complex effects including absorption and scattering. Recovery of the original images is not a trivial task. The underwater image formation varies according to many dependencies. Also, underwater image enhancement models have been affected by the lack of large-scale underwater datasets with reference-enhanced images. Therefore, in this paper, we use cycleGAN to generate a new underwater dataset from existing in-air datasets. We train our data with a U-shape Transformer, which is one of the state-of-the-art models. We add a depth estimation module as object depth is one of the most vital dependencies to the underwater image formation model. The estimated depth and the underwater image are the input to our depth-aware U-transformer. Our experiments show that our model achieves state-of-the-art results in objective and subjective evaluations compared with the original U-shape Transformer and other state-of-the-art methods. the code and dataset are available at https://github.com/MBadran2000/Depth-Aware-U-shape-Transformer.git Mohamed Badran, Marwan Torki |
ICIP | 2 |
| 2023 | The Effect of Non-Reference Point Cloud Quality Assessment (NR-PCQA) Loss on 3D Scene Reconstruction from a Single ImageabstractThis paper proposes a two-stage approach for 3D scene reconstruction from a single image. The first stage involves a monocular depth estimation model, and the second stage involves a point cloud model that recovers depth shift and the focal length from the generated depth map. The paper investigates the use of various pre-trained state-of-the-art transformer models and compares them to existing work without transformers. The loss function is improved by adding a No-Reference point cloud quality assessment (NR-PCQA) to account for the quality of the generated point cloud structure. The paper reports results on four datasets using Locally Scale Invariant RMSE (LSIV) as the metric of evaluation. The paper shows that transformer models outperform previous methods, and transformer models that took into account NR-PCQA outperformed those that did not. Mohamed Zaytoon, Marwan Torki |
ISCC | 2 |
| 2022 | Ablation-CAM++: Grouped Recursive Visual Explanations for Deep Convolutional NetworksabstractRecently, providing explainable deep learning models has sparked a lot of attention. In this paper, we take a further step in this direction. We introduce a time-efficient method, called Ablation-CAM++, which can generate smooth visual explanations of CNN model predictions. Our approach uses the concept of studying the ablation analysis to determine the importance of activation maps w.r.t. the target class, similar to Ablation-CAM. However, instead of focusing on the individual importance of each activation map, we group activation maps using a clustering technique. Then, we construct a binary tree for each group by recursively splitting these groups, studying the ablation of each subgroup, and applying tree pruning. We perform qualitative and quantitative evaluations of our visual explanations against Ablation-CAM and Grad-CAM. Our approach can provide visual explanations in less than half of the time of Ablation-CAM. Using average drop and average increase evaluation metrics on 2000 images of the ImageNet validation set, we provide a comparison of the effect of applying different clustering techniques in our method. Ahmed Salama, Noha Adly, Marwan Torki |
ICIP | 3 |
| 2022 | Vision Transformers Based Classification for Glaucomatous Eye ConditionabstractGlaucoma is an eye condition marked by apoptotic ganglion cell death. This condition is primarily caused by increased intraocular pressure (IOP). The clinical examination of the ganglion cell layer is difficult. However, its demise results in distinctive optic nerve alterations. Several clinical examinations are present to diagnose the glaucoma condition. The fundus image captures the optic nerve. Hence, we can apply deep learning models to automatically diagnose the Glaucoma condition in a fundus image. In this paper, we study the classification of the fundus image using a vision transformer-based ensemble. Vision transformers employ self-attention to capture global characteristics of the fundus image. This makes vision transformers a perfect choice for classification problems. We formed one large merged dataset of six publicly available full fundus images for Glaucoma detection. We provide a comprehensive evaluation for more than seven different vision transformer baseline models. Also, we propose an ensemble of the best vision transformer models. We report areas under the curve, sensitivity, and specificity measures due to the class imbalance in the available data. We report the best standalone model achieves 92.57 in sensitivity, 96.94 in specificity, and 97.9 in AUC. Moustafa Wassel, Ahmed M. Hamdi, Noha Adly, Marwan Torki |
ICPR | 4 |
| 2020 | AQAD: 17, 000+ Arabic Questions for Machine Comprehension of TextabstractCurrent Arabic Machine Reading for Question Answering datasets suffer from important shortcomings. The available datasets are either small-sized high-quality collections or large-sized low-quality datasets. To address the aforementioned problems we present our Arabic Question-Answer dataset (AQAD). AQAD is a new Arabic reading comprehension large-sized high-quality dataset consisting of 17,000+ questions and answers. To collect the AQAD dataset, we present a fully automated data collector. Our collector works on a set of Arabic Wikipedia articles for the extractive question answering task. The chosen articles match the articles used in the well-known Stanford Question Answering Dataset (SQuAD). We provide evaluation results on the AQAD dataset using two state-of-the-art models for machine-reading question answering problems. Namely, BERT and BIDAF models which result in 0.37 and 0.32 F-1 measure on AQAD dataset. Adel Atef, Bassam Mattar, Sandra Sherif, Eman Elrefai, Marwan Torki |
AICCSA | 5 |
| 2020 | A Character Aware Gated Convolution Model for Cloze-style Medical Machine ComprehensionabstractThe machine comprehension research in the medical field has not received significant attention despite its importance in practical applications. This paper is concerned with the cloze-style problem which is one of the tasks of machine comprehension on the medical data. We propose a new model that combines character-level embedding, pre-trained biomedical word embedding, and convolution with attention. The model adopted by this paper achieves better performance than the state-of-the-art results on the BioMedical Knowledge Comprehension Title (BMKC_T) and BioMedical Knowledge Comprehension Last Sentence (BMKC_LS) datasets. We report 84.7% and 77.4% accuracy on BMKC_T and BMKC_LS datasets, respectively. Reham Omar, Nagwa M. El-Makky, Marwan Torki |
AICCSA | 3 |
| 2020 | DeepCReg: Improving Cellular-based Outdoor Localization using CNN-based RegressorsabstractIn this paper, we propose DeepCReg, a convolutional neural network based regressor, that leverages the ubiquitous cellular data to estimate the location of the user in an outdoor environment. We formulate the problem of outdoor localization of a user as a regression problem. This formulation overcomes the limitations of other neural network based classification methods which estimates the position using a grid cell of pre-specified dimensions. We regress on the position directly which leads to better scalability when the testbed area is increased. Moreover, we introduce the usage of convolutional neural networks instead of fully connected neural networks to add more robustness to small changes in the environment. We evaluate our system on two different datasets to emphasize on the scalability of our regression approach. The testbeds are of size 0.147 km2and 1.469 km2. Our system achieves median localization error of 2.06m and 2.82m on each dataset respectively, outperforming current state-of-the-art outdoor cellular based systems by at least 877% improvement in the median localization error. Karim El-Awaad, Mohamed Ezzeldin, Marwan Torki |
WCNC | 3 |
| 2019 | Spectrometer as an Ubiquitous Sensor for IoT Applications Targeting Food QualityabstractIn this paper, we present an IoT application targeting food quality. We utilize the spectrometer as a ubiquitous sensor to provide a Near-infrared (NIR) spectrum that can be used to develop regression models targeting milk quality. We solve a regression problem for different milk components such as Fat, Protein, Lactose, and Solids-Non-Fat. We apply a variety of comparative configurations of regression models, pre-processing techniques, wavelength selection, and feature extraction methods. Our study shows that for the milk data, non-linear models (GPR and SVR) are better choices when compared to linear models. Dina Samak, Sarah Jad, Nour Nabil, Amr G. Wassal, Marwan Torki |
AICCSA | 5 |
| 2019 | Deep Convolutional Neural Network with Multi-Task Learning Scheme for Modulations RecognitionabstractOne of the main characteristics in cognitive radios is situation awareness. By classifying the modulation schemes used in surrounding transmissions, a secondary user (SU) can identify the existing users in the system and adjust his/her transmission parameters accordingly. In this paper, we propose a multi-task learning (MTL) approach to recognize the modulation scheme used among a specific set of analog and digital modulations. This approach uses a deep convolutional neural network (CNN) to extract the necessary features in order to classify the different modulation schemes. The MTL is used to separately train the modulation classes that normally cause a considerable confusion and therefore improve the overall classification accuracy. Our results on the RadioML dataset show that the suggested architecture achieves higher overall classification accuracy compared to the recently proposed Convolutional, Long Short Term Memory (LSTM), Deep Neural Network (CLDNN). Our classification accuracy of 86.97% at 18 dB SNR outperforms the state-of-the-art with 5% relative improvement. Omar S. Mossad, Mustafa ElNainay, Marwan Torki |
IWCMC | 3 |
| 2019 | WiDeep: WiFi-based Accurate and Robust Indoor Localization System using Deep LearningabstractRobust and accurate indoor localization has been the goal of several research efforts over the past decade. Due to the ubiquitous availability of WiFi indoors, many indoor localization systems have been proposed relying on WiFi fingerprinting. However, due to the inherent noise and instability of the wireless signals, the localization accuracy usually degrades and is not robust to dynamic changes in the environment.We present WiDeep, a deep learning-based indoor localization system that achieves a fine-grained and robust accuracy in the presence of noise. Specifically, WiDeep combines a stacked denoising autoencoders deep learning model and a probabilistic framework to handle the noise in the received WiFi signal and capture the complex relationship between the WiFi APs signals heard by the mobile phone and its location. WiDeep also introduces a number of modules to address practical challenges such as avoiding over-training and handling heterogeneous devices.We evaluate WiDeep in two testbeds of different sizes and densities of access points. The results show that it can achieve a mean localization accuracy of 2.64m and 1.21m for the larger and the smaller testbeds, respectively. This accuracy outperforms the state-of-the-art techniques in all test scenarios and is robust to heterogeneous devices. Moustafa Abbas, Moustafa Elhamshary, Hamada Rizk, Marwan Torki, Moustafa Youssef 0001 |
PerCom | 4 |
| 2018 | A Supervised Learning Approach using the Combination of Semantic and Lexical Features for Arabic Community Question AnsweringabstractIn this paper we address the problem of Community Question Answering (CQA) for Arabic language. We mainly explore the direction of combining both lexical and semantic features for enhancing the retrieval task of possible answers to a posted question. We show a comprehensive evaluation on the SEMEval2017 CQA dataset for Arabic language. Our Mean Average Precision (MAP) achieves 62.85% when using a supervised machine learning approach (linear SVM). We outperformed best reported results on such dataset. This is achieved by defining a mix of word embedding, latent semantic similarity features and other lexical similarity features. Mahmoud Abdel-Latif, Mohamed Samir, Shady Abdel-Aziz, Mohamed Heeba, Ahmed Elmasry, Marwan Torki |
AICCSA | 6 |
| 2018 | DeepLoc: a ubiquitous accurate and low-overhead outdoor cellular localization systemabstractRecent years have witnessed fast growth in outdoor location-based services. While GPS is considered a ubiquitous localization system, it is not supported by low-end phones, requires direct line of sight to the satellites, and can drain the phone battery quickly. Ahmed Shokry, Marwan Torki, Moustafa Youssef 0001 |
SIGSPATIAL/GIS | 2 |
| 2018 | Imbalanced Toxic Comments Classification Using Data Augmentation and Deep LearningabstractRecently cyber-bullying and online harassment have become two of the most serious issues in many public online communities. In this paper, we use data from Wikipedia talk page edits to train multi-label classifier that detects different types of toxicity in online user-generated content. We present different data augmentation techniques to overcome the data imbalance problem in the Wikipedia dataset. The proposed solution is an ensemble of three models: convolutional neural network (CNN), bidirectional long short-term memory (LSTM) and bidirectional gated recurrent units (GRU). We divide the classification problem into two steps, first we determine whether or not the input is toxic then we find the types of toxicity present in the toxic content. The evaluation results show that the proposed ensemble approach provides the highest accuracy among all considered algorithms. It achieves 0.828 F1-score for toxic/non-toxic classification and 0.872 for toxicity types prediction. Mai Ibrahim, Marwan Torki, Nagwa M. El-Makky |
ICMLA | 2 |
| 2018 | Interactive Image Segmentation Using Multimodal Regularized Kernel EmbeddingabstractInteractive image segmentation is a classification problem that involves feature extraction like color features. However, in most cases the color features alone cannot discriminate foreground and background. This is due to the fact that both foreground and background can have possibly overlapping color modalities. Using kernels, different notions of pixel similarities can be incorporated. In this paper, we present an interactive image segmentation approach based on kernel embedding using Kernel Local Fisher Discriminant Analysis (KLFDA). We use KLFDA to transform pixel features into a new discriminative feature space. In this feature space the between-class separability is maximized and the locality within each class is preserved. One difficulty for using KLFDA is the setting for the regularization parameter ε. We propose different strategies to overcome this limitation. Our proposed strategies achieve better qualitative and quantitative results compared to state-of-the-art algorithms on the well known ISEG data set for interactive image segmentation. El Moatasem Madani, Marwan Torki |
ICMLA | 2 |
| 2018 | CNN based Indoor Localization using RSS Time-SeriesabstractIndoor localization of mobile nodes is receiving great interest due to the recent advances in mobile devices and the increasing number of location-based services. Fingerprinting based on Wifi received signal strength (RSS) is widely used for indoor localization due to its simplicity and low hardware requirements. However, its positioning accuracy is significantly affected by random fluctuations of RSS values caused by fading and multi-path phenomena. This paper presents a convolutional neural network (CNN) based approach for indoor localization using RSS time-series from wireless local area network (WLAN) access points. Applying CNN on a time-series of RSS readings is expected to reduce the noise and randomness present in separate RSS values and hence improve the localization accuracy. The proposed model is implemented and evaluated on a multi-building and multi-floor dataset, UJIIndoorLoc dataset. The proposed approach provides 100% accuracy for building prediction, 100% accuracy for floor prediction and the mean error in coordinates estimation is 2.77 m. Mai Ibrahim, Marwan Torki, Mustafa ElNainay |
ISCC | 2 |
| 2017 | QweetFinder: Real-Time Finding and Filtering of Question Tweets
Ameer Albahem, Maram Hasanain, Marwan Torki, Tamer Elsayed |
ECIR | 3 |
| 2016 | A multi-modal feature fusion framework for kinect-based facial expression recognition using Dual Kernel Discriminant Analysis (DKDA)abstractWe present a multi-modal feature fusion framework for Kinect-based Facial Expression Recognition (FER). The framework extracts and pre-processes 2D and 3D features separately. The types of 2D and 3D features are selected to maximize the accuracy of the system, with the Histogram of Oriented Gradient (HOG) features for 2D data and statistically selected angles for 3D data giving the best performance. The sets of 2D features and 3D features are reduced and later combined using a novel Dual Kernel Discriminant Analysis (DKDA) approach. Final classification is done using SVMs. The framework is benchmarked on a public Kinect-based FER dataset which includes data for 32 subjects (in both frontal and non-frontal poses and two expression intensities) and 6 basic expressions (plus neutral), namely: happiness, sadness, anger, disgust, fear, and surprise. The framework shows that the proposed combination of 2D and 3D features outperforms simpler existing combinations of 2D and 3D features, as well as systems that use either 2D or 3D features only. The proposed system also outperforms Linear Discriminant Analysis (LDA)-transformed and traditional Kernel Discriminant Analysis (KDA)-transformed systems, with an average accuracy improving of 10%. It also outperforms the state of the art by more than 13% in frontal poses. Sherin Aly, A. Lynn Abbott, Marwan Torki |
WACV | 3 |
| 2016 | Linear-time online action detection from 3D skeletal data using bags of gestureletsabstractSliding window is one direct way to extend a successful recognition system to handle the more challenging detection problem. While action recognition decides only whether or not an action is present in a pre-segmented video sequence, action detection identifies the time interval where the action occurred in an unsegmented video stream. Sliding window approaches can however be slow as they maximize a classifier score over all possible sub-intervals. Even though new schemes utilize dynamic programming to speed up the search for the optimal sub-interval, they require offline processing on the whole video sequence. In this paper, we propose a novel approach for online action detection based on 3D skeleton sequences extracted from depth data. It identifies the sub-interval with the maximum classifier score in linear time. Furthermore, it is suitable for real-time applications with low latency. Moustafa Meshry, Mohamed E. Hussein 0001, Marwan Torki |
WACV | 3 |
| 2016 | Learning representations from multiple manifolds
Chan-Su Lee, Ahmed M. Elgammal, Marwan Torki |
Pattern Recognit. | 3 |
| 2015 | Multi-Modality Feature Transform: An Interactive Image Segmentation ApproachabstractIntroducing suitable features in the scribble-based foreground-background (Fg/Bg) segmentation problem is crucial. In many cases, the object of interest has different re-gions with different color modalities. The same applies to a non-uniform background. Fg/Bg color modalities can even overlap when the appearance is solely modeled using color spaces like RGB or Lab. In this paper, we purposefully discriminate Fg scribbles from Bg scribbles for a better representation. This is achieved by learning a discrimina-tive embedding space from the user-provided scribbles. The transformation between the original features and the embedded features is calculated. This transformation is used to project unlabeled features onto the same embedding space. The transformed features are then used in a supervised classification manner to solve the Fg/Bg segmentation problem. We further refine the results using a self-learning strategy, by expanding scribbles and re-computing the embedding and transformations. Finally, we evaluate our algorithms and compare their performance against the state-of-the-art methods on the ISEG dataset with clear improvements over competing methods. 1 Moustafa Meshry, Ahmed Taha 0001, Marwan Torki |
BMVC | 3 |
| 2015 | Seeded Laplacian: An interactive image segmentation approach using eigenfunctionsabstractIn this paper, we cast the scribbled-based interactive image segmentation as a semi-supervised learning problem. Our novel approach alleviates the need to solve an expensive generalized eigenvector problem by approximating the eigenvectors using a more efficiently computed eigenfunctions. The smoothness operator defined on feature densities at the limit n → ∞ recovers the exact eigenvectors of the graph Laplacian, where n is the number of nodes in the graph. In our experiments scribble annotation is applied, where users label few pixels as foreground and background to guide the foreground/background segmentation. Experiments are carried out on standard data-sets which contain a wide variety of natural images. We achieve better qualitative and quantitative results compared to state-of-the-art algorithms. Ahmed Taha 0001, Marwan Torki |
ICIP | 2 |
| 2015 | Real-Time Multi-scale Action Detection from 3D Skeleton DataabstractIn this paper we introduce a real-time system for action detection. The system uses a small set of robust features extracted from 3D skeleton data. Features are effectively described based on the probability distribution of skeleton data. The descriptor computes a pyramid of sample covariance matrices and mean vectors to encode the relationship between the features. For handling the intra-class variations of actions, such as action temporal scale variations, the descriptor is computed using different window scales for each action. Discriminative elements of the descriptor are mined using feature selection. The system achieves accurate detection results on difficult unsegmented sequences. Experiments on MSRC-12 and G3D datasets show that the proposed system outperforms the state-of-the-art in detection accuracy with very low latency. To the best of our knowledge, we are the first to propose using multi-scale description in action detection from 3D skeleton data. Amr Sharaf, Marwan Torki, Mohamed E. Hussein 0001, Motaz El-Saban |
WACV | 2 |
| 2014 | Spatial-Visual Label Propagation for Local Feature ClassificationabstractIn this paper we present a novel approach to integrate feature similarity and spatial consistency of local features to achieve the goal of localizing an object of interest in an image. The goal is to achieve coherent and accurate labeling of feature points in a simple and effective way. We introduced our Spatial-Visual Label Propagation algorithm to infer the labels of local features in a test image from known labels. This is done in a transductive manner to provide spatial and feature smoothing over the learned labels. We show the value of our novel approach by a diverse set of experiments with successful improvements over previous methods and baseline classifiers. Tarek El-Gaaly, Marwan Torki, Ahmed M. Elgammal |
ICPR | 2 |
| 2013 | Histogram of Oriented Displacements (HOD): Describing Trajectories of Human Joints for Action Recognition
Mohammad Abdelaziz Gowayyed, Marwan Torki, Mohamed E. Hussein 0001, Motaz El-Saban |
IJCAI | 2 |
| 2013 | Human Action Recognition Using a Temporal Hierarchy of Covariance Descriptors on 3D Joint Locations
Mohamed E. Hussein 0001, Marwan Torki, Mohammad Abdelaziz Gowayyed, Motaz El-Saban |
IJCAI | 2 |
| 2012 | RGBD object pose recognition using local-global multi-kernel regression
Tarek El-Gaaly, Marwan Torki |
ICPR | 2 |
| 2011 | Regression from local features for viewpoint and pose estimationabstractIn this paper we propose a framework for learning a regression function form a set of local features in an image. The regression is learned from an embedded representation that reflects the local features and their spatial arrangement as well as enforces supervised manifold constraints on the data. We applied the approach for viewpoint estimation on a Multiview car dataset, a head pose dataset and arm posture dataset. The experimental results show that this approach has superior results (up to 67% improvement) to the state-of-the-art approaches in very challenging datasets. Marwan Torki, Ahmed M. Elgammal |
ICCV | 1 |
| 2010 | Putting local features on a manifoldabstractLocal features have proven very useful for recognition. Manifold learning has proven to be a very powerful tool in data analysis. However, manifold learning application for images are mainly based on holistic vectorized representations of images. The challenging question that we address in this paper is how can we learn image manifolds from a punch of local features in a smooth way that captures the feature similarity and spatial arrangement variability between images. We introduce a novel framework for learning a manifold representation from collections of local features in images. We first show how we can learn a feature embedding representation that preserves both the local appearance similarity as well as the spatial structure of the features. We also show how we can embed features from a new image by introducing a solution for the out-of-sample that is suitable for this context. By solving these two problems and defining a proper distance measure in the feature embedding space, we can reach an image manifold embedding space. Marwan Torki, Ahmed M. Elgammal |
CVPR | 1 |
| 2010 | One-shot multi-set non-rigid feature-spatial matchingabstractWe introduce a novel framework for nonrigid feature matching among multiple sets in a way that takes into consideration both the feature descriptor and the features spatial arrangement. We learn an embedded representation that combines both the descriptor similarity and the spatial arrangement in a unified Euclidean embedding space. This unified embedding is reached by minimizing an objective function that has two sources of weights; the feature spatial arrangement and the feature descriptor similarity scores across the different sets. The solution can be obtained directly by solving one Eigen-value problem that is linear in the number of features. Therefore, the framework is very efficient and can scale up to handle a large number of features. Experimental evaluation is done using different sets showing outstanding results compared to the state of the art; up to 100% accuracy is achieved in the case of the well known `Hotel' sequence. Marwan Torki, Ahmed M. Elgammal |
CVPR | 1 |
| 2010 | Learning a Joint Manifold Representation from Multiple Data SetsabstractThe problem we address in the paper is how to learn a joint representation from data lying on multiple manifolds. We are given multiple data sets and there is an underlying common manifold among the different data set. We propose a framework to learn an embedding of all the points on all the manifolds in a way that preserves the local structure on each manifold and, in the same time, collapses all the different manifolds into one manifold in the embedding space, while preserving the implicit correspondences between the points across different data sets. The proposed solution works as extensions to current state of the art spectral-embedding approaches to handle multiple manifolds. Marwan Torki, Ahmed M. Elgammal, Chan-Su Lee |
ICPR | 1 |