EDBT 2026 Demo / reviewers in the wild / expert
Hatem A. Rashwan
dblp:54/10781 · also Hatem Abd Ellatif FatahAllah Ibrahim Mahmoud Rashwan
· DBLP profile ↗
33ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0001-5421-1637ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoAtXNet: A dual-stream hybrid transformer based on relative cross-attention for end-to-end camera localization from RGB-D imagesabstractCamera localization is the process of automatically determining the position and orientation of a camera with respect to its 3D environment based on the images it captures. Camera localization methods, including classical structure-based techniques and modern deep neural networks, often encounter limitations in visually complex environments when relying only on RGB images. Some of these limitations can be overcome by processing RGB-D images, provided that the available color and depth cues are properly integrated. Building upon CoAtNet, a successful hybrid Transformer neural model, this paper introduces CoAtXNet: a dual-stream architecture for jointly processing RGB color and depth channels. CoAtXNet interrelates two CoAtNet networks working in tandem through a straightforward cross-attention mechanism that intertwines the final Transformer stages. This integration of independent data streams yields enhanced feature representations. Experiments on the well-known 7-Scenes and 12-Scenes RGB-D camera localization datasets show that CoAtXNet achieves highly competitive performance, improving average translation and orientation accuracy over several recent localization methods. In addition to quantitative comparisons, we analyze the behavior of CoAtXNet under progressively degraded RGB appearance. The results show that the proposed model is less sensitive to appearance degradation than simpler RGB-D fusion baselines, suggesting that the cross-attention mechanism allows the model to place greater emphasis on geometric structures when RGB reliability decreases. An implementation of the proposed method is publicly available. Hussein Hasan, Miguel Ángel García, Hatem A. Rashwan, Domenec Puig |
Comput. Vis. Image Underst. | 3 |
| 2026 | M-TabNet: A Transformer-Based Multi-Encoder for Early Neonatal Birth Weight Prediction Using Multimodal DataabstractBirth weight (BW) is a key indicator of neonatal health, and low birth weight (LBW) is linked to increased mortality and morbidity. Early prediction of BW facilitates timely prevention of impaired foetal growth. However, available techniques such as ultrasonography have limitations, including less accuracy when applied before 20 weeks of gestation and operator-dependent variability. Existing BW prediction models often neglect nutritional and genetic influences, and focus mainly on physiological and lifestyle factors. This study presents an attention-based transformer model with a multi-encoder architecture for early ($< 12$ weeks) BW prediction. Our model effectively integrates diverse maternal data, including physiological, lifestyle, nutritional, and genetic data, addressing limitations seen in previous attention-based models such as TabNet. The model achieves a Mean Absolute Error (MAE) of 122 grams and an $R^{2}$ value of 0.94, showing its high predictive accuracy and interoperability with our in-house private dataset. Independent validation confirms generalizability (MAE: 105 grams, $R^{2}$: 0.95) with the IEEE children dataset. To enhance clinical utility, predicted BW is classified into low and normal categories, achieving a sensitivity of 97.55% and a specificity of 94.48%, facilitating early risk stratification. Model interpretability is reinforced through feature importance and SHAP analysis, highlighting significant influences of maternal age, tobacco exposure, and vitamin B12 status, with genetic factors playing a secondary role. Our results emphasize the potential of advanced deep learning models to improve early BW prediction, offering a robust, interpretable, and personalized tool to identify pregnancies at risk and optimize neonatal outcomes. Muhammad Mursil, Hatem A. Rashwan, Luis Santos-Calderon, Pere Cavallé-Busquets, Michelle M. Murphy, Domenec Puig |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Transformer-Based Bi-encoder for Early Neonatal Birth Weight Prediction Using Maternal Nutritional and Health Insights
Muhammad Mursil, Hatem A. Rashwan, Adnan Khalid 0001, Luis Santos-Calderon, Pere Cavallé-Busquets, Michelle M. Murphy, Domenec Puig |
AIME (2) | 2 |
| 2025 | Correction to: Implicit regularization of a deep augmented neural network model for human motion prediction
Gaurav Kumar Yadav, Mohamed Abdel-Nasser, Hatem A. Rashwan, Domenec Puig, Gora Chand Nandi |
Appl. Intell. | 3 |
| 2025 | Efficient crack segmentation with multi-decoder networks and enhanced feature fusionabstractIn infrastructure management, intelligent crack detection is vital, particularly for maintaining crucial elements such as road networks in urban areas. Detecting pavement defects promptly and accurately is essential for timely repairs and hazard prevention. However, this task is challenging due to factors, such as complex backgrounds, micro defects, diverse defect shapes and sizes, and class imbalance issues. Innovative approaches and advanced technologies are needed to address these challenges and effectively manage infrastructure complexities. In this study, we propose a novel framework for crack segmentation, called CrackMaster. CrackMaster utilizes advanced neural network architectures, leveraging the next generation of convolutional networks (ConvNeXt) as an encoder and dual decoders customized for distinct tasks. The first decoder adopts a self-supervised learning paradigm to reconstruct images, thereby enhancing feature extraction capabilities. Meanwhile, the second decoder combines deep labelling network for semantic image segmentation (Deeplabv3+) with a light deep neural network (LinkNet) to facilitate precise segmentation. Notably, we introduce an Enhanced Feature Fusion (EFF) block to improve features quality, enhancing information flow and context preservation, thus boosting segmentation performance. Experimental results conducted on three diverse datasets, including our in-house Road Crack Dataset (RCD), DeepCrack537, and Yang Crack Dataset (YCD) datasets, demonstrate the effectiveness of our framework achieving outstanding Intersection over Union scores (IoU) of 86.0%, 87.8%, and 76.9%, respectively, showing superior accuracy and robustness in crack segmentation tasks. These findings underscore the potential applicability of our framework in real-world infrastructure management scenarios. The code is publicly available at: https://github.com/AmmarOkran/CrackMaster . • CrackMaster Framework : ConvNeXt-based model with dual decoders for automated crack segmentation. • Image Reconstruction : Utilizes a self-supervised learning paradigm in the first decoder to boost feature extraction capabilities. • Precise Segmentation : Integrates Deeplabv3+ and LinkNet networks in the second decoder for accurate crack image segmentation. • Enhanced Feature Fusion : EFF block to improve feature quality, information flow, and context preservation. Ammar M. Okran, Hatem A. Rashwan, Adel Saleh, Domenec Puig |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | CoHAtNet: An integrated convolutional-transformer architecture with hybrid self-attention for end-to-end camera localizationabstractCamera localization refers to the process of automatically determining the position and orientation of a camera within its 3D environment from the images it captures. Traditional camera localization methods often rely on Convolutional Neural Networks, which are effective at extracting local visual features but struggle to capture long-range dependencies critical for accurate localization. In contrast, Transformer-based approaches model global contextual relationships appropriately, although they often lack precision in fine-grained spatial representations. To bridge this gap, we introduce CoHAtNet, a novel Convolutional Hybrid-Attention Network that tightly integrates convolutional and self-attention mechanisms. Unlike previous hybrid models that stack convolutional and attention layers separately, CoHAtNet embeds local features extracted via Mobile Inverted Bottleneck Convolution blocks directly into the Value component of the self-attention mechanism of Transformers. This yields a hybrid self-attention block capable of dynamically capturing both local spatial detail and global semantic context within a single attention layer. Additionally, CoHAtNet enables modality-level fusion by processing RGB and depth data jointly in a unified pipeline, allowing the model to leverage complementary appearance and geometric cues throughout. Extensive evaluations have been conducted on two widely-used camera localization datasets: 7-Scenes (RGB-D) and Cambridge Landmarks (RGB). Experimental results show that CoHAtNet achieves state-of-the-art performance in both translation and orientation accuracy. These results highlight the effectiveness of our hybrid design in challenging indoor and outdoor environments. This makes CoHAtNet a strong candidate for end-to-end camera localization tasks. Hussein Hasan, Miguel Ángel García, Hatem A. Rashwan, Domenec Puig |
Image Vis. Comput. | 3 |
| 2025 | Interpretable deep neural networks for advancing early neonatal birth weight prediction using multimodal maternal factorsabstractBACKGROUND: Neonatal low birth weight (LBW) is a significant predictor of increased morbidity and mortality among newborns. Predominantly, traditional prediction methods depend heavily on ultrasonography, which does not consider risk factors affecting birth weight (BW). OBJECTIVE: This study introduces a robust deep neural network for a clinical decision-support system designed to early predict neonatal BW, using data available during early pregnancy, with enhanced precision. This innovative system incorporates a comprehensive array of maternal factors, placing particular emphasis on nutritional elements alongside physiological and lifestyle variables. METHODS: We employed and validated various traditional machine learning models as well as an interpretable deep learning model using the TabNet architecture, noted for its proficient handling of tabular data and high level of interpretability. The efficacy of these models was evaluated against extensive datasets that encompass a broad spectrum of maternal health indicators. RESULTS: The TabNet model exhibited outstanding predictive capabilities, achieving an accuracy of 96% and an area under the curve (AUC) of 0.96. Significantly, maternal vitamin B12 and folate status emerged as pivotal predictors of BW, emphasizing the crucial role of nutritional factors in influencing neonatal health outcomes. CONCLUSIONS: Our results demonstrate the substantial benefits of integrating multimodal maternal factors into predictive models for neonatal BW, markedly enhancing the precision over traditional AI methods. The developed decision-support system not only has a possible application in prenatal care but also provides actionable insights that can be leveraged to mitigate the risks associated with LBW, thereby improving clinical decision-making processes and outcomes. Muhammad Mursil, Hatem A. Rashwan, Adnan Khalid 0001, Pere Cavallé-Busquets, Luis Santos-Calderon, Michelle M. Murphy, Domenec Puig |
J. Biomed. Informatics | 2 |
| 2025 | Adaptive weighted multi-teacher distillation for efficient medical imaging segmentation with limited dataabstractAdvances in deep learning models have significantly improved performance in medical tasks, but their complex structures and high computational requirements pose challenges for clinical implementation. Additionally, data privacy concerns limit the availability of comprehensive datasets needed to train accurate models. To address these issues, we propose a novel adaptive knowledge distillation (KD) framework for medical imaging segmentation that integrates intermediate and high-level feature pairwise relationships between teacher and student models. Our framework features adaptive multi-teacher distillation, where multiple teacher models, each trained on limited data from different sites and hospitals with various scanning protocols, distill their knowledge to a student model using adaptive weighting. This method allows each teacher to convey deep feature representations to the student’s intermediate layers, enhancing performance without increasing complexity. To validate the efficacy of our framework, we conducted extensive experiments on two publicly available medical datasets, focusing on prostate and spleen tumor segmentation tasks. Our adaptive KD approach significantly improved dice scores by up to 9%, surpassing all tested baseline models. These results highlight the potential of our KD framework to enhance medical imaging segmentation while ensuring data privacy and security. • Implement both single and multi-teacher distillation scenarios. • Apply adaptive multi-teacher distillation using limited data. • Distill intermediate and high-level features. • Conduct CT and MRI medical imaging segmentation experiments using public datasets. Eddardaa Ben Loussaief, Hatem A. Rashwan, Mohammed Ayad, Adnan Khalid 0001, Domemec Puig |
Knowl. Based Syst. | 2 |
| 2024 | FGR-Net: Interpretable fundus image gradeability classification based on deep reconstruction learningabstractThe performance of diagnostic Computer-Aided Design (CAD) systems for retinal diseases depends on the quality of the retinal images being screened. Thus, many studies have been developed to evaluate and assess the quality of such retinal images. However, most of them did not investigate the relationship between the accuracy of the developed models and the quality of the visualization of interpretability methods for distinguishing between gradable and non-gradable retinal images. Consequently, this paper presents a novel framework called “FGR-Net” to automatically asses and interpret underlying fundus image quality by merging an autoencoder network with a classifier network. The FGR-Net model also provides an interpretable quality assessment through visualizations. In particular, FGR-Net uses a deep autoencoder to reconstruct the input image in order to extract the visual characteristics of the input fundus images based on self-supervised learning. The extracted features by the autoencoder are then fed into a deep classifier network to distinguish between gradable and ungradable fundus images. FGR-Net is evaluated with different interpretability methods, which indicates that the autoencoder is a key factor in forcing the classifier to focus on the relevant structures of the fundus images, such as the fovea, optic disc, and prominent blood vessels. Additionally, the interpretability methods can provide visual feedback for ophthalmologists to understand how our model evaluates the quality of fundus images. The experimental results showed the superiority of FGR-Net over the state-of-the-art quality assessment methods, with an accuracy of >89% and an F1-score of >87%. The code is publicly available at https://github.com/saifalkh/FGR-Net. Saif Khalid, Hatem A. Rashwan, Saddam Abdulwahab, Mohamed Abdel-Nasser, Facundo Manuel Quiroga, Domenec Puig |
Expert Syst. Appl. | 2 |
| 2024 | VISTA: vision improvement via split and reconstruct deep neural network for fundus image quality assessmentabstractAbstract Widespread eye conditions such as cataracts, diabetic retinopathy, and glaucoma impact people worldwide. Ophthalmology uses fundus photography for diagnosing these retinal disorders, but fundus images are prone to image quality challenges. Accurate diagnosis hinges on high-quality fundus images. Therefore, there is a need for image quality assessment methods to evaluate fundus images before diagnosis. Consequently, this paper introduces a deep learning model tailored for fundus images that supports large images. Our division method centres on preserving the original image’s high-resolution features while maintaining low computing and high accuracy. The proposed approach encompasses two fundamental components: an autoencoder model for input image reconstruction and image classification to classify the image quality based on the latent features extracted by the autoencoder, all performed at the original image size, without alteration, before reassembly for decoding networks. Through post hoc interpretability methods, we verified that our model focuses on key elements of fundus image quality. Additionally, an intrinsic interpretability module has been designed into the network that allows decomposing class scores into underlying concepts quality such as brightness or presence of anatomical structures. Experimental results in our model with EyeQ, a fundus image dataset with three categories (Good, Usable, and Rejected) demonstrate that our approach produces competitive outcomes compared to other deep learning-based methods with an overall accuracy of 0.9066, a precision of 0.8843, a recall of 0.8905, and an impressive F1-score of 0.8868. The code is publicly available at https://github.com/saifalkhaldiurv/VISTA_-Image-Quality-Assessment . Saif Khalid, Saddam Abdulwahab, Oscar Stanchi, Facundo Manuel Quiroga, Franco Ronchetti, Domenec Puig, Hatem A. Rashwan |
Neural Comput. Appl. | 7 |
| 2023 | Implicit regularization of a deep augmented neural network model for human motion predictionabstractAbstract Predicting human motion based on past observed motion is one of the challenging issues in computer vision and graphics. Existing research works are dealing with this issue by using discriminative models and showing the results for cases that follow a homogeneous distribution (in distribution) and not discussing the issues of the domain shift problem, where training and testing data follow a heterogeneous (out of distribution) problem, which is the reality when such models are used in practice. However, recent research proposed addressing domain shift issues by augmenting the discriminative model with a generative model and obtained better results. In the present investigation, we propose regularizing the extended network by inserting linear layers to minimize the rank of the latent space and train the entire end-to-end network. We regularize the network to strengthen the model to deal effectively with domain shift scenarios. Both training and testing data come from different distribution sets; to deal with this, we toughen our network by adding the extra linear layers to the network encoder. We tested our model with the benchmark datasets, CMU Motion Capture and Human3.6M, and proved that our model outperforms 14 OoD actions of H3.6M and 7 OoD actions of CMU MoCap in terms of the Euclidean distance calculated between predicted and ground truth joint angle values. Our average results of 14 OoD actions for short-term (80, 160, 320, 400) are 0.34, 0.6, 0.96, 1.07, and for CMU MoCap of 7 OoD actions for short-term and long term (80, 160, 320, 400, 1000) are 0.28, 0.45, 0.77, 0.89, 1.46. All these results are much better than the other state-of-the-art results. Gaurav Kumar Yadav, Mohamed Abdel-Nasser, Hatem A. Rashwan, Domenec Puig, Gora Chand Nandi |
Appl. Intell. | 3 |
| 2023 | GCNDepth: Self-supervised monocular depth estimation based on graph convolutional networkabstractDepth estimation is a challenging task of 3D reconstruction to enhance the accuracy sensing of environment awareness. This work brings a new solution with improvements, which increases the quantitative and qualitative understanding of depth maps compared to existing methods. Recently, convolutional neural networks (CNN) have demonstrated their extraordinary ability to estimate depth maps from monocular videos. However, traditional CNN does not support a topological structure, and they can work only on regular image regions with determined sizes and weights. On the other hand, graph convolutional networks (GCN) can handle the convolution of non-Euclidean data, and they can be applied to irregular image regions within a topological structure. Therefore, to preserve object geometric appearances and objects locations in the scene, in this work, we aim to exploit GCN for a self-supervised monocular depth estimation model. Our model consists of two parallel auto-encoder networks: the first is an auto-encoder that will depend on ResNet-50 and extract the feature from the input image and on multi-scale GCN to estimate the depth map. In turn, the second network will be used to estimate the ego-motion vector (i.e., 3D pose) between two consecutive frames based on ResNet-18. The estimated 3D pose and depth map will be used to construct the target image. A combination of loss functions related to photometric, reprojection, and smoothness is used to cope with bad depth prediction and preserve the discontinuities of the objects. Our method and performance are improved quantitatively and qualitatively. In particular, our method provided comparable and promising results with a high prediction accuracy of 89% on the publicly available KITTI dataset. Our method also offers 40% reduction in the number of trainable parameters compared to the state of the art solutions.In addition, we tested our trained model with Make3D dataset to evaluate the trained model on a new dataset with low resolution images. The source code is publicly available at (https://github.com/ArminMasoumian/GCNDepth.git) Armin Masoumian, Hatem A. Rashwan, Saddam Abdulwahab, Julián Cristiano, Muhammad Salman Asif, Domenec Puig |
Neurocomputing | 2 |
| 2022 | Effective Deep Learning-Based Ensemble Model for Road Crack DetectionabstractThis paper proposes an effective deep learning-based model for crack detection in images acquired by different acquisition systems (e.g., cameras mounted on vehicles and drones and smartphone cameras) in six countries. By utilizing successful training procedures and including multi-scale feature extraction models, the ensemble model is built using effective variations of the cutting-edge object detection technique, Yolov7. The top crack detection models are fused using the non-maximum suppression method to create the proposed ensemble model. The proposed crack detection model is trained and validated using the crowdsensing-based road damage detection challenge (CRDDC2022). With the test set, the proposed model produced an average F1 score of 0.65 with all leaderboards of CRDDC2022 (all countries, India, Japan, Norway, and the United States leaderboards). Our approach is ranked in the 5thposition in the CRDDC2022 challenge1. The source code is available at https://github.com/AmmarOkran/CRDD2022. Ammar M. Okran, Mohamed Abdel-Nasser, Hatem A. Rashwan, Domenec Puig |
IEEE Big Data | 3 |
| 2022 | Monocular depth map estimation based on a multi-scale deep architecture and curvilinear saliency feature boosting
Saddam Abdulwahab, Hatem A. Rashwan, Miguel Ángel García, Armin Masoumian, Domenec Puig |
Neural Comput. Appl. | 2 |
| 2022 | Efficient deep learning-based semantic mapping approach using monocular vision for resource-limited mobile robots
Raghav Narula, Hatem A. Rashwan, Mohamed Abdel-Nasser, Domenec Puig, Gora Chand Nandi |
Neural Comput. Appl. | 3 |
| 2021 | SLSNet: Skin lesion segmentation using a lightweight generative adversarial networkabstractThe determination of precise skin lesion boundaries in dermoscopic images using automated methods faces many challenges, most importantly, the presence of hair, inconspicuous lesion edges and low contrast in dermoscopic images, and variability in the color, texture and shapes of skin lesions. Existing deep learning-based skin lesion segmentation algorithms are expensive in terms of computational time and memory. Consequently, running such segmentation algorithms requires a powerful GPU and high bandwidth memory, which are not available in dermoscopy devices. Thus, this article aims to achieve precise skin lesion segmentation with minimum resources: a lightweight, efficient generative adversarial network (GAN) model called SLSNet, which combines 1-D kernel factorized networks, position and channel attention, and multiscale aggregation mechanisms with a GAN model. The 1-D kernel factorized network reduces the computational cost of 2D filtering. The position and channel attention modules enhance the discriminative ability between the lesion and non-lesion feature representations in spatial and channel dimensions, respectively. A multiscale block is also used to aggregate the coarse-to-fine features of input skin images and reduce the effect of the artifacts. SLSNet is evaluated on two publicly available datasets: ISBI 2017 and the ISIC 2018. Although SLSNet has only 2.35 million parameters, the experimental results demonstrate that it achieves segmentation results on a par with the state-of-the-art skin lesion segmentation methods with an accuracy of 97.61%, and Dice and Jaccard similarity coefficients of 90.63% and 81.98%, respectively. SLSNet can run at more than 110 frames per second (FPS) in a single GTX1080Ti GPU, which is faster than well-known deep learning-based image segmentation models, such as FCN. Therefore, SLSNet can be used for practical dermoscopic applications. Md. Mostafa Kamal Sarker, Hatem A. Rashwan, Farhan Akram, Vivek Kumar Singh 0008, Syeda Furruka Banu, Forhad U. H. Chowdhury, Kabir Ahmed Choudhury, Sylvie Chambon, Petia Radeva, Domenec Puig, Mohamed Abdel-Nasser |
Expert Syst. Appl. | 2 |
| 2020 | Breast tumor segmentation in ultrasound images using contextual-information-aware deep adversarial learning framework
Vivek Kumar Singh 0008, Mohamed Abdel-Nasser, Farhan Akram, Hatem A. Rashwan, Md. Mostafa Kamal Sarker, Nidhi Pandey, Santiago Romaní, Domenec Puig |
Expert Syst. Appl. | 4 |
| 2020 | Breast tumor segmentation and shape classification in mammograms using generative adversarial and convolutional neural network
Vivek Kumar Singh 0008, Hatem A. Rashwan, Santiago Romaní, Farhan Akram, Nidhi Pandey, Md. Mostafa Kamal Sarker, Adel Saleh, Meritxell Arenas, Miguel Arquez, Domenec Puig, Jordina Torrents-Barrena |
Expert Syst. Appl. | 2 |
| 2020 | Action representation and recognition through temporal co-occurrence of flow fields and convolutional neural networks
Hatem A. Rashwan, Miguel Ángel García, Saddam Abdulwahab, Domenec Puig |
Multim. Tools Appl. | 1 |
| 2020 | Adversarial Learning for Depth and Viewpoint Estimation From a Single ImageabstractEstimating a depth map and, at the same time, predicting the 3D pose of an object from a single 2D color image is a very challenging task. Depth estimation is typically performed through stereo vision by following several time-consuming stages, such as epipolar geometry, rectification and matching. Alternatively, when stereo vision is not useful or applicable, depth relations can be inferred from a single image as studied in this paper. More precisely, deep learning is applied in order to solve the problem of estimating a depth map from a single image. Then, that map is used for predicting the 3D pose of the main object depicted in the image. The proposed model consists of two successive neural networks. The first network is based on a Generative Adversarial Neural network (GAN). It estimates a dense depth map from the given color image. A Convolutional Neural Network (CNN) is then used to predict the 3D pose from the generated depth map through regression. The main difficulty to jointly estimate depth maps and 3D poses using deep networks is the lack of training data with both depth and viewpoint annotations. This contribution assumes a cross-domain training procedure with 3D CAD models corresponding to objects appearing in real images in order to render depth images from different viewpoints. These rendered images are then used to guide the GAN network to learn the mapping from the image domain to the depth domain. By exploiting the dataset as a source of training data, the proposed model outperforms state-of-the-art models on the PASCAL 3D+ dataset. The code of the proposed model is publicly available at https://github.com/SaddamAbdulrhman/Depth-and-Viewpoint-Estimation/tree/master. Saddam Abdulwahab, Hatem A. Rashwan, Miguel Ángel García, Mohammed Jabreel, Sylvie Chambon, Domenec Puig |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Gait representation and recognition from temporal co-occurrence of flow fields
Hatem A. Rashwan, Miguel Ángel García, Sylvie Chambon, Domenec Puig |
Mach. Vis. Appl. | 1 |
| 2019 | Using Curvilinear Features in Focus for Registering a Single Image to a 3D ObjectabstractIn the context of 2D/3D registration, this paper introduces an approach that allows for matching features detected in two different modalities, photographs, and 3D models, by using a common 2D representation. More precisely, 2D images are matched with a set of depth images representing the 3D model. After introducing the concept of Curvilinear Saliency, which is related to curvature estimation, we propose a new ridge and valley detector for depth images rendered from 3D models. A variant of this detector is adapted to photographs, first by considering multi-scale features and second by integrating the focus curve principle. Finally, a registration algorithm determines the correct view of the 3D model and, thus, the pose of the photograph. This approach relies on the Histogram of Curvilinear Saliency (HCS), an adaptation of the Histogram of Oriented Gradients (HOG) to the proposed features in 2D and 3D. The presented results highlight both the quality of the features detected in terms of repeatability and the interest of the approach for registration and pose estimation. Hatem A. Rashwan, Sylvie Chambon, Pierre Gurdjos, Géraldine Morin, Vincent Charvillat |
IEEE Trans. Image Process. | 1 |
| 2018 | BSCGAN: Deep Background Subtraction with Conditional Generative Adversarial NetworksabstractThis paper proposes a deep background subtraction method based on conditional Generative Adversarial Network (cGAN). The proposed model consists of two successive networks: generator and discriminator. The generator learns the mapping from the observing input (i.e., image and background), to the output (i.e., foreground mask). Then, the discriminator learns a loss function to train this mapping by comparing real foreground (i.e., ground-truth) and fake foreground (i.e., predicted output) with observing the input image and background. Evaluating the model performance with two public datasets, CDnet 2014 and BMC, shows that the proposed model outperforms the state-of-the-art methods. Mohamed Chafik Bakkay, Hatem A. Rashwan, Houssam Salmane, Louahdi Khoudour, D. Puigtt, Yassine Ruichek |
ICIP | 2 |
| 2018 | SLSDeep: Skin Lesion Segmentation Based on Dilated Residual and Pyramid Pooling Networks
Md. Mostafa Kamal Sarker, Hatem A. Rashwan, Farhan Akram, Syeda Furruka Banu, Adel Saleh, Vivek Kumar Singh 0008, Forhad U. H. Chowdhury, Saddam Abdulwahab, Santiago Romaní, Petia Radeva, Domenec Puig |
MICCAI (2) | 2 |
| 2018 | Conditional Generative Adversarial and Convolutional Networks for X-ray Breast Mass Segmentation and Shape Classification
Vivek Kumar Singh 0008, Santiago Romaní, Hatem A. Rashwan, Farhan Akram, Nidhi Pandey, Md. Mostafa Kamal Sarker, Saddam Abdulwahab, Jordina Torrents-Barrena, Adel Saleh, Miguel Arquez, Meritxell Arenas, Domenec Puig |
MICCAI (2) | 3 |
| 2018 | Automatic detection of individual and touching moths from trap images by combining contour-based and region-based segmentationabstractInsect detection is one of the most challenging problems of biometric image processing. This study focuses on developing a method to detect both individual insects and touching insects from trap images in extreme conditions. This method is able to combine recent approaches on contour‐based and region‐based segmentation. More precisely, the two contributions are: an adaptive k ‐means clustering approach by using the contour's convex hull and a new region merging algorithm. Quantitative evaluations show that the proposed method can detect insects with higher accuracy than that of the most used approaches. Mohamed Chafik Bakkay, Sylvie Chambon, Hatem A. Rashwan, Christian Lubat, Sébastien Barsotti |
IET Comput. Vis. | 3 |
| 2017 | Analyzing the evolution of breast tumors through flow fields and strain tensors
Mohamed Abdel-Nasser, Antonio Moreno, Hatem A. Rashwan, Domenec Puig |
Pattern Recognit. Lett. | 3 |
| 2016 | Towards multi-scale feature detection repeatable over intensity and depth imagesabstractObject recognition based on local features computed at multiple locations is robust to occlusions, strong viewpoint changes and object deformations. These features should be repeatable, precise and distinctive. We present an operator for repeatable feature detection on depth images (relative to 3D models) as well as 2D intensity images. The proposed detector is based on estimating the curviness saliency at multiple scales in each kind of image. We also propose quality measures that evaluate the repeatability of the features between depth and intensity images. The experiments show that the proposed detector outperforms both the most powerful, classical point detectors (e.g., SIFT) and edge detection techniques. Hatem A. Rashwan, Sylvie Chambon, Pierre Gurdjos, Géraldine Morin, Vincent Charvillat |
ICIP | 1 |
| 2016 | Defeating face de-identification methods based on DCT-block scrambling
Hatem A. Rashwan, Miguel Ángel García, Antoni Martínez-Ballesté, Domenec Puig |
Mach. Vis. Appl. | 1 |
| 2015 | Analysis of tissue abnormality and breast density in mammographic images using a uniform local directional pattern
Mohamed Abdel-Nasser, Hatem A. Rashwan, Domenec Puig, Antonio Moreno |
Expert Syst. Appl. | 2 |
| 2014 | Illumination-Robust Optical Flow Using a Local Directional PatternabstractMost of the variational optical flow methods are based on the well-known brightness constancy assumption or high-order constancy assumptions to implement the data term in the optimization energy function. Unfortunately, any variation in the lighting within the scene violates the brightness constancy constraint; in turn, the gradient constancy assumption does not work properly with large illumination changes. This paper proposes an illumination-robust constancy based on a robust texture descriptor rather than the brightness constancy. Thus, the similarity function used as a data term was obtained from extracting texture features through the local directional pattern descriptor for two consecutive frames within the duality total variational optical flow algorithm. In addition, a weighted nonlocal term that depends on both the color similarity and the occlusion state of pixels is integrated during the optimization process to increase the accuracy of the resulting flow field. The experimental results show a qualitative comparison with the proposed approach and yield state-of-the-art results on the KITTI, Midleburry, and MPI-sintel data sets. Mahmoud A. Mohamed, Hatem A. Rashwan, Bärbel Mertsching, Miguel Ángel García, Domenec Puig |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Variational Optical Flow Estimation Based on Stick Tensor VotingabstractVariational optical flow techniques allow the estimation of flow fields from spatio-temporal derivatives. They are based on minimizing a functional that contains a data term and a regularization term. Recently, numerous approaches have been presented for improving the accuracy of the estimated flow fields. Among them, tensor voting has been shown to be particularly effective in the preservation of flow discontinuities. This paper presents an adaptation of the data term by using anisotropic stick tensor voting in order to gain robustness against noise and outliers with significantly lower computational cost than (full) tensor voting. In addition, an anisotropic complementary smoothness term depending on directional information estimated through stick tensor voting is utilized in order to preserve discontinuity capabilities of the estimated flow fields. Finally, a weighted non-local term that depends on both the estimated directional information and the occlusion state of pixels is integrated during the optimization process in order to denoise the final flow field. The proposed approach yields state-of-the-art results on the Middlebury benchmark. Hatem A. Rashwan, Miguel Ángel García, Domenec Puig |
IEEE Trans. Image Process. | 1 |
| 2012 | Improving the robustness of variational optical flow through tensor voting
Hatem A. Rashwan, Domenec Puig, Miguel Ángel García |
Comput. Vis. Image Underst. | 1 |