EDBT 2026 Demo / reviewers in the wild / expert
Domenec Puig
dblp:76/447 · also Domenec Puig-Valls
· DBLP profile ↗
88ranked-venue papers
4as first author
26since 2021 · last 2026
0000-0002-0562-4205ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoAtXNet: A dual-stream hybrid transformer based on relative cross-attention for end-to-end camera localization from RGB-D imagesabstractCamera localization is the process of automatically determining the position and orientation of a camera with respect to its 3D environment based on the images it captures. Camera localization methods, including classical structure-based techniques and modern deep neural networks, often encounter limitations in visually complex environments when relying only on RGB images. Some of these limitations can be overcome by processing RGB-D images, provided that the available color and depth cues are properly integrated. Building upon CoAtNet, a successful hybrid Transformer neural model, this paper introduces CoAtXNet: a dual-stream architecture for jointly processing RGB color and depth channels. CoAtXNet interrelates two CoAtNet networks working in tandem through a straightforward cross-attention mechanism that intertwines the final Transformer stages. This integration of independent data streams yields enhanced feature representations. Experiments on the well-known 7-Scenes and 12-Scenes RGB-D camera localization datasets show that CoAtXNet achieves highly competitive performance, improving average translation and orientation accuracy over several recent localization methods. In addition to quantitative comparisons, we analyze the behavior of CoAtXNet under progressively degraded RGB appearance. The results show that the proposed model is less sensitive to appearance degradation than simpler RGB-D fusion baselines, suggesting that the cross-attention mechanism allows the model to place greater emphasis on geometric structures when RGB reliability decreases. An implementation of the proposed method is publicly available. Hussein Hasan, Miguel Ángel García, Hatem A. Rashwan, Domenec Puig |
Comput. Vis. Image Underst. | 4 |
| 2026 | M-TabNet: A Transformer-Based Multi-Encoder for Early Neonatal Birth Weight Prediction Using Multimodal DataabstractBirth weight (BW) is a key indicator of neonatal health, and low birth weight (LBW) is linked to increased mortality and morbidity. Early prediction of BW facilitates timely prevention of impaired foetal growth. However, available techniques such as ultrasonography have limitations, including less accuracy when applied before 20 weeks of gestation and operator-dependent variability. Existing BW prediction models often neglect nutritional and genetic influences, and focus mainly on physiological and lifestyle factors. This study presents an attention-based transformer model with a multi-encoder architecture for early ($< 12$ weeks) BW prediction. Our model effectively integrates diverse maternal data, including physiological, lifestyle, nutritional, and genetic data, addressing limitations seen in previous attention-based models such as TabNet. The model achieves a Mean Absolute Error (MAE) of 122 grams and an $R^{2}$ value of 0.94, showing its high predictive accuracy and interoperability with our in-house private dataset. Independent validation confirms generalizability (MAE: 105 grams, $R^{2}$: 0.95) with the IEEE children dataset. To enhance clinical utility, predicted BW is classified into low and normal categories, achieving a sensitivity of 97.55% and a specificity of 94.48%, facilitating early risk stratification. Model interpretability is reinforced through feature importance and SHAP analysis, highlighting significant influences of maternal age, tobacco exposure, and vitamin B12 status, with genetic factors playing a secondary role. Our results emphasize the potential of advanced deep learning models to improve early BW prediction, offering a robust, interpretable, and personalized tool to identify pregnancies at risk and optimize neonatal outcomes. Muhammad Mursil, Hatem A. Rashwan, Luis Santos-Calderon, Pere Cavallé-Busquets, Michelle M. Murphy, Domenec Puig |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Transformer-Based Bi-encoder for Early Neonatal Birth Weight Prediction Using Maternal Nutritional and Health Insights
Muhammad Mursil, Hatem A. Rashwan, Adnan Khalid 0001, Luis Santos-Calderon, Pere Cavallé-Busquets, Michelle M. Murphy, Domenec Puig |
AIME (2) | 7 |
| 2025 | Unsupervised Domain Adaptation with Contrastive Learning for Classifying Breast Lesions in MammogramsabstractDeep learning models enhance breast cancer detection in mammograms but struggle with domain shifts, where test data differ from training data. Domain adaptation (DA) helps address this issue but often relies on unstable adversarial techniques. Breast lesion classification in mammograms also faces challenges like data scarcity and overfitting. Mixup mitigates these by generating synthetic samples, increasing variability, and improving robustness. Meanwhile, contrastive learning enhances feature alignment, boosting generalization and classification accuracy across domains. This paper proposes a DA model that integrates mixup and contrastive learning to improve feature alignment and generalization, leading to more accurate breast lesion classification. Our approach outperforms standard DA methods, achieving 82.5% accuracy, 0.774 F1 score, and 0.7868 AUC on INbreast (target dataset), surpassing DANN (63.6%) and Deep CORAL (67.7%). It also generalizes well, reaching 70.4% accuracy on CMMD and 63.64% on CDD-CESM, demonstrating its effectiveness in addressing domain shifts. Mariam M. Hassan, Mohamed Ragab 0002, Mohamed Abdel-Nasser, Domenec Puig |
CBMS | 4 |
| 2025 | Correction to: Implicit regularization of a deep augmented neural network model for human motion prediction
Gaurav Kumar Yadav, Mohamed Abdel-Nasser, Hatem A. Rashwan, Domenec Puig, Gora Chand Nandi |
Appl. Intell. | 4 |
| 2025 | Efficient crack segmentation with multi-decoder networks and enhanced feature fusionabstractIn infrastructure management, intelligent crack detection is vital, particularly for maintaining crucial elements such as road networks in urban areas. Detecting pavement defects promptly and accurately is essential for timely repairs and hazard prevention. However, this task is challenging due to factors, such as complex backgrounds, micro defects, diverse defect shapes and sizes, and class imbalance issues. Innovative approaches and advanced technologies are needed to address these challenges and effectively manage infrastructure complexities. In this study, we propose a novel framework for crack segmentation, called CrackMaster. CrackMaster utilizes advanced neural network architectures, leveraging the next generation of convolutional networks (ConvNeXt) as an encoder and dual decoders customized for distinct tasks. The first decoder adopts a self-supervised learning paradigm to reconstruct images, thereby enhancing feature extraction capabilities. Meanwhile, the second decoder combines deep labelling network for semantic image segmentation (Deeplabv3+) with a light deep neural network (LinkNet) to facilitate precise segmentation. Notably, we introduce an Enhanced Feature Fusion (EFF) block to improve features quality, enhancing information flow and context preservation, thus boosting segmentation performance. Experimental results conducted on three diverse datasets, including our in-house Road Crack Dataset (RCD), DeepCrack537, and Yang Crack Dataset (YCD) datasets, demonstrate the effectiveness of our framework achieving outstanding Intersection over Union scores (IoU) of 86.0%, 87.8%, and 76.9%, respectively, showing superior accuracy and robustness in crack segmentation tasks. These findings underscore the potential applicability of our framework in real-world infrastructure management scenarios. The code is publicly available at: https://github.com/AmmarOkran/CrackMaster . • CrackMaster Framework : ConvNeXt-based model with dual decoders for automated crack segmentation. • Image Reconstruction : Utilizes a self-supervised learning paradigm in the first decoder to boost feature extraction capabilities. • Precise Segmentation : Integrates Deeplabv3+ and LinkNet networks in the second decoder for accurate crack image segmentation. • Enhanced Feature Fusion : EFF block to improve feature quality, information flow, and context preservation. Ammar M. Okran, Hatem A. Rashwan, Adel Saleh, Domenec Puig |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | CoAtUNet: A symmetric encoder-decoder with hybrid transformers for semantic segmentation of breast ultrasound images
Nadeem Zaidkilani, Miguel Ángel García, Domenec Puig |
Neurocomputing | 3 |
| 2025 | CoHAtNet: An integrated convolutional-transformer architecture with hybrid self-attention for end-to-end camera localizationabstractCamera localization refers to the process of automatically determining the position and orientation of a camera within its 3D environment from the images it captures. Traditional camera localization methods often rely on Convolutional Neural Networks, which are effective at extracting local visual features but struggle to capture long-range dependencies critical for accurate localization. In contrast, Transformer-based approaches model global contextual relationships appropriately, although they often lack precision in fine-grained spatial representations. To bridge this gap, we introduce CoHAtNet, a novel Convolutional Hybrid-Attention Network that tightly integrates convolutional and self-attention mechanisms. Unlike previous hybrid models that stack convolutional and attention layers separately, CoHAtNet embeds local features extracted via Mobile Inverted Bottleneck Convolution blocks directly into the Value component of the self-attention mechanism of Transformers. This yields a hybrid self-attention block capable of dynamically capturing both local spatial detail and global semantic context within a single attention layer. Additionally, CoHAtNet enables modality-level fusion by processing RGB and depth data jointly in a unified pipeline, allowing the model to leverage complementary appearance and geometric cues throughout. Extensive evaluations have been conducted on two widely-used camera localization datasets: 7-Scenes (RGB-D) and Cambridge Landmarks (RGB). Experimental results show that CoHAtNet achieves state-of-the-art performance in both translation and orientation accuracy. These results highlight the effectiveness of our hybrid design in challenging indoor and outdoor environments. This makes CoHAtNet a strong candidate for end-to-end camera localization tasks. Hussein Hasan, Miguel Ángel García, Hatem A. Rashwan, Domenec Puig |
Image Vis. Comput. | 4 |
| 2025 | Interpretable deep neural networks for advancing early neonatal birth weight prediction using multimodal maternal factorsabstractBACKGROUND: Neonatal low birth weight (LBW) is a significant predictor of increased morbidity and mortality among newborns. Predominantly, traditional prediction methods depend heavily on ultrasonography, which does not consider risk factors affecting birth weight (BW). OBJECTIVE: This study introduces a robust deep neural network for a clinical decision-support system designed to early predict neonatal BW, using data available during early pregnancy, with enhanced precision. This innovative system incorporates a comprehensive array of maternal factors, placing particular emphasis on nutritional elements alongside physiological and lifestyle variables. METHODS: We employed and validated various traditional machine learning models as well as an interpretable deep learning model using the TabNet architecture, noted for its proficient handling of tabular data and high level of interpretability. The efficacy of these models was evaluated against extensive datasets that encompass a broad spectrum of maternal health indicators. RESULTS: The TabNet model exhibited outstanding predictive capabilities, achieving an accuracy of 96% and an area under the curve (AUC) of 0.96. Significantly, maternal vitamin B12 and folate status emerged as pivotal predictors of BW, emphasizing the crucial role of nutritional factors in influencing neonatal health outcomes. CONCLUSIONS: Our results demonstrate the substantial benefits of integrating multimodal maternal factors into predictive models for neonatal BW, markedly enhancing the precision over traditional AI methods. The developed decision-support system not only has a possible application in prenatal care but also provides actionable insights that can be leveraged to mitigate the risks associated with LBW, thereby improving clinical decision-making processes and outcomes. Muhammad Mursil, Hatem A. Rashwan, Adnan Khalid 0001, Pere Cavallé-Busquets, Luis Santos-Calderon, Michelle M. Murphy, Domenec Puig |
J. Biomed. Informatics | 7 |
| 2025 | Online dimensionality reduction through stacked generalization of spectral methods with deep networksabstractAbstract Analyzing large volumes of high-dimensional data poses significant challenges. Dimensionality reduction aims to reveal the most prominent properties of data by embedding them into a low-dimensional representation. Spectral dimensionality reduction methods using kernel matrices have been proven to yield optimal results. Online versions of those methods are desirable to incrementally project new data without recomputing the whole embedding from the complete dataset. In addition, integrating different spectral methods may have a synergistic effect. This paper presents an online dimensionality reduction method based on deep neural networks that integrates embeddings optimized by statistical approximation of neighborhoods and induced by different spectral methods through stacking ensemble learning. In particular, the proposed method first applies a self-supervised stage in order to train a set of deep encoders based on the embeddings induced by different spectral methods applied to a given input dataset. Those basis encoders are optimized and then integrated through a metamodel constituted by a fully connected network. A supervised and an unsupervised approach have been designed depending on whether the final aim is to enforce topological preservation or cluster induction. The proposed method has been experimentally validated on well-known image datasets and compared to some of the most relevant dimensionality reduction techniques by using widely-used quality measures. Juan C. Alvarado-Pérez, Miguel Ángel García, Domenec Puig |
Mach. Learn. | 3 |
| 2025 | A two-stage progressive deep segmentation network for tumor detection in breast ultrasound images
Nadeem Zaidkilani, Mohamed Abdel-Nasser, Miguel Ángel García, Domenec Puig |
Multim. Tools Appl. | 4 |
| 2024 | FGR-Net: Interpretable fundus image gradeability classification based on deep reconstruction learningabstractThe performance of diagnostic Computer-Aided Design (CAD) systems for retinal diseases depends on the quality of the retinal images being screened. Thus, many studies have been developed to evaluate and assess the quality of such retinal images. However, most of them did not investigate the relationship between the accuracy of the developed models and the quality of the visualization of interpretability methods for distinguishing between gradable and non-gradable retinal images. Consequently, this paper presents a novel framework called “FGR-Net” to automatically asses and interpret underlying fundus image quality by merging an autoencoder network with a classifier network. The FGR-Net model also provides an interpretable quality assessment through visualizations. In particular, FGR-Net uses a deep autoencoder to reconstruct the input image in order to extract the visual characteristics of the input fundus images based on self-supervised learning. The extracted features by the autoencoder are then fed into a deep classifier network to distinguish between gradable and ungradable fundus images. FGR-Net is evaluated with different interpretability methods, which indicates that the autoencoder is a key factor in forcing the classifier to focus on the relevant structures of the fundus images, such as the fovea, optic disc, and prominent blood vessels. Additionally, the interpretability methods can provide visual feedback for ophthalmologists to understand how our model evaluates the quality of fundus images. The experimental results showed the superiority of FGR-Net over the state-of-the-art quality assessment methods, with an accuracy of >89% and an F1-score of >87%. The code is publicly available at https://github.com/saifalkh/FGR-Net. Saif Khalid, Hatem A. Rashwan, Saddam Abdulwahab, Mohamed Abdel-Nasser, Facundo Manuel Quiroga, Domenec Puig |
Expert Syst. Appl. | 6 |
| 2024 | Designing an adaptive cost function for dynamic human pose predictions
Gaurav Kumar Yadav, Domenec Puig, Gora Chand Nandi |
Multim. Tools Appl. | 2 |
| 2024 | VISTA: vision improvement via split and reconstruct deep neural network for fundus image quality assessmentabstractAbstract Widespread eye conditions such as cataracts, diabetic retinopathy, and glaucoma impact people worldwide. Ophthalmology uses fundus photography for diagnosing these retinal disorders, but fundus images are prone to image quality challenges. Accurate diagnosis hinges on high-quality fundus images. Therefore, there is a need for image quality assessment methods to evaluate fundus images before diagnosis. Consequently, this paper introduces a deep learning model tailored for fundus images that supports large images. Our division method centres on preserving the original image’s high-resolution features while maintaining low computing and high accuracy. The proposed approach encompasses two fundamental components: an autoencoder model for input image reconstruction and image classification to classify the image quality based on the latent features extracted by the autoencoder, all performed at the original image size, without alteration, before reassembly for decoding networks. Through post hoc interpretability methods, we verified that our model focuses on key elements of fundus image quality. Additionally, an intrinsic interpretability module has been designed into the network that allows decomposing class scores into underlying concepts quality such as brightness or presence of anatomical structures. Experimental results in our model with EyeQ, a fundus image dataset with three categories (Good, Usable, and Rejected) demonstrate that our approach produces competitive outcomes compared to other deep learning-based methods with an overall accuracy of 0.9066, a precision of 0.8843, a recall of 0.8905, and an impressive F1-score of 0.8868. The code is publicly available at https://github.com/saifalkhaldiurv/VISTA_-Image-Quality-Assessment . Saif Khalid, Saddam Abdulwahab, Oscar Stanchi, Facundo Manuel Quiroga, Franco Ronchetti, Domenec Puig, Hatem A. Rashwan |
Neural Comput. Appl. | 6 |
| 2024 | Dual-Stream CoAtNet models for accurate breast ultrasound image segmentation
Nadeem Zaidkilani, Miguel Ángel García, Domenec Puig |
Neural Comput. Appl. | 3 |
| 2024 | 3D-IncNet: Head and Neck (H&N) Primary Tumors Segmentation and Survival PredictionabstractCancer begins when healthy cells change and grow out of control, forming a mass called a tumor. Head and neck (H&N) cancers usually develop in or around the head and neck, including the mouth (oral cavity), nose and sinuses, throat (pharynx), and voice box (larynx). 4% of all cancers are H&N cancers with a very low survival rate (a five-year survival rate of 64.7%). FDG-PET/CT imaging is often used for early diagnosis and staging of H&N tumors, thus improving these patients' survival rates. This work presents a novel 3D-Inception-Residual aided with 3D depth-wise convolution and squeeze and excitation block. We introduce a 3D depth-wise convolution-inception encoder consisting of an additional 3D squeeze and excitation block and a 3D depth-wise convolution-based residual learning decoder (3D-IncNet), which not only helps to recalibrate the channel-wise features but adaptively through explicit inter-dependencies modeling but also integrate the coarse and fine features resulting in accurate tumor segmentation. We further demonstrate the effectiveness of inception-residual encoder-decoder architecture in achieving better dice scores and the impact of depth-wise convolution in lowering the computational cost. We applied random forest for survival prediction on deep, clinical, and radiomics features. Experiments are conducted on the benchmark HECKTOR21 challenge, which showed significantly better performance by surpassing the state-of-the-artwork and achieved 0.836 and 0.811 concordance index and dice scores, respectively. We made the model and code publicly available. Abdul Qayyum 0002, Abdessalam Benzinou, Muhammad Imran Razzak, Moona Mazher, Thanh Thi Nguyen 0001, Domenec Puig, Fatemeh Vafaee |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Implicit regularization of a deep augmented neural network model for human motion predictionabstractAbstract Predicting human motion based on past observed motion is one of the challenging issues in computer vision and graphics. Existing research works are dealing with this issue by using discriminative models and showing the results for cases that follow a homogeneous distribution (in distribution) and not discussing the issues of the domain shift problem, where training and testing data follow a heterogeneous (out of distribution) problem, which is the reality when such models are used in practice. However, recent research proposed addressing domain shift issues by augmenting the discriminative model with a generative model and obtained better results. In the present investigation, we propose regularizing the extended network by inserting linear layers to minimize the rank of the latent space and train the entire end-to-end network. We regularize the network to strengthen the model to deal effectively with domain shift scenarios. Both training and testing data come from different distribution sets; to deal with this, we toughen our network by adding the extra linear layers to the network encoder. We tested our model with the benchmark datasets, CMU Motion Capture and Human3.6M, and proved that our model outperforms 14 OoD actions of H3.6M and 7 OoD actions of CMU MoCap in terms of the Euclidean distance calculated between predicted and ground truth joint angle values. Our average results of 14 OoD actions for short-term (80, 160, 320, 400) are 0.34, 0.6, 0.96, 1.07, and for CMU MoCap of 7 OoD actions for short-term and long term (80, 160, 320, 400, 1000) are 0.28, 0.45, 0.77, 0.89, 1.46. All these results are much better than the other state-of-the-art results. Gaurav Kumar Yadav, Mohamed Abdel-Nasser, Hatem A. Rashwan, Domenec Puig, Gora Chand Nandi |
Appl. Intell. | 4 |
| 2023 | GCNDepth: Self-supervised monocular depth estimation based on graph convolutional networkabstractDepth estimation is a challenging task of 3D reconstruction to enhance the accuracy sensing of environment awareness. This work brings a new solution with improvements, which increases the quantitative and qualitative understanding of depth maps compared to existing methods. Recently, convolutional neural networks (CNN) have demonstrated their extraordinary ability to estimate depth maps from monocular videos. However, traditional CNN does not support a topological structure, and they can work only on regular image regions with determined sizes and weights. On the other hand, graph convolutional networks (GCN) can handle the convolution of non-Euclidean data, and they can be applied to irregular image regions within a topological structure. Therefore, to preserve object geometric appearances and objects locations in the scene, in this work, we aim to exploit GCN for a self-supervised monocular depth estimation model. Our model consists of two parallel auto-encoder networks: the first is an auto-encoder that will depend on ResNet-50 and extract the feature from the input image and on multi-scale GCN to estimate the depth map. In turn, the second network will be used to estimate the ego-motion vector (i.e., 3D pose) between two consecutive frames based on ResNet-18. The estimated 3D pose and depth map will be used to construct the target image. A combination of loss functions related to photometric, reprojection, and smoothness is used to cope with bad depth prediction and preserve the discontinuities of the objects. Our method and performance are improved quantitatively and qualitatively. In particular, our method provided comparable and promising results with a high prediction accuracy of 89% on the publicly available KITTI dataset. Our method also offers 40% reduction in the number of trainable parameters compared to the state of the art solutions.In addition, we tested our trained model with Make3D dataset to evaluate the trained model on a new dataset with low resolution images. The source code is publicly available at (https://github.com/ArminMasoumian/GCNDepth.git) Armin Masoumian, Hatem A. Rashwan, Saddam Abdulwahab, Julián Cristiano, Muhammad Salman Asif, Domenec Puig |
Neurocomputing | 6 |
| 2023 | Fetal brain tissue annotation and segmentation challenge resultsabstractIn-utero fetal MRI is emerging as an important tool in the diagnosis and analysis of the developing human brain. Automatic segmentation of the developing fetal brain is a vital step in the quantitative analysis of prenatal neurodevelopment both in the research and clinical context. However, manual segmentation of cerebral structures is time-consuming and prone to error and inter-observer variability. Therefore, we organized the Fetal Tissue Annotation (FeTA) Challenge in 2021 in order to encourage the development of automatic segmentation algorithms on an international level. The challenge utilized FeTA Dataset, an open dataset of fetal brain MRI reconstructions segmented into seven different tissues (external cerebrospinal fluid, gray matter, white matter, ventricles, cerebellum, brainstem, deep gray matter). 20 international teams participated in this challenge, submitting a total of 21 algorithms for evaluation. In this paper, we provide a detailed analysis of the results from both a technical and clinical perspective. All participants relied on deep learning methods, mainly U-Nets, with some variability present in the network architecture, optimization, and image pre- and post-processing. The majority of teams used existing medical imaging deep learning frameworks. The main differences between the submissions were the fine tuning done during training, and the specific pre- and post-processing steps performed. The challenge results showed that almost all submissions performed similarly. Four of the top five teams used ensemble learning methods. However, one team's algorithm performed significantly superior to the other submissions, and consisted of an asymmetrical U-Net network architecture. This paper provides a first of its kind benchmark for future automatic multi-tissue segmentation algorithms for the developing human brain in utero. Kelly Payette, Hongwei Li 0004, Priscille de Dumast, Roxane Licandro, Md Mahfuzur Rahman Siddiquee, Daguang Xu, Andriy Myronenko, Yuchen Pei, Lisheng Wang, Juanying Xie, Huiquan Zhang, Guiming Dong, Hao Fu 0014, Guotai Wang, ZunHyan Rieu, Hyun Gi Kim, Davood Karimi, Ali Gholipour, Helena R. Torres, Bruno Oliveira 0002, João L. Vilaça, Netanell Avisdris, Ori Ben-Zvi, Dafna Ben-Bashat, Lucas Fidon, Michael Aertsen, Tom Vercauteren, Daniel Sobotka, Georg Langs, Mireia Alenyà, Maria Inmaculada Villanueva, Oscar Camara 0001, Bella Specktor-Fadida, Leo Joskowicz, Liao Weibin, Lv Yi, Xuesong Li 0003, Moona Mazher, Abdul Qayyum 0002, Domenec Puig, Hamza Kebiri, KuanLun Liao, YiXuan Wu, JinTai Chen, Yunzhi Xu, Lana Vasung, Bjoern Menze, Meritxell Bach Cuadra, András Jakab |
Medical Image Anal. | 45 |
| 2023 | Deep Learning Segmentation of the Right Ventricle in Cardiac MRI: The M&Ms ChallengeabstractIn recent years, several deep learning models have been proposed to accurately quantify and diagnose cardiac pathologies. These automated tools heavily rely on the accurate segmentation of cardiac structures in MRI images. However, segmentation of the right ventricle is challenging due to its highly complex shape and ill-defined borders. Hence, there is a need for new methods to handle such structure's geometrical and textural complexities, notably in the presence of pathologies such as Dilated Right Ventricle, Tricuspid Regurgitation, Arrhythmogenesis, Tetralogy of Fallot, and Inter-atrial Communication. The last MICCAI challenge on right ventricle segmentation was held in 2012 and included only 48 cases from a single clinical center. As part of the 12th Workshop on Statistical Atlases and Computational Models of the Heart (STACOM 2021), the M&Ms-2 challenge was organized to promote the interest of the research community around right ventricle segmentation in multi-disease, multi-view, and multi-center cardiac MRI. Three hundred sixty CMR cases, including short-axis and long-axis 4-chamber views, were collected from three Spanish hospitals using nine different scanners from three different vendors, and included a diverse set of right and left ventricle pathologies. The solutions provided by the participants show that nnU-Net achieved the best results overall. However, multi-view approaches were able to capture additional information, highlighting the need to integrate multiple cardiac diseases, views, scanners, and acquisition protocols to produce reliable automatic cardiac segmentation algorithms. Carlos Martín-Isla, Víctor M. Campello, Cristian Izquierdo, Kaisar Kushibar, Carla Sendra-Balcells, Polyxeni Gkontra, Alireza Sojoudi, Mitchell J. Fulton, Tewodros Weldebirhan Arega, Kumaradevan Punithakumar, Lei Li 0020, Xiaowu Sun, Yasmina Alkhalil, Di Liu 0003, Sana Jabbar, Sandro F. Queiros, Francesco Galati, Moona Mazher, Zheyao Gao, Marcel Beetz, Lennart Tautz, Christoforos Galazis, Marta Varela, Markus Hüllebrand, Vicente Grau, Xiahai Zhuang, Domenec Puig, Maria A. Zuluaga, Hassan Mohy-ud-Din, Dimitris N. Metaxas, Marcel Breeuwer, Rob J. van der Geest, Michelle Noga, Stéphanie Bricq, Mark Rentschler, Andrea Guala 0002, Steffen E. Petersen, Sergio Escalera, Jose Rodriguez-Palomares, Karim Lekadir |
IEEE J. Biomed. Health Informatics | 27 |
| 2022 | Effective Deep Learning-Based Ensemble Model for Road Crack DetectionabstractThis paper proposes an effective deep learning-based model for crack detection in images acquired by different acquisition systems (e.g., cameras mounted on vehicles and drones and smartphone cameras) in six countries. By utilizing successful training procedures and including multi-scale feature extraction models, the ensemble model is built using effective variations of the cutting-edge object detection technique, Yolov7. The top crack detection models are fused using the non-maximum suppression method to create the proposed ensemble model. The proposed crack detection model is trained and validated using the crowdsensing-based road damage detection challenge (CRDDC2022). With the test set, the proposed model produced an average F1 score of 0.65 with all leaderboards of CRDDC2022 (all countries, India, Japan, Norway, and the United States leaderboards). Our approach is ranked in the 5thposition in the CRDDC2022 challenge1. The source code is available at https://github.com/AmmarOkran/CRDD2022. Ammar M. Okran, Mohamed Abdel-Nasser, Hatem A. Rashwan, Domenec Puig |
IEEE Big Data | 4 |
| 2022 | Monocular depth map estimation based on a multi-scale deep architecture and curvilinear saliency feature boosting
Saddam Abdulwahab, Hatem A. Rashwan, Miguel Ángel García, Armin Masoumian, Domenec Puig |
Neural Comput. Appl. | 5 |
| 2022 | Efficient deep learning-based semantic mapping approach using monocular vision for resource-limited mobile robots
Raghav Narula, Hatem A. Rashwan, Mohamed Abdel-Nasser, Domenec Puig, Gora Chand Nandi |
Neural Comput. Appl. | 5 |
| 2021 | A color fusion model based on Markowitz portfolio optimization for optic disc segmentation in retinal images
José Escorcia-Gutierrez, Jordina Torrents-Barrena, A. Margarita R. Gamarra, Pedro Romero-Aroca, Aïda Valls, Domenec Puig |
Expert Syst. Appl. | 6 |
| 2021 | SLSNet: Skin lesion segmentation using a lightweight generative adversarial networkabstractThe determination of precise skin lesion boundaries in dermoscopic images using automated methods faces many challenges, most importantly, the presence of hair, inconspicuous lesion edges and low contrast in dermoscopic images, and variability in the color, texture and shapes of skin lesions. Existing deep learning-based skin lesion segmentation algorithms are expensive in terms of computational time and memory. Consequently, running such segmentation algorithms requires a powerful GPU and high bandwidth memory, which are not available in dermoscopy devices. Thus, this article aims to achieve precise skin lesion segmentation with minimum resources: a lightweight, efficient generative adversarial network (GAN) model called SLSNet, which combines 1-D kernel factorized networks, position and channel attention, and multiscale aggregation mechanisms with a GAN model. The 1-D kernel factorized network reduces the computational cost of 2D filtering. The position and channel attention modules enhance the discriminative ability between the lesion and non-lesion feature representations in spatial and channel dimensions, respectively. A multiscale block is also used to aggregate the coarse-to-fine features of input skin images and reduce the effect of the artifacts. SLSNet is evaluated on two publicly available datasets: ISBI 2017 and the ISIC 2018. Although SLSNet has only 2.35 million parameters, the experimental results demonstrate that it achieves segmentation results on a par with the state-of-the-art skin lesion segmentation methods with an accuracy of 97.61%, and Dice and Jaccard similarity coefficients of 90.63% and 81.98%, respectively. SLSNet can run at more than 110 frames per second (FPS) in a single GTX1080Ti GPU, which is faster than well-known deep learning-based image segmentation models, such as FCN. Therefore, SLSNet can be used for practical dermoscopic applications. Md. Mostafa Kamal Sarker, Hatem A. Rashwan, Farhan Akram, Vivek Kumar Singh 0008, Syeda Furruka Banu, Forhad U. H. Chowdhury, Kabir Ahmed Choudhury, Sylvie Chambon, Petia Radeva, Domenec Puig, Mohamed Abdel-Nasser |
Expert Syst. Appl. | 10 |
| 2021 | Feature Extraction and Selection for Emotion Recognition from Electrodermal ActivityabstractElectrodermal activity (EDA) is indicative of psychological processes related to human cognition and emotions. Previous research has studied many methods for extracting EDA features; however, their appropriateness for emotion recognition has been tested using a small number of distinct feature sets and on different, usually small, data sets. In the current research, we reviewed 25 studies and implemented 40 different EDA features across time, frequency and time-frequency domains on the publicly available AMIGOS dataset. We performed a systematic comparison of these EDA features using three feature selection methods, Joint Mutual Information (JMI), Conditional Mutual Information Maximization (CMIM) and Double Input Symmetrical Relevance (DISR) and machine learning techniques. We found that approximately the same numbers of features are required to obtain the optimal accuracy for the arousal recognition and the valence recognition. Also, the subject-dependent classification results were significantly higher than the subject-independent classification for both arousal and valence recognition. Statistical features related to the Mel-Frequency Cepstral Coefficients (MFCC) were explored for the first time for the emotion recognition from EDA signals and they outperformed all other feature groups, including the most commonly used Skin Conductance Response (SCR) related features. Jainendra Shukla, Miguel Barreda-Ángeles, Joan Oliver, Gora Chand Nandi, Domenec Puig |
IEEE Trans. Affect. Comput. | 5 |
| 2020 | Breast tumor segmentation in ultrasound images using contextual-information-aware deep adversarial learning framework
Vivek Kumar Singh 0008, Mohamed Abdel-Nasser, Farhan Akram, Hatem A. Rashwan, Md. Mostafa Kamal Sarker, Nidhi Pandey, Santiago Romaní, Domenec Puig |
Expert Syst. Appl. | 8 |
| 2020 | Breast tumor segmentation and shape classification in mammograms using generative adversarial and convolutional neural network
Vivek Kumar Singh 0008, Hatem A. Rashwan, Santiago Romaní, Farhan Akram, Nidhi Pandey, Md. Mostafa Kamal Sarker, Adel Saleh, Meritxell Arenas, Miguel Arquez, Domenec Puig, Jordina Torrents-Barrena |
Expert Syst. Appl. | 10 |
| 2020 | A deep learning interpretable classifier for diabetic retinopathy disease grading
Jordi de La Torre, Aïda Valls, Domenec Puig |
Neurocomputing | 3 |
| 2020 | Action representation and recognition through temporal co-occurrence of flow fields and convolutional neural networks
Hatem A. Rashwan, Miguel Ángel García, Saddam Abdulwahab, Domenec Puig |
Multim. Tools Appl. | 4 |
| 2020 | Adversarial Learning for Depth and Viewpoint Estimation From a Single ImageabstractEstimating a depth map and, at the same time, predicting the 3D pose of an object from a single 2D color image is a very challenging task. Depth estimation is typically performed through stereo vision by following several time-consuming stages, such as epipolar geometry, rectification and matching. Alternatively, when stereo vision is not useful or applicable, depth relations can be inferred from a single image as studied in this paper. More precisely, deep learning is applied in order to solve the problem of estimating a depth map from a single image. Then, that map is used for predicting the 3D pose of the main object depicted in the image. The proposed model consists of two successive neural networks. The first network is based on a Generative Adversarial Neural network (GAN). It estimates a dense depth map from the given color image. A Convolutional Neural Network (CNN) is then used to predict the 3D pose from the generated depth map through regression. The main difficulty to jointly estimate depth maps and 3D poses using deep networks is the lack of training data with both depth and viewpoint annotations. This contribution assumes a cross-domain training procedure with 3D CAD models corresponding to objects appearing in real images in order to render depth images from different viewpoints. These rendered images are then used to guide the GAN network to learn the mapping from the image domain to the depth domain. By exploiting the dataset as a source of training data, the proposed model outperforms state-of-the-art models on the PASCAL 3D+ dataset. The code of the proposed model is publicly available at https://github.com/SaddamAbdulrhman/Depth-and-Viewpoint-Estimation/tree/master. Saddam Abdulwahab, Hatem A. Rashwan, Miguel Ángel García, Mohammed Jabreel, Sylvie Chambon, Domenec Puig |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Hierarchical Approach to Classify Food Scenes in Egocentric Photo-StreamsabstractRecent studies have shown that the environment where people eat can affect their nutritional behavior [1]. In this paper, we provide automatic tools for personalized analysis of a person's health habits by the examination of daily recorded egocentric photo-streams. Specifically, we propose a new automatic approach for the classification of food-related environments, that is able to classify up to 15 such scenes. In this way, people can monitor the context around their food intake in order to get an objective insight into their daily eating routine. We propose a model that classifies food-related scenes organized in a semantic hierarchy. Additionally, we present and make available a new egocentric dataset composed of more than 33 000 images recorded by a wearable camera, over which our proposed model has been tested. Our approach obtains an accuracy and F-score of 56% and 65%, respectively, clearly outperforming the baseline methods. Estefanía Talavera, Maria Leyva-Vallina, Md. Mostafa Kamal Sarker, Domenec Puig, Nicolai Petkov, Petia Radeva |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | Stakeholders Acceptance and Expectations of Robot-Assisted Therapy for Children with Autism Spectrum DisorderabstractRobot assisted therapy for children with Autism Spectrum Disorder (ASD) should take into account the stakeholders expectations about their potential benefits. Any disparity between the stakeholders expectations and the gained benefits may negatively impact the acceptance and adoption of the robot assisted therapy. In this research, we conducted an observational study with eleven parents and five clinical professionals related with the children with ASD who were preselected to undergo robot assisted therapeutic sessions. The aim was to investigate and identify the potential impact regarding the interventions delivered by the social robots during the interventions, roles of the social robots and benefit offered by them. Specifically, the social robot Cozmo was used for this study. Their opinions were collected using questionnaires and were analyzed quantitatively and qualitatively.The results of the study confirm a positive attitude towards the adoption of these technologies, both among the caregivers and the professionals. Joan Oliver, Rebeca Oliván, Jainendra Shukla, Annabel Folch, Rafael Martínez-Leal, Mireia Castellá, Domenec Puig |
RO-MAN | 7 |
| 2019 | Gait representation and recognition from temporal co-occurrence of flow fields
Hatem A. Rashwan, Miguel Ángel García, Sylvie Chambon, Domenec Puig |
Mach. Vis. Appl. | 4 |
| 2018 | Comparative Study of the Behavior of Feature Reduction Methods in Person Re-identification TaskabstractOne of the goals of person re-identification systems is to support video-surveillance operators and forensic investigators to find an individual of interest in videos acquired by a network of non-overlapping cameras.This is attained by sorting images of previously observed individuals for decreasing values of their similarity with a given probe individual.Existing appearance descriptors, together with their similarity measures, are mostly aimed at improving ranking quality.Many of these descriptors generate a high feature vector represented as an image signature.To tackle person re-identification in real-world scenario the processing time will be crucial, so an individual of interest within a network camera should be found out swiftly.We therefore study some feature reduction methods to achieve a significant trade-off between processing time and ranking quality.Although, observing some redundancies on the generated patterns of a given descriptor are not deniable, we suggest to employ a feature reduction method before use of it in real-world scenarios.In particular, we have tested three reduction methods: PCA, KPCA, and Isomap.We then evaluate our study on two benchmark data sets (VIPeR, and i-LIDS), by using two state-of-the-art descriptors on person re-identification task.The results presented in this paper, after applying the feature reduction step, are very promising in terms of recognition rate. Bahram Lavi Sefidgari, Mehdi Fatan, Domenec Puig |
ICPRAM | 3 |
| 2018 | SLSDeep: Skin Lesion Segmentation Based on Dilated Residual and Pyramid Pooling Networks
Md. Mostafa Kamal Sarker, Hatem A. Rashwan, Farhan Akram, Syeda Furruka Banu, Adel Saleh, Vivek Kumar Singh 0008, Forhad U. H. Chowdhury, Saddam Abdulwahab, Santiago Romaní, Petia Radeva, Domenec Puig |
MICCAI (2) | 11 |
| 2018 | Conditional Generative Adversarial and Convolutional Networks for X-ray Breast Mass Segmentation and Shape Classification
Vivek Kumar Singh 0008, Santiago Romaní, Hatem A. Rashwan, Farhan Akram, Nidhi Pandey, Md. Mostafa Kamal Sarker, Saddam Abdulwahab, Jordina Torrents-Barrena, Adel Saleh, Miguel Arquez, Meritxell Arenas, Domenec Puig |
MICCAI (2) | 12 |
| 2018 | An optimized convolutional neural network with bottleneck and spatial pyramid pooling layers for classification of foods
Elnaz Jahani Heravi, Hamed H. Aghdam, Domenec Puig |
Pattern Recognit. Lett. | 3 |
| 2018 | Aggregating the temporal coherent descriptors in videos using multiple learning kernel for action recognition
Adel Saleh, Mohamed Abdel-Nasser, Miguel Ángel García, Domenec Puig |
Pattern Recognit. Lett. | 4 |
| 2018 | Weighted kappa loss function for multi-class classification of ordinal data in deep learning
Jordi de La Torre, Domenec Puig, Aïda Valls |
Pattern Recognit. Lett. | 2 |
| 2017 | Effectiveness of socially assistive robotics during cognitive stimulation interventions: Impact on caregiversabstractExecution of cognitive stimulation interventions for cognitive training of individuals in need represents significant burden on caregivers in time and labor costs. Recent advancements in Socially Assistive Robotics (SAR) research can be exploited to reduce caregivers burden by work sharing with robots and supplementing/complementing human resources in execution of interventions. Current research evaluates the effectiveness of the SAR empowered cognitive training activity of Bingo Musical among thirty individuals with ID in multi-center trials. A multidimensional evaluation of caregivers workload was conducted; including subjective workload, time spent on users personalized interventions, and qualitative interviews with caregivers. The results of the research confirm a significant reduction in caregivers burden and raise a concern about the need of a specific training of the caregivers to take maximum advantage of SAR in health care. Jainendra Shukla, Miguel Barreda-Ángeles, Joan Oliver, Domenec Puig |
RO-MAN | 4 |
| 2017 | Breast tumor classification in ultrasound images using texture analysis and super-resolution methods
Mohamed Abdel-Nasser, Jaime Melendez, Antonio Moreno, Osama Ahmed Omer, Domenec Puig |
Eng. Appl. Artif. Intell. | 5 |
| 2017 | A Practical and Highly Optimized Convolutional Neural Network for Classifying Traffic Signs in Real-Time
Hamed H. Aghdam, Elnaz Jahani Heravi, Domenec Puig |
Int. J. Comput. Vis. | 3 |
| 2017 | A Closed-Form Focus Profile Model for Conventional Digital Cameras
Said Pertuz, Miguel Ángel García, Domenec Puig, Henry Arguello |
Int. J. Comput. Vis. | 3 |
| 2017 | Analyzing the evolution of breast tumors through flow fields and strain tensors
Mohamed Abdel-Nasser, Antonio Moreno, Hatem A. Rashwan, Domenec Puig |
Pattern Recognit. Lett. | 4 |
| 2016 | Classification of foods by transferring knowledge from ImageNet datasetabstractAutomatic classification of foods is a way to control food intake and tackle with obesity. However, it is a challenging problem since foods are highly deformable and complex objects. Results on ImageNet dataset have revealed that Convolutional Neural Network has a great expressive power to model natural objects. Nonetheless, it is not trivial to train a ConvNet from scratch for classification of foods. This is due to the fact that ConvNets require large datasets and to our knowledge there is not a large public dataset of food for this purpose. Alternative solution is to transfer knowledge from trained ConvNets to the domain of foods. In this work, we study how transferable are state-of-art ConvNets to the task of food classification. We also propose a method for transferring knowledge from a bigger ConvNet to a smaller ConvNet by keeping its accuracy similar to the bigger ConvNet. Our experiments on UECFood256 datasets show that Googlenet, VGG and residual networks produce comparable results if we start transferring knowledge from appropriate layer. In addition, we show that our method is able to effectively transfer knowledge to the smaller ConvNet using unlabeled samples. Elnaz Jahani Heravi, Hamed H. Aghdam, Domenec Puig |
ICMV | 3 |
| 2016 | Automatic nipple detection in breast thermograms
Mohamed Abdel-Nasser, Adel Saleh, Antonio Moreno, Domenec Puig |
Expert Syst. Appl. | 4 |
| 2016 | Towards cost reduction of breast cancer diagnosis using mammography texture analysisabstractIn this paper we analyse the performance of various texture analysis methods for the purpose of reducing the number of false positives in breast cancer detection; as a result, the cost of breast cancer diagnosis would be reduced. We consider well-known methods such as local binary patterns, histogram of oriented gradients, co-occurrence matrix features and Gabor filters. Moreover, we propose the use of local directional number patterns as a new feature extraction method for breast mass detection. For each method, different classifiers are trained on the extracted features to predict the class of unknown instances. In order to improve the mass detection capability of each individual method, we use feature combination techniques and classifier majority voting. Some experiments were performed on the images obtained from a public breast cancer database, achieving promising levels of sensitivity and specificity. Mohamed Abdel-Nasser, Antonio Moreno, Domenec Puig |
J. Exp. Theor. Artif. Intell. | 3 |
| 2016 | Computer-aided diagnosis of breast cancer via Gabor wavelet bank and binary-class SVM in mammographic imagesabstractBreast cancer is one of the most dangerous diseases that attack women in their 40s worldwide. Due to this fact, it is estimated that one in eight women will develop a malignant carcinoma during their life. In addition, the carelessness of performing regular screenings is an important reason for the increase of mortality. However, computer-aided diagnosis systems attempt to enhance the quality of mammograms as well as the detection of early signs related to the disease. In this paper we propose a bank of Gabor filters to calculate the mean, standard deviation, skewness and kurtosis features by four-sized evaluation windows. Therefore, an active strategy is used to select the most relevant pixels. Finally, a supervised classification stage using two-class support vector machines is utilised through an accurate estimation of kernel parameters. In order to show the development of our methodology based on mammographic image analysis, two main experiments are fulfilled: abnormal/normal breast tissue classification and the ability to detect the different breast cancer types. Moreover, the public screen–film mini-MIAS database is compared with a digitised breast cancer database to evaluate the method robustness. The area under the receiver operating characteristic curve is used to measure the performance of the method. Furthermore, both confusion matrix and accuracy are calculated to assess the results of the proposed algorithm. Jordina Torrents-Barrena, Domenec Puig, Jaime Melendez, Aïda Valls |
J. Exp. Theor. Artif. Intell. | 2 |
| 2016 | Defeating face de-identification methods based on DCT-block scrambling
Hatem A. Rashwan, Miguel Ángel García, Antoni Martínez-Ballesté, Domenec Puig |
Mach. Vis. Appl. | 4 |
| 2015 | A New One Class Classifier Based on Ensemble of Binary Classifiers
Hamed H. Aghdam, Elnaz Jahani Heravi, Domenec Puig |
CAIP (2) | 3 |
| 2015 | Toward an optimal convolutional neural network for traffic sign recognitionabstractConvolutional Neural Networks (CNN) beat the human performance on German Traffic Sign Benchmark competition. Both the winner and the runner-up teams trained CNNs to recognize 43 traffic signs. However, both networks are not computationally efficient since they have many free parameters and they use highly computational activation functions. In this paper, we propose a new architecture that reduces the number of the parameters 27% and 22% compared with the two networks. Furthermore, our network uses Leaky Rectified Linear Units (ReLU) as the activation function that only needs a few operations to produce the result. Specifically, compared with the hyperbolic tangent and rectified sigmoid activation functions utilized in the two networks, Leaky ReLU needs only one multiplication operation which makes it computationally much more efficient than the two other functions. Our experiments on the Gertman Traffic Sign Benchmark dataset shows 0:6% improvement on the best reported classification accuracy while it reduces the overall number of parameters 85% compared with the winner network in the competition. Hamed H. Aghdam, Elnaz Jahani Heravi, Domenec Puig |
ICMV | 3 |
| 2015 | An unsupervised method for summarizing egocentric sport videosabstractPeople are getting more interested to record their sport activities using head-worn or hand-held cameras. This type of videos which is called egocentric sport videos has different motion and appearance patterns compared with life-logging videos. While a life-logging video can be defined in terms of well-defined human-object interactions, notwithstanding, it is not trivial to describe egocentric sport videos using well-defined activities. For this reason, summarizing egocentric sport videos based on human-object interaction might fail to produce meaningful results. In this paper, we propose an unsupervised method for summarizing egocentric videos by identifying the key-frames of the video. Our method utilizes both appearance and motion information and it automatically finds the number of the key-frames. Our blind user study on the new dataset collected from YouTube shows that in 93:5% cases, the users choose the proposed method as their first video summary choice. In addition, our method is within the top 2 choices of the users in 99% of studies. Hamed H. Aghdam, Elnaz Jahani Heravi, Domenec Puig |
ICMV | 3 |
| 2015 | Weighting video information into a multikernel SVM for human action recognitionabstractAction classification using a Bag of Words (BoW) representation has shown computational simplicity and good performance, but the increasing number of categories, including actions with high confusion, and the addition of significant contextual information has led most authors to focus their efforts on the combination of image descriptors. In this approach we code the action videos using a BoW representation with diverse image descriptors and introduce them to the optimal SVM kernel as a linear combination of learning weighted single kernels. Experiments have been carried out on the action database HMDB and the upturn achieved with our approach is much better than the state of the art, reaching an improvement of 14.63% of accuracy. Jordi Bautista-Ballester, Jaume Vergés-Llahí, Domenec Puig |
ICMV | 3 |
| 2015 | A deep convolutional neural network for recognizing foodsabstractControlling the food intake is an efficient way that each person can undertake to tackle the obesity problem in countries worldwide. This is achievable by developing a smartphone application that is able to recognize foods and compute their calories. State-of-art methods are chiefly based on hand-crafted feature extraction methods such as HOG and Gabor. Recent advances in large-scale object recognition datasets such as ImageNet have revealed that deep Convolutional Neural Networks (CNN) possess more representation power than the hand-crafted features. The main challenge with CNNs is to find the appropriate architecture for each problem. In this paper, we propose a deep CNN which consists of 769; 988 parameters. Our experiments show that the proposed CNN outperforms the state-of-art methods and improves the best result of traditional methods 17%. Moreover, using an ensemble of two CNNs that have been trained two different times, we are able to improve the classification performance 21:5%. Elnaz Jahani Heravi, Hamed H. Aghdam, Domenec Puig |
ICMV | 3 |
| 2015 | Focus-aided scene segmentation
Said Pertuz, Miguel Ángel García, Domenec Puig |
Comput. Vis. Image Underst. | 3 |
| 2015 | Analysis of tissue abnormality and breast density in mammographic images using a uniform local directional pattern
Mohamed Abdel-Nasser, Hatem A. Rashwan, Domenec Puig, Antonio Moreno |
Expert Syst. Appl. | 3 |
| 2015 | Efficient Focus Sampling Through Depth-of-Field Calibration
Said Pertuz, Miguel Ángel García, Domenec Puig |
Int. J. Comput. Vis. | 3 |
| 2015 | A k-anonymous approach to privacy preserving collaborative filtering
Fran Casino, Josep Domingo-Ferrer, Constantinos Patsakis, Domenec Puig, Agusti Solanas |
J. Comput. Syst. Sci. | 4 |
| 2014 | Adaptive Probabilistic Thresholding Method for Accurate Breast Region Segmentation in MammogramsabstractSegmentation of the breast region is usually the first step in the analysis of mammograms. Due to the non-uniformity of the background, breast segmentation presents several difficulties especially for film based mammograms. Our experimental results show that 50% of digitized film based mammograms in the mini-MIAS database do not have uniform intensity in the background. For this reason, applying a global thresholding method produces inaccurate results. In addition, finding the optimal global threshold value by only using histogram information requires a reliable objective function that characterizes the statistics of the background and the mammogram regions in the digitized mammograms. A second way to find the boundary of the breast consists in fitting a deformable model, such as snakes, on the mammogram. However, this method has three main shortcomings. First, the model must be initialized near the boundary. Second, using gradient information in the objective function can push the boundary toward the tissues inside the breast rather than the actual boundary. Third, in some mammograms the breast region is occluded by artifacts, such as labels, that have high gradient values on their boundary and cause the deformable model to be fitted on the artifact. To address these problems we propose a probabilistic adaptive thresholding method that uses texture information and its probability to find the most probable threshold values for specific parts of the mammogram. The experimental results on mini-MIAS database show that our proposed method outperforms the state-of-art methods and improves the accuracy at least 37% in comparison with the best results obtained by contour growing methods. Hamed H. Aghdam, Domenec Puig, Agusti Solanas |
ICPR | 2 |
| 2014 | A Novel Mammography Image Representation Framework with Application to Image RegistrationabstractX-ray mammography is a fundamental tool for breast cancer detection and diagnosis. A difficult problem arises when analyzing, integrating and comparing the information from different mammograms due to intensity changes and distortions induced by breast deformations. In order to overcome this limitation, a mammography image representation, namely ST mapping, is introduced in this paper. The proposed method consists of mapping the image intensities according to a curvilinear coordinate system that adapts to the breast geometry in order to yield a deformation-robust representation of the image features. In a practical application, the ST mapping is exploited for performing image registration. To our knowledge, this approach is completely novel since it does not require neither computing global or local geometric transformations nor finding point correspondences between images. In contrast, the registration is performed only based on the breast contour. Experiments with synthetic image deformations of real mammography images are provided in order to show the robustness of the proposed method to general deformations. Said Pertuz, Carme Julià, Domenec Puig |
ICPR | 3 |
| 2014 | Illumination-Robust Optical Flow Using a Local Directional PatternabstractMost of the variational optical flow methods are based on the well-known brightness constancy assumption or high-order constancy assumptions to implement the data term in the optimization energy function. Unfortunately, any variation in the lighting within the scene violates the brightness constancy constraint; in turn, the gradient constancy assumption does not work properly with large illumination changes. This paper proposes an illumination-robust constancy based on a robust texture descriptor rather than the brightness constancy. Thus, the similarity function used as a data term was obtained from extracting texture features through the local directional pattern descriptor for two consecutive frames within the duality total variational optical flow algorithm. In addition, a weighted nonlocal term that depends on both the color similarity and the occlusion state of pixels is integrated during the optimization process to increase the accuracy of the resulting flow field. The experimental results show a qualitative comparison with the proposed approach and yield state-of-the-art results on the KITTI, Midleburry, and MPI-sintel data sets. Mahmoud A. Mohamed, Hatem A. Rashwan, Bärbel Mertsching, Miguel Ángel García, Domenec Puig |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2013 | Reliability measure for shape-from-focus
Said Pertuz, Domenec Puig, Miguel Ángel García |
Image Vis. Comput. | 2 |
| 2013 | Analysis of focus measure operators for shape-from-focus
Said Pertuz, Domenec Puig, Miguel Ángel García |
Pattern Recognit. | 2 |
| 2013 | Generation of All-in-Focus Images by Noise-Robust Selective Fusion of Limited Depth-of-Field ImagesabstractThe limited depth-of-field of some cameras prevents them from capturing perfectly focused images when the imaged scene covers a large distance range. In order to compensate for this problem, image fusion has been exploited for combining images captured with different camera settings, thus yielding a higher quality all-in-focus image. Since most current approaches for image fusion rely on maximizing the spatial frequency of the composed image, the fusion process is sensitive to noise. In this paper, a new algorithm for computing the all-in-focus image from a sequence of images captured with a low depth-of-field camera is presented. The proposed approach adaptively fuses the different frames of the focus sequence in order to reduce noise while preserving image features. The algorithm consists of three stages: 1) focus measure; 2) selectivity measure; 3) and image fusion. An extensive set of experimental tests has been carried out in order to compare the proposed algorithm with state-of-the-art all-in-focus methods using both synthetic and real sequences. The obtained results show the advantages of the proposed scheme even for high levels of noise. Said Pertuz, Domenec Puig, Miguel Ángel García, Andrea Fusiello |
IEEE Trans. Image Process. | 2 |
| 2013 | Variational Optical Flow Estimation Based on Stick Tensor VotingabstractVariational optical flow techniques allow the estimation of flow fields from spatio-temporal derivatives. They are based on minimizing a functional that contains a data term and a regularization term. Recently, numerous approaches have been presented for improving the accuracy of the estimated flow fields. Among them, tensor voting has been shown to be particularly effective in the preservation of flow discontinuities. This paper presents an adaptation of the data term by using anisotropic stick tensor voting in order to gain robustness against noise and outliers with significantly lower computational cost than (full) tensor voting. In addition, an anisotropic complementary smoothness term depending on directional information estimated through stick tensor voting is utilized in order to preserve discontinuity capabilities of the estimated flow fields. Finally, a weighted non-local term that depends on both the estimated directional information and the occlusion state of pixels is integrated during the optimization process in order to denoise the final flow field. The proposed approach yields state-of-the-art results on the Middlebury benchmark. Hatem A. Rashwan, Miguel Ángel García, Domenec Puig |
IEEE Trans. Image Process. | 3 |
| 2012 | Improving the robustness of variational optical flow through tensor voting
Hatem A. Rashwan, Domenec Puig, Miguel Ángel García |
Comput. Vis. Image Underst. | 2 |
| 2011 | Supervised texture segmentation through a multi-level pixel-based classifier based on specifically designed filtersabstractThis paper presents a new, efficient technique for supervised texture segmentation based on a set of specifically designed filters and a multi-level pixel-based classifier. Filter design is carried out by means of a neural network, which is trained to maximize the filters' discrimination power among the texture classes under consideration. Texture features obtained with these filters are then processed by a classification scheme that utilizes multiple evaluation window sizes following a top-down approach, which iteratively refines the resulting segmentation. The proposed technique is compared to previous supervised texture segmenters by using both synthetic compositions and real outdoor textured images. Jaime Melendez, Xavier Gironés, Domenec Puig |
ICIP | 3 |
| 2011 | Shape-based image segmentation through photometric stereo
Carme Julià, Rodrigo Moreno, Domenec Puig, Miguel Ángel García |
Comput. Vis. Image Underst. | 3 |
| 2011 | Unsupervised texture-based image segmentation through pattern discovery
Jaime Melendez, Miguel Ángel García, Domenec Puig, Maria Petrou |
Comput. Vis. Image Underst. | 3 |
| 2011 | Edge-preserving color image denoising through tensor voting
Rodrigo Moreno, Miguel Ángel García, Domenec Puig, Carme Julià |
Comput. Vis. Image Underst. | 3 |
| 2011 | On Improving the Efficiency of Tensor VotingabstractThis paper proposes two alternative formulations to reduce the high computational complexity of tensor voting, a robust perceptual grouping technique used to extract salient information from noisy data. The first scheme consists of numerical approximations of the votes, which have been derived from an in-depth analysis of the plate and ball voting processes. The second scheme simplifies the formulation while keeping the same perceptual meaning of the original tensor voting: The stick tensor voting and the stick component of the plate tensor voting must reinforce surfaceness, the plate components of both the plate and ball tensor voting must boost curveness, whereas junctionness must be strengthened by the ball component of the ball tensor voting. Two new parameters have been proposed for the second formulation in order to control the potentially conflictive influence of the stick component of the plate vote and the ball component of the ball vote. Results show that the proposed formulations can be used in applications where efficiency is an issue since they have a complexity of order O(1). Moreover, the second proposed formulation has been shown to be more appropriate than the original tensor voting for estimating saliencies by appropriately setting the two new parameters. Rodrigo Moreno, Miguel Ángel García, Domenec Puig, Luis Pizarro, Bernhard Burgeth, Joachim Weickert |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | On Adapting Pixel-based Classification to Unsupervised Texture SegmentationabstractAn inherent problem of unsupervised texture segmentation is the absence of previous knowledge regarding the texture patterns present in the images to be segmented. A new efficient methodology for unsupervised image segmentation based on texture is proposed. It takes advantage of a supervised pixel-based texture classifier trained with feature vectors associated with a set of texture patterns initially extracted through a clustering algorithm. Therefore, the final segmentation is achieved by classifying each image pixel into one of the patterns obtained after the previous clustering process. Multi-sized evaluation windows following a top-down approach are applied during pixel classification in order to improve accuracy. The proposed technique has been experimentally validated on MeasTex, VisTex and Brodatz compositions, as well as on complex ground and aerial outdoor images. Comparisons with state-of the-art unsupervised texture segmenters are also provided. Jaime Melendez, Domenec Puig, Miguel Ángel García |
ICPR | 2 |
| 2010 | Robust Color Image Segmentation through Tensor VotingabstractThis paper presents a new method for robust color image segmentation based on tensor voting, a robust perceptual grouping technique used to extract salient information from noisy data. First, an adaptation of tensor voting to both image denoising and robust edge detection is applied. Second, pixels in the filtered image are classified into likely-homogeneous and likely-inhomogeneous by means of the edginess maps generated in the first step. Third, the likely-homosgeneous pixels are segmented through an efficient graph-based segmenter. Finally, a modified version of the same graph-based segmenter is applied to the likely-inhomogeneous pixels in order to obtain the final segmentation. Experiments show that the proposed algorithm has a better performance than the state-of-the-art. Rodrigo Moreno, Miguel Ángel García, Domenec Puig |
ICPR | 3 |
| 2010 | Improving Shape-from-Focus by Compensating for Image Magnification ShiftabstractImages taken with different focus settings are used in shape-from-focus to reconstruct the depth map of a scene. A problem when acquiring images with different focus settings is the shift of image features due to changes in magnification. This paper shows that those changes affect the shape-from-focus performance and that the final reconstruction can be improved by compensating for that shift. The proposed scheme takes into account the effects due to magnification changes between near and far focused images and it is able to determine the depth of the scene points with higher accuracy than traditional techniques. Experimental results of the application of the proposed method are shown. Said Pertuz, Domenec Puig, Miguel Ángel García |
ICPR | 2 |
| 2010 | Multi-level pixel-based texture classification through efficient prototype selection via normalized cut
Jaime Melendez, Domenec Puig, Miguel Ángel García |
Pattern Recognit. | 2 |
| 2010 | Application-independent feature selection for texture classification
Domenec Puig, Miguel Ángel García, Jaime Melendez |
Pattern Recognit. | 1 |
| 2009 | On Adapting the Tensor Voting Framework to Robust Color Image Denoising
Rodrigo Moreno, Miguel Ángel García, Domenec Puig, Carme Julià |
CAIP | 3 |
| 2009 | Gabor-based texture classification through efficient prototype selection via normalized cutabstractThis paper presents a new efficient technique for supervised pixel-based texture classification. The proposed scheme first performs a selection process that automatically determines a subset of prototypes that characterize each texture class based on the outcome of a multichannel Gabor wavelet filter bank. Then, every image pixel is classified into one of the given texture classes by using a K-NN classifier fed with the prototypes determined previously. The proposed technique is compared to previous texture classifiers by using both Brodatz and real outdoor textured images. Jaime Melendez, Domenec Puig, Miguel Ángel García |
ICIP | 2 |
| 2009 | Robust color edge detection through tensor votingabstractThis paper presents a new method for color edge detection based on the tensor voting framework, a robust perceptual grouping technique used to extract salient information from noisy data. The tensor voting framework is adapted to encode color information via tensors in order to propagate them into a neighborhood through a voting process specifically designed for color edge detection by taking into account perceptual color differences, region uniformity and edginess according to a set of intuitive perceptual criteria. Perceptual color differences are estimated by means of an optimized version of the CIEDE2000 formula, while uniformity and edginess are estimated by means of saliency maps obtained from the tensor voting process. Experiments show that the proposed algorithm is more robust and has a similar performance in precision when compared with the state-of-the-art. Rodrigo Moreno, Miguel Ángel García, Domenec Puig, Carme Julià |
ICIP | 3 |
| 2009 | A new methodology for evaluation of edge detectorsabstractThis paper defines a new methodology for evaluating edge detectors through measurements on edginess maps instead of on binary edge maps as previous methodologies do. These measurements avoid possible bias introduced by the application-dependent process of generating binary edge maps from edginess maps. The features of completeness, discriminability, precision and robustness, which a general-purpose edge detector must comply with, are introduced. The R, DS, P and FAR-measurements in addition to PSNR applied to the edginess maps are defined to assess the performance of edge detection. The R, DS, P and FAR-measurements can be seen as generalizations of previously proposed measurements on binary edge maps. Well-known and state-of-the-art edge detectors have been compared by means of the new proposed metrics. Results show that it is difficult for an edge detector to comply with all the proposed features. Rodrigo Moreno, Domenec Puig, Carme Julià, Miguel Ángel García |
ICIP | 2 |
| 2008 | Efficient distance-based per-pixel texture classification with Gabor wavelet filters
Jaime Melendez, Miguel Ángel García, Domenec Puig |
Pattern Anal. Appl. | 3 |
| 2007 | Comparative Evaluation of Classical Methods, Optimized Gabor Filters and LBP for Texture Feature Selection and Classification
Jaime Melendez, Domenec Puig, Miguel Ángel García |
CAIP | 2 |
| 2007 | Pixel-Based Texture Classification by Integration of Multiple Feature Extraction Methods Evaluated over Multisized WindowsabstractThis paper presents a pixel-based texture classifier oriented to the identification of texture models that can be present in an input image, given a set of models known in advance. The proposed methodology is based on the integration of texture features generated by texture methods that belong to different families, which are evaluated over multiple windows of different sizes. This is a novelty with respect to the current texture classifiers, which are based on specific families of texture methods evaluated over single windows of a size defined empirically. Experiments show that this integration strategy produces better results than classical texture classifiers based on specific families of texture methods. Domenec Puig, Miguel Ángel García |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2007 | Supervised texture classification by integration of multiple texture methods and evaluation windows
Miguel Ángel García, Domenec Puig |
Image Vis. Comput. | 2 |
| 2006 | Automatic texture feature selection for image pixel classification
Domenec Puig, Miguel Ángel García |
Pattern Recognit. | 1 |
| 2003 | Pixel classification through divergence-based integration of texture methods with conflict resolutionabstractThis paper presents a new technique for combining multiple texture feature extraction methods in order to classify the pixels of an input image into a set of texture models of interest. The problem of integrating multiple texture methods for classification purposes is cast as a collaborative decision making problem. Each texture method is considered to be an expert that gives an opinion about the membership of every input image pixel to each texture model, along with a conviction about that judgement. A conviction measure based on the Kullback J-divergence between texture models is proposed, along with an arbitration mechanism that combines those convictions by taking into account conflicts that may occur when different experts disagree with a similar strength. The proposed technique is compared to previous pixel-based texture classifiers by using real textured images. Domenec Puig, Miguel Ángel García |
ICIP (2) | 1 |
| 2002 | Recognizing specific texture patterns by integration of multiple texture methodsabstractA supervised pixel-based classifier for identifying the presence of a given set of texture patterns of interest in a complex textured image is described. The proposed technique integrates the outcome of multiple texture feature extraction methods belonging to different families. In this way, it yields lower classification rates than previous texture classifiers based on specific families of texture methods. Experimental results with real outdoor images are presented. Miguel Ángel García, Domenec Puig |
ICIP (1) | 2 |