EDBT 2026 Demo / reviewers in the wild / expert
Eduardo Fidalgo
dblp:94/8013 · also Eduardo F. Fidalgo, Eduardo Fidalgo Fernandez
· DBLP profile ↗
28ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0003-1202-5232ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 2 first-author · 15 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Age Estimation Robustness Under Realistic Facial Occlusions
Waqar Tanveer, Annalisa Franco, Guido Borghi, Laura Fernández-Robles, Eduardo Fidalgo |
ICPR (13) | 5 |
| 2026 | Building a multi-class Short Message Service dataset for smishing detection using agglomerative clustering and dataset fusion
Alicia Martínez-Mendoza, Eduardo Fidalgo, Enrique Alegre, Laura Fernández-Robles |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Underage detection through a multi-task and MultiAge approach for screening minors in unconstrained imagery
Christopher Gaul, Eduardo Fidalgo, Enrique Alegre, Rocío Alaíz-Rodríguez, Eri Pérez Corral |
Pattern Recognit. | 2 |
| 2026 | Overcoming occlusions in the wild: A multi-task age head approach to age estimationabstractFacial age estimation has achieved considerable success under controlled conditions. However, in unconstrained real-world scenarios, which are often referred to as ”in the wild”, age estimation remains challenging, especially when faces are partially occluded, which may obscure their visibility. To address this limitation, we propose a new approach integrating generative adversarial networks (GANs) and transformer architectures to enable robust age estimation from occluded faces. We employ an SN-Patch GAN to effectively remove occlusions, while an Attentive Residual Convolution Module (ARCM), paired with a Swin Transformer, enhances feature representation. Additionally, we introduce a Multi-Task Age Head (MTAH) that combines regression and distribution learning, further improving age estimation under occlusion. Experimental results on the FG-NET, UTKFace, and MORPH datasets demonstrate that our proposed approach surpasses existing state-of-the-art techniques for occluded facial age estimation by achieving an MAE of 3.00, 4.54, and 2.53 years, respectively. Waqar Tanveer, Laura Fernández-Robles, Eduardo Fidalgo, Víctor González-Castro, Enrique Alegre |
Pattern Recognit. | 3 |
| 2025 | Classifying the content of online notepad services using active learningabstractAbstract Pastebin is an online notepad service to share text anonymously. However, it could be misused to propagate suspicious or even illegal activities, like leaking sensitive information or sharing hyperlinks to child sexual abuse material. Due to the high rate of daily upload pastes, manual inspection of this material is not feasible. Conversely, an automatic classifier could identify such activities with little or no human intervention. However, a supervised model may require a significant number of training samples and have to handle distinct text typologies presented in Pastebin. This paper presents a classification approach composed of three cascading supervised classifiers that use Active Learning to select and label the most informative samples from Pastebin. The modularity of the proposed design allows each classifier to adapt to a specific text typology. The first classifier determines whether the text is a code snippet, and the second is to identify whether it is readable. The third classification level is twofold: (i) a binary classifier to say whether the text is suspicious and (ii) a multiclass classifier with seven predefined categories of possibly illegal activities. The average class recall of the binary and multiclass classifiers is $$95.24\%$$ 95.24 % and $$80.33\%$$ 80.33 % , respectively. Additionally, this paper presents a dataset of 3.8 million Pastebin samples, called onlIne Notepad Services PastEbin aCtiviTies (INSPECT-3.8M), along with their labels using our classification framework. Our classifier recognised that $$7.54\%$$ 7.54 % of the collected samples are correlated with presumably criminal activities. Law enforcement agencies may benefit from the insights shared in our research when aiming to investigate or automate the monitoring of Pastebin or other Online Notepad Services. This would allow responsible authorities to block illegal content before it spreads to the public. Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Sarah Jane Delany, Francisco Jáñez-Martino |
J. Intell. Inf. Syst. | 2 |
| 2025 | Improving audio embeddings with squeeze-and-excitation: Introducing SaEENet
Roberto Andrés Vasco Carofilis, Laura Fernández-Robles, Enrique Alegre, Eduardo Fidalgo |
Knowl. Based Syst. | 4 |
| 2025 | Spam email classification based on cybersecurity potential risk using natural language processing
Francisco Jáñez-Martino, Rocío Alaíz-Rodríguez, Víctor González-Castro, Eduardo Fidalgo, Enrique Alegre |
Knowl. Based Syst. | 4 |
| 2024 | DeepHSAR: Semi-supervised fine-grained learning for multi-label human sexual activity recognition
Abhishek Gangwar, Víctor González-Castro, Enrique Alegre, Eduardo Fidalgo, Alicia Martínez-Mendoza |
Inf. Process. Manag. | 4 |
| 2023 | Supervised ranking approach to identify infLuential websites in the darknetabstractAbstract The anonymity and high security of the Tor network allow it to host a significant amount of criminal activities. Some Tor domains attract more traffic than others, as they offer better products or services to their customers. Detecting the most influential domains in Tor can help detect serious criminal activities. Therefore, in this paper, we present a novel supervised ranking framework for detecting the most influential domains. Our approach represents each domain with 40 features extracted from five sources: text, named entities, HTML markup, network topology, and visual content to train the learning-to-rank (LtR) scheme to sort the domains based on user-defined criteria. We experimented on a subset of 290 manually ranked drug-related websites from Tor and obtained the following results. First, among the explored LtR schemes, the listwise approach outperforms the benchmarked methods with an NDCG of 0.93 for the top-10 ranked domains. Second, we quantitatively proved that our framework surpasses the link-based ranking techniques. Third, we observed that using the user-visible text feature can obtain comparable performance to all the features with a decrease of 0.02 at NDCG@5. The proposed framework might support law enforcement agencies in detecting the most influential domains related to possible suspicious activities. Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Deisy Chaves |
Appl. Intell. | 2 |
| 2023 | DeepSumm: Exploiting topic models and sequence to sequence networks for extractive text summarization
Akanksha Joshi, Eduardo Fidalgo, Enrique Alegre, Laura Fernández-Robles |
Expert Syst. Appl. | 2 |
| 2023 | Triple-BigGAN: Semi-supervised generative adversarial networks for image synthesis and classification on sexual facial expression recognition
Abhishek Gangwar, Víctor González-Castro, Enrique Alegre, Eduardo Fidalgo |
Neurocomputing | 4 |
| 2023 | Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation MappingabstractAutomatic accent classification is an active research field concerning speech processing. It can be useful to identify a speaker's region of origin, which can be applied in police investigations carried out by Law Enforcement Agencies, as well as for the improvement of current speech recognition systems. This paper presents a novel descriptor called Grad-Transfer, extracted using the Gradient-weighted Class Activation Mapping (Grad-CAM) method based on convolutional neural network (CNN) interpretability. Additionally, we propose a methodology for accent classification that implements Grad-Transfer, which is based on transferring the knowledge acquired by a CNN to a classical machine learning algorithm. The paper works on two hypotheses: the coarse localization maps produced by Grad-CAM on spectrograms are able to highlight the regions of the spectrograms that are important for predicting accents, and Grad-Transfer descriptors computed from audios represent distinctive descriptions of the target accents. These hypotheses were demonstrated experimentally, clustering the generated Grad-Transfer descriptors according to the original accent of the audios using Birch and$k$-means algorithms. We carried out experiments on the Voice Cloning Toolkit dataset, seeing an increase of macro average accuracy, and unweighted average recall in the results obtained by a Gaussian Naive Bayes classifier up to$23.00\%$, and$23.58\%$, respectively, compared to a model trained with spectrograms. This demonstrates that Grad-Transfer is able to improve the performance of accent classification models and opens the door to new implementations in similar tasks. Roberto Andrés Vasco Carofilis, Enrique Alegre, Eduardo Fidalgo, Laura Fernández-Robles |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Detecting malware using text documents extracted from spam email through machine learningabstractSpam has become an effective way for cybercriminals to spread malware. Although cybersecurity agencies and companies develop products and organise courses for people to detect malicious spam email patterns, spam attacks are not totally avoided yet. In this work, we present and make publicly available "Spam Email Malware Detection - 600" (SEMD-600), a new dataset, based on Bruce Guenter's, for malware detection in spam using only the text of the email. We also introduce a pipeline for malware detection based on traditional Natural Language Processing (NLP) techniques. Using SEMD-600, we compare the text representation techniques Bag of Words and Term Frequency-Inverse Document Frequency (TF-IDF), in combination with three different supervised classifiers: Support Vector Machine, Naive Bayes and Logistic Regression, to detect malware in plain text documents. We found that combining TF-IDF with Logistic Regression achieved the best performance, with a macro F1 score of 0.763. Luis Ángel Redondo-Gutierrez, Francisco Jáñez-Martino, Eduardo Fidalgo, Enrique Alegre, Víctor González-Castro, Rocío Alaíz-Rodríguez |
DocEng | 3 |
| 2022 | RankSum - An unsupervised extractive text summarization based on rank fusion
Akanksha Joshi, Eduardo Fidalgo, Enrique Alegre, Rocío Alaíz-Rodríguez |
Expert Syst. Appl. | 2 |
| 2022 | Phishing websites detection using a novel multipurpose dataset and web technologies featuresabstractPhishing attacks are one of the most challenging social engineering cyberattacks due to the large amount of entities involved in online transactions and services. In these attacks, criminals deceive users to hijack their credentials or sensitive data through a login form which replicates the original website and submits the data to a malicious server. Many anti-phishing techniques have been developed in recent years, using different resource such as the URL and HTML code from legitimate index websites and phishing ones. These techniques have some limitations when predicting legitimate login websites, since, usually, no login forms are present in the legitimate class used for training the proposed model. Hence, in this work we present a methodology for phishing website detection in real scenarios, which uses URL, HTML, and web technology features. Since there is not any updated and multipurpose dataset for this task, we crafted the Phishing Index Login Websites Dataset (PILWD), an offline phishing dataset composed of 134,000 verified samples, that offers to researchers a wide variety of data to test and compare their approaches. Since approximately three-quarters of collected phishing samples request the introduction of credentials, we decided to crawl legitimate login websites to match the phishing standpoint. The developed approach is independent of third party services and the method relies on a new set of features used for the very first time in this problem, some of them extracted from the web technologies used by the on each specific website. Experimental results show that phishing websites can be detected with 97.95% accuracy using a LightGBM classifier and the complete set of the 54 features selected, when it was evaluated on PILWD dataset. Manuel Sánchez-Paniagua, Eduardo Fidalgo, Enrique Alegre, Rocío Alaíz-Rodríguez |
Expert Syst. Appl. | 2 |
| 2022 | A survey on methods, datasets and implementations for scene text spottingabstractAbstract Text Spotting is the union of the tasks of detection and transcription of the text that is present in images. Due to the various problems often found when retrieving text, such as orientation, aspect ratio, vertical text or multiple languages in the same image, this can be a challenging task. In this paper, the most recent methods and publications in this field are analysed and compared. Apart from presenting features already seen in other surveys, such as their architectures and performance on different datasets, novel perspectives for comparison are also included, such as the hardware, software, backbone architectures, main problems to solve, or programming languages of the algorithms. The review highlights information often omitted in other studies, providing a better understanding of the current state of research in Text Spotting, from 2016 to 2022, current problems and future trends, as well as establishing a baseline for future methods development, comparison of results and serving as guideline for choosing the most appropriate method to solve a particular problem. Pablo Blanco-Medina, Eduardo Fidalgo, Enrique Alegre, Víctor González-Castro |
IET Image Process. | 2 |
| 2021 | Trustworthiness of spam email addresses using machine learningabstractCybercriminals have increasingly used spam email to send scams, phishing, malware and other frauds to organisations and people. They design sophisticated and contextualised emails to make them look trustworthy for users, being the sender addresses an essential part. Although cybersecurity agencies and companies develop products and organise courses for people to detect emails patterns, spam attacks are not totally avoided yet. Francisco Jáñez-Martino, Rocío Alaíz-Rodríguez, Víctor González-Castro, Eduardo Fidalgo |
DocEng | 4 |
| 2021 | AttM-CNN: Attention and metric learning based CNN for pornography, age and Child Sexual Abuse (CSA) Detection in images
Abhishek Gangwar, Víctor González-Castro, Enrique Alegre, Eduardo Fidalgo |
Neurocomputing | 4 |
| 2021 | A new perceptual hashing method for verification and identity classification of occluded faces
Rubel Biswas, Víctor González-Castro, Eduardo Fidalgo, Enrique Alegre |
Image Vis. Comput. | 3 |
| 2021 | Image retrieval based on texture using latent space representation of discrete Fourier transformed maps
Surajit Saikia, Laura Fernández-Robles, Enrique Alegre, Eduardo Fidalgo |
Neural Comput. Appl. | 4 |
| 2020 | File Name Classification Approach to Identify Child Sexual Abuse
Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Rocío Alaíz-Rodríguez |
ICPRAM | 2 |
| 2020 | Perceptual image hashing based on frequency dominant neighborhood structure applied to Tor domains recognition
Rubel Biswas, Víctor González-Castro, Eduardo Fidalgo, Enrique Alegre |
Neurocomputing | 3 |
| 2020 | Improving named entity recognition in noisy user-generated text with local distance neighbor feature
Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Laura Fernández-Robles |
Neurocomputing | 2 |
| 2019 | SummCoder: An unsupervised framework for extractive text summarization based on deep auto-encoders
Akanksha Joshi, Eduardo Fidalgo, Enrique Alegre, Laura Fernández-Robles |
Expert Syst. Appl. | 2 |
| 2019 | ToRank: Identifying the most influential suspicious domains in the Tor network
Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Laura Fernández-Robles |
Expert Syst. Appl. | 2 |
| 2018 | Boosting image classification through semantic attention filtering strategies
Eduardo Fidalgo, Enrique Alegre, Víctor González-Castro, Laura Fernández-Robles |
Pattern Recognit. Lett. | 1 |
| 2017 | Classifying Illegal Activities on Tor Network Based on Web Textual ContentsabstractMhd Wesam Al Nabki, Eduardo Fidalgo, Enrique Alegre, Ivan de Paz. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Ivan de Paz |
EACL (1) | 2 |
| 2016 | Compass radius estimation for improved image classification using Edge-SIFT
Eduardo Fidalgo, Enrique Alegre, Víctor González-Castro, Laura Fernández-Robles |
Neurocomputing | 1 |