EDBT 2026 Demo / reviewers in the wild / expert
Enrique Alegre
dblp:15/3114 · also Enrique Alegre-Gutiérrez
· DBLP profile ↗
42ranked-venue papers
2as first author
21since 2021 · last 2026
0000-0003-2081-774XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 1 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Building a multi-class Short Message Service dataset for smishing detection using agglomerative clustering and dataset fusion
Alicia Martínez-Mendoza, Eduardo Fidalgo, Enrique Alegre, Laura Fernández-Robles |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Underage detection through a multi-task and MultiAge approach for screening minors in unconstrained imagery
Christopher Gaul, Eduardo Fidalgo, Enrique Alegre, Rocío Alaíz-Rodríguez, Eri Pérez Corral |
Pattern Recognit. | 3 |
| 2026 | Overcoming occlusions in the wild: A multi-task age head approach to age estimationabstractFacial age estimation has achieved considerable success under controlled conditions. However, in unconstrained real-world scenarios, which are often referred to as ”in the wild”, age estimation remains challenging, especially when faces are partially occluded, which may obscure their visibility. To address this limitation, we propose a new approach integrating generative adversarial networks (GANs) and transformer architectures to enable robust age estimation from occluded faces. We employ an SN-Patch GAN to effectively remove occlusions, while an Attentive Residual Convolution Module (ARCM), paired with a Swin Transformer, enhances feature representation. Additionally, we introduce a Multi-Task Age Head (MTAH) that combines regression and distribution learning, further improving age estimation under occlusion. Experimental results on the FG-NET, UTKFace, and MORPH datasets demonstrate that our proposed approach surpasses existing state-of-the-art techniques for occluded facial age estimation by achieving an MAE of 3.00, 4.54, and 2.53 years, respectively. Waqar Tanveer, Laura Fernández-Robles, Eduardo Fidalgo, Víctor González-Castro, Enrique Alegre |
Pattern Recognit. | 5 |
| 2025 | Classifying the content of online notepad services using active learningabstractAbstract Pastebin is an online notepad service to share text anonymously. However, it could be misused to propagate suspicious or even illegal activities, like leaking sensitive information or sharing hyperlinks to child sexual abuse material. Due to the high rate of daily upload pastes, manual inspection of this material is not feasible. Conversely, an automatic classifier could identify such activities with little or no human intervention. However, a supervised model may require a significant number of training samples and have to handle distinct text typologies presented in Pastebin. This paper presents a classification approach composed of three cascading supervised classifiers that use Active Learning to select and label the most informative samples from Pastebin. The modularity of the proposed design allows each classifier to adapt to a specific text typology. The first classifier determines whether the text is a code snippet, and the second is to identify whether it is readable. The third classification level is twofold: (i) a binary classifier to say whether the text is suspicious and (ii) a multiclass classifier with seven predefined categories of possibly illegal activities. The average class recall of the binary and multiclass classifiers is $$95.24\%$$ 95.24 % and $$80.33\%$$ 80.33 % , respectively. Additionally, this paper presents a dataset of 3.8 million Pastebin samples, called onlIne Notepad Services PastEbin aCtiviTies (INSPECT-3.8M), along with their labels using our classification framework. Our classifier recognised that $$7.54\%$$ 7.54 % of the collected samples are correlated with presumably criminal activities. Law enforcement agencies may benefit from the insights shared in our research when aiming to investigate or automate the monitoring of Pastebin or other Online Notepad Services. This would allow responsible authorities to block illegal content before it spreads to the public. Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Sarah Jane Delany, Francisco Jáñez-Martino |
J. Intell. Inf. Syst. | 3 |
| 2025 | Improving audio embeddings with squeeze-and-excitation: Introducing SaEENet
Roberto Andrés Vasco Carofilis, Laura Fernández-Robles, Enrique Alegre, Eduardo Fidalgo |
Knowl. Based Syst. | 3 |
| 2025 | Spam email classification based on cybersecurity potential risk using natural language processing
Francisco Jáñez-Martino, Rocío Alaíz-Rodríguez, Víctor González-Castro, Eduardo Fidalgo, Enrique Alegre |
Knowl. Based Syst. | 5 |
| 2024 | PushPull-Net: Inhibition-Driven ResNet Robust to Image Corruptions
Guru Swaroop Bennabhaktula, Enrique Alegre, Nicola Strisciuglio, George Azzopardi |
ICPR (8) | 2 |
| 2024 | DeepHSAR: Semi-supervised fine-grained learning for multi-label human sexual activity recognition
Abhishek Gangwar, Víctor González-Castro, Enrique Alegre, Eduardo Fidalgo, Alicia Martínez-Mendoza |
Inf. Process. Manag. | 3 |
| 2023 | Supervised ranking approach to identify infLuential websites in the darknetabstractAbstract The anonymity and high security of the Tor network allow it to host a significant amount of criminal activities. Some Tor domains attract more traffic than others, as they offer better products or services to their customers. Detecting the most influential domains in Tor can help detect serious criminal activities. Therefore, in this paper, we present a novel supervised ranking framework for detecting the most influential domains. Our approach represents each domain with 40 features extracted from five sources: text, named entities, HTML markup, network topology, and visual content to train the learning-to-rank (LtR) scheme to sort the domains based on user-defined criteria. We experimented on a subset of 290 manually ranked drug-related websites from Tor and obtained the following results. First, among the explored LtR schemes, the listwise approach outperforms the benchmarked methods with an NDCG of 0.93 for the top-10 ranked domains. Second, we quantitatively proved that our framework surpasses the link-based ranking techniques. Third, we observed that using the user-visible text feature can obtain comparable performance to all the features with a decrease of 0.02 at NDCG@5. The proposed framework might support law enforcement agencies in detecting the most influential domains related to possible suspicious activities. Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Deisy Chaves |
Appl. Intell. | 3 |
| 2023 | DeepSumm: Exploiting topic models and sequence to sequence networks for extractive text summarization
Akanksha Joshi, Eduardo Fidalgo, Enrique Alegre, Laura Fernández-Robles |
Expert Syst. Appl. | 3 |
| 2023 | Triple-BigGAN: Semi-supervised generative adversarial networks for image synthesis and classification on sexual facial expression recognition
Abhishek Gangwar, Víctor González-Castro, Enrique Alegre, Eduardo Fidalgo |
Neurocomputing | 3 |
| 2023 | Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation MappingabstractAutomatic accent classification is an active research field concerning speech processing. It can be useful to identify a speaker's region of origin, which can be applied in police investigations carried out by Law Enforcement Agencies, as well as for the improvement of current speech recognition systems. This paper presents a novel descriptor called Grad-Transfer, extracted using the Gradient-weighted Class Activation Mapping (Grad-CAM) method based on convolutional neural network (CNN) interpretability. Additionally, we propose a methodology for accent classification that implements Grad-Transfer, which is based on transferring the knowledge acquired by a CNN to a classical machine learning algorithm. The paper works on two hypotheses: the coarse localization maps produced by Grad-CAM on spectrograms are able to highlight the regions of the spectrograms that are important for predicting accents, and Grad-Transfer descriptors computed from audios represent distinctive descriptions of the target accents. These hypotheses were demonstrated experimentally, clustering the generated Grad-Transfer descriptors according to the original accent of the audios using Birch and$k$-means algorithms. We carried out experiments on the Voice Cloning Toolkit dataset, seeing an increase of macro average accuracy, and unweighted average recall in the results obtained by a Gaussian Naive Bayes classifier up to$23.00\%$, and$23.58\%$, respectively, compared to a model trained with spectrograms. This demonstrates that Grad-Transfer is able to improve the performance of accent classification models and opens the door to new implementations in similar tasks. Roberto Andrés Vasco Carofilis, Enrique Alegre, Eduardo Fidalgo, Laura Fernández-Robles |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Detecting malware using text documents extracted from spam email through machine learningabstractSpam has become an effective way for cybercriminals to spread malware. Although cybersecurity agencies and companies develop products and organise courses for people to detect malicious spam email patterns, spam attacks are not totally avoided yet. In this work, we present and make publicly available "Spam Email Malware Detection - 600" (SEMD-600), a new dataset, based on Bruce Guenter's, for malware detection in spam using only the text of the email. We also introduce a pipeline for malware detection based on traditional Natural Language Processing (NLP) techniques. Using SEMD-600, we compare the text representation techniques Bag of Words and Term Frequency-Inverse Document Frequency (TF-IDF), in combination with three different supervised classifiers: Support Vector Machine, Naive Bayes and Logistic Regression, to detect malware in plain text documents. We found that combining TF-IDF with Logistic Regression achieved the best performance, with a macro F1 score of 0.763. Luis Ángel Redondo-Gutierrez, Francisco Jáñez-Martino, Eduardo Fidalgo, Enrique Alegre, Víctor González-Castro, Rocío Alaíz-Rodríguez |
DocEng | 4 |
| 2022 | Camera model identification based on forensic traces extracted from homogeneous patchesabstractA crucial challenge in digital image forensics is to identify the source camera model used to generate given images. This is of prime importance, especially for Law Enforcement Agencies in their investigations of Child Sexual Abuse Material found in darknets or seized storage devices. In this work, we address this challenge by proposing a solution that is characterized by two main contributions. It relies on the extraction of rather small homogeneous regions that we extract very efficiently from the integral image, and on a hierarchical classification approach with convolutional neural networks as the underlying models. We rely on homogeneous regions as they contain camera traces that are less distorted than regions with high-level scene content. The hierarchical approach that we propose is important for scaling up and making minimal modifications when new cameras are added. Furthermore, this scheme performs better than the traditional single classifier approach. By means of thorough experimentation on the publicly available Dresden data set, we achieve an accuracy of 99.01% with 5-fold cross-validation on the ‘natural’ subset of this data set. To the best of our knowledge, this is the best result ever reported for Dresden data set. Guru Swaroop Bennabhaktula, Enrique Alegre, Dimka Karastoyanova, George Azzopardi |
Expert Syst. Appl. | 2 |
| 2022 | RankSum - An unsupervised extractive text summarization based on rank fusion
Akanksha Joshi, Eduardo Fidalgo, Enrique Alegre, Rocío Alaíz-Rodríguez |
Expert Syst. Appl. | 3 |
| 2022 | Phishing websites detection using a novel multipurpose dataset and web technologies featuresabstractPhishing attacks are one of the most challenging social engineering cyberattacks due to the large amount of entities involved in online transactions and services. In these attacks, criminals deceive users to hijack their credentials or sensitive data through a login form which replicates the original website and submits the data to a malicious server. Many anti-phishing techniques have been developed in recent years, using different resource such as the URL and HTML code from legitimate index websites and phishing ones. These techniques have some limitations when predicting legitimate login websites, since, usually, no login forms are present in the legitimate class used for training the proposed model. Hence, in this work we present a methodology for phishing website detection in real scenarios, which uses URL, HTML, and web technology features. Since there is not any updated and multipurpose dataset for this task, we crafted the Phishing Index Login Websites Dataset (PILWD), an offline phishing dataset composed of 134,000 verified samples, that offers to researchers a wide variety of data to test and compare their approaches. Since approximately three-quarters of collected phishing samples request the introduction of credentials, we decided to crawl legitimate login websites to match the phishing standpoint. The developed approach is independent of third party services and the method relies on a new set of features used for the very first time in this problem, some of them extracted from the web technologies used by the on each specific website. Experimental results show that phishing websites can be detected with 97.95% accuracy using a LightGBM classifier and the complete set of the 54 features selected, when it was evaluated on PILWD dataset. Manuel Sánchez-Paniagua, Eduardo Fidalgo, Enrique Alegre, Rocío Alaíz-Rodríguez |
Expert Syst. Appl. | 3 |
| 2022 | A survey on methods, datasets and implementations for scene text spottingabstractAbstract Text Spotting is the union of the tasks of detection and transcription of the text that is present in images. Due to the various problems often found when retrieving text, such as orientation, aspect ratio, vertical text or multiple languages in the same image, this can be a challenging task. In this paper, the most recent methods and publications in this field are analysed and compared. Apart from presenting features already seen in other surveys, such as their architectures and performance on different datasets, novel perspectives for comparison are also included, such as the hardware, software, backbone architectures, main problems to solve, or programming languages of the algorithms. The review highlights information often omitted in other studies, providing a better understanding of the current state of research in Text Spotting, from 2016 to 2022, current problems and future trends, as well as establishing a baseline for future methods development, comparison of results and serving as guideline for choosing the most appropriate method to solve a particular problem. Pablo Blanco-Medina, Eduardo Fidalgo, Enrique Alegre, Víctor González-Castro |
IET Image Process. | 3 |
| 2021 | Video Camera Identification from Sensor Pattern Noise with a Constrained ConvNetabstractThe identification of source cameras from videos, though it is a highly relevant forensic analysis topic, has been studied much less than its counterpart that uses images. In this work we propose a method to identify the source camera of a video based on camera specific noise patterns that we extract from video frames. For the extraction of noise pattern features, we propose an extended version of a constrained convolutional layer capable of processing color inputs. Our system is designed to classify individual video frames which are in turn combined by a majority vote to identify the source camera. We evaluated this approach on the benchmark VISION data set consisting of 1539 videos from 28 different cameras. To the best of our knowledge, this is the first work that addresses the challenge of video camera identification on a device level. The experiments show that our approach is very promising, achieving up to 93.1% accuracy while being robust to the WhatsApp and YouTube compression techniques. This work is part of the EU-funded project 4NSEEK focused on forensics against child sexual abuse. Derrick Timmerman, Guru Swaroop Bennabhaktula, Enrique Alegre, George Azzopardi |
ICPRAM | 3 |
| 2021 | AttM-CNN: Attention and metric learning based CNN for pornography, age and Child Sexual Abuse (CSA) Detection in images
Abhishek Gangwar, Víctor González-Castro, Enrique Alegre, Eduardo Fidalgo |
Neurocomputing | 3 |
| 2021 | A new perceptual hashing method for verification and identity classification of occluded faces
Rubel Biswas, Víctor González-Castro, Eduardo Fidalgo, Enrique Alegre |
Image Vis. Comput. | 4 |
| 2021 | Image retrieval based on texture using latent space representation of discrete Fourier transformed maps
Surajit Saikia, Laura Fernández-Robles, Enrique Alegre, Eduardo Fidalgo |
Neural Comput. Appl. | 3 |
| 2020 | Device-based Image Matching with Similarity Learning by Convolutional Neural Networks that Exploit the Underlying Camera Sensor Pattern NoiseabstractOne of the challenging problems in digital image forensics is the capability to identify images that are captured by the same camera device. This knowledge can help forensic experts in gathering intelligence about suspects by analyzing digital images. In this paper, we propose a two-part network to quantify the likelihood that a given pair of images have the same source camera, and we evaluated it on the benchmark Dresden data set containing 1851 images from 31 different cameras. To the best of our knowledge, we are the first ones addressing the challenge of device-based image matching. Though the proposed approach is not yet forensics ready, our experiments show that this direction is worth pursuing, achieving at this moment 85 percent accuracy. This ongoing work is part of the EU-funded project 4NSEEK concerned with forensics against child sexual abuse. Guru Swaroop Bennabhaktula, Enrique Alegre, Dimka Karastoyanova, George Azzopardi |
ICPRAM | 2 |
| 2020 | File Name Classification Approach to Identify Child Sexual Abuse
Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Rocío Alaíz-Rodríguez |
ICPRAM | 3 |
| 2020 | Perceptual image hashing based on frequency dominant neighborhood structure applied to Tor domains recognition
Rubel Biswas, Víctor González-Castro, Eduardo Fidalgo, Enrique Alegre |
Neurocomputing | 4 |
| 2020 | Improving named entity recognition in noisy user-generated text with local distance neighbor feature
Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Laura Fernández-Robles |
Neurocomputing | 3 |
| 2019 | SummCoder: An unsupervised framework for extractive text summarization based on deep auto-encoders
Akanksha Joshi, Eduardo Fidalgo, Enrique Alegre, Laura Fernández-Robles |
Expert Syst. Appl. | 3 |
| 2019 | ToRank: Identifying the most influential suspicious domains in the Tor network
Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Laura Fernández-Robles |
Expert Syst. Appl. | 3 |
| 2018 | Boosting image classification through semantic attention filtering strategies
Eduardo Fidalgo, Enrique Alegre, Víctor González-Castro, Laura Fernández-Robles |
Pattern Recognit. Lett. | 2 |
| 2017 | Classifying Illegal Activities on Tor Network Based on Web Textual ContentsabstractMhd Wesam Al Nabki, Eduardo Fidalgo, Enrique Alegre, Ivan de Paz. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Ivan de Paz |
EACL (1) | 3 |
| 2016 | Increased generalization capability of trainable COSFIRE filters with application to machine visionabstractThe recently proposed trainable COSFIRE filters are highly effective in a wide range of computer vision applications, including object recognition, image classification, contour detection and retinal vessel segmentation. A COSFIRE filter is selective for a collection of contour parts in a certain spatial arrangement. These contour parts and their spatial arrangement are determined in an automatic configuration procedure from a single user-specified pattern of interest. The traditional configuration, however, does not guarantee the selection of the most distinctive contour parts. We propose a genetic algorithm-based optimization step in the configuration of COSFIRE filters that determines the minimum subset of contour parts that best characterize the pattern of interest. We use a public dataset of images of an edge milling head machine equipped with multiple cutting tools to demonstrate the effectiveness of the proposed optimization step for the detection and localization of such tools. The optimization process that we propose yields COSFIRE filters with substantially higher generalization capability. With an average of only six COSFIRE filters we achieve high precision P and recall R rates (P = 91.99%; R = 96.22%). This outperforms the original COSFIRE filter approach (without optimization) mostly in terms of recall. The proposed optimization procedure increases the efficiency of COSFIRE filters with little effect on the selectivity. George Azzopardi, Laura Fernández-Robles, Enrique Alegre, Nicolai Petkov |
ICPR | 3 |
| 2016 | Compass radius estimation for improved image classification using Edge-SIFT
Eduardo Fidalgo, Enrique Alegre, Víctor González-Castro, Laura Fernández-Robles |
Neurocomputing | 2 |
| 2015 | Cutting Edge Localisation in an Edge Profile Milling Head
Laura Fernández-Robles, George Azzopardi, Enrique Alegre, Nicolai Petkov |
CAIP (2) | 3 |
| 2014 | Local Oriented Statistics Information Booster (LOSIB) for Texture ClassificationabstractLocal oriented statistical information booster (LOSIB) is a descriptor enhancer based on the extraction of the gray level differences along several orientations. Specifically, the mean of the differences along particular orientations is considered. In this paper we have carried out some experiments using several classical texture descriptors to show that classification results are better when they are combined with LOSIB, than without it. Both parametric and non-parametric classifiers, Support Vector Machine and k-Nearest Neighbourhoods respectively, were applied to assess this new method. Furthermore, two different texture dataset were evaluated: KTH-Tips-2a and Brodatz32 to prove the robustness of LOSIB. Global descriptors such as WCF4 (Wavelet Co-occurrence Features), that extracts Haralick features from the Wavelet Transform, have been combined with LOSIB obtaining an improvement of 16.94% on KTH and 7.55% on Brodatz when classifying with SVM. Moreover, LOSIB was used together with state-of-the-art local descriptors such as LBP (Local Binary Pattern) and several of its recent variants. Combined with CLBP (Complete LBP), the LOSIB booster results were improved in 5.80% on KTH-Tips 2a and 7.09% on the Brodatz dataset. For all the tested descriptors, we have observed that a higher performance has been achieved, with the two classifiers on both datasets, when using some LOSIB settings. Oscar García-Olalla, Enrique Alegre, Laura Fernández-Robles, Víctor González-Castro |
ICPR | 2 |
| 2014 | aZIBO: A New Descriptor Based in Shape Moments and Rotational Invariant FeaturesabstractIn this work, a descriptor called a ZIBO (absolute Zernike moments with Invariant Boundary Orientation) that describes the shape of objects using the module of Zernike moments and the edge features obtained from an almost rotational invariant version of the Edge Gradient Co-occurrence Matrix (EGCM) is proposed. The two descriptors obtained, the Zernike module as global descriptor and the new version of EGCM as local one, are used to characterize images from three different datasets, Kimia99, MPEG2 and MPEG7. Later on, the concatenation of both local and global descriptors was evaluated using kNN with City block and Chi-square distance metrics. Also, the descriptors are assessed separately with a weight-based method, being the results obtained compared with the ones reached by the baseline method, ZMEG (Zernike Moment Edge Gradient). Using MPEG7, which is the most challenging dataset, and the weight-based classifier, this proposal obtained a success rate of 78.29%, outperforming the 75.86% achieved by ZMEG method. With the MPEG2 dataset, results were even better with an 81.00% of success rate against 77.25% of ZMEG. María Teresa García-Ordás, Enrique Alegre, Víctor González-Castro, Diego García-Ordás |
ICPR | 2 |
| 2013 | Evaluation of LBP Variants Using Several Metrics and kNN Classifiers
Oscar García-Olalla, Enrique Alegre, María Teresa García-Ordás, Laura Fernández-Robles |
SISAP | 2 |
| 2013 | Evaluation of Different Metrics for Shape Based Image Retrieval Using a New Contour Points Descriptor
María Teresa García-Ordás, Enrique Alegre, Oscar García-Olalla, Diego García-Ordás |
SISAP | 2 |
| 2013 | Class distribution estimation based on the Hellinger distance
Víctor González-Castro, Rocío Alaíz-Rodríguez, Enrique Alegre |
Inf. Sci. | 3 |
| 2011 | Boar Spermatozoa Classification Using Longitudinal and Transversal Profiles (LTP) Descriptor in Digital Images
Enrique Alegre, Oscar García-Olalla, Víctor González-Castro, Swapna Joshi |
IWCIA | 1 |
| 2010 | Estimating Class Proportions in Boar Semen Analysis Using the Hellinger Distance
Víctor González-Castro, Rocío Alaíz-Rodríguez, Laura Fernández-Robles, Roberto Guzmán-Martínez, Enrique Alegre |
IEA/AIE (1) | 5 |
| 2005 | Statistical Approach to Boar Semen Head Classification Based on Intracellular Intensity Distribution
Lidia Sánchez-González, Nicolai Petkov, Enrique Alegre |
CAIP | 3 |
| 2005 | Tool Insert Wear Classification Using Statistical Descriptors and Neuronal Networks
Enrique Alegre, Rocío Alaíz-Rodríguez, Joaquín Barreiro, M. Viñuela |
CIARP | 1 |
| 2005 | Classification of Boar Spermatozoid Head Images Using a Model Intracellular Density Distribution
Lidia Sánchez-González, Nicolai Petkov, Enrique Alegre |
CIARP | 3 |