VLDB 2026 Research / reviewers in the wild / expert
Francisco Jáñez-Martino
dblp:247/7965
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0001-7665-6418ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On persuasion in spam email: A multi-granularity text analysisabstractData will be made available on request Francisco Jáñez-Martino, Alberto Barrón-Cedeño, Rocío Alaíz-Rodríguez, Víctor González-Castro, Arianna Muti |
Expert Syst. Appl. | 1 |
| 2025 | Classifying the content of online notepad services using active learningabstractAbstract Pastebin is an online notepad service to share text anonymously. However, it could be misused to propagate suspicious or even illegal activities, like leaking sensitive information or sharing hyperlinks to child sexual abuse material. Due to the high rate of daily upload pastes, manual inspection of this material is not feasible. Conversely, an automatic classifier could identify such activities with little or no human intervention. However, a supervised model may require a significant number of training samples and have to handle distinct text typologies presented in Pastebin. This paper presents a classification approach composed of three cascading supervised classifiers that use Active Learning to select and label the most informative samples from Pastebin. The modularity of the proposed design allows each classifier to adapt to a specific text typology. The first classifier determines whether the text is a code snippet, and the second is to identify whether it is readable. The third classification level is twofold: (i) a binary classifier to say whether the text is suspicious and (ii) a multiclass classifier with seven predefined categories of possibly illegal activities. The average class recall of the binary and multiclass classifiers is $$95.24\%$$ 95.24 % and $$80.33\%$$ 80.33 % , respectively. Additionally, this paper presents a dataset of 3.8 million Pastebin samples, called onlIne Notepad Services PastEbin aCtiviTies (INSPECT-3.8M), along with their labels using our classification framework. Our classifier recognised that $$7.54\%$$ 7.54 % of the collected samples are correlated with presumably criminal activities. Law enforcement agencies may benefit from the insights shared in our research when aiming to investigate or automate the monitoring of Pastebin or other Online Notepad Services. This would allow responsible authorities to block illegal content before it spreads to the public. Mhd Wesam Al-Nabki, Eduardo Fidalgo, Enrique Alegre, Sarah Jane Delany, Francisco Jáñez-Martino |
J. Intell. Inf. Syst. | 5 |
| 2025 | Spam email classification based on cybersecurity potential risk using natural language processing
Francisco Jáñez-Martino, Rocío Alaíz-Rodríguez, Víctor González-Castro, Eduardo Fidalgo, Enrique Alegre |
Knowl. Based Syst. | 1 |
| 2022 | Detecting malware using text documents extracted from spam email through machine learningabstractSpam has become an effective way for cybercriminals to spread malware. Although cybersecurity agencies and companies develop products and organise courses for people to detect malicious spam email patterns, spam attacks are not totally avoided yet. In this work, we present and make publicly available "Spam Email Malware Detection - 600" (SEMD-600), a new dataset, based on Bruce Guenter's, for malware detection in spam using only the text of the email. We also introduce a pipeline for malware detection based on traditional Natural Language Processing (NLP) techniques. Using SEMD-600, we compare the text representation techniques Bag of Words and Term Frequency-Inverse Document Frequency (TF-IDF), in combination with three different supervised classifiers: Support Vector Machine, Naive Bayes and Logistic Regression, to detect malware in plain text documents. We found that combining TF-IDF with Logistic Regression achieved the best performance, with a macro F1 score of 0.763. Luis Ángel Redondo-Gutierrez, Francisco Jáñez-Martino, Eduardo Fidalgo, Enrique Alegre, Víctor González-Castro, Rocío Alaíz-Rodríguez |
DocEng | 2 |
| 2021 | Trustworthiness of spam email addresses using machine learningabstractCybercriminals have increasingly used spam email to send scams, phishing, malware and other frauds to organisations and people. They design sophisticated and contextualised emails to make them look trustworthy for users, being the sender addresses an essential part. Although cybersecurity agencies and companies develop products and organise courses for people to detect emails patterns, spam attacks are not totally avoided yet. Francisco Jáñez-Martino, Rocío Alaíz-Rodríguez, Víctor González-Castro, Eduardo Fidalgo |
DocEng | 1 |