VLDB 2026 Research / reviewers in the wild / expert
Alejandro Martín
dblp:70/2062
· DBLP profile ↗
24ranked-venue papers
10as first author
14since 2021 · last 2027
0000-0002-0800-7632ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 9 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Beyond semantics: Content leakage mitigation using synthetic hard negatives for style embeddingsabstractPurpose: Authorship can be defined as a combination of content and style . Modern open-source transformer foundational authorship models apply contrastive learning techniques. When naively contrasting texts to an authorship task some amount of semantic leakage is present, as authors frequently repeat their topic preferences. Our aim is to reduce spurious correlations due to topic leakage born from contrastive objectives. Methodology: We present a technique to modify a well-established contrastive learning objective (InfoNCE) using synthetic hard negative examples for in-domain topic leakage improvements, while preserving competitive out of domain performance. This topic leakage mitigation technique aims to distance the content embedding space from the style embedding space. Our experiments aim to demonstrate this technique in detail, using exclusively affordable encoder-only models instead of costly hard negative mining. Results: We showcase the performance with ablations on two different datasets and compare them on out-of-domain challenges. We improve on challenging evaluations with prolific authors, with up to 10% increase in accuracy on highly diverse bodies of work . Trials with standard challenges also demonstrate the preservation of zero-shot capabilities of this method as fine-tuning. Javier Huertas-Tato, Adrian Giron-Jimenez, Alejandro Martín, David Camacho |
Expert Syst. Appl. | 3 |
| 2025 | Toxic Discourse in the Digital Battlefield: Analysing Telegram Channels During the Russia-Ukraine 'Conflict'abstractABSTRACT Instant messenger Telegram has emerged as a favoured platform for far‐right activism, conspiracy theories, political propaganda, and misinformation, which has its own target audience. This study explores the application of multilingual pre‐trained language models to detect and measure toxicity in political content on Telegram channels. The proposed techniques have shown notable advancements in identifying toxic information using a fine‐tuned RoBERTa model. Through the combination of data analysis, time‐series analysis, and BERTopic modelling, the research demonstrates how toxicity varies by topic, country, and time period, using metadata. The study identified key topics in the dataset, which includes 23.6 million messages from 1491 Telegram channels, including the Russian–Ukrainian conflict and political tensions in Europe and the United States from 2016 to 1 July 2023. Despite these achievements, challenges such as the dominance of Russian language content and a focus on specific topics were highlighted. This research advances the understanding of how toxic language and propaganda are disseminated across different languages and political narratives, contributing to the study of digital communication and information warfare. Arsenii Tretiakov, Sergio D'Antonio-Maceiras, Áurea Anguera de Sojo-Hernández, Alejandro Martín |
Expert Syst. J. Knowl. Eng. | 4 |
| 2025 | Deep-Sync: A novel deep learning-based tool for semantic-aware subtitling synchronisation
Alejandro Martín, Israel González-Carrasco, Víctor Rodríguez-Fernández, Monica Souto-Rico, David Camacho, Belén Ruíz-Mezcua |
Neural Comput. Appl. | 1 |
| 2024 | Topic Modeling in Telegram Channels During the Russia-Ukraine Conflict
Arsenii Tretiakov, Sergio D'Antonio-Maceiras, Alejandro Martín |
IDEAL (1) | 3 |
| 2024 | Using Contrastive Learning to Map Stylistic Similarities in Narrative Writers
María Valero-Redondo, Javier Huertas-Tato, Sergio D'Antonio-Maceiras, Alejandro Martín, David Camacho |
IDEAL (1) | 4 |
| 2024 | Understanding writing style in social media with a supervised contrastively pre-trained transformerabstractWe introduce the Style Transformer for Authorship Representations (STAR) to detect and characterize writing style in social media. The model is trained on a heterogeneous large corpus derived from public sources with 4.5⋅106 authored texts from 70k authors leveraging Supervised Contrastive Loss to minimize the distance between texts authored by the same individual. This pretext pre-training task yields competitive performance at zero-shot with PAN challenges on attribution and clustering. We attain promising results on PAN verification challenges using STAR as a feature extractor. Finally, we present results from our test partition on Reddit, where using a support base of 8 documents of 512 tokens, we can discern authors from sets of up to 1616 authors with at least 80% accuracy. We share our pre-trained model at huggingface AIDA-UPM/star and our code is available at jahuerta92/star. Javier Huertas-Tato, Alejandro Martín, David Camacho |
Knowl. Based Syst. | 2 |
| 2023 | BERTuit: Understanding Spanish language in Twitter with transformersabstractAbstract The appearance of complex attention‐based language models such as BERT, RoBERTa or GPT‐3 has allowed to address highly complex tasks in a plethora of scenarios. However, when applied to specific domains, these models encounter considerable difficulties. This is the case of Social Networks such as Twitter, an ever‐changing stream of information written with informal and complex language, where each message requires careful evaluation to be understood even by humans given the important role that context plays. Addressing tasks in this domain through Natural Language Processing involves severe challenges. When powerful state‐of‐the‐art multilingual language models are applied to this scenario, language specific nuances get lost in translation. To face these challenges we present BERTuit, the largest transformer proposed so far for Spanish language, pre‐trained on a massive dataset of 230 M Spanish tweets using RoBERTa optimization. Our motivation is to provide a powerful resource to better understand Spanish Twitter and to be used on applications focused on this social network, with special emphasis on solutions devoted to tackle the spreading of misinformation in this platform. BERTuit is evaluated on several tasks and compared against M‐BERT, XLM‐RoBERTa and XLM‐T, very competitive multilingual transformers. The utility of our approach is shown with applications, in this case: an unsupervised methodology to visualize groups of hoaxes; and supervised profiling of authors spreading disinformation. Javier Huertas-Tato, Alejandro Martín, David Camacho |
Expert Syst. J. Knowl. Eng. | 2 |
| 2023 | Evolving Generative Adversarial Networks to improve image steganographyabstractImages have been repeatedly used as the perfect environment to hide information through the use of steganography techniques. Whether messages, documents or even other images, the bitmap of an digital picture provides a place where hidden data can be embedded without human notice. So far, a plethora of steganography methods can be found in the state-of-the-art literature, together with steganalysis techniques, devoted to detect the presence of hidden information in files. Recent steganography techniques rely on Convolutional Neural Networks, trying to embed as information as possible while minimising visual changes in the image. Following this trend, this article tries to demonstrate that a Generative Adversarial Network (GAN) can be used to improve the ability of a spatial domain steganalysis method and to insert secret information with minimal image alteration. Through a training process, the GAN learns how to adapt an image to later introduce a message using the Least Significant Bit steganography algorithm. The results evidence that the approach is successful at avoiding detection by a state-of-the-art Deep Learning steganalysis architecture. Alejandro Martín, Alfonso Hernández, Moutaz Alazab, Jason J. Jung, David Camacho |
Expert Syst. Appl. | 1 |
| 2022 | Detection of False Information in Spanish Using Machine Learning Techniques
Arsenii Tretiakov, Alejandro Martín, David Camacho |
IDEAL | 2 |
| 2022 | Generating Authorship Embeddings with TransformersabstractAuthorship attribution and profiling tools provide useful instruments with wide areas of application, such as disinformation spreaders detection. Models developed until now to fulfil these tasks usually rely on manually crafted features or on a training process restricted and limited by the number of authors involved. Besides, current methods have limited capacity to generate a broad representation of the author, without considering a great variety of features that can be extracted from their texts. In this paper, we propose a contrastive training method to generate representative embeddings of the authorship of a text and able to generalize to unseen authors. The core of this method is the Transformer architecture, which is known to generate very powerful semantically-aware text representations. Using a pretrained RoBERTa-large model, we evaluate our method on the environment of the standardized Gutenberg corpus, detecting the authorship of literary works. Representations generated by our proposal can be used to extract representative authorship embeddings and to visualize meaningful relationships between authors, genres and books. Furthermore, the embedding method achieves zero-shot 79% accuracy and 94% top-5 accuracy when tasked to distinguish a text piece from a set of 100 authored texts. Our code is readily available on GitHub11https://github.com/jahuerta92/authorship-embedding Javier Huertas-Tato, Alejandro Martín, Álvaro Huertas-García, David Camacho |
IJCNN | 2 |
| 2022 | SILT: Efficient transformer training for inter-lingual inference
Javier Huertas-Tato, Alejandro Martín, David Camacho |
Expert Syst. Appl. | 2 |
| 2022 | FacTeR-Check: Semi-automated fact-checking through semantic similarity and natural language inferenceabstractOur society produces and shares overwhelming amounts of information through Online Social Networks (OSNs). Within this environment, misinformation and disinformation have proliferated, becoming a public safety concern in most countries. Allowing the public and professionals to efficiently find reliable evidence about the factual veracity of a claim is a crucial step to mitigate this harmful spread. To this end, we propose FacTeR-Check, a multilingual architecture for semi-automated fact-checking and hoaxes propagation analysis that can be used to implement applications designed for both the general public and for fact-checking organisations. FacTeR-Check implements three different modules relying on the XLM-RoBERTa Transformer architecture to evaluate semantic similarity, to calculate natural language inference and to build search queries through automatic keywords extraction and Named-Entity Recognition. The three modules have been validated using state-of-the-art benchmark datasets, exhibiting good performance in all of them. Besides, FacTeR-Check is employed to collect and label a dataset, called NLI19-SP, composed of more than 40,000 tweets supporting or denying 60 hoaxes related to COVID-19, released publicly. Finally, an analysis of the data collected in this dataset is provided, which allows to obtain a deep insight of how disinformation operated during the COVID-19 pandemic in Spanish-speaking countries. Alejandro Martín, Javier Huertas-Tato, Álvaro Huertas-García, Guillermo Villar-Rodríguez, David Camacho |
Knowl. Based Syst. | 1 |
| 2022 | Recent advances on effective and efficient deep learning-based solutions
Alejandro Martín, David Camacho |
Neural Comput. Appl. | 1 |
| 2021 | Countering Misinformation Through Semantic-Aware Multilingual Models
Álvaro Huertas-García, Javier Huertas-Tato, Alejandro Martín, David Camacho |
IDEAL | 3 |
| 2020 | Statistically-driven Coral Reef metaheuristic for automatic hyperparameter setting and architecture design of Convolutional Neural NetworksabstractThe adjustment of the hyperparameters and network structure of Convolutional Neural Networks (CNNs) composes an important step towards building effective, but still efficient learning models. The selection of the best configuration is a problem-dependent task that involves to explore an enormous and complex search space. Due to this reason, the use of heuristic-based search fits perfectly within this task, seeking to obtain a near to optimal solution in a complex and large exploratory space. This paper presents SCRODeep, a self-adapting algorithm based on a statistically-driven Coral Reef Optimisation algorithm (SCRO), for the selection of the most adequate CNNs architecture in a particular domain. This metaheuristic has been designed to navigate through a search space where the architecture (defining the particular set of layers, including convolutional or pooling layers), and the hyperparameters of the network (i.e. activation functions, number of units or the kernel initializer, among others) are represented, but where the connections weights and bias are inferred using typical CNNs optimisation algorithms. In contrast to other approaches, where the use of a metaheuristic implies in turn to fix a series of hyperparameters (i.e. the mutation probability in a genetic algorithm), our approach follows a self-parametrisation perspective, thus removing the necessity of fixing these values. The method has been tested in the design of CNNs for image classification, showing that SCRODeep is able to find competitive solutions, while the complexity of the architectures found is constrained. Alejandro Martín, Raúl Lara-Cabrera, Víctor Manuel Vargas Yun, Pedro Antonio Gutiérrez, César Hervás-Martínez, David Camacho |
CEC | 1 |
| 2020 | Cloud Type Identification Using Data Fusion and Ensemble Learning
Javier Huertas-Tato, Alejandro Martín, David Camacho |
IDEAL (2) | 2 |
| 2019 | Dynamic emphatical narration for reduced authorial burden and increased user freedom in interactive storytellingabstractInteractive storytelling systems have become very popular as they engage users in the creation of narrative. A fundamental challenge for such systems is that the users feel unconstrained in their exploration of the environment and yet retain for the author some control of what the user does. Traditional solutions address this challenge by hardening the authoring task, since much effort has to be devoted to create the story and the necessary mechanisms to provide the user with enough freedom that, at the same time, keeps the story consistent. We propose a solution based on dynamic narration that combines a set of authored materials associated with particular points in the environment and the exploratory movements of the user, and includes statements of varying emphasis to guide the user towards the desired elements. This allows the system to generate a narration in real time while giving the user the possibility to decide what to do. The solution has been field tested and results are reported on three different aspects: the freedom of the user in terms of combinations in which the authored material is traversed; the divergence of user explorations from author's intentions; and user response to guidance statements of differing emphasis. Gonzalo Méndez 0001, Raquel Hervás, Pablo Gervás, Alejandro Martín, Frank Julca |
Connect. Sci. | 4 |
| 2019 | The impact of class imbalance in classification performance metrics based on the binary confusion matrixabstractA major issue in the classification of class imbalanced datasets involves the determination of the most suitable performance metrics to be used. In previous work using several examples, it has been shown that imbalance can exert a major impact on the value and meaning of accuracy and on certain other well-known performance metrics. In this paper, our approach goes beyond simply studying case studies and develops a systematic analysis of this impact by simulating the results obtained using binary classifiers. A set of functions and numerical indicators are attained which enables the comparison of the behaviour of several performance metrics based on the binary confusion matrix when they are faced with imbalanced datasets. Throughout the paper, a new way to measure the imbalance is defined which surpasses the Imbalance Ratio used in previous studies. From the simulation results, several clusters of performance metrics have been identified that involve the use of Geometric Mean or Bookmaker Informedness as the best null-biased metrics if their focus on classification successes (dismissing the errors) presents no limitation for the specific application where they are used. However, if classification errors must also be considered, then the Matthews Correlation Coefficient arises as the best choice. Finally, a set of null-biased multi-perspective Class Balance Metrics is proposed which extends the concept of Class Balance Accuracy to other performance metrics. Amalia Luque, Alejandro Carrasco, Alejandro Martín, Ana de las Heras |
Pattern Recognit. | 3 |
| 2018 | CANDYMAN: Classifying Android malware families by modelling dynamic traces with Markov chains
Alejandro Martín, Víctor Rodríguez-Fernández, David Camacho |
Eng. Appl. Artif. Intell. | 1 |
| 2018 | Picking on the family: Disrupting android malware triage by forcing misclassificationabstractMachine learning classification algorithms are widely applied to different malware analysis problems because of their proven abilities to learn from examples and perform relatively well with little human input. Use cases include the labelling of malicious samples according to families during triage of suspected malware. However, automated algorithms are vulnerable to attacks. An attacker could carefully manipulate the sample to force the algorithm to produce a particular output. In this paper we discuss one such attack on Android malware classifiers. We design and implement a prototype tool, called IagoDroid, that takes as input a malware sample and a target family, and modifies the sample to cause it to be classified as belonging to this family while preserving its original semantics. Our technique relies on a search process that generates variants of the original sample without modifying their semantics. We tested IagoDroid against RevealDroid, a recent, open source, Android malware classifier based on a variety of static features. IagoDroid successfully forces misclassification for 28 of the 29 representative malware families present in the DREBIN dataset. Remarkably, it does so by modifying just a single feature of the original malware. On average, it finds the first evasive sample in the first search iteration, and converges to a 100% evasive population within 4 iterations. Finally, we introduce RevealDroid*, a more robust classifier that implements several techniques proposed in other adversarial learning domains. Our experiments suggest that RevealDroid* can correctly detect up to 99% of the variants generated by IagoDroid. Alejandro Calleja, Alejandro Martín, Héctor D. Menéndez 0001, Juan Tapiador, David Clark 0001 |
Expert Syst. Appl. | 2 |
| 2018 | EvoDeep: A new evolutionary approach for automatic Deep Neural Networks parametrisation
Alejandro Martín, Raúl Lara-Cabrera, Félix Fuentes-Hurtado, Valery Naranjo, David Camacho |
J. Parallel Distributed Comput. | 1 |
| 2017 | Evolving Deep Neural Networks architectures for Android malware classificationabstractDeep Neural Networks (DNN) have become a powerful, widely used, and successful mechanism to solve problems of different nature and varied complexity. Their ability to build models adapted to complex non-linear problems, have made them a technique widely applied and studied. One of the fields where this technique is currently being applied is in the malware classification problem. The malware classification problem has an increasing complexity, due to the growing number of features needed to represent the behaviour of the application as exhaustively as possible. Although other classification methods, as those based on SVM, have been traditionally used, the DNN pose a promising tool in this field. However, the parameters and architecture setting of these DNNs present a serious restriction, due to the necessary time to find the most appropriate configuration. This paper proposes a new genetic algorithm designed to evolve the parameters, and the architecture, of a DNN with the goal of maximising the malware classification accuracy, and minimizing the complexity of the model. This model is tested against a dataset of malware samples, which are represented using a set of static features, so the DNN has been trained to perform a static malware classification task. The experiments carried out using this dataset show that the genetic algorithm is able to select the parameters and the DNN architecture settings, achieving a 91% accuracy. Alejandro Martín, Félix Fuentes-Hurtado, Valery Naranjo, David Camacho |
CEC | 1 |
| 2017 | MOCDroid: multi-objective evolutionary classifier for Android malware detection
Alejandro Martín, Héctor D. Menéndez 0001, David Camacho |
Soft Comput. | 1 |
| 2016 | Genetic boosting classification for malware detectionabstractIn the last few years virus writers have made use of new obfuscation techniques with the aim of hindering malware in order to difficult their detection by Anti-Virus engines. Strategies to reverse this trend involve executing potentially malicious programs and monitor the actions they perform in runtime, what is known as dynamic analysis. In this paper we present a method able to reach a high accuracy rate without using this kind of analysis. Instead we use a static analysis approach, which discards those samples that cannot be classified with enough certainty and need, certainly, a dynamic analysis. The K-means clustering algorithm has been used to group samples into regions according to their features. Then a boosting process, guided by a genetic algorithm, is executed in each region that are evaluated using a test dataset discarding those regions which do not reach a minimum accuracy threshold. Alejandro Martín, Héctor D. Menéndez 0001, David Camacho |
CEC | 1 |