EDBT 2026 Demo / reviewers in the wild / expert
Sergio Pastrana
dblp:41/8823 · also Sergio Pastrana Portillo
· DBLP profile ↗
26ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0003-1036-6359ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 13 · 2 first-author · 9 since 2021Computer networks · 8 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Snorkeling in Dark Waters: A Longitudinal Surface Exploration of Unique Tor Hidden ServicesabstractThe Onion Router (Tor) is a controversial network whose utility is constantly under scrutiny. On the one hand, it allows for anonymous interaction and cooperation of users seeking untraceable navigation on the Internet. This freedom also attracts criminals who aim to thwart law enforcement investigations, e.g., trading illegal products or services such as drugs or weapons. Tor allows delivering content without revealing the actual hosting address, by means of.onion (or hidden) services. Different from regular domains, these services can not be resolved by traditional name services, are not indexed by regular search engines, and they frequently change. This generates uncertainty about the extent and size of the Tor network and the type of content offered. In this work, we present a large-scale analysis of the Tor Network. We leverage our crawler, dubbed Mimir, which automatically collects and visits content linked within the pages to collect a dataset of pages from more than 25k sites. We analyze the topology of the Tor Network, including its depth and reachability from the surface web. We define a set of heuristics to detect the presence of replicated content (mirrors) and show that most of the analyzed content in the Dark Web (≈82%) is a replica of other content. Also, we train a custom Machine Learning classifier to understand the type of content the hidden services offer. Overall, our study provides new insights into the Tor network, highlighting the importance of initial seeding for focus on specific topics, and optimize the crawling process. We show that previous work on large-scale Tor measurements does not consider the presence of mirrors, which biases their understanding of the Dark Web topology and the distribution of content. Alfonso Rodriguez Barredo-Valenzuela, Sergio Pastrana, Guillermo Suarez-Tangil |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | IoC Stalker: Early detection of Indicators of CompromiseabstractOnline underground forums are used by cybercriminals to share information and knowledge related to malicious activities. Participants exchange "Indicators of Compromise" (IoCs) within the discussions. These may include Hashes, Domains, URLs, or IPs with potential malicious intent. While Open Source Intelligence (OSINT) eventually identifies these malicious IoCs, it may take an extensive amount of time, sometimes up to years, before they are identified as threats. However, the context in which these IoCs appear, and the information provided through the posts’ and authors’ context can already offer valuable insights about their malicious nature. Unfortunately, the large amount of unstructured noisy forum data presents a hurdle for automation. In this paper, we address the challenge of automatically distinguishing between posts containing IoCs posing a threat and those being harmless. We design a learning pipeline that does not use features derived from IoCs, enabling a timely identification of novel threats. We operate over a temporal representation of forum data and offer valuable insights into the optimal time window that tracks concept drift. We also study which types of IoCs are harder to predict (e.g., IPs) and how transfer learning from other types can help to improve their identification. We conduct our analysis on a prominent hacking forum, spanning over 18 years of data, and find that our model can detect IoCs ≈490 days before they appear in OSINT. Mariella Mischinger, Sergio Pastrana, Guillermo Suarez-Tangil |
ACSAC | 2 |
| 2024 | Reviewing War: Unconventional User Reviews as a Side Channel to Circumvent Information ControlsabstractDuring the first days of the 2022 Russian invasion of Ukraine, Russia's media regulator blocked access to many global social media platforms and news sites, including Twitter, Facebook, and the BBC. To bypass the information controls set by Russian authorities, pro-Ukrainian groups explored unconventional ways to reach out to the Russian population, such as posting war-related content in the user reviews of Russian businesses available on Google Maps or Tripadvisor. This paper provides a first analysis of this new phenomenon by analyzing the unconventional strategies used to avoid state censorship in the Russian Federation during the conflict. Specifically, we analyze reviews posted on these platforms from the beginning of the war to September 2022. We measure the channeling of war-related messages through user reviews on Tripadvisor and Google Maps. Our analysis of the content posted on these services reveals that users leveraged these platforms to seek and exchange humanitarian and travel advice, but also to disseminate disinformation and polarized messages. Finally, we analyze the response of platforms in terms of content moderation and their impact. José Miguel Moreno, Sergio Pastrana, Jens Helge Reelfs, Pelayo Vallina, Savvas Zannettou, Andriy Panchenko 0001, Georgios Smaragdakis, Oliver Hohlfeld, Narseo Vallina-Rodriguez, Juan Tapiador |
ICWSM | 2 |
| 2023 | An analysis of fake social media engagement servicesabstractFake engagement services allow users of online social media and other web platforms to illegitimately increase their online reach and boost their perceived popularity. Driven by socio-economic and even political motivations, the demand for fake engagement services has increased in the last years, which has incentivized the rise of a vast underground market and support infrastructure. Prior research in this area has been limited to the study of the infrastructure used to provide these services (e.g., botnets) and to the development of algorithms to detect and remove fake activity in online targeted platforms. Yet, the platforms in which these services are sold (known as panels) and the underground markets offering these services have not received much research attention. To fill this knowledge gap, this paper studies Social Media Management (SMM) panels, i.e., reselling platforms—often found in underground forums—in which a large variety of fake engagement services are offered. By daily crawling 86 representative SMM panels for 4 months, we harvest a dataset with 2.8 M forum entries grouped into 61k different services. This dataset allows us to build a detailed catalog of the services for sale, the platforms they target, and to derive new insights on fake social engagement services and its market. We then perform an economic analysis of fake engagement services and their trading activities by automatically analyzing 7k threads in underground forums. Our analysis reveals a broad range of offered services and levels of customization, where buyers can acquire fake engagement services by selecting features such as the quality of the service, the speed of delivery, the country of origin, and even personal attributes of the fake account (e.g., gender). The price analysis also yields interesting empirical results, showing significant disparities between prices of the same product across different markets. These observations suggest that the market is still undeveloped and sellers do not know the real market value of the services that they offer, leading them to underprice or overprice their services. David Nevado Catalán, Sergio Pastrana, Narseo Vallina-Rodriguez, Juan Tapiador |
Comput. Secur. | 2 |
| 2023 | Towards automated homomorphic encryption parameter selection with fuzzy logic and linear programmingabstractHomomorphic Encryption (HE) is a set of powerful properties of certain cryptosystems that allow privacy-preserving operation over the encrypted text. Still, HE is not widespread due to limitations in terms of efficiency and usability. Among the challenges of HE, scheme parametrization (i.e., the selection of appropriate parameters within the algorithms) is a relevant multi-faced problem. First, the parametrization needs to comply with a set of properties to guarantee the security of the underlying scheme. Second, parametrization requires a deep understanding of the low-level primitives since the parameters have a confronting impact on the scheme’s precision, performance, and security. Finally, the circuit to be executed influences, and it is influenced by, the parametrization. Thus, there is no general optimal selection of parameters, and this selection depends on the circuit and the scenario of the application. Currently, most existing HE frameworks require cryptographers to address these considerations manually. It requires a minimum of expertise acquired through a steep learning curve. In this paper, we propose a unified solution for the aforementioned challenges. Concretely, we present an expert system combining Fuzzy Logic and Linear Programming. The Fuzzy Logic Modules receive a user selection of high-level priorities for the security, efficiency, and performance of the cryptosystem. Based on these preferences, the expert system generates a Linear Programming Model that obtains optimal combinations of parameters by considering those priorities while preserving a minimum level of security for the cryptosystem. We conduct an extended evaluation showing that an expert system generates optimal parameter selections that maintain user preferences without undergoing the inherent complexity of analyzing the circuit. José Cabrero-Holgueras, Sergio Pastrana |
Expert Syst. Appl. | 2 |
| 2022 | Towards Improving Code Stylometry Analysis in Underground ForumsabstractAbstract Code Stylometry has emerged as a powerful mechanism to identify programmers. While there have been significant advances in the field, existing mechanisms underperform in challenging domains. One such domain is studying the provenance of code shared in underground forums, where code posts tend to have small or incomplete source code fragments. This paper proposes a method designed to deal with the idiosyncrasies of code snippets shared in these forums. Our system fuses a forum-specific learning pipeline with Conformal Prediction to generate predictions with precise confidence levels as a novelty. We see that identifying unreliable code snippets is paramount to generate high-accuracy predictions, and this is a task where traditional learning settings fail. Overall, our method performs as twice as well as the state-of-the-art in a constrained setting with a large number of authors (i.e., 100). When dealing with a smaller number of authors (i.e., 20), it performs at high accuracy (89%). We also evaluate our work on an open-world assumption and see that our method is more effective at retaining samples. Michal Tereszkowski-Kaminski, Sergio Pastrana, Jorge Blasco Alís, Guillermo Suarez-Tangil |
Proc. Priv. Enhancing Technol. | 2 |
| 2021 | Detecting Video-Game Injectors Exchanged in Game Cheating Communities
Panicos Karkallis, Jorge Blasco Alís, Guillermo Suarez-Tangil, Sergio Pastrana |
ESORICS (1) | 4 |
| 2021 | Trouble Over-The-Air: An Analysis of FOTA Apps in the Android EcosystemabstractAndroid firmware updates are typically managed by the so-called FOTA (Firmware Over-the-Air) apps. Such apps are highly privileged and play a critical role in maintaining devices secured and updated. The Android operating system offers standard mechanisms—available to Original Equipment Manufacturers (OEMs)—to implement their own FOTA apps but such vendor-specific implementations could be a source of security and privacy issues due to poor software engineering practices. This paper performs the first large-scale and systematic analysis of the FOTA ecosystem through a dataset of 2,013 FOTA apps detected with a tool designed for this purpose over 422,121 pre-installed apps. We classify the different stakeholders developing and deploying FOTA apps on the Android update ecosystem, showing that 43% of FOTA apps are developed by third parties. We report that some devices can have as many as 5 apps implementing FOTA capabilities. By means of static analysis of the code of FOTA apps, we show that some apps present behaviors that can be considered privacy intrusive, such as the collection of sensitive user data (e.g., geolocation linked to unique hardware identifiers), and a significant presence of third-party trackers. We also discover implementation issues leading to critical vulnerabilities, such as the use of public AOSP test keys both for signing FOTA apps and for update verification, thus allowing any update signed with the same key to be installed. Finally, we study telemetry data collected from real devices by a commercial security tool. We demonstrate that FOTA apps are responsible for the installation of non-system apps (e.g., entertainment apps and games), including malware and Potentially Unwanted Programs (PUP). Our findings suggest that FOTA development practices are misaligned with Google’s recommendations. Eduardo Blázquez, Sergio Pastrana, Álvaro Feal, Julien Gamba, Platon Kotzias, Narseo Vallina-Rodriguez, Juan Tapiador |
SP | 2 |
| 2021 | A Methodology For Large-Scale Identification of Related Accounts in Underground Forums
José Cabrero-Holgueras, Sergio Pastrana |
Comput. Secur. | 2 |
| 2021 | Avaddon ransomware: An in-depth analysis and decryption of infected systems
Javier Yuste, Sergio Pastrana |
Comput. Secur. | 2 |
| 2021 | SoK: Privacy-Preserving Computation Techniques for Deep LearningabstractAbstract Deep Learning (DL) is a powerful solution for complex problems in many disciplines such as finance, medical research, or social sciences. Due to the high computational cost of DL algorithms, data scientists often rely upon Machine Learning as a Service (MLaaS) to outsource the computation onto third-party servers. However, outsourcing the computation raises privacy concerns when dealing with sensitive information, e.g., health or financial records. Also, privacy regulations like the European GDPR limit the collection, distribution, and use of such sensitive data. Recent advances in privacy-preserving computation techniques (i.e., Homomorphic Encryption and Secure Multiparty Computation) have enabled DL training and inference over protected data. However, these techniques are still immature and difficult to deploy in practical scenarios. In this work, we review the evolution of the adaptation of privacy-preserving computation techniques onto DL, to understand the gap between research proposals and practical applications. We highlight the relative advantages and disadvantages, considering aspects such as efficiency shortcomings, reproducibility issues due to the lack of standard tools and programming interfaces, or lack of integration with DL frameworks commonly used by the data science community. José Cabrero-Holgueras, Sergio Pastrana |
Proc. Priv. Enhancing Technol. | 2 |
| 2021 | Blocklist Babel: On the Transparency and Dynamics of Open Source BlocklistingabstractBlocklists constitute a widely-used Internet security mechanism to filter undesired network traffic based on IP/domain reputation and behavior. Many blocklists are distributed in open source form by threat intelligence providers who aggregate and process input from their own sensors, but also from third-party feeds or providers. Despite their wide adoption, many open-source blocklist providers lack clear documentation about their structure, curation process, contents, dynamics, and inter-relationships with other providers. In this paper, we perform a transparency and content analysis of 2,093 free and open source blocklists with the aim of exploring those questions. To that end, we perform a longitudinal 6-month crawling campaign yielding more than 13.5M unique records. This allows us to shed light on their nature, dynamics, inter-provider relationships, and transparency. Specifically, we discuss how the lack of consensus on distribution formats, blocklist labeling taxonomy, content focus, and temporal dynamics creates a complex ecosystem that complicates their combined crawling, aggregation and use. We also provide observations regarding their generally low overlap as well as acute differences in terms of liveness (i.e., how frequently records get indexed and removed from the list) and the lack of documentation about their data collection processes, nature and intended purpose. We conclude the paper with recommendations in terms of transparency, accountability, and standardization. Álvaro Feal, Pelayo Vallina, Julien Gamba, Sergio Pastrana, Antonio Nappa, Oliver Hohlfeld, Narseo Vallina-Rodriguez, Juan Tapiador |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2019 | Understanding eWhoringabstractIn this paper, we describe a new type of online fraud, referred to as 'eWhoring' by offenders. This crime script analysis provides an overview of the 'eWhoring' business model, drawing on more than 6,500 posts crawled from an online underground forum. This is an unusual fraud type, in that offenders readily share information about how it is committed in a way that is almost prescriptive. There are economic factors at play here, as providing information about how to make money from 'eWhoring' can increase the demand for the types of images that enable it to happen. We find that sexualised images are typically stolen and shared online. While some images are shared for free, these can quickly become 'saturated', leading to the demand for (and trade in) more exclusive 'packs'. These images are then sold to unwitting customers who believe they have paid for a virtual sexual encounter. A variety of online services are used for carrying out this fraud type, including email, video, dating sites, social media, classified advertisements, and payment platforms. This analysis reveals potential interventions that could be applied to each stage of the crime commission process to prevent and disrupt this crime type. Alice Hutchings, Sergio Pastrana |
EuroS&P | 2 |
| 2019 | Measuring eWhoringabstracteWhoring is the term used by offenders to refer to a type of online fraud in which cybersexual encounters are simulated for financial gain. Perpetrators use social engineering techniques to impersonate young women in online communities, e.g., chat or social networking sites. They engage potential customers in conversation with the aim of selling misleading sexual material -- mostly photographs and interactive video shows -- illicitly compiled from third-party sites. eWhoring is a popular topic in underground communities, with forums acting as a gateway into offending. Users not only share knowledge and tutorials, but also trade in goods and services, such as packs of images and videos. In this paper, we present a processing pipeline to quantitatively analyse various aspects of eWhoring. Our pipeline integrates multiple tools to crawl, annotate, and classify material in a semi-automatic way. It builds in precautions to safeguard against significant ethical issues, such as avoiding the researchers' exposure to pornographic material, and legal concerns, which were justified as some of the images were classified as child exploitation material. We use it to perform a longitudinal measurement of eWhoring activities in 10 specialised underground forums from 2008 to 2019. Our study focuses on three of the main eWhoring components: (i) the acquisition and provenance of images; (ii) the financial profits and monetisation techniques; and (iii) a social network analysis of the offenders, including their relationships, interests, and pathways before and after engaging in this fraudulent activity. We provide recommendations, including potential intervention approaches. Sergio Pastrana, Alice Hutchings, Daniel R. Thomas, Juan Tapiador |
Internet Measurement Conference | 1 |
| 2019 | A First Look at the Crypto-Mining Malware Ecosystem: A Decade of Unrestricted WealthabstractIllicit crypto-mining leverages resources stolen from victims to mine cryptocurrencies on behalf of criminals. While recent works have analyzed one side of this threat, i.e.: web-browser cryptojacking, only commercial reports have partially covered binary-based crypto-mining malware. Sergio Pastrana, Guillermo Suarez-Tangil |
Internet Measurement Conference | 1 |
| 2018 | Characterizing Eve: Analysing Cybercrime Actors in a Large Underground Forum
Sergio Pastrana, Alice Hutchings, Andrew Caines, Paula Buttery |
RAID | 1 |
| 2018 | CrimeBB: Enabling Cybercrime Research on Underground Forums at ScaleabstractUnderground forums allow criminals to interact, exchange knowledge, and trade in products and services. They also provide a pathway into cybercrime, tempting the curious to join those already motivated to obtain easy money. Analysing these forums enables us to better understand the behaviours of offenders and pathways into crime. Prior research has been valuable, but limited by a reliance on datasets that are incomplete or outdated. More complete data, going back many years, allows for comprehensive research into the evolution of forums and their users. We describe CrimeBot, a crawler designed around the particular challenges of capturing data from underground forums. CrimeBot is used to update and maintain CrimeBB, a dataset of more than 48m posts made from 1m accounts in 4 different operational forums over a decade. This dataset presents a new opportunity for large-scale and longitudinal analysis using up-to-date information. We illustrate the potential by presenting a case study using CrimeBB, which analyses which activities lead new actors into engagement with cybercrime. CrimeBB is available to other academic researchers under a legal agreement, designed to prevent misuse and provide safeguards for ethical research. Sergio Pastrana, Daniel R. Thomas, Alice Hutchings, Richard Clayton 0001 |
WWW | 1 |
| 2017 | Ethical issues in research using datasets of illicit originabstractWe evaluate the use of data obtained by illicit means against a broad set of ethical and legal issues. Our analysis covers both the direct collection, and secondary uses of, data obtained via illicit means such as exploiting a vulnerability, or unauthorized disclosure. We extract ethical principles from existing advice and guidance and analyse how they have been applied within more than 20 recent peer reviewed papers that deal with illicitly obtained datasets. We find that existing advice and guidance does not address all of the problems that researchers have faced and explain how the papers tackle ethical issues inconsistently, and sometimes not at all. Our analysis reveals not only a lack of application of safeguards but also that legitimate ethical justifications for research are being overlooked. In many cases positive benefits, as well as potential harms, remain entirely unidentified. Few papers record explicit Research Ethics Board (REB) approval for the activity that is described and the justifications given for exemption suggest deficiencies in the REB process. Daniel R. Thomas, Sergio Pastrana, Alice Hutchings, Richard Clayton 0001, Alastair R. Beresford |
Internet Measurement Conference | 2 |
| 2016 | AVRAND: A Software-Based Defense Against Code Reuse Attacks for AVR Embedded Devices
Sergio Pastrana, Juan Tapiador, Guillermo Suarez-Tangil, Pedro Peris-Lopez |
DIMVA | 1 |
| 2016 | PAgIoT - Privacy-preserving Aggregation protocol for Internet of Things
Lorena González-Manzano, José María de Fuentes, Sergio Pastrana, Pedro Peris-Lopez, Luis Hernández Encinas |
J. Netw. Comput. Appl. | 3 |
| 2015 | Probabilistic yoking proofs for large scale IoT systems
José María de Fuentes, Pedro Peris-Lopez, Juan Tapiador, Sergio Pastrana |
Ad Hoc Networks | 4 |
| 2015 | DEFIDNET: A framework for optimal allocation of cyberdefenses in Intrusion Detection Networks
Sergio Pastrana, Juan Tapiador, Agustín Orfila, Pedro Peris-Lopez |
Comput. Networks | 1 |
| 2015 | Power-aware anomaly detection in smartphones: An analysis of on-platform versus externalized operation
Guillermo Suarez-Tangil, Juan Tapiador, Pedro Peris-Lopez, Sergio Pastrana |
Pervasive Mob. Comput. | 4 |
| 2014 | Randomized Anagram revisited
Sergio Pastrana, Agustín Orfila, Juan Tapiador, Pedro Peris-Lopez |
J. Netw. Comput. Appl. | 1 |
| 2012 | Evaluation of classification algorithms for intrusion detection in MANETs
Sergio Pastrana, Aikaterini Mitrokotsa, Agustín Orfila, Pedro Peris-Lopez |
Knowl. Based Syst. | 1 |
| 2011 | Artificial Immunity-based Correlation System
Guillermo Suarez-Tangil, Esther Palomar, Sergio Pastrana, Arturo Ribagorda |
SECRYPT | 3 |