Pelayo Vallina

dblp:204/9456 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0001-7551-7176ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
4 papers
Privacy and data protection · 66% Web and mobile security · 20% Network security · 15%
Computer networks
1 paper
Network measurement and analytics · 50% Content delivery and video streaming · 50%
Databases, data mining, and information retrieval
2 papers
Web and social media mining · 85% Data mining · 15%

Topics — the 5 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Privacy and data protection › privacy compliance
GDPR compliance
0.412019
Tales from the Porn: A Comprehensive Privacy Analysis of the Web Porn Ecosystem · Internet Measurement Conference 2019
Privacy and data protection
regulatory compliance
0.412019
Tales from the Porn: A Comprehensive Privacy Analysis of the Web Porn Ecosystem · Internet Measurement Conference 2019
Privacy and data protection › web tracking
third-party tracking
0.412019
Tales from the Porn: A Comprehensive Privacy Analysis of the Web Porn Ecosystem · Internet Measurement Conference 2019
Privacy and data protection
web tracking
0.412019
Tales from the Porn: A Comprehensive Privacy Analysis of the Web Porn Ecosystem · Internet Measurement Conference 2019
Data mining
anomaly detection
0.112019
Nameles: An intelligent system for Real-Time Filtering of Invalid Ad Traffic · WWW 2019

Methods — techniques the papers use, named apart from their topics

ecosystem characterization · 1.3cross-platform measurement · 1.3large-scale empirical analysis · 0.9label agreement analysis · 0.9scalable system design · 0.8real-time classification · 0.8large-scale web measurement · 0.4fingerprinting detection · 0.4cookie analysis · 0.4
YearPublicationVenuePosition
2024 Reviewing War: Unconventional User Reviews as a Side Channel to Circumvent Information Controls
abstract
During the first days of the 2022 Russian invasion of Ukraine, Russia's media regulator blocked access to many global social media platforms and news sites, including Twitter, Facebook, and the BBC. To bypass the information controls set by Russian authorities, pro-Ukrainian groups explored unconventional ways to reach out to the Russian population, such as posting war-related content in the user reviews of Russian businesses available on Google Maps or Tripadvisor. This paper provides a first analysis of this new phenomenon by analyzing the unconventional strategies used to avoid state censorship in the Russian Federation during the conflict. Specifically, we analyze reviews posted on these platforms from the beginning of the war to September 2022. We measure the channeling of war-related messages through user reviews on Tripadvisor and Google Maps. Our analysis of the content posted on these services reveals that users leveraged these platforms to seek and exchange humanitarian and travel advice, but also to disseminate disinformation and polarized messages. Finally, we analyze the response of platforms in terms of content moderation and their impact.
José Miguel Moreno, Sergio Pastrana, Jens Helge Reelfs, Pelayo Vallina, Savvas Zannettou, Andriy Panchenko 0001, Georgios Smaragdakis, Oliver Hohlfeld, Narseo Vallina-Rodriguez, Juan Tapiador
ICWSM4
2023 Cashing in on Contacts: Characterizing the OnlyFans Ecosystem
abstract
Adult video-sharing has undergone dramatic shifts. New platforms that directly interconnect (often amateur) producers and consumers now allow content creators to promote material across the web and directly monetize the content they produce. OnlyFans is the most prominent example of this new trend. OnlyFans is a content subscription service where creators earn money from users who subscribe to their material. In contrast to prior adult platforms, OnlyFans emphasizes creator-consumer interaction for audience accumulation and maintenance. This results in a wide cross-platform ecosystem geared towards bringing consumers to creators’ accounts. In this paper, we inspect this emerging ecosystem, focusing on content creators and the third-party platforms they connect to.
Pelayo Vallina, Ignacio Castro, Gareth Tyson
WWW1
2021 Blocklist Babel: On the Transparency and Dynamics of Open Source Blocklisting
abstract
Blocklists constitute a widely-used Internet security mechanism to filter undesired network traffic based on IP/domain reputation and behavior. Many blocklists are distributed in open source form by threat intelligence providers who aggregate and process input from their own sensors, but also from third-party feeds or providers. Despite their wide adoption, many open-source blocklist providers lack clear documentation about their structure, curation process, contents, dynamics, and inter-relationships with other providers. In this paper, we perform a transparency and content analysis of 2,093 free and open source blocklists with the aim of exploring those questions. To that end, we perform a longitudinal 6-month crawling campaign yielding more than 13.5M unique records. This allows us to shed light on their nature, dynamics, inter-provider relationships, and transparency. Specifically, we discuss how the lack of consensus on distribution formats, blocklist labeling taxonomy, content focus, and temporal dynamics creates a complex ecosystem that complicates their combined crawling, aggregation and use. We also provide observations regarding their generally low overlap as well as acute differences in terms of liveness (i.e., how frequently records get indexed and removed from the list) and the lack of documentation about their data collection processes, nature and intended purpose. We conclude the paper with recommendations in terms of transparency, accountability, and standardization.
Álvaro Feal, Pelayo Vallina, Julien Gamba, Sergio Pastrana, Antonio Nappa, Oliver Hohlfeld, Narseo Vallina-Rodriguez, Juan Tapiador
IEEE Trans. Netw. Serv. Manag.2
2020 Mis-shapes, Mistakes, Misfits: An Analysis of Domain Classification Services
abstract
Domain classification services have applications in multiple areas, including cybersecurity, content blocking, and targeted advertising. Yet, these services are often a black box in terms of their methodology to classifying domains, which makes it difficult to assess their strengths, aptness for specific applications, and limitations. In this work, we perform a large-scale analysis of 13 popular domain classification services on more than 4.4M hostnames. Our study empirically explores their methodologies, scalability limitations, label constellations, and their suitability to academic research as well as other practical applications such as content filtering. We find that the coverage varies enormously across providers, ranging from over 90% to below 1%. All services deviate from their documented taxonomy, hampering sound usage for research. Further, labels are highly inconsistent across providers, who show little agreement over domains, making it difficult to compare or combine these services. We also show how the dynamics of crowd-sourced efforts may be obstructed by scalability and coverage aspects as well as subjective disagreements among human labelers. Finally, through case studies, we showcase that most services are not fit for detecting specialized content for research or content-blocking purposes. We conclude with actionable recommendations on their usage based on our empirical insights and experience. Particularly, we focus on how users should handle the significant disparities observed across services both in technical solutions and in research.
Pelayo Vallina, Victor Le Pochat, Álvaro Feal, Marius Paraschiv, Julien Gamba, Tim Burke, Oliver Hohlfeld, Juan Tapiador, Narseo Vallina-Rodriguez
Internet Measurement Conference1
2019 Tales from the Porn: A Comprehensive Privacy Analysis of the Web Porn Ecosystem
abstract
Modern privacy regulations, including the General Data Protection Regulation (GDPR) in the European Union, aim to control user tracking activities in websites and mobile applications. These privacy rules typically contain specific provisions and strict requirements for websites that provide sensitive material to end users such as sexual, religious, and health services. However, little is known about the privacy risks that users face when visiting such websites, and about their regulatory compliance. In this paper, we present the first comprehensive and large-scale analysis of 6,843 pornographic websites. We provide an exhaustive behavioral analysis of the use of tracking methods by these websites, and their lack of regulatory compliance, including the absence of age-verification mechanisms and methods to obtain informed user consent. The results indicate that, as in the regular web, tracking is prevalent across pornographic sites: 72% of the websites use third-party cookies and 5% leverage advanced user fingerprinting technologies. Yet, our analysis reveals a third-party tracking ecosystem semi-decoupled from the regular web in which various analytics and advertising services track users across, and outside, pornographic websites. We complete the paper with a regulatory compliance analysis in the context of the EU GDPR, and newer legal requirements to implement verifiable access control mechanisms (e.g., UK's Digital Economy Act). We find that only 16% of the analyzed websites have an accessible privacy policy and only 4% provide a cookie consent banner. The use of verifiable access control mechanisms is limited to prominent pornographic websites.
Pelayo Vallina, Álvaro Feal, Julien Gamba, Narseo Vallina-Rodriguez, Antonio Fernández 0001
Internet Measurement Conference1
2019 Nameles: An intelligent system for Real-Time Filtering of Invalid Ad Traffic
abstract
Invalid ad traffic is an inherent problem of programmatic advertising that has not been properly addressed so far. Traditionally, it has been considered that invalid ad traffic only harms the interests of advertisers, which pay for the cost of invalid ad impressions while other industry stakeholders earn revenue through commissions regardless of the quality of the impression. Our first contribution consists of providing evidence that shows how the Demand Side Platforms (DSPs), one of the most important intermediaries in the programmatic advertising supply chain, may be suffering from economic losses due to invalid ad traffic. Addressing the problem of invalid traffic at DSPs requires a highly scalable solution that can identify invalid traffic in real time at the individual bid request level. The second and main contribution is the design and implementation of a solution for the invalid traffic problem, a system that can be seamlessly integrated into the current programmatic ecosystem by the DSPs. Our system has been released under an open source license, becoming the first auditable solution for invalid ad traffic detection. The intrinsic transparency of our solution along with the good results obtained in industrial trials have led the World Federation of Advertisers to endorse it.
Antonio Pastor 0002, Matti Antero Parssinen, Patricia Callejo, Pelayo Vallina, Rubén Cuevas Rumín, Ángel Cuevas, Mikko Kotila, Arturo Azcorra
WWW4