Rubén Cuevas Rumín

dblp:79/5634 · also Rubén Cuevas · DBLP profile ↗
← Back
52ranked-venue papers
12as first author
10since 2021 · last 2025
0000-0002-1440-8360ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 29 · 8 first-author · 4 since 2021Systems, architecture and hardware · 5 · 2 first-author · 1 since 2021Security and privacy · 5 · 2 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2025 Unveiling Network Performance in the Wild: An Ad-Driven Analysis of Mobile Download Speeds
abstract
Accurate measurement of mobile network performance is crucial for optimizing user experience and ensuring regulatory compliance. Traditional methods like crowdsourcing approaches, though effective, depend heavily on user participation and extensive infrastructure. In this paper, we introduce adNPM, a novel technique for measuring download speed by embedded measurement code in ads displayed across web browsers and mobile apps, without requiring user participation. Through controlled lab tests and real-world deployments in 15 countries, we demonstrate that adNPM achieves accuracy comparable to well-established tools like Speedtest by Ookla and Opensignal while significantly reducing data consumption.
Miguel A. Bermejo-Agueda, Patricia Callejo, Rubén Cuevas Rumín, Ángel Cuevas, Ramakrishnan Durairajan, Reza Rejaie, Álvaro Mayol
WWW3
2024 Analysis and Implementation of Nanotargeting on LinkedIn Based on Publicly Available Non-PII
abstract
The literature has shown that combining a few non-Personal Identifiable Information (non-PII) is enough to make a user unique in a dataset including millions of users. This work demonstrates that a combination of a few non-PII items can be activated to nanotarget users. We demonstrate that the combination of the location and 5 rare (13 random) skills in a LinkedIn profile is enough to become unique in a user base of ∼ 970M users with a probability of 75%. The novelty is that these attributes are publicly accessible to anyone registered on LinkedIn and can be activated through advertising campaigns. We ran an experiment configuring ad campaigns using the location and skills of three of the paper’s authors, demonstrating how all the ads using ≥ 13 skills were delivered exclusively to the targeted user. We reported this vulnerability to LinkedIn, which initially ignored the problem, but fixed it as of November 2023.
Ángel Merino, José González Cabañas, Ángel Cuevas, Rubén Cuevas Rumín
CHI4
2024 Time series clustering with random convolutional kernels
abstract
Abstract Time series data, spanning applications ranging from climatology to finance to healthcare, presents significant challenges in data mining due to its size and complexity. One open issue lies in time series clustering, which is crucial for processing large volumes of unlabeled time series data and unlocking valuable insights. Traditional and modern analysis methods, however, often struggle with these complexities. To address these limitations, we introduce R-Clustering, a novel method that utilizes convolutional architectures with randomly selected parameters. Through extensive evaluations, R-Clustering demonstrates superior performance over existing methods in terms of clustering accuracy, computational efficiency and scalability. Empirical results obtained using the UCR archive demonstrate the effectiveness of our approach across diverse time series datasets. The findings highlight the significance of R-Clustering in various domains and applications, contributing to the advancement of time series data mining.
Jorge Marco-Blanco, Rubén Cuevas Rumín
Data Min. Knowl. Discov.2
2024 FP-tracer: Fine-grained Browser Fingerprinting Detection via Taint-tracking and Entropy-based Thresholds
abstract
Browser fingerprinting is an effective technique to track web users by building a fingerprint from their browser attributes. It is also stealthy because the tracker uses legitimate JavaScript API calls offered by the browser engine, which can be obfuscated before they are sent to a (third-party) server. Current browser fingerprinting methodologies employ coarse-grained collection and classification techniques, such as binary classification of fingerprinters based on the number of non-obfuscated exfiltrated attributes. As a result, they produce inconsistent findings. Meanwhile, the privacy of millions of web users is at risk daily. We address this gap by presenting FP-tracer, a novel methodology to detect and classify browser fingerprinters based on dynamic taint tracking and joint entropy classification. Our methodology enables detecting first- and third-party fingerprinters even when they use obfuscation by tainting attributes, propagating them, and logging when they are leaked (via 62 sources and 25 sinks). Moreover, it discriminates the invasiveness of fingerprinting activities, even from the same service, by measuring the joint entropy of the collected attributes and clustering them. We implement FP-tracer by extending Foxhound, a privacy-oriented Firefox fork with numeric type tainting, more taint tracking sources and sinks, support for multiple sources, and better logging capabilities. We embed our implementation in our automated crawling infrastructure, which is capable of testing websites in parallel using programmable and reproducible logic. We will open-source our implementation. We evaluate FP-tracer by performing a large-scale crawl over the Tranco Top 100K, and detect, amongst others, audio, canvas, and storage fingerprinting on the web. Among others, we find high fingerprinting activities in 8% of domains, with more moderate activity reaching 75%. Notably, fingerprinting is almost five times more likely to be performed by third-party scripts for high activity levels. In addition, we measure that the most severe category of fingerprinting obfuscates 46% of transmitted attributes, and 38% of fingerprinters involve two or more domains. Finally, we find that existing consent banners do not provide an effective defense against browser fingerprinting
Soumaya Boussaha, Lukas Hock, Miguel Bermejo, Rubén Cuevas Rumín, Ángel Cuevas, David Klein 0001, Martin Johns, Luca Compagna, Daniele Antonioli, Thomas Barber
Proc. Priv. Enhancing Technol.4
2024 Overprofiling Analysis on Major Internet Players
abstract
Many Internet services obtain their revenue through the delivery of online advertisements based on the commercial exploitation of users’ profiles. The accuracy and size of these profiles have important implications in terms of advertisers’ campaign performance and users’ privacy. Despite the importance of auditing the profiling accuracy, very little effort has been devoted both in industry and academia. This paper presents the most comprehensive auditing effort to understand the profiling accuracy of four major online advertising platforms: Google, Facebook, Twitter, and LinkedIn. Our work unveils that less than 50% of the assigned interests are relevant. Moreover, platforms can distinguish what interests within the assigned ones are more relevant but hide this information from users and advertisers. Finally, we have proposed a very simple solution that only uses 25 general interests per user. This proposal outperforms all the analyzed platforms in terms of profile accuracy while improving users’ privacy.
Francisco Caravaca, José González Cabañas, Ángel Cuevas, Rubén Cuevas Rumín
Proc. Priv. Enhancing Technol.4
2023 Poster: Analysis of User Uniqueness on LinkedIn Based on Publicly Available Non-PII
abstract
The literature has shown combining a few non-Personal Identifiable Information (non-PII) is enough to make a user unique in a dataset including millions of users. In this work, we demonstrate that the combination of the location and 6 rare (14 random) skills in a LinkedIn profile is enough to become unique in a user base of ~800M users with a probability of 75%. The novelty is these attributes are publicly accessible to anyone registered on LinkedIn and could be activated through advertising campaigns.
Ángel Merino, José González Cabañas, Ángel Cuevas, Rubén Cuevas Rumín
IMC4
2023 A Deep Dive into the Accuracy of IP Geolocation Databases and its Impact on Online Advertising
abstract
The quest for every time more personalized Internet experience relies on the enriched contextual information about each user. Online advertising also follows this approach. Among the context information that advertising stakeholders leverage, location information is certainly one of them. However, when this information is not directly available from the end users, advertising stakeholders infer it using geolocation databases, matching IP addresses to a position on earth. The accuracy of this approach has often been questioned in the past: however, the reality check on an advertising stakeholder shows that this technique accounts for a large fraction of the served advertisements. In this paper, we revisit the work in the field, that is mostly from almost one decade ago, through the lenses of big data. More specifically, we, i) benchmark two commercial Internet geolocation databases, evaluate the quality of their information using a ground-truth database of user positions containing over 2 billion samples, ii) analyze the internals of these databases, devising a theoretical upper bound for the quality of the Internet geolocation approach, and iii) we run an empirical study that unveils the monetary impact of this technology by considering the costs associated with a real-world ad impressions dataset.
Patricia Callejo, Marco Gramaglia, Rubén Cuevas Rumín, Ángel Cuevas
IEEE Trans. Mob. Comput.3
2023 CarbonTag: A Browser-Based Method for Approximating Energy Consumption of Online Ads
abstract
Energy is today the most critical environmental challenge. The amount of carbon emissions contributing to climate change is significantly influenced by both the production and consumption of energy. Measuring and reducing the energy consumption of services is a crucial step toward reducing adverse environmental effects caused by carbon emissions. Millions of websites rely on online advertisements to generate revenue, with most websites earning most or all of their revenues from ads. As a result, hundreds of billions of online ads are delivered daily to internet users to be rendered in their browsers. Both the delivery and rendering of each ad consume energy. This study investigates how much energy online ads use in the rendering process and offers a way for predicting it as part of rendering the ad. To the best of the authors’ knowledge, this is the first study to calculate the energy usage of single advertisements in the rendering process. Our research further introduces different levels of consumption by which online ads can be classified based on energy efficiency. This classification will allow advertisers to add energy efficiency metrics and optimize campaigns towards consuming less possible.
José González Cabañas, Patricia Callejo, Rubén Cuevas Rumín, Steffen Svartberg, Tommy Torjesen, Ángel Cuevas, Antonio Pastor 0002, Mikko Kotila
IEEE Trans. Sustain. Comput.3
2022 Measuring DoH with web ads
abstract
In this paper we present a large measurement study of the impact on the performance of the adoption of HTTPS as a transport for the DNS protocol (DoH) with public resolvers compared to the existent approach of using non-encrypted transport of DNS queries with the resolver services locally provided by ISPs. Using on web-ads as the mean to execute our tests, we perform over 42 million measurements from more than 4 million vantage points distributed in 32 countries and served by over 2,500 ISPs. We find that, the median resolution time increased 17 ms when using DoH with Cloudflare, 41 ms when using DoH with Quad9, 68 ms when using DoH with Google and 170 ms when using DoH with DNS.SB, compared to using Do53 with the local resolver for a non-cached name. We find similar increases even when using caching. The results presented in the paper contribute to the ongoing discussion of the tradeoffs involved in the combined adoption of public resolvers and DoH.
Patricia Callejo, Marcelo Bagnulo, Jaime González Ruiz, Andra Lutu, Alberto García-Martínez, Rubén Cuevas Rumín
Comput. Networks6
2021 Unique on Facebook: formulation and evidence of (nano)targeting individual users with non-PII data
abstract
The privacy of an individual is bounded by the ability of a third party to reveal their identity. Certain data items such as a passport ID or a mobile phone number may be used to uniquely identify a person. These are referred to as Personal Identifiable Information (PII) items. Previous literature has also reported that, in datasets including millions of users, a combination of several non-PII items (which alone are not enough to identify an individual) can uniquely identify an individual within the dataset. In this paper, we define a data-driven model to quantify the number of interests from a user that make them unique on Facebook. To the best of our knowledge, this represents the first study of individuals' uniqueness at the world population scale. Besides, users' interests are actionable non-PII items that can be used to define ad campaigns and deliver tailored ads to Facebook users. We run an experiment through 21 Facebook ad campaigns that target three of the authors of this paper to prove that, if an advertiser knows enough interests from a user, the Facebook Advertising Platform can be systematically exploited to deliver ads exclusively to a specific user. We refer to this practice as nanotargeting. Finally, we discuss the harmful risks associated with nanotargeting such as psychological persuasion, user manipulation, or blackmailing, and provide easily implementable countermeasures to preclude attacks based on nanotargeting campaigns on Facebook.
José González Cabañas, Ángel Cuevas, Rubén Cuevas Rumín, Juan López-Fernández, David García 0001
Internet Measurement Conference3
2019 Q-Tag: a transparent solution to measure ads viewability rate in online advertising campaigns
abstract
Viewability is one of the most important metrics used in ad-tech to measure the performance quality of ad campaigns. The viewability standard defines the visibility conditions an ad impression must meet to achieve a sufficient marketing effect to be considered viewed. The ad-tech industry offers opaque measures of viewability whose performance is questionable. To address this issue, we propose a novel methodology for measuring viewability in ad campaigns. The disclosure of the functional details of this technique makes it reproducible and auditable. Our solution has been deployed in production by a Demand Side Platform (DSP) to measure the viewability rate of the ad campaigns. Leveraging the infrastructure of this DSP, we compare the performance of our methodology with a commercial solution. Both techniques report a similar overall viewability rate of 50%. However, our solution measured the viewability in 93% of the ads served by the DSP, unlike to 74% of the ads measured by the commercial solution. A rough estimation indicates that this increase in the measured rate may lead to a revenue increase of $3.5 million per year for a mid-sized DSP serving 100M of ads per day.
Patricia Callejo, Antonio Pastor 0002, Rubén Cuevas Rumín, Ángel Cuevas
CoNEXT3
2019 Beyond content analysis: detecting targeted ads via distributed counting
abstract
Being able to check whether an online advertisement has been targeted is essential for resolving privacy controversies and implementing in practice data protection regulations like GDPR, CCPA, and COPPA. In this paper we describe the design, implementation, and deployment of an advertisement auditing system called eyeWnder that uses crowdsourcing to reveal in real time whether a display advertisement has been targeted or not. Crowdsourcing simplifies the detection of targeted advertising, but requires reporting to a central repository the impressions seen by different users, thereby jeopardizing their privacy. We break this deadlock with a privacy preserving data sharing protocol that allows eyeWnder to compute global statistics required to detect targeting, while keeping the advertisements seen by individual users and their browsing history private. We conduct a simulation study to explore the effect of different parameters and a live validation to demonstrate the accuracy of our approach. Unlike previous solutions, eyeWnder can even detect indirect targeting, i.e. , marketing campaigns that promote a product or service whose description bears no semantic overlap with its targeted audience.
Costas Iordanou, Nicolas Kourtellis, Juan Miguel Carrascosa, Claudio Soriente, Rubén Cuevas Rumín, Nikolaos Laoutaris
CoNEXT5
2019 Nameles: An intelligent system for Real-Time Filtering of Invalid Ad Traffic
abstract
Invalid ad traffic is an inherent problem of programmatic advertising that has not been properly addressed so far. Traditionally, it has been considered that invalid ad traffic only harms the interests of advertisers, which pay for the cost of invalid ad impressions while other industry stakeholders earn revenue through commissions regardless of the quality of the impression. Our first contribution consists of providing evidence that shows how the Demand Side Platforms (DSPs), one of the most important intermediaries in the programmatic advertising supply chain, may be suffering from economic losses due to invalid ad traffic. Addressing the problem of invalid traffic at DSPs requires a highly scalable solution that can identify invalid traffic in real time at the individual bid request level. The second and main contribution is the design and implementation of a solution for the invalid traffic problem, a system that can be seamlessly integrated into the current programmatic ecosystem by the DSPs. Our system has been released under an open source license, becoming the first auditable solution for invalid ad traffic detection. The intrinsic transparency of our solution along with the good results obtained in industrial trials have led the World Federation of Advertisers to endorse it.
Antonio Pastor 0002, Matti Antero Parssinen, Patricia Callejo, Pelayo Vallina, Rubén Cuevas Rumín, Ángel Cuevas, Mikko Kotila, Arturo Azcorra
WWW5
2018 Unveiling and Quantifying Facebook Exploitation of Sensitive Personal Data for Advertising Purposes
José González Cabañas, Ángel Cuevas, Rubén Cuevas Rumín
USENIX Security Symposium3
2017 FDVT: Data Valuation Tool for Facebook Users
abstract
The OECD, the European Union and other public and private initiatives are claiming for the necessity of tools that create awareness among Internet users about the monetary value associated to the commercial exploitation of their online personal information. This paper presents the first tool addressing this challenge, the Facebook Data Valuation Tool (FDVT). The FDVT provides Facebook users with a personalized and real-time estimation of the revenue they generate for Facebook. Relying on the FDVT, we are able to shed light into several relevant HCI research questions that require a data valuation tool in place. The obtained results reveal that (i) there exists a deep lack of awareness among Internet users regarding the monetary value of personal information, (ii) data valuation tools such as the FDVT are useful means to reduce such knowledge gap, (iii) 1/3 of the users testing the FDVT show a substantial engagement with the tool.
José González Cabañas, Ángel Cuevas, Rubén Cuevas Rumín
CHI3
2017 Opportunities and Challenges of Ad-based Measurements from the Edge of the Network
abstract
For many years, the research community, practitioners, and regulators have used myriad methods and tools to understand the complex structure and behavior of ISPs from the edge of the network. Unfortunately, the nature of these techniques forces the researcher to find a balance between ISP-coverage, user scale, and accuracy. In this paper we present AdTag, a network measurement paradigm that leverages the opportunistic nature of online targeted advertising to measure the Internet from the edge of the network. We discuss and formalize AdTag's design space---including technical, ethical, deployability and economic factors---and its potential to analyze a wide spectrum of Internet connectivity aspects from the browser. We run several experiments to demonstrate that AdTag can be tailored towards geographic and device-based user groups, finding also several challenges to be faced in order to maximize the number of samples. In a 7-day campaign, AdTag could access more than 20K ISPs at a global scale (185 countries) using millions of edge nodes.
Patricia Callejo, Conor Kelton, Narseo Vallina-Rodriguez, Rubén Cuevas Rumín, Oliver Gasser, Christian Kreibich, Florian Wohlfart, Ángel Cuevas
HotNets4
2017 Energy-optimal collaborative file distribution in wired networks
Kshitiz Verma, Gianluca Rizzo, Antonio Fernández 0001, Rubén Cuevas Rumín, Arturo Azcorra, Shmuel Zaks, Alberto García-Martínez
Peer-to-Peer Netw. Appl.4
2016 Your Data in the Eyes of the Beholders: Design of a Unified Data Valuation Portal to Estimate Value of Personal Information from Market Perspective
abstract
Nowadays Internet companies that offer valuable services "for free" are becoming ubiquitous. Users benefiting from these services have to expose their personal information through these services as they utilize them. On the other hand, personal information is becoming a merchandisable commodity, venues that sell personal information by auction are emerging. One of these markets is in the form of advertising systems. Despite being a lucrative business, the hoarding of user personal information by commercial companies is a growing issue primarily because of its non-transparent nature. In this paper we present a data valuation portal that shades light on what kinds of personal information is on market and the financial value of it.
Yonas Mitike Kassa, José González Cabañas, Ángel Cuevas, Rubén Cuevas Rumín, Miriam Marciel, Roberto Gonzalez
ARES4
2016 Independent Auditing of Online Display Advertising Campaigns
abstract
The reported lack of transparency of the online advertising market may seriously affect the interests of advertisers. In this paper, we present a novel methodology that allows advertisers to independently assess the quality of display advertising campaigns. This methodology also serves to audit the accuracy and completeness of reports delivered by the vendor responsible for running a campaign. We have applied our methodology in 8 display ad campaigns configured in Google AdWords, which overall produced 160K ad impressions displayed in more than 7K publishers. Our results reveal that AdWords seems to provide incomplete information to advertisers. Specifically, we found that: (i) AdWords did not report 57% of publishers where ad impressions from our campaigns were delivered, (ii) AdWords reports a large fraction of contextually meaningful impressions based on (non-disclosed) criteria different from the publisher’s theme, (iii) higher CPM investment does not lead to get impressions delivered to more popular publishers, (iv) AdWords does not offer default control of frequency cap, (v) around 10% ad impressions in two of our campaigns were delivered to IP’s from Data Centers. The industry considers these IPs to be likely related to fraud. These findings should contribute to open a debate between advertisers and Ad Tech vendors to standardize the utilization of independent auditing methodologies as the one presented in this work.
Patricia Callejo, Rubén Cuevas Rumín, Ángel Cuevas, Mikko Kotila
HotNets2
2016 Understanding the Detection of View Fraud in Video Content Portals
abstract
While substantial effort has been devoted to understand fraudulent activity in traditional online advertising (search and banner), more recent forms such as video ads have received little attention. The understanding and identification of fraudulent activity (i.e., fake views) in video ads for advertisers, is complicated as they rely exclusively on the detection mechanisms deployed by video hosting portals. In this context, the development of independent tools able to monitor and audit the fidelity of these systems are missing today and needed by both industry and regulators.
Miriam Marciel, Rubén Cuevas Rumín, Albert Banchs, Roberto Gonzalez, Stefano Traverso, Mohamed Ahmed 0001, Arturo Azcorra
WWW2
2016 CSD: A multi-user similarity metric for community recommendation in online social networks
Xiao Han 0001, Leye Wang, Reza Farahbakhsh, Ángel Cuevas, Rubén Cuevas Rumín, Noël Crespi
Expert Syst. Appl.5
2016 Assessing the Evolution of Google+ in Its First Two Years
abstract
In the era when Facebook and Twitter dominate the market for social media, Google has introduced Google+ (G+) and reported a significant growth in its size while others called it a ghost town. This begs the question of whether G+ can really attract a significant number of connected and active users despite the dominance of Facebook and Twitter. This paper presents a detailed longitudinal characterization of G+ based on large-scale measurements. We identify the main components of G+ structure and characterize the key feature of their users and their evolution over time. We then conduct detailed analysis on the evolution of connectivity and activity among users in the largest connected component (LCC) of G+ structure, and compare their characteristics to other major online social networks (OSNs). We show that despite the dramatic growth in the size of G+, the relative size of the LCC has been decreasing and its connectivity has become less clustered. While the aggregate user activity has gradually increased, only a very small fraction of users exhibit any type of activity, and an even smaller fraction of these users attracts any reaction. The identity of users with most followers and reactions reveal that most of them are related to high-tech industry. To our knowledge, this study offers the most comprehensive characterization of G+ based on the largest collected datasets.
Roberto Gonzalez, Rubén Cuevas Rumín, Reza Motamedi, Reza Rejaie, Ángel Cuevas
IEEE/ACM Trans. Netw.2
2015 I always feel like somebody's watching me: measuring online behavioural advertising
abstract
Online Behavioural targeted Advertising (OBA) has risen in prominence as a method to increase the effectiveness of online advertising. OBA operates by associating tags or labels to users based on their online activity and then using these labels to target them. This rise has been accompanied by privacy concerns from researchers, regulators and the press. In this paper, we present a novel methodology for measuring and understanding OBA in the online advertising market. We rely on training artificial online personas representing behavioural traits like 'cooking', 'movies', 'motor sports', etc. and build a measurement system that is automated, scalable and supports testing of multiple configurations. We observe that OBA is a frequent practice and notice that categories valued more by advertisers are more intensely targeted. In addition, we provide evidences showing that the advertising market targets sensitive topics (e.g, religion or health) despite the existence of regulation that bans such practices. We also compare the volume of OBA advertising for our personas in two different geographical locations (US and Spain) and see little geographic bias in terms of intensity of OBA targeting. Finally, we check for targeting with do-not-track (DNT) enabled and discover that DNT is not yet enforced in the web.
Juan Miguel Carrascosa, Jakub Mikians, Rubén Cuevas Rumín, Vijay Erramilli, Nikolaos Laoutaris
CoNEXT3
2015 Deploying Large-Scale Datasets on-Demand in the Cloud: Treats and Tricks on Data Distribution
abstract
Public clouds have democratised the access to analytics for virtually any institution in the world. Virtual machines (VMs) can be provisioned on demand to crunch data after uploading into the VMs. While this task is trivial for a few tens of VMs, it becomes increasingly complex and time consuming when the scale grows to hundreds or thousands of VMs crunching tens or hundreds of TB. Moreover, the elapsed time comes at a price: the cost of provisioning VMs in the cloud and keeping them waiting to load the data. In this paper we present a big data provisioning service that incorporates hierarchical and peer-to-peer data distribution techniques to speed-up data loading into the VMs used for data processing. The system dynamically mutates the sources of the data for the VMs to speed-up data loading. We tested this solution with 1000 VMs and 100 TB of data, reducing time by at least 30 percent over current state of the art techniques. This dynamic topology mechanism is tightly coupled with classic declarative machine configuration techniques (the system takes a single high-level declarative configuration file and configures both software and data loading). Together, these two techniques simplify the deployment of big data in the cloud for end users who may not be experts in infrastructure management.
Luis Miguel Vaquero González, Antonio Celorio, Félix Cuadrado, Rubén Cuevas Rumín
IEEE Trans. Cloud Comput.4
2014 TorrentGuard: Stopping scam and malware distribution in the BitTorrent ecosystem
Rubén Cuevas Rumín, Michal Kryczka, Roberto Gonzalez, Ángel Cuevas, Arturo Azcorra
Comput. Networks1
2014 Computer communications special issue on Green Networking
Michela Meo, Esther Le Rouzic, Rubén Cuevas Rumín, Carmen Guerrero
Comput. Commun.3
2014 Research challenges on energy-efficient networking design
Michela Meo, Esther Le Rouzic, Rubén Cuevas Rumín, Carmen Guerrero
Comput. Commun.3
2014 On exploiting social relationship and personal background for content discovery in P2P networks
abstract
Content discovery is a critical issue in unstructured Peer-to-Peer (P2P) networks as nodes maintain only local network information. However, similarly without global information about human networks, one still can find specific persons via his/her friends by using social information. Therefore, in this paper, we investigate the problem of how social information (i.e., friends and background information) could benefit content discovery in P2P networks. We collect social information of 384,494 user profiles from Facebook, and build a social P2P network model based on the empirical analysis. In this model, we enrich nodes in P2P networks with social information and link nodes via their friendships. Each node extracts two types of social features–Knowledge and Similarity–and assigns more weight to the friends that have higher similarity and more knowledge. Furthermore, we present a novel content discovery algorithm which can explore the latent relationships among a node’s friends. A node computes stable scores for all its friends regarding their weight and the latent relationships. It then selects the top friends with higher scores to query content. Extensive experiments validate performance of the proposed mechanism. In particular, for personal interests searching, the proposed mechanism can achieve 100% of Search Success Rate by selecting the top 20 friends within two-hop. It also achieves 6.5 Hits on average, which improves 8x the performance of the compared methods.
Xiao Han 0001, Ángel Cuevas, Noël Crespi, Rubén Cuevas Rumín, Xiaodi Huang 0001
Future Gener. Comput. Syst.4
2014 Understanding the locality effect in Twitter: measurement and analysis
Rubén Cuevas Rumín, Roberto Gonzalez, Ángel Cuevas, Carmen Guerrero
Pers. Ubiquitous Comput.1
2014 BitTorrent Locality and Transit TrafficReduction: When, Why, and at What Cost?
abstract
A substantial amount of work has recently gone into localizing BitTorrent traffic within an ISP in order to avoid excessive and often times unnecessary transit costs. Several architectures and systems have been proposed and the initial results from specific ISPs and a few torrents have been encouraging. In this work we attempt to deepen and scale our understanding of locality and its potential. Looking at specific ISPs, we consider tens of thousands of concurrent torrents, and thus capture ISP-wide implications that cannot be appreciated by looking at only a handful of torrents. Second, we go beyond individual case studies and present results for few thousands ISPs represented in our data set of up to 40K torrents involving more than 3.9M concurrent peers and more than 20M in the course of a day spread in 11K ASes. Finally, we develop scalable methodologies that allow us to process this huge data set and derive accurate traffic matrices of torrents. Using the previous methods we obtain the following main findings: i) Although there are a large number of very small ISPs without enough resources for localizing traffic, by analyzing the 100 largest ISPs we show that Locality policies are expected to significantly reduce the transit traffic with respect to the default random overlay construction method in these ISPs; ii) contrary to the popular belief, increasing the access speed of the clients of an ISP does not necessarily help to localize more traffic; iii) by studying several real ISPs, we have shown that soft speed-aware locality policies guarantee win-win situations for ISPs and end users. Furthermore, the maximum transit traffic savings that an ISP can achieve without limiting the number of inter-ISP overlay links is bounded by “unlocalizable” torrents with few local clients. The application of restrictions in the number of inter-ISP links leads to a higher transit traffic reduction but the QoS of clients downloading “unlocalizable” torrents would be severely harmed.
Rubén Cuevas Rumín, Nikolaos Laoutaris, Xiaoyuan Yang 0001, Georgos Siganos, Pablo Rodriguez 0001
IEEE Trans. Parallel Distributed Syst.1
2014 Dynamic Data-Centric Storage for long-term storage in Wireless Sensor and Actor Networks
Ángel Cuevas, Manuel Urueña, Gustavo de Veciana, Rubén Cuevas Rumín, Noël Crespi
Wirel. Networks4
2013 Investigating the reaction of BitTorrent content publishers to antipiracy actions
abstract
During recent years, a few countries have put in place online antipiracy laws and there has been some major enforcement actions against violators. This raises the question that to what extent antipiracy actions have been effective in deterring online piracy? This is a challenging issue to explore because of the difficulty to capture user behavior, and to identify the subtle effect of various underlying (and potentially opposing) causes. In this paper, we tackle this question by examining the impact of two major antipiracy actions, the closure of Megaupload and the implementation of the French antipiracy law, on publishers in the largest BitTorrent portal who are major providers of copyrighted content online. We capture snapshots of BitTorrent publishers at proper times relative to the targeted antipiracy event and use the trends in the number and the level of activity of these publishers to assess their reaction to these events. Our investigation illustrates the importance of examining the impact of antipiracy events on different groups of publishers and provides valuable insights on the effect of selected major antipiracy actions on publishers' behavior.
Reza Farahbakhsh, Ángel Cuevas, Rubén Cuevas Rumín, Reza Rejaie, Michal Kryczka, Roberto Gonzalez, Noël Crespi
P2P3
2013 On Weather and Internet Traffic Demand
Juan Camilo Cardona Restrepo, Rade Stanojevic, Rubén Cuevas Rumín
PAM3
2013 Google+ or Google-?: dissecting the evolution of the new OSN in its first year
abstract
In the era when Facebook and Twitter dominate the market for social media, Google has introduced Google+ (G+) and reported a significant growth in its size while others called it a ghost town. This begs the question that "whether G+ can really attract a significant number of connected and active users despite the dominance of Facebook and Twitter?".
Roberto Gonzalez, Rubén Cuevas Rumín, Reza Motamedi, Reza Rejaie, Ángel Cuevas
WWW2
2013 Unveiling the Incentives for Content Publishing in Popular BitTorrent Portals
abstract
BitTorrent is the most popular peer-to-peer (P2P) content delivery application where individual users share various types of content with tens of thousands of other users. The growing popularity of BitTorrent is primarily due to the availability of valuable content without any cost for the consumers. However, apart from the required resources, publishing valuable (and often copyrighted) content has serious legal implications for the users who publish the material. This raises the question that whether (at least major) content publishers behave in an altruistic fashion or have other motives such as financial incentives. In this paper, we identify the content publishers of more than 55 K torrents in two major BitTorrent portals and examine their characteristics. We discover that around 100 publishers are responsible for publishing 67% of the content, which corresponds to 75% of the downloads. Our investigation reveals several key insights about major publishers. First, antipiracy agencies and malicious users publish “fake” files to protect copyrighted content and spread malware, respectively. Second, excluding the fake publishers, content publishing in major BitTorrent portals appears to be largely driven by companies that try to attract consumers to their own Web sites for financial gain. Finally, we demonstrate that profit-driven publishers attract more loyal consumers than altruistic top publishers, whereas the latter have a larger fraction of loyal consumers with a higher degree of loyalty than the former.
Rubén Cuevas Rumín, Michal Kryczka, Ángel Cuevas, Sebastian Kaune, Carmen Guerrero, Reza Rejaie
IEEE/ACM Trans. Netw.1
2012 Greening the Internet: Energy-Optimal File Distribution
abstract
Despite file distribution applications are responsible for a major portion of the current Internet traffic, so far little effort has been dedicated to study file distribution from the point of view of energy efficiency. In this paper, we present the first extensive and detailed theoretical study for the problem of energy efficiency in file distribution. Specifically, we first demonstrate that the general problem of minimizing energy consumption in file distribution is NP-hard. For restricted versions of the problem, we derive tight lower bounds on energy consumption, and we design a family of algorithms that achieve these bounds. Our results prove that through collaborative p2p schemes up to 50% energy savings are achievable with respect to the best available centralized file distribution scheme. Through simulation, we show that even in heterogeneous settings (e.g., considering network congestion, and link variability across hosts) our collaborative algorithms always achieve significant energy savings with respect to the power consumption of centralized file distribution systems.
Kshitiz Verma, Gianluca Rizzo, Antonio Fernández 0001, Rubén Cuevas Rumín, Arturo Azcorra
NCA4
2011 Deep diving into BitTorrent locality
abstract
A substantial amount of work has recently gone into localizing BitTorrent traffic within an ISP in order to avoid excessive and often times unnecessary transit costs. Several architectures and systems have been proposed and the initial results from specific ISPs and a few torrents have been encouraging. In this work we attempt to deepen and scale our understanding of locality and its potential. Looking at specific ISPs, we consider tens of thousands of concurrent torrents, and thus capture ISP-wide implications that cannot be appreciated by looking at only a handful of torrents. Secondly, we go beyond individual case studies and present results for the top 100 ISPs in terms of number of users represented in our dataset of up to 40K torrents involving more than 3.9M concurrent peers and more than 20M in the course of a day spread in 11K ASes. We develop scalable methodologies that allow us to process this huge dataset and get concrete quantitative answers rather than qualitative speculations to questions like: “what is the minimum and the maximum transit traffic reduction across hundreds of ISPs?”, “what are the win-win boundaries for ISPs and their users?”, “what is the maximum amount of transit traffic that can be localized without requiring fine-grained control of inter-AS overlay connections?”.
Rubén Cuevas Rumín, Nikolaos Laoutaris, Xiaoyuan Yang 0001, Georgos Siganos, Pablo Rodriguez 0001
INFOCOM1
2011 Unrevealing the structure of live BitTorrent swarms: Methodology and analysis
abstract
BitTorrent is one of the most popular application in the current Internet. However, we still have little knowledge about the topology of real BitTorrent swarms and how the traffic is actually exchanged among peers. This paper addresses fundamental questions regarding the topology of live BitTorrent swarms. For this purpose we have collected the evolution of the graph topology of 250 real torrents from its birth during a period of 15 days. Using this dataset we first demonstrate that real BitTorrent swarms are neither random graphs nor small world networks. Furthermore, we will see how some factors such as the torrent popularity affect the swarm topology. Secondly, the paper proposes a novel methodology in order to infer the clustered peers in real BitTorrent swarms, something that was not possible so far. Finally, we dedicate special effort to demonstrate that current BitTorrent swarms are experiencing a marked locality phenomenon at the overlay construction level (or connectivity graph). This locality effect is even more pronounced when we consider the exchange traffic relationships between peers. This suggests that an important portion of the BitTorrent traffic is currently confined within the ISPs. This opens a discussion regarding the relative gain of the locality solution proposed so far.
Michal Kryczka, Rubén Cuevas Rumín, Carmen Guerrero, Arturo Azcorra
Peer-to-Peer Computing2
2011 Modelling data-aggregation in multi-replication data centric storage systems for wireless sensor and actor networks
abstract
This paper studies data-centric storage (DCS) as a suitable system to perform data aggregation on wireless sensor and actor networks (WSANs), in which sensor and actor nodes collaborate together in a fully distributed way without any central base station that manages the network or provides connectivity to the outside world. The authors compare different multi-replication DCS proposals and choose the best one to be applied when studying data aggregation. In addition, the authors provide mathematical models for the production, consumption and overall network traffic for different application profiles. Those application profiles are based on the ability of a particular application to perform data aggregation and on what type of traffic is dominant, either the consumption or the production one. Furthermore, the authors provide closed formulas for each application profile that defines the optimal number of replicas that minimise the overall network traffic. Finally, the authors validate the proposed models via simulation.
Ángel Cuevas, Manuel Urueña, Rubén Cuevas Rumín, Ricardo Romeral
IET Commun.3
2010 Is content publishing in BitTorrent altruistic or profit-driven?
abstract
BitTorrent is the most popular P2P content delivery application where individual users share various type of content with tens of thousands of other users. The growing popularity of BitTorrent is primarily due to the availability of valuable content without any cost for the consumers. However, apart from required resources, publishing (sharing) valuable (and often copyrighted) content has serious legal implications for users who publish the material (or publishers). This raises a question that whether (at least major) content publishers behave in an altruistic fashion or have other incentives such as financial. In this study, we identify the content publishers of more than 55K torrents in two major BitTorrent portals and examine their behavior. We demonstrate that a small fraction of publishers is responsible for 67 % of the published content and 75 % of the downloads. Our investigations reveal that these major publishers respond to two different profiles. On the one hand, antipiracy agencies and malicious publishers publish a large amount of fake files to protect copyrighted content and spread malware respectively. On the other hand, content publishing in BitTorrent is largely driven by companies with financial incentives. Therefore, if these companies lose their interest or are unable to publish content, BitTorrent traffic/portals may disappear or at least their associated traffic will be significantly reduced.
Rubén Cuevas Rumín, Michal Kryczka, Ángel Cuevas, Sebastian Kaune, Carmen Guerrero, Reza Rejaie
CoNEXT1
2010 Unraveling BitTorrent's File Unavailability: Measurements and Analysis
abstract
BitTorrent suffers from one fundamental problem: the long-term availability of content. This occurs on a massive-scale with 38% of torrents becoming unavailable within the first month. In this paper we explore this problem by performing two large-scale measurement studies including 46K torrents and 29M users. The studies go significantly beyond any previous work by combining per-node, per-torrent and system-wide observations to ascertain the causes, characteristics and repercussions of file unavailability. The study confirms the conclusion from previous works that seeders have a significant impact on both performance and availability. However, we also present some crucial new findings: (i) the presence of seeders is not the sole factor involved in file availability, (ii) 23.5% of nodes that operate in seedless torrents can finish their downloads, and (iii) BitTorrent availability is discontinuous, operating in cycles of temporary unavailability.
Sebastian Kaune, Rubén Cuevas Rumín, Gareth Tyson, Andreas Mauthe, Carmen Guerrero, Ralf Steinmetz
Peer-to-Peer Computing2
2010 Deep diving into BitTorrent locality
abstract
A substantial amount of work has recently gone into localizing BitTorrent traffic within an ISP in order to avoid excessive and often times unnecessary transit costs. In this work we aim to answer yet unanswered questions such as: what is the minimum and the maximum transit traffic reduction across hundreds of ISPs?, what are the win-win boundaries for ISPs and their users?, what is the maximum amount of transit traffic that can be localized without requiring fine-grained control of inter-AS overlay connections?, what is the impact to transit traffic from upgrades of residential broadband speeds?.
Rubén Cuevas Rumín, Nikolaos Laoutaris, Xiaoyuan Yang 0001, Georgos Siganos, Pablo Rodriguez 0001
SIGMETRICS1
2010 A collaborative P2P scheme for NAT Traversal Server discovery based on topological information
Rubén Cuevas Rumín, Ángel Cuevas, Albert Cabellos-Aparicio, Loránd Jakab, Carmen Guerrero
Comput. Networks1
2009 fP2P-HN: A P2P-Based Route Optimization Solution for Mobile IP and NEMO Clients
abstract
Wireless technologies are rapidly evolving and the users are demanding the possibility of changing its point of attachment to the Internet (i.e. default router) without breaking the IP communications. This can be achieved by using Mobile IP or NEMO, however mobile clients must forward its data packets through its Home Agent (HA) in order to communicate with its peers. This sub-optimal route (lack of route optimization) reduces considerably the communications performance, increases the delay and the infrastructure load. Additionally, since the HA must forward all the mobile clients' data packets, it can become the bottleneck of such networks. In this paper we present the fP2P-HN architecture, a P2P-based solution that allows deploying several HAes throughout the Internet. With this architecture a mobile client can select a closer HA to its topological position in order to reduce the delay of the paths towards its peers. Furthermore it incorporates flexible HAes that, as we will see, reduce the load at these entities. The main challenge of our solution is signaling the location of the HAes in Internet. We provide an analytical model that evaluates the costs and the benefits of the fP2P-HN architecture. The model shows that the signaling grows logarithmically with the number of HAes and that the reduction is, at least, 20% (lower bound).
Albert Cabellos-Aparicio, Rubén Cuevas Rumín, Jordi Domingo-Pascual, Ángel Cuevas, Carmen Guerrero
ICC2
2009 Routing Fairness in Chord: Analysis and Enhancement
abstract
In Peer-to-Peer (P2P) systems where stored objects are small, routing dominates the cost of publishing and retrieving an object. In such systems, the issue of fairly balancing the routing load among all nodes becomes critical. In this paper we address this issue for Chord-based P2P systems. We first present an analytical model to evaluate the routing fairness of Chord based on the well accepted Jain's Fairness Index (FI). Our model shows that Chord performs poorly, with a FI around 0.6, mainly due to the different sizes of the zones between nodes. Following this observation, we propose a simple enhancement to the Chord finger selection algorithm with the goal of mitigating this effect. The key advantage of our proposal as compared to previous approaches is that it does not add any overhead to the basic Chord algorithm. The proposed approach is evaluated analytically showing a very substantial improvement over Chord, with a FI around 0.9. We conduct an extensive large-scale simulation study to evaluate our proposal and validate the analysis. The simulation study includes, among other aspects, churn conditions, heterogeneous nodes and Zipf-like object popularity.
Rubén Cuevas Rumín, Manuel Urueña, Albert Banchs
INFOCOM1
2009 A Hierarchical P2PSIP Architecture to Support Skype-like Services
abstract
A hierarchical DHT overlay architecture based on P2PSIP is proposed to support a skype-like service. The IETF P2PSIP working group is standardizing a protocol to support any DHT in order to deploy services inside a domain. We extend its functionality to allow the interaction between peers of different domains. Furthermore, we perform an analysis of the routing performance and resource consumption under a skype-like scenario where VoIP calls are more likely to happen among users of the same domain.
Isaías Martinez-Yelmo, Carmen Guerrero, Rubén Cuevas Rumín, Andreas Mauthe
PDP3
2009 H-P2PSIP: Interconnection of P2PSIP domains for global multimedia services based on a hierarchical DHT overlay network
Isaías Martinez-Yelmo, Alex Bikfalvi, Rubén Cuevas Rumín, Carmen Guerrero, Jaime García-Reinoso
Comput. Networks3
2009 fP2P-HN: A P2P-based route optimization architecture for mobile IP-based community networks
Rubén Cuevas Rumín, Albert Cabellos-Aparicio, Ángel Cuevas, Jordi Domingo-Pascual, Arturo Azcorra
Comput. Networks1
2008 OnMove: a protocol for content distribution in wireless delay tolerant networks based on social information
abstract
We present OnMove, a protocol for content distribution in wireless delay tolerant networks for use by handheld devices. To improve content distribution, OnMove exploits social characteristics (social similarities and physical encounters) between individuals. We motivate the problem and describe a content sharing protocol based on a ranking algorithm that exploits the social and networking characteristics of individuals.
Rubén Cuevas Rumín, Eva Jaho, Carmen Guerrero, Ioannis Stavrakakis
CoNEXT1
2008 Routing Performance in a Hierarchical DHT-based Overlay Network
abstract
The scalability properties of DHT based overlay networks is considered satisfactory. However, in large scale systems this might still cause a problem since they have a logarithmic complexity depending. Further, they only provide a one dimensional structure and do not make use on inherent clustering properties of some applications (e.g. P2PVoIP or locality aware overlays). Thus, structures based on a hierarchical approach can have performance as well as structural advantages. In this paper, a generic hierarchical architecture based on super-peers is presented where a peer ID is composed by a prefix ID and a suffix ID. Prefix ID is only routed at the super-peer level and the Suffix ID at the peer level. We specifically analyse the Routing Performance of this approach within the context of two specific overlays, viz. CAN and Kademlia.
Isaías Martinez-Yelmo, Rubén Cuevas Rumín, Carmen Guerrero, Andreas Mauthe
PDP2
2008 Multilevel network characterization using regular topologies
José M. Gutiérrez López, Mohamed Imine, Rubén Cuevas Rumín, Jens Myrup Pedersen, Ole Brun Madsen
Comput. Networks3
2007 P2P Based Architecture for Global Home Agent Dynamic Discovery in IP Mobility
abstract
Mobility in packet networks has become a critical issue in the last years. Mobile IP and the network mobility basic support protocol are the IETF proposals to provide mobility. However, both of them introduce performance limitations, due to the presence of an entity (home agent) in the communication path. Those problems have been tried to be solved in different ways. A family of solutions has been proposed in order to mitigate those problems by allowing mobile devices to use several geographically distributed home agents (thus making shorter the communication path). These techniques require a method to discover a close home agent, among those geographically distributed, to the mobile device. This paper proposes a peer-to-peer based solution, called peer-to-peer home agent network, in order to discover a close home agent. The proposed solution is simple, fully global, dynamic and it can be developed in IPv4 and IPv6.
Rubén Cuevas Rumín, Carmen Guerrero, Ángel Cuevas, María Calderón, Carlos J. Bernardos
VTC Spring1