Marc Sánchez Artigas

dblp:69/1023 · DBLP profile ↗
← Back
72ranked-venue papers
22as first author
15since 2021 · last 2026
0000-0002-9700-7318ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 8 first-author · 9 since 2021Computer networks · 28 · 10 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorSecurity and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Batchmon: Exploiting Serverless Functions for Cost-Effective, SLO-Driven Batch Inference Serving
Josep Calero-Santo, Marc Sánchez Artigas
ICDCS2
2025 Serverless Data Analytics (Finally) Bridging the Gap: Introducing the Ortzi DataFrame
abstract
Serverless technologies have simplified distributed computing by streamlining resource management and offering out-of-the-box usability of cloud resources. However, serverless computing has yet to fully permeate the broader data analytics community. One of the major reasons causing this slow adoption is the lack of a handy serverless interface for seamlessly running recurrent workloads in the cloud. To fill this gap, we introduce in this work the missing piece in serverless analytics: the Ortzi Dataframe, a practical and intuitive programming abstraction that mirrors pandas DataFrames, so that users can effortlessly run their local, single-threaded Python code at scale in the cloud. Needless to say, such a powerful abstraction is certainly useless if not backed by a serverless analytics system that can operate over it in parallel. For this reason, another major contribution of this paper is a fully-fledged system that can run jobs in parallel across the cloud continuum using the novel Ortzi Dataframes. The new system leverages the specific capabilities of each serverless backend without user intervention. Our evaluation demonstrates that Ortzi enables exploration of the nuanced trade-offs of heterogeneous backends with min-imal programming changes and overhead. By harnessing the seamless nature of Ortzi, we optimize jobs through strategic backend selection, still delivering a user-friendly open source framework for programmers without cloud expertise.
Germán T. Eizaguirre, Marc Hostau, Marc Sánchez Artigas
CLOUD3
2025 Let It Unthread: The Good, the Bad and the Ugly Within Webassembly Portable Multithreading
abstract
With a well-defined low-level virtual instruction set, low memory footprint and fast start-up times, WebAssembly has been strongly positioned as a lightweight alternative to containers. For a long time, one missing piece from WebAssembly standalone runtimes has been the incapacity to run shared-memory parallel code and benefit from multicore execution. With the recent WASI threads proposal, the tantalizing promise of cross-platform highperformance execution is more real than ever. In this article, we explore to what extent this promise is fulfilled by investigating the translation of POSIX threads applications to WebAssembly, and how their execution compares to native code. Using standardized benchmarks and a deep analysis of a popular standalone runtime, we reveal interesting findings on the lack of performance in multithreaded code cross-compiled to WebAssembly. We elaborate on the difficulty of correcting these inefficiencies, and even provide a mitigation to excessive thread locking caused by the default WASI libe memory allocator. Overall, we see WASI threads as a good starting point for the efficient execution of multithreaded code.
Marc Sánchez Artigas, Julen Bohoyo Bengoetxea
CCGrid1
2025 Optimizing WebAssembly Garbage Collection in Go: Performance Insights, Tuning Tips, and Batch Execution Strategies
abstract
WebAssembly is gaining popularity these days as a portable intermediate binary format for programming languages. With a well-specified low-level virtual instruction set, low memory footprint and high-performance virtual machines, it has arisen as a lightweight alternative to containers. Even the tiniest container typically sits somewhere around tens of MBs and requires several hundred MBs of memory to run, which is infeasible in resource-constrained edge devices. Though WebAssembly has been battle-tested for memory unmanaged languages such as C++, there exists a systematic lack of investigation on managed languages such as Golang that utilize a garbage collector (GC). Taking TinyGo as a paradigmatic example, we study in this paper the implications of garbage collection for WebAssembly. We show that WebAssembly introduces some inefficiencies compared to native execution and share tips on how to raise its performance. Yet more compelling, we demonstrate that it is possible to dynamically adjust the GC by rewriting the WebAssembly binaries, and in this way, control the impact of garbage collection on the execution time by trading off extra memory. In edge contexts, this ability is essential to fine-tune the execution of a batch of WebAssembly jobs over the same physical device as shown here for the first time.
Safia Guellil, Marc Sánchez Artigas
ICDCS2
2024 A Seer knows best: Auto-tuned object storage shuffling for serverless analytics
Germán T. Eizaguirre, Marc Sánchez Artigas
J. Parallel Distributed Comput.2
2024 Exploiting inherent elasticity of serverless in algorithms with unbalanced and irregular workloads
abstract
Function-as-a-Service execution model in serverless computing has been successful in running large-scale computations like MapReduce, linear algebra, and machine learning. However, little attention has been given to executing highly-dynamic parallel applications with unbalanced and irregular workloads. These algorithms are difficult to execute with good parallel efficiency due to the challenge of provisioning the required computing resources in time, leading to resource over- and under-provisioning in clusters of static size. We propose that the elasticity and fine-grained “pay-as-you-go model” of the FaaS model can be a key enabler for effectively running these algorithms in the cloud. We use a simple serverless executor pool abstraction, and evaluate it using three algorithms with unbalanced and irregular workloads. Results show that their serverless implementation can outperform a static Spark cluster of large virtual machines by up to 55% with the same cost, and can even outperform a single large virtual machine running locally.
Gerard Finol, Gerard París, Pedro García López, Marc Sánchez Artigas
J. Parallel Distributed Comput.4
2024 MLLess: Achieving cost efficiency in serverless machine learning training
abstract
Function-as-a-Service (FaaS) has raised a growing interest in how to “tame” serverless computing to enable domain-specific use cases such as data-intensive applications and machine learning (ML), to name a few. Recently, several systems have been implemented for training ML models. Certainly, these research articles are significant steps in the correct direction. However, they do not completely answer the nagging question of when serverless ML training can be more cost-effective compared to traditional “serverful” computing. To help in this endeavor, we propose MLLess, a FaaS-based ML training prototype built atop IBM Cloud Functions. To boost cost-efficiency, MLLess implements two innovative optimizations tailored to the traits of serverless computing: on one hand, a significance filter, to make indirect communication more effective, and on the other hand, a scale-in auto-tuner, to reduce cost by benefiting from the FaaS sub-second billing model (often per 100 ms). Our results certify that MLLess can be 15X faster than serverful ML systems [27] at a lower cost for sparse ML models that exhibit fast convergence such as sparse logistic regression and matrix factorization. Furthermore, our results show that MLLess can easily scale out to increasingly large fleets of serverless workers.
Pablo Gimeno Sarroca, Marc Sánchez Artigas
J. Parallel Distributed Comput.2
2024 Corrigendum to "MLLess: Achieving Cost Efficiency in Serverless Machine Learning Training" [Journal of Parallel and Distributed Computing 183 (2024) 104764]
Pablo Gimeno Sarroca, Marc Sánchez Artigas
J. Parallel Distributed Comput.2
2023 Is Performance of Object Storage Predictable for Serverless I/O Workloads? A Comparative Study
abstract
Serverless architectures abstract resource provisioning away from the user. However, this property may be at odds with performance. One example of this is Function as a Service (FaaS), where the lack of network addressability compels developers to resort to serverless storage services such as AWS S3 to share (intermediate) data between the functions. For IO-bound workflows, the literature has shown that the performance of parallel reads and writes highly depends on the level of parallelism. Simply put, both an excess or a deficiency in the number of functions may lead to longer IO times. The good news is that the provisioning of functions is fast. Consequently, it is feasible to auto-provision the serverless functions to the optimal number to minimize IO latency. For this, the performance of object storage must be predictable and consistent. We confirmed this in the past for IBM COS. And in this paper, we show that the same occurs to AWS S3. Concretely, we prove that the optimal level of parallelism for parallel reads and writes can be approximated analytically for AWS S3.
Germán T. Eizaguirre, Marc Sánchez Artigas
ICNP2
2023 On Data Processing through the Lenses of S3 Object Lambda
abstract
Despite that Function-as-a-Service (FaaS) has settled down as one of the fundamental cloud programming models, it is still evolving quickly. Recently, Amazon has introduced S3 Object Lambda, which allows a user-defined function to be automatically invoked to process an object as it is being downloaded from S3. As with any new feature, careful study thereof is the key to elucidate if S3 Object Lambda, or more generally, if inline serverless data processing, is a valuable addition to the cloud. For this reason, we conduct an extensive measurement study of this novel service, in order to characterize its architecture and performance (in terms of coldstart latency, TTFB times, and more). We particularly put an eye on the streaming capabilities of this new form of function, as it may open the door to empower existing serverless systems with stream processing capacities. We discuss the pros and cons of this new capability through several workloads, concluding that S3 Object Lambda can go far beyond its original purpose and be leveraged as a building block for more complex abstractions.
Pablo Gimeno Sarroca, Marc Sánchez Artigas
INFOCOM2
2023 Outsourcing Data Processing Jobs With Lithops
abstract
Unexpectedly, the rise of serverless computing has also collaterally started the “democratization” of massive-scale data parallelism. This new trend heralded by PyWren pursues to enable untrained users to execute single-machine code in the cloud at massive scale through platforms like AWS Lambda. Driven by this vision, this article presentsLithops, which carries forward the pioneering work of PyWren to better exploit the innate parallelism of à la MapReduce tasks atop several Functions-as-a-Service platforms such as AWS Lambda, IBM Cloud Functions, Google Cloud Functions or Knative. Instead of waiting for a cluster to be up and running in the cloud,Lithopsmakes easy the task of spawning hundreds and thousands of cloud functions to execute a large job in a few seconds from start. With Lithops, for instance, users can painlessly perform exploratory data analysis from within a Jupyter notebook, while it is the Lithops’s engine which takes care of launching the parallel cloud functions, loading dependencies, automatically partitioning the data, etc. In this article, we describe the design and innovative features of Lithops and evaluate it using several representative applications, including sentiment analysis, Monte Carlo simulations, and hyperparameter tuning. These applications manifest the Lithops’ ability to scale single-machine code computations to thousands of cores. And very importantly, without the need of booting a cold cluster or keeping a warm cluster for occasional tasks.
Josep Sampé, Marc Sánchez Artigas, Gil Vernik, Ido Yehekzel, Pedro García López
IEEE Trans. Cloud Comput.2
2022 Egeon: Software-Defined Data Protection for Object Storage
abstract
With the growth in popularity of cloud computing, object storage systems (e.g., Amazon S3, OpenStack Swift, Ceph) have gained momentum for their relatively low per-G B costs and high availability. However, as increasingly more sensitive data is being accrued, the need to natively integrate privacy controls into the storage is growing in relevance. Today, due to the poor object storage interface, privacy controls are enforced by data curators with full access to data in the clear. This motivates the need for a new approach to data privacy that can provide strong assurance and control to data owners. To fulfill this need, this paper presents Egeon, a novel software-defined data protection framework for object storage. Egeon enables users to declaratively set privacy policies on how their data can be shared. In the privacy policies, the users can build complex data protection services through the composition of data transformations, which are invoked inline by Egeon upon a read request. As a result, data owners can trivially display multiple views from the same data piece, and modify these views by only updating the policies. And all without restructuring the internals of the underlying object storage system. The Egeon prototype has been built atop OpenStack Swift. Evaluation results shows promise in developing data protection services with little overhead directly into the object store. Further, depending on the amount of data filtered out in the transformed views, end-to-end latency can be low due to the savings in network communication.
Raul Saiz-Laudo, Marc Sánchez Artigas
CCGRID2
2022 A seer knows best: optimized object storage shuffling for serverless analytics
abstract
Serverless platforms offer high resource elasticity and pay-as-you-go billing, making them a compelling choice for data analytics. To craft a "pure" serverless solution, the common practice is to transfer intermediate data between serverless functions via serverless object storage (IBM COS; AWS S3). However, prior works have led to inconclusive results about the performance of object storage, since they have left large margin for optimization. To verify that object storage has been underrated, we design a novel shuffle manager for serverless data analytics termed Seer. Specifically, Seer dynamically chooses between two shuffle algorithms to maximize performance. The algorithm choice is based on some predictive models, and very importantly, without users having to specify intermediate data sizes at the time of the job submission. We integrate Seer with PyWren-IBM [31], a serverless analytics framework, and evaluate it against both serverful (e.g., Spark) and serverless systems (e.g., Google BigQuery). Our results certify that our new shuffle manager can deliver performance improvements over them.
Marc Sánchez Artigas, Germán T. Eizaguirre
Middleware1
2022 Stateful Serverless Computing with Crucial
abstract
Serverless computing greatly simplifies the use of cloud resources. In particular, Function-as-a-Service (FaaS) platforms enable programmers to develop applications as individual functions that can run and scale independently. Unfortunately, applications that require fine-grained support for mutable state and synchronization, such as machine learning (ML) and scientific computing, are notoriously hard to build with this new paradigm. In this work, we aim at bridging this gap. We present Crucial , a system to program highly-parallel stateful serverless applications. Crucial retains the simplicity of serverless computing. It is built upon the key insight that FaaS resembles to concurrent programming at the scale of a datacenter. Accordingly, a distributed shared memory layer is the natural answer to the needs for fine-grained state management and synchronization. Crucial allows to port effortlessly a multi-threaded code base to serverless, where it can benefit from the scalability and pay-per-use model of FaaS platforms. We validate Crucial with the help of micro-benchmarks and by considering various stateful applications. Beyond classical parallel tasks (e.g., a Monte Carlo simulation), these applications include representative ML algorithms such as k -means and logistic regression. Our evaluation shows that Crucial obtains superior or comparable performance to Apache Spark at similar cost (18%–40% faster). We also use Crucial to port (part of) a state-of-the-art multi-threaded ML library to serverless. The ported application is up to 30% faster than with a dedicated high-end server. Finally, we attest that Crucial can rival in performance with a single-machine, multi-threaded implementation of a complex coordination problem. Overall, Crucial delivers all these benefits with less than 6% of changes in the code bases of the evaluated applications.
Daniel Barcelona Pons, Pierre Sutra, Marc Sánchez Artigas, Gerard París, Pedro García López
ACM Trans. Softw. Eng. Methodol.3
2021 Experience Paper: Towards enhancing cost efficiency in serverless machine learning training
abstract
Function-as-a-Service (FaaS) has raised a growing interest in how to "tame" serverless to enable domain-specific use cases such as data-intensive applications and machine learning (ML), to name a few. Recently, several systems have been implemented for training ML models. Certainly, these research articles are significant steps in the correct direction. However, they do not completely answer the nagging question of when serverless ML training can be more cost-effective compared to traditional "serverful" computing. To help in this task, we propose MLLess, a FaaS-based ML training prototype built atop IBM Cloud Functions. To boost cost-efficiency, MLLess implements two key optimizations: a significance filter and a scale-in auto-tuner, and leverages them to specialize model training to the FaaS model. Our results certify that MLLess can be 15X faster than serverful ML systems [24] at a lower cost for ML models (such as sparse logistic regression and matrix factorization) that exhibit fast convergence.
Marc Sánchez Artigas, Pablo Gimeno Sarroca
Middleware1
2020 Serverless Elastic Exploration of Unbalanced Algorithms
abstract
In recent years, serverless computing and, in particular the Function-as-a-Service (Faas) execution model, has proven to be efficient for running parallel computing tasks. However, little attention has been paid to highly-parallel applications with unbalanced and irregular workloads. The main challenge of executing this type of algorithms in the cloud is the difficulty to account for the computing requirements beforehand. This places a burden on scientific users who very often make bad decisions by either overprovisioning resources or inadvertently limiting the parallelism of these algorithms due to resource contention. Our hypothesis is that the elasticity and ease of management of serverless computing can help users avoid such decisions, which may lead to undesirable cost-performance consequences for unbalanced problem spaces. In this work, we show that with a simple serverless executor pool abstraction one can achieve a better cost-performance trade-off than a Spark cluster of static size and large EC2 VMs. To support this conclusion, we evaluate two unbalanced algorithms: the Unbalanced Tree Search (UTS) and the Mandelbrot Set using the Mariani-Silver algorithm. For instance, our serverless implementation of UTS is able to outperform Spark by up to 55% with the same cost. This provides the first concrete evidence that highly-parallel, irregular workloads can be efficiently executed using purely stateless functions with almost zero burden on users - i.e., no need for users to understand non-obvious system-level parameters and optimizations.
Gerard París, Pedro García López, Marc Sánchez Artigas
CLOUD3
2019 Lamda-Flow: Automatic Pushdown of Dataflow Operators Close to the Data
abstract
Modern data analytics infrastructures are composed of physically disaggregated compute and storage clusters. Thus, dataflow analytics engines, such as Apache Spark or Flink, are left with no choice but to transfer datasets to the compute cluster prior to their actual processing. For large data volumes, this becomes problematic, since it involves massive data transfers that exhaust network bandwidth, that waste compute cluster memory, and that may become a performance barrier. To overcome this problem, we present λFlow: a framework for automatically pushing dataflow operators (e.g., map, flatMap, filter, etc.) down onto the storage layer. The novelty of λFlow is that it manages the pushdown granularity at the operator level, which makes it a unique problem. To wit, it requires addressing several challenges, such as how to encapsulate dataflow operators and execute them on the storage cluster, and how to keep track of dependencies such that operators can be pushed down safely onto the storage layer. Our evaluation reports significant reductions in resource usage for a large variety of IO-bound jobs. For instance, λFlow was able to reduce both network bandwidth and memory requirements by 90% in Spark. Our Flink experiments also prove the extensibility of λFlow to other engines.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López, Yosef Moatti, Filip Gluszak
CCGRID2
2019 On the FaaS Track: Building Stateful Distributed Applications with Serverless Architectures
abstract
Serverless computing is an emerging paradigm that greatly simplifies the usage of cloud resources and suits well to many tasks. Most notably, Function-as-a-Service (FaaS) enables programmers to develop cloud applications as individual functions that can run and scale independently. Yet, due to the disaggregation of storage and compute resources in FaaS, applications that require fine-grained support for mutable state and synchronization, such as machine learning and scientific computing, are hard to build.
Daniel Barcelona Pons, Marc Sánchez Artigas, Gerard París, Pierre Sutra, Pedro García López
Middleware2
2019 Software-defined object storage in multi-tenant environments
Raúl Gracia Tinedo, Josep Sampé, Gerard París, Marc Sánchez Artigas, Pedro García López, Yosef Moatti
Future Gener. Comput. Syst.4
2018 StackSync: Attribute-based data sharing in file synchronization services
abstract
Summary Personal Cloud (PC) storage services such as Dropbox or Google Drive have become increasingly popular in the last few years. Unfortunately, these services are still not secure. Even assuming “perfect” data confidentiality, securely sharing a folder in these services is still an issue, and this without mentioning the fact that the existing sharing mechanisms are typically too coarse‐grained by operating at the folder level. This is insufficient in many real situations where it is more natural to grant or deny access to files based on arbitrary user attributes and selected attributes of the file. In this research, we explore these issues and show that fine‐grained access control with strong privacy guarantees is practical in the PC. To investigate the potential practicality, we took our fully fledged open source implementation of a PC system and extended it to support fine‐grained private data sharing through attribute‐based encryption. The result was the first design and implementation of a file synchronization service, called StackSync, where the user retains complete control over his or her own data when sharing it. Never before, attribute‐based encryption had been implemented and tested on a real PC service. Our results show that StackSync is both secure and efficient for the PC.
Marc Sánchez Artigas, Cristian Cotes, Marc Ruiz Rodríguez, Pedro García López
Concurr. Comput. Pract. Exp.1
2018 Giving wings to your data: A first experience of Personal Cloud interoperability
Raúl Gracia Tinedo, Cristian Cotes, Edgar Zamora-Gómez, Genís Ortiz, Adrián Moreno-Martínez, Marc Sánchez Artigas, Pedro García López, Raquel Sánchez, Alberto Gómez 0003, Anastasio Illana
Future Gener. Comput. Syst.6
2018 Enhancing Tree-Based ORAM Using Batched Request Reordering
abstract
We explore a new design space for tree-based oblivious RAM (ORAM) constructions, which has not received much attention from the research community. Concretely, our approach is to dynamically reorder the sequence of input requests into batches, such that the portion of the paths shared by the requests in the batch is maximized. In this way, the amount of data fetched per ORAM access can be significantly diminished, thus saving I/O bandwidth. Our results show that the average performance gain is of between 5%-35% over the baseline ORAM, even in real workloads with causal dependencies, which confirms the practical utility of dynamic batching.
Marc Sánchez Artigas
IEEE Trans. Inf. Forensics Secur.1
2018 BenchBox: A User-Driven Benchmarking Framework for Fat-Client Storage Systems
abstract
In many online storage services, end-users mainly interact with the system via “fat” storage clients that integrate complex functionality. This means that to obtain a complete performance evaluation of one of such systems we may need to generate workloads on the client side that reproduce the behavior of real users. Unfortunately, this remains as an open research challenge today. We present BenchBox: A distributed performance evaluation framework for fat-client storage systems. On the one hand, BenchBox can generate workloads directly in storage clients that mimic users exhibiting a certain behavior, namely, user stereotypes. To this end, the framework enables to plug-in workload models and feed them with compact recipes that capture the behavior of user stereotypes (e.g., storage activity, type of file contents, data sharing links). On the other hand, BenchBox provides researchers with management and monitoringfacilities to deploy experiments and analyze the performance of groups of storage clients. To demonstrate our framework, we equipped BenchBox with a 2-layer workload modelthat reproduces both the activity-e.g., types of operations, frequency- and data-e.g., file sizes, data types-of users in a Personal Cloud. We used this model to generate workloads based on user stereotypes that we identified in real traces (UbuntuOne). Our experiments with public providers show how distinct types of users impact on the performance and efficiency of Personal Clouds, which may guide their optimization.
Raúl Gracia Tinedo, Chenglong Zou, Marc Sánchez Artigas, Pedro García López
IEEE Trans. Parallel Distributed Syst.3
2017 Practical Service Placement Approach for Microservices Architecture
abstract
Community networks (CNs) have gained momentum in the last few years with the increasing number of spontaneously deployed WiFi hotspots and home networks. These networks, owned and managed by volunteers, offer various services to their members and to the public. To reduce the complexity of service deployment, community micro-clouds have recently emerged as a promising enabler for the delivery of cloud services to community users. By putting services closer to consumers, micro-clouds pursue not only a better service performance, but also a low entry barrier for the deployment of mainstream Internet services within the CN. Unfortunately, the provisioning of the services is not so simple. Due to the large and irregular topology, high software and hardware diversity of CNs, it requires of a "careful" placement of micro-clouds and services over the network. To achieve this, this paper proposes to leverage state information about the network to inform service placement decisions, and to do so through a fast heuristic algorithm, which is vital to quickly react to changing conditions. To evaluate its performance, we compare our heuristic with one based on random placement in Guifi.net, the biggest CN worldwide. Our experimental results show that our heuristic consistently outperforms random placement by 211% in terms of bandwidth gain. We quantify the benefits of our heuristic on a real live video-streaming service, and demonstrate that video chunk losses decrease significantly, attaining a 37% decrease in the loss packet rate. Further, using a popular Web 2.0 service, we demonstrate that the client response times decrease up to an order of magnitude when using our heuristic.
Mennan Selimi, Llorenç Cerdà-Alabern, Marc Sánchez Artigas, Felix Freitag, Luís Veiga
CCGrid3
2017 Crystal: Software-Defined Storage for Multi-Tenant Object Stores
Raúl Gracia Tinedo, Josep Sampé, Edgar Zamora-Gómez, Marc Sánchez Artigas, Pedro García López, Yosef Moatti, Eran Rom
FAST4
2017 CRESON: Callable and Replicated Shared Objects over NoSQL
abstract
In a Cloud environment, the ability to share and persist objects simplifies the design of applications. Storing objects in a NoSQL database ensures their availability and provides scalability to applications. When Object-NoSQL Mapping is performed at the client side, objects that are accessed by several clients are repeatedly converted between their in-memory and serialized representations. This negatively impacts performance and increases replication costs. In this paper, we describe the design of CRESON, a system supporting callable objects over NoSQL, in which application objects are mapped and instantiated directly on the storage nodes. CRESON supports composition by reference and ensures strong consistency. Objects are replicated and maintained coherent using State Machine Replication. The implementation of CRESON leverages the support of a listenable key-value store (LKVS), a novel NoSQL storage abstraction that we introduce in this paper. We discuss the performance and complexity of CRESON with the example of the portage of a personal cloud storage service, initially developed using an object-relational mapping over a sharded PostgreSQL database. Our results show that CRESON offers a simpler programming experience both in terms of learning time and lines of code, while performing better on average and being more scalable.
Pierre Sutra, Etienne Rivière, Cristian Cotes, Marc Sánchez Artigas, Pedro García López, Emmanuel Bernard, William Burns, Galder Zamarreno
ICDCS4
2017 Too Big to Eat: Boosting Analytics Data Ingestion from Object Stores with Scoop
abstract
Extracting value from data stored in object stores,such as OpenStack Swift and Amazon S3, can be problematicin common scenarios where analytics frameworks and objectstores run in physically disaggregated clusters. One of the mainproblems is that analytics frameworks must ingest large amountsof data from the object store prior to the actual computation;this incurs a significant resources and performance overhead. Toovercome this problem, we present Scoop. Scoop enables analyticsframeworks to benefit from the computational resources of objectstores to optimize the execution of analytics jobs. Scoop achievesthis by enabling the addition of ETL-type actions to the dataupload path and by offloading querying functions to the objectstore through a rich and extensible active object storage layer. Asa proof-of-concept, Scoop enables Apache Spark SQL selectionsand projections to be executed close to the data in OpenStackSwift for accelerating analytics workloads of a smart energy gridcompany (GridPocket). Our experiments in a 63-machine clusterwith real IoT data and SQL queries from GridPocket show thatScoop exhibits query execution times up to 30x faster than thetraditional “ingest-then-compute” approach.
Yosef Moatti, Eran Rom, Raúl Gracia Tinedo, Dalit Naor, Doron Chen, Josep Sampé, Marc Sánchez Artigas, Pedro García López, Filip Gluszak, Eric Deschdt, Francesco Pace, Daniele Venzano, Pietro Michiardi
ICDE7
2017 Data-driven serverless functions for object storage
abstract
Traditionally, active storage techniques have been proposed to move computation tasks to storage nodes in order to exploit data locality. However, we argue in this paper that active storage is ill-suited for cloud storage for two reasons: 1. Lack of elasticity: Computing can only scale out with the number of storage nodes; and 2. Resource Contention: Sharing compute resources can produce interferences in the storage system. Serverless computing is now emerging as a promising alternative for ensuring painless scalability, and also, for simplifying the development of disaggregated computing tasks.
Josep Sampé, Marc Sánchez Artigas, Pedro García López, Gerard París
Middleware2
2016 Vertigo: Programmable Micro-controllers for Software-Defined Object Storage
abstract
Software-defined storage (SDS) aims to minimize the complexity of data management in the Cloud. SDS decouples the control plane from the data plane and simplifies the management of the storage system via automated storage policy enforcement. In this paper, we propose a novel SDS framework for Object Storage that allows to decentralize policy enforcement through the deployment of per-object management policies in the storage nodes. As in active storage systems, we leverage the underutilized CPU time in the storage nodes. But our framework goes one step further. It provides a new management abstraction called micro-controllers which operate on objects depending on their state and content, thereby permitting the implementation of sophisticated management policies, such as the automated deletion of an object based on its access history, and even allowing the orchestration of active storage tasks. Our SDS system avoids the massive interception of data flows by moving that logic to the appropriate objects. Furthermore, our extensible model simplifies the customization of Object Storage services. We present in the validation several interesting use cases such as automated deletion, content level access control, and Web prefetching.
Josep Sampé, Pedro García López, Marc Sánchez Artigas
CLOUD3
2016 Improving the QoE in Personal Clouds with Cross-Swarm Bundling
abstract
Personal cloud storage systems, like Dropbox, are revolutionizing the way people think about and access their files. As the prevailing model, these systems use unicast to push file changes to each of the "unsynced" devices. And as a result, they transmit multiple times the same information, once per unsynced device. This puts an unnecessary strain on outgoing bandwidth at the datacenters. One way to address this is to leverage P2P-like content distribution to benefit from user resources at the edges of the Internet. Although protocols like BitTorrent have proven to be effective in this scenario, we go a step further in this work and propose cross-swarm bundling as a mechanism for file distribution. One key contribution of this work is that, instead of using bundling as means to extend the lifetime of swarms, we show that it can be useful to improve the Quality of Experience (QoE). We validate our proposal using a trace of Ubuntu One, a real personal cloud system, obtaining significant improvements on the QoE levels.
Rahma Chaabouni 0002, Marc Sánchez Artigas, Ala Chaabouni, Pedro García López
LCN2
2016 The power of swarming in personal clouds under bandwidth budget
Rahma Chaabouni 0002, Marc Sánchez Artigas, Pedro García López, Lluis Pamies-Juarez
J. Netw. Comput. Appl.2
2015 Dissecting UbuntuOne: Autopsy of a Global-scale Personal Cloud Back-end
abstract
Personal Cloud services, such as Dropbox or Box, have been widely adopted by users. Unfortunately, very little is known about the internal operation and general characteristics of Personal Clouds since they are proprietary services.
Raúl Gracia Tinedo, Yongchao Tian, Josep Sampé, Hamza Harkous, John Lenton, Pedro García López, Marc Sánchez Artigas, Marko Vukolic
Internet Measurement Conference7
2015 Activity Stereotypes, or How to Cope with Disconnection during Trust Bootstrapping
abstract
Trust-based systems have been proposed as means to fight against malicious agents in peer-to-peer networks, volunteer and grid computing systems, among others. However, there still exist some issues that have been generally overlooked in the literature. One of them is the question of whether punishing disconnecting agents is effective. In this paper, we investigate this question for these initial cases where prior direct and reputational evidence is unavailable, what is referred in the literature as trust bootstrapping. First, we demonstrate that there is not a universally optimal penalty for disconnection and that the effectiveness of this punishment is markedly dependent on the uptime and downtime session lengths. Second, to minimize the effects of an improper selection of the disconnection penalty, we propose to incorporate predictions into the trust bootstrapping process. These predictions based on the current activity of the agents shorten the trust bootstrapping time when direct and reputational information is lacking.
Marc Sánchez Artigas, Blas Herrera
IEEE Trans. Parallel Distributed Syst.1
2014 StackSync: bringing elasticity to dropbox-like file synchronization
abstract
The design of elastic file synchronization services like Dropbox is an open and complex issue yet not unveiled by the major commercial providers, as it includes challenges like fine-grained programmable elasticity and efficient change notification to millions of devices. In this paper, we propose a novel architecture for file synchronization which aims to solve the above two major challenges. At the heart of our proposal lies ObjectMQ, a lightweight framework for providing programmatic elasticity to distributed objects using messaging. The efficient use of indirect communication: i) enables programmatic elasticity based on queue message processing, ii) simplifies change notifications offering simple unicast and multicast primitives; and iii) provides transparent load balancing based on queues.
Pedro García López, Marc Sánchez Artigas, Sergi Toda, Cristian Cotes, John Lenton
Middleware2
2014 Reducing costs in the personal cloud: Is bittorrent a better bet?
abstract
Lately Personal Cloud storage services, like Drop-box, have emerged as user-centric solutions that provide easy management of the users' data. To meet the requirements of their clients, such services require a huge amount of storage and bandwidth. In an attempt to reduce these costs, we focus on maximizing the benefit that can be driven from the interest of users in the same content by the introduction of the BitTorrent protocol. In general, it is assumed that BitTorrent is only effective for large files and/or large swarms, while the client-server approach is more suited for small files and/or small swarms. However, there is no concrete study on the comparative efficiency of both protocols for small files yet. In this paper, we study the download time and offload ratio in BitTorrent compared to HTTP. Based on this study, we propose an algorithm for the management of these protocols. The choice of the protocol is made based on the prediction of the efficiency of BitTorrent and HTTP for each case. We validate our algorithm on a real trace of the Ubuntu One file service, achieving important savings in the cloud bandwidth without degrading the download time.
Rahma Chaabouni 0002, Marc Sánchez Artigas, Pedro García López
P2P2
2014 eWave: Leveraging Energy-Awareness for In-line Deduplication Clusters
abstract
In-line deduplication clusters provide high throughput and scalable storage/archival services to enterprises and organizations. Unfortunately, high throughput comes at the cost of activating several storage nodes on each request, due to the parallel nature of superchunk routing. This may prevent storage nodes from exploiting disk standby times to preserve energy, even for low load periods. We aim to enable deduplication clusters to exploit load valleys to save up disk energy. To this end, we explore the feasibility of deferred writes, diverted access and workload consolidation in this setting.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López
SYSTOR2
2014 On the interplay between data redundancy and retrieval times in P2P storage systems
Lluis Pamies-Juarez, Marc Sánchez Artigas, Pedro García López, Rubén Mondéjar, Rahma Chaabouni 0002
Comput. Networks2
2014 Giving form to social cloud storage through experimentation: Issues and insights
Raúl Gracia Tinedo, Marc Sánchez Artigas, Aleix Ramírez, Adrián Moreno-Martínez, Xavier León, Pedro García López
Future Gener. Comput. Syst.2
2013 Actively Measuring Personal Cloud Storage
abstract
The Personal Cloud model is a mainstream service that meets the growing demand of millions of users for reliable off-site storage. However, despite their broad adoption, very little is known about the quality of service (QoS) of Personal Clouds. In this paper, we present a measurement study of three major Personal Clouds: DropBox, Box and SugarSync. Actively accessing to free accounts through their REST APIs, we analyzed important aspects to characterize their QoS, such as transfer speed, variability and failure rate. Our measurement, conducted during two months, is the first to deeply analyze many facets of these popular services and reveals new insights, such as important performance differences among providers, the existence of transfer speed daily patterns or sudden service breakdowns. We believe that the present analysis of Personal Clouds is of interest to researchers and developers with diverse concerns about Cloud storage, since our observations can help them to understand and characterize the nature of these services.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Adrián Moreno-Martínez, Cristian Cotes, Pedro García López
IEEE CLOUD2
2013 Cloud-as-a-Gift: Effectively Exploiting Personal Cloud Free Accounts via REST APIs
abstract
Personal Clouds, such as DropBox and Box, provide open REST APIs for developers to create clever applications that make their service even more attractive. These APIs are a powerful abstraction that makes it possible for applications to transparently manage data from user accounts, blurring the lines between a Personal Cloud service and storage IaaS. Jointly, Personal Clouds also offer free accounts to lure new users, that normally include reduced storage space and unlimited transfers. However, the unintended consequence of combining open APIs and free accounts is that these companies are exposing automated access to a free storage infrastructure, which may lead to abuse by malicious parties. By exploiting the freemium API service, users may fraudulently consume resources or they can use free accounts as a Cloud storage layer to support abusive applications. We call this vulnerability the storage leeching problem. In this paper, we show how easy it is to implement a file-sharing application able to distribute digital content by abusing Personal Clouds. Making use of open APIs, this application transparently aggregates the limited-space free accounts from multiple providers into a single larger storage layer, while achieving better transfer speed than that received from one provider alone. This demonstrates that free accounts can be easily exploited to obtain a practical Cloud storage service, and therefore, the potential impact of storage leeching.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López
IEEE CLOUD2
2013 Boosting content delivery with BitTorrent in online cloud storage services
abstract
In classic storage services, the transfer protocol used is usually HTTP. This means that all download requests are handled by a central server which sends the requested files in a single stream. But, such transfer is limited by the narrowest network condition along the way, or by the server being overloaded by requests from many clients. In this context, a number of studies have tried to combine BitTorrent content distribution technologies with Cloud environments. In fact, the efficiency of the BitTorrent protocol makes it especially suitable for massive content distribution while reducing bandwidth costs in the Cloud.
Rahma Chaabouni 0002, Pedro García López, Marc Sánchez Artigas, Sandra Ferrer-Celma, Carlos Cebrian
P2P3
2013 Understanding the effects of P2P dynamics on trust bootstrapping
Marc Sánchez Artigas, Blas Herrera
Inf. Sci.1
2012 F2Box: Cloudifying F2F Storage Systems with High Availability Correlation
abstract
The increasing popularity of Cloud storage services is leading end-users to store their digital lives (including photos, videos, work documents, etc.) in the Cloud. However, many users are still reluctant to move their data to the Cloud due to the amount of control ceded to Cloud vendors. To let users retain the control over their data, Friend-to-Friend (F2F) storage systems have been presented in the literature as a promising alternative. However, as we show in this paper, pure F2F storage systems present a poor QoS, mainly due to availability correlations, which results in a loss of attractiveness by end users. To overcome this limitation, we propose a hybrid architecture that combines F2F storage systems and the availability of Cloud storage services to let users infer the right balance between user control and quality of service. This architecture, we called it F2BOX, is able to deliver such a balance thanks to the development of a new suite of data transfer scheduling strategies and a new redundancy calculation algorithm. The main feature of this algorithm is that allow users to adjust the amount of redundancy according to the availability patterns exhibited by friends. Our simulation and experimental results (in Amazon S3) demonstrate the high benefits experienced by end users as a result of the "cloudification" of F2F systems.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López
IEEE CLOUD2
2012 FriendBox: A Hybrid F2F Personal Storage Application
abstract
Personal storage is a mainstream service used by millions of users. Among the existing alternatives, Friend-to-Friend (F2F) systems are nowadays an interesting research topic aimed to leverage a secure and private off-site storage service. However, the specific characteristics of F2F storage systems (reduced node degree, correlated availabilities) represent a hard obstacle to their performance. Actually, it is extremely difficult for a F2F system to guarantee an acceptable storage service quality in terms of transference times and data availability to end-users. In this landscape, we propose to resort to the Cloud for improving the storage service of a F2F system. We present FriendBox: a hybrid F2F personal storage system. FriendBox is the first F2F system that efficiently combines resources of trusted friends with Cloud storage for improving the service quality achievable by pure F2F systems. We evaluated FriendBox through a real deployment in our university campus. We demonstrated that FriendBox achieves high transfer performance and flexible user-defined data availability guarantees. Furthermore, we analyzed the costs of FriendBox demonstrating its economic feasibility.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Adrián Moreno-Martínez, Pedro García López
IEEE CLOUD2
2012 Disconnection punishment in trust bootstrapping: Benefits of activity stereotypes
abstract
Trust-based systems have been proposed as means to fight against malicious agents in peer-to-peer networks. However, there still exist some issues that have been generally overlooked in the literature. One of them is the question of whether punishing disconnecting agents is effective. In this paper, we investigate this question for these initial cases where prior direct and reputational evidence is unavailable, what is referred in the literature as trust bootstrapping. First, we demonstrate that there is not a universally optimal penalty for disconnection and that the effectiveness of this punishment is markedly dependent on the uptime and downtime session lengths. Second, to minimize the effects of an inadequate selection of the disconnection penalty, we propose to incorporate predictions into the trust bootstrapping process. These predictions based on the current activity of the agents enhance the selection of potentially long-lived trustees, shortening the trust bootstrapping time when direct and reputational information is lacking.
Marc Sánchez Artigas
P2P1
2012 FRIENDBOX: A cloudified F2F storage application
abstract
Personal storage is a mainstream service used by millions of users. Among the existing alternatives, Friend-to-Friend (F2F) systems are aimed to leverage a secure and private off-site storage service. However, the specific characteristics of these systems (reduced node degree, correlated availabilities) represent a hard obstacle to their performance. We present FriendBox: a hybrid F2F personal storage system that combines resources of trusted friends with Cloud storage for improving the service quality achievable by pure F2F systems.
Adrián Moreno-Martínez, Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López
P2P3
2012 Analysis of data availability in F2F storage systems: When correlations matter
abstract
Nowadays, the growing necessity for secure and private off-site storage motivates the appearance of novel storage infrastructures. In this sense, it is increasingly common to find storage systems where users interact just with a set of trustworthy participants, such as in Friend-to-Friend (F2F) networks.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López
P2P2
2012 Sophia: A local trust system to secure key-based routing in non-deterministic DHTs
Raúl Gracia Tinedo, Pedro García López, Marc Sánchez Artigas
J. Parallel Distributed Comput.3
2011 Sophia: Local Trust for Securing Routing in DHTs
abstract
Distributed Hash Tables (DHTs) have been used as a common building block in many distributed applications, including Cloud and Grid. However, there are still important security vulnerabilities that hinder their adoption in today'slarge-scale computing platforms. For instance, routing vulnerabilities have been a subject of intensive research but existing solutions rely on redundancy in lieu of improving the quality of routing paths. In this paper, we present Sophia, a novel generic security technique which combines iterative routing with local trust to fortify routing in DHTs. Sophia strictly benefits from first-hand observations about the success/failure of a node's own lookups to improve forwarding paths. Moreover, unlike redundant routing, Sophia dynamically protects routing without introducing additional network overhead. To the best of our knowledge, this is the first work which exploits a local trust system to fortify routing in DHTs. We compared the performance of Sophia with redundant routing in Kademlia DHT. We obtained significant improvements regarding routing resilience, self-adjustment and network traffic reduction.
Raúl Gracia Tinedo, Pedro García López, Marc Sánchez Artigas
CCGRID3
2011 Evaluation of P2P Systems under Different Churn Models: Why We Should Bother
Marc Sánchez Artigas, Enrique Fernández-Casado
Euro-Par (1)1
2011 SocialHelpers: Introducing social trust to ameliorate churn in P2P reputation systems
abstract
Reputation systems rely on historical information to account for uncertainty about the intention of users to cooperate. In peer-to-peer (P2P) systems, however, accumulating experience tends to be slow due to the high rates of churn - the continuous process of arrival and departure of peers. The flow of transactions is continuously interrupted by departures, which can significantly affect the convergence of reputation systems. To shed light on this, this paper presents an accurate model for capturing the influence of churn on the process of building reputations. Using our model, system architects can determine the minimal transaction rate that guarantees fast convergence and design their systems accordingly. Unfortunately, the natural transaction rate of users is sometimes too low (e.g., due to physical constraints like network bandwidth, etc.) that many of them are likely to experience significant delays in the process of building reputations for their neighbors. We face this problem by leveraging the inherent trust in social networks. The basic idea is that users ask their social links to transact with strangers and together generate reputation ratings in a short time scale. Our simulation results report reductions of 50% or greater in the convergence time in environments with high churn rates.
Marc Sánchez Artigas, Blas Herrera
Peer-to-Peer Computing1
2011 Enforcing fairness in P2P storage systems using asymmetric reciprocal exchanges
abstract
In P2P storage systems peers need to contribute some local storage resources in order to obtain a certain online and reliable storage capacity. To guarantee that the storage service works, P2P storage systems have to meet two main requirements. First, the storage system needs to maintain fairness among peers by ensuring that peers consuming more online storage capacity contribute more local storage resources. And second, to reduce redundancy costs and improve reliability, the storage system must incentivize low-available peers to improve their online availability. Traditionally, P2P storage systems achieved these two requirements by (i) using symmetric reciprocal exchanges between peers, and by (ii) allowing peers to selfishly select their set of storage partners. However, in this paper we show that these two mechanisms are suboptimal in terms of the overall storage resources contributed by all peers. To minimize this amount of contributed resources, we design a novel incentive mechanism based on asymmetric reciprocal exchanges between peers. Our mechanism incentivizes peers to select storage partners uniformly at random, and to establish asymmetric exchange relationships with them. These asymmetric exchange relationships allow low-available peers to compensate the increase of redundancy of high-available peers by giving them more storage capacity. We show that our solution reduces the overall amount of contributed resources as well as the resources contributed by each peer individually. Using real P2P availability traces, we show that our incentive mechanism can reduce the overall savings up to 60%, and individual savings from 2% up to 75%, depending on peers' availabilities.
Lluis Pamies-Juarez, Pedro García López, Marc Sánchez Artigas
Peer-to-Peer Computing3
2011 Towards the design of optimal data redundancy schemes for heterogeneous cloud storage infrastructures
Lluis Pamies-Juarez, Pedro García López, Marc Sánchez Artigas, Blas Herrera
Comput. Networks3
2010 Distributed Access Enforcement in P2P Networks: When Privacy Comes into Play
abstract
In open environments such as peer-to-peer networks, the decision to collaborate with multiple users (e.g., by granting access to a resource) is hard to achieve in practice due to extreme decentralization. The literature contains a plethora of examples where a scalable solution for access control is basic to spur their adoption.Motivated by this need, we introduce a novel protocol to enforce access control in a distributed but also privacy-preserving manner - i.e., so as to minimize the disclosure of privileges and of access policies. Privacy is rather scarce in peer-to-peer systems, for which we believe that our protocol is a valuable contribution. Using extensive simulations on top of real Internet topologies, we illustrate the applicability of our protocol, which is efficient both in terms of communication and rounds of interaction.
Marc Sánchez Artigas
Peer-to-Peer Computing1
2010 Affluenza: Towards Universal Churn Generation
abstract
Churn is an inherent property of peer-to-peer (P2P) networks. Despite its relevance, yet, there is not a universal tool to bring researchers the opportunity to compare their contributions under the same general conditions. To fill this gap, we present the first open-source, simulator-independent tool for churn modeling.
Enrique Fernández-Casado, Marc Sánchez Artigas, Pedro García López
Peer-to-Peer Computing2
2010 Availability and Redundancy in Harmony: Measuring Retrieval Times in P2P Storage Systems
abstract
Peer-to-peer (P2P) storage systems are strongly affected by churn - temporal and permanent peer failures. Because of this churn, the main requirement of such systems is to guarantee that stored objects can always be retrieved. This requirement is specially needed in two main situations: when users want to access the stored objects or when data maintenance processes have to repair lost information. To meet this requirement, exiting P2P storage systems introduce large amounts of redundancy that maintain data availability close to 100%. Unfortunately, these large amounts of redundancy increase the storage costs, either by reducing the overall net capacity or by increasing the communication required for data maintenance. In order to minimize storage costs, P2P storage systems can reduce data redundancy. However, less redundancy means lower data availability, which leads to increase object retrieval times. Unfortunately, longer retrieval times could compromise data maintenance processes and could penalize user's retrieval times. It is crucial then for P2P storage systems to predict the effects of a redundancy reduction. In order to provide this information, we present a novel analytical framework to measure object retrieval times under different redundancy and churn circumstances. Our framework can be directly used by backup applications aiming to maintain durability at the lower cost, or by data sharing applications that seek to reduce costs by penalizing user retrieval times. We validate our framework by simulation using real P2P traces (Skype and eMule's KAD).
Lluis Pamies-Juarez, Pedro García López, Marc Sánchez Artigas
Peer-to-Peer Computing3
2010 p2pWeb: An open, decentralized infrastructure of Web servers for sharing ephemeral Web content
Marc Sánchez Artigas, Jordi Pujol Ahulló, Lluis Pamies-Juarez, Pedro García López
Comput. Networks1
2010 Guest editorial for the special issue collaborative P2P systems
Pedro García López, Michael W. Sobolewski, Marc Sánchez Artigas
Comput. Networks3
2010 eSciGrid: A P2P-based e-science Grid for scalable and efficient data sharing
Marc Sánchez Artigas, Pedro García López
Future Gener. Comput. Syst.1
2010 Echo: A peer-to-peer clustering framework for improving communication in DHTs
Marc Sánchez Artigas, Pedro García López
J. Parallel Distributed Comput.1
2009 Exploring the Feasibility of Reputation Models for Improving P2P Routing under Churn
Marc Sánchez Artigas, Pedro García López, Blas Herrera
Euro-Par1
2009 Heterogeneity-Aware Erasure Codes for Peer-to-Peer Storage Systems
abstract
Peer-to-peer (P2P) storage systems rely on data redundancy to obtain high levels of data availability. Among the existing data redundancy schemes, erasure coding is a widely adopted scheme in existing P2P storage systems. By properly tunning its parameters, erasure codes can minimize the required data redundancy, which reduces both the storage and the network overheads. However, to perform this optimization, storage systems need to measured the obtained data availability. Existing P2P storage systems assume homogeneous node availabilities in order to simplify this measurement. As we will prove, this assumption entails efficiency losses when real node availabilities are highly heterogeneous. In this work, we analyze how erasure codes can be optimized in an availability-aware fashion. We propose an analytical framework able to measure data availability more precisely than existing works. As a result, we can optimize the erasure code deployment while reducing its associated overheads. Our experiments show how by considering real node availabilities it is possible to reduce data redundancy about 50% and up to 80% in some specific scenarios.
Lluis Pamies-Juarez, Pedro García López, Marc Sánchez Artigas
ICPP3
2009 On Routing in Distributed Hash Tables: Is Reputation a Shelter from Malicious Behavior and Churn?
abstract
Recently, it has been argued that reputation mechanisms could be used to improve routing by conditioning next-hop decisions to the past behavior of peers. However, churn may severely hinder the applicability of reputations mechanisms. In particular, short peer lifetimes imply that reputations are typically generated from a small number of transactions and are few reliable. To examine how high rates of churn affect reputation systems, we present an analytical model to study the potential damage done by malicious peers together with churn. With our model, we show that it cannot be expected in general that reputations are reliable. We then analyze the impact of this result by proposing a new routing protocol for Chord. Mainly, the protocol exploits reputation to improve the decision about which neighbor select as next-hop peer. Our experimental results show that routing algorithms can obtain important benefits from reputation - even when peer lifetimes are short and the fraction of bad users is moderate.
Marc Sánchez Artigas, Pedro García López
Peer-to-Peer Computing1
2008 Secure Forwarding in DHTs - Is Redundancy the Key to Robustness?
Marc Sánchez Artigas, Pedro García López, Antonio F. Skarmeta
Euro-Par1
2008 Bypass: Providing secure DHT routing through bypassing malicious peers
abstract
Much research in the last years has been devoted to the development of efficient Distributed Hash Tables (DHTs). While many works have studied DHT systems, few have examined their security issues. For example, Chord and other DHT implementations rely on the cooperation of individual peers to route requests. Consequently, any malicious node can drop and misroute messages at will, censoring the access of honest peers to content. In this paper, we introduce Bypass, a novel DHT routing protocol designed to mitigate routing attacks. A key distinguishing feature of Bypass from other implementations is a feedback-based filtering protocol that allows peers to avoid adversarial nodes when routing to the correct holders of a key. Our experimental results show that in principle Bypass can achieve a lookup success rate close to theoretical bounds.
Marc Sánchez Artigas, Pedro García López, Antonio F. Skarmeta
ISCC1
2008 Supporting geographical queries onto DHTs
abstract
Location-based services (LBS) are currently receiving world-wide attention as a consequence of the massive usage of mobile devices, but such location services require scalable distributed infrastructures in order to resolve spatial queries efficiently. We propose a novel methodology to enable geographical query support to distributed hash tables (DHTs). The contributions of our methodology are the followings: a) our technique is DHT-generic, b) it makes an effective clusterization of nodes and information into geographical areas, c) providing data locality without sacrificing routing and data load balancing, d) it is able to answer classical spatial range queries, as well as e) a new kind of queries we call geocast, all of them in a distributed, scalable way. We demonstrate the feasibility of our approach through representative simulations.
Jordi Pujol Ahulló, Pedro García López, Marc Sánchez Artigas, Antonio F. Skarmeta
LCN3
2008 On the Feasibility of Dynamic Superpeer Ratio Maintenance
abstract
The notion of "superpeer" has been shown to be very effective to increase the scalability of P2P applications. For superpeer systems to work, it is critical to preserve the optimal ratio between the number of superpeers and normal peers participating in the overlay. This requires that peers change dynamically their role (i.e., from su-perpeer to normal peer and vice versa) in the presence of node arrivals and departures, a problem that is hard to solve if no peer has global knowledge of the network. In this article, we first investigate the feasibility of superpeer ratio maintenance when each peer can decide to be a superpeer independently of each other. We then show how this problem can be treated as an optimization problem, and we propose a distributed algorithm, based on particle swarm optimization (PSO), to solve it. Our simulation results prove the viability of a PSO-based approach for this problem.
Marc Sánchez Artigas, Pedro García López, Antonio F. Skarmeta
Peer-to-Peer Computing1
2008 TR-clustering: Alleviating the impact of false clustering on P2P overlay networks
Marc Sánchez Artigas, Pedro García López, Antonio F. Skarmeta, José Santa
Comput. Networks1
2008 Architecture and evaluation of a unified V2V and V2I communication system based on cellular networks
José Santa, Antonio F. Skarmeta, Marc Sánchez Artigas
Comput. Commun.3
2007 SQS: Similarity Query Scheme for Peer-to-Peer Databases
abstract
Similarity search is a hot research topic on peer-to-peer systems. In this paper we present SQS, a similarity query scheme for peer-to-peer databases. In this work we provide a novel linearization mechanism that enables structured queries without the burden of a global information maintenance scheme. The system offers exact match and range searches to multidimensional data. SQS employs Cyclone, a hierarchical overlay that is able to build disjoint clusters in terms of network latency and enables data search load balancing by caching per cluster scheme. Finally, we show the good properties of SQS through representative simulation results.
Jordi Pujol Ahulló, Pedro García López, Marc Sánchez Artigas, Antonio F. Skarmeta
ISCC3
2007 A Comparative Study of Hierarchical DHT Systems
abstract
Much research in the last few years has been devoted to development of efficient structured peer-to-peer (P2P) overlay networks, which offer distributed hash table (DHT) functionality. Most of these systems have been devised as flat, non-hierarchical structures, in contrast to the most scalable distributed systems of the past. To cope with this, a significant number of hierarchical DHT designs have been proposed in the literature. Unfortunately, no design is "universally" better. Actually, what is lacking is an analytic framework to identify the good hierarchical design for a given workload. In this paper, we provide such a framework, and we use it to compare the two main hierarchical DHT designs: The homogenous design, in which all nodes act equal roles, against the superpeer design, in which a small subset of peers (i.e., the most powerful and stable), behave as proxies, interconnecting clusters with highly dynamic membership. Our analysis reveals that, on the contrary to what was initially expected, the costs incurred by hierarchical superpeer design are not necessarily minimized.
Marc Sánchez Artigas, Pedro García López, Antonio F. Skarmeta
LCN1
2005 Cyclone: A Novel Design Schema for Hierarchical DHTs
abstract
Recent research efforts have improved the existing flat distributed hash tables to accommodate hierarchical structure. Nevertheless, many problems still remain to be solved regarding scalability issues, autonomous systems, connection degree, and network proximity. In this paper, we present a new hierarchical DHT called Cyclone that aims to solve the aforementioned issues with a near-optimal architecture. Cyclone provides optimal logarithmic routing hops without establishing unnecessary connection links to other nodes. Our approach follows a horizontal and uniform leaf-based approach that considerably reduces the overall number of links per node. Furthermore, Cyclone also offers a disjoint multipath routing scheme that benefits from network proximity and thus creates a more robust overlay infrastructure.
Marc Sánchez Artigas, Pedro García López, Jordi Pujol Ahulló, Antonio F. Skarmeta
Peer-to-Peer Computing1