EDBT 2026 Demo / reviewers in the wild / expert
Pedro García López
dblp:63/4811
· DBLP profile ↗
78ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-9848-1492ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 1 first-author · 8 since 2021Computer networks · 28 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MetaPaper: Massively Scaling Minecraft Worlds With Dynamic Load Balancing
Marina López Alet, Stepan Klymonchuk, Daniel Barcelona Pons, Pedro García López |
IPDPS | 4 |
| 2025 | Quantifying Serverless Elasticity: The gumeter Benchmark Suite
Germán T. Eizaguirre, Enrique Molina-Giménez, Gerard Finol, Carlos Molina 0004, Pedro García López |
ICSOC (1) | 5 |
| 2025 | Burst Computing: Quick, Sudden, Massively Parallel Processing on Serverless Resources
Daniel Barcelona Pons, Aitor Arjona, Pedro García López, Enrique Molina-Giménez, Stepan Klymonchuk |
USENIX ATC | 3 |
| 2025 | Building Stateless Serverless Vector DBs via Block-based Data PartitioningabstractRetrieval-Augmented Generation (RAG) and other AI/ML workloads rely on vector databases (DBs) for efficient analysis of unstructured data. However, cluster (or serverful ) vector DB architectures, such as Milvus, lack the elasticity to handle high workload fluctuations, sparsity, and burstiness. Serverless vector DBs-- i.e., vector DBs built on top of cloud functions--have emerged as a promising alternative architecture, but they are still in their infancy. This paper presents the first experimental study comparing data partitioning strategies in vector DBs built atop stateless Function-as-a-Service (FaaS). Through extensive benchmarks, we reveal key limitations of clustering-based data partitioning when applied to dynamic datasets ( e.g. , complexity, load balancing). We then evaluate a block-based alternative that addresses such limitations ( e.g. , up to 5.8× faster data partitioning, up to 63% lower costs, similar querying times). Moreover, our results show that a stateless serverless vector DB using block-based data partitioning achieves competitive performance with Milvus in several aspects ( e.g. , up to 65.6× faster data partitioning, similar recall), while reducing costs for sparse workloads (up to 99%). Our empirical insights aim to guide the design of next-generation serverless vector DBs. Daniel Barcelona Pons, Raúl Gracia Tinedo, Albert Cañadilla-Domingo, Xavier Roca-Canals, Pedro García López |
Proc. ACM Manag. Data | 5 |
| 2024 | A Cloud-Agnostic Serverless Architecture for Distributed Machine LearningabstractServerless computing has shown vast potential for big data analytics applications, especially involving machine learning algorithms. Nevertheless, little consideration has been given in the literature to cloud-agnostic serverless architectures that leverage existing parallel implementations of machine learning algorithms. This work bridges this gap by proposing a multicloud serverless architecture for distributed machine learning, that enables machine learning engineers without cloud computing expertise to effortlessly port already implemented parallel machine learning algorithms to serverless, whilst overcoming vendor lock-in. In this work, two stateful machine learning algorithms have been ported to serverless, k-means clustering and logistic regression. The serverless implementation of k-means provided superior performance and scalability compared to a serverful implementation when using a number of workers that is equal to or slightly lower than the total number of vCPUs available on the VM running the serverful implementation. Additionally, it achieved an 87-fold speedup compared to a sequential implementation. Moreover, two storage designs of the shared state will be proposed for the serverless implementations, one that requires locks for updating the shared state, and another that is lock-free. Our experimental evaluation demonstrates that the performance of the lock-free serverless implementation of k-means declines with the increase in the number of clusters. Ionut Predoaia, Pedro García López |
BDCAT | 2 |
| 2024 | Dataplug: Unlocking extreme data analytics with on-the-fly dynamic partitioning of unstructured dataabstractThe elasticity of the Cloud is very appealing for processing large scientific data. However, enormous volumes of unstructured research data, totaling petabytes, remain untapped in data repositories due to the lack of efficient parallel data access. Even-sized partitioning of these data to enable its parallel processing requires a complete re-write to storage, becoming prohibitively expensive for high volumes. In this article we present Dataplug, an extensible framework that enables fine-grained parallel data access to unstructured scientific data in object storage. Dataplug employs read-only, format-aware indexing, allowing to define dynamically-sized partitions using various partitioning strategies. This approach avoids writing the partitioned dataset back to storage, enabling distributed workers to fetch data partitions on-the-fly directly from large data blobs, efficiently leveraging the high bandwidth capability of object storage. Validations on genomic (FASTQGZip) and geospatial (LiDAR) data formats demonstrate that Dataplug considerably lowers pre-processing compute costs (between 65.5% — 71.31% less) without imposing significant overheads. Aitor Arjona, Pedro García López, Daniel Barcelona Pons |
CCGrid | 2 |
| 2024 | Enhancing HPC with Serverless Computing: Lithops on MareNostrum5abstractServerless computing offers a novel alternative for developing and deploying applications. By abstracting backend architecture from the user, developers are encouraged to write code without worrying about server management, scaling, or maintenance. The Function-as-a-Service (FaaS) model optimizes this by allowing developers to deploy discrete functions that scale automatically in response to demand without deep concerns about the scalability of the execution infrastructure or platform. Lithops, a multi-cloud serverless computing framework, follows this trend and enables developers to execute Python code across thousands of cloud cores without modifying local scripts. Despite its potential to deploy parallel jobs, Lithops has been designed for running big data jobs over cloud environments, and its applicability to High-Performance Computing (HPC) systems has been unexplored. This paper introduces a novel architecture enabling Lithops deployment in HPC systems like the MareNostrum 5 supercomputer. By leveraging the immense computational power of the HPC-MN5 supercomputer and the FaaS model of Lithops, the architecture aims to offer high performance and scalability while simplifying application coding and deployment. Our evaluations display Lithops' benchmarks over the MareNostrum 5 HPC scale with the number of nodes, outperforming other commercial cloud platforms in terms of Floating Point Operations Per Second (FLOPS) and read-write bandwidth, and avoiding CPU wastage. Andres Benavides Arevalo, Daniel Coll Tejeda, Aaron Call, Pedro García López, Ramon Nou Castell |
ICNP | 4 |
| 2024 | Exploiting inherent elasticity of serverless in algorithms with unbalanced and irregular workloadsabstractFunction-as-a-Service execution model in serverless computing has been successful in running large-scale computations like MapReduce, linear algebra, and machine learning. However, little attention has been given to executing highly-dynamic parallel applications with unbalanced and irregular workloads. These algorithms are difficult to execute with good parallel efficiency due to the challenge of provisioning the required computing resources in time, leading to resource over- and under-provisioning in clusters of static size. We propose that the elasticity and fine-grained “pay-as-you-go model” of the FaaS model can be a key enabler for effectively running these algorithms in the cloud. We use a simple serverless executor pool abstraction, and evaluate it using three algorithms with unbalanced and irregular workloads. Results show that their serverless implementation can outperform a static Spark cluster of large virtual machines by up to 55% with the same cost, and can even outperform a single large virtual machine running locally. Gerard Finol, Gerard París, Pedro García López, Marc Sánchez Artigas |
J. Parallel Distributed Comput. | 3 |
| 2023 | Glider: Serverless Ephemeral Stateful Near-Data ComputationabstractServerless data analytics generate a large amount of intermediate data during computation stages. However, serverless functions, which are short-lived and lack direct communication, face significant challenges in managing this data effectively. The traditional approach of using object storage to carry the data proves to be slow and costly, as it involves constant movement of data back and forth. Although specialized ephemeral storage solutions have been developed to address this issue, they fail to tackle the fundamental challenge of minimizing data movements. This work focuses on incorporating near-data computation into an ephemeral storage system to reduce the volume of transferred data in serverless analytics. We present Glider with the aim to enhance communication between serverless compute stages, allowing data to smoothly "glide" through the processing pipeline instead of bouncing between different services. Glider achieves this by leveraging stateful near-data execution of complex data-bound operations and an efficient I/O streaming interface. Under evaluation, it reduces data transfers by up to 99.7%, improves storage utilization by up to 99.8%, and enhances performance by up to 2.7×. In sum, Glider improves serverless data analytics by optimizing data movement, streamlining processing, and avoiding redundant transfers. Daniel Barcelona Pons, Pedro García López, Bernard Metzler |
Middleware | 2 |
| 2023 | Transparent serverless execution of Python multiprocessing applicationsabstractAccess transparency means that both local and remote resources are accessed using identical operations. With transparency, unmodified single-machine applications could run over disaggregated compute, storage, and memory resources. Hiding the complexity of distributed systems through transparency would have great benefits, like scaling-out local-parallel scientific applications over flexible disaggregated resources in the Cloud. This paper presents a performance evaluation where we assess the feasibility of access transparency over state-of-the-art Cloud disaggregated resources for Python multiprocessing applications. We have interfaced the multiprocessing module with an implementation that transparently runs processes on serverless functions and uses an in-memory data store for shared state. To evaluate transparency, we run in the Cloud four unmodified applications: Uber Research’s Evolution Strategies, Baselines-AI’s Proximal Policy Optimization, Pandaral.lel’s dataframe, and Scikit Learn’s Hyperparameter tuning. We compare execution time and scalability of the same application running over disaggregated resources using our library, with the single-machine Python multiprocessing libraries in a large VM. For equal resources, applications efficiently using message-passing abstractions achieve comparable results despite the significant overheads of remote communication. Other shared-memory intensive applications do not perform due to high remote memory latency. The results show that Python’s multiprocessing library design is an enabler towards transparency: legacy applications using efficient disaggregated abstractions can transparently scale beyond VM limited resources for increased parallelism without changing the underlying code or architecture. Aitor Arjona, Gerard Finol, Pedro García López |
Future Gener. Comput. Syst. | 3 |
| 2023 | Outsourcing Data Processing Jobs With LithopsabstractUnexpectedly, the rise of serverless computing has also collaterally started the “democratization” of massive-scale data parallelism. This new trend heralded by PyWren pursues to enable untrained users to execute single-machine code in the cloud at massive scale through platforms like AWS Lambda. Driven by this vision, this article presentsLithops, which carries forward the pioneering work of PyWren to better exploit the innate parallelism of à la MapReduce tasks atop several Functions-as-a-Service platforms such as AWS Lambda, IBM Cloud Functions, Google Cloud Functions or Knative. Instead of waiting for a cluster to be up and running in the cloud,Lithopsmakes easy the task of spawning hundreds and thousands of cloud functions to execute a large job in a few seconds from start. With Lithops, for instance, users can painlessly perform exploratory data analysis from within a Jupyter notebook, while it is the Lithops’s engine which takes care of launching the parallel cloud functions, loading dependencies, automatically partitioning the data, etc. In this article, we describe the design and innovative features of Lithops and evaluate it using several representative applications, including sentiment analysis, Monte Carlo simulations, and hyperparameter tuning. These applications manifest the Lithops’ ability to scale single-machine code computations to thousands of cores. And very importantly, without the need of booting a cold cluster or keeping a warm cluster for occasional tasks. Josep Sampé, Marc Sánchez Artigas, Gil Vernik, Ido Yehekzel, Pedro García López |
IEEE Trans. Cloud Comput. | 5 |
| 2022 | Stateful Serverless Computing with CrucialabstractServerless computing greatly simplifies the use of cloud resources. In particular, Function-as-a-Service (FaaS) platforms enable programmers to develop applications as individual functions that can run and scale independently. Unfortunately, applications that require fine-grained support for mutable state and synchronization, such as machine learning (ML) and scientific computing, are notoriously hard to build with this new paradigm. In this work, we aim at bridging this gap. We present Crucial , a system to program highly-parallel stateful serverless applications. Crucial retains the simplicity of serverless computing. It is built upon the key insight that FaaS resembles to concurrent programming at the scale of a datacenter. Accordingly, a distributed shared memory layer is the natural answer to the needs for fine-grained state management and synchronization. Crucial allows to port effortlessly a multi-threaded code base to serverless, where it can benefit from the scalability and pay-per-use model of FaaS platforms. We validate Crucial with the help of micro-benchmarks and by considering various stateful applications. Beyond classical parallel tasks (e.g., a Monte Carlo simulation), these applications include representative ML algorithms such as k -means and logistic regression. Our evaluation shows that Crucial obtains superior or comparable performance to Apache Spark at similar cost (18%–40% faster). We also use Crucial to port (part of) a state-of-the-art multi-threaded ML library to serverless. The ported application is up to 30% faster than with a dedicated high-end server. Finally, we attest that Crucial can rival in performance with a single-machine, multi-threaded implementation of a complex coordination problem. Overall, Crucial delivers all these benefits with less than 6% of changes in the code bases of the evaluated applications. Daniel Barcelona Pons, Pierre Sutra, Marc Sánchez Artigas, Gerard París, Pedro García López |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2021 | Triggerflow: Trigger-based orchestration of serverless workflowsabstractAs more applications are being moved to the Cloud thanks to serverless computing, it is increasingly necessary to support the native life cycle execution of those applications in the data center. But existing cloud orchestration systems either focus on short-running workflows (like IBM Composer or Amazon Step Functions Express Workflows) or impose considerable overheads for synchronizing massively parallel jobs (Azure Durable Functions, Amazon Step Functions). None of them are open systems enabling extensible interception and optimization of custom workflows. We present Triggerflow: an extensible Trigger-based Orchestration architecture for serverless workflows. We demonstrate that Triggerflow is a novel serverless building block capable of constructing different reactive orchestrators (State Machines, Directed Acyclic Graphs, Workflow as code, Federated Learning orchestrator). We also validate that it can support high-volume event processing workloads, auto-scale on demand with scale down to zero when not used, and transparently guarantee fault tolerance and efficient resource usage when orchestrating long running scientific workflows. Aitor Arjona, Pedro García López, Josep Sampé, Aleksander Slominski, Lionel Villard |
Future Gener. Comput. Syst. | 2 |
| 2021 | Benchmarking parallelism in FaaS platformsabstractServerless computing has seen a myriad of work exploring its potential. Some systems tackle Function-as-a-Service (FaaS) properties on automatic elasticity and scale to run highly-parallel computing jobs. However, they focus on specific platforms and convey that their ideas can be extrapolated to any FaaS runtime. An important question arises: do all FaaS platforms fit parallel computations? In this paper, we argue that not all of them provide the necessary means to host highly-parallel applications. To validate our hypothesis, we create a comparative framework and categorize the architectures of four cloud FaaS offerings, emphasizing parallel performance. We attest and extend this description with an empirical experiment that consists in plotting in deep detail the evolution of a parallel computing job on each service. The analysis of our results evinces that FaaS is not inherently good for parallel computations and architectural differences across platforms are decisive to categorize their performance. A key insight is the importance of virtualization technologies and the scheduling approach of FaaS platforms. Parallelism improves with lighter virtualization and proactive scheduling due to finer resource allocation and faster elasticity. This causes some platforms like AWS and IBM to perform well for highly-parallel computations, while others such as Azure present difficulties to achieve the required parallelism degree. Consequently, the information in this paper becomes of special interest to help users choose the most adequate infrastructure for their parallel applications. Daniel Barcelona Pons, Pedro García López |
Future Gener. Comput. Syst. | 2 |
| 2020 | Serverless Elastic Exploration of Unbalanced AlgorithmsabstractIn recent years, serverless computing and, in particular the Function-as-a-Service (Faas) execution model, has proven to be efficient for running parallel computing tasks. However, little attention has been paid to highly-parallel applications with unbalanced and irregular workloads. The main challenge of executing this type of algorithms in the cloud is the difficulty to account for the computing requirements beforehand. This places a burden on scientific users who very often make bad decisions by either overprovisioning resources or inadvertently limiting the parallelism of these algorithms due to resource contention. Our hypothesis is that the elasticity and ease of management of serverless computing can help users avoid such decisions, which may lead to undesirable cost-performance consequences for unbalanced problem spaces. In this work, we show that with a simple serverless executor pool abstraction one can achieve a better cost-performance trade-off than a Spark cluster of static size and large EC2 VMs. To support this conclusion, we evaluate two unbalanced algorithms: the Unbalanced Tree Search (UTS) and the Mandelbrot Set using the Mariani-Silver algorithm. For instance, our serverless implementation of UTS is able to outperform Spark by up to 55% with the same cost. This provides the first concrete evidence that highly-parallel, irregular workloads can be efficiently executed using purely stateless functions with almost zero burden on users - i.e., no need for users to understand non-obvious system-level parameters and optimizations. Gerard París, Pedro García López, Marc Sánchez Artigas |
CLOUD | 2 |
| 2019 | Lamda-Flow: Automatic Pushdown of Dataflow Operators Close to the DataabstractModern data analytics infrastructures are composed of physically disaggregated compute and storage clusters. Thus, dataflow analytics engines, such as Apache Spark or Flink, are left with no choice but to transfer datasets to the compute cluster prior to their actual processing. For large data volumes, this becomes problematic, since it involves massive data transfers that exhaust network bandwidth, that waste compute cluster memory, and that may become a performance barrier. To overcome this problem, we present λFlow: a framework for automatically pushing dataflow operators (e.g., map, flatMap, filter, etc.) down onto the storage layer. The novelty of λFlow is that it manages the pushdown granularity at the operator level, which makes it a unique problem. To wit, it requires addressing several challenges, such as how to encapsulate dataflow operators and execute them on the storage cluster, and how to keep track of dependencies such that operators can be pushed down safely onto the storage layer. Our evaluation reports significant reductions in resource usage for a large variety of IO-bound jobs. For instance, λFlow was able to reduce both network bandwidth and memory requirements by 90% in Spark. Our Flink experiments also prove the extensibility of λFlow to other engines. Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López, Yosef Moatti, Filip Gluszak |
CCGRID | 3 |
| 2019 | Please, do not Decentralize the Internet with (Permissionless) Blockchains!abstractThe old mantra of decentralizing the Internet is coming again with fanfare, this time around the blockchain technology hype. We have already seen a technology supposed to change the nature of the Internet: peer-to-peer. The reality is that peer-to-peer naming systems failed, peer-to-peer social networks failed, and yes, peer-to-peer storage failed as well. In this paper, we will review the research on distributed systems in the last few years to identify the limits of open peer-to-peer networks. We will address issues like system complexity, security and frailty, instability and performance. We will show how many of the aforementioned problems also apply to the recent breed of permissionless blockchain networks. The applicability of such systems to mature industrial applications is undermined by the same properties that make them so interesting for a libertarian audience: namely, their openness, their pseudo-anonymity and their unregulated cryptocurrencies. As such, we argue that permissionless blockchain networks are unsuitable to be the substrate for a decentralized Internet. Yet, there is still hope for more decentralization, albeit in a form somewhat limited with respect to the libertarian view of decentralized Internet: in cooperation rather than in competition with the superpowerful datacenters that dominate the world today. This is derived from the recent surge in interest in byzantine fault tolerance and permissioned blockchains, which opens the door to a world where use of trusted third parties is not the only way to arbitrate an ensemble of entities. The ability of establish trust through permissioned blockchains enables to move the control from the datacenters to the edge, truly realizing the promises of edge-centric computing. Pedro García López, Alberto Montresor, Anwitaman Datta |
ICDCS | 1 |
| 2019 | On the FaaS Track: Building Stateful Distributed Applications with Serverless ArchitecturesabstractServerless computing is an emerging paradigm that greatly simplifies the usage of cloud resources and suits well to many tasks. Most notably, Function-as-a-Service (FaaS) enables programmers to develop cloud applications as individual functions that can run and scale independently. Yet, due to the disaggregation of storage and compute resources in FaaS, applications that require fine-grained support for mutable state and synchronization, such as machine learning and scientific computing, are hard to build. Daniel Barcelona Pons, Marc Sánchez Artigas, Gerard París, Pierre Sutra, Pedro García López |
Middleware | 5 |
| 2019 | Software-defined object storage in multi-tenant environments
Raúl Gracia Tinedo, Josep Sampé, Gerard París, Marc Sánchez Artigas, Pedro García López, Yosef Moatti |
Future Gener. Comput. Syst. | 5 |
| 2018 | StackSync: Attribute-based data sharing in file synchronization servicesabstractSummary Personal Cloud (PC) storage services such as Dropbox or Google Drive have become increasingly popular in the last few years. Unfortunately, these services are still not secure. Even assuming “perfect” data confidentiality, securely sharing a folder in these services is still an issue, and this without mentioning the fact that the existing sharing mechanisms are typically too coarse‐grained by operating at the folder level. This is insufficient in many real situations where it is more natural to grant or deny access to files based on arbitrary user attributes and selected attributes of the file. In this research, we explore these issues and show that fine‐grained access control with strong privacy guarantees is practical in the PC. To investigate the potential practicality, we took our fully fledged open source implementation of a PC system and extended it to support fine‐grained private data sharing through attribute‐based encryption. The result was the first design and implementation of a file synchronization service, called StackSync, where the user retains complete control over his or her own data when sharing it. Never before, attribute‐based encryption had been implemented and tested on a real PC service. Our results show that StackSync is both secure and efficient for the PC. Marc Sánchez Artigas, Cristian Cotes, Marc Ruiz Rodríguez, Pedro García López |
Concurr. Comput. Pract. Exp. | 4 |
| 2018 | Giving wings to your data: A first experience of Personal Cloud interoperability
Raúl Gracia Tinedo, Cristian Cotes, Edgar Zamora-Gómez, Genís Ortiz, Adrián Moreno-Martínez, Marc Sánchez Artigas, Pedro García López, Raquel Sánchez, Alberto Gómez 0003, Anastasio Illana |
Future Gener. Comput. Syst. | 7 |
| 2018 | BenchBox: A User-Driven Benchmarking Framework for Fat-Client Storage SystemsabstractIn many online storage services, end-users mainly interact with the system via “fat” storage clients that integrate complex functionality. This means that to obtain a complete performance evaluation of one of such systems we may need to generate workloads on the client side that reproduce the behavior of real users. Unfortunately, this remains as an open research challenge today. We present BenchBox: A distributed performance evaluation framework for fat-client storage systems. On the one hand, BenchBox can generate workloads directly in storage clients that mimic users exhibiting a certain behavior, namely, user stereotypes. To this end, the framework enables to plug-in workload models and feed them with compact recipes that capture the behavior of user stereotypes (e.g., storage activity, type of file contents, data sharing links). On the other hand, BenchBox provides researchers with management and monitoringfacilities to deploy experiments and analyze the performance of groups of storage clients. To demonstrate our framework, we equipped BenchBox with a 2-layer workload modelthat reproduces both the activity-e.g., types of operations, frequency- and data-e.g., file sizes, data types-of users in a Personal Cloud. We used this model to generate workloads based on user stereotypes that we identified in real traces (UbuntuOne). Our experiments with public providers show how distinct types of users impact on the performance and efficiency of Personal Clouds, which may guide their optimization. Raúl Gracia Tinedo, Chenglong Zou, Marc Sánchez Artigas, Pedro García López |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | Crystal: Software-Defined Storage for Multi-Tenant Object Stores
Raúl Gracia Tinedo, Josep Sampé, Edgar Zamora-Gómez, Marc Sánchez Artigas, Pedro García López, Yosef Moatti, Eran Rom |
FAST | 5 |
| 2017 | CRESON: Callable and Replicated Shared Objects over NoSQLabstractIn a Cloud environment, the ability to share and persist objects simplifies the design of applications. Storing objects in a NoSQL database ensures their availability and provides scalability to applications. When Object-NoSQL Mapping is performed at the client side, objects that are accessed by several clients are repeatedly converted between their in-memory and serialized representations. This negatively impacts performance and increases replication costs. In this paper, we describe the design of CRESON, a system supporting callable objects over NoSQL, in which application objects are mapped and instantiated directly on the storage nodes. CRESON supports composition by reference and ensures strong consistency. Objects are replicated and maintained coherent using State Machine Replication. The implementation of CRESON leverages the support of a listenable key-value store (LKVS), a novel NoSQL storage abstraction that we introduce in this paper. We discuss the performance and complexity of CRESON with the example of the portage of a personal cloud storage service, initially developed using an object-relational mapping over a sharded PostgreSQL database. Our results show that CRESON offers a simpler programming experience both in terms of learning time and lines of code, while performing better on average and being more scalable. Pierre Sutra, Etienne Rivière, Cristian Cotes, Marc Sánchez Artigas, Pedro García López, Emmanuel Bernard, William Burns, Galder Zamarreno |
ICDCS | 5 |
| 2017 | Too Big to Eat: Boosting Analytics Data Ingestion from Object Stores with ScoopabstractExtracting value from data stored in object stores,such as OpenStack Swift and Amazon S3, can be problematicin common scenarios where analytics frameworks and objectstores run in physically disaggregated clusters. One of the mainproblems is that analytics frameworks must ingest large amountsof data from the object store prior to the actual computation;this incurs a significant resources and performance overhead. Toovercome this problem, we present Scoop. Scoop enables analyticsframeworks to benefit from the computational resources of objectstores to optimize the execution of analytics jobs. Scoop achievesthis by enabling the addition of ETL-type actions to the dataupload path and by offloading querying functions to the objectstore through a rich and extensible active object storage layer. Asa proof-of-concept, Scoop enables Apache Spark SQL selectionsand projections to be executed close to the data in OpenStackSwift for accelerating analytics workloads of a smart energy gridcompany (GridPocket). Our experiments in a 63-machine clusterwith real IoT data and SQL queries from GridPocket show thatScoop exhibits query execution times up to 30x faster than thetraditional “ingest-then-compute” approach. Yosef Moatti, Eran Rom, Raúl Gracia Tinedo, Dalit Naor, Doron Chen, Josep Sampé, Marc Sánchez Artigas, Pedro García López, Filip Gluszak, Eric Deschdt, Francesco Pace, Daniele Venzano, Pietro Michiardi |
ICDE | 8 |
| 2017 | Data-driven serverless functions for object storageabstractTraditionally, active storage techniques have been proposed to move computation tasks to storage nodes in order to exploit data locality. However, we argue in this paper that active storage is ill-suited for cloud storage for two reasons: 1. Lack of elasticity: Computing can only scale out with the number of storage nodes; and 2. Resource Contention: Sharing compute resources can produce interferences in the storage system. Serverless computing is now emerging as a promising alternative for ensuring painless scalability, and also, for simplifying the development of disaggregated computing tasks. Josep Sampé, Marc Sánchez Artigas, Pedro García López, Gerard París |
Middleware | 3 |
| 2016 | Vertigo: Programmable Micro-controllers for Software-Defined Object StorageabstractSoftware-defined storage (SDS) aims to minimize the complexity of data management in the Cloud. SDS decouples the control plane from the data plane and simplifies the management of the storage system via automated storage policy enforcement. In this paper, we propose a novel SDS framework for Object Storage that allows to decentralize policy enforcement through the deployment of per-object management policies in the storage nodes. As in active storage systems, we leverage the underutilized CPU time in the storage nodes. But our framework goes one step further. It provides a new management abstraction called micro-controllers which operate on objects depending on their state and content, thereby permitting the implementation of sophisticated management policies, such as the automated deletion of an object based on its access history, and even allowing the orchestration of active storage tasks. Our SDS system avoids the massive interception of data flows by moving that logic to the appropriate objects. Furthermore, our extensible model simplifies the customization of Object Storage services. We present in the validation several interesting use cases such as automated deletion, content level access control, and Web prefetching. Josep Sampé, Pedro García López, Marc Sánchez Artigas |
CLOUD | 2 |
| 2016 | Understanding Data Sharing in Private Personal CloudsabstractData sharing in Personal Clouds blurs the lines between on-line storage and content distribution with a strong social component. Such social information may be exploited by researchers to devise optimized data management techniques for Personal Clouds. Unfortunately, due their proprietary nature, data sharing is one of the least studied facets of these systems. In this work, we present the first study of data sharing in a private Personal Cloud. Concretely, we contribute a dataset collected at the metadata back-end of NEC: an enterprise oriented Personal Cloud. First, our analysis provides a deep inspection of the storage layer of NEC, comparing it with a well-known public vendor (UbuntuOne). Second, we study the social structure of NEC user communities, as well as the storage characteristics of user sharing links via multiplex network techniques. Finally, we discuss a battery of data management optimizations for NEC derived from our findings, which may be of independent interest for other similar systems. Our proposals include content distribution, caching and data placement. We believe that both our study and dataset will foster further research in this field. Raúl Gracia Tinedo, Pedro García López, Alberto Gómez 0003, Anastasio Illana |
CLOUD | 2 |
| 2016 | Improving the QoE in Personal Clouds with Cross-Swarm BundlingabstractPersonal cloud storage systems, like Dropbox, are revolutionizing the way people think about and access their files. As the prevailing model, these systems use unicast to push file changes to each of the "unsynced" devices. And as a result, they transmit multiple times the same information, once per unsynced device. This puts an unnecessary strain on outgoing bandwidth at the datacenters. One way to address this is to leverage P2P-like content distribution to benefit from user resources at the edges of the Internet. Although protocols like BitTorrent have proven to be effective in this scenario, we go a step further in this work and propose cross-swarm bundling as a mechanism for file distribution. One key contribution of this work is that, instead of using bundling as means to extend the lifetime of swarms, we show that it can be useful to improve the Quality of Experience (QoE). We validate our proposal using a trace of Ubuntu One, a real personal cloud system, obtaining significant improvements on the QoE levels. Rahma Chaabouni 0002, Marc Sánchez Artigas, Ala Chaabouni, Pedro García López |
LCN | 4 |
| 2016 | The power of swarming in personal clouds under bandwidth budget
Rahma Chaabouni 0002, Marc Sánchez Artigas, Pedro García López, Lluis Pamies-Juarez |
J. Netw. Comput. Appl. | 3 |
| 2015 | Dissecting UbuntuOne: Autopsy of a Global-scale Personal Cloud Back-endabstractPersonal Cloud services, such as Dropbox or Box, have been widely adopted by users. Unfortunately, very little is known about the internal operation and general characteristics of Personal Clouds since they are proprietary services. Raúl Gracia Tinedo, Yongchao Tian, Josep Sampé, Hamza Harkous, John Lenton, Pedro García López, Marc Sánchez Artigas, Marko Vukolic |
Internet Measurement Conference | 6 |
| 2014 | Implicit BPM: A Business Process Platform for Transparent Workflow Weaving
Rubén Mondéjar, Pedro García López, Carles Pairot, Enric Brull |
BPM | 2 |
| 2014 | StackSync: bringing elasticity to dropbox-like file synchronizationabstractThe design of elastic file synchronization services like Dropbox is an open and complex issue yet not unveiled by the major commercial providers, as it includes challenges like fine-grained programmable elasticity and efficient change notification to millions of devices. In this paper, we propose a novel architecture for file synchronization which aims to solve the above two major challenges. At the heart of our proposal lies ObjectMQ, a lightweight framework for providing programmatic elasticity to distributed objects using messaging. The efficient use of indirect communication: i) enables programmatic elasticity based on queue message processing, ii) simplifies change notifications offering simple unicast and multicast primitives; and iii) provides transparent load balancing based on queues. Pedro García López, Marc Sánchez Artigas, Sergi Toda, Cristian Cotes, John Lenton |
Middleware | 1 |
| 2014 | Reducing costs in the personal cloud: Is bittorrent a better bet?abstractLately Personal Cloud storage services, like Drop-box, have emerged as user-centric solutions that provide easy management of the users' data. To meet the requirements of their clients, such services require a huge amount of storage and bandwidth. In an attempt to reduce these costs, we focus on maximizing the benefit that can be driven from the interest of users in the same content by the introduction of the BitTorrent protocol. In general, it is assumed that BitTorrent is only effective for large files and/or large swarms, while the client-server approach is more suited for small files and/or small swarms. However, there is no concrete study on the comparative efficiency of both protocols for small files yet. In this paper, we study the download time and offload ratio in BitTorrent compared to HTTP. Based on this study, we propose an algorithm for the management of these protocols. The choice of the protocol is made based on the prediction of the efficiency of BitTorrent and HTTP for each case. We validate our algorithm on a real trace of the Ubuntu One file service, achieving important savings in the cloud bandwidth without degrading the download time. Rahma Chaabouni 0002, Marc Sánchez Artigas, Pedro García López |
P2P | 3 |
| 2014 | eWave: Leveraging Energy-Awareness for In-line Deduplication ClustersabstractIn-line deduplication clusters provide high throughput and scalable storage/archival services to enterprises and organizations. Unfortunately, high throughput comes at the cost of activating several storage nodes on each request, due to the parallel nature of superchunk routing. This may prevent storage nodes from exploiting disk standby times to preserve energy, even for low load periods. We aim to enable deduplication clusters to exploit load valleys to save up disk energy. To this end, we explore the feasibility of deferred writes, diverted access and workload consolidation in this setting. Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López |
SYSTOR | 3 |
| 2014 | On the interplay between data redundancy and retrieval times in P2P storage systems
Lluis Pamies-Juarez, Marc Sánchez Artigas, Pedro García López, Rubén Mondéjar, Rahma Chaabouni 0002 |
Comput. Networks | 3 |
| 2014 | Giving form to social cloud storage through experimentation: Issues and insights
Raúl Gracia Tinedo, Marc Sánchez Artigas, Aleix Ramírez, Adrián Moreno-Martínez, Xavier León, Pedro García López |
Future Gener. Comput. Syst. | 6 |
| 2013 | Actively Measuring Personal Cloud StorageabstractThe Personal Cloud model is a mainstream service that meets the growing demand of millions of users for reliable off-site storage. However, despite their broad adoption, very little is known about the quality of service (QoS) of Personal Clouds. In this paper, we present a measurement study of three major Personal Clouds: DropBox, Box and SugarSync. Actively accessing to free accounts through their REST APIs, we analyzed important aspects to characterize their QoS, such as transfer speed, variability and failure rate. Our measurement, conducted during two months, is the first to deeply analyze many facets of these popular services and reveals new insights, such as important performance differences among providers, the existence of transfer speed daily patterns or sudden service breakdowns. We believe that the present analysis of Personal Clouds is of interest to researchers and developers with diverse concerns about Cloud storage, since our observations can help them to understand and characterize the nature of these services. Raúl Gracia Tinedo, Marc Sánchez Artigas, Adrián Moreno-Martínez, Cristian Cotes, Pedro García López |
IEEE CLOUD | 5 |
| 2013 | Cloud-as-a-Gift: Effectively Exploiting Personal Cloud Free Accounts via REST APIsabstractPersonal Clouds, such as DropBox and Box, provide open REST APIs for developers to create clever applications that make their service even more attractive. These APIs are a powerful abstraction that makes it possible for applications to transparently manage data from user accounts, blurring the lines between a Personal Cloud service and storage IaaS. Jointly, Personal Clouds also offer free accounts to lure new users, that normally include reduced storage space and unlimited transfers. However, the unintended consequence of combining open APIs and free accounts is that these companies are exposing automated access to a free storage infrastructure, which may lead to abuse by malicious parties. By exploiting the freemium API service, users may fraudulently consume resources or they can use free accounts as a Cloud storage layer to support abusive applications. We call this vulnerability the storage leeching problem. In this paper, we show how easy it is to implement a file-sharing application able to distribute digital content by abusing Personal Clouds. Making use of open APIs, this application transparently aggregates the limited-space free accounts from multiple providers into a single larger storage layer, while achieving better transfer speed than that received from one provider alone. This demonstrates that free accounts can be easily exploited to obtain a practical Cloud storage service, and therefore, the potential impact of storage leeching. Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López |
IEEE CLOUD | 3 |
| 2013 | Boosting content delivery with BitTorrent in online cloud storage servicesabstractIn classic storage services, the transfer protocol used is usually HTTP. This means that all download requests are handled by a central server which sends the requested files in a single stream. But, such transfer is limited by the narrowest network condition along the way, or by the server being overloaded by requests from many clients. In this context, a number of studies have tried to combine BitTorrent content distribution technologies with Cloud environments. In fact, the efficiency of the BitTorrent protocol makes it especially suitable for massive content distribution while reducing bandwidth costs in the Cloud. Rahma Chaabouni 0002, Pedro García López, Marc Sánchez Artigas, Sandra Ferrer-Celma, Carlos Cebrian |
P2P | 2 |
| 2013 | CloudSNAP: A transparent infrastructure for decentralized web deployment using distributed interception
Rubén Mondéjar, Pedro García López, Carles Pairot, Lluis Pamies-Juarez |
Future Gener. Comput. Syst. | 2 |
| 2012 | F2Box: Cloudifying F2F Storage Systems with High Availability CorrelationabstractThe increasing popularity of Cloud storage services is leading end-users to store their digital lives (including photos, videos, work documents, etc.) in the Cloud. However, many users are still reluctant to move their data to the Cloud due to the amount of control ceded to Cloud vendors. To let users retain the control over their data, Friend-to-Friend (F2F) storage systems have been presented in the literature as a promising alternative. However, as we show in this paper, pure F2F storage systems present a poor QoS, mainly due to availability correlations, which results in a loss of attractiveness by end users. To overcome this limitation, we propose a hybrid architecture that combines F2F storage systems and the availability of Cloud storage services to let users infer the right balance between user control and quality of service. This architecture, we called it F2BOX, is able to deliver such a balance thanks to the development of a new suite of data transfer scheduling strategies and a new redundancy calculation algorithm. The main feature of this algorithm is that allow users to adjust the amount of redundancy according to the availability patterns exhibited by friends. Our simulation and experimental results (in Amazon S3) demonstrate the high benefits experienced by end users as a result of the "cloudification" of F2F systems. Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López |
IEEE CLOUD | 3 |
| 2012 | FriendBox: A Hybrid F2F Personal Storage ApplicationabstractPersonal storage is a mainstream service used by millions of users. Among the existing alternatives, Friend-to-Friend (F2F) systems are nowadays an interesting research topic aimed to leverage a secure and private off-site storage service. However, the specific characteristics of F2F storage systems (reduced node degree, correlated availabilities) represent a hard obstacle to their performance. Actually, it is extremely difficult for a F2F system to guarantee an acceptable storage service quality in terms of transference times and data availability to end-users. In this landscape, we propose to resort to the Cloud for improving the storage service of a F2F system. We present FriendBox: a hybrid F2F personal storage system. FriendBox is the first F2F system that efficiently combines resources of trusted friends with Cloud storage for improving the service quality achievable by pure F2F systems. We evaluated FriendBox through a real deployment in our university campus. We demonstrated that FriendBox achieves high transfer performance and flexible user-defined data availability guarantees. Furthermore, we analyzed the costs of FriendBox demonstrating its economic feasibility. Raúl Gracia Tinedo, Marc Sánchez Artigas, Adrián Moreno-Martínez, Pedro García López |
IEEE CLOUD | 4 |
| 2012 | FRIENDBOX: A cloudified F2F storage applicationabstractPersonal storage is a mainstream service used by millions of users. Among the existing alternatives, Friend-to-Friend (F2F) systems are aimed to leverage a secure and private off-site storage service. However, the specific characteristics of these systems (reduced node degree, correlated availabilities) represent a hard obstacle to their performance. We present FriendBox: a hybrid F2F personal storage system that combines resources of trusted friends with Cloud storage for improving the service quality achievable by pure F2F systems. Adrián Moreno-Martínez, Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López |
P2P | 4 |
| 2012 | Analysis of data availability in F2F storage systems: When correlations matterabstractNowadays, the growing necessity for secure and private off-site storage motivates the appearance of novel storage infrastructures. In this sense, it is increasingly common to find storage systems where users interact just with a set of trustworthy participants, such as in Friend-to-Friend (F2F) networks. Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López |
P2P | 3 |
| 2012 | TaKo: Providing transparent collaboration on single-user applications
Rubén Mondéjar, Pedro García López, Enrique Fernández-Casado, Carles Pairot |
Comput. Lang. Syst. Struct. | 2 |
| 2012 | Damon: A distributed AOP middleware for large-scale scenarios
Rubén Mondéjar, Pedro García López, Carles Pairot, Lluis Pamies-Juarez |
Inf. Softw. Technol. | 2 |
| 2012 | Sophia: A local trust system to secure key-based routing in non-deterministic DHTs
Raúl Gracia Tinedo, Pedro García López, Marc Sánchez Artigas |
J. Parallel Distributed Comput. | 2 |
| 2012 | Energy- and Delay-Efficient Routing in Mobile Ad Hoc Networks
Nicola Costagliola, Pedro García López, Francesco Oliviero, Simon Pietro Romano |
Mob. Networks Appl. | 2 |
| 2011 | Sophia: Local Trust for Securing Routing in DHTsabstractDistributed Hash Tables (DHTs) have been used as a common building block in many distributed applications, including Cloud and Grid. However, there are still important security vulnerabilities that hinder their adoption in today'slarge-scale computing platforms. For instance, routing vulnerabilities have been a subject of intensive research but existing solutions rely on redundancy in lieu of improving the quality of routing paths. In this paper, we present Sophia, a novel generic security technique which combines iterative routing with local trust to fortify routing in DHTs. Sophia strictly benefits from first-hand observations about the success/failure of a node's own lookups to improve forwarding paths. Moreover, unlike redundant routing, Sophia dynamically protects routing without introducing additional network overhead. To the best of our knowledge, this is the first work which exploits a local trust system to fortify routing in DHTs. We compared the performance of Sophia with redundant routing in Kademlia DHT. We obtained significant improvements regarding routing resilience, self-adjustment and network traffic reduction. Raúl Gracia Tinedo, Pedro García López, Marc Sánchez Artigas |
CCGRID | 2 |
| 2011 | Improving BitTorrent download times using community partnersabstractIn this paper we present the concept of community downloads as a mechanism to improve the overall performance of BitTorrent clients. A community is a group of nodes interested in the same content working cooperatively inside a swarm. To reinforce this cooperation among community nodes, we designed two new algorithms: Group Rarest-First (piece selection) and Group-Balanced Tit-for-tat (unchoke policy). Our algorithms treat the group as a big node, prioritizing community members and helping them to improve their download ratios. Our validation shows improvements in download time around 20% and up to 71% in different swarm scenarios. Marc Espelt Palau, Pedro García López |
LCN | 2 |
| 2011 | Enforcing fairness in P2P storage systems using asymmetric reciprocal exchangesabstractIn P2P storage systems peers need to contribute some local storage resources in order to obtain a certain online and reliable storage capacity. To guarantee that the storage service works, P2P storage systems have to meet two main requirements. First, the storage system needs to maintain fairness among peers by ensuring that peers consuming more online storage capacity contribute more local storage resources. And second, to reduce redundancy costs and improve reliability, the storage system must incentivize low-available peers to improve their online availability. Traditionally, P2P storage systems achieved these two requirements by (i) using symmetric reciprocal exchanges between peers, and by (ii) allowing peers to selfishly select their set of storage partners. However, in this paper we show that these two mechanisms are suboptimal in terms of the overall storage resources contributed by all peers. To minimize this amount of contributed resources, we design a novel incentive mechanism based on asymmetric reciprocal exchanges between peers. Our mechanism incentivizes peers to select storage partners uniformly at random, and to establish asymmetric exchange relationships with them. These asymmetric exchange relationships allow low-available peers to compensate the increase of redundancy of high-available peers by giving them more storage capacity. We show that our solution reduces the overall amount of contributed resources as well as the resources contributed by each peer individually. Using real P2P availability traces, we show that our incentive mechanism can reduce the overall savings up to 60%, and individual savings from 2% up to 75%, depending on peers' availabilities. Lluis Pamies-Juarez, Pedro García López, Marc Sánchez Artigas |
Peer-to-Peer Computing | 2 |
| 2011 | Towards the design of optimal data redundancy schemes for heterogeneous cloud storage infrastructures
Lluis Pamies-Juarez, Pedro García López, Marc Sánchez Artigas, Blas Herrera |
Comput. Networks | 2 |
| 2010 | Affluenza: Towards Universal Churn GenerationabstractChurn is an inherent property of peer-to-peer (P2P) networks. Despite its relevance, yet, there is not a universal tool to bring researchers the opportunity to compare their contributions under the same general conditions. To fill this gap, we present the first open-source, simulator-independent tool for churn modeling. Enrique Fernández-Casado, Marc Sánchez Artigas, Pedro García López |
Peer-to-Peer Computing | 3 |
| 2010 | Availability and Redundancy in Harmony: Measuring Retrieval Times in P2P Storage SystemsabstractPeer-to-peer (P2P) storage systems are strongly affected by churn - temporal and permanent peer failures. Because of this churn, the main requirement of such systems is to guarantee that stored objects can always be retrieved. This requirement is specially needed in two main situations: when users want to access the stored objects or when data maintenance processes have to repair lost information. To meet this requirement, exiting P2P storage systems introduce large amounts of redundancy that maintain data availability close to 100%. Unfortunately, these large amounts of redundancy increase the storage costs, either by reducing the overall net capacity or by increasing the communication required for data maintenance. In order to minimize storage costs, P2P storage systems can reduce data redundancy. However, less redundancy means lower data availability, which leads to increase object retrieval times. Unfortunately, longer retrieval times could compromise data maintenance processes and could penalize user's retrieval times. It is crucial then for P2P storage systems to predict the effects of a redundancy reduction. In order to provide this information, we present a novel analytical framework to measure object retrieval times under different redundancy and churn circumstances. Our framework can be directly used by backup applications aiming to maintain durability at the lower cost, or by data sharing applications that seek to reduce costs by penalizing user retrieval times. We validate our framework by simulation using real P2P traces (Skype and eMule's KAD). Lluis Pamies-Juarez, Pedro García López, Marc Sánchez Artigas |
Peer-to-Peer Computing | 2 |
| 2010 | p2pWeb: An open, decentralized infrastructure of Web servers for sharing ephemeral Web content
Marc Sánchez Artigas, Jordi Pujol Ahulló, Lluis Pamies-Juarez, Pedro García López |
Comput. Networks | 4 |
| 2010 | Guest editorial for the special issue collaborative P2P systems
Pedro García López, Michael W. Sobolewski, Marc Sánchez Artigas |
Comput. Networks | 1 |
| 2010 | eSciGrid: A P2P-based e-science Grid for scalable and efficient data sharing
Marc Sánchez Artigas, Pedro García López |
Future Gener. Comput. Syst. | 2 |
| 2010 | Enabling portability in advanced information-centric services over structured peer-to-peer systems
Jordi Pujol Ahulló, Pedro García López |
J. Netw. Comput. Appl. | 2 |
| 2010 | Moving routing protocols to the user space in MANET middleware
Pedro García López, Raúl Gracia Tinedo, Josep M. Banús Alsina |
J. Netw. Comput. Appl. | 1 |
| 2010 | Echo: A peer-to-peer clustering framework for improving communication in DHTs
Marc Sánchez Artigas, Pedro García López |
J. Parallel Distributed Comput. | 2 |
| 2009 | Exploring the Feasibility of Reputation Models for Improving P2P Routing under Churn
Marc Sánchez Artigas, Pedro García López, Blas Herrera |
Euro-Par | 2 |
| 2009 | Heterogeneity-Aware Erasure Codes for Peer-to-Peer Storage SystemsabstractPeer-to-peer (P2P) storage systems rely on data redundancy to obtain high levels of data availability. Among the existing data redundancy schemes, erasure coding is a widely adopted scheme in existing P2P storage systems. By properly tunning its parameters, erasure codes can minimize the required data redundancy, which reduces both the storage and the network overheads. However, to perform this optimization, storage systems need to measured the obtained data availability. Existing P2P storage systems assume homogeneous node availabilities in order to simplify this measurement. As we will prove, this assumption entails efficiency losses when real node availabilities are highly heterogeneous. In this work, we analyze how erasure codes can be optimized in an availability-aware fashion. We propose an analytical framework able to measure data availability more precisely than existing works. As a result, we can optimize the erasure code deployment while reducing its associated overheads. Our experiments show how by considering real node availabilities it is possible to reduce data redundancy about 50% and up to 80% in some specific scenarios. Lluis Pamies-Juarez, Pedro García López, Marc Sánchez Artigas |
ICPP | 2 |
| 2009 | PlanetSim: An Extensible Simulation Tool for Peer-to-Peer Networks and ServicesabstractWe introduce PlanetSim, a discrete event-based simulation framework for peer-to-peer overlay networks and services. It is implemented in Java and provides good qualities for both researchers and developers, such as a strong system development background, as well as modularity, flexibility and clarity on its design and implementation. All this is corroborated by the important community using and supporting PlanetSim. Jordi Pujol Ahulló, Pedro García López |
Peer-to-Peer Computing | 2 |
| 2009 | On Routing in Distributed Hash Tables: Is Reputation a Shelter from Malicious Behavior and Churn?abstractRecently, it has been argued that reputation mechanisms could be used to improve routing by conditioning next-hop decisions to the past behavior of peers. However, churn may severely hinder the applicability of reputations mechanisms. In particular, short peer lifetimes imply that reputations are typically generated from a small number of transactions and are few reliable. To examine how high rates of churn affect reputation systems, we present an analytical model to study the potential damage done by malicious peers together with churn. With our model, we show that it cannot be expected in general that reputations are reliable. We then analyze the impact of this result by proposing a new routing protocol for Chord. Mainly, the protocol exploits reputation to improve the decision about which neighbor select as next-hop peer. Our experimental results show that routing algorithms can obtain important benefits from reputation - even when peer lifetimes are short and the fraction of bad users is moderate. Marc Sánchez Artigas, Pedro García López |
Peer-to-Peer Computing | 2 |
| 2009 | POPEYE: providing collaborative services for ad hoc and spontaneous communities
Juan A. Botía Blaya, Isabelle M. Demeure, Paolo Gianrossi, Pedro García López, Juan A. Martínez 0001, Eike Michael Meyer, Patrizio Pelliccione, Frédérique Tastet-Cherel |
Serv. Oriented Comput. Appl. | 4 |
| 2008 | LightPS: Lightweight Content-Based Publish/Subscribe for Peer-to-Peer SystemsabstractIn this paper we present a technique for content-based publish/subscribe (pub/sub) systems that works without (distributed) explicit multicast group management. This technique, so-called LightPS, employs a rendezvous-based approach and, thus, maps subscription and event information into the peer-to-peer node Id keyspace. On the contrary to what could be expected, LightPS suits for high-dimensional pub/sub domains, requiring very low memory capacity and time to run subscription and event notification processes. We present its good performance through a formal theoretical analysis. Jordi Pujol Ahulló, Pedro García López, Antonio F. Skarmeta |
CISIS | 2 |
| 2008 | Secure Forwarding in DHTs - Is Redundancy the Key to Robustness?
Marc Sánchez Artigas, Pedro García López, Antonio F. Skarmeta |
Euro-Par | 2 |
| 2008 | Bypass: Providing secure DHT routing through bypassing malicious peersabstractMuch research in the last years has been devoted to the development of efficient Distributed Hash Tables (DHTs). While many works have studied DHT systems, few have examined their security issues. For example, Chord and other DHT implementations rely on the cooperation of individual peers to route requests. Consequently, any malicious node can drop and misroute messages at will, censoring the access of honest peers to content. In this paper, we introduce Bypass, a novel DHT routing protocol designed to mitigate routing attacks. A key distinguishing feature of Bypass from other implementations is a feedback-based filtering protocol that allows peers to avoid adversarial nodes when routing to the correct holders of a key. Our experimental results show that in principle Bypass can achieve a lookup success rate close to theoretical bounds. Marc Sánchez Artigas, Pedro García López, Antonio F. Skarmeta |
ISCC | 2 |
| 2008 | Supporting geographical queries onto DHTsabstractLocation-based services (LBS) are currently receiving world-wide attention as a consequence of the massive usage of mobile devices, but such location services require scalable distributed infrastructures in order to resolve spatial queries efficiently. We propose a novel methodology to enable geographical query support to distributed hash tables (DHTs). The contributions of our methodology are the followings: a) our technique is DHT-generic, b) it makes an effective clusterization of nodes and information into geographical areas, c) providing data locality without sacrificing routing and data load balancing, d) it is able to answer classical spatial range queries, as well as e) a new kind of queries we call geocast, all of them in a distributed, scalable way. We demonstrate the feasibility of our approach through representative simulations. Jordi Pujol Ahulló, Pedro García López, Marc Sánchez Artigas, Antonio F. Skarmeta |
LCN | 2 |
| 2008 | On the Feasibility of Dynamic Superpeer Ratio MaintenanceabstractThe notion of "superpeer" has been shown to be very effective to increase the scalability of P2P applications. For superpeer systems to work, it is critical to preserve the optimal ratio between the number of superpeers and normal peers participating in the overlay. This requires that peers change dynamically their role (i.e., from su-perpeer to normal peer and vice versa) in the presence of node arrivals and departures, a problem that is hard to solve if no peer has global knowledge of the network. In this article, we first investigate the feasibility of superpeer ratio maintenance when each peer can decide to be a superpeer independently of each other. We then show how this problem can be treated as an optimization problem, and we propose a distributed algorithm, based on particle swarm optimization (PSO), to solve it. Our simulation results prove the viability of a PSO-based approach for this problem. Marc Sánchez Artigas, Pedro García López, Antonio F. Skarmeta |
Peer-to-Peer Computing | 2 |
| 2008 | TR-clustering: Alleviating the impact of false clustering on P2P overlay networks
Marc Sánchez Artigas, Pedro García López, Antonio F. Skarmeta, José Santa |
Comput. Networks | 2 |
| 2007 | SQS: Similarity Query Scheme for Peer-to-Peer DatabasesabstractSimilarity search is a hot research topic on peer-to-peer systems. In this paper we present SQS, a similarity query scheme for peer-to-peer databases. In this work we provide a novel linearization mechanism that enables structured queries without the burden of a global information maintenance scheme. The system offers exact match and range searches to multidimensional data. SQS employs Cyclone, a hierarchical overlay that is able to build disjoint clusters in terms of network latency and enables data search load balancing by caching per cluster scheme. Finally, we show the good properties of SQS through representative simulation results. Jordi Pujol Ahulló, Pedro García López, Marc Sánchez Artigas, Antonio F. Skarmeta |
ISCC | 2 |
| 2007 | A Comparative Study of Hierarchical DHT SystemsabstractMuch research in the last few years has been devoted to development of efficient structured peer-to-peer (P2P) overlay networks, which offer distributed hash table (DHT) functionality. Most of these systems have been devised as flat, non-hierarchical structures, in contrast to the most scalable distributed systems of the past. To cope with this, a significant number of hierarchical DHT designs have been proposed in the literature. Unfortunately, no design is "universally" better. Actually, what is lacking is an analytic framework to identify the good hierarchical design for a given workload. In this paper, we provide such a framework, and we use it to compare the two main hierarchical DHT designs: The homogenous design, in which all nodes act equal roles, against the superpeer design, in which a small subset of peers (i.e., the most powerful and stable), behave as proxies, interconnecting clusters with highly dynamic membership. Our analysis reveals that, on the contrary to what was initially expected, the costs incurred by hierarchical superpeer design are not necessarily minimized. Marc Sánchez Artigas, Pedro García López, Antonio F. Skarmeta |
LCN | 2 |
| 2005 | The Planet Project: collaborative educational content repositories on structured peer-to-peer gridsabstractIn this paper we present the Planet Project. Its main goal is focused on educational content generation and wide-area distribution. For that matter, we have designed a distributed content repository (PlanetDR) which has been built on top of a structured peer-to-peer grid middleware called Dermi. The system has been made with interoperability in mind and it thus follows the IMS Digital Repositories Interoperability standard through an implementation of the eduSource Communication Language protocol. PlanetDR has been extended to support a federation mode, which to the best of our knowledge constitutes the first attempt in providing an alliance of content repositories throughout a structured peer-to-peer grid. Moreover, we have also created several collaborative tools which will he integrated in the content's life cycle in order to promote knowledge communities around educational content hierarchies. Subsequently, we have developed PlanetDR Communities which allow researchers to easily locate themselves by their own keywords of interest, which are stored and looked up in a peer-to-peer grid decentralized infrastructure. Carles Pairot, Pedro García López, Robert Rallo, Josep Blat, Antonio F. Skarmeta |
CCGRID | 2 |
| 2005 | Cyclone: A Novel Design Schema for Hierarchical DHTsabstractRecent research efforts have improved the existing flat distributed hash tables to accommodate hierarchical structure. Nevertheless, many problems still remain to be solved regarding scalability issues, autonomous systems, connection degree, and network proximity. In this paper, we present a new hierarchical DHT called Cyclone that aims to solve the aforementioned issues with a near-optimal architecture. Cyclone provides optimal logarithmic routing hops without establishing unnecessary connection links to other nodes. Our approach follows a horizontal and uniform leaf-based approach that considerably reduces the overall number of links per node. Furthermore, Cyclone also offers a disjoint multipath routing scheme that benefits from network proximity and thus creates a more robust overlay infrastructure. Marc Sánchez Artigas, Pedro García López, Jordi Pujol Ahulló, Antonio F. Skarmeta |
Peer-to-Peer Computing | 2 |
| 2005 | Towards new load-balancing schemes for structured peer-to-peer grids
Carles Pairot, Pedro García López, Antonio F. Skarmeta, Rubén Mondéjar |
Future Gener. Comput. Syst. | 2 |
| 2004 | DERMI: A Decentralized Peer-to-Peer Event-Based Object MiddlewareabstractWe present DERMI, a decentralized wide-area event-based object middleware built on top of a peer-to-peer substrate. Its main building block is the underlying publish/subscribe event notification system provided by the peer-to-peer layer. By using this methodology, innovative benefits like distributed interception, high performance synchronous/asynchronous one-to-one/one-to-many notifications and decentralized object location services are provided. Moreover, new programming abstractions (anycall and manycall) are introduced, which allow the programmer to make calls to groups of objects without taking care of which of them responds until a determinate condition is met. We believe that such middleware is a solid building block for future wide-area distributed component infrastructures. Carles Pairot, Pedro García López, Antonio F. Skarmeta |
ICDCS | 2 |