Raúl Gracia Tinedo

dblp:80/8752 · DBLP profile ↗
← Back
24ranked-venue papers
18as first author
4since 2021 · last 2025
0000-0003-1842-2976ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 10 first-authorComputer networks · 6 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 first-authorDatabases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Cloud and datacenter computing · 36% Storage systems · 30% Performance modeling and evaluation · 23%
Computer networks
3 papers
Edge and fog computing · 96% Network measurement and analytics · 4%
Databases, data mining, and information retrieval
3 papers
Data stream processing · 58% Distributed and cloud data management · 22% Information retrieval · 20%
Software engineering, system software, and programming languages
1 paper
Services computing and microservices · 100%

Topics — the 23 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
data distribution
0.912025
Building Stateless Serverless Vector DBs via Block-based Data Partitioning · Proc. ACM Manag. Data 2025
Cloud and datacenter computing › serverless computing
function-as-a-service
0.912025
Building Stateless Serverless Vector DBs via Block-based Data Partitioning · Proc. ACM Manag. Data 2025
Cloud and datacenter computing
serverless computing
0.912025
Building Stateless Serverless Vector DBs via Block-based Data Partitioning · Proc. ACM Manag. Data 2025
Storage systems › data management › storage for machine learning
vector database
0.912025
Building Stateless Serverless Vector DBs via Block-based Data Partitioning · Proc. ACM Manag. Data 2025
Edge and fog computing
cloud-edge continuum
0.812024
Towards Multi-Tier Stream Data Tiering in the Cloud-Edge Continuum · ICNP 2024
Edge and fog computing
resource management
0.712023
The Nanoservices Framework: Co-Locating Microservices in the Cloud-Edge Continuum · ICNP 2023
Services computing and microservices › microservice architecture
microservice deployment
0.712023
The Nanoservices Framework: Co-Locating Microservices in the Cloud-Edge Continuum · ICNP 2023
Storage systems › distributed storage
personal cloud storage
0.312018
BenchBox: A User-Driven Benchmarking Framework for Fat-Client Storage Systems · IEEE Trans. Parallel Distributed Syst. 2018
Performance modeling and evaluation
storage performance evaluation
0.312018
BenchBox: A User-Driven Benchmarking Framework for Fat-Client Storage Systems · IEEE Trans. Parallel Distributed Syst. 2018
Performance modeling and evaluation › workload characterization
workload generation
0.312018
BenchBox: A User-Driven Benchmarking Framework for Fat-Client Storage Systems · IEEE Trans. Parallel Distributed Syst. 2018
Distributed and cloud data management
query offloading
0.312017
Too Big to Eat: Boosting Analytics Data Ingestion from Object Stores with Scoop · ICDE 2017
Storage systems
object storage
0.312017
Crystal: Software-Defined Storage for Multi-Tenant Object Stores · FAST 2017
Storage systems › storage architecture
software-defined storage
0.312017
Crystal: Software-Defined Storage for Multi-Tenant Object Stores · FAST 2017
Information retrieval
retrieval-augmented generation
0.312025
Building Stateless Serverless Vector DBs via Block-based Data Partitioning · Proc. ACM Manag. Data 2025
Cloud and datacenter computing › serverless computing
ephemeral storage
0.212024
Towards Multi-Tier Stream Data Tiering in the Cloud-Edge Continuum · ICNP 2024
Storage systems › storage hierarchy
tiered storage
0.212024
Towards Multi-Tier Stream Data Tiering in the Cloud-Edge Continuum · ICNP 2024
Performance modeling and evaluation
benchmarking
0.212015
SDGen: Mimicking Datasets for Content Generation in Storage Benchmarks · FAST 2015
Performance modeling and evaluation › benchmarking
storage benchmarking
0.212015
SDGen: Mimicking Datasets for Content Generation in Storage Benchmarks · FAST 2015
Performance modeling and evaluation
synthetic data generation
0.212015
SDGen: Mimicking Datasets for Content Generation in Storage Benchmarks · FAST 2015
Performance modeling and evaluation
workload characterization
0.212015
SDGen: Mimicking Datasets for Content Generation in Storage Benchmarks · FAST 2015
Cloud and datacenter computing › virtualization
containerization
0.212023
The Nanoservices Framework: Co-Locating Microservices in the Cloud-Edge Continuum · ICNP 2023
Performance modeling and evaluation › benchmarking
benchmarking framework
0.112018
BenchBox: A User-Driven Benchmarking Framework for Fat-Client Storage Systems · IEEE Trans. Parallel Distributed Syst. 2018
Performance modeling and evaluation › performance monitoring
distributed system performance measurement
0.112018
BenchBox: A User-Driven Benchmarking Framework for Fat-Client Storage Systems · IEEE Trans. Parallel Distributed Syst. 2018

Methods — techniques the papers use, named apart from their topics

system integration · 2.3performance evaluation · 2.3container co-location · 2.0experimental benchmarking · 1.7active object storage layer · 0.6ETL-type actions · 0.6reverse engineering · 0.4measurement study · 0.4workload modeling · 0.3user stereotype modeling · 0.3
YearPublicationVenuePosition
2025 Building Stateless Serverless Vector DBs via Block-based Data Partitioning
abstract
Retrieval-Augmented Generation (RAG) and other AI/ML workloads rely on vector databases (DBs) for efficient analysis of unstructured data. However, cluster (or serverful ) vector DB architectures, such as Milvus, lack the elasticity to handle high workload fluctuations, sparsity, and burstiness. Serverless vector DBs-- i.e., vector DBs built on top of cloud functions--have emerged as a promising alternative architecture, but they are still in their infancy. This paper presents the first experimental study comparing data partitioning strategies in vector DBs built atop stateless Function-as-a-Service (FaaS). Through extensive benchmarks, we reveal key limitations of clustering-based data partitioning when applied to dynamic datasets ( e.g. , complexity, load balancing). We then evaluate a block-based alternative that addresses such limitations ( e.g. , up to 5.8× faster data partitioning, up to 63% lower costs, similar querying times). Moreover, our results show that a stateless serverless vector DB using block-based data partitioning achieves competitive performance with Milvus in several aspects ( e.g. , up to 65.6× faster data partitioning, similar recall), while reducing costs for sparse workloads (up to 99%). Our empirical insights aim to guide the design of next-generation serverless vector DBs.
Daniel Barcelona Pons, Raúl Gracia Tinedo, Albert Cañadilla-Domingo, Xavier Roca-Canals, Pedro García López
Proc. ACM Manag. Data2
2024 Towards Multi-Tier Stream Data Tiering in the Cloud-Edge Continuum
abstract
Event streaming systems (e.g., Apache Kafka, Apache Pulsar) are a popular substrate for ingesting data with low latency from continuous data sources, such as sensors, cameras, or server logs. Due to the sheer amount of data being stored as data streams, several systems are incorporating storage tiering as a core feature. However, in some cases, the design of the streaming system assumes a reliable connection with the external storage to offload data. This may not be the case when deploying streaming pipelines in the Cloud-Edge Continuum. In this paper, we evaluate deploying a streaming storage system with integrated data tiering (Pravega) in the CloudEdge Continuum. We identify that while Pravega provides good 10 performance, extended unavailability of the long-term storage service may impact stream data ingestion. This can be problematic in Edge use cases with stringent streaming ingestion and processing requirements. To mitigate this problem, we explore the concept of multi-tier long-term storage in Pravega. We implement this concept by integrating an ephemeral tiered storage system (GEDS) to augment Pravega with advanced data tiering mechanisms. Our preliminary results show that GEDS can exploit multiple storage tiers that increase by$3.8 x$the tolerance of the streaming system to long-term storage unavailability.
Omar Jundi, Raúl Gracia Tinedo, Sean Ahearne, Pascal Spörri, Bernard Metzler
ICNP2
2023 The Nanoservices Framework: Co-Locating Microservices in the Cloud-Edge Continuum
abstract
Today, the microservices paradigm has emerged as one of the most widely adopted patterns to develop, package, and deploy software on a large scale. However, as they were originally designed for the cloud, the direct application of microservices DevOps practices to resource-constrained environments like the Edge may not be optimal. Specifically, deploying each piece of software as an individual microservice may result in a significant resource footprint (e.g., storage space and network bandwidth related to managing base images, CPU, and memory). In this work, we explore the concept of opportunistically grouping microservice code within the same container to reduce resource footprint when deploying multiple microservices at the Edge. To materialize this concept, we present the N anoservices framework: a framework that formalizes and provides practical means for developers to build and deploy groups of microservices on the same container (a.k.a., Nanoservices). Our early results show that with Nanoservices we can achieve a significant resource footprint reduction (base image storage, CPU, memory) with minimal effort from the developer's viewpoint.
Eric Caron, Raúl Gracia Tinedo
ICNP2
2023 Pravega: A Tiered Storage System for Data Streams
abstract
The growing popularity of the data stream abstraction entails new challenging requirements when it comes to data ingestion and storage. Many organizations expect to retain data streams for extended periods of time and to store such stream data in a cost-effective manner. It is also crucial to reconcile apparently opposite properties, like data durability and consistency, along with high performance. Furthermore, data streams should not only deal with a high degree of parallelism, but also adapt to fluctuating workloads with little or no admin intervention. To our knowledge, no storage system for data streams fully copes with all these requirements.
Raúl Gracia Tinedo, Flavio Paiva Junqueira, Tom Kaitchuck, Sachin Joshi
Middleware1
2019 Lamda-Flow: Automatic Pushdown of Dataflow Operators Close to the Data
abstract
Modern data analytics infrastructures are composed of physically disaggregated compute and storage clusters. Thus, dataflow analytics engines, such as Apache Spark or Flink, are left with no choice but to transfer datasets to the compute cluster prior to their actual processing. For large data volumes, this becomes problematic, since it involves massive data transfers that exhaust network bandwidth, that waste compute cluster memory, and that may become a performance barrier. To overcome this problem, we present λFlow: a framework for automatically pushing dataflow operators (e.g., map, flatMap, filter, etc.) down onto the storage layer. The novelty of λFlow is that it manages the pushdown granularity at the operator level, which makes it a unique problem. To wit, it requires addressing several challenges, such as how to encapsulate dataflow operators and execute them on the storage cluster, and how to keep track of dependencies such that operators can be pushed down safely onto the storage layer. Our evaluation reports significant reductions in resource usage for a large variety of IO-bound jobs. For instance, λFlow was able to reduce both network bandwidth and memory requirements by 90% in Spark. Our Flink experiments also prove the extensibility of λFlow to other engines.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López, Yosef Moatti, Filip Gluszak
CCGRID1
2019 Software-defined object storage in multi-tenant environments
Raúl Gracia Tinedo, Josep Sampé, Gerard París, Marc Sánchez Artigas, Pedro García López, Yosef Moatti
Future Gener. Comput. Syst.1
2018 Giving wings to your data: A first experience of Personal Cloud interoperability
Raúl Gracia Tinedo, Cristian Cotes, Edgar Zamora-Gómez, Genís Ortiz, Adrián Moreno-Martínez, Marc Sánchez Artigas, Pedro García López, Raquel Sánchez, Alberto Gómez 0003, Anastasio Illana
Future Gener. Comput. Syst.1
2018 BenchBox: A User-Driven Benchmarking Framework for Fat-Client Storage Systems
abstract
In many online storage services, end-users mainly interact with the system via “fat” storage clients that integrate complex functionality. This means that to obtain a complete performance evaluation of one of such systems we may need to generate workloads on the client side that reproduce the behavior of real users. Unfortunately, this remains as an open research challenge today. We present BenchBox: A distributed performance evaluation framework for fat-client storage systems. On the one hand, BenchBox can generate workloads directly in storage clients that mimic users exhibiting a certain behavior, namely, user stereotypes. To this end, the framework enables to plug-in workload models and feed them with compact recipes that capture the behavior of user stereotypes (e.g., storage activity, type of file contents, data sharing links). On the other hand, BenchBox provides researchers with management and monitoringfacilities to deploy experiments and analyze the performance of groups of storage clients. To demonstrate our framework, we equipped BenchBox with a 2-layer workload modelthat reproduces both the activity-e.g., types of operations, frequency- and data-e.g., file sizes, data types-of users in a Personal Cloud. We used this model to generate workloads based on user stereotypes that we identified in real traces (UbuntuOne). Our experiments with public providers show how distinct types of users impact on the performance and efficiency of Personal Clouds, which may guide their optimization.
Raúl Gracia Tinedo, Chenglong Zou, Marc Sánchez Artigas, Pedro García López
IEEE Trans. Parallel Distributed Syst.1
2017 Crystal: Software-Defined Storage for Multi-Tenant Object Stores
Raúl Gracia Tinedo, Josep Sampé, Edgar Zamora-Gómez, Marc Sánchez Artigas, Pedro García López, Yosef Moatti, Eran Rom
FAST1
2017 Too Big to Eat: Boosting Analytics Data Ingestion from Object Stores with Scoop
abstract
Extracting value from data stored in object stores,such as OpenStack Swift and Amazon S3, can be problematicin common scenarios where analytics frameworks and objectstores run in physically disaggregated clusters. One of the mainproblems is that analytics frameworks must ingest large amountsof data from the object store prior to the actual computation;this incurs a significant resources and performance overhead. Toovercome this problem, we present Scoop. Scoop enables analyticsframeworks to benefit from the computational resources of objectstores to optimize the execution of analytics jobs. Scoop achievesthis by enabling the addition of ETL-type actions to the dataupload path and by offloading querying functions to the objectstore through a rich and extensible active object storage layer. Asa proof-of-concept, Scoop enables Apache Spark SQL selectionsand projections to be executed close to the data in OpenStackSwift for accelerating analytics workloads of a smart energy gridcompany (GridPocket). Our experiments in a 63-machine clusterwith real IoT data and SQL queries from GridPocket show thatScoop exhibits query execution times up to 30x faster than thetraditional “ingest-then-compute” approach.
Yosef Moatti, Eran Rom, Raúl Gracia Tinedo, Dalit Naor, Doron Chen, Josep Sampé, Marc Sánchez Artigas, Pedro García López, Filip Gluszak, Eric Deschdt, Francesco Pace, Daniele Venzano, Pietro Michiardi
ICDE3
2016 Understanding Data Sharing in Private Personal Clouds
abstract
Data sharing in Personal Clouds blurs the lines between on-line storage and content distribution with a strong social component. Such social information may be exploited by researchers to devise optimized data management techniques for Personal Clouds. Unfortunately, due their proprietary nature, data sharing is one of the least studied facets of these systems. In this work, we present the first study of data sharing in a private Personal Cloud. Concretely, we contribute a dataset collected at the metadata back-end of NEC: an enterprise oriented Personal Cloud. First, our analysis provides a deep inspection of the storage layer of NEC, comparing it with a well-known public vendor (UbuntuOne). Second, we study the social structure of NEC user communities, as well as the storage characteristics of user sharing links via multiplex network techniques. Finally, we discuss a battery of data management optimizations for NEC derived from our findings, which may be of independent interest for other similar systems. Our proposals include content distribution, caching and data placement. We believe that both our study and dataset will foster further research in this field.
Raúl Gracia Tinedo, Pedro García López, Alberto Gómez 0003, Anastasio Illana
CLOUD1
2015 SDGen: Mimicking Datasets for Content Generation in Storage Benchmarks
Raúl Gracia Tinedo, Danny Harnik, Dalit Naor, Dmitry Sotnikov, Sivan Toledo, Aviad Zuck
FAST1
2015 Dissecting UbuntuOne: Autopsy of a Global-scale Personal Cloud Back-end
abstract
Personal Cloud services, such as Dropbox or Box, have been widely adopted by users. Unfortunately, very little is known about the internal operation and general characteristics of Personal Clouds since they are proprietary services.
Raúl Gracia Tinedo, Yongchao Tian, Josep Sampé, Hamza Harkous, John Lenton, Pedro García López, Marc Sánchez Artigas, Marko Vukolic
Internet Measurement Conference1
2014 eWave: Leveraging Energy-Awareness for In-line Deduplication Clusters
abstract
In-line deduplication clusters provide high throughput and scalable storage/archival services to enterprises and organizations. Unfortunately, high throughput comes at the cost of activating several storage nodes on each request, due to the parallel nature of superchunk routing. This may prevent storage nodes from exploiting disk standby times to preserve energy, even for low load periods. We aim to enable deduplication clusters to exploit load valleys to save up disk energy. To this end, we explore the feasibility of deferred writes, diverted access and workload consolidation in this setting.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López
SYSTOR1
2014 Giving form to social cloud storage through experimentation: Issues and insights
Raúl Gracia Tinedo, Marc Sánchez Artigas, Aleix Ramírez, Adrián Moreno-Martínez, Xavier León, Pedro García López
Future Gener. Comput. Syst.1
2013 Actively Measuring Personal Cloud Storage
abstract
The Personal Cloud model is a mainstream service that meets the growing demand of millions of users for reliable off-site storage. However, despite their broad adoption, very little is known about the quality of service (QoS) of Personal Clouds. In this paper, we present a measurement study of three major Personal Clouds: DropBox, Box and SugarSync. Actively accessing to free accounts through their REST APIs, we analyzed important aspects to characterize their QoS, such as transfer speed, variability and failure rate. Our measurement, conducted during two months, is the first to deeply analyze many facets of these popular services and reveals new insights, such as important performance differences among providers, the existence of transfer speed daily patterns or sudden service breakdowns. We believe that the present analysis of Personal Clouds is of interest to researchers and developers with diverse concerns about Cloud storage, since our observations can help them to understand and characterize the nature of these services.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Adrián Moreno-Martínez, Cristian Cotes, Pedro García López
IEEE CLOUD1
2013 Cloud-as-a-Gift: Effectively Exploiting Personal Cloud Free Accounts via REST APIs
abstract
Personal Clouds, such as DropBox and Box, provide open REST APIs for developers to create clever applications that make their service even more attractive. These APIs are a powerful abstraction that makes it possible for applications to transparently manage data from user accounts, blurring the lines between a Personal Cloud service and storage IaaS. Jointly, Personal Clouds also offer free accounts to lure new users, that normally include reduced storage space and unlimited transfers. However, the unintended consequence of combining open APIs and free accounts is that these companies are exposing automated access to a free storage infrastructure, which may lead to abuse by malicious parties. By exploiting the freemium API service, users may fraudulently consume resources or they can use free accounts as a Cloud storage layer to support abusive applications. We call this vulnerability the storage leeching problem. In this paper, we show how easy it is to implement a file-sharing application able to distribute digital content by abusing Personal Clouds. Making use of open APIs, this application transparently aggregates the limited-space free accounts from multiple providers into a single larger storage layer, while achieving better transfer speed than that received from one provider alone. This demonstrates that free accounts can be easily exploited to obtain a practical Cloud storage service, and therefore, the potential impact of storage leeching.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López
IEEE CLOUD1
2012 F2Box: Cloudifying F2F Storage Systems with High Availability Correlation
abstract
The increasing popularity of Cloud storage services is leading end-users to store their digital lives (including photos, videos, work documents, etc.) in the Cloud. However, many users are still reluctant to move their data to the Cloud due to the amount of control ceded to Cloud vendors. To let users retain the control over their data, Friend-to-Friend (F2F) storage systems have been presented in the literature as a promising alternative. However, as we show in this paper, pure F2F storage systems present a poor QoS, mainly due to availability correlations, which results in a loss of attractiveness by end users. To overcome this limitation, we propose a hybrid architecture that combines F2F storage systems and the availability of Cloud storage services to let users infer the right balance between user control and quality of service. This architecture, we called it F2BOX, is able to deliver such a balance thanks to the development of a new suite of data transfer scheduling strategies and a new redundancy calculation algorithm. The main feature of this algorithm is that allow users to adjust the amount of redundancy according to the availability patterns exhibited by friends. Our simulation and experimental results (in Amazon S3) demonstrate the high benefits experienced by end users as a result of the "cloudification" of F2F systems.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López
IEEE CLOUD1
2012 FriendBox: A Hybrid F2F Personal Storage Application
abstract
Personal storage is a mainstream service used by millions of users. Among the existing alternatives, Friend-to-Friend (F2F) systems are nowadays an interesting research topic aimed to leverage a secure and private off-site storage service. However, the specific characteristics of F2F storage systems (reduced node degree, correlated availabilities) represent a hard obstacle to their performance. Actually, it is extremely difficult for a F2F system to guarantee an acceptable storage service quality in terms of transference times and data availability to end-users. In this landscape, we propose to resort to the Cloud for improving the storage service of a F2F system. We present FriendBox: a hybrid F2F personal storage system. FriendBox is the first F2F system that efficiently combines resources of trusted friends with Cloud storage for improving the service quality achievable by pure F2F systems. We evaluated FriendBox through a real deployment in our university campus. We demonstrated that FriendBox achieves high transfer performance and flexible user-defined data availability guarantees. Furthermore, we analyzed the costs of FriendBox demonstrating its economic feasibility.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Adrián Moreno-Martínez, Pedro García López
IEEE CLOUD1
2012 FRIENDBOX: A cloudified F2F storage application
abstract
Personal storage is a mainstream service used by millions of users. Among the existing alternatives, Friend-to-Friend (F2F) systems are aimed to leverage a secure and private off-site storage service. However, the specific characteristics of these systems (reduced node degree, correlated availabilities) represent a hard obstacle to their performance. We present FriendBox: a hybrid F2F personal storage system that combines resources of trusted friends with Cloud storage for improving the service quality achievable by pure F2F systems.
Adrián Moreno-Martínez, Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López
P2P2
2012 Analysis of data availability in F2F storage systems: When correlations matter
abstract
Nowadays, the growing necessity for secure and private off-site storage motivates the appearance of novel storage infrastructures. In this sense, it is increasingly common to find storage systems where users interact just with a set of trustworthy participants, such as in Friend-to-Friend (F2F) networks.
Raúl Gracia Tinedo, Marc Sánchez Artigas, Pedro García López
P2P1
2012 Sophia: A local trust system to secure key-based routing in non-deterministic DHTs
Raúl Gracia Tinedo, Pedro García López, Marc Sánchez Artigas
J. Parallel Distributed Comput.1
2011 Sophia: Local Trust for Securing Routing in DHTs
abstract
Distributed Hash Tables (DHTs) have been used as a common building block in many distributed applications, including Cloud and Grid. However, there are still important security vulnerabilities that hinder their adoption in today'slarge-scale computing platforms. For instance, routing vulnerabilities have been a subject of intensive research but existing solutions rely on redundancy in lieu of improving the quality of routing paths. In this paper, we present Sophia, a novel generic security technique which combines iterative routing with local trust to fortify routing in DHTs. Sophia strictly benefits from first-hand observations about the success/failure of a node's own lookups to improve forwarding paths. Moreover, unlike redundant routing, Sophia dynamically protects routing without introducing additional network overhead. To the best of our knowledge, this is the first work which exploits a local trust system to fortify routing in DHTs. We compared the performance of Sophia with redundant routing in Kademlia DHT. We obtained significant improvements regarding routing resilience, self-adjustment and network traffic reduction.
Raúl Gracia Tinedo, Pedro García López, Marc Sánchez Artigas
CCGRID1
2010 Moving routing protocols to the user space in MANET middleware
Pedro García López, Raúl Gracia Tinedo, Josep M. Banús Alsina
J. Netw. Comput. Appl.2