Ricardo Jiménez-Peris

dblp:40/4806 · DBLP profile ↗
← Back
63ranked-venue papers
17as first author
3since 2021 · last 2023
0000-0001-9174-6569ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 19 · 3 first-author · 3 since 2021Systems, architecture and hardware · 12 · 2 first-authorSecurity and privacy · 9 · 3 first-authorHuman-computer interaction and ubiquitous computing · 8 · 5 first-authorSoftware engineering, systems software and programming languages · 7 · 1 first-authorArtificial intelligence and machine learning · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021
YearPublicationVenuePosition
2023 MobiSpaces: An Architecture for Energy-Efficient Data Spaces for Mobility Data
abstract
In this paper, we present an architecture for mobility data spaces enabling trustworthy and reliable data operations along with its main constituent parts. The architecture makes use of a data lake for scalable storage of diverse mobility data sets, on top of which separate computing and storage layers are implemented to allow independent scaling with a data operations toolbox providing all data operations. Furthermore, to cater for mobility analytics, machine learning and artificial intelligence support, an edge analytics suite is provided that encompasses distributed algorithms for mobility analytics and federated learning, thereby exploiting edge computing technologies. In turn, this is supported by a resource allocator that monitors the energy consumption of data-intensive operations and provides this information to the platform for intelligent task placement in edge devices, aiming at energy-efficient operations. As a result, an end-to-end platform is proposed that combines data services and infrastructure services towards supporting mobility application domains, such as urban and maritime.
Christos Doulkeridis, Georgios M. Santipantakis, Nikolaos Koutroumanis, George Makridis, Vasilis Koukos, George S. Theodoropoulos, Yannis Theodoridis, Dimosthenis Kyriazis, Pavlos Kranas, Diego Burgos, Ricardo Jiménez-Peris, Mariana M. G. Duarte, Mahmoud Attia Sakr, Esteban Zimányi, Anita Graser, Clemens Heistracher, Kristian Torp, Ioannis Chrysakis, Theofanis Orphanoudakis, Evgenia Kapassa, Marios Touloupou, Jürgen Neises, Petros Petrou, Sophia Karagiorgou, Rosario Catelli, Domenico Messina, Marcelo Corrales Compagnucci, Matteo Falsetta
IEEE Big Data11
2022 Elastic scalable transaction processing in LeanXcale
abstract
International audience
Ricardo Jiménez-Peris, Diego Burgos-Sancho, Francisco J. Ballesteros, Marta Patiño-Martínez, Patrick Valduriez
Inf. Syst.1
2021 Parallel query processing in a polystore
Pavlos Kranas, Boyan Kolev, Oleksandra Levchenko, Esther Pacitti, Patrick Valduriez, Ricardo Jiménez-Peris, Marta Patiño-Martínez
Distributed Parallel Databases6
2019 NUMA-aware Deployments for LeanXcale Database Appliance
Ricardo Jiménez-Peris, Francisco J. Ballesteros, Pavlos Kranas, Diego Burgos, Patricio Martínez
CLOSER1
2019 Parallel Streaming Implementation of Online Time Series Correlation Discovery on Sliding Windows with Regression Capabilities
abstract
International audience
Boyan Kolev, Reza Akbarinia, Ricardo Jiménez-Peris, Oleksandra Levchenko, Florent Masseglia, Marta Patiño-Martínez, Patrick Valduriez
CLOSER3
2019 Parallel Efficient Data Loading
abstract
In this paper we discuss how we architected and developed a parallel data loader for LeanXcale database. The loader is characterized for its efficiency and parallelism. LeanXcale can scale up and scale out to very large numbers and loading data in the traditional way it is not exploiting its full potential in terms of the loading rate it can reach. For this reason, we have created a parallel loader that can reach the maximum insertion rate LeanXcale can handle. LeanXcale also exhibits a dual interface, key-value and SQL, that has been exploited by the parallel loader. Basically, the loading leverages the key-value API and results in a highly efficient process that avoids the overhead of SQL processing. Finally, in order to guarantee the parallelism we have developed a data sampler that samples data to generate a histogram of data distribution and use it to pre-split the regions across LeanXcale instances to guarantee that all instances get an even amount of data during loading, thus g uaranteeing the peak processing loading capability of the deployment.
Ricardo Jiménez-Peris, Francisco J. Ballesteros, Ainhoa Azqueta-Alzúaz, Pavlos Kranas, Diego Burgos, Patricio Martínez
DATA1
2019 Pipelined Implementation of a Parallel Streaming Method for Time Series Correlation Discovery on Sliding Windows
Boyan Kolev, Reza Akbarinia, Ricardo Jiménez-Peris, Oleksandra Levchenko, Florent Masseglia, Marta Patiño-Martínez, Patrick Valduriez
DATA3
2019 Modular FPGA Acceleration of Data Analytics in Heterogenous Computing
abstract
Emerging cloud applications like machine learning, AI and big data analytics require high performance computing systems that can sustain the increased amount of data processing without consuming excessive power. Towards this end, many cloud operators have started deploying hardware accelerators, like FPGAs, to increase the performance of computationally intensive tasks but increasing the programming complexity to utilize these accelerators. VINEYARD has developed an efficient framework that allows the seamless deployment and utilization of hardware accelerators in the cloud without increasing the programming complexity and offering the flexibility of software packages. This paper presents a modular approach for the acceleration of data analytics using FPGAs. The modular approach allows the automatic development of integrated hardware designs for the acceleration of data analytics. The proposed framework shows the data analytics modules can be used to achieve up to 3.5x speedup compared to high performance general-purpose processors.
Elias Koromilas, Christoforos Kachris, Dimitrios Soudris, Francisco J. Ballesteros, Patricio Martínez, Ricardo Jiménez-Peris
DATE6
2018 Parallel Polyglot Query Processing on Heterogeneous Cloud Data Stores with LeanXcale
abstract
The blooming of different cloud data stores has turned polystore systems to a major topic in the nowadays cloud landscape. Especially, as the amount of processed data grows rapidly each year, much attention is being paid on taking advantage of the parallel processing capabilities of the underlying data stores. To provide data federation, a typical polystore solution defines a common data model and query language with translations to API calls or queries to each data store. However, this may lead to losing important querying capabilities. The polyglot approach of the CloudMdsQL query language allows data store native queries to be expressed as inline scripts and combined with regular SQL statements in ad-hoc integration queries. Moreover, efficient optimization techniques, such as bind join, can still take place to improve the performance of selective joins. In this paper, we introduce the distributed architecture of the LeanXcale query engine that processes polyglot queries in the CloudMdsQL query language, yet allowing native scripts to be handled in parallel at data store shards, so that efficient and scalable parallel joins take place at the query engine level. The experimental evaluation of the LeanXcale parallel query engine on various join queries illustrates well the performance benefits of exploiting the parallelism of the underlying data management technologies in combination with the high expressivity provided by their scripting/querying frameworks.
Boyan Kolev, Oleksandra Levchenko, Esther Pacitti, Patrick Valduriez, Ricardo Vilaça, Rui C. Gonçalves, Ricardo Jiménez-Peris, Pavlos Kranas
IEEE BigData7
2017 Massive Data Load on Distributed Database Systems over HBase
abstract
Big Data has become a pervasive technology to manage the ever-increasing volumes of data. Among Big Data solutions, scalable data stores play an important role, especially, key-value data stores due to their large scalability (thousands of nodes). The typical workflow for Big Data applications include two phases. The first one is to load the data into the data store typically as part of an ETL (Extract-Transform-Load) process. The second one is the processing of the data itself. BigTable and HBase are the preferred key-value solutions based on range-partitioned data stores. However, the loading phase is inefficient and creates a single node bottleneck. In this paper, we identify and quantify this bottleneck and propose a tool for parallel massive data loading that solves satisfactorily the bottleneck enabling all the parallelism and throughput of the underlying key-value data store during the loading phase as well. The proposed solution has been implemented as a tool for parallel massive data loading over HBase, the key-value data store of the Hadoop ecosystem.
Ainhoa Azqueta-Alzúaz, Marta Patiño-Martínez, Ivan Brondino, Ricardo Jiménez-Peris
CCGrid4
2017 Load balancing for Key Value Data Stores
Ainhoa Azqueta-Alzúaz, Ivan Brondino, Marta Patiño-Martínez, Ricardo Jiménez-Peris
EDBT4
2016 Benchmarking polystores: The CloudMdsQL experience
abstract
The CloudMdsQL polystore provides integrated access to multiple heterogeneous data stores, such as RDBMS, NoSQL or even HDFS through a big data analytics framework such as MapReduce or Spark. The CloudMdsQL language is a functional SQL-like query language with a flexible nested data model. A major capability is to exploit the full power of each of the underlying data stores by allowing native queries to be expressed as functions and involved in SQL statements. The CloudMdsQL polystore has been validated with a good number of different data stores: HDFS, key-value, document, graph, RDBMS and OLAP engine. In this paper, we introduce the benchmarking of the CloudMdsQL polystore and evaluate the performance benefits of important features enabled by the query language and engine.
Boyan Kolev, Raquel Pau, Oleksandra Levchenko, Patrick Valduriez, Ricardo Jiménez-Peris, José Pereira 0001
IEEE BigData5
2016 Design of an RDMA Communication Middleware for Asynchronous Shuffling in Analytical Processing
abstract
A key component in a distributed parallel analytical processing engine is shuffling, the distribution of data to multiple nodes such that the computation can be done in parallel. In this paper we describe the initial design of a communication middleware to support asynchronous shuffling of data among multiple processes on a distributed memory environment. The proposed middleware relies on RDMA (Remote Direct Memory Access) operations to transfer data, and provides basic operations to send and queue data on remote machines, and to retrieve this queued data. Preliminary results show that the RDMA-based middleware can provide a 75% reduction on communication costs, when compared with a traditional sockets implementation.
Rui C. Gonçalves, José Pereira 0001, Ricardo Jiménez-Peris
CLOSER (1)3
2016 PaaS-CEP - A Query Language for Complex Event Processing and Databases
abstract
Nowadays many applications must process events at a very high rate. These events are processed on the fly, without being stored. Complex Event Processing technology (CEP) is used to implement such applications. Some of the CEP systems, like Apache Storm the most popular CEPs, lack a query language and operators to program queries as done in traditional relational databases. This paper presents PaaS-CEP, a CEP language that provides a SQL-like language to program queries for CEP and its integration with data stores (database or key-value store). Our current implementation is done on top of Apache Storm however, the CEP language can be used with any CEP. The paper describes the architecture of the PaaS-CEP, its query language and the algebraic operators. The paper also details the integration of the CEP with traditional data stores that allows the correlation of live streaming data with the stored data.
Ricardo Jiménez-Peris, Valerio Vianello, Marta Patiño-Martínez
CLOSER (1)1
2016 Design and Implementation of the CloudMdsQL Multistore System
abstract
International audience
Boyan Kolev, Carlyna Bondiombouy, Oleksandra Levchenko, Patrick Valduriez, Ricardo Jiménez-Peris, Raquel Pau, José Pereira 0001
CLOSER (1)5
2016 An RDMA Middleware for Asynchronous Multi-stage Shuffling in Analytical Processing
abstract
A key component in large scale distributed analytical processing is shuffling , the distribution of data to multiple nodes such that the computation can be done in parallel. In this paper we describe the design and implementation of a communication middleware to support data shuffling for executing multi-stage analytical processing operations in parallel. The middleware relies on RDMA (Remote Direct Memory Access) to provide basic operations to asynchronously exchange data among multiple machines. Experimental results show that the RDMA-based middleware developed can provide a 75 % reduction of the costs of communication operations on parallel analytical processing tasks, when compared with a sockets middleware.
Rui C. Gonçalves, José Pereira 0001, Ricardo Jiménez-Peris
DAIS3
2016 Snapshot Isolation for Neo4j
Marta Patiño-Martínez, Ricardo Jiménez-Peris, Diego Burgos-Sancho, Ivan Brondino, Valerio Vianello, Rohit Dhamane
EDBT2
2016 The CloudMdsQL Multistore System
abstract
The blooming of different cloud data management infrastructures has turned multistore systems to a major topic in the nowadays cloud landscape. In this demonstration, we present a Cloud Multidatastore Query Language (CloudMdsQL), and its query engine. CloudMdsQL is a functional SQL-like language, capable of querying multiple heterogeneous data stores (relational and NoSQL) within a single query that may contain embedded invocations to each data store's native query interface. The major innovation is that a CloudMdsQL query can exploit the full power of local data stores, by simply allowing some local data store native queries (e.g. a breadth-first search query against a graph database) to be called as functions, and at the same time be optimized. Within our demonstration, we focus on two use cases each involving four diverse data stores (graph, document, relational, and key-value) with its corresponding CloudMdsQL queries. The query execution flows are visualized by an embedded real-time monitoring subsystem. The users can also try out different ad-hoc queries, not necessarily in the context of the use cases.
Boyan Kolev, Carlyna Bondiombouy, Patrick Valduriez, Ricardo Jiménez-Peris, Raquel Pau, José Pereira 0001
SIGMOD Conference4
2016 CloudMdsQL: querying heterogeneous cloud data stores with a common language
Boyan Kolev, Patrick Valduriez, Carlyna Bondiombouy, Ricardo Jiménez-Peris, Raquel Pau, José Pereira 0001
Distributed Parallel Databases4
2015 STONE: A streaming DDoS defense framework
Vincenzo Gulisano, Mar Callau-Zori, Zhang Fu, Ricardo Jiménez-Peris, Marina Papatriantafilou, Marta Patiño-Martínez
Expert Syst. Appl.4
2014 Performance evaluation of database replication systems
abstract
One of the most demanding needs in cloud computing is that of having scalable and highly available databases. One of the ways to attend these needs is to leverage the scalable replication techniques developed in the last decade. These techniques allow increasing both the availability and scalability of databases. Many replication protocols have been proposed during the last decade. The main research challenge was how to scale under the eager replication model, the one that provides consistency across replicas. In this paper, we examine three eager database replication systems available today: Middle-R, C-JDBC and MySQL Cluster using TPC-W benchmark. We analyze their architecture, replication protocols and compare the performance both in the absence of failures and when there are failures.
Rohit Dhamane, Marta Patiño-Martínez, Valerio Vianello, Ricardo Jiménez-Peris
IDEAS4
2013 A Scalable SIEM Correlation Engine and Its Application to the Olympic Games IT Infrastructure
abstract
The security event correlation scalability has become a major concern for security analysts and IT administrators when considering complex IT infrastructures that need to handle gargantuan amounts of events or wide correlation window spans. The current correlation capabilities of Security Information and Event Management (SIEM), based on a single node in centralized servers, have proved to be insufficient to process large event streams. This paper introduces a step forward in the current state of the art to address the aforementioned problems. The proposed model takes into account the two main aspects of this field: distributed correlation and query parallelization. We present a case study of a multiple-step attack on the Olympic Games IT infrastructure to illustrate the applicability of our approach.
Valerio Vianello, Vincenzo Gulisano, Ricardo Jiménez-Peris, Marta Patiño-Martínez, Rubén Torres, Elsa Prieto
ARES3
2013 Transactional Failure Recovery for a Distributed Key-Value Store
Muhammad Yousuf Ahmad, Bettina Kemme, Ivan Brondino, Marta Patiño-Martínez, Ricardo Jiménez-Peris
Middleware5
2012 StreamCloud: An Elastic and Scalable Data Streaming System
abstract
Many applications in several domains such as telecommunications, network security, large-scale sensor networks, require online processing of continuous data flows. They produce very high loads that requires aggregating the processing capacity of many nodes. Current Stream Processing Engines do not scale with the input load due to single-node bottlenecks. Additionally, they are based on static configurations that lead to either under or overprovisioning. In this paper, we present StreamCloud, a scalable and elastic stream processing engine for processing large data stream volumes. StreamCloud uses a novel parallelization technique that splits queries into subqueries that are allocated to independent sets of nodes in a way that minimizes the distribution overhead. Its elastic protocols exhibit low intrusiveness, enabling effective adjustment of resources to the incoming load. Elasticity is combined with dynamic load balancing to minimize the computational resources used. The paper presents the system design, implementation, and a thorough evaluation of the scalability and elasticity of the fully implemented system.
Vincenzo Gulisano, Ricardo Jiménez-Peris, Marta Patiño-Martínez, Claudio Soriente, Patrick Valduriez
IEEE Trans. Parallel Distributed Syst.2
2011 Distributed data management in 2020?
abstract
Work on distributed data management commenced shortly after the introduction of the relational model in the mid-1970's. 1970's and 1980's were very active periods for the development of distributed relational database technology, and claims were made that in the following ten years centralized databases will be an “antique curiosity” and most organizations will move toward distributed database managers [1]. That prediction has certainly become true, and all commercial DBMSs today are distributed.
M. Tamer Özsu, Patrick Valduriez, Serge Abiteboul, Bettina Kemme, Ricardo Jiménez-Peris, Beng Chin Ooi
ICDE5
2011 Elastic SI-Cache: consistent and scalable caching in multi-tier architectures
Francisco Perez-Sorrosal, Marta Patiño-Martínez, Ricardo Jiménez-Peris, Bettina Kemme
VLDB J.3
2010 Distributed Systems and Algorithms
Pascal Felber, Ricardo Jiménez-Peris, Giovanni Schmid, Pierre Sens 0001
Euro-Par (1)2
2010 StreamCloud: A Large Scale Data Streaming System
abstract
Data streaming has become an important paradigm for the real-time processing of continuous data flows in domains such as finance, telecommunications, networking, Some applications in these domains require to process massive data flows that current technology is unable to manage, that is, streams that, even for a single query operator, require the capacity of potentially many machines. Research efforts on data streaming have mainly focused on scaling in the number of queries or query operators, but overlooked the scalability issue with respect to the stream volume. In this paper, we present StreamCloud a large scale data streaming system for processing large data stream volumes. We focus on how to parallelize continuous queries to obtain a highly scalable data streaming infrastructure. StreamCloud goes beyond the state of the art by using a novel parallelization technique that splits queries into subqueries that are allocated to independent sets of nodes in a way that minimizes the distribution overhead. StreamCloud is implemented as a middleware and is highly independent of the underlying data streaming engine. We explore and evaluate different strategies to parallelize data streaming and tackle with the main bottlenecks and overheads to achieve scalability. The paper presents the system design, implementation and a thorough evaluation of the scalability of the fully implemented system.
Vincenzo Gulisano, Ricardo Jiménez-Peris, Marta Patiño-Martínez, Patrick Valduriez
ICDCS2
2009 PolyVaccine: Protecting Web Servers against Zero-Day, Polymorphic and Metamorphic Exploits
abstract
Today Web servers are ubiquitous having become critical infrastructures of many organizations. However, they are still one of the most vulnerable parts of organizations infrastructure. Exploits are many times used by worms to fast propagate across the full Internet being Web servers one of their main targets. New exploit techniques have arouse in the last few years that have rendered useless traditional IDS techniques based on signature identification. Exploits use polymorphism (code encryption) and metamorphism (code obfuscation) to evade detection from signature-based IDSs. In this paper, we address precisely the topic of how to protect Web servers against zero-day (new), polymorphic, and metamorphic malware embedded in data streams (requests) that target Web servers. We rely on a novel technique to detect harmful binary code injection (i.e., exploits) in HTTP requests that is more efficient than current techniques based on binary code emulation or instrumentation of virtual engines. The detection of exploits is done through sandbox processes. The technique is complemented by another set of techniques such as caching, and pooling, to reduce its cost to neglectable levels.Our technique has little assumptions regarding the exploit unlike previous approaches that assume the existence of sled or getPC code, loops, read of the payload, writes to different addresses, etc. The evaluation shows that caching is highly effective and that the average latency introduced by our system is neglectable.
Luis Campo-Giralte, Ricardo Jiménez-Peris, Marta Patiño-Martínez
SRDS2
2009 Snapshot isolation and integrity constraints in replicated databases
abstract
Database replication is widely used for fault tolerance and performance. However, it requires replica control to keep data copies consistent despite updates. The traditional correctness criterion for the concurrent execution of transactions in a replicated database is 1-copy-serializability. It is based on serializability, the strongest isolation level in a nonreplicated system. In recent years, however, Snapshot Isolation (SI), a slightly weaker isolation level, has become popular in commercial database systems. There exist already several replica control protocols that provide SI in a replicated system. However, most of the correctness reasoning for these protocols has been rather informal. Additionally, most of the work so far ignores the issue of integrity constraints. In this article, we provide a formal definition of 1-copy-SI using and extending a well-established definition of SI in a nonreplicated system. Our definition considers integrity constraints in a way that conforms to the way integrity constraints are handled in commercial systems. We discuss a set of necessary and sufficient conditions for a replicated history to be producible under 1-copy-SI. This makes our formalism a convenient tool to prove the correctness of replica control algorithms.
Yi Lin 0005, Bettina Kemme, Ricardo Jiménez-Peris, Marta Patiño-Martínez, José Enrique Armendáriz-Iñigo
ACM Trans. Database Syst.3
2008 Maximizing quorum availability in multi-clustered systems
abstract
Quorum-based schemes are one of the main abstractions in the design of data-sharing replicated systems. One of the most salient characteristics of a quorum system is its availability for operation, i.e., the probability that there exists a network component in the current system state that contains a quorum. The advent of highly-available global electronic services leads to ubiquitous deployment of multi-clustered replicated architectures with sites from multiple clusters connected by inherently unreliable wide-area networks. Yet, the traditional methods for analyzing availability have not been taking link failures into account.
Roman Vitenberg, Ricardo Jiménez-Peris
PODC2
2008 An Autonomic Approach for Replication of Internet-based Services
abstract
As more and more applications are deployed as Internet-based services, they have to be available anytime anywhere in a seamless manner. This requires the underlying infrastructure to provide scalability, fault tolerance and fast response times. While replicating the services and the data they access across sites that are located in different geographic regions is a promising means to achieve these requirements, data consistency is challenging if data continuously changes and queries are dynamic by nature, as is typical for e-commerce applications.Thus, current WAN replication solutions either trade performance for data consistency or are notable to scale in wide-area settings. In this paper, we present a novel approach to provide performance and consistency for Internet services. One of the main contributions is an autonomic replica placement module that places data copies only on servers close to clients that actually need them. The goal is to find the right trade-off between fast local access and the overhead of keeping data copies consistent. As data access patterns might change over time, reconfiguration is done periodically and online, i.e., allowing sites to receive new data copies or drop data copies while at the same time transaction processing continues in the system.
Damián Serrano, Marta Patiño-Martínez, Ricardo Jiménez-Peris, Bettina Kemme
SRDS3
2008 Scalable and topology-aware reconciliation on P2P networks
Vidal Martins, Esther Pacitti, Manal El Dick, Ricardo Jiménez-Peris
Distributed Parallel Databases4
2008 Dynamic quorums for DHT-based enterprise infrastructures
Roberto Baldoni, Ricardo Jiménez-Peris, Marta Patiño-Martínez, Leonardo Querzoni, Antonino Virgillito
J. Parallel Distributed Comput.2
2007 Aspect Separation in Web Service Orchestration: A Reflective Approach and its Application to Decentralized Execution
abstract
Web service orchestration is becoming widely spread for the creation of composite Web services using standard specifications such as BPEL4WS. The myriad of specifications and aspects that should be considered in orchestrated Web services are resulting in increasing complexity. This complexity leads to software infrastructures difficult to maintain with interwoven code involving different aspects such as security, fault tolerance, distribution, etc. In this paper, we present ZenFlow a reflective BPEL engine that enables to separate the implementation of different aspects among them and from the implementation of the regular orchestration functionality of the BPEL engine. We illustrate its capabilities and performance exercising the reflective interface through a decentralized orchestration use case.
Ricardo Jiménez-Peris, Marta Patiño-Martínez, Ernestina Martel-Jordán, R. Naranjo-Izquierdo
ICWS1
2007 Consistent and Scalable Cache Replication for Multi-tier J2EE Applications
Francisco Perez-Sorrosal, Marta Patiño-Martínez, Ricardo Jiménez-Peris, Bettina Kemme
Middleware3
2007 Boosting Database Replication Scalability through Partial Replication and 1-Copy-Snapshot-Isolation
abstract
Databases have become a crucial component in modern information systems. At the same time, they have become the main bottleneck in most systems. Database replication protocols have been proposed to solve the scalability problem by scaling out in a cluster of sites. Current techniques have attained some degree of scalability, however there are two main limitations to existing approaches. Firstly, most solutions adopt a full replication model where all sites store a full copy of the database. The coordination overhead imposed by keeping all replicas consistent allows such approaches to achieve only medium scalability. Secondly, most replication protocols rely on the traditional consistency criterion, 1-copy-serializability, which limits concurrency, and thus scalability of the system. In this paper, we first analyze analytically the performance gains that can be achieved by various partial replication configurations, i.e., configurations where not all sites store all data. From there, we derive a partial replication protocol that provides 1-copy-snapshot isolation as correctness criterion. We have evaluated the protocol with TPC-W and the results show better scalability than full replication.
Damián Serrano, Marta Patiño-Martínez, Ricardo Jiménez-Peris, Bettina Kemme
PRDC3
2007 Enhancing Edge Computing with Database Replication
abstract
As the use of the Internet continues to grow explosively, edge computing has emerged as an important technique for delivering Web content over the Internet. Edge computing moves data and computation closer to end-users for fast local access and better load distribution. Current approaches use caching, which does not work well with highly dynamic data. In this paper, we propose a different approach to enhance edge computing. Our approach lies in a wide area data replication protocol that enables the delivery of dynamic content with full consistency guarantees and with all the benefits of edge computing, such as low latency and high scalability. What is more, the proposed solution is fully transparent to the applications that are brought to the edge. Our extensive evaluations in a real wide area network using TPC-W show promising results.
Yi Lin 0005, Bettina Kemme, Marta Patiño-Martínez, Ricardo Jiménez-Peris
SRDS4
2007 Enterprise Grids: Challenges Ahead
Ricardo Jiménez-Peris, Marta Patiño-Martínez, Bettina Kemme
J. Grid Comput.1
2006 Highly Available Long Running Transactions and Activities for J2EE Applications
abstract
Today’s business applications are typically built on top of middleware platforms such as J2EE and use transactions that have evolved into long running activities able to adapt to different circumstances. Specifications, such as the J2EE Activity Service, have arised for applications requiring that support. These applications also demand high availability to prevent financial losses and/or service level agreements (SLAs) violations due to service unavailability or crashes. Replication is a means to attain high availability but current middleware does not provide highly available transactions. In the advent of crashes, running transactions abort and the application is forced to re-execute them, what results in a loss of availability and transparency. Most approaches using J2EE consider the replication of either the application server or the database. This results in poor availability when the non-replicated tier crashes. This paper presents a novel J2EE replication support for both, application server and database layers providing highly available transactions and long running activities. Failure masking is transparent to client applications. A prototype has been implemented and evaluated.
Francisco Perez-Sorrosal, Marta Patiño-Martínez, Ricardo Jiménez-Peris, Jaksa Vuckovic
ICDCS3
2006 Lightweight Reflection for Middleware-based Database Replication
abstract
Middleware-based database replication approaches have emerged in the last few years as an alternative to traditional database replication implemented within the database kernel. A middleware approach enables third party vendors to provide high availability solutions, a growing practice nowadays in the software industry. However, middleware solutions often lack scalability and exhibit a number of consistency and performance issues. The reason is that in most cases the middleware has to handle the database as a black box, and hence, cannot take advantage of the many optimizations implemented in the database kernel. Thus, middleware solutions often reimplement key functionality but cannot achieve the same efficiency as a kernel implementation. Reflection has been proposed during the last decade as a fruitful paradigm to separate non-functional aspects from functional ones, simplifying software development and maintenance whilst fostering reuse. However, fully reflective databases are not feasible due to the high cost of reflection. Our claim is that by exposing some minimal database functionality through a lightweight reflective interface, efficient and scalable middleware database replication can be attained. In this paper we explore a wide variety of such lightweight reflective interfaces and discuss what kind of replication algorithms they enable. We also discuss implementation alternatives for some of these interfaces and evaluate their performance
Jorge Salas, Ricardo Jiménez-Peris, Marta Patiño-Martínez, Bettina Kemme
SRDS2
2006 WS-replication: a framework for highly available web services
abstract
Due to the rapid acceptance of web services and its fast spreading, a number of mission-critical systems will be deployed as web services in next years. The availability of those systems must be guaranteed in case of failures and network disconnections. An example of web services for which availability will be a crucial issue are those belonging to coordination web service infrastructure, such as web services for transactional coordination (e.g., WS-CAF and WS-Transaction). These services should remain available despite site and connectivity failures to enable business interactions on a 24x7 basis. Some of the common techniques for attaining availability consist in the use of a clustering approach. However, in an Internet setting a domain can get partitioned from the network due to a link overload or some other connectivity problems. The unavailability of a coordination service impacts the availability of all the partners in the business process. That is, coordination services are an example of critical components that need higher provisions for availability. In this paper, we address this problem by providing an infrastructure, WS-Replication, for WAN replication of web services. The infrastructure is based on a group communication web service, WS-Multicast, that respects the web service autonomy. The transport of WS-Multicast is based on SOAP and relies exclusively on web service technology for interaction across organizations. We have replicated WS-CAF using our WS-Replication framework and evaluated its performance.
Jorge Salas, Francisco Perez-Sorrosal, Marta Patiño-Martínez, Ricardo Jiménez-Peris
WWW4
2005 Consistent Data Replication: Is It Feasible in WANs?
Yi Lin 0005, Bettina Kemme, Marta Patiño-Martínez, Ricardo Jiménez-Peris
Euro-Par4
2005 Dynamic Quorums for DHT-based P2P Networks
abstract
Peer-to-peer systems (P2P) have become a popular technique to architect decentralized systems. However, despite its popularity most P2P systems consist in simple applications such as file sharing or chat systems. The main reason is that more complex applications require levels of consistency that nowadays are not offered by P2P systems. In this paper, we explore how to provide consistency based on distributed mutual exclusion via quorum systems in DHT-based P2P networks. Our results show that quorum systems applied directly to such networks are not scalable due to the high traffic imposed onto the underlying network. The paper introduces some design principles for both quorum systems and protocols using them that help to boost their performance. These design principles consist in dynamic and decentralized selection of quorums and in the exposition and exploitation of internals of the DHT such as the finger table. We show that by combining both design principles it is possible to minimize the number of visited sites and the latency needed to obtain a quorum
Roberto Baldoni, Leonardo Querzoni, Antonino Virgillito, Ricardo Jiménez-Peris, Marta Patiño-Martínez
NCA4
2005 Middleware based Data Replication providing Snapshot Isolation
abstract
Many cluster based replication solutions have been proposed providing scalability and fault-tolerance. Many of these solutions perform replica control in a middleware on top of the database replicas. In such a setting concurrency control is a challenge and is often performed on a table basis. Additionally, some systems put severe requirements on transaction programs (e.g., to declare all objects to be accessed in advance). This paper addresses these issues and presents a middleware-based replication scheme which provides the popular snapshot isolation level at the same tuple-level granularity as database systems like PostgreSQL and Oracle, without any need to declare transaction properties in advance. Both read-only and update transactions can be executed at any replica while providing data consistency at all times. Our approach provides what we call "1-copy-snapshot-isolation" as long as the underlying database replicas provide snapshot isolation. We have implemented our approach as a replicated middleware on top of PostgreSQL replicas. By providing a standard JDBC interface, the middleware is completely transparent to the client program. Fault-tolerance is provided by automatically reconnecting clients in case of crashes. Our middleware shows good performance in terms of response times and scalability.
Yi Lin 0005, Bettina Kemme, Marta Patiño-Martínez, Ricardo Jiménez-Peris
SIGMOD Conference4
2005 ZenFlow: A Visual Web Service Composition Tool for BPEL4WS
abstract
Web services have become a very powerful technology to build service oriented architectures and standardize the access to legacy services. Through Web service composition new added value Web services can be created out of existing ones. Examples of these compositions are virtual organizations, outsourcing, enterprise application integration, business process definitions and business to business inter/intra-enterprise relationships. In order to enable the construction of business processes as composite Web services, a number of composition languages has been proposed by the software industry. However, the handiwork of specifying a business process with these languages through simple text or XML editors is tough, complex and error prone. Visual support can ease the definition of business processes. In this paper, we describe ZenFlow, a visual composition tool for Web services written in BPEL4WS. ZenFlow provides several visual facilities to ease the definition of a business process such as multiple views of a process, syntactic and semantic awareness, filtering, logical zooming capabilities and hierarchical representations.
Alberto Martínez, Marta Patiño-Martínez, Ricardo Jiménez-Peris, Francisco Perez-Sorrosal
VL/HCC3
2005 MIDDLE-R: Consistent database replication at the middleware level
abstract
The widespread use of clusters and Web farms has increased the importance of data replication. In this article, we show how to implement consistent and scalable data replication at the middleware level. We do this by combining transactional concurrency control with group communication primitives. The article presents different replication protocols, argues their correctness, describes their implementation as part of a generic middleware, Middle-R, and proves their feasibility with an extensive performance evaluation. The solution proposed is well suited for a variety of applications including Web farms and distributed object platforms.
Marta Patiño-Martínez, Ricardo Jiménez-Peris, Bettina Kemme, Gustavo Alonso
ACM Trans. Comput. Syst.2
2004 Adaptive Middleware for Data Replication
Jesús M. Milán-Franco, Ricardo Jiménez-Peris, Marta Patiño-Martínez, Bettina Kemme
Middleware2
2003 Are quorums an alternative for data replication?
abstract
Data replication is playing an increasingly important role in the design of parallel information systems. In particular, the widespread use of cluster architectures often requires to replicate data for performance and availability reasons. However, maintaining the consistency of the different replicas is known to cause severe scalability problems. To address this limitation, quorums are often suggested as a way to reduce the overall overhead of replication. In this article, we analyze several quorum types in order to better understand their behavior in practice. The results obtained challenge many of the assumptions behind quorum based replication. Our evaluation indicates that the conventional read-one/write-all-available approach is the best choice for a large range of applications requiring data replication. We believe this is an important result for anybody developing code for computing clusters as the read-one/write-all-available strategy is much simpler to implement and more flexible than quorum-based approaches. In this article, we show that, in addition, it is also the best choice using a number of other selection criteria.
Ricardo Jiménez-Peris, Marta Patiño-Martínez, Gustavo Alonso, Bettina Kemme
ACM Trans. Database Syst.1
2002 Improving the Scalability of Fault-Tolerant Database Clusters
abstract
Replication has become a central element in modem information systems playing a dual role: increase availability and enhance scalability. Unfortunately, most existing protocols increase availability at the cost of scalability; This paper presents architecture, implementation and performance of a middleware based replication tool that provides both availability and better scalability than existing systems. Main characteristics are the usage of specialized broadcast primitives and efficient data propagation.
Ricardo Jiménez-Peris, Marta Patiño-Martínez, Bettina Kemme, Gustavo Alonso
ICDCS1
2002 Non-Intrusive, Parallel Recovery of Replicated Data
abstract
The increasingly widespread use of cluster architectures has resulted in many new application scenarios for data replication. While data replication is, in principle, a well understood problem. recovery of replicated systems has not yet received enough attention. In the case of clusters, recovery procedures are particularly important since they have to keep a high level of availability even during recovery. In fact, recovery is part of the normal operations of any cluster as the cluster is expected to continue working while sites leave or join the system. However, traditional recovery techniques usually require stopping processing. Once a quiescent state has been reached, the system proceeds to synchronize the state of failed or new replicas. In this paper. we concentrate on how to perform recovery in a replication middleware without having to stop processing. The proposed protocol focuses on how to minimize the redundancies that take place during concurrent recovery of several sites.
Ricardo Jiménez-Peris, Marta Patiño-Martínez, Gustavo Alonso
SRDS1
2001 How to Select a Replication Protocol According to Scalability, Availability, and Communication Overhead
abstract
Data replication is playing an increasingly important role in the design of parallel information systems. In particular, the widespread use of cluster architectures in high-performance computing has created many opportunities for applying data replication techniques in new areas. For instance, as part of work related to cluster computing in bioinformatics, we have been confronted with the problem of having to choose an optimal replication strategy in terms of scalability, availability and communication overhead. Thus, we have evaluated several representative replication protocols in order to better understand their behavior in practice. The results obtained are surprising in that they challenge many of the assumptions behind existing protocols. Our evaluation indicates that the conventional read-one/write-all approach is the best choice for a large range of applications requiring data replication. We believe this is an important result for anybody developing code for computing clusters as the read-one/write-all strategy is much simpler to implement and more flexible than quorum-based approaches. In this paper we show that, in addition, it is also the best choice using a number of other selection criteria.
Ricardo Jiménez-Peris, Marta Patiño-Martínez, Bettina Kemme, Gustavo Alonso
SRDS1
2001 A Low-Latency Non-blocking Commit Service
Ricardo Jiménez-Peris, Marta Patiño-Martínez, Gustavo Alonso, Sergio Arévalo
DISC1
2000 Deterministic Scheduling for Transactional Multithreaded Replicas
abstract
One way to implement a fault-tolerant service is by replicating it at sites that fail independently. One of the replication techniques is active replication where each request is executed by all the replicas. Thus, the effects of failures can be completely masked, resulting in an increase of service availability. In order to preserve consistency among replicas, replicas must exhibit a deterministic behavior, which has traditionally been achieved by restricting replicas to being single-threaded. However, this approach cannot be applied in some setups like transactional systems, where it is not admissible to process transactions sequentially. The authors present a deterministic scheduling algorithm for multithreaded replicas in a transactional framework. To ensure replica determinism, requests to replicated servers are submitted by means of reliable and totally ordered multicast. Internally, a deterministic scheduler ensures that all threads are scheduled in the same way at all replicas which guarantees replica consistency.
Ricardo Jiménez-Peris, Marta Patiño-Martínez, Sergio Arévalo
SRDS1
2000 Scalable Replication in Database Clusters
Marta Patiño-Martínez, Ricardo Jiménez-Peris, Bettina Kemme, Gustavo Alonso
DISC2
2000 Using interpreted CompositeCalls to improve operating system services
abstract
A large number of protection domain crossings and context switches is often the cause of bad performance in complex object-oriented systems. We have identified the CompositeCall pattern which has been used to address this problem for decades. The pattern modifies the traditional client/server interaction model so that clients are able to build compound requests that are evaluated in the server domain. We implemented CompositeCalls for both a traditional OS, Linux, and an experimental object-oriented μkernel, Off++. In the first case, we learned about implications of applying CompositeCall to a non-object-oriented ‘legacy’ system. In both experiments, we learned when CompositeCalls help improving system performance and when they do not help. In addition, our experiments gave us important insights about some pernicious design traditions extensively used in OS construction. Copyright © 2000 John Wiley & Sons, Ltd.
Francisco J. Ballesteros, Ricardo Jiménez-Peris, Marta Patiño-Martínez, Fabio Kon, Sergio Arévalo, Roy H. Campbell
Softw. Pract. Exp.2
1999 Adding breadth to CS1 and CS2 courses through visual and interactive programming projects
abstract
The aim of programming projects in CS1/CS2 is to put in practice concepts and techniques learnt during lectures. Programming projects serve a dual purpose: first, the students get to practice the programming concepts taught in class, and second, they are introduced to an array of topics that they will cover later in their computer science education.In this work, we present programming projects we have successfully used in CS1/CS2. These topics have added breadth to CS1/CS2 as well as whetted our students' appetite by exposing them to concurrent programming, event-driven programming, graphics management and human-computer interfaces, data compression, image processing and genetic algorithms.We also include the background material, such as tools and libraries we have provided our students to render the more difficult projects amenable to our introductory computer science classes.
Ricardo Jiménez-Peris, Sami Khuri, Marta Patiño-Martínez
SIGCSE1
1997 The locker metaphor to teach dynamic memory
abstract
Some students experience difficulties when first introduced to dynamic memory. The goal of this paper is to present an analogy between dynamic memory programming and a real-world example that will help students in understanding the underlying concepts behind dynamic memory: a left-luggage room with lockers.
Ricardo Jiménez-Peris, Cristóbal Pareja-Flores, Marta Patiño-Martínez, J. Ángel Velázquez-Iturbide
SIGCSE1
1997 AnLex and AnSin: a compiler generator system for beginners
abstract
The study of compiler generators is an integral part of compiler construction, and for this reason it is customary to have a programming project entirely devoted to it in compiler courses. There are many compilers generators, but their use in a compiler course presents several problems (e.g. the parsers generated are difficult to understand and to debug). In this paper, we describe such problems and present a compiler generator system, AnLex-AnSin, that solves these problems, and can thus be used in compiler programming projects.
Marta Patiño-Martínez, J. Ignacio Castelló-Gómez, Ricardo Jiménez-Peris
SIGCSE3
1996 An overview of visualization: its use and design: report of the working group in visualization
abstract
article Free Access Share on An overview of visualization: its use and design: report of the working group in visualization Authors: Joe Bergin Pace University Pace UniversityView Profile , Ken Brodie University of Leeds, UK University of Leeds, UKView Profile , Marta Patiño-Martínez Universitat Politecnica de Madrid, Spain Universitat Politecnica de Madrid, SpainView Profile , Myles McNally Alma College Alma CollegeView Profile , Tom Naps Lawrence University Lawrence UniversityView Profile , Susan Rodger Duke University Duke UniversityView Profile , Judith Wilson Temple University Temple UniversityView Profile , Michael Goldweber Beloit College Beloit CollegeView Profile , Sami Khuri San Jose State University San Jose State UniversityView Profile , Ricardo Jiménez-Peris Universitat Politecnica de Madrid, Spain Universitat Politecnica de Madrid, SpainView Profile Authors Info & Claims ACM SIGCUE OutlookVolume 24Issue 1-3Jan.-July, 1996 pp 192–200https://doi.org/10.1145/1013718.237647Online:01 January 1996Publication History 30citation851DownloadsMetricsTotal Citations30Total Downloads851Last 12 Months41Last 6 weeks6 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Joseph Bergin, Ken Brodlie, Marta Patiño-Martínez, Myles F. McNally, Thomas L. Naps, Susan H. Rodger, Judith Wilson, Michael Goldweber, Sami Khuri, Ricardo Jiménez-Peris
ITiCSE10
1996 A Modula-2 interpreter/visualizer
abstract
No abstract available.
Ricardo Jiménez-Peris, Marta Patiño-Martínez
ITiCSE1
1996 DD-Mod: a library for teaching distributed programming
abstract
No abstract available.
Ricardo Jiménez-Peris, Marta Patiño-Martínez, Jesús M. Milán-Franco
ITiCSE1
1996 Graphical visualization of the evaluation of functional programs
abstract
article Graphical visualization of the evaluation of functional programs Share on Authors: Ricardo Jiménez-Peris Depto. de Lenguajes y Sistemas Informáticos e Ingeniería del Software, Facultad de Informática. Universidad Politécnica de Madrid, Campus de Montegancedo, s/n. 28060-Madrid, Spain Depto. de Lenguajes y Sistemas Informáticos e Ingeniería del Software, Facultad de Informática. Universidad Politécnica de Madrid, Campus de Montegancedo, s/n. 28060-Madrid, SpainView Profile , Cristóbal Pareja-Flores Depto. de Informática y Automática, Escuela Universitaria de Estadística. Universidad Complutense de Madrid, Avda. Puerta de Hierro, sin. 28040-Madrid. Spain Depto. de Informática y Automática, Escuela Universitaria de Estadística. Universidad Complutense de Madrid, Avda. Puerta de Hierro, sin. 28040-Madrid. SpainView Profile , Marta Patiño-Martínez Depto. de Lenguajes y Sistemas Informáticos e Ingeniería del Software, Facultad de Informática. Universidad Politécnica de Madrid, Campus de Montegancedo, s/n. 28060-Madrid, Spain Depto. de Lenguajes y Sistemas Informáticos e Ingeniería del Software, Facultad de Informática. Universidad Politécnica de Madrid, Campus de Montegancedo, s/n. 28060-Madrid, SpainView Profile , J. Ángel Velázquez-Iturbide Depto. de Lenguajes y Sistemas Informáticos e Ingeniería del Software, Facultad de Informática. Universidad Politécnica de Madrid, Campus de Montegancedo, s/n. 28060-Madrid, Spain Depto. de Lenguajes y Sistemas Informáticos e Ingeniería del Software, Facultad de Informática. Universidad Politécnica de Madrid, Campus de Montegancedo, s/n. 28060-Madrid, SpainView Profile Authors Info & Claims ACM SIGCUE OutlookVolume 24Issue 1-3Jan.-July, 1996 pp 36–38https://doi.org/10.1145/1013718.237520Published:01 January 1996 7citation191DownloadsMetricsTotal Citations7Total Downloads191Last 12 Months4Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Ricardo Jiménez-Peris, Cristóbal Pareja-Flores, Marta Patiño-Martínez, J. Ángel Velázquez-Iturbide
ITiCSE1