EDBT 2026 Demo / reviewers in the wild / expert
Nuno M. Preguiça
dblp:18/4843
· DBLP profile ↗
52ranked-venue papers
8as first author
9since 2021 · last 2025
0000-0002-1513-1527ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-authorSecurity and privacy · 5 · 1 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Computer networks · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ensuring Convergence and Invariants Without CoordinationabstractThe CAP theorem demonstrates a trade-off between consistency and availability (and, by extension, latency) in systems where network partitions are unavoidable, such as in cloud computing and local-first software. While adopting weak consistency can preserve availability, it may result in inconsistencies that compromise application correctness. Replicated data types provide a principled, coordination-free approach to guarantee convergence but do not consider application invariants. Existing methods for maintaining invariants in replicated systems either rely on coordination - undermining the benefits of weak consistency - or suffer from limited applicability. This paper introduces the No-Op framework, a generic approach for enforcing consistency without coordination while guaranteeing both convergence and invariant preservation. The core idea of the No-Op approach is to resolve conflicts among concurrent operations by prioritising one operation over the other according to programmer-defined conflict resolution policies. This prioritisation transforms the less-preferred operation into a no-side-effect operation, ensuring conflict-free execution. We formalise the model underlying the No-Op framework and introduce a replication protocol built upon it, accompanied by a formal proof of correctness for both the framework and the protocol. Furthermore, we demonstrate the framework’s applicability by showcasing the design of widely used replicated data types and the preservation of a wide range of application invariants. Dina Borrego, Nuno M. Preguiça, Elisa Gonzalez Boix, Carla Ferreira 0001 |
ECOOP | 2 |
| 2024 | Large-Scale Causal Data Replication for Stateful Edge ApplicationsabstractEdge computing is becoming an increasingly popular paradigm, with modern Internet services leveraging hundreds of edge locations to serve their users. However, existing data replication solutions are not designed to operate in this environment, which restricts the edge components of Internet services to operate as read-only caches and entry points for accessing data centers, severely limiting the benefits extracted from the edge. This paper presents Arboreal, a novel distributed data management system for cloud and edge infrastructures that enables stateful edge applications to be deployed with full (read and write) local access to application data, overcoming the limitations of existing solutions. Arboreal's data replication protocol allows it to automatically and dynamically replicate data across edge locations according to application needs, while providing global causal+ consistency. By relying on a hierarchical topology, Arboreal scales to hundreds of edge locations, while recovering from failures in a decentralized and localized manner, without compromising consistency or durability guarantees. Evaluation shows that the scalability of Arboreal heavily outperforms state-of-the-art solutions, while the dynamic replication mechanism allows to effectively support a wide variety of edge scenarios including mobile clients. Pedro Fouto, Nuno M. Preguiça, João Leitão 0001 |
ICDCS | 2 |
| 2024 | Efficient Write Operations in Event Sourcing with ReplicationabstractEvent Sourcing (ES) is an architectural pattern where applications maintain a record of all events that alter their state. While replication of the event $\log$ can enhance dependability, existing ES frameworks do not exploit replication to accelerate write operations. The challenge lies in enabling write operations on replicas without disrupting the sequential ordering of the log or incurring significant communication overhead between geographically distributed replicas. Our approach addresses this by partitioning resources across replicas, ensuring that most operations require only local ordering, reserving interreplica communication for scenarios where resources are limited. We implemented this approach using the Axon framework and evaluated its performance against MongoDB Replica Sets, where writes must go through the primary replica. The results demonstrate that our method significantly improves the speed of local write operations, thereby fully leveraging the benefits of replication. Tiago Rolo, Nuno M. Preguiça, Filipe Araújo |
NCA | 2 |
| 2023 | PS-CRDTs: CRDTs in highly volatile environmentsabstractThe implementation of collaborative applications in highly volatile environments, such as the ones composed of mobile devices, requires low coordination mechanisms. The replication without coordination semantics of Conflict-Free Replicated Data Types (CRDTs) makes them a natural solution for these execution contexts. However, the current CRDT models require each replica to know all other replicas beforehand or to discover them on-the-fly. Such solutions are not compatible with the dynamic ingress and egress of nodes in volatile environments. To cope with this limitation, we propose the Publish/Subscribe Conflict-Free Replicated Data Type (PS-CRDT) model that combines CRDTs with the publish/subscribe interaction model, and, with that, enable the spatial and temporal decoupling of update propagation. We implemented PS-CRDTs in Thyme, a reactive storage system for mobile edge computing. Our experimental results show that PS-CRDTs require less communication than other CRDT-based solutions in volatile environments. António Barreto, Hervé Paulino, João A. Silva, Nuno M. Preguiça |
Future Gener. Comput. Syst. | 4 |
| 2022 | Engage: Session Guarantees for the EdgeabstractEdge computing offers support for latencyconstrained applications, by replicating data in the edge. Edge storage systems need to adopt both partial replication, as only data of interest needs to be replicated, and weak consistency models, to avoid the overhead and latency induced by the coordination mechanisms of strong consistency models. In this context, session guarantees are a powerful tool that can be used to simplify the design of edge applications. This paper presents Engage, a storage system that offers efficient support for session guarantees in a partially replicated edge setting. To achieve this, Engage combines the use of vector clocks and distributed metadata propagation services with a payload propagation scheme tailored for the edge. We have implemented Engage and evaluated its performance experimentally. The results show that, when compared with previous proposals, the combination of techniques employed by Engage reduce both the number of false dependencies, that can slow down the system, and the signaling overhead, while improving the freshness of data exposed to clients. Miguel Belém, Pedro Fouto, Taras Lykhenko, João Leitão 0001, Nuno M. Preguiça, Luís E. T. Rodrigues |
ICCCN | 5 |
| 2022 | Babel: A Framework for Developing Performant and Dependable Distributed ProtocolsabstractPrototyping and implementing distributed algorithms, particularly those that address challenges related with fault-tolerance and dependability, is a time consuming task. This is, in part, due to the need of addressing low level aspects such as management of communication channels, controlling timeouts or periodic tasks, and dealing with concurrency issues. This has a significant impact for researchers that want to build prototypes for conducting experimental evaluation; practitioners that want to compare different design alternatives/solutions; and even for practical teaching activities on distributed algorithms courses. In this paper we present Babel, a novel framework to develop, implement, and execute distributed protocols and systems. Babel promotes an event driven programming and execution model that simplifies the task of translating typical specifications or descriptions of algorithms into performant prototypes, while allowing the programmer to focus on the relevant challenges of these algorithms by transparently handling time consuming low level aspects. Furthermore, Babel provides, and allows the definition of, networking components that can capture different network capabilities (e.g., P2P, Client/Server, p-accrual Failure Detector), making the code mostly independent from the underlying communication aspects. Babel was built to be generic and can be used to implement a wide variety of different classes of distributed protocols. We conduct our experimental work with two relevant case studies, a Peer-to-Peer application and a State Machine Replication application, that show the generality and ease of use of Babel and present competitive performance when compared with significantly more complex implementations. Pedro Fouto, Pedro Ákos Costa, Nuno M. Preguiça, João Leitão 0001 |
SRDS | 3 |
| 2022 | High Throughput Replication with Integrated Membership Management
Pedro Fouto, Nuno M. Preguiça, João Leitão 0001 |
USENIX ATC | 2 |
| 2021 | It's about Thyme: On the design and implementation of a time-aware reactive storage system for pervasive edge computing environmentsabstractNowadays, smart mobile devices generate huge amounts of data in all sorts of gatherings. Much of that data has localized and ephemeral interest, but can be of great use if shared among co-located devices. However, mobile devices often experience poor connectivity, leading to availability issues if application storage and logic are fully delegated to a remote cloud infrastructure. In turn, the edge computing paradigm pushes computations and storage beyond the data center, closer to end-user devices where data is generated and consumed, enabling the execution of certain components of edge-enabled systems directly and cooperatively on edge devices. In this article, we address the challenge of supporting reliable and efficient data storage and dissemination among co-located wireless mobile devices without resorting to centralized services or network infrastructures. We propose Thyme, a novel time-aware reactive data storage system for pervasive edge computing environments, that exploits synergies between the storage substrate and the publish/subscribe paradigm. We present the design of Thyme and elaborate a three-fold evaluation, through an analytical study, and both simulation and real world experimentations, characterizing the scenarios best suited for its use. The evaluation shows that Thyme allows the notification and retrieval of relevant data with low overhead and latency, and also with low energy consumption, proving to be a practical solution in a variety of situations. João A. Silva, Filipe Cerqueira, Hervé Paulino, João Lourenço, João Leitão 0001, Nuno M. Preguiça |
Future Gener. Comput. Syst. | 6 |
| 2021 | ECROs: building global scale systems from sequential codeabstractTo ease the development of geo-distributed applications, replicated data types (RDTs) offer a familiar programming interface while ensuring state convergence, low latency, and high availability. However, RDTs are still designed exclusively by experts using ad-hoc solutions that are error-prone and result in brittle systems. Recent works statically detect conflicting operations on existing data types and coordinate those at runtime to guarantee convergence and preserve application invariants. However, these approaches are too conservative, imposing coordination on a large number of operations. In this work, we propose a principled approach to design and implement efficient RDTs taking into account application invariants. Developers extend sequential data types with a distributed specification, which together form an RDT. We statically analyze the specification to detect conflicts and unravel their cause. This information is then used at runtime to serialize concurrent operations safely and efficiently. Our approach derives a correct RDT from any sequential data type without changes to the data type's implementation and with minimal coordination. We implement our approach in Scala and develop an extensive portfolio of RDTs. The evaluation shows that our approach provides performance similar to conflict-free replicated data types for commutative operations, and considerably improves the performance of non-commutative operations, compared to existing solutions. Kevin De Porre, Carla Ferreira 0001, Nuno M. Preguiça, Elisa Gonzalez Boix |
Proc. ACM Program. Lang. | 3 |
| 2020 | Causality Tracking Trade-offs for Distributed StorageabstractAfter the seminal paper by L. Lamport, which introduced (scalar) logical clocks, several other data structures for keeping track of causality in distributed systems have been proposed, including vector and matrix clocks. These are able to capture causal dependencies with more detail but, unfortunately, also consume a substantially larger amount of network bandwidth and storage space than Lamport clocks. This raises the question of whether the benefits of these more complex structures are worth their cost. We address this question in the context of partially replicated systems. We show that for some workloads the use of more expensive clocks does bring significant benefits and that for other workloads no visible benefits can be observed. The paper provides a characterization of the scenarios where each type of clock is more beneficial, helping designers to develop more efficient distributed storage systems. Hugo Guerreiro, Luís E. T. Rodrigues, Nuno M. Preguiça, Nívia Cruz Quental |
NCA | 3 |
| 2020 | Practical Client-side Replication: Weak Consistency Semantics for Insecure Settings
Albert van der Linde, João Leitão 0001, Nuno M. Preguiça |
Proc. VLDB Endow. | 3 |
| 2019 | Enabling Fog Computing using Self-Organizing Compute NodesabstractThe emergence of fog computing has led to the design of multi-layer fog computing models which are organized hierarchically. These models commonly dictate the hierarchical structure to all the participating compute nodes. However, organizing the compute nodes by adding customized connections that do not abide by the hierarchical approach, may result in improved performance due to the network’s properties i.e., latency or bandwidth between the nodes. For this reason, in this paper we propose an alternative to the hierarchical approach, which is the self-organizing compute nodes. These nodes organize themselves into a flat model which leverages on the network’s properties to provide improved performance. The results of the evaluation show that this approach reduces bandwidth utilization (~30%) by using optimized messaging instead of direct messaging. Furthermore, we show that following a flat model, enables the design of mechanisms for fault tolerance which has been mostly neglected in existing hierarchical models. Vasileios Karagiannis, Stefan Schulte 0002, João Leitão 0001, Nuno M. Preguiça |
ICFEC | 4 |
| 2019 | Time-aware reactive storage in wireless edge environmentsabstractNowadays, smart mobile devices generate huge amounts of data in all sorts of gatherings. Much of that data has localized and ephemeral interest, but can be of great use if shared among co-located devices. However, these devices often experience poor connectivity, leading to availability issues if applications' storage and logic are fully delegated to a remote cloud infrastructure. In turn, the edge computing paradigm pushes computations and storage beyond the data center, closer to end-user devices where data is generated and consumed. Thus, enabling the execution of certain components of edge-enabled systems directly and cooperatively on edge devices. In this paper, we address the challenge of supporting reliable and efficient data storage and dissemination among co-located wireless mobile devices without resorting to centralized services or network infrastructures. We propose Thyme, a novel time-aware reactive data storage system for wireless edge networks, that exploits synergies between the storage substrate and the publish/subscribe paradigm. We present the design of Thyme and evaluate it through simulation, characterizing the scenarios best suited for its use. The evaluation shows that Thyme allows for reliable notification and retrieval of relevant data with low overhead and latency. João A. Silva, Hervé Paulino, João Lourenço, João Leitão 0001, Nuno M. Preguiça |
MobiQuitous | 5 |
| 2018 | Software Rejuvenation in Computer Systems: An Automatic Forecasting Approach Based on Time SeriesabstractDistributed computing is bringing many advantages in cost, flexibility and availability. However, it increases the demand for performance and reliability. Resources such as CPU, memory, storage and network bandwidth, are very susceptible of presenting software aging issues. Therefore, proactive actions, also known as software rejuvenation must be performed to avoid these issues. The identification of the best moment to perform software rejuvenation is not a simple task, mostly because it may affect the system's availability and reliability. To overcome this problem, we propose an automatic forecasting strategy to support the system administrators to choose the best moment to perform software rejuvenation. Our strategy uses six time series techniques: Drift, Simple Exponential Smoothing, Holt, Holt-Winters, Linear Regression, and ARIMA. In our proposal, the most suitable one is chosen automatically as the best fit for a particular scenario. Three case studies were performed to evaluate the efficiency of our automatic strategy. Our proposal aims to increase the system's availability while decreasing the QoS violation probability. In one of our experiments, we can observe a reduction of 92.3% in the system's downtime. This research supports decision making activities and opens possibilities to foster the usage of forecasting strategies when dealing with software aging phenomenon. Jean Araujo 0001, Rúbens de Souza Matos Júnior, Nuno M. Preguiça, Paulo Romero Martins Maciel |
IPCCC | 4 |
| 2018 | Practical and Fast Causal Consistent Partial Geo-ReplicationabstractDistributed storage systems are a fundamental component of large-scale Internet services. To keep up with the increasing expectations of users regarding availability and latency, the design of data storage systems has evolved to achieve these properties by exploiting techniques such as partial replication, geo-replication, and weaker consistency models. How to combine all these techniques in a single solution in a practical and efficient way is highly challenging. In this paper we propose a novel replication scheme that can offer causal+ consistency in a geo-distributed scenario with partial replication, where datacenters replicate different portions of the entire database. We leverage on a recently proposed methodology that decouples the propagation of data and causality-tracking metadata. Our solution presents a novel causal consistency tracking and enforcing algorithm, focusing on maximizing parallelism in the execution of remote operations which, as we show, has a significant influence on the performance of a partially replicated system. We also propose and implement a design to integrate our solution in the popular Cassandra database. Experimental results show that, by exploring a new position in the trade-off between throughput and data visibility (by balancing the execution of local and remote operations, respectively), our solution presents overall good performance. Pedro Fouto, João Leitão 0001, Nuno M. Preguiça |
NCA | 3 |
| 2018 | Fine-grained consistency for geo-replicated systems
Cheng Li 0001, Nuno M. Preguiça, Rodrigo Rodrigues 0001 |
USENIX ATC | 2 |
| 2018 | IPA: Invariant-preserving Applications for Weakly consistent Replicated DatabasesabstractIt is common to use weakly consistent replication to achieve high availability and low latency at a global scale. In this setting, concurrent updates may lead to states where application invariants do not hold. Some systems coordinate the execution of (conflicting) operations to avoid invariant violations, leading to high latency and reduced availability for those operations. This problem is worsened by the difficulty in identifying precisely which operations conflict. In this paper we propose a novel approach to preserve application invariants without coordinating the execution of operations. The approach consists of modifying operations in a way that application invariants are maintained in the presence of concurrent updates. When no conflicting updates occur, the modified operations present their original semantics. Otherwise, we use sensible and deterministic conflict resolution policies that preserve the invariants of the application. To implement this approach, we developed a static analysis, IPA, that identifies conflicting operations and proposes the necessary modifications to operations. Our analysis shows that IPA can avoid invariant violations in many applications, including typical database applications. Our evaluation reveals that the offline static analysis runs fast enough for being used with large applications. The overhead introduced in the modified operations is low and it leads to lower latency and higher throughput when compared with other approaches that enforce invariants. Valter Balegas, Sérgio Duarte, Carla Ferreira 0001, Rodrigo Rodrigues 0001, Nuno M. Preguiça |
Proc. VLDB Endow. | 5 |
| 2017 | Non-Uniform ReplicationabstractReplication is a key technique in the design of efficient and reliable distributed systems. As information grows, it becomes difficult or even impossible to store all information at every replica. A common approach to deal with this problem is to rely on partial replication, where each replica maintains only a part of the total system information. As a consequence, a remote replica might need to be contacted for computing the reply to some given query, which leads to high latency costs particularly in geo-replicated settings. In this work, we introduce the concept of non-uniform replication, where each replica stores only part of the information, but where all replicas store enough information to answer every query. We apply this concept to eventual consistency and conflict-free replicated data types. We show that this model can address useful problems and present two data types that solve such problems. Our evaluation shows that non-uniform replication is more efficient than traditional replication, using less storage space and network bandwidth. Gonçalo Cabrita, Nuno M. Preguiça |
OPODIS | 2 |
| 2017 | Fine-Grained Consistency Upgrades for Online ServicesabstractOnline services such as Facebook or Twitter have public APIs to enable an easy integration of these services with third party applications. However, the developers who design these applications have no information about the consistency provided by these services, which exacerbates the complexity of reasoning about the semantics of the applications they are developing. In this paper, we show that is possible to deploy a transparent middleware between the application and the service, which enables a fine-grained control over the session guarantees that comprise the consistency semantics provided by these APIs, without having to gain access to the implementation of the underlying services. We evaluated our middleware using the Facebook public API and the Redis datastore, and our results show that we are able to provide fine-grained control of the consistency semantics incurring in a small local storage and modest latency overhead. Filipe Freitas, João Leitão 0001, Nuno M. Preguiça, Rodrigo Rodrigues 0001 |
SRDS | 3 |
| 2017 | Legion: Enriching Internet Services with Peer-to-Peer InteractionsabstractMany web applications are built around direct interactions among users, from collaborative applications and social networks to multi-user games. Despite being user-centric, these applications are usually supported by services running on servers that mediate all interactions among clients. When users are in close vicinity of each other, relying on a centralized infrastructure for mediating user interactions leads to unnecessarily high latency while hampering fault-tolerance and scalability. Albert van der Linde, Pedro Fouto, João Leitão 0001, Nuno M. Preguiça, Santiago J. Castiñeira, Annette Bieniusa |
WWW | 4 |
| 2017 | Blotter: Low Latency Transactions for Geo-Replicated StorageabstractMost geo-replicated storage systems use weak consistency to avoid the performance penalty of coordinating replicas in different data centers. This departure from strong semantics poses problems to application programmers, who need to address the anomalies enabled by weak consistency. In this paper we use a recently proposed isolation level, called Non-Monotonic Snapshot Isolation, to achieve ACID transactions with low latency. To this end, we present Blotter, a geo-replicated system that leverages these semantics in the design of a new concurrency control protocol that leaves a small amount of local state during reads to make commits more efficient, which is combined with a configuration of Paxos that is tailored for good performance in wide area settings. Read operations always run on the local data center, and update transactions complete in a small number of message steps to a subset of the replicas. We implemented Blotter as an extension to Cassandra. Our experimental evaluation shows that Blotter has a small overhead at the data center scale, and performs better across data centers when compared with our implementations of the core Spanner protocol and of Snapshot Isolation on the same codebase. Henrique Moniz, João Leitão 0001, Ricardo J. Dias, Johannes Gehrke, Nuno M. Preguiça, Rodrigo Rodrigues 0001 |
WWW | 5 |
| 2016 | Characterizing the Consistency of Online Services (Practical Experience Report)abstractWhile several proposals for the specification and implementation of various consistency models exist, little is known about what is the consistency currently offered by online services with millions of users. Such knowledge is important, not only because it allows for setting the right expectations and justifying the behavior observed by users, but also because it can be used for improving the process of developing applications that use APIs offered by such services. To fill this gap, this paper presents a measurement study of the consistency of the APIs exported by four widely used Internet services, the Facebook Feed, Facebook Groups, Blogger, and Google+. To conduct this study, our work (1) proposes definitions for a set of relevant consistency properties, (2) develops a simple, yet generic methodology comprising a small number of tests, which probe these services from a user perspective, and try to uncover consistency anomalies that are key to our definitions, and (3) reports on the analysis of the data obtained from running these tests for a period of several weeks. Our measurement study shows that some of these services do exhibit consistency anomalies, including some behaviors that may appear counter-intuitive for users, such as the lack of session guarantees for write monotonicity. Filipe Freitas, João Leitão 0001, Nuno M. Preguiça, Rodrigo Rodrigues 0001 |
DSN | 3 |
| 2016 | Cure: Strong Semantics Meets High Availability and Low LatencyabstractDevelopers of cloud-scale applications face a difficult decision of which kind of storage to use, summarised by the CAP theorem. Currently the choice is between classical CP databases, which provide strong guarantees but are slow, expensive, and unavailable under partition, and NoSQL-style AP databases, which are fast and available, but too hard to program against. We present an alternative: Cure provides the highest level of guarantees that remains compatible with availability. These guarantees include: causal consistency (no ordering anomalies), atomicity (consistent multi-key updates), and support for high-level data types (developer friendly API) with safe resolution of concurrent updates (guaranteeing convergence). These guarantees minimise the anomalies caused by parallelism and distribution, thus facilitating the development of applications. This paper presents the protocols for highly available transactions, and an experimental evaluation showing that Cure is able to achieve scalability similar to eventually-consistent NoSQL databases, while providing stronger guarantees. Deepthi Devaki Akkoorath, Alejandro Z. Tomsic, Manuel Bravo, Zhongmiao Li, Tyler Crain, Annette Bieniusa, Nuno M. Preguiça, Marc Shapiro 0001 |
ICDCS | 7 |
| 2015 | Putting consistency back into eventual consistencyabstractGeo-replicated storage systems are at the core of current Internet services. The designers of the replication protocols used by these systems must choose between either supporting low-latency, eventually-consistent operations, or ensuring strong consistency to ease application correctness. We propose an alternative consistency model, Explicit Consistency, that strengthens eventual consistency with a guarantee to preserve specific invariants defined by the applications. Given these application-specific invariants, a system that supports Explicit Consistency identifies which operations would be unsafe under concurrent execution, and allows programmers to select either violation-avoidance or invariant-repair techniques. We show how to achieve the former, while allowing operations to complete locally in the common case, by relying on a reservation system that moves coordination off the critical path of operation execution. The latter, in turn, allows operations to execute without restriction, and restore invariants by applying a repair operation to the database state. We present the design and evaluation of Indigo, a middleware that provides Explicit Consistency on top of a causally-consistent data store. Indigo guarantees strong application invariants while providing similar latency to an eventually-consistent system in the common case. Valter Balegas, Sérgio Duarte, Carla Ferreira 0001, Rodrigo Rodrigues 0001, Nuno M. Preguiça, Mahsa Najafzadeh, Marc Shapiro 0001 |
EuroSys | 5 |
| 2015 | Write Fast, Read in the Past: Causal Consistency for Client-Side ApplicationsabstractClient-side apps (e.g., mobile or in-browser) need cloud data to be available in a local cache, for both reads and updates. For optimal user experience and developer support, the cache should be consistent and fault-tolerant. In order to scale to high numbers of unreliable and resource-poor clients, and large database, the system needs to use resources sparingly. The SwiftCloud distributed object database is the first to provide fast reads and writes via a causally-consistent client-side local cache backed by the cloud. It is thrifty in resources and scales well, thanks to consistent versioning provided by the cloud, using small and bounded metadata. It remains available during faults, switching to a different data centre when the current one is not responsive, while maintaining its consistency guarantees. This paper presents the SwiftCloud algorithms, design, and experimental evaluation. It shows that client-side apps enjoy the high performance and availability, under the same guarantees as a remote cloud data store, at a small cost. Marek Zawirski, Nuno M. Preguiça, Sérgio Duarte, Annette Bieniusa, Valter Balegas, Marc Shapiro 0001 |
Middleware | 2 |
| 2015 | Extending Eventually Consistent Cloud Databases for Enforcing Numeric InvariantsabstractGeo-replicated databases often offer high availability and low latency by relying on weak consistency models. The inability to enforce invariants across all replicas remains a key shortcoming that prevents the adoption of such databases in several applications. In this paper we show how to extend an eventually consistent cloud database for enforcing numeric invariants. Our approach builds on ideas from escrow transactions, but our novel design overcomes the limitations of previous works. First, by relying on a new replicated data type, our design has no central authority and uses pairwise asynchronous communication only. Second, by layering our design on top of a fault-tolerant database, our approach exhibits better availability during network partitions and data center faults. The evaluation of our prototype, built on top of Riak, shows much lower latency and better scalability than the traditional approach of using strong consistency to enforce numeric invariants. Valter Balegas, Diogo Serra, Sérgio Duarte, Carla Ferreira 0001, Marc Shapiro 0001, Rodrigo Rodrigues 0001, Nuno M. Preguiça |
SRDS | 7 |
| 2015 | PIXIDA: Optimizing Data Parallel Jobs in Wide-Area Data AnalyticsabstractIn the era of global-scale services, big data analytical queries are often required to process datasets that span multiple data centers (DCs). In this setting, cross-DC bandwidth is often the scarcest, most volatile, and/or most expensive resource. However, current widely deployed big data analytics frameworks make no attempt to minimize the traffic traversing these links. In this paper, we present P ixida , a scheduler that aims to minimize data movement across resource constrained links. To achieve this, we introduce a new abstraction called S ilo , which is key to modeling P ixida 's scheduling goals as a graph partitioning problem. Furthermore, we show that existing graph partitioning problem formulations do not map to how big data jobs work, causing their solutions to miss opportunities for avoiding data movement. To address this, we formulate a new graph partitioning problem and propose a novel algorithm to solve it. We integrated P ixida in Spark and our experiments show that, when compared to existing schedulers, P ixida achieves a significant traffic reduction of up to ~ 9x on the aforementioned links. Konstantinos Kloudas, Rodrigo Rodrigues 0001, Nuno M. Preguiça, Margarida Mamede |
Proc. VLDB Endow. | 3 |
| 2014 | Scalable and Accurate Causality Tracking for Eventually Consistent Stores
Paulo Sérgio Almeida, Carlos Baquero, Ricardo Gonçalves 0002, Nuno M. Preguiça, Victor Fonte |
DAIS | 4 |
| 2014 | Automating the Choice of Consistency Levels in Replicated Systems
Cheng Li 0001, João Leitão 0001, Allen Clement, Nuno M. Preguiça, Rodrigo Rodrigues 0001, Viktor Vafeiadis |
USENIX ATC | 4 |
| 2013 | Concurrency control and awareness support for multi-synchronous collaborative editingabstractCollaborative editing tools have become increasingly popular in the last decade, with some systems being used by massive numbers of users. While traditionally collaborative editing systems would either target synchronous or asynchronous collaboration settings, some recent systems support both types Mehdi Ahmed-Nacer, Pascal Urso, Valter Balegas, Nuno M. Preguiça |
CollaborateCom | 4 |
| 2013 | On the Scalability of Snapshot Isolation
Masoud Saeida Ardekani, Pierre Sutra, Marc Shapiro 0001, Nuno M. Preguiça |
Euro-Par | 4 |
| 2013 | MacroDB: Scaling Database Engines on Multicores
João Soares 0003, João Lourenço, Nuno M. Preguiça |
Euro-Par | 3 |
| 2013 | Scalable Data Processing for Community Sensing Applications
Sérgio Duarte, David Navalho, Heitor Ferreira, Nuno M. Preguiça |
Mob. Networks Appl. | 4 |
| 2012 | Making Geo-Replicated Systems Fast as Possible, Consistent when Necessary
Cheng Li 0001, Daniel Porto 0002, Allen Clement, Johannes Gehrke, Nuno M. Preguiça, Rodrigo Rodrigues 0001 |
OSDI | 5 |
| 2012 | Brief announcement: efficient causality tracking in distributed storage systems with dotted version vectorsabstractVersion vectors (VV) are used pervasively to track dependencies between replica versions in multi-version distributed storage systems. In these systems, VV tend to have a dual functionality: identify a version and encode causal dependencies. In this paper, we show that by maintaining the identifier of the version separate from the causal past, it is possible to verify causality in constant time (instead of O(n) for VV) and to precisely track causality with information with size bounded by the degree of replication, and not by the number of concurrent writers. Nuno M. Preguiça, Carlos Baquero, Paulo Sérgio Almeida, Victor Fonte, Ricardo Gonçalves 0002 |
PODC | 1 |
| 2012 | Brief Announcement: Semantics of Eventually Consistent Replicated Sets
Annette Bieniusa, Marek Zawirski, Nuno M. Preguiça, Marc Shapiro 0001, Carlos Baquero, Valter Balegas, Sérgio Duarte |
DISC | 3 |
| 2011 | Combining Mobile and Cloud Storage for Providing Ubiquitous Data Access
João Soares 0003, Nuno M. Preguiça |
Euro-Par (1) | 2 |
| 2011 | Efficient middleware for byzantine fault tolerant database replicationabstractByzantine fault tolerance (BFT) enhances the reliability and availability of replicated systems subject to software bugs, malicious attacks, or other unexpected events. This paper presents Byzantium, a BFT database replication middleware that provides snapshot isolation semantics. It is the first BFT database system that allows for concurrent transaction execution without relying on a centralized component, which is essential for having both performance and robustness. Byzantium builds on an existing BFT library but extends it with a set of techniques for increasing concurrency in the execution of operations, for optimistically executing operations in a single replica, and for striping and load-balancing read operations across replicas. Experimental results show that our replication protocols introduce only a modest performance overhead for read-write dominated workloads and perform better than a non-replicated database system for read-only workloads. Rui Garcia, Rodrigo Rodrigues 0001, Nuno M. Preguiça |
EuroSys | 3 |
| 2011 | Scalable Data Processing for Community Sensing Applications
Heitor Ferreira, Sérgio Duarte, Nuno M. Preguiça, David Navalho |
MobiQuitous | 3 |
| 2011 | Conflict-Free Replicated Data Types
Marc Shapiro 0001, Nuno M. Preguiça, Carlos Baquero, Marek Zawirski |
SSS | 2 |
| 2010 | Collaborative Cellular-Based Location System
David Navalho, Nuno M. Preguiça |
Euro-Par (2) | 2 |
| 2010 | 4Sensing -- Decentralized Processing for Participatory Sensing DataabstractParticipatory Sensing is an emerging application paradigm that leverages the growing ubiquity of sensor-capable smart phones to allow communities carry out wide-area sensing tasks, as a side-effect of people's everyday lives and movements. This paper proposes a decentralized infrastructure for supporting Participatory Sensing applications. It describes an architecture and a domain specific programming language for modeling, prototyping and developing the distributed processing of participatory sensing data with the goal of allowing faster and easier development of these applications. Moreover, a case-study application is also presented as the basis for an experimental evaluation. Heitor Ferreira, Sérgio Duarte, Nuno M. Preguiça |
ICPADS | 3 |
| 2009 | A Commutative Replicated Data Type for Cooperative EditingabstractA commutative replicated data type (CRDT) is one where all concurrent operations commute. The replicas of a CRDT converge automatically, without complex concurrency control. This paper describes Treedoc, a novel CRDT design for cooperative text editing. An essential property is that the identifiers of Treedoc atoms are selected from a dense space. We discuss practical alternatives for implementing the identifier space based on an extended binary tree. We also discuss storage alternatives for data and meta-data, and mechanisms for compacting the tree. In the best case, Treedoc incurs no overhead with respect to a linear text buffer. We validate the results with traces from existing edit histories. Nuno M. Preguiça, Joan Manuel Marquès, Marc Shapiro 0001, Mihai Letia |
ICDCS | 1 |
| 2007 | Topic 14 Mobile and Ubiquitous Computing
Nuno M. Preguiça, Eric Fleury, Holger Karl, Gerd Kortuem |
Euro-Par | 1 |
| 2007 | Scalable Bloom Filters
Paulo Sérgio Almeida, Carlos Baquero, Nuno M. Preguiça, David Hutchison 0001 |
Inf. Process. Lett. | 3 |
| 2006 | Supporting Multi-synchronous Groupware: Data Management Problems and a SolutionabstractIt is common that, in a long-term asynchronous collaborative activity, groups of users engage in occasional synchronous sessions. In this paper, we analyze the data management requirements for supporting this common work practice in typical collaborative activities and applications. We call the applications that support such work practice multi-synchronous applications. This analysis shows that, as users interact in different ways in each setting, some applications have different requirements and need to rely on different data sharing techniques in synchronous and asynchronous settings. We present a data management system that allows to integrate a synchronous session in the context of a long-term asynchronous interaction, using the suitable data sharing techniques in each setting and an automatic mechanism to convert the long sequence of small updates produced in a synchronous session into a large asynchronous contribution. We exemplify the use of our approach with two multi-synchronous applications. Nuno M. Preguiça, José Legatheaux Martins, Henrique João L. Domingos, Sérgio Duarte |
Int. J. Cooperative Inf. Syst. | 1 |
| 2005 | Topic 14 - Mobile and Ubiquitous Computing
Evaggelia Pitoura, Marios D. Dikaiakos, Valérie Issarny, Nuno M. Preguiça |
Euro-Par | 4 |
| 2004 | Rufis: Mobile Data Sharing Using a Generic Constraint-Oriented ReconcilerabstractExisting systems for disconnected data access and reconciliation are monolithic, complex and somewhat ad-hoc. In contrast, we demonstrate here a principled approach based on a general-purpose reconciliation engine. We describe the Reconcilable and Undoable File System, Rufis, implemented on top of the IceCube reconciler. IceCube is generic but supports application-specific reconciliation invariants. Consequently, the code for Rufis is quite small and simple, and the reconciliation logic is well separated from the main file system code. Furthermore, Rufis supports specialised reconciliation for files containing data of known types and enables ad-hoc user scenarios involving multiple applications. Marc Shapiro 0001, Nuno M. Preguiça, James O'Brien |
Mobile Data Management | 2 |
| 2003 | Reservations for Conflict Avoidance in a Mobile Database SystemabstractMobile computing characteristics demand data management systems to support independent operation. However, the execution of updates in a mobile client usually need to be considered tentative because uncoordinated updates that conflict need to be reconciled. In this paper we present a mechanism to independently guarantee that updates can be executed in the server without conflicts. To this end, clients obtain leased reservations upon the database state. Updates are specified as common small PL/SQL programs, dubbed mobile transactions, that execute both in the mobile client and in the server. Using the available reservations, the client transparently verifies that a transaction can be executed in the same way both in the mobile client and in the server, thus leading to the same final result. Mobile transactions may specify conflict detection and resolution rules to be used when transactions cannot be locally guaranteed. Nuno M. Preguiça, José Legatheaux Martins, Miguel Cunha, Henrique João L. Domingos |
MobiSys | 1 |
| 2001 | Supporting Disconnected Operation in DOORSabstractThe increasing popularity of portable computers opens the possibility of collaboration among multiple distributed and disconnected users. In such environments, collaboration is often achieved through the concurrent modification of shared data. DOORS is a distributed object store to support asynchronous collaboration in distributed systems that may contain disconnected computers. In this summary we focus on the mechanisms to support disconnected operation. The DOORS architecture is composed by servers that replicate objects using an epidemic propagation model. Clients cache key objects to support disconnected operation. Users run applications to read and modify the shared data (independently from other users)-a read any/write any model of data access is used. Modifications are propagated from clients to servers and among servers as sequences of operations-the system is log-based. Objects are structured according to an object framework that decomposes object operation in several components. Each component manages a different aspect of object execution. Each object represents a data type (e.g. a structured document) and it is composed by a set of sub-objects. Each sub-object represents a subpart of the data type (e.g. sections). A new object is created composing the set of subobjects that store the type-specific data with the adequate implementations of the other components. The following main characteristics are the base to support disconnected operation in DOORS. Nuno M. Preguiça, José Legatheaux Martins, Sérgio Duarte, Henrique João L. Domingos |
HotOS | 1 |
| 2001 | Revisiting Hierarchical Quorum SystemsabstractIn distributed systems, it is often necessary to provide coordination among the multiple concurrent processes. Quorum systems provide a decentralized approach to provide such coordination that is resilient to node and communication link failures. Quorum systems are highly available and may be used to balance the load among the elements of the system. In this paper, we propose a modification to the hierarchical grid quorum system that leads to a smaller quorum size and better availability and load. We also propose a new hierarchical quorum construction based on the organization of elements in a triangular shape that presents better average quorum size, availability and load than other highly-available systems with almost optimal load. Nuno M. Preguiça, José Legatheaux Martins |
ICDCS | 1 |
| 2000 | Data management support for asynchronous groupwareabstractIn asynchronous collaborative applications, users usually collaborate accessing and modifying shared information independently. We have designed and implemented a replicated object store to support such applications in distributed environments that include mobile computers. Unlike most data management systems, awareness support is integrated in the system. To improve the chance for new contributions, the system provides high data availability. The development of applications is supported by an object framework that decomposes objects in several components, each one managing a different aspect of object "execution". New data types may be created relying on pre-defined components to handle concurrent updates, awareness information, etc. Nuno M. Preguiça, José Legatheaux Martins, Henrique João L. Domingos, Sérgio Duarte |
CSCW | 1 |