VLDB 2026 Research / reviewers in the wild / expert
Carlos Baquero
dblp:42/2941 · also Carlos Baquero-Moreno
· DBLP profile ↗
29ranked-venue papers
2as first author
1since 2021 · last 2021
0000-0002-3933-6850ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 1 since 2021Security and privacy · 3Databases, data management, data science and information retrieval · 3Theory of computation · 2Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Distributed systems · 96% Storage systems · 4% | |
| Computer networks
1 paper |
Internet of things and sensor networks · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems › replication
state machine replication |
0.9 | 2 | 2021 | Efficient replication via timestamp stability · EuroSys 2021 State-machine replication for planet-scale systems · EuroSys 2020 |
Distributed systems › replication › replicated data types
conflict-free replicated data types |
0.4 | 1 | 2019 | Efficient Synchronization of State-Based CRDTs · ICDE 2019 |
Distributed systems
distributed data structures |
0.4 | 1 | 2019 | Efficient Synchronization of State-Based CRDTs · ICDE 2019 |
Distributed systems
replication |
0.4 | 1 | 2019 | Efficient Synchronization of State-Based CRDTs · ICDE 2019 |
Distributed systems
causality tracking |
0.1 | 1 | 2012 | Brief announcement: efficient causality tracking in distributed storage systems with dotted version vectors · PODC 2012 |
Distributed systems
data aggregation |
0.1 | 1 | 2012 | Extrema Propagation: Fast Distributed Estimation of Sums and Network Sizes · IEEE Trans. Parallel Distributed Syst. 2012 |
Distributed systems
distributed coordination and fault tolerance |
0.1 | 1 | 2012 | Brief announcement: efficient causality tracking in distributed storage systems with dotted version vectors · PODC 2012 |
Storage systems
distributed storage |
0.1 | 1 | 2012 | Brief announcement: efficient causality tracking in distributed storage systems with dotted version vectors · PODC 2012 |
Distributed systems › replication
replica consistency |
0.1 | 1 | 2012 | Brief announcement: efficient causality tracking in distributed storage systems with dotted version vectors · PODC 2012 |
Distributed systems
fault tolerance |
0.1 | 1 | 2020 | State-machine replication for planet-scale systems · EuroSys 2020 |
Distributed systems
quorum systems |
0.1 | 1 | 2020 | State-machine replication for planet-scale systems · EuroSys 2020 |
Internet of things and sensor networks › wireless sensor network › distributed algorithms for sensor networks
distributed estimation |
0.0 | 1 | 2012 | Extrema Propagation: Fast Distributed Estimation of Sums and Network Sizes · IEEE Trans. Parallel Distributed Syst. 2012 |
Internet of things and sensor networks
wireless sensor network |
0.0 | 1 | 2012 | Extrema Propagation: Fast Distributed Estimation of Sums and Network Sizes · IEEE Trans. Parallel Distributed Syst. 2012 |
Methods — techniques the papers use, named apart from their topics
timestamping · 0.5stability detection · 0.5single round trip processing · 0.4quorum minimization · 0.4join decomposition · 0.4delta-based synchronization · 0.4epidemic protocols · 0.3duplicate-insensitive message exchange · 0.3probabilistic estimation · 0.1dotted version vectors · 0.1causality tracking · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Efficient replication via timestamp stabilityabstractModern web applications replicate their data across the globe and require strong consistency guarantees for their most critical data. These guarantees are usually provided via state-machine replication (SMR). Recent advances in SMR have focused on leaderless protocols, which improve the availability and performance of traditional Paxos-based solutions. We propose Tempo - a leaderless SMR protocol that, in comparison to prior solutions, achieves superior throughput and offers predictable performance even in contended workloads. To achieve these benefits, Tempo timestamps each application command and executes it only after the timestamp becomes stable, i.e., all commands with a lower timestamp are known. Both the timestamping and stability detection mechanisms are fully decentralized, thus obviating the need for a leader replica. Our protocol furthermore generalizes to partial replication settings, enabling scalability in highly parallel workloads. We evaluate the protocol in both real and simulated geo-distributed environments and demonstrate that it outperforms state-of-the-art alternatives. Vitor Enes, Carlos Baquero, Alexey Gotsman, Pierre Sutra |
EuroSys | 2 |
| 2020 | State-machine replication for planet-scale systemsabstractOnline applications now routinely replicate their data at multiple sites around the world. In this paper we present Atlas, the first state-machine replication protocol tailored for such planet-scale systems. Atlas does not rely on a distinguished leader, so clients enjoy the same quality of service independently of their geographical locations. Furthermore, client-perceived latency improves as we add sites closer to clients. To achieve this, Atlas minimizes the size of its quorums using an observation that concurrent data center failures are rare. It also processes a high percentage of accesses in a single round trip, even when these conflict. We experimentally demonstrate that Atlas consistently outperforms state-of-the-art protocols in planet-scale scenarios. In particular, Atlas is up to two times faster than Flexible Paxos with identical failure assumptions, and more than doubles the performance of Egalitarian Paxos in the YCSB benchmark. Vitor Enes, Carlos Baquero, Tuanir F. Rezende, Alexey Gotsman, Matthieu Perrin, Pierre Sutra |
EuroSys | 2 |
| 2019 | Efficient Synchronization of State-Based CRDTsabstractTo ensure high availability in large scale distributed systems, Conflict-free Replicated Data Types (CRDTs) relax consistency by allowing immediate query and update operations at the local replica, with no need for remote synchronization. State-based CRDTs synchronize replicas by periodically sending their full state to other replicas, which can become extremely costly as the CRDT state grows. Delta-based CRDTs address this problem by producing small incremental states (deltas) to be used in synchronization instead of the full state. However, current synchronization algorithms for delta-based CRDTs induce redundant wasteful delta propagation, performing worse than expected, and surprisingly, no better than state-based. In this paper we: 1) identify two sources of inefficiency in current synchronization algorithms for delta-based CRDTs; 2) bring the concept of join decomposition to state-based CRDTs; 3) exploit join decompositions to obtain optimal deltas and 4) improve the efficiency of synchronization algorithms; and finally, 5) experimentally evaluate the improved algorithms. Vitor Enes, Paulo Sérgio Almeida, Carlos Baquero, João Leitão 0001 |
ICDE | 3 |
| 2019 | Scalable eventually consistent counters over unreliable networks
Paulo Sérgio Almeida, Carlos Baquero |
Distributed Comput. | 2 |
| 2018 | Global-Local View: Scalable Consistency for Concurrent Data Types
Deepthi Devaki Akkoorath, José Brandão 0001, Annette Bieniusa, Carlos Baquero |
Euro-Par | 4 |
| 2018 | Delta state replicated data types
Paulo Sérgio Almeida, Ali Shoker, Carlos Baquero |
J. Parallel Distributed Comput. | 3 |
| 2017 | Aggregation protocols in light of reliable communicationabstractAggregation protocols allow for distributed lightweight computations deployed on ad-hoc networks in a peer-to-peer fashion. Due to reliance on wireless technology, the communication medium is often hostile which makes such protocols susceptible to correctness and performance issues. In this paper, we study the behavior of aggregation protocols when subject to communication failures: message loss, duplication, and network partitions. We show that resolving communication failures at the communication layer, through a simple reliable communication layer, reduces the overhead of using alternative fault tolerance techniques at upper layers, and also preserves the original accuracy and simplicity of protocols. The empirical study we drive shows that tradeoffs exist across various aggregation protocols, and there is no one-size-fits-all protocol. Ziad Kassam, Ali Shoker, Paulo Sérgio Almeida, Carlos Baquero |
NCA | 4 |
| 2017 | Practical evaluation of the Lasp programming model at large scale: an experience reportabstractProgramming models for building large-scale distributed applications assist the developer in reasoning about consistency and distribution. However, many of the programming models for weak consistency, which promise the largest scalability gains, have little in the way of evaluation to demonstrate the promised scalability. We present an experience report on the implementation and large-scale evaluation of one of these models, Lasp, originally presented at PPDP '15, which provides a declarative, functional programming style for distributed applications. We demonstrate the scalability of Lasp's prototype runtime implementation up to 1024 nodes in the Amazon cloud computing environment. It achieves high scalability by uniquely combining hybrid gossip with a programming model based on convergent computation. We report on the engineering challenges of this implementation and its evaluation, specifically related to operating research prototypes in a production cloud environment. Christopher Meiklejohn, Vitor Enes, Junghun Yoo, Carlos Baquero, Peter Van Roy, Annette Bieniusa |
PPDP | 4 |
| 2017 | DottedDB: Anti-Entropy without Merkle Trees, Deletes without TombstonesabstractTo achieve high availability in the face of network partitions, many distributed databases adopt eventual consistency, allow temporary conflicts due to concurrent writes, and use some form of per-key logical clock to detect and resolve such conflicts. Furthermore, nodes synchronize periodically to ensure replica convergence in a process called anti-entropy, normally using Merkle Trees. We present the design of DottedDB, a Dynamo-like key-value store, which uses a novel node-wide logical clock framework, overcoming three fundamental limitations of the state of the art: (1) minimize the metadata per key necessary to track causality, avoiding its growth even in the face of node churn; (2) correctly and durably delete keys, with no need for tombstones; (3) offer a lightweight anti-entropy mechanism to converge replicated data, avoiding the need for Merkle Trees. We evaluate DottedDB against MerkleDB, an otherwise identical database, but using per-key logical clocks and Merkle Trees for anti-entropy, to precisely measure the impact of the novel approach. Results show that: causality metadata per object always converges rapidly to only one id-counter pair; distributed deletes are correctly achieved without global coordination and with constant metadata; divergent nodes are synchronized faster, with less memory-footprint and with less communication overhead than using Merkle Trees. Ricardo Jorge Tome Goncalves, Paulo Sérgio Almeida, Carlos Baquero, Victor Fonte |
SRDS | 3 |
| 2017 | Fault-tolerant aggregation: Flow-Updating meets Mass-Distribution
Paulo Sérgio Almeida, Carlos Baquero, Martin Farach-Colton, Paulo Jesus, Miguel A. Mosteiro |
Distributed Comput. | 2 |
| 2015 | Concise Server-Wide Causality Management for Eventually Consistent Data Stores
Ricardo Gonçalves 0002, Paulo Sérgio Almeida, Carlos Baquero, Victor Fonte |
DAIS | 3 |
| 2015 | Flow updating: Fault-tolerant aggregation for dynamic networks
Paulo Jesus, Carlos Baquero, Paulo Sérgio Almeida |
J. Parallel Distributed Comput. | 2 |
| 2014 | Scalable and Accurate Causality Tracking for Eventually Consistent Stores
Paulo Sérgio Almeida, Carlos Baquero, Ricardo Gonçalves 0002, Nuno M. Preguiça, Victor Fonte |
DAIS | 2 |
| 2014 | Making Operation-Based CRDTs Operation-Based
Carlos Baquero, Paulo Sérgio Almeida, Ali Shoker |
DAIS | 1 |
| 2013 | Topic 8: Distributed Systems and Algorithms - (Introduction)
Achour Mostéfaoui, Andreas Polze, Carlos Baquero, Paul D. Ezhilchelvan, Lars Lundberg |
Euro-Par | 3 |
| 2012 | Spectra: Robust Estimation of Distribution Functions in Networks
Miguel Borges, Paulo Jesus, Carlos Baquero, Paulo Sérgio Almeida |
DAIS | 3 |
| 2012 | Brief announcement: efficient causality tracking in distributed storage systems with dotted version vectorsabstractVersion vectors (VV) are used pervasively to track dependencies between replica versions in multi-version distributed storage systems. In these systems, VV tend to have a dual functionality: identify a version and encode causal dependencies. In this paper, we show that by maintaining the identifier of the version separate from the causal past, it is possible to verify causality in constant time (instead of O(n) for VV) and to precisely track causality with information with size bounded by the degree of replication, and not by the number of concurrent writers. Nuno M. Preguiça, Carlos Baquero, Paulo Sérgio Almeida, Victor Fonte, Ricardo Gonçalves 0002 |
PODC | 2 |
| 2012 | Brief Announcement: Semantics of Eventually Consistent Replicated Sets
Annette Bieniusa, Marek Zawirski, Nuno M. Preguiça, Marc Shapiro 0001, Carlos Baquero, Valter Balegas, Sérgio Duarte |
DISC | 5 |
| 2012 | Extrema Propagation: Fast Distributed Estimation of Sums and Network SizesabstractAggregation of data values plays an important role on distributed computations, in particular, over peer-to-peer and sensor networks, as it can provide a summary of some global system property and direct the actions of self-adaptive distributed algorithms. Examples include using estimates of the network size to dimension distributed hash tables or estimates of the average system load to direct load balancing. Distributed aggregation using nonidempotent functions, like sums, is not trivial as it is not easy to prevent a given value from being accounted for multiple times; this is especially the case if no centralized algorithms or global identifiers can be used. This paper introduces Extrema Propagation, a probabilistic technique for distributed estimation of the sum of positive real numbers. The technique relies on the exchange of duplicate insensitive messages and can be applied in flood and/or epidemic settings, where multipath routing occurs; it is tolerant of message loss; it is fast, as the number of message exchange steps can be made just slightly above the theoretical minimum; and it is fully distributed, with no single point of failure and the result produced at every node. Carlos Baquero, Paulo Sérgio Almeida, Raquel Menezes, Paulo Jesus |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2011 | Fault-Tolerant Aggregation: Flow-Updating Meets Mass-Distribution
Paulo Sérgio Almeida, Carlos Baquero, Martin Farach-Colton, Paulo Jesus, Miguel A. Mosteiro |
OPODIS | 2 |
| 2011 | Ant Colony Optimization with Markov Random Walk for Community Detection in Graphs
Di Jin 0001, Dayou Liu, Bo Yang 0002, Carlos Baquero, Dongxiao He |
PAKDD (2) | 4 |
| 2011 | Conflict-Free Replicated Data Types
Marc Shapiro 0001, Nuno M. Preguiça, Carlos Baquero, Marek Zawirski |
SSS | 3 |
| 2010 | Genetic Algorithm with Local Search for Community Mining in Complex NetworksabstractDetecting communities from complex networks has triggered considerable attention in several application domains. Targeting this problem, a local search based genetic algorithm (GALS) which employs a graph-based representation (LAR) has been proposed in this work. The core of the GALS is a local search based mutation technique. Aiming to overcome the drawbacks of the existing mutation methods, a concept called marginal gene has been proposed, and then an effective and efficient mutation method, combined with a local search strategy which is based on the concept of marginal gene, has also been proposed by analyzing the modularity function. Moreover, in this paper the percolation theory on ER random graphs is employed to further clarify the effectiveness of LAR presentation; A Markov random walk based method is adopted to produce an accurate and diverse initial population; the solution space of GALS will be significantly reduced by using a graph based mechanism. The proposed GALS has been tested on both computer-generated and real-world networks, and compared with some competitive community mining algorithms. Experimental result has shown that GALS is highly effective and efficient for discovering community structure. Di Jin 0001, Dongxiao He, Dayou Liu, Carlos Baquero |
ICTAI (1) | 4 |
| 2010 | Fault-Tolerant Aggregation for Dynamic NetworksabstractData aggregation is a fundamental building block of modern distributed systems. Averaging based approaches, commonly designated gossip-based, are an important class of aggregation algorithms as they allow all nodes to produce a result, converge to any required accuracy, and work independently from the network topology. However, existing approaches exhibit many dependability issues when used in faulty and dynamic environments. This paper extends our own technique, Flow Updating, which is immune to message loss, to operate in dynamic networks, improving its fault tolerance characteristics. Experimental results show that the novel version of Flow Updating vastly outperforms previous averaging algorithms, it self adapts to churn without requiring any periodic restart, supporting node crashes and high levels of message loss. Paulo Jesus, Carlos Baquero, Paulo Sérgio Almeida |
SRDS | 2 |
| 2009 | Fault-Tolerant Aggregation by Flow Updating
Paulo Jesus, Carlos Baquero, Paulo Sérgio Almeida |
DAIS | 2 |
| 2008 | Interval Tree Clocks
Paulo Sérgio Almeida, Carlos Baquero, Victor Fonte |
OPODIS | 2 |
| 2007 | Scalable Bloom Filters
Paulo Sérgio Almeida, Carlos Baquero, Nuno M. Preguiça, David Hutchison 0001 |
Inf. Process. Lett. | 2 |
| 2004 | Bounded Version Vectors
José Bacelar Almeida, Paulo Sérgio Almeida, Carlos Baquero |
DISC | 3 |
| 2002 | Version Stamps - Decentralized Version VectorsabstractVersion vectors and their variants play a central role in update tracking in optimistic distributed systems. Existing mechanisms for a variable number of participants use a mapping from identities to integers, and rely on some form of global configuration or distributed naming protocol to assign unique identifiers to each participant. These approaches are incompatible with replica creation under arbitrary partitions, a typical mode of operation in mobile or poorly connected environments. We present an update tracking mechanism that overcomes this limitation; it departs from the traditional mapping and avoids the use of integer counters, while providing all the functionality of version vectors in what concerns version tracking. Paulo Sérgio Almeida, Carlos Baquero, Victor Fonte |
ICDCS | 2 |