EDBT 2026 Demo / reviewers in the wild / expert
João Barreto 0001
dblp:b/JoaoPBarreto2 · also João Pedro Barreto 0002
· DBLP profile ↗
31ranked-venue papers
10as first author
8since 2021 · last 2026
0000-0002-0726-2025ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FUR: Fast and Unlimited Reads on Persistent Memory TransactionsabstractDespite the recent improvements in supporting Persistent Hardware Transactions (PHTs) on emerging persistent memories (PM), they have largely overlooked the poor performance of Read-Only (RO) transactions, which suffer from two crucial bottlenecks: i) the considerable post-commit delays required to ensure consistency with concurrent update transactions; and ii) the well-known tight read capacity limits of the commercially available HTM implementations. João Barreto 0001, Daniel Castro 0004, Paolo Romano 0002, Alexandro Baldassin |
EuroSys | 1 |
| 2025 | Better Memory Tiering, Right from the First Placement
João Póvoas, João Barreto 0001, Bartosz Chominski, Fedar Karabeinikau, Maciej Maciejewski, Jakub Schmiegel, Kostiantyn Storozhuk |
ICPE | 2 |
| 2023 | NimbleChain: Speeding up Cryptocurrencies in General-purpose Permissionless BlockchainsabstractNakamoto’s seminal work gave rise to permissionless blockchains, as well as a wide range of proposals to mitigate their performance shortcomings. Despite substantial throughput and energy efficiency achievements, most proposals only bring modest (or marginal) gains in transaction commit latency. Consequently, commit latencies in today’s permissionless blockchain landscape remain prohibitively high. This article proposes NimbleChain, a novel algorithm that extends permissionless blockchains based on Nakamoto consensus with a fast path that delivers causal promises of commitment or simply promises . Since promises only partially order transactions, their latency is only a small fraction of the totally ordered commitment latency of Nakamoto consensus. Still, the weak consistency guarantees of promises are strong enough to correctly implement cryptocurrencies. To the best of our knowledge, NimbleChain is the first system to bring together fast, partially ordered transactions with consensus-based, totally ordered transactions in a permissionless setting. This hybrid consistency model is able to speed up cryptocurrency transactions while still supporting smart contracts, which typically have (strong) sequential consistency needs. We implement NimbleChain as an extension of Ethereum and evaluate it in a 500-node geo-distributed deployment. The results show NimbleChain can promise a cryptocurrency transactions up to an order of magnitude faster than a vanilla Ethereum implementation, with marginal overheads. Miguel Matos, João Barreto 0001 |
Distributed Ledger Technol. Res. Pract. | 3 |
| 2022 | Towards OmpSs-2 and OpenACC interoperationabstractThe increasing demand in HPC to utilize accelerators has motivated the development of pragma-based directives to target these devices. OmpSs-2 and OpenACC are both directive-based solutions that allow application programmers to utilize accelerators. The two leverage distinct types of parallelism: task parallelism and data parallelism, respectively. Non-trivial scientific applications can benefit from both types of available parallelism. However, the combination of pragma-based models is difficult to coordinate, as both assume full control and are unaware of each other at runtime. We propose an interoperation mechanism to enable novel composability across pragma-based programming models. We study and propose a clear separation of duties and implement our approach by augmenting the OmpSs-2 programming model, compiler and runtime to support OmpSs-2 + OpenACC programming. Orestis Korakitis, Simon Garcia de Gonzalo, Nicolas L. Guidotti, João Barreto 0001, José Monteiro 0001, Antonio J. Peña |
PPoPP | 4 |
| 2021 | Generalizing QoS-Aware Memory Bandwidth Allocation to Multi-Socket Cloud ServersabstractAlthough the problem of QoS-aware resource allocation is not new, novel hardware-based resource allocation mechanisms have recently become available in commodity cloud servers and enabled a new generation of QoS-aware resource allocation approaches. Unfortunately, to the best of our knowledge, existing proposals are by design tailored to single-socket architectures only. In many warehouse scale data centers, dualsocket (or even larger) machines already constitute the largest share of hosts. This paper presents the full design and implementation of BALM, a QoS-aware memory bandwidth allocation technique for multi-socket architectures. BALM combines commodity bandwidth allocation mechanisms originally designed for single-socket with a novel adaptive cross-socket page migration scheme. Our evaluation with a large and dynamic set of real applications co-located on a dual-socket machine shows that BALM can overcome the efficiency limitations of state-of-the-art. BALM delivers substantial throughput gains to bandwidth-intensive best-effort applications, while ensuring marginal SLO violation windows to latency-critical applications. David Gureya, João Barreto 0001, Vladimir Vlassov |
CLOUD | 2 |
| 2021 | Particle-In-Cell Simulation Using Asynchronous Tasking
Nicolas L. Guidotti, Pedro Ceyrat, João Barreto 0001, José Monteiro 0001, Rodrigo Rodrigues 0001, Ricardo Fonseca, Xavier Martorell, Antonio J. Peña |
Euro-Par | 3 |
| 2021 | SPHT: Scalable Persistent Hardware Transactions
Daniel Castro 0004, Alexandro Baldassin, João Barreto 0001, Paolo Romano 0002 |
FAST | 3 |
| 2021 | BALM: QoS-Aware Memory Bandwidth Partitioning for Multi-Socket Cloud NodesabstractThe recent emergence of novel hardware-based resource partitioning mechanisms has unveiled the opportunity for a new generation of QoS-aware resource allocation approaches for workload consolidation. Still, to the best of our knowledge, existing proposals are, by design, not tailored to the growing prevalence of multi-socket systems in contemporary warehouse-scale data centers. We propose BALM, a QoS-aware memory bandwidth allocation technique for multi-socket architectures that combines commodity bandwidth allocation mechanisms with a novel adaptive cross-socket page migration scheme. Our experimental evaluation with real applications on a dual-socket machine shows that BALM can overcome the efficiency limitations of state-of-the-art. BALM can ensure marginal SLO violation windows while delivering up to 87% throughput gains to bandwidth-intensive best-effort applications when compared to state-of-the-art alternatives. David Gureya, Vladimir Vlassov, João Barreto 0001 |
SPAA | 3 |
| 2020 | Impact of Geo-Distribution and Mining Pools on Blockchains: A Study of EthereumabstractGiven the large adoption and economical impact of permissionless blockchains, the complexity of the underlying systems and the adversarial environment in which they operate, it is fundamental to properly study and understand the emergent behavior and properties of these systems. We describe our experience on a detailed, one-month study of the Ethereum network from several geographically dispersed observation points. We leverage multiple geographic vantage points to assess the key pillars of Ethereum, namely geographical dispersion, network efficiency, blockchain efficiency and security, and the impact of mining pools. Among other new findings, we identify previously undocumented forms of selfish behavior and show that the prevalence of powerful mining pools exacerbates the geographical impact on block propagation delays. Furthermore, we provide a set of open measurement and processing tools, as well as the data set of the collected measurements, in order to promote further research on understanding permissionless blockchains. David Vavricka, João Barreto 0001, Miguel Matos |
DSN | 3 |
| 2020 | NV-PhTM: An Efficient Phase-Based Transactional System for Non-volatile Memory
Alexandro Baldassin, Rafael Murari, João P. L. de Carvalho, Guido Araujo, Daniel Castro 0004, João Barreto 0001, Paolo Romano 0002 |
Euro-Par | 6 |
| 2020 | Bandwidth-Aware Page Placement in NUMAabstractPage placement is a critical problem for memory-intensive applications running on a shared-memory multiprocessor with a non-uniform memory access (NUMA) architecture. State-of-the-art page placement mechanisms interleave pages evenly across NUMA nodes. However, this approach fails to maximize memory throughput in modern NUMA systems, characterized by asymmetric bandwidths and latencies, and sensitive to memory contention and interconnect congestion phenomena.We propose BWAP, a novel page placement mechanism based on asymmetric weighted page interleaving. BWAP combines an analytical performance model of the target NUMA system with on-line iterative tuning of page distribution for a given memory-intensive application. Our experimental evaluation with representative memory-intensive workloads shows that BWAP performs up to 66% better than state-of-the-art techniques. These gains are particularly relevant when multiple co-located applications run in disjoint partitions of a large NUMA machine or when applications do not scale up to the total number of cores. David Gureya, João Neto 0001, Reza Karimi, João Barreto 0001, Pramod Bhatotia, Vivien Quéma, Rodrigo Rodrigues 0001, Paolo Romano 0002, Vladimir Vlassov |
IPDPS | 4 |
| 2019 | Bicycle Mode Activity Detection with Bluetooth Low Energy BeaconsabstractIn a growing number of cities, cycling is being seriously considered to help solving traffic congestion, parking, etc. It is also a cleaner and healthier mean of urban transportation. However, changing users behavior (e.g., using a bicycle instead of a car) is not simple. Thus, to promote and motivate cycling we propose Biklio, a cycling rewarding system that, based on the use of a smartphone, detects when a user starts cycling and makes her/him eligible for rewards. This solution uses a smartphone application that includes a bicycle usage detection component. This component is both highly accurate and cheap, while respecting other requirements, and is based on the use of a Bluetooth Low Energy (BLE) sensor, installed on each bicycle, which is detected by smartphones. The system is implemented and running, and the results obtained are very encouraging. Paulo Ferreira 0001, Andriy Zabolotny, João Barreto 0001 |
NCA | 3 |
| 2019 | Stretching the capacity of hardware transactional memory in IBM POWER architecturesabstractThe hardware transactional memory (HTM) implementations in commercially available processors are significantly hindered by their tight capacity constraints. In practice, this renders current HTMs unsuitable to many real-world workloads of in-memory databases. Ricardo Filipe, Shady Issa, Paolo Romano 0002, João Barreto 0001 |
PPoPP | 4 |
| 2019 | Hardware Transactional Memory meets memory persistency
Daniel Castro 0004, Paolo Romano 0002, João Barreto 0001 |
J. Parallel Distributed Comput. | 3 |
| 2018 | Hardware Transactional Memory Meets Memory PersistencyabstractPersistent Memory (PM) and Hardware Transactional Memory (HTM) are two recent architectural developments whose joint usage promises to drastically accelerate the performance of concurrent, data-intensive applications. Unfortunately, combining these two mechanisms using existing architectural supports is far from being trivial. This paper presents NV-HTM, a system that allows the execution of transactions over PM using unmodified commodity HTM implementations. NV-HTM relies on a hardware-software co-design technique, which is based on three key ideas: i) relying on software to persist transactional modifications after they have been committed via HTM; ii) postponing the externalization of commit events to applications until it is ensured, via software, that any data version produced and observed by committed transactions is first logged in PM; ii) pruning the commit logs via checkpointing schemes that not only bound the log space and recovery time, but also implement wear levelling techniques to enhance PM's endurance. By means of an extensive experimental evaluation, we show that NV-HTM can achieve up to 10× speed-ups and up to 11.6× reduced flush operations with respect to state of the art solutions, which, unlike NV-HTM, require custom modifications to existing HTM systems. Daniel Castro 0004, Paolo Romano 0002, João Barreto 0001 |
IPDPS | 3 |
| 2018 | Online Tuning of Parallelism Degree in Parallel Nesting Transactional MemoryabstractThis paper addresses the problem of self-tuning the parallelism degree in Transactional Memory (TM) systems that support parallel nesting (PN-TM). This problem has been long investigated for TMs not supporting nesting, but, to the best of our knowledge, has never been studied in the context of PN-TMs. Indeed, the problem complexity is inherently exacerbated in PN-TMs, since these require to identify the optimal parallelism degree not only for top-level transactions but also for nested sub-transactions. The increase of the problem dimensionality raises new challenges (e.g., increase of the search space, and proneness to suffer from local maxima), which are unsatisfactorily addressed by self-tuning solutions conceived for flat nesting TMs. We tackle these challenges by proposing AUTOPN, an on-line self-tuning system that combines model-driven learning techniques with localized search heuristics in order to pursue a twofold goal: i) enhance convergence speed by identifying the most promising region of the search space via model-driven techniques, while ii) increasing robustness against modeling errors, via a final local search phase aimed at refining the model's prediction. We further address the problem of tuning the duration of the monitoring windows used to collect feedback on the system's performance, by introducing novel, domain-specific, mechanisms aimed to strike an optimal trade-off between latency and accuracy of the self-tuning process. We integrated AUTOPN with a state of the art PN-TM (JVSTM) and evaluated it via an extensive experimental study. The results of this study highlight that AUTOPN can achieve gains of up to 45× in terms of increased accuracy and 4× faster convergence speed, when compared with several on-line optimization techniques (gradient descent, simulated annealing and genetic algorithm), some of which were already successfully used in the context of flat nesting TMs. Jingna Zeng, Paolo Romano 0002, João Barreto 0001, Luís E. T. Rodrigues, Seif Haridi |
IPDPS | 3 |
| 2016 | The Future(s) of Transactional MemoryabstractThis work investigates how to combine two powerful abstractions to manage concurrent programming: Transactional Memory (TM) and futures. The former hides from programmers the complexity of synchronizing concurrent access to shared data, via the familiar abstraction of atomic transactions. The latter serves to schedule and synchronize the parallel execution of computations whose results are not immediately required. While TM and futures are two widely investigated topics, the problem of how to exploit these two abstractions in synergy is still largely unexplored in the literature. This paper fills this gap by introducing Java Transactional Futures (JTF), a Java-based TM implementation that allows programmers to use futures to coordinate the execution of parallel tasks, while leveraging transactions to synchronize accesses to shared data. JTF provides a simple and intuitive semantic regarding the admissible serialization orders of the futures spawned by transactions, by ensuring that the results produced by a future are always consistent with those that one would obtain by executing the future sequentially. Our experimental results show that the use of futures in a TM allows not only to unlock parallelism within transactions, but also to reduce the cost of conflicts among top-level transactions in high contention workloads. Jingna Zeng, João Barreto 0001, Seif Haridi, Luís E. T. Rodrigues, Paolo Romano 0002 |
ICPP | 2 |
| 2016 | RUBIC: Online Parallelism Tuning for Co-located Transactional Memory ApplicationsabstractWith the advent of Chip-Multiprocessors, Transactional Memory (TM) emerged as a powerful paradigm to simplify parallel programming. Unfortunately, as more cores become available in commodity systems, the scalability limits of a wide class of TM applications become more evident. Hence, online parallelism tuning techniques were proposed to adapt the optimal number of threads of TM applications. However, state-of-the-art solutions are exclusively tailored to single-process systems with relatively static workloads, exhibiting pathological behaviors in scenarios where multiple multi-threaded TM processes contend for the shared hardware resources. Amin Mohtasham, João Barreto 0001 |
SPAA | 2 |
| 2015 | FRAME: Fair Resource Allocation in Multi-process EnvironmentsabstractAs the technology trend moves toward manufacturing many-core systems with hundreds of processing cores, the problem of efficiently managing multiple parallel jobs on such massively parallel systems becomes increasingly important. With traditional time-sharing each process assumes it is the only running process. This assumption can easily lead to system oversubscription and thereby, losing overall performance due to frequent context switches. Space-sharing techniques allocate a certain number of hardware cores to each process and by that malleable processes can set their parallelism level to their allocated number of cores, hence avoiding oversubscription. However, finding the optimal spatial allocation is not a trivial task. In this paper we propose FRAME, a resource allocation technique to maximize a system's overall utility and fairness, running multiple malleable processes with CPU-bound workloads. First, we formalize the resource allocation problem as an NP-hard problem. Then, we use approximation techniques and convex optimization theory to find the optimal solution to the formulated problem, in pseudo-polynomial time. Our evaluation results show that our method is very fast and efficient in finding the optimal solution to the resource allocation problem. Also, the results suggests that the found solution increases the system's overall utility by 48%, in average, with regard to the best alternative allocation policy. Amin Mohtasham, Ricardo Filipe, João Barreto 0001 |
ICPADS | 3 |
| 2015 | Brief Announcement: Fair Adaptive Parallelism for Concurrent Transactional Memory ApplicationsabstractModern parallel machines are likely to run multiple parallel processes together. However, collocating parallel processes in a single machine can easily result in cross-process and cross-thread interferences that can dramatically degrade the system's performance. Such interferences can be mitigated by dynamically adjusting each process' parallelism towards a fair and efficient configuration. We propose a decentralized method for adaptive parallelism for collocated transactional multi-threaded processes. Inspired by well-known results from flow/congestion control mechanisms in communication networks, our technique adopts a hill-climbing strategy that was previously unexplored in the context of adaptive parallelism. Amin Mohtasham, João Barreto 0001 |
SPAA | 2 |
| 2013 | Leveraging Web Prefetching Systems with Data DeduplicationabstractThe continued rise in Internet users and the ever-greater complexity of Web content degrade user-perceived latency in the Web. Web caching, prefetching and data deduplication are techniques that are used to try to mitigate this effect. Regardless of the amount of bandwidth available for network traffic, Web prefetching has the potential to consume free bandwidth to its limits. Deduplication explores data redundancies to reduce the amount of data transferred through the network, thereby freeing occupied bandwidth. To the best of our knowledge no previous work has been done which applies these two techniques combined to Web traffic. The motivation of this work is to ask if, by combining these two techniques, it is possible to significantly reduce the amount of bytes transmitted per user request, thereby improving the user-perceived latency in the Web. In the present work, we developed and implemented a system that combines the use of Web prefetching and deduplication techniques. By testing our system using real-world user navigation traces, we show that deduplication is able to substantially reduce the network costs of prefetching (up to 34% savings in transferred bytes), without degrading the latency gains due to prefetching. By adjusting deduplication parameters it is possible to improve the latency relative to when using only prefetching. Pedro Neves 0001, Paulo Ferreira 0001, João Barreto 0001 |
NCA | 3 |
| 2012 | Unifying Thread-Level Speculation and Transactional Memory
João Barreto 0001, Aleksandar Dragojevic, Paulo Ferreira 0001, Ricardo Filipe, Rachid Guerraoui |
Middleware | 1 |
| 2012 | Hash challenges: Stretching the limits of compare-by-hash in distributed data deduplication
João Barreto 0001, Luís Veiga, Paulo Ferreira 0001 |
Inf. Process. Lett. | 1 |
| 2011 | End-to-End Data Deduplication for the Mobile WebabstractThe emergence of affordable mobile devices with rich interfaces and high-bandwidth wireless connectivity has revolutionized the mobile Web. However, such new trends also imply downloading larger data volumes from the Web, with considerable battery and, often, monetary costs that inevitably degrade user experience. The mobile Web calls for end-to-end data deduplication that is able to achieve both high precision and negligible computational cost on the battery-constrained client side. We propose dedupHTTP, a novel deduplication solution that leverages the generic approach of Cache-Based Compaction to achieve the above requirements. Using a full-fledged implementation of dedupHTTP with real workloads from popular Web sites, we obtained savings in traffic consumption of up to 94,5% when comparing to plain HTTP transfer. Ricardo Filipe, João Barreto 0001 |
NCA | 2 |
| 2010 | Meaningful Metrics for Evaluating Eventual Consistency
João Barreto 0001, Paulo Ferreira 0001 |
Euro-Par (2) | 1 |
| 2010 | Leveraging parallel nesting in transactional memoryabstractExploiting the emerging reality of affordable multi-core architectures goes through providing programmers with simple abstractions that would enable them to easily turn their sequential programs into concurrent ones that expose as much parallelism as possible. While transactional memory promises to make concurrent programming easy to a wide programmer community, current implementations either disallow nested transactions to run in parallel or do not scale to arbitrary parallel nesting depths. This is an important obstacle to the central goal of transactional memory, as programmers can only start parallel threads in restricted parts of their code. João Barreto 0001, Aleksandar Dragojevic, Paulo Ferreira 0001, Rachid Guerraoui, Michal Kapalka |
PPoPP | 1 |
| 2009 | Efficient Locally Trackable Deduplication in Replicated Systems
João Barreto 0001, Paulo Ferreira 0001 |
Middleware | 1 |
| 2007 | Exploiting Our Computational Surroundings for Better Mobile CollaborationabstractMobile collaborative environments, being naturally loosely-coupled, call for optimistic replication solutions in order to attain the requirement of decentralized highly available access to data. However, such connectivity assumptions are also a decisive hindrance to the ability of optimistic replication protocols to rapidly guarantee consistency among the set of loosely-coupled replicas. This paper proposes the extension of conventional optimistic replication protocols to exploit the presence of extraneous nodes surrounding the group of replica nodes in an increasingly ubiquitous computational universe. In particular, we show that using such extra nodes as temporary carriers of lightweight consistency meta-data may significantly improve the efficiency of a replicated system; notably, it reduces commitment delay and conflicts, and allows more network-efficient propagation of updates. We support such a statement with experimental results obtained from a simulated environment. João Barreto 0001, Paulo Ferreira 0001, Marc Shapiro 0001 |
MDM | 1 |
| 2007 | Version Vector Weighted Voting protocol: efficient and fault-tolerant commitment for weakly connected replicasabstractAbstract Mobile and other loosely coupled environments call for decentralized optimistic replication protocols that provide highly available access to shared objects, while ensuring eventual consistency. We propose a protocol based on epidemic weighted voting for achieving such a goal with better availability than traditional primary commit approaches. We improve previous epidemic weighted voting solutions by allowing commitment of multiple, happened‐before related updates at a single distributed election round. We demonstrate that our protocol, in contrast to basic weighted voting solutions, achieves similar update commitment ratios to the primary commit alternative. The improvement over basic weighted voting is especially amplified with weaker replica connectivity, as in mobile and other loosely coupled environments. We support such claims by presenting comparison performance results obtained from side‐by‐side execution of reference protocols in a simulated environment. Copyright © 2007 John Wiley & Sons, Ltd. João Barreto 0001, Paulo Ferreira 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2005 | An Efficient and Fault-Tolerant Update Commitment Protocol for Weakly Connected Replicas
João Barreto 0001, Paulo Ferreira 0001 |
Euro-Par | 1 |
| 2005 | Efficient file storage using content-based indexingabstractContent-based indexing [MCM01] is a technique of proven effectiveness for efficient transference of file contents over low bandwidth network links. Departing from this context, the natural step of extending the application of this technique to local file storage has been proposed by a number of storage solutions [CN02, QD02, BF04]. To some extent, all these solutions share a core storage model. File contents are divided into disjoint chunks of data, each of which is individually stored, along with a unique hash of its contents, in a repository of chunks. The actual files are then stored as sequences of possibly shared references to chunks in the repository. João Barreto 0001, Paulo Ferreira 0001 |
SOSP | 1 |