Diogo Behrens

dblp:131/4370 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-6463-3005ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Enabling Efficient Mobile Tracing with BTrace
abstract
With the growing complexity of smartphone systems, effective tracing becomes vital for enhancing their stability and optimizing the user experience. Unfortunately, existing tracing tools are inefficient in smartphone scenarios. Their distributed designs (with either per-core or per-thread buffers) prioritize performance but lead to missing crucial clues with high probability. While these problems can be overlooked in previous scenarios (e.g., servers), they drastically limit the usefulness of tracing on smartphones.
Arnau Casadevall-Saiz, Diogo Behrens, Ming Fu, Ning Jia 0004, Hermann Härtig, Haibo Chen 0001
ASPLOS (2)5
2023 AtoMig: Automatically Migrating Millions Lines of Code from TSO to WMM
abstract
CPUs with weak memory-consistency models (WMMs), such as Arm and RISC-V, are rapidly increasing their market share. Porting legacy x86 applications to such CPUs requires introducing extra synchronization to prevent WMM-related concurrency bugs---a task often left to human experts.
Martin Beck, Koustubha Bhat, Lazar Stricevic, Diogo Behrens, Ming Fu, Viktor Vafeiadis, Haibo Chen 0001, Hermann Härtig
ASPLOS (2)5
2023 BWoS: Formally Verified Block-based Work Stealing for Parallel Processing
Bohdan Trach, Ming Fu, Diogo Behrens, Jonathan Schwender, Jitang Lei, Viktor Vafeiadis, Hermann Härtig, Haibo Chen 0001
OSDI4
2022 BBQ: A Block-based Bounded Queue for Exchanging Data and Profiling
Diogo Behrens, Ming Fu, Lilith Oberhauser, Jonas Oberhauser, Jitang Lei, Hermann Härtig, Haibo Chen 0001
USENIX ATC2
2021 VSync: push-button verification and optimization for synchronization primitives on weak memory models
abstract
Implementing highly efficient and correct synchronization primitives on modern Weak Memory Model (WMM) architectures, such as ARM and RISC-V, is very difficult even for human experts. We introduce VSync, a framework to assist in optimizing and verifying synchronization primitives on WMM architectures. VSync automatically detects missing and overly-constrained barriers, while ensuring essential safety and liveness properties. VSync relies on two novel techniques: 1) Adaptive Linear Relaxation (ALR), which utilizes barrier monotonicity and speculation to quickly find a correct maximally-relaxed barrier combination; and 2) Await Model Checking (AMC), which for the first time makes it possible to check termination of await loops on WMMs.
Jonas Oberhauser, Rafael Lourenco de Lima Chehab, Diogo Behrens, Ming Fu, Antonio Paolillo, Lilith Oberhauser, Koustubha Bhat, Yuzhong Wen, Haibo Chen 0001, Viktor Vafeiadis
ASPLOS3
2021 CLoF: A Compositional Lock Framework for Multi-level NUMA Systems
abstract
Efficient locking mechanisms are extremely important to support large-scale concurrency and exploit the performance promises of many-core servers. Implementing an efficient, generic, and correct lock is very challenging due to the differences between various NUMA architectures. The performance impact of architectural/NUMA hierarchy differences between x86 and Armv8 are not yet fully explored, leading to unexpected performance when simply porting NUMA-aware locks from x86 to Armv8. Moreover, due to the Armv8 Weak Memory Model (WMM), correctly implementing complicated NUMA-aware locks is very difficult.
Rafael Lourenco de Lima Chehab, Antonio Paolillo, Diogo Behrens, Ming Fu, Hermann Härtig, Haibo Chen 0001
SOSP3
2015 Scalable Error Isolation for Distributed Systems
Diogo Behrens, Marco Serafini, Flavio Paiva Junqueira, Sergei Arnautov, Christof Fetzer
NSDI1
2014 HardPaxos: Replication Hardened against Hardware Errors
abstract
State Machine Replication (SMR) is a common technique to make services fault-tolerant. Practical SMR systems tolerate process crashes, but no hardware errors such as bit flips. Still, hardware errors can cause major service outages, and their rate is expected to increase in the future. Current approaches either incur a high overhead by hardening large parts of the system in software, or increase the cost of ownership by introducing additional hardware components. This work presents HardPaxos, an atomic broadcast algorithm for SMR that enables services to tolerate hardware errors, while incurring little performance and state overhead. HardPaxos requires no additional hardware and has only a small part of its functionality hardened using a combination of AN-encoding and duplicated execution. Our evaluation shows a throughput overhead of at most 5% for typical payload sizes. Moreover, fault injection experiments show that our hardening decreases the number of undetected errors from 15% to 0.02%.
Diogo Behrens, Dmitrii Kuvaiskii, Christof Fetzer
SRDS1
2013 Improving Wide-Area Replication Performance through Informed Leader Election and Overlay Construction
abstract
Replication is an important building block to achieve high availability in the presence of failures. Until recently, wide-area replication with strong consistency guarantees was regarded as impractical due to performance constraints. We investigate how informed leader election combined with a network overlay can improve the performance of distributed consensus, which is at the heart of every replicated data store. Leader election and overlay construction are particularly relevant when replicating data at global scale where network links exhibit diverse performance characteristics. We propose to incorporate knowledge about the link quality and network overlay topology into the leader election algorithm. In particular, we show how optimizing only for a quorum, instead of all replicas, we can increase replication throughput or decrease the request latency. Our measurements show a throughput increase of 1.5x when optimizing for throughput of all replicas and a 3x improvement when the throughput is optimized only for a quorum.
Syed Kewaan Ejaz, Diogo Behrens, Thomas Knauth, Christof Fetzer
IEEE CLOUD2