Marius Poke

dblp:136/7975 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-authorSecurity and privacy · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 82% Performance modeling and evaluation · 8% Parallel and multicore computing · 8%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
consensus
0.522017
AllConcur: Leaderless Concurrent Atomic Broadcast · HPDC 2017
DARE: High-Performance State Machine Replication on RDMA Networks · HPDC 2015
Distributed systems › group communication
atomic broadcast
0.312017
AllConcur: Leaderless Concurrent Atomic Broadcast · HPDC 2017
Distributed systems
fault tolerance
0.312017
AllConcur: Leaderless Concurrent Atomic Broadcast · HPDC 2017
Distributed systems › consensus
fault-tolerant consensus
0.212015
DARE: High-Performance State Machine Replication on RDMA Networks · HPDC 2015
Distributed systems
replication
0.212015
DARE: High-Performance State Machine Replication on RDMA Networks · HPDC 2015
Distributed systems › replication
state machine replication
0.212015
DARE: High-Performance State Machine Replication on RDMA Networks · HPDC 2015
Performance modeling and evaluation › performance model construction
automated performance modeling
0.212013
Using automated performance modeling to find scalability bugs in complex codes · SC 2013
Parallel and multicore computing › parallel computing
parallel applications
0.212013
Using automated performance modeling to find scalability bugs in complex codes · SC 2013
Distributed systems › peer-to-peer systems
overlay networks
0.112017
AllConcur: Leaderless Concurrent Atomic Broadcast · HPDC 2017
Storage systems
key-value storage
0.112015
DARE: High-Performance State Machine Replication on RDMA Networks · HPDC 2015

Methods — techniques the papers use, named apart from their topics

infiniband verbs · 0.3early termination · 0.3TCP sockets · 0.3remote direct memory access · 0.2protocol design · 0.2
YearPublicationVenuePosition
2019 A Dual Digraph Approach for Leaderless Atomic Broadcast
abstract
Many distributed systems work on a common shared state; in such systems, distributed agreement is necessary for consistency. With an increasing number of servers, these systems become more susceptible to single-server failures, increasing the relevance of fault-tolerance. Atomic broadcast enables fault-tolerant distributed agreement, yet it is costly to solve. Most practical algorithms entail linear work per broadcast message. AllConcur - a leaderless approach - reduces the work, by connecting the servers via a sparse resilient overlay network; yet, this resiliency entails redundancy, limiting the reduction of work. In this paper, we propose AllConcur+, an atomic broadcast algorithm that lifts this limitation: During intervals with no failures, it achieves minimal work by using a redundancy-free overlay network. When failures do occur, it automatically recovers by switching to a resilient overlay network. In our performance evaluation of non-failure scenarios, AllConcur+ achieves comparable throughput to AllGather - a non-fault-tolerant distributed agreement algorithm - and outperforms AllConcur, LCR and Libpaxos both in terms of throughput and latency. Furthermore, our evaluation of failure scenarios shows that AllConcur+'s expected performance is robust with regard to occasional failures. Thus, for realistic use cases, leveraging redundancy-free distributed agreement during intervals with no failures improves performance significantly.
Marius Poke, Colin W. Glass
SRDS1
2017 AllConcur: Leaderless Concurrent Atomic Broadcast
abstract
Many distributed systems require coordination between the components involved. With the steady growth of such systems, the probability of failures increases, which necessitates scalable fault-tolerant agreement protocols. The most common practical agreement protocol, for such scenarios, is leader-based atomic broadcast. In this work, we propose AllConcur, a distributed system that provides agreement through a leaderless concurrent atomic broadcast algorithm, thus, not suffering from the bottleneck of a central coordinator. In AllConcur, all components exchange messages concurrently through a logical overlay network that employs early termination to minimize the agreement latency. Our implementation of AllConcur supports standard sockets-based TCP as well as high-performance InfiniBand Verbs communications. AllConcur can handle up to 135 million requests per second and achieves 17x higher throughput than today's standard leader-based protocols, such as Libpaxos. Thus, AllConcur is highly competitive with regard to existing solutions and, due to its decentralized approach, enables hitherto unattainable system designs in a variety of fields.
Marius Poke, Torsten Hoefler, Colin W. Glass
HPDC1
2015 DARE: High-Performance State Machine Replication on RDMA Networks
abstract
The increasing amount of data that needs to be collected and analyzed requires large-scale datacenter architectures that are naturally more susceptible to faults of single components. One way to offer consistent services on such unreliable systems are replicated state machines (RSMs). Yet, traditional RSM protocols cannot deliver the needed latency and request rates for future large-scale systems. In this paper, we propose a new set of protocols based on Remote Direct Memory Access (RDMA) primitives. To asses these mechanisms, we use a strongly consistent key-value store; the evaluation shows that our simple protocols improve RSM performance by more than an order of magnitude. Furthermore, we show that RDMA introduces various new options, such as log access management. Our protocols enable operators to fully utilize the new capabilities of the quickly growing number of RDMA-capable datacenter networks.
Marius Poke, Torsten Hoefler
HPDC1
2013 Using automated performance modeling to find scalability bugs in complex codes
abstract
Many parallel applications suffer from latent performance limitations that may prevent them from scaling to larger machine sizes. Often, such scalability bugs manifest themselves only when an attempt to scale the code is actually being made---a point where remediation can be difficult. However, creating analytical performance models that would allow such issues to be pinpointed earlier is so laborious that application developers attempt it at most for a few selected kernels, running the risk of missing harmful bottlenecks. In this paper, we show how both coverage and speed of this scalability analysis can be substantially improved. Generating an empirical performance model automatically for each part of a parallel program, we can easily identify those parts that will reduce performance at larger core counts. Using a climate simulation as an example, we demonstrate that scalability bugs are not confined to those routines usually chosen as kernels.
Alexandru Calotoiu, Torsten Hoefler, Marius Poke, Felix Wolf 0001
SC3