EDBT 2026 Demo / reviewers in the wild / expert
Riza O. Suminto
dblp:164/8050
· DBLP profile ↗
9ranked-venue papers
1as first author
0since 2021 · last 2019
0000-0003-2248-5961ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 3Databases, data management, data science and information retrieval · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Distributed systems · 47% Hardware reliability and fault tolerance · 20% Cloud and datacenter computing · 12% | |
| Software engineering, system software, and programming languages
3 papers |
Operating systems · 54% Software testing · 46% |
Topics — the 7 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software testing › system testing
distributed system testing |
0.5 | 2 | 2019 | FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019 ScaleCheck: A Single-Machine Approach for Discovering Scalability Bugs in Large Distributed Systems · FAST 2019 |
Distributed systems
distributed system testing |
0.4 | 1 | 2019 | FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019 |
Distributed systems
fault tolerance |
0.4 | 1 | 2019 | FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019 |
Operating systems › i/o › i/o subsystem
i/o scheduling |
0.3 | 1 | 2017 | MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface · SOSP 2017 |
Storage systems › i/o architecture › i/o subsystem
i/o stack |
0.3 | 1 | 2017 | MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface · SOSP 2017 |
Cloud and datacenter computing › quality of service
tail latency |
0.3 | 1 | 2017 | MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface · SOSP 2017 |
Parallel and multicore computing › data parallelism
data-parallel applications |
0.1 | 1 | 2017 | MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface · SOSP 2017 |
Methods — techniques the papers use, named apart from their topics
state symmetry · 0.8parallel flips · 0.8event independence · 0.8fast rejection · 0.6SLO prediction · 0.6incident report analysis · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed SystemsabstractWe present a fast and scalable testing approach for datacenter/cloud systems such as Cassandra, Hadoop, Spark, and ZooKeeper. The uniqueness of our approach is in its ability to overcome the path/state-space explosion problem in testing workloads with complex interleavings of messages and faults. We introduce three powerful algorithms: state symmetry, event independence, and parallel flips, which collectively makes our approach on average 16x (up to 78x) faster than other state-of-the-art solutions. We have integrated our techniques with 8 popular datacenter systems, successfully reproduced 12 old bugs, and found 10 new bugs --- all were done without random walks or manual checkpoints. Jeffrey F. Lukman, Huan Ke, Cesar A. Stuardo, Riza O. Suminto, Daniar Heri Kurniawan, Dikaimin Simon, Satria Priambada, Chen Tian 0002, Tanakorn Leesatapornwongsa, Aarti Gupta, Shan Lu 0001, Haryadi S. Gunawi |
EuroSys | 4 |
| 2019 | ScaleCheck: A Single-Machine Approach for Discovering Scalability Bugs in Large Distributed Systems
Cesar A. Stuardo, Tanakorn Leesatapornwongsa, Riza O. Suminto, Huan Ke, Jeffrey F. Lukman, Wei-Chiu Chuang, Shan Lu 0001, Haryadi S. Gunawi |
FAST | 3 |
| 2018 | Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Systems
Haryadi S. Gunawi, Riza O. Suminto, Russell Sears, Casey Golliher, Swaminathan Sundararaman, Tim Emami, Weiguang Sheng, Nematollah Bidokhti, Caitie McCaffrey, Gary Grider, Parks M. Fields, Kevin Harms, Robert B. Ross, Andree Jacobson, Robert Ricci, Kirk Webb, Peter Alvaro, H. Birali Runesha, Mingzhe Hao, Huaicheng Li |
FAST | 2 |
| 2018 | Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production SystemsabstractFail-slow hardware is an under-studied failure mode. We present a study of 114 reports of fail-slow hardware incidents, collected from large-scale cluster deployments in 14 institutions. We show that all hardware types such as disk, SSD, CPU, memory, and network components can exhibit performance faults. We made several important observations such as faults convert from one form to another, the cascading root causes and impacts can be long, and fail-slow faults can have varying symptoms. From this study, we make suggestions to vendors, operators, and systems designers. Haryadi S. Gunawi, Riza O. Suminto, Russell Sears, Casey Golliher, Swaminathan Sundararaman, Tim Emami, Weiguang Sheng, Nematollah Bidokhti, Caitie McCaffrey, Deepthi Srinivasan, Biswaranjan Panda, Andrew Baptist, Gary Grider, Parks M. Fields, Kevin Harms, Robert B. Ross, Andree Jacobson, Robert Ricci, Kirk Webb, Peter Alvaro, H. Birali Runesha, Mingzhe Hao, Huaicheng Li |
ACM Trans. Storage | 2 |
| 2017 | PBSE: a robust path-based speculative execution for degraded-network tail tolerance in data-parallel frameworksabstractWe reveal loopholes of Speculative Execution (SE) implementations under a unique fault model: node-level network throughput degradation. This problem appears in many data-parallel frameworks such as Hadoop MapReduce and Spark. To address this, we present PBSE, a robust, path-based speculative execution that employs three key ingredients: path progress, path diversity, and path-straggler detection and speculation. We show how PBSE is superior to other approaches such as cloning and aggressive speculation under the aforementioned fault model. PBSE is a general solution, applicable to many data-parallel frameworks such as Hadoop/HDFS+QFS, Spark and Flume. Riza O. Suminto, Cesar A. Stuardo, Alexandra Clark, Huan Ke, Tanakorn Leesatapornwongsa, Daniar Heri Kurniawan, Vincentius Martin, Maheswara Rao G. Uma, Haryadi S. Gunawi |
SoCC | 1 |
| 2017 | Scalability Bugs: When 100-Node Testing is Not EnoughabstractWe highlight the problem of scalability bugs, a new class of bugs that appear in "cloud-scale" distributed systems. Scalability bugs are latent bugs that are cluster-scale dependent, whose symptoms typically surface in large-scale deployments, but not in small or medium-scale deployments. The standard practice to test large distributed systems is to deploy them on a large number of machines ("real-scale testing"), which is difficult and expensive. New methods are needed to reduce developers' burdens in finding, reproducing, and debugging scalability bugs. We propose "scale check," an approach that helps developers find and replay scalability bugs at real scales, but do so only on one machine and still achieve a high accuracy (i.e., similar observed behaviors as if the nodes are deployed in real-scale testing). Tanakorn Leesatapornwongsa, Cesar A. Stuardo, Riza O. Suminto, Huan Ke, Jeffrey F. Lukman, Haryadi S. Gunawi |
HotOS | 3 |
| 2017 | Rivulet: a fault-tolerant platform for smart-home applicationsabstractRivulet is a fault-tolerant distributed platform for running smart-home applications; it can tolerate failures typical for a home environment (e.g., link losses, network partitions, sensor failures, and device crashes). In contrast to existing cloud-centric solutions, which rely exclusively on a home gateway device, Rivulet leverages redundant smart consumer appliances (e.g., TVs, Refrigerators) to spread sensing and actuation across devices local to the home, and avoids making the Smart-Home Hub a single point of failure. Rivulet ensures event delivery in the presence of link loss, network partitions and other failures in the home, to enable applications with reliable sensing in the case of sensor failures, and event processing in the presence of device crashes. In this paper, we present the design and implementation of Rivulet, and evaluate its effective handling of failures in a smart home. Masoud Saeida Ardekani, Rayman Preet Singh, Nitin Agrawal 0001, Douglas B. Terry, Riza O. Suminto |
Middleware | 5 |
| 2017 | MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS InterfaceabstractMittOS provides operating system support to cut millisecond-level tail latencies for data-parallel applications. In MittOS, we advocate a new principle that operating system should quickly reject IOs that cannot be promptly served. To achieve this, MittOS exposes a fast rejecting SLO-aware interface wherein applications can provide their SLOs (e.g., IO deadlines). If MittOS predicts that the IO SLOs cannot be met, MittOS will promptly return EBUSY signal, allowing the application to failover (retry) to another less-busy node without waiting. We build MittOS within the storage stack (disk, SSD, and OS cache managements), but the principle is extensible to CPU and runtime memory managements as well. MittOS' no-wait approach helps reduce IO completion time up to 35% compared to wait-then-speculate approaches. Mingzhe Hao, Huaicheng Li, Michael Hao Tong, Chrisma Pakha, Riza O. Suminto, Cesar A. Stuardo, Andrew A. Chien, Haryadi S. Gunawi |
SOSP | 5 |
| 2016 | Why Does the Cloud Stop Computing? Lessons from Hundreds of Service OutagesabstractWe conducted a cloud outage study (COS) of 32 popular Internet services. We analyzed 1247 headline news and public post-mortem reports that detail 597 unplanned outages that occurred within a 7-year span from 2009 to 2015. We analyzed outage duration, root causes, impacts, and fix procedures. This study reveals the broader availability landscape of modern cloud services and provides answers to why outages still take place even with pervasive redundancies. Haryadi S. Gunawi, Mingzhe Hao, Riza O. Suminto, Agung Laksono, Anang D. Satria, Jeffry Adityatama, Kurnia J. Eliazar |
SoCC | 3 |