Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Riza O. Suminto

dblp:164/8050 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
0since 2021 · last 2019
0000-0003-2248-5961ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 3Databases, data management, data science and information retrieval · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Distributed systems · 47% Hardware reliability and fault tolerance · 20% Cloud and datacenter computing · 12%
Software engineering, system software, and programming languages
3 papers
Operating systems · 54% Software testing · 46%

Topics — the 7 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software testing › system testing
distributed system testing
0.522019
FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019
ScaleCheck: A Single-Machine Approach for Discovering Scalability Bugs in Large Distributed Systems · FAST 2019
Distributed systems
distributed system testing
0.412019
FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019
Distributed systems
fault tolerance
0.412019
FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019
Operating systems › i/o › i/o subsystem
i/o scheduling
0.312017
MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface · SOSP 2017
Storage systems › i/o architecture › i/o subsystem
i/o stack
0.312017
MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface · SOSP 2017
Cloud and datacenter computing › quality of service
tail latency
0.312017
MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface · SOSP 2017
Parallel and multicore computing › data parallelism
data-parallel applications
0.112017
MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface · SOSP 2017

Methods — techniques the papers use, named apart from their topics

state symmetry · 0.8parallel flips · 0.8event independence · 0.8fast rejection · 0.6SLO prediction · 0.6incident report analysis · 0.3
YearPublicationVenuePosition
2019 FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems
abstract
We present a fast and scalable testing approach for datacenter/cloud systems such as Cassandra, Hadoop, Spark, and ZooKeeper. The uniqueness of our approach is in its ability to overcome the path/state-space explosion problem in testing workloads with complex interleavings of messages and faults. We introduce three powerful algorithms: state symmetry, event independence, and parallel flips, which collectively makes our approach on average 16x (up to 78x) faster than other state-of-the-art solutions. We have integrated our techniques with 8 popular datacenter systems, successfully reproduced 12 old bugs, and found 10 new bugs --- all were done without random walks or manual checkpoints.
Jeffrey F. Lukman, Huan Ke, Cesar A. Stuardo, Riza O. Suminto, Daniar Heri Kurniawan, Dikaimin Simon, Satria Priambada, Chen Tian 0002, Tanakorn Leesatapornwongsa, Aarti Gupta, Shan Lu 0001, Haryadi S. Gunawi
EuroSys4
2019 ScaleCheck: A Single-Machine Approach for Discovering Scalability Bugs in Large Distributed Systems
Cesar A. Stuardo, Tanakorn Leesatapornwongsa, Riza O. Suminto, Huan Ke, Jeffrey F. Lukman, Wei-Chiu Chuang, Shan Lu 0001, Haryadi S. Gunawi
FAST3
2018 Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Systems
Haryadi S. Gunawi, Riza O. Suminto, Russell Sears, Casey Golliher, Swaminathan Sundararaman, Tim Emami, Weiguang Sheng, Nematollah Bidokhti, Caitie McCaffrey, Gary Grider, Parks M. Fields, Kevin Harms, Robert B. Ross, Andree Jacobson, Robert Ricci, Kirk Webb, Peter Alvaro, H. Birali Runesha, Mingzhe Hao, Huaicheng Li
FAST2
2018 Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Systems
abstract
Fail-slow hardware is an under-studied failure mode. We present a study of 114 reports of fail-slow hardware incidents, collected from large-scale cluster deployments in 14 institutions. We show that all hardware types such as disk, SSD, CPU, memory, and network components can exhibit performance faults. We made several important observations such as faults convert from one form to another, the cascading root causes and impacts can be long, and fail-slow faults can have varying symptoms. From this study, we make suggestions to vendors, operators, and systems designers.
Haryadi S. Gunawi, Riza O. Suminto, Russell Sears, Casey Golliher, Swaminathan Sundararaman, Tim Emami, Weiguang Sheng, Nematollah Bidokhti, Caitie McCaffrey, Deepthi Srinivasan, Biswaranjan Panda, Andrew Baptist, Gary Grider, Parks M. Fields, Kevin Harms, Robert B. Ross, Andree Jacobson, Robert Ricci, Kirk Webb, Peter Alvaro, H. Birali Runesha, Mingzhe Hao, Huaicheng Li
ACM Trans. Storage2
2017 PBSE: a robust path-based speculative execution for degraded-network tail tolerance in data-parallel frameworks
abstract
We reveal loopholes of Speculative Execution (SE) implementations under a unique fault model: node-level network throughput degradation. This problem appears in many data-parallel frameworks such as Hadoop MapReduce and Spark. To address this, we present PBSE, a robust, path-based speculative execution that employs three key ingredients: path progress, path diversity, and path-straggler detection and speculation. We show how PBSE is superior to other approaches such as cloning and aggressive speculation under the aforementioned fault model. PBSE is a general solution, applicable to many data-parallel frameworks such as Hadoop/HDFS+QFS, Spark and Flume.
Riza O. Suminto, Cesar A. Stuardo, Alexandra Clark, Huan Ke, Tanakorn Leesatapornwongsa, Daniar Heri Kurniawan, Vincentius Martin, Maheswara Rao G. Uma, Haryadi S. Gunawi
SoCC1
2017 Scalability Bugs: When 100-Node Testing is Not Enough
abstract
We highlight the problem of scalability bugs, a new class of bugs that appear in "cloud-scale" distributed systems. Scalability bugs are latent bugs that are cluster-scale dependent, whose symptoms typically surface in large-scale deployments, but not in small or medium-scale deployments. The standard practice to test large distributed systems is to deploy them on a large number of machines ("real-scale testing"), which is difficult and expensive. New methods are needed to reduce developers' burdens in finding, reproducing, and debugging scalability bugs. We propose "scale check," an approach that helps developers find and replay scalability bugs at real scales, but do so only on one machine and still achieve a high accuracy (i.e., similar observed behaviors as if the nodes are deployed in real-scale testing).
Tanakorn Leesatapornwongsa, Cesar A. Stuardo, Riza O. Suminto, Huan Ke, Jeffrey F. Lukman, Haryadi S. Gunawi
HotOS3
2017 Rivulet: a fault-tolerant platform for smart-home applications
abstract
Rivulet is a fault-tolerant distributed platform for running smart-home applications; it can tolerate failures typical for a home environment (e.g., link losses, network partitions, sensor failures, and device crashes). In contrast to existing cloud-centric solutions, which rely exclusively on a home gateway device, Rivulet leverages redundant smart consumer appliances (e.g., TVs, Refrigerators) to spread sensing and actuation across devices local to the home, and avoids making the Smart-Home Hub a single point of failure. Rivulet ensures event delivery in the presence of link loss, network partitions and other failures in the home, to enable applications with reliable sensing in the case of sensor failures, and event processing in the presence of device crashes. In this paper, we present the design and implementation of Rivulet, and evaluate its effective handling of failures in a smart home.
Masoud Saeida Ardekani, Rayman Preet Singh, Nitin Agrawal 0001, Douglas B. Terry, Riza O. Suminto
Middleware5
2017 MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface
abstract
MittOS provides operating system support to cut millisecond-level tail latencies for data-parallel applications. In MittOS, we advocate a new principle that operating system should quickly reject IOs that cannot be promptly served. To achieve this, MittOS exposes a fast rejecting SLO-aware interface wherein applications can provide their SLOs (e.g., IO deadlines). If MittOS predicts that the IO SLOs cannot be met, MittOS will promptly return EBUSY signal, allowing the application to failover (retry) to another less-busy node without waiting. We build MittOS within the storage stack (disk, SSD, and OS cache managements), but the principle is extensible to CPU and runtime memory managements as well. MittOS' no-wait approach helps reduce IO completion time up to 35% compared to wait-then-speculate approaches.
Mingzhe Hao, Huaicheng Li, Michael Hao Tong, Chrisma Pakha, Riza O. Suminto, Cesar A. Stuardo, Andrew A. Chien, Haryadi S. Gunawi
SOSP5
2016 Why Does the Cloud Stop Computing? Lessons from Hundreds of Service Outages
abstract
We conducted a cloud outage study (COS) of 32 popular Internet services. We analyzed 1247 headline news and public post-mortem reports that detail 597 unplanned outages that occurred within a 7-year span from 2009 to 2015. We analyzed outage duration, root causes, impacts, and fix procedures. This study reveals the broader availability landscape of modern cloud services and provides answers to why outages still take place even with pervasive redundancies.
Haryadi S. Gunawi, Mingzhe Hao, Riza O. Suminto, Agung Laksono, Anang D. Satria, Jeffry Adityatama, Kurnia J. Eliazar
SoCC3