Jeffrey F. Lukman

dblp:154/0977 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 4Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Distributed systems · 89% Cloud and datacenter computing · 11%
Software engineering, system software, and programming languages
5 papers
Software testing · 35% Concurrent programming · 18% Program analysis · 18%

Topics — the 10 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › concurrency
distributed concurrency bugs
0.522017
DCatch: Automatically Detecting Distributed Concurrency Bugs in Cloud Systems · ASPLOS 2017
TaxDC: A Taxonomy of Non-Deterministic Concurrency Bugs in Datacenter Distributed Systems · ASPLOS 2016
Software testing › system testing
distributed system testing
0.522019
FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019
ScaleCheck: A Single-Machine Approach for Discovering Scalability Bugs in Large Distributed Systems · FAST 2019
Distributed systems
bug detection
0.522017
DCatch: Automatically Detecting Distributed Concurrency Bugs in Cloud Systems · ASPLOS 2017
SAMC: Semantic-Aware Model Checking for Fast Discovery of Deep Bugs in Cloud Systems · OSDI 2014
Distributed systems
distributed system testing
0.412019
FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019
Distributed systems
fault tolerance
0.412019
FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019
Concurrent programming
concurrency bug detection
0.312017
DCatch: Automatically Detecting Distributed Concurrency Bugs in Cloud Systems · ASPLOS 2017
Program analysis › dynamic analysis
happens-before analysis
0.312017
DCatch: Automatically Detecting Distributed Concurrency Bugs in Cloud Systems · ASPLOS 2017
Empirical software engineering › mining software repositories › bug report analysis
bug characterization
0.212016
TaxDC: A Taxonomy of Non-Deterministic Concurrency Bugs in Datacenter Distributed Systems · ASPLOS 2016
Program verification
model checking
0.212014
SAMC: Semantic-Aware Model Checking for Fast Discovery of Deep Bugs in Cloud Systems · OSDI 2014
Cloud and datacenter computing › cloud deployment model
distributed cloud
0.112017
DCatch: Automatically Detecting Distributed Concurrency Bugs in Cloud Systems · ASPLOS 2017

Methods — techniques the papers use, named apart from their topics

state symmetry · 0.8parallel flips · 0.8event independence · 0.8trace analysis · 0.6runtime tracing · 0.6happens-before rules · 0.6model checking · 0.4
YearPublicationVenuePosition
2019 FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems
abstract
We present a fast and scalable testing approach for datacenter/cloud systems such as Cassandra, Hadoop, Spark, and ZooKeeper. The uniqueness of our approach is in its ability to overcome the path/state-space explosion problem in testing workloads with complex interleavings of messages and faults. We introduce three powerful algorithms: state symmetry, event independence, and parallel flips, which collectively makes our approach on average 16x (up to 78x) faster than other state-of-the-art solutions. We have integrated our techniques with 8 popular datacenter systems, successfully reproduced 12 old bugs, and found 10 new bugs --- all were done without random walks or manual checkpoints.
Jeffrey F. Lukman, Huan Ke, Cesar A. Stuardo, Riza O. Suminto, Daniar Heri Kurniawan, Dikaimin Simon, Satria Priambada, Chen Tian 0002, Tanakorn Leesatapornwongsa, Aarti Gupta, Shan Lu 0001, Haryadi S. Gunawi
EuroSys1
2019 ScaleCheck: A Single-Machine Approach for Discovering Scalability Bugs in Large Distributed Systems
Cesar A. Stuardo, Tanakorn Leesatapornwongsa, Riza O. Suminto, Huan Ke, Jeffrey F. Lukman, Wei-Chiu Chuang, Shan Lu 0001, Haryadi S. Gunawi
FAST5
2017 DCatch: Automatically Detecting Distributed Concurrency Bugs in Cloud Systems
abstract
In big data and cloud computing era, reliability of distributed systems is extremely important. Unfortunately, distributed concurrency bugs, referred to as DCbugs, widely exist. They hide in the large state space of distributed cloud systems and manifest non-deterministically depending on the timing of distributed computation and communication. Effective techniques to detect DCbugs are desired. This paper presents a pilot solution, DCatch, in the world of DCbug detection. DCatch predicts DCbugs by analyzing correct execution of distributed systems. To build DCatch, we design a set of happens-before rules that model a wide variety of communication and concurrency mechanisms in real-world distributed cloud systems. We then build runtime tracing and trace analysis tools to effectively identify concurrent conflicting memory accesses in these systems. Finally, we design tools to help prune false positives and trigger DCbugs. We have evaluated DCatch on four representative open-source distributed cloud systems, Cassandra, Hadoop MapReduce, HBase, and ZooKeeper. By monitoring correct execution of seven workloads on these systems, DCatch reports 32 DCbugs, with 20 of them being truly harmful.
Guangpu Li, Jeffrey F. Lukman, Shan Lu 0001, Haryadi S. Gunawi, Chen Tian 0002
ASPLOS3
2017 Scalability Bugs: When 100-Node Testing is Not Enough
abstract
We highlight the problem of scalability bugs, a new class of bugs that appear in "cloud-scale" distributed systems. Scalability bugs are latent bugs that are cluster-scale dependent, whose symptoms typically surface in large-scale deployments, but not in small or medium-scale deployments. The standard practice to test large distributed systems is to deploy them on a large number of machines ("real-scale testing"), which is difficult and expensive. New methods are needed to reduce developers' burdens in finding, reproducing, and debugging scalability bugs. We propose "scale check," an approach that helps developers find and replay scalability bugs at real scales, but do so only on one machine and still achieve a high accuracy (i.e., similar observed behaviors as if the nodes are deployed in real-scale testing).
Tanakorn Leesatapornwongsa, Cesar A. Stuardo, Riza O. Suminto, Huan Ke, Jeffrey F. Lukman, Haryadi S. Gunawi
HotOS5
2016 TaxDC: A Taxonomy of Non-Deterministic Concurrency Bugs in Datacenter Distributed Systems
abstract
We present TaxDC, the largest and most comprehensive taxonomy of non-deterministic concurrency bugs in distributed systems. We study 104 distributed concurrency (DC) bugs from four widely-deployed cloud-scale datacenter distributed systems, Cassandra, Hadoop MapReduce, HBase and ZooKeeper. We study DC-bug characteristics along several axes of analysis such as the triggering timing condition and input preconditions, error and failure symptoms, and fix strategies, collectively stored as 2,083 classification labels in TaxDC database. We discuss how our study can open up many new research directions in combating DC bugs.
Tanakorn Leesatapornwongsa, Jeffrey F. Lukman, Shan Lu 0001, Haryadi S. Gunawi
ASPLOS2
2014 What Bugs Live in the Cloud? A Study of 3000+ Issues in Cloud Systems
abstract
We conduct a comprehensive study of development and deployment issues of six popular and important cloud systems (Hadoop MapReduce, HDFS, HBase, Cassandra, ZooKeeper and Flume). From the bug repositories, we review in total 21,399 submitted issues within a three-year period (2011-2014). Among these issues, we perform a deep analysis of 3655 "vital" issues (i.e., real issues affecting deployments) with a set of detailed classifications. We name the product of our one-year study Cloud Bug Study database (CbsDB) [9], with which we derive numerous interesting insights unique to cloud systems. To the best of our knowledge, our work is the largest bug study for cloud systems to date.
Haryadi S. Gunawi, Mingzhe Hao, Tanakorn Leesatapornwongsa, Tiratat Patana-anake, Thanh Do, Jeffry Adityatama, Kurnia J. Eliazar, Agung Laksono, Jeffrey F. Lukman, Vincentius Martin, Anang D. Satria
SoCC9
2014 SAMC: Semantic-Aware Model Checking for Fast Discovery of Deep Bugs in Cloud Systems
Tanakorn Leesatapornwongsa, Mingzhe Hao, Pallavi Joshi, Jeffrey F. Lukman, Haryadi S. Gunawi
OSDI4