Kirk Rodrigues

dblp:188/9963 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2024 μSlope: High Compression and Fast Search on Semi-Structured Logs
Devin Gibson, Kirk Rodrigues, Yu Luo 0006, Kaibo Wang, Yupeng Fu, Ding Yuan 0004
OSDI3
2022 Hubble: Performance Debugging with In-Production, Just-In-Time Method Tracing on Android
Yu Luo 0006, Kirk Rodrigues, Cuiqin Li, Lijin Jiang, David Lion, Ding Yuan 0004
OSDI2
2021 CLP: Efficient and Scalable Search on Compressed Text Logs
Kirk Rodrigues, Yu Luo 0006, Ding Yuan 0004
OSDI1
2021 Understanding and Detecting Software Upgrade Failures in Distributed Systems
abstract
Upgrade is one of the most disruptive yet unavoidable maintenance tasks that undermine the availability of distributed systems. Any failure during an upgrade is catastrophic, as it further extends the service disruption caused by the upgrade. The increasing adoption of continuous deployment further increases the frequency and burden of the upgrade task. In practice, upgrade failures have caused many of today's high-profile cloud outages. Unfortunately, there has been little understanding of their characteristics.
Yongle Zhang 0007, Zhuqi Jin, Utsav Sethi, Kirk Rodrigues, Shan Lu 0001, Ding Yuan 0004
SOSP5
2019 An analysis of performance evolution of Linux's core operations
abstract
This paper presents an analysis of how Linux's performance has evolved over the past seven years. Unlike recent works that focus on OS performance in terms of scalability or service of a particular workload, this study goes back to basics: the latency of core kernel operations (e.g., system calls, context switching, etc.). To our surprise, the study shows that the performance of many core operations has worsened or fluctuated significantly over the years. For example, the select system call is 100% slower than it was just two years ago. An in-depth analysis shows that over the past seven years, core kernel subsystems have been forced to accommodate an increasing number of security enhancements and new features. These additions steadily add overhead to core kernel operations but also frequently introduce extreme slowdowns of more than 100%. In addition, simple misconfigurations have also severely impacted kernel performance. Overall, we find most of the slowdowns can be attributed to 11 changes.
Xiang Ren 0003, Kirk Rodrigues, Luyuan Chen, Juan Camilo Vega, Michael Stumm, Ding Yuan 0004
SOSP2
2019 The inflection point hypothesis: a principled debugging approach for locating the root cause of a failure
abstract
The end goal of failure diagnosis is to locate the root cause. Prior root cause localization approaches almost all rely on statistical analysis. This paper proposes taking a different approach based on the observation that if we model an execution as a totally ordered sequence of instructions, then the root cause can be identified by the first instruction where the failure execution deviates from the non-failure execution that has the longest instruction sequence prefix in common with that of the failure execution. Thus, root cause analysis is transformed into a principled search problem to identify the non-failure execution with the longest common prefix. We present Kairux, a tool that does just that. It is, in most cases, capable of pinpointing the root cause of a failure in a distributed system, in a fully automated way. Kairux uses tests from the system's rich unit test suite as building blocks to construct the non-failure execution that has the longest common prefix with the failure execution in order to locate the root cause. By evaluating Kairux on some of the most complex, real-world failures from HBase, HDFS, and ZooKeeper, we show that Kairux can accurately pinpoint each failure's respective root cause.
Yongle Zhang 0007, Kirk Rodrigues, Yu Luo 0006, Michael Stumm, Ding Yuan 0004
SOSP2
2017 The Game of Twenty Questions: Do You Know Where to Log?
abstract
A production system's printed logs are often the only source of runtime information available for postmortem debugging, performance analysis and profiling, security auditing, and user behavior analytics. Therefore, the quality of this data is critically important. Recent work has attempted to enhance log quality by recording additional variable values, but logging statement placement, i.e., where to place a logging statement, which is the most challenging and fundamental problem for improving log quality, has not been adequately addressed so far. This position paper proposes we automate the placement of logging statements by measuring how much uncertainty, i.e., the expected number of possible execution code paths taken by the software, can be removed by adding a logging statement to a basic block. Guided by ideas from information theory, we describe a simple approach that automates logging statement placement. Preliminary results suggest that our algorithm can effectively cover, and further improve, the existing logging statement placements selected by developers. It can compute an optimal logging statement placement that disambiguates the entire function call path with only 0.218% of slowdown.
Xu Zhao 0004, Kirk Rodrigues, Yu Luo 0006, Michael Stumm, Ding Yuan 0004, Yuanyuan Zhou 0001
HotOS2
2017 Log20: Fully Automated Optimal Placement of Log Printing Statements under Specified Overhead Threshold
abstract
When systems fail in production environments, log data is often the only information available to programmers for postmortem debugging. Consequently, programmers' decision on where to place a log printing statement is of crucial importance, as it directly affects how effective and efficient postmortem debugging can be. This paper presents Log20, a tool that determines a near optimal placement of log printing statements under the constraint of adding less than a specified amount of performance overhead. Log20 does this in an automated way without any human involvement. Guided by information theory, the core of our algorithm measures how effective each log printing statement is in disambiguating code paths. To do so, it uses the frequencies of different execution paths that are collected from a production environment by a low-overhead tracing library. We evaluated Log20 on HDFS, HBase, Cassandra, and ZooKeeper, and observed that Log20 is substantially more efficient in code path disambiguation compared to the developers' manually placed log printing statements. Log20 can also output a curve showing the trade-off between the informativeness of the logs and the performance slowdown, so that a developer can choose the right balance.
Xu Zhao 0004, Kirk Rodrigues, Yu Luo 0006, Michael Stumm, Ding Yuan 0004, Yuanyuan Zhou 0001
SOSP2
2016 Non-Intrusive Performance Profiling for Entire Software Stacks Based on the Flow Reconstruction Principle
Xu Zhao 0004, Kirk Rodrigues, Yu Luo 0006, Ding Yuan 0004, Michael Stumm
OSDI2