VLDB 2026 Research / reviewers in the wild / expert
Yongle Zhang 0007
dblp:120/1722-7
· DBLP profile ↗
12ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0001-5350-5182ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 3 first-author · 3 since 2021Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CSnake: Detecting Self-Sustaining Cascading Failure via Causal Stitching of Fault PropagationsabstractRecent studies have revealed that self-sustaining cascading failures in distributed systems frequently lead to widespread outages, which are challenging to contain and recover from. Existing failure detection techniques struggle to expose such failures prior to deployment, as they typically require a complex combination of specific conditions to be triggered. This challenge stems from the inherent nature of cascading failures, as they typically involve a sequence of fault propagations, each activated by distinct conditions. Shangshu Qian, Lin Tan 0001, Yongle Zhang 0007 |
EuroSys | 3 |
| 2026 | UpFuzz: Detecting Data Format Incompatibility Bugs during Distributed Storage System Upgrade
P. C. Sruthi, Yayu Wang, Yaoxu Song, Bishal Basak Papan, Pedro Fonseca 0001, Yongle Zhang 0007 |
NSDI | 8 |
| 2025 | Understanding and Detecting Fail-Slow Hardware Failure Bugs in Cloud Systems
Gen Dong, Yu Hua 0001, Yongle Zhang 0007, Zhangyu Chen, Menglei Chen |
USENIX ATC | 3 |
| 2024 | Demystifying the Fight Against Complexity: A Comprehensive Study of Live Debugging Activities in Production Cloud SystemsabstractDebugging in production cloud systems (or live debugging) is a critical yet challenging task for on-call developers due to the financial impact of cloud service downtime and the inherent complexity of cloud systems. Unfortunately, how debugging is performed, and the unique challenges faced in the production cloud environment have not been investigated in detail. P. C. Sruthi, Deming Chu, Zhengyan Chen, Yongle Zhang 0007 |
SoCC | 5 |
| 2023 | Fail through the Cracks: Cross-System Interaction Failures in Modern Cloud SystemsabstractModern cloud systems are orchestrations of independent and interacting (sub-)systems, each specializing in important services (e.g., data processing, storage, resource management, etc.). Hence, cloud system reliability is affected not only by the reliability of each individual system, but also by the interplay between these systems. We observe that many recent production incidents of cloud systems are manifested through interactions across the system boundaries. However, there is a lack of systematic understanding of this emerging mode of failures, which we term as cross-system interaction failures (or CSI failures). This hinders the development of better design, integration practices, and new tooling. Lilia Tang, Chaitanya Bhandari, Yongle Zhang 0007, Anna Karanika, Shuyang Ji, Indranil Gupta, Tianyin Xu |
EuroSys | 3 |
| 2023 | Vicious Cycles in Distributed Software SystemsabstractA major threat to distributed software systems' reliability is vicious cycles, which are observed when an event in the distributed software system's execution causes a system degradation, and the degradation, in turn, causes more of such events. Vicious cycles often result in large-scale cloud outages that are hard to recover from due to their self-reinforcing nature. This paper formally defines Vicious Cycle, and conducts the first in-depth study of 33 real-world vicious cycles in 13 widely-used open-source distributed software systems, shedding light on the root causes, triggering conditions, and fixing strategies of vicious cycles, with over a dozen concrete implications to combat them. Our findings show that the majority of the vicious cycles are caused by incorrect error handlers, where the handlers do not obtain enough information to distinguish between 1) an error induced by incoming requests and 2) an error induced by an unexpected interference from another error handler. This paper further performs a feasibility study by 1) building a monitoring tool that prevents one type of vicious cycle by collecting information to make a more informed decision in error handling, and 2) investigating the effectiveness of one commonly suggested practice-injecting exponential backoff-to prevent vicious cycles induced by unconstrained retry. Shangshu Qian, Lin Tan 0001, Yongle Zhang 0007 |
ASE | 4 |
| 2022 | Efficiently detecting concurrency bugs in persistent memory programsabstractDue to the salient DRAM-comparable performance, TB-scale capacity, and non-volatility, persistent memory (PM) provides new opportunities for large-scale in-memory computing with instant crash recovery. However, programming PM systems is error-prone due to the existence of crash-consistency bugs, which are challenging to diagnose especially with concurrent programming widely adopted in PM applications to exploit hardware parallelism. Existing bug detection tools for DRAM-based concurrency issues cannot detect PM crash-consistency bugs because they are oblivious to PM operations and PM consistency. On the other hand, existing PM-specific debugging tools only focus on sequential PM programs and cannot effectively detect crash-consistency issues hidden in concurrent executions. Zhangyu Chen, Yu Hua 0001, Yongle Zhang 0007, Luochangqi Ding |
ASPLOS | 3 |
| 2021 | Understanding and Detecting Software Upgrade Failures in Distributed SystemsabstractUpgrade is one of the most disruptive yet unavoidable maintenance tasks that undermine the availability of distributed systems. Any failure during an upgrade is catastrophic, as it further extends the service disruption caused by the upgrade. The increasing adoption of continuous deployment further increases the frequency and burden of the upgrade task. In practice, upgrade failures have caused many of today's high-profile cloud outages. Unfortunately, there has been little understanding of their characteristics. Yongle Zhang 0007, Zhuqi Jin, Utsav Sethi, Kirk Rodrigues, Shan Lu 0001, Ding Yuan 0004 |
SOSP | 1 |
| 2019 | The inflection point hypothesis: a principled debugging approach for locating the root cause of a failureabstractThe end goal of failure diagnosis is to locate the root cause. Prior root cause localization approaches almost all rely on statistical analysis. This paper proposes taking a different approach based on the observation that if we model an execution as a totally ordered sequence of instructions, then the root cause can be identified by the first instruction where the failure execution deviates from the non-failure execution that has the longest instruction sequence prefix in common with that of the failure execution. Thus, root cause analysis is transformed into a principled search problem to identify the non-failure execution with the longest common prefix. We present Kairux, a tool that does just that. It is, in most cases, capable of pinpointing the root cause of a failure in a distributed system, in a fully automated way. Kairux uses tests from the system's rich unit test suite as building blocks to construct the non-failure execution that has the longest common prefix with the failure execution in order to locate the root cause. By evaluating Kairux on some of the most complex, real-world failures from HBase, HDFS, and ZooKeeper, we show that Kairux can accurately pinpoint each failure's respective root cause. Yongle Zhang 0007, Kirk Rodrigues, Yu Luo 0006, Michael Stumm, Ding Yuan 0004 |
SOSP | 1 |
| 2017 | Pensieve: Non-Intrusive Failure Reproduction for Distributed Systems using the Event Chaining ApproachabstractComplex and unforeseen failures in distributed systems must be diagnosed and replicated in a development environment so that developers can understand the underlying problem and verify the resolution. System logs often form the only source of diagnostic information, and developers reconstruct a failure using manual guesswork. This is an unpredictable and time-consuming process which can lead to costly service outages while a failure is repaired. Yongle Zhang 0007, Serguei Makarov, Xiang Ren 0003, David Lion, Ding Yuan 0004 |
SOSP | 1 |
| 2014 | Simple Testing Can Prevent Most Critical Failures: An Analysis of Production Failures in Distributed Data-Intensive Systems
Ding Yuan 0004, Yu Luo 0006, Xin Zhuang, Guilherme Renna Rodrigues, Xu Zhao 0004, Yongle Zhang 0007, Pranay Jain, Michael Stumm |
OSDI | 6 |
| 2014 | lprof: A Non-intrusive Request Flow Profiler for Distributed Systems
Xu Zhao 0004, Yongle Zhang 0007, David Lion, Muhammad Faizan Ullah, Yu Luo 0006, Ding Yuan 0004, Michael Stumm |
OSDI | 2 |