Ding Yuan 0004

dblp:230/9043-4 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0001-6322-0295ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 23 · 5 first-author · 6 since 2021Systems, architecture and hardware · 7 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Deriving Semantic Checkers from Tests to Detect Silent Failures in Production Distributed Systems
Chang Lou, Dimas Shidqi Parikesit, Yujin Huang, Zhewen Yang, Senapati Diwangkara, Yuzhuo Jing, Achmad I. Kistijantoro, Ding Yuan 0004, Suman Nath, Peng Huang 0005
OSDI8
2024 μSlope: High Compression and Fast Search on Semi-Structured Logs
Devin Gibson, Kirk Rodrigues, Yu Luo 0006, Kaibo Wang, Yupeng Fu, Ding Yuan 0004
OSDI9
2023 Relational Debugging - Pinpointing Root Causes of Performance Problems
Xiang Ren 0003, Sitao Wang, Zhuqi Jin, David Lion, Adrian Chiu, Tianyin Xu, Ding Yuan 0004
OSDI7
2022 ctFS: Replacing File Indexing with Hardware Memory Translation through Contiguous File Allocation for Persistent Memory
Ruibin Li, Xiang Ren 0003, Xu Zhao 0004, Siwei He, Michael Stumm, Ding Yuan 0004
FAST6
2022 Hubble: Performance Debugging with In-Production, Just-In-Time Method Tracing on Android
Yu Luo 0006, Kirk Rodrigues, Cuiqin Li, Lijin Jiang, David Lion, Ding Yuan 0004
OSDI8
2022 Investigating Managed Language Runtime Performance: Why JavaScript and Python are 8x and 29x slower than C++, yet Java and Go can be Faster?
David Lion, Adrian Chiu, Michael Stumm, Ding Yuan 0004
USENIX ATC4
2022 ctFS: Replacing File Indexing with Hardware Memory Translation through Contiguous File Allocation for Persistent Memory
abstract
Persistent byte-addressable memory (PM) is poised to become prevalent in future computer systems. PMs are significantly faster than disk storage, and accesses to PMs are governed by the Memory Management Unit (MMU) just as accesses with volatile RAM. These unique characteristics shift the bottleneck from I/O to operations such as block address lookup—for example, in write workloads, up to 45% of the overhead in ext4-DAX is due to building and searching extent trees to translate file offsets to addresses on persistent memory. We propose a novel contiguous file system, ctFS, that eliminates most of the overhead associated with indexing structures such as extent trees in the file system. ctFS represents each file as a contiguous region of virtual memory, hence a lookup from the file offset to the address is simply an offset operation, which can be efficiently performed by the hardware MMU at a fraction of the cost of software-maintained indexes. Evaluating ctFS on real-world workloads such as LevelDB shows it outperforms ext4-DAX and SplitFS by 3.6× and 1.8×, respectively.
Ruibin Li, Xiang Ren 0003, Xu Zhao 0004, Siwei He, Michael Stumm, Ding Yuan 0004
ACM Trans. Storage6
2021 M3: end-to-end memory management in elastic system software stacks
abstract
This paper proposes M3, an end-to-end system that dynamically distributes memory resources among competing applications to maximize their overall performance. Today's data center workloads, can adapt to a wide range of memory sizes, and they are built on complex software stacks.
David Lion, Adrian Chiu, Ding Yuan 0004
EuroSys3
2021 CLP: Efficient and Scalable Search on Compressed Text Logs
Kirk Rodrigues, Yu Luo 0006, Ding Yuan 0004
OSDI3
2021 Understanding and Detecting Software Upgrade Failures in Distributed Systems
abstract
Upgrade is one of the most disruptive yet unavoidable maintenance tasks that undermine the availability of distributed systems. Any failure during an upgrade is catastrophic, as it further extends the service disruption caused by the upgrade. The increasing adoption of continuous deployment further increases the frequency and burden of the upgrade task. In practice, upgrade failures have caused many of today's high-profile cloud outages. Unfortunately, there has been little understanding of their characteristics.
Yongle Zhang 0007, Zhuqi Jin, Utsav Sethi, Kirk Rodrigues, Shan Lu 0001, Ding Yuan 0004
SOSP7
2019 An analysis of performance evolution of Linux's core operations
abstract
This paper presents an analysis of how Linux's performance has evolved over the past seven years. Unlike recent works that focus on OS performance in terms of scalability or service of a particular workload, this study goes back to basics: the latency of core kernel operations (e.g., system calls, context switching, etc.). To our surprise, the study shows that the performance of many core operations has worsened or fluctuated significantly over the years. For example, the select system call is 100% slower than it was just two years ago. An in-depth analysis shows that over the past seven years, core kernel subsystems have been forced to accommodate an increasing number of security enhancements and new features. These additions steadily add overhead to core kernel operations but also frequently introduce extreme slowdowns of more than 100%. In addition, simple misconfigurations have also severely impacted kernel performance. Overall, we find most of the slowdowns can be attributed to 11 changes.
Xiang Ren 0003, Kirk Rodrigues, Luyuan Chen, Juan Camilo Vega, Michael Stumm, Ding Yuan 0004
SOSP6
2019 The inflection point hypothesis: a principled debugging approach for locating the root cause of a failure
abstract
The end goal of failure diagnosis is to locate the root cause. Prior root cause localization approaches almost all rely on statistical analysis. This paper proposes taking a different approach based on the observation that if we model an execution as a totally ordered sequence of instructions, then the root cause can be identified by the first instruction where the failure execution deviates from the non-failure execution that has the longest instruction sequence prefix in common with that of the failure execution. Thus, root cause analysis is transformed into a principled search problem to identify the non-failure execution with the longest common prefix. We present Kairux, a tool that does just that. It is, in most cases, capable of pinpointing the root cause of a failure in a distributed system, in a fully automated way. Kairux uses tests from the system's rich unit test suite as building blocks to construct the non-failure execution that has the longest common prefix with the failure execution in order to locate the root cause. By evaluating Kairux on some of the most complex, real-world failures from HBase, HDFS, and ZooKeeper, we show that Kairux can accurately pinpoint each failure's respective root cause.
Yongle Zhang 0007, Kirk Rodrigues, Yu Luo 0006, Michael Stumm, Ding Yuan 0004
SOSP5
2017 The Game of Twenty Questions: Do You Know Where to Log?
abstract
A production system's printed logs are often the only source of runtime information available for postmortem debugging, performance analysis and profiling, security auditing, and user behavior analytics. Therefore, the quality of this data is critically important. Recent work has attempted to enhance log quality by recording additional variable values, but logging statement placement, i.e., where to place a logging statement, which is the most challenging and fundamental problem for improving log quality, has not been adequately addressed so far. This position paper proposes we automate the placement of logging statements by measuring how much uncertainty, i.e., the expected number of possible execution code paths taken by the software, can be removed by adding a logging statement to a basic block. Guided by ideas from information theory, we describe a simple approach that automates logging statement placement. Preliminary results suggest that our algorithm can effectively cover, and further improve, the existing logging statement placements selected by developers. It can compute an optimal logging statement placement that disambiguates the entire function call path with only 0.218% of slowdown.
Xu Zhao 0004, Kirk Rodrigues, Yu Luo 0006, Michael Stumm, Ding Yuan 0004, Yuanyuan Zhou 0001
HotOS5
2017 Pensieve: Non-Intrusive Failure Reproduction for Distributed Systems using the Event Chaining Approach
abstract
Complex and unforeseen failures in distributed systems must be diagnosed and replicated in a development environment so that developers can understand the underlying problem and verify the resolution. System logs often form the only source of diagnostic information, and developers reconstruct a failure using manual guesswork. This is an unpredictable and time-consuming process which can lead to costly service outages while a failure is repaired.
Yongle Zhang 0007, Serguei Makarov, Xiang Ren 0003, David Lion, Ding Yuan 0004
SOSP5
2017 Log20: Fully Automated Optimal Placement of Log Printing Statements under Specified Overhead Threshold
abstract
When systems fail in production environments, log data is often the only information available to programmers for postmortem debugging. Consequently, programmers' decision on where to place a log printing statement is of crucial importance, as it directly affects how effective and efficient postmortem debugging can be. This paper presents Log20, a tool that determines a near optimal placement of log printing statements under the constraint of adding less than a specified amount of performance overhead. Log20 does this in an automated way without any human involvement. Guided by information theory, the core of our algorithm measures how effective each log printing statement is in disambiguating code paths. To do so, it uses the frequencies of different execution paths that are collected from a production environment by a low-overhead tracing library. We evaluated Log20 on HDFS, HBase, Cassandra, and ZooKeeper, and observed that Log20 is substantially more efficient in code path disambiguation compared to the developers' manually placed log printing statements. Log20 can also output a curve showing the trade-off between the informativeness of the logs and the performance slowdown, so that a developer can choose the right balance.
Xu Zhao 0004, Kirk Rodrigues, Yu Luo 0006, Michael Stumm, Ding Yuan 0004, Yuanyuan Zhou 0001
SOSP5
2016 Don't Get Caught in the Cold, Warm-up Your JVM: Understand and Eliminate JVM Warm-up Overhead in Data-Parallel Systems
David Lion, Adrian Chiu, Hailong Sun 0001, Xin Zhuang, Nikola Grcevski, Ding Yuan 0004
OSDI6
2016 Non-Intrusive Performance Profiling for Entire Software Stacks Based on the Flow Reconstruction Principle
Xu Zhao 0004, Kirk Rodrigues, Yu Luo 0006, Ding Yuan 0004, Michael Stumm
OSDI4
2014 Simple Testing Can Prevent Most Critical Failures: An Analysis of Production Failures in Distributed Data-Intensive Systems
Ding Yuan 0004, Yu Luo 0006, Xin Zhuang, Guilherme Renna Rodrigues, Xu Zhao 0004, Yongle Zhang 0007, Pranay Jain, Michael Stumm
OSDI1
2014 lprof: A Non-intrusive Request Flow Profiler for Distributed Systems
Xu Zhao 0004, Yongle Zhang 0007, David Lion, Muhammad Faizan Ullah, Yu Luo 0006, Ding Yuan 0004, Michael Stumm
OSDI6
2013 Do not blame users for misconfigurations
abstract
Similar to software bugs, configuration errors are also one of the major causes of today's system failures. Many configuration issues manifest themselves in ways similar to software bugs such as crashes, hangs, silent failures. It leaves users clueless and forced to report to developers for technical support, wasting not only users' but also developers' precious time and effort. Unfortunately, unlike software bugs, many software developers take a much less active, responsible role in handling configuration errors because "they are users' faults."
Tianyin Xu, Peng Huang 0005, Tianwei Sheng, Ding Yuan 0004, Yuanyuan Zhou 0001, Shankar Pasupathy
SOSP6
2012 Characterizing logging practices in open-source software
abstract
Software logging is a conventional programming practice. While its efficacy is often important for users and developers to understand what have happened in the production run, yet software logging is often done in an arbitrary manner. So far, there have been little study for understanding logging practices in real world software. This paper makes the first attempt (to the best of our knowledge) to provide a quantitative characteristic study of the current log messages within four pieces of large open-source software. First, we quantitatively show that software logging is pervasive. By examining developers' own modifications to the logging code in the revision history, we find that they often do not make the log messages right in their first attempts, and thus need to spend a significant amount of efforts to modify the log messages as after-thoughts. Our study further provides several interesting findings on where developers spend most of their efforts in modifying the log messages, which can give insights for programmers, tool developers, and language and compiler designers to improve the current logging practice. To demonstrate the benefit of our study, we built a simple checker based on one of our findings and effectively detected 138 pieces of new problematic logging code from studied software (24 of them are already confirmed and fixed by developers).
Ding Yuan 0004, Yuanyuan Zhou 0001
ICSE1
2012 Be Conservative: Enhancing Failure Diagnosis with Proactive Logging
Ding Yuan 0004, Peng Huang 0005, Yang Liu 0044, Michael Mihn-Jong Lee, Xiaoming Tang, Yuanyuan Zhou 0001, Stefan Savage
OSDI1
2012 Improving Software Diagnosability via Log Enhancement
abstract
Diagnosing software failures in the field is notoriously difficult, in part due to the fundamental complexity of troubleshooting any complex software system, but further exacerbated by the paucity of information that is typically available in the production setting. Indeed, for reasons of both overhead and privacy, it is common that only the run-time log generated by a system (e.g., syslog) can be shared with the developers. Unfortunately, the ad-hoc nature of such reports are frequently insufficient for detailed failure diagnosis. This paper seeks to improve this situation within the rubric of existing practice. We describe a tool, LogEnhancer that automatically “enhances” existing logging code to aid in future post-failure debugging. We evaluate LogEnhancer on eight large, real-world applications and demonstrate that it can dramatically reduce the set of potential root failure causes that must be considered while imposing negligible overheads.
Ding Yuan 0004, Yuanyuan Zhou 0001, Stefan Savage
ACM Trans. Comput. Syst.1
2011 Improving software diagnosability via log enhancement
abstract
Diagnosing software failures in the field is notoriously difficult, in part due to the fundamental complexity of trouble-shooting any complex software system, but further exacerbated by the paucity of information that is typically available in the production setting. Indeed, for reasons of both overhead and privacy, it is common that only the run-time log generated by a system (e.g., syslog) can be shared with the developers. Unfortunately, the ad-hoc nature of such reports are frequently insufficient for detailed failure diagnosis. This paper seeks to improve this situation within the rubric of existing practice. We describe a tool, LogEnhancer that automatically enhances existing logging code to aid in future post-failure debugging. We evaluate LogEnhancer on eight large, real-world applications and demonstrate that it can dramatically reduce the set of potential root failure causes that must be considered during diagnosis while imposing negligible overheads.
Ding Yuan 0004, Yuanyuan Zhou 0001, Stefan Savage
ASPLOS1
2011 How do fixes become bugs?
abstract
Software bugs affect system reliability. When a bug is exposed in the field, developers need to fix them. Unfortunately, the bug-fixing process can also introduce errors, which leads to buggy patches that further aggravate the damage to end users and erode software vendors' reputation.
Zuoning Yin, Ding Yuan 0004, Yuanyuan Zhou 0001, Shankar Pasupathy, Lakshmi N. Bairavasundaram
SIGSOFT FSE2
2010 SherLog: error diagnosis by connecting clues from run-time logs
abstract
Computer systems often fail due to many factors such as software bugs or administrator errors. Diagnosing such production run failures is an important but challenging task since it is difficult to reproduce them in house due to various reasons: (1) unavailability of users' inputs and file content due to privacy concerns; (2) difficulty in building the exact same execution environment; and (3) non-determinism of concurrent executions on multi-processors.
Ding Yuan 0004, Haohui Mai, Weiwei Xiong, Lin Tan 0001, Yuanyuan Zhou 0001, Shankar Pasupathy
ASPLOS1
2008 CISpan: Comprehensive Incremental Mining Algorithms of Closed Sequential Patterns for Multi-Versional Software Mining
abstract
Recently, frequent sequential pattern mining algorithms have been widely used in software engineering field to mine various source code or specification patterns.In practice, software evolves from one version to another in its life span.The effort of mining frequent sequential patterns across multiple versions of a software can be substantially reduced by efficient incremental mining.This problem is challenging in this domain since the databases are usually updated in all kinds of manners including insertion, various modifications as well as removal of sequences.Also, different mining tools may have various mining constraints, such as low minimum support.None of the existing work can be applied effectively due to various limitations of such work.For example, our recent work, IncSpan, failed solving the problem because it could neither handle low minimum support nor removal of sequences from database.In this paper, we propose a novel, comprehensive incremental mining algorithm for frequent sequential pattern, CISpan (Comprehensive Incremental Sequential Pattern mining).CISpan supports both closed and complete incremental frequent sequence mining, with all kinds of updates to the database.Compared to IncSpan, CISpan tolerates a wide range for minimum support threshold (as low as 2).Our performance study shows that in addition to handling more test cases on which IncSpan fails, CISpan outperforms IncSpan in all test cases which IncSpan could handle, including various sequence length, number of sequences, modification ratio, etc., with an average of 3.4 times speedup.We also tested CISpan's performance on databases transformed from 20 consecutive versions of Linux Kernel source code.On average, CISpan outperforms the non-incremental CloSpan by 42 times.
Ding Yuan 0004, Kyuhyung Lee, Hong Cheng 0001, Zhenmin Li, Xiao Ma 0014, Yuanyuan Zhou 0001, Jiawei Han 0001
SDM1
2007 HotComments: How to Make Program Comments More Useful?
Lin Tan 0001, Ding Yuan 0004, Yuanyuan Zhou 0001
HotOS2
2007 /*icomment: bugs or bad comments?*/
abstract
Commenting source code has long been a common practice in software development. Compared to source code, comments are more direct, descriptive and easy-to-understand. Comments and sourcecode provide relatively redundant and independent information regarding a program's semantic behavior. As software evolves, they can easily grow out-of-sync, indicating two problems: (1) bugs -the source code does not follow the assumptions and requirements specified by correct program comments; (2) bad comments - comments that are inconsistent with correct code, which can confuse and mislead programmers to introduce bugs in subsequent versions. Unfortunately, as most comments are written in natural language, no solution has been proposed to automatically analyze commentsand detect inconsistencies between comments and source code. This paper takes the first step in automatically analyzing commentswritten in natural language to extract implicit program rulesand use these rules to automatically detect inconsistencies between comments and source code, indicating either bugs or bad comments. Our solution, iComment, combines Natural Language Processing(NLP), Machine Learning, Statistics and Program Analysis techniques to achieve these goals. We evaluate iComment on four large code bases: Linux, Mozilla, Wine and Apache. Our experimental results show that iComment automatically extracts 1832 rules from comments with 90.8-100% accuracy and detects 60 comment-code inconsistencies, 33 newbugs and 27 bad comments, in the latest versions of the four programs. Nineteen of them (12 bugs and 7 bad comments) have already been confirmed by the corresponding developers while the others are currently being analyzed by the developers.
Lin Tan 0001, Ding Yuan 0004, Yuanyuan Zhou 0001
SOSP2