VLDB 2026 Research / reviewers in the wild / expert
Peng Huang 0005
dblp:29/1726-5
· DBLP profile ↗
39ranked-venue papers
5as first author
16since 2021 · last 2025
0000-0001-6315-0848ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 20 · 3 first-author · 10 since 2021Systems, architecture and hardware · 10 · 1 first-author · 3 since 2021Computer networks · 8 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | One-Size-Fits-None: Understanding and Enhancing Slow-Fault Tolerance in Modern Distributed Systems
Ruiming Lu, Yunchi Lu, Yuxuan Jiang 0016, Guangtao Xue, Peng Huang 0005 |
NSDI | 5 |
| 2025 | Training with Confidence: Catching Silent Errors in Deep Learning Training with Automated Proactive Checks
Yuxuan Jiang 0016, Ziming Zhou, Boyu Xu 0005, Beijie Liu, Runhui Xu, Peng Huang 0005 |
OSDI | 6 |
| 2025 | Deriving Semantic Checkers from Tests to Detect Silent Failures in Production Distributed Systems
Chang Lou, Dimas Shidqi Parikesit, Yujin Huang, Zhewen Yang, Senapati Diwangkara, Yuzhuo Jing, Achmad I. Kistijantoro, Ding Yuan 0004, Suman Nath, Peng Huang 0005 |
OSDI | 10 |
| 2025 | Mitigating Application Resource Overload with Targeted Task CancellationabstractModern software inevitably encounters periods of resource overload, during which it must still sustain high servicelevel objective (SLO) attainment while minimizing request loss. However, achieving this balance is challenging due to subtle and unpredictable internal resource contention among concurrently executing requests. Traditional overload control mechanisms, which rely on global signals, such as queuing delays, fail to handle application resource overload effectively because they cannot accurately predict which requests will monopolize critical resources. Yigong Hu, Zeyin Zhang, Yile Gu, Shuangyu Lei, Baris Kasikci, Peng Huang 0005 |
SOSP | 7 |
| 2025 | Optimistic Recovery for High-Availability Software via Partial Process State PreservationabstractAchieving high availability for modern software requires fast and correct recovery from inevitable faults. This is notoriously difficult. Existing techniques either guarantee correctness by discarding all state but suffer from long downtime, or preserve all state to recover quickly but reintroduce the fault. Yuzhuo Jing, Yuqi Mai, Angting Cai, Wanning He, Xiaoyang Qian, Peter M. Chen, Peng Huang 0005 |
SOSP | 8 |
| 2025 | TrainVerify: Equivalence-Based Verification for Distributed LLM TrainingabstractTraining large language models (LLMs) at scale requires parallel execution across thousands of devices, incurring enormous computational costs. Yet, these costly distributed trainings are prone to correctness bugs, causing silent errors and potentially wasting millions of GPU hours. These bugs are challenging to expose through testing. Yunchi Lu, Youshan Miao, Cheng Tan 0005, Peng Huang 0005, Xian Zhang 0001, Fan Yang 0024 |
SOSP | 4 |
| 2024 | Efficient Exposure of Partial Failure Bugs in Distributed Systems with Inferred Abstract States
Peng Huang 0005 |
NSDI | 3 |
| 2024 | Efficient Reproduction of Fault-Induced Failures in Distributed Systems with Feedback-Driven Fault InjectionabstractDebugging a failure usually requires reproducing it first. This can be hard for failures in production distributed systems, where bugs are exposed only by some unusual faulty events. While fault injection testing becomes popular, existing solutions are designed for bug finding. They are ineffective and inefficient to reproduce a specific failure during debugging. Tanakorn Leesatapornwongsa, Suman Nath, Peng Huang 0005 |
SOSP | 5 |
| 2023 | Effective Performance Issue Diagnosis with Value-Assisted Cost ProfilingabstractDiagnosing performance issues is often difficult, especially when they occur only during some program executions. Profilers can help with performance debugging, but are ineffective when the most costly functions are not the root causes of performance issues. To address this problem, we introduce a new profiling methodology, value-assisted cost profiling, and a tool vProf. Our insight is that capturing the values of variables can greatly help diagnose performance issues. vProf continuously records values while profiling normal and buggy program executions. It identifies anomalies in the values and the functions where they occur to pinpoint the real root causes of performance issues. Using a set of 15 real-world performance bugs in four widely used applications, we show that vProf is effective at diagnosing all of the issues while other state-of-the-art tools diagnose only a few of them. We further use vProf to diagnose longstanding performance issues in these applications that have been unresolved for over four years. Lingmei Weng, Yigong Hu, Peng Huang 0005, Jason Nieh |
EuroSys | 3 |
| 2023 | Simplifying Cloud Management with Cloudless ComputingabstractCloud computing has transformed the IT industry, but managing cloud infrastructures remains a difficult task. We make a case for putting today's management practices, known as "Infrastructure-as-Code," on a firmer ground via a principled design. We call this end goal Cloudless Computing: it aims to simplify cloud infrastructure management tasks by supporting them "as-a-service," analogous to serverless computing that relieves users of the burden of managing server instances. By assisting tenants with these tasks, cloud resources will be presented to their users more readily without the undue burden of complex control. We describe the research problems by examining the typical lifecycle of today's cloud infrastructure management, and identify places where a cloudless approach will advance the state of the art. Yiming Qiu 0001, Patrick Tser Jern Kon, Jiarong Xing, Yibo Huang 0005, Xinyu Wang 0006, Peng Huang 0005, Mosharaf Chowdhury, Ang Chen 0001 |
HotNets | 7 |
| 2023 | Pushing Performance Isolation Boundaries into Application with pBoxabstractModern applications are highly concurrent with a diverse mix of activities. One activity can adversely impact the performance of other activities in an application, leading to intra-application interference. Providing fine-grained performance isolation is desirable. Unfortunately, the extensive performance isolation solutions today focus on mitigating coarse-grained interference among multiple applications. They cannot well address intra-app interference, because such issues are typically not caused by contention on hardware resources. Yigong Hu, Gongqi Huang, Peng Huang 0005 |
SOSP | 3 |
| 2023 | Hybrid Block Storage for Efficient Cloud Volume ServiceabstractThe migration of traditional desktop and server applications to the cloud brings challenge of high performance, high reliability, and low cost to the underlying cloud storage. To satisfy the requirement, this article proposes a hybrid cloud-scale block storage system called Ursa . Trace analysis shows that the I/O patterns served by block storage have only limited locality to exploit. Therefore, instead of using solid state drives (SSDs) as a cache layer, Ursa proposes hybrid storage structure that directly stores primary replicas on SSDs and replicates backup replicas on hard disk drives (HDDs) . At the core of Ursa ’s hybrid storage design is an adaptive journal that can bridge the performance gap between primary SSDs and backup HDDs for random writes by transforming small backup writes into journal appends, which are then asynchronously replayed and merged to backup HDDs. To efficiently index the journal, we design a novel range-optimized merge-tree structure that combines a continuous range of keys into a single composite key {offset,length} . Ursa integrates the hybrid structure with designs for high reliability, scalability, and availability. Experiments show that Ursa in its hybrid mode achieves almost the same performance as in its SSD-only mode (storing all replicas on SSDs), and outperforms other block stores (Ceph and Sheepdog) even in their SSD-only mode while achieving much higher CPU efficiency (IOPS and throughput per core). Yiming Zhang 0003, Huiba Li, Shengyun Liu, Peng Huang 0005 |
ACM Trans. Storage | 4 |
| 2022 | Operating System Support for Safe and Efficient Auxiliary Execution
Yuzhuo Jing, Peng Huang 0005 |
OSDI | 2 |
| 2022 | RESIN: A Holistic Service for Dealing with Memory Leaks in Production Cloud Infrastructure
Chang Lou, Peng Huang 0005, Yingnong Dang, Si Qin, Xinsheng Yang, Xukun Li, Qingwei Lin, Murali Chintalapati |
OSDI | 3 |
| 2022 | Demystifying and Checking Silent Semantic Violations in Large Distributed Systems
Chang Lou, Yuzhuo Jing, Peng Huang 0005 |
OSDI | 3 |
| 2021 | Argus: Debugging Performance Issues in Modern Desktop Applications with Annotated Causal Tracing
Lingmei Weng, Peng Huang 0005, Jason Nieh |
USENIX ATC | 2 |
| 2020 | Gandalf: An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud Infrastructure
Ze Li 0005, Ken Hsieh, Yingnong Dang, Peng Huang 0005, Pankaj Singh, Xinsheng Yang, Qingwei Lin, Youjiang Wu, Sebastien Levy, Murali Chintalapati |
NSDI | 5 |
| 2020 | Understanding, Detecting and Localizing Partial Failures in Large System Software
Chang Lou, Peng Huang 0005 |
NSDI | 2 |
| 2020 | Automated Reasoning and Detection of Specious Configuration in Large Systems with Symbolic Execution
Yigong Hu, Gongqi Huang, Peng Huang 0005 |
OSDI | 3 |
| 2020 | Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM Interruptions
Sebastien Levy, Randolph Yao, Youjiang Wu, Yingnong Dang, Peng Huang 0005, Zheng Mu, Pu Zhao 0004, Tarun Ramani, Naga K. Govindaraju, Xukun Li, Qingwei Lin, Gil Lapid Shafriri, Murali Chintalapati |
OSDI | 5 |
| 2019 | A Case for Lease-Based, Utilitarian Resource Management on Mobile DevicesabstractMobile apps have become indispensable in our daily lives, but many apps are not designed to be energy-aware that they may consume the constrained resources on mobile devices in a wasteful manner. Blindly throttling heavy resource usage, while helps reducing energy consumption, prohibits apps from taking advantages of the resources to do useful work. We argue that addressing this issue requires mobile OS to continuously assess if a resource is still truly needed even after it is granted to an app. Yigong Hu, Suyi Liu, Peng Huang 0005 |
ASPLOS | 3 |
| 2019 | URSA: Hybrid Block Storage for Cloud-Scale Virtual DisksabstractThis paper presents URSA, a hybrid block store that provides virtual disks for various applications to run efficiently on cloud VMs. Trace analysis shows that the I/O patterns served by block storage have limited locality to exploit. Therefore, instead of using SSDs as a cache layer, URSA proposes an SSD-HDD-hybrid storage structure that directly stores primary replicas on SSDs and replicates backup replicas on HDDs, using journals to bridge the performance gap between SSDs and HDDs. URSA integrates the hybrid structure with designs for high reliability, scalability, and availability. Experiments show that URSA in its hybrid mode achieves almost the same performance as in its SSD-only mode (storing all replicas on SSDs), and outperforms other block stores (Ceph and Sheepdog) even in their SSD-only mode while achieving much higher CPU efficiency (performance per core). We also discuss some practical issues in our deployment. Huiba Li, Yiming Zhang 0003, Dongsheng Li 0001, Shengyun Liu, Peng Huang 0005, Zheng Qin 0002, Kai Chen 0005, Yongqiang Xiong |
EuroSys | 6 |
| 2019 | Comprehensive and Efficient Runtime Checking in System Software through WatchdogsabstractSystems software today is composed of numerous modules and exhibits complex failure modes. Existing failure detectors focus on catching simple, complete failures and treat programs uniformly at the process level. In this paper, we argue that modern software needs intrinsic failure detectors that are tailored to individual systems and can detect anomalies within a process at finer granularity. We particularly advocate a notion of intrinsic software watchdogs and propose an abstraction for it. Among the different styles of watchdogs, we believe watchdogs that imitate the main program can provide the best combination of completeness, accuracy and localization for detecting gray failures. But, manually constructing such mimic-type watchdogs is challenging and time-consuming. To close this gap, we present an early exploration for automatically generating mimic-type watchdogs. Chang Lou, Peng Huang 0005 |
HotOS | 2 |
| 2018 | End-to-End Automated Exploit Generation for Validating the Security of Processor DesignsabstractThis paper presents Coppelia, an end-to-end tool that, given a processor design and a set of security-critical invariants, automatically generates complete, replayable exploit programs to help designers find, contextualize, and assess the security threat of hardware vulnerabilities. In Coppelia, we develop a hardware-oriented backward symbolic execution engine with a new cycle stitching method and fast validation technique, along with several optimizations for exploit generation. We then add program stubs to complete the exploit. We evaluate Coppelia on three CPUs of different architectures. Coppelia is able to find and generate exploits for 29 of 31 known vulnerabilities in these CPUs, including 11 vulnerabilities that commercial and academic model checking tools can not find. All of the generated exploits are successfully replayable on an FPGA board. Moreover, Coppelia finds 4 new vulnerabilities along with exploits in these CPUs. We also use Coppelia to verify whether a security patch indeed fixed a vulnerability, and to refine a set of assertions. Rui Zhang 0068, Calvin Deutschbein, Peng Huang 0005, Cynthia Sturton |
MICRO | 3 |
| 2018 | Capturing and Enhancing In Situ System Observability for Failure Detection
Peng Huang 0005, Chuanxiong Guo, Jacob R. Lorch, Lidong Zhou, Yingnong Dang |
OSDI | 1 |
| 2018 | TerseCades: Efficient Data Compression in Stream Processing
Gennady Pekhimenko, Chuanxiong Guo, Myeongjae Jeon, Peng Huang 0005, Lidong Zhou |
USENIX ATC | 4 |
| 2017 | Gray Failure: The Achilles' Heel of Cloud-Scale SystemsabstractCloud scale provides the vast resources necessary to replace failed components, but this is useful only if those failures can be detected. For this reason, the major availability breakdowns and performance anomalies we see in cloud environments tend to be caused by subtle underlying faults, i.e., gray failure rather than fail-stop failure. In this paper, we discuss our experiences with gray failure in production cloud-scale systems to show its broad scope and consequences. We also argue that a key feature of gray failure is differential observability: that the system's failure detectors may not notice problems even when applications are afflicted by them. This realization leads us to believe that, to best deal with them, we should focus on bridging the gap between different components' perceptions of what constitutes failure. Peng Huang 0005, Chuanxiong Guo, Lidong Zhou, Jacob R. Lorch, Yingnong Dang, Murali Chintalapati, Randolph Yao |
HotOS | 1 |
| 2017 | Early Detection of Configuration Errors to Reduce Failure Damage
Tianyin Xu, Xinxin Jin, Peng Huang 0005, Yuanyuan Zhou 0001, Shan Lu 0001, Shankar Pasupathy |
USENIX ATC | 3 |
| 2016 | NChecker: saving mobile app developers from network disruptionsabstractMost of today's mobile apps rely on the underlying networks to deliver key functions such as web browsing, file synchronization, and social networking. Compared to desktop-based networks, mobile networks are much more dynamic with frequent connectivity disruptions, network type switches, and quality changes, posing unique programming challenges for mobile app developers. Xinxin Jin, Peng Huang 0005, Tianyin Xu, Yuanyuan Zhou 0001 |
EuroSys | 2 |
| 2016 | DefDroid: Towards a More Defensive Mobile OS Against Disruptive App BehaviorabstractThe mobile app market is enjoying an explosive growth. Many people, including teenagers, set about developing mobile apps. Unfortunately, due to developers' inexperience and the unique mobile programming paradigms, a growing number of immature apps are released to users. Despite having useful functionalities, these apps exhibit disruptive behaviors that are inconsiderate to the mobile system as a whole, e.g.,, retrying network connections too aggressively, waking up the device too frequently, or holding resources for unnecessarily long. These behaviors adversely affect other apps running on the same device and frustrate users with battery drain, excessive cellular data consumption, storage overuse, etc. Peng Huang 0005, Tianyin Xu, Xinxin Jin, Yuanyuan Zhou 0001 |
MobiSys | 1 |
| 2016 | Early Detection of Configuration Errors to Reduce Failure Damage
Tianyin Xu, Xinxin Jin, Peng Huang 0005, Yuanyuan Zhou 0001, Shan Lu 0001, Shankar Pasupathy |
OSDI | 3 |
| 2015 | ConfValley: a systematic configuration validation framework for cloud servicesabstractStudies and many incidents in the headlines suggest misconfigurations remain a major cause of unavailability in large systems despite the large amount of work put into detecting, diagnosing and repairing them. In part, this is because many of the solutions are either post-mortem or too expensive to use in production cloud-scale systems. Configuration validation is the process of explicitly defining specifications and proactively checking configurations against those specifications to prevent misconfigurations from entering production. Peng Huang 0005, William J. Bolosky, Yuanyuan Zhou 0001 |
EuroSys | 1 |
| 2014 | Performance regression testing target prioritization via performance risk analysisabstractAs software evolves, problematic changes can significantly degrade software performance, i.e., introducing performance regression. Performance regression testing is an effective way to reveal such issues in early stages. Yet because of its high overhead, this activity is usually performed infrequently. Consequently, when performance regression issue is spotted at a certain point, multiple commits might have been merged since last testing. Developers have to spend extra time and efforts narrowing down which commit caused the problem. Existing efforts try to improve performance regression testing efficiency through test case reduction or prioritization. Peng Huang 0005, Xiao Ma 0014, Dongcai Shen, Yuanyuan Zhou 0001 |
ICSE | 1 |
| 2013 | eDoctor: Automatically Diagnosing Abnormal Battery Drain Issues on Smartphones
Xiao Ma 0014, Peng Huang 0005, Xinxin Jin, Dongcai Shen, Yuanyuan Zhou 0001, Lawrence K. Saul, Geoffrey M. Voelker |
NSDI | 2 |
| 2013 | Do not blame users for misconfigurationsabstractSimilar to software bugs, configuration errors are also one of the major causes of today's system failures. Many configuration issues manifest themselves in ways similar to software bugs such as crashes, hangs, silent failures. It leaves users clueless and forced to report to developers for technical support, wasting not only users' but also developers' precious time and effort. Unfortunately, unlike software bugs, many software developers take a much less active, responsible role in handling configuration errors because "they are users' faults." Tianyin Xu, Peng Huang 0005, Tianwei Sheng, Ding Yuan 0004, Yuanyuan Zhou 0001, Shankar Pasupathy |
SOSP | 3 |
| 2013 | Understanding latent interactions in online social networksabstractPopular online social networks (OSNs) like Facebook and Twitter are changing the way users communicate and interact with the Internet. A deep understanding of user interactions in OSNs can provide important insights into questions of human social behavior and into the design of social platforms and applications. However, recent studies have shown that a majority of user interactions on OSNs are latent interactions , that is, passive actions, such as profile browsing, that cannot be observed by traditional measurement techniques. In this article, we seek a deeper understanding of both active and latent user interactions in OSNs. For quantifiable data on latent user interactions, we perform a detailed measurement study on Renren, the largest OSN in China with more than 220 million users to date. All friendship links in Renren are public, allowing us to exhaustively crawl a connected graph component of 42 million users and 1.66 billion social links in 2009. Renren also keeps detailed, publicly viewable visitor logs for each user profile. We capture detailed histories of profile visits over a period of 90 days for users in the Peking University Renren network and use statistics of profile visits to study issues of user profile popularity, reciprocity of profile visits, and the impact of content updates on user popularity. We find that latent interactions are much more prevalent and frequent than active events, are nonreciprocal in nature, and that profile popularity is correlated with page views of content rather than with quantity of content updates. Finally, we construct latent interaction graphs as models of user browsing behavior and compare their structural properties, evolution, community structure, and mixing times against those of both active interaction graphs and social graphs. Jing Jiang 0005, Christo Wilson, Xiao Wang 0018, Wenpeng Sha, Peng Huang 0005, Yafei Dai, Ben Y. Zhao |
ACM Trans. Web | 5 |
| 2012 | Be Conservative: Enhancing Failure Diagnosis with Proactive Logging
Ding Yuan 0004, Peng Huang 0005, Yang Liu 0044, Michael Mihn-Jong Lee, Xiaoming Tang, Yuanyuan Zhou 0001, Stefan Savage |
OSDI | 3 |
| 2010 | Understanding latent interactions in online social networksabstractPopular online social networks (OSNs) like Facebook and Twitter are changing the way users communicate and interact with the Internet. A deep understanding of user interactions in OSNs can provide important insights into questions of human social behavior, and into the design of social platforms and applications. However, recent studies have shown that a majority of user interactions on OSNs are latent interactions, passive actions such as profile browsing that cannot be observed by traditional measurement techniques. In this paper, we seek a deeper understanding of both visible and latent user interactions in OSNs. For quantifiable data on latent user interactions, we perform a detailed measurement study on Renren, the largest OSN in China with more than 150 million users to date. All friendship links in Renren are public, allowing us to exhaustively crawl a connected graph component of 42 million users and 1.66 billion social links in 2009. Renren also keeps detailed visitor logs for each user profile, and counters for each photo and diary/blog entry. We capture detailed histories of profile visits over a period of 90 days for more than 61,000 users in the Peking University Renren network, and use statistics of profile visits to study issues of user profile popularity, reciprocity of profile visits, and the impact of content updates on user popularity. We find that latent interactions are much more prevalent and frequent than visible events, non-reciprocal in nature, and that profile popularity are uncorrelated with the frequency of content updates. Finally, we construct latent interaction graphs as models of user browsing behavior, and compare their structural properties against those of both visible interaction graphs and social graphs. Jing Jiang 0005, Christo Wilson, Xiao Wang 0018, Peng Huang 0005, Wenpeng Sha, Yafei Dai, Ben Y. Zhao |
Internet Measurement Conference | 4 |
| 2010 | A multiple user sharing behaviors based approach for fake file detection in P2P environments
Jing Jiang 0005, Qinyuan Feng, Peng Huang 0005, Yafei Dai |
Sci. China Inf. Sci. | 4 |