Peng Huang 0005

dblp:29/1726-5 · DBLP profile ↗
← Back
39ranked-venue papers
5as first author
16since 2021 · last 2025
0000-0001-6315-0848ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 20 · 3 first-author · 10 since 2021Systems, architecture and hardware · 10 · 1 first-author · 3 since 2021Computer networks · 8 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 One-Size-Fits-None: Understanding and Enhancing Slow-Fault Tolerance in Modern Distributed Systems
Ruiming Lu, Yunchi Lu, Yuxuan Jiang 0016, Guangtao Xue, Peng Huang 0005
NSDI5
2025 Training with Confidence: Catching Silent Errors in Deep Learning Training with Automated Proactive Checks
Yuxuan Jiang 0016, Ziming Zhou, Boyu Xu 0005, Beijie Liu, Runhui Xu, Peng Huang 0005
OSDI6
2025 Deriving Semantic Checkers from Tests to Detect Silent Failures in Production Distributed Systems
Chang Lou, Dimas Shidqi Parikesit, Yujin Huang, Zhewen Yang, Senapati Diwangkara, Yuzhuo Jing, Achmad I. Kistijantoro, Ding Yuan 0004, Suman Nath, Peng Huang 0005
OSDI10
2025 Mitigating Application Resource Overload with Targeted Task Cancellation
abstract
Modern software inevitably encounters periods of resource overload, during which it must still sustain high servicelevel objective (SLO) attainment while minimizing request loss. However, achieving this balance is challenging due to subtle and unpredictable internal resource contention among concurrently executing requests. Traditional overload control mechanisms, which rely on global signals, such as queuing delays, fail to handle application resource overload effectively because they cannot accurately predict which requests will monopolize critical resources.
Yigong Hu, Zeyin Zhang, Yile Gu, Shuangyu Lei, Baris Kasikci, Peng Huang 0005
SOSP7
2025 Optimistic Recovery for High-Availability Software via Partial Process State Preservation
abstract
Achieving high availability for modern software requires fast and correct recovery from inevitable faults. This is notoriously difficult. Existing techniques either guarantee correctness by discarding all state but suffer from long downtime, or preserve all state to recover quickly but reintroduce the fault.
Yuzhuo Jing, Yuqi Mai, Angting Cai, Wanning He, Xiaoyang Qian, Peter M. Chen, Peng Huang 0005
SOSP8
2025 TrainVerify: Equivalence-Based Verification for Distributed LLM Training
abstract
Training large language models (LLMs) at scale requires parallel execution across thousands of devices, incurring enormous computational costs. Yet, these costly distributed trainings are prone to correctness bugs, causing silent errors and potentially wasting millions of GPU hours. These bugs are challenging to expose through testing.
Yunchi Lu, Youshan Miao, Cheng Tan 0005, Peng Huang 0005, Xian Zhang 0001, Fan Yang 0024
SOSP4
2024 Efficient Exposure of Partial Failure Bugs in Distributed Systems with Inferred Abstract States
Peng Huang 0005
NSDI3
2024 Efficient Reproduction of Fault-Induced Failures in Distributed Systems with Feedback-Driven Fault Injection
abstract
Debugging a failure usually requires reproducing it first. This can be hard for failures in production distributed systems, where bugs are exposed only by some unusual faulty events. While fault injection testing becomes popular, existing solutions are designed for bug finding. They are ineffective and inefficient to reproduce a specific failure during debugging.
Tanakorn Leesatapornwongsa, Suman Nath, Peng Huang 0005
SOSP5
2023 Effective Performance Issue Diagnosis with Value-Assisted Cost Profiling
abstract
Diagnosing performance issues is often difficult, especially when they occur only during some program executions. Profilers can help with performance debugging, but are ineffective when the most costly functions are not the root causes of performance issues. To address this problem, we introduce a new profiling methodology, value-assisted cost profiling, and a tool vProf. Our insight is that capturing the values of variables can greatly help diagnose performance issues. vProf continuously records values while profiling normal and buggy program executions. It identifies anomalies in the values and the functions where they occur to pinpoint the real root causes of performance issues. Using a set of 15 real-world performance bugs in four widely used applications, we show that vProf is effective at diagnosing all of the issues while other state-of-the-art tools diagnose only a few of them. We further use vProf to diagnose longstanding performance issues in these applications that have been unresolved for over four years.
Lingmei Weng, Yigong Hu, Peng Huang 0005, Jason Nieh
EuroSys3
2023 Simplifying Cloud Management with Cloudless Computing
abstract
Cloud computing has transformed the IT industry, but managing cloud infrastructures remains a difficult task. We make a case for putting today's management practices, known as "Infrastructure-as-Code," on a firmer ground via a principled design. We call this end goal Cloudless Computing: it aims to simplify cloud infrastructure management tasks by supporting them "as-a-service," analogous to serverless computing that relieves users of the burden of managing server instances. By assisting tenants with these tasks, cloud resources will be presented to their users more readily without the undue burden of complex control. We describe the research problems by examining the typical lifecycle of today's cloud infrastructure management, and identify places where a cloudless approach will advance the state of the art.
Yiming Qiu 0001, Patrick Tser Jern Kon, Jiarong Xing, Yibo Huang 0005, Xinyu Wang 0006, Peng Huang 0005, Mosharaf Chowdhury, Ang Chen 0001
HotNets7
2023 Pushing Performance Isolation Boundaries into Application with pBox
abstract
Modern applications are highly concurrent with a diverse mix of activities. One activity can adversely impact the performance of other activities in an application, leading to intra-application interference. Providing fine-grained performance isolation is desirable. Unfortunately, the extensive performance isolation solutions today focus on mitigating coarse-grained interference among multiple applications. They cannot well address intra-app interference, because such issues are typically not caused by contention on hardware resources.
Yigong Hu, Gongqi Huang, Peng Huang 0005
SOSP3
2023 Hybrid Block Storage for Efficient Cloud Volume Service
abstract
The migration of traditional desktop and server applications to the cloud brings challenge of high performance, high reliability, and low cost to the underlying cloud storage. To satisfy the requirement, this article proposes a hybrid cloud-scale block storage system called Ursa . Trace analysis shows that the I/O patterns served by block storage have only limited locality to exploit. Therefore, instead of using solid state drives (SSDs) as a cache layer, Ursa proposes hybrid storage structure that directly stores primary replicas on SSDs and replicates backup replicas on hard disk drives (HDDs) . At the core of Ursa ’s hybrid storage design is an adaptive journal that can bridge the performance gap between primary SSDs and backup HDDs for random writes by transforming small backup writes into journal appends, which are then asynchronously replayed and merged to backup HDDs. To efficiently index the journal, we design a novel range-optimized merge-tree structure that combines a continuous range of keys into a single composite key {offset,length} . Ursa integrates the hybrid structure with designs for high reliability, scalability, and availability. Experiments show that Ursa in its hybrid mode achieves almost the same performance as in its SSD-only mode (storing all replicas on SSDs), and outperforms other block stores (Ceph and Sheepdog) even in their SSD-only mode while achieving much higher CPU efficiency (IOPS and throughput per core).
Yiming Zhang 0003, Huiba Li, Shengyun Liu, Peng Huang 0005
ACM Trans. Storage4
2022 Operating System Support for Safe and Efficient Auxiliary Execution
Yuzhuo Jing, Peng Huang 0005
OSDI2
2022 RESIN: A Holistic Service for Dealing with Memory Leaks in Production Cloud Infrastructure
Chang Lou, Peng Huang 0005, Yingnong Dang, Si Qin, Xinsheng Yang, Xukun Li, Qingwei Lin, Murali Chintalapati
OSDI3
2022 Demystifying and Checking Silent Semantic Violations in Large Distributed Systems
Chang Lou, Yuzhuo Jing, Peng Huang 0005
OSDI3
2021 Argus: Debugging Performance Issues in Modern Desktop Applications with Annotated Causal Tracing
Lingmei Weng, Peng Huang 0005, Jason Nieh
USENIX ATC2
2020 Gandalf: An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud Infrastructure
Ze Li 0005, Ken Hsieh, Yingnong Dang, Peng Huang 0005, Pankaj Singh, Xinsheng Yang, Qingwei Lin, Youjiang Wu, Sebastien Levy, Murali Chintalapati
NSDI5
2020 Understanding, Detecting and Localizing Partial Failures in Large System Software
Chang Lou, Peng Huang 0005
NSDI2
2020 Automated Reasoning and Detection of Specious Configuration in Large Systems with Symbolic Execution
Yigong Hu, Gongqi Huang, Peng Huang 0005
OSDI3
2020 Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM Interruptions
Sebastien Levy, Randolph Yao, Youjiang Wu, Yingnong Dang, Peng Huang 0005, Zheng Mu, Pu Zhao 0004, Tarun Ramani, Naga K. Govindaraju, Xukun Li, Qingwei Lin, Gil Lapid Shafriri, Murali Chintalapati
OSDI5
2019 A Case for Lease-Based, Utilitarian Resource Management on Mobile Devices
abstract
Mobile apps have become indispensable in our daily lives, but many apps are not designed to be energy-aware that they may consume the constrained resources on mobile devices in a wasteful manner. Blindly throttling heavy resource usage, while helps reducing energy consumption, prohibits apps from taking advantages of the resources to do useful work. We argue that addressing this issue requires mobile OS to continuously assess if a resource is still truly needed even after it is granted to an app.
Yigong Hu, Suyi Liu, Peng Huang 0005
ASPLOS3
2019 URSA: Hybrid Block Storage for Cloud-Scale Virtual Disks
abstract
This paper presents URSA, a hybrid block store that provides virtual disks for various applications to run efficiently on cloud VMs. Trace analysis shows that the I/O patterns served by block storage have limited locality to exploit. Therefore, instead of using SSDs as a cache layer, URSA proposes an SSD-HDD-hybrid storage structure that directly stores primary replicas on SSDs and replicates backup replicas on HDDs, using journals to bridge the performance gap between SSDs and HDDs. URSA integrates the hybrid structure with designs for high reliability, scalability, and availability. Experiments show that URSA in its hybrid mode achieves almost the same performance as in its SSD-only mode (storing all replicas on SSDs), and outperforms other block stores (Ceph and Sheepdog) even in their SSD-only mode while achieving much higher CPU efficiency (performance per core). We also discuss some practical issues in our deployment.
Huiba Li, Yiming Zhang 0003, Dongsheng Li 0001, Shengyun Liu, Peng Huang 0005, Zheng Qin 0002, Kai Chen 0005, Yongqiang Xiong
EuroSys6
2019 Comprehensive and Efficient Runtime Checking in System Software through Watchdogs
abstract
Systems software today is composed of numerous modules and exhibits complex failure modes. Existing failure detectors focus on catching simple, complete failures and treat programs uniformly at the process level. In this paper, we argue that modern software needs intrinsic failure detectors that are tailored to individual systems and can detect anomalies within a process at finer granularity. We particularly advocate a notion of intrinsic software watchdogs and propose an abstraction for it. Among the different styles of watchdogs, we believe watchdogs that imitate the main program can provide the best combination of completeness, accuracy and localization for detecting gray failures. But, manually constructing such mimic-type watchdogs is challenging and time-consuming. To close this gap, we present an early exploration for automatically generating mimic-type watchdogs.
Chang Lou, Peng Huang 0005
HotOS2
2018 End-to-End Automated Exploit Generation for Validating the Security of Processor Designs
abstract
This paper presents Coppelia, an end-to-end tool that, given a processor design and a set of security-critical invariants, automatically generates complete, replayable exploit programs to help designers find, contextualize, and assess the security threat of hardware vulnerabilities. In Coppelia, we develop a hardware-oriented backward symbolic execution engine with a new cycle stitching method and fast validation technique, along with several optimizations for exploit generation. We then add program stubs to complete the exploit. We evaluate Coppelia on three CPUs of different architectures. Coppelia is able to find and generate exploits for 29 of 31 known vulnerabilities in these CPUs, including 11 vulnerabilities that commercial and academic model checking tools can not find. All of the generated exploits are successfully replayable on an FPGA board. Moreover, Coppelia finds 4 new vulnerabilities along with exploits in these CPUs. We also use Coppelia to verify whether a security patch indeed fixed a vulnerability, and to refine a set of assertions.
Rui Zhang 0068, Calvin Deutschbein, Peng Huang 0005, Cynthia Sturton
MICRO3
2018 Capturing and Enhancing In Situ System Observability for Failure Detection
Peng Huang 0005, Chuanxiong Guo, Jacob R. Lorch, Lidong Zhou, Yingnong Dang
OSDI1
2018 TerseCades: Efficient Data Compression in Stream Processing
Gennady Pekhimenko, Chuanxiong Guo, Myeongjae Jeon, Peng Huang 0005, Lidong Zhou
USENIX ATC4
2017 Gray Failure: The Achilles' Heel of Cloud-Scale Systems
abstract
Cloud scale provides the vast resources necessary to replace failed components, but this is useful only if those failures can be detected. For this reason, the major availability breakdowns and performance anomalies we see in cloud environments tend to be caused by subtle underlying faults, i.e., gray failure rather than fail-stop failure. In this paper, we discuss our experiences with gray failure in production cloud-scale systems to show its broad scope and consequences. We also argue that a key feature of gray failure is differential observability: that the system's failure detectors may not notice problems even when applications are afflicted by them. This realization leads us to believe that, to best deal with them, we should focus on bridging the gap between different components' perceptions of what constitutes failure.
Peng Huang 0005, Chuanxiong Guo, Lidong Zhou, Jacob R. Lorch, Yingnong Dang, Murali Chintalapati, Randolph Yao
HotOS1
2017 Early Detection of Configuration Errors to Reduce Failure Damage
Tianyin Xu, Xinxin Jin, Peng Huang 0005, Yuanyuan Zhou 0001, Shan Lu 0001, Shankar Pasupathy
USENIX ATC3
2016 NChecker: saving mobile app developers from network disruptions
abstract
Most of today's mobile apps rely on the underlying networks to deliver key functions such as web browsing, file synchronization, and social networking. Compared to desktop-based networks, mobile networks are much more dynamic with frequent connectivity disruptions, network type switches, and quality changes, posing unique programming challenges for mobile app developers.
Xinxin Jin, Peng Huang 0005, Tianyin Xu, Yuanyuan Zhou 0001
EuroSys2
2016 DefDroid: Towards a More Defensive Mobile OS Against Disruptive App Behavior
abstract
The mobile app market is enjoying an explosive growth. Many people, including teenagers, set about developing mobile apps. Unfortunately, due to developers' inexperience and the unique mobile programming paradigms, a growing number of immature apps are released to users. Despite having useful functionalities, these apps exhibit disruptive behaviors that are inconsiderate to the mobile system as a whole, e.g.,, retrying network connections too aggressively, waking up the device too frequently, or holding resources for unnecessarily long. These behaviors adversely affect other apps running on the same device and frustrate users with battery drain, excessive cellular data consumption, storage overuse, etc.
Peng Huang 0005, Tianyin Xu, Xinxin Jin, Yuanyuan Zhou 0001
MobiSys1
2016 Early Detection of Configuration Errors to Reduce Failure Damage
Tianyin Xu, Xinxin Jin, Peng Huang 0005, Yuanyuan Zhou 0001, Shan Lu 0001, Shankar Pasupathy
OSDI3
2015 ConfValley: a systematic configuration validation framework for cloud services
abstract
Studies and many incidents in the headlines suggest misconfigurations remain a major cause of unavailability in large systems despite the large amount of work put into detecting, diagnosing and repairing them. In part, this is because many of the solutions are either post-mortem or too expensive to use in production cloud-scale systems. Configuration validation is the process of explicitly defining specifications and proactively checking configurations against those specifications to prevent misconfigurations from entering production.
Peng Huang 0005, William J. Bolosky, Yuanyuan Zhou 0001
EuroSys1
2014 Performance regression testing target prioritization via performance risk analysis
abstract
As software evolves, problematic changes can significantly degrade software performance, i.e., introducing performance regression. Performance regression testing is an effective way to reveal such issues in early stages. Yet because of its high overhead, this activity is usually performed infrequently. Consequently, when performance regression issue is spotted at a certain point, multiple commits might have been merged since last testing. Developers have to spend extra time and efforts narrowing down which commit caused the problem. Existing efforts try to improve performance regression testing efficiency through test case reduction or prioritization.
Peng Huang 0005, Xiao Ma 0014, Dongcai Shen, Yuanyuan Zhou 0001
ICSE1
2013 eDoctor: Automatically Diagnosing Abnormal Battery Drain Issues on Smartphones
Xiao Ma 0014, Peng Huang 0005, Xinxin Jin, Dongcai Shen, Yuanyuan Zhou 0001, Lawrence K. Saul, Geoffrey M. Voelker
NSDI2
2013 Do not blame users for misconfigurations
abstract
Similar to software bugs, configuration errors are also one of the major causes of today's system failures. Many configuration issues manifest themselves in ways similar to software bugs such as crashes, hangs, silent failures. It leaves users clueless and forced to report to developers for technical support, wasting not only users' but also developers' precious time and effort. Unfortunately, unlike software bugs, many software developers take a much less active, responsible role in handling configuration errors because "they are users' faults."
Tianyin Xu, Peng Huang 0005, Tianwei Sheng, Ding Yuan 0004, Yuanyuan Zhou 0001, Shankar Pasupathy
SOSP3
2013 Understanding latent interactions in online social networks
abstract
Popular online social networks (OSNs) like Facebook and Twitter are changing the way users communicate and interact with the Internet. A deep understanding of user interactions in OSNs can provide important insights into questions of human social behavior and into the design of social platforms and applications. However, recent studies have shown that a majority of user interactions on OSNs are latent interactions , that is, passive actions, such as profile browsing, that cannot be observed by traditional measurement techniques. In this article, we seek a deeper understanding of both active and latent user interactions in OSNs. For quantifiable data on latent user interactions, we perform a detailed measurement study on Renren, the largest OSN in China with more than 220 million users to date. All friendship links in Renren are public, allowing us to exhaustively crawl a connected graph component of 42 million users and 1.66 billion social links in 2009. Renren also keeps detailed, publicly viewable visitor logs for each user profile. We capture detailed histories of profile visits over a period of 90 days for users in the Peking University Renren network and use statistics of profile visits to study issues of user profile popularity, reciprocity of profile visits, and the impact of content updates on user popularity. We find that latent interactions are much more prevalent and frequent than active events, are nonreciprocal in nature, and that profile popularity is correlated with page views of content rather than with quantity of content updates. Finally, we construct latent interaction graphs as models of user browsing behavior and compare their structural properties, evolution, community structure, and mixing times against those of both active interaction graphs and social graphs.
Jing Jiang 0005, Christo Wilson, Xiao Wang 0018, Wenpeng Sha, Peng Huang 0005, Yafei Dai, Ben Y. Zhao
ACM Trans. Web5
2012 Be Conservative: Enhancing Failure Diagnosis with Proactive Logging
Ding Yuan 0004, Peng Huang 0005, Yang Liu 0044, Michael Mihn-Jong Lee, Xiaoming Tang, Yuanyuan Zhou 0001, Stefan Savage
OSDI3
2010 Understanding latent interactions in online social networks
abstract
Popular online social networks (OSNs) like Facebook and Twitter are changing the way users communicate and interact with the Internet. A deep understanding of user interactions in OSNs can provide important insights into questions of human social behavior, and into the design of social platforms and applications. However, recent studies have shown that a majority of user interactions on OSNs are latent interactions, passive actions such as profile browsing that cannot be observed by traditional measurement techniques. In this paper, we seek a deeper understanding of both visible and latent user interactions in OSNs. For quantifiable data on latent user interactions, we perform a detailed measurement study on Renren, the largest OSN in China with more than 150 million users to date. All friendship links in Renren are public, allowing us to exhaustively crawl a connected graph component of 42 million users and 1.66 billion social links in 2009. Renren also keeps detailed visitor logs for each user profile, and counters for each photo and diary/blog entry. We capture detailed histories of profile visits over a period of 90 days for more than 61,000 users in the Peking University Renren network, and use statistics of profile visits to study issues of user profile popularity, reciprocity of profile visits, and the impact of content updates on user popularity. We find that latent interactions are much more prevalent and frequent than visible events, non-reciprocal in nature, and that profile popularity are uncorrelated with the frequency of content updates. Finally, we construct latent interaction graphs as models of user browsing behavior, and compare their structural properties against those of both visible interaction graphs and social graphs.
Jing Jiang 0005, Christo Wilson, Xiao Wang 0018, Peng Huang 0005, Wenpeng Sha, Yafei Dai, Ben Y. Zhao
Internet Measurement Conference4
2010 A multiple user sharing behaviors based approach for fake file detection in P2P environments
Jing Jiang 0005, Qinyuan Feng, Peng Huang 0005, Yafei Dai
Sci. China Inf. Sci.4