Yang Wang 0009

dblp:w/YangWang9 · DBLP profile ↗
← Back
41ranked-venue papers
5as first author
20since 2021 · last 2025
0000-0002-9721-4923ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 13 · 7 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 5 since 2021Computer networks · 3 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Flash-oriented Coded Storage: Research Status and Future Directions
abstract
Flash-based solid-state drives (SSDs) have been widely adopted in various storage systems, manifesting better performance than their forerunner HDDs. However, the characteristics of flash media post some drawbacks when deploying SSD-based storage systems. First, flash media have limited program/erase cycles, making them vulnerable to media failures. Second, SSD foreground I/Os can suffer from inconsistent performance due to interference from background operations like garbage collection (GC). The major solution to the above problems is to introduce data redundancy. Redundant data can not only detect raw bit errors and recover lost data but also enable I/O scheduling to sidestep SSDs that are under performance degradation. Compared with multi-replica, data coding is a more space-efficient way to provide redundancy. However, it is more challenging to simultaneously achieve low access latency, consistent performance, and fast recovery. This article examines the design of coded storage in existing storage systems, with a focus on flash storage systems, and how they address these challenges. The coded storage techniques are categorized into in-device coding, cross-device coding, and cross-machine coding. They are designed for different scenarios and purposes but share some design rationales in common. For each type of coded storage, we begin by presenting the theoretical bases, followed by an overview of how existing studies address the performance and endurance issues of coded storage from a systemic perspective. Finally, we review the history of coded storage, list several key insights from existing works, and speculate some promising directions for flash-oriented coded storage systems.
Zhiyue Li, Guangyan Zhang, Yang Wang 0009
ACM Trans. Storage3
2024 Scaler: Efficient and Effective Cross Flow Analysis
abstract
Performance analysis is challenging as different components (e.g., different libraries, and applications) of a complex system can interact with each other. However, few existing tools focus on understanding such interactions. To bridge this gap, we propose a novel analysis method-"Cross Flow Analysis (XFA)"- that monitors the interactions/flows across these components. We also built the Scaler profiler that provides a holistic view of the time spent on each component (e.g., library or application) and every API inside each component. This paper proposes multiple new techniques, such as Universal Shadow Table, and Relation-Aware Data Folding. These techniques enable Scaler to achieve low runtime overhead, low memory overhead, and high profiling accuracy. Based on our extensive experimental results, Scaler detects multiple unknown performance issues inside widely-used applications, and therefore will be a useful complement to existing work.
Steven (Jiaxun) Tang, Mingcan Xiang, Yang Wang 0009, Bo Wu 0002, Jianjun Chen 0001, Tongping Liu
ASE3
2024 MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at Hyperscale
Arnab Choudhury, Yang Wang 0009, Tuomas Pelkonen, Kutta Srinivasan, Abha Jain, Shenghao Lin, Delia David, Siavash Soleimanifard, Ritesh Tijoriwala, Denis Samoylov, Chunqiang Tang
OSDI2
2024 ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production Testing
Mike Chow, Yang Wang 0009, Ayichew Hailu, Rohan Bopardikar, Jialiang Qu, David Meisner, Santosh Sonawane, Rodrigo Paim, Mack Ward, Ivor Huang, Matt McNally, Daniel Hodges, Zoltan Farkas, Caner Gocmen, Elvis Huang, Chunqiang Tang
OSDI2
2024 FBDetect: Catching Tiny Performance Regressions at Hyperscale through In-Production Monitoring
abstract
This paper presents Meta's FBDetect system, which advances the state of the art in performance regression detection by catching regressions as small as 0.005% in noisy production environments. FBDetect monitors around 800,000 time series covering various types of metrics (e.g., throughput, latency, CPU and memory usage) to detect regressions caused by code or configuration changes in hundreds of services running on millions of servers. FBDetect introduces advanced techniques to capture stack traces fleet-wide, measure fine-grained subroutine-level performance differences, filter out deceptive false-positive regressions, deduplicate correlated regressions, and analyze root causes. Beyond these individual techniques, a key strength of FBDetect over prior work is its battle-tested robustness, proven by seven years of production use, and each year catching regressions that would have wasted millions of servers if left undetected.
Dong Young Yoon, Yang Wang 0009, Miao Yu 0023, Elvis Huang, Juan Ignacio Jones, Abhinay Kukkadapu, Osman Kocas, Jonathan Wiepert, Kapil Goenka, Sherry Chen, Yanjun Lin, Jocelyn Kong, Michael Chow, Chunqiang Tang
SOSP2
2024 On the Feasibility and Benefits of Extensive Evaluation
abstract
Benchmark and system parameters often have a significant impact on performance evaluation, which raises a long-lasting question about which settings we should use. This paper studies the feasibility and benefits of extensive evaluation. A full extensive evaluation, which tests all possible settings, is usually too expensive. This work investigates whether it is possible to sample a subset of the settings and, upon them, generate observations that match those from a full extensive evaluation. Towards this goal, we have explored the incremental sampling approach, which starts by measuring a small subset of random settings, builds a prediction model on these samples using the popular ANOVA approach, adds more samples if the model is not accurate enough, and terminates otherwise. To summarize our findings: 1) Enhancing a research prototype to support extensive evaluation mostly involves changing hard-coded configurations, which does not take much effort. 2) Some systems are highly predictable, which means that they can achieve accurate predictions with a low sampling rate, but some systems are less predictable. 3) We have not found a method that can consistently outperform random sampling + ANOVA. Based on these findings, we provide recommendations to improve artifact predictability and strategies for selecting parameter values during evaluation.
Yujie Hui, Miao Yu 0023, Hao Qi 0008, Yifan Gan, Tianxi Li, Yuke Li 0003, Xueyuan Ren, Sixiang Ma, Xiaoyi Lu 0001, Yang Wang 0009
Proc. ACM Manag. Data10
2024 IsoPredict: Dynamic Predictive Analysis for Detecting Unserializable Behaviors in Weakly Isolated Data Store Applications
abstract
Distributed data stores typically provide weak isolation levels, which are efficient but can lead to unserializable behaviors, which are hard for programmers to understand and often result in errors. This paper presents the first dynamic predictive analysis for data store applications under weak isolation levels, called IsoPredict . Given an observed serializable execution of a data store application, IsoPredict generates and solves SMT constraints to find an unserializable execution that is a feasible execution of the application. IsoPredict introduces novel techniques that handle divergent application behavior; solve mutually recursive sets of constraints; and balance coverage, precision, and performance. An evaluation on four transactional data store benchmarks shows that IsoPredict often predicts unserializable behaviors, 99 % of which are feasible.
Chujun Geng, Spyros Blanas, Michael D. Bond, Yang Wang 0009
Proc. ACM Program. Lang.4
2024 Understanding Silent Data Corruption in Processors for Mitigating its Effects
abstract
Silent Data Corruption (SDC) in processors can lead to various application-level issues, such as incorrect calculations and even data loss. Since traditional techniques are not effective in detecting these errors, it is very hard to address problems caused by SDCs in processors. For the same reason, knowledge about these SDCs in the wild is limited. In this article, we conduct an extensive study on CPU SDCs in a large production CPU population, encompassing over one million processors. In addition to collecting overall statistics, we perform a detailed study to understand (1) whether certain processor features are particularly vulnerable and their potential impacts on applications; (2) the reproducibility of CPU SDCs and the triggering conditions (e.g., temperature) of those less reproducible SDCs; and (3) the challenges to mitigate and handle CPU SDCs. We further investigate the implications that our observations obtained from the above researches have on the SDC fault models, SDC mitigation strategies, and the future research fields. In addition, we design an efficient SDC mitigation approach called Farron, which uses prioritized testing to detect highly reproducible SDCs and temperature control to mitigate less-reproducible SDCs. Our experimental results indicate that Farron can achieve better coverage of CPU SDCs with lower overall overhead, compared to the baseline used in Alibaba Cloud. This demonstrates that our observations are able to assist in SDC mitigation.
Shaobu Wang, Guangyan Zhang, Junyu Wei, Yang Wang 0009, Jiesheng Wu, Qingchao Luo
ACM Trans. Archit. Code Optim.4
2024 Exploiting Data-pattern-aware Vertical Partitioning to Achieve Fast and Low-cost Cloud Log Storage
abstract
Cloud logs can be categorized into on-line, off-line, and near-line logs based on the access frequency. Among them, near-line logs are mainly used for debugging, which means they prefer a low query latency for better user experience. Besides, the storage system for near-line logs prefers a low overall cost including the storage cost to store compressed logs, and the computation cost to compress logs and execute queries. These requirements pose challenges to achieving fast and cheap cloud log storage. This article proposes LogGrep, the first log compression and query tool that exploits both static and runtime patterns to properly structurize and organize log data in fine-grained units. The key idea of LogGrep is “vertical partitioning”: it stores each log entry into multiple partitions by first parsing logs into variable vectors according to static patterns and then extracting runtime pattern(s) automatically within each variable vector. Based on such runtime patterns, LogGrep further decomposes the variable vectors into fine-grained units called “Capsules” and stamps each Capsule with a summary of its values. During the query process, LogGrep can avoid decompressing and scanning Capsules that cannot match the keywords, with the help of the extracted runtime patterns and the Capsule stamps. We further show that the interactive debugging can well utilize the advantages of the vertical-partitioning-based method and mitigate its weaknesses as well. To this end, LogGrep integrates incremental locating and partial reconstruction to mitigate the read amplification incurred by vertical-partitioning-based method. We evaluate LogGrep on 37 cloud logs from the production environment of Alibaba Cloud and the public datasets. The results show that LogGrep can reduce the query latency and the overall cost by an order of magnitude compared with state-of-the-art works. Such results have confirmed that it is worthwhile applying a more sophisticated vertical-partitioning-based method to accelerate queries on compressed cloud logs.
Junyu Wei, Guangyan Zhang, Junchao Chen 0005, Yang Wang 0009, Tingtao Sun, Jiesheng Wu, Jiangwei Jiang
ACM Trans. Storage4
2023 Developer's Responsibility or Database's Responsibility? Rethinking Concurrency Control in Databases
Chaoyi Cheng, Mingzhe Han, Spyros Blanas, Michael D. Bond, Yang Wang 0009
CIDR6
2023 LogGrep: Fast and Cheap Cloud Log Storage by Exploiting both Static and Runtime Patterns
abstract
In cloud systems, near-line logs are mainly used for debugging, which means they prefer a low query latency for a better user experience, and like any other logs, they also prefer a low overall cost including storage cost to store compressed logs and computation cost to compress logs and execute queries.
Junyu Wei, Guangyan Zhang, Junchao Chen 0005, Yang Wang 0009, Tingtao Sun, Jiesheng Wu, Jiangwei Jiang
EuroSys4
2023 EdgeCut: Fast and Low-Overhead Access of User-Associated Contents from Edge Servers
abstract
User-associated contents play an increasingly important role in modern network applications. With growing deployments of edge servers, the capacity of content storage in edge clusters significantly increases, which provides great potential to satisfy content requests with much shorter latency. However, the large number of contents also causes the difficulty of searching contents on edge servers in different locations because indexing contents costs huge DRAM on each edge server. In this work, we explore the opportunity of efficiently indexing user-associated contents and propose a scalable content-sharing mechanism for edge servers, called EdgeCut, that significantly reduces content access latency by allowing many edge servers to share their cached contents. We design a compact and dynamic data structure called Ludo Locator that returns the IP address of the edge server that stores the requested user-associated content. We have implemented a prototype of EdgeCut in a real network environment running in a public geo-distributed cloud. The experiment results show that EdgeCut reduces content access latency by up to 50% and reduces cloud traffic by up to 50% compared to existing solutions. The memory cost is less than 50MB for 10 million mobile users. The simulations using real network latency data show EdgeCut's advantages over existing solutions on a large scale.
Yi Liu 0115, Minmei Wang, Shouqian Shi, Yang Wang 0009, Chen Qian 0001
SEC4
2023 Conveyor: One-Tool-Fits-All Continuous Software Deployment at Meta
Boris Grubic, Yang Wang 0009, Tyler Petrochko, Ran Yaniv, Brad Jones, David Callies, Matt Clarke-Lauer, Dan Kelley, Soteris Demetriou, Kenny Yu, Chunqiang Tang
OSDI2
2023 Understanding Silent Data Corruptions in a Large Production CPU Population
abstract
Silent Data Corruption (SDC) in processors can lead to various application-level issues, such as incorrect calculations and even data loss. Since traditional techniques are not effective in detecting processor SDCs, it is very hard to address problems caused by SDCs. For the same reason, knowledge about SDCs in the wild is limited.
Shaobu Wang, Guangyan Zhang, Junyu Wei, Yang Wang 0009, Jiesheng Wu, Qingchao Luo
SOSP4
2022 BigDL 2.0: Seamless Scaling of AI Pipelines from Laptops to Distributed Cluster
abstract
Most AI projects start with a Python notebook running on a single laptop; however, one usually needs to go through a mountain of pains to scale it to handle larger dataset (for both experimentation and production deployment). These usually entail many manual and error-prone steps for the data scientists to fully take advantage of the available hardware resources (e.g., SIMD instructions, multi-processing, quantization, memory allocation optimization, data partitioning, distributed computing, etc.). To address this challenge, we have open sourced BigDL 2.0 at https://github.com/intel-analytics/BigDL/ under Apache 2.0 license (combining the original BigDL [19] and Analytics Zoo [18] projects); using BigDL 2.0, users can simply build conventional Python notebooks on their laptops (with possible AutoML support), which can then be transparently accelerated on a single node (with up-to 9.6x speedup in our experiments), and seamlessly scaled out to a large cluster (across several hundreds servers in real-world use cases). BigDL 2.0 has already been adopted by many real-world users (such as Mastercard, Burger King, Inspur, etc.) in production.
Jason Jinquan Dai, Dongjie Shi, Shengsheng Huang, Xin Qiu 0006, Guoqiong Song, Yang Wang 0009, Qiyuan Gong, Jiaming Song, Shan Yu 0001, Le Zheng, Yina Chen, Junwei Deng
CVPR9
2022 IsoBugView: Interactively Debugging Isolation Bugs in Database Applications
abstract
Database applications frequently use weaker isolation levels, such as Read Committed, for better performance, which may lead to bugs that do not happen under Serializable. Although a number of works have proposed methods to identify such isolation-related bugs, the difficulty of analyzing reported bugs is often underestimated, since these bugs often involve multiple complicated transactions interleaved in a specific order and they often require users' feedback to improve the accuracy of bug analysis. This paper presents IsoBugView, a tool to visualize isolation bugs and incorporate users' feedback: to address the challenge that a complicated bug may include much information and thus is hard to present, IsoBugView displays a high-level overview of the bug first and displays further information of individual pieces if the developer needs further investigation. To incorporate users' feedback, IsoBugView embeds hook functions into the backend analysis tool to preprocess a dependency graph and postprocess a found cycle and further allows a user to apply predefined hook functions in its graphic user interface. Our experience shows that IsoBugView has greatly improved our productivity of analyzing isolation bugs.
Drew Ripberger, Yifan Gan, Xueyuan Ren, Spyros Blanas, Yang Wang 0009
Proc. VLDB Endow.5
2022 A Study of Database Performance Sensitivity to Experiment Settings
abstract
To allow performance comparison across different systems, our community has developed multiple benchmarks, such as TPC-C and YCSB, which are widely used. However, despite such effort, interpreting and comparing performance numbers is still a challenging task, because one can tune benchmark parameters, system features, and hardware settings, which can lead to very different system behaviors. Such tuning creates a long-standing question of whether the conclusion of a work can hold under different settings. This work tries to shed light on this question by reproducing 11 works evaluated under TPC-C and YCSB, measuring their performance under a wider range of settings, and investigating the reasons for the change of performance numbers. By doing so, this paper tries to motivate the discussion about whether and how we should address this problem. While this paper does not give a complete solution---this is beyond the scope of a single paper, it proposes concrete suggestions we can take to improve the state of the art.
Yang Wang 0009, Miao Yu 0023, Yujie Hui, Xueyuan Ren, Tianxi Li, Xiaoyi Lu 0001
Proc. VLDB Endow.1
2022 PostMan: Rapidly Mitigating Bursty Traffic via On-Demand Offloading of Packet Processing
abstract
Unexpected bursty traffic brought by certain sudden events, such as news in the spotlight on a social network or discounted items on sale, can cause severe load imbalance in backend services. Migrating hot data - the standard approach to achieve load balance - meets a challenge when handling such unexpected load imbalance, because migrating data will slow down the server that is already under heavy pressure. This article proposes PostMan, an alternative approach to rapidly mitigate load imbalance for services processing small requests. Motivated by the observation that processing large packets incurs far less CPU overhead than processing small ones, PostMan deploys a number of middleboxes called helpers to assemble small packets into large ones for the heavily-loaded server. This approach essentially offloads the overhead of packet processing from the heavily-loaded server to helpers. To minimize the overhead, PostMan activates helpers on demand, only when bursty traffic is detected. The heavily-loaded server determines when clients connect/disconnect to/from helpers based on the real-time load statistics. To tolerate helper failures, PostMan can migrate connections across helpers and can ensure packet ordering despite such migration. Driven by real-world workloads, our evaluation shows that, with the help of PostMan, a Memcached server can mitigate bursty traffic within hundreds of milliseconds, while migrating data takes tens of seconds and increases the latency during migration.
Yipei Niu, Panpan Jin, Yikai Xiao, Rong Shi, Fangming Liu, Chen Qian 0001, Yang Wang 0009
IEEE Trans. Parallel Distributed Syst.8
2021 Finding heterogeneous-unsafe configuration parameters in cloud systems
abstract
With the increasing prevalence of heterogeneous hardware and the increasing need for online reconfiguration, there is increasing demand for heterogeneous configurations. However, allowing different nodes to have different configurations may cause errors when these nodes communicate, even if the configuration of each node uses valid values.
Sixiang Ma, Michael D. Bond, Yang Wang 0009
EuroSys4
2021 On the Feasibility of Parser-based Log Compression in Large-Scale Cloud Systems
Junyu Wei, Guangyan Zhang, Yang Wang 0009, Zhanyang Zhu, Junchao Chen 0005, Tingtao Sun, Qi Zhou 0001
FAST3
2020 FlatStore: An Efficient Log-Structured Key-Value Storage Engine for Persistent Memory
abstract
Emerging hardware like persistent memory (PM) and high-speed NICs are promising to build efficient key-value stores. However, we observe that the small-sized access pattern in key-value stores doesn't match with the persistence granularity in PMs, leaving the PM bandwidth underutilized. This paper proposes an efficient PM-based key-value storage engine named FlatStore. Specifically, it decouples the role of a KV store into a persistent log structure for efficient storage and a volatile index for fast indexing. Upon it, FlatStore further incorporates two techniques: 1) compacted log format to maximize the batching opportunity in the log; 2) pipelined horizontal batching to steal log entries from other cores when creating a batch, thus delivering low-latency and high-throughput performance. We implement FlatStore with the volatile index of both a hash table and Masstree. We deploy FlatStore on Optane DC Persistent Memory, and our experiments show that FlatStore achieves up to 35 Mops/s with a single server node, 2.5 - 6.3 times faster than existing systems.
Youmin Chen, Youyou Lu, Fan Yang 0134, Qing Wang 0031, Yang Wang 0009, Jiwu Shu
ASPLOS5
2020 IsoDiff: Debugging Anomalies Caused by Weak Isolation
Yifan Gan, Xueyuan Ren, Drew Ripberger, Spyros Blanas, Yang Wang 0009
Proc. VLDB Endow.5
2020 SICS: Secure and Dynamic Middlebox Outsourcing
abstract
There is an increasing trend that enterprises outsource their middlebox processing to a cloud for lower cost and easier management. However, outsourcing middleboxes brings threats to the enterprise's private information, including the traffic and rules of middleboxes, all of which are visible within the cloud. Existing solutions for secure middlebox outsourcing either incur significant performance overhead or do not support incremental updates. In this article, we present a secure and dynamic middlebox outsourcing framework, SICS, short for Secure In-Cloud Service. SICS encrypts each packet header and uses a label for in-cloud rule matching, which enables the cloud to perform its functionalities correctly with minimum header information leakage. Evaluation results show that SICS achieves higher throughput, faster construction and update speed, and lower resource overhead at the enterprise and in the cloud when compared with existing solutions.
Huazhe Wang, Xin Li 0057, Yang Wang 0009, Yu Zhao 0010, Ye Yu 0001, Hongkun Yang, Chen Qian 0001
IEEE/ACM Trans. Netw.3
2019 BigDL: A Distributed Deep Learning Framework for Big Data
abstract
ThispaperpresentsBigDL (adistributeddeeplearning framework for Apache Spark), which has been used by a variety of users in the industry for building deep learning applications on production big data platforms. It allows deep learning applications to run on the Apache Hadoop/Spark cluster so as to directly process the production data, and as a part of the end-to-end data analysis pipeline for deployment and management. Unlike existing deep learning frameworks, BigDL implements distributed, data parallel training directly on top of the functional compute model (with copy-on-write and coarse-grained operations) of Spark. We also share real-world experience and "war stories" of users that havead-optedBigDLtoaddresstheirchallenges(i.e., howtoeasilybuildend-to-enddataanalysisanddeep learning pipelines for their production data).
Jason Jinquan Dai, Xin Qiu 0006, Yanzhang Wang, Xianyan Jia, Cherry Li Zhang, Shengsheng Huang, Zhongyuan Wu, Yang Wang 0009, Bowen She, Dongjie Shi, Guoqiong Song
SoCC14
2019 Henosis: workload-driven small array consolidation and placement for HDF5 applications on heterogeneous data stores
abstract
Scientific data analysis pipelines face scalability bottlenecks when processing massive datasets that consist of millions of small files. Such datasets commonly arise in domains as diverse as detecting supernovae and post-processing computational fluid dynamics simulations. Furthermore, applications often use inference frameworks such as TensorFlow and PyTorch whose naive I/O methods exacerbate I/O bottlenecks. One solution is to use scientific file formats, such as HDF5 and FITS, to organize small arrays in one big file. However, storing everything in one file does not fully leverage the heterogeneous data storage capabilities of modern clusters.
Donghe Kang, Vedang Patel, Ashwati Nair, Spyros Blanas, Yang Wang 0009, Srinivasan Parthasarathy 0001
ICS5
2019 PostMan: Rapidly Mitigating Bursty Traffic by Offloading Packet Processing
Panpan Jin, Yikai Xiao, Rong Shi, Yipei Niu, Fangming Liu, Chen Qian 0001, Yang Wang 0009
USENIX ATC8
2019 Dayu: Fast and Low-interference Data Recovery in Very-large Storage Systems
Zhufan Wang, Guangyan Zhang, Yang Wang 0009, Qinglin Yang, Jiaji Zhu
USENIX ATC3
2019 Redio: Accelerating Disk-Based Graph Processing by Reducing Disk I/Os
abstract
Disk-based graph systems store part or all of graph data on external devices like hard drives or SSDs, achieving scalability without excessive hardware. However, massive expensive disk I/Os remain the major performance bottleneck of disk-based graph processing. In this paper, we propose Redio, a new approach to accelerating disk-based graph processing by reducing disk I/Os. First, Redio observes that it is feasible to accommodate all vertex states in main memory and this can eliminate almost all vertex-related disk I/Os. Second, Redio introduces a dynamic selective scheduling scheme to identify inactive edges in each iteration and skip them when and only when such skipping can bring performance benefit. To improve its effectiveness, Redioin corporates a compact edge storage to improve data locality and an indexed bitmap to minimize its memory and computation overheads. We have implemented a single-node prototype for Redio under the edge-centric computation model. Extensive experiments show that Redio consistently outperforms well-known edge-centric disk-based systems in all experiments, delivering an average speedup of$4.33\times$on HDDs and$5.33\times$on SSDs over the fastest among them (i.e., GridGraph). Experimental results also show that Redio delivers an average speedup of$3.13\times$on HDDs and$1.28\times$on SSDs over the fastest among representative vertex-centric disk-based systems (i.e., FlashGraph).
Chengwen Wu, Guangyan Zhang, Yang Wang 0009, Xinyang Jiang
IEEE Trans. Computers3
2018 Evaluating Scalability Bottlenecks by Workload Extrapolation
abstract
Testing a scalability bottleneck requires a large system to generate sufficient load, which is usually not accessible to researchers. To address this problem, this paper extrapolates the workload to a bottleneck node. The key observation that motivates our approach is that systems at a large scale are often repeating their behaviors at small scales, by running a job more times, running more nodes of the same type, or running more iterations of the same loop. Following this observation, we record a node's workloads at small scales and extrapolate such workload at a large scale. Towards this goal, we have developed PatternMiner, a semi-automatic tool to identify how workload patterns change with scale. We have tested our method on HDFS NameNode and YARN's Resource Manager. Our evaluation shows that PatternMiner is able to predict 98% of the workloads for NameNode and 83% of the workloads for the Resource Manager. Furthermore, by utilizing the extrapolated workload, we are able to emulate a cluster of up to 60,000 nodes with only 8 physical machines to evaluate NameNode and Resource Manager.
Rong Shi, Yifan Gan, Yang Wang 0009
MASCOTS3
2018 wPerf: Generic Off-CPU Analysis to Identify Bottleneck Waiting Events
Yifan Gan, Sixiang Ma, Yang Wang 0009
OSDI4
2018 Accurate Timeout Detection Despite Arbitrary Processing Delays
Sixiang Ma, Yang Wang 0009
USENIX ATC2
2016 Cheap and Available State Machine Replication
Rong Shi, Yang Wang 0009
USENIX ATC2
2015 High-performance ACID via modular concurrency control
abstract
This paper describes the design, implementation, and evaluation of Callas, a distributed database system that offers to unmodified, transactional ACID applications the opportunity to achieve a level of performance that can currently only be reached by rewriting all or part of the application in a BASE/NoSQL Style. The key to combining performance and ease of programming is to decouple the ACID abstraction---which Callas offers identically for all transactions---from the mechanism used to support it. MCC, the new Modular approach to Concurrency Control at the core of Callas, makes it possible to partition transactions in groups with the guarantee that, as long as the concurrency control mechanism within each group upholds a given isolation property, that property will also hold among transactions in different groups. Because of their limited and specialized scope, these group-specific mechanisms can be customized for concurrency with unprecedented aggressiveness. In our MySQL Cluster-based prototype, Callas yields an 8.2x throughput gain for TPC-C with no programming effort.
Chunzhi Su, Cody Littley, Lorenzo Alvisi, Manos Kapritsos, Yang Wang 0009
SOSP6
2014 Exalt: Empowering Researchers to Evaluate Large-Scale Storage Systems
Yang Wang 0009, Manos Kapritsos, Lara Schmidt, Lorenzo Alvisi, Michael Dahlin
NSDI1
2014 Salt: Combining ACID and BASE in a Distributed Database
Chunzhi Su, Manos Kapritsos, Yang Wang 0009, Navid Yaghmazadeh, Lorenzo Alvisi, Prince Mahajan
OSDI4
2014 Lazy Means Smart: Reducing Repair Bandwidth Costs in Erasure-coded Distributed Storage
abstract
Erasure coding schemes provide higher durability at lower storage cost, and thus constitute an attractive alternative to replication in distributed storage systems, in particular for storing rarely accessed "cold" data. These schemes, however, require an order of magnitude higher recovery bandwidth for maintaining a constant level of durability in the face of node failures. In this paper we propose lazy recovery, a technique to reduce recovery bandwidth demands down to the level of replicated storage. The key insight is that a careful adjustment of recovery rate substantially reduces recovery bandwidth, while keeping the impact on read performance and data durability low. We demonstrate the benefits of lazy recovery via extensive simulation using a realistic distributed storage configuration and published component failure parameters. For example, when applied to the commonly used RS(14, 10) code, lazy recovery reduces repair bandwidth by up to 76% even below replication, while increasing the amount of degraded stripes by 0.1 percentage points. Lazy recovery works well with a variety of erasure coding schemes, including the recently introduced bandwidth efficient codes, achieving up to a factor of 2 additional bandwidth savings.
Mark Silberstein, Lakshmi Ganesh, Yang Wang 0009, Lorenzo Alvisi, Michael Dahlin
SYSTOR3
2013 Robustness in the Salus Scalable Block Store
Yang Wang 0009, Manos Kapritsos, Zuocheng Ren, Prince Mahajan, Jeevitha Kirubanandam, Lorenzo Alvisi, Michael Dahlin
NSDI1
2012 All about Eve: Execute-Verify Replication for Multi-Core Servers
Manos Kapritsos, Yang Wang 0009, Vivien Quéma, Allen Clement, Lorenzo Alvisi, Michael Dahlin
OSDI2
2012 Gnothi: Separating Data and Metadata for Efficient and Available Storage Replication
Yang Wang 0009, Lorenzo Alvisi, Michael Dahlin
USENIX ATC1
2010 SOPA: Selecting the optimal caching policy adaptively
abstract
With the development of storage technology and applications, new caching policies are continuously being introduced. It becomes increasingly important for storage systems to be able to select the matched caching policy dynamically under varying workloads. This article proposes SOPA, a cache framework to adaptively select the matched policy and perform policy switches in storage systems. SOPA encapsulates the functions of a caching policy into a module, and enables online policy switching by policy reconstruction. SOPA then selects the policy matched with the workload dynamically by collecting and analyzing access traces. To reduce the decision-making cost, SOPA proposes an asynchronous decision making process. The simulation experiments show that no single caching policy performed well under all of the different workloads. With SOPA, a storage system could select the appropriate policy for different workloads. The real-system evaluation results show that SOPA reduced the average response time by up to 20.3% and 11.9% compared with LRU and ARC, respectively.
Yang Wang 0009, Jiwu Shu, Guangyan Zhang, Wei Xue 0003
ACM Trans. Storage1
2009 Upright cluster services
abstract
The UpRight library seeks to make Byzantine fault tolerance (BFT) a simple and viable alternative to crash fault tolerance for a range of cluster services. We demonstrate UpRight by producing BFT versions of the Zookeeper lock service and the Hadoop Distributed File System (HDFS). Our design choices in UpRight favor simplifying adoption by existing applications; performance is a secondary concern. Despite these priorities, our BFT Zookeeper and BFT HDFS implementations have performance comparable with the originals while providing additional robustness.
Allen Clement, Manos Kapritsos, Yang Wang 0009, Lorenzo Alvisi, Michael Dahlin, Taylor L. Riché
SOSP4