EDBT 2026 Demo / reviewers in the wild / expert
Hao Wen 0001
dblp:57/77-1
· DBLP profile ↗
11ranked-venue papers
4as first author
4since 2021 · last 2023
0000-0003-3516-4454ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | SMRTS: A Performance and Cost-Effectiveness Optimized SSD-SMR Tiered File System with Data DeduplicationabstractStorage tiering (e.g., SSD+HDD) is designed to achieve a better tradeoff between performance and cost-effectiveness for storage systems. With the development of Shingled Magnetic Recording (SMR) drives, replacing conventional HDD with a higher density of SMR drives in tiered storage can further improve cost-effectiveness. However, with data tracks overlapped in SMR drives, the "non-sequential' writes in SMR drives cause explicit performance penalties, which is the most challenging issue of using SMR drives in storage tiering.In this paper, we present SMRTS, a file system for SSD-SMR tiered storage with data deduplication. First, SMRTS deduplicates the files being migrated from SSD to SMR to solve the non-sequential write issue of SMR drives and further optimize the space utilization. Second, to address the performance overhead caused by deduplication, we propose file recipe reuse and refresh, hints-based container allocations, and fast container validation to address the penalties caused by data fragmentations. We conduct experimental evaluations of SMRTS using both benchmarks and real-world workloads. The evaluation results show that compared with a compatible file system on SSD+HDD tiered storage, SMRTS achieves a similar performance but provides a much larger space (at least 1.25X). The proposed optimizations improve migration performance up to 17X. Zhichao Cao 0002, Hao Wen 0001, Fenggang Wu, David Hung-Chang Du |
ICCD | 2 |
| 2023 | K8sES: Optimizing Kubernetes with Enhanced Storage Service-Level ObjectivesabstractKubernetes (k8s) is a system for managing containerized applications across multiple hosts. It offers automatic deployment, maintenance, scaling, and resource management for applications. Applications in k8s usually have different storage requirements in the form of service-level objectives (SLOs). However, the current k8s storage management has several limitations which cause explicit performance and cost overhead. K8s administrators have to configure storage in advance manually, and users must know configurations and capabilities of provided storage. Users' storage SLOs can be easily violated in k8s.In this paper, we design and implement k8s Enhanced Storage (k8sES) which efficiently supports applications with various storage SLOs along with all other requirements in the Kubernetes environment. We design and incorporate storage scheduling as part of the node scheduling process in k8s. Applications will be scheduled onto the correct nodes and storage without intervention from either users or administrators. Proper storage resources will be dynamically carved based on users' storage SLOs. In addition, we provide a tool to monitor the I/O activities of both applications and storage devices in k8sES. The evaluation shows that k8sES can better meet users' storage SLOs along with other requirements. Also, k8sES can achieve higher resource utilization efficiency with overhead similar to that of the current k8s. Hao Wen 0001, Zhichao Cao 0002, Bingzhe Li, David Hung-Chang Du, Ayman Abouelwafa, Doug Voigt, Shiyong Liu, Jim Diehl, Fenggang Wu |
ICCD | 1 |
| 2021 | TrackLace: Data Management for Interlaced Magnetic RecordingabstractInterlaced Magnetic Recording (IMR) is a promising technology which achieves higher data density and lower write amplification (WA) than Shingled Magnetic Recording (SMR). In IMR, top tracks and bottom tracks are interlaced so each bottom track is partially overlapped with two adjacent top tracks. Top tracks can be updated without any WA, but bottom track updates require reading and rewriting of affected valid data on the two neighboring top tracks. There are few published studies discussing WA in IMR drives. We propose TrackLace to reduce WA for IMR. TrackLace consists of three techniques: Z-Alloc allocates user data to the tracks in alternating directions and spreads unallocated tracks among allocated tracks; Top-Buffer opportunistically utilizes unallocated top tracks to buffer bottom track updates; and Block-Swap progressively swaps bottom track hot data with top track cold data during high space utilization. To further optimize TrackLace performance, we propose a virtual frame design that can keep the relocated block (due to Top-Buffer or Block-Swap) close to its original location and an adaptive buffering mechanism that can avoid unnecessary redirections depending on the write locality. Evaluations show that TrackLace can reduce WA by 45 percent and lower average latency by 31percent compared with baseline schemes. Fenggang Wu, Bingzhe Li, Baoquan Zhang, Zhichao Cao 0002, Jim Diehl, Hao Wen 0001, David Hung-Chang Du |
IEEE Trans. Computers | 6 |
| 2021 | Guaranteed Bang for the Buck: Modeling VDI Applications to Identify Storage RequirementsabstractIn the cloud environment, most services are provided by virtual machines (VMs). Identifying storage requirements of VMs is challenging, but it is essential for good user experiences while optimizing use of storage resources. Determining the storage configuration necessary to support and satisfy VMs first requires an accurate description of the VM configurations, and the problem is further exacerbated by the diversity and special characteristics of the VMs. In this paper, we study Virtual Desktop Infrastructure (VDI), a prevalent and complicated VM application, to identify and characterize storage requirements of VMs and determine how to meet such requirements with minimal storage resources and cost. We first create a model to describe the behavior of VDI, and we collect real VDI traces to populate this model. The model allows us to identify the storage requirements of VDI and determine the potential bottlenecks of a given storage configuration. Based on this information, we can tell what capacity and minimum capability a storage configuration needs in order to support and satisfy given VDI configurations. We show that our model can describe more fine-grained VM behavior varying with time and virtual disk types compared with the rules of thumb currently used in industry. Hao Wen 0001, David Hung-Chang Du, Milan Shetti, Doug Voigt, Shanshan Li 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2019 | ZoneAlloy: Elastic Data and Space Management for Hybrid SMR Drives
Fenggang Wu, Bingzhe Li, Zhichao Cao 0002, Baoquan Zhang, Ming-Hong Yang, Hao Wen 0001, David Hung-Chang Du |
HotStorage | 6 |
| 2019 | NetStorage: A synchronized trace-driven replayer for network-storage system evaluation
Bingzhe Li, Hao Wen 0001, Farnaz Toussi, Clark Anderson, Bernard A. King-Smith, David J. Lilja, David Hung-Chang Du |
Perform. Evaluation | 2 |
| 2019 | TDDFS: A Tier-Aware Data Deduplication-Based File SystemabstractWith the rapid increase in the amount of data produced and the development of new types of storage devices, storage tiering continues to be a popular way to achieve a good tradeoff between performance and cost-effectiveness. In a basic two-tier storage system, a storage tier with higher performance and typically higher cost (the fast tier) is used to store frequently-accessed (active) data while a large amount of less-active data are stored in the lower-performance and low-cost tier (the slow tier). Data are migrated between these two tiers according to their activity. In this article, we propose a Tier-aware Data Deduplication-based File System, called TDDFS, which can operate efficiently on top of a two-tier storage environment. Specifically, to achieve better performance, nearly all file operations are performed in the fast tier. To achieve higher cost-effectiveness, files are migrated from the fast tier to the slow tier if they are no longer active, and this migration is done with data deduplication. The distinctiveness of our design is that it maintains the non-redundant (unique) chunks produced by data deduplication in both tiers if possible. When a file is reloaded (called a reloaded file) from the slow tier to the fast tier, if some data chunks of the file already exist in the fast tier, then the data migration of these chunks from the slow tier can be avoided. Our evaluation shows that TDDFS achieves close to the best overall performance among various file-tiering designs for two-tier storage systems. Zhichao Cao 0002, Hao Wen 0001, Xiongzi Ge, Jim Diehl, David Hung-Chang Du |
ACM Trans. Storage | 2 |
| 2018 | ALACC: Accelerating Restore Performance of Data Deduplication Systems Using Adaptive Look-Ahead Window Assisted Chunk Caching
Zhichao Cao 0002, Hao Wen 0001, Fenggang Wu, David Hung-Chang Du |
FAST | 2 |
| 2018 | Data Management Design for Interlaced Magnetic Recording
Fenggang Wu, Baoquan Zhang, Zhichao Cao 0002, Hao Wen 0001, Bingzhe Li, Jim Diehl, David Hung-Chang Du |
HotStorage | 4 |
| 2018 | JoiNS: Meeting Latency SLO with Integrated Control for Networked StorageabstractMeeting latency SLOs (Service Level Objectives) in a networked storage environment is essential while challenging. In this environment, a storage request has to go through client I/O stacks, dynamically changing networks, and the storage system attached to a server. Its response also has to traverse all the way back to the client. Along this long I/O path, any of these components can become congested. The behavior of one component may affect the performance of the others. Isolated control on each component is not effective to meet latency SLOs of storage requests. In this paper, we propose and implement JoiNS, a system trying to guarantee latency SLO for applications that access data on a remote networked storage. JoiNS carefully considers all the components along the I/O path and controls them in a coordinated fashion. JoiNS has both global network and storage visibilities with a logically centralized controller which keeps monitoring the status of each involved component. JoiNS coordinates these components and adjusts the priority of I/O packets in each component based on the latency SLO, network and storage status, time estimation, and characteristics of each I/O request. We integrate Software Defined Network(SDN) into our system to coordinate with storage. Our evaluation shows JoiNS can achieve up to 6X speedup in this networked storage environment with various loads of background traffic. Hao Wen 0001, Zhichao Cao 0002, Ziqi Fan, Doug Voigt, David Hung-Chang Du |
MASCOTS | 1 |
| 2016 | Guaranteed Bang for the Buck: Modeling VDI Applications with Guaranteed Quality of ServiceabstractIn cloud environment, most services are provided by virtual machines (VMs). Providing storage quality of service (QoS) for VMs is essential to user experiences while challenging. It first requires an accurate estimate and description of VM requirements, however, people usually describe this via rules of thumb. The problems are exacerbated by the diversity and special characteristics of VMs in a computing environment. This paper chooses Virtual Desktop Infrastructure (VDI), a prevalent and complicated VM application, to characterize QoS requirements of VMs and to guarantee QoS with minimal required resources. We create a model to describe QoS requirements of VDI. We have collected real VDI traces from HP to validate the correctness of the model. Then we generate QoS requirements of VDI and determine bottlenecks. Based on this, we can tell what minimum capability a storage appliance needs in order to satisfy a given VDI configuration and QoS requirements. By comparing with industry experience, we validate our model. And our model can describe more fine-grained VM requirements varying with time and virtual disk types, and provide more confidence on sizing storage for VDI as well. Hao Wen 0001, David Hung-Chang Du, Milan Shetti, Doug Voigt, Shanshan Li 0001 |
ICPP | 1 |