EDBT 2026 Demo / reviewers in the wild / expert
Xiao Zhang 0014
dblp:49/4478-14
· DBLP profile ↗
15ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-7680-1179ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | wdCP: Windowed Incremental Checkpointing for Efficient and Bounded LLM RecoveryabstractCheckpointing is essential for fault tolerance in large-scale LLM training, yet periodic full-state checkpoints bring heavy I/O overhead and training stalls. Prior work suggests that differential checkpointing is ineffective for LLMs, since most parameters updates every iteration, leading to dense updates. This paper revisit this assumption and observe that parameter updates are naturally generated inside optimizer execution and exhibit significant temporal and layer-wise heterogeneity. Guided by this, we present wdCP, a lightweight runtime that captures optimizer-level parameter deltas and asynchronously persists them using a windowed buffering mechanism. wdCP further introduces lightweight anchor snapshots to bound recovery cost. We implement wdCP and evaluate it on several representative models. Results show that wdCP introduces less than 5% training overhead while achieving up to 69.2× reduction in checkpoint size and enabling fast, bounded recovery. Wendi Cheng, Xiao Zhang 0014, Xiaonan Zhao, Xiaoling Shu, Jinjiang Wang, Shujie Han 0001 |
CF | 2 |
| 2025 | BVLSM: Write-Efficient LSM-Tree Storage via WAL-Time Key-Value SeparationabstractModern data-intensive applications increasingly store and process big-value items, such as multimedia objects and machine learning embeddings, which exacerbate storage inefficiencies in Log-Structured Merge-Tree (LSM)-based key-value stores. This paper presents BVLSM, a Write-Ahead Log (WAL)time key-value separation mechanism designed to address three key challenges in LSM-Tree storage systems: write amplification, poor memory utilization, and I/O jitter under big-value workloads. Unlike state-of-the-art approaches that delay key-value separation until the flush stage, leading to redundant data in MemTables and repeated writes. BVLSM proactively decouples keys and values during the WAL phase. The MemTable stores only lightweight metadata, allowing multi-queue parallel store for big value. Benchmark results demonstrate that BVLSM significantly outperforms both RocksDB and BlobDB across various write-intensive workloads. Specifically, in asynchronous WAL mode, its 64KB write throughput exceeds that of RocksDB by a factor of 7.6 and that of BlobDB by a factor of 1.9. Wendi Cheng, Jiahe Wei, Xueqiang Shan, Weikai Liu, Xiaonan Zhao, Xiao Zhang 0014 |
HPCC | 7 |
| 2025 | Scheduling virtual machines and containers: A comparative review of techniques, performance, and future trends
Jiameng Zhang, Ruofei Wu, Taoyu Zhong, Shujie Han 0001, Xiao Zhang 0014, Xiaonan Zhao |
J. Syst. Archit. | 9 |
| 2025 | A Load-Balanced Collaborative Repair Algorithm for Single-Disk Failures in Erasure Coded Storage SystemsabstractIn large-scale cloud data centers and distributed storage systems, erasure coding is usually employed to enhance data availability and storage efficiency. However, with the explosive growth of data volume and the continuous expansion of storage system scale, traditional erasure coding techniques face significant challenges in handling single-disk failures. These challenges are primarily reflected in low data recovery efficiency and imbalanced system load distribution, which ultimately result in excessive I/O load and network bandwidth consumption, severely limiting the overall performance of the system. To address these issues, this article proposes a load-balanced data repair algorithm for single disk failures in erasure coded storage systems, called MNCR (Multi-Node Cooperative Repair). This algorithm improves data recovery efficiency in single-disk failure scenarios by minimizing data reading and inter-disk data transmission, using a cooperative repair strategy among disks. In addition, the algorithm designs a dynamic load balancing mechanism, which effectively resolves the issue of imbalanced data load distribution among disks during the repair process, thus avoiding performance bottlenecks caused by overloaded disks. Experimental results show that the MNCR algorithm significantly outperforms traditional methods in terms of repair efficiency and load balancing, providing an effective solution for single disk failure recoveries in erasure coding based large-scale storage systems. Yulong Shi, Chengjia Zhao, Shujie Han 0001, Xiao Zhang 0014 |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2024 | TraceGen: A Block-level Storage System Performance Evaluation Tool for Analyzing and Generating I/O TracesabstractPerformance measurement is essential for detecting potential performance issues and guiding optimization efforts. However, acquiring I/O traces of real applications can be costly in production environments. Also, existing performance measurement tools, such as FIO and Iometer, often oversimplify real-world application characteristics. In this paper, we introduce TraceGen, a block-level performance measurement tool for storage systems that consists of a trace analyzer and a trace generator. The trace analyzer produces two categories of traces: (i) new traces with specified characteristics designed to accurately simulate a range of applications, and (ii) extended traces that maintain similar workload characteristics to the input traces, thereby improving measurement accuracy during trace replay. We evaluate TraceGen using traces from an enterprise production environment and demonstrate its capability to generate new traces with an error margin of less than 1%. Jiahe Wei, Huiru Xie, Jinjiang Wang, Xiaonan Zhao, Shujie Han 0001, Xiao Zhang 0014 |
HPCC | 7 |
| 2024 | Erasure Coding Based Optimization in Decentralized Distributed Storage SystemsabstractNode failures in decentralized distributed storage systems are common. To ensure data availability, these systems employ data redundancy mechanisms, typically relying on repli-cas. This paper proposes a decentralized data redundancy scheme based on erasure coding, with Reed-Solomon and IPFS as exam-ples. Compared with replication, the proposed scheme reduces storage space and enhances fault tolerance. Files are sharded and encoded across multiple nodes, avoiding the high redundancy of replicas. Users can retrieve any K shards from N nodes to reconstruct the original files. This erasure coding optimization combines efficient data exchange among decentralized nodes with erasure coding technology, significantly reducing storage space compared with the replica mechanism. The implementation involves truncating and sharding the blocks within the Merkle DAG generated by files, enabling flexible adjustments to the code rate of erasure codes and the allocation of storage nodes based on user needs and available resources. This method achieves a balance between storage efficiency and data availability. Yiwei Gan, Yulong Shi, Xiao Zhang 0014 |
NAS | 4 |
| 2024 | Design and Implementation of a Turbulence Data Sharing Platform for Scientific Big DataabstractTurbulence is a three-dimensional fluid motion state with multiple scales and mutual coupling in time and space. Studying turbulence phenomena is crucial for the design and drag reduction of aircraft and engines. In recent years, with the powerful computing power of computers, data-based turbulence research methods have become a hot topic, such as the Direct Numerical Simulation (DNS) method, which is receiving attention. However, compared to Large Eddy Simulation (LES), the amount of data generated by direct numerical simulation methods far exceeds that of traditional LES, which brings difficulties in data storage. Abroad, multiple research institutions in Europe, the United States, and Japan have established multiple turbulence data sharing platforms, but the construction of turbulence data sharing platforms in China is still in a blank stage. This article first analyzes the characteristics of turbulence science big data, and then starts from the practical problems faced by domestic research institutions when sharing and using data, designs and implements the first independently controllable turbulence data sharing platform in China. Currently, the platform has integrated 180TB of data from five domestic universities and research institutions, providing an important resource sharing platform and data research tool for turbulence researchers in China. Youjun Zhao, Xiao Zhang 0014, Wendi Cheng, Zhaohui Pan, Chenguang Sun, Xueqiang Shan |
NAS | 2 |
| 2023 | MyWAL: performance optimization by removing redundant input/output stack in key-value storeabstractBased on a log-structured merge (LSM) tree, the key-value (KV) storage system can provide high reading performance and optimize random writing performance. It is widely used in modern data storage systems like e-commerce, online analytics, and real-time communication. An LSM tree stores new KV data in the memory and flushes to disk in batches. To prevent data loss in memory if there is an unexpected crash, RocksDB appends updating data in the write-ahead log (WAL) before updating the memory. However, synchronous WAL significantly reduces writing performance. In this paper, we present a new WAL mechanism named MyWAL. It directly manages raw devices (or partitions) instead of saving data on a traditional file system. These can avoid useless metadata updating and write data sequentially on disks. Experimental results show that MyWAL can significantly improve the data writing performance of RocksDB compared to the traditional WAL for small KV data on solid-state disks (SSDs), as much as five to eight times faster. On non-volatile memory express soild-state drives (NVMe SSDs) and non-volatile memory (NVM), MyWAL can improve data writing performance by 10%–30%. Furthermore, the results of YCSB (Yahoo! Cloud Serving Benchmark) show that the latency decreased by 50% compared with SpanDB. Xiao Zhang 0014, Michael Ngulube, Yonghao Chen, Yiping Zhao |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2022 | An Adaptive Elastic Multi-model Big Data Analysis and Information Extraction SystemabstractAbstract With the diverse applications to industry and domain-specific context, multi-source information extraction on semi-structured and unstructured data, as well as across data models, is becoming more common. However, multi-model information extraction often requires the deployment of multiple data model management, storage, and analysis subsystems on the cloud, many subsystems are not high-resource utilization at the same time, and the resource waste phenomenon is often serious. Therefore, an adaptive scalable multi-model big data analysis and information extraction system is designed and implemented in this paper, which can support data maintenance and cross-model query of relational, graph, document, key and other data models, and can provide efficient cross-model information extraction. On this basis, we can achieve the system resource allocation on demand and fast scaling mechanism, according to the real-time requirements of multi-model big data analysis, and dynamic adjustment of each subsystem resource allocation. Therefore, our solution not only guarantees multi-model query and information extraction performance and quality of service, but also significantly reduces the total consumption of system resources and cost. Qiang Yin 0004, Sheng Du, Jianquan Leng, Yinhao Hong, Feng Zhang 0007, Yunpeng Chai, Xiao Zhang 0014, Xiaonan Zhao, Wei Lu 0015 |
Data Sci. Eng. | 9 |
| 2021 | WOBTree: a write-optimized B+-tree for non-volatile memory
Zhanhuai Li, Xiao Zhang 0014, Xiaonan Zhao, Song Jiang 0001 |
Frontiers Comput. Sci. | 3 |
| 2018 | OC-Cache: An Open-channel SSD Based Cache for Multi-Tenant SystemsabstractIn a multi-tenant cloud environment, tenants are usually hosted by virtual machines. Cloud providers deploy multiple virtual machines on a physical server to better utilize physical resources including CPU, memory, and storage devices. SSDs are often used as an I/O cache shared among the tenants for large storage systems using hard disk drives (HDDs) as their main storage devices, which can receive much of SSD's performance benefit and HDD's cost advantage. A key challenge in the use of the shared cache is to ensure strong performance isolation and maintain its high utilization at the same time. However, conventional SSD cache management approaches cannot effectively address this challenge. In this paper, we propose OC-Cache, an open-channel SSD cache framework which utilizes SSD'd internal parallelism to adaptively allocate cache to tenants for both good performance isolation and high SSD utilization. In particular, OC-Cache uses a tenant's miss ratio curve to determine the amount of cache space allocation and where the allocation is (in dedicated or shared SSD channels) and dynamically manages cache space according to the workload characteristics. Experiments show that OC-Cache significantly reduces interference among tenants, and maintains high utilization of the SSD cache. Zhanhuai Li, Xiao Zhang 0014, Xiaonan Zhao, Xingsheng Zhao, Song Jiang 0001 |
IPCCC | 3 |
| 2018 | FSObserver: A Performance Measurement and Monitoring Tool for Distributed Storage Systems
Xiao Zhang 0014, Lanxin Kong, Shunyi Zhu, Zhanhuai Li, Xiaonan Zhao |
NPC | 1 |
| 2017 | A Hash-Based Space-Efficient Page-Level FTL for Large-Capacity SSDsabstractWith increasing demands on high-performance and large-capacity SSDs in the enterprise-scale storage, the concern about the inefficient use of the DRAM space in SSDs rises, especially for those using page-level FTL (Flash Translation Layer). In such an FTL, the address mapping scheme allows a logical page address (LPA) to be mapped to any physical page address (PPA) in the disk. Though it provides flexible address management and minimizes internal data movements, it requires a large address mapping table whose size is proportional to the capacity of the disk. With the increase of SSD's capacity, the table can be too large to be held entirely in the DRAM buffer of the SSD, causing constantly accessing to the flash for the address translation. This performance penalty due to the buffer misses is particularly high with workloads of weak access locality and large working sets. In this paper, we propose a space- efficient page- level FTL using hash functions in the address translation, named Hash-based Page- level FTL, or HP-FTL in short, to address the concern. HP-FTL trades mapping flexibility with limited performance impact for high space efficiency allowing the entire table to fit in the buffer and eliminating translation misses. The experiment results show that HP-FTL can provide up to 2.6X throughput compared to DFTL, a representative page-level FTL, using the same amount of DRAM for buffering the table. Meanwhile, HP-FTL reduces the mapping table size to about 25% of the table space required by page- level mapping schemes, including DFTL, without having any buffer misses. Fan Ni, Chunyi Liu, Yang Wang 0006, Cheng-Zhong Xu 0001, Xiao Zhang 0014, Song Jiang 0001 |
NAS | 5 |
| 2017 | Freewrite: creating (almost) zero-cost writes to SSD in applicationsabstractWhile flash-based SSDs have much higher access speed than hard disks, they have an Achilles heel, which is the service of write requests. Not only is writing slower than reading, but also it can incur expensive garbage collection operations and reduce SSDs' lifetime. The deduplication technique can help to avoid writing data objects whose contents have been on the disk. A typical object is the disk block, for which a block-level deduplication scheme can help identify duplicate ones and avoid their writing. For the technique to be effective, data written to the disk must not only be the same as those currently on the disk but also be block-aligned. Chunyi Liu, Fan Ni, Xingbo Wu, Xiao Zhang 0014, Song Jiang 0001 |
SYSTOR | 4 |
| 2013 | BFEPM: Best Fit Energy Prediction Modeling Based on CPU UtilizationabstractEnergy cost becomes a major part of data center operational cost. Computer system consume more power when it runs under high workload. Many past studies focused on how to predict power consumption by performance counters. Some models retrieve performance counters from chips. Some models query performance counters from OS. Most of these researches were verified on several machines and claimed their models were accurate under the test. We found different servers have different energy consumption characters even with same CPU. In this paper, we present BFEPM, a best fit energy prediction model. It choose best model based on the power consumption benchmark result. We illustrate how to use benchmark result to find a best fit model. Then we validate the viability and effectiveness of model on all published results. At last, we apply the best fit model on two different machines to estimate the real-time energy consumption. The results show our model can get better results than single model. Xiao Zhang 0014, Jian-Jun Lu, Xiao Qin 0001 |
NAS | 1 |