Shujie Han 0001

dblp:207/3499 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-5311-5782ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 7 since 2021Security and privacy · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 wdCP: Windowed Incremental Checkpointing for Efficient and Bounded LLM Recovery
abstract
Checkpointing is essential for fault tolerance in large-scale LLM training, yet periodic full-state checkpoints bring heavy I/O overhead and training stalls. Prior work suggests that differential checkpointing is ineffective for LLMs, since most parameters updates every iteration, leading to dense updates. This paper revisit this assumption and observe that parameter updates are naturally generated inside optimizer execution and exhibit significant temporal and layer-wise heterogeneity. Guided by this, we present wdCP, a lightweight runtime that captures optimizer-level parameter deltas and asynchronously persists them using a windowed buffering mechanism. wdCP further introduces lightweight anchor snapshots to bound recovery cost. We implement wdCP and evaluate it on several representative models. Results show that wdCP introduces less than 5% training overhead while achieving up to 69.2× reduction in checkpoint size and enabling fast, bounded recovery.
Wendi Cheng, Xiao Zhang 0014, Xiaonan Zhao, Xiaoling Shu, Jinjiang Wang, Shujie Han 0001
CF6
2025 Scheduling virtual machines and containers: A comparative review of techniques, performance, and future trends
Jiameng Zhang, Ruofei Wu, Taoyu Zhong, Shujie Han 0001, Xiao Zhang 0014, Xiaonan Zhao
J. Syst. Archit.7
2025 A Load-Balanced Collaborative Repair Algorithm for Single-Disk Failures in Erasure Coded Storage Systems
abstract
In large-scale cloud data centers and distributed storage systems, erasure coding is usually employed to enhance data availability and storage efficiency. However, with the explosive growth of data volume and the continuous expansion of storage system scale, traditional erasure coding techniques face significant challenges in handling single-disk failures. These challenges are primarily reflected in low data recovery efficiency and imbalanced system load distribution, which ultimately result in excessive I/O load and network bandwidth consumption, severely limiting the overall performance of the system. To address these issues, this article proposes a load-balanced data repair algorithm for single disk failures in erasure coded storage systems, called MNCR (Multi-Node Cooperative Repair). This algorithm improves data recovery efficiency in single-disk failure scenarios by minimizing data reading and inter-disk data transmission, using a cooperative repair strategy among disks. In addition, the algorithm designs a dynamic load balancing mechanism, which effectively resolves the issue of imbalanced data load distribution among disks during the repair process, thus avoiding performance bottlenecks caused by overloaded disks. Experimental results show that the MNCR algorithm significantly outperforms traditional methods in terms of repair efficiency and load balancing, providing an effective solution for single disk failure recoveries in erasure coding based large-scale storage systems.
Yulong Shi, Chengjia Zhao, Shujie Han 0001, Xiao Zhang 0014
ACM Trans. Embed. Comput. Syst.6
2024 TraceGen: A Block-level Storage System Performance Evaluation Tool for Analyzing and Generating I/O Traces
abstract
Performance measurement is essential for detecting potential performance issues and guiding optimization efforts. However, acquiring I/O traces of real applications can be costly in production environments. Also, existing performance measurement tools, such as FIO and Iometer, often oversimplify real-world application characteristics. In this paper, we introduce TraceGen, a block-level performance measurement tool for storage systems that consists of a trace analyzer and a trace generator. The trace analyzer produces two categories of traces: (i) new traces with specified characteristics designed to accurately simulate a range of applications, and (ii) extended traces that maintain similar workload characteristics to the input traces, thereby improving measurement accuracy during trace replay. We evaluate TraceGen using traces from an enterprise production environment and demonstrate its capability to generate new traces with an error margin of less than 1%.
Jiahe Wei, Huiru Xie, Jinjiang Wang, Xiaonan Zhao, Shujie Han 0001, Xiao Zhang 0014
HPCC6
2024 Enhancing LSM-Tree Key-Value Stores for Read-Modify-Writes via Key-Delta Separation
abstract
Read-modify-writes (RMWs) are increasingly observed in practical key-value (KV) storage workloads to support fine-grained updates. To make RMWs efficient, one approach is to write deltas (i.e., changes to current values) to the log-structured merge-tree (LSM-tree), yet it increases the read overhead caused by retrieving and combining a chain of deltas. We propose a notion called key-delta (KD) separation to support efficient reads and RMWs in LSM-tree KV stores under RMW-intensive workloads. KD separation aims to store deltas in separate storage areas and group deltas into storage units called buckets, such that all deltas of a key are kept in the same bucket and can be all together accessed in subsequent reads. To this end, we build KDSep, a middleware layer that realizes KD separation and integrates KDSep into state-of-the-art LSM-tree KV stores (e.g., RocksDB and BlobDB). We show that KDSep achieves significant I/O throughput gains and read latency reduction under RMW-intensive workloads while preserving the efficiency in general workloads.
Yanjing Ren, Shujie Han 0001, Patrick P. C. Lee
ICDE3
2024 PP-Stream: Toward High-Performance Privacy-Preserving Neural Network Inference via Distributed Stream Processing
abstract
Privacy preservation is critical for neural network inference, which often involves collaborative execution of different parties to make predictions on sensitive data based on sensitive neural network models. However, the expensive cryptographic operations of privacy preservation also pose performance chal-lenges to neural network inference. We address this performance-security tension by designing PP-Stream, a distributed stream processing system for high-performance privacy-preserving neural network inference. PP-Stream adopts hybrid privacy-preserving mechanisms for linear and non-linear operations of neural network inference. It treats inference data as real-time data streams, and parallelizes the inference operations across multiple pipelined stages that are executed by multiple servers and threads. It also solves the load-balanced resource allocation across servers and threads as an optimization problem. We prototype PP-Stream and show via testbed experiments that it achieves low inference latencies on various neural network models.
Qingxiu Liu, Qun Huang 0001, Xiang Chen 0017, Sa Wang, Shujie Han 0001, Patrick P. C. Lee
ICDE6
2023 StreamDFP: A General Stream Mining Framework for Adaptive Disk Failure Prediction
abstract
We explore machine learning for accurately predicting imminent disk failures and hence providing proactive fault tolerance for modern large-scale storage systems. Current disk failure prediction approaches are mostly offline and assume that the disk logs required for training learning models are available a priori. However, disk logs are often continuously generated as an evolving data stream, in which the statistical patterns vary over time (also known as concept drift). Such a challenge motivates the need of online techniques that perform training and prediction on the incoming stream of disk logs in real time, while being adaptive to concept drift. We first measure and demonstrate the existence of concept drift on various disk models in production. Motivated by our study, we designStreamDFP, a general stream mining framework for disk failure prediction with concept-drift adaptation based on three key techniques, namely online labeling, concept-drift-aware training, and general prediction, with a primary objective of supporting various machine learning algorithms. We extendStreamDFPto support online transfer learning for minority disk models with concept-drift adaptation. Our evaluation shows thatStreamDFPimproves the prediction accuracy significantly compared to without concept-drift adaptation under various settings, and achieves reasonably high stream processing performance.
Shujie Han 0001, Patrick P. C. Lee, Zhirong Shen
IEEE Trans. Computers1
2022 An In-Depth Correlative Study Between DRAM Errors and Server Failures in Production Data Centers
abstract
Dynamic Random Access Memory (DRAM) errors are prevalent and lead to server failures in production data centers. However, little is known about the correlation between DRAM errors and server failures in state-of-the-art field studies on DRAM error measurement. To fill this void, we present an in-depth data-driven correlative analysis between DRAM errors and server failures, with the primary goal of predicting server failures based on DRAM error characterization and hence enabling proactive reliability maintenance for production data centers. Our analysis is based on an eight-month dataset collected from over three million memory modules in the production data centers at Alibaba. We find that the correctable DRAM errors of most server failures only manifest within a short time before the failures happen, implying that server failure prediction should be conducted regularly at short time intervals for accurate prediction. We also study various impacting factors (including component failures in the memory subsystem, DRAM configurations, types of correctable DRAM errors) on server failures. Furthermore, we design a machine-learning-based server failure prediction workflow and demonstrate the feasibility of server failure prediction based on DRAM error characterization. To this end, we report 14 findings from our measurement and prediction studies.
Zhinan Cheng, Shujie Han 0001, Patrick P. C. Lee, Jiongzhou Liu
SRDS2
2021 General Feature Selection for Failure Prediction in Large-scale SSD Deployment
abstract
Solid-state drive (SSD) failures are likely to cause system-level failures leading to downtime, enabling SSD failure prediction to be critical to large-scale SSD deployment. Existing SSD failure prediction studies are mostly based on customized SSDs with proprietary monitoring metrics, which are difficult to reproduce. To support general SSD failure prediction of different drive models and vendors, this paper proposes Wear-out-updating Ensemble Feature Ranking (WEFR) to select the SMART attributes as learning features in an automated and robust manner. WEFR combines different feature ranking results and automatically generates the final feature selection based on the complexity measures and the change point detection of wear-out degrees. We evaluate our approach using a dataset of nearly 500K working SSDs at Alibaba. Our results show that the proposed approach is effective and outperforms related approaches. We have successfully applied the proposed approach to improve the reliability of cloud storage systems in production SSD-based data centers. We release our dataset for public use.
Shujie Han 0001, Patrick P. C. Lee, Jiongzhou Liu
DSN2
2021 An In-Depth Study of Correlated Failures in Production SSD-Based Data Centers
Shujie Han 0001, Patrick P. C. Lee, Jiongzhou Liu
FAST1
2020 Toward Adaptive Disk Failure Prediction via Stream Mining
abstract
We explore machine learning for accurately predicting imminent disk failures and hence providing proactive fault tolerance for modern storage systems. Current disk failure prediction approaches are mostly offline and assume that the disk logs required for training learning models are available a priori. However, in large-scale disk deployment, disk logs are often continuously generated as an evolving data stream, in which the statistical patterns vary over time (also known as concept drift). Such a challenge motivates the need of online techniques that perform training and prediction on the incoming stream of disk logs in real time, while being adaptive to concept drift.We present StreamDFP, a general stream mining framework for disk failure prediction with concept-drift adaptation. We start with a measurement study and demonstrate the existence of concept drift on various disk models based on the datasets from Backblaze and Alibaba Cloud. Motivated by our study, we design StreamDFP with three key techniques, namely (i) online labeling, (ii) concept-drift-aware training, and (iii) general prediction, with a primary objective of making StreamDFP support various machine learning algorithms as a general frame-work. Our evaluation shows that StreamDFP improves the prediction accuracy significantly compared to without concept-drift adaptation under various settings, and achieves reasonably high stream processing performance.
Shujie Han 0001, Patrick P. C. Lee, Zhirong Shen
ICDCS1
2019 SimEDC: A Simulator for the Reliability Analysis of Erasure-Coded Data Centers
abstract
Modern data centers employ erasure coding to protect data storage against failures. Given the hierarchical nature of data centers, characterizing the effects of erasure coding and redundancy placement on the reliability of erasure-coded data centers is critical yet unexplored. This paper presents a discrete-event simulator called SimEDC, which enables us to conduct a comprehensive simulation analysis of reliability on erasure-coded data centers. SimEDC reports reliability metrics of an erasure-coded data center based on the configurable inputs of the data center topology, erasure codes, redundancy placement, and failure/repair patterns of different subsystems obtained from statistical models or production traces. It can further accelerate the simulation analysis via importance sampling. Our simulation analysis based on SimEDC shows that placing erasure-coded data in fewer racks generally improves reliability by reducing cross-rack repair traffic, even though it sacrifices rack-level fault tolerance in the face of correlated failures.
Mi Zhang 0007, Shujie Han 0001, Patrick P. C. Lee
IEEE Trans. Parallel Distributed Syst.2
2018 A Simulation Analysis of Redundancy and Reliability in Primary Storage Deduplication
abstract
Deduplication has been widely used to improve storage efficiency in modern primary and secondary storage systems, yet how deduplication fundamentally affects storage system reliability remains debatable. This paper aims to analyze and compare storage system reliability with and without deduplication in primary workloads using public file system snapshots from two research groups. We first study the redundancy characteristics of the file system snapshots. We then propose a trace-driven, deduplication-aware simulation framework to analyze data loss in both chunk and file levels due to sector errors and whole-disk failures. Compared to without deduplication, our analysis shows that deduplication consistently reduces the damage of sector errors due to intra-file redundancy elimination, but potentially increases the damages of whole-disk failures if the highly referenced chunks are not carefully placed on disk. To improve reliability, we examine a deliberate copy technique that stores and repairs first the most referenced chunks in a small dedicated physical area (e.g., 1 percent of the physical capacity), and demonstrate its effectiveness through our simulation framework.
Min Fu 0002, Shujie Han 0001, Patrick P. C. Lee, Dan Feng 0001, Zuoning Chen
IEEE Trans. Computers2
2017 A Simulation Analysis of Reliability in Erasure-Coded Data Centers
abstract
Erasure coding has been widely adopted to protect data storage against failures in production data centers. Given the hierarchical nature of data centers, characterizing the effects of erasure coding and redundancy placement on the reliability of erasure-coded data centers is critical yet largely unexplored. This paper presents a comprehensive simulation analysis of reliability on erasure-coded data centers. We conduct the analysis by building a discrete-event simulator called SIMEDC, which reports reliability metrics of an erasure-coded data center based on the configurable inputs of the data center topology, erasure codes, redundancy placement, and failure/repair patterns of different subsystems obtained from statistical models or production traces. Our simulation results show that placing erasure-coded data in fewer racks generally improves reliability by reducing cross-rack repair traffic, even though it sacrifices rack-level fault tolerance in the face of correlated failures.
Mi Zhang 0007, Shujie Han 0001, Patrick P. C. Lee
SRDS2