Xuchao Xie

dblp:138/1792 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0003-4722-7619ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 DPIO: A Unified I/O Architecture for Heterogeneous CPU and DPU NVMeoF
Wenhao Gu, Xuchao Xie, Yujuan Tan, Dezun Dong
HPDC2
2025 Bridging Metadata Service and CXL: A Metadata-Grained and Directory-Aware Storage Engine for Distributed Storage Systems
abstract
The AI training and inference workloads are particularly metadata-intensive and drive an urgent need for distributed file systems (DFS) with high IOPS metadata service. Meanwhile, the emerging Compute Express Link (CXL) protocol introduces memory semantics to the PCIe-attached storage devices and is compelling for building high-performance metadata storage. However, a fundamental mismatch exists between the metadata access granularity and the internal storage granularity of typical CXL-enabled storage devices. Besides, existing metadata storage engines lack the perception of the DFS directory structures and metadata semantics in distributed storage systems. In this paper, we investigate the way to employ CXL-enabled devices as the storage backend for DFS metadata and propose MDSec, a metadata-grained and directory-aware metadata storage engine that bridges the semantic gaps between the DFS metadata service and CXL-enabled storage devices. MDSec unifies the granularity of all kinds of metadata for better metadata placement across CXL-enabled persistent storage, designs a directory-aware metadata grouping and placement strategy to improve the spatial locality of metadata access, and employs fully parallel metadata handlers to enhance metadata processing parallelism for parallel DFS client accesses. We evaluate MDSec on a cluster with 25 nodes. MDSec improves the throughput of Ext4 and NOVA by 258% and 53%, while reducing their latency by 46% and 11%, respectively. These results indicate that MDSec efficiently integrates CXL storage with DFS metadata service.
Xuchao Xie, Xinghan Qiao, Qiulin Wu, Wenhao Gu, Liquan Xiao
CLUSTER2
2025 Sumeru: An Efficient Hybrid-Granularity Cache Management Scheme for CXL-SSDs
Xuchao Xie, Qiulin Wu, Xingyun Qi, Zhenlong Song
ICA3PP (1)2
2023 CLMS: Configurable and Lightweight Metadata Service for Parallel File Systems on NVMe SSDs
Shuaizhe Lv, Xuchao Xie, Zhenlong Song
APPT3
2023 UrsaX: Integrating Block I/O and Message Transfer for Ultrafast Block Storage on Supercomputers
abstract
It is increasingly important for the next-generation exascale supercomputers to extend its applications beyond traditional high-performance computing (HPC) scenarios, so as to achieve high social and economic benefit. Similar to Amazon Web Services (AWS) and Alibaba Cloud, cloud-style virtual HPC service is a promising application scenario on supercomputers, for which remote block storage is the key to provide tenants with supercomputers’ extremely high storage performance. Unfortunately, the state-of-the-art block storage software systems (such as URSA and Ceph) cannot adapt to the advanced hardware features of supercomputers. This article presents UrsaX, an efficient block storage service for our next-generation Tianhe exascale supercomputer that is equipped with the high-performance global express (GLEX) network and nonvolatile memory express (NVMe) SSDs. UrsaX’s virtual disks, which can be mounted like normal physical ones, enable not only traditional HPC applications but also supercomputer-oblivious POSIX applications to enjoy the high performance of supercomputers. At the core of UrsaX is with a novel design of the efficient integration of on-disk block I/O and in-network message transfer on supercomputers. UrsaX utilizes the NVMe Fabrics kernel module to expand the NVMe standard on the supercomputer network, and separates metadata I/O and data I/O of blocks, respectively, being handled over the mini packet (MP) and remote direct memory access (RDMA) protocols. We thoroughly explore the design space for remote block storage on supercomputers, including parallelism, scalability, fault tolerance, and consistency. We conduct an extensive evaluation on a subset of our exascale supercomputer consisting of 44 storage machines (each with four NVMe SSDs). The result shows that UrsaX achieves local-storage-level I/O latency (tens of microseconds) while being able to linearly increase the aggregate performance (IOPS and throughput) as the system scale increases, an order of magnitude higher than the state-of-the-art block storage systems.
Shun Gai, Yiming Zhang 0003, Xuchao Xie, Yong Dong, Zhenlong Song
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 LTNoT: Realizing the Trade-Offs Between Latency and Throughput in NVMe over TCP
Wenhao Gu, Xuchao Xie, Dezun Dong
ICA3PP2
2022 A Transformable NVMeoF Queue Design for Better Differentiating Read and Write Request Processing
abstract
NVMeoF is the latest extension of NVMe for remote storage access which allows remote access to NVMe controllers through high-speed RDMA, FC, and TCP networks. NVMe over TCP (NoT) can build on the basis of large-scale common network infrastructure in datacenters and standard TCP/IP software protocol stack, enabling a wide availability compared with RDMA-enabled specific network infrastructure for NVMe-overRDMA. However, the processing of read/write I/O at the host and target prominently shows significantly different characteristics and requirements, where one side sends the NVMeoF instruction of the request, while the other side sends the requested data. The existing NoT implementation can not meet the different characteristics of requests in the datacenter, which eventually results in the I/O performance being limited by the common processing pipeline and sending strategy. In this paper, we propose RNoT, a transformable queue that can meet the differentiated processing scheme of read or write request characteristics respectively in NoT implementation. Specifically, RNoT defines a switchable working attribute and separates resources for read and write I/O to achieve intra-queue long-term exclusivity, delivers read and write requests into other RNoT queue pairs to achieve inter-queue I/O scheduling, and transfers request command and data with targeted approaches to achieve short and long flow optimization. We implemented RNoT in Linux Kernel and evaluated it using realistic benchmarks and applications. Our experimental results demonstrate that RNoT can achieve 30.39% and 29.27% lower latency than i10 and NoT respectively, increase IOPS by up to 41.34% than NoT on average, thus RNoT can effectively optimize the read and write I/O performance in NoT with dedicated processing scheme.
Wenhao Gu, Xuchao Xie, Dezun Dong
ICPADS2
2022 Alleviating Performance Interference Through Intra-Queue I/O Isolation for NVMe-over-Fabrics
Wenhao Gu, Xuchao Xie, Dezun Dong
NPC2
2020 Idler : I/O Workload Controlling for Better Responsiveness on Host-Aware Shingled Magnetic Recording Drives
abstract
Host-Aware/Drive-Managed Shingled Magnetic Recording (SMR) drives can accept non-sequential writes using a buffer called media cache. Data in the media cache will be migrated to its designated location by a cleaning process if the buffer is full (blocking cleaning) or the drive is idle (idle cleaning). However, blocking cleanings can severely extend the I/O response time. Therefore, it is crucial to fully understand the cleaning process and find ways of mitigating the caused performance degradation. In this article we further evaluate the cleaning process and propose a potential remedy scheme called Idler on Host-Aware SMR drives. Idler adaptively induces idle cleanings based on dynamic workload characteristics and media cache usages to reduce the severity of blocking cleanings. Our evaluations show that in the workloads with a small non-sequential write ratio (about 10 percent), Idler can reduce the tail response time and the workload finish time by 56-88 and 10-23 percent, respectively, compared with those without such control. With the help of an external write buffer on an SSD, the tail response time of SMR drives with Idler can be closer to that of conventional disk drives.
Baoquan Zhang, Ming-Hong Yang, Xuchao Xie, David Hung-Chang Du
IEEE Trans. Computers3
2019 Pinpointing and scheduling access conflicts to improve internal resource utilization in solid-state drives
Xuchao Xie, Liquan Xiao, Dengping Wei, Zhenlong Song, Xiongzi Ge
Frontiers Comput. Sci.1
2019 ZoneTier: A Zone-based Storage Tiering and Caching Co-design to Integrate SSDs with SMR Drives
abstract
Integrating solid-state drives (SSDs) and host-aware shingled magnetic recording (HA-SMR) drives can potentially build a cost-effective high-performance storage system. However, existing SSD tiering and caching designs in such a hybrid system are not fully matched with the intrinsic properties of HA-SMR drives due to their lacking consideration of how to handle non-sequential writes (NSWs). We propose ZoneTier, a zone-based storage tiering and caching co-design, to effectively control all the NSWs by leveraging the host-aware property of HA-SMR drives. ZoneTier exploits real-time data layout of SMR zones to optimize zone placement, reshapes NSWs generated from zone demotions to SMR preferred sequential writes, and transforms the inevitable NSWs to cleaning-friendly write traffics for SMR zones. ZoneTier can be easily extended to match host-managed SMR drives using proactive cleaning policy. We implemented a prototype of ZoneTier with user space data management algorithms and real SSD and HA-SMR drives, which are manipulated by the functions provided by libzbc and libaio. Our experiments show that ZoneTier can reduce zone relocation overhead by 29.41% on average, shorten performance recovery time of HA-SMR drives from cleaning by up to 33.37%, and improve performance by up to 32.31% than existing hybrid storage designs.
Xuchao Xie, Liquan Xiao, David Hung-Chang Du
ACM Trans. Storage1
2018 Duchy: Achieving Both SSD Durability and Controllable SMR Cleaning Overhead in Hybrid Storage Systems
abstract
Integrating solid-state drives (SSDs) and shingled magnetic recording (SMR) drives can build cost-effective hybrid storage systems. However, both SSD and SMR drives endure inherent defects that are mutually exclusive. The write endurance of SSD is limited while SMR drives should prevent from writes due to the cleaning-caused performance degradation. In this paper, we propose Duchy, an endurable SSD caching scheme that simultaneously respects SMR constraints. Duchy filters ineffectual write traffic out of SSD without exacerbating the performance degradation of SMR drives. Meanwhile, Duchy leverages SSD to regulate the written zones in SMR drives to achieve controllable cleaning duration. Our experimental results indicate that compared with legacy SSD caching designs, only Duchy can achieve both system performance improvement and SSD write traffic reduction.
Xuchao Xie, Tianye Yang, Dengping Wei, Liquan Xiao
ICPP1
2018 ChewAnalyzer: Workload-Aware Data Management Across Differentiated Storage Pools
abstract
In multi-tier storage systems, moving data from one tier to the next can be inefficient. And because each type of storage device has its own idiosyncrasies with respect to the workloads that it can best support, unnecessary data movement might result. In this paper, we explore a fully connected storage architecture in which data can move from any storage pool to another. We propose a Chunk-level storage-aware workload Analyzer framework, abbreviated as ChewAnalyzer, to facilitate efficient data placement. Access patterns are characterized in a flexible way by a collection of I/O accesses to a data chunk. ChewAnalyzer employs a Hierarchical Classifier [1] to analyze the chunk patterns step by step. In each classification step, the Chunk Placement Recommender suggests new data placement policies according to the device properties. Based on the analysis of access pattern changes, the Storage Manager can adequately distribute or migrate the data chunks across different storage pools. Our experimental results show that ChewAnalyzer improves the initial data placement and that it migrates data into the proper pools directly and efficiently.
Xiongzi Ge, Xuchao Xie, David Hung-Chang Du, Pradeep Ganesan, Dennis Hahn
MASCOTS2
2015 CER-IOS: Internal Resource Utilization Optimized I/O Scheduling for Solid State Drives
abstract
Modern Solid State Drives (SSDs) integrate more internal resources to get higher performance and capacity. Improving internal resource utilization by exploiting internal parallelism is important to enhance the performance of SSDs. Unfortunately, the internal resource utilization of SSDs is limited at runtime in practice because of the practical access conflicts to internal resources. In this paper, we propose a Conflict Eliminated Requests Based I/O Scheduler (CER-IOS) to better utilize internal parallelism of flash chips by scheduling I/O requests in a more fine-grained way. We introduce Conflict Eliminated Requests (CERs) in which parallelizable memory requests are grouped during the process of address translation in Flash Translation Layer. To schedule conflicting requests, we propose a small CER size prioritized resource distribution scheme, that ensures internal resources can always be distributed to valuable conflicting requests to further improve the efficiency of resource utilization. Our extensive experimental evaluation results show that CER-IOS provides significant improvement of resource utilization at runtime and reduces average I/O latency largely compared to state-of-the-art I/O schedulers implemented in operating systems.
Xuchao Xie, Dengping Wei, Zhenlong Song, Liquan Xiao
ICPADS1
2013 ECAM: An Efficient Cache Management Strategy for Address Mappings in Flash Translation Layer
Xuchao Xie, Dengping Wei, Zhenlong Song, Liquan Xiao
APPT1