VLDB 2026 Research / reviewers in the wild / expert
Xiangrui Yang 0001
dblp:213/1007-1
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0008-7655-1948ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rearchitecting Buffered I/O in the Era of High-Bandwidth SSDs
Yekang Zhan, Tianze Wang, Zheng Peng 0017, Haichuan Hu, Xiangrui Yang 0001, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001 |
FAST | 6 |
| 2025 | Rethinking the Request-to-IO Transformation Process of File Systems for Full Utilization of High-Bandwidth SSDs
Yekang Zhan, Haichuan Hu, Xiangrui Yang 0001, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001 |
FAST | 3 |
| 2025 | HeteroGNN: A Heterogeneous Stage Division Based GNN Training Framework to Maximize CPU-GPU ParallelismabstractGraph Neural Networks (GNNs) have become inevitable tools for extracting knowledge from massive topological structure data. However, experimental observation shows that existing GNN training frameworks exhibit low efficiency when performing memory-access-intensive data preparation stage upon CPU and computation-intensive model training stage upon GPU. This is largely due to data dependency restriction between the two stages and sequential execution of computation in training iterations. Based on these observations, this paper proposes HeteroGNN, an efficient GNN training framework, to maximize parallelism of GNN training upon heterogeneous CPU-GPU architecture. Specifically, HeteroGNN first proposes a data-dependency-aware stage division policy, which divides the two stages to six phases to offer inter-stage parallelism. Further, HeteroGNN establishes a fine-grained computation partition schema, which actively partitions typical GNN computing operations into multiple schedulable tasks suitable for CPU and GPU. Finally, HeteroGNN designs an adaptive task scheduler, which adaptively schedules the tasks upon six phases to maximize GPU efficiency. Experiment results demonstrate that HeteroGNN speeds up end-to-end training time up to 1.29× comparing to state-of-the-art GNN framework, PiPAD, and 1.30× to 2.06× comparing to DGL and PyG. Xiangrui Yang 0001, Yekang Zhan, Qiang Cao 0001, Jie Yao 0001 |
ICME | 1 |
| 2025 | AIS: An Active Idleness I/O Scheduler to Reduce Buffer-Exhausted Degradation of Solid-State DrivesabstractModern solid-state drives (SSDs) continue to boost storage density and I/O bandwidth at the cost of flash-access I/O latency, especially for write, hence they prevalently deploy a build-in buffer to absorb incoming writes. However, when the buffer is used up, the applications suffer from a sudden and long performance decline, i.e., buffer-exhausted degradation (BED). To holistically understand BED and recovery, we design an automated testing toolset (SSDTest) to measure six commodity NVMe SSDs and find: (1) the occurrence of the BED strictly relies on the written-data amount, (2) BED dramatically increases I/O latency of SSDs, especially write and read-after-write, (3) BED can be conditionally reduced and recovered only after a period of idle time, and (4) a read without preceding writes is largely immune to BED, but prolongs the required idle time to recover the available buffer. Furthermore, we build a black-box SSD buffer-recovery model to quantitatively characterize the idleness-recovery behaviors and design an SSD BED predictor to make BED occurrence and buffer recovery predictable. Leveraging this model, we further design an Active Idleness I/O Scheduler (AIS) with small-sized auxiliary storage to actively regulate the I/O idle-intervals to maximize the internal buffer recovery of SSD. AIS adaptively steers incoming data to the auxiliary storage to (1) strategically keep SSD idle to reduce the occurrence of BED and (2) mitigate the tail latency of SSDs caused by read-after-writes during BED. We perform extensive evaluations under a variety of workloads. The results show that AIS improves average, 99th, 99.9th, and 99.99th-percentile latencies of SSDs by up to 29.3%, 37.3%, 78.7%, and 67.2% respectively, with up to 512MB auxiliary storage. Yekang Zhan, Xiangrui Yang 0001, Haichuan Hu, Qiang Cao 0001, Yifan Zhang 0012, Jie Yao 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | RomeFS: A CXL-SSD Aware File System Exploiting Synergy of Memory-Block Dual PathsabstractCompute eXpress Link (CXL) based Solid-State Drives (CXL-SSDs), such as the Samsung CMM-H model, promise to offer CXL.mem memory and CXL.io block dual-mode interfaces. Nonetheless, whether and how cloud applications with diverse and varying access patterns benefit from such dual-mode CXL-SSD remains an open question for academia and industry. Yekang Zhan, Haichuan Hu, Xiangrui Yang 0001, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001 |
SoCC | 3 |
| 2024 | HEncode: A Highly Modularized and Efficient FPGA QC-LDPC Encoder using High Level SynthesisabstractQC-LDPC (Quasi Cyclic Low-Density Parity-Check) codes, as a regular block-based code, have been preva-lently adopted in communication and storage fields to ensure high reliability and bandwidth of data channels. However, existing Field-Programmable Gate Array (FPGA) QC-LDPC encoders designed by RTL experts are generally dedicated to specialized LDPC codes and hardware platforms without flexibility and scalability. Recently, High-Level Synthesis (HLS) was introduced to compile a high-level encoding logic into Register Transfer Level (RTL) implementations, which are low performance and hardware efficiency due to the overlarge HLS-to-RTL design space, especially for large-scale FPGA hardware. This paper proposes a highly modularized and efficient FPGA QC-LDPC encoder, HEncoder, to fully leverage HLS to achieve high bandwidth, flexibility in both code parameters, and hardware efficiency. Firstly, HEncode presents an efficient Encode Block (EB) fully exploiting the FPGA LUT characteristic. Second, HEncode designs a low-level subword-encoding pipeline using multiple EBs and subword-parallel Encode Units (EU). Third, HEncoder designs an encode module with a pipelined data stream consecutively passing Input, EU array, and Output to balance bandwidths of accessing and encoding words. Finally, HEncode develops a design space analyzer to automatically determine the encoder parameters under constrained conditions to achieve high bandwidth. We implemented and evaluated HEncode on the Xilinx U50. The results show that compared to existing encoders, HEncode gains an increase of approximately 154.5× in the peak throughput and about 5.89 × in hardware efficiency to achieve the encoding throughput of 922.66 Gbps. Xiangrui Yang 0001, Yifan Zhang 0012, Qiang Cao 0001, Jie Yao 0001, Xiaodi Tan |
ICCD | 3 |
| 2023 | PMLDS: An LSM-Tree Direct Managed Storage for Key-Value Stores on Byte-Addressable DevicesabstractExisting key-value stores (KVSs) based on log-structured merge-tree (LSM-tree) have been broadly deployed in practice to leverage characteristics of conventional block storage via file system, but lack effective exploitation for emerging byte-addressed persistent memory (PM). We reveal that these KVSs running upon existing PM-aware File systems cause inefficient PM I/O behaviors, including 1) numerous page faults, 2) I/O misaligned with cacheline, and 3) bandwidth wastage of concurrent I/O threads. To make full use of PM without major modification for existing LSM-based KVSs, this paper proposes PMLDS, a direct managed storage for LSM-tree-based KVSs directly running upon PM. PMLDS acts as a unified I/O layer to handle all requests from KVS to PM. PMLDS designs an LSM-tree-aware data layout to directly map the KVS’s persistent objects to the storage slots with fixed location and size, thus simplifying and replacing the file system’s functionality with a minor modification. To improve I/O efficiency, PMLDS further presents three key techniques: 1) pre-allocating reusable data slots to avoid page faults, 2) forcing cacheline-alignment for small requests, and 3) scheduling asynchronous I/O threads to harness PM’s limited parallelism. We implement PMLDS and evaluate it with popular RocksDB under a variety of workloads. The results show that compared to representative PM-aware file systems such as Ext4-DAX, XFS-DAX, NOVA, and WineFS, PMLDS improves the write performance of RocksDB by up to 2.1 × while reducing the read latency by 20%~50%. Ziyi Lu, Qiang Cao 0001, Shucheng Wang, Jie Yao 0001, Xiangrui Yang 0001 |
ICPP | 5 |