Xiaojian Liao

dblp:275/5340 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 3 first-author · 18 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Soft Conflict-Resolution Decision Transformer for Offline Multi-Task Reinforcement Learning
abstract
Multi-task reinforcement learning (MTRL) seeks to learn a unified policy for diverse tasks, but often suffers from gradient conflicts across tasks. Existing masking-based methods attempt to mitigate such conflicts by assigning task-specific parameter masks. However, our empirical study shows that coarse-grained binary masks have the problem of over-suppressing key conflicting parameters, hindering knowledge sharing across tasks. Moreover, different tasks exhibit varying conflict levels, yet existing methods use a one-size-fits-all fixed sparsity strategy to keep training stability and performance, which proves inadequate. These limitations hinder the model’s generalization and learning efficiency. To address these issues, we propose SoCo-DT, a Soft Conflict-resolution method based by parameter importance. By leveraging Fisher information, mask values are dynamically adjusted to retain important parameters while suppressing conflicting ones. In addition, we introduce a dynamic sparsity adjustment strategy based on the Interquartile Range (IQR), which constructs task-specific thresholding schemes using the distribution of conflict and harmony scores during training. To enable adaptive sparsity evolution throughout training, we further incorporate an asymmetric cosine annealing schedule to continuously update the threshold. Experimental results on the Meta-World benchmark show that SoCo-DT outperforms the state-of-the-art method by 7.6% on MT50 and by 10.5% on the suboptimal dataset, demonstrating its effectiveness in mitigating gradient conflicts and improving overall multi-task performance.
Haiyuan Gui, Wenhao Ji, Xiaojian Liao
AAAI7
2026 Scalable RDMA-accelerated Distributed Locks with Shared Stream Abstraction
abstract
Blazing fast RDMA technology revolutionizes modern distributed systems and propels them to offload performance-critical data paths onto this network fabric. Designing an RDMA-optimized data path needs to clear a main hurdle—non-scalable distributed locks. Through a performance dissection of existing lock schemes, we find that software-based lock request ordering and polling-based lock ownership transfer scale poorly, leading to high NIC contention and heavy network congestion. To resolve these bottlenecks, this paper proposes StreamLock, a scalable lock primitive that co-designs the distributed lock protocol with fast RDMA networks. The core of StreamLock is a novel shared stream abstraction with two mechanisms: (i) scalable request ordering by repurposing the line-speed packet receiving provided by modern NICs; (ii) peer-to-peer notification to achieve one-round-trip-time lock ownership transfer. We implement StreamLock with off-the-shelf RDMA NICs and compare it with state-of-the-art distributed locks. Comprehensive experimental results showcase that StreamLock outperforms them significantly.
Miao Cai 0001, Junru Shen, Xiaojian Liao, Rong Gu 0001, Yanchao Zhao, Bing Chen 0002
EuroSys3
2026 RL-Paxos: Relieving the Leader's Burden with Efficient Task Offloading in Distributed Consensus
Jinquan Wang, Bing Wei 0002, Xiaojian Liao, Limin Xiao 0001
ICDE5
2026 Accelerating LLM Inference via Low-Bit Fine-Grained Quantization Algorithm and Bit-Level Accelerator Co-Design
abstract
Large language models (LLMs) have emerged as one of the most impactful and transformative paradigms in natural language processing. Despite their remarkable success, the intensive computational demands and substantial memory footprint impose a significant barrier to efficient LLM inference.In this paper, we present a comprehensive solution to improve LLM inference performance under ultra-low weight precision, meticulously optimized through algorithm and architecture co-design. To achieve this, we first propose a fine-grained intra-cluster bit allocation method that partitions the weights into small clusters and explicitly considers the distribution of outliers and salient points within each cluster. Then, an intra-cluster protection mechanism is proposed to selectively preserve important weights during quantization, where an extended integer format and group-wise scale factor search are further introduced to mitigate accuracy degradation caused by aggressive bit-width reduction. Furthermore, we develop a memory-aligned encoding scheme to facilitate efficient memory access while enabling flexible identification of mixed-precision representations. Finally, we design a lightweight bit-level accelerator for low-bit LLM inference, offering simplified hardware design and enhanced adaptability through parallel bit-level computation. Compared to existing state-of-the-art quantization algorithms, our algorithm achieves higher model accuracy under ultra-low weight precision. Meanwhile, the proposed bit-level accelerator delivers speedups of 1.59×, 1.38×, and 1.61×, along with energy efficiency improvements of 1.52×, 1.42×, and 1.22× over ANT, OliVe, and FineQ, respectively.
Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Tairan Zhang, Jinquan Wang, Yongyue Wang, Xiaojian Liao
IEEE Trans. Computers8
2025 CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
abstract
Large language models like GPT-4 are resource-intensive, but recent advancements suggest that smaller, specialized experts can outperform the monolithic models on specific tasks. The Collaboration-of-Experts (CoE) approach integrates multiple expert models, improving the accuracy of generated results and offering great potential for precision-critical applications, such as automatic circuit board quality inspection. However, deploying CoE serving systems presents challenges to memory capacity due to the large number of experts required, which can lead to significant performance overhead from frequent expert switching across different memory and storage tiers.
Jiashun Suo, Xiaojian Liao, Limin Xiao 0001, Jinquan Wang, Xiao Su 0002, Zhisheng Huo
ASPLOS (2)2
2025 Offloading File System Clients onto a Data Processing Unit (DPU) with DPUFS
abstract
Distributed File System (DFS) is a fundamental system service on public clouds. As DFS has developed and storage disaggregation technology has been applied, DFS clients have increasingly consumed CPU resources, burdening other applications. Offloading file system clients onto a Data Processing Unit (DPU) is a promising solution. We propose DPUFS, a high-performance DPU-based file system. DPUFS offloads file system clients onto a DPU, allowing CPU and memory resources on the host to be reserved for applications. To diminish link overhead, we exploit Scalable Metadata Cache to offload metadata processing and provide scalable performance. To reduce datapath overhead, we introduce RDMA-Based Datapath, which utilizes the software features of the DPU to achieve zero-copy datapath by bypassing DPU memory. Our evaluation shows that DPUFS effectively offloads the file system clients while maximizing performance with the limited resources of the DPU. Compared with the state-of-theart DPFS, DPUFS achieved 37 % to 61 % latency reduction in file operations and up to$3.8 \times$throughput in Filebench.
Qingjie Zeng, Xiaojian Liao, Xianqiang Luo, Jiwu Shu
CCGrid2
2025 CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
abstract
With the rapid development of DNN applications, multi-tenant execution, where multiple DNNs are co-located on a single SoC, is becoming a prevailing trend. Although many methods are proposed in prior works to improve multi-tenant performance, the impact of shared cache is not well studied. This paper proposes CaMDN, an architecture-scheduling co-design to enhance cache efficiency for multi-tenant DNNs on integrated NPUs. Specifically, a lightweight architecture is proposed to support model-exclusive, NPU-controlled regions inside shared cache to eliminate unexpected cache contention. Moreover, a cache scheduling method is proposed to improve shared cache utilization. In particular, it includes a cache-aware mapping method for adaptability to the varying available cache capacity and a dynamic allocation algorithm to adjust the usage among co-located DNNs at runtime. Compared to prior works, CaMDN reduces the memory access by 33.4% on average and achieves a model speedup of up to 2.56 × (1.88 × on average).
Tianhao Cai, Liang Wang 0020, Limin Xiao 0001, Xiaojian Liao
DAC7
2025 Zebra: Efficient Redundant Array of Zoned Namespace SSDs Enabled by Zone Random Write Area (ZRWA)
abstract
Zoned Namespace (ZNS) SSDs have emerged as a promising solution for eliminating device-level garbage collection and achieving predictable performance through the zone abstraction, which supports append-only writes. However, integrating ZNS SSDs into RAID configurations presents challenges, particularly concerning partial parity updates (PPUs) due to the no-overwrite property of ZNS SSDs. Existing solutions introduce dedicated metadata zones to manage PPUs, but these zones become performance bottlenecks, especially under high PPU workloads.In this paper, we introduce Zebra, a novel architecture for ZNS RAID that leverages the Zone Random Write Area (ZRWA) feature of modern ZNS SSDs. Our key insight is that the SSD’s write buffer (or ZRWA) in front of each zone is ideal for PPUs that exhibit a repeated sequential overwrite I/O pattern. By ensuring that the parity chunk fits within the ZRWA window, Zebra enables efficient PPUs within the ZRWA. Combined with a set of techniques, Zebra ensures both high performance and reliability. We evaluate Zebra on both large-zone and small-zone ZNS SSDs, two typical ZNS models. Our experiments with micro and application benchmarks show that Zebra improves the throughput by 1.1×-4.2× compared to RAIZN, a state-of-the-art ZNS RAID system.
Tianyang Jiang, Guangyan Zhang, Xiaojian Liao
HPCA3
2025 Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data Type
abstract
The quantization of Large Language Models (LLMs) poses significant challenges due to the heterogeneous nature of feature point distributions in low-bit quantization scenarios, including salient points, normal outliers, and massive outliers.These challenges are particularly pronounced in supporting both weight-only and weight-activation quantization modes, as existing methods often focus on a single mode and fail to address the diverse feature characteristics holistically, resulting in suboptimal model accuracy and hardware efficiency trade-offs.To tackle these limitations, we introduce Amove, a novel codesign framework that synergistically integrates data type and hardware architecture design for efficient LLM quantization.Our approach is threefold: First, we conduct a comprehensive analysis of quantization granularity and propose a residual approximation mechanism that balances model accuracy and memory overhead under fine-grained quantization.Second, we design a flexible finegrained grouped vectorized data type, enabling seamless support for both weight-activation and low-bit weight-only quantization modes within a unified framework.Third, we implement the hardware architecture of Amove on both GPU tensor core and systolic arraybased architectures.The Amove-enhanced tensor core achieves an average speedup of 2.13× and a 1.70× reduction in energy consumption over the state-of-the-art OliVe design.Furthermore, an Amove-based accelerator achieves up to 2.67× speedup and 1.68× energy reduction over the state-of-the-art accelerator.
Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Xiangrong Xu 0002, Jinquan Wang, Xiaojian Liao
MICRO9
2025 Exploiting intra-chip locality for multi-chip GPUs via two-level shared L1 cache
Xiangrong Xu 0002, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Yuanqiu Lv, Xilong Xie, Xiaojian Liao
J. Syst. Archit.10
2025 Efficiently Enlarging RDMA-Attached Memory with SSD
abstract
RDMA-based in-memory storage systems offer high performance but are restricted by the capacity of physical memory. In this article, we propose TeRM to extend RDMA-attached memory with SSD. TeRM achieves fast remote access on the SSD-extended memory by eliminating page faults of RDMA NIC and CPU from the critical path. We also introduce a set of techniques to reduce the consumption of CPU and network resources. Evaluation shows that TeRM performs close to the performance of the ideal upper bound where all pages are pinned in the physical memory. Compared with existing approaches, TeRM significantly improves the performance of unmodified RDMA-based storage systems, including a file system and a key-value system.
Zhe Yang 0012, Qing Wang 0031, Xiaojian Liao, Youyou Lu, Keji Huang, Jiwu Shu
ACM Trans. Storage3
2024 Volley: Accelerating Write-Read Orders in Disaggregated Storage
abstract
Modern data centers deploy disaggregated storage systems (e.g., NVMe over Fabrics, NVMe-oF) for fine-grained resource elasticity and high resource utilization. A client-side writeback cache is used to absorb writes and buffer frequently accessed data, thereby eliminating unnecessary remote storage accesses and improving performance. Yet, a cache miss on the full cache triggers an evict-and-fetch operation which evicts the old entries before new data blocks are fetched. Existing systems perform the evict-and-fetch operation by sequentially executing write and read I/O operations, which reduces the concurrency and makes it challenging to fully utilize the fast network and storage devices.
Shaoxun Zeng, Xiaojian Liao, Youyou Lu
EuroSys2
2024 TeRM: Extending RDMA-Attached Memory with SSD
Zhe Yang 0012, Qing Wang 0031, Xiaojian Liao, Youyou Lu, Keji Huang, Jiwu Shu
FAST3
2024 Accelerating Virtual Machine File Systems with TimeFS
abstract
Virtualized storage systems leverage high-performance storage devices (e.g., NVMe SSDs), log-structured block devices, and low-latency virtualization techniques (e.g., vhost) to deliver exceptional performance and meet the growing demands of modern cloud applications. However, as volume utilization increases, performance can degrade significantly due to frequent garbage collection, which undermines the full potential of vhost and NVMe SSDs.We propose a cross-layered I/O stack called TimeFS. TimeFS introduces an expiration time estimation technique, which addresses the issue of semantic isolation between the virtual machine file system and the log-structured block device. Additionally, TimeFS implements a hotness-aware data layout strategy, which organizes the data layout based on estimated expiration timestamps, thereby reducing the overhead during garbage collection. Furthermore, TimeFS introduces a bucket sort-based garbage collection algorithm, which optimizes the garbage collection process in coordination with the data layout adjustments to achieve greater efficiency. Our evaluations demonstrate that TimeFS effectively reduces garbage collection overhead, enhances garbage collection efficiency, and improves system I/O performance. We compared TimeFS with the F2FS over vhost virtualized I/O stack. TimeFS achieved performance improvements of 18%–25% in FIO write-intensive workloads and 5%–21% in Filebench, YCSB, and Mobibench workloads.
Jiaxuan Kang, Xiaojian Liao, Jiwu Shu
HPCC2
2023 RIO: Order-Preserving and CPU-Efficient Remote Storage Access
abstract
Modern NVMe SSDs and RDMA networks provide dramatically higher bandwidth and concurrency. Existing networked storage systems (e.g., NVMe over Fabrics) fail to fully exploit these new devices due to inefficient storage ordering guarantees. Severe synchronous execution for storage order in these systems stalls the CPU and I/O devices and lowers the CPU and I/O performance efficiency of the storage system.
Xiaojian Liao, Zhe Yang 0012, Jiwu Shu
EuroSys1
2023 λ-IO: A Unified IO Stack for Computational Storage
Zhe Yang 0012, Youyou Lu, Xiaojian Liao, Youmin Chen, Siyu He, Jiwu Shu
FAST3
2023 SingularFS: A Billion-Scale Distributed File System Using a Single Metadata Server
Youyou Lu, Wenhao Lv, Xiaojian Liao, Shaoxun Zeng, Jiwu Shu
USENIX ATC4
2023 A convolutional spiking neural network with adaptive coding for motor imagery classification
Xiaojian Liao, Yuli Wu 0002, Deheng Wang, Hongmiao Zhang
Neurocomputing1
2023 Efficient Crash Consistency for NVMe over PCIe and RDMA
abstract
This article presents crash-consistent Non-Volatile Memory Express (ccNVMe), a novel extension of the NVMe that defines how host software communicates with the non-volatile memory (e.g., solid-state drive) across a PCI Express bus and RDMA-capable networks with both crash consistency and performance efficiency. Existing storage systems pay a huge tax on crash consistency, and thus cannot fully exploit the multi-queue parallelism and low latency of the NVMe and RDMA interfaces. ccNVMe alleviates this major bottleneck by coupling the crash consistency to the data dissemination. This new idea allows the storage system to achieve crash consistency by taking the free rides of the data dissemination mechanism of NVMe, using only two lightweight memory-mapped I/Os (MMIOs), unlike traditional systems that use complex update protocol and synchronized block I/Os. ccNVMe introduces a series of techniques including transaction-aware MMIO/doorbell and I/O command coalescing to reduce the PCIe traffic as well as to provide atomicity. We present how to build a high-performance and crash-consistent file system named MQFS atop ccNVMe. We experimentally show that MQFS increases the IOPS of RocksDB by 36% and 28% compared to a state-of-the-art file system and Ext4 without journaling, respectively.
Xiaojian Liao, Youyou Lu, Zhe Yang 0012, Jiwu Shu
ACM Trans. Storage1
2022 Efficient Atomic Durability on eADR-Enabled Persistent Memory
abstract
Applications atop persistent memory (PM) require atomic durability to ensure crash consistency. However, existing atomic durability techniques designed for PM systems are based on volatile cache and incur non-negligible performance overhead. Recently, Intel introduces a new feature called eADR (enhanced Asynchronous DRAM Refresh) for Optane PM, which brings an opportunity to build a much more efficient atomic durability system for PM.
Taiyu Zhou, Yajuan Du, Fan Yang 0134, Xiaojian Liao, Youyou Lu
PACT4
2021 Crash Consistent Non-Volatile Memory Express
abstract
This paper presents crash consistent Non-Volatile Memory Express (ccNVMe), a novel extension of the NVMe that defines how host software communicates with the non-volatile memory (e.g., solid-state drive) across a PCI Express bus with both crash consistency and performance efficiency. Existing storage systems pay a huge tax on crash consistency, and thus can not fully exploit the multi-queue parallelism and low latency of the NVMe interface. ccNVMe alleviates this major bottleneck by coupling the crash consistency to the data dissemination. This new idea allows the storage system to achieve crash consistency by taking the free rides of the data dissemination mechanism of NVMe, using only two lightweight memory-mapped I/Os (MMIO), unlike traditional systems that use complex update protocol and heavyweight block I/Os. ccNVMe introduces transaction-aware MMIO and doorbell to reduce the PCIe traffic as well as to provide atomicity. We present how to build a high-performance and crash-consistent file system namely MQFS atop ccNVMe. We experimentally show that MQFS increases the IOPS of RocksDB by 36% and 28% compared to a state-of-the-art file system and Ext4 without journaling, respectively.
Xiaojian Liao, Youyou Lu, Zhe Yang 0012, Jiwu Shu
SOSP1
2021 Max: A Multicore-Accelerated File System for Flash Storage
Xiaojian Liao, Youyou Lu, Erci Xu, Jiwu Shu
USENIX ATC1
2020 OCVM: Optimizing the Isolation of Virtual Machines with Open-Channel SSDs
Xiaojian Liao, Zhe Yang 0012, Youyou Lu, Jiwu Shu
ICA3PP (1)2
2020 Write Dependency Disentanglement with HORAE
Xiaojian Liao, Youyou Lu, Erci Xu, Jiwu Shu
OSDI1