Liangliang Xu

dblp:119/3453 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Towards Fast Erasure Coding at Register Efficiency
abstract
To reduce the high computation overhead induced by erasure coding, an effective way is to convert multiplications in finite fields intoXORs. However, the existing coding libraries adopt standard binaryXORand ignore the register efficiency, which inevitably induces too many extraLOADs/STOREsbetween registers and cache/memory, contributing to the main coding latency. From the view of register efficiency, we redesign the diagram of executingXORsand propose a new coding procedure, Coding with Adaptation to Registers (CAR), which keeps the temporal parities in registers until their constructions are completed. We further propose an enhanced coding procedure, CAR+, which further reduces the number ofLOADsby leveraging multiple registers. By integrating multiple optimizations into CAR and CAR+, we implement an erasure coding library, which increases the encoding throughput by up to 203.1% compared with the state-of-the-art erasure coding libraries.
Wei Wang 0502, Min Lyu, Yongkun Li 0001, Tianyang Niu, Liangliang Xu, Qiliang Li, Yinlong Xu 0001
IEEE Trans. Computers5
2025 ShuffleInfer: Disaggregate LLM Inference for Mixed Downstream Workloads
abstract
Transformer-based large language model (LLM) inference serving is now the backbone of many cloud services. LLM inference consists of a prefill phase and a decode phase. However, existing LLM deployment practices often overlook the distinct characteristics of these phases, leading to significant interference. To mitigate interference, our insight is to carefully schedule and group inference requests based on their characteristics. We realize this idea in ShuffleInfer through three pillars. First, it partitions prompts into fixed-size chunks so that the accelerator always runs close to its computation-saturated limit. Second, it disaggregates prefill and decode instances so each can run independently. Finally, it uses a smart two-level scheduling algorithm augmented with predicted resource usage to avoid decode scheduling hotspots. Results show that ShuffleInfer improves time-to-first-token (TTFT), job completion time (JCT), and inference efficiency in terms of performance per dollar by a large margin, e.g., it uses 38% less resources all the while lowering average TTFT and average JCT by 97% and 47%, respectively.
Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Chenxi Wang 0005, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan
ACM Trans. Archit. Code Optim.3
2025 MetaEC: An Efficient and Resilient Erasure-Coded KV Store on Disaggregated Memory
abstract
In-memory KV stores have recently been migrated from traditional monolithic servers to disaggregated memory (DM) for higher resource utilization and elasticity. These works use replication-based schemes for fault tolerance, which can be replaced with erasure coding (EC) for space efficiency. However, existing EC schemes designed in KV stores on traditional monolithic architectures encounter performance constraints when directly implemented in DM due to the challenges in EC metadata management and consistent parity updating. This article proposes MetaEC, an erasure-coded KV store on DM with high efficiency and resilience. First, for organizing KV pairs to stripes, MetaEC logically forms data chunks and leverages lazy coding to remove the accumulating and coding latency from the critical path. Second, for efficient EC metadata management, MetaEC designs EC metadata structures based on accessing features, and employs a hybrid redundancy schema with deterministic distribution to provide fault tolerance with high storage efficiency. Third, for consistent parity updating, we design a parity updating protocol based on parity logging and co-design EC metadata structures to handle concurrent conflicts by allowing only concurrent reads or writes. Experimental results show that compared with the state-of-the-art replication-based KV stores on DM, MetaEC achieves up to 53.33% latency reduction, up to 31.01% throughput improvement, and 58.17% memory consumption savings.
Qiliang Li, Min Lyu, Liangliang Xu, Wei Wang 0502, Yinlong Xu 0001
ACM Trans. Archit. Code Optim.4
2025 Fast Acceleration Strategies for XOR-Based Erasure Codes
abstract
Erasure coding is a common redundancy scheme for tolerating failures in storage systems. Compared with replication, erasure coding saves a large amount of storage space, but incurs heavy computation overhead and, is more time consuming. In this article, we accelerate the coding speed with three techniques. First, we propose an algorithm to search coding bitmatrices with fewer 1’s from Vandermonde and Cauchy matrices, and further optimize the coding bitmatrices by greedily reducing the number of 1’s in the bitmatrices. So we can find near-optimal coding bitmatrices with the number of 1’s only up to 1% more than the lower bound. Next, we redesign the process of building pointers and reuse the pointers to access data for coding, which obtains a better tradeoff between spatial locality and computation efficiency. Finally, we smartly decompose the coding procedure of wide stripes into multiple subprocedures, to improve spatial locality and reduce the number of XORs. Based on the proposed techniques, we implement an erasure coding library, Cerasure. Extensive experiments show that Cerasure significantly improves the coding throughput. Compared with the state-of-the-art erasure coding libraries, Zerasure and SLPEC, Cerasure increases the encoding throughput by up to 200.2%.
Wei Wang 0502, Min Lyu, Tianyang Niu, Qiliang Li, Liangliang Xu, Yinlong Xu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 AdLeaf: Quantitative Leaf Reconstruction From TLS Point Clouds
abstract
Quantitatively reconstructing the 3D structure of individual leaves within tree canopies is critical for understanding forest function and environmental responses to climate change. While quantitative structure models (QSMs) using terrestrial laser scanning (TLS) effectively capture woody structures, they lack the capability to accurately reconstruct non-woody leaf components. This study proposes AdLeaf (Accurate and Detailed Leaf), a novel approach for fine-scale reconstruction of individual leaves using TLS point clouds. AdLeaf combines wood-leaf separation, individual leaf segmentation, detection and repair of incomplete leaves, explicit reconstruction, and parameter extraction. It automates semantic segmentation at the tree scale to separate woody and leafy components. Instance segmentation is refined through similarity graphs. Incomplete leaves are detected and repaired using shape concavity analysis and symmetry-based mirroring. AdLeaf enables direct measurement of leaf attributes, including count, area, inclination, volume, and azimuth. Validation using field scans, synthetic data, and both in-situ and destructive measurements shows high accuracy: leaf counting errors ranged from 0.58% to 8.23% for trees with 201-4,000 leaves. Reconstructed leaf geometries had mean and standard deviations below 0.83 cm and 0.70 cm, respectively. Leaf area measurements (10–180 cm2) achieved a coefficient of determination (R²) of 0.95, bias of -0.20 cm², and root mean square error of 5.63 cm2. Incomplete leaf detection errors were below 28%, with the repaired area relative RMSE reduced by 9.4%. By addressing QSM limitations, AdLeaf enables explicit 3D leaf reconstructions that support detailed analysis of canopy light interception, spatial heterogeneity, and photosynthesis. It provides a robust framework for linking leaf structure to function at the tree level, advancing forest structure and radiative transfer research.
Guangpeng Fan, Liangliang Xu, Jiani Guo, Ruoyoulan Wang, Hao Lu 0004, Jinhu Wang, Di Wang 0006, Feixiang Chen, Liangliang Nan
IEEE Trans. Geosci. Remote. Sens.2
2025 Toward Efficient Repair for Wide-Stripe Erasure Coding With High Reliability
abstract
Erasure coding is a common redundancy scheme to provide higher reliability with much lower storage overhead compared to replication. It prevents data loss due to failures but induces high repair costs. As data volumes grow exponentially, wide stripes are proposed for extreme storage savings. Wide-stripe erasure codes face the challenges of higher repair costs for single and multiple failures. Our extensive analysis shows that existing repair-efficient erasure codes, such as locally repairable codes (LRCs) and minimum storage regenerating (MSR) codes, are insufficient to meet all the requirements of wide stripes: low storage overhead, low repair cost for both single and multiple failures, and high reliability. In this article, we explore an alternative code scheme, locally repairable with zigzag code (LRZC), which combines the advantages of LRCs and zigzag codes. LRZC divides data blocks and global parity blocks into evenly sized local groups, and generates two local parity blocks by a zigzag code in each group. Under the limit of storage overhead of wide stripes, LRZC reduces the repair cost for single and multiple failures and provides higher reliability compared with existing wide-stripe codes. Experiments show that LRZC reduces the repair cost of single and multiple failures by up to 41.9% and 41.7% compared with the state-of-the-art LRCs.
Wei Wang 0502, Zhipeng Li 0005, Min Lyu, Liangliang Xu, Yinlong Xu 0001
IEEE Trans. Reliab.4
2025 An MDS Code Construction for Optimal Update and Efficient Repair With Linear Subpacketization Level and Small Field Size
abstract
Maximum Distance Separable(MDS) codes can provide the optimal storage efficiency with the same fault tolerance. From the practical considerations, the systematic and optimal update properties of codes are crucial, where the former affects the workflow of read/write operations while the latter impacts the write amplification costs in update intensive scenarios. Moreover, the repair bandwidth, subpacketization level, and finite field size are three important performance metrics to evaluate the effectiveness of codes, which impact the network traffic, I/O performance and computational complexity, respectively. However, various code constructions with the optimal update property were devised to minimize repair bandwidth with high subpacketization levels or huge finite field sizes. While other constructions that reach a good trade-off among these three performance metrics always lack the optimal update property. In this paper, to address the above challenges of constructing practical MDS codes, we presentPermutation Transformation(PT) codesthat excel in the following respects: The systematic and optimal update properties can be both guaranteed; the code reaches nearly optimal repair bandwidth when repairing any single systematic node; the subpacketization level achieves a linear scale of the fault-tolerance capacity; the required size of the finite field to ensure the MDS property is small.
Min Lyu, Liangliang Xu, Zhipeng Li 0005, Yinlong Xu 0001
IEEE Trans. Reliab.3
2024 Fast recovery for large disk enclosures based on RAID2.0: Algorithms and evaluation
Qiliang Li, Min Lyu, Liangliang Xu
J. Parallel Distributed Comput.3
2024 Enabling Efficient Erasure Coding in Disaggregated Memory Systems
abstract
Disaggregated memory (DM) separates compute and memory resources to build a huge memory pool. Erasure coding (EC) is expected to provide fault tolerance in DM with low memory cost. In DM with EC, objects are first coded in compute servers, then directly written to memory servers via high-speed networks like one-sided RDMA. However, as the one-sided RDMA latency goes down to the microsecond level, coding overhead degrades the performance in DM with EC. To enable efficient EC in DM, we thoroughly analyze the coding stack from the perspective of cache efficiency and RDMA transmission. We develop MicroEC, which optimizes the coding workflow by reusing the auxiliary coding data and coordinates the coding and RDMA transmission with an exponential pipeline, as well as carefully adjusting the coding and transmission threads to minimize the latency. We implement a prototype supporting common basic operations, such as write/read/degraded read/recovery. Experiments show that MicroEC reduces the write latency by up to 44.35% and 42.14% and achieves up to$1.80\times$and$1.73\times$write throughput, compared with the state-of-the-art DM systems with EC and 3-way replication for objects not smaller than 1 MB, respectively. For small objects, MicroEC also evidently reduces the variation of latency, e.g., it reduces the P99 latency of writing 1 KB objects by 27.81%.
Qiliang Li, Liangliang Xu, Yongkun Li 0001, Min Lyu, Wei Wang 0502, Pengfei Zuo, Yinlong Xu 0001
IEEE Trans. Parallel Distributed Syst.2
2023 Skadi: Building a Distributed Runtime for Data Systems in Disaggregated Data Centers
abstract
Data-intensive systems are the backbone of today's computing and are responsible for shaping data centers. Over the years, cloud providers have relied on three principles to maintain cost-effective data systems: use disaggregation to decouple scaling, use domain-specific computing to battle waning laws, and use serverless to lower costs. Although they work well individually, they fail to work in harmony: an issue amplified by emerging data system workloads.
Cunchen Hu, Chenxi Wang 0005, Sa Wang, Ninghui Sun, Yungang Bao, Jieru Zhao, Sanidhya Kashyap, Pengfei Zuo, Xusheng Chen, Liangliang Xu, Yizhou Shan
HotOS10
2023 Towards the Diagnosis of Heart Disease Using an Ensemble Learning Approach
abstract
Heart disease is one of the leading causes of death worldwide. Understanding the presence of heart disease is crucial for timely intervention and effective management. However, it is still a challenging task to accurately diagnose and treat heart disease. In this paper, we propose an ensemble machine learning model that combines the predictive power of three state-of-the-art machine learning algorithms: BERT (Bidirectional Encoder Representations from Transformers), FT-transformer, and XG-Boost. Through extensive training and evaluation, we assess the performance of our model using established metrics such as accuracy, precision, recall, F1-score, and receiver operating char-acteristic (ROC) curve analysis. The results obtained demonstrate the effectiveness and validity of our machine learning system in predicting heart disease accurately. This paper presents a promising approach that can assist healthcare professionals in making informed decisions and improving patient outcomes, enhances our understanding of heart disease patterns, and contributes to the development of effective diagnostic and treatment approaches.
Liangliang Xu, Yida Bao, Sanjeev Baskiyar
ICMLA2
2022 A Data Layout and Fast Failure Recovery Scheme for Distributed Storage Systems With Mixed Erasure Codes
abstract
Erasure coding becomes increasingly popular in distributed storage systems (DSSes) for providing high reliability with low storage overhead. However, traditional random data placement induces massive cross-rack traffic and severely imbalanced load during failure recovery, which degrades the recovery performance significantly. In addition, various erasure codes coexisting in a DSS exacerbates the above problems. In this paper, we propose PDL, a PBD-based Data Layout, to optimize failure recovery performance in DSSes. PDL is constructed based on Pairwise Balanced Design, a combinatorial design scheme with uniform mathematical properties, and thus presents a uniform data layout for mixed erasure codes. Then we propose rPDL, a failure recovery scheme based on PDL. rPDL reduces cross-rack traffic effectively and provides nearly balanced cross-rack traffic distribution by uniformly choosing replacement nodes and retrieving determined available blocks to recover the lost blocks. We implemented PDL and rPDL in Hadoop 3.1.1. Compared with the existing data layout and recovery scheme in HDFS, experimental results show that rPDL achieves much higher recovery throughput, 6.27x for single-node failures, 5.14x for multi-node failures and 1.48x for single-rack failures, respectively. It also reduces degraded read latency by 62.83%, and provides evidently better support to front-end applications in case of component failures.
Liangliang Xu, Min Lyu, Zhipeng Li 0005, Cheng Li 0001, Yinlong Xu 0001
IEEE Trans. Computers1
2022 SelectiveEC: Towards Balanced Recovery Load on Erasure-Coded Storage Systems
abstract
Erasure coding (EC) has been commonly used to offer high data reliability with low storage cost. Upon failures, the lost blocks are recovered in batches. Due to the limited number of stripes, the data layout within a batch is non-uniform. Together with the random selection of source and replacement nodes for recovery tasks, the recovery workload among live nodes is skewed within a batch, which severely slows down failure recovery. To solve this problem, We present SelectiveEC, a new recovery task scheduling module that provides provable network traffic and recovery load balancing for large-scale EC-based storage systems. It relies on bipartite graphs to model the recovery traffic among live nodes. Then, it intelligently selects tasks to form batches and carefully determines where to read source blocks or to store recovered ones, using theories such as a perfect or maximum matching and$k$-regular spanning subgraph. SelectiveEC supports single-node failure and multi-node failure recovery, and can be deployed in both homogeneous and heterogeneous network environments. We implement SelectiveEC in HDFS, and evaluate its recovery performance in a local cluster of 18 nodes and AWS EC2 of 50 virtual machine instances. SelectiveEC increases the recovery throughput by up to$30.68\%$compared with state-of-the-art baselines in homogeneous network environments. It further achieves$1.32\times$recovery throughput and$1.23\times$benchmark throughput of HDFS on average in heterogeneous network environments, due to the straggler avoidance by the balanced scheduling.
Liangliang Xu, Min Lyu, Qiliang Li, Lingjiang Xie, Cheng Li 0001, Yinlong Xu 0001
IEEE Trans. Parallel Distributed Syst.1
2021 Fast Reconstruction for Large Disk Enclosures Based on RAID2.0
abstract
In the era of explosive data growth, RAID2.0 architecture with dozens or even hundreds of disks is commonly used to provide large capacity data storage. Due to limited resources, such as memory and CPU, the reconstruction for disk failures in RAID2.0 is executed in batches. Traditional random data placement and recovery scheme make the I/O access highly skewed within a batch, which slows down the reconstruction speed.
Qiliang Li, Min Lyu, Liangliang Xu, Yinlong Xu 0001, Wei Wang 0502
ICPP3
2021 Interactive Pose Attention Network for Human Pose Transfer
Guipeng Zhang, Zhenguo Yang, Minzheng Yuan, Liangliang Xu, Qing Li 0001, Wenyin Liu
WISE (2)6
2020 SelectiveEC: Selective Reconstruction in Erasure-coded Storage Systems
Liangliang Xu, Min Lyu, Qiliang Li, Lingjiang Xie
HotStorage1
2020 PDL: A Data Layout towards Fast Failure Recovery for Erasure-coded Distributed Storage Systems
abstract
Erasure coding becomes increasingly popular in distributed storage systems (DSSes) for providing high reliability with low storage overhead. However, traditional random data placement causes massive cross-rack traffic and severely unbalanced load during failure recovery, degrading the recovery performance significantly. In addition, various erasure coding policies coexisting in a DSS exacerbates the above problem. In this paper, we propose PDL, a PBD-based Data Layout, to optimize failure recovery performance in DSSes. PDL is constructed based on Pairwise Balanced Design, a combinatorial design scheme with uniform mathematical properties, and thus presents a uniform data layout. Then we propose rPDL, a failure recovery scheme based on PDL. rPDL reduces cross-rack traffic effectively and provides nearly balanced cross-rack traffic distribution by uniformly choosing replacement nodes and retrieving determined available blocks to recover the lost blocks. We implemented PDL and rPDL in Hadoop 3.1.1. Compared with existing data layout of HDFS, experimental results show that rPDL reduces degraded read latency by an average of 62.83%, delivers 6.27× data recovery throughput, and provides evidently better support for front-end applications.
Liangliang Xu, Min Lv, Zhipeng Li 0005, Cheng Li 0001, Yinlong Xu 0001
INFOCOM1
2020 Deterministic Data Distribution for Efficient Recovery in Erasure-Coded Storage Systems
abstract
Due to individual unreliable commodity components, failures are common in large-scale distributed storage systems. Erasure codes are widely deployed in practical storage systems to provide fault tolerance with low storage overhead. However, random data distribution (RDD), commonly used in erasure-coded storage systems, induces heavy cross-rack traffic, load imbalance, and random access, which adversely affects failure recovery. In this article, with orthogonal arrays, we define a Deterministic Data Distribution (D3) to uniformly distribute data/parity blocks among nodes, and propose an efficient failure recovery approach based on D3, which minimizes the cross-rack repair traffic against a single node failure. Thanks to the uniformity of D3, the proposed recovery approach balances the repair traffic not only among nodes within a rack but also among racks. We implement D3over Reed-Solomon codes and Locally Repairable Codes in Hadoop Distributed File System (HDFS) with a cluster of 28 machines. Compared with RDD, our experiments show that D3 significantly speeds up the failure recovery up to 2.49 times for RS codes and 1.38 times for LRCs. Moreover, D3supports front-end applications better than RDD in both of normal and recovery states.
Liangliang Xu, Min Lyu, Zhipeng Li 0005, Yongkun Li 0001, Yinlong Xu 0001
IEEE Trans. Parallel Distributed Syst.1
2019 D3: Deterministic Data Distribution for Efficient Data Reconstruction in Erasure-Coded Distributed Storage Systems
abstract
Due to individual unreliable commodity components, failures are common in large-scale distributed storage systems. Erasure codes are widely deployed in practical storage systems to provide fault tolerance with low storage overhead. However, the commonly used random data placement in storage systems based on erasure codes induces to heavy crossrack traffic, load imbalance, and random access, which slow down the recovery process upon failures. In this paper, with orthogonal arrays, we define a Deterministic Data Distribution (D3) of blocks to nodes and racks, and propose an efficient failure recovery approach based on D3. D3not only uniformly distributes data/parity blocks among storage servers, but also balances the repair traffic among racks and storage servers for failure recovery. Furthermore, D3also minimizes the cross-rack repair traffic for data layouts against a single rack failure and provides sequential access for failure recovery. We implement D3in Hadoop Distributed File System (HDFS) with a cluster of 28 machines. Our experiments show that D3significantly speeds up the failure recovery process compared with random data distribution, e.g., 2.21 times for (6, 3)-RS code in a system consisting of eight racks and three nodes in each rack.
Zhipeng Li 0005, Min Lv, Yinlong Xu 0001, Yongkun Li 0001, Liangliang Xu
IPDPS5
2017 A Note on One Weight and Two Weight Projective ℤ4-Codes
abstract
In this paper, we solve the open problems raised in [8] and present some examples to illustrate the obtained results. Moreover, we work out the diophantine problem by Shi and Wang and then give the sufficient conditions for the nonexistence of two-Lee weight projective codes over Z4with type 4k12k2.
Minjia Shi, Liangliang Xu
IEEE Trans. Inf. Theory2
2012 Joint Source-Channel Coding Based on P-LDPC Codes for Radiography Images Transmission
abstract
As the demand for e-Health care increases quickly, the transmission of medical images has been a crucial problem which needs to be solved as soon as possible. In the last decade, joint source-channel coding (JSCC), which combines the source coding with the channel coding reasonably to build an integral system so as to obtain significant improvement of system performance has attracted much attention. A framework of transmitting medical images by a P-JSCC scheme constructed from protograph low-density parity-check (P-LDPC) codes is proposed in this paper. Without loss of generality, we exploit a typical radiography image for simulation, and experimental results show that the receiver can recover the transmitted radiography image with good quality even at a very low signal-to-noise ratio (SNR); the P-JSCC scheme outperforms irregular-JSCC and regular-JSCC.
Huihui Wu, Jiguang He, Liangliang Xu, Lin Wang 0003
TrustCom3