Qiliang Li

dblp:130/4594 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Fast Erasure Coding at Register Efficiency
abstract
To reduce the high computation overhead induced by erasure coding, an effective way is to convert multiplications in finite fields intoXORs. However, the existing coding libraries adopt standard binaryXORand ignore the register efficiency, which inevitably induces too many extraLOADs/STOREsbetween registers and cache/memory, contributing to the main coding latency. From the view of register efficiency, we redesign the diagram of executingXORsand propose a new coding procedure, Coding with Adaptation to Registers (CAR), which keeps the temporal parities in registers until their constructions are completed. We further propose an enhanced coding procedure, CAR+, which further reduces the number ofLOADsby leveraging multiple registers. By integrating multiple optimizations into CAR and CAR+, we implement an erasure coding library, which increases the encoding throughput by up to 203.1% compared with the state-of-the-art erasure coding libraries.
Wei Wang 0502, Min Lyu, Yongkun Li 0001, Tianyang Niu, Liangliang Xu, Qiliang Li, Yinlong Xu 0001
IEEE Trans. Computers6
2026 A 50 μW/Gbps/Lane Power-Efficient MIPI D-PHY Receiver With Architecture-Level Adaptive and Structural Optimizations for Micro-Displays
abstract
Achieving high power efficiency in Mobile Industry Processor Interface (MIPI) D-PHY receivers is crucial for micro-display chips in AR/VR systems, where stringent power constraints exist. However, existing designs often sacrifice power efficiency for higher data rates due to architectural limitations, neglecting optimization for low-power applications. To address this issue, we propose a receiver architecture that substantially enhances power efficiency through three key techniques. First, we improve the gain-bandwidth product (GBW) by employing an autonomous gain scheduling analog front-end (AFE) that dynamically tunes the gain while reducing drive current. Second, we reduce clocking overhead by introducing a self-monitoring interferometric deserializer that enables clock-free pre-scaling and halves the DDR sampling frequency. Third, we increase transition speed and minimize short-circuit power by utilizing a chaotic topological flow actuator (CTFA) with multi-path current feedthrough. Compared to prior state-of-the-art designs, the proposed receiver achieves a power efficiency of$50~\mu $W/Gbps/lane ($42~\mu $A/Gbps/lane), reducing power and current consumption by 46% and 45%, respectively, using a standard 180-nm process.
Haoran Zeng, Yingqi Feng, Tianai Li, Hang Ye 0007, Zunkai Huang, Hui Wang 0036, Yongxin Zhu 0001, Qiliang Li, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.9
2025 Transformer-based material recognition via short-time contact sensing
Zhenyang Liu, Yitian Shao, Qiliang Li, Jingyong Su
Pattern Recognit.3
2025 MetaEC: An Efficient and Resilient Erasure-Coded KV Store on Disaggregated Memory
abstract
In-memory KV stores have recently been migrated from traditional monolithic servers to disaggregated memory (DM) for higher resource utilization and elasticity. These works use replication-based schemes for fault tolerance, which can be replaced with erasure coding (EC) for space efficiency. However, existing EC schemes designed in KV stores on traditional monolithic architectures encounter performance constraints when directly implemented in DM due to the challenges in EC metadata management and consistent parity updating. This article proposes MetaEC, an erasure-coded KV store on DM with high efficiency and resilience. First, for organizing KV pairs to stripes, MetaEC logically forms data chunks and leverages lazy coding to remove the accumulating and coding latency from the critical path. Second, for efficient EC metadata management, MetaEC designs EC metadata structures based on accessing features, and employs a hybrid redundancy schema with deterministic distribution to provide fault tolerance with high storage efficiency. Third, for consistent parity updating, we design a parity updating protocol based on parity logging and co-design EC metadata structures to handle concurrent conflicts by allowing only concurrent reads or writes. Experimental results show that compared with the state-of-the-art replication-based KV stores on DM, MetaEC achieves up to 53.33% latency reduction, up to 31.01% throughput improvement, and 58.17% memory consumption savings.
Qiliang Li, Min Lyu, Liangliang Xu, Wei Wang 0502, Yinlong Xu 0001
ACM Trans. Archit. Code Optim.1
2025 Fast Acceleration Strategies for XOR-Based Erasure Codes
abstract
Erasure coding is a common redundancy scheme for tolerating failures in storage systems. Compared with replication, erasure coding saves a large amount of storage space, but incurs heavy computation overhead and, is more time consuming. In this article, we accelerate the coding speed with three techniques. First, we propose an algorithm to search coding bitmatrices with fewer 1’s from Vandermonde and Cauchy matrices, and further optimize the coding bitmatrices by greedily reducing the number of 1’s in the bitmatrices. So we can find near-optimal coding bitmatrices with the number of 1’s only up to 1% more than the lower bound. Next, we redesign the process of building pointers and reuse the pointers to access data for coding, which obtains a better tradeoff between spatial locality and computation efficiency. Finally, we smartly decompose the coding procedure of wide stripes into multiple subprocedures, to improve spatial locality and reduce the number of XORs. Based on the proposed techniques, we implement an erasure coding library, Cerasure. Extensive experiments show that Cerasure significantly improves the coding throughput. Compared with the state-of-the-art erasure coding libraries, Zerasure and SLPEC, Cerasure increases the encoding throughput by up to 200.2%.
Wei Wang 0502, Min Lyu, Tianyang Niu, Qiliang Li, Liangliang Xu, Yinlong Xu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 Toward Accurate Odor Identification and Effective Feature Learning With an AI-Empowered Electronic Nose
abstract
The development of Internet-of-Thing technology and robotics can be significantly promoted to a higher level with a precise digital sense of odors. However, the detection and identification of diverse odors using electronic sensor systems pose significant challenges. Electronic nose (E-Nose) based on gas sensors provides a cost-effective solution to detect odors. While previous research has primarily focused on enhancing the discriminative capability of E-Noses through machine learning techniques, limited attention has been given to the quality of features extracted or learned from E-Nose data. This paper presents a comprehensive study of E-nose design, digital odor measurement, and feature analysis methods in classifying many different odors into distinctive digital signatures. Experimental investigations involved the construction of a novel E-Nose system with automated data processing capabilities, enabling the study and discrimination of various odors, including essential oils, coffee, and whiskey. In the realm of data analysis, multiple feature extraction methods were compared and evaluated including 1Dconvolutional autoencoder (C-AE) and 1D convolutional neural network (1D-CNN). Motivated by recent progress in representation learning within the realm of face recognition, particularly its proficiency in generating distinctive features within the angular space, an algorithm incorporating a 1D-CNN model complemented by ArcLoss was developed. This innovative approach resulted in an exceptional classification accuracy of 97.76%, demonstrating its robustness in identifying and distinguishing a complex array of odors originating from a variety of essential oils.
Zhenyi Ye, Yaonian Li, Ruth Jin, Qiliang Li
IEEE Internet Things J.4
2024 Fast recovery for large disk enclosures based on RAID2.0: Algorithms and evaluation
Qiliang Li, Min Lyu, Liangliang Xu
J. Parallel Distributed Comput.1
2024 Enabling Efficient Erasure Coding in Disaggregated Memory Systems
abstract
Disaggregated memory (DM) separates compute and memory resources to build a huge memory pool. Erasure coding (EC) is expected to provide fault tolerance in DM with low memory cost. In DM with EC, objects are first coded in compute servers, then directly written to memory servers via high-speed networks like one-sided RDMA. However, as the one-sided RDMA latency goes down to the microsecond level, coding overhead degrades the performance in DM with EC. To enable efficient EC in DM, we thoroughly analyze the coding stack from the perspective of cache efficiency and RDMA transmission. We develop MicroEC, which optimizes the coding workflow by reusing the auxiliary coding data and coordinates the coding and RDMA transmission with an exponential pipeline, as well as carefully adjusting the coding and transmission threads to minimize the latency. We implement a prototype supporting common basic operations, such as write/read/degraded read/recovery. Experiments show that MicroEC reduces the write latency by up to 44.35% and 42.14% and achieves up to$1.80\times$and$1.73\times$write throughput, compared with the state-of-the-art DM systems with EC and 3-way replication for objects not smaller than 1 MB, respectively. For small objects, MicroEC also evidently reduces the variation of latency, e.g., it reduces the P99 latency of writing 1 KB objects by 27.81%.
Qiliang Li, Liangliang Xu, Yongkun Li 0001, Min Lyu, Wei Wang 0502, Pengfei Zuo, Yinlong Xu 0001
IEEE Trans. Parallel Distributed Syst.1
2023 Cerasure: Fast Acceleration Strategies For XOR-Based Erasure Codes
abstract
Erasure coding is a common redundancy scheme for tolerating failures in storage systems. Compared with replication, erasure coding saves a large amount of storage space, but incurs heavy computation overhead and thus is more time-consuming. To this end, we design an algorithm to find a better parity coding matrix to reduce the number of XORs in coding based on Vandermonde matrices instead of Cauchy matrices. In addition, we optimize the coding process, to accelerate the computation speed of XOR and obtain a better tradeoff between spatial locality and computation efficiency. For wide stripes which becomes increasingly interesting, we propose to decompose the coding procedure into multiple subprocedures for better utilization of spatial locality. We integrate these methods into coding procedure and implement an erasure coding library, Cerasure. Extensive experiments show that Cerasure significantly improves the coding speed. Compared with the state-of-the-art erasure coding libraries, Zerasure and SLPEC, Cerasure increases the encoding throughput by up to 109.47%.
Tianyang Niu, Min Lyu, Wei Wang 0502, Qiliang Li, Yinlong Xu 0001
ICCD4
2022 SelectiveEC: Towards Balanced Recovery Load on Erasure-Coded Storage Systems
abstract
Erasure coding (EC) has been commonly used to offer high data reliability with low storage cost. Upon failures, the lost blocks are recovered in batches. Due to the limited number of stripes, the data layout within a batch is non-uniform. Together with the random selection of source and replacement nodes for recovery tasks, the recovery workload among live nodes is skewed within a batch, which severely slows down failure recovery. To solve this problem, We present SelectiveEC, a new recovery task scheduling module that provides provable network traffic and recovery load balancing for large-scale EC-based storage systems. It relies on bipartite graphs to model the recovery traffic among live nodes. Then, it intelligently selects tasks to form batches and carefully determines where to read source blocks or to store recovered ones, using theories such as a perfect or maximum matching and$k$-regular spanning subgraph. SelectiveEC supports single-node failure and multi-node failure recovery, and can be deployed in both homogeneous and heterogeneous network environments. We implement SelectiveEC in HDFS, and evaluate its recovery performance in a local cluster of 18 nodes and AWS EC2 of 50 virtual machine instances. SelectiveEC increases the recovery throughput by up to$30.68\%$compared with state-of-the-art baselines in homogeneous network environments. It further achieves$1.32\times$recovery throughput and$1.23\times$benchmark throughput of HDFS on average in heterogeneous network environments, due to the straggler avoidance by the balanced scheduling.
Liangliang Xu, Min Lyu, Qiliang Li, Lingjiang Xie, Cheng Li 0001, Yinlong Xu 0001
IEEE Trans. Parallel Distributed Syst.3
2021 Fast Reconstruction for Large Disk Enclosures Based on RAID2.0
abstract
In the era of explosive data growth, RAID2.0 architecture with dozens or even hundreds of disks is commonly used to provide large capacity data storage. Due to limited resources, such as memory and CPU, the reconstruction for disk failures in RAID2.0 is executed in batches. Traditional random data placement and recovery scheme make the I/O access highly skewed within a batch, which slows down the reconstruction speed.
Qiliang Li, Min Lyu, Liangliang Xu, Yinlong Xu 0001, Wei Wang 0502
ICPP1
2020 SelectiveEC: Selective Reconstruction in Erasure-coded Storage Systems
Liangliang Xu, Min Lyu, Qiliang Li, Lingjiang Xie
HotStorage3
2020 LBBESA: An efficient software-defined networking load-balancing scheme based on elevator scheduling algorithm
abstract
Summary Elevator scheduling algorithms generally denote methods used to calculate how to use the elevator. These algorithms can distribute elevators to various floors of a building, thereby achieving efficient transportation. From the perspective of the elevator scheduling problem, we address the load‐balancing problem for software‐defined networking (SDN) architecture and propose a load‐balancing method based on the elevator scheduling algorithm, LBBESA. We take advantage of the flexibility of the SDN architecture, obtain the real‐time load of the server through real‐time statistical analyses of the SDN switch port traffic by the controller, and combine this with the idea of regional elevator allocation to coordinate the connection of the client's requests and realize the load balancing of each server in the cluster. Simulation experiments show that, compared with the round‐robin algorithm, LBBESA is more effective in the load balancing of the server pool and can improve the throughput of the server pool to a certain extent. In addition, our scheme is easy to implement and has high scalability.
Qiliang Li, Jie Cui 0004, Hong Zhong 0001, Yichao Du, Yonglong Luo, Lu Liu 0001
Concurr. Comput. Pract. Exp.1