VLDB 2026 Research / reviewers in the wild / expert
Rui Liu 0045
dblp:42/469-45
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-6515-652XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 11 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MACAM: A Flexible Computing-in-Memory Accelerator for Sparse Matrix-Dense Vector MultiplicationabstractSparse Matrix-Dense Vector Multiplication (SpMV) is an important computational primitive which is bounded by memory bandwidth. Computing-in-memory (CIM) is regarded as an effective approach to reduce data movement. Due to the lack of flexibility in architectural design, current CIM-based SpMV accelerators struggle to simultaneously support high-parallelism computations and the storage of irregular sparse data. We propose a flexible CIM-based accelerator named MACAM for high-precision SpMV. Each array of MACAM can be configured into sparse or dense modes according to the local-sparsity of the sparse matrix. We propose a unified data layout approach that enables MACAM to meet the data storage requirements of different modes. We also propose a sparse storage format and a workload-balancing approach to further improve the performance of MACAM. Experiments show that MACAM achieves 167.26× speedup and 286.04× energy saving over the GPU baseline. MACAM also achieves 97.41× and 6.56× speedup and 213.65× and 10.06× energy saving compared with two state-of-the-art CIM-based SpMV accelerators. Xiaoyu Zhang 0009, Rui Liu 0045, Zerun Li, Yinhe Han 0001, Xiaoming Chen 0003 |
DATE | 2 |
| 2026 | GPA: A General-Purpose In-Memory Computing Accelerator
Xiaoyu Zhang 0009, Zerun Li, Rui Liu 0045, Libo Shen, Boyu Long, Xueqi Li 0001, Yinhe Han 0001, Xiaoming Chen 0003 |
ISCAS | 3 |
| 2025 | CIM-BLAS: Computing-in-Memory Accelerator for BLASabstractBasic Linear Algebra Subprograms (BLAS) is a foundational software library for linear algebra kernels, which is widely used in scientific and engineering computing. Existing BLAS accelerations mainly rely on CPUs and GPUs. Many operations in BLAS are data intensive, so they are constrained by the limited memory bandwidth of CPUs and GPUs. The computing-in-memory (CIM) technology can effectively alleviate the memory wall bottleneck and is particularly suitable for accelerating BLAS. In this paper, we propose the first CIM accelerator for BLAS, CIM-BLAS, based on non-volatile memory. CIM-BLAS includes a unified floating-point pipeline to support high-precision arithmetics. High efficiency of the accelerator is achieved by developing configurable data flows to support various BLAS functions. Compared with GPU implementations, CIMBLAS demonstrates several orders of magnitude performance and energy efficiency improvements for executing level-1 and level-2 BLAS functions, and can achieve an energy efficiency improvement of 2.6-24.1 $\times$ for executing level-3 BLAS functions. The improvement increases with the size of the matrix, indicating excellent scalability of CIM-BLAS. Application-level evaluations also demonstrate the potential of CIM for accelerating BLAS. Rui Liu 0045, Zerun Li, Xiaoyu Zhang 0009, Xiaoming Chen 0003, Yinhe Han 0001, Minghua Tang |
DAC | 1 |
| 2024 | MemSort: In-Memory Sorting ArchitectureabstractSorting is one of the most fundamental operations in computer programming and used in countless algorithms. The performance of traditional von Neumann computers running sorting is limited by the bandwidth between memories and processors. Computing-in-memory (CiM) is a promising technology which has the potential to solve the “memory wall” bottleneck. CiM is suitable for data-intensive applications, and it is ideal for accelerating large-scale data sorting. In this paper, we propose a novel in-memory sorting accelerator, named MemSort, based on a proposed in-memory comparison array design based on emerging non-volatile devices. MemSort supports three sort operations including counting sort, merging sort, and the combination of counting sort and merging sort. We build a performance model for the combination sort which enables flexible allocation of resources under given constraints to meet the requirements of various applications for sorting. The evaluation results show that MemSort shows significant performance improvement and energy efficiency at both the system level and application level when processing large-scale data sorting. Compared with the CPU implementation, MemSort achieves energy savings of 19.69-72.75x and speedups of 24.48-38.58 x with the same power constraint. MemSort's throughput is at least 4.86 x higher than that of the recent FPGA-based sorting accelerator FANS. MemSort exhibits more than 11 x throughput and 4.03 x area efficiency, compared with the recent CiM - based sorting accelerator, RIME. Rui Liu 0045, Xiaoyu Zhang 0009, Xinyu Wang 0040, Feng Min, Zhejian Luo, Xiaoming Chen 0003, Yinhe Han 0001, Minghua Tang |
ICCD | 1 |
| 2024 | GAS: General-Purpose In-Memory-Computing Accelerator for Sparse Matrix MultiplicationabstractSparse matrix multiplication is widely used in various practical applications. Different accelerators have been proposed to speed up sparse matrix-dense vector multiplication (SpMV), sparse matrix-sparse vector multiplication (SpMSpV), sparse matrix-dense matrix multiplication (SpMM), and sparse matrix-sparse matrix multiplication (SpMSpM). The performance of traditional sparse matrix multiplication accelerators is typically bounded by memory access due to the poor data locality and irregular memory access. In-memory computing (IMC) is a promising technique to alleviate the memory bottleneck. Previous IMC studies are mostly focused on accelerating a single sparse matrix multiplication function. In this paper, we propose GAS, a general-purpose IMC accelerator for sparse matrix multiplication. GAS integrates non-volatile memory based content-addressable memory (CAM) arrays and multiply-add computation (MAC) arrays to support sparse matrices represented in the double-precision floating-point format. Using a unified outer product based multiplication methodology, GAS supports the acceleration of SpMV, SpMSpv, SpMM, and SpMSpM. We further propose four optimization techniques to speed up the computation of GAS. GAS achieves significant speedups and energy savings over central processing unit (CPU) and graphics processing unit (GPU) implementations. Compared with state-of- the-art traditional and IMC-based accelerators, GAS not only supports more functions, but also achieves higher performance and energy efficiency. Xiaoyu Zhang 0009, Zerun Li, Rui Liu 0045, Xiaoming Chen 0003, Yinhe Han 0001 |
IEEE Trans. Computers | 3 |
| 2023 | FSPA: An FeFET-based Sparse Matrix-Dense Vector Multiplication AcceleratorabstractSparse matrix-dense vector multiplication (SpMV) is widely used in various applications. The performance of traditional SpMV accelerators is bounded by memory. In-memory computing (IMC) is a promising technique to alleviate the memory bottleneck. The current IMC accelerator cannot support sparse storage format and in-situ floating-point multiplication at the same time. In this paper, we propose FSPA, an ferroelectric field-effect transistor (FeFET) based SpMV accelerator. FSPA integrates novel content-addressable memory (CAM) arrays and multiply-add computation (MAC) arrays to support sparse matrices represented in the floating-point format. FSPA achieves significant speedups and energy savings over CPU, GPU and two state-of-the-art IMC accelerators. Xiaoyu Zhang 0009, Zerun Li, Rui Liu 0045, Xiaoming Chen 0003, Yinhe Han 0001 |
DAC | 3 |
| 2023 | LIM-GEN: A Data-Guided Framework for Automated Generation of Heterogeneous Logic-in-Memory ArchitectureabstractMemristor-based logic-in-memory (LIM) is an emerging technology that enables logic operations within memory, making it a promising solution for data-intensive applications. LIM architectures have different types according to where computations are executed, with each type being suitable for specific design objectives and application domains. However, mapping applications to a single LIM mode restricts the full utilization of different LIM modes. In this paper, we propose LIM-GEN, a data-guided framework for automated generation of heterogeneous LIM architectures. To take advantages of different LIM modes, three LIM modes are combined and used as building blocks to create heterogeneous architectures. Given the data-centric nature and large design space, there is an urgent need of developing new EDA tools for synthesizing such LIM architectures. LIM-GEN includes an automatic hardware synthesis flow, which takes behavior-level descriptions as input to generate application-specific architectures and dataflows. During synthesis, data distribution, task allocation and crossbar mapping are optimized through a design space exploration process. We evaluate LIM-GEN in several data-intensive applications and compare the generated heterogeneous architectures with synthesized architectures with a single LIM mode. The experimental results demonstrate significant improvements in latency, area and power consumption, brought by the heterogeneous architectures generated by LIM-GEN. Libo Shen, Boyu Long, Rui Liu 0045, Xiaoyu Zhang 0009, Yinhe Han 0001, Xiaoming Chen 0003 |
ICCAD | 3 |
| 2023 | Hardware-Software Co-Design for Content-Based Sparse AttentionabstractAttention-based pre-trained large models have demonstrated impressive performance in many domains such as natural language processing and computer vision. Unfortunately, due to the quadratic complexity incurred by calculating pairwise correlations across the entire input sequence, processing the attention mechanism becomes the arguably major bottleneck of the whole inference execution. To accelerate the attention mechanism with no loss of accuracy, we present a novel algorithm-architecture co-design that can substantially save runtime as well as energy spent on the attention mechanism. Inspired by the observation that only a small subset of content highly correlates with the others under attention, we devise a hardware-friendly content-based sparsity scheme to eliminate unnecessary relations, thus reducing computation complexity effectively. Furthermore, we develop a tailored hardware for this content-based sparse attention mechanism to best utilize this algorithm innovation. Experiments show that, compared with the implementation based on an Nvidia V100-SXM2 GPU, on average, our design achieves 63× speedup and 505× energy saving with no accuracy loss. Xiaoyu Zhang 0009, Rui Liu 0045, Zhejian Luo, Xiaoming Chen 0003, Yinhe Han 0001 |
ICCD | 3 |
| 2023 | FeCrypto: Instruction Set Architecture for Cryptographic Algorithms Based on FeFET-Based In-Memory ComputingabstractRecently, computing-in-memory (CiM) becomes a promising technology for alleviating the memory wall bottleneck. CiM is suitable for data-intensive applications, especially, cryptographic algorithms. Most current cryptographic accelerators are specific to a single function. It is expensive to accelerate different cryptographic algorithms with different accelerators. In this work, we first introduce a CiM architecture FeMIC that supports multioperand CiM operations, by exploring advantages of state-of-the-art ferroelectric field-effect transistors. Based on that, we propose a novel instruction set together with an accelerator architecture named FeCrypto which supports the acceleration of various cryptographic algorithms. Evaluation results show that FeCrypto has better performance and energy efficiency than software implementations. The energy-delay product (EDP) of FeCrypto is$118.4\times $and$1.93\times $lower than that of the dedicated AES accelerator AIM that is built based on phase-change memories (PCMs) and magnetic random-access memories (MRAMs), respectively. EDP is reduced by$44.7\times $compared with PCM-based EIM, a recent AES accelerator. Compared with MRAM-based EIM, the EDP overhead of FeCrypto for supporting multiple functions is 23.2%. Rui Liu 0045, Xiaoyu Zhang 0009, Zhiwen Xie, Xinyu Wang 0040, Zerun Li, Xiaoming Chen 0003, Yinhe Han 0001, Minghua Tang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | FeMIC: Multi-Operands in-Memory Computing Based on FeFETsabstractThe “memory wall” bottleneck caused by the performance gap between processors and memories is getting worse. Computing-in-memory (CiM), a promising technology to alleviate the “memory wall” bottleneck, has recently attracted much attention. Conventional CiM architectures based on emerging nonvolatile devices have a major drawback that they need${N\,-\,1}$clock cycles to complete a CiM operation with${N}$operands, as they are natively designed for processing two operands. In this work, we propose FeMIC, a new CiM architecture based on ferroelectric field-effect transistors (FeFETs), which natively supports the computation of multiple operands. For a CiM operation with${N}$operands, FeMIC only needs$\left\lfloor {N/2} \right\rfloor$clock cycles. The simulation results based on a calibrated FeFET model reveal that FeMIC can significantly reduce the energy consumption when processing multi-operand CiM operations, compared with state-of-the-arts that use conventional CiM mechanisms. Rui Liu 0045, Xiaoyu Zhang 0009, Xiaoming Chen 0003, Yinhe Han 0001, Minghua Tang |
ASP-DAC | 1 |
| 2022 | Re-FeMAT: A Reconfigurable Multifunctional FeFET-Based Memory ArchitectureabstractMost of current processing-in-memory (PIM) architectures are application specific, that is, they can only accelerate particular functions, e.g., matrix-vector dot product for neural network acceleration. However, practical applications usually involve various functions. In order to accelerate different functions, various accelerators, and dedicated circuits have been proposed. In this work, by exploring the similarities among some commonly used dedicated circuits, we adopt ferroelectric field-effect transistors (FeFETs) to build a reconfigurable multifunctional memory architecture named Re-FeMAT. Re-FeMAT is composed of multiple processing elements (PEs). Each PE is not only a nonvolatile memory array, but also can perform logic operations (i.e., the PIM mode), convolutions (i.e., the binary convolutional neural network and the convolutional neural network (CNN) acceleration mode) and content search (i.e., the ternary content-addressable memory (TCAM) mode) without changing the circuit structure. Re-FeMAT can support applications that require multiple functions. As an example, by configuring different PEs to different working modes and using a simulated annealing algorithm or a tabu search algorithm to optimize the task-PE assignment, Re-FeMAT can completely accelerate few-shot learning applications. Our simulation results based on a calibrated FeFET model show that the proposed Re-FeMAT architecture achieves better performance and power efficiency than the previous FeMAT architecture. Compared with FeFET-based single-functional circuits, though the power dissipation of Re-FeMAT is higher in some modes, the power-delay product is still smaller. Compared with a state-of-the-art FeFET-based multifunctional accelerator named attention-in-memory, Re-FeMAT achieves lower power, latency, and energy when accelerating a complete few-shot learning task. Xiaoyu Zhang 0009, Rui Liu 0045, Yuxin Yang 0002, Yinhe Han 0001, Xiaoming Chen 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |