EDBT 2026 Demo / reviewers in the wild / expert
Zerun Li
dblp:227/0341
· DBLP profile ↗
15ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-7672-0667ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Diffusion Planning with Temporal DiffusionabstractDiffusion planning is a promising method for learning high-performance policies from offline data. To avoid the impact of discrepancies between planning and reality on performance, previous works generate new plans at each time step. However, this incurs significant computational overhead and leads to lower decision frequencies, and frequent plan switching may also affect performance. In contrast, humans might create detailed short-term plans and more general, sometimes vague, long-term plans, and adjust them over time. Inspired by this, we propose the Temporal Diffusion Planner (TDP) which improves decision efficiency by distributing the denoising steps across the time dimension. TDP begins by generating an initial plan that becomes progressively more vague over time. At each subsequent time step, rather than generating an entirely new plan, TDP updates the previous one with a small number of denoising steps. This reduces the average number of denoising steps, improving decision efficiency. Additionally, we introduce an automated replanning mechanism to prevent significant deviations between the plan and reality. Experiments on D4RL show that, compared to previous works that generate new plans every time step, TDP significantly improves the decision-making frequency by 11-24.8 times while achieving higher or comparable performance. Jiaming Guo, Rui Zhang 0040, Zerun Li, Yunkai Gao 0001, Shaohui Peng, Siming Lan, Xing Hu 0001, Zidong Du, Xishan Zhang, Ling Li 0001 |
AAAI | 3 |
| 2026 | MACAM: A Flexible Computing-in-Memory Accelerator for Sparse Matrix-Dense Vector MultiplicationabstractSparse Matrix-Dense Vector Multiplication (SpMV) is an important computational primitive which is bounded by memory bandwidth. Computing-in-memory (CIM) is regarded as an effective approach to reduce data movement. Due to the lack of flexibility in architectural design, current CIM-based SpMV accelerators struggle to simultaneously support high-parallelism computations and the storage of irregular sparse data. We propose a flexible CIM-based accelerator named MACAM for high-precision SpMV. Each array of MACAM can be configured into sparse or dense modes according to the local-sparsity of the sparse matrix. We propose a unified data layout approach that enables MACAM to meet the data storage requirements of different modes. We also propose a sparse storage format and a workload-balancing approach to further improve the performance of MACAM. Experiments show that MACAM achieves 167.26× speedup and 286.04× energy saving over the GPU baseline. MACAM also achieves 97.41× and 6.56× speedup and 213.65× and 10.06× energy saving compared with two state-of-the-art CIM-based SpMV accelerators. Xiaoyu Zhang 0009, Rui Liu 0045, Zerun Li, Yinhe Han 0001, Xiaoming Chen 0003 |
DATE | 3 |
| 2026 | GPA: A General-Purpose In-Memory Computing Accelerator
Xiaoyu Zhang 0009, Zerun Li, Rui Liu 0045, Libo Shen, Boyu Long, Xueqi Li 0001, Yinhe Han 0001, Xiaoming Chen 0003 |
ISCAS | 2 |
| 2025 | CIM-BLAS: Computing-in-Memory Accelerator for BLASabstractBasic Linear Algebra Subprograms (BLAS) is a foundational software library for linear algebra kernels, which is widely used in scientific and engineering computing. Existing BLAS accelerations mainly rely on CPUs and GPUs. Many operations in BLAS are data intensive, so they are constrained by the limited memory bandwidth of CPUs and GPUs. The computing-in-memory (CIM) technology can effectively alleviate the memory wall bottleneck and is particularly suitable for accelerating BLAS. In this paper, we propose the first CIM accelerator for BLAS, CIM-BLAS, based on non-volatile memory. CIM-BLAS includes a unified floating-point pipeline to support high-precision arithmetics. High efficiency of the accelerator is achieved by developing configurable data flows to support various BLAS functions. Compared with GPU implementations, CIMBLAS demonstrates several orders of magnitude performance and energy efficiency improvements for executing level-1 and level-2 BLAS functions, and can achieve an energy efficiency improvement of 2.6-24.1 $\times$ for executing level-3 BLAS functions. The improvement increases with the size of the matrix, indicating excellent scalability of CIM-BLAS. Application-level evaluations also demonstrate the potential of CIM for accelerating BLAS. Rui Liu 0045, Zerun Li, Xiaoyu Zhang 0009, Xiaoming Chen 0003, Yinhe Han 0001, Minghua Tang |
DAC | 2 |
| 2025 | Statistical significance of cluster membership for categorical data
Lianyu Hu 0001, Zerun Li, Mudi Jiang, Zengyou He |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Conjunction subspaces test for conformal and selective classification
Zengyou He, Zerun Li, Mudi Jiang, Lianyu Hu 0001 |
Inf. Sci. | 2 |
| 2025 | A Data-Centric Software-Hardware Co-Designed Architecture for Large-Scale Graph ProcessingabstractGraph processing plays an important role in many practical applications. However, the inherent characteristics of graph processing, including random memory access and the low computation-to-communication ratio, make it difficult to efficiently execute on traditional computing architectures, such as CPUs and GPUs. Near-memory computing has the characteristics of low latency and high bandwidth. It is widely regarded as a promising direction for designing graph processing accelerators. However, the storage space of a single device cannot meet the demand of large-scale graph processing. Using multiple devices will bring lots of inter-device data transmission, which may counteract the benefits of near-memory computing. To fundamentally reduce the data transmission overhead, we propose a data-centric graph processing framework for systems with multiple near-memory computing devices. The framework uses a data-centric programming model as the software hardware interface. For software, we propose an optimized data flow and a heuristic multi-step weighted maximum matching algorithm to achieve efficient inter-device communication and ensure load balancing. For hardware, we design a data reuse driven task controller and a data type-aware on-chip memory, which can effectively improve the utilization of the on-chip memory. Compared with the two most recent near-memory graph accelerators, our framework significantly reduces energy consumption and inter-device communication. Zerun Li, Xiaoming Chen 0003, Yuxin Yang 0002, Feng Min, Xiaoyu Zhang 0009, Yinhe Han 0001 |
IEEE Trans. Computers | 1 |
| 2024 | TMiner: A Vertex-Based Task Scheduling Architecture for Graph Pattern MiningabstractGraph pattern mining discovers important patterns in graphs. It is both computation-and memory-intensive, characterized by numerous set operations and irregular memory access. Graph pattern mining inherently involves a large number of independent tasks, helping to alleviate its computational bottleneck through parallel processing. However, after exploiting parallelism, memory access will become the primary bottleneck. Existing parallelism strategies severely result in redundant and inefficient memory access, making the performance memory bounded. This paper proposes TMiner, a graph pattern mining architecture with optimized memory performance through a systematically designed software-hardware stack. TMiner fundamentally reduces redundant memory access of parallel graph pattern mining in three aspects. (1) TMiner leverages a task partitioning approach based on disjoint neighbor vertex set access, reducing redundant memory access between PEs. (2) TMiner utilizes the global neighbor vertex information to coalesce the access from different neighbor vertex subsets at compilation time, which not only reduces redundant memory access within a task but also improves the data locality. (3) TMiner adopts a data reuse-oriented task scheduling mechanism, which dynamically migrates and merges tasks with similar memory access patterns, reducing redundant memory access within a PE at runtime. A DIMM-based near-memory architecture that exploits the DRAM's internal bandwidth is elaborated for high-performance graph pattern mining, which incorporates the proposed memory access optimization techniques and an extended ISA. Compared with the state-of-the-art software and hardware baselines, TMiner significantly improves the performance. Zerun Li, Xiaoming Chen 0003, Yinhe Han 0001 |
MICRO | 1 |
| 2024 | GAS: General-Purpose In-Memory-Computing Accelerator for Sparse Matrix MultiplicationabstractSparse matrix multiplication is widely used in various practical applications. Different accelerators have been proposed to speed up sparse matrix-dense vector multiplication (SpMV), sparse matrix-sparse vector multiplication (SpMSpV), sparse matrix-dense matrix multiplication (SpMM), and sparse matrix-sparse matrix multiplication (SpMSpM). The performance of traditional sparse matrix multiplication accelerators is typically bounded by memory access due to the poor data locality and irregular memory access. In-memory computing (IMC) is a promising technique to alleviate the memory bottleneck. Previous IMC studies are mostly focused on accelerating a single sparse matrix multiplication function. In this paper, we propose GAS, a general-purpose IMC accelerator for sparse matrix multiplication. GAS integrates non-volatile memory based content-addressable memory (CAM) arrays and multiply-add computation (MAC) arrays to support sparse matrices represented in the double-precision floating-point format. Using a unified outer product based multiplication methodology, GAS supports the acceleration of SpMV, SpMSpv, SpMM, and SpMSpM. We further propose four optimization techniques to speed up the computation of GAS. GAS achieves significant speedups and energy savings over central processing unit (CPU) and graphics processing unit (GPU) implementations. Compared with state-of- the-art traditional and IMC-based accelerators, GAS not only supports more functions, but also achieves higher performance and energy efficiency. Xiaoyu Zhang 0009, Zerun Li, Rui Liu 0045, Xiaoming Chen 0003, Yinhe Han 0001 |
IEEE Trans. Computers | 2 |
| 2023 | FSPA: An FeFET-based Sparse Matrix-Dense Vector Multiplication AcceleratorabstractSparse matrix-dense vector multiplication (SpMV) is widely used in various applications. The performance of traditional SpMV accelerators is bounded by memory. In-memory computing (IMC) is a promising technique to alleviate the memory bottleneck. The current IMC accelerator cannot support sparse storage format and in-situ floating-point multiplication at the same time. In this paper, we propose FSPA, an ferroelectric field-effect transistor (FeFET) based SpMV accelerator. FSPA integrates novel content-addressable memory (CAM) arrays and multiply-add computation (MAC) arrays to support sparse matrices represented in the floating-point format. FSPA achieves significant speedups and energy savings over CPU, GPU and two state-of-the-art IMC accelerators. Xiaoyu Zhang 0009, Zerun Li, Rui Liu 0045, Xiaoming Chen 0003, Yinhe Han 0001 |
DAC | 2 |
| 2023 | FeCrypto: Instruction Set Architecture for Cryptographic Algorithms Based on FeFET-Based In-Memory ComputingabstractRecently, computing-in-memory (CiM) becomes a promising technology for alleviating the memory wall bottleneck. CiM is suitable for data-intensive applications, especially, cryptographic algorithms. Most current cryptographic accelerators are specific to a single function. It is expensive to accelerate different cryptographic algorithms with different accelerators. In this work, we first introduce a CiM architecture FeMIC that supports multioperand CiM operations, by exploring advantages of state-of-the-art ferroelectric field-effect transistors. Based on that, we propose a novel instruction set together with an accelerator architecture named FeCrypto which supports the acceleration of various cryptographic algorithms. Evaluation results show that FeCrypto has better performance and energy efficiency than software implementations. The energy-delay product (EDP) of FeCrypto is$118.4\times $and$1.93\times $lower than that of the dedicated AES accelerator AIM that is built based on phase-change memories (PCMs) and magnetic random-access memories (MRAMs), respectively. EDP is reduced by$44.7\times $compared with PCM-based EIM, a recent AES accelerator. Compared with MRAM-based EIM, the EDP overhead of FeCrypto for supporting multiple functions is 23.2%. Rui Liu 0045, Xiaoyu Zhang 0009, Zhiwen Xie, Xinyu Wang 0040, Zerun Li, Xiaoming Chen 0003, Yinhe Han 0001, Minghua Tang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Optimal Data Allocation for Graph Processing in Processing-in-Memory SystemsabstractGraph processing involves lots of irregular memory accesses and increases demands on high memory bandwidth, making it difficult to execute efficiently on compute-centric architectures. Dedicated graph processing accelerators based on the processing-in-memory (PIM) technique have recently been proposed. Despite they achieved higher performance and energy efficiency than conventional architectures, the data allocation problem for communication minimization in PIM systems (e.g., hybrid memory cubes (HMCs)) has still not been well solved. In this paper, we demonstrate that the conventional “graph data allocation = graph partitioning” assumption is not true, and the memory access patterns of graph algorithms should also be taken into account when partitioning graph data for communication minimization. For this purpose, we classify graph algorithms into two representative classes from a memory access pattern point of view and propose different graph data partitioning strategies for them. We then propose two algorithms to optimize the partition-to-HMC mapping to minimize the inter-HMC communication. Evaluations have proved the superiority of our data allocation framework and the data movement energy efficiency is improved by 4.2-5 × on average than the state-of-the-art GraphP approach. Zerun Li, Xiaoming Chen 0003, Yinhe Han 0001 |
ASP-DAC | 1 |
| 2022 | GraphRing: an HMC-ring based graph processing framework with optimized data movementabstractDue to the irregular memory access and high bandwidth demanding, graph processing is usually inefficient on conventional computer architectures. The recent development of the processing-in-memory (PIM) technique such as hybrid memory cube (HMC) has provided a feasible design direction for graph processing accelerators. Although PIM provides high internal bandwidth, inter-node memory access is inevitable in large-scale graph processing, which greatly affects the performance. In this paper, we propose an HMC-based graph processing framework, GraphRing. GraphRing is a software-hardware codesign framework that optimizes inter-HMC communication. It contains a regularity- and locality-aware graph execution model and a ring-based multi-HMC architecture. The evaluation results based on 5 graph datasets and 4 graph algorithms show that GraphRing achieves on average 2.14× speedup and 3.07× inter-HMC communication energy saving, compared with GraphQ, a state-of-the-art graph processing architecture. Zerun Li, Xiaoming Chen 0003, Yinhe Han 0001 |
DAC | 1 |
| 2022 | Convolutional Neural Network Accelerator for Compression Based on Simon k-meansabstractConvolutional Neural Networks (CNN) are popular models widely used in image classification, target recognition, and other fields. FPGA-based accelerators for CNN are a standard method in recent years to reduce CNN's inference time and energy efficiency. However, the limitations of on-chip storage space and computing resources introduce deep compression. Contrary to most compression algorithms that pay no attention to the underlying hardware acceleration strategy and hardware-only accelerators, this paper introduces a novel model compression scheme with software and hardware collaboration for accelerating inference. First, we propose a pre-processing algorithm named Simon k-means based on clustering to quantify trained weight to speed up inference. Next, we propose a new encoding method for the quantized weight, significantly reducing the model's storage size. Finally, we give the architecture design of the accelerator using the quantized weight to accelerate the convolution. We have evaluated many popular CNNs in image classification tasks on various data sets. Experiments show that the number of multiply-accumulate operations on the convolutional layer can be reduced 66.6% with a slight loss of precision. Yunping Zhao, Jianzhuang Lu, Zerun Li |
IJCNN | 6 |
| 2020 | Practical AMC model based on SAE with various optimisation methods under different noise environmentsabstractAutomatic modulation classification (AMC) has recently attracted widespread attention nowadays due to its desirable features of generalisability and requirement of little prior knowledge through artificial intelligence (AI) technology. The authors propose a stacked auto‐encoder (SAE) based on various optimisation methods structure to intelligently process a feature space that includes spectral‐based features and high‐order cumulants. To unify the dimensionality of the features, they apply different normalisation methods to the feature space before training the SAE model to decide corresponding normalisations under different noise environments. Linear normalisation is superior when signal‐to‐noise ratio (SNR) is low, and standardisation is superior when SNR is between ‐1 and 4 dB. Regularisation works best when SNR is greater than 5 dB. To increase the recognition accuracy of the proposed model, they introduce the unconstrained optimisation theory to adjust the proposed SAE model, including Nelder‐Mead method, Newton optimisation method, conjugate gradient method and quasi‐Newton method. They observe that the quasi‐Newton method offers desirable performance when optimising SAE model. It is the first time to compare these data normalisation methods and discuss unconstrained optimisation theory together to recognise modulation types. The recognition accuracy of this model for eight modulation types can reach 99.8% when SNR ranges from to 10 dB. Zerun Li, Weisong Liu, Zuocheng Xing, Yongzhong Li |
IET Commun. | 1 |