VLDB 2026 Research / reviewers in the wild / expert
Ziru Li
dblp:147/6384
· DBLP profile ↗
11ranked-venue papers
4as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ModSRAM: Algorithm-Hardware Co-Design for Large Number Modular Multiplication in SRAMabstractElliptic curve cryptography (ECC) is widely used in security applications such as public key cryptography (PKC) and zero-knowledge proofs (ZKP). ECC is composed of modular arithmetic, where modular multiplication takes most of the processing time. Computational complexity and memory constraints of ECC limit the performance. Therefore, hardware acceleration on ECC is an active field of research. Processing-in-memory (PIM) is a promising approach to tackle this problem. In this work, we design ModSRAM, the first 8T SRAM PIM architecture to compute large-number modular multiplication efficiently. In addition, we propose R4CSA-LUT, a new algorithm that reduces the cycles for an interleaved algorithm and eliminates carry propagation for addition based on look-up tables (LUT). ModSRAM is co-designed with R4CSA-LUT to support modular multiplication and data reuse in memory with 52% cycle reduction compared to prior works with only 32% area overhead. Jonathan Hao-Cheng Ku, Junyao Zhang 0003, Haoxuan Shan, Saichand Samudrala, Jiawen Wu 0006, Qilin Zheng, Ziru Li, Jeyavijayan Rajendran, Yiran Chen 0001 |
DAC | 7 |
| 2024 | Improving the Efficiency of In-Memory-Computing Macro with a Hybrid Analog-Digital Computing Mode for Lossless Neural Network InferenceabstractAnalog in-memory-computing (IMC) is an attractive technique with a higher energy efficiency to process machine learning workloads. However, the analog computing scheme suffers from large interface circuit overhead. In this work, we propose a macro with a hybrid analog-digital mode computation to reduce the precision requirement of the interface circuit. Considering the distribution of the multiplication and accumulation (MAC) value, we propose a nonlinear transfer function of the computing circuits by only accurately computing low MAC value in the analog domain with a digital mode to deal with the high MAC value with smaller possibility. Silicon measurement results show that the proposed macro could achieve 160 GOPS/mm2 area efficiency and 25.5 TOPS/W for 8b/8b matrix computation. The architectural-level evaluation for real workloads shows that the proposed macro can achieve up to 2.92× higher energy efficiency than conventional analog IMC designs. Qilin Zheng, Ziru Li, Jonathan Hao-Cheng Ku, Yitu Wang, Brady Taylor, Deliang Fan, Yiran Chen 0001 |
DAC | 2 |
| 2024 | Hybrid Digital/Analog Memristor-based Computing Architecture for Sparse Deep Learning AccelerationabstractFine-grained sparsity in recent bio-inspired models such as attention-based model could reduce the computation complexity dramatically. However, the unique sparsity pattern challenges the mapping efficiency of the conventional pure analog memristor-based computing architecture, as the conventional one uses a vector-matrix-multiplication primitives. To fill the gap between the memristor-based architecture and the sparse processing, in this paper, we would like to present our recent progress by using a hybrid digital/analog memristor-based computing architecture to improve the mapping efficiency. Our evaluation result shows that, over previous pure analog memristor-based architecture, our design could deliver up to 8.32× performance improvement and 3.4× energy efficiency improvement on a range of vision and language tasks for the recent attention-based bio-inspired model. Qilin Zheng, Shiyu Li 0001, Yitu Wang, Ziru Li, Yiran Chen 0001, Hai Li 0001 |
ISCAS | 4 |
| 2023 | Accelerating Sparse Attention with a Reconfigurable Non-volatile Processing-In-Memory ArchitectureabstractAttention-based neural networks have shown superior performance in a wide range of tasks. Non-volatile processing-in-memory (NVPIM) architecture shows its great potential to accelerate the dense attention model. However, the unique unstructured and dynamic sparsity pattern in the sparse attention model challenges the mapping efficiency of the NVPIM architecture, as the conventional NVPIM architecture uses a vector-matrix-multiplication primitives. In this paper, we propose a NVPIM architecture to accelerate a dynamic and unstructured sparse computation in the sparse attention. We aim to improve the mapping efficiency for both SDDMM and SpMM by introducing two vector-based primitives with a reconfigurable NVPIM bank. Further, based on our reconfigurable NVPIM bank, we further propose a hybrid stationary data flow to hide the latency. Our evaluation result shows that, over previous NVPIM accelerators, our design could deliver up to 12.36× performance improvement and 3.4× energy efficiency improvement on a range of vision and language tasks. Qilin Zheng, Shiyu Li 0001, Yitu Wang, Ziru Li, Yiran Chen 0001, Hai Li 0001 |
DAC | 4 |
| 2022 | ASTERS: adaptable threshold spike-timing neuromorphic design with twin-column ReRAM synapsesabstractComplex event-driven neuron dynamics was an obstacle to implementing efficient brain-inspired computing architectures with VLSI circuits. To solve this problem and harness the event-driven advantage, we propose ASTERS, a resistive random-access memory (ReRAM) based neuromorphic design to conduct the time-to-first-spike SNN inference. In addition to the fundamental novel axon and neuron circuits, we also propose two techniques through hardware-software co-design: "Multi-Level Firing Threshold Adjustment" to mitigate the impact of ReRAM device process variations, and "Timing Threshold Adjustment" to further speed up the computation. Experimental results show that our cross-layer solution ASTERS achieves more than 34.7% energy savings compared to the existing spiking neuromorphic designs, meanwhile maintaining 90.1% accuracy under the process variations with a 20% standard deviation. Ziru Li, Qilin Zheng, Bonan Yan, Ru Huang 0001, Bing Li 0005, Yiran Chen 0001 |
DAC | 1 |
| 2021 | Neuromorphic Algorithm-hardware Codesign for Temporal Pattern LearningabstractNeuromorphic computing and spiking neural networks (SNN) mimic the behavior of biological systems and have drawn interest for their potential to perform cognitive tasks with high energy efficiency. However, some factors such as temporal dynamics and spike timings prove critical for information processing but are often ignored by existing works, limiting the performance and applications of neuromorphic computing. On one hand, due to the lack of effective SNN training algorithms, it is difficult to utilize the temporal neural dynamics. Many existing algorithms still treat neuron activation statistically. On the other hand, utilizing temporal neural dynamics also poses challenges to hardware design. Synapses exhibit temporal dynamics, serving as memory units that hold historical information, but are often simplified as a connection with weight. Most current models integrate synaptic activations in some storage medium to represent membrane potential and institute a hard reset of membrane potential after the neuron emits a spike. This is done for its simplicity in hardware, requiring only a “clear” signal to wipe the storage medium, but destroys temporal information stored in the neuron.In this work, we derive an efficient training algorithm for Leaky Integrate and Fire neurons, which is capable of training a SNN to learn complex spatial temporal patterns. We achieved competitive accuracy on two complex datasets. We also demonstrate the advantage of our model by a novel temporal pattern association task. Codesigned with this algorithm, we have developed a CMOS circuit implementation for a memristor-based network of neuron and synapses which retains critical neural dynamics with reduced complexity. This circuit implementation of the neuron model is simulated to demonstrate its ability to react to temporal spiking patterns with an adaptive threshold. Haowen Fang, Brady Taylor, Ziru Li, Zaidao Mei, Hai Li 0001, Qinru Qiu |
DAC | 3 |
| 2021 | A model of two-zoned network for platform competition
Ziru Li, Benjamin B. M. Shao, Michael Goul |
Decis. Support Syst. | 1 |
| 2021 | Reducing SRAM Reading Power With Column Data Segment and Weights Correlation Enhancement for CNN ProcessingabstractConvolutional neural network (CNN) has been widely deployed in various processors for intelligent visual signal processing. However, the large amount of activations and weights in CNN causes huge power consumption on SRAM access. Data-adaptive SRAM design is a widely studied method to reduce SRAM reading power based on the utilization of data patterns, while current designs only exploit data patterns in a coarse granularity, and have no advantages when faced with randomly distributed weight data. In this article, we propose a hardware–software co-design scheme to reduce SRAM reading power for CNN processing. First, we propose a reconfigurable data-adaptive SRAM architecture with column data segmentation (CDS-RSRAM) to utilize data patterns. Data in one column is partitioned into several segments, and finer-grained data patterns are exploited within each segment for further reading power reduction. Then, a novel training method—minimum segmented neighbor difference (miniSND)—is proposed for enhancing the correlation of weights. MiniSND improves the similarity of weights without classification accuracy degradation, thus weights could benefit from CDS-RSRAM and be read out with less power consumption. Simulation results demonstrate that the co-design scheme saves up to 66%(8b)/89%(2b) power consumption compared with 8T SRAM. Han Xu 0006, Ziru Li, Deliang Fan, Fei Qiao, Qi Wei 0001, Huazhong Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | ReSiPE: ReRAM-based Single-Spiking Processing-In-Memory EngineabstractProcessing-in-memory (PIM) designs that leverage emerging nanotechnologies like resistive random access memory (ReRAM) have demonstrated enormous potential in accelerating deep learning applications due to high energy efficiency and integration density. The common approach of existing ReRAM-based PIM designs is to encode data into either voltage levels with the assist of power-thirsty analog/digital conversion circuits or spike series by sacrificing computing latency. In this paper, we introduce ReSiPE, a ReRAM-based Single-spiking PIM Engine, which uses the arrival time of a single spike to represent the data. We analyze how to encode data into a set of single spikes and develop the circuit to realize the matrix-based computation. The proposed design can minimize the spike numbers, shorten the computation period, and thus improve energy efficiency dramatically. Our simulation results show that ReSiPE achieves 67.1% power reduction and 1.97× power efficiency improvement compared to rate-coding based ReRAM PIM designs under the comparable area, throughput, and accuracy. Ziru Li, Bonan Yan, Hai Li 0001 |
DAC | 1 |
| 2020 | RED: A ReRAM-Based Efficient Accelerator for Deconvolutional ComputationabstractDeconvolution is a key component in contemporary neural networks, especially, generative adversarial networks (GANs) and fully convolutional networks (FCNs). Due to extra operations of deconvolution compared to convolution, considerable degradation of performance, as well as energy efficiency is incurred when implementing deconvolution on the existing resistive random access memory (ReRAM)-based processing-in-memory (PIM) accelerators. In this article, we propose an ReRAM-based accelerator design, RED, for providing high-performance and low-energy deconvolution. We analyze the deconvolution execution on the existing ReRAM-based PIMs and utilize its interior computation pattern for design optimization. RED includes two major contributions: 1) pixel-wise mapping scheme and 2) zero-skipping data flow. Pixel-wise mapping scheme removes the zero insertion and performs convolutions over several ReRAM arrays and thus enables parallel computations with nonzero inputs. Zero-skipping data flow, assisted with customized input buffers design, enhances the computation parallelism and input data reuse. In evaluation, we compare RED against the existing ReRAM-based PIMs and CMOS-based counterpart with a variety of GAN and FCN models, each of which contains multiple deconvolution layers. The experimental results show that RED achieves a$4.0\times $–$56.16\times $speedup and a$1.05\times $–$18.17\times $energy efficiency improvement over previous related accelerator designs. Ziru Li, Bing Li 0017, Zichen Fan, Hai Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | RED: A ReRAM-based Deconvolution AcceleratorabstractDeconvolution has been widespread in neural networks. For example, it is essential for performing unsupervised learning in generative adversarial networks or constructing fully convolutional networks for semantic segmentation. Resistive RAM (ReRAM)-based processing-in-memory architecture has been widely explored in accelerating convolutional computation and demonstrates good performance. Performing deconvolution on existing ReRAM-based accelerator designs, however, suffers from long latency and high energy consumption because deconvolutional computation includes not only convolution but also extra add-on operations. To realize the more efficient execution for deconvolution, we analyze its computation requirement and propose a ReRAM-based accelerator design, namely, RED. More specific, RED integrates two orthogonal methods, the pixel-wise mapping scheme for reducing redundancy caused by zero-inserting operations and the zero-skipping data flow for increasing the computation parallelism and therefore improving performance. Experimental evaluations show that compared to the state-of- the-art ReRAM-based accelerator, RED can speed up operation 3.69~31.15× and reduce 8%~88.36% energy consumption. Zichen Fan, Ziru Li, Bing Li 0017, Yiran Chen 0001, Hai Li 0001 |
DATE | 2 |