EDBT 2026 Demo / reviewers in the wild / expert
Zhipeng Liao
dblp:333/3778
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI ModelsabstractStochastic computing (SC) offers hardware simplicity but suffers from low throughput, while high-throughput Digital Computing-in-Memory (DCIM) is bottlenecked by costly adder logic for matrix-vector multiplication (MVM). To address this trade-off, this paper introduces a digital stochastic CIM (DS-CIM) architecture that achieves both high accuracy and efficiency. We implement signed multiply-accumulation (MAC) in a compact, unsigned OR-based circuit by modifying the data representation. Throughput is enhanced by replicating this low-cost circuit 64 times with only a 1× area increase. Our core strategy, a shared Pseudo Random Number Generator (PRNG) with 2D partitioning, enables single-cycle mutually exclusive activation to eliminate OR-gate collisions. We also resolve the 1s saturation issue via stochastic process analysis and data remapping, significantly improving accuracy and resilience to input sparsity. Our high-accuracy DS-CIM1 variant achieves 94.45% accuracy for INT8 ResNet18 on CIFAR-10 with a root-mean-squared error (RMSE) of just 0.74%. Meanwhile, our high-efficiency DS-CIM2 variant attains an energy efficiency of 3566.1 TOPS/W and an area efficiency of 363.7 TOPS/mm2, while maintaining a low RMSE of 3.81%. The DS-CIM capability with larger models is further demonstrated through experiments with INT8 ResNet50 on ImageNet and the FP8 LLaMA-7B model. Kunming Shao, Jiangnan Yu, Zhipeng Liao, Yi Zou 0001, Kwang-Ting Cheng, Chi-Ying Tsui |
DATE | 4 |
| 2026 | A Heterogeneous Decision Spiking Transformer Accelerator with Locality-dependent KV Product Cache and Compute Pattern Reconfigurable Engine
Ziyang Shen, Zhipeng Liao, Sitan Shen, Chaoming Fang, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 2 |
| 2026 | Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-Fly Aligned-Mantissa Bitwidth PredictionabstractFP8 low-precision formats have gained significant adoption in transformer inference and training. However, existing digital compute-in-memory (DCIM) architectures face challenges in supporting variable FP8 aligned-mantissa bitwidths, as unified alignment strategies and fixed-precision multiply accumulate (MAC) units struggle to handle input data with diverse distributions. This work presents a flexible FP8 DCIM accelerator with three innovations: 1) a dynamic shift-aware bitwidth prediction (DSBP) with on-the-fly input prediction that adaptively adjusts weight (2/4/6/8b) and input ($2\sim 12$b) aligned-mantissa precision; 2) a FIFO-based input alignment unit (FIAU) replacing complex barrel shifters with pointer-based control; and 3) a precision-scalable INT MAC array achieving flexible weight precision with minimal overhead. Implemented in 28-nm CMOS with a$64~\times ~96$CIM array, the design achieves 20.4 TFLOPS/W for fixed E5M7, demonstrating$2.8\times $higher FP8 efficiency than previous work while supporting all FP8 formats. Results on Llama-7b show that the DSBP achieves higher efficiency than fixed bitwidth mode at the same accuracy level on both BoolQ and Winogrande datasets, with configurable parameters enabling flexible accuracy–efficiency tradeoffs. Kunming Shao, Zhipeng Liao, Xijie Huang, Kwang-Ting Cheng, Chi-Ying Tsui, Yi Zou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM ComputationabstractRetrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating external knowledge retrieval but faces challenges on edge devices due to high storage, energy, and latency demands. Computing-in-Memory (CIM) offers a promising solution by storing document embeddings in CIM macros and enabling in-situ parallel retrievals but is constrained by either low memory density or limited computational accuracy. To address these challenges, we present DIRC-RAG, a novel edge RAG acceleration architecture leveraging Digital In-ReRAM Computation (DIRC). DIRC integrates a high-density multi-level ReRAM subarray with an SRAM cell, utilizing SRAM and differential sensing for robust ReRAM readout and digital multiply-accumulate (MAC) operations. By storing all document embeddings within the CIM macro, DIRC achieves ultra-low-power, single-cycle data loading, substantially reducing both energy consumption and latency compared to off-chip DRAM. A query-stationary (QS) dataflow is supported for RAG tasks, minimizing on-chip data movement and reducing SRAM buffer requirements. We introduce error optimization for the DIRC ReRAM-SRAM cell by extracting the bit-wise spatial error distribution of the ReRAM subarray and applying targeted bit-wise data remapping. An error detection circuit is also implemented to enhance readout resilience against device-and circuit-level variations.Simulation results demonstrate that DIRC-RAG under TSMC 40nm process achieves an on-chip non-volatile memory density of 5.18Mb/mm2and a throughput of 131 TOPS. It delivers a 4MB retrieval latency of 5.6μs/query and an energy consumption of 0.956μJ/query, while maintaining the retrieval precision. Kunming Shao, Zhipeng Liao, Jiangnan Yu, Xijie Huang, Jingyu He, Fengshi Tian, Yi Zou 0001, Kwang-Ting Cheng, Chi-Ying Tsui |
ISLPED | 2 |
| 2024 | Designing Constant Modulus Approximate Binary Phase Waveforms for Multitarget Detection in MIMO Radar Using LSTM NetworksabstractMultiple-input multiple-output (MIMO) radar can improve target detection capability due to its increased spatial degrees of freedom compared with traditional phased array radar. However, in practical applications, perfectly orthogonal waveforms are virtually unattainable. Thus, the orthogonality among the transmitted waveforms becomes a critical factor for determining the operational state in orthogonal MIMO radar. In typical target detection scenarios characterized by complex environments, the received radar echoes are often superimposed by reflections of multiple distinct objects. Such echo signals inevitably impact the orthogonality performance of non-orthogonal waveforms, thereby affecting the detection performance of the MIMO radar. To address this, the article initially undertakes a rigorous analysis of echo signals for non-orthogonal waveforms under both single-target and multi-target conditions. Subsequently, based on the analytical findings, we propose an objective function for the design of constant modulus code division multiple access waveforms tailored for multi-target scenarios. This objective function is then employed as a loss function in an unsupervised training scheme using a recurrent neural network with a long short-term memory architecture. Furthermore, we utilize a nonlinear function to discretize the network output, making the output phase coding approximate to binary coding. This makes the resulting phase coding more practically applicable. Finally, normalization is performed on the output amplitude of each matched filter to further enhance waveform orthogonality. Simulation results are provided to demonstrate the performance of the proposed method. Zizhou Qiu, Keqing Duan, Zhipeng Liao, Jinjun He |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Range-Ambiguous Clutter Suppression for Space-Based Early Warning Radar Using Vertical FDA and Horizontal EPCabstractThe range ambiguous clutter in space-based early warning radar is more severe compared with that in airborne radar due to the faster movement velocity. Meanwhile, the clutter angle-Doppler characteristics vary with range owing to Earth’s rotation. As a result, the mainbeam clutter returned from different ambiguous range occupies most of the Doppler spectrum, which leads to the degradation of the traditional space-time adaptive processing (STAP) performance in canceling clutter. Though frequency diverse array (FDA) based on multiple-input multiple-output can provide the ability to identify the range ambiguous clutter, it cannot perform well for space-based early warning radar because of the finite available vertical elements. In this paper, a novel STAP method, which utilizes the element-pulse coding (EPC) and FDA technique to mitigate the range ambiguous clutter, is proposed. Firstly, EPC is used to pre-whiten the most range ambiguous clutter via coding and decoding in horizontal elements and coherent pulses. Then the interval in the vertical frequency of the residual range-ambiguous clutter is expanded by FDA, and thus the targets and clutter from the vertical mainlobe are easily extracted using a spatial filter with a few vertical elements. Finally, the extracted clutter can be effectively suppressed through the traditional STAP method. Simulation results are provided to demonstrate the performance of the proposed method. Zizhou Qiu, Zhipeng Liao, Jingwei Xu 0002, Keqing Duan |
IEEE Geosci. Remote. Sens. Lett. | 2 |