VLDB 2026 Research / reviewers in the wild / expert
Jinshan Zhang 0006
dblp:284/7963
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2024
0000-0002-7486-9898ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | HARDSEA: Hybrid Analog-ReRAM Clustering and Digital-SRAM In-Memory Computing Accelerator for Dynamic Sparse Self-Attention in TransformerabstractSelf-attention-based transformers have outperformed recurrent and convolutional neural networks (RNN/ CNNs) in many applications. Despite the effectiveness, calculating self-attention is prohibitively costly due to quadratic computation and memory requirements. To solve this challenge, this article proposes a hybrid analog-ReRAM and digital-SRAM in-memory computing accelerator (HARDSEA), a computing-in-memory (CIM) accelerator supporting self-attention in transformer applications. To trade off between energy efficiency and algorithm accuracy, HARDSEA features an algorithm-architecture-circuit codesign. A product-quantization-based scheme dynamically facilitates self-attention sparsity by predicting lightweight token relevance. A hybrid in-memory computing architecture employs both high-efficiency analog ReRAM-CIM and high-precision digital SRAM-CIM to implement the proposed new scheme. The ReRAM-CIM, whose precision is sensitive to circuit nonidealities, takes charge of token relevance prediction where only computing monotonicity is demanded. The SRAM-CIM, utilized for exact sparse attention computing, is reorganized as an on-memory-boundary computing scheme, thus adapting to irregular sparsity patterns. In addition, we propose a time-domain winner-take-all (WTA) circuit to replace the expensive ADCs in ReRAM-CIM macros. Experimental results show that HARDSEA prunes BERT and GPT-2 models to 12%–33% sparsity without accuracy loss, achieving$13.5\times $–$28.5\times $speedup and$291.6\times $–$1894.3\times $energy efficiency over GPU. Compared to state-of-the-art transformer accelerators, HARDSEA has$1.2\times $–$14.9\times $better energy efficiency at the same level of throughput. Shiwei Liu 0002, Chen Mu, Hao Jiang 0024, Yunzhengmao Wang, Jinshan Zhang 0006, Keji Zhou, Qi Liu 0010, Chixiao Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | A $2.53 \mu \mathrm{W}/\text{channel}$ Event-Driven Neural Spike Sorting Processor with Sparsity-Aware Computing-In-Memory MacrosabstractSpike sorting processors with high energy efficiency are widely used in large-scale neural signal processing tasks to monitor the activity of neurons in brains. This paper presents a low-power processor for high-accuracy spike sorting and on-chip incremental learning using an algorithm-hardware co-design approach. The processor introduces an event-driven mechanism with adaptive-threshold detection to conditionally activate the system in order to reduce power consumption. Sparsity-aware computing-in-memory (CIM) macros are also developed in our design to store templates and perform complicated computations efficiently. The prototype is designed using 28nm technology with an area of 0.018 mm2/channel and an overall power efficiency of$\mathbf{2.53} \mu \mathbf{W}/\mathbf{channel}$and 84nW/(channel.cluster) at the voltage of 0.72V. Moreover, the accuracy of the whole design can reach 94.5% in a 32-channel scenario. Hao Jiang 0024, Jiapei Zheng, Yunzhengmao Wang, Jinshan Zhang 0006, Haozhe Zhu, Liangjian Lyu, Yingping Chen, Chixiao Chen, Qi Liu 0010 |
ISCAS | 4 |
| 2023 | A Scalable Die-to-Die Interconnect with Replay and Repair Schemes for 2.5D/3D IntegrationabstractChiplet is a critical technology in the post-Moore era, and the die-to-die (D2D) interconnect is essential for communication between chiplets. Meanwhile, several edge-computing devices based on 2.5D/3D chiplet have recently emerged. However, a lightweight D2D interconnect for 2.5D/3D edge-computing systems is lacking. Given the differences between 2.5D/3D integration, a scalable D2D interconnect with replay and repair schemes is presented in this paper. A credit-based flow control scheme and a custom replay scheme are presented for high efficiency. An effective detection and repair scheme is proposed to enhance fault tolerance for the D2D interconnect. Compared with a previous D2D interconnect design, the proposed D2D interconnect delivers 1.07/1.09Gbps throughput ($\sim 2.4\times/\sim 3.9\times \text{for}\ \text{write}/\text{read}$) and significantly reduced energy/bit with only ∼1.7× increased hardware cost. Additionally, compared with a previous chip-to-chip interconnect design, the proposed D2D interconnect can be configured down to power consumption as low as 0.55pJ/bit and 38.40Gbps throughput, achieving ∼2.5× throughput and significantly reduced latency with a negligible increase in hardware cost. Bo Jiao 0003, Jinshan Zhang 0006, Shiwei Liu 0002, Hao Jiang 0024, Jun Tao 0001, Wenning Jiang, Qi Liu 0010, Lihua Zhang 0002, Haozhe Zhu, Chixiao Chen |
ISCAS | 3 |
| 2022 | A 11.6μ W Computing-on-Memory-Boundary Keyword Spotting Processor with Joint MFCC-CNN Ternary QuantizationabstractThis paper presents an ultra-low-power keyword spotting processor using an algorithm-architecture co-design approach. Joint MFCC-CNN ternary weight quantization is proposed to reduce power consumption. The Mel filter and the DCT module are merged into one matrix multiplication. The merged coefficients and the weights of the rest NN classifier are ternary-quantized, causing less than 3% accuracy loss but 39 × energy efficiency improvement. Moreover, a Computing-on-Memory-Boundary macro is adopted to store the quantized coefficients and weights, and perform matrix multiplications. Compared to the existing computing-in-memory technology, the proposed technique can reduce power consumption due to the higher utilization ratio. To verify the proposed techniques, a keyword spotting processor prototype is designed with 28nm CMOS technology. Simulation results show that the prototype achieves power consumption of 11.6μ W under a power supply of 0.72V and a clock frequency of 250KHz. Xinru Jia, Haozhe Zhu, Yunzheng Wang, Jinshan Zhang 0006, Xiankui Xiong, Dong Xu 0015, Chixiao Chen, Qi Liu 0010 |
ISCAS | 4 |
| 2021 | A 0.57-GOPS/DSP Object Detection PIM Accelerator on FPGAabstractThe paper presents an object detection accelerator featuring a processing-in-memory (PIM) architecture on FPGAs. PIM architectures are well known for their energy efficiency and avoidance of the memory wall. In the accelerator, a PIM unit is developed using BRAM and LUT based counters, which also helps to improve the DSP performance density. The overall architecture consists of 64 PIM units and three memory buffers to store inter-layer results. A shrunk and quantized Tiny-YOLO network is mapped to the PIM accelerator, where DRAM access is fully eliminated during inference. The design achieves a throughput of 201.6 GOPs at 100MHz clock rate and correspondingly, a performance density of 0.57 GOPS/DSP. Bo Jiao 0003, Jinshan Zhang 0006, Yuanyuan Xie, Shunli Wang 0001, Haozhe Zhu, Xiaoyang Kang 0001, Zhiyan Dong, Lihua Zhang 0002, Chixiao Chen |
ASP-DAC | 2 |
| 2021 | Computing Utilization Enhancement for Chiplet-based Homogeneous Processing-in-Memory Deep Learning ProcessorsabstractThis paper presents a design strategy of chiplet-based processing-in-memory systems for deep neural network applications. Monolithic silicon chips are area and power limited, failing to catch the recent rapid growth of deep learning algorithms. The paper first demonstrates a straightforward layer-wise method that partitions the workload of a monolithic accelerator to a multi-chiplet pipeline. A quantitative analysis shows that the straightforward separation degrades the overall utilization of computing resources due to the reduced on-chiplet memory size, thus introducing a higher memory wall. A tile interleaving strategy is proposed to overcome such degradation. This strategy can segment one layer to different chiplets which maximizes the computing utilization. To facilitate the strategy, the modification of the chiplet system hardware is also discussed. To validate the proposed strategy, a nine-chiplet processing-in-memory system is evaluated with a custom-designed object detection network. Each chiplet can achieve a peak performance of 204.8GOPS at a 100-MHz rate. The peak performance of the overall system is 1.711TOPS, where no off-chip memory access is needed. By the tile interleaving strategy, the utilization is improved from 53.9 to 92.8 Bo Jiao 0003, Haozhe Zhu, Jinshan Zhang 0006, Shunli Wang 0001, Xiaoyang Kang 0001, Lihua Zhang 0002, Mingyu Wang 0001, Chixiao Chen |
ACM Great Lakes Symposium on VLSI | 3 |
| 2021 | ALPINE: An Agile Processing-in-Memory Macro Compilation FrameworkabstractProcessing-in-Memory architectures and circuit designs are playing significant roles in the recent energy-efficient machine learning chips. This paper proposes a PIM macro compilation framework called ALPINE to speed up previously tedious and error-prone PIM design flow, paving the way towards open-source and process-portable PIM chips. Relying on an extensible PIM standard cell library, ALPINE can generate the corresponding topology according to the specification, and process placement and routing. The proposed PIM macro is compatible with different storage devices such as SRAM and RRAM, and can support various quantization bit-widths and dataflows. To verify the effectiveness, a 128×128 SRAM-based PIM macro instance is implemented, and the simulation results show that it can achieve an energy efficiency of 19.05TOPS/W under 65nm CMOS technology. The macro performance is not inferior to the state-of-the-art custom PIM designs. Jinshan Zhang 0006, Bo Jiao 0003, Yunzhengmao Wang, Haozhe Zhu, Lihua Zhang 0002, Chixiao Chen |
ACM Great Lakes Symposium on VLSI | 1 |