Yunzhengmao Wang

dblp:295/3505 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2024
0009-0002-4125-4181ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2024 A 6.4-Gbps 0.41-pJ/b fully-digital die-to-die interconnect PHY for silicon interposer based 2.5D integration
Yinglin Yang, Yunzhengmao Wang, Tengyue Yi, Chixiao Chen, Qi Liu 0010
Integr.2
2024 HARDSEA: Hybrid Analog-ReRAM Clustering and Digital-SRAM In-Memory Computing Accelerator for Dynamic Sparse Self-Attention in Transformer
abstract
Self-attention-based transformers have outperformed recurrent and convolutional neural networks (RNN/ CNNs) in many applications. Despite the effectiveness, calculating self-attention is prohibitively costly due to quadratic computation and memory requirements. To solve this challenge, this article proposes a hybrid analog-ReRAM and digital-SRAM in-memory computing accelerator (HARDSEA), a computing-in-memory (CIM) accelerator supporting self-attention in transformer applications. To trade off between energy efficiency and algorithm accuracy, HARDSEA features an algorithm-architecture-circuit codesign. A product-quantization-based scheme dynamically facilitates self-attention sparsity by predicting lightweight token relevance. A hybrid in-memory computing architecture employs both high-efficiency analog ReRAM-CIM and high-precision digital SRAM-CIM to implement the proposed new scheme. The ReRAM-CIM, whose precision is sensitive to circuit nonidealities, takes charge of token relevance prediction where only computing monotonicity is demanded. The SRAM-CIM, utilized for exact sparse attention computing, is reorganized as an on-memory-boundary computing scheme, thus adapting to irregular sparsity patterns. In addition, we propose a time-domain winner-take-all (WTA) circuit to replace the expensive ADCs in ReRAM-CIM macros. Experimental results show that HARDSEA prunes BERT and GPT-2 models to 12%–33% sparsity without accuracy loss, achieving$13.5\times $–$28.5\times $speedup and$291.6\times $–$1894.3\times $energy efficiency over GPU. Compared to state-of-the-art transformer accelerators, HARDSEA has$1.2\times $–$14.9\times $better energy efficiency at the same level of throughput.
Shiwei Liu 0002, Chen Mu, Hao Jiang 0024, Yunzhengmao Wang, Jinshan Zhang 0006, Keji Zhou, Qi Liu 0010, Chixiao Chen
IEEE Trans. Very Large Scale Integr. Syst.4
2023 A $2.53 \mu \mathrm{W}/\text{channel}$ Event-Driven Neural Spike Sorting Processor with Sparsity-Aware Computing-In-Memory Macros
abstract
Spike sorting processors with high energy efficiency are widely used in large-scale neural signal processing tasks to monitor the activity of neurons in brains. This paper presents a low-power processor for high-accuracy spike sorting and on-chip incremental learning using an algorithm-hardware co-design approach. The processor introduces an event-driven mechanism with adaptive-threshold detection to conditionally activate the system in order to reduce power consumption. Sparsity-aware computing-in-memory (CIM) macros are also developed in our design to store templates and perform complicated computations efficiently. The prototype is designed using 28nm technology with an area of 0.018 mm2/channel and an overall power efficiency of$\mathbf{2.53} \mu \mathbf{W}/\mathbf{channel}$and 84nW/(channel.cluster) at the voltage of 0.72V. Moreover, the accuracy of the whole design can reach 94.5% in a 32-channel scenario.
Hao Jiang 0024, Jiapei Zheng, Yunzhengmao Wang, Jinshan Zhang 0006, Haozhe Zhu, Liangjian Lyu, Yingping Chen, Chixiao Chen, Qi Liu 0010
ISCAS3
2021 ALPINE: An Agile Processing-in-Memory Macro Compilation Framework
abstract
Processing-in-Memory architectures and circuit designs are playing significant roles in the recent energy-efficient machine learning chips. This paper proposes a PIM macro compilation framework called ALPINE to speed up previously tedious and error-prone PIM design flow, paving the way towards open-source and process-portable PIM chips. Relying on an extensible PIM standard cell library, ALPINE can generate the corresponding topology according to the specification, and process placement and routing. The proposed PIM macro is compatible with different storage devices such as SRAM and RRAM, and can support various quantization bit-widths and dataflows. To verify the effectiveness, a 128×128 SRAM-based PIM macro instance is implemented, and the simulation results show that it can achieve an energy efficiency of 19.05TOPS/W under 65nm CMOS technology. The macro performance is not inferior to the state-of-the-art custom PIM designs.
Jinshan Zhang 0006, Bo Jiao 0003, Yunzhengmao Wang, Haozhe Zhu, Lihua Zhang 0002, Chixiao Chen
ACM Great Lakes Symposium on VLSI3