Zhou Wang 0005

dblp:28/2295-5 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0002-0203-6854ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 6 first-author · 8 since 2021
YearPublicationVenuePosition
2026 MPE: A Power-Efficient Edge-Device Mamba Processor with Multi-Dimensional Calculation-Compression Scheme
Zhou Wang 0005, Haochen Du, Jiuren Zhou, Xiguang Wu, Qiankun Li 0004, Yanqing Xu 0003, Hanqi Feng, Xiaonan Tang, Shushan Qiao, Yongke Wang, Anil A. Bharath, Emm Mic Drakakis
ISCAS1
2026 GTPE: A 28nm 33.12 TFLOPS/W GNN Training Processor with Unstructured Multi Threshold Pruning, Hybrid Multi-mode Approximate Computing and QUIRE Number System Support
Zhou Wang 0005, Haochen Du, Jiuren Zhou, Xiguang Wu, Qiankun Li 0004, Yanqing Xu 0003, Hanqi Feng, Xiaonan Tang, Shushan Qiao, Tian-Chun Ye 0001, Anil A. Bharath, Emm Mic Drakakis
ISCAS1
2026 GATPE: A High-Performance Edge-Device GAT Processor with Multi-Layer Data-Variation Mechanism
Zhou Wang 0005, Haochen Du, Jiuren Zhou, Xiguang Wu, Qiankun Li 0004, Yanqing Xu 0003, Hanqi Feng, Xiaonan Tang, Shushan Qiao, Anil A. Bharath, Emm Mic Drakakis
ISCAS1
2026 SPICA: Energy-Efficient SRAM-Based Multi-Precision Multi-Mode Compute-in-Memory Accelerator for AI Inference
Xiguang Wu, Zhou Wang 0005, Jiuren Zhou, Genquan Han
ISCAS6
2025 STPE: An Energy-Efficient Edge-Device Transformer Inference Processor with Multi-Mode Data-Compression Scheme
abstract
Transformer-Based models have turned out to be very successful in many artificial intelligence (AI) tasks, outperforming traditional convolutional neural networks (CNNs), especially in the field of Natural Language Processing (NLP). Their success relies upon a self-attention mechanism which, when compared to CNNs, has a global rather than a local receptive domain. This article proposes an energy-efficient edge-device Transformer inference processor termed Smart Transformer Processing Element (STPE). Firstly, STPE sets up a Multi-Mode Indexing and Sparsity Scheme (MISS) for token association, and further reduces the computational load through in-situ computation; secondly, STPE exploits the Local Properties of Attention Mechanism (LPAM) to further reduce redundant and repetitive calculations in Transformer operations by means of a search band calculation and error correction mechanism; thirdly, STPE has designed a Quantization and Compression Parallel Method (QCPM) to improve the computing speed and hardware utilization under weak related (WR) token. Employing 28nm CMOS synthesis tools, the area of the proposed STPE processor is 7.33 mm2. Its peak energy efficiency is 84.15TOPS/W, which is 14.7 times higher than that of the H100 graphics processing unit (GPU) and 3.06 times higher than that of the most advanced Transformer processor.
Zhou Wang 0005, Haochen Du, Vivek Mohan, Jiuren Zhou, Yanqing Xu 0003, Baoyi Han, Xiaonan Tang, Shushan Qiao, Shouyi Yin, Anil A. Bharath, Emmanuel M. Drakakis
ISCAS1
2025 Architectural Exploration of Hybrid Neural Decoders for Neuromorphic Implantable BMI
abstract
This work presents an efficient decoding pipeline for neuromorphic implantable brain-machine interfaces (Neu-iBMI), leveraging sparse neural event data from an event-based neural sensing scheme. We introduce a tunable event filter (EvFilter), which also functions as a spike detector (EvFilter-SPD), significantly reducing the number of events processed for decoding by 192× and 554×, respectively. The proposed pipeline achieves high decoding performance, up to R2= 0.73, with ANN- and SNN-based decoders, eliminating the need for signal recovery, spike detection, or sorting, commonly performed in conventional iBMI systems. The SNN-Decoder reduces computations and memory required by 5 − 23× compared to NN-, and LSTM-Decoders, while the ST-NN-Decoder delivers similar performance to an LSTM-Decoder requiring 2.5× fewer resources. This streamlined approach significantly reduces computational and memory demands, making it ideal for low-power, on-implant, or wearable iBMIs.
Vivek Mohan, Biyan Zhou, Zhou Wang 0005, Anil A. Bharath, Emmanuel M. Drakakis, Arindam Basu
ISCAS3
2025 GPE: A High-Performance Edge GNN Inference Processor with Multi-Parallelism Format-Variation Mechanism
abstract
Recently, Graph Neural Networks (GNNs) have shown great potential in terms of accuracy for problems that are well-described by graph representations, such as problems of path planning. However, implementing GNNs on mobile platforms is challenging as it requires a significant amount of computation and large memory. This article proposes a High-Performance Edge GNN Inference Processor termed GPE (GNN Processing Element). Firstly, GPE sets up Multi-Dimensional Indexing and Dynamic Pruning Schemes (MIDPS) for GNN networks, and achieves cross layer interconnection of multiple neighboring nodes via NOC (Network on Chip); secondly, GPE utilizes Graph Structure Adjacency Table Information (GSATI) of a GNN to further reduce redundant and repetitive calculations by means of repeated matching and difference transfer mechanisms; thirdly, GPE has a graph-based Multi Parallelism Simplification and Operation Method (MPSOM) to improve computing speed and hardware utilization under small data volumes. Using 28nm CMOS synthesis tools, the area of the proposed GPE processor is 5.37 square millimeters. Its peak energy efficiency is 21.5TOPS/W, which is 3.76 times higher than that of the H100 GPU (Graphics Processing Unit), while the energy consumption of GNN is 80.9% lower than the previous SOTA (State of Art) work.
Zhou Wang 0005, Haochen Du, Jiuren Zhou, Yanqing Xu 0003, Vivek Mohan, Baoyi Han, Xiaonan Tang, Shushan Qiao, Shouyi Yin, Anil A. Bharath, Emmanuel M. Drakakis
ISCAS1
2023 CPE: An Energy-Efficient Edge-Device Training with Multi-dimensional Compression Mechanism
abstract
Recently, the edge-device DNN training has become of high importance, while the computation and access energy consumption of are too large. This paper proposes a CPE (Compress Process Element) with three characteristics. Firstly, CPE has a method of Reordering and Reusing Data (RRD) by controlling the output to reorder data. Secondly, CPE owns a Multi-directional Redundant Skip (MRS) mechanism, which anticipates all zeros and duplicate fields in advance. Thirdly, CPE contains a scheme to transform The Calculation Format (TCF), which transforms the input into another form. Evaluated with 28nm CMOS process, using CPE achieves 2.02 × energy reduction and offer 1.73 × speed up outperforming state-of-the-art trainable processor GANPU.
Zhou Wang 0005, Jingchuan Wei, Boxiao Han, Hongjun He, Leibo Liu, Shaojun Wei, Shouyi Yin
DAC1