Qirui Ren

dblp:230/2846 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2024
0000-0002-2155-7057ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2024 A 28-nm Computing-in-Memory-Based Super-Resolution Accelerator Incorporating Macro-Level Pipeline and Texture/Algebraic Sparsity
abstract
Super-resolution (SR) task using the convolutional neural network is a crucial task in improving image and video quality. The introduction of the residual block (RB) raises the depth of the algorithm to perform better reconstruction. The processing of the RB leads to a decrease in hardware utilization and frequent off-chip communications. It is hard to apply such algorithms on edge devices with limited performance. Computing-in-memory (CiM) is one promising method to reduce high power caused by massive data movement in multiply-accumulation computation. The algebraic sparsity (AS) is the structured sparsity (SS) optimization for imaging computing. However, it is an unsolved problem to simultaneously realize the texture sparsity (TS) of the image and the SS of the algorithm in the CiM scheme while maintaining high hardware utilization. Thus, we propose a CiM-based SR task accelerator. There are three key contributions: first, a texture-aware workflow and a dynamic grouping CiM engine can concurrently support TS coupling with AS. Second, a macro-level pipeline scheme together with two custom-sized CiM macros and a high reuse-rate Hadamard transformation circuit reaches 91% hardware utilization. Third, a novel weight update strategy is devised to reduce the performance loss induced by the weight updating. The accelerator prototype is fabricated in a 28-nm CMOS. It scores a 22.8-44.3-TOPS/W peak energy efficiency at the voltage supply of 0.54-1.1 V and the operating frequency of 50-200 MHz, indicating 1.8-6.8x higher compared to the state-of-the-art CiM processors.
Hao Wu 0084, Yong Chen 0005, Yiyang Yuan, Jinshan Yue, Xiangqu Fu, Qirui Ren, Pui-In Mak, Xinghua Wang 0005, Feng Zhang 0014
IEEE Trans. Circuits Syst. I Regul. Pap.6
2023 P3 ViT: A CIM-Based High-Utilization Architecture With Dynamic Pruning and Two-Way Ping-Pong Macro for Vision Transformer
abstract
Transformers have made remarkable contributions to natural language processing (NLP) and many other fields. Recently, transformer-based models have achieved state-of-the-art (SOTA) performance on computer vision tasks compared with traditional convolutional neural networks (CNNs). Unfortunately, existing CNN accelerators cannot efficiently support transformer due to the high computational overhead and redundant data accesses associated with the ‘KQV’ matrix operations in the transformer models. If the recently-developed NLP transformer accelerators are applied to the vision transformer (ViT) models, their efficiency would decrease due to three challenges. 1) Redundant data storage and access still exist in ViT data flow scheduling. 2) For matrix transposition in transformer models, the previous transpose-operation schemes lack flexibility, resulting in extra area overhead. 3) The sparse acceleration schemes for NLP in prior transformer accelerators cannot efficiently accelerate ViT with relatively fewer tokens. To overcome these challenges, we propose$P^{3}$ViT, a computing-in-memory (CIM)-based architecture, to efficiently accelerate ViT, achieving high utilization on data flow scheduling. There are three key contributions: 1) P3ViT architecture supports three ping-pong pipeline scheduling modes, involving inter-core parallel and intra-core ping-pong pipeline mode (IEP-IAP3), inter-core pipeline and parallel mode (IEP2), and full parallel mode, to eliminate redundant memory accesses. 2) A two-way ping-pong CIM macro is proposed, which can be configured to regular calculation mode and transpose calculation mode to adapt to both$\text{Q}\times \text{K}^{\mathrm {T}}$and$\text{A}\times \text{V}$tasks. 3) P3ViT also runs a small prediction network. It prunes redundant tokens to be a standard number hierarchically and dynamically, enabling high-throughput and high-utilization attention computation. Measurements show that P3ViT achieves$1.13\times $higher energy efficiency than the state-of-the-art transformer accelerator and achieves$30.8\times $and$14.6\times $speedup compared to CPU and GPU.
Xiangqu Fu, Qirui Ren, Hao Wu 0084, Feibin Xiang, Jinshan Yue, Yong Chen 0005, Feng Zhang 0014
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 A Security-Enhanced, Charge-Pump-Free, ISO14443-A-/ISO10373-6-Compliant RFID Tag With 16.2-μW Embedded RRAM and Reconfigurable Strong PUF
abstract
Radio frequency identification technology (RFID) has empowered a wide variety of automation industries, such as logistics and freight transportation. To further promote RFID tags adoption, security, power consumption, and cost have always been issues of general concern. This article presents the first synergy of the RFID tag with embedded resistive RAM (RRAM) array and RRAM-based reconfigurable strong physical unclonable function (R-SPUF). The RRAM not only meets the mass storage and technology downscaling but also renders the ultralow-cost “1-cent RFID tag” more feasible. Moreover, the R-SPUF facilitates multiple initializations until a satisfactory distribution and has strong secure keys benefiting from its reconfigurability that improves both safety and reliability. The complete system operates at 13.56 MHz and is compliant with the ISO14443-A and ISO10373-6 (test) protocols. The RFID tag was fabricated on a 1.1-mm2 die based on the 0.18-$\mu \text{m}$CMOS process. Without resorting to the charge pumps for RRAM read–write operations, the total power consumption is as low as 52.3$\mu \text{W}$, of which the RRAM dissipates$16.2~\mu \text{W}$under a wireless power supply.
Qirui Ren, Qiang Huo, Hao Wu 0084, Xiangqu Fu, Xiaoxin Xu, Jianfeng Gao 0005, Xiaojin Zhao, Dengyun Lei, Xinghua Wang 0005, Feng Zhang 0014, Yong Chen 0005, Pui-In Mak
IEEE Trans. Very Large Scale Integr. Syst.1