EDBT 2026 Demo / reviewers in the wild / expert
Yulong Qiu
dblp:143/0270
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0008-8765-3130ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Retention and Bit-Level Approximate STT-MRAM for High-Efficiency AI Applications
Yulong Qiu, Chao Wang 0094, Weimeng Zhao, Zhongzhen Tong, Zhaohao Wang |
ISCAS | 1 |
| 2026 | A Fully-Parallel Digital MRAM Computing-in-Memory Macro Featuring a High-Efficient Dynamic Adder Tree and Bit-Splitting MAC
Zhongzhen Tong, Jiye Yao, Shaohui Ma, Yulong Qiu, Zhaohao Wang, Amara Amara, Xiaoyang Lin |
ISCAS | 4 |
| 2026 | BaM-CIM: A High Throughput Booth Algorithm-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFETabstractAs artificial intelligence (AI) and computational models grow in scale, the demand for computational power and storage has significantly increased. The computing-in-memory (CIM) architecture addresses this challenge by performing computations directly within the memory array, reducing data transfer between the processor and memory. This paper introduces a Booth algorithm-based In-MRAM computing architecture (BaM-CIM) using a hybrid voltage-gated spin-orbit torque MTJ (VGSOT-MTJ) and gate-all-around carbon nanotube field-effect transistors (GAA-CNTFETs) for efficient multiply-and-accumulate (MAC) computing. The key contributions of BaM-CIM are as follows: 1) A Voltage divider reference (VDR) cell is proposed, which enables read operations using only a 2T1M cell structure. Compared to complementary read cells, the VDR reduces the area by half and achieves robust data sensing without requiring precharge/discharge operations. 2) The BaM-CIM circuit is proposed to complete 8b-W/8b-IN/21b-OUT computations in only two cycles (1.6 ns), reducing the number of cycles by 75% compared to single-bit input serial operations and by 50% compared to two-bit serial operations. 3) A three-input 8b Booth computing adder (BCA), along with Modified computing shift adder (MCSA) and Modified computing post adder (MCPA), which can achieve higher energy efficiency. BaM-CIM with 128 Kb is simulated, achieving throughput and energy efficiency of 0.93 TOPS and 258.4 TOPS/W, respectively, at a 0.6 V supply voltage and 1.28 TOPS and 169.5 TOPS/W, respectively, at a 0.8 V supply voltage with 8b-IN, 8b-W, and 21b-OUT. Chenghang Li, Zhongzhen Tong, Yulong Qiu, Jiye Yao, Chao Wang 0094, Zhaohao Wang, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | ASP: A daptive Sparse LUT-based Bit-Slice Accelerator for Efficient CNN InferenceabstractRecent advances in convolutional neural network (CNN) accelerators have leveraged data sparsity, significantly enhancing energy efficiency in resource- and energy-constrained hardware platforms. However, fully exploiting data sparsity and further enhancing energy efficiency require the deployment of specialized hardware accelerators. In this work, we present a bit-slice accelerator that integrates an LUT-based multiplier and a bit manipulation unit to achieve efficient resource utilization and low-latency multi-bit-width operations. Furthermore, the bit-slice processing element (PE) and dynamic PE array utilize adaptive subdivision for irregular matrices, optimizing data handling across varying levels of sparsity. Our synthesized RTL implementation demonstrates a significant improvement in hardware utilization, achieving a throughput of 66.23 GOPS/W on AlexNet and 61.13 GOPS/W on VGG16. This design achieves 2.05x higher energy efficiency and 2.75x lower latency compared to state-of-the-art designs, outperforming them on both the AlexNet and VGG16 benchmarks. Yulong Qiu, Amara Amara |
ISCAS | 2 |
| 2025 | Approximate SOT-MRAM for Neural Network Acceleration with Superior Read PerformanceabstractMagnetoresistive random-access memory (MRAM) has been demonstrated to be a suitable memory technology for neural network (NN) acceleration due to its non-volatility, high density, and fast access speed. However, compared to the widely used static random-access memory (SRAM), MRAM still exhibits a notable disparity in speed and energy. In this paper, we propose a read-related approximation computation (RAC) strategy based on spin-orbit torque MRAM (SOT-MRAM) to enhance the computational speed of NN, and then we introduce a reference reconfigurable array (RRA) architecture to further decrease the read latency and energy consumption, significantly improving the speed and energy efficiency of weight retrieval during computations. Furthermore, we propose an algorithm to verify and optimize NN model performance. The proposed architecture is evaluated using a 28 nm process combined with a SPICE model of the SOT-MRAM. Simulation results indicate that the read speed increases by 4.71X, the read energy consumption is reduced by 70.7%, while the model accuracy loss remains below 1%. Yulong Qiu, Chao Wang 0094, Zhongzhen Tong, Siyuan Cheng 0021, Zhaohao Wang |
ISCAS | 1 |
| 2025 | A Self-Decryption Pass Transistor Logic-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFETabstractSpintronic devices and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFETs)-based computing in-memory architecture are competitive candidates for applications in battery-powered tiny artificial intelligence (AI) edge devices. Meanwhile, data encryption and decryption are also necessary to protect AI model weights and the customized data used to guarantee neural network (NN) inference accuracy. In this study, we propose a self-decryption pass transistor logic (PTL)-based in-MRAM computing macro (SP-CIM) that utilizes hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ)/GAA-CNTFET. The proposed SP-CIM macro enables simultaneous data access, decryption, and full-accuracy multiply-and-accumulate (MAC) operations using the newly introduced voltage-divider self-decryption cell, without the need for additional decryption logic. Compared to existing in-memory decryption strategies, this design reduces energy consumption by 45.7% and decreases decryption delay by 87.2%. To enhance area efficiency and reduce computing latency, we propose a PTL-based multiplication cell that achieves full-accuracy local 2b-IN TEXPRESERVE0 2b-W operations with only 20 transistors (20T). Additionally, novel PTL-based full-swing output half adders (10T-HA) and full adders (14T-FA) are proposed to construct the local adder tree, achieving reductions of 31.8%, 76.4%, and 41.4% in energy, delay, and area, respectively, compared to conventional adder trees in CIM macros. Simulations of the 288 kb SP-CIM macro demonstrated throughput and energy efficiency of 2.25 TOPS and 226.6 TOPS/W, respectively, at a 0.6 V supply voltage, and 2.97 TOPS and 154.1 TOPS/W, respectively, at a 0.8 V supply voltage, with 8b-IN, 8b-W, and 24b-OUT. Zhongzhen Tong, Sifan Sun, Chenghang Li, Jiye Yao, Yulong Qiu, Chao Wang 0094, Zhaohao Wang, Amara Amara, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |