Yufeng Xie 0001

dblp:38/6002-1 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-6541-2925ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 8 since 2021
YearPublicationVenuePosition
2026 BLCIM: An Efficient Radix-16 Booth LUT-Based SRAM-CIM Architecture with Algorithm-Hardware Co-Optimization for NTT
abstract
Lattice-based cryptography relies heavily on the Number Theoretic Transform (NTT), whose performance is dominated by modular multiplication and data movement. This paper proposes BLCIM, the first Radix-16 Booth LUT-based compute-in-memory (CIM) NTT accelerator. We propose an algorithm that precomputes partial modular multiplications with fixed rotation factors and uses input Booth-encoded search results. Then, based on the Booth encoding, it decides whether to perform shifting and inversion to simplify modular multiplication. In addition to the proposed sparsity-aware and stage-skipping schemes, the Radix-16 Booth LUT-based algorithm significantly improves NTT performance at low energy cost. By implementing the algorithm on the SRAM-CIM architecture with a lightweight pipeline design, we achieved nearly 100% utilization for the NTT circuit during acceleration. Simulated in 28 nm CMOS technology, the proposed BLCIM achieves only 4.53% latency and 58.87% energy consumption of the latest work.
Qianhua Li, Hongrui Meng, Chunshan Wang, Shengchao Zhou, Teng Zou, Yufeng Xie 0001
ACM Great Lakes Symposium on VLSI8
2026 A 0.562 mm2 MTJ-Based Ising Machine for 48-bit Integer Factorization in 40 nm CMOS
Kunpeng Gao, Jiadong Chen, Shaohao Wang, Tai Min, Yufeng Xie 0001
ISCAS6
2026 Horizontal-Parallel ADC-less Sparsity-Clock-Aware RRAM CIM Macro for edge AI devices
Teng Zou, Shengchao Zhou, Hongrui Meng, Yufeng Xie 0001
ISCAS4
2026 A 40-nm Training-Inference STT-MRAM Near-Memory Computing Macro for Memory-Augmented Neural Network Acceleration
abstract
Recently, memory-augmented neural networks (MANNs) have gained significant attention as a critical solution for few-shot learning (FSL). These networks leverage external memory to store prior knowledge, thereby enhancing classification efficiency. Spin-transfer torque magnetic random access memory (STT-MRAM) is particularly suited for this application due to its compact cell size, excellent data retention, and scalability. In this article, we introduce a STT-MRAM-based near-memory computing (NMC) macro specifically designed for MANNs. Our approach incorporates several key innovations aimed at overcoming challenges in hardware implementation while improving MANN performance as follows: 1) a parallel computing architecture within the NMC to expedite$L1$distance computations; 2) a memory invert coding (MIC) and self-termination write (STW) scheme that reduce write operations and energy consumption, addressing the issues of frequent writes and high write currents during the training phase of MANNs; 3) a dynamic offset-compensation sense amplifier (DOC-SA) and high-throughput switch-capacitor (HTSC) readout scheme to improve read accuracy and throughput, tackling low read margins and limited readout bandwidth; 4) an exploration of MANN architectures validates the reusability of the NMC macro. The optimized matching-networks (MCHnets)-based structure achieves an accuracy exceeding 90% in five-way and eight-way Omniglot classification tasks. Fabricated with a 40-nm CMOS technology, our design achieves classification accuracies of 96.37% for eight-way-five-shot tasks and 93.72% for 16-way-five-shot tasks on the Omniglot dataset utilizing the optimized MCHnet, showcasing an impressive energy efficiency of 6.47 TOPS/W at the basis of 16-bit$L1$distance computing in the classification tasks of MANN.
Shengchao Zhou, Hongrui Meng, Yajun Wu, Zizhao Ma, Teng Zou, Tai Min, Shaohao Wang, Yufeng Xie 0001
IEEE Trans. Very Large Scale Integr. Syst.9
2025 A 40nm STT-MRAM Near-Memory Computing Macro for Memory-Augmented Neural Network Acceleration
abstract
Memory-augmented neural network (MANN) has gained attention as a pivotal solution for few-shot learning (FSL). Among the candidates for associative memory in MANN accelerators, spin-transfer torque magnetic random-access memory (STT-MRAM) stands out for its compact cell area, long data retention time, and excellent scalability. In this paper, we propose an STT-MRAM near-memory computing (NMC) macro for MANN acceleration. The macro contains following innovations: 1) An array-level parallel computing architecture for L1 distance calculation. 2) A low-area-overhead memory-invert coding technique to reduce write energy consumption. 3) A configurable dynamic offset-compensation sense amplifier (CDOC-SA) to improve classification accuracy. Fabricated in 40nm CMOS process, our macro demonstrates an energy efficiency of 6.47 TOPS/W, achieving the classification accuracy of 98.3% and 93% for 8-way-5-shot tasks and 16-way-5-shot tasks on the Omniglot dataset.
Hongrui Meng, Yajun Wu, Shengchao Zhou, Zizhao Ma, Tai Min, Shaohao Wang, Yufeng Xie 0001
ISCAS7
2025 High sensing margin and parallelism 6T-2MTJ SOT-MRAM based TCAM for energy-Efficient Similarity Priority calculation in MANNs
abstract
With the development of AI applications, there is a demand for high parallel similarity priority calculation, like in MANNs. TCAM, as a type of memory for high parallel searching work, is suitable to perform the task. However, the CMOS based TCAMs suffer from area overhead, static power consumption and data-loss while power failure. The SOT-MRAM presents a promising alternative for TCAM design due to its nonvolatility, low area overhead and no static power consumption. But, its low on/off ratio results in low sensing margin which limits its processing speed. To address the challenges, we propose a novel TCAM structure with high sensing margin, parallelism and low energy consumption. This work proposes the followings: (1) A novel 6T-2MTJ sot-mram based TCAM structure is proposed. Simulations show it has a 2-3x improvement in sensing margin over other MRAM based ones, 1.37x improvement in searching energy (fJ/bit) over other nonvolatile technologies based ones, and 23% area reduction over other CMOS based ones. (2) A segmented power supply scheme is presented to meet the parallelism need of MANNs which improves the parallelism by 4-8x compared to no segmented one.
Shengchao Zhou, Teng Zou, Zeming Wang, Xianwu Hu, Hongrui Meng, Yajun Wu, Chuxin Zhang, Caihua Wan, Yufeng Xie 0001
ISCAS9
2023 A 40nm 150 TOPS/W High Row-Parallel MRAM Compute-in-Memory Macro with Series 3T1MTJ Bitcell for MAC Operation
abstract
Non-volatile Compute-in-Memory (CIM), especially high-speed MRAM CIM, promises to be a solution of “Memory Wall” problem in power-sensitive artificial intelligence edge devices. However, the low resistance and low on/off ratio limit the row parallelism and efficiency of MRAM CIM macros. To overcome these challenges, this work proposes the following: 1) a series 3T1MTJ bit-cell CIM architecture; 2) an input-aware and self-generated dynamic reference array; 3) a high-speed readout pipeline circuit. The proposed macro eliminates errors of high Row-Parallel multiply-and-accumulate (MAC) operation with 150 TOPS/W peak energy efficiency simulated using 40nm process and STT-MTJ.
Zizhao Ma, Xianwu Hu, Gan Wen, Xiaoyang Zeng, Yufeng Xie 0001
ISCAS6
2022 A High Area-Efficiency RRAM-Based Strong PUF with Multi-Entropy Source and Configurable Double-Read Process
abstract
Physically Unclonable Functions (PUFs) are emerging security primitives for authentication due to its high physical security. Especially for those with excellent area-efficiency and reliable immunity against attacks, the demand is larger. In order to achieve higher security and area-efficiency, this paper proposes a strong PUF based on resistive random-access memory (RRAM). We exploit both the switching randomness and intrinsic resistance distribution of RRAM to increase the randomness of entropy source, and design a novel strong PUF structure with double-read process to increase the challenge-response pairs (CRPs) and area-efficiency. A double XOR process is proposed to enhance the immunity against machine learning attack (MLA) with low area-overhead. Compared with the state of the art, the number of CRP has been greatly improved, demonstrating a better area-utilization. Simulation results show that the CRP generation time is 1. 8us, mean intra-HD of1.93%, inter-HD of 49.95% and uniformity of 49.02%. The above features of the proposed strong PUF make it a promising candidate for Internet of Things (IoTs) authentication applications.
Xianwu Hu, Jiayun Feng, Zizhao Ma, Xiaoyang Zeng, Yufeng Xie 0001
ISCAS6