EDBT 2026 Demo / reviewers in the wild / expert
Yufeng Xie 0001
dblp:38/6002-1
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-6541-2925ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BLCIM: An Efficient Radix-16 Booth LUT-Based SRAM-CIM Architecture with Algorithm-Hardware Co-Optimization for NTTabstractLattice-based cryptography relies heavily on the Number Theoretic Transform (NTT), whose performance is dominated by modular multiplication and data movement. This paper proposes BLCIM, the first Radix-16 Booth LUT-based compute-in-memory (CIM) NTT accelerator. We propose an algorithm that precomputes partial modular multiplications with fixed rotation factors and uses input Booth-encoded search results. Then, based on the Booth encoding, it decides whether to perform shifting and inversion to simplify modular multiplication. In addition to the proposed sparsity-aware and stage-skipping schemes, the Radix-16 Booth LUT-based algorithm significantly improves NTT performance at low energy cost. By implementing the algorithm on the SRAM-CIM architecture with a lightweight pipeline design, we achieved nearly 100% utilization for the NTT circuit during acceleration. Simulated in 28 nm CMOS technology, the proposed BLCIM achieves only 4.53% latency and 58.87% energy consumption of the latest work. Qianhua Li, Hongrui Meng, Chunshan Wang, Shengchao Zhou, Teng Zou, Yufeng Xie 0001 |
ACM Great Lakes Symposium on VLSI | 8 |
| 2026 | A 0.562 mm2 MTJ-Based Ising Machine for 48-bit Integer Factorization in 40 nm CMOS
Kunpeng Gao, Jiadong Chen, Shaohao Wang, Tai Min, Yufeng Xie 0001 |
ISCAS | 6 |
| 2026 | Horizontal-Parallel ADC-less Sparsity-Clock-Aware RRAM CIM Macro for edge AI devices
Teng Zou, Shengchao Zhou, Hongrui Meng, Yufeng Xie 0001 |
ISCAS | 4 |
| 2026 | A 40-nm Training-Inference STT-MRAM Near-Memory Computing Macro for Memory-Augmented Neural Network AccelerationabstractRecently, memory-augmented neural networks (MANNs) have gained significant attention as a critical solution for few-shot learning (FSL). These networks leverage external memory to store prior knowledge, thereby enhancing classification efficiency. Spin-transfer torque magnetic random access memory (STT-MRAM) is particularly suited for this application due to its compact cell size, excellent data retention, and scalability. In this article, we introduce a STT-MRAM-based near-memory computing (NMC) macro specifically designed for MANNs. Our approach incorporates several key innovations aimed at overcoming challenges in hardware implementation while improving MANN performance as follows: 1) a parallel computing architecture within the NMC to expedite$L1$distance computations; 2) a memory invert coding (MIC) and self-termination write (STW) scheme that reduce write operations and energy consumption, addressing the issues of frequent writes and high write currents during the training phase of MANNs; 3) a dynamic offset-compensation sense amplifier (DOC-SA) and high-throughput switch-capacitor (HTSC) readout scheme to improve read accuracy and throughput, tackling low read margins and limited readout bandwidth; 4) an exploration of MANN architectures validates the reusability of the NMC macro. The optimized matching-networks (MCHnets)-based structure achieves an accuracy exceeding 90% in five-way and eight-way Omniglot classification tasks. Fabricated with a 40-nm CMOS technology, our design achieves classification accuracies of 96.37% for eight-way-five-shot tasks and 93.72% for 16-way-five-shot tasks on the Omniglot dataset utilizing the optimized MCHnet, showcasing an impressive energy efficiency of 6.47 TOPS/W at the basis of 16-bit$L1$distance computing in the classification tasks of MANN. Shengchao Zhou, Hongrui Meng, Yajun Wu, Zizhao Ma, Teng Zou, Tai Min, Shaohao Wang, Yufeng Xie 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2025 | A 40nm STT-MRAM Near-Memory Computing Macro for Memory-Augmented Neural Network AccelerationabstractMemory-augmented neural network (MANN) has gained attention as a pivotal solution for few-shot learning (FSL). Among the candidates for associative memory in MANN accelerators, spin-transfer torque magnetic random-access memory (STT-MRAM) stands out for its compact cell area, long data retention time, and excellent scalability. In this paper, we propose an STT-MRAM near-memory computing (NMC) macro for MANN acceleration. The macro contains following innovations: 1) An array-level parallel computing architecture for L1 distance calculation. 2) A low-area-overhead memory-invert coding technique to reduce write energy consumption. 3) A configurable dynamic offset-compensation sense amplifier (CDOC-SA) to improve classification accuracy. Fabricated in 40nm CMOS process, our macro demonstrates an energy efficiency of 6.47 TOPS/W, achieving the classification accuracy of 98.3% and 93% for 8-way-5-shot tasks and 16-way-5-shot tasks on the Omniglot dataset. Hongrui Meng, Yajun Wu, Shengchao Zhou, Zizhao Ma, Tai Min, Shaohao Wang, Yufeng Xie 0001 |
ISCAS | 7 |
| 2025 | High sensing margin and parallelism 6T-2MTJ SOT-MRAM based TCAM for energy-Efficient Similarity Priority calculation in MANNsabstractWith the development of AI applications, there is a demand for high parallel similarity priority calculation, like in MANNs. TCAM, as a type of memory for high parallel searching work, is suitable to perform the task. However, the CMOS based TCAMs suffer from area overhead, static power consumption and data-loss while power failure. The SOT-MRAM presents a promising alternative for TCAM design due to its nonvolatility, low area overhead and no static power consumption. But, its low on/off ratio results in low sensing margin which limits its processing speed. To address the challenges, we propose a novel TCAM structure with high sensing margin, parallelism and low energy consumption. This work proposes the followings: (1) A novel 6T-2MTJ sot-mram based TCAM structure is proposed. Simulations show it has a 2-3x improvement in sensing margin over other MRAM based ones, 1.37x improvement in searching energy (fJ/bit) over other nonvolatile technologies based ones, and 23% area reduction over other CMOS based ones. (2) A segmented power supply scheme is presented to meet the parallelism need of MANNs which improves the parallelism by 4-8x compared to no segmented one. Shengchao Zhou, Teng Zou, Zeming Wang, Xianwu Hu, Hongrui Meng, Yajun Wu, Chuxin Zhang, Caihua Wan, Yufeng Xie 0001 |
ISCAS | 9 |
| 2023 | A 40nm 150 TOPS/W High Row-Parallel MRAM Compute-in-Memory Macro with Series 3T1MTJ Bitcell for MAC OperationabstractNon-volatile Compute-in-Memory (CIM), especially high-speed MRAM CIM, promises to be a solution of “Memory Wall” problem in power-sensitive artificial intelligence edge devices. However, the low resistance and low on/off ratio limit the row parallelism and efficiency of MRAM CIM macros. To overcome these challenges, this work proposes the following: 1) a series 3T1MTJ bit-cell CIM architecture; 2) an input-aware and self-generated dynamic reference array; 3) a high-speed readout pipeline circuit. The proposed macro eliminates errors of high Row-Parallel multiply-and-accumulate (MAC) operation with 150 TOPS/W peak energy efficiency simulated using 40nm process and STT-MTJ. Zizhao Ma, Xianwu Hu, Gan Wen, Xiaoyang Zeng, Yufeng Xie 0001 |
ISCAS | 6 |
| 2022 | A High Area-Efficiency RRAM-Based Strong PUF with Multi-Entropy Source and Configurable Double-Read ProcessabstractPhysically Unclonable Functions (PUFs) are emerging security primitives for authentication due to its high physical security. Especially for those with excellent area-efficiency and reliable immunity against attacks, the demand is larger. In order to achieve higher security and area-efficiency, this paper proposes a strong PUF based on resistive random-access memory (RRAM). We exploit both the switching randomness and intrinsic resistance distribution of RRAM to increase the randomness of entropy source, and design a novel strong PUF structure with double-read process to increase the challenge-response pairs (CRPs) and area-efficiency. A double XOR process is proposed to enhance the immunity against machine learning attack (MLA) with low area-overhead. Compared with the state of the art, the number of CRP has been greatly improved, demonstrating a better area-utilization. Simulation results show that the CRP generation time is 1. 8us, mean intra-HD of1.93%, inter-HD of 49.95% and uniformity of 49.02%. The above features of the proposed strong PUF make it a promising candidate for Internet of Things (IoTs) authentication applications. Xianwu Hu, Jiayun Feng, Zizhao Ma, Xiaoyang Zeng, Yufeng Xie 0001 |
ISCAS | 6 |