Hongrui Meng

dblp:409/9502 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0002-6103-9593ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 BLCIM: An Efficient Radix-16 Booth LUT-Based SRAM-CIM Architecture with Algorithm-Hardware Co-Optimization for NTT
abstract
Lattice-based cryptography relies heavily on the Number Theoretic Transform (NTT), whose performance is dominated by modular multiplication and data movement. This paper proposes BLCIM, the first Radix-16 Booth LUT-based compute-in-memory (CIM) NTT accelerator. We propose an algorithm that precomputes partial modular multiplications with fixed rotation factors and uses input Booth-encoded search results. Then, based on the Booth encoding, it decides whether to perform shifting and inversion to simplify modular multiplication. In addition to the proposed sparsity-aware and stage-skipping schemes, the Radix-16 Booth LUT-based algorithm significantly improves NTT performance at low energy cost. By implementing the algorithm on the SRAM-CIM architecture with a lightweight pipeline design, we achieved nearly 100% utilization for the NTT circuit during acceleration. Simulated in 28 nm CMOS technology, the proposed BLCIM achieves only 4.53% latency and 58.87% energy consumption of the latest work.
Qianhua Li, Hongrui Meng, Chunshan Wang, Shengchao Zhou, Teng Zou, Yufeng Xie 0001
ACM Great Lakes Symposium on VLSI2
2026 Horizontal-Parallel ADC-less Sparsity-Clock-Aware RRAM CIM Macro for edge AI devices
Teng Zou, Shengchao Zhou, Hongrui Meng, Yufeng Xie 0001
ISCAS3
2026 A 40-nm Training-Inference STT-MRAM Near-Memory Computing Macro for Memory-Augmented Neural Network Acceleration
abstract
Recently, memory-augmented neural networks (MANNs) have gained significant attention as a critical solution for few-shot learning (FSL). These networks leverage external memory to store prior knowledge, thereby enhancing classification efficiency. Spin-transfer torque magnetic random access memory (STT-MRAM) is particularly suited for this application due to its compact cell size, excellent data retention, and scalability. In this article, we introduce a STT-MRAM-based near-memory computing (NMC) macro specifically designed for MANNs. Our approach incorporates several key innovations aimed at overcoming challenges in hardware implementation while improving MANN performance as follows: 1) a parallel computing architecture within the NMC to expedite$L1$distance computations; 2) a memory invert coding (MIC) and self-termination write (STW) scheme that reduce write operations and energy consumption, addressing the issues of frequent writes and high write currents during the training phase of MANNs; 3) a dynamic offset-compensation sense amplifier (DOC-SA) and high-throughput switch-capacitor (HTSC) readout scheme to improve read accuracy and throughput, tackling low read margins and limited readout bandwidth; 4) an exploration of MANN architectures validates the reusability of the NMC macro. The optimized matching-networks (MCHnets)-based structure achieves an accuracy exceeding 90% in five-way and eight-way Omniglot classification tasks. Fabricated with a 40-nm CMOS technology, our design achieves classification accuracies of 96.37% for eight-way-five-shot tasks and 93.72% for 16-way-five-shot tasks on the Omniglot dataset utilizing the optimized MCHnet, showcasing an impressive energy efficiency of 6.47 TOPS/W at the basis of 16-bit$L1$distance computing in the classification tasks of MANN.
Shengchao Zhou, Hongrui Meng, Yajun Wu, Zizhao Ma, Teng Zou, Tai Min, Shaohao Wang, Yufeng Xie 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2025 A 40nm STT-MRAM Near-Memory Computing Macro for Memory-Augmented Neural Network Acceleration
abstract
Memory-augmented neural network (MANN) has gained attention as a pivotal solution for few-shot learning (FSL). Among the candidates for associative memory in MANN accelerators, spin-transfer torque magnetic random-access memory (STT-MRAM) stands out for its compact cell area, long data retention time, and excellent scalability. In this paper, we propose an STT-MRAM near-memory computing (NMC) macro for MANN acceleration. The macro contains following innovations: 1) An array-level parallel computing architecture for L1 distance calculation. 2) A low-area-overhead memory-invert coding technique to reduce write energy consumption. 3) A configurable dynamic offset-compensation sense amplifier (CDOC-SA) to improve classification accuracy. Fabricated in 40nm CMOS process, our macro demonstrates an energy efficiency of 6.47 TOPS/W, achieving the classification accuracy of 98.3% and 93% for 8-way-5-shot tasks and 16-way-5-shot tasks on the Omniglot dataset.
Hongrui Meng, Yajun Wu, Shengchao Zhou, Zizhao Ma, Tai Min, Shaohao Wang, Yufeng Xie 0001
ISCAS1
2025 High sensing margin and parallelism 6T-2MTJ SOT-MRAM based TCAM for energy-Efficient Similarity Priority calculation in MANNs
abstract
With the development of AI applications, there is a demand for high parallel similarity priority calculation, like in MANNs. TCAM, as a type of memory for high parallel searching work, is suitable to perform the task. However, the CMOS based TCAMs suffer from area overhead, static power consumption and data-loss while power failure. The SOT-MRAM presents a promising alternative for TCAM design due to its nonvolatility, low area overhead and no static power consumption. But, its low on/off ratio results in low sensing margin which limits its processing speed. To address the challenges, we propose a novel TCAM structure with high sensing margin, parallelism and low energy consumption. This work proposes the followings: (1) A novel 6T-2MTJ sot-mram based TCAM structure is proposed. Simulations show it has a 2-3x improvement in sensing margin over other MRAM based ones, 1.37x improvement in searching energy (fJ/bit) over other nonvolatile technologies based ones, and 23% area reduction over other CMOS based ones. (2) A segmented power supply scheme is presented to meet the parallelism need of MANNs which improves the parallelism by 4-8x compared to no segmented one.
Shengchao Zhou, Teng Zou, Zeming Wang, Xianwu Hu, Hongrui Meng, Yajun Wu, Chuxin Zhang, Caihua Wan, Yufeng Xie 0001
ISCAS5