Tai Min

dblp:08/9547 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0003-4682-4091ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A 0.562 mm2 MTJ-Based Ising Machine for 48-bit Integer Factorization in 40 nm CMOS
Kunpeng Gao, Jiadong Chen, Shaohao Wang, Tai Min, Yufeng Xie 0001
ISCAS5
2026 A 40-nm Training-Inference STT-MRAM Near-Memory Computing Macro for Memory-Augmented Neural Network Acceleration
abstract
Recently, memory-augmented neural networks (MANNs) have gained significant attention as a critical solution for few-shot learning (FSL). These networks leverage external memory to store prior knowledge, thereby enhancing classification efficiency. Spin-transfer torque magnetic random access memory (STT-MRAM) is particularly suited for this application due to its compact cell size, excellent data retention, and scalability. In this article, we introduce a STT-MRAM-based near-memory computing (NMC) macro specifically designed for MANNs. Our approach incorporates several key innovations aimed at overcoming challenges in hardware implementation while improving MANN performance as follows: 1) a parallel computing architecture within the NMC to expedite$L1$distance computations; 2) a memory invert coding (MIC) and self-termination write (STW) scheme that reduce write operations and energy consumption, addressing the issues of frequent writes and high write currents during the training phase of MANNs; 3) a dynamic offset-compensation sense amplifier (DOC-SA) and high-throughput switch-capacitor (HTSC) readout scheme to improve read accuracy and throughput, tackling low read margins and limited readout bandwidth; 4) an exploration of MANN architectures validates the reusability of the NMC macro. The optimized matching-networks (MCHnets)-based structure achieves an accuracy exceeding 90% in five-way and eight-way Omniglot classification tasks. Fabricated with a 40-nm CMOS technology, our design achieves classification accuracies of 96.37% for eight-way-five-shot tasks and 93.72% for 16-way-five-shot tasks on the Omniglot dataset utilizing the optimized MCHnet, showcasing an impressive energy efficiency of 6.47 TOPS/W at the basis of 16-bit$L1$distance computing in the classification tasks of MANN.
Shengchao Zhou, Hongrui Meng, Yajun Wu, Zizhao Ma, Teng Zou, Tai Min, Shaohao Wang, Yufeng Xie 0001
IEEE Trans. Very Large Scale Integr. Syst.7
2025 A 40nm STT-MRAM Near-Memory Computing Macro for Memory-Augmented Neural Network Acceleration
abstract
Memory-augmented neural network (MANN) has gained attention as a pivotal solution for few-shot learning (FSL). Among the candidates for associative memory in MANN accelerators, spin-transfer torque magnetic random-access memory (STT-MRAM) stands out for its compact cell area, long data retention time, and excellent scalability. In this paper, we propose an STT-MRAM near-memory computing (NMC) macro for MANN acceleration. The macro contains following innovations: 1) An array-level parallel computing architecture for L1 distance calculation. 2) A low-area-overhead memory-invert coding technique to reduce write energy consumption. 3) A configurable dynamic offset-compensation sense amplifier (CDOC-SA) to improve classification accuracy. Fabricated in 40nm CMOS process, our macro demonstrates an energy efficiency of 6.47 TOPS/W, achieving the classification accuracy of 98.3% and 93% for 8-way-5-shot tasks and 16-way-5-shot tasks on the Omniglot dataset.
Hongrui Meng, Yajun Wu, Shengchao Zhou, Zizhao Ma, Tai Min, Shaohao Wang, Yufeng Xie 0001
ISCAS5
2024 A 3D MCAM architecture based on flash memory enabling binary neural network computing for edge AI
Maoying Bai, Shuhao Wu, Yueran Qi, Tai Min, Xuepeng Zhan, Jiezhi Chen
Sci. China Inf. Sci.9
2014 Hardware implementation of KLMS algorithm using FPGA
abstract
Fast and accurate machine learning algorithms are needed in many physical applications. However, the learning efficiency is badly subjected to the intensive computation. Knowing that hardware implementation could speed up computation effectively, we use a FPGA hardware platform to implement an on-line kernel learning algorithm, namely the kernel least mean square (KLMS) which adopts the simple survival kernel as the Mercer kernel. By using an on-line quantization method and pipeline technology, the requirement of hardware resources and computation burden can be reduced significantly and the data processing speed can be accelerated apparently without losing accuracy. Finally, a 128-way parallel FPGA platform which works at 200MHz is implemented. It could achieve an average speedup of 6553 versus Matlab running on a 3GHz Intel(R) Core(TM) i5-2320 CPU.
Xiaowei Ren, Pengju Ren, Badong Chen, Tai Min, Nanning Zheng 0001
IJCNN4
2011 Design techniques to improve the device write margin for MRAM-based cache memory
abstract
As one promising non-volatile memory technology, magnetoresistive RAM (MRAM) based on magnetic tunneling junctions (MTJs) has recently attracted much attention. However, latest device research has discovered that, in order to maintain sufficient MTJ write margin to prevent device breakdown, MTJs will be subject to unconventionally high random write error rates (e.g., 10-3 and above) as memory cell size is being scaled down. This new discovery seriously threatens the scalability of MRAM, and the material/device research community is actively searching for solutions to largely reduce MTJ write error rates and meanwhile maintain sufficient device write margin. In this paper, we attempt to address this challenge from the architecture level when using MRAM to implement cache memory. In particular, we show that two simple cache architecture design techniques can be used to effectively tolerate high MTJ write error rates at small performance and implementation cost, which makes it much easier to maintain sufficient MTJ write margin and hence push the MRAM scalability envelope. Using the full system simulator PTLsim and a variety of benchmarks, we show that the proposed design techniques can readily accommodate MTJ write error rate up to 0.75% at the penalty of less than 4% processor performance degradation, less than 10% silicon area overhead, and 6% energy consumption overhead.
Hongbin Sun 0001, Chuanyin Liu, Nanning Zheng 0001, Tai Min, Tong Zhang 0002
ACM Great Lakes Symposium on VLSI4