Yunho Jang

dblp:49/2889 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-5848-1287ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Big-Computing and Little-Storing STT-MRAM PIM Architecture With Charge Domain Based MAC Operation
abstract
Spin transfer torque magnetic random access memory (STT-MRAM) is a promising memory technology for processing in memory (PIM) thanks to its high endurance and relatively low device-to-device and cycle-to-cycle variations. However, the low OFF/ON ratio of STT device limits the number of active row-lines during multiply-accumulate (MAC) operations, degrading energy efficiency and computation speed. In this paper, we present an energy efficient and high speed Big-computing and Little-storing STT-MRAM PIM (BCLS-SP) architecture, which can increase the number of active row-lines with almost no area overhead. In the BCLS-SP architecture, a charge domain-based STT-MRAM PIM (CD-SP) structure is employed to concurrently activate many row-lines by improving MAC operation reliability. Filter-wise weight compression (FWC) and weight sharing (WS) are also devised to compress the weights stored in CD-SP, thus reducing area cost. In addition, the proposed architecture performs MAC operations with skipping zero-valued input (SZI) and zero-conversion scheme (ZCS) for better energy efficiency and performance. The simulations using 28nm CMOS process show that the BCLS-SP architecture shows energy reduction of 29% and performance improvement of 3.6 compared to the recent memristive device-based PIM using weight compression and input skipping.
Yunho Jang, Yeseul Kim, Jongsun Park 0001
IEEE Trans. Computers1
2024 STT-MRAM-based Near-Memory Computing Architecture with Read Scheme and Dataflow Co-Design for High-Throughput and Energy-Efficiency
abstract
Spin transfer torque magnetic random access memory (STT-MRAM)-based near-memory computing (NMC) architecture has been actively studied due to its potential for high-throughput and energy-efficient processing of AI algorithms. However, the low read bandwidth of STT-MRAM limits improvements in throughput of NMC architecture by failing to provide enough data to digital logic located near memory arrays. In this paper, we propose reference voltage recycling read scheme (RRS) and hybrid stationary dataflow scheme (HDS) to address the low read bandwidth issue in STT-MRAM. The proposed RRS allows for the recycling of previously developed reference bit-line voltage in a memory array during sequential read operations, enhancing read bandwidth as well as energy-efficiency of STT-MRAM. In addition, HDS balances the demand for input and weight data in digital logics located near memory arrays, facilitating full utilization of digital logics despite the limited read bandwidth of STT-MRAM. As a result, this approach can improve throughput of NMC architecture. The simulations using a 28nm CMOS process show that the proposed STT-MRAM-based NMC architecture can achieve TOPS/W of 8.45 and TOPS/mm2 of 1.37 with 8-bit input and weight, regardless of the data patterns.
Yunho Jang, Yeseul Kim, Jongsun Park 0001
ISLPED1
2022 Stochastic SOT Device Based SNN Architecture for On-Chip Unsupervised STDP Learning
abstract
Emerging device based spiking neural network (SNN) hardware design has been actively studied. Especially, energy and area efficient synapse crossbar has been of particular interest, but processing units for weight summations in synapse crossbar are still a main bottleneck for energy and area efficient hardware design. In this paper, we propose an efficient SNN architecture with stochastic spin-orbit torque (SOT) device based multi-bit synapses. First, we present SOT device based synapse array using modified gray code. The modified gray code based synapse needs only N devices to represent 2^N levels of synapse weights. Accumulative spike technique is also adopted in the proposed synapse array, to improve ADC utilization and reduce the number of neuron updates. In addition, we propose hardware friendly algorithmic techniques to improve classification accuracies as well as energy efficiencies. Non-spike depression based stochastic spike-timing-dependent plasticity is used to reduce the overlapping input representation and classification error. Early read termination is also employed to reduce energy consumption by turning off less associated neurons. The proposed SNN processor has been implemented using 65nm CMOS process, and it shows 90% classification accuracy in MNIST dataset consuming 0.78J/image (training) and 0.23J/image (inference) of energy with an area of 1.12mm2.
Yunho Jang, Gyuseong Kang, Yeongkyo Seo, Kyung-Jin Lee, Byong-Guk Park, Jongsun Park 0001
IEEE Trans. Computers1
2022 SOT-MRAM Digital PIM Architecture With Extended Parallelism in Matrix Multiplication
abstract
Emerging device-based digital processing-in-memory (PIM) architectures have been actively studied due to their energy and area efficiency derived from analog to digital converter (ADC)-less PIM hardware. However, digital PIM architectures generally need large extra memories to copy parameters, and they also suffer from low computation per memory-cycle efficiencies. In this paper, we present a novel spin-orbit torque magnetic random access memory (SOT-MRAM) based digital PIM architecture to alleviate the extra memory size burden and computation cycle issues. First, we propose the spintronics-assisted logic-in-memory (SLIM) cells to support efficient digital logic operations inside memories, where the voltage-controlled magnetic anisotropy (VCMA) is exploited to enhance the computation per memory-cycle efficiencies. In addition, crossed input source PIM (CRISP) architecture is proposed to extend the merits of SLIM cells by eliminating the extra memories for parameter copying while significantly improving the degree of parallel processing. An intra-memory pipelining scheme is also considered to further increase the throughput of CRISP. The proposed CRISP architecture has been implemented using 28 nm CMOS process, and it presents 1.10 TOPS/W and 0.95 TOPS/mm2, showing considerable improvements of energy efficiency and throughput per area, compared to the state-of-the-art digital PIM architecture. Finally, to evaluate the impact of computation errors induced from the SOT devices and circuits in CRISP architecture, classification accuracy simulations have been performed while applying computation errors.
Yunho Jang, Min-Gu Kang, Byong-Guk Park, Kyung-Jin Lee, Jongsun Park 0001
IEEE Trans. Computers2
2022 A Dual-Domain Dynamic Reference Sensing for Reliable Read Operation in SOT-MRAM
abstract
Although spin orbit torque magnetic random access memory (SOT-MRAM) is one of the strong candidates for next-generation embedded memories, the degradation of read margin due to low tunnel magnetoresistance ratio (TMR) with process variations has been a large concern. In this paper, we present the dual-domain dynamic reference (DDDR) sensing scheme, where the reference voltage can be dynamically changed based on the combined voltage and time domain sensing to increase the sensing margin. The Half Schmitt trigger and sample & hold circuits are efficiently employed to generate data-dependent reference voltages and to store the sampled voltage levels at different times, respectively. According to the simulations using 28nm CMOS technology with 128 by 128 SOT-MRAM array, the proposed DDDR approach achieves a 243mV of sensing margin under 6.08E-8 bit-error-rate (BER) at 1.76ns, which is 2X larger margin with more than 100 times lower BER compared to the conventional read scheme. When scaling down the pre-charge voltage, the proposed scheme achieves more than 50% of the read energy savings under 1E-5 target BER condition.
Jooyoon Kim, Yunho Jang, Jongsun Park 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 Low Energy Domain Wall Memory Based Convolution Neural Network Design with Optimizing MAC Architecture
abstract
Running a convolutional neural network (CNN) algorithm using dedicated integrated circuits (ICs) on real-time portable applications is mainly restricted by slow performance and large power consumption. The power and delay are mainly due to external memory access, which incurs considerable energy consumption and bandwidth issues. In this paper, we propose an efficient convolution layer design using domain wall memory (DWM) for eliminating external memory access in image sensor embedded applications. A low energy access scheme using tag is employed to further reduce power consumption. The experimental results show that the proposed CNN architecture achieves 11.2% memory energy savings and 21.8% of MAC operation reduction compared to conventional architecture.
Jooyoon Kim, Yunho Jang, Jongsun Park 0001
ISCAS2
2020 Dynamic-Reference Based Early Write Termination for Low Energy SOT-MRAM
abstract
Although considered as one of the most viable emerging non-volatile memory, the spin-transfer-torque magnetic random access memory (STT-MRAM) suffers from its weakness in the write operation. Spin-orbit torque magnetic random access memory (SOT-MRAM) has been recently proposed to provide lower write energy consumption. Nevertheless, additional write energy reduction is still on demand for embedded memory purposes. In this paper, we propose an early write termination (EWT) technique for SOT-MRAM, which can greatly reduce the write energy consumption by efficiently removing the unnecessarily long write pulse. The proposed Dynamic Reference Early Termination (DRET) scheme provides energy savings in all write operations while guaranteeing reliable operation. Simulation results using 65nm CMOS technology show that 76.6% of write energy can be saved on average compared to the conventional SOT-MRAM.
Eunjong Yeo, Yunho Jang, Yeongkyo Seo, Jongsun Park 0001
ISCAS3
2018 Charge-Recycling based Redundant Write Prevention Technique for Low Power SOT-MRAM
abstract
While the spin transfer torque magnetic memory (STT-MRAM) suffers from its shortcomings such as high write power, slow write operation and reliability issues, spin orbit torque magnetic random access memory (SOT-MRAM) can offer relatively faster write operation with low power based on giant spin hall effect. Although SOT-MRAM provides low power write operation, to meet the power level of current embedded memories, significant reduction of write power is highly required. In this paper, we present a low power write technique for SOT-MRAM. In order to prevent redundant write operation, read-compare-write operation is adopted. As a result, only the SOT cells having different data are written, and write power is saved in the cells with the same data. For further optimization, bitline switching scheme is used to reduce bitline and source line swing in write operation. The negative bitline scheme is also exploited by re-cycling the charge from read operation to increase write current. Simulation results using 65nm CMOS technology show that up to 40.1 % of write energy can be saved compared to the conventional unnecessary write avoidance approach.
Gyuseong Kang, Yunho Jang, Jongsun Park 0001
ISCAS2
2018 Spin Orbit Torque Device based Stochastic Multi-bit Synapses for On-chip STDP Learning
abstract
As a large number of neurons and synapses are needed in spike neural network (SNN) design, emerging devices have been employed to implement synapses and neurons. In this paper, we present a stochastic multi-bit spin orbit torque (SOT) memory based synapse, where only one SOT device is switched for potentiation and depression using modified Gray code. The modified Gray code based approach needs only N devices to represent 2N levels of synapse weights. Early read termination scheme is also adopted to reduce the power consumption of training process by turning off less associated neurons and its ADCs. For MNIST dataset, with comparable classification accuracy, the proposed SNN architecture using 3-bit synapse achieves 68.7% reduction of ADC overhead compared to the conventional 8-level synapse.
Gyuseong Kang, Yunho Jang, Jongsun Park 0001
ISLPED2