Yue Zhang 0010

dblp:47/722-10 · DBLP profile ↗
← Back
31ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0001-6893-7199ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 27 · 3 first-author · 14 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A multi-population co-evolutionary algorithm based on dual-space division for dynamic multi-objective optimization problems
Yaru Hu, Jiaru Xia, Junwei Ou, Yanjie Song 0001, Jinhua Zheng, Gaige Wang, Yue Zhang 0010
Inf. Sci.8
2026 Closed-Loop CBRAM Crossbar System Toward Hardware Acceleration of Quantum Algorithms
abstract
Quantum computing is a compelling new technology that is becoming increasingly practical as time progresses. Quantum computers have the potential to solve problems of great complexity and magnitude across many different industries, utilizing appropriate quantum algorithms. Given the evolving stage of quantum computing, characterized by a limited number of operational quantum computers that demand extensive cooling, face decoherence issues, and incur significant fabrication costs, there exists a pronounced need for quantum computer simulators. In this work, a reconfigurable closed-loop CBRAM crossbar system has been developed for the execution and acceleration of quantum algorithm simulation, leveraging analog in-memory computing. The circuit supports a universal set of quantum gates representation, and through its reprogramming capabilities and feedback loop, can compute any quantum algorithm. To demonstrate its functionality, the 3-qubit Grover algorithm is executed on the proposed circuit. Building on the extensive circuit simulation, a comparative analysis of power and speed between the proposed nanoelectronic circuit and conventional hardware has been conducted, demonstrating the high efficiency and performance gains, followed by a scalability analysis. Furthermore, a framework has been developed that supports the design of custom closed-loop memristive crossbars through a graphical user interface (GUI), providing the user with the capability to execute quantum algorithms and examine the programming and computations of the circuit, assisting with the realization of a hardware prototype.
Iosif-Angelos Fyrigos, Theodoros Panagiotis Chatzinikolaou, Konstantinos Rallis, Vasileios G. Ntinas, Panagiotis Bousoulas, Dimitris Tsoukalas, Panagiotis Dimitrakis, Yue Zhang 0010, Georgios Ch. Sirakoulis
IEEE Trans. Circuits Syst. I Regul. Pap.8
2026 An Evolutionary Algorithm With Memory Guidance for Data Transmission Scheduling Optimization in Communication Satellite Network
abstract
With the rapid development of satellite technology, communication satellites have become an indispensable part of modern infrastructure. They serve as a key pillar for the future integrated communication satellite network (CSN). However, the increasing number of communication satellites presents significant challenges for data transmission between the satellite and ground station. This article focuses on transmitting communication data by scheduling resources for communication tasks. The goal of data transmission scheduling optimization in CSN (DTSOCSN) is to design a scheduling scheme that maximizes task profit across satellite–ground links, considering the constraints of the two working modes of communication satellites. To solve DTSOCSN, a mixed-integer programming model is developed, which incorporates various constraints such as the conditions for feed switching operation and the limitations of task execution windows. Based on the complexity of the problem, we propose an evolutionary algorithm with memory guidance (MGEA). The algorithm takes into account the memory dependence of Caputo fractional-order differential and innovatively designs a crossover operator, called Caputo crossover (CX). This crossover method uses the genetic information stored in memory to guide the crossover operation of the next generation, thereby forming a smooth optimization path and improving the search efficiency of the algorithm. In addition, an elite opposition-based heuristic initialization method and a tracking variation strategy are designed to enhance the algorithm’s ability to find high-quality initial solutions and perform local optimization. Experimental validation proceeds in two stages: first, multiscale simulations demonstrate MGEA’s superior performance over existing mainstream algorithms in task profit, convergence speed, resource utilization, and search efficiency. Second, to verify the contribution of the CX operator, it is integrated into several classical algorithms for comparative testing on benchmark problems. The results consistently show that algorithms using the CX operator achieve significant performance advantages compared with those relying on traditional crossover operators. This study not only provides an effective solution for DTSOCSN but also offers new idea for solving other types of satellite scheduling problems.
Qiuli Li, Yue Zhang 0010, Jiting Li, Witold Pedrycz, Ponnuthurai N. Suganthan, Rammohan Mallipeddi, Yanjie Song 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2025 HRAMTran: A Hybrid-RAM Transformer Accelerator With Dynamic Sparsity Floating-Point CIM and Written-Back Transpose Array
abstract
Transformer model performs outstandingly in various tasks involving artificial intelligence. In this work, we propose a hybrid-RAM Transformer accelerator (HRAMTran) utilizing computing-in-memory (CIM) based on spin-orbit torque magnetic random access memory (SOT-MRAM) and static RAM (SRAM), which supports dynamic sparsity in floating-point (FP) matrix multiplication (MM) and written-back transpose, thereby realizing efficient attention mechanism. First, a dynamic sparsity-based MM scheme is proposed, which dynamically ignores low-impact elements during vector multiplication, thereby effectively reducing the latency and energy consumption of MM. Second, a data-reuse multiply-and-accumulate (MAC) scheme for mantissa is designed to further optimize MM, which shares partial operation result to reduce redundant computation. The SOT-MRAM and SRAM based CIM architectures with dynamic sparsity and data-reuse schemes are constructed to perform weight (Query (Q), Key (K), and Value (V)) and dynamic MM, respectively. This hybrid-RAM CIM method can realize the optimization of energy and latency during attention mechanism computation. Moreover, written-back transpose SRAM array that can write multiple bits into a column simultaneously is designed to significantly reduce write-back cycles for KT. Finally, the HRAMTran accelerator is built to evaluate the performance of transformer implementation through performing machine translation for the WMT14 dataset. Results show that this accelerator realizes 3.6 µJ/Token and 68.77 TFLOPS/W, achieving 4.33× and 2.39× improvement compared with the state-of-the-art transformer accelerator.
Xianan Zhu, Zhengkun Gu, Zhizhong Zhang 0004, Kun Zhang 0030, Weisheng Zhao 0001, Yue Zhang 0010
ICCAD10
2025 A Heterogeneous System With Computing in Memory Processing Elements to Accelerate CNN Inference
abstract
Computing in memory (CIM) is one of the promising solutions to improve computing performance by integrating logic in memory. This work presents an efficient heterogeneous system based on the ultrafast CIM architecture (HS-CIM) to accelerate convolutional neural network (CNN) inference. First, an ultrafast CIM architecture is proposed based on the static random access memory (SRAM) by utilizing the novel total input and full digital scheme to implement multiply-and-accumulate (MAC) operation, which effectively addresses the high delay issue caused by high-precision computing in CIM architecture. Second, a heterogeneous system based on the proposed CIM architecture (HS-CIM) has been constructed with an aligned global cache and an adaptive pruning scheme to eliminate performance degradation and accuracy loss caused by input data bandwidth limitations. Meanwhile, efficient input data and weight data mapping schemes are proposed to minimize the delay and energy caused by input data transmission from the cache to the CIM architecture, thus realizing efficient CNN inference in the HS-CIM system. Finally, we analyze the performance of HS-CIM at the layout level by implementing the LeNet-5 and VGG models of CNN to recognize the image of MNIST and CIFAR-10 datasets, respectively. Results show that the energy efficiency of the proposed CIM architecture achieves 58.32 TOPS/W with 8-bit input/weight precision. Meanwhile, the inference accuracy of the HS-CIM system for MNIST and CIFAR-10 is 99.25% and 92.34%, respectively.
Youxiang Chen, Zhengkun Gu, Haiming Qiu, Kun Zhang 0030, Weisheng Zhao 0001, Yue Zhang 0010
IEEE Trans. Circuits Syst. I Regul. Pap.9
2025 A High-Speed, Low-Power, High-Reliability and Fully Single Event Double Node Upset Tolerant Design for Magnetic Random Access Memory
abstract
Magnetic Random Access Memory (MRAM) has enormous application potential in the aerospace field due to its nonvolatile, high speed, low power, and inherent radiation resistance characteristics. Due to its high sensing reliability, pre-charge differential sense amplifier (PCDSA) has been proposed and widely used in MRAM products. However, such PCDSA is based on traditional CMOS technology, and as the size of CMOS technology continues to shrink, its sensing result is easily affected by single event upset (SEU) or even the single event double node upset (SEDU). Recently, a TSC-PCDSA has been proposed to fully tolerate SEDU. However, it still suffers from slow speed, high power consumption and low reliability during normal sense operation. To address these issues, this paper proposes a novel PCDSA circuit that uses 6 three-input approximate C-elements (TACs) and 2 three-input standard C-elements (TSCs) to provide SEDU-tolerance. By reducing the number of transistors on the discharge path and increasing the difference in discharge current, the proposed PCDSA can achieve high speed, low power and high reliability. By using a physics-based STT-MTJ compact model and a commercial CMOS 40 nm design kit, hybrid simulations have been performed to demonstrate its functionality and evaluate its performance. Simulation results show that when the TMR is 150%, the width of N1-N12 is 480 nm and the$\text {V}_{\text {DD}}$is 1.1 V, the proposed PCDSA sensing error rate (SER) is close to 0% during normal sense operation, achieving a high sense speed of 123.6 ps and a low sense energy of 1.6533 fJ. Compared with the previously proposed TSC-PCDSA, the sense reliability is greatly improved, and the sense time and sense energy are reduced by 1.84 times and 1.27 times, respectively. Moreover, the proposed PCDSA can fully tolerate SEDU by optimizing the layout design. In the worst case where deposited charge$Q_{\text {inj}}$is 2 pC, it can achieve a shorter recover time of 1.28244 ns and a lower recover energy dissipation of 2.1604 pJ than the previously proposed TSC-PCDSA.
Shixuan Wang, Yue Zhang 0010, Weisheng Zhao 0001, Lang Zeng, Deming Zhang
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 A 32 kb 55 nm Radiation-Hardened SRAM Chip With SEU ≤1.1 E-11 Upsets/Bit-Day, SEL >107.1 MeV ⋅ cm²/mg, and TID >100 Krad(Si) for Space Applications
abstract
In this paper, a 32kb radiation-hardened (RH) static random access memory (SRAM) chip, named BH55RHSRAM32K, is proposed and fabricated for space applications. The chip is hardened from the view of the circuit level, layout level, and system level and is fabricated using a 55 nm CMOS process design kit with an RH cell library. At the circuit level, the proposed RH-14T SRAM cell and radiation-hardened pre-charged sense amplifier (RH-PCSA) cell adopt a polarity hardening method, making them fully tolerant of single event upset (SEU). At the layout level, the sensitive nodes in the proposed RH-14T SRAM cell and RH-PCSA cell layouts are isolated. Furthermore, the proposed RH-14T SRAM array adopts a bit-interleaved design, effectively reducing single event double upsets (SEDU). At the system level, an error correction coding (ECC) circuit is implemented to enhance SEU tolerance. Experimental results show that the proposed 32kb RH-SRAM chip can not only obtains superior radiation tolerance, i.e., the SEU ≤ 1.1E-11 upsets/bit-day, the SEL > 107.1 MeV⋅cm2/mg, and the TID > 100 Krad(Si), but also a faster access speed of < 10 ns and a lower write power consumption of 14.664 mW in comparison with the related products.
Deming Zhang, Dingyi Luo, Lang Zeng, Bi Wang 0002, Yue Zhang 0010, Weisheng Zhao 0001
IEEE Trans. Circuits Syst. I Regul. Pap.7
2025 Dynamic Challenge Cross-Selection Physical Unclonable Function Based on MRAM
abstract
The rapid development of Internet of Things (IoT) devices has triggered massive data transmission. Meanwhile, advances in artificial intelligence (AI) introduce new security vulnerabilities in device interactions. These challenges demand lightweight yet robust security solutions. In this context, physical unclonable functions (PUFs) serve as critical hardware security primitives, enabling reliable authentication for edge devices. Nevertheless, PUF is increasingly susceptible to novel threats, notably machine learning attacks. To address this security vulnerability to attacks, we propose a novel double-layer dynamic challenge cross-selection magnetoresistive random access memory PUF (MPUF). This design leverages the inherent process variation in spin-transfer torque magnetoresistive random access memory (STT-MRAM) as an entropy source. The proposed structure incorporates an obfuscation decode circuit (ODC) that combinesxorgates and shift registers. It dynamically obfuscates interlayer relationships between two PUF arrays to enhance circuit nonlinearity. The simulation results demonstrate uniformity of 50.16%, uniqueness of 49.94%, a worst bit error rate (BER) of 2.34% for$- 25~^{\circ } $C to$125~^{\circ } $C and 1.56% for$0.5\sim 1.1$V. In addition, four common machine learning models are used to attack this PUF, achieving accuracies of 50.49%, 50.49%, 50.48%, and 58.41%, which are close to a random guess. Compared with traditional PUF implementations, this work exhibits higher reliability and enhanced security while maintaining low power consumption of approximately 9.975 fJ/bit.
Siying Wu, Yu Gong 0002, Jiaao Dai, Shouzhong Peng, Yue Zhang 0010, You Wang 0002, Weiqiang Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2024 FRM-CIM: Full-Digital Recursive MAC Computing in Memory System Based on MRAM for Neural Network Applications
abstract
Computing in memory (CIM) realizes energy-efficient neural network algorithms by implementing highly parallel multiply-and-accumulate (MAC) operation. However, the MAC delay of CIM will sharply increase with the improvement of computing precision, which restricts its development. In this work, we propose a full-digital recursive MAC (FRM) operation based on spin-transfer-torque magnetic random access memory (STT-MRAM) CIM system to enable fast and energy-efficient image recognition application. First, the fast FRM scheme is proposed by utilizing the recursive operations of read and addition in segmented bit-line array, which effectively reduces the delay of MAC operations to 3.5ns and 4ns for 8-bit and 16-bit input and weight precision, respectively. Second, we design an image recognition system using FRM-CIM architecture as the processing element (PE), where the adaptive pruning method for layers is proposed to improve the compatibility of it with the neural network. By performing image recognition for the MNIST and CIFAR-10 datasets, results show that the throughput and energy efficiency of the FRM-CIM system are 58.51TOPS/mm2 and 11.3--56.72 TOPS/W under 8--16-bit precision, which are improved by 4.3 times and 2.6 times compared with the state-of-the-art works. Finally, the recognition accuracy can reach 96.65% and 82.7% on MNIST and CIFAR-10, respectively.
Zhengkun Gu, Youxiang Chen, Weisheng Zhao 0001, Yue Zhang 0010
DAC7
2024 Data-driven dynamic pricing and inventory management of an omni-channel retailer in an uncertain demand environment
Rui Wang 0017, Yue Zhang 0010, Yanjie Song 0001, Lining Xing 0001
Expert Syst. Appl.4
2024 APIM: An Antiferromagnetic MRAM-Based Processing-In-Memory System for Efficient Bit-Level Operations of Quantized Convolutional Neural Networks
abstract
Quantized Convolutional Neural Network (QCNN) is an attractive approach that reduces hardware overheads, especially for energy-constrained systems. However, existing QCNNs still require non-trivial hardware resources and memory capacity in order not to compromise model accuracy. To address this issue, we propose an antiferromagnetic magnetic random-access memory (ARAM)-based processing-in-memory (PIM) system, leveraging bit-level sparsity. Three optimization techniques are proposed to optimize hardware resource utilization while preserving CNN accuracy. Firstly, the ARAM-based memory subsystem allows dynamic adaptation of variable bit-width across CNN layers. Secondly, the bit-level accelerator employs the bit-fusion format engineered for processing data from the ARAM subsystem. Thirdly, a customized data path within the RISC-V core guarantees efficient instruction processing to the ARAM-based memory subsystem and bit-level accelerator, enabling optimal bit-level data transmission and computation. Experimental results demonstrate that this design remarkably reduces data movement by 50%-83% across existing CNNs. Compared to state-of-the-art designs, it enhances throughput and latency by an average of 5x and 10x, respectively. In addition, this design achieves speedups between 1.63x and 2.96x, outstripping other designs in AlexNet, VGG16, and ResNet18 benchmarks.
Yueting Li 0001, Daoqian Zhu, Jinhao Li 0007, Ao Du, Yue Zhang 0010, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2024 RSACIM: Resistance Summation Analog Computing in Memory With Accuracy Optimization Scheme Based on MRAM
abstract
Computing in memory (CIM) has become a promising candidate to address the Von Neumann bottleneck in processors designed for data-intensive applications. In this article, we propose a resistance summation analog computing in memory (RSACIM) with accuracy optimization scheme in spin transfer torque magnetic random access memory (STT-MRAM), in order to realize energy-efficient and highly reliable analog multiply-and-accumulation (MAC) operation. Firstly, we construct a resistance summation array by serial magnetic tunnel junctions (MTJs) to perform analog MAC operation utilizing time domain technology. Secondly, in order to reduce the impact of position-dependent error caused by resistance summation mechanism, we propose an accuracy optimization scheme to maximize the sensing margin (SM) and computation accuracy. Finally, we design a power-gated reconfigurability control scheme to implement power saving corresponding to different precisions for both input and weight. Evaluation on a 2 Kb RSACIM architecture shows an energy efficiency of 92.9 TOPS/W. System level simulation shows that comparing to existing CIMs based on MRAM, RSACIM architecture saves the inference energy by 4.2 times with 8.4 times lower latency in CIFAR10 image classification task.
Zhengkun Gu, Youxiang Chen, Kun Zhang 0030, Youguang Zhang, Yue Zhang 0010
IEEE Trans. Circuits Syst. I Regul. Pap.8
2024 Generalized Model and Deep Reinforcement Learning-Based Evolutionary Method for Multitype Satellite Observation Scheduling
abstract
Multitype satellite observation, including optical observation satellites, synthetic aperture radar (SAR) satellites, and electromagnetic satellites, has become an important direction in integrated satellite applications due to its ability to cope with various complex situations. In the multitype satellite observation scheduling problem (MTSOSP), the constraints involved in different types of satellites make the problem challenging. This article proposes a mixed-integer programming model and a generalized profit representation method in the model to effectively cope with the situation of multiple types of satellite observations. To obtain a suitable observation plan, a deep reinforcement learning-based genetic algorithm (DRL-GA) is proposed by combining the learning method and genetic algorithm. The DRL-GA adopts a solution generation method to obtain the initial population and assist with local search. In this method, a set of statistical indicators that consider resource utilization and task arrangement performance are regarded as states. By using deep neural networks to estimate the$Q$value of each action, this method can determine the preferred order of task scheduling. An individual update strategy and an elite strategy are used to enhance the search performance of DRL-GA. Simulation results verify that DRL-GA can effectively solve the MTSOSP and outperforms the state-of-the-art algorithms in several aspects. This work reveals the advantages of the proposed generalized model and scheduling method, which exhibit good scalability for various types of observation satellite scheduling problems.
Yanjie Song 0001, Junwei Ou, Witold Pedrycz, Ponnuthurai N. Suganthan, Xinwei Wang 0006, Lining Xing 0001, Yue Zhang 0010
IEEE Trans. Syst. Man Cybern. Syst.7
2023 TAM: A Computing in Memory based on Tandem Array within STT-MRAM for Energy-Efficient Analog MAC Operation
abstract
Computing in memory (CIM) has been demonstrated promising for energy efficient computing. However, the dramatic growth of the data scale in neural network processors has aroused a demand for CIM architecture of higher bit density, for which the spin transfer torque magnetic RAM (STT-MRAM) with high bit density and performance arises as an up-and-coming candidate solution. In this work, we propose an analog CIM scheme based on tandem array within STT-MRAM (TAM) to further improve energy efficiency while achieving high bit density. First, the resistance summation based analog MAC operation minimizes the effect of low tunnel magnetoresistance (TMR) by the serial magnetic tunnel junctions (MTJs) structure in the proposed tandem array with smaller area overhead. Moreover, a read scheme of resistive-to-binary is designed to achieve the MAC results accurately and reliably. Besides, the data-dependent error caused by MTJs in series has been eliminated with a proposed dynamic selection circuit. Simulation results of a 2Kb TAM architecture show 113.2 TOPS/W and 63.7 TOPS/W for 4-bit and 8-bit input/weight precision, respectively, and reduction by 39.3% for bit-cell area compared with existing array of MTJs in series.
Zhengkun Gu, Zuolei Hao, Weisheng Zhao 0001, Yue Zhang 0010
DATE7
2023 A Novel 9T1C-SRAM Compute-In-Memory Macro With Count-Less Pulse-Width Modulation Input and ADC-Less Charge-Integration-Count Output
abstract
This paper presents a novel compute-in-memory (CIM) macro, which mainly consists of three modules: input generator, 9T1C-SRAM CIM array and charge-integration-count output (CICO) circuit. For the input generator, it can achieve the pulse-width modulation mapping scheme without counts, leading to a small area overhead. For the CIM array, one row-cascade current mirror circuit instead of a bias voltage source is shared by all CIM cells in one row. And in each CIM cell, its multiply result is characterized by the charge on its capacitor. Based on the charge-sharing principle, the accumulated result is represented by the charge ($\text{Q}_{\text {CBL}}$) on the charge-bit-line (CBL). In this way, the voltage on the CBL is limited regardless of the number of rows of the CIM array, allowing the large-scale CIM array. For the CICO circuit, it is proposed to quantify the$\text{Q}_{\text {CBL}}$without the ADC, aiming to achieve high area efficiency. With the 14nm FinFET design kit, the design specification of the proposed CIM macro is introduced in detail and its performance is evaluated. Simulation results show that the proposed CIM macro can achieve 4–1370 TOPS/W energy efficiency with IN/W/OUT precision of 6/1/6b and 98.48%/84.56% test accuracy on MNIST and CIFAR-10.
Deming Zhang, Zhipeng Guo 0006, You Wang 0002, Yue Zhang 0010, Lang Zeng
IEEE Trans. Circuits Syst. I Regul. Pap.7
2022 Memristor Crossbar Arrays Performing Quantum Algorithms
abstract
There is a growing interest in quantum computers and quantum algorithm development. It has been proved that ideal quantum computers, with zero error rates and large decoherence times, can solve problems that are intractable for today’s classical computers. Quantum computers use two resources, superposition and entanglement, that have no classical analog. Since quantum computer platforms that are currently available comprise only a few dozen of qubits, the use of quantum simulators is essential in developing and testing new quantum algorithms. We present a novel quantum simulator based on memristor crossbar circuits and use them to simulate well-known quantum algorithms, namely the Deutsch and Grover quantum algorithms. In quantum computing the dominant algebraic operations are matrix-vector multiplications. The execution time grows exponentially with the simulated number of qubits, causing an exponential slowdown in quantum algorithm execution using classical computers. In this work, we show that the inherent characteristics of memristor arrays can be used to overcome this problem and that memristor arrays can be used not only as independent quantum simulators but also as a part of a quantum computer stack where classical computers accelerators are connected. Our memristive crossbar circuits are re-configurable and can be programmed to simulate any quantum algorithm.
Iosif-Angelos Fyrigos, Vasileios G. Ntinas, Nikolaos Vasileiadis, Georgios Ch. Sirakoulis, Panagiotis Dimitrakis, Yue Zhang 0010, Ioannis Karafyllidis
IEEE Trans. Circuits Syst. I Regul. Pap.6
2022 Reconfigurable Bit-Serial Operation Using Toggle SOT-MRAM for High-Performance Computing in Memory Architecture
abstract
Computing in memory (CIM) is a promising candidate for high throughput and energy-efficient data-driven applications, which mitigates the well-known memory bottleneck in Von Neumann architecture. In this paper, we present a reconfigurable bit-serial operation using toggle spin-orbit torque magnetic random access memory (TSOT-MRAM) to perform the computation completely in the bit-cell array instead of in a peripheral circuit. This bit-serial CIM (BSCIM) scheme achieves higher throughput and energy efficiency in CIM. First, basic Boolean logic operations are realized by utilizing the feature of TSOT device. A bit-cell array that implements the bit-serial operation is then built to provide the communication between column and row necessary for arithmetic operations, such as the carry propagation of addition and multiplication. Finally, we analyze the reliability of BSCIM scheme and demonstrate the performance advantage by performing convolution operations for$28\times 28$handwritten digit images in a BSCIM architecture. The results show that the delay and energy of BSCIM architecture are respectively reduced by 1.16-5.49 times and 1.12-1.43 times compared with the existing digital CIM architectures. Besides, its throughput and energy efficiency are also enhanced to 51.2 GOPS and 9.9 TOPS/W respectively.
Yining Bai, Zuolei Hao, Guanda Wang, Kun Zhang 0030, Youguang Zhang, Weifeng Lv, Yue Zhang 0010
IEEE Trans. Circuits Syst. I Regul. Pap.9
2021 Time-Domain Computing in Memory Using Spintronics for Energy-Efficient Convolutional Neural Network
abstract
The data transfer bottleneck in Von Neumann architecture owing to the separation between processor and memory hinders the development of high-performance computing. The computing in memory (CIM) concept is widely considered as a promising solution for overcoming this issue. In this article, we present a time-domain CIM (TD-CIM) scheme using spintronics, which can be applied to construct the energy-efficient convolutional neural network (CNN). Basic Boolean logic operations are implemented through recording the bit-line output at different moments. A multi-addend addition mechanism is then introduced based on the TD-CIM circuit, which can eliminate the cascaded full adders. To further optimize the compatibility of TD-CIM circuit for CNN, we also propose a quantization method that transforms floating-point parameters of pre-trained CNN models into fixed-point parameters. Finally, we build a TD-CIM architecture integrating with a highly reconfigurable array of field-free spin-orbit torque magnetic random access memory (SOT-MRAM) and evaluate its benefits for the quantized CNN. By performing digit recognition with the MNIST dataset, we find that the delay and energy are respectively reduced by 1.22.7 times and 2.4×103-1.1×104times compared with STT-CIM and CRAM based on spintronic memory. Finally, the recognition accuracy can reach 98.65% and 91.11% on MNIST and CIFAR10, respectively.
Yue Zhang 0010, Chenyu Lian, Yining Bai, Guanda Wang, Zhizhong Zhang 0004, Zhenyi Zheng, Kun Zhang 0030, Georgios Ch. Sirakoulis, Youguang Zhang
IEEE Trans. Circuits Syst. I Regul. Pap.1
2020 A Novel In-memory Computing Scheme Based on Toggle Spin Torque MRAM
abstract
This paper proposes a novel in-memory computing (IMC) scheme based on toggle spin torque magnetic random access memory (TST-MRAM), called TST-IMC, which makes full use of the unique TST writing mechanism. In this scheme, all of the computing results are directly written in bit-cells without transferring data out of the memory array. Varied Boolean logic operations, such as, NAND, NOR and XOR, can be achieved by specially configuring decision cells. We can also implement three-input majority logic through replacing a decision cell with a datum cell, which can further be used to realize the carry of full-adder. By using 28 nm CMOS technology node and 50 nm-diameter TST-MRAM, we perform mixed simulations to validate the functionality of the proposed TST-IMC scheme. Simulation results show that XOR logic operation can be carried out within 4 ns at 1.8 V supply voltage while the other basic logic operations can be faster, i.e. within 2 ns. In addition, TST-IMC 33% less time and 44% energy saved comparing with existing IMC schemes.
Yining Bai, Yue Zhang 0010, Guanda Wang, Zhizhong Zhang 0004, Zhenyi Zheng, Kun Zhang 0030, Weisheng Zhao 0001
ACM Great Lakes Symposium on VLSI2
2020 An In-memory Highly Reconfigurable Logic Circuit Based on Diode-assisted Enhanced Magnetoresistance Device
abstract
In the post-Moore era, in order to solve the problem of von Neumann bottleneck and memory wall caused by separation of memory and processor, in-memory-processing (IMP) technique has aroused great attention. Novel non-volatile memory (NVM) based on spintronic devices shows promise for satisfying the needs of low-power consumption and high speed for IMP. However, most spintronic memories based on magnetic tunnel junctions (MTJs) can only implement simple and specific logic functions due to the limits of single device and circuit structure. Otherwise, performing logic functions in memory generates vast dynamic power consumption during frequent reading and writing processes because of the high resistance of miniaturized MTJ. In this paper, we propose an in-memory highly reconfigurable logic circuit based on diode-assisted enhanced magnetoresistance (DEMR) device. Our circuit can realize 16 different logic functions with extremely limited circuit area benefiting from the special structure of DEMR device. With appropriate adjustment of control bit and current, the proposed circuit can further implement complex functions like full adder. The proposed reconfigurable circuit can flexibly meet the performance requirements in different scenarios and will contribute a lot for future in-memory chip design.
Yue Zhang 0010, Kun Zhang 0030, Zhizhong Zhang 0004, Youguang Zhang, Weisheng Zhao 0001
ACM Great Lakes Symposium on VLSI2
2020 Efficient Time-Domain In-Memory Computing Based on TST-MRAM
abstract
In-memory computing is highly promising to address the processor-memory data transfer bottleneck in current computational paradigm. We firstly propose a timedomain in-memory computing (TIMC) scheme based on highspeed low-power toggle spin torque random access memory (TST-MRAM). The difference of voltage drops of bitline caused by simultaneously-activated bit-cells is reflected to time domain. Reconfigurable logic operations can be performed by utilizing D flip-flops (DFFs) to record the outputs at different moments. In order to demonstrate the advantages of this scheme in terms of speed and energy consumption, an efficient multi-digit addition circuit has been designed and analyzed. Compared with existing IMC schemes, such as spin-transfer torque computing-in-memory (STT-CiM) structure, up to 67% energy saving and 10 times delay improvement can be achieved in the case of four-digit addition by using TIMC scheme.
Yue Zhang 0010, Chenyu Lian, Yining Bai, Guanda Wang, Kun Zhang 0030, Youguang Zhang, Weisheng Zhao 0001
ISCAS2
2018 Design Space Exploration of Magnetic Tunnel Junction based Stochastic Computing in Deep Learning
abstract
Magnetic tunnel junction (MTJ) is considered as a promising memory candidate in the more than Moore era because of high power efficiency, fast access speed, nearly infinite endurance and easy 3D integration. The nondeterministic switching behavior has been profited to exploit new directions for computing methods, such as stochastic computing. In this paper, the application of stochastic switching behavior in stochastic computing is explored for deep neural network (DNN). Stochastic computing method features low logic complexity, low energy consumption and fine-grained parallelism, boosting the performance of DNN system by combining MTJ. As a key block of stochastic computing, MTJ based true random number generator design is presented in details. The functionality has been validated by combining the hardware design and post-processing in software. Simulation results are demonstrated visibly by handwritten digits recognition test to show the accuracy. Furthermore, the performance is investigated in terms of accuracy, energy consumption and memory occupation to find more efficient techniques.
You Wang 0002, Yue Zhang 0010, Youguang Zhang, Weisheng Zhao 0001, Hao Cai 0001, Lirida A. B. Naviner
ACM Great Lakes Symposium on VLSI2
2017 A true random number generator based on parallel STT-MTJs
abstract
Random number generators are an essential part of cryptographic systems. For the highest level of security, true random number generators (TRNG) are needed instead of pseudorandom number generators. In this paper, the stochastic behavior of the spin transfer torque magnetic tunnel junction (STT-MTJ) is utilized to produce a TRNG design. A parallel structure with multiple MTJs is proposed that minimizes device variation effects. The design is validated in a 28-nm CMOS process with Monte Carlo simulation using a compact model of the MTJ. The National Institute of Standards and Technology (NIST) statistical test suite is used to verify the randomness quality when generating encryption keys for the Transport Layer Security or Secure Sockets Layer (TLS/SSL) cryptographic protocol. This design has a generation speed of 177.8 Mbit/s, and an energy of 0.64 pJ is consumed to set up the state in one MTJ.
Yuanzhuo Qu, Jie Han 0001, Bruce F. Cockburn, Witold Pedrycz, Yue Zhang 0010, Weisheng Zhao 0001
DATE5
2015 From device to system: cross-layer design exploration of racetrack memory
Guangyu Sun 0003, Chao Zhang 0007, Hehe Li, Yue Zhang 0010, Yizi Gu, Jacques-Olivier Klein, Dafine Ravelosona, Yongpan Liu, Weisheng Zhao 0001, Huazhong Yang
DATE4
2015 Perspectives of racetrack memory based on current-induced domain wall motion: From device to system
abstract
Current-induced domain wall motion (CIDWM) is regarded as a promising way towards achieving emerging high-density, high-speed and low-power non-volatile devices. Racetrack memory is an attractive concept based on this phenomenon, which can store and transfer a series of data along a magnetic nanowire. Although the first prototype has been successfully fabricated, its advancement is relatively arduous caused by certain technique and material limitations. Particularly, the storage capacity issue is one of the most serious bottlenecks hindering its application for practical systems. In this paper, we present two alternative solutions to improve the capacity of racetrack memory: magnetic field assistance and chiral domain wall (DW) motion. The former one can lower the current density for DW shifting; the latter one can utilize materials with low resistivity. Both of them are able to increase the nanowire length and allow higher feasibility of large-capacity racetrack memory. Furthermore, system level simulation shows that a racetrack memory based cache can improve system performance by about 15.8% and significantly reduces the energy consumption, compared to the SRAM counterpart.
Yue Zhang 0010, Chao Zhang 0007, Jacques-Olivier Klein, Dafine Ravelosona, Guangyu Sun 0003, Weisheng Zhao 0001
ISCAS1
2015 Spintronics: Emerging Ultra-Low-Power Circuits and Systems beyond MOS Technology
abstract
Conventional MOS integrated circuits and systems suffer serve power and scalability challenges as technology nodes scale into ultra-deep-micron technology nodes (e.g., below 40nm). Both static and dynamic power dissipations are increasing, caused mainly by the intrinsic leakage currents and large data traffic. Alternative approaches beyond charge-only-based electronics, and in particular, spin-based devices, show promising potential to overcome these issues by adding the spin freedom of electrons to electronic circuits. Spintronics provides data non-volatility, fast data access, and low-power operation, and has now become a hot topic in both academia and industry for achieving ultra-low-power circuits and systems. The ITRS report on emerging research devices identified themagnetic tunnel junction(MTJ) nanopillar (one of the Spintronics nanodevices) as one of the most promising technologies to be part of future micro-electronic circuits. In this review we will give an overview of the status and prospects of spin-based devices and circuits that are currently under intense investigation and development across the world, and address particularly their merits and challenges for practical applications. We will also show that, with a rapid development of Spintronics, some novel computing architectures and paradigms beyond classic Von-Neumann architecture have recently been emerging for next-generation ultra-low-power circuits and systems.
Wang Kang 0001, Yue Zhang 0010, Zhaohao Wang, Jacques-Olivier Klein, Claude Chappert, Dafine Ravelosona, Gefei Wang, Youguang Zhang, Weisheng Zhao 0001
ACM J. Emerg. Technol. Comput. Syst.2
2014 An overview of spin-based integrated circuits
abstract
Conventional CMOS integrated circuits suffer from serve power and scalability challenges as technology node scales into ultra-deep-micron technology nodes. Alternative approaches beyond charge-only based circuits. In particular, spin-based devices or integrated circuits show promising merits to overcome these issues by adding the spin freedom of electrons to the electronic circuits. Spintronics has now become a hot topic in both academics and industrials. This paper overviews the status and prospects of spin-based integrated circuits under intense investigation and address particularly their merits and challenges for practical applications.
Wang Kang 0001, Weisheng Zhao 0001, Zhaohao Wang, Jacques-Olivier Klein, Yue Zhang 0010, Djaafar Chabi, Youguang Zhang, Dafine Ravelosona, Claude Chappert
ASP-DAC5
2014 Spintronics for low-power computing
abstract
Microelectronics has been following Moore's law for almost 40 years. However this trend tends to run out of steam in recent technology nodes. The continuous improvements in the size of the transistors and in the operating frequencies result in serious power consumption, heat dissipation and reliability issues. Spintronics (Nobel Prize of Physics, 2007 awarded to Prof. Fert from Univ. Paris-Sud and Peter Grünberg from Forschungszentrum Jülich) nanodevices can reduce significantly the power, improve the reliability or allow new functionalities. The 2010 ITRS report on emerging research devices identified Magnetic Tunnel Junction (MTJ) nanopillar (the preeminent spintronics nanodevice) as one of the most promising technologies to be part of the future microelectronics circuits. It provides data non-volatility, hardness to radiations, fast data access and low-power operations. Magnetic memories become the most promising candidate for both low power logic computing and the data storage. This tutorial paper presents multi-discipline questions (Device, Circuit, Architecture, System and CAD) related to this topic to share the most recent results and discuss the future challenges.
Yue Zhang 0010, Weisheng Zhao 0001, Jacques-Olivier Klein, Wang Kang 0001, Damien Querlioz, Youguang Zhang, Dafine Ravelosona, Claude Chappert
DATE1
2014 Design and analysis of crossbar architecture based on complementary resistive switching non-volatile memory cells
Weisheng Zhao 0001, Jean-Michel Portal, Wang Kang 0001, Mathieu Moreau, Yue Zhang 0010, Hassen Aziza, Jacques-Olivier Klein, Zhaohao Wang, Damien Querlioz, Damien Deleruyelle, Marc Bocquet, Dafine Ravelosona, Christophe Muller, Claude Chappert
J. Parallel Distributed Comput.5
2013 Spin-electronics based logic fabrics
abstract
Advanced computing ICs in ultra deep-micron technology nodes (e.g. 40 nm) suffer from high power issues, which become one of the major bottlenecks for the future performance progress. Both static and dynamic power dissipation are increasing, caused mainly by the intrinsic leakage currents and large data traffic. Alternative approaches beyond charge-based logic circuits become hot research topics to overcome these issues definitively. By integrating the spin freedom of electrons to electronic devices, spin-electronics is promising for ultra-low power computing as it can provide non-volatility, fast data control and high logic density etc. Today, most of large microelectronics industries investigate this emerging field. In this invited paper for the special session “Nanoscale logic fabrics”, we overview spin-electronics based logic fabrics under intense investigation and address particularly the impact of this technology on logic architectures and new computing paradigms.
Weisheng Zhao 0001, Jacques-Olivier Klein, Zhaohao Wang, Yue Zhang 0010, Nesrine Ben Romdhane, Damien Querlioz, Dafine Ravelosona, Claude Chappert
VLSI-SoC4
2011 Embedded MRAM for high-speed computing
abstract
As the fabrication technology node shrinks down to 90nm or below, high standby power becomes one of the major critical issues for CMOS high-speed computing circuits (e.g. logic and cache memory) due to the high leakage currents. A number of non-volatile storage technologies such as FeRAM, MRAM, PCRAM and RRAM and so on, are under investigation to bring the non-volatility into the logic circuits and then eliminate completely the standby power issue. Thanks to its infinite endurance, high switching/sensing speed and easy 3D integration after CMOS process, MRAM is considered as the most promising one. Numerous logic circuits based on MRAM technology have been proposed and prototyped in the last years. In this paper, we present an overview and current status of these logic circuits and discuss their potential applications in the future from both the physics and architecture points of view.
Weisheng Zhao 0001, Yue Zhang 0010, Yahya Lakys, Jacques-Olivier Klein, Daniel Etiemble, D. Revelosona, Claude Chappert, Lionel Torres, Vitorio Cargnini, Raphael Martins Brum, Yoann Guillemenet, Gilles Sassatelli
VLSI-SoC2