VLDB 2026 Research / reviewers in the wild / expert
Youguang Zhang
dblp:43/8567
· DBLP profile ↗
76ranked-venue papers
0as first author
15since 2021 · last 2025
0009-0008-0928-4210ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 3 since 2021Computer networks · 7Software engineering, systems software and programming languages · 5 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AM-CIM: Approximate Memory Based Near Sensor Compute-in-Memory Architecture for Keyword SpottingabstractCompute-In-Memory (CIM) has emerged as a promising solution to address the von-Neumann bottleneck, making it a key technology for intelligent computing in edge IoT devices, particularly for real-time applications like keyword spotting (KWS). However, traditional CIM architectures face challenges such as high resource consumption, especially in data conversion, which can significantly impact chip area and energy efficiency. To address these challenges, this work proposes a computational CIM architecture utilizing multilevel analog memory, named AM-CIM, tailored for near-sensor (NS) computation of real-time KWS applications. Additionally, approximate memory technology is integrated into the AM-CIM architecture, employing data resilience scheduling for analog memory which contributes to significant reductions in hardware overhead. This integration facilitates a hardware-software co-design approach. To deploy KWS tasks in AM-CIM, a gated recurrent unit (GRU) network, referred to as MAC-GRU, is implemented. By employing Mel-energy as the input feature at the near-sensor end, the system achieves a 93.13% reduction in feature extraction power consumption. Evaluation results based on TSMC 180-nm technology demonstrate that the AM-CIM architecture achieves an accuracy of 88.51% for 10-keyword classification with a power consumption of$546~\mu W$, while reducing analog memory area by 43.32%. Xiaotao Jia, Guangcai Yuan, Jianyi Yu, Cong Shi 0003, Qi Wei 0001, Youguang Zhang, Weisheng Zhao 0001, Fei Qiao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | CiTST-AdderNets: Computing in Toggle Spin Torques MRAM for Energy-Efficient AdderNetsabstractRecently, Adder Neural Networks (AdderNets) have gained widespread attention as an alternative to traditional Convolutional Neural Networks (CNNs) for deep learning tasks. AdderNets use lightweight addition operations to replace multiplication and accumulation (MAC) operations, but can keep almost the same accuracy compared to other CNNs. Nevertheless, challenges still exist with regards to hardware resources, power consumption, and communication bandwidth, primarily due to the ‘Von-Neumann bottlenecks’. However, computing-in-memory (CIM) architecture based on magnetic random-access memory (MRAM) has great potential for edge DNN implementation. In this paper, we propose a novel CIM paradigm using a novel Toggle-Spin-Torques (TST) driven MRAM for energy-efficient AdderNets (called CiTST_AdderNets). In CiTST_AdderNets, MRAM is driven by the interplay of the field-free spin orbit torque (SOT) effect and the spin transfer torque (STT) effect, which offers a fascinating prospect for energy efficiency and speed. Furthermore, a novel CIM paradigm is proposed to implement the dominating subtraction and sum operations in AdderNets, reducing data transfer and the related energy. Meanwhile, a highly parallel array structure integrating computation and storage is designed to support CiTST_AdderNets. In addition, a mapping strategy is proposed to efficiently map the convolution layer on the array. Fully connected layers can also be efficiently computed. The CiTST-AdderNets macro is designed by using a 65-nm CMOS process. Results show that our CiTST-AdderNets consumes about 1.65 mJ, 9.29 mJ, and 42.46 mJ for running VGG8, ResNet-50, and ResNet-18 respectively at 8-bit fixed-point precision. Compared to state-of-the-art platforms, our macro achieves an energy efficiency improvement of 1.45 x to 66.78 x. Lichuan Luo, Erya Deng, Dijun Liu, Zhen Wang 0070, Weiliang Huang, He Zhang 0011, Xiao Liu 0051, Jinyu Bai, Junzhan Liu, Youguang Zhang, Wang Kang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2024 | RSACIM: Resistance Summation Analog Computing in Memory With Accuracy Optimization Scheme Based on MRAMabstractComputing in memory (CIM) has become a promising candidate to address the Von Neumann bottleneck in processors designed for data-intensive applications. In this article, we propose a resistance summation analog computing in memory (RSACIM) with accuracy optimization scheme in spin transfer torque magnetic random access memory (STT-MRAM), in order to realize energy-efficient and highly reliable analog multiply-and-accumulation (MAC) operation. Firstly, we construct a resistance summation array by serial magnetic tunnel junctions (MTJs) to perform analog MAC operation utilizing time domain technology. Secondly, in order to reduce the impact of position-dependent error caused by resistance summation mechanism, we propose an accuracy optimization scheme to maximize the sensing margin (SM) and computation accuracy. Finally, we design a power-gated reconfigurability control scheme to implement power saving corresponding to different precisions for both input and weight. Evaluation on a 2 Kb RSACIM architecture shows an energy efficiency of 92.9 TOPS/W. System level simulation shows that comparing to existing CIMs based on MRAM, RSACIM architecture saves the inference energy by 4.2 times with 8.4 times lower latency in CIFAR10 image classification task. Zhengkun Gu, Youxiang Chen, Kun Zhang 0030, Youguang Zhang, Yue Zhang 0010 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | Variation Aware Evaluation Approach and Design Methodology for SOT-MRAMabstractSpin-orbit torque magnetic random access memory (SOT-MRAM), which exhibits sub-nanosecond write speed and high reliability, is a promising candidate for the future high-level cache. However, SOT-MRAM faces the problem of large bit-cell layout area due to its structural characteristics and write performance requirements, therefore it is necessary to explore the bit-cell design with optimal overall performance under the unified bit-cell area. In this paper, we propose a comprehensive variation aware evaluation approach for the area, latency, and energy of SOT-MRAM under the uniform yield standard. Based on this, the mainstream SOT-MRAM bit-cell designs with high-density method and multi-finger configuration are evaluated, meanwhile bit-cell designs with excellent write performance and their optimum area ranges are identified. Moreover, the source line read (SLR) mode with higher robustness against transistor variation is proposed to improve the read performance, and the dual SL (DSL) method is proposed to further reduce the read latency and write energy. With the DSL method, the read latency and write energy of 2-word-line (WL)-type bit-cells can be reduced by up to 36.5% and 12.6%, respectively. In addition, the DSL method can solve the shunt current issue of 1WL-type bit-cells and reduce the read latency and write energy by up to 43.6% and 17.4%, respectively. Chao Wang 0094, Zhaohao Wang, Shixing Li, Zhongkui Zhang, Youguang Zhang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | Sea Surface Temperature Retrievals Using K- and Ka-Bands With Weak Brightness Temperature Response Residual Neural NetworksabstractSea surface temperature (SST) measurements are crucial in the context of climate change. Microwave SST measurements are currently provided by radiometers operating in the C- and X-bands. In-orbit K- and Ka-band payloads lack the commonly used C- and X-bands for SST retrieval. We present the K-KaSSTNet, a residual neural network (NN) that, for the first time, uses the K and Ka microwave bands with much weaker SST response than C- and X-bands for SST retrieval. Despite training on a limited dataset from 2020 to 2021, K-KaSSTNet consistently achieves reasonable accuracy SST retrievals for data spanning 2017–2022. Moreover, by using deep learning (DL) interpretability methods, we have unveiled the underlying mechanisms driving K-KaSSTNet. When extended to the Special Sensor Microwave Imager/Sounder (SSMIS) and Calibration Microwave Radiometers (CMRs)—payloads typically not used for SST retrieval—the K-KaSSTNet model maintains SST retrievals with reasonable accuracy compared with Advanced Microwave Scanning Radiometer-2 (AMSR-2). This extension broadens the spatiotemporal coverage of microwave SST products and enhances the temporal sampling frequency and continuity of microwave SST measurements. Peng Mao, Xiaobin Yin, Youguang Zhang, Ning Wang 0100, Yan Li 0119, Qing Xu 0009, Xingwei Jiang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | An Energy-Efficient Bayesian Neural Network Implementation Using Stochastic Computing MethodabstractThe robustness of Bayesian neural networks (BNNs) to real-world uncertainties and incompleteness has led to their application in some safety-critical fields. However, evaluating uncertainty during BNN inference requires repeated sampling and feed-forward computing, making them challenging to deploy in low-power or embedded devices. This article proposes the use of stochastic computing (SC) to optimize the hardware performance of BNN inference in terms of energy consumption and hardware utilization. The proposed approach adopts bitstream to represent Gaussian random number and applies it in the inference phase. This allows for the omission of complex transformation computations in the central limit theorem-based Gaussian random number generating (CLT-based GRNG) method and the simplification of multipliers as AND operations. Furthermore, an asynchronous parallel pipeline calculation technique is proposed in computing block to enhance operation speed. Compared with conventional binary radix-based BNN, SC-based BNN (StocBNN) realized by FPGA with 128-bit bitstream consumes much less energy consumption and hardware resources with less than 0.1% accuracy decrease when dealing with MNIST/Fashion-MNIST datasets. Xiaotao Jia, Huiyi Gu, Jianlei Yang 0001, Weitao Pan, Youguang Zhang, Sorin Cotofana, Weisheng Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | NAND-SPIN-based processing-in-MRAM architecture for convolutional neural network acceleration
Yinglin Zhao, Jianlei Yang 0001, Bing Li 0017, Xingzhou Cheng, Xucheng Ye, Xiaotao Jia, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 9 |
| 2023 | Rain Rate Retrieval Algorithm for Dual-Polarized Sentinel-1 SAR in Tropical CycloneabstractHeavy rain is associated with strong winds and extreme waves in a tropical cyclone (TC). In this paper, a practical algorithm for rain rate retrieval in TCs is proposed through 24 dual-polarized (vertical-vertical (VV) and vertical-horizontal (VH)) Sentinel-1 (S-1) synthetic aperture radar (SAR) images acquired in interferometric-wide (IW) swath mode, in which 13 images are collocated with the observations from stepped-frequency microwave radiometers (SFMRs). TC winds are directly obtained from VH-polarized images utilizing the geophysical model function (GMF) S-1 IW mode wind speed retrieval model after noise removal (S1IW.NR). The normalized radar cross section (NRCS) at VV-polarization channel is simulated using GMF CMOD5N and VH-polarized SAR wind. It is found that the difference between the simulated NRCSs and measurements from SAR is linearly related to the rain rate and oscillates with the incidence angle. Following this finding, an empirical algorithm for SAR rain rate retrieval is developed, denoted as CRAIN2_S1, which considers the influence of the radius of the maximum wind speed. The proposed algorithm is applied to 11 images in the dataset, and the validation of the rain rate (up to 35 mm/hr) against the products from global precipitation measurements (GPMs) has a root mean square error of 1.74 mm/hr, a correlation coefficient of 0.92 and a scatter index of 0.29. Collectively, it is concluded that the algorithm CRAIN2_S1 can be practically applied for dual-polarized SAR rain rate retrieval without any external information. Weizeng Shao, Yuyi Hu, Zhengzhong Lai, Youguang Zhang, Xingwei Jiang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Layout Aware Optimization Methodology for SOT-MRAM Based on Technically Feasible Top-Pinned Magnetic Tunnel Junction ProcessabstractThe emerging spin-orbit torque magnetic random-access memory (SOT-MRAM) shows promising prospects in high-level cache applications due to its subnanosecond switching speed and high reliability. However, SOT-MRAM faces the issue of large bit-cell layout area, which is currently the focus of attention. Although many design and evaluation works have emerged, the lack of a unified standard for realistic SOT process has hindered the development of relevant research toward practicality. In this article, the bit-cell area of the SOT-MRAM will be evaluated and optimized based on the technically feasible process. First of all, based on the state-of-the-art top-pinned SOT nanopillar process, the SOT-MRAM design rules are proposed. On this basis, this article systematically summarizes four basic device layout modes and provides optimized layout suggestions for conventional SOT bit-cells with different types and sizes of devices. In addition, a series of area-efficient SOT bit-cell designs based on the common area (CA) and dual common (DC) solutions are proposed, which can reduce the layout area of SOT bit-cells by up to 38.4% with reasonable write latency and energy overhead. Chao Wang 0094, Zhaohao Wang, Zhongkui Zhang, Jiagao Feng, Youguang Zhang, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | SpinCIM: spin orbit torque memory for ternary neural networks based on the computing-in-memory architecture
Lichuan Luo, Dijun Liu, He Zhang 0011, Youguang Zhang, Jinyu Bai, Wang Kang 0001 |
CCF Trans. High Perform. Comput. | 4 |
| 2022 | Reconfigurable Bit-Serial Operation Using Toggle SOT-MRAM for High-Performance Computing in Memory ArchitectureabstractComputing in memory (CIM) is a promising candidate for high throughput and energy-efficient data-driven applications, which mitigates the well-known memory bottleneck in Von Neumann architecture. In this paper, we present a reconfigurable bit-serial operation using toggle spin-orbit torque magnetic random access memory (TSOT-MRAM) to perform the computation completely in the bit-cell array instead of in a peripheral circuit. This bit-serial CIM (BSCIM) scheme achieves higher throughput and energy efficiency in CIM. First, basic Boolean logic operations are realized by utilizing the feature of TSOT device. A bit-cell array that implements the bit-serial operation is then built to provide the communication between column and row necessary for arithmetic operations, such as the carry propagation of addition and multiplication. Finally, we analyze the reliability of BSCIM scheme and demonstrate the performance advantage by performing convolution operations for$28\times 28$handwritten digit images in a BSCIM architecture. The results show that the delay and energy of BSCIM architecture are respectively reduced by 1.16-5.49 times and 1.12-1.43 times compared with the existing digital CIM architectures. Besides, its throughput and energy efficiency are also enhanced to 51.2 GOPS and 9.9 TOPS/W respectively. Yining Bai, Zuolei Hao, Guanda Wang, Kun Zhang 0030, Youguang Zhang, Weifeng Lv, Yue Zhang 0010 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2021 | SpinLiM: Spin Orbit Torque Memory for Ternary Neural Networks Based on the Logic-in-Memory ArchitectureabstractLogic-in-memory architecture based on spintronic memories shows fascinating prospects in neural networks (NNs) for its high energy efficiency and good endurance. In this work, we leveraged two magnetic tunnel junctions (MTJs), which are driven by the interplay of field-free spin orbit torque (SOT) and spin transfer torque (STT) effects, to achieve a novel statefullogic-in-memory paradigm for ternary multiplication operations. Based on this paradigm, we further proposed a highly parallel array structure to serve for ternary neural networks (TNNs). Our results demonstrate the advantage of our design in power consumption compared with CPU, GPU and other state-of-the-art works. Lichuan Luo, He Zhang 0011, Jinyu Bai, Youguang Zhang, Wang Kang 0001, Weisheng Zhao 0001 |
DATE | 4 |
| 2021 | Computing-in-Memory Paradigm Based on STT-MRAM with Synergetic Read/Write-Like ModesabstractWith the surge in demand for data storage and processing in emerging applications, the traditional CMOS-based Von-Neumann architecture is facing challenges such as memory wall and static power consumption. In order to conquer the above-mentioned bottlenecks in computing systems, computing in-memory (CiM) architectures based on non-volatile memory (NVM) have been widely researched. In this paper, we propose a CiM paradigm based on spin-transfer torque magnetic random access memory (STT-MRAM), which combines common read-like mode (RLM) and write-like mode (WLM). On the basis of realizing the basic functions AND/OR/NAND/NOR, our design coordinates the high speed of RLM and the integrity of WLM to perform complex operations like full-adder (FA) and XOR/XNOR. In addition, the high speed and low power consumption of the proposed CiM paradigm are established by circuit-level simulation with a 40 nm design kit. Chao Wang 0094, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 3 |
| 2021 | Fully Single Event Double Node Upset Tolerant Design for Magnetic Random Access MemoryabstractBenefitting from its non-volatility, high speed, low power and inherent radiation hardened characteristic, magnetic random access memory (MRAM) has been used in aerospace and avionic electronics. Owing to its high sensing reliability, precharge differential sense amplifier (PCDSA) has been proposed and widely used in MRAM products. However, such PCDSA is based on the conventional CMOS technology and its sensing result is prone to be affected by the single event upset (SEU) and even the single event double node upset (SEDU) when the CMOS technology node shrinks into the nanometer scale. In this paper, we propose a novel PCDSA to tolerate the SEDU, in which the special three-input C-element that behaves as an inverter when its inputs have the same logic value and holds its previous value when its inputs have the different logic values is employed. By using a physics-based STT-MTJ compact model and a commercial CMOS 40 nm design kit, hybrid simulations have been performed to demonstrate its functionality and evaluate its performance. Simulation results show that it can fully tolerate the SEDU when the amount of the deposited charge (Qinj) reaches up to 2 pC. In the worst case where the Qinjis 2 pC, it can achieve a small recover time of 1.3368 ns and low recover energy dissipation of 1.967 pJ with the optimized VDDof 1 V. Deming Zhang, Lang Zeng, You Wang 0002, Bi Wang 0002, Erya Deng, Chuanjie Wang, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 10 |
| 2021 | Time-Domain Computing in Memory Using Spintronics for Energy-Efficient Convolutional Neural NetworkabstractThe data transfer bottleneck in Von Neumann architecture owing to the separation between processor and memory hinders the development of high-performance computing. The computing in memory (CIM) concept is widely considered as a promising solution for overcoming this issue. In this article, we present a time-domain CIM (TD-CIM) scheme using spintronics, which can be applied to construct the energy-efficient convolutional neural network (CNN). Basic Boolean logic operations are implemented through recording the bit-line output at different moments. A multi-addend addition mechanism is then introduced based on the TD-CIM circuit, which can eliminate the cascaded full adders. To further optimize the compatibility of TD-CIM circuit for CNN, we also propose a quantization method that transforms floating-point parameters of pre-trained CNN models into fixed-point parameters. Finally, we build a TD-CIM architecture integrating with a highly reconfigurable array of field-free spin-orbit torque magnetic random access memory (SOT-MRAM) and evaluate its benefits for the quantized CNN. By performing digit recognition with the MNIST dataset, we find that the delay and energy are respectively reduced by 1.22.7 times and 2.4×103-1.1×104times compared with STT-CIM and CRAM based on spintronic memory. Finally, the recognition accuracy can reach 98.65% and 91.11% on MNIST and CIFAR10, respectively. Yue Zhang 0010, Chenyu Lian, Yining Bai, Guanda Wang, Zhizhong Zhang 0004, Zhenyi Zheng, Kun Zhang 0030, Georgios Ch. Sirakoulis, Youguang Zhang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 11 |
| 2020 | An In-memory Highly Reconfigurable Logic Circuit Based on Diode-assisted Enhanced Magnetoresistance DeviceabstractIn the post-Moore era, in order to solve the problem of von Neumann bottleneck and memory wall caused by separation of memory and processor, in-memory-processing (IMP) technique has aroused great attention. Novel non-volatile memory (NVM) based on spintronic devices shows promise for satisfying the needs of low-power consumption and high speed for IMP. However, most spintronic memories based on magnetic tunnel junctions (MTJs) can only implement simple and specific logic functions due to the limits of single device and circuit structure. Otherwise, performing logic functions in memory generates vast dynamic power consumption during frequent reading and writing processes because of the high resistance of miniaturized MTJ. In this paper, we propose an in-memory highly reconfigurable logic circuit based on diode-assisted enhanced magnetoresistance (DEMR) device. Our circuit can realize 16 different logic functions with extremely limited circuit area benefiting from the special structure of DEMR device. With appropriate adjustment of control bit and current, the proposed circuit can further implement complex functions like full adder. The proposed reconfigurable circuit can flexibly meet the performance requirements in different scenarios and will contribute a lot for future in-memory chip design. Yue Zhang 0010, Kun Zhang 0030, Zhizhong Zhang 0004, Youguang Zhang, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2020 | Deep Neural Network accelerator with Spintronic MemoryabstractUtilizing emerging nonvolatile memories to accelerate deep neural network (DNN) has been considered as one of the promising approaches to solve the bottleneck of data transfer during the multiplication and accumulation (MAC). Among them, spintronic memories show tempting prospect due to their low access power, fast access speed, high density, and relatively mature process. As shown in fig.1, according to the principle to achieve DNN computing, it can be mainly divided into three different technical routes. The first one is an "analog" method [1, 2], as shown in fig.1(a). By transforming the digital input signals into multi-level voltage signals, and applying them to different columns of the memory array, the MAC results can be obtained in different columns with current integrator and analog to digital converter (ADC). Besides, the WL drivers can control the pulse width of different rows, to achieve the effect of multi-bit weights. This method can theoretically achieve high energy efficiency and computing speed. However, the variation of magnetic tunnel junction (MTJ) may have influence on the computing accuracy. Besides, the power consumption and area overhead of the ADC are also challenging. The other two methods are in a "digital" way, and they realize MAC computing through row-by-row read/write operation. Fig.1(b) shows the second reading-based method [3]. The weights of the neural network are stored in the memory cell. By putting the input signal to the modified sensing amplifier (SA), it can also achieve XOR function, which is the core of binary NN, with the content stored in the memory cell. Nevertheless, the modification to the SA is usually to add extra transistors in the read path, which will increase the bit error rate. Fig.1(c) shows the diagram of the last one, which is based on the "stateful logic" [4]. The input data is sent to the modified write driver when the WL receiving weight signals from outside I/O. Based on a unique logic paradigm, it can realize XOR function for BNN within 1 or several memory cells during a write cycle. In this talk, we will review the main research status of DNN accelerators based on spintronic memories. Particularly, our recent work on DNN accelerating will be introduced, which can be implemented with different spintronic memories. He Zhang 0011, Wang Kang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | PRISM: Energy-Efficient Polymorphic Operation Based on Spin-Orbit Torque Memory for Reconfigurable ComputingabstractEmerging Non-Volatile Memories (NVMs) including resistive RAM (ReRAM), phase-change memory (PCM), and magnetic RAM (MRAM), have opened up new pathways for the NVM-based reconfigurable computing. Those NVMs technologies can achieve significant energy-efficient computational operations with only minor modification of the peripheral circuits. However, the supported operations are limited by the array structure and low energy-efficiency of implementing the computation using the memory array. In this paper, the Spin Orbit torque-MRAM based polymorphic circuits are proposed to support the reconfigurable computation for reducing the power consumption and improving the functionalities of the single memory array. With the high speed and energy-efficiency write operation, the proposed memory array support both read-out and write-in reconfigurable operations. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001, Jun Zhou 0017 |
ISCAS | 4 |
| 2020 | Efficient Time-Domain In-Memory Computing Based on TST-MRAMabstractIn-memory computing is highly promising to address the processor-memory data transfer bottleneck in current computational paradigm. We firstly propose a timedomain in-memory computing (TIMC) scheme based on highspeed low-power toggle spin torque random access memory (TST-MRAM). The difference of voltage drops of bitline caused by simultaneously-activated bit-cells is reflected to time domain. Reconfigurable logic operations can be performed by utilizing D flip-flops (DFFs) to record the outputs at different moments. In order to demonstrate the advantages of this scheme in terms of speed and energy consumption, an efficient multi-digit addition circuit has been designed and analyzed. Compared with existing IMC schemes, such as spin-transfer torque computing-in-memory (STT-CiM) structure, up to 67% energy saving and 10 times delay improvement can be achieved in the case of four-digit addition by using TIMC scheme. Yue Zhang 0010, Chenyu Lian, Yining Bai, Guanda Wang, Kun Zhang 0030, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 8 |
| 2020 | Computing-in-Memory Architecture Based on Field-Free SOT-MRAM with Self-Reference MethodabstractOn the current computing platforms, the memory wall between processor and memory has become the toughest challenge for the traditional Von-Neumann computer architecture. Computing-in-Memory (CIM) is taken as a promising approach to solving the above bottleneck in computing systems. In this paper, we propose a CIM platform with field-free spinorbit torque magnetic random access memory (SOT-MRAM). The self-reference (SelfRef) method is designed to enhance the read reliability and directly obtain logic results through memory-like read operations without adding logic cells. Memory read/write and logic operations, including NOT, AND/NAND and OR/NOR, can be implemented in the same SOT-MRAM chip. The speed and power penalties caused by SelfRef scheme are acceptable thanks to the ultrafast switching of the SOT. The read reliability and logic correctness of the proposed CIM are demonstrated by hybrid simulation on a 40 nm technology node. Chao Wang 0094, Zhaohao Wang, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 5 |
| 2020 | A Comparative Cross-layer Study on Racetrack Memories: Domain Wall vs SkyrmionabstractRacetrack memory (RM), a new storage scheme in which information flows along a nanotrack, has been considered as a potential candidate for future high-density storage device instead of hard disk drive (HDD). The first RM technology, which was proposed in 2008 by IBM, relies on a train of opposite magnetic domains separated by domain walls (DWs), named DW-RM. After 10 years of intensive research, a variety of fundamental advancements has been achieved; unfortunately, no product has been available until now. With increasing effort and resources dedicated to the development of DW-RM, it is likely that new materials and mechanisms will soon be discovered for practical applications. However, new concepts might also be on the horizon. Recently, an alternative information carrier, magnetic skyrmion, which was experimentally discovered in 2009, has been regarded as a promising replacement of DW for RM, named skyrmion-based RM (SK-RM). Intensive effort has been involved and amazing advances have been made in observing, writing, manipulating, and deleting individual skyrmions. So, what is the relationship between DW and skyrmion? What are the key differences between DW and skyrmion, or between DW-RM and SK-RM? What benefits could SK-RM bring and what challenges need to be addressed before application? In this review article, we intend to answer these questions through a comparative cross-layer study between DW-RM and SK-RM. This work will provide guidelines, especially for circuit and architecture researchers on RM. Wang Kang 0001, Bi Wu 0002, Xing Chen 0012, Daoqian Zhu, Zhaohao Wang, Xichao Zhang, Youguang Zhang, Weisheng Zhao 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 8 |
| 2020 | Evaluating Chinese HY-2B HSCAT Ocean Wind Products Using Buoys and Other ScatterometersabstractThis letter preliminarily assesses the accuracy of ocean wind products from the Ku-band scatterometer (HSCAT) onboard the recently launched Chinese satellite HY-2B. The wind vectors derived from HSCAT during the period from November 15, 2018 to April 30, 2019 are evaluated. The reference wind data include in situ measurements from the offshore meteorological buoys of the National Data Buoy Center, USA, and westerlies mooring of the National Ocean Technology Center, China, and scatterometer winds from the European Advanced Scatterometer (ASCAT) and Indian difference is limited to 25/V2 km, while the HSCAT winds are SCATSAT-1 Oceansat Scatterometer (OSCAT). The spatial temporally collocated with buoys and other scatterometers by less than 0.5 and 1.5 h, respectively. The comparison results show that the HSCAT data have a root-mean-square error (RMSE) of 0.95-1.20 m/s (14.7°-25.7°) regarding the wind speed and direction, respectively, indicating consistency between the HSCAT winds and the reference. Furthermore, better agreement is found for the HY-2B HSCAT winds processed using the Pencil-beam Wind Processor (PWP) algorithm, regarding wind speed (RMSE of 0.95-1.07 m/s) and particularly with respective to the wind direction (RMSE of 14.7°-19.6°), both satisfying the mission specification (<; 2 m/s and <; 20 for wind speed and direction, respectively). The encouraging validation results over the first 5 months demonstrate that the HY-2B HSCAT wind products will be useful for the scientific community. He Wang 0005, Mingsen Lin, Youguang Zhang, Yiting Chang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2019 | CORN: In-Buffer Computing for Binary Neural NetworkabstractBinary Neural Networks (BNNs) have obtained great attention since they reduce memory usage and power consumption as well as achieve a satisfying recognition accuracy on Image Classification. In particular to the computation of BNNs, the multiply-accumulate operations of convolution-layer are replaced with the bit-wise operations (XNOR and pop-count). Such bit-wise operations are well suited for the hardware accelerator such as in-memory computing (IMC). However, an additional digital processing unit (DPU) is required for the pop-count operation, which induces considerable data movement between the Process Engines (PEs) and data buffers reducing the efficiency of the IMC. In this paper, we present a BNN computing accelerator, namely CORN, which consists of a Spin-Orbit-Torque Magnetic RAM (SOT-MRAM) based data buffer to perform the majority operation (to replace the pop-count process) with the SOT-MRAM-based IMC to accelerate the computing of BNNs. CORN can naturally implement the XNOR operation in the NVM memory array, and feed results to the computing data buffer for the majority write operation. Such a design removes the pop-counter implemented by the DPU and reduces data movement between the data buffer and the memory array. Based on the evaluation results, CORN achieves 61% and 14% power saving with 1.74× and 2.12× speedup, compared to the FPGA and DPU based IMC architecture, respectively. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001, Yuan Xie 0001 |
DATE | 4 |
| 2019 | A Skyrmion Racetrack Memory based Computing In-memory Architecture for Binary Neural Convolutional NetworkabstractA Skyrmion Racetrack Memory (SRM) based Computing In-Memory Architecture (SRM-CIM) was proposed in this paper. Both data and computing operation can be achieved in SRM-CIM. SRM-CIM is used to support convolutional computing in Binary Convolutional Neural Network (BCNN). Experimental results show that SRM-CIM achieves 98.7% and 82% energy reduction when compared with RRAM and SOT-MRAM based counterparts. Yinglin Zhao, Shouyi Yin, Youguang Zhang, Shaojun Wei, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2019 | Current Status of the HY-2B Satellite Radar Altimeter and its ProspectabstractThe HY-2B satellite is the second dynamic environment satellite in China. It was successfully launched on October 24th 2018 with a sun-synchronous orbit at an altitude of~970km. Repeat cycles of 14 days are planned for the first two years with oceanographic purpose and 168 days geodetic cycles will follow for the third year of the mission. The satellite is equipped with a Ku/C bands altimeter and the orbit is determined thanks to SLR, GPS and DORIS systems. Yongjun Jia, Mingsen Lin, Youguang Zhang, Wentao An, Xiaoqing Lu |
IGARSS | 3 |
| 2019 | Magnetic Skyrmion-Based Neural Recording System Design for Brain Machine InterfaceabstractNext-generation brain machine interface demand a high-channel-count neural recording system to wirelessly monitor activities of thousands of neurons. In order to achieve high-density neural recording, further development of single recording channel comprised of a neural amplifier front-end (AFE) and an analog-to-digit converter (ADC) is critical. Despite the great progress made in CMOS implementation of custom-designed neural recording system, hybrid limitations of increasing area and power consumption in line with Moore's law drove great demand for post-CMOS substitutes. Magnetic skyrmion with nano particle-like and non-volatile properties are of both fundamental and applied interests for future bio-inspired electronics. In this work, we propose a compact model including both AFE and ADC based on current-induced skyrmion motion. The proposed system achieved a power consumption of 0.63 pJ/channel with an area overhead of 0.14 μm2. The purpose of this work is to explore the feasibility of magnetic skyrmion for building large-scale, dense neuronal recording system which could pave a new way for future brain machine interface application. Biao Pan, Wang Kang 0001, Xing Chen 0012, Jinyu Bai, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 7 |
| 2019 | SR-WTA: Skyrmion Racing Winner-Takes-All Module for Spiking Neural ComputingabstractSpiking neural network (SNN) has emerged as one of the popular architectures in complex pattern recognition and classification tasks. However, hardware implementation of such algorithms using conventional CMOS based neuron consume resources and power that are orders of magnitude higher than that in human brain. This can be attributed to the mismatch of the computational architecture between biological brain and the current Boolean logic computing platform. Magnetic skyrmions have been intensively studied as a prospective information carrier in neuromorphic computing hardware design. In this work, a compact time-domain skyrmion-racing winner-takes-all (SR-WTA) leaky-integrate-fire (LIF) spiking neuron network is presented for the first time. The skyrmion motion dynamics in the LIF neuron and the behaviors of the neuron network was investigated comprehensively. Both SPICE and micromagnetic simulations are performed to evaluate the functionality and performance of the proposed SR-WTA based SNN. Biao Pan, Wang Kang 0001, Xing Chen 0012, Jinyu Bai, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 6 |
| 2019 | Modulation and Demodulation of Digital Frequency Shift Keying System Based on Spin Torque Nano Oscillator with Voltage Controlled Magnetic Anisotropy EffectabstractIn this work, a spin torque nano oscillator (STNO) device whose frequency can be tuned by Voltage Controlled Magnetic Anisotropy effect (VCMA) is proposed. The requirement of magnetic bias field in previous STNO devices is eliminated by the introduction of VCMA effect. Based on VCMA-STNO, a novel architecture is proposed which can compose of a modulation/demodulation digital frequency shift keying (DFSK) communication system. The proposed architecture utilizes VCMA-STNO as core devices and is much simpler comparing with its CMOS counterpart. The proposed VCMA-STNO modulation/demodulation architecture will help to design next generation spintronics DFSK communication system. Lang Zeng, Zuodong Zhang, Haoxuan Chen, Tianqi Gao, Deming Zhang, Mingzhi Long, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 7 |
| 2019 | An STT-MRAM Based in Memory Architecture for Low Power Integral ComputingabstractThe integral histogram image plays an important role in accelerating the feature computation in vision algorithms. However, the computational process of the integral histogram, called integral computation, has high computational complexity and numerous memory access operations, which limit its wide application. This brief proposes an in-memory computational architecture based on Spin Transfer Torque Magnetic Random Access Memory (STT-MRAM) to solve these problems. The architecture can work in two different modes depending on the requirements: the integral computation mode and the memory mode. The architecture can figure out the integral histogram when in the integral computation mode, and just store the data directly when in the memory mode. Utilizing the non-volatile, high density and low power characteristics of STT-MRAM, we integrate the computational units into the memory array to achieve parallel computation. Reduced number of data transmission between storage units and computation units contributes to cut down the latency and energy consumption. The evaluation results show that, comparing with the state-of-the-art work, our architecture provides$1.1\times \sim 9\times$performance improvements and reduces 87.4$\sim$97.3 percent energy consumption for$64\times 64\sim 512\times 512$size images, just with a 8 percent area overhead. Yinglin Zhao, Wang Kang 0001, Shouyi Yin, Youguang Zhang, Shaojun Wei, Weisheng Zhao 0001 |
IEEE Trans. Computers | 5 |
| 2019 | DASM: Data-Streaming-Based Computing in Nonvolatile Memory Architecture for Embedded SystemabstractEmerging nonvolatile memories (NVMs), including resistive RAM (RRAM), phase-change memory (PCM), and magnetic RAM (MRAM), have opened up new pathways for Computing-In-Memory (CIM). Those NVM technologies can achieve energy-efficient computational operations with only minor modification of the peripheral circuits. Despite many advantages provided by computational NVMs, parallelism is not sufficiently explored in such CIM designs. To break through this limitation on performance gain, we propose a data-streaming design for the NVM-based CIM (e.g., DASM) by leveraging the underlying parallelism in the hardware. DASM benefits from the massive parallelism of data-streaming computing, reduction in data movement of the CIM, and the nonvolatility of memory arrays. Specifically, data streaming operations can be implemented with CIM bitwise operations in both read-out and write-in procedures. In addition, we use the multilevel power gating for the memory array and connections to further boost the performance. Finally, we study a case of inference process for the quantized deep-neural-network-based on the DASM design. DASM architecture achieves 47.8×, 5.1×, 2.1× speedup compared to the NVIDIA Jetson TK1 embedded GPU board, Intel Xeon E5-2640 CPU, the state-of-the-art field-programmable gate array (FPGA) design, with much lower power consumption. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Yufei Ding 0001, Weisheng Zhao 0001, Yuan Xie 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | PXNOR-BNN: In/With Spin-Orbit Torque MRAM Preset-XNOR Operation-Based Binary Neural NetworksabstractConvolution neural networks (CNNs) have demonstrated superior capability in computer vision, speech recognition, autonomous driving, and so forth, which are opening up an artificial intelligence (AI) era. However, conventional CNNs require significant matrix computation and memory usage leading to power and memory issues for mobile deployment and embedded chips. On the algorithm side, the emerging binary neural networks (BNNs) promise portable intelligence by replacing the costly massive floating-point compute-andaccumulate operations with lightweight bit-wise XNOR and popcount operations. On the hardware side, the computingin-memory (CIM) architectures developed by the non-volatile memory (NVM) present outstanding performance regarding high speed and good power efficiency. In this paper, we propose an NVM-based CIM architecture employing a Preset-XNOR operation in/with the spin-orbit torque magnetic random access memory (SOT-MRAM) to accelerate the computation of BNNs (PXNOR-BNN). PXNOR-BNN performs the XNOR operation of BNNs inside the computing-buffer array with only slight modifications of the peripheral circuits. Based on the layer evaluation results, PXNOR-BNN can achieve similar performance compared with the read-based SOT-MRAM counterpart. Finally, the end-to-end estimation demonstrates 12.3× speedup compared with the baseline with 96.6-image/s/W throughput efficiency. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Yuan Xie 0001, Weisheng Zhao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Magnetic skyrmions for future potential memory and logic applications: Alternative information carriersabstractMagnetic skyrmions are swirling topological configurations, which are mostly induced by chiral interactions between atomic spins in non-centrosymmetric magnetic bulks or in thin films with broken inversion symmetry. They hold promise as information carriers in future ultra-dense, low-power memory and logic devices owing to the nanocale size and extremely low spin-polarized currents needed to move them. To date, an intense research effort has led to the identification, creation/annihilation, motion and manipulation of skyrmions at room temperature. Meanwhile, a rich variety of skyrmion-based device concepts and prototypes have been proposed, indicating the considerable potential of magnetic skyrmions in future electronic applications. However, current studies mainly focus on physical or principle investigations, whereas the electrical design methodology, implementation and evaluations are still lacking. In this paper, we will bring the readers in the “design, automation and test (DAT) society” the current status and outlook of skyrmions in relation to future potential racetrack memory and neuromorphic computing applications. Most importantly, we also want to evoke the effort from the DAT society to address the challenges, e.g., all-electrical manipulation of skyrmions at room temperature, for the research and development of practical skyrmion-based electronics. Wang Kang 0001, Xing Chen 0012, Daoqian Zhu, Yangqi Huang, Youguang Zhang, Weisheng Zhao 0001 |
DATE | 6 |
| 2018 | Design Space Exploration of Magnetic Tunnel Junction based Stochastic Computing in Deep LearningabstractMagnetic tunnel junction (MTJ) is considered as a promising memory candidate in the more than Moore era because of high power efficiency, fast access speed, nearly infinite endurance and easy 3D integration. The nondeterministic switching behavior has been profited to exploit new directions for computing methods, such as stochastic computing. In this paper, the application of stochastic switching behavior in stochastic computing is explored for deep neural network (DNN). Stochastic computing method features low logic complexity, low energy consumption and fine-grained parallelism, boosting the performance of DNN system by combining MTJ. As a key block of stochastic computing, MTJ based true random number generator design is presented in details. The functionality has been validated by combining the hardware design and post-processing in software. Simulation results are demonstrated visibly by handwritten digits recognition test to show the accuracy. Furthermore, the performance is investigated in terms of accuracy, energy consumption and memory occupation to find more efficient techniques. You Wang 0002, Yue Zhang 0010, Youguang Zhang, Weisheng Zhao 0001, Hao Cai 0001, Lirida A. B. Naviner |
ACM Great Lakes Symposium on VLSI | 3 |
| 2018 | The Wind Speed Inversion and In-Orbit Assessment of Imaging Altimeter on Tiangong-2 Space StationabstractImaging ALTimeter (IALT) is a new type of radar altimeter system, which observes the earth from 2° to 7° incident angles. In comparison to the conventional altimeters such as HY-2A altimeter, Jason-1/2, TOPEX/Poseidon, which observe the ocean at nadir, the swath of IALT is much wider and its spatial resolution is much higher. The IALT on board Tiangong-2 space station is launched on 15th September, 2016 at Jiuquan Satellite Launch Center. The in-orbit assessment of IALT is done until 30th April, 2017. In this paper, the ocean surface wind speed inversion method based on IALT is established. The neural network algorithm is used for ocean surface wind speed retrieval, and the spatial resolution of retrieved wind speed is 25km. The wind speed inversion accuracy is evaluated by comparing with the ECMWF reanalysis wind speed, buoy wind speed, and boat measurement wind speed. The results show that the Root-Mean-Square (RMS) of retrieved wind speed is 1.85m/s, and the Bias of retrieved wind speed is about -0.21m/s. The wind speed inversion accuracy satisfies performance requirement. Qingliu Bao, Xiaobin Yin, Juhong Zou, Mingsen Lin, Youguang Zhang |
IGARSS | 5 |
| 2018 | The Simulation of Ocean Surface Wind Measured by Polarimetric ScatterometerabstractOcean surface wind field is a very important marine dynamic parameter in marine environment forecasting and climatological studies. Spaceborne scatterometer is one of the most efficient remote sensors than can provide global ocean surface wind measurement. Polarimetric scatterometer (PolScat) simultaneously measures co-polarized and cross-polarized backscattering coefficient and the correlation coefficient of the co- and cross-polarized component of radar echoes which can significantly improve the performance of the sea surface wind field measurements. In this paper, we derive the error model of correlation scattering coefficient from radar echo signals. The effect of antenna polaxis deflection on correlation scattering coefficient error is analyzed. Moreover, an “end-to-end” system simulation model of PolScat is established. A modified Maximum Likelihood Estimation (MLE) is used for ocean surface wind field inversion. Both the co-polarized backscattering coefficient and correlation scattering coefficient are included in the objective function of MLE. The simulation results show that PolScat can effectively reduce the probability of ambiguous solutions. The wind speed and wind direction inversion accuracy of PolScat is better than 1m/s and 15° respectively. Juhang Zau, Shuyan Lang, Yarang Zau, Mingsen Lin, Youguang Zhang, Xiaobin Yin, Qingliu Bao |
IGARSS | 5 |
| 2018 | NEAR: A Novel Energy Aware Replacement Policy for STT-MRAM LLCsabstractAs the technology node shrinks, leakage power becomes a bottleneck for processor performance and memory capacity scalings. Spin Torque Transfer Magnetic Random Access Memory (STT-MRAM) has negligible leakage power, fast access speed, high integration density and non-volatility. Therefore, it is a promising candidate for the last level cache design. However, it suffers from high write energy and slow write speed. In the paper, we observe that the traditional cache replacement policy is not optimal when applied to STT-MRAM from the energy consumption perspective. So we propose a novel write energy aware cache replacement policy, which utilizes a MinHash function to identify the similarities between the cache line to be written back and candidates for the replacement. The cache line with the highest similarity is chosen as the victim. In addition, we propose a new metric for cache replacement considering both performance and write energy to improve the replacement policy further. The experimental results show that our proposed policy can reduce write energy by 33.6% on average compared to the state-of-the-art Least Recently Used (LRU) replacement policy with only 0.5% performance penalty and negligible hardware overhead. Yuanqing Cheng, Ying Wang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 4 |
| 2018 | Progresses and challenges of spin orbit torque driven magnetization switching and application (Invited)abstractSpin orbit torque (SOT) has been proposed as a potential alternative mechanism to the conventional spin transfer torque (STT) for the magnetization switching. Recently, theoretical and experimental works revealed the novel factors influencing the SOT-driven magnetization switching. Emerging SOT-based spintronics memories and circuits were explored to implement fast and energy-efficient write operation. However, the perspective of the SOT mechanism is still challenged by some serious shortcomings, such as area penalty, relatively large switching current density and undesirable use of external magnetic field. Here, we review the progresses in the SOT mechanism involving the magnetization dynamics, device design and circuit development. Key issues to be addressed in optimizing the SOT devices are pointed out. In particular, we discuss the potential solutions to develop high-density SOT-based memories and circuits. Zhaohao Wang, Zuwei Li, Liang Chang 0002, Wang Kang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 7 |
| 2018 | Radiation hardening design for spin-orbit torque magnetic random access memoryabstractAlthough the magnetic tunnel junction (MTJ) is intrinsically immune to radiation, the read/write operations of magnetic random access memory (MRAM) may be vulnerable to radiation-induced current. In this paper, we investigate the radiation hardening design for spin orbit torque based MRAM (SOT-MRAM). The hardening technique is firstly studied at the device level by optimizing the dimension and magnetic parameters. Then we propose radiation hardening read and write circuits addressing the influence of single event upset (SEU). Based on a physics-based SOT-MTJ compact model and a 65nm CMOS design kit, simulation results show that the proposed MOS-stacked read sensing amplifier and write circuits of six PMOS transistors as a feed-back structure to charge/discharge sensitive nodes can correct soft errors. Bi Wang 0002, Zhaohao Wang, Kaihua Cao, Youguang Zhang, Yuanfu Zhao, Weisheng Zhao 0001 |
ISCAS | 4 |
| 2017 | Voltage-controlled MRAM for working memory: Perspectives and challengesabstractMagnetic random access memory (MRAM) has been widely studied for future nonvolatile working memory candidate. However, the mainstream current (spin transfer torque, STT or spin Hall effect, SHE) driven MRAMs (STT-MRAM or SHE-MRAM) face intrinsic problems in terms of high write power and long latency, significantly limiting the applications for low-power and high-speed working memories. The recently-developed new-generation MRAM, named VCMA-MRAM, which exploits the voltage-controlled magnetic anisotropy (VCMA) effect to write (or assist to write) data information into magnetic tunnel junctions (MTJs), holds the promise to efficiently overcome these problems. Despite the impressive possibility of improving write power and speed, this technology, however, is currently under intensive research and development (R&D), and some challenges still await answers. In this paper, we investigate the perspectives and challenges of VCMA-MRAM for working memories from a cross-layer (device/circuit/architecture) design point of view. We demonstrate that VCMA-MRAM outperforms STT-MRAM and SHE-MRAM in terms of area, speed, energy consumption and instruction-per-cycle (IPC) performance, benefiting from the low-power and high-speed VCMA-driven data writing mechanism. On the other hand, challenges in terms of device fabrication and circuit design should be efficiently addressed before practical applications. Wang Kang 0001, Liang Chang 0002, Youguang Zhang, Weisheng Zhao 0001 |
DATE | 3 |
| 2017 | Advanced Low Power Spintronic Memories beyond STT-MRAMabstractUntil now, spin transfer torque magnetic random access memory (STT-MRAM) has drawn considerable R&D interest worldwide. A number of companies and universities are currently involved in this promising technology. In 2016, Everspin released the first 256M STT-MRAM chip, indicating the commercialization and application of STT-MRAM. Nevertheless, STT-MRAM still has some intrinsic limitations, such as dynamic write power and speed, compared with CMOS-based memory technologies. Following the technical evolution process from toggle-MRAM to STT-MRAM, the continuous pursuit of high performance, high density, low power and scalability, drives the intensive R&D of new memory technologies. In this paper, we will show the recent progress in advanced spintronic memories beyond STT-MRAM, such as the spin Hall effect (SHE)-driven and voltage-driven MRAMs. These advanced MRAM technologies do have some unique advantages compared with STT-MRAM, but they also suffer from new design and fabrication challenges. In addition, we will present the latest research in emerging spintronic devices, e.g., magnetic skyrmions, which are potential as information carriers in future spintronic memories, e.g., racetrack memory. Wang Kang 0001, Zhaohao Wang, He Zhang 0011, Youguang Zhang, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2017 | PRESCOTT: Preset-based cross-point architecture for spin-orbit-torque magnetic random access memoryabstractDue to nearly zero leakage power consumption, non-volatile magnetoresistive random access memory (MRAM) is becoming one of the promising candidates for replacing conventional volatile memories (e.g. SRAM and DRAM). In particular, emerging spin-orbit torque (SOT) MRAM is considered to outperform spin-transfer torque (STT) MRAM due to its fast switching, separate read/write paths, and lower energy dissipation. However, the SOT-MRAM technology is still in its infancy; one key design challenge is that the control of SOT-MRAM, which involves three terminals, is more complicated compared with STT-MRAM. In this paper, we propose a novel MRAM write scheme called PRESCOTT1, where the “1” and “0” data values can be written into memory cells through the SOT and STT, respectively. As a result, the write current is unidirectional rather than bi-directional, which addresses the control complexity. Using this unidirectional write scheme, we design a PreSET-based cross-point (CP) MRAM to improve programing speed, write energy dissipation and storage density compared to conventional MRAM. Circuit simulation results demonstrate that our PreSET-based CP MRAM can achieve around 67.14% average write energy reduction and 50.86% improvement in programming speed, compared with CP STT-MRAM. Liang Chang 0002, Zhaohao Wang, Alvin Oliver Glova, Jishen Zhao, Youguang Zhang, Yuan Xie 0001, Weisheng Zhao 0001 |
ICCAD | 5 |
| 2017 | Thermosiphon: A thermal aware NUCA architecture for write energy reduction of the STT-MRAM based LLCsabstractAs the speed gap of the modern processor and the off-chip main memory enlarges, on-chip cache capacity increases to sustain the performance scaling. As a result, the cache power occupies a large portion of the total power budget. STT-MRAM (Spin Transfer Torque Magnetic Memory) is proposed as a promising solution for the low power cache design due to its high integration density and ultra-low leakage. Nevertheless, the high write power and latency of STT-MRAM become new barriers for the commercialization of this emerging technology. In this paper, we investigate the thermal effect on the access performance of STT-MRAM and observe that the temperature can affect the write delay and energy significantly. Then, we explore the NUCA (Non-Uniform Cache Access) design of the CMPs (Chip-Multi-Processors)with STT-MRAM based LLC (Last Level Cache). A thermal aware data migration policy, called “Thermosiphon”, which takes advantage of the thermal property of STT-MRAM, is proposed to reduce the LLC write energy. This policy splits the LLC into different regions based on the thermal distribution and adaptively migrate write intensive data considering the temperature gradient among different thermal regions. Compared to the conventional NUCA design, our proposed design can save 22.5% write energy with negligible hardware overhead. Bi Wu 0002, Yuanqing Cheng, Pengcheng Dai, Jianlei Yang 0001, Youguang Zhang, Dijun Liu, Ying Wang 0001, Weisheng Zhao 0001 |
ICCAD | 5 |
| 2017 | Programmable Stateful In-Memory Computing Paradigm via a Single Resistive DeviceabstractData transfer bandwidth and the related energy consumption has become two of the most critical bottlenecks in conventional von-Newman architecture, owing to the separation of the processor and memory units and the performance mismatch between the two. Realization of the unity of logic computing and data storage in the same die has opened up a promising research direction of in-memory computing (IMC). Meanwhile nonvolatile memory (NVM) based programmable (or reconfigurable) logic architecture has always been a hot topic in the circuit and system societies. To date, lots of interest has been attracted and amazing advance has been made in the two fields, yet none can fully exploit the advantages of both. This paper takes a major step forward by introducing a novel nonvolatile programmable stateful IMC architecture via a single resistive device, which is a completely different design paradigm from previous studies. Each memory cell can perform different Boolean logic functions by dynamically programming the input signals. The computing output result is insitu stored in the memory cell itself and can be readout with a memory-like operation. We will first give a brief review on this filed and then introduce our recent work. We will illustrate how the programmable stateful IMC operations can be implemented via a single resistive device and how the logic computing and data storage can be united within a memory chip. Wang Kang 0001, He Zhang 0011, Youguang Zhang, Weisheng Zhao 0001 |
ICCD | 4 |
| 2017 | The ocean surface current inversion mehtod of Doppler scatterometerabstractOcean surface current is a very important parameter of ocean dynamic environment. The observation and prediction of ocean surface current has attracted more and more concern. Doppler Scatterometer (DopScat) is a new type of radar for ocean surface wind and current field remote sensing. The ocean surface current inversion method of DopScat impacts the measurement accuracy directly. In this paper, we establishes the Maximum Likelihood Estimation(MLE) method to retrieve the ocean surface current and wind simultaneously. The retrieval accuracy for different position in cross-track, wind speed, and current speed are analyzed. The retrieval results show that the RMS of inversion current speed and direction can be smaller than 0.18m/s and 25°respectively, for medium wind speed condition. This is submitted for the special session of “New Developments of Chinese Oceanographic and Meteorological Satellites”. Qingliu Bao, Mingsen Lin, Youguang Zhang, Xiaolong Dong, Shuyan Lang, Peng Gong 0002 |
IGARSS | 3 |
| 2017 | The error transfer of Doppler spectrum model in ocean surface current direct inversionabstractMicrowave remote sensing is one of the most useful methods for observing the ocean parameters. The Doppler frequency of the radar echoes can be used for ocean surface current speed retrieval. While the effect of the ocean currents and waves are interactional. In this paper the suitable ocean wave elevation spectrum and directional distribution function are selected by comparing the ocean Doppler spectrum in C band with the empirical geophysical model function (CDOP). The simulation results show that the ocean surface current speed error is sensitive to the wind speed and wind direction error. With VV polarization, the ocean surface current speed error is about 0.15 m/s when the wind speed error is 2 m/s, and the ocean surface current speed error is smaller than 0.3 m/s when the wind direction error is within 20° in the cross wind direction. This is submitted for the special session of “New Developments of Chinese Oceanographic and Meteorological Satellites”. Mingsen Lin, Youguang Zhang, Qingliu Bao, Peng Gong 0002 |
IGARSS | 2 |
| 2017 | Ocean Surface Current Inversion Method for a Doppler ScatterometerabstractThe ocean surface current is a very important parameter of ocean dynamic environment. It is connected to global climate change, marine environment forecasting, marine navigation, engineering security, and so on. The observation and prediction of ocean surface current have attracted more and more concern. Doppler Scatterometer (DopScat) is a new type of radar for ocean surface wind and current field remote sensing. The ocean surface current inversion method of DopScat impacts the measurement accuracy directly. In this paper, we establish the simulation model of a DopScat and provide the radial velocity error model. The numerical ocean surface Doppler spectrum model is also introduced and validated with the empirical geophysical model function in C-band (CDOP). The suitable ocean wave elevation spectrum and directional distribution function are selected. What is more, this paper establishes the maximum likelihood estimation (MLE) method to retrieve the ocean surface current and wind simultaneously. The retrieval accuracy for different positions in cross track, different wind speeds, and different current speeds are analyzed. At last, the global ocean current field is observed by DopScat and the ocean current is retrieved. In our simulation, the orbit parameters and observation geometry of DopScat are the same as that of HY-2A scatterometer. The retrieval results show that global current speed standard deviation can be smaller than 0.18 m/s for five days and$0.5 {^{\circ }} \times 0.5 {^{\circ }}$grid average. Qingliu Bao, Mingsen Lin, Youguang Zhang, Xiaolong Dong, Shuyan Lang, Peng Gong 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | PDS: pseudo-differential sensing scheme for STT-MRAMabstractSTT-MRAM has been considered as one of the most promising nonvolatile memory candidates in the next-generation of computer architecture. However, the read reliability and dynamic write power concerns greatly hinder its practical application. In this paper, we propose a synergistic solution, namely pseudo-differential sensing (PDS), to jointly address these two concerns. Three techniques, including cell cluster, asymmetric sensing amplifier (ASA) and self-error-detection-correction (SEDC), are proposed to implement the PDS concept. Our experimental results show that the PDS scheme with the 3T3MTJ cell cluster can reduce the area (~21.7%) and write power (~25.6%) of the differential sensing (DS) scheme while improve the read reliability (read margin, ~35.6%) of the typical sensing (TS) scheme for a 16 Mbit cache. Furthermore, the PDS scheme with the 1T3MTJ cell cluster can outperform both the TS and DS schemes in terms of area (~40.0%, ~66.1%), read latency (~16.6%, ~32.1%), read power (~16.7%, ~37.1%), write latency (~5.4%, 16.3%) and write power (~18.6%, ~43.4%). Wang Kang 0001, Tingting Pang, Bi Wu 0002, Weifeng Lv, Youguang Zhang, Guangyu Sun 0003, Weisheng Zhao 0001 |
DAC | 5 |
| 2016 | Sea surface wind speed inversion using low incident NRCSabstractAs the launch of radars, such as the Precipitation Radar (PR) on Tropical Rainfall Measuring Mission (TRMM)[1] satellite and the Surface Wave Investigation and Monitoring (SWIM) on China France Oceanography SATellite (CFOSAT)[2], that operate in low incident angles, more and more NRCS data at low incident angle will be obtained. In order to retrieve the sea surface wind speed using low incident angle NRCS, the empirical GMF of NRCS with wind speed is established. The empirical nadir reflection coefficient |R(0)2| are calculated and the empirical relationship between mean square slop s(u) and wind speed is established. The mean square slop s(u) can be retrieved by fitting the NRCS at certain incident angles with the theoretical Gaussian GMF model. Then the wind speeds are calculated using the empirical corresponding relation between mean square slop and wind speed. The retrieved wind speeds are compared with Tao and NDBC buoy. The results show that the standard deviation (STD) and bias of retrieved wind speeds are smaller than 1.7m/s and 0.1m/s respectively. Qingliu Bao, Youguang Zhang, Wentao An, Limin Cui, Shuyan Lang, Mingsen Lin, Peng Gong 0002 |
IGARSS | 2 |
| 2016 | On the improvement of the HY-2A scatterometer wind quality controlabstractThis paper reviews several wind quality-sensitive parameters derived from HY-2A scatterometer data, such as the wind-inversion residual (or Maximum Likelihood Estimator, MLE) and its spatially averaged value, and the singularity exponent (SE) derived from an image processing technique, called singularity analysis. Their sensitivity to data quality is evaluated using the collocated European Centre for Medium-range Weather Forecasting (ECMWF) model output and satellite radiometer rain data. It shows that SE is the best quality indicator, followed by the spatially averaged MLE and the conventional MLE. A set of MLE and SE thresholds are derived from the sensitivity analysis in order to optimize the quality control (QC) for the HY-2A scatterometer. Wenming Lin, Marcos Portabella, Ad Stoffelen, Anton Verhoef, Shuyan Lang, Youguang Zhang, Mingsen Lin |
IGARSS | 6 |
| 2016 | A preliminary in situ calibration for HY-2A satellite altimeterabstractSince the launch, HY-2A has observing global ocean dynamic environment data over the last 5 years including sea surface height (SSH), significant wave height (SWH) and corresponding auxiliary data (wind fields and water vapor). The in-orbit validation and crossover comparison of which suggest that the various technical indicators cater for basic needs in practicality. However, for the most important ocean dynamic environment parameter SSH, the accurate, reliable and standard processing for absolute calibration has not achieved yet, significantly limiting quantitative application of altimetry production. To evaluate the accuracy of HY-2A, we perform absolute calibration at California offshore Harvest Oil Platform. HY-2A interim geophysical dataset records (IGDR) are used to perform comparison. All corrections are estimated independently avoiding instruments aging. According to results, there are two significant linear trends for HY-2A derived SSHs consisting of -0.001083m/day for cycle 1~99 and 0.001475m/day afterward. The results indicate that HY-2A observations are low by -40.5 ±1.2 cm. Yalong Liu, Ke Xu 0012, Youguang Zhang |
IGARSS | 4 |
| 2016 | Data fusion of sea surface height anomaly from HY-2A AND Jason-2abstractAltimeter provides a unique perspective to observe global ocean frequently which is never happened before. Whereas the altimeter measures the ocean in a narrow track, it is inconvenient in studying the ocean structures such as Kuroshio and Gulf Stream, let alone the mesoscale structures which are important ways to transport mass and energy. These ocean structures determined by sea surface topography (SST) can be estimated using altimeter data. In this study, we present a fusion result of altimetry data by combining HY-2A and Jason-2. The merged data, generated using optimal interpolation (OI), present a detailed description for mesoscale structures in Kuroshio area. In this study, (1) altimetry data for both HY-2A and Jason-2 employed to determine sea surface height anomaly (SSHA) are corrected using collected independent source including European Centre for Medium-Range Weather Forecasts (ECMWF) reanalysis data, TPXO ocean tide model etc. (2) in order to further enrich the details of mesoscale structures, altimetry data from other ongoing mission including Cryosate-2 and SARAL/Altika are take into account. (3) to ensure accuracy of the fusion, observation bias of sea surface heights from HY-2A are corrected to align with the reference mission Jason-2 using cross-calibration, since Jason-2 have been calibrated stringently. The relative biases for HY-2A, SARAL/Altika and Cryosate-2 are 37.36cm, -4.54cm and - 72.5cm and the corresponding standard deviations are 5.58cm, 5.26cm and 5.02 cm respectively, which suggest accurate observations and reliable reprocessing. (4) optimal interpolation is used to perform the interpolation, thus generate SSHA with grid format which are easy to distinguish the mesoscale structures. (5) more details are introduced from multi-missions rather than one. Yalong Liu, Ke Xu 0012, Youguang Zhang |
IGARSS | 3 |
| 2016 | Spin wave based synapse and neuron for ultra low power neuromorphic computation systemabstractIn this work, we have proposed that the neural synapses and neurons can be realized by utilizing spin waves (SWs) as information carrier. The SWs is excited by spin torque nano-oscillator (STNO), and detected with several different physical mechanisms: 1) tunneling magnetic-resistance 2) spin pumping and 3) inverse spin hall effect. The proposed SWs based synapses and neurons can be further combined together to form a neuromorphic computation system with crossbar structure. Possible ultra low power consumption and ultra high speed are the advantage of our proposed SWs based synapses and neurons. Lang Zeng, Deming Zhang, Youguang Zhang, Fanghui Gong, Tianqi Gao, Sa Tu, Haiming Yu, Weisheng Zhao 0001 |
ISCAS | 3 |
| 2016 | Quantitative evaluation of reliability and performance for STT-MRAMabstractDue to its non-volatility, high access speed, ultra low power consumption and unlimited writing/reading cycles, STT-MRAM (Spin Transfer Torque Magnetic Random Access Memory) has emerged as the most promising candidate for the next generation universal memory. However, the process of commercialization of STT-MRAM is hampered by its poor reliability. Generally, these reliability issues are caused by the PVT (Process Variations, Voltage, and Temperature) of both MTJ (Magnetic Tunneling Junction) and transistor. Mitigation and alleviating the impacts of the intrinsic properties and PVT on STT-MRAM is a challenging work. This paper discusses the errors occurring in STT-MRAM resulting from its poor reliability, and analyzes the causes of such errors. To obtain a quantitative assessment of PVT impact on STT-MRAM reliability, we investigate three aspects: writing/reading operation error rate, power consumption and access delay of a single cell. This study is carried out on Cadence platform for 45 nm technology node and the PMA (Perpendicular Magnetic Anisotropy) MTJ model used in the investigation comes from SP INLIB. These quantitative information would be helpful for designing reliability enhancing strategies of STT-MRAM. Liuyang Zhang, Aida Todri, Wang Kang 0001, Youguang Zhang, Lionel Torres, Yuanqing Cheng, Weisheng Zhao 0001 |
ISCAS | 4 |
| 2016 | Read disturbance issue and design techniques for nanoscale STT-MRAM
Yi Ran, Wang Kang 0001, Youguang Zhang, Jacques-Olivier Klein, Weisheng Zhao 0001 |
J. Syst. Archit. | 3 |
| 2016 | Application-Level Scheduling With Probabilistic Deadline ConstraintsabstractOpportunistic scheduling of delay-tolerant traffic has been shown to substantially improve spectrum efficiency. To encourage users to adopt delay-tolerant scheduling for capacity -improvement, it is critical to provide guarantees in terms of completion time. In this paper, we study application-level scheduling with deadline constraints, where the deadline is pre-specified by users/applications and is associated with a deadline violation probability. To address the exponentially-high complexity due to temporally-varying channel conditions and deadline constraints, we develop a novel asymptotic approach that exploits the largeness of the network to our advantage. Specifically, we identify a lower bound on the deadline violation probability, and propose simple policies that achieve the lower bound in the large-system regime. The results in this paper thus provide a rigorous analytical framework to develop and analyze policies for application-level scheduling under very general settings of channel models and deadline requirements. Further, based on the asymptotic approach , we propose the notion of Application-Level Effective Capacity region, i.e., the throughput region that can be supported subject to deadline constraints, which allows us to quantify the potential gain of application-level scheduling. Simulation results show that application-level scheduling can improve the system capacity significantly while guaranteeing the deadline constraints. Huasen Wu, Xiaojun Lin 0001, Xin Liu 0002, Youguang Zhang |
IEEE/ACM Trans. Netw. | 4 |
| 2015 | The latest assessment for the reprocessed GDR product of HY-2A altimeterabstractThis paper is aimed to assess the accuracy of HY-2A altimetry system. To further improve the accuracy and performance of HY-2A observed SSHs, several new treatments including 4 parameters maximum likelihood estimation (MLE4) retracking for Ku and C band, non-parameter sea state bias (NPSSB) model, and reprocessed dual-frequency altimeter ionospheric correction are included. The evaluation from dual-crossover comparison of the fully reprocessed level-2 sensor geophysical dataset records (SGDRs) data suggests that the algorithm and model improvements mentioned above give rise to remarkable promotion to the accuracy of the forthcoming version GDRs. In this study we conclude that the standard deviation of 5.79 cm is achieved from crossover comparison between HY-2A and Jason-2 sea surface height (SSHs), suggesting a promising situation of HY-2A altimetry datasets. Yalong Liu, Ke Xu 0012, Youguang Zhang, Xi-Yu Xu |
IGARSS | 3 |
| 2015 | Energy-efficient neuromorphic computation based on compound spin synapse with stochastic learningabstractRecently, magnetic tunnel junction with in-plane magnetization (i-MTJ) has been exploited to behave as a binary stochastic synapse. However, it suffers from its limited level of synaptic weight, resulting in an inaccurate learning. In this work, a compound synapse that employs multiple perpendicular MTJs (p-MTJs) in series is proposed. It possesses an analog-like synaptic weight under weak programming conditions, which leads to a stochastic learning rule and low power consumption per synaptic event. By performing system-level simulations on the MNIST database, it has been demonstrated that such compound spin synapses can realize stochastic neuromorphic computation with high accuracy and low energy consumption. Deming Zhang, Lang Zeng, Yuanzhuo Qu, Youguang Zhang, Mengxing Wang 0001, Weisheng Zhao 0001, Tianqi Tang 0001, Yu Wang 0002 |
ISCAS | 4 |
| 2015 | Spintronics: Emerging Ultra-Low-Power Circuits and Systems beyond MOS TechnologyabstractConventional MOS integrated circuits and systems suffer serve power and scalability challenges as technology nodes scale into ultra-deep-micron technology nodes (e.g., below 40nm). Both static and dynamic power dissipations are increasing, caused mainly by the intrinsic leakage currents and large data traffic. Alternative approaches beyond charge-only-based electronics, and in particular, spin-based devices, show promising potential to overcome these issues by adding the spin freedom of electrons to electronic circuits. Spintronics provides data non-volatility, fast data access, and low-power operation, and has now become a hot topic in both academia and industry for achieving ultra-low-power circuits and systems. The ITRS report on emerging research devices identified themagnetic tunnel junction(MTJ) nanopillar (one of the Spintronics nanodevices) as one of the most promising technologies to be part of future micro-electronic circuits. In this review we will give an overview of the status and prospects of spin-based devices and circuits that are currently under intense investigation and development across the world, and address particularly their merits and challenges for practical applications. We will also show that, with a rapid development of Spintronics, some novel computing architectures and paradigms beyond classic Von-Neumann architecture have recently been emerging for next-generation ultra-low-power circuits and systems. Wang Kang 0001, Yue Zhang 0010, Zhaohao Wang, Jacques-Olivier Klein, Claude Chappert, Dafine Ravelosona, Gefei Wang, Youguang Zhang, Weisheng Zhao 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 8 |
| 2014 | An overview of spin-based integrated circuitsabstractConventional CMOS integrated circuits suffer from serve power and scalability challenges as technology node scales into ultra-deep-micron technology nodes. Alternative approaches beyond charge-only based circuits. In particular, spin-based devices or integrated circuits show promising merits to overcome these issues by adding the spin freedom of electrons to the electronic circuits. Spintronics has now become a hot topic in both academics and industrials. This paper overviews the status and prospects of spin-based integrated circuits under intense investigation and address particularly their merits and challenges for practical applications. Wang Kang 0001, Weisheng Zhao 0001, Zhaohao Wang, Jacques-Olivier Klein, Yue Zhang 0010, Djaafar Chabi, Youguang Zhang, Dafine Ravelosona, Claude Chappert |
ASP-DAC | 7 |
| 2014 | Spintronics for low-power computingabstractMicroelectronics has been following Moore's law for almost 40 years. However this trend tends to run out of steam in recent technology nodes. The continuous improvements in the size of the transistors and in the operating frequencies result in serious power consumption, heat dissipation and reliability issues. Spintronics (Nobel Prize of Physics, 2007 awarded to Prof. Fert from Univ. Paris-Sud and Peter Grünberg from Forschungszentrum Jülich) nanodevices can reduce significantly the power, improve the reliability or allow new functionalities. The 2010 ITRS report on emerging research devices identified Magnetic Tunnel Junction (MTJ) nanopillar (the preeminent spintronics nanodevice) as one of the most promising technologies to be part of the future microelectronics circuits. It provides data non-volatility, hardness to radiations, fast data access and low-power operations. Magnetic memories become the most promising candidate for both low power logic computing and the data storage. This tutorial paper presents multi-discipline questions (Device, Circuit, Architecture, System and CAD) related to this topic to share the most recent results and discuss the future challenges. Yue Zhang 0010, Weisheng Zhao 0001, Jacques-Olivier Klein, Wang Kang 0001, Damien Querlioz, Youguang Zhang, Dafine Ravelosona, Claude Chappert |
DATE | 6 |
| 2014 | Current status of the HY-2A satellite radar altimeter and its prospectabstractHY-2 satellite was successfully launched on 16 August, 2011. It carried three main microwave instruments into space for operationally observing dynamic ocean environment parameters on a global scale. HY-2 satellite altimeter provides sea surface height, significant wave height, sea surface wind speed. Current status of HY-2 satellite altimeter is put forward in this study. By comparison with the other satellite data, NDBC data and other data, the accuracy of the HY-2's data products is evaluated in this work. Yongjun Jia, Mingsen Lin, Youguang Zhang |
IGARSS | 3 |
| 2014 | A preliminary crossover calibration result for HY-2abstractThis paper presents a result of crossover analysis between HY-2 and Jason-2 mission over ocean. The major objectives of this paper are to assess HY-2 IGDR (Interim Geophysical Data Record) derived SSHs by comparing HY-2 measurements with the reference mission Jason-2, and further to illustrate the potential of HY-2 data in monitoring global sea level variability. The instrument-independent models and data are applied for both HY-2 and Jason-2 to correct the errors including range delays and geophysical effects, thus provide an objective assessment. All the results indicate a good performance of HY-2 measurements. A standard deviation of 6.6 cm and a bias of 0.26 cm between HY-2 and Jason-2 SSHs (Sea Surface Height) is acquired from 50°S to 50°N which is close to the Jason-1 and Jason-2 performance. The results suggest that a promising situation in terms of HY-2 observations. Yalong Liu, Youguang Zhang, Mingsen Lin, Junwu Tang |
IGARSS | 2 |
| 2014 | Ferroelectric tunnel memristor-based neuromorphic network with 1T1R crossbar architectureabstractEmerging ferroelectric tunnel memristors show large OFF/ON resistance ratio (>100) and high operation speed (~10ns), promising to be widely applied in the future synapse-like systems. In this paper we propose a neuromorphic network with ferroelectric tunnel memristor. This network is arranged with classical crossbar topology, in which each crosspoint forms a synapse consisting of a MOS transistor and a memristor. Based on this architecture, we design a spike-timing dependent plasticity (STDP) scheme and a parallel supervised learning circuit. Using a compact model of ferroelectric tunnel memristor and CMOS 40nm design kit, we perform transient simulation to validate the functionality of the proposed STDP and learning circuit. Simulation results show the potential of our neuromorphic network in low power (~100nA or ~1μA) and high speed (μs or ~100ns) computing system. Zhaohao Wang, Weisheng Zhao 0001, Wang Kang 0001, Youguang Zhang, Jacques-Olivier Klein, Claude Chappert |
IJCNN | 4 |
| 2014 | Application-level scheduling with deadline constraintsabstractOpportunistic scheduling of delay-tolerant traffic has been shown to substantially improve spectrum efficiency. To encourage users to adopt delay-tolerant scheduling for capacity-improvement, it is critical to provide guarantees in terms of completion time. In this paper, we study application-level scheduling with deadline constraints, where the deadline is pre-specified by users/applications and is associated with a deadline violation probability. To address the exponentially-high complexity due to temporally-varying channel conditions and deadline constraints, we develop a novel asymptotic approach that exploits the largeness of the network to our advantage. Specifically, we identify a lower bound on the deadline violation probability, and propose simple policies that achieve the lower bound in the large-system regime. The results in this paper thus provide a rigorous analytical framework to develop and analyze policies for application-level scheduling under very general settings of channel models and deadline requirements. Further, based on the asymptotic approach, we propose the notion of Application-Level Effective Capacity region, i.e., the throughput region that can be supported subject to deadline constraints, which allows us to quantify the potential gain of application-level scheduling. Huasen Wu, Xiaojun Lin 0001, Xin Liu 0002, Youguang Zhang |
INFOCOM | 4 |
| 2014 | Cyclic delay transmission for unique word OFDM systems
Mang Liao, Xiang-Gen Xia 0001, Youguang Zhang |
Sci. China Inf. Sci. | 3 |
| 2013 | Decision-Feedback Multiuser Detection in Multicell Multicarrier DS-CDMA Systems with/without BS CooperationabstractIn this contribution, we investigate the signal detection in the multicarrier direct-sequence code-division multiple-access (DS-CDMA) systems employing both time (T)-domain and frequency (F)-domain spreading, which are referred to as the TF/MC DS-CDMA systems. A multicell scenario with universal frequency reuse is considered, where signals experience frequency-selective Rayleigh fading. A decision-feedback multiuser detector (MUD), which represents the extension of a so-called receiver multiuser diversity aided multi-stage minimum mean-square error MUD (RMD/MS-MMSE MUD), is employed to cope with both the intracell multiuser interference (MUI) and intercell interference (ICI), when base-station (BS) cooperation is or is not assumed. Furthermore, when BS cooperation is assumed, symbols detected at different BSs are interchanged and combined to enhance the detection performance. Our studies and performance results demonstrate that the RMD/MS-MMSE MUD constitutes one of the promising multicell processing (MCP) schemes, making it possible for each cell to support heavily overloaded users, even when a frequency-reuse factor one is applied. Xiaojie Ju, Youguang Zhang, Lie-Liang Yang |
VTC Spring | 2 |
| 2013 | Diversity analysis for space-time-frequency (STF) coded MIMO system with a general correlation modelabstractPrevious works on space-time-frequency (STF) codes have focused on frequency-selective but independent fading channels for MIMO-OFDM systems. However, insufficient antenna space, mobile scenarios, and multipath in practice cause correlation in the spatial, temporal, and frequency domains. This paper studies the effect of STF coding on the performance of MIMO-OFDM systems with general spatial, temporal, and frequency/path correlated channels. Specifically, we first derive an upper bound on the maximum achievable diversity and then prove achievability by giving an STF code design example. Finally, we show that our general diversity result recovers those in the existing literature for special correlation structures. Mang Liao, Youguang Zhang, Zixiang Xiong |
WCNC | 2 |
| 2013 | Laxity-based opportunistic scheduling with flow-level dynamics and deadlinesabstractMany data applications in the next generation cellular networks, such as content precaching and video progressive downloading, require flow-level quality of service (QoS) guarantees. One such requirement is deadline, where the transmission task needs to be completed before the application-specific time. To minimize the number of uncompleted transmission tasks, we study laxity-based scheduling policies in this paper. We propose a Less-Laxity-Higher-Possible-Rate (L2HPR) policy and prove its asymptotic optimality in underloaded identical-deadline systems. The asymptotic optimality of L2HPR can be applied to estimate the schedulability of a system and provide insights on the design of scheduling policies for general systems. Based on it, we propose a framework and three heuristic policies for practical systems. Simulation results demonstrate the asymptotic optimality of L2HPR and performance improvement of proposed policies over greedy policies. Huasen Wu, Youguang Zhang, Xin Liu 0002 |
WCNC | 2 |
| 2013 | Capacity of generalised network multiple-input-multiple-output systems with multicell cooperationabstractIn network multiple‐input–multiple‐output (MIMO) systems, cooperative base stations (BSs) and mobile terminals (MTs) are in general at different geographic locations. Correspondingly, the channel gains with respect to different BSs and/or MTs may obey different distributions, which resulted from, different pathlosses, different strength of line‐of‐sight (LOS) components and so on. Furthermore, future wireless communication systems are expected to be equipped with multiple antennas for transmission/receiving. In these multiantenna systems, signals received by the antennas of one BS from one MT may be correlated. Against the above‐mentioned scenarios, in this contribution, the authors study the capacity of the generalised uplink network MIMO systems with multicell cooperation (MCoP), when propagation pathlosses, fast Rician fading with various LOS components and spatial correlation are simultaneously considered. Specifically, the theory of operator‐valued free probability is introduced to derive the approximate eigenvalue distribution (AED) of the correlation matrix of the equivalent channels. Based on the AED, the approximate capacity of uplink network MIMO systems is then studied. Furthermore, a range of special cases are analysed by specialise the authors model and modifying their results. Finally, both numerical and simulation results are provided for characterising the capacity of network MIMO systems with MCoP. Peng Pan 0003, Youguang Zhang, Xiaojie Ju, Lie-Liang Yang |
IET Commun. | 2 |
| 2013 | FASA: Accelerated S-ALOHA Using Access History for Event-Driven M2M CommunicationsabstractSupporting massive device transmission is challenging in machine-to-machine (M2M) communications. Particularly, in event-driven M2M communications, a large number of devices become activated within a short period of time, which in turn causes high radio congestions and severe access delay. To address this issue, we propose a Fast Adaptive S-ALOHA (FASA) scheme for random access control of M2M communication systems with bursty traffic. Instead of the observation in a single slot, the statistics of consecutive idle and collision slots are used in FASA to accelerate the tracking process of network status that is critical for optimizing S-ALOHA systems. With a design based on drift analysis, the estimate of the number of the active devices under FASA converges fast to the true value. Furthermore, by examining the T-slot drifts, we prove that the proposed FASA scheme is stable as long as the average arrival rate is smaller than e-1, in the sense that the Markov chain derived from the scheme is geometrically ergodic. Simulation results demonstrate that under highly bursty traffic, the proposed FASA scheme outperforms traditional additive schemes such as PB-ALOHA and achieves near-optimal performance in reducing access delays. Moreover, compared to multiplicative schemes, FASA shows its robustness under heavy traffic load in addition to better delay performance. Huasen Wu, Richard J. La, Xin Liu 0002, Youguang Zhang |
IEEE/ACM Trans. Netw. | 5 |
| 2012 | A study on wind vector retrieval algorithm for rotating fan-beam scatterometerabstractRotating fan-beam scatterometer (RFSCAT) is a new type of satellite scatterometer that was proposed about one decade ago.However, just as other rotating scatterometers, relatively larger wind retrieval errors occur in the nadir and outer regions than in the middle regions of the swath. In order to address this problem, a modified wind vector retrieval algorithm for RFSCAT is presented in this paper. The new algorithm is featured with adaptively extending the range of wind direction for each wind vector cell position across the whole swath according to the distribution histogram of the retrieved wind direction bias. Simulation experiments demonstrated that the new established algorithm can effectively improve the wind direction retrieval accuracy in the nadir and outer regions of the RFSCAT swath. Xuetong Xie, Shi Huan, Jianqiang Liu 0001, Shuyan Lang, Youguang Zhang, Di Zhu 0001, Kehai Chen, Juhong Zou, Zhou Huang 0002, Weijun Tao |
IGARSS | 5 |
| 2012 | Fast Adaptive S-ALOHA Scheme for Event-Driven Machine-to-Machine CommunicationsabstractMachine-to-Machine (M2M) communication is now playing a market-changing role in a wide range of business world. However, in event-driven M2M communications, a large number of devices activate within a short period of time, which in turn causes high radio congestions and severe access delay. To address this issue, we propose a Fast Adaptive S- ALOHA (FASA) scheme for M2M communication systems with bursty traffic. The statistics of consecutive idle and collision slots, rather than the observation in a single slot, are used in FASA to accelerate the tracking process of network status. Furthermore, the fast convergence property of FASA is guaranteed by using drift analysis. Simulation results demonstrate that the proposed FASA scheme achieves near-optimal performance in reducing access delay, which outperforms that of traditional additive schemes such as PB-ALOHA. Moreover, compared to multiplicative schemes, FASA shows its robustness even under heavy traffic load in addition to better delay performance. Huasen Wu, Richard J. La, Xin Liu 0002, Youguang Zhang |
VTC Fall | 5 |
| 2011 | Asymptotic Spectral-Efficiency of MIMO-CDMA Systems with Arbitrary Spatial CorrelationabstractIn this contribution, we analyze the asymptotic spectral-efficiency (ASE) of multiuser MIMO-CDMA systems, when assuming communications over flat fading channels with arbitrary spatial correlation. Our analysis is built on the operator-valued free probability theory, which is applied to obtain the limit distribution of the correlation matrix's eigenvalues, as the MIMO-CDMA systems' size tends to infinity. The spectral-efficiency (SE) performance of the MIMO-CDMA systems is investigated via both analysis and simulations. Our simulation and numerical results show that the ASE is capable of providing a good measure of the SE achieved by the corresponding realistic MIMO-CDMA systems. Peng Pan 0003, Youguang Zhang, Yuquan Sun, Lie-Liang Yang |
GLOBECOM | 2 |
| 2011 | Redundant Residue Number System Based Multicarrier DS-CDMA for Dynamic Multiple-Access in Cognitive RadiosabstractRedundant residue number system (RRNS)-based multicarrier DS-CDMA (MC/DS-CDMA) is proposed for dynamic multiple-access (DMA) in cognitive radios (CRs). The proposed RRNS-based MC/DS-CDMA DMA has the merits of low-complexity for implementation, high-flexibility for reconfiguration and spectrum handoff, robustness to spectrum varying, and fault-tolerance to errors. Specifically, in our RRNS-based MC/DS-CDMA DMA system, RRNS-based orthogonal modulation aided by MC/DS-CDMA is employed for information transmission. At the receiver, signals are detected subcarrier-by-subcarrier independently based on suboptimum MMSE interference cancellation (SMMSE-IC). In performance study, we model the arrival process of primary users (PUs) in primary radios (PRs) as a Poisson process. Both the bit error rate (BER) performance and throughput performance are investigated. Our studies and performance results show that the RRNS-based MC/DS-CDMA constitutes one of the highly promising DMA schemes for application in CRs. It is capable of achieving a substantial throughput with required quality for the CR systems, while without degrading the quality-of-services (QoS) of the PR systems. Youguang Zhang, Lie-Liang Yang |
VTC Spring | 2 |
| 2010 | Asymptotic Performance Analysis of Time-Frequency-Domain Spread MC DS-CDMA Systems Employing MMSE Multiuser DetectionabstractIn this contribution the asymptotic signal-to-interference-plus-noise ratio (SINR) performance of multicarrier direct-sequence code-division multiple-access systems employing time-frequency-domain spreading, i.e., of the TF/MC DS-CDMA systems, is studied, when separate minimum mean-square error multiuser detection (MMSE-MUD) is considered. The separate MMSE-MUD detects signals first in the time (T)-domain and then in the frequency (F)-domain. Based on random matrix theory, closed-form expressions for the asymptotic SINR of the TF/MC DS-CDMA systems using separate MMSE-MUD is derived, when communicating over additive white Gaussian noise (AWGN) channels. The closed-form expressions show that the asymptotic SINR performance is only depended on the T- and F-domain user load factors as well as noise variance. Hence, they are beneficial to evaluation. Furthermore, our simulation and numerical results show that in most cases the asymptotic SINR can provide a good approximation to the SINR achieved by realistic TF/MC DS-CDMA systems employing separate MMSE-MUD. Peng Pan 0003, Youguang Zhang, Lie-Liang Yang |
VTC Spring | 2 |
| 2008 | Spectral-Efficiency of Time-Frequency-Domain Spread Multicarrier DS-CDMA in Frequency-Selective Nakagami-m Fading ChannelsabstractIn this contribution we study the spectral-efficiency performance of a multicarrier direct-sequence code-division multiple-access system using both time (T)-domain and frequency (F)-domain spreading, which is referred to as TF/MC DS-CDMA for convenience, when communicating over Nakagami-m fading channels. We consider the TF/MC DS-CDMA scheme, since it is a generalized multiple-access scheme that can be readily configured to some other CDMA schemes, so that we can conveniently compare their spectral-efficiency performance. Furthermore, the Nakagami-m fading channel model makes it possible to study the impact of channel quality on the achievable spectral-efficiency. In this contribution, the spectral-efficiency of the TF/MC DS-CDMA is investigated, when various single-user and multiuser detection schemes are invoked, which include optimum detection, minimum mean-square error (MMSE) detection, zero-forcing (ZF) detection and the matched-filter (MF) based detection. One of our conclusions is that channel quality has insignificant impact on the achievable spectral-efficiency of the TF/MC DS-CDMA, when the number of T-domain resolvable paths of the frequency-selective fading channels is sufficiently high. Peng Pan 0003, Lie-Liang Yang, Youguang Zhang |
VTC Fall | 3 |