VLDB 2026 Research / reviewers in the wild / expert
Yongliang Zhou
dblp:71/8591
· DBLP profile ↗
18ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-7327-6759ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MTJ-back-gate SRAM CIM with replica quantization for temperature robust
Yongliang Zhou, Chengxing Dai, Xiulong Wu, Chunyu Peng |
Integr. | 1 |
| 2026 | Low-temperature-drift voltage reference design using magnetic tunnel junctions
Yongliang Zhou, Yingxue Sun, Wangyong Si, Jingxue Zhong, Weizhe Tan, Chunyu Peng, Xiulong Wu |
Integr. | 1 |
| 2026 | Analysis and Design of Memory Testing Algorithm for Computing-in-Memory Using MBISTabstractComputing-in-memory (CIM), as a novel computing architecture for the future, effectively overcomes the bottlenecks in the von Neumann architecture. The CIM architecture embeds logic into the memory array to reduce the data transfer between the processor and memory. However, embedding logic into the memory array increases the test complexity. In this study, we offer a comprehensive examination of the challenges associated with CIM and introduce a novel March-like test algorithm, named March CC, tailored for CIM chips. Computational elements are added to the read/write operation sequences, combining the tests in memory mode and computing mode into one step, which significantly improves the test efficiency. In comparison to the traditional March C− test algorithm, the proposed March CC test algorithm, with a complexity of only 10 N , enhances the fault coverage from 66.7% to 79.8% for six common single-cell fault (SCF) models and nine common double-cell fault (DCF) models. Furthermore, the March CC algorithm demonstrates good compatibility and is applicable to various memory configurations, such as SRAM, RRAM, and MRAM CIM architectures. Zhi-Ting Lin, Siyan Li, Qiushi Feng, Changxin Yue, Yuanyang Wang, Yunlong Liu 0006, Yu Liu 0113, Licai Hao, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Xiulong Wu |
ACM J. Emerg. Technol. Comput. Syst. | 13 |
| 2026 | A Capacitor Discharge-Based SRAM CIM Macro Based on Hybrid-Domain for Convolutional Neural NetworksabstractCompute-in-memory (CIM) is increasingly recognized as an effective hardware accelerator for convolutional neural networks (CNNs). This work proposes a hybrid-domain CIM design using: 1) a multibit compute unit (MBCU) structure that realizes the multiplication operation of 2-bit input and 4-bit weight through the transistor-size-weighted capacitor discharge on the bitline; 2) a hybrid-domain quantization scheme (HDQS) of “time-domain + voltage-domain,” which integrates the high energy efficiency of time-domain quantization with the low-delay advantages of the voltage-domain quantization, and enhances the quantization accuracy through the combined effect of the process tracking module and the reference signal module; 3) the CIM circuit design, layout drawing and simulation verification of hybrid-domain static random access memory (SRAM) were realized by 28-nm CMOS technology, results show that the circuit supports 8-bit multiply–accumulate (MAC) operation, and full-precision quantization in the hybrid-domain form can achieve the optimal energy efficiency of 249.7 TOPS/W per bit at 0.7 V, and area efficiency of 4.29 TOPS/mm2per bit. Furthermore, the integration of the circuits with the VGG-16 network has been demonstrated to yield an inference accuracy of 90.52% in the CIFAR-10 dataset. Bin Qiang, Yongliang Zhou, Xiulong Wu, Chunyu Peng |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | MTJ based temperature compensated beta multiplier Voltage ReferenceabstractThis article mainly explores the interaction between CMOS devices and Magnetic Tunnel Junction (MTJ) devices in terms of temperature characteristics, aiming to achieve a CMOS beta multiplier circuit that combines low power consumption and wide temperature adaptability, making it a stable reference voltage source. The proposed design utilizes the Tunneling Magneto Resistance (TMR) effect of MTJ to compensate for the performance mismatch caused by temperature changes in CMOS. This design adopts TSMC 28nm CMOS craft, which can generate a reference voltage with a linearity of 0.57%/V, a temperature coefficient of 43.6ppm/°C, and a stable voltage of 441.6mV at a minimum supply voltage of 0.6V and a temperature range of 10~110 °C. The noise of this voltage at a frequency of 10Hz is 69.7uV/sqrt (Hz), the power suppression ratio is -62.1dB, and the power consumption is 4.092nW. Yongliang Zhou, Yingxue Sun, Jingxue Zhong, Chengxing Dai, Weizhe Tan, Chunyu Peng, Xin Li 0099, Zhi-Ting Lin, Xiulong Wu |
ISCAS | 1 |
| 2025 | MTJ based Temperature-Adaptive VCO (TAVCO) for Compensating CP-PLL Frequency DriftabstractThe Charge Pump Phase-Locked Loop (CP-PLL) is a commonly utilized component in contemporary mixed-signal electronic systems. It is widely employed for clock generation, synchronization, and frequency synthesis in both digital and wireless functionalities. However, the frequency accuracy of oscillators can be adversely affected by variations in frequency across a broad temperature range. To address this issue, the voltage-controlled oscillator designed in this study employs a four-stage differential delay structure, chosen for its simple circuit architecture, favorable control linearity, and low noise characteristics. This research integrates the temperature behaviors of Complementary Metal-Oxide-Semiconductor (CMOS) and Magnetic Tunnel Junction (MTJ) technologies, utilizing 28nm CMOS technology to enhance the frequency stability of ring oscillators effectively. Simulation results indicate that frequency drift is reduced by 92% within the temperature range of -80°C to 125°C. Yongliang Zhou, Jingxue Zhong, Yingxue Sun, Chengxing Dai, Weizhe Tan, Chunyu Peng, Wenjuan Lu, Xin Li 0099, Zhi-Ting Lin, Xiulong Wu |
ISCAS | 1 |
| 2025 | Full-Array Boolean Logic CIM Macro With Self-Recycling 10T-SRAM Cell for AES SystemsabstractComputing in memory (CIM), which alleviates the need to transfer a large amount of data between processor and memory, significantly reducing latency and energy consumption, is a promising new computing architecture for addressing the von Neumann bottleneck problem. This article proposes a CIM array structure composed of self-recycling 10T static random access memory (SRAM) cells, which can realize orthogonal data writing, and multiple Boolean logical operations for the entire array. The self-recycling and full-array activation characteristics are extremely suitable for accelerating diverse data processing algorithms such as the Advanced Encryption Standard (AES). A 4-kb SRAM is implemented in 55-nm CMOS technology to verify the effectiveness of the design. Compared with other state-of-the-art architectures, the throughput and the operating frequency of the proposed CIM macro are increased to 843 GOPS/kb ($2.64\times $) and 823.7 MHz ($2.6\times $), respectively. The energy efficiency reaches 246.9 TOPS/W. When applied to the AES, the energy consumption is 35.77% less than the digital CIM architecture that is not self-recycling. Xin Li 0099, Lintao Chen, Yang Lou, Baofa Wu, Jiajun Long, Yongliang Zhou, Chunyu Peng, Xiulong Wu, Zhi-Ting Lin |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2024 | A Timing-Shared Adaptive Sensing Methodology for Low-Voltage SRAMabstractLowing static random access memory (SRAM) supply voltage could highly improve energy efficiency, yet energy efficiency still not attains optimal point due to the constraint of the weakest bit-cell, especially in low-voltage SRAM. Adaptive sensing methodology is proposed for the challenge and consists of four elements: Switch unit is built to implement cross-sensing operation, Timing-Shared Decoupling Latch Sense Amplifier (TS-DLSA) allows rapid and successive sensing, the judging module is utilized to trigger the FLAG signal which is for adaptive timing controller to cut off word line (WL). Energy efficiency was obtained by compressing the activation delay of WL compared to global timing scheme. The proposed adaptive sensing methodology was performed with TSMC 28-nm CMOS process, and evaluated in SRAM array of 128x128, 256x256, 512x512, and 1024x1024. Monte Carlo simulation results are formed to confirm that, compared to the global timing scheme, the proposed adaptive sensing methodology has reduced WL activation delay by 75.1%∼41.8% and read operation energy overhead by 77.5%∼30.9% from 0.6V to 1.2V. Compared to the current-latched sense amplifier with a footswitch (FS-CLSA) with proposed adaptive sensing methodology, TS-DLSA with proposed scheme have reduced the WL activation by 4.1%∼10.5% and read operation energy overhead by 15.4%∼23.8% from 256x256 to 2048x2048 at 0.6V. The more cells mounted on the BL, the higher energy revenue gains. Yongliang Zhou, Saiai Wu, Wenjuan Lu, Chunyu Peng, Xin Li 0099, Xiulong Wu |
ISCAS | 1 |
| 2024 | Ultra8T: A sub-threshold 8T SRAM with leakage detection
Shan Shen, Yongliang Zhou, Wenjian Yu |
Integr. | 3 |
| 2024 | Timing Optimization Model and PVT Tracked Scheme for STT-MRAM Voltage-Mode SenseabstractThe impact of process variations on the read operation of low-voltage STT-MRAM becomes severe, posing a challenge in determining the optimal sensing timing of the sense amplifier. This study investigates techniques for refining the timing scheme of sensing circuits in order to improve the sensing reliability of the STT-MRAM. The supply voltage$V_{DD}$, the Tunneling Magnetoresistance Ratio TMR, the low resistance state of bit-cell$R_{P}$, and the parasitic capacitance of bit-line$C_{BL}$are analyzed along with the voltage sense amplifier (VSA) involved in sensing yield. We develop a timing model through theoretical analysis to determine the optimal VSA enable signal (SAE). In addition, an innovative Process-Voltage-Temperature (PVT) tracking scheme is proposed that can track the optimal VSA enable signal (SAE) and suppress timing variations. Monte-Carol simulation in the 28nm CMOS and magnetic tunnel junction (MTJ) process confirms that the combined scheme significantly enhances the robustness of sensing operation. The proposed scheme improves yield by 20% to 35%, reduces power consumption by 43% to 63%, and reduces read access delay by 47% to 59% compared to conventional sensing schemes at 0.6V supply voltage. Yongliang Zhou, Yingxue Sun, Chengxing Dai, Jingxue Zhong, Xiulong Wu, Chunyu Peng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | A CFMB STT-MRAM-Based Computing-in-Memory Proposal With Cascade Computing Unit for Edge AI DevicesabstractThe application of non-volatile memory technology is increasingly attractive for Computing-in-memory (CIM) owing to high integration density and negligible standby power consumption. This study proposes an spin-transfer-torque (STT) magnetic random access memory (MRAM) based CIM macro which incorporates following innovative features: 1) cross-feedback margin-boost (CFMB) scheme to enable robust and fast reading operations against process variation and limited Tunneling Magnetoresistance Ratio (TMR); 2) cascade computing units (CCU) and related design method for efficient and stable multi-bit multiply-and-accumulate (MAC) operation; and 3) dual computing mode scheme and resolution adjustable quantization module to optimize energy efficiency and operating speed. The post-simulations are performed under 28nm CMOS&MTJ technology. The results demonstrate the achievement in energy efficiency of 36.4 TOPS/W while performing MAC operations with up to 16-bit weights, 4-bit inputs, and 22-bit outputs. Yongliang Zhou, Chenghu Dai, Licai Hao, Chunyu Peng, Hao Cai 0001, Xiulong Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | Low-Cost and Highly Robust Quadruple Node Upset Tolerant Latch DesignabstractThis article proposes an exceptionally reliable and low-cost quadruple node upset tolerant latch ($LC$-QNUTL) suitable for the 65 nm CMOS technology. The innovative$LC$-QNUTL latch is primarily composed of three soft-error-immune (SEI) static random-access memory (SRAM) cells and a triple-level C-element (CE) unit, which includes five two-input CE and a clock-gating (CG)-based two-input CE. The SEI SRAM cell utilizes polarity hardening technology and source-isolation technology, significantly reducing the number of sensitive nodes and enhancing the latch’s stability. By using the high-speed transmission gate (TG) technology and stacked structures, the proposed latch offers minimal overhead in terms of delay and power consumption, yielding an improved power delay area product (PDAP). When compared to contemporary quadruple node upset (QNU)-tolerant latch designs (including HLMR, 4NUHL, and LDAVPM), the new design offers substantial improvements—29.53% less delay, 80.09% reduced power consumption, 58.52% smaller silicon area, and 433.43% improved comprehensive PDAP on average. Furthermore, simulation results demonstrate that the$LC$-QNUTL latch exhibits reduced sensitivity to process, voltage, and temperature (PVT) variations, thus providing superior reliability, which makes it an ideal choice for safety-critical applications. Licai Hao, Yaling Wang, Yunlong Liu 0006, Shiyu Zhao 0004, Wenjuan Lu, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 10 |
| 2024 | Soft-Error-Immune Quadruple-Node-Upset Tolerant Latch Based on Polarity Design and Source-Isolation TechnologiesabstractA soft-error-immune quadruple-node-upset tolerant latch (SEI-QNUTL) with a low delay and high performance is proposed using 65-nm CMOS technology. The proposed SEI-QNUTL design consists of three soft-error-immune static random access memory (SEI-SRAM) cells. Furthermore, each SEI-SRAM cell employs polarity design and source-isolation technology to reduce the number of sensitive nodes and enhance the reliability of the latch. Compared with state-of-the-art quadruple-node-upset (QNU) tolerant latches [including high-performance and low-cost single-event multiple-node-upsets resilient (HLMR), QNU tolerant latch (QNUTL), and Latch Design and Algorithm-based Verification Protected against Multiple-Node-Upsets (LDAVPM)], the proposed SEI-QNUTL design reduces (on average) the area, delay, and area-power-delay-product (APDP) by 47.0%, 25.0%, 46.5%, and 66.3%, respectively. Extensive variation analysis validates that the SEI-QNUTL design is less sensitive to process, voltage, and temperature (PVT) variations regarding power consumption and delay. Furthermore, Monte Carlo (MC) simulations show that the proposed latch exhibits high reliability when performing data storage. Compared with the existing latches, the SEI-QNUTL design makes a good tradeoff among delay, power, and area, and it can thus be used in safety-critical applications. Licai Hao, Chenghu Dai, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng, Yongliang Zhou, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2022 | ShareFloat CIM: A Compute-In-Memory Architecture with Floating-Point Multiply-and-Accumulate OperationsabstractCompute-in-memory (CIM) has been widely explored to overcome “Von-Neumann bottleneck” for its high throughput and energy efficiency. However, recent compute-in-memory works can only support integer (INT)-type multiply-and-accumulate (MAC) operations. Floating point MACs (FP-MAC) are highly required to achieve both high performance training and high accuracy inference. In this paper, we proposed a ShareFloat CIM architecture which can support FP-MAC operations. Neural networks with ShareFloat MAC can achieve almost the same accuracy as that with FP64 MAC. A 28nm 64Kb ShareFloat CIM macro was further implemented with an energy efficiency of 18.8 TFLOPS/W and 73.11% accuracy when applied to a VGG-16 network with ShareFloat MAC and CIFAR-100 dataset. An Guo 0001, Yongliang Zhou, Bo Wang 0023, Tianzhu Xiong, Xin Si, Jun Yang 0006 |
ISCAS | 2 |
| 2022 | SNNIM: A 10T-SRAM based Spiking-Neural-Network-In-Memory architecture with capacitance computationabstractSpiking-Neural-Networks (SNN) have natural advantages in high-speed signal processing and big data operation. However, due to the complex implementation of synaptic arrays, SNN based accelerators may face low area utilization and high energy consumption. Computing-In-Memory (CIM) shows great potential in performing intensive and high energy efficient computations. In this work, we proposed a JOT-SRAM based Spiking-Neural-Network-In-Memory architecture (SNNIM) with 28nm CMOS technology node. A compact JOT-SRAM bit-cell was developed to realize signed 5bit synapses arrays and configurable bias arrays (SYBIA). The soma array based standard 8T-SRAM (SMTA) stores the soma membrane voltage and the threshold value. A capacitance computation scheme (CCA) between them was proposed to support various SNN operations. The proposed SNNIM achieved energy efficiency of 25.18 TSyOPSI. And the proposed SNNIM achieved 1.79+× better array efficiency compared with previous works. Bo Wang 0023, Xiang Li 0147, Anran Yin, Zhongyuan Feng, Yuyao Kong, Tianzhu Xiong, Haiming Hsu, Yongliang Zhou, An Guo 0001, Jun Yang 0006, Xin Si |
ISCAS | 10 |
| 2021 | A survey of in-spin transfer torque MRAM computing
Hao Cai 0001, Bo Liu 0019, Juntong Chen, Lirida A. B. Naviner, Yongliang Zhou, Zhen Wang 0019, Jun Yang 0006 |
Sci. China Inf. Sci. | 5 |
| 2020 | Interplay Bitwise Operation in Emerging MRAM for Efficient In-memory Computing
Hao Cai 0001, Honglan Jiang, Yongliang Zhou, Menglin Han, Bo Liu 0019 |
CCF Trans. High Perform. Comput. | 3 |
| 2011 | Rank-two residue iteration method for nonnegative matrix factorization
Hongwei Liu 0001, Yongliang Zhou |
Neurocomputing | 2 |