EDBT 2026 Demo / reviewers in the wild / expert
Qiang Zhao 0007
dblp:00/2166-7
· DBLP profile ↗
14ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-0278-5804ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Analysis and Design of Memory Testing Algorithm for Computing-in-Memory Using MBISTabstractComputing-in-memory (CIM), as a novel computing architecture for the future, effectively overcomes the bottlenecks in the von Neumann architecture. The CIM architecture embeds logic into the memory array to reduce the data transfer between the processor and memory. However, embedding logic into the memory array increases the test complexity. In this study, we offer a comprehensive examination of the challenges associated with CIM and introduce a novel March-like test algorithm, named March CC, tailored for CIM chips. Computational elements are added to the read/write operation sequences, combining the tests in memory mode and computing mode into one step, which significantly improves the test efficiency. In comparison to the traditional March C− test algorithm, the proposed March CC test algorithm, with a complexity of only 10 N , enhances the fault coverage from 66.7% to 79.8% for six common single-cell fault (SCF) models and nine common double-cell fault (DCF) models. Furthermore, the March CC algorithm demonstrates good compatibility and is applicable to various memory configurations, such as SRAM, RRAM, and MRAM CIM architectures. Zhi-Ting Lin, Siyan Li, Qiushi Feng, Changxin Yue, Yuanyang Wang, Yunlong Liu 0006, Yu Liu 0113, Licai Hao, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Xiulong Wu |
ACM J. Emerg. Technol. Comput. Syst. | 12 |
| 2026 | A CIM Macro Embedded With Sign Operations for Parallel Signed Multibit Multiplication-and-Accumulation Using Hybrid Cell Array
Jin Zhang 0036, Zhongzhen Tong, Qiang Zhao 0007, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | A Low-Cost and Triple-Node-Upset Self-Recoverable Latch Design With Low Soft Error RateabstractWith the decrease in feature size of transistors, latches are more sensitive to single-event multiple node upset (MNU), including double node upset (DNU) and triple node upset (TNU). However, the reported TNU self-recoverable (TNUR) latches are facing problems with large areas and power consumption. Based on the polarity design, this article proposes a low-cost TNUR latch (LCTRL) with a low soft error rate (SER) in 28-nm CMOS technology. The proposed LCTRL mainly consists of four interlocked modules and a clock-gated inverter. Compared with the state-of-the-art TNUR latches, including LCTNURL, IHTRL, FATNU, and TRLW, the power consumption, D-Q delay, CLK-to-Q delay, area, and the power-delay–area product (PDAP) of the proposed LCTRL are reduced by 55.09%, 38.64%, 42.93%, 44.65%, and 83.50%, respectively. Due to the polarity design, the SER of the proposed LCTRL is the smallest among compared latches, which suggests that the proposed LCTRL is suitable for use in radiation environments. Licai Hao, Lang Tian, Hao Wang 0239, Shiyu Zhao 0004, Qiang Zhao 0007, Chunyu Peng, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | A 28-nm 9T1C SRAM-Based CIM Macro With Hierarchical Capacitance Weighting and Two-Step Capacitive Comparison ADCs for CNNsabstractIn the realm of charge-domain computing-in-memory (CIM) macros, reducing the area of capacitor ladder and analog-to-digital converter (ADC) while maintaining high throughput remains a significant challenge. This brief introduces an adjustable-weight CIM macro designed to enhance both energy efficiency and area efficiency for convolutional neural networks (CNNs). The proposed architecture uses: 1) a customized 9T1C bit cell for sensing margin improvement and bidirectional decoupled read ports; 2) a hierarchical capacitance weighting (HCW) structure that achieves a weight accumulation of 1/2/4 bits with less capacitance area and weighting time; and 3) a two-step capacitive comparison ADCs (TC-ADCs) readout scheme to improve area efficiency and throughput. The proposed 8-kb static random address memory (SRAM) CIM macro is implemented using 28-nm CMOS technology. It can achieve an energy efficiency of 224.4 TOPS/W and an area efficiency of 21.894 TOPS/mm2, and the accuracies on MNIST, CIFAR-10, and CIFAR-100 datasets are 99.67%, 89.13%, and 67.58% with a 4-b input and 4-b weight. Zhi-Ting Lin, Runru Yu, Miao Long, Yu Liu 0113, Jianxing Zhou, Qingchuan Zhu, Yue Zhao 0029, Lintao Chen, Chunyu Peng, Qiang Zhao 0007, Xin Li 0099, Chenghu Dai, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 12 |
| 2025 | A High-Performance and High-Robustness Triple-Node-Upset Tolerant Latch Based on Redundant-Node HardeningabstractIn response to the issues of high cost, large overhead, and limited node fault tolerance in current latch hardening techniques, this article proposes a latch circuit resistant to triple-node-upset (TNU) based on redundant-node hardening technology. This latch comprises eight 1P2N modules interlocked, with its output isolated by two levels of C-elements (CEs), achieving tolerance to TNU. The performance of the redundant-node reinforcement TNU tolerant latch (RNRTTL) was simulated and verified using CMOS 65 nm technology. The simulation results indicate that the RNRTTL circuit has a D-Q delay of 14.14 ps, static power consumption of$4.03~\mu $w, an area of$32.87~\mu $m2, and an area-static power-D–Q delay-product (APDP) of 1873, respectively. Compared to the triple-node upset tolerant latches TTLL, TNU-latch, TNURL, and HLTNURL reported in the current literature, the proposed latch demonstrates an average reduction of 219.9%, 164.9%, 150.7%, and 2464.8% in D-Q delay, static power consumption, area, and APDP, respectively, indicating that the RNRTTL latch has superior comprehensive performance; furthermore, a series of 2000 Monte Carlo (MC) simulations on the node group$\langle $Q, X0, X$8\rangle $reveal that the proposed latch circuit possesses good stability, making it suitable for harsh radiation environments. Qiang Zhao 0007, Qingyi Liu, Licai Hao, Xin Li 0099, Shengyue Zhang, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | A Computing In-Memory Multibit Multiplication Based on Decoupling and In-Array StoringabstractMultiplications are basic operations of neural networks. Therefore, multiplication results are crucial in analyzing the operating process of neural networks. However, the multiplication strategies are generally based on analog-domain circuits, and the results are in a multiply-and-accumulate (MAC) form. The result of each multiplication in MAC cannot be distinguished accurately using these strategies. Therefore, we proposed an in-memory multibit multiplication based on the decoupling and in-array storage strategy to overcome this problem, and the core module is the 10T1C SRAM cell. Multibit multiplications are decoupled by a series of logical operations. Therefore, in the analysis mode, multiplication results can be saved and outputted in the normal read mode without requiring additional storage. When executing the neural network, the operation results are stored in the cells. Hence, the operands stored in the array are retained. Accumulation operations are completed based on the charge-sharing technology; thus, the linearity of accumulation is high. We simulated and analyzed the performance of the proposed circuit in a 28 nm CMOS process. The absolute value of integral nonlinearity is at most 0.29. Further, due to high data operation parallelism, the throughputs of the logical operation and MAC are up to 6307.8 and 802.8 GOPS, respectively. Jin Zhang 0036, Zhongzhen Tong, Hao Wang 0239, Qiang Zhao 0007, Jiaqun Wang, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | Low-Cost and Highly Robust Quadruple Node Upset Tolerant Latch DesignabstractThis article proposes an exceptionally reliable and low-cost quadruple node upset tolerant latch ($LC$-QNUTL) suitable for the 65 nm CMOS technology. The innovative$LC$-QNUTL latch is primarily composed of three soft-error-immune (SEI) static random-access memory (SRAM) cells and a triple-level C-element (CE) unit, which includes five two-input CE and a clock-gating (CG)-based two-input CE. The SEI SRAM cell utilizes polarity hardening technology and source-isolation technology, significantly reducing the number of sensitive nodes and enhancing the latch’s stability. By using the high-speed transmission gate (TG) technology and stacked structures, the proposed latch offers minimal overhead in terms of delay and power consumption, yielding an improved power delay area product (PDAP). When compared to contemporary quadruple node upset (QNU)-tolerant latch designs (including HLMR, 4NUHL, and LDAVPM), the new design offers substantial improvements—29.53% less delay, 80.09% reduced power consumption, 58.52% smaller silicon area, and 433.43% improved comprehensive PDAP on average. Furthermore, simulation results demonstrate that the$LC$-QNUTL latch exhibits reduced sensitivity to process, voltage, and temperature (PVT) variations, thus providing superior reliability, which makes it an ideal choice for safety-critical applications. Licai Hao, Yaling Wang, Yunlong Liu 0006, Shiyu Zhao 0004, Wenjuan Lu, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2024 | Soft-Error-Immune Quadruple-Node-Upset Tolerant Latch Based on Polarity Design and Source-Isolation TechnologiesabstractA soft-error-immune quadruple-node-upset tolerant latch (SEI-QNUTL) with a low delay and high performance is proposed using 65-nm CMOS technology. The proposed SEI-QNUTL design consists of three soft-error-immune static random access memory (SEI-SRAM) cells. Furthermore, each SEI-SRAM cell employs polarity design and source-isolation technology to reduce the number of sensitive nodes and enhance the reliability of the latch. Compared with state-of-the-art quadruple-node-upset (QNU) tolerant latches [including high-performance and low-cost single-event multiple-node-upsets resilient (HLMR), QNU tolerant latch (QNUTL), and Latch Design and Algorithm-based Verification Protected against Multiple-Node-Upsets (LDAVPM)], the proposed SEI-QNUTL design reduces (on average) the area, delay, and area-power-delay-product (APDP) by 47.0%, 25.0%, 46.5%, and 66.3%, respectively. Extensive variation analysis validates that the SEI-QNUTL design is less sensitive to process, voltage, and temperature (PVT) variations regarding power consumption and delay. Furthermore, Monte Carlo (MC) simulations show that the proposed latch exhibits high reliability when performing data storage. Compared with the existing latches, the SEI-QNUTL design makes a good tradeoff among delay, power, and area, and it can thus be used in safety-critical applications. Licai Hao, Chenghu Dai, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng, Yongliang Zhou, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | Four-branch Siamese network based on sketch-specific data augmentation for sketch recognitionabstractAbstract The shortened abstract is as follows: Sketch recognition has become an important hotspot issue because of sketch's intuitiveness and visualization. The existing sketch‐recognition methods based on handcrafted features and deep features are insufficient in the recognition of the local information of sketches, and the recognition accuracy is not ideal. Accordingly, this paper proposes a four‐branch Siamese network based on sketch‐specific data augmentation to generate discriminative feature representations and improve the sketch‐recognition accuracy. A sketch is an ordered list of strokes, we adopt the semantic information of strokes as the decomposition criteria to divide a sketch into three disjoint local blocks, and then combine the local blocks in pairs to form three new sketches. In order to give full play to the positive effect of local blocks on category prediction and enhance the fine‐grained capability of the network, three newly generated sketches and the original sketch are combined to construct a four‐branch Siamese network. Each branch network adopts the Sketch‐A‐Net architecture with the fully connected layer removed as the basic network, and we improve it by adding shortcut connection layer and multi‐scale weighted bilinear coding (MWBC) modules. Compared with the state‐of‐the‐art methods, the experimental results on the TU‐Berlin dataset demonstrate the excellent performance of our model. Xiuying Wang 0004, Qiang Zhao 0007, Shoubiao Tan |
IET Image Process. | 2 |
| 2023 | High Restore Yield NVSRAM Structures With Dual Complementary RRAM Devices for High-Speed ApplicationsabstractStatic random access memory (SRAM) plays a key role in the overall performance of electronic systems because of its rapid data processing and transmission speed; however, when the system power supply is cut off, the data stored in the nodes are lost. Thus, this article proposes four nonvolatile SRAM (NVSRAM) cells that use unilateral or bilateral structures with dual complementary series resistive random access memory (RRAM) devices. It is found that the read, write, and hold static noise margins (HSNMs) are comparable with those of the standard 6T-SRAM. Moreover, the store and restore operations operate in parallel at high speed. The store operation delay is only 6 ns for unilateral structures and 5 ns for bilateral structures, and the restore delay is only 10 ns for unilateral structures and 6 ns for bilateral structures. The maximum power consumption among the four structures for storing and restoring a “1” are 1.545 pJ/bit and 134.5 fJ/bit, respectively. Furthermore, the dual complementary series resistor structures can achieve a high restore yield at a resistance ratio of 1.5. Therefore, a high restore yield can be achieved even with large resistance fluctuations caused by the voltage, time, and process. Zhi-Ting Lin, Xiulong Wu, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | A Fully Digital SRAM-Based Four-Layer In-Memory Computing Unit Achieving Multiplication Operations and Results StoreabstractThe separation of memory and arithmetic logic unit (ALU) in the von Neumann computing architecture hinders the development of big data and high-performance computing. In-memory computing (IMC) as a new computation method significantly reduces the latency and power consumption of data processing. In this study, we propose a fully digital static random access memory (SRAM)-based IMC architecture, which has the following advantages: 1) it simplifies multiplication to multicycle addition operations, reuses logic cells, and reduces hardware overhead; 2) by adding a pair of nMOS transistors to achieve internal write-back, the computational efficiency is improved, and at the same time, the final result of the multiplication can be stored locally, eliminating the need to read the computational result immediately; and 3) this scheme can be easily expanded to multiplication operations with different bit widths, which provides good scalability. A 4-kb SRAM-IMC macro chip is manufactured using the SMIC 55-nm technology to realize 4-bit multiplication, with an energy efficiency of 51.4 TOPS/W (0.9 V) and a throughput of 234.3 GOPS/mm2. The proposed multiplication–accumulation architecture is applied to a neural network, which achieves 98.7% accuracy with the Mixed National Institute of Standards and Technology database (MNIST) dataset. Zhi-Ting Lin, Shaoying Zhang, Jianping Xia, Yunwei Liu, Kefeng Yu, Zhongzhen Tong, Xiulong Wu, Wenjuan Lu, Chunyu Peng, Qiang Zhao 0007 |
IEEE Trans. Very Large Scale Integr. Syst. | 15 |
| 2022 | Configurable Memory With a Multilevel Shared Structure Enabling In-Memory ComputingabstractFrequent to-and-from data transfers in the von Neumann architecture limit the overall throughput. One of the promising approaches used to overcome von Neumann bottleneck is in-memory computing (IMC) that aims to embed computing in memory to reduce the transfer of memory-processor data. This study proposes a configurable 6-transistor (6T) static random access memory (SRAM) array with a multilevel shared structure for IMC. A multilevel shared structure can effectively improve the utilization rate of the module. In addition to the conventional SRAM operation, the configurable structure can also perform the sum of absolute differences (SAD) and Hamming distance (HD) calculations. To quickly identify the minimum value among multiple calculation results, a four-input sense amplifier (SA) is proposed. The performance of the proposed memory is simulated in a 65-nm CMOS process. The post-layout simulation results show good linearity of the multirow read in the SAD and HD modes. The mean time required by the four-input SA to obtain the result is 190 ps. The SAD and HD calculations yield consumptions of 67.44 fJ/byte and 0.64 fJ/bit, respectively, at 0.8 V. Furthermore, a single column-sharing comparator consumes 2.78 and 3.41 pJ at 0.8 V in the SAD and HD modes, respectively. Yue Zhao 0029, Zhi-Ting Lin, Xiulong Wu, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng, Zhongzhen Tong, Junning Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | Novel Write-Enhanced and Highly Reliable RHPD-12T SRAM Cells for Space ApplicationsabstractIn this brief, we proposed, based on the polarity upset mechanism of single-event transient voltage of n-channel metal-oxide-semiconductor (nMOS) transistors, a novel radiation hardened by polar design (RHPD) 12T SRAM cell to enhance the reliability and operation speed for space applications. Simulation results in Semiconductor Manufacturing International Corporation (SMIC) 65-nm CMOS commercial standard process show that the proposed RHPD-12T cell can tolerate all single-node upsets. Meanwhile, compared with We-QUATRO, QUATRO, and dual interlocked storage cell (DICE), the write speed of the proposed cell can be reduced by ~41.8 and ~35.3%, and the static power consumption is reduced by ~41.6 and ~46.3%, respectively. Monte Carlo (MC) simulation has proved that under high frequency and low supply (0.6 V) voltage, RHPD-12T has the minimum write failure probability compared with five other SRAM cells. Qiang Zhao 0007, Chunyu Peng, Junning Chen, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | Radiation-Hardened 14T SRAM Bitcell With Speed and Power Optimized for Space ApplicationabstractIn this paper, a novel radiation-hardened 14-transistor SRAM bitcell with speed and power optimized [radiation-hardened with speed and power optimized (RSP)-14T] for space application is proposed. By circuit- and layout-level optimization design in a 65-nm CMOS technology, the 3-D TCAD mixed-mode simulation results show that the novel structure is provided with increased resilience to single-event upset as well as single-event-multiple-node upsets due to the charge sharing among OFF-transistors. Moreover, the HSPICE simulation results show that the write speed and power consumption of the proposed RSP-14T are improved by ~65% and ~50%, respectively, compared with those of the radiation hardened design (RHD)-12T memory cell. Chunyu Peng, Jiati Huang, Changyong Liu, Qiang Zhao 0007, Songsong Xiao, Xiulong Wu, Zhi-Ting Lin, Junning Chen, Xuan Zeng 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |