EDBT 2026 Demo / reviewers in the wild / expert
Chenghu Dai
dblp:343/2782
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
0009-0000-3347-3056ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 3 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 2RW Dual-Port 8T-SRAM Macro with Bitline Leakage Current Tracking and Read-Write Arbitration
Chenghu Dai, Junbo Chen, Zaihang Zhang, Licai Hao, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu |
ISCAS | 1 |
| 2026 | An offset-compensation capacitor-coupled DRAM sense amplifier with symmetric sensing and high PVT stability
Chenghu Dai, Jing Lv, Yongqi Qin, Chunyu Peng, Xin Li 0099, Yu Liu 0113, Xiulong Wu, Zhi-Ting Lin |
Integr. | 1 |
| 2026 | Analysis and Design of Memory Testing Algorithm for Computing-in-Memory Using MBISTabstractComputing-in-memory (CIM), as a novel computing architecture for the future, effectively overcomes the bottlenecks in the von Neumann architecture. The CIM architecture embeds logic into the memory array to reduce the data transfer between the processor and memory. However, embedding logic into the memory array increases the test complexity. In this study, we offer a comprehensive examination of the challenges associated with CIM and introduce a novel March-like test algorithm, named March CC, tailored for CIM chips. Computational elements are added to the read/write operation sequences, combining the tests in memory mode and computing mode into one step, which significantly improves the test efficiency. In comparison to the traditional March C− test algorithm, the proposed March CC test algorithm, with a complexity of only 10 N , enhances the fault coverage from 66.7% to 79.8% for six common single-cell fault (SCF) models and nine common double-cell fault (DCF) models. Furthermore, the March CC algorithm demonstrates good compatibility and is applicable to various memory configurations, such as SRAM, RRAM, and MRAM CIM architectures. Zhi-Ting Lin, Siyan Li, Qiushi Feng, Changxin Yue, Yuanyang Wang, Yunlong Liu 0006, Yu Liu 0113, Licai Hao, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Xiulong Wu |
ACM J. Emerg. Technol. Comput. Syst. | 14 |
| 2026 | Time-Domain SRAM-CIM Macro With Dual-Edge Temporal Fused Accumulation for Signed 8-bit Precision MAC
Wenjuan Lu, Xiaobo Gong, Kang Meng, Xiaohang Chen, Jiating Guo, Lijun Guan, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu, Chunyu Peng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2026 | A T8T-SRAM Computing-in-Memory Macro for Ternary Deep Neural Networks and Boolean Logic ComputationsabstractDeep neural networks (DNNs) play important roles in artificial intelligence applications and show hungry computility and power demands. Compared with binary neural networks (BNNs), ternary neural networks (TNNs) have higher representation and adaptive abilities and balance the inference accuracy and computing efficiency between DNNs and BNNs. This article proposed a T8T-SRAM computing-in-memory (CIM) macro to achieve Boolean logic operations and MAC operation of ternary activation and ternary weight. The proposed T8T-SRAM bitcell has a separate read and write path, and can avoid the read disturb issue. In Boolean logic operation mode, the T8T-SRAM macro can achievenand,nor,xnor, andxoroperations with redundant rows, reducing the additional reference voltage generation circuit. In the MAC mode, the result is quantized by an embedded column analog-to-digital converter (ADC), which uses activation refresh to reduce weight changing. In 28-nm CMOS technology, under 0.5-V array supply voltage and 0.9-V peripheral supply voltage, simulation results manifest that the MAC results have good linearity, and feasibility of Boolean logic operation. The proposed T8T-SRAM macro realizes MAC operation of 16 ternary activations and 16 ternary weights with 333.99–816.1-TOPS/W energy efficiency and 61.9-TOPS/mm2area efficiency. Using an ResNet-18 network for the inference of MNIST, and CIFAR-10 datasets, the accuracies were 99.06% and 85.76% with a ternary activation and ternary weight. Chenghu Dai, Zihua Ren, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | TSCIM: A 28nm Transposed Stochastic CIM Macro for On-Chip Training and InferenceabstractThis work introduces a novel Transposed Stochastic Computing-in-Memory (TSCIM) macro designed to enhance the efficiency of on-chip training and inference. The macro incorporates a novel stochastic quantization strategy and utilizes a transposed separated wordline SRAM to enable multi-bit signed MAC operations. Furthermore, a stochastic adder tree is utilized to minimize area and power consumption overhead. The design includes a 4Kb SRAM CIM macro implemented in 28 nm CMOS technology. Simulation results show that the power consumption of the stochastic accumulation circuit (SAC) is reduced by 63.6%, while the area overhead is decreased by a factor of 7.73 compared to designs using full adder (FA) adder trees. Additionally, the computation latency is decreased by 16× compared to traditional stochastic circuits. The TSCIM macro can achieve a peak energy efficiency of 63.02 TOPS/W and an area efficiency of 15.54 TOPS/mm2. Yu Liu 0113, Yang Lou, Kangkang Mao, Xin Li 0099, Chenghu Dai, Xiulong Wu, Zhi-Ting Lin |
ISCAS | 5 |
| 2025 | A Low-Cost and Triple-Node-Upset Self-Recoverable Latch Design With Low Soft Error RateabstractWith the decrease in feature size of transistors, latches are more sensitive to single-event multiple node upset (MNU), including double node upset (DNU) and triple node upset (TNU). However, the reported TNU self-recoverable (TNUR) latches are facing problems with large areas and power consumption. Based on the polarity design, this article proposes a low-cost TNUR latch (LCTRL) with a low soft error rate (SER) in 28-nm CMOS technology. The proposed LCTRL mainly consists of four interlocked modules and a clock-gated inverter. Compared with the state-of-the-art TNUR latches, including LCTNURL, IHTRL, FATNU, and TRLW, the power consumption, D-Q delay, CLK-to-Q delay, area, and the power-delay–area product (PDAP) of the proposed LCTRL are reduced by 55.09%, 38.64%, 42.93%, 44.65%, and 83.50%, respectively. Due to the polarity design, the SER of the proposed LCTRL is the smallest among compared latches, which suggests that the proposed LCTRL is suitable for use in radiation environments. Licai Hao, Lang Tian, Hao Wang 0239, Shiyu Zhao 0004, Qiang Zhao 0007, Chunyu Peng, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | A 28-nm 9T1C SRAM-Based CIM Macro With Hierarchical Capacitance Weighting and Two-Step Capacitive Comparison ADCs for CNNsabstractIn the realm of charge-domain computing-in-memory (CIM) macros, reducing the area of capacitor ladder and analog-to-digital converter (ADC) while maintaining high throughput remains a significant challenge. This brief introduces an adjustable-weight CIM macro designed to enhance both energy efficiency and area efficiency for convolutional neural networks (CNNs). The proposed architecture uses: 1) a customized 9T1C bit cell for sensing margin improvement and bidirectional decoupled read ports; 2) a hierarchical capacitance weighting (HCW) structure that achieves a weight accumulation of 1/2/4 bits with less capacitance area and weighting time; and 3) a two-step capacitive comparison ADCs (TC-ADCs) readout scheme to improve area efficiency and throughput. The proposed 8-kb static random address memory (SRAM) CIM macro is implemented using 28-nm CMOS technology. It can achieve an energy efficiency of 224.4 TOPS/W and an area efficiency of 21.894 TOPS/mm2, and the accuracies on MNIST, CIFAR-10, and CIFAR-100 datasets are 99.67%, 89.13%, and 67.58% with a 4-b input and 4-b weight. Zhi-Ting Lin, Runru Yu, Miao Long, Yu Liu 0113, Jianxing Zhou, Qingchuan Zhu, Yue Zhao 0029, Lintao Chen, Chunyu Peng, Qiang Zhao 0007, Xin Li 0099, Chenghu Dai, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 14 |
| 2025 | High-Reliability and High-Throughput CIM 10T-SRAM for Multiplication and Accumulation Operations With 274.3 GOPS and 200-237.5 TOPS/WabstractArtificial intelligence (AI) is extensively applied in natural language processing, image matching, and image recognition, with convolutional neural networks (CNNs) being crucial. Computing-in-memory (CIM) utilizing static random access memory (SRAM) can enhance the CNN performance. However, this faces issues such as multibit signed data processing, read corruption of traditional SRAM arrays, and increased area overhead due to increased capacitor weighting. This article proposes a 10T-SRAM macro tailored for CNN multiply-accumulate calculation (MAC) computation in image processing. It enables high-throughput full-array operations, with added dual ports facilitating input of multibit data with signed bits. The 10T-SRAM cell features a read-write separation channel, mitigating read disturbance issues seen in dual-port 8T-SRAM arrays or 6T-SRAM arrays. Incorporating redundant columns in the array for charge sharing and weighting conserves area and boosts circuit reliability. In the 28-nm CMOS simulation environment, the proposed architecture achieves a throughput of 274.3 GOPS and an energy efficiency of 200–237.5 TOPS/W, surpassing literature-reported figures by several times. Wenjuan Lu, Lubin Xiang, Chunyu Peng, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | A 28 nm Dual-Mode SRAM-CIM Macro With Local Computing Cell for CNNs and Grayscale Edge DetectionabstractWith the rise of artificial intelligence (AI), neural network applications are growing in demand for efficient data transmission. The traditional von Neumann architecture can no longer keep pace with modern technological needs. Computing-in-memory (CIM) is proposed as a promising solution to address this bottleneck. This work introduces a local computing cell (LCC) scheme based on compact 6T-SRAM cells. The proposed circuit aims to enhance energy efficiency and reduce power consumption by reusing the LCC. The LCC circuit can perform the multiplication of a 2-bit input with a 1-bit weight, which can be applied to convolutional neural networks (CNNs) with the multiply-accumulate (MAC) operations. Through circuit reuse, it can also be used for multibit multiply operations, performing 2-bit input multiplication and 1-bit weight addition, which can be applied to grayscale edge detection in images. The energy efficiency of the SRAM-CIM macro achieves an energy efficiency of 46.3 TOPS/W under MAC operations with input precision of 8-bits and weight precision of 8-bits, and up to 389.1–529.1 TOPS/W under the calculation in one subarray with an input precision of 2-bits and a weight precision of 1-bit. The estimated inference accuracy on CIFAR-10 datasets is 90.21%. Chunyu Peng, Xiaohang Chen, Mengya Gao, Jiating Guo, Lijun Guan, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | A 28-nm Cascode Current Mirror-Based Inconsistency-Free Charging-and-Discharging SRAM-CIM Macro for High-Efficient Convolutional Neural NetworksabstractComputing-in-memory (CIM) is an emerging approach to alleviate the von Neumann bottleneck and enhance energy efficiency and throughput. This brief introduces a 16-Kb static random access memory (SRAM) CIM macro for convolutional neural networks (CNNs), featuring a cascode current mirror-based inconsistency-free computing circuits (CICCs). The bias voltage of CICC is provided by a cascode current mirror (CCM) circuit. The proposed architecture improves the consistency and linearity of bitline (BL) charge and discharge rates in the analog current domain, enhancing computational accuracy. Additionally, the charge and discharge on the BLs represent the positive or negative calculation result, eliminating the need for extra encoding and logic circuits to handle sign bits. The SRAM-CIM macro achieves an energy efficiency of 59.1–134.0 TOPS/W and a throughput of 0.41 TOPS in a 28-nm CMOS technology, and the estimated inference accuracy on MNIST and CIFAR-10 datasets is 96.5% and 91.4%, respectively, with 5-bit input precision and 1-bit weight precision. Chunyu Peng, Jiating Guo, Shengyuan Yan, Xiaohang Chen, Wenjuan Lu, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | A CFMB STT-MRAM-Based Computing-in-Memory Proposal With Cascade Computing Unit for Edge AI DevicesabstractThe application of non-volatile memory technology is increasingly attractive for Computing-in-memory (CIM) owing to high integration density and negligible standby power consumption. This study proposes an spin-transfer-torque (STT) magnetic random access memory (MRAM) based CIM macro which incorporates following innovative features: 1) cross-feedback margin-boost (CFMB) scheme to enable robust and fast reading operations against process variation and limited Tunneling Magnetoresistance Ratio (TMR); 2) cascade computing units (CCU) and related design method for efficient and stable multi-bit multiply-and-accumulate (MAC) operation; and 3) dual computing mode scheme and resolution adjustable quantization module to optimize energy efficiency and operating speed. The post-simulations are performed under 28nm CMOS&MTJ technology. The results demonstrate the achievement in energy efficiency of 36.4 TOPS/W while performing MAC operations with up to 16-bit weights, 4-bit inputs, and 22-bit outputs. Yongliang Zhou, Chenghu Dai, Licai Hao, Chunyu Peng, Hao Cai 0001, Xiulong Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | Low-Cost and Highly Robust Quadruple Node Upset Tolerant Latch DesignabstractThis article proposes an exceptionally reliable and low-cost quadruple node upset tolerant latch ($LC$-QNUTL) suitable for the 65 nm CMOS technology. The innovative$LC$-QNUTL latch is primarily composed of three soft-error-immune (SEI) static random-access memory (SRAM) cells and a triple-level C-element (CE) unit, which includes five two-input CE and a clock-gating (CG)-based two-input CE. The SEI SRAM cell utilizes polarity hardening technology and source-isolation technology, significantly reducing the number of sensitive nodes and enhancing the latch’s stability. By using the high-speed transmission gate (TG) technology and stacked structures, the proposed latch offers minimal overhead in terms of delay and power consumption, yielding an improved power delay area product (PDAP). When compared to contemporary quadruple node upset (QNU)-tolerant latch designs (including HLMR, 4NUHL, and LDAVPM), the new design offers substantial improvements—29.53% less delay, 80.09% reduced power consumption, 58.52% smaller silicon area, and 433.43% improved comprehensive PDAP on average. Furthermore, simulation results demonstrate that the$LC$-QNUTL latch exhibits reduced sensitivity to process, voltage, and temperature (PVT) variations, thus providing superior reliability, which makes it an ideal choice for safety-critical applications. Licai Hao, Yaling Wang, Yunlong Liu 0006, Shiyu Zhao 0004, Wenjuan Lu, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 11 |
| 2024 | Soft-Error-Immune Quadruple-Node-Upset Tolerant Latch Based on Polarity Design and Source-Isolation TechnologiesabstractA soft-error-immune quadruple-node-upset tolerant latch (SEI-QNUTL) with a low delay and high performance is proposed using 65-nm CMOS technology. The proposed SEI-QNUTL design consists of three soft-error-immune static random access memory (SEI-SRAM) cells. Furthermore, each SEI-SRAM cell employs polarity design and source-isolation technology to reduce the number of sensitive nodes and enhance the reliability of the latch. Compared with state-of-the-art quadruple-node-upset (QNU) tolerant latches [including high-performance and low-cost single-event multiple-node-upsets resilient (HLMR), QNU tolerant latch (QNUTL), and Latch Design and Algorithm-based Verification Protected against Multiple-Node-Upsets (LDAVPM)], the proposed SEI-QNUTL design reduces (on average) the area, delay, and area-power-delay-product (APDP) by 47.0%, 25.0%, 46.5%, and 66.3%, respectively. Extensive variation analysis validates that the SEI-QNUTL design is less sensitive to process, voltage, and temperature (PVT) variations regarding power consumption and delay. Furthermore, Monte Carlo (MC) simulations show that the proposed latch exhibits high reliability when performing data storage. Compared with the existing latches, the SEI-QNUTL design makes a good tradeoff among delay, power, and area, and it can thus be used in safety-critical applications. Licai Hao, Chenghu Dai, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng, Yongliang Zhou, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |