EDBT 2026 Demo / reviewers in the wild / expert
Chengshuo Yu
dblp:264/0432
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-0897-7871ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fully Parallel Ising Machine with MRAM based p-bit Probabilistic Simulated Annealing
Tianhao Chen, Tianqu Hu, Shan Yao, Chengshuo Yu, Tianxiao Nie, Xufeng Kou |
ISCAS | 4 |
| 2025 | A Refresh-Reduction Digital eDRAM CIM Macro using Asymmetric Error Tolerance SchemeabstractIn this work, we propose a refresh-reduction 3T1C eDRAM-based digital Compute-In-Memory (CIM) macro with an asymmetric error tolerance scheme to improve the energy consumption caused by frequent refresh operations. The novel asymmetric error tolerance scheme is first implemented in an eDRAM-based CIM macro to mitigate the effects of the wrong weights, extending the refresh period of the eDRAM cell. 3T1C eDRAM cell exhibits precise asymmetric error behavior, perfectly adapting the proposed asymmetric error tolerance scheme. Additionally, a reconfigurable refresh controller is introduced to eliminate refresh performance overhead during CIM operations and further reduce power consumption by reusing CIM readout data. The proposed 128Kb eDRAM-based CIM macro, simulated at 28nm, achieves a 98-100% improvement in the refresh period and a 21.5% improvement in average energy efficiency at 25°C and 60°C. Keji Zhou, Chengshuo Yu, Tianci Cai |
ISCAS | 4 |
| 2025 | A High-Density RRAM-Based Ising Machine with Analog In-Memory Operation for Solving Combinatorial Optimization ProblemsabstractThis work presents a Resistive RAM (RRAM)-based Ising machine characterized by high spin density and efficient analog in-memory computing. The proposed design enables compact spin representations that support interactions with up to eight neighboring spins at 1-bit precision. The functionality and performance of the proposed Ising machine is evaluated using comprehensive simulation in 28nm CMOS technology, incorporating measured RRAM variation characteristics to validate its robustness and efficiency. Additionally, software-based simulations are employed to solve classical combinatorial problems, such as the Max-Cut problem, while accounting for the non-idealities inherent in the proposed RRAM-based analog in-memory computing approach. Each proposed spin occupies an area of 13.5 μm2, achieving an area reduction of 16.7% to 92.1% compared to recent works, based on feature size normalization. Jingxin Deng, Keji Zhou, Honghu Yang, Chengshuo Yu |
ISCAS | 4 |
| 2025 | A High Performance Dual-Wordline RRAM Macro with Replica Bitline Delay Control CircuitabstractIn the conventional RRAM design, differential reading for odd and even bitlines (BLs) facilitates high-speed data retrieval, but this comes at the cost of area overhead due to additional multiplexers, and is prone to causing performance degradation with inaccurate timing signals. In this work, a high performance RRAM macro is presented to solve the above problems through: 1) a compact dual-wordline (WL) array structure that splits the WLs into odd and even pairs, the inactive half of the BLs can be used as differential input without the extra multiplexers, thereby improving the storage density and reducing WL switching power consumption by 45%; 2) a replica BL control circuit to effectively track the BL discharge characteristics and accurately control WL pulse width, thus lowering read energy consumption and latency. A 1Mb RRAM macro with 64-bit bandwidth and 8.85Mb/mm2storage density is implemented using 28nm process, achieving a 3.5ns read cycle and consuming 9fJ per bit during reads, with the FoM (read throughput / area) at least 3.6× higher than prior works. Honghu Yang, Yongkang Han, Tianci Cai, Chengshuo Yu, Keji Zhou |
ISCAS | 4 |
| 2025 | Enhancing All-to-All RRAM Ising Machines With Randomized Granular Update Strategies for Solving Combinatorial Optimization ProblemsabstractIn recent years, Ising machines have emerged as a promising hardware solution for tackling combinatorial optimization problems (COPs). However, existing Ising solvers, whether based on discrete-time or continuous-time approaches, often face challenges in balancing solution quality, scalability, and computational speed. Discrete-time solvers typically suffer from slow convergence due to the sequential nature of spin updates, while continuous-time solvers often lack effective annealing mechanisms, limiting their solution accuracy. To address these limitations, this work proposes a novel architecture that integrates a differential Resistive Random Access Memory (RRAM) cell-based Ising design with a Randomized Granular Update (RAGU) method. This approach enhances scalability to larger spin systems while maintaining robust performance against circuit non-idealities and device variations. Additionally, an adaptive bitline (BL) voltage clamper is incorporated into the read path to limit current magnitudes, significantly improving power efficiency. A key feature of the RAGU method is its ability to naturally introduce randomness during the update process through coarse-grained updates, serving as an imprecise but effective sampling mechanism. This innovation not only accelerates convergence and improves the system’s ability to escape local minima but also ensures high solution quality. Extensive simulations and experiments on randomly weighted graphs with varying densities demonstrate that the proposed architecture consistently achieves near-optimal solutions ($>$96%) while drastically reducing the time-to-solution to as low as 0.6$\mu$s. Qiqiao Wu, Honghu Yang, Chengshuo Yu, Keji Zhou, Haijun Jiang, Hailan Yi, Xiaoyong Xue, Xiaoyang Zeng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | A Dual 7T SRAM-Based Zero-Skipping Compute- In-Memory Macro With 1-6b Binary Searching ADCs for Processing Quantized Neural NetworksabstractThis article presents a novel dual 7T static random-access memory (SRAM)-based compute-in-memory (CIM) macro for processing quantized neural networks. The proposed SRAM-based CIM macro decouples read/write operations and employs a zero-input/weight skipping scheme. A 65nm test chip with$528\times 128$integrated dual 7T bitcells demonstrated reconfigurable precision multiply and accumulate operations with$384\times $binary inputs (0/1) and$384\times 128$programmable multi-bit weights (3/7/15-levels). Each column comprises$384\times $bitcells for a dot product,$48\times $bitcells for offset calibration, and$96\times $bitcells for binary-searching analog-to-digital conversion. The analog-to-digital converter (ADC) converts a voltage difference between two read bitlines (i.e., an analog dot-product result) to a 1-6b digital output code using binary searching in 1-6 conversion cycles using replica bitcells. The test chip with 66Kb embedded dual SRAM bitcells was evaluated for processing neural networks, including the MNIST image classifications using a multi-layer perceptron (MLP) model with its layer configuration of 784-256-256-256-10. The measured classification accuracies are 97.62%, 97.65%, and 97.72% for the 3, 7, and 15 level weights, respectively. The accuracy degradations are only 0.58 to 0.74% off the baseline with software simulations. For the VGG6 model using the CIFAR-10 image dataset, the accuracies are 88.59%, 88.21%, and 89.07% for the 3, 7, and 15 level weights, with degradations of only 0.6 to 1.32% off the software baseline. The measured energy efficiencies are 258.5, 67.9, and 23.9 TOPS/W for the 3, 7, and 15 level weights, respectively, measured at 0.45/0.8V supplies. Chengshuo Yu, Haoge Jiang, Junjie Mu, Kevin Tshun Chuan Chai, Tony Tae-Hyoung Kim, Bongjin Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | 282-to-607 TOPS/W, 7T-SRAM Based CiM with Reconfigurable Column SAR ADC for Neural Network ProcessingabstractCompute in memory ($C$iM) is a promising solution for solving the bottleneck of frequent data interface between memory and processor in Von-Neumann architecture. In this work, a hybrid current/charge domain 7T-SRAM based CiM architecture is proposed to mitigate the PVT-induced RBL variation during computation and thus offer a better linearity without significant impact on the operating frequency and area efficiency. Additionally, a column-referenced 1b to 5b reconfigurable SAR ADC is proposed to support multi-bit output. The proposed design is verified by the Monte-Carlo simulations using 40nm CMOS technology. The 5b mode ADC transferred MAC curve's DNL (LSB) ranges from −0.025 to 0.02 and INL (LSB) ranges from −0.13 to 0.25. The largest RBL variation$(\sigma)$from MAC value −64 to MAC value +64 is 2.08 mV, resulting in a MNIST classification accuracy of 97.5%, which is only 0.1% degradation and Google Speech Command classification accuracy of 80.5%, which is only 0.5% degradation compared to the software baseline, respectively. The whole architecture offers energy efficiency of 282-to-607 TOPS/W for 1-5b output in the MAC operation, which is competitive when compared to other state-of-art$C$iM architectures. Qibang Zang, Wang Ling Goh, Lu Lu 0013, Chengshuo Yu, Junjie Mu, Tony Tae-Hyoung Kim, Bongjin Kim, Dongrui Li, Anh-Tuan Do |
ISCAS | 4 |
| 2023 | A 1-16b Reconfigurable 80Kb 7T SRAM-Based Digital Near-Memory Computing Macro for Processing Neural NetworksabstractThis work introduces a digital SRAM-based near-memory compute macro for DNN inference, improving on-chip weight memory capacity and area efficiency compared to state-of-the-art digital computing-in-memory (CIM) macros. A$20\times 256.1$-16b reconfigurable digital computing near-memory (NM) macro is proposed, supporting a reconfigurable 1-16b precision through the bit-serial computing scheme and the weight and input gating architecture for sparsity-aware operations. Each reconfigurable column MAC comprises$16\times $custom-designed 7T SRAM bitcells to store 1-16b weights, a conventional 6T SRAM for zero weight skip control, a bitwise multiplier, and a full adder with a register for partial-sum accumulations.$20\times $parallel partial-sum outputs are post-accumulated to generate a sub-partitioned output feature map, which will be concatenated to produce the final convolution result. Besides, pipelined array structure improves the throughput of the proposed macro. The proposed near-memory computing macro implements an 80Kb binary weight storage in a 0.473mm2 die area using 65nm. It presents the area/energy efficiency of 4329-270.6 GOPS/mm2 and 315.07-1.23TOPS/W at 1-16b precision. Junjie Mu, Chengshuo Yu, Tony Tae-Hyoung Kim, Bongjin Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | A Logic-Compatible eDRAM Compute-In-Memory With Embedded ADCs for Processing Neural NetworksabstractA novel 4T2C ternary embedded DRAM (eDRAM) cell is proposed for computing a vector-matrix multiplication in the memory array. The proposed eDRAM-based compute-in-memory (CIM) architecture addresses a well-known Von Neumann bottle-neck in the traditional computer architecture and improves both latency and energy in processing neural networks. The proposed ternary eDRAM cell takes a smaller area than prior SRAM-based bitcells using 6-12 transistors. Nevertheless, the compact eDRAM cell stores a ternary state (-1, 0, or +1), while the SRAM bitcells can only store a binary state. We also present a method to mitigate the compute accuracy degradation issue due to device mismatches and variations. Besides, we extend the eDRAM cell retention time to 200μs by adding a custom metal capacitor at the storage node. With the improved retention time, the overall energy consumption of eDRAM macro, including a regular refresh operation, is lower than most of prior SRAM-based CIM macros. A 128×128 ternary eDRAM macro computes a vector-matrix multiplication between a vector with 64 binary inputs and a matrix with 64 × 128 ternary weights. Hence, 128 outputs are generated in parallel. Note that both weight and input bit-precisions are programmable for supporting a wide range of edge computing applications with different performance requirements. The bit-precisions are readily tunable by assigning a variable number of eDRAM cells per weight or adding multiple pulses to input. An embedded column ADC based on replica cells sweeps the reference level for 2N-1 cycles and converts the analog accumulated bitline voltage to a 1-5bit digital output. A critical bitline accumulate operation is simulated (Monte-Carlo, 3K runs). It shows the standard deviation of 2.84% that could degrade the classification accuracy of the MNIST dataset by 0.6% and the CIFAR-10 dataset by 1.3% versus a baseline with no variation. The simulated energy is 1.81fJ/operation, and the energy efficiency is 552.5-17.8TOPS/W (for 1-5bit ADC) at 200MHz using 65nm technology. Chengshuo Yu, Taegeun Yoo, Tony Tae-Hyoung Kim, Kevin Tshun Chuan Chai, Bongjin Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |