EDBT 2026 Demo / reviewers in the wild / expert
Keji Zhou
dblp:182/0762
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-3078-4542ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MICLEAR: Intelligent molecular cytology for intraoperative margin assessment of pancreatic ductal adenocarcinomaabstractPancreatic ductal adenocarcinoma (PDAC) is a highly mortal cancer whose only potentially curative treatment is surgical resection. Intraoperative assessment of its surgical margins is vital for patient survival. Frozen section biopsy is routinely performed for this purpose. However, its heavy reliance on pathologists' expertise often leads to diagnostic discrepancies. The inherent invasiveness of PDAC also leads to sampling errors. This study developed an intelligent molecular cytology approach that improves diagnostic objectivity and broadens sampling coverage. Our method, Multi-Instance Cytology with LEArned Raman-embedding (MICLEAR), leverages compositional information from label-free Raman imaging. First, 4085 cells were brushed off from the pancreases of 41 patients and imaged using stimulated Raman scattering microscopy. Then, a contrastive learning-based cell embedding model was developed to compress each cell's morphological and compositional information into a compact cell vector. Finally, a multi-instance learning-based diagnostic model using cell vectors was employed to predict the likelihood that a patient's margin is positive. MICLEAR achieved 80% sensitivity, 100% specificity, and an area under the receiver operating characteristic curve of 0.86 in 27 patients for validation, comprising 10 with positive margins and 17 with negative margins, in approximately 8 minutes per patient. It may hold promise for more efficient and accurate intraoperative assessment of PDAC surgical margins. Tinghe Fang, Dao-Ning Liu, Keji Zhou, Chunyi Hao, Shuhua Yue |
Medical Image Anal. | 4 |
| 2025 | SRIF: Energy- and Area-Efficient Soft Reset Integrate-and-Fire Neuron for High-Accuracy Spiking Neural NetworkabstractIn this work, we present an energy- and area-efficient soft reset integrate-and-fire (SRIF) neuron circuit for spiking neural networks (SNNs), designed to improve inference accuracy by retaining the residual membrane potential above the firing threshold. The conventional current-subtraction soft reset neuron introduces significant performance overhead. In contrast, the proposed SRIF neuron employs the charge-sharing mechanism, characterized by its simple structure and minimal energy overhead. The 28-nm CMOS SRIF neuron exhibits an energy consumption of 14.7 fJ per spike and occupies an area of 107 µm2. Additionally, the soft reset mechanism achieves a 39.5% improvement in inference accuracy compared to the hard reset baseline when benchmarked with VGG-11 network trained on the CIFAR-10 dataset. Moreover, a capacitor-sharing strategy is proposed to improve the utilization efficiency of reset capacitors, providing 1.24x smaller neuron area. Tianci Cai, Qiqiao Wu, Keji Zhou |
ISCAS | 6 |
| 2025 | A Refresh-Reduction Digital eDRAM CIM Macro using Asymmetric Error Tolerance SchemeabstractIn this work, we propose a refresh-reduction 3T1C eDRAM-based digital Compute-In-Memory (CIM) macro with an asymmetric error tolerance scheme to improve the energy consumption caused by frequent refresh operations. The novel asymmetric error tolerance scheme is first implemented in an eDRAM-based CIM macro to mitigate the effects of the wrong weights, extending the refresh period of the eDRAM cell. 3T1C eDRAM cell exhibits precise asymmetric error behavior, perfectly adapting the proposed asymmetric error tolerance scheme. Additionally, a reconfigurable refresh controller is introduced to eliminate refresh performance overhead during CIM operations and further reduce power consumption by reusing CIM readout data. The proposed 128Kb eDRAM-based CIM macro, simulated at 28nm, achieves a 98-100% improvement in the refresh period and a 21.5% improvement in average energy efficiency at 25°C and 60°C. Keji Zhou, Chengshuo Yu, Tianci Cai |
ISCAS | 3 |
| 2025 | A High-Density RRAM-Based Ising Machine with Analog In-Memory Operation for Solving Combinatorial Optimization ProblemsabstractThis work presents a Resistive RAM (RRAM)-based Ising machine characterized by high spin density and efficient analog in-memory computing. The proposed design enables compact spin representations that support interactions with up to eight neighboring spins at 1-bit precision. The functionality and performance of the proposed Ising machine is evaluated using comprehensive simulation in 28nm CMOS technology, incorporating measured RRAM variation characteristics to validate its robustness and efficiency. Additionally, software-based simulations are employed to solve classical combinatorial problems, such as the Max-Cut problem, while accounting for the non-idealities inherent in the proposed RRAM-based analog in-memory computing approach. Each proposed spin occupies an area of 13.5 μm2, achieving an area reduction of 16.7% to 92.1% compared to recent works, based on feature size normalization. Jingxin Deng, Keji Zhou, Honghu Yang, Chengshuo Yu |
ISCAS | 2 |
| 2025 | A High Performance Dual-Wordline RRAM Macro with Replica Bitline Delay Control CircuitabstractIn the conventional RRAM design, differential reading for odd and even bitlines (BLs) facilitates high-speed data retrieval, but this comes at the cost of area overhead due to additional multiplexers, and is prone to causing performance degradation with inaccurate timing signals. In this work, a high performance RRAM macro is presented to solve the above problems through: 1) a compact dual-wordline (WL) array structure that splits the WLs into odd and even pairs, the inactive half of the BLs can be used as differential input without the extra multiplexers, thereby improving the storage density and reducing WL switching power consumption by 45%; 2) a replica BL control circuit to effectively track the BL discharge characteristics and accurately control WL pulse width, thus lowering read energy consumption and latency. A 1Mb RRAM macro with 64-bit bandwidth and 8.85Mb/mm2storage density is implemented using 28nm process, achieving a 3.5ns read cycle and consuming 9fJ per bit during reads, with the FoM (read throughput / area) at least 3.6× higher than prior works. Honghu Yang, Yongkang Han, Tianci Cai, Chengshuo Yu, Keji Zhou |
ISCAS | 5 |
| 2025 | Enhancing All-to-All RRAM Ising Machines With Randomized Granular Update Strategies for Solving Combinatorial Optimization ProblemsabstractIn recent years, Ising machines have emerged as a promising hardware solution for tackling combinatorial optimization problems (COPs). However, existing Ising solvers, whether based on discrete-time or continuous-time approaches, often face challenges in balancing solution quality, scalability, and computational speed. Discrete-time solvers typically suffer from slow convergence due to the sequential nature of spin updates, while continuous-time solvers often lack effective annealing mechanisms, limiting their solution accuracy. To address these limitations, this work proposes a novel architecture that integrates a differential Resistive Random Access Memory (RRAM) cell-based Ising design with a Randomized Granular Update (RAGU) method. This approach enhances scalability to larger spin systems while maintaining robust performance against circuit non-idealities and device variations. Additionally, an adaptive bitline (BL) voltage clamper is incorporated into the read path to limit current magnitudes, significantly improving power efficiency. A key feature of the RAGU method is its ability to naturally introduce randomness during the update process through coarse-grained updates, serving as an imprecise but effective sampling mechanism. This innovation not only accelerates convergence and improves the system’s ability to escape local minima but also ensures high solution quality. Extensive simulations and experiments on randomly weighted graphs with varying densities demonstrate that the proposed architecture consistently achieves near-optimal solutions ($>$96%) while drastically reducing the time-to-solution to as low as 0.6$\mu$s. Qiqiao Wu, Honghu Yang, Chengshuo Yu, Keji Zhou, Haijun Jiang, Hailan Yi, Xiaoyong Xue, Xiaoyang Zeng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | A Novel Neuromorphic Hardware Using Area-Efficient Chain RRAM-Based Synapses and Compact Neurons With (Anti-) Integration SchemeabstractThe neuromorphic system aims to implement large-scale spiking neural networks (SNN) through hardware, ultimately achieving human-level intelligence. At this stage, it is difficult for a single silicon-based chip to reach the density of the human brain, so it is important to improve the area efficiency of neuromorphic chips. We present a novel neuromorphic hardware that employs high area-efficiency chain RRAM synaptic array, and compact RRAM-based neurons. The proposed chain RRAM structure achieves a small cell size of 41.5F2 at 28 nm logic process. Compared to the conventional 1T1R structure, the chain structure has a 22.2% area reduction and a 58.8% parasitic capacitance reduction. As for the neuron design, we use RRAM instead of the capacitor to integrate membrane voltage. To alleviate the endurance issue of RRAM, we propose an (anti-) integration scheme that removes the operation of membrane voltage reset. The simulation results demonstrate an average energy consumption of 1.15pJ per spike and an area of$15.9\mu $m2. With the (anti-)integration scheme, the RRAM’s lifetime is demonstrated to extend by$2\times $and the energy of programming RRAM is demonstrated to be reduced by$2\times $. Qiqiao Wu, Honghu Yang, Yongkang Han, Haijun Jiang, Keji Zhou, Hailan Yi, Qi Liu 0010 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | HARDSEA: Hybrid Analog-ReRAM Clustering and Digital-SRAM In-Memory Computing Accelerator for Dynamic Sparse Self-Attention in TransformerabstractSelf-attention-based transformers have outperformed recurrent and convolutional neural networks (RNN/ CNNs) in many applications. Despite the effectiveness, calculating self-attention is prohibitively costly due to quadratic computation and memory requirements. To solve this challenge, this article proposes a hybrid analog-ReRAM and digital-SRAM in-memory computing accelerator (HARDSEA), a computing-in-memory (CIM) accelerator supporting self-attention in transformer applications. To trade off between energy efficiency and algorithm accuracy, HARDSEA features an algorithm-architecture-circuit codesign. A product-quantization-based scheme dynamically facilitates self-attention sparsity by predicting lightweight token relevance. A hybrid in-memory computing architecture employs both high-efficiency analog ReRAM-CIM and high-precision digital SRAM-CIM to implement the proposed new scheme. The ReRAM-CIM, whose precision is sensitive to circuit nonidealities, takes charge of token relevance prediction where only computing monotonicity is demanded. The SRAM-CIM, utilized for exact sparse attention computing, is reorganized as an on-memory-boundary computing scheme, thus adapting to irregular sparsity patterns. In addition, we propose a time-domain winner-take-all (WTA) circuit to replace the expensive ADCs in ReRAM-CIM macros. Experimental results show that HARDSEA prunes BERT and GPT-2 models to 12%–33% sparsity without accuracy loss, achieving$13.5\times $–$28.5\times $speedup and$291.6\times $–$1894.3\times $energy efficiency over GPU. Compared to state-of-the-art transformer accelerators, HARDSEA has$1.2\times $–$14.9\times $better energy efficiency at the same level of throughput. Shiwei Liu 0002, Chen Mu, Hao Jiang 0024, Yunzhengmao Wang, Jinshan Zhang 0006, Keji Zhou, Qi Liu 0010, Chixiao Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |