EDBT 2026 Demo / reviewers in the wild / expert
Peng Huang 0004
dblp:29/1726-4
· DBLP profile ↗
26ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0003-3280-0099ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Re-RIS: A Reconfigurable 3D RRAM In-Sensor Architecture for Low-Latency Machine VisionabstractCutting-edge machine vision applications impose stringent latency and energy efficiency demands on edge devices. To address these demands, In-Sensor Computing (ISC) architectures aim to eliminate data movement overhead, while 3D RRAM technology provides the hardware foundation of high memory density and massive computing parallelism. However, existing ISC architectures rely on static resource allocation, failing to address the dynamic "shifting bottleneck" in CNNs— where early layers are compute-bound and later layers are readout-bound. To address this, we propose Re-RIS, a Reconfigurable 3D RRAM In-Sensor architecture. By dynamically switching hardware granularity between high-parallelism and high-throughput modes, Re-RIS optimizes resource utilization for varying layer characteristics. Experimental results on VGG-16 demonstrate an end-to-end latency of 0.93 ms, achieving a 75% reduction compared to static baselines, with an energy efficiency of 244.6 TOPS/W and an area efficiency of 1.85 TOPS/mm2. Lixia Han, Lifeng Liu, Peng Huang 0004 |
DATE | 5 |
| 2026 | A 28nm 143.4-322.5TOPS/W INT8 time-domain CIM macro featuring zero-weight skipping and shift-and-add embedded TDC for Deep Neural Network
Ao Shi, Lianliang Wu, Haobin Shang, Kexun Li, Lifeng Liu, Jinfeng Kang, Peng Huang 0004 |
ISCAS | 13 |
| 2026 | RRAM based CIM, PUF and True Random Number Generator for High-Security AES Encryption System
Kefan Tao, Shiyue Song, Yading Yi, Ao Shi, Haokai Guan, Lianliang Wu, Hao Ai, Lifeng Liu, Yulin Feng, Peng Huang 0004 |
ISCAS | 10 |
| 2026 | RRAM-based CAM for Energy-Efficient In-Memory Text Compression System
Lianliang Wu, Hao Ai, Ao Shi, Haokai Guan, Kexun Li, Kefan Tao, Yulin Feng, Zongwei Wang 0001, Yimao Cai, Peng Huang 0004 |
ISCAS | 14 |
| 2026 | Instant-CIM: An Instant Neural Radiance Field Computing-In-Memory Architecture for Low-Power and Real-Time AR/VR RenderingabstractNovel View Synthesis is a foundational technique for creating immersive Augmented and Virtual Reality (AR/VR) experiences, aiming to generate photorealistic images of a scene from arbitrary camera viewpoints using only a limited set of source images, with Neural Radiance Fields (NeRF) emerging as the state-of-the-art solution. However, real-time NeRF rendering on low-power devices remains challenging due to its memory-intensive hash encoding and compute-intensive Multilayer Perception (MLP). In this work, we propose Instant-CIM, the fully on-chip Computing-in-Memory (CIM) architecture for efficient NeRF rendering. At the algorithm level, Instant-CIM proposes a spatially-adaptive framework that dynamically selects the number of active hash encoding levels per spatial region based on a composite importance score derived from density and gradient. The approach replaces uniform level allocation with a threshold-based strategy that activates finer encoding levels only in regions with high representation complexity. At the hardware level, Instant-CIM proposes an in-situ hash engine that implements in-memory hash query and interpolation through 3D scene grid decomposition and Z-order based mapping schemes. Meanwhile, Instant-CIM proposes a sparse MLP engine that leverages differential-based input complemented by a precision-adjustable skipping mechanism to fully exploit spatial similarities. Comprehensive evaluation across synthetic datasets demonstrates that Instant-CIM achieves 3.0×~4.9× improvement in rendering speed and 8.6×~33× enhancement in energy efficiency compared to state-of-the-art NeRF architecture. Lixia Han, Hui Chen 0015, Xueming Fu, Ke Chen 0018, Peng Huang 0004, Yijun Cui, Weiqiang Liu 0001 |
IEEE Trans. Computers | 7 |
| 2025 | Mitigating methodology of hardware non-ideal characteristics for non-volatile memory based neural networks
Lixia Han, Peng Huang 0004, Yijiao Wang, Haozhang Yang, Jinfeng Kang |
Sci. China Inf. Sci. | 2 |
| 2025 | CIMUS: 3D-Stacked Computing-in-Memory Under Image Sensor Architecture for Efficient Machine VisionabstractComputational image sensors with CNN processing capabilities are emerging to alleviate the energy-intensive and time-consuming data movement between sensors and external processors. However, deploying CNN models onto these computational image sensors faces challenges from the limited on-chip memory resources and insufficient image processing throughput. This work proposes a 3D-stacked NAND flash-based computing-in-memory under image sensor architecture (CIMUS) to facilitate the complete deployment of CNN model. To fully leverage the potential of high bandwidth from the 3D-stacked integration, we design a novel distributed CNN mapping and dataflow to process the full focal plane image in parallel, which senses and recognizes ImageNet tasks with >1000fps. To tackle the computational error of inputs “0” in 3D NAND flash-based CIM, we propose an input-independent offset compensation method, which reduces the average vector-matrix multiplication (VMM) error by 48%. Evaluation results indicate that CIMUS architecture achieves a 9.8× improvement in CNN inference speed and a 33× boost in energy efficiency compared to the state-of-the-art computational image sensor in the ImageNet recognition task. Lixia Han, Haozhang Yang, Ao Shi, Guihai Yu, Yijiao Wang, Yanzhi Wang 0001, Jinfeng Kang, Peng Huang 0004 |
IEEE Trans. Computers | 13 |
| 2025 | An Efficient Flash-Based Computing-in-Memory (CIM) Demonstration of High-Precision (32-bit) Nonlinear Partial Differential Equation (PDE) Solver With Ultra-High Endurance and ReliabilityabstractSolving partial differential equations (PDEs) requires precise numerical iterations that impose significant demands on computational resources and memory capacities, which can be addressed by adopting computing-in-memory (CIM) architecture to reduce the latency and power consumption during data transmission. Among PDEs, nonlinear PDEs present heightened complexities in both analytical investigations and numerical simulations as the presence of nonlinear terms introduces intricate dynamics and mathematical intricacies. The high-precision requirements of PDE solvers, particularly for nonlinear PDE solvers, pose challenges in constructing CIM PDE solvers. In this work, a flash-based high-precision PDE solver has been demonstrated to solve the intractable nonlinear partial differential equation. It’s based on 55nm NOR flash technology with well-optimized Program/Erase (PE) schemes. Utilizing the proposed optimization scheme, the PE endurance can be largely enhanced up to$10^{10}$cycles, which is a record high with suppressed cell degradation and robust reliabilities. Then, applying the Fourier neural operator (FNO) to the optimized flash-based high-precision CIM (32-bit) in the hardware system, a series of nonlinear PDEs can be solved with ~2TOPS/W high energy efficiency, which is$\sim 110\times $higher than CPU. Our optimization strategies make it feasible to use flash-based CIM for high-precision computing with frequent weight updating and the demonstrated PDE solver provides an energy-saving solution to implement general-purpose computation tasks. Zhaohui Sun, Junyao Mei, Yueran Qi, Jing Liu 0035, Xuepeng Zhan, Peng Huang 0004, Jiezhi Chen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 11 |
| 2024 | Pipeline Design of Nonvolatile-based Computing in Memory for Convolutional Neural Networks Inference AcceleratorsabstractNonvolatile-based computing-in-memory inference chips show great potential to accelerate convolutional neural networks. The intrinsic weight stationary characteristic makes pipeline design a crucial solution to further enhance throughput. In this work, we propose a balanced pipeline design and establish performance/area evaluation models for the optimal pipeline solution. The evaluation results indicate that our pipeline design achieves$30\times$computational efficiency improvement. Lixia Han, Peng Huang 0004, Haozhang Yang, Jinfeng Kang |
DATE | 2 |
| 2024 | Low Quantization Error Readout Circuit with Fully Charge-Domain Calculation for Computation-in-Memory Deep Neural NetworkabstractThis work presents a low quantization error readout circuit with fully-charge-domain calculation for quantization and post-process of computation-in-memory (CIM)-based neural network. The contributions include: (1) A novel residual charge accumulation function is designed to achieve charge-domain summation of quantized partial sum, and reduces 38% quantization error; (2) Charge reset is introduced in the integrate & fire circuit to realize <1 LSB INL at ±7 bits and speed of 285MHz/LSB; (3) Sample & hold, current subtraction and bidirectional counter are designed to improve 3.95× energy efficiency and 2.48× area efficiency. Ao Shi, Lixia Han, Lifeng Liu, Linxiao Shen, Peng Huang 0004, Jinfeng Kang |
ISCAS | 8 |
| 2024 | CoMN: Algorithm-Hardware Co-Design Platform for Nonvolatile Memory-Based Convolutional Neural Network AcceleratorsabstractComputing in memory (CIM) convolutional neural network (CNN) accelerators based on nonvolatile memory (NVM) show great potential to improve energy efficiency and throughput, while the multiple design levels and huge design space of CIM-based CNN acceleration system make cross-level co-design methodology and platforms extremely desired. In this work, an algorithm-hardware co-design platform CoMN with the graphic user interface is proposed for designers to fast verify and further optimize the designments. In the platform, 1) a mapper is developed to automatically map CNN models to CIM chips through optimizing pipeline, weight transformation, partition, and placement; 2) accuracy evaluator and performance evaluator are built to jointly estimate accuracy, energy, latency, and area overheads considering the design dependencies across multiple levels; 3) algorithm adapter is exploited to retrain CNN weights for higher hardware accuracy within limited energy budget through nonidealities aware training and energy aware training; 4) hardware optimizer is developed to search hardware microarchitecture and circuit design space in the early design stage. We conduct several case studies to verify the effectiveness of the CoMN platform. Results indicate that CoMN platform can enable algorithm-hardware mapping, hardware-aware algorithm adaption, hardware configuration exploration, and overall algorithm-hardware co-design efficiently. The CoMN platform can be accessed online at http://101.42.97.22:8081/index.html with username “tcad” and password “comnuser”. Lixia Han, Renjie Pan 0003, Hairuo Lu, Haozhang Yang, Peng Huang 0004, Guangyu Sun 0003, Jinfeng Kang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | Specific ADC of NVM-Based Computation-in-Memory for Deep Neural NetworksabstractNon-volatile memory (NVM)-based Computation-in-memory has demonstrated a significant advantage in high-efficiency neural networks. However, the requirement of analog-to-digital converter (ADC) and post-processing circuits not only cost high energy and area but also results in high computation errors, which tradeoffs the performance boost brought by CIM. Here, we present a specific ADC and post-processing circuit of the NVM-based CIM neural network to address these issues. The main contributions include: (1) A novel residual charge accumulation function (RCA) is designed to achieve charge-domain summation of quantized partial sum and reduces 38% quantization error; (2) Charge reset is introduced in the integrate & fire circuit to realize$3.95\times $energy efficiency and$2.48\times $area efficiency. Evaluation based on the measured results of the fabricated chip shows that the VGG-11 neural network with the proposed ADC circuit can achieve a 3.28-time improvement in energy efficiency while maintaining the same network recognition rate. Ao Shi, Lixia Han, Haozhang Yang, Lifeng Liu, Linxiao Shen, Jinfeng Kang, Peng Huang 0004 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 11 |
| 2023 | A Convolution Neural Network Accelerator Design with Weight Mapping and Pipeline OptimizationabstractThe pipeline is an efficient solution to boost performance in non-volatile memory based computing in memory (nvCIM) convolution neural network (CNN) accelerators. However, the previous works seldom focus on pipeline optimization from the perspective of the whole system, especially overlooking the effect of buffer access. In this work, we propose a high-performance NVM-based CNN accelerator with a balanced pipeline design, which takes account of both the macro computing and the buffer access. At the operator level, a matrix-based weight mapping method is proposed to reduce buffer access delay. At the macro level, decoupled access and execution design is introduced to shorten the single-layer latency. At the system level, a hybrid inter/intra-tile design is presented to balance the overall latency across CNN layers. With the collaboration among three methods, we construct a well-balanced pipeline for the nvCIM accelerator at a smaller hardware cost. Experiments show that our pipeline design can achieve 3.7х, 7.5х, and 3.5х throughput improvement for recognition of ImageNet with ResNet18, VGG19, and ResNet34 models, respectively. Lixia Han, Peng Huang 0004, Jinfeng Kang |
DAC | 2 |
| 2023 | A 3.3-Mbit/s true random number generator based on resistive random access memory
Shiyue Song, Peng Huang 0004, Wensheng Shen, Lifeng Liu, Jinfeng Kang |
Sci. China Inf. Sci. | 2 |
| 2023 | An ultra-high-density and energy-efficient content addressable memory design based on 3D-NAND flash
Haozhang Yang, Peng Huang 0004, Runze Han, Jinfeng Kang |
Sci. China Inf. Sci. | 2 |
| 2023 | Co-optimization strategy between array operation and weight mapping for flash computing arrays to achieve high computing efficiency and accuracy
Guihai Yu, Peng Huang 0004, Runze Han, Lixia Han, Jinfeng Kang |
Sci. China Inf. Sci. | 2 |
| 2022 | Efficient Discrete Temporal Coding Spike-Driven In-Memory Computing Macro for Deep Neural Network Based on Nonvolatile MemoryabstractNonvolatile memory (NVM) based neural network can directly perform in situ computation in memory to significantly reduce energy consumption resulting from the data movement. However, the energy consumption by the analog-to-digital converter (ADC) restricts the efficiency of the mixed-signal in-memory computing macro. The rate coding spike-driven in-memory computing macro can increase the energy efficiency via eliminating the ADC, but the improvement is limited because substantial energy is consumed for the coding of multiple spikes. In this work, we propose a discrete temporal coding spike-driven in-memory computing macro, including input coding scheme, weight mapping method, and improved leaky integrate-and-fire (LIF) neuron circuit, to perform the efficient forward inference of deep neural networks based on NVM array. We then optimize the designment of the proposed in-memory computing macro to mitigate the neural network accuracy loss due to the nonlinearity of the LIF neuron and voltage drop caused by interconnect resistance. Because the temporal coding scheme reduces spike numbers and the improved-LIF circuit simultaneously integrates two bit-lines current corresponding to positive and negative weight, the proposed macro achieves 46.63TOPS/W energy efficiency and 1.92TOPS throughput for 3bit temporal coding precision. Lixia Han, Peng Huang 0004, Yijiao Wang, Jinfeng Kang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2019 | Parallel Stateful Logic in RRAM: Theoretical Analysis and Arithmetic DesignabstractProcessing-in-memory (PIM) provides massive parallelism with high energy efficiency and becomes a promising solution to the memory wall problem. Recently, the emerging metal-oxide resistive random access memory (RRAM) has shown its potential to design a PIM architecture. Several stateful logic operations, e.g., NOR and NAND, can be executed in parallel in an RRAM crossbar. Although previous works have designed some algorithms using the stateful logic, it is still under exploration how to fully exploit its potential high parallelism and design an asymptotically fast algorithm for a given function. In this work, we theoretically analyze the parallelism in an RRAM crossbar and design several asymptotically optimal arithmetic algorithms. In detail, we first propose the Single Instruction Multiple Lines (SIML) model to unify the stateful logic families and prove three lower bounds on the time complexity of a parallel RRAM algorithm. Then, we design three algorithms for integer addition functions with the stateful logic, guided by the lower bound analysis. All of them reach the time complexity lower bound. Finally, We make two extensions of the integer addition algorithms, supporting multiplication functions by decomposing them to additions and supporting the flex-point data type by proposing an exponent and mantissa update flow. Experimental evaluation shows that our integer algorithms achieves a speedup up to 13.79x over the previous RRAM algorithms. Our flex-point implementation achieves a 26.60x speedup and saves 73.68% energy compared to an ARM. Feng Wang 0046, Guojie Luo, Guangyu Sun 0003, Jiaxi Zhang 0001, Peng Huang 0004, Jinfeng Kang |
ASAP | 5 |
| 2019 | Analog Deep Neural Network Based on NOR Flash Computing Array for High Speed/Energy Efficiency ComputationabstractIn this paper, a novel hardware implementation of analog deep neural network (DNN) based on NOR Flash Computing Array (NFCA) is presented. The approach eliminates additional analog-to-digital/digital-to-analog (AD/DA) conversion between adjacent layers. Applied to the MNIST recognition task, the simulations indicate that the designed DNN based on the novel implementation approach has the excellent performance such as the time delay of 3×10-7s and energy consumption of 1.97×10-8J per image, which brings 8× and 123× enhancements compared to the conventional digital scheme. The NFCA based analog DNN also saves 86.4% of area. All of the improvements benefit from the analog-signal based scheme. The proposed high speed and energy efficient hardware implementation would be promising in terms of artificial intelligence (AI) at the edge. Yachen Xiang, Peng Huang 0004, Runze Han, Yuning Jiang 0004, Q. M. Shu, Zhiqiang Su, Yongbo Liu, Jinfeng Kang |
ISCAS | 2 |
| 2019 | Efficient evaluation model including interconnect resistance effect for large scale RRAM crossbar array matrix computing
Runze Han, Peng Huang 0004, Yudi Zhao, Xiaole Cui, Jinfeng Kang |
Sci. China Inf. Sci. | 2 |
| 2019 | Circuit design of RRAM-based neuromorphic hardware systems for classification and modified Hebbian learning
Yuning Jiang 0004, Peng Huang 0004, Jinfeng Kang |
Sci. China Inf. Sci. | 2 |
| 2018 | A Novel Convolution Computing Paradigm Based on NOR Flash Array with High Computing Speed and Energy EfficientabstractA novel convolution computing paradigm based on the NOR Flash Array is proposed. Significant improvements both in computing speed and energy consumption are achieved compared to CMOS-based logic computing paradigms. Regarding to the feature extraction task from a 256×256 image, the computing speed of 3.9×104frame per second (fps) and the energy consumption of 0.057nJ/pixel are achieved using the proposed computing paradigm. Runze Han, Peng Huang 0004, Yachen Xiang, Chen Liu 0009, Zhekang Dong, Zhiqiang Su, Yongbo Liu, Jinfeng Kang |
ISCAS | 2 |
| 2015 | Modeling and design optimization of ReRAMabstractResistive switching memories (ReRAM) have been widely studied for applications in next-generation data storage and neurormorphic computing systems. To enable device-circuit-system co-design and optimization, a SPICE model of ReRAM that can reproduce the device characteristics in circuit simulations is needed. In this paper, we present a novel tool for ReRAM design including a physics-based SPICE model, the model parameters extraction strategy, as well as the system assessment method. This physics-based SPICE model can capture all the essential features of HfOx-based ReRAM including the DC/AC and multi-level switching behaviors, switching reliability, and intrinsic device variations. A strategy is developed to extract the critical model parameters from the fabricated ReRAM devices. A variety of electrical measurements on various ReRAMs are performed to verify and calibrate the model. The assessment method based on the experimentally verified SPICE model can be applied to explore a wide range of applications including: 1) variation-aware and reliability-emphasized system design; 2) system performance evaluation; 3) array architecture optimization. This verified design tool not only enables system design but also enables system optimization that capitalizes on device/circuit interaction for both data storage and neuromorphic computing applications. Jinfeng Kang, Haitong Li, Peng Huang 0004, Bin Gao 0006, Zizhen Jiang, H.-S. Philip Wong |
ASP-DAC | 3 |
| 2015 | Variation-aware, reliability-emphasized design and optimization of RRAM using SPICE model
Haitong Li, Zizhen Jiang, Peng Huang 0004, Hong-Yu Chen, Bin Gao 0006, Jinfeng Kang, H.-S. Philip Wong |
DATE | 3 |
| 2014 | Scaling and operation characteristics of HfOx based vertical RRAM for 3D cross-point architectureabstractStacked HfOxbased vertical RRAM with interface engineering for 3D cross-point architecture is fabricated using a cost-effective fabrication process. The excellent performances such as low reset current, fast switching speed, high switching endurance and disturbance immunity, good retention and self-selectivity are demonstrated in the fabricated HfOxbased vertical RRAM devices. The scaling limit and the functionality along with a viable write/read scheme of the presented vertical RRAM are investigated. The experiments show that the pillar electrode thickness and the plane electrode thickness of the vertical RRAM can be scaled down to 3nm and 5nm without significant performance degradation, respectively. Jinfeng Kang, Bin Gao 0006, Peng Huang 0004, Hong-Yu Chen, Zizhen Jiang, H.-S. Philip Wong, Shimeng Yu |
ISCAS | 4 |
| 2014 | Design guidelines for 3D RRAM cross-point architectureabstractDesign guidelines were proposed to evaluate and optimize the 3D RRAM cross-point architecture by a full-size 3D circuit simulation in SPICE. The performance metrics that were evaluated include the write/read margin, access latency, energy consumption per programming, and the density per bit. Different 3D cross-point architecture including the horizontally stacked or the vertically stacked structure were compared in terms of these metrics, revealing the advantages of the vertical RRAM structure. Then the scaling trend of the vertical RRAM based 3D array with respect to the scaling of lateral feature size, vertical electrode thickness and vertical isolation layer thickness were evaluated. The design parameters that affect the scaling trend include the metal interconnect resistance, RRAM on-state cell resistance (or the nonlinearity of the I-V). The design trade-offs are discussed considering those parameters constraints. Shimeng Yu, Yexin Deng, Bin Gao 0006, Peng Huang 0004, Jinfeng Kang, Hong-Yu Chen, Zizhen Jiang, H.-S. Philip Wong |
ISCAS | 4 |