Jinfeng Kang

dblp:97/7420 · DBLP profile ↗
← Back
28ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2026 A 28nm 143.4-322.5TOPS/W INT8 time-domain CIM macro featuring zero-weight skipping and shift-and-add embedded TDC for Deep Neural Network
Ao Shi, Lianliang Wu, Haobin Shang, Kexun Li, Lifeng Liu, Jinfeng Kang, Peng Huang 0004
ISCAS12
2025 Mitigating methodology of hardware non-ideal characteristics for non-volatile memory based neural networks
Lixia Han, Peng Huang 0004, Yijiao Wang, Haozhang Yang, Jinfeng Kang
Sci. China Inf. Sci.8
2025 CIMUS: 3D-Stacked Computing-in-Memory Under Image Sensor Architecture for Efficient Machine Vision
abstract
Computational image sensors with CNN processing capabilities are emerging to alleviate the energy-intensive and time-consuming data movement between sensors and external processors. However, deploying CNN models onto these computational image sensors faces challenges from the limited on-chip memory resources and insufficient image processing throughput. This work proposes a 3D-stacked NAND flash-based computing-in-memory under image sensor architecture (CIMUS) to facilitate the complete deployment of CNN model. To fully leverage the potential of high bandwidth from the 3D-stacked integration, we design a novel distributed CNN mapping and dataflow to process the full focal plane image in parallel, which senses and recognizes ImageNet tasks with >1000fps. To tackle the computational error of inputs “0” in 3D NAND flash-based CIM, we propose an input-independent offset compensation method, which reduces the average vector-matrix multiplication (VMM) error by 48%. Evaluation results indicate that CIMUS architecture achieves a 9.8× improvement in CNN inference speed and a 33× boost in energy efficiency compared to the state-of-the-art computational image sensor in the ImageNet recognition task.
Lixia Han, Haozhang Yang, Ao Shi, Guihai Yu, Yijiao Wang, Yanzhi Wang 0001, Jinfeng Kang, Peng Huang 0004
IEEE Trans. Computers12
2024 Pipeline Design of Nonvolatile-based Computing in Memory for Convolutional Neural Networks Inference Accelerators
abstract
Nonvolatile-based computing-in-memory inference chips show great potential to accelerate convolutional neural networks. The intrinsic weight stationary characteristic makes pipeline design a crucial solution to further enhance throughput. In this work, we propose a balanced pipeline design and establish performance/area evaluation models for the optimal pipeline solution. The evaluation results indicate that our pipeline design achieves$30\times$computational efficiency improvement.
Lixia Han, Peng Huang 0004, Haozhang Yang, Jinfeng Kang
DATE7
2024 Low Quantization Error Readout Circuit with Fully Charge-Domain Calculation for Computation-in-Memory Deep Neural Network
abstract
This work presents a low quantization error readout circuit with fully-charge-domain calculation for quantization and post-process of computation-in-memory (CIM)-based neural network. The contributions include: (1) A novel residual charge accumulation function is designed to achieve charge-domain summation of quantized partial sum, and reduces 38% quantization error; (2) Charge reset is introduced in the integrate & fire circuit to realize <1 LSB INL at ±7 bits and speed of 285MHz/LSB; (3) Sample & hold, current subtraction and bidirectional counter are designed to improve 3.95× energy efficiency and 2.48× area efficiency.
Ao Shi, Lixia Han, Lifeng Liu, Linxiao Shen, Peng Huang 0004, Jinfeng Kang
ISCAS10
2024 CoMN: Algorithm-Hardware Co-Design Platform for Nonvolatile Memory-Based Convolutional Neural Network Accelerators
abstract
Computing in memory (CIM) convolutional neural network (CNN) accelerators based on nonvolatile memory (NVM) show great potential to improve energy efficiency and throughput, while the multiple design levels and huge design space of CIM-based CNN acceleration system make cross-level co-design methodology and platforms extremely desired. In this work, an algorithm-hardware co-design platform CoMN with the graphic user interface is proposed for designers to fast verify and further optimize the designments. In the platform, 1) a mapper is developed to automatically map CNN models to CIM chips through optimizing pipeline, weight transformation, partition, and placement; 2) accuracy evaluator and performance evaluator are built to jointly estimate accuracy, energy, latency, and area overheads considering the design dependencies across multiple levels; 3) algorithm adapter is exploited to retrain CNN weights for higher hardware accuracy within limited energy budget through nonidealities aware training and energy aware training; 4) hardware optimizer is developed to search hardware microarchitecture and circuit design space in the early design stage. We conduct several case studies to verify the effectiveness of the CoMN platform. Results indicate that CoMN platform can enable algorithm-hardware mapping, hardware-aware algorithm adaption, hardware configuration exploration, and overall algorithm-hardware co-design efficiently. The CoMN platform can be accessed online at http://101.42.97.22:8081/index.html with username “tcad” and password “comnuser”.
Lixia Han, Renjie Pan 0003, Hairuo Lu, Haozhang Yang, Peng Huang 0004, Guangyu Sun 0003, Jinfeng Kang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.10
2024 Specific ADC of NVM-Based Computation-in-Memory for Deep Neural Networks
abstract
Non-volatile memory (NVM)-based Computation-in-memory has demonstrated a significant advantage in high-efficiency neural networks. However, the requirement of analog-to-digital converter (ADC) and post-processing circuits not only cost high energy and area but also results in high computation errors, which tradeoffs the performance boost brought by CIM. Here, we present a specific ADC and post-processing circuit of the NVM-based CIM neural network to address these issues. The main contributions include: (1) A novel residual charge accumulation function (RCA) is designed to achieve charge-domain summation of quantized partial sum and reduces 38% quantization error; (2) Charge reset is introduced in the integrate & fire circuit to realize$3.95\times $energy efficiency and$2.48\times $area efficiency. Evaluation based on the measured results of the fabricated chip shows that the VGG-11 neural network with the proposed ADC circuit can achieve a 3.28-time improvement in energy efficiency while maintaining the same network recognition rate.
Ao Shi, Lixia Han, Haozhang Yang, Lifeng Liu, Linxiao Shen, Jinfeng Kang, Peng Huang 0004
IEEE Trans. Circuits Syst. I Regul. Pap.10
2023 A Convolution Neural Network Accelerator Design with Weight Mapping and Pipeline Optimization
abstract
The pipeline is an efficient solution to boost performance in non-volatile memory based computing in memory (nvCIM) convolution neural network (CNN) accelerators. However, the previous works seldom focus on pipeline optimization from the perspective of the whole system, especially overlooking the effect of buffer access. In this work, we propose a high-performance NVM-based CNN accelerator with a balanced pipeline design, which takes account of both the macro computing and the buffer access. At the operator level, a matrix-based weight mapping method is proposed to reduce buffer access delay. At the macro level, decoupled access and execution design is introduced to shorten the single-layer latency. At the system level, a hybrid inter/intra-tile design is presented to balance the overall latency across CNN layers. With the collaboration among three methods, we construct a well-balanced pipeline for the nvCIM accelerator at a smaller hardware cost. Experiments show that our pipeline design can achieve 3.7х, 7.5х, and 3.5х throughput improvement for recognition of ImageNet with ResNet18, VGG19, and ResNet34 models, respectively.
Lixia Han, Peng Huang 0004, Jinfeng Kang
DAC6
2023 A 3.3-Mbit/s true random number generator based on resistive random access memory
Shiyue Song, Peng Huang 0004, Wensheng Shen, Lifeng Liu, Jinfeng Kang
Sci. China Inf. Sci.5
2023 An ultra-high-density and energy-efficient content addressable memory design based on 3D-NAND flash
Haozhang Yang, Peng Huang 0004, Runze Han, Jinfeng Kang
Sci. China Inf. Sci.5
2023 Co-optimization strategy between array operation and weight mapping for flash computing arrays to achieve high computing efficiency and accuracy
Guihai Yu, Peng Huang 0004, Runze Han, Lixia Han, Jinfeng Kang
Sci. China Inf. Sci.6
2022 Efficient Discrete Temporal Coding Spike-Driven In-Memory Computing Macro for Deep Neural Network Based on Nonvolatile Memory
abstract
Nonvolatile memory (NVM) based neural network can directly perform in situ computation in memory to significantly reduce energy consumption resulting from the data movement. However, the energy consumption by the analog-to-digital converter (ADC) restricts the efficiency of the mixed-signal in-memory computing macro. The rate coding spike-driven in-memory computing macro can increase the energy efficiency via eliminating the ADC, but the improvement is limited because substantial energy is consumed for the coding of multiple spikes. In this work, we propose a discrete temporal coding spike-driven in-memory computing macro, including input coding scheme, weight mapping method, and improved leaky integrate-and-fire (LIF) neuron circuit, to perform the efficient forward inference of deep neural networks based on NVM array. We then optimize the designment of the proposed in-memory computing macro to mitigate the neural network accuracy loss due to the nonlinearity of the LIF neuron and voltage drop caused by interconnect resistance. Because the temporal coding scheme reduces spike numbers and the improved-LIF circuit simultaneously integrates two bit-lines current corresponding to positive and negative weight, the proposed macro achieves 46.63TOPS/W energy efficiency and 1.92TOPS throughput for 3bit temporal coding precision.
Lixia Han, Peng Huang 0004, Yijiao Wang, Jinfeng Kang
IEEE Trans. Circuits Syst. I Regul. Pap.7
2021 A physics-based electromigration reliability model for interconnects lifetime prediction
Linlin Cai, Wangyong Chen, Jinfeng Kang, Gang Du
Sci. China Inf. Sci.3
2021 STAR: Synthesis of Stateful Logic in RRAM Targeting High Area Utilization
abstract
Processing-in-memory (PIM) exploits massive parallelism with high energy efficiency and becomes a promising solution to the von Neumann bottleneck. Recently, the emerging metal-oxide resistive random access memory (RRAM) shows its potential to construct a PIM architecture, because several stateful logic operations, e.g., IMP and NOR, can be executed in an RRAM crossbar in parallel. Previous synthesis flows focus on improving latency with stateful logic operations, but they ignore that the memory should be used primarily for storage. i.e., most of the area in the crossbar is used for computation but not storage. In this situation, storage and computation still have to be separated into different crossbars, which leads to considerable data transfer overhead and limited parallelism. In this work, we define the ratio of storage in a crossbar as area utilization. We aim to improve the area utilization without throughput loss by proposing STAR, a novel synthesis flow for the stateful logic. We present two optimization strategies to reduce the computation area in STAR. First, we reduce the area for redundant inputs. For the shared constants among different rows (or columns), we encode them as immediate values into the control signals without writing them into the crossbar at runtime. For the other inputs, we only store one copy of them in the crossbar. Second, we reduce the area for intermediate variables by reusing invalid cells. And we design a scheduling algorithm to find a computation sequence with the minimal variable erasing cycles. Invalid primary inputs can also be erased in this algorithm. Furthermore, we present a case study of the image convolution to demonstrate the effectiveness of STAR. Experimental evaluation shows that STAR achieves 33.03% more area utilization and a 1.43x throughput compared to SIMPLER, the state-of-the-art stateful logic synthesis flow. Our image convolution implementation also provides 78.36% more area utilization and a 1.48x throughput compared with IMAGING, the state-of-the-art stateful logic-based image processing accelerator.
Feng Wang 0046, Guojie Luo, Guangyu Sun 0003, Jiaxi Zhang 0001, Jinfeng Kang, Yuhao Wang 0002, Dimin Niu, Hongzhong Zheng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 The synthesis method of logic circuits based on the iMemComp gates
Xiaole Cui, Qiujun Lin, Xiaoxin Cui, Jinfeng Kang
Integr.6
2019 Parallel Stateful Logic in RRAM: Theoretical Analysis and Arithmetic Design
abstract
Processing-in-memory (PIM) provides massive parallelism with high energy efficiency and becomes a promising solution to the memory wall problem. Recently, the emerging metal-oxide resistive random access memory (RRAM) has shown its potential to design a PIM architecture. Several stateful logic operations, e.g., NOR and NAND, can be executed in parallel in an RRAM crossbar. Although previous works have designed some algorithms using the stateful logic, it is still under exploration how to fully exploit its potential high parallelism and design an asymptotically fast algorithm for a given function. In this work, we theoretically analyze the parallelism in an RRAM crossbar and design several asymptotically optimal arithmetic algorithms. In detail, we first propose the Single Instruction Multiple Lines (SIML) model to unify the stateful logic families and prove three lower bounds on the time complexity of a parallel RRAM algorithm. Then, we design three algorithms for integer addition functions with the stateful logic, guided by the lower bound analysis. All of them reach the time complexity lower bound. Finally, We make two extensions of the integer addition algorithms, supporting multiplication functions by decomposing them to additions and supporting the flex-point data type by proposing an exponent and mantissa update flow. Experimental evaluation shows that our integer algorithms achieves a speedup up to 13.79x over the previous RRAM algorithms. Our flex-point implementation achieves a 26.60x speedup and saves 73.68% energy compared to an ARM.
Feng Wang 0046, Guojie Luo, Guangyu Sun 0003, Jiaxi Zhang 0001, Peng Huang 0004, Jinfeng Kang
ASAP6
2019 Analog Deep Neural Network Based on NOR Flash Computing Array for High Speed/Energy Efficiency Computation
abstract
In this paper, a novel hardware implementation of analog deep neural network (DNN) based on NOR Flash Computing Array (NFCA) is presented. The approach eliminates additional analog-to-digital/digital-to-analog (AD/DA) conversion between adjacent layers. Applied to the MNIST recognition task, the simulations indicate that the designed DNN based on the novel implementation approach has the excellent performance such as the time delay of 3×10-7s and energy consumption of 1.97×10-8J per image, which brings 8× and 123× enhancements compared to the conventional digital scheme. The NFCA based analog DNN also saves 86.4% of area. All of the improvements benefit from the analog-signal based scheme. The proposed high speed and energy efficient hardware implementation would be promising in terms of artificial intelligence (AI) at the edge.
Yachen Xiang, Peng Huang 0004, Runze Han, Yuning Jiang 0004, Q. M. Shu, Zhiqiang Su, Yongbo Liu, Jinfeng Kang
ISCAS10
2019 Efficient evaluation model including interconnect resistance effect for large scale RRAM crossbar array matrix computing
Runze Han, Peng Huang 0004, Yudi Zhao, Xiaole Cui, Jinfeng Kang
Sci. China Inf. Sci.6
2019 Circuit design of RRAM-based neuromorphic hardware systems for classification and modified Hebbian learning
Yuning Jiang 0004, Peng Huang 0004, Jinfeng Kang
Sci. China Inf. Sci.4
2018 A Novel Convolution Computing Paradigm Based on NOR Flash Array with High Computing Speed and Energy Efficient
abstract
A novel convolution computing paradigm based on the NOR Flash Array is proposed. Significant improvements both in computing speed and energy consumption are achieved compared to CMOS-based logic computing paradigms. Regarding to the feature extraction task from a 256×256 image, the computing speed of 3.9×104frame per second (fps) and the energy consumption of 0.057nJ/pixel are achieved using the proposed computing paradigm.
Runze Han, Peng Huang 0004, Yachen Xiang, Chen Liu 0009, Zhekang Dong, Zhiqiang Su, Yongbo Liu, Jinfeng Kang
ISCAS10
2017 Testing of 1TnR RRAM array with sneak path technique
Xiaole Cui, Xiaoxin Cui, Xin'an Wang, Jinfeng Kang
Sci. China Inf. Sci.5
2015 Modeling and design optimization of ReRAM
abstract
Resistive switching memories (ReRAM) have been widely studied for applications in next-generation data storage and neurormorphic computing systems. To enable device-circuit-system co-design and optimization, a SPICE model of ReRAM that can reproduce the device characteristics in circuit simulations is needed. In this paper, we present a novel tool for ReRAM design including a physics-based SPICE model, the model parameters extraction strategy, as well as the system assessment method. This physics-based SPICE model can capture all the essential features of HfOx-based ReRAM including the DC/AC and multi-level switching behaviors, switching reliability, and intrinsic device variations. A strategy is developed to extract the critical model parameters from the fabricated ReRAM devices. A variety of electrical measurements on various ReRAMs are performed to verify and calibrate the model. The assessment method based on the experimentally verified SPICE model can be applied to explore a wide range of applications including: 1) variation-aware and reliability-emphasized system design; 2) system performance evaluation; 3) array architecture optimization. This verified design tool not only enables system design but also enables system optimization that capitalizes on device/circuit interaction for both data storage and neuromorphic computing applications.
Jinfeng Kang, Haitong Li, Peng Huang 0004, Bin Gao 0006, Zizhen Jiang, H.-S. Philip Wong
ASP-DAC1
2015 Variation-aware, reliability-emphasized design and optimization of RRAM using SPICE model
Haitong Li, Zizhen Jiang, Peng Huang 0004, Hong-Yu Chen, Bin Gao 0006, Jinfeng Kang, H.-S. Philip Wong
DATE8
2015 Doping profile modification approach of the optimization of HfO x based resistive switching device by inserting AlO x layer
Lifeng Liu, Jinfeng Kang
Sci. China Inf. Sci.6
2014 Scaling and operation characteristics of HfOx based vertical RRAM for 3D cross-point architecture
abstract
Stacked HfOxbased vertical RRAM with interface engineering for 3D cross-point architecture is fabricated using a cost-effective fabrication process. The excellent performances such as low reset current, fast switching speed, high switching endurance and disturbance immunity, good retention and self-selectivity are demonstrated in the fabricated HfOxbased vertical RRAM devices. The scaling limit and the functionality along with a viable write/read scheme of the presented vertical RRAM are investigated. The experiments show that the pillar electrode thickness and the plane electrode thickness of the vertical RRAM can be scaled down to 3nm and 5nm without significant performance degradation, respectively.
Jinfeng Kang, Bin Gao 0006, Peng Huang 0004, Hong-Yu Chen, Zizhen Jiang, H.-S. Philip Wong, Shimeng Yu
ISCAS1
2014 Design guidelines for 3D RRAM cross-point architecture
abstract
Design guidelines were proposed to evaluate and optimize the 3D RRAM cross-point architecture by a full-size 3D circuit simulation in SPICE. The performance metrics that were evaluated include the write/read margin, access latency, energy consumption per programming, and the density per bit. Different 3D cross-point architecture including the horizontally stacked or the vertically stacked structure were compared in terms of these metrics, revealing the advantages of the vertical RRAM structure. Then the scaling trend of the vertical RRAM based 3D array with respect to the scaling of lateral feature size, vertical electrode thickness and vertical isolation layer thickness were evaluated. The design parameters that affect the scaling trend include the metal interconnect resistance, RRAM on-state cell resistance (or the nonlinearity of the I-V). The design trade-offs are discussed considering those parameters constraints.
Shimeng Yu, Yexin Deng, Bin Gao 0006, Peng Huang 0004, Jinfeng Kang, Hong-Yu Chen, Zizhen Jiang, H.-S. Philip Wong
ISCAS7
2010 A novel voltage-type sense amplifier for low-power nonvolatile memories
Jinfeng Kang, Yangyuan Wang
Sci. China Inf. Sci.2
2009 Challenges of 22 nm and beyond CMOS technology
Ru Huang 0001, HanMing Wu, Jinfeng Kang, Deyuan Xiao, XueLong Shi, Xia An, Runsheng Wang, Xing Zhang 0002, Yangyuan Wang
Sci. China Ser. F Inf. Sci.3