VLDB 2026 Research / reviewers in the wild / expert
Dengfeng Wang
dblp:01/10190
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reinforcement learning enhanced evolutionary algorithms for inverse design of electric cargo truck frames
Dengfeng Wang, Zihao Meng |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Multi-axial vibration fatigue optimization strategy based on the artificial intelligence algorithm
Boqiang Zhang, Dengfeng Wang, Zihao Meng, Fengmin Lian, Jialin Dong, Haijun Ruan |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | DBP-CIM: Energy-Efficient 8T SRAM-Based Diagonal-Block Parallel Computing-in-Memory With Compact Data Layout for Arithmetic OperationsabstractIn this work, an energy-efficient bit-parallel static random-access memory (SRAM)-based computing-in-memory (SRAM-CIM) is proposed for general-purpose in-memory arithmetic operations to adapt diverse computing tasks. A compact diagonal-block parallel (DBP) mapping scheme and a novel arithmetic flow are proposed to address the hardware underutilization issue in conventional two-sided stationary bit-parallel CIM architectures. Specifically, the DBP mapping method is implemented to enhance the throughput by reorganizing the intermediate and final results into diagonal memory blocks, effectively reducing the vacant CIM cells caused by the dynamic bit width during computing. In addition, the proposed hardware-efficient arithmetic flows employ: 1) a pipelined ADD scheme to reduce the critical path latency in near-memory computing units; and 2) shift-based arithmetic operations that halve the hardware resources required for multiplication and division while reducing energy consumption. The post-layout simulations on 28-nm CMOS technology show that the proposed DBP-CIM achieves higher energy efficiency and throughput for general-purpose arithmetic operations, compared with state-of-the-art works. Furthermore, evaluations on the general-purpose benchmarks demonstrate that the DBP-CIM reduces energy consumption and computing cycles by up to 55.9% and 60.9%, compared to the conventional bit-parallel CIM. Dengfeng Wang, Chengjun Chang, Weifeng He, Guanghui He 0002, Yanan Sun 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2026 | A Heterogeneous CIM Architecture With Splittable Nonvolatile Computing-in-SRAM Cell Pairs Enabling Efficient On-Chip Neural Network InferenceabstractHeterogeneous computing-in-memory (CIM) offers a promising solution for efficient neural network (NN) accelerations by leveraging the characteristics of different types of memories. However, the previous heterogeneous CIM either requires additional data transfer between different types of isolated memories or solely relies on in situ embedded single-level (SL) or three-level (TL) nonvolatile memories (NVMs), making it difficult to trade off between storage density and robustness benefits. In this article, a new heterogeneous CIM architecture (NVS-SPT) with enhanced storage density and inference robustness is proposed to enable full on-chip acceleration of practical-scale NNs while scalable to larger models. A splittable nonvolatile computing-in-static random access memory (SRAM) cell pairs (nvS2RAM-CIM) is proposed with hybrid in situ embedded SL and TL resistive random access memory (ReRAM) groups, allowing flexible configuration as split or linked state to enhance storage density and restore yield. A layerwise hybrid-coding search (LHCS) algorithm with bitwise and tritwise data-aware mapping (BTM) method is proposed to determine the optimal weight coding patterns with high array utilizations. In addition, a merged hybrid-coding block (MHCB) generation scheme is employed to enable high computing parallelism by merging the dense computing patterns. The proposed NVS-SPT demonstrates up to$4.2\times $higher storage density compared with previous heterogeneous CIM with pure SL-ReRAMs and achieves up to 44.7% enhanced NN accuracy, compared with previous unified ternary coding. Furthermore, the proposed NVS-SPT exhibits up to$1.72\times $and$1.44\times $enhanced energy efficiency with$3.10\times $and$1.42\times $higher computing density, compared with previous heterogeneous CIM based on pure SL- or TL-ReRAMs, respectively. Dengfeng Wang, Liukai Xu, Weifeng He, Guanghui He 0002, Xueqing Li 0002, Yanan Sun 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | Corrigendum to "Multi-objective optimization of lightweight and crashworthiness of automotive front-end structures based on stacked regression surrogate" [Eng. Appl. Artific. Intellig. 155 (2025) 111138]
Fengmin Lian, Dengfeng Wang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Multi-objective optimization of lightweight and crashworthiness of automotive front-end structures based on stacked regression surrogate
Fengmin Lian, Dengfeng Wang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Hop-CIM: An all-digital two-level approximate SRAM-CIM macro for high energy-efficient HNN acceleration with data-aware early exit and column-wise partial-sum reuse
Shunqin Cai, Liukai Xu, Dengfeng Wang, Keqing Ouyang, Weizhong Wu, Zhi Li 0058, Yanan Sun 0003 |
Integr. | 4 |
| 2025 | Robust Monolithic 3D Carbon-Based Computing-in-SRAM With Variation-Aware Bit-Wise Data-Mapping for High-Performance and Integration DensityabstractBit-serial computing-in memory with SRAM cells (SRAM-CIM) enables a full set of integer and floating-point arithmetic operations and various data-intensive computations. Carbon nanotube field-effect transistors (CN-MOSFETs) with high scalability, energy-efficiency, and low process thermal budget are attractive to realize high-dense monolithic three-dimensional (M3D) SRAM-CIM. However, CN-MOSFETs possess unique process variations with asymmetric spatial correlations which can significantly influence the performance and reliability of carbon-based SRAM-CIM. In this paper, new M3D-4N4P SRAM-CIM cells with CN-MOSFETs are proposed with optimized profiles for achieving ultra-high integration density while preserving robustness of data-access and computation. Furthermore, the variation-aware bit-wise data-mapping method is proposed for enhancing the performance of carbon-based SRAM-CIM by leveraging the spatial correlations of CN-MOSFETs. By minimizing the area skew of vertically-stacked layers, the areas of proposed M3D-4N4P SRAM-CIM cells are reduced by up to 50.32% compared to the previous 6N2P SRAM-CIM cells assuming carbon nanotube transistor technology. The proposed M3D-4N4P SRAM-CIM array also achieves by up to$2.17\times $higher throughput on arithmetic operations and 18.34% lower computing latency with 25.36% reduced energy consumptions on MAC-based benchmarks, respectively, compared to the previous 2D-6N2P SRAM-CIM array. Dengfeng Wang, Weifeng He, Qin Wang 0009, Hailong Jiao, Yanan Sun 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | HEIRS: Hybrid Three-Dimension RRAM- and SRAM-CIM Architecture for Multi-task Transformer AccelerationabstractLarge-scale transformer with millions of weights achieves great success in multiple natural language processing (NLP) tasks. To release the memory bottleneck of multi-task model deployment, transfer learning tunes part of weights with shared parameters among tasks. Moreover, computing-in-memory (CIM) emerges as an efficient solution for neural network acceleration. With higher storage density, RRAM-CIM can store the large-scale model without costly weight loading, compared with another mainstream SRAM-CIM. However, the RRAM rewrite for tuned and dynamic weight matrix-vector-multiplication (MVM) in transformers requires high-cost RRAM writing in RRAM-CIM. Current hybrid CIM can compensate for the weakness of RRAM-CIM by adding SRAM-CIM with independent MVM. However, the tuned weights in transfer learning cannot be implemented due to the demand for the cooperative addition of MVM results from both shared and tuned weights. In this paper, a hybrid three-dimension RRAM-CIM and SRAM-CIM architecture (HEIRS) is proposed for multi-task transformer acceleration, with monolithically 3D integration of high-density RRAM-CIM and high-performance SRAM-CIM. The 3D RRAM-CIM with ultra-high density stores the whole model with mitigated off-chip weight loading. The SRAM-CIM is employed for efficiently performing dynamic weight MVM without RRAM rewrite. Moreover, a novel hybrid-CIM paradigm is proposed with an input selective adder tree, to support cooperative addition in transfer learning. Experiments show that, compared with RRAM-CIM and SRAM-CIM, the proposed HEIRS improves the energy efficiency by up to 7.83x and 2.29x on BERT, respectively. Meanwhile, the latency is also reduced by up to 85.5% and the storage density is enhanced by 7.2x, compared to RRAM-CIM. Liukai Xu, Shuai Yuan 0016, Dengfeng Wang, Xueqing Li 0002, Yanan Sun 0003 |
DAC | 3 |
| 2023 | TL-nvSRAM-CIM: Ultra-High-Density Three-Level ReRAM-Assisted Computing-in-nvSRAM with DC-Power Free Restore and Ternary MAC OperationsabstractAccommodating all the weights on-chip for large-scale NNs remains a great challenge for SRAM based computing-in-memory (SRAM-CIM) with limited on-chip capacity. Previous non-volatile SRAM-CIM (nvSRAM-CIM) addresses this issue by integrating high-density single-level ReRAMs on the top of high-efficiency SRAM-CIM for weight storage to eliminate the off-chip memory access. However, previous SL-nvSRAM-CIM suffers from poor scalability for an increased number of SL-ReRAMs and limited computing efficiency. To overcome these challenges, this work proposes an ultra-high-density three-level ReRAMs-assisted computing-in-nonvolatile-SRAM (TL-nvSRAM-CIM) scheme for large NN models. The clustered n-selector-n-ReRAM (cluster-nSnRs) is employed for reliable weight-restore with eliminated DC power. Furthermore, a ternary SRAM-CIM mechanism with differential computing scheme is proposed for energy-efficient ternary MAC operations while preserving high NN accuracy. The proposed TL-nvSRAM-CIM achieves 7.8x higher storage density, compared with the state-of-art works. Moreover, TL-nvSRAM-CIM shows up to 2.9x and 2.0x enhanced energy efficiency, respectively, compared to the baseline designs of SRAM-CIM and ReRAM-CIM, respectively. Dengfeng Wang, Liukai Xu, Songyuan Liu, Zhi Li 0058, Weifeng He, Xueqing Li 0002, Yanan Sun 0003 |
ICCAD | 1 |
| 2023 | CREAM: Computing in ReRAM-Assisted Energy- and Area-Efficient SRAM for Reliable Neural Network AccelerationabstractSRAM-based computing-in-memory (CIM) has been widely explored to accelerate neural networks (NNs). However, it is challenging to store all weights of many modern NNs due to limited on-chip SRAM capacity. This bottleneck induces a large amount of off-chip DRAM accesses and impedes the improvement of performance and energy efficiency. This paper proposes a new approach of computing in resistive random-access memory (ReRAM)-assisted energy- and area-efficient SRAM (CREAM) for accelerating large-scale NNs while eliminating the DRAM access. The NN weights are all stored in high-density on-chip ReRAMs and restored to the proposed non-volatile SRAM (nvSRAM) CIM cells with array-level parallelism. Furthermore, to deal with the influence of ReRAM and CMOS variations, a novel layer-wise and bit-wise weight-configuration search algorithm is proposed by leveraging different sensitivity of each layer in NN models. A data-aware weight-mapping method is also presented to efficiently map NN models to ReRAMs in CREAM for high computation parallelism. The experiment results show$10.3\times $weight storage density over the standard 6T SRAM array. Evaluations of ResNet-18 and VGG-9 on CIFAR-10/CIFAR-100 datasets show up to$3.47\times $and$1.70\times $energy efficiency over two baseline designs of SRAM-CIM and ReRAM-CIM, respectively, in addition to 15.6% higher accuracy than ReRAM-CIM under device variations. Yanan Sun 0003, Dengfeng Wang, Liukai Xu, Zhi Li 0058, Songyuan Liu, Weifeng He, Yongpan Liu, Huazhong Yang, Xueqing Li 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | CREAM: computing in ReRAM-assisted energy and area-efficient SRAM for neural network accelerationabstractComputing-in-memory has been widely explored to accelerate DNN. However, most existing CIM cannot store all NN weights due to limited SRAM capacity for edge AI devices, inducing a large amount off-chip DRAM access. In this paper, a new computing in ReRAM-assisted energy and area-efficient SRAM (CREAM) is proposed for implementing large-scale NNs while eliminating off-chip DRAM access. The weights of DNN are all stored in the high-dense on-chip ReRAM devices and restored to the proposed nvSRAM-CIM cells with array-level parallelism. A data-aware weight-mapping method is also proposed to enhance the CIM performance while fully exploiting the hardware utilization. Experiment results show that the proposed CREAM scheme enhances the storage density by up to 7.94x compared to the traditional SRAM arrays. The energy-efficiency of proposed CREAM is also enhanced by 2.14x and 1.99x, compared to the traditional SRAM-CIM with off-chip DRAM access and ReRAM-CIM circuits, respectively. Liukai Xu, Songyuan Liu, Zhi Li 0058, Dengfeng Wang, Yanan Sun 0003, Xueqing Li 0002, Weifeng He |
DAC | 4 |