VLDB 2026 Research / reviewers in the wild / expert
Wente Yi
dblp:372/5033
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0002-2949-3577ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Operator-Circuit Co-design Digital SOT-MRAM Computing-in-Memory Accelerator with Double Bit Density and Full-Utilized Bandwidth/ThroughputabstractComputing-in-Memory (CIM) demonstrates exceptional performance on edge AI applications, owing to its in-situ computation capability with minimal data transfer consumption. However, volatile CIMs suffer from inevitable data retention power overhead, while non-volatile MRAM-CIMs still necessitate periodic weight updates constrained by limited memory space, diminishing the intrinsic advantage of CIMs. In this work, we propose a digital SOT-MRAM CIM accelerator with circuit-architecture-operator cross-layer design, achieving double bit density and full utilization of both data transmission bandwidth and computing throughput, thereby satisfying the stringent hardware demands for edge AI applications. Firstly, we propose a refined 2T-1MTJ non-complementary memory cell with an XOR-integrated pre-charged sense amplifier (X-SA), which significantly promotes the storage density and consumes only 6.284 fJ per read-based XOR operation. Then, we devise a channel-flatten data mapping (CFDM) scheme and an operator-aware residual fusion (OARF) structure to full utilize the storage and computing resources. Furthermore, an operator fusion method towards non-linear layers is proposed, achieving an 89.84% size reduction in non-binary parameters. System-level simulations at 40nm demonstrate that our work achieves 284.25 TOPS/W energy efficiency and 5.41 TOPS/mm2area efficiency with an accuracy of 98.72% (87.78%) on MNIST (CIFAR-10) dataset. Tianshuo Bai, Jingcheng Gu, Lehao Tan, Wente Yi, Haolin Ge, Zhenyu Xue, He Zhang 0011, Na Lei, Biao Pan |
DATE | 4 |
| 2026 | High-Efficiency and Low-Deviation Analog-Digital Hybrid Compute-in-Memory Architecture With Dynamic Weight DivisionabstractCompute-in-memory (CIM) reduces data movement but suffers from an accuracy–efficiency trade-off: Analog CIM (ACIM) is energy-efficient but loses accuracy and incurs higher cost at large bit-widths, while digital CIM (DCIM) supports high precision but is inefficient for low-precision tasks. To overcome these challenges, we propose an analog–digital hybrid CIM (HCIM) architecture to address this trade-off, including 1) an analog–digital hybrid 10T SRAM cell without additional transistors and a dual-capacitor-based multicycle weighting module to reduce area; 2) a successive-approximation-register (SAR) ADC with a pseudo C-2C capacitor array that can be reconfigured from an 8-bit ADC into two parallel 4-bit ADCs to improve configurability; 3) configurable weight division and computing resource allocation strategies. Simulations in a 28-nm process show that HCIM achieves 15.56 TOPS/W at 12-bit ($8+4$) with$1.33\times $and$2.35\times $efficiency improvement over DCIM and ACIM and$16\times $lower error. It achieves 27.87 TOPS/W at 8-bit and 78.13 TOPS/W at 4-bit, demonstrating superior energy efficiency, computational accuracy, and flexibility. Linjun Jiang, Sifan Sun, Wente Yi, Dengwen Li, Wang Kang 0001, He Zhang 0011, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | A 24.65 TOPS/W@INT8 Hybrid Analog-Digital Multi-core SRAM CIM Macro with Optimal Weight Dividing and Resource Allocation StrategiesabstractCompute-in-memory (CIM) technology integrates memory and computation to reduce memory bottlenecks in modern systems. However, current CIM architectures face challenges in balancing accuracy and energy efficiency. Analog-CIM (ACIM) is energy-efficient but less accurate, while Digital-CIM (DCIM) is accurate but consumes more energy. In this paper, we propose a novel multi-core hybrid analog-digital CIM macro that effectively addresses this trade-off. Our approach intelligently allocates computation tasks to ACIM and DCIM cores based on their accuracy requirements, achieving a balance of accuracy and efficiency. Additionally, we developed an optimization framework to determine the optimal weight divide ratio and computing resource allocation for the hybrid CIM. Experimental results demonstrate the efficacy of our approach. The proposed hybrid CIM achieves an outstanding energy efficiency of 24.65 TOPS/W at 8-bit precision, surpassing DCIM by a factor of 1.33 while maintaining a low error rate of only 0.4%, which is 30 times better than ACIM at the same precision. Wente Yi, Sifan Sun, Wenjia Wang 0011, Jinyu Bai, He Zhang 0011, Wang Kang 0001 |
ASP-DAC | 2 |
| 2025 | PAR-CIM: A Precise/Approximate Reconfigurable Digital CIM Macro with 0.35-4b Fractional Mixed-Bitwidth QuantizationabstractDigital computing-in-memory (DCIM) enables efficient deep neural networks (DNNs) acceleration but faces limitations in resource overhead, energy efficiency, and architectural flexibility. Existing approximate or reconfigurable DCIM solutions tackle these issues partially without achieving a holistic balance. To address this, we propose PAR-CIM, a highly energy-efficient reconfigurable CIM macro that integrates precise and approximate paradigm. First, we introduce layer/gate-level approximate computation (LGAC) into the adder tree (AT) of the DCIM core, achieving full operation with only 0.35× the area of traditional implementations. Then, we develop a 0.35-4b fractional mixed-bitwidth quantization (FMBQ) algorithm, combining second-order Taylor sensitivity analysis with DoReFa-Net. This is complemented by a high-precision low-approximation (HPLA) mapping scheme to enhance energy efficiency. Additionally, a multi-bit reconfigurable computation mode (MBRM) strategy further improves architectural flexibility and enables the implementation of the proposed design. Under 40nm technology, PAR-CIM achieves 3048 TOPS/W at 1b/1b operations. With FMBQ, ResNet18 and our custom V-FuseMBA trained on CIFAR-10 achieve over 86.61% compression with accuracy loss under 0.74%, reaching classification accuracies of 93.67% and 92.86%, respectively. Zhenyu Xue, Wente Yi, Tianshuo Bai, Lehao Tan, Jingcheng Gu, Weijie Ding, Wang Kang 0001, Biao Pan |
ICCAD | 3 |
| 2025 | An FPGA Processor Combining Point Cloud and SNN for DVS-based ADAS ApplicationabstractAutomatic Emergency Braking (AEB) has become an important component in Advanced Driver Assistance Systems (ADAS) and a potential solution for AEB lies in the integration of Dynamic Vision Sensor (DVS) with Spiking Neural Network (SNN). A high-precision behavioural recognition algorithm called Spikepoint has been proposed by us, which combines Point Cloud with SNN to enable recognition of DVS event data. This work concentrates on the FPGA implementation of Spikepoint, aiming to improve real-time recognition capabilities. The deployment of Spikepoint on FPGA encounters two challenges: 1) Point Cloud processing introduces additional latency 2) Storing parameters that require to be accessed frequently from DDR introduces a significant time overhead. In order to address challenges aforementioned, a novel reference point-based filtering technique for Point Cloud is introduced. Meanwhile, a fine-grained quantization method and other optimization strategies are used on the neuron model. The Xilinx UltraScale+ is employed in the experiments conducted in this work. Our Point-based SNN Processor achieves a recognition frame rate of 92.08 FPS through the novel algorithm and corresponding hardware optimization, while achieving an accuracy of 94.3% on the DVS128 Gesture dataset. Wente Yi, Kexun Cheng, Lehao Tan, Bojun Cheng, Biao Pan |
ISCAS | 3 |
| 2024 | PipeCIM: A High-Throughput Computing-In-Memory Microprocessor With Nested Pipeline and RISC-V Extended InstructionsabstractThe large number of multiply accumulate (MAC) operations in Convolutional Neural Network (CNN) leads to substantial data migration and computation. Although computing-in-memory (CIM) proves to be a promising paradigm for MAC operations, high throughput CNN accelerator still confronts bottlenecks from: the low MAC utilization and the uncessary off-chip memory access. In this paper, we propose a high throughput CIM-based CNN accelerator PipeCIM with three hierarchies of pipelines: Intra-Macro, Near-Memory and Tile-Level. The Intra-Macro Pipeline parallelly executes data transfer and in-memory-computing (IMC) operations. The Near-Memory Pipeline alleviates memory access for pooling and data reshaping. The Tile-Level Pipeline establishes a layer-wise pipeline to further improve the throughput while reducing control complexity. PipeCIM introduces the nested scheme and a Unidirectional Divergent Connection Protocol (UDTCP) to simplify the control of data flow with the help of customized RISC-V instructions. To validate our design, PipeCIM was prototyped in 55 nm process node, achieving energy efficiency of 133.8 TOPS/W and peak throughput of 819 GOPS with a 16KB CIM array, which can accelerate VGG-16 to 128.56$\times$or Inception to 19.754$\times$compared to the baseline. Tingran Chen, Wenjia Wang 0011, Haotian Fu, Wente Yi, Bojun Cheng, He Zhang 0011, Biao Pan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | RDCIM: RISC-V Supported Full-Digital Computing-in-Memory Processor With High Energy Efficiency and Low Area OverheadabstractDigital computing-in-memory (DCIM) that merges computing logic into memory has been proven to be an efficient architecture for accelerating multiply-and-accumulates (MACs). However, low energy efficiency and high area overhead pose a primary restriction for integrating DCIM in re-configurable processors required for multi-functional workloads. To alleviate this dilemma, a novel RISC-V supported full-digital computing-in-memory processor (RDCIM) is designed and fabricated with 55nm CMOS technology. In RDCIM, an adding-on-memory-boundary (AOMB) scheme is adopted to improve the energy efficiency of DCIM. Meanwhile, a multi-precision adaptive accumulator (MPAA) and a serial-parallel conversion supported SRAM buffer (SPBUF) are employed to reduce the area overhead caused by the peripheral circuits and the intermediate buffer for multi-precision support. The results show that the energy efficiency in our design is 16.6 TOPS/W (8-bit) and 66.3 TOPS/W (4-bit). Compared to related works, the proposed RDCIM macro shows a maximum energy efficiency improvement of 1.22$\times$in a continuous computing scenario, an area saving of 1.22$\times$in the accumulator, and an area saving of 3.12$\times$in the input buffer. Moreover, in RDCIM, 5 fine-grained RISC-V extended instructions are designed to dynamically adjust the state of DCIM, reaching 1.2$\times$computation efficiency. Wente Yi, Kefan Mo, Wenjia Wang 0011, Yejun Zeng, Zihan Yuan, Bojun Cheng, Biao Pan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |