VLDB 2026 Research / reviewers in the wild / expert
Biao Pan
dblp:23/8379
· DBLP profile ↗
18ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0002-9524-7617ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Operator-Circuit Co-design Digital SOT-MRAM Computing-in-Memory Accelerator with Double Bit Density and Full-Utilized Bandwidth/ThroughputabstractComputing-in-Memory (CIM) demonstrates exceptional performance on edge AI applications, owing to its in-situ computation capability with minimal data transfer consumption. However, volatile CIMs suffer from inevitable data retention power overhead, while non-volatile MRAM-CIMs still necessitate periodic weight updates constrained by limited memory space, diminishing the intrinsic advantage of CIMs. In this work, we propose a digital SOT-MRAM CIM accelerator with circuit-architecture-operator cross-layer design, achieving double bit density and full utilization of both data transmission bandwidth and computing throughput, thereby satisfying the stringent hardware demands for edge AI applications. Firstly, we propose a refined 2T-1MTJ non-complementary memory cell with an XOR-integrated pre-charged sense amplifier (X-SA), which significantly promotes the storage density and consumes only 6.284 fJ per read-based XOR operation. Then, we devise a channel-flatten data mapping (CFDM) scheme and an operator-aware residual fusion (OARF) structure to full utilize the storage and computing resources. Furthermore, an operator fusion method towards non-linear layers is proposed, achieving an 89.84% size reduction in non-binary parameters. System-level simulations at 40nm demonstrate that our work achieves 284.25 TOPS/W energy efficiency and 5.41 TOPS/mm2area efficiency with an accuracy of 98.72% (87.78%) on MNIST (CIFAR-10) dataset. Tianshuo Bai, Jingcheng Gu, Lehao Tan, Wente Yi, Haolin Ge, Zhenyu Xue, He Zhang 0011, Na Lei, Biao Pan |
DATE | 10 |
| 2026 | SuperCIM: A Charge-domain CIM with 3D Parallelism Reconfigurability based on Asynchronous NoC Protocol
Yuxuan Ran, Tingran Chen, Bohan Xu, Biao Pan |
ISCAS | 4 |
| 2026 | STM-CIM: A 2427 TOPS/W Signed Compute-in-Memory with Analog-Domain Top-k and Matrix Transpose for CNN & Transformer
Tingran Chen, Shuo Liang, Weijie Ding, Biao Pan |
ISCAS | 7 |
| 2026 | Event-Based Motion Deblurring via Multi-Temporal Granularity FusionabstractConventional frame-based cameras inevitably produce blurry effects due to motion occurring during the exposure time. Event camera, a bio-inspired sensor offering continuous visual information could enhance the deblurring performance. Effectively utilizing the high-temporal-resolution event data is crucial for extracting precise motion information and enhancing deblurring performance. However, existing event-based image deblurring methods usually utilize voxel-based event representations, losing the fine-grained temporal details that are mathematically essential for fast motion deblurring. In this paper, we first introduce point cloud-based event representation into the image deblurring task and propose a Multi-Temporal Granularity Network (MTGNet). It combines the spatially dense but temporally coarse-grained voxel-based event representation and the temporally fine-grained but spatially sparse point cloud-based event. To seamlessly integrate such complementary representations, we design a Fine-grained Point Branch. An Aggregation and Mapping Module (AMM) is proposed to align the low-level point-based features with frame-based features and an Adaptive Feature Diffusion Module (AFDM) is designed to manage the resolution discrepancies between event data and image data by enriching the sparse point feature. Extensive subjective and objective evaluations demonstrate that our method outperforms current state-of-the-art approaches on both synthetic and real-world datasets. Our code is available at: https://github.com/xplin13/MTGNet. Xiaopeng Lin, Yulong Huang 0001, Zunchang Liu, Yue Zhou 0010, Haotian Fu, Biao Pan, Bojun Cheng |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | PAR-CIM: A Precise/Approximate Reconfigurable Digital CIM Macro with 0.35-4b Fractional Mixed-Bitwidth QuantizationabstractDigital computing-in-memory (DCIM) enables efficient deep neural networks (DNNs) acceleration but faces limitations in resource overhead, energy efficiency, and architectural flexibility. Existing approximate or reconfigurable DCIM solutions tackle these issues partially without achieving a holistic balance. To address this, we propose PAR-CIM, a highly energy-efficient reconfigurable CIM macro that integrates precise and approximate paradigm. First, we introduce layer/gate-level approximate computation (LGAC) into the adder tree (AT) of the DCIM core, achieving full operation with only 0.35× the area of traditional implementations. Then, we develop a 0.35-4b fractional mixed-bitwidth quantization (FMBQ) algorithm, combining second-order Taylor sensitivity analysis with DoReFa-Net. This is complemented by a high-precision low-approximation (HPLA) mapping scheme to enhance energy efficiency. Additionally, a multi-bit reconfigurable computation mode (MBRM) strategy further improves architectural flexibility and enables the implementation of the proposed design. Under 40nm technology, PAR-CIM achieves 3048 TOPS/W at 1b/1b operations. With FMBQ, ResNet18 and our custom V-FuseMBA trained on CIFAR-10 achieve over 86.61% compression with accuracy loss under 0.74%, reaching classification accuracies of 93.67% and 92.86%, respectively. Zhenyu Xue, Wente Yi, Tianshuo Bai, Lehao Tan, Jingcheng Gu, Weijie Ding, Wang Kang 0001, Biao Pan |
ICCAD | 9 |
| 2025 | ACSNN: A 61.25 TOPS/W, 1.65 ns delay SNN Processor that combines CIM-inspired Synapse and Asynchronous ArchitectureabstractSpiking Neural Networks (SNNs) offer their biological plausibility and dynamic sensitivity which have gained significant attention in real-time systems. However, designing SNN processors with high throughput and real-time response remains challenging due to high power consumption, high area costs and routing competition. In this paper, we present ACSNN, a processor that integrates a CIM-inspired synapse array, Leaky Integrate-and-Fire (LIF) neurons and asynchronous architecture to achieve high parallelism, high energy and area efficiency while declining routing competition. Experimental results demonstrate an impressive power efficiency of 61.25 TOPS/W and a high peak throughput of 2,415 GOPS with 1.65 ns minimum compute delay, highlighting superior performance and power efficiency compared to state-of-the-art designs. Tingran Chen, Yuxuan Ran, Yueting Li 0001, Wang Kang 0001, Biao Pan |
ISCAS | 6 |
| 2025 | An FPGA Processor Combining Point Cloud and SNN for DVS-based ADAS ApplicationabstractAutomatic Emergency Braking (AEB) has become an important component in Advanced Driver Assistance Systems (ADAS) and a potential solution for AEB lies in the integration of Dynamic Vision Sensor (DVS) with Spiking Neural Network (SNN). A high-precision behavioural recognition algorithm called Spikepoint has been proposed by us, which combines Point Cloud with SNN to enable recognition of DVS event data. This work concentrates on the FPGA implementation of Spikepoint, aiming to improve real-time recognition capabilities. The deployment of Spikepoint on FPGA encounters two challenges: 1) Point Cloud processing introduces additional latency 2) Storing parameters that require to be accessed frequently from DDR introduces a significant time overhead. In order to address challenges aforementioned, a novel reference point-based filtering technique for Point Cloud is introduced. Meanwhile, a fine-grained quantization method and other optimization strategies are used on the neuron model. The Xilinx UltraScale+ is employed in the experiments conducted in this work. Our Point-based SNN Processor achieves a recognition frame rate of 92.08 FPS through the novel algorithm and corresponding hardware optimization, while achieving an accuracy of 94.3% on the DVS128 Gesture dataset. Wente Yi, Kexun Cheng, Lehao Tan, Bojun Cheng, Biao Pan |
ISCAS | 9 |
| 2025 | Non-Local Mean Denoising of Multi-Layer With Adaptive Filtering Strength Based on ASIC ImplementationabstractImage denoising is an important algorithm in ASIC real-time image processing. Research has found that after cascaded spatial and temporal denoising, video images still exhibit patches and structural noise. To reduce the noise of this type while considering factors such as hardware resource overhead in ASIC implementation, this paper proposes a multi-layer adaptive threshold denoising method based on Non-Local Mean algorithm and pyramid framework. This algorithm performs spatial noise removal on Y channel image data in the YUV domain. Y component is firstly gaussian downsampled into 2 additional layers which is 1/4 and 1/2 scale of the original. Secondly, adaptive filtering strengths are estimated using DCT to further improve the NLM performance within each layer; Finally, the denoised results from the three-layer filtering are fused and output. Through comparison, the proposal can significantly suppress structural noise such as plagues. Meanwhile, our ASIC implementation of MRNLM can achieve real-time video denoising on-the-fly without outer memory access and the on-chip ASIC is kept minimal. Bingzhang Zhou, Biao Pan, Zhangming Huang, Xianghai Wei |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | CLIF: Complementary Leaky Integrate-and-Fire Neuron for Spiking Neural NetworksabstractSpiking neural networks (SNNs) are promising brain-inspired energy-efficient models. Compared to conventional deep Artificial Neural Networks (ANNs), SNNs exhibit superior efficiency and capability to process temporal information. However, it remains a challenge to train SNNs due to their undifferentiable spiking mechanism. The surrogate gradients method is commonly used to train SNNs, but often comes with an accuracy disadvantage over ANNs counterpart. We link the degraded accuracy to the vanishing of gradient on the temporal dimension through the analytical and experimental study of the training process of Leaky Integrate-and-Fire (LIF) Neuron-based SNNs. Moreover, we propose the Complementary Leaky Integrate-and-Fire (CLIF) Neuron. CLIF creates extra paths to facilitate the backpropagation in computing temporal gradient while keeping binary output. CLIF is hyperparameter-free and features broad applicability. Extensive experiments on a variety of datasets demonstrate CLIF's clear performance advantage over other neuron models. Furthermore, the CLIF's performance even slightly surpasses superior ANNs with identical network structure and training conditions. The code is available at https://github.com/HuuYuLong/Complementary-LIF. Xiaopeng Lin, Haotian Fu, Zunchang Liu, Biao Pan, Bojun Cheng |
ICML | 7 |
| 2024 | An End-to-End In-Memory Computing System Based on a 40-nm eFlash-Based IMC SoC: Circuits, Toolchains, and Systems Co-Design FrameworkabstractDespite its promising potential for Artificial Intelligence (AI) applications, current In-Memory Computing (IMC) technology faces a variety of challenges before mass production. One of the major challenges we face is the absence of efficient toolchains for deploying canonical networks on IMC chips. To address this issue, we propose a co-designed framework that integrates circuit, toolchain, and system elements specifically for IMC. More specifically, our framework consists of several key techniques to improve the key performance including (a) an 8-bit hardware-friendly Quantization-Aware Training (QAT) approach to quantify the deep learning network from floating-point data to fixed-point data, (b) a novel operator optimization technique to increase the computing precision when running the algorithm models on the IMC chips, and (c) an efficient mapping strategy based on the Integer Linear Programming (ILP) approach to improve the computation resource utilization of the IMC array. We assess our method on our 40nm eFlash-based IMC SoC chip with voice recognition, speech noise reduction, and person detection tasks. Our experimental results show an accuracy over 94.60% in a quiet environment and 87.27% in a white noise environment and a false recognition rate below 1 time per 24 hours for voice recognition, a 21.53% improvement for the Perceptual Evaluation of Speech Quality (PESQ) for noise reduction, and a 97.80% accuracy in person detection. Tianshuo Bai, Wanru Mao, Guangyao Wang, Aifei Zhang, Shihang Fu, Shuaikai Liu, Jianchao Hu, Xitong Yang, Biao Pan, Wei W. Xing, Wang Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2024 | PipeCIM: A High-Throughput Computing-In-Memory Microprocessor With Nested Pipeline and RISC-V Extended InstructionsabstractThe large number of multiply accumulate (MAC) operations in Convolutional Neural Network (CNN) leads to substantial data migration and computation. Although computing-in-memory (CIM) proves to be a promising paradigm for MAC operations, high throughput CNN accelerator still confronts bottlenecks from: the low MAC utilization and the uncessary off-chip memory access. In this paper, we propose a high throughput CIM-based CNN accelerator PipeCIM with three hierarchies of pipelines: Intra-Macro, Near-Memory and Tile-Level. The Intra-Macro Pipeline parallelly executes data transfer and in-memory-computing (IMC) operations. The Near-Memory Pipeline alleviates memory access for pooling and data reshaping. The Tile-Level Pipeline establishes a layer-wise pipeline to further improve the throughput while reducing control complexity. PipeCIM introduces the nested scheme and a Unidirectional Divergent Connection Protocol (UDTCP) to simplify the control of data flow with the help of customized RISC-V instructions. To validate our design, PipeCIM was prototyped in 55 nm process node, achieving energy efficiency of 133.8 TOPS/W and peak throughput of 819 GOPS with a 16KB CIM array, which can accelerate VGG-16 to 128.56$\times$or Inception to 19.754$\times$compared to the baseline. Tingran Chen, Wenjia Wang 0011, Haotian Fu, Wente Yi, Bojun Cheng, He Zhang 0011, Biao Pan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2024 | DS-CIM: A 40nm Asynchronous Dual-Spike Driven, MRAM Compute-In-Memory Macro for Spiking Neural NetworkabstractCompute-in-memory (CIM) based on emerging nonvolatile memory (eNVM) is an effective way to deploy neural networks to low-power edge devices for both storage and computation. NVMs such as ReRAM have been widely used in CIM. Meanwhile, MRAM has higher read and write cycles, lower device and cycle variation and a lower bit error rate, making it equally attractive for storage. However, the high read current and low on/off ratio result in large energy consumption in MRAM read limiting its large-scale application in CIM. The spiking neural network (SNN) represents the information as sparse spike sequences and facilitates hardware to achieve low-power computing by taking advantage of its spatial-temporal sparsity. To further increase the input sparsity of SNN and reduce the read energy consumption, this paper proposes ADC-free, dual-spike (DS) -CIM macro, a spiking MRAM CIM macro driven by asynchronous dual spikes. Compared to the conventional rate coding, our dual-spike coding method uses only 2 spikes to encode the information without losing accuracy. Moreover, the event-driven feature allows the macro to have sub-nW static power consumption. Our DS-CIM macro achieves comparable or higher accuracy while maintaining very low energy consumption. Specifically, it achieves accuracies of 96.99%, 82.87%, 90.00%, and 85.97% for digit classification, image classification, gesture recognition, and action recognition tasks, with energy consumption of only 8.07nJ, 71.26nJ, 729.3nJ, and 369.82nJ, respectively. These results emphasize the significance of DS-CIM and provide ideas for low-power inference on edge devices. Haotian Fu, Yulong Huang 0001, Tingran Chen, Chenyi Fu, Yue Zhou 0010, Shouzhong Peng, Zhirui Zong, Biao Pan, Bojun Cheng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2024 | RDCIM: RISC-V Supported Full-Digital Computing-in-Memory Processor With High Energy Efficiency and Low Area OverheadabstractDigital computing-in-memory (DCIM) that merges computing logic into memory has been proven to be an efficient architecture for accelerating multiply-and-accumulates (MACs). However, low energy efficiency and high area overhead pose a primary restriction for integrating DCIM in re-configurable processors required for multi-functional workloads. To alleviate this dilemma, a novel RISC-V supported full-digital computing-in-memory processor (RDCIM) is designed and fabricated with 55nm CMOS technology. In RDCIM, an adding-on-memory-boundary (AOMB) scheme is adopted to improve the energy efficiency of DCIM. Meanwhile, a multi-precision adaptive accumulator (MPAA) and a serial-parallel conversion supported SRAM buffer (SPBUF) are employed to reduce the area overhead caused by the peripheral circuits and the intermediate buffer for multi-precision support. The results show that the energy efficiency in our design is 16.6 TOPS/W (8-bit) and 66.3 TOPS/W (4-bit). Compared to related works, the proposed RDCIM macro shows a maximum energy efficiency improvement of 1.22$\times$in a continuous computing scenario, an area saving of 1.22$\times$in the accumulator, and an area saving of 3.12$\times$in the input buffer. Moreover, in RDCIM, 5 fine-grained RISC-V extended instructions are designed to dynamically adjust the state of DCIM, reaching 1.2$\times$computation efficiency. Wente Yi, Kefan Mo, Wenjia Wang 0011, Yejun Zeng, Zihan Yuan, Bojun Cheng, Biao Pan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2024 | Toward Energy-efficient STT-MRAM-based Near Memory Computing Architecture for Embedded SystemsabstractConvolutional Neural Networks (CNNs) have significantly impacted embedded system applications across various domains. However, this exacerbates the real-time processing and hardware resource-constrained challenges of embedded systems. To tackle these issues, we propose spin-transfer torque magnetic random-access memory (STT-MRAM)-based near memory computing (NMC) design for embedded systems. We optimize this design from three aspects: Fast-pipelined STT-MRAM readout scheme provides higher memory bandwidth for NMC design, enhancing real-time processing capability with a non-trivial area overhead. Direct index compression format in conjunction with digital sparse matrix-vector multiplication (SpMV) accelerator supports various matrices of practical applications that alleviate computing resource requirements. Custom NMC instructions and stream converter for NMC systems dynamically adjust available hardware resources for better utilization. Experimental results demonstrate that the memory bandwidth of STT-MRAM achieves 26.7 GB/s. Energy consumption and latency improvement of digital SpMV accelerator are up to 64× and 1,120× across sparsity matrices spanning from 10% to 99.8%. Single-precision and double-precision elements transmission increased up to 8× and 9.6×, respectively. Furthermore, our design achieves a throughput of up to 15.9× over state-of-the-art designs. Yueting Li 0001, He Zhang 0011, Biao Pan, Keni Qiu, Wang Kang 0001, Jun Wang 0041, Weisheng Zhao 0001 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2021 | HSC: A Hybrid Spin/CMOS Logic Based In-Memory Engine with Area-Efficient Mapping StrategyabstractRecent advances in deep learning have shown that binary neural networks (BNNs) can provide a satisfying accuracy on various tasks with significant reduction in computation power and memory cost. Theoretically, the multiply-and-accumulate (MAC) operations of BNNs can be replaced by in-memory XNOR operations, thereby avoiding frequent data transfer between the buffer and the processor. However, devices supporting in-memory implementation of XNOR operations together with efficient weight-matrix mapping strategy is still an open research area. In this paper, a hybrid spin/CMOS cell (HSC) structure is proposed in which the XNOR operation can be simply realized in an in-memory computing manner by the non-volatile data from the spin component and the volatile data from the CMOS component. Given the time/spatial trade-off, a novel weight mapping method to break the large memory array and unroll the 3D kernel into 2D weight matrix is designed to cooperate with the proposed HSC structure in a time-division way. System-level simulation results show that the proposed BNN processor can achieve a 3.32* speedup and 11.9* improvement in throughput and energy efficiency, which could be attributed to the device and mapping method co-design. Erya Deng, Jinyu Bai, Wang Kang 0001, Biao Pan |
ISCAS | 6 |
| 2021 | Recent progress of integrated circuits and optoelectronic chips
Yue Hao 0001, Genquan Han, Jincheng Zhang 0001, Xiaohua Ma 0001, Zhangming Zhu, Yanan Han, Ling Yang 0003, Jiangyi Shi, Wei Zhang 0343, Biao Pan, Yangqi Huang, Qi Liu 0010, Yimao Cai, Xin Ou, Tiangui You, Huaqiang Wu, Bin Gao 0006, Guoping Guo, Yonghua Chen, Xiangfei Chen, Chunlai Xue, Lixia Zhao, Xihua Zou, Lianshan Yan |
Sci. China Inf. Sci. | 18 |
| 2019 | Magnetic Skyrmion-Based Neural Recording System Design for Brain Machine InterfaceabstractNext-generation brain machine interface demand a high-channel-count neural recording system to wirelessly monitor activities of thousands of neurons. In order to achieve high-density neural recording, further development of single recording channel comprised of a neural amplifier front-end (AFE) and an analog-to-digit converter (ADC) is critical. Despite the great progress made in CMOS implementation of custom-designed neural recording system, hybrid limitations of increasing area and power consumption in line with Moore's law drove great demand for post-CMOS substitutes. Magnetic skyrmion with nano particle-like and non-volatile properties are of both fundamental and applied interests for future bio-inspired electronics. In this work, we propose a compact model including both AFE and ADC based on current-induced skyrmion motion. The proposed system achieved a power consumption of 0.63 pJ/channel with an area overhead of 0.14 μm2. The purpose of this work is to explore the feasibility of magnetic skyrmion for building large-scale, dense neuronal recording system which could pave a new way for future brain machine interface application. Biao Pan, Wang Kang 0001, Xing Chen 0012, Jinyu Bai, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 1 |
| 2019 | SR-WTA: Skyrmion Racing Winner-Takes-All Module for Spiking Neural ComputingabstractSpiking neural network (SNN) has emerged as one of the popular architectures in complex pattern recognition and classification tasks. However, hardware implementation of such algorithms using conventional CMOS based neuron consume resources and power that are orders of magnitude higher than that in human brain. This can be attributed to the mismatch of the computational architecture between biological brain and the current Boolean logic computing platform. Magnetic skyrmions have been intensively studied as a prospective information carrier in neuromorphic computing hardware design. In this work, a compact time-domain skyrmion-racing winner-takes-all (SR-WTA) leaky-integrate-fire (LIF) spiking neuron network is presented for the first time. The skyrmion motion dynamics in the LIF neuron and the behaviors of the neuron network was investigated comprehensively. Both SPICE and micromagnetic simulations are performed to evaluate the functionality and performance of the proposed SR-WTA based SNN. Biao Pan, Wang Kang 0001, Xing Chen 0012, Jinyu Bai, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 1 |