EDBT 2026 Demo / reviewers in the wild / expert
Pinfeng Jiang
dblp:401/4419
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0002-3242-1574ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPEAR: Spike-Aware Point Pruning and Fused FPS-KNN Sampling for Spiking Point Transformer Acceleration Co-Design
Yilong Fang, Pinfeng Jiang, Ming-De Zhu, Xiangshui Miao, Xingsheng Wang |
ISCAS | 2 |
| 2025 | SPARTA: Spike-Aware Token Skipping Co-Optimization with Heterogeneous ReRAM-CIM Architecture for Spiking Transformer AccelerationabstractSpiking neural networks (SNNs) have emerged as a promising paradigm for energy-efficient neural computation, offering advantages over artificial neural networks (ANNs) by using sparse spike activations. Among SNN models, spiking transformers have shown great potential in achieving high accuracy while maintaining low energy consumption, making them ideal for resource-constrained applications. However, a significant challenge in accelerating Spiking Transformers lies in managing the inherent unstructured sparsity in spike activations. This sparsity introduces substantial hardware overhead, limiting the efficiency of existing accelerators.To address these challenges, we propose SPARTA, an algorithm-hardware co-optimization framework designed specifically for spiking transformers. SPARTA utilizes a spike-aware dynamic token skipping algorithm, which applies reinforcement learning to selectively skip less informative tokens, achieving structured token-level sparsity in both spatial and temporal domains. Additionally, we introduce the spike-aware token prediction algorithm to predict and eliminate inactive tokens, further improving efficiency. Meanwhile, we present a dedicated heterogeneous ReRAM-based Compute-in-Memory (CIM) hardware architecture tailored to support token-level sparsity, which integrates a ReRAM analog CIM engine for linear layer and a token-spike fusion engine for optimized token routing and spike attention. Experimental results show that SPARTA achieves up to 543.1× and 10.2× speedup with 308.0× and 5.2× energy efficiency improvement compared to GPU and the state-of-the-art SNN accelerator "COMPASS", while preserving high model accuracy. Pinfeng Jiang, Yilong Fang, Ming-De Zhu, Xiangshui Miao, Xingsheng Wang |
ICCAD | 1 |
| 2025 | ReBA: A Hybrid Sparse Reconfigurable Butterfly Accelerator for Solving Partial Differential Equations via Hardware and Algorithm Co-DesignabstractPartial differential equations (PDEs) are widely used in many scientific and engineering fields. Traditional numerical methods for solving PDEs often struggle with complex equations and require extensive computation. Recent advances in neural operators, particularly the Galerkin Transformer (GT-FNO), offer faster solutions but present unique challenges for hardware due to complex neural network operations such as attention, Fourier transforms, and convolution. To address these challenges, we propose ReBA, a dedicated hardware and algorithm co-design framework specified for GT-FNO. Our design implements hybrid sparsity schemes for efficient workload balance, significantly reducing hardware overhead while maintaining high accuracy. ReBA introduces a reconfigurable dual-engine architecture, featuring an adaptable butterfly processing element (BPE), which efficiently supports FFT and other DNN operations through flexible BPE allocation, offering high parallelism and low-latency computation. Experiments validate that ReBA delivers significant speedups of 34.57×, 2.70×, and 1.82 51.26× over conventional CPUs, GPUs, and prior state-of-the-art accelerators in solving PDEs. Pinfeng Jiang, Yilong Fang, Ming-De Zhu, Xiangshui Miao, Xingsheng Wang |
ISCAS | 1 |
| 2025 | Optimizing hardware-software co-design based on non-ideality in memristor crossbars for in-memory computing
Pinfeng Jiang, Danzhe Song, Menghua Huang, Fan Yang 0148, Xiangshui Miao, Xingsheng Wang |
Sci. China Inf. Sci. | 1 |
| 2025 | ISARA: An Island-Style Systolic Array Reconfigurable Accelerator Based on Memristors for Deep Neural NetworksabstractThe demand for edge artificial intelligence (AI) is significant, particularly in revolutionary technological areas such as the Internet of Things, autonomous driving, and industrial control. However, reliable and high-performance edge AI is still constrained by computing hardware, and improving the performance and reliability of edge AI accelerators remains a key focus for researchers. This work proposes a memristor/resistive random access memory (RRAM)-based island-style systolic array reconfigurable accelerator (ISARA) that meets the reliability and performance requirements of edge AI. Inspired by the island-style architecture of FPGAs, this work proposes a flexible-tile architecture based on RRAM processing element (PE) islands, optimizing the data flow within the systolic array. The design of network-on-chip reduces data processing latency. In addition, to enhance computational efficiency, this work incorporates a bit-fusion scheme within the flexible tile, which reduces analog-to-digital converter (ADC) power consumption and addresses the conductance variation of RRAM. To date, only a few works have completed the entire process from simulation, design, and fabrication to hardware testing. This work fully realizes the design and validation of a new accelerator based on RRAM chips, demonstrating the reliability of RRAM-based systolic array accelerators for the first time. After deploying algorithms, the hardware accelerator achieved recognition rates comparable to software. Compared to similar works, ISARA’s computational efficiency exceeds theirs and has flexible reconfigurability. The same deep neural network (DNN) models are adopted for evaluation and compared to other accelerators, and ISARA’s processing latency is reduced by 200 times. Fan Yang 0148, Pinfeng Jiang, Xiangshui Miao, Xingsheng Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |