EDBT 2026 Demo / reviewers in the wild / expert
Ming-De Zhu
dblp:196/4186 · also Mingde Zhu
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPEAR: Spike-Aware Point Pruning and Fused FPS-KNN Sampling for Spiking Point Transformer Acceleration Co-Design
Yilong Fang, Pinfeng Jiang, Ming-De Zhu, Xiangshui Miao, Xingsheng Wang |
ISCAS | 5 |
| 2025 | SPARTA: Spike-Aware Token Skipping Co-Optimization with Heterogeneous ReRAM-CIM Architecture for Spiking Transformer AccelerationabstractSpiking neural networks (SNNs) have emerged as a promising paradigm for energy-efficient neural computation, offering advantages over artificial neural networks (ANNs) by using sparse spike activations. Among SNN models, spiking transformers have shown great potential in achieving high accuracy while maintaining low energy consumption, making them ideal for resource-constrained applications. However, a significant challenge in accelerating Spiking Transformers lies in managing the inherent unstructured sparsity in spike activations. This sparsity introduces substantial hardware overhead, limiting the efficiency of existing accelerators.To address these challenges, we propose SPARTA, an algorithm-hardware co-optimization framework designed specifically for spiking transformers. SPARTA utilizes a spike-aware dynamic token skipping algorithm, which applies reinforcement learning to selectively skip less informative tokens, achieving structured token-level sparsity in both spatial and temporal domains. Additionally, we introduce the spike-aware token prediction algorithm to predict and eliminate inactive tokens, further improving efficiency. Meanwhile, we present a dedicated heterogeneous ReRAM-based Compute-in-Memory (CIM) hardware architecture tailored to support token-level sparsity, which integrates a ReRAM analog CIM engine for linear layer and a token-spike fusion engine for optimized token routing and spike attention. Experimental results show that SPARTA achieves up to 543.1× and 10.2× speedup with 308.0× and 5.2× energy efficiency improvement compared to GPU and the state-of-the-art SNN accelerator "COMPASS", while preserving high model accuracy. Pinfeng Jiang, Yilong Fang, Ming-De Zhu, Xiangshui Miao, Xingsheng Wang |
ICCAD | 5 |
| 2025 | ReBA: A Hybrid Sparse Reconfigurable Butterfly Accelerator for Solving Partial Differential Equations via Hardware and Algorithm Co-DesignabstractPartial differential equations (PDEs) are widely used in many scientific and engineering fields. Traditional numerical methods for solving PDEs often struggle with complex equations and require extensive computation. Recent advances in neural operators, particularly the Galerkin Transformer (GT-FNO), offer faster solutions but present unique challenges for hardware due to complex neural network operations such as attention, Fourier transforms, and convolution. To address these challenges, we propose ReBA, a dedicated hardware and algorithm co-design framework specified for GT-FNO. Our design implements hybrid sparsity schemes for efficient workload balance, significantly reducing hardware overhead while maintaining high accuracy. ReBA introduces a reconfigurable dual-engine architecture, featuring an adaptable butterfly processing element (BPE), which efficiently supports FFT and other DNN operations through flexible BPE allocation, offering high parallelism and low-latency computation. Experiments validate that ReBA delivers significant speedups of 34.57×, 2.70×, and 1.82 51.26× over conventional CPUs, GPUs, and prior state-of-the-art accelerators in solving PDEs. Pinfeng Jiang, Yilong Fang, Ming-De Zhu, Xiangshui Miao, Xingsheng Wang |
ISCAS | 5 |