Qiwei Dong

dblp:348/0435 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0005-6848-0375ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Lightweight Algorithm-Hardware Co-design for Real-Time Video Frame Interpolation
Jisheng Zhang, Qiwei Dong, Wendong Mao, Zhongfeng Wang 0001
ISCAS2
2026 Semi-supervised medical image segmentation method via dual-view graph contrastive learning and latent space uncertainty rectification
Dongxu Cheng, Qiwei Dong, Ruian Zhu, Yuhui Zheng
Eng. Appl. Artif. Intell.2
2025 SSMA: A Memory-Efficient Accelerator for State Space Model in the Mamba
abstract
Mamba has exhibited great potential across various tasks, achieving the powerful capability of long-sequence modeling with linear complexity. Selective State Space Models (SSMs), the core component of Mamba, possess unique computational flow and massive memory requirements, which pose new challenges for deployment on edge devices and are difficult to be efficiently supported by existing deep learning accelerators. Therefore, we develop the first memory-efficient Selective SSM Accelerator, namely SSMA. Specifically, we design a reconfigurable hardware architecture and low-complexity nonlinear units to efficiently execute the rearranged and fused operations in Selective SSM. Moreover, we introduce a novel tile-stationary recurrent dataflow to achieve on-chip layer fusion by recurrently updating the tiled latent state and SSM parameters, dramatically decreasing memory usage and access. Experimental results show that the proposed SSMA, evaluated on Xilinx ZCU102 FPGA, achieves up to 3.33× speedup and 41.8 × higher energy efficiency than the optimized GPU implementation with the same setting.
Qiwei Dong, Zhongfeng Wang 0001
ISCAS1
2025 An Efficient Window-Based Vision Transformer Accelerator via Mixed-Granularity Sparsity
abstract
Vision Transformers (ViTs) have achieved excellent performance on various computer vision tasks, while their high computation and memory costs pose challenges for practical deployment. To address this issue, token-level pruning is used as an effective method to compress ViTs, discarding unimportant image tokens that contribute little to predictions. However, directly applying unstructured token pruning to window-based ViTs damages their regular feature map structure, resulting in load imbalance when deployed on mobile devices. In this work, we propose an efficient algorithm-hardware co-optimized framework to accelerate window-based ViTs via adaptive Mixed-Granularity Sparsity (MGS). At the algorithm level, a hardware-friendly MGS algorithm is developed by integrating the inherent sparsity, global window pruning, and local N:M token pruning to balance model accuracy and its computational complexity. At the hardware level, we present a dedicated accelerator equipped with a sparse computing core and two lightweight auxiliary processing units to execute window-based calculations efficiently using MGS. Additionally, we devise a dynamic pipeline interleaving dataflow to achieve on-chip layer fusion, which reduces the processing latency and maximizes data reuse. Experimental results demonstrate that, with similar computational complexity, our highly structured MGS algorithm can achieve comparable or even better accuracy than previous compression methods. Moreover, compared to existing FPGA-based accelerators for Transformers, our design can achieve$1.80~\sim ~6.52\times $and$1.16~\sim ~12.05\times $improvements in terms of throughput and energy efficiency, respectively.
Qiwei Dong, Zhongfeng Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 A Unified Accelerator for All-in-One Image Restoration Based on Prompt Degradation Learning
abstract
All-in-one image restoration (IR) recovers images from various unknown distortions by a single model, such as rain, haze, and blur. Transformer-based IR methods have significantly improved the visual effects of the restored images. However, deploying complex IR models on edge devices is challenging due to massive parameters and intensive computations. Moreover, existing accelerators are typically customized for a single task, resulting in severe resource underutilization when executing multiple tasks. Therefore, this paper develops an algorithm-hardware co-design framework to accelerate a novel CNN-Transformer cooperative model for multiple IR tasks. Firstly, on the algorithm level, an Efficient Restoration Foundational Model (ERFM) is proposed to recover corrupted images from various degradations with low model complexity. Secondly, to guide adaptive corruption removal, a novel prompt learning scheme is introduced to fuse context-related degradation cues and boost high-quality reconstruction. Thirdly, on the hardware level, an integer approximation method is proposed to avoid expensive hardware overhead caused by complex nonlinear operations, such as layer normalization and softmax while maintaining comparable IR quality. Moreover, a head stationary dataflow and softmax fusion mechanism are designed to reduce data movement and enhance on-chip resource utilization. Finally, an overall hardware architecture is developed and implemented in TSMC 28 nm CMOS technology. Experimental results show that our ERFM achieves better visual perception than other baselines on seven challenging IR tasks without task-specific fine-tuning. Moreover, compared to other accelerators for vision Transformers, our design can achieve 3.3$\times$and 3.7$\times$improvements in throughput and energy efficiency.
Qiwei Dong, Wendong Mao, Zhongfeng Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 SWAT: An Efficient Swin Transformer Accelerator Based on FPGA
abstract
Swin Transformer achieves greater efficiency than Vision Transformer by utilizing local self-attention and shifted windows. However, existing hardware accelerators designed for Transformer have not been optimized for the unique computation flow and data reuse property in Swin Transformer, resulting in lower hardware utilization and extra memory accesses. To address this issue, we develop SWAT, an efficient Swin Transformer Accelerator based on FPGA. Firstly, to eliminate the redundant computations in shifted windows, a novel tiling strategy is employed, which helps the developed multiplier array to fully utilize the sparsity. Additionally, we deploy a dynamic pipeline interleaving dataflow, which not only reduces the processing latency but also maximizes data reuse, thereby decreasing access to memories. Furthermore, customized quantization strategies and approximate calculations for non-linear calculations are adopted to simplify the hardware complexity with negligible network accuracy loss. We implement SWAT on the Xilinx Alveo U50 platform and evaluate it with Swin-T on the ImageNet dataset. The proposed architecture can achieve improvements of $2.02 \times \sim 3.11 \times$ in power efficiency compared to existing Transformer accelerators on FPGAs.
Qiwei Dong, Xiaoru Xie, Zhongfeng Wang 0001
ASPDAC1