EDBT 2026 Demo / reviewers in the wild / expert
Yunji Qin
dblp:334/1874
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0005-9404-581XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoE-Sched: Enabling Efficient FPGA Deployment of Mixture-of-Experts Vision Transformers via Coordinated SchedulingabstractCompared to traditional Vision Transformers (ViT), Mixture-of-Experts ViTs (MoE-ViTs) are introduced to scale model size without a proportional increase in computational complexity, making them a new research focus. Given the high performance and reconfigurability, field-programmable gate array (FPGA)-based accelerators for MoE-ViTs emerge, delivering substantial gains over general-purpose processors. However, existing accelerators often fall short of efficiently managing the highly dynamic and sparse computation patterns, resulting in suboptimal tradeoffs between resource utilization and performance. To address the inefficiencies in deploying MoE-ViTs on FPGAs, we present MoE-Sched, a novel end-to-end accelerator that embraces a scheduling-centric design philosophy. Rather than optimizing isolated kernels, MoE-Sched coordinates multilevel scheduling, from fine-grained intrakernel streaming to module reuse and multidie mapping, to holistically balance latency, bandwidth (BW), and resource usage. We further integrate a hardware-aware quantization scheme tailored for streaming attention and sparse expert execution, preserving accuracy while minimizing overhead. Experimental results demonstrate that our accelerator achieves nearly 100 frames/s on M3ViT-tiny, a$3.13\times $improvement in throughput, and over 75% energy reduction compared to state-of-the-art (SOTA) FPGA MoE accelerators, while maintaining less than 1% accuracy loss across vision benchmarks. Our implementation will be open-sourced. Jiale Dong, Wenqi Lou, Zhendong Zheng, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | UniCoS: A Unified Neural and Accelerator Co-Search Framework for CNNs and ViTsabstractCurrent algorithm-hardware co-search works often suffer from lengthy training times and inadequate exploration of hardware design spaces, leading to suboptimal performance. This work introduces UniCoS, a unified framework for co-optimizing neural networks and accelerators for CNNs and Vision Transformers (ViTs). By introducing a novel training-free proxy that evaluates accuracy within seconds and a clustering-based algorithm for exploring heterogeneous dataflows, UniCoS efficiently navigates the design spaces of both architectures. Experimental results demonstrate that the solutions generated by UniCoS consistently surpass state-of-the-art (SOTA) methods (e.g., $3.54 \times$ energy-delay product (EDP) improvement with a $1.76 \%$ higher accuracy on ImageNet) while requiring notably reduced search time (up to $48 \times, \sim 3$ hours). The code is available at https://github.com/mine7777/Unicos.git. Wenqi Lou, Cheng Tang 0004, Hongbing Wen, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
DAC | 5 |
| 2025 | Automated FPGA Accelerator Generation Framework for Transformers with Dataflow OptimizationabstractTransformers have revolutionized natural language processing (NLP) and computer vision (CV) tasks, yet their deployment remains constrained by the quadratic complexity of self-attention. While FPGA-based accelerators offer promising solutions, existing designs struggle with fixed dataflow patterns and limited hardware specialization across diverse model architectures and sequence lengths. This paper presents AutoTrans, an automated framework for generating optimized FPGA accelerators tailored to Transformers. To improve dataflow flexibility, we propose configurable fusion granularities at the tensor, row, and block levels for self-attention layers, enabling fine-grained trade-offs between computational efficiency and resource utilization. For hardware specialization, we develop a parameterized heterogeneous multi-core architecture featuring dedicated compute engines for attention and linear layers, guided by a genetic algorithm-based design space exploration (DSE) strategy. Experimental results demonstrate the effectiveness of AutoTrans across varied application scenarios. On the ZCU102 board, AutoTrans achieves up to 621 GOPS when accelerating BERT-base with 4 K-token inputs, yielding a 1.27 × to 1.91 × improvement in energy efficiency over previous designs. For ViT-Base with 256-token inputs on the Alveo U50, AutoTrans attains up to 1548 GOPS, achieving up to 1.96 × higher compute density, highlighting its scalability and adaptability across NLP and vision domains. Wenqi Lou, Yunji Qin, Chao Wang 0003, Lei Gong 0003, Xuehai Zhou |
ICPP | 2 |
| 2025 | UbiMoE: A Ubiquitous Mixture-of-Experts Vision Transformer Accelerator With Hybrid Computation Pattern on FPGAabstractCompared to traditional Vision Transformers (ViT), Mixture-of-Experts Vision Transformers (MoE-ViT) are introduced to scale model size without a proportional increase in computational complexity, making them a new research focus. Given the high performance and reconfigurability, FPGA-based accelerators for MoE-ViT emerge, delivering substantial gains over general-purpose processors. However, existing accelerators often fall short of fully exploring the design space, leading to suboptimal trade-offs between resource utilization and performance. To overcome this problem, we introduce UbiMoE, a novel end-to-end FPGA accelerator tailored for MoE-ViT. Leveraging the unique computational and memory access patterns of MoE-ViTs, we develop a latency-optimized streaming attention kernel and a resource-efficient reusable linear kernel, effectively balancing performance and resource consumption. To further enhance design efficiency, we propose a two-stage heuristic search algorithm that optimally tunes hardware parameters for various FPGA resource constraints. Compared to state-of-the-art (SOTA) FPGA designs, UbiMoE achieves 1.34× and 3.35× throughput improvements for MoE-ViT on Xilinx ZCU102 and Alveo U280 platforms, respectively, while enhancing energy efficiency by 1.75× and 1.54×. Our implementation is available at https://github.com/DJ000011/UbiMoE. Jiale Dong, Wenqi Lou, Zhendong Zheng, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
ISCAS | 4 |
| 2024 | Enhancing Long Sequence Input Processing in FPGA-Based Transformer Accelerators through Attention FusionabstractAttention-based transformers have achieved significant performance breakthroughs in natural language processing (NLP) and computer vision (CV) tasks. Meanwhile, the ever-increasing length of today’s input sequences puts much pressure on computing devices. FPGAs are widely used to accelerate Transformer inference due to their high energy efficiency and flexibility. However, most of the existing FPGA-based Transformer accelerators are oriented to small input lengths, making it hard to accelerate long input sequences. To this end, we design an efficient Transformer accelerator for FPGA and long-sequence input scenarios. We use the tiling softmax algorithm to fuse attention computation, eliminating the memory and bandwidth bottleneck in the attention layer and allowing our accelerator to support arbitrary input sequence lengths. We use BERT-Base on the Alveo U50 board for evaluation, and our implementation achieves computational efficiency improvements of 1.09 ∼ 2.48 × over prior FPGA accelerators. Besides, our accelerator can support up to 175K input sequence length when running BERT-like structures, far more than previous designs. Yunji Qin, Wenqi Lou, Chao Wang 0003, Lei Gong 0003, Xuehai Zhou |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | MFNAS: Multi-fidelity Exploration in Neural Architecture Search with Stable Zero-Shot Proxy
Wenqi Lou, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
PRICAI (1) | 3 |
| 2024 | FlexBCM: Hybrid Block-Circulant Neural Network and Accelerator Co-Search on FPGAsabstractBlock-circulant matrix (BCM) compression has garnered much attention in the hardware acceleration of convolutional neural networks (CNNs) due to its regularity and efficiency. However, constrained by the difficulty of exploring the compression parameter space, existing BCM-based methods often apply a uniform compression parameter to all CNN models’ layers, losing the compression’s flexibility. Additionally, independently optimizing models or accelerators makes achieving the optimal tradeoff between model accuracy and hardware efficiency challenging. To this end, we propose FlexBCM, a joint exploration framework that efficiently explores both the parameter compression and hardware parameter space to generate customized hybrid BCM-compressed CNN and field-programmable gate array (FPGA) accelerator solutions. On the algorithmic side, leveraging the idea of neural architecture search (NAS), we design an efficient differentiable sampling method to rapidly evaluate the accuracy of candidate subnets. Additionally, we devise a hardware-friendly frequency domain quantization scheme for BCM computation. On the hardware side, we develop the efficient and parameter-configurable convolutional core (ConvPU) alongside the BCM computing core (BCMPU). The BCMPU can flexibly accommodate different compression parameters at runtime, incorporate complex-number DSP packing and conjugate symmetry optimizations. For model-to-hardware evaluation, we construct accurate latency and resource consumption models. Moreover, we design a fast hardware generation algorithm based on the coarse-grained search to provide prompt feedback on the hardware evaluation of the current subnet. Finally, we validate FlexBCM on the Xilinx ZCU102 FPGA and compare its compressed CNN-accelerator solutions with previous state-of-the-art works. Experimental results demonstrate that FlexBCM achieves 1.21–3.02 times higher-computational efficiency for ResNet18 and ResNet34 models while maintaining an acceptable accuracy loss on the ImageNet dataset. Wenqi Lou, Yunji Qin, Xuan Wang 0020, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Work-in-Progress: BloCirNN: An Efficient Software/hardware Codesign Approach for Neural Network Accelerators with Block-Circulant MatrixabstractNowadays, the scale of deep neural networks is getting larger and larger. These large-scale deep neural networks are both compute and memory intensive. To overcome these problems, we use block-circulant weight matrices and Fast Fourier Transform (FFT) to compress model and optimize computation. Compared to weight pruning, this method does not suffer from irregular networks. The main contributions of this paper include the implementation of a convolution module and a fully-connected module with High-Level Synthesis (HLS), deployment and performance test on FPGA platform. We use AlexNet as a case study, which demonstrates our design is more efficient than the FPGA2016. Yunji Qin, Lei Gong 0003, Zhendong Zheng, Chao Wang 0003 |
CODES+ISSS | 1 |