Chunyu Qi

dblp:305/2676 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0004-9814-2468ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AQuant: Repurposing CODEC for VLM Acceleration via Adaptive Quantization
Zhuoran Song, Chunyu Qi, Jian Weng, Xiaoyao Liang, Haibing Guan
ISCA2
2025 SAGA: A Memory-Efficient Accelerator for GANN Construction via Harnessing Vertex Similarity
abstract
Graph-traversal-based Approximate Nearest Neighbor (GANN) search and construction have become key retrieval techniques in various domains, such as recommendation systems and social networks. However, deploying GANN in real-world scenarios faces significant challenges, as high-dimensional vertices within the graph can lead to intensive memory demands. Although architectures like NDSearch have been proposed to accelerate GANN search, they are hard to deploy for GANN construction, as their pre-processing methods introduce massive overhead in dynamic graphs. In this paper, given the observation that neighboring vertices in a dynamic graph exhibit feature similarity, we propose SAGA, the first accelerator that alleviates memory bound in GANN construction. To capture this similarity, we directly leverage the first step of construction to gather vertices with the same starting point into a cluster to minimize the similarity detection overhead. Next, we decompose vertices into key and non-key ones, where their deltas fall in a narrow range, which is suitable to be quantized to lower bit widths. Building upon this approach, we design a specialized architecture, which efficiently implements the GANN construction by twolevel scheduling and a mixed-precision supported bit-serial unit. Through comprehensive evaluation, we demonstrate that SAGA can achieve an average speedup of $9.30 \times 4.87 \times 4.15 \times$ and $35.46 \times 7.60 \times 5.15 \times$ energy savings over CPU, GPU and NDSearch, respectively, while retaining task accuracy.
Xueyuan Liu 0001, Chunyu Qi, Yuanzheng Yao, Yanan Sun 0003, Xiaoyao Liang, Zhuoran Song
DAC3
2025 MHDiff: Memory- and Hardware-Efficient Diffusion Acceleration via Focal Pixel Aware Quantization
abstract
Diffusion models have demonstrated superior performance in image generation tasks, thus becoming the mainstream model for generative visual tasks. Diffusion models need to execute multiple timesteps sequentially, resulting in a dramatic increase in workload. Existing accelerators leverage the data similarity between adjacent timesteps and perform mixed-precision differential quantization to accelerate diffusion models. However, merging differential values with raw inputs in each layer of each timestep to ensure computational correctness requires significant memory access for loading raw inputs, which creates a heavy memory burden. Moreover, mixed-precision computations may lead to low hardware utilization if not well designed. Unlike these works, we propose MHDiff, a tailored framework that identifies the focal pixels at the first layer and finetunes them to fit all layers, then represents focal pixels with high-precision while using low-precision for others, thereby accelerating diffusion models while minimizing memory burden. To improve hardware utilization, MHDiff employs a packing module that merges low-precision values into high-precision values to create full high-precision matrices and designs a processing element (PE) array to efficiently process the packed matrices. Extensive experiment results demonstrate that MHDiff can achieve satisfactory performance with negligible quality loss.
Chunyu Qi, Xuhang Wang, Yuanzheng Yao, Naifeng Jing, Chen Zhang 0001, Jun Wang 0001, Zhihui Fu, Xiaoyao Liang, Zhuoran Song
DAC1
2025 SynGPU: Synergizing CUDA and Bit-Serial Tensor Cores for Vision Transformer Acceleration on GPU
abstract
Vision Transformers (ViTs) have demonstrated remarkable performance in computer vision tasks by effectively extracting global features. However, their self-attention mechanism suffers from quadratic time and memory complexity as image resolution or video duration increases, leading to inefficiency on GPUs. To accelerate ViTs, existing works mainly focus on pruning tokens based on value-level sparsity. However, they miss the chance to achieve peak performance as they overlook the bit-level sparsity. Instead, we propose Inter-token Bit-sparsity Awareness (IBA) algorithm to accelerate ViTs by exploring bit-sparsity from similar tokens. Next, we implement IBA on GPUs that synergize CUDA and Tensor Cores by addressing two issues: firstly, the bandwidth congestion of the Register File hinders the parallel ability of CUDA and Tensor Cores. Secondly, due to the varying exponent of floating-point vectors, it is hard to accelerate bitsparse matrix multiplication and accumulation (MMA) in Tensor Core through fixed-point-based bit-level circuits. Therefore, we present SynGPU, an algorithm-hardware co-design framework, to accelerate ViTs. SynGPU enhances data reuse by a novel data mapping to enable full parallelism of CUDA and Tensor Cores. Moreover, it introduces Bit-Serial Tensor Core (BSTC) that supports fixed- and floating-point MMA by combining the fixedpoint Bit-Serial Dot Product (BSDP) and exponent alignment techniques. Extensive experiments show that SynGPU achieves an average of $2.15 \times \sim 3.95 \times$ speedup and $2.49 \times \sim 3.81 \times$ compute density over A100 GPU.
Yuanzheng Yao, Chen Zhang 0001, Chunyu Qi, Jun Wang 0001, Zhihui Fu, Naifeng Jing, Xiaoyao Liang, Zhuoran Song
DAC3
2025 RTSA: A Run-Through Sparse Attention Framework for Video Transformer
abstract
In the realm of video understanding tasks, Video Transformer models (VidT) have recently exhibited impressive accuracy improvements in numerous edge devices. However, their deployment poses significant computational challenges for hardware. To address this, pruning has emerged as a promising approach to reduce computation and memory requirements by eliminating unimportant elements from the attention matrix. Unfortunately, existing pruning algorithms face a limitation in that they only optimize one of the two key modules on VidT's critical path: linear projection or self-attention. Regrettably, due to the variation in battery power in edge devices, the video resolution they generate will also change, which causes both linear projection and self-attention stages to potentially become bottlenecks, the existing approaches lack generality. Accordingly, we establish a Run-Through Sparse Attention (RTSA) framework that simultaneously sparsifies and accelerates two stages. On the algorithm side, unlike current methodologies conducting sparse linear projection by exploring redundancy within each frame, we extract extra redundancy naturally existing between frames. Moreover, for sparse self-attention, as existing pruning algorithms often provide either too coarse-grained or fine-grained sparsity patterns, these algorithms face limitations in simultaneously achieving high sparsity, low accuracy loss, and high speedup, resulting in either compromised accuracy or reduced efficiency. Thus, we prune the attention matrix at a medium granularity—sub-vector. The sub-vectors are generated by isolating each column of the attention matrix. On the hardware side, we observe that the use of distinct computational units for sparse linear projection and self-attention results in pipeline imbalances because of the bottleneck transformation between the two stages. To effectively eliminate pipeline stall, we design a RTSA architecture that supports sequential execution of both sparse linear projection and self-attention. To achieve this, we devised an atomic vector-scalar product computation underpinning all calculations in parse linear projection and self-attention, as well as evolving a spatial array architecture with augmented processing elements (PEs) tailored for the vector-scalar product. Experiments on VidT models show that RTSA can save 2.71$\boldsymbol{\times}$to 5.32$\boldsymbol{\times}$ideal computation with$ \lt 1\%$accuracy loss, achieving 105$\boldsymbol{\times}$, 56.8$\boldsymbol{\times}$, 3.59$\boldsymbol{\times}$, and 3.31$\boldsymbol{\times}$speedup compared to CPU, GPU, as well as the state-of-the-art ViT accelerators ViTCoD and HeatViT.
Xuhang Wang, Zhuoran Song, Chunyu Qi, Fangxin Liu, Naifeng Jing, Li Jiang 0002, Xiaoyao Liang
IEEE Trans. Computers3
2024 CMC: Video Transformer Acceleration via CODEC Assisted Matrix Condensing
abstract
Video Transformers (VidTs) have reached the forefront of accuracy in various video understanding tasks. Despite their remarkable achievements, the processing requirements for a large number of video frames still present a significant performance bottleneck, impeding their deployment to resource-constrained platforms. While accelerators meticulously designed for Vision Transformers (ViTs) have emerged, they may not be the optimal solution for VidTs, primarily due to two reasons. These accelerators tend to overlook the inherent temporal redundancy that characterizes VidTs, limiting their chance for further performance enhancement. Moreover, incorporating a sparse attention prediction module within these accelerators incurs a considerable overhead.
Zhuoran Song, Chunyu Qi, Fangxin Liu, Naifeng Jing, Xiaoyao Liang
ASPLOS (2)2
2024 TSAcc: An Efficient \underline{T}empo-\underline{S}patial Similarity Aware \underline{Acc}elerator for Attention Acceleration
abstract
Attention-based models provide significant accuracy improvement to Natural Language Processing (NLP) and computer vision (CV) fields at the cost of heavy computational and memory demands. Previous works seek to alleviate the performance bottleneck by removing useless relations for each position. However, their attempts only focus on intra-sentence optimization and overlook the opportunity in the temporal domain. In this paper, we accelerate attention by leveraging the tempo-spatial similarity across successive sentences, given the observation that successive sentences tend to bear high similarity. This is rational owing to many semantic similar words (namely tokens) in the attention-based models. We first propose an online-offline prediction algorithm to identify similar tokens/heads. We then design a recovery algorithm so that we can skip the computation on similar tokens/heads in succeeding sentences and recover their results by copying other tokens/heads features in preceding sentences to reserve accuracy. From the hardware aspect, we propose a specialized architecture TSAcc that includes a prediction engine and recovery engine to translate the computational saving in the algorithm to real speedup. Experiments show that TSAcc can achieve 8.5X, 2.7X, 14.1X, and 64.9X speedup compared to SpAtten, Sanger, 1080TI GPU, and Xeon CPU, with negligible accuracy loss.
Zhuoran Song, Chunyu Qi, Yuanzheng Yao, Peng Zhou 0030, Yanyi Zi, Xiaoyao Liang
DAC2
2024 An Innovative Application of Isolation-Based Nearest Neighbor Ensembles on Hyperspectral Anomaly Detection
abstract
This letter presents an innovative application of the isolation using nearest neighbor ensemble (iNNE) method for hyperspectral anomaly detection (HAD). iNNE is an efficient anomaly detector based on nearest neighbors (NNs) and isolation. Based on iNNE, we propose a novel isolation-based HAD framework that detects anomalies as follows. First, the Gabor filter is applied to extract spatial information from the PCA-projected subspace. Gabor features are then employed as the input to the ReMass-iForest to obtain the spatial anomaly score. Then, the iNNE is employed to derive the spectral anomaly score. Finally, we integrate the detection outcomes by linearly combining the acquired spatial and spectral anomaly scores to predict anomaly pixels. Compared with the existing tree-based isolation method, such as the iForest, the proposed iNNE-based method addresses the poor performance of iForest in detecting local anomalies and anomaly detection in high-dimensional data. Experimental analysis across five hyperspectral datasets indicates that the proposed detector consistently attains the highest AUC scores and exhibits improved ROC curves in comparison to established benchmarks.
Guiwei Liu, Guohe Li, Ye Zhu 0002, Guangmao Zhao, Chunyu Qi
IEEE Geosci. Remote. Sens. Lett.7
2023 ViTframe: Vision Transformer Acceleration via Informative Frame Selection for Video Recognition
abstract
Vision Transformer (ViT) has achieved state-of-the-art performance in the computer vision field, showcasing the remarkable potential to become a dominant model in the future. However, the self-attention mechanism within ViT presents significant challenges in terms of computational requirements and storage demands. This limitation becomes particularly pronounced in video recognition tasks, where the computational complexity escalates proportionally with the number of input frames. Current efforts to enhance ViT’s efficiency mainly concentrate on exploiting sparsity within individual frames, which neglects the temporal redundancy across frames, leading to an unsatisfactory solution for video tasks. Alternatively, in this paper, we propose a Vision Transformer acceleration framework called ViTframe, which aims to omit temporal redundancy in the video by dynamically selecting informative frames fed into ViT for fast video recognition. We first introduce an informative frame selection algorithm to pick out the most representative frames from the video for a quick ViT inference and a result compensation mechanism to compensate for the accuracy loss incurred by the informative frame selection. Moreover, we offer a customized architecture to efficiently implement the ViTframe algorithm. We implement ViTframe in a 28nm technology node. Extensive evaluations verify the effectiveness of our proposal on speedup, energy, and accuracy.
Chunyu Qi, Zhuoran Song, Xiaoyao Liang
ICCD1
2023 Multi-step Prediction of LTE-R Communication Quality based on CA-TCN and Differential Evolution
abstract
With the continuous development of heavy-haul railway technology, the demand for the faster transmission and higher bandwidth capacity of wireless communication network is also increasing. Therefore, Long Term Evolution for Railway (LTE-R) begin to replace Global System for Mobile Communications for Railway (GSM-R) to burden the core wireless communication services. In order to improve the operation and maintenance efficiency of LTE-R network, this paper proposes a multi-step LTE-R communication quality prediction method based on Differential Evolution algorithm and Temporal Convolutional Network with Coordinate Attention (TCNCA). Firstly, this method uses differential evolution and permutation importance index to filter the features of multivariate LTE-R communication quality data. Then, the coordinate attention mechanism and TCN network were fused to build the prediction model. Finally, this method still used Differential Evolution to adjust the model parameters based on Mann—Kendall test, so that the model could take into account the prediction accuracy and the trend of data change, thus realize the multi-step prediction for LTE-R communication quality. Experimental results on real data set show that the proposed method can provide decision support for the active maintenance of LTE-R network, and has high application value.
Jiantao Qu, Chunyu Qi, Gaoyun An, He
TrustCom2