Qichu Sun

dblp:380/6683 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0008-1148-4297ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RAPID: Accelerating Point Cloud Diffusion Models via Space-Aware Mix-Precision Quantization
abstract
Point cloud diffusion models, as an emerging 3D generation method, hold broad prospects in 3D modeling, AR/VR, and so on. However, their reliance on costly full-precision neural network computations during extended denoising process limits their practical application. To address this challenge, we propose RAPID, an accelerator co-designed with a space-aware quantization method. First, RAPID uses K-means to partition points into groups and computes scaling factors in each, mitigating accuracy issues caused by uneven distribution. Second, it employs a mixed-precision quantization scheme that uses low precision for internal point groups and high precision for detail-rich edge groups, ensuring generation quality while minimizing bit-width. Third, it reuses computation results for groups with little change between timesteps, reducing redundant calculations. Moreover, RAPID’s hardware features a mixed-precision PE array for efficient computations at various bit-widths, and a filter for dynamic bit-width allocation and result reuse. Evaluations show that, compared to the NVIDIA RTX A5000 GPU and state-of-the-art accelerators, RAPID achieves average speedups of 9.22×, 4.66×, 3.69×, and 3.01×, and energy savings of 61.74×, 4.30×, 3.94×, and 2.76×, with negligible accuracy loss.
Qichu Sun, Linxi Lu, Haishuang Fan, Jingya Wu, Huawei Li 0001, Xiaowei Li 0001, Guihai Yan
DATE1
2025 APTO: Accelerating Serialization-Based Point Cloud Transformers with Position-Aware Pruning
abstract
Point cloud processing has broad applications in autonomous driving and robotics. Serialization-based point cloud transformers map unordered point clouds onto directed curves, use sparse convolution for down-sampling and apply attention in local windows to capture spatial relationships. Despite achieving great accuracy, these models face inference latency challenges: neighbor search in sparse convolution exhibits low parallelism; attention computation remains complex, especially with larger window sizes; softmax introduces data dependencies. This paper proposes APTO, an accelerator for serialization-based models. It uses voxels' z-curve indices to perform neighbor searches in parallel, employs a position-aware pruning strategy using neighboring voxel counts to eliminate useless attention computations, and adopts a fine-grained attention dataflow for parallel processes with minimal data dependencies. Besides, its hardware has dedicated computation cores for efficient processing. Evaluations show that APTO achieves average 10.22×, 3.53× and 2.70× speedups over RTX 4090 GPU, PointAcc, and SpOctA, with 153.59×, 8.57× and 7.25× energy savings.
Qichu Sun, Haishuang Fan, Fangqiang Ding, Linxi Lu, Jingya Wu, Xiaowei Li 0001, Guihai Yan
ASP-DAC1
2025 GRACE: An End-to-End Graph Processing Accelerator on FPGA With Graph Reordering Engine
abstract
Graphs play an important role in various applications. With the rapid expansion of vertices in real life, existing large-scale graph processing frameworks on CPUs and GPUs encounter challenges in optimizing cache usage due to irregular memory access patterns. To address this, graph reordering has been proposed to improve the locality of the graph, but introduces significant overhead without delivering substantial end-to-end performance improvement. While there have been many FPGA-based accelerators for graph processing, achieving high throughput often requires complex graph prepossessing on CPUs. Therefore, implementing an efficient end-to-end graph processing system remains challenging. This article introduces GRACE, an end-to-end FPGA-based graph processing accelerator with a graph reordering engine and a pull-based vertex-centric programming model (PL-VCPM) Engine. First, GRACE employs a customized high-degree vertex cache (HDC) to improve memory access efficiency. Second, GRACE offloads the graph preprocessing to FPGA. We customize an efficient graph reordering engine to complete preprocessing. Third, GRACE adopts a graph pruning strategy to remove the activation and computation redundancy in graph processing. Finally, GRACE introduces a graph conflict board (GCB) to resolve data conflicts and a multiport cache to enhance parallel efficiency. Experimental results demonstrate that GRACE achieves$7.1 \times $end-to-end performance speedup over CPU and$1.8 \times $over GPU, as well as$27.3 \times $and$8.7 \times $energy efficiency over CPU and GPU. Moreover, GRACE delivers up to$34.9 \times $performance speedup compared to the state-of-the-art FPGA accelerator.
Haishuang Fan, Qichu Sun, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 Co-ViSu: Accelerating Video Super-Resolution With Codec Information Reuse
abstract
High-resolution (HR) videos have gained popularity with the widespread adoption of high-definition displays. Super-resolution (SR) techniques aim to recover HR frames from low-resolution (LR) frames. While deep neural network (DNN)-based SR methods have outperformed traditional techniques in quality, they face performance challenges. FPGA-based SR accelerators have been developed to optimize the performance and power efficiency. However, most of these accelerators process only uncompressed video frames and perform per-frame DNN inference, overlooking the temporal-spatial information inherent in compressed video bitstreams. We propose a novel compressed video SR workflow that includes a codec information reuse algorithm and a dedicated FPGA accelerator named Co-ViSu. Our approach leverages the observation that non-key frames can be reconstructed using codec information and HR key-frames, significantly reducing DNN computations. The Co-ViSu algorithm employs subpixel interpolation to enhance high-frequency details and an MV-aware method to improve SR reconstruction quality. The Co-ViSu hardware integrates decoder, SR, and encoder engines within a parallel pipeline architecture, utilizing codec information reuse to bypass non-key frame decoding, eliminate complex DNN computations, and accelerate encoding processes. Experimental results demonstrate that Co-ViSu achieves performance improvements ranging from$3.6\times $to$9.4\times $and a$4.2\times $gain in energy efficiency with minimal quality loss compared to traditional flow. Additionally, Co-ViSu offers a$2.1\times $increase in throughput compared to state-of-the-art solutions.
Haishuang Fan, Qichu Sun, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 Co-Via: A Video Frame Interpolation Accelerator Exploiting Codec Information Reuse
abstract
Video Frame Interpolation (VFI) aims to generate intermediate frames between consecutive frames. Recent DNN-based VFI offers superior quality but suffers from performance issues. However, very few studies have focused on VFI hardware acceleration and existing work overlooks temporal information from compressed video bitstreams. In this paper, we propose a novel compressed VFI workflow and an accelerator, Co-Via. Co-Via exploits codec information reuse to reduce complex DNN computations and alleviate hardware pressure. FPGA-based Co-Via outperforms an RTX 4090 GPU 10.31X, offering a 43.08X energy efficiency boost. Its ASIC version achieves 2.4X higher throughput and 3.6X energy efficiency than the state-of-the-art solution.
Haishuang Fan, Qichu Sun, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan
DAC2
2024 AMST: Accelerating Large-Scale Graph Minimum Spanning Tree Computation on FPGA
abstract
The minimum spanning tree (MST) plays an important role in variant fields, such as chip design and network analysis. With the rapid expansion of vertices in real-life graphs, the bottleneck problem of MST algorithms in large-scale graphs grows more prominent. While there have been many FPGA-based accelerators for large-scale graph algorithms such as Graph Random Walk, and various algorithms to accelerate MST on CPUs and GPUs, effectively implementing MST algorithms for large-scale graphs on FPGAs remains quite challenging. This is due to several reasons: The neighbor vertices in the graph require extensive random memory access and the memory access characteristics vary across different stages and iterations. There are a large number of useless computations due to the existence of internal edges within a component (intra-edge). Parallel MST algorithm suffers from significant communication overhead due to the minimum edge data update conflicts and memory read-write conflicts.This paper proposes AMST to accelerate large-scale graph MST computation on FPGA. First, AMST employs a customized hash-based high-degree vertex cache (HDC) to improve memory access efficiency. Second, AMST adopts a graph pruning strategy that skips intra-edge and sorts edges by weight to eliminate useless computation and memory access. Finally, AMST utilizes a sorting networking module and a multi-port HDC to improve parallel efficiency. The experimental results demonstrate that AMST achieves an average performance speedup of 17.52× over CPU and 1.89× over GPU, as well as 74.96× over CPU and 10.45× over GPU on energy efficiency.
Haishuang Fan, Qichu Sun, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan
IPDPS3