EDBT 2026 Demo / reviewers in the wild / expert
Haishuang Fan
dblp:348/4989
· DBLP profile ↗
12ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-4633-4044ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 7 first-author · 12 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RAPID: Accelerating Point Cloud Diffusion Models via Space-Aware Mix-Precision QuantizationabstractPoint cloud diffusion models, as an emerging 3D generation method, hold broad prospects in 3D modeling, AR/VR, and so on. However, their reliance on costly full-precision neural network computations during extended denoising process limits their practical application. To address this challenge, we propose RAPID, an accelerator co-designed with a space-aware quantization method. First, RAPID uses K-means to partition points into groups and computes scaling factors in each, mitigating accuracy issues caused by uneven distribution. Second, it employs a mixed-precision quantization scheme that uses low precision for internal point groups and high precision for detail-rich edge groups, ensuring generation quality while minimizing bit-width. Third, it reuses computation results for groups with little change between timesteps, reducing redundant calculations. Moreover, RAPID’s hardware features a mixed-precision PE array for efficient computations at various bit-widths, and a filter for dynamic bit-width allocation and result reuse. Evaluations show that, compared to the NVIDIA RTX A5000 GPU and state-of-the-art accelerators, RAPID achieves average speedups of 9.22×, 4.66×, 3.69×, and 3.01×, and energy savings of 61.74×, 4.30×, 3.94×, and 2.76×, with negligible accuracy loss. Qichu Sun, Linxi Lu, Haishuang Fan, Jingya Wu, Huawei Li 0001, Xiaowei Li 0001, Guihai Yan |
DATE | 4 |
| 2025 | APTO: Accelerating Serialization-Based Point Cloud Transformers with Position-Aware PruningabstractPoint cloud processing has broad applications in autonomous driving and robotics. Serialization-based point cloud transformers map unordered point clouds onto directed curves, use sparse convolution for down-sampling and apply attention in local windows to capture spatial relationships. Despite achieving great accuracy, these models face inference latency challenges: neighbor search in sparse convolution exhibits low parallelism; attention computation remains complex, especially with larger window sizes; softmax introduces data dependencies. This paper proposes APTO, an accelerator for serialization-based models. It uses voxels' z-curve indices to perform neighbor searches in parallel, employs a position-aware pruning strategy using neighboring voxel counts to eliminate useless attention computations, and adopts a fine-grained attention dataflow for parallel processes with minimal data dependencies. Besides, its hardware has dedicated computation cores for efficient processing. Evaluations show that APTO achieves average 10.22×, 3.53× and 2.70× speedups over RTX 4090 GPU, PointAcc, and SpOctA, with 153.59×, 8.57× and 7.25× energy savings. Qichu Sun, Haishuang Fan, Fangqiang Ding, Linxi Lu, Jingya Wu, Xiaowei Li 0001, Guihai Yan |
ASP-DAC | 3 |
| 2025 | Flame: A Multiplier-Free LLM Accelerator with Dynamic Block Floating PointabstractRecently, large language models (LLMs) have achieved remarkable success across various machine learning tasks. However, deploying LLM inference still remains challenging due to the high computing overhead and substantial memory footprint. In this paper, we propose Flame, a software-hardware co-design accelerator that enables efficient LLM inference with dynamic block floating point (BFP). Flame employs a layer-wise adaptive precision search algorithm to optimize BFP mantissa bitwidth to balance model accuracy and inference speedup. To mitigate accuracy degradation introduced by BFP conversion, we propose a channel reorder approach that adjusts the value distribution within each tensor group. Finally, we leverage a novel hardware-friendly linear complexity multiplication algorithm to implement an efficient hardware accelerator featuring multiplierfree processing units. Evaluation results show that Flame achieves up to$6.32 \times$speedup and$4.04 \times$energy efficiency compared to GPU, while maintaining superior model accuracy. Ao Lyu, Haishuang Fan, Guihai Yan |
ICCD | 2 |
| 2025 | KPU: Kernel Processing Unit for in-Memory Analytical Query ProcessingabstractDomain-specific architecture has greatly improved performance and energy efficiency in in-memory databases, especially for accelerating single-functional computing logic in analytic query processing, such as sort, join and aggregation. However, as data volumes surge exponentially, these dedicated accelerators are struggling to satisfy the burgeoning demand for handling intricate and multifaceted workloads. A major challenge lies in establishing a flexible framework that engages these ‘coarse-grained’ units without incurring extra overheads from hardware integration, programming, compilation, runtime and operating systems.In this paper, the kernel processing unit (KPU) is proposed to optimize CPU-accelerator heterogeneous systems for in-memory databases. KPU provides a unified interface to consolidate all database query operators. In terms of KPU hardware architecture, kernel customization and data transmission are two critical bottlenecks. To address the challenges, multiple independently designed homogeneous table cores are integrated to support flexible high-performance SQL queries, and a customized efficient data management system (DMS) works collaboratively to maximize the utilization of on-chip memory bandwidth. Additionally, a database application-specific KPU instruction set architecture (KISA) dedicated to parallel analytical query processing is proposed to enable parallel KPU programming. To trade off between accelerator computing capacity and data transfer latency, KPU designs an offloading mechanism to map SQL queries between the CPU and accelerator adaptively based on a performance model and a function simulator. The experiments demonstrate that KPU surpasses the general-purpose CPU and GPU by an average of 24.5× and 8.75×, respectively. Jingya Wu, Wenyan Lu, Haishuang Fan, Hao Kong 0005, Xiaowei Li 0001, Guihai Yan |
IEEE Trans. Computers | 3 |
| 2025 | GRACE: An End-to-End Graph Processing Accelerator on FPGA With Graph Reordering EngineabstractGraphs play an important role in various applications. With the rapid expansion of vertices in real life, existing large-scale graph processing frameworks on CPUs and GPUs encounter challenges in optimizing cache usage due to irregular memory access patterns. To address this, graph reordering has been proposed to improve the locality of the graph, but introduces significant overhead without delivering substantial end-to-end performance improvement. While there have been many FPGA-based accelerators for graph processing, achieving high throughput often requires complex graph prepossessing on CPUs. Therefore, implementing an efficient end-to-end graph processing system remains challenging. This article introduces GRACE, an end-to-end FPGA-based graph processing accelerator with a graph reordering engine and a pull-based vertex-centric programming model (PL-VCPM) Engine. First, GRACE employs a customized high-degree vertex cache (HDC) to improve memory access efficiency. Second, GRACE offloads the graph preprocessing to FPGA. We customize an efficient graph reordering engine to complete preprocessing. Third, GRACE adopts a graph pruning strategy to remove the activation and computation redundancy in graph processing. Finally, GRACE introduces a graph conflict board (GCB) to resolve data conflicts and a multiport cache to enhance parallel efficiency. Experimental results demonstrate that GRACE achieves$7.1 \times $end-to-end performance speedup over CPU and$1.8 \times $over GPU, as well as$27.3 \times $and$8.7 \times $energy efficiency over CPU and GPU. Moreover, GRACE delivers up to$34.9 \times $performance speedup compared to the state-of-the-art FPGA accelerator. Haishuang Fan, Qichu Sun, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | Co-ViSu: Accelerating Video Super-Resolution With Codec Information ReuseabstractHigh-resolution (HR) videos have gained popularity with the widespread adoption of high-definition displays. Super-resolution (SR) techniques aim to recover HR frames from low-resolution (LR) frames. While deep neural network (DNN)-based SR methods have outperformed traditional techniques in quality, they face performance challenges. FPGA-based SR accelerators have been developed to optimize the performance and power efficiency. However, most of these accelerators process only uncompressed video frames and perform per-frame DNN inference, overlooking the temporal-spatial information inherent in compressed video bitstreams. We propose a novel compressed video SR workflow that includes a codec information reuse algorithm and a dedicated FPGA accelerator named Co-ViSu. Our approach leverages the observation that non-key frames can be reconstructed using codec information and HR key-frames, significantly reducing DNN computations. The Co-ViSu algorithm employs subpixel interpolation to enhance high-frequency details and an MV-aware method to improve SR reconstruction quality. The Co-ViSu hardware integrates decoder, SR, and encoder engines within a parallel pipeline architecture, utilizing codec information reuse to bypass non-key frame decoding, eliminate complex DNN computations, and accelerate encoding processes. Experimental results demonstrate that Co-ViSu achieves performance improvements ranging from$3.6\times $to$9.4\times $and a$4.2\times $gain in energy efficiency with minimal quality loss compared to traditional flow. Additionally, Co-ViSu offers a$2.1\times $increase in throughput compared to state-of-the-art solutions. Haishuang Fan, Qichu Sun, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Co-Via: A Video Frame Interpolation Accelerator Exploiting Codec Information ReuseabstractVideo Frame Interpolation (VFI) aims to generate intermediate frames between consecutive frames. Recent DNN-based VFI offers superior quality but suffers from performance issues. However, very few studies have focused on VFI hardware acceleration and existing work overlooks temporal information from compressed video bitstreams. In this paper, we propose a novel compressed VFI workflow and an accelerator, Co-Via. Co-Via exploits codec information reuse to reduce complex DNN computations and alleviate hardware pressure. FPGA-based Co-Via outperforms an RTX 4090 GPU 10.31X, offering a 43.08X energy efficiency boost. Its ASIC version achieves 2.4X higher throughput and 3.6X energy efficiency than the state-of-the-art solution. Haishuang Fan, Qichu Sun, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
DAC | 1 |
| 2024 | AMST: Accelerating Large-Scale Graph Minimum Spanning Tree Computation on FPGAabstractThe minimum spanning tree (MST) plays an important role in variant fields, such as chip design and network analysis. With the rapid expansion of vertices in real-life graphs, the bottleneck problem of MST algorithms in large-scale graphs grows more prominent. While there have been many FPGA-based accelerators for large-scale graph algorithms such as Graph Random Walk, and various algorithms to accelerate MST on CPUs and GPUs, effectively implementing MST algorithms for large-scale graphs on FPGAs remains quite challenging. This is due to several reasons: The neighbor vertices in the graph require extensive random memory access and the memory access characteristics vary across different stages and iterations. There are a large number of useless computations due to the existence of internal edges within a component (intra-edge). Parallel MST algorithm suffers from significant communication overhead due to the minimum edge data update conflicts and memory read-write conflicts.This paper proposes AMST to accelerate large-scale graph MST computation on FPGA. First, AMST employs a customized hash-based high-degree vertex cache (HDC) to improve memory access efficiency. Second, AMST adopts a graph pruning strategy that skips intra-edge and sorts edges by weight to eliminate useless computation and memory access. Finally, AMST utilizes a sorting networking module and a multi-port HDC to improve parallel efficiency. The experimental results demonstrate that AMST achieves an average performance speedup of 17.52× over CPU and 1.89× over GPU, as well as 74.96× over CPU and 10.45× over GPU on energy efficiency. Haishuang Fan, Qichu Sun, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
IPDPS | 1 |
| 2023 | Co-ViSu: a Video Super-Resolution Accelerator Exploiting Codec Information ReuseabstractHigh-resolution (HR) videos have become popular due to the widespread adoption of high-definition displays. Super-resolution (SR) techniques aim to recover HR frames from low-resolution (LR) frames. Recently, deep neural network (DNN)-based SR methods have achieved superior quality compared to traditional methods. FPGA-based SR accelerators have been proposed to optimize performance and power efficiency. However, most accelerators tailored for video SR only accept uncompressed video frames and operate per-frame DNN inference, ignoring the temporal-spatial information in compressed video bitstreams. In contrast, we observe that non-key frames can be directly constructed using codec information and HR key-frames, saving a significant amount of DNN computing. In this paper, we propose a novel compressed video SR flow and a specific FPGA accelerator called Co-ViSu that integrates decoder, SR, and encoder engines. Co-ViSu exploits codec information reuse scheme to skip non-key frame decoding, avoid complex DNN computation and speed up encoding. Our experimental results show that Co-ViSu achieves 3.6x to 9.4x performance, 4.2x energy efficiency gain with only 0.17dB quality loss compared to the traditional flow, and 2.1x throughput than state-of-the-art. Haishuang Fan, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
FPL | 1 |
| 2023 | M2VT: A Multi-Output Encoder Accelerator for Multiple-Way Video TranscodingabstractVideo transcoding is a general but compute-intensive technology in video streaming services. Traditional single-encoder accelerators transcode multiple streams independently in the multi-output scenario. However, this mode neglects redundant computation and introduces high hardware complexity. To solve these issues, we propose a multi-encoder accelerator supporting reuse scheme. We introduce four fast algorithms based on parameter sharing to simplify encoding complexity. To further optimize the architecture, we also propose the standalone stream insertion (SSI) to increase the pipeline efficiency, and co-optimize memory access. Implementation results show that multi-encoder can reduce 68.03% computation complexity. Moreover, the area and power efficiency improve 3.05x and 2.62x. Haishuang Fan, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | KPU-SQL: Kernel Processing Unit for High-Performance SQL AccelerationabstractApplication-specific accelerator is a prominent way for analytic query processing. To achieve a substantial improvement over the state-of-the-art in performance while maintaining programmability, we propose a kernel processing unit (KPU) framework and apply it to SQL acceleration. Kernel customization and data transmission are two critical bottlenecks, we separately optimize them in the key core and shadow core with a self-designed data management system. A software stack named RACE with a performance model and function simulator is also introduced. The experiments demonstrate that KPU-SQL outperforms the CPU and GPU by 24.5x and 8.75x on average, respectively. Hao Kong 0005, Haishuang Fan, Jingya Wu, Liyun Cheng, Wenyan Lu, Guihai Yan, Xiaowei Li 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | BitColor: Accelerating Large-Scale Graph Coloring on FPGA with Parallel Bit-Wise EnginesabstractThe graph coloring algorithm plays a crucial role in many applications such as social network analysis. However, since the minimal graph coloring problem is NP-complete, which is increasingly computationally and memory-intensive as the number of vertices in the graph grows rapidly. Despite numerous FPGA-based works proposed to accelerate large-scale graph processing algorithms, such as Single Source Shortest Path, and various coloring algorithms, such as linear programming algorithms, efficiently implementing the greedy coloring algorithm for large-scale graphs on FPGA still remains highly challenging due to several reasons: ① The coloring algorithm requires color state traversal to determine the final color after traversing neighbor vertices. The time complexity of color traversal is equal to the neighbor vertices traversal, which is inefficient. ② Neighbor vertices traversal requires extensive random memory accesses on vertex color data. ③ Coloring different vertices in parallel is difficult due to potential color update conflicts between adjacent vertices. Haishuang Fan, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
ICPP | 1 |