VLDB 2026 Research / reviewers in the wild / expert
Haikuo Shao
dblp:300/5604
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0009-0008-6965-3436ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM AccelerationabstractLarge language models (LLMs) have revolutionized AI applications, yet their enormous computational demands severely limit deployment and real-time performance. Quantization methods can help reduce computational costs, however, attaining the extreme efficiency associated with ultra-low-bit quantized LLMs at arbitrary precision presents challenges on GPUs. This is primarily due to the limited support for GPU Tensor Cores, inefficient memory management, and inflexible kernel optimizations. To tackle these challenges, we propose a comprehensive acceleration scheme for arbitrary precision LLMs, namely APT-LLM. Firstly, we introduce a novel data format, bipolar-INT, which allows for efficient and lossless conversion with signed INT, while also being more conducive to parallel computation. We also develop a matrix multiplication (MatMul) method allowing for arbitrary precision by dismantling and reassembling matrices at the bit level. This method provides flexible precision and optimizes the utilization of GPU Tensor Cores. In addition, we propose a memory management system focused on data recovery, which strategically employs fast shared memory to substantially increase kernel execution speed and reduce memory access latency. Finally, we develop a kernel mapping method that dynamically selects the optimal configurable hyperparameters of kernels for varying matrix sizes, enabling optimal performance across different LLM architectures and precision settings. In LLM inference, APT-LLM achieves up to a 3.99× speedup compared to FP16 baselines and a 2.16× speedup over NVIDIA CUTLASS INT4 acceleration on RTX 3090. On RTX 4090 and H800, APT-LLM achieves up to 2.44× speedup over FP16 and 1.65× speedup over CUTLASS integer baselines. Shaobo Ma, Chao Fang 0005, Haikuo Shao, Zhongfeng Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | AccLLM: Accelerating Long-Context LLM Inference via Algorithm-Hardware Co-Design
Yanbiao Liang, Huihong Shi, Haikuo Shao, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor CoresabstractLarge language models (LLMs) have been widely applied but face challenges in efficient inference. While quantization methods reduce computational demands, ultra-low bit quantization with arbitrary precision is hindered by limited GPU Tensor Core support and inefficient memory management, leading to suboptimal acceleration. To address these challenges, we propose a comprehensive acceleration scheme for arbitrary precision LLMs. At its core, we introduce a novel bipolar-INT data format that facilitates parallel computing and supports symmetric quantization, effectively reducing data redundancy. Building on this, we implement an arbitrary precision matrix multiplication scheme that decomposes and recovers matrices at the bit level, enabling flexible precision while maximizing GPU Tensor Core utilization. Furthermore, we develop an efficient matrix preprocessing method that optimizes data layout for subsequent computations. Finally, we design a data recovery-oriented memory management system that strategically utilizes fast shared memory, significantly enhancing kernel execution speed and minimizing memory access latency. Experimental results demonstrate our approach's effectiveness, with up to 2.4× speedup in matrix multiplication compared to NVIDIA's CUTLASS. When integrated into LLMs, we achieve up to 6.7× inference acceleration. These improvements significantly enhance LLM inference efficiency, enabling broader and more responsive applications of LLMs. Shaobo Ma, Chao Fang 0005, Haikuo Shao, Zhongfeng Wang 0001 |
ASP-DAC | 3 |
| 2025 | An Efficient Training Architecture for Nonlinear Softmax Function in TransformersabstractTransformers have achieved significant success in modern deep learning. However, their intensive and complicated computations challenge Transformers’ deployment on resource-constrained devices, especially the training process. Softmax, the crucial nonlinear function in Transformers, features complex exponential and division operations. Unfortunately, existing methods cannot effectively reduce the hardware complexity of Softmax while maintaining training accuracy. In this paper, we propose an efficient architecture for Softmax training. Specifically, we present a quantized Softmax training algorithm based on integer arithmetic with sufficient training accuracy. Additionally, we design a reconfigurable architecture to efficiently support various operations during different training phases, along with a staged parallel and pipelined (SPAP) dataflow to reduce latency and improve hardware efficiency. Experimental results show that our architecture achieves up to 1.0 Softmax GinS in throughput and 4.42 GinS/W in energy efficiency at 500MHz on the Xilinx ZCU102 FPGA, which significantly outperforms prior works. Haikuo Shao, Zhongfeng Wang 0001 |
ISCAS | 1 |
| 2025 | Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision TransformerabstractMotivated by the huge success of Transformers in the field of natural language processing (NLP), Vision Transformers (ViTs) have been rapidly developed and achieved remarkable performance in various computer vision tasks. However, their huge model sizes and intensive computations hinder ViTs’ deployment on embedded devices, calling for effective model compression methods, such as quantization. Unfortunately, due to the existence of hardware-unfriendly and quantization-sensitive non-linear operations, particularly Softmax, it is non-trivial to completely quantize all operations in ViTs, yielding either significant accuracy drops or non-negligible hardware costs. In response to challenges associated with standard ViTs, we focus our attention towards the quantization and acceleration for efficient ViTs, which not only eliminate the troublesome Softmax but also integrate linear attention with low computational complexity, and propose Trio-ViT accordingly. Specifically, at the algorithm level, we develop a tailored post-training quantization engine taking the unique activation distributions of Softmax-free efficient ViTs into full consideration, aiming to boost quantization accuracy. Furthermore, at the hardware level, we build an accelerator dedicated to the specific Convolution-Transformer hybrid architecture of efficient ViTs, thereby enhancing hardware efficiency. Extensive experimental results consistently prove the effectiveness of our Trio-ViT framework. Particularly, we can gain up to$\uparrow {3.6}\times $,$\uparrow {5.0}\times $, and$\uparrow {7.3}\times $FPS under comparable accuracy over state-of-the-art ViT accelerators, as well as$\uparrow {6.0}\times $,$\uparrow {1.5}\times $, and$\uparrow {2.1}\times $DSP efficiency. Codes are available athttps://github.com/shihuihong214/Trio-ViT. Huihong Shi, Haikuo Shao, Wendong Mao, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | ASTRA: Reconfigurable Training Architecture Design for Nonlinear Softmax and Activation Functions in TransformersabstractThe efficient training of Transformer-based neural networks on resource-constrained personal devices is attracting continuous attention due to domain adaptions and privacy concerns. However, Transformers’ intensive and complicated computations, especially the crucial nonlinear Softmax and activation (Act) functions, pose challenges for training deployment on edge. This brief proposes an efficient training architecture for both Softmax and Act functions. Specifically, we present a quantized nonlinear training algorithm based on fully integer (int) arithmetic with sufficient training accuracy. Then, we develop a reconfigurable hardware architecture to efficiently support various operations during the training of these nonlinear functions. Furthermore, a staged parallel and pipelined (SPAP) dataflow is presented to reduce latency and improve hardware efficiency. Experimental results show that our architecture achieves up to 1.0 Softmax GinS in throughput and 4.42 GinS/W in energy efficiency at 500 MHz on Xilinx ZCU102 FPGA, outperforming prior works. Haikuo Shao, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | Co-Designing Binarized Transformer and Hardware Accelerator for Efficient End-to-End Edge DeploymentabstractTransformer models have revolutionized AI tasks, but their large size hinders real-world deployment on resource-constrained and latency-critical edge devices. While binarized Transformers offer a promising solution by significantly reducing model size, existing approaches suffer from algorithm-hardware mismatches with limited co-design exploration, leading to suboptimal performance on edge devices. Hence, we propose a co-design method for efficient end-to-end edge deployment of Transformers from three aspects: algorithm, hardware, and joint optimization. First, we propose BMT, a novel hardware-friendly binarized Transformer with optimized quantization methods and components, and we further enhance its model accuracy by leveraging the weighted ternary weight splitting training technique. Second, we develop a streaming processor mixed binarized Transformer accelerator, namely BAT, which is equipped with specialized units and scheduling pipelines for efficient inference of binarized Transformers. Finally, we co-optimize the algorithm and hardware through a design space exploration approach to achieve a global trade-off between accuracy, latency, and robustness for real-world deployments. Experimental results show our co-design achieves up to 2.14~49.37× throughput gains and 3.72~88.53× better energy efficiency over state-of-the-art Transformer accelerators, enabling efficient end-to-end edge deployment. Yuhao Ji, Chao Fang 0005, Shaobo Ma, Haikuo Shao, Zhongfeng Wang 0001 |
ICCAD | 4 |
| 2024 | A Flexible FPGA-Based Accelerator for Efficient Inference of Multi-Precision CNNsabstractMulti-precision (MP) convolutional neural networks (CNNs) have exploited quantization techniques to achieve notable computation reductions while maintaining accuracy. However, most existing accelerators lack sufficient support for MP multiplications, hindering their ability to satisfy the diverse precision requirements of different layers in MP CNNs. To address this issue, we propose a flexible FPGA-based accelerator that efficiently processes CNN inference, supporting both symmetric and asymmetric bit-width computation. Specifically, a reconfigurable computing unit called MP-MAC is specially designed for efficient execution of MP computations to maximize computation utilization within a single DSP. Additionally, an optimized computing data arrangement, based on our flexible parallelism scheme, is presented to further enhance the performance of MP CNN deployment. Moreover, an MP performance model is introduced to estimate the transmission and computation latency, providing valuable guidance for efficient hardware design. The proposed accelerator achieves a throughput of up to 660.60 GOPS on Intel Arria 10 SoC FPGA, with 3.77× better DSP efficiency compared to prior work when evaluated on the same network. Xinyan Liu 0001, Haikuo Shao, Zhongfeng Wang 0001 |
ISCAS | 3 |
| 2024 | An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViTabstractVision Transformers (ViTs) have achieved significant success in computer vision. However, their intensive computations and massive memory footprint challenge ViTs’ deployment on embedded devices, calling for efficient ViTs. Among them, EfficientViT, the state-of-the-art one, features a Convolution-Transformer hybrid architecture, enhancing both accuracy and hardware efficiency. Unfortunately, existing accelerators cannot fully exploit the hardware benefits of EfficientViT due to its unique architecture. In this paper, we propose an FPGA-based accelerator for EfficientViT to advance the hardware efficiency frontier of ViTs. Specifically, we design a reconfigurable architecture to efficiently support various operation types, including lightweight convolutions and attention, boosting hardware utilization. Additionally, we present a time-multiplexed and pipelined dataflow to facilitate both intra- and inter-layer fusions, reducing off-chip data access costs. Experimental results show that our accelerator achieves up to 780.2 GOPS in throughput and 105.1 GOPS/W in energy efficiency at 200MHz on the Xilinx ZCU102 FPGA, which significantly outperforms prior works. Haikuo Shao, Huihong Shi, Wendong Mao, Zhongfeng Wang 0001 |
ISCAS | 1 |
| 2024 | A Low Complexity Online Learning Approximate Message Passing Detector for Massive MIMOabstractRecent research has shown that in massive multiple-input multiple-output (MIMO) detection, the model-driven machine learning (ML) detection algorithm with online training method can adapt to channel variations in real application scenarios and has high detection performance. However, the on-device training hardware of the ML-enhanced detector has not yet been addressed in the current literature. In this article, the architecture for the targeted hardware is designed through optimization on both the algorithms and the hardware. We first introduce a magnitude-based pruning strategy and then propose efficient MMNet (EMMNet)-type algorithms. In the improved algorithms, several algorithmic transformations or approximations are incorporated to reduce computational complexity. For instance, the exponential operation is replaced by a linear fitting function, and division is converted into hardware-efficient shift and subtraction operations. Moreover, to improve energy efficiency, the EMMNet-type algorithms are quantized with fixed-point (FXP) data formats and adopt a hardware-friendly stochastic gradient descent (SGD)-momentum optimizer. Based on the proposed algorithms, a low-complexity and high-throughput training architecture with reusable units is developed, which can support modulations from QPSK to QAM64. Compared with the MMNet, the presented EMMNet detector exhibits a 48% reduction in the number of multiplications without sacrificing detection performance, demonstrating remarkable hardware efficiency. Baoling Hong, Haikuo Shao, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | An Efficient Training Accelerator for Transformers With Hardware-Algorithm Co-OptimizationabstractTransformers have achieved significant success in deep learning, and training Transformers efficiently on resource-constrained platforms has been attracting continuous attention for domain adaptions and privacy concerns. However, deploying Transformers training on these platforms is still challenging due to its dynamic workloads, intensive computations, and massive memory accesses. To address these issues, we propose an Efficient Training Accelerator for TRansformers (TRETA) through a hardware-algorithm co-optimization strategy. First, a hardware-friendly mixed-precision training algorithm is presented based on a compact and efficient data format, which significantly reduces the computation and memory requirements. Second, a flexible and scalable architecture is proposed to achieve high utilization of computing resources when processing arbitrary irregular general matrix multiplication (GEMM) operations during training. These irregular GEMMs lead to severe under-utilization when simply mapped on traditional systolic architectures. Third, we develop training-oriented architectures for the crucial Softmax and layer normalization functions in Transformers, respectively. These area-efficient modules have unified and flexible microarchitectures to meet various computation requirements of different training phases. Finally, TRETA is implemented under Taiwan Semiconductor Manufacturing Company (TSMC) 28-nm technology and evaluated on multiple benchmarks. The experimental results show that our training framework achieves the same accuracy as the full precision baseline. Moreover, TRETA can achieve 14.71 tera operations per second (TOPS) and 3.31 TOPS/W in terms of throughput and energy efficiency, respectively. Compared with prior arts, the proposed design shows 1.4–$24.5\times $speedup and 1.5–$25.4\times $energy efficiency improvement. Haikuo Shao, Jinming Lu, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |