EDBT 2026 Demo / reviewers in the wild / expert
Haidong Yao
dblp:155/6556
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BLOOM: Bit-Slice Framework for DNN Acceleration with Mixed-PrecisionabstractDeep neural networks (DNNs) have revolutionized numerous AI applications, but their vast model sizes and limited hardware resources present significant deployment challenges. Model quantization offers a promising solution to bridge the gap between DNN size and hardware capacity. While INT8 quantization has been widely used, recent research has pushed for even lower precision, such as INT4. However, the presence of outliers-values with unusually large magnitudes-limits the effectiveness of current quantization techniques. Previous compression-based acceleration methods that incorporate outlieraware encoding introduce complex logic. A critical issue we have identified is that serialization and deserialization dominate the encoding/decoding time in these compression workflows, leading to substantial performance penalties during workflow execution. To address this challenge, we introduce a novel computing approach and a compatible architecture design named “BLOOM”. BLOOM leverages the strengths of the “bit-slicing” method, effectively combining structured mixed-precision and bit-level sparsity with adaptive dataflow techniques. The key insight of BLOOM is that outliers require higher precision, while normal values can be processed at lower precision. By interleaving 4-bit values, we efficiently exploit the inherent sparsity in the highprecision components. As a result, the BLOOM-based accelerator outperforms the existing outlier-aware accelerators by an average $1.2 \sim 4.0 \times$ speedup and $24.6 \% \sim 71.3 \%$ energy reduction, respectively, without model accuracy loss. Fangxin Liu, Ning Yang 0012, Zongwu Wang, Xuanpeng Zhu, Haidong Yao, Xiankui Xiong, Li Jiang 0002, Haibing Guan |
DAC | 5 |
| 2025 | OPS: Outlier-Aware Precision-Slice Framework for LLM AccelerationabstractLarge language models (LLMs) have transformed numerous AI applications, with on-device deployment becoming increasingly important for reducing cloud computing costs and protecting user privacy. However, the astronomical model size and limited hardware resources pose significant deployment challenges. Model quantization is a promising approach to mitigate this gap, but the presence of outliers in LLMs reduces its effectiveness. Previous efforts addressed this issue by employing compression-based encoding for mixed-precision quantization. These approaches struggle to balance model accuracy with hard-ware efficiency due to their value-wise outlier granularity and complex encoding/decoding hardware logic. To address this, we propose OPS (Outlier-aware Precision-Slicing), an acceleration framework that exploits massive sparsity in the higher-order part of LLMs by splitting 16-bit values into a 4-bit/12-bit format. Crucially, OPS introduces an early bird mechanism that leverages the high-order 4-bit computation to predict the importance of the full calculation result. This mechanism enables efficient computational skips by continuing execution only for important computations and using preset values for less significant ones. This scheme can be efficiently integrated with existing hardware accelerators like systolic arrays without complex encoding/decoding. As a result, OPS outperforms state-of-the-art outlier-aware accelerators, achieving a 1.3 − 4.3× performance boost with minimal model accuracy loss. This approach enables more efficient on-device LLM deployment, effectively balancing computational efficiency and model accuracy. Fangxin Liu, Ning Yang 0012, Zongwu Wang, Xuanpeng Zhu, Haidong Yao, Xiankui Xiong, Li Jiang 0002 |
DATE | 5 |
| 2024 | FullSparse: A Sparse-Aware GEMM Accelerator with Online Sparsity PredictionabstractLeveraging sparsity optimizes storage and computation for resource-constrained devices in Deep Learning Neural Networks (DNNs). While neural networks naturally incorporate sparsity through operations like ReLU and quantization, diverse sparsity levels (0.2% to 99%) pose challenges for the design of computational units. In this paper, we provide an energy-efficient GEMM accelerator named FullSparse which is designed for diverse applications, accommodating varying sparsity levels in matrix multiplication (0.2% to 99%). This paper introduces three features for nuanced sparsity support: multi-sparsity control, predictive result sparsity, and a multi-sparsity-compatible PE array. Experimental evaluations affirm that our implementation while ensuring adaptability to sparsity, exhibits superior computational power comparable to the existing designs. Jiangnan Yu, Fan Yang 0001, Yuxuan Qiao, Xiankui Xiong, Haidong Yao, Yecheng Zhang |
CF | 8 |
| 2024 | FSMM: An Efficient Matrix Multiplication Accelerator Supporting Flexible SparsityabstractSparse matrix multiplication is a critical operation in deep learning. However, matrix sparsity leads to irregular data flow, which would degrade the efficiency of matrix multiplication. Traditional accelerators, equipped with additional hardware units to address this issue, often experience the issue of low hardware utilization. Furthermore, N : M structured sparsity and corresponding hardware architectures face challenges such as accuracy degradation, limited flexibility, and restricted applicability. In this paper, we propose a Flexible Sparse Matrix Multiplication Accelerator (FSMM), which can improve the efficiency of sparse matrix multiplication through both algorithmic-level and hardware-level optimizations. At the algorithmic level, we propose the matrix-matrix multiplication with block-level outer production and fine-grained matrix reordering algorithm. The algorithm balances the sparsity of each column of a matrix block, which improves matrix compression, balances the load, and speeds up computation. This algorithm reduces storage by 8.2% ~ 85.9%. At the hardware-level, we introduce a flexible architecture for matrix multiplication. It selects the most suitable data path to complete matrix multiplication based on the sparsity of the reordered matrix. FSMM achieves a speedup of 1.90× ~ 16.18× over Systolic Array and 1.70× ~ 2.87× over the existing TSTC approach. Yuxuan Qiao, Fan Yang 0001, Yecheng Zhang, Xiankui Xiong, Haidong Yao |
ICCAD | 6 |
| 2022 | An Automated Compiler for RISC-V Based DNN AcceleratorabstractMultifarious hardware accelerators are developed for the widely used Deep neural networks (DNN). Nowadays the SoCs composed of a general processor and a coupled accelerator are becoming prevalent. Compared to the specialized DNN accelerator for one specific DNN, this kind of coupled architecture is programmable and supports diverse DNNs. However, for the low-level programming interface of the co-processor-like accelerator and the multi-hierarchy memory structure, programming for the DNN accelerator is not easy work. Meanwhile, there are a couple of tensor compilers that deploy the DNN on various hardware. In this work, we combine the flexibility of the tensor compiler and the high efficiency of the hardware accelerator by proposing an automated compiler that can compile tensor programs and generate high-performance programs for programmable DNN accelerators. Our compiler is based on TVM [1] and target at Rocket Chip Coprocessor (RoCC) [2]. The compiler is flexible and supports many kinds of RISC-V instructions. The programmer can define the hardware constraints in the proposed compiler which makes the generated code more efficient. Our compiler can lower the program with the ping-pong strategy and the generated code can achieve 26% speed up compared to the baseline. Wuzhen Xie, Xiaoling Yi, Ruiyao Pu, Xiankui Xiong, Haidong Yao, Chixiao Chen, Jun Tao 0001, Fan Yang 0001 |
ISCAS | 7 |