EDBT 2026 Demo / reviewers in the wild / expert
Yanhuan Liu
dblp:261/4261
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BitRed: Taming Non-Uniform Bit-Level Sparsity with a Programmable RISC-V ISA for DNN AccelerationabstractThe non-uniform and dynamic nature of Bit-Level Sparsity (BLS) poses a critical load-imbalance challenge for parallel hardware accelerators. While the Bit-Interleaving paradigm, represented by state-of-the-art accelerators like Bitlet, shows promise, it is fundamentally constrained by a rigid datapath and severe inter-channel load imbalance. This paper introduces BitRed, an accelerator that embodies a new ''programmable adaptive bit-interleaving'' philosophy. Rather than a monolithic design, BitRed's core is an Adaptive-Sparse Processing Unit (ASPU) that deconstructs the acceleration process into a set of orthogonal RISC-V ISA extensions for pre-processing (cal.pre), adaptive distillation with dynamic load balancing (cal.adis), and PDP-optimal reduction (cal.red). By transforming a rigid hardware problem into a flexible scheduling problem, this ISA-based approach provides a fundamentally more adaptable and extensible solution. Empirical studies on a broad set of benchmarks highlight the following results (normalized to a SCNN baseline): (1) up to 9.4× speedup over the Bitlet, and 5.6× over the latest bit-serial SOTA BitWave; (2) up to 7.6× higher inference efficiency than Bitlet on representative models; (3) 5.072mm2 area and scalable power consumption from 550.43mW (float32) to 495.12mW (16b) and 457.90mW (8b) @28nm TSMC; and (4) high versatility across precisions, and up to 18.9×13.5× higher than NVIDIA A100 and Jetson Orin 32GB, demonstrating significant competitiveness against GPUs. Yanhuan Liu, Kunming Zhang, Yuqun Liu, Siao Wen, Lexin Wang, Tianyu Liu 0007, Zhihua Fan, Xiaochun Ye, Dongrui Fan, Xuejun An |
ASPLOS (2) | 1 |
| 2026 | MLX: Multi-Layer Execution for Structured LLM Workload Acceleration on Spatial Architectures
Zhihua Fan, Zirui Ma, Yuqun Liu, Tengfei Xia, Yanhuan Liu, Kunming Zhang, Xiaochun Ye, Dongrui Fan, Jian Weng 0002 |
ISCA | 7 |
| 2026 | HALO: A heterogeneous accelerator for low-latency and energy-efficient edge LLM inference
Kunming Zhang, Zhihua Fan, Yanhuan Liu, Lexin Wang, Yuqun Liu, Xiaochun Ye |
Future Gener. Comput. Syst. | 3 |
| 2025 | LightCacheRL: A Lightweight Reinforcement Learning Framework for Unified Cache Management
Kunming Zhang, Zhihua Fan, Yingchun Fu, Yanhuan Liu, Lexin Wang, Yuqun Liu |
NPC (1) | 4 |
| 2025 | DFU-E: A Dataflow Architecture for Edge DSP and AI ApplicationsabstractEdge computing aims to enable swift, real-time data processing, analysis, and storage close to the data source. However, edge computing platforms are often constrained by limited processing power and efficiency. This paper presents DFU-E, a dataflow-based accelerator specifically designed to meet the demands of edge digital signal processing (DSP) and artificial intelligence (AI) applications. Our design addresses real-world requirements with three main innovations. First, to accommodate the diverse algorithms utilized at the edge, we propose a multi-layer dataflow mechanism capable of exploiting task-level, instruction block-level, instruction-level, and data-level parallelism. Second, we develop an edge dataflow architecture that includes a customized processing element (PE) array, memory, and on-chip network microarchitecture optimized for the multi-layer dataflow mechanism. Third, we design an edge dataflow software stack that enables automatic optimizations through operator fusion, dataflow graph mapping, and task scheduling. We utilize representative real-world DSP and AI applications for evaluation. Comparing with Nvidia's state-of-the-art edge computing processor, DFU-E achieves up to 1.42× geometric mean performance improvement and 1.27× energy efficiency improvement. Zhihua Fan, Tianyu Liu 0007, Zhen Wang 0045, Meng Wu 0006, Kunming Zhang, Yanhuan Liu, Ninghui Sun, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2024 | OBSD: On-The-Fly Block-Wise Sparse Distillation Accelerating SpGEMMs in DNN ApplicationsabstractThe mainstream deep neural networks (DNNs) widely adopt pruning techniques to alleviate network overfitting and computational complexity. Simultaneously, the weight and activation data exhibit increasingly noticeable block-wise sparsity. Current DNN-specific accelerators mainly focus on element-wise while neglecting sparse features of block granularity. Meanwhile, one of the state-of-the-art work, HIRAC, combining matrix tiling with the fast packing algorithm, SorPack, to achieve effective acceleration of Sparse General Matrix Multiplications (SpGEMMs). However, the SorPack algorithm executes on the host CPU, and its runtime lies on the critical path of the overall execution. In this work, a specific SpGEMM accelerator named OBSD is proposed which achieves high performance and energy efficiency. An on-the-fly block-wise distillation approach is proposed for leveraging the block-wise sparsity in both weight and activation data, which is implemented during the data loading process, consuming minimal additional time. Moreover, a data-flow architecture is devised to improve efficiency of data exchange among data processing elements (DPs) with investigating the relationship for different partition sizes, block-wise sparsity, and block-wise distance. The evaluation results demonstrate that: (1) OBSD achieves average of 1.89× speedup as compared to HIRAC for representative matrices in DNN workloads. (2) An end-to-end evaluation on a DNN model shows a 2.41× speedup and energy efficiency improvement of 1.99× over the HIRAC. (3) OBSD possesses a power consumption 12.9W with 56.6mm2area @28nm TSMC. Yanhuan Liu, Kunming Zhang, Zhihua Fan, Lexin Wang, Tianyu Liu 0007, Zhen Wang 0045, Xiaochun Ye, Dongrui Fan, Xuejun An |
HPCC | 1 |
| 2023 | Alleviating Transfer Latency in DataFlow Accelerator for DSP ApplicationsabstractTowards multiple domains, dataflow accelerators show superiority for their flexible programmability and high efficiency. This efficiency relies highly on data communication between processing elements (PEs), which is sensitive to PE location, array scale and workload size. Laying out instructions as a dataflow graph on the PE array creates more instruction-level parallelism. However, the farther distance between remote PEs and memory banks introduces extra transfer latency, bringing performance degradation to high real-time applications. This paper examines the workloads of digital signal processing across different data scales and classifies latency problems related to data transfers and kernel switching. Specifically, we propose a novel forwarding network on chip to alleviate transfer latency and improve multi-destination sharing in the dataflow execution. Moreover, we devise bandwidth reusing mechanism to speedup kernel switching. The experiment results show that our scalable design achieves up to 2.19× (1.45× on average) speedup while reducing switching overhead by 9.85×, with an area overhead of 10.82% over the conventional dataflow accelerator. Zhihua Fan, Zhen Wang 0045, Tianyu Liu 0007, Junying Huang, Shengzhong Tang, Yanhuan Liu, Kunming Zhang, Xiaochun Ye, Dongrui Fan |
ICCD | 8 |
| 2023 | Accelerating Convolutional Neural Networks by Exploiting the Sparsity of Output ActivationabstractDeep Convolutional Neural Networks (CNNs) are the most widely used family of machine learning methods that have had a transformative effect on a wide range of applications. Previous studies have made great breakthroughs in accelerating CNNs, but they only target on the input sparsity of activation and weight, thus do not eliminate the unnecessary computations due to the fact that more zeros in the output results are not directly caused by the zero-valued positions of the input data. In this paper, we take advantage of the output activation sparsity to reduce the execution time and energy consumption of CNNs. First, we propose an effective prediction method that leverages the output activation sparsity. Our method first predicts the output activation polarity of convolutional layers based on the singular value decomposition (SVD) approach. Then, it uses the predicted negative value to skip invalid computations. Second, an effective accelerator is designed to take advantage of sparsity to achieve CNN inference acceleration. Each PE is equipped with a prediction unit and a non-zero value detection unit to remove invalid computation blocks. And an instruction bypass technique is proposed which further exploits the sparsity of the weights. The efficient dataflow graph mapping approach and pipeline execution ensure high computational resource utilization. Experiments show that our approach achieves up to 1.63× speedup and 55.30% energy reduction compared with dense networks with a slight loss of accuracy. Compared with Eyeriss, our accelerator achieves on average 1.31 × performance improvement and 54% energy reduction. Our accelerator also achieves a similar performance to SnaPEA, but with a better energy efficiency. Zhihua Fan, Zhen Wang 0045, Tianyu Liu 0007, Yanhuan Liu, Meng Wu 0006, Xinxin Wu, Xiaochun Ye, Dongrui Fan, Ninghui Sun, Xuejun An |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | Distilling Bit-level Sparsity Parallelism for General Purpose Deep Learning AccelerationabstractAlong with the rapid evolution of deep neural networks, the ever-increasing complexity imposes formidable computation intensity to the hardware accelerator. In this paper, we propose a novel computing philosophy called “bit interleaving” and the associate accelerator design called “Bitlet” to maximally exploit the bit-level sparsity. Apart from existing bit-serial/parallel accelerators, Bitlet leverages the abundant “sparsity parallelism” in the parameters to enforce the inference acceleration. Bitlet is versatile by supporting diverse precisions on a single platform, including floating-point 32 and fixed-point from 1b to 24b. The versatility enables Bitlet feasible for both efficient inference and training. Empirical studies on 12 domain-specific deep learning applications highlight the following results: (1) up to 81 × /21 × energy efficiency improvement for training/inference over recent high performance GPUs; (2) up to 15 × /8 × higher speedup/efficiency over state-of-the-art fixed-point accelerators; (3) 1.5mm2 area and scalable power consumption from 570mW (float32) to 432mW (16b) and 365mW (8b) @28nm TSMC; (4) highly configurable justified by ablation and sensitivity studies. Liang Chang 0002, Zixuan Zhu 0001, Shengjian Lu, Yanhuan Liu, Mingzhe Zhang 0005 |
MICRO | 6 |