Tianyu Liu 0007

dblp:134/1099-7 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0003-3905-6936ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-author · 12 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
YearPublicationVenuePosition
2026 BitRed: Taming Non-Uniform Bit-Level Sparsity with a Programmable RISC-V ISA for DNN Acceleration
abstract
The non-uniform and dynamic nature of Bit-Level Sparsity (BLS) poses a critical load-imbalance challenge for parallel hardware accelerators. While the Bit-Interleaving paradigm, represented by state-of-the-art accelerators like Bitlet, shows promise, it is fundamentally constrained by a rigid datapath and severe inter-channel load imbalance. This paper introduces BitRed, an accelerator that embodies a new ''programmable adaptive bit-interleaving'' philosophy. Rather than a monolithic design, BitRed's core is an Adaptive-Sparse Processing Unit (ASPU) that deconstructs the acceleration process into a set of orthogonal RISC-V ISA extensions for pre-processing (cal.pre), adaptive distillation with dynamic load balancing (cal.adis), and PDP-optimal reduction (cal.red). By transforming a rigid hardware problem into a flexible scheduling problem, this ISA-based approach provides a fundamentally more adaptable and extensible solution. Empirical studies on a broad set of benchmarks highlight the following results (normalized to a SCNN baseline): (1) up to 9.4× speedup over the Bitlet, and 5.6× over the latest bit-serial SOTA BitWave; (2) up to 7.6× higher inference efficiency than Bitlet on representative models; (3) 5.072mm2 area and scalable power consumption from 550.43mW (float32) to 495.12mW (16b) and 457.90mW (8b) @28nm TSMC; and (4) high versatility across precisions, and up to 18.9×13.5× higher than NVIDIA A100 and Jetson Orin 32GB, demonstrating significant competitiveness against GPUs.
Yanhuan Liu, Kunming Zhang, Yuqun Liu, Siao Wen, Lexin Wang, Tianyu Liu 0007, Zhihua Fan, Xiaochun Ye, Dongrui Fan, Xuejun An
ASPLOS (2)7
2026 HGNNMap: Heterogeneous Graph Neural Network-Based Mapping for Spatial Accelerators
Shengzhong Tang, Zhihua Fan, Tianyu Liu 0007, Xuejun An, Xiaochun Ye
Euro-Par (1)3
2025 Accelerating Authenticated Block Ciphers via RISC-V Custom Cryptography Instructions
abstract
As one of the standardized encryption algorithms, authenticated block ciphers based on Galois/Counter Mode (GCM) is a widely-used method to guarantee the accuracy and reliability in data transmission. The profiling work demonstrates that across the execution process of GCM mode, the authentication operation is the main performance bottleneck because it introduces operations in high-dimensional Galois field (GF), which could not be efficiently executed via existing ISA. To overcome this problem, we propose a custom ISA extension and cooperate it with RISC-V cryptography extension to accelerate the whole process of authenticated block ciphers. Besides, we design a specific crypto core including a fully-pipelined GF(2128) multiplier to support the extended instructions and integrate it into the multi-issue out-of-order core XT910 without introducing any clock frequency overhead. The proposed design significantly reduces the the number of instructions required in the main operations of authenticated block ciphers. We compare the performance of our designs to other existing acceleration scheme based on RISC-V ISA extension. Experimental result shows that our design outperforms other related work and achieves up to 17 × speedup with a lightweight hardware overhead.
Tianyu Liu 0007, Zhen Wang 0045, Zhihua Fan, Xiaochun Ye, Dongrui Fan
DATE3
2025 TSCNN: Compressing and Accelerating Sparse CNNs Using Sign-Reserved Toeplitz Filters
abstract
Exploiting the sparsity in convolutional neural networks (CNNs) is crucial to accelerate computing and reduce energy consumption. However, unstructured sparsity often introduces irregularity in convolutional operations, which complicates the control logic and undermines the benefits of sparsification. Structured sparsity alleviates these problems but sacrifices the adaptability to arbitrary sparse patterns. In this paper, we propose TSCNN, an algorithm-hardware co-design solution that aims to compress and accelerate sparse CNNs while balancing both adaptability to sparsity and computational efficiency. In terms of algorithm, TSCNN adopts pruned filters compressed with sign-reserved Toeplitz matrix format (Tfilters), which systematically enhances the regularity of data reuse and flexibly reduces network parameters by$44 \%-86 \%$while maintaining accuracy. In terms of hardware, TSCNN accelerator employs custom computing components to adapt to the structure of Tfilters and support the adaptive dataflow, further optimizing the computational efficiency. Experiments show that TSCNN outperforms a dense accelerator, SCNN and CSCNN, achieving$4.49 \times, 2.29 \times, 2.08 \times$and$1.29 \times$speedup and reducing energy consumption by$74.65 \%, 41.04 \%, 49.29 \%$and 43.66%, respectively.
Zhen Wang 0045, Tianyu Liu 0007, Zhihua Fan, Xiaochun Ye, Dongrui Fan
HPCC2
2025 DFU-E: A Dataflow Architecture for Edge DSP and AI Applications
abstract
Edge computing aims to enable swift, real-time data processing, analysis, and storage close to the data source. However, edge computing platforms are often constrained by limited processing power and efficiency. This paper presents DFU-E, a dataflow-based accelerator specifically designed to meet the demands of edge digital signal processing (DSP) and artificial intelligence (AI) applications. Our design addresses real-world requirements with three main innovations. First, to accommodate the diverse algorithms utilized at the edge, we propose a multi-layer dataflow mechanism capable of exploiting task-level, instruction block-level, instruction-level, and data-level parallelism. Second, we develop an edge dataflow architecture that includes a customized processing element (PE) array, memory, and on-chip network microarchitecture optimized for the multi-layer dataflow mechanism. Third, we design an edge dataflow software stack that enables automatic optimizations through operator fusion, dataflow graph mapping, and task scheduling. We utilize representative real-world DSP and AI applications for evaluation. Comparing with Nvidia's state-of-the-art edge computing processor, DFU-E achieves up to 1.42× geometric mean performance improvement and 1.27× energy efficiency improvement.
Zhihua Fan, Tianyu Liu 0007, Zhen Wang 0045, Meng Wu 0006, Kunming Zhang, Yanhuan Liu, Ninghui Sun, Xiaochun Ye, Dongrui Fan
IEEE Trans. Parallel Distributed Syst.3
2024 OBSD: On-The-Fly Block-Wise Sparse Distillation Accelerating SpGEMMs in DNN Applications
abstract
The mainstream deep neural networks (DNNs) widely adopt pruning techniques to alleviate network overfitting and computational complexity. Simultaneously, the weight and activation data exhibit increasingly noticeable block-wise sparsity. Current DNN-specific accelerators mainly focus on element-wise while neglecting sparse features of block granularity. Meanwhile, one of the state-of-the-art work, HIRAC, combining matrix tiling with the fast packing algorithm, SorPack, to achieve effective acceleration of Sparse General Matrix Multiplications (SpGEMMs). However, the SorPack algorithm executes on the host CPU, and its runtime lies on the critical path of the overall execution. In this work, a specific SpGEMM accelerator named OBSD is proposed which achieves high performance and energy efficiency. An on-the-fly block-wise distillation approach is proposed for leveraging the block-wise sparsity in both weight and activation data, which is implemented during the data loading process, consuming minimal additional time. Moreover, a data-flow architecture is devised to improve efficiency of data exchange among data processing elements (DPs) with investigating the relationship for different partition sizes, block-wise sparsity, and block-wise distance. The evaluation results demonstrate that: (1) OBSD achieves average of 1.89× speedup as compared to HIRAC for representative matrices in DNN workloads. (2) An end-to-end evaluation on a DNN model shows a 2.41× speedup and energy efficiency improvement of 1.99× over the HIRAC. (3) OBSD possesses a power consumption 12.9W with 56.6mm2area @28nm TSMC.
Yanhuan Liu, Kunming Zhang, Zhihua Fan, Lexin Wang, Tianyu Liu 0007, Zhen Wang 0045, Xiaochun Ye, Dongrui Fan, Xuejun An
HPCC7
2023 A High-accurate Multi-objective Exploration Framework for Design Space of CPU
abstract
To accelerate time-consuming multi-objective design space exploration of CPU, previous work trains prediction models using a set of performance metrics derived from few simulations, then predicts the rest. Unfortunately, the low accuracy of models limits the exploration effect, and how to achieve a good trade-off between multiple objectives is challenging.In this paper, we investigate various prediction models and find out the most accurate basic model. We enhance the model by ensemble learning to improve prediction accuracy. A hypervolume-improvement-based optimization method to trade off between multiple objectives is proposed together with a uniformity-aware selection algorithm to jump out of the local optimum. Experiments demonstrate that our open-source framework can reduce the distance to the Pareto optimal set by 76% and prediction error by 97% compared with the state-of-the-art work.
Mingyu Yan, Xin Liu 0073, Mo Zou, Tianyu Liu 0007, Xiaochun Ye, Dongrui Fan
DAC5
2023 DFGC: DFG-aware NoC Control based on Time Stamp Prediction for Dataflow Architecture
abstract
Coarse-grained reconfigurable architectures (CGRAs) have been regarded as promising accelerating paradigm for the ever-evolving algorithms in multi domains. Obtaining high energy-efficiency on CGRAs relies heavily on the combination of mapping, timing (issuing) and routing to decrease run-time idle. Statically configuration-driven designs are widely adopted, but the leak of hardware flexibility leads to a heavy reliance on burdensome scheduling of compiler to avoid over-serialization. This draw a trade-off between hardware/software co-design in CGRA scheduling.Unlike a static-schedule oriented approach, we propose DFGC (DFG-aware NoC Control), a dataflow-driven CGRA which takes the advantages of fully exploring data parallelism in different kernels by using dataflow dynamic firing with low overhead. The DFGC compiler is responsible for analyzing critical data paths and generating TimeStamp predictions instead of per cycle configuration. The rough predicted results enable router and PE to be sensitive to the entire dataflow graph, thereby accelerating the whole computation process. The DFG-aware NoC design realize a combined scheduling technique of hardware dynamic decision-making together with static prediction. DFGC represents a paradigm of CGRA that is worth exploring, achieving hardware/software co-design without relying on sophisticated designed compiler. Experiments show that DFGC achieves 1.32× energy efficiency improvement over a dataflow architecture and 1.8× energy efficiency improvement over a state-of-the-art static configured CGRA.
Tianyu Liu 0007, Zhihua Fan
ICCD1
2023 Alleviating Transfer Latency in DataFlow Accelerator for DSP Applications
abstract
Towards multiple domains, dataflow accelerators show superiority for their flexible programmability and high efficiency. This efficiency relies highly on data communication between processing elements (PEs), which is sensitive to PE location, array scale and workload size. Laying out instructions as a dataflow graph on the PE array creates more instruction-level parallelism. However, the farther distance between remote PEs and memory banks introduces extra transfer latency, bringing performance degradation to high real-time applications. This paper examines the workloads of digital signal processing across different data scales and classifies latency problems related to data transfers and kernel switching. Specifically, we propose a novel forwarding network on chip to alleviate transfer latency and improve multi-destination sharing in the dataflow execution. Moreover, we devise bandwidth reusing mechanism to speedup kernel switching. The experiment results show that our scalable design achieves up to 2.19× (1.45× on average) speedup while reducing switching overhead by 9.85×, with an area overhead of 10.82% over the conventional dataflow accelerator.
Zhihua Fan, Zhen Wang 0045, Tianyu Liu 0007, Junying Huang, Shengzhong Tang, Yanhuan Liu, Kunming Zhang, Xiaochun Ye, Dongrui Fan
ICCD5
2023 Accelerating Convolutional Neural Networks by Exploiting the Sparsity of Output Activation
abstract
Deep Convolutional Neural Networks (CNNs) are the most widely used family of machine learning methods that have had a transformative effect on a wide range of applications. Previous studies have made great breakthroughs in accelerating CNNs, but they only target on the input sparsity of activation and weight, thus do not eliminate the unnecessary computations due to the fact that more zeros in the output results are not directly caused by the zero-valued positions of the input data. In this paper, we take advantage of the output activation sparsity to reduce the execution time and energy consumption of CNNs. First, we propose an effective prediction method that leverages the output activation sparsity. Our method first predicts the output activation polarity of convolutional layers based on the singular value decomposition (SVD) approach. Then, it uses the predicted negative value to skip invalid computations. Second, an effective accelerator is designed to take advantage of sparsity to achieve CNN inference acceleration. Each PE is equipped with a prediction unit and a non-zero value detection unit to remove invalid computation blocks. And an instruction bypass technique is proposed which further exploits the sparsity of the weights. The efficient dataflow graph mapping approach and pipeline execution ensure high computational resource utilization. Experiments show that our approach achieves up to 1.63× speedup and 55.30% energy reduction compared with dense networks with a slight loss of accuracy. Compared with Eyeriss, our accelerator achieves on average 1.31 × performance improvement and 54% energy reduction. Our accelerator also achieves a similar performance to SnaPEA, but with a better energy efficiency.
Zhihua Fan, Zhen Wang 0045, Tianyu Liu 0007, Yanhuan Liu, Meng Wu 0006, Xinxin Wu, Xiaochun Ye, Dongrui Fan, Ninghui Sun, Xuejun An
IEEE Trans. Parallel Distributed Syst.4
2022 LRP: Predictive output activation based on SVD approach for CNN s acceleration
abstract
Convolutional Neural Networks (CNNs) achieve state-of-the-art performance in a wide range of applications. CNNs contain millions of parameters, and a large number of computations challenge hardware design. In this paper, we take advantage of the output activation sparsity of CNNs to reduce the execution time and energy consumption of the network. We propose Low Rank Prediction (LRP), an effective prediction method that leverages the output activation sparsity. LRP first predicts the output activation polarity of the convolutional layer based on the singular value decomposition (SVD) approach of the convolution kernel. And then it uses the predicted negative value to skip invalid computation in the original convolution. In addition, an effective accelerator, LRPPU, is proposed to take advantage of sparsity to achieve network inference acceleration. Experiments show that our LRPPU achieves 1.48 x speedup and 2.02 x energy reduction compared with dense networks with slight loss of accuracy. Also, it achieves on average 2.57 x speedup over Eyeriss and has similar performance and less accuracy loss compared with SnaPFA.
Xinxin Wu, Zhihua Fan, Tianyu Liu 0007, Xiaochun Ye, Dongrui Fan
DATE3
2022 A Routing-Aware Mapping Method for Dataflow Architectures
Zhihua Fan, Tianyu Liu 0007, Xuejun An, Xiaochun Ye, Dongrui Fan
NPC3