Yue Liang 0004

dblp:26/3264-4 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0005-7916-1812ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DynaX: Sparse Attention Acceleration with Dynamic X: M Fine-Grained Structured Pruning
Xiao Xiong, Zhaorui Chen, Yue Liang 0004, Minghao Tian, Jiaxing Shang, Dajiang Liu
ASPLOS (2)3
2025 CoSpMV: Towards Agile Software and Hardware Co-Design for SpMV Computation
abstract
Sparse Matrix-Vector multiplication (SpMV) is a widely used kernel in scientific or engineering applications and it is commonly implemented in FPGAs for acceleration. Existing works on FPGA usually pre-process the sparse matrix for data compression from the software perspective, and then design a unified architecture from the hardware perspective. However, as different SpMV kernels expose different levels of data parallelism after software processing, a unified architecture may not efficiently tap the underlying parallelism exposed in a specific kernel, leading to poor bandwidth utilization (BU) or poor resource utilization. To this end, this paper proposes an agile software and hardware co-design framework, CoSpMV, that employs design space exploration on both software and hardware for a specific kernel. Specifically, by providing a scalable compressed data format and a highly pipelined hardware template, CoSpMV can select the most suitable software and hardware configurations for different kernels and generate the accelerator instantly. The experimental results show that CoSpMV can achieve 3.91$\times$speedup on GFLOPs, and 1.31$\times$speedup on BU compared to the state-of-the-art work.
Minghao Tian, Yue Liang 0004, Dajiang Liu
IEEE Trans. Computers2
2024 DISC: Exploiting Data Parallelism of Non-Stencil Computations on CGRAs via Dynamic Iteration Scheduling
abstract
Memory partitioning is commonly used to enhance data parallelism of Coarse-grain reconfigurable arrays (CGRAs), typically targeting stencil computation with regular memory access patterns. However, many important workloads, such as linear algebra and signal processing, include non-stencil computation with irregular memory access patterns where memory partitioning is not feasible, leading to memory access conflicts and poor data parallelism. In this paper, we propose a Dynamic Iteration Scheduling CGRA (DISC) that can dynamically exploit data parallelism from non-stencil computations. Via dynamic scheduling on loop iterations, DISC can select conflict-free iterations from an iteration buffer for parallel data access, while holding data dependence. To further enhance the ability to find conflict-free iterations, dynamic data reuse is also introduced to reduce the number of memory references. The experimental results show that DISC can achieve 1.41× performance and 2.75× energy efficiency while consuming much less area and power overhead, as compared to dynamic-scheduling CGRA which supports dynamic operator scheduling.
Yue Liang 0004, Di Mou, Dajiang Liu
ICCAD1