Yu Liu 0007

dblp:97/2274-7 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-1170-3858ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 A 10.60 μW 150 GOPS Mixed-Bit-Width Sparse CNN Accelerator for Life-Threatening Ventricular Arrhythmia Detection
abstract
This paper proposes an ultra-low power, mixed-bit-width sparse convolutional neural network (CNN) accelerator to accelerate ventricular arrhythmia (VA) detection. The chip achieves 50% sparsity in a quantized 1D CNN using a sparse processing element (SPE) architecture. Measurement on the prototype chip TSMC 40nm CMOS low-power (LP) process for the VA classification task demonstrates that it consumes 10.60 μW of power while achieving a performance of 150 GOPS and a diagnostic accuracy of 99.95%. The computation power density is only 0.57 μW/mm2, which is 14.23× smaller than state-of-the-art works, making it highly suitable for implantable and wearable medical devices.
Zhenge Jia, Zheyu Yan, Jay Mok, Manto Yung, Yu Liu 0007, Wujie Wen, Luhong Liang, Kwang-Ting Cheng, Xiaobo Sharon Hu, Yiyu Shi 0001
ASP-DAC6
2025 APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
abstract
DNN accelerators, significantly advanced by model compression and specialized dataflow techniques, have marked considerable progress. However, the frequent access of highprecision partial sums (PSUMs) leads to excessive memory demands in architectures utilizing input/weight stationary dataflows. Traditional compression strategies have typically overlooked PSUM quantization, which may account for 69% of power consumption. This study introduces a novel Additive Partial Sum Quantization (APSQ) method, seamlessly integrating PSUM accumulation into the quantization framework. A grouping strategy that combines APSQ with PSUM quantization enhanced by a reconfigurable architecture is further proposed. The APSQ performs nearly lossless on NLP and CV tasks across BERT, Segformer, and EfficientViT models while compressing PSUMs to INT8. This leads to a notable reduction in energy costs by $\mathbf{2 8-8 7 \%}$. Extended experiments on LLaMA2-7B demonstrate the potential of APSQ for large language models. Code is available at https://github.com/Yonghao-Tan/APSQ.
Yonghao Tan, Pingcheng Dong, Yongkun Wu, Yu Liu 0007, Shih-Yang Liu, Xijie Huang, Luhong Liang, Kwang-Ting Cheng
DAC4
2025 CoXplorer: Multi-Staged Co-Exploration Framework for AI Model Compression and Accelerator Design
abstract
The rapid evolution of artificial intelligence (AI) algorithms demands efficient computing chips, positioning algorithm-hardware co-design as a crucial optimization strategy. However, automating the co-design process remains challenging due to the lack of a unified exploration framework for both algorithmic and hardware domains, as existing tools - hardware design space exploration (DSE) and compression neural architecture search (Compression NAS) - operate independently, relying entirely on manual collaboration. This paper presents CoXplorer, a co-exploration framework that connects model-compression optimization space and architecture design space. We make three key contributions: (1) a multi-staged co-design space decomposition method that enables systematic exploration of compression-hardware design choices with reduced complexity, (2) an AC-Copilot toolchain enhanced with multi-grained performance modeling driven by hardware simulation-compilation hierarchical cooperation to fulfill various evaluation requirements of co-exploration, enabling balanced simulation accuracy-efficiency trade-offs, and (3) a co-exploration workflow with hierarchical and bottleneck-guided search for harmonizing optimization objectives of both model and hardware design spaces, resulting in improved search efficiency. We validate the CoXplorer on two edge chips, which achieve 53.7% throughput and 45.8% energy efficiency improvements for the CNN acceleration, and 7.5× speedup with 9.9× energy efficiency boost for the Transformer acceleration. A case study on large language model acceleration shows CoXplorer’s extensibility to emerging workloads, enhancing LLAMA2-7B inference throughput from 6.75 to 25.46 tokens/s via co-optimization with compression and near-memory computing architecture.
Songchen Ma, Yonghao Tan, Pingcheng Dong, Di Pang, Yu Liu 0007, Luhong Liang, Kwang-Ting Cheng, Fengbin Tu
ICCAD7
2024 Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers
abstract
Non-linear functions are prevalent in Transformers and their lightweight variants, incurring substantial and frequently underestimated hardware costs. Previous state-of-the-art works optimize these operations by piece-wise linear approximation and store the parameters in look-up tables (LUT), but most of them require unfriendly high-precision arithmetics such as FP/INT 32 and lack consideration of integer-only INT quantization. This paper proposed a genetic LUT-Approximation algorithm namely GQA-LUT that can automatically determine the parameters with quantization awareness. The results demonstrate that GQA-LUT achieves negligible degradation on the challenging semantic segmentation task for both vanilla and linear Transformer models. Besides, proposed GQA-LUT enables the employment of INT8-based LUT-Approximation that achieves an area savings of 81.3~81.7% and a power reduction of 79.3~80.2% compared to the high-precision FP/INT 32 alternatives. Code is available at https://github.com/PingchengDong/GQA-LUT.
Pingcheng Dong, Yonghao Tan, Tianwei Ni, Yu Liu 0007, Luhong Liang, Shih-Yang Liu, Xijie Huang, Huaiyu Zhu 0004, Fengwei An, Kwang-Ting Cheng
DAC6
2022 ReAAP: A Reconfigurable and Algorithm-Oriented Array Processor With Compiler-Architecture Co-Design
abstract
Parallelism and data reuse are the most critical issues for the design of hardware acceleration in a deep learning processor. Besides, abundant on-chip memories and precise data management are intrinsic design requirements because most of deep learning algorithms are data-driven and memory-bound. In this paper, we propose a compiler-architecture co-design scheme targeting a reconfigurable and algorithm-oriented array processor, named ReAAP. Given specific deep neural networks, the proposed co-design scheme is effective to perform parallelism and data reuse optimization on compute-intensive layers for guiding reconfigurable computing in hardware. Especially, the systemic optimization is performed in our proposed domain-specific compiler to deal with the intrinsic tensions between parallelism and data locality, for the purpose of automatically mapping diverse layer-level workloads onto our proposed reconfigurable array architecture. In this architecture, abundant on-chip memories are software-controlled and its massive data access is precisely handled by compiler-generated instructions. In our experiments, the ReAAP is implemented on an embedded FPGA platform. Experimental results demonstrate that our proposed co-design scheme is effective to integrate software flexibility with hardware parallelism for accelerating diverse deep learning workloads. As a whole system, ReAAP achieves a consistently high utilization of hardware resource for accelerating all the diverse compute-intensive layers in ResNet, MobileNet, and BERT.
Jianwei Zheng 0002, Yu Liu 0007, Luhong Liang, Deming Chen, Kwang-Ting Cheng
IEEE Trans. Computers2
2020 Optimize FPGA-Based Neural Network Accelerator with Bit-Shift Quantization
abstract
Well-programmed Field Programmable Gate Arrays (FPGAs) can accelerate Deep Neural Network (DNN) with high power efficiency. The dominant workloads of DNNs are Multiply Accumulates (MACs), which can be directly mapped to Digital Signal Processors (DSPs) in the FPGA. A DNN accelerator pursuing high performance can consume almost all the DSPs, but with a considerable amount of Look-up Tables (LUTs) in the FPGA unused or performing MACs inefficiently. To solve this problem, we present a Bit-Shift method for FPGA-based DNN accelerator to fully utilize the resources in the FPGA. The MAC is converted to a limited number of shift-and-add operations, which can be implemented by LUTs with significant improvement of efficiency. A quantization method based on Minimum Mean Absolute Error (MMAE) is proposed to preserve the accuracy of the DNN inference in the conversion of DNN parameters without re-training. The quantized parameters can be compressed to a fixed and fewer number of bits to reduce the memory bandwidth. Accordingly, a Bit-Shift architecture is designed to load the compressed parameters and perform the converted MAC calculations without extra decompression module. A large scale DNN accelerator with the proposed Bit-Shift architecture is implemented in a Xilinx VU095 FPGA. Experimental results show that the proposed method can boost the processing speed by 32% and reach 331 GOPS at 190MHz clock frequency for ResNet-34.
Yu Liu 0007, Luhong Liang
ISCAS1
2017 A Dynamic-Bayesian-Network-Based Fault Diagnosis Methodology Considering Transient and Intermittent Faults
abstract
Transient fault (TF) and intermittent fault (IF) of complex electronic systems are difficult to diagnose. As the performance of electronic products degrades over time, the results of fault diagnosis could be different at different times for the given identical fault symptoms. A dynamic Bayesian network (DBN)-based fault diagnosis methodology in the presence of TF and IF for electronic systems is proposed. DBNs are used to model the dynamic degradation process of electronic products, and Markov chains are used to model the transition relationships of four states, i.e., no fault, TF, IF, and permanent fault. Our fault diagnosis methodology can identify the faulty components and distinguish the fault types. Four fault diagnosis cases of the Genius modular redundancy control system are investigated to demonstrate the application of this methodology.
Baoping Cai, Yu Liu 0007, Min Xie 0001
IEEE Trans Autom. Sci. Eng.2
2016 Modeling and analysis of reliability of multi-release open source software incorporating both fault detection and correction processes
Yu Liu 0007, Min Xie 0001
J. Syst. Softw.2
2015 A New Framework and Application of Software Reliability Estimation Based on Fault Detection and Correction Processes
abstract
Software reliability growth modeling plays an important role in software reliability evaluation. To incorporate more information and provide more accurate analysis, modeling software fault detection and correction processes has attracted widespread research attention recently. However, the assumption of the stochastic fault correction time delay brings more difficulties in modeling and estimating the parameters. In practice, other than the grouped fault data, software test records often include some more detailed information, such as the rough time when one fault is detected or corrected. Such semi-grouped dataset contains more information about fault removal processes than commonly used grouped dataset. Using the semi-grouped datasets can improve the accuracy of time delayed models. In this paper, a fault removal modelling framework for software reliability with semi-grouped data is studied and extended into multi-released software. Also, the corresponding parameter estimation is carried out with Maximum Likelihood estimation method. One test dataset with three releases from a practical software project is applied with the proposed framework, which shows satisfactory performance with the results.
Yu Liu 0007, Min Xie 0001
QRS1