Qingcheng Xiao

dblp:202/5867 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
3since 2021 · last 2022
0000-0001-6230-5342ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2022 Towards Agile DNN Accelerator Design Using Incremental Synthesis on FPGAs
abstract
Hardware-software co-design is the new trend for deep neural network and FPGA accelerator development, which iteratively revises and tunes the full system. The bottleneck of the approach lies in the time-consuming hardware synthesis. In this paper, we propose an incremental synthesis framework Acoda to rapidly design DNN accelerators on FPGAs. Based on the observation that most revisions to DNNs are minor and local, Acoda reuses existing hardware modules and incrementally modifies the accelerator. It first detects the software revisions using a graph edit distance algorithm. Then, it maps the software revisions to hardware revisions through a multi-level reuse hierarchy. As a result, Acoda speeds up the design process by 9.31X to 34.17X and achieves performance results comparable to off-the-shelf accelerators.
Qingcheng Xiao, Yun Liang 0001
FPGA1
2022 FCNNLib: A Flexible Convolution Algorithm Library for Deep Learning on FPGAs
abstract
Convolution features huge complexity and demands high computation capability. Among hardware platforms, field programmable gate array (FPGA) emerges as a promising solution for its substantial available parallelism and energy efficiency. Besides, convolution can be implemented with different algorithms, including conventional, general matrix–matrix multiplication (GEMM), Winograd, and fast Fourier transformation (FFT) algorithms, which are diverse in arithmetic complexity, resource requirement, etc. Different convolutional neural network (CNN) models have different topologies and structures, favoring different convolution algorithms. In response, software libraries such as cuDNN provide a variety of computational primitives to support these algorithms. However, supporting such libraries on FPGAs is challenging. First, multiple algorithms can share the FPGA resources spatially as well as temporally, introducing either reconfiguration overhead or resource underutilization. Second, FPGA implementation remains a significant challenge for library developers. It typically requires significant specialized hardware knowledge. In this article, we proposeFCNNLib, an efficient and scalable convolution algorithm library on FPGAs. To coordinate multiple convolution algorithms on FPGAs, we develop three schedulings: 1) spatial; 2) temporal; and 3) hybrid, which exhibit different tradeoffs in latency and throughput. We explore these schedulings by balancing the reconfiguration overhead, resource utilization, and optimization objectives of the CNNs. Then, we provide efficient and tunable algorithm templates that allow performance tuning through performance and resource models. To arm the users,FCNNLibexposes a set of interfaces to support high-level application designs. We demonstrate the usability ofFCNNLibwith state-of-the-art CNNs.FCNNLibachieves up to$44.6\times $and$1.76\times $energy efficiency in various scenarios compared with software libraries for CPUs and GPUs, respectively.
Yun Liang 0001, Qingcheng Xiao, Liqiang Lu, Jiaming Xie
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 HASCO: Towards Agile HArdware and Software CO-design for Tensor Computation
abstract
Tensor computations overwhelm traditional general-purpose computing devices due to the large amounts of data and operations of the computations. They call for a holistic solution composed of both hardware acceleration and software mapping. Hardware/software (HW/SW) co-design optimizes the hardware and software in concert and produces high-quality solutions. There are two main challenges in the co-design flow. First, multiple methods exist to partition tensor computation and have different impacts on performance and energy efficiency. Besides, the hardware part must be implemented by the intrinsic functions of spatial accelerators. It is hard for programmers to identify and analyze the partitioning methods manually. Second, the overall design space composed of HW/SW partitioning, hardware optimization, and software optimization is huge. The design space needs to be efficiently explored. To this end, we propose an agile co-design approach HASCO that provides an efficient HW/SW solution to dense tensor computation. We use tensor syntax trees as the unified IR, based on which we develop a two-step approach to identify partitioning methods. For each method, HASCO explores the hardware and software design spaces. We propose different algorithms for the explorations, as they have distinct objectives and evaluation costs. Concretely, we develop a multi-objective Bayesian optimization algorithm to explore hardware optimization. For software optimization, we use heuristic and Q-learning algorithms. Experiments demonstrate that HASCO achieves a 1.25X to 1.44X latency reduction through HW/SW co-design compared with developing the hardware and software separately.
Qingcheng Xiao, Size Zheng 0001, Bingzhe Wu, Pengcheng Xu 0005, Xuehai Qian, Yun Liang 0001
ISCA1
2020 FCNNLib: An Efficient and Flexible Convolution Algorithm Library on FPGAs
abstract
Convolutions can be implemented with different algorithms, which are diverse in arithmetic complexity, resource requirement, etc. Multiple algorithms can share the FPGA resources spatially as well as temporally, introducing either reconfiguration overhead or resource underutilization. In this paper, we propose an efficient library FCNNLib to coordinate multiple convolution algorithms on FPGAs. We develop three scheduling techniques: spatial, temporal, and hybrid, which exhibit different trade-offs in latency and throughput. We also expose a set of interfaces to arm the users. Experiments using modern CNNs demonstrate FCNNLib achieves up to 1.315X latency improvement compared with dedicated accelerators and 1.755X energy efficiency improvement compared with cuDNN.
Qingcheng Xiao, Liqiang Lu, Jiaming Xie, Yun Liang 0001
DAC1
2020 Evaluating Fast Algorithms for Convolutional Neural Networks on FPGAs
abstract
In recent years, convolutional neural networks (CNNs) have become widely adopted for computer vision tasks. Field-programmable gate arrays (FPGAs) have been adequately explored as a promising hardware accelerator for CNNs due to its high performance, energy efficiency, and reconfigurability. However, prior FPGA solutions based on the conventional convolutional algorithm is often bounded by the computational capability of FPGAs (e.g., the number of DSPs). To address this problem, the feature maps are transformed to a special domain using fast algorithms to reduce the arithmetic complexity. Winograd and fast Fourier transformation (FFT), as fast algorithm representatives, first transform input data and filter to Winograd or frequency domain, then perform element-wise multiplication, and apply inverse transformation to get the final output. In this paper, we propose a novel architecture for implementing fast algorithms on FPGAs. Our design employs line buffer structure to effectively reuse the feature map data among different tiles. We also effectively pipeline the Winograd/FFT processing element (PE) engine and initiate multiple PEs through parallelization. Meanwhile, there exists a complex design space to explore. We propose an analytical model to predict the resource usage and the performance. Then, we use the model to guide a fast design space exploration. Experiments using the state-of-the-art CNNs demonstrate the best performance and energy efficiency on FPGAs. We achieve 854.6 and 2479.6 GOP/s for AlexNet and VGG16 on Xilinx ZCU102 platform using Winograd. We achieve 130.4 GOP/s for Resnet using Winograd and 201.1 GOP/s for YOLO using FFT on Xilinx ZC706 platform.
Yun Liang 0001, Liqiang Lu, Qingcheng Xiao, Shengen Yan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 Zac: Towards Automatic Optimization and Deployment of Quantized Deep Neural Networks on Embedded Devices
abstract
With the development toward commercial and civil use, the need for deploying Deep neural network (DNN) models on resource-constrained embedded devices is growing. Quantization has a para-mount impact on the performance, storage, and energy efficiency. However, to fully realize these benefits, programmers need to manually utilize low precision operations while maintaining accuracy, which is very challenging. Hence, we present a framework Zac to automatically optimize and deploy quantized DNN models on embedded devices. In order to do this, Zac performs necessary data type conversion and chooses proper data types for the intermediate data. Then it automatically customizes the operations according to the chosen types. Experiments demonstrate that by utilizing quantized models, Zac offers up to 19.18X and 25.44X improvement for throughput and energy efficiency compared with full precision designs, respectively. The automatically generated hardware designs from Zac achieve comparable performance to the highly optimized state-of-the-art accelerators which are designed manually.
Qingcheng Xiao, Yun Liang 0001
ICCAD1
2017 Exploring Heterogeneous Algorithms for Accelerating Deep Convolutional Neural Networks on FPGAs
abstract
Convolutional neural network (CNN) finds applications in a variety of computer vision applications ranging from object recognition and detection to scene understanding owing to its exceptional accuracy. There exist different algorithms for CNNs computation. In this paper, we explore conventional convolution algorithm with a faster algorithm using Winograd's minimal filtering theory for efficient FPGA implementation. Distinct from the conventional convolution algorithm, Winograd algorithm uses less computing resources but puts more pressure on the memory bandwidth. We first propose a fusion architecture that can fuse multiple layers naturally in CNNs, reusing the intermediate data. Based on this fusion architecture, we explore heterogeneous algorithms to maximize the throughput of a CNN. We design an optimal algorithm to determine the fusion and algorithm strategy for each layer. We also develop an automated toolchain to ease the mapping from Caffe model to FPGA bitstream using Vivado HLS. Experiments using widely used VGG and AlexNet demonstrate that our design achieves up to 1.99X performance speedup compared to the prior fusion-based FPGA accelerator for CNNs.
Qingcheng Xiao, Yun Liang 0001, Liqiang Lu, Shengen Yan, Yu-Wing Tai
DAC1
2017 Evaluating Fast Algorithms for Convolutional Neural Networks on FPGAs
abstract
In recent years, Convolutional Neural Networks (CNNs) have become widely adopted for computer vision tasks. FPGAs have been adequately explored as a promising hardware accelerator for CNNs due to its high performance, energy efficiency, and reconfigurability. However, prior FPGA solutions based on the conventional convolutional algorithm is often bounded by the computational capability of FPGAs (e.g., the number of DSPs). In this paper, we demonstrate that fast Winograd algorithm can dramatically reduce the arithmetic complexity, and improve the performance of CNNs on FPGAs. We first propose a novel architecture for implementing Winograd algorithm on FPGAs. Our design employs line buffer structure to effectively reuse the feature map data among different tiles. We also effectively pipeline the Winograd PE engine and initiate multiple PEs through parallelization. Meanwhile, there exists a complex design space to explore. We propose an analytical model to predict the resource usage and reason about the performance. Then, we use the model to guide a fast design space exploration. Experiments using the state-of-the-art CNNs demonstrate the best performance and energy efficiency on FPGAs. We achieve an average 1006.4 GOP/s for the convolutional layers and 854.6 GOP/s for the overall AlexNet and an average 3044.7 GOP/s for the convolutional layers and 2940.7 GOP/s for the overall VGG16 on Xilinx ZCU102 platform.
Liqiang Lu, Yun Liang 0001, Qingcheng Xiao, Shengen Yan
FCCM3
2017 Enabling high performance deep learning networks on embedded systems
abstract
Deep learning is nowadays one of the most popular research topics in computer science. In recent years, the extensive application of convolutional neural network has made it become a new direction for the computer architecture research that is developing rapidly. Currently, there is a growing demand on off-line deploying deep learning network on top of embedded mobile systems. However, how to balance the limited computing and storage resources on embedded platforms, and the huge storage requirements with the increase of network complexity, has become the core problem of current research. In this paper, we explore the optimization technology to enable high-performance deep learning network for embedded systems from two aspects: the neural network design and the acceleration on embedded platforms. We focus on convolutional neural networks. First, we combine several technologies and propose a set of pruning mechanisms to save storage resources. We also explore the concept of block-wise sparsity. Second, from the perspective of mobile deployment, we propose a method to automatically select the optimal convolution/matrix multiplication approach based on the sparsity of the matrix and its sparse structure. Our experiments on NVIDIA TX1 show that our approach can be used together to promote each other and achieve the goal of improving the performance of computation while reducing the storage consumption.
Qian Li 0027, Qingcheng Xiao, Yun Liang 0001
IECON2