VLDB 2026 Research / reviewers in the wild / expert
Guangdong Liu
dblp:53/2009
· DBLP profile ↗
4ranked-venue papers
1as first author
2since 2021 · last 2024
0009-0004-5515-4741ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 100% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator |
0.4 | 1 | 2020 | LPAC: A Low-Precision Accelerator for CNN on FPGAs · FPGA 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
FPGA-based CNN accelerator |
0.4 | 1 | 2020 | LPAC: A Low-Precision Accelerator for CNN on FPGAs · FPGA 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
low-bit quantization accelerator |
0.4 | 1 | 2020 | LPAC: A Low-Precision Accelerator for CNN on FPGAs · FPGA 2020 |
Machine learning › Efficient and distributed learning
model compression |
0.1 | 1 | 2020 | LPAC: A Low-Precision Accelerator for CNN on FPGAs · FPGA 2020 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.1 | 1 | 2020 | LPAC: A Low-Precision Accelerator for CNN on FPGAs · FPGA 2020 |
Methods — techniques the papers use, named apart from their topics
DSP mapping · 0.94a4w quantization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | XVDPU: A High-Performance CNN Accelerator on the Versal Platform Powered by the AI EngineabstractToday, convolutional neural networks (CNNs) are widely used in computer vision applications. However, the trends of higher accuracy and higher resolution generate larger networks. The requirements of computation or I/O are the key bottlenecks. In this article, we propose XVDPU: the AI Engine (AIE)-based CNN accelerator on Versal chips to meet heavy computation requirements. To resolve the IO bottleneck, we adopt several techniques to improve data reuse and reduce I/O requirements. An arithmetic logic unit is further proposed that can better balance resource utilization, new feature support, and efficiency of the whole system. We have successfully deployed more than 100 CNN models with our accelerator. Our experimental results show that the 96-AIE-core implementation can achieve 1,653 frames per second (FPS) for ResNet50 on VCK190, which is 9.8× faster than the design on ZCU102 running at 168.5 FPS. The 256-AIE-core implementation can further achieve 4,050 FPS. We propose a tilling strategy to achieve feature-map-stationary for high-definition CNN with the accelerator, achieving 3.8× FPS improvement on the residual channel attention network and 3.1× on super-efficient super-resolution. This accelerator can also solve the 3D convolution task in disparity estimation, achieving end-to-end performance of 10.1 FPS with all the optimizations. Xijie Jia, Guangdong Liu, Xinlin Yang, Zhuohuan Liu, Mengke Liu, Xiaoyang Yan, Rongzhang Zheng, Dong Li 0025, Satyaprakash Pareek, Jian Weng 0012, Dongliang Xie |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2022 | XVDPU: A High Performance CNN Accelerator on the Versal Platform Powered by the AI EngineabstractThe convolution neural networks (CNNs) are widely used in computer vision applications nowadays. However, the trends of higher accuracy and higher resolution generate larger networks, indicating that computation and I/O bandwidth are key bottlenecks to reach performance. The Xilinx's latest 7nm Versal ACAP platform with AI-Engine (AIE) cores can deliver up-to 8x silicon compute density at 50% the power consumption compared with the traditional FPGA solutions. In this paper, we propose XVDPU: the AIE-based int8-precision CNN accelerator on Versal chips, scaling from 16-AIE-core (C16B1) to 320-AIE-core (C64B5, Peak:109.2 TOPs) to meet computation requirements. To resolve IO bottleneck, we adopt several techniques such as multi-batch (MB), shared-weights (SHRWGT), feature-map-stationary (FMS) and long-load-weights (LLW) to improve data-reuse and reduce I/O requirements. An Arithmetic Logic Unit (ALU) design is further proposed into the accelerator which mainly performs non-convolution layers such as Depthwise-Conv layer, Pooling layer and Non-linear function layers using the same logic resources, which can better balance resource utilization, new feature support and efficiency of the whole system. We have successfully deployed more than 100 CNN models with our accelerator. Our experimental results show that the 96-AIE-core (C32B3, Peak: 32.76 TOPs) implementation can achieve 1653 FPS for ResNet50 on VCK190, which is 9.8x faster than the design on ZCU102 running at 168.5 FPS with peak 3.6 TOPs. The 256-AIE-core (C32B8, Peak: 87.36 TOPs) implementation can further achieve 4050 FPS which better leverages the computing power of Versal AIE devices. The powerful XVDPU will help enable many applications on the embedded system, such as low-latency data center, high level ADAS and complex robotics. Xijie Jia, Guangdong Liu, Xinlin Yang, Rongzhang Zheng, Satyaprakash Pareek, Dongliang Xie |
FPL | 3 |
| 2020 | LPAC: A Low-Precision Accelerator for CNN on FPGAsabstractLow bit quantization of neural network is required on edge devices to achieve lower power consumption and higher performance. 8bit or binary network either consumes a lot of resources or has accuracy degradation. Thus, a full-process hardware-friendly quantization solution of 4A4W (activations 4bit and weights 4bit) is proposed to achieve better accuracy/resource trade-off. It doesn't contain any additional floating operations and achieve accuracy comparable to full-precision. We also implement a low-precision accelerator for CNN (LPAC) on the Xilinx FPGA, which takes full advantage of its DSP by efficiently mapping convolutional computations. Through on-chip reassign management and resource-saving analysis, high performance can be achieved on small chips. Our 4A4W solution achieves 1.8x higher performance than 8A8W and 2.42x increase in power efficiency under the same resource. On ImageNet classification, the accuracy has a gap less than 1% to full-precision in Top-5. On the human pose estimation, we achieve 261 frames per second on ZU2EG, which is 1.78x speed up compared to 8A8W and the accuracy has only 1.62% gap to full-precision. This proves that our solution has better universality. Tiantian Han, Xijie Jia, Guangdong Liu, Pingbo An, Yingran Tan, Lingzhi Sui, Shaoxia Fang, Dongliang Xie, Michaela Blott |
FPGA | 6 |
| 2014 | Partitioned multiprocessor scheduling of mixed-criticality parallel jobsabstractMotivated by the increasing trend in embedded systems towards platform integration, there has been an increasing research interest in scheduling mixed-criticality systems. However, most existing efforts have concentrated on scheduling sequential tasks and ignored intra-task parallelism. In this paper, we study the scheduling of mixed-criticality parallel jobs on multiprocessor platforms. We propose a synchronous mixed-criticality job model, where each job consists of segments, each segment having an arbitrary number of parallel threads that synchronize at the end of the segment. A novel MinLoad algorithm is developed to decompose mixed-criticality parallel jobs into mixed-criticality sequential jobs. This decomposition enables us to leverage existing mixed-criticality scheduling algorithms and schedulability analysis to the multiprocessor scheduling of mixed-criticality parallel jobs. In addition, our MinLoad job decomposition algorithm is designed to make the decomposed mixed-criticality sequential tasks easier to schedule, and thus requires smaller-sized multiprocessor platforms for the mixed-criticality systems. Guangdong Liu, Ying Lu 0002, Shige Wang, Zonghua Gu 0001 |
RTCSA | 1 |