VLDB 2026 Research / reviewers in the wild / expert
Lingzhi Sui
dblp:157/9925
· DBLP profile ↗
7ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Hardware accelerators and domain-specific architectures · 82% Reconfigurable computing and FPGAs · 18% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% | |
| Artificial intelligence
2 papers |
Efficient and distributed learning · 70% Deep learning architectures and training · 30% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
FPGA-based CNN accelerator |
1.1 | 3 | 2020 | LPAC: A Low-Precision Accelerator for CNN on FPGAs · FPGA 2020 DNNVM: End-to-End Compiler Leveraging Operation Fusion on FPGA-based CNN Accelerators · FPGA 2019 Angel-Eye: A Complete Design Flow for Mapping CNN Onto Embedded FPGA · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator |
0.9 | 2 | 2020 | DNNVM: End-to-End Compiler Leveraging Heterogeneous Optimizations on FPGA-Based CNN Accelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 LPAC: A Low-Precision Accelerator for CNN on FPGAs · FPGA 2020 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.8 | 2 | 2020 | DNNVM: End-to-End Compiler Leveraging Heterogeneous Optimizations on FPGA-Based CNN Accelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 Angel-Eye: A Complete Design Flow for Mapping CNN Onto Embedded FPGA · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.7 | 2 | 2019 | DNNVM: End-to-End Compiler Leveraging Operation Fusion on FPGA-based CNN Accelerators · FPGA 2019 Angel-Eye: A Complete Design Flow for Mapping CNN Onto Embedded FPGA · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Compilers and program optimization
deep learning compiler |
0.4 | 1 | 2020 | DNNVM: End-to-End Compiler Leveraging Heterogeneous Optimizations on FPGA-Based CNN Accelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
low-bit quantization accelerator |
0.4 | 1 | 2020 | LPAC: A Low-Precision Accelerator for CNN on FPGAs · FPGA 2020 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2020 | DNNVM: End-to-End Compiler Leveraging Heterogeneous Optimizations on FPGA-Based CNN Accelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Compilers and program optimization › deep learning compiler
operator fusion |
0.4 | 1 | 2019 | DNNVM: End-to-End Compiler Leveraging Operation Fusion on FPGA-based CNN Accelerators · FPGA 2019 |
Machine learning › Efficient and distributed learning
model compression |
0.1 | 1 | 2020 | LPAC: A Low-Precision Accelerator for CNN on FPGAs · FPGA 2020 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.1 | 1 | 2020 | LPAC: A Low-Precision Accelerator for CNN on FPGAs · FPGA 2020 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2019 | DNNVM: End-to-End Compiler Leveraging Operation Fusion on FPGA-based CNN Accelerators · FPGA 2019 |
Methods — techniques the papers use, named apart from their topics
subgraph isomorphism · 2.0heuristic shortest-path search · 1.1loop optimization · 0.9graph optimization · 0.9data layout optimization · 0.9DSP mapping · 0.94a4w quantization · 0.9data quantization · 0.3compilation tool · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | LPAC: A Low-Precision Accelerator for CNN on FPGAsabstractLow bit quantization of neural network is required on edge devices to achieve lower power consumption and higher performance. 8bit or binary network either consumes a lot of resources or has accuracy degradation. Thus, a full-process hardware-friendly quantization solution of 4A4W (activations 4bit and weights 4bit) is proposed to achieve better accuracy/resource trade-off. It doesn't contain any additional floating operations and achieve accuracy comparable to full-precision. We also implement a low-precision accelerator for CNN (LPAC) on the Xilinx FPGA, which takes full advantage of its DSP by efficiently mapping convolutional computations. Through on-chip reassign management and resource-saving analysis, high performance can be achieved on small chips. Our 4A4W solution achieves 1.8x higher performance than 8A8W and 2.42x increase in power efficiency under the same resource. On ImageNet classification, the accuracy has a gap less than 1% to full-precision in Top-5. On the human pose estimation, we achieve 261 frames per second on ZU2EG, which is 1.78x speed up compared to 8A8W and the accuracy has only 1.62% gap to full-precision. This proves that our solution has better universality. Tiantian Han, Xijie Jia, Guangdong Liu, Pingbo An, Yingran Tan, Lingzhi Sui, Shaoxia Fang, Dongliang Xie, Michaela Blott |
FPGA | 9 |
| 2020 | DNNVM: End-to-End Compiler Leveraging Heterogeneous Optimizations on FPGA-Based CNN Acceleratorsabstractstate-of-the-art method for several artificial intelligence domains in recent years. The increasingly complex CNN models are both computation-bound and I/O-bound. Fieldprogrammable gate array-based accelerators driven by custom instruction set architecture (ISA) achieve a balance between generality and efficiency, but there is much on them left to be optimized. We propose the full-stack compiler deep neural network virtual machine (DNNVM), which is an integration of optimizers for graphs, loops and data layouts, an assembler, a runtime supporter, and a validation environment. The DNNVM works in the context of deep learning frameworks and transforms CNN models into the directed acyclic graph: XGraph. Based on XGraph, we transform the optimization challenges for both data layout and pipeline into graph-level problems. DNNVM enumerates all potentially profitable fusion opportunities by a heuristic subgraph isomorphism algorithm to leverage pipeline and data layout optimizations, and searches for the best choice of execution strategies of the whole computing graph. On the Xilinx ZU2@330 MHz and ZU9@330 MHz, we achieve equivalently state-of-the-art performance on our benchmarks by naïve implementations without optimizations, and the throughput is further improved up to 1.26× by leveraging heterogeneous optimizations in DNNVM. Finally, with ZU9@330 MHz, we achieve state-of-the-art performance for VGG and ResNet50. We achieve a throughput of 2.82 TOPs/s and an energy efficiency of 123.7 GOPs/s/W for VGG. Additionally, we achieve 1.38 TOPs/s for ResNet50 and 1.41 TOPs/s for GoogleNet. Shuang Liang 0010, Lingzhi Sui, Xijie Jia, Jiantao Qiu, Yushun Wang, Yu Wang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | DNNVM: End-to-End Compiler Leveraging Operation Fusion on FPGA-based CNN AcceleratorsabstractIn recent years, Convolutional Neural Network(CNN) is becoming the state-of-the-art method in a wide range of Artificial Intelligence(AI) domains. The increasingly large and complex CNN models are both computation bound and I/O bound. FPGA-based accelerators driven by custom Instruction Set Architecture(ISA) achieve a balance between generality and efficiency, and leave much room for optimization. Operation fusion which fuses adjacent operations without saving intermediate results back to off-chip DDR can greatly alleviate bandwidth pressure, operations can be executed by different computation engines concurrently for latency hiding. To leverage optimizations, especially operation fusion on custom instruction-based accelerators, we propose a full-stack compiler DNNVM(Deep Neural Network Virtual Machine). DNNVM is an integration of optimizers for framework-independent computing graph, loops and data layouts, an assembler, a runtime supporter and a validation environment. DNNVM works in the context of deep learning frameworks and transforms CNN models into a directed acyclic graph, XGraph. After analyzing the interaction among fusion depth, tiling across multiple stages and on-chip memory capacity, DNNVM enumerates all potentially profitable fusion opportunities according to custom fusion templates upon XGraph, by a subgraph isomorphism algorithm. In addition, DNNVM searches for the optimal execution strategies by a heuristic shortest-path algorithm. On Xilinx [email protected], we achieve up to 1.26x speedup than naïve implementations without fusion on GoogLeNet. On Xilinx [email protected], we achieve the throughput of 2.82 TOPs/s for VGG, 1.38 TOPs/s for ResNet50 - he fastest ever reported on comparable FPGAs. Shuang Liang 0010, Lingzhi Sui, Jiantao Qiu, Xijie Jia, Yushun Wang, Yu Wang 0002 |
FPGA | 3 |
| 2019 | A High-Performance CNN Processor Based on FPGA for MobileNetsabstractConvolution neural networks (CNNs) have been widely applied in the fields of computer vision tasks. However, it is hard to deploy those standard neural networks into embedded devices because of their large amount of operations and parameters. MobileNet, the state-of-the-art CNN which adopts depthwise separable convolution to replace the standard convolution has significantly reduced operations and parameters with only limited loss in accuracy. A high-performance CNN processor based on FPGA is proposed in this paper. To improve the efficiency, two dedicated computing engines named Conv Engine and Dwcv Engine were designed for pointwise convolution and depthwise convolution respectively. The schedule for Conv Engine and Dwcv Engine has significantly improved the efficiency of our accelerator. Furthermore, we designed a special architecture called Channel Augmentation to improve the efficiency in the first layer of MobileNets. The accelerator can be flexibly deployed to various devices with different configurations to balance hardware resources and computational performance. We implemented our accelerator on ZU2 and ZU9 MPSoC FPGAs. The classification on ImageNet achieved 205.3 frames per second(fps) on ZU2 and 809.8 fps on ZU9, which is 15.4x speedup on ZU2 and 60.7x speedup on ZU9 compared to CPU. We also deployed MobileNet + SSD network on our accelerator for object detection, and achieved 31.0 fps on ZU2 and 124.3 fps on ZU9. Di Wu 0013, Xijie Jia, Tianping Li, Lingzhi Sui, Dongliang Xie |
FPL | 6 |
| 2018 | Real-Time Object Detection and Semantic Segmentation Hardware System with Deep Learning NetworksabstractAdvanced Driver Assistance Systems (ADAS) help the driver in the driving process by detecting objects, doing basic classification, implementing safety guards and so on. Convolution Neural Networks (CNN) has been proved to be an essential to support ADAS. We designed an architecture named Aristotle to execute neural networks for both object detection and semantic segmentation on FPGA. DNNDK (Deep Learning Development Toolkit), a full-stack software tool, with tens of compilation optimization techniques is proposed to improve the energy efficiency and make it easy to develop. The Aristotle architecture is implemented on Xilinx ZU9 FPGA, and two networks are deployed on it to execute object detection and semantic segmentation, respectively. Shaoxia Fang, Shuang Liang 0010, Dongliang Xie, Zhongmin Chen, Lingzhi Sui, Yu Wang 0002 |
FPT | 7 |
| 2018 | Angel-Eye: A Complete Design Flow for Mapping CNN Onto Embedded FPGAabstractConvolutional neural network (CNN) has become a successful algorithm in the region of artificial intelligence and a strong candidate for many computer vision algorithms. But the computation complexity of CNN is much higher than traditional algorithms. With the help of GPU acceleration, CNN-based applications are widely deployed in servers. However, for embedded platforms, CNN-based solutions are still too complex to be applied. Various dedicated hardware designs on field-programmable gate arrays (FPGAs) have been carried out to accelerate CNNs, while few of them explore the whole design flow for both fast deployment and high power efficiency. In this paper, we investigate state-of-the-art CNN models and CNN-based applications. Requirements on memory, computation and the flexibility of the system are summarized for mapping CNN on embedded FPGAs. Based on these requirements, we propose Angel-Eye, a programmable and flexible CNN accelerator architecture, together with data quantization strategy and compilation tool. Data quantization strategy helps reduce the bit-width down to 8-bit with negligible accuracy loss. The compilation tool maps a certain CNN model efficiently onto hardware. Evaluated on Zynq XC7Z045 platform, Angel-Eye is 6× faster and 5× better in power efficiency than peer FPGA implementation on the same platform. Applications of VGG network, pedestrian detection and face alignment are used to evaluate our design on Zynq XC7Z020. NIVIDA TK1 and TX1 platforms are used for comparison. Angel-Eye achieves similar performance and delivers up to 16× better energy efficiency. Kaiyuan Guo, Lingzhi Sui, Jiantao Qiu, Song Han 0003, Yu Wang 0002, Huazhong Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | From model to FPGA: Software-hardware co-design for efficient neural network accelerationabstractPresents a collection of slides covering the following topics: FPGA; software-hardware co-design; neural network acceleration; DeePhi Tech; deep learning; CNN acceleration; efficient inference engine; processing element architecture; and LSTM. Kaiyuan Guo, Lingzhi Sui, Jiantao Qiu, Song Han 0003, Yu Wang 0002, Huazhong Yang |
Hot Chips Symposium | 2 |