EDBT 2026 Demo / reviewers in the wild / expert
FuHai Yu
dblp:236/7072
· DBLP profile ↗
1ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 75% Reconfigurable computing and FPGAs · 25% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
CNN inference accelerator |
0.4 | 1 | 2019 | A Deep Learning Inference Accelerator Based on Model Compression on FPGA · FPGA 2019 |
Reconfigurable computing and FPGAs › FPGA accelerator
FPGA-based CNN inference |
0.4 | 1 | 2019 | A Deep Learning Inference Accelerator Based on Model Compression on FPGA · FPGA 2019 |
Hardware accelerators and domain-specific architectures
model compression |
0.4 | 1 | 2019 | A Deep Learning Inference Accelerator Based on Model Compression on FPGA · FPGA 2019 |
Hardware accelerators and domain-specific architectures
quantization |
0.4 | 1 | 2019 | A Deep Learning Inference Accelerator Based on Model Compression on FPGA · FPGA 2019 |
Methods — techniques the papers use, named apart from their topics
shift-and-accumulate · 0.4incremental network quantization · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | A Deep Learning Inference Accelerator Based on Model Compression on FPGAabstractConvolutional neural networks (CNN) have demonstrated state-of-the-art accuracy in image classification and object detection owing to the increase in data and computation capacity of hardware. However, this state-of-the-art achievement depends heavily on the DSP floating-point computing capability of the device, which increases the power dissipation and cost of the device. In order to solve the problem, we made the first attempt to implement a CNN computing accelerator based on shift operation on FPGA. In this accelerator, an efficient Incremental Network Quantization (INQ) method was applied to compress the CNN model from full precision to 4-bit integer, which represents values of either zero or power of two. Then the multiply and accumulate (MAC) operations for convolution layer and fully-connected layer was converted to shift and accumulation (SAC) operations, and SAC could be easily implemented by the logic elements of FPGA. Consequently, parallelism of CNN inference process can be further expanded. For the SqueezeNet model, single image processing latency was 0.673ms on Intel Arria 10 FPGA (Inspur F10A board) showing a slightly better result than on NVIDIA Tesla P4, and the compute capacity of FPGA increased by 1.77 times at least. Lu Jing, FuHai Yu |
FPGA | 3 |