VLDB 2026 Research / reviewers in the wild / expert
Zhi Gang Liu
dblp:235/2785
· DBLP profile ↗
3ranked-venue papers
3as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 85% Energy-efficient computing · 8% Integrated circuit design · 6% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.6 | 1 | 2022 | S2TA: Exploiting Structured Sparsity for Energy-Efficient Mobile CNN Acceleration · HPCA 2022 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
sparse CNN accelerator |
0.6 | 1 | 2022 | S2TA: Exploiting Structured Sparsity for Energy-Efficient Mobile CNN Acceleration · HPCA 2022 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator › convolution acceleration
convolution accelerator |
0.4 | 1 | 2020 | Efficient Residue Number System Based Winograd Convolution · ECCV (19) 2020 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 1 | 2019 | Learning Low-precision Neural Networks without Straight-Through Estimator (STE) · IJCAI 2019 |
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator |
0.2 | 1 | 2022 | S2TA: Exploiting Structured Sparsity for Energy-Efficient Mobile CNN Acceleration · HPCA 2022 |
Hardware accelerators and domain-specific architectures › sparsity exploitation
structured sparsity |
0.2 | 1 | 2022 | S2TA: Exploiting Structured Sparsity for Energy-Efficient Mobile CNN Acceleration · HPCA 2022 |
Integrated circuit design
residue number system arithmetic |
0.1 | 1 | 2020 | Efficient Residue Number System Based Winograd Convolution · ECCV (19) 2020 |
Methods — techniques the papers use, named apart from their topics
systolic array · 0.6structured sparsity · 0.6density bound block · 0.6winograd convolution · 0.4residue number system · 0.4stochastic gradient descent · 0.4alpha blending · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | S2TA: Exploiting Structured Sparsity for Energy-Efficient Mobile CNN AccelerationabstractExploiting sparsity is a key technique in accelerating quantized convolutional neural network (CNN) inference on mobile devices. Prior sparse CNN accelerators largely exploit unstructured sparsity and achieve significant speedups. Due to the unbounded, largely unpredictable sparsity patterns, however, exploiting unstructured sparsity requires complicated hardware design with significant energy and area overhead, which is particularly detrimental to mobile/IoT inference scenarios where energy and area efficiency are crucial.We propose to exploit structured sparsity, more specifically, Density Bound Block (DBB) sparsity for both weights and activations. DBB block tensors bound the maximum number of non-zeros per block. DBB thus exposes statically predictable sparsity patterns that enable lean sparsity-exploiting hardware and efficient memory access. We propose new hardware primitives to implement DBB sparsity for (static) weights and (dynamic) activations, respectively, with very low overheads.Building on top of the primitives, we describe S2TA, a systolic array-based CNN accelerator that exploits joint weight and activation DBB sparsity and new dimensions of data reuse unavailable on the traditional systolic array. S2TA in 16nm achieves more than 2× speedup and energy reduction compared to a strong baseline of a systolic array with zero-value clock gating, over five popular CNN benchmarks. Compared to two recent non-systolic sparse accelerators, Eyeriss v2 (65nm) and SparTen (45nm), S2TA in 65nm uses about 2.2× per and 3.1× less energy inference, respectively. Zhi Gang Liu, Paul N. Whatmough, Yuhao Zhu 0001, Matthew Mattina |
HPCA | 1 |
| 2020 | Efficient Residue Number System Based Winograd Convolution
Zhi Gang Liu, Matthew Mattina |
ECCV (19) | 1 |
| 2019 | Learning Low-precision Neural Networks without Straight-Through Estimator (STE)abstractThe Straight-Through Estimator (STE) is widely used for back-propagating gradients through the quantization function, but the STE technique lacks a complete theoretical understanding. We propose an alternative methodology called alpha-blending (AB), which quantizes neural networks to low precision using stochastic gradient descent (SGD). Our AB method avoids STE approximation by replacing the quantized weight in the loss function by an affine combination of the quantized weight w_q and the corresponding full-precision weight w with non-trainable scalar coefficient alpha and (1- alpha). During training, alpha is gradually increased from 0 to 1; the gradient updates to the weights are through the full precision term, (1-alpha) * w, of the affine combination; the model is converted from full-precision to low precision progressively. To evaluate the AB method, a 1-bit BinaryNet on CIFAR10 dataset and 8-bits, 4-bits MobileNet v1, ResNet_50 v1/2 on ImageNet are trained using the alpha-blending approach, and the evaluation indicates that AB improves top-1 accuracy by 0.9\%, 0.82\% and 2.93\% respectively compared to the results of STE based quantization. Zhi Gang Liu, Matthew Mattina |
IJCAI | 1 |