Deng Liu

dblp:86/6653 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
2since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 72% Memory systems · 12% Reconfigurable computing and FPGAs · 7%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
0.812024
Bit-Balance: Model-Hardware Codesign for Accelerating NNs by Exploiting Bit-Level Sparsity · IEEE Trans. Computers 2024
Machine learning › Efficient and distributed learning › model compression
quantization
0.812024
Bit-Balance: Model-Hardware Codesign for Accelerating NNs by Exploiting Bit-Level Sparsity · IEEE Trans. Computers 2024
Hardware accelerators and domain-specific architectures › sparsity exploitation
bit-level sparsity exploitation
0.812024
Bit-Balance: Model-Hardware Codesign for Accelerating NNs by Exploiting Bit-Level Sparsity · IEEE Trans. Computers 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN accelerator
bit-serial DNN accelerator
0.812024
Bit-Balance: Model-Hardware Codesign for Accelerating NNs by Exploiting Bit-Level Sparsity · IEEE Trans. Computers 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.812024
Bit-Balance: Model-Hardware Codesign for Accelerating NNs by Exploiting Bit-Level Sparsity · IEEE Trans. Computers 2024
Parallel and multicore computing
load balancing
0.212024
Bit-Balance: Model-Hardware Codesign for Accelerating NNs by Exploiting Bit-Level Sparsity · IEEE Trans. Computers 2024
Reconfigurable computing and FPGAs › coarse-grained reconfigurable architecture
processing element array
0.212024
Bit-Balance: Model-Hardware Codesign for Accelerating NNs by Exploiting Bit-Level Sparsity · IEEE Trans. Computers 2024
Memory systems › cache management › storage caching
cache space management
0.212014
vCacheShare: Automated Server Flash Cache Space Management in a Virtualization Environment · USENIX ATC 2014
Memory systems › cache management › storage caching
flash cache
0.212014
vCacheShare: Automated Server Flash Cache Space Management in a Virtualization Environment · USENIX ATC 2014
Cloud and datacenter computing
virtualization
0.112014
vCacheShare: Automated Server Flash Cache Space Management in a Virtualization Environment · USENIX ATC 2014

Methods — techniques the papers use, named apart from their topics

model-hardware co-design · 1.5channel-wise quantization · 1.5
YearPublicationVenuePosition
2024 Bit-Balance: Model-Hardware Codesign for Accelerating NNs by Exploiting Bit-Level Sparsity
abstract
Bit-serial architectures can handle Neural Networks (NNs) with different weight precision, achieving higher resource efficiency compared with bit-parallel architectures. Besides, the weights contain abundant zero bits owing to the fault tolerance of NNs, indicating that bit sparsity of NNs can be further exploited for performance improvement. However, the irregular proportion of zero bits in each weight causes imbalanced workloads in the Processing Element (PE) array, which degrades performance or induces overhead for sparse processing. Thus, this article proposed a channel-wise bit-sparsity quantization method that keeps the non-zero bit number of each weight in each channel from exceeding a certain threshold and clusters the channels with the same threshold to balance the workloads in PE array with little accuracy loss. Then, we co-designed a sparse bit-serial architecture, called Bit-balance, to improve overall performance, supporting weight-bit sparsity and adaptive bitwidth computation. The whole design was implemented with 65 nm technology at 1 GHz and performs at 447-, 37-, 59-, 240-, and 19-frame/s for AlexNet, VGG-16, ResNet-50, GoogleNet, and Yolo-v3 respectively. Compared with sparse bit-serial accelerator, Bitlet, Bit-balance achieves 1.6$\boldsymbol{\times}$2.1$\boldsymbol{\times}$energy efficiency (frame/J) and 2.3$\boldsymbol{\times}$3.6$\boldsymbol{\times}$resource efficiency (frame/mm${}^{\mathbf{2}}$).
Zhiwei Zou, Deng Liu, Wendi Sun, Song Chen 0001, Yi Kang
IEEE Trans. Computers3
2023 Sense: Model-Hardware Codesign for Accelerating Sparse CNNs on Systolic Arrays
abstract
Sparsity is an intrinsic property of convolutional neural networks (CNNs), worth exploiting for CNN accelerators. However, the extra processing involved comes with hardware overhead, resulting in only marginal profits for most architectures. Meanwhile, systolic arrays have become increasingly competitive on CNN acceleration for its high spatiotemporal locality and low hardware overhead. However, the irregularity of sparsity induces imbalanced workloads under the rigid systolic dataflow, causing performance degradation. Thus, this article proposed a systolic-array-based architecture, called Sense, for sparse CNN acceleration by model-hardware codesign, enabling large performance gains. To balance input feature map (IFM) and weight loads across the processing element (PE) array, we applied channel clustering to gather IFMs with approximate sparsity for array computation and codesigned a load-balancing weight pruning method to keep the sparsity ratio of each kernel at a certain value with little accuracy loss, improving PE utilization and overall performance. In addition, adaptive dataflow configuration was applied to determine the computing strategy based on the storage ratio of IFMs and weights, lowering$1.17\times $–$1.8\times $dynamic random access memory (DRAM) access compared with Swallow and further reducing system energy consumption. The whole design was implemented on ZynqZCU102 with 200 MHz and performs at 471, 34, 53, and 191 image/s for AlexNet, VGG-16, ResNet-50, and GoogleNet, respectively. Compared with sparse systolic-array-based accelerators, Swallow, fusion-enabled systolic architecture (FESA), and SPOTS, Sense achieves$0.97\times $–$2.18\times $,$1.3\times $–$1.67\times $, and$0.94\times $–$1.82\times $energy efficiency (image/J) on these CNNs, respectively.
Deng Liu, Zhiwei Zou, Wendi Sun, Song Chen 0001, Yi Kang
IEEE Trans. Very Large Scale Integr. Syst.2
2019 A diurnal flux balance model of Synechocystis sp. PCC 6803 metabolism
abstract
Phototrophic organisms such as cyanobacteria utilize the sun's energy to convert atmospheric carbon dioxide into organic carbon, resulting in diurnal variations in the cell's metabolism. Flux balance analysis is a widely accepted constraint-based optimization tool for analyzing growth and metabolism, but it is generally used in a time-invariant manner with no provisions for sequestering different biomass components at different time periods. Here we present CycleSyn, a periodic model of Synechocystis sp. PCC 6803 metabolism that spans a 12-hr light/12-hr dark cycle by segmenting it into 12 Time Point Models (TPMs) with a uniform duration of two hours. The developed framework allows for the flow of metabolites across TPMs while inventorying metabolite levels and only allowing for the utilization of currently or previously produced compounds. The 12 TPMs allow for the incorporation of time-dependent constraints that capture the cyclic nature of cellular processes. Imposing bounds on reactions informed by temporally-segmented transcriptomic data enables simulation of phototrophic growth as a single linear programming (LP) problem. The solution provides the time varying reaction fluxes over a 24-hour cycle and the accumulation/consumption of metabolites. The diurnal rhythm of metabolic gene expression driven by the circadian clock and its metabolic consequences is explored. Predicted flux and metabolite pools are in line with published studies regarding the temporal organization of phototrophic growth in Synechocystis PCC 6803 paving the way for constructing time-resolved genome-scale models (GSMs) for organisms with a circadian clock. In addition, the metabolic reorganization that would be required to enable Synechocystis PCC 6803 to temporally separate photosynthesis from oxygen-sensitive nitrogen fixation is also explored using the developed model formalism.
Debolina Sarkar, Thomas J. Mueller, Deng Liu, Himadri B. Pakrasi, Costas D. Maranas
PLoS Comput. Biol.3
2017 Improving Flash Resource Utilization at Minimal Management Cost in Virtualized Flash-Based Storage Systems
abstract
Effectively leveraging Flash resources has emerged as a highly important problem in enterprise storage systems. One of the popular techniques today is to use Flash as a secondary-level host-side cache in the virtual machine environment. Although this approach delivers IO acceleration for VMs' IO workloads, it might not be able to fully exploit the outstanding performance of Flash and justify the high cost-per-GB of Flash resources. In this paper, we design new VMware Flash Resource Managers (VFRM and GLB-VFRM) under the consideration of both performance and the incurred cost for managing Flash resources. Specifically, VFRM and GLB-VFRM aim to maximize the utilization of Flash resources with minimal CPU, memory and IO cost in managing and operating Flash for a dedicated enterprise workload and multiple heterogeneous enterprise workloads, respectively. Our new Flash resource managers adopt the ideas of thermodynamic heating and cooling to identify data blocks that can benefit the most from being put on Flash and migrate data blocks between Flash and magnetic disks in a lazy and asynchronous mode. Experimental evaluation of the prototype shows that both VFRM and GLB-VFRM achieve better cost-effectiveness than traditional caching solutions, i.e., obtaining IO hit ratios even slightly better than some of the conventional algorithms as Flash size increases yet costing orders of magnitude less IO bandwidth.
Jianzhe Tai, Deng Liu, Zhengyu Yang 0001, Xiaoyun Zhu, Jack Lo, Ningfang Mi
IEEE Trans. Cloud Comput.2
2017 Edge-Aware Label Propagation for Mobile Facial Enhancement on the Cloud
abstract
This paper proposes a facial enhancement framework with mask generation for cloud-based mobile applications. We mathematically analyze and unify the mask generation, as well as the state-of-the-art region-aware mask and edit propagation techniques, from a graph-based semi supervised learning perspective. Then we propose a label propagation model with a new edge-aware structure and guided feature for mask generation. The limit analysis of the model leads to a fast algorithm, which reduces the intrinsic computation cost. Then we develop a flexible and efficient cloud-based PaaS system, called FaceMore, for intelligent mobile face enhancement. The flexibility and extendibility of the cloud-based architecture facilitates intelligent facial enhancement applications and the parallel processing effectively improves the efficiency of the algorithm. Qualitative and quantitative evaluations were performed for mask propagation and face enhancement. Comparisons with the previous methods and five representative commercial systems, including PicTreat, Portraiture, Portrait+, Meitu, and Baidu Motu, illustrate the robustness and effectiveness of our method.
Lingyu Liang, Deng Liu
IEEE Trans. Circuits Syst. Video Technol.3
2015 FaceMore: A Face Beautification Platform on the Cloud
abstract
Face More, a cloud-based face beautification platform for intelligent face manipulation, is developed in this work. It provides flexible and efficient cloud API to develop automatic or interactive face retouching applications. A web-site, www.facemore.net, is built on Face More, where user can upload images and obtain various online face beautification services. To obtain automatic inhomogeneous editing effects in a natural and efficient manner, we analyze the novel tool, called region-aware mask, from semi-supervised learning perspective. We reformulate the optimized-based model of region-aware mask using label propagation, and propose a fast approximate algorithm for mask generation, which leads to about 40% speed improvement while maintain high visual quality for face beatification. Qualitative and quantitative evaluations were performed for mask generation and face beatification. Comparisons with five representative commercial systems, including PicTreat, Portraiture, Portrait+, Meitu and Baidu Motu, illustrate the effectiveness of our system for image enhancement of facial lighting, smoothness and color.
Lingyu Liang, Deng Liu
SMC2
2014 VFRM: Flash Resource Manager in VMware ESX Server
abstract
One popular approach of leveraging Flash technology in the virtual machine environment today is using it as a secondary-level host-side cache. Although this approach delivers I/O acceleration for a single VM workload, it might not be able to fully exploit the outstanding performance of Flash and justify the high cost-per-GB of Flash resources. In this paper, we present the design for VMware Flash Resource Manager (VFRM), which aims to maximize the utilization of Flash resources with minimal CPU, memory and I/O cost for managing and operating Flash. It borrows the ideas of heating and cooling from thermodynamics to identify the data blocks that benefit most from being put on Flash, and lazily and asynchronously migrates the data blocks between Flash and spinning disks. Experimental evaluation of the prototype shows that VFRM achieves better cost-effectiveness than traditional caching solutions, and costs orders of magnitude less memory and I/O bandwidth.
Deng Liu, Ningfang Mi, Jianzhe Tai, Xiaoyun Zhu, Jack Lo
NOMS1
2014 vCacheShare: Automated Server Flash Cache Space Management in a Virtualization Environment
Xiaosong Ma, Sandeep Uttamchandani, Deng Liu
USENIX ATC5
2013 S-CAVE: Effective SSD caching to improve virtual machine storage performance
abstract
A unique challenge for SSD storage caching management in a virtual machine (VM) environment is to accomplish the dual objectives: maximizing utilization of shared SSD cache devices and ensuring performance isolation among VMs. In this paper, we present our design and implementation of S-CAVE, a hypervisor-based SSD caching facility, which effectively manages a storage cache in a Multi-VM environment by collecting and exploiting runtime information from both VMs and storage devices. Due to a hypervisor's unique position between VMs and hardware resources, S-CAVE does not require any modification to guest OSes, user applications, or the underlying storage system. A critical issue to address in S-CAVE is how to allocate limited and shared SSD cache space among multiple VMs to achieve the dual goals. This is accomplished in two steps. First, we propose an effective metric to determine the demand for SSD cache space of each VM. Next, by incorporating this cache demand information into a dynamic control mechanism, S-CAVE is able to efficiently provide a fair share of cache space to each VM while achieving the goal of best utilizing the shared SSD cache device. In accordance with the constraints of all the functionalities of a hypervisor, S-CAVE incurs minimum overhead in both memory space and computing time. We have implemented S-CAVE in vSphere ESX, a widely used commercial hypervisor from VMWare. Our extensive experiments have shown its strong effectiveness for various data-intensive applications.
Rubao Lee, Xiaodong Zhang 0001, Deng Liu
PACT5