Wantao Liu

dblp:26/2428 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
6since 2021 · last 2022
0000-0002-8733-2459ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2022 MakeItSmile: Detail-Enhanced Smiling Face Reenactment
abstract
Given a target face and a driving face, face reenactment aims to transfer attributes from the driving face to the target face. In the last decade, a great number of methods have been proposed to generate realistic reenacted faces. However, when these methods are applied to generate a smiling face, most of them can only get a mouth with blurry teeth, making the reenacted face unrealistic. This problem is mainly caused by incomplete tooth structure in the target face image under the setting of one-shot reenactment. In order to obtain smiling reenacted faces with detailed tooth structure, our method uses the tooth information from the driving face rather than the target face. Furthermore, to better represent the tooth structure and expressions of the driving face, we extract the texture with a carefully designed geometry-aware encoder. By training the encoder with tooth segmentation task and non-identity classification task, we acquire refined tooth representations and meanwhile derive the non-identity part of the driving face. We also design a specific generator to fuse the tooth texture features into the target face. Moreover, we add a mouth loss function to further ensure the high definition of the smiling reenacted face. We compare our method to existing state-of-the-art approaches. The experiments show that our method gets comparable results on non-smiling face reenactment and has superior performance on smiling face reenactment.
Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Wantao Liu, Jiao Dai, Jizhong Han
IJCNN4
2022 EAIS: Energy-aware adaptive scheduling for CNN inference on high-performance GPUs
Chunrong Yao, Wantao Liu, Weiqing Tang, Songlin Hu 0001
Future Gener. Comput. Syst.2
2022 EALI: Energy-aware layer-level scheduling for convolutional neural network inference services on GPUs
Chunrong Yao, Wantao Liu, Songlin Hu 0001, Weiqing Tang
Neurocomputing2
2021 Workload Prediction and VM Clustering Based Server Energy Optimization in Enterprise Cloud Data Center
Wantao Liu, Biyu Zhou, Congfeng Jiang, Ruixuan Li 0001, Songlin Hu 0001
ICA3PP (3)2
2021 Deep Reinforcement Learning based Optimization of Battery Charging and Discharging Management for Data Center
abstract
With the progress of power energy storage technologies in capacity, cycle life and reliability, data center can optimize its utilization of energy storage battery to reduce its Total Cost of Ownership (TCO). Data center can cut the peak and fill the valley of their power consumption graphs with proper management of battery charging and discharging, while maintaining uninterruptible power system (UPS) capacity. This paper studies the control technology of data center battery charging and discharging based on Deep Reinforcement Learning (DRL). According to the electricity price and status and cycle life of batteries, the appropriate time is selected to charge and discharge batteries in order to maximize the electricity bill savings. To achieve higher benefits, the system state, charging and discharging actions, reward function and a neural network structure were designed in detail. According to the simulation results, the proposed algorithm can infer the best savings strategy in both USA and Beijing electricity price systems. Compared with the baseline algorithm, the priority experience playback Deep Q-network (DQN) can increase the energy storage savings up to 47% and 55% with the electricity prices of the USA and Beijing respectively.
Wantao Liu, Wei Jiang 0028, Ruixuan Li 0001, Songlin Hu 0001
IJCNN2
2021 Evaluating and analyzing the energy efficiency of CNN inference on high-performance GPU
abstract
Summary Convolutional neural network (CNN) inference usually runs on high‐performance graphic processing units (GPUs). Since GPU is a high power consumption unit, that makes the energy consumption increases sharply due to the deep learning tasks. The energy efficiency of CNN inference is not only related to the software and hardware configurations, but also closely related to the application requirements of inference tasks. However, it is not clear on GPUs at present. In this paper, we conduct a comprehensive study on the model‐level and layer‐level energy efficiency of popular CNN models. The results point out several opportunities for further optimization. We also analyze the parameter settings (i.e., batch size, dynamic voltage and frequency scaling) and propose a revenue model to allow an optimal trade‐off between energy efficiency and latency. Compared with the default settings, the optimal settings can improve revenue by up to 15.31 × . We obtain the following main findings: (i) GPUs do not exploit the parallelism from the model depth and small convolution kernels, resulting in low energy efficiency. (ii) Convolutional layers are the most energy‐consuming CNN layers. However, due to the cache, the power consumption of all layers is relatively balanced. (iii) The energy efficiency of TensorRT is 1.53 × than that of TensorFlow.
Chunrong Yao, Wantao Liu, Weiqing Tang, Jinrong Guo, Songlin Hu 0001, Wei Jiang 0028
Concurr. Comput. Pract. Exp.2
2020 Accelerating Distributed Deep Learning By Adaptive Gradient Quantization
abstract
To accelerate distributed deep learning, gradient quantization technique is widely used to reduce the communication cost. However, the existing quantization schemes suffer from either model accuracy degradation or low compression ratio (arisen from a redundant setting of quantization level or high overhead in determining the level). In this work, we propose a novel adaptive quantization scheme (AdaQS) to explore the balance between model accuracy and quantization level. AdaQS determines the quantization level automatically according to gradient's mean to standard deviation ratio (MSDR). Then, to reduce the quantization overhead, we employ a computationally-friendly way of moment estimation to calculate the MSDR. Finally, theoretical analysis of AdaQS's convergence is conducted for non-convex objectives. Experiments demonstrate that AdaQS performs excellently on very deep model GoogleNet with 2.55% accuracy improvement relative to vanilla SGD and achieves 1.8x end-to-end speedup on AlexNet in a distributed cluster with 4*4 GPUs.
Jinrong Guo, Wantao Liu, Wang Wang, Jizhong Han, Ruixuan Li 0001, Songlin Hu 0001
ICASSP2
2019 AccUDNN: A GPU Memory Efficient Accelerator for Training Ultra-Deep Neural Networks
abstract
With the implementation of mainstream DL frameworks, scarce GPU memory resource is the primary bottleneck that hinders the trainability and training efficiency of ultra-deep neural networks (UDNN). Prior memory optimization works focus on removing the trainability restriction but leave the training efficiency out of consideration. To fill the gap, we present "AccUDNN", an accelerator that aims to make full use of finite GPU memory resource to speed up the training process of UDNN in this paper. AccUDNN mainly includes two modules: memory optimizer and hyperparameter tuner. Memory optimizer develops a novel performance-model guided dynamic swap out/in strategy to meet trainability first and further remedy the efficiency degradation in other swapping strategies. Then, a hyperparameter tuner is designed to explore the efficiency-optimal minibatch size and the matched learning rate after applying the dynamic swapping strategy. Evaluations demonstrate that AccUDNN cuts down the GPU memory requirement of ResNet-152 from more than 24GB to 8GB. In turn, given 12GB GPU memory budget, the efficiency-optimal minibatch size can reach 4.2x larger than Caffe and finally improve the scaling efficiency (speedup) of 8 GPUs' cluster by 1.9x.
Jinrong Guo, Wantao Liu, Wang Wang, Chunrong Yao, Jizhong Han, Ruixuan Li 0001, Songlin Hu 0001
ICCD2
2019 A GPU memory efficient speed-up scheme for training ultra-deep neural networks: poster
abstract
Ultra-deep neural network(UDNN) tends to yield higher-quality model but its training process is often difficult to handle. Scarce GPU DRAM capacity is the primary bottleneck that limits the depth of neural network and the range of trainable minibatch size. In this paper, we present a scheme that dedicates to make the utmost use of finite GPU memory resource to speed up the training process for UDNN. Firstly, a performance-model guided dynamic swap out/in strategy between GPU and host memory is carefully orchestrated to tackle the out-of-memory problem without introducing performance penalty. Then, a hyperparameter (minibatch size, learning rate) tuning policy is designed to explore the optimal configuration after applying the swap strategy from the perspectives of training time and final accuracy simultaneously. Finally, we verify the effectiveness of our scheme in both single and distributed GPU mode.
Jinrong Guo, Wantao Liu, Wang Wang, Qu Lu, Songlin Hu 0001, Jizhong Han, Ruixuan Li 0001
PPoPP2
2018 CGAN Based Cloud Computing Server Power Curve Generating
Wantao Liu, Songlin Hu 0001
ICA3PP (4)2
2018 Multi-stage Gradient Compression: Overcoming the Communication Bottleneck in Distributed Deep Learning
Qu Lu, Wantao Liu, Jizhong Han, Jinrong Guo
ICONIP (1)2
2018 Temperature and Power Aware Server Placement Optimization for Enterprise Data Center
abstract
In data center, the optimal server location selection is related to many factors, such as power, temperature and space location. In order to reduce the local hot point in computer room, decrease the energy consumption of cooling system and eliminate the power fragment of rack, this paper studies the optimization issue of server placement for the enterprise data center based on the joint perception of temperature and power. First, based on the deep learning algorithm, the model of the power, temperature and cooling power of equipment in computer room is established. The power curve model of the server is estimated by using the Conditional Generative Adversarial Network (CGAN), and then the optimal placement position of server is calculated by genetic algorithm. Through simulation experiments, it is proved that the method proposed in this paper is stable and feasible. Compared with the three placement algorithms, this method can reduce the cooling energy consumption by 4-6% and reduce the new power demand of 6-9%, which can effectively improve the energy consumption efficiency of data center.
Wantao Liu, Dongxia Bai
ICPADS2
2015 DualTable: A hybrid storage model for update optimization in Hive
abstract
Hive is the most mature and prevalent data warehouse tool providing SQL-like interface in the Hadoop ecosystem. It is successfully used in many Internet companies and shows its value for big data processing in traditional industries. However, enterprise big data processing systems as in Smart Grid applications usually require complicated business logics and involve many data manipulation operations like updates and deletes. Hive cannot offer sufficient support for these while preserving high query performance. Hive using the Hadoop Distributed File System (HDFS) for storage cannot implement data manipulation efficiently and Hive on HBase suffers from poor query performance even though it can support faster data manipulation. There is a project based on Hive issue Hive-5317 to support update operations, but it has not been finished in Hive's latest version. Since this ACID compliant extension adopts same data storage format on HDFS, the update performance problem is not solved. In this paper, we propose a hybrid storage model called DualTable, which combines the efficient streaming reads of HDFS and the random write capability of HBase. Hive on DualTable provides better data manipulation support and preserves query performance at the same time. Experiments on a TPC-H data set and on a real smart grid data set show that Hive on DualTable is up to 10 times faster than Hive when executing update and delete operations.
Songlin Hu 0001, Wantao Liu, Tilmann Rabl, Hans-Arno Jacobsen, Xubin Pei, Jiye Wang
ICDE2
2014 DGFIndex for Smart Grid: Enhancing Hive with a Cost-Effective Multidimensional Range Index
abstract
In Smart Grid applications, as the number of deployed electric smart meters increases, massive amounts of valuable meter data is generated and collected every day. To enable reliable data collection and make business decisions fast, high throughput storage and high-performance analysis of massive meter data become crucial for grid companies. Considering the advantage of high efficiency, fault tolerance, and price-performance of Hadoop and Hive systems, they are frequently deployed as underlying platform for big data processing. However, in real business use cases, these data analysis applications typically involve multidimensional range queries (MDRQ) as well as batch reading and statistics on the meter data. While Hive is high-performance at complex data batch reading and analysis, it lacks efficient indexing techniques for MDRQ. In this paper, we propose DGFIndex, an index structure for Hive that efficiently supports MDRQ for massive meter data. DGFIndex divides the data space into cubes using the grid file technique. Unlike the existing indexes in Hive, which stores all combinations of multiple dimensions, DGFIndex only stores the information of cubes. This leads to smaller index size and faster query processing. Furthermore, with pre-computing user-defined aggregations of each cube, DGFIndex only needs to access the boundary region for aggregation query. Our comprehensive experiments show that DGFIndex can save significant disk space in comparison with the existing indexes in Hive and the query performance with DGFIndex is 2-50 times faster than existing indexes in Hive and HadoopDB for aggregation query, 2-5 times faster than both for non-aggregation query, 2-75 times faster than scanning the whole table in different query selectivity.
Yue Liu 0006, Songlin Hu 0001, Tilmann Rabl, Wantao Liu, Hans-Arno Jacobsen, Kaifeng Wu, Jintao Li 0001
Proc. VLDB Endow.4
2011 Moving huge scientific datasets over the Internet
abstract
SUMMARY Modern scientific experiments can generate hundreds of gigabytes to terabytes or even petabytes of data that may be maintained in large numbers of relatively small files. Frequently, these data must be disseminated to remote collaborators or computational centers for data analysis. Moving this dataset with high performance and strong robustness and providing a simple interface for users are challenging tasks. We present a data transfer framework comprising a high‐performance data transfer library based on GridFTP, an extensible data scheduler with four data scheduling policies, and a GUI that allows users to transfer their dataset easily, reliably, and securely. This system incorporates automatic tuning mechanisms to select at runtime the number of concurrent threads to be used for transfers. Also included are restart mechanisms for handling client, network, and server failures. Experimental results indicate that our data transfer system can significantly improve data transfer performance and can recover well from failures. Copyright © 2011 John Wiley & Sons, Ltd.
Wantao Liu, Brian Tieman, Rajkumar Kettimuthu, Ian T. Foster
Concurr. Comput. Pract. Exp.1
2010 A data transfer framework for large-scale science experiments
abstract
Modern scientific experiments can generate hundreds of gigabytes to terabytes or even petabytes of data that may furthermore be maintained in large numbers of relatively small files. Frequently, this data must be disseminated to remote collaborators or computational centers for data analysis. Moving this data with high performance and strong robustness and providing a simple interface for users are challenging tasks. We present a data transfer framework comprising a high-performance data transfer library based on GridFTP, a data scheduler, and a graphical user interface that allows users to transfer their data easily, reliably, and securely. This system incorporates automatic tuning mechanisms to select at runtime the number of concurrent threads to be used for transfers. Also included are restart mechanisms capable of dealing with client, network, and server failures. Experimental results indicate that our data transfer system can significantly improve data transfer performance and can recover well from failures.
Wantao Liu, Brian Tieman, Rajkumar Kettimuthu, Ian T. Foster
HPDC1
2010 Resilient Virtual Network Service Provision in Network Virtualization Environments
abstract
Network Virtualization has recently emerged to provide scalable, customized and on-demand virtual network services over a shared substrate network. How to provide VN services with resiliency guarantees against network failures has become a critical issue, meanwhile the service resource usages should be minimized under the strict constraints such as link bandwidth capability and service resiliency guarantees etc. In this paper, we present a resource allocation algorithm to balance the tradeoff between service resource consumptions and service resiliency. By exploiting a heuristic VN mapping scheme and a restoration path selection scheme based on intelligent bandwidth sharing, the algorithm simultaneously makes cost-effective usage of network resources and protects VN services against network failures. We perform evaluations and find that the algorithm is near optimal in terms of network resource usage, especially the additional restoration bandwidth cost for resiliency protection.
Jianxin Li 0002, Tianyu Wo, Chunming Hu, Wantao Liu
ICPADS5
2008 Communicating Security Assertions over the GridFTP Control Channel
abstract
The GridFTP by Allcock, W. (2003) protocol defines a general- purpose mechanism for secure, reliable, high-performance data movement. GridFTP has been widely used for efficiently transferring large volumes of data. It is based on the Internet FTP protocol and thus involves two communication channels: a control channel and a data channel. The commands and responses flow over the control channel, and the data is transmitted over the data channel.
Rajkumar Kettimuthu, Wantao Liu, Frank Siebenlist, Ian T. Foster
eScience2