Heng Wu 0001

dblp:89/5836-1 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
14since 2021 · last 2025
0000-0001-7903-5879ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 9 since 2021Software engineering, systems software and programming languages · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Scheduling based on Block Features for Concurrent Inference with Unseen DNN Models on GPU
abstract
Efficiently scheduling concurrent deep neural network (DNN) inference on the same GPU can significantly optimize resource utilization. Such scheduling requires accurate prediction of concurrent inference time. Existing approaches primarily rely on model-level features for prediction and scheduling, which necessitates retraining and resampling when encountering unseen models to ensure prediction accuracy. However, in MLOps pipelines, the rapid iteration of models introduces numerous unseen models, making accurate predictions highly challenging and increasing the risk of SLA violations. To address these challenges posed by unseen models, we present SKADI, a scheduling framework based on block-level feature extraction and two-stage greedy scheduling. First, SKADI introduces block-level feature extraction, decomposing DNN models into homogeneous blocks (contiguous operator sequences) to enable zero-shot inference time prediction for unseen models. Second, it proposes a round-based and two-stage greedy scheduling strategy that rapidly selects optimal model pairs and overlaps their critical operators. Experimental results show that for unseen models, SKADI reduces the MAPE of concurrent inference time prediction by 59.25% and 57.55% compared to DeInfer and Abacus. Additionally, SKADI reduces SLA violation rates by 70.6% while increasing throughput by 7.3%.
Diaohan Luo, Heran Gao, Yuewen Wu, Heng Wu 0001, Wenbo Zhang 0006
ICPP5
2024 ETS: Deep Learning Training Iteration Time Prediction based on Execution Trace Sliding Window
abstract
Deep learning (DL) has become essential across various computer science domains. Accurately predicting iteration time for DL models in diverse cloud data center environments is critical for making high-quality scheduling decisions. Existing approaches neglect the sequential features inherent in the runtime execution, leading to issues such as overlooking DL framework overhead and struggling to handle diverse sizes of DL models, resulting in either low accuracy or slow convergence of the prediction model. This paper introduces ETS, a novel iteration time prediction method utilizing execution trace sliding windows. Our observation reveals that DL models exhibit a highly sequential runtime execution nature. Building upon this insight, we leverage sliding windows to extract a novel type of sequential features from the runtime execution trace. These features comprehensively capture DL framework overhead and address the diversity challenge in DL model sizes. By combining a best-practice method to train a prediction model, we achieve high accuracy and rapid convergence simultaneously. Experimental validation on over 14,000 DL model configurations demonstrates ETS's effectiveness in predicting the iteration time of DL models, achieving a mere 5.9% prediction error with a training time at the 10-minute level, and improving scheduling outcomes by reducing job completion time by 17%.
Heng Wu 0001, Yuewen Wu, Hua Zhong 0007, Wenbo Zhang 0006, Yan Liu 0102
HPDC3
2023 InstantChain: Enhancing Order-Execute Blockchain Systems for Latency-Sensitive Applications
Heng Wu 0001, Diaohan Luo, Heran Gao, Wenbo Zhang 0006
DASFAA (1)2
2023 SPLIT: QoS-Aware DNN Inference on Shared GPU via Evenly-Sized Model Splitting
abstract
Improving QoS by simultaneously reducing the latency violation rate and jitter in the presence of multiple deep learning inference (DLI) tasks sharing a single edge computing processor remains a challenge. However, existing DLI systems at the edge, designed to maximize throughput, face performance challenges when confronted with requests with varying QoS.
Diaohan Luo, Yuewen Wu, Heng Wu 0001, Tao Wang 0030, Wenbo Zhang 0006
ICPP4
2023 2DPChain: Orchestrating Transactions in Order-Execute Blockchain to Exploit Intra-batch and Inter-batch Parallelism
Heng Wu 0001, Heran Gao, Wenbo Zhang 0006
ICSOC (1)2
2023 Topology-Aware Self-Adaptive Resource Provisioning for Microservices
abstract
Microservice architecture is a popular technology for deploying services in cloud computing, with benefits like loose coupling, high fault tolerance, and scalability. The heterogeneous resource requirements and complex interaction relations have increased the difficulty in provisioning resources for microservices with intricacy topology. Existing approaches allocate resources for different microservices separately, and thus cannot achieve optimal global performance. Moreover, these approaches extract features from specific microservice topologies. We propose a topology-aware self-adaptive resource provisioning approach for microservices. Firstly, we propose a microservice state graph to characterize the status of each microservice in an application. Then, we use graph neural networks and attention to extract the resource requirements and correlation features of microservices. Thirdly, we use a reinforcement learning-based approach to allocate resources for microservices uniformly. Finally, we evaluate our approach by conducting a series of experiments on three typical microservice applications deployed in a heterogeneous cluster. The results show that our approach is efficient in extracting resource and correlation features of microservices, and can guarantee QoS with efficient resource utilization. Our approach can reduce the End-to-End latency by 22%, and can improve resource utilization by 18% with guaranteed latency.
Tao Wang 0030, Yuewen Wu, Heng Wu 0001, Wenbo Zhang 0006
ICWS5
2023 AgileShard: Turning the Sharded Blockchain into a Real-Time Transaction Processing System
abstract
Blockchain, as an emerging transaction processing system, suffers from low throughput and high latency. Sharded blockchains can significantly increase throughput by dividing nodes into groups (i.e., shards) to process disjoint transactions in parallel. However, the diverse latency requirements of transactions in current sharded blockchains are not well met. Three challenges prevent the sharded blockchain from becoming a real-time transaction processing system, namely, the static block size, the first-come-first-served transaction packing strategy, and the load imbalance. Therefore, this paper proposes 3 methods to help turn the sharded blockchain into a real-time transaction processing system. First, we propose an inter-shard dynamic block size negotiation method that enables shards to adaptively determine the globally optimal block size based on the deadlines of pending transactions. Then, we propose a DAG-based transaction packing method for reducing the number of deadline violations and improving parallelism. Finally, we propose a minimum-cost-flow-based shard reconfiguration method to address load imbalance. Under real datasets on Ethereum, experimental results show that AgileShard using the above three methods can effectively improve the transaction deadline satisfaction rate.
Heng Wu 0001, Heran Gao, Wenbo Zhang 0006
RTSS2
2023 Hydra: Deadline-Aware and Efficiency-Oriented Scheduling for Deep Learning Jobs on Heterogeneous GPUs
abstract
With the rapid proliferation of deep learning (DL) jobs running on heterogeneous GPUs, scheduling DL jobs to meet various scheduling requirements, such as meeting deadlines and reducing job completion time (JCT), is critical. Unfortunately, existing efficiency-oriented and deadline-aware efforts are still rudimentary. They lack the capability of scheduling jobs to meet deadline requirements while reducing total JCT, especially when the jobs have various execution times on heterogeneous GPUs. Therefore, we present Hydra, a novel quantitative cost comparison approach, to address this scheduling issue. Here, the cost represents the total JCT plus a dynamic penalty calculated from the total tardiness (i.e., the delay time of exceeding the deadline) of all jobs. Hydra adopts a sampling approach that exploits the inherent iterative periodicity of DL jobs to estimate job execution times accurately on heterogeneous GPUs. Then, Hydra considers various combinations of job sequences and GPUs to obtain the minimized cost by leveraging an efficient branch-and-bound algorithm. Finally, the results of evaluation experiments on Alibaba traces show that Hydra can reduce total tardiness by 85.8% while reducing total JCT as much as possible, compared with state-of-the-art efforts.
Heng Wu 0001, Yuanjia Xu, Yuewen Wu, Hua Zhong 0001, Wenbo Zhang 0006
IEEE Trans. Computers2
2022 Serving unseen deep learning models with near-optimal configurations: a fast adaptive search approach
abstract
Public clouds provide a bewildering choice of configurations for Deep Learning (DL) models, and the choice of configuration will significantly impact the performance and budget. However, it is an obvious challenge to recommend a near-optimal configuration for a particular DL model from a wide range of candidates. The huge search overhead of finding such a configuration is the notorious cold start problem in state-of-the-art efforts, and this problem becomes more severe when they are faced with unseen DL models.
Yuewen Wu, Heng Wu 0001, Diaohan Luo, Yuanjia Xu, Wenbo Zhang 0006, Hua Zhong 0007
SoCC2
2022 EOP: efficient operator partition for deep learning inference over edge servers
abstract
Recently, Deep Learning (DL) models have demonstrated great success for its attractive ability of high accuracy used in artificial intelligence Internet of Things applications. A common deployment solution is to run such DL inference tasks on edge servers. In a DL inference, each operator takes tensors as input and run in a tensor virtual machine, which isolates resource usage among operators. Nevertheless, existing edge-based DL inference approaches can not efficiently use heterogeneous resources (e.g., CPU and low-end GPU) on edge servers and result in sub-optimal DL inference performance, since they can only partition operators in a DL inference with equal or fixed ratios. It is still a big challenge to support partition optimizations over edge servers for a wide range of DL models, such as Convolution Neural Network (CNN), Recurrent Neural Network (RNN) and Transformers.
Yuanjia Xu, Heng Wu 0001, Wenbo Zhang 0006
VEE2
2021 Talos: A Weighted Speedup-Aware Device Placement of Deep Learning Models
abstract
Efficient device placement of deep learning (DL) models, which consist of many operations, is a big challenge when heterogeneous devices (e.g., CPU, GPU) are considered. Existing average speedup and transient speedup approaches do not make full use of operation-level speedups, and the Total Operation Completion Time (TOCT) cannot be optimized efficiently.To address this challenge, we present Talos, a weighted speedup-awareness approach to optimize device placement of multiple DL models. Talos reveals operations within or across DL models have diverse speedups (from 10−1to 102) on heterogeneous devices. In addition, the execution time of operations are widely ranged (from 0.1ms to 100ms). Talos considers the two features simultaneously as weighted speedups, and treats them as costs in an incremental minimum-cost flow. Compared with state-of-the-art efforts, experiment results show that Talos can reduce TOCT by up to 50%.
Yuanjia Xu, Heng Wu 0001, Wenbo Zhang 0006, Yuewen Wu, Heran Gao, Tao Wang 0030
ASAP2
2021 Evaluating the Parallel Execution Schemes of Smart Contract Transactions in Different Blockchains: An Empirical Study
Chengzhi Li, Heng Wu 0001, Heran Gao, Songchang Jin, Tao Huang 0001, Wenbo Zhang 0006
ICA3PP (3)3
2021 Best VM Selection for Big Data Applications across Multiple Frameworks by Transfer Learning
abstract
Cloud providers are presented with a bewildering choice of VM types for a range of contemporary data processing frameworks today. However, existing performance modeling and machine learning efforts cannot pick optimal VM types for multiple frameworks simultaneously, since they are difficult to balance model accuracy and model training cost.
Yuewen Wu, Heng Wu 0001, Yuanjia Xu, Wenbo Zhang 0006, Hua Zhong 0007, Tao Huang 0001
ICPP2
2021 Apollo: Rapidly Picking the Optimal Cloud Configurations for Big Data Analytics Using a Data-Driven Approach
Yuewen Wu, Yuanjia Xu, Heng Wu 0001, Lin-Gang Su, Wenbo Zhang 0006, Hua Zhong 0007
J. Comput. Sci. Technol.3
2020 Hermes: Efficient Cache Management for Container-based Serverless Computing
abstract
Serverless computing systems are shifting towards shorter function durations and larger degrees of parallelism to eliminate intolerable latency. For container-based serverless computing, the state-of-the-art efforts fail to ensure low latency because on-demand container images reloading from remote storage can increase the data transmission rate and downgrades system performance.
Heran Gao, Heng Wu 0001, Wenbo Zhang 0006, Tao Huang 0001
Internetware3
2020 Nuka: A Generic Engine with Millisecond Initialization for Serverless Computing
abstract
Serverless computing is becoming one of the mainstream trends in cloud computing due to its advantages of simplified programming and cost saving. However, existing serverless platforms still adopt Docker container as its execution engine, which has the cold start problem and causes high common-case invocation latency. In this work, we analyze the lifecycle of common-case serverless invocation on existing serverless platforms and find that current container startup and pulling remote images are the two main reasons causing cold start so slow. Based on the study, we implement Nuka, a generic engine with millisecond initialization for serverless computing. Nuka is fully compatible with Docker interface and can smoothly re-place Docker as the execution engine of existing serverless plat-forms. Through the isolation pool that reuses Linux's isolation configurations, Nuka avoids the high cost of container startup's scalability bottleneck, which reduces container's startup time with high concurrency scale. Nuka also avoids pulling remote images through dynamically resolving and importing required software packages from local package caching. A self-adaptive container reuse strategy dynamically controls container's pause time and replica numbers, which effectively reduces the frequency of cold start. Compared with Docker, Nuka can get a millisecond initialization with high concurrency and significantly reduces average time cost of cold startup by 6× on existing serverless platforms.
Shijun Qin, Heng Wu 0001, Yuewen Wu, Yuanjia Xu, Wenbo Zhang 0006
JCC2
2020 A framework to support multi-cloud collaboration
abstract
With the rapid development of cloud computing, major cloud providers have launched various cloud services with different functions to meet customer's needs. Therefore, flexibility is extremely important when developers use these cloud services. However, APIs of cloud services change dozens of times annually without backward compatibility. It means developers have to adapt these clouds with manual efforts. Such efforts make the multi-cloud collaboration extremely complex and cannot meet the demand of flexibility. This paper describes a configuration-based multi-cloud collaboration framework, which can support new clouds with comprehensible configurations. Meanwhile, if cloud APIs are updated without backward compatibility, it can restore services during runtime with minimized configuration. The main technologies used in this article include automatic discovery, unified abstraction, dynamic mapping and incremental update. We tested the virtual machine and container services of seven well-known cloud providers. The system can support heterogeneous clouds well. When the APIs are updated, the system can restore services in less than 200 milliseconds. At the same time, the extra cost of our framework is acceptable to cloud users.
Ting Tang, Heng Wu 0001, Yuewen Wu, Yuanjia Xu, Wenbo Zhang 0006
SERVICES3
2019 Aladdin: Optimized Maximum Flow Management for Shared Production Clusters
abstract
The rise in popularity of long-lived applications (LLAs), such as deep learning and latency-sensitive online Web services, has brought new challenges for cluster schedulers in shared production environments. Scheduling LLAs needs to support complex placement constraints (e.g., to run multiple containers of an application on different machines) and larger degrees of parallelism to provide global optimization. But existing schedulers usually suffer severe constraint violations, high latency and low resource efficiency. This paper describes Aladdin, a novel cluster scheduler that can maximize resource efficiency while avoiding constraint violations: (i) it proposes a multidimensional and nonlinear capacity function to support constraint expressions; (ii) it applies an optimized maximum flow algorithm to improve resource efficiency. Experiments with an Alibaba workload trace from a 10,000-machine cluster show that Aladdin can reduce violated constraints by as mush as 20%. Meanwhile, it improves resource efficiency by 50% compared with state-of-the-art schedulers.
Heng Wu 0001, Wenbo Zhang 0006, Yuanjia Xu, Tao Huang 0001, Haiyang Ding
IPDPS1
2018 HW3C: A Heuristic based Workload Classification and Cloud Configuration Approach for Big Data Analytics
abstract
It is a big challenge to pick up the best cloud configuration for recurring big data analytics jobs running in clouds. Prior efforts may get in a sub-optimal configuration due to a broad spectrum of cloud configurations with a few test runs, such as CherryPick. We present HW3C which is a heuristic based workload classification and cloud configuration system for big data analytics jobs, our insight is classifying a job by comparing its resource preference and usage informantion with other jobs, and then using heuristic rules to distinguish bad samples from good ones in Bayesian Optimization algorithm. Our experiments on HiBench and SparkBench in Aliyun ECS show that the performance of job had been improved by 53% in average comparing with CherryPick, meanwhile the resource cost had been reduced by 40% in average.
Yuewen Wu, Heng Wu 0001, Wenbo Zhang 0006, Yuanjia Xu, Jun Wei 0001, Hua Zhong 0001
Internetware2
2018 IO dependent SSD cache allocation for elastic Hadoop applications
Wei Wang 0049, Yu Huang 0002, Heng Wu 0001, Jun Wei 0001, Tao Huang 0001
Sci. China Inf. Sci.5
2017 Application-centric SSD Cache Allocation for Hadoop Applications
abstract
Flash-based Solid State Drive (SSD) is widely used in the virtualization environment, usually as the cache of the hard disk drive-based Virtual Machine (VM) storage, to improve the IO performance. Existing SSD caching schemes are mainly driven by VM-centric metrics. They treat the VMs as independent units and focus on critical low-level performance metrics of individual VMs, such as the working set, the IO latency, or the throughput. However, for elastic Hadoop applications consisting of multiple VMs, the workload is rapidly changing, and the importance of differnet VMs may be different even if they have the same low-level IO pattern. In this situation, the VM-centric SSD caching schemes may not lead to the best performance, i.e., the shortest job completion time. Considering the importance of VMs and relationships among VMs inside the application may potentially better improve the performance, which we regard as the application-centric metrics. We propose the Application-Centric SSD caching for Hadoop applications (ACSSD), which reduces the job completion time from the application level. AC-SSD uses the genetic algorithm based approach to calculate the nearly optimal weights of virtual machines for allocating SSD cache space and controlling the I/O Operations Per Second (IOPS) based on the importance of the VMs. Moreover, AC-SSD introduces the closed-loop adaptation to face the rapidly changing workload. The evaluation shows that AC-SSD reduces the job completion time by up to 39% for IO sensitive workloads, and up to 29% for rapidly changing workloads.
Wei Wang 0049, Yu Huang 0002, Heng Wu 0001, Jun Wei 0001, Tao Huang 0001
Internetware4
2016 Determine Configuration Entry Correlations for Web Application Systems
abstract
Web application systems, comprising of heterogeneous and loosely coupled components, are usually highly-configurable due to the large number of configuration entries scattering in the components. The dependencies between components lead their entries correlate to one another, which makes the system deployment and migration daunting and error-prone. For two correlated entries, changing value of one entry requires the value change of the other. Otherwise, some implied constraints would be violated and the system failure will occur. Keeping track of entry correlations, which is essential to system reliabilities, is not a simple work as it often crosses products and requires in-depth domain knowledge. This paper proposes a method to automate the process of determining entry correlations. The method first narrows down the exploring scale to those frequently-set entries based on crawled sample data. Then, it generates a correlation score for each entry pair, which is calculated according to entry names, values and inferred types. Thirdly, a set of heuristics are provided to determine a candidate set of the likely correlations. Finally, a rank-ordered list of entry correlations is output so that system administrators can consult it to check system configuration systematically. Based on the method, we implement a tool, Correlation Explorer, and make experiments and evaluations with some real world systems. The result shows that Correlation Explorer is effective in finding a large portion of entry correlations.
Wei Chen 0018, Heng Wu 0001, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
COMPSAC2
2013 A benefit-aware on-demand provisioning approach for multi-tier applications in cloud computing
Heng Wu 0001, Wenbo Zhang 0006, Jun Wei 0001, Tao Huang 0001
Frontiers Comput. Sci.1
2012 An I/O optimizing approach for virtualization-based Internetwares
abstract
Virtualization is a very popular support environment for Internetware deployment. However, the virtual machine can access the hardware only by virtue of virtual machine monitor which can result in a big overhead, especially for I/O sensitive virtual machine. In order to reduce this kind of overhead, this paper gives a research on I/O virtualization and proposes a cache mechanism. In benefit of the cache mechanism build in virtual machine monitor, the data package switching operations would drop dramatically and the overhead is lowered too. It is proved that our method is efficient and effective in decreasing the Internetware I/O overhead in virtualization environment through the experiment.
Wenbo Zhang 0006, Heng Wu 0001
Internetware3