Wenbo Zhang 0006

dblp:31/966-6 · DBLP profile ↗
← Back
64ranked-venue papers
4as first author
22since 2021 · last 2025
0000-0002-0237-5100ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 30 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 5 since 2021Systems, architecture and hardware · 17 · 1 first-author · 10 since 2021Computer networks · 3 · 1 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Scheduling based on Block Features for Concurrent Inference with Unseen DNN Models on GPU
abstract
Efficiently scheduling concurrent deep neural network (DNN) inference on the same GPU can significantly optimize resource utilization. Such scheduling requires accurate prediction of concurrent inference time. Existing approaches primarily rely on model-level features for prediction and scheduling, which necessitates retraining and resampling when encountering unseen models to ensure prediction accuracy. However, in MLOps pipelines, the rapid iteration of models introduces numerous unseen models, making accurate predictions highly challenging and increasing the risk of SLA violations. To address these challenges posed by unseen models, we present SKADI, a scheduling framework based on block-level feature extraction and two-stage greedy scheduling. First, SKADI introduces block-level feature extraction, decomposing DNN models into homogeneous blocks (contiguous operator sequences) to enable zero-shot inference time prediction for unseen models. Second, it proposes a round-based and two-stage greedy scheduling strategy that rapidly selects optimal model pairs and overlaps their critical operators. Experimental results show that for unseen models, SKADI reduces the MAPE of concurrent inference time prediction by 59.25% and 57.55% compared to DeInfer and Abacus. Additionally, SKADI reduces SLA violation rates by 70.6% while increasing throughput by 7.3%.
Diaohan Luo, Heran Gao, Yuewen Wu, Heng Wu 0001, Wenbo Zhang 0006
ICPP7
2024 ETS: Deep Learning Training Iteration Time Prediction based on Execution Trace Sliding Window
abstract
Deep learning (DL) has become essential across various computer science domains. Accurately predicting iteration time for DL models in diverse cloud data center environments is critical for making high-quality scheduling decisions. Existing approaches neglect the sequential features inherent in the runtime execution, leading to issues such as overlooking DL framework overhead and struggling to handle diverse sizes of DL models, resulting in either low accuracy or slow convergence of the prediction model. This paper introduces ETS, a novel iteration time prediction method utilizing execution trace sliding windows. Our observation reveals that DL models exhibit a highly sequential runtime execution nature. Building upon this insight, we leverage sliding windows to extract a novel type of sequential features from the runtime execution trace. These features comprehensively capture DL framework overhead and address the diversity challenge in DL model sizes. By combining a best-practice method to train a prediction model, we achieve high accuracy and rapid convergence simultaneously. Experimental validation on over 14,000 DL model configurations demonstrates ETS's effectiveness in predicting the iteration time of DL models, achieving a mere 5.9% prediction error with a training time at the 10-minute level, and improving scheduling outcomes by reducing job completion time by 17%.
Heng Wu 0001, Yuewen Wu, Hua Zhong 0007, Wenbo Zhang 0006, Yan Liu 0102
HPDC6
2024 CSTL: Compositional Signal Temporal Logic for Adaptive Edge Service Monitoring
abstract
Edge service monitoring is essential to guarantee the healthy of service compositions at runtime. Current techniques focus mostly on the monitoring of atomic edge services, but they are inadequate for that of inter- and composite services. Besides, constraints to be monitored are usually pre-specified, although certain parameters may have to be adapted online according to the execution context. To address these challenges, this paper formulates the problem of edge service monitoring as the interpretation of temporal constraints and time-dependent QoS constraints upon intra-, inter-, and composite services. Leveraging our proposed Compositional Signal Temporal Logic (CSTL) with extended compositional modalities and online parameter settings, an adaptive monitoring mechanism is developed, where constraints are converted to CSTL formulae, and QoS variations and temporal violations are interpreted qualitatively and quantitatively at runtime. Extensive experiments are conducted upon publicly-available datasets, and evaluation results show that CSTL performs better than baseline techniques in terms of expressiveness, applicability, and robustness.
Deng Zhao, Zhangbing Zhou, Wenbo Zhang 0006, Shuiguang Deng, Xiao Xue 0001, Walid Gaaloul
IEEE Trans. Serv. Comput.3
2023 InstantChain: Enhancing Order-Execute Blockchain Systems for Latency-Sensitive Applications
Heng Wu 0001, Diaohan Luo, Heran Gao, Wenbo Zhang 0006
DASFAA (1)5
2023 SPLIT: QoS-Aware DNN Inference on Shared GPU via Evenly-Sized Model Splitting
abstract
Improving QoS by simultaneously reducing the latency violation rate and jitter in the presence of multiple deep learning inference (DLI) tasks sharing a single edge computing processor remains a challenge. However, existing DLI systems at the edge, designed to maximize throughput, face performance challenges when confronted with requests with varying QoS.
Diaohan Luo, Yuewen Wu, Heng Wu 0001, Tao Wang 0030, Wenbo Zhang 0006
ICPP6
2023 2DPChain: Orchestrating Transactions in Order-Execute Blockchain to Exploit Intra-batch and Inter-batch Parallelism
Heng Wu 0001, Heran Gao, Wenbo Zhang 0006
ICSOC (1)5
2023 Topology-Aware Self-Adaptive Resource Provisioning for Microservices
abstract
Microservice architecture is a popular technology for deploying services in cloud computing, with benefits like loose coupling, high fault tolerance, and scalability. The heterogeneous resource requirements and complex interaction relations have increased the difficulty in provisioning resources for microservices with intricacy topology. Existing approaches allocate resources for different microservices separately, and thus cannot achieve optimal global performance. Moreover, these approaches extract features from specific microservice topologies. We propose a topology-aware self-adaptive resource provisioning approach for microservices. Firstly, we propose a microservice state graph to characterize the status of each microservice in an application. Then, we use graph neural networks and attention to extract the resource requirements and correlation features of microservices. Thirdly, we use a reinforcement learning-based approach to allocate resources for microservices uniformly. Finally, we evaluate our approach by conducting a series of experiments on three typical microservice applications deployed in a heterogeneous cluster. The results show that our approach is efficient in extracting resource and correlation features of microservices, and can guarantee QoS with efficient resource utilization. Our approach can reduce the End-to-End latency by 22%, and can improve resource utilization by 18% with guaranteed latency.
Tao Wang 0030, Yuewen Wu, Heng Wu 0001, Wenbo Zhang 0006
ICWS6
2023 AgileShard: Turning the Sharded Blockchain into a Real-Time Transaction Processing System
abstract
Blockchain, as an emerging transaction processing system, suffers from low throughput and high latency. Sharded blockchains can significantly increase throughput by dividing nodes into groups (i.e., shards) to process disjoint transactions in parallel. However, the diverse latency requirements of transactions in current sharded blockchains are not well met. Three challenges prevent the sharded blockchain from becoming a real-time transaction processing system, namely, the static block size, the first-come-first-served transaction packing strategy, and the load imbalance. Therefore, this paper proposes 3 methods to help turn the sharded blockchain into a real-time transaction processing system. First, we propose an inter-shard dynamic block size negotiation method that enables shards to adaptively determine the globally optimal block size based on the deadlines of pending transactions. Then, we propose a DAG-based transaction packing method for reducing the number of deadline violations and improving parallelism. Finally, we propose a minimum-cost-flow-based shard reconfiguration method to address load imbalance. Under real datasets on Ethereum, experimental results show that AgileShard using the above three methods can effectively improve the transaction deadline satisfaction rate.
Heng Wu 0001, Heran Gao, Wenbo Zhang 0006
RTSS5
2023 DRA-MQoS: An MQoS scheduling algorithm based on resource feature matching in federated edge cloud
abstract
Summary Federated edge cloud (FEC) is an edge computing environment where servers in the same edge management domain could collaborate to handle latency‐sensitive services, thus better guaranteeing users' requirements on multiple quality of service (MQoS). Traditional scheduling methods only consider whether the server meets the resource requirements of the service, without paying attention to whether their resource characteristics match. In scenarios where server's resources are dynamically changing, this may reduce the resource utilization and the efficiency of service execution. To address this challenge, a dynamic resource adaptation‐multiple quality of service (DRA‐MQoS) algorithm is proposed for service scheduling in this environment. DRA‐MQoS could dynamically evaluate the resource characteristics of servers and services from the perspectives of “individual” and “overall” by combining the historical scheduling data of services and the utilization of different resources of server clusters. By scheduling the services to servers with the same resource characteristics for execution, the proposed policy fusion algorithm efficiently responds to the dynamically changing quality of service (QoS) demands of users by changing the weight parameters of policies. Simulation results in CloudSimSDN show that the energy consumption and execution time of DRA‐MQoS are reduced by 23% and 12%, respectively, compared with existing methods.
Yujin Li, Bo Liu 0024, Enju Wu, Jianqiang Li 0002, Zhangbing Zhou, Wenbo Zhang 0006
Concurr. Comput. Pract. Exp.6
2023 Hydra: Deadline-Aware and Efficiency-Oriented Scheduling for Deep Learning Jobs on Heterogeneous GPUs
abstract
With the rapid proliferation of deep learning (DL) jobs running on heterogeneous GPUs, scheduling DL jobs to meet various scheduling requirements, such as meeting deadlines and reducing job completion time (JCT), is critical. Unfortunately, existing efficiency-oriented and deadline-aware efforts are still rudimentary. They lack the capability of scheduling jobs to meet deadline requirements while reducing total JCT, especially when the jobs have various execution times on heterogeneous GPUs. Therefore, we present Hydra, a novel quantitative cost comparison approach, to address this scheduling issue. Here, the cost represents the total JCT plus a dynamic penalty calculated from the total tardiness (i.e., the delay time of exceeding the deadline) of all jobs. Hydra adopts a sampling approach that exploits the inherent iterative periodicity of DL jobs to estimate job execution times accurately on heterogeneous GPUs. Then, Hydra considers various combinations of job sequences and GPUs to obtain the minimized cost by leveraging an efficient branch-and-bound algorithm. Finally, the results of evaluation experiments on Alibaba traces show that Hydra can reduce total tardiness by 85.8% while reducing total JCT as much as possible, compared with state-of-the-art efforts.
Heng Wu 0001, Yuanjia Xu, Yuewen Wu, Hua Zhong 0001, Wenbo Zhang 0006
IEEE Trans. Computers6
2022 Serving unseen deep learning models with near-optimal configurations: a fast adaptive search approach
abstract
Public clouds provide a bewildering choice of configurations for Deep Learning (DL) models, and the choice of configuration will significantly impact the performance and budget. However, it is an obvious challenge to recommend a near-optimal configuration for a particular DL model from a wide range of candidates. The huge search overhead of finding such a configuration is the notorious cold start problem in state-of-the-art efforts, and this problem becomes more severe when they are faced with unseen DL models.
Yuewen Wu, Heng Wu 0001, Diaohan Luo, Yuanjia Xu, Wenbo Zhang 0006, Hua Zhong 0007
SoCC6
2022 Adaptive Auto-Scaling of Delay-Sensitive Serverless Services with Reinforcement Learning
abstract
Serverless services such as image recognition and natural language processing have strict response-time constraints. The incoming workloads and resource requirements of a newly deployed serverless service are always unpredictable due to the lack of available historical tracing data. Therefore, making effective auto-scaling decisions for these services is challenging. Open source serverless platforms often work in a best-effort manner, which cannot guarantee the response delay. Moreover, existing studies usually adopt threshold-based methods by configuring additional resource, which cannot well balance the trade-off between the quality of service and resource efficiency. To address the above issues, we propose an adaptive auto-scaling approach for delay-sensitive serverless services with reinforcement learning. First, we characterize the service's resource profile by exploring the performance improvement of different resource allocations with the reinforcement learning method. Then, we propose an adaptive auto-scaling method combining both horizontal and vertical scaling strategies based on the characterized profile to dynamically adjust the resource allocation. Finally, we select three typical services to validate our approach by comparing with two existing state-of-the-art auto-scaling methods. The experimental results show that our approach can accurately characterize services' resource profile, and effectively ensure the response delay constraints while achieving about 10.50% reduction of cost on average.
Tao Wang 0030, Wenbo Zhang 0006
COMPSAC4
2022 EOP: efficient operator partition for deep learning inference over edge servers
abstract
Recently, Deep Learning (DL) models have demonstrated great success for its attractive ability of high accuracy used in artificial intelligence Internet of Things applications. A common deployment solution is to run such DL inference tasks on edge servers. In a DL inference, each operator takes tensors as input and run in a tensor virtual machine, which isolates resource usage among operators. Nevertheless, existing edge-based DL inference approaches can not efficiently use heterogeneous resources (e.g., CPU and low-end GPU) on edge servers and result in sub-optimal DL inference performance, since they can only partition operators in a DL inference with equal or fixed ratios. It is still a big challenge to support partition optimizations over edge servers for a wide range of DL models, such as Convolution Neural Network (CNN), Recurrent Neural Network (RNN) and Transformers.
Yuanjia Xu, Heng Wu 0001, Wenbo Zhang 0006
VEE3
2022 Service Configuration Optimization in Edge-Cloud Networks Leveraging Log Analysis
abstract
The edge–cloud collaboration network is promising to support complex requirements with temporal constraints, where a requirement can be achieved through the composition of computation-demanding and delay-sensitive services. In this setting, most services should be optimally configured at the network edge, in order to decrease service response latency and reducing network resource consumption. To address this challenge, this article proposes an optimal service configuration mechanism, where temporal constraints between services are mined from event logs through our temporal interval discovery mechanism. Service configuration is formulated as a constrained multiobjective optimization problem, which is solved by our improved nondominated sorting geneticalgorithm II. Extensive experiments are conducted, and evaluation results demonstrate that our approach can find the close-to-optimal service configuration in comparison with the state-of-the-art techniques in terms of delay sensitivity and energy efficiency, especially when edge nodes can co-host a relatively large number of services.
Mengyu Sun, Zhangbing Zhou, Xiao Xue 0001, Wenbo Zhang 0006, Patrick C. K. Hung
IEEE Internet Things J.4
2022 Adaptive Configuration of Service-Based Smart Sensors in Edge Networks
abstract
Edge computing promises to facilitate the collaboration of smart sensors at the network edge, in order to satisfy the delay constraints of certain requests, and decrease the transmission of large-volume sensory data from the edge to the cloud. Generally, the functionalities provided by smart sensors are encapsulated as services, and the satisfaction of certain requests is reduced to the composition of services configured upon smart sensors in edge networks. Considering the dynamics and nonpredictability of incoming requests, an adaptive and online service configuration mechanism is essential, especially when various temporal constraints are prescribed by requests and satisfied by configured services. In this article, we formulate this problem in terms of a continuous-time Markov decision process model based on the state–action–reward mechanism. A temporal-difference learning approach is developed to optimize the service configuration while taking long-term delay sensitivity and energy efficiency into consideration. Extensive experiments are conducted, and evaluation results show that our approach outperforms the state-of-art's techniques for achieving close-to-optimal service configuration, and improving the temporal satisfaction of user requests.
Mengyu Sun, Zhangbing Zhou, Xiao Xue 0001, Wenbo Zhang 0006, Walid Gaaloul
IEEE Trans. Ind. Informatics4
2021 Talos: A Weighted Speedup-Aware Device Placement of Deep Learning Models
abstract
Efficient device placement of deep learning (DL) models, which consist of many operations, is a big challenge when heterogeneous devices (e.g., CPU, GPU) are considered. Existing average speedup and transient speedup approaches do not make full use of operation-level speedups, and the Total Operation Completion Time (TOCT) cannot be optimized efficiently.To address this challenge, we present Talos, a weighted speedup-awareness approach to optimize device placement of multiple DL models. Talos reveals operations within or across DL models have diverse speedups (from 10−1to 102) on heterogeneous devices. In addition, the execution time of operations are widely ranged (from 0.1ms to 100ms). Talos considers the two features simultaneously as weighted speedups, and treats them as costs in an incremental minimum-cost flow. Compared with state-of-the-art efforts, experiment results show that Talos can reduce TOCT by up to 50%.
Yuanjia Xu, Heng Wu 0001, Wenbo Zhang 0006, Yuewen Wu, Heran Gao, Tao Wang 0030
ASAP3
2021 Trace-based Intelligent Fault Diagnosis for Microservices with Deep Learning
abstract
Due to the scalability, fault tolerance, and high availability, distributed microservice-based applications gradually replace traditional monolithic applications as one of the main forms of Internet applications. However, current fault diagnosis methods for distributed applications have drawbacks in coarse-grained fault location and inaccurate root-cause analysis. To address the above issues, this paper proposes a trace-based intelligent fault diagnosis approach for microservices with deep learning. First, we build a request weighted directed graph and a request string to characterize the behaviors of microservices with collected historical traces. Then, we build a normal trace dataset in normal status and a faulty dataset by injecting faults, and then calculate the expected intervals of microservices’ response time and the call sequences. After that, we train the fault diagnosis model based on the deep neural network with the trace datasets to diagnose faulty microservices. Finally, we have deployed a typical open-source microservice-based application TrainTicket to validate our approach by injecting various typical faults. The results show that our approach can effectively characterize the behavior of microservices when processing requests and effectively detect faults. For fault detection, our approach achieves 91.5% accuracy in detecting faults, and has the accuracy of 85.2% in locating root causes.
Kegang Wei, Tao Wang 0030, Wenbo Zhang 0006
COMPSAC5
2021 Evaluating the Parallel Execution Schemes of Smart Contract Transactions in Different Blockchains: An Empirical Study
Chengzhi Li, Heng Wu 0001, Heran Gao, Songchang Jin, Tao Huang 0001, Wenbo Zhang 0006
ICA3PP (3)7
2021 Best VM Selection for Big Data Applications across Multiple Frameworks by Transfer Learning
abstract
Cloud providers are presented with a bewildering choice of VM types for a range of contemporary data processing frameworks today. However, existing performance modeling and machine learning efforts cannot pick optimal VM types for multiple frameworks simultaneously, since they are difficult to balance model accuracy and model training cost.
Yuewen Wu, Heng Wu 0001, Yuanjia Xu, Wenbo Zhang 0006, Hua Zhong 0007, Tao Huang 0001
ICPP5
2021 Data & Computation-Intensive Service Re-Scheduling In Edge Networks
abstract
The collaboration of Internet of Things (IoT) devices is promising nowadays to achieve complex requests in edge networks. In this setting, the functionalities of IoT devices are usually encapsulated as IoT services. A request can be fulfilled by the composition of data- or computation-intensive IoT services, which require to either consume a relatively large amount of sensory data or mandate a heavy computation capacity. Discovering functionally complementary IoT services, while satisfying their pre-specified spatial constraints, is a challenge, since certain IoT services may non-exist with respect to current IoT services deployment situation. To remedy this issue, we propose an energy-aware Data- and Computation-intensive service Migration and Scheduling mechanism (DCMS) to re-schedule certain services from their hosting devices to the ones within the geographical region prescribed by the request. Extensive experiments are conducted and evaluation results show that our DCMS is promising in reducing the energy consumption and average delay, in comparison with the state of the art's techniques.
Zhangbing Zhou, Zhuofeng Zhao, Sami Yangui, Wenbo Zhang 0006
ICWS5
2021 CTL-Based Dynamic IoT Service Composition
abstract
The collaboration of contiguous Internet of Things (IoT) devices is envisioned to satisfy complex applications which are beyond the capacity of single devices. The functionalities of IoT devices are encapsulated as IoT services, and their collaboration is implemented in terms of IoT service composition. Considering the capacity occupancy, release, and consumption caused by the implementation of IoT services, their composition is challenging in capacity-dynamically fluctuating IoT networks. This paper proposes a dynamic IoT service composition mechanism with inter-service dependencies adopted to capture the dynamic changes of IoT devices, and this change is specified by various Quality-of-Service factors. IoT service composition is formalized under Computation Tree Logic specification with certain composite structures and dynamic dependencies, and this composition is formally achieved by an optimized model checking method. Extensive experiments are conducted on publicly available datasets, and evaluation results show that our technique outperforms the state-of-the-art's approaches in relevant performance metrics.
Deng Zhao, Zhangbing Zhou, Xiao Xue 0001, Zhuofeng Zhao, Walid Gaaloul, Wenbo Zhang 0006
ICWS6
2021 Apollo: Rapidly Picking the Optimal Cloud Configurations for Big Data Analytics Using a Data-Driven Approach
Yuewen Wu, Yuanjia Xu, Heng Wu 0001, Lin-Gang Su, Wenbo Zhang 0006, Hua Zhong 0007
J. Comput. Sci. Technol.5
2020 Hermes: Efficient Cache Management for Container-based Serverless Computing
abstract
Serverless computing systems are shifting towards shorter function durations and larger degrees of parallelism to eliminate intolerable latency. For container-based serverless computing, the state-of-the-art efforts fail to ensure low latency because on-demand container images reloading from remote storage can increase the data transmission rate and downgrades system performance.
Heran Gao, Heng Wu 0001, Wenbo Zhang 0006, Tao Huang 0001
Internetware4
2020 Nuka: A Generic Engine with Millisecond Initialization for Serverless Computing
abstract
Serverless computing is becoming one of the mainstream trends in cloud computing due to its advantages of simplified programming and cost saving. However, existing serverless platforms still adopt Docker container as its execution engine, which has the cold start problem and causes high common-case invocation latency. In this work, we analyze the lifecycle of common-case serverless invocation on existing serverless platforms and find that current container startup and pulling remote images are the two main reasons causing cold start so slow. Based on the study, we implement Nuka, a generic engine with millisecond initialization for serverless computing. Nuka is fully compatible with Docker interface and can smoothly re-place Docker as the execution engine of existing serverless plat-forms. Through the isolation pool that reuses Linux's isolation configurations, Nuka avoids the high cost of container startup's scalability bottleneck, which reduces container's startup time with high concurrency scale. Nuka also avoids pulling remote images through dynamically resolving and importing required software packages from local package caching. A self-adaptive container reuse strategy dynamically controls container's pause time and replica numbers, which effectively reduces the frequency of cold start. Compared with Docker, Nuka can get a millisecond initialization with high concurrency and significantly reduces average time cost of cold startup by 6× on existing serverless platforms.
Shijun Qin, Heng Wu 0001, Yuewen Wu, Yuanjia Xu, Wenbo Zhang 0006
JCC6
2020 A framework to support multi-cloud collaboration
abstract
With the rapid development of cloud computing, major cloud providers have launched various cloud services with different functions to meet customer's needs. Therefore, flexibility is extremely important when developers use these cloud services. However, APIs of cloud services change dozens of times annually without backward compatibility. It means developers have to adapt these clouds with manual efforts. Such efforts make the multi-cloud collaboration extremely complex and cannot meet the demand of flexibility. This paper describes a configuration-based multi-cloud collaboration framework, which can support new clouds with comprehensible configurations. Meanwhile, if cloud APIs are updated without backward compatibility, it can restore services during runtime with minimized configuration. The main technologies used in this article include automatic discovery, unified abstraction, dynamic mapping and incremental update. We tested the virtual machine and container services of seven well-known cloud providers. The system can support heterogeneous clouds well. When the APIs are updated, the system can restore services in less than 200 milliseconds. At the same time, the extra cost of our framework is acceptable to cloud users.
Ting Tang, Heng Wu 0001, Yuewen Wu, Yuanjia Xu, Wenbo Zhang 0006
SERVICES7
2020 Pattern-Based Personalized Workflow Fragment Discovery
abstract
The workflow fragment discovery is essential to facilitate the reuse and repurposing of the best-practices evidenced by legacy workflows. A novel scientific experiment may be satisfied by the composition of (i) fragments that correspond to general functionalities, and (ii) other fragments that are personalized somehow, in the domain. We denote these types of fragments as basic and personalized patterns, respectively. Based on this observation, this paper proposes a novel workflow fragments discovery mechanism. Evaluation results demonstrate that this technique is more accurate in discovering personalized workflow fragments than the state of art's techniques.
Jinfeng Wen, Zhangbing Zhou, Wenbo Zhang 0006
SERVICES3
2020 Workflow-Aware Automatic Fault Diagnosis for Microservice-Based Applications With Statistics
abstract
Microservice architectures bring many benefits, e.g., faster delivery, improved scalability, and greater autonomy, so they are widely adopted to develop and operate Internet-based applications. How to effectively diagnose the faults of applications with lots of dynamic microservices has become a key to guarantee applications' performance and reliability. As a microservice performs various behaviors in different workflows of processing requests, existing approaches often cannot accurately locate the root cause of an application with interactive microservices in a dynamic deployment environment. We propose a workflow-aware automatic fault diagnosis approach for microservice-based applications with statistics. We characterize traces across microservices with calling trees, and then learn trace patterns as baselines. For the faults affecting the workflows of processing requests, we estimate the workflows' anomaly degrees, and then locate the microservices causing anomalies by comparing the difference between current traces and learned baselines with tree edit distance. For performance anomalies causing significantly increased response time, we employ principal component analysis to extract suspicious microservices with large fluctuation in response time. Finally, we evaluate our approach on three typical microservice-based applications with a series of experiments. The results show that our approach can accurately locate the microservices causing anomalies.
Tao Wang 0030, Wenbo Zhang 0006, Zeyu Gu
IEEE Trans. Netw. Serv. Manag.2
2019 IoT Service Composition for Concurrent Timed Applications
abstract
Concurrent applications may share certain components which can be conducted once for all, while mandating the satisfaction of their spatial-temporal constraints. A mechanism is proposed in this paper to identify common components, and to integrate and optimize concurrent service requests, where a component corresponds to a snippet of IoT service compositions. Consequently, composing IoT services with respect to concurrent requests can be reduced to a constrained multi-objective optimization problem, which can be solved by heuristic algorithms. Experimental results demonstrate the efficiency of this technique in comparison with the state of art's techniques, especially when the number of IoT nodes and functionality-overlapping are relatively large.
Mengyu Sun, Zhangbing Zhou, Wenbo Zhang 0006, Patrick C. K. Hung
ICWS3
2019 Aladdin: Optimized Maximum Flow Management for Shared Production Clusters
abstract
The rise in popularity of long-lived applications (LLAs), such as deep learning and latency-sensitive online Web services, has brought new challenges for cluster schedulers in shared production environments. Scheduling LLAs needs to support complex placement constraints (e.g., to run multiple containers of an application on different machines) and larger degrees of parallelism to provide global optimization. But existing schedulers usually suffer severe constraint violations, high latency and low resource efficiency. This paper describes Aladdin, a novel cluster scheduler that can maximize resource efficiency while avoiding constraint violations: (i) it proposes a multidimensional and nonlinear capacity function to support constraint expressions; (ii) it applies an optimized maximum flow algorithm to improve resource efficiency. Experiments with an Alibaba workload trace from a 10,000-machine cluster show that Aladdin can reduce violated constraints by as mush as 20%. Meanwhile, it improves resource efficiency by 50% compared with state-of-the-art schedulers.
Heng Wu 0001, Wenbo Zhang 0006, Yuanjia Xu, Tao Huang 0001, Haiyang Ding
IPDPS2
2018 HW3C: A Heuristic based Workload Classification and Cloud Configuration Approach for Big Data Analytics
abstract
It is a big challenge to pick up the best cloud configuration for recurring big data analytics jobs running in clouds. Prior efforts may get in a sub-optimal configuration due to a broad spectrum of cloud configurations with a few test runs, such as CherryPick. We present HW3C which is a heuristic based workload classification and cloud configuration system for big data analytics jobs, our insight is classifying a job by comparing its resource preference and usage informantion with other jobs, and then using heuristic rules to distinguish bad samples from good ones in Bayesian Optimization algorithm. Our experiments on HiBench and SparkBench in Aliyun ECS show that the performance of job had been improved by 53% in average comparing with CherryPick, meanwhile the resource cost had been reduced by 40% in average.
Yuewen Wu, Heng Wu 0001, Wenbo Zhang 0006, Yuanjia Xu, Jun Wei 0001, Hua Zhong 0001
Internetware3
2018 Self-adaptive cloud monitoring with online anomaly detection
Tao Wang 0030, Wenbo Zhang 0006, Zeyu Gu, Hua Zhong 0007
Future Gener. Comput. Syst.3
2017 RefCRE: A Reference Count Based Rewriting Framework for VM Image Restoration
abstract
Deduplication is widely adopted in virtual machine(VM) backup to save storage space. However, the deduplication storage could cause serious fragmentation, which severely affects the performance of restoring VMs. Current studies mainly focus on backups from a single data source, whereas the backup of VM images is usually a group of behaviors. Exploiting the block reference helps to defragment deduplication storage. In this paper, we propose RefCRE, a framework for accelerating VM image restoration. The framework implements two reference count based methods (RCR and ATL) to reduce the dispersion degree of blocks and reduce the replacement frequency. Experimental results show that the framework can efficiently reduce the dispersion degree of blocks and effectively improve the restoration performance, when the cache size is properly set.
Xiaozhao Xing, Tao Wang 0030, Zhigang Lu 0004, Wenbo Zhang 0006
COMPSAC (1)5
2017 Efficient image restoration of virtual machines with reference count based rewriting and caching
Tao Wang 0030, Xiaozhao Xing, Wenbo Zhang 0006, Hua Zhong 0007
Future Gener. Comput. Syst.4
2017 ReSeer: Efficient search-based replay for multiprocessor virtual machines
Tao Wang 0030, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001
J. Syst. Softw.3
2016 Clustering-based acceleration for virtual machine image deduplication in the cloud environment
Wenbo Zhang 0006, Tao Wang 0030, Tao Huang 0001
J. Syst. Softw.2
2016 FD4C: Automatic Fault Diagnosis Framework for Web Applications in Cloud Computing
abstract
The large-scale dynamic cloud computing environment has raised great challenges for fault diagnosis in Web applications: First, fluctuating workloads cause traditional application models to change over time; second, modeling the behaviors of complex applications usually requires domain knowledge which is difficult to obtain; third, managing large-scale applications manually is impractical for operators. To address these issues, this paper proposes an automatic fault (F) diagnosis (D) framework for (4) Web applications in cloud (C) computing (FD4C). In this paper, we propose an online incremental clustering method to recognize access behavior patterns. We also use correlation analysis to model the correlations between the workloads and application performance/resource utilization metrics in a specific access behavior pattern. FD4C detects faults by discovering the abrupt changes of correlation coefficients with control charts. Then, FD4C identifies the fault-related metrics using a feature selection method. To evaluate our proposal, we inject typical faults into TPC-W benchmark and apply FD4C to diagnose the injected faults. The experimental results show that FD4C can effectively detect the typical faults and accurately locate the metrics related to the faults.
Tao Wang 0030, Wenbo Zhang 0006, Chunyang Ye, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2015 Efficient Search-Based Automatic Execution Replay for Virtual Machines
Tao Wang 0030, Wenbo Zhang 0006, Jun Wei 0001
APSCC3
2015 VMon: Monitoring and Quantifying Virtual Machine Interference via Hardware Performance Counter
abstract
Virtualization greatly improves resource utilization in IaaS platforms, but it also introduces potential interference between virtual machines (VMs). For example, VMs may suffer from performance degradation, when they are located in one host and compete for sharing physical resources. Thus, how to efficiently monitor and quantify the VMs interference becomes a key challenge for IaaS providers. In this paper, we present Vmon, a system to transparently monitor and quantify the interference between VMs with the hardware performance counters (HPCs). By collecting the HPCs of different VMs and exploring the LLC miss rates within HPCs, Vmon analyzes the relationship between the LLC miss rates and VM performance degradation to predict the interference between different resource-intensive VMs, and mitigate the VMs interference. The experimental results show that Vmon predicts the performance degradation in the accuracy of more than 90% with less than 10% performance overhead.
Sa Wang, Wenbo Zhang 0006, Tao Wang 0030, Chunyang Ye, Tao Huang 0001
COMPSAC2
2015 Fault detection for cloud computing systems with correlation analysis
abstract
The large-scale dynamic cloud computing environment has raised great challenges for fault diagnosis in Web applications. First, fluctuating workloads cause traditional application models to change over time. Moreover, modeling the behaviors of complex applications always requires domain knowledge which is difficult to obtain. Finally, managing large-scale applications manually is impractical for operators. This paper addresses these issues and proposes an automatic fault diagnosis method for Web applications in cloud computing. We propose an online incremental clustering method to recognize access behavior patterns, and uses CCA to model the correlation between workloads and the metrics of application performance/resource utilization in a specific access behavior pattern. Our method detects anomalies by discovering the abrupt change of correlation coefficients with a EWMA control chart, and then locates suspicious metrics using a feature selection method combining ReliefF and SVM-RFE. We validate our method by injecting typical faults in TPC-W an industry-standard benchmark, and the experimental results demonstrate that it can effectively detect typical faults.
Tao Wang 0030, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001
IM2
2014 Profit-driven resource scheduling for virtualized cloud systems
abstract
Virtualized resource renting is a key issue in IaaS (Infrastructure-as-a-Service) cloud systems. Suitable resource allocation improves resource utilization and increases profit for application providers. The application provider will obtain better revenues according to the Service Level Agreement (SLA), if they rent more virtual resources. However, they will invest much more capital for renting these virtual resources. How many resources a provider rents has become a key thing for cloud applications. This paper addresses the reconciliation objectives by proposing a profit-driven resource scheduling method for virtualized cloud systems. Compared with traditional methods, our method aims at maximizing the revenues by introducing SLA and the cost of renting cloud resource, instead of increasing resource utilization or decreasing early finishing time. We model the performance of applications with queueing theory; calculate the revenues according to the SLA and renting cost; adjust the amount of virtual resources to adapt to dynamic workloads in period. We have implemented a framework for scheduling virtual resources, and applied it in our IaaS cloud platform OnceCloud. The experimental results demonstrate that our method has advantages over existing ones in revenues.
Shiyang Ye, Tao Wang 0030, Wenbo Zhang 0006, Hua Zhong 0007
ICIS3
2014 A Lightweight Virtual Machine Image Deduplication Backup Approach in Cloud Environment
abstract
As most clouds are based on virtualization technology, more and more virtual machine images are created within data centers. Depending on the need of disaster recovery, the storage space used for backup would easily sprawl to a TB or PB level with the growth of images. Unfortunately, different images have a large amount of same data segments. Those duplicated data segments will lead to serious waste of storage resource. Although there is a lot of work focus on deduplication storage and could achieve a good result in removing duplicate copies, they are not very suitable for virtual machine image deduplication in a cloud environment. Because huge resource usage of deduplication operations could lead to serious performance interference to the hosting virtual machines. This paper propose a local deduplication method which can speed up the operation progress of virtual machine image deduplication and reduce the operation time. The method is based on an improved k-means clustering algorithm, which could classify the metadata of backup image to reduce the search space of index lookup and improve the index lookup performance. Experiments show that our approach is robust and effective. It can significantly reduce the performance interference to hosting virtual machine with an acceptable increase in disk space usage.
Wenbo Zhang 0006, Shiyang Ye, Jun Wei 0001, Tao Huang 0001
COMPSAC2
2014 A class loading sensitive approach to detection of runtime type errors in component-based Java programs
Wenbo Zhang 0006, Hua Zhong 0007
Inf. Softw. Technol.1
2014 Workload-aware anomaly detection for Web applications
Tao Wang 0030, Jun Wei 0001, Wenbo Zhang 0006, Hua Zhong 0001, Tao Huang 0001
J. Syst. Softw.3
2013 VM image update notification mechanism based on pub/sub paradigm in cloud
abstract
Virtual machine image encapsulates the whole software stack including operating system, middleware, user application and other software products. Failure occurred in any layer of the software stack will be treated as image failure. However, virtual machine image with potential failures can be convert to template and spread to a wide range by means of template replication. And this paper refer to this phenomenon as "image failure propagation". Usually, patching is a widely adopted solution to resolve software failures. Nevertheless, virtual machine image patches are difficult to deliver to the final users in cloud computing environment for its openness and multi-tenancy features. This paper described image failure propagation model for the first time and proposed a promoting mechanism based on pub/sub computing paradigm to combat with the patching delivery problem.
Shiyang Ye, Wenbo Zhang 0006, Jun Wei 0001, Tao Huang 0001
Internetware3
2013 Detecting performance anomaly with correlation analysis for Internetware
Tao Wang 0030, Jun Wei 0001, Wenbo Zhang 0006, Hua Zhong 0001, Tao Huang 0001
Sci. China Inf. Sci.4
2013 A benefit-aware on-demand provisioning approach for multi-tier applications in cloud computing
Heng Wu 0001, Wenbo Zhang 0006, Jun Wei 0001, Tao Huang 0001
Frontiers Comput. Sci.2
2012 Application-Level CPU Consumption Estimation: Towards Performance Isolation of Multi-tenancy Web Applications
abstract
Performance isolation is a key requirement for application-level multi-tenant sharing hosting environments. It requires knowledge of the resource consumption of the various tenants. It is of great importance not only to be aware of the resource consumption of a tenant's given kind of transaction mix, but also to be able to be aware of the resource consumption of a given transaction type. However, direct measurement of CPU resource consumption requires instrumentation and incurs overhead. Recently, regression analysis has been applied to indirectly approximate resource consumption, but challenges still remain for cases with non-determinism and multicollinearity. In this work, we adapts Kalman filter to estimate CPU consumptions from easily observed data. We also propose techniques to deal with the non-determinism and the multicollinearity issues. Experimental results show that estimation results are in agreement with the corresponding measurements with acceptable estimation errors, especially with appropriately tuned filter settings taken into account. Experiments also demonstrate the utility of the approach in avoiding performance interference and CPU overloading.
Wei Wang 0049, Xiang Huang 0005, Xiulei Qin, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001
IEEE CLOUD4
2012 Semi-static Detection of Runtime Type Errors in Component-Based Java Programs
abstract
The using of multiple custom class loaders in component-based Java programs may lead to more runtime type errors. These errors can happen at various program statements and may be wrapped in different types of exceptions by JVM, therefore posing difficulties for dealing with them. Traditional static analysis approaches only consider static types and thus cannot detect many of them. We propose a semi-static approach based on points-to analysis and dynamically gathered behavior information of Java class loaders to detect runtime type errors in component-based Java programs without running them. We also implement a prototype tool for OSGi-based programs, where OSGi is a typical Java component framework.
Wenbo Zhang 0006
APSEC2
2012 Optimizing data migration for cloud-based key-value stores
abstract
As one database offloading strategy, elastic key-value stores are often introduced to speed up the application performance with dynamic scalability. Since the workload is varied, efficient data migration with minimal impact in service is critical for the issue of elasticity and scalability. However, due to the new virtualization technology, real-time and low-latency requirements, data migration within cloud-based key-value stores has to face new challenges: effects of VM interference, and the need to trade off between the two ingredients of migration cost, namely migration time and performance impact. To fulfill these challenges, in this paper we explore a new approach to optimize the data migration. Explicitly, we build two interference-aware models to predict the migration time and performance impact for each migration action using statistical machine learning, and then create a cost model to strike a balance between the two ingredients. Using the load rebalancing scenario as a case study, we have designed one cost-aware migration algorithm that utilizes the cost model to guide the choice of possible migration actions. Finally, we demonstrate the effectiveness of the approach using Yahoo! Cloud Serving Benchmark (YCSB).
Xiulei Qin, Wenbo Zhang 0006, Wei Wang 0049, Jun Wei 0001, Tao Huang 0001
CIKM2
2012 Towards a Cost-Aware Data Migration Approach for Key-Value Stores
abstract
Live data migration is an important technique for key-value stores. However, due to the stateful feature, new virtualization technology, stringent low latency requirements and unexpected workload changes, key-value stores deployed in cloud environment have to face new challenges for data migration: effects of VM interference, and the need to trade off between the two ingredients of migration cost, say migration time and performance impact. To address these challenges, we focus on the data migration problem in a load rebalancing scenario and build a new framework that aims to rebalance load while minimizing migration costs. We build two interference-aware prediction models to predict the migration time and performance impact for each action using statistical machine learning and then create a cost model to strike a right balance between the two ingredients of cost. A cost-aware migration algorithm is designed to utilize the cost model and balance rate to guide the choice of possible migration actions. We demonstrate the effectiveness of the data migration approach as well as the cost model and two prediction models using YCSB.
Xiulei Qin, Wenbo Zhang 0006, Wei Wang 0049, Jun Wei 0001, Tao Huang 0001
CLUSTER2
2012 Workload-Aware Online Anomaly Detection in Enterprise Applications with Local Outlier Factor
abstract
Detecting anomalies are essential for improving the reliability of enterprise applications. Current approaches set thresholds for metrics or model correlations between metrics, and anomalies are detected when the thresholds are violated or the correlations are broken. However, we have found that the dynamic workload fluctuating over multiple time scales causes system metrics and their correlations to change. Moreover, it is difficult to model various metric correlations in complex applications. This paper addresses these problems and proposes an online anomaly detection approach for enterprise applications. A method is presented for recognizing workload patterns with an incremental clustering algorithm. The Local Outlier Factor (LOF) based on the specific workload pattern is adopted for detecting anomalies. Our approach is evaluated on a testbed running the TPC-W benchmark. The experimental results show that our approach can capture workload fluctuations accurately and detect the typical faults effectively.
Tao Wang 0030, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001
COMPSAC2
2012 PaaS-Oriented Performance Modeling for Cloud Computing
abstract
PaaS is one of the most popular paradigms of cloud computing and the performance guarantee of PaaS-oriented applications has been critically concerned. Performance models, such as a Layer Queue Network (LQN) model, are efficient at performance guaranteeing in a highly dynamic computing environment for their capabilities of capacity planning. However, it is not a trivial work to build and employ such models for a PaaS platform because of the lack of designs and the transaction-intensive feature of the PaaS-oriented applications. In this paper, a PaaS-oriented performance modeling approach is proposed. The LQN model of the PaaS-oriented application can be dynamically built through tracing their interactions with the PaaS platform. And the CPU consumptions, which are important to the accuracy of the LQN model, can be refined through a Kalman-filter method. A modeling tool is implemented and experimental results have shown the effectiveness of our approach.
Wenbo Zhang 0006, Xiang Huang 0005, Ningjiang Chen, Wei Wang 0049, Hua Zhong 0007
COMPSAC1
2012 Elasticat: A load rebalancing framework for cloud-based key-value stores
abstract
The problem of load rebalancing is an important issue for cloud-based key-value stores. However, the new virtualization environment and the store's stateful feature make this classical issue more challenging. In this paper, we build a new load rebalancing framework for cloud-based key-value stores, namely ElastiCat. It can be used for auto reconfiguring the store system with minimal costs and no disruption to the availability of the service. To evaluate and minimize the rebalancing costs, we firstly build two interference-aware prediction models to predict the data migration time and performance impact for each action using statistical machine learning and then create a cost model to strike a right balance between them. A cost-aware rebalancing algorithm is designed to utilize the cost model and balance rate to create a rebalancing plan and guide the choice of possible rebalancing actions. To maintain the availability of storage service, we propose a lightweight piggy-back based data access protocol. Finally, we demonstrate the effectiveness of the framework as well as the cost model using YCSB.
Xiulei Qin, Wei Wang 0049, Wenbo Zhang 0006, Jun Wei 0001, Tao Huang 0001
HiPC3
2012 An I/O optimizing approach for virtualization-based Internetwares
abstract
Virtualization is a very popular support environment for Internetware deployment. However, the virtual machine can access the hardware only by virtue of virtual machine monitor which can result in a big overhead, especially for I/O sensitive virtual machine. In order to reduce this kind of overhead, this paper gives a research on I/O virtualization and proposes a cache mechanism. In benefit of the cache mechanism build in virtual machine monitor, the data package switching operations would drop dramatically and the overhead is lowered too. It is proved that our method is efficient and effective in decreasing the Internetware I/O overhead in virtualization environment through the experiment.
Wenbo Zhang 0006, Heng Wu 0001
Internetware2
2012 Online Anomaly Detection for Components in OSGi-based Software
Tao Wang 0030, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001
SEKE2
2011 An Adaptive Performance Modeling Approach to Performance Profiling of Multi-service Web Applications
abstract
The performance of multi-service applications are known to be determined mainly by the interactions between workload and behaviors of the application. The change of workload can lead to dynamic service demands on system resources, and even cause dynamic bottleneck switches between services inside the application. In this paper, to profiling large-applications' behaviors, and help to locate the bottleneck and optimize their capacities, we focus on modeling their behavior according to the workload. Although this topic has been well studied at testing stage, building such a model under live workload remains a challenge, because the workload and application behaviors are time-varying. To tackle this problem, we propose an adaptive approach to build and rebuild performance model according to log files. Both the user behaviors and their corresponding internal service relations are modeled, and the CPU time consumed by each service is also obtained through Kalman filter, which can "absorb" some level of noise in real-world data. Our model can explain the behaviors of both the whole application and the individual services, and provide valuable information for capacity planning and bottleneck detection. At last, our work is evaluated with TPC-W bench mark, whose results can demonstrate the effectiveness of our approach.
Xiang Huang 0005, Wei Wang 0049, Wenbo Zhang 0006, Jun Wei 0001, Tao Huang 0001
COMPSAC3
2011 On-line Cache Strategy Reconfiguration for Elastic Caching Platform: A Machine Learning Approach
abstract
Cloud computing provide scalability and high availability for web applications using such techniques as distributed caching and clustering. As one database offloading strategy, elastic caching platforms (ECPs) are introduced to speed up the performance or handle application state management with fault tolerance. Several cache strategies for ECPs have been proposed, say replicated strategy, partitioned strategy and near strategy. We first evaluate the impact of the three cache strategies using the TPC-W benchmark and find that there is no single cache strategy suitable for all conditions, the selection of the best strategy is related with workload patterns, cluster size and the number of concurrent users. This raises the question of when and how the cache strategy should be reconfigured as the condition varies which has received comparatively less attention. In this paper, we present a machine learning based approach to solving this problem. The key features of the approach are off-line training coupled with on-line system monitoring and robust synchronization process after triggering a reconfiguration, at the same time the performance model is periodically updated. More explicitly, first a rule set used to identify which cache strategy is optimal under the current condition are trained with the system statistics and performance results. We then introduce a framework to switch the cache strategy on-line as the workload varies and keep its overhead to acceptable levels. Finally, we illustrate the advantages of this approach by carrying out a set of experiments.
Xiulei Qin, Wenbo Zhang 0006, Wei Wang 0049, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
COMPSAC2
2011 A Statistical Approach for Estimating CPU Consumption in Shared Java Middleware Server
abstract
Middleware sharing is one of the important resource sharing approaches which enables sharing of costs across a large pool of users. However, the shared Java middleware server easily causes interference on performance between concurrent user requests. A key requirement to an effective performance isolation is the knowledge of the resource consumption of the various kinds of use requests classified according to different application context information. Direct measurement of resource consumption requires instrumentation which is impractical. In this paper, we demonstrate that CPU consumptions of various kinds of user requests on a given hardware can be approximated by a proposed Kalman filter based approach. Experimental results derived from testing the approach by using the TPC-W e-commerce suite deployed on a widely-used Java middleware server (Tomcat) illustrate the potential of this approach.
Wei Wang 0049, Xiang Huang 0005, Yunkui Song, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
COMPSAC4
2011 Bench4Q: A QoS-Oriented E-Commerce Benchmark
abstract
E-commerce systems are typically QoS-sensitive, so QoS-oriented tunings of e-commerce servers are very important for such systems. However, existing e-commerce benchmarks are insufficient for supporting QoS-oriented tunings, because some critical QoS features of e-commerce systems cannot be precisely evaluated by them. One example of these features is the integrality of service, which is usually expressed as a session, provided to customers. This paper presents a QoS-oriented e-commerce benchmark, which is named Bench4Q and is an extension of TPC-W supporting QoS-oriented tuning of e-commerce servers. The main features of Bench4Q include: (1) supporting session-based metrics analysis and (2) simulating QoS-sensitive load for QoS-oriented capacity analysis. We illustrate the promising benefits of these features for QoS-oriented tuning of an e-commerce server by a series of Bench4Q benchmarking on a typical e-commerce server.
Wenbo Zhang 0006, Sa Wang, Wei Wang 0049, Hua Zhong 0007
COMPSAC1
2010 A Two-Phase Approach to Subscription Subsumption Checking for Content-Based Publish/Subscribe Systems
abstract
The efficiency of subscription subsumption checking remains a key issue for content-based publish/subscribe systems. In this paper, we propose an efficient data structure called subscription subsumption graph (SSG). This data structure could differentiate the two types of subsumption relationships and help speed up the process of subsumption checking and subscription cancellation. We then present a two-phase approach to subscription subsumption checking. Phase one is mainly about checking of non-numeric constraints by using an index structure which could help filter out most of irrelevant subscriptions while phase two is about checking of remaining numeric constraints where SSG is employed. Finally, we introduce an efficient SSG-based unsubscription algorithm that could find out which subscriptions need to be forwarded without any redundant computing. We illustrate the advantages of this approach by carrying out extensive experiments.
Xiulei Qin, Jun Wei 0001, Wenbo Zhang 0006, Hua Zhong 0001, Tao Huang 0001
AINA3
2010 An adaptive fine-grained performance modeling approach for internetware
abstract
With the great success of internet technology, internetware has become one of the most important software paradigms. But the open, dynamic and uncertain network makes it difficult to guarantee the performance of internetwares. Feed forward control method has been proved to be an effective mechanism for performance guarantee in advance, but it is difficult to work well in such a dynamic environment, in which performance aspects are highly changeable because for the load fluctuation and software updates. In this paper, we proposed an adaptive performance modeling approach to adapt the environment and provide fine-grained performance guarantee. In our approach, the service invocation sequences corresponding to the load of internetware are constructed adaptively. And the service time of each service, which is the most performance parameter of our performance tool, is accurately acquired through Kalman filter.
Xiang Huang 0005, Wei Wang 0049, Wenbo Zhang 0006, Jun Wei 0001, Tao Huang 0001
Internetware3
2010 Towards PaaS using service-oriented component model
abstract
PaaS (Platform-as-a-Service) which provides execution environment for applications deployed in cloud is a key technology of cloud computing. However, to meet various requirements of different users on software platform under changing application environment, it's crucial to design and develop extensible, customizable and dynamic PaaS systems. In this paper, we analyze the software requirements of PaaS. Then a PaaS system architecture using R-OSGi as software foundation is designed. Finally, the presented approach is validated by a case study involving a web-based application E-bookstore.
Tao Wang 0030, Wenbo Zhang 0006, Jun Wei 0001
Internetware3
2009 Impacts Separation Framework for Performance Prediction of Middleware-Based Systems
abstract
Model-based performance prediction is a promising cost effective approach to analyze performance of software in early development stages. The key of this approach is to build a performance model suitably coupled with the software artifacts. But this approach is very difficult to be employed in middleware based applications because middleware bring lots of expense to the system performance. Building a performance model directly including details of middleware is somewhat impractical for these details are far from clear to application designers. This paper describes a performance impacts separation framework which separates an application into isolate aspects. Designers only need to care about the application level model design, our framework will automatically compose each aspect together to form a complete model, using an extended aspect-oriented modeling (AOM) technologies. This AOM method is designed particularly for middleware based system designs. A case study has been presented to illustrate that our framework can deal with different deployment scenarios of middleware based application easily.
Xiang Huang 0005, Wenbo Zhang 0006, Jun Wei 0001
COMPSAC (2)2
2004 Performance Tuning for Application Server OnceAS
Wenbo Zhang 0006, Beihong Jin, Ningjiang Chen, Tao Huang 0001
ISPA1