EDBT 2026 Demo / reviewers in the wild / expert
Wei Wang 0054
dblp:35/7092-54
· DBLP profile ↗
36ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0003-2262-2508ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Debugging Support for Students with Blindness and Visual Impairments on Notebook-based Programming EnvironmentsabstractA primary challenge for students with blindness or visual impairments (BVI) in programming education lies in the limited accessibility of debugging tools. While debugging is crucial, students with BVI often face significant barriers when trying to navigate to error messages, locate errors in the code, and verify fixes. This challenge is particularly problematic for Jupyter Notebook and its variants, as error messages are typically displayed in a separate block beneath the code. This separation makes it difficult for screen readers to effectively report the location and nature of errors. God'Salvation F. Oguibe, Lauryn Castro Hernandez, Katherine Holloway, Kathy B. Ewoldt, Leslie Cockerill Neely, Taslima Akter, Wei Wang 0054 |
SIGCSE (1) | 7 |
| 2024 | Improving Resource and Energy Efficiency for Cloud 3D through Excessive Rendering ReductionabstractThe rise of cloud gaming makes interactive 3D applications an emerging type of data center workload. However, the excessive rendering in current cloud 3D systems leads to large gaps between the cloud and client frame rates (FPS, frames per second), thus wasting resources and power. Although FPS regulation can remove excessive rendering, due to the highly-varying frame processing time and the use of rendering delays, existing cloud FPS regulation solutions have low FPS and slow motion-to-photon (MtP) latency, causing violations of Quality-of-Service (QoS) requirements. Jerry Lucas, Sen He 0002, Tongping Liu, Xiaoyin Wang, Wei Wang 0054 |
EuroSys | 6 |
| 2024 | Python Programming Education with Semantics-oriented Screen Reading for K-12 Students with Vision ImpairmentsabstractBecause of the high pay and high flexibility, computer science careers can be highly viable for people with blindness or vision impairments (BVI). However, in our programming education, we observed that existing screen readers used by students with BVI usually cannot properly handle computer programs, which mix English letters, digits, and punctuation marks. When applied to computer programs, current screen readers either ignore the punctuation marks, or mix English words, digits, and punctuation marks, making the screen reading either incorrect or hard to understand. The resulting difficulty in understanding program statements significantly hinders students' ability to locate incorrect code and independent coding. God'Salvation F. Oguibe, Lauryn M. Castro, Katherine Cantrell, Kathy B. Ewoldt, Leslie Cockerill Neely, Wei Wang 0054 |
SIGCSE (2) | 6 |
| 2024 | Real-time intelligent on-device monitoring of heart rate variability with PPG sensors
Jingye Xu, Yuntong Zhang 0001, Mimi Xie, Wei Wang 0054, Dakai Zhu 0001 |
J. Syst. Archit. | 4 |
| 2023 | NUMAlloc: A Faster NUMA Memory AllocatorabstractThe NUMA architecture accommodates the hardware trend of an increasing number of CPU cores. It requires the cooperation of memory allocators to achieve good performance for multithreaded applications. Unfortunately, existing allocators do not support NUMA architecture well. This paper presents a novel memory allocator – NUMAlloc, that is designed for the NUMA architecture. is centered on a binding-based memory management. On top of it, proposes an “origin-aware memory management” to ensure the locality of memory allocations and deallocations, as well as a method called “incremental sharing” to balance the performance benefits and memory overhead of using transparent huge pages. According to our extensive evaluation, NUMAlloc has the best performance among all evaluated allocators, running 15.7% faster than the second-best allocator (mimalloc), and 20.9% faster than the default Linux allocator with reasonable memory overhead. NUMAlloc is also scalable to 128 threads and is ready for deployment. Hanmei Yang, Wei Wang 0054, Sandip Kundu, Bo Wu 0002, Hui Guan 0001, Tongping Liu |
ISMM | 4 |
| 2023 | Virtual Summer Camp for High School Students with Disabilities - An Experience ReportabstractIn the past years, the authors held computer programming and machine learning summer camp for high-school students with disabilities. Due to the pandemic, the summer camp was offered virtually in 2020 and 2021. This paper reports our experience of teaching this summer camp. The main goal of the summer camp was to let students with disabilities get first-hand experience of working in STEM fields to encourage them to pursue STEM careers. The curriculum was primarily composed of hands-on activities for Python programming and Computer Vision. Besides lectures and programming tasks, there were also guest speakers and an external panel to offer their personal experiences of working in STEM fields. Wei Wang 0054, Kathy B. Ewoldt, Mimi Xie, Alberto M. Mestas-Nuñez, Sean Soderman, Jeffrey Wang |
SIGCSE (1) | 1 |
| 2023 | Online Performance Modeling and Prediction for Single-VM Applications in Multi-Tenant CloudsabstractClouds have been adopted widely by many organizations for their supports of flexible resource demands and low cost, which is normally achieved through sharing the underlying hardware among multiple cloud tenants. However, such sharing with the changes in resource contentions in virtual machines (VMs) can result in large variations for the performance of cloud applications, which makes it difficult for ordinary cloud users to estimate the run-time performance of their applications. In this article, we propose online learning methodologies for performance modeling and prediction of applications that run repetitively on multi-tenant clouds (such as on-line data analytic tasks). Here, a few micro-benchmarks are utilized to probe the in-situ perceivable performance of CPU, memory and I/O components of the target VM. Then, based on such profiling information and in-place measured application’s performance, the predictive models can be derived with either Regression or Neural-Network techniques. In particular, to address the changes in the intensity of resource contentions of a VM over time and its effects on the target application, we proposedperiodic model retrainingwhere the sliding-window technique was exploited to control the frequency and historical data used for model retraining. Moreover, aprogressive modelingapproach has been devised where the Regression and Neural-Network models are gradually updated for better adaptation to recent changes in resource contention. With 17 representative applications from PARSEC, NAS Parallel and CloudSuite benchmarks being considered, we have extensively evaluated the proposed online schemes for the prediction accuracy of the resulting models and associated overheads on both a private and public clouds. The evaluation results show that, even on the private cloud with high and radically changed resource contention, the average prediction errors of the considered models can be less than 20 percent with periodic retraining. The prediction errors generally decrease with higher retraining frequencies and more historical data points but incurring higher run-time overheads. Furthermore, with the neural-network progressive models, the average prediction errors can be reduced by about 7 percent with much reduced run-time overheads (up to 265X) on the private cloud. For public clouds with less resource contentions, the average prediction errors can be less than 4 percent for the considered models with our proposed online schemes. Hamidreza Moradi, Wei Wang 0054, Dakai Zhu 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2022 | Wasserstein Adversarial Transformer for Cloud Workload PredictionabstractPredictive VM (Virtual Machine) auto-scaling is a promising technique to optimize cloud applications’ operating costs and performance. Understanding the job arrival rate is crucial for accurately predicting future changes in cloud workloads and proactively provisioning and de-provisioning VMs for hosting the applications. However, developing a model that accurately predicts cloud workload changes is extremely challenging due to the dynamic nature of cloud workloads. Long- Short-Term-Memory (LSTM) models have been developed for cloud workload prediction. Unfortunately, the state-of-the-art LSTM model leverages recurrences to predict, which naturally adds complexity and increases the inference overhead as input sequences grow longer. To develop a cloud workload prediction model with high accuracy and low inference overhead, this work presents a novel time-series forecasting model called WGAN-gp Transformer, inspired by the Transformer network and improved Wasserstein-GANs. The proposed method adopts a Transformer network as a generator and a multi-layer perceptron as a critic. The extensive evaluations with real-world workload traces show WGAN- gp Transformer achieves 5× faster inference time with up to 5.1% higher prediction accuracy against the state-of-the-art. We also apply WGAN-gp Transformer to auto-scaling mechanisms on Google cloud platforms, and the WGAN-gp Transformer-based auto-scaling mechanism outperforms the LSTM-based mechanism by significantly reducing VM over-provisioning and under-provisioning rates. Shivani Arbat, Vinodh Kumaran Jayakumar, Wei Wang 0054, In Kee Kim |
AAAI | 4 |
| 2022 | KneeScale: Efficient Resource Scaling for Serverless Computing at the EdgeabstractServerless computing is a promising paradigm for delivering services to the Internet of Things (IoT) applications at the edge of the network. Its event-triggered computation, as well as fine-grained and agile resource scaling, is well-suited for a resource-constrained edge computing environment. However, general-purpose auto-scalers that are predominant in the cloud settings perform poorly for serverless computing at the Edge. This is mainly due to the difficulty in quickly determining the optimal resource allocation under resource-budget constraints and dynamic workloads. In this paper, we present an adaptive auto-scaler, KneeScale, that dynamically adjusts the number of replicas for serverless functions to reach a point at which the relative cost to increase resource allocation is no longer worth the corresponding performance benefit. We have designed and implemented KneeScale as lightweight system software that utilizes Kubernetes for resource management. Experimental results with a function-as-a-service (FaaS) benchmark, FunetionBeneh, and an open-source serverless computing platform, OpenFaaS, demonstrate the superior performance and resource efficiency of KneeScale. It outperforms Kubernetes Horizontal Pod AutoScaler (HPA) and OpenFaaS built-in scheduler in terms of cumulative performance under a given resource budget by up to 32 % and 106 % respectively. KneeScale achieves higher cumulative throughput than both competing techniques, lower latencies than OpenFaaS built-in scheduler, and similar latencies compared to HPA for a variety of serverless functions. Jordan Molone, Wei Wang 0054, Palden Lama |
CCGRID | 4 |
| 2022 | A Cloud 3D Dataset and Application-Specific Learned Image Compression in Cloud 3D
Sen He 0002, Vinodh Kumaran Jayakumar, Wei Wang 0054 |
ECCV (38) | 4 |
| 2022 | CloudBruno: A Low-Overhead Online Workload Prediction Framework for Cloud ComputingabstractAccurate prediction of future incoming workloads to cloud applications, such as future user request count, is critical to proactive auto-scaling, and in general, critical to the cost-effectiveness of cloud deployments. However, designing a generic predictive framework that can accurately predict for any types of workloads is difficult, especially when the workload is dynamic and can change to a pattern that has not been observed in the training data sets. However, existing workload prediction solutions typically rely on complex machine learning models, which require comprehensive training data, making it difficult for them to handle dynamic workloads. Moreover, the training of existing workload prediction solutions are also expensive in terms of both time and computing resources. This paper presents a generic and low-cost online workload prediction framework, called Cloud Bruno, which combines the more accurate LSTM models with less expensive but fast SVM models to achieve high accuracy and low training overhead. When compared to existing predictors, CloudBruno had at least 8.8 % lower error than existing deep learning-based predictors for a highly-dynamic workload that does not have comprehensive training data (i.e, has changes unknown to training data). For workloads with comprehensive training data, Cloud Bruno's error was at most 2.5 % higher than optimized deep learning-based predictors. More importantly, Cloud Bruno can effectively execute on a free cloud CPU, allowing it to be used as an online workload predictor without additional cost. Vinodh Kumaran Jayakumar, Shivani Arbat, In Kee Kim, Wei Wang 0054 |
IC2E | 4 |
| 2022 | Guaranteeing Performance SLAs of Cloud Applications Under Resource StormsabstractIn modern data centers, enterprise cloud instances run not only foreground applications like web and databases, but also different background services (e.g., backup, virus/compliance scan, batch) to manage the cloud instances securely and improve the overall resource utilization. These background services often incur resource storms that suddenly consume a lot of shared resources on cloud instances. The resource storms significantly degrade the performance of foreground applications by interfering in the preemption of the shared resources, resulting in frequent SLA violations. However, stock OS schedulers are not designed to handle these situations, and prior works are insufficient to address such resource storms under highly dynamic cloud workloads. This article presents Orchestra, a cloud-specific framework for controlling multiple applications in the user space, aiming at meeting corresponding SLAs. Orchestra takes an online approach with lightweight monitoring and performance models for both applications on the fly. It optimizes the resource allocations to meet corresponding SLAs. We evaluate the performance of Orchestra on a production cloud with a diverse range of SLAs. Orchestra guarantees the foreground application's performance SLAs at all times. At the same time, Orchestra maintains the background's performance by minimizing its performance penalty with proper allocation of the shared resources. In Kee Kim, Jinho Hwang, Wei Wang 0054, Marty Humphrey |
IEEE Trans. Cloud Comput. | 3 |
| 2022 | Forecasting Cloud Application Workloads With CloudInsight for Predictive Resource ManagementabstractPredictive cloud resource management has been widely adopted to overcome the limitations of reactive cloud autoscaling. The predictive resource management is highly relying on workload predictors, which estimate short-/long-term fluctuations of cloud application workloads. These predictors tend to be pre-optimized for specific workload patterns. However, such predictors are still insufficient to handle real-world cloud workloads whose patterns may be unknown a priori, may dynamically change over time and may be irregular. As a result, these predictors often cause over-/under-provisioning of cloud resources. To address this problem, we have created CloudInsight, a novel cloud workload prediction framework, leveraging the combined power of multiple workload predictors. CloudInsight creates an ensemble model using multiple predictors to make accurate predictions for real workloads. The weights of the predictors in CloudInsight are determined at runtime with their accuracy for the current workload using multi-class regression. The ensemble model is periodically optimized to handle sudden changes in the workload. We evaluated CloudInsight with various real workload traces. The results show that CloudInsight has 13–27 percent higher accuracy than state-of-the-art predictors. Moreover, the results from trace-based simulations with a cloud resource manager show that CloudInsight has 15–20 percent less under-/over-provisioning periods, resulting in high cost-efficiency and low SLA violations. In Kee Kim, Wei Wang 0054, Yanjun Qi, Marty Humphrey |
IEEE Trans. Cloud Comput. | 2 |
| 2021 | An Empirical Analysis of VM Startup Times in Public IaaS CloudsabstractVM startup time is an essential factor in designing elastic cloud applications. VM autoscaling can reduce the under-/over-provisioning period of VMs with a precise estimation of VM startup time, and in turn, it can guarantee the application's SLOs with improved cost-efficiency. Unfortunately, VM startup time has been little studied, and previous measurement results did not consider various configurations of VMs. This work performs a thorough analysis of VM startup times in two IaaS clouds (AWS, GCP). Specifically, we collected 300,000 data points from each provider by applying diverse VM configurations. i.e., different VM types, image sizes, location, purchase models. With extensive analysis, we found several important factors that can change VM startup time significantly. Moreover, by comparing with a previous study, we confirm that AWS made significant improvements for reducing VM startup times and currently has much quicker VM startup times than in the past. Jianwei Hao, Wei Wang 0054, In Kee Kim |
CLOUD | 3 |
| 2021 | NumaPerf: predictive NUMA profilingabstractIt is extremely challenging to achieve optimal performance of parallel applications on a NUMA architecture, which necessitates the assistance of profiling tools. However, existing NUMA-profiling tools share some similar shortcomings, such as portability, effectiveness, and helpfulness issues. This paper proposes a novel profiling tool–NumaPerf–that overcomes these issues. NumaPerf aims to identify potential performance issues for any NUMA architecture, instead of only on the current hardware. To achieve this, NumaPerf focuses on memory sharing patterns between threads, instead of real remote accesses. NumaPerf further detects potential thread migrations and load imbalance issues that could significantly affect the performance but are omitted by existing profilers. NumaPerf also identifies cache coherence issues separately that may require different fix strategies. Based on our extensive evaluation, NumaPerf can identify more performance issues than any existing tool, while fixing them leads to significant performance speedup. Hui Guan 0001, Wei Wang 0054, Xu Liu 0001, Tongping Liu |
ICS | 4 |
| 2021 | Performance Testing for Cloud Computing with Dependent Data BootstrappingabstractTo effectively utilize cloud computing, cloud practice and research require accurate knowledge of the performance of cloud applications. However, due to the random performance fluctuations, obtaining accurate performance results in the cloud is extremely difficult. To handle this random fluctuation, prior research on cloud performance testing relied on a non-parametric statistic tool called bootstrapping to design their stop criteria. However, in this paper, we show that the basic bootstrapping employed by prior work overlooks the internal dependency within cloud performance test data, which leads to inaccurate performance results.We then present Metior, a novel automated cloud performance testing methodology, which is designed based on statistical tools of block bootstrapping, the law of large numbers, and autocorrelation. These statistical tools allow Metior to properly consider the internal dependency within cloud performance test data. They also provide better coverage of cloud performance fluctuation and reduce the testing cost. Experimental evaluation on two public clouds showed that 98% of Metior’s tests could provide performance results with less than 3% error. Metior also significantly outperformed existing cloud performance testing methodologies in terms of accuracy and cost – with up to 14% increase in the accurate test count and up to 3.1 times reduction in testing cost. Sen He 0002, Palden Lama, In Kee Kim, Wei Wang 0054 |
ASE | 6 |
| 2020 | uPredict: A User-Level Profiler-Based Predictive Framework in Multi-Tenant CloudsabstractAccurate performance prediction for cloud applications is an essential component to support many cloud resource management and auto-scaling policies. However, most existing studies on performance prediction for cloud applications in multitenant clouds are at the system level and may require access to performance counters in hypervisors. In this work, we propose uPredict, a user-level profiler-based performance predictive framework for single-VM (virtual machine) applications in multitenant clouds. We designed three micro-benchmarks to assess the contention of CPUs, memory and disks in a VM, respectively. Based on the measured performance of an application and micro-benchmarks, the application and VM-specific predictive models are derived by exploiting various regression and neural network based techniques. These models can then be used to predict the application's performance using the in-situ profiled resource contention with the micro-benchmarks. We evaluated uPredict extensively with representative benchmarks from PARSEC, NAS Parallel Benchmarks and CloudSuite, on a private cloud and two public clouds. The results show that the average prediction errors are between 10.4% to 17% for various predictive models on the private cloud with high resource contention, while the errors are within 4% on public clouds. A smart load-balancing scheme powered by uPredict is presented and can effectively reduce the execution and turnaround times of the considered application by 19% and 10%, respectively. Hamidreza Moradi, Wei Wang 0054, Amanda S. Fernandez, Dakai Zhu 0001 |
IC2E | 2 |
| 2020 | A Self-Optimized Generic Workload Prediction Framework for Cloud ComputingabstractThe accurate prediction of the future workload, such as the job arrival rate and the user request rate, is critical to the efficiency of resource management and elasticity in the cloud. However, designing a generic workload predictor that works properly for various types of workload is very challenging due to the large variety of workload patterns and the dynamic changes within a workload. Because of these challenges, existing workload predictors are usually hand-tuned for specific (types of) workloads for maximum accuracy. This necessity to individually tune the predictors also makes it very difficult to reproduce the results from prior research, as the predictor designs have a strong dependency on the workloads.In this paper, we present a novel generic workload prediction framework, LoadDynamics, that can provide high accuracy predictions for any workloads. LoadDynamics employs Long-Short-Term-Memory models and can automatically optimize its internal parameters for an individual workload to achieve high prediction accuracy. We evaluated LoadDynamics with a mixture of workload traces representing public cloud applications, scientific applications, data center jobs and web applications. The evaluation results show that LoadDynamics have only 18% prediction error on average, which is at least 6.7% lower than state-of-the-art workload prediction techniques. The error of LoadDynamics was also only 1% higher than the best predictor found by exhaustive search for each workload. When applied in the Google Cloud, LoadDynamics-enabled auto-scaling policy also outperformed the state-of-the-art predictors by reducing the job turnaround time by at least 24.6% and reducing virtual machine over-provisioning by at least 4.8%. Vinodh Kumaran Jayakumar, In Kee Kim, Wei Wang 0054 |
IPDPS | 4 |
| 2020 | A Benchmarking Framework for Interactive 3D Applications in the CloudabstractWith the growing popularity of cloud gaming and cloud virtual reality (VR), interactive 3D applications have become a major class of workloads for the cloud. However, despite their growing importance, there is limited public research on how to design cloud systems to efficiently support these applications due to the lack of an open and reliable research infrastructure, including benchmarks and performance analysis tools. The challenges of generating human-like inputs under various system/application nondeterminism and dissecting the performance of complex graphics systems make it very difficult to design such an infrastructure. In this paper, we present the design of a novel research infrastructure, Pictor, for cloud 3D applications and systems. Pictor employs AI to mimic human interactions with complex 3D applications. It can also track the processing of user inputs to provide in-depth performance measurements for the complex software and hardware stack used for cloud 3D-graphics rendering. With Pictor, we designed a benchmark suite with six interactive 3D applications. Performance analyses were conducted with these benchmarks, which show that cloud system designs, including both system software and hardware designs, are crucial to the performance of cloud 3D applications. The analyses also show that energy consumption can be reduced by at least 37% when two 3D applications share a could server. To demonstrate the effectiveness of Pictor, we also implemented two optimizations to address two performance bottlenecks discovered in a state-of-the-art cloud 3D-graphics rendering system. These two optimizations improved the frame rate by 57.7% on average. Sen He 0002, Sunzhou Huang, Danny H. K. Tsang, Lingjia Tang, Jason Mars, Wei Wang 0054 |
MICRO | 7 |
| 2020 | TIMER-Cloud: Time-Sensitive VM Provisioning in Resource-Constrained CloudsabstractResource management is a vital factor for better performance in cloud systems and many resource allocation algorithms have been studied. In this work, focusing on applications with timing constraints (i.e., deadlines) running on resource-constrained clouds that have multiple heterogeneous nodes of computing resources (e.g., CPU cores and memory), we propose TIMER-Cloud, a time-sensitive resource allocation and virtual machine (VM) provisioning framework. As a key component of the framework, user requests (of running certain applications) are prioritized according to their deadlines and resource demands (in the form of VM and its operation time). Specifically, in addition to the intuitive Earliest Deadline First (EDF) ordering of requests, we propose three prioritization heuristics: a) one based on the Time-Sensitive Resource Factor (TSRF) that incorporates a request's deadline and usage efficiency of all its resources; b) the Dominant Share (DS) extension of TSRF that emphasizes the most demanded resource of a request aiming at obtaining balanced resource usage among the nodes; and c) a unified k-EDFscheme that integrates the ideas of EDF and TSRF/DS to balance the needs of meeting imminent deadlines of requests and improving resource usage efficiency. Then, for the mapping of the prioritized user requests to the heterogeneous nodes, we propose a novel request-to-node mapping algorithm based on the idea of euclidean Distance that finds the node with the best match of its resource requirements for each request. TIMER-Cloud has been implemented and validated on a cloud testbed powered by OpenStack with a few heterogeneous nodes. The proposed VM provisioning schemes are further evaluated through extensive simulations using the execution data of benchmark applications. The results show that the proposed schemes can outperform the state-of-the-art deadline oblivious scheme by serving up to 12 percent more user requests and achieving up to 8 percent more system rewards for the over-loaded scenario with 140 percent system load. Rehana Begam, Wei Wang 0054, Dakai Zhu 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2019 | A statistics-based performance testing methodology for cloud applicationsabstractThe low cost of resource ownership and flexibility have led users to increasingly port their applications to the clouds. To fully realize the cost benefits of cloud services, users usually need to reliably know the execution performance of their applications. However, due to the random performance fluctuations experienced by cloud applications, the black box nature of public clouds and the cloud usage costs, testing on clouds to acquire accurate performance results is extremely difficult. In this paper, we present a novel cloud performance testing methodology called PT4Cloud. By employing non-parametric statistical approaches of likelihood theory and the bootstrap method, PT4Cloud provides reliable stop conditions to obtain highly accurate performance distributions with confidence bands. These statistical approaches also allow users to specify intuitive accuracy goals and easily trade between accuracy and testing cost. We evaluated PT4Cloud with 33 benchmark configurations on Amazon Web Service and Chameleon clouds. When compared with performance data obtained from extensive performance tests, PT4Cloud provides testing results with 95.4% accuracy on average while reducing the number of test runs by 62%. We also propose two test execution reduction techniques for PT4Cloud, which can reduce the number of test runs by 90.1% while retaining an average accuracy of 91%. We compared our technique to three other techniques and found that our results are much more accurate. Sen He 0002, Glenna Manns, John Saunders, Wei Wang 0054, Lori L. Pollock, Mary Lou Soffa |
ESEC/SIGSOFT FSE | 4 |
| 2018 | Flexible VM Provisioning for Time-Sensitive Applications with Multiple Execution OptionsabstractSeveral recent studies have investigated the virtual machine (VM) provisioning problem for requests with time constraints (deadlines) in cloud systems. These studies typically assumed that a request is associated with a single execution time when running on VMs with a given resource demand. In this paper, we consider modern applications that are normally implemented with generic frameworks that allow them to execute with various numbers of threads on VMs with different resource demands. For such applications, it is possible for the users to specify multiple execution options (MEOs) for a request where each execution option is represented by a certain number of VMs with some resources to run the application and its corresponding execution time. We investigate the problem of virtual machine provisioning for such time-sensitive requests with MEOs in resource-constrained clouds. By incorporating the MEOs of requests, we propose several novel and flexible VM provisioning schemes that carefully balance resource usage efficiency, input workloads and request deadlines with the objective of achieving higher resource utilization and system benefits. We evaluated the proposed MEO-aware schemes on various workloads with both benchmark requests and synthetic requests. The results show that our MEO-aware algorithms outperform the state-of-the-art schemes that consider only a single execution option of requests by serving up to 38% more requests and achieving up to 27% more benefits. Rehana Begam, Hamidreza Moradi, Wei Wang 0054, Dakai Zhu 0001 |
IEEE CLOUD | 3 |
| 2018 | CloudInsight: Utilizing a Council of Experts to Predict Future Cloud Application WorkloadsabstractMany predictive approaches have been proposed to overcome the limitations of reactive autoscaling on clouds. These approaches leverage workload predictors that are usually targeted for a particular workload pattern and can fail to handle real-world cloud workloads whose patterns may be unknown a priori, may dynamically change over time, or may be irregular. The result is that resources are frequently under-and overprovisioned. To address this problem, we create a novel cloud workload prediction framework called CloudInsight, leveraging the combined power of multiple workload predictors that collectively provide a "council of experts". The weights of the predictors in this ensemble model are determined in real-time based on their accuracy for current workload using multi-class regression. Under real workload traces, CloudInsight has 13% - 27% better accuracy than state-of-the-art predictors. It also has low overhead for predicting future workload changes (<; 100 ms) and creating a new ensemble workload predictor (<; 1.1 sec.). In Kee Kim, Wei Wang 0054, Yanjun Qi, Marty Humphrey |
IEEE CLOUD | 2 |
| 2018 | Testing Cloud Applications under Cloud-Uncertainty Performance EffectsabstractThe paradigm shift of deploying applications to the cloud has introduced both opportunities and challenges. Although clouds use elasticity to scale resource usage at runtime to help meet an application's performance requirements, developers are still challenged by unpredictable performance, little control of execution environment, and differences among cloud service providers, all while being charged for their cloud usages. Application performance stability is particularly affected by multi-tenancy in which the hardware is shared among varying applications and virtual machines. Developers porting their applications need to meet performance requirements, but testing on the cloud under the effects of performance uncertainty is difficult and expensive, due to high cloud usage costs. This paper presents a first approach to testing an application with typical inputs for how its performance will be affected by performance uncertainty, without incurring undue costs of brute force testing in the cloud. We specify cloud uncertainty testing criteria, design a test-based strategy to characterize the black box cloud's performance distributions using these testing criteria, and support execution of tests to characterize the resource usage and cloud baseline performance of the application to be deployed. Importantly, we developed a smart test oracle that estimates the application's performance with certain confidence levels using the above characterization test results and determines whether it will meet its performance requirements. We evaluated our testing approach on both the Chameleon cloud and Amazon web services; results indicate that this testing strategy shows promise as a cost-effective approach to test for performance effects of cloud uncertainty when porting an application to the cloud. Wei Wang 0054, Ningjing Tian, Sunzhou Huang, Sen He 0002, Abhijeet Srivastava, Mary Lou Soffa, Lori L. Pollock |
ICST | 1 |
| 2018 | Orchestra: Guaranteeing Performance SLAs for Cloud Applications by Avoiding Resource StormsabstractThis paper presents Orchestra, a cloud-specific framework for managing both foreground applications (e.g., Web, DBMS) and background services (e.g., backup, security check, batch jobs) in the user space. Orchestra is designed to address "resource storms" caused by sudden executions of the background services on the cloud instances. The resource storms significantly degrade the performance of foreground applications by interfering in the preemption of the shared resources, resulting in frequent SLA violations and poor user experience. Orchestra takes an online approach using lightweight monitoring and creates performance models for multiple cloud applications on the fly. It then optimizes the allocations of shared resources to meet SLAs. We evaluate the performance of Orchestra on a production cloud (Amazon EC2) with a diverse range of SLA requirements. The experiment results show that Orchestra successfully guarantees the foreground application's performance to meet its SLA targets at all times. Moreover, Orchestra maintains the background's performance by minimizing its performance penalty with proper allocation of the shared resources. In Kee Kim, Jinho Hwang, Wei Wang 0054, Marty Humphrey |
ISPDC | 3 |
| 2018 | iReplayer: in-situ and identical record-and-replay for multithreaded applicationsabstractReproducing executions of multithreaded programs is very challenging due to many intrinsic and external non-deterministic factors. Existing RnR systems achieve significant progress in terms of performance overhead, but none targets the in-situ setting, in which replay occurs within the same process as the recording process. Also, most existing work cannot achieve identical replay, which may prevent the reproduction of some errors. Hongyu Liu 0005, Sam Silvestro, Wei Wang 0054, Chen Tian 0002, Tongping Liu |
PLDI | 3 |
| 2016 | Empirical Evaluation of Workload Forecasting Techniques for Predictive Cloud Resource ScalingabstractMany predictive resource scaling approaches have been proposed to overcome the limitations of the conventional reactive approaches most often used in clouds today. In general, due to the complexity of clouds, these reactive approaches were often forced to make significant limiting assumptions in either the operating conditions/requirements or expected workload patterns. As such, it is extremely difficult for cloud users to know which - if any - existing workload predictor will work best for their particular cloud activity, especially when considering highly-variable workload patterns, non-trivial billing models, variety of resources to add/subtract, etc. To solve this problem, we conduct comprehensive evaluations for a variety of workload predictors under real-world cloud configurations. The workload predictors cover four classes of 21 predictors: naive, regression, temporal, and non-temporal methods. We simulate a cloud application under four realistic workload patterns, two different cloud billing models, and three different styles of predictive scaling. Our evaluation confirms that no workload predictor is universally best for all workload patterns, and shows that Predictive Scaling-out + Predictive Scaling-in has the best cost efficiency and the lowest job deadline miss rate in cloud resource management, on average providing 30% better cost efficiency and 80% less job deadline miss rate compared to other styles of predictive scaling. In Kee Kim, Wei Wang 0054, Yanjun Qi, Marty Humphrey |
CLOUD | 2 |
| 2016 | Predicting the memory bandwidth and optimal core allocations for multi-threaded applications on large-scale NUMA machinesabstractModern NUMA platforms offer large numbers of cores to boost performance through parallelism and multi-threading. However, because performance scalability is limited by available memory bandwidth, the strategy of allocating all cores can result in degraded performance. Consequently, accurately predicting optimal (best performing) core allocations, and executing applications with these allocations are crucial for achieving the best performance. Previous research focused on the prediction of optimal numbers of cores. However, in this paper, we show that, because of the asymmetric NUMA memory configuration and the asymmetric application memory behavior, optimal core allocations are not merely optimal numbers of cores. Additionally, previous studies do not adequately consider NUMA memory resources, which further limits their ability to accurately predict optimal core allocations. In this paper, we present a model, NuCore, which predicts both memory bandwidth usage and optimal core allocations. NuCore considers various memory resources and NUMA asymmetry, and employs Integer Programming to achieve high accuracy and low overhead. Experimental results from real NUMA machines show that the core allocations predicted by NuCore provide 1.27x average speedup over using all cores with only 75.6% cores allocated. NuCore also provides 1.18x and 1.21x average speedups over two state-of-the-art techniques. Our results also show that NuCore faithfully models NUMA memory systems and predicts memory bandwidth usages with only 10% average error. Wei Wang 0054, Jack W. Davidson, Mary Lou Soffa |
HPCA | 1 |
| 2015 | PICS: A Public IaaS Cloud SimulatorabstractPublic clouds become essential for many organizations to run their applications because they provide huge financial benefits and great flexibility. However, it is very challenging to accurately evaluate the performance and cost of applications without actual deployment on the clouds. Existing cloud simulators are generally designed from the perspective of cloud service providers, thus they can be under-developed for answering questions for the perspective of cloud users. To solve this prediction and evaluation problem, we created a Public Cloud IaaS Simulator (PICS). PICS enables the cloud user to evaluate the cost and performance of public IaaS clouds along with such dimensions like VM and storage service, resource scaling, job scheduling, and diverse workload patterns. We extensively validated PICS by comparing its results with the data acquired from real public IaaS cloud using real cloud-applications. We show that PICS provides highly accurate simulation results (less than 5% of average errors) under a variety of use cases. Moreover, we evaluated PICS' sensitivity with imprecise simulation parameters. The results show that PICS still provides very reliable simulation results with imprecise simulation parameters and performance uncertainty. In Kee Kim, Wei Wang 0054, Marty Humphrey |
CLOUD | 2 |
| 2014 | DraMon: Predicting memory bandwidth usage of multi-threaded programs with high accuracy and low overheadabstractMemory bandwidth severely limits the scalability and performance of today's multi-core systems. Because of this limitation, many studies that focused on improving multi-core scalability rely on bandwidth usage predictions to achieve the best results. However, existing bandwidth prediction models have low accuracy, causing these studies to have inaccurate conclusions or perform sub-optimally. Most of these models make predictions based on the bandwidth usage samples of a few trial runs. Many factors that affect bandwidth usage and the complex DRAM operations are overlooked. This paper presents DraMon, a model that predicts bandwidth usages for multi-threaded programs with low overhead. It achieves high accuracy through highly accurate predictions of DRAM contention and DRAM concurrency, as well as by considering a wide range of hardware and software factors that impact bandwidth usage. We implemented two versions of DraMon: DraMon-T, a memory-trace based model, and DraMon-R, a run-time model which uses hardware performance counters. When evaluated on a real machine with memory-intensive benchmarks, DraMon-T has average accuracies of 99.17% and 94.70% for DRAM contention predictions and bandwidth predictions, respectively. DraMon-R has average accuracies of 98.55% and 93.37% for DRAM contention and bandwidth predictions respectively, with only 0.50% overhead on average. Wei Wang 0054, Tanima Dey, Jack W. Davidson, Mary Lou Soffa |
HPCA | 1 |
| 2013 | ReQoS: reactive static/dynamic compilation for QoS in warehouse scale computersabstractAs multicore processors with expanding core counts continue to dominate the server market, the overall utilization of the class of datacenters known as warehouse scale computers (WSCs) depends heavily on colocation of multiple workloads on each server to take advantage of the computational power provided by modern processors. However, many of the applications running in WSCs, such as websearch, are user-facing and have quality of service (QoS) requirements. When multiple applications are co-located on a multicore machine, contention for shared memory resources threatens application QoS as severe cross-core performance interference may occur. WSC operators are left with two options: either disregard QoS to maximize WSC utilization, or disallow the co-location of high-priority user-facing applications with other applications, resulting in low machine utilization and millions of dollars wasted. Lingjia Tang, Jason Mars, Wei Wang 0054, Tanima Dey, Mary Lou Soffa |
ASPLOS | 3 |
| 2013 | ReSense: Mapping dynamic workloads of colocated multithreaded applications using resource sensitivityabstractTo utilize the full potential of modern chip multiprocessors and obtain scalable performance improvements, it is critical to mitigate resource contention created by multithreaded workloads. In this article, we describe ReSense, the first runtime system that uses application characteristics to dynamically map multithreaded applications from dynamic workloads—workloads where multithreaded applications arrive, execute, and terminate continuously in unpredictable ways. ReSense mitigates contention for the shared resources in the memory hierarchy by applying a novel thread-mapping algorithm that dynamically adjusts the mapping of threads from dynamic workloads using a precalculated sensitivity score. The sensitivity score quantifies an application's sensitivity to sharing a particular memory resource and is calculated by an efficient characterization process that involves running the multithreaded application by itself on the target platform. To measure ReSense's effectiveness, sensitivity scores were determined for 21 benchmarks from PARSEC-2.1 and NPB-OMP-3.3 for the shared resources in the memory hierarchy on four different platforms. Using three different-sized dynamic workloads composed of randomly selected two, four, and eight corunning benchmarks with randomly selected start times, ReSense was able to improve the average response time of the three workloads by up to 27.03%, 20.89%, and 29.34% and throughput by up to 19.97%, 46.56%, and 29.86%, respectively, over the native OS on real hardware. By estimating and comparing ReSense's effectiveness with the optimal thread mapping for two different workloads, we found that the maximum average difference with the experimentally determined optimal performance was 1.49% for average response time and 2.08% for throughput. Tanima Dey, Wei Wang 0054, Jack W. Davidson, Mary Lou Soffa |
ACM Trans. Archit. Code Optim. | 2 |
| 2012 | Performance analysis of thread mappings with a holistic view of the hardware resourcesabstractWith the shift to chip multiprocessors, managing shared resources has become a critical issue in realizing their full potential. Previous research has shown that thread mapping is a powerful tool for resource management. However, the difficulty of simultaneously managing multiple hardware resources and the varying nature of the workloads have impeded the efficiency of thread mapping algorithms. To overcome the difficulties of simultaneously managing multiple resources with thread mapping, the interaction between various microarchitectural resources and thread characteristics must be well understood. This paper presents an in-depth analysis of PARSEC benchmarks running under different thread mappings to investigate the interaction of various thread mappings with microarchitectural resources including, L1 I/D-caches, I/D TLBs, L2 caches, hardware prefetchers, off-chip memory interconnects, branch predictors, memory disambiguation units and the cores. For each resource, the analysis provides guidelines for how to improve its utilization when mapping threads with different characteristics. We also analyze how the relative importance of the resources varies depending on the workloads. Our experiments show that when only memory resources are considered, thread mapping improves an application's performance by as much as 14% over the default Linux scheduler. In contrast, when both memory and processor resources are considered the mapping algorithm achieves performance improvements by as much as 28%. Additionally, we demonstrate that thread mapping should consider L2 caches, prefetchers and off-chip memory interconnects as one resource, and we present a new metric called L2-misses-memory-latency-product (L2MP) for evaluating their aggregated performance impact. Wei Wang 0054, Tanima Dey, Jason Mars, Lingjia Tang, Jack W. Davidson, Mary Lou Soffa |
ISPASS | 1 |
| 2012 | REEact: a customizable virtual execution manager for multicore platformsabstractWith the shift to many-core chip multiprocessors (CMPs), a critical issue is how to effectively coordinate and manage the execution of applications and hardware resources to overcome performance, power consumption, and reliability challenges stemming from hardware and application variations inherent in this new computing environment. Effective resource and application management on CMPs requires consideration of user/application/hardware-specific requirements and dynamic adaption of management decisions based on the actual run-time environment. However, designing an algorithm to manage resources and applications that can dynamically adapt based on the run-time environment is difficult because most resource and application management and monitoring facilities are only available at the operating system level. This paper presents REEact, an infrastructure that provides the capability to specify user-level management policies with dynamic adaptation. REEact is a virtual execution environment that provides a framework and core services to quickly enable the design of custom management policies for dynamically managing resources and applications. To demonstrate the capabilities and usefulness of REEact, this paper describes three case studies--each illustrating the use of REEact to apply a specific dynamic management policy on a real CMP. Through these case studies, we demonstrate that REEact can effectively and efficiently implement policies to dynamically manage resources and adapt application execution. Wei Wang 0054, Tanima Dey, Ryan W. Moore, Mahmut Aktasoglu, Bruce R. Childers, Jack W. Davidson, Mary Jane Irwin, Mahmut T. Kandemir, Mary Lou Soffa |
VEE | 1 |
| 2011 | Characterizing multi-threaded applications based on shared-resource contentionabstractFor higher processing and computing power, chip multiprocessors (CMPs) have become the new mainstream architecture. This shift to CMPs has created many challenges for fully utilizing the power of multiple execution cores. One of these challenges is managing contention for shared resources. Most of the recent research address contention for shared resources by single-threaded applications. However, as CMPs scale up to many cores, the trend of application design has shifted towards multi-threaded programming and new parallel models to fully utilize the underlying hardware. There are differences between how single- and multi-threaded applications contend for shared resources. Therefore, to develop approaches to reduce shared resource contention for emerging multi-threaded applications, it is crucial to understand how their performances are affected by contention for a particular shared resource. In this research, we propose and evaluate a general methodology for characterizing multi-threaded applications by determining the effect of shared-resource contention on performance. To demonstrate the methodology, we characterize the applications in the widely used PARSEC benchmark suite for shared-memory resource contention. The characterization reveals several interesting aspects of the benchmark suite. Three of twelve PARSEC benchmarks exhibit no contention for cache resources. Nine of the benchmarks exhibit contention for the L2-cache. Of these nine, only three exhibit contention between their own threads-most contention is because of competition with a co-runner. Interestingly, contention for the Front Side Bus is a major factor with all but two of the benchmarks and degrades performance by more than 11%. Tanima Dey, Wei Wang 0054, Jack W. Davidson, Mary Lou Soffa |
ISPASS | 2 |
| 1997 | Multispace Search for Minimizing the Maximum Nodal DegreeabstractHajek and Sasaki (1988) showed that, for continuous traffic and packet radio network, the selection of paths that minimize the maximum nodal degree generates schedules of minimum-length. This result suggests that minimization of the maximum nodal degree provides good (although not necessarily optimal) performance in slotted networks with fixed-length packets. We give a multispace search algorithm that interplays structural operations in conjunction with a local search algorithm for the minimization of the maximum nodal degree. Structural operations disturb the environment of forming local minima, which makes multispace search a very natural approach to the problem. Experimental results indicate that this method has improved local search in terms of the solution quality and its sensitivity to the initial random assignment. Wei Wang 0054, Danny H. K. Tsang |
ICCCN | 3 |