VLDB 2026 Research / reviewers in the wild / expert
David McKee 0001
dblp:151/3386 · also David Wesley McKee
· DBLP profile ↗
10ranked-venue papers
4as first author
0since 2021 · last 2019
0000-0002-9047-7990ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Cloud and datacenter computing · 31% Distributed systems · 31% Performance modeling and evaluation · 26% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
anomaly detection |
0.4 | 1 | 2019 | Straggler Root-Cause and Impact Analysis for Massive-scale Virtualized Cloud Datacenters · IEEE Trans. Serv. Comput. 2019 |
Distributed systems
root cause analysis |
0.4 | 1 | 2019 | Straggler Root-Cause and Impact Analysis for Massive-scale Virtualized Cloud Datacenters · IEEE Trans. Serv. Comput. 2019 |
Cloud and datacenter computing
straggler detection |
0.4 | 1 | 2019 | Straggler Root-Cause and Impact Analysis for Massive-scale Virtualized Cloud Datacenters · IEEE Trans. Serv. Comput. 2019 |
Performance modeling and evaluation
workload characterization |
0.4 | 1 | 2019 | Straggler Root-Cause and Impact Analysis for Massive-scale Virtualized Cloud Datacenters · IEEE Trans. Serv. Comput. 2019 |
Embedded and real-time systems
cyber-physical systems |
0.2 | 1 | 2016 | SEED: A Scalable Approach for Cyber-Physical System Simulation · IEEE Trans. Serv. Comput. 2016 |
Performance modeling and evaluation › simulation › parallel and distributed simulation
distributed simulation |
0.2 | 1 | 2016 | SEED: A Scalable Approach for Cyber-Physical System Simulation · IEEE Trans. Serv. Comput. 2016 |
High-performance computing
large-scale simulation |
0.1 | 1 | 2016 | SEED: A Scalable Approach for Cyber-Physical System Simulation · IEEE Trans. Serv. Comput. 2016 |
Methods — techniques the papers use, named apart from their topics
online analytic agents · 0.4offline execution pattern modeling · 0.4event messaging · 0.2automated simulation partitioning · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Straggler Root-Cause and Impact Analysis for Massive-scale Virtualized Cloud DatacentersabstractIncreased complexity and scale of virtualized distributed systems has resulted in the manifestation of emergent phenomena substantially affecting overall system performance. This phenomena is known as “Long Tail”, whereby a small proportion of task stragglers significantly impede job completion time. While work focuses on straggler detection and mitigation, there is limited work that empirically studies straggler root-cause and quantifies its impact upon system operation. Such analysis is critical to ascertain in-depth knowledge of straggler occurrence for focusing developmental and research efforts towards solving the Long Tail challenge. This paper provides an empirical analysis of straggler root-cause within virtualized Cloud datacenters; we analyze two large-scale production systems to quantify the frequency and impact stragglers impose, and propose a method for conducting root-cause analysis. Results demonstrate approximately 5 percent of task stragglers impact 50 percent of total jobs for batch processes, and 53 percent of stragglers occur due to high server resource utilization. We leverage these findings to propose a method for extreme straggler detection through a combination of offline execution patterns modeling and online analytic agents to monitor tasks at runtime. Experiments show the approach is capable of detecting stragglers less than 11 percent into their execution lifecycle with 95 percent accuracy for short duration jobs. Peter Garraghan, Xue Ouyang 0003, Renyu Yang, David McKee 0001, Jie Xu 0007 |
IEEE Trans. Serv. Comput. | 4 |
| 2018 | Adaptive Speculation for Efficient Internetware Application Execution in CloudsabstractModern Cloud computing systems are massive in scale, featuring environments that can execute highly dynamic Internetware applications with huge numbers of interacting tasks. This has led to a substantial challenge—the straggler problem, whereby a small subset of slow tasks significantly impede parallel job completion. This problem results in longer service responses, degraded system performance, and late timing failures that can easily threaten Quality of Service (QoS) compliance. Speculative execution (or speculation) is the prominent method deployed in Clouds to tolerate stragglers by creating task replicas at runtime. The method detects stragglers by specifying a predefined threshold to calculate the difference between individual tasks and the average task progression within a job. However, such a static threshold debilitates speculation effectiveness as it fails to capture the intrinsic diversity of timing constraints in Internetware applications, as well as dynamic environmental factors, such as resource utilization. By considering such characteristics, different levels of strictness for replica creation can be imposed to adaptively achieve specified levels of QoS for different applications. In this article, we present an algorithm to improve the execution efficiency of Internetware applications by dynamically calculating the straggler threshold, considering key parameters including job QoS timing constraints, task execution progress, and optimal system resource utilization. We implement this dynamic straggler threshold into the YARN architecture to evaluate it’s effectiveness against existing state-of-the-art solutions. Results demonstrate that the proposed approach is capable of reducing parallel job response time by up to 20% compared to the static threshold, as well as a higher speculation success rate, achieving up to 66.67% against 16.67% in comparison to the static method. Xue Ouyang 0003, Peter Garraghan, Bernhard Primas, David McKee 0001, Paul Townend, Jie Xu 0007 |
ACM Trans. Internet Techn. | 4 |
| 2017 | A Framework and Task Allocation Analysis for Infrastructure Independent Energy-Efficient Scheduling in Cloud Data CentersabstractCloud computing represents a paradigm shift in provisioning on-demand computational resources underpinned by data center infrastructure, which now constitutes 1.5% of worldwide energy consumption. Such consumption is not merely limited to operating IT devices, but encompasses cooling systems representing 40% total data center energy usage. Given the substantive complexity and heterogeneity of data center operation spanning both computing and cooling components, obtaining analytical models for optimizing data center energy-efficiency is an inherently difficult challenge. Specifically, difficulties arise pertaining to the non-intuitive relationship between computing and cooling energy in the data center, computationally complex energy modeling, as well as cooling models restricted to a specific class of data center facility geometry - all of which arise from the interdisciplinary nature of this research domain.In this paper we propose a framework for energy-efficient scheduling to alleviate these challenges. It is applicable to any type of data center infrastructure and does not require complex modeling of energy.Instead, the concept of a target workload distribution is proposed. If the workload is assigned to nodes according to the target workload distribution, then the energy consumption is minimized. The exact target workload distribution is unknown, but an approximated distribution is delivered by the framework. The scheduling objective is to assign workload to nodes such that the workload distribution becomes as similar as possible to the target distribution in order to reduce energy consumption.Several mathematically sound algorithms have been designed to address this novel type of scheduling problem. Simulation results demonstrate that our algorithms reduce the relative deviation by at least 16.9% and the relative variance by at least 22.67% in comparison to (asymmetric) load balancing algorithms. Bernhard Primas, Peter Garraghan, David McKee 0001, Jon Summers, Jie Xu 0007 |
CloudCom | 3 |
| 2017 | Massive-Scale Automation in Cyber-Physical Systems: Vision & ChallengesabstractThe next era of computing is the evolution of the Internet of Things (IoT) and Smart Cities with development of the Internet of Simulation (IoS). The existing technologies of Cloud, Edge, and Fog computing as well as HPC being applied to the domains of Big Data and deep learning are not adequate to handle the scale and complexity of the systems required to facilitate a fully integrated and automated smart city. This integration of existing systems will create an explosion of data streams at a scale not yet experienced. The additional data can be combined with simulations as services (SIMaaS) to provide a shared model of reality across all integrated systems, things, devices, and individuals within the city. There are also numerous challenges in managing the security and safety of the integrated systems. This paper presents an overview of the existing state-of-the-art in automating, augmenting, and integrating systems across the domains of smart cities, autonomous vehicles, energy efficiency, smart manufacturing in Industry 4.0, and healthcare. Additionally the key challenges relating to Big Data, a model of reality, augmentation of systems, computation, and security are examined. David McKee 0001, Stephen J. Clement, Jaber Almutairi, Jie Xu 0007 |
ISADS | 1 |
| 2016 | Straggler Detection in Parallel Computing Systems through Dynamic Threshold CalculationabstractCloud computing systems face the substantial challenge of the Long Tail problem: a small subset of straggling tasks significantly impede parallel jobs completion. This behavior results in longer service response times and degraded system utilization. Speculative execution, which create task replicas at runtime, is a typical method deployed in large-scale distributed systems to tolerate stragglers. This approach defines stragglers by specifying a static threshold value, which calculates the temporal difference between an individual task and the average task progression for a job. However, specifying static threshold debilitates speculation effectiveness as it fails to consider the intrinsic diversity of job timing constraints within modern day Cloud computing systems. Capturing such heterogeneity enables the ability to impose different levels of strictness for replica creation while achieving specified levels of QoS for different application types. Furthermore, a static threshold also fails to consider system environmental constraints in terms of replication overheads and optimal system resource usage. In this paper we present an algorithm for dynamically calculating a threshold value to identify task stragglers, considering key parameters including job QoS timing constraints, task execution characteristics, and optimal system resource utilization. We study and demonstrate the effectiveness of our algorithm through simulating a number of different operational scenarios based on real production cluster data against state-of-the-art solutions. Results demonstrate that our approach is capable of creating 58.62% less replicas under high resource utilization while reducing response time up to 17.86% for idle periods compared to a static threshold. Xue Ouyang 0003, Peter Garraghan, David McKee 0001, Paul Townend, Jie Xu 0007 |
AINA | 3 |
| 2016 | Tolerating Transient Late-Timing Faults in Cloud-Based Real-Time Stream ProcessingabstractReal-time stream processing is a frequently deployed application within Cloud datacenters that is required to provision high levels of performance and reliability. Numerous fault-tolerant approaches have been proposed to effectively achieve this objective in the presence of crash failures. However, such systems struggle with transient late-timing faults - a fault classification challenging to effectively tolerate - that manifests increasingly within large-scale distributed systems. Such faults represent a significant threat towards minimizing soft real-time execution of streaming applications in the presence of failures. This work proposes a fault-tolerant approach for QoS-aware data prediction to tolerate transient late-timing faults. The approach is capable of determining the most effective data prediction algorithm for imposed QoS constraints on a failed stream processor at run-time. We integrated our approach into Apache Storm with experiment results showing its ability to minimize stream processor end-to-end execution time by 61% compared to other fault-tolerant approaches. The approach incurs 12% additional CPU utilization while reducing network usage by 44%. Peter Garraghan, Stuart Perks, Xue Ouyang 0003, David McKee 0001, Ismael Solís Moreno |
ISORC | 4 |
| 2016 | SEED: A Scalable Approach for Cyber-Physical System SimulationabstractSimulation is critical when studying real operational behavior of increasingly complex Cyber-Physical Systems, forecasting future behavior, and experimenting with hypothetical scenarios. A critical aspect of simulation is the ability to evaluate large-scale systems within a reasonable time frame while modeling complex interactions between millions of components. However, modern simulations face limitations in provisioning this functionality for CPSs in terms of balancing simulation complexity with performance, resulting in substantial operational costs required for completing simulation execution. Moreover, users are required to have expertise in modeling and configuring simulations to infrastructure which is time consuming. In this paper we present Simulation EnvironmEnt Distributor (SEED), a novel approach for simulating large-scale CPSs across a loosely-coupled distributed system requiring minimal user configuration. This is achieved through automated simulation partitioning and instantiation while enforcing tight event messaging across the system. SEED operates efficiently within both small and large-scale OTS hardware, agnostic of cluster heterogeneity and OS running, and is capable of simulating the full system and network stack of a CPS. Our approach is validated through experiments conducted in a cluster to simulate CPS operation. Results demonstrate that SEED is capable of simulating CPSs containing 2,000,000 tasks across 2,000 nodes with only 6.89× slow down relative to real time, and executes effectively across distributed infrastructure. Peter Garraghan, David McKee 0001, Xue Ouyang 0003, David Webster, Jie Xu 0007 |
IEEE Trans. Serv. Comput. | 2 |
| 2015 | DIVIDER: Modelling and Evaluating Real-Time Service-Oriented Cyberphysical Co-SimulationsabstractThe ability to reliably distribute simulations across a distributed system and seamlessly integrate them as a workflow regardless of their level of abstraction is critical to improving the quality of product manufacturing. This paper presents the DIVIDER architecture for managing and maintaining real-time performance simulations integrated through SOAs. The described approach captures features present in complex workflow patterns such as asynchronous arbitrary cycles and estimates the worst case execution time in the context of the interfering execution environment. David McKee 0001, David Webster, Jie Xu 0007, David Battersby |
ISORC | 1 |
| 2014 | M-VCR: Multi-View Consensus Recognition for Real-Time ExperimentationabstractA major application area in the computer vision domain is gesture recognition, requiring real-time image classification to respond to human interactions. However, current state-of-the-art high-quality algorithms for image classification do not meet many dynamic real-time requirements. This paper presents the development of M-VCR - a novel approach for improving the reliability of real-time image classification. M-VCR increases the quality of classifications under real-time constraints through the adoption of fast classification algorithms, although these algorithms individually produce lower quality results, utilisation under a 'consensus' approach can achieve results equivalent to those of much higher-quality algorithms. The proposed approach also allows for different algorithms to be utilised in parallel, building on the fault tolerance technique of N-versioning. A significant improvement in image classification is experimentally demonstrated for both the SURF and MSER feature detectors through our integration consensus approach. This improvement is delivered entirely through the integration method without requiring modification of the source algorithms being used. David McKee 0001, Paul Townend, David Webster, Jie Xu 0007 |
ISORC | 1 |
| 2014 | Towards a Virtual Integration Design and Analysis Enviroment for Automotive EngineeringabstractAs the automotive industry moves towards reduced physical prototyping it is becoming more dependent on distributed simulations. However, the current technologies do not fully enable real-time distributed simulations which involve both virtual and physical components. This paper considers the current approaches to real-time distributed simulation and proposes the use of service-orientation. The highlights and current shortfalls of current research in real-time service orientation are then identified. Finally a key area of research is focused upon requiring a fundamental change of the understanding of quality of service and capability that is necessary to enable dependable real-time service orientated architectures. David McKee 0001, David Webster, Paul Townend, Jie Xu 0007, David Battersby |
ISORC | 1 |