VLDB 2026 Research / reviewers in the wild / expert
Jonathan Mace
dblp:154/0937
· DBLP profile ↗
21ranked-venue papers
6as first author
10since 2021 · last 2024
0000-0002-3701-9296ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 5 since 2021Computer networks · 3 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Building AI Agents for Autonomous Clouds: Challenges and Design PrinciplesabstractThe rapid growth in the use of Large Language Models (LLMs) and AI Agents as part of software development and deployment is revolutionizing the information technology landscape. While code generation receives significant attention, a higher-impact application lies in using agents for the operational resilience of cloud services, which currently require significant human effort and domain knowledge. There is a growing interest in AI for IT Operations (AIOps), which aims to automate complex operational tasks, like fault localization and root cause analysis, reducing human intervention and customer impact. However, achieving the vision of autonomous and self-healing clouds through AIOps is hampered by the lack of standardized frameworks for building, evaluating, and improving AIOps agents. This vision paper lays the groundwork for such a framework by framing the requirements and then discussing design decisions that satisfy them. We also propose AIOpsLab, a prototype implementation leveraging agent-cloud-interface that orchestrates an application, injects real-time faults using chaos engineering, and interfaces with an agent to localize and resolve the faults. We report promising results and lay the groundwork to build a modular and robust framework for building, evaluating, and improving agents for autonomous clouds. Manish Shetty, Yinfang Chen, Gagan Somashekar, Minghua Ma, Yogesh L. Simmhan, Xuchao Zhang, Jonathan Mace, Dax Vandevoorde, Pedro Henrique B. Las-Casas, Shachee Mishra Gupta, Suman Nath, Chetan Bansal, Saravan Rajmohan |
SoCC | 7 |
| 2024 | If At First You Don't Succeed, Try, Try, Again...? Insights and LLM-informed Tooling for Detecting Retry Bugs in Software SystemsabstractRetry---the re-execution of a task on failure---is a common mechanism to enable resilient software systems. Yet, despite its commonality and long history, retry remains difficult to implement and test. Bogdan Alexandru Stoica, Utsav Sethi, Yiming Su, Cyrus Zhou, Shan Lu 0001, Jonathan Mace, Madan Musuvathi, Suman Nath |
SOSP | 6 |
| 2024 | A Qualitative Interview Study of Distributed Tracing Visualisation: A Characterisation of Challenges and OpportunitiesabstractDistributed tracing tools have emerged in recent years to enable operators of modern internet applications to troubleshoot cross-component problems in deployed applications. Due to the rich, detailed diagnostic data captured by distributed tracing tools, effectively presenting this data is important. However, use of visualisation to enable sensemaking of this complex data in distributed tracing tools has received relatively little attention. Consequently, operators struggle to make effective use of existing tools. In this article we present the first characterisation of distributed tracing visualisation through a qualitative interview study with six practitioners from two large internet companies. Across two rounds of 1-on-1 interviews we use grounded theory coding to establish users, extract concrete use cases and identify shortcomings of existing distributed tracing tools. We derive guidelines for development of future distributed tracing tools and expose several open research problems that have wide reaching implications for visualisation research and other domains. Thomas James Davidson, Emily Wall 0001, Jonathan Mace |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Groundhog: Efficient Request Isolation in FaaSabstractSecurity is a core responsibility for Function-as-a-Service (FaaS) providers. The prevailing approach isolates concurrent executions of functions in separate containers. However, successive invocations of the same function commonly reuse the runtime state of a previous invocation in order to avoid container cold-start delays. Although efficient, this container reuse has security implications for functions that are invoked on behalf of differently privileged users or administrative domains: bugs in a function's implementation --- or a third-party library/runtime it depends on --- may leak private data from one invocation of the function to a subsequent one. Mohamed Alzayat, Jonathan Mace, Peter Druschel, Deepak Garg 0001 |
EuroSys | 2 |
| 2023 | The Benefit of Hindsight: Tracing Edge-Cases in Distributed Systems
Lei Zhang 0223, Vaastav Anand, Ymir Vigfusson, Jonathan Mace |
NSDI | 5 |
| 2023 | Detection Is Better Than Cure: A Cloud Incidents PerspectiveabstractCloud providers use automated watchdogs or monitors to continuously observe service availability and to proactively report incidents when system performance degrades. Improper monitoring can lead to delays in the detection and mitigation of production incidents, which can be extremely expensive in terms of customer impacts and manual toil from engineering resources. Therefore, a systematic understanding of the pitfalls in current monitoring practices and how they can lead to production incidents is crucial for ensuring continuous reliability of cloud services. Vaibhav Ganatra, Anjaly Parayil, Supriyo Ghosh, Yu Kang 0006, Minghua Ma, Chetan Bansal, Suman Nath, Jonathan Mace |
ESEC/SIGSOFT FSE | 8 |
| 2023 | Blueprint: A Toolchain for Highly-Reconfigurable Microservice ApplicationsabstractResearchers and practitioners care deeply about the performance and correctness of microservice applications. To investigate problematic application behavior and prototype potential improvements, researchers and practitioners experiment with different designs, implementations, and deployment configurations. We argue that a key requirement for microservice experimentation is the ability to rapidly reconfigure applications and to iteratively Configure, Build, and Deploy (CBD) new variants of an application that alter or improve its design. We focus on three core experimentation use-cases: (1) updating the design to use different components, libraries, and mechanisms; (2) identifying and reproducing problematic behaviors caused by different designs; and (3) prototyping and evaluating potential solutions to such behaviors. We present Blueprint, a microservice development toolchain that enables rapid CBD. With a few lines of code, users can easily reconfigure an application's design; Blueprint then generates a fully-functioning variant of the application under the new design. Blueprint is open-source and extensible; it supports a wide variety of reconfigurable design dimensions. We have ported all major microservice benchmarks to it. Our evaluation demonstrates how Blueprint simplifies experimentation use-cases with orders-of-magnitude less code change. Vaastav Anand, Deepak Garg 0001, Antoine Kaufmann, Jonathan Mace |
SOSP | 4 |
| 2023 | Antipode: Enforcing Cross-Service Causal Consistency in Distributed ApplicationsabstractModern internet-scale applications suffer from cross-service inconsistencies, arising because applications combine multiple independent and mutually-oblivious datastores. The end-to-end execution flow of each user request spans many different services and datastores along the way, implicitly establishing ordering dependencies among operations at different datastores. Readers should observe this ordering and, in today's systems, they do not. João Loff, Daniel Porto 0002, João Garcia 0001, Jonathan Mace, Rodrigo Rodrigues 0001 |
SOSP | 4 |
| 2022 | See it to believe it?: the role of visualisation in systems researchabstractA common fixture of computer systems research are Visualisation-in-the-Loop Tools: tools that produce complex output data and require a human user to interpret the data visually. However, systems research frequently omits or sidelines details of the visualisation components that were necessary for the tool. In a survey of 1,274 recent systems papers we find that at least 7.7% (98) of them present a visualisation-in-the-loop tool. We also find that the majority of these publications pay no attention and give little explanation to implemented visualisations. We propose that the impact and reach of visualisation-in-the-loop systems research can be greatly enhanced when exposition is given to visualisation, and propose a concrete checklist of steps for authors to realise this opportunity. Thomas James Davidson, Jonathan Mace |
SoCC | 2 |
| 2021 | Systems trivia nightabstractThe past year has been mentally and physically testing because of the COVID19 pandemic. Because of the pandemic, conferences have moved to a remote style which has really hampered the social aspect of conferences. To increase the social aspect of conferences in a remote time, we propose doing an online Trivia as an event at HotOS. Vaastav Anand, Roberta De Viti, Jonathan Mace |
HotOS | 3 |
| 2020 | Serving DNNs like Clockwork: Performance Predictability from the Bottom Up
Arpan Gujarati, Reza Karimi, Safya Alzayat, Antoine Kaufmann, Ymir Vigfusson, Jonathan Mace |
OSDI | 7 |
| 2019 | Sifter: Scalable Sampling for Distributed Traces, without Feature EngineeringabstractDistributed tracing is a core component of cloud and datacenter systems, and provides visibility into their end-to-end runtime behavior. To reduce computational and storage overheads, most tracing frameworks do not keep all traces, but sample them uniformly at random. While effective at reducing overheads, uniform random sampling inevitably captures redundant, common-case execution traces, which are less useful for analysis and troubleshooting tasks. In this work we present Sifter, a general-purpose framework for biased trace sampling. Sifter captures qualitatively more diverse traces, by weighting sampling decisions towards edge-case code paths, infrequent request types, and anomalous events. Sifter does so by using the incoming stream of traces to build an unbiased low-dimensional model that approximates the system's common-case behavior. Sifter then biases sampling decisions towards traces that are poorly captured by this model. We have implemented Sifter, integrated it with several open-source tracing systems, and evaluate with traces from a range of open-source and production distributed systems. Our evaluation shows that Sifter effectively biases towards anomalous and outlier executions, is robust to noisy and heterogeneous traces, is efficient and scalable, and adapts to changes in workloads over time. Pedro Henrique B. Las-Casas, Giorgi Papakerashvili, Vaastav Anand, Jonathan Mace |
SoCC | 4 |
| 2018 | Weighted Sampling of Execution Traces: Capturing More Needles and Less HayabstractEnd-to-end tracing has emerged recently as a valuable tool to improve the dependability of distributed systems, by performing dynamic verification and diagnosing correctness and performance problems. Contrary to logging, end-to-end traces enable coherent sampling of the entire execution of specific requests, and this is exploited by many deployments to reduce the overhead and storage requirements of tracing. This sampling, however, is usually done uniformly at random, which dedicates a large fraction of the sampling budget to common, 'normal' executions, while missing infrequent, but sometimes important, erroneous or anomalous executions. In this paper we define the representative trace sampling problem, and present a new approach, based on clustering of execution graphs, that is able to bias the sampling of requests to maximize the diversity of execution traces stored towards infrequent patterns. In a preliminary, but encouraging work, we show how our approach chooses to persist representative and diverse executions, even when anomalous ones are very infrequent. Pedro Henrique B. Las-Casas, Jonathan Mace, Dorgival O. Guedes, Rodrigo Fonseca |
SoCC | 2 |
| 2018 | Universal context propagation for distributed system instrumentationabstractMany tools for analyzing distributed systems propagate contexts along the execution paths of requests, tasks, and jobs, in order to correlate events across process, component and machine boundaries. There is a wide range of existing and proposed uses for these tools, which we call cross-cutting tools, such as tracing, debugging, taint propagation, provenance, auditing, and resource management, but few of them get deployed pervasively in large systems. When they do, they are brittle, hard to evolve, and cannot coexist with each other. While they use very different context metadata, the way they propagate the information alongside execution is the same. Nevertheless, in existing tools, these aspects are deeply intertwined, causing most of these problems. Jonathan Mace, Rodrigo Fonseca |
EuroSys | 1 |
| 2017 | Canopy: An End-to-End Performance Tracing And Analysis SystemabstractThis paper presents Canopy, Facebook's end-to-end performance tracing infrastructure. Canopy records causally related performance data across the end-to-end execution path of requests, including from browsers, mobile applications, and backend services. Canopy processes traces in near real-time, derives user-specified features, and outputs to performance datasets that aggregate across billions of requests. Using Canopy, Facebook engineers can query and analyze performance data in real-time. Canopy addresses three challenges we have encountered in scaling performance analysis: supporting the range of execution and performance models used by different components of the Facebook stack; supporting interactive ad-hoc analysis of performance data; and enabling deep customization by users, from sampling traces to extracting and visualizing features. Canopy currently records and processes over 1 billion traces per day. We discuss how Canopy has evolved to apply to a wide range of scenarios, and present case studies of its use in solving various performance challenges. Jonathan Kaldor, Jonathan Mace, Michal Bejda, Edison Gao, Wiktor Kuropatwa, Joe O'Neill, Kian Win Ong, Bill Schaller, Pingjia Shan, Brendan Viscomi, Vinod Venkataraman, Kaushik Veeraraghavan, Yee Jiun Song |
SOSP | 2 |
| 2017 | Pivot Tracing: Dynamic Causal Monitoring for Distributed SystemsabstractMonitoring and troubleshooting distributed systems is notoriously difficult; potential problems are complex, varied, and unpredictable. The monitoring and diagnosis tools commonly used today—logs, counters, and metrics—have two important limitations: what gets recorded is defined a priori , and the information is recorded in a component- or machine-centric way, making it extremely hard to correlate events that cross these boundaries. This article presents Pivot Tracing, a monitoring framework for distributed systems that addresses both limitations by combining dynamic instrumentation with a novel relational operator: the happened-before join. Pivot Tracing gives users, at runtime, the ability to define arbitrary metrics at one point of the system, while being able to select, filter, and group by events meaningful at other parts of the system, even when crossing component or machine boundaries. We have implemented a prototype of Pivot Tracing for Java-based systems and evaluate it on a heterogeneous Hadoop cluster comprising HDFS, HBase, MapReduce, and YARN. We show that Pivot Tracing can effectively identify a diverse range of root causes such as software bugs, misconfiguration, and limping hardware. We show that Pivot Tracing is dynamic, extensible, and enables cross-tier analysis between inter-operating applications, with low execution overhead. Jonathan Mace, Ryan Roelke, Rodrigo Fonseca |
ACM Trans. Comput. Syst. | 1 |
| 2016 | Principled workflow-centric tracing of distributed systemsabstractWorkflow-centric tracing captures the workflow of causally-related events (e.g., work done to process a request) within and among the components of a distributed system. As distributed systems grow in scale and complexity, such tracing is becoming a critical tool for understanding distributed system behavior. Yet, there is a fundamental lack of clarity about how such infrastructures should be designed to provide maximum benefit for important management tasks, such as resource accounting and diagnosis. Without research into this important issue, there is a danger that workflow-centric tracing will not reach its full potential. To help, this paper distills the design space of workflow-centric tracing and describes key design choices that can help or hinder a tracing infrastructures utility for important tasks. Our design space and the design choices we suggest are based on our experiences developing several previous workflow-centric tracing infrastructures. Raja R. Sambasivan, Ilari Shafer, Jonathan Mace, Benjamin H. Sigelman, Rodrigo Fonseca, Gregory R. Ganger |
SoCC | 3 |
| 2016 | 2DFQ: Two-Dimensional Fair Queuing for Multi-Tenant Cloud ServicesabstractIn many important cloud services, different tenants execute their requests in the thread pool of the same process, requiring fair sharing of resources. However, using fair queue schedulers to provide fairness in this context is difficult because of high execution concurrency, and because request costs are unknown and have high variance. Using fair schedulers like WFQ and WF²Q in such settings leads to bursty schedules, where large requests block small ones for long periods of time. In this paper, we propose Two-Dimensional Fair Queueing (2DFQ), which spreads requests of different costs across di erent threads and minimizes the impact of tenants with unpredictable requests. In evaluation on production workloads from Azure Storage, a large-scale cloud system at Microsoft, we show that 2DFQ reduces the burstiness of service by 1-2 orders of magnitude. On workloads where many large requests compete with small ones, 2DFQ improves 99th percentile latencies by up to 2 orders of magnitude. Jonathan Mace, Peter Bodík, Madan Musuvathi, Rodrigo Fonseca, Krishnan Varadarajan |
SIGCOMM | 1 |
| 2016 | Pivot Tracing: Dynamic Causal Monitoring for Distributed Systems
Jonathan Mace, Ryan Roelke, Rodrigo Fonseca |
USENIX ATC | 1 |
| 2015 | Retro: Targeted Resource Management in Multi-tenant Distributed Systems
Jonathan Mace, Peter Bodík, Rodrigo Fonseca, Madan Musuvathi |
NSDI | 1 |
| 2015 | Pivot tracing: dynamic causal monitoring for distributed systemsabstractMonitoring and troubleshooting distributed systems is notoriously difficult; potential problems are complex, varied, and unpredictable. The monitoring and diagnosis tools commonly used today -- logs, counters, and metrics -- have two important limitations: what gets recorded is defined a priori, and the information is recorded in a component- or machine-centric way, making it extremely hard to correlate events that cross these boundaries. This paper presents Pivot Tracing, a monitoring framework for distributed systems that addresses both limitations by combining dynamic instrumentation with a novel relational operator: the happened-before join. Pivot Tracing gives users, at runtime, the ability to define arbitrary metrics at one point of the system, while being able to select, filter, and group by events meaningful at other parts of the system, even when crossing component or machine boundaries. We have implemented a prototype of Pivot Tracing for Java-based systems and evaluate it on a heterogeneous Hadoop cluster comprising HDFS, HBase, MapReduce, and YARN. We show that Pivot Tracing can effectively identify a diverse range of root causes such as software bugs, misconfiguration, and limping hardware. We show that Pivot Tracing is dynamic, extensible, and enables cross-tier analysis between inter-operating applications, with low execution overhead. Jonathan Mace, Ryan Roelke, Rodrigo Fonseca |
SOSP | 1 |