Nick Mitchell

dblp:25/3824 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 7 first-authorSystems, architecture and hardware · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
6 papers
Program analysis · 62% Software testing · 12% Runtime systems and virtual machines · 11%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Performance modeling and evaluation · 60% Cloud and datacenter computing · 30% Processor architecture and microarchitecture · 5%

Topics — the 16 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program analysis
dynamic analysis
0.432014
Scalable Runtime Bloat Detection Using Abstract Dynamic Slicing · ACM Trans. Softw. Eng. Methodol. 2014
Finding low-utility data structures · PLDI 2010
Go with the flow: profiling copies to find runtime bloat · PLDI 2009
Program analysis › dynamic analysis
runtime bloat detection
0.432014
Scalable Runtime Bloat Detection Using Abstract Dynamic Slicing · ACM Trans. Softw. Eng. Methodol. 2014
Finding low-utility data structures · PLDI 2010
Go with the flow: profiling copies to find runtime bloat · PLDI 2009
Program analysis › dynamic analysis
profiling
0.222010
Finding low-utility data structures · PLDI 2010
Go with the flow: profiling copies to find runtime bloat · PLDI 2009
Empirical software engineering › software economics
cost-benefit analysis
0.212014
Scalable Runtime Bloat Detection Using Abstract Dynamic Slicing · ACM Trans. Softw. Eng. Methodol. 2014
Program analysis › dynamic analysis
dynamic slicing
0.212014
Scalable Runtime Bloat Detection Using Abstract Dynamic Slicing · ACM Trans. Softw. Eng. Methodol. 2014
Performance modeling and evaluation
capacity planning
0.212013
On-the-fly capacity planning · OOPSLA 2013
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.212013
On-the-fly capacity planning · OOPSLA 2013
Performance modeling and evaluation
workload characterization
0.212013
On-the-fly capacity planning · OOPSLA 2013
Performance modeling and evaluation
bottleneck analysis
0.112010
Performance analysis of idle programs · OOPSLA 2010
Runtime systems and virtual machines › runtime memory management
memory bloat
0.112007
The causes of bloat, the limits of health · OOPSLA 2007
Software maintenance and evolution
technical debt
0.112007
The causes of bloat, the limits of health · OOPSLA 2007
Cloud and datacenter computing › cloud deployment
cloud application deployment
0.112015
CanaryAdvisor: a statistical-based tool for canary testing (demo) · ISSTA 2015
Runtime systems and virtual machines
garbage collection
0.012010
Performance analysis of idle programs · OOPSLA 2010
Processor architecture and microarchitecture
memory latency tolerance
0.011998
Multi-processor Performance on the Tera MTA · SC 1998
Processor architecture and microarchitecture
multithreading
0.011998
Multi-processor Performance on the Tera MTA · SC 1998
Performance modeling and evaluation
parallel performance evaluation
0.011998
Multi-processor Performance on the Tera MTA · SC 1998

Methods — techniques the papers use, named apart from their topics

statistical analysis · 0.4continuous monitoring · 0.4sampling · 0.2expert system · 0.2declarative rules · 0.2profiling · 0.2runtime dependence graph · 0.2bounded abstract domains · 0.2call stack histogram · 0.2amdahl's law analysis · 0.0NAS benchmarks · 0.0
YearPublicationVenuePosition
2025 The Cloud, Like Building and Running Go Binaries
abstract
We document the UX challenges of targeting distributed batch-processing code to Cloud resources. The challenges largely stem from the need for every team member to know everything about the mechanics of achieving scale. Thus, both running and writing code become overwhelming tasks. We present Lunchpail, a tool designed to address these challenges.We present two case studies, one of a team running AI/ML workloads and one of the code they wrote to make it happen. We quantify the challenges with two novel UX metrics: multiplicity and divergence. We show that the code base manifests a multitude of concerns, including distribution, packaging, and automation; 64–98% of the team’s code diverges from the main goal of the application. The story is paralleled when running workloads. Users switch between 3–7 types of tasks on a daily basis (high multiplicity). The nature of these tasks differ greatly from the users’ core competencies (high divergence). In particular, we show that all users assume the daily burdens of cluster operators.We demonstrate that four angles of attack, combined, can yield significant reductions in complexity: 1) Adopt a Serverless approach, allowing code to focus on that core "2%". 2) Treat application packaging like building a Golang binary via go build. This binary embeds source, configuration, deployment logic, and a lightweight runtime that channels data to workers with fan-out and queuing. 3) Treat running distributed applications pipelines against Cloud resources like launching said binaries, with simple bash "|" syntax; 4) When possible, avoid multi-tenancy, and instead target Cloud virtual machines directly.We present a large experimental study to quantify the viability of obtaining a dedicated "burst" of cloud resources for every job run. We show VMs can be ready in well under a minute, which is 10-20x faster than scaling a Kubernetes cluster.We embody this approach in Lunchpail. Lunchpail itself is small, weighing in at 12k lines of code (10% of the size of Kubeflow, 2.5% of Ray, 1% of Kueue). We validate Lunchpail against AI/ML code, legacy chip design workloads, and show that it adds little overhead on top of acquiring Cloud VMs.
Diana Arroyo, Paul Castro, Thuan Doan, Nick Mitchell, Sara Kokkila Schumacher, Ed Seabolt, Aleksander Slominski, Ansu Varghese, Lionel Villard, Cora Coleman
ICDCS4
2015 CanaryAdvisor: a statistical-based tool for canary testing (demo)
abstract
Canary testing is an emerging technique that offers to minimize the risk of deploying a new version of software. It does so by slowly transferring load from the current to the new ("canary") version. As this ramp-up progresses, a human compares the performance and correctness of the two versions, and assesses whether to abort the canary version. For canary testing to be effective, a plethora of metrics must be analyzed, including CPU utilization and logged errors, across hundreds to thousands of machines. Performing this analysis manually is both time consuming and error prone. In this paper, we present CanaryAdvisor, a tool for automatic canary testing of cloud-based applications. CanaryAdvisor continuously monitors the deployed versions of an application and detects degradations in correctness, performance, and/or scalability. We describe our design and implementation of the CanaryAdvisor and outline open challenges.
Alexander Tarvo, Peter F. Sweeney, Nick Mitchell, V. T. Rajan, Matthew Arnold, Ioana Baldini
ISSTA3
2014 Scalable Runtime Bloat Detection Using Abstract Dynamic Slicing
abstract
Many large-scale Java applications suffer from runtime bloat. They execute large volumes of methods and create many temporary objects, all to execute relatively simple operations. There are large opportunities for performance optimizations in these applications, but most are being missed by existing optimization and tooling technology. While JIT optimizations struggle for a few percent improvement, performance experts analyze deployed applications and regularly find gains of 2× or more. Finding such big gains is difficult, for both humans and compilers, because of the diffuse nature of runtime bloat. Time is spread thinly across calling contexts, making it difficult to judge how to improve performance. Our experience shows that, in order to identify large performance bottlenecks in a program, it is more important to understand its dynamic dataflow than traditional performance metrics, such as running time. This article presents a general framework for designing and implementing scalable analysis algorithms to find causes of bloat in Java programs. At the heart of this framework is a generalized form of runtime dependence graph computed by abstract dynamic slicing , a semantics-aware technique that achieves high scalability by performing dynamic slicing over bounded abstract domains. The framework is instantiated to create two independent dynamic analyses, copy profiling and cost-benefit analysis , that help programmers identify performance bottlenecks by identifying, respectively, high-volume copy activities and data structures that have high construction cost but low benefit for the forward execution. We have successfully applied these analyses to large-scale and long-running Java applications. We show that both analyses are effective at detecting inefficient operations that can be optimized for better performance. We also demonstrate that the general framework is flexible enough to be instantiated for dynamic analyses in a variety of application domains.
Guoqing Harry Xu, Nick Mitchell, Matthew Arnold, Atanas Rountev, Edith Schonberg, Gary Sevitsky
ACM Trans. Softw. Eng. Methodol.2
2013 On-the-fly capacity planning
abstract
When resolving performance problems, a simple histogram of hot call stacks does not cut it, especially given the highly fluid nature of modern deployments. Why bother tuning, when adding a few CPUs via the management console will quickly resolve the problem? The findings of these tools are also presented without any sense of context: e.g. string conversion may be expensive, but only matters if it contributes greatly to the response time of user logins.
Nick Mitchell, Peter F. Sweeney
OOPSLA1
2011 Patterns of Memory Inefficiency
Adriana E. Chis, Nick Mitchell, Edith Schonberg, Gary Sevitsky, Patrick O'Sullivan, Trevor Parsons, John Murphy 0001
ECOOP2
2010 The big pileup
abstract
Programmers no longer write monolithic applications, they assemble code from a sea of reusable libraries and frameworks. This layered process of construction has a magnifying effect on local coding decisions. Piece by innocent piece, seemingly harmless constant factors pile up. They become part of an interstitial excess, marbled throughout the code and APIs, and difficult to remove. It is not uncommon for large applications to miss their performance targets by an order of magnitude. We commonly see web requests create objects and invoke methods by the hundreds of thousands to retrieve and format a few database records. Current Java optimizers and garbage collectors don't address many of these systemic problems. This talk discusses these issues, via many examples, with a goal of motivating research on the programming of large-scale artifacts in a way that local, often ad hoc, decisions can be unwound, rather then pile up, in the large.
Nick Mitchell
ISPASS1
2010 Performance analysis of idle programs
abstract
This paper presents an approach for performance analysis of modern enterprise-class server applications. In our experience, performance bottlenecks in these applications differ qualitatively from bottlenecks in smaller, stand-alone systems. Small applications and benchmarks often suffer from CPU-intensive hot spots. In contrast, enterprise-class multi-tier applications often suffer from problems that manifest not as hot spots, but as idle time, indicating a lack of forward motion. Many factors can contribute to undesirable idle time, including locking problems, excessive system-level activities like garbage collection, various resource constraints, and problems driving load.We present the design and methodology for WAIT, a tool to diagnosis the root cause of idle time in server applications. Given lightweight samples of Java activity on a single tier, the tool can often pinpoint the primary bottleneck on a multi-tier system. The methodology centers on an informative abstraction of the states of idleness observed in a running program. This abstraction allows the tool to distinguish, for example, between hold-ups on a database machine, insufficient load, lock contention in application code, and a conventional bottleneck due to a hot method. To compute the abstraction, we present a simple expert system based on an extensible set of declarative rules.WAIT can be deployed on the fly, without modifying or even restarting the application. Many groups in IBM have applied the tool to diagnosis performance problems in commercial systems, and we present a number of examples as case studies.
Erik R. Altman, Matthew Arnold, Stephen J. Fink, Nick Mitchell
OOPSLA4
2010 Finding low-utility data structures
abstract
Many opportunities for easy, big-win, program optimizations are missed by compilers. This is especially true in highly layered Java applications. Often at the heart of these missed optimization opportunities lie computations that, with great expense, produce data values that have little impact on the program's final output. Constructing a new date formatter to format every date, or populating a large set full of expensively constructed structures only to check its size: these involve costs that are out of line with the benefits gained. This disparity between the formation costs and accrued benefits of data structures is at the heart of much runtime bloat.
Guoqing Harry Xu, Nick Mitchell, Matthew Arnold, Atanas Rountev, Edith Schonberg, Gary Sevitsky
PLDI2
2009 Making Sense of Large Heaps
Nick Mitchell, Edith Schonberg, Gary Sevitsky
ECOOP1
2009 Go with the flow: profiling copies to find runtime bloat
abstract
Many large-scale Java applications suffer from runtime bloat. They execute large volumes of methods, and create many temporary objects, all to execute relatively simple operations. There are large opportunities for performance optimizations in these applications, but most are being missed by existing optimization and tooling technology. While JIT optimizations struggle for a few percent, performance experts analyze deployed applications and regularly find gains of 2x or more.
Guoqing Harry Xu, Matthew Arnold, Nick Mitchell, Atanas Rountev, Gary Sevitsky
PLDI3
2007 The causes of bloat, the limits of health
abstract
Applications often have large runtime memory requirements. In some cases, large memory footprint helps accomplish an important functional, performance, or engineering requirement. A large cache,for example, may ameliorate a pernicious performance problem. In general, however, finding a good balance between memory consumption and other requirements is quite challenging. To do so, the development team must distinguish effective from excessive use of memory.
Nick Mitchell, Gary Sevitsky
OOPSLA1
2006 The Runtime Structure of Object Ownership
Nick Mitchell
ECOOP1
2006 Modeling Runtime Behavior in Framework-Based Applications
Nick Mitchell, Gary Sevitsky, Harini Srinivasan
ECOOP1
2003 LeakBot: An Automated and Lightweight Tool for Diagnosing Memory Leaks in Large Java Applications
Nick Mitchell, Gary Sevitsky
ECOOP1
1998 Multi-processor Performance on the Tera MTA
abstract
The Tera MTA is a revolutionary commercial computer based on a multithreaded processor architecture. In contrast to many other parallel architectures, the Tera MTA can effectively use high amounts of parallelism on a single processor. By running multiple threads on a single processor, it can tolerate memory latency and to keep the processor saturated. If the computation is sufficiently large, it can benefit from running on multiple processors. A primary architectural goal of the MTA is that it provide scalable performance over multiple processors. This paper is a preliminary investigation of the first multi-processor Tera MTA. In a previous paper [1] we reported that on the kernel NAS 2 benchmarks [2], a single-processor MTA system running at the architected clock speed would be similar in performance to a single processor of the Cray T90. We found that the compilers of both machines were able to find the necessary threads or vector operations, after making standard changes to the random number generator. In this paper we update the single-processor results in two ways: we use only actual clock speeds, and we report improvements given by further tuning of the MTA codes. We then investigate the performance of the best single-processor codes when run on a two-processor MTA, making no further tuning effort. The parallel efficiency of the codes range from 77% to 99%. An analysis shows that the "serial bottlenecks" -- unparallelized code sections and the cost of allocating and freeing the parallel hardware resources -- account for less than a percent of the runtimes. Thus, Amdahl's Law needn't take effect on the NAS benchmarks until there are hundreds of processors running thousands of threads. Instead, the major source of inefficiency appears to be an imperfect network connecting the processors to the memory. Ideally, the network can support one memory reference per instruction. The current hardware has defects that reduce the throughput to about 85% of this rate. Except for the EP benchmark, the tuned codes issue memory references at nearly the peak rate of one per instruction. Consequently, the network can support the memory references issued by one, but not two, processors. As a result, the parallel efficiency of EP is near- perfect, but the others are reduced accordingly. Another reason for imperfect speedup pertains to the compiler. While the definition of a thread in a single processor or multi-processor mode is essentially the same, there is a different implementation and an associated overhead with running on multiple processors. We characterize the overhead of running "frays" (a collection of threads running on a single processor) and "crews" (a collection of frays, one per processor.)
Allan Snavely, Larry Carter, Jay Boisseau, Amitava Majumdar 0001, Kang Su Gatlin, Nick Mitchell, John Feo, Brian D. Koblenz
SC6