Matthew Arnold

dblp:78/6495 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
2since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 19 · 8 first-authorSystems, architecture and hardware · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
17 papers
Program analysis · 44% Runtime systems and virtual machines · 17% Compilers and program optimization · 13%
Artificial intelligence
1 paper
Trustworthy machine learning · 67% Efficient and distributed learning · 33%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Performance modeling and evaluation · 40% Parallel and multicore computing · 33% Cloud and datacenter computing · 18%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › automated machine learning
model performance prediction
0.512021
Learning Prediction Intervals for Model Performance · AAAI 2021
Machine learning › Trustworthy machine learning › uncertainty estimation
prediction intervals
0.512021
Learning Prediction Intervals for Model Performance · AAAI 2021
Machine learning › Trustworthy machine learning
uncertainty estimation
0.512021
Learning Prediction Intervals for Model Performance · AAAI 2021
Program analysis
dynamic analysis
0.432014
Scalable Runtime Bloat Detection Using Abstract Dynamic Slicing · ACM Trans. Softw. Eng. Methodol. 2014
Finding low-utility data structures · PLDI 2010
Go with the flow: profiling copies to find runtime bloat · PLDI 2009
Program analysis › dynamic analysis
runtime bloat detection
0.432014
Scalable Runtime Bloat Detection Using Abstract Dynamic Slicing · ACM Trans. Softw. Eng. Methodol. 2014
Finding low-utility data structures · PLDI 2010
Go with the flow: profiling copies to find runtime bloat · PLDI 2009
Compilers and program optimization › dynamic optimization
profile-guided optimization
0.372007
Using hpm-sampling to drive dynamic compilation · OOPSLA 2007
Online performance auditing: using hot optimizations without getting burned · PLDI 2006
A Survey of Adaptive Optimization in Virtual Machines · Proc. IEEE 2005
Program analysis › dynamic analysis
profiling
0.342010
Finding low-utility data structures · PLDI 2010
Go with the flow: profiling copies to find runtime bloat · PLDI 2009
Online feedback-directed optimization of Java · OOPSLA 2002
Program analysis › dynamic analysis
runtime monitoring
0.222011
QVM: An Efficient Runtime for Detecting Defects in Deployed Systems · ACM Trans. Softw. Eng. Methodol. 2011
QVM: an efficient runtime for detecting defects in deployed systems · OOPSLA 2008
Empirical software engineering › software economics
cost-benefit analysis
0.212014
Scalable Runtime Bloat Detection Using Abstract Dynamic Slicing · ACM Trans. Softw. Eng. Methodol. 2014
Program analysis › dynamic analysis
dynamic slicing
0.212014
Scalable Runtime Bloat Detection Using Abstract Dynamic Slicing · ACM Trans. Softw. Eng. Methodol. 2014
Runtime systems and virtual machines
dynamic compilation
0.122007
Using hpm-sampling to drive dynamic compilation · OOPSLA 2007
A Survey of Adaptive Optimization in Virtual Machines · Proc. IEEE 2005
Programming languages and type systems
aspect-oriented programming
0.122006
Efficient control flow quantification · OOPSLA 2006
Adapting virtual machine techniques for seamless aspect support · OOPSLA 2006
Software testing
fault detection
0.112011
QVM: An Efficient Runtime for Detecting Defects in Deployed Systems · ACM Trans. Softw. Eng. Methodol. 2011
Performance modeling and evaluation
bottleneck analysis
0.112010
Performance analysis of idle programs · OOPSLA 2010
Parallel and multicore computing
dynamic analysis
0.112009
A concurrent dynamic analysis framework for multicore hardware · OOPSLA 2009
Operating systems › resource management
memory management
0.112008
Jolt: lightweight dynamic analysis and removal of object churn · OOPSLA 2008
Program verification › dynamic verification
runtime verification
0.112008
QVM: an efficient runtime for detecting defects in deployed systems · OOPSLA 2008
Runtime systems and virtual machines › runtime memory management
stack allocation
0.112008
Jolt: lightweight dynamic analysis and removal of object churn · OOPSLA 2008
Runtime systems and virtual machines › virtual machine implementation
java virtual machine
0.142011
QVM: An Efficient Runtime for Detecting Defects in Deployed Systems · ACM Trans. Softw. Eng. Methodol. 2011
Improving virtual machine performance using a cross-run profile repository · OOPSLA 2005
Online feedback-directed optimization of Java · OOPSLA 2002
Cloud and datacenter computing › cloud deployment
cloud application deployment
0.112015
CanaryAdvisor: a statistical-based tool for canary testing (demo) · ISSTA 2015
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation
0.112006
Efficient control flow quantification · OOPSLA 2006
Software maintenance and evolution › software performance engineering
performance anomaly detection
0.112006
Online performance auditing: using hot optimizations without getting burned · PLDI 2006
Program analysis › dynamic analysis › profiling
edge profiling
0.012002
Online feedback-directed optimization of Java · OOPSLA 2002
Runtime systems and virtual machines
garbage collection
0.012010
Performance analysis of idle programs · OOPSLA 2010
Program analysis › dynamic analysis
instrumentation
0.012001
A Framework for Reducing the Cost of Instrumented Code · PLDI 2001
Compilers and program optimization
dynamic optimization
0.012000
Adaptive optimization in the Jalapeño JVM · OOPSLA 2000
Compilers and program optimization › interprocedural optimization
profile-guided inlining
0.012000
Adaptive optimization in the Jalapeño JVM · OOPSLA 2000
Debugging and program repair
bug reproduction
0.012008
QVM: an efficient runtime for detecting defects in deployed systems · OOPSLA 2008
Program analysis › static analysis › pointer analysis
escape analysis
0.012008
Jolt: lightweight dynamic analysis and removal of object churn · OOPSLA 2008
Performance modeling and evaluation › performance monitoring
hardware performance monitoring
0.012007
Using hpm-sampling to drive dynamic compilation · OOPSLA 2007

Methods — techniques the papers use, named apart from their topics

transfer learning · 0.5statistical analysis · 0.4continuous monitoring · 0.4declarative rules · 0.2profiling · 0.2runtime dependence graph · 0.2bounded abstract domains · 0.2typestate checking · 0.1runtime monitoring · 0.1overhead management · 0.1sampling · 0.1online profiling · 0.1expert system · 0.1dynamic analysis · 0.1binary rewriting · 0.1hardware performance monitor sampling · 0.1online optimization evaluation · 0.1
YearPublicationVenuePosition
2023 Game Physics Engine Using Optimised Geometric Algebra RISC-V Vector Extensions Code Using Fourier Series Data
Ed Saribatir, Niko Zurstraßen, Dietmar Hildenbrand, Florian Stock, Atilio Morillo Piña, Frederic von Wegner, Zheng Yan 0001, Shiping Wen 0001, Matthew Arnold
CGI (4)9
2021 Learning Prediction Intervals for Model Performance
abstract
Understanding model performance on unlabeled data is a fundamental challenge of developing, deploying, and maintaining AI systems. Model performance is typically evaluated using test sets or periodic manual quality assessments, both of which require laborious manual data labeling. Automated performance prediction techniques aim to mitigate this burden, but potential inaccuracy and a lack of trust in their predictions has prevented their widespread adoption. We address this core problem of performance prediction uncertainty with a method to compute prediction intervals for model performance. Our methodology uses transfer learning to train an uncertainty model to estimate the uncertainty of model performance predictions. We evaluate our approach across a wide range of drift conditions and show substantial improvement over competitive baselines. We believe this result makes prediction intervals, and performance prediction in general, significantly more practical for real-world use.
Benjamin Elder, Matthew Arnold, Anupama Murthi, Jirí Navrátil 0001
AAAI2
2015 CanaryAdvisor: a statistical-based tool for canary testing (demo)
abstract
Canary testing is an emerging technique that offers to minimize the risk of deploying a new version of software. It does so by slowly transferring load from the current to the new ("canary") version. As this ramp-up progresses, a human compares the performance and correctness of the two versions, and assesses whether to abort the canary version. For canary testing to be effective, a plethora of metrics must be analyzed, including CPU utilization and logged errors, across hundreds to thousands of machines. Performing this analysis manually is both time consuming and error prone. In this paper, we present CanaryAdvisor, a tool for automatic canary testing of cloud-based applications. CanaryAdvisor continuously monitors the deployed versions of an application and detects degradations in correctness, performance, and/or scalability. We describe our design and implementation of the CanaryAdvisor and outline open challenges.
Alexander Tarvo, Peter F. Sweeney, Nick Mitchell, V. T. Rajan, Matthew Arnold, Ioana Baldini
ISSTA5
2014 Scalable Runtime Bloat Detection Using Abstract Dynamic Slicing
abstract
Many large-scale Java applications suffer from runtime bloat. They execute large volumes of methods and create many temporary objects, all to execute relatively simple operations. There are large opportunities for performance optimizations in these applications, but most are being missed by existing optimization and tooling technology. While JIT optimizations struggle for a few percent improvement, performance experts analyze deployed applications and regularly find gains of 2× or more. Finding such big gains is difficult, for both humans and compilers, because of the diffuse nature of runtime bloat. Time is spread thinly across calling contexts, making it difficult to judge how to improve performance. Our experience shows that, in order to identify large performance bottlenecks in a program, it is more important to understand its dynamic dataflow than traditional performance metrics, such as running time. This article presents a general framework for designing and implementing scalable analysis algorithms to find causes of bloat in Java programs. At the heart of this framework is a generalized form of runtime dependence graph computed by abstract dynamic slicing , a semantics-aware technique that achieves high scalability by performing dynamic slicing over bounded abstract domains. The framework is instantiated to create two independent dynamic analyses, copy profiling and cost-benefit analysis , that help programmers identify performance bottlenecks by identifying, respectively, high-volume copy activities and data structures that have high construction cost but low benefit for the forward execution. We have successfully applied these analyses to large-scale and long-running Java applications. We show that both analyses are effective at detecting inefficient operations that can be optimized for better performance. We also demonstrate that the general framework is flexible enough to be instantiated for dynamic analyses in a variety of application domains.
Guoqing Harry Xu, Nick Mitchell, Matthew Arnold, Atanas Rountev, Edith Schonberg, Gary Sevitsky
ACM Trans. Softw. Eng. Methodol.3
2011 A different approach to teaching Chinese through serious games
abstract
From the days of computer assisted language learning (CALL) [10], using computers as a means to acquire a new language has been a long standing research field. Chinese is a notoriousloy difficult language to learn, and teaching methods that are used take a long time and can be tedious[7], but these techniques have been shown to be effective[11]. We believe that by creating a serious game that uses newer second language acauisition techniques, the learning process can be expidited and enjoyable. Our game exposes the player to an abundance of simple, comprehensible target language input, which provides an interesting, motivating, and low stress setting. Also Lost in the Middle Kingdom utilizes total immersion, focusing on the language's culture to create a holistic experience.
Jeremiah J. Shepherd, Renaldo J. Doe, Matthew Arnold, Jijun Tang
FDG3
2011 QVM: An Efficient Runtime for Detecting Defects in Deployed Systems
abstract
Coping with software defects that occur in the post-deployment stage is a challenging problem: bugs may occur only when the system uses a specific configuration and only under certain usage scenarios. Nevertheless, halting production systems until the bug is tracked and fixed is often impossible. Thus, developers have to try to reproduce the bug in laboratory conditions. Often, the reproduction of the bug takes most of the debugging effort. In this paper we suggest an approach to address this problem by using a specialized runtime environment called Quality Virtual Machine (QVM). QVM efficiently detects defects by continuously monitoring the execution of the application in a production setting. QVM enables the efficient checking of violations of user-specified correctness properties, that is, typestate safety properties, Java assertions, and heap properties pertaining to ownership. QVM is markedly different from existing techniques for continuous monitoring by using a novel overhead manager which enforces a user-specified overhead budget for quality checks. Existing tools for error detection in the field usually disrupt the operation of the deployed system. QVM, on the other hand, provides a balanced trade-off between the cost of the monitoring process and the maintenance of sufficient accuracy for detecting defects. Specifically, the overhead cost of using QVM instead of a standard JVM, is low enough to be acceptable in production environments. We implemented QVM on top of IBM’s J9 Java Virtual Machine and used it to detect and fix various errors in real-world applications.
Matthew Arnold, Martin T. Vechev, Eran Yahav
ACM Trans. Softw. Eng. Methodol.1
2010 Performance analysis of idle programs
abstract
This paper presents an approach for performance analysis of modern enterprise-class server applications. In our experience, performance bottlenecks in these applications differ qualitatively from bottlenecks in smaller, stand-alone systems. Small applications and benchmarks often suffer from CPU-intensive hot spots. In contrast, enterprise-class multi-tier applications often suffer from problems that manifest not as hot spots, but as idle time, indicating a lack of forward motion. Many factors can contribute to undesirable idle time, including locking problems, excessive system-level activities like garbage collection, various resource constraints, and problems driving load.We present the design and methodology for WAIT, a tool to diagnosis the root cause of idle time in server applications. Given lightweight samples of Java activity on a single tier, the tool can often pinpoint the primary bottleneck on a multi-tier system. The methodology centers on an informative abstraction of the states of idleness observed in a running program. This abstraction allows the tool to distinguish, for example, between hold-ups on a database machine, insufficient load, lock contention in application code, and a conventional bottleneck due to a hot method. To compute the abstraction, we present a simple expert system based on an extensible set of declarative rules.WAIT can be deployed on the fly, without modifying or even restarting the application. Many groups in IBM have applied the tool to diagnosis performance problems in commercial systems, and we present a number of examples as case studies.
Erik R. Altman, Matthew Arnold, Stephen J. Fink, Nick Mitchell
OOPSLA2
2010 Finding low-utility data structures
abstract
Many opportunities for easy, big-win, program optimizations are missed by compilers. This is especially true in highly layered Java applications. Often at the heart of these missed optimization opportunities lie computations that, with great expense, produce data values that have little impact on the program's final output. Constructing a new date formatter to format every date, or populating a large set full of expensively constructed structures only to check its size: these involve costs that are out of line with the benefits gained. This disparity between the formation costs and accrued benefits of data structures is at the heart of much runtime bloat.
Guoqing Harry Xu, Nick Mitchell, Matthew Arnold, Atanas Rountev, Edith Schonberg, Gary Sevitsky
PLDI3
2009 A concurrent dynamic analysis framework for multicore hardware
abstract
Software has spent the bounty of Moore's law by solving harder problems and exploiting abstractions, such as high-level languages, virtual machine technology, binary rewriting, and dynamic analysis. Abstractions make programmers more productive and programs more portable, but usually slow them down. Since Moore's law is now delivering multiple cores instead of faster processors, future systems must either bear a relatively higher cost for abstractions or use some cores to help tolerate abstraction costs.
Matthew Arnold, Steve Blackburn, Kathryn S. McKinley
OOPSLA2
2009 Go with the flow: profiling copies to find runtime bloat
abstract
Many large-scale Java applications suffer from runtime bloat. They execute large volumes of methods, and create many temporary objects, all to execute relatively simple operations. There are large opportunities for performance optimizations in these applications, but most are being missed by existing optimization and tooling technology. While JIT optimizations struggle for a few percent, performance experts analyze deployed applications and regularly find gains of 2x or more.
Guoqing Harry Xu, Matthew Arnold, Nick Mitchell, Atanas Rountev, Gary Sevitsky
PLDI2
2008 QVM: an efficient runtime for detecting defects in deployed systems
abstract
Coping with software defects that occur in the post-deployment stage is a challenging problem: bugs may occur only when the system uses a specific configuration and only under certain usage scenarios. Nevertheless, halting production systems until the bug is tracked and fixed is often impossible. Thus, developers have to try to reproduce the bug in laboratory conditions. Often the reproduction of the bug consists of the lion share of the debugging effort.
Matthew Arnold, Martin T. Vechev, Eran Yahav
OOPSLA1
2008 Jolt: lightweight dynamic analysis and removal of object churn
abstract
It has been observed that component-based applications exhibit object churn, the excessive creation of short-lived objects, often caused by trading performance for modularity. Because churned objects are short-lived, they appear to be good candidates for stack allocation. Unfortunately, most churned objects escape their allocating function, making escape analysis ineffective.
Ajeet Shankar, Matthew Arnold, Rastislav Bodík
OOPSLA2
2007 A Loop Correlation Technique to Improve Performance Auditing
Jeremy Lau, Matthew Arnold, Michael Hind, Brad Calder
PACT2
2007 Using hpm-sampling to drive dynamic compilation
abstract
All high-performance production JVMs employ an adaptive strategy for program execution. Methods are first executed unoptimized and then an online profiling mechanism is used to find a subset of methods that should be optimized during the same execution. This paper empirically evaluates the design space of several profilers for initiating dynamic compilation and shows that existing online profiling schemes suffer from several limitations. They provide an insufficient number of samples, are untimely, and have limited accuracy at determining the frequently executed methods. We describe and comprehensively evaluate HPM-sampling, a simple but effective profiling scheme for finding optimization candidates using hardware performance monitors (HPMs) that addresses the aforementioned limitations. We show that HPM-sampling is more accurate; has low overhead; and improves performance by 5.7% on average and up to 18.3% when compared to the default system in Jikes RVM, without changing the compiler.
Dries Buytaert, Andy Georges, Michael Hind, Matthew Arnold, Lieven Eeckhout, Koen De Bosschere
OOPSLA4
2007 Dynamic compilation: the benefits of early investing
abstract
Dynamic compilation is typically performed in a separate thread, asynchronously with the remaining application threads. This compilation thread is often scheduled for execution in a simple round-robin fashion either by the operating system or by the virtual machine itself. Despite the popularity of this approach in production virtual machines, it has a number of shortcomings that can lead to suboptimal performance.
Prasad A. Kulkarni, Matthew Arnold, Michael Hind
VEE2
2006 Adapting virtual machine techniques for seamless aspect support
abstract
Current approaches to compiling aspect-oriented programs are in-efficient. This inefficiency has negative effects on the productiv-ity of the development process and is especially prohibitive for dynamic aspect deployment. In this work, we present how well-known virtual machine techniques can be used with only slight modifications to support fast aspect deployment while retaining runtime performance. Our implementation accelerates dynamic as-pect deployment by several orders of magnitude relative to main-stream aspect-oriented environments. We also provide a detailed comparison of alternative implementations of execution environ-ments with support for dynamic aspect deployment.
Christoph Bockisch, Matthew Arnold, Tom Dinkelaker, Mira Mezini
OOPSLA2
2006 Efficient control flow quantification
abstract
Aspect-oriented programming (AOP) is increasingly gaining in popularity. However, the focus of aspect-oriented language research has been mostly on language design issues; efficient implementation techniques have been less popular. As a result, the performance of certain AOP constructs is still poor. This is in particular true for constructs that rely on dynamic properties of the execution (e.g., the cflow construct).In this paper, we present efficient implementation techniques for cflow that exploit direct access to internal structures of the virtual machine running an application, such as the call stack, as well as the integration of these techniques into the just-in-time compiler code generation process.Our results show that AOP has the potential to make programs that need to define control flow-dependent behavior not only more modular but also more efficient. By making means for control flow-dependent behavior part of the language, AOP opens the possibility of applying sophisticated compiler optimizations that are out of reach for application programmers.
Christoph Bockisch, Sebastian Kanthak, Michael Haupt 0003, Matthew Arnold, Mira Mezini
OOPSLA4
2006 Online performance auditing: using hot optimizations without getting burned
abstract
As hardware complexity increases and virtualization is added at more layers of the execution stack, predicting the performance impact of optimizations becomes increasingly difficult. Production compilers and virtual machines invest substantial development effort in performance tuning to achieve good performance for a range of benchmarks. Although optimizations typically perform well on average, they often have unpredictable impact on running time, sometimes degrading performance significantly. Today's VMs perform sophisticated feedback-directed optimizations, but these techniques do not address performance degradations, and they actually make the situation worse by making the system more unpredictable.This paper presents an online framework for evaluating the effectiveness of optimizations, enabling an online system to automatically identify and correct performance anomalies that occur at runtime. This work opens the door for a fundamental shift in the way optimizations are developed and tuned for online systems, and may allow the body of work in offline empirical optimization search to be applied automatically at runtime. We present our implementation and evaluation of this system in a product Java VM.
Jeremy Lau, Matthew Arnold, Michael Hind, Brad Calder
PLDI2
2005 Collecting and Exploiting High-Accuracy Call Graph Profiles in Virtual Machines
abstract
Due to the high dynamic frequency of virtual method calls in typical object-oriented programs, feedback-directed devirtualization and inlining is one of the most important optimizations performed by high-performance virtual machines. A critical input to effective feedback-directed inlining is an accurate dynamic call graph. In a virtual machine, the dynamic call graph is computed online during program execution. Therefore, to maximize overall system performance, the profiling mechanism must strike a balance between profile accuracy, the speed at which the profile becomes available to the optimizer, and profiling overhead. This paper introduces a new low-overhead sampling-based technique that rapidly converges on a high-accuracy dynamic call graph. We have implemented the technique in two high-performance virtual machines: Jikes RVM and J9. We empirically assess our profiling technique by reporting on the accuracy of the dynamic call graphs it computes and by demonstrating that increasing the accuracy of the dynamic call graph results in more effective feedback-directed inlining.
Matthew Arnold, David Grove
CGO1
2005 Improving virtual machine performance using a cross-run profile repository
abstract
Virtual machines for languages such as the Java programming language make extensive use of online profiling and dynamic optimization to improve program performance. But despite the important role that profiling plays in achieving high performance, current virtual machines discard a program's profile data at the end of execution, wasting the opportunity to use past knowledge to improve future performance. In this paper, we present a fully automated architecture for exploiting cross-run profile data in virtual machines. Our work addresses a number of challenges that previously limited the practicality of such an approach.We apply this architecture to address the problem of selective optimization, and describe our implementation in IBM's J9 Java virtual machine. Our results demonstrate substantial performance improvements on a broad suite of Java programs, with the average performance ranging from 8.8% -- 16.6% depending on the execution scenario.
Matthew Arnold, Adam Welc, V. T. Rajan
OOPSLA1
2005 A Survey of Adaptive Optimization in Virtual Machines
abstract
Virtual machines face significant performance challenges beyond those confronted by traditional static optimizers. First, portable program representations and dynamic language features, such as dynamic class loading, force the deferral of most optimizations until runtime, inducing runtime optimization overhead. Second, modular program representations preclude many forms of whole-program interprocedural optimization. Third, virtual machines incur additional costs for runtime services such as security guarantees and automatic memory management. To address these challenges, vendors have invested considerable resources into adaptive optimization systems in production virtual machines. Today, mainstream virtual machine implementations include substantial infrastructure for online monitoring and profiling, runtime compilation, and feedback-directed optimization. As a result, adaptive optimization has begun to mature as a widespread production-level technology. This paper surveys the evolution and current state of adaptive optimization technology in virtual machines.
Matthew Arnold, Stephen J. Fink, David Grove, Michael Hind, Peter F. Sweeney
Proc. IEEE1
2002 Thin Guards: A Simple and Effective Technique for Reducing the Penalty of Dynamic Class Loading
Matthew Arnold, Barbara G. Ryder
ECOOP1
2002 Online feedback-directed optimization of Java
abstract
This paper describes the implementation of an online feedback-directed optimization system. The system is fully automatic; it requires no prior (offline) profiling run. It uses a previously developed low-overhead instrumentation sampling framework to collect control flow graph edge profiles. This profile information is used to drive several traditional optimizations, as well as a novel algorithm for performing feedback-directed control flow graph node splitting. We empirically evaluate this system and demonstrate improvements in peak performance of up to 17% while keeping overhead low, with no individual execution being degraded by more than 2% because of instrumentation.
Matthew Arnold, Michael Hind, Barbara G. Ryder
OOPSLA1
2001 A Framework for Reducing the Cost of Instrumented Code
abstract
Instrumenting code to collect profiling information can cause substantial execution overhead. This overhead makes instrumentation difficult to perform at runtime, often preventing many known offline feedback-directed optimizations from being used in online systems. This paper presents a general framework for performing instrumentation sampling to reduce the overhead of previously expensive instrumentation. The framework is simple and effective, using code-duplication and counter-based sampling to allow switching between instrumented and non-instrumented code.
Matthew Arnold, Barbara G. Ryder
PLDI1
2000 Adaptive optimization in the Jalapeño JVM
abstract
Future high-performance virtual machines will improve performance through sophisticated online feedback-directed optimizations. this paper presents the architecture of the Jalapeño Adaptive Optimization System, a system to support leading-edge virtual machine technology and enable ongoing research on online feedback-directed optimizations. We describe the extensible system architecture, based on a federation of threads with asynchronous communication. We present an implementation of the general architecture that supports adaptive multi-level optimization based purely on statistical sampling. We empirically demonstrate that this profiling technique has low overhead and can improve startup and steady-state performance, even without the presence of online feedback-directed optimizations. The paper also describes and evaluates an online feedback-directed inlining optimization based on statistical edge sampling. The system is written completely in Java, applying the described techniques not only to application code and standard libraries, but also to the virtual machine itself.
Matthew Arnold, Stephen J. Fink, David Grove, Michael Hind, Peter F. Sweeney
OOPSLA1