VLDB 2026 Research / reviewers in the wild / expert
Carl Friedrich Bolz-Tereick
dblp:50/2706 · also Carl Friedrich Bolz
· DBLP profile ↗
19ranked-venue papers
7as first author
4since 2021 · last 2026
0000-0003-4562-1356ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 19 · 7 first-author · 4 since 2021Theory of computation · 2 · 2 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generating Interpreter-Specific Tracers for Meta-tracing JIT CompilersabstractThe RPython framework’s meta-tracing JIT compiler uses a single generic tracer for all interpreters written in the RPython language, but this genericity causes overhead: the tracer dispatches every operation through an opcode lookup and performs redundant bookkeeping even for pure operations. We propose GenExtension, which specializes the tracer to a given interpreter at translation time, eliminating dynamic dispatch and skipping recording and heap-cache invalidation for pure operations. We implemented a prototype in RPython and applied it to the PyPy interpreter. On 28 benchmarks from the PyPy benchmark suite, GenExtension reduces tracing time by 11% (geometric mean) and resume data generation (the metadata needed to deoptimize back to the interpreter) by 28%, while producing nearly equivalent machine code and leaving the steady-state performance unchanged (0.99 ×). Turning the tracing speedup into a visible warmup improvement is future work. Yusuke Izawa, Carl Friedrich Bolz-Tereick, Nico Rittinghaus, Hidehiko Masuhara |
MPLR | 2 |
| 2025 | Pydrofoil: Accelerating Sail-Based Instruction Set Simulators
Carl Friedrich Bolz-Tereick, Luke Panayi, Ferdia McKeogh, Tom Spink, Martin Berger 0001 |
ECOOP | 1 |
| 2025 | A Lightweight Method for Generating Multi-Tier JIT Compilation Virtual Machine in a Meta-Tracing Compiler Framework
Yusuke Izawa, Hidehiko Masuhara, Carl Friedrich Bolz-Tereick |
ECOOP | 3 |
| 2025 | Tracing Just-in-Time Compilation for Effects and HandlersabstractEffect handlers are a programming language feature that has recently gained popularity. They allow for nonlocal yet structured control flow and subsume features like generators, exceptions, asynchronicity, etc. However, implementations of effect handlers currently often sacrifice features to enable efficient implementations. Meta-tracing just-in-time (JIT) compilers promise to yield the performance of a compiler by implementing an interpreter. They record execution in a trace, dynamically detect hot loops, and aggressively optimize those using information available at runtime. They excel at optimizing dynamic control flow, which is exactly what effect handlers introduce. We present the first evaluation of tracing JIT compilation specifically for effect handlers. To this end, we developed RPython-based tracing JIT implementations for Eff, Effekt, and Koka by compiling them to a common bytecode format. We evaluate the performance, discuss which classes of effectful programs are optimized well and how our additional optimizations influence performance. We also benchmark against a baseline of state-of-the-art mainstream language implementations. Marcial Gaißert, Carl Friedrich Bolz-Tereick, Jonathan Immanuel Brachthäuser |
Proc. ACM Program. Lang. | 2 |
| 2020 | Type freezing: exploiting attribute type monomorphism in tracing JIT compilersabstractDynamic programming languages continue to increase in popularity. While just-in-time (JIT) compilation can improve the performance of dynamic programming languages, a significant performance gap remains with respect to ahead-of-time compiled languages. Existing JIT compilers exploit type monomorphism through type specialization, and use runtime checks to ensure correctness. Unfortunately, these checks can introduce non-negligible overhead. In this paper, we present type freezing, a novel software solution for exploiting attribute type monomorphism. Type freezing "freezes" type monomorphic attributes of user-defined types, and eliminates the necessity of runtime type checks when performing reads from these attributes. Instead, runtime type checks are done when writing these attributes to validate type monomorphism. We implement type freezing as an extension to PyPy, a state-of-the-art tracing JIT compiler for Python. Our evaluation shows type freezing can improve performance and reduce dynamic instruction count for those applications with a significant number of attribute accesses. Berkin Ilbeyi, Carl Friedrich Bolz-Tereick, Christopher Batten |
CGO | 3 |
| 2017 | Virtual machine warmup blows hot and coldabstractVirtual Machines (VMs) with Just-In-Time (JIT) compilers are traditionally thought to execute programs in two phases: the initial warmup phase determines which parts of a program would most benefit from dynamic compilation, before JIT compiling those parts into machine code; subsequently the program is said to be at a steady state of peak performance. Measurement methodologies almost always discard data collected during the warmup phase such that reported measurements focus entirely on peak performance. We introduce a fully automated statistical approach, based on changepoint analysis, which allows us to determine if a program has reached a steady state and, if so, whether that represents peak performance or not. Using this, we show that even when run in the most controlled of circumstances, small, deterministic, widely studied microbenchmarks often fail to reach a steady state of peak performance on a variety of common VMs. Repeating our experiment on 3 different machines, we found that at most 43.5% of pairs consistently reach a steady state of peak performance. Edd Barrett, Carl Friedrich Bolz-Tereick, Rebecca Killick, Sarah Mount, Laurence Tratt |
Proc. ACM Program. Lang. | 2 |
| 2017 | Sound gradual typing: only mostly deadabstractWhile gradual typing has proven itself attractive to programmers, many systems have avoided sound gradual typing due to the run time overhead of enforcement. In the context of sound gradual typing, both anecdotal and systematic evidence has suggested that run time costs are quite high, and often unacceptable, casting doubt on the viability of soundness as an approach. We show that these overheads are not fundamental, and that with appropriate improvements, just-in-time compilers can greatly reduce the overhead of sound gradual typing . Our study takes benchmarks published in a recent paper on gradual typing performance in Typed Racket (Takikawa et al., POPL 2016) and evaluates them using a experimental tracing JIT compiler for Racket, called Pycket. On typical benchmarks, Pycket is able to eliminate more than 90% of the gradual typing overhead. While our current results are not the final word in optimizing gradual typing, we show that the situation is not dire, and where more work is needed. Pycket's performance comes from several sources, which we detail and measure individually. First, we apply a sophisticated tracing JIT compiler and optimizer, automatically generated in Pycket using the RPython framework originally created for PyPy. Second, we focus our optimization efforts on the challenges posed by run time checks, implemented in Racket by chaperones and impersonators . We introduce representation improvements, including a novel use of hidden classes to optimize these data structures. Spenser Bauman, Carl Friedrich Bolz-Tereick, Jeremy G. Siek, Sam Tobin-Hochstadt |
Proc. ACM Program. Lang. | 2 |
| 2017 | Adaptive just-in-time value class optimization for lowering memory consumption and improving execution time performance
Tobias Pape, Carl Friedrich Bolz-Tereick, Robert Hirschfeld |
Sci. Comput. Program. | 2 |
| 2016 | Fine-grained Language Composition: A Case StudyabstractAlthough run-time language composition is common, it normally takes the form of a crude Foreign Function Interface (FFI). While useful, such compositions tend to be coarse-grained and slow. In this paper we introduce a novel fine-grained syntactic composition of PHP and Python which allows users to embed each language inside the other, including referencing variables across languages. This composition raises novel design and implementation challenges. We show that good solutions can be found to the design challenges; and that the resulting implementation imposes an acceptable performance overhead of, at most, 2.6x. Edd Barrett, Carl Friedrich Bolz-Tereick, Lukas Diekmann, Laurence Tratt |
ECOOP | 2 |
| 2016 | Making an Embedded DBMS JIT-friendlyabstractThis artifact contains: the SQPyte prototype, a JIT for executing SQLite queries; and PyPy-SQPyte, a version of the PyPy Python VM which embeds SQPyte. In addition, a benchmark suite is included, which allows performance comparison against standard SQLite and the Java embedded database H2. Carl Friedrich Bolz-Tereick, Darya Kurilova, Laurence Tratt |
ECOOP | 1 |
| 2015 | Language-independent storage strategies for tracing-JIT-based virtual machinesabstractStorage strategies have been proposed as a run-time optimization for the PyPy Python implementation and have shown promising results for optimizing execution speed and memory requirements. However, it remained unclear whether the approach works equally well in other dynamic languages. Furthermore, while PyPy is based on RPython, a language to write VMs with reusable components such as a tracing just-in-time compiler and garbage collection, the strategies design itself was not generalized to be reusable across languages implemented using that same toolchain. In this paper, we present a general design and implementation for storage strategies and show how they can be reused across different RPython-based languages. We evaluate the performance of our implementation for RSqueak, an RPython-based VM for Squeak/Smalltalk and show that storage strategies may indeed offer performance benefits for certain workloads in other dynamic programming languages.We furthermore evaluate the generality of our implementation by applying it to Topaz, a Ruby VM, and Pycket, a Racket implementation. Tobias Pape, Tim Felgentreff, Robert Hirschfeld, Anton Gulenko, Carl Friedrich Bolz-Tereick |
DLS | 5 |
| 2015 | Pycket: a tracing JIT for a functional languageabstractWe present Pycket, a high-performance tracing JIT compiler for Racket. Pycket supports a wide variety of the sophisticated features in Racket such as contracts, continuations, classes, structures, dynamic binding, and more. On average, over a standard suite of benchmarks, Pycket outperforms existing compilers, both Racket's JIT and other highly-optimizing Scheme compilers. Further, Pycket provides much better performance for Racket proxies than existing systems, dramatically reducing the overhead of contracts and gradual typing. We validate this claim with performance evaluation on multiple existing benchmark suites. The Pycket implementation is of independent interest as an application of the RPython meta-tracing framework (originally created for PyPy), which automatically generates tracing JIT compilers from interpreters. Prior work on meta-tracing focuses on bytecode interpreters, whereas Pycket is a high-level interpreter based on the CEK abstract machine and operates directly on abstract syntax trees. Pycket supports proper tail calls and first-class continuations. In the setting of a functional language, where recursion and higher-order functions are more prevalent than explicit loops, the most significant performance challenge for a tracing JIT is identifying which control flows constitute a loop---we discuss two strategies for identifying loops and measure their impact. Spenser Bauman, Carl Friedrich Bolz-Tereick, Robert Hirschfeld, Vasily Kirilichev, Tobias Pape, Jeremy G. Siek, Sam Tobin-Hochstadt |
ICFP | 2 |
| 2015 | Approaches to interpreter compositionabstractIn this paper, we compose six different Python and Prolog VMs into 4 pairwise compositions: one using C interpreters, one running on the JVM, one using meta-tracing interpreters, and one using a C interpreter and a meta-tracing interpreter. We show that programs that cross the language barrier frequently execute faster in a meta-tracing composition, and that meta-tracing imposes a significantly lower overhead on composed programs relative to mono-language programs. Edd Barrett, Carl Friedrich Bolz-Tereick, Laurence Tratt |
Comput. Lang. Syst. Struct. | 2 |
| 2015 | The impact of meta-tracing on VM design and implementationabstractMost modern languages are implemented using Virtual Machines (VMs). While the best VMs use Just-In-Time (JIT) compilers to achieve good performance, JITs are costly to implement, and few VMs therefore come with one. The RPython language allows tracing JIT VMs to be automatically created from an interpreter, changing the economics of VM implementation. In this paper, we explain, through two concrete VMs, how meta-tracing RPython VMs can be designed and optimised, and, experimentally, the performance levels one might reasonably expect from them. Carl Friedrich Bolz-Tereick, Laurence Tratt |
Sci. Comput. Program. | 1 |
| 2013 | Storage strategies for collections in dynamically typed languagesabstractDynamically typed language implementations often use more memory and execute slower than their statically typed cousins, in part because operations on collections of elements are unoptimised. This paper describes storage strategies, which dynamically optimise collections whose elements are instances of the same primitive type. We implement storage strategies in the PyPy virtual machine, giving a performance increase of 18% on wide-ranging benchmarks of real Python programs. We show that storage strategies are simple to implement, needing only 1500LoC in PyPy, and have applicability to a wide range of virtual machines. Carl Friedrich Bolz-Tereick, Lukas Diekmann, Laurence Tratt |
OOPSLA | 1 |
| 2012 | Loop-aware optimizations in PyPy's tracing JITabstractOne of the nice properties of a tracing just-in-time compiler (JIT) is that many of its optimizations are simple, requiring one forward pass only. This is not true for loop-invariant code motion which is a very important optimization for code with tight kernels. Especially for dynamic languages that typically perform quite a lot of loop invariant type checking, boxed value unwrapping and virtual method lookups. In this paper we explain a scheme pioneered within the context of the LuaJIT project for making basic optimizations loop-aware by using a simple pre-processing step on the trace without changing the optimizations themselves. Håkan Ardö, Carl Friedrich Bolz-Tereick, Maciej Fijalkowski |
DLS | 2 |
| 2011 | Allocation removal by partial evaluation in a tracing JITabstractThe performance of many dynamic language implementations suffers from high allocation rates and runtime type checks. This makes dynamic languages less applicable to purely algorithmic problems, despite their growing popularity. In this paper we present a simple compiler optimization based on online partial evaluation to remove object allocations and runtime type checks in the context of a tracing JIT. We evaluate the optimization using a Python VM and find that it gives good results for all our (real-life) benchmarks. Carl Friedrich Bolz-Tereick, Antonio Cuni, Maciej Fijalkowski, Michael Leuschel, Samuele Pedroni, Armin Rigo |
PEPM | 1 |
| 2010 | Towards a jitting VM for prolog executionabstractMost Prolog implementations are implemented in low-level languages such as C and are based on a variation of the WAM instruction set, which enhances their performance but makes them hard to write. In addition, many of the more dynamic features of Prolog (like assert), despite their popularity, are not well supported. We present a high-level continuation-based Prolog interpreter based on the PyPy project. The PyPy project makes it possible to easily and efficiently implement dynamic languages. It provides tools that automatically generate a just-in-time compiler for a given interpreter of the target language, by using partial evaluation techniques. The resulting Prolog implementation is surprisingly efficient: it clearly outperforms existing interpreters of Prolog in high-level languages such as Java. Moreover, on some benchmarks, our system outperforms state-of-the-art WAM-based Prolog implementations. Our paper aims to show that declarative languages such as Prolog can indeed benefit from having a just-in-time compiler and that PyPy can form the basis for implementing programming languages other than Python. Carl Friedrich Bolz-Tereick, Michael Leuschel, David Schneider 0001 |
PPDP | 1 |
| 2009 | Towards Just-In-Time Partial Evaluation of Prolog
Carl Friedrich Bolz-Tereick, Michael Leuschel, Armin Rigo |
LOPSTR | 1 |