VLDB 2026 Research / reviewers in the wild / expert
Michal Cierniak
dblp:36/4240
· DBLP profile ↗
12ranked-venue papers
8as first author
0since 2021 · last 2005
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-authorSoftware engineering, systems software and programming languages · 6 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
5 papers |
Runtime systems and virtual machines · 60% Compilers and program optimization · 40% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Memory systems · 61% Parallel and multicore computing · 35% Interconnection networks and networks-on-chip · 5% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation |
0.1 | 3 | 2000 | Practicing JUDO: Java under dynamic optimizations · PLDI 2000 Fast, Effective Code Generation in a Just-In-Time Java Compiler · PLDI 1998 Support for Garbage Collection at Every Instruction in a Java Compiler · PLDI 1999 |
Runtime systems and virtual machines
garbage collection |
0.1 | 2 | 2000 | Practicing JUDO: Java under dynamic optimizations · PLDI 2000 Support for Garbage Collection at Every Instruction in a Java Compiler · PLDI 1999 |
Runtime systems and virtual machines › virtual machine implementation
java virtual machine |
0.0 | 1 | 2000 | Practicing JUDO: Java under dynamic optimizations · PLDI 2000 |
Compilers and program optimization
compiler optimization |
0.0 | 1 | 1998 | Fast, Effective Code Generation in a Just-In-Time Java Compiler · PLDI 1998 |
Memory systems
cache coherence |
0.0 | 1 | 1997 | VM-Based Shared Memory on Low-Latency, Remote-Memory-Access Networks · ISCA 1997 |
Memory systems › shared memory
distributed shared memory |
0.0 | 1 | 1997 | VM-Based Shared Memory on Low-Latency, Remote-Memory-Access Networks · ISCA 1997 |
Memory systems › memory consistency › memory consistency model
release consistency |
0.0 | 1 | 1997 | VM-Based Shared Memory on Low-Latency, Remote-Memory-Access Networks · ISCA 1997 |
Memory systems › shared memory › distributed shared memory
software distributed shared memory |
0.0 | 1 | 1997 | VM-Based Shared Memory on Low-Latency, Remote-Memory-Access Networks · ISCA 1997 |
Compilers and program optimization › instruction scheduling
compile-time scheduling |
0.0 | 1 | 1995 | Loop Scheduling for Heterogeneity · HPDC 1995 |
Compilers and program optimization › memory optimization
data locality optimization |
0.0 | 1 | 1995 | Unifying Data and Control Transformations for Distributed Shared Memory Machines · PLDI 1995 |
Compilers and program optimization
loop transformation |
0.0 | 1 | 1995 | Unifying Data and Control Transformations for Distributed Shared Memory Machines · PLDI 1995 |
Parallel and multicore computing
load balancing |
0.0 | 1 | 1995 | Loop Scheduling for Heterogeneity · HPDC 1995 |
Parallel and multicore computing
locality optimization |
0.0 | 1 | 1995 | Unifying Data and Control Transformations for Distributed Shared Memory Machines · PLDI 1995 |
Parallel and multicore computing › parallel scheduling
parallel loop scheduling |
0.0 | 1 | 1995 | Loop Scheduling for Heterogeneity · HPDC 1995 |
Methods — techniques the papers use, named apart from their topics
data transformation · 0.0control transformation · 0.0compiler algorithm · 0.0just-in-time compilation · 0.0dynamic patching · 0.0register allocation · 0.0common subexpression elimination · 0.0array bounds check elimination · 0.0simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2005 | The Open Runtime Platform: a flexible high-performance managed runtime environmentabstractAbstract The Open Runtime Platform (ORP) is a high‐performance managed runtime environment (MRTE) that features exact generational garbage collection, fast thread synchronization, and multiple coexisting just‐in‐time compilers (JITs). ORP was designed for flexibility in order to support experiments in dynamic compilation, garbage collection, synchronization, and other technologies. It can be built to run either Java or Common Language Infrastructure (CLI) applications, to run under the Windows or Linux operating systems, and to run on the IA‐32 or Itanium processor family (IPF) architectures. Achieving high performance in a MRTE presents many challenges, particularly when flexibility is a major goal. First, to enable the use of different garbage collectors and JITs, each component must be isolated from the rest of the environment through a well‐defined software interface. Without careful attention, this isolation could easily harm performance. Second, MRTEs have correctness and safety requirements that traditional languages such as C++ lack. These requirements, including null pointer checks, array bounds checks, and type checks, impose additional runtime overhead. Finally, the dynamic nature of MRTEs makes some traditional compiler optimizations, such as devirtualization of method calls, more difficult to implement or more limited in applicability. To get full performance, JITs and the core virtual machine (VM) must cooperate to reduce or eliminate (where possible) these MRTE‐specific overheads. In this paper, we describe the structure of ORP in detail, paying particular attention to how it supports flexibility while preserving high performance. We describe the interfaces between the garbage collector, the JIT, and the core VM; how these interfaces enable multiple garbage collectors and JITs without sacrificing performance; and how they allow the JIT and the core VM to reduce or eliminate MRTE‐specific performance issues. Copyright © 2005 John Wiley & Sons, Ltd. Michal Cierniak, Marsha Eng, Neal Glew, Brian T. Lewis, James M. Stichnoth |
Concurr. Pract. Exp. | 1 |
| 2004 | Improving 64-Bit Java IPF Performance by Compressing Heap Referencesabstract64-bit processor architectures like the Intel/spl reg/ Itanium/spl reg/ processor family are designed for large applications that need large memory addresses. When running applications that fit within a 32-bit address space, 64-bit CPUs are at a disadvantage compared to 32-bit CPUs because of the larger memory footprints for their data. This results in worse cache and TLB utilization, and consequently lower performance because of increased miss ratios. This paper considers software techniques for virtual machines that allow 32-bit pointers to be used on 64-bit CPUs for managed runtime applications that do not need the full 64-bit address space. We describe our pointer compression techniques and discuss our experience implementing these for Java applications. In addition, we give performance results with our techniques for both the SPEC JVM98 and SPEC JBB2000 benchmarks. We demonstrate a 12% performance improvement on SPEC JBB2000 and a reduction in the number of garbage collections required for a given heap size. Ali-Reza Adl-Tabatabai, Jay Bharadwaj, Michal Cierniak, Marsha Eng, Jesse Fang, Brian T. Lewis, Brian R. Murphy, James M. Stichnoth |
CGO | 3 |
| 2000 | Practicing JUDO: Java under dynamic optimizationsabstractA high-performance implementation of a Java Virtual Machine (JVM) consists of efficient implementation of Just-In-Time (JIT) compilation, exception handling, synchronization mechanism, and garbage collection (GC). These components are tightly coupled to achieve high performance. In this paper, we present some static anddynamic techniques implemented in the JIT compilation and exception handling of the Microprocessor Research Lab Virtual Machine (MRL VM), i.e., lazy exceptions, lazy GC mapping, dynamic patching, and bounds checking elimination. Our experiments used IA-32 as the hardware platform, but the optimizations can be generalized to other architectures. Michal Cierniak, Guei-Yuan Lueh, James M. Stichnoth |
PLDI | 1 |
| 1999 | Support for Garbage Collection at Every Instruction in a Java CompilerabstractA high-performance implementation of a Java Virtual Machine1 requires a compiler to translate Java bytecodes into native instructions, as well as an advanced garbage collector (e.g., copying or generational). When the Java heap is exhausted and the garbage collector executes, the compiler must report to the garbage collector all live object references contained in physical registers and stack locations. Typical compilers only allow certain instructions (e.g., call instructions and backward branches) to be GC-safe; if GC happens at some other instruction, the compiler may need to advance execution to the next GC-safe point. Until now, no one has ever attempted to make every compiler-generated instruction GC-safe, due to the perception that recording this information would require too much space. This kind of support could improve the GC performance in multithreaded applications. We show how to use simple compression techniques to reduce the size of the GC map to about 20% of the generated code size, a result that is competitive with the best previously published results. In addition, we extend the work of Agesen, Detlefs, and Moss, regarding the so-called "JSR Problem" (the single exception to Java's type safety property), in a way that eliminates the need for extra runtime overhead in the generated code. James M. Stichnoth, Guei-Yuan Lueh, Michal Cierniak |
PLDI | 3 |
| 1998 | Fast, Effective Code Generation in a Just-In-Time Java CompilerabstractA Just-In-Time (JIT) Java compiler produces native code from Java byte code instructions during program execution. As such, compilation speed is more important in a Java JIT compiler than in a traditional compiler, requiring optimization algorithms to be lightweight and effective. We present the structure of a Java JIT compiler for the Intel Architecture, describe the lightweight implementation of JIT compiler optimizations (e.g., common subexpression elimination, register allocation, and elimination of array bounds checking), and evaluate the performance benefits and tradeoffs of the optimizations. This JIT compiler has been shipped with version 2.5 of Intel's VTune for Java product. Ali-Reza Adl-Tabatabai, Michal Cierniak, Guei-Yuan Lueh, Vishesh M. Parikh, James M. Stichnoth |
PLDI | 2 |
| 1997 | VM-Based Shared Memory on Low-Latency, Remote-Memory-Access NetworksabstractRecent technological advances have produced network interfaces that provide users with very low-latency access to the memory of remote machines. We examine the impact of such networks on the implementation and performance of software DSM. Specifically, we compare two DSM systems---Cashmere and TreadMarks---on a 32-processor DEC Alpha cluster connected by a Memory Channel network.Both Cashmere and TreadMarks use virtual memory to maintain coherence on pages, and both use lazy, multi-writer release consistency. The systems differ dramatically, however, in the mechanisms used to track sharing information and to collect and merge concurrent updates to a page, with the result that Cashmere communicates much more frequently, and at a much finer grain.Our principal conclusion is that low-latency networks make DSM based on fine-grain communication competitive with more coarse-grain approaches, but that further hardware improvements will be needed before such systems can provide consistently superior performance. In our experiments, Cashmere scales slightly better than TreadMarks for applications with false sharing. At the same time, it is severely constrained by limitations of the current Memory Channel hardware. In general, performance is better for TreadMarks. Leonidas I. Kontothanassis, Galen C. Hunt, Robert Stets, Nikos Hardavellas, Michal Cierniak, Srinivasan Parthasarathy 0001, Wagner Meira Jr., Sandhya Dwarkadas, Michael L. Scott |
ISCA | 5 |
| 1997 | Compile-Time Scheduling Algorithms for a Heterogeneous Network of WorkstationsabstractIn this paper, we study the problem of scheduling parallel loops at compile time for a heterogeneous network of workstations. We consider heterogeneity in various aspects of parallel programming: program, processor, memory and network. A heterogeneous program has parallel loops with different amounts of work being done in each iteration; heterogeneous processors have different speeds; heterogeneous memory refers to the different amounts of user-available memory on the machines and a heterogeneous network has different communication costs between processors. We propose a simple yet comprehensive model for use in compiling for a network of processors, and develop compiler algorithms for generating optimal and near-optimal schedules of loops for load balancing, communication optimizations, network contention and memory heterogeneity. Experiments show that a significant performance improvement is achieved using our techniques. Michal Cierniak, Mohammed J. Zaki, Wei Li 0015 |
Comput. J. | 1 |
| 1997 | Just-in-Time Pptimizations for High-Performance Java ProgramsabstractOur previous experience with an off-line Java optimizer has shown that some traditional algorithms used in compilers are too slow for a JIT compiler. In this paper we propose and implement faster ways of performing analyses needed for our optimizations. For instance, we have replaced reaching definitions with constant values and loop induction variables with loop-defined variables. As a result, our JIT compiler, Briki, is very fast, so that its running time is negligible even when the data sets used with our benchmarks result in execution times of only a few seconds. The impact for the same benchmarks running on more realistic problem sizes would be even smaller. Currently the speedups resulting from applying our optimizations are between 10% and 20%, but when the JIT compiler performs standard optimizations which are absent in the current version of the JIT compiler used by us, thespeedups should be similar to the ones observed for Fortran programs – up to 50%. © 1997 John Wiley & Sons, Ltd. Michal Cierniak, Wei Li 0015 |
Concurr. Pract. Exp. | 1 |
| 1997 | Optimizing Java bytecodesabstractWe have developed a research compiler for Java class files. The compiler, which we call Briki, is designed to test new compilation techniques. We focus on optimizations which are only possible or much easier to perform on a high-level intermediate representation. We have designed such a representation, JavaIR, and have written a front-end which recovers the high-level structure from the information from the class file. Some of the high-level optimizations can be performed by the Java compiler which produces the class file. There is, however, a set of machine-dependent optimizations which have to be customized for the specific architecture and so can only be performed when the machine code is generated from the bytecodes, e.g. in a just-in-time (JIT) compiler. We choose memory hierarchy optimizations as an example of machine-dependent techniques. We show that there is an intersection of the set of machine-dependent optimizations and the set of high-level optimizations. One such example is array remapping, which requires multi-dimensional array references which are not present in the bytecodes and at the same time requires information about memory organization and the mapping of bytecodes to machine instructions. We develop a set of optimizations for accessing array elements and object fields and show their impact on a set of benchmarks which we run on two machines with a JIT compiler. The execution times are reduced by as much as 50% and we argue that the improvement could be even higher with a more mature JIT technology. © 1997 John Wiley & Sons, Ltd. Michal Cierniak, Wei Li 0015 |
Concurr. Pract. Exp. | 1 |
| 1997 | A Portable Browser for Performance ProgrammingabstractWe present jCITE, a performance tuning tool for scientific applications. By combining the static information produced by the compiler with the profile data from real program execution, jCITE can be used to quickly understand the performance bottlenecks. The compiler information allows great understanding of what optimizations have been performed. The user can also find out which optimizations have not been applied and why. Platform independence makes Java the ideal implementation platform for our tool. SGI users can have the same performance analysis tool on all platforms. You can run jCITE on an SGI desktop machine (O2, Octane, Indy) or on a PC running Windows to analyze and optimize the performance of a scientific code running on an SGI Challenge or an SGI Origin machine. In our experiments we were able to significantly speed up some SPEC95 applications in a few days or even a few hours without any prior knowledge of those applications. A longer version of this paper describes all our experiments with jCITE. © 1997 John Wiley & Sons, Ltd. Michal Cierniak, Suresh Srinivas |
Concurr. Pract. Exp. | 1 |
| 1995 | Loop Scheduling for HeterogeneityabstractIn this paper we study the problem of scheduling parallel loops at compile-time for a heterogeneous network of machines. We consider heterogeneity in three aspects of parallel programming: program, processor and network. A heterogeneous program has parallel loops with different amount of work in each iteration; heterogeneous processors have different speeds; and a heterogeneous network has different cost of communication between processors. We propose a simple yet comprehensive model for use in compiling for a network of processors, and develop compiler algorithms for generating optimal and sub-optimal schedules of loops for load balancing, communication optimizations and network contention. Experiments show that a significant improvement of performance is achieved using our techniques. Michal Cierniak, Wei Li 0015, Mohammed J. Zaki |
HPDC | 1 |
| 1995 | Unifying Data and Control Transformations for Distributed Shared Memory MachinesabstractWe present a unified approach to locality optimization that employs both data and control transformations. Data transformations include changing the array layout in memory. Control transformations involve changing the execution order of programs. We have developed new techniques for compiler optimizations for distributed shared-memory machines, although the same techniques can be used for sequential machines with a memory hierarchy. Michal Cierniak, Wei Li 0015 |
PLDI | 1 |