VLDB 2026 Research / reviewers in the wild / expert
Artur Klauser
dblp:51/6212
· DBLP profile ↗
8ranked-venue papers
2as first author
0since 2021 · last 2006
0000-0002-9465-1239ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Processor architecture and microarchitecture · 74% Energy-efficient computing · 10% Performance modeling and evaluation · 6% | |
| Software engineering, system software, and programming languages
1 paper |
Program analysis · 100% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program analysis › dynamic analysis › instrumentation
binary instrumentation |
0.1 | 1 | 2005 | Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005 |
Program analysis › dynamic analysis
dynamic instrumentation |
0.1 | 1 | 2005 | Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005 |
Processor architecture and microarchitecture
branch prediction |
0.1 | 3 | 1999 | Instruction Fetch Mechanisms for Multipath Execution Processors · MICRO 1999 Confidence Estimation for Speculation Control · ISCA 1998 Pipeline Gating: Speculation Control for Energy Reduction · ISCA 1998 |
Processor architecture and microarchitecture › speculative execution
multi-path execution |
0.0 | 2 | 1999 | Instruction Fetch Mechanisms for Multipath Execution Processors · MICRO 1999 Selective Eager Execution on the PolyPath Architecture · ISCA 1998 |
Processor architecture and microarchitecture › speculation
speculation control |
0.0 | 2 | 1998 | Pipeline Gating: Speculation Control for Energy Reduction · ISCA 1998 Confidence Estimation for Speculation Control · ISCA 1998 |
Processor architecture and microarchitecture
speculative execution |
0.0 | 2 | 1998 | Pipeline Gating: Speculation Control for Energy Reduction · ISCA 1998 Selective Eager Execution on the PolyPath Architecture · ISCA 1998 |
Processor architecture and microarchitecture › branch prediction
branch misprediction |
0.0 | 1 | 1999 | Instruction Fetch Mechanisms for Multipath Execution Processors · MICRO 1999 |
Processor architecture and microarchitecture
instruction fetch |
0.0 | 1 | 1999 | Instruction Fetch Mechanisms for Multipath Execution Processors · MICRO 1999 |
Processor architecture and microarchitecture › branch prediction
branch confidence estimation |
0.0 | 1 | 1998 | Confidence Estimation for Speculation Control · ISCA 1998 |
Parallel and multicore computing
eager execution |
0.0 | 1 | 1998 | Selective Eager Execution on the PolyPath Architecture · ISCA 1998 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 1 | 1998 | Selective Eager Execution on the PolyPath Architecture · ISCA 1998 |
Energy-efficient computing › power management › fine-grain power management
pipeline gating |
0.0 | 1 | 1998 | Pipeline Gating: Speculation Control for Energy Reduction · ISCA 1998 |
Energy-efficient computing › power management
processor power reduction |
0.0 | 1 | 1998 | Pipeline Gating: Speculation Control for Energy Reduction · ISCA 1998 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 1 | 1998 | Selective Eager Execution on the PolyPath Architecture · ISCA 1998 |
Performance modeling and evaluation
profiling |
0.0 | 1 | 2005 | Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005 |
Storage systems › storage management
storage reclamation |
0.0 | 1 | 1996 | Semi-automatic, Self-adaptive Control of Garbage Collection Rates in Object Databases · SIGMOD Conference 1996 |
Processor architecture and microarchitecture
multithreading |
0.0 | 1 | 1999 | Instruction Fetch Mechanisms for Multipath Execution Processors · MICRO 1999 |
Performance modeling and evaluation › simulation › processor simulation
pipeline simulation |
0.0 | 1 | 1998 | Confidence Estimation for Speculation Control · ISCA 1998 |
Storage systems › data management
object database |
0.0 | 1 | 1996 | Semi-automatic, Self-adaptive Control of Garbage Collection Rates in Object Databases · SIGMOD Conference 1996 |
Methods — techniques the papers use, named apart from their topics
liveness analysis · 0.1instruction scheduling · 0.1inlining · 0.1dynamic compilation · 0.1register reallocation · 0.1register re-allocation · 0.1simulation · 0.0register renaming · 0.0pipeline simulation · 0.0instruction tagging · 0.0execution-driven simulation · 0.0trace-driven simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2006 | A dynamic binary instrumentation engine for the ARM architectureabstractDynamic binary instrumentation (DBI) is a powerful technique for analyzing the runtime behavior of software. While numerous DBI frameworks have been developed for general-purpose architectures, work on DBI frameworks for embedded architectures has been fairly limited. In this paper, we describe the design, implementation, and applications of the ARM version of Pin, a dynamic instrumentation system from Intel. In particular, we highlight the design decisions that are geared toward the space and processing limitations of embedded systems. Pin for ARM is publicly available and is shipped with dozens of sample plug-in instrumentation tools. It has been downloaded over 500 times since its release. Kim M. Hazelwood, Artur Klauser |
CASES | 2 |
| 2005 | Pin: building customized program analysis tools with dynamic instrumentationabstractRobust and powerful software instrumentation tools are essential for program analysis tasks such as profiling, performance evaluation, and bug detection. To meet this need, we have developed a new instrumentation system called Pin. Our goals are to provide easy-to-use, portable, transparent, and efficient instrumentation. Instrumentation tools (called Pintools) are written in C/C++ using Pin's rich API. Pin follows the model of ATOM, allowing the tool writer to analyze an application at the instruction level without the need for detailed knowledge of the underlying instruction set. The API is designed to be architecture independent whenever possible, making Pintools source compatible across different architectures. However, a Pintool can access architecture-specific details when necessary. Instrumentation with Pin is mostly transparent as the application and Pintool observe the application's original, uninstrumented behavior. Pin uses dynamic compilation to instrument executables while they are running. For efficiency, Pin uses several techniques, including inlining, register re-allocation, liveness analysis, and instruction scheduling to optimize instrumentation. This fully automated approach delivers significantly better instrumentation performance than similar tools. For example, Pin is 3.3x faster than Valgrind and 2x faster than DynamoRIO for basic-block counting. To illustrate Pin's versatility, we describe two Pintools in daily use to analyze production software. Pin is publicly available for Linux platforms on four architectures: IA32 (32-bit x86), EM64T (64-bit x86), Itanium®, and ARM. In the ten months since Pin 2 was released in July 2004, there have been over 3000 downloads from its website. Chi-Keung Luk, Robert S. Cohn, Robert Muth, Harish Patil, Artur Klauser, P. Geoffrey Lowney, Steven Wallace, Vijay Janapa Reddi, Kim M. Hazelwood |
PLDI | 5 |
| 2000 | An infrastructure for generating and sharing experimental workloads for persistent object systemsabstractPerformance evaluation of persistent object system implementations requires the use and evaluation of experimental workloads. Such workloads include a schema describing how the data are related, and application behaviors that capture how the data are manipulated over time. In this paper, we describe an infrastructure for generating and sharing experimental workloads to be used in evaluating the performance of persistent object system implementations. The infrastructure consists of a toolkit that aids the analyst in modeling and instrumenting experimental workloads, and a trace format that allows the analyst to easily reuse and share the workloads. Our infrastructure provides the following benefits: the process of building new experiments for analysis is made easier; experiments to evaluate the performance of implementations can be conducted and reproduced with less effort; and pertinent information can be gathered in a cost-effective manner. We describe the two major components of this infrastructure, the trace format and the toolkit. We also describe our experiences using these components to model, instrument, and experiment with the OO7 benchmark. Copyright © 2000 John Wiley & Sons, Ltd. Thorna O. Humphries, Artur Klauser, Alexander L. Wolf, Benjamin G. Zorn |
Softw. Pract. Exp. | 2 |
| 1999 | Instruction Fetch Mechanisms for Multipath Execution ProcessorsabstractBranch mispredictions can have a major performance impact on high-performance processors. Multipath execution has recently been introduced to help limit the misprediction penalties incurred by branches that are difficult to predict. This paper presents efficient instruction fetch architecture designs for these multipath processor execution cores. We evaluate a number of design trade-offs for the first-level instruction cache and the multipath PC fetch arbiter. Furthermore we evaluate the effect of additional bandwidth limitations imposed by the processor frontend pipeline. Our results show that instruction fetch support for efficient multipath execution can be achieved with realizable hardware implementations. In addition, we show that the best performing instruction fetch designs for multipath execution and multithreaded processors are likely to differ, since both designs optimize the processor for different performance goals (minimal execution time vs maximal throughput). Artur Klauser, Dirk Grunwald |
MICRO | 1 |
| 1998 | Confidence Estimation for Speculation ControlabstractModern processors improve instruction level parallelism by speculation. The outcome of data and control decisions is predicted, and the operations are speculatively executed and only committed if the original predictions were correct. There are a number of other ways that processor resources could be used, such as threading or eager execution. As the use of speculation increases, we believe more processors will need some form of speculation control to balance the benefits of speculation against other possible activities. Confidence estimation is one technique that can be exploited by architects for speculation control. In this paper, we introduce performance metrics to compare confidence estimation mechanisms, and argue that these metrics are appropriate for speculation control. We compare a number of confidence estimation mechanisms, focusing on mechanisms that have a small implementation cost and gain benefit by exploiting characteristics of branch predictors, such as clustering of mispredicted branches. We compare the performance of the different confidence estimation methods using detailed pipeline simulations. Using these simulations, we show how to improve some confidence estimators, providing better insight for future investigations comparing and applying confidence estimators. Dirk Grunwald, Artur Klauser, Srilatha Manne, Andrew R. Pleszkun |
ISCA | 2 |
| 1998 | Selective Eager Execution on the PolyPath ArchitectureabstractControl-flow misprediction penalties are a major impediment to high performance in wide-issue superscalar processors. In this paper we present Selective Eager Execution (SEE), an execution model to overcome mis-speculation penalties by executing both paths after diffident branches. We present the micro-architecture of the PolyPath processor which is an extension of an aggressive superscalar out-of-order architecture. The PolyPath architecture uses a novel instruction tagging and register renaming mechanism to execute instructions from multiple paths simultaneously in the same processor pipeline, while retaining maximum resource availability for single-path code sequences. Results of our execution-driven, pipeline-level simulations show that SEE can improve performance by as much as 36% for the go benchmark, and an average of 14% on SPECint95, when compared to a normal superscalar, out-of-order speculative execution, monopath processor. Moreover our architectural model is both elegant and practical to implement, using a small amount of additional state and control logic. Artur Klauser, Abhijit Paithankar, Dirk Grunwald |
ISCA | 1 |
| 1998 | Pipeline Gating: Speculation Control for Energy ReductionabstractBranch prediction has enabled microprocessors to increase instruction level parallelism (ILP) by allowing programs to speculatively execute beyond control boundaries. Although speculative execution is essential for increasing the instructions per cycle (IPC), it does come at a cost. A large amount of unnecessary work results from wrong-path instructions entering the pipeline due to branch misprediction. Results generated with the SimpleScalar tool set using a 4-way issue pipeline and various branch predictors show an instruction overhead of 16% to 105% for event instruction committed. The instruction overhead will increase in the future as processors use more aggressive speculation and wider issue widths. In this paper we present an innovative method for power reduction ,which, unlike previous work that sacrificed flexibility or performance reduces power in high-performance microprocessors without impacting performance. In particular we introduce a hardware mechanism called pipeline gating to control rampant speculation in the pipeline. We present inexpensive mechanisms for determining when a branch is likely to mispredict, and for stopping wrong-path instructions from entering the pipeline. Results show up to a 38% reduction in wrong-path instructions with a negligible performance loss (/spl ap/1%). Best of all, even in programs with a high branch prediction accuracy, performance does not noticeable degrade. Our analysis indicates that there is little risk in implementing this method in existing processors since it does not impact performance and can benefit energy reduction. Srilatha Manne, Artur Klauser, Dirk Grunwald |
ISCA | 2 |
| 1996 | Semi-automatic, Self-adaptive Control of Garbage Collection Rates in Object DatabasesabstractA fundamental problem in automating object database storage reclamation is determining how often to perform garbage collection. We show that the choice of collection rate can have a significant impact on application performance and that the "best" rate depends on the dynamic behavior of the application, tempered by the particular performance goals of the user. We describe two semi-automatic, selfadaptive policies for controlling collection rate that we have developed to address the problem. Using tracedriven simulations, we evaluate the performance of the policies on a test database application that demonstrates two distinct reclustering behaviors. Our results show that the policies are effective at achieving user-specified levels of I/O operations and database garbage percentage. We also investigate the sensitivity of the policies over a range of object connectivities. The evaluation demonstrates that semi-automatic, self-adaptive policies are a practical means for flexibly controllin... Jonathan E. Cook 0001, Artur Klauser, Alexander L. Wolf, Benjamin G. Zorn |
SIGMOD Conference | 2 |