EDBT 2026 Demo / reviewers in the wild / expert
P. Geoffrey Lowney
dblp:07/3011
· DBLP profile ↗
9ranked-venue papers
2as first author
0since 2021 · last 2005
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 6 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
5 papers |
Program analysis · 55% Compilers and program optimization · 41% Software maintenance and evolution · 2% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Processor architecture and microarchitecture · 63% Performance modeling and evaluation · 20% Memory systems · 8% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program analysis › dynamic analysis › instrumentation
binary instrumentation |
0.1 | 1 | 2005 | Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005 |
Program analysis › dynamic analysis
dynamic instrumentation |
0.1 | 1 | 2005 | Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005 |
Processor architecture and microarchitecture › instruction set architecture
vector extension |
0.0 | 1 | 2002 | Tarantula: A Vector Extension to the Alpha Architecture · ISCA 2002 |
Processor architecture and microarchitecture
vector processor |
0.0 | 1 | 2002 | Tarantula: A Vector Extension to the Alpha Architecture · ISCA 2002 |
Compilers and program optimization
code layout optimization |
0.0 | 1 | 2001 | Code layout optimizations for transaction processing workloads · ISCA 2001 |
Performance modeling and evaluation
profiling |
0.0 | 1 | 2005 | Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005 |
Compilers and program optimization
compiler optimization |
0.0 | 1 | 1996 | Hot Cold Optimization of Large Windows/NT Applications · MICRO 1996 |
Compilers and program optimization
code motion |
0.0 | 1 | 1994 | Avoidance and Suppression of Compensation Code in a Trace Scheduling Compiler · ACM Trans. Program. Lang. Syst. 1994 |
Compilers and program optimization
instruction scheduling |
0.0 | 1 | 1994 | Avoidance and Suppression of Compensation Code in a Trace Scheduling Compiler · ACM Trans. Program. Lang. Syst. 1994 |
Compilers and program optimization › instruction scheduling
trace scheduling |
0.0 | 1 | 1994 | Avoidance and Suppression of Compensation Code in a Trace Scheduling Compiler · ACM Trans. Program. Lang. Syst. 1994 |
High-performance computing
scientific computing systems |
0.0 | 1 | 2002 | Tarantula: A Vector Extension to the Alpha Architecture · ISCA 2002 |
Transaction processing and concurrency control
OLTP |
0.0 | 1 | 2001 | Code layout optimizations for transaction processing workloads · ISCA 2001 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2001 | Code layout optimizations for transaction processing workloads · ISCA 2001 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 1 | 1994 | Avoidance and Suppression of Compensation Code in a Trace Scheduling Compiler · ACM Trans. Program. Lang. Syst. 1994 |
Processor architecture and microarchitecture › instruction-level parallelism
VLIW |
0.0 | 1 | 1994 | Avoidance and Suppression of Compensation Code in a Trace Scheduling Compiler · ACM Trans. Program. Lang. Syst. 1994 |
Programming languages and type systems › programming paradigms
array programming |
0.0 | 1 | 1981 | Carrier Arrays: An Idiom-Preserving Extension to APL · POPL 1981 |
Programming languages and type systems
language design |
0.0 | 1 | 1981 | Carrier Arrays: An Idiom-Preserving Extension to APL · POPL 1981 |
Methods — techniques the papers use, named apart from their topics
liveness analysis · 0.1instruction scheduling · 0.1inlining · 0.1dynamic compilation · 0.1register reallocation · 0.1register re-allocation · 0.1simulation · 0.0trace scheduling · 0.0loop unrolling · 0.0global flow analysis · 0.0profile-guided optimization · 0.0partition-based parallel application · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2005 | Pin: building customized program analysis tools with dynamic instrumentationabstractRobust and powerful software instrumentation tools are essential for program analysis tasks such as profiling, performance evaluation, and bug detection. To meet this need, we have developed a new instrumentation system called Pin. Our goals are to provide easy-to-use, portable, transparent, and efficient instrumentation. Instrumentation tools (called Pintools) are written in C/C++ using Pin's rich API. Pin follows the model of ATOM, allowing the tool writer to analyze an application at the instruction level without the need for detailed knowledge of the underlying instruction set. The API is designed to be architecture independent whenever possible, making Pintools source compatible across different architectures. However, a Pintool can access architecture-specific details when necessary. Instrumentation with Pin is mostly transparent as the application and Pintool observe the application's original, uninstrumented behavior. Pin uses dynamic compilation to instrument executables while they are running. For efficiency, Pin uses several techniques, including inlining, register re-allocation, liveness analysis, and instruction scheduling to optimize instrumentation. This fully automated approach delivers significantly better instrumentation performance than similar tools. For example, Pin is 3.3x faster than Valgrind and 2x faster than DynamoRIO for basic-block counting. To illustrate Pin's versatility, we describe two Pintools in daily use to analyze production software. Pin is publicly available for Linux platforms on four architectures: IA32 (32-bit x86), EM64T (64-bit x86), Itanium®, and ARM. In the ten months since Pin 2 was released in July 2004, there have been over 3000 downloads from its website. Chi-Keung Luk, Robert S. Cohn, Robert Muth, Harish Patil, Artur Klauser, P. Geoffrey Lowney, Steven Wallace, Vijay Janapa Reddi, Kim M. Hazelwood |
PLDI | 6 |
| 2004 | Ispike: A Post-link Optimizer for the Intel®Itanium®ArchitectureabstractIspike is a post-link optimizer developed for the Intel/spl reg/ Itanium Processor Family (IPF) processors. The IPF architecture poses both opportunities and challenges to post-link optimizations. IPF offers a rich set of performance counters to collect detailed profile information at a low cost, which is essential to post-link optimization being practical. At the same time, the predication and bundling features on IPF make post-link code transformation more challenging than on other architectures. In Ispike, we have implemented optimizations like code layout, instruction prefetching, data layout, and data prefetching that exploit the IPF advantages, and strategies that cope with the IPF-specific challenges. Using SPEC CINT2000 as benchmarks, we show that Ispike improves performance by as much as 40% on the ltanium/spl reg/2 processor, with average improvement of 8.5% and 9.9% over executables generated by the Intel/spl reg/ Electron compiler and by the Gcc compiler, respectively. We also demonstrate that statistical profiles collected via IPF performance counters and complete profiles collected via instrumentation produce equal performance benefit, but the profiling overhead is significantly lower for performance counters. Chi-Keung Luk, Robert Muth, Harish Patil, Robert S. Cohn, P. Geoffrey Lowney |
CGO | 5 |
| 2002 | Profile-guided post-link stride prefetchingabstractData prefetching is an e ective approach to addressing the memory latency problem. While a few processors have implemented hardware-based data prefetching, the majority of modern processors support data-prefetch instructions and rely on compilers to automatically insert prefetches. However, most prefetching schemes in commercial compilers suffer from two limitations: (1) the source code must be available before prefetching can be applied, and (2) these prefetching schemes target only loops with statically-known strided accesses. In this study, we broaden the scope of softwarecontrolled prefetching by addressing the above two limitations. We use pro ling to discover strided accesses that frequently occur during program execution but are not determinable by the compiler. We then use the strides discovered to insert prefetches into the executable directly, without the need for re-compilation. Performance evaluation was done on an Alpha 21264-based system with a 64KB data cache and an 8MB secondary cache. We nd that even with such large caches, our technique o ers speedups ranging from 3% to 56 % in 11 out of the 26 SPEC2000 benchmarks. Our technique has been incorporated into Pixie and Spike, two products in Compaq's Tru64 Unix. Chi-Keung Luk, Robert Muth, Harish Patil, Richard Weiss 0001, P. Geoffrey Lowney, Robert S. Cohn |
ICS | 5 |
| 2002 | Tarantula: A Vector Extension to the Alpha ArchitectureabstractTarantula is an aggressive floating point machine targeted at technical, scientific and bioinformatics workloads, originally planned as a follow-on candidate to the EV8 processor. Tarantula adds to the EV8 core a vector unit capable of 32 double-precision flops per cycle. The vector unit fetches data directly from a 16 MByte second level cache with a peak bandwidth of sixty four 64-bit values per cycle. The whole chip is backed by a memory controller capable of delivering over 64 GBytes/s of raw bandwidth. Tarantula extends the Alpha ISA with new vector instructions that operate on new architectural state. Salient features of the architecture and implementation are: (1) it fully integrates into a virtual-memory cache-coherent system without changes to its coherency protocol, (2) provides high bandwidth for non-unit stride memory accesses, (3) supports gather/scatter instructions efficiently, (4) fully integrates with the EV8 core with a narrow, streamlined interface, rather than acting as a co-processor (5) can achieve a peak of 104 operations per cycle, and (6) achieves excellent "real-computation" per transistor and per watt ratios. Our detailed simulations show that Tarantula achieves an average speedup of 5X over EV8, out of a peak speedup in terms of flops of 8X. Furthermore, performance on gather/scatter intensive benchmarks such as Radix Sort is also remarkable: a speedup of almost 3X over EV8 and 15 sustained operations per cycle. Several benchmarks exceed 20 operations per cycle. Roger Espasa, Federico Ardanaz, Julio Gago, Roger Gramunt, Isaac Hernandez, Toni Juan, Joel S. Emer, Stephen Felix, P. Geoffrey Lowney, Matthew Mattina, André Seznec |
ISCA | 9 |
| 2001 | Code layout optimizations for transaction processing workloadsabstractCommercial applications such as databases and Web servers constitute the most important market segment for high-performance servers. Among these applications, on-line transaction processing (OLTP) workloads provide a challenging set of requirements for system designs since they often exhibit inefficient executions dominated by a large memory stall component. This behavior arises from large instruction and data footprints and high communication miss rates. A number of recent studies have characterized the behavior of commercial workloads and proposed architectural features to improve their performance. However, there has been little research on the impact of software and compiler-level optimizations for improving the behavior of such workloads. Alex Ramírez, Luiz André Barroso, Kourosh Gharachorloo, Robert S. Cohn, Josep Lluís Larriba-Pey, P. Geoffrey Lowney, Mateo Valero |
ISCA | 6 |
| 1996 | Hot Cold Optimization of Large Windows/NT ApplicationsabstractA dynamic instruction trace often contains many unnecessary instructions that are required only by the unexecuted portion of the program. Hot-cold optimization (HCO) is a technique that realizes this performance opportunity. HCO uses profile information to partition each routine into frequently executed (hot) and infrequently executed (cold) parts. Unnecessary operations in the hot portion are removed and compensation code is added on transitions from hot to cold as needed. We evaluate HCO on a collection of large Windows/NT applications. HCO is most effective on the programs that are call intensive and have flat profiles, providing a 3-8% reduction in path length beyond conventional optimization. Robert S. Cohn, P. Geoffrey Lowney |
MICRO | 2 |
| 1994 | Avoidance and Suppression of Compensation Code in a Trace Scheduling CompilerabstractTrace scheduling is an optimization technique that selects a sequence of basic blocks as a trace and schedules the operations from the trace together. If an operation is moved across basic block boundaries, one or more compensation copies may be required in the off-trace code. This article discusses the generation of compensation code in a trace scheduling compiler and presents techniques for limiting the amount of compensation code: avoidance (restricting code motion so that no compensation code is required) and suppression (analyzing the global flow of the program to detect when a copy is redundant). We evaluate the effectiveness of these techniques based on measurements for the SPEC89 suite and the Livermore Fortran Kernels, using our implementation of trace scheduling for a Multiflow Trace 7/300. The article compares different compiler models contrasting the performance of trace scheduling with the performance obtained from typical RISC compilation techniques. There are two key results of this study. First, the amount of compensation code generated is not large. For the SPEC89 suite, the average code size increase due to trace scheduling is 6%. Avoidance is more important than suppression, although there are some kernels that benefit significantly from compensation code suppression. Since compensation code is not a major issue, a compiler can be more aggressive in code motion and loop unrolling. Second, compensation code is not critical to obtain the benefits of trace scheduling. Our implementation of trace scheduling improves the SPEC mark rating by 30% over basic block scheduling, but restricting trace scheduling so that no compensation code is required improves the rating by 25%. This indicates that most basic block scheduling techniques can be extended to trace scheduling without requiring any complicated compensation code bookkeeping. Stefan M. Freudenberger, Thomas R. Gross, P. Geoffrey Lowney |
ACM Trans. Program. Lang. Syst. | 3 |
| 1993 | The multiflow trace scheduling compiler
P. Geoffrey Lowney, Stefan M. Freudenberger, Thomas J. Karzes, Woody Lichtenstein, Robert P. Nix, John S. O'Donnell, John C. Ruttenberg |
J. Supercomput. | 1 |
| 1981 | Carrier Arrays: An Idiom-Preserving Extension to APLabstractThe idiomatic APL programming style is limited by the constraints of a rectangular, homogeneous array as a data structure. Non-scalar data is difficult to represent and manipulate, and the non-scalar APL functions have no uniform extension to higher rank arrays. The carrier array is an extension to APL which addresses these limitations while preserving the economical APL style. A carrier array is a ragged array with an associated partition which allows functions to be applied to subarrays in parallel. The primitive functions are given base definitions on scalars and vectors, and they are extended to higher rank arrays by uniform application mechanisms. Carrier arrays also allow the last dimensions of an array to be treated as a single datum; the primitive functions are given extended definitions on scalars and vectors of this non-scalar data.This paper defines the carrier array and gives the accompanying changes to the definitions of the APL primitive functions. Examples of programming with carrier arrays are presented, and implementation issues are discussed. P. Geoffrey Lowney |
POPL | 1 |