VLDB 2026 Research / reviewers in the wild / expert
Daniel M. Lavery
dblp:73/4160
· DBLP profile ↗
13ranked-venue papers
2as first author
0since 2021 · last 2007
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-authorSoftware engineering, systems software and programming languages · 6Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Performance modeling and evaluation · 36% Processor architecture and microarchitecture · 35% Memory systems · 22% | |
| Software engineering, system software, and programming languages
7 papers |
Compilers and program optimization · 82% Program analysis · 14% Operating systems · 5% |
Topics — the 23 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2007 | Comparative characterization of SPEC CPU2000 and CPU2006 on Itanium architecture · SIGMETRICS 2007 |
Performance modeling and evaluation › system-level analysis › architecture evaluation
processor performance characterization |
0.1 | 1 | 2007 | Comparative characterization of SPEC CPU2000 and CPU2006 on Itanium architecture · SIGMETRICS 2007 |
Memory systems
cache |
0.1 | 2 | 2007 | Speculative precomputation: long-range prefetching of delinquent loads · ISCA 2001 Comparative characterization of SPEC CPU2000 and CPU2006 on Itanium architecture · SIGMETRICS 2007 |
Compilers and program optimization
instruction scheduling |
0.0 | 3 | 1996 | Modulo Scheduling of Loops in Control-intensive Non-numeric Programs · MICRO 1996 The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors · IEEE Trans. Computers 1995 Unrolling-based optimizations for modulo scheduling · MICRO 1995 |
Memory systems › memory access optimization
memory prefetching |
0.0 | 1 | 2002 | Post-Pass Binary Adaptation for Software-Based Speculative Precomputation · PLDI 2002 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 3 | 1996 | Modulo Scheduling of Loops in Control-intensive Non-numeric Programs · MICRO 1996 Unrolling-based optimizations for modulo scheduling · MICRO 1995 Compiler technology for future microprocessors · Proc. IEEE 1995 |
Compilers and program optimization
compiler analysis |
0.0 | 1 | 2001 | On the Importance of Points-to Analysis and Other Memory Disambiguation Methods for C Programs · PLDI 2001 |
Compilers and program optimization › dependence analysis
memory disambiguation |
0.0 | 1 | 2001 | On the Importance of Points-to Analysis and Other Memory Disambiguation Methods for C Programs · PLDI 2001 |
Program analysis › static analysis
pointer analysis |
0.0 | 1 | 2001 | On the Importance of Points-to Analysis and Other Memory Disambiguation Methods for C Programs · PLDI 2001 |
Processor architecture and microarchitecture
multithreading |
0.0 | 1 | 2001 | Speculative precomputation: long-range prefetching of delinquent loads · ISCA 2001 |
Compilers and program optimization › instruction scheduling › software pipelining
modulo scheduling |
0.0 | 2 | 1996 | Modulo Scheduling of Loops in Control-intensive Non-numeric Programs · MICRO 1996 Unrolling-based optimizations for modulo scheduling · MICRO 1995 |
Reconfigurable computing and FPGAs
modulo scheduling |
0.0 | 2 | 1996 | Modulo Scheduling of Loops in Control-intensive Non-numeric Programs · MICRO 1996 Unrolling-based optimizations for modulo scheduling · MICRO 1995 |
Processor architecture and microarchitecture
branch prediction |
0.0 | 1 | 2007 | Comparative characterization of SPEC CPU2000 and CPU2006 on Itanium architecture · SIGMETRICS 2007 |
Processor architecture and microarchitecture
speculation |
0.0 | 1 | 1996 | Modulo Scheduling of Loops in Control-intensive Non-numeric Programs · MICRO 1996 |
Compilers and program optimization
dependence analysis |
0.0 | 1 | 1995 | Compiler technology for future microprocessors · Proc. IEEE 1995 |
Compilers and program optimization › instruction scheduling
global instruction scheduling |
0.0 | 1 | 1995 | The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors · IEEE Trans. Computers 1995 |
Compilers and program optimization › instruction scheduling
instruction-level parallelism |
0.0 | 1 | 1995 | Compiler technology for future microprocessors · Proc. IEEE 1995 |
Compilers and program optimization › loop transformation
loop unrolling |
0.0 | 1 | 1995 | Unrolling-based optimizations for modulo scheduling · MICRO 1995 |
Compilers and program optimization
predicated compilation |
0.0 | 1 | 1995 | Compiler technology for future microprocessors · Proc. IEEE 1995 |
Processor architecture and microarchitecture
instruction scheduling |
0.0 | 1 | 1995 | The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors · IEEE Trans. Computers 1995 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 1 | 1995 | The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors · IEEE Trans. Computers 1995 |
Program analysis › dynamic analysis › instrumentation
binary instrumentation |
0.0 | 1 | 2002 | Post-Pass Binary Adaptation for Software-Based Speculative Precomputation · PLDI 2002 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.0 | 1 | 2001 | Speculative precomputation: long-range prefetching of delinquent loads · ISCA 2001 |
Methods — techniques the papers use, named apart from their topics
workload characterization · 0.1speculative prefetch threads · 0.1post-pass compilation · 0.1IPC analysis · 0.1preemption modeling · 0.0compiler optimization · 0.0superblock scheduling · 0.0modulo variable expansion · 0.0hyperblock scheduling · 0.0thread-level speculation · 0.0simulation · 0.0points-to analysis · 0.0acyclic scheduling · 0.0IMPACT compiler · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2007 | Comparative characterization of SPEC CPU2000 and CPU2006 on Itanium architectureabstractRecently SPEC1 released the next generation of its CPU benchmark, widely used by compiler writers and architects for measuring processor performance. This calls for characterization of the applications in SPEC CPU2006 to guide the design of future microprocessors. In addition, it necessitates assessing the change in the characteristics of the applications from one suite to another. Although similar studies using the retired SPEC CPU benchmark suites have been done in the past, to the best of our knowledge, a thorough characterization of CPU2006 and its comparison with CPU2000 has not been done so far. In this paper, we present the above; specifically, we analyze IPC (instructions per cycle), L1, L2 data cache misses and branch prediction, especially in CPU2006. Arun Kejariwal, Gerolf Hoflehner, Darshan Desai, Daniel M. Lavery, Alexandru Nicolau, Alexander V. Veidenbaum |
SIGMETRICS | 4 |
| 2004 | Compiler Optimizations for Transaction Processing Workloads on Itanium® Linux SystemsabstractThis paper discusses a repertoire of well-known and new compiler optimizations that help produce excellent server application performance and investigates their performance contributions. These optimizations combined produce a 40% speed-up in on-line transaction processing (OLTP) performance and have been implemented in the Intel C/C++ Itanium compiler. In particular, the paper presents compiler optimizations that take advantage of the Itanium register stack, proposes an enhanced Linux preemption model and demonstrates their performance potential for server applications. Gerolf Hoflehner, Knud Kirkegaard, Rod Skinner, Daniel M. Lavery, Yong-Fong Lee, Wei Li 0015 |
MICRO | 4 |
| 2003 | Optimizations to Prevent Cache Penalties for the Intel ® Itanium 2 ProcessorabstractThis paper describes scheduling optimizations in the Intel/spl reg/ Itanium/spl reg/ compiler to prevent cache penalties due to various micro-architectural effects on the Itanium 2 processor. This paper does not try to improve cache hit rates but to avoid penalties, which probably all processors have in one form or another, even in the case of cache hits. These optimizations make use of sophisticated methods for disambiguation of memory references, and this paper examines the performance improvement obtained by integrating these methods into the cache optimizations. Jean-Francois Collard, Daniel M. Lavery |
CGO | 2 |
| 2003 | Optimization for the Intel® Itanium ®Architectur Register StackabstractThe Intel/spl reg/ Itanium/spl reg/ architecture contains a number of innovative compiler-controllable features designed to exploit instruction level parallelism. New code generation and optimization techniques are critical to the application of these features to improve processor performance. For instance, the Itanium/spl reg/ architecture provides a compiler-controllable virtual register stack to reduce the penalty of memory accesses associated with procedure calls. The Itanium/spl reg/ Register Stack Engine (RSE) transparently manages the register stack and saves and restores physical registers to and from memory as needed. Existing code generation techniques for the register stack aggressively allocate virtual registers without regard to the register pressure on different control-flow paths. As such, applications with large data sets may stress the RSE, and cause substantial execution delays due to the high number of register saves and restores. Since the Itanium/spl reg/ architecture is developed around Explicitly Parallel Instruction Computing (EPIC) concepts, solutions to increasing the register stack efficiency favor code generation techniques rather than hardware approaches. Alex Settle, Daniel A. Connors, Gerolf Hoflehner, Daniel M. Lavery |
CGO | 4 |
| 2002 | Post-Pass Binary Adaptation for Software-Based Speculative PrecomputationabstractRecently, a number of thread-based prefetching techniques have been proposed. These techniques aim at improving the latency of single-threaded applications by leveraging multithreading resources to perform memory prefetching via speculative prefetch threads. Software-based speculative precomputation (SSP) is one such technique, proposed for multithreaded Itanium models. SSP does not require expensive hardware support-instead it relies on the compiler to adapt binaries to perform prefetching on otherwise idle hardware thread contexts at run time. This paper presents a post-pass compilation tool for generating SSP-enhanced binaries. The tool is able to: (1) analyze a single-threaded application to generate prefetch threads; (2) identify and embed trigger points in the original binary; and (3) produce a new binary that has the prefetch threads attached. The execution of the new binary spawns the speculative prefetch threads, which are executed concurrently with the main thread. Our results indicate that for a set of pointer-intensive benchmarks, the prefetching performed by the speculative threads achieves an average of 87% speedup on an in-order processor and 5% speedup on an out-of-order processor. Shih-Wei Liao, Perry H. Wang, Hong Wang 0003, John Paul Shen, Gerolf Hoflehner, Daniel M. Lavery |
PLDI | 6 |
| 2001 | Speculative precomputation: long-range prefetching of delinquent loadsabstractThis paper explores Speculative Precomputation, a technique that uses idle thread context in a multithreaded architecture to improve performance of single-threaded applications. It attacks program stalls from data cache misses by pre-computing future memory accesses in available thread contexts, and prefetching these data. This technique is evaluated by simulating the performance of a research processor based on the Itanium™ ISA supporting Simultaneous Multithreading. Two primary forms of Speculative Precomputation are evaluated. If only the non-speculative thread spawns speculative threads, performance gains of up to 30% are achieved when assuming ideal hardware. However, this speedup drops considerably with more realistic hardware assumptions. Permitting speculative threads to directly spawn additional speculative threads reduces the overhead associated with spawning threads and enables significantly more aggressive speculation, overcoming this limitation. Even with realistic costs for spawning threads, speedups as high as 169% are achieved, with an average speedup of 76%. Jamison D. Collins, Hong Wang 0003, Dean M. Tullsen, Christopher J. Hughes, Yong-Fong Lee, Daniel M. Lavery, John Paul Shen |
ISCA | 6 |
| 2001 | On the Importance of Points-to Analysis and Other Memory Disambiguation Methods for C ProgramsabstractIn this paper, we evaluate the benefits achievable from pointer analysis and other memory disambiguation techniques for C/C++ programs, using the framework of the production compiler for the Intel ® Itanium TM processor. Most of the prior work on memory disambiguation has primarily focused on pointer analysis, and either presents only static estimates of the accuracy of the analysis (such as average points-to set size), or provides performance data in the context of certain individual optiraizations. In contrast, our study is based on a complete memory disambiguation framework that uses a whole set of techniques including pointer analysis. Further, it presents how various compiler analyses and optimizations interact with the memory disambiguator, evaluates how much they benefit from disambiguation, and measures the eventual impact on the performance of the program. The paper also analyzes the types of disambiguation queries that are typically received by the disambiguator, which disambiguation techniques prove most effective in resolving them, and what type of queries prove difficult to be resolved. The study is based on empirical data collected for the SPEC CINT2000 C/C++ programs, running on the Itanium processor. 1. Rakesh Ghiya, Daniel M. Lavery, David C. Sehr |
PLDI | 2 |
| 1996 | Modulo Scheduling of Loops in Control-intensive Non-numeric ProgramsabstractMuch of the previous work on modulo scheduling has targeted numeric programs, in which, often, the majority of the loops are well-behaved loop-counter-based loops without early exits. In control-intensive non-numeric programs, the loops frequently have characteristics that make it more difficult to effectively apply modulo scheduling. These characteristics include multiple control flow paths, loops that are not based on a loop counter, and multiple exits. In these loops, the presence of unimportant paths with high resource usage or long dependence chains can penalize the important paths. A path that contains a hazard such as another nested loop can prohibit modulo scheduling of the loop. Control dependences can severely restrict the overlap of the blocks within and across iterations. This paper describes a set of methods that allow effective modulo scheduling of loops with multiple exits. The techniques include removal of control dependences to enable speculation, extensions to modulo variable expansion, and a new epilogue generation scheme. These methods can be used with superblock and hyperblock techniques to allow modulo scheduling of the selected paths of loops with arbitrary control flow. A case study is presented to show how these methods, combined with superblock techniques, enable module scheduling to be effectively applied to control-intensive non-numeric programs. Performance results for several SPEC CINT92 benchmarks and Unix utility programs are reported and demonstrate the applicability of modulo scheduling to this class of programs. Daniel M. Lavery, Wen-Mei W. Hwu |
MICRO | 1 |
| 1995 | Unrolling-based optimizations for modulo schedulingabstractModulo scheduling is a method for overlapping successive iterations of a loop in order to find sufficient instruction-level parallelism to fully utilize high-issue-rate processors. The achieved throughput module scheduled loop depends on the resource requirements, the dependence pattern, and the register requirements of the computation in the loop body. Traditionally, unrolling followed by acyclic scheduling of the unrolled body has been an alternative to module scheduling. However, there are benefits to unrolling even if the loop is to be module scheduled. Unrolling can improve the throughput by allowing a smaller non-integral effective initiation interval to be achieved. After unrolling, optimizations can be applied to the loop that reduce both the resource requirements and the height of the critical paths. Together, unrolling and unrolling-based optimizations can enable the completion of multiple iterations per cycle in some cases. This paper describes the benefits of unrolling and a set of optimizations for unrolled loops which have been implemented in the IMPACT compiler. The performance benefits of unrolling for five of the SPEC CFP92 programs are reported. Daniel M. Lavery, Wen-Mei W. Hwu |
MICRO | 1 |
| 1995 | Compiler technology for future microprocessorsabstractAdvances in hardware technology have made it possible for microprocessors to execute a large number of instructions concurrently (i.e., in parallel). These microprocessors take advantage of the opportunity to execute instructions in parallel to increase the execution speed of a program. As in other forms of parallel processing, the performance of these microprocessors can vary greatly depending on the qualify of the software. In particular the quality of compilers can make an order of magnitude difference in performance. This paper presents a new generation of compiler technology that has emerged to deliver the large amount of instruction-level-parallelism that is already required by some current state-of-the-art microprocessors and will be required by more future microprocessors. We introduce critical components of the technology which deal with difficult problems that are encountered when compiling programs for a high degree of instruction-level-parallelism. We present examples to illustrate the functional requirements of these components. To provide more insight into the challenges involved, we present in-depth case studies on predicated compilation and maintenance of dependence information, two of the components that are largely missing from most current commercial compilers. Wen-Mei W. Hwu, Richard E. Hank, David M. Gallagher, Scott A. Mahlke, Daniel M. Lavery, Grant E. Haab, John C. Gyllenhaal, David I. August |
Proc. IEEE | 5 |
| 1995 | The Importance of Prepass Code Scheduling for Superscalar and Superpipelined ProcessorsabstractSuperscalar and superpipelined processors utilize parallelism to achieve peak performance that can be several times higher than that of conventional scalar processors. In order for this potential to be translated into the speedup of real program, the compiler must be able to schedule instructions so that the parallel hardware is effectively utilized. Previous work has shown that prepass code scheduling helps to produce a better schedule for scientific programs, but the importance of prescheduling has never been demonstrated for control-intensive non-numeric programs. These programs are significantly different from the scientific programs because they contain frequent branches. The compiler must do global scheduling in order to find enough independent instructions. In this paper, the code optimizer and scheduler of the IMPACT-I C compiler is described. Within this framework, we study the importance of prepass code scheduling for a set of production C programs. It is shown that, in contrast to the results previously obtained for scientific programs, prescheduling is not important for compiling control-intensive programs to the current generation of superscalar and superpipelined processors. However, if some of the current restrictions on upward code motion can be removed in future architectures, prescheduling would substantially improve the execution time of this class of programs on both superscalar and superpipelined processors.> Pohua P. Chang, Daniel M. Lavery, Scott A. Mahlke, William Y. Chen, Wen-Mei W. Hwu |
IEEE Trans. Computers | 2 |
| 1993 | The superblock: An effective technique for VLIW and superscalar compilation
Wen-Mei W. Hwu, Scott A. Mahlke, William Y. Chen, Pohua P. Chang, Nancy J. Warter, Roger A. Bringmann, Roland G. Ouellette, Richard E. Hank, Tokuzo Kiyohara, Grant E. Haab, John G. Holm, Daniel M. Lavery |
J. Supercomput. | 12 |
| 1991 | The Organization of the Cedar System
Jeff Konicek, Tracy Tilton, Alexander V. Veidenbaum, Chuanqi Zhu, Edward S. Davidson, Ruppert A. Downing, Michael J. Haney, Pen-Chung Yew, P. Michael Farmwald, David J. Kuck, Daniel M. Lavery, Robert A. Lindsey, D. Pointer, John T. Andrews, T. Murphy, Stephen W. Turner, Nancy J. Warter |
ICPP (1) | 12 |