VLDB 2026 Research / reviewers in the wild / expert
Steven Wallace
dblp:77/5724
· DBLP profile ↗
10ranked-venue papers
7as first author
1since 2021 · last 2025
0009-0000-3777-1587ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 7 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-authorComputer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
1 paper |
Routing and switching · 77% Network measurement and analytics · 23% | |
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Processor architecture and microarchitecture · 87% Memory systems · 8% Performance modeling and evaluation · 5% | |
| Software engineering, system software, and programming languages
1 paper |
Program analysis · 100% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
instruction fetch |
0.1 | 3 | 1999 | Instruction Recycling on a Multiple-Path Processor · HPCA 1999 Modeled and Measured Instruction Fetching Performance for Superscalar Microprocessors · IEEE Trans. Parallel Distributed Syst. 1998 Multiple Branch and Block Prediction · HPCA 1997 |
Program analysis › dynamic analysis › instrumentation
binary instrumentation |
0.1 | 1 | 2005 | Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005 |
Program analysis › dynamic analysis
dynamic instrumentation |
0.1 | 1 | 2005 | Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005 |
Processor architecture and microarchitecture › speculative execution
multi-path execution |
0.0 | 2 | 1999 | Instruction Recycling on a Multiple-Path Processor · HPCA 1999 Threaded Multiple Path Execution · ISCA 1998 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.0 | 2 | 1999 | Threaded Multiple Path Execution · ISCA 1998 Instruction Recycling on a Multiple-Path Processor · HPCA 1999 |
Processor architecture and microarchitecture
branch prediction |
0.0 | 2 | 1998 | Multiple Branch and Block Prediction · HPCA 1997 Threaded Multiple Path Execution · ISCA 1998 |
Processor architecture and microarchitecture › dynamic optimization
instruction reuse |
0.0 | 1 | 1999 | Instruction Recycling on a Multiple-Path Processor · HPCA 1999 |
Processor architecture and microarchitecture
speculative execution |
0.0 | 1 | 1999 | Instruction Recycling on a Multiple-Path Processor · HPCA 1999 |
Processor architecture and microarchitecture › branch prediction
branch target buffer |
0.0 | 1 | 1998 | Modeled and Measured Instruction Fetching Performance for Superscalar Microprocessors · IEEE Trans. Parallel Distributed Syst. 1998 |
Memory systems › cache
prefetching |
0.0 | 1 | 1998 | Modeled and Measured Instruction Fetching Performance for Superscalar Microprocessors · IEEE Trans. Parallel Distributed Syst. 1998 |
Processor architecture and microarchitecture › multithreading
speculative multithreading |
0.0 | 1 | 1998 | Threaded Multiple Path Execution · ISCA 1998 |
Processor architecture and microarchitecture › instruction fetch
fetch bandwidth |
0.0 | 1 | 1997 | Multiple Branch and Block Prediction · HPCA 1997 |
Processor architecture and microarchitecture › branch prediction
multiple branch prediction |
0.0 | 1 | 1997 | Multiple Branch and Block Prediction · HPCA 1997 |
Performance modeling and evaluation
profiling |
0.0 | 1 | 2005 | Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005 |
Processor architecture and microarchitecture › branch prediction
branch misprediction |
0.0 | 1 | 1998 | Threaded Multiple Path Execution · ISCA 1998 |
Memory systems
cache |
0.0 | 1 | 1998 | Modeled and Measured Instruction Fetching Performance for Superscalar Microprocessors · IEEE Trans. Parallel Distributed Syst. 1998 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 1 | 1997 | Multiple Branch and Block Prediction · HPCA 1997 |
Processor architecture and microarchitecture › superscalar processor
wide-issue |
0.0 | 1 | 1997 | Multiple Branch and Block Prediction · HPCA 1997 |
Methods — techniques the papers use, named apart from their topics
active probing · 0.9BGP analysis · 0.9liveness analysis · 0.1instruction scheduling · 0.1inlining · 0.1dynamic compilation · 0.1register reallocation · 0.1register re-allocation · 0.1simulation · 0.0trace injection · 0.0performance simulation · 0.0mathematical modeling · 0.0two-level adaptive prediction · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | R&E Routing Policy: Inference and ImplicationabstractBGP hides information that is crucial for building accurate routing models. In this paper, we combine BGP and active probing to infer relative route preference policies of research and education R&E connected ASes. We inferred that systems in ≈88% of <12K prefixes that 2,578 ASes announced in the R&E ecosystem were insensitive to AS path length when selecting provider routes -- only ≈8-9% appeared to assign the same local preference to available R&E and commodity routes. We validate our method, and discuss broader application of the method to infer relative route preference, a crucial step in being able to accurately model routing policies. Matthew J. Luckie, Steven Wallace, Karl Newell, Jeff Bartig, Sadi Koçak, Niels den Otter, Kaj Koole, James Deaton, K. C. Claffy |
IMC | 2 |
| 2007 | SuperPin: Parallelizing Dynamic Instrumentation for Real-Time PerformanceabstractDynamic instrumentation systems have proven to be extremely valuable for program introspection, architectural simulation, and bug detection. Yet a major drawback of modern instrumentation systems is that the instrumented applications often execute several orders of magnitude slower than native application performance. In this paper, we present a novel approach to dynamic instrumentation where several non-overlapping slices of an application are launched as separate instrumentation threads and executed in parallel in order to approach real-time performance. A direct implementation of our technique in the Pin dynamic instrumentation system results in dramatic speedups for various instrumentation tasks - often resulting in order-of-magnitude performance improvements. Our implementation is available as part of the Pin distribution, which has been downloaded over 10,000 times since its release Steven Wallace, Kim M. Hazelwood |
CGO | 1 |
| 2005 | Pin: building customized program analysis tools with dynamic instrumentationabstractRobust and powerful software instrumentation tools are essential for program analysis tasks such as profiling, performance evaluation, and bug detection. To meet this need, we have developed a new instrumentation system called Pin. Our goals are to provide easy-to-use, portable, transparent, and efficient instrumentation. Instrumentation tools (called Pintools) are written in C/C++ using Pin's rich API. Pin follows the model of ATOM, allowing the tool writer to analyze an application at the instruction level without the need for detailed knowledge of the underlying instruction set. The API is designed to be architecture independent whenever possible, making Pintools source compatible across different architectures. However, a Pintool can access architecture-specific details when necessary. Instrumentation with Pin is mostly transparent as the application and Pintool observe the application's original, uninstrumented behavior. Pin uses dynamic compilation to instrument executables while they are running. For efficiency, Pin uses several techniques, including inlining, register re-allocation, liveness analysis, and instruction scheduling to optimize instrumentation. This fully automated approach delivers significantly better instrumentation performance than similar tools. For example, Pin is 3.3x faster than Valgrind and 2x faster than DynamoRIO for basic-block counting. To illustrate Pin's versatility, we describe two Pintools in daily use to analyze production software. Pin is publicly available for Linux platforms on four architectures: IA32 (32-bit x86), EM64T (64-bit x86), Itanium®, and ARM. In the ten months since Pin 2 was released in July 2004, there have been over 3000 downloads from its website. Chi-Keung Luk, Robert S. Cohn, Robert Muth, Harish Patil, Artur Klauser, P. Geoffrey Lowney, Steven Wallace, Vijay Janapa Reddi, Kim M. Hazelwood |
PLDI | 7 |
| 2003 | The rationale of the current optical networking initiatives
Cees T. A. M. de Laat, Erik Radius, Steven Wallace |
Future Gener. Comput. Syst. | 3 |
| 1999 | Instruction Recycling on a Multiple-Path ProcessorabstractProcessors that can simultaneously execute multiple paths of execution will only exacerbate the fetch bandwidth problem already plaguing conventional processors. On a multiple-path processor which speculatively executes less likely paths of hard-to-predict branches, the work done along a speculative path is normally discarded if that path is found to be incorrect. Instead, it can be beneficial to keep these instruction traces stored in the processor for possible future use. This paper introduces instruction recycling, where previously decoded instructions from recently executed paths are injected back into the rename stage. This increases the supply of instructions to the execution pipeline and decreases fetch latency. In addition, if the operands have not changed for a recycled instruction, the instruction can bypass the issue and execution stages, benefiting from instruction reuse. Instruction recycling and reuse are examined for a simultaneous multithreading architecture with multiple path execution. It is shown to increase performance by 7% for single-program workloads and by 12% on multiple-program workloads. Steven Wallace, Dean M. Tullsen, Brad Calder |
HPCA | 1 |
| 1998 | Threaded Multiple Path ExecutionabstractThis paper presents Threaded Multi-Path Execution (TME), which exploits existing hardware on a Simultaneous Multithreading (SMT) processor to speculatively execute multiple paths of execution. When there are fewer threads in an SMT processor than hardware contexts, threaded multi-path execution uses spare contexts to fetch and execute code along the less likely path of hard-to-predict branches. This paper describes the hardware mechanisms needed to enable an SMT processor to efficiently spawn speculative threads for threaded multi-path execution. The Mapping Synchronization Bus is described which enables the spawning of these multiple paths. Policies are examined for deciding which branches to fork, and for managing competition between primary and alternate path threads for critical resources. Our results show that TME increases the single program performance of an SMT with eight thread contexts by 14%-23% on average, depending on the misprediction penalty, for programs with a high misprediction rate. Steven Wallace, Brad Calder, Dean M. Tullsen |
ISCA | 1 |
| 1998 | Modeled and Measured Instruction Fetching Performance for Superscalar MicroprocessorsabstractInstruction fetching is critical to the performance of a superscalar microprocessor. We develop a mathematical model for three different cache techniques and evaluate its performance both in theory and in simulation using the SPEC95 suite of benchmarks. In all the techniques, the fetching performance is dramatically lower than ideal expectations. To help remedy the situation, we also evaluate its performance using prefetching. Nevertheless, fetching performance is fundamentally limited by control transfers. To solve this problem, we introduce a new fetching mechanism called a dual branch target buffer. The dual branch target buffer enables fetching performance to leap beyond the limitation imposed by conventional methods and achieve a high instruction fetching rate. Steven Wallace, Nader Bagherzadeh |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1997 | Multiple Branch and Block PredictionabstractAccurate branch prediction and instruction fetch prediction of a microprocessor are critical to achieve high performance. For a processor which fetches and executes multiple instructions per cycle, an accurate and high bandwidth instruction fetching mechanism becomes increasingly important to performance. Unfortunately, the relatively small basic block size exhibited in many general-purpose applications severely limits instruction fetching. In order to achieve a high fetching rate for wide-issue superscalars, a scalable method to predict multiple branches per block of sequential instructions is presented. Its accuracy is equivalent to a scalar two-level adaptive prediction. Also, to overcome the limitation imposed by control transfers, a scalable method to predict multiple blocks is presented. As a result, a two black, multiple branch prediction mechanism for a block width of 8 instructions achieves an effective fetching rate of 8 instructions per cycle on the SPEC95 benchmark suite. Steven Wallace, Nader Bagherzadeh |
HPCA | 1 |
| 1995 | Design and implementation of a 100 MHz centralized instruction window for a superscalar microprocessorabstractThe maxim of the superscalar architecture is that higher performance can be achieved by executing multiple instructions simultaneously. This can be realized on hardware by using a centralized instruction window. We present the design and implementation of a centralized instruction window capable of out-of-order issue and completion of four instructions per cycle. A compact layout (6.4 mm by 2.2 mm) of a 32-entry instruction window resulted from a full-custom design in 1.0 /spl mu/m (drawn) 3-layer metal CMOS technology. The layout was verified by simulation and shown to operate at a clock frequency over 100 MHz. Steven Wallace, Nirav Dagli, Nader Bagherzadeh |
ICCD | 1 |
| 1994 | Performance Issues of a Superscalar MicroprocessorabstractCache, dynamic scheduling, bypassing, branch prediction, and fetch efficiency are primary issues concerning performance of a superscalar microprocessor. This paper considers all these issues and shows their impact on performance by running our simulator on seventeen different programs. Our approach in handling branch prediction is shown to significantly decrease the bad branch penalty. Furthermore, results show that the average instruction fetch places an upper bound on speedup and is the most critical factor in determining overall performance. Its performance impact is greater than all other factors combined. Steven Wallace, Nader Bagherzadeh |
ICPP (1) | 1 |