Steven Wallace

dblp:77/5724 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
1since 2021 · last 2025
0009-0000-3777-1587ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 7 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-authorComputer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
1 paper
Routing and switching · 77% Network measurement and analytics · 23%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Processor architecture and microarchitecture · 87% Memory systems · 8% Performance modeling and evaluation · 5%
Software engineering, system software, and programming languages
1 paper
Program analysis · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
instruction fetch
0.131999
Instruction Recycling on a Multiple-Path Processor · HPCA 1999
Modeled and Measured Instruction Fetching Performance for Superscalar Microprocessors · IEEE Trans. Parallel Distributed Syst. 1998
Multiple Branch and Block Prediction · HPCA 1997
Program analysis › dynamic analysis › instrumentation
binary instrumentation
0.112005
Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005
Program analysis › dynamic analysis
dynamic instrumentation
0.112005
Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005
Processor architecture and microarchitecture › speculative execution
multi-path execution
0.021999
Instruction Recycling on a Multiple-Path Processor · HPCA 1999
Threaded Multiple Path Execution · ISCA 1998
Processor architecture and microarchitecture › multithreading
simultaneous multithreading
0.021999
Threaded Multiple Path Execution · ISCA 1998
Instruction Recycling on a Multiple-Path Processor · HPCA 1999
Processor architecture and microarchitecture
branch prediction
0.021998
Multiple Branch and Block Prediction · HPCA 1997
Threaded Multiple Path Execution · ISCA 1998
Processor architecture and microarchitecture › dynamic optimization
instruction reuse
0.011999
Instruction Recycling on a Multiple-Path Processor · HPCA 1999
Processor architecture and microarchitecture
speculative execution
0.011999
Instruction Recycling on a Multiple-Path Processor · HPCA 1999
Processor architecture and microarchitecture › branch prediction
branch target buffer
0.011998
Modeled and Measured Instruction Fetching Performance for Superscalar Microprocessors · IEEE Trans. Parallel Distributed Syst. 1998
Memory systems › cache
prefetching
0.011998
Modeled and Measured Instruction Fetching Performance for Superscalar Microprocessors · IEEE Trans. Parallel Distributed Syst. 1998
Processor architecture and microarchitecture › multithreading
speculative multithreading
0.011998
Threaded Multiple Path Execution · ISCA 1998
Processor architecture and microarchitecture › instruction fetch
fetch bandwidth
0.011997
Multiple Branch and Block Prediction · HPCA 1997
Processor architecture and microarchitecture › branch prediction
multiple branch prediction
0.011997
Multiple Branch and Block Prediction · HPCA 1997
Performance modeling and evaluation
profiling
0.012005
Pin: building customized program analysis tools with dynamic instrumentation · PLDI 2005
Processor architecture and microarchitecture › branch prediction
branch misprediction
0.011998
Threaded Multiple Path Execution · ISCA 1998
Memory systems
cache
0.011998
Modeled and Measured Instruction Fetching Performance for Superscalar Microprocessors · IEEE Trans. Parallel Distributed Syst. 1998
Processor architecture and microarchitecture
superscalar processor
0.011997
Multiple Branch and Block Prediction · HPCA 1997
Processor architecture and microarchitecture › superscalar processor
wide-issue
0.011997
Multiple Branch and Block Prediction · HPCA 1997

Methods — techniques the papers use, named apart from their topics

active probing · 0.9BGP analysis · 0.9liveness analysis · 0.1instruction scheduling · 0.1inlining · 0.1dynamic compilation · 0.1register reallocation · 0.1register re-allocation · 0.1simulation · 0.0trace injection · 0.0performance simulation · 0.0mathematical modeling · 0.0two-level adaptive prediction · 0.0
YearPublicationVenuePosition
2025 R&E Routing Policy: Inference and Implication
abstract
BGP hides information that is crucial for building accurate routing models. In this paper, we combine BGP and active probing to infer relative route preference policies of research and education R&E connected ASes. We inferred that systems in ≈88% of <12K prefixes that 2,578 ASes announced in the R&E ecosystem were insensitive to AS path length when selecting provider routes -- only ≈8-9% appeared to assign the same local preference to available R&E and commodity routes. We validate our method, and discuss broader application of the method to infer relative route preference, a crucial step in being able to accurately model routing policies.
Matthew J. Luckie, Steven Wallace, Karl Newell, Jeff Bartig, Sadi Koçak, Niels den Otter, Kaj Koole, James Deaton, K. C. Claffy
IMC2
2007 SuperPin: Parallelizing Dynamic Instrumentation for Real-Time Performance
abstract
Dynamic instrumentation systems have proven to be extremely valuable for program introspection, architectural simulation, and bug detection. Yet a major drawback of modern instrumentation systems is that the instrumented applications often execute several orders of magnitude slower than native application performance. In this paper, we present a novel approach to dynamic instrumentation where several non-overlapping slices of an application are launched as separate instrumentation threads and executed in parallel in order to approach real-time performance. A direct implementation of our technique in the Pin dynamic instrumentation system results in dramatic speedups for various instrumentation tasks - often resulting in order-of-magnitude performance improvements. Our implementation is available as part of the Pin distribution, which has been downloaded over 10,000 times since its release
Steven Wallace, Kim M. Hazelwood
CGO1
2005 Pin: building customized program analysis tools with dynamic instrumentation
abstract
Robust and powerful software instrumentation tools are essential for program analysis tasks such as profiling, performance evaluation, and bug detection. To meet this need, we have developed a new instrumentation system called Pin. Our goals are to provide easy-to-use, portable, transparent, and efficient instrumentation. Instrumentation tools (called Pintools) are written in C/C++ using Pin's rich API. Pin follows the model of ATOM, allowing the tool writer to analyze an application at the instruction level without the need for detailed knowledge of the underlying instruction set. The API is designed to be architecture independent whenever possible, making Pintools source compatible across different architectures. However, a Pintool can access architecture-specific details when necessary. Instrumentation with Pin is mostly transparent as the application and Pintool observe the application's original, uninstrumented behavior. Pin uses dynamic compilation to instrument executables while they are running. For efficiency, Pin uses several techniques, including inlining, register re-allocation, liveness analysis, and instruction scheduling to optimize instrumentation. This fully automated approach delivers significantly better instrumentation performance than similar tools. For example, Pin is 3.3x faster than Valgrind and 2x faster than DynamoRIO for basic-block counting. To illustrate Pin's versatility, we describe two Pintools in daily use to analyze production software. Pin is publicly available for Linux platforms on four architectures: IA32 (32-bit x86), EM64T (64-bit x86), Itanium®, and ARM. In the ten months since Pin 2 was released in July 2004, there have been over 3000 downloads from its website.
Chi-Keung Luk, Robert S. Cohn, Robert Muth, Harish Patil, Artur Klauser, P. Geoffrey Lowney, Steven Wallace, Vijay Janapa Reddi, Kim M. Hazelwood
PLDI7
2003 The rationale of the current optical networking initiatives
Cees T. A. M. de Laat, Erik Radius, Steven Wallace
Future Gener. Comput. Syst.3
1999 Instruction Recycling on a Multiple-Path Processor
abstract
Processors that can simultaneously execute multiple paths of execution will only exacerbate the fetch bandwidth problem already plaguing conventional processors. On a multiple-path processor which speculatively executes less likely paths of hard-to-predict branches, the work done along a speculative path is normally discarded if that path is found to be incorrect. Instead, it can be beneficial to keep these instruction traces stored in the processor for possible future use. This paper introduces instruction recycling, where previously decoded instructions from recently executed paths are injected back into the rename stage. This increases the supply of instructions to the execution pipeline and decreases fetch latency. In addition, if the operands have not changed for a recycled instruction, the instruction can bypass the issue and execution stages, benefiting from instruction reuse. Instruction recycling and reuse are examined for a simultaneous multithreading architecture with multiple path execution. It is shown to increase performance by 7% for single-program workloads and by 12% on multiple-program workloads.
Steven Wallace, Dean M. Tullsen, Brad Calder
HPCA1
1998 Threaded Multiple Path Execution
abstract
This paper presents Threaded Multi-Path Execution (TME), which exploits existing hardware on a Simultaneous Multithreading (SMT) processor to speculatively execute multiple paths of execution. When there are fewer threads in an SMT processor than hardware contexts, threaded multi-path execution uses spare contexts to fetch and execute code along the less likely path of hard-to-predict branches. This paper describes the hardware mechanisms needed to enable an SMT processor to efficiently spawn speculative threads for threaded multi-path execution. The Mapping Synchronization Bus is described which enables the spawning of these multiple paths. Policies are examined for deciding which branches to fork, and for managing competition between primary and alternate path threads for critical resources. Our results show that TME increases the single program performance of an SMT with eight thread contexts by 14%-23% on average, depending on the misprediction penalty, for programs with a high misprediction rate.
Steven Wallace, Brad Calder, Dean M. Tullsen
ISCA1
1998 Modeled and Measured Instruction Fetching Performance for Superscalar Microprocessors
abstract
Instruction fetching is critical to the performance of a superscalar microprocessor. We develop a mathematical model for three different cache techniques and evaluate its performance both in theory and in simulation using the SPEC95 suite of benchmarks. In all the techniques, the fetching performance is dramatically lower than ideal expectations. To help remedy the situation, we also evaluate its performance using prefetching. Nevertheless, fetching performance is fundamentally limited by control transfers. To solve this problem, we introduce a new fetching mechanism called a dual branch target buffer. The dual branch target buffer enables fetching performance to leap beyond the limitation imposed by conventional methods and achieve a high instruction fetching rate.
Steven Wallace, Nader Bagherzadeh
IEEE Trans. Parallel Distributed Syst.1
1997 Multiple Branch and Block Prediction
abstract
Accurate branch prediction and instruction fetch prediction of a microprocessor are critical to achieve high performance. For a processor which fetches and executes multiple instructions per cycle, an accurate and high bandwidth instruction fetching mechanism becomes increasingly important to performance. Unfortunately, the relatively small basic block size exhibited in many general-purpose applications severely limits instruction fetching. In order to achieve a high fetching rate for wide-issue superscalars, a scalable method to predict multiple branches per block of sequential instructions is presented. Its accuracy is equivalent to a scalar two-level adaptive prediction. Also, to overcome the limitation imposed by control transfers, a scalable method to predict multiple blocks is presented. As a result, a two black, multiple branch prediction mechanism for a block width of 8 instructions achieves an effective fetching rate of 8 instructions per cycle on the SPEC95 benchmark suite.
Steven Wallace, Nader Bagherzadeh
HPCA1
1995 Design and implementation of a 100 MHz centralized instruction window for a superscalar microprocessor
abstract
The maxim of the superscalar architecture is that higher performance can be achieved by executing multiple instructions simultaneously. This can be realized on hardware by using a centralized instruction window. We present the design and implementation of a centralized instruction window capable of out-of-order issue and completion of four instructions per cycle. A compact layout (6.4 mm by 2.2 mm) of a 32-entry instruction window resulted from a full-custom design in 1.0 /spl mu/m (drawn) 3-layer metal CMOS technology. The layout was verified by simulation and shown to operate at a clock frequency over 100 MHz.
Steven Wallace, Nirav Dagli, Nader Bagherzadeh
ICCD1
1994 Performance Issues of a Superscalar Microprocessor
abstract
Cache, dynamic scheduling, bypassing, branch prediction, and fetch efficiency are primary issues concerning performance of a superscalar microprocessor. This paper considers all these issues and shows their impact on performance by running our simulator on seventeen different programs. Our approach in handling branch prediction is shown to significantly decrease the bad branch penalty. Furthermore, results show that the average instruction fetch places an upper bound on speedup and is the most critical factor in determining overall performance. Its performance impact is greater than all other factors combined.
Steven Wallace, Nader Bagherzadeh
ICPP (1)1