Jamison D. Collins

dblp:82/4748 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 7 first-authorSoftware engineering, systems software and programming languages · 6 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
14 papers
Processor architecture and microarchitecture · 29% Memory systems · 16% Reconfigurable computing and FPGAs · 15%
Software engineering, system software, and programming languages
4 papers
Programming languages and type systems · 74% Operating systems · 14% Compilers and program optimization · 11%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing › heterogeneous programming models
heterogeneous multicore programming
0.222008
Merge: a programming model for heterogeneous multi-core systems · ASPLOS 2008
EXOCHI: architecture and programming environment for a heterogeneous multi-core multithreaded system · PLDI 2007
Processor architecture and microarchitecture
multithreading
0.242006
Multiple Instruction Stream Processor · ISCA 2006
Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004
Speculative precomputation: long-range prefetching of delinquent loads · ISCA 2001
Memory systems
cache
0.142004
Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004
Runtime identification of cache conflict misses: The adaptive miss buffer · ACM Trans. Comput. Syst. 2001
Speculative precomputation: long-range prefetching of delinquent loads · ISCA 2001
Memory systems › cache
prefetching
0.142004
Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004
Pointer cache assisted prefetching · MICRO 2002
Dynamic speculative precomputation · MICRO 2001
Reconfigurable computing and FPGAs
FPGA-based emulation
0.112010
Intel nehalem processor core made FPGA synthesizable · FPGA 2010
Reconfigurable computing and FPGAs
FPGA prototyping
0.112009
Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009
Reconfigurable computing and FPGAs
processor emulation
0.112009
Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009
Parallel and multicore computing › data-parallel programming
mapreduce
0.112008
Merge: a programming model for heterogeneous multi-core systems · ASPLOS 2008
Performance modeling and evaluation › parallel system performance
multiprocessor performance modeling
0.112008
CPR: Composable performance regression for scalable multiprocessor models · MICRO 2008
Performance modeling and evaluation › simulation › parallel architecture simulation
multiprocessor simulation
0.112008
CPR: Composable performance regression for scalable multiprocessor models · MICRO 2008
Parallel and multicore computing
parallel programming models
0.112008
Merge: a programming model for heterogeneous multi-core systems · ASPLOS 2008
Performance modeling and evaluation
performance prediction
0.112008
CPR: Composable performance regression for scalable multiprocessor models · MICRO 2008
Parallel and multicore computing
programming models
0.112008
Merge: a programming model for heterogeneous multi-core systems · ASPLOS 2008
Performance modeling and evaluation
simulation
0.112008
CPR: Composable performance regression for scalable multiprocessor models · MICRO 2008
Processor architecture and microarchitecture
multiple instruction stream
0.112006
Multiple Instruction Stream Processor · ISCA 2006
Memory systems › cache › cache miss
cache miss classification
0.122001
Runtime identification of cache conflict misses: The adaptive miss buffer · ACM Trans. Comput. Syst. 2001
Hardware Identification of Cache Conflict Misses · MICRO 1999
Processor architecture and microarchitecture
branch prediction
0.012004
Control Flow Optimization Via Dynamic Reconvergence Prediction · MICRO 2004
Processor architecture and microarchitecture › multithreading
helper threads
0.012004
Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004
Processor architecture and microarchitecture
speculative execution
0.012004
Control Flow Optimization Via Dynamic Reconvergence Prediction · MICRO 2004
Processor architecture and microarchitecture
memory latency tolerance
0.012002
Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation · HPCA 2002
Processor architecture and microarchitecture
out-of-order execution
0.012002
Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation · HPCA 2002
Processor architecture and microarchitecture
value prediction
0.012002
Pointer cache assisted prefetching · MICRO 2002
Memory systems › cache
cache optimization
0.022001
Hardware Identification of Cache Conflict Misses · MICRO 1999
Runtime identification of cache conflict misses: The adaptive miss buffer · ACM Trans. Comput. Syst. 2001
Reconfigurable computing and FPGAs › FPGA-based emulation
multi-FPGA emulation
0.012010
Intel nehalem processor core made FPGA synthesizable · FPGA 2010
Processor architecture and microarchitecture › multithreading
simultaneous multithreading
0.032002
Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation · HPCA 2002
Dynamic speculative precomputation · MICRO 2001
Speculative precomputation: long-range prefetching of delinquent loads · ISCA 2001
Electronic design automation › hardware verification and test
hardware verification
0.012009
Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009
Electronic design automation › hardware verification and test › functional verification
pre-silicon verification
0.012009
Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009
Processor architecture and microarchitecture › multithreading
speculative multithreading
0.022004
Control Flow Optimization Via Dynamic Reconvergence Prediction · MICRO 2004
Pointer cache assisted prefetching · MICRO 2002
Programming languages and type systems › method dispatch
dynamic dispatch
0.012008
Merge: a programming model for heterogeneous multi-core systems · ASPLOS 2008
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous multicore architecture
0.012007
EXOCHI: architecture and programming environment for a heterogeneous multi-core multithreaded system · PLDI 2007

Methods — techniques the papers use, named apart from their topics

predicate dispatch · 0.2library-based parallel language · 0.2dynamic function selection · 0.2fat binary compilation · 0.1simulation · 0.1latch mapping · 0.1clock gating conversion · 0.1FPGA synthesis · 0.1regression modeling · 0.1contention modeling · 0.1OpenMP pragma extension · 0.1sequencer architecture · 0.1cache-coherent shared memory · 0.1asynchronous control transfer · 0.1user-level multithreading · 0.0prefetching · 0.0
YearPublicationVenuePosition
2010 Intel nehalem processor core made FPGA synthesizable
abstract
We present a FPGA-synthesizable version of the Intel Nehalem processor core, synthesized, partitioned and mapped to a multi-FPGA emulation system consisting of Xilinx Virtex-4 and Virtex-5 FPGAs. To our knowledge, this is the first time a modern state-of-the-art x86 design with the out-of-order micro-architecture is made FPGA synthesizable and capable of high-speed cycle-accurate emulation. Unlike the Intel Atom core which was made FPGA synthesizable on a single Xilinx Virtex-5 in a previous endeavor, the Nehalem core is a more complex design with aggressive clock-gating, double phase latch RAMs, and RTL constructs that have no true equivalent in FPGA architectures. Despite these challenges, we are successful in making the RTL synthesizable with only 5% RTL code modifications, partitioning the design across five FPGAs, and emulating the core at 520 KHz. The synthesizable Nehalem core is able to boot Linux and execute standard x86 workloads with all architectural features enabled.
Graham Schelle, Jamison D. Collins, Ethan Schuchman, Perry H. Wang, Gautham N. Chinya, Ralf Plate, Thorsten Mattner, Franz Olbrich, Per Hammarlund, Ronak Singhal, Jim Brayton, Sebastian Steibl, Hong Wang 0003
FPGA2
2009 Intel® atomTM processor core made FPGA-synthesizable
abstract
We present an FPGA-synthesizable version of the Intel Atom processor core, synthesized to a Virtex-5 based FPGA emulation system. To make the production Atom design in SystemVerilog synthesizable through industry standard EDA tool flow, we transformed and mapped latches in the design, converted clock gating, and replaced nonsynthesizable constructs with FPGA-synthesizable counterparts. Additionally, as the target FPGA emulator is hosted on a PC platform with the Pentium-based CPU socket that supports a significantly different front side bus (FSB) protocol from that of the Atom processor, we replaced the existing bus control logic in the Atom core with an alternate FSB protocol to communicate with the rest of the PC platform. With these efforts, we succeeded in synthesizing the entire Atom processor core to fit within a single Virtex-5 LX330 FPGA. The synthesizable Atom core runs at 50Mhz on the Pentium PC motherboard with fully functional I/O peripherals. It is capable of booting off-the-shelf MS-DOS, Windows XP and Linux operating systems, and executing standard x86 workloads.
Perry H. Wang, Jamison D. Collins, Christopher T. Weaver, Belliappa Kuttanna, Shahram Salamian, Gautham N. Chinya, Ethan Schuchman, Oliver Schilling, Thorsten Doil, Sebastian Steibl, Hong Wang 0003
FPGA2
2008 Pangaea: a tightly-coupled IA32 heterogeneous chip multiprocessor
abstract
Moore's Law and the drive towards performance efficiency have led to the on-chip integration of general-purpose cores with special-purpose accelerators. Pangaea is a heterogeneous CMP design for non-rendering workloads that integrates IA32 CPU cores with non-IA32 GPU-class multi-cores, extending the current state-of-the-art CPU-GPU integration that physically "fuses" existing CPU and GPU designs. Pangaea introduces (1) a resource repartitioning of the GPU, where the hardware budget dedicated for 3D-specific graphics processing is used to build more general-purpose GPU cores, and (2) a 3-instruction extension to the IA32 ISA that supports tighter architectural integration and fine-grain shared memory collaborative multithreading between the IA32 CPU cores and the non-IA32 GPU cores. We implement Pangaea and the current CPU-GPU designs in fully-functional synthesizable RTL based on the production quality RTL of an IA32 CPU and an Intel GMA X4500 GPU. On a 65 nm ASIC process technology, the legacy graphics-specific fixed-function hardware has the area of 9 GPU cores and total power consumption of 5 GPU cores. With the ISA extensions, the latency from the time an IA32 core spawns a GPU thread to the time the thread begins execution is reduced from thousands of cycles to fewer than 30 cycles. Pangaea is synthesized on a FPGA-based prototype and runs off-the-shelf IA32 OSes. A set of general-purpose non-graphics workloads demonstrate speedups of up to 8.8x.
Henry Wong, Anne Bracy, Ethan Schuchman, Tor M. Aamodt, Jamison D. Collins, Perry H. Wang, Gautham N. Chinya, Ankur Khandelwal Groen, Hong Wang 0003
PACT5
2008 Merge: a programming model for heterogeneous multi-core systems
abstract
In this paper we propose the Merge framework, a general purpose programming model for heterogeneous multi-core systems. The Merge framework replaces current ad hoc approaches to parallel programming on heterogeneous platforms with a rigorous, library-based methodology that can automatically distribute computation across heterogeneous cores to achieve increased energy and performance efficiency. The Merge framework provides (1) a predicate dispatch-based library system for managing and invoking function variants for multiple architectures; (2) a high-level, library-oriented parallel language based on map-reduce; and (3) a compiler and runtime which implement the map-reduce language pattern by dynamically selecting the best available function implementations for a given input and machine configuration. Using a generic sequencer architecture interface for heterogeneous accelerators, the Merge framework can integrate function variants for specialized accelerators, offering the potential for to-the-metal performance for a wide range of heterogeneous architectures, all transparent to the user. The Merge framework has been prototyped on a heterogeneous platform consisting of an Intel Core 2 Duo CPU and an 8-core 32-thread Intel Graphics and Media Accelerator X3000, and a homogeneous 32-way Unisys SMP system with Intel Xeon processors. We implemented a set of benchmarks using the Merge framework and enhanced the library with X3000 specific implementations, achieving speedups of 3.6x -- 8.5x using the X3000 and 5.2x -- 22x using the 32-way system relative to the straight C reference implementation on a single IA32 core.
Michael D. Linderman, Jamison D. Collins, Hong Wang 0003, Teresa H. Meng
ASPLOS2
2008 Processor Performance Modeling using Symbolic Simulation
abstract
We propose a method of analytically characterizing processor performance as a function of circuit latencies. In our approach, we modify traditional simulation to use variables instead of fixed latencies for the internal functional units. The simulation engine then algebraically computes execution times, and the result is a mathematical equation which characterizes the performance space across numerous processor configurations. We discuss the computational complexity issues of this approach and show that instruction chunking and simple equation redundancy checking can make this approach feasible-we can model a large multi-dimensional design space with thousands to millions of design parameter combinations for about 10times the simulation time of a single conventional simulation run. We demonstrate our approach by exploring two different machines: a traditional MlPS-style in-order pipeline and the Intel Graphics Media Accelerator X3000.
Omid Azizi, Jamison D. Collins, Dinesh Patil, Hong Wang 0003, Mark Horowitz
ISPASS2
2008 CPR: Composable performance regression for scalable multiprocessor models
abstract
Uniprocessor simulators track resource utilization cycle by cycle to estimate performance. Multiprocessor simulators, however, must account for synchronization events that increase the cost of every cycle simulated and shared resource contention that increases the total number of cycles simulated. These effects cause multiprocessor simulation times to scale superlinearly with the number of cores. Composable performance regression (CPR) fundamentally addresses these intractable multiprocessor simulation times, estimating multiprocessor performance with a combination of uniprocessor, contention, and penalty models. The uniprocessor model predicts baseline performance of each core while the contention models predict interfering accesses from other cores. Uniprocessor and contention model outputs are composed by a penalty model to produce the final multiprocessor performance estimate. Trained with a production quality simulator, CPR is accurate with median errors of 6.63, 4.83 percent for dual-, quad-core multiprocessors. Furthermore, composable regression is scalable, requiring 0.33x the simulations required by prior regression strategies.
Benjamin C. Lee, Jamison D. Collins, Hong Wang 0003, David Brooks 0001
MICRO2
2007 Sequencer virtualization
abstract
The Multiple Instruction Stream Processor (MISP) architecture introduces the sequencer as a new class of architectural resource, and provides a minimalist user-level MIMD instruction set extension for application programs to directly control execution of concurrent instruction streams on these sequencers. As with classic architectural resources, namely, registers and memory, the sequencer architectural resource can be subject to virtualization. This paper details the idea of Sequencer Virtualization (SV), a foundational architectural support to decouple architectural virtual sequencers from physical sequencers. SV enables more efficient utilization of sequencer resources at the microarchitectural level while maintaining a consistent programming interface at the architectural level. To evaluate the key tradeoffs for SV, we conduct extensive experiments by implementing a prototype SV system using a custom firmware on a large-scale multiprocessor system. Using the prototype SV system, we demonstrate that SV improves efficiency in sequencer utilization while incurring little performance overhead. In particular, for a set of real multithreaded workloads, SV can significantly improve sequencer utilization, achieving an average of 32% better wall-clock performance than MISP without SV support in a multi-programming environment.
Perry H. Wang, Jamison D. Collins, Gautham N. Chinya, Bernard Lint, Asit Mallick, Koichi Yamada, Hong Wang 0003
ICS2
2007 EXOCHI: architecture and programming environment for a heterogeneous multi-core multithreaded system
abstract
Future mainstream microprocessors will likely integrate specialized accelerators, such as GPUs, onto a single die to achieve better performance and power efficiency. However, it remains a keen challenge to program such a heterogeneous multicore platform, since these specialized accelerators feature ISAs and functionality that are significantly different from the general purpose CPU cores. In this paper, we present EXOCHI: (1) Exoskeleton Sequencer(EXO), an architecture to represent heterogeneous acceleratorsas ISA-based MIMD architecture resources, and a shared virtual memory heterogeneous multithreaded program execution model that tightly couples specialized accelerator cores with generalpurpose CPU cores, and (2) C for Heterogeneous Integration(CHI), an integrated C/C++ programming environment that supports accelerator-specific inline assembly and domain-specific languages. The CHI compiler extends the OpenMP pragma for heterogeneous multithreading programming, and produces a single fat binary with code sections corresponding to different instruction sets. The runtime can judiciously spread parallel computation across the heterogeneous cores to optimize performance and power.
Perry H. Wang, Jamison D. Collins, Gautham N. Chinya, Xinmin Tian, Milind Girkar, Nick Y. Yang, Guei-Yuan Lueh, Hong Wang 0003
PLDI2
2006 Multiple Instruction Stream Processor
abstract
Microprocessor design is undergoing a major paradigm shift towards multi-core designs, in anticipation that future performance gains will come from exploiting threadlevel parallelism in the software. To support this trend, we present a novel processor architecture called the Multiple Instruction Stream Processing (MISP) architecture. MISP introduces the sequencer as a new category of architectural resource, and defines a canonical set of instructions to support user-level inter-sequencer signaling and asynchronous control transfer. MISP allows an application program to directly manage user-level threads without OS intervention. By supporting the classic cache-coherent shared-memory programming model, MISP does not require a radical shift in the multithreaded programming paradigm. This paper describes the design and evaluation of the MISP architecture for the IA-32 family of microprocessors. Using a research prototype MISP processor built on an IA-32-based multiprocessor system equipped with special firmware, we demonstrate the feasibility of implementing the MISP architecture. We then examine the utility of MISP by (1) assessing the key architectural tradeoffs of the MISP architecture design and (2) showing how legacy multithreaded applications can be migrated to MISP with relative ease.
Richard A. Hankins, Gautham N. Chinya, Jamison D. Collins, Perry H. Wang, Ryan N. Rakvic, Hong Wang 0003, John Paul Shen
ISCA3
2004 Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform
abstract
Helper threading is a technology to accelerate a program by exploiting a processor's multithreading capability to run ``assist'' threads. Previous experiments on hyper-threaded processors have demonstrated significant speedups by using helper threads to prefetch hard-to-predict delinquent data accesses. In order to apply this technique to processors that do not have built-in hardware support for multithreading, we introduce virtual multithreading (VMT), a novel form of switch-on-event user-level multithreading, capable of fly-weight multiplexing of event-driven thread executions on a single processor without additional operating system support. The compiler plays a key role in minimizing synchronization cost by judiciously partitioning register usage among the user-level threads. The VMT approach makes it possible to launch dynamic helper thread instances in response to long-latency cache miss events, and to run helper threads in the shadow of cache misses when the main thread would be otherwise stalled.The concept of VMT is prototyped on an Itanium ® 2 processor using features provided by the Processor Abstraction Layer (PAL) firmware mechanism already present in currently shipping processors. On a 4-way MP physical system equipped with VMT-enabled Itanium 2 processors, helper threading via the VMT mechanism can achieve significant performance gains for a diverse set of real-world workloads, ranging from single-threaded workstation benchmarks to heavily multithreaded large scale decision support systems (DSS) using the IBM DB2 Universal Database. We measure a wall-clock speedup of 5.8% to 38.5% for the workstation benchmarks, and 5.0% to 12.7% on various queries in the DSS workload.
Perry H. Wang, Jamison D. Collins, Hong Wang 0003, Dongkeun Kim, Bill Greene, Kai-Ming Chan, Aamir B. Yunus, Terry Sych, Stephen F. Moore, John Paul Shen
ASPLOS2
2004 Clustered Multithreaded Architectures - Pursuing both IPC and Cycle Time
abstract
Summary form only given. Clustering is an architectural technique that allows the design of wide superscalar processors without sacrificing cycle time, but at the cost of longer communication latencies. Simultaneous multithreading architectures effectively tolerate instruction latency, but put even more pressure on timing-critical processor resources. We show that the synergistic combination of the two techniques minimizes the IPC impact of the clustered architecture, and even permits more aggressive clustering of the processor than is possible with a single-threaded processor. Additionally, we show that multithreading enables effective instruction steering policies unavailable to a single-threaded clustered architecture. We explore the impact of aggressively clustering four complex processor structures, (1) instruction window wakeup and functional unit bypass logic, (2) register renaming logic, (3) the fetch unit, and (4) the integer register file, on a simultaneous multithreading processor.
Jamison D. Collins, Dean M. Tullsen
IPDPS1
2004 Control Flow Optimization Via Dynamic Reconvergence Prediction
abstract
This paper presents a novel microarchitecture technique for accurately predicting control flow reconvergence dynamically. A reconvergence point is the earliest dynamic instruction in the program where we can expect program paths to reconverge regardless of the outcome or target of the current branch. Thus, even if the immediate control flow after a branch is uncertain, execution following the reconvergence point is certain. This paper proposes a novel hardware re-convergence predictor which is both implementable and accurate, with a 4KB predictor achieving more than 95% accuracy for SPEC INT, and larger implementations achieving greater than 99% accuracy. The information provided from reconvergence prediction can increase the effectiveness of a range of previously proposed performance optimizations, including speculative multithreading, control independence, and squash reuse. This paper also demonstrates a new technique that takes advantage of the dynamic reconvergence prediction information in order to predict a wrong path excursion ahead of branch resolution. On average, 34% of wrong path fetches are eliminated.
Jamison D. Collins, Dean M. Tullsen, Hong Wang 0003
MICRO1
2002 Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation
abstract
The performance of in-order execution Itanium/sup TM/ processors can suffer significantly due to cache misses. Two memory latency tolerance approaches can be applied for the Itanium processors. One uses an out-of-order (OOO) execution core; the other assumes multithreading support and exploits cache prefetching via speculative precomputation (SP). This paper evaluates and contrasts these two approaches. In addition, this paper assesses the effectiveness of combining the two approaches. For a select set of memory-intensive programs, an in-order SMT Itanium processor using speculative precomputation can achieve performance improvement (92%) comparable to that of an out-of-order design (87%). Applying both 000 and SP yields a total performance improvement of 141% over the baseline in-order machine. OOO tends to be effective in prefetching-for L1 misses; whereas SP is primarily good at covering L2 and L3 misses. Our analysis indicates that the two approaches can be redundant or complementary depending on the type of delinquent loads that each targets. Both approaches are effective on delinquent loads in the loop body; however only SP is effective on delinquent loads found in loop control code.
Perry H. Wang, Hong Wang 0003, Jamison D. Collins, Ed Grochowski, Ralph-Michael Kling, John Paul Shen
HPCA3
2002 Pointer cache assisted prefetching
abstract
Data prefetching effectively reduces the negative effects of long load latencies on the performance of modern processors. Hardware prefetchers employ hardware structures to predict future memory addresses based on previous patterns. Thread-based prefetchers use portions of the actual program code to determine future load addresses for prefetching. This paper proposes the use of a pointer cache, which tracks pointer transitions, to aid prefetching. The pointer cache provides, for a given pointer's effective address, the base address of the object pointed to by the pointer. We examine using the pointer cache in a wide issue superscalar processor as a value predictor and to aid prefetching when a chain of pointers is being traversed. When a load misses in the L1 cache, but hits in the pointer cache, the first two cache blocks of the pointed to object are prefetched. In addition, the load's dependencies are broken by using the pointer cache hit as a value prediction. We also examine using the pointer cache to allow speculative precomputation to run farther ahead of the main thread of execution than in prior studies. Previously proposed thread-based prefetchers are limited in how far they can run ahead of the main thread when traversing a chain of recurrent dependent loads. When combined with the pointer cache, a speculative thread can make better progress ahead of the main thread, rapidly traversing data structures in the face of cache misses caused by pointer transitions.
Jamison D. Collins, Suleyman Sair, Brad Calder, Dean M. Tullsen
MICRO1
2001 Speculative precomputation: long-range prefetching of delinquent loads
abstract
This paper explores Speculative Precomputation, a technique that uses idle thread context in a multithreaded architecture to improve performance of single-threaded applications. It attacks program stalls from data cache misses by pre-computing future memory accesses in available thread contexts, and prefetching these data. This technique is evaluated by simulating the performance of a research processor based on the Itanium™ ISA supporting Simultaneous Multithreading. Two primary forms of Speculative Precomputation are evaluated. If only the non-speculative thread spawns speculative threads, performance gains of up to 30% are achieved when assuming ideal hardware. However, this speedup drops considerably with more realistic hardware assumptions. Permitting speculative threads to directly spawn additional speculative threads reduces the overhead associated with spawning threads and enables significantly more aggressive speculation, overcoming this limitation. Even with realistic costs for spawning threads, speedups as high as 169% are achieved, with an average speedup of 76%.
Jamison D. Collins, Hong Wang 0003, Dean M. Tullsen, Christopher J. Hughes, Yong-Fong Lee, Daniel M. Lavery, John Paul Shen
ISCA1
2001 Dynamic speculative precomputation
abstract
A large number of memory accesses in memory-bound applications are irregular, such as pointer dereferences, and can be effectively targeted by thread-based prefetching techniques like Speculative Precomputation. These techniques execute instructions, for example on an available SMT thread context, that have been extracted directly from the program they are trying to accelerate. Proposed techniques typically require manual user intervention to extract and optimize instruction sequences. This paper proposes Dynamic Speculative Precomputation, which performs all necessary instruction analysis, extraction, and optimization through the use of back-end instruction analysis hardware, located off the processor's critical path. For a set of memory limited benchmarks an average speedup of 14% is achieved when constructing simple p-slices, and this gain grows to 33% when making use of aggressive optimizations.
Jamison D. Collins, Dean M. Tullsen, Hong Wang 0003, John Paul Shen
MICRO1
2001 Runtime identification of cache conflict misses: The adaptive miss buffer
abstract
This paper describes the miss classification table, a simple mechanism that enables the processor or memory controller to identify each cache miss as either a conflict miss or a capacity (non-conflict) miss. The miss classification table works by storing part of the tag of the most recently evicted line of a cache set. If the next miss to that cache set has a matching tag, it is identified as a conflict miss. This technique correctly identifies 88% of misses.Several applications of this information are demonstrated, including improvements to victim caching, next-line prefetching, cache exclusion, and a pseudo-associative cache. This paper also presents the adaptive miss buffer (AMB), which combines several of these techniques, targeting each miss with the most appropriate optimization, all within a single small miss buffer. The AMB's combination of techniques achieves 16% better performance than any single technique alone.
Jamison D. Collins, Dean M. Tullsen
ACM Trans. Comput. Syst.1
1999 Hardware Identification of Cache Conflict Misses
abstract
This paper describes the Miss Classification Table, a simple mechanism that enables the processor or memory controller to identify each cache miss as either a conflict miss or a capacity (non-conflict) miss. The miss classification table works by storing part of the tag of the most recently evicted line of a cache set. If the next miss to that cache set has a matching tag, it is identified as a conflict miss. This technique correctly identifies 87% of misses in the worst case. Several applications of this information are demonstrated, including improvements to victim caching, next-line prefetching, cache exclusion, and a pseudo-associative cache. This paper also presents the Adaptive Miss Buffer (AMB), which combines several of these techniques, targeting each miss with the most appropriate optimization, all within a single small miss buffer. The AMB's combination of techniques achieves 16% better performance than any single technique alone.
Jamison D. Collins, Dean M. Tullsen
MICRO1