EDBT 2026 Demo / reviewers in the wild / expert
Jamison D. Collins
dblp:82/4748
· DBLP profile ↗
18ranked-venue papers
7as first author
0since 2021 · last 2010
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 7 first-authorSoftware engineering, systems software and programming languages · 6 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
14 papers |
Processor architecture and microarchitecture · 29% Memory systems · 16% Reconfigurable computing and FPGAs · 15% | |
| Software engineering, system software, and programming languages
4 papers |
Programming languages and type systems · 74% Operating systems · 14% Compilers and program optimization · 11% |
Topics — the 30 heaviest of 37, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing › heterogeneous programming models
heterogeneous multicore programming |
0.2 | 2 | 2008 | Merge: a programming model for heterogeneous multi-core systems · ASPLOS 2008 EXOCHI: architecture and programming environment for a heterogeneous multi-core multithreaded system · PLDI 2007 |
Processor architecture and microarchitecture
multithreading |
0.2 | 4 | 2006 | Multiple Instruction Stream Processor · ISCA 2006 Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004 Speculative precomputation: long-range prefetching of delinquent loads · ISCA 2001 |
Memory systems
cache |
0.1 | 4 | 2004 | Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004 Runtime identification of cache conflict misses: The adaptive miss buffer · ACM Trans. Comput. Syst. 2001 Speculative precomputation: long-range prefetching of delinquent loads · ISCA 2001 |
Memory systems › cache
prefetching |
0.1 | 4 | 2004 | Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004 Pointer cache assisted prefetching · MICRO 2002 Dynamic speculative precomputation · MICRO 2001 |
Reconfigurable computing and FPGAs
FPGA-based emulation |
0.1 | 1 | 2010 | Intel nehalem processor core made FPGA synthesizable · FPGA 2010 |
Reconfigurable computing and FPGAs
FPGA prototyping |
0.1 | 1 | 2009 | Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009 |
Reconfigurable computing and FPGAs
processor emulation |
0.1 | 1 | 2009 | Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009 |
Parallel and multicore computing › data-parallel programming
mapreduce |
0.1 | 1 | 2008 | Merge: a programming model for heterogeneous multi-core systems · ASPLOS 2008 |
Performance modeling and evaluation › parallel system performance
multiprocessor performance modeling |
0.1 | 1 | 2008 | CPR: Composable performance regression for scalable multiprocessor models · MICRO 2008 |
Performance modeling and evaluation › simulation › parallel architecture simulation
multiprocessor simulation |
0.1 | 1 | 2008 | CPR: Composable performance regression for scalable multiprocessor models · MICRO 2008 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2008 | Merge: a programming model for heterogeneous multi-core systems · ASPLOS 2008 |
Performance modeling and evaluation
performance prediction |
0.1 | 1 | 2008 | CPR: Composable performance regression for scalable multiprocessor models · MICRO 2008 |
Parallel and multicore computing
programming models |
0.1 | 1 | 2008 | Merge: a programming model for heterogeneous multi-core systems · ASPLOS 2008 |
Performance modeling and evaluation
simulation |
0.1 | 1 | 2008 | CPR: Composable performance regression for scalable multiprocessor models · MICRO 2008 |
Processor architecture and microarchitecture
multiple instruction stream |
0.1 | 1 | 2006 | Multiple Instruction Stream Processor · ISCA 2006 |
Memory systems › cache › cache miss
cache miss classification |
0.1 | 2 | 2001 | Runtime identification of cache conflict misses: The adaptive miss buffer · ACM Trans. Comput. Syst. 2001 Hardware Identification of Cache Conflict Misses · MICRO 1999 |
Processor architecture and microarchitecture
branch prediction |
0.0 | 1 | 2004 | Control Flow Optimization Via Dynamic Reconvergence Prediction · MICRO 2004 |
Processor architecture and microarchitecture › multithreading
helper threads |
0.0 | 1 | 2004 | Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004 |
Processor architecture and microarchitecture
speculative execution |
0.0 | 1 | 2004 | Control Flow Optimization Via Dynamic Reconvergence Prediction · MICRO 2004 |
Processor architecture and microarchitecture
memory latency tolerance |
0.0 | 1 | 2002 | Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation · HPCA 2002 |
Processor architecture and microarchitecture
out-of-order execution |
0.0 | 1 | 2002 | Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation · HPCA 2002 |
Processor architecture and microarchitecture
value prediction |
0.0 | 1 | 2002 | Pointer cache assisted prefetching · MICRO 2002 |
Memory systems › cache
cache optimization |
0.0 | 2 | 2001 | Hardware Identification of Cache Conflict Misses · MICRO 1999 Runtime identification of cache conflict misses: The adaptive miss buffer · ACM Trans. Comput. Syst. 2001 |
Reconfigurable computing and FPGAs › FPGA-based emulation
multi-FPGA emulation |
0.0 | 1 | 2010 | Intel nehalem processor core made FPGA synthesizable · FPGA 2010 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.0 | 3 | 2002 | Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation · HPCA 2002 Dynamic speculative precomputation · MICRO 2001 Speculative precomputation: long-range prefetching of delinquent loads · ISCA 2001 |
Electronic design automation › hardware verification and test
hardware verification |
0.0 | 1 | 2009 | Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009 |
Electronic design automation › hardware verification and test › functional verification
pre-silicon verification |
0.0 | 1 | 2009 | Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009 |
Processor architecture and microarchitecture › multithreading
speculative multithreading |
0.0 | 2 | 2004 | Control Flow Optimization Via Dynamic Reconvergence Prediction · MICRO 2004 Pointer cache assisted prefetching · MICRO 2002 |
Programming languages and type systems › method dispatch
dynamic dispatch |
0.0 | 1 | 2008 | Merge: a programming model for heterogeneous multi-core systems · ASPLOS 2008 |
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous multicore architecture |
0.0 | 1 | 2007 | EXOCHI: architecture and programming environment for a heterogeneous multi-core multithreaded system · PLDI 2007 |
Methods — techniques the papers use, named apart from their topics
predicate dispatch · 0.2library-based parallel language · 0.2dynamic function selection · 0.2fat binary compilation · 0.1simulation · 0.1latch mapping · 0.1clock gating conversion · 0.1FPGA synthesis · 0.1regression modeling · 0.1contention modeling · 0.1OpenMP pragma extension · 0.1sequencer architecture · 0.1cache-coherent shared memory · 0.1asynchronous control transfer · 0.1user-level multithreading · 0.0prefetching · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2010 | Intel nehalem processor core made FPGA synthesizableabstractWe present a FPGA-synthesizable version of the Intel Nehalem processor core, synthesized, partitioned and mapped to a multi-FPGA emulation system consisting of Xilinx Virtex-4 and Virtex-5 FPGAs. To our knowledge, this is the first time a modern state-of-the-art x86 design with the out-of-order micro-architecture is made FPGA synthesizable and capable of high-speed cycle-accurate emulation. Unlike the Intel Atom core which was made FPGA synthesizable on a single Xilinx Virtex-5 in a previous endeavor, the Nehalem core is a more complex design with aggressive clock-gating, double phase latch RAMs, and RTL constructs that have no true equivalent in FPGA architectures. Despite these challenges, we are successful in making the RTL synthesizable with only 5% RTL code modifications, partitioning the design across five FPGAs, and emulating the core at 520 KHz. The synthesizable Nehalem core is able to boot Linux and execute standard x86 workloads with all architectural features enabled. Graham Schelle, Jamison D. Collins, Ethan Schuchman, Perry H. Wang, Gautham N. Chinya, Ralf Plate, Thorsten Mattner, Franz Olbrich, Per Hammarlund, Ronak Singhal, Jim Brayton, Sebastian Steibl, Hong Wang 0003 |
FPGA | 2 |
| 2009 | Intel® atomTM processor core made FPGA-synthesizableabstractWe present an FPGA-synthesizable version of the Intel Atom processor core, synthesized to a Virtex-5 based FPGA emulation system. To make the production Atom design in SystemVerilog synthesizable through industry standard EDA tool flow, we transformed and mapped latches in the design, converted clock gating, and replaced nonsynthesizable constructs with FPGA-synthesizable counterparts. Additionally, as the target FPGA emulator is hosted on a PC platform with the Pentium-based CPU socket that supports a significantly different front side bus (FSB) protocol from that of the Atom processor, we replaced the existing bus control logic in the Atom core with an alternate FSB protocol to communicate with the rest of the PC platform. With these efforts, we succeeded in synthesizing the entire Atom processor core to fit within a single Virtex-5 LX330 FPGA. The synthesizable Atom core runs at 50Mhz on the Pentium PC motherboard with fully functional I/O peripherals. It is capable of booting off-the-shelf MS-DOS, Windows XP and Linux operating systems, and executing standard x86 workloads. Perry H. Wang, Jamison D. Collins, Christopher T. Weaver, Belliappa Kuttanna, Shahram Salamian, Gautham N. Chinya, Ethan Schuchman, Oliver Schilling, Thorsten Doil, Sebastian Steibl, Hong Wang 0003 |
FPGA | 2 |
| 2008 | Pangaea: a tightly-coupled IA32 heterogeneous chip multiprocessorabstractMoore's Law and the drive towards performance efficiency have led to the on-chip integration of general-purpose cores with special-purpose accelerators. Pangaea is a heterogeneous CMP design for non-rendering workloads that integrates IA32 CPU cores with non-IA32 GPU-class multi-cores, extending the current state-of-the-art CPU-GPU integration that physically "fuses" existing CPU and GPU designs. Pangaea introduces (1) a resource repartitioning of the GPU, where the hardware budget dedicated for 3D-specific graphics processing is used to build more general-purpose GPU cores, and (2) a 3-instruction extension to the IA32 ISA that supports tighter architectural integration and fine-grain shared memory collaborative multithreading between the IA32 CPU cores and the non-IA32 GPU cores. We implement Pangaea and the current CPU-GPU designs in fully-functional synthesizable RTL based on the production quality RTL of an IA32 CPU and an Intel GMA X4500 GPU. On a 65 nm ASIC process technology, the legacy graphics-specific fixed-function hardware has the area of 9 GPU cores and total power consumption of 5 GPU cores. With the ISA extensions, the latency from the time an IA32 core spawns a GPU thread to the time the thread begins execution is reduced from thousands of cycles to fewer than 30 cycles. Pangaea is synthesized on a FPGA-based prototype and runs off-the-shelf IA32 OSes. A set of general-purpose non-graphics workloads demonstrate speedups of up to 8.8x. Henry Wong, Anne Bracy, Ethan Schuchman, Tor M. Aamodt, Jamison D. Collins, Perry H. Wang, Gautham N. Chinya, Ankur Khandelwal Groen, Hong Wang 0003 |
PACT | 5 |
| 2008 | Merge: a programming model for heterogeneous multi-core systemsabstractIn this paper we propose the Merge framework, a general purpose programming model for heterogeneous multi-core systems. The Merge framework replaces current ad hoc approaches to parallel programming on heterogeneous platforms with a rigorous, library-based methodology that can automatically distribute computation across heterogeneous cores to achieve increased energy and performance efficiency. The Merge framework provides (1) a predicate dispatch-based library system for managing and invoking function variants for multiple architectures; (2) a high-level, library-oriented parallel language based on map-reduce; and (3) a compiler and runtime which implement the map-reduce language pattern by dynamically selecting the best available function implementations for a given input and machine configuration. Using a generic sequencer architecture interface for heterogeneous accelerators, the Merge framework can integrate function variants for specialized accelerators, offering the potential for to-the-metal performance for a wide range of heterogeneous architectures, all transparent to the user. The Merge framework has been prototyped on a heterogeneous platform consisting of an Intel Core 2 Duo CPU and an 8-core 32-thread Intel Graphics and Media Accelerator X3000, and a homogeneous 32-way Unisys SMP system with Intel Xeon processors. We implemented a set of benchmarks using the Merge framework and enhanced the library with X3000 specific implementations, achieving speedups of 3.6x -- 8.5x using the X3000 and 5.2x -- 22x using the 32-way system relative to the straight C reference implementation on a single IA32 core. Michael D. Linderman, Jamison D. Collins, Hong Wang 0003, Teresa H. Meng |
ASPLOS | 2 |
| 2008 | Processor Performance Modeling using Symbolic SimulationabstractWe propose a method of analytically characterizing processor performance as a function of circuit latencies. In our approach, we modify traditional simulation to use variables instead of fixed latencies for the internal functional units. The simulation engine then algebraically computes execution times, and the result is a mathematical equation which characterizes the performance space across numerous processor configurations. We discuss the computational complexity issues of this approach and show that instruction chunking and simple equation redundancy checking can make this approach feasible-we can model a large multi-dimensional design space with thousands to millions of design parameter combinations for about 10times the simulation time of a single conventional simulation run. We demonstrate our approach by exploring two different machines: a traditional MlPS-style in-order pipeline and the Intel Graphics Media Accelerator X3000. Omid Azizi, Jamison D. Collins, Dinesh Patil, Hong Wang 0003, Mark Horowitz |
ISPASS | 2 |
| 2008 | CPR: Composable performance regression for scalable multiprocessor modelsabstractUniprocessor simulators track resource utilization cycle by cycle to estimate performance. Multiprocessor simulators, however, must account for synchronization events that increase the cost of every cycle simulated and shared resource contention that increases the total number of cycles simulated. These effects cause multiprocessor simulation times to scale superlinearly with the number of cores. Composable performance regression (CPR) fundamentally addresses these intractable multiprocessor simulation times, estimating multiprocessor performance with a combination of uniprocessor, contention, and penalty models. The uniprocessor model predicts baseline performance of each core while the contention models predict interfering accesses from other cores. Uniprocessor and contention model outputs are composed by a penalty model to produce the final multiprocessor performance estimate. Trained with a production quality simulator, CPR is accurate with median errors of 6.63, 4.83 percent for dual-, quad-core multiprocessors. Furthermore, composable regression is scalable, requiring 0.33x the simulations required by prior regression strategies. Benjamin C. Lee, Jamison D. Collins, Hong Wang 0003, David Brooks 0001 |
MICRO | 2 |
| 2007 | Sequencer virtualizationabstractThe Multiple Instruction Stream Processor (MISP) architecture introduces the sequencer as a new class of architectural resource, and provides a minimalist user-level MIMD instruction set extension for application programs to directly control execution of concurrent instruction streams on these sequencers. As with classic architectural resources, namely, registers and memory, the sequencer architectural resource can be subject to virtualization. This paper details the idea of Sequencer Virtualization (SV), a foundational architectural support to decouple architectural virtual sequencers from physical sequencers. SV enables more efficient utilization of sequencer resources at the microarchitectural level while maintaining a consistent programming interface at the architectural level. To evaluate the key tradeoffs for SV, we conduct extensive experiments by implementing a prototype SV system using a custom firmware on a large-scale multiprocessor system. Using the prototype SV system, we demonstrate that SV improves efficiency in sequencer utilization while incurring little performance overhead. In particular, for a set of real multithreaded workloads, SV can significantly improve sequencer utilization, achieving an average of 32% better wall-clock performance than MISP without SV support in a multi-programming environment. Perry H. Wang, Jamison D. Collins, Gautham N. Chinya, Bernard Lint, Asit Mallick, Koichi Yamada, Hong Wang 0003 |
ICS | 2 |
| 2007 | EXOCHI: architecture and programming environment for a heterogeneous multi-core multithreaded systemabstractFuture mainstream microprocessors will likely integrate specialized accelerators, such as GPUs, onto a single die to achieve better performance and power efficiency. However, it remains a keen challenge to program such a heterogeneous multicore platform, since these specialized accelerators feature ISAs and functionality that are significantly different from the general purpose CPU cores. In this paper, we present EXOCHI: (1) Exoskeleton Sequencer(EXO), an architecture to represent heterogeneous acceleratorsas ISA-based MIMD architecture resources, and a shared virtual memory heterogeneous multithreaded program execution model that tightly couples specialized accelerator cores with generalpurpose CPU cores, and (2) C for Heterogeneous Integration(CHI), an integrated C/C++ programming environment that supports accelerator-specific inline assembly and domain-specific languages. The CHI compiler extends the OpenMP pragma for heterogeneous multithreading programming, and produces a single fat binary with code sections corresponding to different instruction sets. The runtime can judiciously spread parallel computation across the heterogeneous cores to optimize performance and power. Perry H. Wang, Jamison D. Collins, Gautham N. Chinya, Xinmin Tian, Milind Girkar, Nick Y. Yang, Guei-Yuan Lueh, Hong Wang 0003 |
PLDI | 2 |
| 2006 | Multiple Instruction Stream ProcessorabstractMicroprocessor design is undergoing a major paradigm shift towards multi-core designs, in anticipation that future performance gains will come from exploiting threadlevel parallelism in the software. To support this trend, we present a novel processor architecture called the Multiple Instruction Stream Processing (MISP) architecture. MISP introduces the sequencer as a new category of architectural resource, and defines a canonical set of instructions to support user-level inter-sequencer signaling and asynchronous control transfer. MISP allows an application program to directly manage user-level threads without OS intervention. By supporting the classic cache-coherent shared-memory programming model, MISP does not require a radical shift in the multithreaded programming paradigm. This paper describes the design and evaluation of the MISP architecture for the IA-32 family of microprocessors. Using a research prototype MISP processor built on an IA-32-based multiprocessor system equipped with special firmware, we demonstrate the feasibility of implementing the MISP architecture. We then examine the utility of MISP by (1) assessing the key architectural tradeoffs of the MISP architecture design and (2) showing how legacy multithreaded applications can be migrated to MISP with relative ease. Richard A. Hankins, Gautham N. Chinya, Jamison D. Collins, Perry H. Wang, Ryan N. Rakvic, Hong Wang 0003, John Paul Shen |
ISCA | 3 |
| 2004 | Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platformabstractHelper threading is a technology to accelerate a program by exploiting a processor's multithreading capability to run ``assist'' threads. Previous experiments on hyper-threaded processors have demonstrated significant speedups by using helper threads to prefetch hard-to-predict delinquent data accesses. In order to apply this technique to processors that do not have built-in hardware support for multithreading, we introduce virtual multithreading (VMT), a novel form of switch-on-event user-level multithreading, capable of fly-weight multiplexing of event-driven thread executions on a single processor without additional operating system support. The compiler plays a key role in minimizing synchronization cost by judiciously partitioning register usage among the user-level threads. The VMT approach makes it possible to launch dynamic helper thread instances in response to long-latency cache miss events, and to run helper threads in the shadow of cache misses when the main thread would be otherwise stalled.The concept of VMT is prototyped on an Itanium ® 2 processor using features provided by the Processor Abstraction Layer (PAL) firmware mechanism already present in currently shipping processors. On a 4-way MP physical system equipped with VMT-enabled Itanium 2 processors, helper threading via the VMT mechanism can achieve significant performance gains for a diverse set of real-world workloads, ranging from single-threaded workstation benchmarks to heavily multithreaded large scale decision support systems (DSS) using the IBM DB2 Universal Database. We measure a wall-clock speedup of 5.8% to 38.5% for the workstation benchmarks, and 5.0% to 12.7% on various queries in the DSS workload. Perry H. Wang, Jamison D. Collins, Hong Wang 0003, Dongkeun Kim, Bill Greene, Kai-Ming Chan, Aamir B. Yunus, Terry Sych, Stephen F. Moore, John Paul Shen |
ASPLOS | 2 |
| 2004 | Clustered Multithreaded Architectures - Pursuing both IPC and Cycle TimeabstractSummary form only given. Clustering is an architectural technique that allows the design of wide superscalar processors without sacrificing cycle time, but at the cost of longer communication latencies. Simultaneous multithreading architectures effectively tolerate instruction latency, but put even more pressure on timing-critical processor resources. We show that the synergistic combination of the two techniques minimizes the IPC impact of the clustered architecture, and even permits more aggressive clustering of the processor than is possible with a single-threaded processor. Additionally, we show that multithreading enables effective instruction steering policies unavailable to a single-threaded clustered architecture. We explore the impact of aggressively clustering four complex processor structures, (1) instruction window wakeup and functional unit bypass logic, (2) register renaming logic, (3) the fetch unit, and (4) the integer register file, on a simultaneous multithreading processor. Jamison D. Collins, Dean M. Tullsen |
IPDPS | 1 |
| 2004 | Control Flow Optimization Via Dynamic Reconvergence PredictionabstractThis paper presents a novel microarchitecture technique for accurately predicting control flow reconvergence dynamically. A reconvergence point is the earliest dynamic instruction in the program where we can expect program paths to reconverge regardless of the outcome or target of the current branch. Thus, even if the immediate control flow after a branch is uncertain, execution following the reconvergence point is certain. This paper proposes a novel hardware re-convergence predictor which is both implementable and accurate, with a 4KB predictor achieving more than 95% accuracy for SPEC INT, and larger implementations achieving greater than 99% accuracy. The information provided from reconvergence prediction can increase the effectiveness of a range of previously proposed performance optimizations, including speculative multithreading, control independence, and squash reuse. This paper also demonstrates a new technique that takes advantage of the dynamic reconvergence prediction information in order to predict a wrong path excursion ahead of branch resolution. On average, 34% of wrong path fetches are eliminated. Jamison D. Collins, Dean M. Tullsen, Hong Wang 0003 |
MICRO | 1 |
| 2002 | Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative PrecomputationabstractThe performance of in-order execution Itanium/sup TM/ processors can suffer significantly due to cache misses. Two memory latency tolerance approaches can be applied for the Itanium processors. One uses an out-of-order (OOO) execution core; the other assumes multithreading support and exploits cache prefetching via speculative precomputation (SP). This paper evaluates and contrasts these two approaches. In addition, this paper assesses the effectiveness of combining the two approaches. For a select set of memory-intensive programs, an in-order SMT Itanium processor using speculative precomputation can achieve performance improvement (92%) comparable to that of an out-of-order design (87%). Applying both 000 and SP yields a total performance improvement of 141% over the baseline in-order machine. OOO tends to be effective in prefetching-for L1 misses; whereas SP is primarily good at covering L2 and L3 misses. Our analysis indicates that the two approaches can be redundant or complementary depending on the type of delinquent loads that each targets. Both approaches are effective on delinquent loads in the loop body; however only SP is effective on delinquent loads found in loop control code. Perry H. Wang, Hong Wang 0003, Jamison D. Collins, Ed Grochowski, Ralph-Michael Kling, John Paul Shen |
HPCA | 3 |
| 2002 | Pointer cache assisted prefetchingabstractData prefetching effectively reduces the negative effects of long load latencies on the performance of modern processors. Hardware prefetchers employ hardware structures to predict future memory addresses based on previous patterns. Thread-based prefetchers use portions of the actual program code to determine future load addresses for prefetching. This paper proposes the use of a pointer cache, which tracks pointer transitions, to aid prefetching. The pointer cache provides, for a given pointer's effective address, the base address of the object pointed to by the pointer. We examine using the pointer cache in a wide issue superscalar processor as a value predictor and to aid prefetching when a chain of pointers is being traversed. When a load misses in the L1 cache, but hits in the pointer cache, the first two cache blocks of the pointed to object are prefetched. In addition, the load's dependencies are broken by using the pointer cache hit as a value prediction. We also examine using the pointer cache to allow speculative precomputation to run farther ahead of the main thread of execution than in prior studies. Previously proposed thread-based prefetchers are limited in how far they can run ahead of the main thread when traversing a chain of recurrent dependent loads. When combined with the pointer cache, a speculative thread can make better progress ahead of the main thread, rapidly traversing data structures in the face of cache misses caused by pointer transitions. Jamison D. Collins, Suleyman Sair, Brad Calder, Dean M. Tullsen |
MICRO | 1 |
| 2001 | Speculative precomputation: long-range prefetching of delinquent loadsabstractThis paper explores Speculative Precomputation, a technique that uses idle thread context in a multithreaded architecture to improve performance of single-threaded applications. It attacks program stalls from data cache misses by pre-computing future memory accesses in available thread contexts, and prefetching these data. This technique is evaluated by simulating the performance of a research processor based on the Itanium™ ISA supporting Simultaneous Multithreading. Two primary forms of Speculative Precomputation are evaluated. If only the non-speculative thread spawns speculative threads, performance gains of up to 30% are achieved when assuming ideal hardware. However, this speedup drops considerably with more realistic hardware assumptions. Permitting speculative threads to directly spawn additional speculative threads reduces the overhead associated with spawning threads and enables significantly more aggressive speculation, overcoming this limitation. Even with realistic costs for spawning threads, speedups as high as 169% are achieved, with an average speedup of 76%. Jamison D. Collins, Hong Wang 0003, Dean M. Tullsen, Christopher J. Hughes, Yong-Fong Lee, Daniel M. Lavery, John Paul Shen |
ISCA | 1 |
| 2001 | Dynamic speculative precomputationabstractA large number of memory accesses in memory-bound applications are irregular, such as pointer dereferences, and can be effectively targeted by thread-based prefetching techniques like Speculative Precomputation. These techniques execute instructions, for example on an available SMT thread context, that have been extracted directly from the program they are trying to accelerate. Proposed techniques typically require manual user intervention to extract and optimize instruction sequences. This paper proposes Dynamic Speculative Precomputation, which performs all necessary instruction analysis, extraction, and optimization through the use of back-end instruction analysis hardware, located off the processor's critical path. For a set of memory limited benchmarks an average speedup of 14% is achieved when constructing simple p-slices, and this gain grows to 33% when making use of aggressive optimizations. Jamison D. Collins, Dean M. Tullsen, Hong Wang 0003, John Paul Shen |
MICRO | 1 |
| 2001 | Runtime identification of cache conflict misses: The adaptive miss bufferabstractThis paper describes the miss classification table, a simple mechanism that enables the processor or memory controller to identify each cache miss as either a conflict miss or a capacity (non-conflict) miss. The miss classification table works by storing part of the tag of the most recently evicted line of a cache set. If the next miss to that cache set has a matching tag, it is identified as a conflict miss. This technique correctly identifies 88% of misses.Several applications of this information are demonstrated, including improvements to victim caching, next-line prefetching, cache exclusion, and a pseudo-associative cache. This paper also presents the adaptive miss buffer (AMB), which combines several of these techniques, targeting each miss with the most appropriate optimization, all within a single small miss buffer. The AMB's combination of techniques achieves 16% better performance than any single technique alone. Jamison D. Collins, Dean M. Tullsen |
ACM Trans. Comput. Syst. | 1 |
| 1999 | Hardware Identification of Cache Conflict MissesabstractThis paper describes the Miss Classification Table, a simple mechanism that enables the processor or memory controller to identify each cache miss as either a conflict miss or a capacity (non-conflict) miss. The miss classification table works by storing part of the tag of the most recently evicted line of a cache set. If the next miss to that cache set has a matching tag, it is identified as a conflict miss. This technique correctly identifies 87% of misses in the worst case. Several applications of this information are demonstrated, including improvements to victim caching, next-line prefetching, cache exclusion, and a pseudo-associative cache. This paper also presents the Adaptive Miss Buffer (AMB), which combines several of these techniques, targeting each miss with the most appropriate optimization, all within a single small miss buffer. The AMB's combination of techniques achieves 16% better performance than any single technique alone. Jamison D. Collins, Dean M. Tullsen |
MICRO | 1 |