VLDB 2026 Research / reviewers in the wild / expert
Gary S. Tyson
dblp:99/123
· DBLP profile ↗
52ranked-venue papers
6as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 37 · 6 first-authorSoftware engineering, systems software and programming languages · 15Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
22 papers |
Processor architecture and microarchitecture · 53% Memory systems · 34% Embedded and real-time systems · 5% | |
| Software engineering, system software, and programming languages
5 papers |
Compilers and program optimization · 100% |
Topics — the 30 heaviest of 58, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
instruction packing |
0.1 | 2 | 2005 | Reducing Instruction Fetch Cost by Packing Instructions into RegisterWindows · MICRO 2005 Improving Program Efficiency by Packing Instructions into Registers · ISCA 2005 |
Processor architecture and microarchitecture › instruction set architecture
instruction set design |
0.1 | 2 | 2005 | An Energy Efficient Instruction Set Synthesis Framework for Low Power Embedded System Designs · IEEE Trans. Computers 2005 FITS: framework-based instruction-set tuning synthesis for embedded application specific processors · DAC 2004 |
Memory systems
cache management |
0.1 | 4 | 2001 | Branch History Guided Instruction Prefetching · HPCA 2001 Eager writeback - a technique for improving bandwidth utilization · MICRO 2000 Active Management of Data Caches by Exploiting Reuse Information · IEEE Trans. Computers 1999 |
Compilers and program optimization › compiler optimization
optimization phase ordering |
0.1 | 1 | 2009 | Practical exhaustive optimization phase order exploration and evaluation · ACM Trans. Archit. Code Optim. 2009 |
Memory systems › cache › cache organization
filter cache |
0.1 | 1 | 2007 | Guaranteeing Hits to Improve the Efficiency of a Small Instruction Cache · MICRO 2007 |
Memory systems › cache › CPU cache
instruction cache |
0.1 | 1 | 2007 | Guaranteeing Hits to Improve the Efficiency of a Small Instruction Cache · MICRO 2007 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.1 | 6 | 2001 | Improving the Accuracy and Performance of Memory Communication Through Renaming · MICRO 1997 The effects of predicated execution on branch prediction · MICRO 1994 Techniques for extracting instruction level parallelism on MIMD architectures · MICRO 1993 |
Processor architecture and microarchitecture
branch prediction |
0.1 | 3 | 2000 | Improving BTB performance in the presence of DLLs · MICRO 2000 Analyzing the Working Set Characteristics of Branch Execution · MICRO 1998 The effects of predicated execution on branch prediction · MICRO 1994 |
Embedded and real-time systems › embedded processor
code compression |
0.1 | 1 | 2005 | Improving Program Efficiency by Packing Instructions into Registers · ISCA 2005 |
Processor architecture and microarchitecture
instruction fetch |
0.1 | 1 | 2005 | Reducing Instruction Fetch Cost by Packing Instructions into RegisterWindows · MICRO 2005 |
Processor architecture and microarchitecture
instruction set architecture |
0.1 | 1 | 2005 | Improving Program Efficiency by Packing Instructions into Registers · ISCA 2005 |
Processor architecture and microarchitecture › instruction set architecture › instruction set design
instruction set synthesis |
0.1 | 1 | 2005 | An Energy Efficient Instruction Set Synthesis Framework for Low Power Embedded System Designs · IEEE Trans. Computers 2005 |
Performance modeling and evaluation
workload characterization |
0.1 | 2 | 2004 | A Prefetch Taxonomy · IEEE Trans. Computers 2004 Analyzing the Working Set Characteristics of Branch Execution · MICRO 1998 |
Processor architecture and microarchitecture › special-purpose processor › application-specific processor design
instruction synthesis |
0.0 | 1 | 2004 | FITS: framework-based instruction-set tuning synthesis for embedded application specific processors · DAC 2004 |
Memory systems › cache
prefetching |
0.0 | 1 | 2004 | A Prefetch Taxonomy · IEEE Trans. Computers 2004 |
Compilers and program optimization
code size reduction |
0.0 | 2 | 2005 | Reducing Instruction Fetch Cost by Packing Instructions into RegisterWindows · MICRO 2005 Improving Program Efficiency by Packing Instructions into Registers · ISCA 2005 |
Compilers and program optimization › register allocation
register pressure reduction |
0.0 | 1 | 2001 | Evaluating the Use of Register Queues in Software Pipelined Loops · IEEE Trans. Computers 2001 |
Compilers and program optimization › instruction scheduling
software pipelining |
0.0 | 1 | 2001 | Evaluating the Use of Register Queues in Software Pipelined Loops · IEEE Trans. Computers 2001 |
Memory systems › cache management
cache partitioning |
0.0 | 1 | 2001 | Stack Value File: Custom Microarchitecture for the Stack · HPCA 2001 |
Memory systems › cache management › instruction cache management
instruction cache miss reduction |
0.0 | 1 | 2001 | Branch History Guided Instruction Prefetching · HPCA 2001 |
Processor architecture and microarchitecture › instruction fetch
instruction prefetching |
0.0 | 1 | 2001 | Branch History Guided Instruction Prefetching · HPCA 2001 |
Processor architecture and microarchitecture › front-end
instruction supply |
0.0 | 1 | 2001 | Branch History Guided Instruction Prefetching · HPCA 2001 |
Memory systems
memory hierarchy |
0.0 | 1 | 2001 | Stack Value File: Custom Microarchitecture for the Stack · HPCA 2001 |
Processor architecture and microarchitecture
register file |
0.0 | 1 | 2001 | Evaluating the Use of Register Queues in Software Pipelined Loops · IEEE Trans. Computers 2001 |
Memory systems
cache design |
0.0 | 2 | 1997 | On High-Bandwidth Data Cache Design for Multi-Issue Processors · MICRO 1997 A Study of Single-Chip Processor/Cache Organizations for Large Numbers of Transistors · ISCA 1994 |
Processor architecture and microarchitecture › branch prediction
branch target buffer |
0.0 | 1 | 2000 | Improving BTB performance in the presence of DLLs · MICRO 2000 |
Memory systems
cache |
0.0 | 1 | 2000 | Eager writeback - a technique for improving bandwidth utilization · MICRO 2000 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.0 | 1 | 2007 | Guaranteeing Hits to Improve the Efficiency of a Small Instruction Cache · MICRO 2007 |
Processor architecture and microarchitecture › memory system microarchitecture
memory renaming |
0.0 | 1 | 1997 | Improving the Accuracy and Performance of Memory Communication Through Renaming · MICRO 1997 |
Memory systems › cache › cache organization
multi-ported cache |
0.0 | 1 | 1997 | On High-Bandwidth Data Cache Design for Multi-Issue Processors · MICRO 1997 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.2loop cache · 0.1instruction register file · 0.1performance simulation · 0.1exhaustive search · 0.1metadata bits · 0.1guaranteed hit tracking · 0.1benchmark evaluation · 0.1prefetch classification · 0.0coverage and accuracy metrics · 0.0trace-driven simulation · 0.0execution-driven simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Experience of Administering Our First S-STEM Program to Broaden Participation in Computer ScienceabstractThis paper documents the findings of our analysis of the implementation of our six-year NSF S-STEM scholarship program. One major finding was that, for underrepresented students to major in computer science, knowing the major existed and understanding the nature of the program were the most important factors. Also, the academic support system and hands-on nature of the major had a significant impact on scholarship recipients' persistence in the major. Evidence demonstrated that scholarship recipients had a 10%+ higher year-to-year persistence rate from their freshmen to sophomore year than that of all computer science students of the same entering classes. For all computer science students, college computer science major GPAs were not strongly correlated with their high school GPAs, financial need, or ACT math scores. This paper also presents lessons learned and resulting recommendations for future new scholarship administrators, as our lessons can likely be applied to other grants that recruit and deal with underrepresented groups. An-I Wang, David B. Whalley, Gary S. Tyson |
SIGCSE | 4 |
| 2019 | Amniote: A User Space Interface to the Android Runtime
Zachary Yannes, Gary S. Tyson |
ENASE | 2 |
| 2016 | Interactive Augmented Reality for Dance
Taylor Brockhoeft, Jennifer Petuch, James Bach, Emil Djerekarov, Margareta Ackerman, Gary S. Tyson |
ICCC | 6 |
| 2016 | Agave: A benchmark suite for exploring the complexities of the Android software stackabstractTraditional suites used for benchmarking high-performance computing platforms or for architectural design space exploration use much simpler virtual memory layouts and multitasking/ multithreading schemes, which means that they cannot be used to study the complex interactions among the layers of the Android software stack. To demonstrate this, we present memory reference and concurrency data showing how Android applications differ from traditional C benchmarks. We propose the Agave suite of open-source applications as the basis for a standard, multipurpose Android benchmark suite. We make all sources and tools available in hopes that the community will adopt and build on this initial version of Agave. Martin K. Brown, Zachary Yannes, Michael Lustig, Mazdak Sanati, Sally A. McKee, Gary S. Tyson, Steven K. Reinhardt |
ISPASS | 6 |
| 2015 | Scheduling instruction effects for a statically pipelined processorabstractStatically pipelined processors have a fully exposed datapath where all portions of the pipeline are directly controlled by effects within an instruction, which simplifies hardware and enables a new level of compiler optimizations. This paper describes an effect scheduling strategy to aggressively compact instructions, which has a critical impact on code size and performance. Unique scheduling challenges include more frequent name dependences and fewer renaming opportunities due to static pipeline (SP) registers being dedicated for specific operations. We also realized the SP in a hardware implementation language (VHDL) to evaluate the real energy benefits. Despite the compiler challenges, we achieve performance, code size, and energy improvements compared to a conventional MIPS processor. B. Davis, Ryan Baird, Peter Gavin, Magnus Själander, Ian Finlayson, F. Rasapour, G. Cook, Gang-Ryung Uh, David B. Whalley, Gary S. Tyson |
CASES | 10 |
| 2015 | Serious 3D gaming research for the vision impairedabstractThis paper describes “A walk in the park” which is a serious game designed and developed for performing auditory-vision sensory substitution (AVSS) research for the visually impaired. Until now, evaluating AVSS solutions required the creation of a sensory substitution device (SSD) to test a strategy for converting images into sounds. This process often takes years to complete. 3D game engines provide a research and training platform that is capable of quickly implementing and testing different image-to-sound conversion strategies in a virtual world environment. Justin B. Marshall, Gary S. Tyson, Juan Llanos, Roberto Miguel Sanchez, Francisca B. Marshall |
HealthCom | 2 |
| 2014 | A journey toward obtaining our first NSF S-STEM (scholarship) grantabstractComputer science Ph.D. training provides numerous opportunities to prepare doctoral graduates to write research grant proposals. However, writing scholarship grant proposals is a very different process, and a newcomer might go through many attempts before obtaining their first awarded grant. This paper documents our four proposal submissions prior to acquiring our first NSF S-STEM grant for the Department of Computer Science at Florida State University. This paper also highlights major issues to consider when writing such proposals. We hope that future newcomers will be able to avoid some of the pitfalls we encountered in obtaining scholarship grants of a similar nature. An-I Wang, Gary S. Tyson, David B. Whalley, Robert A. van Engelen |
SIGCSE | 2 |
| 2013 | Improving processor efficiency by statically pipelining instructionsabstractA new generation of applications requires reduced power consumption without sacrificing performance. Instruction pipelining is commonly used to meet application performance requirements, but some implementation aspects of pipelining are inefficient with respect to energy usage. We propose static pipelining as a new instruction set architecture to enable more efficient instruction flow through the pipeline, which is accomplished by exposing the pipeline structure to the compiler. While this approach simplifies hardware pipeline requirements, significant modifications to the compiler are required. This paper describes the code generation and compiler optimizations we implemented to exploit the features of this architecture. We show that we can achieve performance and code size improvements despite a very low-level instruction representation. We also demonstrate that static pipelining of instructions reduces energy usage by simplifying hardware, avoiding many unnecessary operations, and allowing the compiler to perform optimizations that are not possible on traditional architectures. Ian Finlayson, Brandon Davis, Peter Gavin, Gang-Ryung Uh, David B. Whalley, Magnus Själander, Gary S. Tyson |
LCTES | 7 |
| 2009 | Guaranteeing instruction fetch behavior with a lookahead instruction fetch engine (LIFE)abstractInstruction fetch behavior has been shown to be very regular and predictable, even for diverse application areas. In this work, we propose the Lookahead Instruction Fetch Engine (LIFE), which is designed to exploit the regularity present in instruction fetch. The nucleus of LIFE is the Tagless Hit Instruction Cache (TH-IC), a small cache that assists the instruction fetch pipeline stage as it efficiently captures information about both sequential and non-sequential transitions between instructions. TH-IC provides a considerable savings in fetch energy without incurring the performance penalty normally associated with small filter instruction caches. LIFE extends TH-IC by making use of advanced control flow metadata to further improve utilization of fetch-associated structures such as the branch predictor, branch target buffer, and return address stack. These structures are selectively disabled by LIFE when it can be determined that they are unnecessary for the following instruction to be fetched. Our results show that LIFE enables further reductions in total processor energy consumption with no impact on application execution times even for the most aggressive power-saving configuration. We also explore the use of LIFE metadata on guiding decisions further down the pipeline. Next sequential line prefetch for the data cache can be enhanced by only prefetching when the triggering instruction has been previously accessed in the TH-IC. This strategy reduces the number of useless prefetches and thus contributes to improving overall processor efficiency. LIFE enables designers to boost instruction fetch efficiency by reducing energy cost without negatively affecting performance. Stephen Roderick Hines, Yuval Peress, Peter Gavin, David B. Whalley, Gary S. Tyson |
LCTES | 5 |
| 2009 | Practical exhaustive optimization phase order exploration and evaluationabstractChoosing the most appropriate optimization phase ordering has been a long-standing problem in compiler optimizations. Exhaustive evaluation of all possible orderings of optimization phases for each function is generally dismissed as infeasible for production-quality compilers targeting accepted benchmarks. In this article, we show that it is possible to exhaustively evaluate the optimization phase order space for each function in a reasonable amount of time for most of the functions in our benchmark suite. To achieve this goal, we used various techniques to significantly prune the optimization phase order search space so that it can be inexpensively enumerated in most cases and reduce the number of program simulations required to evaluate program performance for each distinct phase ordering. The techniques described are applicable to other compilers in which it is desirable to find the best phase ordering for most functions in a reasonable amount of time. We also describe some interesting properties of the optimization phase order space, which will prove useful for further studies of related problems in compilers. Prasad A. Kulkarni, David B. Whalley, Gary S. Tyson, Jack W. Davidson |
ACM Trans. Archit. Code Optim. | 3 |
| 2008 | Archer: A Community Distributed Computing Infrastructure for Computer Architecture Research and Education
Renato J. O. Figueiredo, P. Oscar Boykin, José A. B. Fortes, Tao Li 0006, Jie-Kwon Peir, David Wolinsky, Lizy Kurian John, David R. Kaeli, David J. Lilja, Sally A. McKee, Gokhan Memik, Alain J. Roy, Gary S. Tyson |
CollaborateCom | 13 |
| 2008 | Enhancing the effectiveness of utilizing an instruction register fileabstractThis paper describes the outcomes of the NSF Grant CNS-0615085: CSR-EHS: enhancing the effectiveness of utilizing an instruction register file. We improved promoting instructions to reside in the IRF and adapted compiler optimizations to better utilize an IRF. We show that an IRF can decrease the execution time penalty of using an LO/filter cache while further reducing energy consumption. Finally, we introduce a tagless hit instruction cache that significantly reduces energy consumption without increasing execution time. David B. Whalley, Gary S. Tyson |
IPDPS | 2 |
| 2007 | Facilitating compiler optimizations through the dynamic mapping of alternate register structuresabstractAggressive compiler optimizations such as software pipelining and loop invariant code motion can significantly improve application performance, but these transformations often require the use of several additional registers to hold data values across one or more loop iterations. Compilers that target embedded systems may often have difficulty exploiting these optimizations since many embedded systems typically do not have as many general purpose registers available. Alternate register structures like register queues can be used to facilitate the application of these optimizations due to common reference patterns. In this paper, we propose a microarchitectural technique that permits these alternate register structures to be efficiently mapped into a given processor architecture and automatically exploited by an optimizing compiler. We show that this minimally invasive technique can be used to facilitate the application of software pipelining and loop invariant code motion for a variety of embedded benchmarks. This leads to performance improvements for the embedded processor, as well as new opportunities for further aggressive optimization of embedded systems software due to a significant decrease in the register pressure of tight loops. Christopher Zimmer 0001, Stephen Roderick Hines, Prasad A. Kulkarni, Gary S. Tyson, David B. Whalley |
CASES | 4 |
| 2007 | Evaluating Heuristic Optimization Phase Order Search AlgorithmsabstractProgram-specific or function-specific optimization phase sequences are universally accepted to achieve better overall performance than any fixed optimization phase ordering. A number of heuristic phase order space search algorithms have been devised to find customized phase orderings achieving high performance for each function. However, to make this approach of iterative compilation more widely accepted and deployed in mainstream compilers, it is essential to modify existing algorithms, or develop new ones that find near-optimal solutions quickly. As a step in this direction, in this paper we attempt to identify and understand the important properties of some commonly employed heuristic search methods by using information collected during an exhaustive exploration of the phase order search space. We compare the performance obtained by each algorithm with all others, as well as with the optimal phase ordering performance. Finally, we show how we can use the features of the phase order space to improve existing algorithms as well as devise new, and better performing search algorithms. Prasad A. Kulkarni, David B. Whalley, Gary S. Tyson |
CGO | 3 |
| 2007 | Leveraging High Performance Data Cache Techniques to Save Power in Embedded Systems
Major Bhadauria, Sally A. McKee, Gary S. Tyson |
HiPEAC | 4 |
| 2007 | Addressing instruction fetch bottlenecks by using an instruction register fileabstractThe Instruction Register File (IRF) is an architectural extension for providing improved access to frequently occurring instructions. An optimizing compiler can exploit an IRF by packing an application's instructions, resulting in decreased code size, reduced energy consumption and improved execution time primarily due to a smaller footprint in the instruction cache. The nature of the IRF also allows the execution of packed instructions to overlap with instruction fetch, thus providing a means for tolerating increased fetch latencies, like those experienced by encrypted ICs as well as the presence of low-power L0 caches. Although previous research has focused on the direct benefits of instruction packing, this paper explores the use of increased fetch bandwidth provided by packed instructions. Small L0 caches improve energy efficiency but can increase execution time due to frequent cache misses. We show that this penalty can be significantly reduced by overlapping the execution of packed instructions with miss stalls. The IRF can also be used to supply additional instructions to a more aggressive execution engine, effectively reducing dependence on instruction cache bandwidth. This can improve energy efficiency, in addition to providing additional flexibility for evaluating various design tradeoffs in a pipeline with asymmetric instruction bandwidth. Thus, we show that the IRF is a complementary technique, operating as a buffer tolerating fetch bottlenecks, as well as providing additional fetch bandwidth for an aggressive pipeline backend. Stephen Roderick Hines, Gary S. Tyson, David B. Whalley |
LCTES | 2 |
| 2007 | Guaranteeing Hits to Improve the Efficiency of a Small Instruction CacheabstractVery small instruction caches have been shown to greatly reduce fetch energy. However, for many applications the use of a small filter cache can lead to an unacceptable increase in execution time. In this paper, we propose the tagless hit instruction cache (TH-IC), a technique for completely eliminating the performance penalty associated with filter caches, as well as a further reduction in energy consumption due to not having to access the tag array on cache hits. Using a few metadata bits per line, we are able to more efficiently track the cache contents and guarantee when hits will occur in our small TH-IC. When a hit is not guaranteed, we can instead fetch directly from the L1 instruction cache, eliminating any additional cycles due to a TH-IC miss. Experimental results show that the overall processor energy consumption can be significantly reduced due to the faster application running time and the elimination of tag comparisons for most of the accesses. Stephen Roderick Hines, David B. Whalley, Gary S. Tyson |
MICRO | 3 |
| 2006 | Adapting compilation techniques to enhance the packing of instructions into registersabstractThe architectural design of embedded systems is becoming increasingly idiosyncratic to meet varying constraints regarding energy consumption, code size, and execution time. Traditional compiler optimizations are often tuned for improving general architectural constraints, yet these heuristics may not be as beneficial to less conventional designs. Instruction packing is a recently developed compiler/architectural approach for reducing energy consumption, code size, and execution time by placing the frequently occurring instructions into an Instruction Register File (IRF). Multiple IRF instructions are made accessible via special packed instruction formats. This paper presents the design and analysis of a compilation framework and its associated optimizations for improving the efficiency of instruction packing. We show that several new heuristics can be developed for IRF promotion, instruction selection, register re-assignment and instruction scheduling, leading to significant reductions in energy consumption, code size, and/or execution time when compared to results using a standard optimizing compiler targeting the IRF. Stephen Roderick Hines, David B. Whalley, Gary S. Tyson |
CASES | 3 |
| 2006 | Exhaustive Optimization Phase Order Space ExplorationabstractThe phase-ordering problem is a long standing issue for compiler writers. Most optimizing compilers typically have numerous different code-improving phases, many of which can be applied in any order. These phases interact by enabling or disabling opportunities for other optimization phases to be applied. As a result, varying the order of applying optimization phases to a program can produce different code, with potentially significant performance variation amongst them. Complicating this problem further is the fact that there is no universal optimization phase order that will produce the best code, since the best phase order depends on the function being compiled, the compiler, and the target architecture characteristics. Moreover, finding the optimal optimization sequence for even a single function is hard as the space of attempted optimization phase sequences is huge and the interactions between different optimizations are poorly understood. Most previous studies performed to search for the most effective optimization phase sequence assume the optimization phase order search space to be extremely large, and hence consider exhaustive exploration of this space infeasible. In this paper we show that even though the attempted search space is extremely large, with careful and aggressive pruning it is possible to limit the actual search space with no loss of information so that it can be completely evaluated in a matter of minutes or a few hours for most functions. We were able to exhaustively enumerate all the possible function instances that can be produced by different phase orderings performed by our compiler for more than 98% of the functions in our benchmark suite. In this paper we describe the algorithm we used to make exhaustive search of the optimization phase order space possible. We then analyze this space to automatically calculate relationships between different phases. Finally, we show that the results of this analysis can be used to reduce the compilation time for a conventional batch compiler. Prasad A. Kulkarni, David B. Whalley, Gary S. Tyson, Jack W. Davidson |
CGO | 3 |
| 2006 | Reducing the cost of conditional transfers of control by using comparison specificationsabstractA significant portion of a program's execution cycles are typically dedicated to performing conditional transfers of control. Much of the research on reducing the costs of these operations has focused on the branch, while the comparison has been largely ignored. In this paper we investigate reducing the cost of comparisons in conditional transfers of control. We decouple the specification of the values to be compared from the actual comparison itself, which now occurs as part of the branch instruction. The specification of the register or immediate values involved in the comparison is accomplished via a new instruction called a comparison specification, which is loop invariant. Decoupling the specification of the comparison from the actual comparison performed before the branch reduces the number of instructions in the loop, which provides performance benefits not possible when using conventional comparison instructions. Results from applying this technique on the ARM processor show that both the number of instructions executed and execution cycles are reduced. William C. Kreahling, Stephen Roderick Hines, David B. Whalley, Gary S. Tyson |
LCTES | 4 |
| 2006 | In search of near-optimal optimization phase orderingsabstractPhase ordering is a long standing challenge for traditional optimizing compilers. Varying the order of applying optimization phases to a program can produce different code, with potentially significant performance variation amongst them. A key insight to addressing the phase ordering problem is that many different optimization sequences produce the same code. In an earlier study, we used this observation to restate the phase ordering problem to concentrate on finding all distinct function instances that can be produced due to different phase orderings, instead of attempting to generate code for all possible optimization sequences. Using a novel search algorithm we were able to show that it is possible to exhaustively enumerate the set of all possible function instances that can be produced by different phase orderings in our compiler for most of the functions in our benchmark suite [1]. Finding the optimal function instance within this set for almost any dynamic measure of performance still appears impractical since that would involve execution/simulation of all generated function instances. To find the dynamically optimal function instance we exploit the observation that the enumeration space for a function typically contains a very small number of distinct control flow paths. We simulate only one function instance from each group of function instances having the identical control flow, and use that information to estimate the dynamic performance of the remaining functions in that group. We further show that the estimated dynamic frequency counts obtained by using our method correlate extremely well to simulated processor cycle counts. Thus, by using our measure of dynamic frequencies to identify a small number of the best performing function instances we can often find the optimal phase ordering for a function within a reasonable amount of time. Finally, we perform a case study to evaluate how adept our genetic algorithm is for finding optimal phase orderings within our compiler, and demonstrate how the algorithm can be improved. Prasad A. Kulkarni, David B. Whalley, Gary S. Tyson, Jack W. Davidson |
LCTES | 3 |
| 2005 | Beyond Basic Region Caching: Specializing Cache Structures for High Performance and Energy Conservation
Michael J. Geiger, Sally A. McKee, Gary S. Tyson |
HiPEAC | 3 |
| 2005 | Improving Program Efficiency by Packing Instructions into RegistersabstractNew processors, both embedded and general purpose, often have conflicting design requirements involving space, power, and performance. Architectural features and compiler optimizations often target one or more design goals at the expense of the others. This paper presents a novel architectural and compiler approach to simultaneously reduce power requirements, decrease code size, and improve performance by integrating an instruction register file (IRF) into the architecture. Frequently occurring instructions are placed in the IRF. Multiple entries in the IRF can be referenced by a single packed instruction in ROM or LI instruction cache. Unlike conventional code compression, our approach allows the frequent instructions to be referenced in arbitrary combinations. The experimental results show significant improvements in space and power, as well as some improvement in execution time when using only 32 entries. These advantages make packing instructions into registers an effective approach for improving overall efficiency. Stephen Roderick Hines, Joshua Green, Gary S. Tyson, David B. Whalley |
ISCA | 3 |
| 2005 | PowerFITS: Reduce Dynamic and Static I-Cache Power Using Application Specific Instruction Set SynthesisabstractPower consumption, performance, area, and cost are critical concerns in designing microprocessors for embedded systems such as portable handheld computing and personal telecommunication devices. In previous work [A. Cheng et al., (2004)], we introduced the concept of framework-based instruction-set tuning synthesis (FITS), which is a new instruction synthesis paradigm that falls between a general-purpose embedded processor and a synthesized application specific processor (ASP). We address these design constraints through FITS by improving the code density. A FITS processor improves code density by tailoring the instruction set to the requirement of a target application to reduce the code size. This is achieved by replacing the fixed instruction and register decoding of general purpose embedded processor with programmable decoders that can achieve ASP performance, low power consumption, and compact chip area with the fabrication advantages of a mass produced single chip solution to amortize the cost. Instruction cache has been recognized as one of the most predominant source of power dissipation in a microprocessor. For instance, in Intel's StrongARMprocessor, 27% of total chip power loss goes into the instruction cache [J. Montanaro et al., (1996)]. In this paper, we demonstrate how FITS can be applied to improve the instruction cache power efficiency. Experimental results show that our synthesized instruction sets result in significant power reduction in the instruction cache compared to ARM instructions. For 21 benchmarks from the MiBench suite [M. Guthaus et al., (2001)], our simulation results indicate on average: a 49.4% saving for switching power; a 43.9% saving for internal power; a 14.9% saving for leakage power; a 46.6% saving for total cache power with up to 60.3% saving for peak power Allen C. Cheng, Gary S. Tyson, Trevor N. Mudge |
ISPASS | 2 |
| 2005 | Reducing Instruction Fetch Cost by Packing Instructions into RegisterWindowsabstractInstruction packing is a combination compiler/architectural approach that allows for decreased code size, reduced power consumption and improved performance. The packing is obtained by placing frequently occurring instructions into an instruction register file (IRF). Multiple IRF entries can then be accessed using special packed instructions. Previous IRF efforts focused on using a single 32-entry register file for the duration of an application. This paper presents software and hardware extensions to the IRF supporting multiple instruction register windows to allow a greater number of relevant instructions to be available for packing in each function. Windows are shared among similar functions to reduce the overall costs involved in such an approach. The results indicate that significant improvements in instruction fetch cost can be obtained by using this simple architectural enhancement. We also show that using an IRF with a loop cache, which is also used to reduce energy consumption, results in much less energy consumption than using either feature in isolation. Stephen Roderick Hines, Gary S. Tyson, David B. Whalley |
MICRO | 2 |
| 2005 | An Energy Efficient Instruction Set Synthesis Framework for Low Power Embedded System DesignsabstractEnergy efficiency, performance, area, and cost are critical concerns in designing microprocessors for embedded systems such as portable handheld computing and personal telecommunication devices. This work introduces the concept of framework-based instruction set synthesis (FITS), which is a new instruction synthesis paradigm that falls between a general-purpose embedded processor and a synthesized application specific processor (ASP). FITS processors reduce code size and energy consumption by tailoring the instruction set to the requirement of a target application. This is achieved by replacing the fixed instruction and register decoding of a general purpose embedded processor with programmable decoders that can achieve ASP performance, low energy consumption, and smaller code size with the fabrication advantages of a mass produced single chip solution. Experimental results show that our synthesized instruction sets result in significant power reduction in the level one instruction cache compared to ARM instructions. The instruction cache is one of the most predominant sources of power dissipation on the processor. For instance, in Intel's StrongARM processor, 27 percent of total chip power loss goes into the instruction cache. For 21 MiBench benchmarks, our simulation results, indicate, on average, a 49.4 percent saving for switching power, a 43.9 percent saving for internal power, a 14.9 percent saving for leakage power, a 46.6 percent saving for total instruction cache power with up to 60.3 percent saving for peak power. Allen C. Cheng, Gary S. Tyson |
IEEE Trans. Computers | 2 |
| 2004 | FITS: framework-based instruction-set tuning synthesis for embedded application specific processorsabstractWe propose a new instruction synthesis paradigm that falls between a general-purpose embedded processor and a synthesized application specific processor (ASP). This is achieved by replacing the fixed instruction and register decoding of general purpose embedded processor with programmable decoders that can achieve ASP performance with the fabrication advantages of a mass produced single chip solution. Allen C. Cheng, Gary S. Tyson, Trevor N. Mudge |
DAC | 2 |
| 2004 | A Prefetch TaxonomyabstractThe growing difference between processor and main memory cycle time demands the use of aggressive prefetch algorithms to reduce the effective memory access latency. However, prefetching can significantly increase memory traffic and unsuccessful prefetches may pollute the cache. Metrics such as coverage and accuracy result from a simplistic classification of individual prefetches as "good" or "bad." They do not capture the full effect of each prefetch and, hence, do not accurately reflect the quality of the prefetch algorithm. Gross statistics such as changes in the number of misses, total traffic, and IPC are not attributable to individual prefetches. Such gross metrics are therefore useful only for ranking existing prefetch algorithms; they do not evaluate the effect of individual prefetches so that an algorithm might be tuned. We introduce a new, accurate, and complete taxonomy, called the Prefetch Traffic and Miss Taxonomy (PTMT), for classifying each prefetch by precisely accounting for the difference in traffic and misses it generates, either directly or indirectly. We illustrate the use of PTMT by evaluating two data prefetch algorithms. Vijayalakshmi Srinivasan, Edward S. Davidson, Gary S. Tyson |
IEEE Trans. Computers | 3 |
| 2001 | Stack Value File: Custom Microarchitecture for the StackabstractAs processor performance increases, there is a corresponding increase in the demands on the memory system, including caches. Research papers have proposed partitioning the cache into instruction/data, temporal/non-temporal, and/or stack/non-stack regions. Each of these designs can improve performance by constructing two separate structures which can be probed in parallel while reducing contention. In this paper, we propose a new memory organization that partitions data references into stack and nonstack regions. Non-stack references are routed to a conventional cache. Stack references, on the other hand, are shown to have several characteristics that can be leveraged to improve performance using a less conventional storage organization. This paper enumerates those characteristics and proposes a new microarchitectural feature, the stack value file (SVF), which exploits them to improve instruction-level parallelism, reduce stack access latencies, reduce demand on the first-level cache, and reduce data bus traffic. Our results show that the SVF can improve execution performance by 29 to 65% while reducing overhead traffic for the stack region by many orders of magnitude over cache structures of the same size. Hsien-Hsin S. Lee, Mikhail Smelyanskiy, Chris J. Newburn, Gary S. Tyson |
HPCA | 4 |
| 2001 | Branch History Guided Instruction PrefetchingabstractInstruction cache misses stall the fetch stage of the processor pipeline and hence affect instruction supply to the processor. Instruction prefetching has been proposed as a mechanism to reduce instruction cache (I-cache) misses. However, a prefetch is effective only if accurate and initiated sufficiently early to cover the miss penalty. This paper presents a new hardware-based instruction prefetching mechanism, Branch History Guided Prefetching (BHGP), to improve the timeliness of instruction prefetches. BHGP correlates the execution of a branch instruction with I-cache misses and uses branch instructions to trigger prefetches of instructions that occur (N-1) branches later in the program execution, for a given N>1. Evaluations on commercial applications, windows-NT applications, and some CPU2000 applications show an average reduction of 66% in miss rate over all applications. BHGP improved the IPC bp 12 to 14% for the CPU2000 applications studied; on average 80% of the BHGP prefetches arrived in cache before their next use, even on a 4-wide issue machine with a 15 cycle L2 access penalty. Vijayalakshmi Srinivasan, Edward S. Davidson, Gary S. Tyson, Mark J. Charney, Thomas R. Puzak |
HPCA | 3 |
| 2001 | Allocation by Conflict: A Simple Effective Multilateral Cache Management SchemeabstractSeveral schemes have been proposed that incorporate an auxiliary buffer to improve the performance of a given size cache. Victim caching, aims to reduce the impact of conflict misses in direct-mapped caches. Victim offers competitive performance benefits, but requires a costly data path for swaps and saves between the main cache and the added buffer. Several multilateral schemes (e.g. NTS, PCS) offer competitive performance with Victim across a wide range of associativities, but require no swap/save data path. While these schemes perform well overall, their overall performance lags that of Victim when the main cache is direct-mapped. Furthermore, they also require costly hardware support, but in the form of history tables for maintaining allocation decision information. The paper introduces a multilateral cache management scheme, allocation by conflict (ABC), which generally outperforms Victim, NTS, and PCS. Furthermore, ABC has the lowest hardware requirements of any multilateral scheme-only a single additional bit per block in the main cache is required to maintain usage information for the allocation decision process, and no swap/save data path is needed. Edward S. Tam, Stevan A. Vlaovic, Gary S. Tyson, Edward S. Davidson |
ICCD | 3 |
| 2001 | Evaluating the Use of Register Queues in Software Pipelined LoopsabstractIn this paper, we examine the effectiveness of a new hardware mechanism, called register queues (RQs), which effectively decouples the architected register space from the physical registers. Using RQs, the compiler can allocate physical registers to store live values in the software pipelined loop while minimizing the pressure placed on architected registers. We show that decoupling the architected register space from the physical register space can greatly increase the applicability of software pipelining, even as memory latencies increase. RQs combine the major aspects of existing rotating register file and register connection techniques to generate efficient software pipeline schedules. Through the use of RQs, we can minimize the register pressure and code expansion caused by software pipelining. We demonstrate the effect of incorporating register queues and software pipelining with 983 loops taken from the Perfect Club, the SPEC suites, and the Livermore Kernels. Gary S. Tyson, Mikhail Smelyanskiy, Edward S. Davidson |
IEEE Trans. Computers | 1 |
| 2000 | Region-based caching: an energy-delay efficient memory architecture for embedded processorsabstractPower consumption has been a major concern in designing microprocessors for portable systems such as notebook computers, hand-held computing and personal telecommunication devices.As these devices increase in popularity and are used in a wider range of applications, a low p o wer design becomes more critical.In this paper, we propose a new microarchitectural data cache design called region-based c aching that can reduce power consumption.Power savings is achieved by re-organizing the the rst level cache to more eÆciently exploit memory reference characteristics produced by programming language semantics.These characteristics enable the cache to be partitioned by memory region (stack, global, heap), reducing power consumption, while retaining comparable performance to a conventional cache design.Applications from the MediaBench b e n c hmark suite indicate that a design with two additional small region-based caches results in 66% reduction in average in energy-delay product.Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page.To copy otherwise, to republish, to post on servers or to redistribute to lists, requires prior specific Hsien-Hsin S. Lee, Gary S. Tyson |
CASES | 2 |
| 2000 | Instruction overhead and data locality effects in superscalar processorsabstractTo reduce software development and maintenance costs, programmers are increasingly using object oriented programming languages, such as C++, and relying on highly flexible data structures, such as linked lists. Object oriented programming languages provide features that help manage complex software systems, but object oriented programs tend to suffer increased instruction counts, e.g. due to generalized class implementations and many more calls to small functions. Using linked data structures increases programming flexibility by allowing easy addition and deletion of nodes, and by dynamically allocating memory to satisfy applications that use large memory space. However, successive elements in linked data structures may be allocated noncontinuously in memory, leading to poor spatial locality for list traversals which in turn increases cache misses and reduces performance. This paper evaluates the impact of both the increased instruction overhead and poor spatial locality on superscalar processor performance as issue width increases. We show that underutilized resources of wide-issue processors can partially alleviate the impact of the instruction overhead. However, poor locality tends to cause more performance degradation as the processor issue width increases. Finally we show that the spatial locality of some programs can be improved by using a vector representation to replace linked list structures. Vectors exhibit better spatial locality during list traversals, but suffer from instruction overhead and memory copy overhead when nodes are added to and deleted from the structure. Murali Annavaram, Gary S. Tyson, Edward S. Davidson |
ISPASS | 2 |
| 2000 | Quantifying instruction-level parallelism limits on an EPIC architectureabstractEPIC architectures rely heavily on state-of-the-art compiler technology to deliver optimal performance while keeping hardware design simple. It is generally believed that an optimizing compiler has an enormous scheduling window to exploit instruction-level parallelism (ILP) since the compiler orchestrates the entire program. Many state-of-the-art compilers typically confine optimizations to loop boundaries (e.g. software pipelining, trace scheduling, and loop unrolling) and function boundaries (e.g. loop peeling, loop exchanges, invariant hoisting, and global optimizations). Although techniques such as function inlining and interprocedural optimizations can alleviate these constraints to a limited extent, loop and function boundaries are often the real scopes of the compiler scheduler. Several previous ILP studies have explored the limits of parallelism on dynamic superscalar machines; however, those results are not applicable to EPIC architectures since they rely on dynamic scheduling, not static code scheduling by the compiler, to reorder instructions. In this paper, we evaluate the limits in ILP obtained through compiler scheduling alone. We quantify these limits as more restrictive scheduling constraints are imposed-starting from inter-procedural code scheduling, to intra-procedural and finally to loop-confined code scheduling. Hsien-Hsin S. Lee, Youfeng Wu, Gary S. Tyson |
ISPASS | 3 |
| 2000 | Eager writeback - a technique for improving bandwidth utilizationabstractModern high-performance processors utilize multi-level cache structures to help tolerate the increasing latency of main memory. Most of these caches employ either a writeback or a write-through strategy to deal with store operations. Write-through caches propagate data to more distant memory levels at the time each store occurs, which requires a very large bandwidth between the memory hierarchy levels. Writeback caches can significantly reduce the bandwidth requirements between caches and memory by marking cache lines as dirty when stores are processed and writing those lines to the memory system only when that dirty line is evicted. Unfortunately, for applications that experience significant numbers of cache misses due to streaming data, writeback cache designs can degrade overall system performance by clustering bus activity when dirty lines contend with data being fetched into the cache. In this paper we present a new technique called Eager Writeback, which re-distributes and balances memory traffic by writing and "cleaning" dirty cache lines prior to their eviction. Eager Writeback can be viewed as a compromise between write-through and writeback policies, in which dirty lines are written later than write-through, but prior to writeback. We will show that this approach can reduce the large number of writes seen in a write-through design, while avoiding the performance degradation caused by clustering bus traffic in a writeback approach. Hsien-Hsin S. Lee, Gary S. Tyson, Matthew K. Farrens |
MICRO | 2 |
| 2000 | Improving BTB performance in the presence of DLLsabstractDynamically Linked Libraries (DLLs) promote software modularity, portability, and flexibility and their use has become widespread. The authors characterize the behavior of five applications that make heavy use of DLLs, with a particular focus on the effects of DLLs on Branch Target Buffer (BTB) performance. DLLs aggravate hot set contention in the BTB. Standard software remedies are ineffective because the DLLs are shared, compiled separately, and dynamically linked to applications. We propose a hardware technique, the DLL BTB, that adds a small second buffer to the BTB and dedicates it to storing DLL target addresses. We show that the DLL BTB performance is similar to a BTB with a victim buffer, but the DLL BTB requires no parallel lookups or datapaths between the original BTB and the added buffer. Stevan A. Vlaovic, Edward S. Davidson, Gary S. Tyson |
MICRO | 3 |
| 1999 | Classifying load and store instructions for memory renamingabstractMemory operations remain a significant bottleneck in dynamically scheduled pipelined processors, due in part to the inability to statically determine the existence of memory address dependencies. Hardware memory renaming techniques have been proposed to predict which stores a load might be dependent upon. These prediction techniques can be used to speculatively forward a value from a predicted store dependency to a load through a value prediction table. However, these techniques require large, timeconsuming hardware tables. In this paper we propose a software-guided approach for identifying dependencies between store and load instructions and the Load Marking (LM) architecture to communicate these dependencies to the hardware. Compiler analysis and profiles are used to find important store/load relationships, and these relationships are identified during execution via hints or an n-bit tag. For those loads that are not marked for renaming, we then use additional profiling information ... Glenn Reinman, Brad Calder, Dean M. Tullsen, Gary S. Tyson, Todd M. Austin |
International Conference on Supercomputing | 4 |
| 1999 | Active Management of Data Caches by Exploiting Reuse InformationabstractAs microprocessor speeds continue to outpace memory subsystems in speed, minimizing average data access time grows in importance. Multilateral caches afford an opportunity to reduce the average data access time by active management of block allocation and replacement decisions. We evaluate and compare the performance of traditional caches and multilateral caches with three active block allocation schemes: MAT, NTS, and PCS. We also compare the performance of NTS and PCS to multilateral caches with a near-optimal, but nonimplementable policy, pseudo-opt, that employs future knowledge to achieve both active allocation and active replacement. NTS and PGS are evaluated relative to pseudo-opt with respect to miss ratio, accuracy of predicting reference locality, actual usage accuracy, and tour lengths of blocks in the cache. Results show that the multilateral schemes do outperform traditional cache management schemes, but fall short of pseudo-opt; increasing their prediction accuracy and incorporating active replacement decisions would allow them to more closely approach pseudo-opt performance. Edward S. Tam, Jude A. Rivers, Vijayalakshmi Srinivasan, Gary S. Tyson, Edward S. Davidson |
IEEE Trans. Computers | 4 |
| 1998 | Evaluating the performance of active cache management schemesabstractIn this paper we examine the performance of two multi-lateral cache schemes; one makes block allocation decisions correlated to the reference behavior of regions of memory (NTS), the other correlated to the reference behavior of memory accessing instructions (PCS). To determine the efficacy of exploiting these reference correlation schemes to improve cache management, we compare the performance of these multi-lateral schemes to a multi-lateral configuration that uses a near-optimal (but non-implementable) replacement policy, pseudo-opt. In addition to miss ratio, three metrics are used to evaluate the performance of these schemes, relative to pseudo-opt: 1) prediction accuracy in determining reference locality, 2) actual usage accuracy, i.e. how likely a block in an implementable scheme and the near-optimal scheme exhibit the same reuse characteristic, and 3) tour length of a line in the cache. Results show that while the NTS and PCS schemes outperform traditional cache management schemes, they fall short of the pseudo-opt performance; this is due to their simple prediction strategies and because their active management addresses only block allocation and not replacement. Edward S. Tam, Jude A. Rivers, Vijayalakshmi Srinivasan, Gary S. Tyson, Edward S. Davidson |
ICCD | 4 |
| 1998 | Utilizing Reuse Information in Data Cache ManagementabstractAs microprocessor speeds continue to outgrow memory subsystem speeds, minimizing the average data access time grows in importance. As current data caches are often poorly and inefficiently managed, a good management technique can improve the average data access time. This paper presents a comparative evaluation of two approaches that utilize reuse information for more efficiently managing the firstlevel cache. While one approach is based on the effective address of the data being referenced, the other uses the program counter of the memory instruction generating the reference. Our evaluations show that using effective address reuse information performs better than using program counter reuse information. In addition, we show that the Victim cache performs best for multi-lateral caches with a direct-mapped main cache and high L2 cache latency, while the NTS (effective-addressbased) approach performs better as the L2 latency decreases or the associativity of the main cache increases. Jude A. Rivers, Edward S. Tam, Gary S. Tyson, Edward S. Davidson, Matthew K. Farrens |
International Conference on Supercomputing | 3 |
| 1998 | mlcache: A Flexible Multi-Lateral Cache SimulatorabstractAs the gap between processor and memory speeds increases, cache performance becomes more critical to overall system performance. Multi-lateral cache designs such as the Assist, Victim, and NTS cache have been shown to perform as well as or better than larger, single structure caches. Unlike current cache simulators, mlcache (an event-driven, timing-sensitive simulator based on the Latency Effects cache timing model) can evaluate a variety of multilateral cache configurations. It was developed to help designers in the middle of the design cycle decide which cache configuration would best meet the performance needs of the target processor. It can easily model various cache configurations by using its library of cache state and data movement routines. We use the SPEC95 benchmarks to illustrate how mlcache can be used to compare the performance of several different data cache configurations. Edward S. Tam, Jude A. Rivers, Gary S. Tyson, Edward S. Davidson |
MASCOTS | 3 |
| 1998 | Analyzing the Working Set Characteristics of Branch ExecutionabstractTo achieve highly accurate branch prediction, it is necessary not only to allocate more resources to branch prediction hardware but also to improve the understanding of branch execution characteristics. In this paper, we present a new profile-based conditional branch analysis technique called branch working set analysis to provide additional information about control flow behavior of general purpose applications. This analysis evaluates the dynamic behavior of branch execution by partitioning either individual branches or pre-classified branch groups into sets based on temporal locality and ordering information. We refer to these sets as the working sets of branches. To demonstrate the usefulness of this form of analysis, we examine the efficiency of current allocation techniques for branch history table (BHT) space and propose a new solution to this allocation process that improves the performance of these tables. In our approach the mapping between branch instructions and BHT entries is specified during compilation to reduce table contention-leading to more relevant histories and improved predictor performance. As a result, even for programs with a large number of static branches, only 100 to 200 history entries are needed to approximate the performance of larger 1024-entry BHT. Furthermore, when the technique is applied to a predictor with 1024-entry BHT its prediction accuracy is improved by 16%-comparable with the performance of a BHT of infinite capacity. Sangwook P. Kim, Gary S. Tyson |
MICRO | 2 |
| 1997 | On High-Bandwidth Data Cache Design for Multi-Issue ProcessorsabstractHighly aggressive multi-issue processor designs of the past few years and projections for the next decade require that we redesign the operation of the cache memory system. The number of instructions that must be processed (including correctly predicted ones) will approach 16 or more per cycle. Since memory operations account for about a third of all instructions executed these systems will have to support multiple data references per cycle. We explore reference stream characteristics to determine how best to meet the need for ever increasing access rates. We identify limitations of existing multi-ported cache designs and propose a new structure, the locality-based interleaved cache (LBIC), to exploit the characteristics of the data reference stream while approaching the economy of traditional multi-bank cache design. Experimental results show that the LBIC structure is capable of outperforming current multi-ported approaches. Jude A. Rivers, Gary S. Tyson, Edward S. Davidson, Todd M. Austin |
MICRO | 2 |
| 1997 | Improving the Accuracy and Performance of Memory Communication Through RenamingabstractAs processors continue to exploit more instruction-level parallelism, a greater demand is placed on reducing the effects of memory access latency. In this paper, we introduce a novel modification of the processor pipeline called memory renaming. Memory renaming applies register access techniques to load instructions, reducing the effect of delays caused by the need to calculate effective addresses for the load and all preceding stores before the data can be fetched. Memory renaming allows the processor to speculatively fetch values when the producer of the data can be reliably determined without the need for an effective address. This work extends previous studies of data value and dependence speculation. When memory renaming is added to the processor pipeline, renaming can be applied to 30% to 50% of all memory references, translating to an overall improvement in execution time of up to 41%. Furthermore, this improvement is seen across all memory segments-including the heap segment, which has often been difficult to manage efficiently. Gary S. Tyson, Todd M. Austin |
MICRO | 1 |
| 1995 | A modified approach to data cache managementabstractAs processor performance continues to improve, more emphasis must be placed on the performance of the memory system. In this paper, a detailed characterization of data cache behavior for individual load instructions is given. We show that by selectively applying cache line allocation according the characteristics of individual load instructions, overall performance can be improved for both the data cache and the memory system. This approach can improve some aspects of memory performance by as much as 60 percent on existing executables. Gary S. Tyson, Matthew K. Farrens, John Matthews, Andrew R. Pleszkun |
MICRO | 1 |
| 1994 | A Study of Single-Chip Processor/Cache Organizations for Large Numbers of TransistorsabstractPresents a trace-driven simulation-based study of a wide range of cache configurations and processor counts. This study was undertaken in an attempt to help answer the question of how best to allocate large numbers of transistors, a question that is rapidly increasing in importance as transistor densities continue to climb. At what point does continuing to increase the size of the on-chip first level cache cease to provide sufficient increases in hit rate and become prohibitively difficult to access in a single cycle? In order to compare different configurations, the concept of an Equivalent Cache Transistor is presented. Results indicate that the access time of the first-level data cache is more important than the size. In addition, it appears that once approximately 15 million transistors become available, a two processor configuration is preferable to a single processor with correspondingly larger caches.> Matthew K. Farrens, Gary S. Tyson, Andrew R. Pleszkun |
ISCA | 2 |
| 1994 | The effects of predicated execution on branch predictionabstractHigh performance architectures have always had to deal with the performance limiting impact of branch operations. Microprocessor designs are going to have to deal with this problem as well, as they move towards deeper pipelines and support for multiple instruction issue. Branch prediction schemes are often used to alleviate the negative impact of branch operations by allowing the speculative execution of instructions after an unresolved branch. Another technique is to eliminate branch instructions altogether. Predication can remove forward branch instructions by translating the instructions following the branch into predicate form. Gary S. Tyson |
MICRO | 1 |
| 1993 | Techniques for extracting instruction level parallelism on MIMD architecturesabstractExtensive research has been done on extracting parallelism from single instruction stream processors. The authors present some results of an investigation into ways to modify MIMD architectures to allow them to extract the instruction level parallelism achieved by current superscalar and VLIW machines. A new architecture is proposed which utilizes the advantages of a multiple instruction stream design while addressing some of the limitations that have prevented MIMD architectures from performing ILP operation. A new code scheduling mechanism is described to support this new architecture by partitioning instructions across multiple processing elements in order to exploit this level of parallelism.> Gary S. Tyson, Matthew K. Farrens |
MICRO | 1 |
| 1992 | A partitioned translation lookaside buffer approach to reducing address bandwithabstractSimulations indicate a simple modification of existing virtual memory hardware can significantly reduce the number of pins required to transmit address information from processor to off-chip memory. This modification consists of partitioning a TLB so that virtual page numbers are stored in a cache on the processor and corresponding real page numbers are sotred in registers at the memory, making it possible to transmit a small register index instead of the entire real page number. Matthew K. Farrens, Arvin Park, Rob Fanfelle, Pius Ng, Gary S. Tyson |
ISCA | 5 |
| 1992 | Modifying VM hardware to reduce address pin requirements
Matthew K. Farrens, Arvin Park, Gary S. Tyson |
MICRO | 3 |
| 1992 | MISC: a Multiple Instruction Stream ComputerabstractThis paper describes a single chip Multiple Instruction Stream Computer (MISC) capable of extracting instruction level parallelism from a broad spectrum of programs. The MISC architecture uses multiple asynchronous processing elements to separate a program into streams that can be executed in parallel, and integrates a conflict-free message passing system into the lowest level of the processor design to facilitate low latency intra-MISC communication. This approach allows for increased machine parallelism with minimal code expansion, and provides an alternative approach to single instruction stream multi-issue machines such as SuperScalar and VLIW. # # 1. Introduction The goal of most high-performance computers is to maximize the amount of work that can be done per unit time. This quantity (work) can be expressed by the following equation [HePa90]: work = clock rate× instruction count 1 ############### × Clocks per Instruction 1 ################### A number of different approa... Gary S. Tyson, Matthew K. Farrens, Andrew R. Pleszkun |
MICRO | 1 |