EDBT 2026 Demo / reviewers in the wild / expert
Thomas M. Conte
dblp:91/1557 · also Tom Conte 0001
· DBLP profile ↗
56ranked-venue papers
15as first author
6since 2021 · last 2025
0000-0001-7037-2377ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 50 · 13 first-author · 5 since 2021Software engineering, systems software and programming languages · 11 · 3 first-author · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RASSM: Residue-based Acceleration of Single Sparse Matrix Computation via Adaptive Tiling
Anirudh Jain, Pulkit Gupta, Thomas M. Conte |
ASPLOS (1) | 3 |
| 2025 | ASDF: A Compiler for Qwerty, a Basis-Oriented Quantum Programming LanguageabstractQwerty is a high-level quantum programming language built on bases and functions rather than circuits. This new paradigm introduces new challenges in compilation, namely synthesizing circuits from basis translations and automatically specializing adjoint or predicated forms of functions. This paper presents ASDF, an open-source compiler for Qwerty that answers these challenges in compiling basis-oriented languages. Enabled with a novel high-level quantum IR implemented in the MLIR framework, our compiler produces OpenQASM 3 or QIR for either simulation or execution on hardware. Our compiler is evaluated by comparing the fault-tolerant resource requirements of generated circuits with other compilers, finding that ASDF produces circuits with comparable cost to prior circuit-oriented compilers. Austin J. Adams, Sharjeel Khan, Arjun S. Bhamra, Ryan R. Abusaada, Anthony M. Cabrera, Cameron C. Hoechst, Travis S. Humble, Jeffrey Young 0001, Thomas M. Conte |
CGO | 9 |
| 2025 | A Blueprint for Q-CS1, an Introductory Quantum Programming CourseabstractDespite the need to build a quantum workforce, current courses that introduce quantum programming are rooted in quantum notation that students may find intimidating. We propose Q-CS1, a quantum equivalent of CS1 that begins with hands-on quantum programming. Q-CS1 is enabled by the Qwerty quantum programming language, which allows for reasoning about qubit behavior without physics notation or quantum circuits. An outline of Q-CS1 is provided along with plans for assessing its effectiveness. Austin J. Adams, Rodrigo Borela, Jeffrey Young 0001, Thomas M. Conte |
SIGCSE (2) | 4 |
| 2022 | "Smarter" NICs for faster molecular dynamics: a case studyabstractThis work evaluates the benefits of using a “smart” network interface card (SmartNIC) as a compute accelerator for the example of the MiniMD molecular dynamics proxy application. The accelerator is NVIDIA's BlueField-2 card, which includes an 8-core Arm processor along with a small amount of DRAM and storage. We test the networking and data movement performance of these cards compared to a standard Intel server host using microbenchmarks and MiniMD. In MiniMD, we identify two distinct classes of computation, namely core computation and maintenance computation, which are executed in sequence. We restructure the algorithm and code to weaken this dependence and increase task parallelism, thereby making it possible to increase utilization of the BlueField-2 concurrently with the host. We evaluate our implementation on a cluster consisting of 16 dual-socket Intel Broadwell host nodes with one BlueField-2 per host-node. Our results show that while the overall compute performance of BlueField-2 is limited, using them with a modified MiniMD algorithm allows for up to 20% speedup over the host CPU baseline with no loss in simulation accuracy. Sara Karamati, Clay Hughes, Karl S. Hemmert, Ryan E. Grant, Whit Schonbein, Scott Levy, Thomas M. Conte, Jeffrey Young 0001, Richard W. Vuduc |
IPDPS | 7 |
| 2022 | Scalable Energy-Efficient Microarchitectures With Computational Error Tolerance Via Redundant Residue Number SystemsabstractDue to high leakage current and threshold voltage, Dennard scaling has reached its limit on conventional semiconductor technology. Energy reduction at the transistor level by simply lowering supply voltage has proven to be infeasible for these devices (e.g., MOSFETs). Some recently proposed millivolt switch techniques aim to mitigate these issues, by maintaining a high on/off ratio of drain currents with a much lower supply voltage. However,$V_{dd}$reduction is constrained by high intermittent error probabilities in millivolt switches. Energy-efficient microarchitectures that are computationally error-tolerant are therefore urgently needed. This article systematically leverages the error correction and checkpointing properties of Redundant Residue Number Systems (RRNS) by varying the number of non-redundant ($n$) and redundant ($r$) residues. The state-of-the-art of RRNS microarchitecture is confined to a fixed configuration point within such a($n$n,$r$r)-RRNSdesign plane, as it supports single error correction alone. Being able to efficiently handle resilience in this($n$n,$r$r)-RRNSplane significantly improves reliability, allowing further${V_{dd}}$reduction to save energy. To this end, first, we propose a scalable RRNS microarchitecture that simultaneously supports both, error-correction, as well as checkpointing with restart capabilities upon detecting uncorrectable errors. Second, we design a novel RRNS-based adaptive checkpointing&restart mechanisms that automatically guarantees reliability while minimizing the energy-delay product (EDP). To the best of our knowledge, these are the first set of checkpointing mechanisms targeting the RRNS infrastructure. Moreover, these mechanisms optimize the usage efficiency of memory capacity. Third, we systematically explore the RRNS design space to find the best ($n$,$r$) configuration point. For similar reliability when compared to a conventional binary core without computationally error-tolerant (runs at high$V_{dd}$), the proposed RRNS scalable microarchitecture reduces EDP by 53 percent on average for memory-intensive workloads and by 67 percent on average for non-memory-intensive workloads. Bobin Deng, Sriseshan Srikanth, Anirudh Jain, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook |
IEEE Trans. Computers | 4 |
| 2021 | SortCache: Intelligent Cache Management for Accelerating Sparse Data WorkloadsabstractSparse data applications have irregular access patterns that stymie modern memory architectures. Although hyper-sparse workloads have received considerable attention in the past, moderately-sparse workloads prevalent in machine learning applications, graph processing and HPC have not. Where the former can bypass the cache hierarchy, the latter fit in the cache. This article makes the observation that intelligent, near-processor cache management can improve bandwidth utilization for data-irregular accesses, thereby accelerating moderately-sparse workloads. We propose SortCache, a processor-centric approach to accelerating sparse workloads by introducing accelerators that leverage the on-chip cache subsystem, with minimal programmer intervention. Sriseshan Srikanth, Anirudh Jain, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook |
ACM Trans. Archit. Code Optim. | 3 |
| 2020 | Special Session: Exploring the Ultimate Limits of Adiabatic CircuitsabstractThe field of adiabatic circuits is rooted in electronics know-how stretching all the way back to the 1960s and has potential applications in vastly increasing the energy efficiency of far-future computing. But now, the field is experiencing an increased level of attention in part due to its potential to reduce the vulnerability of systems to side-channel attacks that exploit, e.g., unwanted EM emissions, power supply fluctuations, and so forth. In this context, one natural question is: Just how low can the energy dissipation from adiabatic circuits, and the associated extraneous signal emissions, be made to go? We argue that the ultimate limits of this approach lie much farther away than is commonly appreciated. Recent advances at Sandia National Laboratories in the design of fully static, fully adiabatic CMOS logic styles and high-quality energy-recovering resonant power-clock drivers offer the potential to reduce dynamic switching losses by multiple orders of magnitude, and, particularly for cryogenic applications, optimization of device structures can reduce the standby power consumption of inactive devices, and the ultimate dissipation limits of the adiabatic approach, by multiple orders of magnitude as well. In this paper, we review the above issues, and give a preliminary overview of our group's activities towards the demonstration of groundbreaking levels of energy efficiency for semiconductor-based logic, together with a broader exploration of the ultimate limits of physically realizable techniques for approaching the theoretical ideal of perfect thermodynamic reversibility in computing, and the study of the implications of this technology direction for practical computing architectures. Michael P. Frank, Robert W. Brocato, Thomas M. Conte, Alexander H. Hsia, Anirudh Jain, Nancy A. Missert, Karpur Shukla, Brian D. Tierney |
ICCD | 3 |
| 2020 | MetaStrider: Architectures for Scalable Memory-centric Reduction of Sparse Data StreamsabstractReduction is an operation performed on the values of two or more key-value pairs that share the same key. Reduction of sparse data streams finds application in a wide variety of domains such as data and graph analytics, cybersecurity, machine learning, and HPC applications. However, these applications exhibit low locality of reference, rendering traditional architectures and data representations inefficient. This article presents MetaStrider, a significant algorithmic and architectural enhancement to the state-of-the-art, SuperStrider. Furthermore, these enhancements enable a variety of parallel, memory-centric architectures that we propose, resulting in demonstrated performance that scales near-linearly with available memory-level parallelism. Sriseshan Srikanth, Anirudh Jain, Joseph M. Lennon, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook |
ACM Trans. Archit. Code Optim. | 4 |
| 2019 | A microbenchmark characterization of the Emu chick
Jeffrey Young 0001, Eric R. Hein, Srinivas Eswar, Patrick Lavin, Jiajia Li 0001, E. Jason Riedy, Richard W. Vuduc, Thomas M. Conte |
Parallel Comput. | 8 |
| 2018 | Memory System Design for Ultra Low Power, Computationally Error Resilient Processor MicroarchitecturesabstractDennard scaling ended a decade ago. Energy reduction by lowering supply voltage has been limited because of guard bands and a subthreshold slope of over 60mV/decade in MOSFETs. On the other hand, newly-proposed logic devices maintain a high on/off ratio for drain currents even at significantly lower operating voltages. However, such ultra low power technology would eventually suffer from intermittent errors in logic as a result of operating close to the thermal noise floor. Computational error correction mitigates this issue by efficiently correcting stochastic bit errors that may occur in computational logic operating at low signal energies, thereby allowing for energy reduction by lowering supply voltage to tens of millivolts. Cores based on a Redundant Residual Number System (RRNS), which represents a number using a tuple of smaller numbers, are a promising candidate for implementing energyefficient computational error correction. However, prior RRNS core microarchitectures abstract away the memory hierarchy and do not consider the power-performance impact of RNS-based memory addressing. When compared with a non-error-correcting core addressing memory in binary, naive RNS-based memory addressing schemes cause a slowdown of over 3x/2x for inorder/out-of-order cores respectively. In this paper, we analyze RNS-based memory access pattern behavior and provide solutions in the form of novel schemes and the resulting design space exploration, thereby, extending and enabling a tangible, ultra low power RRNS based architecture. Sriseshan Srikanth, Paul G. Rabbat, Eric R. Hein, Bobin Deng, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook, Michael P. Frank |
HPCA | 5 |
| 2018 | Extending Moore's Law via Computationally Error-Tolerant ComputingabstractDennard scaling has ended. Lowering the voltage supply ( V dd ) to sub-volt levels causes intermittent losses in signal integrity, rendering further scaling (down) no longer acceptable as a means to lower the power required by a processor core. However, it is possible to correct the occasional errors caused due to lower V dd in an efficient manner and effectively lower power. By deploying the right amount and kind of redundancy, we can strike a balance between overhead incurred in achieving reliability and energy savings realized by permitting lower V dd . One promising approach is the Redundant Residue Number System (RRNS) representation. Unlike other error correcting codes, RRNS has the important property of being closed under addition, subtraction and multiplication, thus enabling computational error correction at a fraction of an overhead compared to conventional approaches. We use the RRNS scheme to design a Computationally-Redundant, Energy-Efficient core, including the microarchitecture, Instruction Set Architecture (ISA) and RRNS centered algorithms. From the simulation results, this RRNS system can reduce the energy-delay-product by about 3× for multiplication intensive workloads and by about 2× in general, when compared to a non-error-correcting binary core. Bobin Deng, Sriseshan Srikanth, Eric R. Hein, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook, Michael P. Frank |
ACM Trans. Archit. Code Optim. | 4 |
| 2015 | Rebooting Computing and Low-Power Image Recognition Challengeabstract“Rebooting Computing” (RC) is an effort in the IEEE to rethink future computers. RC started in 2012 by the co-chairs, Elie Track (IEEE Council on Superconductivity) and Tom Conte (Computer Society). RC takes a holistic approach, considering revolutionary as well as evolutionary solutions needed to advance computer technologies. Three summits have been held in 2013 and 2014, discussing different technologies, from emerging devices to user interface, from security to energy efficiency, from neuromorphic to reversible computing. The first part of this paper introduces RC to the design automation community and solicits revolutionary ideas from the community for the directions of future computer research. Energy efficiency is identified as one of the most important challenges in future computer technologies. The importance of energy efficiency spans from miniature embedded sensors to wearable computers, from individual desktops to data centers. To gauge the state of the art, the RC Committee organized the first Low Power Image Recognition Challenge (LPIRC). Each image contains one or multiple objects, among 200 categories. A contestant has to provide a working system that can recognize the objects and report the bounding boxes of the objects. The second part of this paper explains LPIRC and the solutions from the top two winners. Yung-Hsiang Lu, Alan M. Kadin, Alexander C. Berg, Thomas M. Conte, Erik DeBenedictis, Ganesh Gingade, Bichlien Hoang, Yongzhen Huang, Boxun Li, Jingyu Liu 0004, Wei Liu 0015, Huizi Mao, Junran Peng, Tianqi Tang 0001, Elie K. Track, Jingqiu Wang, Tao Wang 0004, Yu Wang 0002 |
ICCAD | 4 |
| 2015 | Contech: Efficiently Generating Dynamic Task Graphs for Arbitrary Parallel ProgramsabstractParallel programs can be characterized by task graphs encoding instructions, memory accesses, and the parallel work’s dependencies, while representing any threading library and architecture. This article presents Contech, a high performance framework for generating dynamic task graphs from arbitrary parallel programs, and a novel representation enabling programmers and compiler optimizations to understand and exploit program aspects. The Contech framework supports a variety of languages (including C, C++, and Fortran), parallelization libraries, and ISAs (including × 86 and ARM). Running natively for collection speed and minimizing program perturbation, the instrumentation shows 4 × improvement over a Pin-based implementation on PARSEC and NAS benchmarks. Brian P. Railing, Eric R. Hein, Thomas M. Conte |
ACM Trans. Archit. Code Optim. | 3 |
| 2014 | Manifold: A parallel simulation framework for multicore systemsabstractThis paper presents Manifold, an open-source parallel simulation framework for multicore architectures. It consists of a parallel simulation kernel, a set of microarchitecture components, and an integrated library of power, thermal, reliability, and energy models. Using the components as building blocks, users can assemble multicore architecture simulation models and perform serial or parallel simulations to study the architectural and/or the physical characteristics of the models. Users can also create new components for Manifold or port existing models. Importantly, Manifold's component-based design provides the user with the ability to easily replace a component with another for efficient explorations of the design space. It also allows components to evolve independently and making it easy for simulators to incorporate new components as they become available. The distinguishing features of Manifold include i) transparent parallel execution, ii) integration of power, thermal, reliability, and energy models, iii) full system simulation, e.g., operating system and system binaries, and iv) component-based design. In this paper we provide a description of the software architecture of Manifold, and its main elements - a parallel multicore emulator front-end and a parallel component-based back-end timing model. We describe a few simulators that are built with Manifold components to illustrate its flexibility, and present test results of the scalability obtained on full-system simulation of coherent shared-memory multicore models with 16, 32, and 64 cores executing PARSEC and SPLASH-2 benchmarks. Jun Wang 0077, Jesse G. Beu, Rishiraj A. Bheda, Thomas M. Conte, Zhenjiang Dong, Chad D. Kersey, Mitchelle Rasquinha, George F. Riley, William J. Song, Sudhakar Yalamanchili |
ISPASS | 4 |
| 2013 | High-speed formal verification of heterogeneous coherence hierarchiesabstractAs more heterogeneous architecture solutions continue to emerge, coherence solutions tailored for these architectures will become mandatory. Coherence hierarchies will likely continue to be prevalent in future large-scale shared memory architectures. However, past experience has shown that hierarchical coherence protocol design is a non-trivial problem, especially when considering the verification effort required to guarantee correctness. While some strategies do exist for verification of homogenous coherence hierarchies, support for reasonable verification of heterogeneous coherence hierarchies is currently unavailable. Ideally, hierarchical coherence protocols composed of `building block' protocols should be able to take advantage of incremental verification to side step the state-space explosion problem which hampers any large-scale verification effort. In this work, we prove this can be accomplished through the use of the Manager-Client Pairing (MCP) framework, which provides encapsulation and permission checking support that enables a form of state-space symmetry. When combined with an inductive proof, this ensures the validation properties of proper permission distribution and livelock/deadlock freedom are enforced by any hierarchical composition of MCP compliant protocols. Demonstration of this methodology through the MurPhi formal verifier shows several orders of magnitude improvement in verification cost compared to full hierarchy verification. Jesse G. Beu, Jason A. Poovey, Eric R. Hein, Thomas M. Conte |
HPCA | 4 |
| 2012 | Extrapolation Pitfalls When Evaluating Limited Endurance MemoryabstractMany new non-volatile memory technologies have been considered as a future scalable alternative to DRAM. Memory technologies such as MRAM, FeRAM, PCM have emerged as the most viable alternatives. But these memories have limited wear endurance. Practically realizable main memory systems employing these memory technologies are possible only if the wear across these memories is reduced as well as uniformly distributed. Limited endurance has resulted in extensive wear leveling research with the goal of uniformly distributing write traffic throughout available physical memory. Basic support for wear leveling is already present in existing systems, in the form of operating system paging. The Operating System (OS) changes virtual to physical translations over time. As a result, write traffic is naturally spread out. Proper evaluation of the need for wear leveling as well as the impact of the corresponding technique must take this phenomenon into account. Ignoring the effect of OS paging mechanism can result in highly inaccurate memory lifetime extrapolations. We demonstrate through simulation results, the effects of inaccurate extrapolations in the absence of OS modeling. Accurate memory lifetime simulation can take from many months to years. Although sampling techniques are commonly employed for speedup, our results show that naïve extrapolation techniques can lead to wildly different lifetime estimates. We show how sampling can be accurately applied by accounting for the different components in the write stream observed by main memory. Finally, we present a heuristic to quickly estimate memory lifetime for a given application. Rishiraj A. Bheda, Jesse G. Beu, Brian P. Railing, Thomas M. Conte |
MASCOTS | 4 |
| 2012 | Accelerating Multi-threaded Application Simulation through Barrier-Interval Time-ParallelismabstractIn the last decade, the microprocessor industry has undergone a dramatic change, ushering in the new era of multi-/manycore processors. As new designs incorporate increasing core counts, simulation technology has not matched pace, resulting in simulation times that increasingly dominate the design cycle. Complexities associated with the execution of code and communication between simulated cores has presented new obstacles for the simulation of manycore designs. Hence, many techniques developed to accelerate uniprocessor simulation cannot be easily adapted to accelerate manycore simulation. In this work, a novel time-parallel barrier-interval simulation methodology is presented to rapidly accelerate the simulation of certain classes of multi-threaded workloads. A program delineated into intervals by barriers may be accurately simulated in parallel. This approach avoids challenges originating from unknown thread progressions, since the program location of each executing thread is known. For the workloads tested, wall-clock speedups range from 1.22× to 596×, with an average of 13.94×. Furthermore, this approach allows the estimation of stable performance metrics such as cycle counts with minimal losses in accuracy (2%, on average, for all tested workloads). The proposed technique provides a fast and accurate mechanism to rapidly accelerate particular classes of manycore simulations. Paul D. Bryan, Jason A. Poovey, Jesse G. Beu, Thomas M. Conte |
MASCOTS | 4 |
| 2011 | Manager-client pairing: a framework for implementing coherence hierarchiesabstractAs technology continues to scale, the need for more sophisticated coherence management is becoming a necessity. The likely solution to this problem is the use of coherence hierarchies, analogous to how cache hierarchies have helped address the memory-wall problem in the past. Previous work in the construction of large-scale coherence protocols, however, demonstrates the complexity inherent to this design space. Jesse G. Beu, Michel C. Rosier, Thomas M. Conte |
MICRO | 3 |
| 2009 | On power and energy trends of IEEE 802.11n PHYabstractThe main contribution of this work is to decipher the power and energy characteristics of IEEE 802.11 PHY. In this work, we implement an IEEE 802.11n receiver and transmitter benchmark and measure power using an HDL-based scalable clustered-processor called CLAW. Second, we show that such a scalable processor can help conserve power and energy without sacrificing performance. We were able to extract 28% to 43% energy reduction with minimal overhead. Balaji V. Iyer, Thomas M. Conte |
MSWiM | 2 |
| 2008 | Energy-aware opcode designabstractEmbedded processors are required to achieve high performance while running on batteries. Thus, they must exploit all the possible means available to reduce energy consumption while not sacrificing performance. In this work, one technique to reduce energy is explored to intelligently design the instruction-opcodes of a processor based on a target-workload. The optimization is done using a heuristic that not-only minimizes switching between adjacent instructions, but also simplifies the decoding to reduce latches to save dynamic energy. On average, an optimized opcode is able to be decoded using 40-60% less latches in the decoder. In addition, it is shown that a decoder optimized for algorithms that had similar program structure, similar data-types or similar behavior exhibited consistent patterns of energy reduction. The techniques presented in this paper yield an average 10% reduction in the total dynamic energy. It is also shown that this heuristic can be used to achieve similar results on different issue-width processors. Balaji V. Iyer, Jason A. Poovey, Thomas M. Conte |
ICCD | 3 |
| 2007 | Keynote: Insight, Not (Random) Numbers: An Embedded Perspective
Thomas M. Conte |
HiPEAC | 1 |
| 2007 | Combining cluster sampling with single pass methods for efficient sampling regimen designabstractMicroarchitectural simulation is orders of magnitude slower than native execution. As more elements are accurately modeled, problems associated with slow simulation are further exacerbated. Given these issues, many researchers have devised sampling techniques to reduce simulation time. When cluster sampling techniques are used, care must be taken to remove sampling and non-sampling biases. Researchers have devised clever methods for effectively reducing non-sampling bias, but little work has been proposed for efficient reduction of sampling bias (sampling regimen design). Traditionally, sampling regimen design has been an iterative process that required a full workload simulation for error comparison. In this study, a single-pass simulation technique for sampling regimen design is proposed. Using this method, thousands of sampling regimen candidates can be simultaneously evaluated. With this technique, simulation speed was increased by an average factor of 17 with a maximum increase of 73 times relative to the total workload simulation. Additionally, this technique allows the user to effectively estimate the sample error without running the entire workload. Paul D. Bryan, Thomas M. Conte |
ICCD | 2 |
| 2007 | Reverse State Reconstruction for Sampled Microarchitectural SimulationabstractFor simulation, a tradeoff exists between speed and accuracy. The more instructions simulated from the workload, the more accurate the results - but at a higher cost. To reduce processor simulation times, a variety of techniques have been introduced. Statistically sampled simulation is one method that mitigates the cost of simulation while retaining high accuracy. A contiguous group of instructions, called a cluster, is simulated and then a fast type of simulation is used to skip to the next group. As instructions are skipped, non-sampling bias is introduced and must be removed for accurate measurements to be taken. In this paper, the reverse state reconstruction warm-up method is introduced. While skipping between clusters, the data necessary for reconstruction are recorded. Later, these data are scanned in reverse order so that processor state can be approximated without functionally applying every skipped instruction. By trading storage for speed, the proposed method introduces the concept of on-demand state reconstruction for sampled simulations. Using this technique, the method isolates ineffectual instructions from the skipped instructions without the use of profiling. Compared to SMARTS, reverse state reconstruction achieves a maximum and average speedup ratio of 2.45 and 1.64, respectively, with minimal sacrifice to accuracy (less than 0.3%) Paul D. Bryan, Michel C. Rosier, Thomas M. Conte |
ISPASS | 3 |
| 2005 | Insight, not (random) numbersabstractSummary form is only given. Hamming said, "The purpose of computing is insight, not numbers," yet this conference, like many today, is awash only in numbers. These numbers are perhaps more strategic than insightful. The numbers are used by designers, who want to prove their invention is better than the status quo. Then there are marketers, who want to prove their product's the one to buy over the competition. And then there are the users, who quite frankly are not getting much insight out of any of this. This talk will step back and discuss two aspects of insightful computing: who speaks for the users, and how much we should trust our numbers Thomas M. Conte |
ISPASS | 1 |
| 2005 | Spectral prefetcher: An effective mechanism for L2 cache prefetchingabstractEffective data prefetching requires accurate mechanisms to predict embedded patterns in the miss reference behavior. This paper proposes a novel prefetching mechanism, called the spectral prefetcher (SP), that accurately identifies the pattern by dynamically adjusting to its frequency. The proposed mechanism divides the memory address space into tag concentration zones (TCzones) and detects either the pattern of tags (higher order bits) or the pattern of strides (differences between consecutive tags) within each TCzone. The prefetcher dynamically determines whether the pattern of tags or strides will increase the effectiveness of prefetching and switches accordingly. To measure the performance of our scheme, we use a cycle-accurate aggressive out-of-order simulator that models bus occupancy, bus protocol, and limited bandwidth. Our experimental results show performance improvement of 1.59, on average, and at best 2.10 for the memory-intensive benchmarks we studied. Further, we show that SP outperforms the previously proposed scheme, with twice the size of SP, by 39% and a larger L2 cache, with equivalent storage area by 31%. Jesse G. Beu, Thomas M. Conte |
ACM Trans. Archit. Code Optim. | 3 |
| 2005 | Enhancing Memory-Level Parallelism via Recovery-Free Value PredictionabstractThe ever-increasing computational power of contemporary microprocessors reduces the execution time spent on arithmetic computations (i.e., the computations not involving slow memory operations such as cache misses) significantly. Therefore, for memory-intensive workloads, it becomes more important to overlap multiple cache misses than to overlap slow memory operations with other computations. In this paper, we propose a novel technique to parallelize sequential cache misses, thereby increasing memory-level parallelism (MLP). Our idea is based on value prediction, which was proposed originally as an instruction-level parallelism (ILP) optimization to break true data dependencies. In this paper, we advocate value prediction in its capability to enhance MLP instead of ILP. We propose using value prediction and value-speculative execution only for prefetching so that not only the complex prediction validation and misprediction recovery mechanisms are avoided, but better performance can also be achieved for memory-intensive workloads. The minor hardware modifications that are required also enable aggressive memory disambiguation for prefetching. The experimental results show that our technique enhances MLP effectively and achieves significant speedups, even with a simple stride value predictor. Huiyang Zhou, Thomas M. Conte |
IEEE Trans. Computers | 2 |
| 2005 | High-Performance and Low-Cost Dual-Thread VLIW Processor Using Weld Architecture ParadigmabstractThis paper presents a cost-effective and high-performance dual-thread VLIW processor model. The dual-thread VLIW processor model is a low-cost subset of the Weld architecture paradigm. It supports one main thread and one speculative thread running simultaneously in a VLIW processor with a register file and a fetch unit per thread along with memory disambiguation hardware for speculative load and store operations. This paper analyzes the performance impact of the dual-thread VLIW processor, which includes analysis of migrating disambiguation hardware for speculative load operations to the compiler and of the sensitivity of the model to the variation of branch misprediction, second-level cache miss penalties, and register file copy time. Up to 34 percent improvement in performance can be attained using the dual-thread VLIW processor when compared to a single-threaded VLIW processor model. Emre Ozer 0001, Thomas M. Conte |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2003 | Enhancing memory level parallelism via recovery-free value predictionabstractThe ever-increasing computational power of contemporary microprocessors reduces the execution time spent on arithmetic computations (i.e., the computations not involving slow memory operations such as cache misses) significantly. Therefore, for memory intensive workloads, it becomes more important to overlap multiple cache misses than to overlap slow memory operations with other computations. In this paper, we propose a novel technique to parallelize sequential cache misses, thereby increasing memory-level parallelism (MLP). Our idea is based on the value prediction, which was proposed originally as an instruction-level-parallelism (ILP) optimization to break true data dependencies. In this paper, we advocate value prediction in its capability to enhance MLP instead of ILP. We propose to use value prediction and value speculative execution only for prefetching so that the complex prediction validation and misprediction recovery mechanisms are avoided and only minor changes in the microarchitecture are needed. The same hardware modifications also enable aggressive memory disambiguation for prefetching. The experimental results show that our technique enhances MLP effectively and achieves significant speedups even with a simple stride value predictor. Huiyang Zhou, Thomas M. Conte |
ICS | 2 |
| 2003 | Detecting Global Stride Locality in Value Streams
Huiyang Zhou, Jill Flanagan, Thomas M. Conte |
ISCA | 3 |
| 2003 | Modeling Value Speculation: An Optimal Edge Selection ProblemabstractTechniques for value speculation have been proposed for dynamically scheduled and statically scheduled machines to increase instruction-level parallelism (ILP) by breaking flow (true) dependences and allowing value-dependent operations to be executed speculatively. The effectiveness of value speculation depends upon the ability to select and break dependences to shorten overall execution time, while encountering penalties for value misprediction. To understand and improve the techniques for value speculation, we model value speculation as an optimal edge selection problem. The optimal edge selection problem involves finding a minimal set of edges (dependences) to break in a data dependence graph that achieves maximal benefits from value speculation, while taking the penalties for value misprediction into account. Based on three properties observed from the optimal edge selection problem, an efficient optimal edge selection algorithm is designed. From the experimental results of running the optimal edge selection algorithm for the 20 most heavily executed paths selected from each SPECint95 benchmark, several insights are shown. The average critical path reduction is 9.61 percent on an average and 25.57 percent at its maximum. Surprisingly, 66 percent of the edges selected by the optimal algorithm have value prediction accuracies over 99 percent. Moreover, most of the selected edges cross the middle of the data dependence graph. The selected producer operations thereby tend to reside in the upper portion of the data dependence graph, while the selected consumer operations appear toward the lower portion. Chao-ying Fu, Jill T. Bodine, Thomas M. Conte |
IEEE Trans. Computers | 3 |
| 2003 | Adaptive mode control: A static-power-efficient cache designabstractLower threshold voltages in deep submicron technologies cause more leakage current, increasing static power dissipation. This trend, combined with the trend of larger/more cache memories dominating die area, has prompted circuit designers to develop SRAM cells with low-leakage operating modes (e.g., sleep mode). Sleep mode reduces static power dissipation, but data stored in a sleeping cell is unreliable or lost. So, at the architecture level, there is interest in exploiting sleep mode to reduce static power dissipation while maintaining high performance.Current approaches dynamically control the operating mode of large groups of cache lines or even individual cache lines. However, the performance monitoring mechanism that controls the percentage of sleep-mode lines, and identifies particular lines for sleep mode, is somewhat arbitrary. There is no way to know what the performance could be with all cache lines active, so arbitrary miss rate targets are set (perhaps on a per-benchmark basis using profile information), and the control mechanism tracks these targets. We propose applying sleep mode only to the data store and not the tag store. By keeping the entire tag store active the hardware knows what the hypothetical miss rate would be if all data lines were active, and the actual miss rate can be made to precisely track it. Simulations show that an average of 73% of I-cache lines and 54% of D-cache lines are put in sleep mode with an average IPC impact of only 1.7%, for 64 KB caches. Huiyang Zhou, Mark C. Toburen, Eric Rotenberg, Thomas M. Conte |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2001 | Weld: A Multithreading Technique Towards Latency-Tolerant VLIW Processors
Emre Ozer 0001, Thomas M. Conte |
HiPC | 2 |
| 2000 | Properties of Rescheduling Size Invariance for Dynamic Rescheduling-Based VLIW Cross-Generation CompatibilityabstractThe object-code compatibility problem in VLIW architectures stems from their statically scheduled nature. Dynamic rescheduling (DR) is a technique to solve the compatibility problem in VLIWs. DR reschedules program code pages at first-time page faults, i.e., when the code pages are accessed for the first time during execution. Treating a page of code as the unit of rescheduling makes it susceptible to the hazards of changes in the page size during the process of rescheduling. This paper shows that the changes in the page size are only due to insertion and/or deletion of NOPs in the code. Further, it presents an ISA encoding, called list encoding, which does not require explicit encoding of the NOPs in the code. Algorithms to perform rescheduling on acyclic code and cyclic code are presented, followed by the discussion of the property of rescheduling-size invariance (RSI) satisfied by list encoding. Thomas M. Conte, Sumedh W. Sathaye |
IEEE Trans. Computers | 1 |
| 2000 | System-level power consumption modeling and tradeoff analysis techniques for superscalar processor designabstractThis paper presents systematic techniques to find low-power high-performance superscalar processors tailored to specific user applications. The model of power is novel because it separates power into architectural and technology components. The architectural component is found via trace-driven simulation, which also produces performance estimates. An example technology model is presented that estimates the technology component, along with critical delay time and real estate usage. This model is based on case studies of actual designs. It is used to solve an important problem: decreasing power consumption in a superscalar processor without greatly impacting performance. Results are presented from runs using simulated annealing to reduce power consumption subject to performance reduction bounds. The major contributions of this paper are the separation of architectural and technology components of dynamic power the use of trace-driven simulation for architectural power measurement, and the use of a near-optimal search to tailor a processor design to a benchmark. Thomas M. Conte, Kishore N. Menezes, Sumedh W. Sathaye, Mark C. Toburen |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1999 | Dynamically Programmable Cache Evaluation and VirtualizationabstractNo abstract available. Mouna Nakkar, David G. Bentlage, John Harding, David Schwartz, Paul D. Franzon, Thomas M. Conte |
FPGA | 6 |
| 1999 | Compiler-Driven Cached Code Compression Schemes for Embedded ILP ProcessorsabstractDuring the last 15 years, embedded systems have grown in complexity and performance to rival desktop systems. The architectures of these systems present unique challenges to processor microarchitecture, including instruction encoding and instruction fetch processes. This paper presents new techniques for reducing embedded system code size without reducing functionality. This approach is to extract the pipeline decoder logic for an embedded VLIW processor in software at system development time. The code size reduction is achieved by Huffman compressing or tailor encoding the ISA of the original program. Some interesting results were found. In particular, the degree of compression for the ROM doesn't translate into an improvement in instructions delivered per cycle. Experiments found that when the missprediction penalty of the added Huffman decoder stage was taken into account, a Tailored ISA approach produced higher performance. Methods that compress the entire operation using Huffman encodings, and decompress at ICache hit time still achieved a median performance advantage, while providing higher ROM size savings. All results were generated by an optimizing compiler and tool suite, and presented for an encoding similar to the Intel/HP IA-64 architecture. Sergei Y. Larin, Thomas M. Conte |
MICRO | 2 |
| 1998 | Value Speculation Scheduling for High Performance ProcessorsabstractRecent research in value prediction shows a surprising amount of predictability for the values produced by register-writing instructions. Several hardware based value predictor designs have been proposed to exploit this predictability by eliminating flow dependencies for highly predictable values. This paper proposed a hardware and software based scheme for value speculation scheduling (VSS). Static VLIW scheduling techniques are used to speculate value dependent instructions by scheduling them above the instructions whose results they are dependent on. Prediction hardware is used to provide value predictions for allowing the execution of speculated instructions to continue. In the case of miss-predicted values, control flow is redirected to patch-up code so that execution can proceed with the correct results. In this paper, experiments in VSS for load operations in the SPECint95 benchmarks are performed. Speedup of up to 17% has been shown for using VSS. Empirical results on the value predictability of loads, based on value profiling data, are also provided. Chao-ying Fu, Matthew D. Jennings, Sergei Y. Larin, Thomas M. Conte |
ASPLOS | 4 |
| 1998 | Treegion Scheduling for Wide Issue ProcessorsabstractInstruction scheduling is one of the most important phases of compilation for high-performance processors. A compiler typically divides a program into multiple regions of code and then schedules each region. Many past efforts have focused on linear regions such as traces and superblocks. The linearity of these regions can limit speculation, leading to under-utilization of processor resources, especially on wide-issue machines. A type of non-linear region called a treegion is presented in this paper. The formation and scheduling of treegions takes into account multiple execution paths, and the larger scope of treegions allows more speculation, leading to higher utilization and better performance. Multiple scheduling heuristics for treegions are compared against scheduling for several types of linear regions. Empirical results illustrate that instruction scheduling using treegions treegion scheduling-holds promise. Treegion scheduling using the global weight heuristic outperforms the next highest performing region-superblocks by up to 20%. William A. Havanki, Sanjeev Banerjia, Thomas M. Conte |
HPCA | 3 |
| 1998 | Unified Assign and Schedule: A New Approach to Scheduling for Clustered Register File MicroarchitecturesabstractRecently, there has been a trend towards clustered microarchitectures to reduce the cycle time for wide issue microprocessors. In such processors, the register file and functional units are partitioned and grouped into clusters. Instruction scheduling for a clustered machine requires assignment and scheduling of operations to the clusters. In this paper, a new scheduling algorithm named unified-assign-and-schedule (UAS) is proposed for clustered, statically-scheduled architectures. UAS merges the cluster assignment and instruction scheduling phases in a natural and straightforward fashion. We compared the performance of UAS with various heuristics to the well-known Bottom-up Greedy (BUG) algorithm and to an optimal cluster scheduling algorithm, measuring the schedule lengths produced by all of the schedulers. Our results show that UAS gives better performance than the BUG algorithm and is quite close to optimal. Emre Ozer 0001, Sanjeev Banerjia, Thomas M. Conte |
MICRO | 3 |
| 1998 | MPS: Miss-Path Scheduling for Multiple-Issue ProcessorsabstractMany contemporary multiple issue processors employ out-of-order scheduling hardware in the processor pipeline. Such scheduling hardware can yield good performance without relying on compile-time scheduling. The hardware can also schedule around unexpected run-time occurrences such as cache misses. As issue widths increase, however, the complexity of such scheduling hardware increases considerably and can have an impact on the cycle time of the processor. This paper presents the design of a multiple issue processor that uses an alternative approach called miss path scheduling or MPS. Scheduling hardware is removed from the processor pipeline altogether and placed on the path between the instruction cache and the next level of memory. Scheduling is performed at cache miss time as instructions are received from memory. Scheduled blocks of instructions are issued to an aggressively clocked in-order execution core. Details of a hardware scheduler that can perform speculation are outlined and shown to be feasible. Performance results from simulations are presented that highlight the effectiveness of an MPS design. Sanjeev Banerjia, Sumedh W. Sathaye, Kishore N. Menezes, Thomas M. Conte |
IEEE Trans. Computers | 4 |
| 1998 | Combining Trace Sampling with Single Pass Methods for Efficient Cache SimulationabstractThe design of the memory hierarchy is crucial to the performance of high performance computer systems. The incorporation of multiple levels of caches into the memory hierarchy is known to increase the performance of high end machines, but the development of architectural prototypes of various memory hierarchy designs is costly and time consuming. In this paper, we will describe a single pass method used in combination with trace sampling techniques to produce a fast and accurate approach for simulating multiple sizes of caches simultaneously. Thomas M. Conte, Mary Ann Hirsch, Wen-Mei W. Hwu |
IEEE Trans. Computers | 1 |
| 1997 | Treegion Scheduling for Highly Parallel Processors
Sanjeev Banerjia, William A. Havanki, Thomas M. Conte |
Euro-Par | 3 |
| 1996 | Reducing State Loss For Effective Trace Sampling of Superscalar ProcessorsabstractThere is a wealth of technological alternatives that can be incorporated into a processor design. These include reservation station designs, functional unit duplication, and processor branch handling strategies. The performance of a given design is measured through the execution of application programs and other workloads. Presently, trace driven simulation is the most popular method of processor performance analysis in the development stage of system design. Current techniques of trace driven simulation, however, are extremely slow and expensive. A fast and accurate method for statistical trace sampling of superscalar processors is proposed. Thomas M. Conte, Mary Ann Hirsch, Kishore N. Menezes |
ICCD | 1 |
| 1996 | Instruction Fetch Mechanisms for VLIW Architectures with Compressed EncodingsabstractVLIW architectures use very wide instruction words in conjunction with high bandwidth to the instruction cache to achieve multiple instruction issue. This report uses the TINKER experimental testbed to examine instruction fetch and instruction cache mechanisms for VLIWs. A compressed instruction encoding for VLIWs is defined and a classification scheme for i-fetch hardware for such an encoding is introduced. Several interesting cache and i-fetch organizations are described and evaluated through trace-driven simulations. A new i-fetch mechanism using a silo cache is found to have the best performance. Thomas M. Conte, Sanjeev Banerjia, Sergei Y. Larin, Kishore N. Menezes, Sumedh W. Sathaye |
MICRO | 1 |
| 1996 | Accurate and Practical Profile-driven Compilation Using the Profile BufferabstractProfiling is a technique of gathering program statistics in order to aid program optimization. In particular, it is an essential component of compiler optimization for the extraction of instruction-level parallelism. Code instrumentation has been the most popular method of profiling. However real-time, interactive, and transaction processing applications suffer from the high execution-time overhead imposed by software instrumentation. This paper suggests the use of hardware dedicated to the task of profiling. The hardware proposed consists of a set of counters, the profile buffer. A profile collection method that combines the use of hardware, the compiler and operating system support is described. Three methods for profile buffer indexing, address-mapping, selective indexing, and compiler indexing are presented that allow this approach to produce accurate profiling information with very little execution slowdown. The profile information obtained is applied to a prominent compiler optimization, namely superblock scheduling. The resulting instruction-level parallelism approaches that obtained through the use of perfect profile information. Thomas M. Conte, Kishore N. Menezes, Mary Ann Hirsch |
MICRO | 1 |
| 1996 | A Persistent Rescheduled-page Cache for Low Overhead Object Code Compatibility in VLIW ArchitecturesabstractObject-code compatibility between processor generations is an open issue for VLIW architectures. A potential solution is a technique termed dynamic rescheduling, which performs run-time software rescheduling at the first-time page faults. The time required for rescheduling the pages constitutes a large portion of the overhead of this method. A disk caching scheme that uses a persistent rescheduled-page cache (PRC) is presented. The scheme reduces the overhead associated with dynamic rescheduling by saving rescheduled pages on disk, across program executions. Operating system support is required for dynamic rescheduling and management of the PRC. The implementation details for the PRC are discussed. Results of simulations used to gauge the effectiveness of PRC indicate that: the PRC is effective in reducing the overhead of dynamic rescheduling; and due to different overhead requirements of programs, a split PRC organization performs better than a unified PRC. The unified PRC was studied for two different page replacement policies: LRU and overhead-based replacement. It was found that with LRU replacement, all the programs consistently perform better with increasing PRC sizes, but the high-overhead programs take a consistent performance hit compared to the low-overhead programs. With overhead-based replacement, the performance of high-overhead programs improves substantially, while the low-overhead programs perform only slightly worse than in the case of the LRU replacement. Thomas M. Conte, Sumedh W. Sathaye, Sanjeev Banerjia |
MICRO | 1 |
| 1995 | Optimization of Instruction Fetch Mechanisms for High Issue RatesabstractRecent superscalar processors issue four instructions per cycle. These processors are also powered by highly-parallel superscalar cores. The potential performance can only be exploited when fed by high instruction bandwidth. This task is the responsibility of the instruction fetch unit. Accurate branch prediction and low I-cache miss ratios are essential for the efficient operation of the fetch unit. Several studies on cache design and branch prediction address this problem. However, these techniques are not sufficient. Even in the presence of efficient cache designs and branch prediction, the fetch unit must continuously extract multiple, non-sequential instructions from the instruction cache, realign these in the proper order, and supply them to the decoder. This paper explores solutions to this problem and presents several schemes with varying degrees of performance and cost. The most-general scheme, the collapsing buffer, achieves near-perfect performance and consistently aligns instructions in excess of 90% of the time, over a wide range of issue rates. The performance boost provided by compiler optimization techniques is also investigated. Results show that compiler optimization can significantly enhance performance across all schemes. The collapsing buffer supplemented by compiler techniques remains the best-performing mechanism. The paper closes with recommendations and suggestions for future. Thomas M. Conte, Kishore N. Menezes, Patrick M. Mills, Burzin A. Patel |
ISCA | 1 |
| 1995 | Dynamic rescheduling: a technique for object code compatibility in VLIW architecturesabstractLack of object code compatibility in VLIW architectures is a severe limit to their adoption as a general-purpose computing paradigm. Previous approaches include hardware and software techniques, both of which have drawbacks. Hardware techniques add to the complexity of the architecture, whereas software techniques require multiple executables. This paper presents a technique called dynamic rescheduling that applies software techniques dynamically, using intervention by the operating system. Results are presented to demonstrate the viability of the technique using the Illinois IMPACT compiler and the TINKER architectural framework. Thomas M. Conte, Sumedh W. Sathaye |
MICRO | 1 |
| 1994 | Using branch handling hardware to support profile-driven optimizationabstractProfile-based optimizations can be used for instruction scheduling, loop scheduling, data preloading, function in-lining, and instruction cache performance enhancement. However, these techniques have not been embraced by software vendors because programs instrumented for profiling run 2-30 times slower, an awkward compile-run-recompile sequence is required, and a test input suite must be collected and validated for each program. This paper proposes using existing branch handling hardware to generate profile information in real time. Techniques are presented for both one-level and two-level branch hardware organizations. The approach produces high accuracy with small slowdown in execution (0.4%-4.6%). This allows a program to be profiled while it is used, eliminating the need for a test input suite. This practically removes the inconvenience of profiling. With contemporary processors driven increasingly by compiler support, hardware-based profiling is important for high-performance systems. Thomas M. Conte, Burzin A. Patel, J. Stan Cox |
MICRO | 1 |
| 1994 | The Susceptibility of Programs to Context SwitchingabstractModern memory systems are composed of several levels of caching. The design of these levels is largely an empirical practice. One highly-effective empirical method is the single-pass method wherein all caches in a broad design space are evaluated in one pass over the trace. Multiprogramming degrades memory system performance since context switching reduces the effectiveness of cache memories. Few single-pass methods exist which account for multiprogramming effects. This paper uses a general model of single-pass algorithms, the recurrence/conflict model, and extends the model for recording the effects due to both voluntary context switches and involuntary context switches. Involuntary context switches are modeled using the distribution of lengths between a reference to an address and the re-reference to the same address. The paper makes the assumptions that involuntary context switches are equally likely to occur between each reference, and that one can independently estimate f/sub CS/, the fraction of a cache's contents flushed between context switches. The case where f/sub CS/=1 is used to measure the effect of worst-case context switch penalty (the susceptibility) of several members of the SPEC89 benchmark set to context switching. Some empirical results of F/sub CS/ are presented to illustrate the case where f/sub CS/> Wen-Mei W. Hwu, Thomas M. Conte |
IEEE Trans. Computers | 2 |
| 1993 | Determining Cost-Effective Multiple Issue Processor DesignsabstractSeveral commercial processors, including the Motorola 88110 and the DEC Alpha, are capable of issuing multiple operations per clock cycle. Optimization of the pipeline depth and number of function units in these processors has been largely ignored due to limited semiconductor resources. Recently, advances in feature size and packaging technologies have removed these limitations. It is possible that next-generation processor designs may benefit from multiple function unit copies and optimize pipeline depths. The paper investigates the feasibility of performing synthesis at the architectural specification level. The design space is optimized for performance constrained by a hardware model of silicon area. The results of this study indicate that cost-effective high performance can be achieved with the addition of small amounts of function unit duplication. These results are also used to comment on the validity of the "benchmark suite" approach to performance evaluation and machine design.> Thomas M. Conte, William H. Mangione-Smith |
ICCD | 1 |
| 1993 | The Effect of Code Expanding Optimizations on Instruction Cache DesignabstractShows that code expanding optimizations have strong and nonintuitive implications on instruction cache design. Three types of code expanding optimizations are studied in this paper: instruction placement, function inline expansion, and superscalar optimizations. Overall, instruction placement reduces the miss ratio of small caches. Function inline expansion improves the performance for small cache sizes, but degrades the performance of medium caches. Superscalar optimizations increase the miss ratio for all cache sizes. However, they also increase the sequentiality of instruction access so that a simple load forwarding scheme effectively cancels the negative effects. Overall, the authors show that with load forwarding, the three types of code expanding optimizations jointly improve the performance of small caches and have little effect on large caches.> William Y. Chen, Pohua P. Chang, Thomas M. Conte, Wen-Mei W. Hwu |
IEEE Trans. Computers | 3 |
| 1992 | Tradeoffs in processor/memory interfaces for superscalar processors
Thomas M. Conte |
MICRO | 1 |
| 1992 | Systematic prototyping of superscalar computer architecturesabstractIt is argued that the correct solution to the computer architecture design process is a prototyping approach. In this approach, the initial design is selected using the workload itself. This design is an architectural prototype, specifying the high-level design decisions that are difficult to acquire via detailed simulation. These design decisions include determining the mix of function units in the execution stage of the processor and the dimensions of the caches in the memory subsystem. Architectural prototypes are selected by trading off accuracy in hardware simulation for an increase in usable workload size.> Thomas M. Conte, Wen-Mei W. Hwu |
RSP | 1 |
| 1989 | Comparing Software and Hardware Schemes For Reducing the Cost of BranchesabstractPipelining has become a common technique to increase throughput of the instruction fetch, instruction decode, and instruction execution portions of modern comput-ers. Branch instructions disrupt the flow of instructions through the the pipeline, increasing the overall execution cost of branch instructions. Three schemes to reduce the cost of branches are presented in the context of a gen-eral pipeline model. Ten realistic Unix domain programs are used to directly compare the cost and performance of the three schemes and the results are in favor of the software-based scheme. For example, the software-based scheme has a cost of 1.65 cycles/branch vs. a cost of 1.68 cycles/branch of the best hardware scheme for a highly pipelined processor (11-stage pipeline). The results are 1.19 (software scheme) vs. 1.23 cycles/branch (best hard-ware scheme) for a moderately pipelined processor (5-stage pipeline). 1 Wen-Mei W. Hwu, Thomas M. Conte, Pohua P. Chang |
ISCA | 2 |
| 1989 | A Simulation Study of Simultaneous Vector Prefetch Performance in Multiprocessor Memory Subsystems (Extended Abstract)
Wen-Mei W. Hwu, Thomas M. Conte |
SIGMETRICS | 2 |