EDBT 2026 Demo / reviewers in the wild / expert
David H. Albonesi
dblp:44/1649
· DBLP profile ↗
46ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 42 · 4 first-authorSoftware engineering, systems software and programming languages · 10 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
24 papers |
Hardware accelerators and domain-specific architectures · 24% Processor architecture and microarchitecture · 21% Energy-efficient computing · 18% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 30 heaviest of 64, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing
power management |
0.5 | 5 | 2017 | Dynamic GPGPU Power Management Using Adaptive Model Predictive Control · HPCA 2017 Flicker: a dynamically adaptive architecture for power limited multicore systems · ISCA 2013 A Dynamically Tunable Memory Hierarchy · IEEE Trans. Computers 2003 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.4 | 5 | 2017 | Dynamic GPGPU Power Management Using Adaptive Model Predictive Control · HPCA 2017 Dynamically Trading Frequency for Complexity in a GALS Microprocessor · MICRO 2004 Profile-Based Dynamic Voltage and Frequency Scaling for a Multiple Clock Domain Microprocessor · ISCA 2003 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.4 | 1 | 2020 | CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores · MICRO 2020 |
Hardware accelerators and domain-specific architectures
sparse matrix multiplication accelerator |
0.4 | 1 | 2020 | MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise Product · MICRO 2020 |
Hardware accelerators and domain-specific architectures › sparse matrix multiplication accelerator
SpGEMM accelerator |
0.4 | 1 | 2020 | MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise Product · MICRO 2020 |
Hardware accelerators and domain-specific architectures
tensor accelerator |
0.4 | 1 | 2020 | Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor Computations · HPCA 2020 |
Performance modeling and evaluation
workload characterization |
0.4 | 1 | 2020 | CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores · MICRO 2020 |
Processor architecture and microarchitecture
multicore design |
0.3 | 2 | 2020 | Flicker: a dynamically adaptive architecture for power limited multicore systems · ISCA 2013 CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores · MICRO 2020 |
GPUs and heterogeneous computing
GPU power management |
0.3 | 1 | 2017 | Dynamic GPGPU Power Management Using Adaptive Model Predictive Control · HPCA 2017 |
Processor architecture and microarchitecture › chip multiprocessor
reconfigurable multicore |
0.2 | 2 | 2020 | CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores · MICRO 2020 ReMAP: A Reconfigurable Heterogeneous Multicore Architecture · MICRO 2010 |
Cloud and datacenter computing
dynamic adaptation |
0.2 | 1 | 2013 | Flicker: a dynamically adaptive architecture for power limited multicore systems · ISCA 2013 |
Memory systems
on-chip memory |
0.1 | 1 | 2020 | MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise Product · MICRO 2020 |
High-performance computing › sparse linear algebra
sparse matrix storage format |
0.1 | 1 | 2020 | MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise Product · MICRO 2020 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.1 | 2 | 2008 | Addressing thermal nonuniformity in SMT workloads · ACM Trans. Archit. Code Optim. 2008 Front-End Policies for Improved Issue Efficiency in SMT Processors · HPCA 2003 |
Processor architecture and microarchitecture › multicore design
heterogeneous multicore |
0.1 | 1 | 2010 | ReMAP: A Reconfigurable Heterogeneous Multicore Architecture · MICRO 2010 |
Processor architecture and microarchitecture › chip multiprocessor
inter-core communication |
0.1 | 1 | 2010 | ReMAP: A Reconfigurable Heterogeneous Multicore Architecture · MICRO 2010 |
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable fabric |
0.1 | 1 | 2010 | ReMAP: A Reconfigurable Heterogeneous Multicore Architecture · MICRO 2010 |
Interconnection networks and networks-on-chip
optical network-on-chip |
0.1 | 1 | 2009 | Phastlane: a rapid transit optical routing network · ISCA 2009 |
Electronic design automation › physical design › routing
optical routing |
0.1 | 1 | 2009 | Phastlane: a rapid transit optical routing network · ISCA 2009 |
Electronic design automation › physical design
routing |
0.1 | 1 | 2009 | Phastlane: a rapid transit optical routing network · ISCA 2009 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2009 | Leveraging Optical Technology in Future Bus-based Chip Multiprocessors · MICRO 2006 Phastlane: a rapid transit optical routing network · ISCA 2009 |
Performance modeling and evaluation
performance prediction |
0.1 | 1 | 2017 | Dynamic GPGPU Power Management Using Adaptive Model Predictive Control · HPCA 2017 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.1 | 3 | 2003 | Dynamic Data Dependence Tracking and its Application to Branch Prediction · HPCA 2003 Dynamically allocating processor resources between nearby and distant ILP · ISCA 2001 Dynamically Managing the Communication-Parallelism Trade-off in Future Clustered Processors · ISCA 2003 |
Energy-efficient computing › thermal management
dynamic thermal management |
0.1 | 1 | 2008 | Addressing thermal nonuniformity in SMT workloads · ACM Trans. Archit. Code Optim. 2008 |
Processor architecture and microarchitecture › out-of-order execution
issue queue |
0.1 | 2 | 2003 | Energy Efficient Co-Adaptive Instruction Fetch and Issue · ISCA 2003 Front-End Policies for Improved Issue Efficiency in SMT Processors · HPCA 2003 |
Energy-efficient computing
thermal management |
0.1 | 1 | 2008 | Addressing thermal nonuniformity in SMT workloads · ACM Trans. Archit. Code Optim. 2008 |
Parallel and multicore computing › parallel programming runtimes
thread management |
0.1 | 1 | 2008 | Addressing thermal nonuniformity in SMT workloads · ACM Trans. Archit. Code Optim. 2008 |
Energy-efficient computing › microprocessor power management
multiple clock domain processor |
0.1 | 2 | 2003 | Profile-Based Dynamic Voltage and Frequency Scaling for a Multiple Clock Domain Microprocessor · ISCA 2003 Energy-Efficient Processor Design Using Multiple Clock Domains with Dynamic Voltage and Frequency Scaling · HPCA 2002 |
Memory systems › memory hierarchy
cache hierarchy |
0.1 | 2 | 2003 | A Dynamically Tunable Memory Hierarchy · IEEE Trans. Computers 2003 Memory hierarchy reconfiguration for energy and performance in general-purpose processor architectures · MICRO 2000 |
Memory systems › memory management › virtual memory › address translation
TLB |
0.1 | 2 | 2003 | A Dynamically Tunable Memory Hierarchy · IEEE Trans. Computers 2003 Memory hierarchy reconfiguration for energy and performance in general-purpose processor architectures · MICRO 2000 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.5sparse storage format · 0.4hardware-software co-design · 0.4gem5 · 0.4dynamically dimensioned search · 0.4data mining · 0.4collaborative filtering · 0.4performance prediction · 0.3model predictive control · 0.3phase detection · 0.1profiling · 0.0binary rewriting · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor ComputationsabstractTensor factorizations are powerful tools in many machine learning and data analytics applications. Tensors are often sparse, which makes sparse tensor factorizations memory bound. In this work, we propose a hardware accelerator that can accelerate both dense and sparse tensor factorizations. We co-design the hardware and a sparse storage format, which allows accessing the sparse data in vectorized and streaming fashion and maximizes the utilization of the memory bandwidth. We extract a common computation pattern that is found in numerous matrix and tensor operations and implement it in the hardware. By designing the hardware based on this common compute pattern, we can not only accelerate tensor factorizations but also mixed sparse-dense matrix operations. We show significant speedup and energy benefit over the state-of-the-art CPU and GPU implementations of tensor factorizations and over CPU, GPU and accelerators for matrix operations. Nitish Kumar Srivastava, Hanchen Jin, Shaden Smith, Hongbo Rong, David H. Albonesi, Zhiru Zhang |
HPCA | 5 |
| 2020 | CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable MulticoresabstractMulti-tenancy for latency-critical applications leads to resource interference and unpredictable performance. Core reconfiguration opens up more opportunities for application colocation, as it allows the hardware to adjust to the dynamic performance and power needs of a specific mix of co-scheduled services. However, reconfigurability also introduces challenges, as even for a small number of reconfigurable cores, exploring the design space becomes more time- and resource-demanding.We present CuttleSys, a runtime for reconfigurable multicores that leverages scalable and lightweight data mining to quickly identify suitable core and cache configurations for a set of co-scheduled applications. The runtime combines collaborative filtering to infer the behavior of each job on every core and cache configuration, with Dynamically Dimensioned Search to efficiently explore the configuration space. We evaluate CuttleSys on multicores with tens of reconfigurable cores and show up to 2.46× and 1.55× performance improvements compared to core-level gating and oracle-like asymmetric multicores respectively, under stringent power constraints. Neeraj Kulkarni, Gonzalo Gonzalez-Pumariega, Amulya Khurana, Christine A. Shoemaker, Christina Delimitrou, David H. Albonesi |
MICRO | 6 |
| 2020 | MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise ProductabstractSparse-sparse matrix multiplication (SpGEMM) is a computation kernel widely used in numerous application domains such as data analytics, graph processing, and scientific computing. In this work we propose MatRaptor, a novel SpGEMM accelerator that is high performance and highly resource efficient. Unlike conventional methods using inner or outer product as the meta operation for matrix multiplication, our approach is based on row-wise product, which offers a better tradeoff in terms of data reuse and on-chip memory requirements, and achieves higher performance for large sparse matrices. We further propose a new hardware-friendly sparse storage format, which allows parallel compute engines to access the sparse data in a vectorized and streaming fashion, leading to high utilization of memory bandwidth. We prototype and simulate our accelerator architecture using gem5 on a diverse set of matrices. Our experiments show that MatRaptor achieves 129.2× speedup over single-threaded CPU, 8.8× speedup over GPU and 1.8× speedup over the state-of-the-art SpGEMM accelerator (OuterSPACE). MatRaptor also has 7.2× lower power consumption and 31.3× smaller area compared to OuterSPACE. Nitish Kumar Srivastava, Hanchen Jin, Jie Liu 0072, David H. Albonesi, Zhiru Zhang |
MICRO | 4 |
| 2019 | T2S-Tensor: Productively Generating High-Performance Spatial Hardware for Dense Tensor ComputationsabstractWe present a language and compilation framework for productively generating high-performance systolic arrays for dense tensor kernels on spatial architectures, including FPGAs and CGRAs. It decouples a functional specification from a spatial mapping, allowing programmers to quickly explore various spatial optimizations for the same function. The actual implementation of these optimizations is left to a compiler. Thus, productivity and performance are achieved at the same time. We used this framework to implement several important dense tensor kernels. We implemented dense matrix multiply for an Arria-10 FPGA and a research CGRA, achieving 88% and 92% of the performance of manually written, and highly optimized expert (ninja") implementations in just 3% of their engineering time. Three other tensor kernels, including MTTKRP, TTM and TTMc, were also implemented with high performance and low design effort, and for the first time on spatial architectures." Nitish Kumar Srivastava, Hongbo Rong, Prithayan Barua, Guanyu Feng, Huanqi Cao, Zhiru Zhang, David H. Albonesi, Vivek Sarkar, Paul Petersen, Geoff Lowney, Adam Herr, Christopher J. Hughes, Timothy G. Mattson, Pradeep Dubey |
FCCM | 7 |
| 2017 | Dynamic GPGPU Power Management Using Adaptive Model Predictive ControlabstractModern processors can greatly increase energy efficiency through techniques such as dynamic voltage and frequency scaling. Traditional predictive schemes are limited in their effectiveness by their inability to plan for the performance and energy characteristics of upcoming phases. To date, there has been little research exploring more proactive techniques that account for expected future behavior when making decisions. This paper proposes using Model Predictive Control (MPC) to attempt to maximize the energy efficiency of GPU kernels without compromising performance. We develop performance and power prediction models for a recent CPU-GPU heterogeneous processor. Our system then dynamically adjusts hardware states based on recent execution history, the pattern of upcoming kernels, and the predicted behavior of those kernels. We also dynamically trade off the performance overhead and the effectiveness of MPC in finding the best configuration by adapting the horizon length at runtime. Our MPC technique limits performance loss by proactively spending energy on the kernel iterations that will gain the most performance from that energy. This energy can then be recovered in future iterations that are less performance sensitive. Our scheme also avoids wasting energy on low-throughput phases when it foresees future high-throughput kernels that could better use that energy. Compared to state-of-the-practice schemes, our approach achieves 24.8% energy savings with a performance loss (including MPC overheads) of 1.8%. Compared to state-of-the-art history-based schemes, our approach achieves 6.6% chip-wide energy savings while simultaneously improving performance by 9.6%. Abhinandan Majumdar, Leonardo Piga, Indrani Paul, Joseph L. Greathouse, Wei Huang 0004, David H. Albonesi |
HPCA | 6 |
| 2017 | DeepRecon: Dynamically reconfigurable architecture for accelerating deep neural networksabstractDeep learning models are computationally expensive and their performance depends strongly on the underlying hardware platform. General purpose compute platforms such as GPUs have been widely used for implementing deep learning techniques. However, with the advent of emerging application domains such as internet of things, developments of custom integrated circuits capable of efficiently implementing deep learning models with low power and form factor are in high demand. In this paper we analyze both the computation and communication costs of common deep networks. We propose a reconfigurable architecture that efficiently utilizes computational and storage resources for accelerating deep learning techniques without loss of algorithmic accuracy. Tayyar Rzayev, Saber Moradi, David H. Albonesi, Rajit Manohar |
IJCNN | 3 |
| 2017 | Toolbox for exploration of energy-efficient event processors for human-computer interactionabstractThe advent of high speed input sensor and display technologies and the drive for faster interactive response suggests that human-computer interaction (HCI) task processing deadlines of a few milliseconds or less may be required in future handheld devices. At the same time, users will expect the same, if not better, battery life than today's devices under these more stringent response requirements. In this paper, we present a toolbox for exploring the design space of HCI event processors. We first describe the simulation platform for interactive environments that runs mobile user interface code with inputs recorded from human users. We validate it against a hardware platform from prior work. Given system-level constraints on latency, we demonstrate how this toolbox can be used to design a custom heterogeneous event processor that maximizes battery life. We show that our toolbox can pick design points that are 1.5-2.5× more energy-efficient than general-purpose big. LITTLE architectures. Tayyar Rzayev, David H. Albonesi, François Guimbretière, Rajit Manohar, Jaeyeon Kihm |
ISPASS | 2 |
| 2013 | Flicker: a dynamically adaptive architecture for power limited multicore systemsabstractFuture microprocessors may become so power constrained that not all transistors will be able to be powered on at once. These systems will be required to nimbly adapt to changes in the chip power that is allocated to general-purpose cores and to specialized accelerators. Paula Petrica, Adam M. Izraelevitz, David H. Albonesi, Christine A. Shoemaker |
ISCA | 3 |
| 2011 | A low-latency, high-throughput on-chip optical router architecture for future chip multiprocessorsabstractTens and eventually hundreds of processing cores are projected to be integrated onto future microprocessors, making the global interconnect a key component to achieving scalable chip performance within a given power envelope. While CMOS-compatible nanophotonics has emerged as a leading candidate for replacing global wires beyond the 16nm timeframe, on-chip optical interconnect architectures are typically limited in scalability or are dependent on comparatively slow electrical control networks. In this article, we present a hybrid electrical/optical router for future large scale, cache coherent multicore microprocessors. The heart of the router is a low-latency optical crossbar that uses predecoded source routing and switch state preconfiguration to transmit cache-line-sized packets several hops in a single clock cycle under contentionless conditions. Overall, our optical router achieves 2X better network performance than a state-of-the-art electrical baseline in a mesh topology while consuming 30% less network power. Mark J. Cianchetti, David H. Albonesi |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2010 | Dynamically managed multithreaded reconfigurable architectures for chip multiprocessorsabstractPrior work has demonstrated that reconfigurable logic can significantly benefit certain applications. However, reconfigurable architectures have traditionally suffered from high area overhead and limited application coverage. We present a dynamically managed multithreaded reconfigurable architecture consisting of multiple clusters of shared reconfigurable fabrics that greatly reduces the area overhead of reconfigurability while still offering the same power efficiency and performance benefits. Like other shared SMT and CMP resources, the dynamic partitioning of the reconfigurable resource among sharing threads, along with the co-scheduling of threads among different reconfigurable clusters, must be intelligently managed for the full benefits of the shared fabrics to be realized. Matthew A. Watkins, David H. Albonesi |
PACT | 2 |
| 2010 | Scalable thread scheduling and global power management for heterogeneous many-core architecturesabstractFuture many-core microprocessors are likely to be heterogeneous, by design or due to variability and defects. The latter type of heterogeneity is especially challenging due to its unpredictability. To minimize the performance and power impact of these hardware imperfections, the runtime thread scheduler and global power manager must be nimble enough to handle such random heterogeneity. With hundreds of cores expected on a single die in the future, these algorithms must provide high power-performance efficiency, yet remain scalable with low runtime overhead. Jonathan A. Winter, David H. Albonesi, Christine A. Shoemaker |
PACT | 2 |
| 2010 | Adaptive Cache Memories for SMT ProcessorsabstractResizable caches can trade-off capacity for access speed to dynamically match the needs of the workload. In Simultaneous Multi-Threaded (SMT) cores, the caching needs can vary greatly across the number of threads and their characteristics, offering opportunities to dynamically adjust cache resources to the workload. In this paper we propose the use of resizable caches in order to improve the performance of SMT cores, and introduce a new control algorithm that provides good results independent of the number of running threads. In workloads with a single thread, the resizable cache control algorithm should optimize for cache miss behavior because misses typically form the critical path. In contrast, with several independent threads running, we show that optimizing for cache hit behavior has more impact, since large SMT workloads have other threads to run during a cache miss. Moreover, we demonstrate that these seemingly diametrically opposed policies can be simultaneously satisfied by using the harmonic mean of the per-thread speedups as the metric to evaluate the system performance, and to smoothly and naturally adjust to the degree of multithreading. Sonia López, Oscar Garnica, David H. Albonesi, Steven G. Dropsho, Juan Lanchares, J. Ignacio Hidalgo |
DSD | 3 |
| 2010 | ReMAP: A Reconfigurable Heterogeneous Multicore ArchitectureabstractThis paper presents ReMAP, a reconfigurable architecture geared towards accelerating and parallelizing applications within a heterogeneous CMP. In ReMAP, threads share a common reconfigurable fabric that can be configured for individual thread computation or fine-grained communication with integrated computation. The architecture supports both fine-grained point-to-point communication for pipeline parallelization and fine-grained barrier synchronization. The combination of communication and configurable computation within ReMAP provides the unique ability to perform customized computation while data is transferred between cores, and to execute custom global functions after barrier synchronization. ReMAP demonstrates significantly higher performance and energy efficiency compared to hard-wired communication-only mechanisms, and over what can ideally be achieved by allocating the fabric area to additional or more powerful cores. Matthew A. Watkins, David H. Albonesi |
MICRO | 2 |
| 2009 | Phastlane: a rapid transit optical routing networkabstractTens and eventually hundreds of processing cores are projected to be integrated onto future microprocessors, making the global interconnect a key component to achieving scalable chip performance within a given power envelope. While CMOS-compatible nanophotonics has emerged as a leading candidate for replacing global wires beyond the 22nm timeframe, on-chip optical interconnect architectures proposed thus far are either limited in scalability or are dependent on comparatively slow electrical control networks. Mark J. Cianchetti, Joseph C. Kerekes, David H. Albonesi |
ISCA | 3 |
| 2008 | Scheduling algorithms for unpredictably heterogeneous CMP architecturesabstractIn future large-scale multi-core microprocessors, hard errors and process variations will create dynamic heterogeneity, causing performance and power characteristics to differ among the cores in an unanticipated manner. Under this scenario, naive assignments of applications to cores degraded by various faults and variations may result in large performance losses and power inefficiencies. We propose scheduling algorithms based on the Hungarian Algorithm and artificial intelligence (AI) search techniques that account for this future uncertainty in core characteristics. These thread assignment policies effectively match the capabilities of each degraded core with the requirements of the applications, achieving an ED2only 3.2% and 3.7% higher, respectively, than a baseline eight core chip multiprocessor with no degradation, compared to over 22% for a round robin policy. Jonathan A. Winter, David H. Albonesi |
DSN | 2 |
| 2008 | Shared reconfigurable architectures for CMPSabstractThis paper investigates reconfigurable architectures suitable for chip multiprocessors (CMPs). Prior research has established that augmenting a conventional processor with reconfigurable logic can dramatically improve the performance of certain application classes, but this comes at non-trivial power and area costs. Given substantial observed time and space differences in fabric usage, we propose that pools of programmable logic should be shared among multiple cores. While a common shared pool is more compact and power efficient, fabric conflicts may lead to large performance losses relative to per-core private fabrics. We identify particular characteristics of past reconfigurable fabric designs that are particularly amenable to fabric sharing. We then propose spatially and temporally shared fabrics in a CMP. The sharing policies that we devise incur negligible performance loss compared to private fabrics, while cutting the area and peak power of the fabric by 4X. Matthew A. Watkins, Mark J. Cianchetti, David H. Albonesi |
FPL | 3 |
| 2008 | Addressing thermal nonuniformity in SMT workloadsabstractWe explore DTM techniques within the context of uniform and nonuniform SMT workloads. While DVS is suitable for addressing workloads with uniformly high temperatures, for nonuniform workloads, performance loss occurs because of the slowdown of the cooler thread. To address this, we propose and evaluate DTM mechanisms that exploit the steering-based thread management mechanisms inherent in a clustered SMT architecture. We show that in contrast to DVS, which operates globally, our techniques are more effective at controlling temperature for nonuniform workloads. Furthermore, we devise a DTM technique that combines steering and DVS to achieve consistently good performance across all workloads. Jonathan A. Winter, David H. Albonesi |
ACM Trans. Archit. Code Optim. | 2 |
| 2007 | Rate-Driven Control of Resizable Caches for Highly Threaded SMT Processors
Sonia López, Steven G. Dropsho, David H. Albonesi, Oscar Garnica, Juan Lanchares |
PACT | 3 |
| 2007 | Dynamic Capacity-Speed Tradeoffs in SMT Processor Caches
Sonia López, Steven G. Dropsho, David H. Albonesi, Oscar Garnica, Juan Lanchares |
HiPEAC | 3 |
| 2007 | Predictions of CMOS compatible on-chip optical interconnect
Mikhail Haurylau, Nicholas Nelson 0001, David H. Albonesi, Philippe M. Fauchet, Eby G. Friedman |
Integr. | 5 |
| 2006 | Compatible phase co-scheduling on a CMP of multi-threaded processorsabstractThe industry is rapidly moving towards the adoption of chip multi-processors (CMPs) of simultaneous multi-threaded (SMT) cores for general purpose systems. The most prominent use of such processors, at least in the near term, is as job servers running multiple independent threads on the different contexts of the various SMT cores. In such an environment, the co-scheduling of phases from different threads plays a significant role in the overall throughput. Less throughput is achieved when phases from different threads that conflict for particular hardware resources are scheduled together, compared with the situation where compatible phases are co-scheduled on the same SMT core. Achieving the latter requires precise per-phase hardware statistics that the scheduler can use to rapidly identify possible incompatibilities among phases of different threads, thereby avoiding the potentially high performance cost of inter-thread contention. In this paper, we devise phase co-scheduling policies for a dual-core CMP of dual-threaded SMT processors. We explore a number of approaches and find that the use of ready and in-flight instruction metrics permits effective co-scheduling of compatible phases among the four contexts. This approach significantly outperforms the worst static grouping of threads, and very closely matches the best static grouping, even outperforming it by as much as 7%. Ali El-Moursy, Rajeev Garg, David H. Albonesi, Sandhya Dwarkadas |
IPDPS | 3 |
| 2006 | Localized microarchitecture-level voltage managementabstractDiminishing voltage margins, coupled with power and temperature constraints, call for microarchitecture-level runtime mechanisms for voltage control. This paper describes a localized approach for dynamic voltage management within the domains of a globally asynchronous, locally synchronous (GALS) processor design. Dynamic voltage scaling at this fine grain level permits effective temperature management with less performance impact than global voltage control. Yongkang Zhu, David H. Albonesi |
ISCAS | 2 |
| 2006 | Synergistic temperature and energy management in GALS processor architecturesabstractWe propose a synergistic temperature and energy management scheme for GALS processors. Localized DVS is applied in domains that contain hotspots, permitting other critical domains to run unabated, thereby reducing performance cost relative to global DVS, and also creating execution slack in peripheral cooler domains that can be exploited to save energy. The reduction in energy in turn creates a steeper temperature gradient between the domains, permitting heat to flow more easily out of the hotspot domain. This symbiotic cyclical relationship between temperature and energy management leads to both significantly better performance, and lower energy, than the use of DTM alone. Yongkang Zhu, David H. Albonesi |
ISLPED | 2 |
| 2006 | Leveraging Optical Technology in Future Bus-based Chip MultiprocessorsabstractAlthough silicon optical technology is still in its formative stages, and the more near-term application is chip-to-chip communication, rapid advances have been made in the development of on-chip optical interconnects. In this paper, we investigate the integration of CMOS-compatible optical technology to on-chip cache-coherent buses in future CMPs. While not exhaustive, our investigation yields a hierarchical opto-electrical system that exploits the advantages of optical technology while abiding by projected limitations. Our evaluation shows that, for the applications considered, compared to an aggressive all-electrical bus of similar power and area, significant performance improvements can be achieved using an opto-electrical bus. This performance improvement is largely dependent on the application's bandwidth demand and on the number of implemented wavelengths per optical waveguide. We also present a number of critical areas for future work that we discover in the course of our research Nevin Kirman, Meyrem Kirman, Rajeev K. Dokania, José F. Martínez, Alyssa B. Apsel, Matthew A. Watkins, David H. Albonesi |
MICRO | 7 |
| 2005 | Partitioning Multi-Threaded Processors with a Large Number of ThreadsabstractToday's general-purpose processors are increasingly using multithreading in order to better leverage the additional on-chip real estate available with each technology generation. Simultaneous multi-threading (SMT) was originally proposed as a large dynamic superscalar processor with monolithic hardware structures shared among all threads. Inters hyper-threaded Pentium 4 processor partitions the queue structures among two threads, demonstrating more balanced performance by reducing the hoarding of structures by a single thread. IBM's Power5 processor is a 2-way chip multiprocessor (CMP) of SMT processors, each supporting 2 threads, which significantly reduces design complexity and can improve power efficiency. This paper examines processor partitioning options for larger numbers of threads on a chip. While growing transistor budgets permit four and eight-thread processors to be designed, design complexity, power dissipation, and wire scaling limitations create significant barriers to their actual realization. We explore the design choices of sharing, or of partitioning and distributing, the front end (instruction cache, instruction fetch, and dispatch), the execution units and associated state, as well as the L1 Dcache banks, in a clustered multi-threaded (CMT) processor. We show that the best performance is obtained by restricting the sharing of the L1 Dcache banks and the execution engines among threads. On the other hand, significant sharing of the front-end resources is the best approach. When compared against large monolithic SMT processors, a CMT processor provides very competitive IPC performance on average, 90-96% of that of partitioned SMT while being more scalable and much more power efficient. In a CMP organization, the gap between SMT and CMT processors shrinks further, making a CMP of CMT processors a highly viable alternative for the future Ali El-Moursy, Rajeev Garg, David H. Albonesi, Sandhya Dwarkadas |
ISPASS | 3 |
| 2005 | A High Performance, Energy Efficient GALS ProcessorMicroarchitecture with Reduced Implementation ComplexityabstractAs the costs and challenges of global clock distribution grow with each new microprocessor generation, a globally asynchronous, locally synchronous (GALS) approach becomes an attractive alternative. One proposed GALS approach, called a multiple clock domain (MCD) processor, achieves impressive energy savings for a relatively low performance cost. However, the approach requires separating the processor into four domains, including separating the integer and memory domains which complicates load scheduling, and the implementation of 32 voltage and frequency levels in each domain. In addition, the hardware-based control algorithm, though effective overall, produces a significant performance degradation for some applications. In this paper, we devise modifications to the MCD design that retain many of its benefits while greatly reducing the implementation complexity. We first determine that the synchronization channels that are most responsible for the MCD performance degradation are those involving cache access, and propose merging the integer and memory domains to virtually eliminate this overhead. We further propose significantly reducing the number of voltage levels, separating the reorder buffer into its own domain to permit front-end frequency scaling, separating the L2 cache to permit standard power optimizations to be used, and a new online algorithm that provides consistent results across our benchmark suite. The overall result is a significant reduction in the performance degradation of the original MCD approach and greater energy savings, with a greatly simplified microarchitecture that is much easier to implement Yongkang Zhu, David H. Albonesi, Alper Buyuktosunoglu |
ISPASS | 2 |
| 2004 | Dynamically Trading Frequency for Complexity in a GALS MicroprocessorabstractMicroprocessors are traditionally designed to provide "best overall" performance across a wide range of applications and operating environments. Several groups have proposed hardware techniques that save energy by "downsizing" hardware resources that are underutilized by the current application phase. Others have proposed a different energy-saving approach: dividing the processor into domains and dynamically changing the clock frequency and voltage within each domain during phases when the full domain frequency is not required. What has not been studied to date is how to exploit the adaptive nature of these approaches to improve performance rather than to save energy. In this paper, we describe an adaptive globally asynchronous, locally synchronous (GALS) microprocessor with a fixed global voltage and four independently clocked domains. Each domain is streamlined with modest hardware structures for very high clock frequency. Key structures can then be upsized on demand to exploit more distant parallelism, improve branch prediction, or increase cache capacity. Although doing so requires decreasing the associated domain frequency, other domain frequencies are unaffected. Our approach, therefore, is to maximize the throughput of each domain by finding the proper balance between the number of clock periods, and the clock frequency, for each application phase. To achieve this objective, we use novel hardware-based control techniques that accurately and efficiently capture the performance of all possible cache and queue configurations within a single interval, without having to resort to exhaustive online exploration or expensive offline profiling. Measuring across a broad suite of application benchmarks, we find that configuring our adaptive GALS processor just once per application yields 17.6% better performance, on average, than that of the "best overall" fully synchronous design. By adapting automatically to application phases, we can increase this advantage to more than 20%. Steven G. Dropsho, Greg Semeraro, David H. Albonesi, Grigorios Magklis, Michael L. Scott |
MICRO | 3 |
| 2003 | Dynamic Data Dependence Tracking and its Application to Branch PredictionabstractTo continue to improve processor performance, microarchitects seek to increase the effective instruction level parallelism (ILP) that can be exploited in applications. A fundamental limit to improving ILP is data dependences among instructions. If data dependence information is available at run-time, there are many uses to improve ILP. Prior published examples include decoupled branch execution architectures and critical instruction detection. In this paper, we describe an efficient hardware mechanism to dynamically track the data dependence chains of the instructions in the pipeline. This information is available on a cycle-by-cycle basis to the microengine for optimizing its performance. We then use this design in a new value-based branch prediction design using available register value information (ARVI). From the use of data dependence information, the ARVI branch predictor has better prediction accuracy over a comparably sized hybrid branch predictor With ARVI used as the second-level branch predictor the improved prediction accuracy results in a 12.6% performance improvement on average across the SPEC95 integer benchmark suite. Lei Chen 0021, Steven G. Dropsho, David H. Albonesi |
HPCA | 3 |
| 2003 | Front-End Policies for Improved Issue Efficiency in SMT ProcessorsabstractThe performance and power optimization of dynamic superscalar microprocessors requires striking a careful balance between exploiting parallelism and hardware simplification. Hardware structures which are needlessly complex may exacerbate critical timing paths and dissipate extra power. One such structure requiring careful design is the issue queue. In a simultaneous multi-threading (SMT) processor it is particularly challenging to achieve issue queue simplification due to the increased utilization of the queue afforded by multi-threading. In this paper we propose new front-end policies that reduce the required integer and floating point issue queue sizes in SMT processors. We explore both general policies as well as those directed towards alleviating a particular cause of issue queue inefficiency. For the same level of performance, the most effective policies reduce the issue queue occupancy by 33% for an SMT processor with appropriately sized issue queue resources. Ali El-Moursy, David H. Albonesi |
HPCA | 2 |
| 2003 | Dynamically Managing the Communication-Parallelism Trade-off in Future Clustered ProcessorsabstractClustered microarchitectures are an attractive alternative to large monolithic superscalar designs due to their potential for higher clock rates in the face of increasingly wire-delay-constrained process technologies. As increasing transistor counts allow an increase in the number of clusters, thereby allowing more aggressive use of instruction-level parallelism (ILP), the inter-cluster communication increases as data values get spread across a wider area. As a result of the emergence of this trade-off between communication and parallelism, a subset of the total on-chip clusters is optimal for performance. To match the hardware to the application's needs, we use a robust algorithm to dynamically tune the clustered architecture. The algorithm, which is based on program metrics gathered at periodic intervals, achieves an 11% performance improvement on average over the best statically defined architecture. We also show that the use of additional hardware and reconfiguration at basic block boundaries can achieve average improvements of 15%. Our results demonstrate that reconfiguration provides an effective solution to the communication and parallelism trade-off inherent in the communication-bound processors of the future. Rajeev Balasubramonian, Sandhya Dwarkadas, David H. Albonesi |
ISCA | 3 |
| 2003 | Energy Efficient Co-Adaptive Instruction Fetch and IssueabstractFront-end instruction delivery accounts for a significant fraction of the energy consumed in a dynamic superscalar processor. The issue queue in these processors serves two crucial roles: it bridges the front and back ends of the processor and serves as the window of instructions for the out-of-order engine. A mismatch between the front end producer rate and back end consumer rate, and between the supplied instruction window from the front end, and the required instruction window to exploit the level of application parallelism, results in additional front-end energy, and increases the issue queue utilization. While the former increases overall processor energy consumption, the latter aggravates the issue queue hot spot problem.We propose a complementary combination of fetch gating and issue queue adaptation to address both of these issues. We introduce an issue-centric fetch gating scheme based on issue queue utilization and application parallelism characteristics. Our scheme attempts to provide an instruction window size that matches the current parallelism characteristics of the application while maintaining enough queue entries to avoid back-end starvation. Compared to a conventional fetch gating scheme based on flow-rate matching, we demonstrate 20% better overall energy-delay with a 44% additional reduction in issue queue energy. We identify Icache energy savings as the largest contributor to the overall savings and quantify the sources of savings in this structure. We then couple this issue-driven fetch gating approach with an issue queue adaptation scheme based on queue utilization. While the fetch gating scheme provides a window of issue queue instructions appropriate to the level of program parallelism, the issue queue adaptation approach shuts down the remaining underutilized issue queue entries. Used in tandem, these complementary techniques yield a 20% greater issue queue energy savings than the addition of the savings from each technique applied in isolation. The result of this combined approach is a 6% overall energy-delay savings coupled with a 54% reduction in issue queue energy. Alper Buyuktosunoglu, Tejas Karkhanis, David H. Albonesi, Pradip Bose |
ISCA | 3 |
| 2003 | Profile-Based Dynamic Voltage and Frequency Scaling for a Multiple Clock Domain MicroprocessorabstractA Multiple Clock Domain (MCD) processor addresses the challenges of clock distribution and power dissipation by dividing a chip into several (coarse-grained) clock domains, allowing frequency and voltage to be reduced in domains that are not currently on the application’s critical path. Given a reconfiguration mechanism capable of choosing appropriate times and values for voltage/frequency scaling, an MCD processor has the potential to achieve significant energy savings with low performance degradation. Early work on MCD processors evaluated the potential for energy savings by manually inserting reconfiguration instructions into applications, or by employing an oracle driven by off-line analysis of (identical) prior program runs. Subsequent work developed a hardware-based on-line mechanism that averages 75–85% of the energy-delay improvement achieved via off-line analysis. In this paper we consider the automatic insertion of reconfiguration instructions into applications, using profiledriven binary rewriting. Profile-based reconfiguration introduces the need for “training runs” prior to production use of a given application, but avoids the hardware complexity of on-line reconfiguration. It also has the potential to yield significantly greater energy savings. Experimental results (training on small data sets and then running on larger, alternative data sets) indicate that the profile-driven approach is more stable than hardware-based reconfiguration, and yields virtually all of the energy-delay improvement achieved via off-line analysis. Grigorios Magklis, Michael L. Scott, Greg Semeraro, David H. Albonesi, Steven G. Dropsho |
ISCA | 4 |
| 2003 | A Dynamically Tunable Memory HierarchyabstractThe widespread use of repeaters in long wires creates the possibility of dynamically sizing regular on-chip structures. We present a tunable cache and translation lookaside buffer (TLB) hierarchy that leverages repeater insertion to dynamically trade off size for speed and power consumption on a per-application phase basis using a novel configuration management algorithm. In comparison to a conventional design that is fixed at a single design point targeted to the average application, the dynamically tunable cache and TLB hierarchy can be tailored to the needs of each application phase. The configuration algorithm dynamically detects phase changes and selects a configuration based on the application's ability to tolerate different hit and miss latencies in order to improve the memory energy-delay product. We evaluate the performance and energy consumption of our approach and project the effects of technology scaling trends on our design. Rajeev Balasubramonian, David H. Albonesi, Alper Buyuktosunoglu, Sandhya Dwarkadas |
IEEE Trans. Computers | 2 |
| 2002 | Energy-Efficient Processor Design Using Multiple Clock Domains with Dynamic Voltage and Frequency ScalingabstractAs clock frequency increases and feature size decreases, clock distribution and wire delays present a growing challenge to the designers of singly-clocked, globally synchronous systems. We describe an alternative approach, which we call a multiple clock domain (MCD) processor, in which the chip is divided into several clock domains, within which independent voltage and frequency scaling can be performed. Boundaries between domains are chosen to exploit existing queues, thereby minimizing inter-domain synchronization costs. We propose four clock domains, corresponding to the front end , integer units, floating point units, and load-store units. We evaluate this design using a simulation infrastructure based on SimpleScalar and Wattch. In an attempt to quantify potential energy savings independent of any particular on-line control strategy, we use off-line analysis of traces from a single-speed run of each of our benchmark applications to identify profitable reconfiguration points for a subsequent dynamic scaling run. Using applications from the MediaBench, Olden, and SPEC2000 benchmark suites, we obtain an average energy-delay product improvement of 20% with MCD compared to a modest 3% savings from voltage scaling a single clock and voltage system. Greg Semeraro, Grigorios Magklis, Rajeev Balasubramonian, David H. Albonesi, Sandhya Dwarkadas, Michael L. Scott |
HPCA | 4 |
| 2002 | Tradeoffs in power-efficient issue queue designabstractA major consumer of microprocessor power is the issue queue. Several microprocessors, including the Alpha 21264 and POWER4TM, use a compacting latch-based issue queue design which has the advantage of simplicity of design and verification. The disadvantage of this structure, however, is its high power dissipation.In this paper, we explore different issue queue power optimization techniques that vary not only in their performance and power characteristics, but in how much they deviate from the baseline implementation. By developing and comparing techniques that build incrementally on the baseline design, as well as those that achieve higher power savings through a more significant redesign effort, we quantify the extra benefit the higher design cost techniques provide over their more straightforward counterparts. Alper Buyuktosunoglu, David H. Albonesi, Pradip Bose, Peter W. Cook, Stanley Schuster |
ISLPED | 2 |
| 2002 | A microarchitectural-level step-power analysis toolabstractClock gating is an effective means for reducing average power consumption. However, clock gating can exacerbate maximum cycle-to-cycle current swings, or the step-power (Ldi/dt) problem. We present a microarchitecture-level step-power simulator and demonstrate its use in exploring how design alternatives impact relative step-power levels. We show how the tool can be used to identify major sources of high microprocessor step-power events. Our experiments indicate that branch mispredictions are a major cause of high step-power occurrences. We also show that high step-power events are infrequent which suggest that architectural techniques may limit step-power at potentially low performance cost. Wael El-Essawy, David H. Albonesi, Balaram Sinharoy |
ISLPED | 2 |
| 2002 | Managing static leakage energy in microprocessor functional unitsabstractStatic energy due to subthreshold leakage current is projected to become a major component of the total energy in high performance microprocessors. Many studies so far have examined and proposed techniques to reduce leakage in on-chip storage structures. In this study, static energy is reduced in the integer functional units by leveraging the unique qualities of dual threshold voltage domino logic. Domino logic has desirable properties that greatly reduce leakage current while providing fast propagation times. However due to the energy cost of entering the low leakage current state (sleep mode), domino logic has thus far been used only for leakage reduction in the longterm standby mode. We examine the utility of the sleep mode (while considering the aforementioned costs) when idle times are relatively short, one to a few hundred cycles, as is often the case for functional units. Using an analytical energy model suitable for architecture-level analysis, we explore the interaction of the application and technology, and the effect on energy and performance as the underlying parameters are varied, on a set of benchmarks. Our results show that if the leakage approaches the magnitude as projected in the literature, even for short idle intervals as few as ten cycles, an aggressive policy of activating the sleep mode at every idle period performs well and a more complex control strategy may not be warranted. We also propose a simple design, called Gradual Sleep, to reduce the energy impact of using the sleep mode for smaller idle periods. Steven G. Dropsho, Volkan Kursun, David H. Albonesi, Sandhya Dwarkadas, Eby G. Friedman |
MICRO | 3 |
| 2002 | Dynamic frequency and voltage control for a multiple clock domain microarchitectureabstractWe describe the design, analysis, and performance of an on-line algorithm to dynamically control the frequency/voltage of a Multiple Clock Domain (MCD) microarchitecture. The MCD microarchitecture allows the frequency/voltage of microprocessor regions to be adjusted independently and dynamically, allowing energy savings when the frequency of some regions can be reduced without significantly impacting performance. Our algorithm achieves on average a 19.0% reduction in Energy Per Instruction (EPI), a 3.2% increase in Cycles Per Instruction (CPI), a 16.7% improvement in Energy-Delay Product, and a Power Savings to Performance Degradation ratio of 4.6. Traditional frequency/voltage scaling techniques which apply reductions globally to a fully synchronous processor achieve a Power Savings to Performance Degradation ratio of only 2-3. Our Energy-Delay Product improvement is 85.5% of what has been achieved using an off-line algorithm. These results were achieved using a broad range of applications from the MediaBench, Olden, and Spec2000 benchmark suites using an algorithm we show to require minimal hardware resources. Greg Semeraro, David H. Albonesi, Steven G. Dropsho, Grigorios Magklis, Sandhya Dwarkadas, Michael L. Scott |
MICRO | 2 |
| 2001 | A circuit level implementation of an adaptive issue queue for power-aware microprocessorsabstractIncreasing power dissipation has become a major constraint for future performance gains in the design of microprocessors. In this paper, we present the circuit design of an issue queue for a superscalar processor that leverages transmission gate insertion to provide dynamic low-cost configurability of size and speed. A novel circuit structure dynamically gathers statistics of issue queue activity over intervals of instruction execution. These statistics are then used to change the size of an issue queue organization on-the-fly to improve issue queue energy and performance. When applied to a fixed, full-size issue queue structure, the result is up to a 70% reduction in energy dissipation. The complexity of the additional circuitry to achieve this result is almost negligible. Furthermore, self-timed techniques embedded in the adaptive scheme can provide a 56% decrease in cycle time of the CAM array read of the issue queue when we change the adaptive issue queue size from 32 entries (largest possible) to 8 entries (smallest possible in our design). Alper Buyuktosunoglu, David H. Albonesi, Stanley Schuster, David Brooks 0001, Pradip Bose, Peter W. Cook |
ACM Great Lakes Symposium on VLSI | 2 |
| 2001 | Dynamically allocating processor resources between nearby and distant ILPabstractModern superscalar processors use wide instruction issue widths and out-of-order execution in order to increase instruction-level parallelism (ILP). Because instructions must be committed in order so as to guarantee precise exceptions, increasing ILP implies increasing the sizes of structures such as the register file, issue queue, and reorder buffer. Simultaneously, cycle time constraints limit the sizes of these structures, resulting in conflicting design requirements. Rajeev Balasubramonian, Sandhya Dwarkadas, David H. Albonesi |
ISCA | 3 |
| 2001 | Reducing the complexity of the register file in dynamic superscalar processorsabstractDynamic superscalar processors execute multiple instructions out-of-order by looking for independent operations within a large window. The number of physical registers within the processor has a direct impact on the size of this window as most in-flight instructions require a new physical register at dispatch. A large multi-ported register file helps improve the instruction-level parallelism (ILP), but may have a detrimental effect on clock speed, especially in future wire-limited technologies. In this paper, we propose a register file organization that reduces register file size and port requirements for a given amount of ILP. We use a two-level register file organization to reduce register file size requirements, and a banked organization to reduce port requirements. We demonstrate empirically that the resulting register file organizations have reduced latency and (in the case of the banked organization) energy requirements for similar instructions per cycle (IPC) performance and improved instructions per second (IPS) performance in comparison to a conventional monolithic register file. The choice of organization is dependent on design goals. Rajeev Balasubramonian, Sandhya Dwarkadas, David H. Albonesi |
MICRO | 3 |
| 2000 | Memory hierarchy reconfiguration for energy and performance in general-purpose processor architecturesabstractConventional microarchitectures choose a single memory hierarchy design point targeted at the average application. In this paper, we propose a cache and TLB layout and design that leverages repeater insertion to provide dynamic low-cost configurability trading off size and speed on a per application phase basis. A novel configuration management algorithm dynamically detects phase changes and reacts to an application's hit and miss intolerance in order to improve memory hierarchy performance while taking energy consumption into consideration. When applied to a two-level cache and TLB hierarchy at 0.1 /spl mu/m technology, the result is an average 15% reduction in cycles per instruction (CPI), corresponding to an average 27% reduction in memory-CPI, across a broad class of applications compared to the best conventional two-level hierarchy of comparable size. Projecting to sub-.1 /spl mu/m technology design considerations that call for a three-level conventional cache hierarchy for performance reasons, we demonstrate that a configurable L2/L3 cache hierarchy coupled with a conventional LI results in an average 43% reduction in memory hierarchy energy in addition to improved performance. Rajeev Balasubramonian, David H. Albonesi, Alper Buyuktosunoglu, Sandhya Dwarkadas |
MICRO | 2 |
| 1999 | Selective Cache Ways: On-Demand Cache Resource AllocationabstractIncreasing levels of microprocessor power dissipation call for new approaches at the architectural level that save energy by better matching of on-chip resources to application requirements. Selective cache ways provides the ability to disable a subset of the ways in a set associative cache during periods of modest cache activity, while the full cache may remain operational for more cache-intensive periods. Because this approach leverages the subarray partitioning that is already present for performance reasons, only minor changes to a conventional cache are required and therefore, full-speed cache operation can be maintained. Furthermore, the tradeoff between performance and energy is flexible, and can be dynamically tailored to meet changing application and machine environmental conditions. We show that trading off a small performance degradation for energy savings can produce a significant reduction in cache energy dissipation using this approach. David H. Albonesi |
MICRO | 1 |
| 1999 | STATS: A framework for microprocessor and system-level design space explorationabstractAs microprocessor-based systems grow in complexity, and the processor-memory speed gap widens further, more emphasis needs to be placed on early design space exploration in order to produce the highest performance systems with minimal schedule impact. We discuss the critical issues associated with architectural evaluation of complex microprocessor-based systems, and present a methodology for the comprehensive and semiautomatic evaluation of processor, cache hierarchy, system interconnect, and main memory architectural and technological alternatives. We discuss the implementation of the methodology, and describe how it can be used in early design space exploration. The unique aspects of the methodology are further illustrated through two architectural investigations performed using the toolset. David H. Albonesi, Israel Koren |
J. Syst. Archit. | 1 |
| 1998 | Dynamic IPC/Clock Rate OptimizationabstractCurrent microprocessor designs set the functionality and clock rate of the chip at design time based on the configuration that achieves the best overall performance over a range of target applications. The result may be poor performance when running applications whose requirements are not well-matched to the particular hardware organization chosen. We present a new approach called Complexity-Adaptive Processors (CAPs) in which the IPC/clock rate tradeoff can be altered at runtime to dynamically match the changing requirements of the instruction stream. By exploiting repeater methodologies used increasingly in deep sub-micron designs, CAPs achieve this flexibility with potentially no cycle time impact compared to a fixed architecture. Our preliminary results in applying this approach to on-chip caches and instruction queues indicate that CAPs have the potential to significantly outperform conventional approaches on workloads containing both general purpose and scientific applications. David H. Albonesi |
ISCA | 1 |
| 1995 | An analytical model of high performance superscalar-based multiprocessors
David H. Albonesi, Israel Koren |
PACT | 1 |