EDBT 2026 Demo / reviewers in the wild / expert
Greg Semeraro
dblp:85/6207
· DBLP profile ↗
4ranked-venue papers
2as first author
0since 2021 · last 2004
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Energy-efficient computing · 56% Integrated circuit design · 20% Electronic design automation · 11% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.2 | 4 | 2004 | Dynamically Trading Frequency for Complexity in a GALS Microprocessor · MICRO 2004 Profile-Based Dynamic Voltage and Frequency Scaling for a Multiple Clock Domain Microprocessor · ISCA 2003 Dynamic frequency and voltage control for a multiple clock domain microarchitecture · MICRO 2002 |
Energy-efficient computing › microprocessor power management
multiple clock domain processor |
0.1 | 2 | 2003 | Profile-Based Dynamic Voltage and Frequency Scaling for a Multiple Clock Domain Microprocessor · ISCA 2003 Energy-Efficient Processor Design Using Multiple Clock Domains with Dynamic Voltage and Frequency Scaling · HPCA 2002 |
Electronic design automation › design for manufacturability
area fill synthesis |
0.0 | 1 | 2004 | Dynamically Trading Frequency for Complexity in a GALS Microprocessor · MICRO 2004 |
Integrated circuit design › asynchronous circuit design
globally asynchronous locally synchronous design |
0.0 | 1 | 2004 | Dynamically Trading Frequency for Complexity in a GALS Microprocessor · MICRO 2004 |
High-performance computing
performance optimization |
0.0 | 1 | 2004 | Dynamically Trading Frequency for Complexity in a GALS Microprocessor · MICRO 2004 |
Integrated circuit design › clocking
multiple clock domain |
0.0 | 1 | 2002 | Dynamic frequency and voltage control for a multiple clock domain microarchitecture · MICRO 2002 |
Compilers and program optimization
binary rewriting |
0.0 | 1 | 2003 | Profile-Based Dynamic Voltage and Frequency Scaling for a Multiple Clock Domain Microprocessor · ISCA 2003 |
Methods — techniques the papers use, named apart from their topics
profiling · 0.1binary rewriting · 0.1phase detection · 0.0hardware-based control · 0.0simulation · 0.0online control algorithm · 0.0off-line trace analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2004 | Dynamically Trading Frequency for Complexity in a GALS MicroprocessorabstractMicroprocessors are traditionally designed to provide "best overall" performance across a wide range of applications and operating environments. Several groups have proposed hardware techniques that save energy by "downsizing" hardware resources that are underutilized by the current application phase. Others have proposed a different energy-saving approach: dividing the processor into domains and dynamically changing the clock frequency and voltage within each domain during phases when the full domain frequency is not required. What has not been studied to date is how to exploit the adaptive nature of these approaches to improve performance rather than to save energy. In this paper, we describe an adaptive globally asynchronous, locally synchronous (GALS) microprocessor with a fixed global voltage and four independently clocked domains. Each domain is streamlined with modest hardware structures for very high clock frequency. Key structures can then be upsized on demand to exploit more distant parallelism, improve branch prediction, or increase cache capacity. Although doing so requires decreasing the associated domain frequency, other domain frequencies are unaffected. Our approach, therefore, is to maximize the throughput of each domain by finding the proper balance between the number of clock periods, and the clock frequency, for each application phase. To achieve this objective, we use novel hardware-based control techniques that accurately and efficiently capture the performance of all possible cache and queue configurations within a single interval, without having to resort to exhaustive online exploration or expensive offline profiling. Measuring across a broad suite of application benchmarks, we find that configuring our adaptive GALS processor just once per application yields 17.6% better performance, on average, than that of the "best overall" fully synchronous design. By adapting automatically to application phases, we can increase this advantage to more than 20%. Steven G. Dropsho, Greg Semeraro, David H. Albonesi, Grigorios Magklis, Michael L. Scott |
MICRO | 2 |
| 2003 | Profile-Based Dynamic Voltage and Frequency Scaling for a Multiple Clock Domain MicroprocessorabstractA Multiple Clock Domain (MCD) processor addresses the challenges of clock distribution and power dissipation by dividing a chip into several (coarse-grained) clock domains, allowing frequency and voltage to be reduced in domains that are not currently on the application’s critical path. Given a reconfiguration mechanism capable of choosing appropriate times and values for voltage/frequency scaling, an MCD processor has the potential to achieve significant energy savings with low performance degradation. Early work on MCD processors evaluated the potential for energy savings by manually inserting reconfiguration instructions into applications, or by employing an oracle driven by off-line analysis of (identical) prior program runs. Subsequent work developed a hardware-based on-line mechanism that averages 75–85% of the energy-delay improvement achieved via off-line analysis. In this paper we consider the automatic insertion of reconfiguration instructions into applications, using profiledriven binary rewriting. Profile-based reconfiguration introduces the need for “training runs” prior to production use of a given application, but avoids the hardware complexity of on-line reconfiguration. It also has the potential to yield significantly greater energy savings. Experimental results (training on small data sets and then running on larger, alternative data sets) indicate that the profile-driven approach is more stable than hardware-based reconfiguration, and yields virtually all of the energy-delay improvement achieved via off-line analysis. Grigorios Magklis, Michael L. Scott, Greg Semeraro, David H. Albonesi, Steven G. Dropsho |
ISCA | 3 |
| 2002 | Energy-Efficient Processor Design Using Multiple Clock Domains with Dynamic Voltage and Frequency ScalingabstractAs clock frequency increases and feature size decreases, clock distribution and wire delays present a growing challenge to the designers of singly-clocked, globally synchronous systems. We describe an alternative approach, which we call a multiple clock domain (MCD) processor, in which the chip is divided into several clock domains, within which independent voltage and frequency scaling can be performed. Boundaries between domains are chosen to exploit existing queues, thereby minimizing inter-domain synchronization costs. We propose four clock domains, corresponding to the front end , integer units, floating point units, and load-store units. We evaluate this design using a simulation infrastructure based on SimpleScalar and Wattch. In an attempt to quantify potential energy savings independent of any particular on-line control strategy, we use off-line analysis of traces from a single-speed run of each of our benchmark applications to identify profitable reconfiguration points for a subsequent dynamic scaling run. Using applications from the MediaBench, Olden, and SPEC2000 benchmark suites, we obtain an average energy-delay product improvement of 20% with MCD compared to a modest 3% savings from voltage scaling a single clock and voltage system. Greg Semeraro, Grigorios Magklis, Rajeev Balasubramonian, David H. Albonesi, Sandhya Dwarkadas, Michael L. Scott |
HPCA | 1 |
| 2002 | Dynamic frequency and voltage control for a multiple clock domain microarchitectureabstractWe describe the design, analysis, and performance of an on-line algorithm to dynamically control the frequency/voltage of a Multiple Clock Domain (MCD) microarchitecture. The MCD microarchitecture allows the frequency/voltage of microprocessor regions to be adjusted independently and dynamically, allowing energy savings when the frequency of some regions can be reduced without significantly impacting performance. Our algorithm achieves on average a 19.0% reduction in Energy Per Instruction (EPI), a 3.2% increase in Cycles Per Instruction (CPI), a 16.7% improvement in Energy-Delay Product, and a Power Savings to Performance Degradation ratio of 4.6. Traditional frequency/voltage scaling techniques which apply reductions globally to a fully synchronous processor achieve a Power Savings to Performance Degradation ratio of only 2-3. Our Energy-Delay Product improvement is 85.5% of what has been achieved using an off-line algorithm. These results were achieved using a broad range of applications from the MediaBench, Olden, and Spec2000 benchmark suites using an algorithm we show to require minimal hardware resources. Greg Semeraro, David H. Albonesi, Steven G. Dropsho, Grigorios Magklis, Sandhya Dwarkadas, Michael L. Scott |
MICRO | 1 |