Dan Ernst

dblp:12/3555 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
0since 2021 · last 2003
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Processor architecture and microarchitecture · 43% Parallel and multicore computing · 20% Hardware reliability and fault tolerance · 14%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › task scheduling
dynamic scheduling
0.122003
Cyclone: A Broadcast-Free Dynamic Instruction Scheduler with Selective Replay · ISCA 2003
Efficient Dynamic Scheduling Through Tag Elimination · ISCA 2002
Processor architecture and microarchitecture
instruction scheduling
0.122003
Cyclone: A Broadcast-Free Dynamic Instruction Scheduler with Selective Replay · ISCA 2003
Efficient Dynamic Scheduling Through Tag Elimination · ISCA 2002
Energy-efficient computing › voltage scaling
dynamic voltage scaling
0.012003
Razor: A Low-Power Pipeline Based on Circuit-Level Timing Speculation · MICRO 2003
Hardware reliability and fault tolerance
error detection and correction
0.012003
Razor: A Low-Power Pipeline Based on Circuit-Level Timing Speculation · MICRO 2003
Integrated circuit design
low-power circuit design
0.012003
Razor: A Low-Power Pipeline Based on Circuit-Level Timing Speculation · MICRO 2003
Processor architecture and microarchitecture › speculation
timing speculation
0.012003
Razor: A Low-Power Pipeline Based on Circuit-Level Timing Speculation · MICRO 2003
Processor architecture and microarchitecture › speculative execution
data dependence speculation
0.012003
Cyclone: A Broadcast-Free Dynamic Instruction Scheduler with Selective Replay · ISCA 2003
Hardware reliability and fault tolerance
error recovery
0.012003
Razor: A Low-Power Pipeline Based on Circuit-Level Timing Speculation · MICRO 2003
Processor architecture and microarchitecture
instruction issue logic
0.012003
Cyclone: A Broadcast-Free Dynamic Instruction Scheduler with Selective Replay · ISCA 2003
Processor architecture and microarchitecture
pipelining
0.012003
Razor: A Low-Power Pipeline Based on Circuit-Level Timing Speculation · MICRO 2003
Energy-efficient computing › low-power design
power optimization
0.012002
Efficient Dynamic Scheduling Through Tag Elimination · ISCA 2002
Processor architecture and microarchitecture › instruction issue logic
wakeup logic
0.012002
Efficient Dynamic Scheduling Through Tag Elimination · ISCA 2002

Methods — techniques the papers use, named apart from their topics

voltage scaling · 0.0timing error detection · 0.0simulation · 0.0list-based single-pass scheduling · 0.0simulation-based analysis · 0.0circuit-level timing analysis · 0.0
YearPublicationVenuePosition
2003 Cyclone: A Broadcast-Free Dynamic Instruction Scheduler with Selective Replay
abstract
To achieve high instruction throughput, instruction schedulers must be capable of producing high-quality schedules that maximize functional unit utilization while at the same time enabling fast instruction issue logic. Many solutions exist to the scheduling problem, ranging from compile-time to run-time approaches. Compile-time solutions feature fast and simple hardware, but at the expense of conservative schedules. Dynamic schedulers produce high-quality schedules that incorporate run-time information and dependence speculation, but implementing these schedulers requires complex circuits that can slow processor clock speeds. We present the Cyclone scheduler, a novel design that captures the benefits of both compile-and run-time scheduling. Our approach utilizes a list-based single-pass instruction scheduling algorithm, implemented by hardware at run-time in the front end of the processor pipeline. Once scheduled, instructions are injected into a timed queue that orchestrates their entry into execution. To accommodate branch and load/store dependence speculation, the Cyclone scheduler supports a simple selective replay mechanism. We implement this technique by overloading instruction register forwarding to also detect instructions dependent on incorrectly scheduled operations. Detailed simulation analyses suggest that with sufficient queue width, the Cyclone scheduler can rival the instruction throughput of similarly wide monolithic dynamic schedulers. Furthermore, the circuit complexity of the Cyclone scheduler is much more favorable than a broadcast-based scheduler, as our approach requires no global control signals.
Dan Ernst, Andrew Hamel, Todd M. Austin
ISCA1
2003 Razor: A Low-Power Pipeline Based on Circuit-Level Timing Speculation
abstract
With increasing clock frequencies and silicon integration, power aware computing has become a critical concern in the design of embedded processors and systems-on-chip. One of the more effective and widely used methods for power-aware computing is dynamic voltage scaling (DVS). In order to obtain the maximum power savings from DVS, it is essential to scale the supply voltage as low as possible while ensuring correct operation of the processor. The critical voltage is chosen such that under a worst-case scenario of process and environmental variations, the processor always operates correctly. However, this approach leads to a very conservative supply voltage since such a worst-case combination of different variabilities is very rare. In this paper, we propose a new approach to DVS, called Razor, based on dynamic detection and correction of circuit timing errors. The key idea of Razor is to tune the supply voltage by monitoring the error rate during circuit operation, thereby eliminating the need for voltage margins and exploiting the data dependence of circuit delay. A Razor flip-flop is introduced that double-samples pipeline stage values, once with a fast clock and again with a time-borrowing delayed clock. A metastability-tolerant comparator then validates latch values sampled with the fast clock. In the event of timing error, a modified pipeline mispeculation recovery mechanism restores correct program state. A prototype Razor pipeline was designed in a 0.18 /spl mu/m technology and was analyzed. Razor energy overhead during normal operation is limited to 3.1%. Analyses of a full-custom multiplier and a SPICE-level Kogge-Stone adder model reveal that substantial energy savings are possible for these devices (up to 64.2%) with little impact on performance due to error recovery (less than 3%).
Dan Ernst, Nam Sung Kim, Shidhartha Das, Sanjay Pant, Rajeev R. Rao, Toan Pham, Conrad H. Ziesler, David T. Blaauw, Todd M. Austin, Krisztián Flautner, Trevor N. Mudge
MICRO1
2002 Efficient Dynamic Scheduling Through Tag Elimination
abstract
An increasingly large portion of scheduler latency is derived from the monolithic content addressable memory arrays accessed during instruction wake-up. The performance of the scheduler can be improved by decreasing the number of tag comparisons necessary to schedule instructions. Using detailed simulation-based analyses, we find that most instructions enter the window with at least one of their input operands already available. By putting these instructions into specialized windows with fewer tag comparators, load capacitance on the scheduler critical path can be reduced, with only very small effects on program throughput. For instructions with multiple unavailable operands, we introduce a last-tag speculation mechanism that eliminates all remaining tag comparators except those for the last arriving input operand. By combining these two tag-reduction schemes, we are able to construct dynamic schedulers with approximately one quarter of the tag comparators found in conventional designs. Conservative circuit-level timing analyses indicate that the optimized designs are 20-45% faster and require 10-20% less power depending on instruction window size.
Dan Ernst, Todd M. Austin
ISCA1
2001 Performance analysis using pipeline visualization
abstract
High-end microprocessors are increasing in complexity to push the limits of speed and performance. As a result, analyzing these complex system can be an arduous task. Architectural simulators, acting as sofrware processors, are able to run programs and give statistics about the performance of the code on the design. While these statistics are valuable for identifying problems, they often do not provide the fidelity necessary to diagnose the cause of sluggish performance. This paper presents a cross-platform tool that can be used to visualize the flow of instructions through an architectural processor pipeline model. The Graphical Pipeline Viewel; GPC: uses a colorized pipeline trace display to deliver an efJicient diagnostic and analysis environment. The resource view of the tool, which can display cycle statistics, aids in distinguishing possible bottlenecks and architectural trade-ogs. As such, the tool is able to suggest code and architectural modifications to increase program performance.’
Christopher T. Weaver, Kenneth C. Barr, Eric D. Marsman, Dan Ernst, Todd M. Austin
ISPASS4