David Ofelt

dblp:39/857 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
0since 2021 · last 2000
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Performance modeling and evaluation · 55% Processor architecture and microarchitecture · 19% Electronic design automation · 14%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
multiprocessor architecture
0.031997
Hardware/software co-design of the Stanford FLASH multiprocessor · Proc. IEEE 1997
The Stanford FLASH Multiprocessor · ISCA 1994
The Performance Impact of Flexibility in the Stanford FLASH Multiprocessor · ASPLOS 1994
Electronic design automation
hardware/software co-design
0.021998
Digital System Simulation: Methodologies and Examples · DAC 1998
Hardware/software co-design of the Stanford FLASH multiprocessor · Proc. IEEE 1997
Performance modeling and evaluation
performance prediction
0.022000
Efficient performance prediction for modern microprocessors · SIGMETRICS 2000
Digital System Simulation: Methodologies and Examples · DAC 1998
Performance modeling and evaluation › simulation
architectural simulation
0.012000
FLASH vs. (Simulated) FLASH: Closing the Simulation Loop · ASPLOS 2000
Performance modeling and evaluation
simulation
0.021998
Digital System Simulation: Methodologies and Examples · DAC 1998
Hardware/software co-design of the Stanford FLASH multiprocessor · Proc. IEEE 1997
Performance modeling and evaluation › simulation
digital system simulation
0.011998
Digital System Simulation: Methodologies and Examples · DAC 1998
Performance modeling and evaluation
performance monitoring
0.011996
Integrating Performance Monitoring and Communication in Parallel Computers · SIGMETRICS 1996
Memory systems › cache coherence
cache-coherent shared memory
0.011994
The Stanford FLASH Multiprocessor · ISCA 1994
Parallel and multicore computing › parallel programming models
message passing
0.011994
The Stanford FLASH Multiprocessor · ISCA 1994
Processor architecture and microarchitecture › out-of-order execution
out-of-order issue
0.012000
Efficient performance prediction for modern microprocessors · SIGMETRICS 2000
Processor architecture and microarchitecture › pipelining
pipeline performance
0.012000
Efficient performance prediction for modern microprocessors · SIGMETRICS 2000
Performance modeling and evaluation › system-level analysis › architecture evaluation
processor performance evaluation
0.012000
FLASH vs. (Simulated) FLASH: Closing the Simulation Loop · ASPLOS 2000
Electronic design automation › hardware verification and test
hardware verification
0.011997
Hardware/software co-design of the Stanford FLASH multiprocessor · Proc. IEEE 1997
Parallel and multicore computing › parallel computing
parallel communication
0.011996
Integrating Performance Monitoring and Communication in Parallel Computers · SIGMETRICS 1996
Memory systems › cache coherence
cache coherence protocol
0.011994
The Performance Impact of Flexibility in the Stanford FLASH Multiprocessor · ASPLOS 1994
Interconnection networks and networks-on-chip
network interface
0.011994
The Stanford FLASH Multiprocessor · ISCA 1994

Methods — techniques the papers use, named apart from their topics

profile-based analysis · 0.0microbenchmarking · 0.0lightweight instrumentation · 0.0hardware gold-standard comparison · 0.0hierarchical simulation · 0.0hardware-software co-design · 0.0verilog · 0.0system-level simulation · 0.0simulation · 0.0
YearPublicationVenuePosition
2000 FLASH vs. (Simulated) FLASH: Closing the Simulation Loop
abstract
Simulation is the primary method for evaluating computer systems during all phases of the design process. One significant problem with simulation is that it rarely models the system exactly, and quantifying the resulting simulator error can be difficult. More importantly, architects often assume without proof that although their simulator may make inaccurate absolute performance predictions, it will still accurately predict architectural trends.This paper studies the source and magnitude of error in a range of architectural simulators by comparing the simulated execution time of several applications and microbenchmarks to their execution time on the actual hardware being modeled. The existence of a hardware gold standard allows us to find, quantify, and fix simulator inaccuracies. We then use the simulators to predict architectural trends and analyze the sensitivity of the results to the simulator configuration. We find that most of our simulators predict trends accurately, as long as they model all of the important performance effects for the application in question. Unfortunately, it is difficult to know what these effects are without having a hardware reference, as they can be quite subtle. This calls into question the value, for architectural studies, of highly detailed simulators whose characteristics are not carefully validated against s real hardware design.
Jeff Gibson, Robert Kunz, David Ofelt, Mark A. Heinrich
ASPLOS3
2000 Efficient performance prediction for modern microprocessors
abstract
Generating an accurate estimate of the performance of a program on a given system is important to a large number of people. Computer architects, compiler writers, and developers all need insight into a machine's performance. There are a number of performance estimation techniques in use, from profile-based approaches to full machine simulation. This paper discusses a profile-based performance estimation technique that uses a lightweight instrumentation phase that runs in order number of dynamic instructions, followed by an analysis phase that runs in roughly order number of static instructions. This technique accurately predicts the performance of the core pipeline of a detailed out-of-order issue processor model while scheduling far fewer instructions than does full simulation. The difference between the predicted execution time and the time obtained from full simulation is only a few percent.
David Ofelt, John L. Hennessy
SIGMETRICS1
1998 Digital System Simulation: Methodologies and Examples
abstract
Two major trends in the digital design industry are the increase insystem complexity and the increasing importance of short designtimes. The rise in design complexity is motivated by consumerdemand for higher performance products as well as increases inintegration density which allow more functionality to be placed ona single chip. A consequence of this rise in complexity is a significantincrease in the amount of simulation required to design digitalsystems. Simulation time typically scales as the square of theincrease in system complexity [4]. Short design times are importantbecause once a design has been conceived there is a limited timewindow in which to bring the system to market while its performanceis competitive.Simulation serves many purposes during the design cycle of a digitalsystem. In the early stages of design, high-level simulation isused for performance prediction and analysis. In the middle of thedesign cycle, simulation is used to develop the software algorithmsand refine the hardware. In the later stages of design, simulation isused make sure performance targets are reached and to verify thecorrectness of the hardware and software. The different simulationobjectives require varying levels of modeling detail. To keep designtime to a minimum, it is critical to structure the simulation environmentto make it possible to trade-off simulation performance formodel detail in a flexible manner that allows concurrent hardwareand software development.In this paper we describe the different simulation methodologies fordeveloping complex digital systems, and give examples of one suchsimulation environment. The rest of this paper is organized as follows.In Section 2 we describe and classify the various simulationmethodologies that are used in digital system design and describehow they are used in the various stages of the design cycle. In Section3 we provide examples of the methodologies. We describe asophisticated simulation environment used to develop a large ASICfor the Stanford FLASH multiprocessor.
Kunle Olukotun, Mark A. Heinrich, David Ofelt
DAC3
1997 Hardware/software co-design of the Stanford FLASH multiprocessor
abstract
Hardware/software co-design is a methodology for solving design problems in systems with processors or embedded controllers where the design requirements mandate a functionality and performance level for the system, independent of the hardware and software boundary. In addition to the challenges of functional correctness and total system performance, design time is often a critical factor. To design MAGIC, the programmable memory and communication controller for the Stanford FLASH multiprocessor, the authors employed a hardware/software co-design methodology. This methodology allowed them to concurrently design the hardware and software thereby reducing design time while simultaneously ensuring that the design would meet ambitious performance goals. Serializing the hardware and software design would have lengthened the design time and significantly increased the amount of redesign when the tradeoffs between the hardware and software implementations became clear late in the design process. The co-design approach led them to build a series of hierarchical simulators that allowed them to begin design verification early and to reduce the level of effort required to ensure a functional design.
Mark A. Heinrich, David Ofelt, Mark Horowitz, John L. Hennessy
Proc. IEEE2
1996 Integrating Performance Monitoring and Communication in Parallel Computers
Margaret Martonosi, David Ofelt, Mark A. Heinrich
SIGMETRICS2
1994 The Performance Impact of Flexibility in the Stanford FLASH Multiprocessor
abstract
A flexible communication mechanism is a desirable feature in multiprocessors because it allows support for multiple communication protocols, expands performance monitoring capabilities, and leads to a simpler design and debug process. In the Stanford FLASH multiprocessor, flexibility is obtained by requiring all transactions in a node to pass through a programmable node controller, called MAGIC. In this paper, we evaluate the performance costs of flexibility by comparing the performance of FLASH to that of an idealized hardwired machine on representative parallel applications and a multiprogramming workload. To measure the performance of FLASH, we use a detailed simulator of the FLASH and MAGIC designs, together with the code sequences that implement the cache-coherence protocol. We find that for a range of optimized parallel applications the performance differences between the idealized machine and FLASH are small. For these programs, either the miss rates are small or the latency of the programmable protocol can be hidden behind the memory access time. For applications that incur a large number of remote misses or exhibit substantial hot-spotting, performance is poor for both machines, though the increased remote access latencies or the occupancy of MAGIC lead to lower performance for the flexible design. In most cases, however, FLASH is only 2%–12% slower than the idealized machine.
Mark A. Heinrich, Jeffrey Kuskin, David Ofelt, John Heinlein, Joel Baxter, Jaswinder Pal Singh, Richard Simoni, Kourosh Gharachorloo, David Nakahira, Mark Horowitz, Anoop Gupta, Mendel Rosenblum, John L. Hennessy
ASPLOS3
1994 The Stanford FLASH Multiprocessor
abstract
The FLASH multiprocessor efficiently integrates support for cache-coherent shared memory and high-performance message passing, while minimizing both hardware and software overhead. Each node in FLASH contains a microprocessor, a portion of the machine's global memory, a port to the interconnection network, The MAGIC chip handles all communication both within the node and among nodes, using hardwired data paths for efficient data movement and a programmable processor optimized for executing protocol operations. The use of the protocol processor makes FLASH very flexible/spl minus/it can support a variety of different communication mechanisms/spl minus/and simplifies the design and implementation. This paper presents the architecture of FLASH and MAGIC, and discusses the base cache-coherence and message-passing protocols. Latency and occupancy numbers, which are derived from our system-level simulator and our Verilog code, are given for several common protocol operations. The paper also describes our software strategy and FLASH's current status.>
Jeffrey Kuskin, David Ofelt, Mark A. Heinrich, John Heinlein, Richard Simoni, Kourosh Gharachorloo, John Chapin, David Nakahira, Joel Baxter, Mark Horowitz, Anoop Gupta, Mendel Rosenblum, John L. Hennessy
ISCA2